06 October 2026

🤖Prompt Engineering: Knowledge Bases (Just the Quotes)

"How can a cognitive system process environmental input and stored knowledge so as to benefit from experience? More specific versions of this question include the following: How can a system organize its experience so that it has some basis for action even in unfamiliar situations? How can a system determine that rules in its knowledge base are inadequate? How can it generate plausible new rules to replace the inadequate ones? How can it refine rules that are useful but non-optimal? How can it use metaphor and analogy to transfer information and procedures from one domain to another?" (John H Holland et al, "Induction: Processes Of Inference, Learning, And Discovery", 1986)

"Inference is the process of matching current facts from the domain space to the existing knowledge and inferring new facts. An inference process is a chain of matchings. The intermediate results obtained during the inference process are matched against the existing knowledge. The length of the chain is different. It depends on the knowledge base and on the inference method applied." (Nikola K Kasabov, "Foundations of Neural Networks, Fuzzy Systems, and Knowledge Engineering", 1996)

"Representation is the process of transforming existing problem knowledge to some of the known knowledge-engineering schemes in order to process it by applying knowledge-engineering methods. The result of the representation process is the problem knowledge base in a computer format." (Nikola K Kasabov, "Foundations of Neural Networks, Fuzzy Systems, and Knowledge Engineering", 1996)

"LLMs are trained on large volumes of data, which inherently provides them with an immense knowledge base and understanding of different languages. Yet, LLMs at their core are complex text completion engines. Since this knowledge and understanding of language is compressed in a very high-dimensional latent space. LLMs end up using these in a very fluid and intelligible way (which often leads to hallucinations). In order to guide LLMs to focus on specific topics or pieces of information to solve certain tasks, (for instance, question-answering from a given piece of text), it is important to provide contextual information explicitly. While most current generations of LLMs have extremely wide context windows, it is recommended to preprocess context into overlapping smaller chunks for better results, reduced latency, and so on. For similar reasons, it is also recommended to preprocess contextual information in clear and task-specific formats. This aspect of context preprocessing is extremely useful in Retrieval-Gugmented Generation (RAG) scenarios." (Joseph Babcock & Raghav Bali, "Generative AI with Python and PyTorch" 2nd. Ed., 2025)

"LLMs excel at understanding context and making associations among words, phrases, and concepts to provide relevant information based on the input query or prompt. While structured knowledge bases rely on humancurated data, LLMs can  automatically extract knowledge from unstructured text. When trained on diverse textual sources, they can process a vast amount of information without explicit human intervention. However, this also introduces a challenge, as the model can learn biased or incorrect information from the training data." (Abi Aryan, "LLMOps: Managing Large Language Models in Production", 2025)

"RAG is a framework that combines the strengths of traditional information retrieval systems with the generative capabilities of LLMs. In this setup, an LLM is augmented with a retrieval component that fetches relevant information from external data sources, such as knowledge bases or databases, to produce more accurate and contextually relevant responses. This method enhances the LLM’s output by grounding it in authoritative, up-to-date information." (Aldo Marzullo et al, "Graph Machine Learning" 2nd Ed., 2025)

"Semantic Kernel is a framework designed to simplify integrating LLMs into applications that require dynamic knowledge, reasoning, and state tracking. It’s particularly useful when you want to build complex, modular AI systems that can interact with external APIs, knowledge bases, or decision-making processes. Semantic Kernel focuses on building more flexible AI systems that can handle a variety of tasks beyond just generating text. It allows for modularity, enabling developers to easily combine different components - such as embeddings, prompt templates, and custom functions - in a cohesive manner." (Abi Aryan, "LLMOps: Managing Large Language Models in Production", 2025)

"Intelligent systems connect users to AI and ML to achieve meaningful objectives. An intelligent system is one in which intelligence evolves and improves over time, particularly when it improves by watching how users interact with the system.[...] The primary objective of the intelligent system is to support users in accomplishing complex tasks - not by replacing them, but by enhancing their decision-making capabilities. [...] An intelligent system must also have the ability to learn from user interactions and explicit feedback, as well as utilize contextual information. The system should contin-uously develop, use, and maintain an evolving knowledge base. This evolution is driven not only by data sources but also by ongoing interactions with users." (Alessandro Negro et al, "Knowledge Graphs and LLMs in Action", 2026)

"Misinformation is false or misleading information, and its generation by AI is particularly dangerous because of the aura of credibility these systems can project. In a RAG system, misinformation primarily arises from two failure points: Retrieval of inaccurate content from the knowledge base and fabrication or distortion by the large language model (LLM) during generation, even when given good context. The first line of defense is ensuring the integrity of the knowledge base. A RAG system is only as reliable as the documents it has access to. If non-credible, manipulated, or satirical sources are ingested, the system will retrieve and use them as fact. This makes rigorous data curation and source validation the most critical step in combating misinformation. The second line of defense is strengthening the connection between retrieval and generation to prevent the LLM from 'going off script'. The LLM, based on its pre-trained knowledge, might confidently generate an answer that contradicts the provided evidence or adds unsupported details – a phenomenon known as 'hallucination'. To mitigate this, thesystem must be designed to strictly adhere to the retrieved context." (Saloni Garg et al, "RAG Artificial Intelligence: Retrieval-Augmented Generation in Generative AI", 2026)

"The foundational form of RAG, often called naive RAG, follows a straightforward pattern. A pipeline retrieves supporting context from external sources such as enterprise documents, knowledge bases, or structured datasets and appends that information to the model’s prompt before inference. In the most common implementation, each document is converted into an embedding, a numerical representation of its semantic meaning, using either the same foundation model or a specialized embedding model. When a user submits a query, the system performs a vector similarity search to find documents whose embeddings most closely match the query’s vector representation, and the retrieved content is concatenated with the user query before being passed to the language model." (Bennie Haelen, "ML and Generative AI in the Data Lakehouse Building and Deploying AI Applications at Scale", 2026)

No comments:

Related Posts Plugin for WordPress, Blogger...

About Me

My photo
Koeln, NRW, Germany
IT Professional with more than 25 years experience in IT in the area of full life-cycle of Web/Desktop/Database Applications Development, Software Engineering, Consultancy, Data Management, Data Quality, Data Migrations, Reporting, ERP implementations & support, Team/Project/IT Management, etc.