Showing posts sorted by date for query Systems Engineering. Sort by relevance Show all posts
Showing posts sorted by date for query Systems Engineering. Sort by relevance Show all posts

10 October 2026

🤖Prompt Engineering: Issues (Just the Quotes)

"Chain-of-thought prompting is a method that forces LLMs to reason through a series of steps, resulting in more structured, transparent, and precise outputs. The goal is to break down complex tasks into smaller, interconnected subtasks, allowing the LLM to address each subtask in a stepby-step manner. This not only helps the model to 'focus' on specific aspects of the problem, but also encourages it to generate intermediate outputs, making it easier to identify and debug potential issues along the way. Another significant advantage of chain-of-thought prompting is the improved interpretability and transparency of the LLM-generated response. By offering insights into the model’s reasoning process, we, as users, can better understand and qualify how the final output was derived, which promotes trust in the model’s decision-making abilities." (Sinan Ozdemir, "Quick Start Guide to Large Language Models: Strategies and Best Practices for Using ChatGPT and Other LLMs", 2024)

"Agentic workflows break when the logic is messy - if, say, the plans don’t decompose or memory is poorly structured. However, infrastructure-level LLM applications introduce even more failure points and complexity. If the protocols don’t sync with each other, or the data flows start leaking, or the model boundaries are unclear... there are far too many failure points to count. While most people have been jumping on the bandwagon to adopt MCPs or A2A, very few are equipped to handle the LLMOps issues these tools introduce." (Abi Aryan, "LLMOps: Managing Large Language Models in Production", 2025)

"As with many other deep learning-based approaches, another major challenge is in interpretability. While knowledge graphs provide a structured and transparent way to store relationships, LLMs operate as a black box, making it difficult to understand how specific outputs are generated. [...] Data alignment is also a key issue, as structured knowledge graphs and unstructured text data must be carefully preprocessed to ensure consistency.  Differences in data formats, ontology mismatches, and information redundancy can create inefficiencies when integrating these two paradigms. Developing robust pipelines that seamlessly connect graph-based insights with LLM-generated text remains an open challenge." (Aldo Marzullo et al, "Graph Machine Learning" 2nd Ed., 2025)

"The art of mega-prompts spanning multiple written pages and looking like essays has become commonplace for complex tasks when building applications to get things `just right'. Unfortunately, they bring with them lots of issues: errors, portability, complexity, and more. The GenAI world didn’t plan for mega-prompts. They have simply evolved into what they’ve become today because practitioners kept wanting to do more and more complex things, and their only way to express those intents was with a prompt. But step back and look at some of these prompts [...] Lurking just below the surface are a bunch of classical computing concepts like data, programming instructions, control flows, memory, and stora - all the components typically associated with classical computing elements." (Rob Thomas et al, "AI Value Creators: Beyond the Generative AI User Mindset", 2025)

"A solid data foundation is critical for AI success. The foundation can grow in phases, but it needs to be there. AI will add pressure to the data foundation. Modern workloads require data that is connected, interpretable, high quality, and governed. Modern workloads utilize diverse data, and this is especially true with the popularity of GenAI. Trustworthy AI depends on knowing where the data came from and how it has changed. Modern architectures including the lakehouse and the data fabric exist because the traditional ways of managing data no longer fit the scale or complexity of what organizations need to do today. In other words, there is no path toward enterprise AI without a strong data foundation. This was echoed in the expert advices. Experts spoke about the dangers of 'vanity projects', particularly with GenAI, when foundational data issues were ignored. Others described how their companies moved faster specifically because they had already invested in governance, semantic layers, lineage, and observability." (Fern Halper, "Data Makes the World Go 'Round", 2026)

"As AI systems become more autonomous and integrated into decision-making processes, the question of accountability grows in importance. Businesses and developers must ensure that AI systems operate within predefined ethical boundaries and that there are mechanisms to identify and correct issues when things go wrong. Because AI models can make decisions that directly impact people’s lives (such as credit approvals, hiring decisions, or healthcare recommendations), organizations must maintain accountability for these outcomes. Clear guidelines should be in place to determine who is responsible when an AI system makes a mistake or causes harm."  (Bennie Haelen, "ML and Generative AI in the Data Lakehouse Building and Deploying AI Applications at Scale", 2026)

"Driven by a fear of missing out (FOMO), many companies have launched AI initiatives. But many failed. For example, customer service chatbots have been introduced to handle inquiries, reduce costs and improve efficiency. But chatbots failed to grasp the complexity of customer issues; they created frustration rather than solving problems. Customers were stuck in repetitive loops, being asked the same questions, and unable to reach a human when needed. Instead of enhancing the customer experience, they delivered spikes in customer complaints instead." (Alan Watkins & G C Cooke, "Smarter than You Winning in Business with Superintelligent AI", 2026)

"With autonomous agents there is the risk that they can take action that isn’t governed. These systems introduce autonomy, interdependence between agents, and possibly emergent behavior. [...] Because agents act autonomously, a single misconfigured or misaligned agent can propagate errors at scale. In multiagent environments, one faulty output can trigger many downstream mistakes. Other risks include tool misuse, conflicting goals among agents (and other interoperability issues), and operational opacity, when teams cannot easily determine which agent took which action or why. Governance must extend to agent registration, version control, permissioning, and simulation testing before deployment. Likewise, accountability may also blur. If an agent takes a dangerous action, who is responsible? There are, of course, cybersecurity risks as agents pose a new attack surface."  (Fern Halper, "Data Makes the World Go 'Round", 2026)

09 October 2026

🪙Business Intelligence: Single Source of Truth [SSoT] (Just the Quotes)

"It is important to remember that the 'single version of the truth' - or enterprise logical data model - is not and should not be built all at once (that would take too long), but that it evolves over time as the project-specific logical data models are merged, one-by-one, a project at a time." (Sid Adelman et al, "Data Strategy", 2005

"Having multiple data lakes replicates the same problems that were created with multiple data warehouses - disparate data siloes and data fiefdoms that don't facilitate sharing of the corporate data assets across the organization. Organizations need to have a single data lake from which they can source the data for their BI/data warehousing and analytic needs. The data lake may never become the 'single version of the truth' for the organization, but then again, neither will the data warehouse. Instead, the data lake becomes the 'single or central repository for all the organization's data' from which all the organization's reporting and analytic needs are sourced." (Billl Schmarzo, "Driving Business Strategies with Data Science: Big Data MBA" 1st Ed., 2015)

"Data warehousing, as we are aware, is the traditional approach of consolidating data from multiple source systems and combining into one store that would serve as the source for analytical and business intelligence reporting. The concept of data warehousing resolved the problems of data heterogeneity and low-level integration. In terms of objectives, a data lake is no different from a data warehouse. Both are primary advocates of terms like 'single source of truth' and 'central data repository'." (Saurabh Gupta et al, "Practical Enterprise Data Lake Insights", 2018)

"Another myth is that we shall have a single source of truth for each concept or entity. […] This is a wonderful idea, and is placed to prevent multiple copies of out-of-date and untrustworthy data. But in reality it’s proved costly, an impediment to scale and speed, or simply unachievable. Data Mesh does not enforce the idea of one source of truth. However, it places multiple practices in place that reduces the likelihood of multiple copies of out-of-date data." (Zhamak Dehghani, "Data Mesh: Delivering Data-Driven Value at Scale", 2021)

"Data mesh [...] reduces points of centralization that act as coordination bottlenecks. It finds a new way of decomposing the data architecture without slowing the organization down with synchronizations. It removes the gap between where the data originates and where it gets used and removes the accidental complexities - aka pipelines - that happen in between the two planes of data. Data mesh departs from data myths such as a single source of truth, or one tightly controlled canonical data model." (Zhamak Dehghani, "Data Mesh: Delivering Data-Driven Value at Scale", 2021)

"Data mesh [...] reduces points of centralization that act as coordination bottlenecks. It finds a new way of decomposing the data architecture without slowing the organization down with synchronizations. It removes the gap between where the data originates and where it gets used and removes the accidental complexities - aka pipelines - that happen in between the two planes of data. Data mesh departs from data myths such as a single source of truth, or one tightly controlled canonical data model." (Zhamak Dehghani, "Data Mesh: Delivering Data-Driven Value at Scale", 2021)

"Unlike other analytical data management paradigms, data mesh does not embrace the concept of the mythical single source of truth. Every data product provides a truthful portion of the reality - for a particular domain - to the best of its ability, a single slice of truth." (Zhamak Dehghani, "Data Mesh: Delivering Data-Driven Value at Scale", 2021)

"Lakehouse is a new architecture and data storage paradigm that combines the characteristics of both data warehouses and data lakes to create a unified basis for all types of use cases to be built on top of it. There is no need to move data around. Data is curated and remains in an open format and serves as the single source of truth (SSOT) for all the consumption layers. A modern data platform has needs that span traditional data warehouses, data lakes, machine learning systems, and streaming systems and there is some overlap among these systems. A Lakehouse offers features that span all four systems [...]" (Anindita Mahapatra, "Simplifying Data Engineering and Analytics with Delta", 2022)

"The Data Fabric architecture needs to guarantee this single version of the truth within the application and transactional landscape, which – depending on the deployment option of an MDM solution – could also mean to assemble this single version of the truth based on core information that is dispersed and maintained in various data stores." (Eberhard Hechler et al, "Data Fabric and Data Mesh Approaches with AI", 2023)

"The main goal of the transaction log is to enable multiple readers and writers to operate on a given version of a dataset file simultaneously and to provide additional information, like data skipping indexes to the execution engine for more performant operations. The Delta Lake transaction log always shows the user a consistent view of the data and serves as a single source of truth. It is the central repository that tracks all changes the user makes to a Delta table." (Bennie Haelen & Dan Davis, "Delta Lake: Up and Running - Modern Data Lakehouse Architectures with Delta Lake", 2023)

"The hub and spoke, or 'star network', is a data architecture model that centralizes data from various sources into a single hub, such as a data warehouse or data lake. The hub serves as the source of truth for data and provides standardized schemas and formats. The spokes are the various applications or services that consume data from the hub for different purposes, such as analytics, reporting, or ma-chine learning. Spokes can also perform transformations or aggregations on data before presenting it to end users. The hub and spoke architecture aims to simplify data integration and management by reducing complexity and redundancy in data pipelines" (Christopher Maneu et al, "The Definitive Guide to Microsoft Fabric From discovery to building a unified, secure, and scalable data platform", 2025)

"Traditional paradigms build systems for specific purposes with structured, homogeneous databases. This approach works for tailored needs but is impractical for complex domains that need to adapt to user characteristics and integrate heterogeneous data. KGs capture connections, enabling relationship discovery through graph pattern matching and traversal. Both the Resource Description Framework (RDF) and Labeled Property Graphs (LPGs) provide machine-readable formats that humans can interpret. KGs emphasize rich, meaningful data representations usable by both humans and machines, enabling a paradigm shift where intelligent behavior is encoded in a unique source of truth." (Alessandro Negro et al, "Knowledge Graphs and LLMs in Action", 2026)

🤖Prompt Engineering: Trust (Just the Quotes)

"Chain-of-thought prompting is a method that forces LLMs to reason through a series of steps, resulting in more structured, transparent, and precise outputs. The goal is to break down complex tasks into smaller, interconnected subtasks, allowing the LLM to address each subtask in a stepby-step manner. This not only helps the model to 'focus' on specific aspects of the problem, but also encourages it to generate intermediate outputs, making it easier to identify and debug potential issues along the way. Another significant advantage of chain-of-thought prompting is the improved interpretability and transparency of the LLM-generated response. By offering insights into the model’s reasoning process, we, as users, can better understand and qualify how the final output was derived, which promotes trust in the model’s decision-making abilities." (Sinan Ozdemir, "Quick Start Guide to Large Language Models: Strategies and Best Practices for Using ChatGPT and Other LLMs", 2024)

"A 'hallucination' in the context of generative AI refers to the phenomenon where a model produces information that is factually incorrect, nonsensical, or not grounded in its input data or pre-existing knowledge. These are not mere typos or minor inaccuracies; they are confident, coherent, and often persuasive fabrications. In high-stakes domains like healthcare, law, or finance, a single hallucination can have severe consequences, eroding user trust and leading to catastrophic decision-making. While all LLMs are prone to this, the RAG architecture is specifically designed to combat it by tethering the model’s output to an external, verifiable knowledge base. Understanding why hallucinations occur is the essential first step to building more reliable and truthful AI systems." (Saloni Garg et al, "RAG Artificial Intelligence: Retrieval-Augmented Generation in Generative AI", 2026)

"A solid data foundation is critical for AI success. The foundation can grow in phases, but it needs to be there. AI will add pressure to the data foundation. Modern workloads require data that is connected, interpretable, high quality, and governed. Modern workloads utilize diverse data, and this is especially true with the popularity of GenAI. Trustworthy AI depends on knowing where the data came from and how it has changed. Modern architectures including the lakehouse and the data fabric exist because the traditional ways of managing data no longer fit the scale or complexity of what organizations need to do today. In other words, there is no path toward enterprise AI without a strong data foundation. This was echoed in the expert advices. Experts spoke about the dangers of 'vanity projects', particularly with GenAI, when foundational data issues were ignored. Others described how their companies moved faster specifically because they had already invested in governance, semantic layers, lineage, and observability." (Fern Halper, "Data Makes the World Go 'Round", 2026

"As AI systems become more integrated into critical decision-making processes, the 'black box' problem – the inability to understand how a model arrived at a specific output – becomes a major barrier to trust and adoption. For RAG systems, this is particularly crucial; a user needs to know not just the answer, but why the system believes that answer to be true. Transparency and explainability (explainable AI, XAI) are the disciplines focused on making AI reasoning understandable to humans. A transparent RAG system allows users to verify the accuracy of its responses, builds trust by demonstrating a logical process, and enables developers to debug and improve the system. It transforms the AI from an oracle that must be blindly trusted into a tool for augmented intelligence, where the human remains the ultimate arbiter of truth." (Saloni Garg et al, "RAG Artificial Intelligence: Retrieval-Augmented Generation in Generative AI", 2026)

"Generative AI has a lot of problems because it needs a lot of data to work. Generative models like GPT and GANs need a lot of training data to learn patterns and make good outputs. This data dependency can cause problems like overfitting, which happens when the model does well on training data but doesn’t work well with new, unseen data. Also, the quality of the content that is generated is directly related to the diversity and representativeness of the training data. This means that biased or incomplete datasets can lead to outputs that are wrong or unfair." (Saloni Garg et al, "RAG Artificial Intelligence: Retrieval-Augmented Generation in Generative AI", 2026)

"If ethical lapses or AI failures occur, the impact on a business can be significant. Misinformation, biases, or harmful content generated by AI can lead to reputational damage, customer distrust, and potential regulatory scrutiny. The public relations fallout from an AI-driven error or ethical misstep can erode consumer confidence, resulting in lost revenue and lasting harm to brand image. Businesses, therefore, need to proactively address ethical considerations in AI implementation, not only to ensure compliance but also to protect and strengthen their reputation in a highly competitive, and increasingly transparent, marketplace." (Bennie Haelen, "ML and Generative AI in the Data Lakehouse Building and Deploying AI Applications at Scale", 2026)

"RAG applications must be built with semantics, metadata, and governance in mind. The retrieved information must be high-quality, secure, and appropriate for the user’s role. Equally important is monitoring and management: checking whether source data has changed, ensuring vector stores remain accurate, and watching for hallucinations or data leakage. Organizations are definitely starting to experiment with RAG models today; some are putting them into production applications. Some believe that using RAG helps mitigate hallucinations because it is grounded in trusted organizational data." (Fern Halper, "Data Makes the World Go 'Round", 2026)

"RAG models offer several advantages over standard generative models, addressing many of their limitations. Traditional generative models, like GPT, rely solely on patterns learned during training, which can lead to issues such as factual inaccuracies, model hallucinations, and lack of contextual relevance. These models generate content based on pre-existing knowledge, often without the ability to verify or update the information, making them less reliable for tasks requiring high accuracy. In contrast, RAG models integrate retrieval mechanisms that allow them to access external knowledge sources in real time. This ensures that the generated content is grounded in verified data, significantly improving factual consistency and relevance. For example, while a standard generative model might produce a plausible but incorrect answer to a factual question, a RAG model can retrieve and incorporate accurate information from a trusted source, reducing errors." (Saloni Garg et al, "RAG Artificial Intelligence: Retrieval-Augmented Generation in Generative AI", 2026)

"The power of RAG systems stems from their ability to access and process vast amounts of data. However, this very capability introduces significant risks regarding the privacy of individuals and the security of sensitive information. Unlike a simple chatbot, a RAG system often has access to proprietary corporate data, internal documentation, and potentially personal user information within its knowledge base. A data breach or misuse of this information can lead to severe financial, legal, and reputational damage. Furthermore, a global patchwork of stringent regulations now governs how personal data must be handled, making compliance a central pillar of AI system design, not an afterthought. Ethical deployment requires an architecture built on privacy by design and by default, ensuring user trust is maintained through robust technical and procedural safeguards." (Saloni Garg et al, "RAG Artificial Intelligence: Retrieval-Augmented Generation in Generative AI", 2026)

"The promise of AI is its ability to process information objectively and at scale. However, this promise is fundamentally threatened by the twin challenges of bias and misinformation. AI systems are not born in a vacuum; they are created by humans and trained on data produced by humans. Consequently, they are prone to inheriting and even amplifying our prejudices, errors, and the systemic inequalities present in that data. In a RAG system, this risk is a two-fold problem: first in the retrieval of information, and second in the generation of a response based on that retrieval. A failure to address these issues doesn’t just lead to technically incorrect outputs; it can perpetuate social harm, erode public trust, and lead to the widespread dissemination of falsehoods. Therefore, understanding and mitigating bias and misinformation is not an optional add-on but a core requirement for any ethically deployed AI system." (Saloni Garg et al, "RAG Artificial Intelligence: Retrieval-Augmented Generation in Generative AI", 2026)

08 October 2026

🪙Business Intelligence: Signals (Just the Quotes)

"Data are generally collected as a basis for action. However, unless potential signals are separated from probable noise, the actions taken may be totally inconsistent with the data. Thus, the proper use of data requires that you have simple and effective methods of analysis which will properly separate potential signals from probable noise." (Donald J Wheeler, "Understanding Variation: The Key to Managing Chaos" 2nd Ed., 2000)

"No matter what the data, and no matter how the values are arranged and presented, you must always use some method of analysis to come up with an interpretation of the data.
While every data set contains noise, some data sets may contain signals. Therefore, before you can detect a signal within any given data set, you must first filter out the noise." (Donald J Wheeler," Understanding Variation: The Key to Managing Chaos" 2nd Ed., 2000)

"Data analysis is not generally thought of as being simple or easy, but it can be. The first step is to understand that the purpose of data analysis is to separate any signals that may be contained within the data from the noise in the data. Once you have filtered out the noise, anything left over will be your potential signals. The rest is just details." (Donald J Wheeler," Myths About Data Analysis", International Lean & Six Sigma Conference, 2012)

"A complete data analysis will involve the following steps: (i) Finding a good model to fit the signal based on the data. (ii) Finding a good model to fit the noise, based on the residuals from the model. (iii) Adjusting variances, test statistics, confidence intervals, and predictions, based on the model for the noise." (DeWayne R Derryberry, "Basic data analysis for time series with R", 2014

"Data catalogs that focus on search bring techniques and methods from information retrieval and web search engines to the data domain within enterprises. Some of those catalogs, such as Facebook Nemo, use advanced machine learning and NLP tools to provide personalized search of data within an enterprise. The search can also use data-specific signals such as usage, popularity, and freshness to rank data assets by usefulness." (Fadi Maali & Jason Lim, "Implementing a Modern Data Catalog to Power Data Intelligence: Make Trustworthy Data Central to Your Organization", 2022)

"In a self-service environment with multiple publishers, it’s impossible to completely avoid data redundancy and overlapping. Multiple data assets with similar content, but possibly with varying quality, will exist. A data catalog can guide users to trusted data that comes from a reliable source and is frequently used. A data catalog can also use various explicit and implicit quality signals when ranking datasets for recommendation. Some of those signals are discussed next. Furthermore, a data catalog can recommend domain experts who are automatically identified based on actual data usage." (Fadi Maali & Jason Lim, "Implementing a Modern Data Catalog to Power Data Intelligence: Make Trustworthy Data Central to Your Organization", 2022)

"Many argue that model drift is best monitored by monitoring the data drift in incoming data and the drift in the generated features. As and when the ground truth is available, it is joined by some primary key criteria with the inference data in a Delta table. Again, the update and merge operation support in Delta makes this a breeze. Now the actual and predicted values of the inference data are computed to see how well the model is doing in terms of the quality of insight generation. The feature engineering pipeline is completely in-house and is easier to monitor for drift. The model interpretability may indicate that some columns contributing to the predictive power are incorrect, and it may be necessary to add or remove features. In such cases, a threshold of tolerance is violated, which signals a need for model retraining." (Anindita Mahapatra, "Simplifying Data Engineering and Analytics with Delta", 2022)

"Establishing a comprehensive observability architecture necessitates a systematic approach that spans the entirety of the data pipeline, from initial telemetry collection to actionable insights accessible by diverse stakeholders. The core objective is to unify distributed data sources - metrics, logs, traces, and quality signals - into a coherent framework that enables rapid diagnosis, continuous monitoring, and strategic decision-making." (William Smith, "Soda Core for Modern Data Quality and Observability: The Complete Guide for Developers and Engineers", 2025)

"A defining attribute of the domain-oriented ownership principle is the focus on context preservation in Data Management. This aspect accentuates the importance of keeping data within its native domain environment, allowing it to retain its original context, value, and meaning. When data is managed close to its source, its contextual richness is preserved. This sharply contrasts centralized models, where data is often abstracted from its source, leading to potential loss of signal or context. When data remains within its generating domain, it retains the nuances and specificities unique to its activities, challenges, and goals." (Pradeep Menon, "Data Mesh Principles, patterns, architecture, and strategies for data-driven decision making", 2024)

 More on Signals in Data Science, Graphical Representation, Systems Engineering

🕸Systems Engineering: Signals (Just the Quotes)

"The theory of communication is partly concerned with the measurement of information content of signals, as their essential property in the establishment of communication links. But the information content of signals is not to be regarded as a commodity; it is more a property or potential of the signals, and as a concept it is closely related to the idea of selection, or discrimination. This mathematical theory first arose in telegraphy and telephony, being developed for the purpose of measuring the information content of telecommunication signals. It concerned only the signals themselves as transmitted along wires, or broadcast through the aether, and is quite abstracted from all questions of 'meaning'. Nor does it concern the importance, the value, or truth to any particular person. As a theory, it lies at the syntactic level of sign theory and is abstracted from the semantic and pragmatic levels. We shall argue [...] that, though the theory does not directly involve biological elements, it is nevertheless quite basic to the study of human communication - basic but insufficient." (Colin Cherry, "On Human Communication", 1957)) 

"The term closed loop-learning process refers to the idea that one learns by determining what s desired and comparing what is actually taking place as measured at the process and feedback for comparison. The difference between what is desired and what is taking place provides an error indication which is used to develop a signal to the process being controlled." (Harold Chestnut, 1984) 

"In a real experiment the noise present in a signal is usually considered to be the result of the interplay of a large number of degrees of freedom over which one has no control. This type of noise can be reduced by improving the experimental apparatus. But we have seen that another type of noise, which is not removable by any refinement of technique, can be present. This is what we have called the deterministic noise. Despite its intractability it provides us with a way to describe noisy signals by simple mathematical models, making possible a dynamical system approach to the problem of turbulence." (David Ruelle, "Chaotic Evolution and Strange Attractors: The statistical analysis of time series for deterministic nonlinear systems", 1989)

"The cybernetics phase of cognitive science produced an amazing array of concrete results, in addition to its long-term" (often underground) influence: the use of mathematical logic to understand the operation of the nervous system; the invention of information processing machines" (as digital computers), thus laying the basis for artificial intelligence; the establishment of the metadiscipline of system theory, which has had an imprint in many branches of science, such as engineering" (systems analysis, control theory), biology" (regulatory physiology, ecology), social sciences" (family therapy, structural anthropology, management, urban studies), and economics" (game theory); information theory as a statistical theory of signal and communication channels; the first examples of self-organizing systems. This list is impressive: we tend to consider many of these notions and tools an integrative part of our life […]" (Francisco Varela, "The Embodied Mind", 1991)

"An artificial neural network is an information-processing system that has certain performance characteristics in common with biological neural networks. Artificial neural networks have been developed as generalizations of mathematical models of human cognition or neural biology, based on the assumptions that: 1. Information processing occurs at many simple elements called neurons. 2. Signals are passed between neurons over connection links. 3. Each connection link has an associated weight, which, in a typical neural net, multiplies the signal transmitted. 4. Each neuron applies an activation function" (usually nonlinear) to its net input" (sum of weighted input signals) to determine its output signal." (Laurene Fausett, "Fundamentals of Neural Networks", 1994)

"I wage war on noise every day as part of my work as a scientist and engineer. We try to maximize signal-to-noise ratios. We try to filter noise out of measurements of sounds or images or anything else that conveys information from the world around us. We code the transmission of digital messages with extra 0s and 1s to defeat line noise and burst noise and any other form of interference. We design sophisticated algorithms to track noise and then cancel it in headphones or in a sonogram. Some of us even teach classes on how to defeat this nemesis of the digital age. Such action further conditions our anti-noise reflexes." (Bart Kosko, "Noise", 2006)

"Perturbations are often regarded as noise. What is the difference? Noise is usually understood from the point of the experimenter. If we measure it from the outside, noise is the fluctuation of the value we measure. However, from the point of view of the network, noise is a series ofperturbations changing its original status. Network perturbations can be called either signals or noise." (Péter Csermely, "Weak Links: The Universal Key to the Stabilityof Networks and Complex Systems", 2009)

"To understand, how noise is related to scale-freeness, we have to do some mathematics again. Noise is usually characterized by a mathematical trick. The seemingly random fluctuation of the signal is regarded as a sum of sinusoidal waves. The components of the million waves giving the final noise structure are characterized by their frequency. To describe noise, we plot the contribution (called spectral density) of the various waves we use to model the noise as a function of their frequency. This transformation is called a Fourier transformation [...]" (Péter Csermely, "Weak Links: The Universal Key to the Stabilityof Networks and Complex Systems", 2009)

"Complex systems seem to have this property, with large periods of apparent stasis marked by sudden and catastrophic failures. These processes may not literally be random, but they are so irreducibly complex" (right down to the last grain of sand) that it just won’t be possible to predict them beyond a certain level. […] And yet complex processes produce order and beauty when you zoom out and look at them from enough distance." (Nate Silver, "The Signal and the Noise: Why So Many Predictions Fail-but Some Don't", 2012)

"Information theory leads to the quantification of the information content of the source, as denoted by entropy, the characterization of the information-bearing capacity of the communication channel, as related to its noise characteristics, and consequently the establishment of the relationship between the information content of the source and the capacity of the channel. In short, information theory provides a quantitative measure of the information contained in message signals and help determine the capacity of a communication system to transfer this information from source to sink over a noisy channel in a reliable fashion." (Ali Grami, "Information Theory", 2016)

More on Signals in Business Intelligence, Data Science, Graphical Representation

07 October 2026

🕸Systems Engineering: Noise (Just the Quotes)

"Higher, directed forms of energy (e.g., mechanical, electric, chemical) are dissipated, that is, progressively converted into the lowest form of energy, i.e., undirected heat movement of molecules; chemical systems tend toward equilibria with maximum entropy; machines wear out owing to friction; in communication channels, information can only be lost by conversion of messages into noise but not vice versa, and so forth." (Ludwig von Bertalanffy, "Robots, Men and Minds", 1967)

"To adapt to a changing environment, the system needs a variety of stable states that is large enough to react to all perturbations but not so large as to make its evolution uncontrollably chaotic. The most adequate states are selected according to their fitness, either directly by the environment, or by subsystems that have adapted to the environment at an earlier stage. Formally, the basic mechanism underlying self-organization is the (often noise-driven) variation which explores different regions in the system’s state space until it enters an attractor. This precludes further variation outside the attractor, and thus restricts the freedom of the system’s components to behave independently. This is equivalent to the increase of coherence, or decrease of statistical entropy, that defines self-organization." (Francis Heylighen, "The Science Of Self-Organization And Adaptivity", 1970)

"The power and beauty of stochastic approximation theory is that it provides simple, easy to implement gain sequences which guarantee convergence without depending (explicitly) on knowledge of the function to be minimized or the noise properties. Unfortunately, convergence is usually extremely slow. This is to be expected, as 'good performance' cannot be expected if no (or very little) knowledge of the nature of the problem is built into the algorithm. In other words, the strength of stochastic approximation (simplicity, little a priori knowledge) is also its weakness." (Fred C Scweppe, "Uncertain dynamic systems", 1973)

"In a real experiment the noise present in a signal is usually considered to be the result of the interplay of a large number of degrees of freedom over which one has no control. This type of noise can be reduced by improving the experimental apparatus. But we have seen that another type of noise, which is not removable by any refinement of technique, can be present. This is what we have called the deterministic noise. Despite its intractability it provides us with a way to describe noisy signals by simple mathematical models, making possible a dynamical system approach to the problem of turbulence." (David Ruelle, "Chaotic Evolution and Strange Attractors: The statistical analysis of time series for deterministic nonlinear systems", 1989)

"Black-noise phenomena govern natural and unnatural catastrophes like floods, droughts, bear markets, and various outrageous outages, such as those of electrical power. Because of their black spectra, such disasters often come in clusters." (Manfred R Schroeder, "Fractals, Chaos, Power Laws", 1991)

"What we now call chaos is a time evolution with sensitive dependence on initial condition. The motion on a strange attractor is thus chaotic. One also speaks of deterministic noise when the irregular oscillations that are observed appear noisy, but the mechanism that produces them is deterministic." (David Ruelle, "Chance and Chaos", 1991)

"An essential element of dynamics systems is a positive feedback that self-enhances the initial deviation from the mean. The avalanche is proverbial. Cities grow since they attract more people, and in the universe, a local accumulation of dust may attract more dust, eventually leading to the birth of a star. Earlier or later, self-enhancing processes evoke an antagonistic reaction. A collapsing stock market stimulates the purchase of shares at a low price, thereby stabilizing the market. The increasing noise, dirt, crime and traffic jams may discourage people from moving into a big city." (Hans Meinhardt, "The Algorithmic Beauty of Sea Shells", 1995)

"Chaos can leave statistical footprints that look like noise. This can arise from simple systems that are deterministic and not random. [...] The surprising mathematical fact is that most systems are chaotic. Change the starting value ever so slightly and soon the system wanders off on a new chaotic path no matter how close the starting point of the new path was to the starting point of the old path. Mathematicians call this sensitivity to initial conditions but many scientists just call it the butterfly effect. And what holds in math seems to hold in the real world - more and more systems appear to be chaotic." (Bart Kosko, "Noise", 2006)

"'Chaos' refers to systems that are very sensitive to small changes in their inputs. A minuscule change in a chaotic communication system can flip a 0 to a 1 or vice versa. This is the so-called butterfly effect: Small changes in the input of a chaotic system can produce large changes in the output. Suppose a butterfly flaps its wings in a slightly different way. can change its flight path. The change in flight path can in time change how a swarm of butterflies migrates." (Bart Kosko, "Noise", 2006)

"I wage war on noise every day as part of my work as a scientist and engineer. We try to maximize signal-to-noise ratios. We try to filter noise out of measurements of sounds or images or anything else that conveys information from the world around us. We code the transmission of digital messages with extra 0s and 1s to defeat line noise and burst noise and any other form of interference. We design sophisticated algorithms to track noise and then cancel it in headphones or in a sonogram. Some of us even teach classes on how to defeat this nemesis of the digital age. Such action further conditions our anti-noise reflexes." (Bart Kosko, "Noise", 2006)

"Linear systems do not benefit from noise because the output of a linear system is just a simple scaled version of the input [...] Put noise in a linear system and you get out noise. Sometimes you get out a lot more noise than you put in. This can produce explosive effects in feedback systems that take their own outputs as inputs." (Bart Kosko, "Noise", 2006)

"This phenomenon, common to chaos theory, is also known as sensitive dependence on initial conditions. Just a small change in the initial conditions can drastically change the long-term behavior of a system. Such a small amount of difference in a measurement might be considered experimental noise, background noise, or an inaccuracy of the equipment." (Greg Rae, Chaos Theory: A Brief Introduction, 2006)

"Neural networks are a popular model for learning, in part because of their basic similarity to neural assemblies in the human brain. They capture many useful effects, such as learning from complex data, robustness to noise or damage, and variations in the data set. " (Peter C R Lane, Order Out of Chaos: Order in Neural Networks, 2007)

"Noise is bad for the network, if high and continuous noise levels disturb all network functions. So far, the take-home message is that we have to stop noise in order to survive. This assumption is wrong. Reducing the noise to zero would mean no interaction of the network with the environment. Isolation is clearly a bad strategy, since such an isolated network will die. However, zero noise is bad for another reason too. Noise can be helpful in many ways. The first documented observations of good noise were sailors’ reports on the peculiar phenomenon that  disordered raindrops falling on the ocean can calm roughseas. Another example of the optimal level of noise is opinion formation. A low noise is not enough for modulation of opinion formation, while strong fluctuations prevent the formation of a definitive collective opinion." (Péter Csermely, "Weak Links: The Universal Key to the Stabilityof Networks and Complex Systems", 2009)

"Perturbations are often regarded as noise. What is the difference? Noise is usually understood from the point of the experimenter. If we measure it from the outside, noise is the fluctuation of the value we measure. However, from the point of view of the network, noise is a series ofperturbations changing its original status. Network perturbations can be called either signals or noise." (Péter Csermely, "Weak Links: The Universal Key to the Stabilityof Networks and Complex Systems", 2009)

"Self-organizing networks suffer various types of random damage. Therefore, if the network remained static, it would soon become dysfunctional. Some networks have developed highly specific screening systems which recognize and repair random damage. On the one hand, this process requires energy, which arrives in the form of perturbations or noise. On the other hand, noise-triggered network restructuring will repeat a few steps of the original self-organization and therefore constitutes a much cheaper way of providing a continuous repair function, with the additional advantage that it is always adaptive with respect to the actual environment of the network." (Péter Csermely, "Weak Links: The Universal Key to the Stabilityof Networks and Complex Systems", 2009)

"To understand, how noise is related to scale-freeness, we have to do some mathematics again. Noise is usually characterized by a mathematical trick. The seemingly random fluctuation of the signal is regarded as a sum of sinusoidal waves. The components of the million waves giving the final noise structure are characterized by their frequency. To describe noise, we plot the contribution (called spectral density) of the various waves we use to model the noise as a function of their frequency. This transformation is called a Fourier transformation [...]" (Péter Csermely, "Weak Links: The Universal Key to the Stabilityof Networks and Complex Systems", 2009)

"When some systems are stuck in a dangerous impasse, randomness and only randomness can unlock them and set them free. You can see here that absence of randomness equals guaranteed death. The idea of injecting random noise into a system to improve its functioning has been applied across fields. By a mechanism called stochastic resonance, adding random noise to the background makes you hear the sounds (say, music) with more accuracy." (Nassim N Taleb, "Antifragile: Things that gain from disorder", 2012)

"It is evident that chaotic behavior, in the new scientific sense of the term, is very different from random, erratic motion. With the help of strange attractors a distinction can be made between mere randomness, or 'noise', and chaos. Chaotic behavior is deterministic and patterned, and strange attractors allow us to transform the seemingly random data into distinct visible shapes." (Fritjof Capra, "The Systems View of Life: A Unifying Vision", 2014)

"Neural networks can model very complex patterns and decision boundaries in the data and, as such, are very powerful. In fact, they are so powerful that they can even model the noise in the training data, which is something that definitely should be avoided. One way to avoid this overfitting is by using a validation set in a similar way as with decision trees.[...] Another scheme to prevent a neural network from overfitting is weight regularization, whereby the idea is to keep the weights small in absolute sense because otherwise they may be fitting the noise in the data. This is then implemented by adding a weight size term (e.g., Euclidean norm) to the objective function of the neural network." (Bart Baesens, "Analytics in a Big Data World: The Essential Guide to Data Science and Its Applications", 2014)

🤖Prompt Engineering: Accuracy (Just the Quotes)

"Attention is a mechanism used in deep learning models (not just Transformers) that assigns different weights to different parts of the input, allowing the model to prioritize and emphasize the most important information while performing tasks like translation or summarization. Essentially, attention allows a model to 'focus' on different parts of the input dynamically, leading to improved performance and more accurate results. Before the popularization of attention, most neural networks processed all inputs equally and the models relied on a fixed representation of the input to make predictions. Modern LLMs that rely on attention can dynamically focus on different parts of input sequences, allowing them to weigh the importance of each part in making predictions." (Sinan Ozdemir, "Quick Start Guide to Large Language Models: Strategies and Best Practices for Using ChatGPT and Other LLMs", 2024)

"Fine-tuning involves training the LLM on a smaller, task-specific dataset to adjust its parameters for the specific task at hand. This allows the LLM to leverage its pre-trained knowledge of the language to improve its accuracy for the specific task. Fine-tuning has been shown to drastically improve performance on domain-specific and task-specific tasks and lets LLMs adapt quickly to a wide variety of NLP applications." (Sinan Ozdemir, "Quick Start Guide to Large Language Models: Strategies and Best Practices for Using ChatGPT and Other LLMs", 2024)

"Large language models (LLMs) are AI models that are usually (but not necessarily) derived from the Transformer architecture and are designed to understand and generate human language, code, and much more. These models are trained on vast amounts of text data, allowing them to capture the complexities and nuances of human language. LLMs can perform a wide range of language-related tasks, from simple text classification to text generation, with high accuracy, fluency, and style." (Sinan Ozdemir, "Quick Start Guide to Large Language Models: Strategies and Best Practices for Using ChatGPT and Other LLMs", 2024) 

"Despite their impressive capabilities, LLMs are not without limitations. One of the most significant challenges is the problem of hallucination, where an LLM generates factually incorrect or misleading information that appears plausible. This is particularly problematic in domains requiring high factual accuracy, such as healthcare, finance, and legal applications. To mitigate hallucinations and enhance the reliability of LLM outputs,  Retrieval-Augmented Generation (RAG) has emerged as a powerful technique. RAG works by dynamically retrieving relevant information from an external knowledge source (such as a knowledge graph) at inference time, rather than just relying on pre-trained knowledge. This approach ensures that the model has access to up-to-date and accurate data, grounding answers in verified information rather than generating content purely from its internal representations." (Aldo Marzullo et al, "Graph Machine Learning" 2nd Ed., 2025)

"Despite their impressive capabilities, LLMs are not without limitations. One of the most significant challenges is the problem of hallucination, where an LLM generates factually incorrect or misleading information that appears plausible. This is particularly problematic in domains requiring high factual accuracy, such as healthcare, finance, and legal applications. To mitigate hallucinations and enhance the reliability of LLM outputs,  Retrieval-Augmented Generation (RAG) has emerged as a powerful technique. RAG works by dynamically retrieving relevant information from an external knowledge source (such as a knowledge graph) at inference time, rather than just relying on pre-trained knowledge. This approach ensures that the model has access to up-to-date and accurate data, grounding answers in verified information rather than generating content purely from its internal representations." (Aldo Marzullo et al, "Graph Machine Learning" 2nd Ed., 2025)

"Generative AI tools for coding are sometimes inaccurate. They can produce results that look good but are wrong. This is common with LLMs. They can write code or chat like a person. And sometimes, they share information that’s just plain wrong. Not just a bit off, but totally backwards or nonsense. And they say it so confidently! We call this 'hallucinating', which is a funny term, but it makes sense." (Jeremy C Morgan, "Coding with AI: Examples in Python", 2025)

"In prompt engineering, we customize the prompts or questions we give the model to get more accurate or insightful responses. The way a prompt is structured has a massive impact on how well a model understands the task at hand and, ultimately, how well it performs. Given LLMs’ versatility, prompt engineering has become an important skill for getting the most out of these models across different domains and tasks. The key is to understand how different prompt structures lead to different model behaviors. There are various strategies - ranging from simple one-shot prompting to more complex techniques like chain-of-thought prompting - that can significantly improve the effectiveness of LLMs." (Abi Aryan, "LLMOps: Managing Large Language Models in Production", 2025)

"RAG is a framework that combines the strengths of traditional information retrieval systems with the generative capabilities of LLMs. In this setup, an LLM is augmented with a retrieval component that fetches relevant information from external data sources, such as knowledge bases or databases, to produce more accurate and contextually relevant responses. This method enhances the LLM’s output by grounding it in authoritative, up-to-date information." (Aldo Marzullo et al, "Graph Machine Learning" 2nd Ed., 2025)

"A 'hallucination' in the context of generative AI refers to the phenomenon where a model produces information that is factually incorrect, nonsensical, or not grounded in its input data or pre-existing knowledge. These are not mere typos or minor inaccuracies; they are confident, coherent, and often persuasive fabrications. In high-stakes domains like healthcare, law, or finance, a single hallucination can have severe consequences, eroding user trust and leading to catastrophic decision-making. While all LLMs are prone to this, the RAG architecture is specifically designed to combat it by tethering the model’s output to an external, verifiable knowledge base. Understanding why hallucinations occur is the essential first step to building more reliable and truthful AI systems." (Saloni Garg et al, "RAG Artificial Intelligence: Retrieval-Augmented Generation in Generative AI", 2026) 

"Another big problem is model hallucinations, which happen when generative models make content that seems real but is actually wrong or made up. For instance, a language model could write a news story or a medical diagnosis that has wrong information. This happens because these models value coherence and fluency more than factual accuracy. When the training data is not enough or is not clear, they often 'fill in the gaps' with made-up information." (Saloni Garg et al, "RAG Artificial Intelligence: Retrieval-Augmented Generation in Generative AI", 2026)

"Beyond intentionally misleading content, GenAI systems can produce inaccurate information unintentionally. LLMs are prone to hallucination, generating plausible but false statements with the same confidence as accurate ones. In enterprise contexts, this poses particular risks: an AI assistant might report incorrect financial figures, fabricate customer details, or misrepresent historical trends. Organizations deploying GenAI must implement validation mechanisms, human oversight, and retrieval-augmented approaches that ground model outputs in verified data sources." (Bennie Haelen, "ML and Generative AI in the Data Lakehouse Building and Deploying AI Applications at Scale", 2026)

"Ensuring that a language model reliably retrieves and presents correct information, often referred to as factual recall, is critical for any production-grade application. Whether you’re building an internal helpdesk assistant, a medical Q&A system, or an automated compliance auditor, users expect concise, accurate answers that align with up-to-date source material. Unfortunately, without explicit context, even the most powerful LLM can hallucinate or omit key facts." (Bennie Haelen, "ML and Generative AI in the Data Lakehouse Building and Deploying AI Applications at Scale", 2026)

"Generative artificial intelligence (GenAI), powered by large language models (LLMs) like Google’s Gemini and OpenAI’s GPT, has transformed how we work and live, revolutionizing business after business. Despite this success, generative AI falls short in domains where specific domain knowledge, high accuracy, and explainability are essential. And it has other significant limitations, including hallucinations and a lack of context and relations. This is where knowledge graphs (KGs) come in, provid-ing contextual information - such as experiences, environmental characteristics, cultural aspects, and social normsneeded to build the 'third wave of AI' for mission-critical applications." (Alessandro Negro et al, "Knowledge Graphs and LLMs in Action", 2026)

"RAG applications must be built with semantics, metadata, and governance in mind. The retrieved information must be high-quality, secure, and appropriate for the user’s role. Equally important is monitoring and management: checking whether source data has changed, ensuring vector stores remain accurate, and watching for hallucinations or data leakage. Organizations are definitely starting to experiment with RAG models today; some are putting them into production applications. Some believe that using RAG helps mitigate hallucinations because it is grounded in trusted organizational data." (Fern Halper, "Data Makes the World Go 'Round", 2026)

"[...] RAG models excel in dynamic environments where information changes frequently, such as news generation or customer support. Standard generative models, constrained by their training data, may provide outdated or irrelevant responses. RAG, however, can pull the latest information, ensuring up-to-date and contextually appropriate outputs. While RAG models may require more computational resources due to the retrieval step, the trade-off is often justified by the substantial improvements in accuracy and reliability, making them a superior choice for many real-world applications." (Saloni Garg et al, "RAG Artificial Intelligence: Retrieval-Augmented Generation in Generative AI", 2026)

"Retrieval mechanisms are essential for addressing some of the key limitations of traditional generative AI models, such as factual inaccuracies, lack of context awareness, and model hallucinations. While generative models excel at creating coherent and fluent content, they often struggle to produce outputs that are factually correct or contextually relevant. This is because these models rely solely on patterns learned during training, without access to real-time or external information. For example, a generative model might generate a plausible-sounding but incorrect answer to a factual question, as it cannot verify the accuracy of its response." (Saloni Garg et al, "RAG Artificial Intelligence: Retrieval-Augmented Generation in Generative AI", 2026)

"The foundational premise of RAG is that the most effective way to reduce hallucinations is to provide the LLM with the correct, explicit information needed to answer a query, thereby minimizing its need to rely on fallible parametric knowledge. Therefore, the quality, relevance, and accuracy of the retrieval step are the most significant factors in determining the truthfulness of the final output. Better retrieval is the most powerful antidote to hallucination. If the retriever fails to find the correct information, the generator is essentially left to guess, making hallucinations almost inevitable. The goal is to create a tight, unambiguous link between the user’s question and the evidence in the knowledge base." (Saloni Garg et al, "RAG Artificial Intelligence: Retrieval-Augmented Generation in Generative AI", 2026)

06 October 2026

🤖Prompt Engineering: Knowledge Bases (Just the Quotes)

"How can a cognitive system process environmental input and stored knowledge so as to benefit from experience? More specific versions of this question include the following: How can a system organize its experience so that it has some basis for action even in unfamiliar situations? How can a system determine that rules in its knowledge base are inadequate? How can it generate plausible new rules to replace the inadequate ones? How can it refine rules that are useful but non-optimal? How can it use metaphor and analogy to transfer information and procedures from one domain to another?" (John H Holland et al, "Induction: Processes Of Inference, Learning, And Discovery", 1986)

"Inference is the process of matching current facts from the domain space to the existing knowledge and inferring new facts. An inference process is a chain of matchings. The intermediate results obtained during the inference process are matched against the existing knowledge. The length of the chain is different. It depends on the knowledge base and on the inference method applied." (Nikola K Kasabov, "Foundations of Neural Networks, Fuzzy Systems, and Knowledge Engineering", 1996)

"Representation is the process of transforming existing problem knowledge to some of the known knowledge-engineering schemes in order to process it by applying knowledge-engineering methods. The result of the representation process is the problem knowledge base in a computer format." (Nikola K Kasabov, "Foundations of Neural Networks, Fuzzy Systems, and Knowledge Engineering", 1996)

"LLMs are trained on large volumes of data, which inherently provides them with an immense knowledge base and understanding of different languages. Yet, LLMs at their core are complex text completion engines. Since this knowledge and understanding of language is compressed in a very high-dimensional latent space. LLMs end up using these in a very fluid and intelligible way (which often leads to hallucinations). In order to guide LLMs to focus on specific topics or pieces of information to solve certain tasks, (for instance, question-answering from a given piece of text), it is important to provide contextual information explicitly. While most current generations of LLMs have extremely wide context windows, it is recommended to preprocess context into overlapping smaller chunks for better results, reduced latency, and so on. For similar reasons, it is also recommended to preprocess contextual information in clear and task-specific formats. This aspect of context preprocessing is extremely useful in Retrieval-Gugmented Generation (RAG) scenarios." (Joseph Babcock & Raghav Bali, "Generative AI with Python and PyTorch" 2nd. Ed., 2025)

"LLMs excel at understanding context and making associations among words, phrases, and concepts to provide relevant information based on the input query or prompt. While structured knowledge bases rely on humancurated data, LLMs can  automatically extract knowledge from unstructured text. When trained on diverse textual sources, they can process a vast amount of information without explicit human intervention. However, this also introduces a challenge, as the model can learn biased or incorrect information from the training data." (Abi Aryan, "LLMOps: Managing Large Language Models in Production", 2025)

"RAG is a framework that combines the strengths of traditional information retrieval systems with the generative capabilities of LLMs. In this setup, an LLM is augmented with a retrieval component that fetches relevant information from external data sources, such as knowledge bases or databases, to produce more accurate and contextually relevant responses. This method enhances the LLM’s output by grounding it in authoritative, up-to-date information." (Aldo Marzullo et al, "Graph Machine Learning" 2nd Ed., 2025)

"Semantic Kernel is a framework designed to simplify integrating LLMs into applications that require dynamic knowledge, reasoning, and state tracking. It’s particularly useful when you want to build complex, modular AI systems that can interact with external APIs, knowledge bases, or decision-making processes. Semantic Kernel focuses on building more flexible AI systems that can handle a variety of tasks beyond just generating text. It allows for modularity, enabling developers to easily combine different components - such as embeddings, prompt templates, and custom functions - in a cohesive manner." (Abi Aryan, "LLMOps: Managing Large Language Models in Production", 2025)

"Intelligent systems connect users to AI and ML to achieve meaningful objectives. An intelligent system is one in which intelligence evolves and improves over time, particularly when it improves by watching how users interact with the system.[...] The primary objective of the intelligent system is to support users in accomplishing complex tasks - not by replacing them, but by enhancing their decision-making capabilities. [...] An intelligent system must also have the ability to learn from user interactions and explicit feedback, as well as utilize contextual information. The system should contin-uously develop, use, and maintain an evolving knowledge base. This evolution is driven not only by data sources but also by ongoing interactions with users." (Alessandro Negro et al, "Knowledge Graphs and LLMs in Action", 2026)

"Misinformation is false or misleading information, and its generation by AI is particularly dangerous because of the aura of credibility these systems can project. In a RAG system, misinformation primarily arises from two failure points: Retrieval of inaccurate content from the knowledge base and fabrication or distortion by the large language model (LLM) during generation, even when given good context. The first line of defense is ensuring the integrity of the knowledge base. A RAG system is only as reliable as the documents it has access to. If non-credible, manipulated, or satirical sources are ingested, the system will retrieve and use them as fact. This makes rigorous data curation and source validation the most critical step in combating misinformation. The second line of defense is strengthening the connection between retrieval and generation to prevent the LLM from 'going off script'. The LLM, based on its pre-trained knowledge, might confidently generate an answer that contradicts the provided evidence or adds unsupported details – a phenomenon known as 'hallucination'. To mitigate this, thesystem must be designed to strictly adhere to the retrieved context." (Saloni Garg et al, "RAG Artificial Intelligence: Retrieval-Augmented Generation in Generative AI", 2026)

"The foundational form of RAG, often called naive RAG, follows a straightforward pattern. A pipeline retrieves supporting context from external sources such as enterprise documents, knowledge bases, or structured datasets and appends that information to the model’s prompt before inference. In the most common implementation, each document is converted into an embedding, a numerical representation of its semantic meaning, using either the same foundation model or a specialized embedding model. When a user submits a query, the system performs a vector similarity search to find documents whose embeddings most closely match the query’s vector representation, and the retrieved content is concatenated with the user query before being passed to the language model." (Bennie Haelen, "ML and Generative AI in the Data Lakehouse Building and Deploying AI Applications at Scale", 2026)

🤖Prompt Engineering: Failure (Just the Quotes)

"Agentic workflows break when the logic is messy - if, say, the plans don’t decompose or memory is poorly structured. However, infrastructure-level LLM applications introduce even more failure points and complexity. If the protocols don’t sync with each other, or the data flows start leaking, or the model boundaries are unclear... there are far too many failure points to count. While most people have been jumping on the bandwagon to adopt MCPs or A2A, very few are equipped to handle the LLMOps issues these tools introduce." (Abi Aryan, "LLMOps: Managing Large Language Models in Production", 2025)

"Data drift manifests in several distinct ways. Input drift typically shows up as an increase in adversarial or malformed queries that deviate from the original training or design expectations. This can stress the system’s robustness and degrade output quality. Retriever drift occurs when the relevance of the documents returned by retrieval components declines, even if the retrieval algorithms and configurations remain unchanged. Similarly, embedding drift arises when the vector representations used to compare semantic similarity become less effective, causing retrieval systems to fail despite stable system parameters." (Abi Aryan, "LLMOps: Managing Large Language Models in Production", 2025)

"LLM deployment failures often trace back not to the model itself, but to the prompts it receives. In production environments, prompts are rarely fixed, handcrafted snippets. Instead, they are dynamically generated, assembled from templates, and parameterized based on upstream data sources or evolving user state. This dynamism introduces complexity and variability that can subtly undermine the system’s performance if not carefully managed." (Abi Aryan, "LLMOps: Managing Large Language Models in Production", 2025)

"The simplest form of an agent is little more than a wrapped prompt. It takes an input, does some local reasoning, returns an output, and exits. There’s no memory, no iteration, no feedback loop. These are useful when the task is bounded, like generating a SQL query, converting a paragraph to a tweet, or answering a direct question. But single-step agents are brittle. They assume everything is known up front. They can’t handle surprises or partial failures. You’ll quickly outgrow them when tasks involve multiple actions or require state tracking." (Abi Aryan, "LLMOps: Managing Large Language Models in Production", 2025)

"If ethical lapses or AI failures occur, the impact on a business can be significant. Misinformation, biases, or harmful content generated by AI can lead to reputational damage, customer distrust, and potential regulatory scrutiny. The public relations fallout from an AI-driven error or ethical misstep can erode consumer confidence, resulting in lost revenue and lasting harm to brand image. Businesses, therefore, need to proactively address ethical considerations in AI implementation, not only to ensure compliance but also to protect and strengthen their reputation in a highly competitive, and increasingly transparent, marketplace." (Bennie Haelen, "ML and Generative AI in the Data Lakehouse Building and Deploying AI Applications at Scale", 2026)

"[...] KGs and LLMs can be the foundation for different types of reasoning, complementing each other in intelligent systems. We can use KGs for tasks that require precise, rule-based reasoning and explicit knowledge representation, and LLMs for tasks involving pattern recognition, context understanding, handling ambiguity or incomplete information, and reasoning about graph structures and their derived metrics. However, neither approach inherently possesses common-sense reasoning capabilities comparable to those of humans, and they often fail to make intuitive leaps or understand the implicit context that would be obvious to a person. These limitations underscore the importance of carefully considering the strengths and weaknesses of each approach when designing intelligent systems and potentially developing a powerful hybrid IAS." (Alessandro Negro et al, "Knowledge Graphs and LLMs in Action", 2026)

"Misinformation is false or misleading information, and its generation by AI is particularly dangerous because of the aura of credibility these systems can project. In a RAG system, misinformation primarily arises from two failure points: Retrieval of inaccurate content from the knowledge base and fabrication or distortion by the large language model (LLM) during generation, even when given good context. The first line of defense is ensuring the integrity of the knowledge base. A RAG system is only as reliable as the documents it has access to. If non-credible, manipulated, or satirical sources are ingested, the system will retrieve and use them as fact. This makes rigorous data curation and source validation the most critical step in combating misinformation. The second line of defense is strengthening the connection between retrieval and generation to prevent the LLM from 'going off script'. The LLM, based on its pre-trained knowledge, might confidently generate an answer that contradicts the provided evidence or adds unsupported details – a phenomenon known as 'hallucination'. To mitigate this, thesystem must be designed to strictly adhere to the retrieved context." (Saloni Garg et al, "RAG Artificial Intelligence: Retrieval-Augmented Generation in Generative AI", 2026)

"The promise of AI is its ability to process information objectively and at scale. However, this promise is fundamentally threatened by the twin challenges of bias and misinformation. AI systems are not born in a vacuum; they are created by humans and trained on data produced by humans. Consequently, they are prone to inheriting and even amplifying our prejudices, errors, and the systemic inequalities present in that data. In a RAG system, this risk is a two-fold problem: first in the retrieval of information, and second in the generation of a response based on that retrieval. A failure to address these issues doesn’t just lead to technically incorrect outputs; it can perpetuate social harm, erode public trust, and lead to the widespread dissemination of falsehoods. Therefore, understanding and mitigating bias and misinformation is not an optional add-on but a core requirement for any ethically deployed AI system." (Saloni Garg et al, "RAG Artificial Intelligence: Retrieval-Augmented Generation in Generative AI", 2026)

05 October 2026

🤖Prompt Engineering: Learning (Just the Quotes)

"There is a plethora of credible scenarios for achieving human-level intelligence in a machine. We will be able to evolve and train a system combining massively parallel neural nets with other paradigms to understand language and model knowledge, including the ability to read and understand written documents. Although the ability of today's computers to extract and learn knowledge from natural-language documents is quite limited, their abilities in this domain are improving rapidly. Computers will be able to read on their own, understanding and modeling what they have read, by the second decade of the twenty-first century. We can then have our computers read all of the world's literature books, magazines, scientific journals, and other available material. Ultimately, the machines will gather knowledge on their own by venturing into the physical world, drawing from the full spectrum of media and information services, and sharing knowledge with each other (which machines can do far more easily than their human creators)." (Ray Kurzweil, "The Age of Spiritual Machines: When Computers Exceed Human Intelligence", 1999)

"The no free lunch theorem for machine learning states that, averaged over all possible data generating distributions, every classification algorithm has the same error rate when classifying previously unobserved points. In other words, in some sense, no machine learning algorithm is universally any better than any other. The most sophisticated algorithm we can conceive of has the same average performance (over all possible tasks) as merely predicting that every point belongs to the same class. [...] the goal of machine learning research is not to seek a universal learning algorithm or the absolute best learning algorithm. Instead, our goal is to understand what kinds of distributions are relevant to the 'real world' that an AI agent experiences, and what kinds of machine learning algorithms perform well on data drawn from the kinds of data generating distributions we care about." (Ian Goodfellow et al, "Deep Learning", 2015)

"Self-attention, sometimes called intra-attention is an attention mechanism relating different positions of a single sequence in order to compute a representation of the sequence. Self-attention has been used successfully in a variety of tasks including reading comprehension, abstractive summarization, textual entailment and learning task-independent sentence representations.  End-to-end memory networks are based on a recurrent attention mechanism instead of sequence-aligned recurrence and have been shown to perform well on simple-language question answering and language modeling tasks. To the best of our knowledge, however, the Transformer is the first transduction model relying entirely on self-attention to compute representations of its input and output without using sequence-aligned RNNs or convolution." (Ashish Vaswani et al, "Attention Is All You Need", 2017)

"[...] building an effective LLM-based application can require more than just plugging in a pre-trained model and retrieving results - what if we want to parse them for a better user experience? We might also want to lean on the learnings of massively large language models to help complete the loop and create a useful end-to-end LLM-based application. This is where prompt engineering comes into the picture." (Sinan Ozdemir, "Quick Start Guide to Large Language Models: Strategies and Best Practices for Using ChatGPT and Other LLMs", 2024) 

"Language modeling is a subfield of NLP that involves the creation of statistical/deep learning models for predicting the likelihood of a sequence of tokens in a specified vocabulary (a limited and known set of tokens). There are generally two kinds of language modeling tasks out there: autoencoding tasks and autoregressive tasks." (Sinan Ozdemir, "Quick Start Guide to Large Language Models: Strategies and Best Practices for Using ChatGPT and Other LLMs", 2024)

"The idea behind transfer learning is that the pre-trained model has already learned a lot of information about the language and relationships between words, and this information can be used as a starting point to improve performance on a new task. Transfer learning allows LLMs to be fine-tuned for specific tasks with much smaller amounts of task-specific data than would be required if the model were trained from scratch. This greatly reduces the amount of time and resources needed to train LLMs." (Sinan Ozdemir, "Quick Start Guide to Large Language Models: Strategies and Best Practices for Using ChatGPT and Other LLMs", 2024)

"Transfer learning is a technique used in machine learning to leverage the knowledge gained from one task to improve performance on another related task. Transfer learning for LLMs involves taking an LLM that has been pre-trained on one corpus of text data and then fine-tuning it for a specific 'downstream' task, such as text classification or text generation, by updating themodel’s parameters with task-specific data." (Sinan Ozdemir, "Quick Start Guide to Large Language Models: Strategies and Best Practices for Using ChatGPT and Other LLMs", 2024)

"As with many other deep learning-based approaches, another major challenge is in interpretability. While knowledge graphs provide a structured and transparent way to store relationships, LLMs operate as a black box, making it difficult to understand how specific outputs are generated. [...] Data alignment is also a key issue, as structured knowledge graphs and unstructured text data must be carefully preprocessed to ensure consistency.  Differences in data formats, ontology mismatches, and information redundancy can create inefficiencies when integrating these two paradigms. Developing robust pipelines that seamlessly connect graph-based insights with LLM-generated text remains an open challenge." (Aldo Marzullo et al, "Graph Machine Learning" 2nd Ed., 2025)

"Generative AI for coding and language tools is based on the LLM concept. A large language model is a type of neural network that processes and generates text in a humanlike way. It does this by being trained on a massive dataset of text, which allows it to learn human language patterns, as described previously. It lets LLMs translate, write, and answer questions with text. LLMs can contain natural language, source code, and  more." (Jeremy C Morgan, "Coding with AI: Examples in Python", 2025)

"LLMs excel at understanding context and making associations among words, phrases, and concepts to provide relevant information based on the input query or prompt. While structured knowledge bases rely on humancurated data, LLMs can  automatically extract knowledge from unstructured text. When trained on diverse textual sources, they can process a vast amount of information without explicit human intervention. However, this also introduces a challenge, as the model can learn biased or incorrect information from the training data." (Abi Aryan, "LLMOps: Managing Large Language Models in Production", 2025)

"Transformers are complex models built like LEGO blocks using multiple smart and specialized components. [...] Briefly, a vanilla transformer model consists of separate stacks of encoders and decoders. Each encoder block includes multi-head self-attention, enabling the model to capture relationships between tokens regardless of their positions. Residual connections help maintain gradient flow, preventing the vanishing gradient problem. Layer normalization ensures training stability, and feed-forward layers introduce non-linearity and learn complex token interactions. Decoder blocks contain the same components but also include an encoder-decoder attention mechanism to incorporate context from the encoder. The model uses embedding layers to convert tokens into a continuous latent space for contextual learning and positional encoding to preserve the order of tokens in the sequence." (Joseph Babcock & Raghav Bali, "Generative AI with Python and PyTorch" 2nd. Ed., 2025)

"With MCP, a model no longer has to guess what’s possible. Instead, it can discover tools, query data sources, and select prompts - all in real time, all through a shared protocol. This means a model doesn’t just generate responses; it acts, it calls tools, it gathers context, and it learns how to interact with the outside world in a modular, controlled way." (Abi Aryan, "LLMOps: Managing Large Language Models in Production", 2025)

"Generative AI has a lot of problems because it needs a lot of data to work. Generative models like GPT and GANs need a lot of training data to learn patterns and make good outputs. This data dependency can cause problems like overfitting, which happens when the model does well on training data but doesn’t work well with new, unseen data. Also, the quality of the content that is generated is directly related to the diversity and representativeness of the training data. This means that biased or incomplete datasets can lead to outputs that are wrong or unfair." (Saloni Garg et al, "RAG Artificial Intelligence: Retrieval-Augmented Generation in Generative AI", 2026)

"In RAG methods, the AI model itself doesn’t actually 'remember' or learn the proprietary data directly. Instead, the proprietary data are stored separately in what’s called a vector database. When the model is asked a question, it first performs a quick search of the proprietary database, finds relevant pieces of information, and then uses these to generate its response. The model’s core parameters remain entirely unchanged and are never updated with this private data. In this sense, the model hasn’t learned or 'seen' your proprietary data in its internal parameters, it only temporarily consults it as a reference to formulate an answer." (Alan Watkins & G C Cooke, "Smarter than You Winning in Business with Superintelligent AI", 2026

"Intelligent systems connect users to AI and ML to achieve meaningful objectives. An intelligent system is one in which intelligence evolves and improves over time, particularly when it improves by watching how users interact with the system.[...] The primary objective of the intelligent system is to support users in accomplishing complex tasks - not by replacing them, but by enhancing their decision-making capabilities. [...] An intelligent system must also have the ability to learn from user interactions and explicit feedback, as well as utilize contextual information. The system should contin-uously develop, use, and maintain an evolving knowledge base. This evolution is driven not only by data sources but also by ongoing interactions with users." (Alessandro Negro et al, "Knowledge Graphs and LLMs in Action", 2026)

"LangGraph handles the perception, reasoning, and action flow while maintaining memory and context across tasks. In this sense, it functions as both the development environment and the orchestration layer - coordinating the steps of perception, reasoning, and action while managing connections to external systems. Underneath this, emerging standards like MCP ensure that agents can connect securely and consistently to tools and data sources, making agentic architectures portable and scalable across platforms. Taken together, the orchestration layer and emerging interoperability standards like MCP form the foundation for scalable agentic AI. They make it possible for agents to perceive, reason, act, and learn in coordinated ways across complex environments, translating autonomous intelligence into practical, enterprise-grade capability." (Fern Halper, "Data Makes the World Go 'Round", 2026)

"RAG models offer several advantages over standard generative models, addressing many of their limitations. Traditional generative models, like GPT, rely solely on patterns learned during training, which can lead to issues such as factual inaccuracies, model hallucinations, and lack of contextual relevance. These models generate content based on pre-existing knowledge, often without the ability to verify or update the information, making them less reliable for tasks requiring high accuracy. In contrast, RAG models integrate retrieval mechanisms that allow them to access external knowledge sources in real time. This ensures that the generated content is grounded in verified data, significantly improving factual consistency and relevance. For example, while a standard generative model might produce a plausible but incorrect answer to a factual question, a RAG model can retrieve and incorporate accurate information from a trusted source, reducing errors." (Saloni Garg et al, "RAG Artificial Intelligence: Retrieval-Augmented Generation in Generative AI", 2026)

"Retrieval mechanisms are essential for addressing some of the key limitations of traditional generative AI models, such as factual inaccuracies, lack of context awareness, and model hallucinations. While generative models excel at creating coherent and fluent content, they often struggle to produce outputs that are factually correct or contextually relevant. This is because these models rely solely on patterns learned during training, without access to real-time or external information. For example, a generative model might generate a plausible-sounding but incorrect answer to a factual question, as it cannot verify the accuracy of its response." (Saloni Garg et al, "RAG Artificial Intelligence: Retrieval-Augmented Generation in Generative AI", 2026)

🤖Prompt Engineering: Errors (Just the Quotes)

"The no free lunch theorem for machine learning states that, averaged over all possible data generating distributions, every classification algorithm has the same error rate when classifying previously unobserved points. In other words, in some sense, no machine learning algorithm is universally any better than any other. The most sophisticated algorithm we can conceive of has the same average performance (over all possible tasks) as merely predicting that every point belongs to the same class. [...] the goal of machine learning research is not to seek a universal learning algorithm or the absolute best learning algorithm. Instead, our goal is to understand what kinds of distributions are relevant to the 'real world' that an AI agent experiences, and what kinds of machine learning algorithms perform well on data drawn from the kinds of data generating distributions we care about." (Ian Goodfellow et al, "Deep Learning", 2015)

"The art of mega-prompts spanning multiple written pages and looking like essays has become commonplace for complex tasks when building applications to get things `just right'. Unfortunately, they bring with them lots of issues: errors, portability, complexity, and more. The GenAI world didn’t plan for mega-prompts. They have simply evolved into what they’ve become today because practitioners kept wanting to do more and more complex things, and their only way to express those intents was with a prompt. But step back and look at some of these prompts [...] Lurking just below the surface are a bunch of classical computing concepts like data, programming instructions, control flows, memory, and stora - all the components typically associated with classical computing elements." (Rob Thomas et al, "AI Value Creators: Beyond the Generative AI User Mindset", 2025)

"The same difficulties that characterize training deep feedforward networks also apply to RNNs; gradients tend to die out over long distances using traditional activation functions (or explode if the gradients become greater than 1). However, unlike feedforward networks, RNNs aren’t trained with traditional backpropagation, but rather a variant known as Backpropagation through Time (BPTT): the network is unrolled, as before, and backpropagation is used, averaging over errors at each time point (since an 'output', the hidden state, occurs at each step). Also, in the case of RNNs, we run into the problem that the network has a very short memory; it only incorporates information from the most recent unit before the current one and has trouble maintaining long-range context. For applications such as translation, this is clearly a problem, as the interpretation of a word at the end of a sentence may depend on terms near the beginning, not just those directly preceding it." (Joseph Babcock & Raghav Bali, "Generative AI with Python and PyTorch" 2nd. Ed., 2025)

"If ethical lapses or AI failures occur, the impact on a business can be significant. Misinformation, biases, or harmful content generated by AI can lead to reputational damage, customer distrust, and potential regulatory scrutiny. The public relations fallout from an AI-driven error or ethical misstep can erode consumer confidence, resulting in lost revenue and lasting harm to brand image. Businesses, therefore, need to proactively address ethical considerations in AI implementation, not only to ensure compliance but also to protect and strengthen their reputation in a highly competitive, and increasingly transparent, marketplace." (Bennie Haelen, "ML and Generative AI in the Data Lakehouse Building and Deploying AI Applications at Scale", 2026)

"[...] LLMs raise serious concerns about ethics, bias and fairness, errors in reasoning, hallucinations, and misuse (e.g., misinformation and disinformation). These concerns are exacerbated by modern LLMs being both literal and figurative 'black boxes': Literal black boxes because many advanced AI systems are proprietary and the weights (trained parameters of the models) are not released to the public; and figurative black boxes because even the open-source AI models are so complicated that understanding them and developing safety guardrails has thus far proven extremely difficult." (Mike X Cohen,"50 ML Projects To Understand LLMs", 2026)

"[...] RAG models excel in dynamic environments where information changes frequently, such as news generation or customer support. Standard generative models, constrained by their training data, may provide outdated or irrelevant responses. RAG, however, can pull the latest information, ensuring up-to-date and contextually appropriate outputs. While RAG models may require more computational resources due to the retrieval step, the trade-off is often justified by the substantial improvements in accuracy and reliability, making them a superior choice for many real-world applications." (Saloni Garg et al, "RAG Artificial Intelligence: Retrieval-Augmented Generation in Generative AI", 2026)

"RAG models offer several advantages over standard generative models, addressing many of their limitations. Traditional generative models, like GPT, rely solely on patterns learned during training, which can lead to issues such as factual inaccuracies, model hallucinations, and lack of contextual relevance. These models generate content based on pre-existing knowledge, often without the ability to verify or update the information, making them less reliable for tasks requiring high accuracy. In contrast, RAG models integrate retrieval mechanisms that allow them to access external knowledge sources in real time. This ensures that the generated content is grounded in verified data, significantly improving factual consistency and relevance. For example, while a standard generative model might produce a plausible but incorrect answer to a factual question, a RAG model can retrieve and incorporate accurate information from a trusted source, reducing errors." (Saloni Garg et al, "RAG Artificial Intelligence: Retrieval-Augmented Generation in Generative AI", 2026)

"The promise of AI is its ability to process information objectively and at scale. However, this promise is fundamentally threatened by the twin challenges of bias and misinformation. AI systems are not born in a vacuum; they are created by humans and trained on data produced by humans. Consequently, they are prone to inheriting and even amplifying our prejudices, errors, and the systemic inequalities present in that data. In a RAG system, this risk is a two-fold problem: first in the retrieval of information, and second in the generation of a response based on that retrieval. A failure to address these issues doesn’t just lead to technically incorrect outputs; it can perpetuate social harm, erode public trust, and lead to the widespread dissemination of falsehoods. Therefore, understanding and mitigating bias and misinformation is not an optional add-on but a core requirement for any ethically deployed AI system." (Saloni Garg et al, "RAG Artificial Intelligence: Retrieval-Augmented Generation in Generative AI", 2026)

"With autonomous agents there is the risk that they can take action that isn’t governed. These systems introduce autonomy, interdependence between agents, and possibly emergent behavior. [...] Because agents act autonomously, a single misconfigured or misaligned agent can propagate errors at scale. In multiagent environments, one faulty output can trigger many downstream mistakes. Other risks include tool misuse, conflicting goals among agents (and other interoperability issues), and operational opacity, when teams cannot easily determine which agent took which action or why. Governance must extend to agent registration, version control, permissioning, and simulation testing before deployment. Likewise, accountability may also blur. If an agent takes a dangerous action, who is responsible? There are, of course, cybersecurity risks as agents pose a new attack surface."  (Fern Halper, "Data Makes the World Go 'Round", 2026)

25 September 2026

⛩️Douglas T Ross - Collected Quotes

"Automatic design has the computer do too much and the human do too little, whereas automatic programming has the human do too much and the computer do too little. Both techniques are important, but are not representative for what we wish to mean by computer-aided design." (Douglas T Ross, "Computer-Aided Design: A Statement of Objectives", 1960)

"Computer-aided design is not automatic design, although it must include many automatic design features. By automatic design we mean design procedures which are capable of being completely specified in a form which a computer can execute without human intervention." (Douglas T Ross, "Computer-Aided Design: A Statement of Objectives", 1960)

"It is very difficult to define what is meant by computer-aided design since the complete definition is, in fact, the sum and substance of the total project effort which has only begun. It is much easier to describe, what is not computer-aided design as we mean it." (Douglas T Ross, "Computer-Aided Design: A Statement of Objectives", 1960)

"The objective of the Computer-Aided Design Project is to evolve a machine systems which will permit the human designer and the computer to work together on creative design problems."  (Douglas T Ross, "Computer-Aided Design: A Statement of Objectives", 1960)

"Mechanical drawings and blueprints are not mere pictures, but a complete and rich language. In blueprint language, scientific, mathematical, and geometric formulations, notations, mensurations, and naming do not merely describe an object or process, they actually model it. Because of broad differences in subject, purpose, roles, and the needs of the people who use them, many forms of blueprint have evolved, but all rigorously present well structured information in understandable form." (Douglas T Ross, "Structured analysis (SA): A language for communicating ideas", IEEE Transactions on Software Engineering Vol. 3 No. 1, 1977)

"Structured analysis (SA) combines blueprint-like graphic language with the nouns and verbs of any other language to provide a hierarchic, top-down, gradual exposition of detail in the form of an SA model. The things and happenings of a subject are expressed in a data decomposition and an activity decomposition, both of which employ the same graphic building block, the SA box, to represent a part of a whole. SA arrows, representing input, output, control, and mechanism, express the relation of each part to the whole." (Douglas T Ross, "Structured analysis (SA): A language for communicating ideas", IEEE Transactions on Software Engineering Vol. 3 No. 1, 1977)

"The natural law of good communications takes the following, quite different, form in SA: Everything worth saying about anything worth saying something about must be expressed in six or fewer pieces." (Douglas T Ross, "Structured analysis (SA): A language for communicating ideas", IEEE Transactions on Software Engineering Vol. 3 No. 1, 1977)

"There are certain basic, known principles about how people's minds go about the business of understanding, and communicating understanding by means of language, which have been known and used for many centuries. No matter how these principles are addressed, they always end up with hierarchic decomposition as being the heart of good storytelling." (Douglas T Ross, "Structured analysis (SA): A language for communicating ideas", IEEE Transactions on Software Engineering Vol. 3 No. 1, 1977)

"We never have any understanding of any subject matter except in terms of our own mental constructs of ‘things’ and ‘happenings’ of that subject matter." (Douglas T Ross, "Structured analysis (SA): A language for communicating ideas", IEEE Transactions on Software Engineering Vol. 3 No. 1, 1977)

"A general theme for what I'm trying to convey and what actually drove me and my very industrious and creative project members over all these years, is… that there is much more to it than pictures. It has to be a picture language. There has to be meaning there, and the meaning is useful. You're trying to solve problems. So it really comes down to man machine problem solving. Better means of communication and expression is what always has driven our work." (Douglas T Ross, "Retrospectives: The Early Years in Computer Graphics at at MIT", Lincoln Lab and Harvard, 1989)

"There is a rigorous science, just waiting to be recognized and developed, which encompasses the whole of 'the software problem,' as defined, including the hardware, software, languages, devices, logic, data, knowledge, users, users, and effectiveness, etc. for end-users, providers, enablers, commissioners, and sponsors, alike." (Douglas T Ross,, 1989)

Related Posts Plugin for WordPress, Blogger...

About Me

My photo
Koeln, NRW, Germany
IT Professional with more than 25 years experience in IT in the area of full life-cycle of Web/Desktop/Database Applications Development, Software Engineering, Consultancy, Data Management, Data Quality, Data Migrations, Reporting, ERP implementations & support, Team/Project/IT Management, etc.