A healthy AI stack can still serve confidently wrong answers
An AI chatbot can spend weeks in tuning, answer accurately enough for stakeholders to approve it, and then become confidently wrong three months after launch even though nobody changed its model or prompts. In the illustrative scenario, roughly a third of questions eventually receive wrong answers because the world changed: pricing moved, a policy was updated, or a product specification received a new version while the knowledge store stayed unchanged. Those figures describe the scenario rather than measured industry prevalence.
The same failure can survive different retrieval architectures because retrieval and correctness answer different questions. Whether information reaches the application through a vector store, document index, or API call, standard retrieval can determine that information is relevant or available without determining whether it remains correct. An obsolete pricing document can still rank highly, while a record whose field silently disappeared can continue downstream. The model then receives apparently authoritative information and answers accordingly while operational dashboards remain green.
This is one of the most common production failure modes in enterprise AI right now. The mechanism behind that assessment is clear: a production system can execute exactly as designed while its information deteriorates. When monitoring treats successful execution as evidence of health, teams can spend their troubleshooting effort on parts of the AI stack that never caused the underlying change.
The usual troubleshooting path can target the wrong layer
That misleading symptom naturally directs attention toward the LLM first. A team sees newly inaccurate answers, tries another model, and modifies prompts because those are the components visibly generating the response. When those changes fail to restore accuracy, attention shifts to retrieval and context infrastructure. The team may then investigate or replace the system that decides what information enters the model’s context.
That second investigation addresses a genuine class of problems, and vendors have a commercial reason to sell technology for it. AWS has entered the “context layer” race with a knowledge graph that learns from agent usage, while Snowflake’s Horizon Context and Cortex Sense address agents producing confident wrong answers when the underlying business logic lacks governance. AWS and Snowflake benefit when enterprises invest in these context and AI services, so their products should be read in that commercial setting. Their presence also shows that engineering teams have tools aimed specifically at controlling the information supplied to agents.
Those context controls still depend on the quality of their inputs. A knowledge graph can organize relationships and a context system can improve what an agent receives, but stale or incorrect inputs retain those defects. Better retrieval can improve selection without proving that the selected fact is still true. Diagnosing a confident wrong answer therefore requires moving further upstream before concluding that the model or retrieval architecture needs to change.
A project in mind?
Schedule a 30-minute meeting with us.
Senior experts helping you move faster across product, engineering, cloud & AI.
Pipeline completion does not establish data correctness
Moving upstream changes what pipeline health needs to mean. A successful job establishes that an expected computation or transfer ran, while data quality requires testing the values the job transported against the requirements of their consumers. Relevance and availability are similarly limited signals because a system can retrieve available information that is highly relevant to a question while supplying a wrong answer from deteriorated information.
A fintech pipeline shows the same mechanism outside AI. An upstream system changed a field without notifying downstream users, yet the downstream pipeline continued to complete normally. Its monitoring checked whether the job finished, so bad values propagated into dashboards without triggering an operational failure. A customer eventually detected an inconsistency, by which point the values had already traveled downstream.
The absence of an error is not the presence of correctness. This principle covers both the fintech field change and the AI examples because a document can become stale without becoming inaccessible, and a field can disappear without stopping a record from moving. Unless validation tests the information itself, execution monitoring has no reason to classify either event as a failure.
Redefine pipeline health around four properties of trustworthy data
Once successful execution stops being the whole definition of health, data observability has to cover what happens to information throughout its path. The practical model has four dimensions: correctness, freshness, consistency, and lineage. Each dimension answers a different failure question and needs its own measurement. Together they make monitoring sensitive to semantic failures that ordinary job-status checks can leave invisible.
Correctness begins at the record level: does incoming data satisfy the structures and rules its consumers require? Validation should check field types, unexpected null values, permitted ranges, and other required conditions as data arrives. Great Expectations and Soda are commercial examples of tools that automate validation at row and column level, so both companies benefit when teams invest in automated data-quality controls. One operating measure is the percentage of records that pass validation on each pipeline run.
Freshness asks a different question because perfectly structured information can still be obsolete. A recent pipeline check says little when the underlying source has not successfully updated within the period its consumers require. Teams can track the elapsed time since each source’s most recent successful update and establish a service-level agreement, or SLA, for each dataset. Dataset-specific SLAs matter because one source may require hourly refreshes while another can remain useful on a slower schedule.
Consistency follows freshness by asking whether copies of a fact still agree after the data reaches multiple destinations or indexes. Two systems fed by the same source can drift apart without either independently looking broken, leaving the discrepancy undiscovered until a user compares them. Periodic cross-system checks make that disagreement measurable. A team can calculate the mismatch rate between downstream destinations and alert when the rate exceeds its defined threshold.
Lineage addresses the investigation that follows a detected problem. Lineage records where an output originated and the transformations it passed through before reaching a consumer, allowing engineers to query that history directly. When an AI answer is wrong, the team can use lineage to locate the source of its supporting information and determine how that information acquired its current form. The investigation then depends less on employees reconstructing the pipeline from memory.
Because lineage has to support that investigation, useful coverage depends on whether engineers can query it for critical datasets. A dataset may be well understood by the engineer who built its pipeline while remaining operationally opaque to everyone else. Teams can therefore measure the fraction of critical datasets whose lineage is queryable. That measure turns provenance into a system capability that remains available when a particular engineer is absent.
These controls can build on infrastructure already present in a data organization. The engineering shift is what those systems test and enforce: records against correctness rules, update times against source-specific SLAs, destinations against each other, and outputs against recorded provenance. Pipeline monitoring then describes the state of the data alongside the execution that moved it.
Uber and Netflix built key pieces of this answer before the RAG era
This broader definition of health predates retrieval-augmented generation, or RAG, in which an application retrieves external information and supplies it to a generative model for answering. Uber built its Unified Data Quality platform before RAG existed, using a dedicated system for data quality and observability. The platform supports more than 2,000 critical datasets and detects around 90% of data-quality incidents before downstream consumers receive them. Those results show what upstream quality controls can do at substantial scale.
Uber’s experience matters because detection happens before consumption. Once incorrect information has reached a dashboard, model, or customer-facing application, the organization has to respond after the problem has propagated. Automated quality controls instead identify a problem while there is still an opportunity to stop affected data from reaching consumers. AI raises the consequences because a generated answer can present supplied information with high confidence.
Where Uber illustrates early quality detection, Netflix developed another part of the reliability model through company-wide data lineage. Its system allows users to determine where a dataset originated and which systems or processes touched it, covering dependencies across Kafka topics, ML models, experimentation, and warehouse tables. The system was designed for human users and has become more important as AI and LLM applications have proliferated. Its practical role is to make provenance answerable across a complex data estate.
Together, the Uber and Netflix examples show that quality validation and queryable provenance were engineering needs before enterprises started placing generative models over their data. RAG gives these established controls another important consumer because AI applications rely on the same underlying information. When that information deteriorates, the model can turn an old data-quality problem into a fluent customer-facing answer.
Socure shows what correctness-first data flow looks like in practice
That need for controls before consumption appears concretely at Socure, where client data could arrive in whatever format each client chose and could occasionally be silently wrong. Socure needed to identify bad incoming information before it propagated into downstream systems while preserving enough provenance to understand where it originated. Great Expectations became part of the foundation for that control; as a data-quality tool provider, Great Expectations has a commercial stake in wider adoption of automated validation. Validation moved closer to ingestion, where a failure could still be contained.
Socure applied each of the four dimensions at a specific point in the data flow. Schema and range checks tested correctness when information entered the system, while per-source SLAs established freshness expectations that reflected each input’s characteristics. Cross-system checks tested consistency after data appeared in multiple places. File-level lineage preserved provenance so affected information could later be traced to what arrived.
Those controls operated behind a write-audit-publish pattern, which makes validation a gate between arrival and downstream use. Incoming data first landed in staging, where required validation ran before anything was released further downstream. Only data that passed the required checks moved from that controlled state into systems that depended on it. The decision to publish therefore depended on evidence about the data as well as successful ingestion.
Because reporting, ML models, and AI retrieval consumed the controlled data, the improvement appeared across all three classes of downstream consumer. Reporting became more accurate, as did ML models and AI retrieval operating over the same data. These applications can look like separate technology problems when inaccurate results surface, but their results can share the same defective input. Validation before publication addresses that common failure point.
Test the data layer before changing the model or context infrastructure
Socure’s implementation provides a practical diagnostic order when retrieval-based AI starts producing confident mistakes. Teams can first establish whether the information reaching the model and context components satisfies the requirements expected of it. The four dimensions translate directly into four questions:
- Is the underlying data validated against the standards required by its consumers?
- What is the oldest piece of content currently being served with high confidence?
- Would two chunks of the same source ever disagree with each other in the same retrieval result?
- Could you trace where it came from if it turned out to be wrong?
The first question tests correctness, while the second exposes freshness in terms that matter to an AI consumer instead of relying on when a pipeline last ran. The third tests consistency where retrieval can surface conflicting representations of the same source. The fourth tests lineage at the point an investigation needs it. A team that cannot answer these questions has a diagnostic gap between its source systems and the information available to its agent.
Finding that gap gives the team a reason to inspect data engineering before committing to a model swap or vendor migration. Model-specific failures can still occur, and a retrieval system can still select poor context; AWS and Snowflake are addressing those real context-layer problems. The diagnostic task is to establish which layer violated its requirements before changing another one. Otherwise a replacement component can inherit the same stale, incomplete, inconsistent, or untraceable information.
Recap
For business leaders, the central risk is not simply that an AI model can be wrong. It is that the entire AI stack can appear operationally healthy while the information behind its answers has already become unreliable. That makes data quality an AI governance and business reliability issue, not only a data engineering concern.
Before approving another model upgrade, retrieval redesign, or context-layer investment, executives should ask whether teams can demonstrate the correctness, freshness, consistency, and lineage of the data their AI systems consume. Those capabilities create evidence about where failures originate and help prevent expensive technology changes from treating the wrong problem.
The organizations best positioned to operate reliable AI will treat trustworthy data as part of the production system itself. Models and retrieval architectures will continue to change, but every generation of AI will remain dependent on the quality of the information it receives.
A project in mind?
Schedule a 30-minute meeting with us.
Senior experts helping you move faster across product, engineering, cloud & AI.


