A company can build a sophisticated data organization and still spend specialist time repairing defects that entered through a web form, CRM record, or migration. An invalid email address makes the problem concrete. Engineers can transform the record, analysts can discover its effect on a metric, and governance teams can define rules for handling it. All of this happens after the company has accepted the defective input.

Some quality problems arise later, during transformation or when datasets are combined. Others require context unavailable when a record first enters a system. The executive decision is about control placement: prevent predictable defects at capture, then detect and remediate problems that require downstream context. This separates where a defect originates from who eventually discovers it.

Organization design determines data quality responsibility

Centralized, decentralized, and federated team structures assign data responsibilities differently. Whatever structure a company chooses, it determines who defines standards, detects problems, and repairs records. Capture quality is a separate system-design decision that determines whether an email address can be validated when a customer submits it.

Specialized roles create the same distinction at a finer level. Engineers may build pipelines, analysts may find inconsistencies in business metrics, and governance teams may define acceptable values and ownership. If a malformed phone number passes through all three groups, each can respond according to its responsibility. Preventing that defect requires a control at or near the point where the number enters the organization.

This distinction helps executives avoid treating “data quality” as one broad responsibility. Data quality here means whether information is accurate, complete, consistent, and usable enough for its intended purpose. Different defects become detectable at different stages. Responsibility can follow the earliest point where a defect can be reliably identified and effectively corrected.

Upstream controls reduce predictable defects

A web form makes the control-placement problem concrete. It can check whether required fields are present and whether values follow rules the organization can evaluate at submission. Similar controls can operate when information enters a CRM or is admitted during a migration. Source validation means applying these checks as data enters a system, before downstream processing begins.

These checks have a clear boundary. A syntactically valid email address can belong to the wrong person, while a correctly formatted CRM field can contain an inaccurate business fact. Information can also become stale after collection. A capture rule can prevent only defects for which enough information exists at capture to make a reliable decision.

The same boundary applies to standardization. Address, email, phone-number, and identity fields can sometimes be normalized into consistent formats before entering downstream systems. That prevents known inconsistencies from becoming repeated cleaning tasks later. Normalization alone cannot establish that a value represents a true or current fact.

Migrations create a related control point because existing records are admitted into a new environment. Formatting inconsistencies, incomplete fields, and identity records may require checks or standardization before admission. Defects that pass through become inputs to later pipelines and applications. The earlier web-form example follows the same logic: accepting a detectable defect transfers work downstream.

Okoone experts
LET'S TALK!

A project in mind?
Schedule a 30-minute meeting with us.

Senior experts helping you move faster across product, engineering, cloud & AI.

Please enter a valid business email address.

Downstream controls handle context-dependent problems

Capture controls cannot evaluate every future use of a record. A problem may appear only after transformations occur or datasets created for different purposes are combined. Downstream controls are checks performed after capture, when more context is available. They include testing and monitoring for defects that emerge later or survive earlier controls.

Lineage can support this work by recording where data came from and how it moved through systems. When a downstream check finds a problem, that history can help identify affected inputs and transformations. Remediation can then address records that could not be judged reliably when they first entered the organization. The result is a layered control model based on when enough information exists to detect each failure.

This division also limits unnecessary complexity at capture. Entry systems should apply rules that can make reliable decisions with the information available there. Later controls can evaluate relationships and transformations visible only downstream. The design question is where each defect first becomes both detectable and actionable.

AI creates another downstream control point

AI systems add another downstream use for organizational data. When an automated workflow consumes a defective record, that defect can affect processing that depends on it. The quality question is the same as for a pipeline or analytical model: could the defect have been identified earlier, or does detection require context available only later? Control placement should follow the answer.

AI-based cleansing or anomaly detection, where an organization chooses to use it, fits within this layered model. Such tools can examine data after capture for patterns an entry rule was not designed to evaluate. A reliable capture rule remains the appropriate control for a malformed field it can identify. Repeatedly repairing the same preventable defect downstream leaves the mechanism that admits it unchanged.

Executives can allocate quality work by tracing recurring defects to the stage where they enter or arise, then finding the earliest stage where each can be judged reliably. Place the control there. A malformed field may be suitable for source validation, while a contradiction revealed only after datasets are combined requires a later check. The operating model then follows the failure points instead of forcing every quality problem onto a single team.

Main highlights

  • Align responsibility with control points: Data team structure determines who defines standards, detects defects, and repairs records, but capture quality remains a system-design decision. Assign responsibility based on where each defect can first be reliably identified and corrected.
  • Prevent predictable defects upstream: Use source validation and standardization for errors that can be reliably detected when data enters web forms, CRMs, or migration workflows. This reduces repeated downstream cleaning without treating capture controls as proof that data is accurate or current.
  • Keep downstream controls for context-dependent defects: Some quality problems emerge only after transformation or integration. Use downstream testing, monitoring, and lineage to identify and remediate defects that require context unavailable at capture.
  • Apply the same control model to AI: AI systems create another downstream point where defective data can affect processing. Use AI-based detection where later context adds value, but fix recurring, preventable defects at their source rather than repeatedly cleansing them downstream.

Alexander Procter

September 7, 2026

5 Min

Okoone experts
LET'S TALK!

A project in mind?
Schedule a 30-minute meeting with us.

Senior experts helping you move faster across product, engineering, cloud & AI.

Please enter a valid business email address.