Cheaper code shifts the engineering bottleneck

The first implementation of a distributed streaming pipeline is becoming easier to produce, which changes where engineering effort matters. Commit histories of modern data platforms over the last two years show a directional pattern: the friction of producing syntax has fallen substantially. Cursor, Claude Code and agentic workflows inside Docker containers and IDEs now let engineers start with natural-language intent and reach working code much sooner. The pattern is directional rather than a measured trend, but it changes the decisions engineering teams have to make.

For a complex API integration, an agent can navigate the repository, produce test coverage, inspect stack traces, suggest refactors and generate an initial implementation. An engineer can describe a Kafka-to-Iceberg sink mapping in plain English and receive a credible starting point before personally inspecting every relevant file. Because the agent can absorb more local implementation work, the engineer can focus more attention on the conditions the implementation must satisfy.

That shift creates a counterintuitive requirement for teams deciding how much autonomy to give agents. As agents become primary authors of more local logic, engineering value moves toward encoding business and system boundaries precisely enough for generated changes to converge on the right result. Useful autonomy can require stronger engineered constraints because producing plausible code is a different problem from determining whether that code is correct for the business and the surrounding system.

Agents converge when the problem is genuinely bounded

The need for stronger boundaries follows from how an agent turns intent into changes. An LLM has computational capacity, but useful work begins when a prompt, business requirement, system instruction or failing test supplies direction. The agent can turn that direction into code, tool calls, queries, tests and modifications to a running system. Its basic development loop is straightforward: propose a change, execute it, observe the result, correct the implementation and retry.

For a bounded task, that repeated feedback can work extremely well because each result reduces uncertainty. Give the agent a known input schema, a known target schema, a small codebase and tests that detect the relevant failures, and each attempt produces information that narrows the next one. The agent inspects the code, changes it, runs the tests and uses the results to revise its work. A visible definition of done then tells the agent when the task has succeeded.

The useful property of that loop is the combination of a narrow search space and informative feedback. Repeated attempts alone do not create convergence because each result must distinguish a better implementation from a worse one according to the requirements that matter. When inputs, outputs, rules and failure conditions are explicit, an agent can operate with considerable autonomy because the environment supplies enough information to correct mistakes. Enterprise systems become harder when they cannot provide those stable conditions.

Okoone experts
LET'S TALK!

A project in mind?
Schedule a 30-minute meeting with us.

Senior experts helping you move faster across product, engineering, cloud & AI.

Please enter a valid business email address.

Complex systems break the assumption that more iteration means more correctness

When stable conditions disappear, a difficult repository can degrade an agent’s feedback loop even as the agent continues producing plausible work. A clear task can begin with a reasonable assumption that proves stale, after which the agent repairs a symptom, interprets an obsolete migration as current behavior and carries those decisions into its working context. Further tool calls can return plausible but conflicting information. Later decisions can become less certain than the first because they depend on an expanding chain of unresolved assumptions.

That accumulation can be described as “operational entropy”: stale assumptions, branching context and unresolved dependencies build up while the loop keeps trying to advance. More generated code, queries and tool calls do not necessarily reduce that uncertainty because the agent may have no signal identifying which premise went wrong. Continued activity can then move the implementation further from the correct result. Iteration produces convergence only when the environment can push the agent toward the intended behavior.

Corrective signals provide that push by making an error observable. A human interruption, a failing test, a precise data contract, a deterministic tool or an evaluation can identify what the agent got wrong, allowing it to eliminate an incorrect path. Their effectiveness depends on whether the actual requirement has been encoded somewhere the agent can observe. A test suite cannot guide an agent toward a business rule that neither the test nor another machine-readable boundary expresses.

That limitation matters in enterprise environments because their state and rules keep changing. A real-time pricing engine can depend simultaneously on mutable operational state, third-party APIs, late-arriving events, regional policy and business rules distributed between software and human knowledge. The correct output therefore depends on information beyond code visible in one repository. An agent making a local change may encounter several sources of truth whose current relationships must be understood before the change is safe.

Those changing relationships also appear across ordinary platform components. Clickstream meaning changes as product behavior changes, while operational databases mutate through customer activity. APIs impose rate limits and move between versions, schemas evolve, and security policies change. Legacy systems add another constraint because rules can remain embedded for years in exception handling without ever becoming explicit documentation.

Because these components interact, changing one system can alter the meaning or behavior of another. A request that looks local may therefore require decisions about several systems at once, forcing the agent to reason across assumptions created at different times, by different teams and for different purposes. Each implicit dependency expands the conditions the generated implementation must satisfy. The apparent size of a code change can consequently hide the real size of the correctness problem.

Data lakehouses make that correctness problem particularly clear because physical consistency does not establish semantic correctness, meaning correctness according to the business concept the data is supposed to represent. A pipeline can process its inputs, satisfy its schema expectations and pass its technical tests while calculating a result that finance does not recognize. From the execution environment’s perspective, nothing has necessarily failed. From the business perspective, the output can still be unusable.

The data-lakehouse case shows why extra iterations cannot solve every failure. A feedback loop can correct only errors that its available signals make observable, while complex systems often hold requirements outside those signals. Mutable state and coupled dependencies increase the chance that an agent follows internally plausible reasoning based on an obsolete premise. Semantic correctness therefore has to become part of the environment in which the agent works.

A green pipeline can still encode the wrong business answer

A hypothetical revenue-model change makes that semantic failure concrete. An agent receives a request to add a customer_tier field, searches the operational database and discovers a field named status. It maps status into the transformation and runs the existing type and nullability tests. Those tests pass, leaving clean code and a green pipeline.

The green result is still semantically wrong because account status and customer tier represent different business concepts. Suppose the platform has a semantic data contract stating that customer_tier must be derived from trailing twelve-month spend. The same contract identifies a business owner and explicitly prohibits populating the field from account status. When the agent submits its transformation, the contract rejects the change before the incorrect value reaches the dashboard.

The contract changes the development loop because the semantic error becomes visible, specific and recoverable. The agent now receives information that its chosen derivation violates an explicit requirement, so it can discard the status mapping and revise the transformation around trailing twelve-month spend. In this case, the engineer’s high-value contribution is the boundary that defines what a valid implementation means. Encoding that boundary gives the automated loop information it previously lacked.

Complete testing can provide the same guidance when every relevant condition has already been captured in tests. An agent can then repeatedly run the suite, repair failures and converge on an acceptable implementation. The customer_tier example shows the limit: type and nullability tests can fully cover their technical properties while carrying no information about the business meaning of the field. A semantic contract supplies that missing correctness criterion.

Once correctness criteria are explicit, testing and iteration can enforce them. A green result answers the questions the pipeline has been configured to ask, while an unstated rule about trailing twelve-month spend remains invisible to that pipeline. The same problem applies to an unrecorded business owner whose approval defines a field’s meaning. Making semantics explicit changes what the automated loop can recognize as failure.

Changing environments make current criteria as important as explicit criteria. An API can change version, a schema can evolve and a security or regional policy can impose a new condition on behavior that worked previously. Mutable operational state can alter the inputs on which a decision depends, while undocumented legacy rules can keep affecting execution through old exception paths. Unless those changes reach the agent as current correctness criteria, additional retries can keep optimizing against an obsolete definition of success.

The resulting boundary for AI adoption lies between bounded and unbounded work. The customer_tier request looks bounded at the level of a field addition, but its semantic definition reaches into a business metric and ownership rule. Once those conditions are explicit, the agent has a tractable implementation problem. While they remain implicit, clean code and successful execution provide weak evidence that the business requirement has been met.

Architecture becomes the mechanism for safer autonomy

Because explicit conditions turn apparently local work into genuinely bounded work, architecture takes on a different role in an agent-heavy development process. The proposed engineering mandate can be called “designing equilibrium”: engineers build conditions that let generated logic be checked and corrected when it is wrong. As agents become faster at implementation, those conditions determine how far a team can safely extend autonomous development. Architecture becomes part of the agent’s corrective environment.

Several familiar architecture practices can supply the boundaries that environment needs: strict semantic layers, immutable event logs, data contracts, idempotent APIs and deterministic state machines. A semantic layer makes a business concept explicit so an agent does not have to infer it from similarly named database fields. An immutable event log preserves a dependable history, while an idempotent API makes repeated requests produce a predictable effect. Deterministic state machines constrain allowed transitions, and contracts specify requirements that generated changes must satisfy.

Those boundaries reduce how many assumptions an agent must resolve at once, turning some coupled problems into bounded domains with defined inputs, explicit rules and dependable feedback. The agent can then write a transformation, execute tests, repair the failures they expose and ship the change without reconstructing the unwritten history behind every table and service. Its autonomy becomes useful because architecture has reduced ambiguity before implementation begins. The architecture is therefore doing correctness work inside the development loop.

As code generation gets cheaper, that correctness work becomes more consequential. A weakly specified platform asks the agent to infer semantics, historical exceptions and current policy while also generating code, while a well-bounded platform turns those decisions into constraints the agent can inspect. When business requirements themselves change quickly, maintaining those constraints becomes part of engineering work because stale boundaries would encode a different kind of error. Stronger constraints can support greater autonomy when they provide current, precise feedback.

That relationship matters especially in complex enterprise systems with mutable state, evolving APIs and schemas, policy changes, coupled dependencies and undocumented rules. In environments where agents increasingly produce local implementations, engineering value moves toward defining the conditions under which those generated changes can be trusted. Keeping those conditions explicit and current allows the agent’s own feedback loop to enforce them.

Key takeaways for decision-makers

  • Shift engineering investment toward correctness: Cheaper AI-generated code moves engineering value toward defining business rules, system boundaries and acceptable outcomes. Engineering leaders can prioritize architecture and specifications that make those conditions explicit.
  • Bound agent work with reliable feedback: AI agents converge more reliably when inputs, outputs, failure conditions and definitions of done are clear. Platform teams can increase agent autonomy where tests and deterministic tools provide specific corrective signals.
  • Make enterprise complexity observable: Mutable state, evolving APIs, coupled systems and undocumented rules create uncertainty that repeated agent iterations cannot resolve by themselves. Architecture owners can expose current dependencies and requirements through machine-readable contracts and controls.
  • Encode business semantics as correctness criteria: Technically valid pipelines can still produce business-invalid results when semantic rules remain implicit. Data owners can define field meaning, derivation rules and ownership so automated tests and agents can detect semantic errors.
  • Design architecture for safer autonomy: Semantic layers, immutable event logs, data contracts, idempotent APIs and deterministic state machines give agents dependable boundaries. Engineering organizations can keep these constraints current so greater agent autonomy remains aligned with business and system requirements.

Alexander Procter

October 7, 2026

11 Min

Okoone experts
LET'S TALK!

A project in mind?
Schedule a 30-minute meeting with us.

Senior experts helping you move faster across product, engineering, cloud & AI.

Please enter a valid business email address.