The wrong way to measure forward-deployed engineering

A working AI deployment on real customer data can still be evidence of a product that is failing to learn. Forward-deployed engineering, or FDE, puts engineers inside customer environments to connect products to operating systems and turn demonstrations into working deployments. The familiar pitch is compelling: an embedded engineer arrives, encodes a workflow within weeks, and gets the system working on the customer’s data. Buyers can reasonably see speed, while investors can see growing FDE headcount and infer growth.

Those visible signals leave the important question unanswered because weak and strong FDE can produce the same successful first deployment. In weak FDE, engineers repeatedly supply capabilities the software cannot provide on its own. In strong FDE, they discover architectural edge cases and turn those discoveries into capabilities that later customers can reuse. The organizations can look similar in both cases while the product economics move in very different directions.

The difference becomes visible with the next comparable customer. Does that customer arrive to more working product and fewer unknowns, or does another engineering team have to translate the environment by hand? First-deployment speed shows that capable people can make the system work. Improvement in the next deployment shows whether their work made the system itself more capable.

Enterprise AI often needs context the model does not contain

Human translation has a legitimate role because enterprise AI often depends on knowledge absent from both the model and formal data. Model choice remains consequential in some domains, but many business workflows depend on rules, exceptions, definitions, and operating logic accumulated over years. Access to a company’s data provides part of the required context, while understanding how its people make decisions from that data requires another layer of knowledge.

A large telecommunications deployment shows where that additional context appears. The initial definition of a “high-intent” customer conflicted with the retention team’s real save-desk criteria: the model produced one signal, while the people responsible for retention acted on another. Their criteria incorporated historical knowledge about which offers had worked for particular tenure bands and regions. No schema recorded those rules because the relevant knowledge lived in the judgment of employees who had been doing the work “for a decade.”

That gap defined the field engineer’s job. The engineer sat with experienced employees, extracted the criteria they applied, and encoded that logic so the system could use knowledge it could not otherwise see. Once that knowledge was explicit, the intelligence layer could be trusted to trigger an action rather than merely produce a score. The intervention changed what the software could safely decide.

The change also affected subsequent work. New acquisition and retention use cases went “from idea to execution in days rather than months” because teams could add decisions to a shared foundation instead of reconstructing the integration for each use case. The first intervention therefore mattered beyond the initial retention workflow. It left behind logic that reduced the work required to turn another business decision into an operational use case.

That persistence is what a system of intelligence requires. Such a system combines software with enterprise context, retains useful knowledge from deployments, and applies it to improve later decisions. Forward-deployed engineers can supply the initial context layer by discovering business rules, exceptions, workflow logic, and definitions that formal systems have never captured. The important transition occurs when knowledge first delivered through a person becomes persistent capability.

Without that transition, the same reason FDE is valuable can also make it expensive. Undocumented organizational knowledge has to be found and encoded somewhere, so engineers may end up reconstructing integrations, workflows, decision logic, and missing functionality for customer after customer. A successful deployment can hide that repetition because the customer still gets a working result. At organizational scale, the question is whether the system accumulates context or the field organization repeatedly recreates it.

Okoone experts
LET'S TALK!

A project in mind?
Schedule a 30-minute meeting with us.

Senior experts helping you move faster across product, engineering, cloud & AI.

Please enter a valid business email address.

The real test is whether field learning survives the deployment

Once tacit context has been uncovered, where it goes determines whether FDE compounds. A useful field discovery might become a semantic mapping, policy module, workflow template, connector, or evaluation. Each artifact captures part of a problem so another deployment can begin with more knowledge already encoded. FDE then becomes a human-delivered context layer whose useful discoveries move into the product.

That destination separates a “sandbox” model from a “mud” model. In the sandbox model, engineers take a general-purpose engine into a difficult customer environment, identify what it lacks, implement the missing piece, and return the resulting learning so the capability can be reused. In the mud model, engineers construct missing functionality manually for a particular customer, while the underlying system has no effective route for absorbing that work as reusable capability. Both can produce an impressive deployment.

Most companies sit somewhere between those endpoints, combining reusable connectors and playbooks for common situations with bespoke judgment where customers differ. That middle ground makes the sandbox-versus-mud distinction a diagnostic for how field work moves through an organization. Customized engineering by itself reveals little about whether a vendor is building a compounding product. The useful question is what happens to the knowledge created by that customization.

From outside the organization, those versions can look almost identical. An engineer can be on-site, working directly with customer data and writing sophisticated code whether the work is reusable, account-specific, or entirely bespoke. Headcount, technical résumés, and polished demonstrations expose the activity while concealing its destination. Stronger evidence appears when another comparable deployment begins.

A repeat deployment should reveal concrete differences if the earlier work compounded. The team should encounter fewer unknowns, write less custom code, and arrive with better tests because previous field discoveries have already been captured. A later engagement that effectively restarts the technical reasoning can still benefit from a company getting better at execution. Product learning requires a further result: the software itself must begin with more of the required understanding.

Producing that gain requires a deliberate learning loop. First, a team observes an exception in the field and codifies the useful part into a reusable artifact. It then validates the artifact through an evaluation and security review, releases the resulting capability into the product, and measures whether the next deployment actually became easier. Each stage matters because reuse is established only after the artifact has been codified, validated, shipped, and shown to affect subsequent work.

The final measurement deserves particular attention because the learning loop has been characterized as the point where “most companies quietly fail.” The diagnostic behind that characterization is concrete: a company has to measure what changed for the next deployment. Without that measurement, “learnings” can accumulate as documentation and internal knowledge while customers continue to require the same amount of human translation. Measuring the next deployment turns the claim of learning into an observable change in work.

That observable change separates two legitimate but economically different capabilities. A company may improve at running deployments because its engineers develop strong execution practices and customer relationships. Those strengths can support a capable services business even when the underlying product changes little. Product compounding occurs when reusable capability remains after the field engineer leaves and materially changes how future work gets done.

The central measure of FDE is therefore the reduction in human translation across comparable deployments. A fast engineer can solve one customer’s immediate problem, while reusable mappings, modules, connectors, templates, evaluations, and supported capabilities can change the starting point for the next customer. The first outcome proves delivery skill. The second proves that field experience is entering the product.

Customization needs a destination

Reducing human translation does not require every piece of customer logic to enter the core product. Some logic is proprietary to an account, some exists only for a temporary requirement, and some is too idiosyncratic to generalize usefully. Shipping those field customizations indiscriminately would mix customer-specific complexity with reusable product knowledge. The harder task is deciding what deserves to persist and where.

That decision puts FDE work into three practical buckets:

  • Product intelligence can compound across customers and should become reusable capability.
  • Configurable customer logic can remain reusable within one account while being inappropriate for broad distribution.
  • One-off services work solves a specific need and remains bespoke.

These buckets make customization itself weak evidence about the quality of an FDE model because enterprise deployment is expected to involve customer-specific work. The more useful test is whether the organization can identify which category a piece of work occupies and preserve discoveries that can benefit future customers. Classification determines whether field work is deliberately retained, locally configured, or treated as a one-time service. Its destination determines what the next deployment can inherit.

The classification also sets a practical boundary for productization. FDE compounds when teams distinguish bespoke work from generalizable learning, then demonstrate that the generalizable portion reduces the human translation needed in later deployments. The aim is disciplined retention of useful knowledge while necessary customer specificity remains where it belongs. With those categories defined, leaders can measure whether the balance changes over time.

Measure whether human translation is declining

Once the categories are clear, leaders can measure whether their FDE model is actually changing. The relevant quantity is human translation per unit of value delivered, which should decline as reusable logic moves into the product. Absolute FDE headcount can rise at the same time. A fast-growing company can employ more field engineers overall while making each deployment lighter.

That combination explains why growing FDE headcount alone says little about product compounding. More engineers may be required because the vendor is serving many more customers and workflows, while each engagement requires less repeated custom work than its predecessors. The field team’s time can then shift toward extending reusable capabilities instead of rebuilding integrations, workflows, and decision logic already encountered elsewhere. Headcount alone cannot reveal whether that shift is happening.

Five operating measures make the change visible. Together, they test deployment effort and whether useful discoveries move quickly enough into reusable capability.

Measure What it reveals Desired direction
Engineers per live workflow Human staffing required for operating deployments Decline as workflows become easier to support
Engineering hours per deployment Amount of custom implementation effort Decline
Time-to-value by vertical Whether accumulated vertical knowledge accelerates delivery Decline
Share of implementation work reused How much prior work survives into later deployments Increase
Productization lag Time from a field discovery to a tested capability available to the next customer Decline

Among those measures, productization lag exposes the organizational path from discovery to use elsewhere. A company may have excellent field insight yet leave it within individual teams for too long to affect subsequent customers. Shorter lag, lower custom engineering, and higher reuse show that discoveries are moving through evaluation, security review, product release, and deployment quickly enough to matter. The measures connect field learning to an observable change in later delivery.

If those measures do not improve, growing FDE capacity may simply increase delivery throughput. The desired human outcome is more precise than reducing the number of engineers: a larger share of what those engineers discover should persist as product capability. Knowledge that remains solely with field teams disappears from the product’s practical starting point for the next customer, even when the organization itself remembers it. That distinction gives buyers, investors, and product leaders a concrete basis for diligence.

Three questions reveal more than an FDE demo

Those mechanics make the first diligence question straightforward: “How is FDE priced?” Pricing is a signal rather than decisive proof. A separate professional-services line can reflect transparent accounting, while bundled FDE can operate as a loss leader funded through utilization. The more revealing commercial test is whether contracts, renewals, and margins distinguish repeatable productization from bespoke delivery.

Because pricing only shows the commercial structure, the second question moves to the learning path: “Where does field learning go?” Strong engineering résumés establish who is doing the work, while the handoff into product shows whether discoveries survive the engagement. Ask who owns that handoff, which artifacts come out of it, and how quickly they become tested and supported capabilities. The organizational interface provides stronger evidence of compounding than the field engineer’s job title.

Once that path is visible, the third question tests its effect: “What got faster on the last repeat deployment?” Require a specific vertical and a measured change, such as fewer engineering hours, fewer weeks to value, fewer custom integrations, or a higher reuse rate. Generic references to “learnings” and “playbooks” leave the crucial causal claim untested. A credible answer identifies what changed between comparable deployments and how the organization measured the difference.

Key highlights

  • Measure repeat deployments: FDE compounds when comparable customers require less human translation over time. Track whether engineering hours, unknowns, custom code, and time-to-value decline after earlier deployments.
  • Turn enterprise context into capability: Field engineers uncover business rules, exceptions, and operating knowledge that enterprise AI cannot access from formal data alone. Product teams that encode this context into reusable capabilities give later deployments a stronger starting point.
  • Build a field-to-product learning loop: Route useful field discoveries through codification, evaluation, security review, and product release. Measure whether those capabilities reduce work on subsequent deployments to verify that field learning survives the engagement.
  • Give customization a destination: Classify field work as reusable product intelligence, configurable customer logic, or one-off services work. This keeps account-specific complexity contained while preserving discoveries that can benefit future customers.
  • Track human translation per unit of value: FDE headcount can grow while individual deployments become more efficient. Monitor engineers per live workflow, engineering hours, time-to-value, reuse rates, and productization lag to determine whether the model is compounding.
  • Test FDE claims with operating evidence: Buyers, investors, and product leaders can ask how FDE is priced, where field learning goes, and what improved on the latest comparable deployment. Specific changes in engineering effort, delivery time, integrations, or reuse provide stronger evidence than demos, résumés, or generic references to playbooks.

Alexander Procter

October 7, 2026

12 Min

Okoone experts
LET'S TALK!

A project in mind?
Schedule a 30-minute meeting with us.

Senior experts helping you move faster across product, engineering, cloud & AI.

Please enter a valid business email address.