A strangler migration is finished when the legacy system is retired
A strangler migration can deliver its replacement capabilities successfully while leaving the modernization program unfinished. Replatforming engagements behind these observations have repeatedly stalled around “80% done”: visible capabilities are live and stakeholders have recognized the delivery, yet the legacy system remains necessary. In one case, it still handled 15% of production traffic. That remaining dependency matters because incremental migration captures its full risk reduction only when the old system can be switched off.
That retirement condition changes what teams should measure. Feature parity shows whether the new system performs the work planned for it, while retirement shows whether production has stopped depending on the old system. A facade can still send a small set of reads or writes to legacy after the replacement appears complete, while batch jobs or downstream consumers can preserve dependencies that never appear in the product interface. Completion is therefore an operational state with specific traffic, data, and ownership conditions.
Those conditions also change how engineering leaders should use the strangler fig pattern. Incremental replacement reduces the risk of a single large cutover because individual capabilities can move and roll back separately. During migration, however, the organization runs a temporary dual-system architecture, so both systems and the machinery connecting them require engineering attention. Shipping migrated capabilities starts that period; retiring legacy ends it.
Routing is the easy part of the pattern
That temporary architecture starts with a routing mechanism established more than two decades ago. Martin Fowler named and defined the strangler fig pattern in 2004 in his bliki: an interception layer sits in front of the system, capabilities move incrementally, and requests go to the appropriate implementation until legacy receives nothing. Wikipedia gives the same core definition, while Microsoft Learn’s Azure Architecture Center and AWS Prescriptive Guidance describe cloud-specific implementations. Architecture-pattern catalogs, modernization decision guides, and monolith-versus-microservices comparisons place the mechanism among broader modernization choices.
In practice, every client request first reaches a facade, usually a reverse proxy or API gateway, which chooses between the old and new implementations. The facade can make that choice from a URL path, request header, or feature flag, so clients can keep the same interface while a capability moves. Microsoft’s Azure guidance uses API Management or Application Gateway for this interception, sending migrated paths to the replacement and other paths to legacy; Microsoft benefits when teams adopt Azure services for that design. AWS, which likewise benefits from use of its cloud services, describes an analogous implementation with Amazon API Gateway or an Application Load Balancer.
The interception point makes capability-by-capability migration possible. A team can send one domain to the new implementation behind a feature flag, observe its behavior, and restore the legacy route if a problem appears without changing other domains. The rollback boundary can therefore be much smaller than the application. Gradual production exposure becomes practical because the program does not depend on one irreversible cutover date.
Because routing supports migration, the facade has a deliberately temporary role. An anti-corruption layer has a different purpose: it isolates one model from another system that the architecture intends to retain. A strangler facade should disappear after the old system is gone unless the organization deliberately gives it a permanent role. Until then, its place on the request path means it must be deployed, patched, monitored, supported, and included in on-call procedures while both application stacks receive the same operational treatment.
That operational burden grows when coexistence lasts longer than planned. Experience reported across replatforming engagements includes old and new systems running together for extended periods, while a separate qualitative observation describes overlaps lasting “sometimes for years.” That duration is an observation rather than the result of a named study. Every additional month carries the cost of two stacks plus routing, synchronization, and reconciliation workloads, so lower cutover risk comes with a continuing parallel-operation cost.
A project in mind?
Schedule a 30-minute meeting with us.
Senior experts helping you move faster across product, engineering, cloud & AI.
Early wins push the hardest dependencies into the final 20%
That parallel-operation cost depends on migration order because sensible early sequencing creates a predictable later problem. Teams naturally start with capabilities that have clear interfaces and limited dependencies, since those slices can prove routing and rollback before the program reaches its hardest data problems. Such migrations tend to go smoothly. Their success can then create more confidence than the remaining estate justifies.
Starting with the biggest or most painful module creates the opposite problem. Teams often choose it because its business value is easiest to justify, yet the module may touch tables written by five other services, batch jobs that nobody fully owns, and fields that cannot be renamed without breaking a downstream report. Moving it first forces the team to confront difficult shared-state, synchronization, and reconciliation problems before establishing a migration process. High business value alone does not create a useful first boundary.
A better first slice follows from that dependency problem: it has few upstream dependencies and a clean interception point. Preferably, it is read-heavy, has a narrow write surface, and can move behind the facade without immediately taking ownership of heavily shared state. Chris Richardson’s pattern and decomposition catalog supports a focus on loosely coupled capabilities with defined boundaries, including decomposition by business capability or subdomain. Such a slice can prove routing, the selected synchronization or dual-write path, reconciliation, observation, and rollback under production conditions.
Once those controls work, the program can apply them to harder slices. A capability that widely shares mutable data should wait until a defined ownership model and tested rollback process exist; coupling that cannot be separated may put the capability outside the strangler sequence. For a monolith, boundary discovery may therefore need to happen first as a separate decomposition exercise. After teams establish business-capability or subdomain boundaries, they can move one bounded-context slice at a time while other domains continue toward the legacy implementation.
Good sequencing still tends to leave the weakest boundaries until late. The remaining estate can include batch jobs with undocumented effects, legacy modules that three teams still write to directly, and reporting pipelines that continue consuming the old schema. These dependencies are operational obligations with owners, or gaps where an owner should exist. Unless each one enters the migration scope, it can survive every successful product cutover.
That is how “80% done” becomes a durable state. The early slices were tractable, while later work was difficult to estimate, politically unclear, or never assigned. Once the flagship capability runs on the new system, business stakeholders can reasonably see the investment as delivered. When the remaining jobs, writers, and consumers have neither owners nor funded line items, nobody has a mandate to remove them.
The earlier case with 15% of traffic still reaching legacy shows the consequence. Most traffic had moved, but production still needed the old implementation, so the facade and both application stacks remained operational dependencies. Completion by traffic volume therefore says little about whether infrastructure can actually be retired. Permanent fallback preserves the old-system dependency even when feature flags made each earlier migration step safely reversible.
Dual-system operation turns data ownership into the critical engineering problem
Those late dependencies often converge on data ownership because request routing cannot decide which stored state is authoritative. During overlap, both implementations may participate in a business domain, so each domain needs a declared source of truth before its first slice goes live. Without that declaration, a mismatch between stores becomes an argument about correctness after a failure occurs. The facade can route calls, but it cannot decide which database wins when two versions disagree.
Dual writes offer a quick way to synchronize those databases. An application receives one update and writes it to both legacy and replacement stores, and the implementation may fit into a single sprint. The failure mode appears when the first write succeeds and the second fails, leaving one store ahead without an inherent signal that identifies the authoritative value. Replatforming experience described here includes teams choosing that short implementation path and then spending months investigating the resulting drift.
Change data capture (CDC), meaning the capture and replication of database changes as they occur, moves synchronization out of the synchronous application write path. With log-based CDC, the migration consumes changes from the legacy database’s write-ahead log and asynchronously applies them to the new store. Confluent’s CDC documentation says a properly tuned log-based pipeline typically keeps replication delay under a second; Confluent has a commercial interest in adoption of its data-streaming technology. By comparison, reported polling-based synchronization in these migrations had minutes of delay, although that figure is not attached to a separately named producer.
Because CDC changes rather than eliminates the failure surface, state still has to be verified. Replication can be delayed or fail elsewhere in the processing chain, mappings can be wrong, and downstream logic can still disagree about ownership. Reconciliation must therefore be a planned migration deliverable: compare relevant states, identify drift, establish which system supplies the correct value, and run the verification before declaring a cutover complete. Leaving reconciliation until cleanup delays discovery of the inconsistencies that determine whether legacy can safely retire.
That verification depends on a specific domain-level ownership rule. At every migration gate, the practical question is “what still writes to the old source, and who owns shifting it,” because an unexpected writer preserves the legacy dependency after ordinary application traffic has moved. Direct database writers are particularly easy to miss because they bypass the facade. Finding one late can invalidate an apparently complete cutover.
Shared mutable state makes the ownership rule even more important because old and new software can update the same records concurrently. Without a clear boundary, those updates can race under production load and create results that a periodic comparison job repeatedly has to repair. Reconciliation then changes from a migration control into a recurring incident workload. Stable coexistence therefore requires exclusive or otherwise explicit write ownership.
As long as coexistence continues, so does its workload. Teams must maintain the old stack, new stack, facade, CDC or dual-write machinery, reconciliation jobs, monitoring, alerting, and on-call coverage together, while infrastructure and licenses for the parallel systems remain in the budget. The temporary migration architecture is consequently technical debt that must be tracked alongside the legacy code being removed. Its cost ends only when its components become unnecessary.
Yet ownership can weaken while that work still needs sustained attention. The replatforming experience behind these observations says reconciliation often receives resources only through the “first quarter,” an operational observation without a named quantitative study. Reconciliation is needed until the data transition is demonstrably complete, so the end of the initial migration team’s assignment is an unsafe substitute for an engineering completion condition. A reconciliation process without a long-lived owner can keep producing results without anyone responsible for acting on them.
A clean result also has to span meaningful business activity. Zero drift on a quiet day says little about monthly processing, scheduled jobs, reporting runs, or other infrequent work. Before final retirement, reconciliation should remain clean across a full business cycle so those paths have an opportunity to execute. Data ownership, synchronization, and reconciliation thus become controls over the whole period in which two production systems coexist.
Fund retirement gates and the last 20%
Because unresolved dependencies keep that operating model alive, funding has to cover their removal. A program organized around feature parity can fund replacement capabilities and exhaust its mandate while old writers and consumers remain. Retirement-linked decision gates put those dependencies inside the funded program. Before money for the next slice is released, the gate should identify every remaining legacy writer, assign an owner, and establish how it will move.
The final phase deserves an explicit budget because early feature plans tend to hide its work. Ownerless jobs have to be investigated, direct writes redirected, downstream reports changed, data drift eliminated, and infrastructure decommissioned after those dependencies disappear. Calling this the last 20% is an observed program pattern rather than a universal measured ratio, so the exact share can vary by estate. The management requirement remains the same: retirement work needs a line item before visible delivery creates pressure to declare success.
Program risk makes that requirement relevant beyond strangler migrations. McKinsey Digital reported in 2012 that large IT projects averaged 45% over budget and 7% over time. Those figures concern large IT projects generally rather than strangler programs specifically. They support budgeting for work that continues after new functionality reaches production instead of assuming the original feature scope and schedule will absorb every late dependency.
A large rewrite with one funding line and one cutover has a different risk profile. Its spending can appear simpler because the organization budgets one replacement and one transition date. Its failures, however, can surface late when much of the program must work together. A strangler approach distributes those decisions across incremental slices and allows per-slice rollback, while its business case must absorb the cost of parallel operation; the choice depends partly on whether the estate can be scoped reliably enough for a focused replacement.
For a strangler program, that funding model needs an operational, testable completion gate. It should require:
- Zero production traffic to the legacy system, with 100% of reads and writes routed to the replacement.
- A named new-system owner for every domain previously owned by legacy.
- Zero reconciliation drift across a complete business cycle.
- Every downstream consumer of legacy data or APIs repointed to the new system or formally deprecated.
Those conditions expose dependencies that feature parity can miss. A replacement can implement every user-facing function while a reporting process still reads the old schema, or while one team writes directly to a legacy database. Switching legacy off under either condition would break a remaining consumer. Keeping it running avoids that immediate break while preserving the cost and dependency the modernization program is intended to remove.
The same gate gives the business case a concrete completion test. Traffic share establishes whether the old execution path remains active, ownership establishes whether responsibility moved with the capability, reconciliation establishes whether state survived the transition, and downstream migration establishes whether hidden consumers have moved. Once those conditions hold, engineers can schedule retirement as an executable change rather than predict that nobody depends on the system.
Temporary migration infrastructure needs its own exit plan
Retiring legacy then exposes one final dependency created by the migration itself: the facade. Replatforming engagements have seen temporary facades outlive the legacy applications they were introduced to help remove because a service on every migration request gradually acquires production responsibilities. Logging, alerts, on-call runbooks, and maintainership accumulate around it. Those responsibilities can make preservation the default even after its original routing job disappears.
One observed facade had been scoped for a “single quarter” yet was still routing traffic “two years later” because removal had no owner. By then, continued operation could feel safer than changing a component carrying production traffic, even though its original migration purpose had expired. Temporary infrastructure therefore needs its own retirement owner and gate from the beginning. Without that ownership, accumulated support obligations can turn another migration component into permanent platform work.
After legacy shutdown, the team has two coherent outcomes. It can remove the facade because the old/new routing decision no longer exists, eliminating the migration-specific layer and its maintenance burden. Alternatively, the organization can decide that the facade now provides enduring integration value and retain it deliberately. A retained facade should receive a new name, an explicit owner, and an SLA so its production status reflects its permanent role.
Some systems should not be strangled
The cost of creating and later removing those temporary components sets a boundary around where the pattern is useful. Strangling requires an interception point through which migration traffic can be divided. Tightly coupled desktop clients, batch processes that read databases directly, and systems with neither an API boundary nor a message bus can lack that control point. For those systems, creating architectural separation may need to become a separate prerequisite project before incremental routing is viable.
Indivisible shared mutable state can create an equally strong blocker. If old and new implementations cannot receive clean write ownership, a facade leaves the core state problem unresolved, so database separation and domain-boundary work may become the substantive modernization effort. Teams should establish business-capability or subdomain boundaries first and then determine whether the resulting slices support a strangler sequence. Some highly coupled modules may remain unsuitable even after other parts of the estate move successfully.
Small estates face a different constraint because parallel operation has a real minimum cost. The replatforming experience here includes teams spending significant engineering effort building a facade around an application that a small team could have rewritten outright. When an estate can be specified completely and replaced in a focused effort, operating two systems together with synchronization and reconciliation may cost more than the incremental approach’s reduction in cutover risk. Suitability therefore depends on the architecture and economics of the individual estate.
That suitability decision turns on whether incremental replacement can create separable ownership and a credible retirement path. An estate without a useful interception boundary may need architectural separation first; mutable state that cannot be divided may require domain or database work; and an estate where dual operation costs more than a focused rewrite may favor direct replacement. The strangler fig pattern earns its risk advantage where each migrated slice moves the organization measurably closer to switching the old system off.
The bottom line
The strangler fig pattern should be judged by what it removes, not only by what it delivers. Moving capabilities incrementally can reduce cutover risk, but the business continues carrying legacy cost and exposure while old traffic, data dependencies, consumers, or migration infrastructure remain in production.
For executives, that makes retirement a program objective rather than a post-launch cleanup task. Funding, ownership, and governance should extend through the point where legacy can be switched off safely. Retirement gates give leaders a practical way to test progress by asking whether traffic has moved, data reconciles, downstream consumers have migrated, and every remaining dependency has an owner.
This also creates a clearer basis for deciding whether to use the pattern at all. Where capabilities can be separated and ownership transferred incrementally, strangling can exchange one large cutover risk for a series of smaller, reversible changes. Where shared state cannot be divided or parallel operation is disproportionately expensive, a focused replacement may be the better investment.
The strategic distinction is simple: a replacement system going live creates new capability; a legacy system going dark captures the full modernization outcome. A strangler migration is complete only when the organization no longer needs the system it set out to replace.
A project in mind?
Schedule a 30-minute meeting with us.
Senior experts helping you move faster across product, engineering, cloud & AI.


