Cloud migration starts with deciding whether the move creates value

An application can reach the cloud on schedule, pass its cutover tests, and still leave the business with higher costs and the same architectural constraints. The 2024 Accelerate State of DevOps Report from DORA found that moving to cloud without adopting its inherent flexibility can produce worse outcomes than remaining on-prem. That finding puts the first migration decision ahead of infrastructure: teams need to decide whether an application should move, how much it should change, and which target architecture can deliver the intended business outcome.

Rehosting, or lift-and-shift, makes that decision concrete because the application moves largely unchanged. It reduces migration effort and risk and gets an application into cloud quickly, but its long-term optimization value is low because architectural debt moves with it and compute may fit cloud pricing poorly. Right-sizing, reserved instances, auto-scaling, and managed services can address cost and performance regressions, so even a relatively simple move requires an architecture review rather than treating cloud as a new location for the existing system.

That architecture review has to connect the move to outcomes such as better scalability, higher reliability, and lower cost at scale. A technically clean migration cannot establish that those outcomes are plausible when the existing architecture or the team’s capabilities work against them. Migration planning therefore starts upstream of infrastructure selection, with evidence about what the application really does and a decision about whether moving it is justified.

Discover the system you actually have before choosing its destination

The evidence for that decision starts in production because diagrams and wikis can describe a system that no longer exists. Techstack, a migration consulting provider that benefits commercially when organizations buy migration services, uses a discovery-first assessment in which automated tools inventory running services, APIs, data flows, and integration points while its team records inter-service dependencies, shared databases, and other integration touchpoints. From that work, Techstack produces a dependency map and a risk-tiered application inventory, giving engineering and business stakeholders a current-state model they can review before migration choices become commitments.

That current-state model also needs enough time-series evidence to reveal workloads that a short observation window could hide. Teams should collect at least 30 days of utilization data before making the migration decision, which can expose batch processing, month-end work, and less frequent traffic patterns. Latency, error rates, and throughput should be baselined over that period so target designs and later migration results can be compared with observed production behavior.

The risk of relying on records alone is clear for a VP of Engineering at a 500-person company responsible for a 12-year-old order-management system. The engineers who originally built it have left, its wiki is outdated, and three undocumented nightly batch jobs send data to finance. A migration based on that wiki can appear complete until a cutover interrupts those jobs, turning a documentation gap into an operational incident with consequences outside the application team.

Because documentation gaps can become operational failures, they belong in discovery as explicit risk inputs. Automated runtime discovery can surface dependencies that static records missed, while a living dependency register records what the team learns and changes as migration waves expose more relationships. Undocumented workflows and shadow IT, meaning technology used outside formal IT ownership or records, also belong in the risk inventory because production behavior establishes what must be preserved regardless of its documentation status.

Those runtime dependencies matter especially around shared databases. Two applications can look independent at the service level while reading and writing the same database, which couples their behavior when one application moves independently. Techstack classifies any application sharing a database with another migration candidate as high-complexity, and the consulting provider says its migration experience identifies shared databases as a frequent reason timelines fail.

The technical map also needs ownership evidence because dependencies alone cannot establish business importance or acceptable risk. A service with unclear ownership or a workflow whose business sponsor cannot be identified creates a decision problem because nobody can reliably establish its importance, acceptable downtime, or retirement conditions. By the end of discovery, the team should have a stakeholder-reviewed dependency map, 30 or more days of utilization data, automated inventory results, current performance baselines, and explicitly identified uncertainty, turning unknown production behavior into risks that can be classified and managed.

Okoone experts
LET'S TALK!

A project in mind?
Schedule a 30-minute meeting with us.

Senior experts helping you move faster across product, engineering, cloud & AI.

Please enter a valid business email address.

The first migration decision is whether to migrate at all

With the current system visible, each application can be assessed on four dimensions: business criticality, technical debt, migration complexity, and cloud-native fit. Engineering and business stakeholders then validate the assessment together, while unclear ownership or missing sponsorship is escalated rather than quietly carried into execution. Shared databases remain explicit complexity factors because classification based on application code alone can substantially understate the work.

Those dimensions become actionable through the 6 R’s framework, which assigns each application a migration treatment instead of assuming all applications need the same cloud path. The effort, risk, and long-term optimization levels below express the practical trade-offs used in that classification.

Strategy Also known as Effort / risk Long-term optimization value Practical consequence
Rehost Lift-and-shift Low Low Moves fastest but carries architectural debt and cloud-pricing inefficiencies forward
Replatform Lift-tinker-and-shift Medium Medium Makes targeted improvements, such as adopting a managed database, but requires enough application knowledge to change it safely
Refactor / Re-architect Re-architect High High Offers the highest stated long-term ROI potential, while teams often underestimate effort and risk
Repurchase Drop-and-shop Medium Medium Replaces commodity functionality with SaaS or another commercial product, with data migration remaining a material risk
Retire Decommission Low Low Switches off an application and is the most cost-effective option, with more candidates often available than organizations recognize
Retain Revisit Low Low Leaves the application in its current environment for now; useful as a temporary decision but weak as permanent avoidance

The classification matters when it changes the work engineering performs. A stable, low-business-value application due for retirement within 18 months is a candidate for Retain or Retire, while a core application facing growing demand can justify the risk and effort of Refactor or Re-architect. When time matters and architectural debt is manageable, Replatform can capture meaningful cloud benefits without committing the team to a full rebuild.

Repurchase addresses workloads whose commodity functionality does not justify continued ownership of a custom implementation. Moving such a workload to SaaS can remove the need to migrate and operate its existing application stack, although the replacement still requires deliberate data migration planning. Retire takes that reduction further by removing an unnecessary workload entirely, so the inventory becomes an opportunity to reduce the estate instead of reproducing all of it in cloud.

Those technical classifications need a financial test because economics can invalidate a technically straightforward strategy. Total cost of ownership, or TCO, should enter the architecture decision before provisioning. If lift-and-shift TCO exceeds current on-prem cost and refactoring cannot create sufficient value, preserving the current environment can be the sounder decision.

The same assessment must account for organizational capability and compliance requirements. Migration may be inappropriate when required skills are unavailable and there is no budget for outside support, or when cloud compliance requirements add complexity without proportional business benefit. An application scheduled to disappear within 18 months is another clear case where migration work can consume resources that will soon be discarded.

Together, those tests make retention, retirement, SaaS replacement, and incremental modernization deliberate outcomes when the evidence supports them. That range matters especially for legacy systems that would otherwise absorb disproportionate migration effort. Applications that pass the technical, business, capability, and TCO assessment can then move to target architecture design.

Modernize only as far as the application and team can support

For applications that proceed, architecture design comes before infrastructure provisioning because the migration strategy determines what must be built. A rehost requires deliberate mapping of existing compute, storage, and networking to cloud equivalents, including documentation of differences. Replatforming and refactoring require an explicit target state because the degree of modernization determines the networking, data movement, deployment, and operational work that follows.

AWS, which commercially benefits when customers adopt its cloud services, provides several conditional examples of those target choices. ECS on Fargate can be a starting point for request/response workloads, while AWS Lambda fits event handlers and EKS fits cases where Kubernetes is actually required. Those choices follow workload requirements because adopting EKS merely because it is available adds Kubernetes operations even when the application has no requirement that justifies them.

The same requirement-first discipline applies to legacy monoliths, where the decomposition pattern should be selected before infrastructure code is written. A strangler fig approach puts a facade or proxy in front of the existing monolith and sends newly built or refactored functions to separate services while older functions continue running in place. As more functions move, the monolith progressively becomes smaller, allowing modernization to proceed without a single wholesale rewrite.

Where a network proxy is impractical, branch by abstraction moves the transition inside the existing codebase. Engineers introduce an abstraction in that codebase, build the new behavior behind it, shift use to the new implementation at code level, and eventually remove the obsolete implementation. AWS Migration Hub Refactor Spaces can support this pattern, a connection cited by the AWS .NET Modernization Blog in 2023; AWS has a commercial interest in adoption of the service.

Whichever decomposition pattern is used, changing code boundaries does not by itself separate data ownership. A monolith split into several services remains tightly coupled if those services still read and write the same tables, and separating a shared database can take months. Splitting schemas before deciding the correct domain boundaries can make integration harder, so data decomposition needs its own incremental plan.

That incremental plan can start with domain-driven design or event storming with domain experts to identify bounded contexts, meaning areas of the business model with clear conceptual and ownership boundaries. Read replicas can reduce load on the primary database, while change data capture, or CDC, records incremental database changes and streams change logs to Kafka. An anti-corruption layer then translates those database changes into domain events so the old database representation does not determine the interfaces of the emerging domains.

Once those boundaries are established, schema ownership can change progressively. Each service can gain its own schema one bounded context at a time, while CDC keeps old and new data paths synchronized during the transition; AWS DMS is one example of a tool for continuous replication and is another AWS service the company benefits from customers adopting. Because full separation may require months, the work can continue alongside early migration waves instead of delaying all movement until the database has been completely reorganized.

The target architecture also has to fit the team that will operate it. A modular monolith preserves explicit domain boundaries within one deployable application and avoids much of the operational complexity introduced by distributed services. When engineers lack distributed-systems experience, or the workload’s scale does not justify microservices, that structure can be a durable target.

A stable application can also gain new behavior at its boundaries rather than through internal decomposition. In Techstack’s 9-dots menu pattern, multiple legacy monoliths were connected through an integration layer so they could operate as a unified product without rewriting each system. Techstack says the individual monoliths stayed stable while the user experience became coherent, an example the migration consultancy uses to show how modernization can occur at system boundaries when rewriting mature applications would add unnecessary risk.

Different workload characteristics can support a very different target architecture. Techstack’s fundraising-platform example involved unpredictable traffic spikes and high throughput, so the company built an AWS Lambda serverless architecture that could scale with demand. Techstack also presents the design as supporting reliable scaling and low-downtime deployment, claims that support the kind of migration and modernization work it sells; the decision for engineering teams is whether traffic, architecture, and operational capability have the same fit.

That fit also includes security, compliance, and operations because each can change the target architecture. Security and compliance requirements should be documented before migration because adding them later creates rework and audit risk, while logging, tracing, and alerting need to be defined as parts of the target architecture. Designing those requirements before provisioning keeps infrastructure choices subordinate to the operating model the application actually needs.

Execute migration as a series of reversible experiments

Once the target is justified and designed, execution can test those design choices with two or three non-critical, low-dependency applications. A pilot exposes how the target environment behaves under the organization’s own networking, security, deployment, and operating practices, and the team should review the pilot retrospectively before broader waves begin. Techstack’s wave-based methodology uses this progression to turn the inventory into a phased roadmap with sequence, owners, and explicit go/no-go criteria, reflecting the migration approach the consultancy sells.

As the pilot expands into waves, cross-wave decisions need a named owner so separate teams do not create incompatible patterns. A migration architect is accountable for consistency in architecture decisions, naming conventions, networking patterns, and security controls as applications move. Capability gaps should also be identified during planning, with external migration support considered for the first one or two waves when internal teams need practical experience alongside experienced practitioners.

That ownership must extend to measurable operating agreements before anyone schedules cutover. Latency, error rates, and throughput baselines provide comparison points and can become rollback triggers, while acceptable downtime, escalation paths, and rollback thresholds should be written down. An incident at 2 AM is the wrong time for stakeholders to discover that “We’ll figure it out if something goes wrong” was the effective agreement.

With rollback conditions agreed, data movement can start before the cutover window. The team should confirm that the source environment still matches the dependency map, start continuous replication through AWS DMS or an equivalent mechanism, and validate replication every day; AWS benefits commercially when its DMS service is selected. After deployment to the target, functional, performance, security, and smoke tests, which are quick checks that essential functions work, must pass before production traffic begins to move.

Passing those checks allows traffic exposure to increase gradually during a low-traffic period, from 10% to 50% to 100%. DNS, Application Load Balancer rules, and API Gateway can control routing; ALB weighted routing and API Gateway canary deployments, which expose a new version to a limited share of traffic first, support those gradual shifts. These are AWS services, so AWS benefits from their adoption. DNS or load-balancer cutover should first be tested in staging so production uses a routing mechanism the team has already verified.

Gradual traffic movement is useful because rollback remains executable while the new environment is being validated. The on-prem environment should stay live for one to two weeks as a rollback target, with explicit triggers such as error-rate thresholds, latency degradation percentages, or signals of data inconsistency. Keeping the old environment available costs resources for a limited period, but it gives the team a prepared response when measured behavior crosses those thresholds.

Those controls give each wave concrete completion conditions without turning cutover into a single irreversible event. Replication is checked daily before traffic moves, the required tests establish target readiness, and the old environment remains available afterward for the agreed rollback period. The next wave can then inherit what earlier migrations established while the migration architect keeps the resulting patterns consistent across the program.

A cutover is not the finish line

Once production traffic has moved, live behavior tests whether the target architecture delivered what its design predicted. Teams should compare latency, errors, and throughput with their pre-migration baselines and investigate regressions immediately. The evidence gathered before migration now provides the reference for determining whether scalability and reliability actually improved under real demand.

Operational validation then feeds into a longer financial check. Cloud spending should be compared with TCO projections at 30 and 90 days, and compute should be right-sized from observed usage rather than assumptions made before migration. On-prem resources remain available through the extended validation period and are decommissioned only after that period passes without rollback conditions being triggered.

Each completed wave also changes what the organization knows about the remaining estate, so the retrospective becomes an input to the next migration rather than a record of the previous one. The team should document what broke, what slowed execution, and which tooling gaps appeared, then address those findings before the next wave starts and adjust its sequence from the lessons learned. The dependency register can continue evolving as production reveals relationships that earlier discovery missed.

Key executive takeaways

  • Start with business value: Application owners should test migration choices against scalability, reliability, cost, and architectural constraints before committing resources. A successful cutover creates limited value when the target preserves the same constraints at higher cost.
  • Map the production system: Engineering teams should build a stakeholder-reviewed dependency map from runtime discovery and at least 30 days of utilization and performance data. Shared databases, undocumented workflows, shadow IT, and unclear ownership belong in the risk assessment.
  • Decide whether each application should migrate: Technology and business owners should assess business criticality, technical debt, migration complexity, cloud fit, TCO, compliance, and available skills. Rehost, replatform, refactor, repurchase, retire, and retain should remain valid outcomes.
  • Match modernization to operating capability: Architecture teams should select target patterns based on workload requirements and the organization’s ability to operate them. Incremental decomposition, explicit data ownership, or a modular monolith may offer better risk and value profiles than wholesale modernization.
  • Make migration reversible: Migration owners should begin with low-risk pilots, move applications in waves, continuously replicate data, and define measurable rollback triggers before cutover. Gradual traffic shifts and a temporary live source environment limit the impact of production regressions.
  • Measure value after cutover: Engineering and finance teams should compare production performance with pre-migration baselines and cloud spending with TCO projections at 30 and 90 days. Each wave’s operational findings should update dependency records, architecture patterns, and plans for subsequent migrations.

Alexander Procter

October 1, 2026

15 Min

Okoone experts
LET'S TALK!

A project in mind?
Schedule a 30-minute meeting with us.

Senior experts helping you move faster across product, engineering, cloud & AI.

Please enter a valid business email address.