AI can make a growing application easier to keep unified, changing an architecture decision many developers were taught to make almost automatically. If you learned to code with AI in the last two years, your application may have grown from AI-generated scaffolding into a substantial Next.js application backed by Postgres before you consciously considered its architecture. When that application becomes slow, fragile, or difficult to change, Kubernetes is not automatically the next step. In 2026, the case for keeping more software together has become stronger.
That stronger case has a clear limit because AI has changed some software costs more than others. Microservices became attractive partly because large codebases and coordinated deployments placed heavy cognitive and organizational demands on humans. AI reduces some of those demands by navigating code, making mechanical changes, and helping create tests. Physical and process constraints still make independent scaling and fault isolation useful, so services continue to make sense when those requirements exist.
AI changes the default architecture decision
For an AI-assisted developer, the important change is when distribution becomes worth its cost. A tangled application can often be improved within its existing deployment boundary before being divided into separately operated systems. AI can make that unified codebase easier to understand and modify, extending the useful life of an architecture that previously became uncomfortable sooner for human developers and teams.
That longer useful life changes the default while preserving clear reasons to split a system. A component with genuinely different scaling requirements may belong in its own service, while a component whose failure must be isolated from the rest of the application may justify a separate process. Those exceptions depend on runtime properties that AI cannot remove. Their importance becomes clearer once the original technical and organizational case for microservices is separated from the human constraints that accompanied it.
Why microservices became the answer to growing software
That original case starts with the deployment boundary of a monolith: one application with one codebase, one build, and one deployment. User accounts, billing, product logic, and administration can all run in the same process, communicating through ordinary function calls. The arrangement keeps interaction within the application and normally gives developers one system to build and operate. As the application grows, however, every part remains within the same deployment boundary.
Microservices change that boundary by separating functions into isolated, independently deployable services. A billing service and a user service can run separately and communicate across a network through APIs or message queues. Each service can have its own database and programming language, and a separate team can own it. The separation creates technical and organizational independence, but network communication and multiple running systems then become part of the application architecture.
That trade became influential through examples such as Netflix. As its platform and engineering organization grew, its monolith became a bottleneck: even a small modification could mean rebuilding, retesting, and redeploying the whole platform. Netflix began splitting the system around 2009, and its migration became a model that many other companies followed. By roughly 2015–2020, microservices had become a common recommendation for software expected to grow.
The Netflix pattern also supported an organizational argument because large engineering groups need to make changes concurrently. Dividing a system into owned services can allow, for example, thirty teams to work at the same time with less interference between their releases and technology decisions. The illustrative thirty-team case reveals an assumption behind much conventional guidance: a company needs enough simultaneous independent work for service-level autonomy to repay the extra operating cost.
That assumption applies unevenly because smaller teams may inherit machinery designed to coordinate many independent groups without receiving the same benefit. Traditional objections to monoliths still matter, but they arise from different causes. AI changes the objections rooted in human effort much more than those rooted in runtime behavior.
A project in mind?
Schedule a 30-minute meeting with us.
Senior experts helping you move faster across product, engineering, cloud & AI.
AI weakens some monolith objections far more than others
The first objection, “Slower development speed.”, follows from an obvious property of a growing monolith: more code creates more relationships for a developer to understand, so changes can become slower and riskier. At an illustrative extreme of a million lines, no human can keep the entire application in working memory. Separating the application can limit how much of that code an individual team has to understand at once.
That cognitive limit is where AI changes the problem substantially. An AI agent does not experience human fatigue from reading through a large repository and can navigate a large codebase while writing or reviewing changes. The underlying code can still be complicated, and poor internal structure still creates dependencies that have to be understood. Repository size by itself becomes a weaker reason to introduce network and deployment boundaries when an AI assistant can inspect relevant code across the application.
The second objection, “Technology lock-in” and “lack of flexibility.”, also depends partly on how much work humans must perform. In a monolith, a shared framework and language can make a broad upgrade expensive, and a subsystem cannot simply adopt another language as an independent service can. Teams historically avoided some migrations because applying a large volume of mechanical changes across one codebase consumed too much engineering time.
That migration burden falls when AI can perform much of the mechanical transformation before humans review the result. A framework or language decision still affects a unified application broadly, so the architectural constraint remains. But the cost of escaping an aging technical choice has changed: migrations that once required extensive repetitive human editing can shift more of that work to AI, with engineers concentrating on review and the parts that require judgment.
The third objection, “Reliability.”, behaves differently because it depends on process isolation. If the payment module in a monolith throws an unhandled exception that crashes its process, the rest of the application goes down with it. Better AI-generated code may reduce the probability of some defects, but two modules running in one process still share a failure domain. Reliability requirements can therefore provide a concrete technical reason to split a system.
The fourth objection, “Scalability.”, reaches a similar physical limit. A monolithic process cannot independently add capacity to one internal component, so a hot workload can force the company to run more copies of the entire application. A video encoder is a clear example: if encoding demand grows independently from the rest of the product, scaling that work as a separate service can avoid scaling unrelated application functions with it.
The importance of that limit depends on whether an application actually reaches it. Most applications never reach the point where scaling individual components beats running additional copies of the monolith. Netflix-level scaling problems remain legitimate engineering problems for Netflix, but applications with different demands should make the decision from their own workload. Independent scaling earns its complexity when a workload gives a concrete reason for it.
The fifth objection, “Deployment.”, lies between these human and physical constraints. A “one-line fix” in a monolith still requires deploying the whole application, and historically that coupling could discourage frequent releases. Less frequent deployments could then become larger deployments, increasing the amount of change exposed at once and raising deployment risk.
Modern tooling reduces part of that burden because CI pipelines and AI-generated test suites can make building and validating a full application faster. One illustrative comparison puts a monolithic deployment at “2 hours” a decade ago and “a matter of minutes” now. The comparison captures the effect of faster tooling: deploying the whole application can require less time and manual effort. The modules still share one release cycle.
These five objections divide according to what creates their cost. AI has the largest effect where developers must absorb large amounts of code, perform mechanical migration work, or handle parts of testing and deployment coordination. Runtime properties remain: processes share failure domains, computing resources have to be allocated somewhere, and one deployable remains one deployable. Architecture decisions in 2026 need to distinguish these categories because each responds differently to AI assistance.
Microservices can turn code complexity into operational complexity
Once a team crosses that deployment boundary, code that was local becomes a distributed-system problem. A function call inside a monolith becomes a network interaction between services, where either side or the connection between them can fail. Following one request can then require distributed tracing, which records its path through multiple services, instead of following a conventional stack trace within one process. A defect that previously produced a stack trace can become “an afternoon of correlating logs across five dashboards.”
Those network failures create additional states that application code and operations must handle. A request can partially succeed when one service completes work and another fails, creating a need for retry behavior and careful handling of duplicate operations. Data can become eventually consistent, meaning separate parts of the system reach the same state after a delay rather than as one immediate transaction. Services can also develop version skew when different versions run or interact at the same time, adding compatibility concerns to ordinary application changes.
Those additional states bring infrastructure expense as well as engineering work because each independent service has to be deployed, observed, communicated with, and managed according to its role in the wider system. For an organization that needs independent ownership and releases, those costs can buy useful autonomy. A team without that requirement can instead end up operating distributed infrastructure because its codebase became difficult to navigate.
That mismatch reaches its worst form in a distributed monolith: multiple deployables that remain so tightly coupled that they have to change together. An illustrative system with fifty deployables gains little independence if a meaningful product change still requires coordinated releases across them. The organization then pays for network calls, distributed debugging, and separate deployments while retaining the coordination pressure of a unified application. Service boundaries deliver their central benefit only when they create meaningful independence.
The cost of unnecessary boundaries has led some organizations to consolidate specific workloads. In 2023, Amazon’s Prime Video team published a case study in which it moved a particular workload from microservices back to a monolith and reported cutting infrastructure costs by around 90%. The result shows that consolidation can produce a large economic gain in a real production workload. Its scope is that particular workload; it does not establish a general rule for Amazon or other production systems.
Broader evidence shows that teams are also reconsidering selected boundaries rather than treating every existing service as permanent. A 2025 CNCF survey reported that 42% of organizations that initially adopted microservices had consolidated at least some services into larger deployable units. Because the figure describes partial consolidation, it does not mean that 42% rejected microservices as an architecture. It shows that teams are willing to revisit service boundaries when the independence they provide no longer repays their operational cost.
The better default is a modular monolith
Reconsidering a service boundary still requires internal structure because consolidation should preserve clear responsibilities. A modular monolith keeps one process and deployment while dividing the application internally into modules with strong interfaces. Billing and users can remain separate modules, for example, with explicit public interfaces through which they interact. The billing code should not casually reach into the user module’s database tables simply because both modules live in the same application.
That internal structure starts with one repository, one deployment process, and one database. This arrangement gives an AI assistant broad visibility across the application, helping it trace dependencies and understand how a requested change interacts with surrounding code. For an application without an independent-scaling or isolation requirement, the same arrangement avoids creating distributed-system work before there is a technical reason to carry it.
Within that unified deployment, the second step is to impose modular boundaries before adding network boundaries. Billing and users can each expose a clear public interface, with internal code and data access kept behind it. An AI assistant can help check those rules during code review by flagging changes that improperly cross a module boundary. Clear internal interfaces preserve much of the boundary discipline associated with microservices while interactions remain ordinary in-process calls.
Those internal boundaries also preserve the option to extract a module later. If a module eventually develops a requirement for independent operation, its established interface can become the starting point for a service contract. Extraction is substantially clearer when callers already depend on a defined boundary than when they reach directly into another module’s internals. Modularity therefore prepares for a future split without imposing its operating cost in advance.
With that preparation in place, the third step is to extract a service when the team can name the requirement that demands it. Video processing is the recurring example because its compute requirements can differ sharply enough from ordinary application work to justify independent scaling and isolation. Once that requirement exists, the operational cost of another service buys a specific property the monolith cannot provide. Until then, keeping the component inside the modular application avoids network behavior the product does not need.
The fourth step is to carry forward the design concepts developed around microservices even while deployment remains unified. Bounded contexts define areas of a system with clear domain responsibilities; API contracts make the permitted interaction between areas explicit; failure isolation forces engineers to consider how one component’s problems affect another. Those concepts improve the structure of a monolith because they establish boundaries and responsibility before a team chooses deployment technology.
That boundary discipline is a durable contribution of the microservices period. Engineering teams learned to define responsibilities carefully, assign ownership, and make interactions explicit. A network can enforce those decisions between processes, while a well-designed application can enforce them internally until separation provides a concrete operational benefit.
Know what should still make you split the system
The limits of a modular monolith follow from the runtime constraints that AI leaves intact. Process isolation remains valuable when one subsystem must be prevented from taking down another, and component-specific scaling remains valuable when one workload needs materially different capacity. Whole-application deployment also remains a genuine constraint even when CI and AI-assisted testing make that deployment cheaper and faster.
Those limits make “which problems require network/process separation?” the useful architecture question. A concrete technical requirement can answer it: independent scaling, failure isolation, or another need that depends on separate execution and deployment. A genuine organizational requirement can answer it too when teams need service-level independence so they can own and release parts of a sufficiently large system separately.
Because those requirements vary by organization and workload, no useful team-size cutoff can substitute for them. Small and mid-sized teams can otherwise adopt operational practices developed for organizations large enough to benefit from many independent services, while most applications never reach the scale where component-level scaling becomes decisive. Team structure, workload, reliability needs, and deployment pressure determine whether independence repays its cost, so the requirement itself is the relevant trigger.
Microservices remain a valid answer to problems that require distribution. In 2026, AI can carry more of the cognitive and mechanical burden of a unified codebase, while modular design can provide strong boundaries before separate processes become necessary. Distribution should begin when a named technical or organizational requirement makes independence worth operating.
Key takeaways for decision-makers
- Make modular monoliths the default: AI makes large codebases easier to navigate, modify, test, and migrate, extending the useful life of a unified application. Architecture owners can keep clear module boundaries and delay distribution until a specific requirement justifies it.
- Separate human costs from runtime constraints: AI reduces cognitive and mechanical work that historically strengthened the case for microservices. CTOs evaluating a split should focus on requirements such as independent scaling, fault isolation, and autonomous deployment that AI cannot remove.
- Account for distributed-system costs: Microservices introduce network failures, retries, eventual consistency, distributed tracing, version compatibility, and additional infrastructure. Engineering organizations gain value from those costs when service boundaries create meaningful operational or organizational independence.
- Design modules for future extraction: Clear interfaces, bounded responsibilities, and controlled data access preserve architectural discipline inside a monolith. Platform and application teams can use these boundaries to make later service extraction simpler when runtime requirements emerge.
- Split services when independence has measurable value: Workloads with distinct scaling, reliability, or deployment needs remain strong candidates for separate services. Architecture decisions should tie each proposed service boundary to a concrete technical or organizational requirement rather than anticipated growth alone.
A project in mind?
Schedule a 30-minute meeting with us.
Senior experts helping you move faster across product, engineering, cloud & AI.


