Modular Monolith or Microservices: A Decision Guide
Choose a modular monolith unless one part of your system genuinely needs to scale, deploy or fail independently, because microservices are a response to specific pain, not a sign of maturity. Small teams that split too early take on distributed systems complexity years before they have the headcount to operate it.
Can a modular monolith replace microservices?
The appeal of microservices is usually independent deployability and clear ownership boundaries, not literally running separate processes. A modular monolith, one codebase with enforced internal boundaries between modules, each with its own owner and a narrow interface, gets you most of that clarity without the operational tax of running, monitoring, and securing a dozen separate services. If your actual complaint is that the codebase feels tangled and nobody's sure what's safe to change, enforcing module boundaries inside a single deployable often fixes the real problem faster than a service split does.
When should you split out a microservice?
The strongest reason to extract a genuine microservice is that one part of your system needs to scale, deploy, or fail independently of the rest: a video processing pipeline with wildly different resource needs than your main API, or a component owned by a team that ships on a different cadence and keeps getting blocked by unrelated changes elsewhere in the codebase. Splitting along organizational convenience alone, without one of those forcing functions, tends to produce services that still deploy together in practice because they're too tightly coupled to release independently.
Every service boundary is also a new zero-trust boundary
Each service you extract needs its own authentication to its neighbors, its own certificate or token lifecycle, and its own place in your network policy, none of which existed when the same code lived inside one process calling a function directly. Teams that split services without budgeting for this tend to end up with either an unauthenticated internal network they quietly trust by default, which undermines a zero-trust posture, or a proliferation of ad hoc auth schemes between services that nobody fully documented. Count this cost explicitly before splitting, not after the first internal security review flags it.
Deploy frequency is the metric that actually tells you if it's working
Independent, frequent releases are the entire point of decoupling services, so track deploy frequency per component after a split, not just once. DORA's research groups engineering teams by deployment frequency, and the spread between clusters is wide: the strongest teams release on demand while the weakest can go as long as 180 days between releases1. If your services are technically separate but still get released together in lockstep because of shared dependencies or a shared release process, you've paid the operational cost of microservices without collecting the benefit, and that gap between measured deploy frequency and the promise of independence is worth reviewing a few months after any split.
Watch for a distributed monolith, the worst of both architectures
A distributed monolith happens when services are split at the process level but still share a database, deploy in lockstep, or call each other synchronously in long chains that break together. It combines the network latency, partial failure modes, and operational overhead of microservices with the tight coupling of a monolith, without the benefit of either. If a single change routinely requires coordinated deploys across three or four services, that's usually the sign the boundary was drawn in the wrong place, not that you need a fifth service.
Decide per boundary, not once for the whole system
Few real systems are cleanly one architecture end to end. It's common and reasonable to run a modular monolith for the bulk of your product logic while extracting one or two genuine microservices for the pieces that have a real independent scaling or ownership need. Make the call boundary by boundary against the criteria above, rather than picking an architecture as a company wide identity and forcing every new feature to fit it.
Questions to answer for each candidate boundary:
- Does this component need to scale, deploy or fail independently, or have very different resource needs from the main API?
- Is a team that ships on a different cadence repeatedly blocked by unrelated changes elsewhere in the codebase?
- Have you budgeted authentication, certificate or token lifecycle, and network policy for the new service boundary?
- Will the new service avoid sharing a database or deploying in lockstep, which would create a distributed monolith?
- Will you track deploy frequency for this component after the split to confirm the change helped?
What Good Looks Like
A good architectural boundary is one that lets a component deploy, scale, or fail independently, with its own authenticated identity, and you can point to the specific forcing function that justified splitting it out.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Is a modular monolith actually easier to maintain than microservices for a small team?
Usually yes, mainly because it removes the operational overhead of running, monitoring, and securing multiple independently deployed services. A well organized monolith with enforced module boundaries gives most small teams clearer ownership without that added cost.
How do we know when it's actually time to extract a microservice?
Look for a genuine forcing function: a component that needs to scale independently, has wildly different resource needs, or is owned by a team whose release cadence keeps getting blocked by unrelated code. Extracting a service for organizational tidiness alone, without one of those reasons, rarely pays off.
What's the biggest security risk that comes with splitting a monolith into services?
Trusting the internal network by default instead of authenticating every service to service call. What used to be an in-process function call now crosses a network boundary, and that boundary needs its own identity and access controls, not an assumption that anything inside your infrastructure is automatically trusted.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.
Related Guides
Continuous Device Verification for a Zero-Trust API
How continuous device and identity verification actually works in a zero-trust architecture, and where to draw the line for a small engineering team.
Rolling Out Zero Trust in Production Without a Broad Outage
A checklist for rolling out stricter API authentication and authorization in production, and the pitfalls that turn a rollout into an incident.
How to Audit Whether Your APIs Actually Enforce Zero Trust
A step-by-step method for testing whether your APIs enforce zero trust in practice, not just on paper, and what to do with what you find.
Monolith or Microservices: How to Tell Which One You Actually Need
How to decide between a monolith and microservices based on your team size and deploy needs, not on which one sounds more modern.
The Real Latency Cost of Zero Trust, and How to Measure It
How to find out how much latency your zero trust controls actually add, which checks are worth the cost, and which ones you can move off the hot path.
Keeping Auth Checks Fast as Your API Traffic Grows
A worked example for keeping zero trust authorization checks fast as request volume grows, and where teams usually add latency without noticing.