When Event-Driven Messaging Is Worth the Complexity
Event-driven architecture gets pitched as an upgrade from direct service calls, and sometimes it is. It also trades one kind of problem for another: instead of a service calling another service directly and getting an immediate answer, you get a message that's eventually processed, somewhere, by something, with a debugging story that's harder to follow when it goes wrong.
That tradeoff is worth it for some systems and not for others, and the difference usually comes down to whether you actually need the decoupling, not whether messaging sounds like the more sophisticated choice.
What You're Actually Buying With Events
The real benefit of an event-driven approach is decoupling: the service publishing an event doesn't need to know who's listening, how many listeners there are, or whether a listener is temporarily down. New consumers can be added without touching the publisher at all, and a struggling consumer can fall behind without taking the publisher down with it.
That's genuinely valuable when you have several independent systems that all care about the same event, an order being placed triggering inventory, billing, and notifications, each of which can succeed or fail independently without blocking the others.
What It Costs You in Return
A direct call fails immediately and visibly; a message can sit in a queue, get retried, land in a dead-letter queue, or process successfully but out of order relative to another related message, and none of that is visible from the code that published it. Debugging "why didn't this thing happen" becomes a search across a message broker's internals instead of reading a stack trace.
You also take on new failure modes that don't exist with direct calls: duplicate processing if a message gets redelivered, and ordering issues if your consumers assume events arrive in the sequence they were published, which message brokers don't always guarantee.
When Direct Calls Are Actually the Better Fit
If you have one service calling one other service, and the caller genuinely needs an immediate answer to proceed, a direct call with normal error handling and retries is simpler, easier to debug, and easier for a new engineer to reason about. Adding a message broker to that relationship adds real operational overhead, another system to run, monitor, and understand, without buying any decoupling benefit you actually needed.
This is the most common overcorrection: a system that's fundamentally a synchronous chain of two or three services gets rebuilt around events because messaging is seen as the more modern default, not because the actual traffic pattern calls for it.
For example, consider a signup screen that creates an account and must show the new profile immediately afterward. The caller needs an answer before it can continue, so a direct call with normal error handling and retries is the simpler fit. Routing that through a message broker would add a queue to monitor and a new way for the account to appear late, with no decoupling benefit. A welcome email, by contrast, doesn't block anything and is a fair candidate for an event. Judge each interaction by whether the caller needs the result to proceed.
A Middle Path: Events for Fan-Out, Calls for Everything Else
Many systems don't need to choose one pattern universally. Use events specifically where one action genuinely triggers multiple independent downstream effects that don't need to block the original request, and keep direct calls for the parts of your system that are a straightforward, synchronous chain with one clear next step.
This mixed approach avoids both failure modes: the operational fragility of forcing everything through a message broker, and the tight coupling of forcing every fan-out scenario through a growing chain of direct calls that all have to succeed for the original request to complete.
If You Do Go Event-Driven: The Non-Negotiables
Design every consumer to handle a message arriving more than once, since at-least-once delivery is the norm for most message brokers, and treating duplicate delivery as an edge case rather than the expected case is where a lot of event-driven bugs come from. Build observability into the pipeline itself, not just the endpoints, so you can actually see where a specific message is stuck rather than only knowing that something, somewhere, didn't happen.
Decide your ordering requirements explicitly per event type, rather than assuming order is preserved by default, and only pay the extra cost of strict ordering where a consumer genuinely can't tolerate events arriving out of sequence.
If you go event-driven, treat these as requirements:
- Design every consumer to handle the same message arriving more than once, since at-least-once delivery is the norm for most brokers.
- Build observability into the pipeline itself, so you can see where a specific message is stuck and not only that something didn't happen.
- Decide ordering requirements explicitly for each event type instead of assuming events arrive in the order they were published.
- Pay for strict ordering only where a consumer genuinely can't tolerate events arriving out of sequence.
What Good Looks Like
Good use of event-driven architecture means it's applied specifically where multiple independent systems need to react to the same event without blocking each other, with every consumer built to handle duplicate delivery, not adopted universally because it sounds more scalable.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Should we use event-driven messaging instead of direct service calls?
Only where you actually need the decoupling: multiple independent systems reacting to the same event, where each should succeed or fail on its own. For a simple chain of one service calling another for an immediate answer, direct calls with normal error handling are simpler, easier to debug, and don't add a message broker's operational overhead for no real benefit.
What makes event-driven systems harder to debug than direct calls?
A direct call fails immediately and visibly in the calling code. A message can be delayed, retried, or stuck in a dead-letter queue with no visible connection back to the code that published it, so debugging becomes searching a message broker's internals instead of reading a straightforward stack trace or error.
Can we mix event-driven and direct-call patterns in the same system?
Yes, and for most systems that's the better default. Use events where one action genuinely triggers multiple independent downstream effects, and keep direct calls for straightforward, synchronous chains with one clear next step. Forcing everything through one pattern universally usually costs more than it's worth in one direction or the other.
What should we handle first if we do build an event-driven system?
Design every consumer to handle receiving the same message more than once, since most message brokers guarantee at-least-once delivery, not exactly-once. Treating duplicate delivery as the expected case rather than a rare edge case avoids a large share of the bugs teams run into after adopting event-driven messaging.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Moving From Direct API Calls to an Event Queue Without Losing Messages
How to move one workflow from direct service calls to an event queue, covering delivery guarantees, dead letter queues, and idempotent consumers.
Verifying Devices Before They Touch Production, Not After
How to build device verification into a zero-trust rollout, what actually counts as a trust signal, and where teams stop checking too early.
Decoupling Services With Events Without Losing Traceability
A worked example of decoupling two services with an event queue, and the specific traceability and ordering problems that show up once you do.
Event-Driven Architecture: The Questions to Answer Before You Adopt It
Message queues decouple services but trade synchronous simplicity for new failure modes. Here are the questions worth answering before you commit.
What Happens When a Message Queue Backs Up, Walked Through Start to Finish
A walkthrough of a message queue backlog building up in production, what caused it, and the specific changes that would have caught it sooner.
A Production Deployment Checklist That Actually Catches Problems
A stage-by-stage deployment checklist for distributed systems, covering rollback readiness, dependency ordering, and the checks teams skip under pressure.