Event-Driven Architecture: The Questions to Answer Before You Adopt It
Event-driven architecture, where services communicate through a message queue instead of calling each other directly, gets pitched as an unambiguous upgrade: more decoupled, more resilient, more scalable. It's a real tradeoff, not a pure upgrade. You trade the simplicity of a synchronous call, where you know immediately whether something succeeded, for a set of new questions about ordering, duplication, and failure that a direct call never had to answer.
The questions below are the ones worth answering honestly before adopting message queuing broadly, not after you've already built several services around it.
What Problem Are You Actually Trying to Solve?
Event-driven architecture solves two specific problems well: decoupling a fast-changing producer from a slower or less reliable consumer, and smoothing a bursty workload so a downstream service processes at its own steady pace instead of being overwhelmed by traffic spikes. If neither of those describes your actual pain, and you're adopting queues because they seem like the modern default, you're taking on new operational complexity without a matching benefit.
Name the specific coupling problem you have today. "Service A calling Service B directly means B's downtime takes down A too" is a real, specific problem a queue can fix. "It feels more scalable" isn't specific enough to justify the tradeoff.
What Happens if a Message Gets Delivered Twice?
Most queuing systems guarantee at-least-once delivery, not exactly-once, which means your consumer will eventually see a duplicate message, whether from a retry after a timeout, a redeploy, or an infrastructure hiccup. If processing that message twice charges a customer twice or sends a duplicate email, that's a bug waiting to happen the first time your infrastructure has a bad day.
Design consumers to be idempotent, processing the same message twice produces the same result as processing it once, typically by tracking a message ID you've already handled. This is worth building before you go live, not after the first duplicate causes a real incident.
Does Ordering Actually Matter for This Event Type?
Some event streams need strict ordering, a sequence of account balance changes has to apply in order or the final balance is wrong. Others don't, a stream of independent analytics events can process in any order without breaking anything. Treating every stream as if it needs strict ordering adds real complexity and often real throughput cost for streams that never needed the guarantee in the first place.
Decide this per event type, not as a blanket policy. A single queue handling both order-sensitive and order-indifferent events is usually a sign the events should be split into separate streams with separate guarantees, rather than everything paying for the strictest requirement in the mix.
What Happens When a Message Can't Be Processed?
A message that fails processing, a malformed payload, a downstream dependency that's down, needs somewhere to go besides silently disappearing or retrying forever and blocking everything behind it. A dead letter queue, where failed messages land after a bounded number of retries, keeps a bad message from stalling the entire stream while still preserving it for investigation.
Someone needs to actually monitor that dead letter queue, though, or it becomes a quiet graveyard of failures nobody looks at until a customer notices something never happened. Treat a growing dead letter queue as an alert-worthy signal, not a place messages go to be forgotten.
Can You Actually Trace a Request Across the Queue?
A synchronous call gives you a single trace: request in, response out, one stack to debug when something breaks. An event that gets published, sits in a queue, and gets picked up by a consumer minutes later breaks that trace unless you deliberately propagate a correlation ID through the message and into whatever logging or tracing system you use downstream.
Build this before you need it during an incident, not while you're trying to figure out why a customer's order never made it through three services later. Losing the ability to trace a request end to end is one of the most common, and most avoidable, regressions teams hit when they move from synchronous calls to queues.
Before adopting queues broadly, confirm you have an answer for each of these:
- The specific problem you're solving, such as decoupling a fast producer from a slower consumer or smoothing a bursty workload.
- Idempotent consumers, because at-least-once delivery means duplicate messages will eventually arrive.
- Which event streams need strict ordering and which can safely process in any order.
- A dead letter queue with a bounded retry count, so failed messages neither vanish nor block everything behind them.
- Trace context propagated through every message, so you can follow one request across the queue.
What Good Looks Like
Good event-driven architecture means every consumer is idempotent against duplicate delivery, ordering guarantees are decided per event type rather than assumed universally, failed messages land in a monitored dead letter queue, and a correlation ID traces a request across the queue.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Do we need Kafka, or is a simpler queue enough?
For most small teams, a simpler managed queue is enough and far easier to operate than running Kafka yourselves. Reach for Kafka specifically when you need very high throughput, long message retention for replay, or multiple independent consumers reading the same stream at different paces, not by default.
How do we handle a consumer that's falling behind the queue?
Find out whether the backlog is a temporary spike or a sustained gap. A short backlog that drains on its own is normal, but a steadily growing one means the consumer needs more capacity or faster processing logic. Alert on a backlog that keeps growing rather than on any backlog existing, since some backlog during bursts is expected.
Is event-driven architecture harder to debug than synchronous calls?
Yes, by default, since a single request can now span a queue and an asynchronous consumer instead of one traceable call stack. Correlation IDs propagated through every message close most of that gap, but only if you build that propagation in from the start rather than retrofitting it after a hard-to-debug incident.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Moving From Direct API Calls to an Event Queue Without Losing Messages
How to move one workflow from direct service calls to an event queue, covering delivery guarantees, dead letter queues, and idempotent consumers.
Decoupling Services With Events Without Losing Traceability
A worked example of decoupling two services with an event queue, and the specific traceability and ordering problems that show up once you do.
What Happens When a Message Queue Backs Up, Walked Through Start to Finish
A walkthrough of a message queue backlog building up in production, what caused it, and the specific changes that would have caught it sooner.
When Event-Driven Messaging Is Worth the Complexity
Where event-driven messaging genuinely earns its added complexity over direct calls, the debugging cost it adds, and a middle path that avoids both extremes.
Moving From Request-Response to Event-Driven Without a Rewrite
A worked example of introducing event-driven messaging into an existing request-response API one workflow at a time, without a full rewrite.
Should Document Ingestion for RAG Be Synchronous or Event-Driven?
A comparison of synchronous and event-driven ingestion patterns for a RAG pipeline, and when the added complexity of message queuing is worth it.