Engineering Leadership & Technical HiringPlaybook3 min readUpdated September 2026

Event-Driven Architecture: The Questions to Answer Before You Adopt It

Event-driven architecture, where services communicate through a message queue instead of calling each other directly, gets pitched as an unambiguous upgrade: more decoupled, more resilient, more scalable. It's a real tradeoff, not a pure upgrade. You trade the simplicity of a synchronous call, where you know immediately whether something succeeded, for a set of new questions about ordering, duplication, and failure that a direct call never had to answer.

The questions below are the ones worth answering honestly before adopting message queuing broadly, not after you've already built several services around it.

What Problem Are You Actually Trying to Solve?

Event-driven architecture solves two specific problems well: decoupling a fast-changing producer from a slower or less reliable consumer, and smoothing a bursty workload so a downstream service processes at its own steady pace instead of being overwhelmed by traffic spikes. If neither of those describes your actual pain, and you're adopting queues because they seem like the modern default, you're taking on new operational complexity without a matching benefit.

Name the specific coupling problem you have today. "Service A calling Service B directly means B's downtime takes down A too" is a real, specific problem a queue can fix. "It feels more scalable" isn't specific enough to justify the tradeoff.

What Happens if a Message Gets Delivered Twice?

Most queuing systems guarantee at-least-once delivery, not exactly-once, which means your consumer will eventually see a duplicate message, whether from a retry after a timeout, a redeploy, or an infrastructure hiccup. If processing that message twice charges a customer twice or sends a duplicate email, that's a bug waiting to happen the first time your infrastructure has a bad day.

Design consumers to be idempotent, processing the same message twice produces the same result as processing it once, typically by tracking a message ID you've already handled. This is worth building before you go live, not after the first duplicate causes a real incident.

Does Ordering Actually Matter for This Event Type?

Some event streams need strict ordering, a sequence of account balance changes has to apply in order or the final balance is wrong. Others don't, a stream of independent analytics events can process in any order without breaking anything. Treating every stream as if it needs strict ordering adds real complexity and often real throughput cost for streams that never needed the guarantee in the first place.

Decide this per event type, not as a blanket policy. A single queue handling both order-sensitive and order-indifferent events is usually a sign the events should be split into separate streams with separate guarantees, rather than everything paying for the strictest requirement in the mix.

What Happens When a Message Can't Be Processed?

A message that fails processing, a malformed payload, a downstream dependency that's down, needs somewhere to go besides silently disappearing or retrying forever and blocking everything behind it. A dead letter queue, where failed messages land after a bounded number of retries, keeps a bad message from stalling the entire stream while still preserving it for investigation.

Someone needs to actually monitor that dead letter queue, though, or it becomes a quiet graveyard of failures nobody looks at until a customer notices something never happened. Treat a growing dead letter queue as an alert-worthy signal, not a place messages go to be forgotten.

Can You Actually Trace a Request Across the Queue?

A synchronous call gives you a single trace: request in, response out, one stack to debug when something breaks. An event that gets published, sits in a queue, and gets picked up by a consumer minutes later breaks that trace unless you deliberately propagate a correlation ID through the message and into whatever logging or tracing system you use downstream.

Build this before you need it during an incident, not while you're trying to figure out why a customer's order never made it through three services later. Losing the ability to trace a request end to end is one of the most common, and most avoidable, regressions teams hit when they move from synchronous calls to queues.

Before adopting queues broadly, confirm you have an answer for each of these:

  • The specific problem you're solving, such as decoupling a fast producer from a slower consumer or smoothing a bursty workload.
  • Idempotent consumers, because at-least-once delivery means duplicate messages will eventually arrive.
  • Which event streams need strict ordering and which can safely process in any order.
  • A dead letter queue with a bounded retry count, so failed messages neither vanish nor block everything behind them.
  • Trace context propagated through every message, so you can follow one request across the queue.
Executive Capability Standard

What Good Looks Like

Good event-driven architecture means every consumer is idempotent against duplicate delivery, ordering guarantees are decided per event type rather than assumed universally, failed messages land in a monitored dead letter queue, and a correlation ID traces a request across the queue.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Name the specific coupling or burst problem you're trying to solve before adopting queuing for a new service.
2. Do Manually:Manually trace one request through your queue end to end and confirm you can actually follow it across services.
3. Delegate:Assign an engineer ownership of idempotency and dead letter queue handling standards that every new consumer must follow.
4. Automate:Add automated alerting on dead letter queue growth and consumer lag so a stalled consumer surfaces before a customer notices.
5. Buy:Move to a managed queuing platform if your team is spending more time operating queue infrastructure than building on top of it.

How to Get Started

Frequently Asked Questions

Do we need Kafka, or is a simpler queue enough?

For most small teams, a simpler managed queue is enough and far easier to operate than running Kafka yourselves. Reach for Kafka specifically when you need very high throughput, long message retention for replay, or multiple independent consumers reading the same stream at different paces, not by default.

How do we handle a consumer that's falling behind the queue?

Find out whether the backlog is a temporary spike or a sustained gap. A short backlog that drains on its own is normal, but a steadily growing one means the consumer needs more capacity or faster processing logic. Alert on a backlog that keeps growing rather than on any backlog existing, since some backlog during bursts is expected.

Is event-driven architecture harder to debug than synchronous calls?

Yes, by default, since a single request can now span a queue and an asynchronous consumer instead of one traceable call stack. Correlation IDs propagated through every message close most of that gap, but only if you build that propagation in from the start rather than retrofitting it after a hard-to-debug incident.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides