Model Context Protocol & Agentic ArchitecturePlaybook3 min readUpdated September 2026

The Questions to Ask Before You Add a Message Queue

Event-driven architecture gets adopted because it promises loose coupling and independent scaling, and it can genuinely deliver both. It also introduces a specific set of new problems, mostly around ordering, duplication, and debugging, that a direct request-response call never had to deal with. Here are the questions worth answering honestly before committing to it for a given piece of your system.

Answer these per interaction, not once for the whole architecture, since the right answer for one part of your system is often the wrong one for another.

Does this interaction actually need to be asynchronous?

Not every interaction between services benefits from going through a queue. If the caller genuinely needs to wait for a result before proceeding, a direct call is simpler to build, trace, and debug than an asynchronous round trip through a queue and a response channel. Reach for a queue when the caller can genuinely move on without waiting, or when you specifically need to buffer against a downstream service being temporarily slower or unavailable, not as a default architectural style.

What happens if the same message is delivered twice?

Most message queues offer at-least-once delivery, meaning your consumers will, eventually, receive a duplicate message, whether from a retry after a slow acknowledgment or an infrastructure hiccup. Design consumers to be idempotent, safe to process the same message more than once without a different outcome, from the start, rather than treating duplicate delivery as a rare edge case to handle later. It is not rare; it is a normal, expected part of how these systems behave.

Does the order messages arrive in actually matter?

If processing order matters for correctness, an update event arriving before the creation event it depends on, you need a queue and a consumer design that actually guarantees ordering for the specific keys where it matters, which usually means partitioning by that key rather than relying on the queue's default behavior. Confirm explicitly whether your use case needs this guarantee, because building and testing ordered delivery is more work than an unordered queue, and paying that cost where it isn't needed is wasted effort.

How will you know when a message has silently failed to process?

A consumer that throws away a message it can't process, without moving it somewhere visible, creates a silent data loss problem that's invisible until someone notices missing data downstream, sometimes weeks later. Set up a dead-letter queue or equivalent for messages that fail processing after retries, and alert on it actively, so a processing failure surfaces immediately instead of showing up as an unexplained gap discovered much later.

Can you actually trace a request across the whole event chain?

A request that flows through three services connected by direct calls is straightforward to trace with a standard request ID. The same logical flow through a queue, where a publisher and its eventual consumer might run seconds or minutes apart, needs deliberate tracing: a correlation ID propagated through every message, and tooling that can reconstruct the chain across that gap. Build this in from the start, because debugging an event-driven flow without it means piecing together timestamps and guesswork during an incident.

What's your plan for changing a message's shape later?

An event schema that changes without a version identifier breaks every consumer still expecting the old shape, and unlike a direct API where you at least control both ends of a single call, an event can have consumers you've lost track of entirely. Version every event schema explicitly from the first one you publish, even when there's currently only one consumer, since a second consumer, built by a different team, tends to show up sooner than expected once the pattern proves useful.

Who actually owns a queue once more than one team publishes to it?

A queue that starts with one publisher and one consumer often ends up with several of each once the pattern proves itself, and at that point an unowned queue becomes a place where a schema change from one team silently breaks a consumer built by another team that nobody remembered to notify. Assign explicit ownership of the queue and its schema once a second team starts depending on it, with a clear process for proposing and reviewing changes that affect other consumers.

Before you move a given interaction onto a queue, confirm that:

  • The caller can genuinely move on without waiting, or you need to buffer against a slow or unavailable downstream service.
  • Consumers are idempotent, so processing the same message twice does not change the outcome.
  • Ordering is guaranteed for the specific keys where correctness depends on it.
  • A dead-letter queue exists for messages that fail after retries, with active alerts on it.
  • A correlation ID travels through every message, so you can trace a request across the whole chain.
  • Event schemas are versioned from the first one, and one team explicitly owns the queue.
Executive Capability Standard

What Good Looks Like

A sound event-driven design uses a queue only where the interaction is genuinely asynchronous or needs decoupling, builds consumers to handle duplicate delivery safely, guarantees ordering only where correctness actually requires it, routes and alerts on failed messages instead of dropping them silently, and propagates a traceable correlation ID through every event.

Building The Capability (5-Stage Skill Ladder)

1. Learn:read through your current or proposed event flows and check each one against the six questions above before committing to the design
2. Do Manually:manually trace one event through its full chain, from publish to final consumer, to confirm you can actually reconstruct the flow during an incident
3. Delegate:assign a specific owner for schema versioning and dead-letter monitoring across your event infrastructure, not per individual consumer
4. Automate:automate alerting on dead-letter queue volume and on any consumer falling behind its queue's growth rate
5. Buy:a managed message queue or event streaming platform removes the operational burden of running the queue itself, though the idempotency, ordering, and tracing decisions above are still yours to make regardless of which platform you choose

How to Get Started

Frequently Asked Questions

Is event-driven architecture always more scalable than direct calls?

Not automatically. It can decouple services so they scale independently, but a poorly designed consumer can still become a bottleneck, and a queue itself needs to be sized and monitored like any other piece of infrastructure. The scalability benefit comes from the design, not from adopting a queue by itself.

How do we handle a message that keeps failing to process?

Route it to a dead-letter queue after a defined number of retries, and alert on anything landing there, rather than retrying indefinitely or silently dropping it. A message stuck in an infinite retry loop can also mask a real bug by hiding it behind repeated, seemingly normal-looking failures.

Do we need a message queue if we're not at a scale where performance is an issue yet?

Scale isn't the only reason to use one; decoupling deploy schedules or isolating a flaky downstream dependency can justify it even at moderate scale. But if neither of those applies, a direct call is simpler to build and debug, and simplicity has real value while the system is still small.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides