Moving From Request-Response to Event-Driven Without a Rewrite
Event-driven architecture usually gets proposed as an all or nothing rewrite, which is exactly why most attempts stall out halfway through. The more workable path is picking one workflow that's genuinely a poor fit for synchronous request-response, moving just that one to a message queue, and letting the rest of the system stay exactly as it is until there's a real reason to change it.
Which workflows should move to events first?
The clearest candidates for an event-driven approach are workflows where the client triggering the action doesn't need to wait for it to finish: sending a welcome email after signup, generating a report, processing an uploaded file. If your current implementation makes the caller wait synchronously for one of these to complete, that's added latency and a fragile dependency for no real benefit, since nothing about the interaction actually requires an immediate response.
Publish an event, keep the old endpoint working during the transition
Introduce a message queue for the chosen workflow by having the existing endpoint publish an event and return immediately, with a separate consumer picking up the event and doing the actual work asynchronously. Keep the old synchronous code path available behind a flag during the transition, so you can compare behavior and roll back cleanly if the new path misbehaves under real traffic, rather than committing to the new path before you've seen it handle production load.
For example, suppose signup currently sends a welcome email inside the request, so a slow mail provider makes signup slow. The endpoint can instead publish a signup event and return right away, while a consumer sends the email. Keep the old synchronous path behind a flag, and compare results on real traffic for a while. The common mistake is removing the old path on the same day the new one ships. If the consumer misbehaves, flipping the flag back should restore the old behavior immediately, with no emergency deploy.
What should happen when a message fails to process?
A message that fails processing needs a defined path: a retry with backoff, a dead letter queue for messages that fail repeatedly, and an alert so someone notices a stuck queue rather than discovering it days later when a customer asks why their report never arrived. Write this behavior down and test it deliberately, by forcing a failure in a test environment, before the first real failure in production forces you to improvise the policy under pressure. Decide how many retries is too many and what happens next, since a message that retries forever against a permanently broken dependency is its own kind of stuck queue.
Message queues are a new zero-trust boundary too
A producer publishing to a queue and a consumer reading from it are two separate identities that need their own authentication and authorization, the same as any other service to service call, not an implicit trust relationship because they're both internal. Decide who can publish to which topics and who can consume from them, and log that access the same way you'd log an API call, since an unauthenticated or overly permissive queue is a quiet way to reintroduce the exact trust assumptions a zero-trust architecture is meant to remove.
Watch for ordering and duplicate delivery assumptions that don't hold
Most message queues either don't guarantee strict ordering or can deliver the same message more than once under specific failure conditions, and code written assuming synchronous, exactly once behavior will misbehave in ways that are hard to reproduce later. Make consumers idempotent, safe to process the same message twice without a bad side effect, from the start, rather than treating it as an edge case to handle later once a duplicate delivery has already caused a real problem. A simple, durable record of which message IDs have already been processed is usually enough to make this safe without redesigning the whole consumer.
Expand to the next workflow only once this one is boring
Once the first migrated workflow has run reliably through a normal traffic cycle, including a failure or two that the retry and dead letter handling caught cleanly, that's the signal to consider moving the next workflow, not a fixed timeline. Treat each migration as its own small project with its own verification, rather than declaring the whole system event-driven after the first success and rushing the rest. A system with a handful of well understood asynchronous workflows and the rest synchronous is a perfectly reasonable stopping point, not a compromise you need to apologize for.
The migration path, one workflow at a time:
- Pick one workflow where the caller does not need an immediate answer, such as a welcome email or report generation.
- Have the existing endpoint publish an event and return, with a separate consumer doing the work, and keep the old path behind a flag.
- Define retries with backoff, a dead letter queue and an alert for stuck messages, and test a forced failure.
- Give producers and consumers their own authentication and authorization, and log that access.
- Make consumers idempotent, then move the next workflow only once this one has become boring.
What Good Looks Like
A good event-driven migration moves one genuinely asynchronous workflow at a time, with defined retry and dead letter handling, idempotent consumers, and the same authentication between producer and consumer as any other internal service call.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How do we know if a workflow is actually a good fit for event-driven messaging?
Look for cases where the caller doesn't need an immediate result: background processing, notifications, report generation. If the caller genuinely needs a synchronous answer to continue, forcing that into an asynchronous event adds complexity without a real benefit.
What's the most common mistake teams make moving their first workflow to a message queue?
Not planning for duplicate delivery and message failure up front. Consumers written assuming a message arrives exactly once, in order, tend to break in ways that are hard to reproduce once real failure conditions and retries start happening in production.
Do message queues need the same authentication as our regular API endpoints?
Yes. A producer and consumer are separate identities making a service to service call, and that connection needs the same authentication, authorization, and logging as any other internal API call, not an implicit trust because the queue lives inside your own infrastructure.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Continuous Device Verification for a Zero-Trust API
How continuous device and identity verification actually works in a zero-trust architecture, and where to draw the line for a small engineering team.
Moving From Direct API Calls to an Event Queue Without Losing Messages
How to move one workflow from direct service calls to an event queue, covering delivery guarantees, dead letter queues, and idempotent consumers.
Rolling Out Zero Trust in Production Without a Broad Outage
A checklist for rolling out stricter API authentication and authorization in production, and the pitfalls that turn a rollout into an incident.
How to Audit Whether Your APIs Actually Enforce Zero Trust
A step-by-step method for testing whether your APIs enforce zero trust in practice, not just on paper, and what to do with what you find.
Decoupling Services With Events Without Losing Traceability
A worked example of decoupling two services with an event queue, and the specific traceability and ordering problems that show up once you do.
Event-Driven Architecture: The Questions to Answer Before You Adopt It
Message queues decouple services but trade synchronous simplicity for new failure modes. Here are the questions worth answering before you commit.