Moving From Direct API Calls to an Event Queue Without Losing Messages
Most teams don't adopt event-driven architecture all at once, they move one workflow at a time from a direct service call to a queue, usually because a chain of synchronous calls started failing together during an outage. The order that migration happens in matters more than the broker you pick.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Which workflow should you move to an event queue first?
Choose a single workflow where a direct call is already causing pain, a notification send that shouldn't block a checkout, an export job that times out the request that triggered it, and move only that one. Trying to redesign the whole system around events at once means every mistake in your delivery guarantees or consumer design shows up everywhere simultaneously.
A single migrated workflow also gives the team a concrete pattern to copy for the next one, instead of a set of abstract principles nobody's actually used yet.
Decide What Delivery Guarantee This Event Actually Needs
At-least-once delivery means a consumer might see the same event twice, and that's fine for most workflows as long as the consumer handles it. Exactly-once delivery is expensive to guarantee end to end and is rarely worth building for a workflow that could just be made idempotent instead.
Write down, before you build anything, what happens if this specific event is processed twice. If the answer is "nothing bad," you don't need exactly-once semantics, you need an idempotent consumer, which is a much simpler thing to build correctly.
Build the Dead Letter Queue Before You Need It
A message that fails processing repeatedly needs somewhere to go that isn't silently dropped and isn't retried forever, blocking every message behind it. Set up a dead letter queue as part of the initial rollout, not as a follow-up task after the first incident where messages went missing.
Alert on messages landing in the dead letter queue, not just on consumer error rates. A consumer can have a healthy error rate while quietly losing a small, consistent trickle of messages that all end up dead-lettered and unexamined.
How do you make event consumers idempotent?
A consumer that processes an event twice and charges a customer twice, or sends a duplicate notification, is a bigger problem than a slow consumer. Give every event a unique identifier and have the consumer check whether it's already handled that identifier before doing anything with side effects.
This check needs to survive a consumer restart, so store it somewhere durable, not in memory. It's a small amount of extra code per consumer and it's what actually makes at-least-once delivery safe to use.
Watch Queue Depth, Not Just Consumer Error Rates
A consumer that's healthy but too slow for incoming volume produces no errors at all, just a growing backlog that eventually becomes stale data or a timeout somewhere downstream. Queue depth and the age of the oldest unprocessed message catch this in a way error rate dashboards don't.
Set an alert threshold on both, tuned to how stale a message is allowed to get before it's actually a problem for that specific workflow. A backlog that's fine for an email queue is not fine for a payment webhook.
Keep the Old Direct Call as a Fallback During Rollout
Run the event-driven path alongside the old direct call for a defined period rather than cutting over in one deploy, so a broker outage or a consumer bug doesn't take down the workflow entirely while the team is still learning how the new path behaves under real load.
Remove the fallback once the event path has run cleanly through a full cycle of your typical traffic pattern, not on a fixed calendar date. A workflow that only spikes once a month needs to see that spike before the fallback comes out.
Decide How Consumers Handle a Schema Change to the Event
An event's shape changes eventually, a field gets added, renamed, or removed, and a consumer built to expect the old shape will either error out or silently misread the new one. Version the event payload from the start, even if it feels premature, so a producer can add fields without breaking every consumer reading the queue.
Deploy consumer changes before producer changes when a field is being removed or renamed, so the consumer can handle both the old and new shape during the rollout window. Doing it the other way around guarantees a window where the consumer is reading events it doesn't understand yet.
A safe migration of one workflow runs in this order:
- Choose one workflow where a direct call already causes pain, such as a notification that shouldn't block checkout.
- Decide the delivery guarantee the event needs; at-least-once is usually fine when the consumer handles duplicates.
- Set up a dead letter queue before launch so failing messages are neither dropped nor retried forever.
- Give each event a unique identifier and have the consumer check whether it has already handled it.
- Watch queue depth and the age of the oldest message, and keep the old direct call as a fallback during rollout.
- Version the event payload from the start so a schema change doesn't break existing consumers.
What Good Looks Like
A sound move to event-driven architecture migrates one workflow at a time, picks a delivery guarantee deliberately instead of by default, and builds dead letter handling and idempotent consumers before the first real failure exposes their absence.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Track the rollout stages, the pilot workflow, the fallback removal date, the next workflow to migrate, as a ClickUp board so the migration doesn't stall after the first success.
Write up the consumer pattern, idempotency check and dead letter handling included, in Trainual once it's proven, so the next engineer copies a working template instead of starting from scratch.
Frequently Asked Questions
Do we need a full message broker, or is a simple job queue enough?
For a single workflow with one or two consumers, a simple job queue is often enough and is much easier to operate. A broker with topics, consumer groups, and replay earns its complexity once you have several independent services consuming the same event.
How do we handle message ordering if it matters for this workflow?
Most brokers only guarantee order within a single partition or queue, keyed by something like a customer or order ID. If order matters, key your events on that identifier so related events land on the same partition, rather than trying to enforce global ordering.
What's the biggest mistake teams make migrating their first workflow to events?
Skipping the dead letter queue and idempotency work because the demo works fine without them. Both only matter once something fails or a message gets redelivered, which won't show up in testing but will show up in production within the first few weeks.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Decoupling Services With Events Without Losing Traceability
A worked example of decoupling two services with an event queue, and the specific traceability and ordering problems that show up once you do.
Event-Driven Architecture: The Questions to Answer Before You Adopt It
Message queues decouple services but trade synchronous simplicity for new failure modes. Here are the questions worth answering before you commit.
Moving From Request-Response to Event-Driven Without a Rewrite
A worked example of introducing event-driven messaging into an existing request-response API one workflow at a time, without a full rewrite.
What Happens When a Message Queue Backs Up, Walked Through Start to Finish
A walkthrough of a message queue backlog building up in production, what caused it, and the specific changes that would have caught it sooner.
When Event-Driven Messaging Is Worth the Complexity
Where event-driven messaging genuinely earns its added complexity over direct calls, the debugging cost it adds, and a middle path that avoids both extremes.
The Message Queue Decision That Determines Your Failure Modes
Choosing between a queue and a stream for event-driven messaging sets your failure modes for years. What each actually guarantees, and where each breaks.