Developer Productivity & Platform EngineeringPlaybook3 min readUpdated September 2026

Moving From Direct API Calls to an Event Queue Without Losing Messages

Most teams don't adopt event-driven architecture all at once, they move one workflow at a time from a direct service call to a queue, usually because a chain of synchronous calls started failing together during an outage. The order that migration happens in matters more than the broker you pick.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Which workflow should you move to an event queue first?

Choose a single workflow where a direct call is already causing pain, a notification send that shouldn't block a checkout, an export job that times out the request that triggered it, and move only that one. Trying to redesign the whole system around events at once means every mistake in your delivery guarantees or consumer design shows up everywhere simultaneously.

A single migrated workflow also gives the team a concrete pattern to copy for the next one, instead of a set of abstract principles nobody's actually used yet.

Decide What Delivery Guarantee This Event Actually Needs

At-least-once delivery means a consumer might see the same event twice, and that's fine for most workflows as long as the consumer handles it. Exactly-once delivery is expensive to guarantee end to end and is rarely worth building for a workflow that could just be made idempotent instead.

Write down, before you build anything, what happens if this specific event is processed twice. If the answer is "nothing bad," you don't need exactly-once semantics, you need an idempotent consumer, which is a much simpler thing to build correctly.

Build the Dead Letter Queue Before You Need It

A message that fails processing repeatedly needs somewhere to go that isn't silently dropped and isn't retried forever, blocking every message behind it. Set up a dead letter queue as part of the initial rollout, not as a follow-up task after the first incident where messages went missing.

Alert on messages landing in the dead letter queue, not just on consumer error rates. A consumer can have a healthy error rate while quietly losing a small, consistent trickle of messages that all end up dead-lettered and unexamined.

How do you make event consumers idempotent?

A consumer that processes an event twice and charges a customer twice, or sends a duplicate notification, is a bigger problem than a slow consumer. Give every event a unique identifier and have the consumer check whether it's already handled that identifier before doing anything with side effects.

This check needs to survive a consumer restart, so store it somewhere durable, not in memory. It's a small amount of extra code per consumer and it's what actually makes at-least-once delivery safe to use.

Watch Queue Depth, Not Just Consumer Error Rates

A consumer that's healthy but too slow for incoming volume produces no errors at all, just a growing backlog that eventually becomes stale data or a timeout somewhere downstream. Queue depth and the age of the oldest unprocessed message catch this in a way error rate dashboards don't.

Set an alert threshold on both, tuned to how stale a message is allowed to get before it's actually a problem for that specific workflow. A backlog that's fine for an email queue is not fine for a payment webhook.

Keep the Old Direct Call as a Fallback During Rollout

Run the event-driven path alongside the old direct call for a defined period rather than cutting over in one deploy, so a broker outage or a consumer bug doesn't take down the workflow entirely while the team is still learning how the new path behaves under real load.

Remove the fallback once the event path has run cleanly through a full cycle of your typical traffic pattern, not on a fixed calendar date. A workflow that only spikes once a month needs to see that spike before the fallback comes out.

Decide How Consumers Handle a Schema Change to the Event

An event's shape changes eventually, a field gets added, renamed, or removed, and a consumer built to expect the old shape will either error out or silently misread the new one. Version the event payload from the start, even if it feels premature, so a producer can add fields without breaking every consumer reading the queue.

Deploy consumer changes before producer changes when a field is being removed or renamed, so the consumer can handle both the old and new shape during the rollout window. Doing it the other way around guarantees a window where the consumer is reading events it doesn't understand yet.

A safe migration of one workflow runs in this order:

  1. Choose one workflow where a direct call already causes pain, such as a notification that shouldn't block checkout.
  2. Decide the delivery guarantee the event needs; at-least-once is usually fine when the consumer handles duplicates.
  3. Set up a dead letter queue before launch so failing messages are neither dropped nor retried forever.
  4. Give each event a unique identifier and have the consumer check whether it has already handled it.
  5. Watch queue depth and the age of the oldest message, and keep the old direct call as a fallback during rollout.
  6. Version the event payload from the start so a schema change doesn't break existing consumers.
Executive Capability Standard

What Good Looks Like

A sound move to event-driven architecture migrates one workflow at a time, picks a delivery guarantee deliberately instead of by default, and builds dead letter handling and idempotent consumers before the first real failure exposes their absence.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Map the synchronous calls in your system that are causing timeouts or coupling problems, and pick the single best candidate to move first.
2. Do Manually:Stand up a queue for that one workflow and run it alongside the existing direct call before removing the fallback.
3. Delegate:Have the consumer's owning engineer write the idempotency check and dead letter handling as part of the same change, not a follow-up.
4. Automate:Add alerting on queue depth and dead letter volume so a stalled consumer surfaces before it becomes a customer-facing incident.
5. Buy:Move to a managed broker once you're running enough queues that operating the broker itself has become its own maintenance burden.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Frequently Asked Questions

Do we need a full message broker, or is a simple job queue enough?

For a single workflow with one or two consumers, a simple job queue is often enough and is much easier to operate. A broker with topics, consumer groups, and replay earns its complexity once you have several independent services consuming the same event.

How do we handle message ordering if it matters for this workflow?

Most brokers only guarantee order within a single partition or queue, keyed by something like a customer or order ID. If order matters, key your events on that identifier so related events land on the same partition, rather than trying to enforce global ordering.

What's the biggest mistake teams make migrating their first workflow to events?

Skipping the dead letter queue and idempotency work because the demo works fine without them. Both only matter once something fails or a message gets redelivered, which won't show up in testing but will show up in production within the first few weeks.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides