Should Document Ingestion for RAG Be Synchronous or Event-Driven?
Document ingestion for RAG should be synchronous when volume is low and one reliable source feeds the corpus, and event-driven once volume, source diversity, or coupling to the embedding step starts causing problems. Synchronous ingestion embeds a document in the same request that changes it; event-driven ingestion queues an event for a separate worker.
The right choice depends on your ingestion volume, how many source systems feed your corpus, and how much latency between a document changing and it becoming searchable is acceptable.
When does synchronous RAG ingestion become a problem?
Embedding a document synchronously, as part of the request that creates or updates it, means the caller waits for that embedding call to finish, and if the embedding provider is slow or briefly unavailable, that latency or failure propagates directly to whatever created the document in the first place. This coupling is fine at low volume with a reliable embedding provider, and it becomes a real liability the moment either the embedding step gets slower or the source system's own reliability shouldn't depend on your RAG pipeline's health.
Event-driven ingestion decouples the source from the pipeline
Dropping a change event onto a queue and letting a separate worker handle embedding and indexing means the source system's own request finishes immediately, regardless of how long embedding takes or whether the vector database is briefly unavailable. The tradeoff is a delay between a document changing and it becoming searchable, and a new piece of infrastructure, the queue itself, that has to be operated, monitored, and have its own failure modes accounted for.
Multiple source systems make the event-driven case stronger
If your corpus draws from several source systems, a support ticket system, a document store, a wiki, each with its own change frequency and reliability profile, a shared queue gives you one consistent ingestion pattern instead of custom synchronous integration code per source. It also means an outage in the embedding pipeline doesn't cascade back to every source system simultaneously; events simply queue up and process once the pipeline recovers.
How do you handle duplicate and out-of-order ingestion events?
A queue-based system needs an answer for what happens if the same document's update event arrives twice, or if two updates to the same document arrive out of order. Design ingestion to be idempotent, so processing the same event twice produces the same end state, and use a document version or timestamp to detect and discard an out-of-order update rather than letting a stale event silently overwrite a newer one.
Monitor queue depth and processing lag as a first-class signal
The event-driven pattern trades immediate consistency for resilience, but only if you're watching for the failure mode it introduces: a growing backlog of unprocessed events, which means your index is falling behind what's actually true in the source systems. Alert on queue depth and processing lag directly, since without that visibility, a stalled ingestion worker can quietly leave your search results stale for hours before anyone notices.
Pick based on volume and coupling tolerance, not a default preference
Low document volume, a single reliable source, and a team still building the pipeline favor synchronous ingestion for its simplicity. Higher volume, multiple sources, or a source system that genuinely can't tolerate coupling to your RAG pipeline's latency favor an event-driven approach. Neither is the objectively better architecture; the right one depends on constraints that are worth stating explicitly rather than assuming.
It's also fine to run both at once, deliberately: synchronous for a low-volume, latency-sensitive source, event-driven for a bulk source that updates in large batches, as long as each pattern's tradeoffs are understood for the specific source it's applied to, rather than the whole pipeline being forced into one uniform pattern for consistency's sake alone. Write down which pattern applies to which source and why, so the mixed approach reads as a deliberate design decision to the next engineer who touches it, not as leftover inconsistency from two different eras of the codebase.
Check these points before moving ingestion onto a queue:
- Is ingestion volume or the number of source systems growing enough that synchronous embedding is slowing down or destabilizing the systems that create documents?
- Can the source system's own reliability be decoupled from your embedding provider's slowness or outages?
- Is ingestion idempotent, so processing the same event twice ends in the same state, with versions or timestamps to discard stale updates?
- Are queue depth and processing lag monitored and alerted on, so a stalled worker cannot quietly leave the index out of date?
- Is the delay between a document changing and becoming searchable acceptable for your users?
What Good Looks Like
Good ingestion architecture for a RAG pipeline means the synchronous-versus-event-driven choice follows your actual volume and source coupling constraints, and if you're event-driven, queue depth and processing lag are actively monitored.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
When should RAG document ingestion move from synchronous to event-driven?
When ingestion volume grows, you add multiple source systems, or a source system's own reliability shouldn't depend on your embedding pipeline's health. At low volume with one reliable source, synchronous ingestion is simpler and the added complexity of a queue isn't yet worth it.
What's the biggest risk in event-driven RAG ingestion?
A stalled or backlogged queue that quietly leaves your search index out of date, with nobody noticing until someone asks why a recently updated document isn't showing up in results. Alert on queue depth and processing lag directly rather than assuming the pipeline is keeping up.
Does event-driven ingestion need to handle duplicate or out-of-order events?
Yes. Design ingestion to be idempotent so processing the same event twice produces the same result, and use a document version or timestamp to discard an out-of-order update rather than letting a stale event overwrite a newer one.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
A Go-Live Checklist for Shipping a RAG Pipeline
The specific steps to check before a RAG pipeline goes live: index warm-up, model version pinning, a canary check, and a real rollback plan.
Moving From Direct API Calls to an Event Queue Without Losing Messages
How to move one workflow from direct service calls to an event queue, covering delivery guarantees, dead letter queues, and idempotent consumers.
Decoupling Services With Events Without Losing Traceability
A worked example of decoupling two services with an event queue, and the specific traceability and ordering problems that show up once you do.
Event-Driven Architecture: The Questions to Answer Before You Adopt It
Message queues decouple services but trade synchronous simplicity for new failure modes. Here are the questions worth answering before you commit.
Zero Trust for a RAG Pipeline Means No Service Gets a Free Pass
A decision framework for applying zero trust to a production RAG pipeline: verifying every service and user call, not just the ones at the edge.
What Happens When a Message Queue Backs Up, Walked Through Start to Finish
A walkthrough of a message queue backlog building up in production, what caused it, and the specific changes that would have caught it sooner.