Production RAG & Vector Data ArchitecturePlaybook3 min readUpdated September 2026

Should Your RAG Pipeline Be One Service or Four?

A RAG pipeline naturally has four stages: ingestion, embedding, retrieval, and generation. Whether those live in one deployable service or four independent ones isn't a question with a universal right answer, and teams that copy whichever pattern a blog post recommends usually end up fighting the tradeoffs of a structure that doesn't match their actual constraints.

The decision comes down to how independently each stage needs to scale, deploy, and fail, not a general preference for one architectural style.

Why start with one service for a RAG pipeline?

A single service handling ingestion through generation is simpler to develop, test, and deploy, especially early on: one codebase, one deployment pipeline, no network calls between your own stages, and no distributed tracing needed to follow a request through your own code. For a team still figuring out its chunking strategy, its retrieval quality, and its prompt design, this simplicity has real value, since the architecture isn't the thing you're trying to get right yet.

When do separate services help a RAG pipeline?

Ingestion runs in batches and is compute-heavy in bursts; retrieval needs to be fast and available continuously; generation is often the most expensive and highest-latency step. Bundled into one service, you scale all of them together even though their load patterns don't match, and a spike in ingestion load can degrade retrieval latency for users actively querying the system. Split into services, each one scales to its own load pattern, and a failure in ingestion doesn't take down retrieval for existing users.

Watch for the split that adds latency without adding value

Splitting retrieval and generation into separate services adds a network hop between two steps that are almost always called together, on the same request, for the same user. If that split doesn't come with an independent scaling or deployment need, it's pure latency overhead: you've paid the cost of a network call without the benefit that justifies it. Split along the boundary where independent scaling or deployment actually matters, not along every conceptual stage just because you can.

A middle path: split ingestion, keep query time together

A common, effective pattern splits ingestion into its own service, since its load pattern and failure modes are genuinely different from the rest, while keeping retrieval and generation together in a single query-time service, since those two are always called sequentially for the same request anyway. This gets the biggest independent-scaling win, ingestion no longer competes with live queries for resources, without paying a network hop for the retrieval-to-generation handoff that happens on every single request.

Decide with your actual load pattern, not a default preference

Look at your own traffic: how often does ingestion run relative to queries, how differently do ingestion and query-time load spike, and how often do you deploy changes to one stage without touching the others? If ingestion and retrieval never spike independently and you deploy the whole pipeline together anyway, a split architecture is buying you flexibility you're not using. If they do spike independently, the split earns its complexity.

Ask these questions about your own traffic before you split anything:

  • How often does ingestion run compared with query traffic, and do the two ever spike at different times?
  • Do you deploy a change to one stage, such as chunking, without touching retrieval or generation?
  • Would splitting retrieval from generation give you an independent scaling or deployment benefit, or only an extra network hop?
  • Can your team absorb separate deployments, on-call surfaces, and distributed tracing without slowing down product work?

Revisit the decision as load patterns change, not just once at launch

A monolith that made sense at launch, when ingestion and query volume were both small, can become the wrong structure once ingestion runs on a much larger schedule than query traffic implies. Revisit this architecture decision when your load patterns actually diverge, not preemptively based on a general belief that splitting services is always the more mature choice.

Weigh the operational cost of a split honestly

Separate services mean separate deployments, separate on-call surfaces, and a real distributed tracing requirement to debug a single user request across service boundaries. None of that is free, and a small team taking on that overhead before it's earned its keep often spends more engineering time maintaining the split than it would have spent living with a monolith's coarser scaling for another few months. Write the operational cost down next to the expected benefit before you commit, the same way you'd evaluate any other infrastructure investment, rather than treating the split as free just because the code itself is straightforward to write.

Executive Capability Standard

What Good Looks Like

Good architectural judgment for a RAG pipeline means the split between services, if any, follows where scaling, deployment, or failure patterns actually diverge, not a default preference for one style.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Look at your own ingestion and query-time load patterns before deciding anything, since the right structure depends on how much they actually diverge.
2. Do Manually:Track deploy frequency and load spikes per pipeline stage by hand for a month before deciding whether a split is worth the complexity.
3. Delegate:Have one senior engineer make and document this architecture decision explicitly, with the load-pattern reasoning behind it, instead of it emerging implicitly from whoever builds each piece.
4. Automate:Once split, automate independent scaling policies per service so the split actually delivers the resource-efficiency benefit it's meant to.
5. Buy:Use your platform's built-in autoscaling per service rather than building custom scaling logic once you've decided a split is worth it.

How to Get Started

Frequently Asked Questions

Should ingestion and query-time retrieval always be separate services?

Not always, but it's the split that earns its complexity most often, since ingestion's load pattern, bursty and batch, is usually genuinely different from query-time retrieval's, continuous and latency-sensitive. Split when those patterns actually diverge in your traffic, not by default.

Does splitting retrieval and generation into separate services help performance?

Usually not, since they're almost always called together on the same request. A split there adds a network hop without an independent scaling or deployment benefit to justify it, unless you have a specific reason those two stages need to scale separately.

When should a RAG pipeline move from a monolith to separate services?

When your actual load patterns diverge: ingestion and query traffic spike independently, or you need to deploy changes to one stage without touching the others. If that's not happening yet, a monolith is simpler and the split can wait until it is.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides