Production RAG & Vector Data ArchitecturePlaybook3 min readUpdated September 2026

What GDPR's Right to Erasure Means for a Vector Index

Under GDPR's right to erasure, deleting someone's data means deleting the vectors embedded from it too, not only the source record in your system of record. Nothing about a vector database looks like it's storing personal data, so that step is easy to miss, and the deletion isn't complete without it.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Does deleting the source document delete the embedding?

Not automatically, in almost every setup. The embedding is a separate artifact stored in a separate system, and unless your deletion workflow explicitly calls the vector database to remove the corresponding vector, it persists after the source record is gone. This is the single most common gap in a RAG pipeline's privacy handling: the application layer's delete button works, but nothing tells the vector store.

Does an embedding itself count as personal data?

Treat it as though it does. An embedding derived from personal data can, in some cases, be partially reconstructed back toward the original text, and even without full reconstruction, it's a representation that's linked to an identifiable person through your own systems. Under this reading, the embedding carries the same erasure obligation as the source text it was derived from, not a lesser one. For anything that depends on your specific jurisdiction or data category, confirm the exact obligation with your attorney rather than relying on a general rule.

What about backups and snapshots that still hold the deleted vector?

An erasure request needs to reach backups and index snapshots too, not just the live collection, within whatever retention window your privacy policy commits to. This is worth planning for explicitly rather than discovering during your first real erasure request: know how far back your snapshots go, and have a process for either purging a specific vector from an old backup or accepting a bounded retention window as your actual erasure timeline.

Who's the data controller and who's the processor?

In most RAG setups, your company is the data controller and your vector database vendor is a processor acting on your instructions. That means the erasure obligation is ultimately yours to fulfill, and you need a data processing agreement with the vendor that specifically covers deletion requests, including how quickly they commit to actually purging data from their own systems, not just marking it as deleted in a way that's still recoverable.

What should the actual deletion workflow look like?

Maintain an explicit mapping from each source record's identifier to the vector IDs derived from it, since without that mapping, finding every vector tied to a person requires a search rather than a direct lookup. When an erasure request comes in, the workflow should delete the source record, delete every mapped vector, and confirm the deletion happened rather than assuming the delete call succeeded. Log the confirmation, since that log is what you'd show if the deletion is ever questioned.

A workable erasure runbook covers these steps in order:

  1. Look up every vector ID mapped to the person's source records, using the mapping you maintain rather than searching the index.
  2. Delete the source record and every mapped vector, including older vectors left behind by earlier ingestion runs of the same content.
  3. Confirm each deletion actually happened instead of assuming the delete call succeeded.
  4. Check backups and index snapshots, and either purge the vector there or record the bounded retention window that applies.
  5. Log the confirmation so you can show it if the deletion is ever questioned.

What if the same content was chunked and embedded more than once?

A document re-ingested after an edit, or processed by two different pipelines feeding the same collection, can leave more than one vector tied to the same source record. If your source-to-vector mapping only tracks the most recent ingestion, an erasure request can miss older, orphaned vectors from a prior run. Track every vector ever derived from a record, not just the current one, or run a periodic reconciliation pass that looks for vectors with no corresponding mapping entry and flags them for review.

For example, a support team edits a customer's record twice, and each edit re-ingests the document, leaving three vectors tied to it. The mapping only lists the newest one. When that customer asks for erasure, the workflow deletes the source record and one vector, reports success, and two orphaned vectors keep the person's information searchable. The fix is a mapping that adds every vector ID at each ingestion, plus a periodic reconciliation pass that flags vectors with no mapping entry, so an erasure request finds everything derived from the record and not just the latest run.

Does a cross-border transfer question apply to the embedding step too?

If your embedding provider processes data outside the region your compliance commitments require, that's a transfer question independent of where the vector database itself is hosted, and it's easy to overlook because the embedding call feels like a stateless computation rather than a data transfer. Confirm where the embedding provider actually processes requests, not just where your vector database lives, and reflect that in your data processing agreement and any transfer mechanism your legal team requires. Ask the same question of any reranking or generation model in the request path, since each hop is a separate processing location to account for.

Executive Capability Standard

What Good Looks Like

The privacy standard is an explicit mapping from source records to their derived vectors, a deletion workflow that reaches the index and its backups, and a confirmed, logged record that each erasure request was actually completed.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Trace one real document through your pipeline and confirm whether deleting it today actually removes its vector from the index.
2. Do Manually:Build the source-to-vector ID mapping for your highest-risk collection by hand, then manually process your next few erasure requests against it.
3. Delegate:Assign a specific engineer to own the deletion workflow end to end, including backups, rather than leaving it split across teams.
4. Automate:Wire vector deletion into your existing account or record deletion flow so it happens automatically instead of requiring a manual follow-up step.
5. Buy:Bring in privacy counsel to confirm your specific retention and deletion timelines once you're handling regulated personal data at scale.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Frequently Asked Questions

Is anonymizing a document before embedding it a safe alternative to deletion later?

It reduces risk but isn't a full substitute for a working erasure process, since anonymization before ingestion doesn't help with content that was already embedded before the anonymization step existed. It's a good practice going forward, alongside a real deletion workflow for what's already indexed, not instead of one.

Do we need to track which vectors came from which source record even for non-regulated data?

It's worth doing regardless of jurisdiction, since the same mapping that supports a GDPR erasure request also makes routine data cleanup, fixing a bad ingestion, removing an outdated document, far easier to do correctly. Building it only once you're legally required to is usually more expensive than building it from the start.

How fast does a vector need to be deleted after an erasure request?

This depends on the specific regulation and your own stated policy commitments, so confirm the exact timeline with your attorney. Whatever the deadline, make sure your workflow can actually meet it: a manual process that depends on someone remembering to check the vector database will eventually miss the window.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides