What GDPR's Right to Erasure Means for a Vector Index
Under GDPR's right to erasure, deleting someone's data means deleting the vectors embedded from it too, not only the source record in your system of record. Nothing about a vector database looks like it's storing personal data, so that step is easy to miss, and the deletion isn't complete without it.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Does deleting the source document delete the embedding?
Not automatically, in almost every setup. The embedding is a separate artifact stored in a separate system, and unless your deletion workflow explicitly calls the vector database to remove the corresponding vector, it persists after the source record is gone. This is the single most common gap in a RAG pipeline's privacy handling: the application layer's delete button works, but nothing tells the vector store.
Does an embedding itself count as personal data?
Treat it as though it does. An embedding derived from personal data can, in some cases, be partially reconstructed back toward the original text, and even without full reconstruction, it's a representation that's linked to an identifiable person through your own systems. Under this reading, the embedding carries the same erasure obligation as the source text it was derived from, not a lesser one. For anything that depends on your specific jurisdiction or data category, confirm the exact obligation with your attorney rather than relying on a general rule.
What about backups and snapshots that still hold the deleted vector?
An erasure request needs to reach backups and index snapshots too, not just the live collection, within whatever retention window your privacy policy commits to. This is worth planning for explicitly rather than discovering during your first real erasure request: know how far back your snapshots go, and have a process for either purging a specific vector from an old backup or accepting a bounded retention window as your actual erasure timeline.
Who's the data controller and who's the processor?
In most RAG setups, your company is the data controller and your vector database vendor is a processor acting on your instructions. That means the erasure obligation is ultimately yours to fulfill, and you need a data processing agreement with the vendor that specifically covers deletion requests, including how quickly they commit to actually purging data from their own systems, not just marking it as deleted in a way that's still recoverable.
What should the actual deletion workflow look like?
Maintain an explicit mapping from each source record's identifier to the vector IDs derived from it, since without that mapping, finding every vector tied to a person requires a search rather than a direct lookup. When an erasure request comes in, the workflow should delete the source record, delete every mapped vector, and confirm the deletion happened rather than assuming the delete call succeeded. Log the confirmation, since that log is what you'd show if the deletion is ever questioned.
A workable erasure runbook covers these steps in order:
- Look up every vector ID mapped to the person's source records, using the mapping you maintain rather than searching the index.
- Delete the source record and every mapped vector, including older vectors left behind by earlier ingestion runs of the same content.
- Confirm each deletion actually happened instead of assuming the delete call succeeded.
- Check backups and index snapshots, and either purge the vector there or record the bounded retention window that applies.
- Log the confirmation so you can show it if the deletion is ever questioned.
What if the same content was chunked and embedded more than once?
A document re-ingested after an edit, or processed by two different pipelines feeding the same collection, can leave more than one vector tied to the same source record. If your source-to-vector mapping only tracks the most recent ingestion, an erasure request can miss older, orphaned vectors from a prior run. Track every vector ever derived from a record, not just the current one, or run a periodic reconciliation pass that looks for vectors with no corresponding mapping entry and flags them for review.
For example, a support team edits a customer's record twice, and each edit re-ingests the document, leaving three vectors tied to it. The mapping only lists the newest one. When that customer asks for erasure, the workflow deletes the source record and one vector, reports success, and two orphaned vectors keep the person's information searchable. The fix is a mapping that adds every vector ID at each ingestion, plus a periodic reconciliation pass that flags vectors with no mapping entry, so an erasure request finds everything derived from the record and not just the latest run.
Does a cross-border transfer question apply to the embedding step too?
If your embedding provider processes data outside the region your compliance commitments require, that's a transfer question independent of where the vector database itself is hosted, and it's easy to overlook because the embedding call feels like a stateless computation rather than a data transfer. Confirm where the embedding provider actually processes requests, not just where your vector database lives, and reflect that in your data processing agreement and any transfer mechanism your legal team requires. Ask the same question of any reranking or generation model in the request path, since each hop is a separate processing location to account for.
What Good Looks Like
The privacy standard is an explicit mapping from source records to their derived vectors, a deletion workflow that reaches the index and its backups, and a confirmed, logged record that each erasure request was actually completed.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Useful for tracking that a documented deletion workflow and vendor data processing agreements are actually kept current, not just written once.
Serves the same purpose as Vanta here: ongoing evidence that your erasure process and vendor agreements stay in place over time.
Frequently Asked Questions
Is anonymizing a document before embedding it a safe alternative to deletion later?
It reduces risk but isn't a full substitute for a working erasure process, since anonymization before ingestion doesn't help with content that was already embedded before the anonymization step existed. It's a good practice going forward, alongside a real deletion workflow for what's already indexed, not instead of one.
Do we need to track which vectors came from which source record even for non-regulated data?
It's worth doing regardless of jurisdiction, since the same mapping that supports a GDPR erasure request also makes routine data cleanup, fixing a bad ingestion, removing an outdated document, far easier to do correctly. Building it only once you're legally required to is usually more expensive than building it from the start.
How fast does a vector need to be deleted after an erasure request?
This depends on the specific regulation and your own stated policy commitments, so confirm the exact timeline with your attorney. Whatever the deadline, make sure your workflow can actually meet it: a manual process that depends on someone remembering to check the vector database will eventually miss the window.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
The Data-Mapping Step Most GDPR Programs Skip
Why GDPR and data-privacy programs stall without a real data map, and a practical process for building one across your actual production systems.
A Practical Data Privacy Checklist for Engineering Teams With EU Users
The concrete engineering work behind data privacy compliance, from data mapping to deletion pipelines, and where to bring in a lawyer instead of guessing.
A Go-Live Checklist for Shipping a RAG Pipeline
The specific steps to check before a RAG pipeline goes live: index warm-up, model version pinning, a canary check, and a real rollback plan.
Mapping SOC 2 Controls to a RAG Pipeline's Real Components
SOC 2 auditors ask about access, change management, and vendors in the abstract. Here's what each control actually maps to in a RAG pipeline.
Handling GDPR Erasure Requests in a Streaming Pipeline
Answers to the privacy questions a real-time pipeline actually raises: erasure across replicated topics, data minimization, and cross-border transfer.
A Practical Data Privacy Checklist If You Have EU Customers
A practical checklist for companies serving EU customers: whether GDPR applies, where personal data actually lives, and what to check in a DPA.