Production RAG & Vector Data ArchitecturePlaybook3 min readUpdated September 2026

Catching Retrieval API Schema Drift Before It Breaks Things

Contract testing catches retrieval API schema drift by publishing a schema that the retrieval service and its callers both test against on every build. Without it, each side can pass its own tests and still stop working together, because the actual contract between them was never written down anywhere both could check.

Here's how to close that gap with contract testing specifically, rather than hoping integration tests happen to catch it.

How do you define the contract as a schema both sides can check?

Tribal knowledge about what a retrieval response looks like, a Slack thread from six months ago, a comment in someone's code, isn't something a build can verify. Publish an actual schema, OpenAPI or JSON Schema both work, that describes every field, its type, and whether it's required. Both the provider and every consumer test against this same schema, which turns a hallway conversation about what changed into an artifact a computer can check on every build.

Keep the schema in a location both teams can see and version alongside the code, not buried in a wiki page that drifts out of sync with what's actually deployed.

Write consumer-driven tests, not just provider tests

A provider test confirms the service does what its own team thinks it should do. A consumer-driven contract test records what a specific consumer actually depends on, then runs that expectation against the provider's real output as part of the provider's own build. This flips the usual order: instead of the provider guessing what might break a caller, each caller states its actual expectations explicitly, and the provider's CI enforces them automatically before a change ships.

This catches something integration tests alone tend to miss: a field a consumer relies on that the provider's own team didn't realize was load-bearing, because nobody on the provider side wrote a test for it.

How do you catch a silent field-type change, not just a missing field?

A missing field usually throws an obvious error somewhere. A field that changes type quietly, a similarity score moving from a float to a string-formatted number, or gaining more decimal precision than a consumer's parser expects, can fail silently or produce subtly wrong behavior instead of an obvious crash. Contract tests that check types, not just field presence, catch this class of change specifically, which is exactly the kind of drift that's hardest to spot through manual review.

For example, a provider team changes the similarity score field from a number to a string so it can control formatting. Every provider test still passes because the field is present, and a consumer that sorts by score keeps running, now sorting text alphabetically and returning subtly wrong rankings. A contract test that checks types fails on the provider's next build and names the field and the consumer. The fix becomes a short conversation in a pull request instead of a confusing bug report from users weeks later.

Run contract tests on every deploy of either side

A contract test that only runs occasionally, or only on the provider's side, misses changes introduced by either party. Run the full suite of consumer expectations against the provider on every provider deploy, and run each consumer's own contract tests against a mock or a pinned version of the provider on every consumer deploy. This closes the loop from both directions instead of relying on one side to remember to check compatibility before shipping.

Version breaking changes explicitly

When a genuinely breaking change is necessary, keep a contract test suite scoped to each supported API version, not just the latest one. This way, an old client's assumptions stay verified against the version it actually calls, even as a new version rolls out alongside it, and you find out immediately if a deprecated version accidentally breaks before its planned retirement rather than discovering it from a confused caller.

Make contract test failures actionable, not just red

A failing contract test that just says "expectation mismatch" sends someone digging through diffs to understand what actually changed. Report the specific field, the expected type or value versus the actual one, and which consumer's expectation failed, directly in the test output. A contract test suite that's fast to diagnose gets fixed quickly; one that requires real investigative work to understand tends to get skipped or disabled the first time it's inconvenient, which quietly erodes the whole point of having it.

A working contract testing setup includes:

  • A published OpenAPI or JSON Schema describing every field, its type, and whether it's required, versioned alongside the code.
  • Consumer-driven tests that record what each caller depends on and run against the provider's real output in the provider's build.
  • Type checks as well as presence checks, so a float quietly becoming a string fails the build.
  • Runs on every deploy of either side, against a mock or pinned provider version for consumers.
  • A test suite per supported API version, so deprecated versions stay verified until retirement.
  • Failure output naming the field, the expected versus actual value, and which consumer's expectation broke.
Executive Capability Standard

What Good Looks Like

The contract testing standard is a published, versioned schema with consumer-driven tests run on every deploy from both the provider and each consumer, catching field and type drift before it reaches production.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Write down the retrieval API's actual response shape as it exists today, including anything undocumented that a consumer currently depends on.
2. Do Manually:Publish a basic schema and manually check a consumer's real usage against it once, before building any automated testing around it.
3. Delegate:Assign an owner on both the provider and at least one major consumer team to maintain their respective contract tests going forward.
4. Automate:Wire contract tests into CI on both sides so a breaking change fails a build instead of reaching a deployed environment.
5. Buy:Bring in a platform engineer to set up shared contract testing tooling once more than a couple of internal consumers depend on the same retrieval service.

How to Get Started

Frequently Asked Questions

Do we need contract tests if we only have one internal consumer today?

It's worth starting even with one consumer, since the discipline of writing down the actual contract catches drift early and the cost of adding a second consumer later drops significantly once the pattern already exists. Retrofitting contract tests after several consumers and years of undocumented assumptions have accumulated is a much bigger project.

How is a contract test different from a normal integration test?

An integration test typically runs both services together and checks that a specific scenario works end to end. A contract test checks the schema and expectations independently, often without running both services at once, which makes it faster and lets each team run it in their own CI without depending on the other service being deployed and reachable.

What happens when a consumer's expectations and the provider's roadmap genuinely conflict?

That's a real conversation to have explicitly, not something to resolve by quietly breaking the contract test. A failing consumer-driven test at that point is doing its job: surfacing the conflict early, in a build, rather than letting it reach production and become an incident instead of a planning discussion.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides