What SOC 2 Actually Expects From a Model-Serving Team
SOC 2 doesn't have a control called AI model serving, which is exactly why teams get confused about what applies. It has controls for access management, change management, vulnerability management, and vendor risk, and every one of them touches your inference stack whether the framework mentions models by name or not.
Map your model-serving specifics onto those existing controls instead of waiting for a framework that names them explicitly; none of the common ones will.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Where model serving maps onto standard SOC 2 controls
- Change management covers model deploys the same way it covers code deploys: document the process, require review, and keep a record of what changed and when.
- Access management covers who can deploy models, view raw prompts, and pull weights, the same categories from any role-based access review.
- Vulnerability management covers your inference server software and container images, not just your web application stack.
- Vendor risk management covers any third-party model API you call, the same as any other subprocessor that touches customer data.
If your auditor doesn't ask about models specifically, that's normal; answer the standard question with your model-serving specifics anyway.
Evidence an auditor will actually ask for
Expect requests for: your model deploy log, showing who approved each production change; your access list for the inference platform, with a recent review date; patch records for your serving software; and your vendor risk assessment for any external model provider, including what happens to data you send it.
None of this is exotic once you've built the underlying practice. Most of the audit friction comes from teams that have the practice informally, in someone's head or a chat thread, but nothing written down that produces evidence on request.
A simple way to prepare is to pick one production model change from the last quarter and trace it end to end as if you were the auditor. For example, ask who requested the change, who approved it, where that approval is recorded, and which access list shows the approver was allowed to approve it. If any answer lives only in a chat thread, move it into your ticketing or deploy system now. Repeat the exercise for one access review and one patch. Gaps found this way are ordinary process gaps, and closing them before the audit window opens costs far less than explaining them afterward.
Patch records: a control auditors check closely
Vulnerability and patch management gets scrutinized closely because it's easy to verify and easy to fake informally. Critical vulnerabilities patched within 15 days and high-severity ones within 30 is a defensible baseline drawn from federal guidance1; hold your serving stack to a written version of that SLA and keep a record showing you actually hit it, not just a policy stating that you should.
A policy with no evidence behind it is worse in an audit than no policy at all; it signals a gap between what you say and what you do.
What Vanta and Drata automate, and what still needs a person
Both platforms pull evidence automatically from connected systems: access lists, patch status, deploy logs where they're integrated. That removes most of the manual screenshot work that used to eat weeks before an audit.
What they don't do is decide whether your model-serving controls are actually adequate for your risk profile, or catch a control that technically exists but doesn't work the way it's documented. Treat the automated evidence as the paperwork layer and keep a person accountable for whether the underlying practice is real.
What changes once you're SOC 2 Type II instead of Type I
A Type I audit confirms your controls are designed correctly at a point in time. A Type II audit, the one most enterprise customers actually ask for, confirms those controls operated correctly over a period, usually six to twelve months.
For model serving specifically, that means your patch records, access reviews, and deploy approvals need to show a consistent pattern across the whole window, not just look correct on the day the auditor shows up. A single missed access review partway through is the kind of gap a Type II audit is designed to catch.
A pre-audit checklist specific to model serving
- Confirm your access list for the inference platform is current, not just for the app.
- Confirm your model deploy log actually captures approvals, not just deploy timestamps.
- Confirm your vendor risk assessment covers every model API you call, not just your main cloud provider.
- Confirm patch records exist for your inference server software specifically, separate from your general app stack.
Any gap here is cheap to fix weeks before an audit and expensive to explain during one.
What Good Looks Like
A compliance-ready model-serving setup maps deploys, access, patching, and vendor risk onto your existing SOC 2 controls, with evidence, deploy logs, access lists, patch records, ready on request rather than assembled after the auditor asks.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Frequently Asked Questions
Does SOC 2 have specific requirements for AI model serving?
No, and it doesn't need to. SOC 2's existing controls for change management, access management, vulnerability management, and vendor risk typically apply to your inference stack if it's in scope for your audit. Map your model deploys, access lists, patch records, and model API vendors onto those categories rather than waiting for AI-specific language that isn't coming.
What evidence should we prepare specifically for the model-serving parts of an audit?
Your model deploy log with approvals, a current access list for who can deploy or view raw prompts, patch records for your inference server software, and a vendor risk assessment covering any third-party model API you use. All four map onto standard controls an auditor already expects to see.
Can compliance automation platforms replace our own review of model-serving controls?
No. They automate evidence collection, access-review reminders, and patch tracking, which cuts the manual work significantly. They don't judge whether a control is actually adequate or catch one that's documented but not really followed. Keep a person accountable for that judgment even after you automate the paperwork.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Security patch remediation SLAs (CISA federal mandates, used as industry norm). CISA Binding Operational Directives 19-02 and 22-01 (CISA briefing hosted at NIST CSRC), 2022.
Related Guides
What a Real Security Audit of Model Serving Should Cover
A practical checklist for auditing AI model serving and inference: endpoint access, weight security, prompt logging, and patch timelines.
A Rollout Checklist for Swapping Models in Production
A rollout checklist for swapping AI models in production: evaluation gates, canary and shadow traffic, and fast rollback paths.
Governing Infrastructure as Code for Your GPU Fleet
Why GPU capacity managed through Terraform or Pulumi needs stricter review and drift detection than ordinary infrastructure, and how to set that up.
Why Inference Latency Creeps Up After You Ship
Where AI model-serving latency actually hides: tokenization, queueing, batching windows, and network hops, plus a worked example fix.
Zero Trust for Machines Calling Your Model Endpoints
Why internal services calling your inference endpoints still need identity checks, and where CrowdStrike-style posture checks and Tenable-style scanning fit.
Audit Logging for Model Serving: Build It or Buy It?
What to log for every inference request, how long to keep it, and when a compliance automation platform is worth it instead of building the pipeline yourself.