AI Model Serving & Inference OptimizationPlaybook3 min readUpdated September 2026

A Network Isolation Checklist for a Production Inference Cluster

A GPU cluster serving live inference traffic is a valuable, sensitive piece of infrastructure: it holds model weights, it processes customer input directly, and it often has network paths to internal systems that a public-facing service should not have. Isolating it properly is less about any single control and more about checking that a handful of them are actually in place together.

None of the four checks below are unusual on their own. What is unusual is a team that has verified all four recently, rather than assuming a setup done correctly once is still correct today.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Check One: Nothing Reaches the GPU Nodes Directly From the Public Internet

Every request should pass through a gateway or load balancer before it reaches a node actually running inference. Confirm this by trying to reach a GPU node's address directly from outside your network rather than trusting a diagram that says it is private. Security group and firewall rules drift over time as new services get added, and a rule opened temporarily for debugging is the most common way this quietly breaks.

Check Two: Peering Connections Are Scoped to What They Actually Need

If your model serving cluster peers with other networks, whether a data pipeline, a customer's own VPC, or another internal service, check that the peering connection only allows the specific ports and directions of traffic it needs, not open access between the two networks. A peering connection set up broadly to get an integration working quickly and never tightened afterward is a common source of unintended exposure a year later.

Check Three: Multi-Tenant Traffic Cannot Cross Between Customers

If one cluster serves multiple customers or business units, confirm that a request tagged for one tenant cannot reach another tenant's data or model configuration, whether through a shared cache, a shared queue, or a routing bug. Test this directly rather than relying on the intended design: send a request as one tenant and confirm nothing in the response or in logs reveals another tenant's information.

Check Four: You Would Notice an Unexpected Connection

Isolation only matters if you would know when it fails. Confirm that connection attempts to your GPU nodes from outside their expected sources generate an alert someone actually sees, not just a log line nobody reads. A compliance automation platform like Vanta can help here by continuously checking your cloud network configuration for drift against the isolation rules you set, rather than relying on someone remembering to re-check manually.

Where This Quietly Breaks Over Time

The most common failure is not a bad initial design. It is drift: a rule opened for a debugging session and never closed, a new service added to the network without going through the same review as the original setup, or a peering connection expanded for one feature and never narrowed back down. Schedule a recheck of all four items above on a fixed calendar, not only when something prompts you to look.

Recheck these four items on a fixed calendar:

  • Confirm GPU nodes cannot be reached directly from the public internet by testing from outside your network, not by trusting a diagram.
  • Confirm each peering connection allows only the specific ports and directions of traffic it needs, and nothing broader.
  • Confirm a request sent as one tenant cannot reveal data or configuration belonging to another tenant through a shared cache or queue.
  • Confirm an unexpected connection attempt to a GPU node triggers an alert that a person actually sees and acts on.

A Worked Example: One Debugging Shortcut, Six Months Later

Say an engineer opens a security group rule to reach a GPU node directly while chasing a hard-to-reproduce latency bug during an incident. The bug gets fixed, the incident closes, and the rule stays open because closing it was never part of the incident checklist. Six months later, a routine external scan finds the node reachable from outside the network, and nobody remembers why. None of the four checks above are hard to run. What makes them work is treating every exception opened during an incident as something with an expiration date, not a permanent fixture nobody circles back to close.

Documenting the Setup So the Next Engineer Doesn't Guess

Write down what your isolation setup is supposed to look like: which subnets should be private, which peering connections exist and why, and what the expected alert behavior is for an unexpected connection attempt. Without that written baseline, a new engineer reviewing the network has no way to tell an intentional configuration from an accidental one, and drift becomes much harder to spot because nobody has a clear picture of what correct looks like.

Executive Capability Standard

What Good Looks Like

GPU nodes serving inference are unreachable from the public internet, peering connections are scoped narrowly, multi-tenant isolation is tested directly, and unexpected connection attempts generate a visible alert.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Diagram your current network paths into the model serving cluster and identify every route that reaches a GPU node.
2. Do Manually:Try reaching a GPU node's address directly from outside your network to confirm it is actually unreachable, rather than trusting the design diagram.
3. Delegate:Assign a platform engineer to run the four isolation checks on a fixed quarterly schedule and document the results.
4. Automate:Set up alerting on unexpected connection attempts to GPU nodes so drift is caught without a manual review.
5. Buy:Use a compliance automation platform such as Vanta to continuously monitor your cloud network configuration for isolation drift.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Vanta

Vanta fits here for continuously checking your cloud network configuration against the isolation rules you set, catching drift you would otherwise only find by manual review.

Visit Vanta→

Frequently Asked Questions

How often should we re-verify network isolation on our inference cluster?

At least quarterly, and after any change to your network configuration such as a new peering connection or a new service added to the same network. Isolation usually fails through drift, not a bad original design, so a fixed recheck schedule matters more than a one-time audit.

Is a private subnet enough to isolate a model serving cluster?

A private subnet is a good start but not sufficient on its own. You still need to confirm peering connections are scoped narrowly, multi-tenant traffic cannot cross between customers, and you would actually be alerted to an unexpected connection attempt.

What is the biggest risk in a multi-tenant inference setup specifically?

Traffic or data crossing between tenants through a shared component such as a cache or queue, rather than through the network layer itself. Test this directly by sending a request as one tenant and confirming nothing in the response reveals another tenant's information.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides