Building an RBAC Model That Survives Contact With Reality
Most role-based access systems start clean and drift into a mess of one-off exceptions within a year: a role created for a single person's edge case, a permission nobody remembers granting, a service account with far more access than it uses. The fix isn't more roles, it's building the model around actions from the start.
This is the sequence that tends to hold up as the team and the product both grow.
How do you start an RBAC model from actions, not job titles?
A role called "Manager" tells you nothing about what it can actually do, and it tends to accumulate permissions over time as different managers ask for different things. Start instead from a list of concrete actions your API exposes: read customer records, refund a payment, delete a workspace, modify another user's role. Build roles by grouping actions that genuinely belong together, then name the role after what it does.
This sounds like more upfront work than copying a title-based structure from your HR system, and it is, but it's the difference between an access model you can actually reason about and one you can only guess at.
Separate roles for humans from roles for services
A service account that calls your API on a schedule doesn't need, and shouldn't have, the same role structure as a human user clicking through a dashboard. Give machine identities their own narrow, explicit permission set scoped to exactly what that specific job does, rather than reusing a human role because it happened to have the right permissions included.
The most common finding in an RBAC audit is a service account with admin-level access because it was faster to grant at setup time than to figure out the minimal set it actually needed. Fix this at creation time; it's much harder to narrow later once other things quietly depend on the excess access.
How do you build a deny-by-default access matrix?
Write the actual matrix, roles down one side, actions across the top, before you build any admin screen to manage it. Every cell starts denied; you only fill in a checkmark where a role genuinely needs that action. This surfaces gaps you won't see by reasoning about roles individually: a role that seems fine in isolation often turns out, once it's in the matrix next to everything else, to have an action nobody can explain why it needs.
Keep the matrix somewhere your whole engineering team can see it, not buried in one person's design doc. A model nobody else can review isn't actually a model, it's one person's memory.
Test role changes the same way you test code
A permission change deployed without a test is a permission change nobody verified. Write tests that assert a given role can perform the actions it should and, just as importantly, cannot perform the actions it shouldn't. The second half gets skipped far more often than the first, and it's the half that actually catches an over-broad grant before it reaches production.
Run these tests on every deploy that touches the permission model, the same way you'd run any other regression suite, rather than relying on someone remembering to check manually.
For example, suppose a support role is meant to view customer records but not export them. A test that only confirms support staff can open a record will pass even if a recent change quietly granted export rights. Add the second assertion, that an export attempt is rejected with the expected status, and run both on every deploy that touches permissions. When a new role is added, copy the closest existing role's tests first, then change the expectations, so the denial cases are never forgotten.
Review the impossible combinations, not just the common ones
Once your matrix exists, look specifically for role combinations that shouldn't be possible together: a single account that can both approve an expense and submit one, or both modify user roles and audit the log of role changes. These separation-of-duties gaps are easy to miss when you're reviewing roles one at a time, and they're exactly the kind of finding an auditor or a careful attacker looks for first.
Schedule this review on a fixed cadence rather than only after an incident. Roles that were fine in isolation when created can become a problem in combination once two separate teams each add their own permission without seeing the other's.
Build the model in this sequence:
- List the concrete actions your API exposes, such as reading customer records, refunding a payment or changing another user's role.
- Group actions that belong together into roles, and name each role after what it does rather than a job title.
- Give machine identities their own narrow permission sets instead of reusing a human role.
- Write the deny-by-default matrix with roles down one side and actions across the top, and fill in only cells with a clear need.
- Test that each role can do what it should and cannot do what it shouldn't, then review combinations that break separation of duties.
What Good Looks Like
Good access control means every role maps to a specific, documented set of actions, service accounts have their own narrow roles separate from human ones, and someone actively checks for separation-of-duties conflicts on a schedule.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How many roles is too many?
There's no fixed number, but if you can't explain the difference between two roles in one sentence, you probably have too many. A common failure mode is a role created for one person's specific need that never gets consolidated once that need turns out to be common; watch for that pattern more than for a raw count.
Should service accounts ever share a role with human users?
Generally no, even if the permissions happen to overlap today. Human and machine access patterns diverge over time, a human role picks up UI-specific permissions a service account never needs, and untangling a shared role later is harder than starting separate ones now.
How often should we audit for separation-of-duties conflicts?
Quarterly for most teams, and immediately after any reorganization that moves permissions between teams. These conflicts accumulate quietly as roles change hands, so a fixed cadence catches them faster than waiting for someone to notice during an audit or an incident.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Continuous Device Verification for a Zero-Trust API
How continuous device and identity verification actually works in a zero-trust architecture, and where to draw the line for a small engineering team.
Rolling Out Zero Trust in Production Without a Broad Outage
A checklist for rolling out stricter API authentication and authorization in production, and the pitfalls that turn a rollout into an incident.
How to Audit Whether Your APIs Actually Enforce Zero Trust
A step-by-step method for testing whether your APIs enforce zero trust in practice, not just on paper, and what to do with what you find.
The Real Latency Cost of Zero Trust, and How to Measure It
How to find out how much latency your zero trust controls actually add, which checks are worth the cost, and which ones you can move off the hot path.
Designing Role-Based Access Control That Survives Your Next Reorg
A worksheet approach to mapping roles to permissions so access control doesn't quietly rot every time your team's structure changes.
Keeping Auth Checks Fast as Your API Traffic Grows
A worked example for keeping zero trust authorization checks fast as request volume grows, and where teams usually add latency without noticing.