Building a CI/CD Hardening Scorecard You Can Show a Client
For a managed security service provider, a client's CI/CD pipeline isn't just infrastructure, it's part of what you're being paid to secure. A pipeline that deploys code without required review, without secret scanning, and without any record of who approved what is a gap in the client's security posture whether or not anyone's found it yet.
This worksheet walks through the rows worth scoring on a client's pipeline and how GitHub Actions and GitLab CI each handle them, so an assessment produces the same answers regardless of which platform the client happens to be running.
Use it the same way you'd use any other security control checklist: as a repeatable structure that turns a subjective impression of a client's pipeline into a specific, defensible list of findings.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Row one: branch protection and required reviews
Score whether the client's default branch actually requires a passing pipeline and at least one approving review before a merge is allowed, not just whether the setting exists somewhere in the repository configuration. Both GitHub Actions and GitLab CI support this natively through branch protection rules and merge request approval rules; the gap is almost always in enforcement, not availability.
A surprising number of client repositories have the setting configured for their default branch but not for release branches, which is exactly where an attacker with stolen credentials would push a change to avoid review.
Row two: secret scanning and dependency gates
Score whether the pipeline actively blocks a merge when it detects a hardcoded credential or a dependency with a known critical vulnerability, rather than just flagging it after the fact. GitHub's secret scanning and dependency review, and GitLab's built-in secret detection and dependency scanning, both do this as pipeline jobs; the difference between clients is almost always whether the finding is configured to fail the build or just log a warning nobody reads.
A warning that doesn't block a merge is close to not existing for the purposes of a security assessment.
Row three: artifact signing and provenance
Score whether the client can prove that the artifact running in production is the exact one their pipeline built and tested, not something swapped in afterward. Both platforms support signing build artifacts and generating provenance attestations tying an artifact back to its source commit and pipeline run.
Change failure rate is a useful cross-check here too: teams with weak deploy controls tend to see more incidents traced back to changes that shouldn't have shipped1. If a client can't answer which pipeline run produced a given production artifact, that's usually a sign the rest of the scorecard needs a closer look too.
Turning the scorecard into a client deliverable
Present the scorecard as a simple status per row, in place, partial, missing, alongside a specific fix for anything short of in place. Clients respond better to a concrete list of gaps and fixes than to a general statement that their pipeline needs hardening, and it gives you a natural follow-on engagement scoped around closing the gaps you found.
Re-run the scorecard on a schedule rather than treating it as a one-time assessment. Pipelines drift as teams add new repositories and new contributors, and a hardening review from a year ago tells you little about the client's current exposure. Attach a rough remediation timeline to each missing or partial row too, since a client's own leadership usually needs that timeline to prioritize the fix against everything else competing for engineering time.
A client-ready scorecard should include:
- A clear status for each row: in place, partial, or missing.
- A specific fix for every row that falls short of in place.
- A re-run schedule so the scorecard keeps tracking gaps over time.
- A proposed follow-on engagement scoped around closing the gaps you found.
What the scorecard doesn't cover
This worksheet focuses on the pipeline itself, not the broader application security program it feeds into. A hardened pipeline still deploys vulnerable application code if nobody's running static analysis or a security review on the code itself; treat CI/CD hardening as one row in a larger security assessment, not the whole assessment.
Similarly, a self-hosted runner introduces its own attack surface that this scorecard doesn't score directly, patch status, network segmentation, and who has access to the machine deserve their own checklist alongside this one. A client who scores well on every row here can still have a poorly patched runner sitting on a flat network, so pair this worksheet with a separate review of the infrastructure the pipeline actually runs on before calling the assessment complete. Selling the two together, pipeline hardening and runner infrastructure review, is usually an easier conversation than presenting a clean pipeline scorecard and a separate, unrelated infrastructure finding weeks apart.
What Good Looks Like
Good looks like every row on this scorecard reading in place for a client's production pipeline, re-checked on a fixed schedule, with specific fixes documented for anything short of that.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
When a client's pipeline deploys to AWS, checking that it authenticates through a scoped role rather than a long-lived access key belongs on the same hardening scorecard as the pipeline's own settings.
For clients on Google Cloud, workload identity federation removes a stored service account key as an attack surface the pipeline would otherwise depend on.
Vanta gives your clients continuous monitoring of the same controls this scorecard checks manually, which turns a quarterly assessment into something closer to real-time coverage.
Frequently Asked Questions
Which platform makes secret scanning easier to enforce, GitHub Actions or GitLab CI?
Both can enforce it equally well once configured correctly; the real difference between clients is almost never the platform, it's whether the finding is set to block a merge or just log a warning. Check the enforcement setting directly rather than assuming it's on.
How often should we re-run a CI/CD hardening assessment for a client?
Quarterly is reasonable for most clients, sooner if they're adding new repositories or contributors frequently. Pipelines drift faster than most other infrastructure because it's easy for a new project to skip the standard someone set up a year earlier.
Does artifact signing actually stop a supply chain attack, or is it just documentation?
It's a detection and accountability control, not a preventive one on its own. It lets you prove after the fact whether a production artifact matches what the pipeline actually built, which matters for incident response even though it doesn't stop a compromised pipeline from producing a bad artifact in the first place.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Change failure rate by DORA performance cluster. DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.
Related Guides
GitHub Actions vs GitLab CI vs CircleCI: Continuous Integration Comparison
Compare GitHub Actions, GitLab CI, and CircleCI: build speeds, runner pricing, matrix testing, Docker orchestration, secret management, and DORA metrics.
Application Security for the Team That Sells Security
Questions an MSSP should ask before choosing Snyk or GitHub Advanced Security to secure its own detection tooling and client-facing platform.
Database Infrastructure for Managed Security Providers
MSSPs storing security event data and audit trails have narrower requirements than most apps. Here's how Supabase and AWS RDS compare.
Cursor vs GitHub Copilot for MSSPs and Security Operations Teams
A managed security provider has to vet an AI coding vendor the way it vets any tool touching client data. What to check before Cursor or Copilot join the SOC.
SOC 2 for MSSPs: Proving Your Own Security, Not Just Selling It
Why a managed security service provider's own SOC 2 audit is different, and how Vanta, Drata and Secureframe fit a security vendor that's already instrumented.
CrowdStrike vs SentinelOne for MSSPs Building a Service
For an MSSP, the CrowdStrike vs SentinelOne choice is about partner economics and differentiation, not just detection quality. A provider side breakdown.