Reviewing AI-Generated Code for Security: A Practical Checklist
AI-generated code isn't inherently insecure, but it is often plausible, and plausible code passes casual review. Treat it like code from a fast new hire: read it, check the risky parts and let automated scanners catch what your eyes won't.
This checklist targets the failure patterns that show up most in assistant output: invented dependencies, missing authorization, unsafe input handling, secrets and outdated idioms. It is meant for a pull request reviewer with about ten minutes.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Why does AI-written code need a different kind of review?
Assistants optimize for code that looks right and runs. That produces distinctive problems:
- The code compiles and passes the happy-path test, so nothing feels wrong.
- It reproduces patterns from older or insecure examples, such as weak hashing or string-built queries.
- It fills gaps confidently. If it doesn't know your authorization model, it may skip the check rather than ask.
- Volume goes up. More code per pull request means less attention per line.
The author's responsibility doesn't change. Whoever accepts the suggestion owns it, and the reviewer shouldn't lower the bar because a tool wrote the diff. Your team's AI assistant policy should say so explicitly.
The ten-minute review checklist
Work through these in order on any pull request with substantial AI-assisted code:
- Dependencies: does every new package exist, is it the intended one, and is it maintained? Assistants sometimes suggest names that don't exist, and attackers can register look-alike names.
- Authentication and authorization: does each new endpoint or action verify who is calling and whether they may do this to this record? Check object-level access, not just login.
- Input handling: is input validated on the server, with parameterized queries and output encoding where data reaches a database, shell, file path or page?
- Secrets: any keys, tokens or connection strings in code, tests or comments? Are new environment variables read from config, not defaulted to a real value?
- Cryptography: is it using a vetted library and current algorithms, and is randomness from a secure source? Hand-rolled crypto is a rejection.
- Error handling and logging: do errors leak internals, and do logs capture tokens or personal data?
- Data exposure: does the response return more fields than the caller needs?
- Tests: do the tests include a failing case and an abuse case, or only the happy path?
Ask the author to explain any block they can't describe. If they can't, it shouldn't merge.
Which patterns fail most often?
Learn these by sight so they jump out in a diff:
- Missing tenant checks: a query that fetches a record by ID with no filter on the owning account.
- Trusting the client: price, role or user ID taken from the request body.
- Permissive defaults: wildcard CORS, debug modes, disabled certificate checks left in "temporarily".
- Outdated API use: deprecated auth flows or hash functions copied from old tutorials.
- Swallowed exceptions: a broad catch that returns success, hiding a failed security check.
Keep a running list of what you find in your own repo. The same three mistakes usually account for most findings, and they belong in your team's rules for the assistant.
Which automated checks should back up the human review?
A reviewer can't catch everything, so layer tools that run on every pull request:
- Secret scanning on commits, ideally blocking the push.
- Dependency scanning for known vulnerabilities and for suspicious new packages.
- Static analysis for the language you use, tuned to fail on high-severity findings only, or engineers will stop reading it.
- Tests that fail when a security rule is broken, such as an endpoint reachable without a token.
A scanning product such as Snyk covers dependency and code scanning in one workflow, but the pattern matters more than the vendor. On timing, CISA's federal directive BOD 19-02 sets 15 days for remediating critical vulnerabilities on internet-accessible systems1, which is a reasonable internal target to borrow.
When should a senior engineer or security specialist look too?
Some code deserves a second, specialized reviewer regardless of who wrote it: authentication and session handling, payments, anything touching encryption, file uploads, data export and deletion, and code that runs with elevated privileges. Require that extra approval by path using code ownership rules, so the process doesn't depend on someone remembering. For everything else, a normal reviewer with this checklist is enough.
What Good Looks Like
Every pull request with AI-assisted code gets a checklist review of dependencies, authorization, input handling and secrets, backed by automated scanning in CI.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Frequently Asked Questions
Is AI-generated code less secure than human code?
Not necessarily, but it fails differently. It tends to look correct while skipping checks it has no context for. Review it with the same standards as any code, and add automated scanning so review isn't the only safeguard.
What is a hallucinated package?
A dependency name an assistant suggests that doesn't exist, or that exists only because someone registered it to trap installers. Verify every new package's name, publisher and maintenance history before installing it.
Should we label AI-generated code in pull requests?
It's optional. Labels help if you'll use them to focus review or to measure adoption. The important rule is that the author owns and can explain every line, whether or not they wrote it by hand.
Can a scanner replace code review of AI output?
No. Scanners find known patterns and vulnerable packages, but miss logic errors such as a missing ownership check. Use both: scanning for breadth, human review for intent and authorization.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Security patch remediation SLAs (CISA federal mandates, used as industry norm). CISA Binding Operational Directives 19-02 and 22-01 (CISA briefing hosted at NIST CSRC), 2022.
Related Guides
Auditing Your AI Code Review Tool for What It's Actually Missing
A thirty-minute audit for finding out what your AI code review tool catches, what it misses, and where it's training your team to stop reading diffs.
Setting Up AI Code Review the Right Way
A rollout order for AI code review: what it catches well, where it misses real risk, and which pull requests still need a second human.
Rolling Out AI Code Review Without Burying Your Team
A practical rollout plan for AI code review: what to let it block, how to tune out false positives, and how to keep a human as the tie-breaker.
Secure Code Review: A Checklist for Everyday Pull Requests
A practical checklist for secure code review that reviewers can apply to pull requests: authorization, input handling, secrets, dependencies, logging and more.
Rolling Out AI Code Review Without Drowning Reviewers in Noise
A staged rollout for AI code review tools: shadow mode first, then advisory comments, then a required check, so it earns trust instead of getting muted.
Where AI Code Review Catches Real Bugs, and Where It Misses
A clear-eyed look at what automated code review reliably catches in pull requests, where it still misses real defects, and how to route the rest to people.