AI codingChecklist3 min readUpdated September 2026

Reviewing AI-Generated Code for Security: A Practical Checklist

AI-generated code isn't inherently insecure, but it is often plausible, and plausible code passes casual review. Treat it like code from a fast new hire: read it, check the risky parts and let automated scanners catch what your eyes won't.

This checklist targets the failure patterns that show up most in assistant output: invented dependencies, missing authorization, unsafe input handling, secrets and outdated idioms. It is meant for a pull request reviewer with about ten minutes.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Why does AI-written code need a different kind of review?

Assistants optimize for code that looks right and runs. That produces distinctive problems:

  • The code compiles and passes the happy-path test, so nothing feels wrong.
  • It reproduces patterns from older or insecure examples, such as weak hashing or string-built queries.
  • It fills gaps confidently. If it doesn't know your authorization model, it may skip the check rather than ask.
  • Volume goes up. More code per pull request means less attention per line.

The author's responsibility doesn't change. Whoever accepts the suggestion owns it, and the reviewer shouldn't lower the bar because a tool wrote the diff. Your team's AI assistant policy should say so explicitly.

The ten-minute review checklist

Work through these in order on any pull request with substantial AI-assisted code:

  1. Dependencies: does every new package exist, is it the intended one, and is it maintained? Assistants sometimes suggest names that don't exist, and attackers can register look-alike names.
  2. Authentication and authorization: does each new endpoint or action verify who is calling and whether they may do this to this record? Check object-level access, not just login.
  3. Input handling: is input validated on the server, with parameterized queries and output encoding where data reaches a database, shell, file path or page?
  4. Secrets: any keys, tokens or connection strings in code, tests or comments? Are new environment variables read from config, not defaulted to a real value?
  5. Cryptography: is it using a vetted library and current algorithms, and is randomness from a secure source? Hand-rolled crypto is a rejection.
  6. Error handling and logging: do errors leak internals, and do logs capture tokens or personal data?
  7. Data exposure: does the response return more fields than the caller needs?
  8. Tests: do the tests include a failing case and an abuse case, or only the happy path?

Ask the author to explain any block they can't describe. If they can't, it shouldn't merge.

Which patterns fail most often?

Learn these by sight so they jump out in a diff:

  • Missing tenant checks: a query that fetches a record by ID with no filter on the owning account.
  • Trusting the client: price, role or user ID taken from the request body.
  • Permissive defaults: wildcard CORS, debug modes, disabled certificate checks left in "temporarily".
  • Outdated API use: deprecated auth flows or hash functions copied from old tutorials.
  • Swallowed exceptions: a broad catch that returns success, hiding a failed security check.

Keep a running list of what you find in your own repo. The same three mistakes usually account for most findings, and they belong in your team's rules for the assistant.

Which automated checks should back up the human review?

A reviewer can't catch everything, so layer tools that run on every pull request:

  • Secret scanning on commits, ideally blocking the push.
  • Dependency scanning for known vulnerabilities and for suspicious new packages.
  • Static analysis for the language you use, tuned to fail on high-severity findings only, or engineers will stop reading it.
  • Tests that fail when a security rule is broken, such as an endpoint reachable without a token.

A scanning product such as Snyk covers dependency and code scanning in one workflow, but the pattern matters more than the vendor. On timing, CISA's federal directive BOD 19-02 sets 15 days for remediating critical vulnerabilities on internet-accessible systems1, which is a reasonable internal target to borrow.

When should a senior engineer or security specialist look too?

Some code deserves a second, specialized reviewer regardless of who wrote it: authentication and session handling, payments, anything touching encryption, file uploads, data export and deletion, and code that runs with elevated privileges. Require that extra approval by path using code ownership rules, so the process doesn't depend on someone remembering. For everything else, a normal reviewer with this checklist is enough.

Executive Capability Standard

What Good Looks Like

Every pull request with AI-assisted code gets a checklist review of dependencies, authorization, input handling and secrets, backed by automated scanning in CI.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Learn the recurring patterns: hallucinated packages, missing object-level authorization, trusted client input and outdated cryptography.
2. Do Manually:Adopt the eight-point checklist in your pull request template and use it for a month.
3. Delegate:Assign code owners for sensitive paths so authentication, payments and data export need a specialist approval.
4. Automate:Run secret scanning, dependency scanning and static analysis in CI, with tests that fail on broken security rules.
5. Buy:Adopt a code and dependency security scanner once triaging findings by hand takes too much time.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Snyk

Fits as a dependency and code scanner that runs on every pull request, including AI-assisted ones.

Visit Snyk→
GitHub Copilot

Fits when your team already uses it and you want this checklist to sit alongside your pull request reviews.

Visit GitHub Copilot→

Frequently Asked Questions

Is AI-generated code less secure than human code?

Not necessarily, but it fails differently. It tends to look correct while skipping checks it has no context for. Review it with the same standards as any code, and add automated scanning so review isn't the only safeguard.

What is a hallucinated package?

A dependency name an assistant suggests that doesn't exist, or that exists only because someone registered it to trap installers. Verify every new package's name, publisher and maintenance history before installing it.

Should we label AI-generated code in pull requests?

It's optional. Labels help if you'll use them to focus review or to measure adoption. The important rule is that the author owns and can explain every line, whether or not they wrote it by hand.

Can a scanner replace code review of AI output?

No. Scanners find known patterns and vulnerable packages, but miss logic errors such as a missing ownership check. Use both: scanning for breadth, human review for intent and authorization.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Security patch remediation SLAs (CISA federal mandates, used as industry norm). CISA Binding Operational Directives 19-02 and 22-01 (CISA briefing hosted at NIST CSRC), 2022.

Related Guides