← Back to Blog
AI QualityMay 28, 2025· 9 min read

Why AI-Generated Code Still Needs Expert Review — A Real-World Breakdown

We audited 1,200 AI-generated functions and categorised every failure. Here's what we found.

Why AI-Generated Code Still Needs Expert Review — A Real-World Breakdown

We reviewed 1,200 AI-generated functions across 23 production projects over 12 months. We categorised every flaw we caught before they could ship. The results are educational.

The breakdown

Security issues (22% of flagged code). The most common: missing input validation on API endpoints that accept user data, overly permissive CORS configurations, JWT tokens stored in localStorage instead of httpOnly cookies, and SQL queries built with string interpolation instead of parameterized statements. None of these are exotic. All of them are serious.

Logic errors (31%). AI code handles the happy path well. It handles edge cases poorly. Off-by-one errors in pagination, incorrect handling of null versus undefined, race conditions in async operations when two API calls resolve out of order. These don't fail visibly in development — they fail in production under real usage patterns.

Performance issues (19%). The most common: N+1 queries (fetching a list, then fetching related data for each item in a loop), missing database indexes on columns used in WHERE clauses, large objects passed by value in tight loops, and blocking operations on the main thread in Node.js services.

Architecture mismatch (17%). AI generates code that works for the feature described in the prompt, but doesn't fit the surrounding system. A new component that reimplements auth logic that already exists elsewhere. A new API endpoint that duplicates a function three files away. A state management approach that conflicts with how the rest of the app handles state.

The remaining 11%. A long tail: deprecated library usage, missing error handling, incorrect timezone assumptions, hardcoded values that should be configuration, and code that's simply unclear to the point of being a maintenance liability.

The important caveat

These issues aren't AI-specific. Human developers introduce all of these same classes of bug. The difference is that AI produces code at a rate that can outpace a team's review capacity if you don't have explicit process controls. Our 47-point review checklist exists precisely because the volume of AI output requires a structured approach to catch what would otherwise slip through.

What good review looks like

Effective AI code review is not reading every line. It's structured: you focus on the boundaries (inputs, outputs, data access, external calls), you run static analysis, you check that the code integrates correctly with the existing system, and you read the parts that look "too clean" — because those are the places AI glosses over complexity it doesn't fully understand.

Done correctly, this review takes roughly 20–30% of the time it would take to write the code manually. That's where the 10× speed multiplier comes from, and why it doesn't come at the cost of quality.

Ready to ship like this?

Get a fixed-cost quote within 48 hours. No obligation.

Start Your Project