Cortex 2026 Benchmark — PRs Per Author Up 20% but Incidents Per PR Up 23.5%. We're Shipping Faster Into More Fires

The Cortex “Engineering in the Age of AI” 2026 benchmark dropped last week, and the headline numbers are going to make a lot of engineering leaders feel great — right until they read the fine print.

The Good News (On the Surface)

PRs per author increased 20% year-over-year. Deployment frequency is up across the board. Code output is accelerating at a pace we haven’t seen since the shift to microservices. Every engineering org I talk to is proudly reporting “we’re shipping faster than ever.” And by the raw throughput metrics, they’re right.

The Bad News (In Production)

Dig one layer deeper and the quality story is genuinely alarming. Incidents per PR increased 23.5% year-over-year. Change failure rates rose 30%. We are generating more code faster, and that code is causing more problems in production. The report’s most damning finding: only 33% of engineering leaders have data proving AI is actually improving outcomes. The other two-thirds? They’re operating on faith, anecdotes, and vibes. “It feels like we’re more productive” is not a measurement strategy.

Meanwhile, 41% of organizations rely on “informal guidelines” for AI governance — no formal policies, no measurement frameworks, no quality gates specifically designed for AI-assisted code. We’ve deployed the most powerful code generation technology in history and decided that a Slack message saying “use your judgment with Copilot” constitutes a governance framework.

What I Saw on My Own Team

I’ll share our experience because I think it’s representative. After rolling out GitHub Copilot across my 45-person engineering org, our PR velocity increased 35% in the first quarter. Leadership loved the charts. Our VP of Engineering presented the numbers at a board meeting. Everyone was thrilled.

Then my on-call team started getting paged more frequently. Not dramatically at first — maybe 15-20% more pages per week. The incidents were subtle: edge cases not handled in new payment processing logic, race conditions in concurrent queue consumers, integration issues where AI-generated code in Service A made assumptions about Service B’s API contract that weren’t quite right. These were bugs that experienced human developers would have caught through contextual awareness — knowing that “oh, this endpoint sometimes returns 204 instead of 200 when the cart is empty” or “this queue consumer can receive duplicate messages during a deployment.”

When I finally correlated our PR merge rates with our incident rates over a 6-month window, the curves moved together almost perfectly. More PRs, more incidents. The ratio was disturbingly consistent.

The Root Cause

AI generates plausible-looking code that passes surface-level review but lacks understanding of three critical things:

  1. System invariants — the unwritten rules about how your distributed system actually behaves under load, during deploys, and in failure modes
  2. Business logic edge cases — the weird customer scenarios that took years of production incidents to encode in your codebase
  3. Cross-service dependencies — the implicit contracts between services that aren’t captured in API specs or documentation

Traditional code review was already struggling to catch these issues when humans wrote the code. With 20% more PRs flooding the pipeline, reviewers are even more overwhelmed. Review depth is inversely proportional to review volume — when you’re reviewing 12 PRs a day instead of 10, each one gets less attention. And AI-generated code often looks clean and well-structured, which makes reviewers less vigilant. It pattern-matches as “good code” even when it’s subtly wrong.

What I’m Implementing

We’ve rolled out three specific countermeasures:

  1. AI-specific code review checklist. Every PR where AI assistance was used (self-reported by the author) gets reviewed against a supplemental checklist: Does this code handle our known edge cases? Does it respect our rate limiting patterns? Does it account for eventual consistency in our data layer? Does it follow our retry/backoff conventions?

  2. Mandatory integration tests for multi-service AI-assisted PRs. If an AI-assisted PR touches more than one service boundary, it requires an integration test that exercises the actual service interaction — not a mock. This has already caught 4 production-bound bugs in 6 weeks.

  3. A “quality tax.” For every 10 PRs shipped, the team spends one sprint day on production hardening — reviewing recent incidents, adding regression tests, strengthening monitoring, and auditing AI-generated code in critical paths.

The Controversial Take

I’d trade the 20% PR velocity increase for the old incident rate any day of the week. Every incident costs us 2-4 engineer-hours in response and remediation, plus customer trust, plus on-call burnout. Speed that generates incidents isn’t speed — it’s debt with compound interest. You’re borrowing from your future reliability budget to pad your current throughput numbers.

The Cortex benchmark paints a picture of an industry that has collectively decided to measure the accelerator while ignoring the engine temperature gauge.

Are your incident rates correlating with your AI-assisted code velocity? What are you doing about it?

The governance gap is what keeps me up at night, Luis. We have detailed, battle-tested processes for security review, accessibility compliance, and database migration approval — multi-step workflows with designated reviewers, documented criteria, and audit trails. But AI-generated code? It just flows through the exact same pipeline with zero additional scrutiny. We’re treating it as if it’s indistinguishable from human-written code, and the data is telling us that’s a bad assumption.

I’ve been pushing for what I’m calling an “AI code flag” on PRs where AI assistance was substantively used (not just autocomplete, but meaningful code generation). The goal is to track quality metrics separately so we can make data-driven decisions instead of arguing about feelings. We rolled this out 8 weeks ago as a voluntary checkbox in our PR template.

Early data is illuminating: AI-flagged PRs have a 1.7x higher defect rate in the first 30 days post-deploy compared to non-flagged PRs. Not catastrophic — we’re not seeing production meltdowns — but statistically significant enough to justify differentiated review processes. The defects cluster in exactly the categories you described: unhandled edge cases, incorrect assumptions about upstream service behavior, and missing error handling for failure modes that aren’t obvious from reading the API docs.

The 41% figure for organizations relying on “informal guidelines” — honestly, I think that’s generous. In my network of VP Engineering peers (about 20 people across Series B through public companies), I’d estimate the real number is closer to 60-70%, and “informal” often means “nonexistent in practice.” Someone wrote a Confluence page 6 months ago. Nobody reads it. There’s no enforcement mechanism. That’s the state of AI governance at most companies I talk to.

What concerns me most is the accountability gap. When a human developer writes buggy code, we have a clear chain: the author, the reviewer, the team lead. When AI generates code and a developer copy-pastes it with light modifications, who owns the quality? The developer who prompted the AI? The reviewer who approved a PR they assumed was human-written? The platform team that provisioned the AI tool? We haven’t answered these questions, and until we do, the incident rate will keep climbing.

I want to push back on the methodology slightly, Luis — not because I disagree with the conclusion, but because the data story is more nuanced than the headline suggests.

The 23.5% increase in incidents per PR could be confounded by several factors that changed simultaneously. Teams are also deploying more frequently (which mechanically increases the surface area for incidents), working on more complex features (as AI handles the simpler tasks, humans tackle harder problems), and expanding to new platforms and markets (which introduces unfamiliar failure modes). You’d need to control for code complexity, deployment target, team composition, and feature novelty to isolate the AI effect specifically. Correlation between PR velocity and incident rates doesn’t establish that AI caused the quality degradation — it’s possible both are driven by the same underlying pressure to ship faster.

That said, the directional signal is clear enough to warrant action even without perfect causal attribution. You don’t need a randomized controlled trial to start implementing quality gates.

My recommendation for anyone who wants to actually measure this: instrument your AI coding tools to log which code blocks were AI-generated vs. human-written, then correlate at the function level with production errors. Most Copilot and Cursor integrations can be configured to emit telemetry about which suggestions were accepted. Pair that with your error tracking (Sentry, Datadog, etc.) and you can build a surprisingly granular picture.

We did this for a quarter at my previous company and the findings were specific and actionable:

  • AI-generated error handling code was 3x more likely to fail in production than AI-generated happy-path code
  • AI-generated database queries had comparable defect rates to human-written ones (the AI is actually good at SQL)
  • AI-generated API integration code had 2.4x higher defect rate when the target API had non-standard conventions
  • AI-generated test code had excellent coverage numbers but missed edge cases at 4x the rate of human-written tests — the tests passed but didn’t test the right things

The pattern is clear: AI is great at the sunny day scenario and terrible at edge cases. It writes beautiful code for the 90% case and silently ignores the 10% that causes 90% of your production incidents. This makes sense when you think about training data — most code examples online demonstrate the happy path, not the gnarly edge cases that live in private codebases.

The practical implication: if you’re going to use AI-generated code, invest disproportionately in edge case testing and failure mode review. That’s where the quality gap is widest.

The security implications of shipping 20% more PRs with 23.5% more incidents per PR are compounding in ways that most engineering leaders aren’t tracking.

Here’s the flywheel effect I’m seeing:

More incidents → more pressure on incident response → less time for security reviews → more vulnerabilities ship → more security incidents → even more pressure on incident response.

It’s a flywheel spinning in the wrong direction, and AI-accelerated shipping velocity is pouring fuel on it.

Let me put specific numbers on this. My security engineering team (4 people) was already at capacity reviewing ~10 PRs per week that touched security-sensitive code paths. After the AI-driven PR velocity increase, we’re now seeing ~15 security-relevant PRs per week — a 50% increase that maps closely to the overall PR velocity bump. But my team didn’t grow by 50%. Something has to give, and what’s giving is review depth.

Last quarter, we caught a classic vulnerability in AI-generated authentication code: it properly validated JWT tokens but didn’t check token expiration in one of three code paths. The AI had generated a clean, well-structured auth middleware that looked correct at a glance. It had proper error handling, good logging, and followed our naming conventions. But it missed the exp claim validation in the refresh token flow — something a human developer familiar with our auth system would have caught because they’d know we had a production incident about this exact issue 18 months ago. That institutional knowledge doesn’t exist in the AI’s training data.

We’ve now implemented a hard policy: any PR touching authentication, authorization, or payment processing must have a security engineer review regardless of whether AI was involved. The overhead is significant — my team reviews about 15 PRs per week now with mandatory turnaround SLAs — but the alternative is letting AI-generated auth code ship with the same review bar as a CSS change.

Additional measures we’ve added:

  • Automated SAST scanning with rules specifically tuned for common AI code generation patterns (missing input validation, overly permissive CORS, hardcoded fallback values)
  • Security-specific integration tests for auth flows that run on every PR, not just nightly
  • A “security hotspot” annotation system in our codebase that flags files and functions where AI-generated code requires extra scrutiny

The Cortex data should be a wake-up call for security teams specifically. If your incident rate is climbing 23.5%, your security incident rate is almost certainly climbing with it — you just might not have discovered those incidents yet. Security bugs are slower to surface than functional bugs but far more expensive when they do.