I’ve been wrestling with a tension that I suspect many of you are facing too.
The Setup: Our board increased our 2025 engineering budget by 18%—part of the 61% of companies that expanded engineering spend. The driver? AI transformation. We bought licenses for GitHub Copilot, invested in LLM infrastructure, hired AI specialists, and freed up “20% time” for AI experimentation.
The Problem: When our CFO asked me last month, “What’s the ROI on our AI investment?”—I couldn’t give her a satisfying answer.
And I’m not alone. According to recent analysis, only 20% of engineering teams are using engineering metrics to measure AI impact, despite widespread adoption.
The Faith-Based AI Budgeting Problem
Here’s what I’m seeing across the industry:
- 84% of developers use AI daily (high adoption
) - 41% of committed code is now AI-generated (real integration
) - But productivity gains remain flat at ~10% since 2023 (outcomes?
)
We’re treating AI like a necessary infrastructure investment—something you fund because “everyone else is doing it” rather than because we can demonstrate clear business impact.
What We’re Actually Measuring (And Why It’s Not Enough)
Most teams I talk to track:
- Adoption rates - “80% of our engineers use Copilot”
- AI code share - “35% of our codebase is AI-generated”
- Subjective surveys - “Developers feel 20% faster”
But we’re not measuring:
- Time-to-validated-customer-value (not just “time to PR”)
- Quality-adjusted velocity (bugs per AI-generated vs human-written code)
- Actual ROI including token costs (not just seat licenses)
The wild optimism of early AI adoption has given way to a more sober reckoning in 2026, where executives demand not just innovation but proof of its worth. As one analysis put it: “2026 is the year the bills come due on two years of AI experiments.”
The Measurement Challenge Nobody Wants to Admit
The hard part: our existing metrics were built for a world where humans write code.
- PR velocity goes up because AI cranks out more code—but does that code ship faster? Solve customer problems better?
- Cycle time looks worse because AI-generated code requires longer reviews (anyone else seeing the 91% longer review times?)
- Defect rates are inflated by AI producing 1.7× more bugs in some studies
We’re using a thermometer to measure distance. The tool doesn’t match the question.
What I’m Trying Instead
I’m experimenting with a multi-dimensional AI impact framework:
- Adoption & Usage (table stakes, not outcomes)
- AI Code Share (40-50% is the healthy range per 2026 benchmarks)
- Complexity-Adjusted Velocity - Did we ship harder problems faster, or just more trivial PRs?
- Quality Metrics - Bug rates, security vulnerabilities, code review feedback
- Business ROI - Token costs + seat licenses vs. measurable business outcomes (revenue, customer satisfaction, support reduction)
But honestly? I’m still figuring this out.
My Questions for You
- How are you demonstrating AI ROI to your CFO/board? What metrics actually convinced them?
- Are you seeing the AI productivity paradox? (Developers feel faster but team velocity is flat)
- What’s your threshold for “this AI investment isn’t working”? How long do you give it before pulling back?
The pressure is real. We increased budgets on the promise of AI productivity. But if we can’t measure it, how do we justify continued investment—or know when to course-correct?
Sources: