Only 20% Measure AI Impact, But 61% Increased Engineering Budgets in 2025—Are We Funding AI on Faith or Data?

I’ve been wrestling with a tension that I suspect many of you are facing too.

The Setup: Our board increased our 2025 engineering budget by 18%—part of the 61% of companies that expanded engineering spend. The driver? AI transformation. We bought licenses for GitHub Copilot, invested in LLM infrastructure, hired AI specialists, and freed up “20% time” for AI experimentation.

The Problem: When our CFO asked me last month, “What’s the ROI on our AI investment?”—I couldn’t give her a satisfying answer.

And I’m not alone. According to recent analysis, only 20% of engineering teams are using engineering metrics to measure AI impact, despite widespread adoption.

The Faith-Based AI Budgeting Problem

Here’s what I’m seeing across the industry:

  • 84% of developers use AI daily (high adoption :white_check_mark:)
  • 41% of committed code is now AI-generated (real integration :white_check_mark:)
  • But productivity gains remain flat at ~10% since 2023 (outcomes? :cross_mark:)

We’re treating AI like a necessary infrastructure investment—something you fund because “everyone else is doing it” rather than because we can demonstrate clear business impact.

What We’re Actually Measuring (And Why It’s Not Enough)

Most teams I talk to track:

  1. Adoption rates - “80% of our engineers use Copilot”
  2. AI code share - “35% of our codebase is AI-generated”
  3. Subjective surveys - “Developers feel 20% faster”

But we’re not measuring:

  • Time-to-validated-customer-value (not just “time to PR”)
  • Quality-adjusted velocity (bugs per AI-generated vs human-written code)
  • Actual ROI including token costs (not just seat licenses)

The wild optimism of early AI adoption has given way to a more sober reckoning in 2026, where executives demand not just innovation but proof of its worth. As one analysis put it: “2026 is the year the bills come due on two years of AI experiments.”

The Measurement Challenge Nobody Wants to Admit

The hard part: our existing metrics were built for a world where humans write code.

  • PR velocity goes up because AI cranks out more code—but does that code ship faster? Solve customer problems better?
  • Cycle time looks worse because AI-generated code requires longer reviews (anyone else seeing the 91% longer review times?)
  • Defect rates are inflated by AI producing 1.7× more bugs in some studies

We’re using a thermometer to measure distance. The tool doesn’t match the question.

What I’m Trying Instead

I’m experimenting with a multi-dimensional AI impact framework:

  1. Adoption & Usage (table stakes, not outcomes)
  2. AI Code Share (40-50% is the healthy range per 2026 benchmarks)
  3. Complexity-Adjusted Velocity - Did we ship harder problems faster, or just more trivial PRs?
  4. Quality Metrics - Bug rates, security vulnerabilities, code review feedback
  5. Business ROI - Token costs + seat licenses vs. measurable business outcomes (revenue, customer satisfaction, support reduction)

But honestly? I’m still figuring this out.

My Questions for You

  1. How are you demonstrating AI ROI to your CFO/board? What metrics actually convinced them?
  2. Are you seeing the AI productivity paradox? (Developers feel faster but team velocity is flat)
  3. What’s your threshold for “this AI investment isn’t working”? How long do you give it before pulling back?

The pressure is real. We increased budgets on the promise of AI productivity. But if we can’t measure it, how do we justify continued investment—or know when to course-correct?

Sources:

This hits close to home. We’re facing the exact same scrutiny from our CFO, and I’ll be honest—the conversation last quarter was uncomfortable.

What Actually Worked for Our ROI Conversation

We shifted the conversation from “productivity gains” to “capability expansion.”

Instead of: “Our developers are 20% faster”
We said: “We’re executing a 40% more ambitious roadmap with the same headcount.”

The Numbers We Used:

  1. Avoided Headcount Costs: We needed to add 3 senior engineers ($450K total) but didn’t because of AI-assisted productivity. That’s our “ROI” right there—money we didn’t spend.

  2. Complexity of Work: We tracked the difficulty of shipped features, not just volume. Our “High Complexity” feature delivery went up 35% while “Low Complexity” stayed flat. AI handles the boring stuff; humans tackle harder problems.

  3. Time-to-Market for Strategic Bets: We shipped our enterprise product 6 weeks ahead of schedule, which unlocked $2M in ARR a quarter early. That’s measurable business impact.

The Uncomfortable Truth About Measurement

You’re absolutely right that traditional metrics break down. But here’s what I learned: executives don’t actually want “engineering metrics”—they want business outcomes.

When I stopped talking about DORA metrics and started talking about revenue enabled, costs avoided, and strategic bets won, the conversation changed completely.

The AI Productivity Paradox Is Real

Are you seeing the AI productivity paradox? (Developers feel faster but team velocity is flat)

Yes. 100%. We’re seeing this too.

Individual engineers crank out 30% more PRs. But sprint velocity? Unchanged. Why?

  1. Code review bottleneck: Seniors spend 4-6 hours/week extra reviewing AI-generated code (it has 1.7× more issues on average)
  2. Integration complexity: More PRs doesn’t mean faster shipping when they conflict or need coordination
  3. Junior engineers compress learning: They finish tasks in 3-4 weeks instead of 6-8, but they’re not learning how the code works—they’re learning to prompt

We’re not “20% faster.” We’re handling 40% more ambitious work with the same timeline. That’s the reframe.

My Threshold for “Not Working”

I gave us 9 months to show measurable business impact (we’re at month 9 now). If we couldn’t point to:

  • Revenue acceleration
  • Cost avoidance
  • Customer satisfaction improvement

…then I’d cut AI investment by 50% and reassess.

But we did hit those metrics. It just took longer than the hype cycle suggested, and the gains are different than we expected.

Bottom line: Stop trying to measure AI productivity the way we measured human productivity. The tool changed; the measurement must change too.

Michelle, I love your reframe around “capability expansion”—that’s exactly the mindset shift we needed.

But I want to push back on one thing: who actually benefits from AI productivity, and are we measuring that honestly?

The Uneven Distribution of AI Gains

We’ve been tracking AI impact across different engineer levels for 6 months. Here’s what we’re seeing:

Junior Engineers (0-2 years):

  • 55% faster task completion :white_check_mark:
  • But 40% drop in “deep understanding” based on code review questions :warning:
  • Heavy reliance on AI = shallow learning

Mid-Level Engineers (3-7 years):

  • 25% faster on greenfield work :white_check_mark:
  • 10% slower on debugging/refactoring (AI-generated code is harder to troubleshoot) :cross_mark:
  • Best ROI group overall

Senior Engineers (8+ years):

  • 15% faster on new features :white_check_mark:
  • 30-40% slower because of code review burden for AI-generated PRs :cross_mark:
  • Burning out from being “AI code quality gatekeepers”

The Hidden Cost: Senior Engineer Burnout

This is the part that keeps me up at night. Our senior engineers are shouldering an invisible tax:

  1. Reviewing 30% more PRs (because juniors ship faster with AI)
  2. Fixing AI-generated technical debt (1.7× more bugs means more cleanup)
  3. Teaching juniors why the AI solution is wrong (not just that it’s wrong)

We’re seeing early signs of senior engineer dissatisfaction—two of my best people mentioned burnout in 1:1s. They feel like they’re doing less “building” and more “quality control for robots.”

Is that a sustainable ROI model? We’re trading junior velocity for senior engagement.

The Measurement Gap We’re Not Talking About

Most AI impact discussions focus on individual productivity. But engineering is a team sport. What we’re not measuring:

  • Team cohesion: Are we building shared understanding, or just parallel tracks of AI-assisted work?
  • Knowledge transfer: When AI generates the code, how do juniors learn the patterns?
  • Long-term maintainability: Who’s going to debug this AI spaghetti in 18 months?

I’m not saying AI is bad—we’re seeing real gains. But I think we’re optimizing for the wrong metrics if we’re not tracking these human costs.

My Controversial Take

I think the “40% more ambitious roadmap” success metric is real—but it might come at the cost of:

  1. Senior engineer retention (who wants to be an AI code reviewer forever?)
  2. Junior engineer skill development (they’re fast but fragile)
  3. Technical debt accumulation (AI code has 60% less refactoring, 48% more copy-paste patterns)

What happens in year 2 when the bill comes due on that technical debt? Does the ROI still hold?

I’d love to hear how others are thinking about the long-term sustainability of AI-driven productivity, not just the quarterly wins.

Luis raises a critical point that I think engineering leaders are underestimating: the product velocity paradox.

From a product perspective, I’m seeing the disconnect between engineering velocity and customer impact—and it’s creating friction in how we prioritize AI investment.

The Engineering-Product Measurement Disconnect

Engineering tells me: “We shipped 35% more features this quarter thanks to AI!”

But when I look at the product metrics:

  • Feature adoption rates: unchanged
  • Customer satisfaction: flat
  • Time-to-validated-learning: actually slower

Why? Because we’re shipping faster, but we’re not learning faster.

The Hidden Bottleneck: Product Discovery

Here’s what I’m seeing in our workflow:

Before AI:

  • 3 weeks to build feature → 2 weeks to validate with customers → iterate

After AI:

  • 1 week to build feature (67% faster! :tada:) → still 2 weeks to validate → oh wait, we built the wrong thing

The bottleneck isn’t engineering execution anymore—it’s product discovery, customer research, and validation. AI sped up the wrong part of the funnel.

The Measurement Question From a Product Lens

When Michelle says “40% more ambitious roadmap,” I have to ask:

Are those ambitious features the right features? Or are we just building more things faster without validating whether they solve customer problems?

In our case, we shipped 3 features ahead of schedule thanks to AI velocity. Great, right?

But 2 of those 3 features hit only 30% of projected usage. We built them so fast we didn’t do proper customer development. We skipped the “slow thinking” part.

What I’d Rather Measure

Instead of “engineering velocity” or “AI code share,” I want to measure:

  1. Time-to-validated-customer-value (not just “time to ship”)
  2. Feature hit rate (what % of shipped features meet adoption/satisfaction targets?)
  3. Iteration cycles per feature (are we learning faster, or just building faster?)

Because here’s the uncomfortable truth: if engineering ships the wrong thing 2× faster, we didn’t save time—we wasted it faster.

The AI Investment Question for Product Leaders

From my seat, the ROI question isn’t “Did AI make engineering faster?”

It’s: “Did AI help us learn what customers need, faster?”

And honestly? For most teams, the answer is no. AI accelerated execution, but it didn’t accelerate discovery.

So when the CFO asks “What’s the ROI on AI investment?”—I’m asking a different question:

Should we invest in AI for product discovery (customer research, experimentation, data analysis) instead of just AI for code generation?

Because right now, we’re speeding up the wrong part of the product development cycle.