Headcount Planning When Compute Writes the Code
Every annual planning cycle in every engineering org runs on the same hidden equation: roadmap ambition divided by engineer output equals requisitions. It has been true for so long that nobody writes it down anymore. You size the work, you divide by what a team can ship in a year, and the remainder becomes a hiring plan. Finance builds the budget around it, recruiting builds pipelines around it, and managers build careers around it.
That equation quietly broke. In an agent-heavy org, the marginal unit of engineering output is no longer another senior hire — it is tokens plus the review bandwidth to absorb what those tokens produce. NVIDIA now hands engineers token budgets worth roughly half their base salary, and Jensen Huang has said he would be "deeply alarmed" if a $500,000 engineer consumed less than $250,000 of tokens a year. Whether or not you take the specific ratio seriously, the structural point stands: a company can now convert dollars into working code through two different doors, and the annual plan only has a form field for one of them.
This post is about what headcount planning looks like when hiring one staff engineer and buying $400,000 of inference are competing line items for the same outcome — and why the honest planning question has shifted from "how many engineers does this take" to "what is our decision and review capacity, and what inference budget saturates it."
The Ritual Assumes Output Scales with Heads
Traditional capacity planning is an artifact of an era when the binding constraint was typing. Engineer-hours were the scarce input, so every planning abstraction — story points per sprint, FTEs per initiative, the loaded cost per head — was a proxy for how much code a human could produce. FTE math exists precisely because finance needs labor input in quantifiable units it can project forward: historical FTE data in, next year's salary and benefits expense out.
Agents break the proxy in both directions at once:
- Production is no longer scarce. A single engineer running a fleet of coding agents can generate the pull-request volume of a small team. Telemetry from Faros AI across more than 10,000 developers found high-AI-adoption teams merging 98% more pull requests than their peers.
- Absorption is now scarce. The same study found PR review time rising 91% and PR size growing 154%. The code arrives faster than the organization can responsibly accept it.
- Org-level outcomes stayed flat. Despite the individual-level surge, organizational DORA metrics — deployment frequency, lead time, change failure rate — showed no measurable improvement.
Read those three findings together and the conclusion is uncomfortable: adding production capacity (whether heads or agents) to a system whose constraint is absorption produces inventory, not throughput. The planning ritual keeps sizing the wrong resource.
The New Binding Constraint Is Review, Not Production
If tokens can produce arbitrarily more code, the question that actually determines your output next year is: how much machine-generated work can your organization verify, integrate, and stand behind?
Call this your absorption capacity. It is made of things the headcount spreadsheet has no column for:
- Senior engineers with enough context to review agent output against architectural intent, not just correctness. One 2025 study found seniors spending 4.3 minutes reviewing AI-generated suggestions versus 1.2 minutes for human-written code — the review is slower per unit precisely because the author cannot be interrogated.
- Decision-makers who can say yes or no to a design without convening a committee. Agents idle on ambiguity; every unresolved product question is a stalled fleet.
- CI, staging, and observability infrastructure that lets you trust a merge without a human tracing every path. LinearB's benchmarks found AI-generated PRs waiting 4.6x longer for a reviewer to even pick them up — a queue that is pure absorption debt.
Here is the ratio nobody has a benchmark for yet: reviewers per unit of agent-fleet throughput. Every org running agents at scale is discovering its own number empirically, usually by watching the PR queue back up. Some teams find one strong reviewer can absorb the output of three or four agent-heavy engineers; others find the ratio closer to one-to-one once the work touches shared infrastructure. The number varies with codebase modularity, test coverage, and how much architectural context lives in reviewable artifacts versus people's heads. What does not vary is that the ratio — not the headcount — is the ceiling on next year's output.
Why Finance Still Forces the Conversation into Headcount Boxes
The obvious rejoinder: if inference is now a production input, just budget for it. The reason this does not happen cleanly is that the two spends live in different financial universes.
Headcount is capitalized organizational knowledge. It comes with a requisition process, a compensation band, an approval chain, and a predictable annual cost. Finance has a century of tooling for it. Inference is an opex line that behaves like a cloud bill with a mind of its own — usage-driven, nonlinear, and uncapped by default. Deloitte documents a healthcare enterprise whose 8–10% monthly token growth compounded into more than $6 million of unplanned annualized cost in six months. CFO frameworks emerging in 2026 all converge on the same anxiety: the people deciding how tokens get spent are not the people accountable for the bill.
So the planning conversation defaults to the instrument finance can govern. Roadmap ambition gets converted into requisitions not because heads are the right unit but because heads are the legible unit. Meanwhile the actual capacity lever — the inference budget and the review structure around it — hides inside "cloud costs," ungoverned and unplanned.
The fix is not to make inference spending look like headcount. It is to give it the same planning citizenship:
- A token line modeled by role and seniority, the way benefits are modeled today. A staff engineer orchestrating agents has a materially different inference profile than a junior doing supervised work.
- A cost model finance can interrogate: average tokens per workflow × workflow volume × cost per token, adjusted for model mix — and re-forecast quarterly, because unit prices and consumption patterns both move fast.
- A defined business metric per agentic initiative. If engineering cannot say what an agent workflow costs per unit of outcome at scale, it is not ready for a budget line — the same discipline you would apply to any capital request.
The Line Items Are Now Substitutes — Sometimes
The provocative version of the shift is the direct comparison: one staff engineer at $450,000 fully loaded, versus $400,000 of inference plus the review capacity you already have. Framed that way, planning becomes a portfolio question rather than a hiring question.
But the substitution only holds under specific conditions, and pretending otherwise is how orgs end up with a token bill and a stalled roadmap:
- Inference substitutes for headcount when the work is well-specified, verifiable by existing tests and reviewers, and parallelizable — migrations, coverage expansion, integration glue, the long tail of well-understood tickets.
- Headcount substitutes for nothing when the work is deciding what to build, resolving ambiguity, owning an outcome, or carrying context across quarters. No token budget produces a person who can be accountable.
- The two are complements, not substitutes, at the boundary: every dollar of inference you add consumes review and decision capacity, which is bought with senior heads. Buy inference without absorption and you have bought a backlog.
This is also why the market data looks contradictory. PwC's numbers show companies most exposed to AI growing wages and productivity fastest, while new software engineering postings fell 15% year over year in early 2026, with junior and mid-level roles absorbing most of the decline. Both are the same phenomenon: orgs are trading production headcount for inference, while bidding up the people who expand absorption capacity.
How to Write Next Year's Plan
If you are drafting a plan this cycle, the shape that matches the new constraint looks something like this:
- Size your absorption capacity first. Count the engineers who can genuinely review agent-scale output in each domain, measure your current PR pickup latency, and be honest about where decisions bottleneck. This number — not roadmap ambition — is your real ceiling.
- Budget inference to saturate it, not exceed it. Work backward from absorption capacity to a token budget. If your reviewers are already the queue, more tokens buy you inventory. Spend the marginal dollar on absorption instead: test infrastructure, better specs, architectural documentation agents can consume.
- Justify every requisition by which constraint it relaxes. "We need three more engineers" is no longer a plan. "We need one staff engineer to expand review capacity in payments, and $300k of inference to saturate the capacity we already have in platform" is a plan. Some reqs will convert to opex under scrutiny; the ones that survive will be for judgment, ownership, and context — the inputs tokens cannot produce.
- Give the inference line the same review cadence as the hiring plan. Quarterly re-forecast, per-workflow unit economics, and a named owner. An ungoverned token budget is not agility; it is a headcount plan you are running by accident.
The orgs that struggle over the next few years will not be the ones that spent too much or too little on inference. They will be the ones that kept planning as if production were the constraint — hiring to type faster while their review queues quietly became the roadmap. The annual ritual survives; the equation inside it has to change. Output no longer scales with heads. It scales with the judgment you can bring to bear on what the machines produce — and that is the capacity worth planning for.
- https://zenvanriel.com/ai-engineer-blog/token-budgets-ai-engineers-compensation-model/
- https://www.deloitte.com/us/en/services/consulting/articles/cfo-guide-ai-token-economics.html
- https://www.forbes.com/councils/forbesfinancecouncil/2026/05/27/a-cfos-five-layer-framework-to-govern-ai-token-spend-before-it-governs-you/
- https://www.bcg.com/publications/2026/how-ceos-can-optimize-ai-token-costs
- https://www.pwc.com/gx/en/services/ai/ai-jobs-barometer.html
- https://blog.logrocket.com/ai-coding-tools-shift-bottleneck-to-review/
- https://blog.codacy.com/ai-breaking-code-review-how-engineering-teams-survive-pr-bottleneck
- https://addyo.substack.com/p/code-review-in-the-age-of-ai
- https://www.thesaascfo.com/a-cfos-guide-to-tracking-digital-labor-and-agentic-ai/
