Blog / 23 Aug 2026
Enterprise AI rollouts are being measured wrong
What a Claude enterprise rollout actually costs, once you consider how much time employees will spend correcting answers.

Anthropic has won the enterprise race but the Claude rollouts are being judged completely wrong.
Nobody hires anybody purely based on their hourly rate. When hiring, you take 3 key things into consideration: how much they cost, how good their work is, and how much of that quality work they can complete. Two people on the same salary can have wildly different impacts on your business.
However, that is not remotely how anyone is buying AI. The evaluation once rolled out seems to be entirely based on price per token. Which has no consideration for the quality of work delivered and how much of this they can produce.
Senior management and board level reports only include their Anthropic bill and departmental token usage, at no point are firms trying to quantify the value added by AI, which is baffling as every other internal project will have an ROI target or expectation attached to it.
The chart at the top of this page is what it looks like when you do measure it properly. Every system below was run against BigFinanceBench, a set of finance research questions written by analysts. The horizontal axis is not simply the token bill, it's what a correct answer really costs once you add the time somebody spends checking the work and redoing whatever came back incorrectly. Primer is our own equity research system, and I'll come back to why it sits where it does.
The rest of this article is how you build that chart, and why the usual way of judging these rollouts misses it.
Anthropic appear to have won in enterprise
Ramp tracks corporate card spend across tens of thousands of businesses, and their numbers have Anthropic winning about 70% of head-to-head deals vs OpenAI when a company buys AI for the first time.
In finance, I'd say it's much closer to 100%, speaking from my own experience of speaking to the world's leading banks and asset managers.
The rollouts and adoption are being held up due to the bills they generate
I've had the same conversation with several banks and large asset managers over the past few months. Each of these institutions rolled out Claude this year and despite great intentions of higher productivity and efficiencies when the project was proposed to an internal committee, all those committees can talk about now is the bill they're getting on a monthly basis.
Token use per person came in at multiple times the expectation of what procurement had modelled, which has led these institutions to be stuck between a rock and a hard place of encouraging more usage but at the same time trying to avoid that famed $500m accidental bill.
The purpose of this article is to give some context on how financial institutions should approach evaluating models and tools as they go through their transformation.
Focus on what a correct answer costs
Evaluating models and tools should come down to a simple question: what is the trade off between getting the right (or at least helpful) answer and how much it cost to get there.
BigFinanceBench is Rogo's finance benchmark: 928 finance research questions written by experts. And in our view, this is the current gold standard within the finance industry. It isn't perfect though: seven of its 50 public questions have reference answers that are provably wrong or that no system has ever passed, something I've written about before, so everything below runs on the remaining 43.
The below chart shows both the accuracy scores of key foundational LLMs, plus Primer, and their respective cost per correct answer. Counting tokens only for now; the checking time comes later.
Claude costs the most per correct answer, and it isn't the most accurate
Claude is at the expensive end and the middle of the pack from an accuracy perspective. Fable 5, the flagship, costs five times what GPT-5.5 costs per correct answer and scores lower than it. On these 43 questions every Claude model we tested is both less accurate and dearer per correct answer than a GPT model, which is a fairly unusual place for a premium product to end up. Whatever the premium is for, it isn't showing up here.
I've left out nine cheaper models Rogo also scores, all of which beat everything here on cost per correct answer. They're out because nobody is rolling out Gemma 4 31B to research departments. I'll come back to that later, when that cheapness gets a proper test, and it doesn't go well for them.
Every Claude model is beaten on price and accuracy at once
All 28 models Rogo scores, plus Primer, on their published cost per question. Cheaper is further left, more accurate is higher.
Opus 5 is the best Claude model yet, and a GPT model gets within two points of it for a ninth of the cost. Plot the whole field and the position is stark: of the 28 models Rogo scores, not one Claude model sits on the efficient frontier. Every one of the six is beaten on price and accuracy at the same time, Sonnet 5 by a dozen other models.
Principal vs agent problem in LLM usage
Why does the same class of model cost four or five times as much to reach roughly the same place? It all comes down to token volume. They simply use more and the reason nothing stops them doing this is a textbook principal-agent problem, running three layers deep.
Start with the analyst. They type the question and somebody else pays for the tokens, so there is no reason for them to economise, and no way for them to do so even if they wanted to. Nothing in the interface tells them that this question will cost forty cents and that one will cost nine dollars.
Then the agent itself. It decides how many searches to run, how long to think, how many times to re-read the same filing. Nothing in how these systems are trained or scored tells them any of that has a price. Quite the opposite: they are graded on whether the answer was good, so more effort is always the safer choice. The agent is spending your money with no information about your budget and no reason to care.
And then the vendor (in this case, Anthropic), which is the layer that really matters. Their revenue is the tokens. An agent that answered your question in half the searches would halve that line of their P&L. I am not suggesting anyone sits in a room deciding to run up your bill; the truth is duller and rather worse than that. Efficiency is simply not the thing anybody is being rewarded for. Every model launch is judged on accuracy, every leaderboard ranks accuracy, and thinking longer generally scores better. So the whole industry has spent two years optimising hard on the one variable that also happens to increase what you pay, and not one vendor publishes cost per correct answer. I had to work out every figure in this piece myself, from a benchmark author's data, because nobody selling these systems reports it (obviously).
Every incentive in the chain points the same way, and it is not in favour of your bank or fund's finance department. The analyst can't see the meter, the agent has no reason to be careful, and the vendor operating this for you couldn't be less incentivised to reduce token spend.
You can see it in the token counts
Here's the proof, using two models that charge exactly the same per token. Same price list, same harness, same questions. One of them spent two and a half times more than the other getting through them.
When the price per token is identical, the whole difference is how much each one decided to use. GPT-5.5 isn't priced any higher and GPT-5.6 Sol isn't cheaper per token, it's just better at getting to the answer faster.
That pattern is everywhere; Claude Sonnet 5 is a tier below Opus 5 and costs less than half as much per token, yet it costs more per question, $1.84 against $1.71. Sonnet 5 burned roughly two and a half times the tokens that Opus 5 did to answer the same questions. Fable 5 is the priciest model per token on the board and only the third priciest per question, because Sonnet 5 outspends it on sheer volume.
Claude models burn far more tokens on the same question
Estimated from each model's published cost per question and its list price per token.
Those token counts are estimates but the spread is the best representation of this phenomenon: about 25x more tokens between the most and least frugal, on the same benchmark with the same tools available.
Don't forget that a wrong answer isn't free either
There is one assumption buried in every cost figure I have quoted so far, and in every price list any of these vendors publishes: a wrong answer costs nothing.
Obviously it doesn't cost nothing, so we're going to put a number on it.
Think about what actually happens when one of these systems hands an analyst a wrong figure. You can't tell a right answer from a wrong one by looking at the confidence of the prose, so somebody has to check it against the filing. That takes time whether the answer was right or not. Then the wrong ones get done again. None of that effort is in the token bill, none of it is on a leaderboard, and all of it is paid for in the salary of the analyst and ultimately the opportunity cost of their time.
So let's attach a real price: let's assume the average analyst is paid $150,000 a year, works 12 hours a day, 5 days a week and takes 20 days holiday a year. This means they are broadly paid $52/hour.
For every wrong answer, it takes them half an hour to fix it, $26. We can argue until the cows come home about the cost per hour and how long it takes to fix an answer, but it doesn't make much difference to the ultimate calculation.
The first panel below shows the token cost only, but the second panel shows the opportunity cost of lower accuracy combined with the token cost.
Pricing the wrong answers reverses the ranking
A wrong answer priced at $26, being half an hour of an analyst on $52 an hour.
The order flips completely. Everything that looked like a bargain on tokens is expensive the moment a person has to read it and correct it. Gemma 4 31B costs six cents a correct answer on tokens but a whopping $67 all in.
Every Claude model here loses to the GPT models even if the cost of checking an answer is zero, because they cost more per question and get less right.
85% more work, same analyst
What actually limits a research team is analyst hours and human capital, not the AI budget. So you need to be asking what an LLM or system gives you in the only unit that counts, which is finished, accurate work per person per week.
If we assume 60 hours a week, 5 mins to check an answer that's right and half an hour to catch and fix one that's wrong, this is what a week of finished work looks like.
What one analyst gets through in a week
A 60 hour week, five minutes to check an answer that is right, half an hour to catch and fix one that is wrong.
This is the number that's actually important for your internal committee. The same tasks and the same time spent doing them: you'd get 352 finished answers with Primer, 217 with Opus 4.7, 190 with Sonnet 4.6. You'd get 57% more work from Primer than you would from GPT-5.5, and a massive 85% vs Sonnet 4.6.
Detail on why Primer outperforms is here. Ultimately it comes down to specialised tools doing specialist tasks, and third party vendors are naturally aligned with the buyer from a token use perspective too.
What this all means
None of this means the frontier models are bad. Opus 5 tops its family and the capability is real. It's just being sold on one measure while your business runs on another.
The fix isn't complicated. Make sure you scope out a series of tasks, mark them, and then use accuracy as a way to estimate how much time it would take to correct the wrong answers. This is the number you should be including in the internal report to establish whether you've made the right choice.