
Enterprise AI rollouts are being measured wrong
What a Claude enterprise rollout actually costs, once you consider how much time employees will spend correcting answers.
Blog
A collection of our thoughts on AI, equity research, and the analyst job.

What a Claude enterprise rollout actually costs, once you consider how much time employees will spend correcting answers.
I read all 220 questions in FrontierFinance. Nearly three quarters of the answer keys have errors.

We ran Primer against Rogo's BigFinanceBench using their own three-trial, two-judge protocol. Primer tops the leaderboard — 79.1% final-answer accuracy against 55.8% for the best frontier model — and hand-checking all fifty public questions surfaced errors in 12% of them.
Earlier this year we shared our results on the UK, US and EU portion of Daloopa’s FinRetrieval benchmark. We’ve now run the whole thing: all 500 questions, every region, and got all of them right. Getting there taught us more about benchmarks than it did about AI.
OpenAI just released GPT‑5.6. Before it could supersede the 'model in situ' at Primer, it had to beat it. Having tested 5.6 on real analyst work, we see a clear step up in quality from GPT‑5.5, and it comes at no extra cost. This is how we tested it, and what we found.

They say you’re the average of the five people you spend the most time with. If that’s true, we’re all slowly becoming a weighted average of our AI agents.
Wall Street Prep tested AI agents on a realistic financial modeling task. We added Primer, then used AI judges to score the actual workbook artifacts side by side.

For fundamental analysts and PMs deciding where to spend time and tooling budget, this comparison matters because Step 1 and Step 2 are different jobs. If you are comparing Primer and AlphaSense, the practical split is simple: AlphaSense is strongest at Step 1, finding and scanning information across a broad universe, while Primer is strongest at Step 2, developing, testing, and maintaining an investment idea until it is actionable.

The first test for any financial AI system is straightforward: can it retrieve the exact number from source disclosures, with the correct period and unit? Daloopa’s published FinRetrieval results placed the public frontier around ~90% retrieval accuracy; on our covered UK/US/EU subset, we reached 100% retrieval accuracy. As a result, the debate and evaluation framework should now shift from raw accuracy to how useful agents are in real workflows.

We are drowning in AI tools that can read, but starving for AI that can actually reason.

How to use it without fooling yourself

Most Equity “Research” Is Table-Stakes Work. Edge Is Choosing What to Do Next. Alistair Smallwood