NewPrimer ranked #1 in financial modelingRead analysis

Blog / 29 Jul 2026

Primer tops BigFinanceBench

We ran Primer against Rogo's BigFinanceBench using their own three-trial, two-judge protocol. Primer tops the leaderboard — 79.1% final-answer accuracy against 55.8% for the best frontier model — and hand-checking all fifty public questions surfaced errors in 12% of them.

For years, "AI can do finance" claims rested on benchmarks that asked for trivia: find last quarter's revenue, name the CFO, done. Models got pretty good at those and every leaderboard compressed into a two-point spread that told you nothing. Then Rogo went and raised the bar.

BigFinanceBench, released at the end of May, gets the idea right. 928 questions of real analyst work: modelling how a private equity fund's profits get split between its investors and its managers, rebuilding a company's divisional numbers after it changed how it reports them, working from a company's total valuation down to what the equity is worth. Each one graded on the working, not just the final number. Fifty questions are public, together with full traces and judge verdicts for ten frontier models. Almost nobody publishes that much, and it's the reason we could write everything below.

We ran Primer against it, using Rogo's own three-trial, two-judge protocol. And for all that Rogo has moved the industry forward on benchmark design, we found errors in 12% of the public questions: details below.

The results

Primer answers 79% of the benchmark's questions correctly, with an average rubric score of 79%. The strongest general-purpose model, GPT-5.5, manages 56% and 64% on the same basis. Against the average frontier model on the leaderboard, Primer's lead is roughly 38 percentage points.

A note on the counting. Fifty questions are public, but seven have faulty reference answers: five are demonstrably wrong, one has never been passed by any system, and one we dispute in part. All seven are documented below. So the figures above are scored on the remaining 43, with the same exclusions applied to every model. Leaderboard: Primer vs frontier models on BigFinanceBench, final-answer accuracy and rubric score

Why Primer outperformed and why the harness is so important

The obvious objection to the leaderboard above is that it isn't a fair fight, and that's correct. The models ran the benchmark's minimal scaffold while Primer ran a full production system.

That objection has a very clean test. At the time of testing, Primer's agent was powered by GPT-5.5, the same model that sits at the top of the frontier table above. Given the benchmark's own tools and left to work alone, GPT-5.5 scores 55.8%. Inside Primer, the same model scores 79.1%. Same brain, same questions, same judges. The model didn't change, so everything in that 23-point gap comes from the harness.

A harness is everything that stands between a raw model and a finished piece of work: what the model gets to read, which tools it can use, and what it has been taught about how the domain actually operates. Building that layer well is most of what we do at Primer.

The gap between generalised intelligence and domain-specific usefulness is exactly the product we're building at Primer. We've written about the hard work and effort we have put into our best-in-class retrieval accuracy and our focus on very high quality financial modelling capabilities, which are key drivers in the outperformance of Primer vs generalised LLMs.

What's interesting is that despite the releases of Fable 5, Opus 4.7 and GPT-5.6, we didn't see any of them beat GPT-5.5 in the benchmark, which further reinforces our view that if you want agents to do work to a standard you recognise in finance, you need a specialist harness.

Waterfall: GPT-5.5 to Primer, with model change −2.3, data & retrieval +16.3, calculation +9.3

The errors in the benchmark

While reviewing Primer's graded runs, we kept hitting questions where the working looked right and the grade said wrong. Our first assumption, every time, was that Primer had a bug. On one question Primer produced the identical answer seven runs in a row while we hunted for our mistake. Eventually we sat down and went through every question by hand.

See below for details of the errors; full workings for every one are in the appendix:

1. The Netflix question asks for marketing spend per net new subscriber. The rubric's own inputs are $1.3B of spend and 9.497M net additions, which works out to $136.74 per subscriber. The reference accepts only $0.14: a units slip, millions divided by thousands. Twenty-nine of the thirty published model trials say $136.74; none say $0.14.

2. The PE waterfall question describes a fund that earns $1B of profit with 20% carried interest. Twenty percent of $1B is $200M, and that is the most the managers can be paid under any waterfall convention. The reference pays them roughly $380M, because its formula counts money that was merely returned to investors as if it were profit.

3. The Shake Shack question asks what restaurant-level profit becomes if pre-opening costs fall and G&A rises. Neither line is part of restaurant-level profit as the company defines it, so the answer is the reported figure, unchanged. The reference moves it anyway, and is off by $0.65K even on its own reading.

4. The CoreWeave question asks for available liquidity at the quarter end. The company's own disclosure, less letters of credit, gives $6,480.8M, which is Primer's answer to the decimal. The reference reaches $8,113.3M by counting a credit-line increase that happened six weeks after the quarter ended, plus borrowing capacity the loan's own terms had already extinguished.

5. The Parsons question requires deducting noncontrolling interest from EBITDA. The rubric deducts $11.2M, a figure Parsons never disclosed as NCI. But the row directly next to NCI in the cash-flow statement, "payments for acquired warrants," is $11.2M. Someone copied the wrong line. The company's actual figure gives a different answer, which is what Primer submitted.

6. The Roku question asks for the gross-margin "uplift" after an adjustment. The rubric's own numbers show costs rising from 45.44% of revenue to 46.46%, so the margin fell, from 54.56% to 53.54%. Everyone agrees on the 102 bps; the reference just reads its own decline as an increase.

7. The Portillo's question (disputed in part) values payments under a tax receivable agreement. The contract pays holders 85% of the tax savings (the company's own filings say so), but the reference values 100% of them. Our answer is the reference's answer times 85%; the rest of the gap is input choices we hold more loosely.

8. The IDEX question is the strangest one. Nearly all of its points depend on matching three historical P/E multiples the question never explains how to construct. No model has ever passed it: not one of the ten Rogo published, not ours, not anything. 136 judge verdicts, zero passes. We don't think it can be passed.

Fifty questions is 5% of this benchmark, but it's the shop window: the subset every outside evaluation runs and every marketing deck quotes, ours now included. Additionally, BigFinanceBench is now included in the launch evals table on OpenAI's GPT-5.6 launch page, the only finance eval for the launch. If we all care about the integrity of benchmarks, we think these errors should be amended ahead of the next OpenAI model launch.

We should note that it's not clear whether OpenAI are using all 928 questions or just the public 50, or indeed the broader setup.

Benchmarks should reflect the judgement in finance

A number of these questions have more than one defensible construction, and the benchmark silently picks one. Basic or diluted EPS when the question just says "EPS." A stock's high over a window: daily closes or intraday extremes. Growth measured on the originally-reported base or the restated one. Total debt at principal or carrying value. And here the benchmark itself can't decide: one leverage question's reference uses principal outstanding, another's uses carrying value.

Ask two good analysts and you'll get two defensible numbers. In finance there is usually more than one way to skin a cat, and choosing between the ways, then saying which you chose and why, is the actual skill. Rogo understands this better than most. Their launch post says it plainly: "correctness in finance is often highly nuanced, and measuring that nuance matters." The rubrics are that principle made real. The final-answer gate just hasn't caught up with it: the gate accepts one construction and zeroes the rest, so part of what any leaderboard on this benchmark measures is whose conventions a model happened to guess.

The fix isn't complicated: accept the small set of standard constructions, or say which convention the question wants. We'd rather lose points to a stated convention than win them from a guessed one.


Appendix: the disputed questions, worked in full

For each: the question, the reference answer, our answer, the working, and what the ten published models said across their thirty recorded trials. This appendix contains every derivation needed to check the claims.

1. Netflix: marketing spend per net new US/Canada subscriber, FY2024

Reference: $0.14. Correct: $136.74.

The reference's own rubric lines record the correct inputs verbatim: "Identifies Netflix FY2024 US and Canada net new users as 9,497,000" and "Calculates Netflix FY2024 sales and marketing expense attributable to net new users in US and Canada as $1,298,606,205." Divide them: $1,298,606,205 ÷ 9,497,000 = $136.74. The rubric then accepts only "$0.14 (acceptable within +/- $0.01)", which is 1,298.6 ÷ 9,497: dollars in millions divided by members in thousands, the member figure rescaled as if it were already in thousands. As a sanity check, $0.14 to acquire a streaming subscriber is off by three orders of magnitude from any plausible customer-acquisition cost.

Field check: 29 of 30 published trials answered $136.74, across all ten models. None answered $0.14.

2. Private equity fund waterfall: LP net IRR and GP IRR

Reference: LP net IRR 12.88%, GP IRR 66.54%.

The question is self-contained math: $1B fund, 5% GP commitment, two $500M investments, each returning 2.0x at exit, 8% preferred return, 20% carried interest, European waterfall with GP catch-up.

The bound that decides the dispute: the fund invests $1,000M and returns $2,000M, so total profit is $1,000M. With a 20% carry rate, carried interest cannot exceed $200M; that is what 20% means, under any waterfall structure. The reference's rubric line "Calculates Year 5 GP Catch-Up $320,962,210" plus a further profit split pays the GP roughly $380M, an effective 38% take, because the formula counts money that was merely returned to investors as if it were profit inside the catch-up base.

We rest the accusation entirely on that bound. The correct answers legitimately vary with GP-participation convention: defensible constructions put the LP net IRR between roughly 16.1% and 16.7% and the GP IRR between roughly 45% and 50%, and the published models cluster in exactly that region. No model in 30 trials produced 12.88% or 66.54%.

3. Shake Shack: restaurant-level profit if pre-opening costs fall 60% and G&A rises 5%

Reference: $259,749.2 thousand. Correct: $257,874.0 thousand, unchanged from reported.

In Shake Shack's own non-GAAP reconciliation, restaurant-level profit is Shack revenue less Shack-level operating expenses; pre-opening costs and G&A sit below the line and are excluded from the measure by construction. Change them however you like and restaurant-level profit, as the company defines it, does not move.

Two supporting observations. First, the reference itself implicitly concedes the G&A half of this: it flows the pre-opening saving into the metric while G&A's increase changes a metric that G&A isn't part of either. Second, even granting the flow-through reading in full: 60% × $15,547K − 5% × $149,047K = +$1,875.85K, giving $259,749.85K; the reference says $259,749.2K, internally off by $0.65K within its own construction.

Field check: the modal published answer is $257,874K unchanged (five models, plus Primer). No model matched the reference except one trial within rounding.

4. CoreWeave: available liquidity

Reference: $8,113.3M. Correct: $6,480.8M.

The reference's rubric builds "(1,894.399 + 47.449 + 1,800 + 4,632.479 − 261.0) = $8,113.327 million", a sum that includes (a) a revolver upsize announced November 10, roughly six weeks after the September 30 balance-sheet date, and (b) $632.5M of a delayed-draw term loan tranche whose availability had already amortized away under the facility's own repayment schedule (its credit agreement's prepayment mechanics, §2.09, are why those commitments don't spring back). The company's own liquidity disclosure as of the balance-sheet date, less outstanding letters of credit, is $6,480.8M.

Primer answered $6,480.8M, the strongest kind of receipt, because the "wrong" answer being graded down is the company's own disclosure.

5. Parsons: pro-forma leverage with EBITDA net of NCI

Reference chain: $345.3M LTM EBITDA − "$11.2M NCI" = $334.0M. Correct: $1,826.3M max debt capacity / $1,226.3M incremental / 2.7x.

The rubric's "Record PSN's LTM NCI as of Q3'22 = $11.2M" does not match Parsons's disclosed NCI. It exactly matches the adjacent row in the cash-flow statement: Payments for acquired warrants, $11,243K. The provenance is arithmetic: 345,251 − 11,243 = 334,008, the rubric's own $334.0M to the digit. Netting the company's actual NCI produces the chain above, which is what Primer's original benchmark run answered, to the decimal, on all three figures.

6. Roku: Platform gross-margin change after adjusting for restructuring

Reference: +102 bps "uplift," with adjusted "margins" of 45.44% and 46.46%. Correct: a decline of 102 bps, margins 54.56% → 53.54%.

Work from the reference's own rubric values. It puts FY2023 adjusted cost of revenue at $1,360,514K and FY2024 at $1,636,816K. Set those against Platform revenue and you get cost-of-revenue ratios of 45.44% and 46.46%; the reference's two "margins" are cost ratios. Gross margin is the complement, and it fell. Everyone agrees on the magnitude, including the reference; the disagreement is purely which ratio the word "margin" names, and standard usage is not ambiguous.

Rogo's own site makes the point for us. The homepage's "Example 2" showcase is this Roku question, and the top-scoring exemplar answer displayed there reads: "the adjusted Platform cost ratio rises from 45.44% (FY23) to 46.46% (FY24) — an uplift of ~102 bps." The showcase answer calls the figures what they are, cost ratios, and still books the rise as an uplift. A rising cost ratio is a falling margin; the site's own front page concedes the naming and repeats the direction error in the same sentence.

Field check: eight of ten models report the change as a decline. No model reports an uplift as its computed result.

7. Portillo's: five-year present value of Tax Receivable Agreement payments (partially disputed)

Reference: $16.9M. Ours: ~$14.4M.

Tax receivable agreements standardly obligate the company to pay pre-IPO holders 85% of realized tax savings, and Portillo's own MD&A states it in both words and numbers: "…approximately 85% of such amount, or $353.2 million." The reference's rubric values 100% of the gross tax shields ($16.9M); applying the contractual payout rate gives ~$14.4M. That factor we claim firmly. The remainder of any gap is discounting and input choices we hold loosely; the modal published answer, $13.4M, uses yet another defensible chain. No model in 30 trials produced $16.9M.

A note on IDEX

We exclude one further question from our verified set without calling it an error: a P/E-band valuation whose rubric gates 33 of 36 points behind three multiples (24.15x / 28.27x / 35.87x) with unstated construction conventions, and whose instruction ("the latest FY2025 guidance… from the Q2'25 report") contradicts itself once later guidance exists. It has never been passed by any model, any trial, any judge: 136 verdicts, zero passes. A question no system can pass isn't measuring the systems.

A note on Chipotle

One more question earns a note not because its reference is wrong but because of what grading did with it. The question projects Chipotle's FY2025 food & beverage revenue assuming 15% more openings than 2024 and constant revenue per location. Primer computed it at full precision: 304 openings × 1.15 = 349.6, rounded conservatively to 349; 4,075 total locations × $3,018,621.58 per location = $12,300,882,930.76. The reference's $12,300.90M comes from pre-rounding revenue to $11,247.4M before dividing. At the precision of the reference answer itself, $12,300.9M, the two are identical; the residual exists only because the reference rounded an input before using it. One judge passed Primer's answer and the other docked the two-decimal intermediate: a $20,000 artifact on a $12.3 billion answer, created by the key's own rounding, splitting the verdict depending on which judge you drew. It's the final-answer gate at its most literal: grading against a rounding artifact the reference itself introduced.