Evals
How Primer actually performs.
Equity research is an output-quality problem, so we measure Primer on the work itself: retrieving the right numbers, building models, following a real methodology, and completing analyst tasks end to end.
01
BigFinanceBench
Can AI complete real analyst work from source to answer?
Read the full analysisA benchmark of real finance work
BigFinanceBench has 928 finance questions. Only 50 are public: private equity waterfalls, segment restatements, valuation bridges, and other real analyst work.
How each score is measured
Frontier scores are the official BigFinanceBench leaderboard: final-answer accuracy across all 928 questions, led by Muse Spark 1.1. Primer is measured on the public set - we found errors in seven reference answers, excluded them, and scored 79.1% answer accuracy across the remaining 43.
Why Primer outperformed
Primer runs on GPT-5.5 - on its own it scores 44.3% on the leaderboard; inside Primer, 79.1%. The gap is the harness: what the model reads, which tools it can use, and what it has been taught about finance.
02
Data retrieval
How accurately can AI pull data?
Read the full analysisA 500-question retrieval test
The benchmark tests whether an AI agent can find the right filing and pull the correct number. Right company, right period, right unit.
Why accuracy matters
Retrieval is the foundation of research. If the data is wrong, the analysis isn't useful.
Why Primer outperformed
Primer reads the full filing, not isolated snippets, and reasons like an analyst. Company, period, unit, and source context stay intact.
03
Financial modelling
Best-in-class modelling capability.
Read the full analysisA real analyst task
The Wall Street Prep benchmark asked AI tools to build Apple's full three-statement model: historicals, forecasts, assumptions, sources, comments, and supporting schedules.
Judged like an actual financial model
The review included a deep review of the formulas, links, structure, formatting, comments, and diagnostics. Not just whether the output looked complete.
Why Primer outperformed
Primer was built by analysts who understand modelling, not just Excel automation. Instead of forcing the agent to work cell by cell in Excel, Primer uses state of the art modelling infrastructure, then converts the finished model to Excel.
04
Underlying cash flow
Can AI follow a real methodology to understand underlying cash conversion?
The task
Each AI was handed the same specific methodology for getting to a company's underlying cash flow. Strip out working capital noise, accounting distortions, and one-offs. End at sustainable free cash flow to equity.
Judged on justification, not just numbers
Every adjustment needs evidence and a clean bridge from P&L to cash. Invented figures or unsupported one-offs fail outright.
Why Primer outperformed
Primer's harness is purpose built to work like a good analyst. It knows what counts as evidence, when an adjustment is justified, and how to show its reasoning.
05
Earnings previews
That actually help you generate alpha.
What we tested
Whether the earnings previews Primer writes are alpha generating. We took a long position when the score was strongly positive, a short position when strongly negative, and held for 14 days.
The result
34% alpha across 2,727 calls over 203 trading days, at a 53.6% hit rate. Each trade was matched against notional SPX over the same window.
Why Primer outperformed
A clear demonstration of a best in class agent harness. Primer's previews reflect real analyst judgement, so the strongest calls carried genuine directional signal.
Alpha generated following Primer's preview call over the last 203 trading days.
The backtest didn't just beat the S&P 500; it did so with risk-adjusted consistency.