Evals

How Primer actually performs.

Equity research is an output-quality problem, so we measure Primer on the work itself: retrieving the right numbers, building models, following a real methodology, and completing analyst tasks end to end.

01

BigFinanceBench

Can AI complete real analyst work from source to answer?

Read the full analysis

A benchmark of real finance work

BigFinanceBench has 928 finance questions. Only 50 are public: private equity waterfalls, segment restatements, valuation bridges, and other real analyst work.

How each score is measured

Frontier scores are the official BigFinanceBench leaderboard: final-answer accuracy across all 928 questions, led by Muse Spark 1.1. Primer is measured on the public set - we found errors in seven reference answers, excluded them, and scored 79.1% answer accuracy across the remaining 43.

Why Primer outperformed

Primer runs on GPT-5.5 - on its own it scores 44.3% on the leaderboard; inside Primer, 79.1%. The gap is the harness: what the model reads, which tools it can use, and what it has been taught about finance.

BigFinanceBenchAnswer accuracy
Primer79.1%
Muse Spark 1.153.4%
Claude Opus 546.1%
Claude Fable 545.2%
GPT-5.544.3%

02

Data retrieval

How accurately can AI pull data?

Read the full analysis

A 500-question retrieval test

The benchmark tests whether an AI agent can find the right filing and pull the correct number. Right company, right period, right unit.

Why accuracy matters

Retrieval is the foundation of research. If the data is wrong, the analysis isn't useful.

Why Primer outperformed

Primer reads the full filing, not isolated snippets, and reasons like an analyst. Company, period, unit, and source context stay intact.

Data retrievalAccuracy
Primer100%
Claude Opus 4.5 + Daloopa MCP90.8%
Gemini 3 Pro + Daloopa MCP90.6%
GPT-5.2 + Daloopa MCP89.2%
ChatGPT 5.2 (Web)70.8%
Gemini 3 Pro (Web)69.2%
Claude Opus 4.5 (Web)19.8%

03

Financial modelling

Best-in-class modelling capability.

Read the full analysis

A real analyst task

The Wall Street Prep benchmark asked AI tools to build Apple's full three-statement model: historicals, forecasts, assumptions, sources, comments, and supporting schedules.

Judged like an actual financial model

The review included a deep review of the formulas, links, structure, formatting, comments, and diagnostics. Not just whether the output looked complete.

Why Primer outperformed

Primer was built by analysts who understand modelling, not just Excel automation. Instead of forcing the agent to work cell by cell in Excel, Primer uses state of the art modelling infrastructure, then converts the finished model to Excel.

Financial modellingScore
Primer81%
Shortcut.ai50%
Claude in Excel48%
ChatGPT43%
Microsoft Copilot40%

04

Underlying cash flow

Can AI follow a real methodology to understand underlying cash conversion?

The task

Each AI was handed the same specific methodology for getting to a company's underlying cash flow. Strip out working capital noise, accounting distortions, and one-offs. End at sustainable free cash flow to equity.

Judged on justification, not just numbers

Every adjustment needs evidence and a clean bridge from P&L to cash. Invented figures or unsupported one-offs fail outright.

Why Primer outperformed

Primer's harness is purpose built to work like a good analyst. It knows what counts as evidence, when an adjustment is justified, and how to show its reasoning.

Underlying cash flowScore
Primer81.6%
Anthropic Fable 5 Max76.3%
OpenAI GPT-5.5 Ext. Thinking68.4%
Anthropic Opus 4.8 Max55.9%
Grok SuperGrok Heavy18.8%
Gemini 3.5 Thinking10.3%

05

Earnings previews

That actually help you generate alpha.

What we tested

Whether the earnings previews Primer writes are alpha generating. We took a long position when the score was strongly positive, a short position when strongly negative, and held for 14 days.

The result

34% alpha across 2,727 calls over 203 trading days, at a 53.6% hit rate. Each trade was matched against notional SPX over the same window.

Why Primer outperformed

A clear demonstration of a best in class agent harness. Primer's previews reflect real analyst judgement, so the strongest calls carried genuine directional signal.

Earnings previewsBacktest
+34%Alpha vs SPX

Alpha generated following Primer's preview call over the last 203 trading days.

1.42Sharpe ratio

The backtest didn't just beat the S&P 500; it did so with risk-adjusted consistency.

Benchmark-relative cumulative P&L
SepNovJanMarMay