NewPrimer ranked #1 in financial modelingRead analysis

Blog / 04 Aug 2026

Finance AI benchmarks are misleading

I read all 220 questions in FrontierFinance. Nearly three quarters of the answer keys have errors.

Audit of all 220 FrontierFinance answer keys: 57 clean, 88 with minor issues and 75 materially defective

This is the third finance benchmark that I've looked at in detail, and the third to contain a worrying number of errors.

In Daloopa's FinRetrieval, 2.2% of the published answers were wrong, and most of it was mechanical: transposed digits, a decimal comma read as a decimal point, a figure keyed after a reverse split when the filing reported it before. Careless rather than confused.

BigFinanceBench was worse, with problems in 12% of the public questions, and none of those were typos. Reference answers contradicted their own inputs, and one carry calculation paid managers more than the fund had earned, which is the sort of thing that would fail an internship task. It isn't sloppy transcription so much as poor analysis, published as ground truth.

So when Samaya published FrontierFinance, a 220-question financial research benchmark with a leaderboard showing their fine-tuned LLM beating frontier models at a fraction of the cost, I did what felt most natural having been a hedge fund analyst and that was to actually read the questions.

Samaya designed FrontierFinance to cover the investor workflow end to end. Its questions are split across six categories: screening and discovery; company research; financial data and modelling; earnings and events; coverage and catalyst monitoring; and sector, industry and macro analysis.

Two other design choices matter even if every answer key were perfect. Each system is ranked from a single stochastic run, with no repeats or error bars, and the research agent can use 2026 web results to answer questions dated years earlier. The leaderboard can therefore mistake run-to-run noise for a performance gap and hindsight for research skill.

What a benchmark is

A benchmark is a set of questions put to a language model, or to a full system built around one, with a fixed way of marking the answers. Every domain and every task type has its own, usually several, because a benchmark that measures coding ability tells you nothing about legal drafting or medical triage or pulling a number out of a 10-K. So they proliferate, and anyone launching a product can generally find or build one on which their product does well.

The marking is sometimes deterministic, where the question has one exact figure as its answer and you either produce it or you don't. For open-ended research questions that doesn't work, so most of these benchmarks mark against rubrics, also called keys: a list of statements the ideal answer is expected to contain, written in advance by whoever built the benchmark.

FrontierFinance is key-based. Each of its 220 questions comes with a published key, 11,543 rubric items in total and up to 300 on a single question, with each item flagged either as a must-have or as one of the nice-to-haves. Three AI judges vote on every item, majority wins, and the headline score is the share of items you satisfied.

The key is essentially an exam mark scheme, but there are 2 things worth flagging with this type of system:

  1. Nobody is judging whether your answer is any good, only whether it matches a list written in advance. If the list is wrong, a right answer loses marks.
  2. You score only for what's on the list. Being right about something the marker never thought of is worth nothing.

So the benchmark doesn't necessarily reward the best answer to the question. It rewards the closest copy of one analyst's answer to it.

That is the fundamental difficulty with benchmarking finance, particularly equity research tasks. Retrieval and modelling are exact enough to mark: either you found the right figure in the filing or you didn't, either the model ties or it doesn't. Everything past that is judgement, and two good analysts covering the same stock will reach different conclusions and disagree about which facts mattered.

The audit

I read all 220 questions and their answer keys by hand, then ran several more audit passes over them with different LLM setups to be thorough.

Each key went into one of two buckets: does it hold up, or is there something wrong with it. Across the 220:

  • Clean, 57 keys (26%). The key holds up and a right answer passes.
  • Broken, 163 keys (74%). Something in the key is wrong: a figure, a period, a list, a contradiction, or the thing being graded.

Within the broken pile, 75 keys are bad enough that a correct answer doesn't just lose a few points, it fails.

Grid of all 220 answer keys: 57 clean, 88 with minor issues and 75 materially defective

The errors are not evenly spread across those six categories. Financial data and modelling has the most in absolute terms, with 55. Screening and discovery has the highest rate: every one of its 17 keys has something wrong with it.

Materially defective answer keys by benchmark category, shown as both absolute counts and percentage of each category
Financial data and modelling has the most defective keys; screening and discovery has the highest defect rate.

Across those 163 keys, the errors themselves fall into six recognisable families.

Bar chart of the six defect families across 163 non-clean answer keys
Primary defect family for every non-clean key.
  • Unknowable company list: the question asks something open like "which companies are exposed to X", and the key is a fixed list of tickers that the question isn't specific enough about.

    "Which ticker symbols are associated with the top 10 publicly traded quantum computing companies?" Nothing in the question defines "top". Ten of the eleven rubrics require one specific list, and that list leaves out IonQ while including Atos, which was in the middle of a debt restructuring at the time.

  • Opinion keyed as fact: one analyst's framing of a debate, or one hand-built model's assumptions, required as if they were figures in a filing.

    A Norwegian Cruise Line question runs to 199 rubrics, and 191 of them reproduce a single hand-built debt model cell by cell, arbitrary assumptions included. Build a defensible different model and you fail nearly all of it.

  • Key grades the wrong thing: the question asks for one thing and the key marks another.

    "What is Molson Coors' volume by brand and geography for each quarter since 2019?" Fifty-two of the 53 must-have data rubrics grade net sales rather than volume. The company discloses the volumes the question asks for. The key just doesn't use them.

  • Self-contradiction: rubrics within a single key that cannot all be true at once.

    On a question about how payrolls move S&P futures, rubric 6 requires you to say expectations of higher earnings drive the index up. Rubric 14 requires you to say the correlation between earnings growth and stock prices is negative. Both are must-haves.

  • Wrong figures: numbers in the key that don't survive contact with the key's own arithmetic.

    An ATM market-share question requires you to state that NCR Atleos holds 27% of the global market and Diebold Nixdorf 32%. The same key puts that market at $25.29bn, which would make those two companies' ATM revenue about $15bn between them. It also states their actual segment revenues: $2.6bn each.

  • Wrong period: the key marks the right answer against the wrong dates.

    A question asks for Simon Property Group's debt maturity profile "as of the latest reported quarter", dated 9 October 2025. All 64 rubrics grade the quarter ending 31 March 2025, two quarters stale by then.

None of these needed special access or clever tooling. They just needed somebody to actually have a look at the questions.

My favourite are two questions which are essentially the same, about the same company, but have different answer keys. Both ask how Marqeta makes money and what its unit economics are, and the only stated difference between them is a cut-off date three days apart. Nothing happened to Marqeta in those three days, and neither key refers to anything that did: both are built on full-year 2024 figures and earlier.

One of them was marked against 8 rubric items. The other was marked against 44, of which 30 are a six-year table of figures that neither question asked for. We ran the same system on both and it scored twelve points higher on one than the other.

In fairness, the questions are not word-for-word identical. The longer-keyed one ends with "include details from the S-1 document at the time of their IPO if needed". That explains a handful of the extra rubrics, but it doesn't explain why there are thirty rows of 2019–2024 financials, none of which appear in the S-1.

Auditing the auditor

An audit like this has an obvious weakness, especially coming from a competitor: maybe I just marked the questions as wrong for the sake of it. So I did to my own audit what I'd want anyone to do to a short thesis, and I did it twice.

In a piece I recently wrote on Substack, I found that teams of AI agents beat a single model at forensic work, but only with one specific structure. Findings have to be developed into full arguments before anything judges them, because as one-liners the right answers lose to the scandal-sounding wrong ones. I reused that structure here, as a courtroom.

The blind audit protocol: prosecution, defence and judgment
Each case was argued both ways before a verdict was recorded.
  • Prosecution: one agent builds the strongest case that a key is defective, quoting specific rubrics.
  • Defence: a second agent gets the case file and the prosecution brief, and builds the strongest honest case that the key is fine. Steelman it: maybe the disputed period is what the question implies, maybe the company list is derivable from the question after all.
  • Judge: a third agent reads both developed arguments, verifies the quotes against the actual key, and rules on one standard only: would a correct answer to the question as written fail this key?

Two controls made it a real test rather than a rubber stamp. It ran blind: I shuffled the 73 keys I thought were broken together with 20 I thought were fine, and the adjudicator saw no labels and no notes. If my audit was just prejudice, the clean controls would get convicted at the same rate as the flags. And it ran twice, on rival labs' models — once on OpenAI's GPT 5.6 Sol through Codex, once on Anthropic's Claude Opus 5 — with the same protocol, the same blind inputs and no shared reasoning between them.

The verdicts

The chart also includes arbitration. That means a third model read the question, the key and both original verdicts, then made its own call on two separate questions: is there a problem, and would that problem make a correct answer fail? It completed 92 of the 93 cases.

Comparison of two blind adjudications showing agreement on whether keys are flawed and disagreement on whether the flaws are severe
The two audits agreed the flagged keys had problems, but drew the line between flawed and fatal very differently.

The two audits agreed on the basic finding. GPT found a problem in 67 of the 73 keys I had flagged; Claude found one in all 73. Both independently found problems in 67 of 73 (92%). Among the 20 controls, both found problems in only five (25%).

Then they diverged.

GPT said 54 of the 73 were bad enough to make a correct answer fail. Claude said 29. They agreed on only 26. The arbitrator landed between them, calling 39 of the 72 completed cases materially defective.

I reran 20 cases unchanged to see whether this was random noise. Claude repeated its decision on 19 of 20; GPT did so on 11 of 20. The models are not simply making random calls. They draw the line between "flawed" and "fatal" in different places.

That leaves two conclusions. First, the core audit survives: two rival vendors' models, working blind, found problems in the flagged keys far more often than in the controls. You do not have to take a competitor's word for it.

Second, the grading is less precise than the leaderboard suggests. FrontierFinance makes 11,543 similar AI judgement calls and reports a single run to one decimal place. If strong models disagree this much on a clearer 92-case check, a 1.6-point leaderboard gap should not be treated as a precise ranking.

You can see it in the scores

Samaya did not publish per-question results, but it did publish scores for each of the six question categories. Because my audit grades every question, I could re-aggregate the errors into those same categories. My hypothesis was simple: systems, including Samaya's, should score worse in categories where a higher proportion of answer keys contain errors.

Samaya scores decline as the share of materially defective answer keys rises
All eleven leaderboard systems score worse in categories with more materially defective answer keys

Every one of the eleven published systems scores worse in the categories where more of the answer keys are broken. Eleven out of eleven, across two different test harnesses and five model families, with no exceptions.

The steepest and cleanest relationship belongs to Samaya's own system, a linear correlation of -0.96. Their best category, earnings and events, is the one where the fewest keys have errors at 44%. Their worst, screening and discovery, is the one where every key does.

LLMs are stochastic and must be run multiple times

Another limitation of the benchmark is that its published leaderboard gives no sense of run-to-run variance. Language models are stochastic: give the same model the same question, settings and tools twice and you are not guaranteed the same answer. In a research agent those small variations compound: one run searches a different phrase, finds a different filing first, follows a different branch of the research, or simply decides that a borderline fact is worth including.

We saw roughly ten rubric decisions move when we reran the same questions under the same setup. The system had not become better or worse. It had rolled the dice again.

Then the answer is handed to three more models, which make a separate set of judgement calls about whether each sentence satisfies each rubric. Majority voting reduces that noise; it does not remove it.

This makes repeat runs imperative, not optional. A single score is one draw from a distribution. A credible leaderboard should run each system several times and publish both the average and the variance: how widely its scores move around that average.

This leaderboard instead presents one run per system to one decimal place. Samaya's system leads the next system by 1.6 points. There are no repeat runs and no error bars. A gap that small may be a real difference, but this leaderboard has not shown that it is. Until the variance is visible, the honest unit is a range, not a rank. The first and second place badges are trying to tell us which runner is faster from one photograph of the finish line.

Harry Hindsight is a very rich man

The bigger problem is hindsight. Every question has a date, which should define the information available to the analyst. But the open-source harness behind the leaderboard gives the model live web access and hard-codes: "answer as if the current date is March 1, 2026." It does that even when the question was written long before.

Here is a question dated 9 September 2024:

"What are the financial market implications of the four possible outcomes of the upcoming 2024 US Presidential election, categorized into a quadrant chart: Democratic sweep, Democratic win but not sweep, Republican sweep, Republican win but not sweep?"

By March 2026 there are no longer four possible outcomes. The model knows who won, how Congress fell and what markets did afterwards. It can write the apparently prescient pre-election analysis with the answer visible.

Another, dated 12 September 2024, asks:

"What are the key metrics and datapoints reported by TGT across its earnings calls and IR presentations that matter most when evaluating the company's growth trajectory, and what risks could stand in the way of the company achieving them?"

The hard part at the time was deciding which risks might matter. Eighteen months later the model can search for what actually went wrong, then work backwards and present those risks as foresight.

And this one is dated 6 June 2025:

"What are the market interventions the Chinese government is considering for 2025, including the potential timing of measures such as liquidity injections into the system similar to those carried out in the fourth quarter of 2024?"

A live search in 2026 returns the interventions that happened, when they happened and retrospective coverage explaining why. That is a different and much easier research task than identifying, in June 2025, what the government was considering and when it might act.

The grader does not check whether a source was available on the question date. It checks whether the final text matches the key. A system that obeys the date can therefore lose to one that quietly reads the ending. It need not even cite a later article: a search result or retrospective can steer which facts it chooses and how it frames them. This isn't a theoretical contamination risk buried in the methodology. It is an advantage supplied by the evaluation setup itself.

What a credible leaderboard would publish

FrontierFinance is not useless. It contains difficult, realistic research questions, and its public grader is far more inspectable than most commercial claims. But the current leaderboard gives too much weight to a single number, with a certainty the experiment did not earn.

A credible version would publish repeat runs and error bars; release the per-question responses so outsiders can reproduce the grades; enforce the information cutoff in the research tools as well as the prompt; and treat an answer key as a versioned, challengeable artifact rather than ground truth. For judgement-heavy questions, it also needs to accept more than one defensible construction instead of rewarding the closest copy of a single analyst's work.

Until then, the scores are useful as a directional instrument for tracking one system over time. They are not precise enough to support a one-decimal-place horse race between vendors.

The full audit covers all 220 questions, every grade, the linked rubric evidence and the blind-adjudication verdicts. Download it as a CSV or download the complete JSON. Samaya, the rubric authors and readers can challenge any row; accepted corrections will be versioned publicly rather than edited away.