TL;DR
A study commissioned by the fintech firm Saturn put 18 models through 121 money questions, each asked five times, and reports that the answers were wrong 57% of the time. Difficulty made it worse: the average error rate on hard questions reached 88%, and the weakest models missed almost every complex query. Saturn is calling for the FCA to bring AI financial advice inside the regulatory perimeter — a position from which it stands to benefit.
What the study says it found
The report, “Artificial Authority: Should you trust AI to deliver financial advice?”, put over 10,000 individual queries through the models. An answer counted as a failure if it contained a factual error, skipped something material, or left out a warning that should have been there.
Paid tools did better than free ones, though neither result is comfortable: 49% error for paid, 63% for free, rising to 93% on the hardest questions for free models. Saturn names the individual performers. Claude Haiku 4.5 came last on 82% wrong, ahead of Gemini 3.1 Pro on 73%, with Grok 4.5 at 59% and ChatGPT 5.6 Luna at 58%. The strongest was Claude Opus 5 in reasoning mode, still wrong 39% of the time.
The specific failures are more useful than the headline percentage. One pension tax answer would have exposed someone to a £17,500 HMRC bill. In debt scenarios, models steered users towards clearing the highest-interest balance first ahead of priority arrears like council tax and rent — advice that can end in eviction or bailiffs. One model fabricated a rule letting graduates suspend student loan repayments by emigrating, which would in fact have raised their monthly payments.
Read the provenance
Saturn designed the methodology, applied the scoring and paid for the work, and has a direct interest in how the FCA rules. That does not make the findings wrong, but it means the precise figures are the firm’s own and should be cited as such rather than treated as an independent benchmark.
Looking forward
The regulator is not starting from nothing. Its Mills Review found 26% of consumers already trust general-purpose AI with money questions, and the FCA has warned that people taking such advice get none of the redress that regulated advice carries. Independent replication is what this area is missing — with 18 named models and a published question set, it is now straightforward for someone without a commercial stake to run it again.