AI agents · OpenClaw · self-hosting · automation

Quick Answer

ChatGPT vs Claude vs Gemini vs Grok: Financial Advice (2026)

Published:

The headline numbers

UK fintech Saturn published Artificial Authority: Should you trust AI to deliver financial advice? in mid-September 2026, and the Financial Times ran it on September 20. The design: 18 AI models, 121 consumer money questions across debt, mortgages, pensions and tax, each asked five times, giving more than 10,000 responses. An answer failed if it contained a factual error, omitted something material, or missed a warning that should have been given.

  • 57% of answers were wrong overall (43% accurate).
  • 49% wrong for paid tiers; 63% for free tiers.
  • 88% wrong on average for the hardest, multi-step questions; 93% for free models; the weakest models missed almost every complex query (up to 99%).

Verified September 21, 2026 against Saturn’s report coverage in the FT, Professional Adviser, IFA Magazine and ResultSense, and PensionBee’s 2026 AI and Your Money report.

Per-model results

Saturn named individual performers. Only the models the coverage identifies are listed; the remaining models in the 18 were not broken out publicly.

Model (as tested, Sep 2026)Wrong answersNotes
Claude Opus 5 (reasoning mode)39%Best in study; Anthropic’s default Opus since July 24, 2026
ChatGPT — GPT-5.6 Luna58%Consumer ChatGPT model tier at test time; GPT-6 Astra (Sep 3) not in the published breakdown
Grok 4.559%xAI
Gemini 3.1 Pro73%Google; Gemini 3.5 Pro and 3.8 Flash not in the published breakdown
Claude Haiku 4.582%Last place; a small, fast tier
Copilot and 13 othersnot publishedIncluded in the 57% average

Two readings follow. First, reasoning mode and model size matter more than brand: Anthropic fielded both the best and the worst model. Second, the spread between “best” and “worst” is 43 points, but even the best is wrong two times in five on questions an adviser must get right every time.

What the failures looked like

The specific errors are more instructive than the percentage:

  • Pension tax. One answer would have exposed the user to a £17,500 HMRC charge — the kind of allowance/threshold mistake a model makes when it applies last year’s rules or the wrong jurisdiction.
  • Debt prioritisation. Models steered users to clear the highest-interest balance first, ahead of priority arrears (council tax, rent). Mathematically tidy, legally dangerous: priority debts carry eviction and bailiff consequences that a credit card does not.
  • Fabricated rules. One model invented a provision letting graduates suspend student-loan repayments by emigrating — the opposite is true, and following it would have raised monthly payments.
  • Omitted risk warnings and failure to account for upcoming tax changes — both count as failures under Saturn’s rubric even when the arithmetic was right.

These map onto known LLM weaknesses: stale rules, jurisdiction blur, confident completion of a plausible-sounding rule, and optimising for the stated objective (minimise interest) rather than the unstated constraint (don’t get evicted).

The demand side: people are acting on it

PensionBee’s US survey of 1,000 adults who use chatbots for personal finance, released the same week, found 57% would act on the advice without independent verification, 23% had already received incorrect financial information from a chatbot, and 5% discovered the error only after acting. Gen Z and Millennials were the most willing to share sensitive financial data with a chatbot. In the UK, the FCA’s Mills Review put consumer trust in general-purpose AI for money questions at 26%, with the regulator warning that such advice carries none of the redress of regulated advice.

The provenance caveat

Saturn designed the test, scored it and paid for it, and it wants the FCA to regulate AI financial advice — a market it operates in. ResultSense’s summary is the right one: the figures are “the firm’s own and should be cited as such rather than treated as an independent benchmark.” The mitigating factor is reproducibility — 18 named models and a published question set make an independent rerun straightforward, and nobody without a commercial stake has done one yet.

How this fits the vendors’ own moves

Both OpenAI and Anthropic launched vertical finance products this month — ChatGPT for Financial Services and Claude for Financial Advisors — aimed at advisers and institutions, with retrieval over the firm’s own rules, compliance logging and human sign-off. The Saturn result is the argument for that architecture: a general chatbot answering a consumer cold is the worst configuration, a model grounded in current, jurisdiction-specific documents with a licensed human in the loop is what the products sell. The same logic runs through the Visa Trust Index on agentic commerce: consumers will delegate money decisions to agents faster than the agents earn it.

Using them anyway: a checklist

  1. Model choice: frontier reasoning tier only — Claude Opus 5 or Fable 5.1 with extended thinking, GPT-6 Astra, Gemini 3.5 Pro. Never a free/fast tier for money.
  2. Pin the facts: jurisdiction, tax year, residency, employment status, exact balances. Ask the model to restate them before answering.
  3. Demand the rule: “Which specific rule or threshold are you applying, and where is it published?” Then open the government page.
  4. Ask for the failure case: “What would make this advice wrong for me?” — this surfaces omitted warnings, the most common failure class.
  5. Red lines: pension withdrawals, tax elections, debt prioritisation, mortgage products — get a regulated adviser. The study’s costliest errors were all in these categories.
  6. Keep the transcript. If you do act on it, you will want the record; you will not have regulatory redress.

Sources