Skip to content
AI.info

The Pulse

Chatbots Miss 57% of Financial Questions in Saturn Study

A Saturn study tested 18 AI models on 121 personal-finance questions and found incorrect answers in 57% of responses. The findings arrive as PensionBee reports that many users would act on chatbot financial advice without checking it.

Chatbots Miss 57% of Financial Questions in Saturn Study

AI.info Team ·

57% of answers were wrong

AI models gave incorrect answers 57% of the time in a test of personal-finance questions, according to a Saturn study covered by Startup Fortune. The research tested 18 models, including ChatGPT, Claude, Copilot, Grok, and Gemini, against 121 questions covering topics such as debt, student loans, mortgages, pensions, taxes, and savings.

Saturn repeated each question five times, producing more than 10,000 responses for evaluation. The study described its findings in a report titled Artificial Authority: Should you trust AI to deliver financial advice? A separate account published by Financial Reporter confirms the overall error rate and says Saturn marked answers as failures when they contained factual errors, missed important points, or omitted warnings.

The results do not mean that every model failed on 57% of the 121 questions. The reported figure refers to the frequency of incorrect answers across the tested responses. That distinction matters because a single question could receive different answers when repeated, exposing a second problem beyond simple accuracy: inconsistency.

Complex money questions produced worse results

Performance deteriorated on the study's harder questions. Saturn reported an average error rate of 88% for the most difficult queries, with some models reaching a 99% failure rate, according to the reports.

The questions covered decisions that consumers routinely face but that depend heavily on individual circumstances. Examples include whether to overpay a mortgage or invest extra money, how much to save for retirement, and what to do after missing a student-loan payment. A fluent explanation can sound useful while still overlooking tax rules, interest rates, income volatility, debt terms, or a household's need for emergency cash.

Financial Reporter said free models produced incorrect answers in 63% of cases, compared with 49% for paid models. The same report said free tools failed on 93% of the hardest questions. Those comparisons come from Saturn's own testing framework and should not be treated as a general ranking of every free and paid AI product.

Users often skip the verification step

The accuracy findings coincide with a survey from PensionBee that points to a gap between model performance and user behavior. In its 2026 AI and Your Money Report, PensionBee surveyed 1,000 U.S. adults who use AI for financial questions and found that 57% would accept a chatbot's answer on a financial decision without checking it.

Twenty-three percent said a chatbot had already given them incorrect financial information, while 5% said they discovered the error only after acting on it. PensionBee also found that 53% of respondents had privacy concerns even as they shared information such as bank statements, salary details, and debt data with AI tools.

The survey indicates that users are not limiting chatbots to low-risk tasks such as explaining a term or organizing a budget. Respondents described using AI for decisions involving investments, retirement timing, and other choices that can be difficult to reverse.

Regulators are watching the advice boundary

The Financial Conduct Authority's Mills Review, published in July, identifies reliance on unregulated AI for financial guidance or advice as a potential consumer harm. The review also raises concerns about mis-selling, model errors, bias, and the difficulty of determining who is accountable when consumers delegate decisions to AI-enabled systems.

The FCA says existing rules provide a framework for firms using AI and does not propose a separate set of AI regulations in the review. It does, however, call for continued attention to the boundary between general guidance and regulated financial advice as automated systems become more personalized.

That boundary is central to the Saturn findings. A chatbot can summarize a pension concept or help a user prepare questions for a licensed adviser without making a regulated recommendation. The risk rises when a user treats a confident answer as a complete plan and acts without checking the underlying calculations or assumptions.

What the study does—and does not—show

Saturn's test does not establish that AI is useless for personal finance. Other research has found that language models can encourage broad habits such as saving more, reducing spending, and diversifying investments. An MIT Sloan study published in May found that the quality of AI-generated guidance varied with the information users supplied and that models struggled with income shocks, portfolio rebalancing, and retirement withdrawals.

The studies measure different things. Saturn tested answers against a set of financial questions and scored errors and omissions. The MIT research examined simulated lifetime financial outcomes based on advice generated from user-written prompts. Together, they show why a polished response is not the same as advice that is accurate, personalized, and suitable for a specific household.

For consumers, the practical lesson is narrow but significant: use a chatbot to clarify terminology, compare questions, or organize information, then verify any recommendation involving borrowing, taxes, retirement, insurance, or investments with authoritative sources or a qualified professional. In Saturn's test, relying on the first answer would have meant accepting an incorrect response more often than not.

Source

Startup Fortune

Explore

More articles