AI chatbots often answer financial queries incorrectly
claude copilot gemini grok
| Source: HN | Original article
Top AI chatbots—including ChatGPT, Claude, Copilot, Grok and Gemini—answered financial questions incorrectly about 57% of the time.
A new study by technology firm Saturn finds that the most widely used generative‑AI chatbots – including ChatGPT, Claude, Copilot, Grok and Gemini – deliver incorrect answers to financial questions 57 percent of the time on average. The analysis, which tested 18 models on a range of queries, shows error rates climbing to as high as 99 percent for the most complex scenarios.
The findings arrive as regulators in the UK and elsewhere intensify scrutiny of AI’s role in financial decision‑making. Unlike licensed advisers, general‑purpose chatbots are not bound by fiduciary duties or consumer‑protection rules, meaning users can be misled without recourse. Saturn’s report highlights two recurring problems: failure to incorporate the latest tax legislation and the generation of fabricated financial rules that appear plausible but have no legal basis.
For consumers, the study underscores the need to treat AI‑generated advice as a starting point rather than a definitive answer, and to verify any guidance against official sources before acting. For the industry, the results raise pressure to embed up‑to‑date regulatory data into models or to develop specialised, compliance‑focused tools.
Watchers should monitor forthcoming regulatory responses, including possible mandates for transparency disclosures or accuracy benchmarks for AI financial services. In parallel, AI developers are likely to face heightened demand for model‑training pipelines that ingest real‑time fiscal updates, and for mechanisms that flag when a query exceeds the system’s reliable knowledge domain. The Saturn study may become a catalyst for tighter oversight and for a clearer demarcation between casual chatbots and regulated financial advice.
Sources
Back to AIPULSEN