ChatGPT testing shows the chatbot struggles with basic arithmetic questions, confidently offering entertaining but wrong answers, an inherent problem with LLMs
‘Large language models’ supply grammatically correct answers but struggle with calculations — Cheating With ChatGPT: Can an AI Chatbot Pass AP Lit?
Context & Ripple Effects
The early tests establish a reliability gap between fluent prose and calculation. Later coverage finds the same pattern in Khan Academy's Khanmigo tutoring bot, where basic arithmetic errors persisted as the organization worked on accuracy.
The issue is broader than math: an analysis of Stack Overflow responses found incorrect information in 52% of answers, while experts described ChatGPT's tendency to fill gaps with plausible-sounding language. Together, that makes confident presentation—not merely isolated wrong answers—the central product risk.
First-order effects
- ChatGPT users cannot treat its arithmetic output as a dependable answer source; calculations require independent checking despite the chatbot's confident tone.
- OpenAI faces a mismatch between ChatGPT's conversational usefulness and the verification burden imposed on users when an answer involves factual or numerical accuracy.
Second-order effects
- Khan Academy's effort to improve Khanmigo's accuracy shows education-oriented AI products must add safeguards around basic-answer generation rather than rely on fluent explanations alone.
- Developers using ChatGPT for technical help face a similar review burden: the later Stack Overflow analysis ties inaccurate answers to verbose output, increasing the time needed to identify errors.
Third-order effects
- If arithmetic errors and confabulation persist across tutoring and programming use, AI assistants will be adopted as drafting interfaces with human validation built into consequential workflows rather than as autonomous answer engines.
- The durable competition shifts toward systems that can make answer reliability legible and constrain unsupported output, because polished language alone does not establish correctness.
The trend: Generative AI is moving from novelty chat toward verified, workflow-bound assistance as recurring accuracy failures expose the cost of unreviewed answers.