Hands-on with Khan Academy's ChatGPT-powered AI tutoring bot Khanmigo, which frequently makes basic arithmetic errors, as Khan Academy works to improve accuracy
Khanmigo, a ChatGPT-powered bot, made frequent calculation errors during a Journal reporter's test X: @natarajsindam , @carnage4life , and @matt_barnum X: @natarajsindam : This is not fully true. LLMs alone are not good at Math. But LLMs are generating ability to call calculators when necessary or other tools where they dont have knowledge. This is not fully productized, so apps like these are making promises they cant keep. Dare Obasanjo / @carnage4life : It's a remarkable trend that the rise of LLMs means that computers have gotten dumber. While they've gotten better at understanding what we're asking, they now can't be trusted to do basic math or answer factual questions. One step forward but how many steps backwards? [image] Matt Barnum / @matt_barnum : 1/ Will AI be able to serve as a personalized tutor for struggling students? Many hope so. But currently there is at least one big problem: AI is just not very good at math itself, so it's not clear it can teach math. 🧵 https://www.wsj.com/... https://www.wsj.com/...
Context & Ripple Effects
Khanmigo was introduced as a GPT-based assistant for math, coding and discussion, then trialed as a one-to-one tutoring tool at Khan Lab School. This test puts pressure on that tutoring premise: the original Khanmigo pilot depended on reliable help without simply handing students answers.
The problem also follows an established limitation rather than an isolated tutoring-app failure: earlier ChatGPT arithmetic testing found confident errors on basic calculations. Khan Academy's work to improve accuracy therefore matters most in the instructional moments where students may treat an answer as authoritative.
First-order effects
- Khanmigo users and educators must verify arithmetic outputs rather than rely on the bot as a dependable math authority while Khan Academy improves accuracy.
- Khan Academy faces a product-quality gap between conversational tutoring and calculation tasks, particularly where an incorrect intermediate step can mislead a learner.
Second-order effects
- AI tutoring providers are pushed to constrain calculation responses or connect models to deterministic tools, because general-purpose language-model fluency does not itself ensure numerical correctness.
- Schools and teachers evaluating AI tutors are likely to place more weight on oversight and validation workflows, reducing the value of unsupervised use for foundational skills.
Third-order effects
- If reliable tutoring products increasingly require models to orchestrate calculators and other specialized tools, competition will shift from the base chatbot to the quality of the integrated instructional workflow.
- The broader market may separate low-stakes conversational assistance from high-trust educational tasks, with demonstrable accuracy becoming a prerequisite for wider classroom deployment.
The trend: AI education products are moving from standalone language-model demonstrations toward tool-assisted, supervised systems designed for tasks where correctness is essential.