Why LLMs, which don't induce an algorithm that computes multiplication, still don't truly understand multiplication no matter how much data they are trained on
Some Reply Guy on X assured me yesteday that “transformers can multiply”. Even pointed me to a paper, allegedly offering proof.
Context & Ripple Effects
The article sits in an early debate over what strong task outputs from Transformer-based systems actually demonstrate. Related coverage traces how the Transformer architecture accelerated language processing, while later work continued to test whether its apparent reasoning reflects something more than pattern-based performance.
The question remains consequential for mathematical use cases: later coverage reports both no evidence of formal reasoning in language models and efforts to use LLMs in mathematics without surrendering direct mathematical understanding.
First-order effects
- The article challenges the use of successful multiplication outputs as proof that a Transformer has learned the underlying procedure; claims of capability must distinguish correct answers from an induced algorithm.
- For developers and evaluators, multiplication becomes an example of why demonstrations on familiar tasks alone may not establish robust reasoning or understanding.
Second-order effects
- Model evaluations are pushed toward tests that probe generalization and procedure, rather than relying on isolated examples of correct arithmetic output—a concern echoed in later tests of LLM limits on classic reasoning problems.
- Users considering LLMs for mathematical work have stronger reason to retain verification and human understanding in the workflow, consistent with the later discussion of preserving direct mathematical experience.
Third-order effects
- If output-level success repeatedly fails to establish algorithmic competence, the industry will need a clearer separation between language-model fluency and dependable symbolic reasoning.
- The broader consequence is a more disciplined capability discourse: scaling and benchmark gains may remain commercially useful, but they will not by themselves settle claims about understanding.
The trend: AI evaluation is shifting from whether models can produce a right answer to whether they can reliably generalize the procedure that produces it.