/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

← → days · ↑ ↓ browse · Enter similar · o open

Why LLMs, which don't induce an algorithm that computes multiplication, still don't truly understand multiplication no matter how much data they are trained on

Some Reply Guy on X assured me yesteday that “transformers can multiply”.  Even pointed me to a paper, allegedly offering proof.

Marcus on AI Gary Marcus

Context & Ripple Effects

The article sits in an early debate over what strong task outputs from Transformer-based systems actually demonstrate. Related coverage traces how the Transformer architecture accelerated language processing, while later work continued to test whether its apparent reasoning reflects something more than pattern-based performance.

The question remains consequential for mathematical use cases: later coverage reports both no evidence of formal reasoning in language models and efforts to use LLMs in mathematics without surrendering direct mathematical understanding.

First-order effects

  • The article challenges the use of successful multiplication outputs as proof that a Transformer has learned the underlying procedure; claims of capability must distinguish correct answers from an induced algorithm.
  • For developers and evaluators, multiplication becomes an example of why demonstrations on familiar tasks alone may not establish robust reasoning or understanding.

Second-order effects

Third-order effects

  • If output-level success repeatedly fails to establish algorithmic competence, the industry will need a clearer separation between language-model fluency and dependable symbolic reasoning.
  • The broader consequence is a more disciplined capability discourse: scaling and benchmark gains may remain commercially useful, but they will not by themselves settle claims about understanding.

The trend: AI evaluation is shifting from whether models can produce a right answer to whether they can reliably generalize the procedure that produces it.

Discussion

  • @garymarcus Gary Marcus on x
    So it turns out the only two things keeping us from AGI are a piece of paper and a pencil. Who knew?
  • @bobehayes Bob E. Hayes on x
    “Math is hard” — if you are an LLM - and why that matters “A calculator... would be at 100%, because it is programmed... with an algorithm that actually computes multiplication. The LLM never induces such an algorithm.” ~ @garymarcus https://garymarcus.substack.com/ ... #AI [imag…
  • @garymarcus Gary Marcus on x
    What LLMs Have in Common with Teen Talk Barbie [image]
  • @bair82 Bair on x
    @cajundiscordian ... It takes like 500 hours. It's grade 4 level math, at 130-140 academic hours a year (100 real hours). But that's missing the point. Transformers can multiply https://arxiv.org/...
  • @nlholdem Paul Miller on x
    @GaryMarcus I think this is a relevant (and possibly important) paper: https://arxiv.org/... Train a deep net on a symbolic reasoning task, obtain near-perfect results on validation data, but find it hasn't learned the problem at all! Only the statistics of the training set
  • @garymarcus Gary Marcus on x
    Zero percent accuracy. On something pocket calculators could do just fine in 1975. #AGI
  • @garymarcus Gary Marcus on x
    @bair82 ... Did you read the paper? No calculator ever would be this bad. Even with a massive amount of relevant data, it doesn't really understand what multiplication is, with accuracy dropping rapidly as problem increases in digits. [image]
  • @andrew_n_carr Andrew Carr on x
    Language models are bad a basic math. GPT-4 has right around 0% accuracy rate on 5 digit multiplication. Most open models can't even add. Why is that? There are a few reasons why numbers are hard. The main one is Tokenization. When training a tokenizer from scratch, you take... […
  • @brianmcc Brian McCullough on x
    Think about it. What's wild about this is, from the very first vacuum tube days, computers were WAY better than us at 2 things: math and remembering data. Now the computers are good at talking to us, but not remembering or doing simple calculator stuff. https://garymarcus.substac…
  • r/singularity r on reddit
    “Math is hard” — if you are an LLM - and why that matters |  No matter how much data you train them on, they still don't truly understand multiplication.