/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Internal document: Google trained PaLM 2 on 3.6T tokens and 340B parameters, compared to 780B tokens and 540B parameters for the original PaLM in 2022

- Google's PaLM 2 large language model is using nearly five times the amount of text data for training as its predecessor LLM, CNBC has learned.

CNBC Jennifer Elias

Context & Ripple Effects

Google debuted PaLM 2 last week as a four-size family powering 25 products (across multilingual, reasoning, and coding workloads), but its technical paper stayed quiet on the training data and hardware setup behind it (forthcoming on many of the model's major limitations). An internal document obtained by CNBC now fills part of that gap.

The numbers mark a deliberate break from the 2022 recipe: where the original 540B-parameter PaLM trained on just 780B tokens, PaLM 2 runs at 340B parameters on 3.6T tokens — roughly five times the text on a substantially smaller network.

First-order effects

  • Google gets a cheaper, faster-serving flagship: cutting parameters by about 200B while multiplying training data lets the same model underpin 25 products without the inference bill of the original 540B PaLM.
  • The internal document partially closes a disclosure gap Google's own paper left open on data and hardware, putting concrete scale figures into public hands that the company chose not to publish.

Second-order effects

  • Rival frontier labs now face a comparison baseline set by leaked internals rather than official papers, raising pressure to either disclose comparable training-scale figures or accept that competitors and press will source them elsewhere.
  • Downstream builds inherit the efficiency: fine-tuned derivatives like Med-PaLM 2 — the basis for the MedLM family Google Cloud later shipped — start from a smaller base model whose serving costs are structurally lower.

Third-order effects

  • If the pattern holds, frontier labs converge on data-heavy, right-sized models over parameter maximalism, making published parameter counts an unreliable proxy for capability and pushing evaluation toward benchmarks and outputs.
  • Training-scale transparency shifts from papers to leaks and internal documents, weakening the technical report as the industry's accountability mechanism and inviting regulator or customer demands for standardized disclosure.

The trend: Frontier labs are trading raw parameter counts for dramatically larger training datasets while keeping the specifics of that training out of their published papers.

Discussion

  • @jenn_elias Jennifer Elias on x
    NEW: Google's newest AI model PaLM 2 is trained on nearly 5x more text data than its predecessor, we've learned. With 3.6 trillion tokens, it's trained on significantly more data than LaMDA, LLaMA and potentially GPT-4. https://www.cnbc.com/...
  • @kifleswing Kif on x
    better data, or more data? google says, why not both foundation model sets are still getting larger https://www.cnbc.com/...
  • @wavesblog @wavesblog on x
    From “you are the product” to “you are the input into a LLM” https://twitter.com/...