Internal document: Google trained PaLM 2 on 3.6T tokens and 340B parameters, compared to 780B tokens and 540B parameters for the original PaLM in 2022
- Google's PaLM 2 large language model is using nearly five times the amount of text data for training as its predecessor LLM, CNBC has learned.
Context & Ripple Effects
Google debuted PaLM 2 last week as a four-size family powering 25 products (across multilingual, reasoning, and coding workloads), but its technical paper stayed quiet on the training data and hardware setup behind it (forthcoming on many of the model's major limitations). An internal document obtained by CNBC now fills part of that gap.
The numbers mark a deliberate break from the 2022 recipe: where the original 540B-parameter PaLM trained on just 780B tokens, PaLM 2 runs at 340B parameters on 3.6T tokens — roughly five times the text on a substantially smaller network.
First-order effects
- Google gets a cheaper, faster-serving flagship: cutting parameters by about 200B while multiplying training data lets the same model underpin 25 products without the inference bill of the original 540B PaLM.
- The internal document partially closes a disclosure gap Google's own paper left open on data and hardware, putting concrete scale figures into public hands that the company chose not to publish.
Second-order effects
- Rival frontier labs now face a comparison baseline set by leaked internals rather than official papers, raising pressure to either disclose comparable training-scale figures or accept that competitors and press will source them elsewhere.
- Downstream builds inherit the efficiency: fine-tuned derivatives like Med-PaLM 2 — the basis for the MedLM family Google Cloud later shipped — start from a smaller base model whose serving costs are structurally lower.
Third-order effects
- If the pattern holds, frontier labs converge on data-heavy, right-sized models over parameter maximalism, making published parameter counts an unreliable proxy for capability and pushing evaluation toward benchmarks and outputs.
- Training-scale transparency shifts from papers to leaks and internal documents, weakening the technical report as the industry's accountability mechanism and inviting regulator or customer demands for standardized disclosure.
The trend: Frontier labs are trading raw parameter counts for dramatically larger training datasets while keeping the specifics of that training out of their published papers.