Primate Labs releases Geekbench 6 with new and updated hardware and app tests that use datasets meant to be more representative of real-world workloads
Context & Ripple Effects
Geekbench 6 updates the test suite around datasets intended to resemble real workloads, extending Primate Labs’ cross-platform benchmarking approach discussed in a creator interview on why cross-platform benchmarks matter. It establishes a new baseline rather than simply adding another score to the existing one.
Later releases show Primate Labs broadening that baseline in two directions: Geekbench AI’s CPU, GPU, and NPU tests and Geekbench 7’s larger CPU and GPU datasets plus media tests. The update matters because the benchmark’s workload choices shape which hardware strengths are made legible in comparisons.
First-order effects
- Hardware reviewers, device makers, and buyers using Geekbench must treat Geekbench 6 results as a separate comparison set from earlier versions, because its tests and datasets have changed.
- Primate Labs gives CPUs and application hardware a test suite designed to emphasize a wider set of real-world-style workloads rather than relying on the prior workload mix.
Second-order effects
- Chip vendors’ performance messaging and reviewers’ test procedures shift toward Geekbench 6 scores, while legacy Geekbench results lose direct comparability for product-to-product claims.
- The new workload mix raises the value of performance characteristics that its updated hardware and app tests capture, influencing how competing devices are framed across platforms.
Third-order effects
- Benchmark suites are becoming broader measurement interfaces: Primate Labs’ subsequent AI and expanded CPU, GPU, video, and audio tests indicate that a single headline CPU score is giving way to workload-specific comparisons.
- As benchmark versions evolve, comparability increasingly depends on the test generation and workload category, concentrating influence in the benchmark designer’s choice of representative datasets.
The trend: Cross-platform benchmarking is moving from narrow processor scoring toward evolving, workload-specific tests spanning conventional compute, AI, graphics, and media tasks.