An interview with Geekbench creator John Poole on making the tool cross-platform from the start, what's new in Geekbench 6, why benchmarking matters, and more
so that you can tell when something is wrong. https://arstechnica.com/...
Context & Ripple Effects
This interview lands days after Primate Labs shipped Geekbench 6 with reworked tests built around real-world workloads, and John Poole uses it to explain the design philosophy behind the tool — notably that Geekbench has been cross-platform from the start rather than bolted onto one architecture after the fact.
That choice is what made Geekbench useful at moments like Apple's M1 debut in the Mac mini, where reviewers needed a common yardstick to compare an ARM Mac against Intel and AMD silicon. The interview matters because it frames how the whole industry should read the scores that launch-day coverage leans on.
First-order effects
- Reviewers and buyers comparing laptops, phones, and desktops immediately get a new baseline: with Geekbench 6's datasets meant to mirror actual workloads, published scores from the old version are no longer comparable to fresh results.
- Chip and device makers whose marketing cites Geekbench numbers have to re-anchor their claims to the v6 scale, since their headline figures reset against the new test suite.
Second-order effects
- As workloads diversify beyond general-purpose CPU tasks, Primate Labs extends the brand into adjacent categories — most concretely with Geekbench AI 1.0, which scores CPUs, GPUs, and NPUs on AI-centric performance across mobile and desktop.
- Vendors optimizing for measured performance now face pressure on two fronts at once: keeping general compute scores competitive while also showing credible NPU numbers, which raises the cost of benchmark-driven marketing.
Third-order effects
- If the pattern holds, benchmarking fragments from one universal score into workload-specific suites — general compute, AI inference, media encode — and 'benchmark-as-market-interface' becomes a portfolio rather than a single number.
- Each suite generation then escalates in demand to stay ahead of hardware, a treadmill visible in Geekbench 7's larger datasets and added video/audio encoding tests; vendors that tune for yesterday's suite lose their comparative advantage with every major release.
The trend: Benchmarks are evolving from single cross-platform CPU scores into portfolios of workload-specific suites, because diversified hardware — ARM Macs, phone NPUs — can no longer be fairly compared on one number.