/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

← → days · ↑ ↓ browse · Enter similar · o open

An interview with Geekbench creator John Poole on making the tool cross-platform from the start, what's new in Geekbench 6, why benchmarking matters, and more

so that you can tell when something is wrong. https://arstechnica.com/...

Ars Technica Andrew Cunningham

Context & Ripple Effects

This interview lands days after Primate Labs shipped Geekbench 6 with reworked tests built around real-world workloads, and John Poole uses it to explain the design philosophy behind the tool — notably that Geekbench has been cross-platform from the start rather than bolted onto one architecture after the fact.

That choice is what made Geekbench useful at moments like Apple's M1 debut in the Mac mini, where reviewers needed a common yardstick to compare an ARM Mac against Intel and AMD silicon. The interview matters because it frames how the whole industry should read the scores that launch-day coverage leans on.

First-order effects

  • Reviewers and buyers comparing laptops, phones, and desktops immediately get a new baseline: with Geekbench 6's datasets meant to mirror actual workloads, published scores from the old version are no longer comparable to fresh results.
  • Chip and device makers whose marketing cites Geekbench numbers have to re-anchor their claims to the v6 scale, since their headline figures reset against the new test suite.

Second-order effects

  • As workloads diversify beyond general-purpose CPU tasks, Primate Labs extends the brand into adjacent categories — most concretely with Geekbench AI 1.0, which scores CPUs, GPUs, and NPUs on AI-centric performance across mobile and desktop.
  • Vendors optimizing for measured performance now face pressure on two fronts at once: keeping general compute scores competitive while also showing credible NPU numbers, which raises the cost of benchmark-driven marketing.

Third-order effects

  • If the pattern holds, benchmarking fragments from one universal score into workload-specific suites — general compute, AI inference, media encode — and 'benchmark-as-market-interface' becomes a portfolio rather than a single number.
  • Each suite generation then escalates in demand to stay ahead of hardware, a treadmill visible in Geekbench 7's larger datasets and added video/audio encoding tests; vendors that tune for yesterday's suite lose their comparative advantage with every major release.

The trend: Benchmarks are evolving from single cross-platform CPU scores into portfolios of workload-specific suites, because diversified hardware — ARM Macs, phone NPUs — can no longer be fairly compared on one number.

Discussion

  • @arstechnica @arstechnica on x
    It's easy to dismiss benchmarking as something that you don't need to care about unless you're a reviewer or a showboating hobbyist. But there's still value in knowing how fast something is supposed to be—so that you can tell when something is wrong. https://arstechnica.com/...