/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Anthropic says Opus 4.5 outscored all humans on a take-home exam it gives to prospective performance engineering candidates, within a prescribed two-hour limit

Michael Nuñez / VentureBeat :

VentureBeat Michael Nuñez

Context & Ripple Effects

This claim sits early in a capability arc centered on Anthropic’s own performance-engineering hiring test. Subsequent coverage says the company redesigned that take-home assessment after Claude repeatedly beat it, underscoring how quickly a static screening exercise can lose discriminatory value.

Later Anthropic reporting emphasized models that devote more attention to difficult task components without being explicitly prompted, extending the same question from a single test result to how models execute complex technical work.

First-order effects

  • Anthropic can point to a concrete, time-bounded internal hiring exercise as evidence for Opus 4.5’s performance-engineering capability, although the result remains a company-reported evaluation.
  • Performance-engineering candidates and hiring teams using similar take-home formats face a more immediate integrity problem: an answer alone is less reliable evidence of an applicant’s unaided skill.

Second-order effects

  • Employers may shift technical assessment toward supervised, interactive, or system-specific work that tests judgment and verification rather than only completed code or analysis.
  • Model vendors will face stronger demand for evaluations tied to real workflows and time limits, not just broad benchmark scores; Anthropic’s later decision to revise its own test illustrates that pressure.

Third-order effects

  • If model capability continues to overtake fixed hiring exercises, technical hiring is likely to place more weight on human oversight, problem framing, and the ability to audit AI-assisted work.
  • The durable competitive measure may shift toward cost and reliability per completed engineering task, with vendors and employers needing assessments that remain useful as models improve.

The trend: AI coding models are moving from tools that assist technical candidates toward agents that can challenge the validity of conventional technical screening itself.

Discussion

  • @shiringhaffary Shirin Ghaffary on x
    Interesting nugget from Anthropic about its Opus 4.5 release today: company says the model got s higher score than any Anthropic interviewee candidate ever has on an engineering take-home assignment from @rachelmetz https://www.bloomberg.com/... [image]
  • @matthewberman Matthew Berman on x
    Absolutely insane stat. Opus 4.5 outperformed EVERY SINGLE HUMAN CANDIDATE EVER in Anthropic's notoriously difficult take-home exam for prospective performance engineering candidates. [image]
  • @bran_don_gell Brandon Gell on x
    it totally and completely blows my mind that this is an incredible feat AND 5 years from now we'll look back at this and laugh at what we thought was impressive. this is a crazy paradigm shift.
  • @rohanpaul_ai Rohan Paul on x
    Unreal result. Anthropic said it tested Claude Opus 4.5 on a notoriously demanding take-home exam that it gives to prospective performance engineers, and the model scored higher than any human candidate ever had. [image]
  • @_sholtodouglas Sholto Douglas on x
    This was a truly eerie threshold for me
  • @trishume Tristan Hume on x
    Every time we train a great new model I need to frantically try to write a new take home that the model can't defeat so we can still hire post-release. This one was tough, many drafts based on real problems fell before Claude Code's “ultrathink” and needed to be scrapped.
  • @apples_jimmy @apples_jimmy on x
    “ Opus 4.5 scored higher than any human candidate ever ” on a take home exam they use as an internal benchmark. Also I feel like Anthropic is holding back, Dario not revealing his full power to avoid too big of a jump in capabilities. [image]
  • @daniel_c0deb0t Daniel Liu on x
    opus 4.5 does better than me on the perf take home, which I took to get my current job