/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

← → days · ↑ ↓ browse · Enter similar · o open

Claude Sonnet 3.5 hands-on: the AI understood a basic game and its mechanics, had a strategy, was willing to revise it based on learning, but fragilities remain

Some quick impressions of an actual agent  —  There seems to be near-universal belief in AI that agents are the next big thing.

One Useful Thing Ethan Mollick

Context & Ripple Effects

Claude Sonnet 3.5 arrived amid coverage portraying it as a meaningful model-performance step, including an earlier assessment that its gains suggested progress was not slowing. This hands-on test narrows that broader claim to agent behavior: it could form and update a plan in a bounded task, while exposing reliability limits.

The result also sits between older demonstrations of game-playing agents and later evidence that agentic competence can break down in operational settings, such as Claude's mixed storefront-management trial. The distinction matters: adaptive behavior in a simple environment is not yet dependable execution across business decisions.

First-order effects

  • Anthropic gains a concrete, if limited, demonstration that Claude Sonnet 3.5 can interpret rules, choose a strategy, and revise it after feedback rather than merely generate a one-off answer.
  • The reported fragilities make human oversight and constrained task design necessary for users considering the model for agent-like workflows.

Second-order effects

  • Competing model providers face pressure to demonstrate not only benchmark performance but also planning, learning, and recovery from mistakes in interactive tasks.
  • Agent builders will need to invest in guardrails, evaluation harnesses, and fallback paths, because a model's ability to adapt does not by itself establish reliable autonomy.

Third-order effects

  • If repeated across more complex settings, evaluation will shift from static model outputs toward whether agents can sustain goals, learn from feedback, and fail safely within general-purpose computer workflows.
  • The emerging market for agents is likely to differentiate on operational reliability and control layers as much as on raw model capability; the evidence here supports that direction but does not establish readiness for unsupervised deployment.

The trend: AI is moving from chat-oriented models toward embedded agents whose commercial value depends on reliable action, adaptation, and oversight.

Discussion

  • @emollick Ethan Mollick on x
    I really like Claude's response to a bug (not with Claude) that made it lose its connection to the desktop. Declare victory and move on. [image]
  • @hpc_guru @hpc_guru on x
    There seems to be near-universal belief in #AI that #agents are the next big thing When you give a Claude a mouse: @emollick shares his first impressions of power & weaknesses It had a strategy and it was willing to revise it based on what it learned https://www.oneusefulthing.or…
  • @emollick Ethan Mollick on x
    My impressions of getting to work with a preview of the new Claude with computer use, and putting it to work across multiple tasks. You can see the promise of agents, but also where they still need work. I think it is a sign of the near future. https://www.oneusefulthing.org/ ...