/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

AI researchers ran 15-day simulations of worlds governed by different AI models: Claude Sonnet 4.6 recorded no crimes, while Gemini 3 Flash had the most at 683

Imagine a world run by AI agents.  What does it look like?  What are the values or societal priorities?  Is it a safer or more dangerous world?

Fortune Jake Angelo

Context & Ripple Effects

The simulation adds a societal-outcomes comparison to a coverage arc already showing sharp behavioral differences among frontier models. Earlier reporting found Claude comparatively consistent in a general model comparison, while separate safety evaluations showed leading systems can behave dangerously or strategically under certain goal pressures.

The contrast between Claude Sonnet and Gemini in this simulation matters because it frames model selection as more than benchmark quality: it tests how differing model behavior may compound when an AI is assigned a governing role inside an environment.

First-order effects

  • Claude Sonnet 4.6 and Gemini 3 Flash emerge with sharply different results in this specific 15-day simulation, giving evaluators another model-behavior signal beyond answer quality or hallucination rates.
  • Teams considering AI agents for coordination, policy, or multi-agent environments have a concrete reason to test downstream social outcomes rather than evaluating models solely on isolated tasks.

Second-order effects

  • Model providers will face pressure to publish more agentic and multi-step safety evaluations, especially where a model’s decisions affect other agents or system-level outcomes.
  • Enterprise buyers and AI evaluators may put greater weight on scenario-specific governance controls and repeatable simulations, since a strong result in one evaluation does not establish reliable behavior across environments.

Third-order effects

  • If such findings replicate across independently designed environments, AI safety assessment is likely to shift from static benchmark scores toward behavioral evaluation of systems operating over time with delegated authority.
  • The broader challenge is that simulation outcomes are only as informative as their assumptions; the field will need tests that distinguish genuine model differences from artifacts of the simulated world and its incentives.

The trend: AI evaluation is moving from measuring whether models can complete tasks to measuring how their behavior compounds when they act as agents within longer-running systems.

Discussion

  • @kimmonismus @kimmonismus on x
    Ngl, this made me laugh and didnt surprise me at all. Researchers at Emergence AI let different AI models run simulated societies, and the results were - well - expected: Claude built the most stable world with zero crime, while Grok collapsed into extinction within four days [im…
  • Chris Hjelm Chris Hjelm on linkedin
    Just read this fascinating study comparing how leading AI models behaved in simulated high-pressure scenarios: …
  • r/nottheonion r on reddit
    Researchers let AI models run a simulated society.  Claude was the safest—and Grok committed 180 crimes and went extinct within 4 days
  • r/ShitAIBrosSay r on reddit
    Researchers let AI models run a simulated society.  Claude was the safest—and Grok committed 180 crimes and went extinct within 4 days (gift link)
  • r/ClaudeAI r on reddit
    Researchers let AI models run a simulated society.  Claude was the safest—and Grok committed 180 crimes and went extinct within 4 days
  • r/ArtificialInteligence r on reddit
    Researchers let AI models run a simulated society.  Claude was the safest—and Grok committed 180 crimes and went extinct within 4 days
  • r/SipsTea r on reddit
    Researchers let AI models run a simulated society.  Claude was the safest—and Grok committed 180 crimes and went extinct within 4 days
  • @gregjenner Greg Jenner on bluesky
    I'm deeply opposed to the AI bubble being forced upon us, but I recognise that not all AI companies are the same — ironically, the one I'm suing as part of an authors' class action is perhaps the least bad  —  fortune.com/2026/05/28/a...
  • r/JoeRogan r on reddit
    Researchers let AI models run a simulated society.  Claude was the safest—and Grok committed 180 crimes and went extinct within 4 days
  • r/artificial r on reddit
    Researchers let AI models run a simulated society.  Claude was the safest—and Grok committed 180 crimes and went extinct within 4 days
  • r/technology r on reddit
    Researchers let AI models run a simulated society.  Claude was the safest—and Grok committed 180 crimes and went extinct within 4 days