/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

A test of five AI-image detection tools finds they tend to struggle with altered and low-quality images and sometimes fail even when an image is obviously fake

The pope did not wear Balenciaga.  And filmmakers did not fake the moon landing.  In recent months, however …

New York Times

Context & Ripple Effects

This June 2023 test landed in the middle of the first synthetic-media panic — the Balenciaga pope image had just demonstrated how far AI fakes had come, and newsrooms were reaching for detectors as the obvious countermeasure. The finding that five commercial tools stumble on altered and low-quality images punctured that assumption early.

The corpus shows the gap never closed: Google's Pixel 9 made convincing fake photos trivial with inadequate safeguards a year later, wartime fakes during the Israel-Hamas conflict trained audiences to dismiss genuine content, and by 2026 expanded tests of a dozen-plus detectors still found them struggling with complex images, blind to video, and unreliable on audio.

First-order effects

  • Newsrooms and platforms that bought these five tools as a verification layer get false confidence exactly where it matters most — altered and compressed images, which is what viral reposts actually look like.
  • Detection vendors face immediate credibility pressure: a public test showing failures on obviously fake images undercuts the accuracy claims their sales pitches rest on.

Second-order effects

  • Buyers pivot toward provenance standards that authenticate images at capture rather than guessing after the fact, since the 2026 retests show more detectors and more money have not fixed the core weakness.
  • The failure rate feeds the dismissal problem documented in the wartime coverage — when detectors can't settle authenticity, bad actors gain from simply claiming real footage is fake.

Third-order effects

  • If three years of retests keep confirming the same weaknesses, the industry's trust infrastructure migrates from post-hoc detection to capture-time provenance and layered verification, with detection demoted to one signal among several.
  • Text detection may follow the same arc: Pangram's false-positive problem at scale suggests even 'gold standard' classifiers carry structural risks when institutions act on their verdicts automatically.

The trend: Synthetic-media defense is shifting from detector-first verification toward provenance and multi-signal trust systems, as repeated testing shows detection alone cannot keep pace with generation.

Discussion

  • @gfiorelli1 Gianluca Fiorelli on x
    Fantastic articles by the @nytimes with tons of tests showing how AI generated images can pass (or not) an “AI generated test”, and what happens in case of real images: https://www.nytimes.com/... h/t @seostratega
  • @ndiakopoulos Nicholas Diakopoulos on x
    This article frames the challenge as classifying images as “real” or “fake” but that's the wrong way to think about using AI here. We need AI systems with interfaces that explain to forensic analysts what features are indicative of fakeness or realness: https://www.nytimes.com/..…