/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Pangram Labs: ~21% of the 75,800 peer reviews submitted for ICLR 2026, a major ML conference, were fully AI-generated, and 50%+ contained signs of AI use

By - Miryam Naddaf 0  —  Miryam Naddaf is a science writer based in London.  —  Search author on:  —  PubMed Google Scholar

Nature Miryam Naddaf

Context & Ripple Effects

The reported ICLR result marks a sharp escalation from a 2024 analysis of computer-science reviews that found LLM-written material at the sentence level, rather than widespread fully generated reviews. It makes AI use a workflow-integrity issue for research evaluation, not just an authorship concern.

The finding also helps explain why conferences later moved to tighten rules on LLM use in paper writing and review. Detection remains consequential but imperfect at scale: later coverage examined concerns around Pangram's claimed accuracy. Earlier evidence of LLM-written review sentences provides the immediate baseline.

First-order effects

  • ICLR organizers and program chairs face a substantially larger review-quality and disclosure problem: a reported 21% of 75,800 reviews were fully AI-generated, while more than half showed some AI-use signal.
  • Authors and reviewers may face greater scrutiny of review provenance, especially where generated critiques could affect accept-or-reject decisions without clear human accountability.

Second-order effects

  • Conference operators are pressured to define enforceable boundaries between permitted assistance and generated reviewing, while adding review audits or disclosure requirements increases administrative work.
  • AI-writing detection becomes a more central governance tool, but its use can create disputes when high-volume screening turns uncertain signals into individual review judgments.

Third-order effects

  • If this pattern persists, peer review may shift from a trust-based volunteer process toward one with formal provenance, disclosure, and enforcement layers for AI-assisted work.
  • The broader challenge is likely to be operational governance rather than a simple ban: conferences must preserve reviewer capacity while maintaining confidence that evaluations reflect accountable expert judgment.

The trend: Generative AI is moving from optional drafting aid to a governance problem in knowledge-production workflows, forcing institutions to build rules and verification around human accountability.

Discussion

  • @southernwintrs Will I Am on x
    A paper that is edited by AI is not necessarily bad! I don't think my papers, which I edit with AI, is made worse by the fact that I use it as a tool (where I supervise and edit inputs and outputs). Kinda funny that at AI conferences, using AI tools is now considered *bad*.
  • @v_shaal Vishal Verma on x
    The real question: what's worse - AI-generated reviews of real research, or human reviews of AI-generated papers? Because if both sides are automating this, we've basically just created an extremely expensive bot-to-bot book club.
  • @b_shrir Sriram B on x
    Just tried this for the papers on which I am a reviewer. For many of them, my review is the only human-generated one😂
  • @pandaashwinee @pandaashwinee on x
    it would be interesting to understand whether the AI reviewers prefer AI writing. my understanding so far is that they do, but i wonder if anyone has looked into this quantitatively?
  • @damienteney Damien Teney on x
    Amazingly the set of people complaining about low-quality reviews/papers seems completely disjoint from those who wrote low-quality reviews/papers. 🤔
  • @nogarot Noga H. Rotman on x
    Super interesting findings, especially on the submissions. Most papers seem to have very little AI-generated content, and lower AI usage actually correlates with higher scores. Some encouraging data in all of this.
  • @lovvge Bernie Wang on x
    Both as an AC and author, i couldn't care less who / what writes the reviews ... the only thing I care about is whether, underneath the tokens, there are valid points.
  • @dcml0714 Cheng Han Chiang on x
    Glad to see that all four reviewers I spent a lot of time writing for are classified as full human-written. Kind of
  • @dawei_li_asu Dawei Li on x
    With ~21% of #ICLR reviews flagged as AI-generated, maybe it's time to ask: who's judging your paper — a human or an LLM? In our recent paper: “Who's Your Judge? On the Detectability of LLM-Generated Judgments” we study how to detect LLM-generated judgments and reviews — even [im…
  • @abhinav95_ Abhinav Shukla on x
    Good news: this tool detects all my ICLR reviews as fully-human written. Bad news: a lot of reviews on the papers in my stack are fully AI generated (including one in which mine is the only one detected as human-written!).
  • @canyuchen3 @canyuchen3 on x
    🚨21% of ICLR 2026 reviews are fully AI-generated, and 35% are AI-edited🚨. We are already living in an era where AI silently shapes scientific peer review. Understanding, detecting, and defining 𝐡𝐮𝐦𝐚𝐧-𝐀𝐈 𝐜𝐨-𝐚𝐮𝐭𝐡𝐨𝐫𝐬𝐡𝐢𝐩 is no longer optional, but is critical [image]
  • @alexolegimas Alex Imas on x
    More than a 5th of reviews at a top CS conference were fully AI-generated. These reviews were materially different both in terms of content and scores assigned. Curious what this looks like for Economics. Easy to do, and I'm happy to facilitate if any editors want to run this.
  • @shuaichenchang Shuaichen Chang on x
    Now we're trusting AI to detect AI-generated content more than we trust AI-generated content itself haha. To be clear, I have full respect for the people who put in the effort to make this analysis possible. From my own experience: I have one submission where all four reviews
  • @xianjun_agi Xianjun Yang on x
    It's worth noting that most AI detectors can only return a probability of whether the text is AI-generated. But a high probability alone can not serve as practical evidence. Luckily, our previous work published at ICLR 2024 can provide strong text-level EVIDENCE to support the
  • @max_spero_ Max Spero on x
    Curious about AI use in paper writing or reviews? We ran every paper and every review through @pangramlabs, and this is what we found. 🧵
  • @danielkhashabi Daniel Khashabi on x
    Fascinating snapshot of AI's impact in scientific writing/reviewing from this ICLR cycle: (1) More AI in the writing correlates with *lower* review scores. (2) More AI in the reviews correlates with *higher* scores. Lots of room to improve our tech tools (and our habits!)
  • @mbodhisattwa Bodhisattwa Majumder on x
    Super interesting! Though I'd be wary of false positives & cases of AI editing when disclosed. E.g., one of my reviews is tagged as “lightly AI-edited,” which is, of course, wrong. :) Though I think @pangramlabs's false positives for fully AI-generation are very, very low.
  • @dogacel0 @dogacel0 on x
    I wanted to uncover some interactions in ICLR data on AI-usage in submissions and reviews, so I analyzed it further. What surprised me is that even the fully AI reviews gave lower scores to submissions with more AI-generated content on average. AI still prefers human-written [ima…
  • @leonpalafox Leon Palafox on x
    @gneubig @pangramlabs I'm afraid this may flag people who are not as fluent in English and just wanted it to help them correct grammar
  • @gneubig Graham Neubig on x
    Obvious caveat: LLM-generated text detection is not perfect so there will be mistakes. Take this as a guide, not as the truth. But the reason why I asked for this was because I suspected several of my reviews were AI generated, and the results matched w/ my intuition.
  • @shuaichenchang Shuaichen Chang on x
    @gneubig @pangramlabs From my own experience: I have one submission where all four reviews were flagged as “Fully AI-generated.” Interestingly, I actually found the reviews quite reasonable, and we're addressing them seriously. I've also reviewed five papers myself. I write the r…
  • @xing_rui12683 @xing_rui12683 on x
    I check my reviews. Result in 2 moderate AI edited, 2 heavily and 1 light. Because I wrote review in Chinese and gpt helps me to translate into English. It is not a surprising result. But I think i am a responsible Reviewer :)
  • @kahnchana Kanchana Ranasinghe on x
    @gneubig @pangramlabs Interesting! One paper I reviewed (which I actually liked) was withdrawn after getting two bad review scores (from other reviewers). This analysis tags two of those reviews as 100% AI generated.
  • @max_spero_ Max Spero on x
    @gneubig @pangramlabs Thanks for the push to do this. People deserve more transparency around AI use. Nobody wants to spend time digging into a nonsensical paper that is largely AI-generated. And nobody wants to spend hours on new experiments to respond to a review that came stra…
  • @haryoaw @haryoaw on x
    Should we add a feature to OpenReview to show the confidence percentage of a paper and its reviews, indicating whether they are AI-generated?
  • @max_spero_ Max Spero on x
    We were curious about our false positive rate, so we ran all ICLR 2022 reviews (pre-ChatGPT) as a baseline. Lightly AI-edited FPR: 1 in 1,000 Moderately AI-edited FPR: 1 in 5,000 Heavily AI-edited FPR: 1 in 10,000 Fully AI-generated: No false positives [image]
  • @gneubig Graham Neubig on x
    ICLR authors, want to check if your reviews are likely AI generated? ICLR reviewers, want to check if your paper is likely AI generated? Here are AI detection results for every ICLR paper and review from @pangramlabs! It seems that ~21% of reviews may be AI? [image]
  • @katjadiehl Katja Diehl on bluesky
    It's a slop.  —  “Major AI conference flooded with peer reviews written by AI.  —  Controversy has erupted after 21% of manuscript reviews for an international AI conference were found to be generated by artificial intelligence.  I don't want mobility to be planned by AI.  —  www…
  • r/nottheonion r on reddit
    Major AI conference flooded with peer reviews written fully by AI