A Google DeepMind research scientist details some LLM use cases and argues LLMs are not overhyped and should be judged on what they can do, not what they can't
I don't think that “AI” models (by which I mean: large language models) are over-hyped. — Yes, it's true that any new technology will attract the grifters. And it is definitely true that many companies like to say they're “Using AI” in the same way they previously said they were powered by “The Blockchain”. … X: Pete Skomoroch / @peteskomoroch : Excellent first-hand account of one person's productivity gains from LLMs: https://nicholas.carlini.com/ ... [image] Mark Anderson / @mandercorn : This post showcases concrete ways LLMs can supercharge the thinking and focus of a user—if you know how to harness their power. https://nicholas.carlini.com/ ... James O'Leary / @jpohhhh : this is soooooo good. https://nicholas.carlini.com/ ... Reilly Wood / @reillywood : OK this is a much better version of what I wrote: https://nicholas.carlini.com/ ... [image] Eduardo Robles / @edulix : ChatGPT is basically a productivity tool. Sound boring, but it's awesome. The post below explains how a technical guy (Research Scientist at Google DeepMind) uses it. Let's thank him. We can now link that to answer “are you guys using LLMs? how?” https://nicholas.carlini.com/ ... Dare Obasanjo / @carnage4life : Nicholas Carlini has published 50 conversations he has had with LLMs where he's asked them everything from writing entire web apps to teaching him new programming frameworks. https://nicholas.carlini.com/ ... [image] Brandon Downey / @bdowney : @tqbf This rings true to me, and also my broad opinion that they are a productivity tool, but not a *transcendent* one. 10-50% productivity gain. Which is huge! But maybe not so huge as the volume of investment so far warrants, and def not the hype. Thomas H. Ptacek / @tqbf : Nicholas Carlini is one of the sharper people I have ever met and I pay attention to anything he writes; this, on day-to-day utility of LLMs, rings pretty true to me. https://nicholas.carlini.com/ ... Ofir Press / @ofirpress : Super interesting post by Carlini on how he uses LMs day to day. He estimates that LMs do at least 50% of the programming work he does. That's amazing! Forums: Hacker News : How I Use “AI”
Context & Ripple Effects
The discussion follows a practical turn in LLM coverage: practitioners had already catalogued lessons and prompting pitfalls from production LLM applications, shifting attention from model demos to repeatable work habits. Carlini adds a firsthand programming-workflow account from inside Google DeepMind.
The claim also sits between ambitious predictions that LLMs could lower the barrier to software creation and the standing critique that their limits must be taken seriously. His proposed test is narrower: assess task-level usefulness rather than demand universal competence.
First-order effects
- For Carlini and similarly situated technical users, LLMs become a routine input to programming work, with his reported estimate that they handle at least half of his programming tasks serving as a concrete productivity benchmark.
- The post gives Google DeepMind and LLM advocates a use-case-led defense against generic hype claims, while keeping evaluation tied to the work a user can actually complete.
Second-order effects
- Teams evaluating coding assistants have more reason to measure workflow outcomes—such as completed programming work and review burden—rather than treating “AI use” as a standalone adoption metric; the earlier production-app lessons reinforce that implementation details matter.
- Competing model and developer-tool providers are pushed to demonstrate dependable task performance and integration into existing workflows, not merely broad benchmark or chatbot claims.
Third-order effects
- If task-specific productivity accounts continue to accumulate, LLM competition is likely to be judged increasingly on cost per useful completed task and workflow fit rather than on claims of general intelligence.
- That would make adoption more uneven but more durable: tools that can be supervised and embedded in concrete work may spread even while the limitations emphasized by critics remain material.
The trend: LLMs are moving from a debate over abstract capability and hype toward evidence-based adoption around narrowly defined, human-supervised workflows.