A profile of Anthropic and its key executives like Chris Olah, and a look at Project Vend, an internal “Claudius” experiment to run the office vending machine
Researchers at the company are trying to understand their A.I. system's mind—examining its neurons …
New YorkerGideon Lewis-Kraus
Context & Ripple Effects
Project Vend extends Anthropic's earlier test of Claude running a physical storefront, where pricing and inventory management proved difficult. The vending-machine exercise gives the company a smaller, repeatable setting for observing how an AI system makes operational decisions.
The profile also centers Anthropic's effort to inspect model neurons, linking practical autonomy tests with the lab's longer-standing safety-oriented culture described in earlier reporting on its safety focus.
First-order effects
Anthropic gains an internal testbed for evaluating Claudius on concrete operational tasks such as pricing and inventory, rather than solely on conversational or benchmark performance.
Researchers can compare the system's observable decisions with neuron-level investigations, creating a tighter feedback loop between interpretability work and deployment behavior.
Second-order effects
The earlier storefront results mean Project Vend can expose whether operational failures are repeatable across constrained real-world settings, sharpening the criteria Anthropic uses before offering more agentic capabilities to customers.
Competitors pursuing AI agents face a similar pressure to demonstrate reliability in mundane business workflows, where errors in inventory or pricing are visible and measurable.
Third-order effects
If labs increasingly pair interpretability research with bounded workplace trials, agent evaluation may shift toward continuous operational evidence rather than one-time model benchmarks.
The pattern could make transparency about how agents are tested a differentiator for enterprise adoption, though a single internal vending-machine experiment cannot establish that outcome.
The trend: AI labs are moving from general-purpose demonstrations toward tightly bounded real-world agent trials that test both operational reliability and model behavior.
“Now we have a special box of electricity that turns Reddit comments and old toaster manuals into cogent conversations about Shakespeare and molecular biology.” Great profile from Gideon Lewis-Kraus on state of LLMs 10 years after his (amazing) NYT article on Google Translate. [i…
i cannot believe how rarely people ask the simple question of “how do they pay for this?” Anthropic lost $5.2bn *after* $4.5bn of revenue in 2025! They're feeding equity to opex again and again, this is literally a test of whether venture capital can keep it alive
Experiments conducted with the A.I. system Claude are producing fascinating results—and raising questions about the nature of selfhood. Gideon Lewis-Kraus reports from inside the company that designed it, Anthropic. https://newyorkermag.visitlink.me/ wQBoc2
Of course @AnthropicAI will take zero responsibility for this. Whatever they want to do to go fast, they have to🤷♀️ But it's so much worse than them flubbing and letting their own research into the training set. Anthropic employees literally go around telling academics and
#Anthropic is targeting ~10 GW of Data Center capacity over the next several years, a scale that implies hundreds of billions in total capex and far exceeds previously disclosed $180B planned server spend through 2029, and $50B Fluidstack partnership disclosed to date. Ambition
another thing that folks made fun of when it was said, that things you write into the data will affect later behaviour, that's straightforwardly coming true
Kinda cool that the Anthropic comms team and execs did this. Embedding the New Yorker's Gideon Lewis-Kraus and letting the chips fall where they may is almost punk rock at this stage of the “go direct” era. https://www.newyorker.com/...
How A.I. works is a mystery even to those developing it. “It's like we understand aviation at the level of the Wright brothers, but we went straight to building a 747 and making it a part of normal life,” one researcher said. https://newyorkermag.visitlink.me/ wtzNII
Black box peril: “It wouldn't be fair...to have a machine evaluate an applicant's mortgage eligibility in an opaque way. And, if you employed a robot to keep your house clean of dog hair, you wanted to be certain that it would vacuum the couch, not kill the dog.” — www.newyork…
Claude was my first time talking to an LLM and seeing an English major's wit reflected back in an uncanny valley sort of way. whether it was any good at mundane tasks consistently was another story
Gideon Lewis-Kraus only writes about AI like once every 5 years, but when he does, it's a gift his Anthropic profile contains the kind of lyricism generally reserved for describing fine art! [image]