2025 LLM Year in Review: shift toward RLVR, Claude Code emerged as the first convincing example of an LLM agent, Nano Banana was paradigm shifting, and more
The year’s LLM discussion had already moved from broad multimodal adoption and falling prices toward practical capability questions. Karpathy’s earlier “Software 3.0” framing positioned natural-language programming and agents as the next interface layer.
RLVR becomes a clearer organizing priority for LLM development, shifting attention toward systems whose reasoning can be checked against verifiable outcomes.
Claude gains a concrete reference point in the agent market: Claude Code is presented as the first persuasive example of an LLM agent, raising the bar from chat assistance to useful task execution.
Second-order effects
Competing model providers and coding-tool vendors face pressure to demonstrate similarly dependable agent workflows, not just stronger general-purpose model outputs.
The reported importance of Nano Banana broadens the competitive focus beyond text reasoning and coding, making differentiated multimodal interaction a more consequential product frontier.
Third-order effects
If verifiable-reward training and useful agents continue to reinforce one another, LLM competition may increasingly hinge on reliability in bounded workflows rather than on conversational fluency alone.
The pattern points toward software products organized around agents that execute work with oversight, though the durability of that shift depends on whether performance generalizes beyond early coding use cases.
The trend: LLMs are moving from broadly capable assistants toward systems differentiated by verifiable reasoning, agentic execution, and new multimodal interfaces.
“In 2025, Reinforcement Learning from Verifiable Rewards (RLVR) emerged as the de facto new major stage” “Supervision bits-wise, human neural nets are optimized for survival ... but LLM neural nets are optimized for imitating humanity's text”
JUST IN: Karpathy no longer trust LLM benchmarks. It's not long ago that LMArena is his go to benchmark to guage how good a model is in real world settings. Now he simplies disregard benchmarks altogether due to the prevalence of benchmaxxing. [image]
Nice article. I can't emphasize the “vibe coding” section enough. The ability to create your own software tools through text and/or speech is so incredibly liberating and empowering. The (digital) world is your oyster!
- Claude have made some big strides in attracting enterprise users with their high quality agents. - OAI have had some public sentiment shift but shouldn't under estimate how sticky the ChatGPT interface is. Plus Jonny Ive's design still to come (!) - Wonder how Gemini and Xai
“The core issue is that benchmarks are almost by construction verifiable environments and are therefore immediately susceptible to RLVR and weaker forms of it via synthetic data generation.”
“I suspect that LLM labs will trend to graduate the generally capable college student, but LLM apps will organize, finetune and actually animate teams of them into deployed professionals in specific verticals by supplying private data, sensors and actuators and feedback loops.”
Benchmarks are a rigged game We've all known it. But have we played out where this nets out? Eventually, we'll ignore them completely For now, we're in an uncanny valley. We think they're strictly better than no approximation of a model's capability. But that's a shaky
Andrej's 2025 LLM Review is the PewDiePie YouTube Rewind of AI. His naming superpower is unmatched. 2 tweets. 2 most influential AI terms of the year: - “Vibe coding” - “Jagged Intelligence” [image]