Cursor releases Composer 2.5, saying it's better at sustained work on long-running tasks and follows complex instructions more reliably; it's built on Kimi K2.5
Composer 2.5 is now available in Cursor. — It's a substantial improvement in intelligence and behavior over Composer 2.
Cursor
Context & Ripple Effects
Cursor’s Composer line moved from the March launch of Composer 2—positioned for autonomous, lengthy coding work—to an updated release shortly afterward. Moonshot and Cursor had separately clarified that Composer 2 began from Kimi K2.5, with access provided through Fireworks AI.
The update arrives as Cursor faces scrutiny over AI-assisted coding reliability, while its CEO has cautioned that “vibe coding” can leave advanced projects on unstable foundations. That makes sustained-task execution and instruction-following central product claims rather than merely incremental model benchmarks.
First-order effects
Cursor users gain Composer 2.5 in the product, with Cursor asserting better performance on long-running coding tasks and complex instructions.
Cursor further ties the Composer offering to Kimi K2.5, reinforcing Moonshot’s model as an underlying component of Cursor’s coding-agent stack.
Second-order effects
Cursor will be judged more directly on whether its coding agents can maintain correctness across extended workflows, particularly given the existing reliability criticism around AI-assisted development.
The release raises the competitive bar for coding-agent providers: product differentiation shifts toward dependable multi-step execution and adherence to detailed requirements, not simply code-generation quality.
Third-order effects
If these improvements prove durable, AI coding tools may increasingly compete as autonomous workflow systems, making reliability, supervision, and maintainability more consequential than isolated model capability claims.
Cursor’s dependence on an external foundation model illustrates a layered market in which end-user agent companies differentiate through training, behavior, and product integration while relying on upstream model providers and inference partners.
The trend: This is part of the shift from AI coding assistants that generate snippets toward agents expected to carry out longer, instruction-heavy software tasks with dependable behavior.
Introducing Composer 2.5, our most powerful model yet. It's more intelligent, better at sustained work on long-running tasks, and more reliable at following complex instructions. For the next week, we're doubling the included usage of the model. [image]
composer 2.5 is really really great. I had it on last week for some testing, forgot that it was on, & totally didn't realize I wasn't on gpt 5.5 (my usual) for a while. the team did a fantastic job!!
Together with SpaceXAI, we're training a significantly larger model from scratch, using 10x more total compute. With Colossus 2's million H100-equivalents and our combined data and training techniques, we expect this to be a major leap in model capability.
We improved Composer by scaling training, generating more complex RL environments, and introducing new learning methods. For example, we use text feedback during RL to learn faster by assigning credit in rollouts spanning hundreds of thousands of tokens.
cursor is at frontier scale, both in terms of performance and compute if composer 2.5's budget was put into a pre-train: ~6.3T total, 200B active trained on ~56T tokens if composer 3 allocates 50% of the budget to pre-training: ~500B active, 15.3T total trained on 135T tokens. [i…
Very cool to see Cursor doubling down on training great models. In my opinion, ultimately all serious companies in AI will want to train models themselves, based on open-source instead of outsourcing AI to others via APIs!