Anthropic says Opus 4.6 supports a 1M context window in beta, scored 90.2% on BigLaw Bench, the highest for any Claude model, and boosts agentic capabilities
David Gewirtz /ZDNET:NEW
Context & Ripple Effects
Anthropic’s Opus line has moved from the company’s November positioning of Opus 4.5 around coding, agents and computer use to a new release centered on longer input capacity and more autonomous task handling. The reported 1M-token beta window is a material step beyond Claude 2.1’s earlier 200K-token window.
The release also pairs capacity with task-specific positioning: Anthropic says Opus 4.6 can work across company data, regulatory filings and market information for detailed financial analysis. Its reported greater focus on difficult parts of a task and BigLaw Bench result make legal and knowledge-work evaluation a prominent part of the launch narrative.
First-order effects
- Claude users and developers can test a 1M-token context window in beta, potentially reducing the need to split large source collections across separate prompts or workflows.
- Anthropic gains a new enterprise-facing proof point for Opus 4.6: a reported 90.2% BigLaw Bench score, alongside claims of stronger agentic behavior and analysis across financial materials.
Second-order effects
- Teams building document-heavy legal, finance and research workflows will need to reassess context design: larger windows can simplify ingestion, but they also make selection, cost control and evaluation of long-running agent tasks more consequential.
- Competing model providers face pressure to match not only headline context capacity but also evidence that large-context models can sustain performance on domain-specific tasks such as analysis of company and regulatory information.
Third-order effects
- If large-context, agent-oriented models prove dependable in production, differentiation may shift from isolated chatbot answers toward systems that can retain and act on a broader working set of enterprise materials.
- That shift would elevate context engineering and domain benchmarks as buying criteria; reported benchmark leadership alone will not settle whether an agent is reliable enough for consequential professional workflows.
The trend: This is part of the shift from conversational models toward agentic work systems competing on usable context, task persistence and domain-specific performance.