Anthropic announces Claude 3 Opus, Sonnet, and Haiku, aiming to reduce AI model hallucinations; Opus and Sonnet are available now, and Haiku in the coming weeks
The startup says new versions of Claude will be twice as likely to answer a question correctly.
BloombergRachel Metz
Context & Ripple Effects
Claude 3 is Anthropic’s initial three-tier release: Opus and Sonnet target capability, while Haiku was positioned shortly afterward for high-volume, latency-sensitive use cases. The common claim across the launch is that reducing incorrect answers is central to making the models more dependable.
Anthropic immediately expands Claude into distinct performance and deployment tiers, giving users a choice between available Opus and Sonnet models while awaiting Haiku.
The company makes answer accuracy a core product claim, raising the practical bar for teams evaluating Claude for tasks where unsupported responses are costly.
Second-order effects
Model providers competing for the same users face pressure to differentiate on both reliability and fit-for-purpose tiers, rather than benchmark capability alone.
Customers can more readily match model capability to workload needs; Haiku’s subsequent positioning around speed and volume makes latency and operating economics part of the selection decision.
Third-order effects
If tiered releases continue to improve reliability, competition is likely to move toward AI portfolios optimized for specific workflows, not one universal flagship model.
The durable measure of progress becomes cost per dependable outcome: headline model gains matter only insofar as users can deploy them with fewer errors and acceptable latency.
The trend: This is one data point in the industrialization of generative AI, where vendors package improving model reliability into workload-specific tiers and product experiences.
News: Anthropic has released Claude 3, a trio of AI models it says can outperform rivals like OpenAI's GPT-4 and Google's Gemini 1 Ultra. @kenrickcai and I spoke to cofounders Dario and Daniela Amodei about the release for @Forbes. https://www.forbes.com/...
@kenrickcai @Forbes Anthropic's new flagship model, Claude 3 Opus, beat GPT-4 and Gemini on a number of benchmarks. But it's pricy, and CEO Amodei admitted it's unknown how it fully stacks up against unreleased models like OpenAI's GPT 4 Turbo or Google's Gemini 1.5 Ultra. https:…
@kenrickcai ... We spoke to Anthropic about perceptions from some that it's models have degraded over time; on the LMSYS leaderboard, Claude 1 ranks higher than Claude 2. Amodei said Claude 3 has been trained to generate far fewer “incorrect refusals” than its predecessor, withou…
Thrilled about these new models - I've been playing around with Claude 3 Opus a lot and it's very capable and useful. Like with most frontier models, it has chewed through a bunch of evals so we need to now build more complicated evals to better understand its capabilities.
And then there were three... I got access to the new Anthropic Claude 3 AI a few days ago, so not enough time for a full review, but it was obvious it was GPT-4 class even before they released the testing stats. At the same time, like Gemini Advanced, it doesn't blow GPT-4 away. …
Haiku is the fastest and most cost-effective model on the market for its intelligence category. For the vast majority of workloads, Sonnet is 2x faster than Claude 2 and Claude 2.1, while Opus is about the same speed as past models.
Opus and Sonnet are accessible in our API which is now generally available, enabling developers to start using these models immediately. Sonnet is powering the free experience on https://claude.ai/, with Opus available for Claude Pro subscribers.
Claude 3 offers sophisticated vision capabilities on par with other leading models. The models can process a wide range of visual formats, including photos, charts, graphs and technical diagrams. [video]
Today, we're announcing Claude 3, our next generation of AI models. The three state-of-the-art models—Claude 3 Opus, Claude 3 Sonnet, and Claude 3 Haiku—set new industry benchmarks across reasoning, math, coding, multilingual understanding, and vision. [image]