[Thread] A US paper shows the best frontier LLM models solve 0% of hard coding problems from Codeforces, ICPC, and IOI, domains where expert humans still excel
This is really BAD news of LLM's coding skill. ☹️ The best Frontier LLM models achieve 0% on hard real-life Programming Contest problems, domains where expert humans still excel. LiveCodeBench Pro, a ...
OpenAI's o1 models aren't a straightforward upgrade to GPT-4o, as they introduce some major cost and performance trade-offs in exchange for improved “reasoning”
delving into OpenAI's new ‘o1’ model PYMNTS.com : OpenAI's ‘Strawberry’ Model Sparks Fresh Discussions on AI Capabilities M.G. Siegler / Spyglass : OpenAI Reasons ‘o1’ is a Better Name than ‘Strawberr...
OpenAI claims that in a qualifying exam for the International Mathematics Olympiad, o1 correctly solved 83.3% of the problems, while GPT-4o solved only 13.4%
Sam Altman says it “doesn't constitute AGI” Poulami Saha / Financial Express : OpenAI makes big AI breakthrough, ChatGPT can now think and reason: Details Emilia David / VentureBeat : How to prompt on...