OpenAI's GPT-4 failed a CFA Institute exam, the finance world's self-styled toughest test, scoring eight out of 24, despite the answers being available online
Artificial intelligence versus arbitrary irrelevance — It's an algorithmic mystery box that inspires fear, awe and derision in equal measure. Tweets: @bgurley , @mbrookerhk , @robinwigg , and @zerobeta Tweets: Bill Gurley / @bgurley : I have never been prouder of my CFA credentials! Apparently it did quite well on other standardized tests... https://twitter.com/... Matthew Brooker / @mbrookerhk : Good news: ChatGPT would probably fail a CFA exam https://www.ft.com/... Robin Wigglesworth / @robinwigg : Rejoyce CFAs! Banger of a post from Bryce. https://twitter.com/... @zerobeta : Time for ChatGPT to start a “family office”, engagement farm on Twitter, and roll it into a shitty substack. https://twitter.com/...
Context & Ripple Effects
The CFA result lands days after GPT-4's launch cycle, in which coverage framed it as more precise and accurate than its predecessor and an "outrageous, if imperfect" step forward — and one week before Microsoft researchers would claim early signs of AGI and near-human performance across coding, medicine, and law. Scoring eight out of 24 on the CFA exam, with answers available online, is the first hard counterweight to that narrative.
The finance commentariat — Bill Gurley among them — seized on it as vindication for the credential, but the longer arc matters more: by the time GPT-5 shipped as an underwhelming, incrementally improved release, the gap between headline benchmarks and durable professional competence looked structural rather than transitional.
First-order effects
- CFA charterholders get a rare public data point that their exam resists large language models, blunting OpenAI's positioning of GPT-4 as a top-tier test-taker across standardized tests.
- OpenAI faces an immediate credibility problem: the same week its model is being described in near-human terms, it fails a professional finance exam whose answers were accessible online.
Second-order effects
- Competing labs' benchmark claims come under sharper scrutiny, since the CFA result shows vendor-selected evaluations can flatter a model while domain-specific exams expose gaps.
- Professional credentialing bodies gain leverage to treat AI performance skeptically — either tightening exam security or resisting pressure to accept AI-assisted candidates.
Third-order effects
- If the pattern holds from GPT-4 through the GPT-4.5 research preview and GPT-5's incremental reception, the industry shifts from benchmark-led marketing toward harder, domain-specific evaluation — and buyers learn to discount vendor-reported scores when making adoption decisions.
The trend: AI capability claims are increasingly tested against real professional exams rather than curated benchmarks, exposing a persistent gap between headline scores and applied competence that later releases have narrowed slowly.