AI training data provider Scale AI releases SEAL Leaderboards, which uses private datasets to rank LLMs in domains like coding, instruction following, and math
SiliconANGLEMike Wheatley
Context & Ripple Effects
Scale AI’s leaderboard release extends its role beyond supplying training data into measuring model performance. Its earlier Pentagon contract to test and evaluate LLMs showed that evaluation itself could be a paid, high-stakes service, not merely a research exercise.
The use of private datasets makes the benchmark a differentiated asset rather than a simple aggregation of public test scores. That approach later broadened into a public SEAL Showdown based on contributor votes, suggesting an effort to cover both controlled domain tests and everyday usability.
First-order effects
Model developers gain another external scorecard for coding, instruction following, and math, while Scale AI gains a product that can showcase the value of its proprietary data and evaluation workflow.
Buyers comparing LLMs can use SEAL results as an additional input, though private test sets limit independent inspection of the underlying tasks and scoring choices.
Second-order effects
Competing model providers face pressure to optimize for a growing set of third-party evaluations, while evaluation vendors and data-curation firms must differentiate on test quality, domain coverage, and trustworthiness.
Private benchmarks increase the commercial value of high-quality held-out data; this aligns with the adjacent push to better curate AI training datasets, where data selection and evaluation become linked markets.
Third-order effects
If proprietary evaluation suites become influential in purchasing and deployment decisions, control of credible test data can become a strategic layer of the AI stack alongside model development and data labeling.
The trade-off will be durable: private datasets may resist benchmark contamination better than public tests, but their influence depends on customers accepting limited transparency and governance around scoring.
The trend: AI data suppliers are expanding into evaluation infrastructure, seeking to turn proprietary datasets into recurring influence over which models enterprises and institutions select.
Why LLM evaluations are a complex challenge. Insights like these are crucial for AI/ML testers to build strong & dependable testing for LLM-powered features. #AI #MachineLearning #LLM #Testing
Better evaluations means faster progress in AI - both on new capabilities and responsible deployment. Excited to launch this new initiative - including important results on how leading models perform in Spanish.
Glad to see scale going in this direction: Fully private LLM leaderboard. Even has the fine print 🤓: “To ensure leaderboard integrity, we require that models can only be featured the FIRST TIME when an organization encounters the prompts.”
Mode model evals ≠ better performance. Remember: More model evals are great but in practice there's a HUGE gap between evaluating a model in the abstract vs. evaluating a model for a specific use case. Usually the question you're asking is not, “Is GPT-4 better overall than
As predicted, scale is entering the LLM eval game, but with a private (read: non trainable on) evals for frontier models! This is great, very trusted resource, in addition to LMSys, Reddit Vibes, X shitpoasting and broken open evals. Will cover tomorrow on @thursdai_pod !
We're going to need a lot more investment in high-quality evals and benchmarks to help us understand the actual comparative utility of the various models. This new set of private evals and leaderboard from Scale are great to see.
Nice, a serious contender to @lmsysorg in evaluating LLMs has entered the chat. LLM evals are improving, but not so long ago their state was very bleak, with qualitative experience very often disagreeing with quantitative rankings. This is because good evals are very difficult [i…
Awesome to have more unbiased & difficult to game evals Hope we start to see more dynamic range to differentiate performance on harder problems. As this shows, evals similar to gsm8k are ~saturated
New public leaderboard from Scale! It looks like a solid set of evals. Mitigates two of the biggest problems in evals today: eval sets contaminated in model training, and rater quality for human evaluation.
🚀 Instruction Following - SEAL Leaderboards are out! IF winners: - GPT-4o and GPT-4 Turbo - Llama 3 70B Instruct - Mistral Large Gemini Pro 1.5 leaps into top 3 in preference rankings, and Claude rockets to #2 in factuality. See https://scale.com/... [image]
🚀 Coding - The first expert evaluated SEAL Leaderboards are out! The coding race is neck and neck, winners: - GPT-4 Turbo and GPT-4o - Gemini Pro 1.5 - Claude 3 Opus See https://scale.com/... for details detailed analysis for each model! [image]
🚀 Math - we released the GSM1k last month. Today, we augmented it with human ratings to account for chatty yet correct responses. Explore the GSM1k leaderboard as part of SEAL Leaderboards. We were glad to see LLMs have mostly nailed grade school math! [image]
very impressive & thorough work, congrats @summeryue0 @alexandr_wang ! i think private datasets *or* dynamically generated evals are the future for LLM benchmarking, and this is a great start on the first of these
🚀 Spanish - The first expert evaluated SEAL Leaderboards are out! Spanish is our first multilingual leaderboard ( https://scale.com/...), winners: - GPT-4o - Gemini 1.5 Pro (post-I/O) - GPT-4 Turbo We plan to roll out more languages, which ones should we build next? [image]
New from Scale: SEAL Leaderboards — a new benchmark arena for frontier LLMs - Private, novel assessments that models can't train on - ELO-scale rankings (via Bradley-Terry) - Domain leaderboards (today: coding, math, instruct, Spanish — more soon!) (Links in reply)
🚀 Introducing the SEAL Leaderboards! We rank LLMs using private datasets that can't be gamed. Vetted experts handle the ratings, and we share our methods in detail openly! Check out our leaderboards at https://scale.com/...! Which evals should we build next? [image]
1/ We are launching SEAL Leaderboards—private, expert evaluations of leading frontier models. Our design principles: 🔒Private + Unexploitable. No overfitting on evals! 🎓Domain Expert Evals 🏆Continuously Updated w/new Data and Models Read more in 🧵 https://scale.com/... [image]
Excited to be launching the first-of-their-kind SEAL Leaderboards for LLMs. Everyone knows evals are broken, and we want to help fix that 🥇⚖️🔒 https://scale.com/... We're also taking the Scale Evaluation platform into GA, and look forward to getting it into more hands! 🚀 [image]
Scale is excited to release the SEAL leaderboards which rank frontier LLMs, kicking off the first truly expert-driven, trustworthy LLM contest open to all. https://scl.ai/... [image]