/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

AI training data provider Scale AI releases SEAL Leaderboards, which uses private datasets to rank LLMs in domains like coding, instruction following, and math

SiliconANGLE Mike Wheatley

Context & Ripple Effects

Scale AI’s leaderboard release extends its role beyond supplying training data into measuring model performance. Its earlier Pentagon contract to test and evaluate LLMs showed that evaluation itself could be a paid, high-stakes service, not merely a research exercise.

The use of private datasets makes the benchmark a differentiated asset rather than a simple aggregation of public test scores. That approach later broadened into a public SEAL Showdown based on contributor votes, suggesting an effort to cover both controlled domain tests and everyday usability.

First-order effects

  • Model developers gain another external scorecard for coding, instruction following, and math, while Scale AI gains a product that can showcase the value of its proprietary data and evaluation workflow.
  • Buyers comparing LLMs can use SEAL results as an additional input, though private test sets limit independent inspection of the underlying tasks and scoring choices.

Second-order effects

  • Competing model providers face pressure to optimize for a growing set of third-party evaluations, while evaluation vendors and data-curation firms must differentiate on test quality, domain coverage, and trustworthiness.
  • Private benchmarks increase the commercial value of high-quality held-out data; this aligns with the adjacent push to better curate AI training datasets, where data selection and evaluation become linked markets.

Third-order effects

  • If proprietary evaluation suites become influential in purchasing and deployment decisions, control of credible test data can become a strategic layer of the AI stack alongside model development and data labeling.
  • The trade-off will be durable: private datasets may resist benchmark contamination better than public tests, but their influence depends on customers accepting limited transparency and governance around scoring.

The trend: AI data suppliers are expanding into evaluation infrastructure, seeking to turn proprietary datasets into recurring influence over which models enterprises and institutions select.

Discussion

  • @akchay_sri Akchay Srivastava on x
    Why LLM evaluations are a complex challenge. Insights like these are crucial for AI/ML testers to build strong & dependable testing for LLM-powered features. #AI #MachineLearning #LLM #Testing
  • @vjkaruna Vijay Karunamurthy on x
    Better evaluations means faster progress in AI - both on new capabilities and responsible deployment. Excited to launch this new initiative - including important results on how leading models perform in Spanish.
  • @arankomatsuzaki Aran Komatsuzaki on x
    ScaleAI just released LLM leaderboards by extending GSM1K to various domains! This can be a great complement to lmsys eval.
  • @natolambert Nathan Lambert on x
    Glad to see scale going in this direction: Fully private LLM leaderboard. Even has the fine print 🤓: “To ensure leaderboard integrity, we require that models can only be featured the FIRST TIME when an organization encounters the prompts.”
  • @0interestrates Rahul on x
    looks promising first leaderboard i have seen with gpt4-turbo > gpt-4o for coding
  • @razrazcle Raza Habib on x
    Mode model evals ≠ better performance. Remember: More model evals are great but in practice there's a HUGE gap between evaluating a model in the abstract vs. evaluating a model for a specific use case. Usually the question you're asking is not, “Is GPT-4 better overall than
  • @altryne Alex Volkov on x
    As predicted, scale is entering the LLM eval game, but with a private (read: non trainable on) evals for frontier models! This is great, very trusted resource, in addition to LMSys, Reddit Vibes, X shitpoasting and broken open evals. Will cover tomorrow on @thursdai_pod !
  • @packym Packy McCormick on x
    My favorite wordcel is also a shape rotator. [image]
  • @natfriedman Nat Friedman on x
    We're going to need a lot more investment in high-quality evals and benchmarks to help us understand the actual comparative utility of the various models. This new set of private evals and leaderboard from Scale are great to see.
  • @karpathy Andrej Karpathy on x
    Nice, a serious contender to @lmsysorg in evaluating LLMs has entered the chat. LLM evals are improving, but not so long ago their state was very bleak, with qualitative experience very often disagreeing with quantitative rankings. This is because good evals are very difficult [i…
  • @madiator Mahesh Sathiamoorthy on x
    Nice, private evals that can't be gamed.
  • @dmdohan David Dohan on x
    Awesome to have more unbiased & difficult to game evals Hope we start to see more dynamic range to differentiate performance on harder problems. As this shows, evals similar to gsm8k are ~saturated
  • @dustinvtran Dustin Tran on x
    New public leaderboard from Scale! It looks like a solid set of evals. Mitigates two of the biggest problems in evals today: eval sets contaminated in model training, and rater quality for human evaluation.
  • @summeryue0 Summer Yue on x
    🚀 Instruction Following - SEAL Leaderboards are out! IF winners: - GPT-4o and GPT-4 Turbo - Llama 3 70B Instruct - Mistral Large Gemini Pro 1.5 leaps into top 3 in preference rankings, and Claude rockets to #2 in factuality. See https://scale.com/... [image]
  • @summeryue0 Summer Yue on x
    🚀 Coding - The first expert evaluated SEAL Leaderboards are out! The coding race is neck and neck, winners: - GPT-4 Turbo and GPT-4o - Gemini Pro 1.5 - Claude 3 Opus See https://scale.com/... for details detailed analysis for each model! [image]
  • @summeryue0 Summer Yue on x
    🚀 Math - we released the GSM1k last month. Today, we augmented it with human ratings to account for chatty yet correct responses. Explore the GSM1k leaderboard as part of SEAL Leaderboards. We were glad to see LLMs have mostly nailed grade school math! [image]
  • @herbiebradley Herbie Bradley on x
    very impressive & thorough work, congrats @summeryue0 @alexandr_wang ! i think private datasets *or* dynamically generated evals are the future for LLM benchmarking, and this is a great start on the first of these
  • @summeryue0 Summer Yue on x
    🚀 Spanish - The first expert evaluated SEAL Leaderboards are out! Spanish is our first multilingual leaderboard ( https://scale.com/...), winners: - GPT-4o - Gemini 1.5 Pro (post-I/O) - GPT-4 Turbo We plan to roll out more languages, which ones should we build next? [image]
  • @goodside Riley Goodside on x
    New from Scale: SEAL Leaderboards — a new benchmark arena for frontier LLMs - Private, novel assessments that models can't train on - ELO-scale rankings (via Bradley-Terry) - Domain leaderboards (today: coding, math, instruct, Spanish — more soon!) (Links in reply)
  • @summeryue0 Summer Yue on x
    🚀 Introducing the SEAL Leaderboards! We rank LLMs using private datasets that can't be gamed. Vetted experts handle the ratings, and we share our methods in detail openly! Check out our leaderboards at https://scale.com/...! Which evals should we build next? [image]
  • @alexandr_wang Alexandr Wang on x
    1/ We are launching SEAL Leaderboards—private, expert evaluations of leading frontier models. Our design principles: 🔒Private + Unexploitable. No overfitting on evals! 🎓Domain Expert Evals 🏆Continuously Updated w/new Data and Models Read more in 🧵 https://scale.com/... [image]
  • @danielxberrios Daniel Berrios on x
    Excited to be launching the first-of-their-kind SEAL Leaderboards for LLMs. Everyone knows evals are broken, and we want to help fix that 🥇⚖️🔒 https://scale.com/... We're also taking the Scale Evaluation platform into GA, and look forward to getting it into more hands! 🚀 [image]
  • @scale_ai @scale_ai on x
    Scale is excited to release the SEAL leaderboards which rank frontier LLMs, kicking off the first truly expert-driven, trustworthy LLM contest open to all. https://scl.ai/... [image]