Stanford unveils the Foundation Model Transparency Index, featuring 100 indicators; Llama 2 led at 54%, GPT-4 placed third at 48%, and PaLM 2 took fifth at 40%
https://www.nytimes.com/... [image] Mark Coggins / @coggins@mastodon.social : This is the kind of needed AI regulation—requiring model makers to reveal how they trained their language models—that #MarcAndreessen railed against in his recent “manifesto.” — https://www.nytimes.com/... #ai #artificialintelligence #llm X: @stanfordhai : Foundation model developers are becoming less transparent, making it difficult for researchers, policymakers, and the public to understand how these models work, as well as their limitations and impact. Scholars assess the transparency of 10 developers. https://hai.stanford.edu/... [image] Percy Liang / @percyliang : As capabilities of foundation models are waxing, *transparency* is waning. How do we quantify transparency? We introduce the Foundation Models Transparency Index (FMTI), evaluating 10 foundation model developers on 100 indicators. https://crfm.stanford.edu/fmti/ [image] Dina Bass / @dinabass : Super-interesting. The report looks at various categories of transparency. The ones were everyone is largely terrible: labor (except for Bloomz), feedback and impact. Louise Matsakis / @lmatsakis : the new Foundation Model Transparency Index is so useful. Did not expect Amazon to score the lowest here tbh https://crfm.stanford.edu/fmti/ [image] @sreekanthmreddy : So, the 2 most transparent models are open souce! @sashamtl : I get that communicating about ‘dimensions of transparency’ is important, but surely we want *actual access* to each model/data/code in order to verify whether what's reported is actually true? (Pics or it didn't happen) Shayne Longpre / @shayneredford : 📢 The🌟Foundation Model Transparency Index🌟 scores 10 developers on 💯 indicators. 1⃣ All 10 score poorly, particularly Data, Labor, Compute 2⃣ Transparency is possible! 82/100 are scored by >1 3⃣ Transparency is a precondition for informed & responsible AI policy. 1/ [image] Clem / @clementdelangue : The scores + ranking look super weird and inconsistent to me & a lot of interesting open models are missing but I love the message from @Stanford @StanfordHAI. More transparency = more safety for AI https://hai.stanford.edu/... [image] Rachel Metz / @rachelmetz : interesting research just dropped from Stanford/MIT/Princeton: the foundation model transparency index, constructed by analyzing public info from 10 key AI companies. nobody did great! @AIatMeta #1 w/ Llama 2, @huggingface #2 w/BLOOMZ, @OpenAI #3 w/GPT-4. https://crfm.stanford.edu/fmti/ [image] Percy Liang / @percyliang : Open developers (Meta, Hugging Face, Stability) are more transparent (all score in the top 4 and well above the average). Much of that margin comes from greater upstream transparency. Closed developers can control downstream use, but this does not transfer to transparency. [image] @josephjacks_ : “No major foundation model developer is close to providing adequate transparency, revealing a fundamental lack of transparency in the AI industry.”... (LLaMa 2) only scores 54% on Stanford's foundation model transparency index: https://crfm.stanford.edu/fmti/ [image] Sayash Kapoor / @sayashk : Foundation models have profound societal impact, but transparency about these models is waning. Today, we are launching the Foundation Model Transparency Index, which offers a deep dive into the transparency practices and standards of key AI developers. https://crfm.stanford.edu/fmti/ [image] Arvind Narayanan / @random_walker : This is a really impressive and thorough effort from Stanford, MIT, and Princeton researchers to document the (lack of) transparency of 10 major foundation models on 100 transparency indicators. 👏 https://crfm.stanford.edu/fmti/ Blog https://www.aisnakeoil.com/... Paper https://crfm.stanford.edu/... [image] Forums: r/singularity : Stanford researchers have ranked 10 foundation A.I. models on how openly they operate.The most transparent model was LLaMA 2, with a score of 54 % …
Context & Ripple Effects
The index turns a broad transparency debate into a comparable scorecard across major foundation-model developers. It follows research finding providers could feasibly meet draft EU AI Act transparency requirements, undercutting the claim that meaningful disclosure is impractical.
The gap also matters against a competing definition of openness: the Allen Institute’s planned open generative model OLMo shows that model access and model-maker disclosure are related but distinct forms of accountability.
First-order effects
- Llama 2, GPT-4, and PaLM 2 receive public, comparable transparency benchmarks, while the researchers’ conclusion that no major developer is adequately transparent puts all evaluated providers under the same scrutiny.
- Researchers, policymakers, and prospective model users gain a 100-indicator framework for identifying what developers do and do not disclose about their systems.
Second-order effects
- Model developers face pressure to improve documentation and disclosure practices, not merely model capability, because relative transparency is now measurable across peers.
- The index gives policy discussions a concrete baseline for testing whether voluntary disclosures approach the level researchers previously found feasible under proposed EU rules.
Third-order effects
- If repeated and adopted, standardized transparency scoring could make disclosure quality a durable competitive and governance dimension for foundation models alongside performance and access.
- The pattern points toward accountability regimes that evaluate model developers through auditable evidence rather than self-described openness; whether they become enforceable standards remains uncertain.
The trend: Foundation-model governance is moving from general calls for openness toward standardized, comparable disclosure requirements.