/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

OpenAI open sources Evals, its framework for automatically evaluating the performance of its AI models, letting users report shortcomings and guide improvements

TechCrunch Kyle Wiggers

Discussion

  • @drjimfan @drjimfan on x
    Want early access to GPT-4? Do it now: https://github.com/... is an official framework for evaluating OpenAI models. They will grant GPT-4 access to those who submit high quality evals. Thanks to my friend Andrew Kondrich @kondrich2 who built this initiative at OpenAI!
  • @e0m Evan Morikawa on x
    If you see something GPT-4 can't do well, or think you can prove a fundamental deficiency, contribute evals! This is by far the best way to help close these skill gaps. Internally we use evals to guide enormous amounts of model development. https://github.com/...
  • @swyx @swyx on x
    as LLMs grow and grow and grow in capabilities, it is getting more impt to have good model evaluation/benchmarking frameworks. OpenAI is also releasing their eval framework, fully MIT licensed: https://github.com/... Used by Stripe and well documented. Runs MMLU in 189 LOC https:…
  • @officiallogank @officiallogank on x
    We are giving priority GPT-4 access to those who contribute evals to our new evals repo: https://github.com/... Here, you can write tests for the model so we can improve things over time.