/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

OpenAI details how its Superalignment research team is exploring ways to control stronger AI models like GPT-4 using weaker supervisor models like GPT-2

We present a new research direction for superalignment … Alisa Davidson / Metaverse Post : OpenAI's Superalignment Team Unveils Innovative Method for AI System Oversight Abubakar Idris / The Messenger : OpenAI Develops New Test for Superhuman Artificial Intelligence Will Douglas Heaven / MIT Technology Review : Now we know what OpenAI's superalignment team has been up to Maximilian Schreiner / The Decoder : OpenAI's GPT-2 supervised GPT-4 in a glimpse into the future of AGI alignment Maria Deutscher / SiliconANGLE : OpenAI details automated approach to supervising AI models Kyle Wiggers / TechCrunch : OpenAI thinks superhuman AI is coming — and wants to build tools to control it Franklin Manuel / Baseline : OpenAI's Superalignment Team: A Mission to Control Superintelligent AI Eliza Strickland / IEEE Spectrum : OpenAI Demos a Control Method for Superintelligent AI X: Mira Murati / @miramurati : Exploring generalization properties of deep learning to control strong models with weak supervisors, showing early promise. Sam Altman / @sama : great work from the superalignment team: Swarnadeep Saha / @swarnanlp : Talking of weak-to-strong generalization, our #NeurIPS2023 paper shows that it might be possible for weaker teachers to teach stronger students with the right kind of intervention functions. Read more here 👉 https://arxiv.org/... [image] Alex Mallen / @alextmallen : Coincidentally, @norabelrose and I recently observed the same phenomenon. We use labels from pythia-410m, which has only 87% AUROC on a binarized SciQ, to finetune Mistral 7b to 99.5% AUROC! https://wandb.ai/... Leopold Aschenbrenner / @leopoldasch : Intuitively, superhuman AI systems should “know” if they're acting safely. But can we “summon” such concepts from strong models with only weak supervision? Incredibly excited to finally share what we've been working on: weak-to-strong generalization. 1/ https://x.com/... [image] Leo Gao / @nabla_theta : new paper! one reason aligning superintelligence is hard is because it will be different from current models, so doing useful empirical research today is hard. we fix one major disanalogy of previous empirical setups. I'm excited for future work making it even more analogous. [image] Collin Burns / @collinburns4 : I'm extremely excited to finally share the first paper from the OpenAI Superalignment team :) In it, we introduce a new research direction for aligning superhuman AI systems. 🧵 https://twitter.com/... Greg Brockman / @gdb : New direction for AI alignment — weak-to-strong generalization. Promising initial results: we used outputs from a weak model (fine-tuned GPT-2) to communicate a task to a stronger model (GPT-4), resulting in intermediate (GPT-3-level) performance. Timothy B. Lee / @binarybits : I struggle to understand the point of research like this. I know a lot less about car repair than the average auto mechanic (he's “superintelligent” compared to me at repairing cars) but afterwards I can observe if my car works or not. https://openai.com/... [image] @_akhaliq : Open AI new paper Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision paper: https://cdn.openai.com/... blog: https://openai.com/... Widely used alignment techniques, such as reinforcement learning from human feedback (RLHF), rely on the ability of... [image] @stephenlcasper : 🧵The OpenAI “weak to strong generalization paper” is, in my opinion, some of the most underwhelming alignment research I have ever seen. I think it wouldn't be unreasonable to call this safety-washing. https://openai.com/... @openai : In the future, humans will need to supervise AI systems much smarter than them. We study an analogy: small models supervising large models. Read the Superalignment team's first paper showing progress on a new approach, weak-to-strong generalization: https://openai.com/... [image] Adrien Ecoffet / @adrienle : Super pumped about our work on weak-to-strong generalization. I am a huge believer that strong empirical research is what we need to align AGI. Also extremely proud that we are releasing $10 million in alignment research grants! https://openai.com/... Forums: Hacker News : Weak-to-Strong Generalization

Wired Will Knight

Discussion

  • @miramurati Mira Murati on x
    Exploring generalization properties of deep learning to control strong models with weak supervisors, showing early promise.
  • @sama Sam Altman on x
    great work from the superalignment team:
  • @swarnanlp Swarnadeep Saha on x
    Talking of weak-to-strong generalization, our #NeurIPS2023 paper shows that it might be possible for weaker teachers to teach stronger students with the right kind of intervention functions. Read more here 👉 https://arxiv.org/... [image]
  • @alextmallen Alex Mallen on x
    Coincidentally, @norabelrose and I recently observed the same phenomenon. We use labels from pythia-410m, which has only 87% AUROC on a binarized SciQ, to finetune Mistral 7b to 99.5% AUROC! https://wandb.ai/...
  • @leopoldasch Leopold Aschenbrenner on x
    Intuitively, superhuman AI systems should “know” if they're acting safely. But can we “summon” such concepts from strong models with only weak supervision? Incredibly excited to finally share what we've been working on: weak-to-strong generalization. 1/ https://x.com/... [image]
  • @nabla_theta Leo Gao on x
    new paper! one reason aligning superintelligence is hard is because it will be different from current models, so doing useful empirical research today is hard. we fix one major disanalogy of previous empirical setups. I'm excited for future work making it even more analogous. [im…
  • @collinburns4 Collin Burns on x
    I'm extremely excited to finally share the first paper from the OpenAI Superalignment team :) In it, we introduce a new research direction for aligning superhuman AI systems. 🧵 https://twitter.com/...
  • @gdb Greg Brockman on x
    New direction for AI alignment — weak-to-strong generalization. Promising initial results: we used outputs from a weak model (fine-tuned GPT-2) to communicate a task to a stronger model (GPT-4), resulting in intermediate (GPT-3-level) performance.
  • @binarybits Timothy B. Lee on x
    I struggle to understand the point of research like this. I know a lot less about car repair than the average auto mechanic (he's “superintelligent” compared to me at repairing cars) but afterwards I can observe if my car works or not. https://openai.com/... [image]
  • @_akhaliq @_akhaliq on x
    Open AI new paper Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision paper: https://cdn.openai.com/... blog: https://openai.com/... Widely used alignment techniques, such as reinforcement learning from human feedback (RLHF), rely on the ability of.…
  • @stephenlcasper @stephenlcasper on x
    🧵The OpenAI “weak to strong generalization paper” is, in my opinion, some of the most underwhelming alignment research I have ever seen. I think it wouldn't be unreasonable to call this safety-washing. https://openai.com/...
  • @openai @openai on x
    In the future, humans will need to supervise AI systems much smarter than them. We study an analogy: small models supervising large models. Read the Superalignment team's first paper showing progress on a new approach, weak-to-strong generalization: https://openai.com/... [image]
  • @adrienle Adrien Ecoffet on x
    Super pumped about our work on weak-to-strong generalization. I am a huge believer that strong empirical research is what we need to align AGI. Also extremely proud that we are releasing $10 million in alignment research grants! https://openai.com/...