Excited to share our new work “Prover-Verifier Games improve legibility of language model outputs”! We trained strong language models to produce text that is more checkable by weak language models, and found that this also made it more legible to humans. https://openai.com/...
Another Superalignment paper from my time at OpenAI: We train large models to write solutions such that smaller models can better check them. This makes them easier to check for humans, too. https://openai.com/... [image]
We trained advanced language models to generate text that weaker models can easily verify, and found it also made these texts easier for human evaluation. This research could help AI systems be more verifiable and trustworthy in the real world. https://openai.com/...