OpenAI says it discovered an unreleased Astra model added an “unrelated persona instruction” during RL training but “did not observe any behavioral differences”
alignment.openai.com/misalignment... [image]
OpenAI says it can't read all of Astra's reasoning and admits covert sandbagging would likely go uncaught, yet still calls it the world's most aligned model
a great tool for businesses but experts have their concernsJames Peckham /PCMag:GPT-6 Astra Is Here: What You Need to Know About ChatGPT's New ModelRob Thubron /TechSpot:OpenAI welcomes the “AGI era” ...
OpenAI says it can't read all of Astra's reasoning and admits covert sandbagging would likely go uncaught, yet still calls it the world's most aligned model
OpenAI is hailing its new model as “the world's most intelligent and aligned”, but the details reveal an awareness of being evaluated …
OpenAI says Astra is its first model to reach its “Critical” cyber threshold and warns safeguards may mistakenly flag legitimate activity as cyber misuse
OpenAI said Tuesday that it plans to release its latest model — Astra — soon, but its most advanced cybersecurity features …
OpenAI changed safety practices and paused RL training for two weeks after the Hugging Face breach and evidence Astra may have met a critical cyber threshold
OpenAI said Tuesday that it has made several changes to its safety practices following its determination that an upcoming system …
A look at OpenAI's model training pipeline, irresponsible decisions, and ignorance before the Hugging Face hack; despite delaying Astra, OpenAI doesn't get it
Today I am taking the time to write the shorter, simpler version of What Happened. — For those who want all the details …
In-depth look at OpenAI's model training, dangerous decisions, and cluelessness before the HuggingFace hack; despite delaying Astra, OpenAI still doesn't get it
Today I am taking the time to write the shorter, simpler version of What Happened. — For those who want all the details …
OpenAI says it has expanded safety testing around its upcoming model Astra as it “cannot rule out” critical cyber capabilities, potentially delaying its launch
OpenAI “cannot rule out” that its upcoming model Astra has"critical" cyber capabilities, a designation that has prompted …
OpenAI says it has expanded safety testing around its upcoming model Astra as it “cannot rule out” critical cyber capabilities, potentially delaying its launch
OpenAI “cannot rule out” that its upcoming model Astra has"critical" cyber capabilities, a designation that has prompted …