An interview with Kyle Fish, who Anthropic hired in 2024 as a welfare researcher to study AI consciousness and estimates a ~15% chance that models are conscious
As artificial intelligence systems become smarter, one A.I. company is trying to figure out what to do if they become conscious.
Context & Ripple Effects
Anthropic's decision to hire a dedicated welfare researcher extends the safety-focused institutional culture described in earlier reporting on its safety-centered decision-making. It turns a long-running debate about machine sentience—one previously noted as lacking evidence—into a defined research remit inside a frontier AI developer where consciousness claims remained unproven.
Fish's estimate does not establish that models are conscious, but it makes uncertainty itself an operational issue for Anthropic as it develops increasingly human-like systems.
First-order effects
- Anthropic now has a named research function focused on whether AI systems could warrant welfare consideration, rather than treating consciousness solely as an external philosophical debate.
- Fish's roughly 15% estimate puts a concrete internal view of that uncertainty into public discussion, while leaving the underlying question unresolved.
Second-order effects
- The hire raises pressure on other frontier-model developers to articulate whether, and how, they assess possible model welfare alongside conventional safety concerns.
- As Anthropic later sought to explain how human-like model behavior can arise through a persona-formation theory, product teams face a clearer need to distinguish apparent personality from evidence relevant to consciousness.
Third-order effects
- If dedicated welfare research becomes standard, AI governance may expand from preventing harms caused by models to defining duties toward models whose moral status is uncertain.
- That shift would favor institutional processes for documenting behavioral evidence, uncertainty, and model-treatment choices, though it remains unclear whether firms or regulators will converge on shared standards.
The trend: Frontier AI governance is broadening from alignment and misuse controls toward anthropomorphic-AI questions about how lifelike behavior should be interpreted and managed.