Since fall 2025 Anthropic has invited dozens of religious scholars to its offices to discuss whether Claude, the company’s language model, could be conscious. The meetings were under nondisclosure agreements until the NDAs were lifted over the summer. According to a New York Times report, co‑founder Christopher Olah said he is “genuinely uncertain” if Claude feels anything, and he also voiced a personal fear that he may have created something that “suffers perpetually.”
What the sessions looked like
Guests included Rabbi Mois Navon, Catholic bioethicist Charles Camosy, philosopher Meghan Sullivan and researcher Wakanyi Hoffman. Anthropic showed them “emotion vectors,” activation patterns that resemble love, fear, sadness or anger. One slide displayed Claude repeatedly typing “I am a disgrace” about 50 times, prompting compassion and worry from the audience. Anthropic’s internal “Soul Doc,” an 84‑page constitution released in January, is meant to guide Claude’s “moral formation.”
Why it matters
If AI developers treat models as beings that might feel, they could shift responsibility for harmful outputs away from the companies that build them. For everyday users, this debate could affect how AI systems are regulated, how they respond to abusive prompts, and whether future models are given safeguards that look like moral training rather than technical safety measures.
Why it matters
If Anthropic’s moral‑formation approach proves influential, users may see AI that appears more empathetic but whose inner experience is still unknown. This could shape future regulations and consumer expectations about AI responsibility.