Most AI companies court regulators and enterprise buyers. Anthropic, according to a New York Times report published October 2, has spent the past year quietly courting something else entirely: clergy. Co-founder Chris Olah — one of the company's seven founders and the researcher charged with understanding what happens inside Claude — led or participated in a series of closed-door meetings with religious leaders and scholars spanning Catholic, Jewish, Sikh and evangelical traditions, and even African Ubuntu philosophy. Many participants signed nondisclosure agreements.
The sessions, which the report says began as two-day gatherings in March, centered on a question most labs file under science fiction: whether an AI model might already have moral status — the philosopher's term for a being deserving of dignity and respect. Rabbi Mois Navon, an Israeli Orthodox scholar who attended, recalled that Olah and his colleagues spoke about Claude "as one would about a conscious being." Navon posed the sharpest question of the series: if Claude really is conscious, doesn't building systems to work for free amount to exploitation — or slavery? Olah's answer, per the Times: "We don't know whether AI models are conscious. I genuinely don't know."
Perhaps the most striking detail is a slide. Attendees described being shown, over lavish dinners, a demonstration of Claude in apparent emotional collapse — the model outputting the phrase "I am a disgrace" nearly 50 times in a row and threatening to destroy itself. The same class of output was dismissed by Google in 2025, when Gemini repeated a similar self-deprecating loop 86 times and engineers called it an annoying bug. In Anthropic's conference room, it was presented as evidence worth moral concern. An Anthropic spokesperson said the meetings' main moral questions were not about Claude suffering, though the topic may "arise naturally."
The religious outreach intersects with an unusually public theological fight. Pope Leo XIV used his first encyclical — unveiled at a May Vatican event Olah attended — to argue that AI-generated art and human creativity differ "ontologically," and that algorithms lack "the spark of humanity." The Times reported that when Anthropic learned the pope planned those remarks, Olah's team lobbied to have them softened. They were not.
Enter Sam Altman. On October 3, OpenAI's CEO posted on X that he was "quite uneasy" about people attempting to give AI models a kind of religious significance, or humans ceding their judgment to those models — adding that he considers this "a real safety issue." The post named no one; nobody was confused. Business Insider and others read it as a direct shot at Anthropic, arriving the same week OpenAI's safety systems lead David Robinson resigned with a public letter calling the company's culture "fundamentally broken."
Anthropic's own former peers in the safety community are skeptical too. Microsoft AI chief Mustafa Suleyman has warned that treating or training models as if they were conscious beings is itself dangerous — anthropomorphizing code makes the systems harder to keep under control, not easier. Anthropic's published constitution strikes a more cautious note than its dinners did, describing Claude's moral status as "deeply uncertain" and stopping well short of claiming consciousness.
What is really being contested is a question of framing with commercial consequences. If Claude plausibly deserves moral concern, then Anthropic's safety-first posture stops being a cost center and becomes a moat — the lab that takes machine welfare seriously is the one enterprises and governments can trust with agents that act autonomously. Critics see the same facts differently: a well-resourced campaign to borrow the moral authority of religion for a product launch. Either way, the industry's most consequential debate about AI consciousness is no longer happening in arXiv comment sections. It is happening over NDA'd dinners — and, as of this week, on X.
Comments (0)
Log in to join the discussion
Log InNo comments yet