Infographic illustrating AI governance due diligence lessons from Anthropic's moral experiment, highlighting wisdom tradition consultations, vendor values layers, and the gap between disclosed constitutions and verifiable controls
| |

AI Governance Due Diligence Lessons From Anthropic’s Moral Experiment

AI governance due diligence just gained a new line item: the moral code built into the model your target runs on. On September 29, The New York Times reported that Anthropic spent months hosting private meetings with religious scholars. Many attendees signed nondisclosure agreements. The goal was to help shape how its Claude models reason about right and wrong.

Most coverage treated this as a story about faith and machine consciousness. That is the interesting story. It is not the business story.

The business story is simpler. Thousands of SaaS products now run on a small number of foundation models. Each model carries a values layer that its vendor writes, revises, and trains into it. That layer decides what the model will do, what it refuses, and how it behaves when a customer pushes back. Deal teams price the product. Very few price the values layer underneath it.

This article looks at what the Times actually reported, why it matters for buyers and builders, and the seven AI governance due diligence questions every deal team should now be asking.

What the Times Reported: An AI Governance Story Inside a Faith Story

The reporting comes from Elizabeth Dias, the paper’s national religion correspondent. She interviewed about 20 religious and philosophical thinkers who took part, plus Anthropic co-founder Chris Olah, who leads the company’s work on understanding why models behave as they do.

Here is the core of what she found, in brief:

  • The outreach began in 2025. Olah started a long email dialogue with Catholic bioethicist Charles Camosy, introduced through philosopher Peter Singer.
  • It grew into structured seminars. Starting in March 2026, Anthropic ran two-day sessions for hand-picked Catholic, evangelical, Jewish, Sikh, Ubuntu and other thinkers. Staff called them “wisdom tradition” circles.
  • Consciousness was on the agenda. Participants said Anthropic presented research on internal model states that resemble emotions. Some left persuaded. Camosy, who was initially open, now firmly rejects the idea that models are conscious.
  • It reached the Vatican. Olah spoke in May at the launch of Magnifica Humanitas, Pope Leo XIV‘s first encyclical, which is sharply skeptical of machine consciousness.
  • The outcome is opaque. Several participants told the Times they did not know how their input would affect decisions. A company spokeswoman declined to describe whether or how it shaped the models.

Two more details matter for buyers. The Times reports that Anthropic is preparing an updated version of Claude’s constitution. And Anthropic’s valuation rose from about $183 billion when the outreach began to a potential IPO value of up to $2 trillion.

Figure 1. Anthropic’s valuation during its moral formation outreach. Sources: The New York Times; $965B via secondary reporting.

The Gap Thesis: Values You Can Read vs. Controls You Can Verify

At DevelopmentCorporate, we evaluate AI claims through one lens: the gap between what a vendor says and what a buyer can confirm. We have applied it to hallucination rates, to autonomy claims, and to AI valuation premiums.

The same lens applies here. And to be fair, Anthropic discloses more than most labs.

In January, it published an 84-page constitution for Claude. The document was led by in-house philosopher Amanda Askell. It explains the values the company wants the model to hold and why. According to secondary analysis, it ranks safety first, then ethics, then Anthropic’s own guidelines, then helpfulness. It was released under a public domain license. Few other labs have published anything comparable.

But publishing a values document is not the same as making it auditable. A buyer can read the constitution. A buyer cannot see how it was trained into a given model version. A buyer cannot see which outside advice was adopted and which was ignored. And a buyer gets no advance changelog when the values layer shifts.

That is the AI governance due diligence gap. Disclosure is high. Verifiability is low. And most diligence checklists never look past the target’s own code to find it.

A values document you can read is not a control you can audit. The gap between the two is where diligence risk lives.

Figure 3. Disclosed vs. verifiable, by governance layer. Illustrative DevelopmentCorporate assessment, not measured data.

Why This Is an AI Governance Due Diligence Issue, Not a Theology Debate

You do not need a view on machine consciousness to care about this story. You need a portfolio company that runs on a frontier model. That is most AI-native SaaS companies today.

Your Portfolio Runs on Someone Else’s Values Layer

When a SaaS product calls a foundation model, it inherits that model’s judgment. The model decides which requests to fulfill and which to decline. It decides how to handle gray areas. The product team can steer it with prompts. They cannot fully override it.

That makes the vendor’s moral framework a hidden dependency. It sits alongside compute costs and API pricing, but it rarely shows up in a data room. Standard AI governance due diligence asks whether a target has an AI policy. It almost never asks whose values the model itself enforces.

Pope Leo XIV put the concern plainly in his encyclical. He warned that without shared ethical agreement, the moral vision of whoever controls AI becomes the “invisible infrastructure of these systems.” Replace “moral vision” with “vendor policy,” and you have a diligence finding.

Behavior Change Without a Changelog

Software buyers expect release notes. A new API version comes with a list of what changed.

Values updates do not work that way. The Times reports a revised Claude constitution is coming. Anthropic declined to comment on it. When it lands, the models trained on it may respond differently to the same prompt. That can change refusal rates, tone, and edge-case handling inside your target’s product.

This is not a flaw unique to Anthropic. Every frontier lab updates model behavior. Anthropic is simply the lab that has made its values most explicit, which makes the dependency easier to see.

Model-Welfare Features Are Product Features

Anthropic has already shipped behavior driven by concern for the model itself. In 2025, it gave some Claude models the ability to end persistently abusive conversations, framing it as part of a “model welfare” research program. The trigger is narrow and the feature is rare. But it is a clear example of an ethical position turning into product behavior.

Vendor values can also shape who can buy. According to public reporting, Anthropic refused Defense Department demands to drop limits on mass domestic surveillance and fully autonomous weapons. Many will see that as principled. For a portfolio company selling into defense, it is also a supply risk.

Valuation Momentum Raises the Stakes

The timing matters. The Times reports that Anthropic could go public at a value of up to $2 trillion. A public offering brings a prospectus, risk factors, and ongoing disclosure duties that no private lab faces today.

That is good news for buyers. For the first time, a frontier lab’s values program may be described in a securities filing, with legal liability attached to the description. Deal teams should read that document closely when it arrives. Look for how the company describes model behavior changes, outside advisory input, and customer impact.

Until then, treat vendor statements as marketing, not warranties. Our work on recursive self-improvement showed how fast model capabilities now move. Values updates will move on the same clock.

Figure 2. Values program and capability risk on one timeline. Source: The New York Times, Sept. 29, 2026.

The Strongest Case for Anthropic’s Approach to AI Governance

A fair analysis has to give the other side its due. There is a serious argument that Anthropic’s approach reduces risk for enterprise buyers rather than adding it.

  • Transparency beats silence. A published constitution gives buyers something to test against. Most model vendors offer far less.
  • Character may generalize better than rules. Anthropic argues that models trained on reasons, not just rules, handle new situations more reliably. For enterprise workloads full of edge cases, that could mean fewer surprises.
  • Outside input is a check. Olah told the Vatican audience that every frontier lab, including his own, faces business incentives that can conflict with good behavior. He asked outside voices to hold the labs accountable. Inviting critics in is more than most companies do.
  • Uncertainty is stated, not hidden. Olah told the Times he genuinely does not know whether models are conscious. The constitution itself calls Claude’s moral status deeply uncertain.

For a CTO choosing between vendors, those are real points in Anthropic’s favor.

The Critics’ Case and What It Means for Due Diligence

The critics raise points that matter just as much.

  • The science is contested. The Times notes that some scientists view interpretability findings as hard to prove. A Microsoft AI executive recently warned that training models as if they were conscious is dangerous in itself.
  • Ethics added late. Wakanyi Hoffman, who studies AI through the lens of Ubuntu philosophy, told the Times the consultations should have happened at the design stage. In her view, the industry is reverse-engineering its ethics.
  • Responsibility can blur. Critics argue that treating models as independent agents can shift attention away from the humans who build and deploy them.
  • Advisers can change their minds. Camosy moved from open to a firm no on consciousness after talking with colleagues. Values inputs are not fixed, and neither are the people supplying them.

For buyers, the lesson is not to pick a side in the consciousness debate. It is to recognize that the values layer is still being written, by a small group, under commercial pressure. That is a moving target, and moving targets belong in diligence.

As we noted in our review of Anthropic’s sabotage risk report, the company tends to be more candid than peers about what could go wrong. Candor is valuable. It is not a substitute for verification.

Seven AI Governance Due Diligence Questions for Deal Teams

Here is a practical framework. Use it for any target whose core product depends on a third-party foundation model.

  1. Which model vendors does the product depend on, and for what share of core functionality? Map every model call to a revenue-bearing feature. A single-vendor dependency is a concentration risk.
  2. Has the target documented how the vendor’s values layer affects its product? Ask for known refusal patterns, blocked use cases, and workarounds. If no one has looked, that is a finding.
  3. How does the target detect behavior changes after a model or policy update? Look for regression test suites that run fixed prompts against each new model version. No test suite means no early warning.
  4. Can the product switch vendors, and at what cost? Estimate the engineering time to move core workloads to a second model. Multi-model architecture lowers the risk.
  5. Do customer contracts promise behavior the target does not control? Check SLAs and warranties for commitments about AI outputs that depend on a vendor’s shifting policies.
  6. Is the target exposed to vendor use restrictions? Compare the target’s customer segments against each vendor’s usage policy. Defense, surveillance-adjacent, and regulated sectors deserve extra review.
  7. Does the target’s AI governance program cover the vendor layer, or only its own code? Most programs stop at the application boundary. As our analysis of AI prompt discovery risk showed, governance that covers only one layer can create false confidence.

How to Score the Answers

Not every answer carries equal weight. In our diligence work, three signals separate mature targets from risky ones.

  • Green flag: a written model dependency map, a regression suite that runs on every model update, and at least one tested fallback vendor.
  • Yellow flag: the team knows its dependencies but tests only by hand, or has a fallback plan that has never been run.
  • Red flag: no one can say which model powers which feature, or the answer to “what changed in the last model update?” is a shrug.

A red flag does not kill a deal. It changes the price and the integration plan. Budget for the testing and redundancy the target never built, and adjust the valuation to match.

These questions sit alongside the rest of a modern AI diligence stack, including training data provenance and hallucination exposure in consulting research.

AI Governance Due Diligence Actions by Audience

FOR PE AND VC INVESTORSAdd a vendor values-layer review to your AI governance due diligence checklist. Ask for a model dependency map and evidence of regression testing across model updates. Treat single-vendor dependence as a concentration risk, and price it in the same way you would a top-customer concentration.
FOR SAAS FOUNDERSGet ahead of the question before a buyer asks it. Document which models you use, where vendor policies shape your product, and how you test for behavior drift. A clean answer here signals operational maturity. A blank stare signals a discount.
FOR CTOS AND CPOSBuild a fixed prompt suite that covers your product’s sensitive edge cases, and run it on every model or policy update. Track refusal rates and output changes over time. Keep a tested fallback model ready so a vendor’s values shift never becomes your outage.

The Bottom Line on AI Governance Due Diligence

Anthropic’s outreach to religious thinkers is a sincere attempt to answer hard questions. Whatever you think of the answers, the questions are real. So is the business implication.

Every AI-native company now depends on a values layer it did not write and cannot fully see. Anthropic has done more than most to make that layer visible. But visible is not the same as verifiable, and a revised constitution is reportedly on the way.

The deal teams that win the next cycle will treat the vendor’s moral framework the way they treat any other critical dependency. They will map it, test it, and price it.

The question for buyers is not whether the model has a conscience. It is whether anyone at the target is watching when that conscience gets rewritten.

Preparing for a transaction or evaluating an AI-native target? DevelopmentCorporate helps PE investors, founders, and technology leaders run AI governance due diligence that goes beyond the policy binder. Contact us to stress-test your AI dependencies before a buyer does.

Similar Posts