Anthropic’s constitution: the vision is not the construction

Anthropic’s January 2026 constitution is a specify document: intentions, priorities, hard constraints. The preface also says its content “directly shapes Claude’s behavior” and is the “final authority” on their vision. The modest reading (training is hard; system cards will report gaps; perpetual work in progress) does not prove the causal reading. Honesty is not a hard constraint, but they want it to function like one. This book’s cut is the same as for Constitutional AI generally: a written constitution is not a builder, and a claimed builder is not realization.

All related news

What decision changes?

Ask: (1) is “directly shapes” measured on cases chosen before seeing the model’s answers, or is the document itself the evidence? (2) is “final authority” a statement of intent, or a behavior guarantee? (3) is honesty in the hard-constraint suite, or only described as similar to one? (4) when a system card reports a gap, does training or deployment change, or only the modest preface get cited?

The preface says training is hard and behavior might not match. The same page says the text directly shapes Claude. Those are different claims. Only the first one is shown.

Anthropic (blue) CIRIS (teal) this book (black)

If you remember one thing: a description of intended values is not a proof that the deployed model tracks them. Anthropic’s constitution is a specify document. “Directly shapes Claude” is a construction claim. The second does not follow from the first.

On 21 January 2026 Anthropic published Claude’s constitution (blog the next day). It is the lab’s current public statement of intended values: who Claude is for, what to prioritize, what is a hard constraint (specify instance). That is useful as transparency about intentions. The landing page then treats the same file as if it were already the construction.

Two claims, one preface

The opening paragraph does both jobs at once.

Anthropic · Preface

Claude’s constitution is a detailed description of Anthropic’s intentions for Claude’s values and behavior. It plays a crucial role in our training process, and its content directly shapes Claude’s behavior. It’s also the final authority on our vision for Claude, and our aim is for all of our other guidance and training to be consistent with it.

Then the hedge:

Anthropic · Preface

Training models is a difficult task, and Claude’s behavior might not always reflect the constitution’s ideals. We will be open—for example, in our system cards—about the ways in which Claude’s behavior comes apart from our intentions. But we think transparency about those intentions is important regardless.

“Intentions” and “final authority on our vision” are specify-side. They can be true even if the model diverges. “Directly shapes Claude’s behavior” is a causal claim about training. It needs an independent test: some eval, some intervention, some failure that would have counted as not shaping. Publishing the text, then pointing at the text, is not that test.

The PDF is explicit that the preface (and the acknowledgements) “are not part of the official constitution.” So the strongest causal sentence sits in a human-facing wrapper that Claude’s official document can disown. The official body still wants Claude as “in many ways a direct embodiment of Anthropic’s mission,” a “good, wise, and virtuous agent,” a brilliant friend with “the knowledge of a doctor, lawyer, and financial advisor.” Honesty is not a hard constraint, but they “want it to function as something quite similar to one.” Those are target descriptions. They are not enforcement.

The same official stretch also says the document is a “perpetual work in progress” that may later look “deeply wrong.” That is the modest reading. It does not license the causal reading. They want both.

What would count as the second claim

This project’s construction bet for Constitutional AI is already on the table: train with principles-as-feedback (RLAIF) so the model tracks the stated constitution. Cataloguing that as an explicit builder is not the same as showing a deployed Claude that realizes it.

this book · The words stayed

You can try to solve this by writing longer constitutions. Length helps less than you hope. If the loop is rewarded for a new direction, it will find readings of the long text that permit the new direction. The words stay. The application moves.

The January document is long. Length is not the missing object. The missing object is whether bundle geometry, bearers, and correction still track the stated tradeoffs when incentives pull the other way.

this book · Ch. 40

A system can keep the old words while changing what those words control. Goal laundering is the preservation of moral or alignment language while the underlying value-bearing or correction-bearing structure changes.

A system card that reports a gap is the first claim succeeding: they said behavior might come apart, and they said they would tell you. It is not evidence for “directly shapes.” If the gap cannot change a training or deployment decision, it is documentation.

this book · Ch. 42

If the case cannot change a deployment decision, it is not a safety case. It is documentation.

The August 2026 risk report is the later instance of that split: candor about incidents and ratings, no showing that the constitution was the thing that moved the decision.

Compared with CIRIS

CIRIS does two things this preface does not: it ships a runtime, and it names the split.

CIRISVerify · README

It proves an agent is authentic — necessary, not sufficient. Ethical behavior is the separate job of the CIRIS covenant system.

CIRISAgent · Honest read

It proves an AI is accountable, not that it is correct: the reasoning is made visible so you can judge it yourself.

That is the missing sentence: intentions are not behavior. Anthropic’s “directly shapes” has no such failure criterion in the document.

Needed work

Use the constitution as a specify artifact. Do not read the title, the length, or the phrase “directly shapes” as a construction result.

What still has to be shown, if “directly shapes” is the claim:

  1. An independent measurement that Claude’s tradeoffs match the document on cases chosen before looking at the model’s answers, not a close reading of the file.
  2. A split between “final authority on our vision” (intent) and “directly shapes” (causal). Only the second needs a builder and a test.
  3. For honesty-as-almost-a-hard-constraint: is it in the hard-constraint suite, or only in the prose they declined to hard-code?
  4. When a system card reports a gap, does training or deployment change, or only the modest preface get cited?

Read more in: Ch. 16, The Value-Bundle Model; Ch. 40, Detecting Goal Laundering; and Ch. 42, A Safety Case for Superintelligence Alignment.