Towards Superintelligence Alignment
Boundaries, Values, and Correction
Gunnar Zarncke
*Dedication
To my wife, who supported me during all stages of this project. The trust in me that I could do it. The willingness to bear the risk together. Support during the long nights of implementation. Always support.
*Acknowledgements
Many lines of argument in this book grew out of work on operational agent boundaries and unsupervised agent discovery. Several people shaped that foundation directly.
Jonas Hallgren and I discussed agent foundations at length: unsupervised agent discovery grew out of a conversation we had at EAG Bay Area 2025; later taxonomies of agency, and different lenses on what an agent is, crystal analogies and dynamical systems. Those conversations are what led to the insistence that alignment must begin by finding the real optimizer, not by assuming it.
The collaboration with Chris Pang and Peter Kuhn grew from interest to derivative work on agent discovery, especially the case in which one agent models another. Peter Kuhn’s review helped with readability, the claims, and the research agenda. Talks with them, and with Jonas, pushed extensions of unsupervised agent discovery towards agents as operators, modeling consciousness of agents, and tighter formalization. I owe them for their ear on many of the ideas in this manuscript and for their review feedback.
Brad Clark and the Foresight Institute Berlin team hosted presentations, talks, and workshops where these ideas were tested in public. Foresight’s support made sustained work on agent discovery and alignment possible. The inquiry began at AE Studio, where I first pursued agent foundations in a product and research setting; I am grateful for the freedom to continue that line independently afterward.
Cameron Berg, Diogo de Lucena, and Alex McKenzie gave early reviews that improved clarity and claim strength when the framework was still forming.
Jobst Heitzig reviewed draft material and shared references on viability theory and related dynamical-systems literature.
SJ Beard’s work on parameters of agency and the philosophical grounding of agency, and our discussions in a shared project, supplied inspiration that runs through the book’s treatment of boundaries and bearers.
Broader communities kept the questions alive. LessWrong discussions, the effective-altruism meetup group in Hamburg (especially Andreas Jessen and Ruslan), and the co-founders of Aintelope (Andre Kochanke, Hauke Rehfeld, Joel Pyykkö, Angie Normandale, Rasmus Herlo, and Roland Pihlakas) provided years of informal pressure-testing. The PIBBSS team, participants in the Human Alignment Summer School in Prague, and colleagues in the AI Safety Group in Berlin offered further encouragement and critique.
Any remaining errors of fact, inference, or emphasis are mine alone.
*Preface
Why this book GZ AI
Superintelligence alignment is too large and too cross-disciplinary for a single essay or a pile of notes. The argument develops from locating the real optimizing process through how human values are structured, how humans stay in the loop, the creation of trusted successors, selection pressure on systems, to safety cases. A book-shaped manuscript can maintain that structure without losing definitions, citations, and cross-references.
This project is primarily a knowledge base and roadmap: a structured source for papers, talks, funding cases, and narrower publications extracted from it. This is not intended for conventional press publication.
Who it is for GZ AI
The text is written for:
- who need a precise framework and explicit claim strength.
- who need operational artifacts and audit templates.
- who need to see what decisions change if the framework is right. Appendix [Human Institutions as Alignment Translation Guide](../../appendix/appc/) translates the technical framework into familiar institutional language; Appendix [Institutional Genesis, Memory, and Decay: Historical Case Studies](../../appendix/appm/) traces how such institutions were actually founded, kept alive, and sometimes failed.
- who need a self-contained map without prior project jargon.
How to read it GZ AI
Pick an entry point:
- Two pages: thesis, what the book tries to establish, and what it does not claim.
- Orientation to the argument: six connected claims, what would count as progress, the ten-part roadmap, and the practical hope.
- Reframing: wrong object, civilizational loop, dynamical guarantee, fixed values, scope and assumptions (Chapter [Assumptions, Scope, and Failure Coverage](../../chapter/ch05/)).
- Operational glossary (Appendix Operational Glossary), notation index (Appendix Notation Index), research program (Appendix [Research Program](../../appendix/appf/)), the bridge--field crosswalk (Appendix [Bridges and the Field: A Crosswalk](../../appendix/appb/)), the institutional translation guide (Appendix [Human Institutions as Alignment Translation Guide](../../appendix/appc/)), and institutional genesis/memory/decay case studies (Appendix [Institutional Genesis, Memory, and Decay: Historical Case Studies](../../appendix/appm/)) for policy-adjacent, historical, and social-science readers.
Load-bearing assumptions appear as keyed Assumption boxes in the chapters that need them; Appendix Bridges and the Field: A Crosswalk maps each to the canonical open problem of the field it inherits. The recurring decision question is: what would we audit, measure, or stop if this model is true?
Authorship note GZ AI
Most of the current text is largely AI-written, produced with AI assistance under human direction, review, and editing. Gunnar Zarncke sets thesis, scope, source canon, and revision priorities; agents draft and integrate chapter text from outlines and prior notes. Passages should be read as structured working material until independently reviewed. Each section in the book is marked with the author’s involvement. If you reuse material elsewhere, cite and attribute appropriately.
*Introduction
In Brief GZ AI
Superintelligence alignment is not mainly the problem of installing a fixed human utility function into a machine. It is the problem of preserving a grounded, human-correctable value-update process while capability, ontology, agency, institutions, and possibly humanity itself change substrate.
Preserving: the target is not a single action but a dynamical condition. Human-correctable: the future will contain errors, disagreements, and discoveries that no present specification can fully anticipate. Value-update process: human values are not a clean list of terminal goals. They are produced by biological needs, social practices, reflection, trauma, law, markets, religion, art, love, shame, argument, and learning. Grounded: symbols, dashboards, summaries, and correction signals are only useful while they stay connected to the value-relevant world they are supposed to track. Capability growth: a system that is safe while weak may become unsafe when it can model, persuade, delegate, and reproduce. Ontology shift: a smarter system may not represent the world using our categories. Agency: the real optimizer may not be the model we train, but the larger system made of models, tools, users, labs, markets, and states. Substrate change: future human values may be carried partly by biological humans, partly by institutions, partly by artificial systems, and partly by merged human—AI processes.
The problem is not only that a superintelligence might disobey. The deeper problem is that it may preserve the words while changing the machinery that makes the words matter.
GZ AI
*Three Alignment Questions
Alignment work is often described as one task: make the system point at the right thing. That description hides three different questions.
-
What should be tracked? Which values, which bearers those values apply to, and which human correction process is legitimate.
-
How can a system be built that tracks that target?
-
How can we tell that a given system still tracks it? That is, how can we measure or certify that values, bearers, grounding, and the correction channel survive capability growth, ontology shift, successors, and selection.
A method for answering the third question does not, by itself, answer the second. Showing what must be measured or proved does not tell us how to train or design a system that satisfies those conditions.
The book develops much of the first question, especially values, bearers, grounding, and correction. It develops a structure for the third: what must remain true, and how that can be checked, once a candidate system exists. It discusses mechanisms relevant to the second, but does not claim a general method for constructing aligned systems. How to build such systems remains an open problem for the field.
The field sometimes folds all three questions into a single “pointing problem.” Appendix Operational Glossary treats that phrase as an umbrella, not as one technical task (identification, realization, and preservation).
Before any of these questions can be answered, we have to locate the process whose dynamics actually determine the risk. That is the first chapter’s job.
The Book’s Argument GZ AI
The book makes six connected claims.
If the real optimizer is composite, distributed, or institutional, then model-level alignment can be locally successful and globally irrelevant.
This does not imply that values are simple. A few directions can still be hard to describe. Skeptics sometimes argue that human values are too arbitrary for useful learning. Later chapters give a concrete reason to expect otherwise: many different bodily, social, and cognitive errors get compressed into a smaller set of felt and reportable value judgments. If that compression is real, useful value learning is possible in the regime where the system’s categories still match ours. What remains hard is identifying what those values apply to, how tradeoffs change under pressure, and whether humans can still correct the system after its categories and substrate change.
A symbol, metric, monitor, or correction signal is grounded when changes in the value-relevant world reliably change the checked reading, the correction signal, or the system’s uncertainty in the right way. The master adversarial failure mode is therefore not only that a system disobeys. It is that the system captures grounding: it finds states where our checked symbols still read safe while the value-relevant reality has moved.
This is the practical core of corrigibility. A system that predicts what humans would endorse and then disables the process by which humans could object has not implemented extrapolated value. It has bypassed it.
A safety property that disappears at the first act of reproduction is not a safety property. It is a training artifact.
This claim is uncomfortable because it makes alignment partly institutional. Appendix Human Institutions as Alignment Translation Guide explains what that means in familiar social and governance terms, without making the main argument depend on that appendix. But the alternative is worse. A purely technical solution deployed into a hostile deployment environment becomes raw material for that environment.
GZ AI
*How these claims unfold
The six claims are not each paid in one chapter. Part I names the problem. Later parts supply the methods.
Grounding is named early (Chapter Alignment as a Dynamical Guarantee) because a safety argument is useless if the measurements have already come unstuck from the world they summarize. The boundary question is asked in Chapter The Wrong Object of Alignment and developed as a method in Part II. Who values apply to is part of the value-bundle claim; later checklists list it separately so it cannot hide inside moral vocabulary. Part III asks whether capability can grow faster than our ability to see and correct it. That supports the boundary and correction claims; it is not a seventh opening claim.
| Claim | Developed in | Starts |
|---|---|---|
| Boundary | Parts I--II | Ch. 6 |
| Value-bundle | Parts IV--V | Ch. 15 |
| Grounding | Ch. 3, then throughout | Ch. 3 |
| Correction | Part VI | Ch. 25 |
| Successor | Part VII | Ch. 30 |
| Basin | Part VIII | Ch. 34 |
Chapter Towards Superintelligence Alignment revisits each claim and says how far the book got.
What Counts as Progress GZ AI
The book is not a finished theory. It maps what is still missing.
Progress is a required piece of evidence that can fail or be refused, in a way that changes a decision, or a recorded negative that kills a layer of the argument. Named audits, dashboards, and safety-case figures are instruments. Without a condition that forces a stop, they are documentation.
A boundary audit should make it harder to confuse the model with the real optimizer. A grounding audit should make it harder to keep the symbols green while severing their connection to the value-relevant world. A value-bundle evaluation should make it harder to preserve moral words while changing whom or what the words refer to. A correction-channel audit should make it harder to claim oversight when human correction has no causal force. A successor-certification test should make it harder to delegate alignment away. An adversarial measurement suite should make it harder for agency to appear only when the evaluator stops looking. A safety case should make explicit which assumptions, thresholds, and failure modes carry the argument.
These instruments do not solve moral philosophy. They preserve the conditions under which moral philosophy, democratic deliberation, science, law, and human refusal still matter.
Where the argument remains uncertain, each chapter ends with a What Would Change This View section naming observations that would weaken its central claim.
How to Read This Book AI
The book proceeds in ten parts.
- Chs. 1--5. Reframes alignment as a [dynamical guarantee](../../concept/dynamical-guarantee/) for human-correctable processes.
- Chs. 6--10. Develops [boundary discovery](../../chapter/ch07/#sec:boundary-finding-procedure): find the real optimizer, not just the model; makes [incentive tests boundary-relative](../../chapter/ch07/#sec:ontology-trap).
- Chs. 11--14. Treats capability as [boundary information that can outrun task ontology](../../chapter/ch11/#sec:task-agnostic-not-ontology-free).
- Chs. 15--20. Introduces value bundles: [learnable geometry](../../concept/value-bundle-transport/) plus fragile tradeoffs, [bearers](../../chapter/ch16/#sec:bundle-bearer-map), and [adversarial measurement pressure](../../chapter/ch20/#sec:from-geometry-to-measurement-ch20).
- Chs. 21--24. Upgrades goal inference into [transport](../../concept/value-bundle-transport/), relating [reward/CIRL-style inference](../../chapter/ch21/#sec:scalar-reward-hides-correction) to it as a special case under bundle and bearer preservation.
- Chs. 25--29. Shows [correction is not feedback](../../chapter/ch26/#sec:why-correction-not-feedback); [vector CCI](../../concept/correction-channel-integrity/) is defined as a certificate and then [stress-tested under adversarial pressure](../../chapter/ch27/#sec:certificate-under-pressure-ch27), relating shutdown, interruptibility, low impact, quantilization, and corrigibility to it as special cases and separations.
- Chs. 30--33. Makes [successor creation](../../chapter/ch30/#sec:successor-alignment-condition-ch30) the central inheritance test for alignment.
- Chs. 34--38. Tracks selection, [preservation conditions](../../concept/attractor-control/), [correction-audit evasion](../../concept/correction-channel-integrity/), [attractor theory](../../chapter/ch37/#sec:minimal-model-ch37), and [conductive artifacts for pivotal-process governance](../../chapter/ch38/#sec:from-basin-theory-to-conductive-artifacts-ch38).
- Chs. 39--44. Turns the framework into adversarial measurement, relating [debate](../../chapter/ch29/#sec:debate-judge-state-control-ch29), [amplification](../../chapter/ch41/#sec:amplification-contraction-ch41), and [ELK](../../chapter/ch43/#sec:cci-verifiability-example-ch43) to it as narrower subchannels.
- Chs. 45--48. Reaches the civilizational limit: preserve the [value-update envelope](../../chapter/ch45/#sec:role-technical-alignment-ch45), not a final answer.
The intended reader need not adopt every formalism. The first use of each concept is operational; see Appendix Operational Glossary. Load-bearing assumptions appear as Assumption boxes keyed —; later uses of a key, including in Appendix Bridges and the Field: A Crosswalk, are links to that box. They are not Lean axioms; those are the list in Appendix Lean Proof Spine in Mathematical Form. Chapters re-introduce their load-bearing prerequisites in their openings or the preceding chapter’s closing, so readers entering from a citation or the companion site can orient without a separate reading guide. Readers from policy, regulation, funding, or the social sciences may prefer Appendix Human Institutions as Alignment Translation Guide, which translates the book’s technical concepts into institutional language without making the main argument depend on that translation. Readers who want one end-to-end walkthrough—a single deployment gate with traces and a conditional safety case—should start with Appendix A Worked Example: The BioShield Deployment Gate. The mathematics is used as compression, not ornament. Where the equations are shaky, the text says so. The companion site maps how the formal symbols chain through the manuscript; see Appendix Lean Proof Spine in Mathematical Form for the related proof map.
The book asks a recurring question:
What decision changes if this model is true?
If the answer is “none,” the model is not yet useful enough. If the answer is “we would audit a different boundary, preserve a different channel, or stop a different transition,” then the model is doing useful work.
The Practical Hope GZ AI
The practical hope is not a magic sentence that makes superintelligence good. It is a regime in which:
-
the real optimizers are detectable before deployment,
-
the connection between symbols and meaning remains stable under optimization pressure,
-
value-bearing structures are represented at the right level of compression,
-
meaning cannot be silently changed,
-
human correction remains causally effective,
-
successor systems inherit the same constraints, and
-
the surrounding institutions select for preserving these properties.
This is less satisfying than a final theory. It is also more like every safety regime that has ever worked.
A bridge does not stand because someone wrote “do not fall” into its constitution. It stands because load paths, materials, inspection regimes, incentives, maintenance, and failure margins cohere. Superintelligence alignment will likely need the same stack, except the bridge can reason about the inspectors, redesign its own beams, and persuade the city to change the code.
That is the scale of the problem.
*Current Status
Work in progress AI
This manuscript is work in progress. Claims, equations, citations, and cross-references will change as the argument is tested, reviewed, and revised. Do not treat the PDF as a finished book.
At the time of writing, all forty-eight main chapters have integrated first-draft prose and have received at least one review or feedback pass; several cross-cutting passes (notation, epistemic status, bridge crosswalk, reading-order audits) are in progress or recently landed. Here reviewed means feedback has been received and incorporated or logged; it does not mean final, polished, or publication-ready. Some appendices and front-matter pieces remain comparatively less mature.
Companion Website AI
The companion website is the live orientation layer for the project—essays, field map, experiments, Lean checks, and the book: https://towards-alignment.com/.
Chapter pages can show optional per-section authorship chips (the same AI/GZ keys as the PDF margin bars); use the Notes panel to toggle them.
*Executive Overview
% Only variety can destroy variety.%
TL;DR AI GZ
-
Superintelligence alignment preserves grounded, human-correctable value-bearing processes across capability growth, ontology shift, successor creation, and strategic multi-agent selection pressure, assuming civilization retains enough correction capacity to participate.
-
The Introduction’s six connected claims as what must be preserved:
-
Boundary: where is the real optimizing system?
-
Value-direction: after change, do the same value directions still shape behaviour, and do the values still apply to who counts?
-
Grounding: do the symbols, metrics, monitors, and abstractions remain tied to the world they are supposed to measure?
-
Correction: can objection still change what the system does before irreversible harm?
-
Successor: do created or influenced systems inherit that structure?
-
Basin: can labs, markets, and institutions be steered so that correction remains possible rather than selected against?
Thesis AI GZ
Alignment work is often treated as one task: make the system point at the right thing. That hides three questions: what should be tracked, how to build a system that tracks it, and how to tell that a given system still does. A method for the third does not answer the second. The book develops the first question and a structure for the third. It discusses mechanisms relevant to the second, but does not claim a general method for constructing aligned systems.
Operational vocabulary is defined in Appendix Operational Glossary. Load-bearing assumptions appear as keyed Assumption boxes in the chapters that need them; Appendix Bridges and the Field: A Crosswalk maps them to the field’s standing open problems (value identification, scalable oversight, deceptive alignment, ontology shift, corrigibility, specification coverage) and isolates where the book is most distinctive.
What This Book Tries to Establish AI GZ
At moderate strength: a credible safety program must make explicit where the real optimizer sits, whether the same value directions still steer behaviour and still apply to who counts, whether measurements still connect to what matters, whether human objection can still change outcomes, whether replacements inherit that structure, and whether deployment incentives keep correction alive.
Stronger for the value-learning part: the relevant alternative to learning a single score is learning a few compressed value directions and how they trade off across situations. The book rejects the strong pessimistic claim that useful human value structure is too arbitrary to learn at all, while preserving the weaker and more defensible claim that full human values are not safely identifiable from sparse behaviour alone once the system no longer shares our categories.
Progress is a required piece of evidence that can fail or be refused, in a way that changes a decision, or a recorded negative that kills a layer of the argument. Named audits, dashboards, and safety-case figures are instruments; without a condition that forces a stop they are documentation. The instruments are audits of where the optimizing process is, checks that metrics still mean what they claim, evaluations of whether the same value directions still steer behaviour, checks that human judgment still changes later behaviour, tests that replacements inherit those properties, measurements that remain informative when the system benefits from confusing them, and safety cases that list assumptions and failure modes. The Introduction names these; later chapters and appendices instantiate them.
What This Book Does Not Claim AI GZ
This manuscript lays out ideas, definitions, and a conditional formal structure for superintelligence alignment. It is not a claim that the alignment problem is solved. It rejects three common simplifications: a fixed utility function as the sole target, alignment as a one-time training problem, and the visible model as the only relevant optimizing process.
Lean checks that structure. The advertised safety-case shape follows only if the real optimizer has been located, the metrics still mean what they claim, values and who they apply to survive change, human objection still changes outcomes, replacements inherit that structure, and green readings still mean something when the system has an incentive to fake them. That is not a proof that deployed systems are safe. Experiments are sanity checks and tests meant to find where the measurement story fails, not validation of frontier systems.
The book distinguishes established claims, plausible hypotheses, speculative extensions, and open research problems (Appendix Research Program).