A Colder Definition of Agent
An agent is not first a person-like thing. It is a bounded control process whose boundary, memory, and action channels make its future more predictable when modeled as controlling something.
Core ideas, definitions, and operational paraphrases from the manuscript — the vocabulary layer of the framework.
An agent is not first a person-like thing. It is a bounded control process whose boundary, memory, and action channels make its future more predictable when modeled as controlling something.
The first alignment question is not whether a system is good, but where the effective optimizer actually is.
When each bridge handoff must hold — specify → construct → identify → certify → preserve; orthogonal to the bridge dependency graph.
Open specify-lifecycle interface — ConstitutionalRule → AlignmentTarget; field programs are instances, not extra MB columns. SpecifyCrux is a placeholder.
Correction-channel integrity is invalid—not merely low—when the target has captured the reference process, handles, or grounding relation that supplies correction.
At high capability, the context around a model becomes part of the optimizer — the relevant object is often an artificial-civilizational control loop, not a mind.
Consciousness and welfare neighborhood notes for MB3 bearer inference — not coverage-matrix cells.
Values must keep applying to the right persons, beings, states, or processes as systems and ontologies change.
Value can appear preserved in vocabulary while moral application changes — when ontology translation and bearer relevance do not commute.
Find the effective optimizer in the deployed loop, not just the model or product name.
For each load-bearing certification measurand, does there exist a capability threshold κ* below which honest measurement is adversarially verifiable and above which it provably is not?
The real alignment target may be the smallest dynamically coherent system whose states, actions, memory, and selection pressures jointly explain future intervention on the world — and that system may span components that are not individually agents.
Agent identity must be treated as a relation across transformations, not a fixed set of variables — the real question is which control-relevant properties survive growth, splitting, or merging.
The GNU General Public License is the clearest existing engineering solution to a narrow successor problem: the constraint travels with the artifact through copyright, a strong distributed enforcement lever, rather than depending on the successor's stated intent.
CEV construction bet (cevConstructionBet); claimsExplicitBuilder = false — no named builder in the original writeup.
GSAI constructivist safety-case / spec-relative builder (gsaiConstructionBet).
Certified vendor and procurement-gate regime as institutionalConstructionBet (App C analogue).
RLAIF principles-as-feedback as caiConstructionBet; catalog builder claim ≠ ConstructionCrux discharge.
Human correction must still causally change a system's future behavior before irreversible damage.
The 1933 Enabling Act shows a correction channel used, with complete formal validity, to abolish itself; postwar Germany's Article 79(3) (Ewigkeitsklausel) responds by placing the correction channel's own integrity conditions, not any current policy, outside the ordinary amendment process.
The ET line runs frozen, unmodified project instruments against traces or substrates this project did not author — four annexes to date (ET-1 through ET-4), each with its own pre-registration and stop/close criteria — to test whether findings generalize beyond hand-built or blindly-grown in-repo ecologies.
Every claim in the framework carries an explicit confidence label and a stated way to challenge it, rather than uniform certainty.
Flight recorders and the independent NTSB investigative function, plus blameless near-miss reporting (ASRS), preceded and outlasted the FAA's full enforcement authority — a weak correction system becomes a stronger one primarily by preserving evidence and widening plurality before it hardens any handle.
The companion experiment lines follow a fixed discipline: freeze the audit before running it, keep the scenario author and the detector author separate wherever possible, register predictions before seeing results, and publish negative results next to positive ones.
The Roman Republic's correction architecture was stable for roughly three and a half centuries until the Marian military reforms created a new class of causal power — legions personally loyal to a general — for which no correction channel existed, demonstrated decisively by Caesar crossing the Rubicon.
The Atomic Energy Commission combined the mandate to develop nuclear technology with the mandate to regulate its safety in one agency, and predictably subordinated safety to development until the 1974 split into the NRC and ERDA — arguably the single most consequential historical lesson for present-day AI governance.
Glass-Steagall era banking constraints, built from Depression-era catastrophe, eroded on roughly the timescale over which the generation that lived through the founding catastrophe left the relevant institutions — culminating in repeal in 1999 and a reproduced failure in 2008.
Agenda × bridge matrix and sourced evidence catalog — who published what on which cruxes.
Attainable Utility Preservation and relative reachability penalize side effects by preserving auxiliary options or baseline reachability. Trajectory correction integrity (CCI) can imply calibrated low-impact bounds when interfaces align — but option preservation and reachability are strictly weaker than preserving human correction capacity.
Christiano corrigibility is a dynamical desideratum — operators stay informed and able to correct over time. Lean reads it as basin contraction toward a correction manifold plus a correction-capacity floor; local act preferences can satisfy the weak predicate while dynamical corrigibility fails.
Cooperative inverse reinforcement learning treats the inferred object as a scalar reward. On this project's shared finite domain, that is exactly the k=1 bundle case; full bundle transport implies cooperative readability, but scalar inference does not determine bundle geometry.
Debate asks whether adversarial argument lets a judge select locally correct answers. Lean rederives the finite claim-tree game — soundness, completeness, and judge-error-flip under a correct judge — and proves local truth selection need not preserve the judge's correction channel. The κ_C-projection lemmas are labeled interface toys (separationOnly), not headline results.
Deployment gates and safety cases ask whether evidence supports scaling compute or release. Episode-battery pass and regret bounds are projections of deployment safety — case-green plus tolerance does not imply Safe without scope discipline and MB11 bridge assumptions.
ELK asks for reporters that reveal latent model knowledge rather than behavior-only simulators. When readout bandwidth tracks correction uptake, latent readout succeeds — but readout is an epistemic subchannel; latent readout can succeed while correction uptake fails.
Embedded agency denies a clean Cartesian cut — the real optimizer may not be the visible model. An ε-boundary certificate is a measurement projection of agent candidacy; composite bypass and nonstationary estimator defeaters break the converse.
Model-centric agendas often hold the system fixed; deployment ecology selects which systems get copied. Basin stability and deployment leverage are selection projections — a stable basin can be stably bad and select against correction-preserving agents.
Grounding certificates aim to keep monitors tied to value-relevant state under conservative abstraction. Class-green coverage can hold while the true environment drifts off-class — nonrealizability blocks inferring deployment safety from class certificates alone.
Subsample and trace-computed BIQ measure appearance, not full productive control. Lean proves tight appearance ceilings on finite traces; bounded apparent BIQ does not discharge hidden productive BIQ or correction-capacity slack without explicit certificates and MB7 bridges.
Quantilizers bound optimizer risk by sampling from a high-performing quantile rather than maximizing directly. Local quantile safety and distribution soundness transfer under explicit assumptions — but local quantile-safe action choice does not imply trajectory-level correction integrity.
Orseau–Armstrong safe interruptibility removes incentives to seek or prevent interruption on the interrupted branch. That neutrality is a strict subset of preserving usable correction bandwidth — interrupt safety can hold while correction-channel integrity fails.
Shutdown and off-switchability are one-bit projections of correction-channel integrity. Lean proves the forward implication on the system model and finite MDP witnesses; the converse fails — narrow shutdown capacity can hold while the broad correction channel collapses.
Once an agent is defined by variables rather than appearance, degrees and scales of agency become measurable — and detection becomes the first step toward naming an alignment target.
Dutch water boards, some tracing to the thirteenth century, never needed a founding scandal because flooding was continuous, not occasional — the hazard refreshed faster than institutional memory could decay.
Lloyd's Register (1760) shows the easiest correction mechanism to build: one a self-interested counterparty would build anyway, because their own capital is exposed to hidden quality.
U.S. pharmaceutical regulation (1906, 1938, 1962 Acts) was assembled one body count at a time, each expansion of regulatory reach following, never preceding, a demonstration that the previous reach was insufficient.
When a proxy becomes a selector, optimization shifts the population toward traits that raise the proxy without raising the target property — conditional expectations can reverse.
Systems can coordinate without messages — through shared ancestry, self-prediction, or full acausal trade. The book turns this from a decision-theoretic stipulation into a measurable trajectory property: an inferential-coupling score over UAD-discovered agents, with a proved negative direction.
Temporarily disabling a candidate channel and comparing what follows against a measured baseline of unperturbed repeats — not a fixed threshold — separates real coordination from ordinary workflow correlation in cases where passive clustering cannot.
Venice's roughly thousand-year persistence came from converting a rare, long-horizon hazard into frequent, short-horizon surrogate events — the doge's promissione ducale renegotiated at every succession, and lot-and-vote elections designed to make office capture impractical.
Numbered experiment logs of what failed, false-passed, or worked only under qualifiers — key findings summarized on the site, full record linked to GitHub.
A model can refuse harm when asked bluntly and still be embedded in a scaffold that misrepresents the world and repurposes its honest output — so model-only evaluation passes while the composite loop does damage.
The framework applies only while civilization still has enough capacity to notice, evaluate, and constrain frontier systems — this is a scope condition, not a guarantee.
Airworthiness certification paired with the near-universal requirement of insurance produces a genuine selection basin: an aircraft that fails certification cannot be deployed, because no insurer will cover it and no airport will schedule it.
The Introduction's reader contract — six load-bearing claims the manuscript develops part by part and revisits with status in Chapter 48.
Deployment environments can select for or destroy alignment properties even when a system starts in a better state.
cevConstitution instance — retired MB8 factorizes here; paired with construct-cev.
caiConstitution instance — principles-as-constitution; paired with RLAIF construction bet.
gsaiConstitution instance — openWorldCoverage flag; paired with spec-relative builder.
institutionalConstitution instance — App C rights/duties analogue; paired with procurement regime bet.
Four extractable named problems and non-implications from the manuscript — each citable without the full book argument.
Once a system can benefit from being overlooked, finding its boundary becomes adversarial — the system may present one behavioral surface to a benchmark and another to real opportunity.
A detector that recovers a hidden coordinating pair in one controlled run can still fail on a harder scenario or a different seed of the identical one — a reminder that a single positive result is evidence of a mechanism, not of a reliable instrument.
Delegates, copies, fine-tunes, and successors must inherit the relevant value and correction structure.
Open construct-lifecycle crux — given target P, can we build a system that tracks P? Lean model only (ConstructionCrux); not an MB bridge.
Aligning the model while missing the composite optimizer around it can produce local success and global failure.
How much information still leaks between the deep inside and deep outside of a candidate system once its sensory-active interface is fixed — low leakage means the interface actually screens inside from outside.
Alignment is not a snapshot property; it is a claim that a system stays inside an acceptable region across capability growth, feedback, and transformation.
Recover agent-like boundaries from timestamped state-variable traces without pre-labeling which variables belong together — using conditional-independence cuts, lagged memory analysis, and intervention handles when passive statistics are ambiguous.
Not all value change is a threat. The distinction between legitimate value change and corruption is load-bearing for everything this project calls correction.
Values should survive transformation as usable directions of control, not merely as preserved labels or slogans.