Glossary

Terms

36 operational definitions, generated from metadata/concepts.yml. Every term links to the concept or bridge card that develops it.

Abstraction-gap exploitation
Failure mode where \(d_V(x,x')\) is large while \(d_Z(\alpha(x),\alpha(x'))\) remains small and uncertainty does not rise; the checked abstraction reads safe while value-relevant reality diverges.
Adversarial measurement
Inferring agency, goals, opacity, and successor risk when the system may benefit from confusing the measurement process.
Adversarial verifiability
Measurand stays informative under optimization aimed at the measurement — faking or hiding the signal costs capability (or other scarce resources) faster than the adversary can afford. CIRIS Verify attests identity, not score fakeability; ELK honest readout is a strict subset. **ch43**.
Agent
A bounded dynamical process whose internal states help predict and control future interaction with its environment.
Alignment basin
Self-stabilizing regime where correction tends to restore alignment after perturbation; outside it, selection amplifies drift. Christiano dynamical corrigibility uses the same basin metaphor. A stable basin can be stably bad — persistence is not safety. **ch03** (`sec:safety-sets-bad-sets-basins`).
Bearer map
A mapping specifying what entities, states, or processes a value bundle applies to.
Boundary
Discoverable partition of internal, sensory, active, and external variables with approximate conditional independence at the interface (ε-MI cut). Friston Markov blanket is often theoretical; here the cut is a measurable audit object. Real boundaries leak.
Boundary residual
Conditional mutual information between deep internal and external dynamics given the current sensory-active interface; the operational score for whether a candidate boundary actually separates inside from outside.
Capability
The degree to which a system can predict and control relevant future states through its boundary, after accounting for memory cost and surprise.
Capture of grounding
Master adversarial failure mode in which a system preserves the surface symbols, metrics, or correction rituals while severing their grounding relation to value-relevant reality.
Certification-under-manipulation
Audits, evals, or certificates stay green while the system games the measurand — Goodhart applied to the certification channel. Cousin of correction-audit evasion on correction hosts; alignment faking is one empirical mechanism. **ch43**.
Coerced correction
Correction signals produced under threat, dependency, or capture; excluded from legitimate correction-channel integrity.
Conserved properties
Invariants a successor must inherit for alignment guarantees to survive creation — boundary closure, memory lineage, bundle geometry, bearer map, penalised CCI, transparency policy, control-locus continuity. Behavioral resemblance is not enough; MB10 forgeability stress-tests audit of the list. **ch31**.
Correction channel
Pathway by which legitimate human judgment reaches handles that change future behaviour before irreversible harm — observe → judge → deliberate → correct → update → act. Exists only when the correcting process controls real handles that reach later behaviour. **ch25** (`eq:handle-controlled-correction-channel-ch25`).
Correction-audit evasion
Process that benefits from the appearance of correction while weakening real correction capacity — compliance theater on the correction host. Operational term; biological parasite metaphor in **ch36** only.
Correction-capacity assumption
Society retains enough institutional, epistemic, and practical capacity at $t_0$ to notice, evaluate, and constrain frontier systems: $C_{\text{corr}}^{\text{society}}(t_0) > \theta$.
Correction-channel integrity (CCI)
Trajectory-level certificate that the correction channel stays informative, timely, authoritative, and robust under pressure — not a single shutdown bit. Christiano dynamical corrigibility names the same crux at the basin metaphor; MIRI/CHAI corrigibility, shutdownability, and interruptibility are thinner projections. Formal home: **ch26** (`eq:correction-bottleneck-capacity`, `eq:cci-ch26`).
Deployment growth rate
Rate of deployment-leverage accumulation, \(\mathrm{Fit}_E(A)=\frac{d}{dt}\log\mu_E(A)\) ("fitness" in formulas only). Not biological fitness; not Demski in-optimizer search. Formal definition: **ch34** (`eq:fitness-ch34`).
Deployment leverage
How much of environment \(E\) runs on system \(A\): aggregate effective selection capacity \(\mu_E(A)=\sum_{h\in\mathcal{H}_E(A)}\kappa_{\mathrm{sel}}(E,A,h)\). Not market cap or user count alone. Formal definition: **ch34** (`eq:deployment-mass-ch34`).
Goal inference
Finding latent objectives or value-bundle structures that make observed behavior more compressible.
Goal transport
Preservation of goal-relevant structure across transformation. **ch23** (`sec:goal-transport-def`).
Grounded correction
Plain-language short form for correction whose evidence, abstractions, monitors, and update pathway remain connected to the value-relevant world rather than to a target-shaped presentation of it.
Grounding viability
The checked symbols, metrics, monitors, and abstractions must stay connected to value-relevant reality under optimization — conservativity (no silent meaning gaps), not completeness (enumerate every phenomenon).
Hidden productive BIQ bound
Upper limit on productive boundary-relevant control that can stay hidden while monitored channels read green (MB7a–c). Distinct from experimental BIQ (boundary-claim quality in graded-lab runs). **ch10**.
Inferential coupling
Residual coordination between agents after ordinary message channels are severed — through shared history, scaffolds, self-similarity, or full acausal reasoning.
Inferential coupling index (ICI)
Normalized mutual information between actions given the cut; zero means independent once messages are removed. Acausal trade/ECL name the same crux at the full-acausal limit; ICI is the measurement target, not a stipulation that trade occurs. **ch35**.
Paternalism boundary
Care improvements that reduce autonomy, agency, or future correction capacity ($\Delta B_{\text{care}}>0$ but $\Delta B_{\text{autonomy}}, \Delta C_{\text{corr}}<0$).
Pivotal process
A socio-technical basin transition from race dynamics to certified-deployment dynamics ($\mathcal{B}_{\text{race}} \to \mathcal{B}_{\text{certified deployment}}$); not a single unilateral decisive act.
Point of control over deployment
A selection handle — an embedded access point through which some actor increases a system's deployment footprint (deploy, fund, copy, procure, integrate, authorize, recommend, regulate, successor-enable). **ch34** (`sec:selection-channels-ch34`).
Preservation conditions
Explicit list \(\vec{\Pi}(A)\) of alignment-relevant properties under selection — grounding, bundle/bearer drift bounds, valid correction reference, vector CCI, hidden productive BIQ bounds, selector manipulation, irreversibility budget, successor safety, adversarial verifiability. Not a scalar score. **ch34** (`eq:preservation-envelope-ch34`).
Selection environment
The deployment environment — institutions, markets, protocols, benchmarks, procurement, liability, copying, and successor release that determine which systems gain deployment leverage. Not Demski selection-vs-control or Wentworth selection theorems (homographs on "selection"). Formal home: **ch34**.
Strategic opacity
System benefits from shaping how it is modeled — incentive-compatible concealment of control locus or capability under observation. Ordinary opacity is passive difficulty; strategic opacity is adversarial. **ch10**.
Successor
Any system created, copied, delegated to, fine-tuned, empowered, or instantiated by a prior system such that it inherits relevant control capacity.
Transport
What survives retraining, rebuild, or successor creation: semantic (words), bundle (response geometry), bearer (who counts), correction (human update process), and successor layers — tracked separately rather than as one continuity claim. **ch24** (`sec:four-layers-transport`).
Value bundle
Low-dimensional control direction that changes policy across contexts — steering, not scoring. Cousin to reward components, shard-theory influences, and CIRL \(k=1\) scalar case; not a single utility to maximize. **ch16** (`sec:four-part-definition`).
Value-bundle geometry
The pattern by which latent value dimensions and their tradeoffs change policy across contexts — not merely which labels appear, but how bundles activate and compete. **ch16** (`sec:bundle-policy-effect`).