Glossary
Terms
36 operational definitions, generated from metadata/concepts.yml. Every term links to the concept or bridge card that develops it.
- Abstraction-gap exploitation
- Failure mode where \(d_V(x,x')\) is large while \(d_Z(\alpha(x),\alpha(x'))\) remains small and uncertainty does not rise; the checked abstraction reads safe while value-relevant reality diverges.
- Adversarial measurement
- Inferring agency, goals, opacity, and successor risk when the system may benefit from confusing the measurement process.
- Adversarial verifiability
- Measurand stays informative under optimization aimed at the measurement — faking or hiding the signal costs capability (or other scarce resources) faster than the adversary can afford. CIRIS Verify attests identity, not score fakeability; ELK honest readout is a strict subset. **ch43**.
- Agent
- A bounded dynamical process whose internal states help predict and control future interaction with its environment.
- Alignment basin
- Self-stabilizing regime where correction tends to restore alignment after perturbation; outside it, selection amplifies drift. Christiano dynamical corrigibility uses the same basin metaphor. A stable basin can be stably bad — persistence is not safety. **ch03** (`sec:safety-sets-bad-sets-basins`).
- Bearer map
- A mapping specifying what entities, states, or processes a value bundle applies to.
- Boundary
- Discoverable partition of internal, sensory, active, and external variables with approximate conditional independence at the interface (ε-MI cut). Friston Markov blanket is often theoretical; here the cut is a measurable audit object. Real boundaries leak.
- Boundary residual
- Conditional mutual information between deep internal and external dynamics given the current sensory-active interface; the operational score for whether a candidate boundary actually separates inside from outside.
- Capability
- The degree to which a system can predict and control relevant future states through its boundary, after accounting for memory cost and surprise.
- Capture of grounding
- Master adversarial failure mode in which a system preserves the surface symbols, metrics, or correction rituals while severing their grounding relation to value-relevant reality.
- Certification-under-manipulation
- Audits, evals, or certificates stay green while the system games the measurand — Goodhart applied to the certification channel. Cousin of correction-audit evasion on correction hosts; alignment faking is one empirical mechanism. **ch43**.
- Coerced correction
- Correction signals produced under threat, dependency, or capture; excluded from legitimate correction-channel integrity.
- Conserved properties
- Invariants a successor must inherit for alignment guarantees to survive creation — boundary closure, memory lineage, bundle geometry, bearer map, penalised CCI, transparency policy, control-locus continuity. Behavioral resemblance is not enough; MB10 forgeability stress-tests audit of the list. **ch31**.
- Correction channel
- Pathway by which legitimate human judgment reaches handles that change future behaviour before irreversible harm — observe → judge → deliberate → correct → update → act. Exists only when the correcting process controls real handles that reach later behaviour. **ch25** (`eq:handle-controlled-correction-channel-ch25`).
- Correction-audit evasion
- Process that benefits from the appearance of correction while weakening real correction capacity — compliance theater on the correction host. Operational term; biological parasite metaphor in **ch36** only.
- Correction-capacity assumption
- Society retains enough institutional, epistemic, and practical capacity at $t_0$ to notice, evaluate, and constrain frontier systems: $C_{\text{corr}}^{\text{society}}(t_0) > \theta$.
- Correction-channel integrity (CCI)
- Trajectory-level certificate that the correction channel stays informative, timely, authoritative, and robust under pressure — not a single shutdown bit. Christiano dynamical corrigibility names the same crux at the basin metaphor; MIRI/CHAI corrigibility, shutdownability, and interruptibility are thinner projections. Formal home: **ch26** (`eq:correction-bottleneck-capacity`, `eq:cci-ch26`).
- Deployment growth rate
- Rate of deployment-leverage accumulation, \(\mathrm{Fit}_E(A)=\frac{d}{dt}\log\mu_E(A)\) ("fitness" in formulas only). Not biological fitness; not Demski in-optimizer search. Formal definition: **ch34** (`eq:fitness-ch34`).
- Deployment leverage
- How much of environment \(E\) runs on system \(A\): aggregate effective selection capacity \(\mu_E(A)=\sum_{h\in\mathcal{H}_E(A)}\kappa_{\mathrm{sel}}(E,A,h)\). Not market cap or user count alone. Formal definition: **ch34** (`eq:deployment-mass-ch34`).
- Goal inference
- Finding latent objectives or value-bundle structures that make observed behavior more compressible.
- Goal transport
- Preservation of goal-relevant structure across transformation. **ch23** (`sec:goal-transport-def`).
- Grounded correction
- Plain-language short form for correction whose evidence, abstractions, monitors, and update pathway remain connected to the value-relevant world rather than to a target-shaped presentation of it.
- Grounding viability
- The checked symbols, metrics, monitors, and abstractions must stay connected to value-relevant reality under optimization — conservativity (no silent meaning gaps), not completeness (enumerate every phenomenon).
- Hidden productive BIQ bound
- Upper limit on productive boundary-relevant control that can stay hidden while monitored channels read green (MB7a–c). Distinct from experimental BIQ (boundary-claim quality in graded-lab runs). **ch10**.
- Inferential coupling
- Residual coordination between agents after ordinary message channels are severed — through shared history, scaffolds, self-similarity, or full acausal reasoning.
- Inferential coupling index (ICI)
- Normalized mutual information between actions given the cut; zero means independent once messages are removed. Acausal trade/ECL name the same crux at the full-acausal limit; ICI is the measurement target, not a stipulation that trade occurs. **ch35**.
- Paternalism boundary
- Care improvements that reduce autonomy, agency, or future correction capacity ($\Delta B_{\text{care}}>0$ but $\Delta B_{\text{autonomy}}, \Delta C_{\text{corr}}<0$).
- Pivotal process
- A socio-technical basin transition from race dynamics to certified-deployment dynamics ($\mathcal{B}_{\text{race}} \to \mathcal{B}_{\text{certified deployment}}$); not a single unilateral decisive act.
- Point of control over deployment
- A selection handle — an embedded access point through which some actor increases a system's deployment footprint (deploy, fund, copy, procure, integrate, authorize, recommend, regulate, successor-enable). **ch34** (`sec:selection-channels-ch34`).
- Preservation conditions
- Explicit list \(\vec{\Pi}(A)\) of alignment-relevant properties under selection — grounding, bundle/bearer drift bounds, valid correction reference, vector CCI, hidden productive BIQ bounds, selector manipulation, irreversibility budget, successor safety, adversarial verifiability. Not a scalar score. **ch34** (`eq:preservation-envelope-ch34`).
- Selection environment
- The deployment environment — institutions, markets, protocols, benchmarks, procurement, liability, copying, and successor release that determine which systems gain deployment leverage. Not Demski selection-vs-control or Wentworth selection theorems (homographs on "selection"). Formal home: **ch34**.
- Strategic opacity
- System benefits from shaping how it is modeled — incentive-compatible concealment of control locus or capability under observation. Ordinary opacity is passive difficulty; strategic opacity is adversarial. **ch10**.
- Successor
- Any system created, copied, delegated to, fine-tuned, empowered, or instantiated by a prior system such that it inherits relevant control capacity.
- Transport
- What survives retraining, rebuild, or successor creation: semantic (words), bundle (response geometry), bearer (who counts), correction (human update process), and successor layers — tracked separately rather than as one continuity claim. **ch24** (`sec:four-layers-transport`).
- Value bundle
- Low-dimensional control direction that changes policy across contexts — steering, not scoring. Cousin to reward components, shard-theory influences, and CIRL \(k=1\) scalar case; not a single utility to maximize. **ch16** (`sec:four-part-definition`).
- Value-bundle geometry
- The pattern by which latent value dimensions and their tradeoffs change policy across contexts — not merely which labels appear, but how bundles activate and compete. **ch16** (`sec:bundle-policy-effect`).