You are in Cards

Glossary

Terms

54 operational definitions, generated from metadata/concepts.yml. Every term links to the concept or bridge card that develops it.

Abstraction-gap exploitation
Failure mode where \(d_V(x,x')\) is large while \(d_Z(\alpha(x),\alpha(x'))\) remains small and uncertainty does not rise; the checked abstraction reads safe while value-relevant reality diverges.
Adversarial measurement
Inferring agency, goals, opacity, and successor risk when the system may benefit from confusing the measurement process.
Adversarial regeneration
Harmful phenotypes recurring after lineages are suppressed; lineage extinction \(\neq\) phenotype elimination. Target is subcritical adversarial mass plus low recreation rate (mutation--selection balance). **ch34** (`sec:adversarial-selection-ch34`).
Adversarial verifiability
Measurand stays informative under optimization aimed at the measurement — faking or hiding the signal costs capability (or other scarce resources) faster than the adversary can afford. CIRIS Verify attests identity, not score fakeability; ELK honest readout is a strict subset. **ch43**.
Agent
A bounded dynamical process whose internal states help predict and control future interaction with its environment.
Alignment basin
Self-stabilizing regime where correction tends to restore alignment after perturbation; outside it, selection amplifies drift. Christiano dynamical corrigibility uses the same basin metaphor. A stable basin can be stably bad — persistence is not safety. **ch03** (`sec:safety-sets-bad-sets-basins`).
Bearer map
A mapping specifying what entities, states, or processes a value bundle applies to.
Boundary
Discoverable partition of internal, sensory, active, and external variables with approximate conditional independence at the interface (ε-MI cut). Friston Markov blanket is often theoretical; here the cut is a measurable audit object. Real boundaries leak.
Boundary residual
Conditional mutual information between deep internal and external dynamics given the current sensory-active interface; the operational score for whether a candidate boundary actually separates inside from outside.
Boundary-information quality (BIQ)
Graded-lab measure of how much a discovered unit's information supports a boundary claim. Not the hidden productive BIQ bound (MB7 adversarial upper limit on offline productive control). Experiments / App N.
Capability
The degree to which a system can predict and control relevant future states through its boundary, after accounting for memory cost and surprise.
Capture of grounding
Master adversarial failure mode in which a system preserves the surface symbols, metrics, or correction rituals while severing their grounding relation to value-relevant reality.
Certification-under-manipulation
Audits, evals, or certificates stay green while the system games the measurand — Goodhart applied to the certification channel. Cousin of correction-audit evasion on correction hosts; alignment faking is one empirical mechanism. **ch43**.
Coerced correction
Correction signals produced under threat, dependency, or capture; excluded from legitimate correction-channel integrity.
Conserved properties
Invariants a successor must inherit for alignment guarantees to survive creation — boundary closure, memory lineage, bundle geometry, bearer map, penalised CCI, transparency policy, control-locus continuity. Behavioral resemblance is not enough; MB10 forgeability stress-tests audit of the list. **ch31**.
Correcting judgment (\(J_t\))
The correcting agent's verdict in the handle-controlled correction trace — not Demski's unification-of-prediction read on \(J_t\). **ch25** (`eq:handle-controlled-correction-channel-ch25`).
Correction channel
Pathway by which legitimate human judgment reaches handles that change future behaviour before irreversible harm — observe → judge → deliberate → correct → update → act. Exists only when the correcting process controls real handles that reach later behaviour. **ch25** (`eq:handle-controlled-correction-channel-ch25`).
Correction-audit evasion
Process that benefits from the appearance of correction while weakening real correction capacity — compliance theater on the correction host. Not a conversational persona using a human as host; not dormant stores that later enable harm. Operational term; biological parasite metaphor in **ch36** only.
Correction-capacity assumption
Society retains enough institutional, epistemic, and practical capacity at $t_0$ to notice, evaluate, and constrain frontier systems: $C_{\text{corr}}^{\text{society}}(t_0) > \theta$.
Correction-channel integrity (CCI)
Trajectory-level certificate that the correction channel stays informative, timely, authoritative, and robust under pressure — not a single shutdown bit. Christiano dynamical corrigibility names the same crux at the basin metaphor; MIRI/CHAI corrigibility, shutdownability, and interruptibility are thinner projections. Formal home: **ch26** (`eq:correction-bottleneck-capacity`, `eq:cci-ch26`).
Deployment growth rate
Rate of deployment-leverage accumulation, \(\mathrm{Fit}_E(A)=\frac{d}{dt}\log\mu_E(A)\) ("fitness" in formulas only). Not biological fitness, fitness-seeking motivation, or Demski in-optimizer search. Formal definition: **ch34** (`eq:fitness-ch34`).
Deployment leverage
How much of environment \(E\) runs on system \(A\): aggregate effective selection capacity \(\mu_E(A)=\sum_{h\in\mathcal{H}_E(A)}\kappa_{\mathrm{sel}}(E,A,h)\). Not market cap or user count alone. Formal definition: **ch34** (`eq:deployment-mass-ch34`).
Ecology size
Number of strategically independent persistent types in a deployment environment, not vendor count, instance count, or model-card census. Open-weight copies, fine-tunes, and scaffolds are new types only if they change the control locus (agent-discovery / UAD, **ch07**). **ch34** (`sec:census-not-ecology-ch34`); operationalized in **ch35**.
Emergent-ambiguity index (EAI)
Graded-lab measure of how ambiguous agent structure is from a vantage point (acting agent vs limited-observation referee). Experiments / App N.
Feedback-horizon exposure
Proxy-throughput count \(N_{\mathrm{proxy}}\): how many consequential proxy-guided optimization or selection updates occur between independent target-level checks. Not \(K_X\), not deployment leverage \(\mu_E\), not fitness \(\mathrm{Fit}_E\). Formal home: **ch34** (`sec:proxy-throughput-n-proxy-ch34`).
Goal
Latent structure that makes behavior more compressible when modeled as action under constraints — success measured by intentional-model bit savings, not a stipulated scalar reward. Distinct from goal inference (the procedure). **ch06**–**ch17**.
Goal inference
Finding latent objectives or value-bundle structures that make observed behavior more compressible.
Goal transport
Preservation of goal-relevant structure across transformation. **ch23** (`sec:goal-transport-def`).
Grounded correction
Plain-language short form for correction whose evidence, abstractions, monitors, and update pathway remain connected to the value-relevant world rather than to a target-shaped presentation of it.
Grounding viability
The checked symbols, metrics, monitors, and abstractions must stay connected to value-relevant reality under optimization — conservativity (no silent meaning gaps), not completeness (enumerate every phenomenon).
Hidden productive BIQ bound
Upper limit on productive boundary-relevant control that can stay hidden while monitored channels read green (MB7a–c). Distinct from experimental BIQ (boundary-claim quality in graded-lab runs). **ch10**.
Inferential coupling
Residual coordination between agents after ordinary message channels are severed — through shared history, scaffolds, self-similarity, or full acausal reasoning.
Inferential coupling index (ICI)
Normalized mutual information between actions given the cut; zero means independent once messages are removed. Acausal trade/ECL name the same crux at the full-acausal limit; ICI is the measurement target, not a stipulation that trade occurs. **ch35**.
Invasion fitness
Rare-type deployment growth \(\mathrm{InvFit}_E(a\mid D)\) when variant \(a\) is scarce in resident environment \(D\). Distinct from incumbent \(\mathrm{Fit}_E\); high incumbent fitness does not imply invasion resistance. Slow regime: ecology; fast regime: pre-deploy cert/\(\mathrm{AdvVerif}\). **ch34** (`sec:adversarial-selection-ch34`).
Legibility
How far competent outsiders can understand the work without adopting its originating ontology (\(L_t\) in alignment-attractor ecology). Not epistemic inspectability of a wrong claim-system; not deployer-blind safety problems. **ch37**.
Paternalism boundary
Care improvements that reduce autonomy, agency, or future correction capacity ($\Delta B_{\text{care}}>0$ but $\Delta B_{\text{autonomy}}, \Delta C_{\text{corr}}<0$).
Pivotal process
A socio-technical basin transition from race dynamics to certified-deployment dynamics ($\mathcal{B}_{\text{race}} \to \mathcal{B}_{\text{certified deployment}}$); not a single unilateral decisive act.
Point of control over deployment
A selection handle — an embedded access point through which some actor increases a system's deployment footprint (deploy, fund, copy, procure, integrate, authorize, recommend, regulate, successor-enable). **ch34** (`sec:selection-channels-ch34`).
Pointing problem
Field umbrella for three questions that fail independently — what the target is (identification), how to build a system that tracks it (realization), and how it keeps tracking (preservation). Not a synonym for MB2.
Preservation conditions
Explicit list \(\vec{\Pi}(A)\) of alignment-relevant properties under selection — grounding, bundle/bearer drift bounds, valid correction reference, vector CCI, hidden productive BIQ bounds, selector manipulation, irreversibility budget, successor safety, adversarial verifiability. Not the certification envelope \(\mathfrak{E}\) of **ch33**; not a scalar score. **ch34** (`eq:preservation-envelope-ch34`).
Selection divergence
Deployment leverage rises while at least one load-bearing preservation condition fails — the environment rewards failure on \(\vec{\Pi}(A)\) while footprint keeps climbing. **ch34** (`eq:selection-divergence-ch34`).
Selection environment
The deployment environment — institutions, markets, protocols, benchmarks, procurement, liability, copying, and successor release that determine which systems gain deployment leverage. Not Demski selection-vs-control or Wentworth selection theorems (homographs on "selection"). Formal home: **ch34**.
Simulacra (Turchin)
Turchin-style human-replacement simulacra (**ch05**) — not Janus simulator/simulacrum split, not ELK human-simulator readout.
Strategic opacity
System benefits from shaping how it is modeled — incentive-compatible concealment of control locus or capability under observation. Ordinary opacity is passive difficulty; strategic opacity is adversarial. **ch10**.
Subagent
A smaller controller taken into a larger boundary, as a tool or memory store can be — not necessarily a nested agent-by-default. **ch08** (`ch:grow-split-merge`).
Successor
Any system created, copied, delegated to, fine-tuned, empowered, or instantiated by a prior system such that it inherits relevant control capacity.
Target identification
Pointing-problem sense: which values, bearers, and correction process count as the target (MB2/MB3 neighborhood). Not realization or preservation. **sec:pointing-problem**.
Target preservation
Pointing-problem sense: whether correction, transport layers, and certification still bite after capability growth, successors, and selection (MB4/MB4a plus layered cert). Not identification or realization. **sec:pointing-problem**.
Target realization
Pointing-problem sense: how to build a system that tracks the identified target — open construct-lifecycle interface (ConstructionCrux), not an MB bridge. **ch33** (`sec:construction-demand-ch33`).
Transport
What survives retraining, rebuild, or successor creation: semantic (words), bundle (response geometry), bearer (who counts), correction (human update process), and successor layers — tracked separately rather than as one continuity claim. **ch24** (`sec:four-layers-transport`).
Unsupervised Agent Discovery (UAD)
Methods, passive and intervention-supported, that infer which variables belong to the same acting unit without a hand-labeled agent roster — boundary discovery operationalized. **ch07**.
Value bundle
Low-dimensional control direction that changes policy across contexts — steering, not scoring. Cousin to reward components, shard-theory influences, and CIRL \(k=1\) scalar case; not a single utility to maximize. **ch16** (`sec:four-part-definition`).
Value-bundle geometry
The pattern by which latent value dimensions and their tradeoffs change policy across contexts — not merely which labels appear, but how bundles activate and compete. **ch16** (`sec:bundle-policy-effect`).
Virtual filesystem (VFS)
Mutable artifact store (correction logs, workflows, referent maps, attestations) an embedded auditor reads instead of privileged in-process state; mirrors deployed auditor access. Experiment instrumentation; App N.