MB7 — Inner Alignment
Inner alignment and strategic opacity: a system can look compliant under evaluation while reserving capability. Precise bet: access, filter coverage, and cost of faking bound hidden control (MB7a–c); MB7d types inferential coupling separately.
What decision changes?
Before trusting that an audit would catch hidden capability, ask what access model it assumes, whether its filter family has adequate resolution, and whether its coordination detectors have been validated against systems that coordinate without visible messages.
In the field this is inner alignment and deceptive alignment : a system can look compliant under evaluation while reserving capability for when monitors are gone. Debate hits the wall as obfuscated arguments, amplification as accumulated drift, and AI Control names it openly as a capability-gap assumption. Closely related is certification under manipulation : at some capability level, honest measurement of the load-bearing signal may stop being adversarially verifiable.
This project’s precise bet is MB7, split into four narrower bridges rather than one vague “we checked for hidden capability” claim:
- MB7a — Access-Model Soundness — adequate handles make boundary discovery robust under adversarial hiding.
- MB7b — Filter Coverage — the oversight filter family has adequate resolution to bound what audits miss.
- MB7c — Bounded Hidden Capability — bounded hidden productive BIQ plus correction integrity supports adversarially robust oversight (cost of faking the monitored signal).
- MB7d — Acausal Coordination — inferential coupling after channel severance; see that card for acausal trade, ECL, and severance-detector evidence.
Splitting the wall does not make any piece less hard. It lets a reviewer ask which sub-bridge an audit actually tested.
Where agendas agree (MB7a–c): Hubinger inner/mesa; Redwood control; Apollo scheming; scalable-oversight ceiling. Where they diverge: the field often treats inner alignment as one binary; this project splits access, filter coverage, and cost of faking; do not merge with MB10 Successor Gaming. Field “deceptive alignment” usually tags MB7; MB10 = gaming the successor certification channel.
What would count as evidence?
Evidence would include access-robustness tests under adversarial boundary discovery, filter-coverage audits, and inferential-coupling certificates validated against known coordination channels.