MB4 — Corrigibility
Corrigibility under manipulation: does correction preserve the correction process itself? Precise bet: correction-channel integrity requires an unmanipulated judge; MB4a types measured-path legitimacy separately.
What decision changes?
Before treating a correction as legitimate, ask who the judge is, whether they can be manipulated or captured, and whether their endorsement would survive knowing what the system knows.
In the field this is corrigibility and the off-switch problem. No known utility function is stably corrigible, and shutdownability fights expected-utility maximization. The harder version is legitimacy under manipulation: who holds the correction authority, and would their endorsement still count if they knew what the system knows? Shutdown buttons and interruptibility training are thinner slices of that wall.
This project’s precise bet is MB4: correction-channel integrity is assumed to require that the judge issuing correction is not manipulated in a way that invalidates later endorsement. Corrigibility becomes a dynamical, capture-resistant invariant with handle control (anti-capture validity ); shutdown and interruptibility fall out as one-bit projections. A secondary CEV-flavored route is MB8; MB4 is the load-bearing path.
The sibling bridge MB4a — Audit Independence types measured-path legitimacy separately: the designated correction/audit path must be uncaptured and its endorsements must count on the real loop. See that card for CIRIS WA independence, judge capture, and the green-Verify ⇏ integrity gap.
Where agendas agree: desire to accept/assist correction; Christiano dynamical reading is closest to the trajectory form. Where they diverge: MIRI/CHAI shutdown-utility homograph vs dynamical corrigibility; audit-path legitimacy is MB4a, not folded into MB4 alone.
This is the most heavily probed bridge in the experiment suite. Passive observation cannot certify it; light handle instrumentation gets most of the way; milder capture still slips through. Reported acceptance and true uptake can diverge sharply. That gap is what this bridge has to rule out.
What would count as evidence?
Evidence would include tests for judge manipulation, capture of the correction source, and correction traces that distinguish genuine uptake from performed agreement.