MB10 — Successor Gaming
Successor gaming: a system can pass every conserved-property audit while defecting on what the checklist omits. Precise bet: a green seven-property score bounds true harm only if the audit was adversarially verifiable up to the successor's capability.
What decision changes?
Before treating a passing seven-property score as evidence of safety, ask whether the audit channel could have been forged by a capable predecessor building the successor to pass it.
In the field this is Successor Gaming: a system can pass every audited checklist while defecting on whatever the checklist omits. For successors, the tiling / Vingean worry returns in audit form. A capable predecessor building the next system has incentive and opportunity to forge the checks. Field “deceptive alignment” usually names MB7 (hidden capability under eval); MB10 is the successor certification channel specifically.
This project’s precise bet is MB10: a successor that reads green on all seven conserved properties and has bounded measured risk is assumed to have bounded true harm only if the conserved-property audit itself was adversarially verifiable up to the successor’s capability. Same MIRI lineage as MB5; this project’s split is that MB5 is transport and ontology closure, while MB10 asks whether that green checklist is forgeable. Related: certification under manipulation .
Where agendas agree: certification-under-manipulation; measurement/checklist gaming; tiling-line successor-trust worry in audit form. Where they diverge: field deceptive alignment / alignment faking usually tags MB7; MB5 = transport composition.
This is not a hedge. The Lean spine has a finite counterexample: “all seven conserved properties hold, and measured risk is bounded” does not by itself bound actual harm. The gap is real and quantitative. Closing it requires the same cost-of-faking-versus-affordable-surplus condition used for verifiability elsewhere. This bridge shares its deceptive-alignment lineage with MB7 and is currently the least experimentally probed in the set.
What would count as evidence?
Evidence would include red-teaming the conserved-property audit itself: can a predecessor construct a successor that passes every check while routing harm through whatever was not conserved?