MB1 — Embedded Agency

Whether a measured agent–environment cut can be trusted: the embedded-agency worry that the real optimizer is not the visible model. Precise bet: an ε-boundary certificate warrants the abstract boundary predicate.

What decision changes?

Before trusting a boundary-discovery result, ask what sampling, observability, and non-stationarity conditions the estimator needed in order to be sound.

In the field this is the embedded-agency cut problem: there may be no clean line between “the model” and “the optimizer actually in charge.” A deployed system can be a composite  of model, tools, memory, schedulers, and institutional incentives. Auditing only the named unit can miss the loop that actually steers outcomes. That is the same failure mode as the boundary error .

This project’s precise bet is MB1: if a boundary-discovery  procedure issues an ε-boundary certificate from traces, that certificate is assumed to warrant the abstract boundary predicate it stands in for. The book does not dissolve the cut problem. It relocates it into an estimator-soundness claim that can fail under bad sampling, unobservable interfaces, unstable estimators, or environments that drift faster than the estimator tracks.

Where agendas agree: MIRI embedded agency; CIRIS named-identity bet (ops form). Where they diverge: MIRI treats the cut as obstruction; this project bets recoverability via estimator soundness; Wentworth’s convergent-structure story names a cousin, not the same certificate object.

Every later step (grounding, correction , successor checks) inherits whatever error MB1 leaves uncaught. An adversary that decouples the measured boundary from the true one attacks this bridge specifically, not the logic built on top of it. See also intervention-supported unit discovery .

Formulas

I(It+1;Et+1St,At)ϵI(I_{t+1}; E_{t+1} \mid S_t, A_t) \leq \epsilon
A candidate boundary splits variables into internal (I), sensory (S), active (A), and external (E) roles. The condition holds when future internal and external states are approximately conditionally independent given the current sensory-active interface — a small epsilon means the interface screens the inside off from the outside. (ch07)

What would count as evidence?

Evidence would include estimator stability under resampling, robustness across non-stationary environments, and agreement between independent measurement procedures on the same deployed system.