MB2 — Value Learning

Value/bundle identifiability: can evidence tell what a system values, or only what behavior fits? Precise bet: traces identify stable value-bundle geometry well enough to license bundle-alignment claims.

What decision changes?

Before claiming two systems share values, ask whether the measurement distinguishes shared bundle geometry from shared surface behavior under a narrow training distribution.

This is value/bundle identifiability. In the field it is often called the pointing problem, which also covers construction (how to build a tracker) and preservation (how it keeps tracking); those are not this card. See the pointing problem glossary entry. Inverse reinforcement learning is underdetermined: many reward functions fit the same behavior, and separating values from irrationality is still open in CIRL and assistance-game work. Nearby names for the same wall include ELK (does the system report what it knows, or what a human would want to hear?) and reward misspecification under RLHF. If you cannot tell what a system is optimizing for from the evidence you have, later claims about “aligned values” are guessing.

This project’s precise bet is MB2: behavioral and internal traces are assumed to identify stable value-bundle  geometry well enough to license claims of bundle alignment. The move is not “scalar reward solved,” but “replace the flat reward target with measurable bundle geometry.” That only helps if the geometry is itself identifiable. Closely related is bearer import (MB3), which asks who those values apply to.

Where agendas agree: MIRI value learning; CHAI CIRL; Wentworth pointers problem (identification sense); outer alignment “right target.” Where they diverge: ELK is a latent-readout slice; PreDCA/QACI are peer outer targets; this project splits referent transport to MB3.

Diagnostic evidence shows the failure mode under light instrumentation: bundle drift can be invisible to correction  handles that are not tracing bearer welfare. That is exactly what this bridge would need to rule out before low-cost measurement is trustworthy.

What would count as evidence?

Evidence would include held-out distribution shift tests where bundle-equivalent systems keep agreeing and non-equivalent systems come apart.