Kosoy / infra-Bayesianism & LTA
Can a learning-theoretic theory of intelligent agents be built tightly enough to license alignment claims under stated assumptions, including misspecification, inner daemons, and recursive self-improvement? For Physicalist Superimitation: can agent detection and user identification pick out the intended user (including under simulation hypotheses), and does superimitation of inferred values stay well-defined when the world-model ontology is wrong?
Introduction
Vanessa Kosoy’s learning-theoretic agenda (LTA) aims to create a general mathematical theory of intelligent agents, so that alignment claims can be proved or rigorously conjectured relative to stated assumptions. Frequentist guarantees (regret bounds) are a standard of understanding in that theory. Nonrealizability (the true environment is not in the hypothesis class) and inner daemons are among the problems the theory is built to treat. Infra-Bayesianism is a major constructive layer for deep uncertainty—not the whole agenda. Infra-Bayesian physicalism (later also Formal Computational Realism) adds a physicalist/computational-realist ontology; the bridge transform is a core construction of that layer, with a role independent of outer alignment. Physicalist Superimitation (PSI; formerly PreDCA; later also Computational Superimitation) is a hypothesized protocol on top of that stack: identify the user as an agent in the physicalist ontology and superimitate their values (learn them and pursue them substantially better). The 2022 PreDCA writeup centered a “precursor” pointer; later PSI formulations do not treat precursor as central.
Who carries it: Vanessa Kosoy (+ Appel; logical-induction neighborhood via Garrabrant)
What they aim to do. Create a general mathematical theory of intelligent agents, and—as one hypothesized application, not the agenda’s main motivation—a protocol that reliably learns and acts on the user’s values (Physicalist Superimitation).
The hard question. Can a learning-theoretic theory of intelligent agents be built tightly enough to license alignment claims under stated assumptions, including misspecification, inner daemons, and recursive self-improvement? For Physicalist Superimitation: can agent detection and user identification pick out the intended user (including under simulation hypotheses), and does superimitation of inferred values stay well-defined when the world-model ontology is wrong?
What they produce. The learning-theoretic agenda, the Infra-Bayesianism sequence, infra-Bayesian physicalism, and the Physicalist Superimitation protocol (formerly PreDCA).
Key terms. Key terms include the learning-theoretic agenda, infra-Bayesianism, infra-Bayesian physicalism / Formal Computational Realism, the bridge transform, regret bounds, nonrealizability, daemons, and Physicalist Superimitation / PreDCA / Computational Superimitation.
Related field cruxes. Embedded Agency; Value Learning; Value Referent; Tiling; Inner Alignment; Grounding Drift
What they contribute. Model-class misspecification and grain-of-truth analysis; regret-bounded agents; inner daemons; infra-Bayesian physicalism and the bridge transform as a physicalist layer; Physicalist Superimitation as a sibling outer-alignment protocol (superimitation after agent detection and user identification), reached through a different formal path than CIRL-style pointing.
How this project treats it. This project maps LTA’s problems onto Embedded Agency, Value Learning, Value Referent, Tiling, Inner Alignment, and Grounding Drift, with Acausal Coordination as a logical-induction neighborhood cousin—it does not treat LTA as a replacement ontology. On the outer endpoint, PSI is a peer proposal for learning and acting on the user’s values. This project names separately three checks it still wants: whether inferred values remain usable directions of control after transformation (not only preserved labels), whether they keep applying to the right persons or processes , and whether a correction process stays open. Kosoy’s protocol is meant to address usable control (superimitation of inferred user values). Whether those three remain separately checkable after PSI, or whether the protocol already discharges them, is an open disagreement—not a claim that PSI is missing those pieces in her terms.
Specify / construct (peer outer target)
Not a ConstitutionalRule instance. PSI is a peer outer-target: identify the user and superimitate inferred values under infra-Bayesian physicalism. The bridge transform is an IBP construction (physicalism / cartesian privilege), not the outer-alignment mechanism. PreDCA’s “precursor” pointer is the earlier formulation. This project tags PSI on MB2/MB3, not as a specify-schema filling. Listed so the table does not hide a major outer-alignment construction bet.
Links
- LTA (2018 overview)
- LTA status (2023)
- Infra-Bayesianism (LessWrong)
- Infra-Bayesian physicalism
- PreDCA (Alignment Forum tag)
- Vanessa Kosoy
See the coverage matrix for evidence tagged to this agenda, and the glossary for shared terms.