Wentworth / natural abstractions

Do natural abstractions—the variables that survive selection—align with value-relevant structure when systems are trained or deployed at scale?

Introduction

John Wentworth’s natural-abstractions lineage develops the Natural Abstraction Hypothesis, selection theorems, and a compression-based theory of agency. It asks whether the abstractions that survive optimization align with value-relevant structure at scale. Convergent macro-variables can inform Value Learning targets but do not discharge Tiling transport on their own.

Who carries it: John Wentworth (+ LessWrong NAH cluster)

What they aim to do. Develop a mathematical theory of abstractions and agency that could constrain what alignment targets are even learnable under optimization pressure.

The hard question. Do natural abstractions—the variables that survive selection—align with value-relevant structure when systems are trained or deployed at scale?

What they produce. The lineage centers on the Natural Abstraction Hypothesis, selection theorems that predict which abstractions survive optimization, and a compression-based theory of agency.

Key terms. Signature terms include natural abstractions, selection theorems, natural latents, and agency as compression—the idea that useful macro-variables emerge reliably from messy micro-dynamics.

Related field cruxes. Value Learning; Value Referent; Embedded Agency

What they contribute. The Natural Abstraction Hypothesis acts as a falsifier for low-dimensional value stories: if the wrong latents are natural, naive Value Learning and Value Referent targets may misfire even when surface training looks successful (see ch17 WWCTV).

How this project treats it. This project types bundle geometry and ontology-shift transport separately; natural-abstraction convergence is informative but is not promoted to a load-bearing spine assumption—Tiling still needs its own discharge argument.

See the coverage matrix for evidence tagged to this agenda, and the glossary for shared terms.