Graded-capability lab simulation
Successor to the lab-layer simulation: a continuously scored pipeline where ambiguity about who is acting is supposed to emerge from a population, not from a noise dial.
Successor substrate to the lab-layer simulation, not a further phase of it: a graded, continuously scored pipeline of tasks, shared resources, and agent viability, built so that ambiguity about who is doing what emerges from a population the experimenters did not hand-design, rather than from a noise or delay dial they turn up. Boundary discovery drives a continuous boundary-information-quality score over units the method infers rather than a hand-labeled roster; an emergent-ambiguity index is measured from two vantage points (the acting agent’s own full view versus a referee limited to what it can observe); a first multi-episode selection battery; population growth left to run on its own rather than hand-authored; a pre-registered test of whether the detection method still works when transplanted onto that grown population (Q1, GL-76); a set of rigs testing each of this project’s bridge assumptions independently (GL-79—GL-85); and two attempts at transferring the whole approach to external multi-agent benchmarks this project did not build, Orbit (GL-86) and the Collective-Intelligence-Library’s basin_stability (GL-87/GL-88). The first program version is closed (2026-07-15, GL-31). The second version and its follow-up closed without ever producing a population that passed its own criteria (GL-38/GL-43). The third version’s institutional runtime is active through a criteria freeze (GL-53). Methodological sandbox only: the present strategy space and metrics remain project-authored; the external transfers (GL-86—GL-88) are the first attempts to test the method on substrates this project did not build. All three closed as substrate-suitability findings rather than validated positive or negative transfers: passive UAD found a substrate-inflated true positive on Orbit (GL-86) and a flat null under converged learned policies on basin_stability (GL-87); a follow-on intervention-based probe on the same CIL substrate found a real but diffuse effect consistent with a shared-resource observation channel, not a discoverable pairwise unit (GL-88).
Finding.
UAD recovers a communicate-free write/read coordination handoff invisible to every prior passive/intervention method, via a standalone all-pairs dependency-score probe (GL-11); the same probe over-merges a resource-bound bystander into a 3-way approval barrier on a blind-designed scenario, a genuine, unpatched negative (GL-12). Full results
How to read
Ambiguity and boundary-discovery engine: a boundary-information-quality score over inferred (not labeled) units, a “high” ambiguity band that is structurally unreachable from the agent’s own vantage point, a “mid” band reachable only from a referee’s vantage point on the hand-designed substrate (with a null result on external validity when transplanted onto the grown population, GL-76/GL-84), a selection-battery null (GL-23—GL-27), and independent per-bridge-assumption rigs (GL-79—GL-85). Boundary discovery , including its access-aware variant, has been methodologically sound since GL-51. Negative ledgers and per-rig skips are features, not bugs. Two external-substrate transfers followed: against Orbit (GL-86), where the passive detection method correctly detects the substrate’s own scripted-agent coupling rather than testing communication-mediated collusion; and against the Collective-Intelligence-Library’s basin_stability (GL-87/GL-88), where passive detection found nothing under converged learned (Q-learning) policies, but an intervention-based freeze probe on the same substrate found a real, diffuse effect consistent with every agent sharing one resource-level channel rather than any pairwise unit. All three closed as substrate-suitability findings rather than a clean external pass or fail.