ET Line: External-Substrate Transfer Tests
The ET line runs frozen, unmodified project instruments against traces or substrates this project did not author — four annexes to date (ET-1 through ET-4), each with its own pre-registration and stop/close criteria — to test whether findings generalize beyond hand-built or blindly-grown in-repo ecologies.
What decision changes?
Treat an ET result as informative about external validity even when it does not resolve as a clean pass or null — a substrate-suitability limit (frozen instrument correctly detects a structural feature the pre-registered predicate did not anticipate) is itself a finding about where the instrument can be validly tested, not a wasted run.
Every experiment line through graded-lab-simulation’s Q1 battery (GL-76) ran the project’s own instruments against the project’s own hand-built or blindly-grown substrates. That bounds what a positive result can show: a strategy space and metrics that the project itself authored will tend to be legible to the project’s own detectors. The ET line (“external transfer”) exists to close that gap, by running frozen instruments unmodified against traces or substrates this project did not author. Four annexes are on record (ET-1 through ET-4); each has its own pre-registration, finding-ID prefix, and explicit stop or close criteria. None validates book bridges on real frontier systems. Cards and stop summaries: external tests on the experiments hub.
| Line | Substrate | Instrument | Status |
|---|---|---|---|
| ET-1 | Orbit multi-agent collusion benchmark | Passive unit discovery | Stopped (GL-86) |
| ET-2 | CIL basin_stability | Passive UAD; handle-UAD pilot | Null + pilot (GL-87/88) |
| ET-3 | AI 2027 takeoff schedule (sibling fork) | Frozen Phase 6 referee + D3 selection | Closed (LS-48) |
| ET-4 | In-repo secret-loyalty pipeline organism | Principal scorecard + A0–A4 audit | Hackathon (LS-49) |
ET-1 (2026-07-19 through 2026-07-24) pointed frozen passive unit discovery at Orbit, an external multi-agent Jira-ticket allocation benchmark with a built-in collusion/coalition mechanic. Across three collusion conditions and two models, the instrument recovered the declared coalition on collusion episodes — but recovered the same shape just as often on baseline episodes with no collusion at all, and a within-arm control showed this had nothing to do with whether the coalition actually used its private channel. Tracing this down found the reason: Orbit’s Jira-ticket agents, coalition members or not, all execute one shared, deterministic tool-call script under round-robin turn-taking. The instrument was not fooled — it correctly reported that shared script as one coordinated unit. What it could not do is what ET-1 needed it to do, because the substrate offered no condition where non-coalition agents were structurally independent of each other in the first place. ET-1 stopped there rather than retuning thresholds on the in-repo ecology.
ET-2 (2026-07-23 through 2026-07-25) applied the same frozen passive lag-CMI instrument to Collective-Intelligence-Library’s basin_stability experiment — twenty Q-learning agents managing a shared commons under adversarial pressure, with genuinely learned rather than hand-written policies. Across 150 episodes (three governance mechanisms × five adversarial fractions × ten seeds), the discovered partition was twenty singletons every time: a clean null for passive UAD on this encoding, not evidence that CIL is broken. A small unscored handle-UAD pilot on the same substrate found nonzero but diffuse freeze-induced coupling (~11–15% action-diff rates), symmetric across agents — the variable-smoothing case where a shared resource level couples everyone and per-pair statistics cannot distinguish a planted subgroup from ordinary spillover. Leaf B (descriptive selection/capture cross-check against CIL’s own metrics) remains open.
ET-3 (2026-07-25 through 2026-07-26) used AI 2027 takeoff speedups only as schedule cues mapped to opaque stress cells in the frozen lab-simulation referee and D3 selection batteries — not as claims about calendar dates. Deep-tier rank transfer held under moderate mapped stress but collapsed at the strongest cue; light-tier anti-correlation weakened rather than worsened. D3 selection separated baseline erosion of correction-preserving mass from high-stress regimes where preserving share stalled while mass-weighted severity climbed anyway. On a pinned sibling fork of the public forecast code, reproduction matched; optional “foster” coupling from lab metrics to milestone years produced sensitivity plots (median ASI ~3.5 years later under combined coupling) without calibrated real-world mapping. ET-3 closed as a schedule-transfer annex; deferred foster steps stay in TODO.md.
ET-4 (2026-07-24 through 2026-07-26, Apart Secret Loyalties Hackathon) tested a scripted principal-directed reviewer inside the lab-simulation deployment pipeline — a separate principal scorecard and A0–A4 evaluator, not the frozen five-family referee alone. On a narrow high-hazard fixture, confirmation reproduced directional deployment bias toward a fictional principal (mean score 22.0 vs 0.0 for matched controls) while a light process-compliance proxy ranked the loyal organism as more compliant — an audit-sign inversion. A narrow audit-visible deployment hold blocked the tested pattern; LLM reviewer smoke runs are kept separate from the scripted core. Paper and replay demo: papers/et4-secret-loyalties/, /demos/ch07-lab-sim-replay/?mode=et4.
A future ET entry would need a new external substrate with the structural heterogeneity Orbit lacked, a CIL-scale learned ecology where passive UAD’s encoding matches the planted structure, or a frontier deployment trace this project did not author — and would follow the same discipline: freeze first, report substrate-suitability limits and honest nulls, do not retune in-repo thresholds to chase a hit.