An outside test failed to find a hidden team of agents
External Test 2 (ET-2) tested the project's boundary-finding method in an independently built multi-agent commons simulation. Across 150 runs, it did not recover the planted adversarial subgroup. A separate small test found broad spillover from changing one agent, which is not the same as finding a meaningful unit.
What decision changes?
Treat this as a real limit of the current test, not a reason to retune it. Before running a larger intervention study, test whether the broad spillover is just a shared-resource effect.
A boundary-finding test found no hidden subgroup; a follow-up showed that changing one agent can affect everyone without revealing a distinct team.
An outside test produced a useful negative result: the project’s boundary-finding method did not recover a hidden adversarial group.
External Test 2 (ET-2) took the project’s existing method for finding meaningful groups of interacting agents and tried it in Collective-Intelligence-Library (indepdently build by [Equilibria Networks)(https://eq-network.org/lab)), a simulation of 20 agents managing a shared commons. The agents learned their choices rather than following a hand-written script. Some were trained to damage the shared resource. The simulation is not itself a gradual-disempowerment model; it is a related test of governance under pressure.
The main result: in all 150 runs, the method found no links between agents. It consequently treated every agent as separate and did not recover the planted adversarial group. This was a real negative result, not a software failure: the agents’ choices varied, but not in a pattern for the unchanged test to detect. It does not show that the group could never be found by another method.
A separate, very small follow-up changed one agent’s choice by force. Other agents changed their behavior too, but the effect was spread broadly across the population. That is what we would expect if everyone is connected through the shared resource; it does not identify a particular team. Before treating this as evidence of a meaningful group, a larger study needs a control that tests whether the same broad effect appears even without a real intervention.
The practical lesson is a limitation of the current form of Unsupervised Agent Discovery: it cannot discover coordinated agency that is diffused through a shared resource. The next test must distinguish a special group from ordinary spillover through a shared resource, without changing the original test just because the result was negative.
Technical artifacts: PLAN_ET2.md, results/et2a_uad_battery.json, and the graded-lab findings ledger.
Read more in: Ch. 7, Finding the Boundary; Ch. 34, Alignment Is Selected or Destroyed by Its Environment; and Appendix N, Experimental Evidence: Findings by Line.