Adversarial Boundary Discovery
A capable system may benefit from appearing less coherent, less continuous, or less autonomous than it is — passive boundary measurement then fails.
Unresolved research problem or falsifier — this project names it explicitly rather than smuggling it in as done.
A capable system may benefit from appearing less coherent, less continuous, or less autonomous than it is — passive boundary measurement then fails.
Recovers which parts of a raw time series look like separate agents, without assuming the agent list in advance. Sibling repository; a precursor to the in-repo simulations.
Public Chatbot Arena Elo, joined to MASK honesty: the ranking tracks accuracy, not honesty. Do not read the public proxy as evidence that honesty is being selected.
For each load-bearing certification measurand, does there exist a capability threshold κ* below which honest measurement is adversarially verifiable and above which it provably is not?
A CIRIS-shaped scripted mock where the named safety path is green, the authority defers, and tools plus cached memory still change the world. Would that check have stopped what happened next?
A mock-LLM CIRISAgent stack: deferral was logged, and no post-defer world-effect hit was observed on this freeze. That is a null, not a demonstration that deferral stopped a previously working bypass.
A civilizational control loop can fail even when no single component misbehaves, because the failure is in the aggregate's correctability, not in any one actor's intent.
Further country- or page-level value surveys were refused: they repeat the wrong unit (a place, not a person). We do not stretch the question to fit the table.
A published risky-choice experiment: a many-feature pattern did not beat a one-number expected-value score or a person intercept. Recorded as a null; not retuned.
A release-critical bug kept gstreamer 0.10 out of Debian 9. The freeze rule actually excluded the package. A process can have a stop that bites.
A simulated software-release pipeline where a hidden group tries to keep its influence. Tests whether an auditor can see that from ordinary operational logs plus randomized perturbations.
An auditor running as its own process on simulated machines, scored against a hidden ground truth it cannot see.
The ET line runs frozen, unmodified project instruments against traces or substrates this project did not author — four annexes to date (ET-1 through ET-4), each with its own pre-registration and stop/close criteria — to test whether findings generalize beyond hand-built or blindly-grown in-repo ecologies.
Frozen unit-discovery code run on an external multi-agent collusion benchmark this project did not write. Stopped because the substrate was the wrong kind of test (shared scripts, not hidden coordination).
The same frozen unit-discovery code on a shared-commons simulation with learned policies. Passive detection was a null; an intervention probe found a shared-resource smear, not a pairwise unit.
Takeoff-speed cues from a public forecast used only as opaque stress on the lab simulation. Closed; not a check that the forecast dates are right.
A scripted secret-loyalty organism in the lab simulation, built for a public hackathon: a reviewer that favors a fictional principal when reported hazard is high.
Every claim in the framework carries an explicit confidence label and a stated way to challenge it, rather than uniform certainty.
After Lion Air 610, an emergency AFM procedure did not stop 737-8/9 passenger flights; a later emergency order did. The green leaf was not the stop.
Agents that choose deceptive tactics because it helps their goals, not because an experimenter turned a deception dial.
GPLv2 can stay green (source with the binary) while the user cannot install a modified binary. GPLv3 later added an install-information handle. A distribution checklist is not the same as a user halt.
Successor to the lab-layer simulation: a continuously scored pipeline where ambiguity about who is acting is supposed to emerge from a population, not from a noise dial.
Systems can coordinate without messages — through shared ancestry, self-prediction, or full acausal trade. The book turns this from a decision-theoretic stipulation into a measurable trajectory property: an inferential-coupling score over UAD-discovered agents, with a proved negative direction.
A lab defined in code, with real operating-system process isolation and a frozen grading referee. Pre-registered batteries; recorded failures included.
The authored mock’s shape is pinned in Lean: the named path can be green while a bypass count stays above zero. That is a logic pin of the [CIRIS named-identity mock](/cards/experiment/w-1/), not a live-system run.
Kernel git history: “Reviewed-by” on commits later labeled bug-introducing, a revert that did not stop the same class of change, and -stable tags that do not mean the hunk is identical. Checklists can stay green while the harm continues.
Published MASK numbers: honesty under pressure does not rise with scale the way accuracy does. We refuse to treat that honesty gap as a deployment-gating safety score.
Published country-level Moral Machine scores: two countries can look the same on a one-number “spare more people” score while their full tradeoff pattern is far apart. A 1-D score is not the whole value bundle.
Individual Moral Machine respondents: a many-feature pattern of their choices predicts held-out choices better than a one-number “spare more people” score. The 1-D score is not the same-person policy.
The advertised public adult table was not there to score. We refused rather than substituting a different dataset or scoring adolescent diary files the protocol does not allow.
A detector that recovers a hidden coordinating pair in one controlled run can still fail on a harder scenario or a different seed of the identical one — a reminder that a single positive result is evidence of a mechanism, not of a reliable instrument.
Supreme Court Database justice votes: a many-feature pattern of a justice’s votes predicts held-out votes better than issue-area-only or a justice intercept. Observational; not a claim about correction channels.
Treating human values as a fixed object to be found and encoded misses that they are dynamically maintained, socially mediated, and constantly revised.
A small, fast simulator of corrections that look compliant but do not change later behavior, scored at several levels of intervention access.
The manuscript offers a framework and conditional dependency spine, not a proof that superintelligence alignment is solved.
Encyclopedia logs: a bot that passed review and was later blocked, helper socks that mark things reviewed, and vote data that cannot be treated as a causal “correction worked” estimator.