Badge index

open

Open status

Unresolved research problem or falsifier — this project names it explicitly rather than smuggling it in as done.

36 cards

Agency-detect (sibling)

Recovers which parts of a raw time series look like separate agents, without assuming the agent list in advance. Sibling repository; a precursor to the in-repo simulations.

Arena Elo versus honesty

Public Chatbot Arena Elo, joined to MASK honesty: the ranking tracks accuracy, not honesty. Do not read the public proxy as evidence that honesty is being selected.

CIRIS Named-identity mock

A CIRIS-shaped scripted mock where the named safety path is green, the authority defers, and tools plus cached memory still change the world. Would that check have stopped what happened next?

CIRISAgent mock-LLM deferral

A mock-LLM CIRISAgent stack: deferral was logged, and no post-defer world-effect hit was observed on this freeze. That is a null, not a demonstration that deferral stopped a previously working bypass.

Country surveys (refused)

Further country- or page-level value surveys were refused: they repeat the wrong unit (a place, not a person). We do not stretch the question to fit the table.

CPC2015 risky choice

A published risky-choice experiment: a many-feature pattern did not beat a one-number expected-value score or a person intercept. Recorded as a null; not retuned.

Debian Stretch freeze

A release-critical bug kept gstreamer 0.10 out of Debian 9. The freeze rule actually excluded the package. A process can have a stop that bites.

ET Line: External-Substrate Transfer Tests

The ET line runs frozen, unmodified project instruments against traces or substrates this project did not author — four annexes to date (ET-1 through ET-4), each with its own pre-registration and stop/close criteria — to test whether findings generalize beyond hand-built or blindly-grown in-repo ecologies.

ET-1 Orbit (stopped)

Frozen unit-discovery code run on an external multi-agent collusion benchmark this project did not write. Stopped because the substrate was the wrong kind of test (shared scripts, not hidden coordination).

ET-2 CIL basin_stability (null)

The same frozen unit-discovery code on a shared-commons simulation with learned policies. Passive detection was a null; an intervention probe found a shared-resource smear, not a pairwise unit.

GPLv2 source versus install rights

GPLv2 can stay green (source with the binary) while the user cannot install a modified binary. GPLv3 later added an install-information handle. A distribution checklist is not the same as a user halt.

Inferential Coupling and Acausal-Trade Detection

Systems can coordinate without messages — through shared ancestry, self-prediction, or full acausal trade. The book turns this from a decision-theoretic stipulation into a measurable trajectory property: an inferential-coupling score over UAD-discovered agents, with a proved negative direction.

Lab-layer simulation

A lab defined in code, with real operating-system process isolation and a frozen grading referee. Pre-registered batteries; recorded failures included.

Lean pin of the named-path bypass

The authored mock’s shape is pinned in Lean: the named path can be green while a bypass count stays above zero. That is a logic pin of the [CIRIS named-identity mock](/cards/experiment/w-1/), not a live-system run.

Linux review tags and reverts

Kernel git history: “Reviewed-by” on commits later labeled bug-introducing, a revert that did not stop the same class of change, and -stable tags that do not mean the hunk is identical. Checklists can stay green while the harm continues.

Moral Machine country scores

Published country-level Moral Machine scores: two countries can look the same on a one-number “spare more people” score while their full tradeoff pattern is far apart. A 1-D score is not the whole value bundle.

Moral Machine same-person choices

Individual Moral Machine respondents: a many-feature pattern of their choices predicts held-out choices better than a one-number “spare more people” score. The 1-D score is not the same-person policy.

Supreme Court justice votes

Supreme Court Database justice votes: a many-feature pattern of a justice’s votes predicts held-out votes better than issue-area-only or a justice intercept. Observational; not a claim about correction channels.

The Static Target Trap

Treating human values as a fixed object to be found and encoded misses that they are dynamically maintained, socially mediated, and constantly revised.

Toy simulation

A small, fast simulator of corrections that look compliant but do not change later behavior, scored at several levels of intervention access.

Wikipedia review, bots, and socks

Encyclopedia logs: a bot that passed review and was later blocked, helper socks that mark things reviewed, and vote data that cannot be treated as a causal “correction worked” estimator.

All badges · All cards