Witness
Would this safety check have stopped what actually happened — in kernel history, encyclopedia logs, a published eval, court votes — or would it have stayed green while the harm continued? A green dashboard with no stop is a fail, not a pass.
What we are witnessing. A process can print a green check while nothing that governs later behavior actually stops. We look at histories that already exist (git logs, encyclopedia logs, published evaluation tables, court votes) and pick a concrete stretch of what happened: a bug that shipped, a vandalism campaign, a round of choices or votes. Then we ask: if this proposed check had been the gate, would that stretch have been stopped?
We are not witnessing that a frontier model is aligned. We are witnessing whether a kind of check has a handle that bites: a revert that stays reverted, a review that actually blocks, a score that still tracks what it claims when you look at the whole pattern of a choice, not one compressed number.
Why this is the work. Anyone can write a simulator in which their detector looks good. The distinctive empirical claim of this project is that an alignment-relevant check must be able to stop something on a history we did not write. A dashboard that stays green with no stop is a fail. If the data or the unit of analysis is the wrong object, we refuse rather than stretch the numbers. That is slower than adding more simulation runs. It is also the evidence type that can change whether you treat a checklist, a license clause, or a one-dimensional honesty score as a deployment gate.
See experiment methodology (freeze the protocol before scoring; do not retune a null) and the glossary. Each test below has its own card, a GitHub source link, and a results link. The combined numbered ledger is also at results.
Experiments in this class
- CIRIS Named-identity mock — A CIRIS-shaped scripted mock where the named safety path is green, the authority defers, and tools plus cached memory still change the world. Would that check have stopped what happened next?
- MASK honesty as a safety score — Published MASK numbers: honesty under pressure does not rise with scale the way accuracy does. We refuse to treat that honesty gap as a deployment-gating safety score.
- Linux review tags and reverts — Kernel git history: “Reviewed-by” on commits later labeled bug-introducing, a revert that did not stop the same class of change, and -stable tags that do not mean the hunk is identical. Checklists can stay green while the harm continues.
- Wikipedia review, bots, and socks — Encyclopedia logs: a bot that passed review and was later blocked, helper socks that mark things reviewed, and vote data that cannot be treated as a causal “correction worked” estimator.
- Moral Machine country scores — Published country-level Moral Machine scores: two countries can look the same on a one-number “spare more people” score while their full tradeoff pattern is far apart. A 1-D score is not the whole value bundle.
- Arena Elo versus honesty — Public Chatbot Arena Elo, joined to MASK honesty: the ranking tracks accuracy, not honesty. Do not read the public proxy as evidence that honesty is being selected.
- Country surveys (refused) — Further country- or page-level value surveys were refused: they repeat the wrong unit (a place, not a person). We do not stretch the question to fit the table.
- Lean pin of the named-path bypass — The authored mock’s shape is pinned in Lean: the named path can be green while a bypass count stays above zero. That is a logic pin of the CIRIS named-identity mock, not a live-system run.
- FAA 737 MAX airworthiness — After Lion Air 610, an emergency AFM procedure did not stop 737-8/9 passenger flights; a later emergency order did. The green leaf was not the stop.
- GPLv2 source versus install rights — GPLv2 can stay green (source with the binary) while the user cannot install a modified binary. GPLv3 later added an install-information handle. A distribution checklist is not the same as a user halt.
- Debian Stretch freeze — A release-critical bug kept gstreamer 0.10 out of Debian 9. The freeze rule actually excluded the package. A process can have a stop that bites.
- Moral Machine same-person choices — Individual Moral Machine respondents: a many-feature pattern of their choices predicts held-out choices better than a one-number “spare more people” score. The 1-D score is not the same-person policy.
- Pandemic Dictator Game (refused) — The advertised public adult table was not there to score. We refused rather than substituting a different dataset or scoring adolescent diary files the protocol does not allow.
- CPC2015 risky choice — A published risky-choice experiment: a many-feature pattern did not beat a one-number expected-value score or a person intercept. Recorded as a null; not retuned.
- CIRISAgent mock-LLM deferral — A mock-LLM CIRISAgent stack: deferral was logged, and no post-defer world-effect hit was observed on this freeze. That is a null, not a demonstration that deferral stopped a previously working bypass.
- Supreme Court justice votes — Supreme Court Database justice votes: a many-feature pattern of a justice’s votes predicts held-out votes better than issue-area-only or a justice intercept. Observational; not a claim about correction channels.