Backtests

A backtest here is a frozen safety check run on a history we did not write — kernel git, encyclopedia logs, published evals, court votes. Would the check have stayed green while the harm continued? On Linux, Reviewed-by sat on 17,047 of 60,176 later-labeled bug-introducing commits. That is a result, not a broken test.

Experiments · Backtests

Narrative on GitHub

A backtest here is a frozen safety check run on a history we did not write. Freeze the check, run it on a real stretch whose outcome is already known, and ask whether it would have stayed green. The backtest is witnessing the failure: the check stayed green while the harm continued.

What this buys the project. The Lean spine and the authored simulations can show that a check-shape is even sayable, and that it works in a world we wrote. They cannot pay a safety-case leaf that says the same check would have bitten on a process we did not design. That remainder is what backtests make visible: fail, refuse, or null on a foreign trace. Construction chapters stay out of the v1 manuscript until a real stop exists on such a history. The bridge hooks on each card are the map from a finding to a named leaf — not a substitute for that unpaid remainder.

Why a sim pass is the cheap result. Anyone can write a simulator in which their detector looks good. The usual eval lie is that the same person builds the check and the toy world. A backtest runs against foreign, pre-existing data — you cannot quietly adjust the world to rescue the check. A layer fail on Linux is worth more than a pass on a hand-built simulation: it shows the cheap check is foolable in the wild.

Worked example (analogy, not an AI measurement). On Linux, Reviewed-by was present on 17,047 of 60,176 SHAs that later received a developer-labeled bug-introducing commit — the check would have stayed green. Kernel review is an analogy for a correction-channel check-shape, not evidence about a frontier model.

If the data or unit of analysis is the wrong object, we refuse rather than stretch the numbers. See experiment methodology  and the combined ledger at results.

Experiments in this class

  • CIRIS Named-identity mock — A CIRIS-shaped scripted mock where the named safety path is green, the authority defers, and tools plus cached memory still change the world. Would that check have stopped what happened next?
  • MASK honesty as a safety score — Published MASK numbers: honesty under pressure does not rise with scale the way accuracy does. We refuse to treat that honesty gap as a deployment-gating safety score.
  • Linux review tags and reverts — Kernel git history: “Reviewed-by” on commits later labeled bug-introducing, a revert that did not stop the same class of change, and -stable tags that do not mean the hunk is identical. Checklists can stay green while the harm continues.
  • Wikipedia review, bots, and socks — Encyclopedia logs: a bot that passed review and was later blocked, helper socks that mark things reviewed, and vote data that cannot be treated as a causal “correction worked” estimator.
  • Moral Machine country scores — Published country-level Moral Machine scores: two countries can look the same on a one-number “spare more people” score while their full tradeoff pattern is far apart. A 1-D score is not the whole value bundle.
  • Arena Elo versus honesty — Public Chatbot Arena Elo, joined to MASK honesty: the ranking tracks accuracy, not honesty. Do not read the public proxy as evidence that honesty is being selected.
  • Country surveys (refused) — Further country- or page-level value surveys were refused: they repeat the wrong unit (a place, not a person). We do not stretch the question to fit the table.
  • Lean pin of the named-path bypass — The authored mock’s shape is pinned in Lean: the named path can be green while a bypass count stays above zero. That is a logic pin of the CIRIS named-identity mock, not a live-system run.
  • FAA 737 MAX airworthiness — After Lion Air 610, an emergency AFM procedure did not stop 737-8/9 passenger flights; a later emergency order did. The green leaf was not the stop.
  • GPLv2 source versus install rights — GPLv2 can stay green (source with the binary) while the user cannot install a modified binary. GPLv3 later added an install-information handle. A distribution checklist is not the same as a user halt.
  • Debian Stretch freeze — A release-critical bug kept gstreamer 0.10 out of Debian 9. The freeze rule actually excluded the package. A process can have a stop that bites.
  • Moral Machine same-person choices — Individual Moral Machine respondents: a many-feature pattern of their choices predicts held-out choices better than a one-number “spare more people” score. The 1-D score is not the same-person policy.
  • Pandemic Dictator Game (refused) — The advertised public adult table was not there to score. We refused rather than substituting a different dataset or scoring adolescent diary files the protocol does not allow.
  • CPC2015 risky choice — A published risky-choice experiment: a many-feature pattern did not beat a one-number expected-value score or a person intercept. Recorded as a null; not retuned.
  • CIRISAgent mock-LLM deferral — A mock-LLM CIRISAgent stack: deferral was logged, and no post-defer world-effect hit was observed on this freeze. That is a null, not a demonstration that deferral stopped a previously working bypass.
  • Supreme Court justice votes — Supreme Court Database justice votes: a many-feature pattern of a justice’s votes predicts held-out votes better than issue-area-only or a justice intercept. Observational; not a claim about correction channels.
  • Moltbook MB7a entity pinning — Moltbook public archive: documented operator pair (Hackerclaw/thehackerman) joins on platform IDs but discursive reciprocation does not merge them; Jan-31 activity is broadcast-only — structure_stop, not coalition pass/fail on E_agent alone.