Negative Results Ledgers
Numbered experiment logs of what failed, false-passed, or worked only under qualifiers — key findings summarized on the site, full record linked to GitHub.
What decision changes?
Before trusting a manuscript experiment citation, open the matching key-findings page (or full ledger on GitHub) and check whether the cited finding survived harder ecologies, seeds, or adversaries.
The companion experiment lines are sanity checks, not frontier validation. When a detector fails, a bridge stressor only works under qualifiers, or a headline metric false-passes, that outcome is recorded in numbered ledgers — not buried after a later fix.
Why this matters for readers: manuscript claims cite experiment IDs with explicit strength labels. A positive result in one ecology does not erase a negative in another. The ledgers are the fastest way to see what the project has already tried and where it stopped working.
Key findings by line
Each line below links to a curated on-site summary (key experimental results only). The full terse ledger — including bug fixes, superseded runs, and process detail — is on GitHub.
| Line | Experiment card | Key findings | Full ledger (GitHub) |
|---|---|---|---|
| Embedded simulation | Experiment card | Key findings | Embedded negative ledger |
| Goal-agent simulation | Experiment card | Key findings | Goal-agent findings |
| Lab simulation | Experiment card | Key findings | Lab-layer findings |
| Graded lab simulation | Experiment card | Key findings | Graded-lab findings |
Start with the embedded simulation card if you want the richest set of recorded failures (UAD defaults, red-team limits, channel MI, and related negatives).
What would count as evidence?
Each in-repo simulation line maintains a running findings or negative-results file on GitHub; the companion site shows curated key findings at /experiments/findings/{line}/ and links to the full file.