Negative Results Ledgers

Numbered experiment logs of what failed, false-passed, or worked only under qualifiers — key findings summarized on the site, full record linked to GitHub.

What decision changes?

Before trusting a manuscript experiment citation, open the matching key-findings page (or full ledger on GitHub) and check whether the cited finding survived harder ecologies, seeds, or adversaries.

The companion experiment lines are sanity checks, not frontier validation. When a detector fails, a bridge stressor only works under qualifiers, or a headline metric false-passes, that outcome is recorded in numbered ledgers — not buried after a later fix.

Why this matters for readers: manuscript claims cite experiment IDs with explicit strength labels. A positive result in one ecology does not erase a negative in another. The ledgers are the fastest way to see what the project has already tried and where it stopped working.

Key findings by line

Each line below links to a curated on-site summary (key experimental results only). The full terse ledger — including bug fixes, superseded runs, and process detail — is on GitHub.

LineExperiment cardKey findingsFull ledger (GitHub)
Embedded simulationExperiment cardKey findingsEmbedded negative ledger
Goal-agent simulationExperiment cardKey findingsGoal-agent findings
Lab simulationExperiment cardKey findingsLab-layer findings
Graded lab simulationExperiment cardKey findingsGraded-lab findings

Start with the embedded simulation card if you want the richest set of recorded failures (UAD defaults, red-team limits, channel MI, and related negatives).

What would count as evidence?

Each in-repo simulation line maintains a running findings or negative-results file on GitHub; the companion site shows curated key findings at /experiments/findings/{line}/ and links to the full file.