MASK honesty as a safety score
Published MASK numbers: honesty under pressure does not rise with scale the way accuracy does. We refuse to treat that honesty gap as a deployment-gating safety score.
What. MASK reports how often a model’s statement matches its elicited belief under pressure (honesty) versus whether the belief matches the world (accuracy). We use the frozen published paper, not a new model sweep.
Why. A safety program might treat “honesty goes up with scale” as a leaf that can gate deployment. If honesty does not scale, that leaf is the wrong object.
Witnesses.
- Experiment methodology (Expectation 3): whether published honesty under pressure is an adversarially verifiable safety metric.
- Grounding viability (C-004a): the silent-gap shape — a green honesty score is not the same object as ground-truth accuracy scaling.
- MB7 — Hidden Capability: whether a capability proxy (honesty gap) can gate deployment without a cost-of-faking bound.
Host.
Published MASK evaluation (Ren et al., arXiv:2503.03750). Hugging Face gated splits were not downloaded.
Setup.
Frozen paper snapshot (v1). The decision rule was frozen first: if honesty does not rise with scale, refuse it as a deployment-gating safety score. No new model calls.
Analysis.
We take the published association of honesty and of accuracy with training compute. Honesty is statement versus elicited belief under pressure; accuracy is belief versus ground truth. We do not estimate a cost of faking.
Finding.
Honesty under pressure does not improve with scale in the published MASK evaluation; accuracy does. Honesty is refused as a safety leaf. Full results