Pihlakas, 2025 — Systematic runaway-optimiser-like {LLM} failure modes on biologically and economically aligned {AI} safety benchmarks
Long-horizon BioBlue benchmarks: language models often revert from homeostasis to unbounded single-objective maximization.
Publication links
Long-horizon BioBlue benchmarks: language models often revert from homeostasis to unbounded single-objective maximization.