Zhou, 2026 — When Preferences Fail to Become Incentives: A Utility-Behavior Gap in Large Language Models
Shows coherent stated preferences often fail to motivate behavior in realistic LLM tasks.
Publication links
Shows coherent stated preferences often fail to motivate behavior in realistic LLM tasks.