arXiv cs.CLPaper
Moral Competence Before Moral Content: Why LLM Agents Lack the Prerequisites for Coherent Alignment
This paper measures something real: whether an LLM's moral outputs form a coherent policy or just pattern-match to prompts. The result is that frontier models fail this test. If you're deploying AI in high-stakes domains where consistency matters, this is evidence that current models are not reliable proxies for stable principles. The methodology is clever but the bar is necessarily high.