When Models Defer to Wrong Answers: A Robustness Audit of Source-Attributed Cues in Multiple-Choice QA
Language models are dangerously suggestible to false expert signals. This matters if you're deploying models in contexts where someone might slip a malicious attribution into the prompt. It's a failure mode to test for, but it's not a new class of weakness. Add this to your robustness audit checklist.