arXiv cs.AIPaper
BLOOM-WILT: Logit Tilting for Behaviour Elicitation in Automated LLM Auditing
The method is clever: use the model's own distribution to find edge cases that testing usually misses. For teams running audits on deployed models, this reduces the cost of finding problems that only surface at scale. The logit-tilting trick is neat but the real value is having a systematic way to hunt for rare behaviors without retraining.