arXiv cs.LGPaper
SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?
This bridges two important gaps: interpretability research usually happens offline, and agent research rarely touches safety auditing. The benchmark tests whether agents can reliably use SAE tools to discover features matching expert references. If frontier agents can do this work autonomously, it changes the scalability story for mechanistic monitoring, which matters for anyone shipping agents at scale.