arXiv cs.AIPaper
Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing
This targets a real and underappreciated failure mode: agents that run correct code but draw statistically invalid conclusions. Anyone deploying LLM agents for research or data analysis workflows should treat P-Bench as a sanity check before trusting agent-generated p-values in production reports.