arXiv cs.LGPaper
Can LLMs Discover Scientific Laws in Real and Parallel Worlds?
The benchmark itself is the contribution here, and it's solid. Published targets pose a contamination risk; real data sidesteps that. This is the right way to measure whether LLMs can do scientific reasoning, not just regurgitate it. If you're building AI-for-science tooling, this benchmark is how you'll soon be judged. Study the evaluation protocol.