arXiv cs.CLPaper
ScienceArena: Benchmarking LLMs on Latest Scientific Olympiad Competitions
This is the benchmark that should ship with every frontier model evals report. It catches real failures: visual grounding, problem decomposition, maintaining global context across multi-step reasoning. For builders using LLMs on scientific workflows, this is the test suite to steal from. For researchers, this closes a gap that data contamination has made urgent.