arXiv cs.AIPaper
Vulnerability Localization Benchmark: Measuring Agentic Security Analysis at Repository Scale
This is the right evaluation for agentic security tools. Localization is harder and more practical than detection or repair, and 500 real vulnerabilities across six ecosystems is solid coverage. The benchmark will likely become standard. Use it to test whether your agent framework can actually navigate and reason over real codebases, not toy examples.