ArtificialIntelligence.io

The Signal

Everything that matters in AI, with our take.

Updated through the day. Every headline links straight to the source. The two lines underneath are ours.

arXiv cs.CLPaper

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring

The real story here is that SWE-bench Verified, the benchmark half the industry cites for coding agent claims, has a nearly 60% flawed-test rate on its unsolved instances and leaks gold patches into training data. Anyone benchmarking or marketing against SWE-bench numbers should treat them with more skepticism starting now. ProMax's refactoring focus is a better proxy for real engineering work than single-file bug fixes, so expect it to get adopted by labs wanting a cleaner leaderboard story.

TechCrunch AIArticle

Meta launches Muse Code, an AI agent for large code bases

Meta entering agentic coding directly competes with Cursor, Devin, and OpenAI's Codex-based tools rather than just shipping another chat assistant. The pitch on large codebase handling is the hard problem every coding agent still struggles with, so the real test is whether Muse Code's context and retrieval actually outperform incumbents on messy enterprise repos, not greenfield demos. Worth a trial run against your actual codebase before switching tooling.

Hacker News (AI, 50+ points)Article

Software development with AI is starting to feel like cooking steak

High engagement on Hacker News signals this touched a nerve about the gap between AI coding demos and the judgment required to use the tools well in practice. The steak metaphor is catchy but the underlying claim, that AI coding tools reward experienced judgment more than they replace it, is now a familiar refrain rather than new evidence. Read the comment thread if you want a temperature check on developer sentiment, not for new information.