ArtificialIntelligence.io

Today

The AI Signal

8 October 2026

arXiv cs.CLPaper

From Expert-Guided Proof Search to Automated Open-Problem Solving

This is a genuine research contribution. Bolzano doesn't just solve problems; it contributes novel results that humans in the field recognize. Four solutions to open questions in theoretical computer science, confirmed by paper authors. For builders, this proves LLM agents can do proof search reliably enough for real work. For investors, this is evidence that the agent layer is mature enough for specialized reasoning tasks. The open-source release matters too.

arXiv cs.AIPaperClaude Watch

AgentTime: Can Agents Estimate and Control Their Own Runtime?

Agent reliability is moving from capability to predictability. If you're shipping agents in production, runtime control is quickly becoming table stakes. Claude's current performance here is a known gap, and builders should test their own models on this benchmark before committing to long-running workflows. This matters more as agents move from prototypes to systems people depend on.

arXiv cs.CLPaper

Training Advisors for LLM Agents from Task Outcomes

This is practically useful. Training a 4B critic that generalizes to larger models and different architectures, with 25+ point improvements on MuSiQue, shows that agent feedback can be factored into a reusable module. For teams building agents, this suggests an efficient path to debugging and iterating on reasoning without touching your base model. Production-ready approach.

Also worth your time

The daily signal, in your inbox.

Coming soon. In the meantime, the Tuesday Brief is free.

Get the free brief