ArtificialIntelligence.io

The Signal

Everything that matters in AI, with our take.

Updated through the day. Every headline links straight to the source. The two lines underneath are ours.

arXiv cs.AIPaper

Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models

This is a serious indictment of current evals: if your molecular model is just memorizing published data, you don't have a molecular model. The authors find verbatim retrieval is widespread and worsens under chain-of-thought reasoning, which is counterintuitive and alarming. For biotech founders using LLM evals to validate molecular property prediction, this means your benchmark scores are likely garbage. If you're a lab reporting that frontier models excel at molecular reasoning, you need to re-run your evals controlling for contamination. This undermines an entire category of claimed capability.