This is a real benchmark score on a published test, which matters more than marketing claims. 92.8 on Terminal-Bench 2.1 is a credible signal that software engineering agents are getting more reliable. If you're evaluating agent models for code generation, this is now data you can't ignore, but benchmark gaming is also getting sophisticated, so validate in your own codebase before betting the pipeline on it.
When someone of Tao's stature weighs in on AI limitations, it carries weight. The title suggests a systematic problem, not a bug, which matters for anyone building math-dependent agents or tools. The low comment count means the post itself is probably dense and requires reading, but it's worth the time if mathematical correctness is part of your stack.
Cognition's $2 billion raise is the real AI story here. Devin proved that autonomous coding has unit economics worth chasing; now the capital is following. The Boring Company noise and Stoke Space dilute this, but AI tooling is drawing the biggest checks. For founders: the window to raise at pre-scale is closing, speed matters, and agents matter more than models right now.
Tan is making a policy argument that distillation should be treated as fair use, not IP violation. The logic is that if frontier models train on public knowledge, derivatives trained on them should be shareable too. This signals where YC portfolio companies want regulatory cover to go: building on top of the big labs without licensing deals.
This is intellectual property friction, not new, but with fresh institutional weight. The mathematicians have a coherent complaint: models trained on arXiv and textbooks reproduce and sometimes regurgitate their proofs. The labs will likely offer data removal processes and call it solved. Neither side moves much.