The interesting number is 10 million users for Codex, which suggests coding agents have crossed from early-adopter tool into mainstream developer habit faster than most expected. The laundry list of ChatGPT Work features, Sites, Subagents, Finance, no-code, reads like OpenAI trying to become the default work OS rather than just a model provider. Anyone building vertical agent products should watch whether OpenAI's horizontal bundle cannibalizes their niche.
Field reports from real domain deployments are more useful than benchmark papers because they show where agents actually save time versus where they create new debugging overhead. Genomics and scientific computing are good stress tests since the codebases are old, messy, and full of domain-specific correctness requirements. Worth reading if you're evaluating coding agents for technical, non-web-app codebases.
Robotics remains the place where model capability meets hard physical constraints, so any claimed leap in whole-body coordination deserves scrutiny for real-world demo versus lab conditions. Google continues to push Gemini beyond chat and code into embodied systems, which matters for anyone tracking where multimodal models end up deployed physically. Watch for third-party hands-on tests before treating this as a capability shift.