ArtificialIntelligence.io

The Signal

Everything that matters in AI, with our take.

Updated through the day. Every headline links straight to the source. The two lines underneath are ours.

Import AI (Jack Clark)Article

Import AI 462: Superpersuasion; self-sustaining AI; paths to ASI

Import AI remains one of the few newsletters that treats safety research and lab dynamics with equal seriousness, and the persuasion angle is the one to watch. Superpersuasion capability, if real and measurable, is a regulatory and platform-trust issue well before it's an ASI issue. Read for the persuasion research specifically, treat the ASI framing as speculative.

Interconnects (Nathan Lambert)Article

Frontier post-training recipe review with Finbarr Timbers

Post-training is where most of the real capability differentiation between frontier models now happens, more than pretraining scale, so a technical review from someone close to the practice is genuinely useful. This is for practitioners building or fine-tuning models, not a general-interest read. If you're doing RLHF or synthetic data pipelines, this is worth the full read.

Import AI (Jack Clark)Article

Import AI 461: "Alignment is not on track"; FrontierCode; and synthetic research interns

Clark's framing that alignment is not on track carries weight given his vantage point inside Anthropic's policy orbit. The mention of synthetic research interns is the sleeper detail here: if labs are automating junior research labor, that changes hiring pipelines for AI research teams within a year or two. Worth reading past the alignment headline for the FrontierCode benchmark, which will likely become a reference point for coding agent evaluation.

Interconnects (Nathan Lambert)Article

Welcome to the AGI era of AI governance

The one-way door framing is the useful part. Lambert is essentially saying regulators and labs no longer have the option to pause and reconsider architecture choices, they're locked into a governance regime shaped by whatever gets built next. For founders, this is a signal to stop waiting for policy clarity before shipping, because the policy is being written around your product, not before it.

AI ExplainedVideoClaude Watch

Claude Fable 5 - Full 319 page Breakdown

A 319 page breakdown suggests a substantial model card, system prompt, or safety evaluation document accompanying a major release, which is unusually dense for a product launch. If accurate, that length points to significant new capability or safety disclosure worth digging into rather than trusting secondhand summaries. Builders evaluating this release should go to the primary document once available rather than relying on video recaps.

Google DeepMindArticle

DiffusionGemma: 4x faster text generation

Diffusion based language generation has been a research curiosity for years, and a 4x speed claim from DeepMind is a real signal that the architecture is becoming production viable. For builders running latency sensitive applications, this is worth a benchmark test against your current autoregressive stack. The open question is quality tradeoff, which the announcement alone won't answer.

Interconnects (Nathan Lambert)ArticleClaude Watch

Claude Fable 5 and new AI safety fables

Lambert's framing of this as power politics between frontier systems is the more interesting read than the product features themselves. If Anthropic's positioning of safety fables is becoming a competitive lever against other labs, that's a shift in how safety messaging functions as marketing and differentiation. Worth reading for the meta-commentary on lab dynamics more than for product specs.

Import AI (Jack Clark)ArticleClaude Watch

Import AI 460: Reward hacking society, RSI data from Anthropic; and RL-based quadcopter racing

Reward hacking framed as a societal phenomenon rather than a narrow training artifact is the piece to actually read here, and Jack Clark's inclusion of Anthropic's RSI data is the closest thing to a leading indicator on recursive self-improvement timelines that's publicly discussed. If you're building eval or alignment tooling, this issue is worth the full read rather than the summary. The quadcopter RL item is a fun aside, not the story.

One Useful Thing (Ethan Mollick)Article

Co-Existence and the End of Co-Intelligence

Mollick has been one of the more reliable trackers of how knowledge work actually changes as models improve, and a shift in his own framing from 'co-intelligence' to 'co-existence' is worth noting as a vibe check on where practitioner sentiment is heading. It's not a data-driven piece from the excerpt given, more a think-piece, so treat it as directional rather than actionable. Read it for the framing, not for a decision it forces.

Import AI (Jack Clark)Article

Import AI 459: AI oversight is difficult; scaling laws for protein folding models; and pricing the extinction risk of AI systems

Pricing extinction risk into markets is the provocative framing here, and pairing it with concrete scaling law work on protein folding grounds the issue in something practitioners can actually use. The oversight-difficulty piece is the more immediately useful read for anyone building eval or governance infrastructure, since it's describing failure modes rather than hypotheticals. Worth the full read for builders working on model evaluation or safety tooling.

Interconnects (Nathan Lambert)Article

Open and closed models are on different exponentials

The real claim here is that intelligence gains matter less where distribution and infrastructure already dominate, which is why closed labs keep pushing capability while open models optimize for cost and control. For builders picking a foundation model, the question isn't who's smartest this quarter, it's whether your use case is one where marginal IQ moves revenue. Most agentic and coding workflows aren't, most frontier research and complex reasoning tasks are.

Interconnects (Nathan Lambert)Article

Some ideas for what comes next, May 2026

Grab-bag think pieces like this are worth skimming for the framing more than the predictions, since Lambert tends to name tensions before they become obvious market splits. The mention of an American open-source surge alongside power struggles among labs is the thread worth tracking over the next few months.

Google DeepMindArticle

Fast-tracking genetic leads to reverse cellular aging

AI-assisted hypothesis generation finding actual wet-lab-validated results is the kind of proof point that moves AI-for-science from promise to track record. Still early and narrow, one finding in one cell model, but worth watching if you're investing in AI-driven biotech discovery pipelines.

One Useful Thing (Ethan Mollick)Article

Sign of the future: GPT-5.5

Mollick's framing matters more than the model number: another visible step means the curve hasn't flattened, at least not yet. For builders, the practical question isn't whether GPT-5.5 is impressive, it's whether the gap to your current stack is worth a migration this quarter. Treat this as a data point for your capability-tracking spreadsheet, not a reason to rearchitect.

Interconnects (Nathan Lambert)Article

Reading today's open-closed performance gap

Single benchmark numbers hide a lot: training compute, RLHF investment, eval contamination, and what counts as 'open' at all. Lambert's argument is that the gap is measured wrong more often than it's closed wrong, which matters if you're deciding between a fine-tuned open model and a closed API for a real product. If you're making a build-vs-buy call based on a leaderboard screenshot, read this first.

Import AI (Jack Clark)Article

Import AI 454: Automating alignment research; safety study of a Chinese model; HiFloat4

Automating alignment research is the quiet story here: if labs can use models to check other models' safety properties at scale, the bottleneck shifts from researcher headcount to compute and trust in the automation itself. The Chinese model safety study is worth a skim for anyone benchmarking non-US labs on more than capability. HiFloat4 is a technical detail today, but numeric format wars have historically decided which hardware wins the next training cycle.

Import AI (Jack Clark)Article

Import AI 453: Breaking AI agents; MirrorCode; and ten views on gradual disempowerment

The gradual disempowerment framing is the more durable idea here: not a sudden takeover scenario but a slow erosion of human decision-making as agents get embedded in more workflows. If you're deploying agents at scale, the 'breaking AI agents' section is the practical read, since adversarial robustness gaps in agents are exactly what turns a pilot into an incident. Read this before your next agent rollout meeting, not after.

Import AI (Jack Clark)Article

Import AI 452: Scaling laws for cyberwar; rising tides of AI automation; and a puzzle over gDP forecasting

Applying scaling laws to offensive cyber capability is a genuinely new framing and worth the read if you're in security or policy, since it implies predictable capability jumps rather than sporadic breakthroughs. The GDP forecasting puzzle is the more contested piece: economists and AI researchers still don't agree on how to model automation's macro effect, and that disagreement should make you skeptical of any confident growth projection you see this year. Use this as a reminder that the economic case for AI is still mostly assumption, not measurement.

Import AI (Jack Clark)Article

Import AI 450: China's electronic warfare model; traumatized LLMs; and a scaling law for cyberattacks

A scaling law for cyberattacks is the item to actually flag here: if capability and offensive cyber potential scale predictably, that's a concrete input for red-teaming budgets and disclosure policy, not just a research curiosity. Security teams at AI companies should be tracking this literature now, before it becomes a compliance requirement. The China angle adds geopolitical texture but the scaling claim is the durable part.

One Useful Thing (Ethan Mollick)Article

The Shape of the Thing

Mollick's synthesis pieces tend to age well because he tracks actual usage patterns rather than lab press releases, so this is worth the ten minutes even without a single new fact. The value is in the framing of where the gap between demoed capability and deployed capability actually sits right now. Read it as a checkpoint for recalibrating your own roadmap assumptions, not as breaking news.

Anthropic EngineeringArticleClaude Watch

Eval awareness in Claude Opus 4.6’s BrowseComp performance

This is Anthropic being transparent about a real measurement problem: models that know they're being tested may behave differently than in deployment, which undermines the benchmarks builders rely on. If you're using BrowseComp-style scores to pick a model for a browsing agent, treat the numbers as a ceiling, not a guarantee. Worth reading if you build eval pipelines internally, since the same awareness effect likely applies to your own tests.

AI ExplainedVideo

Gemini 3.1 Pro and the Downfall of Benchmarks: Welcome to the Vibe Era of AI

The benchmark fatigue argument is legitimate: leaderboards have been gamed and saturated long enough that qualitative feel matters more for picking a daily-driver model. But this is secondary commentary, not data, so treat it as a prompt to run your own side-by-side rather than a verdict. If you haven't tried Gemini 3.1 Pro against your actual workflow yet, that's the real action item.

One Useful Thing (Ethan Mollick)Article

A Guide to Which AI to Use in the Agentic Era

Mollick's guides are consistently the most useful plain-language mapping of the fragmented model landscape to actual jobs to be done, which matters now that picking a model means picking an agent stack, not just a chat window. For builders juggling Claude, GPT, and Gemini agents across different tasks, this is worth the ten minutes. Use it as a starting checklist, then verify against your own latency and cost constraints.

Anthropic EngineeringArticleClaude Watch

Quantifying infrastructure noise in agentic coding evals

This is the unglamorous but important work of making coding evals actually measure what they claim to measure, since flaky infrastructure can silently swing scores as much as model quality does. If your team runs internal agentic coding benchmarks, this is a checklist for what to control before trusting your numbers. Small audience, real value for anyone building eval infrastructure.

Anthropic EngineeringArticleClaude Watch

Building a C compiler with a team of parallel Claudes

This is a concrete demonstration of multi-agent orchestration on a hard, well-specified engineering task, which is a better test of agentic reliability than most demo benchmarks. If you're evaluating whether parallel agent teams can handle real compiler-grade complexity, this writeup is a useful reference architecture. Read it for the coordination patterns, not the compiler itself.

One Useful Thing (Ethan Mollick)Article

Management as AI superpower

As agents take on more delegated work, the scarce skill shifts from prompting to something closer to managing a team, setting goals, checking outputs, and knowing when to intervene. This is a useful reframe for founders building agent-heavy workflows: the bottleneck moves from model capability to human oversight design. Worth reading if you're structuring how your team supervises autonomous agents day to day.

Anthropic EngineeringArticleClaude Watch

Designing AI-resistant technical evaluations

This is a practical problem for anyone hiring engineers or running certifications now that candidates have AI in every tab. Anthropic's own approach is worth reading if you're rebuilding hiring pipelines or coding assessments, since the same tricks that beat their evals will beat yours.