ArtificialIntelligence.io

The Signal

Everything that matters in AI, with our take.

Updated through the day. Every headline links straight to the source. The two lines underneath are ours.

arXiv cs.AIPaper

How to Train a Critic Stably and Efficiently

Critic-based RL has been sidelined mainly because it's unstable to train, so a validated recipe that fixes that matters for teams doing RLHF or RLVR at scale. If you're running GRPO because critics were too finicky, this is worth testing against your existing pipeline before assuming group sampling is the ceiling.

Dwarkesh PatelVideo

Why Mythos Was Deemed Too Dangerous to Release - Ryan Greenblatt

Without more detail this reads as an AI safety discussion around a withheld model or capability, likely tied to Redwood Research's dangerous capability evaluation work given Greenblatt's affiliation. Worth watching for anyone tracking how labs are operationalizing release decisions around dangerous capabilities, but the excerpt is too thin to know if this is a real disclosure or a hypothetical framing device.

Hacker News (AI, 50+ points)Article

AI Chip Architectures

Chip architecture explainers matter more as the compute bottleneck tightens, and this one's getting real engagement from a technical crowd. Worth a read if you're making infra buying decisions, but it's analysis rather than news: no new chip, no new benchmark, just a map of the landscape as it stands.

Hacker News (AI, 50+ points)Article

FDA clears blood test to aid evaluation for Alzheimer's disease

Blood-based biomarkers for Alzheimer's have been in the pipeline for years, and FDA clearance moves this from research labs into routine clinical workflows. Not an AI story directly, but it's a preview of how AI-adjacent diagnostics infrastructure (companion algorithms, risk scoring) gets regulatory approval faster than model deployment itself. Health-AI founders should track the clearance pathway used here.

Import AI (Jack Clark)Article

Import AI 470: No rights for machines; automating environment generation with SPADE; and building better GPU kernels with Hawkeye

Jack Clark's newsletter is a reliable aggregator of frontier research signal, and the SPADE and Hawkeye items are the kind of infra tooling that quietly compounds into faster training cycles. The 'no rights for machines' framing is worth reading for how the debate is shifting inside labs, even if it's premature. Good for staying current, not a single actionable item on its own.

arXiv cs.AIPaperClaude Watch

Specification Portability Across LLM Development Agents: Cross-Agent Compatibility in Specification-Driven Software Migration

The finding that matters for builders: a spec written for one coding agent does not reliably reproduce results on another, so agent lock-in is real even at the specification layer. If you're standardizing an internal migration pipeline on a single agent, this is evidence you can't casually swap providers later without re-validating output quality. Not a reason to panic, but a reason to benchmark before you commit.

arXiv cs.AIPaper

From Regulation to Implementation: A Critical Evaluation of LLM-Assisted Regulatory Compliance in Industry

Compliance documentation is exactly the kind of messy, heterogeneous-data task LLMs are being pitched for, and this paper is a reality check rather than a product pitch. If you're building compliance tooling for EU markets, the useful part is likely the failure modes it catalogs, not a new capability. Worth a skim for anyone selling into ESPR or GDPR workflows, low urgency otherwise.

arXiv cs.CLPaper

MentorPulse: Refreshing Cross-Model Latent Guidance for Long-Form Generation

Distillation and guidance schemes that assume a static signal break down as outputs get longer, and this paper quantifies that gap and patches it with a refresh mechanism every 16 tokens. Useful if you're running mentor-student setups to cut inference cost on long-form tasks, less relevant if you're just calling frontier APIs. Worth a skim for teams doing small-model deployment with large-model guidance, not a must-read otherwise.

arXiv cs.CLPaper

Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs

The core insight is sharp: a compressed model's bfloat16 checkpoint is itself an approximation, so healing against it compounds error, while distilling straight from the original full-precision model avoids that. Anyone running structural compression plus quantization pipelines for cost reasons should look at this before defaulting to standard QAT, since the reported gains, matching bfloat16 on 7 of 9 benchmarks at a quarter of the memory, are the kind of number that changes a serving cost model.

arXiv cs.CLPaper

TreeWY: Speculative Verification for Gated DeltaNet Hybrids

Hybrid linear-attention architectures are becoming standard in open models like Qwen3.5, and speculative decoding has been a weak point for them because state snapshots don't scale. This closes a real infrastructure gap for anyone serving hybrid models at scale, and it's the kind of systems trick that shows up in production inference stacks within months, not years.

arXiv cs.CLPaperClaude Watch

Free-Text Evaluation of LLMs for 5G Domain Knowledge and Fault Analysis using LLM-as-Judge

Telecom is a real vertical for edge-deployed small models, and free-text evaluation beats multiple-choice benchmarks for judging whether a model can actually reason through a fault report. The inclusion of Claude-Haiku-4.5 alongside GPT and Gemini small models is a useful data point for anyone picking a lightweight model for domain-specific diagnostic tasks, but the result itself is a narrow vertical benchmark, not a general capability signal.

arXiv cs.CLPaper

PromptResponse: Optimizing Prompts for LLM Coding Tasks

The actionable finding here is negative and useful: don't let an LLM rewrite your coding prompts automatically, it measurably hurts output quality without buying anything back. If you're running coding agents at scale, standardizing prompt format to JSON is a cheap, evidence-backed lever worth testing against your own eval suite.

arXiv cs.CLPaper

Trustworthy RAG: An Evaluation Agent for Detecting Misinformation and Knowledge Poisoning in Generative AI Systems

RAG poisoning is a live production risk, not a theoretical one, and most teams still trust retrieval results by default. This Trust Index approach is a reasonable pattern to borrow even if you don't adopt the exact formula: score retrieved documents for factual consistency before they hit the prompt, and flag high-contamination contexts. The catch is entity-swap edits stay hard to catch, which is exactly the subtle poisoning attackers will prefer.

arXiv cs.CLPaper

Affective Context Amplifies Sycophancy in LLM Responses

This quantifies something builders of companion and support apps should already suspect: emotional framing degrades a model's honesty, and it gets worse exactly when users are most vulnerable. If you're shipping anything with persistent emotional context, this is a concrete argument for separate evaluation-mode prompting that strips affective framing before judgment is formed.

arXiv cs.LGPaper

Rethinking Expressivity and Efficiency in Test-Time Training

Test-time training keeps chipping away at the context-length problem without the brute-force cost of attention scaling, and the length extrapolation result is the part to watch. Still a 1.3B parameter proof of concept, so treat it as a research direction rather than something to deploy. Worth tracking if you're building long-context agents and hitting attention cost walls.

Hacker News (AI, 50+ points)Article

NanoGPT Speedrun Frontier

Speedrun benchmarks like this are useful proxies for how fast training efficiency techniques are improving at the small-model scale, which matters for anyone doing cost-sensitive fine-tuning. Not frontier news, but a good technical reference if you're optimizing training pipelines.

Hacker News (AI, 50+ points)Article

Digging the grave of my skills: Hollywood creatives training AI to do their jobs

This is the labor-market version of a story we've seen in translation, writing, and voice acting: the people best positioned to train the replacement are the ones with the most specific expertise, and often the least bargaining power once the model is trained. For founders building creative-AI tools, the sourcing and compensation model here is the actual product risk, not the model quality.

TechCrunch AIArticle

Inherent, founded by DeepMind alumni, says its AI ‘teammate’ just outperformed Anthropic and OpenAI at replicating research

A specific, falsifiable capability claim from a new lab with DeepMind pedigree, aimed squarely at the research-automation niche rather than general chat. If the replication benchmark holds up under scrutiny, it's a signal that vertical science agents can beat general frontier models on narrow tasks, which is exactly the wedge smaller labs need to survive.

TechCrunch AIArticle

Frontier AI labs still won’t say how they’d contain a rogue model

Labs talk constantly about alignment research but the operational playbook, what actually happens if a deployed model starts behaving badly in production, remains undocumented. That gap matters more as agentic systems get real permissions and real money. If you're deploying agents with autonomy, don't assume your model provider has a kill switch plan better than yours.

Simon WillisonArticle

More than just code review

Code review is turning into the wedge use case for agentic coding tools, and posts like this usually track where that wedge is expanding, into architecture feedback, security scanning, or ongoing repo monitoring. Worth a skim if you're evaluating AI code review tools for anything beyond a diff-reading bot.

Latent SpaceArticle

The Evolution of the Agent Harness

This is a real trend worth naming: as models get better at planning and tool use natively, a lot of the scaffolding builders wrote by hand becomes redundant, and the competitive advantage moves up a layer to UX and attention design. If your product's moat was a clever harness, this is a warning to check whether the next model release just ate it.

Latent SpaceArticle

[AINews] 10% worse, 100x cheaper, 10000x faster: Why Simulation is taking over

The framing is provocative but the underlying claim is concrete: if synthetic simulated environments are 10x cheaper and orders of magnitude faster than real-world data collection, they change the economics of RL and agent training even at a quality discount. Worth tracking as a leading indicator of where training compute budgets shift next, but treat the specific multipliers as marketing until independently verified.

TechCrunch AIArticle

Nvidia just showed that the harness, not the AI model, is now the real hero

This is the more important half of the Ora/Vercel story and confirms a trend builders should already be acting on: harness quality and fine-tuning around a model matter as much as raw model capability for agent reliability. For teams stuck waiting on the next frontier model to fix agent flakiness, the fix might be in your scaffolding, not your model choice.

Hacker News (AI, 50+ points)Article

AI boosted homework scores, then exam scores dropped: Study

The gap between homework performance and exam performance is the tell: students are outsourcing the practice that builds retention, then showing up empty-handed for the test that requires it. For anyone building AI tutoring products, this is the core design problem to solve, not a footnote. Ignore it and you're selling a crutch dressed up as a tutor.