ArtificialIntelligence.io

Video

The channels worth your time.

Watch here, or jump to the channel. The take underneath each one is ours.

Matthew Berman

DeepSeek Fails the Rubik’s Cube Test

DeepSeek's agent performance is still flaky on spatial reasoning tasks. If you're evaluating DeepSeek for agent workflows, this is a concrete data point to run your own tests on rather than assume it handles physical simulation or complex multi-step spatial problems. Tool-use doesn't mean reasoning.

Matthew Berman

Why Hyperagent is serious

Without the video itself, this reads as mid-tier commentary on an emerging agent tool. Berman's an influential voice in the builder community, so if he's flagging Hyperagent as serious, it's worth a look if you're building multi-step workflows. Context would tell us whether this is a framework innovation or just good marketing.

Matthew Berman

Anthropic went CRAZY (Mythos/Fable 5.1)

The title is hype, but if there's a real Fable 5.1 release with material improvements, builders need to know. We can't score this properly without the full story. Go to item 5 for actual substance instead of enthusiasm.

Matthew Berman

I've had early access to Astra... it's INSANE

This is a YouTuber impression, not a technical assessment. Berman has an audience that values speed-to-opinion, so this will drive early adoption discourse. But 'INSANE' tells you nothing about where Astra actually wins. Use this to know what builders will try first, not what they should try.

Matthew Berman

xAI actually did it... (Grok 4.6)

Third-party reaction videos are a weak signal on their own, but a Grok release landing days after other frontier updates keeps the pressure on the model layer's pricing and benchmark race. Worth a skim for capability claims, but wait for independent evals before shifting any production workload toward Grok.