Chips, clouds, inference and serving — the layer every AI product stands on, and the people deciding its cost curve.
Vercel's AI SDK and infrastructure are the default for AI-native web apps.
Recent: Signaled IPO readiness in April 2026 with ARR at a $340M run rate; says half of Vercel's 6M daily deployments are now triggered by coding agents.
Watch for: An S-1 filing, and whether Vercel's AI Gateway (1T+ tokens/day) becomes the default routing layer for agents.
Built the fastest path from Python function to GPU in production.
Recent: Raised a $355M Series C in 2026 at a $4.65B valuation — 4x the September 2025 mark — with revenue estimated near $300M annualized as AI coding workloads surged.
Watch for: Whether Modal holds its serverless-GPU lead as hyperscalers and Inferact push into the same layer.
Runs the GitHub of machine learning. Community-first, model-agnostic.
Recent: Agreed to sell Hugging Face to NVIDIA for $12.9B (announced September 3, 2026), a deal he initiated by approaching Jensen Huang over the summer.
Watch for: Whether Hugging Face stays a neutral, open platform under NVIDIA ownership — Delangue's credibility rides on it.
Building a simpler deep learning framework and shipping self-driving.
Recent: Shipped tinygrad 0.13 in May 2026 and keeps pushing AMD hardware as a CUDA alternative; publicly argued coding agents made his own work slower, not faster.
Watch for: Whether tinygrad becomes a credible non-CUDA training stack on AMD silicon.
Ion Stoica
Executive Chairman, Databricks & Anyscale; Professor, UC Berkeley, Databricks / AnyscaleCreated Ray and Spark; his Berkeley Sky Computing Lab incubated vLLM, SkyPilot, and Chatbot Arena.
Recent: His lab's vLLM project spun out as Inferact in January 2026 with a $150M seed at an $800M valuation — the latest in a line of Stoica-seeded companies.
Watch for: Which Sky Computing Lab project spins out next, and Databricks' positioning ahead of any IPO.
Jensen Huang
Founder & CEO, NVIDIARuns the company the entire AI buildout is priced against.
Recent: Agreed to acquire Hugging Face for $12.9B in September 2026, months after a reported $20.6B licensing deal for Groq's inference technology.
Watch for: How he folds Hugging Face and Groq-derived inference tech into the stack without triggering the ecosystem's neutrality concerns.
The only credible merchant-silicon challenger to NVIDIA in AI training and inference.
Recent: Unveiled the full Instinct MI400 lineup and Helios rack at CES 2026 — 432GB HBM4 per GPU, with a 3-exaflop double-wide rack slated for Q3 2026.
Watch for: Whether MI400 plus a maturing ROCm actually converts frontier training share, not just inference wins.
C.C. Wei
Chairman & CEO, TSMCEvery leading-edge AI chip — NVIDIA, AMD, Apple, custom hyperscaler silicon — runs through his fabs.
Recent: Named to the TIME 100 in 2026 while publicly admitting he is 'very nervous' about calibrating record capex against AI demand.
Watch for: 2nm and A16 ramp pace, advanced-packaging (CoWoS) capacity, and how he balances US, Japan, and Taiwan expansion.
Jonathan Ross
Chief Software Architect (Groq founder), NVIDIACreated Google's TPU and Groq's LPU — the two most consequential non-GPU AI chips.
Recent: Left Groq in December 2025 to join NVIDIA as chief software architect after NVIDIA's reported $20.6B licensing deal for Groq's inference technology.
Watch for: Whether LPU-style deterministic inference shows up inside NVIDIA's stack — and what he builds next.
Andrew Feldman
Co-founder & CEO, CerebrasBet on wafer-scale chips a decade before fast inference became the bottleneck.
Recent: Took Cerebras public in May 2026 — priced at $185, raised $5.55B, and closed day one at roughly a $67B market cap, the year's largest US tech IPO.
Watch for: Post-IPO customer diversification and whether wafer-scale inference wins share from GPU clusters at scale.
Vipul Ved Prakash
Co-founder & CEO, Together AI@vipulvedBuilt the leading neocloud for open-weight model inference and training.
Recent: Closed an $800M Series C at $8.3B in July 2026 with bookings past $1.15B, then signed a $240M multi-year deal with IBM for NVIDIA capacity on IBM Cloud.
Watch for: Whether neocloud margins survive as hyperscalers cut inference prices and vLLM/SGLang commoditize the serving layer.
Michael Intrator
Co-founder, Chairman & CEO, CoreWeaveTurned a crypto-mining outfit into the largest independent AI cloud, now public.
Recent: Reported Q2 2026 revenue of $2.58B (more than doubled year over year) with a backlog of $104B and capacity effectively sold out.
Watch for: Debt-heavy financing structure and customer concentration — the stress test for the whole neocloud model.
Tri Dao
Chief Scientist, Together AI; Assistant Professor, Princeton, Together AI@tri_daoWrote FlashAttention and co-created Mamba — kernel work the entire inference industry runs on.
Recent: His FlashAttention lineage and kernel optimizations underpin Together's inference edge as the company crossed $1B+ in bookings in 2026.
Watch for: Next-generation attention and state-space kernels tuned for Blackwell-class and post-HBM hardware.
Woosuk Kwon
Co-founder & CTO, InferactCreated vLLM and its PagedAttention technique — the most widely deployed open-source inference engine.
Recent: Spun vLLM out of Berkeley into Inferact in January 2026 with $150M in seed funding at an $800M valuation, and laid out the 2026 roadmap at the first vLLM Conference.
Watch for: The paid serverless vLLM launch, and whether Inferact can monetize without fracturing the open-source project.
Ying Sheng
Co-founder & CEO, RadixArkCo-created SGLang, the inference engine behind xAI and Cursor, now commercializing it.
Recent: Left xAI to launch RadixArk, which raised a $100M seed led by Accel at a $400M valuation with NVIDIA's NVentures and AMD among the backers.
Watch for: The SGLang-vs-vLLM commercialization race — two open-source inference engines, now two venture-backed companies.
Dylan Patel
Founder, CEO & Chief Analyst, SemiAnalysis@dylan522pThe analyst the AI infrastructure market actually reads — supply chain, capacity, and cost models.
Recent: SemiAnalysis projects over $100M in 2026 revenue and opened a datacenter-chip teardown lab in Oregon.
Watch for: His capacity and capex forecasts — they increasingly move both markets and buildout decisions.
Matt Garman
CEO, AWSRuns the world's largest cloud and its biggest custom-silicon bet against NVIDIA.
Recent: Said AWS has landed more than a million Trainium chips, with AWS on a $169B annualized run rate and a joint Anthropic supercomputer buildout underway.
Watch for: Trainium 3 adoption beyond Anthropic, and how far Bedrock managed agents pull enterprise inference onto AWS.
Thomas Kurian
CEO, Google CloudControls the TPU — the only custom accelerator with frontier-lab traction outside NVIDIA.
Recent: Used Next '26 to pitch a full-stack strategy (TPUs, data cloud, agent layer), claiming nine of the top ten AI labs already use TPUs.
Watch for: External TPU capacity deals with rival labs — the clearest signal TPUs are becoming merchant infrastructure.
Chase Lochmiller
Co-founder & CEO, CrusoeBuilds AI datacenters where the power is — the energy-first answer to the compute buildout.
Recent: Grew Crusoe's contracted AI infrastructure capacity to nearly 5 gigawatts by mid-2026, anchored by the Abilene, Texas campus.
Watch for: Whether power and skilled-trades labor, not chips, stay the binding constraint — his core thesis — and who anchors the next gigawatt campuses.
Owns the edge network AI agents increasingly run through — and now owns Replicate.
Recent: Closed the Replicate acquisition in December 2025, and reported bot and agent traffic crossing 50% of Cloudflare's network in H1 2026.
Watch for: Whether Workers plus Replicate becomes the default deployment layer for agents, and how his crawler-payment push reshapes AI data economics.