Your AI inference costs are about to get a challenger.
Positron AI closed an $875M Series C at a $5B valuation today.
The play: memory-first inference silicon that doesn't need HBM or CoWoS.
While Nvidia's Rubin Ultra roadmap just scaled back from 1TB HBM4E to 192GB because supply doesn't exist — Positron built around the constraint entirely.
Their Asimov chip uses commodity LPDDR5X memory.
288GB to 2,304GB per chip.
Tapes out on TSMC N3P at end of 2026.
Production in H2 2027.
The Titan system packs 4-8 Asimov chips into a single node.
Serves models beyond 16 trillion parameters.
Context windows beyond 10 million tokens.
No HBM dependency. No CoWoS bottleneck.
And this isn't vaporware.
50+ racks of Atlas already deployed at Oracle Cloud Infrastructure.
Jump Trading and Parasail are production customers.
Investors: NEA, Atreudes, SemiAnalysis Capital's Dylan Patel, Jim Clark.
Strategic money from Hudson River Trading, Cisco, Naver.
Here's what this means for your budget.
Inference is eating AI spend.
Every agent, copilot, and assistant runs on it.
And the constraint isn't compute — it's memory bandwidth and power.
If Positron delivers tokens per dollar at scale, your CFO will ask why you're paying Nvidia premiums.
Audit your inference infrastructure contracts.
If you're locked into HBM-dependent silicon, you're paying for a supply chain bottleneck that Positron just sidestepped.
The inference cost war just started. Your procurement team should be paying attention.
Positron AI just raised $875M to kill your Nvidia inference bill.
AI-Assisted Content — Produced with AI assistance and human editorial review.
Learn more
0 Comments