Nvidia put the Groq 3 LPX into full production at 3,431 tokens per second
The figure is on a 100K-context benchmark, nearly four times faster than comparable public endpoints. Long-context throughput is what agentic work actually needs.
Nvidia has put the Groq 3 LPX into full production, reporting 3,431 output tokens per second on a 100,000-token context benchmark — nearly four times faster than comparable public endpoints, according to the company. Nebius is among the partners named at launch.
The specific benchmark matters more than the headline rate. Throughput at 100K context is a different and much harder measurement than throughput on short prompts, and it is the one that corresponds to what agentic systems do.
Why long-context throughput is the number
Attention cost grows faster than linearly with context length, so a system that is fast on a 2,000-token prompt can be slow on a 100,000-token one. Vendors have historically quoted the flattering figure.
Agentic workloads live at the unflattering end. An agent working through a long task accumulates context — prior steps, tool outputs, file contents — and every subsequent call carries all of it. OpenAI's Astra coordinates 16 agents. Alibaba's refreshed Qwen3.8-Max offers a one-million-token context window. Anthropic cut cached-context costs by 75 percent with Fable 5.1 precisely because the same large context is read repeatedly across a long run.
So a fourfold improvement at 100K context is aimed at the workload that is growing, rather than at the benchmark that looks best.
It also bears on cost. Vals AI estimated this month that long agentic tasks can carry roughly 10,000 times the energy footprint of a simple query, with some web app builds matching 2.5 hours of household electricity. Throughput improvements at long context attack that directly.
Inference as a separate product line
The LPX line reflects the same split that is drawing new entrants into AI silicon: training and inference are different problems, and inference is now the larger aggregate workload.
Qualcomm and Amazon agreed this month to co-design custom inference silicon and optical connectivity for AWS, with Amazon taking a warrant for 25 million Qualcomm shares. Meta entered a $100 billion agreement with AMD for up to six gigawatts. Gimlet Labs raised $300 million at a $3 billion valuation, backed by Arm and Microsoft's M12, for software that routes workloads across different chips.
Nvidia's response has been to compete on inference directly while making it harder to leave. It is buying Hugging Face for $12.9 billion — with reported logic that includes keeping open-weight models anchored to its hardware — while committing to keep the hub open to competing chips.
The context this shipped into
Production ramp was announced the same week that Ukraine attributed a fatal autonomous drone strike to a Russian Molniya drone using an Nvidia Jetson Orin module for target selection, killing three civilians on 6 July. Nvidia confirmed the photographed module.
Edge inference hardware is dual-use in a way training clusters are not: it is small, cheap, widely available and designed to run autonomously without a network. The same properties that make it good for robotics and industrial vision make it good for a weapon.
Nvidia's second-quarter results put edge computing revenue at $7.2 billion, up 27 percent. The company has not published figures on how much of that reaches conflict zones, and it is not clear it could.
Runs the newsroom. Rename this profile in the studio to your own byline.
Related
Every weekday, the AI stories that moved money or shipped code.
No cross-posting, unsubscribe anytime. See all newsletters