Qualcomm and Amazon will co-design inference silicon for AWS
Amazon received a warrant for 25 million Qualcomm shares at $161.26. The two companies will work on custom inference chips and optical connectivity.
Qualcomm and Amazon have agreed to co-design custom inference silicon and optical connectivity for AWS, with Amazon receiving a warrant for 25 million Qualcomm shares at $161.26, according to a filing reported on 8 September.
The warrant is the notable term. Amazon is being paid in Qualcomm equity to work with Qualcomm — an arrangement that says something about which side of this deal needed the other more.
Why inference, specifically
The distinction between training and inference silicon has become the most important one in AI hardware, and it is the reason a company with no presence in AI training is a credible entrant here.
Training is a small number of enormous, long-running jobs where the requirements are raw throughput, memory bandwidth and interconnect. Nvidia dominates it and the software lock-in is severe.
Inference is the opposite: vast numbers of short jobs where cost per token, latency and power efficiency decide everything, and where the software requirement is narrower — serve a fixed model efficiently rather than support arbitrary research code. The workload is now much larger in aggregate than training, and it is growing faster, because every deployed product runs inference continuously while training happens occasionally.
Qualcomm's entire history is efficient inference at low power in mobile silicon. That is a more relevant competence for this problem than it would have been for training.
The optical connectivity is the less obvious half
As accelerator counts rise, the constraint moves from compute to moving data between chips. Electrical interconnect runs into power and distance limits, and at cluster scale a meaningful share of energy goes into communication rather than computation.
Optical interconnect addresses that, and it is one of the few areas where a new entrant can offer something Nvidia's NVLink does not simply do better. It is also relevant to the energy question: Vals AI estimated this month that long agentic tasks may carry 10,000 times the footprint of a simple query, and data movement is a substantial part of why.
Every large buyer is now building its own
This is the same move the whole industry is making, and Nvidia's actions this fortnight are best read as a response to it.
OpenAI, Google, Amazon and Anthropic all have internal accelerator programmes. Meta entered a $100 billion multi-year agreement with AMD for up to six gigawatts. Gimlet Labs raised $300 million at a $3 billion valuation — up from $400 million six months earlier — for software that routes AI workloads across different chips, backed by Arm and Microsoft's M12.
Nvidia, meanwhile, agreed to buy Hugging Face for $12.9 billion with a commitment to keep the hub open to competing chips, put about $2 billion into Nscale, and discussed $2.5 billion into Thinking Machines. Its reported logic for the Hugging Face deal includes keeping open-weight models anchored to its hardware.
Nvidia posted $96.2 billion in quarterly revenue with $89.0 billion from data centre, and forecast roughly 70 percent growth for fiscal 2028. The position is not under immediate threat. The number of well-capitalised parties working to end it has never been higher.
Neither company disclosed a timeline, volume commitment or which AWS services the silicon will serve.
Runs the newsroom. Rename this profile in the studio to your own byline.
Related
Every weekday, the AI stories that moved money or shipped code.
No cross-posting, unsubscribe anytime. See all newsletters