en.Wedoany.com Reported - On August 11, IBM signed a multi-year agreement valued at $240 million with AI cloud platform Together AI, under which IBM will build a large-scale AI inference cluster based on NVIDIA HGX B300 systems in the United States to provide cloud computing power for Together AI's open model inference services. The cluster is scheduled to become operational in the first quarter of 2027.

The cluster will be deployed on IBM Cloud, with the initial phase expected to be equipped with approximately 2,000 NVIDIA Blackwell Ultra B300 chips and will adopt the NVIDIA Spectrum-X Ethernet platform to connect compute nodes. IBM stated that this is the first compute cluster on its cloud platform built specifically for large-scale AI inference workloads, utilizing HGX B300 and Spectrum-X technologies.
The HGX B300 platform is used by server manufacturers to build high-density AI systems. Each HGX B300 configuration features 8 Blackwell Ultra GPUs with a total GPU memory capacity of 2.1TB, interconnected via fifth-generation NVLink, achieving a total GPU-to-GPU interconnect bandwidth of 14.4TB/s and network bandwidth of 1.6TB/s. Based on approximately 2,000 chips, the initial deployment scale is equivalent to roughly 250 eight-GPU HGX B300 systems, though the actual number of servers and rack configurations will be subject to the final engineering plan.
Unlike compute clusters primarily used for model training, this project focuses on model inference—that is, using already-trained models to process user requests and generate text, code, or other content. Together AI will run open models such as DeepSeek, MiniMax, and Kimi on this cluster to provide enterprise customers with model invocation and production-grade inference services.
Kai Mak, Chief Revenue Officer of Together AI, stated that the company expects the relevant compute capacity to be fully reserved two to three months before the official launch. Together AI currently processes approximately 400 trillion tokens per month. As enterprise customers expand their deployment of open models, the company is increasing GPU resources through self-built infrastructure, leasing, and long-term cloud compute agreements.
In July 2026, Together AI completed a $800 million Series C funding round, achieving a post-investment valuation of $8.3 billion. The round was led by Prosperity7 Ventures, a subsidiary of Saudi Aramco, with participation from NVIDIA, Vista Equity Partners, General Catalyst, and other institutions. The company plans to use the raised funds to expand its products and compute infrastructure, aiming to increase its overall infrastructure scale by approximately 50 times over the next five years.
This agreement will also expand IBM Cloud's Blackwell Ultra compute deployment. IBM had previously announced the availability of NVIDIA Blackwell Ultra GPUs on its cloud platform for large model training, inference, and complex reasoning workloads. Once the new cluster goes live, Together AI will become the primary user of this dedicated compute capacity. Subsequent milestones for both parties include server deployment, Spectrum-X network construction, software platform integration, cluster debugging, and commercial operation in the first quarter of 2027.





















