en.Wedoany.com Reported - At the WAIC 2026 conference, China's PPIO officially launched its Agentic Cloud positioning and a new intelligent model gateway product, proposing the creation of an "Intelligent Token Factory" for the Agent era. As the global AI agent market is expected to grow from $5.1 billion in 2024 to $47.1 billion in 2030, the demand for Tokens in Agent tasks is surging sharply.

Yao Xin, co-founder and CEO of PPIO, pointed out that Token customers are shifting from humans to Agents, and cloud service providers need to offer "Token factories" with lower latency and higher throughput. According to data from China Insights Consultancy, although inference costs are declining by more than 10 times annually, the number of Tokens consumed per single Agent task has grown from hundreds in the early days to tens of thousands or even hundreds of thousands, causing overall cost pressure to rise rather than fall. Additionally, a single model cannot cover the diverse needs of Agents, and Agents have much higher requirements for latency and stability than traditional AI applications.
The core product launched by PPIO, the Intelligent Model Gateway, adopts a "Mixture of Models" (MoM) mechanism. In critical Agent steps, the gateway distributes problems to multiple expert models, generating answers through cross-validation to avoid task failures caused by model specialization biases. The gateway also features intelligent scheduling and cost optimization capabilities, routing tasks based on type and controlling costs through mechanisms such as compressing redundant context and budget guardrails. According to the DRACO Deep Research benchmark, by integrating Mimo-V2.5-Pro, Kimi-K2.7, and GLM-5.2 for hybrid inference, PPIO can boost performance to near the level of Claude Fable 5, improving intelligence by 20% and reducing costs by 50%-60%. PPIO's self-developed inference acceleration engine is deeply optimized for Agent scenarios, achieving up to a 10x improvement in inference performance. Combined with Prompt Cache technology, Token costs can be reduced by up to 80%. PPIO has been selected as one of the first batch of "Enterprise-Level Token Service Performance Climbing Baselines" by the China Academy of Information and Communications Technology, achieving TPS ≥ 55/second, TTFT ≤ 0.9 seconds, and a call success rate ≥ 99.9% in general scenarios.

In terms of the Agent runtime environment, PPIO has launched the Agent Sandbox, which achieves millisecond-level cold starts (<200ms) based on microVM technology. Each task runs in an independent virtual machine environment, supporting the simultaneous creation of tens of thousands of sandboxes. The Auto Pause/Resume mechanism reduces costs by up to 90% or more compared to similar products. Since its launch at last year's WAIC, this first Agent Sandbox in China compatible with the E2B interface has seen business scale grow by over 123 times. The PPIO Agent Harness integrates capabilities such as Browser Use, Computer Use, Code Interpreter, and MCP Server, covering the complete chain from information acquisition to execution for Agents. The platform comes pre-loaded with multiple agent templates like PPClaw and PPHermes, allowing developers to complete cloud deployment in as little as 10 minutes, while also supporting 7×24 continuous hosting and scheduled task scheduling.

In terms of data, in June 2026, PPIO's average daily Token calls exceeded 1.2 trillion, ranking first among independent AI cloud computing service providers in China, representing an over 8-fold increase compared to the same period in 2025. The company's revenue grew from 358.4 million yuan in 2023 to 770.3 million yuan in 2025, with over 670,000 registered developers globally. PPIO's average GPU utilization rate remained above 75% in 2025, far exceeding the industry average of 40%-50%. Yao Xin explained that PPIO leverages the "peak shaving and valley filling" capability of its global scheduling network, utilizing the time difference between the Eastern and Western Hemispheres to push GPU utilization to 70%-80%. Currently, PPIO has deployed distributed multi-clusters in regions including Europe, North America, South America, and Southeast Asia, covering six continents globally.

Yao Xin stated that the ceiling of intelligence exists not only within models but also in Agent engineering and model fusion beyond the model itself, which can significantly enhance intelligence levels. PPIO is committed to making Tokens cheaper and Agent operations more stable, laying the groundwork for the underlying infrastructure in anticipation of the explosion of the agent economy.

Based on the Harness layer, the platform comes pre-loaded with out-of-the-box agent templates such as PPClaw and PPHermes. Developers can deploy them to the cloud with a single click via the console, going live in as little as 10 minutes. After deployment, it supports 7×24 continuous hosting and scheduled task scheduling. When not in use, billing can be paused with one click, and it can be resumed in seconds. The platform also provides AI-native interfaces for Agents, supporting the MCP standard protocol, allowing Agents to call infrastructure such as GPU computing power and sandbox environments through natural language. The platform is compatible with mainstream Agent frameworks like LangChain, CrewAI, and AutoGen.

PPIO's distributed computing network and high-density computing nodes provide underlying support for the Intelligent Token Factory. By leveraging cross-regional time zone differences and the staggered demands of diverse customer groups, it dynamically allocates global idle computing power to demand peaks, ensuring effective utilization of GPU resources. This scheduling capability directly reduces the actual cost per unit Token, serving as the underlying guarantee for PPIO's "front shop, back factory" model.










