en.Wedoany.com Reported - On July 19, iFLYTEK officially launched the Spark Token Factory, aimed at helping enterprises improve the efficiency and management of large model applications, and build a future-oriented large model infrastructure foundation.
Enterprises face a series of operational management challenges during the large-scale deployment of AI. The variety of models continues to increase, calling costs keep rising, and system stability requirements are heightened, while capabilities in cross-model management, cost control, security governance, and operational monitoring lag behind. Enterprises often need to simultaneously access multiple model services to support different business scenarios such as code assistants, intelligent customer service, knowledge Q&A, content generation, and document processing. The complexity of different tasks varies significantly, but in practice, a unified model is often used for processing, leading to inefficient resource utilization and continuously growing Token costs. Running multiple models in parallel also imposes higher demands on system stability, security governance, and operational management. Enterprise AI construction is gradually transitioning from "model capability building" to "AI infrastructure building," where the focus is no longer solely on the capabilities of the models themselves, but on how to truly integrate model capabilities into business systems to achieve stable operation, unified governance, and continuous optimization.
The Spark Token Factory is positioned as an enterprise-level AI model intelligent routing platform, providing core capabilities such as unified model access, intelligent routing and scheduling, Token cost optimization, and security compliance governance. It forms a complete closed loop covering "access—governance—observation—operations," offering enterprises a unified large model service entry point to achieve manageable, controllable, and traceable large model resources.

The intelligent routing function comprehensively determines task complexity based on multi-dimensional features such as Prompt length, preset rules, custom rules, conversation rounds, and session affinity, implementing intelligent L1 to L3 hierarchical management for requests. The platform matches suitable model resources by considering five factors: quality, cost, latency, availability, and security level, allowing simple tasks to call more cost-effective models and complex tasks to prioritize high-capability models. The average decision latency can be controlled within 100 milliseconds. Through intelligent routing capabilities, the platform effectively reduces the call pressure on large models, unlocking more cost optimization space for enterprises.

In localized private deployment scenarios, the Spark Token Factory performs end-to-end engineering optimization for mainstream open-source large models based on the characteristics of Ascend hardware. On the memory side, various model compression and precision optimization techniques are used to reduce memory usage while maintaining controllable precision. On the computation and communication side, core computation operators are fused, and multi-card communication is scheduled and reorganized to achieve parallel overlap of computation and communication. On the cache and scheduling side, cross-node distributed KV Cache management and multi-level storage scheduling are employed, introducing context-aware cache routing and context reuse reorganization strategies. On the decoding side, a parallel decoding acceleration mechanism tailored for Ascend hardware is introduced. Measured data shows that under the same hardware conditions, inference efficiency is improved by approximately 5 times compared to the open-source vLLM-Ascend framework, and the first Token latency is reduced by 30%–40%.

In terms of security, the Spark Token Factory adopts a three-power separation architecture: management does not touch data, operations do not touch permissions, and auditing does not touch business—each is independent and mutually constraining. The platform automatically generates immutable, full-chain audit records for all key operations, covering six major categories: data access, permission changes, model calls, and data exports, meeting information security level protection requirements. Sensitive information protection covers personal privacy and business data, with hierarchical routing to ensure sensitive data does not leave its domain.

The Spark Token Factory provides end-to-end full-chain log auditing capabilities, recording the entire process from when a request enters the platform to when the final result is returned. Through log tracking, routing decision records, Token consumption statistics, and cost attribution analysis, enterprises can grasp the operational status and resource consumption of each model call. Combined with budget management, ROI analysis, and visual reports, it provides data support for subsequent operational optimization and resource planning. The launch of the Spark Token Factory further enhances iFLYTEK's capability layout in the enterprise AI infrastructure field and marks a new stage where enterprise large model applications are transitioning from point-based exploration to large-scale operations.










