The platform is aimed at companies managing multiple AI models and includes an inference engine optimized for private deployments using Huawei Ascend hardware.

iFlytek on July 19 launched Spark Token Factory, an enterprise platform designed to provide a unified access layer for AI models, route requests among models, support governance and optimize token costs.

The company positions the product as a middle layer between enterprise applications and large language models. It is intended for deployments spanning coding assistants, customer service, knowledge retrieval, content generation and document processing, where companies may need to connect to multiple model services.

According to iFlytek, Spark Token Factory combines unified model access, intelligent routing, token-cost optimization and security and compliance governance. The company said the platform is designed to create a closed loop covering access, governance, observability and operations, giving enterprises a centralized entry point for managing model services.

iFlytek said its routing system classifies requests into three levels, L1 through L3, using factors including prompt length, predefined and customized rules, number of dialogue turns and session affinity. The platform then matches requests to model resources based on quality, cost, latency, availability and security level.

The approach is intended to send simpler tasks, such as text classification and information extraction, to more cost-effective models, while assigning long-document analysis and complex reasoning to higher-capability models. iFlytek said the average routing-decision latency can be kept below 100 milliseconds.

For private deployments using domestic hardware, Spark Token Factory also includes an inference engine optimized for Huawei Ascend hardware. iFlytek said it has conducted end-to-end engineering optimization for mainstream open-source large models, covering areas from underlying operators to upper-layer scheduling.

The company said its optimization work includes model compression and precision techniques intended to reduce memory use while maintaining controllable accuracy. It also cited work on fusing core inference operators and reorganizing multi-card communication scheduling to improve computing efficiency.

The announcement addresses a growing operational challenge for enterprises moving from isolated AI pilots to multi-business, multi-scenario and multi-team deployments. As organizations use more models, iFlytek argues they need stronger tools for model management, cost control, system stability and security governance.

iFlytek did not disclose pricing, customer deployments, supported model coverage or independent performance testing.