Models

Tencent launches Hy ASR 3.0 speech recognition

Tencent has launched the preview of Hy ASR 3.0 on Tencent Cloud, integrating its 295B MoE language model to deliver context-aware speech recognition with low word error rates.

AlphaSignal4 Aug 2026Models
Illustration generated for this story

Tencent has launched the preview of its Hy ASR 3.0 speech recognition system on Tencent Cloud, offering API access targeted at customer service, meeting transcription, and voice search. Unlike traditional systems that transcribe audio strictly verbatim, this new iteration connects the speech pipeline directly to the semantic capabilities of Tencent's latest large language model, Hy3. This integration allows the system to perform context-aware transcriptions, effectively shifting the technology from simple acoustic translation to deep semantic comprehension.

The architecture of Hy ASR 3.0 relies on a Mixture of Experts design, utilizing a proprietary unsupervised speech encoder trained on tens of millions of hours of unlabeled audio data. It inherits semantic reasoning capabilities from the 295-billion-parameter Hy3 MoE model. On open benchmarks, the system achieves a multilingual word error rate of approximately 3 percent, specifically scoring a 3.34 percent word error rate in Mandarin, 2.62 percent in English, and 3.12 percent in Cantonese. For comparison, NVIDIA's Canary Qwen 2.5B leads the Hugging Face Open ASR Leaderboard with a 5.63 percent word error rate, though that benchmark focuses primarily on English.

For developers and enterprise practitioners, this semantic approach solves several persistent real-world speech-to-text challenges. The system introduces four primary upgrades: improved general accuracy, context-based homophone correction, hotword injection, and enhanced robustness when processing whispered or noisy audio. By leveraging the underlying language model, the system can infer the speaker's actual intent to correct homophones that sound identical but differ in meaning based on context.

The technology is already active in Yuanbao, Tencent's consumer AI assistant, where users can access dialect recognition and context correction features for free. With the preview now available on Tencent Cloud, enterprise developers can begin integrating these context-aware transcription capabilities into their own applications to improve user experience in complex acoustic environments.

This is our own summary of reporting by AlphaSignal

More in Models