Models

NVIDIA Adapts Nemotron 3.5 to Saudi Arabic Dialects

NVIDIA has adapted its Nemotron 3.5 speech model to Saudi Arabic dialects, demonstrating a fine-tuning workflow that slashes local error rates while preserving general language skills.

NVIDIA Developer Blog12 hrs agoModels
Image: NVIDIA Developer Blog

NVIDIA developers have demonstrated a workflow to adapt the NVIDIA Nemotron 3.5 ASR model for regional dialects, specifically targeting Saudi Najdi and Hijazi. Using the NVIDIA NeMo framework, the team curated a 133.7-hour corpus of Saudi speech, retaining 103,559 out of 125,490 utterances. To prevent the model from forgetting its existing multilingual capabilities, they implemented a weighted replay mix consisting of 90 percent Saudi speech and 10 percent FLEURS data, which was split into 7 percent English and 3 percent Arabic. Training was performed over 12,000 steps on two NVIDIA RTX PRO 6000 Blackwell Workstation Edition GPUs.

This targeted training dramatically improved transcription quality. The word error rate on the Najdi and Hijazi test split dropped from 55.05 percent to 29.96 percent, while the character error rate fell from 31.63 percent to 12.18 percent. On the full SADA test set, the word error rate decreased from 58.84 percent to 35.61 percent. Crucially, the model did not lose its grasp of other languages; English performance on FLEURS actually improved from 11.04 percent to 10.42 percent word error rate, and FLEURS Arabic improved from 12.67 percent to 11.41 percent.

For practitioners, the workflow highlights important trade-offs between computational cost and accuracy. Nemotron 3.5 features a 24-layer encoder. While updating all 24 layers yielded the best results, freezing some layers can save memory and compute. For instance, updating only the top eight layers left 230.4 million parameters trainable and 407.6 million frozen, resulting in a word error rate of 32.32 percent. Updating only the top six layers resulted in a 33.42 percent word error rate.

Practitioners can also boost accuracy at inference time without retraining. By expanding the attention context to 13 lookahead frames and using beam-8 MALSD decoding, the team cut the word error rate by an additional 2.71 absolute points to 27.25 percent. This adjustment adds about 800 milliseconds of latency, making it ideal for batch workloads like call archives but less suitable for live captioning. Additionally, the newly released NVIDIA Nemotron 3 Diarization model can be integrated to provide speaker-attributed transcription for up to eight speakers.

This is our own summary of reporting by NVIDIA Developer Blog

More in Models