Now Reading
NVIDIA Boosts Saudi Arabic Speech Recognition

NVIDIA Boosts Saudi Arabic Speech Recognition

NVIDIA Arabic speech recognition AI

NVIDIA has fine-tuned its Nemotron 3.5 ASR model using Saudi Arabia’s SADA dataset. The work significantly improves recognition of Najdi and Hijazi Arabic dialects.

According to NVIDIA, the fine-tuned model cut the word error rate for the targeted dialects from 55.05% to 29.96%.

SADA, or the Saudi Audio Dataset for Arabic, is provided by the Saudi Data and AI Authority (SDAIA). SDAIA developed the dataset with the Saudi Broadcasting Authority to support Arabic speech and language research.

SADA Dataset Improves Dialect Recognition

NVIDIA used 133.7 hours of Najdi and Hijazi audio during fine-tuning. The company used NVIDIA NeMo to adapt the multilingual speech recognition model while retaining its broader capabilities.

As a result, the targeted test set showed a major improvement in transcription accuracy. Word error rate fell from 55.05% to 29.96%, while character error rate dropped from 31.63% to 12.18%.

Moreover, the full SADA test set improved substantially. Word error rate declined from 58.84% to 35.61%, while character error rate fell from 35.40% to 15.97%.

The model also maintained its performance outside the target dialects. For example, English word error rate improved from 11.04% to 10.42%. Arabic performance also improved on the FLEURS evaluation set.

NVIDIA completed the reported training run in about 4.5 hours using two GPUs. The experiment involved 12,000 training steps.

Nemotron 3.5 Targets Real-Time Voice Applications

Nemotron 3.5 ASR supports streaming transcription across 40 language locales. In addition, the model includes automatic language detection and supports low-latency speech recognition.

The model uses a Cache-Aware FastConformer-RNNT architecture. This approach reuses previously processed context, reducing redundant computation during streaming transcription.

Consequently, the improved Saudi Arabic model could support several voice-based applications. These include virtual assistants, conversational systems, call-centre transcription and live captioning.

NVIDIA also tested different decoding configurations after fine-tuning. With wider context and beam-search decoding, the reported word error rate fell further to 27.25% on the Najdi and Hijazi test split. However, the wider context increased latency.

See Also
Accenture Anthropic AI safety partnership

The results show why regional speech data matters for multilingual AI. A model can perform well across languages while still struggling with local dialects, accents and recording conditions.

Saudi Data Supports Localised AI

The SADA dataset contains roughly 667 hours of transcribed audio. More than 600 hours came from 57 Saudi television programs and series covering more than 10 Saudi dialects.

The dataset includes more than 125,000 categorised audio clips. SDAIA released SADA through Kaggle to support researchers and developers working on Arabic speech technologies.

Therefore, the NVIDIA experiment demonstrates how national datasets can improve global AI models for regional use. It also provides a workflow that developers can adapt to other languages and dialects.

The broader significance extends beyond transcription accuracy. Better dialect recognition can improve voice assistants, customer service systems, media archives, and searchable audio content.

For Saudi Arabia, the development also strengthens the case for locally representative training data. Meanwhile, NVIDIA’s results suggest that targeted fine-tuning can improve regional performance without rebuilding a model from scratch.

View Comments (0)

Leave a Reply

Your email address will not be published.

© 2024 The Technology Express. All Rights Reserved.