Saudi — Najdi (ar-najdi), Hejazi (ar-hijazi), Qassimi (ar-qassimi)
Speed Sweet Spot
0.9 – 1.2x
Sherine (Arabic — Egyptian)
Property
Value
Language
Arabic
Gender
Female
Tone
Warm, conversational
Best For
Egyptian-market conversational agents, customer service
Dialect
Egyptian — select via ar-egyptian
Speed Sweet Spot
0.9 – 1.2x
Myriam (Arabic — Levantine)
Property
Value
Language
Arabic
Gender
Female
Tone
Warm, conversational
Best For
Levantine-market conversational agents, customer service
Dialect
Levantine — select via ar-levantine
Speed Sweet Spot
0.9 – 1.2x
Yara_en (English)
Property
Value
Language
English
Gender
Female
Tone
Professional, neutral
Best For
English prompts, bilingual IVR, international audiences
Accent
Neutral English
Speed Sweet Spot
0.9 – 1.3x
Voice Selection Guide
Scenario
Recommended Voice
Why
Banking IVR (Saudi)
Yara (ar-najdi / ar-hijazi / ar-qassimi)
Clear, professional Saudi voice for structured prompts
Egyptian-market bot
Sherine (ar-egyptian)
Native Egyptian dialect, warm and conversational
Levantine-market bot
Myriam (ar-levantine)
Native Levantine dialect
Bilingual system (AR/EN)
Yara + Yara_en
Same voice family, consistent brand experience
English-only service
Yara_en
Professional neutral English
Speech Recognition Engine
SM-STT-V1
Neural ASR engine providing high-accuracy Arabic transcription with real-time streaming support.
Property
Value
Engine
SM-STT-V1
Primary Language
Arabic
Secondary Language
English
Accuracy
High accuracy on clean audio (contact sales for benchmarks)
Max Audio Duration
300 seconds (5 min) per REST request
Max File Size
100 MB
Supported Formats
FLAC, MP3, WAV, OGG, WebM
Optimal Sample Rate
16 kHz mono
REST Endpoint
https://api.withsm.ai/v1/asr/audio/transcriptions
Diarization Endpoint
https://api.withsm.ai/v1/asr/transcribe
WebSocket Endpoint
wss://api.withsm.ai/v1/asr/stream
gRPC Endpoint
api.withsm.ai:9101
Authentication
X-API-Key header (required)
ASR Capabilities
Feature
Supported
Notes
Arabic transcription
✅
Najdi, Hejazi, Qassimi, Egyptian, Levantine
English transcription
✅
Optimized as secondary language
Mixed Arabic/English
✅
Auto-detects language switches
Streaming (real-time)
✅
gRPC bidirectional streaming
Word timestamps
✅
Via gRPC response fields
Confidence scores
✅
Per-utterance confidence (0.0 – 1.0)
Speaker diarization
✅
Available via /v1/asr/transcribe endpoint
Punctuation
✅
Automatic punctuation insertion
Number formatting
✅
Spoken numbers converted to digits
ASR Accuracy Tips
Factor
Impact
Recommendation
Audio quality
High
Use 16kHz+ sample rate, minimize background noise
Speaker distance
High
Microphone within 30cm of speaker
Audio format
Medium
FLAC or WAV preferred over lossy formats
Segment length
Medium
5-30 seconds per segment for best accuracy
Speaking pace
Low
Normal speaking pace (120-160 words/min Arabic)
Diacritics context
Low
Model infers diacritics from context
Language Support
Language Support Matrix
Text-to-Speech
Language
Voice(s)
Quality
Dialect
Arabic
Yara, Sherine, Myriam
⭐⭐⭐⭐⭐
Najdi, Hejazi, Qassimi, Egyptian, Levantine
English
Yara_en
⭐⭐⭐⭐
Neutral
Speech Recognition
Language
Quality
Notes
Arabic (Najdi)
✅ Full support
ar-najdi
Arabic (Hejazi)
✅ Full support
ar-hijazi
Arabic (Qassimi)
✅ Full support
ar-qassimi
Arabic (Egyptian)
✅ Full support
ar-egyptian
Arabic (Levantine)
✅ Full support
Syrian, Lebanese, Jordanian — ar-levantine
English
✅ Full support
en
Mixed AR/EN
✅ Full support
Code-switching handled
Arabic Dialect Details
TTS Dialect Behavior
Select the dialect via the dialect parameter or the X-TTS-Dialect header. Yara covers the Saudi dialects — Najdi (ar-najdi), Hejazi (ar-hijazi), Qassimi (ar-qassimi); Sherine covers Egyptian (ar-egyptian); Myriam covers Levantine (ar-levantine):
Voice × dialect support:
Voice
ar-najdi
ar-hijazi
ar-qassimi
ar-egyptian
ar-levantine
Yara
✓
✓
✓
✗
✗
Sherine
✗
✗
✗
✓
✗
Myriam
✗
✗
✗
✗
✓
Requesting a dialect a voice doesn't cover falls back to the default voice.
Feature
Behavior
ق (Qaf)
Pronounced as /q/ (MSA standard)
ج (Jeem)
Pronounced as /dʒ/ (standard)
ث (Tha)
Pronounced as /θ/ (standard)
ذ (Dhal)
Pronounced as /ð/ (standard)
Numbers
Arabic-style reading (e.g., خمسة وعشرون not عشرون وخمسة)
Date format
Day-Month-Year convention
Currency
Saudi Riyal (ريال) recognized by default
ASR Dialect Handling
The ASR engine is trained on multi-dialect Arabic data and automatically adapts to the speaker's dialect (Najdi, Hejazi, Qassimi, Egyptian, Levantine) without explicit configuration. Transcription output preserves dialectal spelling — e.g. Najdi "وش لونك" is transcribed as spoken, not converted to MSA.
Dialect handling (ASR): You never select a dialect for transcription — on REST, WebSocket, or gRPC. Pass ar (or omit for auto-detect) and the model detects the spoken dialect automatically (Najdi, Hejazi, Qassimi, Egyptian, Levantine). The dialect codes (ar-najdi … ar-levantine) are used only for TTS voice/dialect selection.
Mixed Language (Code-Switching)
SM-AI-MODELS handles Arabic/English code-switching — common in Gulf business communication.
TTS Code-Switching
Code
{ "input": "يرجى إرسال الـ report إلى قسم الـ HR قبل نهاية اليوم", "voice": "Yara"}
The engine automatically:
Detects English words within Arabic text
Switches pronunciation model for English segments
Maintains natural prosody across language boundaries
Tips for best code-switching results:
Pattern
Quality
Example
Arabic sentence with English terms
⭐⭐⭐⭐⭐
"أرسل الـ email الآن"
Full sentence switch
⭐⭐⭐⭐
"شكراً. Thank you for calling."
Word-level alternation
⭐⭐⭐
Complex mixing may reduce naturalness
English sentence with Arabic name
⭐⭐⭐⭐
Use Yara_en for primarily English content
ASR Code-Switching
The ASR engine detects language switches at the word level:
Code
// Input audio: "أبغى أحجز appointment يوم الخميس"// Output:{ "text": "أبغى أحجز appointment يوم الخميس", "language": "ar"}
Character Sets
Supported Arabic Characters
Range
Characters
Support
Arabic letters
ا ب ت ث ج ح خ د ذ ر ز س ش ص ض ط ظ ع غ ف ق ك ل م ن ه و ي
✅
Hamza variants
أ إ آ ء ئ ؤ
✅
Diacritics
َ ِ ُ ّ ْ ً ٍ ٌ
✅
Punctuation
، ؛ ؟ ! .
✅
Arabic numerals
٠ ١ ٢ ٣ ٤ ٥ ٦ ٧ ٨ ٩
✅
Western numerals
0 1 2 3 4 5 6 7 8 9
✅
Tatweel
ـ (kashida)
✅ Ignored
Language Detection (ASR)
The ASR engine automatically detects the spoken language and returns it in the response:
Code
{ "text": "مرحباً بكم في يونيكود", "language": "ar"}
Code
{ "text": "Welcome to Unicode", "language": "en"}
For mixed-language audio, the language field reflects the dominant language.
Hardware Requirements
SM-AI-MODELS is self-hosted and requires GPU acceleration for optimal performance.
Hardware Requirements
GPU acceleration is required. Contact your account manager for detailed hardware sizing based on your expected workload.