Nabrah STT V1
STT-V1
في هذه الصفحة
انتقل إلى قسم
- Nabrah STT V1
- Model overview
- Available variants
- Default
- EOT
- Compare variants
- Recommended use cases
- Default
- EOT
- Languages and code-switching
- Real-time streaming
- Streaming audio requirements
- Performance
- Transcription accuracy
- Streaming latency
- End-of-turn detection
- End-of-utterance control
- Priority words
- Text normalization
- Punctuation
- Async file transcription
- Capabilities
- Using Nabrah...
- Default model
- EOT variant
Nabrah STT V1
Stable
Nabrah STT V1 is a real-time speech-to-text model designed for streaming speech recognition.
It supports Arabic, English, and code-switching between Arabic and English within the same conversation.
Nabrah STT V1 is available in two variants:
Default: optimized for higher transcription accuracy.
EOT: optimized for conversational applications that require end-of-turn detection.
Model overview
Specification | Value |
|---|---|
Model family | Nabrah STT V1 |
Version | V1 |
Status | Stable |
Release date |
|
Primary use case | Real-time streaming transcription |
Languages | Arabic, English |
Code-switching | Arabic-English supported |
Streaming | Supported |
Streaming latency | As low as 160 ms |
Streaming sample rate | 16 kHz |
Streaming encoding | LINEAR_PCM |
Channels | Mono |
Priority words | Supported |
Available variants
Default
The default Nabrah STT V1 model provides the highest transcription accuracy within the current Nabrah STT model family.
It is recommended when accurate transcription is more important than built-in end-of-turn detection.
The default model is selected automatically when no alternative recognition_model is specified.
EOT
eot_nabrah
The EOT variant is designed for real-time conversational AI and voice agent applications.
It emits a special <eot> token when it detects that the speaker has reached the end of their conversational turn.
Compared with the default model, the EOT variant provides lower transcription accuracy but adds native end-of-turn signaling and applies Inverse Text Normalization (ITN).
Compare variants
Capability | Default | EOT |
|---|---|---|
Transcription accuracy | Higher | Lower |
Real-time streaming | Yes | Yes |
Arabic | Yes | Yes |
English | Yes | Yes |
Arabic-English code-switching | Yes | Yes |
End-of-turn detection | No | Yes |
Text normalization | Yes | Yes |
Automatic punctuation | No | Yes |
Recommended for | General real-time transcription | Voice agents and conversational AI |
Recommended use cases
Default
Use the default model for applications such as:
Real-time transcription
Customer support calls
Live speech recognition
Voice applications where transcription accuracy is the priority
Arabic-English conversations
Applications that use external turn-detection logic
EOT
Use the EOT variant for applications such as:
Voice agents
Conversational AI
AI assistants
Real-time interactive voice systems
Applications that require native end-of-turn detection
Languages and code-switching
Nabrah STT V1 supports:
Arabic
English
Arabic-English code-switching
Code-switching allows the model to recognize Arabic and English within the same sentence or conversation.
Example:
على الإيميل Report أرسل لي ال
Real-time streaming
Nabrah STT V1 is primarily designed for real-time streaming speech recognition.
Streaming is available over WebSocket:
WS /api/ext/stt/ws
Audio is streamed as raw PCM binary frames.
Streaming audio requirements
Specification | Value |
|---|---|
Encoding | LINEAR_PCM |
Sample rate | 16 kHz |
Channels | Mono |
Transport | WebSocket |
Streaming latency | As low as 160 ms |
Performance
Transcription accuracy
The default Nabrah STT V1 model achieved:
WER: 31.84%
Benchmark: SADA
Environment: Cloud-based streaming use case
For more info about the SADA click here
35.19% for eot_nabrah model
Streaming latency
Streaming latency can be as low as:
160 ms
Actual latency may vary depending on infrastructure, network conditions, audio input, application configuration, and the first token it self.
End-of-turn detection
The EOT variant can detect when the speaker has completed their conversational turn.
Select it using:
eot_nabrah
When an end-of-turn is detected, the model emits:
<eot>
Example transcription:
```text \ \
ابي احجز موعد بكرة
```Applications can use the <eot> token as a signal to trigger the next stage of the conversational pipeline, such as sending the transcript to an LLM.
The token is returned inline with the transcript and should be removed by the client if it should not be displayed to the end user.
End-of-utterance control
For streaming transcription, the end_of_utterance_silence_ms setting controls how much trailing silence is required before the current utterance is treated as complete.
Example:
```json \
{ \
"end_of_utterance_silence_ms": 500 \
}
```Lower values can make conversational responses faster, while higher values allow speakers to pause for longer without ending the utterance.
This setting affects both preliminary and final transcription behavior.
Priority words
Nabrah STT V1 supports recognition biasing using priority words and phrases.
This is useful for improving recognition of:
- Brand names
- Product names
- Industry terminology
- Names
- Domain-specific vocabulary
Example:
```json \
{ \
"priority_words": [ \
"نبرة", \
"ذكاء اصطناعي", \
"خدمة العملاء" \
], \
"priority_words_strength": 0.5 \
}
```The priority_words_strength parameter controls how strongly the recognizer favors the supplied vocabulary.
Text normalization
The default model applies stronger Arabic orthographic normalization to some characters in the returned transcript.
Examples include normalization of some Alef forms and converting ة to ه.
The EOT variant applies lighter normalization and preserves the original text form more closely.
For conversational applications where the transcript is passed directly to an LLM, this normalization generally preserves the intended meaning.
Punctuation
The default Nabrah STT V1 model does not currently generate punctuation automatically.
Variant | Automatic punctuation |
|---|---|
Default | Not supported |
EOT | Supported |
Async file transcription
Nabrah also supports asynchronous transcription for longer audio files.
Maximum audio duration: 2 hours
File transcription is available through:
POST /api/ext/stt/async-transcribe
The transcription runs in the background and moves through the following states:
queued → in_progress → completed
Completed results include:
- Transcribed text
- Audio duration
- Timestamps
For file-based transcription, WAV is recommended.
Capabilities
Capability | Support |
|---|---|
Real-time streaming | Yes |
Arabic | Yes |
English | Yes |
Arabic-English code-switching | Yes |
Priority words | Yes |
End-of-turn detection | EOT variant |
Automatic punctuation | No on default |
Async file transcription | Yes |
Timestamps | Yes |
Speaker diarization | No |
Word-level timestamps | Yes |
Language detection | No |
Noise robustness | Yes |
Using Nabrah STT V1
Default model
The default model is selected when recognition_model is omitted or left empty.
```json \
{ \
"api_key": "YOUR_API_KEY", \
"recognition_model": "", \
"end_of_utterance_silence_ms": 500 \
}
```EOT variant
To enable native end-of-turn detection:
```json \
{ \
"api_key": "YOUR_API_KEY", \
"recognition_model": "eot_nabrah", \
"end_of_utterance_silence_ms": 500 \
}
```