NabrahNabrahNabrah
  • التجربة
  • المميزات
  • الأسعار
تسجيل الدخولاتصل بنا
ModelsSTT-V1

بحث في المستندات

ابحث في المستندات حسب العنوان أو التصنيف.

ModelsSTT-V1

بحث في المستندات

ابحث في المستندات حسب العنوان أو التصنيف.

المستندات
Introduction
Core concepts
Quick Start
Nabrah Agents overview
Workspace overview
Calls analytics
WhatsApp analytics
Agents
Custom tools
Analysis groups
Follow-ups
Inbound services
Outbound services
Contacts
WhatsApp
Call logs
Ticketing
Nabrah Studio overview
Text to speech
Speech to text
Voice cloning
Voices
Plugins overview
Web-call widget
Direct link
Account
Workspace & teams
TTS-V1
STT-V1

Nabrah STT V1

STT-V1

10/01/2026

في هذه الصفحة

انتقل إلى قسم
  • Nabrah STT V1
  • Model overview
  • Available variants
  • Default
  • EOT
  • Compare variants
  • Recommended use cases
  • Default
  • EOT
  • Languages and code-switching
  • Real-time streaming
  • Streaming audio requirements
  • Performance
  • Transcription accuracy
  • Streaming latency
  • End-of-turn detection
  • End-of-utterance control
  • Priority words
  • Text normalization
  • Punctuation
  • Async file transcription
  • Capabilities
  • Using Nabrah...
  • Default model
  • EOT variant

Nabrah STT V1

Stable

Nabrah STT V1 is a real-time speech-to-text model designed for streaming speech recognition.

It supports Arabic, English, and code-switching between Arabic and English within the same conversation.

Nabrah STT V1 is available in two variants:

Default: optimized for higher transcription accuracy.

EOT: optimized for conversational applications that require end-of-turn detection.

Model overview

Specification

Value

Model family

Nabrah STT V1

Version

V1

Status

Stable

Release date

2026-03

Primary use case

Real-time streaming transcription

Languages

Arabic, English

Code-switching

Arabic-English supported

Streaming

Supported

Streaming latency

As low as 160 ms

Streaming sample rate

16 kHz

Streaming encoding

LINEAR_PCM

Channels

Mono

Priority words

Supported

Available variants

Default

The default Nabrah STT V1 model provides the highest transcription accuracy within the current Nabrah STT model family.

It is recommended when accurate transcription is more important than built-in end-of-turn detection.

The default model is selected automatically when no alternative recognition_model is specified.

EOT

eot_nabrah

The EOT variant is designed for real-time conversational AI and voice agent applications.

It emits a special <eot> token when it detects that the speaker has reached the end of their conversational turn.

Compared with the default model, the EOT variant provides lower transcription accuracy but adds native end-of-turn signaling and applies Inverse Text Normalization (ITN).

Compare variants

Capability

Default

EOT

Transcription accuracy

Higher

Lower

Real-time streaming

Yes

Yes

Arabic

Yes

Yes

English

Yes

Yes

Arabic-English code-switching

Yes

Yes

End-of-turn detection

No

Yes

Text normalization

Yes

Yes

Automatic punctuation

No

Yes

Recommended for

General real-time transcription

Voice agents and conversational AI

Recommended use cases

Default

Use the default model for applications such as:

  • Real-time transcription

  • Customer support calls

  • Live speech recognition

  • Voice applications where transcription accuracy is the priority

  • Arabic-English conversations

  • Applications that use external turn-detection logic

EOT

Use the EOT variant for applications such as:

  • Voice agents

  • Conversational AI

  • AI assistants

  • Real-time interactive voice systems

  • Applications that require native end-of-turn detection

Languages and code-switching

Nabrah STT V1 supports:

  • Arabic

  • English

  • Arabic-English code-switching

Code-switching allows the model to recognize Arabic and English within the same sentence or conversation.

Example:

على الإيميل Report أرسل لي ال

Real-time streaming

Nabrah STT V1 is primarily designed for real-time streaming speech recognition.

Streaming is available over WebSocket:

WS /api/ext/stt/ws

Audio is streamed as raw PCM binary frames.

Streaming audio requirements

Specification

Value

Encoding

LINEAR_PCM

Sample rate

16 kHz

Channels

Mono

Transport

WebSocket

Streaming latency

As low as 160 ms

Performance

Transcription accuracy

The default Nabrah STT V1 model achieved:

WER: 31.84%

Benchmark: SADA
Environment: Cloud-based streaming use case

For more info about the SADA click here

35.19% for eot_nabrah model

Streaming latency

Streaming latency can be as low as:

160 ms

Actual latency may vary depending on infrastructure, network conditions, audio input, application configuration, and the first token it self.

End-of-turn detection

The EOT variant can detect when the speaker has completed their conversational turn.

Select it using:

eot_nabrah

When an end-of-turn is detected, the model emits:

<eot>

Example transcription:

```text \ \
ابي احجز موعد بكرة

```

Applications can use the <eot> token as a signal to trigger the next stage of the conversational pipeline, such as sending the transcript to an LLM.

The token is returned inline with the transcript and should be removed by the client if it should not be displayed to the end user.

End-of-utterance control

For streaming transcription, the end_of_utterance_silence_ms setting controls how much trailing silence is required before the current utterance is treated as complete.

Example:

```json \
{ \
"end_of_utterance_silence_ms": 500 \
}

```

Lower values can make conversational responses faster, while higher values allow speakers to pause for longer without ending the utterance.

This setting affects both preliminary and final transcription behavior.

Priority words

Nabrah STT V1 supports recognition biasing using priority words and phrases.

This is useful for improving recognition of:

- Brand names

- Product names

- Industry terminology

- Names

- Domain-specific vocabulary

Example:

```json \
{ \
"priority_words": [ \
"نبرة", \
"ذكاء اصطناعي", \
"خدمة العملاء" \
], \
"priority_words_strength": 0.5 \
}

```

The priority_words_strength parameter controls how strongly the recognizer favors the supplied vocabulary.

Text normalization

The default model applies stronger Arabic orthographic normalization to some characters in the returned transcript.

Examples include normalization of some Alef forms and converting ة to ه.

The EOT variant applies lighter normalization and preserves the original text form more closely.

For conversational applications where the transcript is passed directly to an LLM, this normalization generally preserves the intended meaning.

Punctuation

The default Nabrah STT V1 model does not currently generate punctuation automatically.

Variant

Automatic punctuation

Default

Not supported

EOT

Supported

Async file transcription

Nabrah also supports asynchronous transcription for longer audio files.

Maximum audio duration: 2 hours

File transcription is available through:

POST /api/ext/stt/async-transcribe

The transcription runs in the background and moves through the following states:

queued → in_progress → completed

Completed results include:

- Transcribed text

- Audio duration

- Timestamps

For file-based transcription, WAV is recommended.

Capabilities

Capability

Support

Real-time streaming

Yes

Arabic

Yes

English

Yes

Arabic-English code-switching

Yes

Priority words

Yes

End-of-turn detection

EOT variant

Automatic punctuation

No on default

Async file transcription

Yes

Timestamps

Yes

Speaker diarization

No

Word-level timestamps

Yes

Language detection

No

Noise robustness

Yes

Using Nabrah STT V1

Default model

The default model is selected when recognition_model is omitted or left empty.

```json \
{ \
"api_key": "YOUR_API_KEY", \
"recognition_model": "", \
"end_of_utterance_silence_ms": 500 \
}

```

EOT variant

To enable native end-of-turn detection:

```json \
{ \
"api_key": "YOUR_API_KEY", \
"recognition_model": "eot_nabrah", \
"end_of_utterance_silence_ms": 500 \
}

```

Generate Text
API Reference

في هذه الصفحة

  • Nabrah STT V1
  • Model overview
  • Available variants
  • Default
  • EOT
  • Compare variants
  • Recommended use cases
  • Default
  • EOT
  • Languages and code-switching
  • Real-time streaming
  • Streaming audio requirements
  • Performance
  • Transcription accuracy
  • Streaming latency
  • End-of-turn detection
  • End-of-utterance control
  • Priority words
  • Text normalization
  • Punctuation
  • Async file transcription
  • Capabilities
  • Using Nabrah...
  • Default model
  • EOT variant