NabrahNabrahNabrah
  • Demos
  • Features
  • Pricing
LoginContact us
ModelsTTS-V1

Search documentation

Search documentation by title or category.

ModelsTTS-V1

Search documentation

Search documentation by title or category.

Docs
Introduction
Core concepts
Quick Start
Nabrah Agents overview
Workspace overview
Calls analytics
WhatsApp analytics
Agents
Custom tools
Analysis groups
Follow-ups
Inbound services
Outbound services
Contacts
WhatsApp
Call logs
Ticketing
Nabrah Studio overview
Text to speech
Speech to text
Voice cloning
Voices
Plugins overview
Web-call widget
Direct link
Account
Workspace & teams
TTS-V1
STT-V1

Nabrah TTS

TTS-V1

10/01/2026

On this page

Jump to a section
  • Nabrah_tts_v1
  • Stable
  • Model overview
  • Recommended use cases
  • Streaming and non-streaming
  • Streaming
  • Non-streaming
  • Audio output
  • MP3
  • WAV
  • PCM
  • Voice cloning
  • Input limits
  • Speech controls
  • Latency
  • Using Nabrah...

Nabrah TTS V1

Nabrah_tts_v1

Stable

Nabrah TTS V1 is a text-to-speech model designed for real-time voice applications. It converts text into natural speech and supports both streaming and non-streaming generation.

The model can be used with Nabrah's available voices as well as custom cloned voices.

Model overview

Specification

Value

Model ID

Nabrah_tts_v1

Version

V1

Status

Stable

Release date

2026-06

Supported languages

Arabic,English,French, German, Hindi, Indonesian, Italian, Turkish, Spanish, Russian

Supported dialects

Saudi Arabic

Streaming

Supported

Non-streaming

Supported

Voice cloning

Supported

Server-side TTFB

120 – 130 ms

Maximum input length

16,384 characters

Audio formats

MP3, WAV, PCM

Recommended use cases

Nabrah TTS V1 is suitable for applications that require generated speech, including:

  • Voice agents

  • Customer service applications

  • IVR systems

  • Conversational AI

  • Interactive voice experiences

  • Content narration

  • Real-time speech generation

Streaming and non-streaming

Nabrah TTS V1 supports both streaming and non-streaming speech generation.

Streaming

In streaming mode, audio is delivered progressively as it is generated. This allows applications to begin receiving audio before the full speech output has been synthesized.

Streaming is recommended for latency-sensitive applications such as voice agents and real-time conversational systems.

Server-side TTFB: 120 –130 ms

TTFB represents the server-side time until the first response bytes are returned. It does not represent full end-to-end application latency.

Non-streaming

In non-streaming mode, the complete audio file is returned after speech generation has finished.

This mode is suitable when progressive playback is not required, such as generating audio files for later playback or storage.

Audio output

Nabrah TTS V1 supports three output formats.

Format

Specification

Recommended for

MP3

192 kbps compressed audio

Web playback and smaller payloads

WAV

48 kHz, mono, 16-bit PCM

High-quality lossless audio

PCM

48 kHz, mono, signed 16-bit little-endian

Real-time audio pipelines

MP3

MP3 provides compressed audio at 192 kbps, making it suitable for applications where smaller payload size and broad playback compatibility are important.

WAV

WAV output uses uncompressed 16-bit PCM audio at 48 kHz mono.

It is suitable for applications that require lossless audio or a self-contained audio file with format information included in the file header.

PCM

PCM output contains raw signed 16-bit little-endian audio samples at 48 kHz mono.

It does not contain a file header and is intended for applications that consume raw audio frames directly.

Voice cloning

Nabrah TTS V1 supports custom cloned voices.

A cloned voice can be used for speech generation after the cloning process has completed and the voice is marked as ready.

The same voice IDs can be used when configuring Nabrah voice agents.

Learn about Voice Cloning

View available Voices

Input limits

The model accepts up to:

16,384 characters per request

The limit is calculated after text normalization.

Requests that exceed the supported input length must be divided into smaller segments before synthesis.

Speech controls

Some speech control parameters are currently exposed through the API but are not yet applied to generated audio.

Capability

Current support

Voice cloning

Supported

Speed control

Not currently applied

Expressiveness control

Not currently applied

Pronunciation control

Not currently applied

SSML

Supported

Support for additional speech controls may change as the model evolves.

Latency

Metric

Result

Server-side TTFB

120–130 ms

Actual end-to-end latency may vary depending on network conditions, infrastructure, request configuration, and client-side audio handling.

Using Nabrah TTS V1

Use the following model ID when generating speech:

Nabrah_tts_v1

Speech generation is available through:

POST /api/ext/tts/generations

Example:

```json \
{ \
"model": "Nabrah_tts_v1", \
"input": "Hello from Nabrah!", \
"voice": "YOUR_VOICE_ID", \
"response_format": "mp3", \
"stream": false \
}

```

Generate Speech

API Reference

On this page

  • Nabrah_tts_v1
  • Stable
  • Model overview
  • Recommended use cases
  • Streaming and non-streaming
  • Streaming
  • Non-streaming
  • Audio output
  • MP3
  • WAV
  • PCM
  • Voice cloning
  • Input limits
  • Speech controls
  • Latency
  • Using Nabrah...