Skip to Content
APIText to Speech

TTS API

Synthesize speech from text using Revolab’s Nada voice models.

POST https://api.revolab.ai/v1/tts

Request body (JSON)

ParameterTypeRequiredDefaultDescription
modelstringoptionalthis environment’s default TTS modelModel to use for synthesis. See Models for the models available to you.
textstringrequired—Text to synthesize. Maximum 10,000 characters. Long text is segmented server-side — do not split it yourself.
voice_idstringrequired—Voice identifier. List the IDs your key accepts with GET /v1/voices.
languagestringoptional"en"BCP-47 language code (e.g., "en", "ms", "zh"). Defaults to English when omitted.
speedfloatoptionalnull (normal speed)Playback speed multiplier. Range: 0.5 (half speed) to 2.0 (double speed).
output_formatstringoptional"wav"Output container: wav, mp3, ogg (Ogg/Opus), aac, flac, or pcm (raw 24 kHz mono s16le). See Output formats.
streambooleanoptionalfalsetrue delivers audio as it is synthesized instead of waiting for the whole generation. Recommended for anything a person is waiting on — see Choosing a delivery mode.

Request examples

Reach for "stream": true first. Audio is delivered as it is synthesized, so playback can begin in about 0.45 seconds whatever the length of the text; a buffered call cannot start until the whole generation is finished, which for a long passage is the difference between half a second and most of a minute. The buffered form below is the right choice when you want the complete file in one piece — see Choosing a delivery mode.

# Streamed (start here) — MP3 arrives progressively and plays as it lands. curl -X POST https://api.revolab.ai/v1/tts \ -H "Authorization: Bearer $REVOLAB_API_KEY" \ -H "Content-Type: application/json" \ --no-buffer \ -d '{ "model": "nada-1.0-pro", "text": "The quick brown fox jumped over the lazy dog.", "voice_id": "<your-voice-id>", "stream": true, "output_format": "mp3" }' \ -o speech.mp3 # Buffered — one complete WAV, kept in your History. curl -X POST https://api.revolab.ai/v1/tts \ -H "Authorization: Bearer $REVOLAB_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "nada-1.0-pro", "text": "The quick brown fox jumped over the lazy dog.", "voice_id": "<your-voice-id>", "language": "en", "speed": 1.0, "output_format": "wav" }'

Response (200 OK)

The response body is the audio — a complete WAV file (Content-Type: audio/wav). Save it or pipe it straight to a player; there is no download URL round-trip. Generation metrics ride as response headers:

HeaderDescription
X-Duration-SDuration of the generated audio in seconds.
X-Latency-MsSynthesis latency in milliseconds.
X-Request-IdRequest id — quote it in support requests.
curl -X POST https://api.revolab.ai/v1/tts \ -H "Authorization: Bearer $REVOLAB_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "nada-1.0-pro", "voice_id": "<your-voice-id>", "text": "Hello!"}' \ -o hello.wav

Generations are still stored to your History (dashboard playback) — the upload happens in the background and never delays the response.

Choosing a delivery mode

Streamed ("stream": true)Buffered (default)
Time to first audio~0.45 s at any text lengthgrows with the text: ~0.8 s at 100 chars, ~56 s at 10,000
Transferchunked, no Content-Lengthone body with a Content-Length
Metrics headersnoneX-Duration-S, X-Latency-Ms
Stored in Historyno — the stream is the deliveryyes
Failure after the first byteconnection ends; treat a short stream as incompletenormal JSON error envelope

Anything a person waits on should stream. Choose buffered when you need the length up front, want the metrics headers, or want the generation kept for dashboard playback.

Output formats

output_format is enforced, not advisory — an unsupported value returns 400 validation_error rather than quietly substituting another container.

output_formatContent-TypeNotes
wav (default)audio/wav24 kHz mono 16-bit RIFF, 384 kbps. Buffered only — a RIFF header is length-prefixed, so under stream: true this delivers raw PCM (see Streaming).
mp3audio/mpeg64 kbps mono. ~6x smaller than WAV. Buffered or streamed.
oggaudio/oggOgg/Opus at 32 kbps. ~13x smaller than WAV and the best quality-per-byte for speech. Buffered or streamed.
aacaudio/aac64 kbps ADTS (not .m4a — an MP4 container puts its index at the end and cannot stream). ~6x smaller than WAV. Buffered or streamed.
flacaudio/flacLossless, ~130 kbps on speech, so about 3x smaller than WAV with identical audio. Buffered or streamed — but see the note below.
pcmaudio/pcmRaw 24 kHz mono 16-bit little-endian, no container. What the engine produces natively — lowest latency, no decode step.

Sizes above are measured on 37 seconds of real speech, not estimated.

A note on streamed FLAC: FLAC records its total sample count in a header written before the audio, so a streamed FLAC cannot state its own length and players will show an unknown duration and seek poorly. Buffered FLAC has the field filled in. If you want lossless for archiving or editing, ask for it buffered; stream it only if you are consuming the samples directly.

curl -X POST https://api.revolab.ai/v1/tts \ -H "Authorization: Bearer $REVOLAB_API_KEY" \ -H "Content-Type: application/json" \ -d '{"voice_id": "<your-voice-id>", "text": "Hello!", "output_format": "mp3"}' \ -o hello.mp3

pcm responses also carry X-Sample-Rate: 24000 and X-Audio-Format: pcm_s16le so the format is never out-of-band.

Long text

Send the whole passage in one request. Text is segmented server-side on sentence boundaries and the segments are synthesized as one continuous generation, so prosody carries across the whole passage.

Splitting client-side and stitching the parts is actively worse: each part is synthesized in isolation, so intonation resets at every seam and the joins are audible. It also serialises the round-trips, which is what makes long replies feel slow.

Above roughly 2,000 characters, prefer stream: true. The buffered path holds the entire generation in memory before it sends anything, so both the wait and the memory grow with the text; the streaming path relays as it goes and its time-to-first-audio does not move.

Streaming

Set "stream": true to receive audio while it is being synthesized, instead of waiting for the full generation. Playback can start after the first chunk — usually well under a second.

Time-to-first-audio holds at roughly 0.4 s whether the text is 100 characters or 10,000 — the first sentence starts playing while the rest is still being generated. The buffered path, by contrast, waits for the whole generation.

The stream honours output_format:

output_formatStreamed Content-Type
pcmaudio/pcm — raw 24 kHz mono 16-bit little-endian
wav (default)audio/pcm — see the note below
mp3audio/mpeg
oggaudio/ogg (Opus)
aacaudio/aac (ADTS)
flacaudio/flac — lossless, but with no duration in the header

wav streams as raw PCM rather than a RIFF file: a WAV header is length-prefixed, so the container cannot be written before the audio exists. Pass output_format: "pcm" to say so explicitly, or mp3/ogg for a container a player can consume directly.

PCM responses carry two extra headers so the format is never out-of-band:

HeaderValue
X-Sample-Rate24000
X-Audio-Formatpcm_s16le
curl --no-buffer -X POST https://api.revolab.ai/v1/tts \ -H "Authorization: Bearer $REVOLAB_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "nada-1.0-pro", "voice_id": "aisyah", "text": "Hello from Revolab!", "stream": true}' \ | ffplay -f s16le -ar 24000 -ch_layout mono -i - -autoexit -nodisp

An MP3 stream needs no decoding on the client — pipe it straight to a player:

curl --no-buffer -X POST https://api.revolab.ai/v1/tts \ -H "Authorization: Bearer $REVOLAB_API_KEY" \ -H "Content-Type: application/json" \ -d '{"voice_id": "aisyah", "text": "Hello from Revolab!", "stream": true, "output_format": "mp3"}' \ | ffplay -i - -autoexit -nodisp

Notes:

  • Streamed generations are not stored; the stream is the delivery. Use the default buffered response when you want the generation kept in your History.
  • Errors that occur before the first chunk (unknown voice, depleted balance, rate limits) return the normal JSON error envelope. Once audio starts flowing, the connection simply ends on failure — treat an unexpectedly short stream as an incomplete generation.
  • Billing meters the audio actually delivered: disconnect early and you are charged only for what you received.

Using an OpenAI SDK instead? The OpenAI-compatible endpoint offers the same stream via response_format: "pcm" (and an SSE framing via stream_format: "sse").

Status codes

200 OK — Success

WAV audio bytes (audio/wav) with X-Duration-S / X-Latency-Ms headers.

400 Bad Request — Unknown model, text too long, invalid voice_id, or unsupported output_format

{ "error": { "code": "validation_error", "message": "Model 'nada-2.0' is not available.", "request_id": "req_01HXYZ..." } }

401 Unauthorized — Missing, invalid, or revoked API key

{ "error": { "code": "unauthorized", "message": "Invalid API key format. Keys must start with 'rvl_live_'.", "request_id": "req_01HXYZ..." } }

422 Unprocessable Entity — Request body fails JSON schema validation (e.g., missing required field)

{ "detail": [ { "loc": ["body", "text"], "msg": "field required", "type": "value_error.missing" } ] }

402 Payment Required — The organization’s balance is depleted

{ "error": { "code": "insufficient_balance", "message": "Insufficient balance to complete this request.", "request_id": "req_01HXYZ..." } }

429 Too Many Requests — Rate limit exceeded

Also quota_exceeded (monthly cap) and concurrency_limit_exceeded (too many in-flight requests); honor Retry-After.

{ "error": { "code": "rate_limited", "message": "Rate limit exceeded.", "request_id": "req_01HXYZ..." } }

502 Bad Gateway — TTS model endpoint returned an error or malformed audio

{ "error": { "code": "endpoint_unavailable", "message": "The upstream TTS service returned an unexpected response.", "request_id": "req_01HXYZ..." } }

503 Service Unavailable — TTS service is under maintenance or endpoint config is missing

{ "error": { "code": "service_unavailable", "message": "TTS service is temporarily unavailable.", "request_id": "req_01HXYZ..." } }