TTS API
Synthesize speech from text using Revolab’s Nada voice models.
POST https://api.revolab.ai/v1/ttsRequest body (JSON)
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
model | string | optional | this environment’s default TTS model | Model to use for synthesis. See Models for the models available to you. |
text | string | required | — | Text to synthesize. Maximum 10,000 characters. Long text is segmented server-side — do not split it yourself. |
voice_id | string | required | — | Voice identifier. List the IDs your key accepts with GET /v1/voices. |
language | string | optional | "en" | BCP-47 language code (e.g., "en", "ms", "zh"). Defaults to English when omitted. |
speed | float | optional | null (normal speed) | Playback speed multiplier. Range: 0.5 (half speed) to 2.0 (double speed). |
output_format | string | optional | "wav" | Output container: wav, mp3, ogg (Ogg/Opus), aac, flac, or pcm (raw 24 kHz mono s16le). See Output formats. |
stream | boolean | optional | false | true delivers audio as it is synthesized instead of waiting for the whole generation. Recommended for anything a person is waiting on — see Choosing a delivery mode. |
Request examples
Reach for "stream": true first. Audio is delivered as it is synthesized, so
playback can begin in about 0.45 seconds whatever the length of the text; a
buffered call cannot start until the whole generation is finished, which for a
long passage is the difference between half a second and most of a minute. The
buffered form below is the right choice when you want the complete file in one
piece — see Choosing a delivery mode.
cURL
# Streamed (start here) — MP3 arrives progressively and plays as it lands.
curl -X POST https://api.revolab.ai/v1/tts \
-H "Authorization: Bearer $REVOLAB_API_KEY" \
-H "Content-Type: application/json" \
--no-buffer \
-d '{
"model": "nada-1.0-pro",
"text": "The quick brown fox jumped over the lazy dog.",
"voice_id": "<your-voice-id>",
"stream": true,
"output_format": "mp3"
}' \
-o speech.mp3
# Buffered — one complete WAV, kept in your History.
curl -X POST https://api.revolab.ai/v1/tts \
-H "Authorization: Bearer $REVOLAB_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "nada-1.0-pro",
"text": "The quick brown fox jumped over the lazy dog.",
"voice_id": "<your-voice-id>",
"language": "en",
"speed": 1.0,
"output_format": "wav"
}'Response (200 OK)
The response body is the audio — a complete WAV file (Content-Type: audio/wav). Save it or pipe it straight to a player; there is no download
URL round-trip. Generation metrics ride as response headers:
| Header | Description |
|---|---|
X-Duration-S | Duration of the generated audio in seconds. |
X-Latency-Ms | Synthesis latency in milliseconds. |
X-Request-Id | Request id — quote it in support requests. |
curl -X POST https://api.revolab.ai/v1/tts \
-H "Authorization: Bearer $REVOLAB_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "nada-1.0-pro", "voice_id": "<your-voice-id>", "text": "Hello!"}' \
-o hello.wavGenerations are still stored to your History (dashboard playback) — the upload happens in the background and never delays the response.
Choosing a delivery mode
Streamed ("stream": true) | Buffered (default) | |
|---|---|---|
| Time to first audio | ~0.45 s at any text length | grows with the text: ~0.8 s at 100 chars, ~56 s at 10,000 |
| Transfer | chunked, no Content-Length | one body with a Content-Length |
| Metrics headers | none | X-Duration-S, X-Latency-Ms |
| Stored in History | no — the stream is the delivery | yes |
| Failure after the first byte | connection ends; treat a short stream as incomplete | normal JSON error envelope |
Anything a person waits on should stream. Choose buffered when you need the length up front, want the metrics headers, or want the generation kept for dashboard playback.
Output formats
output_format is enforced, not advisory — an unsupported value returns
400 validation_error rather than quietly substituting another container.
output_format | Content-Type | Notes |
|---|---|---|
wav (default) | audio/wav | 24 kHz mono 16-bit RIFF, 384 kbps. Buffered only — a RIFF header is length-prefixed, so under stream: true this delivers raw PCM (see Streaming). |
mp3 | audio/mpeg | 64 kbps mono. ~6x smaller than WAV. Buffered or streamed. |
ogg | audio/ogg | Ogg/Opus at 32 kbps. ~13x smaller than WAV and the best quality-per-byte for speech. Buffered or streamed. |
aac | audio/aac | 64 kbps ADTS (not .m4a — an MP4 container puts its index at the end and cannot stream). ~6x smaller than WAV. Buffered or streamed. |
flac | audio/flac | Lossless, ~130 kbps on speech, so about 3x smaller than WAV with identical audio. Buffered or streamed — but see the note below. |
pcm | audio/pcm | Raw 24 kHz mono 16-bit little-endian, no container. What the engine produces natively — lowest latency, no decode step. |
Sizes above are measured on 37 seconds of real speech, not estimated.
A note on streamed FLAC: FLAC records its total sample count in a header written before the audio, so a streamed FLAC cannot state its own length and players will show an unknown duration and seek poorly. Buffered FLAC has the field filled in. If you want lossless for archiving or editing, ask for it buffered; stream it only if you are consuming the samples directly.
curl -X POST https://api.revolab.ai/v1/tts \
-H "Authorization: Bearer $REVOLAB_API_KEY" \
-H "Content-Type: application/json" \
-d '{"voice_id": "<your-voice-id>", "text": "Hello!", "output_format": "mp3"}' \
-o hello.mp3pcm responses also carry X-Sample-Rate: 24000 and
X-Audio-Format: pcm_s16le so the format is never out-of-band.
Long text
Send the whole passage in one request. Text is segmented server-side on sentence boundaries and the segments are synthesized as one continuous generation, so prosody carries across the whole passage.
Splitting client-side and stitching the parts is actively worse: each part is synthesized in isolation, so intonation resets at every seam and the joins are audible. It also serialises the round-trips, which is what makes long replies feel slow.
Above roughly 2,000 characters, prefer stream: true. The buffered path holds
the entire generation in memory before it sends anything, so both the wait and
the memory grow with the text; the streaming path relays as it goes and its
time-to-first-audio does not move.
Streaming
Set "stream": true to receive audio while it is being synthesized,
instead of waiting for the full generation. Playback can start after the
first chunk — usually well under a second.
Time-to-first-audio holds at roughly 0.4 s whether the text is 100 characters or 10,000 — the first sentence starts playing while the rest is still being generated. The buffered path, by contrast, waits for the whole generation.
The stream honours output_format:
output_format | Streamed Content-Type |
|---|---|
pcm | audio/pcm — raw 24 kHz mono 16-bit little-endian |
wav (default) | audio/pcm — see the note below |
mp3 | audio/mpeg |
ogg | audio/ogg (Opus) |
aac | audio/aac (ADTS) |
flac | audio/flac — lossless, but with no duration in the header |
wav streams as raw PCM rather than a RIFF file: a WAV header is
length-prefixed, so the container cannot be written before the audio exists.
Pass output_format: "pcm" to say so explicitly, or mp3/ogg for a
container a player can consume directly.
PCM responses carry two extra headers so the format is never out-of-band:
| Header | Value |
|---|---|
X-Sample-Rate | 24000 |
X-Audio-Format | pcm_s16le |
curl --no-buffer -X POST https://api.revolab.ai/v1/tts \
-H "Authorization: Bearer $REVOLAB_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "nada-1.0-pro", "voice_id": "aisyah", "text": "Hello from Revolab!", "stream": true}' \
| ffplay -f s16le -ar 24000 -ch_layout mono -i - -autoexit -nodispAn MP3 stream needs no decoding on the client — pipe it straight to a player:
curl --no-buffer -X POST https://api.revolab.ai/v1/tts \
-H "Authorization: Bearer $REVOLAB_API_KEY" \
-H "Content-Type: application/json" \
-d '{"voice_id": "aisyah", "text": "Hello from Revolab!", "stream": true, "output_format": "mp3"}' \
| ffplay -i - -autoexit -nodispNotes:
- Streamed generations are not stored; the stream is the delivery. Use the default buffered response when you want the generation kept in your History.
- Errors that occur before the first chunk (unknown voice, depleted balance, rate limits) return the normal JSON error envelope. Once audio starts flowing, the connection simply ends on failure — treat an unexpectedly short stream as an incomplete generation.
- Billing meters the audio actually delivered: disconnect early and you are charged only for what you received.
Using an OpenAI SDK instead? The OpenAI-compatible endpoint
offers the same stream via response_format: "pcm" (and an SSE framing via
stream_format: "sse").
Status codes
200 OK — Success
WAV audio bytes (audio/wav) with X-Duration-S / X-Latency-Ms headers.
400 Bad Request — Unknown model, text too long, invalid voice_id, or unsupported output_format
{
"error": {
"code": "validation_error",
"message": "Model 'nada-2.0' is not available.",
"request_id": "req_01HXYZ..."
}
}401 Unauthorized — Missing, invalid, or revoked API key
{
"error": {
"code": "unauthorized",
"message": "Invalid API key format. Keys must start with 'rvl_live_'.",
"request_id": "req_01HXYZ..."
}
}422 Unprocessable Entity — Request body fails JSON schema validation (e.g., missing required field)
{
"detail": [
{
"loc": ["body", "text"],
"msg": "field required",
"type": "value_error.missing"
}
]
}402 Payment Required — The organization’s balance is depleted
{
"error": {
"code": "insufficient_balance",
"message": "Insufficient balance to complete this request.",
"request_id": "req_01HXYZ..."
}
}429 Too Many Requests — Rate limit exceeded
Also quota_exceeded (monthly cap) and concurrency_limit_exceeded (too many
in-flight requests); honor Retry-After.
{
"error": {
"code": "rate_limited",
"message": "Rate limit exceeded.",
"request_id": "req_01HXYZ..."
}
}502 Bad Gateway — TTS model endpoint returned an error or malformed audio
{
"error": {
"code": "endpoint_unavailable",
"message": "The upstream TTS service returned an unexpected response.",
"request_id": "req_01HXYZ..."
}
}503 Service Unavailable — TTS service is under maintenance or endpoint config is missing
{
"error": {
"code": "service_unavailable",
"message": "TTS service is temporarily unavailable.",
"request_id": "req_01HXYZ..."
}
}Related
- Voices API — list the
voice_idvalues your key accepts - Models — compare Nada models
- Error catalog — full list of error codes
- Authentication — how to send your API key