OpenAI-compatible API
Point any OpenAI SDK at Revolab with only a base_url and API-key swap. Same
keys, same quotas, same metering as the native API — the compatible routes are
thin adapters over the same engine.
Endpoints
POST https://api.revolab.ai/v1/audio/speech
POST https://api.revolab.ai/v1/audio/transcriptionsSet base_url="https://api.revolab.ai/v1" and pass your rvl_live_ key as the
OpenAI API key. Errors use the OpenAI envelope, so SDK exception types
(BadRequestError, RateLimitError, …) behave as expected.
Text to speech
Python
from openai import OpenAI
client = OpenAI(base_url="https://api.revolab.ai/v1", api_key="rvl_live_...")
# Start here: mp3 is streamed as it is synthesized, so playback can begin
# in ~0.45 s no matter how long the text is.
with client.audio.speech.with_streaming_response.create(
model="nada-1.0-pro",
voice="<your-voice-id>",
input="Streaming starts before synthesis finishes.",
response_format="mp3",
) as response:
for chunk in response.iter_bytes():
play(chunk) # first chunk arrives long before the last
# Buffered WAV — our default. Use it when you want the whole file at once
# (a Content-Length up front, and the generation kept in your History).
speech = client.audio.speech.create(
model="nada-1.0-pro",
voice="<your-voice-id>",
input="Hello from Revolab.",
response_format="wav",
)
speech.write_to_file("speech.wav")Speech to text
Python
from openai import OpenAI
client = OpenAI(base_url="https://api.revolab.ai/v1", api_key="rvl_live_...")
audio_file = open("speech.wav", "rb")
transcript = client.audio.transcriptions.create(
model="aisyah-1.0-pro",
file=audio_file,
)
print(transcript.text)
# stream=True works too (SDK-compatibility emulation: one full-text
# delta + done — latency is unchanged, our engines return the final
# transcript in one shot)
stream = client.audio.transcriptions.create(
model="aisyah-1.0-pro", file=audio_file, stream=True
)
for event in stream:
print(event.type, getattr(event, "delta", getattr(event, "text", "")))Support matrix
| Area | Support |
|---|---|
| TTS output | Every response_format OpenAI names: wav, mp3, opus (Ogg-encapsulated), aac (ADTS), flac and pcm (24 kHz mono s16le). The default is wav, not OpenAI’s mp3 — pass response_format explicitly if that matters to you. Delivery returns the audio bytes directly, as the native /v1/tts does. |
| TTS streaming | Follows the container: mp3, opus, aac and pcm stream chunks as the model produces them. wav and flac are buffered, because both finish their header only once the audio is known — that is what gives them a real, seekable duration. stream_format="sse" additionally emits speech.audio.delta / done events and works with any streamed format. Streamed generations are not stored to your History. |
| STT formats | json (default), text, verbose_json. srt / vtt and word timestamps are not available (engines return whole-utterance text). |
| STT streaming | stream=true is an emulation: one transcript.text.delta with the full text, then transcript.text.done. No latency benefit. |
| Models | Revolab model IDs only (not OpenAI’s) — see Models for the IDs available to you. OpenAI names like tts-1 or whisper-1 return a 400 listing valid models. |
| Model listing | client.models.list() works — GET /v1/models returns the active catalog in the OpenAI list shape. |
| Speed | 0.5–2.0 (OpenAI allows 0.25–4.0; out-of-range values are rejected, never silently clamped). |
Voice IDs come from the Voice Library in your dashboard; quotas and error codes match the native API.