Audio API
The Audio API provides speech-to-text transcription and text-to-speech synthesis. It is the backend for voice input in Lakehousecat chat sessions and can also be called directly for audio processing tasks.
Base URL: http://<host>:42012
What you can do
- Transcribe audio files to text (speech-to-text)
- Optionally post-process transcriptions with an LLM to improve accuracy
- Convert text to speech (text-to-speech)
Authentication
All endpoints require a Bearer token. Generate your API key in the Lakehousecat UI:
Account Settings → Security → API Keys → Generate API Key
Authorization: Bearer <your-api-key>
Quick Start
Transcribe an audio file
curl -X POST "http://localhost:42012/api/v1/audio/transcriptions" \
-H "Authorization: Bearer <your-api-key>" \
-F "file=@/path/to/audio.mp3"
import requests
BASE_URL = "http://localhost:42012"
HEADERS = {"Authorization": "Bearer <your-api-key>"}
with open("/path/to/audio.mp3", "rb") as f:
response = requests.post(
f"{BASE_URL}/api/v1/audio/transcriptions",
headers=HEADERS,
files={"file": f},
)
transcription = response.json()["text"]
print(transcription)
Transcribe with LLM optimization
Pass optimize=true to run the transcription result through an LLM for improved punctuation and accuracy:
curl -X POST "http://localhost:42012/api/v1/audio/transcriptions?optimize=true" \
-H "Authorization: Bearer <your-api-key>" \
-F "file=@/path/to/audio.mp3"
with open("/path/to/audio.mp3", "rb") as f:
response = requests.post(
f"{BASE_URL}/api/v1/audio/transcriptions",
params={"optimize": True},
headers=HEADERS,
files={"file": f},
)
Text-to-speech
curl -X POST "http://localhost:42012/api/v1/audio/speech" \
-H "Authorization: Bearer <your-api-key>" \
-H "Content-Type: application/json" \
-d '{"input": "Hello, how can I help you today?", "voice": "alloy"}' \
--output speech.mp3
response = requests.post(
f"{BASE_URL}/api/v1/audio/speech",
json={"input": "Hello, how can I help you today?", "voice": "alloy"},
headers=HEADERS,
)
with open("speech.mp3", "wb") as f:
f.write(response.content)
Supported Audio Formats
The transcription endpoint accepts common audio formats including MP3, MP4, WAV, M4A, and WebM. Maximum file size depends on the configured provider.
Endpoint Groups
- Transcriptions —
/api/v1/audio/transcriptions— speech-to-text - Speech —
/api/v1/audio/speech— text-to-speech - Config —
/api/v1/audio/config— read active audio configuration
UI Equivalent
Voice input in chat sessions uses this API. The microphone button in the session input toolbar triggers a transcription request automatically.