Audio API
The Audio API provides speech-to-text transcription. It is the backend for voice input in Lakehousecat sessions and can also be called directly for audio processing tasks.
Base URL: http://<host>:42012
What you can do
- Transcribe audio files to text (speech-to-text)
- Optionally post-process transcriptions with an LLM to improve accuracy
Authentication
All endpoints require a Bearer token. Generate your API key in the Lakehousecat UI:
Account Settings → Security → API Keys → Generate API Key
Authorization: Bearer <your-api-key>
Quick Start
Transcribe an audio file
curl -X POST "http://localhost:42012/api/v1/audio/transcriptions" \
-H "Authorization: Bearer <your-api-key>" \
-F "file=@/path/to/audio.mp3"
import requests
BASE_URL = "http://localhost:42012"
HEADERS = {"Authorization": "Bearer <your-api-key>"}
with open("/path/to/audio.mp3", "rb") as f:
response = requests.post(
f"{BASE_URL}/api/v1/audio/transcriptions",
headers=HEADERS,
files={"file": f},
)
transcription = response.json()["text"]
print(transcription)
Transcribe with LLM optimization
Pass optimize=true to run the transcription result through an LLM for improved punctuation and accuracy:
curl -X POST "http://localhost:42012/api/v1/audio/transcriptions?optimize=true" \
-H "Authorization: Bearer <your-api-key>" \
-F "file=@/path/to/audio.mp3"
with open("/path/to/audio.mp3", "rb") as f:
response = requests.post(
f"{BASE_URL}/api/v1/audio/transcriptions",
params={"optimize": True},
headers=HEADERS,
files={"file": f},
)
Supported Audio Formats
The transcription endpoint accepts common audio formats including MP3, MP4, WAV, M4A, and WebM. Maximum file size depends on the configured provider.
Endpoint Groups
- Transcriptions —
/api/v1/audio/transcriptions— speech-to-text - Config —
/api/v1/audio/config— read active Speech-to-Text configuration
UI Equivalent
Voice input in sessions uses this API. The microphone button in the session input toolbar triggers a transcription request automatically.