Skip to main content
Version: 0.0.41

Audio API

The Audio API provides speech-to-text transcription and text-to-speech synthesis. It is the backend for voice input in Lakehousecat chat sessions and can also be called directly for audio processing tasks.

Base URL: http://<host>:42012

What you can do​

  • Transcribe audio files to text (speech-to-text)
  • Optionally post-process transcriptions with an LLM to improve accuracy
  • Convert text to speech (text-to-speech)

Authentication​

All endpoints require a Bearer token. Generate your API key in the Lakehousecat UI:

Account Settings → Security → API Keys → Generate API Key

Authorization: Bearer <your-api-key>

Quick Start​

Transcribe an audio file​

curl -X POST "http://localhost:42012/api/v1/audio/transcriptions" \
-H "Authorization: Bearer <your-api-key>" \
-F "file=@/path/to/audio.mp3"
import requests

BASE_URL = "http://localhost:42012"
HEADERS = {"Authorization": "Bearer <your-api-key>"}

with open("/path/to/audio.mp3", "rb") as f:
response = requests.post(
f"{BASE_URL}/api/v1/audio/transcriptions",
headers=HEADERS,
files={"file": f},
)

transcription = response.json()["text"]
print(transcription)

Transcribe with LLM optimization​

Pass optimize=true to run the transcription result through an LLM for improved punctuation and accuracy:

curl -X POST "http://localhost:42012/api/v1/audio/transcriptions?optimize=true" \
-H "Authorization: Bearer <your-api-key>" \
-F "file=@/path/to/audio.mp3"
with open("/path/to/audio.mp3", "rb") as f:
response = requests.post(
f"{BASE_URL}/api/v1/audio/transcriptions",
params={"optimize": True},
headers=HEADERS,
files={"file": f},
)

Text-to-speech​

curl -X POST "http://localhost:42012/api/v1/audio/speech" \
-H "Authorization: Bearer <your-api-key>" \
-H "Content-Type: application/json" \
-d '{"input": "Hello, how can I help you today?", "voice": "alloy"}' \
--output speech.mp3
response = requests.post(
f"{BASE_URL}/api/v1/audio/speech",
json={"input": "Hello, how can I help you today?", "voice": "alloy"},
headers=HEADERS,
)

with open("speech.mp3", "wb") as f:
f.write(response.content)

Supported Audio Formats​

The transcription endpoint accepts common audio formats including MP3, MP4, WAV, M4A, and WebM. Maximum file size depends on the configured provider.

Endpoint Groups​

  • Transcriptions — /api/v1/audio/transcriptions — speech-to-text
  • Speech — /api/v1/audio/speech — text-to-speech
  • Config — /api/v1/audio/config — read active audio configuration

UI Equivalent​

Voice input in chat sessions uses this API. The microphone button in the session input toolbar triggers a transcription request automatically.