Provider Models
Provider Models connect Lakehousecat to an external AI provider. They encapsulate the credentials and configuration needed to call a specific AI service. Creating and managing Provider Models is an Administrator responsibility, as it involves API keys and incurs costs on the provider's side.
Supported Providers
Lakehousecat supports the following AI providers:
- OpenAI — Chat, Speech-to-Text (STT)
- Anthropic — Chat
- Google — Chat, Speech-to-Text (STT)
- Azure (Azure OpenAI) — Chat, Speech-to-Text (STT)
- AWS (Amazon Bedrock / Transcribe) — Chat, Speech-to-Text (STT)
- xAI — Chat, Speech-to-Text (STT)
Creating a Provider Model
Navigate to Workspace → Models → Provider Models and click the + icon.
Required Fields
| Field | Description |
|---|---|
| Provider Type | Select the provider (OpenAI, Anthropic, Google, Azure, AWS, xAI). |
| Model Type | Select the task type: Chat or STT (Speech-to-Text). |
| Configuration Name | A unique, descriptive name for this configuration (e.g. anthropic-claude-chat). |
Provider-specific credential fields are shown after selecting the Provider Type. See the individual provider pages for details.
Administrator Default Settings
After filling in the required fields, administrators can optionally set the saved configuration as:
- Default UI model — used as the pre-selected model in the chat interface for the selected model type
- Default Backend model — used by the backend for internal operations (Chat type only)
These defaults are applied after saving the configuration.
The Default Backend Model
One Provider Model can be designated as the Default Backend Model. This model is used by the system for all semantic extraction operations — the background process that analyzes data sources and builds the semantic layer that enables natural language querying.
The Default Backend Model is a system-level configuration, distinct from the Provider Model assigned to a Custom Model for chat queries. It is set by the Administrator and applies globally across all data source and Custom Model operations.
To set the Default Backend Model, open a saved Provider Model configuration and toggle Default Backend model in the Administrator Default Settings section.
Multiple Providers
Multiple Provider Models can be active in parallel. The relationship between a Provider Model and a Custom Model is 1:1 — each Custom Model uses exactly one Provider Model for chat queries.
However, this assignment can be changed after the Custom Model is created. For example, you can start with a cost-effective model during development and switch to a more capable model for production — without recreating the Custom Model.
Different Custom Models can each use a different Provider Model, allowing you to run multiple providers simultaneously for different use cases.
Model Compatibility
The model identifiers shown in Lakehousecat's dropdowns are models that have been actively tested with the platform. Lakehousecat validates correct API integration, response formatting, and analytical behavior for each listed model.
The AI model landscape evolves rapidly — new models are released nearly every week. We cannot test every model version as it appears. As a result, the listed models represent a subset of what is technically compatible. In practice, most chat-capable models from the supported providers work correctly with Lakehousecat even if not explicitly listed.
If you need to use a model that is not in the dropdown, you can enter the model ID manually. Lakehousecat will attempt to use it, but compatibility is not guaranteed for models that have not been tested.
We continuously expand our test coverage and update the model lists as new versions are validated.
Sharing
After creating a Provider Model, an Administrator can share it with individual users or groups using the Share function.
Sharing with Builders: Builders who have been granted access to a Provider Model can use it as the base for Custom Models. This is the standard workflow.
Sharing directly with Users: A Provider Model can also be shared with regular Users (without a Custom Model in between). In this case, the User interacts directly with the LLM provider — this is pure chatbot behavior with no system prompt, no data connections, and no Lakehousecat skills. This is a valid but limited use case, essentially a raw API proxy. For full data-driven analytics, share a Custom Model instead.
API Rate Limits
The effective API rate limits depend on your chosen AI provider and the specific plan or contract you have with them. Lakehousecat itself does not impose additional rate limits on top of what the provider enforces.
Some features — in particular chart generation — use internal agent workflows that issue multiple sequential API calls to complete a single user request. Depending on your provider plan, this may cause rate limits to be reached faster than with simple single-turn queries.
When a rate limit is exceeded, the provider typically returns HTTP status 429 (Too Many Requests). This is expected behavior from the provider side and not a bug in Lakehousecat.
If you encounter rate limit errors:
- Check the application logs for HTTP 429 entries or similar rate limit indicators.
- Review which features or workflows triggered the limit (e.g. chart generation, semantic extraction).
- Contact your AI provider to discuss plan upgrades or rate limit adjustments.
Model Quality Recommendations
Lakehousecat's analytics and chart generation workflows are agentic: the model must follow multi-step instructions, reason over structured data, and ground its responses in query results. Not every AI model is suited for this.
Frontier models are strongly recommended for production use. Models from the leading providers — such as the latest GPT-4 class models via OpenAI or Azure, Claude Sonnet/Opus via Anthropic or AWS Bedrock, and Gemini Pro/Ultra via Google — have been validated against Lakehousecat's full feature set and consistently deliver correct, data-grounded responses.
Open-weight models are experimental. Many open-weight models, even recent and capable ones, can produce lower-quality responses in agentic workflows:
- They may ignore provided data context and generate generic or fictitious content.
- They may fail to follow structured tool-call instructions correctly.
- Response quality varies significantly between model families and versions.
Lakehousecat cannot validate every open-weight model that becomes available. If you choose to configure an open-weight model:
- Treat it as experimental, without quality guarantees.
- Test it against your actual data before rolling it out to end users.
- Be aware that prompt behavior tuned for frontier models may not transfer.
For analytics-heavy workloads, we recommend staying with frontier models unless you have a specific reason to use open-weight alternatives.
Next Steps
For provider-specific configuration details, see the dedicated pages: