Quickstart
This guide walks you through the full setup from a fresh deployment to your first session. The complete process takes 30–60 minutes depending on your infrastructure and the complexity of your data.
An interactive onboarding wizard is planned and will be available after MVP. For now, follow the steps below.
Before You Begin
- Lakehousecat must be deployed and accessible (see Deployment)
- You must have an Administrator account
- You need API credentials for at least one AI provider (OpenAI, Anthropic, Google, Azure, or AWS)
- You need connection details for at least one database or data source
Your Infrastructure, Your Data
Lakehousecat runs entirely within your infrastructure. Your data never leaves your environment — it is loaded into the ClickHouse instance that is part of your deployment. All AI processing uses your own provider API keys, called against the provider directly from your cluster.
Step 1 — Configure a Provider Model
Who: Administrator Where: Workspace → Models → Provider Models
Create a Provider Model by entering your API credentials for the AI provider of your choice. Set one model as the Default Backend Model — this is the model the system uses for semantic extraction operations.
See Provider Models for details.
Step 2 — Connect a Data Source
Who: Administrator or Builder Where: Workspace → Data → Data Sources
Connect your first data source. Supported types include relational databases (PostgreSQL, MySQL, Microsoft SQL Server), object storage, lake formats, file uploads, and Timeline Data Sources.
If you are new to the platform, a Timeline Data Source is the easiest starting point — it uses a simple CSV-like structure and requires no database connection.
For databases, you will need connection credentials and network access from your Kubernetes cluster to the database host.
See Data Sources for details.
Step 3 — Load the Data Source
Who: Administrator or Builder Where: Data Source → Operations tab
Run a Full Load on the data source. This copies the data from the source system into ClickHouse, where it is stored for analytics and semantic processing.
After the load completes, the data is available inside your Lakehousecat instance. The original source system is not queried again until you run another load.
Step 4 — Run Semantic Extraction
Who: Administrator or Builder Where: Data Source → Operations tab
Enable the semantic extraction toggles and trigger the process. This creates Job Definitions in the Operations section and starts a background Airflow process.
Semantic extraction is the intelligence layer that analyzes your data — its relationships, data types, content, structures, and hierarchies. The output is semantic metadata in PostgreSQL and supporting structures in ClickHouse that describe and classify the data source. This work is done once per data source and can be reused across multiple Custom Models.
Duration: Depends on the complexity of the data source (number of tables, columns, and relationships). A small data source may complete in a few minutes; a large, complex schema may take ~15 minutes or more.
See Semantic Extraction for a full explanation.
Step 5 — Create a Custom Model
Who: Administrator or Builder Where: Workspace → Models → Custom Models
Create a Custom Model and connect it to:
- One or more trained data sources (data sources that have completed semantic extraction)
- A Provider Model (the AI model used for chat queries in sessions)
The Custom Model defines the data scope and AI configuration for a group of users. Different combinations of data sources and filters produce different Custom Models — for example, a Sales model using the sales database, and an HR model using HR data only.
See Custom Models for details.
Step 6 — Run Semantic Model Creation
Who: Administrator or Builder Where: Custom Model → Operations tab
Trigger the semantic model creation process. This is a background Airflow process that builds the Star Views in ClickHouse — the denormalized analytical entities that sessions and charts query. No data is loaded at this stage; it is a pure semantic layer that combines the knowledge from the trained data sources.
Because the data source extraction work has already been done, Custom Model training is significantly faster — typically ~5 minutes. It must complete before users can run sessions against the model.
Step 7 — Open a Session and Ask a Question
Who: Any user with access to the Custom Model Where: Workspace → Sessions
Create a new session, select the Custom Model you configured, and ask your first analytical question in natural language. The system queries the semantic layer, generates SQL, executes it against ClickHouse, and returns a chart or answer.
What Kinds of Questions Can I Ask?
The questions you can ask depend entirely on the data your Custom Model was trained on. Lakehousecat does not have built-in knowledge of your domain — its intelligence comes from the semantic layer built from your data sources. Questions that match your data structure and terminology will produce accurate answers; questions outside the model's data scope will not.
Question types by chart intent
The LLM automatically selects the most appropriate chart type based on your question. You can also steer it explicitly:
| Intent | Example phrasing |
|---|---|
| Bar or column chart | "Show me X broken down by Y", "Compare X across categories" |
| Trend over time | "Trend of X over the last 6 months", "How has X changed quarter over quarter?" |
| Comparison | "Compare X vs Y", "Difference between A and B" |
| Ranking | "Top 10 by X", "Which Y has the highest Z?" |
| Map visualization | "Show X on a map by region" (requires Mapbox to be configured) |
| Summary / KPI | "What is the total X for this year?", "Current value of Y" |
Overriding the chart type
If the AI selects a chart type you don't want, you can override it directly in your question: "Show me the top 10 customers by revenue as a horizontal bar chart." The AI will follow the instruction.
A note on accuracy
All responses are probabilistic — they are generated by a language model, not deterministic SQL queries you write manually. More capable provider models generally produce more accurate results, but even the best models can occasionally misinterpret ambiguous questions or produce subtly incorrect SQL. Always validate critical results against your source data.
Two-Stage Architecture: Reuse and Composition
The setup above reveals the two-stage architecture at the core of Lakehousecat:
Stage 1 — Data Source level: Load + Semantic Extraction This work is done once per data source. The result — semantic metadata describing the data's structure, relationships, and meaning — is stored and reused as a building block.
Stage 2 — Custom Model level: Compose + Train Multiple Custom Models can be built from the same trained data sources. Each Custom Model must still be trained (to build its Star Views — the analytical entities charts query), but this is fast (~5 min) because it reuses the DS-level knowledge. No data is loaded at this stage.
Different combinations of data sources produce different Custom Models for different domains or user groups. Semantic extraction at the DS level is expensive — run it sparingly, and re-run only when the underlying data structure changes significantly.
Next Steps
- Semantic Extraction — understand what happens under the hood
- Sharing — share models and charts with your team
- Provider Models — manage AI provider configurations
- Administration — invite users and manage your deployment