Skip to main content
Version: 0.0.42

Core Concepts

Lakehousecat is built on five fundamental concepts that work together to create an AI-powered analytics platform for structured data.

Data​

Foundation: Connect and integrate structured data sources

The Data layer is the entry point for all data in Lakehousecat. Connect relational databases, cloud object storage, or open table formats. Lakehousecat is designed for tabular, structured data — it analyzes rows and columns, not unstructured content such as documents, PDFs, or free text.

File uploads are supported for structured formats (CSV). Unstructured file types are not supported.

Supported source types:

  • Relational databases (PostgreSQL, MySQL, Microsoft SQL Server)
  • Object storage (Amazon S3)
  • Open table formats (Delta Lake, Apache Hudi, Apache Iceberg)
  • File upload (CSV)

Semantic​

Meaning: The bridge between raw data and natural language

The Semantic Layer sits between your data sources and the analytics engine. It captures the meaning of your data — table and column descriptions, business terminology, and dimensional hierarchies — so that the AI can generate accurate SQL from natural language questions.

The Semantic Layer is built automatically through Semantic Extraction, which reads the schema of a connected data source and generates descriptions using an LLM. Builders can review and refine these descriptions manually to improve accuracy.

Key capabilities:

  • Automated schema analysis and description generation
  • Business terminology and hierarchy modeling
  • Foundation for all natural language queries in Analytics

Star Views are the analytical output of Custom Model training. When a Custom Model is trained on one or more Data Sources, the system generates denormalized view structures in ClickHouse that connect data across the linked sources into a unified, queryable form. These Star Views are what sessions and charts actually query — they are the bridge between the semantic metadata and the SQL that the LLM generates.

Models​

Intelligence: Provider and Custom Models for your domain

Models power Lakehousecat's AI capabilities through two distinct types. Provider Models are the underlying LLM connections (OpenAI, Anthropic, Google, Azure, AWS Bedrock). Custom Models are built on top of a Provider Model and linked to one or more data sources with their Semantic Layer — making the model domain-specific and ready for natural language queries against your data.

System prompt customization allows Builders to further shape how the model responds.

Key capabilities:

  • Provider model integration (5 major LLM providers)
  • Custom model creation with data source binding
  • System prompt customization
  • Domain-specific model adaptation

Analytics​

Insights: Natural language queries and automated visualization

The Analytics component is where users interact with their data through natural language. Ask a question, and Lakehousecat queries the Semantic Layer to generate SQL, executes it against the data source, and returns an answer with an automatically generated chart or dashboard.

Key capabilities:

  • Natural language data queries
  • Automated chart and dashboard generation
  • Session-based conversational exploration
  • Chart sharing and embedding

Operations​

Automation: Data ingestion and workflow management

Operations manages the automated processes that keep your data current. Data sources are loaded incrementally or in full through scheduled or manually triggered jobs. The workflow engine (Apache Airflow) handles execution, dependency management, and monitoring. Builders and Administrators can define custom workflows using packaged task code.

Key capabilities:

  • Incremental and full-load data ingestion
  • Scheduled and manual job execution
  • Custom workflow definition with task packaging
  • Operational monitoring

Access​

Security: Role-based permissions and audit

The Access layer controls who can see and interact with every resource in Lakehousecat — data sources, models, charts, dashboards, sessions, and prompts. Permissions are managed at user and group level. All platform activity is recorded in an audit log.

Key capabilities:

  • Role-based access control (Administrator, Builder, User)
  • User and group management
  • Resource-level sharing and permissions
  • Audit logging for compliance

These concepts build on each other: Data provides the raw material, the Semantic Layer gives it meaning, Models connect AI to that meaning, Analytics turns questions into answers, Operations keeps data fresh, and Access ensures everything stays secure.