Data
Data Sources are the foundation of every analytics workflow in Lakehousecat. They define the connection to your data — whether it lives in a relational database, a cloud data lake, object storage, or as uploaded files. Once connected, the data becomes available for semantic extraction, query generation, and AI-assisted analysis.
Prerequisites
To create or manage data sources, you need one of the following roles:
- Admin
- Builder
Navigating to Data Sources
- Open the workspace.
- In the top navigation, click Data.
The Data page displays all data sources you own, with options to switch between your own, data sources shared by you, and data sources shared with you.
Datasource Types
Lakehousecat supports a range of data source types. Select the type that matches your data infrastructure:
| Type | Description |
|---|---|
| PostgreSQL | Connect to a PostgreSQL database via connection URI |
| MySQL | Connect to a MySQL database via connection URI |
| MSSQL | Connect to a Microsoft SQL Server database via connection URI |
| Delta Lake | Connect to a Delta Lake table stored in object storage (e.g., SeaweedFS, S3) |
| Apache Hudi | Connect to an Apache Hudi table stored in object storage |
| Apache Iceberg | Connect to an Apache Iceberg table stored in object storage |
| S3 File Storage | Connect to a CSV or Parquet file hosted in an S3-compatible bucket |
| File Upload | Upload files (CSV, PDF, DOCX, etc.) directly from your local machine |
| Timeline | Generate a virtual time dimension table for date-based analytics |
The list of supported data source types continues to grow. New connector types may be added in future releases.
Creating a Datasource
Step 1 — Select the Datasource Type
On the Data Sources page, click the + (Plus) button to open the creation form.
A type selection panel appears. Browse the available types and click the one that matches your data source. The icon for your chosen type appears at the top of the form.
Step 2 — Enter the Name
Enter a unique, descriptive Name for the data source. The name is used to identify the data source throughout the workspace, including in model configurations and operations.
Step 3 — Fill in Connection Details
The required fields depend on the selected type.
SQL Databases (PostgreSQL, MySQL, MSSQL)
| Field | Required | Description |
|---|---|---|
| Connection URI | Yes | Full connection string including credentials, host, port, and database name. Example: postgresql://user:password@host:5432/mydb |
| Use SSH | No | Enable if your database is behind an SSH bastion. When enabled, an SSH Settings field appears where you enter your SSH configuration as JSON. |
Delta Lake, Apache Hudi, Apache Iceberg
| Field | Required | Description |
|---|---|---|
| URL | Yes | The full URL of the table in object storage. Example: http://seaweedfs-s3:8333/deltalake/tablename |
| Access Key ID | Yes | The access key for object storage authentication |
| Secret Access Key | Yes | The secret key for object storage authentication |
S3 File Storage
| Field | Required | Description |
|---|---|---|
| URL | Yes | The full URL of the file in S3. Example: https://bucketname.s3.region.amazonaws.com/csv/filename.csv |
| Access Key ID | Yes | AWS access key ID |
| Secret Access Key | Yes | AWS secret access key |
| Format | Yes | The file format (e.g., CSV, PARQUET) |
| Compression | No | The compression type if applicable (e.g., gzip) |
File Upload
Click the Click here to select files button to open a file picker. You can select one or multiple tabular files. Supported formats are CSV and Excel (XLSX, XLS).
The files are uploaded when you click Create.
Timeline
The Timeline type generates a virtual date dimension table — a calendar table often used in BI tools for time-based analysis.
| Field | Required | Description |
|---|---|---|
| Start Date | Yes | The earliest date in the generated table |
| End Date | Yes | The latest date in the generated table |
| Granularity | No | The smallest time unit: Second, Minute, Hour, Day, Week, Month, Quarter, or Year (default: Day) |
| First Day of Week | No | Monday or Sunday |
| Weekend Days | No | Select which days count as weekend days |
| Is Active | No | Toggle to enable or disable the timeline datasource |
| Enable Fiscal Year Attributes | No | When enabled, add fiscal calendar fields to the table |
When fiscal year attributes are enabled, two additional fields appear:
| Field | Description |
|---|---|
| First Month of Fiscal Year | A number from 1 (January) to 12 (December) |
| Fiscal Variant | The fiscal week pattern, e.g., 4-4-5 |
Step 4 — Add a Description
The Description field is optional but recommended. A clear description helps the AI semantic layer understand the purpose of the data source and generate better queries.
You can type the description manually or use the microphone icon in the bottom-right corner of the description field to record a voice input, which is automatically transcribed and optimized.
Step 5 — Advanced Settings (Optional)
Click Show Advanced to expand the advanced configuration section. This allows you to override the default AI models used for semantic processing of this data source:
| Field | Description |
|---|---|
| Override Chat Model | Select a specific chat model for backend analysis of this datasource. If left empty, the system default is used. |
Step 6 — Create or Validate
Two action buttons appear at the bottom of the form:
- Validate — Creates the data source and immediately runs a semantic extraction validation. Use this to verify that the connection and configuration are correct before saving.
- Create — Creates the data source and opens the Edit view, where you can configure advanced settings such as schema filters, column filters, and data load schedules.
After clicking Create, you are automatically redirected to the datasource editor for further configuration.
The Datasource List
The Data Sources list displays all data sources for your account. Each entry shows the datasource type, name, status, owner, and optionally the chat model override.
Views
Toggle between two display modes using the icons in the toolbar:
- List View — A compact table layout
- Cards View — A tile-based grid layout
Search and Filter
- Use the search bar to filter data sources by name.
- Use the color tags in the search bar to filter by tag. Click a colored dot to activate or deactivate that filter.
Actions
Hover over a data source entry to reveal the available actions:
| Action | Description |
|---|---|
| Edit (pencil icon) | Opens the data source editor for advanced configuration |
| Clone | Creates a shallow copy of the data source with _clone appended to the name |
| Deep Clone | Creates a full copy including associated ClickHouse data |
| Share | Opens the sharing dialog to grant access to users or groups |
| Delete | Permanently removes the data source |
After Creation: The Datasource Editor
After creating a data source, you are automatically taken to the editor. The editor contains multiple tabs, including:
- Connection Settings — Review and update the connection configuration
- Filters — Define schema, table, or column-level filters to limit which data is included in semantic extraction
- Operations — Trigger or schedule semantic extraction jobs
- Sharing — Manage access permissions
Filters are particularly important for large datasets. By filtering down to only the relevant schemas, tables, or columns, you improve the quality of AI-generated queries and reduce processing time.
Next Steps
Once your data source is created and configured, the next step is to trigger the semantic extraction in the Operations tab. This process analyzes the data structure and creates the semantic layer that the AI uses to understand and query your data.
After that, you can link the data source to a Custom Model to make it available in the chat interface.