Skip to main content
Version: Next

Data

Data Sources are the foundation of every analytics workflow in Lakehousecat. They define the connection to your data — whether it lives in a relational database, a cloud data lake, object storage, or as uploaded files. Once connected, the data becomes available for semantic extraction, query generation, and AI-assisted analysis.


Prerequisites​

To create or manage data sources, you need one of the following roles:

  • Admin
  • Builder

  1. Open the workspace.
  2. In the top navigation, click Data.

The Data page displays all data sources you own, with options to switch between your own, data sources shared by you, and data sources shared with you.


Datasource Types​

Lakehousecat supports a range of data source types. Select the type that matches your data infrastructure. For the current list of supported types and their details, see Datasource Types.


Creating a Datasource​

Step 1 — Select the Datasource Type​

On the Data Sources page, click the + (Plus) button to open the creation form.

A type selection panel appears. Browse the available types and click the one that matches your data source. The icon for your chosen type appears at the top of the form.

Step 2 — Enter the Name​

Enter a unique, descriptive Name for the data source. The name is used to identify the data source throughout the workspace, including in model configurations and operations.

Step 3 — Fill in Connection Details​

The required fields depend on the selected type.

SQL Databases (PostgreSQL, MySQL, MSSQL)​

FieldRequiredDescription
Connection URIYesFull connection string including credentials, host, port, and database name. Example: postgresql://user:password@host:5432/mydb
Use SSHNoEnable if your database is behind an SSH bastion. When enabled, an SSH Settings field appears where you enter your SSH configuration as JSON.

Delta Lake, Apache Hudi, Apache Iceberg​

FieldRequiredDescription
URLYesThe full URL of the table in object storage. Example: http://seaweedfs-s3:8333/deltalake/tablename
Access Key IDYesThe access key for object storage authentication
Secret Access KeyYesThe secret key for object storage authentication

S3 File Storage​

FieldRequiredDescription
URLYesThe full URL of the file in S3. Example: https://bucketname.s3.region.amazonaws.com/csv/filename.csv
Access Key IDYesAWS access key ID
Secret Access KeyYesAWS secret access key
FormatYesThe file format (e.g., CSV, PARQUET)
CompressionNoThe compression type if applicable (e.g., gzip)

File Upload​

Click the Click here to select files button to open a file picker. You can select one or multiple tabular files. Supported formats are CSV and Excel (XLSX, XLS).

The files are uploaded when you click Create.

Timeline​

The Timeline type generates a virtual date dimension table — a calendar table often used in BI tools for time-based analysis.

FieldRequiredDescription
Start DateYesThe earliest date in the generated table
End DateYesThe latest date in the generated table
GranularityNoThe smallest time unit: Second, Minute, Hour, Day, Week, Month, Quarter, or Year (default: Day)
First Day of WeekNoMonday or Sunday
Weekend DaysNoSelect which days count as weekend days
Is ActiveNoToggle to enable or disable the timeline datasource
Enable Fiscal Year AttributesNoWhen enabled, add fiscal calendar fields to the table

When fiscal year attributes are enabled, two additional fields appear:

FieldDescription
First Month of Fiscal YearA number from 1 (January) to 12 (December)
Fiscal VariantThe fiscal week pattern, e.g., 4-4-5

Step 4 — Add a Description​

The Description field is optional but recommended. A clear description helps the AI semantic layer understand the purpose of the data source and generate better queries.

You can type the description manually or use the microphone icon in the bottom-right corner of the description field to record a voice input, which is automatically transcribed and optimized.

Step 5 — Advanced Settings (Optional)​

Click Show Advanced to expand the advanced configuration section. This allows you to override the default AI models used for semantic processing of this data source:

FieldDescription
Override Chat ModelSelect a specific chat model for backend analysis of this datasource. If left empty, the system default is used.

Step 6 — Create or Validate​

Two action buttons appear at the bottom of the form:

  • Validate — Creates the data source and immediately runs a semantic extraction validation. Use this to verify that the connection and configuration are correct before saving.
  • Create — Creates the data source and opens the Edit view, where you can configure advanced settings such as schema filters, column filters, and data load schedules.

After clicking Create, you are automatically redirected to the datasource editor for further configuration.


The Datasource List​

The Data Sources list displays all data sources for your account. Each entry shows the datasource type, name, status, owner, and optionally the chat model override.

Views​

Toggle between two display modes using the icons in the toolbar:

  • List View — A compact table layout
  • Cards View — A tile-based grid layout

Search and Filter​

  • Use the search bar to filter data sources by name.
  • Use the color tags in the search bar to filter by tag. Click a colored dot to activate or deactivate that filter.

Actions​

Hover over a data source entry to reveal the available actions:

ActionDescription
Edit (pencil icon)Opens the data source editor for advanced configuration
CloneCreates a shallow copy of the data source with _clone appended to the name
Deep CloneCreates a full copy including associated ClickHouse data
ShareOpens the sharing dialog to grant access to users or groups
DeletePermanently removes the data source

After Creation: The Datasource Editor​

After creating a data source, you are automatically taken to the editor. The editor contains multiple tabs, including:

  • Connection Settings — Review and update the connection configuration
  • Filters — Define schema, table, or column-level filters to limit which data is included in semantic extraction
  • Operations — Trigger or schedule semantic extraction jobs
  • Sharing — Manage access permissions

Filters are particularly important for large datasets. By filtering down to only the relevant schemas, tables, or columns, you improve the quality of AI-generated queries and reduce processing time.


Next Steps​

Once your data source is created and configured, the next step is to trigger the semantic extraction in the Operations tab. This process analyzes the data structure and creates the semantic layer that the AI uses to understand and query your data.

After that, you can link the data source to a Custom Model to make it available in the chat interface.