Descriptions
The Descriptions tab shows the semantic metadata generated for this data source during semantic extraction. It lists all extracted tables and their columns, along with the descriptions, data types, synonyms, and other attributes the AI uses to understand and query the data.
What Is Shown
Each entry in the Descriptions tab represents a table in the datasource. Expanding a table shows its column-level details.
Table-Level Information
| Field | Description |
|---|---|
| Table name | The name of the table as it exists in the source system |
| Schema name | The schema the table belongs to (for databases that have schemas) |
| Description | A human-readable description of what the table contains. Editable. |
| Sync status | Whether the table's metadata is current — reflects the state after the last semantic extraction |
| Locked | When locked, the table's metadata is protected from being overwritten by future extractions |
| Enabled | When disabled, the table is excluded from AI query generation |
Column-Level Information
For each column in a table:
| Field | Description |
|---|---|
| Column name | The column name as it exists in the source |
| Data type | The column's data type (e.g., VARCHAR, INTEGER, TIMESTAMP) |
| Description | A human-readable description of the column's business meaning. Editable. |
| Synonyms | Alternative names for the column that the AI can recognize in natural language queries |
| Is primary key | Whether the column is a primary key |
Declaring Measure Semantics
For numeric columns, you can go beyond a text description and explicitly declare what the column means as a measure, instead of leaving it to inference from the description text.
Expand a column and set:
| Field | Description |
|---|---|
| Role | Measure, Dimension, Identifier, Temporal, or Status — what kind of value this column holds |
| Aggregation | The default aggregation for a measure column: Sum, Average, Count, Count Distinct, Min, Max, or None |
| Unit | A short unit label, e.g. EUR, %, pcs |
| Display name | The business name shown to users in place of the raw column name |
| Synonyms | Business synonyms for the measure itself (in addition to the column-level synonyms above) |
Click Apply to save the column's measure attributes. Each column saves independently — you don't need to re-save the whole table.
Once a column has been curated this way, its collapsed row shows a Curated badge together with the current aggregation in plain text (e.g. "Sum"), so you can see at a glance which columns in a wide table already have explicit measure semantics without expanding each one.
Columns you haven't curated keep using the aggregation the AI infers automatically — declaring measure semantics is an advanced option for cases where the automatic inference picks the wrong aggregation or a business term the AI should use instead.
Provenance
Each column's measure semantics carry a provenance label — Manual (you or a previous editor curated it), LLM, or Annotator (set automatically during semantic extraction). Only manually curated columns are protected the way locked descriptions are — see Locking.
Curating Key Roles
Expanding a column of a source table shows its Key role — Primary Key, Identifier, or none — together with where that role comes from:
| Origin | Meaning |
|---|---|
| Declared | The source database declares it (for example a primary key constraint). It cannot be changed here |
| Curated | You set it |
| Inferred | Lakehousecat derived it from the data profile and relationships |
| Name | Lakehousecat guessed it from the column name |
If an inferred role is wrong, pick Primary Key, Identifier or No key in the key role editor and click Apply. No key also stops name-based guesses, so a column such as customer_id is then no longer treated as a key. Select Not curated to withdraw your choice and let the role be derived again. Curations survive re-extraction and schema scans.
The key role editor is only available on source tables, not on generated views. Views inherit curated roles after the next semantic update and model update — a banner reminds you of this after you apply a change.
Foreign keys are curated on the connection rather than the column: confirm the relationship in the Relationships tab.
Derived Measures
Some business metrics are not a single column but a formula across several columns — for example, gross margin as revenue minus cost. The Derived Measures section, below the column list, lets you define these directly on a table's description.
For each derived measure, provide:
| Field | Description |
|---|---|
| Name | The measure's identifier, e.g. gross_margin |
| Expression | An aggregated SQL expression over the table's columns, e.g. SUM(extended_amount) - SUM(total_product_cost) |
| Unit | A short unit label, e.g. EUR |
| Synonyms | Business synonyms the AI can recognize |
| Description | A short business description of what the formula represents |
Click + New to add a derived measure, and Save derived measures to persist the whole list. Saving replaces the full list, so make all the changes you want in one pass before saving.
Editing Descriptions
You can manually edit any description or synonym directly in the Descriptions tab:
- Click the description field for a table or column.
- Edit the text.
- Save the change.
Enriching descriptions is one of the most impactful ways to improve query accuracy. The AI uses these descriptions to understand what each table and column means in business terms.
Examples of useful descriptions:
- Table:
orders→ "Contains all sales orders placed by customers. Each row is one order." - Column:
order_amount→ "The total order value in EUR, excluding VAT. Negative values represent refunds." - Column:
cust_id→ "Foreign key to the customers table (customers.id). Synonym: customer_id, customer number."
Locking
When a table or column description is locked, future semantic extractions will not overwrite it. Use locking to protect descriptions you have manually crafted after confirming they are accurate.
Relationship Hints
Cross-table and cross-datasource relationship hints can be added in column descriptions or at the table level. Since the AI cannot automatically infer foreign key relationships, describing them explicitly improves join-based query generation.
Example: "orders.customer_id links to customers.customers.id"
Best Practices
- Run semantic extraction first — descriptions are generated automatically from the schema.
- After extraction, review and enrich descriptions for tables and columns that are central to your analysis.
- Lock descriptions you have carefully customized to prevent them from being reset.
- Add synonyms for columns with technical or abbreviated names (e.g.,
amt→ synonym:amount,value).