Skip to main content
Version: Next

Descriptions

The Descriptions tab shows the semantic metadata generated for this data source during semantic extraction. It lists all extracted tables and their columns, along with the descriptions, data types, synonyms, and other attributes the AI uses to understand and query the data.


What Is Shown​

Each entry in the Descriptions tab represents a table in the datasource. Expanding a table shows its column-level details.

Table-Level Information​

FieldDescription
Table nameThe name of the table as it exists in the source system
Schema nameThe schema the table belongs to (for databases that have schemas)
DescriptionA human-readable description of what the table contains. Editable.
Sync statusWhether the table's metadata is current — reflects the state after the last semantic extraction
LockedWhen locked, the table's metadata is protected from being overwritten by future extractions
EnabledWhen disabled, the table is excluded from AI query generation

Column-Level Information​

For each column in a table:

FieldDescription
Column nameThe column name as it exists in the source
Data typeThe column's data type (e.g., VARCHAR, INTEGER, TIMESTAMP)
DescriptionA human-readable description of the column's business meaning. Editable.
SynonymsAlternative names for the column that the AI can recognize in natural language queries
Is primary keyWhether the column is a primary key

Declaring Measure Semantics​

For numeric columns, you can go beyond a text description and explicitly declare what the column means as a measure, instead of leaving it to inference from the description text.

Expand a column and set:

FieldDescription
RoleMeasure, Dimension, Identifier, Temporal, or Status — what kind of value this column holds
AggregationThe default aggregation for a measure column: Sum, Average, Count, Count Distinct, Min, Max, or None
UnitA short unit label, e.g. EUR, %, pcs
Display nameThe business name shown to users in place of the raw column name
SynonymsBusiness synonyms for the measure itself (in addition to the column-level synonyms above)

Click Apply to save the column's measure attributes. Each column saves independently — you don't need to re-save the whole table.

Once a column has been curated this way, its collapsed row shows a Curated badge together with the current aggregation in plain text (e.g. "Sum"), so you can see at a glance which columns in a wide table already have explicit measure semantics without expanding each one.

Columns you haven't curated keep using the aggregation the AI infers automatically — declaring measure semantics is an advanced option for cases where the automatic inference picks the wrong aggregation or a business term the AI should use instead.

Provenance​

Each column's measure semantics carry a provenance label — Manual (you or a previous editor curated it), LLM, or Annotator (set automatically during semantic extraction). Only manually curated columns are protected the way locked descriptions are — see Locking.


Curating Key Roles​

Expanding a column of a source table shows its Key role — Primary Key, Identifier, or none — together with where that role comes from:

OriginMeaning
DeclaredThe source database declares it (for example a primary key constraint). It cannot be changed here
CuratedYou set it
InferredLakehousecat derived it from the data profile and relationships
NameLakehousecat guessed it from the column name

If an inferred role is wrong, pick Primary Key, Identifier or No key in the key role editor and click Apply. No key also stops name-based guesses, so a column such as customer_id is then no longer treated as a key. Select Not curated to withdraw your choice and let the role be derived again. Curations survive re-extraction and schema scans.

The key role editor is only available on source tables, not on generated views. Views inherit curated roles after the next semantic update and model update — a banner reminds you of this after you apply a change.

Foreign keys are curated on the connection rather than the column: confirm the relationship in the Relationships tab.


Derived Measures​

Some business metrics are not a single column but a formula across several columns — for example, gross margin as revenue minus cost. The Derived Measures section, below the column list, lets you define these directly on a table's description.

For each derived measure, provide:

FieldDescription
NameThe measure's identifier, e.g. gross_margin
ExpressionAn aggregated SQL expression over the table's columns, e.g. SUM(extended_amount) - SUM(total_product_cost)
UnitA short unit label, e.g. EUR
SynonymsBusiness synonyms the AI can recognize
DescriptionA short business description of what the formula represents

Click + New to add a derived measure, and Save derived measures to persist the whole list. Saving replaces the full list, so make all the changes you want in one pass before saving.


Editing Descriptions​

You can manually edit any description or synonym directly in the Descriptions tab:

  1. Click the description field for a table or column.
  2. Edit the text.
  3. Save the change.

Enriching descriptions is one of the most impactful ways to improve query accuracy. The AI uses these descriptions to understand what each table and column means in business terms.

Examples of useful descriptions:

  • Table: orders → "Contains all sales orders placed by customers. Each row is one order."
  • Column: order_amount → "The total order value in EUR, excluding VAT. Negative values represent refunds."
  • Column: cust_id → "Foreign key to the customers table (customers.id). Synonym: customer_id, customer number."

Locking​

When a table or column description is locked, future semantic extractions will not overwrite it. Use locking to protect descriptions you have manually crafted after confirming they are accurate.


Relationship Hints​

Cross-table and cross-datasource relationship hints can be added in column descriptions or at the table level. Since the AI cannot automatically infer foreign key relationships, describing them explicitly improves join-based query generation.

Example: "orders.customer_id links to customers.customers.id"


Best Practices​

  • Run semantic extraction first — descriptions are generated automatically from the schema.
  • After extraction, review and enrich descriptions for tables and columns that are central to your analysis.
  • Lock descriptions you have carefully customized to prevent them from being reset.
  • Add synonyms for columns with technical or abbreviated names (e.g., amt → synonym: amount, value).