Skip to main content
Version: 0.0.36

Operations

Datasource Operations within Lakehousecat enable powerful data manipulation and management, facilitating various processes from data loading to semantic structuring. This guide outlines available operations and provides recommended practices for efficient use.

To access Datasource Operations, follow these steps:

  1. Workspace Access

    • Navigate to your Lakehousecat workspace.
  2. Data Section

    • Click on 'Data' to explore datasource options.
  3. Select Datasources

    • Choose your desired datasource and proceed to the Operations tab in the left sidebar.

Available Operations​

Full Data Load​

  • Purpose: Perform a complete data load according to predefined filters. This process is crucial for initializing new datasources with comprehensive data sets.
  • Recommendations: Conduct Full Load once, especially if dealing with large volumes of data. Adjust job parameters like port size and timeouts as needed.

Incremental Data Load​

  • Purpose: Focus on updates by loading only changed data. This operation keeps data current without reloading the full dataset.
  • Usage: Activate Incremental Load following Full Load to ensure continued data currency.

Create Semantic Layer​

  • Purpose: Establish semantic layers on the datasource level, enhancing data interpretation and interaction.
  • Procedure: Pre-load the data completely before creating this layer. Typically a one-time setup.

Update Semantic Layer​

  • Purpose: Synchronize structures when schema changes occur at the datasource level.
  • Operationalization: Recommended whenever data schema adjustments are made to keep semantic layers reflective of current structures.

Delete Semantic Layer​

  • Purpose: Remove semantic layers in instances of testing or re-modeling processes.
  • Application: Useful for experimentation with different configurations and parameters.

Clear Data​

  • Purpose: Remove all loaded data. Allows for datasource reconfiguration or a fresh start.
  • Context: Ideal for resetting data settings and operations.

Typical Workflow​

Follow these steps for effective datasource operations management:

  1. Initial Setup: Conduct a Full Data Load.
  2. Continuous Updates: Enable Incremental Data Load for ongoing changes.
  3. Semantic Structuring: Create semantic layers one-time early in the process.
  4. Maintenance: Use Update Semantic Layer as schemas evolve.

Each operation requires a specific job definition, configurable by those with Builder or Admin roles. This includes setting schedules and other operational parameters.

By adhering to these guidelines, users can optimize datasource operations within Lakehousecat, ensuring structured, efficient, and dynamic data management.