Filters
Filters are a powerful feature within Lakehousecat's Datasource management, enabling precise control over which data is loaded and analyzed. This guide details the filter options available to optimize data sources for efficient use.
Importance of Filtering
When modeling data sources, it's essential to define which data is relevant for analysis and user interaction. Given the often extensive nature of data sources, utilizing filters helps focus on pertinent data while ignoring the rest.
Types of Filters
Schema-Level Filtering
- Schema Selection: If your datasource contains multiple schemas, initial filtering can be done at the schema level by selecting specific schemas that contain relevant data for analysis.
Table-Level Filtering
- Table Selection: Further refinement occurs within schemas by filtering on tables. Select specific tables from chosen schemas to narrow down the scope of data analysis.
Column-Level Filtering
- Column Selection: Enhance granularity by filtering on columns within tables. Focus only on columns essential to your analysis, ignoring irrelevant data points.
Row-Level Filtering
- Detailed Data Segmentation: Row-level filtering allows for highly flexible and detailed data selection. Define conditions to include qualitative distinctions within data sets.
Filtering Strategies
- Include: Explicitly pick schemas, tables, or columns to include, ensuring that only selected data is considered.
- Exclude: Opt to exclude certain schemas, tables, or columns, effectively disregarding them during data loading and analysis.
Row-Level Security
Utilize row-level security by applying diverse filter options to the same datasource. This ensures tailored data access, providing specific data views to different user groups.
Note: Schema Filters are only available with datasources that support schema-level filtering.
Benefits of Using Filters
Applying filters strategically ensures:
- Granular Control: Cover schema-, table-, column-, and row-level use-cases effectively.
- Efficient Data Use: Load only relevant data, enhancing the economic use of resources within the Lakehousecat framework.
By setting datasource filters appropriately, users can achieve precise data management, facilitating meaningful analysis while optimizing resource utilization.