Skip to main content
Version: 0.0.38

Parameters

Understanding and configuring parameters within Lakehousecat's datasource settings is key to optimizing semantic processing and vector creation. This guide details the available settings for parameter customization and enhancement.

Accessing Model Settings​

To access parameter settings for your datasource, follow these navigation steps:

  1. Workspace Navigation

    • Start in your Lakehousecat workspace.
  2. Data Section

    • Click on 'Data' to explore available data configurations.
  3. Custom Model Selection

    • Select your Custom Model and proceed to the Parameter settings menu.

Parameter Configuration Options​

Model Optimization​

  • Semantic Enhancement: Parameters in the Model Settings allow for optimization of semantics concerning a specific datasource.
  • Vector Creation: Customize vector generation based on specific project needs.
  • Optional Model Selection: If desired, substitute the default backend model with an alternative LLM or embedding model tailored to the datasource.

Recommendation: Generally, retain default backend models unless specific customization is required. Proper configuration of Lakehousecat CAD ensures default models are equipped for optimal performance.

Retrieval Settings (RAG Settings)​

  • RAG Customization: Adjust RAG and RAG-Template settings for tailored data retrieval.
  • Parameters Included:
    • Chunk-Size: Determines the size of data chunks during processing.
    • Chunk-Overlap: Defines overlap between chunks to manage data continuity.
    • Top-K-Results: Sets how many top results are retrieved, impacting vector accuracy.
  • OCR PDF Extraction: Control PDF data extraction through OCR settings, ensuring text is accurately captured.

Best Practices​

  • Default Values: In most cases, maintain the default values for reliable and efficient parameter functions.
  • Customization: Opt for specific adjustments when unique data challenges or project requirements dictate necessity.

By using these settings judiciously, users can refine data-source processes, enhancing semantic interactions and vector management within the Lakehousecat environment.