Skip to main content
Version: 0.0.37

Documents in Datasources

Managing documents within Lakehousecat's datasources is crucial for organizing and processing different data types effectively. This guide outlines how document handling is integrated, focusing on upload processes and document interaction.

Document Support in Datasources​

Not all datasources support document integration. Documents are typically generated during specific upload processes:

  • File Uploads: Documents can be created when files are uploaded directly to the datasource.
  • S3 Datasource: When handling S3 datasources, documents are generated if case-based uploads occur.

Document Processing​

Depending on the nature and structure of documents, they are processed as follows:

  • Structured Documents: Formats like CSV are loaded directly into the analytical pool, where they can be utilized for structured data analysis.
  • Semi-Structured/Unstructured Documents: These are uploaded into the vector database and made accessible via model interaction, enabling nuanced processing and analysis.

Document Management Features​

Within the document management area, users have the capability to:

  • View Documents: Access a list of all documents associated with the datasource, including those uploaded automatically.
  • Tenant Data Retrieval: Gather data pertinent to specific tenants or organizational units.
  • Delete Documents: Remove documents that are no longer needed or relevant to ensure data consistency and accuracy.

Automated Document Availability​

All documents linked to the datasource through automated processes are readily available in this section, and they are considered during user queries and interactions.

Best Practices​

  • Utilize Structure: Prefer structured documents for regular data analysis to ensure compatibility and ease of use.
  • Manage Unstructured Content: Leverage semi-structured and unstructured documents for complex analytical tasks, enhancing model interactions.
  • Regular Housekeeping: Periodically delete irrelevant documents to maintain data quality and system efficiency.

By understanding and managing document features within Lakehousecat, users can optimize data processes, driving effective analysis and informed decision-making.