Performance
Lakehousecat's performance is primarily determined by three factors: the infrastructure available in the customer's Kubernetes cluster, how data sources are configured, and the LLM provider in use.
Infrastructure Scaling
Every Lakehousecat service is independently scalable via the Operator configuration. Horizontal scaling (additional replicas) and vertical scaling (increased CPU/memory) are both supported. See Scaling for details.
Performance bottlenecks are typically resolved by scaling the relevant service. Resource-intensive operations — data loading, semantic layer creation, backup jobs — run as isolated Kubernetes pods and do not block user-facing services.
Data Management
The volume and quality of data loaded into Lakehousecat directly affects query performance and resource consumption:
- Load only what is needed — loading large datasets that are not relevant to the analytical use case increases processing time and resource usage
- Use incremental loads — prefer incremental datasource loads over full reloads when only new or updated records are needed
- Filter at the source — narrow the data set as early as possible (in the source query or connector configuration) rather than loading everything and filtering later
Builders and Administrators are responsible for deciding which data sources to connect and how to scope them.
LLM Provider Performance
Response times for AI-powered features (chat, analytics queries) depend on the configured LLM provider:
- Model complexity — reasoning-oriented models take longer to respond than faster lightweight models
- Provider infrastructure geography — latency varies depending on where the provider's inference infrastructure is located relative to the deployment
- Rate limits and queuing — at high usage, provider-side rate limiting can increase wait times
- Provider outages — LLM-dependent features are unavailable if the configured provider has an outage
Customers can configure multiple providers and models, and select the most appropriate one per use case. See Provider Models for details.
Monitoring
Lakehousecat services produce structured logs that are written to the configured object storage. These logs can be ingested into monitoring tools such as Grafana + Loki, Datadog, Elastic Stack, or any S3-compatible log consumer.
Lakehousecat does not include a built-in performance monitoring dashboard. Infrastructure-level metrics (pod CPU/memory, node utilization) are available through standard Kubernetes monitoring tools (kubectl top, Prometheus, cloud provider dashboards).