Skip to main content
Version: Next

Troubleshooting

This page covers the most common issues encountered when running Lakehousecat and how to diagnose them.


Pods Not Starting​

Symptoms: One or more pods are stuck in Pending, CrashLoopBackOff, or Error state after deployment or upgrade.

Diagnose:

# Describe the pod to see events and error messages
kubectl describe pod <pod-name> -n <namespace>

# View pod logs
kubectl logs <pod-name> -n <namespace>

# Check Operator logs
kubectl logs -f deployment/lhc-operator -n lhc-operator

Common causes:

CauseSignsResolution
Insufficient cluster resourcesInsufficient cpu or Insufficient memory in pod eventsIncrease node capacity or reduce resource requests in the Operator config
Image pull errorsImagePullBackOff, ErrImagePull in pod eventsCheck network access to AWS ECR; verify the Operator has image pull credentials
Configuration errorsError or CreateContainerConfigError in pod eventsCheck secrets and ConfigMaps referenced by the pod; verify the Lakehousecat CR spec

Kubernetes-level pod issues are the responsibility of the Kubernetes administrator. Lakehousecat support can assist with application-level issues once pods are running.


Semantic Extraction Fails​

Symptoms: Semantic extraction job does not complete, times out, or shows errors in Job Runs.

Where to look:

Navigate to Workspace → Operations → Job Runs and select the failed job run to view the Airflow task logs.

Common causes:

CauseSignsResolution
Airflow worker resource limitsTasks marked as killed or timeout errors in logsIncrease Airflow worker CPU and/or memory in the Operator configuration; increase the job timeout
Poor data qualityErrors or unexpected output in extraction logsImprove column names and add descriptions to the data source before re-running
Data source not fully loadedExtraction runs but produces poor resultsEnsure a full load completed successfully before triggering extraction

See Semantic Extraction for best practices and a full explanation of what can go wrong.


Charts Show Unexpected Data​

Symptoms: A chart displays data that seems incorrect, incomplete, or does not match expectations.

What to consider:

LLM outputs are probabilistic — the model generates SQL based on its semantic understanding of the data, which may not always match what you intended. This is a normal characteristic of LLM-based systems, not a bug.

Steps to improve results:

  1. Evaluate the chart critically — use the Load Chart Data button in the chart detail view to inspect the underlying data and SQL
  2. Improve data quality at the source — clean column names, consistent values, and meaningful content improve semantic understanding
  3. Add descriptions to tables and columns in the Data Source configuration — these are used by the semantic layer to improve query generation
  4. Re-run semantic extraction after making significant improvements to data quality or descriptions
  5. Rephrase your question — more specific language often yields better results

API Key Errors from LLM Provider​

Symptoms: Sessions fail with provider errors; the system reports API authentication failures.

Diagnose:

Check the backend logs for detailed error messages. Logs are stored in SeaweedFS at:

backend_logs/

Access via the SeaweedFS console or kubectl exec into the SeaweedFS pod.

Common causes:

CauseResolution
Key expired or rotatedUpdate the API key in Workspace → Models → Provider Models
Incorrect key formatVerify the key format matches what the provider expects
Rate limits exceededCheck your provider dashboard for rate limit status; consider upgrading your provider tier
Regional restrictionsSome providers restrict API access by region; verify your cluster's egress IP is allowed

Login Fails After Restore​

Symptoms: Users cannot log in after a restore operation; authentication errors in the backend.

Most likely cause: The encryption key (stored in Kubernetes secrets) was not restored before restoring the database data. The encrypted values in the database cannot be decrypted with a different key — causing authentication and configuration errors.

Resolution:

  1. Verify that secrets were restored before database data (see Restore for the correct restore order)
  2. If the restore order was incorrect, restore from backup again in the correct sequence: secrets first, then PVCs
  3. Check backend logs for specific decryption or authentication error messages

Still stuck? Open a support ticket via portal.lakehousecat.com — available on every tier, including FREE. Include the backend log output to speed up diagnosis. Paid tiers receive prioritized handling; see Support by Subscription Tier.


Getting More Help​