Skip to main content
Version: 0.0.41

Monitoring

Scope of the Lakehousecat Operator​

The Lakehousecat Operator is designed to operate within a single Kubernetes namespace and requires only namespace-scoped RBAC permissions. The Operator can deploy Grafana and Prometheus into the target namespace as optional components, providing a baseline for metrics collection within that namespace.

However, cluster-level monitoring is outside the scope of the Operator. Most customers already have an established monitoring infrastructure — Datadog, AWS CloudWatch, Google Cloud Monitoring, Grafana with a cluster-wide Prometheus installation, or another solution. Lakehousecat is designed to integrate into whatever system is already in place rather than replacing it.

Customers are responsible for cluster-level monitoring integration

The Operator manages the application namespace. Log collection, forwarding, and retention at the cluster or infrastructure level must be configured by the customer's operations team according to their own monitoring standards and tooling.


Application Logs​

Backend Service Logs​

All Lakehousecat backend services — including the API server, Airflow workers, and other platform components — produce structured application logs. These logs are:

  • Written to stdout/stderr (standard Kubernetes log output, accessible via kubectl logs)
  • Continuously buffered and flushed to SeaweedFS at regular intervals (every 60 seconds, or when the buffer reaches 1,000 records)

Log Storage in SeaweedFS​

Logs are stored in the internal SeaweedFS instance under the following path structure:

backend_logs/logs/<timestamp>_<uuid>.log

Each log file uses the format:

<timestamp> - <logger_name> - <level> - <message>

Log files are generated on a rolling basis — a new file is created every flush cycle or on service shutdown. This means logs are always available in SeaweedFS for inspection and support, even without an external log aggregation system connected.

Airflow / Job Logs​

Task execution logs from Airflow jobs (data loads, semantic extraction, custom tasks) are also stored in SeaweedFS as part of the job run output. These are accessible from the Job Runs detail view in the Operations area of the workspace.


Audit Logs​

In addition to application logs, Lakehousecat maintains structured audit logs for all user-initiated changes to platform entities:

EntityLogged events
Data SourcesCreated, updated, filters changed, operations triggered
Custom ModelsCreated, configured, semantic model operations, shared
ChartsCreated, edited, shared
DashboardsCreated, edited, shared
Job DefinitionsCreated, updated, triggered
UsersInvited, role changed, removed

Audit log entries are available within the application (each entity's Audit Log tab) and are also persisted to SeaweedFS as part of the structured log output.


Connecting Your Monitoring System​

Lakehousecat produces logs through two channels that your monitoring system should consume:

1. Kubernetes Pod Logs (stdout/stderr)​

Standard Kubernetes log output is available via:

kubectl logs -n <namespace> <pod-name>
kubectl logs -n <namespace> -l app=<label> --all-containers

Most monitoring agents (Datadog Agent, Fluent Bit, Logstash, AWS CloudWatch Agent, etc.) can be configured to collect pod logs from the namespace automatically. Configure your log collector to forward logs from the Lakehousecat namespace to your central log management system.

2. SeaweedFS Log Files​

Logs stored in SeaweedFS under backend_logs/ are available for retrieval via the SeaweedFS S3 API or the SeaweedFS Filer UI. You can:

  • Connect an S3-compatible log sink or ETL pipeline to read from the SeaweedFS bucket
  • Schedule periodic exports to your log management system
  • Access log files directly for incident investigation or support

SeaweedFS connection details (endpoint, access key, bucket name) are available in the deployment configuration under the seaweedfs: section of the Custom Resource.


Grafana and Prometheus (Optional)​

If no existing metrics infrastructure is available, the Operator can deploy Grafana and Prometheus into the application namespace. This provides:

  • Kubernetes resource metrics for pods in the namespace (CPU, memory, restarts)
  • A baseline Grafana dashboard for operational visibility

This is a lightweight, namespace-scoped setup and is not a replacement for a cluster-wide monitoring solution. For production environments, integrating with an existing cluster-level monitoring stack is recommended.


Recommendations​

TaskRecommendation
Collect pod logsConfigure your log agent to scrape the Lakehousecat namespace
Retain application logsSet a retention policy on the SeaweedFS backend_logs/ bucket
Alert on errorsForward ERROR-level log entries to your alerting system
Audit complianceUse the per-entity Audit Log tabs in the UI; raw logs available in SeaweedFS
Job execution monitoringUse the Job Runs section in the Operations area for task-level status