Monitoring
Scope of the Lakehousecat Operator
The Lakehousecat Operator is designed to operate within a single Kubernetes namespace and requires only namespace-scoped RBAC permissions. The Operator deploys the application services and exposes the data customers need for observability — structured logs in object storage and Prometheus-format /metrics endpoints on each FastAPI service.
The Operator does not deploy a monitoring stack of its own. Cluster-level and infrastructure monitoring is the responsibility of the customer's operations team. Most customers already have an established monitoring infrastructure — Datadog, AWS CloudWatch, Google Cloud Monitoring, Grafana with a cluster-wide Prometheus installation, or another solution — and Lakehousecat is designed to integrate into whatever system is already in place rather than replacing it.
The Operator manages the application namespace. Log collection, forwarding, and retention at the cluster or infrastructure level must be configured by the customer's operations team according to their own monitoring standards and tooling.
Application Logs
Backend Service Logs
All Lakehousecat backend services — including the API server, Airflow workers, and other platform components — produce structured application logs. These logs are:
- Written to stdout/stderr (standard Kubernetes log output, accessible via
kubectl logs) - Continuously buffered and flushed to SeaweedFS at regular intervals (every 60 seconds, or when the buffer reaches 1,000 records)
Log Storage in SeaweedFS
Logs are stored in the internal SeaweedFS instance under the following path structure:
backend_logs/logs/<timestamp>_<uuid>.log
Each log file uses the format:
<timestamp> - <logger_name> - <level> - <message>
Log files are generated on a rolling basis — a new file is created every flush cycle or on service shutdown. This means logs are always available in SeaweedFS for inspection and support, even without an external log aggregation system connected.
Airflow / Job Logs
Task execution logs from Airflow jobs (data loads, semantic extraction, custom tasks) are also stored in SeaweedFS as part of the job run output. These are accessible from the Job Runs detail view in the Operations area of the workspace.
Audit Logs
In addition to application logs, Lakehousecat maintains structured audit logs for all user-initiated changes to platform entities:
| Entity | Logged events |
|---|---|
| Data Sources | Created, updated, filters changed, operations triggered |
| Custom Models | Created, configured, semantic model operations, shared |
| Charts | Created, edited, shared |
| Dashboards | Created, edited, shared |
| Job Definitions | Created, updated, triggered |
| Users | Invited, role changed, removed |
Audit log entries are available within the application (each entity's Audit Log tab) and are also persisted to SeaweedFS as part of the structured log output.
Connecting Your Monitoring System
Lakehousecat produces logs through two channels that your monitoring system should consume:
1. Kubernetes Pod Logs (stdout/stderr)
Standard Kubernetes log output is available via:
kubectl logs -n <namespace> <pod-name>
kubectl logs -n <namespace> -l app=<label> --all-containers
Most monitoring agents (Datadog Agent, Fluent Bit, Logstash, AWS CloudWatch Agent, etc.) can be configured to collect pod logs from the namespace automatically. Configure your log collector to forward logs from the Lakehousecat namespace to your central log management system.
2. SeaweedFS Log Files
Logs stored in SeaweedFS under backend_logs/ are available for retrieval via the SeaweedFS S3 API or the SeaweedFS Filer UI. You can:
- Connect an S3-compatible log sink or ETL pipeline to read from the SeaweedFS bucket
- Schedule periodic exports to your log management system
- Access log files directly for incident investigation or support
SeaweedFS connection details (endpoint, access key, bucket name) are available in the deployment configuration under the seaweedfs: section of the Custom Resource.
Application Metrics — Prometheus /metrics Endpoint
Each FastAPI backend service (lhc, llm, analytics, semantic, audio) exposes a /metrics endpoint in the standard Prometheus exposition format. This endpoint is intended for scraping by your existing monitoring system (Prometheus, Datadog Agent, Grafana Mimir, AWS Managed Prometheus, etc.).
What is exposed
- HTTP request metrics — request counts and latency histograms per route template, method, and status code
- Python process metrics — resident memory, CPU seconds, GC statistics, open file descriptors
Path labels are normalized to route templates (for example /api/v1/charts/{id} instead of concrete IDs) so cardinality stays bounded and no user-identifying values leak into labels.
Endpoint properties
| Property | Value |
|---|---|
| Path | /metrics |
| Format | Prometheus text exposition (OpenMetrics-compatible) |
| Authentication | None by design — follows the Prometheus ecosystem convention |
| Access control | Cluster-internal only (ClusterIP service + your NetworkPolicy) |
| Toggle | CRD field monitoring.metricsEnabled (default true) |
To disable the endpoints, set monitoring.metricsEnabled: false in the Lakehousecat Custom Resource. The Operator propagates the flag as the environment variable LHC_METRICS_ENABLED into each service.
Example: ServiceMonitor for Prometheus Operator
Customers running prometheus-operator can scrape Lakehousecat services with a ServiceMonitor such as:
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: lakehousecat-services
namespace: <your-lakehousecat-namespace>
spec:
selector:
matchLabels:
app.kubernetes.io/part-of: lakehousecat
endpoints:
- port: http
path: /metrics
interval: 30s
Datadog, New Relic and other agents support equivalent auto-discovery via Kubernetes service annotations — consult your agent's documentation.
Cluster-level Metrics
Pod- and node-level metrics (CPU, memory, restarts, kubelet stats) are produced by Kubernetes itself via kube-state-metrics, cAdvisor, and the kubelet — the customer's cluster monitoring tooling collects these directly. The Operator does not duplicate this.
Recommendations
| Task | Recommendation |
|---|---|
| Collect pod logs | Configure your log agent to scrape the Lakehousecat namespace |
| Retain application logs | Set a retention policy on the SeaweedFS backend_logs/ bucket |
| Collect application metrics | Scrape /metrics on each backend service with your existing Prometheus / Datadog / etc. |
| Collect cluster metrics | Use your cluster-wide tooling (kube-state-metrics, agent-based APM, etc.) |
| Alert on errors | Forward ERROR-level log entries to your alerting system |
| Audit compliance | Use the per-entity Audit Log tabs in the UI; raw logs available in SeaweedFS |
| Job execution monitoring | Use the Job Runs section in the Operations area for task-level status |