Skip to main content
Version: Next

Monitoring

Scope of the Lakehousecat Operator​

The Lakehousecat Operator is designed to operate within a single Kubernetes namespace and requires only namespace-scoped RBAC permissions. The Operator deploys the application services and exposes the data customers need for observability — structured logs in object storage and Prometheus-format /metrics endpoints on each FastAPI service.

The Operator does not deploy a monitoring stack of its own. Cluster-level and infrastructure monitoring is the responsibility of the customer's operations team. Most customers already have an established monitoring infrastructure — Datadog, AWS CloudWatch, Google Cloud Monitoring, Grafana with a cluster-wide Prometheus installation, or another solution — and Lakehousecat is designed to integrate into whatever system is already in place rather than replacing it.

Customers are responsible for cluster-level monitoring integration

The Operator manages the application namespace. Log collection, forwarding, and retention at the cluster or infrastructure level must be configured by the customer's operations team according to their own monitoring standards and tooling.


Application Logs​

Backend Service Logs​

All Lakehousecat backend services — including the API server, Airflow workers, and other platform components — produce structured application logs. These logs are:

  • Written to stdout/stderr (standard Kubernetes log output, accessible via kubectl logs)
  • Continuously buffered and flushed to SeaweedFS at regular intervals (every 60 seconds, or when the buffer reaches 1,000 records)

Log Storage in SeaweedFS​

Logs are stored in the internal SeaweedFS instance under the following path structure:

backend_logs/logs/<timestamp>_<uuid>.log

Each log file uses the format:

<timestamp> - <logger_name> - <level> - <message>

Log files are generated on a rolling basis — a new file is created every flush cycle or on service shutdown. This means logs are always available in SeaweedFS for inspection and support, even without an external log aggregation system connected.

Airflow / Job Logs​

Task execution logs from Airflow jobs (data loads, semantic extraction, custom tasks) are also stored in SeaweedFS as part of the job run output. These are accessible from the Job Runs detail view in the Operations area of the workspace.


Audit Logs​

In addition to application logs, Lakehousecat maintains structured audit logs for all user-initiated changes to platform entities:

EntityLogged events
Data SourcesCreated, updated, filters changed, operations triggered
Custom ModelsCreated, configured, semantic model operations, shared
ChartsCreated, edited, shared
DashboardsCreated, edited, shared
Job DefinitionsCreated, updated, triggered
UsersInvited, role changed, removed

Audit log entries are available within the application (each entity's Audit Log tab) and are also persisted to SeaweedFS as part of the structured log output.


Connecting Your Monitoring System​

Lakehousecat produces logs through two channels that your monitoring system should consume:

1. Kubernetes Pod Logs (stdout/stderr)​

Standard Kubernetes log output is available via:

kubectl logs -n <namespace> <pod-name>
kubectl logs -n <namespace> -l app=<label> --all-containers

Most monitoring agents (Datadog Agent, Fluent Bit, Logstash, AWS CloudWatch Agent, etc.) can be configured to collect pod logs from the namespace automatically. Configure your log collector to forward logs from the Lakehousecat namespace to your central log management system.

2. SeaweedFS Log Files​

Logs stored in SeaweedFS under backend_logs/ are available for retrieval via the SeaweedFS S3 API or the SeaweedFS Filer UI. You can:

  • Connect an S3-compatible log sink or ETL pipeline to read from the SeaweedFS bucket
  • Schedule periodic exports to your log management system
  • Access log files directly for incident investigation or support

SeaweedFS connection details (endpoint, access key, bucket name) are available in the deployment configuration under the seaweedfs: section of the Custom Resource.


Application Metrics — Prometheus /metrics Endpoint​

Each FastAPI backend service (lhc, llm, analytics, semantic, audio) exposes a /metrics endpoint in the standard Prometheus exposition format. This endpoint is intended for scraping by your existing monitoring system (Prometheus, Datadog Agent, Grafana Mimir, AWS Managed Prometheus, etc.).

What is exposed​

  • HTTP request metrics — request counts and latency histograms per route template, method, and status code
  • Python process metrics — resident memory, CPU seconds, GC statistics, open file descriptors

Path labels are normalized to route templates (for example /api/v1/charts/{id} instead of concrete IDs) so cardinality stays bounded and no user-identifying values leak into labels.

Endpoint properties​

PropertyValue
Path/metrics
FormatPrometheus text exposition (OpenMetrics-compatible)
AuthenticationNone by design — follows the Prometheus ecosystem convention
Access controlCluster-internal only (ClusterIP service + your NetworkPolicy)
ToggleCRD field monitoring.metricsEnabled (default true)

To disable the endpoints, set monitoring.metricsEnabled: false in the Lakehousecat Custom Resource. The Operator propagates the flag as the environment variable LHC_METRICS_ENABLED into each service.

Example: ServiceMonitor for Prometheus Operator​

Customers running prometheus-operator can scrape Lakehousecat services with a ServiceMonitor such as:

apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: lakehousecat-services
namespace: <your-lakehousecat-namespace>
spec:
selector:
matchLabels:
app.kubernetes.io/part-of: lakehousecat
endpoints:
- port: http
path: /metrics
interval: 30s

Datadog, New Relic and other agents support equivalent auto-discovery via Kubernetes service annotations — consult your agent's documentation.

Cluster-level Metrics​

Pod- and node-level metrics (CPU, memory, restarts, kubelet stats) are produced by Kubernetes itself via kube-state-metrics, cAdvisor, and the kubelet — the customer's cluster monitoring tooling collects these directly. The Operator does not duplicate this.


Recommendations​

TaskRecommendation
Collect pod logsConfigure your log agent to scrape the Lakehousecat namespace
Retain application logsSet a retention policy on the SeaweedFS backend_logs/ bucket
Collect application metricsScrape /metrics on each backend service with your existing Prometheus / Datadog / etc.
Collect cluster metricsUse your cluster-wide tooling (kube-state-metrics, agent-based APM, etc.)
Alert on errorsForward ERROR-level log entries to your alerting system
Audit complianceUse the per-entity Audit Log tabs in the UI; raw logs available in SeaweedFS
Job execution monitoringUse the Job Runs section in the Operations area for task-level status