Security
Lakehousecat treats security as an ongoing process, not a fixed state. Enterprise data processing requires a defense-in-depth approach: encryption at rest, enforced TLS in transit, minimal infrastructure permissions, per-user MFA, and comprehensive audit trails across all platform entities. Each layer is designed with the principle that no single control should be the last line of defense.
Security Areas
| Area | Description |
|---|---|
| Authentication | User accounts, roles, groups, and optional identity provider (OAuth/OIDC) integration |
| Audit Logging | Per-entity change history and service-level operational logs |
| Compliance | Current compliance status and certification roadmap |
Encryption at Rest
All sensitive data stored by Lakehousecat is encrypted using a symmetric key before being written to the database:
- API keys — LLM and embedding provider API keys
- Provider configuration — model connection credentials and parameters
- Data source configuration — connection strings, passwords, and SSH credentials for connected databases
- User-sensitive data — credential material stored at the application database level
The symmetric encryption key is stored as a Kubernetes Secret (lhc-encryption-key) in the application namespace. If this secret is lost, encrypted data in the database cannot be recovered, even if the backup files are intact. See Backup for details on why secret backup is critical.
Transport Security (TLS)
All external traffic is encrypted via TLS. The Operator configures TLS at the ingress level based on the deployment environment:
| Ingress Mode | TLS Mechanism |
|---|---|
| NGINX (on-premises, IKS, dev) | TLS via Kubernetes Secret; if no customer-provided certificate is configured, the Operator generates a self-signed certificate |
| AWS ALB (EKS) | TLS via AWS ACM — Certificate ARN is configured in the Custom Resource |
TLS is enforced at the ingress level. All HTTP traffic is redirected to HTTPS. Internal service communication within the Kubernetes namespace uses standard cluster networking.
GCE (Google Cloud Load Balancer) and Azure Application Gateway ingress modes are planned for future releases.
Kubernetes RBAC — Least Privilege
The Lakehousecat Operator is designed to run with the minimum permissions required. The Operator operates entirely within namespace-scoped RBAC and does not require cluster-admin or any cluster-wide permissions.
Permissions are limited to the resources needed to manage the application namespace:
- Workload resources: Deployments, StatefulSets, DaemonSets, Jobs, CronJobs, Pods
- Configuration resources: ConfigMaps, Secrets, ServiceAccounts
- Networking resources: Services, Ingresses, NetworkPolicies, EndpointSlices
- Storage resources: PersistentVolumeClaims (provisioning only — not storage classes or PVs)
- Namespace-scoped Roles and RoleBindings (no ClusterRole creation)
- Leader election: Leases
The Operator has no access to cluster-level resources such as Nodes, StorageClasses, ClusterRoles, or PersistentVolumes. This ensures the Operator cannot escalate privileges or affect other namespaces in a shared cluster.
Access Control
All access in Lakehousecat is governed by a combination of roles and groups:
- Roles define a user's baseline capabilities: Administrator, Builder, or User
- Groups control access to specific objects: Custom Models, Data Sources, Dashboards, and Charts
A user's role determines what actions they can perform (e.g., create data sources, manage users). Group membership determines which resources they can access. Both dimensions must be satisfied. See Roles and Groups for details.
Multi-Factor Authentication
MFA is available to all users and configured per-account from Profile → Account → Security. Lakehousecat uses TOTP-based MFA, compatible with any standard authenticator app (Google Authenticator, Authy, etc.). Backup codes are generated during setup and can be regenerated at any time.
See Users for the full MFA setup procedure.
Audit Logging
Lakehousecat produces two categories of audit data:
1. Per-entity audit trail — Every platform entity (Data Sources, Custom Models, Charts, Dashboards, Job Definitions, Users) maintains a structured change history: creation, updates, configuration changes, and sharing events. These are accessible within the application on the Audit Log tab of each entity.
2. Service-level operational logs — All backend services write structured logs to SeaweedFS at regular intervals. These cover API activity, authentication events, job execution, and errors. Log level is configurable. Customers are responsible for forwarding and retaining these logs using their own tooling.
See Audit Logging for log format, configuration, and integration guidance.
Scope and Responsibilities
Lakehousecat is deployed inside the customer's own Kubernetes cluster. The boundary between platform and infrastructure responsibility:
| Area | Responsibility |
|---|---|
| Encryption of sensitive application data at rest | Lakehousecat (built-in) |
| TLS for external ingress traffic | Lakehousecat Operator (self-signed fallback if no cert provided) |
| Per-entity audit trail | Lakehousecat (built-in) |
| Service-level operational logging | Lakehousecat (built-in, SeaweedFS-stored) |
| MFA for user accounts | Lakehousecat (built-in, per-user opt-in) |
| Log retention, forwarding, and alerting | Customer |
| Cluster network security and network policies | Customer |
| Node hardening and infrastructure access controls | Customer |
| Public cloud IAM and security posture | Customer |
| Container image scanning | Customer |
| Backup strategy for PVCs and secrets | Shared — see Backup |
Topics such as cluster-level network policies, Kubernetes admission controllers, node security, container scanning, and public cloud IAM are outside the scope of the Lakehousecat platform.
Compliance
Lakehousecat is working toward SOC 2 certification after initial market launch. No formal certifications are currently held. See Compliance for the full roadmap.