Skip to main content
Version: Next

Limits

Concurrent Users​

Lakehousecat's asynchronous backend architecture is designed for high concurrency. Tested capacity:

  • ~1,000 concurrent Users — validated with standard query and analytics workloads
  • ~100 concurrent Builders — tested with parallel semantic model training and data loading operations

The actual capacity of any deployment depends on the resources available in the customer's Kubernetes cluster. Every service in the Lakehousecat framework can be scaled horizontally (replicas) and vertically (CPU/memory) via the Operator configuration. There is no hard architectural ceiling — performance scales with available resources.

Resource-intensive workflows (e.g., full data source loads, semantic layer creation) run as isolated Kubernetes pods via the workflow engine and do not compete directly with user-facing request handling.

LLM Provider Limits​

LLM providers impose their own constraints that can affect Lakehousecat operation. These limits are defined by the customer's contract and tier with the provider — they are not controlled by Lakehousecat.

Limit TypeDescription
Rate LimitsMaximum requests per minute/hour/day — defined by the provider tier and the customer's contract. Multiple users sharing a single API key consume from the same rate limit budget.
Context WindowMaximum tokens per request — affects how much data or conversation history can be processed in a single call. Varies by model.
Response LatencyVaries by provider, model complexity, and geographic proximity to the provider's inference infrastructure.

Shared API Key Concurrency​

All users of a Lakehousecat instance share the configured API key for each provider. If a large number of users send queries simultaneously, the combined request rate can hit provider rate limits, resulting in delayed or queued responses.

Rate limits depend on the customer's provider contract

A provider API key on a free or low-tier plan may be insufficient for multi-user deployments. Customers should configure their API keys to match the expected concurrent usage of their organization. Upgrading to a higher provider tier or requesting rate limit increases from the provider is the customer's responsibility.

Customers are responsible for selecting appropriate provider tiers, monitoring their API usage, and managing rate limits through the provider's own tooling.

Lakehousecat API Rate Limiting​

Separate from LLM provider limits above, Lakehousecat's own API ingress applies rate limiting to the /api/v1/... routes, by default 60 requests/second and 20 concurrent connections per client IP address.

This is a DoS-protection default, not a per-subscription-tier capacity number. Seat count (how many users your subscription allows) and concurrency (how many requests can be in flight from one apparent source at once) are two different things — upgrading your subscription tier does not, by itself, change this ingress setting. It rarely matters for deployments where users connect from distinct IP addresses; it is more likely to matter when many users share one apparent source IP, for example behind a corporate NAT gateway or a site-to-site VPN into the cluster's network.

Administrators can raise these limits or exempt specific source IP ranges via spec.ingress.rateLimit on the Lakehousecat Custom Resource — see API Ingress Rate Limiting for the configuration fields and an example.

Infrastructure and Storage​

Lakehousecat is deployed on the customer's Kubernetes cluster. The following constraints are determined by the customer's infrastructure, not by Lakehousecat:

  • Storage capacity — determined by PVC sizes and the cluster's storage class
  • Network bandwidth — determined by the cluster's network configuration
  • Compute limits — determined by node pool sizes and quotas
  • Cloud provider quotas — account-level limits on vCPUs, instances, and services

If the cluster runs out of storage or compute, this is the cluster administrator's responsibility to resolve by adjusting the infrastructure configuration or Operator resource allocations.