Load Storage Sizing
When Lakehousecat loads a datasource into its warehouse, the data does not stream straight from the source into ClickHouse. It is materialised on local disk first: the load runs in three stages — extract, normalize, load — and each stage writes intermediate files before the next one picks them up.
This page explains how much local space a load needs, where that space is taken from, and which knobs exist if a load runs out of it.
Where the intermediate data goes
Loads run inside a task pod, in a scratch directory pinned to /tmp/dlt. That directory is
backed by an emptyDir volume, which lives on the node's disk — the same disk that holds
container images and the logs of every other pod on that node.
If the node disk runs full, the kubelet starts evicting pods to reclaim space. It selects them by priority and QoS class — not by which pod caused the pressure. An oversized load can therefore take down a database pod scheduled on the same node. See Task Pod Storage Limits for how to contain this.
How much space a load needs
Loads are processed one table at a time. Each table is extracted, normalised, written into the warehouse, and its intermediate files are then deleted before the next table starts.
The practical consequence:
The peak local disk usage of a load is driven by the largest single table, not by the total size of the schema.
A reference measurement over a 23-table schema showed a peak of 0.53 MB against 9.46 MB of source data — the intermediate files of already-loaded tables no longer accumulate.
Two caveats when sizing:
- Intermediate files are written as compressed Parquet. They are usually smaller than the source table, but the order of magnitude follows it — plan against your largest table's size, not against an average row count.
- If a table fails to load, its files are deliberately kept for troubleshooting. A load with failing tables therefore consumes more space than a clean one.
Limiting what a single load can consume
By default there is no upper bound other than the free space on the node. If you operate mixed workloads on shared nodes, set an explicit limit so that an oversized load fails on its own instead of pressuring its neighbours.
Both values are set on the Lakehousecat resource and apply to every job task pod:
apiVersion: lhc.lakehousecat.com/v1alpha2
kind: Lakehousecat
spec:
platform:
taskPodStorage:
request: "5Gi" # what the scheduler reserves
limit: "20Gi" # what the kubelet enforces
requestis what makes the scheduler account for the space. Without it, task pods get placed on nodes that never had room for them.limitis the point at which the kubelet evicts the pod. Leave it empty to keep the current unbounded behaviour.
Both are standard Kubernetes quantities (Gi, Mi, G, M, …). Decimal values such as
1.5Gi are rejected — use 1536Mi.
A task pod runs two containers. The operator splits the configured budget between them, so
limit: "20Gi" means the pod as a whole may use 20 GiB — not 20 GiB per container. You size
against the load, not against the pod's internal structure.
There is no universally correct default, which is why none is applied. The value has to be at least the size of your largest table's intermediate representation — set it too low and loads that used to succeed will start failing. When a limit is exceeded, the pod is evicted and the job is reported as a generic task failure.
Trading throughput for a bounded footprint
Loading table by table costs some throughput: each table carries a fixed per-table overhead. A reference run over 23 small tables took about 60 % longer than loading the whole schema in one pass. The relative overhead shrinks as tables get larger, since the fixed cost is amortised over more data.
If a specific source misbehaves under table-wise loading, the behaviour can be turned off with an environment variable on the semantic service:
LHC_DLT_CHUNKED_LOAD=false
Disabling it restores the previous behaviour: the entire schema is materialised in one pass and the intermediate files are kept until the pod terminates. Local disk usage then scales with the full size of the schema. Use it to work around a problem, and re-enable it afterwards.