Kubernetes Operator#
The LMCache Kubernetes operator automates the deployment and lifecycle
management of LMCache multiprocess servers. Instead of hand-writing
DaemonSets, Services, and ConfigMaps (as described in the manual
Deployment Guide guide), you declare a single LMCacheEngine custom
resource and the operator reconciles all underlying Kubernetes objects.
Why Use the Operator#
The manual DaemonSet approach works, but it has sharp edges the operator eliminates:
Auto-injected pod settings – The operator always mounts the host’s
/dev/shm(hostPath) and sets--host 0.0.0.0. Forgetting the shared/dev/shmin a hand-written manifest causes silent CUDA IPC failures (cudaErrorMapBufferObjectFailed) that are hard to debug.Node-local service discovery – The operator creates a ClusterIP Service with
internalTrafficPolicy=Localand a connection ConfigMap that vLLM pods simply mount. NohostNetwork, no Downward API, no shell variable substitution.Auto-computed resource sizing – Memory requests and limits are derived from
l1.sizeGB, avoiding OOM kills (under-provisioned) or wasted node capacity (over-provisioned).Declarative Prometheus integration – Set
prometheus.serviceMonitor.enabled: trueand the operator creates aServiceMonitorCR that the Prometheus Operator discovers automatically.CRD validation – OpenAPI schema validation catches misconfigurations (e.g.,
l1.sizeGB <= 0, invalid port range) atkubectl applytime, before any pods are created.
Prerequisites#
Kubernetes 1.20+
kubectlconfigured to access your cluster(Optional) Prometheus Operator for ServiceMonitor support
Installing the Operator#
Option A: One-line install from release (recommended)
# Latest stable release
kubectl apply -f https://github.com/LMCache/LMCache/releases/download/operator-latest/install.yaml
# Or nightly build from the dev branch
kubectl apply -f https://github.com/LMCache/LMCache/releases/download/operator-nightly-latest/install.yaml
Option B: Build from source
cd operator
make build
make install
make deploy IMG=<your-registry>/lmcache-operator:latest
Deploying an LMCacheEngine#
A minimal CR deploys a DaemonSet with 60 GB L1 cache on every GPU node:
apiVersion: lmcache.lmcache.ai/v1alpha1
kind: LMCacheEngine
metadata:
name: my-cache
spec:
l1:
sizeGB: 60
kubectl apply -f lmcache-engine.yaml
The operator automatically:
Creates a DaemonSet running one LMCache server pod per matched node
Mounts the host’s
/dev/shm(hostPath) for CUDA IPC and passes--host 0.0.0.0to the serverCreates a node-local ClusterIP Service for vLLM discovery
Creates a connection ConfigMap (
my-cache-connection) with thekv-transfer-configJSON that vLLM needsAuto-computes resource requests/limits from the L1 cache size
Defaults
nodeSelectortonvidia.com/gpu.present: "true"
Note
The operator defaults the container image to lmcache/vllm-openai:latest.
Override with spec.image.repository and spec.image.tag to pin a
specific version.
Connecting vLLM#
The operator creates a ConfigMap named <engine-name>-connection containing
the kv-transfer-config JSON. You can either let the operator’s mutating
webhook inject it for you (recommended – keeps your vLLM manifest clean) or
mount it by hand. See Connection Injection (Webhook) below for the
webhook flow; the rest of this section describes the manual mount that is its
equivalent.
Mount it in your vLLM Deployment:
apiVersion: apps/v1
kind: Deployment
metadata:
name: vllm
spec:
replicas: 1
selector:
matchLabels:
app: vllm
template:
metadata:
labels:
app: vllm
spec:
containers:
- name: vllm
image: lmcache/vllm-openai:latest
env:
# Deterministic hashing required by LMCache
- name: PYTHONHASHSEED
value: "0"
command: ["/bin/sh", "-c"]
args:
- |
exec python3 -m vllm.entrypoints.openai.api_server \
--model <your-model> \
--port 8000 \
--gpu-memory-utilization 0.8 \
--kv-transfer-config "$(cat /etc/lmcache/kv-transfer-config.json)"
ports:
- name: http
containerPort: 8000
volumeMounts:
- name: kv-transfer-config
mountPath: /etc/lmcache
readOnly: true
# Required for CUDA IPC between vLLM and LMCache: both pods
# must see the same /dev/shm tmpfs
- name: lmcache-dev-shm
mountPath: /dev/shm
resources:
limits:
nvidia.com/gpu: "1"
volumes:
- name: kv-transfer-config
configMap:
name: my-cache-connection # <engine-name>-connection
- name: lmcache-dev-shm
hostPath:
path: /dev/shm
type: Directory
Key requirements for vLLM pods:
Shared /dev/shm – CUDA IPC (
cudaIpcOpenMemHandle) needs vLLM and LMCache to see the same/dev/shmtmpfs (PyTorch’s CUDA IPC handles reference a shared-memory ref-counter file there). Mount the host’s/dev/shmvia hostPath as above, or setspec.hostIPC: trueon the engine andhostIPC: trueon the vLLM pod to share the host IPC namespace instead.PYTHONHASHSEED=0 – Ensures deterministic token hashing so vLLM and LMCache produce consistent cache keys.
ConfigMap mount – The
$(cat ...)pattern reads the connection JSON inline. The ConfigMap name is always<LMCacheEngine name>-connection.No hostNetwork needed – The operator’s node-local Service handles routing via
internalTrafficPolicy=Local.
Connection Injection (Webhook)#
Hand-wiring the ConfigMap mount and the $(cat ...) argument substitution
above is repetitive across vLLM Deployments. A mutating admission webhook
shipped with the operator can do it for you so the vLLM manifest stays clean.
It mirrors the CacheBlend webhook (see CacheBlend) with an
lmcache- annotation/label discriminator so the two injectors never
cross-fire on the same pod.
When invoked on an opted-in pod whose <engine>-connection ConfigMap exists,
the webhook mutates the pod at admission time to add:
--kv-transfer-config <JSON>– theLMCacheMPConnectorconfig, read verbatim from the engine’s<engine>-connectionConfigMap and inlined onto the vLLM container’sargs(no volume mount needed);a hostPath mount of the host’s
/dev/shm(CUDA IPC with the node-local server;hostIPC: trueinstead when the engine setsspec.hostIPC);PYTHONHASHSEED=0on the vLLM container env, set-if-absent – it preserves a value you already set.
The connector config lives in the connection ConfigMap; the webhook also
reads the LMCacheEngine CR (when present) to mirror its spec.hostIPC
mode and its optional injection sub-spec. It fails open
(failurePolicy: Ignore) and is idempotent (re-admitted pods carrying the
lmcache.ai/lmcache-injected stamp are allowed unchanged).
Prerequisites#
cert-manager +
make deploy(notmake run, which is controller-only and disables the webhook viaENABLE_WEBHOOKS=false) – same as the CacheBlend webhook; install once per cluster (see CacheBlend “Additional Prerequisites”).Pod Security Standards – the injected hostPath
/dev/shmmount (andhostIPC, when the engine opts in) is rejected by thebaseline/restrictedPSS profiles, so the vLLM pod’s namespace must be labeledpod-security.kubernetes.io/enforce=privileged.Engine reconciled in the same namespace – the webhook reads the
<engine>-connectionConfigMap directly, so theLMCacheEnginemust already exist in the vLLM pod’s namespace.
Opting a vLLM Pod In#
Add the opt-in label and the engine-binding annotation to the pod template, and
launch vLLM via the image ENTRYPOINT (args only) – a
command: ["/bin/sh", "-c", ...] wrapper is skipped (the webhook stamps
lmcache.ai/lmcache-skip-reason=command-override because appended args would
not reach vllm serve):
apiVersion: apps/v1
kind: Deployment
metadata:
name: vllm-lmcache
spec:
replicas: 1
selector:
matchLabels:
app: vllm-lmcache
template:
metadata:
labels:
app: vllm-lmcache
lmcache.ai/lmcache-inject: "true" # opt-in (webhook objectSelector)
annotations:
lmcache.ai/lmcache-engine: "my-cache" # bind to the engine (same namespace)
# Optional -- name the vLLM container if it is not the first one:
# lmcache.ai/lmcache-container: "vllm"
spec:
runtimeClassName: nvidia
# Do NOT mount an emptyDir at /dev/shm -- the webhook injects a
# hostPath mount of the host's /dev/shm; an emptyDir would shadow the
# host's /dev/shm and break cudaIpcOpenMemHandle.
containers:
- name: vllm
image: lmcache/vllm-openai:latest
# Args-only launch (image ENTRYPOINT is ["vllm", "serve"]). The
# webhook appends --kv-transfer-config; do NOT add it yourself
# (a user-supplied one stamps skip-reason=kv-transfer-config-present).
args: ["<your-model>", "--port", "8000", "--gpu-memory-utilization", "0.8"]
ports:
- name: http
containerPort: 8000
resources:
limits:
nvidia.com/gpu: "1"
A ready-to-edit manifest lives at
operator/config/samples/vllm_lmcache_deployment.yaml.
Verifying Injection#
The webhook mutates Pods, not the Deployment, so inspect a pod (not the Deployment spec):
kubectl get pod -l app=vllm-lmcache -o yaml | \
grep -E "lmcache-dev-shm|hostIPC|kv-transfer-config|lmcache-injected|lmcache-skip-reason"
If nothing was injected, check the pod’s lmcache.ai/lmcache-skip-reason
annotation:
command-override– the pod uses ash -cwrapper, so injected args would not reachvllm serve.kv-transfer-config-present– the user already supplied--kv-transfer-config; the webhook does not clobber it.engine-not-found– the<engine>-connectionConfigMap is missing (engine not yet reconciled, or wrong namespace, or wrong name).target-container-not-found– thelmcache.ai/lmcache-containerannotation names a container the pod does not have.
With failurePolicy: Ignore a webhook / cert problem also leaves the pod
un-mutated silently – confirm the operator pod is Running and the
MutatingWebhookConfiguration exists.
Using the Latest (or a Pinned) lmcache#
By default a vLLM pod runs whatever lmcache is baked into its image. To run
a different lmcache build instead – e.g. ship the latest lmcache onto an
older, stable vLLM image, or keep the vLLM client on the exact build its
LMCacheEngine server runs – set spec.injection.payloadImage on the
engine. The webhook then additionally stages that image’s lmcache tree into
each opted-in pod: an emptyDir + an init container that copies the tree in, a
read-only mount, and PYTHONPATH=/lmcache-payload so vLLM imports the staged
lmcache instead of the baked-in one. No vLLM image rebuild.
1. Build the payload image. It ships the unpacked lmcache tree under
/payload and copies it to $SHARED_DIR on start. docker/Dockerfile.payload
builds it by extracting an ABI-matched lmcache from an lmcache image (the
SOURCE_IMAGE build-arg selects the version):
docker build -f docker/Dockerfile.payload \
--build-arg SOURCE_IMAGE=lmcache/vllm-openai:latest-nightly \
-t <registry>/lmcache-payload:latest .
docker push <registry>/lmcache-payload:latest
2. Point the engine at it. payloadImage.repository has no valid default
(the inherited image default is not a payload), so set it explicitly; leaving
injection unset keeps connection-only wiring.
apiVersion: lmcache.lmcache.ai/v1alpha1
kind: LMCacheEngine
metadata:
name: my-cache-versioned
spec:
l1:
sizeGB: 60
injection:
payloadImage:
repository: <registry>/lmcache-payload
tag: latest
pullPolicy: Always # :latest moves -- re-pull for the current build
# imagePullSecrets: # private payload registry only
# - name: my-registry-secret
Opted-in pods bound to this engine (label + annotation as above) need no
changes – the webhook stages the payload automatically. Ready-to-apply
samples: config/samples/lmcache_v1alpha1_lmcacheengine_injection.yaml and
config/samples/vllm_lmcache_injection_deployment.yaml.
Note
The payload’s lmcache must be ABI-compatible (same Python minor
version and a compatible torch) with the vLLM image that imports it – it
ships compiled extensions. If they differ, import lmcache fails with an
undefined symbol error in the vLLM pod. Building the payload from an
lmcache image close to your vLLM image keeps them compatible.
3. Verify the swap on a running pod – contrast the normal import with one
that ignores the injected PYTHONPATH:
POD=$(kubectl get pod -l app=vllm-lmcache-versioned -o name | head -1)
# imports the STAGED build (from /lmcache-payload):
kubectl exec $POD -c vllm -- python3 -c \
"import lmcache; print(lmcache.__version__, lmcache.__file__)"
# PYTHONPATH stripped -> the image's baked-in build (site-packages):
kubectl exec $POD -c vllm -- env -u PYTHONPATH python3 -c \
"import lmcache; print(lmcache.__version__, lmcache.__file__)"
Two different sources for the same module confirms the swap. If nothing was
staged, check lmcache.ai/lmcache-skip-reason on the pod.
Verifying the Deployment#
# Check LMCacheEngine status
kubectl get lmc
Expected output:
NAME PHASE READY DESIRED AGE
my-cache Running 3 3 5m
# Check the connection ConfigMap
kubectl get configmap my-cache-connection -o yaml
# Check LMCache pods
kubectl get pods -l app.kubernetes.io/managed-by=lmcache-operator
# Check detailed status with endpoints
kubectl describe lmc my-cache
CRD Spec Reference#
Image#
Field |
Default |
Description |
|---|---|---|
|
|
Container image repository. |
|
|
Container image tag. |
|
|
|
|
– |
Image pull secret references. |
Server#
Field |
Default |
Description |
|---|---|---|
|
|
ZMQ listening port (1024–65535). |
|
|
Token chunk size. |
|
|
Worker threads for ZMQ requests. |
|
|
|
|
|
HTTP frontend port for health checks and cache admin (1024–65535). |
L1 Cache#
Field |
Default |
Description |
|---|---|---|
|
required |
L1 cache size in GB. Must be > 0. |
Eviction#
Field |
Default |
Description |
|---|---|---|
|
|
|
|
|
Usage ratio (0.0–1.0] to trigger eviction. |
|
|
Fraction to evict (0.0–1.0]. |
Prometheus#
Field |
Default |
Description |
|---|---|---|
|
|
Expose Prometheus metrics. |
|
|
|
|
|
Create a ServiceMonitor CR. |
|
|
Scrape interval. |
|
– |
Extra labels on the ServiceMonitor. |
L2 Storage#
Field |
Default |
Description |
|---|---|---|
|
– |
List of L2 backends ( |
|
– |
Optional serde transform on KV bytes to/from the L2 adapter
(rendered as the |
|
– |
At-rest encryption of L2 KV bytes (AES-GCM, keyed per
|
|
– |
Required. User-created Secret in the engine’s namespace holding
the master key under the |
|
|
How per- |
|
|
AES key size: |
GPU & Security#
Field |
Default |
Description |
|---|---|---|
|
|
GPU vendor: |
|
|
Run the pod in the host IPC namespace instead of mounting the host’s
|
|
|
Run the engine container in privileged mode. On most clusters
|
Scheduling#
Field |
Default |
Description |
|---|---|---|
|
GPU nodes |
Defaults to |
|
– |
Pod affinity rules. |
|
– |
Pod tolerations. |
|
– |
Priority class for pods. |
Overrides & Extras#
Field |
Default |
Description |
|---|---|---|
|
|
|
|
– |
Override auto-computed resources. |
|
– |
Extra environment variables. |
|
– |
Extra volumes. |
|
– |
Extra volume mounts. |
|
– |
Additional init containers run before the lmcache container starts,
in order. Use to prepare state the engine can’t create itself, e.g.
pre-sizing a |
|
– |
Extra pod annotations. |
|
– |
Extra pod labels. |
|
– |
ServiceAccount for pods. |
|
– |
Extra CLI flags (appended last, can override). |
Auto-Computed Resources#
When spec.resourceOverrides is not set, the operator derives resources from
l1.sizeGB:
CPU request:
4coresMemory request:
ceil(l1.sizeGB + 5)GiMemory limit:
ceil(memoryRequest * 1.5)Gi
For example, l1.sizeGB: 60 produces a 65 Gi request and 98 Gi limit.
Auto-Injected Pod Settings#
The operator always injects these into the pod spec:
Host /dev/shm mount (hostPath) – Required for CUDA IPC between LMCache and vLLM: both processes must see the same
/dev/shmtmpfs (PyTorch’s CUDA IPC handles reference a shared-memory ref-counter file there). Setspec.hostIPC: trueto share the host IPC namespace instead (needed on clusters that block hostPath volumes, and forgpuVendor: amd); the mount is then omitted.–host 0.0.0.0 – Binds the server to all interfaces so the node-local Service can route to it.
NVIDIA_VISIBLE_DEVICES=all – Ensures GPU access for IPC-based memory transfers.
NVIDIA_DRIVER_CAPABILITIES=all – Exposes all driver capabilities (compute, utility, etc.) to the container.
TCP socket probes – Startup (5s initial, 30 failures), liveness (10s), and readiness (5s) probes on the server port.
Note
The operator never mounts an emptyDir at /dev/shm. Mounting an emptyDir
there would shadow the host’s /dev/shm with a private tmpfs and break
CUDA IPC.
Resources Created#
For an LMCacheEngine named my-cache:
Resource |
Name |
Purpose |
|---|---|---|
DaemonSet |
|
Runs LMCache server pods. |
Service (ClusterIP) |
|
Node-local discovery ( |
Service (headless) |
|
Prometheus scrape target. |
ConfigMap |
|
|
ServiceMonitor |
|
Prometheus Operator integration (when enabled). |
The connection ConfigMap contains:
{
"kv_connector": "LMCacheMPConnector",
"kv_role": "kv_both",
"kv_connector_extra_config": {
"lmcache.mp.host": "tcp://my-cache.default.svc.cluster.local",
"lmcache.mp.port": "5555"
}
}
Status & Conditions#
kubectl describe lmc my-cache
The status section includes:
phase:
Pending,Running,Degraded, orFailed.readyInstances / desiredInstances: Instance counts.
endpoints: Per-node connection info (node name, host IP, pod name, port, readiness).
conditions:
Available– At least one instance is ready.AllInstancesReady– All desired instances are ready.ConfigValid– Spec validation passed.
Validation Rules#
The operator validates the CR spec at apply time:
Field |
Rule |
|---|---|
|
Required, must be > 0. |
|
Must be |
|
Must be in (0.0, 1.0]. |
|
Must be in (0.0, 1.0]. |
|
Must be in [1024, 65535]. |
|
Exactly one serde type must be set (only |
|
Required, must be non-empty. |
|
Rejected if the raw adapter config already sets a |
Examples#
Target Only GPU Nodes#
Use nodeSelector to run LMCache only on GPU nodes. New GPU nodes
automatically get an LMCache pod:
apiVersion: lmcache.lmcache.ai/v1alpha1
kind: LMCacheEngine
metadata:
name: my-cache
spec:
nodeSelector:
nvidia.com/gpu.present: "true"
l1:
sizeGB: 60
Note
The operator defaults nodeSelector to nvidia.com/gpu.present: "true"
when not specified, so a minimal CR already targets GPU nodes.
Custom Server Port#
If the default port (5555) conflicts with other services:
apiVersion: lmcache.lmcache.ai/v1alpha1
kind: LMCacheEngine
metadata:
name: my-cache
spec:
server:
port: 6555
l1:
sizeGB: 60
The connection ConfigMap updates automatically – vLLM pods pick up the new port on restart.
Encrypted L2 Backend#
Encrypt KV bytes at rest in the L2 tier with the aesgcm serde. Create
the master-key Secret first, in the engine’s namespace (the operator never
generates keys):
head -c 16 /dev/urandom > master.key # 16 bytes for AES-128, 32 for AES-256
kubectl create secret generic lmcache-l2-master-key \
--from-file=master=master.key
apiVersion: lmcache.lmcache.ai/v1alpha1
kind: LMCacheEngine
metadata:
name: my-cache
spec:
l1:
sizeGB: 60
l2Backend:
raw:
type: fs
config:
base_path: /data/lmcache/l2
serde:
aesgcm:
masterKeySecretRef:
name: lmcache-l2-master-key
The Secret is mounted directly into the engine pods (read-only, only the
master data key). A missing Secret or missing key surfaces as a pod
mount event and self-heals once the Secret is fixed. See KV Cache Compression for
the key model and threat model.
Production with Prometheus Monitoring#
apiVersion: lmcache.lmcache.ai/v1alpha1
kind: LMCacheEngine
metadata:
name: production-cache
namespace: llm-serving
spec:
nodeSelector:
nvidia.com/gpu.present: "true"
image:
repository: lmcache/standalone
tag: v0.1.0
server:
port: 6555
chunkSize: 256
maxWorkers: 4
l1:
sizeGB: 60
eviction:
triggerWatermark: 0.8
evictionRatio: 0.2
prometheus:
enabled: true
port: 9090
serviceMonitor:
enabled: true
labels:
release: kube-prometheus-stack
podAnnotations:
prometheus.io/scrape: "true"
prometheus.io/port: "9090"
priorityClassName: system-node-critical
See Observability for metric names and Grafana configuration.
Override Auto-Computed Resources#
apiVersion: lmcache.lmcache.ai/v1alpha1
kind: LMCacheEngine
metadata:
name: my-cache
spec:
l1:
sizeGB: 60
resourceOverrides:
requests:
memory: "70Gi"
cpu: "8"
limits:
memory: "100Gi"
Pre-Sizing an L2 Raw Block Device#
The raw_block L2 adapter opens its device_path file expecting it to
already exist at the configured size – it does not create or grow the file
itself. On a fresh node, that file is missing and the engine crashes with a
FileNotFoundError. Use initContainers to fallocate (or truncate) it
before the lmcache container starts:
apiVersion: lmcache.lmcache.ai/v1alpha1
kind: LMCacheEngine
metadata:
name: my-cache
spec:
l1:
sizeGB: 60
volumes:
- name: kv-cache-root
hostPath:
path: /path/to/local/disk # e.g. a mounted NVMe scratch disk
type: DirectoryOrCreate
volumeMounts:
- name: kv-cache-root
mountPath: /mnt/kv-cache-root
initContainers:
- name: preallocate-raw-block
image: busybox
command:
- sh
- -c
- |
test -f /mnt/kv-cache-root/lmcache-l2.raw || \
fallocate -l 10000000000 /mnt/kv-cache-root/lmcache-l2.raw || \
truncate -s 10000000000 /mnt/kv-cache-root/lmcache-l2.raw
volumeMounts:
- name: kv-cache-root
mountPath: /mnt/kv-cache-root
l2Backend:
raw:
type: raw_block
config:
device_path: "/mnt/kv-cache-root/lmcache-l2.raw"
capacity_bytes: 10000000000 # 10 GB, must match the size above
Note
capacity_bytes in l2Backend.raw.config must match the size the
init container allocates. The test -f ... || guard makes the
fallocate idempotent – a pod restart never truncates a file that
already has data, so a populated L2 cache survives engine restarts.
Init containers listed here run in order, before the lmcache container,
and share the engine’s volumes/volumeMounts – mount the same
volume in both if the init container needs to touch it.
CacheBlend#
CacheBlend reuses cached KV at shifted (non-prefix) positions by recomputing a
small subset of tokens. The operator manages it as a second CRD,
CacheBlendEngine, plus a mutating admission webhook that injects the
pure-Python lmcache-cacheblend vLLM plugin into your serving pods – so you
do not rebuild the vLLM image. See Blending
for the technique itself.
It has two halves the operator runs together:
a GPU-resident CacheBlend V3 engine (
lmcache server --engine-type blend), deployed as a DaemonSet with the same GPU model asLMCacheEngine(runtimeClassName: nvidia+NVIDIA_VISIBLE_DEVICES=all+ the host/dev/shmmount – orhostIPCwhenspec.hostIPCis set – plusprivilegedwhenspec.privilegedis set, and nonvidia.com/gpuclaim) so it shares the vLLM GPU for same-device CUDA IPC; andthe vLLM-side plugin, injected into opted-in pods by the webhook.
Additional Prerequisites#
Beyond the operator prerequisites above:
cert-manager – the webhook’s serving certificate is issued by a cert-manager
Issuer+Certificate. Install it beforemake deploy:kubectl apply -f https://github.com/cert-manager/cert-manager/releases/latest/download/cert-manager.yaml kubectl -n cert-manager wait --for=condition=Available deploy --all --timeout=180sDeploy with the webhook – use
make deploy(notmake run, which is controller-only and disables the webhook viaENABLE_WEBHOOKS=false).Pod Security Standards – the webhook injects a hostPath
/dev/shmmount (orhostIPC/privilegedwhen the engine opts in), which thebaseline/restrictedprofiles reject, so label the engine’s and the vLLM pod’s namespacespod-security.kubernetes.io/enforce=privileged.
Deploying a CacheBlendEngine#
apiVersion: lmcache.lmcache.ai/v1alpha1
kind: CacheBlendEngine
metadata:
name: my-cacheblend
spec:
l1:
sizeGB: 60
injection:
# The (private) cacheblend-plugin init-container image -- repository/tag/
# pullPolicy, like spec.image. Set repository to YOUR image; the
# inherited engine-image default is not a valid payload.
payloadImage:
repository: <registry>/cacheblend-plugin
tag: <tag>
# Appended to the vLLM pod so the private payload image can pull; the
# Secret must exist in the vLLM pod's namespace.
imagePullSecrets:
- name: my-registry-secret
The engine runs lmcache server --engine-type blend as a DaemonSet and
emits a my-cacheblend-connection ConfigMap with the CBKVConnector
kv-transfer-config (the operator wires the node-local Service host/port and
the cb.* tunables).
Opting a vLLM Pod In#
Label the pod template for the webhook and bind it to an engine by name. Launch
vLLM via the image ENTRYPOINT (args only) – a
command: ["/bin/sh", "-c", ...] wrapper is skipped, since appended args would
not reach vllm serve:
apiVersion: apps/v1
kind: Deployment
metadata:
name: vllm-cacheblend
spec:
replicas: 1
selector:
matchLabels:
app: vllm-cacheblend
template:
metadata:
labels:
app: vllm-cacheblend
lmcache.ai/cacheblend-inject: "true" # opt-in (webhook objectSelector)
annotations:
lmcache.ai/cacheblend-engine: "my-cacheblend" # bind to the engine
spec:
runtimeClassName: nvidia
containers:
- name: vllm
image: lmcache/vllm-openai:<pinned-tag>
args: ["<your-model>", "--port", "8000", "--gpu-memory-utilization", "0.8"]
resources:
limits:
nvidia.com/gpu: "1"
The webhook injects the plugin init container, PYTHONPATH, the /dev/shm
mount, the
private-image pull secret, and the required CacheBlend vLLM flags
(--kv-transfer-config from the engine’s connection ConfigMap,
--pipeline-parallel-size 1, --no-enable-chunked-prefill,
--enforce-eager). You supply only the model and your non-CacheBlend flags.
Verifying Injection#
The webhook mutates Pods, not the Deployment, so inspect a pod:
kubectl get pod -l app=vllm-cacheblend -o yaml | \
grep -E "initContainers|cb-plugin|PYTHONPATH|kv-transfer-config|cacheblend-injected|skip-reason"
If nothing was injected, check the pod’s lmcache.ai/cacheblend-skip-reason
annotation: command-override (a sh -c wrapper was used),
kv-transfer-config-present (you set your own), engine-not-found (the
<name>-connection ConfigMap is missing), payload-image-unset (the
engine’s injection.payloadImage has no repository), or
target-container-not-found (the requested targetContainer /
cacheblend-container annotation names a container the pod does not have).
With failurePolicy: Ignore a
webhook/cert problem also leaves the pod un-mutated silently – confirm the
operator pod is Running and the MutatingWebhookConfiguration exists.
CacheBlendEngine Fields#
CacheBlendEngineSpec mirrors LMCacheEngineSpec (every field in the CRD
Spec Reference above) and adds:
Field |
Default |
Description |
|---|---|---|
|
|
Layer at which token importance is scored ( |
|
|
Fraction of non-prefix-hit tokens recomputed ( |
|
unset |
Pads PARTIAL row counts to a multiple of this bucket
( |
|
required |
The (private) cacheblend-plugin init-container image
( |
|
– |
Pull secrets appended to the vLLM pod for the private payload image. |
|
first container |
Name of the vLLM container to inject into. |
|
|
|
server.chunkSize defaults to 256 and must equal 256 (the blend matcher
requires it).
PD Disaggregation#
The operator has first-class support for PD (Prefill-Decode) disaggregation.
Adding a pd block to an LMCacheEngine spec switches the engine’s
connection ConfigMap to include MultiConnector configs (NixlConnector +
LMCacheMPConnector) alongside the standard bare connector, and tells the
webhook to inject the NIXL side-channel environment variables into opted-in
vLLM pods automatically.
See Disaggregated Prefill for background on what PD disaggregation is and how the pieces fit together.
Note
PD disaggregation requires NIXL in the vLLM environment (lmcache[nixl]
extra) and vllm-project/vllm#46865 merged.
How it works#
Deploy a single LMCacheEngine CR with a pd block. One DaemonSet
instance per node serves both prefiller and decoder vLLM pods. The operator:
Builds a multi-key ConfigMap – the
<engine>-connectionConfigMap emits three keys so the same engine can serve all pod types:kv-transfer-config.json– bareLMCacheMPConnector(fallback for pods without apd-roleannotation; no NIXL).kv-transfer-config-prefiller.json–MultiConnectorwithkv_role=kv_producer.kv-transfer-config-decoder.json–MultiConnectorwithkv_role=kv_consumer.
Injects the correct config via the webhook – the webhook reads the
lmcache.ai/pd-roleannotation on each vLLM pod and injects the matching ConfigMap key as--kv-transfer-config.Injects NIXL env vars – opted-in PD pods receive two extra env vars:
VLLM_NIXL_SIDE_CHANNEL_HOST– set to the pod’s own IP via the downward API (status.podIP).VLLM_NIXL_SIDE_CHANNEL_PORT– taken fromspec.pd.nixlSideChannelPort(default5558). If the pod pre-sets this env var, the webhook will leave it unchanged – useful when both roles run on the same host and need distinct ports (e.g. prefiller5557, decoder5558).
PDSpec Fields#
Field |
Default |
Description |
|---|---|---|
|
|
Port the NIXL agent advertises for handshake negotiation. Injected by
the webhook unless the pod pre-sets |
|
|
|
|
(omitted) |
When set to |
Deploying a PD Engine#
A single LMCacheEngine handles both roles:
apiVersion: lmcache.lmcache.ai/v1alpha1
kind: LMCacheEngine
metadata:
name: lmcache-engine
spec:
l1:
sizeGB: 100
server:
port: 5555
chunkSize: 256
pd:
nixlLoadFailurePolicy: fail
kubectl apply -f engine.yaml
kubectl get lmc # wait for Running
Opting vLLM Pods In#
Use the standard label + annotation pattern (see
Connection Injection (Webhook)), adding the lmcache.ai/pd-role
annotation to select the prefiller or decoder config:
# prefiller vLLM pod template
metadata:
labels:
lmcache.ai/lmcache-inject: "true"
annotations:
lmcache.ai/lmcache-engine: "lmcache-engine"
lmcache.ai/pd-role: "prefiller"
# decoder vLLM pod template
metadata:
labels:
lmcache.ai/lmcache-inject: "true"
annotations:
lmcache.ai/lmcache-engine: "lmcache-engine"
lmcache.ai/pd-role: "decoder"
The webhook injects --kv-transfer-config (the role-specific MultiConnector
JSON), hostIPC: true, PYTHONHASHSEED=0, VLLM_NIXL_SIDE_CHANNEL_HOST,
and VLLM_NIXL_SIDE_CHANNEL_PORT into each opted-in pod. Do not mount
the ConfigMap or add --kv-transfer-config yourself.
Note
NIXL RDMA and hostNetwork – NIXL’s UCX backend requires valid RDMA GIDs, which are derived from the host’s network interfaces. Under standard overlay CNI (each pod has its own network namespace) the GID table inside the pod is empty and UCX backend initialization fails. The recommended workarounds are:
hostNetwork: true+dnsPolicy: ClusterFirstWithHostNet(quick test). With hostNetwork both roles share the host IP, so set distinctVLLM_NIXL_SIDE_CHANNEL_PORTvalues per role (e.g. 5557 / 5558) in the pod env before the webhook runs – the webhook will not override a pre-set value.SR-IOV with Multus (production): assign each pod a dedicated VF with its own GID, no
hostNetworkrequired.
Pods without a lmcache.ai/pd-role annotation that are bound to a PD engine
fall back to the bare LMCacheMPConnector config (no NIXL) – they still
benefit from the LMCache KV cache without participating in disaggregation.
Router#
The vllm-router is not managed by the operator. Deploy it as a plain
Kubernetes Deployment pointing at the prefiller and decoder vLLM Services:
vllm-router \
--policy round_robin \
--vllm-pd-disaggregation \
--prefill http://pd-prefiller.<namespace>.svc.cluster.local:8001 \
--decode http://pd-decoder.<namespace>.svc.cluster.local:8002 \
--host 0.0.0.0 --port 30000
A ready-to-edit manifest is at
operator/config/samples/vllm_pd_disaggregation.yaml.
Note
Name your prefiller and decoder vLLM Services with a prefix other than
vllm- (e.g., pd-prefiller, pd-decoder). Kubernetes injects
<SERVICE_NAME>_* env vars into every pod in the namespace; a vllm-
prefix generates VLLM_* vars that vLLM’s env-var validator flags as
unknown.
LMCacheCoordinator#
The LMCacheCoordinator CRD runs the mp coordinator – a fleet-wide HTTP
service that tracks mp server instances, evicts those whose heartbeats lapse,
performs L2 quota eviction, and hosts the global CacheBlend fingerprint
directory. It is a plain (non-GPU) Deployment exposed through a ClusterIP
Service; engines reach it via coordinator.ref or coordinator.url.
Deploying a Coordinator#
A ready-to-edit manifest lives at
config/samples/lmcache_v1alpha1_lmcachecoordinator.yaml in the operator
repo. A minimal coordinator:
apiVersion: lmcache.lmcache.ai/v1alpha1
kind: LMCacheCoordinator
metadata:
name: my-coordinator
spec:
port: 9300
kubectl get lmcc my-coordinator # shortName: lmcc
Connecting an Engine#
Point an LMCacheEngine / CacheBlendEngine at the coordinator through its
coordinator block. Use ref to name a coordinator in the same namespace
(the operator resolves it to the in-cluster Service URL), or url for an
explicit endpoint:
spec:
coordinator:
ref:
name: my-coordinator # or: url: http://my-coordinator.default.svc:9300
heartbeatInterval: 5 # seconds; must be > 0
l2EventReporting: false # report L2 store/lookup events for fleet eviction
Coordinator CRD Spec Reference#
Topology#
Field |
Default |
Description |
|---|---|---|
|
|
Coordinator pods. The registry is per-process in-memory, so >1 only makes sense behind a shared durable backend. Must be >= 0. |
|
shared engine image |
Runs the same lmcache binary as the engines. |
|
– |
Image pull secret references. |
HTTP Server#
Field |
Default |
Description |
|---|---|---|
|
|
Address the coordinator’s HTTP server binds to. |
|
|
HTTP port (1–65535). |
Membership & Health#
Field |
Default |
Description |
|---|---|---|
|
|
Seconds without a heartbeat after which an instance is evicted. Set
comfortably above the engines’ |
|
|
Seconds between health-check sweeps; |
L2 Quota Eviction#
Field |
Default |
Description |
|---|---|---|
|
|
Seconds between L2 eviction sweeps; |
|
|
Fraction of tracked keys (by count) to evict per cycle, [0.0, 1.0]. |
|
|
Usage fraction of the quota that fires eviction, (0.0, 1.0]. |
Global CacheBlend Directory#
Field |
Default |
Description |
|---|---|---|
|
|
Tokens per chunk for the global CacheBlend directory (the match unit). Must equal the LMCache chunk size the blend servers use. Must be > 0. |
|
|
Positions between match probes. |
Prometheus, Scheduling & Overrides#
Field |
Default |
Description |
|---|---|---|
|
|
Expose the metrics container port. See the note below. |
|
|
Metrics port. |
|
|
Create a ServiceMonitor CR (and headless metrics Service). |
|
|
Scrape interval. |
|
|
|
|
– |
Pod resource requests/limits (no auto-compute; the coordinator is CPU/memory light). |
|
– |
Pod scheduling controls. |
|
– |
Standard pod-shaping fields. |
|
– |
Extra CLI flags (appended last, can override any auto-generated flag). |
Note
The coordinator process does not yet expose a /metrics endpoint. The
Prometheus wiring is present for parity but is only useful once metrics are
added; serviceMonitor.enabled defaults to false.
Coordinator Resources Created#
For an LMCacheCoordinator named my-coordinator:
Resource |
Name |
Purpose |
|---|---|---|
Deployment |
|
Runs the coordinator HTTP server pods. |
Service (ClusterIP) |
|
Fleet-wide discovery on the HTTP port. |
Service (headless) |
|
Prometheus scrape target (when |
ServiceMonitor |
|
Prometheus Operator integration (when |
The status endpoint other components use to reach the coordinator is
http://<name>.<namespace>.svc:<port> (e.g.
http://my-coordinator.default.svc:9300).
Coordinator Status & Conditions#
The status section includes:
phase:
Pending,Running,Degraded, orFailed.replicas / readyReplicas: Pod counts from the Deployment.
endpoint: In-cluster URL for reaching the coordinator.
observedGeneration: Most recent reconciled generation.
conditions:
Available– At least one replica is ready.AllInstancesReady– All desired replicas are ready.ConfigValid– Spec validation passed.
Coordinator Validation Rules#
Field |
Rule |
|---|---|
|
Must be in [1, 65535]. |
|
Must be >= 0. |
|
Must be > 0. |
|
Must be >= 0. |
|
Must be in [0.0, 1.0]. |
|
Must be in (0.0, 1.0]. |
|
Must be > 0. |
Operator vs Manual Deployment#
Concern |
Manual DaemonSet |
LMCacheEngine Operator |
|---|---|---|
|
Must set manually |
Auto-injected (hostPath mount; |
|
Must set manually |
Auto-injected |
Service discovery |
|
Node-local ClusterIP Service + ConfigMap |
vLLM config |
Copy JSON into Deployment |
Mount |
Resource sizing |
Manual calculation |
Auto-computed from |
Prometheus |
Manual ServiceMonitor |
|
Validation |
Runtime errors only |
|
New GPU nodes |
DaemonSet handles it |
DaemonSet handles it (same) |
Security Considerations#
The hostPath /dev/shm mount exposes the host’s shared-memory tmpfs to the
container – a narrower grant than the whole host IPC namespace, but still
host-level access. spec.hostIPC: true (opt-in, default false) instead
exposes the host’s IPC namespace (System V IPC, POSIX message queues): any
process in the container can interact with IPC resources from other processes
on the same host.
Deploy only in trusted environments.
Clusters using Pod Security Standards must allow the
privilegedprofile for the LMCache namespace – thebaselineandrestrictedprofiles reject hostPath volumes andhostIPCalike.spec.privilegeddefaults tofalse. When enabled (required forgpuVendor: amd), the engine container additionally runs privileged, granting it full device access – enable it only where GPU visibility requires it.
Development#
make generate # Generate DeepCopy methods
make manifests # Generate CRD YAML + RBAC
make build # Compile operator binary
make fmt # go fmt
make vet # go vet
make test # Run unit tests
make lint # Run golangci-lint
Pushing a custom operator image:
# Docker Hub
make docker-build docker-push IMG=docker.io/<your-user>/lmcache-operator:latest
make deploy IMG=docker.io/<your-user>/lmcache-operator:latest
# Multi-platform (amd64 + arm64)
make docker-buildx IMG=<your-registry>/lmcache-operator:latest
If your cluster needs pull credentials:
kubectl create secret docker-registry regcred \
--docker-server=<your-registry> \
--docker-username=<username> \
--docker-password=<password> \
-n lmcache-operator-system