Adding a New Device Backend#
This guide explains how to add a new accelerator to LMCache in Multiprocess (MP) mode.
For basic users: read Part 1 only — adding a
DeviceSpec is all you need to get your device working with the
built-in Python fallback ops.
For advanced users: continue to Part 2 for native ops and advanced transfer modes.
Part 1 — Basic Function Enabling#
For the majority of devices, a single DeviceSpec class is
sufficient. LMCache ships with a complete Python fallback ops
(lmcache/python_ops_fallback.py) that works on any device
supporting standard PyTorch tensor operations — no custom kernels
required.
Prerequisites#
Your PyTorch backend should support:
Device Discovery & Status:
torch.<device>.is_available()→booltorch.<device>.device_count()→int
Device Context & Synchronization:
torch.<device>.set_device(device)→Nonetorch.<device>.current_device()→inttorch.<device>.synchronize()→None
Data Movement:
tensor.to(device)/tensor.cpu()(host↔device transfers)
Add a FooDeviceSpec#
Create lmcache/v1/platform/foo/__init__.py:
# SPDX-License-Identifier: Apache-2.0
"""foo-specific platform primitives."""
from lmcache.v1.platform.base_device_spec import DeviceSpec
class FooDeviceSpec(DeviceSpec):
"""foo device specification for LMCache registry discovery."""
@property
def device_type(self) -> str:
return "foo"
@property
def torch_module_name(self) -> str:
return "foo"
@property
def ops_module(self) -> str | None:
return None
def is_available(self) -> bool:
"""Check backend availability without importing lmcache.__init__."""
try:
import torch
return hasattr(torch, "foo") and torch.foo.is_available()
except Exception:
return False
Key properties:
Property |
Required |
Purpose |
|---|---|---|
|
yes |
Device type string (e.g. |
|
yes |
Attribute on the |
|
no |
Ops module path; leave the base-class default ( |
|
yes |
Returns |
Note
The hasattr(torch, "foo") guard shown above is only needed for
out-of-tree PyTorch extensions (e.g. torch.musa, torch.xpu
when installed as a plug-in). For accelerators shipped inside
PyTorch itself (like torch.cuda) a plain
torch.foo.is_available() is enough.
That’s it. Defining this class is enough for auto-discovery — no
global list or manual registration call is required. All ops
automatically route through lmcache.python_ops_fallback, which is
not performant but is functionally applicable to any device that
supports the standard PyTorch tensor surface. Later, each device can
provide its own ops (Python, C, CUDA, Rust, or any language exposing
Python bindings) to override the fallback function-by-function.
Verification#
Start LMCache server:
lmcache server --l1-size-gb 10 --eviction-policy LRU --port 5555
Run vLLM with MP connector. If multiple accelerators are visible on
the host, set DEVICE_TYPE to force LMCache to pick the new backend
instead of auto-detecting:
export DEVICE_TYPE=foo # optional; only when auto-detection picks the wrong device
vllm serve <your-model> \
--kv-transfer-config '{
"kv_connector": "LMCacheMPConnector",
"kv_connector_module_path": "lmcache.integration.vllm.lmcache_mp_connector",
"kv_role": "kv_both",
"kv_connector_extra_config": {
"lmcache.mp.host": "tcp://localhost",
"lmcache.mp.port": "5555",
"lmcache.mp.mp_transfer_mode": "engine_driven"
}
}' \
--no-enable-prefix-caching \
--port 8000
Note
The default transfer mode is auto (CUDA → LMCache-driven, other
devices → engine-driven). The example above explicitly sets
engine_driven so that a new non-CUDA device works without
additional capability checks. For the lmcache_driven mode
(IPC zero-copy), see Advanced transfer mode.
Check the LMCache logs — with ops_module = None you should see:
torch_dev=..., torch_device_type=foo
Custom ops not supported for device: foo, using fallback ops.
This confirms your DeviceSpec was discovered and the Python
fallback is active. (Once you set ops_module in
Part 2, the second line becomes
Using backend: foo_device.ops instead.)
Debugging checklist:
[ ]
torch.foo.is_available()returnsTrue.[ ] Set
DEVICE_TYPE=footo force selection if not picked up automatically.[ ] Log shows either
Custom ops not supported for device: foo, using fallback ops.(pure fallback) orUsing backend: foo_device.ops(custom ops loaded).[ ] Engine-driven transfer works end-to-end (check the LMCache logs to confirm whether the SHM or Pickle sub-path is chosen — both should succeed).
[ ] Store/retrieve correctness is verified.
[ ] TP>1 / multi-worker behavior is verified.
Part 2 — Performance Optimization#
Once basic functionality is verified, add device-specific optimizations.
Device-specific ops#
The generic ops interface is defined in
lmcache/python_ops_fallback.py. Each vendor may replace any
subset of these functions with a device-specific implementation; the
choice of language (C, CUDA, SYCL, Python, Rust, …) and packaging is
entirely up to the vendor. LMCache only cares that the resulting
symbols are importable from Python.
How it works. At import time, get_backend(device_type) imports
the module named by your DeviceSpec.ops_module and merges its
symbols over python_ops_fallback:
callers → lmcache.c_ops → foo_device.ops (functions you defined)
→ lmcache.python_ops_fallback (everything else)
Integration contract#
Regardless of how you build your ops module, the following contract must hold:
Same symbol names. Every function you override must be exposed under the exact name used in
python_ops_fallback(e.g.multi_layer_block_kv_transfer).Same call signature. Positional/keyword arguments, argument order and semantics must match the fallback; callers invoke the merged module without knowing which backend answered.
Importable Python module.
ops_moduleis a fully-qualified Python module path (importlib.import_modulemust succeed). How the module gets there — a pure-Python file, a pybind11 extension, a ctypes wrapper, a RustPyO3module, etc. — is your choice.Partial override is allowed. You do not have to reimplement every function. Anything you leave out keeps using the Python fallback, so incremental optimization is supported.
Point ``DeviceSpec.ops_module`` at the module once it is reachable on
sys.path:class FooDeviceSpec(DeviceSpec): @property def ops_module(self) -> str | None: return "foo_device.ops" # any importable module path
Implementation notes#
multi_layer_block_kv_transferandlmcache_memcpy_asyncare the hot entry points for both engine-driven and LMCache-driven transfer; other functions inpython_ops_fallbackcan be overridden as needed.If you fall back to the generic path from inside a device-specific wrapper (e.g. when inputs are unsupported), call the corresponding
lmcache.python_ops_fallbackfunction directly to preserve semantics.Keep your ops module free of side effects at import time — it is imported eagerly during
get_backend().
For concrete reference implementations, see
lmcache/v1/platform/cuda/ and lmcache/v1/platform/musa/.
Both are examples of what a vendor may do, not templates every new
backend has to follow.
Advanced transfer mode#
By default, the transfer mode is AUTO: the router dispatches
strictly by device_type — device_type == "cuda" goes to
LMCacheDrivenTransferContext (IPC zero-copy), everything else to
EngineDrivenTransferContext. A non-CUDA device that supports IPC
handle transfer can still opt into LMCache-driven explicitly (below).
Note
ROCm is not a separate backend: under PyTorch a ROCm GPU reports
device_type == "cuda" and reuses CudaIPCWrapper (CUDA IPC)
for the LMCache-driven path, so it works in AUTO mode with no extra
setup. There is currently no dedicated platform/rocm package.
When the caller (or LMCACHE_MP_TRANSFER_MODE) explicitly requests
lmcache_driven, _build_lmcache_driven_context performs two hard
checks — both must succeed, otherwise the factory raises
ValueError (no silent fallback):
Your
DeviceSpecsubclass must bind aDeviceIPCWrappersubclass (exposing awrapclassmethod) viaipc_wrapper_cls.resolve_kv_wrapper_factory()reads that binding off the registered spec — no separate registry / auto-scan.DeviceSpec.is_handle_transfer_available()must returnTrue(the base-class default; override toFalseonly if your device lacks IPC handle transfer).
Separately, the LMCache-driven server module also requires a
BaseCacheContext subclass under
lmcache/v1/platform/foo/cache_context.py and a matching
DeviceSpec.create_cache_context override that lazy-imports and
instantiates it. The platform-agnostic factory
lmcache.v1.platform.cache_context.create_cache_context dispatches
by device_type through the DeviceSpec registry and invokes
that hook; the default DeviceSpec.create_cache_context raises
NotImplementedError so a missing override surfaces loudly instead
of silently falling back. The cache context itself manages the KV
cache layout and pointers used for IPC transfer.
Host-side pinning via pin_memory_backend is optional and only
affects staging-buffer performance; it is not required to enable
LMCache-driven mode.
Override these methods in your DeviceSpec:
class FooDeviceSpec(DeviceSpec):
@property
def ipc_wrapper_cls(self):
"""Bind the DeviceIPCWrapper subclass for this device.
Lazy import so the accelerator-specific module is only
pulled in when the LMCache-driven path is actually used.
"""
from lmcache.v1.platform.foo.ipc_wrapper import FooIPCWrapper
return FooIPCWrapper
def is_handle_transfer_available(self) -> bool:
"""Return True if your device supports IPC handle transfer."""
return True # base-class default; override to False if unsupported
@property
def pin_memory_backend(self):
"""Return a PinMemoryBackend subclass, or None.
Optional; only affects host staging performance.
"""
return None # default
def create_cache_context(self, *args, **kwargs):
"""Lazy-import and instantiate the BaseCacheContext for this device.
Required for LMCache-driven mode; the base-class default
raises ``NotImplementedError``.
"""
from lmcache.v1.platform.foo.cache_context import FooCacheContext
return FooCacheContext(*args, **kwargs)
Opt into LMCache-driven mode by setting lmcache.mp.mp_transfer_mode
to lmcache_driven in the vLLM kv_connector_extra_config shown
in Part 1, or by exporting
LMCACHE_MP_TRANSFER_MODE=lmcache_driven. If either hard check
fails, the factory raises ValueError and refuses to construct the
context — switch back to engine_driven or auto.
References#
Topic |
Path |
|---|---|
Device spec base |
|
Backend loading |
|
Python fallback |
|
Cache context base |
|
Cache context factory |
|
Reference |
|
Reference |
|
Engine-driven call site |
|