Raw Block (Rust)#
A built-in L2 adapter that stores KV objects in fixed-size slots on a raw block device or pre-sized file using the Rust raw-device I/O bindings. It reuses the existing raw-block metadata checkpoint model and writes directly into the caller-provided load buffers during prefetch.
Required fields:
device_path: Raw device path or pre-sized file path.slot_bytes: Fixed slot size in bytes. Must be aligned toblock_align.
Optional fields:
capacity_bytes: Optional cap on the usable device bytes. Default0means use the full device/file size.use_odirect:trueorfalse(defaulttrue).block_align: Device alignment in bytes (default4096). Must be a power of two.header_bytes: Per-slot header reservation (default4096).meta_total_bytes: Reserved metadata checkpoint region (default256MiB).meta_magic/meta_version: Metadata checkpoint identity/version knobs.meta_checkpoint_interval_sec/meta_idle_quiet_ms/meta_enable_periodic/meta_verify_on_load: Checkpoint and recovery controls carried over from the legacy raw-block backend.load_checkpoint_on_init: Load an existing on-device metadata checkpoint during startup (defaulttrue). Set tofalseto start with an empty in-memory index instead.enable_zero_copy: Try aligned direct-buffer I/O when possible.io_engine: Rust raw-block I/O engine. Valid values are"posix"(default synchronouspread/pwritepath),"io_uring"(direct Rust io_uring syscall path).use_uring_cmd: Enable NVMe passthrough via io_uring command interface for direct device access. Requiresio_engine="io_uring"and NVMe character device node (e.g.,/dev/ng0n1).iouring_queue_depth: Queue depth forio_engine="io_uring".max_data_transfer_size: Maximum data transfer size foruse_uring_cmd=true. Large transfers are split into smaller chunks that fit within device limits. When left unset (or<= 0), it is auto-detected from the device’s sysfs queue limits, bounded by bothmax_hw_sectors_kbandmax_segments * page_size(the NVMe passthrough path uses one scatter-gather segment per page).fdp_enabled: Enables NVMe Flexible Data Placement (FDP) discovery and non-zero placement identifier registration.cache_saltvalues with":"bucket prefixes use FDP placement by default. Requiresio_engine="io_uring"anduse_uring_cmd=true.fdp_placement_ids: Optional non-zero placement identifier list for KV data placement. If omitted, the adapter uses all device-reported non-zero identifiers exceptmeta_checkpoint_placement_id.fdp_data_placement_policy: KV data placement policy."none"omits FDP directives for KV data writes."cache_salt_prefix"assigns case-insensitivecache_saltprefixes to FDP placement identifiers."cache_salt_rank"keeps that bucket isolation and further separates rank-local write streams within each bucket when identifiers are available. The default is"cache_salt_prefix"whenfdp_enabled=true.fdp_slot_reuse_policy: Free-slot reuse policy."pid_affinity"prefers a free slot last assigned to the same FDP placement identifier before falling back to any free slot."none"disables affinity-based reuse. The default is"pid_affinity"whenfdp_enabled=trueand"none"otherwise.meta_checkpoint_placement_id: Optional non-zero placement identifier for metadata checkpoint payload/header writes. Omit it to keep checkpoint writes on default NVMe placement.num_store_workers/num_lookup_workers/num_load_workers: Worker-thread counts for each operation type.
Notes:
raw_blockis a server-owned MP adapter. It does not support per-TP device-path mappings in MP mode.raw_blockremains"type": "raw_block"for all supported engines.raw_blockowns on-device slot allocation, checkpointing, and recovery throughRawBlockCore. Slot reclamation is driven by the shared/global L2 eviction controller or explicitdelete()calls.slot_bytes,header_bytes, andmeta_total_bytesmust be multiples ofblock_align.If
use_odirectis enabled, the server’s--l1-align-bytesshould be at leastblock_align.With
O_DIRECT, raw-block I/O rejects offsets and total I/O lengths that are not multiples ofblock_align. Misaligned write buffers use an aligned bounce buffer.persist_enabledmust remaintruefor this adapter.For
use_uring_cmd=true,device_pathmust use the NVMe character device node (e.g.,/dev/ng0n1) instead of the block device node (/dev/nvme0n1). The character device provides direct NVMe command passthrough.For
use_uring_cmd=true,block_alignmust be a multiple of the NVMe namespace LBA size. An incompatible value is rejected when the device opens.use_uring_cmdrequiresio_engine="io_uring"to be set.When
use_uring_cmd=true,use_odirectis ignored for NVMe namespace character devices. FDP examples setuse_odirect=falsebecauseio_uring_cmduses NVMe passthrough rather than the POSIX write path.FDP registers only non-zero placement identifiers.
fdp_placement_idsis the KV data placement pool: if omitted, all discovered non-zero identifiers exceptmeta_checkpoint_placement_idare used; if provided, every identifier must be reported by the device and must not contain 0.meta_checkpoint_placement_idmust not overlap withfdp_placement_ids. Keeping metadata checkpoints and KV data on separate placement identifiers avoids mixing long-lived raw-block metadata with cache data buckets.With
fdp_data_placement_policy="cache_salt_prefix",cache_saltvalues containing":"are bucketed by the prefix before":". The bucket name is case-insensitive, soRAG:app1,rag:app2, andrag:share the same bucket. Values without":"and values with an empty prefix use no FDP directive. Buckets are assigned exclusive FDP placement identifiers in first-seen order while identifiers are available.If the number of discovered cache-salt buckets exceeds the number of registered FDP placement identifiers, additional buckets continue to store with no FDP directive and the adapter emits one warning.
report_status()exposes a fallback count and a bounded bucket sample instead of retaining the complete fallback bucket set. Emptycache_saltvalues also use default NVMe placement with no directive.With
fdp_data_placement_policy="cache_salt_rank", placement is keyed by the case-insensitivecache_saltbucket prefix and the local rank encoded inkv_rank. The suffix after":"remains part of the object key but does not affect FDP placement. For example,batch:app-aandbatch:app-bshare thebatchbucket: local rank 0 from both applications shares one placement identifier, local rank 1 shares another, and so on. Each bucket can use at most the detected visible GPU count worth of placement identifiers. Extra ranks, or buckets that cannot obtain an identifier, store with no FDP directive.Sharing placement identifiers within a bucket can be intentional because the available FDP placement-identifier pool is finite. Operators can assign the same prefix to applications with similar cache lifetimes, or use a distinct prefix for each application when placement isolation is required. For example,
app01:<suffix>throughapp16:<suffix>can provide 16 application buckets with eight rank streams each, consuming up to 128 placement identifiers when at least 128 identifiers are available and the adapter detects at least eight visible GPUs. Applications using the same prefix share the placement identifiers for equal local ranks instead of receiving separate sets.Bucket-to-placement assignments are process-local. Restart recovery does not need them for correctness because
cache_saltis part of the object key, but first-seen FDP placement assignments may change after adapter restart.Slot affinity is process-local and is not stored in metadata checkpoints. After recovery, free slots have no recorded affinity until they are reused.
Store and retrieve requests must use the same
cache_saltto address the same object. This is the normal LMCache key identity rule; FDP placement is a write directive and is not used to locate data on reads.Metadata checkpoint writes use
meta_checkpoint_placement_idwhen configured, otherwise they use default NVMe placement with no directive.
Configuration examples:
# Basic raw_block with posix I/O
--l2-adapter '{"type": "raw_block", "device_path": "/dev/nvme0n1", "slot_bytes": 1048576, "block_align": 4096, "header_bytes": 4096, "meta_total_bytes": 268435456, "use_odirect": true, "num_store_workers": 2, "num_lookup_workers": 1, "num_load_workers": 4}'
# With io_uring
--l2-adapter '{"type": "raw_block", "device_path": "/dev/nvme0n1", "slot_bytes": 1048576, "io_engine": "io_uring", "iouring_queue_depth": 256, "use_odirect": true}'
# With io_uring_cmd (NVMe passthrough)
--l2-adapter '{"type": "raw_block", "device_path": "/dev/ng0n1", "slot_bytes": 1048576, "io_engine": "io_uring", "use_uring_cmd": true, "iouring_queue_depth": 256, "max_data_transfer_size": 131072, "use_odirect": false}'
# With FDP discovery and cache_salt prefix placement enabled
--l2-adapter '{"type": "raw_block", "device_path": "/dev/ng0n1", "slot_bytes": 1048576, "io_engine": "io_uring", "use_uring_cmd": true, "fdp_enabled": true, "use_odirect": false}'
# With FDP discovery and cache_salt/rank placement enabled
--l2-adapter '{"type": "raw_block", "device_path": "/dev/ng0n1", "slot_bytes": 1048576, "io_engine": "io_uring", "use_uring_cmd": true, "fdp_enabled": true, "fdp_data_placement_policy": "cache_salt_rank", "use_odirect": false}'
# With FDP discovery only, keeping KV data writes on default NVMe placement
--l2-adapter '{"type": "raw_block", "device_path": "/dev/ng0n1", "slot_bytes": 1048576, "io_engine": "io_uring", "use_uring_cmd": true, "fdp_enabled": true, "fdp_data_placement_policy": "none", "use_odirect": false}'
# With eviction
--l2-adapter '{"type": "raw_block", "device_path": "/dev/nvme0n1", "slot_bytes": 1048576, "load_checkpoint_on_init": false, "eviction": {"eviction_policy": "LRU", "trigger_watermark": 0.9, "eviction_ratio": 0.1}}'
Hardware-gated FDP status validation:
FDP live-device validation is opt-in because it requires an FDP-capable NVMe
namespace character device. The status probe opens the character device through
the Rust raw-block binding and calls fetch_fdp_status() with a read-only file
descriptor. It does not issue writes, initialize the MP adapter layout, write KV
data, or verify the MP adapter’s cache-salt placement policy.
LMCACHE_TEST_FDP_CHAR_DEVICE=/dev/ng0n1 \
pytest -q tests/v1/storage_backend/test_raw_block_fdp_status_probe.py
When the variable is not set, the test skips. If the configured device, kernel, or controller does not support the FDP status query, the test skips with the underlying capability error. A passing status probe only confirms that the live device can answer the FDP status query; full adapter initialization and KV write placement on hardware are separate validations.