Figure: the cover page for the Technical Reference "KV Cache Offloading for LLM Inference", version 4.0, August 2026, by Sam Pooni of San Jose, CA. It presents a distributed endpoint architecture with intelligent caching, attention-aware eviction and CXL.mem acceleration, built on three standards โ CXL 3.0, UCIe and UEC โ and three techniques: per-head eviction, EMA attention scoring and RoPE-aware prefetch. The animated rings, nodes and flow lines behind the title are decorative. All values are an analytical model, not measured.
CXL 3.0
UCIe
UEC
Technical Reference
KV Cache Offloading
for LLM Inference
Distributed endpoint architecture with intelligent caching,
attention-aware eviction, and CXL.mem acceleration
Sam Pooni
San Jose, CA
Version 5.0 ยท August 2026
All quantitative results in this document are an analytical model, not measured benchmarks.