ยฉ 2025โ€“2026 Subramaniyam (Sam) Pooni
All Rights Reserved
Proprietary & Confidential
Technical Reference Documentation

KV-Cache Offloading for LLM Inference
Over CXL Memory

A hardware-measured decomposition of decode-phase KV bandwidth, and a protocol-level proposal for persisting eviction-policy metadata across memory tiers.

3.4ร—
KV byte cost vs weight bytes
Measured, model-derived ยท DGX Spark GB10
87%
Bandwidth achieved vs spec (236.5 of 273 GB/s)
Measured
+34.15 pts
Eviction-metadata hit-rate gain, LFU @ 32 GiB
Simulated, provisional
6ร—
Memory expansion via CXL tiering
Analytical
๐Ÿ’ก The Pitch
Protocol-level persistence of eviction-policy metadata across memory tiers โ†’ Read the pitch
๐Ÿ“– Chapters
CH 0 Executive Summary

The core findings condensed: measured KV bandwidth and cost ratios, the metadata-persistence proposal, and what remains open.

Key Results Scope Evidence Tiers
CH 1 Introduction

Why KV-cache bandwidth and capacity matter for LLM serving, and what this document does and does not claim.

Motivation Document Scope Terminology
CH 2 Background

LLM inference fundamentals โ€” prefill vs. decode, the memory-bandwidth gap, and how existing serving stacks handle KV state.

Prefill vs Decode Memory Wall Existing Approaches
CH 3 System Architecture

The tiered memory system under study: CXL attachment topology, controller roles, and where KV state physically lives.

CXL Topology Controller Design Tier Layout
CH 4 Hardware Measurement Methodology and Results

How the DGX Spark GB10 measurements were collected, the decode-latency model that was fit, and its accuracy and limits.

Methodology GB10 Results Model Fit
CH 5 Bandwidth and Tier Economics

What fraction of spec bandwidth each tier actually delivers, and the cost/capacity tradeoffs across the memory hierarchy.

Achieved Bandwidth Tier Cost Capacity Tradeoffs
CH 6 Preprocessing and Prefill Considerations

How prefill-phase work and input preprocessing interact with a tiered KV memory system, and where they don't.

Prefill Phase Preprocessing Interaction Effects
CH 7 KV State Management โ€” The Protocol Argument

The central proposal: persisting eviction-policy metadata across memory tiers at the protocol level, and the simulated hit-rate case for it.

Eviction Metadata Protocol Design Simulated Results
CH 8 Mixture-of-Experts and KV Cache Interaction

How MoE routing behavior complicates KV-cache residency and eviction decisions relative to dense models.

Expert Routing Cache Interaction MoE-Specific Effects
CH 9 GPU and Controller Integration

Integration points between GPU memory management and a CXL memory controller โ€” mapping, hints, and fault paths.

Memory Mapping Controller Hints Fault Handling
CH 10 The 2026 Landscape

Where commercial CXL products and inference-serving software stand as of 2026, and how this work relates to them.

CXL Products Serving Stacks Positioning
CH 11 Results Summary

All measured, simulated, and analytical results gathered in one place, each labeled with its evidence tier.

Consolidated Results Evidence Tiers Caveats
CH 12 A Decision Procedure

A practical procedure for deciding when a CXL-tiered KV-cache design is worth adopting for a given deployment.

Decision Criteria Deployment Fit Tradeoff Checklist
CH 13 Conclusion

What was established, what remains provisional, and the open questions for future measurement work.

Summary Open Questions Future Work
๐Ÿ“š Technical Appendix