Figure: bandwidth aggregation across a CXL 3.0 switch topology. A single x16 PCIe 5.0 endpoint is capped at 64 GB/s, so a B200 GPU is fanned out through a CXL 3.0 switch to four endpoints of 64 GB/s each, a theoretical aggregate of 256 GB/s and a combined 1 TB of DRAM plus 16 TB of flash. Against a Llama-70B per-layer transfer of 2.22 GB โ attention 302 MB, FFN 1.41 GB, KV slice 512 MB โ inside a 5 to 10 millisecond layer compute budget, the required bandwidth is 222 to 445 GB/s: one endpoint takes 34.8 ms and stalls the GPU, two endpoints take 17.4 ms and still stall, four endpoints take 8.7 ms and fit a 10 ms budget. The counterpoint is that steady-state decode demand is only 11.4 GB/s for 16 users at 10 tokens per second with a 3 percent miss tail, which is 4.5 percent fabric utilization, so the four endpoints are chosen for capacity, 1 TB DRAM and 16 TB flash, not for the 256 GB/s ceiling. The 140 GB of FP16 weights stay HBM-resident and never cross CXL. All values are an analytical model, not measured.
Bandwidth Aggregation
Linear scaling through CXL 3.0 switch topology
Single Endpoint Limitation
CXL bandwidth per endpoint is constrained by PCIe lane count. A single x16 Gen5 endpoint maxes out regardless of internal capability.
Analytical model โ all bandwidths, capacities and times are derived from link budgets and vendor specifications, not measured on hardware. Aggregate figures are theoretical ceilings. Canonical numbers v4.0.