Figure: the single-node distributed endpoint architecture. A host system pairs a CPU control plane with a B200 GPU holding 192 GB of HBM3e at 8 TB/s and a CXL root port on x16 Gen5; that root port feeds a CXL 3.0 switch, which fans out over four links to four identical CXL Type-3 endpoints. The switch is specified by what the design requires of it rather than by a backplane number: it must be non-blocking across four by-16 Gen5 downstream ports plus the host upstream port, so that all four endpoints can be driven concurrently at their full 64 GB per second each. Each endpoint carries 8 Cortex-A78 cores, 256 GB of DDR5 at 200 nanoseconds and 4 TB of NVMe flash at 25 microseconds, so the four together add 1 TB of DRAM, 16 TB of flash and 256 GB/s of aggregate bandwidth. The point is that steady-state demand for 16 users at 10 tokens per second is only 11.4 GB/s, 4.5 percent of what is supplied, so the fabric is capacity-driven rather than bandwidth-driven. All values are an analytical model, not measured.
Distributed Endpoint Architecture
Single-Node Configuration: GPU + CXL Switch + 4 Type-3 Endpoints
4 endpoints x 256 GB DDR5 = 1 TB; x 4 TB NVMe = 16 TB; x 64 GB/s = 256 GB/s aggregate.
Steady-state CXL demand for 16 users at 10 tok/s is 11.4 GB/s — 4.5% of what is supplied,
so the fabric is capacity-driven, not bandwidth-driven.
On the switch: this design states no backplane bandwidth figure, because none follows
from the topology and none was sourced. The requirement is a port-count and blocking property, not a
headline number: the switch must be non-blocking across its four ×16 Gen5 downstream ports
plus the host upstream port, so that all four endpoints can be driven concurrently at 64 GB/s
each — the 4 × 64 GB/s = 256 GB/s aggregate quoted above. Any commercially available
CXL 3.0 switch that meets that port and blocking requirement satisfies the design; the internal
switching capacity needed to achieve it is a vendor implementation matter, not a claim this package
makes. [Design target]
Analytical model — capacities, bandwidths and latencies are modeled from vendor specifications. Nothing here is measured.