Figure: the single-node distributed endpoint architecture. A host system pairs a CPU control plane with a B200 GPU holding 192 GB of HBM3e at 8 TB/s and a CXL root port on x16 Gen5; that root port feeds a CXL 3.0 switch, which fans out over four links to four identical CXL Type-3 endpoints. The switch is specified by what the design requires of it rather than by a backplane number: it must be non-blocking across four by-16 Gen5 downstream ports plus the host upstream port, so that all four endpoints can be driven concurrently at their full 64 GB per second each. Each endpoint carries 8 Cortex-A78 cores, 256 GB of DDR5 at 200 nanoseconds and 4 TB of NVMe flash at 25 microseconds, so the four together add 1 TB of DRAM, 16 TB of flash and 256 GB/s of aggregate bandwidth. The point is that steady-state demand for 16 users at 10 tokens per second is only 11.4 GB/s, 4.5 percent of what is supplied, so the fabric is capacity-driven rather than bandwidth-driven. All values are an analytical model, not measured.

Distributed Endpoint Architecture

Single-Node Configuration: GPU + CXL Switch + 4 Type-3 Endpoints

HOST SYSTEM
CPU
Control Plane
B200 GPU
192 GB HBM3e
8 TB/s
CXL Root
x16 Gen5
CXL 3.0 Switch
Non-blocking across 4 × ×16 Gen5 downstream ports + host upstream port
Endpoint 0
CXL Type-3
ARM Cores
8× Cortex-A78
DDR5 DRAM
256 GB @ 200 ns
NVMe Flash
4 TB @ 25 μs
Endpoint 1
CXL Type-3
ARM Cores
8× Cortex-A78
DDR5 DRAM
256 GB @ 200 ns
NVMe Flash
4 TB @ 25 μs
Endpoint 2
CXL Type-3
ARM Cores
8× Cortex-A78
DDR5 DRAM
256 GB @ 200 ns
NVMe Flash
4 TB @ 25 μs
Endpoint 3
CXL Type-3
ARM Cores
8× Cortex-A78
DDR5 DRAM
256 GB @ 200 ns
NVMe Flash
4 TB @ 25 μs
1 TB
Total DRAM
16 TB
Total Flash
256 GB/s
Aggregate BW
200 ns
DRAM Latency

4 endpoints x 256 GB DDR5 = 1 TB; x 4 TB NVMe = 16 TB; x 64 GB/s = 256 GB/s aggregate. Steady-state CXL demand for 16 users at 10 tok/s is 11.4 GB/s — 4.5% of what is supplied, so the fabric is capacity-driven, not bandwidth-driven.
On the switch: this design states no backplane bandwidth figure, because none follows from the topology and none was sourced. The requirement is a port-count and blocking property, not a headline number: the switch must be non-blocking across its four ×16 Gen5 downstream ports plus the host upstream port, so that all four endpoints can be driven concurrently at 64 GB/s each — the 4 × 64 GB/s = 256 GB/s aggregate quoted above. Any commercially available CXL 3.0 switch that meets that port and blocking requirement satisfies the design; the internal switching capacity needed to achieve it is a vendor implementation matter, not a claim this package makes. [Design target]
Analytical model — capacities, bandwidths and latencies are modeled from vendor specifications. Nothing here is measured.