Figure: the internal architecture of one CXL 3.0 Type-3 memory-compute endpoint, one of four identical units. A host CPU or GPU issues load and store instructions over PCIe 5.0 x8 or x16 into the endpoint package, where a CXL protocol engine terminates CXL.mem for HDM-D and HDM-DB, CXL.io for mailbox and configuration, and CXL.cache for coherency. A UCIe 1.1 die-to-die link carrying over 1 TB/s joins three chiplets: a memory controller driving 8 DDR5-5600 channels at 44.8 GB/s each for 358.4 GB/s, a compute chiplet of 8 Cortex-A78 cores at 3 GHz with 8 MB of L3, and a control and monitor block holding the policy engine, access tracker and prefetch queue. Alongside them sit 64 MB of on-package eSRAM at 8 nanoseconds for metadata and index storage, and an NVMe controller on PCIe 4.0 x4. Externally the endpoint attaches 256 GB of DDR5 across 8 DIMMs and 4 TB of NVMe flash as the cold tier. The bandwidth hierarchy falls from over 1 TB/s inside the package, to 358.4 GB/s of local DDR5, to 64 GB/s across the external CXL link, to about 14 GB/s of flash. Four such endpoints give 1 TB of DRAM at 200 nanoseconds, 16 TB of flash at 25 microseconds and 256 GB/s aggregate, at $5,000 per endpoint. All values are an analytical model, not measured.

Distributed Endpoint Architecture

CXL 3.0 Type-3 Memory-Compute Node

HOST (CPU/GPU)
Load/Store Instructions
PCIe 5.0 x8/x16
ENDPOINT PACKAGE (UCIe Integrated)
CXL PROTOCOL ENGINE
CXL.mem
HDM-D/HDM-DB
CXL.io
Mailbox/Config
CXL.cache
Coherency
UCIe 1.1 — 1+ TB/s Die-to-Die
MEMORY CONTROLLER
CH 0-3
DDR5-5600
CH 4-7
DDR5-5600
8ch × 44.8 GB/s = 358.4 GB/s
COMPUTE CHIPLET
Core 0-3
A78 @ 3GHz
Core 4-7
A78 @ 3GHz
L3 Cache — 8 MB
CONTROL & MONITOR
Policy Engine
Access Tracker
Prefetch Queue
On-Package eSRAM
64 MB @ 8ns — Metadata & Index Storage
NVMe Controller
PCIe 4.0 x4 to Flash
DDR5 DRAM
256 GB
8× DDR5-5600 DIMMs
NVMe Flash
4 TB
Cold tier — 16 TB across 4 endpoints
DATA FLOW LEGEND
CXL.mem (Data)
CXL.io (Control)
UCIe (Internal)
DDR5 (Memory)
BANDWIDTH HIERARCHY

One of 4 identical endpoints: 256 GB DDR5 @ 200 ns and 4 TB NVMe @ 25 µs each, 64 GB/s each (1 TB DRAM, 16 TB flash, 256 GB/s aggregate). Endpoint unit cost $5,000.
Analytical model — component specifications and derived bandwidths are modeled, not measured.