CXL · Module 2
Evolution of CXL
The CXL revision chronology from 1.0 through 4.0 organised around the architectural pressure each release answered, plus the hardware of staying compatible — semantic version comparison, feature gating on the negotiated revision, capability dependencies and clean fallback, all simulated.
Chapter 2.2 established the constraint every CXL revision has had to satisfy: parts already in the field do not change. New capability must be additive and gated, because a device shipped in 2021 will meet hosts built before and after it.
This chapter is the chronology under that constraint. It is deliberately not a feature list — a reader who memorises which bullet belongs to which revision has learned the least durable part. What survives is what problem remained at each step.
1. The One-Sentence Model
CXL's revisions track one widening scope: from a device attached to one host, to devices shared through a switch, to resources composed across a fabric, to refinement of how that memory is operated — with bandwidth doubling underneath the whole sequence and backward compatibility never broken.
The scope, in one line:
direct attach → switching and pooling → fabric scale → operational refinement
(bandwidth doubles underneath, twice)2. What This Chapter Owns
| Question | Owned by |
|---|---|
| What CXL is | 2.1 |
| Why it is a standard | 2.2 |
| What each revision solved | this chapter |
| What is PCIe supplies | 2.4 |
| What the architecture is for | 2.5 |
Switching, pooling and fabric mechanics are named here and taught in Modules 12, 15 and 16. This chapter explains why each arrived when it did.
3. The Chronology
Dates and headline capabilities below are from CXL Consortium announcements. Where a claim is a summary rather than a quotation, it is written as one.
| Revision | Released | Rate |
|---|---|---|
| 1.0 | Mar 2019 | 32 GT/s |
| 1.1 | 2019 | 32 GT/s |
| 2.0 | Nov 2020 | 32 GT/s |
| 3.0 | 2 Aug 2022 | 64 GT/s |
| 3.1 | Nov 2023 | 64 GT/s |
| 3.2 | 3 Dec 2024 | 64 GT/s |
| 4.0 | 18 Nov 2025 | 128 GT/s |
The PCIe generation underneath tracks the rate: 1.x and 2.0 on PCIe 5.0, the 3.x line on PCIe 6.x, and 4.0 on PCIe 7.0. And the pressure each release answered, in order — 1.0 made coherent attach exist at all; 1.1 refined direct attach; 2.0 asked how one device serves many hosts; 3.0 asked how many devices serve many hosts; 3.1 extended fabric reach and added host-to-host; 3.2 addressed operating the memory in practice; 4.0 addressed bandwidth and port aggregation.
Two things are worth noticing before any individual row.
Bandwidth doubled twice and the architecture changed four times. The rate steps are 1.x/2.0 → 3.x → 4.0. The capability steps are more frequent, which is the empirical form of Chapter 1.6's argument that semantics and bandwidth are independent axes.
Backward compatibility is stated in every release. The CXL 3.0 announcement lists "full backward compatibility with CXL 2.0, CXL 1.1, and CXL 1.0"; the 4.0 material carries the same commitment forward. That is the constraint from Chapter 2.2 visible in the record.
4. 1.0 and 1.1 — Coherent Attach Exists
CXL made its public debut in March 2019. The problem it addressed is the whole of Module 1: a device could move bulk data to and from host memory but could not hold coherent cached copies of it, and memory attached to a device could not be system memory.
The scope was deliberately narrow: one device, one host, a direct link. No switching, no sharing, no fabric. That narrowness is why it shipped — the three protocols and the coherence model were already a large enough problem without topology on top.
What remained after 1.x is the obvious next question. If a memory device is useful to one host, and hosts do not all need memory at the same time, why is the device bound to one host for its lifetime?
5. 2.0 — One Device, Many Hosts
Released November 2020, CXL 2.0 introduced switching, memory pooling and support for persistent memory, along with security capabilities including encryption and device authentication.
The architectural content is the first item, and it is bigger than it looks. A switch means a device is no longer bound to the host it is cabled to:
1.x host ──────── device binding fixed at build time
2.0 host ──┐
├── switch ── device binding becomes an assignment
host ──┘This is the release where Chapter 1.1's stranding argument became addressable. Memory sitting unused in one machine while another is capacity-bound is a topology problem; pooling is the response, and pooling requires something between hosts and devices that can be reconfigured.
Note also what pooling introduces that direct attach never needed: something must decide who gets what. Allocation, partitioning and failure domains all enter the architecture here — the system-management cost Chapter 2.1 §11 named as part of the trade.
6. 3.0 — Many Devices, Many Hosts
Released 2 August 2022, CXL 3.0 doubled the data rate to 64 GT/s "with no added latency over CXL 2.0", and expanded scope from switching to fabric.
The announcement's own highlights group into two arcs:
Fabric capabilities — multi-headed and fabric-attached devices, enhanced fabric management, and composable disaggregated infrastructure.
Scalability and utilisation — enhanced memory pooling, multi-level switching, new enhanced coherency capabilities, and improved software capabilities.
The step from 2.0 to 3.0 is the step from a switch to a fabric, and the difference is worth being precise about. A single switch lets several hosts share a set of devices. Multi-level switching and fabric management let a topology be built, managed and reconfigured at rack scale — with peer-to-peer communication so devices can exchange data without staging through a host, which is Chapter 1.3 §14's point about removing a staging hop.
The latency claim deserves emphasis because it is unusual. Doubling a data rate normally costs latency somewhere; stating that it does not is a design goal being met, not a marketing line, and it matters because coherent traffic is latency-sensitive in a way bulk I/O is not.
7. 3.1 and 3.2 — Reach, Then Operations
Two releases that add less headline capability and more of what makes a deployment survivable.
CXL 3.1, announced November 2023, extended fabric capabilities further and added host-to-host communication — a relationship that had not previously existed in the model, where the participants are two hosts rather than a host and a device.
CXL 3.2, released 3 December 2024, is the operations release. Its additions include the CXL Hot-Memory Unit (CHMU) for monitoring memory access behaviour, host-only coherent host-managed device memory (HDM-H), and further security work.
The pattern in 3.2 is worth naming because it recurs in every maturing standard: once a capability is deployed, the next requirement is to observe and manage it. A memory device that works is a starting point; a memory device whose access behaviour can be monitored is one an operator can actually tier, place and troubleshoot. That is a different kind of feature from "pooling", and its arrival is a signal that the earlier capabilities are in real use.
8. 4.0 — Bandwidth, and Ports That Aggregate
Released 18 November 2025, CXL 4.0 doubles the data rate to 128 GT/s on the foundation of PCIe 7.0, achieved by doubling the Nyquist frequency while preserving PAM4 signalling and the flit-based structure with FEC and CRC introduced in CXL 3.0.
Its architectural addition is Bundled Ports — aggregating multiple physical device ports into a single logical entity, so a device can connect to one or more host root ports or switch upstream ports while remaining compatible with existing software models. The Consortium's description notes that bundling lets an internal design double bandwidth by doubling stacks rather than doubling datapath widths or internal frequency.
That last clause is the interesting engineering content. Bundling is a way of buying bandwidth that costs replication rather than frequency or width — a trade a chip designer recognises immediately, because doubling an internal datapath width or clock is expensive in ways that instantiating a second stack is not.
CXL 4.0 also introduces a native x2 width for increased platform fan-out and support for up to four retimers for extended channel reach, and it maintains full backward compatibility with 3.x, 2.0, 1.1 and 1.0.
9. The Hardware of Staying Compatible
A chronology is only interesting to a hardware engineer if it constrains a design, and this one does. Every revision above shipped into a world containing all the previous ones, so four mechanisms have to work: compare revisions correctly, gate features on the negotiated revision, respect dependencies between capabilities, and fall back cleanly when a peer will not agree.
10. RTL 1 — Comparing Revisions
Purpose
Deciding which end is older sounds trivial and is the source of a classic bug.
// Semantic revision comparison, and the wrong way to do it.
//
// `older_ok` compares MAJOR first and uses MINOR only as a tiebreak.
// `older_bad` compares MINOR alone -- the "compared the wrong field first"
// bug, which is correct for every pair that shares a major number and wrong
// for every pair that does not.
module version_compare (
input logic [3:0] a_major,
input logic [3:0] a_minor,
input logic [3:0] b_major,
input logic [3:0] b_minor,
output logic a_older_ok,
output logic a_older_bad,
output logic compare_disagrees
);
assign a_older_ok = (a_major < b_major)
|| ((a_major == b_major) && (a_minor < b_minor));
// BUG shape, kept here deliberately so a testbench can compare the two.
assign a_older_bad = (a_minor < b_minor);
assign compare_disagrees = (a_older_ok != a_older_bad);
endmodulePlacement. Wherever two advertisements meet — host software or host bridge logic after discovery, per Chapter 2.2 §8.
Invariant. The effective revision is never above either endpoint's. Synthesis is two comparators and a small amount of select logic; there is no state.
Simulation evidence
Both comparisons instantiated on the same pairs. Verbatim from the Icarus run:
=== EXP1: semantic revision compare vs comparing minor first ===
2.0 against 3.0 semantic=1 minor-only=0 <-- DISAGREE
3.1 against 3.2 semantic=1 minor-only=1
2.5 against 3.0 semantic=1 minor-only=0 <-- DISAGREE
3.2 against 4.0 semantic=1 minor-only=0 <-- DISAGREEThree of four pairs disagree, and the one that agrees is the dangerous one. 3.1 against 3.2 shares a major number, so comparing minor alone happens to be right — which is exactly the case a bench built around one revision family will exercise. The bug is invisible while every part on the bench is a 3.x part, and it inverts the moment a 2.x or 4.0 peer appears.
Read the third row as the failure in production terms: a 2.5 host meeting a 3.0 device would be judged newer by the broken comparison, so the pair would operate at 2.5's peer's terms incorrectly — enabling behaviour one end does not define.
11. RTL 2 — Gating a Feature on the Negotiated Revision
Purpose
A capability introduced at revision N must not activate against a peer that predates it.
// A feature that only exists from a given revision onwards must not activate
// against a peer that predates it. The revision thresholds here are teaching
// values, not CXL specification requirements.
module feature_by_version #(
parameter int unsigned MIN_MAJOR = 3 // feature introduced at rev 3.0
) (
input logic clk,
input logic rst_n,
input logic [3:0] eff_major,
input logic requested,
input logic locally_supported,
output logic enabled,
output logic enable_denied,
output logic premature_enable_err
);
logic version_ok;
assign version_ok = (eff_major >= MIN_MAJOR[3:0]);
// BOTH terms are required: a feature must be wanted, implemented locally,
// AND legal at the revision the pair negotiated down to.
assign enabled = requested && locally_supported && version_ok;
assign enable_denied = requested && !(locally_supported && version_ok);
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) premature_enable_err <= 1'b0;
else if (enabled && !version_ok) premature_enable_err <= 1'b1;
end
endmoduleThe critical input is eff_major, not this endpoint's own revision. A 4.0 device talking to a 2.0 host has eff_major = 2, and every 3.x-and-later feature must be off — not because the device cannot do it, but because the host has no definition for what it would receive.
Simulation evidence
=== EXP2: a rev-3 feature meeting different effective revisions ===
effective 2.x: enabled=0 denied=1
effective 3.x: enabled=1 denied=0
effective 4.x: enabled=1 denied=0
effective 4.x, not implemented locally: enabled=0 denied=1Three of the four rows are refusals, from three different causes: too old, too old, and not built. Only the conjunction of wanted, implemented and legal at this revision produces an enable, and a design that drops any one term has a bug that appears only against a specific class of peer.
12. RTL 3 — Capability Dependencies
Purpose
Capabilities are not independent. Enabling a dependent one without its prerequisite produces a configuration discovery accepts and the datapath cannot honour.
// Some capabilities require others. GENERIC model: the dependency edges here
// are teaching values.
// bit0 base
// bit1 requires bit0
// bit2 requires bit1 (and therefore bit0, transitively)
// bit3 independent
module cap_dependency #(
parameter int unsigned NFEAT = 4
) (
input logic clk,
input logic rst_n,
input logic [NFEAT-1:0] requested,
output logic [NFEAT-1:0] granted,
output logic [NFEAT-1:0] missing_prereq,
output logic dependency_err
);
// Grant a bit only when its prerequisite chain is also requested.
assign granted[0] = requested[0];
assign granted[1] = requested[1] && requested[0];
assign granted[2] = requested[2] && requested[1] && requested[0];
assign granted[3] = requested[3];
assign missing_prereq = requested & ~granted;
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) dependency_err <= 1'b0;
else if (missing_prereq != '0) dependency_err <= 1'b1;
end
endmoduleSimulation evidence
=== EXP3: capability dependencies ===
request 0111 -> granted 0111 missing 0000
request 0100 -> granted 0000 missing 0100 <-- prereqs absent
request 1000 -> granted 1000 missing 0000 <-- independentThe middle row is the lesson: requesting bit 2 alone grants nothing, because its prerequisite chain is absent. Note that the transitive case is written explicitly — bit 2 requires bit 1 and bit 0 — rather than relying on bit 1's own guard. Chaining guards implicitly works until somebody reorders the assignments.
Verification note. The attack here is every subset of the request vector, which is 2^NFEAT cases and entirely tractable at these widths. Exhaustive is the right strategy when the space is small, and it is a strategy people skip because the module looks obvious.
13. RTL 4 — Falling Back Cleanly
Purpose
When a peer will not agree, back down in defined steps and stop at the floor.
// Negotiate down to a revision both ends can honour, then stay there.
// This is NOT CXL link training, alternate-protocol negotiation, or any
// specification-defined procedure.
//
// 00 PROBE no agreement yet
// 01 TRY proposing a revision
// 10 SETTLED both ends accepted; this is the operating revision
// 11 BASE fell all the way back to the base revision
module fallback_fsm (
input logic clk,
input logic rst_n,
input logic start,
input logic [3:0] own_max_major,
input logic [3:0] peer_max_major,
input logic peer_reject,
output logic [1:0] state_q,
output logic [3:0] proposed_q,
output logic settled,
output logic at_base,
output logic underflow_err
);
localparam logic [1:0] PROBE = 2'b00, TRY = 2'b01, SETTLED = 2'b10, BASE = 2'b11;
localparam logic [3:0] BASE_MAJOR = 4'd1;
assign settled = (state_q == SETTLED) || (state_q == BASE);
assign at_base = (state_q == BASE);
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
state_q <= PROBE; proposed_q <= 4'd0; underflow_err <= 1'b0;
end else begin
case (state_q)
PROBE: if (start) begin
// Never propose above what the peer advertised: the first proposal
// is already the intersection, not our own maximum.
proposed_q <= (own_max_major < peer_max_major) ? own_max_major : peer_max_major;
state_q <= TRY;
end
TRY: begin
if (peer_reject) begin
if (proposed_q <= BASE_MAJOR) begin
// Cannot go lower than the base revision.
state_q <= BASE;
proposed_q <= BASE_MAJOR;
end else begin
proposed_q <= proposed_q - 4'd1;
end
end else begin
state_q <= (proposed_q <= BASE_MAJOR) ? BASE : SETTLED;
end
end
default: ; // SETTLED and BASE are terminal until reset
endcase
if (proposed_q > 4'd0 && proposed_q < BASE_MAJOR) underflow_err <= 1'b1;
end
end
endmoduleBackpressure. There is none in the usual sense — the peer's rejection is the backpressure, and the design's obligation is to make progress toward a terminal state rather than retry the same proposal.
Simulation evidence
=== EXP4: fallback when the peer keeps rejecting ===
own max 4, peer max 3 -> first proposal = 3 (already the intersection)
peer rejects -> proposal 2
peer rejects -> proposal 1
peer rejects -> proposal 1 state=11 at_base=1
further rejects -> proposal 1 at_base=1 underflow_err=0Two properties are visible. The first proposal is already the intersection — a design that opens at its own maximum wastes a round trip and, worse, may propose something the peer has no way to refuse meaningfully. And the descent terminates: at the base revision it stops, enters a terminal state, and further rejections change nothing. underflow_err stays 0 because the floor held.
A fallback that can descend below its floor is worse than one that fails, because it proposes something neither end defines and the resulting behaviour is undefined rather than merely unavailable.
14. Assertions
Icarus does not execute concurrent SVA, so these were not run; the table gives the procedural check.
// V1 — the effective revision never exceeds either endpoint.
a_eff_not_above_either: assert property (@(posedge clk) disable iff (!rst_n)
(eff_major <= host_major) && (eff_major <= dev_major));
// V2 — a feature never activates below the revision that introduced it.
// This is the property Debug Lab 1 violates.
a_no_premature_feature: assert property (@(posedge clk) disable iff (!rst_n)
enabled |-> (eff_major >= MIN_MAJOR));
// V3 — a capability is never granted without its prerequisite chain.
a_prereq_respected: assert property (@(posedge clk) disable iff (!rst_n)
granted[2] |-> (granted[1] && granted[0]));
// V4 — nothing is granted that was not requested.
a_granted_subset: assert property (@(posedge clk) disable iff (!rst_n)
(granted & ~requested) == '0);
// V5 — the fallback never proposes below the base revision.
a_no_underflow: assert property (@(posedge clk) disable iff (!rst_n)
(proposed_q != '0) |-> (proposed_q >= BASE_MAJOR));
// V6 — LIVENESS: the negotiation terminates. Without a floor and a terminal
// state, a rejecting peer produces an unbounded descent.
a_negotiation_terminates: assert property (@(posedge clk) disable iff (!rst_n)
$rose(start) |-> ##[1:MAX_ROUNDS] settled);| SVA | Testbench check | Result |
|---|---|---|
| V1 | effective vs both, every pairing | held |
| V2 | rev-3 feature at effective 2.x | denied; flag stayed 0 |
| V3, V4 | every request subset driven | bit 2 alone granted nothing |
| V5 | peer rejects past the floor | proposal stopped at 1 |
| V6 | sustained rejection | reached BASE and stayed |
V6 is the one people omit. Every other property here is safety — nothing illegal is enabled. V6 is liveness, and a negotiation that never terminates is a link that never comes up, which presents as a dead device rather than as a misbehaving one.
15. Debug Lab
A fabric feature activates against a host that predates it
GATED-ON-OWN-REVISION// This device is 4.0, so its 3.x-and-later features are available.
assign enabled = requested && locally_supported && (own_major >= MIN_MAJOR);Works perfectly against every modern host on the bench. Against an older host, the device emits behaviour the host has no definition for — the host reports an unsupported or malformed request, and the device reports nothing at all. The correct gate refuses it:
effective 2.x: enabled=0 denied=1
effective 3.x: enabled=1 denied=0The gate reads own_major instead of eff_major. Those are the same number whenever both ends are current, which is the entire bring-up environment — so the bug is structurally invisible until an older peer appears, typically at a customer.
This is Chapter 2.2's Debug Lab 1 in its revision form: correct discovery, correct intersection, and then a decision made from local state rather than from the agreement.
Gate on the negotiated revision, so the device behaves as the older end requires:
assign enabled = requested && locally_supported && (eff_major >= MIN_MAJOR);Prevention. Assert enabled |-> (eff_major >= MIN_MAJOR), and run the regression against older peer profiles deliberately. A bench containing only current-revision models cannot find this, and a current-revision bench is the default.
Version comparison inverts as soon as a peer from another major revision appears
COMPARED-MINOR-FIRST// Lower minor means older.
assign a_is_older = (a_minor < b_minor);Correct for every pair on the bench, wrong for most pairs in the field. Both comparisons instantiated on the same inputs:
2.0 against 3.0 semantic=1 minor-only=0 <-- DISAGREE
3.1 against 3.2 semantic=1 minor-only=1
2.5 against 3.0 semantic=1 minor-only=0 <-- DISAGREE
3.2 against 4.0 semantic=1 minor-only=0 <-- DISAGREEThe consequence is that the pair negotiates to the wrong revision — usually the newer one — and then enables behaviour the older end does not implement, so the failure surfaces as Debug Lab 1's symptom one layer down.
A revision is an ordered pair, not a number, and comparing one component of an ordered pair is only valid when the other component is equal. Row two shows exactly that case succeeding by coincidence — 3.1 against 3.2 shares a major, so the shortcut is right, and that is the family a single-generation bench is built from.
The general shape is worth recognising because it recurs: a comparison that is correct on the subset you test and wrong on the set you ship into.
Compare major first, minor as a tiebreak:
assign a_is_older = (a_major < b_major)
|| ((a_major == b_major) && (a_minor < b_minor));Prevention. Test cross-major pairs in both directions — this is one directed test that eliminates the entire bug class. Note also the trap in packing: comparing {major, minor} as one word is safe only while both fields keep their widths and both ends encode them identically, so an assertion written that way is correct under today's parameters and silently wrong under tomorrow's.
A capability is advertised and negotiated, and the datapath behind it was never built
ADVERTISED-WITHOUT-DEPENDENCY// Grant whatever was requested; the feature bits are independent.
assign granted = requested; // BUG: ignores the prerequisite chainDiscovery succeeds, negotiation succeeds, and the first transaction that exercises the dependent capability fails inside the device. The correct model refuses the configuration outright:
request 0100 -> granted 0000 missing 0100 <-- prereqs absentCapabilities were treated as an independent bitmap when they are a dependency graph. A capability whose implementation is built on another cannot function without it, so granting it alone produces a configuration that every discovery step accepts and no datapath can honour.
This is the compliance failure class from Chapter 2.2 §10: the claim and the implementation disagree, and nothing in a self-consistent testbench compares them.
Encode the dependency edges explicitly, including transitively, and report the shortfall:
assign granted[2] = requested[2] && requested[1] && requested[0];
assign missing_prereq = requested & ~granted;Prevention. Assert granted[2] |-> (granted[1] && granted[0]) and (granted & ~requested) == 0. Then drive every subset of the request vector — the space is 2^NFEAT and small enough to be exhaustive, which is the right strategy precisely because the module looks too obvious to need it.
16. How This Appears in Real Engineering
Architecture and product
The recurring decision is which revision to target and which optional capabilities to build, knowing both that older peers exist and that newer ones will. The chronology in Section 3 is the input to that decision: a part aimed at pooled deployments needs the switching-era capabilities; one aimed at direct attach may not. Every capability added is validation surface that never shrinks.
RTL engineer
Four mechanisms, all in Sections 10 to 13: compare revisions semantically, gate on the negotiated revision, encode dependency edges explicitly, and give the fallback a floor and a terminal state. The habit that prevents most of the failures is asking, for every capability gate, "what does this read — my revision or the agreed one?"
Verification engineer
The configuration space is revisions × feature subsets × which end is older, and the bugs live at the boundaries. Three stimulus requirements follow: cross-major pairs in both directions, older-peer profiles deliberately included, and exhaustive subsets of small capability vectors. A single-generation bench cannot find Debug Labs 1 or 2, and a single-generation bench is what a project naturally builds.
Firmware and system software
Never infer capability from a revision number. Chapter 2.2's measurement showed two endpoints at the same revision with different usable features; the revision bounds what is possible and the capability bits say what is there. The other durable rule is that software must handle the fallback outcome rather than assuming the maximum was achieved.
Validation and compliance
Older-generation interoperability is the part that gets squeezed when schedules slip, and it is where the field failures come from. A device that has only been tested against current hosts has been tested against the easiest population it will ever meet.
17. Common Misconceptions
18. Interview Reasoning
19. Summary
The revision history is one widening scope. Direct attach in 1.x made coherent attach exist. 2.0 (November 2020) unbound the device from its host with switching and pooling, which is where stranded capacity first became addressable. 3.0 (2 August 2022) doubled the rate to 64 GT/s with no added latency and turned switching into a fabric — multi-level switching, fabric management, peer-to-peer, composability. 3.1 (November 2023) extended fabric reach and added host-to-host communication. 3.2 (3 December 2024) made the deployed memory observable and manageable. 4.0 (18 November 2025) doubled the rate again to 128 GT/s on PCIe 7.0 and added Bundled Ports.
Two observations survive the dates. The rate doubled twice while the architecture changed at nearly every release — the bandwidth-is-not-semantics argument, visible in a release history. And every release states backward compatibility, which is why the capability layers are cumulative rather than successive.
That compatibility commitment is what reaches into RTL. Four mechanisms carry it, and each has a characteristic failure: compare revisions semantically or a cross-major peer inverts the result (measured: three of four pairs disagree); gate capabilities on the negotiated revision or a feature activates against a host that predates it; encode dependency edges explicitly or a configuration passes discovery and fails in the datapath; and give the fallback a floor or a rejecting peer produces an unbounded descent.
All four failures share a property worth carrying: they are invisible on a bench where every part is current, and a bench where every part is current is what a project naturally builds.
20. What Comes Next
This chapter tracked what CXL added over time. Chapter 2.4 goes underneath it, to what was never re-invented: the PCIe infrastructure every revision above was built on, and the precise sense in which CXL "uses" it. That relationship is the most commonly mis-stated fact about CXL, and getting it right is what makes the rest of the architecture legible.
For adjacent material: What Is CXL? has the definition and device classes, The CXL Consortium has the interoperability contract this chapter's compatibility rules implement, and the PCIe track covers the foundation. The path is on the CXL tutorials index.
Standards & specifications
- Governing standard
- CXL Specification (CXL Consortium)(opens CXL Consortium in a new tab)
Defines CXL.io, CXL.cache and CXL.mem, and the coherence and memory-pooling behaviour built on them. System design and deployment topology are not mandated.
This page also covers RTL structure, verification approach and debugging technique. Those are engineering practice built on the standard, not requirements the standard itself imposes.
Where this fits
Part of the CXL curriculum.