Skip to content

UCIe · Module 23

UCIe vs NVLink

Why comparing UCIe with NVLink is usually a scope error — NVIDIA’s own material distinguishes a multi-GPU system interconnect from a chip-to-chip variant, and only one of those sits at UCIe’s layer. The valid comparison, the two-level topology where both coexist, and the retry-at-two-layers failure that turns one operation into two.

Chapter 23.3's error was vertical — a fabric compared against a link. This one is a scope error, and it has a twist: "NVLink" is not one thing, and NVIDIA's own material says so.

1. The One-Sentence Model

A die-to-die link connects dies inside a package. A scale-up interconnect connects accelerators across a system. These are different scopes with different reach, topology, failure and addressing requirements — and NVIDIA's own documentation distinguishes a multi-GPU system interconnect from a chip-to-chip variant. Only the second sits at UCIe's layer.

So most "UCIe vs NVLink" comparisons compare the wrong pair. §5 disambiguates, §7 states the valid comparison — and the valid one turns out to be unusually well-sourced, because NVIDIA frames the two as alternatives itself.

2. What This Chapter Owns

QuestionWhere
Layer normalisation as a method23.1 — UCIe vs PCIe
Standard vs private boundary; the organisational argument23.2 — UCIe vs Proprietary D2D
Fabric vs link; obligation survival across recovery23.3 — UCIe vs Infinity Fabric
Accelerator packages as bandwidth graphs; traffic classes22.3 — AI Accelerators on UCIe
Identity, generation, ordering across a die boundary22.4 — Data-Centre Processors
Physical vs semantic events; duplicate delivery21.6 §19
Open chiplet ecosystems24.1 (next)

Three things are new here:

Scope disambiguation (§5) — the distinction that makes the rest of the chapter possible, taken from NVIDIA's own material rather than asserted.

The two-level topology (§9–§11), where a package-internal link and a system-scale interconnect are complementary layers on one path, and a router must know which is which.

And retry at two layers (§13) — an advanced failure where two independently correct retry mechanisms produce one duplicated operation, and the fix is an identity hierarchy rather than a fix to either retry.

3. Sourcing

4. Claim-vs-Evidence — NVIDIA's Own Framing

ClaimExact evidenceLevelSource / dateEstablishesDoes not establish
NVLink-C2C is a chip-to-chip and die-to-die interconnect"an ultra-fast chip-to-chip and die-to-die interconnect that will allow custom dies to coherently interconnect to the company's GPUs, CPUs, DPUs, NICs and SOCs"CNVIDIA newsroom, 22 Mar 2022it targets the die/chip boundarynothing about the scale-up link
NVLink-C2C extends NVLink to chip-to-chip"extends the industry-leading NVLink technology to a chip-to-chip interconnect"CNVLink-C2C page, Aug 2026the two are distinct and related — §5that they are interchangeable
NVIDIA states support for UCIe"NVIDIA will also support the developing Universal Chiplet Interconnect Express (UCIe) standard announced earlier this month."CNVIDIA newsroom, 22 Mar 2022a public position, datedimplementation, schedule, or any product
NVIDIA frames UCIe and NVLink-C2C as alternatives"Custom silicon integration with NVIDIA chips can either use the UCIe standard or NVLink-C2C, which is optimized for lower latency, higher bandwidth and greater power efficiency."CNVIDIA newsroom, 22 Mar 2022the two occupy the SAME decision point — §7the comparative claim is NVIDIA's, unquantified against UCIe
Bandwidth"coherent interconnect bandwidth of 900 gigabytes per second or higher"CNVIDIA newsroom, 22 Mar 2022a 2022 figure for NVLink-C2Cnot current-generation; not comparable to any UCIe figure (§8)
Efficiency vs PCIe"up to 6x more energy efficiency and 3.5x more area efficiency than a PCIe Gen 6 PHY on NVIDIA chips"CNVLink-C2C page, Aug 2026a PHY-level, generation-stated vendor comparisonnothing about UCIe — the baseline is PCIe Gen 6
Protocols carried"works with Arm's AMBA CHI or CXL industry-standard protocols"CNVIDIA newsroom, 2022; page, 2026it is a transport carrying named protocols — §6that it defines those protocols
Coherence and atomics"coherent data transfers between processors and accelerators"; "supports atomics"CNVLink-C2C page, Aug 2026memory-semantic capability is claimedthe mechanism
Physical extensibility"extensible from PCB-level integration, multi-chip modules (MCM), and silicon interposer or wafer-level connections"CNVLink-C2C page, Aug 2026it spans several packaging scopesa specific product's usage

Four readings, and the fourth is why this chapter can be sharper than 23.3.

Row 4 is remarkable and it is the chapter's foundation. NVIDIA itself places UCIe and NVLink-C2C at the same decision point — custom silicon integration. That is a vendor validating the layer alignment, which is much stronger than my asserting it.

Row 4's comparative clause is a vendor claim and is reported as one. "Optimized for lower latency, higher bandwidth and greater power efficiency" is NVIDIA's characterisation of its own product relative to the standard, dated 2022, and unquantified. I neither endorse nor dispute it; I attribute it.

Row 6 is a usable numeric claim precisely because it is fully qualified: a PHY-level comparison, against a named PCIe generation, on NVIDIA chips, with both quantities stated. It is also not about UCIe at all — which is why it cannot be repurposed into the comparison this chapter is nominally about (§8).

And row 7 shows the structural parallel. NVLink-C2C carries AMBA CHI or CXL. UCIe carries PCIe and CXL natively. Both are transports beneath named protocols — which is exactly the alignment §7 needs.

The disambiguation the whole chapter rests on, and it comes from NVIDIA's own material.

NVLinkNVLink-C2C
described asmulti-GPU communication at system scale"chip-to-chip and die-to-die interconnect"
scopeacross a systemchip and die boundaries
typical roleconnecting accelerators, with switching in larger topologiesconnecting custom dies to NVIDIA GPUs, CPUs, DPUs, NICs, SoCs
at UCIe's layer?noyes — and NVIDIA frames them as alternatives (§4 row 4)

Three properties.

The distinction is NVIDIA's, not mine. The product page says C2C "extends" NVLink "to a chip-to-chip interconnect"extension implies the base was something else.

Nearly every popular comparison uses the wrong member of the pair. Comparing UCIe against the system-scale link produces exactly 23.3 §9's error — a graph property against an edge property — and no arithmetic relates them.

And the right member makes the comparison legitimate, which is the unusual and pleasant result: there is a real, same-layer comparison here, and a vendor has already told you where it sits.

6. Normalising the Stacks

RowNVLink-C2CUCIe
protocol carriedAMBA CHI or CXL (§4 row 7)PCIe and CXL natively mapped
transportvendor-definedD2D Adapter — CRC + retry, optional
physicalvendor-definedspecified signalling
scopechip-to-chip / die-to-diedie-to-die, in-package
who must agreeNVIDIA and a partnerany two conforming implementations
specificationnot publicmembership-gated, but a shared document exists

Two readings.

Rows 1 and 4 align, which is what makes §7's comparison valid. Both are transports beneath a named protocol, at a chip/die boundary.

And rows 5–6 are 23.2's trade, restated with a specific vendor. This is a standard versus a vendor interface comparison — and NVIDIA's own framing (§4 row 4) presents exactly that choice to a customer. The chapter is therefore 23.2 §8's organisational question made concrete.

7. The Valid Comparison

Compare UCIe with NVLink-C2C, not with NVLink. Both are transports at a chip/die boundary beneath a named protocol, and NVIDIA presents them as alternatives for the same decision (§4 row 4).

DimensionNVLink-C2CUCIe
the decision it servesintegrating custom silicon with NVIDIA chipsintegrating dies with any conforming die
who you must agree withNVIDIAthe specification
performance positioningNVIDIA claims lower latency, higher bandwidth, greater power efficiency (2022, unquantified vs UCIe)no counter-claim made here
ecosystem reachwithin NVIDIA's integration programmeany conforming implementation
specification accessnot publicmembership-gated

Three readings.

Row 1 is the substantive difference and it is not a performance difference. One connects you to a specific vendor's silicon; the other connects you to a class of conforming implementations. 23.2 §8's organisational boundary decides which you want.

Row 3 is reported, not evaluated. I have no independent basis to confirm or dispute NVIDIA's 2022 comparative claim, and inventing one would be worse than leaving it attributed.

And row 2 is the real question a customer faces. "Agree with NVIDIA" is not a criticism — it is a precise statement of the contract's counterparty, and for a partner building specifically to attach to NVIDIA silicon it may be exactly the right one.

8. Why Bandwidth Numbers Cannot Be Compared Here

Condition (23.1 §13)Met?
both numbers measure comparable quantitiesno — I have no UCIe figure at all (§3)
generation stated for bothno — the NVIDIA figure is 2022
topology and scope stated for bothpartly
raw line rate vs delivered payload explicitno
both from comparable source levelsno — Level C vs Level B

So no comparison is made, and three specific traps are worth naming:

Aggregate against per-link. A system-scale accelerator figure aggregates many links across a topology; a die-to-die figure is one boundary. 23.3 §9's graph-versus-edge error.

Generation mixing. The 900 GB/s figure is dated March 2022 and described as "or higher". Quoting it as a current-generation capability, or against a current UCIe revision, builds a table from two eras — which is very common in practice and indefensible.

And borrowing the PCIe baseline. §4 row 6 is a real, well-qualified comparison — against a PCIe Gen 6 PHY. It says nothing about UCIe, and transitively concluding anything about UCIe from it is an inference with no support.

9. Two Levels, One Path

A conceptual diagram of a two-level topology. On the left, package A contains two compute dies and a bridge die connected by die-to-die links. On the right, package B has the same structure. Between the two packages, the bridge dies are connected by a system-scale interconnect, optionally through a switch. A label marks the die-to-die links as package-internal scope and the interconnect between packages as system scope. A note states that an operation crossing both needs a semantic identity that outlives either transport. The whole diagram is marked conceptual.Compute die A1package ASystem scopescale-up / switchCompute die B1package BBridge die AD2D inside (§11)Bridge die BD2D inside (§11)Compute die A2package ACompute die B2package BIdentity spansboth hops (§12)12
A conceptual two-level topology. Inside each package, dies are connected by a die-to-die link; between packages, accelerators are connected by a system-scale interconnect. The two are complementary layers on one end-to-end path, not alternatives — and a transaction that crosses both needs an identity that outlives either transport. This arrangement is conceptual and is not attributed to any vendor's product.

Three things to read.

The diagram is CONCEPTUAL and is not attributed to any vendor's product (§3). It is the arrangement the two scopes imply.

The two technologies appear on the same path, not in competition. An operation from A1 to B1 crosses a die-to-die link, a system-scale hop, and another die-to-die link. Neither could do the other's job, and §10 is why.

And the muted node is §12's problem. A transaction crossing both hops needs an identity that outlives either transport — which is the setup for §13's failure.

10. Why the Scopes Cannot Substitute

RequirementPackage-internalSystem scale
reachmillimetres, controlled substrateacross a chassis or rack
topologyfixed at design timeswitched, reconfigurable
addressinga small, fixed die grapha larger, discoverable space
failure modeldies fail with the packagea peer can fail or be absent alone
error handlingshort channel, low error ratelonger channel; different budget
hot-plug / servicenot applicablemay be required

And rows 2, 4 and 6 are the ones that make substitution incoherent. A die-to-die link has no switching, no peer-absence model and no service story — not as deficiencies but because a fixed in-package graph never needed them. 23.1 §10's argument, one scope further out.

11. Illustrative — Scope-Aware Routing

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// ILLUSTRATIVE ONLY. A router that selects a transport by DESTINATION SCOPE.
// The scope is a property of the destination, not of the traffic's contents.
typedef enum logic [1:0] {
  SCOPE_LOCAL_DIE,      // stays on this die — no transport at all
  SCOPE_LOCAL_PACKAGE,  // another die in this package -> D2D
  SCOPE_REMOTE_DEVICE   // another package/accelerator -> scale-up
} dest_scope_e;
 
typedef struct packed {
  logic [SEM_W-1:0]   sem_id;
  logic [GEN_W-1:0]   generation;
  logic [NODE_W-1:0]  dst_node;
  dest_scope_e        scope;        // resolved ONCE, at issue
  logic [EPOCH_W-1:0] route_epoch;
} routed_msg_t;
 
// Scope is resolved from the destination node, once, and then CARRIED.
// Re-deriving it downstream is how two blocks come to disagree (22.3 §11).
function automatic dest_scope_e resolve_scope(logic [NODE_W-1:0] dst);
  if (dst == LOCAL_NODE_ID)                    return SCOPE_LOCAL_DIE;
  else if (node_in_this_package(dst))          return SCOPE_LOCAL_PACKAGE;
  else                                         return SCOPE_REMOTE_DEVICE;
endfunction
 
// Transport selection is a pure function of scope. Note there is NO default
// that silently routes an unknown scope somewhere — §12's orphan starts here.
always_comb begin
  d2d_valid    = 1'b0;
  scaleup_valid= 1'b0;
  local_valid  = 1'b0;
  scope_error  = 1'b0;
  unique case (msg.scope)
    SCOPE_LOCAL_DIE:     local_valid   = msg_valid;
    SCOPE_LOCAL_PACKAGE: d2d_valid     = msg_valid;
    SCOPE_REMOTE_DEVICE: scaleup_valid = msg_valid;
    default:             scope_error   = msg_valid;   // observable, not silent
  endcase
end

Architecture. Scope resolved once at issue, carried with the message, and used as the sole input to transport selection.

State. None here; the message carries it.

Cycle/event behaviour. resolve_scope runs at issue. Downstream blocks read msg.scope and never re-derive it22.3 §11's captured-not-recomputed rule.

Contract. The default arm raises scope_error rather than routing somewhere. A silent default is how a message with an unresolvable destination becomes an orphan21.6 §31's causeless event, seen from the transmitting side.

Failure. §12.

DV/debug. scope in the trace event (21.7 §16) makes a mis-routing visible directly. A SCOPE_LOCAL_PACKAGE message observed on the scale-up interface is a one-line finding instead of a latency mystery.

12. Wrong RTL — One Transport for Every Destination

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// WRONG — one transport abstraction, chosen for "simplicity". Every
// destination goes through the same path and scope is inferred later.
assign uniform_valid   = msg_valid;
assign uniform_payload = pack(msg);
// ... the transport layer figures out where it goes from the address.

Two failures, in opposite directions.

FailureMechanismCost
local traffic takes the external patha package-internal control message is framed and issued for the scale-up transportenormous latency for a message that never needed to leave the package — and it consumes scarce external bandwidth
external traffic inherits internal assumptionstimeouts, error handling and buffering sized for a short in-package hopspurious timeouts and inadequate error handling on a longer, higher-error path

Four properties.

Both failures come from the same design decision, which is what makes it seductive: one abstraction looks cleaner than three.

The first is a performance disaster with a correctness-shaped symptom. A control message that retires work (22.3 §10) taking a system-scale round trip can starve the local engine — the completion path is exactly the class that must not be delayed.

The second is a correctness problem. 23.1 §16's completion_bound differs by scope; a timeout derived for an in-package hop fires spuriously on an external one, and a spurious timeout is 22.4 §11's identity-reuse trigger.

And the fix is §11's: scope is a property of the destination, resolved once, carried, and used to select among transports with different contracts — not one transport with a runtime guess.

13. Failure — Retry at Two Layers

The most advanced failure in Module 23, and neither retry mechanism is wrong.

StepEvent
1semantic operation X issued from package A to package B
2crosses the D2D hop; the D2D transport retransmits once — normal, correct
3crosses the scale-up hop; that transport also retransmits once — normal, correct
4B's receive path sees physical attempts and must decide what they mean
5without a per-boundary attempt identity, the two retries look like two operations
6X is executed twice — or counted twice, or completed twice
7the system reports a duplicate semantic delivery, and blames one of the transports

Five readings.

Both retries are the reliability mechanisms working (21.6 §19). Neither is a defect, and "fixing" either weakens the link.

The defect is an identity gap. There is one semantic operation, and it has two independent physical attempt histories — one per transport boundary. A single flat attempt counter cannot represent that.

The fix is a hierarchy (§14): sem_id + generation identify the operation, and each boundary carries its own attempt identity. A retry increments the attempt at its boundary and never touches the semantic identity.

Step 7 is why this is hard to debug. The evidence points at a transport, and the transport is innocent — 21.6 §20's duplicate-delivery checklist, with an extra layer.

And it only occurs when both boundaries retry on the same operation, which needs load and a marginal condition on both — so it survives testing and appears at scale.

14. Corrected RTL — an Identity Hierarchy

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// CORRECTED. One semantic identity; a SEPARATE attempt identity per transport
// boundary. This is 21.6 §18's physical-vs-semantic distinction extended to
// two transports.
typedef struct packed {
  // SEMANTIC — one per operation, stable across every hop and every retry
  logic [SEM_W-1:0]  sem_id;
  logic [GEN_W-1:0]  generation;
  // PER-BOUNDARY attempt identities — independent, and neither is semantic
  logic [ATT_W-1:0]  d2d_attempt;      // retries on the package-internal hop
  logic [ATT_W-1:0]  scaleup_attempt;  // retries on the system-scale hop
  dest_scope_e       scope;
  logic [EPOCH_W-1:0] route_epoch;
} two_level_msg_t;
 
// Each transport increments ONLY its own attempt field. The semantic identity
// is never touched by a transport event.
always_ff @(posedge clk or negedge rst_n) begin
  if (!rst_n) begin
    d2d_attempt_q     <= '0;
    scaleup_attempt_q <= '0;
  end else begin
    if (d2d_retx_fire)     d2d_attempt_q     <= d2d_attempt_q + ATT_W'(1);
    if (scaleup_retx_fire) scaleup_attempt_q <= scaleup_attempt_q + ATT_W'(1);
    // NOTE: no branch here writes sem_id or generation. That absence IS the fix.
  end
end
 
// Delivery is de-duplicated on the SEMANTIC pair, never on attempt fields.
logic already_delivered;
assign already_delivered = delivered_set_contains(msg.sem_id, msg.generation);
 
assign do_execute = rx_valid && !already_delivered;

Architecture. A three-level identity — operation, and one attempt counter per transport boundary — so a retry is representable without being confusable with a new operation.

State. Two attempt counters plus the receiver's delivered set.

Cycle/event behaviour. Each transport increments only its own field. No branch writes sem_id or generation — and that absence is the correctness argument, not an omission.

Contract. De-duplication keys on sem_id + generation only. Including an attempt field in the key defeats the entire mechanism, because each retry then presents a different key and every attempt executes.

Failure. A single shared attempt counter incremented by both transports makes the two histories indistinguishable — §13. And de-duplicating on the wrong key is worse than not de-duplicating, because it looks implemented.

DV/debug. Both attempt fields belong in the trace (21.7 §16). A trace showing d2d_attempt = 2, scaleup_attempt = 2, sem_id unchanged, and one delivery, is the system working — and being able to see that at a glance is what stops an engineer blaming a transport.

15. End-to-End Latency

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
ILLUSTRATIVE symbolic decomposition. No numeric value from either technology
is used (§3).
 
  t_end_to_end  =  t_src_die                (protocol processing, queueing)
                +  t_d2d_A                  (package-internal hop, package A)
                +  t_bridge_A               (scope transition — §11)
                +  t_scaleup                (system-scale hop, incl. switching)
                +  t_bridge_B
                +  t_d2d_B                  (package-internal hop, package B)
                +  t_dst_die
                +  ...the return path
 
  Improving t_d2d_A alone changes the total by at most t_d2d_A's share.
  In a multi-package operation, t_scaleup and the bridges frequently dominate.

Three readings.

No single technology owns this number, so attributing end-to-end latency to "NVLink" or "UCIe" is meaningless for a path that crosses both.

The bridge terms are real and are often forgotten. A scope transition is protocol work (23.1 §11's t_bridge), paid twice on a round trip through the topology of §9.

And this is 22.3 §19's min() argument in the latency domain: improving a non-dominant term produces no useful change, so measure the decomposition before choosing what to optimise.

16. Assertions

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// MANDATORY. Illustrative architectural properties (§3) — not normative for
// either technology.
 
// (1) A LOCAL DESTINATION NEVER LEAVES THE PACKAGE.
// English: a message whose resolved scope is package-local is never issued on
// the scale-up interface. Catches §12's first failure at the issue cycle.
a_local_stays_local: assert property (
  @(posedge clk) disable iff (!rst_n)
    (msg_valid && (msg.scope == SCOPE_LOCAL_PACKAGE)) |-> !scaleup_valid
);
 
// (2) SEMANTIC IDENTITY SURVIVES A SCOPE TRANSITION.
// English: crossing a bridge does not change sem_id or generation. Catches a
// bridge that re-tags, which would make de-duplication impossible downstream.
a_identity_survives_bridge: assert property (
  @(posedge clk) disable iff (!rst_n)
    bridge_forward_fire
      |-> (bridge_out_sem_id  == bridge_in_sem_id)
       && (bridge_out_gen     == bridge_in_gen)
);
 
// (3) A RETRY AT EITHER BOUNDARY CANNOT CREATE A SECOND COMPLETION.
// English: however many attempts occur at either transport, one semantic
// operation completes once. This is §13, bounded by retirement rather than by
// an invented window (21.6 §29).
a_one_completion_per_operation: assert property (
  @(posedge clk) disable iff (!rst_n)
    (complete_fire && (complete_id == TEST_ID))
      |=> !(complete_fire && (complete_id == TEST_ID)
            && (complete_gen == $past(complete_gen)))
          throughout (1'b1 [*1:$] ##0 (retire_fire && (retire_id == TEST_ID)))
);
 
// (4) A TRANSPORT EVENT NEVER WRITES THE SEMANTIC IDENTITY.
// English: retries touch only their own attempt field. Structural
// non-interference — best discharged formally (21.7 §17).
a_transport_does_not_touch_identity: assert property (
  @(posedge clk) disable iff (!rst_n)
    (d2d_retx_fire || scaleup_retx_fire)
      |=> ($stable(sem_id_q) && $stable(generation_q))
);
 
// (5) ROUTE SCOPE IS STABLE FOR A LIVE OPERATION.
// English: an in-flight operation's scope cannot change underneath it.
// Catches a route-table update mid-flight (22.3 §14).
a_scope_stable_while_live: assert property (
  @(posedge clk) disable iff (!rst_n)
    (op_live && $changed(route_epoch_q)) |-> in_quiesce
);

Architecture. One routing property, one bridge-integrity property, one exactly-once property, one non-interference property, one lifetime property.

State. The identity registers, attempt counters, and route epoch.

Sampled timing. Property (3) is bounded by an observed retirement, not by a fixed window — 21.6 §29's three forms, where only the retirement-bounded one both terminates and reports a real pass. Property (4) uses |=> so it compares the retry cycle against the next cycle's identity, the correct phase for a registered value.

Contract. Property (3) requires that retirement is guaranteed to occur, or it is vacuously true forever (22.5 §14's pairing requirement). That guarantee must be written alongside it.

Failure if omitted. Without (1), §12's local-message-on-external-path ships and presents as a latency mystery. Without (3) and (4), §13 ships and is blamed on a transport that is working correctly.

DV/debug. Property (4) belongs in a formal flow — non-interference is a proof obligation, and simulation shows only that it held for the stimulus you ran.

17. Common Misconceptions

"UCIe competes directly with NVLink." §5: NVIDIA's own material distinguishes a system-scale link from a chip-to-chip variant, and only the second is at UCIe's layer.

"The link with higher bandwidth is better." §8: aggregate system figures and per-link die-to-die figures measure different objects, and the numbers available here differ by generation and source level.

"NVLink is just a PHY." §4 row 7: NVLink-C2C is described as carrying AMBA CHI or CXL and supporting coherent transfers and atomics — a transport beneath named protocols, not a signalling layer alone.

"UCIe is an external GPU scale-up network." §10: no switching, no peer-absence model, no service story — because a fixed in-package graph never needed them.

"One transport should connect every scope." §12: local traffic takes an enormous detour, and external traffic inherits timeouts sized for a short hop.

"Retry at two levels means two transactions." §13: two correct retries, one semantic operation — and the fix is an identity hierarchy, not a change to either retry.

"Internal package bandwidth determines multi-GPU application performance." §15: for a path crossing both scopes, no single term owns the total, and the scale-up hop and bridges frequently dominate.

"A UCIe-connected GPU die would remove the need for external scale-up." §10: a die-to-die link cannot reach across a chassis, switch, or tolerate a peer failing alone. The scopes are complementary.

18. Understanding Check

19. Module 23 Synthesis

One framework, five technologies, and only defensible wording.

TechnologyPrimary architectural scopePhysical scopeSemantic ownershipStandard or vendorWhere it sitsWhat it is not
PCIecomponent-to-component across a platformboard / systemdefines its own transaction semanticsstandard (PCI-SIG)platform interconnectnot a die-to-die link
UCIedie-to-die inside a packagepackagecarries protocols; defines none of themstandard (consortium)D2D boundarynot a fabric, not a platform link, not a coherence protocol
proprietary D2Ddie-to-die, one co-designed pairpackageoften fused with the interfacevendorD2D boundarynot reusable across an organisation
a system fabricsemantics, routing, topology across many nodesspans the systemowns coherence, ordering, identityvendorabove transportsnot a link
a scale-up interconnectaccelerator-to-accelerator across a systemchassis / rackcarries named protocolsvendorsystem scopenot an in-package D2D link

And the question to carry out of Module 23:

"What architectural boundary am I trying to standardise?"not "which interconnect is fastest?"

Because every chapter reached the same shape: the comparison only became answerable after the layer and scope were pinned, and once they were, the deciding variable was usually organisational (23.2 §8) rather than technical.

20. Summary

Five things.

"NVLink" names two things (§5), and only the chip-to-chip one is at UCIe's layer — a distinction from NVIDIA's own material, not an assertion of mine.

The valid comparison is unusually well-sourced (§7). NVIDIA places UCIe and NVLink-C2C at the same decision point, and the substantive difference is who you must agree with, not a performance figure.

No numeric comparison is possible from these sources (§8): no UCIe figure, a 2022 generation on one side, and a well-qualified NVIDIA number whose baseline is PCIe Gen 6, not UCIe.

The scopes are complementary on one path (§9–§10), and scope is a property of the destination — resolved once, carried, and used to select among transports with different contracts (§11–§12).

And retry at two layers is one operation, not two (§13). Both retries are correct; the gap is identity, and the fix is a hierarchy where no transport event ever writes the semantic identity (§14).