UCIe · Module 22
AI Accelerators on UCIe
An AI accelerator package read as a bandwidth graph with heterogeneous ownership — which edges a standard die-to-die link can actually change, why standardising the wrong boundary produces no speedup and a wrong conclusion, how bulk tensor traffic starves the completion path that retires work, and what the public record on UCIe optical chiplets does and does not establish.
Chapters 22.1 and 22.2 read two vendors' records. This chapter asks a harder question: where in an AI accelerator package would a standard die-to-die link actually change anything?
1. The One-Sentence Model
An AI accelerator package is a bandwidth graph with heterogeneous ownership. Every edge has a required rate, an available rate, a latency, and an owner — and end-to-end throughput is set by the tightest edge on the path the workload actually uses. UCIe changes an edge's ownership and interoperability; it does not change any other edge's capacity.
Which is why "AI needs bandwidth, therefore UCIe" is not an argument. It names a requirement and a technology without establishing that the technology sits on the binding edge — and §9 is the worked case where it does not.
2. What This Chapter Owns
| Question | Where it is answered |
|---|---|
| Evidence levels, the four vendor claims, tense discipline | 22.1 — Intel Chiplets on UCIe |
| The incumbent-fabric problem, layering vs replacement | 22.2 — AMD Chiplets on UCIe |
| The three-layer stack, FDI and RDI | 19.1 — Link Architecture |
| Protocol mapping — native, Streaming, Raw Mode | 19.2 — Protocol Engines |
| Bandwidth, latency and package-level performance | 15.1 · 15.2 · 15.4 |
| Buffering, flow control, backpressure | 13.1 · 13.3 · 19.4 |
| Throughput diagnosis and lost opportunity | 21.5 — Throughput Issues |
| Server-class processors | 22.4 — Data-Centre Processors (next) |
22.1 and 22.2 established the evidence method; this chapter uses it without restating it. What is new is architectural:
The bandwidth graph (§5–§8), which makes "where would this help?" a question with an answer rather than an opinion.
The wrong-boundary failure (§9), where a link is upgraded, nothing improves, and the team concludes the standard is slow.
Traffic-class starvation (§11–§14) — the mechanism by which bulk tensor traffic stops an accelerator from retiring work, with the RTL that causes it and the RTL that fixes it.
And a genuinely different evidence shape (§4). The clearest UCIe evidence in this space is neither a CPU vendor nor a shipping accelerator — it is a component supplier, and reading that correctly is its own exercise.
3. Sourcing and Evidence Date
4. Claim-vs-Evidence — UCIe and AI Accelerators
| Claim | Evidence | Level | Source / date | Proves | Does not prove |
|---|---|---|---|---|---|
| A UCIe optical I/O chiplet has been publicly announced | TeraPHY™, described as "the industry's first Universal Chiplet Interconnect Express (UCIe™) optical interconnect chiplet" | C/D | Ayar Labs, 31 Mar 2025 | a UCIe-facing optical chiplet was announced and showcased | not shipping — the release gives no availability, sampling or GA date |
| Its stated bandwidth is 8 Tbps | "8 Tbps bandwidth" for the chiplet | C | Ayar Labs, 31 Mar 2025 | the supplier's figure for its own part | nothing about any accelerator's D2D bandwidth, and it is not a UCIe figure |
| It was shown publicly | showcased at OFC, 30 Mar – 3 Apr 2025 | D | Ayar Labs, 2025 | a public demonstration occurred | production deployment |
| Multiple ecosystem participants endorsed it | supporting statements from AMD, ASE, Alphawave Semi, d-Matrix, GlobalFoundries, TSMC, Jabil and the UCIe Consortium | E | Ayar Labs, 31 Mar 2025 | breadth of ecosystem interest | that any of them uses it in a product |
| A partner may integrate it | Alphawave Semi referenced potential integration with an I/O chiplet | C (future capability) | Ayar Labs, 31 Mar 2025 | stated intent by a third party | current deployment |
| A shipping AI accelerator uses UCIe for die-to-die | none found (§3) | — | — | — | — |
| No UCIe revision is named for the part | the release names UCIe with no revision | — | Ayar Labs, 31 Mar 2025 | — | which revision it implements — 21.6 §32 makes this matter |
Four readings, and the first reframes the whole space.
The strongest UCIe evidence in AI is a component supplier, not an accelerator vendor. That is not a weakness in the record — it is what an ecosystem looks like early. A supplier builds a UCIe-facing part precisely because it wants to sell into packages it did not design (22.2 §9's decisive pressure).
Row 4 is the row most likely to be misread. Seven named organisations endorsing a component is real ecosystem signal and is not evidence that any of them ships it. A supporting quote in a partner's press release is Level E.
Row 7 matters more than it looks. A part described only as "UCIe" without a revision cannot be assessed for interoperability with a specific peer — which is 22.2 §12's 1.0-versus-1.1 Streaming problem as a purchasing question.
And the missing row is the useful finding. The claim everyone wants — a shipping accelerator with UCIe die-to-die — is the one with no evidence behind it in what I could reach.
5. The Bandwidth Graph
Stop thinking about "the interconnect" and start thinking about edges.
| Edge | Typically carries | Usually limited by |
|---|---|---|
| compute ↔ local memory | tensor operands | memory service rate |
| compute ↔ compute | activations, partial results | link capacity or topology |
| compute ↔ memory-side die | bulk reads and writes | whichever of the two is slower |
| compute ↔ host/control | commands, completions | latency, not bandwidth |
| package ↔ external scale-up | cross-package traffic | the external network |
| any ↔ management | telemetry, configuration | nothing — it is tiny |
Four properties, and the fourth is the chapter's method.
Each edge has four attributes, not one. Required rate, available rate, latency, and owner — and the owner determines whether you can change it at all.
Directionality and burstiness matter as much as rate. An edge sized for average demand stalls on the burst; 21.5 §25's two cases apply unchanged, and a training step's traffic is emphatically not uniform.
The binding edge depends on the workload, not the hardware. A model whose weights fit locally exercises different edges from one that streams them — so "the bottleneck" is not a property of the package.
And UCIe is a candidate for the owner column on some edges and none of the others. It can make an edge interoperable, standard-characterised, and sourceable from someone else. It cannot make memory faster.
6. Conceptual Topologies
Three things to read, and the label matters as much as the picture.
The whole diagram is CONCEPTUAL (§3). It is the architecture an accelerator could have, drawn so the bandwidth-graph argument is concrete — no public source reviewed here documents a product built this way, and presenting a speculative topology as fact is the error 22.1 exists to prevent.
The two red nodes are where the binding constraint usually sits, and neither is a candidate boundary. Local memory service and the external network are the classic limiters — which is §9's entire point.
The three accent edges are where a standard interface changes ownership. Each is a boundary where the die on the far side might plausibly come from somewhere else (§7) — and that, not bandwidth, is the argument for standardising it.
7. Five Boundaries, Five Different Arguments
| Topology | The boundary | Why standardise it? |
|---|---|---|
| A | compute ↔ IO/control die | the IO die may be reused across a product family, or sourced |
| B | compute ↔ memory-side chiplet | only if that die comes from elsewhere — §9 warns about the rest |
| C | compute tile ↔ compute tile | usually not — co-design wins (22.2 §10) |
| D | accelerator ↔ coherent host/control die | the host side may be a different vendor entirely |
| E | accelerator ↔ optical / scale-up bridge | the strongest case — §4's evidence is exactly here |
Two readings.
Row E is where the actual public evidence lives, and the reason is structural: an optical I/O supplier is by definition selling into packages it did not design. 22.2 §9's decisive pressure — a die crossing an organisational boundary — is the normal case for a component vendor, which is why the first UCIe-facing parts appear there rather than inside a vertically integrated accelerator.
And row C is the one to resist. The highest-bandwidth, lowest-latency, most co-designed boundary in the package is exactly where a standard interface has the least to offer and the most constraint to impose.
8. What UCIe Does Not Solve
A standard die-to-die link changes one edge's ownership. It does not touch any of these.
| Not solved | Why |
|---|---|
| memory bandwidth | set by the memory technology and its controller, not the link to it |
| compute throughput | set by the compute dies |
| global coherence | a protocol-layer semantic problem (22.2 §15) |
| the memory device protocol | UCIe is a die-to-die link, not a memory interface |
| package power and thermal | frequently the real ceiling |
| scheduler and software quality | the largest lever in most real systems |
| the external scale-up network | outside the package entirely |
And the fourth row deserves emphasis because the confusion is common. "UCIe replaces HBM" is a category error: HBM is a memory technology with its own interface; UCIe is a link between dies. A memory-side chiplet reached over UCIe still has to talk to memory devices over a memory interface, and that interface's service rate is unchanged by how the chiplet is reached.
9. Wrong Architecture — Standardising the Wrong Boundary
A worked case, illustrative throughout.
| Step | What happened |
|---|---|
| decision | put a standard link between compute and the memory-side chiplet |
| link capacity | comfortably above the previous interface |
| measured application throughput | unchanged |
| conclusion drawn | "UCIe is too slow / adds too much overhead" |
| what the counters said | link utilisation low, memory-side queue permanently full |
| actual limiter | the memory service rate behind the chiplet |
Five readings.
The link was never the constraint, so raising its capacity could not raise throughput. 21.5 §51's consumer-bottleneck trace, arriving as an architecture decision rather than a debug finding.
The conclusion drawn is exactly backwards and very hard to dislodge, because it is supported by a real observation: the link was changed and nothing improved.
The discriminating evidence was available and cheap. Low link utilisation with a full downstream queue is the signature (21.5 §23's waterfall) — and it says the constraint is behind the link, not in it.
The correct reading of the same result is a success, not a failure: the boundary now has standard characterisation and could be sourced elsewhere, at no throughput cost. Whether that is worth it is a commercial question — but it is not "too slow".
And the general rule is 21.5 §42's denominator discipline applied to architecture: before standardising an edge for performance reasons, prove that edge is binding. If it is not, standardise it for ownership reasons or not at all.
10. Traffic Classes
An accelerator's package traffic is not one thing, and treating it as one thing is §12's bug.
| Class | Volume | Latency sensitivity | If starved |
|---|---|---|---|
| control / commands | tiny | high | the engine runs out of work |
| coherent metadata | small | high | ordering stalls |
| bulk reads | enormous | low | throughput drops |
| bulk writes | enormous | low | throughput drops |
| completions / progress | small | critical | work is never retired — §12 |
| telemetry | tiny | none | you lose visibility |
Two properties, and the second is the failure mechanism.
Volume and importance are inversely correlated. The classes that matter most for progress are the smallest; the classes that dominate the wire matter least per byte.
And they compete for the same resources unless something prevents it. 19.5 §8's domain question: if two classes draw from one physical pool, one can consume all of it — and §12 is that, worked through to a deadlock.
11. Illustrative — Traffic Classification
// ILLUSTRATIVE ONLY. Traffic classes, and the two attributes that drive every
// resource decision downstream: whether a class is PROGRESS-CRITICAL and
// whether it is BULK. Those two bits are what §13's reserve keys on.
typedef enum logic [2:0] {
TC_CONTROL = 3'd0, // commands into the engine
TC_COHERENCE = 3'd1, // coherent metadata
TC_BULK_RD = 3'd2, // tensor operand reads
TC_BULK_WR = 3'd3, // result writes
TC_COMPLETION = 3'd4, // retirement — the class that frees resources
TC_TELEMETRY = 3'd5
} accel_tc_e;
// Two derived properties, defined ONCE, so every consumer agrees (19.5 §38's
// single-definition argument). A class that is progress-critical must never be
// blocked behind a bulk class.
function automatic logic tc_is_progress(accel_tc_e tc);
return (tc == TC_CONTROL) || (tc == TC_COHERENCE) || (tc == TC_COMPLETION);
endfunction
function automatic logic tc_is_bulk(accel_tc_e tc);
return (tc == TC_BULK_RD) || (tc == TC_BULK_WR);
endfunction
// Classification happens ONCE, at admission, and travels with the object.
// Re-deriving it downstream is how two blocks come to disagree about what a
// packet is (21.4 §38).
typedef struct packed {
logic [OBJ_W-1:0] obj_id;
accel_tc_e tc;
logic is_progress; // captured, not recomputed
logic is_bulk;
logic [LEN_W-1:0] length;
logic [EPOCH_W-1:0] cfg_epoch; // 21.6 §10 — which configuration
} accel_object_t;
always_comb begin
admit_obj.obj_id = src_obj_id;
admit_obj.tc = src_tc;
admit_obj.is_progress = tc_is_progress(src_tc);
admit_obj.is_bulk = tc_is_bulk(src_tc);
admit_obj.length = src_length;
admit_obj.cfg_epoch = cfg_epoch_q;
endArchitecture. A closed class enumeration plus two derived predicates defined once, so the reserve logic, the queue selector and the arbiter cannot disagree about what a class is.
State. None here — the classification is captured into the object record at admission.
Cycle/event behaviour. Classified once, at admission, and carried. Not recomputed downstream.
Contract. is_progress and is_bulk must be captured, not re-derived. 21.4 §38's single-definition rule: two blocks that each call tc_is_progress are fine today and diverge the moment someone adds a class and updates one call site.
Failure. If classification is recomputed at the arbiter from a field that was rewritten in between, a completion can be arbitrated as bulk — and §12's starvation happens to the one class that must never be starved.
DV/debug. The two bits belong in the trace event (21.7 §16). A trace showing a TC_COMPLETION object with is_progress clear is a classification bug caught directly, rather than inferred from a hang.
12. Wrong RTL — One Queue for Everything
// WRONG — one queue, one credit pool, strict FIFO order.
// Plausible, simple, and it deadlocks under exactly the load it was built for.
logic [QW-1:0] q_mem [QDEPTH];
logic [PW-1:0] q_wr_ptr_q, q_rd_ptr_q;
logic [CW-1:0] credit_q; // ONE pool for all classes
assign q_push = obj_valid && (credit_q != '0);
assign q_pop = q_nonempty && dn_ready; // strict FIFO — no class awarenessThe failure chain, cycle by cycle in ownership terms:
| Step | State |
|---|---|
| 1 | a large tensor read burst is admitted; the queue fills with TC_BULK_RD |
| 2 | credit reaches zero — the shared pool is entirely owned by bulk traffic |
| 3 | a TC_COMPLETION arrives for finished work |
| 4 | the compute engine cannot retire the finished work |
| 5 | with output buffers full, the engine cannot accept new bulk data |
| 6 | so the queued bulk reads cannot be consumed |
| 7 | the queue never drains → credit never returns → step 3 never resolves |
Five properties.
It is a genuine cyclic wait, not a slowdown. 13.3's wait-for cycle: completion waits on credit, credit waits on drain, drain waits on the engine, the engine waits on completion. Nothing times out because nothing is broken.
It only happens under load, which is why it survives directed testing and appears in the first large training run.
The symptom points at the wrong place. The observable is a full queue and zero credit — so the investigation starts at flow control, where the accounting is perfectly correct (21.3 §36's distinction between an accounting failure and genuine congestion).
The head-of-line component makes it worse, not better: even with credit, strict FIFO order puts the completion behind every queued bulk object (21.5 §45) — and raising an arbiter weight cannot fix it, because the completion is not a requester until it reaches the head.
And this is 19.5 §8's physical-pool test failing. Two classes with different progress properties sharing one physical pool is the defect; everything above is a consequence.
13. Corrected RTL — Class Queues With a Progress Reserve
// CORRECTED. Separate queues per class group, and — critically — a RESERVED
// portion of the shared resource that bulk traffic can never consume.
localparam int PROGRESS_RESERVE = 4; // units bulk may never take
logic [CW-1:0] credit_total_q; // total units currently available
logic [CW-1:0] bulk_outstanding_q;
// Bulk may only be admitted while leaving the reserve intact. Progress traffic
// may use everything. This single expression is the whole fix.
logic bulk_may_admit, progress_may_admit;
assign bulk_may_admit = (credit_total_q > CW'(PROGRESS_RESERVE));
assign progress_may_admit = (credit_total_q != '0);
assign admit_ok = admit_obj.is_progress ? progress_may_admit : bulk_may_admit;
// Per-class-group queues, so a blocked bulk head cannot hide an eligible
// completion behind it (21.5 §46).
logic q_push_prog, q_push_bulk;
assign q_push_prog = obj_valid && admit_obj.is_progress && progress_may_admit;
assign q_push_bulk = obj_valid && !admit_obj.is_progress && bulk_may_admit;
// Arbitration: progress traffic wins whenever it has work. Its volume is tiny
// (§10), so this cannot starve bulk in practice — but the assumption must be
// stated, because it is what makes strict priority safe here.
logic serve_progress;
assign serve_progress = q_prog_nonempty && dn_ready;
assign serve_bulk = !q_prog_nonempty && q_bulk_nonempty && dn_ready;
// The credit update, as ONE signed next-state expression so a simultaneous
// admit and return nets correctly rather than losing one (19.5 §14).
logic signed [CW:0] credit_next;
assign credit_next = $signed({1'b0, credit_total_q})
+ $signed({1'b0, ret_units_this_cycle})
- $signed({1'b0, admit_units_this_cycle});
always_ff @(posedge clk or negedge por_n) begin
if (!por_n) credit_total_q <= CW'(CREDIT_INIT);
else credit_total_q <= credit_next[CW-1:0];
endArchitecture. Two queues by progress class, plus a reserve that bulk admission may not encroach on. The reserve is the mechanism; the separate queues prevent head-of-line blocking.
State. Two queues, a total credit counter, and a bulk-outstanding counter.
Cycle/event behaviour. Admission is gated per class; credit updates as one signed next-state expression, so a return and an admit landing together net correctly instead of one being lost to two sequential assignments.
Contract. PROGRESS_RESERVE must be at least the maximum number of progress objects that can be outstanding at once, or the reserve is decorative — a reserve of 4 with 8 possible concurrent completions still deadlocks, more rarely and less reproducibly. Sizing it requires knowing that maximum, which is a design-level obligation, not a tuning knob.
Failure. Two. Separate queues without a reserve fixes head-of-line and still deadlocks at step 2, because bulk can still take every credit. A reserve without separate queues fixes the credit deadlock and leaves the completion stuck behind bulk objects in FIFO order. Both are needed, and each alone looks like a fix in testing.
DV/debug. The invariant is checkable directly (§17): credit available to progress traffic never reaches zero while bulk is outstanding. And credit_min_seen for the progress class (21.7 §21) is the silicon counter that proves the reserve held.
14. Illustrative — Routing Across Dies
// ILLUSTRATIVE ONLY (§11). A route table keyed by destination and class. The
// class is part of the key because different classes may legitimately take
// different paths — and because a route change must not apply to an object
// already in flight (§15).
typedef struct packed {
logic [DIE_W-1:0] dest_die;
logic [PORT_W-1:0] egress_port;
logic [PRIO_W-1:0] priority;
logic valid;
logic [EPOCH_W-1:0] route_epoch; // which routing configuration
} route_entry_t;
route_entry_t route_tbl [NUM_DIES][NUM_TC];
logic [EPOCH_W-1:0] route_epoch_q;
// Lookup carries the epoch OUT with the decision, so a response can later be
// checked against the configuration that was live when the request was routed
// (21.6 §11's stale-event problem).
function automatic route_entry_t lookup(logic [DIE_W-1:0] die, accel_tc_e tc);
route_entry_t e = route_tbl[die][tc];
if (!e.valid) begin
// An invalid route is NOT a silent drop. A dropped object with no record
// is 21.6 §31's causeless event — it looks like the far end invented a
// missing request.
report_route_miss(die, tc);
end
return e;
endfunction
// MANDATORY. English: the routing configuration may not change while any
// object routed under the previous configuration is still in flight. Fires on
// the mid-flight route mutation that makes a response un-attributable.
property p_no_route_change_with_inflight;
@(posedge clk) disable iff (!por_n)
$changed(route_epoch_q) |-> (inflight_count_q == '0);
endproperty
a_no_route_change_inflight: assert property (p_no_route_change_with_inflight);Architecture. A route table indexed by destination and class, with an epoch that ties a routing decision to the configuration that produced it.
State. NUM_DIES × NUM_TC entries plus the current epoch.
Cycle/event behaviour. Looked up at admission; the epoch travels with the object, not with the table.
Contract. A route change requires quiescence of in-flight objects — 21.6 §26's configuration-commit contract. Changing the table under live traffic means a response arrives attributable to a route that no longer exists, and the checker then reports a violation against rules that did not apply when the request was issued.
Failure. A route miss handled as a silent drop is the worst outcome: the object never arrives, no error is recorded, and the far end's model sees a completion with no request — 21.6 §31's causeless event, which reads as the peer inventing traffic.
DV/debug. route_epoch in the trace makes a stale-route failure a two-second check. Without it, the same failure is indistinguishable from a protocol violation by the peer.
15. Illustrative — Capability Descriptor
// ILLUSTRATIVE ONLY (§11). What one chiplet must tell another before traffic
// flows. Generic — NOT a UCIe field layout and NOT a vendor structure (§3).
typedef struct packed {
logic [15:0] vendor_id;
logic [7:0] chiplet_role; // compute / memory-side / io / bridge
logic [7:0] spec_major; // which revision — 21.6 §32
logic [7:0] spec_minor;
logic [7:0] tc_supported; // one bit per traffic class it accepts
logic [7:0] tc_required; // classes it REQUIRES the peer to accept
logic [15:0] max_outstanding; // per class group
logic reliable_delivery; // does it expect the Adapter's retry?
logic coherent_capable;
logic [7:0] feature_bits;
} accel_capability_t;Architecture. Identity, revision, class support, class requirements, and resource limits — the minimum for two independently designed chiplets to establish whether they can work together at all.
State. A constant per chiplet plus a register holding the peer's copy.
Cycle/event behaviour. Exchanged once per link agreement, re-exchanged on renegotiation.
Contract. tc_required is the field that makes a mismatch detectable at bring-up rather than at runtime. Without it, two chiplets negotiate successfully and fail the first time a required class is used — 21.1's hardest shape: a link that came up and should not have.
Failure. §16.
DV/debug. Both capability records belong in the snapshot (21.7 §9). In an interoperability matrix, the failing pairing is usually distinguished by exactly one bit — and without the record that comparison cannot be made.
16. Failure — Two Chiplet Suppliers Disagree
Generic, and deliberately without vendor names.
| Chiplet A | Chiplet B | |
|---|---|---|
| feature F in capability bits | advertised | advertised |
| feature F active after negotiation | assumed yes | no — disabled by configuration |
| behaviour when F-dependent traffic arrives | sends it | rejects or misinterprets |
| each side's self-check | passes | passes |
Four readings.
Both sides are internally consistent, which is 21.4 §17's layer-local correctness across an organisational boundary — and neither supplier's verification environment can find it, because each tested against its own assumption.
Advertising a capability is not the same as it being active. 21.1 §29's requested-versus-active distinction, at the feature level: tc_supported says "I can"; only the negotiated agreement says "we do."
The fix is an explicit active feature set, returned by the negotiation and readable by both sides — not each side inferring it from the other's advertisement.
And the blame question is the wrong question. 21.7 §31's peer matrix: this eliminates "A is broken" and "B is broken" and leaves an interaction, whose resolution is a specification-applicability argument (21.6 §4) rather than a defect in either part.
17. Assertions
// MANDATORY. Illustrative architectural properties (§3) — not UCIe requirements.
// (1) THE PROGRESS RESERVE HOLDS. English: bulk admission never consumes the
// units reserved for progress traffic. Sampled every cycle. Fires at the
// admission that would deadlock §12, before anything is stuck.
a_progress_reserve_held: assert property (
@(posedge clk) disable iff (!por_n)
(admit_fire && !admit_obj.is_progress)
|-> (credit_total_q > CW'(PROGRESS_RESERVE))
);
// (2) AN ACCEPTED COMMAND OWNS A LIVE TRANSACTION. English: every admitted
// object has a live tracking entry until it completes. Fires when admission
// and tracking disagree — the bug that makes outstanding counts meaningless.
a_accepted_is_tracked: assert property (
@(posedge clk) disable iff (!por_n)
admit_fire |=> live_entry[$past(admit_obj.obj_id)]
);
// (3) A RESPONSE BELONGS TO THE CURRENT CONFIG EPOCH. English: a response
// carrying a stale epoch must be rejected explicitly, not applied. This is
// 21.6 §11 as a property — without it a straggler is silently misinterpreted.
a_response_epoch_current: assert property (
@(posedge clk) disable iff (!por_n)
(resp_fire && (resp_epoch != cfg_epoch_q)) |-> resp_rejected
);
// (4) A DUPLICATE PHYSICAL ATTEMPT DOES NOT DUPLICATE SEMANTIC COMPLETION.
// English: however many times an object is transmitted, it completes once.
// This is the property that separates a working retry mechanism from a
// reliability violation (21.6 §19).
a_completion_once: assert property (
@(posedge clk) disable iff (!por_n)
(complete_fire && (complete_id == TEST_ID))
|=> !(complete_fire && (complete_id == TEST_ID))
throughout (1'b1 [*1:$] ##0 (retire_fire && (retire_id == TEST_ID)))
);Architecture. Four properties, each attached to a specific failure above rather than restating an assignment.
State. The tracking table and the epoch register.
Sampled timing. Property (2) uses |=> with $past, so it checks this cycle's admission against next cycle's tracking state — the correct phase for a registered update. Property (4) is bounded by retirement, an observed event, rather than by an invented window (21.6 §29's three forms).
Contract. Property (1) is the reserve's specification. It must be written against the admission event, not against a steady-state occupancy — checking occupancy after the fact reports the deadlock rather than preventing it.
Failure if omitted. Without (1), §12's deadlock is found in a training run. Without (3), stale responses are applied and corrupt live state. Without (4), a retry mechanism doing its job is indistinguishable from a reliability violation — and the "fix" weakens the retry.
DV/debug. These four map onto §19's verification questions, and three of them survive into silicon as counters (21.7 §21) rather than assertions.
18. Reliability Overhead Is Not Payload
An accelerator's headline metric is bytes moved. A retry moves bytes and delivers nothing new.
| Counter | Answers |
|---|---|
bytes_wire | how busy the link was |
bytes_unique | how much new data actually arrived |
bytes_unique / bytes_wire | the retry overhead |
Two readings.
A metric built on wire bytes improves as the link degrades (21.5 §16) — the worst possible property for a performance indicator, and especially damaging here because tensor traffic is so dominant that a few per cent of retry is a large absolute number.
And the correct response to high retry overhead is not in the transfer path. 21.5 §38: the scheduler and the arbiter are doing their jobs; the investigation belongs to signal integrity (21.7 §32's physical-suspicion branch).
19. Worked Performance Model
ILLUSTRATIVE. All values are internal to this example and are NOT UCIe,
vendor or product figures (§3). Units are symbolic "bytes per unit time."
Compute demand on this edge X = 100
Link available rate Y = 140
Memory service rate behind it Z = 60
Naive upper bound min(X, Y, Z) = 60
Now the efficiency terms, applied to the LINK only:
unique / wire fraction u = 0.90 (10% retry — §18)
payload / wire fraction p = 0.85 (framing overhead)
effective link rate Y' = 140 x 0.90 x 0.85 = 107.1
Revised bound min(100, 107.1, 60) = 60
CONCLUSION: the link's efficiency terms do not matter here at all.
Z binds by a wide margin, and it still binds after the link is derated.Four readings.
The efficiency terms changed nothing, because the link was not the binding term and remains non-binding after derating. That is the arithmetic behind §9's wrong architecture.
The check worth running is how much slack the non-binding terms have. Y' = 107.1 against Z = 60 is a wide margin; at Y' = 62 the answer would be fragile, and a small increase in retry would make the link binding.
Applying the efficiency terms to the wrong quantity is a common error. They derate the link, not the memory service — multiplying Z by u would understate memory by 10% for no reason.
And min() is an upper bound, not a prediction. Burstiness, latency-hiding, queue depth and scheduling all sit below it (21.5 §26) — a real system reaches min() only if every edge is kept busy, which is a separate problem.
20. Debugging a Package Bottleneck
The counters, and the order to read them (21.5 §56's tree, scoped to a package).
| Read | Says |
|---|---|
| per-link utilisation | low + downstream full → the constraint is behind the link (§9) |
| queue occupancy per stage | which boundary is limiting (21.5 §23) |
| memory service rate | usually the answer |
unique / wire | retry overhead (§18) |
| completion latency | whether progress traffic is being starved (§12) |
| per-class transfer share | whether one class dominates |
And the single most diagnostic pairing is the first row. Low link utilisation with a full downstream queue eliminates the link as a candidate in one read — which is the observation §9's team did not make.
21. Verification Questions
| Question | Why |
|---|---|
| does progress traffic advance under maximum bulk load? | §12's deadlock |
| are the classes genuinely independent, or sharing a pool? | 19.5 §8 |
| what happens to outstanding commands across a recovery? | 14.2 |
| is semantic completion exactly once under retry? | §17 property 4 |
| is a stale route or config epoch rejected? | §14, §17 property 3 |
| is host/accelerator ordering preserved where required? | 21.6 §14 |
| is the progress reserve sized against the true maximum? | §13's contract |
And the first is the one to write first, because it is the test §12's design passes in directed runs and fails in the first real workload.
22. What Public Material Cannot Tell You
Unless a source explicitly discloses it, none of this is knowable — and none of it appears above.
| Not knowable | Module framing the question |
|---|---|
| any accelerator's RTL | — |
| queue depths, credit counts, reserve sizing | 13.2 · 19.4 |
| exact die topology and which dies exist | §6 is conceptual |
| arbitration policy between classes | 19.3 |
| retry buffer sizing and replay policy | 14.2 |
| internal protocol carried over any link | 19.2 |
| per-link latency or bandwidth | 15.1 · 15.2 |
| exact UCIe widths, rates or modes used | §4 — not even the revision is public for the one announced part |
| internal FDI/RDI structure | 19.1 |
And row 8 is worth dwelling on. The clearest UCIe part in this space is described publicly as "UCIe" with no revision named (§4) — so even for the component with the best documentation, the interoperability question 21.6 §32 demands cannot be answered from public material.
23. Common Misconceptions
"AI accelerators need UCIe because AI needs bandwidth." §5, §9: bandwidth is a requirement, not an argument. UCIe changes an edge's ownership; it changes no edge's capacity.
"A chiplet accelerator automatically uses UCIe." 22.1 §12: multi-die predates the standard, and co-designed interfaces have real advantages at internal boundaries.
"UCIe replaces HBM." §8: a category error. HBM is a memory technology with its own interface; UCIe is a die-to-die link. A memory-side chiplet still talks to memory over a memory interface.
"UCIe solves scale-up networking." §8: the external network is outside the package. A UCIe-facing bridge chiplet is a boundary to that network, not a replacement for it.
"A demonstration proves a shipping accelerator." §4: an announced component with no availability date, endorsed by seven organisations, is Level C/D.
"A standard PHY means standard memory semantics." §8, §16: the link is standard; what crosses it and what the far side does with it are separate agreements.
"More links always increase throughput." §19: adding capacity to a non-binding edge changes min() by nothing.
"One shared queue maximises utilisation." §12: it maximises utilisation right up to the deadlock, and the deadlock is structural rather than a slowdown.
"Physical retry traffic is useful AI payload." §18: a metric built on wire bytes improves as the link degrades.
24. Understanding Check
25. Summary
Five things.
Model the package as a bandwidth graph (§5). Four attributes per edge — required rate, available rate, latency, owner — and the binding edge depends on the workload, not the hardware.
A standard link changes ownership, not capacity (§8). It cannot make memory faster, compute wider, coherence simpler, or the external network larger.
So prove the edge binds before standardising it for performance (§9, §19). Standardising a non-binding edge is fine for ownership reasons and produces no speedup — and the conclusion "the standard is slow" is the predictable, wrong outcome.
Progress traffic and bulk traffic must not share a pool (§12–§13). Separate queues fix head-of-line; a reserve fixes credit exhaustion; both are required, and the deadlock is structural rather than a slowdown.
And the public record here is a component supplier, not an accelerator (§4). One dated, primary-sourced announcement with no availability date, seven Level-E endorsements, and no revision named — which is real ecosystem evidence and is not a shipping claim.