Skip to content

UCIe · Module 22

AI Accelerators on UCIe

An AI accelerator package read as a bandwidth graph with heterogeneous ownership — which edges a standard die-to-die link can actually change, why standardising the wrong boundary produces no speedup and a wrong conclusion, how bulk tensor traffic starves the completion path that retires work, and what the public record on UCIe optical chiplets does and does not establish.

Chapters 22.1 and 22.2 read two vendors' records. This chapter asks a harder question: where in an AI accelerator package would a standard die-to-die link actually change anything?

1. The One-Sentence Model

An AI accelerator package is a bandwidth graph with heterogeneous ownership. Every edge has a required rate, an available rate, a latency, and an owner — and end-to-end throughput is set by the tightest edge on the path the workload actually uses. UCIe changes an edge's ownership and interoperability; it does not change any other edge's capacity.

Which is why "AI needs bandwidth, therefore UCIe" is not an argument. It names a requirement and a technology without establishing that the technology sits on the binding edge — and §9 is the worked case where it does not.

2. What This Chapter Owns

QuestionWhere it is answered
Evidence levels, the four vendor claims, tense discipline22.1 — Intel Chiplets on UCIe
The incumbent-fabric problem, layering vs replacement22.2 — AMD Chiplets on UCIe
The three-layer stack, FDI and RDI19.1 — Link Architecture
Protocol mapping — native, Streaming, Raw Mode19.2 — Protocol Engines
Bandwidth, latency and package-level performance15.1 · 15.2 · 15.4
Buffering, flow control, backpressure13.1 · 13.3 · 19.4
Throughput diagnosis and lost opportunity21.5 — Throughput Issues
Server-class processors22.4 — Data-Centre Processors (next)

22.1 and 22.2 established the evidence method; this chapter uses it without restating it. What is new is architectural:

The bandwidth graph (§5–§8), which makes "where would this help?" a question with an answer rather than an opinion.

The wrong-boundary failure (§9), where a link is upgraded, nothing improves, and the team concludes the standard is slow.

Traffic-class starvation (§11–§14) — the mechanism by which bulk tensor traffic stops an accelerator from retiring work, with the RTL that causes it and the RTL that fixes it.

And a genuinely different evidence shape (§4). The clearest UCIe evidence in this space is neither a CPU vendor nor a shipping accelerator — it is a component supplier, and reading that correctly is its own exercise.

3. Sourcing and Evidence Date

4. Claim-vs-Evidence — UCIe and AI Accelerators

ClaimEvidenceLevelSource / dateProvesDoes not prove
A UCIe optical I/O chiplet has been publicly announcedTeraPHY™, described as "the industry's first Universal Chiplet Interconnect Express (UCIe™) optical interconnect chiplet"C/DAyar Labs, 31 Mar 2025a UCIe-facing optical chiplet was announced and showcasednot shipping — the release gives no availability, sampling or GA date
Its stated bandwidth is 8 Tbps"8 Tbps bandwidth" for the chipletCAyar Labs, 31 Mar 2025the supplier's figure for its own partnothing about any accelerator's D2D bandwidth, and it is not a UCIe figure
It was shown publiclyshowcased at OFC, 30 Mar – 3 Apr 2025DAyar Labs, 2025a public demonstration occurredproduction deployment
Multiple ecosystem participants endorsed itsupporting statements from AMD, ASE, Alphawave Semi, d-Matrix, GlobalFoundries, TSMC, Jabil and the UCIe ConsortiumEAyar Labs, 31 Mar 2025breadth of ecosystem interestthat any of them uses it in a product
A partner may integrate itAlphawave Semi referenced potential integration with an I/O chipletC (future capability)Ayar Labs, 31 Mar 2025stated intent by a third partycurrent deployment
A shipping AI accelerator uses UCIe for die-to-dienone found (§3)
No UCIe revision is named for the partthe release names UCIe with no revisionAyar Labs, 31 Mar 2025which revision it implements21.6 §32 makes this matter

Four readings, and the first reframes the whole space.

The strongest UCIe evidence in AI is a component supplier, not an accelerator vendor. That is not a weakness in the record — it is what an ecosystem looks like early. A supplier builds a UCIe-facing part precisely because it wants to sell into packages it did not design (22.2 §9's decisive pressure).

Row 4 is the row most likely to be misread. Seven named organisations endorsing a component is real ecosystem signal and is not evidence that any of them ships it. A supporting quote in a partner's press release is Level E.

Row 7 matters more than it looks. A part described only as "UCIe" without a revision cannot be assessed for interoperability with a specific peer — which is 22.2 §12's 1.0-versus-1.1 Streaming problem as a purchasing question.

And the missing row is the useful finding. The claim everyone wants — a shipping accelerator with UCIe die-to-die — is the one with no evidence behind it in what I could reach.

5. The Bandwidth Graph

Stop thinking about "the interconnect" and start thinking about edges.

EdgeTypically carriesUsually limited by
compute ↔ local memorytensor operandsmemory service rate
compute ↔ computeactivations, partial resultslink capacity or topology
compute ↔ memory-side diebulk reads and writeswhichever of the two is slower
compute ↔ host/controlcommands, completionslatency, not bandwidth
package ↔ external scale-upcross-package trafficthe external network
any ↔ managementtelemetry, configurationnothing — it is tiny

Four properties, and the fourth is the chapter's method.

Each edge has four attributes, not one. Required rate, available rate, latency, and owner — and the owner determines whether you can change it at all.

Directionality and burstiness matter as much as rate. An edge sized for average demand stalls on the burst; 21.5 §25's two cases apply unchanged, and a training step's traffic is emphatically not uniform.

The binding edge depends on the workload, not the hardware. A model whose weights fit locally exercises different edges from one that streams them — so "the bottleneck" is not a property of the package.

And UCIe is a candidate for the owner column on some edges and none of the others. It can make an edge interoperable, standard-characterised, and sourceable from someone else. It cannot make memory faster.

6. Conceptual Topologies

A conceptual block diagram of an AI accelerator package drawn as a bandwidth graph. Compute dies sit at the centre, connected to local memory stacks on one side and a memory-side chiplet on the other. An input output and control die connects the compute dies to a host interface. An optical or scale-up bridge chiplet connects the package to an external network. A management block connects to the compute dies. Three edges are marked as candidate standard-interface boundaries: compute to input output die, compute to memory-side chiplet, and compute to the optical bridge. The compute to local memory edge and the bridge to external network edge are marked as the edges where the binding constraint usually sits. The entire diagram is labelled conceptual.Local memoryCONSTRAINT (§9)Compute diesconceptualMemory-side diecandidate boundaryIO / control diecandidate boundaryManagementtiny, never the limitHost interfacelatency, not rateOptical bridgecandidate boundaryExternal scale-upCONSTRAINT (§8)12
A conceptual AI accelerator package drawn as a bandwidth graph. Every block and edge here is illustrative — this is not any vendor's product topology, and no public source reviewed for this chapter documents an accelerator built this way. The three edges marked as candidates are the boundaries where a standard die-to-die link changes ownership; the memory and external edges are where the binding constraint usually sits.

Three things to read, and the label matters as much as the picture.

The whole diagram is CONCEPTUAL (§3). It is the architecture an accelerator could have, drawn so the bandwidth-graph argument is concrete — no public source reviewed here documents a product built this way, and presenting a speculative topology as fact is the error 22.1 exists to prevent.

The two red nodes are where the binding constraint usually sits, and neither is a candidate boundary. Local memory service and the external network are the classic limiters — which is §9's entire point.

The three accent edges are where a standard interface changes ownership. Each is a boundary where the die on the far side might plausibly come from somewhere else (§7) — and that, not bandwidth, is the argument for standardising it.

7. Five Boundaries, Five Different Arguments

TopologyThe boundaryWhy standardise it?
Acompute ↔ IO/control diethe IO die may be reused across a product family, or sourced
Bcompute ↔ memory-side chipletonly if that die comes from elsewhere — §9 warns about the rest
Ccompute tile ↔ compute tileusually not — co-design wins (22.2 §10)
Daccelerator ↔ coherent host/control diethe host side may be a different vendor entirely
Eaccelerator ↔ optical / scale-up bridgethe strongest case — §4's evidence is exactly here

Two readings.

Row E is where the actual public evidence lives, and the reason is structural: an optical I/O supplier is by definition selling into packages it did not design. 22.2 §9's decisive pressure — a die crossing an organisational boundary — is the normal case for a component vendor, which is why the first UCIe-facing parts appear there rather than inside a vertically integrated accelerator.

And row C is the one to resist. The highest-bandwidth, lowest-latency, most co-designed boundary in the package is exactly where a standard interface has the least to offer and the most constraint to impose.

8. What UCIe Does Not Solve

A standard die-to-die link changes one edge's ownership. It does not touch any of these.

Not solvedWhy
memory bandwidthset by the memory technology and its controller, not the link to it
compute throughputset by the compute dies
global coherencea protocol-layer semantic problem (22.2 §15)
the memory device protocolUCIe is a die-to-die link, not a memory interface
package power and thermalfrequently the real ceiling
scheduler and software qualitythe largest lever in most real systems
the external scale-up networkoutside the package entirely

And the fourth row deserves emphasis because the confusion is common. "UCIe replaces HBM" is a category error: HBM is a memory technology with its own interface; UCIe is a link between dies. A memory-side chiplet reached over UCIe still has to talk to memory devices over a memory interface, and that interface's service rate is unchanged by how the chiplet is reached.

9. Wrong Architecture — Standardising the Wrong Boundary

A worked case, illustrative throughout.

StepWhat happened
decisionput a standard link between compute and the memory-side chiplet
link capacitycomfortably above the previous interface
measured application throughputunchanged
conclusion drawn"UCIe is too slow / adds too much overhead"
what the counters saidlink utilisation low, memory-side queue permanently full
actual limiterthe memory service rate behind the chiplet

Five readings.

The link was never the constraint, so raising its capacity could not raise throughput. 21.5 §51's consumer-bottleneck trace, arriving as an architecture decision rather than a debug finding.

The conclusion drawn is exactly backwards and very hard to dislodge, because it is supported by a real observation: the link was changed and nothing improved.

The discriminating evidence was available and cheap. Low link utilisation with a full downstream queue is the signature (21.5 §23's waterfall) — and it says the constraint is behind the link, not in it.

The correct reading of the same result is a success, not a failure: the boundary now has standard characterisation and could be sourced elsewhere, at no throughput cost. Whether that is worth it is a commercial question — but it is not "too slow".

And the general rule is 21.5 §42's denominator discipline applied to architecture: before standardising an edge for performance reasons, prove that edge is binding. If it is not, standardise it for ownership reasons or not at all.

10. Traffic Classes

An accelerator's package traffic is not one thing, and treating it as one thing is §12's bug.

ClassVolumeLatency sensitivityIf starved
control / commandstinyhighthe engine runs out of work
coherent metadatasmallhighordering stalls
bulk readsenormouslowthroughput drops
bulk writesenormouslowthroughput drops
completions / progresssmallcriticalwork is never retired — §12
telemetrytinynoneyou lose visibility

Two properties, and the second is the failure mechanism.

Volume and importance are inversely correlated. The classes that matter most for progress are the smallest; the classes that dominate the wire matter least per byte.

And they compete for the same resources unless something prevents it. 19.5 §8's domain question: if two classes draw from one physical pool, one can consume all of it — and §12 is that, worked through to a deadlock.

11. Illustrative — Traffic Classification

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// ILLUSTRATIVE ONLY. Traffic classes, and the two attributes that drive every
// resource decision downstream: whether a class is PROGRESS-CRITICAL and
// whether it is BULK. Those two bits are what §13's reserve keys on.
typedef enum logic [2:0] {
  TC_CONTROL    = 3'd0,   // commands into the engine
  TC_COHERENCE  = 3'd1,   // coherent metadata
  TC_BULK_RD    = 3'd2,   // tensor operand reads
  TC_BULK_WR    = 3'd3,   // result writes
  TC_COMPLETION = 3'd4,   // retirement — the class that frees resources
  TC_TELEMETRY  = 3'd5
} accel_tc_e;
 
// Two derived properties, defined ONCE, so every consumer agrees (19.5 §38's
// single-definition argument). A class that is progress-critical must never be
// blocked behind a bulk class.
function automatic logic tc_is_progress(accel_tc_e tc);
  return (tc == TC_CONTROL) || (tc == TC_COHERENCE) || (tc == TC_COMPLETION);
endfunction
 
function automatic logic tc_is_bulk(accel_tc_e tc);
  return (tc == TC_BULK_RD) || (tc == TC_BULK_WR);
endfunction
 
// Classification happens ONCE, at admission, and travels with the object.
// Re-deriving it downstream is how two blocks come to disagree about what a
// packet is (21.4 §38).
typedef struct packed {
  logic [OBJ_W-1:0]   obj_id;
  accel_tc_e          tc;
  logic               is_progress;   // captured, not recomputed
  logic               is_bulk;
  logic [LEN_W-1:0]   length;
  logic [EPOCH_W-1:0] cfg_epoch;     // 21.6 §10 — which configuration
} accel_object_t;
 
always_comb begin
  admit_obj.obj_id      = src_obj_id;
  admit_obj.tc          = src_tc;
  admit_obj.is_progress = tc_is_progress(src_tc);
  admit_obj.is_bulk     = tc_is_bulk(src_tc);
  admit_obj.length      = src_length;
  admit_obj.cfg_epoch   = cfg_epoch_q;
end

Architecture. A closed class enumeration plus two derived predicates defined once, so the reserve logic, the queue selector and the arbiter cannot disagree about what a class is.

State. None here — the classification is captured into the object record at admission.

Cycle/event behaviour. Classified once, at admission, and carried. Not recomputed downstream.

Contract. is_progress and is_bulk must be captured, not re-derived. 21.4 §38's single-definition rule: two blocks that each call tc_is_progress are fine today and diverge the moment someone adds a class and updates one call site.

Failure. If classification is recomputed at the arbiter from a field that was rewritten in between, a completion can be arbitrated as bulk — and §12's starvation happens to the one class that must never be starved.

DV/debug. The two bits belong in the trace event (21.7 §16). A trace showing a TC_COMPLETION object with is_progress clear is a classification bug caught directly, rather than inferred from a hang.

12. Wrong RTL — One Queue for Everything

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// WRONG — one queue, one credit pool, strict FIFO order.
// Plausible, simple, and it deadlocks under exactly the load it was built for.
logic [QW-1:0] q_mem [QDEPTH];
logic [PW-1:0] q_wr_ptr_q, q_rd_ptr_q;
logic [CW-1:0] credit_q;                 // ONE pool for all classes
 
assign q_push = obj_valid && (credit_q != '0);
assign q_pop  = q_nonempty && dn_ready;  // strict FIFO — no class awareness

The failure chain, cycle by cycle in ownership terms:

StepState
1a large tensor read burst is admitted; the queue fills with TC_BULK_RD
2credit reaches zero — the shared pool is entirely owned by bulk traffic
3a TC_COMPLETION arrives for finished work
4the compute engine cannot retire the finished work
5with output buffers full, the engine cannot accept new bulk data
6so the queued bulk reads cannot be consumed
7the queue never drains → credit never returnsstep 3 never resolves

Five properties.

It is a genuine cyclic wait, not a slowdown. 13.3's wait-for cycle: completion waits on credit, credit waits on drain, drain waits on the engine, the engine waits on completion. Nothing times out because nothing is broken.

It only happens under load, which is why it survives directed testing and appears in the first large training run.

The symptom points at the wrong place. The observable is a full queue and zero credit — so the investigation starts at flow control, where the accounting is perfectly correct (21.3 §36's distinction between an accounting failure and genuine congestion).

The head-of-line component makes it worse, not better: even with credit, strict FIFO order puts the completion behind every queued bulk object (21.5 §45) — and raising an arbiter weight cannot fix it, because the completion is not a requester until it reaches the head.

And this is 19.5 §8's physical-pool test failing. Two classes with different progress properties sharing one physical pool is the defect; everything above is a consequence.

13. Corrected RTL — Class Queues With a Progress Reserve

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// CORRECTED. Separate queues per class group, and — critically — a RESERVED
// portion of the shared resource that bulk traffic can never consume.
localparam int PROGRESS_RESERVE = 4;     // units bulk may never take
 
logic [CW-1:0] credit_total_q;           // total units currently available
logic [CW-1:0] bulk_outstanding_q;
 
// Bulk may only be admitted while leaving the reserve intact. Progress traffic
// may use everything. This single expression is the whole fix.
logic bulk_may_admit, progress_may_admit;
assign bulk_may_admit     = (credit_total_q > CW'(PROGRESS_RESERVE));
assign progress_may_admit = (credit_total_q != '0);
 
assign admit_ok = admit_obj.is_progress ? progress_may_admit : bulk_may_admit;
 
// Per-class-group queues, so a blocked bulk head cannot hide an eligible
// completion behind it (21.5 §46).
logic q_push_prog, q_push_bulk;
assign q_push_prog = obj_valid &&  admit_obj.is_progress && progress_may_admit;
assign q_push_bulk = obj_valid && !admit_obj.is_progress && bulk_may_admit;
 
// Arbitration: progress traffic wins whenever it has work. Its volume is tiny
// (§10), so this cannot starve bulk in practice — but the assumption must be
// stated, because it is what makes strict priority safe here.
logic serve_progress;
assign serve_progress = q_prog_nonempty && dn_ready;
assign serve_bulk     = !q_prog_nonempty && q_bulk_nonempty && dn_ready;
 
// The credit update, as ONE signed next-state expression so a simultaneous
// admit and return nets correctly rather than losing one (19.5 §14).
logic signed [CW:0] credit_next;
assign credit_next = $signed({1'b0, credit_total_q})
                   + $signed({1'b0, ret_units_this_cycle})
                   - $signed({1'b0, admit_units_this_cycle});
 
always_ff @(posedge clk or negedge por_n) begin
  if (!por_n) credit_total_q <= CW'(CREDIT_INIT);
  else        credit_total_q <= credit_next[CW-1:0];
end

Architecture. Two queues by progress class, plus a reserve that bulk admission may not encroach on. The reserve is the mechanism; the separate queues prevent head-of-line blocking.

State. Two queues, a total credit counter, and a bulk-outstanding counter.

Cycle/event behaviour. Admission is gated per class; credit updates as one signed next-state expression, so a return and an admit landing together net correctly instead of one being lost to two sequential assignments.

Contract. PROGRESS_RESERVE must be at least the maximum number of progress objects that can be outstanding at once, or the reserve is decorative — a reserve of 4 with 8 possible concurrent completions still deadlocks, more rarely and less reproducibly. Sizing it requires knowing that maximum, which is a design-level obligation, not a tuning knob.

Failure. Two. Separate queues without a reserve fixes head-of-line and still deadlocks at step 2, because bulk can still take every credit. A reserve without separate queues fixes the credit deadlock and leaves the completion stuck behind bulk objects in FIFO order. Both are needed, and each alone looks like a fix in testing.

DV/debug. The invariant is checkable directly (§17): credit available to progress traffic never reaches zero while bulk is outstanding. And credit_min_seen for the progress class (21.7 §21) is the silicon counter that proves the reserve held.

14. Illustrative — Routing Across Dies

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// ILLUSTRATIVE ONLY (§11). A route table keyed by destination and class. The
// class is part of the key because different classes may legitimately take
// different paths — and because a route change must not apply to an object
// already in flight (§15).
typedef struct packed {
  logic [DIE_W-1:0]    dest_die;
  logic [PORT_W-1:0]   egress_port;
  logic [PRIO_W-1:0]   priority;
  logic                valid;
  logic [EPOCH_W-1:0]  route_epoch;   // which routing configuration
} route_entry_t;
 
route_entry_t route_tbl [NUM_DIES][NUM_TC];
logic [EPOCH_W-1:0] route_epoch_q;
 
// Lookup carries the epoch OUT with the decision, so a response can later be
// checked against the configuration that was live when the request was routed
// (21.6 §11's stale-event problem).
function automatic route_entry_t lookup(logic [DIE_W-1:0] die, accel_tc_e tc);
  route_entry_t e = route_tbl[die][tc];
  if (!e.valid) begin
    // An invalid route is NOT a silent drop. A dropped object with no record
    // is 21.6 §31's causeless event — it looks like the far end invented a
    // missing request.
    report_route_miss(die, tc);
  end
  return e;
endfunction
 
// MANDATORY. English: the routing configuration may not change while any
// object routed under the previous configuration is still in flight. Fires on
// the mid-flight route mutation that makes a response un-attributable.
property p_no_route_change_with_inflight;
  @(posedge clk) disable iff (!por_n)
    $changed(route_epoch_q) |-> (inflight_count_q == '0);
endproperty
a_no_route_change_inflight: assert property (p_no_route_change_with_inflight);

Architecture. A route table indexed by destination and class, with an epoch that ties a routing decision to the configuration that produced it.

State. NUM_DIES × NUM_TC entries plus the current epoch.

Cycle/event behaviour. Looked up at admission; the epoch travels with the object, not with the table.

Contract. A route change requires quiescence of in-flight objects — 21.6 §26's configuration-commit contract. Changing the table under live traffic means a response arrives attributable to a route that no longer exists, and the checker then reports a violation against rules that did not apply when the request was issued.

Failure. A route miss handled as a silent drop is the worst outcome: the object never arrives, no error is recorded, and the far end's model sees a completion with no request21.6 §31's causeless event, which reads as the peer inventing traffic.

DV/debug. route_epoch in the trace makes a stale-route failure a two-second check. Without it, the same failure is indistinguishable from a protocol violation by the peer.

15. Illustrative — Capability Descriptor

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// ILLUSTRATIVE ONLY (§11). What one chiplet must tell another before traffic
// flows. Generic — NOT a UCIe field layout and NOT a vendor structure (§3).
typedef struct packed {
  logic [15:0] vendor_id;
  logic [7:0]  chiplet_role;        // compute / memory-side / io / bridge
  logic [7:0]  spec_major;          // which revision — 21.6 §32
  logic [7:0]  spec_minor;
  logic [7:0]  tc_supported;        // one bit per traffic class it accepts
  logic [7:0]  tc_required;         // classes it REQUIRES the peer to accept
  logic [15:0] max_outstanding;     // per class group
  logic        reliable_delivery;   // does it expect the Adapter's retry?
  logic        coherent_capable;
  logic [7:0]  feature_bits;
} accel_capability_t;

Architecture. Identity, revision, class support, class requirements, and resource limits — the minimum for two independently designed chiplets to establish whether they can work together at all.

State. A constant per chiplet plus a register holding the peer's copy.

Cycle/event behaviour. Exchanged once per link agreement, re-exchanged on renegotiation.

Contract. tc_required is the field that makes a mismatch detectable at bring-up rather than at runtime. Without it, two chiplets negotiate successfully and fail the first time a required class is used — 21.1's hardest shape: a link that came up and should not have.

Failure. §16.

DV/debug. Both capability records belong in the snapshot (21.7 §9). In an interoperability matrix, the failing pairing is usually distinguished by exactly one bit — and without the record that comparison cannot be made.

16. Failure — Two Chiplet Suppliers Disagree

Generic, and deliberately without vendor names.

Chiplet AChiplet B
feature F in capability bitsadvertisedadvertised
feature F active after negotiationassumed yesno — disabled by configuration
behaviour when F-dependent traffic arrivessends itrejects or misinterprets
each side's self-checkpassespasses

Four readings.

Both sides are internally consistent, which is 21.4 §17's layer-local correctness across an organisational boundary — and neither supplier's verification environment can find it, because each tested against its own assumption.

Advertising a capability is not the same as it being active. 21.1 §29's requested-versus-active distinction, at the feature level: tc_supported says "I can"; only the negotiated agreement says "we do."

The fix is an explicit active feature set, returned by the negotiation and readable by both sides — not each side inferring it from the other's advertisement.

And the blame question is the wrong question. 21.7 §31's peer matrix: this eliminates "A is broken" and "B is broken" and leaves an interaction, whose resolution is a specification-applicability argument (21.6 §4) rather than a defect in either part.

17. Assertions

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// MANDATORY. Illustrative architectural properties (§3) — not UCIe requirements.
 
// (1) THE PROGRESS RESERVE HOLDS. English: bulk admission never consumes the
// units reserved for progress traffic. Sampled every cycle. Fires at the
// admission that would deadlock §12, before anything is stuck.
a_progress_reserve_held: assert property (
  @(posedge clk) disable iff (!por_n)
    (admit_fire && !admit_obj.is_progress)
      |-> (credit_total_q > CW'(PROGRESS_RESERVE))
);
 
// (2) AN ACCEPTED COMMAND OWNS A LIVE TRANSACTION. English: every admitted
// object has a live tracking entry until it completes. Fires when admission
// and tracking disagree — the bug that makes outstanding counts meaningless.
a_accepted_is_tracked: assert property (
  @(posedge clk) disable iff (!por_n)
    admit_fire |=> live_entry[$past(admit_obj.obj_id)]
);
 
// (3) A RESPONSE BELONGS TO THE CURRENT CONFIG EPOCH. English: a response
// carrying a stale epoch must be rejected explicitly, not applied. This is
// 21.6 §11 as a property — without it a straggler is silently misinterpreted.
a_response_epoch_current: assert property (
  @(posedge clk) disable iff (!por_n)
    (resp_fire && (resp_epoch != cfg_epoch_q)) |-> resp_rejected
);
 
// (4) A DUPLICATE PHYSICAL ATTEMPT DOES NOT DUPLICATE SEMANTIC COMPLETION.
// English: however many times an object is transmitted, it completes once.
// This is the property that separates a working retry mechanism from a
// reliability violation (21.6 §19).
a_completion_once: assert property (
  @(posedge clk) disable iff (!por_n)
    (complete_fire && (complete_id == TEST_ID))
      |=> !(complete_fire && (complete_id == TEST_ID))
          throughout (1'b1 [*1:$] ##0 (retire_fire && (retire_id == TEST_ID)))
);

Architecture. Four properties, each attached to a specific failure above rather than restating an assignment.

State. The tracking table and the epoch register.

Sampled timing. Property (2) uses |=> with $past, so it checks this cycle's admission against next cycle's tracking state — the correct phase for a registered update. Property (4) is bounded by retirement, an observed event, rather than by an invented window (21.6 §29's three forms).

Contract. Property (1) is the reserve's specification. It must be written against the admission event, not against a steady-state occupancy — checking occupancy after the fact reports the deadlock rather than preventing it.

Failure if omitted. Without (1), §12's deadlock is found in a training run. Without (3), stale responses are applied and corrupt live state. Without (4), a retry mechanism doing its job is indistinguishable from a reliability violation — and the "fix" weakens the retry.

DV/debug. These four map onto §19's verification questions, and three of them survive into silicon as counters (21.7 §21) rather than assertions.

18. Reliability Overhead Is Not Payload

An accelerator's headline metric is bytes moved. A retry moves bytes and delivers nothing new.

CounterAnswers
bytes_wirehow busy the link was
bytes_uniquehow much new data actually arrived
bytes_unique / bytes_wirethe retry overhead

Two readings.

A metric built on wire bytes improves as the link degrades (21.5 §16) — the worst possible property for a performance indicator, and especially damaging here because tensor traffic is so dominant that a few per cent of retry is a large absolute number.

And the correct response to high retry overhead is not in the transfer path. 21.5 §38: the scheduler and the arbiter are doing their jobs; the investigation belongs to signal integrity (21.7 §32's physical-suspicion branch).

19. Worked Performance Model

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
ILLUSTRATIVE. All values are internal to this example and are NOT UCIe,
vendor or product figures (§3). Units are symbolic "bytes per unit time."
 
  Compute demand on this edge      X = 100
  Link available rate              Y = 140
  Memory service rate behind it    Z =  60
 
  Naive upper bound                min(X, Y, Z) = 60
 
  Now the efficiency terms, applied to the LINK only:
    unique / wire fraction         u = 0.90   (10% retry — §18)
    payload / wire fraction        p = 0.85   (framing overhead)
    effective link rate            Y' = 140 x 0.90 x 0.85 = 107.1
 
  Revised bound                    min(100, 107.1, 60) = 60
 
  CONCLUSION: the link's efficiency terms do not matter here at all.
  Z binds by a wide margin, and it still binds after the link is derated.

Four readings.

The efficiency terms changed nothing, because the link was not the binding term and remains non-binding after derating. That is the arithmetic behind §9's wrong architecture.

The check worth running is how much slack the non-binding terms have. Y' = 107.1 against Z = 60 is a wide margin; at Y' = 62 the answer would be fragile, and a small increase in retry would make the link binding.

Applying the efficiency terms to the wrong quantity is a common error. They derate the link, not the memory service — multiplying Z by u would understate memory by 10% for no reason.

And min() is an upper bound, not a prediction. Burstiness, latency-hiding, queue depth and scheduling all sit below it (21.5 §26) — a real system reaches min() only if every edge is kept busy, which is a separate problem.

20. Debugging a Package Bottleneck

The counters, and the order to read them (21.5 §56's tree, scoped to a package).

ReadSays
per-link utilisationlow + downstream full → the constraint is behind the link (§9)
queue occupancy per stagewhich boundary is limiting (21.5 §23)
memory service rateusually the answer
unique / wireretry overhead (§18)
completion latencywhether progress traffic is being starved (§12)
per-class transfer sharewhether one class dominates

And the single most diagnostic pairing is the first row. Low link utilisation with a full downstream queue eliminates the link as a candidate in one read — which is the observation §9's team did not make.

21. Verification Questions

QuestionWhy
does progress traffic advance under maximum bulk load?§12's deadlock
are the classes genuinely independent, or sharing a pool?19.5 §8
what happens to outstanding commands across a recovery?14.2
is semantic completion exactly once under retry?§17 property 4
is a stale route or config epoch rejected?§14, §17 property 3
is host/accelerator ordering preserved where required?21.6 §14
is the progress reserve sized against the true maximum?§13's contract

And the first is the one to write first, because it is the test §12's design passes in directed runs and fails in the first real workload.

22. What Public Material Cannot Tell You

Unless a source explicitly discloses it, none of this is knowable — and none of it appears above.

Not knowableModule framing the question
any accelerator's RTL
queue depths, credit counts, reserve sizing13.2 · 19.4
exact die topology and which dies exist§6 is conceptual
arbitration policy between classes19.3
retry buffer sizing and replay policy14.2
internal protocol carried over any link19.2
per-link latency or bandwidth15.1 · 15.2
exact UCIe widths, rates or modes used§4 — not even the revision is public for the one announced part
internal FDI/RDI structure19.1

And row 8 is worth dwelling on. The clearest UCIe part in this space is described publicly as "UCIe" with no revision named (§4) — so even for the component with the best documentation, the interoperability question 21.6 §32 demands cannot be answered from public material.

23. Common Misconceptions

"AI accelerators need UCIe because AI needs bandwidth." §5, §9: bandwidth is a requirement, not an argument. UCIe changes an edge's ownership; it changes no edge's capacity.

"A chiplet accelerator automatically uses UCIe." 22.1 §12: multi-die predates the standard, and co-designed interfaces have real advantages at internal boundaries.

"UCIe replaces HBM." §8: a category error. HBM is a memory technology with its own interface; UCIe is a die-to-die link. A memory-side chiplet still talks to memory over a memory interface.

"UCIe solves scale-up networking." §8: the external network is outside the package. A UCIe-facing bridge chiplet is a boundary to that network, not a replacement for it.

"A demonstration proves a shipping accelerator." §4: an announced component with no availability date, endorsed by seven organisations, is Level C/D.

"A standard PHY means standard memory semantics." §8, §16: the link is standard; what crosses it and what the far side does with it are separate agreements.

"More links always increase throughput." §19: adding capacity to a non-binding edge changes min() by nothing.

"One shared queue maximises utilisation." §12: it maximises utilisation right up to the deadlock, and the deadlock is structural rather than a slowdown.

"Physical retry traffic is useful AI payload." §18: a metric built on wire bytes improves as the link degrades.

24. Understanding Check

25. Summary

Five things.

Model the package as a bandwidth graph (§5). Four attributes per edge — required rate, available rate, latency, owner — and the binding edge depends on the workload, not the hardware.

A standard link changes ownership, not capacity (§8). It cannot make memory faster, compute wider, coherence simpler, or the external network larger.

So prove the edge binds before standardising it for performance (§9, §19). Standardising a non-binding edge is fine for ownership reasons and produces no speedup — and the conclusion "the standard is slow" is the predictable, wrong outcome.

Progress traffic and bulk traffic must not share a pool (§12–§13). Separate queues fix head-of-line; a reserve fixes credit exhaustion; both are required, and the deadlock is structural rather than a slowdown.

And the public record here is a component supplier, not an accelerator (§4). One dated, primary-sourced announcement with no availability date, seven Level-E endorsements, and no revision named — which is real ecosystem evidence and is not a shipping claim.