Skip to content

UCIe · Module 23

UCIe vs Proprietary D2D

When a co-designed private die-to-die interface is still the right engineering answer, and what an open standard actually buys — the organisational boundary that changes the calculation, why private-interface cost is paid per pairing rather than per product, and the hidden timing assumption that works for years and fails the first time a peer changes.

Chapter 23.1 compared a standard against a standard, with balanced evidence on both sides. This chapter compares a standard against something that, by construction, publishes very little — and the honest answer is genuinely two-sided.

1. The One-Sentence Model

A proprietary die-to-die interface optimises one relationship. An open standard standardises the boundary so that the relationship can change. The trade is not quality against quality — it is co-design freedom against a stable contract, and which one wins depends on whether the two dies will always be owned by the same people.

So this is not a chapter about open being better. A private interface co-designed for exactly one die pair can be better at that job than any standard, and §6 says so plainly. The question is what happens when the pair changes.

2. What This Chapter Owns

QuestionWhere
Layer normalisation as a comparison method23.1 — UCIe vs PCIe
The incumbent-fabric problem; layering vs replacement22.2 — AMD Chiplets on UCIe
The seven-layer chiplet integration contract22.5 — Future SoCs
Reusable IP — parameterisation, configuration, collateral19.6 — Reusable UCIe IP
Interoperability and what compliance establishes20.7 — Compliance Testing
Verification environments and independent models20.2 · 20.4
Fabric vs link — the layer-mismatch chapter23.3 (next)

Three things are new here:

The organisational boundary (§8–§9) — the variable that actually decides this question, and it is not a technical one.

The change-cost model (§10). A private interface's cost is quadratic in pairings, not linear in products, and that is the whole economic argument in one line.

And the hidden-assumption failure (§13–§15): a timing assumption that is true, undocumented, correct for four years, and fatal the first time the peer's generation changes.

3. Sourcing and a Severe Evidence Asymmetry

4. Normalising the Comparison

A proprietary D2D interface and UCIe are not automatically at the same layer either23.1 §2's rule applies here too.

LayerA private D2D interfaceUCIe
semanticsfrequently private and fused with the interfacecarries a protocol; several are defined
reliabilityprivate, or absent by design if the channel is trustedAdapter CRC + retry, described as optional
physicalprivate signalling, co-designed with the packagespecified signalling
scopedie to die, in-packagedie to die, in-package
who must agreetwo teams, privatelyany two conforming implementations

Two readings.

Row 4 is the one row where they genuinely match — both are in-package die-to-die. That is what makes this the fairest comparison in Module 23, and why it can be argued on engineering merit rather than on scope.

Row 1 is where private interfaces often differ structurally. A private interface may not have a separable protocol layer at all — semantics and transport can be fused, because there was never a reason to separate them. That fusion is efficient and is exactly what makes the interface unreusable (23.1 §15's coupling, as a deliberate design choice rather than an accident).

5. Two Packages, Two Agreements

A block diagram comparing two package arrangements. On the left, a co-designed package shows die A and die B connected by a private die-to-die interface, with a note that width, clocking, framing, reset, reliability and semantics are all settled privately inside one team. On the right, an ecosystem package shows chiplet A and chiplet B connected by a standard contract boundary, with a note that the same variables must be expressed as capabilities, negotiated at bring-up and verified against every legal peer. A centre label states that the variables move from private knowledge to explicit agreement rather than disappearing.Die Aone teamSAME VARIABLESprivate vs explicitChiplet Asupplier 1Private D2Dsettled internallyStandard contractnegotiated (§11)Die Bsame teamChiplet Bsupplier 212
The same physical arrangement under two different agreements. On the left, two co-designed dies share a private interface and every architectural variable is settled inside one team. On the right, the boundary is a published contract and each of those variables must be expressed, negotiated and verified. The variables do not disappear — they move from private knowledge to explicit agreement.

Three things to read.

The physical arrangement is identical. Two dies, one package, one boundary. Nothing about the silicon requires one agreement or the other.

The centre node is the argument. Width, clocking, framing, reset, reliability, versioning and semantics exist in both cases. On the left they are private knowledge; on the right they must be expressed, negotiated and verified.

And that is a real cost, honestly. Expressing a variable costs wire, logic, bring-up time and verification. The standard column is not free — §7 is what it buys in exchange.

6. What a Private Interface Genuinely Buys

Stated without hedging, because pretending otherwise makes the chapter useless.

AdvantageWhy it is real
exact width and clocking for this pairno need to support widths nobody will use
framing tuned to the actual trafficoverhead sized to the payload that exists
no negotiation surfaceno capability exchange, no version handling, no mismatch cases
co-design with the packagethe interface and the substrate evolve together
freedom to change every generationno compatibility obligation to anyone
narrow verification matrixone peer, one interpretation (20.2)
omit unused features entirelya feature nobody needs costs nothing if it does not exist
potentially lower overheadfewer layers, fewer generic mechanisms

Three readings.

Row 3 is larger than it looks. A negotiated interface must handle every combination of capabilities two conforming implementations might present, including all the ones that never occur in practice (22.5 §8). A private interface has one combination.

Row 5 is the deepest advantage. Compatibility obligations accumulate; a private interface can be redesigned each generation with no legacy, which over several generations is a large amount of freedom.

And "potentially lower overhead" is deliberately hedged. It is a plausible consequence of fewer generic mechanisms — whether it produces measurable application benefit depends entirely on whether that boundary is the binding constraint (22.3 §9), which §16 develops.

7. What Standardisation Buys

AdvantageCondition
a die can come from elsewherethe decisive one — §8
a die can be sold into someone else's packagesame argument, other direction
independent development against a stable boundarythe boundary must actually be stable
reusable verification collateralif it is produced (22.5 §14)
cross-foundry potentialdemonstrated (22.1 §10)
supplier independence and sourcing flexibilitycommercial, not technical
common debug and observability expectationspartially — the specification does not mandate collateral

And the critical caveat, which 22.5 established in full: standardisation does not produce plug-and-play. Two chiplets can both conform at the wire and remain unintegrable on protocol semantics, management, reset sequencing, power states, security and version profile. The standard moves the boundary from private to explicit; it does not remove the other six agreements.

8. The Organisational Boundary

The flagship idea of this chapter, and it is not a technical variable.

The value of a stable contract is a function of how likely the relationship is to change. Two dies owned by one team, on one schedule, in one product family, have a relationship that will not change without both sides knowing. Two dies that cross a team, a business unit, a company, a process node or a product generation have a relationship that will change without asking.

The boundary crossesContract value
nothing — one team, one schedulelow — private co-design likely wins
two teams in one companymoderate and routinely underestimated — §12
a business unit or product-line boundaryhigh
a companydecisive — a private interface cannot cross it at all
a process node or foundryhigh
a product generationhigh — the peer you integrate with in three years is not the one you designed against

Three properties.

Row 4 is not a preference — it is structural. A private interface requires a shared internal specification, and there is no mechanism for that across a commercial boundary. 22.2 §9: the decisive pressures all involve a die crossing an organisation.

Row 6 is the one architects miss. "Both dies are ours" is true at design time and says nothing about the integration three years later, when one die has been respun and the other has not.

And row 2 is §12's whole subject. Same company is not the same as same assumptions — and a hidden assumption inside one company fails exactly as hard as one across two.

9. Where Each Wins

BoundaryLikely answer
two tightly coupled compute dies, one team, one productprivate — §6 dominates
compute ↔ IO die reused across a familyeither — depends on whether the family diverges
compute ↔ a die from another business unitstandard — §8 row 3
compute ↔ a third-party chipletstandard — no alternative exists
selling a die into someone else's packagestandard — the customer decides
a boundary expected to survive several generationsstandard — §8 row 6

And the honest summary is that both will persist, in the same package, at different boundaries — 22.2 §15's conclusion, reached from the other direction.

10. The Change-Cost Model

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
ILLUSTRATIVE model. Symbolic units — no vendor cost data is used or implied.
 
PRIVATE INTERFACES, N die types that must interoperate:
 
    each PAIRING needs its own agreement, bring-up and verification
    pairings  =  N(N-1)/2
 
    N = 2  ->   1 pairing      (private is clearly cheapest)
    N = 4  ->   6 pairings
    N = 6  ->  15 pairings
    N = 8  ->  28 pairings
 
    total_cost  ~=  N(N-1)/2  x  C_pair
 
STANDARD BOUNDARY:
 
    each die implements the contract ONCE
    total_cost  ~=  N x C_conform  +  P x C_interop_test
 
    where P is the pairings you actually SHIP and test — normally far
    fewer than N(N-1)/2, because you only test what you sell.
 
CROSSOVER: private wins while N is small. The curves cross because one
side is QUADRATIC in die types and the other is LINEAR, plus the pairings
you choose to test.

Four readings.

The shape is the argument, not the numbers. N(N-1)/2 against Nquadratic against linear — and no cost estimate is needed to see where that goes.

C_conform is genuinely higher than a private interface's per-pair cost, which is why private wins at small N. Conforming means implementing capability exchange, version handling and mismatch cases that a private interface simply does not have (§6 row 3).

But P is the honest term that keeps this from being propaganda. A standard does not eliminate pairwise interoperability testing (20.7, 22.5 §15) — it reduces the design work to linear while leaving testing proportional to the pairings you ship.

And N grows for reasons unrelated to interconnect strategy — process disaggregation, product variants, third-party components, generational respins (22.1 §18). A team that chose private at N=2 may be at N=6 four years later without ever having decided to be.

11. Illustrative — A Private Interface

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// ILLUSTRATIVE ONLY. A tightly co-designed private interface. Everything here
// is efficient BECAUSE it is fused — and unreusable for the same reason.
module private_d2d_tx (
  input  logic                clk,
  input  logic                rst_n,
  // Internal opcodes: the SEMANTICS are the interface. There is no separable
  // protocol layer, because there was never a second protocol (§4 row 1).
  input  logic [3:0]          op,          // meaning known only to A and B
  input  logic [255:0]        payload,     // width fixed to THIS pair
  input  logic                req,
  output logic                gnt,
  output logic [W_PRIV-1:0]   wire_out
);
  // Framing is minimal because both sides know the format. No header, no
  // length, no version, no capability — nothing that is not needed.
  assign wire_out = { op, payload[W_PRIV-5:0] };
 
  // Grant is combinational because die B is known to always accept within a
  // fixed pipeline depth. That knowledge lives in a design review, an email
  // thread, and nobody's RTL. §13 is what that costs.
  assign gnt = req && !local_stall;
endmodule

Architecture. Semantics, framing and flow control fused into one narrow interface, sized exactly for one die pair.

State. Almost none — which is precisely the efficiency.

Cycle/event behaviour. req/gnt in a single cycle, on the assumption that the peer accepts within a known depth.

Contract. The contract exists and is not written down anywhere in the design. Opcode meanings, the 256-bit payload width, the acceptance-depth assumption, the reset relationship — all are agreements between two teams held outside the RTL.

Failure. §13. Every unexpressed agreement is a failure waiting for the peer to change.

DV/debug. The verification matrix is one peer, one interpretation — genuinely cheaper — and the monitor written for it is unusable against anything else, because it decodes a format that only exists here.

12. Wrong — "Same Company, So the Assumption Is Safe"

The most expensive false belief in this chapter.

Why it fails
teams changethe people who agreed it have moved on
generations divergedie A respins; die B does not
schedules decoupleone side ships against an older understanding
the assumption was never written downso a reviewer cannot check it
the failure is silicon-only§13 — simulation used the same assumption on both sides

And the last row is the killer. If both models embed the same assumption, the testbench proves the assumption is self-consistent, not that it is true21.6 §24's shared-model failure, arriving through an organisational path rather than a technical one.

13. Wrong RTL — the Hidden Timing Assumption

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// WRONG — die A assumes die B always accepts within N cycles. The assumption
// is TRUE for four years, is never expressed, and is not checkable.
localparam int ASSUMED_ACCEPT_CYCLES = 3;
 
always_ff @(posedge clk or negedge rst_n) begin
  if (!rst_n) outstanding_q <= '0;
  else if (req && gnt) begin
    outstanding_q <= outstanding_q + 1'b1;
    // The buffer is sized for ASSUMED_ACCEPT_CYCLES of outstanding work.
    // There is no back-pressure path, because none was ever needed.
  end
end

The failure, generation by generation.

GenerationDie B behaviourResult
1–3accepts within 3 cycles, alwaysworks perfectly for years
4B adds a pipeline stage; worst case is now 6A's buffer overflows
symptomdata loss under load only
simulationpasses — B's model was updated, A's assumption was not expressed to check
first evidencesilicon, under a workload nobody ran in simulation

Five properties.

Nothing was broken by anyone. Die B made a legitimate design change and violated no stated contract, because there was no stated contract.

The assumption is invisible to review. A reviewer reading die A's RTL sees a buffer size; nothing says why it is that size, so nobody flags it when B's timing changes.

It fails only under load, so it survives directed testing and appears in the first sustained workload — 22.3 §12's pattern.

And the fix is not a bigger buffer. A bigger buffer moves the failure to a longer stall; the defect is that a timing property is being relied on and not expressed. §14.

This is also why §12's "same company" belief is expensive: the organisational boundary here is time, not a company — and time crosses every boundary.

14. Corrected RTL — Express the Contract

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// CORRECTED. The assumption becomes an explicit capability plus real
// back-pressure. Both changes are needed: the capability makes the assumption
// REVIEWABLE, and the back-pressure makes it UNNECESSARY.
typedef struct packed {
  logic [7:0]  version;              // explicit — §17
  logic [15:0] max_accept_latency;   // was ASSUMED_ACCEPT_CYCLES, now STATED
  logic [15:0] max_outstanding;      // what the peer can hold
  logic [15:0] max_payload_bytes;    // was a hard-coded width
  logic        supports_backpressure;
  logic        retains_across_reset; // 23.1 §17's field
} d2d_contract_t;
 
d2d_contract_t peer_caps_q;
 
// Buffer depth is DERIVED from the peer's stated capability, not from a
// remembered number. Change the peer, and the derivation follows.
logic [15:0] required_depth;
assign required_depth = peer_caps_q.max_accept_latency; // + margin per design
 
// And real back-pressure, so exceeding the estimate STALLS instead of losing
// data. Acceptance is qualified — offering is not acceptance (21.5 §20).
assign req = have_work && (outstanding_q < peer_caps_q.max_outstanding);
 
always_ff @(posedge clk or negedge rst_n) begin
  if (!rst_n) outstanding_q <= '0;
  else
    // ONE signed next-state expression so a simultaneous issue and completion
    // nets correctly rather than one update being lost (19.5 §14).
    outstanding_q <= outstanding_q + 16'(req && gnt) - 16'(completion_fire);
end
 
// MANDATORY. English: outstanding work never exceeds what the peer stated it
// can hold. Fires at the issue that §13 would have lost data on — before any
// loss, and in simulation, because the bound is now a VALUE not a belief.
a_outstanding_within_peer_capability: assert property (
  @(posedge clk) disable iff (!rst_n)
    (outstanding_q <= peer_caps_q.max_outstanding)
);
 
// MANDATORY. English: no hidden fixed-latency assumption remains — the design
// must tolerate the peer taking its full stated latency. Fires if any path
// still assumes faster acceptance.
a_tolerates_stated_latency: assert property (
  @(posedge clk) disable iff (!rst_n)
    (req && !gnt) |-> ##[0:MAX_STATED_LATENCY] (gnt || !req)
);

Architecture. A capability record plus derived sizing plus real back-pressure — three changes, and each covers a different failure mode.

State. The peer's capability record and an outstanding counter.

Cycle/event behaviour. outstanding_q is a single signed next-state expression, so a same-cycle issue and completion net to zero rather than one being lost to two sequential assignments.

Contract. max_accept_latency is now a stated value the peer publishes, so a reviewer can check the buffer derivation and a regression can check the bound. The point is not that the number changed — it is that it became checkable.

Failure. Deriving depth from the capability but omitting back-pressure leaves the design correct only while the capability record is accurate. Adding back-pressure but keeping the hard-coded depth stalls unnecessarily and hides the mismatch. Both are needed.

DV/debug. peer_caps_q belongs in a debug snapshot (21.7 §9). In an interoperability matrix, the failing pairing is usually distinguished by exactly one capability field — and without the record, that comparison cannot be made at all.

15. What the Refactor Actually Changed

Before (§11, §13)After (§14)
the assumptionexisted, unwrittenstated as a value
reviewabilitynonea reviewer can check the derivation
failure modesilent data loss in silicona stall, and an assertion in simulation
peer changebreaks silentlydetected at negotiation
costzero wires, zero logica capability exchange and a back-pressure path

And the last row is the honest one. The refactor costs real hardware and real bring-up time — which is exactly §6's argument for a private interface at small N, and exactly why the decision belongs to §8's organisational question rather than to taste.

16. The Performance Myth

"Proprietary is always faster" is not established, and the reasoning matters more than the verdict.

ClaimStatus
a private interface has more optimisation freedomtrue — §6
more freedom can produce a better boundarytrue
a better boundary produces better applicationsonly if that boundary binds

Three readings.

The third row is 22.3 §9's entire lesson. If the binding constraint is memory service, the consumer, or the software scheduler, an improvement at a non-binding boundary changes min() by nothing (22.3 §19).

And there is a second-order effect that runs the other way: a standard boundary can enable sourcing a better die than the one you would have built. A slightly less efficient interface to a much better component is a net win, and no interface-level comparison captures it.

So the defensible statement is narrow: a private interface can be optimised in ways a standard cannot, and whether that matters is a system question, not an interconnect question. 23.1 §13's discipline: without a stated scope and a binding-constraint analysis, a speed claim is not a measurement.

17. Versioning Is Not Only a Standards Problem

A private interface versions implicitly: "both dies are revision 7."

That works whileIt fails when
both dies ship togetherone respins independently — §13
one team owns bothownership splits
there is one productvariants diverge

And the correction is the same either way: an explicit version field, an intersection of capabilities, and the minimum of the two revisions (22.5 §10's max bug applies unchanged). A reusable private chiplet needs version contracts as much as a standard one does — the standard just forces the issue.

18. Verification Cost, Honestly

PrivateStandard
interpretations to verify againstoneevery legal peer (20.2)
model reusenone across pairscollateral is reusable if it exists (22.5 §14)
pairwise interop testingrequired per pairstill required per shipped pair
conformance testingnot applicablevalidates against a reference (20.7)

Two readings, and both cut against easy answers.

Row 1 favours private and is real. Verifying against one known peer is genuinely cheaper than verifying against a space of conforming implementations.

Row 3 is the claim to resist in both directions. A standard does not remove pairwise testing (22.5 §15) — and a private interface does not avoid it either. What the standard changes is that the collateral can be shared and conformance can be checked against a reference; the pairing still has to be tested.

19. Common Misconceptions

"Open is always better." §6, §9: a co-designed private interface wins at a boundary that will not change, and §10's crossover is real.

"Proprietary is always faster." §16: more optimisation freedom is real; whether it produces application benefit depends on whether that boundary binds.

"A standard link makes chiplets plug-and-play." §7, 22.5 §4: six other agreements remain, and a conforming link can still fail to integrate.

"A private interface needs less verification." §18: less interpretation space, yes. Pairwise testing, no — and §13's failure is exactly what the narrow matrix missed.

"If two teams are in the same company, hidden assumptions are safe." §12: teams change, generations diverge, and both models embed the same assumption so simulation proves nothing.

"Versioning matters only for standards." §17: implicit versioning fails the moment one die respins independently.

"Standardisation eliminates pairwise interoperability testing." §18: it makes collateral reusable and conformance checkable. The pairing still ships and still gets tested.

"There is little public detail on proprietary interfaces, so they must be simple." §3: publishing is a requirement of interoperability, not a measure of sophistication. Absence of documentation says nothing about the design.

20. Understanding Check

21. Summary

Five things.

Both options are legitimate and the trade is co-design freedom against a stable contract (§1, §6). A private interface can be better at the one job it was built for.

The variables do not disappear — they move (§5). Width, clocking, framing, reset, reliability, versioning and semantics exist either way; the standard makes them explicit, at real cost.

The organisational boundary decides it (§8), and time is one of those boundaries — which is why "both dies are ours" is a statement about today (§12).

Private cost is quadratic in die types; standard design cost is linear plus the pairings you ship (§10). The shape is the argument; no cost data is needed.

And unexpressed assumptions are the characteristic private-interface failure (§13): true for years, invisible to review, self-consistent in simulation, and fatal in silicon under load. The fix is to make the assumption a value.

On evidence: this comparison is structurally asymmetric (§3). Standards publish because interoperability requires it; private interfaces do not because it does not. I state no electrical parameter, width, framing detail or performance figure for any specific proprietary interface, and the registry's "AIB / IFOP / etc." is a curriculum hint rather than evidence.