UCIe · Module 23
UCIe vs Proprietary D2D
When a co-designed private die-to-die interface is still the right engineering answer, and what an open standard actually buys — the organisational boundary that changes the calculation, why private-interface cost is paid per pairing rather than per product, and the hidden timing assumption that works for years and fails the first time a peer changes.
Chapter 23.1 compared a standard against a standard, with balanced evidence on both sides. This chapter compares a standard against something that, by construction, publishes very little — and the honest answer is genuinely two-sided.
1. The One-Sentence Model
A proprietary die-to-die interface optimises one relationship. An open standard standardises the boundary so that the relationship can change. The trade is not quality against quality — it is co-design freedom against a stable contract, and which one wins depends on whether the two dies will always be owned by the same people.
So this is not a chapter about open being better. A private interface co-designed for exactly one die pair can be better at that job than any standard, and §6 says so plainly. The question is what happens when the pair changes.
2. What This Chapter Owns
| Question | Where |
|---|---|
| Layer normalisation as a comparison method | 23.1 — UCIe vs PCIe |
| The incumbent-fabric problem; layering vs replacement | 22.2 — AMD Chiplets on UCIe |
| The seven-layer chiplet integration contract | 22.5 — Future SoCs |
| Reusable IP — parameterisation, configuration, collateral | 19.6 — Reusable UCIe IP |
| Interoperability and what compliance establishes | 20.7 — Compliance Testing |
| Verification environments and independent models | 20.2 · 20.4 |
| Fabric vs link — the layer-mismatch chapter | 23.3 (next) |
Three things are new here:
The organisational boundary (§8–§9) — the variable that actually decides this question, and it is not a technical one.
The change-cost model (§10). A private interface's cost is quadratic in pairings, not linear in products, and that is the whole economic argument in one line.
And the hidden-assumption failure (§13–§15): a timing assumption that is true, undocumented, correct for four years, and fatal the first time the peer's generation changes.
3. Sourcing and a Severe Evidence Asymmetry
4. Normalising the Comparison
A proprietary D2D interface and UCIe are not automatically at the same layer either — 23.1 §2's rule applies here too.
| Layer | A private D2D interface | UCIe |
|---|---|---|
| semantics | frequently private and fused with the interface | carries a protocol; several are defined |
| reliability | private, or absent by design if the channel is trusted | Adapter CRC + retry, described as optional |
| physical | private signalling, co-designed with the package | specified signalling |
| scope | die to die, in-package | die to die, in-package |
| who must agree | two teams, privately | any two conforming implementations |
Two readings.
Row 4 is the one row where they genuinely match — both are in-package die-to-die. That is what makes this the fairest comparison in Module 23, and why it can be argued on engineering merit rather than on scope.
Row 1 is where private interfaces often differ structurally. A private interface may not have a separable protocol layer at all — semantics and transport can be fused, because there was never a reason to separate them. That fusion is efficient and is exactly what makes the interface unreusable (23.1 §15's coupling, as a deliberate design choice rather than an accident).
5. Two Packages, Two Agreements
Three things to read.
The physical arrangement is identical. Two dies, one package, one boundary. Nothing about the silicon requires one agreement or the other.
The centre node is the argument. Width, clocking, framing, reset, reliability, versioning and semantics exist in both cases. On the left they are private knowledge; on the right they must be expressed, negotiated and verified.
And that is a real cost, honestly. Expressing a variable costs wire, logic, bring-up time and verification. The standard column is not free — §7 is what it buys in exchange.
6. What a Private Interface Genuinely Buys
Stated without hedging, because pretending otherwise makes the chapter useless.
| Advantage | Why it is real |
|---|---|
| exact width and clocking for this pair | no need to support widths nobody will use |
| framing tuned to the actual traffic | overhead sized to the payload that exists |
| no negotiation surface | no capability exchange, no version handling, no mismatch cases |
| co-design with the package | the interface and the substrate evolve together |
| freedom to change every generation | no compatibility obligation to anyone |
| narrow verification matrix | one peer, one interpretation (20.2) |
| omit unused features entirely | a feature nobody needs costs nothing if it does not exist |
| potentially lower overhead | fewer layers, fewer generic mechanisms |
Three readings.
Row 3 is larger than it looks. A negotiated interface must handle every combination of capabilities two conforming implementations might present, including all the ones that never occur in practice (22.5 §8). A private interface has one combination.
Row 5 is the deepest advantage. Compatibility obligations accumulate; a private interface can be redesigned each generation with no legacy, which over several generations is a large amount of freedom.
And "potentially lower overhead" is deliberately hedged. It is a plausible consequence of fewer generic mechanisms — whether it produces measurable application benefit depends entirely on whether that boundary is the binding constraint (22.3 §9), which §16 develops.
7. What Standardisation Buys
| Advantage | Condition |
|---|---|
| a die can come from elsewhere | the decisive one — §8 |
| a die can be sold into someone else's package | same argument, other direction |
| independent development against a stable boundary | the boundary must actually be stable |
| reusable verification collateral | if it is produced (22.5 §14) |
| cross-foundry potential | demonstrated (22.1 §10) |
| supplier independence and sourcing flexibility | commercial, not technical |
| common debug and observability expectations | partially — the specification does not mandate collateral |
And the critical caveat, which 22.5 established in full: standardisation does not produce plug-and-play. Two chiplets can both conform at the wire and remain unintegrable on protocol semantics, management, reset sequencing, power states, security and version profile. The standard moves the boundary from private to explicit; it does not remove the other six agreements.
8. The Organisational Boundary
The flagship idea of this chapter, and it is not a technical variable.
The value of a stable contract is a function of how likely the relationship is to change. Two dies owned by one team, on one schedule, in one product family, have a relationship that will not change without both sides knowing. Two dies that cross a team, a business unit, a company, a process node or a product generation have a relationship that will change without asking.
| The boundary crosses | Contract value |
|---|---|
| nothing — one team, one schedule | low — private co-design likely wins |
| two teams in one company | moderate and routinely underestimated — §12 |
| a business unit or product-line boundary | high |
| a company | decisive — a private interface cannot cross it at all |
| a process node or foundry | high |
| a product generation | high — the peer you integrate with in three years is not the one you designed against |
Three properties.
Row 4 is not a preference — it is structural. A private interface requires a shared internal specification, and there is no mechanism for that across a commercial boundary. 22.2 §9: the decisive pressures all involve a die crossing an organisation.
Row 6 is the one architects miss. "Both dies are ours" is true at design time and says nothing about the integration three years later, when one die has been respun and the other has not.
And row 2 is §12's whole subject. Same company is not the same as same assumptions — and a hidden assumption inside one company fails exactly as hard as one across two.
9. Where Each Wins
| Boundary | Likely answer |
|---|---|
| two tightly coupled compute dies, one team, one product | private — §6 dominates |
| compute ↔ IO die reused across a family | either — depends on whether the family diverges |
| compute ↔ a die from another business unit | standard — §8 row 3 |
| compute ↔ a third-party chiplet | standard — no alternative exists |
| selling a die into someone else's package | standard — the customer decides |
| a boundary expected to survive several generations | standard — §8 row 6 |
And the honest summary is that both will persist, in the same package, at different boundaries — 22.2 §15's conclusion, reached from the other direction.
10. The Change-Cost Model
ILLUSTRATIVE model. Symbolic units — no vendor cost data is used or implied.
PRIVATE INTERFACES, N die types that must interoperate:
each PAIRING needs its own agreement, bring-up and verification
pairings = N(N-1)/2
N = 2 -> 1 pairing (private is clearly cheapest)
N = 4 -> 6 pairings
N = 6 -> 15 pairings
N = 8 -> 28 pairings
total_cost ~= N(N-1)/2 x C_pair
STANDARD BOUNDARY:
each die implements the contract ONCE
total_cost ~= N x C_conform + P x C_interop_test
where P is the pairings you actually SHIP and test — normally far
fewer than N(N-1)/2, because you only test what you sell.
CROSSOVER: private wins while N is small. The curves cross because one
side is QUADRATIC in die types and the other is LINEAR, plus the pairings
you choose to test.Four readings.
The shape is the argument, not the numbers. N(N-1)/2 against N — quadratic against linear — and no cost estimate is needed to see where that goes.
C_conform is genuinely higher than a private interface's per-pair cost, which is why private wins at small N. Conforming means implementing capability exchange, version handling and mismatch cases that a private interface simply does not have (§6 row 3).
But P is the honest term that keeps this from being propaganda. A standard does not eliminate pairwise interoperability testing (20.7, 22.5 §15) — it reduces the design work to linear while leaving testing proportional to the pairings you ship.
And N grows for reasons unrelated to interconnect strategy — process disaggregation, product variants, third-party components, generational respins (22.1 §18). A team that chose private at N=2 may be at N=6 four years later without ever having decided to be.
11. Illustrative — A Private Interface
// ILLUSTRATIVE ONLY. A tightly co-designed private interface. Everything here
// is efficient BECAUSE it is fused — and unreusable for the same reason.
module private_d2d_tx (
input logic clk,
input logic rst_n,
// Internal opcodes: the SEMANTICS are the interface. There is no separable
// protocol layer, because there was never a second protocol (§4 row 1).
input logic [3:0] op, // meaning known only to A and B
input logic [255:0] payload, // width fixed to THIS pair
input logic req,
output logic gnt,
output logic [W_PRIV-1:0] wire_out
);
// Framing is minimal because both sides know the format. No header, no
// length, no version, no capability — nothing that is not needed.
assign wire_out = { op, payload[W_PRIV-5:0] };
// Grant is combinational because die B is known to always accept within a
// fixed pipeline depth. That knowledge lives in a design review, an email
// thread, and nobody's RTL. §13 is what that costs.
assign gnt = req && !local_stall;
endmoduleArchitecture. Semantics, framing and flow control fused into one narrow interface, sized exactly for one die pair.
State. Almost none — which is precisely the efficiency.
Cycle/event behaviour. req/gnt in a single cycle, on the assumption that the peer accepts within a known depth.
Contract. The contract exists and is not written down anywhere in the design. Opcode meanings, the 256-bit payload width, the acceptance-depth assumption, the reset relationship — all are agreements between two teams held outside the RTL.
Failure. §13. Every unexpressed agreement is a failure waiting for the peer to change.
DV/debug. The verification matrix is one peer, one interpretation — genuinely cheaper — and the monitor written for it is unusable against anything else, because it decodes a format that only exists here.
12. Wrong — "Same Company, So the Assumption Is Safe"
The most expensive false belief in this chapter.
| Why it fails | |
|---|---|
| teams change | the people who agreed it have moved on |
| generations diverge | die A respins; die B does not |
| schedules decouple | one side ships against an older understanding |
| the assumption was never written down | so a reviewer cannot check it |
| the failure is silicon-only | §13 — simulation used the same assumption on both sides |
And the last row is the killer. If both models embed the same assumption, the testbench proves the assumption is self-consistent, not that it is true — 21.6 §24's shared-model failure, arriving through an organisational path rather than a technical one.
13. Wrong RTL — the Hidden Timing Assumption
// WRONG — die A assumes die B always accepts within N cycles. The assumption
// is TRUE for four years, is never expressed, and is not checkable.
localparam int ASSUMED_ACCEPT_CYCLES = 3;
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) outstanding_q <= '0;
else if (req && gnt) begin
outstanding_q <= outstanding_q + 1'b1;
// The buffer is sized for ASSUMED_ACCEPT_CYCLES of outstanding work.
// There is no back-pressure path, because none was ever needed.
end
endThe failure, generation by generation.
| Generation | Die B behaviour | Result |
|---|---|---|
| 1–3 | accepts within 3 cycles, always | works perfectly for years |
| 4 | B adds a pipeline stage; worst case is now 6 | A's buffer overflows |
| — | symptom | data loss under load only |
| — | simulation | passes — B's model was updated, A's assumption was not expressed to check |
| — | first evidence | silicon, under a workload nobody ran in simulation |
Five properties.
Nothing was broken by anyone. Die B made a legitimate design change and violated no stated contract, because there was no stated contract.
The assumption is invisible to review. A reviewer reading die A's RTL sees a buffer size; nothing says why it is that size, so nobody flags it when B's timing changes.
It fails only under load, so it survives directed testing and appears in the first sustained workload — 22.3 §12's pattern.
And the fix is not a bigger buffer. A bigger buffer moves the failure to a longer stall; the defect is that a timing property is being relied on and not expressed. §14.
This is also why §12's "same company" belief is expensive: the organisational boundary here is time, not a company — and time crosses every boundary.
14. Corrected RTL — Express the Contract
// CORRECTED. The assumption becomes an explicit capability plus real
// back-pressure. Both changes are needed: the capability makes the assumption
// REVIEWABLE, and the back-pressure makes it UNNECESSARY.
typedef struct packed {
logic [7:0] version; // explicit — §17
logic [15:0] max_accept_latency; // was ASSUMED_ACCEPT_CYCLES, now STATED
logic [15:0] max_outstanding; // what the peer can hold
logic [15:0] max_payload_bytes; // was a hard-coded width
logic supports_backpressure;
logic retains_across_reset; // 23.1 §17's field
} d2d_contract_t;
d2d_contract_t peer_caps_q;
// Buffer depth is DERIVED from the peer's stated capability, not from a
// remembered number. Change the peer, and the derivation follows.
logic [15:0] required_depth;
assign required_depth = peer_caps_q.max_accept_latency; // + margin per design
// And real back-pressure, so exceeding the estimate STALLS instead of losing
// data. Acceptance is qualified — offering is not acceptance (21.5 §20).
assign req = have_work && (outstanding_q < peer_caps_q.max_outstanding);
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) outstanding_q <= '0;
else
// ONE signed next-state expression so a simultaneous issue and completion
// nets correctly rather than one update being lost (19.5 §14).
outstanding_q <= outstanding_q + 16'(req && gnt) - 16'(completion_fire);
end
// MANDATORY. English: outstanding work never exceeds what the peer stated it
// can hold. Fires at the issue that §13 would have lost data on — before any
// loss, and in simulation, because the bound is now a VALUE not a belief.
a_outstanding_within_peer_capability: assert property (
@(posedge clk) disable iff (!rst_n)
(outstanding_q <= peer_caps_q.max_outstanding)
);
// MANDATORY. English: no hidden fixed-latency assumption remains — the design
// must tolerate the peer taking its full stated latency. Fires if any path
// still assumes faster acceptance.
a_tolerates_stated_latency: assert property (
@(posedge clk) disable iff (!rst_n)
(req && !gnt) |-> ##[0:MAX_STATED_LATENCY] (gnt || !req)
);Architecture. A capability record plus derived sizing plus real back-pressure — three changes, and each covers a different failure mode.
State. The peer's capability record and an outstanding counter.
Cycle/event behaviour. outstanding_q is a single signed next-state expression, so a same-cycle issue and completion net to zero rather than one being lost to two sequential assignments.
Contract. max_accept_latency is now a stated value the peer publishes, so a reviewer can check the buffer derivation and a regression can check the bound. The point is not that the number changed — it is that it became checkable.
Failure. Deriving depth from the capability but omitting back-pressure leaves the design correct only while the capability record is accurate. Adding back-pressure but keeping the hard-coded depth stalls unnecessarily and hides the mismatch. Both are needed.
DV/debug. peer_caps_q belongs in a debug snapshot (21.7 §9). In an interoperability matrix, the failing pairing is usually distinguished by exactly one capability field — and without the record, that comparison cannot be made at all.
15. What the Refactor Actually Changed
| Before (§11, §13) | After (§14) | |
|---|---|---|
| the assumption | existed, unwritten | stated as a value |
| reviewability | none | a reviewer can check the derivation |
| failure mode | silent data loss in silicon | a stall, and an assertion in simulation |
| peer change | breaks silently | detected at negotiation |
| cost | zero wires, zero logic | a capability exchange and a back-pressure path |
And the last row is the honest one. The refactor costs real hardware and real bring-up time — which is exactly §6's argument for a private interface at small N, and exactly why the decision belongs to §8's organisational question rather than to taste.
16. The Performance Myth
"Proprietary is always faster" is not established, and the reasoning matters more than the verdict.
| Claim | Status |
|---|---|
| a private interface has more optimisation freedom | true — §6 |
| more freedom can produce a better boundary | true |
| a better boundary produces better applications | only if that boundary binds |
Three readings.
The third row is 22.3 §9's entire lesson. If the binding constraint is memory service, the consumer, or the software scheduler, an improvement at a non-binding boundary changes min() by nothing (22.3 §19).
And there is a second-order effect that runs the other way: a standard boundary can enable sourcing a better die than the one you would have built. A slightly less efficient interface to a much better component is a net win, and no interface-level comparison captures it.
So the defensible statement is narrow: a private interface can be optimised in ways a standard cannot, and whether that matters is a system question, not an interconnect question. 23.1 §13's discipline: without a stated scope and a binding-constraint analysis, a speed claim is not a measurement.
17. Versioning Is Not Only a Standards Problem
A private interface versions implicitly: "both dies are revision 7."
| That works while | It fails when |
|---|---|
| both dies ship together | one respins independently — §13 |
| one team owns both | ownership splits |
| there is one product | variants diverge |
And the correction is the same either way: an explicit version field, an intersection of capabilities, and the minimum of the two revisions (22.5 §10's max bug applies unchanged). A reusable private chiplet needs version contracts as much as a standard one does — the standard just forces the issue.
18. Verification Cost, Honestly
| Private | Standard | |
|---|---|---|
| interpretations to verify against | one | every legal peer (20.2) |
| model reuse | none across pairs | collateral is reusable if it exists (22.5 §14) |
| pairwise interop testing | required per pair | still required per shipped pair |
| conformance testing | not applicable | validates against a reference (20.7) |
Two readings, and both cut against easy answers.
Row 1 favours private and is real. Verifying against one known peer is genuinely cheaper than verifying against a space of conforming implementations.
Row 3 is the claim to resist in both directions. A standard does not remove pairwise testing (22.5 §15) — and a private interface does not avoid it either. What the standard changes is that the collateral can be shared and conformance can be checked against a reference; the pairing still has to be tested.
19. Common Misconceptions
"Open is always better." §6, §9: a co-designed private interface wins at a boundary that will not change, and §10's crossover is real.
"Proprietary is always faster." §16: more optimisation freedom is real; whether it produces application benefit depends on whether that boundary binds.
"A standard link makes chiplets plug-and-play." §7, 22.5 §4: six other agreements remain, and a conforming link can still fail to integrate.
"A private interface needs less verification." §18: less interpretation space, yes. Pairwise testing, no — and §13's failure is exactly what the narrow matrix missed.
"If two teams are in the same company, hidden assumptions are safe." §12: teams change, generations diverge, and both models embed the same assumption so simulation proves nothing.
"Versioning matters only for standards." §17: implicit versioning fails the moment one die respins independently.
"Standardisation eliminates pairwise interoperability testing." §18: it makes collateral reusable and conformance checkable. The pairing still ships and still gets tested.
"There is little public detail on proprietary interfaces, so they must be simple." §3: publishing is a requirement of interoperability, not a measure of sophistication. Absence of documentation says nothing about the design.
20. Understanding Check
21. Summary
Five things.
Both options are legitimate and the trade is co-design freedom against a stable contract (§1, §6). A private interface can be better at the one job it was built for.
The variables do not disappear — they move (§5). Width, clocking, framing, reset, reliability, versioning and semantics exist either way; the standard makes them explicit, at real cost.
The organisational boundary decides it (§8), and time is one of those boundaries — which is why "both dies are ours" is a statement about today (§12).
Private cost is quadratic in die types; standard design cost is linear plus the pairings you ship (§10). The shape is the argument; no cost data is needed.
And unexpressed assumptions are the characteristic private-interface failure (§13): true for years, invisible to review, self-consistent in simulation, and fatal in silicon under load. The fix is to make the assumption a value.
On evidence: this comparison is structurally asymmetric (§3). Standards publish because interoperability requires it; private interfaces do not because it does not. I state no electrical parameter, width, framing detail or performance figure for any specific proprietary interface, and the registry's "AIB / IFOP / etc." is a curriculum hint rather than evidence.