DDR · Module 31
DDR vs HBM
Per signal wire DDR carries more bandwidth; HBM wins because it reaches a surface that can host far more wires. The deciding quantity is connections per unit of capability, and a perimeter's ratio falls as one over the side length.
Start with the number that contradicts the expectation, because it fixes what this comparison is actually about.
DERIVED below, and recomputed in §5: per signal wire, a DDR channel carries more bandwidth than an HBM stack does. Not less. HBM does not move more data per connection — it moves slightly less, and it wins anyway, decisively, for a reason that is not about memory at all.
The deciding quantity is not bandwidth per connection. It is how many connections the chosen surface can host per unit of the capability being fed — and a package perimeter's ratio falls as one over the side length, monotonically, from the first millimetre.
So this is the chapter where axis A4 moves further than any axis in the module, and the reason is geometry. CURRICULUM-DERIVED from 26.1 §1: a die's capability scales with its area L² and its edge connections with its perimeter 4L, so connections per unit of capability is 4/L — and the usual answer to wanting more of something, build it bigger, is exactly the wrong move.
Axis A1 is identical for the second chapter running. The same destructive 1T1C cell, the same thirteen obligations of 31.1 §5. So once again no structural argument separates the two, and once again the decision is entirely stage two — except that here stage two has a feasibility half before its cost half, which is the shape 31.1 §1 gave to stage one.
1. The Comparison That Is Not Available
Both technologies are DRAM with the same cell, so the usual differentiators are unavailable and it is worth saying which.
| Claim you might reach for | Why it is not available here |
|---|---|
| “HBM has lower latency” | the same destructive read, restore, row buffer and timing classes — 31.1 §4's chain runs identically |
| “HBM has a simpler controller” | §7: the obligations are multiplied, not reduced |
| “HBM is faster per wire” | §5 computes it and the result goes the other way |
| “HBM is the newer technology” | not an engineering argument, and both continue to advance |
| “HBM gives more capacity” | CURRICULUM-DERIVED from 26.4 §1: capacity and bandwidth are two separate purchases, and conflating them is that chapter's subject |
So what remains is a geometry argument and a cost argument, in that order — and the geometry argument is a feasibility test that can close the question before any cost is considered, exactly as 31.1 §1's admissibility test can.
The two-stage structure, restated for this comparison:
STAGE ONE -- FEASIBILITY (a geometry question)
can the chosen SURFACE host enough connections to carry the
required bandwidth, at a manufacturable pitch, for a die of
the size the capability requires?
if no: the technology is not slow for this application. It is
IMPOSSIBLE for it, and no signalling improvement changes that
-- §4 is why the exponent cannot be argued away.
STAGE TWO -- COST (everything else)
interposer, assembly, yield, repairability, capacity
granularity, and the controller replication of §7.2. What Is the Same
Axis A1 is identical, so 31.2 §2's thirteen-for-thirteen table holds again and is not repeated. The same destructive read, the same restore, the same row-buffer concept, the same four constraint classes of 13.3 with different values.
One shared property deserves emphasis because it is often assumed away. CURRICULUM-DERIVED from 26.1 §5 and 26.1 §6, which own why the two halves of a channel share a command bus and what semi-independent actually means: HBM's channels are not fully independent memories. So a design that models them as N unrelated DDR channels has over-stated their independence, and 26.1 §10's pseudo-channel command arbiter exists precisely because of what is shared.
And one more, which §12's defect turns on. Both technologies transfer in bursts, so both have a minimum useful transfer that spans a contiguous run of addresses. That burst span is what interacts with an interleave granularity, and it is the same in kind for both — which is why the hazard is a configuration hazard rather than a technology property.
3. The Axes, for This Comparison
| Axis | DDR | HBM | Who owns the detail |
|---|---|---|---|
| A1 cell | destructive | destructive | identical — §2 |
| A2 obligations | thirteen, once | thirteen, times the channel count | §7 |
| A3 granularity | burst-shaped | burst-shaped, but finer per channel | 26.1 §4 |
| A4 channel shape | few, wide | many, narrow, semi-independent | 26.1 §3, 26.1 §6 |
| A5 latency composition | the same terms | the same terms, shorter transit | 23.1 |
| A6 bandwidth rung | rungs 1 and 2 | rung 1 moves by an order | 30.8 §2 |
| A7 physical coupling | socketed or soldered, replaceable | on-package, assembled once, not replaceable | §16 |
A4 is the axis that moves, and A2 and A7 are its consequences. That is the chapter's structure: one axis moves for a geometric reason, and two others follow.
A6 deserves its qualification immediately. Rung 1 — peak — moves by an order of magnitude. Rungs 3, 4 and 5 do not follow automatically, and 30.8 §4 owns the consequence: three of the four gaps in the bandwidth ladder lie outside the memory, so a technology that multiplies rung 1 can leave application-observed bandwidth almost unchanged. CURRICULUM-DERIVED from 26.4 §7, which owns the result directly: one requester cannot fill them.
4. Why the Surface Decides, and Why Cleverness Cannot
CURRICULUM-DERIVED from 26.1 §1, whose derivation is the load-bearing one for this whole chapter. A square die of side L has capability scaling as L² and edge connections scaling as 4L, so:
connections per unit of capability = 4L / L² = 4 / LCURRICULUM-DERIVED, with that chapter's ILLUSTRATIVE side lengths and its recomputation: 5 mm gives 0.800, 10 mm gives 0.400, 20 mm gives 0.200. Doubling the side halves the connections available per unit of capability.
Three properties of that expression decide the comparison, and 26.1 §1 owns all three.
It is a gradient, not a threshold. There is no size at which perimeter routing suddenly fails; it degrades continuously from the first millimetre. So there is no pin-count wall cycle to point at, and a design cannot wait for a discrete event before acting.
Building bigger makes it worse. More capability to feed, proportionally fewer edges to feed it through. CURRICULUM-DERIVED, and it is the inversion that makes the wall interesting: the usual response to wanting more of something is the wrong move here.
And no signalling improvement changes the exponent. A better interface multiplies the numerator by a constant. CURRICULUM-DERIVED from 26.1 §1 and 4.8: it cannot turn 4/L into anything that does not fall.
So the class of solution is selected by arithmetic before any engineering is done. CURRICULUM-DERIVED from 26.1 §1's callout: if the binding quantity is perimeter / area, there are exactly two ways out — make the numerator scale like the denominator, or stop using the perimeter — and connecting through the die's face does both at once, because a face is an area, so the available count scales as L² and the ratio becomes a constant instead of falling.
That is the entire deciding quantity, and it is worth stating in the form an engineer can apply:
Ask which surface the connections land on. A perimeter gives a ratio that falls as
4/L. A face gives a constant. Every other difference between these two technologies is a consequence or a cost.
5. Per Connection, the Result Goes the Other Way
Now compute the thing everybody assumes, because the answer is instructive.
HBM side — CURRICULUM-DERIVED, from 26.4 §1, which attributes its figures to 4.8 §2 and 26.2 §8: the interface is 1024 bits wide at every stack height, and a stack delivers 307.2 GB/s.
DERIVED, recomputed:
307.2 GB/s / 1024 signals = 0.300 GB/s per signal
and per-pin rate:
307.2e9 B/s x 8 bits/B / 1024 = 2.40 Gbit/s per signalDDR side — ILLUSTRATIVE and parametric, because this chapter has no verified DDR pin count of its own and §Scope forbids inventing one. Take a channel W bits wide at R transfers per second:
bandwidth = W x R / 8 bytes per second
per signal = (W x R / 8) / W = R / 8 bytes per second
DERIVED, and read it before substituting anything: the per-signal
figure does not depend on W at all. It is the TRANSFER RATE
divided by eight, and nothing else.
ILLUSTRATIVE R = 3200 MT/s:
per signal = 3200e6 / 8 = 0.400 GB/s per signal
per-pin rate = 3.20 Gbit/s per signalDERIVED comparison:
| per signal | per-pin rate | ratio | |
|---|---|---|---|
| HBM stack, CURRICULUM-DERIVED | 0.300 GB/s | 2.40 Gbit/s | — |
DDR channel, ILLUSTRATIVE R = 3200 MT/s | 0.400 GB/s | 3.20 Gbit/s | 1.33× better |
So per wire, the perimeter technology is ahead — by a third, on these figures. And the reason is structural rather than accidental: a signal driven across a board at a socket must clear a harder electrical problem than one driven across an interposer, so it is driven faster precisely because there are so few of it. Module 22 owns the electrical half of that, and 26.2 §2 owns the interposer property — density — that makes the many-slow-wires arrangement possible at all.
Two results follow, and the second is the one to carry.
The bandwidth advantage is entirely a count advantage. DERIVED: 1024 × 0.300 = 307.2 against W × 0.400. For the perimeter technology to match one stack it needs W = 307.2 / 0.400 = 768 signals on one channel — and §4 is the reason a package perimeter cannot host that per unit of the capability being fed. So the comparison is not which is faster per wire but which surface can host 768-plus wires.
And that reframes what a “better interface” can buy. Improving the per-signal rate multiplies one column of the table above. CURRICULUM-DERIVED from 26.1 §1: a constant factor cannot turn 4/L into something that does not fall, so signalling improvements postpone the wall and never remove it — and 31.4 is the chapter about a technology that pushes that column as far as it goes and pays for it elsewhere.
6. The Feasibility Test, Stated
Stage one from §1, as a procedure. It uses 26.2's conversions and adds nothing to them.
INPUTS
B_req required bandwidth
L die side implied by the capability to be fed
pitch manufacturable connection pitch on the chosen surface
r_sig achievable per-signal rate on that surface
f_sig fraction of connections that carry a SIGNAL rather than
power or ground -- 26.2 §6 owns that this fraction is
well below one, and this chapter does not invent it
PERIMETER SURFACE
available = (4 x L / pitch) x f_sig
carried = available x r_sig
FEASIBLE iff carried >= B_req
FACE SURFACE
available = (L^2 / pitch^2) x f_sig
carried = available x r_sig
FEASIBLE iff carried >= B_req
DERIVED, and the whole content is the exponent: the perimeter's
available count is LINEAR in L and the face's is QUADRATIC. So
the face's advantage GROWS with the die, and the technology
becomes more attractive exactly as the problem gets harder.The f_sig term is why a naive count overstates both sides. CURRICULUM-DERIVED from 26.2 §6, which owns the point that not every bump carries a signal and that the split is what determines how many bumps a given interface actually costs. A feasibility estimate that divides an area by a pitch and stops has computed an upper bound on an upper bound.
And 26.2 §5 owns the related correction, which matters for the same reason: the usable bump field is smaller than expected. So stage one should be run with the pessimistic figures, because a feasibility test that passes optimistically and fails in assembly has cost the whole schedule.
The honest limit of this test: it decides whether a surface can carry the traffic. It says nothing about whether the requester can supply it — 26.4 §7's one requester cannot fill them — so a design that clears stage one can still deliver rung-5 bandwidth barely different from what it had. §16 puts that measurement before the packaging decision rather than after it.
7. Obligation Multiplication — the Module's Third Pattern
Three chapters, three patterns, and naming the third completes the set.
| Chapter | What happens to the obligation set |
|---|---|
| 31.1 | deleted — twelve of thirteen vanish, because axis A1 changed |
| 31.2 | added — thirteen become twenty-two, none removed |
| This chapter | multiplied — thirteen become thirteen per channel |
And multiplication is worse than addition in a specific, measurable way: it changes what scales.
CURRICULUM-DERIVED from 26.1 §3 and 26.1 §6, which own the channel and pseudo-channel structure and what semi-independent means. Each channel needs its own row state, its own elapsed counters, its own rolling activate window and its own refresh accounting, because those are per-resource obligations and the resources are distinct.
DERIVED, reusing 31.1 §10's per-bank state figure of 30 bits
(row_open, row_known, row_idx, since_act, since_pre) and an
ILLUSTRATIVE 8 banks per channel:
per channel = 8 banks x 30 bits = 240 bits
1 channel = 240 bits
8 channels = 1,920 bits
16 channels = 3,840 bits
plus, and this is the term that is NOT linear:
a cross-channel distributor and its balance accounting, whose
cost grows with the number of destinations it must choose
among -- and 26.1 §10's pseudo-channel arbiter exists because
the channels are SEMI-independent, so the shared parts need
arbitrating rather than merely replicating.Two consequences worth stating, and the second is a design-review argument.
Verification surface multiplies too. CURRICULUM-DERIVED from 30.9 §6's taxonomy: each replicated instance carries the same hazard population — stale state, convention off-by-one, count-versus-index — so sixteen channels is sixteen chances for each. And a bug in the replicated block appears sixteen times or in one instance only, which is a diagnostic: 19.1 §8's per-lane narrowing applied to channels.
And the distributor becomes the new hard problem, displacing the scheduler. With one wide channel, the interesting logic is the scheduler. With sixteen narrow ones, each scheduler's job is easier — fewer requests to choose among — and the difficult decision moves upstream into which channel a request belongs to. That is the address-decode question, 18.2 owns mapping for performance, and §11's defect lives exactly there.
8. The Interleave Hazard, in Principle
§2 noted that both technologies transfer in bursts spanning a contiguous run of addresses. Now combine that with many channels.
a request's burst covers BURST_SPAN contiguous bytes.
the channel decode picks a channel from some address field.
the field's weight determines the INTERLEAVE GRANULARITY --
how many contiguous bytes land on one channel before the
next channel takes over.
the requirement, stated as an inequality:
INTERLEAVE_GRAN >= BURST_SPAN
if it holds : every burst lies entirely within one channel.
if it fails : a single burst STRADDLES two channels, and one
request becomes two channel transactions that
must be split, tracked and rejoined.STRUCTURAL, and the important part is which configuration hides it: with few wide channels the granularity is naturally large and the inequality holds without anyone checking it. With many narrow channels the granularity is chosen small — deliberately, to spread a stream across channels — and the inequality becomes a real constraint that must be enforced.
So the hazard is a configuration hazard rather than a technology property, which is exactly the shape 31.1 §12 warned about: code correct under one configuration, reused under another without re-deriving the requirement. §11 is that code.
And the failure is not a performance loss. A straddling burst that is not split returns bytes from the wrong channel for part of its span — a data-correctness failure with no error signal, the shape 31.2 §7 identified for the PASR mask and 30.4 §5 for a missed write deadline.
9. Two Surfaces, as a Stack
Read outwards from the sixth layer. The surface is the cause; the multiplication above it is the consequence; and the distributor above that is the new hard problem the consequence creates. The cell, at the bottom, is not the difference for the second chapter in a row — which is the module's recurring lesson about where comparisons actually live.
10. RTL — Decode and Replicate
The parameterised comparative block for this chapter. One NUM_CH parameter takes it from the few-wide configuration to the many-narrow one, and the decode and the replication are in the same module on purpose — because §7's argument is that the replication is cheap and the decode is where the difficulty moved.
// ---------------------------------------------------------------------
// channel_decode_replicate -- the comparative block of §10.
//
// CLASSIFICATION: synthesisable, ILLUSTRATIVE parameter values, and
// CORRECT as written. The intentionally defective block is §11.
//
// WHAT IT IS: an address decoder plus per-channel obligation state,
// parameterised on the channel count. NUM_CH = 1 or 2 is the
// few-wide configuration; 8 or 16 the many-narrow one. §7's
// multiplication is the `for` loop over channels; §8's inequality is
// the elaboration guard.
//
// WHY IT EXISTS HERE: §7 claims the difficult decision MOVES
// UPSTREAM from the scheduler to the distributor as the channel
// count rises. This block is both halves of that claim in one place,
// so the reader can see that the per-channel state is a loop and the
// decode is a design decision.
//
// HOW TO RUN IT: elaborate at NUM_CH = 1 and at NUM_CH = 16 and
// compare state-element counts; then present a request whose burst
// span exceeds the interleave granularity.
// EXPECTED RESULT: the second case is REFUSED AT ELABORATION by the
// guard below, which is the correct place to refuse it -- §8's
// inequality is a static property of the configuration, not a
// runtime condition.
//
// SYNTHESIS: NUM_CH x BANKS_PER_CH per-bank records, plus a decode
// that is combinational and a one-hot channel select.
//
// LIMITATIONS: models the DECODE and the per-channel OBLIGATION
// STATE. It does not model the shared command bus between the two
// halves of a channel -- 26.1 §5 owns why the halves share it and
// 26.1 §10 owns the arbiter that follows, and duplicating that here
// would rebuild another chapter's subject. Balance across channels
// is measured, not enforced: §13.
// ---------------------------------------------------------------------
module channel_decode_replicate #(
parameter int NUM_CH = 16,
parameter int BANKS_PER_CH = 8,
parameter int ADDR_W = 34,
parameter int ROW_W = 15,
// §8's two quantities. BURST_SPAN is the contiguous byte run one
// burst covers; INTERLEAVE_GRAN is how many contiguous bytes land
// on one channel before the next takes over.
parameter int BURST_SPAN = 64,
parameter int INTERLEAVE_GRAN = 256,
// ILLUSTRATIVE timing. 14.1 to 14.3 own the real values.
parameter int TRCD = 14,
parameter int TRP = 14,
parameter int TRAS = 34,
// COUNT, not INDEX: each counter must REPRESENT its threshold, so
// it needs $clog2(threshold + 1) bits. TRAS is the largest.
parameter int CNT_W = $clog2(TRAS + 1),
parameter int CH_W = (NUM_CH > 1) ? $clog2(NUM_CH) : 1,
parameter int BANK_W = $clog2(BANKS_PER_CH),
// The bit position the channel field starts at. DERIVED from the
// granularity rather than chosen, so the two cannot disagree.
parameter int CH_SHIFT = $clog2(INTERLEAVE_GRAN)
)(
input logic clk,
input logic rst_n,
input logic req_valid,
input logic [ADDR_W-1:0] req_addr,
input logic [ROW_W-1:0] req_row,
output logic [CH_W-1:0] sel_ch,
output logic [NUM_CH-1:0] ch_onehot,
output logic straddles,
output logic [NUM_CH-1:0] ch_legal,
output logic [15:0] state_bits_reported
);
initial begin
if (NUM_CH < 1) $fatal(1, "channel_decode_replicate: NUM_CH >= 1");
if (NUM_CH > 1 && (NUM_CH & (NUM_CH - 1)) != 0)
$fatal(1, "channel_decode_replicate: NUM_CH must be a power of two for a bit-field decode");
if (BURST_SPAN < 1) $fatal(1, "channel_decode_replicate: BURST_SPAN >= 1");
if ((INTERLEAVE_GRAN & (INTERLEAVE_GRAN - 1)) != 0)
$fatal(1, "channel_decode_replicate: INTERLEAVE_GRAN must be a power of two");
// §8's inequality, enforced AT ELABORATION. This is the correct
// place for it: the relationship between the burst span and the
// interleave granularity is a static property of the
// configuration, so a design that violates it should not
// elaborate rather than failing a runtime check on the one
// access that happens to straddle.
if (INTERLEAVE_GRAN < BURST_SPAN)
$fatal(1, "channel_decode_replicate: INTERLEAVE_GRAN (%0d) < BURST_SPAN (%0d): a burst would straddle two channels (§8)",
INTERLEAVE_GRAN, BURST_SPAN);
if (CH_SHIFT + CH_W > ADDR_W)
$fatal(1, "channel_decode_replicate: channel field falls outside the address");
end
// ---- The decode. One line of logic, and §7's argument is that
// this line is where the difficulty went.
generate
if (NUM_CH > 1) begin : g_multi
assign sel_ch = req_addr[CH_SHIFT + CH_W - 1 : CH_SHIFT];
end else begin : g_single
assign sel_ch = '0;
end
endgenerate
always_comb begin
ch_onehot = '0;
if (req_valid) ch_onehot[sel_ch] = 1'b1;
end
// ---- The straddle check, retained as a RUNTIME output even though
// the elaboration guard makes it unreachable. It is not
// redundant: it is the observable that proves the guard was
// the right guard, and 30.5 §11's lesson is that a guard which
// another guard always shadows is UNTESTED rather than
// working. §14's cover watches this signal for exactly that
// reason.
logic [ADDR_W-1:0] span_end;
always_comb begin
span_end = req_addr + BURST_SPAN[ADDR_W-1:0] - 1;
straddles = req_valid &&
(span_end[CH_SHIFT + CH_W - 1 : CH_SHIFT] != sel_ch);
end
// ===================================================================
// §7's MULTIPLICATION. The obligations are identical in kind to
// 31.1 §5's thirteen; what changes is that there are NUM_CH of
// every one of them. Written as a loop precisely to make the point
// that replication is CHEAP TO WRITE and expensive to verify.
// ===================================================================
logic row_open [NUM_CH][BANKS_PER_CH];
logic row_known [NUM_CH][BANKS_PER_CH];
logic [ROW_W-1:0] row_idx [NUM_CH][BANKS_PER_CH];
logic [CNT_W-1:0] since_act [NUM_CH][BANKS_PER_CH];
logic [CNT_W-1:0] since_pre [NUM_CH][BANKS_PER_CH];
logic [BANK_W-1:0] req_bank;
assign req_bank = req_addr[CH_SHIFT + CH_W + BANK_W - 1 : CH_SHIFT + CH_W];
always_ff @(posedge clk) begin
if (!rst_n) begin
for (int c = 0; c < NUM_CH; c++)
for (int b = 0; b < BANKS_PER_CH; b++) begin
row_open[c][b] <= 1'b0;
row_known[c][b] <= 1'b0;
row_idx[c][b] <= '0;
since_act[c][b] <= '0;
since_pre[c][b] <= '0;
end
end else begin
for (int c = 0; c < NUM_CH; c++)
for (int b = 0; b < BANKS_PER_CH; b++) begin
if (since_act[c][b] != {CNT_W{1'b1}}) since_act[c][b] <= since_act[c][b] + 1'b1;
if (since_pre[c][b] != {CNT_W{1'b1}}) since_pre[c][b] <= since_pre[c][b] + 1'b1;
end
// Only the selected channel's record advances on a command. The
// per-channel independence of the OBLIGATIONS is real even
// though the channels are only SEMI-independent at the command
// bus (26.1 §6) -- the timing state is per resource and the
// resources are distinct.
if (req_valid && ch_legal[sel_ch]) begin
row_open[sel_ch][req_bank] <= 1'b1;
row_known[sel_ch][req_bank] <= 1'b1;
row_idx[sel_ch][req_bank] <= req_row;
since_act[sel_ch][req_bank] <= '0;
end
end
end
// ---- Per-channel legality. NUM_CH independent copies of the same
// maximum-over-rules test that 13.3 owns.
always_comb begin
for (int c = 0; c < NUM_CH; c++) begin
ch_legal[c] = row_open[c][req_bank]
? (since_act[c][req_bank] >= TRCD[CNT_W-1:0])
: (since_pre[c][req_bank] >= TRP[CNT_W-1:0]);
end
end
// §7's count, reported so the comparison is measured and not
// claimed. 30 bits per bank, from 31.1 §10's figure.
assign state_bits_reported = NUM_CH * BANKS_PER_CH * 30;
endmoduleThe elaboration guard is the block's most important line, and it is worth defending. §8's inequality is a static property of the configuration, so refusing to elaborate is strictly better than checking at runtime — a runtime check fails on the one access that happens to straddle, which is data-dependent and may not occur in a regression at all. CURRICULUM-DERIVED from 30.9 §3's tool table: lint and elaboration prove structural facts, and a structural fact should be proved by the cheapest tool that can prove it.
11. RTL Review — The Channel Distributor
The intended contract:
- Decode each request to exactly one channel.
- A request's burst must lie entirely within the channel it was decoded to — §8's inequality.
- Where the configuration cannot satisfy clause 2, the request must be split into per-channel sub-requests, each with its own length, and rejoined on return.
- The channel select must be one-hot when a request is valid, and zero otherwise.
- The decode must be stable for the duration of a request, so that a return can be attributed to the channel that served it.
// ---------------------------------------------------------------------
// channel_distributor -- INTENTIONALLY DEFECTIVE, for review (§11).
//
// CLASSIFICATION: synthesisable, ILLUSTRATIVE, and CONTAINS A BUG.
//
// WHAT IT IS MEANT TO DO: the five-clause contract above -- spread a
// request stream across NUM_CH channels, splitting any request whose
// burst would straddle a channel boundary.
//
// WHY IT EXISTS HERE: §8 establishes that the interleave-versus-burst
// inequality holds WITHOUT ANYONE CHECKING IT at the few-wide
// configuration and becomes a real constraint at the many-narrow
// one. This block is that configuration dependence, and it is the
// third instance in this module of the same bug class: code correct
// under one parameterisation, reused under another without
// re-deriving the requirement (31.1 §12, 31.2 §12).
//
// HOW TO RUN IT: elaborate with NUM_CH = 16 and INTERLEAVE_GRAN = 32
// against BURST_SPAN = 64, then present an aligned request.
// EXPECTED RESULT under clause 3: the request is SPLIT, and
// `split_req` is asserted.
// EXPECTED TRACE: a burst covering bytes 0 to 63 with a 32-byte
// granularity must produce two sub-requests, to channels 0 and 1.
//
// SYNTHESIS: one combinational decode plus a one-hot expander.
//
// LIMITATIONS: models the DECODE only. Return-path rejoining is the
// other half of clause 3 and is out of scope here -- 10.5 §9 owns
// the association problem -- and that omission is STATED rather than
// hidden. It is not what the bug is.
// ---------------------------------------------------------------------
module channel_distributor #(
parameter int NUM_CH = 16,
parameter int ADDR_W = 34,
parameter int BURST_SPAN = 64,
parameter int INTERLEAVE_GRAN = 32,
parameter int CH_W = (NUM_CH > 1) ? $clog2(NUM_CH) : 1,
parameter int CH_SHIFT = $clog2(INTERLEAVE_GRAN)
)(
input logic clk,
input logic rst_n,
input logic req_valid,
input logic [ADDR_W-1:0] req_addr,
output logic [CH_W-1:0] sel_ch,
output logic [NUM_CH-1:0] ch_onehot,
output logic split_req,
output logic [7:0] sub_req_count
);
initial begin
if (NUM_CH < 1) $fatal(1, "channel_distributor: NUM_CH >= 1");
if ((INTERLEAVE_GRAN & (INTERLEAVE_GRAN - 1)) != 0)
$fatal(1, "channel_distributor: INTERLEAVE_GRAN must be a power of two");
if (CH_SHIFT + CH_W > ADDR_W)
$fatal(1, "channel_distributor: channel field falls outside the address");
// NOTE: §10's guard on INTERLEAVE_GRAN >= BURST_SPAN is ABSENT
// here, and deliberately so -- this block is supposed to SUPPORT
// the straddling configuration by splitting (clause 3), so
// refusing to elaborate would defeat its purpose. §12 is about
// what was removed along with it.
end
// Clause 1 and 4: decode to one channel, one-hot.
assign sel_ch = (NUM_CH > 1) ? req_addr[CH_SHIFT + CH_W - 1 : CH_SHIFT]
: '0;
always_comb begin
ch_onehot = '0;
if (req_valid) ch_onehot[sel_ch] = 1'b1;
end
// Clause 3: decide whether a split is needed.
assign split_req = 1'b0; // <-- THE DEFECT
assign sub_req_count = req_valid ? 8'd1 : 8'd0;
endmoduleBefore reading on: which clause, and why did dropping §10's elaboration guard look like the right move?
12. The Defect — The Guard Was Load-Bearing
The violated clause is 3, and the defect is that split_req is tied low: the block never splits, whatever the configuration.
And the reason it is interesting is the comment that removes the guard. §10 refuses to elaborate when INTERLEAVE_GRAN < BURST_SPAN. This block deliberately drops that guard, with a defensible justification — this block is supposed to support the straddling configuration by splitting — and then does not implement the splitting.
ILLUSTRATIVE NUM_CH = 16, INTERLEAVE_GRAN = 32, BURST_SPAN = 64.
CH_SHIFT = 5, CH_W = 4, so the channel field is addr[8:5].
a request at address 0, burst covering bytes 0..63:
byte 0..31 -> addr[8:5] = 0 -> channel 0
byte 32..63 -> addr[8:5] = 1 -> channel 1
contract clause 3 : SPLIT into two sub-requests
this block : sel_ch = 0, sub_req_count = 1, split_req = 0
so channel 0 is asked for 64 bytes. It has 32. It returns 64 --
the second 32 coming from whatever is at the next address WITHIN
channel 0, which is byte offset 512 of the interleaved space.
the requester receives 64 contiguous-looking bytes of which the
second half belongs to a different address entirely.Three properties of this failure make it the worst in the chapter.
It is a data-correctness failure with no error signal. Nothing is illegal; no timing rule is violated; the channel served a perfectly legal request for data it does not hold. Chapter 31.2 §7 identified this shape for the PASR mask and 30.4 §5 for a missed write deadline — an obligation whose violation returns something rather than an error.
It is configuration-silent. At INTERLEAVE_GRAN = 256 and BURST_SPAN = 64 the inequality holds, no burst ever straddles, and split_req tied low is correct. So the block is right in the few-wide configuration and in any many-channel configuration with a coarse granularity — and it becomes wrong only when someone reduces the granularity to spread a stream more finely, which is a performance tuning change that nobody expects to affect correctness.
And the guard that would have caught it was removed for a good reason. That is the review lesson, and it generalises:
A guard removed because a block is supposed to handle the case it guarded against creates an obligation to implement the handling. Removing the guard and not implementing the handling leaves the design strictly worse than before the guard existed — because the configuration is now reachable and unprotected.
So the correct review question is not why is this guard missing but what obligation did its removal create, and where is that obligation discharged? The answer here is nowhere, and the commit that removed the guard is where to look.
The correction, and it needs both the split and the count:
// CORRECTED. Clause 3: compute the channel of the burst's LAST
// byte as well as its first, and split when they differ. Note that
// the general case can span MORE than two channels when the burst
// span exceeds twice the granularity, so the count is derived
// rather than fixed at two -- a two-way split would be a second
// bug of the same family, right for one ratio and wrong for
// another.
logic [ADDR_W-1:0] span_end;
logic [CH_W-1:0] end_ch;
always_comb begin
span_end = req_addr + BURST_SPAN[ADDR_W-1:0] - 1;
end_ch = span_end[CH_SHIFT + CH_W - 1 : CH_SHIFT];
split_req = req_valid && (end_ch != sel_ch);
// The number of channels the burst touches. DERIVED from the
// aligned span rather than from the channel indices, because the
// indices WRAP and a difference computed on them is wrong at the
// wrap point -- which would be a third bug of the same family.
sub_req_count = req_valid
? 8'( ((req_addr[CH_SHIFT-1:0] + BURST_SPAN - 1) >> CH_SHIFT) + 1 )
: 8'd0;
endAnd the structural recommendation, which is stronger than the fix: keep §10's elaboration guard and add a second parameter that explicitly enables splitting. A configuration that needs splitting then has to say so, the guard fires for every configuration that did not ask, and the obligation created by disabling the guard is visible in the parameter list rather than in a comment. CURRICULUM-DERIVED from 30.9 §3: a structural fact should be proved by the cheapest tool that can prove it, and making the dangerous configuration opt-in keeps elaboration as that tool.
13. RTL — Measuring the Multiplication
§7 claims the difficulty moves from the scheduler to the distributor. This block measures whether it did, by reporting the two quantities that distinguish a working distribution from a broken one: balance and per-channel stall attribution.
// ---------------------------------------------------------------------
// channel_balance_monitor -- verification/telemetry, CORRECT as written.
//
// CLASSIFICATION: synthesisable telemetry. Drives nothing.
//
// WHAT IT DOES: counts requests per channel over a window and
// reports the imbalance, plus how many cycles each channel was
// blocked while another had work available.
//
// WHY IT EXISTS HERE: §7 argues the distributor is the new hard
// problem, and §11's defect is one way it goes wrong. This block
// catches the OTHER way, which produces no wrong data at all: a
// decode that is correct per request and systematically unbalanced
// across the stream, so the many-channel memory delivers the
// bandwidth of a few-channel one. 30.8's ladder names that exactly
// -- rung 1 moved and rung 4 did not.
//
// HOW TO RUN IT: run the real access pattern and read `imbalance`.
// EXPECTED RESULT: a stride that is a multiple of NUM_CH x
// INTERLEAVE_GRAN concentrates on one channel and imbalance
// approaches its maximum -- 18.2's mapping-for-performance problem
// at channel granularity.
//
// SYNTHESIS: NUM_CH counters plus a max/min tracker. No divider.
//
// LIMITATIONS: measures DISTRIBUTION, not correctness. §11's defect
// produces a perfectly balanced stream of wrong data, so this block
// cannot see it -- which is why §14 asserts correctness separately
// and this block's LIMITATIONS say so rather than implying coverage
// it does not give (27.3's independence discipline).
// ---------------------------------------------------------------------
module channel_balance_monitor #(
parameter int NUM_CH = 16,
parameter int WIN = 65536,
// COUNT, not INDEX: each per-channel counter must be able to
// represent a window in which EVERY request went to one channel,
// so it needs $clog2(WIN + 1) bits. Sized $clog2(WIN / NUM_CH) --
// the "fair share" -- it would saturate in exactly the imbalanced
// case it exists to detect, and report perfect balance.
parameter int CNT_W = $clog2(WIN + 1)
)(
input logic clk,
input logic rst_n,
input logic req_fire,
input logic [$clog2(NUM_CH > 1 ? NUM_CH : 2)-1:0] req_ch,
input logic [NUM_CH-1:0] ch_has_work,
input logic [NUM_CH-1:0] ch_blocked,
input logic win_tick,
output logic [CNT_W-1:0] r_max_ch,
output logic [CNT_W-1:0] r_min_ch,
output logic [CNT_W-1:0] r_imbalance,
output logic [CNT_W-1:0] r_starved_cycles,
output logic result_valid
);
initial begin
if (NUM_CH < 1) $fatal(1, "channel_balance_monitor: NUM_CH >= 1");
if (WIN < NUM_CH)
$fatal(1, "channel_balance_monitor: WIN must be >= NUM_CH or balance is meaningless");
end
logic [CNT_W-1:0] cnt [NUM_CH];
logic [CNT_W-1:0] starved;
always_ff @(posedge clk) begin
if (!rst_n) begin
for (int c = 0; c < NUM_CH; c++) cnt[c] <= '0;
starved <= '0;
r_max_ch <= '0;
// AND-accumulator discipline: a minimum tracker must start at
// all-ones or the first sample never wins and the reported
// minimum stays zero forever -- which would report maximum
// imbalance on every window regardless of the traffic.
r_min_ch <= '1;
r_imbalance <= '0;
r_starved_cycles <= '0;
result_valid <= 1'b0;
end else if (win_tick) begin
begin
logic [CNT_W-1:0] mx, mn;
mx = '0;
mn = '1;
for (int c = 0; c < NUM_CH; c++) begin
if (cnt[c] > mx) mx = cnt[c];
if (cnt[c] < mn) mn = cnt[c];
end
r_max_ch <= mx;
r_min_ch <= mn;
// Reported as a DIFFERENCE with both endpoints, not as a
// ratio. 27.5 refuses a single percentage and 30.8 §11
// requires a derived statistic's range to be asserted; a
// difference with its endpoints lets the consumer see the
// denominator and own the rounding.
r_imbalance <= mx - mn;
end
r_starved_cycles <= starved;
result_valid <= 1'b1;
for (int c = 0; c < NUM_CH; c++) cnt[c] <= '0;
starved <= '0;
end else begin
result_valid <= 1'b0;
if (req_fire && cnt[req_ch] != {CNT_W{1'b1}})
cnt[req_ch] <= cnt[req_ch] + 1'b1;
// A starved cycle: some channel had work and was blocked while
// at least one other channel was idle with nothing to do. That
// is the distributor's failure signature rather than the
// scheduler's -- 30.5 §3's four-stage attribution, applied
// across channels instead of within one.
if (((ch_has_work & ch_blocked) != '0)
&& ((~ch_has_work) != '0)
&& starved != {CNT_W{1'b1}})
starved <= starved + 1'b1;
end
end
endmoduler_starved_cycles is the measurement that distinguishes the two failures §7 predicts. A high imbalance with low starvation means the pattern concentrates — 18.2's mapping problem at channel granularity. A low imbalance with high starvation means the distribution is fine and something downstream is blocking, which sends the investigation back to the per-channel schedulers. Two opposite conclusions from one pair of counters, and 30.8 §13 owns that shape of instrument.
14. SVA Review — Correct Per Request, Wrong Across the Stream
The property written for §11's block, and it is a real property that passes:
// Offered as "proves the distributor decodes correctly".
// Clause 1 and 4. Correct, non-vacuous, and PASSES on the
// defective block.
property p_channel_select_is_onehot;
@(posedge clk) disable iff (!rst_n)
req_valid |-> $onehot(ch_onehot);
endproperty
assert property (p_channel_select_is_onehot)
else $error("channel select was not one-hot for a valid request");Q. It passes. What does it prove?
That the decode picks exactly one channel — which is precisely what the defect does. The defective block's failure is not that it picks the wrong number of channels; it is that picking one channel is the wrong answer for this request, and one-hotness cannot express that.
This is 30.9 §6's variety 2 — the property does not name the contract's key quantity. The key quantity here is not ch_onehot at all; it is the burst's span, which appears nowhere in the property and nowhere in the defective block. A property whose signal list does not include the burst length cannot constrain a claim about the burst's extent.
And it is the module's third instance of variety 10 as well, because at the coarse-granularity configuration this property plus split_req tied low together constitute a correct design — so the property is sound there and, at the fine granularity, certifies the absence of splitting as the intended behaviour.
What actually covers clauses 2 and 3:
// Clause 2. It names BURST_SPAN, which is the quantity the
// defective block never computes. Variety 2's repair: make the
// property mention the thing whose corruption it must catch.
property p_burst_lies_within_one_channel;
@(posedge clk) disable iff (!rst_n)
(req_valid && !split_req) |->
(((req_addr + BURST_SPAN - 1) >> CH_SHIFT) ==
(req_addr >> CH_SHIFT));
endproperty
assert property (p_burst_lies_within_one_channel)
else $error("an unsplit burst crosses a channel boundary");
// Clause 3, the other direction: when the span DOES cross, a split
// must be signalled. Stating both directions separately matters --
// a design that always splits satisfies the property above and
// wastes half the channel bandwidth, and a design that never
// splits satisfies this one vacuously.
property p_crossing_burst_is_split;
@(posedge clk) disable iff (!rst_n)
(req_valid &&
(((req_addr + BURST_SPAN - 1) >> CH_SHIFT) != (req_addr >> CH_SHIFT)))
|-> split_req;
endproperty
assert property (p_crossing_burst_is_split)
else $error("a crossing burst was not split");
// Clause 3's arithmetic. The sub-request count must equal the
// number of channels the span actually touches -- a COMPARATIVE
// obligation, so it is checked against an independently computed
// expectation rather than against the design's own derivation.
// 30.6 §11: "best" and "correct count" obligations need a model,
// and 27.3's independence requirement says the model must not
// reuse the design's expression.
property p_sub_req_count_matches_span;
@(posedge clk) disable iff (!rst_n)
req_valid |-> (sub_req_count == ref_channels_touched(req_addr, BURST_SPAN));
endproperty
assert property (p_sub_req_count_matches_span)
else $error("sub-request count disagrees with the reference model");
// Clause 5: the decode must be stable while a request is in
// flight, or a return cannot be attributed. 10.5 §9's association
// problem, at channel granularity.
property p_decode_stable_in_flight;
@(posedge clk) disable iff (!rst_n)
(req_valid && !req_fire) |=> $stable(sel_ch);
endproperty
assert property (p_decode_stable_in_flight)
else $error("channel decode changed while a request was pending");
// The elaboration-time fact, restated as a runtime invariant for
// the block that KEEPS the guard (§10). It must be unreachable
// there, and 30.5 §11's lesson applies: a guard another guard
// always shadows is untested, so the cover below must be watched.
property p_no_straddle_when_guarded;
@(posedge clk) disable iff (!rst_n)
(INTERLEAVE_GRAN >= BURST_SPAN) -> !straddles;
endproperty
assert property (p_no_straddle_when_guarded)
else $error("a straddle occurred in a configuration that forbids it");
// ---- Covers. The configuration covers are mandatory here for the
// same reason as in 31.1 §14 and 31.2 §14: the defect lives
// in a configuration, so coverage must be per configuration.
// The GRANULARITY regimes -- the two sides of §8's inequality.
cover property (@(posedge clk) disable iff (!rst_n)
req_valid && (INTERLEAVE_GRAN >= BURST_SPAN));
cover property (@(posedge clk) disable iff (!rst_n)
req_valid && (INTERLEAVE_GRAN < BURST_SPAN));
// An actual crossing burst, which is the antecedent of
// p_crossing_burst_is_split. A zero here means the fine-grained
// configuration ran and never presented the case -- which random
// aligned traffic will do, because aligned bursts do not cross.
cover property (@(posedge clk) disable iff (!rst_n)
req_valid &&
(((req_addr + BURST_SPAN - 1) >> CH_SHIFT) != (req_addr >> CH_SHIFT)));
// A burst spanning MORE than two channels, which is where the
// naive two-way split in §12's discussion would itself fail.
cover property (@(posedge clk) disable iff (!rst_n)
req_valid && (sub_req_count > 8'd2));
// §10's shadowed guard, watched. This cover must go DEAD in the
// guarded configuration and be LIVE in the unguarded one -- two
// different instruments, and 30.10 §12's point that knowing which
// kind you are writing is the senior move.
cover property (@(posedge clk) disable iff (!rst_n) straddles);
// The channel-concentration case, so §13's imbalance measurement
// is known to have been exercised at its extreme.
cover property (@(posedge clk) disable iff (!rst_n)
result_valid && (r_min_ch == '0) && (r_max_ch != '0));The third cover is the one that matters most, and it is not obvious. Aligned bursts do not cross channel boundaries. So a random-address regression on the fine-grained configuration will exercise crossings, but a regression using naturally aligned traffic — which is most realistic traffic — will not, and p_crossing_burst_is_split passes vacuously. CURRICULUM-DERIVED from 27.2 §6: the fix for a vacuous property is not a better property but a cover on the antecedent, and here the antecedent requires deliberately misaligned stimulus.
Follow-up an interviewer should ask: should the system-level environment generate misaligned requests? Usually not — if the upstream interconnect guarantees alignment, injecting misalignment tests a case the system cannot produce. At the block level, yes, because this block's contract is about spans and not about what the interconnect promises. Naming that boundary is the answer, and it is the same distinction 30.4 §9 draws about violating column spacing at the block level but never at the system level.
15. What the Replicated Block's Assertions Prove
§14 reviewed the distributor and left §10's replicated block unasserted, and that omission is worse here than it was in 31.1 §14 — because replication multiplies the hazard population. CURRICULUM-DERIVED from 30.9 §6: each instance carries the same stale-state, convention and count-versus-index hazards, so sixteen channels is sixteen chances for each, and the bug that appears in one instance only is a different diagnosis from the bug that appears in all sixteen.
The property that matters most in a replicated design is the one nobody writes for a single instance, because for a single instance it is trivially true.
// ---- ISOLATION. The property that only exists once a design is
// replicated, and the one that catches the classic
// replication bug: an index computed from the wrong variable.
// A command to one channel must not disturb any other
// channel's obligation state. For NUM_CH = 1 this is
// vacuously true, which is exactly why a design scaled up
// from one channel has never had it checked.
property p_channel_state_isolation;
@(posedge clk) disable iff (!rst_n)
(req_valid && ch_legal[sel_ch]) |=>
(foreach_other_channel_stable(sel_ch));
endproperty
assert property (p_channel_state_isolation)
else $error("a command to one channel disturbed another channel's state");
// Written concretely for a bounded channel count, because the
// helper above hides the thing worth seeing -- that the check is
// over every OTHER index, and that "every other" is what the
// buggy version gets wrong.
generate
for (genvar c = 0; c < NUM_CH; c++) begin : g_iso
property p_other_channel_untouched;
@(posedge clk) disable iff (!rst_n)
(req_valid && ch_legal[sel_ch] && (sel_ch != c)) |=>
($stable(row_open[c]) && $stable(row_idx[c]) && $stable(row_known[c]));
endproperty
assert property (p_other_channel_untouched)
else $error("channel %0d's state changed on a command to another channel", c);
end
endgenerate
// ---- Per-channel legality must be computed from that channel's
// OWN state. A single shared counter feeding every channel's
// test would pass every timing property and be wrong for
// fifteen of sixteen channels -- and it would pass because
// the timing arithmetic is right, which is 30.5 §11's
// "shape versus content" distinction at replication scale.
generate
for (genvar c = 0; c < NUM_CH; c++) begin : g_leg
property p_legality_uses_own_state;
@(posedge clk) disable iff (!rst_n)
ch_legal[c] |-> (row_open[c][req_bank]
? (since_act[c][req_bank] >= TRCD)
: (since_pre[c][req_bank] >= TRP));
endproperty
assert property (p_legality_uses_own_state)
else $error("channel %0d's legality was computed from foreign state", c);
end
endgenerate
// ---- The elaboration guard's runtime shadow. §10 keeps the guard,
// so `straddles` must be unreachable there. This is the
// assertion form of 30.5 §11's warning: a guard that another
// guard always shadows is UNTESTED, so the claim that it is
// unreachable should be asserted rather than assumed.
property p_straddle_unreachable_when_guarded;
@(posedge clk) disable iff (!rst_n)
!straddles;
endproperty
assert property (p_straddle_unreachable_when_guarded)
else $error("a straddle occurred despite the elaboration guard");
// ---- The reported state count must match the configuration, so
// that §7's measured multiplication cannot silently disagree
// with the parameters. An INVARIANT, and the cheapest
// possible check that the configuration which elaborated is
// the one intended -- 31.1 §13's argument, reused.
property p_state_count_matches_config;
@(posedge clk) disable iff (!rst_n)
state_bits_reported == (NUM_CH * BANKS_PER_CH * 30);
endproperty
assert property (p_state_count_matches_config)
else $error("reported state bits disagree with the channel configuration");
// ---- Covers. The replication regimes, because a property proved
// at one channel count says nothing about another -- variety
// 10 again, with the channel count as the parameter.
cover property (@(posedge clk) disable iff (!rst_n)
req_valid && (NUM_CH == 1));
cover property (@(posedge clk) disable iff (!rst_n)
req_valid && (NUM_CH > 1));
// Every channel must have been SELECTED at least once, or the
// isolation properties above held only for the indices the
// stimulus happened to reach. This is the replication analogue of
// 31.2 §14's duration cover: the coverage item must be on the
// dimension the defect scales with, and here that dimension is
// the index space.
generate
for (genvar c = 0; c < NUM_CH; c++) begin : g_cov
cover property (@(posedge clk) disable iff (!rst_n)
req_valid && (sel_ch == c));
end
endgeneratep_channel_state_isolation is the section's deliverable, and the reason is a general one about scaling a design up.
A replicated design needs a class of property that a single instance does not: that operating on one instance leaves the others alone. For one instance that property is vacuously true, so a design grown from a single-channel predecessor has never had it checked — and the index arithmetic that broke is new code.
And the per-channel cover generate loop is the third coverage-reachability problem this module has produced. Chapter 31.1 §14 needed a cover per configuration; 31.2 §14 needed one per duration; this needs one per index. The common rule is worth stating once:
The coverage item must be on the dimension the defect scales with. A defect that scales with a configuration needs a configuration cover, one that scales with a duration needs a duration cover, and one that scales with an index space needs a cover per index.
Follow-up an interviewer should ask: is a per-index cover affordable at sixteen channels? Yes, and it is the cheap case — sixteen cover points is nothing, and the stimulus to reach them is a strided address sweep. The expensive version is the one that must be argued down: at 1024 banks across 16 channels, a cover per bank per channel is 16,384 points, and CURRICULUM-DERIVED from 27.5 §4 — which owns the reachability arithmetic and the finding that most of a naive cross is unhittable — that cross should be reduced deliberately rather than declared closed. Covering the index space of the thing that is replicated, and not the cross of everything inside it, is the distinction.
16. The Capacity Decision You Have Not Made Yet
Axis A7 is where this comparison becomes irreversible, and it interacts with a decision most projects take later.
CURRICULUM-DERIVED from 26.4 §1, which owns the two purchases and proves their independence: the interface is the same width at every stack height, so height changes capacity and leaves bandwidth alone. A die adds capacity only; a stack adds bandwidth, streams and capacity together. So one purchase is strictly more powerful and strictly more expensive.
Now add the packaging consequence, which is this chapter's contribution:
| Decision | DDR, socketed or soldered | HBM, on-package |
|---|---|---|
| Change capacity after design | possible — a different module or population | fixed at assembly |
| Change capacity after shipping | possible for socketed | impossible |
| Replace a failed device | possible for socketed | the package is scrap |
| Probe the interface during bring-up | possible at the socket or the board | no access — 30.10 §2's observability collapse, worsened |
| Change bandwidth | add channels if the perimeter allows — §6 | add a stack, which is a new package |
Two results follow, and they are decision inputs rather than trivia.
The capacity decision is pulled forward into the packaging decision. A project choosing the on-package technology has committed to a capacity range at the point it commits to a stack count, because 26.4 §1's two purchases are made simultaneously in one assembly. So a capacity requirement that is still uncertain must be resolved earlier than it otherwise would be, and that schedule consequence belongs in the comparison.
And the repairability loss compounds a cost this module has otherwise ignored. CURRICULUM-DERIVED from 26.3 §4 and 26.3 §6, which own why a stack needs spares and that repair is per group rather than per via: the technology has internal repair precisely because assembly yield on a stack is a real risk. That repair happens at manufacture. There is none in the field, and a design whose service model assumed module replacement has to change it.
The honest counterweight, because §Scope forbids a winner. For a system whose capability genuinely requires the bandwidth, §6's feasibility test has already closed the question — the perimeter cannot carry it, so replaceability was never on offer. Axis A7 is a cost to be acknowledged, not a tiebreaker to be applied after feasibility has already decided.
17. What Would You Measure?
Q. You must decide between these technologies for a new accelerator. What do you measure, in what order?
| Measurement | What it settles | Cost | Owner |
|---|---|---|---|
| Required bandwidth against §6's perimeter feasibility | whether the question is a choice at all | hours | §6, 26.1 §1 |
| The requester's concurrency — how many streams it can sustain | whether a many-channel memory can be filled | days | 26.4 §7 |
| Arithmetic intensity of the workload | whether the compute rate is servable at all | days | 26.4 §5 |
Stride against NUM_CH × INTERLEAVE_GRAN | channel concentration — §13's imbalance | days | 18.2, §13 |
r_imbalance and r_starved_cycles together | distributor problem versus downstream problem | days | §13 |
| Is the capacity requirement settled? | §16 — the on-package choice fixes it at assembly | hours | 26.4 §1 |
| The service model — is field replacement assumed? | §16, and it can close the question | hours | §16 |
| The full five-rung ladder on a prototype | whether rung 1 moving moves rung 5 | weeks | 30.8 §3 |
Row one is first because it is the only row that can prove the question is not a choice. If the perimeter cannot carry the required bandwidth for a die of the necessary size, the comparison is over and everything below row one is a cost analysis of the only feasible option.
Row two is the row that reverses the conclusion most often. CURRICULUM-DERIVED from 26.4 §7: one requester cannot fill them. A design that clears feasibility and cannot supply the concurrency has bought rung 1 and will measure rung 5 unchanged — which is 30.8 §4's result that three of the four gaps lie outside the memory, arriving as a packaging decision.
And rows six and seven cost hours and are usually asked last. A settled capacity requirement and a service model are answerable in an afternoon, and either can close the question — so asking them after a week of bandwidth modelling is the ordering error 30.8 §1 warns about.
18. Common Wrong Answers
“HBM has more bandwidth per pin.” §5. DERIVED: 0.300 GB/s per signal against an illustrative 0.400 for DDR. Per wire the perimeter technology is ahead, and the advantage is entirely a count advantage.
“HBM is faster, so it has lower latency.” §1, §2. Identical cell, identical restore, identical constraint classes. Transit is shorter; the access is not.
“The pin-count wall is a threshold we have not reached.” §4, and 26.1 §1 owns the correction: it is a gradient that degrades continuously from the first millimetre. There is no event to wait for.
“A better interface postpones the wall indefinitely.” §4, §5. A constant factor cannot turn 4/L into something that does not fall. It postpones and never removes.
“Make the die bigger to fit more channels.” §4. A larger die has more capability to feed and proportionally fewer edges to feed it through — the inversion that makes the wall interesting.
“Divide the bump area by the pitch to get the connection count.” §6. That is an upper bound on an upper bound: 26.2 §5 owns that the usable field is smaller than expected, and 26.2 §6 that not every bump carries a signal.
“HBM channels are independent memories.” §2, and 26.1 §6 owns what semi-independent means: the two halves share a command bus, which is why a pseudo-channel arbiter exists.
“More channels means a simpler controller because each one is smaller.” §7. Each scheduler's job is easier and there are sixteen of them, plus a distributor that did not exist before. Thirteen obligations become thirteen per channel.
“HBM gives you more capacity.” §16, and 26.4 §1 owns the correction: capacity and bandwidth are two separate purchases, and a die adds only the first.
“More bandwidth will make the application faster.” §3, §16. Rung 1 moving does not move rung 5 — 30.8 §4 — and 26.4 §7 owns the reason: one requester cannot fill them.
“Interleave as finely as possible to spread the traffic.” §8. Below the burst span, a single burst straddles two channels — and if it is not split, the data is wrong with no error signal.
“The straddle case never happens with aligned traffic.” §14. Correct, and that is exactly why the property passes vacuously. Aligned bursts do not cross, so the antecedent needs deliberately misaligned block-level stimulus.
“The decode is one-hot, so the distributor is correct.” §14. One-hotness is what the defect does. The property never names the burst span, so it cannot constrain a claim about the burst's extent.
“The guard was unnecessary because this block handles the case.” §12. Removing a guard because a block is supposed to handle the case creates an obligation to implement it. Here it was never discharged, which leaves the design worse than before the guard existed.
“Splitting into two sub-requests handles it.” §12. Only when the span crosses one boundary. A burst spanning more than twice the granularity touches more than two channels, and a fixed two-way split is the same bug family again.
“The channels are balanced, so the distribution is working.” §13, §14. Balance and correctness are independent: §11's defect produces a perfectly balanced stream of wrong data, and §13's block says so in its own limitations.
“We can add capacity later.” §16. On-package capacity is fixed at assembly, so the capacity decision is pulled forward into the packaging decision.
“A failed device gets replaced.” §16. The stack's repair — 26.3 §6's per-group spares — happens at manufacture. There is none in the field, and the package is scrap.
19. Self-Check
-
Compute bandwidth per signal for both technologies from §5's inputs, showing the DDR side's independence from channel width. State which is ahead and by how much.
-
Given that result, explain in two sentences where HBM's bandwidth advantage actually comes from.
-
Write
4/Lfrom memory and state its three properties from §4. For each, say what design conclusion it forbids. -
Explain why connecting through a face changes the exponent rather than the constant, and why that distinction decides the class of solution before any engineering.
-
Run §6's feasibility test symbolically for both surfaces. Identify the term that makes a naive area-divided-by-pitch estimate an upper bound on an upper bound.
-
Name the module's three obligation patterns, one per chapter, and say which is worst for verification surface and why.
-
State §8's inequality. Then explain why it holds unchecked in one configuration and becomes a real constraint in the other, and what kind of change makes a working design violate it.
-
Find the defect in §11 without reading §12. Then answer the harder question: what obligation did removing §10's guard create, and where should it have been discharged?
-
Explain why
p_channel_select_is_onehotpasses on the defective block, and why no strengthening of a one-hotness property reaches the contract. -
Explain why
p_crossing_burst_is_splitpasses vacuously under realistic traffic, and state where the misaligned stimulus belongs and where it does not. -
Explain why the isolation property of §15 is vacuously true for a single instance, and why that makes a design scaled up from one channel especially exposed. Then state the general rule about which dimension a coverage item must be on.
-
Give two of §16's irreversibility consequences and, for each, the project decision it pulls forward or forecloses.
20. Where This Goes
Per wire, the perimeter technology is ahead; the on-package technology wins because it reaches a surface that can host far more wires. The deciding quantity is connections per unit of capability, a perimeter's ratio falls as 4/L monotonically from the first millimetre, no signalling improvement changes that exponent, and the feasibility test can close the question before any cost is considered. Axis A1 is identical for the second chapter running; the obligation set is multiplied rather than changed; the difficult decision moves from the scheduler to the distributor; and axis A7 pulls the capacity decision forward into the packaging decision and forecloses field replacement.
Three results carry forward. A guard removed because a block is supposed to handle the guarded case creates an obligation, and the commit that removed it is where to look for the missing handling. Balance and correctness are independent measurements, and a perfectly balanced stream of wrong data is a real failure mode. And the antecedent of a span-crossing property requires deliberately misaligned stimulus, because realistic aligned traffic makes it vacuous — the third distinct coverage-reachability problem this module has produced, after 31.1 §14's configuration covers and 31.2 §14's duration covers.
Chapter 31.4 closes the module with the axis this chapter deliberately left alone. §5 showed that per-signal rate is one column of the table and that improving it postpones the wall without removing it. The final comparison is with a technology that pushes that column as far as a soldered point-to-point channel allows — and the interesting part is not the rate. It is what becomes affordable to give up once the consumer stops caring about latency, and what must be added back once the per-pin rate rises far enough that the channel stops being reliable on its own.
Continue learning
Related tutorials
- Related topic
HBM Overview
HBM reaches hundreds of GB/s with a per-pin rate lower than DDR5's. It wins on width, not speed — and getting that width required changing the packaging, which adds a fourth design layer to array physics, device architecture and the interface.
- Related topic
Channels
A channel has its own command bus, data bus and scheduling, so two channels contend for nothing. That independence means the address alone decides which channel a request uses, making the mapping — not the hardware — the thing that determines whether the parallelism is real.
- Related topic
Why HBM Exists
Connections scale with perimeter while capability scales with area, and the ratio falls as 4/L. The width that answers it divides into sixteen streams that preserve the 64-byte granule exactly.
- Related topic
2.5D Packaging for HBM
A bump pitch turns an area into a count, and the face-versus-edge advantage has the closed form L/4p. The surprise is that a full 1024-bit bump field uses about 2% of a die's face.
Standards & specifications
- Governing standard
- JEDEC JESD79 (DDR SDRAM)(opens JEDEC Solid State Technology Association in a new tab)
Defines the DDR SDRAM device itself — signals, command encoding, mode registers, timing parameters and the initialisation sequence — one document per generation. Memory-controller microarchitecture, address-mapping policy, PHY training algorithms and board-level design are not specified by it.
This page also covers RTL structure, verification approach and debugging technique. Those are engineering practice built on the standard, not requirements the standard itself imposes.
Where this fits
Part of the DDR curriculum.
