Skip to content
VLSI Mentor

DDR · Module 31

DDR vs HBM

Per signal wire DDR carries more bandwidth; HBM wins because it reaches a surface that can host far more wires. The deciding quantity is connections per unit of capability, and a perimeter's ratio falls as one over the side length.

Start with the number that contradicts the expectation, because it fixes what this comparison is actually about.

DERIVED below, and recomputed in §5: per signal wire, a DDR channel carries more bandwidth than an HBM stack does. Not less. HBM does not move more data per connection — it moves slightly less, and it wins anyway, decisively, for a reason that is not about memory at all.

The deciding quantity is not bandwidth per connection. It is how many connections the chosen surface can host per unit of the capability being fed — and a package perimeter's ratio falls as one over the side length, monotonically, from the first millimetre.

So this is the chapter where axis A4 moves further than any axis in the module, and the reason is geometry. CURRICULUM-DERIVED from 26.1 §1: a die's capability scales with its area and its edge connections with its perimeter 4L, so connections per unit of capability is 4/Land the usual answer to wanting more of something, build it bigger, is exactly the wrong move.

Axis A1 is identical for the second chapter running. The same destructive 1T1C cell, the same thirteen obligations of 31.1 §5. So once again no structural argument separates the two, and once again the decision is entirely stage two — except that here stage two has a feasibility half before its cost half, which is the shape 31.1 §1 gave to stage one.

1. The Comparison That Is Not Available

Both technologies are DRAM with the same cell, so the usual differentiators are unavailable and it is worth saying which.

Claim you might reach forWhy it is not available here
“HBM has lower latency”the same destructive read, restore, row buffer and timing classes — 31.1 §4's chain runs identically
“HBM has a simpler controller”§7: the obligations are multiplied, not reduced
“HBM is faster per wire”§5 computes it and the result goes the other way
“HBM is the newer technology”not an engineering argument, and both continue to advance
“HBM gives more capacity”CURRICULUM-DERIVED from 26.4 §1: capacity and bandwidth are two separate purchases, and conflating them is that chapter's subject

So what remains is a geometry argument and a cost argument, in that order — and the geometry argument is a feasibility test that can close the question before any cost is considered, exactly as 31.1 §1's admissibility test can.

The two-stage structure, restated for this comparison:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   STAGE ONE -- FEASIBILITY (a geometry question)
     can the chosen SURFACE host enough connections to carry the
     required bandwidth, at a manufacturable pitch, for a die of
     the size the capability requires?

     if no: the technology is not slow for this application. It is
     IMPOSSIBLE for it, and no signalling improvement changes that
     -- §4 is why the exponent cannot be argued away.

   STAGE TWO -- COST (everything else)
     interposer, assembly, yield, repairability, capacity
     granularity, and the controller replication of §7.

2. What Is the Same

Axis A1 is identical, so 31.2 §2's thirteen-for-thirteen table holds again and is not repeated. The same destructive read, the same restore, the same row-buffer concept, the same four constraint classes of 13.3 with different values.

One shared property deserves emphasis because it is often assumed away. CURRICULUM-DERIVED from 26.1 §5 and 26.1 §6, which own why the two halves of a channel share a command bus and what semi-independent actually means: HBM's channels are not fully independent memories. So a design that models them as N unrelated DDR channels has over-stated their independence, and 26.1 §10's pseudo-channel command arbiter exists precisely because of what is shared.

And one more, which §12's defect turns on. Both technologies transfer in bursts, so both have a minimum useful transfer that spans a contiguous run of addresses. That burst span is what interacts with an interleave granularity, and it is the same in kind for both — which is why the hazard is a configuration hazard rather than a technology property.

3. The Axes, for This Comparison

AxisDDRHBMWho owns the detail
A1 celldestructivedestructiveidentical — §2
A2 obligationsthirteen, oncethirteen, times the channel count§7
A3 granularityburst-shapedburst-shaped, but finer per channel26.1 §4
A4 channel shapefew, widemany, narrow, semi-independent26.1 §3, 26.1 §6
A5 latency compositionthe same termsthe same terms, shorter transit23.1
A6 bandwidth rungrungs 1 and 2rung 1 moves by an order30.8 §2
A7 physical couplingsocketed or soldered, replaceableon-package, assembled once, not replaceable§16

A4 is the axis that moves, and A2 and A7 are its consequences. That is the chapter's structure: one axis moves for a geometric reason, and two others follow.

A6 deserves its qualification immediately. Rung 1 — peak — moves by an order of magnitude. Rungs 3, 4 and 5 do not follow automatically, and 30.8 §4 owns the consequence: three of the four gaps in the bandwidth ladder lie outside the memory, so a technology that multiplies rung 1 can leave application-observed bandwidth almost unchanged. CURRICULUM-DERIVED from 26.4 §7, which owns the result directly: one requester cannot fill them.

4. Why the Surface Decides, and Why Cleverness Cannot

CURRICULUM-DERIVED from 26.1 §1, whose derivation is the load-bearing one for this whole chapter. A square die of side L has capability scaling as and edge connections scaling as 4L, so:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   connections per unit of capability  =  4L / L²  =  4 / L

CURRICULUM-DERIVED, with that chapter's ILLUSTRATIVE side lengths and its recomputation: 5 mm gives 0.800, 10 mm gives 0.400, 20 mm gives 0.200. Doubling the side halves the connections available per unit of capability.

Three properties of that expression decide the comparison, and 26.1 §1 owns all three.

It is a gradient, not a threshold. There is no size at which perimeter routing suddenly fails; it degrades continuously from the first millimetre. So there is no pin-count wall cycle to point at, and a design cannot wait for a discrete event before acting.

Building bigger makes it worse. More capability to feed, proportionally fewer edges to feed it through. CURRICULUM-DERIVED, and it is the inversion that makes the wall interesting: the usual response to wanting more of something is the wrong move here.

And no signalling improvement changes the exponent. A better interface multiplies the numerator by a constant. CURRICULUM-DERIVED from 26.1 §1 and 4.8: it cannot turn 4/L into anything that does not fall.

So the class of solution is selected by arithmetic before any engineering is done. CURRICULUM-DERIVED from 26.1 §1's callout: if the binding quantity is perimeter / area, there are exactly two ways out — make the numerator scale like the denominator, or stop using the perimeter — and connecting through the die's face does both at once, because a face is an area, so the available count scales as and the ratio becomes a constant instead of falling.

That is the entire deciding quantity, and it is worth stating in the form an engineer can apply:

Ask which surface the connections land on. A perimeter gives a ratio that falls as 4/L. A face gives a constant. Every other difference between these two technologies is a consequence or a cost.

5. Per Connection, the Result Goes the Other Way

Now compute the thing everybody assumes, because the answer is instructive.

HBM side — CURRICULUM-DERIVED, from 26.4 §1, which attributes its figures to 4.8 §2 and 26.2 §8: the interface is 1024 bits wide at every stack height, and a stack delivers 307.2 GB/s.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   DERIVED, recomputed:
       307.2 GB/s  /  1024 signals   =  0.300 GB/s per signal
   and per-pin rate:
       307.2e9 B/s x 8 bits/B / 1024 =  2.40 Gbit/s per signal

DDR side — ILLUSTRATIVE and parametric, because this chapter has no verified DDR pin count of its own and §Scope forbids inventing one. Take a channel W bits wide at R transfers per second:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   bandwidth       =  W x R / 8   bytes per second
   per signal      =  (W x R / 8) / W  =  R / 8   bytes per second

   DERIVED, and read it before substituting anything: the per-signal
   figure does not depend on W at all. It is the TRANSFER RATE
   divided by eight, and nothing else.

   ILLUSTRATIVE R = 3200 MT/s:
       per signal  =  3200e6 / 8  =  0.400 GB/s per signal
       per-pin rate = 3.20 Gbit/s per signal

DERIVED comparison:

per signalper-pin rateratio
HBM stack, CURRICULUM-DERIVED0.300 GB/s2.40 Gbit/s
DDR channel, ILLUSTRATIVE R = 3200 MT/s0.400 GB/s3.20 Gbit/s1.33× better

So per wire, the perimeter technology is ahead — by a third, on these figures. And the reason is structural rather than accidental: a signal driven across a board at a socket must clear a harder electrical problem than one driven across an interposer, so it is driven faster precisely because there are so few of it. Module 22 owns the electrical half of that, and 26.2 §2 owns the interposer property — density — that makes the many-slow-wires arrangement possible at all.

Two results follow, and the second is the one to carry.

The bandwidth advantage is entirely a count advantage. DERIVED: 1024 × 0.300 = 307.2 against W × 0.400. For the perimeter technology to match one stack it needs W = 307.2 / 0.400 = 768 signals on one channel — and §4 is the reason a package perimeter cannot host that per unit of the capability being fed. So the comparison is not which is faster per wire but which surface can host 768-plus wires.

And that reframes what a “better interface” can buy. Improving the per-signal rate multiplies one column of the table above. CURRICULUM-DERIVED from 26.1 §1: a constant factor cannot turn 4/L into something that does not fall, so signalling improvements postpone the wall and never remove it — and 31.4 is the chapter about a technology that pushes that column as far as it goes and pays for it elsewhere.

6. The Feasibility Test, Stated

Stage one from §1, as a procedure. It uses 26.2's conversions and adds nothing to them.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   INPUTS
     B_req    required bandwidth
     L        die side implied by the capability to be fed
     pitch    manufacturable connection pitch on the chosen surface
     r_sig    achievable per-signal rate on that surface
     f_sig    fraction of connections that carry a SIGNAL rather than
              power or ground -- 26.2 §6 owns that this fraction is
              well below one, and this chapter does not invent it

   PERIMETER SURFACE
     available  =  (4 x L / pitch) x f_sig
     carried    =  available x r_sig
     FEASIBLE iff carried >= B_req

   FACE SURFACE
     available  =  (L^2 / pitch^2) x f_sig
     carried    =  available x r_sig
     FEASIBLE iff carried >= B_req

   DERIVED, and the whole content is the exponent: the perimeter's
   available count is LINEAR in L and the face's is QUADRATIC. So
   the face's advantage GROWS with the die, and the technology
   becomes more attractive exactly as the problem gets harder.

The f_sig term is why a naive count overstates both sides. CURRICULUM-DERIVED from 26.2 §6, which owns the point that not every bump carries a signal and that the split is what determines how many bumps a given interface actually costs. A feasibility estimate that divides an area by a pitch and stops has computed an upper bound on an upper bound.

And 26.2 §5 owns the related correction, which matters for the same reason: the usable bump field is smaller than expected. So stage one should be run with the pessimistic figures, because a feasibility test that passes optimistically and fails in assembly has cost the whole schedule.

The honest limit of this test: it decides whether a surface can carry the traffic. It says nothing about whether the requester can supply it26.4 §7's one requester cannot fill themso a design that clears stage one can still deliver rung-5 bandwidth barely different from what it had. §16 puts that measurement before the packaging decision rather than after it.

7. Obligation Multiplication — the Module's Third Pattern

Three chapters, three patterns, and naming the third completes the set.

ChapterWhat happens to the obligation set
31.1deleted — twelve of thirteen vanish, because axis A1 changed
31.2added — thirteen become twenty-two, none removed
This chaptermultiplied — thirteen become thirteen per channel

And multiplication is worse than addition in a specific, measurable way: it changes what scales.

CURRICULUM-DERIVED from 26.1 §3 and 26.1 §6, which own the channel and pseudo-channel structure and what semi-independent means. Each channel needs its own row state, its own elapsed counters, its own rolling activate window and its own refresh accounting, because those are per-resource obligations and the resources are distinct.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   DERIVED, reusing 31.1 §10's per-bank state figure of 30 bits
   (row_open, row_known, row_idx, since_act, since_pre) and an
   ILLUSTRATIVE 8 banks per channel:

     per channel   =  8 banks x 30 bits  =  240 bits

     1 channel     =    240 bits
     8 channels    =  1,920 bits
    16 channels    =  3,840 bits

   plus, and this is the term that is NOT linear:

     a cross-channel distributor and its balance accounting, whose
     cost grows with the number of destinations it must choose
     among -- and 26.1 §10's pseudo-channel arbiter exists because
     the channels are SEMI-independent, so the shared parts need
     arbitrating rather than merely replicating.

Two consequences worth stating, and the second is a design-review argument.

Verification surface multiplies too. CURRICULUM-DERIVED from 30.9 §6's taxonomy: each replicated instance carries the same hazard population — stale state, convention off-by-one, count-versus-index — so sixteen channels is sixteen chances for each. And a bug in the replicated block appears sixteen times or in one instance only, which is a diagnostic: 19.1 §8's per-lane narrowing applied to channels.

And the distributor becomes the new hard problem, displacing the scheduler. With one wide channel, the interesting logic is the scheduler. With sixteen narrow ones, each scheduler's job is easier — fewer requests to choose among — and the difficult decision moves upstream into which channel a request belongs to. That is the address-decode question, 18.2 owns mapping for performance, and §11's defect lives exactly there.

8. The Interleave Hazard, in Principle

§2 noted that both technologies transfer in bursts spanning a contiguous run of addresses. Now combine that with many channels.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   a request's burst covers  BURST_SPAN  contiguous bytes.
   the channel decode picks a channel from some address field.
   the field's weight determines the INTERLEAVE GRANULARITY --
   how many contiguous bytes land on one channel before the
   next channel takes over.

   the requirement, stated as an inequality:

       INTERLEAVE_GRAN  >=  BURST_SPAN

   if it holds : every burst lies entirely within one channel.
   if it fails : a single burst STRADDLES two channels, and one
                 request becomes two channel transactions that
                 must be split, tracked and rejoined.

STRUCTURAL, and the important part is which configuration hides it: with few wide channels the granularity is naturally large and the inequality holds without anyone checking it. With many narrow channels the granularity is chosen small — deliberately, to spread a stream across channels — and the inequality becomes a real constraint that must be enforced.

So the hazard is a configuration hazard rather than a technology property, which is exactly the shape 31.1 §12 warned about: code correct under one configuration, reused under another without re-deriving the requirement. §11 is that code.

And the failure is not a performance loss. A straddling burst that is not split returns bytes from the wrong channel for part of its span — a data-correctness failure with no error signal, the shape 31.2 §7 identified for the PASR mask and 30.4 §5 for a missed write deadline.

9. Two Surfaces, as a Stack

A stacked layer diagram comparing the two technologies from the requester downwards. The top layer is the requester, and it is the layer that must supply enough concurrent demand to fill whatever is below it; chapter twenty-six point four establishes that one requester cannot fill a many-channel memory. The second layer is the channel distributor, which is present in a meaningful form only for the many-channel technology: it decodes which channel a request belongs to and must enforce that the interleave granularity is at least the burst span, and this is the layer where chapter eleven's defect lives. The third layer is the replicated controller, and it is the multiplication layer: thirteen obligations per channel rather than thirteen in total, so row state, elapsed counters, the rolling activate window and refresh accounting all exist once per channel. The fourth layer is the semi-independent channel structure, where the two halves of a channel share a command bus, so the shared parts need arbitrating rather than merely replicating. The fifth layer is the thirteen shared obligations themselves, identical in kind for both technologies because the cell is identical. The sixth layer is the connection surface, and this is the causal layer of the comparison: a package perimeter offers connections proportional to four times the side length, giving a ratio that falls as four over the side length, while a die face offers connections proportional to the side length squared, giving a constant ratio. The bottom layer is the cell, identical destructive single-transistor storage in both. The diagram is read from the sixth layer outwards: the surface choice is the cause, the multiplication in layer three is its consequence, and the distributor in layer two is the new hard problem that consequence creates.The same cell, two surfaces — and where the multiplication landsRequester — and it must FILL what is below26.4 §7: one requester cannot fill a many-channel memory. Rung 1 moving does not move rung 5.26.4 §7: one requester cannot fill a many-channel memory. Rung 1 moving does not move rung 5.Channel distributor — MANY-CHANNEL ONLYDecodes the channel and must enforce INTERLEAVE_GRAN >= BURST_SPAN (§8). §11's defect is here.Decodes the channel and must enforce INTERLEAVE_GRAN >= BURST_SPAN (§8). §11's defect is here.Replicated controller — THE MULTIPLICATIONThirteen obligations PER CHANNEL. Row state, counters, rolling window and refresh, once each (§7).Thirteen obligations PER CHANNEL. Row state, counters, rolling window and refresh, once each (§7).Semi-independent channels26.1 §5, §6: the two halves share a command bus, so shared parts are arbitrated, not replicated.26.1 §5, §6: the two halves share a command bus, so shared parts are arbitrated, not replicated.The thirteen shared obligations — IDENTICAL IN KINDSame four constraint classes of 13.3, different values. Axis A1 is the same, so this layer is too.Same four constraint classes of 13.3, different values. Axis A1 is the same, so this layer is too.THE CONNECTION SURFACE — THE CAUSAL LAYER (A4)Perimeter: 4L connections, ratio falls as 4/L. Face: L-squared connections, ratio CONSTANT (26.1 §1).Perimeter: 4L connections, ratio falls as 4/L. Face: L-squared connections, ratio CONSTANT (26.1 §1).The cell (A1) — IDENTICALDestructive 1T1C in both. Second chapter running where the bottom layer is not the difference.Destructive 1T1C in both. Second chapter running where the bottom layer is not the difference.

Read outwards from the sixth layer. The surface is the cause; the multiplication above it is the consequence; and the distributor above that is the new hard problem the consequence creates. The cell, at the bottom, is not the difference for the second chapter in a row — which is the module's recurring lesson about where comparisons actually live.

10. RTL — Decode and Replicate

The parameterised comparative block for this chapter. One NUM_CH parameter takes it from the few-wide configuration to the many-narrow one, and the decode and the replication are in the same module on purpose — because §7's argument is that the replication is cheap and the decode is where the difficulty moved.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// ---------------------------------------------------------------------
// channel_decode_replicate -- the comparative block of §10.
//
// CLASSIFICATION: synthesisable, ILLUSTRATIVE parameter values, and
// CORRECT as written. The intentionally defective block is §11.
//
// WHAT IT IS: an address decoder plus per-channel obligation state,
// parameterised on the channel count. NUM_CH = 1 or 2 is the
// few-wide configuration; 8 or 16 the many-narrow one. §7's
// multiplication is the `for` loop over channels; §8's inequality is
// the elaboration guard.
//
// WHY IT EXISTS HERE: §7 claims the difficult decision MOVES
// UPSTREAM from the scheduler to the distributor as the channel
// count rises. This block is both halves of that claim in one place,
// so the reader can see that the per-channel state is a loop and the
// decode is a design decision.
//
// HOW TO RUN IT: elaborate at NUM_CH = 1 and at NUM_CH = 16 and
// compare state-element counts; then present a request whose burst
// span exceeds the interleave granularity.
// EXPECTED RESULT: the second case is REFUSED AT ELABORATION by the
// guard below, which is the correct place to refuse it -- §8's
// inequality is a static property of the configuration, not a
// runtime condition.
//
// SYNTHESIS: NUM_CH x BANKS_PER_CH per-bank records, plus a decode
// that is combinational and a one-hot channel select.
//
// LIMITATIONS: models the DECODE and the per-channel OBLIGATION
// STATE. It does not model the shared command bus between the two
// halves of a channel -- 26.1 §5 owns why the halves share it and
// 26.1 §10 owns the arbiter that follows, and duplicating that here
// would rebuild another chapter's subject. Balance across channels
// is measured, not enforced: §13.
// ---------------------------------------------------------------------
module channel_decode_replicate #(
  parameter int NUM_CH          = 16,
  parameter int BANKS_PER_CH    = 8,
  parameter int ADDR_W          = 34,
  parameter int ROW_W           = 15,

  // §8's two quantities. BURST_SPAN is the contiguous byte run one
  // burst covers; INTERLEAVE_GRAN is how many contiguous bytes land
  // on one channel before the next takes over.
  parameter int BURST_SPAN      = 64,
  parameter int INTERLEAVE_GRAN = 256,

  // ILLUSTRATIVE timing. 14.1 to 14.3 own the real values.
  parameter int TRCD            = 14,
  parameter int TRP             = 14,
  parameter int TRAS            = 34,

  // COUNT, not INDEX: each counter must REPRESENT its threshold, so
  // it needs $clog2(threshold + 1) bits. TRAS is the largest.
  parameter int CNT_W           = $clog2(TRAS + 1),
  parameter int CH_W            = (NUM_CH > 1) ? $clog2(NUM_CH) : 1,
  parameter int BANK_W          = $clog2(BANKS_PER_CH),

  // The bit position the channel field starts at. DERIVED from the
  // granularity rather than chosen, so the two cannot disagree.
  parameter int CH_SHIFT        = $clog2(INTERLEAVE_GRAN)
)(
  input  logic                clk,
  input  logic                rst_n,

  input  logic                req_valid,
  input  logic [ADDR_W-1:0]   req_addr,
  input  logic [ROW_W-1:0]    req_row,

  output logic [CH_W-1:0]     sel_ch,
  output logic [NUM_CH-1:0]   ch_onehot,
  output logic                straddles,
  output logic [NUM_CH-1:0]   ch_legal,
  output logic [15:0]         state_bits_reported
);
  initial begin
    if (NUM_CH < 1) $fatal(1, "channel_decode_replicate: NUM_CH >= 1");
    if (NUM_CH > 1 && (NUM_CH & (NUM_CH - 1)) != 0)
      $fatal(1, "channel_decode_replicate: NUM_CH must be a power of two for a bit-field decode");
    if (BURST_SPAN < 1) $fatal(1, "channel_decode_replicate: BURST_SPAN >= 1");
    if ((INTERLEAVE_GRAN & (INTERLEAVE_GRAN - 1)) != 0)
      $fatal(1, "channel_decode_replicate: INTERLEAVE_GRAN must be a power of two");

    // §8's inequality, enforced AT ELABORATION. This is the correct
    // place for it: the relationship between the burst span and the
    // interleave granularity is a static property of the
    // configuration, so a design that violates it should not
    // elaborate rather than failing a runtime check on the one
    // access that happens to straddle.
    if (INTERLEAVE_GRAN < BURST_SPAN)
      $fatal(1, "channel_decode_replicate: INTERLEAVE_GRAN (%0d) < BURST_SPAN (%0d): a burst would straddle two channels (§8)",
             INTERLEAVE_GRAN, BURST_SPAN);

    if (CH_SHIFT + CH_W > ADDR_W)
      $fatal(1, "channel_decode_replicate: channel field falls outside the address");
  end

  // ---- The decode. One line of logic, and §7's argument is that
  //      this line is where the difficulty went.
  generate
  if (NUM_CH > 1) begin : g_multi
    assign sel_ch = req_addr[CH_SHIFT + CH_W - 1 : CH_SHIFT];
  end else begin : g_single
    assign sel_ch = '0;
  end
  endgenerate

  always_comb begin
    ch_onehot = '0;
    if (req_valid) ch_onehot[sel_ch] = 1'b1;
  end

  // ---- The straddle check, retained as a RUNTIME output even though
  //      the elaboration guard makes it unreachable. It is not
  //      redundant: it is the observable that proves the guard was
  //      the right guard, and 30.5 §11's lesson is that a guard which
  //      another guard always shadows is UNTESTED rather than
  //      working. §14's cover watches this signal for exactly that
  //      reason.
  logic [ADDR_W-1:0] span_end;
  always_comb begin
    span_end  = req_addr + BURST_SPAN[ADDR_W-1:0] - 1;
    straddles = req_valid &&
                (span_end[CH_SHIFT + CH_W - 1 : CH_SHIFT] != sel_ch);
  end

  // ===================================================================
  // §7's MULTIPLICATION. The obligations are identical in kind to
  // 31.1 §5's thirteen; what changes is that there are NUM_CH of
  // every one of them. Written as a loop precisely to make the point
  // that replication is CHEAP TO WRITE and expensive to verify.
  // ===================================================================
  logic             row_open  [NUM_CH][BANKS_PER_CH];
  logic             row_known [NUM_CH][BANKS_PER_CH];
  logic [ROW_W-1:0] row_idx   [NUM_CH][BANKS_PER_CH];
  logic [CNT_W-1:0] since_act [NUM_CH][BANKS_PER_CH];
  logic [CNT_W-1:0] since_pre [NUM_CH][BANKS_PER_CH];

  logic [BANK_W-1:0] req_bank;
  assign req_bank = req_addr[CH_SHIFT + CH_W + BANK_W - 1 : CH_SHIFT + CH_W];

  always_ff @(posedge clk) begin
    if (!rst_n) begin
      for (int c = 0; c < NUM_CH; c++)
        for (int b = 0; b < BANKS_PER_CH; b++) begin
          row_open[c][b]  <= 1'b0;
          row_known[c][b] <= 1'b0;
          row_idx[c][b]   <= '0;
          since_act[c][b] <= '0;
          since_pre[c][b] <= '0;
        end
    end else begin
      for (int c = 0; c < NUM_CH; c++)
        for (int b = 0; b < BANKS_PER_CH; b++) begin
          if (since_act[c][b] != {CNT_W{1'b1}}) since_act[c][b] <= since_act[c][b] + 1'b1;
          if (since_pre[c][b] != {CNT_W{1'b1}}) since_pre[c][b] <= since_pre[c][b] + 1'b1;
        end
      // Only the selected channel's record advances on a command. The
      // per-channel independence of the OBLIGATIONS is real even
      // though the channels are only SEMI-independent at the command
      // bus (26.1 §6) -- the timing state is per resource and the
      // resources are distinct.
      if (req_valid && ch_legal[sel_ch]) begin
        row_open[sel_ch][req_bank]  <= 1'b1;
        row_known[sel_ch][req_bank] <= 1'b1;
        row_idx[sel_ch][req_bank]   <= req_row;
        since_act[sel_ch][req_bank] <= '0;
      end
    end
  end

  // ---- Per-channel legality. NUM_CH independent copies of the same
  //      maximum-over-rules test that 13.3 owns.
  always_comb begin
    for (int c = 0; c < NUM_CH; c++) begin
      ch_legal[c] = row_open[c][req_bank]
                  ? (since_act[c][req_bank] >= TRCD[CNT_W-1:0])
                  : (since_pre[c][req_bank] >= TRP[CNT_W-1:0]);
    end
  end

  // §7's count, reported so the comparison is measured and not
  // claimed. 30 bits per bank, from 31.1 §10's figure.
  assign state_bits_reported = NUM_CH * BANKS_PER_CH * 30;
endmodule

The elaboration guard is the block's most important line, and it is worth defending. §8's inequality is a static property of the configuration, so refusing to elaborate is strictly better than checking at runtime — a runtime check fails on the one access that happens to straddle, which is data-dependent and may not occur in a regression at all. CURRICULUM-DERIVED from 30.9 §3's tool table: lint and elaboration prove structural facts, and a structural fact should be proved by the cheapest tool that can prove it.

11. RTL Review — The Channel Distributor

The intended contract:

  1. Decode each request to exactly one channel.
  2. A request's burst must lie entirely within the channel it was decoded to — §8's inequality.
  3. Where the configuration cannot satisfy clause 2, the request must be split into per-channel sub-requests, each with its own length, and rejoined on return.
  4. The channel select must be one-hot when a request is valid, and zero otherwise.
  5. The decode must be stable for the duration of a request, so that a return can be attributed to the channel that served it.
Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// ---------------------------------------------------------------------
// channel_distributor -- INTENTIONALLY DEFECTIVE, for review (§11).
//
// CLASSIFICATION: synthesisable, ILLUSTRATIVE, and CONTAINS A BUG.
//
// WHAT IT IS MEANT TO DO: the five-clause contract above -- spread a
// request stream across NUM_CH channels, splitting any request whose
// burst would straddle a channel boundary.
//
// WHY IT EXISTS HERE: §8 establishes that the interleave-versus-burst
// inequality holds WITHOUT ANYONE CHECKING IT at the few-wide
// configuration and becomes a real constraint at the many-narrow
// one. This block is that configuration dependence, and it is the
// third instance in this module of the same bug class: code correct
// under one parameterisation, reused under another without
// re-deriving the requirement (31.1 §12, 31.2 §12).
//
// HOW TO RUN IT: elaborate with NUM_CH = 16 and INTERLEAVE_GRAN = 32
// against BURST_SPAN = 64, then present an aligned request.
// EXPECTED RESULT under clause 3: the request is SPLIT, and
// `split_req` is asserted.
// EXPECTED TRACE: a burst covering bytes 0 to 63 with a 32-byte
// granularity must produce two sub-requests, to channels 0 and 1.
//
// SYNTHESIS: one combinational decode plus a one-hot expander.
//
// LIMITATIONS: models the DECODE only. Return-path rejoining is the
// other half of clause 3 and is out of scope here -- 10.5 §9 owns
// the association problem -- and that omission is STATED rather than
// hidden. It is not what the bug is.
// ---------------------------------------------------------------------
module channel_distributor #(
  parameter int NUM_CH          = 16,
  parameter int ADDR_W          = 34,
  parameter int BURST_SPAN      = 64,
  parameter int INTERLEAVE_GRAN = 32,
  parameter int CH_W            = (NUM_CH > 1) ? $clog2(NUM_CH) : 1,
  parameter int CH_SHIFT        = $clog2(INTERLEAVE_GRAN)
)(
  input  logic               clk,
  input  logic               rst_n,

  input  logic               req_valid,
  input  logic [ADDR_W-1:0]  req_addr,

  output logic [CH_W-1:0]    sel_ch,
  output logic [NUM_CH-1:0]  ch_onehot,
  output logic               split_req,
  output logic [7:0]         sub_req_count
);
  initial begin
    if (NUM_CH < 1) $fatal(1, "channel_distributor: NUM_CH >= 1");
    if ((INTERLEAVE_GRAN & (INTERLEAVE_GRAN - 1)) != 0)
      $fatal(1, "channel_distributor: INTERLEAVE_GRAN must be a power of two");
    if (CH_SHIFT + CH_W > ADDR_W)
      $fatal(1, "channel_distributor: channel field falls outside the address");
    // NOTE: §10's guard on INTERLEAVE_GRAN >= BURST_SPAN is ABSENT
    // here, and deliberately so -- this block is supposed to SUPPORT
    // the straddling configuration by splitting (clause 3), so
    // refusing to elaborate would defeat its purpose. §12 is about
    // what was removed along with it.
  end

  // Clause 1 and 4: decode to one channel, one-hot.
  assign sel_ch = (NUM_CH > 1) ? req_addr[CH_SHIFT + CH_W - 1 : CH_SHIFT]
                               : '0;

  always_comb begin
    ch_onehot = '0;
    if (req_valid) ch_onehot[sel_ch] = 1'b1;
  end

  // Clause 3: decide whether a split is needed.
  assign split_req     = 1'b0;                          // <-- THE DEFECT
  assign sub_req_count = req_valid ? 8'd1 : 8'd0;
endmodule

Before reading on: which clause, and why did dropping §10's elaboration guard look like the right move?

12. The Defect — The Guard Was Load-Bearing

The violated clause is 3, and the defect is that split_req is tied low: the block never splits, whatever the configuration.

And the reason it is interesting is the comment that removes the guard. §10 refuses to elaborate when INTERLEAVE_GRAN < BURST_SPAN. This block deliberately drops that guard, with a defensible justification — this block is supposed to support the straddling configuration by splittingand then does not implement the splitting.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   ILLUSTRATIVE NUM_CH = 16, INTERLEAVE_GRAN = 32, BURST_SPAN = 64.
   CH_SHIFT = 5, CH_W = 4, so the channel field is addr[8:5].

   a request at address 0, burst covering bytes 0..63:

       byte  0..31  ->  addr[8:5] = 0  ->  channel 0
       byte 32..63  ->  addr[8:5] = 1  ->  channel 1

   contract clause 3 : SPLIT into two sub-requests
   this block        : sel_ch = 0, sub_req_count = 1, split_req = 0

   so channel 0 is asked for 64 bytes. It has 32. It returns 64 --
   the second 32 coming from whatever is at the next address WITHIN
   channel 0, which is byte offset 512 of the interleaved space.

   the requester receives 64 contiguous-looking bytes of which the
   second half belongs to a different address entirely.

Three properties of this failure make it the worst in the chapter.

It is a data-correctness failure with no error signal. Nothing is illegal; no timing rule is violated; the channel served a perfectly legal request for data it does not hold. Chapter 31.2 §7 identified this shape for the PASR mask and 30.4 §5 for a missed write deadline — an obligation whose violation returns something rather than an error.

It is configuration-silent. At INTERLEAVE_GRAN = 256 and BURST_SPAN = 64 the inequality holds, no burst ever straddles, and split_req tied low is correct. So the block is right in the few-wide configuration and in any many-channel configuration with a coarse granularity — and it becomes wrong only when someone reduces the granularity to spread a stream more finely, which is a performance tuning change that nobody expects to affect correctness.

And the guard that would have caught it was removed for a good reason. That is the review lesson, and it generalises:

A guard removed because a block is supposed to handle the case it guarded against creates an obligation to implement the handling. Removing the guard and not implementing the handling leaves the design strictly worse than before the guard existed — because the configuration is now reachable and unprotected.

So the correct review question is not why is this guard missing but what obligation did its removal create, and where is that obligation discharged? The answer here is nowhere, and the commit that removed the guard is where to look.

The correction, and it needs both the split and the count:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
  // CORRECTED. Clause 3: compute the channel of the burst's LAST
  // byte as well as its first, and split when they differ. Note that
  // the general case can span MORE than two channels when the burst
  // span exceeds twice the granularity, so the count is derived
  // rather than fixed at two -- a two-way split would be a second
  // bug of the same family, right for one ratio and wrong for
  // another.
  logic [ADDR_W-1:0] span_end;
  logic [CH_W-1:0]   end_ch;
  always_comb begin
    span_end = req_addr + BURST_SPAN[ADDR_W-1:0] - 1;
    end_ch   = span_end[CH_SHIFT + CH_W - 1 : CH_SHIFT];

    split_req = req_valid && (end_ch != sel_ch);

    // The number of channels the burst touches. DERIVED from the
    // aligned span rather than from the channel indices, because the
    // indices WRAP and a difference computed on them is wrong at the
    // wrap point -- which would be a third bug of the same family.
    sub_req_count = req_valid
      ? 8'( ((req_addr[CH_SHIFT-1:0] + BURST_SPAN - 1) >> CH_SHIFT) + 1 )
      : 8'd0;
  end

And the structural recommendation, which is stronger than the fix: keep §10's elaboration guard and add a second parameter that explicitly enables splitting. A configuration that needs splitting then has to say so, the guard fires for every configuration that did not ask, and the obligation created by disabling the guard is visible in the parameter list rather than in a comment. CURRICULUM-DERIVED from 30.9 §3: a structural fact should be proved by the cheapest tool that can prove it, and making the dangerous configuration opt-in keeps elaboration as that tool.

13. RTL — Measuring the Multiplication

§7 claims the difficulty moves from the scheduler to the distributor. This block measures whether it did, by reporting the two quantities that distinguish a working distribution from a broken one: balance and per-channel stall attribution.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// ---------------------------------------------------------------------
// channel_balance_monitor -- verification/telemetry, CORRECT as written.
//
// CLASSIFICATION: synthesisable telemetry. Drives nothing.
//
// WHAT IT DOES: counts requests per channel over a window and
// reports the imbalance, plus how many cycles each channel was
// blocked while another had work available.
//
// WHY IT EXISTS HERE: §7 argues the distributor is the new hard
// problem, and §11's defect is one way it goes wrong. This block
// catches the OTHER way, which produces no wrong data at all: a
// decode that is correct per request and systematically unbalanced
// across the stream, so the many-channel memory delivers the
// bandwidth of a few-channel one. 30.8's ladder names that exactly
// -- rung 1 moved and rung 4 did not.
//
// HOW TO RUN IT: run the real access pattern and read `imbalance`.
// EXPECTED RESULT: a stride that is a multiple of NUM_CH x
// INTERLEAVE_GRAN concentrates on one channel and imbalance
// approaches its maximum -- 18.2's mapping-for-performance problem
// at channel granularity.
//
// SYNTHESIS: NUM_CH counters plus a max/min tracker. No divider.
//
// LIMITATIONS: measures DISTRIBUTION, not correctness. §11's defect
// produces a perfectly balanced stream of wrong data, so this block
// cannot see it -- which is why §14 asserts correctness separately
// and this block's LIMITATIONS say so rather than implying coverage
// it does not give (27.3's independence discipline).
// ---------------------------------------------------------------------
module channel_balance_monitor #(
  parameter int NUM_CH = 16,
  parameter int WIN    = 65536,
  // COUNT, not INDEX: each per-channel counter must be able to
  // represent a window in which EVERY request went to one channel,
  // so it needs $clog2(WIN + 1) bits. Sized $clog2(WIN / NUM_CH) --
  // the "fair share" -- it would saturate in exactly the imbalanced
  // case it exists to detect, and report perfect balance.
  parameter int CNT_W  = $clog2(WIN + 1)
)(
  input  logic                clk,
  input  logic                rst_n,

  input  logic                req_fire,
  input  logic [$clog2(NUM_CH > 1 ? NUM_CH : 2)-1:0] req_ch,
  input  logic [NUM_CH-1:0]   ch_has_work,
  input  logic [NUM_CH-1:0]   ch_blocked,
  input  logic                win_tick,

  output logic [CNT_W-1:0]    r_max_ch,
  output logic [CNT_W-1:0]    r_min_ch,
  output logic [CNT_W-1:0]    r_imbalance,
  output logic [CNT_W-1:0]    r_starved_cycles,
  output logic                result_valid
);
  initial begin
    if (NUM_CH < 1) $fatal(1, "channel_balance_monitor: NUM_CH >= 1");
    if (WIN < NUM_CH)
      $fatal(1, "channel_balance_monitor: WIN must be >= NUM_CH or balance is meaningless");
  end

  logic [CNT_W-1:0] cnt [NUM_CH];
  logic [CNT_W-1:0] starved;

  always_ff @(posedge clk) begin
    if (!rst_n) begin
      for (int c = 0; c < NUM_CH; c++) cnt[c] <= '0;
      starved          <= '0;
      r_max_ch         <= '0;
      // AND-accumulator discipline: a minimum tracker must start at
      // all-ones or the first sample never wins and the reported
      // minimum stays zero forever -- which would report maximum
      // imbalance on every window regardless of the traffic.
      r_min_ch         <= '1;
      r_imbalance      <= '0;
      r_starved_cycles <= '0;
      result_valid     <= 1'b0;
    end else if (win_tick) begin
      begin
        logic [CNT_W-1:0] mx, mn;
        mx = '0;
        mn = '1;
        for (int c = 0; c < NUM_CH; c++) begin
          if (cnt[c] > mx) mx = cnt[c];
          if (cnt[c] < mn) mn = cnt[c];
        end
        r_max_ch    <= mx;
        r_min_ch    <= mn;
        // Reported as a DIFFERENCE with both endpoints, not as a
        // ratio. 27.5 refuses a single percentage and 30.8 §11
        // requires a derived statistic's range to be asserted; a
        // difference with its endpoints lets the consumer see the
        // denominator and own the rounding.
        r_imbalance <= mx - mn;
      end
      r_starved_cycles <= starved;
      result_valid     <= 1'b1;
      for (int c = 0; c < NUM_CH; c++) cnt[c] <= '0;
      starved          <= '0;
    end else begin
      result_valid <= 1'b0;
      if (req_fire && cnt[req_ch] != {CNT_W{1'b1}})
        cnt[req_ch] <= cnt[req_ch] + 1'b1;

      // A starved cycle: some channel had work and was blocked while
      // at least one other channel was idle with nothing to do. That
      // is the distributor's failure signature rather than the
      // scheduler's -- 30.5 §3's four-stage attribution, applied
      // across channels instead of within one.
      if (((ch_has_work & ch_blocked) != '0)
          && ((~ch_has_work) != '0)
          && starved != {CNT_W{1'b1}})
        starved <= starved + 1'b1;
    end
  end
endmodule

r_starved_cycles is the measurement that distinguishes the two failures §7 predicts. A high imbalance with low starvation means the pattern concentrates18.2's mapping problem at channel granularity. A low imbalance with high starvation means the distribution is fine and something downstream is blocking, which sends the investigation back to the per-channel schedulers. Two opposite conclusions from one pair of counters, and 30.8 §13 owns that shape of instrument.

14. SVA Review — Correct Per Request, Wrong Across the Stream

The property written for §11's block, and it is a real property that passes:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
  // Offered as "proves the distributor decodes correctly".
  // Clause 1 and 4. Correct, non-vacuous, and PASSES on the
  // defective block.
  property p_channel_select_is_onehot;
    @(posedge clk) disable iff (!rst_n)
      req_valid |-> $onehot(ch_onehot);
  endproperty
  assert property (p_channel_select_is_onehot)
    else $error("channel select was not one-hot for a valid request");

Q. It passes. What does it prove?

That the decode picks exactly one channel — which is precisely what the defect does. The defective block's failure is not that it picks the wrong number of channels; it is that picking one channel is the wrong answer for this request, and one-hotness cannot express that.

This is 30.9 §6's variety 2 — the property does not name the contract's key quantity. The key quantity here is not ch_onehot at all; it is the burst's span, which appears nowhere in the property and nowhere in the defective block. A property whose signal list does not include the burst length cannot constrain a claim about the burst's extent.

And it is the module's third instance of variety 10 as well, because at the coarse-granularity configuration this property plus split_req tied low together constitute a correct design — so the property is sound there and, at the fine granularity, certifies the absence of splitting as the intended behaviour.

What actually covers clauses 2 and 3:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
  // Clause 2. It names BURST_SPAN, which is the quantity the
  // defective block never computes. Variety 2's repair: make the
  // property mention the thing whose corruption it must catch.
  property p_burst_lies_within_one_channel;
    @(posedge clk) disable iff (!rst_n)
      (req_valid && !split_req) |->
        (((req_addr + BURST_SPAN - 1) >> CH_SHIFT) ==
          (req_addr >> CH_SHIFT));
  endproperty
  assert property (p_burst_lies_within_one_channel)
    else $error("an unsplit burst crosses a channel boundary");

  // Clause 3, the other direction: when the span DOES cross, a split
  // must be signalled. Stating both directions separately matters --
  // a design that always splits satisfies the property above and
  // wastes half the channel bandwidth, and a design that never
  // splits satisfies this one vacuously.
  property p_crossing_burst_is_split;
    @(posedge clk) disable iff (!rst_n)
      (req_valid &&
       (((req_addr + BURST_SPAN - 1) >> CH_SHIFT) != (req_addr >> CH_SHIFT)))
      |-> split_req;
  endproperty
  assert property (p_crossing_burst_is_split)
    else $error("a crossing burst was not split");

  // Clause 3's arithmetic. The sub-request count must equal the
  // number of channels the span actually touches -- a COMPARATIVE
  // obligation, so it is checked against an independently computed
  // expectation rather than against the design's own derivation.
  // 30.6 §11: "best" and "correct count" obligations need a model,
  // and 27.3's independence requirement says the model must not
  // reuse the design's expression.
  property p_sub_req_count_matches_span;
    @(posedge clk) disable iff (!rst_n)
      req_valid |-> (sub_req_count == ref_channels_touched(req_addr, BURST_SPAN));
  endproperty
  assert property (p_sub_req_count_matches_span)
    else $error("sub-request count disagrees with the reference model");

  // Clause 5: the decode must be stable while a request is in
  // flight, or a return cannot be attributed. 10.5 §9's association
  // problem, at channel granularity.
  property p_decode_stable_in_flight;
    @(posedge clk) disable iff (!rst_n)
      (req_valid && !req_fire) |=> $stable(sel_ch);
  endproperty
  assert property (p_decode_stable_in_flight)
    else $error("channel decode changed while a request was pending");

  // The elaboration-time fact, restated as a runtime invariant for
  // the block that KEEPS the guard (§10). It must be unreachable
  // there, and 30.5 §11's lesson applies: a guard another guard
  // always shadows is untested, so the cover below must be watched.
  property p_no_straddle_when_guarded;
    @(posedge clk) disable iff (!rst_n)
      (INTERLEAVE_GRAN >= BURST_SPAN) -> !straddles;
  endproperty
  assert property (p_no_straddle_when_guarded)
    else $error("a straddle occurred in a configuration that forbids it");

  // ---- Covers. The configuration covers are mandatory here for the
  //      same reason as in 31.1 §14 and 31.2 §14: the defect lives
  //      in a configuration, so coverage must be per configuration.
  // The GRANULARITY regimes -- the two sides of §8's inequality.
  cover property (@(posedge clk) disable iff (!rst_n)
                  req_valid && (INTERLEAVE_GRAN >= BURST_SPAN));
  cover property (@(posedge clk) disable iff (!rst_n)
                  req_valid && (INTERLEAVE_GRAN < BURST_SPAN));
  // An actual crossing burst, which is the antecedent of
  // p_crossing_burst_is_split. A zero here means the fine-grained
  // configuration ran and never presented the case -- which random
  // aligned traffic will do, because aligned bursts do not cross.
  cover property (@(posedge clk) disable iff (!rst_n)
                  req_valid &&
                  (((req_addr + BURST_SPAN - 1) >> CH_SHIFT) != (req_addr >> CH_SHIFT)));
  // A burst spanning MORE than two channels, which is where the
  // naive two-way split in §12's discussion would itself fail.
  cover property (@(posedge clk) disable iff (!rst_n)
                  req_valid && (sub_req_count > 8'd2));
  // §10's shadowed guard, watched. This cover must go DEAD in the
  // guarded configuration and be LIVE in the unguarded one -- two
  // different instruments, and 30.10 §12's point that knowing which
  // kind you are writing is the senior move.
  cover property (@(posedge clk) disable iff (!rst_n) straddles);
  // The channel-concentration case, so §13's imbalance measurement
  // is known to have been exercised at its extreme.
  cover property (@(posedge clk) disable iff (!rst_n)
                  result_valid && (r_min_ch == '0) && (r_max_ch != '0));

The third cover is the one that matters most, and it is not obvious. Aligned bursts do not cross channel boundaries. So a random-address regression on the fine-grained configuration will exercise crossings, but a regression using naturally aligned traffic — which is most realistic traffic — will not, and p_crossing_burst_is_split passes vacuously. CURRICULUM-DERIVED from 27.2 §6: the fix for a vacuous property is not a better property but a cover on the antecedent, and here the antecedent requires deliberately misaligned stimulus.

Follow-up an interviewer should ask: should the system-level environment generate misaligned requests? Usually not — if the upstream interconnect guarantees alignment, injecting misalignment tests a case the system cannot produce. At the block level, yes, because this block's contract is about spans and not about what the interconnect promises. Naming that boundary is the answer, and it is the same distinction 30.4 §9 draws about violating column spacing at the block level but never at the system level.

15. What the Replicated Block's Assertions Prove

§14 reviewed the distributor and left §10's replicated block unasserted, and that omission is worse here than it was in 31.1 §14 — because replication multiplies the hazard population. CURRICULUM-DERIVED from 30.9 §6: each instance carries the same stale-state, convention and count-versus-index hazards, so sixteen channels is sixteen chances for each, and the bug that appears in one instance only is a different diagnosis from the bug that appears in all sixteen.

The property that matters most in a replicated design is the one nobody writes for a single instance, because for a single instance it is trivially true.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
  // ---- ISOLATION. The property that only exists once a design is
  //      replicated, and the one that catches the classic
  //      replication bug: an index computed from the wrong variable.
  //      A command to one channel must not disturb any other
  //      channel's obligation state. For NUM_CH = 1 this is
  //      vacuously true, which is exactly why a design scaled up
  //      from one channel has never had it checked.
  property p_channel_state_isolation;
    @(posedge clk) disable iff (!rst_n)
      (req_valid && ch_legal[sel_ch]) |=>
        (foreach_other_channel_stable(sel_ch));
  endproperty
  assert property (p_channel_state_isolation)
    else $error("a command to one channel disturbed another channel's state");

  // Written concretely for a bounded channel count, because the
  // helper above hides the thing worth seeing -- that the check is
  // over every OTHER index, and that "every other" is what the
  // buggy version gets wrong.
  generate
  for (genvar c = 0; c < NUM_CH; c++) begin : g_iso
    property p_other_channel_untouched;
      @(posedge clk) disable iff (!rst_n)
        (req_valid && ch_legal[sel_ch] && (sel_ch != c)) |=>
          ($stable(row_open[c]) && $stable(row_idx[c]) && $stable(row_known[c]));
    endproperty
    assert property (p_other_channel_untouched)
      else $error("channel %0d's state changed on a command to another channel", c);
  end
  endgenerate

  // ---- Per-channel legality must be computed from that channel's
  //      OWN state. A single shared counter feeding every channel's
  //      test would pass every timing property and be wrong for
  //      fifteen of sixteen channels -- and it would pass because
  //      the timing arithmetic is right, which is 30.5 §11's
  //      "shape versus content" distinction at replication scale.
  generate
  for (genvar c = 0; c < NUM_CH; c++) begin : g_leg
    property p_legality_uses_own_state;
      @(posedge clk) disable iff (!rst_n)
        ch_legal[c] |-> (row_open[c][req_bank]
                          ? (since_act[c][req_bank] >= TRCD)
                          : (since_pre[c][req_bank] >= TRP));
    endproperty
    assert property (p_legality_uses_own_state)
      else $error("channel %0d's legality was computed from foreign state", c);
  end
  endgenerate

  // ---- The elaboration guard's runtime shadow. §10 keeps the guard,
  //      so `straddles` must be unreachable there. This is the
  //      assertion form of 30.5 §11's warning: a guard that another
  //      guard always shadows is UNTESTED, so the claim that it is
  //      unreachable should be asserted rather than assumed.
  property p_straddle_unreachable_when_guarded;
    @(posedge clk) disable iff (!rst_n)
      !straddles;
  endproperty
  assert property (p_straddle_unreachable_when_guarded)
    else $error("a straddle occurred despite the elaboration guard");

  // ---- The reported state count must match the configuration, so
  //      that §7's measured multiplication cannot silently disagree
  //      with the parameters. An INVARIANT, and the cheapest
  //      possible check that the configuration which elaborated is
  //      the one intended -- 31.1 §13's argument, reused.
  property p_state_count_matches_config;
    @(posedge clk) disable iff (!rst_n)
      state_bits_reported == (NUM_CH * BANKS_PER_CH * 30);
  endproperty
  assert property (p_state_count_matches_config)
    else $error("reported state bits disagree with the channel configuration");

  // ---- Covers. The replication regimes, because a property proved
  //      at one channel count says nothing about another -- variety
  //      10 again, with the channel count as the parameter.
  cover property (@(posedge clk) disable iff (!rst_n)
                  req_valid && (NUM_CH == 1));
  cover property (@(posedge clk) disable iff (!rst_n)
                  req_valid && (NUM_CH > 1));
  // Every channel must have been SELECTED at least once, or the
  // isolation properties above held only for the indices the
  // stimulus happened to reach. This is the replication analogue of
  // 31.2 §14's duration cover: the coverage item must be on the
  // dimension the defect scales with, and here that dimension is
  // the index space.
  generate
  for (genvar c = 0; c < NUM_CH; c++) begin : g_cov
    cover property (@(posedge clk) disable iff (!rst_n)
                    req_valid && (sel_ch == c));
  end
  endgenerate

p_channel_state_isolation is the section's deliverable, and the reason is a general one about scaling a design up.

A replicated design needs a class of property that a single instance does not: that operating on one instance leaves the others alone. For one instance that property is vacuously true, so a design grown from a single-channel predecessor has never had it checked — and the index arithmetic that broke is new code.

And the per-channel cover generate loop is the third coverage-reachability problem this module has produced. Chapter 31.1 §14 needed a cover per configuration; 31.2 §14 needed one per duration; this needs one per index. The common rule is worth stating once:

The coverage item must be on the dimension the defect scales with. A defect that scales with a configuration needs a configuration cover, one that scales with a duration needs a duration cover, and one that scales with an index space needs a cover per index.

Follow-up an interviewer should ask: is a per-index cover affordable at sixteen channels? Yes, and it is the cheap case — sixteen cover points is nothing, and the stimulus to reach them is a strided address sweep. The expensive version is the one that must be argued down: at 1024 banks across 16 channels, a cover per bank per channel is 16,384 points, and CURRICULUM-DERIVED from 27.5 §4 — which owns the reachability arithmetic and the finding that most of a naive cross is unhittable — that cross should be reduced deliberately rather than declared closed. Covering the index space of the thing that is replicated, and not the cross of everything inside it, is the distinction.

16. The Capacity Decision You Have Not Made Yet

Axis A7 is where this comparison becomes irreversible, and it interacts with a decision most projects take later.

CURRICULUM-DERIVED from 26.4 §1, which owns the two purchases and proves their independence: the interface is the same width at every stack height, so height changes capacity and leaves bandwidth alone. A die adds capacity only; a stack adds bandwidth, streams and capacity together. So one purchase is strictly more powerful and strictly more expensive.

Now add the packaging consequence, which is this chapter's contribution:

DecisionDDR, socketed or solderedHBM, on-package
Change capacity after designpossible — a different module or populationfixed at assembly
Change capacity after shippingpossible for socketedimpossible
Replace a failed devicepossible for socketedthe package is scrap
Probe the interface during bring-uppossible at the socket or the boardno access30.10 §2's observability collapse, worsened
Change bandwidthadd channels if the perimeter allows — §6add a stack, which is a new package

Two results follow, and they are decision inputs rather than trivia.

The capacity decision is pulled forward into the packaging decision. A project choosing the on-package technology has committed to a capacity range at the point it commits to a stack count, because 26.4 §1's two purchases are made simultaneously in one assembly. So a capacity requirement that is still uncertain must be resolved earlier than it otherwise would be, and that schedule consequence belongs in the comparison.

And the repairability loss compounds a cost this module has otherwise ignored. CURRICULUM-DERIVED from 26.3 §4 and 26.3 §6, which own why a stack needs spares and that repair is per group rather than per via: the technology has internal repair precisely because assembly yield on a stack is a real risk. That repair happens at manufacture. There is none in the field, and a design whose service model assumed module replacement has to change it.

The honest counterweight, because §Scope forbids a winner. For a system whose capability genuinely requires the bandwidth, §6's feasibility test has already closed the question — the perimeter cannot carry it, so replaceability was never on offer. Axis A7 is a cost to be acknowledged, not a tiebreaker to be applied after feasibility has already decided.

17. What Would You Measure?

Q. You must decide between these technologies for a new accelerator. What do you measure, in what order?

MeasurementWhat it settlesCostOwner
Required bandwidth against §6's perimeter feasibilitywhether the question is a choice at allhours§6, 26.1 §1
The requester's concurrency — how many streams it can sustainwhether a many-channel memory can be filleddays26.4 §7
Arithmetic intensity of the workloadwhether the compute rate is servable at alldays26.4 §5
Stride against NUM_CH × INTERLEAVE_GRANchannel concentration — §13's imbalancedays18.2, §13
r_imbalance and r_starved_cycles togetherdistributor problem versus downstream problemdays§13
Is the capacity requirement settled?§16 — the on-package choice fixes it at assemblyhours26.4 §1
The service model — is field replacement assumed?§16, and it can close the questionhours§16
The full five-rung ladder on a prototypewhether rung 1 moving moves rung 5weeks30.8 §3

Row one is first because it is the only row that can prove the question is not a choice. If the perimeter cannot carry the required bandwidth for a die of the necessary size, the comparison is over and everything below row one is a cost analysis of the only feasible option.

Row two is the row that reverses the conclusion most often. CURRICULUM-DERIVED from 26.4 §7: one requester cannot fill them. A design that clears feasibility and cannot supply the concurrency has bought rung 1 and will measure rung 5 unchanged — which is 30.8 §4's result that three of the four gaps lie outside the memory, arriving as a packaging decision.

And rows six and seven cost hours and are usually asked last. A settled capacity requirement and a service model are answerable in an afternoon, and either can close the question — so asking them after a week of bandwidth modelling is the ordering error 30.8 §1 warns about.

18. Common Wrong Answers

“HBM has more bandwidth per pin.” §5. DERIVED: 0.300 GB/s per signal against an illustrative 0.400 for DDR. Per wire the perimeter technology is ahead, and the advantage is entirely a count advantage.

“HBM is faster, so it has lower latency.” §1, §2. Identical cell, identical restore, identical constraint classes. Transit is shorter; the access is not.

“The pin-count wall is a threshold we have not reached.” §4, and 26.1 §1 owns the correction: it is a gradient that degrades continuously from the first millimetre. There is no event to wait for.

“A better interface postpones the wall indefinitely.” §4, §5. A constant factor cannot turn 4/L into something that does not fall. It postpones and never removes.

“Make the die bigger to fit more channels.” §4. A larger die has more capability to feed and proportionally fewer edges to feed it through — the inversion that makes the wall interesting.

“Divide the bump area by the pitch to get the connection count.” §6. That is an upper bound on an upper bound: 26.2 §5 owns that the usable field is smaller than expected, and 26.2 §6 that not every bump carries a signal.

“HBM channels are independent memories.” §2, and 26.1 §6 owns what semi-independent means: the two halves share a command bus, which is why a pseudo-channel arbiter exists.

“More channels means a simpler controller because each one is smaller.” §7. Each scheduler's job is easier and there are sixteen of them, plus a distributor that did not exist before. Thirteen obligations become thirteen per channel.

“HBM gives you more capacity.” §16, and 26.4 §1 owns the correction: capacity and bandwidth are two separate purchases, and a die adds only the first.

“More bandwidth will make the application faster.” §3, §16. Rung 1 moving does not move rung 5 — 30.8 §4 — and 26.4 §7 owns the reason: one requester cannot fill them.

“Interleave as finely as possible to spread the traffic.” §8. Below the burst span, a single burst straddles two channels — and if it is not split, the data is wrong with no error signal.

“The straddle case never happens with aligned traffic.” §14. Correct, and that is exactly why the property passes vacuously. Aligned bursts do not cross, so the antecedent needs deliberately misaligned block-level stimulus.

“The decode is one-hot, so the distributor is correct.” §14. One-hotness is what the defect does. The property never names the burst span, so it cannot constrain a claim about the burst's extent.

“The guard was unnecessary because this block handles the case.” §12. Removing a guard because a block is supposed to handle the case creates an obligation to implement it. Here it was never discharged, which leaves the design worse than before the guard existed.

“Splitting into two sub-requests handles it.” §12. Only when the span crosses one boundary. A burst spanning more than twice the granularity touches more than two channels, and a fixed two-way split is the same bug family again.

“The channels are balanced, so the distribution is working.” §13, §14. Balance and correctness are independent: §11's defect produces a perfectly balanced stream of wrong data, and §13's block says so in its own limitations.

“We can add capacity later.” §16. On-package capacity is fixed at assembly, so the capacity decision is pulled forward into the packaging decision.

“A failed device gets replaced.” §16. The stack's repair — 26.3 §6's per-group spares — happens at manufacture. There is none in the field, and the package is scrap.

19. Self-Check

  1. Compute bandwidth per signal for both technologies from §5's inputs, showing the DDR side's independence from channel width. State which is ahead and by how much.

  2. Given that result, explain in two sentences where HBM's bandwidth advantage actually comes from.

  3. Write 4/L from memory and state its three properties from §4. For each, say what design conclusion it forbids.

  4. Explain why connecting through a face changes the exponent rather than the constant, and why that distinction decides the class of solution before any engineering.

  5. Run §6's feasibility test symbolically for both surfaces. Identify the term that makes a naive area-divided-by-pitch estimate an upper bound on an upper bound.

  6. Name the module's three obligation patterns, one per chapter, and say which is worst for verification surface and why.

  7. State §8's inequality. Then explain why it holds unchecked in one configuration and becomes a real constraint in the other, and what kind of change makes a working design violate it.

  8. Find the defect in §11 without reading §12. Then answer the harder question: what obligation did removing §10's guard create, and where should it have been discharged?

  9. Explain why p_channel_select_is_onehot passes on the defective block, and why no strengthening of a one-hotness property reaches the contract.

  10. Explain why p_crossing_burst_is_split passes vacuously under realistic traffic, and state where the misaligned stimulus belongs and where it does not.

  11. Explain why the isolation property of §15 is vacuously true for a single instance, and why that makes a design scaled up from one channel especially exposed. Then state the general rule about which dimension a coverage item must be on.

  12. Give two of §16's irreversibility consequences and, for each, the project decision it pulls forward or forecloses.

20. Where This Goes

Per wire, the perimeter technology is ahead; the on-package technology wins because it reaches a surface that can host far more wires. The deciding quantity is connections per unit of capability, a perimeter's ratio falls as 4/L monotonically from the first millimetre, no signalling improvement changes that exponent, and the feasibility test can close the question before any cost is considered. Axis A1 is identical for the second chapter running; the obligation set is multiplied rather than changed; the difficult decision moves from the scheduler to the distributor; and axis A7 pulls the capacity decision forward into the packaging decision and forecloses field replacement.

Three results carry forward. A guard removed because a block is supposed to handle the guarded case creates an obligation, and the commit that removed it is where to look for the missing handling. Balance and correctness are independent measurements, and a perfectly balanced stream of wrong data is a real failure mode. And the antecedent of a span-crossing property requires deliberately misaligned stimulus, because realistic aligned traffic makes it vacuous — the third distinct coverage-reachability problem this module has produced, after 31.1 §14's configuration covers and 31.2 §14's duration covers.

Chapter 31.4 closes the module with the axis this chapter deliberately left alone. §5 showed that per-signal rate is one column of the table and that improving it postpones the wall without removing it. The final comparison is with a technology that pushes that column as far as a soldered point-to-point channel allows — and the interesting part is not the rate. It is what becomes affordable to give up once the consumer stops caring about latency, and what must be added back once the per-pin rate rises far enough that the channel stops being reliable on its own.

Continue learning

Standards & specifications

Governing standard
JEDEC JESD79 (DDR SDRAM)(opens JEDEC Solid State Technology Association in a new tab)

Defines the DDR SDRAM device itself — signals, command encoding, mode registers, timing parameters and the initialisation sequence — one document per generation. Memory-controller microarchitecture, address-mapping policy, PHY training algorithms and board-level design are not specified by it.

This page also covers RTL structure, verification approach and debugging technique. Those are engineering practice built on the standard, not requirements the standard itself imposes.

Where this fits

Part of the DDR curriculum.