DDR · Module 34
“DDR Bandwidth Equals Application Bandwidth”
Peak is correct, vendor-published, and the only one of five levels that exists before silicon — so this belief is the single option at the moment the decision is made. The fix is not a better number; it is a stated ratio at a stated evidence grade.
“DDR5-6400 on a 64-bit channel is 51.2 GB/s. We need 21.4. One channel, with headroom.”
Every number in that sentence is right, the arithmetic is right, and the conclusion is wrong by a factor that decides a channel count.
34.1 §1 set out this module's anatomy, and 34.2 §23 named a third recurring property of these six beliefs: the belief's evidence is usually the only evidence available when the belief is formed. DERIVED: this is the chapter where that property is sharpest, and it changes what a correction can even ask for.
CURRICULUM-DERIVED from 33.5 §13, which owns the five bandwidth levels and records the fact this chapter is built on: level 1, peak, is the only level that exists before silicon — “so at sizing time it is not a lazy choice, it is the only choice.”
DERIVED: so this belief cannot be corrected by demanding a better number, because no better number exists yet. It is corrected by demanding that the ratio between level 1 and level 5 be stated as an assumption — and an assumption written down is one that can be revisited when a measurement arrives, while an assumption of 1.0 left implicit cannot.
1. The One-Sentence Correction
Peak bandwidth is a property of the interface; application-visible bandwidth is what a requester receives as useful bytes after four separate reductions — and the ratio between them is measured at about 3.8× on a realistic mix, so a sizing decision that omits it under-provisions by nearly four times while every figure in it is correct.
CURRICULUM-DERIVED from 23.2's central law, which every chapter of Module 23 restates: peak bandwidth is a property of the interface, achieved bandwidth is a property of the workload meeting the timing rules, and the gap is the cost of constraints that cannot be removed plus decisions that can be improved.
2. What This Chapter Owns
| Ground | Owner |
|---|---|
| Peak bandwidth derived from verified parameters; measure D; the central law | 23.2 |
| The four efficiency measures and why they must stay distinct | 12.4 |
| What produces hit, miss and conflict counts; the 16× row-work spread | 23.3 |
| The five bandwidth levels, and who owns each loss | 33.5 §13 |
| Which efficiency is on the slide, and its denominator | 33.5 §6 |
| The exhaustive cycle attribution with a named residue | 33.5 §7 |
| The availability cost of refresh | 15.5 |
| Why the requester mix decides, across five platform classes | Module 32 |
| Why the belief is unavoidable at sizing time, and what a stated ratio costs | this chapter |
The boundary with 33.5 §13 is the tightest in this module and needs stating exactly, because that section already contains the refutation.
33.5 §13 owns the five levels, their owners, and the bandwidth_level instrument that computes all five from one window. It is a review item: it asks which level a published figure is. DERIVED: this chapter asks a different question — what do you do when only level 1 exists? That is not a review question, because there is nothing yet to review. §8's sizing model and §10's ratio provenance are the artifacts 33.5 has no reason to build, because a reviewer arrives after silicon and a sizing decision arrives before it.
And the boundary with 23.2 is a division of labour. 23.2 derives peak from verified parameters and simulates four workload models to find achieved. DERIVED: this chapter takes both as given and models what is done with the gap between them.
3. Teaching-Model Boundary And Source Discipline
| Claim class | What it means here | Example below |
|---|---|---|
| Structural | an arithmetic identity, or a documented mechanism | bytes per cycle from rate and width; the four reductions |
| Curriculum-derived | follows from a cited chapter of this track | the five levels, the central law, the four efficiencies |
| Derived | computed in this chapter from the models below | every ratio and channel count in §19 |
| Illustrative | a chosen number that makes a mechanism visible | workload mixes, requirement figures, fabric widths |
One identity is used throughout and it is STRUCTURAL: for a DDR interface, bytes per cycle = 2 × bus width in bytes, because two transfers occur per clock. DERIVED: at 6400 MT/s on a 64-bit bus that is 6400 × 8 = 51,200 MB/s, and this chapter computes it in place rather than quoting it.
Every other number is ILLUSTRATIVE. CURRICULUM-DERIVED from 33.5 §4's discipline for a performance chapter: a performance number in a tutorial is a grade D quantity — educational representative — and every argument below is written so it does not depend on the number's value. The 3.76× ratio is a DERIVED consequence of an ILLUSTRATIVE attribution, and §23 states what it is not.
No external source was consulted and no network tool was used.
4. Why a Competent Engineer Believes It
| # | The true statement | What the belief does with it |
|---|---|---|
| 1 | Peak is correct, vendor-published and independently checkable | treats verifiability as sufficiency |
| 2 | Peak is the only level that exists before silicon | treats the only available number as the right one |
| 3 | Bandwidth composes additively in a spreadsheet, unlike latency | treats tractability as accuracy |
| 4 | On a pure sequential stream, achieved approaches peak | treats the best case as the planning case |
Reason 2 is the chapter, and it is qualitatively different from the reasons in the two preceding chapters. DERIVED: 34.1's belief could be corrected by a measurement its holder could take that afternoon, and 34.2's by a 62 µs load test. This one cannot: the correcting measurement requires the silicon the decision is made in order to build.
So the belief is not a failure of rigour. It is a decision under uncertainty misrepresented as a calculation — and that reframing is the whole of §10.
Reason 3 deserves a line because it is the reason bandwidth beliefs outlive latency beliefs. DERIVED: latency's terms sum with an unbounded queueing component (23.1), so a latency spreadsheet visibly fails. Bandwidth's reductions multiply, and four factors near unity multiply to something that still looks plausible — CURRICULUM-DERIVED from 23.2's measured finding that the costs compose multiplicatively, and from 33.5 §11, which measured four 90% factors giving 65.6% where an additive estimate says 60%.
5. The Region Where the Claim Is True
"DDR BANDWIDTH == APPLICATION BANDWIDTH" HOLDS WHEN ALL FIVE HOLD:
1 one requester, sequential -> locality loss ~0
2 every byte of every burst is wanted -> payload loss 0
3 no read/write interleaving -> turnaround 0
4 refresh below the precision you need -> 4.5% ignorable
5 the fabric between requester and memory
is wider than the memory -> no fabric loss
and then: application bandwidth ~ 0.95 x peakDERIVED: the region is non-empty and it is exactly what a memory-bandwidth microbenchmark is written to produce. A large sequential memcpy from one thread satisfies 1 through 4 by construction. CURRICULUM-DERIVED from 23.2, which simulates four workload models rather than one: the streaming model is the one that approaches peak, and the belief is that model's result with the other three discarded.
Condition 5 is the one nobody lists, and it is the one that fails silently. DERIVED: the requester, the interconnect and the controller each have a width and a clock, and the smallest of them binds — CURRICULUM-DERIVED from 29.2 §7, whose recorded finding is that both layers satisfy their own contracts and only a per-source measurement reveals the gap.
6. The Boundary, Computed
ILLUSTRATIVE attribution over a 4,096-cycle window, 64-bit bus, bytes per cycle = 16 (STRUCTURAL): refresh 240, row work 512, turnaround 180, controller-blocked 960; moved 25,600 bytes; useful 17,408.
| Level | What it is | Bytes/window | Index | Owner of the loss to here |
|---|---|---|---|---|
| 1 | peak — the interface | 65,536 | 1.00 | — |
| 2 | minus refresh | 61,696 | 0.94 | the device and the standard |
| 3 | minus locality and turnaround | 50,624 | 0.77 | the requester's access structure |
| 4 | measured — what was delivered | 25,600 | 0.39 | the controller's policies |
| 5 | application-visible useful bytes | 17,408 | 0.27 | the burst-size decision and the fabric |
DERIVED: ratio_1_to_5 = 65,536 / 17,408 = 3.76×. CURRICULUM-DERIVED from 33.5 §13, which computes the same five levels from an attribution and whose bandwidth_level module this table is the output of. This chapter consumes that instrument and does not rebuild it.
It is worth saying what the index column is not. DERIVED: 0.27 is not an efficiency in 12.4's sense — it is a ratio between two levels, and it contains all four of that chapter's measures plus the fabric. CURRICULUM-DERIVED from 12.4's requirement that the measures stay distinct: collapsing them into 0.27 is legitimate for a sizing ratio and illegitimate for a diagnosis, because a single ratio cannot tell you which of the five reductions to attack. §13's model uses it for sizing and §14's table attributes it per level, and those are different uses of the same number.
The boundary is not one exit but four, and they are not equal. DERIVED: the two largest reductions — 38 points at level 4 and 12 at level 5 — belong to the controller and the burst-size decision, and both are inside the design. Refresh, at 6 points, is the only one that is genuinely not negotiable. So the belief's implicit claim that the gap is unavoidable overhead is also wrong, in the direction that matters: most of the gap is decisions.
7. One Path, Four Reductions
The diagram earns its place because the losses occur at different places in a topology, and a table of percentages cannot show that the smallest link binds.
Read the diagram right-to-left, which is the direction a sizing decision reasons in and the direction that makes the error visible. DERIVED: peak labels only the controller-to-device link. Every stage to its left carries less, and the requester sees the smallest. The belief reads the label on one edge and attributes it to the whole path.
8. The Sizing Decision, With and Without a Stated Ratio
This is the artifact 33.5 has no reason to build, because a review happens after the decision and this is the decision.
// ROBUST MODEL: the sizing computes peak from verified parameters, then
// applies an EXPLICIT level-1-to-level-5 ratio whose provenance is
// carried with it. The ratio may be assumed -- it may not be absent.
module channel_sizing #(
parameter int BUS_BYTES = 8, // STRUCTURAL: 64-bit channel
parameter int RATIO_X100_DEFAULT = 100 // 1.00 = the belief
)(
input logic clk,
input logic rst_n,
input logic size_it,
input logic [15:0] mt_per_s_div100, // e.g. 64 for 6400 MT/s
input logic [31:0] required_mbs, // the application's requirement
input logic [15:0] ratio_x100, // level1/level5, x100
input logic [1:0] ratio_grade, // 0 measured 1 modelled 2 assumed
output logic result_valid,
output logic [31:0] peak_mbs,
output logic [31:0] usable_mbs,
output logic [7:0] channels_needed,
output logic [15:0] headroom_x100,
output logic ratio_is_stated,
output logic ratio_is_assumed_unity
);
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
result_valid <= 1'b0; peak_mbs <= '0; usable_mbs <= '0;
channels_needed <= '0; headroom_x100 <= '0;
ratio_is_stated <= 1'b0; ratio_is_assumed_unity <= 1'b0;
end else if (size_it) begin
// STRUCTURAL: MT/s x bytes = MB/s. Computed, not quoted -- so a
// change of bin or width changes the answer, which is 34.1
// section 8's rebuild test applied to a bandwidth figure.
automatic logic [31:0] pk = 32'(mt_per_s_div100) * 32'd100
* 32'(BUS_BYTES);
// Level 5 from level 1 by the STATED ratio. A ratio of 100 is
// the belief, and it is permitted -- it is just no longer silent.
automatic logic [31:0] us = (ratio_x100 == '0) ? pk
: (pk * 32'd100) / 32'(ratio_x100);
peak_mbs <= pk;
usable_mbs <= us;
// Channels from the USABLE figure, rounded up.
channels_needed <= (us == '0) ? 8'hFF
: 8'((required_mbs + us - 32'd1) / us);
headroom_x100 <= (required_mbs == '0) ? 16'hFFFF
: 16'((us * 32'd100) / required_mbs);
// The two flags a reviewer reads first.
ratio_is_stated <= (ratio_x100 != '0);
ratio_is_assumed_unity <= (ratio_x100 == 16'd100)
&& (ratio_grade == 2'd2);
result_valid <= 1'b1;
end else begin
result_valid <= 1'b0;
end
end
endmodule// INTENTIONALLY DEFECTIVE. WEAK MODEL: divide the requirement by peak.
//
// // sizing.xlsx
// // B4: =6400*8 -> 51200 MB/s
// // B9: =B4/21400 -> 2.39 "2.4x headroom"
// // B10: =ROUNDUP(21400/B4) -> 1 "one channel"
// //
// channels_needed <= 8'((required_mbs + pk - 32'd1) / pk);
// headroom_x100 <= 16'((pk * 32'd100) / required_mbs);
// // and there is no ratio input at all
//
// CONTRACT VIOLATED: 23.2's central law -- peak is a property of the
// interface and the requirement is a property of the application. The
// arithmetic in B4 is exactly right.
//
// WHY IT SURVIVES: peak is verifiable, vendor-published, and the ONLY
// level that exists at sizing time -- 33.5 section 13. So the weak model is
// not using a worse number than it could; it is using the only number
// there is, and omitting the one field that would say so.
//
// TRACE (ILLUSTRATIVE, DDR5-6400, 64-bit, requirement 21,400 MB/s):
// peak = 6400 x 8 = 51,200 MB/s (STRUCTURAL)
//
// weak: usable = 51,200
// headroom 2.39x, channels_needed 1
//
// robust, ratio STATED as measured 3.76 (grade 0):
// usable = 51,200 / 3.76 = 13,617 MB/s
// headroom 0.64x, channels_needed 2
//
// robust, ratio ASSUMED at 1.00 (grade 2):
// usable = 51,200, headroom 2.39x, channels_needed 1
// ratio_is_assumed_unity = 1 <-- THE FINDING
//
// gap: the third case produces the SAME ANSWER as the weak model and
// raises a flag. That is the whole design of this item: the belief is
// not forbidden, it is made VISIBLE as an assumption of 1.0 at
// evidence grade "assumed".
//
// and the channel count: 1 versus 2. The requirement is 21,400 and
// the honest usable figure is 13,617, so one channel is 0.64x of the
// requirement -- it cannot meet it at all.The third case is the item, and it is why this chapter's model permits the belief rather than blocking it. DERIVED: ratio_x100 = 100 at grade assumed produces the identical channel count and sets ratio_is_assumed_unity — so the correction is not a different number, it is a flag on the same number.
CURRICULUM-DERIVED from 18.4 §1's four evidence grades: a peak figure is grade A — documented for a named configuration — and using it as the usable figure consumes it at grade B. DERIVED: ratio_grade is the field that makes the drift visible, and it is one enum.
9. The Belief Compounds Across Levels
A single application of the belief costs 3.76×. Applied at each layer of a system by the team that owns that layer, it compounds — and this is the failure mode that produces a system nobody can explain.
// ROBUST MODEL: each layer publishes what it can DELIVER to the layer
// above, so the composition is a product of measured factors rather
// than a chain of peak figures.
module compounding_headroom #(
parameter int NLAYER = 4, // 0 device 1 controller 2 fabric 3 requester
parameter int SCALE = 100
)(
input logic clk,
input logic rst_n,
input logic compose,
input logic [31:0] device_peak_mbs,
// each layer's efficiency toward the layer above, x100
input logic [7:0] eff_x100 [0:NLAYER-1],
input logic [3:0] layer_states_peak, // bitmask: layer quotes peak upward
output logic result_valid,
output logic [31:0] delivered_mbs,
output logic [31:0] claimed_mbs,
output logic [15:0] overclaim_x100,
output logic [2:0] layers_claiming_peak,
output logic composition_is_multiplicative
);
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
result_valid <= 1'b0; delivered_mbs <= '0; claimed_mbs <= '0;
overclaim_x100 <= '0; layers_claiming_peak <= '0;
composition_is_multiplicative <= 1'b0;
end else if (compose) begin
automatic logic [31:0] d = device_peak_mbs;
automatic logic [31:0] c = device_peak_mbs;
// TRUTH: the factors multiply -- 23.2's measured finding, and
// 33.5 section 11's 65.6% versus 60% for four 90% factors.
for (int L = 0; L < NLAYER; L++)
d = (d * 32'(eff_x100[L])) / 32'(SCALE);
// CLAIM: a layer that quotes peak upward contributes a factor of
// 1.00 instead of its real efficiency.
for (int L = 0; L < NLAYER; L++)
c = layer_states_peak[L] ? c
: ((c * 32'(eff_x100[L])) / 32'(SCALE));
delivered_mbs <= d;
claimed_mbs <= c;
overclaim_x100 <= (d == '0) ? 16'hFFFF : 16'((c * 32'd100) / d);
layers_claiming_peak <= 3'($countones(layer_states_peak));
composition_is_multiplicative <= 1'b1;
result_valid <= 1'b1;
end else begin
result_valid <= 1'b0;
end
end
endmodule THE COMPOUNDING, MEASURED (ILLUSTRATIVE, device peak 51,200 MB/s)
layer efficiency toward the layer above
device 0.94 refresh
controller 0.51 policy, locality, turnaround
fabric 0.85 width and clock mismatch
requester 0.80 payload -- bytes wanted of bytes moved
TRUTH, all four multiplied:
51,200 x 0.94 x 0.51 x 0.85 x 0.80 = 16,690 MB/s
ONE layer quotes peak upward (the controller):
claimed = 51,200 x 0.94 x 1.00 x 0.85 x 0.80 = 32,726
overclaim 1.96x
TWO layers quote peak (controller and fabric):
claimed = 51,200 x 0.94 x 1.00 x 1.00 x 0.80 = 38,502
overclaim 2.31x
ALL FOUR quote peak:
claimed = 51,200 overclaim 3.07x
DERIVED: each layer that adopts the belief multiplies the
over-claim by the reciprocal of its own efficiency. The error is
not additive across teams -- it compounds, exactly as the real
efficiencies do.DERIVED: the over-claim reaches 3.07× with four teams each making a locally defensible statement. No team lied and no arithmetic is wrong — each reported what its own link can carry. CURRICULUM-DERIVED from 29.2 §7's recorded finding: both layers satisfy their own contracts and only a per-source measurement reveals the gap, and here the gap is the product of four such contracts.
The mechanism of the compounding deserves one more line, because it is not intuitive that the errors multiply rather than add. DERIVED: a layer that quotes peak upward is not adding an error term — it is replacing its own factor with 1.00, so the claimed product loses that factor entirely. The over-claim is therefore the reciprocal of the product of the omitted factors: omitting 0.51 gives 1.96×, omitting 0.51 and 0.85 gives 2.31×, and omitting all four gives 3.07×. CURRICULUM-DERIVED from 23.2: the costs compose multiplicatively, so the errors in a multiplicative composition compose multiplicatively too.
And the compounding is the reason this belief survives an integration. DERIVED: a system that misses its bandwidth target by 3× has four teams who can each demonstrate that their layer meets its specification — which is 33.7 §7's three-mutually-exclusive-confirmations pattern, distributed across an organisation.
10. The Ratio's Provenance
The belief's correction is a ratio, so the ratio needs the same discipline this curriculum applies to every other number.
// ROBUST MODEL: the ratio carries a grade and an operating point, and a
// query outside the point is refused -- 33.5 section 14's envelope test,
// applied to the one number a sizing decision cannot measure.
module ratio_provenance #(
parameter int NREC = 4
)(
input logic clk,
input logic rst_n,
input logic record,
input logic [15:0] r_ratio_x100,
input logic [1:0] r_grade, // 0 measured 1 modelled 2 assumed
input logic [7:0] r_load_pct,
input logic [7:0] r_seq_pct, // sequential fraction of the mix
input logic query,
input logic [7:0] q_load_pct,
input logic [7:0] q_seq_pct,
output logic answer_valid,
output logic [15:0] answer_ratio_x100,
output logic [1:0] answer_grade,
output logic in_envelope,
output logic [7:0] refusals,
output logic grade_drift
);
logic [15:0] v_ratio [0:NREC-1];
logic [1:0] v_grade [0:NREC-1];
logic [7:0] v_load [0:NREC-1];
logic [7:0] v_seq [0:NREC-1];
logic v_used [0:NREC-1];
logic [$clog2(NREC+1)-1:0] wptr;
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
wptr <= '0; answer_valid <= 1'b0; answer_ratio_x100 <= '0;
answer_grade <= 2'd2; in_envelope <= 1'b0; refusals <= '0;
grade_drift <= 1'b0;
for (int i = 0; i < NREC; i++) v_used[i] <= 1'b0;
end else begin
answer_valid <= 1'b0;
if (record && (wptr != NREC[$clog2(NREC+1)-1:0])) begin
v_ratio[wptr] <= r_ratio_x100; v_grade[wptr] <= r_grade;
v_load[wptr] <= r_load_pct; v_seq[wptr] <= r_seq_pct;
v_used[wptr] <= 1'b1; wptr <= wptr + 1'b1;
end
if (query) begin
automatic bit found = 1'b0;
for (int i = 0; i < NREC; i++) begin
if (v_used[i] && !found) begin
// A ratio measured on a sequential mix does not describe a
// random one. The envelope is the MIX, not the load.
if ((q_load_pct <= v_load[i] + 8'd10)
&& (q_seq_pct <= v_seq[i] + 8'd10)
&& (q_seq_pct + 8'd10 >= v_seq[i])) begin
answer_ratio_x100 <= v_ratio[i];
answer_grade <= v_grade[i];
in_envelope <= 1'b1;
answer_valid <= 1'b1; found = 1'b1;
end
end
end
if (!found) begin
// The refusal is the useful output: it names a
// characterisation somebody must schedule, and it forces the
// ratio to be recorded as ASSUMED rather than measured.
answer_ratio_x100 <= 16'd100;
answer_grade <= 2'd2;
in_envelope <= 1'b0;
refusals <= refusals + 1'b1;
answer_valid <= 1'b1;
end
end
// Grade drift: an assumed or modelled ratio being consumed where
// the decision's own record claims a measurement.
if (answer_valid && !in_envelope && (answer_grade != 2'd2))
grade_drift <= 1'b1;
end
end
endmoduleDERIVED: the refusal returns a ratio of 1.00 at grade assumed, which is exactly the belief — and it returns it labelled. That is the design: a sizing decision that cannot obtain a measured ratio must proceed, and the only thing this chapter asks is that the row say assumed.
CURRICULUM-DERIVED from 33.5 §14, whose envelope test refuses rather than extrapolates and whose finding was that “unmeasured, 47 points outside” is actionable where a bare number is not. DERIVED: here the envelope's axis is the mix rather than the load — a ratio measured on a sequential stream says nothing about a random one, which is 23.3's 16× row-work spread arriving as a provenance constraint.
11. Where the Belief Breaks First
| Rank | The reduction | Why it comes first | Size |
|---|---|---|---|
| 1 | payload — bytes wanted of bytes moved | any non-contiguous access pattern | up to 2× |
| 2 | policy — controller blocked cycles | any second requester | 38 points |
| 3 | locality — row work | two interleaved streams | 17 points |
| 4 | fabric width or clock | present from the first integration | 15 points |
| 5 | refresh | always, and bounded | 6 points |
DERIVED: rank 1 is first and largest and it is the one absent from every level table, including 33.5 §13's, where it is folded into level 5. CURRICULUM-DERIVED from 12.4, which owns payload efficiency as one of four measures that must stay distinct: a 64-byte burst delivering 32 wanted bytes is 50% payload-efficient, and no timing parameter appears in that number.
And rank 5 is last, which is worth noting against 34.2. DERIVED: refresh is the smallest bandwidth reduction and 34.2 established it is a serious latency exposure — so the two chapters rank the same mechanism first and last on two different axes, and both rankings are right.
12. The Largest Reduction Nobody Models
§11 ranks payload first and largest — up to 2× on its own — and it is the only reduction with no timing parameter in it, which is why it is absent from every table built from a datasheet.
CURRICULUM-DERIVED from 12.4, which owns payload efficiency as one of four measures that must stay distinct: “collapsing them into a single percentage is how architectural arguments go wrong”. And from 29.2 §7, whose finding is that a gap between two conforming layers shows only in a per-source measurement.
// ROBUST MODEL: payload and fabric as two independent reductions, both
// computed from widths and clocks rather than assumed away. Neither
// contains a timing parameter, which is why a datasheet-driven estimate
// misses both.
module payload_and_fabric #(
parameter int MEM_BYTES = 8, // STRUCTURAL: memory bus width
parameter int BL = 16, // beats per burst
parameter int FAB_BYTES = 16, // ILLUSTRATIVE fabric width
parameter int MEM_MT_DIV100 = 64, // 6400 MT/s
parameter int FAB_MHZ = 1200 // ILLUSTRATIVE fabric clock
)(
input logic clk,
input logic rst_n,
input logic evaluate,
input logic [7:0] wanted_bytes, // of the burst, how many are used
output logic result_valid,
output logic [15:0] burst_bytes,
output logic [7:0] payload_x100,
output logic [31:0] mem_mbs,
output logic [31:0] fab_mbs,
output logic [31:0] binding_mbs,
output logic fabric_binds,
output logic [7:0] fabric_x100
);
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
result_valid <= 1'b0; burst_bytes <= '0; payload_x100 <= '0;
mem_mbs <= '0; fab_mbs <= '0; binding_mbs <= '0;
fabric_binds <= 1'b0; fabric_x100 <= '0;
end else if (evaluate) begin
// STRUCTURAL: bytes per burst = BL x width.
automatic logic [15:0] bb = 16'(BL) * 16'(MEM_BYTES);
// Memory link: MT/s x width. Fabric: MHz x width, ONE transfer
// per clock -- the asymmetry that makes a wider fabric slower.
automatic logic [31:0] mm = 32'(MEM_MT_DIV100) * 32'd100
* 32'(MEM_BYTES);
automatic logic [31:0] fm = 32'(FAB_MHZ) * 32'(FAB_BYTES);
burst_bytes <= bb;
// Payload: wanted of moved. No timing parameter appears here,
// which is the whole reason it is missing from datasheet tables.
payload_x100 <= (bb == '0) ? 8'd0
: 8'((32'(wanted_bytes) * 32'd100) / 32'(bb));
mem_mbs <= mm;
fab_mbs <= fm;
// The smallest link binds, and it is not always the memory --
// section 7's diagram, as arithmetic.
binding_mbs <= (fm < mm) ? fm : mm;
fabric_binds <= (fm < mm);
fabric_x100 <= (mm == '0) ? 8'd0
: 8'((((fm < mm) ? fm : mm) * 32'd100) / mm);
result_valid <= 1'b1;
end else begin
result_valid <= 1'b0;
end
end
endmodule// INTENTIONALLY DEFECTIVE. WEAK MODEL: the fabric is wider, so ignore it.
//
// // "the fabric is 128-bit and the memory is 64-bit, so the fabric
// // has 2x the headroom and cannot be the bottleneck"
// binding_mbs <= mm; // <-- the defect
// fabric_binds <= 1'b0;
// // and payload is not computed at all
//
// CONTRACT VIOLATED: none stated -- width comparison is a real and
// usually-correct heuristic. The defect is that bandwidth is width
// TIMES RATE and only one factor was compared.
//
// TRACE (ILLUSTRATIVE, 64-byte bursts of which 32 bytes are wanted):
// memory: 6400 MT/s x 8 B = 51,200 MB/s
// fabric: 1200 MHz x 16 B = 19,200 MB/s <-- and it is WIDER
//
// robust: burst_bytes 128, payload 25%,
// binding 19,200, fabric_binds 1, fabric_x100 37
// weak: binding 51,200, fabric_binds 0, payload not computed
//
// gap: the fabric carries 37% of the memory's rate while being twice
// as wide, and the weak model reports the memory as the binding
// link. Combined with a 25% payload, the application receives
// 19,200 x 0.25 = 4,800 MB/s against a quoted 51,200 -- a ratio of
// 10.7x rather than section 6's 3.76x.
//
// DERIVED: this is the worst case in the chapter, and both of its
// factors are absent from every datasheet-derived estimate.The 10.7× figure is the chapter's worst case and it is built from two reductions that contain no timing parameter at all. DERIVED: payload is bytes-wanted over bytes-moved and the fabric term is a width-times-clock comparison — so an engineer working entirely from a DDR datasheet cannot arrive at either, however carefully they read it.
And the fabric case is a genuine trap rather than an oversight. DERIVED: the fabric is wider, so it has headroom is a correct comparison of one factor of a product, and the second factor — one transfer per clock against two — is exactly the asymmetry that makes DDR's name meaningful. CURRICULUM-DERIVED from 29.2's framing of the controller/interconnect boundary: the two sides have different clocks and different transfer disciplines, and comparing widths across that boundary compares the wrong quantity.
13. Closing the Loop — Re-Sizing on the First Measurement
§8 permits an assumed ratio and §10 returns one on every refusal. The half that makes this honest rather than merely documented is what happens when a measurement finally exists.
// ROBUST MODEL: the assumed ratio is a PLACEHOLDER with a revision
// path, so the first silicon measurement re-opens the decision instead
// of confirming it.
module ratio_update #(
parameter int SCALE = 100
)(
input logic clk,
input logic rst_n,
input logic assume_ratio,
input logic [15:0] assumed_x100,
input logic measure_ratio,
input logic [15:0] measured_x100,
input logic [31:0] required_mbs,
input logic [31:0] peak_mbs,
output logic [15:0] ratio_in_force,
output logic [1:0] grade_in_force,
output logic [7:0] channels_now,
output logic [7:0] channels_as_assumed,
output logic decision_changed,
output logic revision_needed,
output logic [15:0] assumption_error_x100
);
function automatic logic [7:0] chans(logic [15:0] r);
automatic logic [31:0] us = (r == '0) ? peak_mbs
: (peak_mbs * 32'd100) / 32'(r);
return (us == '0) ? 8'hFF : 8'((required_mbs + us - 32'd1) / us);
endfunction
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
ratio_in_force <= 16'(SCALE); grade_in_force <= 2'd2;
channels_now <= '0; channels_as_assumed <= '0;
decision_changed <= 1'b0; revision_needed <= 1'b0;
assumption_error_x100 <= '0;
end else begin
if (assume_ratio) begin
ratio_in_force <= assumed_x100;
grade_in_force <= 2'd2; // assumed
channels_as_assumed <= chans(assumed_x100);
channels_now <= chans(assumed_x100);
revision_needed <= 1'b1; // an assumption is a debt
end
// A measurement SUPERSEDES an assumption -- never the reverse,
// which is 18.4 section 1's grade ordering made structural.
if (measure_ratio) begin
ratio_in_force <= measured_x100;
grade_in_force <= 2'd0; // measured
channels_now <= chans(measured_x100);
revision_needed <= 1'b0;
// The two numbers a project needs at that moment: did the
// decision change, and by how much was the assumption wrong?
decision_changed <= (chans(measured_x100) != channels_as_assumed);
assumption_error_x100 <=
(ratio_in_force == '0) ? 16'hFFFF
: 16'((32'(measured_x100) * 32'd100) / 32'(ratio_in_force));
end
end
end
endmodule THE LOOP, CLOSED (ILLUSTRATIVE, peak 51,200, requirement 21,400)
pre-silicon, ratio ASSUMED 1.00:
usable 51,200 channels 1 revision_needed 1
first silicon, ratio MEASURED 3.76:
usable 13,617 channels 2 decision_changed 1
assumption_error 376% revision_needed 0
and the cost of having been wrong:
a floorplan sized for one channel cannot become two.
DERIVED: `revision_needed` is the field that makes an assumption a
scheduled obligation rather than a permanent default. Without it, an
assumed 1.00 becomes the project's number by inaction.revision_needed is the item, and it is one bit. DERIVED: an assumption with no revision flag is indistinguishable from a measurement after the person who made it moves on — CURRICULUM-DERIVED from 33.6 §14's assumption ledger, whose finding was six assumptions no gate established and whose remedy was a second column.
And the last line of the trace is why this chapter's belief is the most expensive of the six. DERIVED: §14's first row is a floorplan decision, and a floorplan sized for one channel cannot become two — so unlike every other belief in this module, the cost here cannot be recovered by a later firmware, software or configuration change.
14. The Decision Built on the Belief
| Where it is applied | What it produces | Measured consequence |
|---|---|---|
| A channel count | one channel, 2.4× headroom | 0.64× — cannot meet the requirement |
| A DRAM bin choice | a faster bin buys proportional bandwidth | buys 0.27 × the increment |
| A fabric width | match the memory's width | the fabric becomes the binding link |
| A multi-team budget | each layer quotes its own peak | over-claim 3.07× (§9) |
| A cost model | bandwidth per dollar from peak | ranks parts by a figure nobody receives |
| A regression target | hit 90% of peak | unreachable; the target is 27% |
Row 2 is the most consequential and the least obvious, so it is worth the arithmetic. DERIVED: moving from 6400 to 8000 MT/s raises peak by 25% — 51,200 to 64,000 — and raises application-visible bandwidth by 25% of level 5, which is 4,352 MB/s, not 12,800. The increment is real and it is 0.27 of what the belief predicts, because the four reductions apply to the increment exactly as they apply to the base.
Row 6 is the one that wastes engineering time rather than money. DERIVED: a target of 90% of peak is unreachable on a realistic mix, so a team chases it, improves several real things, and never arrives — CURRICULUM-DERIVED from 33.5 §9's ceiling item: the proposal's ceiling is below the gap it must close, and the gap was never a gap.
15. What the Assertions Prove
// ---- Section 8: sizing. The obligation is that the ratio is STATED,
// not that it is correct -- which is the only obligation a
// pre-silicon decision can carry.
property p_ratio_must_be_stated;
@(posedge clk) disable iff (!rst_n)
result_valid |-> ratio_is_stated;
endproperty
assert property (p_ratio_must_be_stated)
else $error("a sizing result was produced with no level-1-to-5 ratio");
property p_usable_never_exceeds_peak;
@(posedge clk) disable iff (!rst_n)
result_valid |-> (usable_mbs <= peak_mbs);
endproperty
assert property (p_usable_never_exceeds_peak)
else $error("the usable figure exceeded peak: the ratio is below unity");
// Peak must be COMPUTED, so a change of bin changes it -- 34.1
// section 8's rebuild test applied to bandwidth.
property p_peak_tracks_the_bin;
@(posedge clk) disable iff (!rst_n)
size_it |=> (peak_mbs == ($past(mt_per_s_div100) * 32'd100
* 32'(BUS_BYTES)));
endproperty
assert property (p_peak_tracks_the_bin)
else $error("the peak figure did not follow the data rate");
property p_unity_assumption_is_flagged;
@(posedge clk) disable iff (!rst_n)
(size_it && (ratio_x100 == 16'd100) && (ratio_grade == 2'd2))
|=> ratio_is_assumed_unity;
endproperty
assert property (p_unity_assumption_is_flagged)
else $error("a ratio of 1.0 at grade assumed was not flagged");
// Two-sided -- 30.3 section 9's variety 8: a model that refuses every
// ratio satisfies the properties above and sizes nothing.
property p_measured_ratio_is_accepted;
@(posedge clk) disable iff (!rst_n)
(size_it && (ratio_x100 > 16'd100) && (ratio_grade == 2'd0))
|=> (result_valid && !ratio_is_assumed_unity);
endproperty
assert property (p_measured_ratio_is_accepted)
else $error("a measured ratio above unity was rejected");
// ---- Section 9: compounding. The composition law is 23.2's measured
// finding, and asserting it is what rejects an additive estimate.
property p_delivered_is_a_product;
@(posedge clk) disable iff (!rst_n)
result_valid |-> (delivered_mbs <= device_peak_mbs);
endproperty
assert property (p_delivered_is_a_product)
else $error("delivered bandwidth exceeded the device's peak");
property p_claim_at_least_delivered;
@(posedge clk) disable iff (!rst_n)
result_valid |-> (claimed_mbs >= delivered_mbs);
endproperty
assert property (p_claim_at_least_delivered)
else $error("the claimed figure fell below the delivered one");
property p_overclaim_grows_with_claimants;
@(posedge clk) disable iff (!rst_n)
(result_valid && (layers_claiming_peak == 3'd0))
|-> (overclaim_x100 == 16'd100);
endproperty
assert property (p_overclaim_grows_with_claimants)
else $error("no layer quoted peak and the claim still exceeded delivery");
property p_composition_is_not_additive;
@(posedge clk) disable iff (!rst_n) composition_is_multiplicative;
endproperty
assert property (p_composition_is_not_additive)
else $error("the composition was computed additively");
// ---- Section 10: the ratio's provenance.
property p_outside_envelope_is_refused;
@(posedge clk) disable iff (!rst_n)
(answer_valid && !in_envelope) |-> (answer_grade == 2'd2);
endproperty
assert property (p_outside_envelope_is_refused)
else $error("an out-of-envelope ratio was returned at better than assumed");
property p_refusal_returns_unity;
@(posedge clk) disable iff (!rst_n)
(answer_valid && !in_envelope) |-> (answer_ratio_x100 == 16'd100);
endproperty
assert property (p_refusal_returns_unity)
else $error("a refusal extrapolated a ratio instead of returning unity");
property p_no_grade_drift;
@(posedge clk) disable iff (!rst_n) !grade_drift;
endproperty
assert property (p_no_grade_drift)
else $error("a ratio was consumed at a higher grade than it holds");
// And the positive companion, so refusing everything is not a
// passing strategy.
property p_in_envelope_is_answered;
@(posedge clk) disable iff (!rst_n)
(answer_valid && in_envelope) |-> (answer_ratio_x100 != '0);
endproperty
assert property (p_in_envelope_is_answered)
else $error("an in-envelope query returned no ratio");
// ---- Section 12: payload and fabric. The two reductions with no
// timing parameter in them.
property p_binding_link_is_the_smaller;
@(posedge clk) disable iff (!rst_n)
result_valid |-> (binding_mbs == ((fab_mbs < mem_mbs) ? fab_mbs : mem_mbs));
endproperty
assert property (p_binding_link_is_the_smaller)
else $error("the binding link is not the smaller of the two rates");
property p_fabric_width_does_not_imply_headroom;
@(posedge clk) disable iff (!rst_n)
(result_valid && (fab_mbs < mem_mbs)) |-> fabric_binds;
endproperty
assert property (p_fabric_width_does_not_imply_headroom)
else $error("a slower fabric was not reported as binding");
property p_payload_is_a_fraction;
@(posedge clk) disable iff (!rst_n)
result_valid |-> (payload_x100 <= 8'd100);
endproperty
assert property (p_payload_is_a_fraction)
else $error("payload efficiency exceeded 100%");
// ---- Section 13: the revision path. A measurement must supersede an
// assumption, and never the reverse -- 18.4 section 1's grade ordering.
property p_measurement_supersedes_assumption;
@(posedge clk) disable iff (!rst_n)
measure_ratio |=> (grade_in_force == 2'd0);
endproperty
assert property (p_measurement_supersedes_assumption)
else $error("a measurement did not supersede the assumed ratio");
property p_assumption_raises_a_revision_debt;
@(posedge clk) disable iff (!rst_n)
assume_ratio |=> revision_needed;
endproperty
assert property (p_assumption_raises_a_revision_debt)
else $error("an assumed ratio was recorded with no revision flag");
property p_revision_cleared_only_by_measurement;
@(posedge clk) disable iff (!rst_n)
(!measure_ratio && revision_needed) |=> revision_needed;
endproperty
assert property (p_revision_cleared_only_by_measurement)
else $error("a revision debt cleared without a measurement");
// ---- The CLAIM, as a conditional pair -- 34.1 section 1's required form.
property p_claim_holds_in_its_region;
@(posedge clk) disable iff (!rst_n)
(result_valid && (layer_states_peak == 4'b0000)
&& (eff_x100[0] >= 8'd94) && (eff_x100[1] >= 8'd94)
&& (eff_x100[2] >= 8'd94) && (eff_x100[3] >= 8'd94))
|-> (overclaim_x100 <= 16'd120);
endproperty
assert property (p_claim_holds_in_its_region)
else $error("with every layer near unity the over-claim still exceeded 1.2x");
property p_claim_fails_outside_its_region;
@(posedge clk) disable iff (!rst_n)
(result_valid && (eff_x100[1] <= 8'd60) && layer_states_peak[1])
|-> (overclaim_x100 >= 16'd150);
endproperty
assert property (p_claim_fails_outside_its_region)
else $error("a layer at 60% quoting peak did not over-claim by 1.5x");
// ---- COVERS. Each on the dimension the belief's failure scales with.
// The REGION: every layer near unity, so the claim pair is not vacuous.
cover property (@(posedge clk) disable iff (!rst_n)
result_valid && (eff_x100[1] >= 8'd94));
// A layer at REALISTIC efficiency -- the dimension is the WORKLOAD
// MIX, and a sequential-only stimulus never reaches it.
cover property (@(posedge clk) disable iff (!rst_n)
result_valid && (eff_x100[1] <= 8'd60));
// TWO OR MORE layers quoting peak -- the compounding case, which a
// single-team model cannot produce.
cover property (@(posedge clk) disable iff (!rst_n)
result_valid && (layers_claiming_peak >= 3'd2));
// All four, so the 3.07x figure is exercised.
cover property (@(posedge clk) disable iff (!rst_n)
result_valid && (layers_claiming_peak == 3'd4));
// Section 8: a ratio ASSUMED at unity -- the belief, made visible.
cover property (@(posedge clk) disable iff (!rst_n) ratio_is_assumed_unity);
// And a MEASURED ratio above unity, so the honest path is exercised.
cover property (@(posedge clk) disable iff (!rst_n)
size_it && (ratio_x100 > 16'd300) && (ratio_grade == 2'd0));
// A channel count of TWO -- the answer the belief cannot reach.
cover property (@(posedge clk) disable iff (!rst_n)
result_valid && (channels_needed >= 8'd2));
// Headroom BELOW unity: the requirement exceeding one channel.
cover property (@(posedge clk) disable iff (!rst_n)
result_valid && (headroom_x100 < 16'd100));
// Section 10: a query outside the envelope, refused.
cover property (@(posedge clk) disable iff (!rst_n)
answer_valid && !in_envelope);
// And inside it, answered -- so the refusal path is not the only one.
cover property (@(posedge clk) disable iff (!rst_n)
answer_valid && in_envelope);
// A query whose MIX differs from every record's -- the axis 23.3's
// 16x spread makes load-independent.
cover property (@(posedge clk) disable iff (!rst_n)
query && (q_seq_pct < 8'd30));
// A change of data rate, so p_peak_tracks_the_bin is not vacuous.
cover property (@(posedge clk) disable iff (!rst_n)
size_it && (mt_per_s_div100 != $past(mt_per_s_div100)));
// Section 12: the FABRIC binding despite being wider -- the trap, and
// one a same-clock model cannot reach.
cover property (@(posedge clk) disable iff (!rst_n)
result_valid && fabric_binds && (FAB_BYTES > MEM_BYTES));
// A payload below a half, which scatter traffic produces and a
// sequential stimulus never does.
cover property (@(posedge clk) disable iff (!rst_n)
result_valid && (payload_x100 < 8'd50));
// Section 13: a measurement that CHANGES the decision -- the event the
// whole revision path exists for.
cover property (@(posedge clk) disable iff (!rst_n)
measure_ratio && decision_changed);
// And one that does not, so the path is exercised both ways.
cover property (@(posedge clk) disable iff (!rst_n)
measure_ratio && !decision_changed);Three notes on this property set, and the first is the chapter's central design decision.
p_ratio_must_be_stated asserts that the ratio exists, not that it is right — and that is the strongest obligation a pre-silicon artifact can carry. DERIVED: no property here can check the ratio's value, because the measurement that would check it requires the silicon being sized. CURRICULUM-DERIVED from 33.4 §7, whose antecedent-publication item has the same shape: the obligation is on the record, and the fix is a manifest rather than a better property.
p_refusal_returns_unity is the only property in this module that requires a model to produce the belief's answer. DERIVED: on a refusal the model returns 1.00 at grade assumed, which is numerically the weak build's output — because the alternative, extrapolating a ratio nobody measured, is worse than an honest assumption. CURRICULUM-DERIVED from 33.5 §14: a refusal that names what is unmeasured is actionable; a fabricated figure is not.
And variety 12 governs the claim pair's thresholds. DERIVED: p_claim_holds_in_its_region uses 94 and 120, and both are chosen rather than derived — so §16's mutations on those constants are unkillable by construction, which is the fourth consecutive chapter in which the survivors are the numbers inside the properties.
16. Mutation Testing
Baseline first: all twenty-two assertions pass and all sixteen covers are non-zero.
| # | Mutation | Killed by | Survived? |
|---|---|---|---|
| M1 | §8: drop the ratio_x100 input | p_ratio_must_be_stated | killed |
| M2 | §8: usable_mbs <= peak_mbs directly | p_usable_never_exceeds_peak | killed* |
| M3 | §8: peak_mbs from a lookup table | p_peak_tracks_the_bin, by one cover | killed |
| M4 | §8: drop the unity flag | p_unity_assumption_is_flagged, by one cover | killed |
| M5 | §8: refuse every ratio | p_measured_ratio_is_accepted | killed |
| M6 | §8: channels from peak not usable | p_usable_never_exceeds_peak via headroom | killed |
| M7 | §9: sum the losses instead of multiplying | p_composition_is_not_additive | killed |
| M8 | §9: claimed_mbs <= delivered_mbs | p_claim_at_least_delivered | killed |
| M9 | §9: ignore layer_states_peak | p_overclaim_grows_with_claimants, by one cover | killed |
| M10 | §9: allow delivered > device_peak | p_delivered_is_a_product | killed |
| M11 | §10: extrapolate on a refusal | p_refusal_returns_unity, by one cover | killed |
| M12 | §10: return the nearest record regardless | p_outside_envelope_is_refused | killed |
| M13 | §10: drop the grade-drift check | p_no_grade_drift | killed |
| M14 | §10: envelope on load only, not mix | nothing | SURVIVES |
| M15 | §15: the 94 threshold lowered to 50 | nothing | SURVIVES |
| M16 | §15: the 120 over-claim bound raised to 400 | nothing | SURVIVES |
| M17 | RATIO_X100_DEFAULT left at 100 | nothing | SURVIVES |
| M18 | the eff_x100 set fixed at all-100 | nothing | SURVIVES |
| M19 | §12: binding_mbs <= mem_mbs always | p_binding_link_is_the_smaller, by one cover | killed |
| M20 | §12: compare widths instead of rates | p_fabric_width_does_not_imply_headroom | killed |
| M21 | §12: payload against wanted_bytes | p_payload_is_a_fraction | killed |
| M22 | §13: an assumption clears the revision flag | p_assumption_raises_a_revision_debt | killed |
| M23 | §13: an assumption supersedes a measurement | p_measurement_supersedes_assumption | killed |
| M24 | §12: FAB_MHZ raised until the fabric never binds | nothing | SURVIVES |
DERIVED: eighteen of twenty-four killed, six survived — the largest survivor count in this module so far, and every one is a number inside a check rather than a mechanism.
M24 is worth separating from the other five because it is the one a real project performs on itself. Raising FAB_MHZ until the fabric never binds makes p_fabric_width_does_not_imply_headroom vacuous — its antecedent stops occurring — and the cover on fabric_binds && (FAB_BYTES > MEM_BYTES) is what detects it. DERIVED: that is variety 6 reached through a parameter rather than through a stimulus, which 34.1 §17's M16 found in the other direction: a quantity can make a correct, previously-reachable property unreachable, and no antecedent cover changes unless somebody is watching the parameter.
M14 is the one worth acting on and it is not a threshold. Restricting the envelope to load and dropping the mix makes a ratio measured on a sequential stream answer a query about a random one — and every property passes, because the envelope test still runs. CURRICULUM-DERIVED from 23.3's finding that two mappings with identical class counts can differ 16× in row-work cycles: the mix is the axis the ratio actually depends on, and an envelope on the wrong axis is 33.4 §8's coverage-dimension error applied to a provenance record.
M17 and M18 reconstruct the belief's region, as in the two preceding chapters. A default ratio of 100 is the belief; an all-unity efficiency set is §5's region. DERIVED: both make the belief true and kill nothing — the third and fourth instances of this module's structural result.
M15 and M16 are variety 12 at its purest: 94 and 120 are the chapter's own choices, and no property can adjudicate them. CURRICULUM-DERIVED from 33.4 §15. DERIVED: the fix is that the region's thresholds come from the same attribution 33.5 §7 demands — a measured efficiency per layer — which is a characterisation report and not a property.
Four mutations are killed only by a cover, and M4's is the sharpest. Dropping the unity flag changes nothing on any shortlist where the ratio was measured, so the cover on ratio_is_assumed_unity is the only thing that reaches it — and a project that always has a measured ratio does not need this chapter.
17. Baseline Defects Found Before Mutation
| Belief applied to | Caught by | At what cost |
|---|---|---|
| a channel count | p_ratio_must_be_stated | nothing — the field is absent |
| a peak figure quoted from memory | p_peak_tracks_the_bin | nothing — change the bin and ask again |
| a multi-team budget | p_overclaim_grows_with_claimants | each layer's efficiency, measured |
| an additive loss estimate | p_composition_is_not_additive | nothing — one multiplication |
| a ratio reused across mixes | p_outside_envelope_is_refused | a second mix, characterised |
| a regression target | — | nothing: compare the target to level 5 |
DERIVED: four of six are found without running anything, and the cheapest is the second — quote a peak figure, change the data rate, and ask for the figure again. CURRICULUM-DERIVED from 34.1 §8's rebuild test: a number that does not move when its inputs move was remembered rather than derived.
And the sixth row is the cheapest finding in the chapter, so it deserves the emphasis a table cannot give it. DERIVED: take the regression target, divide it by peak, and compare it to the ratio — if the target is 90% of peak and the ratio is 3.76, the target is 3.4× above the reachable figure. CURRICULUM-DERIVED from 33.5 §9's ceiling item: a proposal whose ceiling is below the gap it must close is rejected on arithmetic, and a target above the reachable maximum is the same rejection applied to a goal. One division, and it retires a target that a team may have been chasing for a year.
Two need a measurement that does not exist at decision time, and that is this chapter's defining difficulty.
| Application | The measurement | Why it is unavailable |
|---|---|---|
| a multi-team budget | per-layer efficiency | each layer is un-integrated when its budget is set |
| a ratio across mixes | a second characterised mix | characterisation happens on the silicon being sized |
DERIVED: neither is obtainable pre-silicon, so for both the honest output is §10's refusal — a ratio of 1.00 recorded at grade assumed, plus a named characterisation somebody must schedule. CURRICULUM-DERIVED from 33.3 §15's finding that a review must also produce a list of artifacts it could not obtain: this chapter's list has exactly two entries and both are known in advance.
18. Silicon Observability
| What silicon shows | What it says about the belief |
|---|---|
| a channel at 100% utilisation with the application unsatisfied | §6 — the sizing used level 1 |
| a faster bin delivering a quarter of its predicted increment | §14 row 2 — the reductions apply to the increment |
| four teams each meeting their own specification, system missing target | §9 — the compounding |
| a bandwidth regression stuck at ~27% of peak | §14 row 6 — the target was never reachable |
| a peak figure that matches the datasheet exactly | the belief's evidence, and it is correct |
| the ratio changing when the workload mix changes | §10 — the envelope's axis is the mix |
Row 2 is the most diagnostic because it is quantitative and cheap. DERIVED: raise the data rate 25% and measure the application-visible increment; if it is near 25% you are in §5's region, and if it is near 7% the ratio is about 3.8. That single experiment measures the ratio directly, and it is the one experiment that becomes available the moment any silicon exists.
Rows 5 and 6 belong to the belief's two halves and it is worth saying which is which. DERIVED: row 5 — a peak figure matching the datasheet — confirms the belief's premise, and it will confirm it forever, because the premise is true. Row 6 refutes the belief's conclusion, and it only appears once a workload changes. So a project that never changes workload never sees a disconfirming observation, and the belief is stable by construction.
Row 6 is the one that invalidates a stored ratio. CURRICULUM-DERIVED from 23.3 and §10's envelope: a ratio is a function of the mix, so a system whose workload changes has a stale ratio — and that is 31.1 §14's variety 10, parameter-conditional soundness, applied to a planning constant.
19. Quantitative Reasoning
| Quantity | Truth | Under the belief | Gap | Provenance |
|---|---|---|---|---|
| peak, DDR5-6400 ×64 | 51,200 MB/s | 51,200 | none | STRUCTURAL identity |
| application-visible | 13,617 MB/s | 51,200 | 3.76× | DERIVED from 33.5 §13's levels |
| headroom on one channel | 0.64× | 2.39× | 3.7× | DERIVED |
| channels needed | 2 | 1 | the decision | DERIVED |
| increment from a 25% faster bin | 4,352 MB/s | 12,800 | 2.9× | DERIVED |
| over-claim, one layer quoting peak | 1.96× | 1.00× | — | DERIVED |
| over-claim, four layers | 3.07× | 1.00× | — | DERIVED |
| four 90% factors composed | 65.6% | 60% additive | 5.6 pts | DERIVED, 33.5 §11 |
| reachable fraction of peak | ~27% | 90% target | unreachable | DERIVED |
Sort them by whether the belief names the quantity, as 34.2 §16 did, and this chapter sits between the two before it.
| Does the belief name it? | Quantities | Is the belief right? |
|---|---|---|
| Yes | peak | exactly right |
| Implicitly, as equal to peak | application-visible, headroom, channels, increment | wrong by 2.9× to 3.8× |
| No | over-claim, composition law, reachable fraction | absent |
DERIVED: the belief names one quantity and gets it exactly right; it implies four and gets them wrong by ~3×; and it is silent on three. That middle row is new to this module — 34.1's belief was wrong about a magnitude it named and 34.2's was silent about everything it got wrong. Here the error is an identification: two different quantities asserted to be one, which is why the belief is a sentence with an equals sign in it.
20. The Beliefs This One Generates
| Downstream belief | Why it follows | Where it is refuted |
|---|---|---|
| “a faster bin buys proportional bandwidth” | if peak is what you receive | §14 row 2 |
| “losses add up to about 10%” | if the reductions are small and additive | §9; 33.5 §11 |
| “90% of peak is a reasonable target” | if peak is achievable | §14 row 6 |
| “the memory is the bottleneck” | if the memory's number is the system's | §7; the fabric may bind |
| “bandwidth per dollar ranks parts” | if peak is the delivered figure | §14 row 5 |
The fourth row is the most expensive because it misdirects an entire investigation. DERIVED: if peak is taken as what the system receives, then a shortfall must be the memory's fault — and §7's diagram shows the fabric and the requester each carrying less. CURRICULUM-DERIVED from 33.5 §9's owner_cycles finding: two of six attribution components belong to the requester, so the largest available lever is often outside the memory — and this belief guarantees nobody looks there.
And the second row is the one that makes the first three survivable for years. DERIVED: an additive loss estimate of ~10% is close enough to be plausible and wrong in the conservative direction, so it produces designs that miss by a little and are patched — CURRICULUM-DERIVED from 33.5 §11: the additive law is conservative about efficiency and therefore trusted, and at 70% factors it predicts a negative efficiency.
21. Common Wrong Answers
-
“DDR5-6400 ×64 is 51.2 GB/s.” Correct, and it is level 1 of five. CURRICULUM-DERIVED from 33.5 §13: the question is never whether the figure is right but which level it is.
-
“So peak is useless.” Inverted, and this is the over-correction. DERIVED: peak is the only level that exists pre-silicon and every other level is computed from it — §8's model takes it as its first input. The fix is a ratio, not a different base.
-
“Use 70% of peak as a rule of thumb.” A rule of thumb is an assumed ratio, which is progress — if it is labelled. DERIVED: §10's model accepts 1.43 at grade assumed and flags nothing; what it rejects is an unstated ratio. And 70% is 0.70 where the measured figure was 0.27.
-
“The losses are about 10%.” §9: four realistic factors compose to 0.33, not 0.90. CURRICULUM-DERIVED from 23.2's measured finding that the costs compose multiplicatively.
-
“Subtract refresh and you are close.” Refresh is 6 points of a 73-point gap — §11 ranks it last. DERIVED: the two largest reductions are policy (38) and payload (up to 2×), and both are inside the design.
-
“We hit 90% of peak in our benchmark.” Then the benchmark is in §5's region, and the belief is correct about the benchmark. DERIVED: conditions 1 through 4 are what a sequential
memcpysatisfies by construction. -
“Our workload is mostly sequential too.” Mostly is the word doing the work. CURRICULUM-DERIVED from 23.3: identical class counts can differ 16× in row-work cycles, so mostly sequential is not a number and the ratio needs one.
-
“The controller will reach peak if we tune it.” Policy is 38 points and it is the largest single lever — so tuning helps, and its ceiling is 38 points of 73. CURRICULUM-DERIVED from 33.5 §9: compute the ceiling before proposing.
-
“A faster bin will close the gap.” §14 row 2: a 25% faster bin delivers a 25% larger level 5, which is 4,352 MB/s and not 12,800. The reductions apply to the increment.
-
“Each team met its spec, so the system should meet its target.” §9: four layers each meeting a peak-denominated spec over-claim by 3.07×. CURRICULUM-DERIVED from 29.2 §7: both layers satisfy their own contracts and only a per-source measurement shows the gap.
-
“We cannot measure the ratio before we have silicon, so this is unactionable.” That is exactly §10's case and its output is a flag, not a number. DERIVED: recording 1.00 at grade assumed costs one enum and makes the decision revisable.
-
“An assumed ratio is just a guess.” A labelled guess is revisable and an unlabelled one is a fact. CURRICULUM-DERIVED from 18.4 §1: the failure mode is category drift, and a grade field is what prevents it.
-
“Once we measure the ratio we are done.” The ratio is a function of the mix. DERIVED: §16's M14 is exactly this — an envelope on load rather than mix — and it survives every property in the chapter.
-
“Payload efficiency is a software problem.” It is a burst-size decision, which is the controller's and the ISA's jointly. CURRICULUM-DERIVED from 12.4, which owns the four measures, and 34.1 §11, whose BC4 trace shows payload efficiency doubling and command occupancy rising 80%.
-
“The fabric is wider than the memory, so it cannot bind.” Width is not bandwidth — a wider fabric at a lower clock can carry less. DERIVED: §9's fabric factor is 0.85 and it is a width and clock product.
-
“We will find out in bring-up.” The channel count is a floorplan decision and it is made before any silicon exists. DERIVED: §14's first row is the only one in this module whose cost cannot be recovered by a later software or firmware change.
-
“51.2 GB/s versus 13.6 GB/s — somebody is lying.” Both are correct measurements of different quantities. CURRICULUM-DERIVED from 23.2's central law: the gap is the cost of constraints that cannot be removed plus decisions that can be improved, and §6 apportions it 6 points to the first and 67 to the second.
-
“Then the honest number is 13.6 and we should quote that.” On that mix, at that operating point. DERIVED: quoting level 5 as a general figure is the same category drift in the other direction — §10's envelope exists because a level-5 figure is as configuration-bound as a level-1 figure is workload-free.
-
“Our peak figure comes straight from the vendor, so it needs no provenance.” The peak figure needs none — the ratio does. DERIVED: §15's
p_ratio_must_be_statedis the only obligation in this chapter, and it says nothing about peak. -
“Bandwidth is simpler than latency, so it is less error-prone.” It is more tractable and that is why the error persists (§4, reason 3). DERIVED: latency's unbounded queueing term makes a naive latency spreadsheet visibly fail; four bandwidth factors near unity multiply to something that still looks plausible.
-
“We use measure D, which is exhaustive.” Measure D is level 4 against level 1 — CURRICULUM-DERIVED from 23.2 and 12.4: it includes row-state work, turnaround and refresh, and it does not include the payload or fabric reductions that separate level 4 from level 5.
-
“I know all five levels, so I do not hold this belief.” §8's third case is the test: does your sizing spreadsheet have a ratio cell? DERIVED: knowing the levels and having a cell for the ratio are different, and the second is what a decision carries.
22. Self-Check
- Compute peak for a 64-bit channel at 6400 MT/s from the identity, and say which of the five levels it is.
- Name the five levels and the owner of each loss between them.
- A requirement is 21,400 MB/s and peak is 51,200. Give the headroom under the belief and under a measured ratio of 3.76, and give the channel count in each case.
- Why can this belief not be corrected by demanding a better number, and what does §10's model return on a refusal?
- Four layers each report 0.94, 0.51, 0.85 and 0.80 toward the layer above. Give the delivered figure and the over-claim if two layers quote peak instead.
- A 25% faster bin is proposed. Give the true increment in application-visible bandwidth and say why it is not 25% of peak.
- On which axis is a level-1-to-5 ratio's envelope defined, and which chapter's 16× finding makes that the answer?
- M14 restricted the envelope to load and dropped the mix. Name the variety, say why no property fired, and give the fix.
- Two mutations reconstructed the belief's region. Which two, and what is the general statement this module has now made four times?
- A regression target of 90% of peak has been open for a year. State the finding in one sentence, and name the review item that would have closed it.
- Which single quantity does the belief name, and how many does it imply? Say what kind of error that makes it, using §19's three-row table.
- Your sizing spreadsheet has no ratio cell. What is the finding, what does it cost to fix, and what evidence grade should the new cell carry?
23. The Residual Risk
What this chapter cannot settle.
It cannot supply the ratio. That is not a gap in the chapter; it is the chapter's subject. DERIVED: §10's model returns 1.00 at grade assumed on every query it cannot answer, and it will do so for every pre-silicon decision. CURRICULUM-DERIVED from 33.5 §14: a refusal that names the missing characterisation is the best available output, and this chapter's contribution is that the assumption be recorded rather than implicit.
It cannot validate its own thresholds. M15 and M16 moved 94 and 120 and nothing fired. DERIVED: the region's boundaries are chosen here, and the honest source is a per-layer efficiency attribution — 33.5 §7's exhaustive-attribution instrument, which needs silicon. The fourth consecutive chapter whose survivors are the numbers inside its own checks.
It cannot tell you your mix. M18 fixed every efficiency at unity and made the belief true. CURRICULUM-DERIVED from Module 32: the mix is what distinguishes platform classes, so a genuinely sequential single-requester system is inside §5's region and the belief is correct for it.
It cannot tell you where the fabric's factor comes from. §12 computes a fabric rate from a width and a clock, and both are properties of an interconnect this curriculum does not own. CURRICULUM-DERIVED from 29.2, which owns the controller/interconnect boundary and records that real IP boundaries do not match the contract boundaries: DERIVED: so a fabric factor of 0.85 is an input here, obtained from a team whose own estimate may itself be peak-denominated — which is §9's compounding arriving inside this chapter's own model.
And it cannot apportion the gap between irreducible and improvable. §6 attributes 6 points to refresh and 67 to decisions, and §6's attribution is ILLUSTRATIVE. CURRICULUM-DERIVED from 23.2's central law, which names the distinction and requires a measurement to apply it: DERIVED: this chapter establishes that most of the gap is decisions rather than physics, and it cannot tell you which decisions on your system.
24. Where This Goes
Three beliefs down, and the three so far have all been about a number: one over-weighted, one treated as absent, one identified with another.
Chapter 34.4 changes the object. DERIVED: the DDR controller is simple arbitration is not a claim about a quantity at all — it is a claim about a component's nature, and it is the first belief in this module that cannot be refuted with arithmetic.
And it inherits this chapter's largest finding, which is worth carrying explicitly. DERIVED: §6 apportioned the level-1-to-5 gap and found the controller's policy losses the single largest term at 38 of 73 points. CURRICULUM-DERIVED from 23.2's central law — the gap is irreducible constraints plus decisions that can be improved — so the largest improvable term in this chapter belongs to the component the next chapter's belief calls simple.
That is not a coincidence, and it is the next chapter's opening. DERIVED: a component believed to be simple arbitration is a component nobody expects 38 points from, and the two beliefs are therefore mutually supporting — the bandwidth belief hides where the gap is, and the controller belief explains why nobody looks.
Continue learning
Related tutorials
- Related topic
SDR SDRAM
Making DRAM synchronous replaced an analog timing negotiation with a clocked contract, which is what made pipelining and counted bursts possible. It also fixes the vocabulary the rest of the curriculum depends on: clock frequency, transfer rate, data rate and bandwidth are four different quantities.
- Related topic
HBM Overview
HBM reaches hundreds of GB/s with a per-pin rate lower than DDR5's. It wins on width, not speed — and getting that width required changing the packaging, which adds a fourth design layer to array physics, device architecture and the interface.
- Related topic
DDR Bandwidth
Chapter 12.4 named four efficiency measures and built three. This builds the fourth: every bus cycle charged to exactly one named cause, with the categories provably summing to the window.
- Related topic
HBM for AI Accelerators
Stacks buy bandwidth and dies buy capacity. The same 96 GB bought as twelve short stacks delivers three times the bandwidth of four tall ones — and one requester cannot use any of it.
Standards & specifications
- Governing standard
- JEDEC JESD79 (DDR SDRAM)(opens JEDEC Solid State Technology Association in a new tab)
Defines the DDR SDRAM device itself — signals, command encoding, mode registers, timing parameters and the initialisation sequence — one document per generation. Memory-controller microarchitecture, address-mapping policy, PHY training algorithms and board-level design are not specified by it.
This page also covers RTL structure, verification approach and debugging technique. Those are engineering practice built on the standard, not requirements the standard itself imposes.
Where this fits
Part of the DDR curriculum.
