DDR · Module 32
AI Accelerators
Both halves of the memory architecture — the stack count and the requester parallelism — are silicon, committed before the workload that will run on them exists. So the deciding quantity is the cost asymmetry between the two provisioning errors, and the controller's job is to produce the evidence that improves the next decision.
Chapter 26.4 §8 established that the stack count and the requester parallelism must be designed together and that neither can be chosen independently — a tighter coupling between memory and compute than any other memory in this curriculum imposes.
This chapter starts from that result and asks the question it leaves open: both halves are silicon, so both are committed before the workload that will run on them exists.
The dominating constraint is that the provisioning decision is taken under uncertainty and cannot be revised. So the deciding quantity is not the expected workload — it is the cost asymmetry between the two provisioning errors, and that asymmetry moves the rational provisioning point away from the expectation.
And the decision it forces on the controller is unlike any other in this module.
The controller's job is not only to serve this generation's workload. It is to produce the evidence that makes the next provisioning decision better — because the current decision is already fixed and the only thing still in play is the next one.
Axis A1 is identical for the fourth chapter in this module. The same destructive read, the same thirteen obligations of 31.1 §5. What differs is that this platform class commits more of its architecture to an unknown workload than any other, and pays for the commitment in silicon.
1. The Shared Baseline, and This Chapter's Question
CURRICULUM-DERIVED from 32.1 §1: all five platform classes carry the same thirteen obligations of 31.1 §5, and the module's question is always the same three steps.
1 WHICH CONSTRAINT
both halves of 26.4 §8's co-design are SILICON, so both are
committed before the workload exists (§3)
2 WHAT DECISION
provision against the COST ASYMMETRY rather than the
expectation (§5, §6), and instrument for the NEXT decision (§7)
3 WHAT GRADE
D for every cost and workload figure. 26.4's stream and
stack arithmetic is A/CURRICULUM-DERIVED and is the only
firm ground available (§9)2. The Contract This Chapter Was Given
Chapter 26.4 named this chapter in its Scope and reserved territory. Honouring that explicitly, rather than assuming it, is part of the work.
| 26.4 reserved | Who owns it | This chapter's position |
|---|---|---|
| Why an accelerator's architecture depends on HBM's properties | 26.4 itself | consumed in §3 and §4; not restated |
| Access-pattern taxonomy | 29.5 | cited in §8; not restated |
| Case studies of named products | nominally this chapter | declined — see below |
| Naming a product | nobody; 26.4 names none | none named here either |
The third row needs stating plainly, because it is a deviation from the chapter's title.
A case study of a named product would require grade-A facts about that product's memory provisioning, requester parallelism and workload, and CURRICULUM-DERIVED from 18.4 §1, a claim without a source and a category is not usable engineering evidence however plausible it looks. None of those facts is publicly documented for accelerators, so the alternatives are to invent them as grade D dressed as grade A — the category drift 18.4 §1 warns about — or to build the chapter around the decision instead.
So this chapter is a case study of a decision rather than of a product, which is what §4 of this module's law asks for anyway: a case study is a design decision, the constraint that forced it, and the evidence grade of every claim. The constraint is real, the decision is real, and the arithmetic is reproducible — and no part of it depends on a fact nobody published.
3. What Must Be Fixed Before the Workload Exists
CURRICULUM-DERIVED from 26.4 §8, which owns both halves and their coupling: an accelerator cannot choose its memory-level parallelism at runtime — the number of independent outstanding requests it can have is set by how many independent execution units, memory pipelines and outstanding-request slots it was designed with, and that is a silicon decision made before any workload runs.
Both halves of the pair are therefore silicon, and that is unusual. Contrast the four preceding platform classes.
| Platform class | What is committed in silicon | What remains adjustable |
|---|---|---|
| 32.1 CPU | the controller | population, interleaving, page policy, firmware — 18.4 §5 |
| 32.2 mobile | the controller and the attach | throttle thresholds, refresh mode, window |
| 32.3 GPU-class | the controller, the attach, the partition | tuning, per-population policy |
| This class | the stack count, the stream count, AND the requester parallelism | almost nothing that matters |
| 32.5 server | the controller | the tier composition, at deployment |
Row four is the chapter's constraint and row five is the contrast that makes it sharp. CURRICULUM-DERIVED from 26.4 §8: the stack count and the requester parallelism must be designed together, and an accelerator designed for one is over-provisioned against the other. So the pair is a single joint decision, taken once, and both halves are expensive to be wrong about.
And the workload is the one input the decision needs and does not have. CURRICULUM-DERIVED from 26.4 §5, which owns arithmetic intensity — the ratio at which a compute rate becomes unservable — so the correct provisioning is a function of the workload's intensity, and a workload's intensity is a property of software that may be written after the silicon ships.
the joint decision, stated as what it needs and what it has
needs : the workload's arithmetic intensity
the workload's achievable outstanding-request count
the workload's address distribution
has : an estimate of each, from workloads that exist NOW
and a design lifetime over which they will change
so the decision is a bet, and §5 costs the two ways to lose it.4. The Four Mismatches, Costed
CURRICULUM-DERIVED from 26.4 §9, which owns the four mismatches and the observation that all of them produce working systems. What that chapter does not do is cost them, and the costs are not equal.
| Mismatch | 26.4's description | Direction | What it costs |
|---|---|---|---|
| M1 tall stacks, bandwidth-bound workload | capacity you cannot feed | under-provisioned bandwidth | a fraction of achievable performance — 26.4 §3 puts it at a third |
| M2 many stacks, capacity-bound workload | bandwidth you cannot use | over-provisioned bandwidth | money and integration risk, at full performance |
| M3 enough stacks, too few outstanding requests | a fraction of the bandwidth | under-provisioned requester | a fraction of achievable performance, permanently |
| M4 enough requests, concentrated distribution | a fraction of the bandwidth | not a provisioning error at all | fixable in software or mapping — §8 |
Three observations, and the third reframes the problem.
M1 and M3 are the same error on the two different halves, and both are under-provisioning. CURRICULUM-DERIVED from 26.4 §9, both produce a fraction of the bandwidth and nothing malfunctions — so neither is detected by any correctness test.
M2 is the only over-provisioning error, and it is the only one that delivers full performance. That asymmetry is the whole of §5. A system with too much memory bandwidth runs the workload at the speed the workload can achieve; a system with too little runs it slower, forever.
And M4 is not a provisioning error at all, which is why it must be separated from the other three before any provisioning conclusion is drawn. CURRICULUM-DERIVED from 26.4 §7: a distribution that reaches the streams is an address-mapping property rather than a concurrency one — so M4 is fixable after the silicon ships and M1 and M3 are not. A measurement that cannot separate M3 from M4 will read a software problem as a silicon one, and 26.4 §12's monitor exists precisely to keep those two apart.
5. The Asymmetry, in Closed Form
Cost both directions against the same denominator — delivered performance — because that is what the system is bought for. Every figure grade D; the form is the transferable part.
DEFINITIONS
V the value of the system running at full workload speed
c the cost of the memory subsystem at the NOMINAL
provisioning point
f the misprovisioning factor, f > 1
OVER-provisioning by f (M2):
extra memory cost = (f - 1) * c
performance = FULL
cost per unit of delivered performance rises by
(1 + (f-1)*c/V_total) where V_total is the whole system's
cost -- a small number when memory
is a fraction of the system.
UNDER-provisioning by f (M1 or M3):
memory cost saved = (1 - 1/f) * c
performance = 1/f of achievable (26.4 §3's mechanism)
value delivered = V / f
cost per unit of delivered performance rises by
f * (1 - (1 - 1/f)*c/V_total) ~ f for small c/V_total
DERIVED, and this is the asymmetry:
over-provision by f -> cost/performance x (1 + (f-1)*c/V)
under-provision by f -> cost/performance x approximately f
the two are equal only when c/V is close to 1, i.e. when the
memory IS the system. Whenever memory is a fraction of the
system's cost, UNDER-provisioning is far worse.Put grade-D numbers on it.
GRADE D. memory subsystem = 30% of system cost, so c/V = 0.30.
misprovisioning factor f = 2.
over-provision 2x : cost/perf x (1 + 1.0 * 0.30) = 1.30
under-provision 2x: cost/perf x 2.00 * (1 - 0.15) = 1.70
DERIVED ratio of penalties: 1.70 / 1.30 = 1.31x worse to
under-provision, at this cost share and this factor.
recompute at c/V = 0.10 (memory a tenth of the system):
over : 1 + 1.0*0.10 = 1.10
under : 2.00 * (1 - 0.05) = 1.90
DERIVED ratio: 1.73x worse to under-provision.
recompute at c/V = 0.60 (memory the dominant cost):
over : 1 + 1.0*0.60 = 1.60
under : 2.00 * (1 - 0.30) = 1.40
DERIVED ratio: 0.88 -- now OVER-provisioning is worse.So the asymmetry has a sign that flips, and the quantity that flips it is the memory's share of system cost. DERIVED from the closed form: the break-even is where (f−1)·c/V = f·(1 − (1−1/f)·c/V) − 1, and for f = 2 that resolves to c/V ≈ 0.5.
When the memory subsystem is less than about half the system's cost, under-provisioning is the more expensive error and the rational provisioning point sits ABOVE the expected workload. When memory dominates the cost, the bias reverses.
And the qualification that keeps this honest. The model charges under-provisioning with a proportional performance loss, which CURRICULUM-DERIVED from 26.4 §3 is the right shape for a bandwidth-bound workload — but a workload that is not bandwidth-bound loses nothing from under-provisioned bandwidth, and then M2's cost dominates trivially. So the model's input is not just the cost share: it is the probability that the workload is bandwidth-bound at all, which is what 26.4 §5's arithmetic intensity decides. Stating that dependency is part of the answer.
6. Provision for the Error You Can Afford
§5's result is a bias, not a number, and applying it needs a bound — which is 32.3 §17's discipline reused.
the decision, as a procedure
1 ESTIMATE the workload's arithmetic intensity, and the
probability it is bandwidth-bound (26.4 §5)
2 COMPUTE the memory's share of system cost, c/V
3 IF c/V < ~0.5 AND the workload is likely bandwidth-bound
-> bias the provisioning UP. The cheaper error is M2.
IF c/V > ~0.5
-> bias DOWN. The cheaper error is M1/M3.
4 COMPUTE THE CEILING of the bias, because provisioning up has
one: 26.4 §8's table gives the minimum outstanding requests
to fill a given stack count, and a requester that cannot
supply them converts M2 into M3.Step 4 is the step that is usually skipped and it is the one that makes the bias bounded rather than a licence.
CURRICULUM-DERIVED from 26.4 §8's table: 1 stack needs 16 outstanding, 4 needs 64, 8 needs 128, 12 needs 192 — just to touch every stream once, and considerably more to keep them busy, since each must be re-issued as it completes.
DERIVED from that table. Suppose the requester was designed for
128 concurrent outstanding accesses.
provisioning up to 8 stacks : needs 128. EXACTLY met.
provisioning up to 12 stacks : needs 192. NOT met.
so provisioning the 12-stack configuration against a 128-request
requester does not buy 12 stacks of bandwidth. It buys 8 stacks
of bandwidth and 12 stacks of cost -- which is M2 and M3 at the
same time, and 26.4 §9 says both produce working systems.
CEILING: the provisioning bias is bounded by the requester's
own parallelism, and 26.4 §8 is explicit that the requester's
parallelism is silicon too.So the bias is bounded by the other half of the same joint decision, which is 26.4 §8's coupling appearing as a constraint on §5's conclusion. CURRICULUM-DERIVED from 30.8 §6: compute the ceiling of your lever before committing to it — and here the ceiling is set by a decision being taken in the same review.
The honest summary of §5 and §6 together: provision above the expectation, by no more than the requester can fill, and record which of the two halves bounded you. That last clause is §7's subject, because the binding half is the thing the next generation most needs to know.
7. What the Controller Owes the Next Generation
This platform class's provisioning decision is already fixed by the time the controller runs. So the controller's most valuable output is not performance — it is evidence about a decision that has not been taken yet.
CURRICULUM-DERIVED from 23.5, which owns the counterfactual measurement that lets a design evaluate a policy it has not deployed. This section applies that technique to provisioning, and the translation is direct.
| Counterfactual question | What must be recorded to answer it | Why the controller is the only place |
|---|---|---|
| Would more stacks have helped? | were the existing streams saturated, or idle? | only the controller sees per-stream occupancy |
| Would more outstanding requests have helped? | was the outstanding count at its ceiling while streams were idle? | only the controller sees both at once |
| Would a better distribution have helped? | were streams idle while the outstanding count was high? | this is M4, and it needs both numbers — 26.4 §12 |
| Was the workload bandwidth-bound at all? | did demand ever exceed what the streams could deliver? | §5's model needs this and nothing else supplies it |
Read the second and third rows together: they are M3 and M4, and they have 26.4 §9's identical symptom. CURRICULUM-DERIVED from 26.4 §12, whose monitor reports the two requirements separately precisely because §9 establishes they are different causes with an identical symptom. §11's block extends that to the counterfactual question, which needs the same two numbers plus their joint history.
And the obligation this creates is a design requirement rather than a nicety. CURRICULUM-DERIVED from 30.10 §13, which owns debug observables as design deliverables with a cost in reproduction hours: here the cost is a generation. A system that ships without counterfactual instrumentation forces the next provisioning decision to be taken on the same estimate as the last one, and the estimate was the thing that was uncertain.
On a platform whose architecture is committed before its workload exists, the controller's instrumentation is not a debug feature. It is the only mechanism by which the next decision is better informed than this one — and its absence guarantees the same bet is placed twice.
8. Degradation Looks Like Distribution
Chapter 26.4 §9 names two causes of a fraction of the bandwidth with an identical symptom. There is a third, and it is not a provisioning error either.
CURRICULUM-DERIVED from 26.3, which owns the vertical path, why a stack needs spares, that one spare does almost all the work, and that repair is per group rather than per via — and that the repair happens at manufacture. Chapter 31.3 §16 owns the consequence: there is none in the field.
a third cause of "a fraction of the bandwidth"
M5 (this section) : one stack degraded or partially repaired,
so its streams deliver less than the others
and why it presents identically to M4:
M4 : streams idle because ADDRESSES did not reach them
M5 : streams slow because that STACK cannot deliver
both show: outstanding count high, aggregate bandwidth low,
some streams under-delivering. A monitor that reports only
"streams occupied" cannot tell them apart.The discriminator is per-stream throughput rather than per-stream occupancy, and that is the distinction §12's block is built on.
| Observation | M3 too few requests | M4 bad distribution | M5 degradation |
|---|---|---|---|
| Outstanding count | low | high | high |
| Streams occupied | few | few | many |
| Per-stream throughput when occupied | normal | normal | low on a subset |
| Which subset | — | — | contiguous within one stack |
| Fixable after shipping? | no — silicon | yes — mapping or software | no — 31.3 §16 |
Row four is the cheapest discriminator and it is structural. CURRICULUM-DERIVED from 26.1 §3 and 26.1 §6: streams are grouped into channels and pseudo-channels within a stack, so a degradation is confined to one stack's stream range while a distribution failure follows the address pattern, which has no reason to align with a stack boundary. A monitor that reports per-stream throughput grouped by stack answers the question in one window.
And row five is why the distinction matters commercially rather than only diagnostically. M4 is a software fix and M5 is a replacement — and CURRICULUM-DERIVED from 31.3 §16, on this attach the package is scrap. So misreading M5 as M4 sends a team to optimise addresses on hardware that will not improve, and misreading M4 as M5 scraps a working part.
9. The Evidence Situation
§Scope's grades, and this platform class has the poorest evidence base in the module — worse even than 32.3 §8's.
| What you might want | Best grade available | Why |
|---|---|---|
| Streams per stack, interface width, per-stack bandwidth | A, cited | 26.4 §1 and 26.4 §6, attributing 4.8 §2 and 26.2 §8 |
| Minimum outstanding requests per stack count | A/DERIVED | 26.4 §8's table, derived from §6 under §7's stated model |
| That a stack needs spares and repair is per group | A | 26.3 §4, 26.3 §6 |
| A named accelerator's outstanding-request capability | none | not published; it is per-engine and per-configuration |
| A named accelerator's memory provisioning rationale | none | a commercial decision, never published |
| The memory's share of system cost | none from any technical source | a commercial figure, and §5 turns on it |
| A workload's arithmetic intensity | C or D | measurable on software that exists; unknowable for software that does not |
Two observations, and the second is why §2 declined the product-tour reading of this chapter's title.
The only firm ground is Module 26's arithmetic, and this chapter leans on it entirely. Every structural claim in §3 to §8 reduces to 26.4's table, 26.4 §7's two requirements, or 26.3's repair model. Nothing here needs a fact about a product.
And §5's deciding quantity has no technical grade at all. c/V — the memory's share of system cost — is a commercial figure, and the sign of §5's asymmetry depends on it. So the memory team can supply the model and cannot supply its most important input, which is structurally the same finding as 32.2 §9's about thermal resistance: the analysis is incomplete by construction until somebody outside the team supplies a number. CURRICULUM-DERIVED from 18.4 §1: that qualification must travel with every conclusion drawn from it.
10. The Platform, as Blocks
Two nodes are marked SILICON and one is marked software, and that split is the chapter. CURRICULUM-DERIVED from 26.4 §8: the engines' parallelism and the stack count are both design-time; CURRICULUM-DERIVED from 26.4 §7, the distribution is an address-mapping property, so it is not. Two of the three inputs to the outcome are unchangeable and the third is not, which is why §8's discriminator matters so much.
And the rightmost lower node is unusual for a diagram in this curriculum: it is not part of the system. The next provisioning decision is the only thing still open, and §7 is the argument that the instrumentation exists to serve it. Both instrumentation blocks point there and nowhere else.
11. RTL — The Counterfactual Reporter
The platform-profile block for this chapter. NUM_STACKS and REQ_CEILING are the two halves of 26.4 §8's joint decision, and the block's job is to report which of them bound the system — because §6's procedure needs exactly that, and §7 argues nothing else can supply it.
// ---------------------------------------------------------------------
// provisioning_counterfactual -- the platform-profile block of §11.
//
// CLASSIFICATION: synthesisable telemetry. Drives nothing. Grade-D
// parameter values except where noted CURRICULUM-DERIVED.
//
// WHAT IT IS: §7's four counterfactual questions, answered from the
// JOINT history of the outstanding-request count and the stream
// occupancy. 26.4 §12's monitor reports the two requirements
// separately; this block additionally records WHICH ONE WAS BINDING
// over time, which is what a provisioning decision needs and a
// point-in-time report does not supply.
//
// WHY IT EXISTS HERE: §3 establishes that both halves are silicon and
// committed before the workload exists, so the current decision is
// closed. §7 concludes the controller's most valuable output is
// evidence for the next one, and 23.5 owns the counterfactual
// technique this applies.
//
// HOW TO RUN IT: run the real workload for a full window and read
// bound_by_requester against bound_by_streams.
// EXPECTED RESULT: exactly one of the two dominates, or both are low
// and the workload was not bandwidth-bound at all -- which is §5's
// most important input and the one nothing else reports.
//
// SYNTHESIS: three counters, two comparators, one min/max tracker.
// No divider: every ratio is reported as a pair.
//
// LIMITATIONS: it reports which half was BINDING, not what a
// different provisioning would have delivered numerically -- that
// requires a model of the alternative, and 23.5's counterfactual is
// explicit that a replay against an alternative is how such a number
// is obtained. Stating the boundary rather than implying a
// prediction. It also cannot see M5: per-stack throughput is §12's,
// and 27.3's independence discipline requires saying so.
// ---------------------------------------------------------------------
module provisioning_counterfactual #(
// The MEMORY half of the joint decision. CURRICULUM-DERIVED from
// 26.4 §6: sixteen streams per stack.
parameter int NUM_STACKS = 8,
parameter int STREAMS_PER_STK = 16,
parameter int TOTAL_STREAMS = NUM_STACKS * STREAMS_PER_STK,
// The REQUESTER half. CURRICULUM-DERIVED from 26.4 §8: an
// accelerator cannot choose this at runtime, so it is a parameter
// here in the same sense that it is silicon there.
parameter int REQ_CEILING = 128,
parameter int WIN = 1048576,
// COUNT, not INDEX: each counter must represent a full window, so
// $clog2(WIN + 1). Sized for an expected rate it would saturate at
// exactly the saturated operating point the counterfactual is
// about, and report the system as unbound when it was bound.
parameter int CNT_W = $clog2(WIN + 1),
parameter int OCC_W = $clog2(REQ_CEILING + 1),
parameter int STR_W = $clog2(TOTAL_STREAMS + 1)
)(
input logic clk,
input logic rst_n,
input logic [OCC_W-1:0] outstanding, // requests in flight
input logic [STR_W-1:0] streams_busy, // streams with work
input logic demand_pending, // somebody wants memory
input logic win_tick,
output logic [CNT_W-1:0] c_bound_requester,
output logic [CNT_W-1:0] c_bound_streams,
output logic [CNT_W-1:0] c_bound_neither,
output logic [CNT_W-1:0] c_demand_cycles,
output logic [STR_W-1:0] streams_busy_max,
output logic [OCC_W-1:0] outstanding_max,
output logic result_valid,
output logic counter_saturated
);
initial begin
if (NUM_STACKS < 1) $fatal(1, "provisioning_counterfactual: NUM_STACKS >= 1");
if (REQ_CEILING < 1) $fatal(1, "provisioning_counterfactual: REQ_CEILING >= 1");
// 26.4 §8's table, as an elaboration WARNING rather than a
// fatal: a requester that cannot supply one access per stream
// converts an over-provisioned memory (M2) into M3 as well, and
// §6 step 4 is exactly this check. It is a warning because the
// configuration is legal and merely unwise, and saying which is
// part of the contract.
if (REQ_CEILING < TOTAL_STREAMS)
$warning("provisioning_counterfactual: REQ_CEILING %0d < TOTAL_STREAMS %0d -- cannot touch every stream once (26.4 §8, §6 step 4)",
REQ_CEILING, TOTAL_STREAMS);
end
logic [CNT_W-1:0] br, bs, bn, dc;
logic [STR_W-1:0] smax;
logic [OCC_W-1:0] omax;
// The three attributions. They are MUTUALLY EXCLUSIVE and jointly
// exhaustive over demand cycles, which is what makes the report a
// partition rather than three loosely related counters -- 23.2 §9's
// exhaustive-and-disjoint standard, applied to provisioning.
logic at_req_ceiling, streams_saturated;
assign at_req_ceiling = (outstanding >= REQ_CEILING[OCC_W-1:0]);
assign streams_saturated = (streams_busy >= TOTAL_STREAMS[STR_W-1:0]);
always_ff @(posedge clk) begin
if (!rst_n) begin
br <= '0; bs <= '0; bn <= '0; dc <= '0;
smax <= '0; omax <= '0;
c_bound_requester <= '0; c_bound_streams <= '0;
c_bound_neither <= '0; c_demand_cycles <= '0;
streams_busy_max <= '0; outstanding_max <= '0;
result_valid <= 1'b0; counter_saturated <= 1'b0;
end else if (win_tick) begin
c_bound_requester <= br;
c_bound_streams <= bs;
c_bound_neither <= bn;
c_demand_cycles <= dc;
streams_busy_max <= smax;
outstanding_max <= omax;
result_valid <= 1'b1;
br <= '0; bs <= '0; bn <= '0; dc <= '0;
smax <= '0; omax <= '0;
end else begin
result_valid <= 1'b0;
// The denominator is DEMAND cycles, not window cycles. 30.8
// §10's denominator error: counting idle cycles would conflate
// "the provisioning was adequate" with "nobody asked", and
// §5's model needs the second case reported SEPARATELY because
// it means the workload was not bandwidth-bound at all.
if (demand_pending) begin
if (dc != {CNT_W{1'b1}}) dc <= dc + 1'b1;
else counter_saturated <= 1'b1;
// §7's rows two and three, as a partition. The joint test is
// the whole point: 26.4 §12 reports the two requirements
// separately and 26.4 §9 says they share a symptom, so only
// the CONJUNCTION attributes a cycle.
if (at_req_ceiling && !streams_saturated) begin
// Requests are the ceiling and streams are idle -> MORE
// OUTSTANDING REQUESTS would have helped. This is M3.
if (br != {CNT_W{1'b1}}) br <= br + 1'b1;
else counter_saturated <= 1'b1;
end else if (streams_saturated) begin
// Streams are full -> MORE STACKS would have helped. M1.
if (bs != {CNT_W{1'b1}}) bs <= bs + 1'b1;
else counter_saturated <= 1'b1;
end else begin
// Neither is at its limit and demand is pending. The
// binding constraint is elsewhere -- distribution (M4),
// degradation (M5, §12), or the workload simply is not
// bandwidth-bound. §5's most important input.
if (bn != {CNT_W{1'b1}}) bn <= bn + 1'b1;
else counter_saturated <= 1'b1;
end
end
if (streams_busy > smax) smax <= streams_busy;
if (outstanding > omax) omax <= outstanding;
end
end
endmoduleThe c_bound_neither counter is the block's most valuable output and the least obvious.
A large c_bound_neither means demand was pending while neither half of the provisioning was at its limit — so more memory would not have helped and more requests would not have helped. CURRICULUM-DERIVED from §5's qualification: a workload that is not bandwidth-bound loses nothing from under-provisioned bandwidth, and then M2's cost dominates trivially. So a high c_bound_neither is the single strongest argument available for provisioning down next time — and it is the one number a bandwidth-focused report would never produce.
And the partition is exhaustive over demand cycles by construction. CURRICULUM-DERIVED from 23.2 §9, whose standard is exhaustive and disjoint attribution charging every cycle to one named cause — applied here to provisioning rather than to bandwidth, so br + bs + bn = dc and §17 asserts it.
12. RTL — Separating Degradation From Distribution
§8 established a third cause of the same symptom, with a structural discriminator: per-stream throughput grouped by stack. This block is that measurement.
// ---------------------------------------------------------------------
// per_stack_throughput -- verification/telemetry, CORRECT as written.
//
// CLASSIFICATION: synthesisable telemetry. Drives nothing. Grade D.
//
// WHAT IT DOES: counts delivered beats per stack over a window and
// reports the spread, plus whether the under-delivering set is
// CONTIGUOUS within one stack.
//
// WHY IT EXISTS HERE: §8's table shows M4 and M5 differ in exactly
// one observable -- whether the under-delivering streams align with
// a stack boundary -- because 26.1 §3 and §6 establish that streams
// are grouped into channels within a stack while an address
// distribution has no reason to follow that grouping. And §8 row
// five is why it matters commercially: M4 is a software fix, M5 is
// scrap (31.3 §16).
//
// HOW TO RUN IT: run the workload and compare per-stack delivery.
// EXPECTED RESULT: a healthy system's stacks deliver within a narrow
// spread. One stack low is M5; a scattered pattern is M4.
//
// SYNTHESIS: one counter per stack plus a min/max tracker.
//
// LIMITATIONS: attributes to a STACK, not to a via or a group --
// 26.3 §6 owns that repair is per group rather than per via, and
// this block cannot see inside a stack. Nor can it prove degradation:
// a stack serving a cold address range legitimately delivers less.
// §18's row five is the disambiguation, and stating that this block
// does not perform it is 27.3's independence discipline.
// ---------------------------------------------------------------------
module per_stack_throughput #(
parameter int NUM_STACKS = 8,
parameter int WIN = 1048576,
// COUNT, not INDEX: a window in which ONE stack served every beat
// must be representable, so $clog2(WIN + 1). Sized to the fair
// share WIN / NUM_STACKS it would saturate in exactly the skewed
// case it exists to detect, and report balance.
parameter int CNT_W = $clog2(WIN + 1),
parameter int SID_W = $clog2(NUM_STACKS > 1 ? NUM_STACKS : 2)
)(
input logic clk,
input logic rst_n,
input logic beat_delivered,
input logic [SID_W-1:0] beat_stack,
input logic win_tick,
output logic [CNT_W-1:0] r_stack [NUM_STACKS],
output logic [CNT_W-1:0] r_max,
output logic [CNT_W-1:0] r_min,
output logic [SID_W-1:0] r_min_stack,
output logic single_stack_low,
output logic result_valid
);
initial begin
if (NUM_STACKS < 1) $fatal(1, "per_stack_throughput: NUM_STACKS >= 1");
if (WIN < NUM_STACKS)
$fatal(1, "per_stack_throughput: WIN < NUM_STACKS makes the spread meaningless");
end
logic [CNT_W-1:0] c [NUM_STACKS];
always_ff @(posedge clk) begin
if (!rst_n) begin
for (int i = 0; i < NUM_STACKS; i++) begin c[i] <= '0; r_stack[i] <= '0; end
r_max <= '0;
// AND-accumulator discipline: a minimum tracker starts at
// all-ones, or the first sample never wins and the reported
// minimum stays zero -- which would flag every window as
// having a dead stack.
r_min <= '1;
r_min_stack <= '0;
single_stack_low <= 1'b0; result_valid <= 1'b0;
end else if (win_tick) begin
begin
logic [CNT_W-1:0] mx, mn;
logic [SID_W-1:0] mni;
logic [31:0] n_low;
mx = '0; mn = '1; mni = '0; n_low = 0;
for (int i = 0; i < NUM_STACKS; i++) begin
if (c[i] > mx) mx = c[i];
if (c[i] < mn) begin mn = c[i]; mni = i[SID_W-1:0]; end
end
// How many stacks are materially below the maximum? §8's
// discriminator: ONE low stack is M5; SEVERAL scattered is
// more consistent with M4, because a distribution failure has
// no reason to spare all but one stack.
for (int i = 0; i < NUM_STACKS; i++)
if ((c[i] * 4) < (mx * 3)) n_low = n_low + 1;
r_max <= mx; r_min <= mn; r_min_stack <= mni;
// Reported as a FLAG plus both endpoints and the index, never
// as a lone ratio -- 27.5 refuses a single percentage and
// 30.8 §11 requires a derived statistic's range to be visible.
single_stack_low <= (n_low == 1) && (NUM_STACKS > 1);
end
for (int i = 0; i < NUM_STACKS; i++) begin
r_stack[i] <= c[i];
c[i] <= '0;
end
result_valid <= 1'b1;
end else begin
result_valid <= 1'b0;
if (beat_delivered && c[beat_stack] != {CNT_W{1'b1}})
c[beat_stack] <= c[beat_stack] + 1'b1;
end
end
endmodulesingle_stack_low is the discriminator and it is deliberately a weak claim. One stack materially below the others is consistent with M5 and does not prove it — a stack serving a cold address range legitimately delivers less, which the block's limitations say explicitly. CURRICULUM-DERIVED from 18.4 §1: this output is grade C about this machine, and the qualification must travel with it. §18 row five is the experiment that promotes it.
13. RTL — The Provisioning Ceiling
§6 step 4 computes a ceiling on the provisioning bias and §11 only warns about it at elaboration. This block computes it, because 26.4 §8's table is arithmetic and a number a design review can read beats a warning nobody sees.
// ---------------------------------------------------------------------
// provisioning_ceiling -- CORRECT as written. Grade-D inputs except
// where noted CURRICULUM-DERIVED.
//
// CLASSIFICATION: synthesisable, but intended as an elaboration-time
// and design-review artifact. It drives nothing.
//
// WHAT IT DOES: given the requester's outstanding-request ceiling,
// reports the largest stack count that ceiling can actually fill,
// and flags a configuration that provisions beyond it.
//
// WHY IT EXISTS HERE: §6 step 4 bounds §5's provisioning bias by the
// requester's own parallelism, and 26.4 §8 supplies the arithmetic:
// sixteen streams per stack, and one outstanding access per stream
// merely to touch each once. A provisioning decision taken without
// this number converts M2 into M2-and-M3 simultaneously (§6), and
// 26.4 §9 says both produce working systems.
//
// HOW TO RUN IT: set REQ_CEILING to the requester's designed
// parallelism and sweep NUM_STACKS.
// EXPECTED RESULT: max_useful_stacks = floor(REQ_CEILING / 16), and
// over_provisioned rises once NUM_STACKS exceeds it.
//
// SYNTHESIS: a shift and two comparators. No divider -- the streams
// per stack is a power of two, which is CURRICULUM-DERIVED from
// 26.4 §6 rather than assumed.
//
// LIMITATIONS: reports the ceiling to TOUCH every stream once. 26.4
// §8 is explicit that KEEPING them busy needs "considerably more,
// since each must be re-issued as it completes" -- so this is a
// LOWER bound on the requirement and an UPPER bound on the useful
// stack count. Saying which direction each bound runs is the point;
// a single number here would be read as the answer.
// ---------------------------------------------------------------------
module provisioning_ceiling #(
// CURRICULUM-DERIVED from 26.4 §6. A power of two, which is why no
// divider is needed below.
parameter int STREAMS_PER_STK = 16,
// The requester half of 26.4 §8's joint decision. Silicon.
parameter int REQ_CEILING = 128,
// The memory half, as proposed.
parameter int NUM_STACKS = 8,
// §6 step 4's honest margin: 26.4 §8 says touching every stream
// once is not enough to keep them busy. This is the multiplier a
// design review must argue for, and it is grade D here because the
// real value depends on the round-trip and the re-issue rate.
parameter int REISSUE_FACTOR = 2,
parameter int LOG_SPS = $clog2(STREAMS_PER_STK),
// COUNT, not INDEX: the reported stack count must represent
// NUM_STACKS inclusive, so $clog2(NUM_STACKS + 1). Sized
// $clog2(NUM_STACKS) the ceiling can never equal the proposal and
// over_provisioned would be unreachable at the exact boundary --
// the guard failing on the one configuration it is for.
parameter int STK_W = $clog2(NUM_STACKS + 1)
)(
input logic clk,
input logic rst_n,
// Runtime versions, so a platform whose requester ceiling is
// discoverable can evaluate this without re-elaborating.
input logic [15:0] cfg_req_ceiling,
input logic [7:0] cfg_num_stacks,
input logic cfg_valid,
output logic [STK_W-1:0] max_useful_stacks,
output logic [STK_W-1:0] max_useful_stacks_busy,
output logic over_provisioned,
output logic under_requested,
output logic cfg_out_of_range
);
initial begin
if (STREAMS_PER_STK < 1 || (STREAMS_PER_STK & (STREAMS_PER_STK - 1)) != 0)
$fatal(1, "provisioning_ceiling: STREAMS_PER_STK must be a power of two (26.4 §6)");
if (REQ_CEILING < 1) $fatal(1, "provisioning_ceiling: REQ_CEILING >= 1");
if (REISSUE_FACTOR < 1)
$fatal(1, "provisioning_ceiling: REISSUE_FACTOR >= 1");
end
// A configured ceiling of zero, or a stack count of zero, is not a
// configuration -- it is an absent one, and 30.7 §9's discipline
// says unknown is not a default.
assign cfg_out_of_range = cfg_valid &&
((cfg_req_ceiling == '0) || (cfg_num_stacks == '0));
always_comb begin
if (!cfg_valid || cfg_out_of_range) begin
max_useful_stacks = '0;
max_useful_stacks_busy = '0;
over_provisioned = 1'b0;
under_requested = 1'b0;
end else begin
// 26.4 §8's table, computed. A shift rather than a divide,
// because STREAMS_PER_STK is a power of two.
max_useful_stacks = STK_W'(cfg_req_ceiling >> LOG_SPS);
// And the honest version, per this block's LIMITATIONS: 26.4
// §8 says touching every stream once is not enough. Reported
// SEPARATELY rather than replacing the first number, so a
// reviewer sees both bounds and cannot mistake one for the
// other.
max_useful_stacks_busy = STK_W'((cfg_req_ceiling / REISSUE_FACTOR) >> LOG_SPS);
// The two directions of §6's bound. Both are flagged, because
// §4 establishes M2 and M3 are different errors with different
// costs and a single "mismatch" flag would conflate them.
over_provisioned = (cfg_num_stacks > max_useful_stacks);
under_requested = (cfg_num_stacks < max_useful_stacks_busy);
end
end
endmoduleThe block reports two ceilings on purpose, and the reason is 26.4 §8's own wording. That chapter's table gives the outstanding requests needed just to touch every stream once, and states that keeping them busy needs considerably more, since each must be re-issued as it completes. So one number is a hard upper bound on useful stacks and the other is a realistic one, and collapsing them would produce a figure a review would read as the answer.
DERIVED, with REQ_CEILING = 128 and the CURRICULUM-DERIVED sixteen streams per stack:
max_useful_stacks = 128 >> 4 = 8 stacks
max_useful_stacks_busy = (128/2) >> 4 = 4 stacks
so a 128-request requester can TOUCH 8 stacks' streams and
realistically KEEP 4 stacks busy at a grade-D re-issue factor
of 2. A 12-stack proposal is over-provisioned against both.And under_requested is the mirror flag that keeps the bound two-sided. CURRICULUM-DERIVED from 30.3 §9: a safety property cannot detect over-restriction, so a design that only flags over-provisioning will never object to a requester built for far more concurrency than its memory can absorb — which is 26.4 §8's an accelerator designed for 192 outstanding accesses is over-provisioned against a single stack, in the other direction.
// ---- The ceiling block's own contract.
// The relationship 26.4 §8's table encodes, as an invariant. If
// this fails, the block and the table disagree and the table wins.
property p_ceiling_matches_table;
@(posedge clk) disable iff (!rst_n)
(cfg_valid && !cfg_out_of_range) |->
(max_useful_stacks == (cfg_req_ceiling >> LOG_SPS));
endproperty
assert property (p_ceiling_matches_table)
else $error("the computed ceiling disagrees with 26.4 §8's arithmetic");
// The realistic ceiling must never EXCEED the touch-once ceiling,
// because a re-issue factor of at least one cannot increase the
// number of stacks a fixed request budget can fill. A derived
// statistic whose range follows from its definition -- 30.8 §11.
property p_busy_ceiling_not_above_touch_ceiling;
@(posedge clk) disable iff (!rst_n)
(cfg_valid && !cfg_out_of_range) |->
(max_useful_stacks_busy <= max_useful_stacks);
endproperty
assert property (p_busy_ceiling_not_above_touch_ceiling)
else $error("the realistic ceiling exceeded the touch-once ceiling");
// An absent configuration produces no verdict rather than a
// default one -- the same discipline 32.2 §15's filter applies,
// and 30.7 §9's "unknown is not a default".
property p_no_verdict_without_config;
@(posedge clk) disable iff (!rst_n)
(!cfg_valid || cfg_out_of_range) |-> (!over_provisioned && !under_requested);
endproperty
assert property (p_no_verdict_without_config)
else $error("a provisioning verdict was issued with no valid configuration");
// ---- Covers: both verdicts and the balanced case, per 31.1 §14's
// configuration rule. A block that has only ever reported one
// verdict has been exercised in one configuration.
cover property (@(posedge clk) disable iff (!rst_n) over_provisioned);
cover property (@(posedge clk) disable iff (!rst_n) under_requested);
cover property (@(posedge clk) disable iff (!rst_n)
cfg_valid && !over_provisioned && !under_requested);14. RTL Review — A Utilisation Report
The intended contract:
- Report utilisation as delivered beats over the beats the stream space could have delivered in the same interval.
- The denominator must be the provisioned capacity,
TOTAL_STREAMS— not the capacity that happened to be occupied, because the question is whether the provisioning was adequate. - Attribute a shortfall to a named cause — requester ceiling, stream saturation, or neither — per §11's partition.
- Report the demand denominator separately, so adequate provisioning is distinguishable from nobody asked.
- Never report a utilisation above 100%, and flag it if the inputs would imply one.
// ---------------------------------------------------------------------
// stream_utilisation -- INTENTIONALLY DEFECTIVE, for review (§14).
//
// CLASSIFICATION: synthesisable telemetry, grade-D values, CONTAINS
// A BUG.
//
// WHAT IT IS MEANT TO DO: the five-clause contract above -- report
// whether the provisioned stream space was used, so §7's
// counterfactual question is answerable.
//
// WHY IT EXISTS HERE: §7 argues the controller's most valuable output
// is evidence for the NEXT provisioning decision. This block is that
// evidence, computed against the wrong denominator -- and because
// the number it produces is plausible and high, it argues for
// exactly the wrong conclusion.
//
// HOW TO RUN IT: run a workload whose addresses concentrate on a
// quarter of the stream space (M4).
// EXPECTED RESULT under clause 2: utilisation is LOW, because three
// quarters of the provisioned streams delivered nothing.
// EXPECTED TRACE: the denominator must be TOTAL_STREAMS x window,
// independent of how many streams were busy.
//
// SYNTHESIS: two counters and a compare.
//
// LIMITATIONS: reports utilisation only; per-stack attribution is
// §12's. That is STATED and is not the bug.
// ---------------------------------------------------------------------
module stream_utilisation #(
parameter int NUM_STACKS = 8,
parameter int STREAMS_PER_STK = 16, // CURRICULUM-DERIVED, 26.4 §6
parameter int TOTAL_STREAMS = NUM_STACKS * STREAMS_PER_STK,
parameter int WIN = 1048576,
parameter int CNT_W = $clog2(WIN + 1),
parameter int STR_W = $clog2(TOTAL_STREAMS + 1),
// COUNT, not INDEX: the capacity accumulator must represent
// TOTAL_STREAMS x WIN, so it needs STR_W + CNT_W bits. One short
// and it wraps on a fully busy window -- reporting a LOW
// utilisation at exactly the saturated point the report is for.
parameter int CAP_W = STR_W + CNT_W
)(
input logic clk,
input logic rst_n,
input logic beat_delivered,
input logic [STR_W-1:0] streams_busy, // on the interface, and...
input logic demand_pending,
input logic win_tick,
output logic [CNT_W-1:0] r_beats,
output logic [CAP_W-1:0] r_capacity,
output logic [CNT_W-1:0] r_demand_cycles,
output logic over_unity,
output logic result_valid
);
logic [CNT_W-1:0] beats, dem;
logic [CAP_W-1:0] cap;
initial begin
if (NUM_STACKS < 1) $fatal(1, "stream_utilisation: NUM_STACKS >= 1");
end
always_ff @(posedge clk) begin
if (!rst_n) begin
beats <= '0; dem <= '0; cap <= '0;
r_beats <= '0; r_capacity <= '0; r_demand_cycles <= '0;
over_unity <= 1'b0; result_valid <= 1'b0;
end else if (win_tick) begin
r_beats <= beats;
r_capacity <= cap;
// Clause 4 is HONOURED: the demand denominator is reported
// separately, so a reviewer checking "can this distinguish
// adequate provisioning from an idle system?" finds that it
// can. That is what conceals the defect.
r_demand_cycles <= dem;
// Clause 5 is HONOURED: an impossible utilisation is flagged.
over_unity <= (beats > cap[CNT_W-1:0]);
result_valid <= 1'b1;
beats <= '0; dem <= '0; cap <= '0;
end else begin
result_valid <= 1'b0;
if (beat_delivered && beats != {CNT_W{1'b1}}) beats <= beats + 1'b1;
if (demand_pending && dem != {CNT_W{1'b1}}) dem <= dem + 1'b1;
// The capacity accumulator.
cap <= cap + {{(CAP_W-STR_W){1'b0}}, streams_busy}; // <-- THE DEFECT
end
end
endmoduleBefore reading on: which clause, and which mismatch does the resulting number point at?
15. The Defect — The Denominator That Moves With the Numerator
The violated clause is 2. The capacity accumulator sums streams_busy — the streams that happened to be occupied — where the contract requires TOTAL_STREAMS, the streams that were provisioned.
So the denominator shrinks whenever the numerator does. A workload that uses a quarter of the stream space produces a quarter of the beats and a quarter of the capacity, and the ratio comes out near 100%.
GRADE D. NUM_STACKS = 8, STREAMS_PER_STK = 16, so
TOTAL_STREAMS = 128 (CURRICULUM-DERIVED structure, 26.4 §6).
a workload whose addresses concentrate on 32 streams (M4):
beats delivered over the window : 32 streams x W beats
contract denominator (clause 2) : 128 x W
DERIVED utilisation = 32/128 = 25.0% <- the TRUTH
this block's denominator : 32 x W (streams_busy)
DERIVED utilisation = 32/32 = 100.0% <- what it reports
and clause 5's over-unity flag never fires, because the two
quantities move together and the ratio is exactly 1.Now the review question: which mismatch does a 100% reading point at?
It points at M1 — the streams are saturated, buy more stacks. CURRICULUM-DERIVED from 26.4 §9's table, M1 is tall stacks, bandwidth-bound workload and the response to saturated streams is more bandwidth. So the next provisioning decision buys stacks.
The mismatch actually present was M4 — a concentrated distribution — which 26.4 §7 establishes is an address-mapping property and §4 row four establishes is not a provisioning error at all and is fixable in software.
DERIVED consequence, and it is the worst outcome in this module:
actual problem : M4, fixable after shipping, costs nothing
reported problem: M1, fixable only in silicon
action taken : provision more stacks next generation
result: the new silicon has MORE streams, the addresses still
concentrate on the same fraction, and the utilisation report
still reads 100%. The next decision buys stacks again.
the defect is self-confirming: the action it recommends does
not change the number it produces.Three properties make this the module's most consequential defect.
It is self-confirming, which none of the module's other defects are. Chapter 32.1 §13's wrong field position produced a channel imbalance that a channel-mix report would reveal; 32.3 §13's shared pool produced a starvation a per-population monitor would reveal. This one produces a number that survives the action it recommends, so each generation confirms the last generation's conclusion.
It honours the clause that makes it look careful. Clause 4's separate demand denominator is present and correct, and clause 5's over-unity flag is present and can never fire — CURRICULUM-DERIVED from 30.5 §11: a guard that another property makes unreachable is untested, not working, and here the unreachability is caused by the very defect the guard would have caught in a different form.
And it is 30.8 §10's denominator error in a new place. That chapter found an efficiency monitor whose denominator counted arrivals rather than outstanding work and therefore reported the opposite of the truth. This is the same class — and the recurrence is the point: a denominator that is derived from the same signal as the numerator produces a ratio that cannot fall.
The correction, and it is one term:
// CORRECTED. Clause 2: the denominator is the PROVISIONED
// capacity, which is a constant per cycle and independent of
// what was busy. The question is whether the provisioning was
// adequate, so the provisioning must be the denominator.
//
// streams_busy remains on the interface and is now genuinely
// unused HERE -- it belongs to §11's block, which needs it for
// the attribution partition. A signal that is right for one
// block and wrong for another is not a signal to delete; it is
// one to route correctly.
cap <= cap + TOTAL_STREAMS[CAP_W-1:0];And the general finding, stated so it can be checked by grep:
A utilisation denominator must be independent of its numerator. If both are derived from the same observation, the ratio cannot fall and the report cannot distinguish a busy system from a small one. Find every ratio in a telemetry block and ask what would have to happen for it to read low — if nothing would, the denominator is wrong.
16. SVA Review — Proving a Ratio Is Well-Formed
The properties written for §14's block, and all three are correct:
// Offered as "proves the utilisation report is sound". All PASS on
// the defective block.
property p_utilisation_never_over_unity;
@(posedge clk) disable iff (!rst_n)
result_valid |-> (r_beats <= r_capacity);
endproperty
assert property (p_utilisation_never_over_unity)
else $error("delivered beats exceeded the reported capacity");
property p_demand_reported_separately;
@(posedge clk) disable iff (!rst_n)
result_valid |-> (r_demand_cycles <= WIN);
endproperty
assert property (p_demand_reported_separately)
else $error("demand cycles exceeded the window");
property p_counters_cleared_each_window;
@(posedge clk) disable iff (!rst_n)
result_valid |=> (r_beats == $past(r_beats));
endproperty
assert property (p_counters_cleared_each_window)
else $error("a latched result changed outside a window tick");Q. All three pass. What have they proved?
That the ratio is well-formed — bounded, with a separately reported demand denominator, and latched cleanly. None of them says anything about what the denominator is. Chapter 30.9 §6's variety 2, and it is the fifth consecutive chapter in which the defect's key quantity is absent from the property set.
And the first property is worse than merely silent — it is guaranteed by the defect. With the denominator derived from the same observation as the numerator, r_beats ≤ r_capacity cannot fail, so p_utilisation_never_over_unity is a property whose antecedent-free claim is made true by the bug. CURRICULUM-DERIVED from 30.5 §11: a guard another mechanism always satisfies is untested, and here the mechanism satisfying it is the defect.
What actually covers clause 2:
// Clause 2. It names the PROVISIONED constant, so a denominator
// derived from streams_busy cannot satisfy it. Variety 2's repair,
// and note that the property is an equality rather than a bound --
// a capacity accumulator has exactly one correct value per cycle.
property p_capacity_is_provisioned_not_observed;
@(posedge clk) disable iff (!rst_n)
result_valid |-> (r_capacity == (TOTAL_STREAMS * $past(WIN_ELAPSED)));
endproperty
assert property (p_capacity_is_provisioned_not_observed)
else $error("the capacity denominator was not the provisioned capacity");
// The INDEPENDENCE property, which is the general form and the one
// worth carrying: the denominator must not move when only the
// numerator's driver moves. This is the property that catches the
// whole class §15 names, not just this instance.
property p_denominator_independent_of_numerator;
@(posedge clk) disable iff (!rst_n)
($changed(streams_busy) && !win_tick) |=> $stable(cap_delta);
endproperty
assert property (p_denominator_independent_of_numerator)
else $error("the capacity denominator moved with the busy-stream count");
// §11's partition, asserted as exhaustive and disjoint -- 23.2 §9's
// standard. Without this the three attribution counters are three
// loosely related numbers rather than a partition, and §7's
// counterfactual cannot be read off them.
property p_attribution_is_a_partition;
@(posedge clk) disable iff (!rst_n)
result_valid |->
((c_bound_requester + c_bound_streams + c_bound_neither) == c_demand_cycles);
endproperty
assert property (p_attribution_is_a_partition)
else $error("the provisioning attribution does not sum to the demand cycles");
// Clause 3's mutual exclusion, stated separately because a
// partition needs both properties: they sum correctly AND no cycle
// is charged twice.
property p_attribution_is_disjoint;
@(posedge clk) disable iff (!rst_n)
demand_pending |-> $onehot({at_req_ceiling && !streams_saturated,
streams_saturated,
!at_req_ceiling && !streams_saturated});
endproperty
assert property (p_attribution_is_disjoint)
else $error("a demand cycle was attributed to more than one cause");
// Saturation must never be reached, because a saturated
// attribution counter understates a cause -- and §15 shows the
// action taken depends on WHICH cause dominates, so an
// understatement changes a silicon decision. 30.8 §11's range
// rule, with the consequence named.
property p_no_attribution_saturation;
@(posedge clk) disable iff (!rst_n)
result_valid |-> !counter_saturated;
endproperty
assert property (p_no_attribution_saturation)
else $error("an attribution counter saturated: a provisioning conclusion is understated");p_denominator_independent_of_numerator is the section's deliverable and it generalises past this block.
For every ratio a telemetry block reports, assert that the denominator does not move when the numerator's driver moves. That single property catches the whole class — 30.8 §10's arrival-counting efficiency monitor, 32.1 §14's channel-mix denominator, and §15's capacity accumulator — and it is cheaper to write than any of the three defects was to find.
// ---- Covers. The dimension this defect scales with is the
// ADDRESS DISTRIBUTION, and 31.3 §15's rule says the cover
// must be on that dimension.
// A CONCENTRATED distribution -- the case that separates the
// correct denominator from the defective one. A uniform-address
// testbench never produces it, and the defect is invisible under
// uniform traffic because busy == provisioned.
cover property (@(posedge clk) disable iff (!rst_n)
demand_pending && (streams_busy * 4 < TOTAL_STREAMS));
// A UNIFORM distribution, so both regimes are known to have run.
cover property (@(posedge clk) disable iff (!rst_n)
demand_pending && (streams_busy >= TOTAL_STREAMS));
// Each of §11's three attributions reached, per the per-dimension
// rule -- an attribution counter that never incremented proves
// nothing about the cause it represents.
cover property (@(posedge clk) disable iff (!rst_n)
demand_pending && at_req_ceiling && !streams_saturated);
cover property (@(posedge clk) disable iff (!rst_n)
demand_pending && streams_saturated);
// The "neither" case, which §11 argues is the strongest available
// argument for provisioning DOWN and which a bandwidth-focused
// report would never produce.
cover property (@(posedge clk) disable iff (!rst_n)
demand_pending && !at_req_ceiling && !streams_saturated);
// And §12's discriminator: exactly one stack materially low.
cover property (@(posedge clk) disable iff (!rst_n) single_stack_low);The first cover is the one that finds the defect and a uniform-address testbench cannot reach it. Under uniform traffic streams_busy equals TOTAL_STREAMS, so the wrong denominator and the right one are numerically identical — the defect is not merely hidden, it is absent in that regime. CURRICULUM-DERIVED from 31.4 §14's rule: when the missing dimension belongs to the environment's model rather than its stimulus, running longer never reaches it, and an address generator that distributes uniformly is exactly such a model.
Follow-up an interviewer should ask: should the testbench generate concentrated addresses? Yes, and it is cheap — CURRICULUM-DERIVED from 26.4 §7, a concentrated distribution is a real workload behaviour that chapter names explicitly, so this is not injecting an illegal stimulus; it is generating a legal one the default generator happens not to produce. That distinction is the one 30.4 §9 draws between block-level and system-level injection, and here both levels may legitimately do it.
17. What the Counterfactual Block's Assertions Prove
§16 reviewed the utilisation report. §11's block is the one whose output drives a silicon decision, and its own contract has not yet been asserted.
// The partition, from both directions. §11's whole value is that
// the three counters partition the demand cycles -- 23.2 §9's
// exhaustive-and-disjoint standard -- so a reader can attribute
// every demand cycle to one named cause.
property p_partition_sums_to_demand;
@(posedge clk) disable iff (!rst_n)
result_valid |->
((c_bound_requester + c_bound_streams + c_bound_neither) == c_demand_cycles);
endproperty
assert property (p_partition_sums_to_demand)
else $error("the attribution counters do not sum to the demand cycles");
// Nothing may be attributed on a cycle with no demand. This is the
// property that keeps §11's denominator honest: 30.8 §10's
// denominator error was counting the wrong cycles, and this
// forbids the same mistake here.
property p_no_attribution_without_demand;
@(posedge clk) disable iff (!rst_n)
(!demand_pending) |=> ($stable(c_bound_requester) &&
$stable(c_bound_streams) &&
$stable(c_bound_neither));
endproperty
assert property (p_no_attribution_without_demand)
else $error("a cycle with no demand was attributed to a provisioning cause");
// The two observed maxima must respect their provisioned ceilings,
// or the block is reporting a state the configuration forbids --
// 30.8 §11's range rule, and here an out-of-range maximum means
// the ceiling parameter disagrees with the hardware.
property p_maxima_within_provisioning;
@(posedge clk) disable iff (!rst_n)
result_valid |-> ((outstanding_max <= REQ_CEILING) &&
(streams_busy_max <= TOTAL_STREAMS));
endproperty
assert property (p_maxima_within_provisioning)
else $error("an observed maximum exceeded its provisioned ceiling");
// 26.4 §8's table as a runtime invariant: if the requester ceiling
// cannot touch every stream once, the streams can never ALL be
// busy. A design reporting otherwise has a counter or a parameter
// that disagrees with the hardware.
property p_ceiling_bounds_stream_occupancy;
@(posedge clk) disable iff (!rst_n)
(REQ_CEILING < TOTAL_STREAMS) |-> (streams_busy <= REQ_CEILING);
endproperty
assert property (p_ceiling_bounds_stream_occupancy)
else $error("more streams were busy than the requester can have outstanding");
// And the invariant that makes the counterfactual TRUSTWORTHY
// rather than merely present: the block must never attribute to
// BOTH halves at once, because §7's whole use of it is to decide
// which half to buy next time.
property p_never_bound_by_both;
@(posedge clk) disable iff (!rst_n)
!(at_req_ceiling && streams_saturated && demand_pending) ||
((c_bound_requester == $past(c_bound_requester)) ||
(c_bound_streams == $past(c_bound_streams)));
endproperty
assert property (p_never_bound_by_both)
else $error("a cycle was attributed to both provisioning halves");
// ---- Covers: the configuration arms, per 31.1 §14's rule, since
// REQ_CEILING versus TOTAL_STREAMS is exactly the
// configuration 26.4 §8's table is about.
cover property (@(posedge clk) disable iff (!rst_n)
result_valid && (REQ_CEILING >= TOTAL_STREAMS));
cover property (@(posedge clk) disable iff (!rst_n)
result_valid && (REQ_CEILING < TOTAL_STREAMS));
// A window in which the requester half dominated, and one in which
// the memory half did -- both must occur or the counterfactual has
// only ever produced one answer.
cover property (@(posedge clk) disable iff (!rst_n)
result_valid && (c_bound_requester > c_bound_streams));
cover property (@(posedge clk) disable iff (!rst_n)
result_valid && (c_bound_streams > c_bound_requester));p_never_bound_by_both is the property that makes the counterfactual usable rather than merely correct. §7's entire purpose is to decide which half to buy, and a cycle attributed to both halves gives no guidance. CURRICULUM-DERIVED from 30.5 §3: an unattributed outcome is a loop that found something rather than a reportable fact — and a doubly-attributed one is worse, because it looks like evidence for whichever half the reader already preferred.
And p_ceiling_bounds_stream_occupancy encodes 26.4 §8's table as a runtime check. If the requester cannot have TOTAL_STREAMS accesses outstanding, it is arithmetically impossible for every stream to be busy — so a report saying otherwise means a counter, a parameter or the hardware disagree. That is the cheapest available check that the block's configuration matches the silicon it is measuring.
18. What Would You Measure?
Q. You own the memory subsystem and the next generation's provisioning is under discussion. What do you measure, in what order?
| Measurement | What it settles | Cost | Owner |
|---|---|---|---|
Ask for c/V — the memory's share of system cost | §5's asymmetry has a sign and this sets it | hours | §9 |
Check whether REQ_CEILING ≥ TOTAL_STREAMS | §6 step 4 — whether provisioning up can even be filled | hours | 26.4 §8 |
| Check the utilisation denominator in the existing telemetry | §15's defect, answerable by reading one line | hours | §15 |
c_bound_requester vs c_bound_streams vs c_bound_neither | which half to buy, and whether to buy at all | days | §11 |
c_bound_neither specifically | the strongest argument for provisioning down | days | §11 |
| Concentration of the address distribution | M4, which is fixable in software and not a provisioning error | days | 26.4 §7, §16 |
single_stack_low from §12 | M5 — and it is grade C until promoted | days | §12 |
| Remap the addresses of the low stack and re-measure | promotes M5 from grade C: if it follows the addresses it was M4 | days | §8, §12 |
| Arithmetic intensity of the candidate workloads | whether bandwidth is the binding resource at all | weeks | 26.4 §5 |
Row one is a number you request rather than measure, and §5's conclusion changes sign on it. That is the second platform class in this module where the deciding quantity comes from outside the memory team — CURRICULUM-DERIVED from 32.2 §9, where theta had the same property. A memory team that does not ask has silently assumed a value, and for c/V the two possible assumptions recommend opposite actions.
Row three costs an hour and protects a generation. §15's defect is self-confirming, so the earlier it is found the better — and it is found by reading the denominator, not by running anything.
And row eight is the experiment that promotes §12's grade-C finding. DERIVED from §8's table: if the under-delivery follows the addresses when they are remapped, it was M4; if it stays with the stack, it is M5. One experiment separates a software fix from a scrapped package, and CURRICULUM-DERIVED from 31.3 §16, on this attach the package really is scrap — so the experiment is worth a week of anyone's time.
19. Common Wrong Answers
“Buy as much memory bandwidth as the budget allows.” §5, §6. The asymmetry has a sign that flips on the memory's share of system cost, and the bias is bounded by the requester's own parallelism.
“Provision for the expected workload.” §5. The expectation is the wrong target when the two errors cost differently — and they do, by a factor that depends on c/V.
“Over-provisioning is the safe choice.” §5. Only while memory is less than about half the system cost. Above that the bias reverses, and the model says so.
“More stacks always help.” §6, and 26.4 §8 owns the table: a 12-stack system needs 192 concurrent outstanding accesses just to touch every stream once. A 128-request requester buys 8 stacks of bandwidth and 12 of cost.
“The requester's parallelism can be raised in firmware.” §3, and 26.4 §8 owns the correction: an accelerator cannot choose its memory-level parallelism at runtime — it is set by the execution units and outstanding-request slots it was designed with.
“A fraction of the bandwidth means we under-provisioned.” §4, §8. Three other causes produce the same symptom: too few outstanding requests, a concentrated distribution, and a degraded stack. Only the first is a provisioning error.
“The distribution is the memory's problem to fix.” §4, and 26.4 §7 owns it as an address-mapping property — which means it is fixable after shipping, unlike the provisioning.
“A degraded stack would show up as an error.” §8. It shows up as a fraction of the bandwidth, identically to a bad distribution, and 26.3's repair happens at manufacture with none in the field.
“One stack low means the stack is degraded.” §12, §18. It is grade C until the remap experiment: if the under-delivery follows the addresses it was the distribution.
“This generation's decision is what matters.” §7. It is already fixed. The only open decision is the next one, and the controller's instrumentation is the only mechanism that informs it.
“Instrumentation is a debug feature we can cut.” §7. Its absence guarantees the next provisioning bet is placed on the same estimate as this one — a cost measured in generations, not in reproduction hours.
“Utilisation is 100%, so we need more bandwidth.” §15. Check the denominator. If it is derived from what was busy, the ratio cannot fall and the report is self-confirming.
“streams_busy is on the interface, so the report is distribution-aware.” §15. It is used, and used as the denominator — which is worse than ignoring it, because it makes the wrong answer look measured.
“The utilisation assertions pass.” §16. They prove the ratio is well-formed and never name what the denominator is. The bounding property is guaranteed by the defect.
“We will find it with a uniform-address regression.” §16. Under uniform traffic the wrong denominator and the right one are numerically identical — the defect is absent, not hidden.
“c_bound_neither is a measurement error.” §11. It is the strongest available argument for provisioning down: demand pending while neither half was at its limit means neither purchase would have helped.
“This chapter should have named real accelerators.” §2, §9. Every fact that would require is unpublished, so the alternatives were grade D dressed as grade A or a case study of the decision. Chapter 18.4 §1's category drift is why the second was chosen.
“Module 26 already covered this.” §2. It owns the two purchases, the arithmetic, the requester-count property and the four mismatches. It does not cost them, does not address the commitment-under-uncertainty problem, and hands the distribution question off as a mapping property.
20. Self-Check
-
State what is fixed in silicon on this platform class that is adjustable on the four others, and name the chapter that owns the coupling.
-
List 26.4 §9's four mismatches, mark each as over- or under-provisioning, and identify the one that is not a provisioning error at all.
-
Derive §5's two cost expressions. State the condition under which under-provisioning is the more expensive error and compute the break-even for
f = 2. -
Recompute §5's penalties at
c/V = 0.40andf = 3. State which error is worse and by how much. -
Give §5's qualification about workloads that are not bandwidth-bound, and say what that adds to the model's required inputs.
-
Give §6's four-step procedure and explain what bounds the provisioning bias, citing the table it comes from.
-
State why the controller's instrumentation is a design deliverable on this platform class, and what its absence guarantees.
-
Give §8's three-cause table and the single structural discriminator that separates degradation from distribution.
-
Find the defect in §14 without reading §15. Then answer the harder question: which mismatch does the resulting number point at, and which was present?
-
Explain what makes §14's defect self-confirming, and why that is worse than the module's other defects.
-
State the general rule §16 derives about ratios in telemetry, and name the three defects across this curriculum that it catches.
21. Where This Goes
This platform class commits both halves of its memory architecture to silicon before the workload exists. The two provisioning errors do not cost the same, and the sign of the asymmetry depends on the memory's share of system cost — below about half, under-provisioning is worse and the rational point sits above the expectation, bounded by what the requester can fill; three further causes produce the same fraction of the bandwidth symptom, only one of which is a provisioning error; and because the current decision is closed, the controller's instrumentation exists to inform the next one.
Three results carry forward. A utilisation denominator must be independent of its numerator, or the ratio cannot fall and the report is self-confirming. A defect whose recommended action does not change the number it produces will be confirmed by every generation that acts on it. And the deciding quantity came from outside the memory team for the second time in this module — theta in 32.2 §9, c/V here — so an analysis that does not name its external input has silently assumed one.
Chapter 32.5 closes the module with the one platform class whose composition is not fixed in silicon. Every chapter so far has had a memory system decided at design time and operated as built. On a server platform the memory composition is chosen at deployment and changed during service — and DDR stops being the whole memory system and becomes one tier among several. The question is what a controller must report when it no longer owns the placement decision, and what being one tier costs on the DDR side.
Continue learning
Related tutorials
- Related topic
HBM Overview
HBM reaches hundreds of GB/s with a per-pin rate lower than DDR5's. It wins on width, not speed — and getting that width required changing the packaging, which adds a fourth design layer to array physics, device architecture and the interface.
- Related topic
Why HBM Exists
Connections scale with perimeter while capability scales with area, and the ratio falls as 4/L. The width that answers it divides into sixteen streams that preserve the 64-byte granule exactly.
- Related topic
2.5D Packaging for HBM
A bump pitch turns an area into a count, and the face-versus-edge advantage has the closed form L/4p. The surprise is that a full 1024-bit bump field uses about 2% of a die's face.
- Related topic
Through-Silicon Vias (TSVs)
At a per-via reliability of one in ten thousand, a twelve-high stack works 2.5% of the time. One spare per group of 64 removes 99.7% of that risk — and 48 spares tolerate between 2 and 48 failures.
Standards & specifications
- Governing standard
- JEDEC JESD79 (DDR SDRAM)(opens JEDEC Solid State Technology Association in a new tab)
Defines the DDR SDRAM device itself — signals, command encoding, mode registers, timing parameters and the initialisation sequence — one document per generation. Memory-controller microarchitecture, address-mapping policy, PHY training algorithms and board-level design are not specified by it.
This page also covers RTL structure, verification approach and debugging technique. Those are engineering practice built on the standard, not requirements the standard itself imposes.
Where this fits
Part of the DDR curriculum.
