SPI · Module 12
SDR, DDR Transfers, and Throughput
Why data on both edges is independent of lane width, why the byte assembly is identical in both modes, why the dummy phase does not halve so DDR gives under two times, and why DDR parts specify more dummy cycles than their SDR modes.
Chapter 12.2 multiplied the bits per edge. There is one more axis, and it multiplies the edges per period.
DDR moves data on both edges of every SCLK period, so it should double the throughput. It does not — it gives about 1.33× on a short transfer. What did it fail to halve?
The same thing width failed to shrink. The dummy phase is counted in clock periods, and DDR does not change how many periods a cycle count occupies.
1. Two Independent Axes
Width and data rate multiply, and it is worth being precise about why they are independent.
bits per SCLK period = lanes × (ddr ? 2 : 1)Width decides how many bits cross simultaneously. Rate decides how many times per period that happens. Neither constrains the other, so all eight combinations exist:
SDR DDR
1 lane 1 bit/period 2 bits/period
2 lanes 2 4
4 lanes 4 8
8 lanes 8 16An octal DDR part moves sixteen bits per SCLK period — two bytes. At 100 MHz that is 200 MB/s from a serial interface, which is the point at which "serial" stops being a useful description of what is happening.
2. The Byte Assembly Does Not Change
Here is the property the whole chapter rests on, and it is easy to miss because it sounds like a technicality.
A group of bits arrives per edge. In SDR, one edge per period carries data; in DDR, both do. But the group itself is the same size, the bit-to-lane mapping is the same (Chapter 12.2), and the order of groups is the same.
So the byte assembly is identical. Given the same sequence of groups, SDR and DDR produce the same bytes. The only difference is how long the sequence takes:
edges = 8 × n_bytes / lanes ← identical in both modes
sclk_periods = ddr ? edges / 2 : edges ← the ONLY thing DDR changesThat framing is worth insisting on because it collapses a confusing feature into one line of arithmetic. DDR is not a different protocol, a different encoding or a different bit order. It is the same groups in the same order occupying half as many periods — and §6's testbench proves it by replaying one group sequence through both modes and requiring identical bytes.
3. Watching Both Edges
One group per edge, or two per period
8 cyclesThe -- cells in the SDR row are the edges DDR reclaims. Note what the figure does not show: the groups are the same size in both rows. Widening them is a separate axis, and doing both fills the DDR row with wider pills rather than more of them.
4. What DDR Does Not Halve
The dummy phase is specified in SCLK cycles (Chapter 10.5), so DDR leaves it exactly as it was. Take four lanes, four bytes, eight dummy cycles:
edges data periods + dummy total
SDR 8 8 8 16
DDR 8 4 8 1216 → 12 periods, which is 1.33× — not 2×. The data periods did halve; the total did not, because eight of the sixteen were never data.
And on a one-byte transfer the effect is starker:
SDR: 2 data periods + 8 dummy = 10
DDR: 1 data period + 8 dummy = 9 → 1.11×1.11× from a feature described as double data rate. This is Chapter 12.1's Amdahl bound arriving from the other direction: there, the fixed phases limited what widening could win; here they limit what doubling the rate can win, and by the same arithmetic.
Removing the dummy phase confirms the diagnosis exactly. With zero dummy cycles on a sixteen-byte transfer, DDR gives 32 → 16 periods — precisely 2×. So the dummy phase is not merely a reason the speedup falls short; on these numbers it is the only one.
5. And DDR Usually Needs MORE Dummy Cycles
This is the part that turns a disappointing result into an actively awkward one.
Halving the time per bit halves the sampling budget. The device's internal path from array to pin has not changed, so a part that needed eight dummy cycles at single rate may need ten or twelve at double rate to present data with adequate setup. Real DDR parts specify exactly that.
So the honest comparison is not SDR-with-8 against DDR-with-8. It is:
SDR, 8 dummy, 4 bytes, 4 lanes: 8 + 8 = 16 periods
DDR, 12 dummy, 4 bytes, 4 lanes: 4 + 12 = 16 periods → 1.00×No gain at all on a four-byte transfer. The extra dummy cycles consumed the entire benefit.
The gain returns with length, because the dummy phase is a fixed cost:
64 bytes: SDR 128 + 8 = 136 DDR 64 + 12 = 76 → 1.79×
256 bytes: SDR 512 + 8 = 520 DDR 256 + 12 = 268 → 1.94×Which gives the rule. DDR is a large-transfer feature. On the long sequential reads that Chapter 12.5's execute-in-place generates it approaches its nominal 2×; on the short scattered accesses a register-polling driver generates it can be worth nothing whatever.
6. The Pairing Constraint
One structural detail with no SDR equivalent.
Groups arrive in pairs in DDR — one per edge of a period. So a transfer needing an odd number of groups leaves a period carrying only one, and what happens on the unused edge is a question the transaction has to answer.
1 byte at 8 lanes = 1 group → odd: one period, one edge used
1 byte at 4 lanes = 2 groups → even: one period, both edges
3 bytes at 8 lanes = 3 groups → oddDevices handle this by requiring even-length transfers in some modes, or by defining the trailing edge as ignored. Either way it is a constraint a controller must know about, and §6's gearbox reports it rather than silently rounding — because a transfer rounded up by one group has read a byte nobody asked for, and rounded down has lost one.
7. Building the DDR Gearbox — Three HDLs
The circuit
Circuit. A byte assembler clocked once per data-carrying edge.
State. An assembly register, a group counter, and an edge counter.
Datapath. Identical in both modes — §2's property expressed as the absence of a mode input anywhere in the assembly path. The mode affects only the reported period count.
Control. Counts groups into bytes and edges to completion.
Clock and reset. One tick per data-carrying edge; asynchronous active-low reset.
Enables. partial_period reports §6's odd group count in DDR.
Timing. One group per tick, byte_valid per completed byte.
Synthesis. One 8-bit register, two counters, a shifter.
Limitations. The edge-per-tick clocking is a modelling choice, not an implementation one. A real DDR receiver captures on both physical edges of one clock, which needs either a double-rate clock or explicit rising- and falling-edge registers — and that is where the timing difficulty of §1 actually lives. This block separates the counting from the electricals so the counting can be reasoned about; it is not a DDR input buffer.
That limitation is worth stating plainly rather than glossing: the block is honest about what DDR changes and says nothing about how to capture on a falling edge reliably, which is a physical-design problem rather than an arithmetic one.
// spi_ddr_gear.sv
//
// Chapter 12.4 -- single and double data rate on the same lanes.
//
// Widening the bus (Chapter 12.2) moves more bits per EDGE. Double data
// rate moves data on BOTH edges of each SCLK period, so it moves the same
// bits per edge and twice as many edges per period. The two are
// independent: an octal DDR part moves 16 bits per SCLK period.
//
// This block is clocked once per data-carrying EDGE, which makes the
// difference between the modes purely a counting relationship:
//
// edges = 8 * n_bytes / lanes -- the same in both modes
// sclk_periods = ddr ? edges / 2 : edges -- the only thing DDR changes
//
// Expressing it that way is the point. The byte assembly is IDENTICAL in
// both modes, so the same group sequence must produce the same bytes -- and
// the testbench proves that by replaying one sequence through both.
//
// THE DUMMY PHASE DOES NOT HALVE. It is specified in SCLK cycles (Chapter
// 10.5), so DDR leaves it untouched and the speedup is again below the
// nominal factor -- the same Amdahl argument as Chapter 12.1, arriving from
// the other direction. In practice DDR parts specify MORE dummy cycles than
// their SDR modes, because halving the time per bit tightens the sampling
// budget and the device needs the extra latency to meet it.
//
// A DDR transfer also needs an EVEN number of groups, because groups arrive
// in pairs -- one per edge of a period. An odd count leaves a period
// carrying a single group, which is reported rather than silently rounded.
module spi_ddr_gear #(
parameter int MAX_LANES = 8,
parameter int CNT_W = 16
) (
input logic clk, // one tick per data-carrying edge
input logic rst_n,
input logic [3:0] lanes, // 1, 2, 4 or 8
input logic ddr, // 0 one edge per period, 1 both
input logic start,
input logic [CNT_W-1:0] n_bytes,
input logic [5:0] dummy_cycles, // SCLK CYCLES -- not halved
// One group arrives per tick while busy.
input logic [MAX_LANES-1:0] group_in,
output logic [7:0] byte_out,
output logic byte_valid,
output logic [CNT_W-1:0] edges_used,
output logic [CNT_W-1:0] sclk_periods, // what DDR actually reduces
output logic [CNT_W-1:0] total_periods, // including the dummy phase
output logic partial_period, // an odd group count in DDR
output logic busy,
output logic done
);
logic [7:0] asm_sr; // assembling byte
logic [3:0] grp_left; // groups remaining in this byte
logic [CNT_W-1:0] edges_left;
logic [3:0] grp_per_byte;
function automatic logic [3:0] group_count(input logic [3:0] n);
case (n)
4'd1: group_count = 4'd8;
4'd2: group_count = 4'd4;
4'd4: group_count = 4'd2;
4'd8: group_count = 4'd1;
default: group_count = 4'd8;
endcase
endfunction
function automatic logic [2:0] lane_shift(input logic [3:0] n);
case (n)
4'd1: lane_shift = 3'd0;
4'd2: lane_shift = 3'd1;
4'd4: lane_shift = 3'd2;
4'd8: lane_shift = 3'd3;
default: lane_shift = 3'd0;
endcase
endfunction
wire [MAX_LANES-1:0] lane_mask = MAX_LANES'((1 << lanes) - 1);
// The total edge count, which does NOT depend on the data rate -- the
// same bits must cross the same lanes either way.
wire [CNT_W-1:0] n_edges = (n_bytes << 3) >> lane_shift(lanes);
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
asm_sr <= 8'h00;
grp_left <= 4'd0;
edges_left <= {CNT_W{1'b0}};
grp_per_byte <= 4'd8;
byte_out <= 8'h00;
byte_valid <= 1'b0;
edges_used <= {CNT_W{1'b0}};
sclk_periods <= {CNT_W{1'b0}};
total_periods <= {CNT_W{1'b0}};
partial_period <= 1'b0;
busy <= 1'b0;
done <= 1'b0;
end else begin
byte_valid <= 1'b0;
done <= 1'b0;
if (start && !busy) begin
asm_sr <= 8'h00;
grp_per_byte <= group_count(lanes);
grp_left <= group_count(lanes);
edges_left <= n_edges;
edges_used <= n_edges;
// The ONLY thing the data rate changes: how many SCLK
// periods those edges occupy.
sclk_periods <= ddr ? (n_edges >> 1) : n_edges;
// ...and the dummy phase is added unchanged, because it is
// specified in SCLK cycles rather than in bits.
total_periods <= (ddr ? (n_edges >> 1) : n_edges)
+ CNT_W'(dummy_cycles);
// A DDR transfer needs an even number of groups: they arrive
// in pairs, one per edge of a period. An odd count leaves a
// period carrying a single group.
partial_period <= ddr && (n_edges[0] == 1'b1);
busy <= (n_edges != {CNT_W{1'b0}});
end else if (busy) begin
// Byte assembly, identical in both modes -- which is the
// property that makes SDR and DDR interchangeable from the
// consumer's point of view.
asm_sr <= (asm_sr << lanes) | 8'(group_in & lane_mask);
if (grp_left == 4'd1) begin
byte_out <= 8'((asm_sr << lanes) | 8'(group_in & lane_mask));
byte_valid <= 1'b1;
grp_left <= grp_per_byte;
asm_sr <= 8'h00;
end else begin
grp_left <= grp_left - 4'd1;
end
if (edges_left == CNT_W'(1)) begin
busy <= 1'b0;
done <= 1'b1;
end else begin
edges_left <= edges_left - 1'b1;
end
end
end
end
endmodule// spi_ddr_gear_tb.sv
//
// The central property is EQUIVALENCE: the same group sequence pushed
// through SDR and DDR must produce the same bytes, because the byte
// assembly does not depend on the data rate. Only the SCLK period count
// differs, and it must differ by exactly two.
//
// The second property is the one the chapter is about: the dummy phase does
// not halve, so the end-to-end speedup is below two -- and well below it on
// a short transfer.
`timescale 1ns/1ps
module spi_ddr_gear_tb;
localparam int MAX_LANES = 8;
localparam int CNT_W = 16;
logic clk = 1'b0;
logic rst_n = 1'b0;
always #5 clk = ~clk;
logic [3:0] lanes = 4'd4;
logic ddr = 1'b0;
logic start = 1'b0;
logic [CNT_W-1:0] n_bytes = 16'd4;
logic [5:0] dummy_cycles = 6'd8;
logic [MAX_LANES-1:0] group_in = 8'h00;
logic [7:0] byte_out;
logic byte_valid;
logic [CNT_W-1:0] edges_used, sclk_periods, total_periods;
logic partial_period, busy, done;
int errors = 0;
spi_ddr_gear #(.MAX_LANES(MAX_LANES), .CNT_W(CNT_W)) dut (
.clk(clk), .rst_n(rst_n),
.lanes(lanes), .ddr(ddr),
.start(start), .n_bytes(n_bytes), .dummy_cycles(dummy_cycles),
.group_in(group_in),
.byte_out(byte_out), .byte_valid(byte_valid),
.edges_used(edges_used), .sclk_periods(sclk_periods),
.total_periods(total_periods), .partial_period(partial_period),
.busy(busy), .done(done)
);
// Received bytes, for the equivalence comparison.
logic [7:0] got [0:63];
int n_got;
always_ff @(posedge clk) begin
if (rst_n && byte_valid && n_got < 64) begin
got[n_got] <= byte_out;
n_got <= n_got + 1;
end
end
// The source pattern, as groups. Built once so both modes see exactly
// the same sequence -- which is what makes the comparison meaningful.
logic [MAX_LANES-1:0] src [0:511];
int n_src;
task automatic build_source(input int nl, input int nb);
int gpb, k;
logic [7:0] v;
begin
gpb = 8 / nl;
n_src = 0;
for (int b = 0; b < nb; b++) begin
v = 8'(8'h31 + b[7:0]); // a pattern with no symmetry
for (k = 0; k < gpb; k++) begin
// Most significant group first, matching Chapter 12.2.
src[n_src] = MAX_LANES'((v >> (8 - nl * (k + 1)))
& ((1 << nl) - 1));
n_src++;
end
end
end
endtask
task automatic run(input int nl, input bit is_ddr, input int nb,
input int dum);
int guard, idx;
begin
@(negedge clk);
lanes = 4'(nl); ddr = is_ddr; n_bytes = CNT_W'(nb);
dummy_cycles = 6'(dum);
n_got = 0;
start = 1'b1;
@(negedge clk);
start = 1'b0;
idx = 0; guard = 0;
while (busy && guard < 1024) begin
group_in = (idx < n_src) ? src[idx] : 8'h00;
idx++;
@(negedge clk);
guard++;
end
group_in = 8'h00;
@(negedge clk);
end
endtask
// Captured SDR results, for comparison against DDR.
logic [7:0] sdr_bytes [0:63];
int sdr_n;
logic [CNT_W-1:0] sdr_edges, sdr_periods, sdr_total;
initial begin
n_got = 0; n_src = 0;
repeat (3) @(negedge clk);
rst_n = 1'b1;
@(negedge clk);
// 1. SDR at four lanes, four bytes. Establish the baseline.
build_source(4, 4);
run(4, 1'b0, 4, 8);
if (n_got != 4) begin
$display(" FAIL: SDR produced %0d bytes, expected 4", n_got);
errors++;
end
for (int i = 0; i < n_got; i++) sdr_bytes[i] = got[i];
sdr_n = n_got;
sdr_edges = edges_used; sdr_periods = sclk_periods;
sdr_total = total_periods;
// The bytes must be the source pattern, not merely self-consistent.
for (int i = 0; i < 4; i++) begin
if (got[i] !== 8'(8'h31 + i[7:0])) begin
$display(" FAIL: SDR byte %0d is 0x%02h, expected 0x%02h",
i, got[i], 8'h31 + i[7:0]);
errors++;
end
end
$display(" 4 lanes SDR, 4 bytes: %0d edges, %0d SCLK periods, %0d with dummy",
edges_used, sclk_periods, total_periods);
// 2. DDR, THE SAME source sequence. Identical bytes, half the SCLK
// periods -- and the same edge count, because the same bits must
// still cross the same lanes.
run(4, 1'b1, 4, 8);
if (n_got != sdr_n) begin
$display(" FAIL: DDR produced %0d bytes, SDR produced %0d",
n_got, sdr_n);
errors++;
end
for (int i = 0; i < n_got; i++) begin
if (got[i] !== sdr_bytes[i]) begin
$display(" FAIL: DDR byte %0d is 0x%02h, SDR gave 0x%02h",
i, got[i], sdr_bytes[i]);
errors++;
end
end
if (edges_used !== sdr_edges) begin
$display(" FAIL: DDR used %0d edges, SDR used %0d -- these must match",
edges_used, sdr_edges);
errors++;
end
if (sclk_periods !== (sdr_periods >> 1)) begin
$display(" FAIL: DDR took %0d SCLK periods, expected half of %0d",
sclk_periods, sdr_periods);
errors++;
end
$display(" 4 lanes DDR, 4 bytes: %0d edges (same), %0d SCLK periods (half), %0d with dummy",
edges_used, sclk_periods, total_periods);
// 3. THE DUMMY PHASE DOES NOT HALVE. The data periods halved; the
// total did not, because the dummy count is in SCLK cycles.
if (total_periods !== (sclk_periods + 16'd8)) begin
$display(" FAIL: DDR total %0d is not data %0d plus dummy 8",
total_periods, sclk_periods);
errors++;
end
$display(" speedup with dummy: %0d -> %0d periods, which is %0d.%02dx -- not 2x",
sdr_total, total_periods,
sdr_total / total_periods,
((sdr_total * 100) / total_periods) % 100);
if ((sdr_total * 100) / total_periods >= 200) begin
$display(" FAIL: DDR achieved 2x or better -- the dummy phase must cost something");
errors++;
end
// 4. A SHORT transfer, where the dummy phase dominates and DDR
// buys even less. The same Amdahl argument as Chapter 12.1.
build_source(4, 1);
run(4, 1'b0, 1, 8);
sdr_total = total_periods;
run(4, 1'b1, 1, 8);
$display(" 1 byte, 4 lanes: %0d -> %0d periods with dummy, %0d.%02dx",
sdr_total, total_periods,
sdr_total / total_periods,
((sdr_total * 100) / total_periods) % 100);
if ((sdr_total * 100) / total_periods >= 130) begin
$display(" FAIL: on one byte DDR gave 1.3x or better");
errors++;
end
// 5. WITHOUT a dummy phase, DDR does give exactly two -- which
// isolates the dummy phase as the whole reason it does not.
build_source(4, 16);
run(4, 1'b0, 16, 0);
sdr_total = total_periods;
run(4, 1'b1, 16, 0);
if (total_periods !== (sdr_total >> 1)) begin
$display(" FAIL: with no dummy phase DDR gave %0d, expected half of %0d",
total_periods, sdr_total);
errors++;
end
$display(" no dummy phase: %0d -> %0d periods, exactly 2x",
sdr_total, total_periods);
// 6. EQUIVALENCE ACROSS EVERY WIDTH. The byte assembly must not
// depend on the data rate at any width.
for (int wi = 0; wi < 4; wi++) begin
int nl;
nl = (wi == 0) ? 1 : (wi == 1) ? 2 : (wi == 2) ? 4 : 8;
build_source(nl, 8);
run(nl, 1'b0, 8, 8);
for (int i = 0; i < n_got; i++) sdr_bytes[i] = got[i];
sdr_n = n_got; sdr_edges = edges_used; sdr_periods = sclk_periods;
run(nl, 1'b1, 8, 8);
if (n_got != sdr_n) begin
$display(" FAIL: %0d lanes: DDR gave %0d bytes, SDR %0d",
nl, n_got, sdr_n);
errors++;
end
for (int i = 0; i < n_got; i++) begin
if (got[i] !== sdr_bytes[i]) begin
$display(" FAIL: %0d lanes byte %0d differs: DDR 0x%02h, SDR 0x%02h",
nl, i, got[i], sdr_bytes[i]);
errors++;
end
if (got[i] !== 8'(8'h31 + i[7:0])) begin
$display(" FAIL: %0d lanes byte %0d is 0x%02h, source was 0x%02h",
nl, i, got[i], 8'h31 + i[7:0]);
errors++;
end
end
if (edges_used !== sdr_edges) begin
$display(" FAIL: %0d lanes: edge counts differ between modes", nl);
errors++;
end
if (sclk_periods !== (sdr_periods >> 1)) begin
$display(" FAIL: %0d lanes: DDR periods %0d, expected half of %0d",
nl, sclk_periods, sdr_periods);
errors++;
end
end
$display(" equivalence at 1, 2, 4 and 8 lanes: identical bytes, identical edges, half the periods");
// 7. AN ODD GROUP COUNT. Groups arrive in pairs in DDR, so an odd
// count leaves a period carrying one group -- reported rather
// than silently rounded. One byte at eight lanes is one group.
build_source(8, 1);
run(8, 1'b1, 1, 8);
if (!partial_period) begin
$display(" FAIL: one group in DDR was not reported as a partial period");
errors++;
end
$display(" 1 byte at 8 lanes in DDR: %0d edge, partial period reported",
edges_used);
// 8. An even count must NOT be reported, and SDR never is -- a
// single-edge mode has no pairing requirement at all.
build_source(8, 2);
run(8, 1'b1, 2, 8);
if (partial_period) begin
$display(" FAIL: two groups in DDR was reported as partial");
errors++;
end
build_source(8, 1);
run(8, 1'b0, 1, 8);
if (partial_period) begin
$display(" FAIL: SDR reported a partial period -- it has no pairing requirement");
errors++;
end
$display(" pairing: odd reported in DDR only, even never, SDR never");
if (errors == 0)
$display("PASS: the same group sequence produces identical bytes in SDR and DDR at every width and uses the same number of edges, DDR halves the SCLK periods the data occupies and nothing else, the dummy phase is unchanged so the end-to-end speedup is below two and far below it on a short transfer, removing the dummy phase restores exactly two, and an odd group count in DDR is reported rather than rounded");
else
$display("FAIL: %0d error(s)", errors);
$finish;
end
endmoduleThe testbench is built around the equivalence claim, and the construction matters as much as the checks.
The source sequence is built once, into an array, before either mode runs. Both modes then see exactly the same groups in the same order. A testbench that generated the stimulus separately for each mode could not distinguish "the two modes agree" from "the two generators agree".
The bytes are checked against the source pattern, not merely against each other. That distinction is the same one Chapter 12.2 drew about round trips: two modes agreeing proves consistency, not correctness, so the absolute check is required alongside.
Then the three quantitative properties. The edge count must be identical between modes — the same bits cross the same lanes either way. The period count must be exactly half. And the total must be the data periods plus the dummy count, which is the arithmetic of §4 stated as a check.
Two bound checks encode the chapter's claims. DDR with a dummy phase must give strictly less than 2×, and on one byte less than 1.3×. Then removing the dummy phase must restore exactly 2× — which is what isolates the dummy phase as the sole cause rather than one of several.
The equivalence is finally swept across all four widths, because the assembly path's mode-independence must hold at every group size and not just at four lanes.
// spi_ddr_gear.v
//
// Chapter 12.4 -- single and double data rate on the same lanes, in
// Verilog-2001.
//
// Widening the bus moves more bits per EDGE. Double data rate moves data on
// BOTH edges of each SCLK period, so it moves the same bits per edge and
// twice as many edges per period. The two are independent: an octal DDR part
// moves 16 bits per SCLK period.
//
// This block is clocked once per data-carrying EDGE, which makes the
// difference between the modes purely a counting relationship:
//
// edges = 8 * n_bytes / lanes -- the same in both modes
// sclk_periods = ddr ? edges / 2 : edges -- the only thing DDR changes
//
// The byte assembly is IDENTICAL in both modes, so the same group sequence
// must produce the same bytes.
//
// THE DUMMY PHASE DOES NOT HALVE. It is specified in SCLK cycles, so DDR
// leaves it untouched and the speedup is below the nominal factor. In
// practice DDR parts specify MORE dummy cycles than their SDR modes, because
// halving the time per bit tightens the sampling budget.
//
// A DDR transfer also needs an EVEN number of groups, because groups arrive
// in pairs -- one per edge of a period. An odd count is reported rather than
// silently rounded.
module spi_ddr_gear #(
parameter MAX_LANES = 8,
parameter CNT_W = 16
) (
input wire clk, // one tick per data-carrying edge
input wire rst_n,
input wire [3:0] lanes, // 1, 2, 4 or 8
input wire ddr, // 0 one edge per period, 1 both
input wire start,
input wire [CNT_W-1:0] n_bytes,
input wire [5:0] dummy_cycles, // SCLK CYCLES -- not halved
// One group arrives per tick while busy.
input wire [MAX_LANES-1:0] group_in,
output reg [7:0] byte_out,
output reg byte_valid,
output reg [CNT_W-1:0] edges_used,
output reg [CNT_W-1:0] sclk_periods, // what DDR actually reduces
output reg [CNT_W-1:0] total_periods, // including the dummy phase
output reg partial_period, // an odd group count in DDR
output reg busy,
output reg done
);
reg [7:0] asm_sr; // assembling byte
reg [3:0] grp_left; // groups remaining in this byte
reg [CNT_W-1:0] edges_left;
reg [3:0] grp_per_byte;
function [3:0] group_count;
input [3:0] n;
begin
case (n)
4'd1: group_count = 4'd8;
4'd2: group_count = 4'd4;
4'd4: group_count = 4'd2;
4'd8: group_count = 4'd1;
default: group_count = 4'd8;
endcase
end
endfunction
function [2:0] lane_shift;
input [3:0] n;
begin
case (n)
4'd1: lane_shift = 3'd0;
4'd2: lane_shift = 3'd1;
4'd4: lane_shift = 3'd2;
4'd8: lane_shift = 3'd3;
default: lane_shift = 3'd0;
endcase
end
endfunction
wire [MAX_LANES-1:0] lane_mask = (1 << lanes) - 1;
// The total edge count, which does NOT depend on the data rate -- the
// same bits must cross the same lanes either way.
wire [CNT_W-1:0] n_edges = (n_bytes << 3) >> lane_shift(lanes);
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
asm_sr <= 8'h00;
grp_left <= 4'd0;
edges_left <= {CNT_W{1'b0}};
grp_per_byte <= 4'd8;
byte_out <= 8'h00;
byte_valid <= 1'b0;
edges_used <= {CNT_W{1'b0}};
sclk_periods <= {CNT_W{1'b0}};
total_periods <= {CNT_W{1'b0}};
partial_period <= 1'b0;
busy <= 1'b0;
done <= 1'b0;
end else begin
byte_valid <= 1'b0;
done <= 1'b0;
if (start && !busy) begin
asm_sr <= 8'h00;
grp_per_byte <= group_count(lanes);
grp_left <= group_count(lanes);
edges_left <= n_edges;
edges_used <= n_edges;
// The ONLY thing the data rate changes: how many SCLK
// periods those edges occupy.
sclk_periods <= ddr ? (n_edges >> 1) : n_edges;
// ...and the dummy phase is added unchanged, because it is
// specified in SCLK cycles rather than in bits.
total_periods <= (ddr ? (n_edges >> 1) : n_edges)
+ dummy_cycles;
// A DDR transfer needs an even number of groups: they arrive
// in pairs, one per edge of a period.
partial_period <= ddr && (n_edges[0] == 1'b1);
busy <= (n_edges != {CNT_W{1'b0}});
end else if (busy) begin
// Byte assembly, identical in both modes -- the property
// that makes SDR and DDR interchangeable to the consumer.
asm_sr <= (asm_sr << lanes) | (group_in & lane_mask);
if (grp_left == 4'd1) begin
byte_out <= (asm_sr << lanes) | (group_in & lane_mask);
byte_valid <= 1'b1;
grp_left <= grp_per_byte;
asm_sr <= 8'h00;
end else begin
grp_left <= grp_left - 4'd1;
end
if (edges_left == 1) begin
busy <= 1'b0;
done <= 1'b1;
end else begin
edges_left <= edges_left - 1'b1;
end
end
end
end
endmodule// spi_ddr_gear_tb.v
//
// The same checks as the SystemVerilog testbench. The central property is
// EQUIVALENCE: the same group sequence pushed through SDR and DDR must
// produce the same bytes, because the byte assembly does not depend on the
// data rate. Only the SCLK period count differs, by exactly two.
//
// The second property is the chapter's point: the dummy phase does not
// halve, so the end-to-end speedup is below two.
`timescale 1ns/1ps
module spi_ddr_gear_tb;
parameter MAX_LANES = 8;
parameter CNT_W = 16;
reg clk;
reg rst_n;
reg [3:0] lanes;
reg ddr;
reg start;
reg [CNT_W-1:0] n_bytes;
reg [5:0] dummy_cycles;
reg [MAX_LANES-1:0] group_in;
wire [7:0] byte_out;
wire byte_valid;
wire [CNT_W-1:0] edges_used, sclk_periods, total_periods;
wire partial_period, busy, done;
integer errors;
integer i, wi, nl, b, k, gpb, idx, guard;
reg [7:0] v;
initial begin
clk = 1'b0; rst_n = 1'b0;
lanes = 4'd4; ddr = 1'b0; start = 1'b0;
n_bytes = 16'd4; dummy_cycles = 6'd8; group_in = 8'h00;
errors = 0;
end
always #5 clk = ~clk;
spi_ddr_gear #(.MAX_LANES(MAX_LANES), .CNT_W(CNT_W)) dut (
.clk(clk), .rst_n(rst_n),
.lanes(lanes), .ddr(ddr),
.start(start), .n_bytes(n_bytes), .dummy_cycles(dummy_cycles),
.group_in(group_in),
.byte_out(byte_out), .byte_valid(byte_valid),
.edges_used(edges_used), .sclk_periods(sclk_periods),
.total_periods(total_periods), .partial_period(partial_period),
.busy(busy), .done(done)
);
// Received bytes, for the equivalence comparison.
reg [7:0] got [0:63];
integer n_got;
initial n_got = 0;
always @(posedge clk) begin
if (rst_n && byte_valid && n_got < 64) begin
got[n_got] <= byte_out;
n_got <= n_got + 1;
end
end
// The source pattern, as groups. Built once so both modes see exactly
// the same sequence -- which is what makes the comparison meaningful.
reg [MAX_LANES-1:0] src [0:511];
integer n_src;
task build_source;
input integer nl_i;
input integer nb;
begin
gpb = 8 / nl_i;
n_src = 0;
for (b = 0; b < nb; b = b + 1) begin
v = 8'h31 + b[7:0]; // a pattern with no symmetry
for (k = 0; k < gpb; k = k + 1) begin
// Most significant group first, matching Chapter 12.2.
src[n_src] = (v >> (8 - nl_i * (k + 1))) & ((1 << nl_i) - 1);
n_src = n_src + 1;
end
end
end
endtask
task run;
input integer nl_i;
input is_ddr;
input integer nb;
input integer dum;
begin
@(negedge clk);
lanes = nl_i[3:0]; ddr = is_ddr; n_bytes = nb[CNT_W-1:0];
dummy_cycles = dum[5:0];
n_got = 0;
start = 1'b1;
@(negedge clk);
start = 1'b0;
idx = 0; guard = 0;
while (busy && guard < 1024) begin
if (idx < n_src) group_in = src[idx];
else group_in = 8'h00;
idx = idx + 1;
@(negedge clk);
guard = guard + 1;
end
group_in = 8'h00;
@(negedge clk);
end
endtask
// Captured SDR results, for comparison against DDR.
reg [7:0] sdr_bytes [0:63];
integer sdr_n;
reg [CNT_W-1:0] sdr_edges, sdr_periods, sdr_total;
initial begin
n_src = 0;
repeat (3) @(negedge clk);
rst_n = 1'b1;
@(negedge clk);
// 1. SDR at four lanes, four bytes. Establish the baseline.
build_source(4, 4);
run(4, 1'b0, 4, 8);
if (n_got != 4) begin
$display(" FAIL: SDR produced %0d bytes, expected 4", n_got);
errors = errors + 1;
end
for (i = 0; i < n_got; i = i + 1) sdr_bytes[i] = got[i];
sdr_n = n_got;
sdr_edges = edges_used; sdr_periods = sclk_periods;
sdr_total = total_periods;
// The bytes must be the source pattern, not merely self-consistent.
for (i = 0; i < 4; i = i + 1) begin
if (got[i] !== (8'h31 + i[7:0])) begin
$display(" FAIL: SDR byte %0d is 0x%02h, expected 0x%02h",
i, got[i], 8'h31 + i[7:0]);
errors = errors + 1;
end
end
$display(" 4 lanes SDR, 4 bytes: %0d edges, %0d SCLK periods, %0d with dummy",
edges_used, sclk_periods, total_periods);
// 2. DDR, THE SAME source sequence.
run(4, 1'b1, 4, 8);
if (n_got != sdr_n) begin
$display(" FAIL: DDR produced %0d bytes, SDR produced %0d",
n_got, sdr_n);
errors = errors + 1;
end
for (i = 0; i < n_got; i = i + 1) begin
if (got[i] !== sdr_bytes[i]) begin
$display(" FAIL: DDR byte %0d is 0x%02h, SDR gave 0x%02h",
i, got[i], sdr_bytes[i]);
errors = errors + 1;
end
end
if (edges_used !== sdr_edges) begin
$display(" FAIL: DDR used %0d edges, SDR used %0d -- these must match",
edges_used, sdr_edges);
errors = errors + 1;
end
if (sclk_periods !== (sdr_periods >> 1)) begin
$display(" FAIL: DDR took %0d SCLK periods, expected half of %0d",
sclk_periods, sdr_periods);
errors = errors + 1;
end
$display(" 4 lanes DDR, 4 bytes: %0d edges (same), %0d SCLK periods (half), %0d with dummy",
edges_used, sclk_periods, total_periods);
// 3. THE DUMMY PHASE DOES NOT HALVE.
if (total_periods !== (sclk_periods + 16'd8)) begin
$display(" FAIL: DDR total %0d is not data %0d plus dummy 8",
total_periods, sclk_periods);
errors = errors + 1;
end
$display(" speedup with dummy: %0d -> %0d periods, which is %0d.%02dx -- not 2x",
sdr_total, total_periods,
sdr_total / total_periods,
((sdr_total * 100) / total_periods) % 100);
if ((sdr_total * 100) / total_periods >= 200) begin
$display(" FAIL: DDR achieved 2x or better -- the dummy phase must cost something");
errors = errors + 1;
end
// 4. A SHORT transfer, where the dummy phase dominates.
build_source(4, 1);
run(4, 1'b0, 1, 8);
sdr_total = total_periods;
run(4, 1'b1, 1, 8);
$display(" 1 byte, 4 lanes: %0d -> %0d periods with dummy, %0d.%02dx",
sdr_total, total_periods,
sdr_total / total_periods,
((sdr_total * 100) / total_periods) % 100);
if ((sdr_total * 100) / total_periods >= 130) begin
$display(" FAIL: on one byte DDR gave 1.3x or better");
errors = errors + 1;
end
// 5. WITHOUT a dummy phase, DDR does give exactly two -- which
// isolates the dummy phase as the whole reason it does not.
build_source(4, 16);
run(4, 1'b0, 16, 0);
sdr_total = total_periods;
run(4, 1'b1, 16, 0);
if (total_periods !== (sdr_total >> 1)) begin
$display(" FAIL: with no dummy phase DDR gave %0d, expected half of %0d",
total_periods, sdr_total);
errors = errors + 1;
end
$display(" no dummy phase: %0d -> %0d periods, exactly 2x",
sdr_total, total_periods);
// 6. EQUIVALENCE ACROSS EVERY WIDTH.
for (wi = 0; wi < 4; wi = wi + 1) begin
if (wi == 0) nl = 1;
else if (wi == 1) nl = 2;
else if (wi == 2) nl = 4;
else nl = 8;
build_source(nl, 8);
run(nl, 1'b0, 8, 8);
for (i = 0; i < n_got; i = i + 1) sdr_bytes[i] = got[i];
sdr_n = n_got; sdr_edges = edges_used; sdr_periods = sclk_periods;
run(nl, 1'b1, 8, 8);
if (n_got != sdr_n) begin
$display(" FAIL: %0d lanes: DDR gave %0d bytes, SDR %0d",
nl, n_got, sdr_n);
errors = errors + 1;
end
for (i = 0; i < n_got; i = i + 1) begin
if (got[i] !== sdr_bytes[i]) begin
$display(" FAIL: %0d lanes byte %0d differs: DDR 0x%02h, SDR 0x%02h",
nl, i, got[i], sdr_bytes[i]);
errors = errors + 1;
end
if (got[i] !== (8'h31 + i[7:0])) begin
$display(" FAIL: %0d lanes byte %0d is 0x%02h, source was 0x%02h",
nl, i, got[i], 8'h31 + i[7:0]);
errors = errors + 1;
end
end
if (edges_used !== sdr_edges) begin
$display(" FAIL: %0d lanes: edge counts differ between modes", nl);
errors = errors + 1;
end
if (sclk_periods !== (sdr_periods >> 1)) begin
$display(" FAIL: %0d lanes: DDR periods %0d, expected half of %0d",
nl, sclk_periods, sdr_periods);
errors = errors + 1;
end
end
$display(" equivalence at 1, 2, 4 and 8 lanes: identical bytes, identical edges, half the periods");
// 7. AN ODD GROUP COUNT in DDR.
build_source(8, 1);
run(8, 1'b1, 1, 8);
if (!partial_period) begin
$display(" FAIL: one group in DDR was not reported as a partial period");
errors = errors + 1;
end
$display(" 1 byte at 8 lanes in DDR: %0d edge, partial period reported",
edges_used);
// 8. An even count must NOT be reported, and SDR never is.
build_source(8, 2);
run(8, 1'b1, 2, 8);
if (partial_period) begin
$display(" FAIL: two groups in DDR was reported as partial");
errors = errors + 1;
end
build_source(8, 1);
run(8, 1'b0, 1, 8);
if (partial_period) begin
$display(" FAIL: SDR reported a partial period");
errors = errors + 1;
end
$display(" pairing: odd reported in DDR only, even never, SDR never");
if (errors == 0)
$display("PASS: the same group sequence produces identical bytes in SDR and DDR at every width and uses the same number of edges, DDR halves the SCLK periods the data occupies and nothing else, the dummy phase is unchanged so the end-to-end speedup is below two and far below it on a short transfer, removing the dummy phase restores exactly two, and an odd group count in DDR is reported rather than rounded");
else
$display("FAIL: %0d error(s)", errors);
$finish;
end
endmodule-- spi_ddr_gear.vhd
--
-- Chapter 12.4 -- single and double data rate on the same lanes, in VHDL.
--
-- Widening the bus moves more bits per EDGE. Double data rate moves data on
-- BOTH edges of each SCLK period, so it moves the same bits per edge and
-- twice as many edges per period. The two are independent: an octal DDR part
-- moves 16 bits per SCLK period.
--
-- This block is clocked once per data-carrying EDGE, which makes the
-- difference between the modes purely a counting relationship:
--
-- edges = 8 * n_bytes / lanes -- the same in both modes
-- sclk_periods = ddr ? edges / 2 : edges -- the only thing DDR changes
--
-- The byte assembly is IDENTICAL in both modes, so the same group sequence
-- must produce the same bytes.
--
-- THE DUMMY PHASE DOES NOT HALVE. It is specified in SCLK cycles, so DDR
-- leaves it untouched and the speedup is below the nominal factor. In
-- practice DDR parts specify MORE dummy cycles than their SDR modes, because
-- halving the time per bit tightens the sampling budget.
--
-- A DDR transfer also needs an EVEN number of groups, because groups arrive
-- in pairs -- one per edge of a period.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_ddr_gear is
generic (
MAX_LANES : positive := 8;
CNT_W : positive := 16
);
port (
clk : in std_logic; -- one tick per data-carrying edge
rst_n : in std_logic;
lanes : in unsigned(3 downto 0); -- 1, 2, 4 or 8
ddr : in std_logic; -- 0 one edge per period, 1 both
start : in std_logic;
n_bytes : in unsigned(CNT_W - 1 downto 0);
dummy_cycles : in unsigned(5 downto 0); -- SCLK CYCLES
group_in : in unsigned(MAX_LANES - 1 downto 0);
byte_out : out unsigned(7 downto 0);
byte_valid : out std_logic;
edges_used : out unsigned(CNT_W - 1 downto 0);
sclk_periods : out unsigned(CNT_W - 1 downto 0);
total_periods : out unsigned(CNT_W - 1 downto 0);
partial_period : out std_logic;
busy : out std_logic;
done : out std_logic
);
end entity;
architecture rtl of spi_ddr_gear is
function group_count(n : unsigned(3 downto 0)) return natural is
begin
case to_integer(n) is
when 1 => return 8;
when 2 => return 4;
when 4 => return 2;
when 8 => return 1;
when others => return 8;
end case;
end function;
function lane_div(n : unsigned(3 downto 0)) return natural is
begin
case to_integer(n) is
when 1 => return 1;
when 2 => return 2;
when 4 => return 4;
when 8 => return 8;
when others => return 1;
end case;
end function;
signal asm_sr : unsigned(7 downto 0) := (others => '0');
signal grp_left : natural := 0;
signal gpb_r : natural := 8;
signal edges_left : natural := 0;
signal bo_r : unsigned(7 downto 0) := (others => '0');
signal bv_r : std_logic := '0';
signal eu_r : unsigned(CNT_W - 1 downto 0) := (others => '0');
signal sp_r : unsigned(CNT_W - 1 downto 0) := (others => '0');
signal tp_r : unsigned(CNT_W - 1 downto 0) := (others => '0');
signal pp_r : std_logic := '0';
signal busy_r : std_logic := '0';
signal done_r : std_logic := '0';
signal lane_mask : unsigned(MAX_LANES - 1 downto 0) := (others => '0');
begin
lane_mask <= to_unsigned(2 ** to_integer(lanes) - 1, MAX_LANES)
when to_integer(lanes) <= MAX_LANES
else to_unsigned(0, MAX_LANES);
gear : process (clk, rst_n)
variable n_edges : natural;
variable nxt : unsigned(7 downto 0);
begin
if rst_n = '0' then
asm_sr <= (others => '0');
grp_left <= 0;
gpb_r <= 8;
edges_left <= 0;
bo_r <= (others => '0');
bv_r <= '0';
eu_r <= (others => '0');
sp_r <= (others => '0');
tp_r <= (others => '0');
pp_r <= '0';
busy_r <= '0';
done_r <= '0';
elsif rising_edge(clk) then
bv_r <= '0';
done_r <= '0';
if start = '1' and busy_r = '0' then
-- The total edge count does NOT depend on the data rate --
-- the same bits must cross the same lanes either way.
n_edges := (to_integer(n_bytes) * 8) / lane_div(lanes);
asm_sr <= (others => '0');
gpb_r <= group_count(lanes);
grp_left <= group_count(lanes);
edges_left <= n_edges;
eu_r <= to_unsigned(n_edges, CNT_W);
-- The ONLY thing the data rate changes: how many SCLK
-- periods those edges occupy.
if ddr = '1' then
sp_r <= to_unsigned(n_edges / 2, CNT_W);
tp_r <= to_unsigned(n_edges / 2 +
to_integer(dummy_cycles), CNT_W);
-- A DDR transfer needs an even number of groups.
if (n_edges mod 2) = 1 then pp_r <= '1';
else pp_r <= '0'; end if;
else
sp_r <= to_unsigned(n_edges, CNT_W);
tp_r <= to_unsigned(n_edges +
to_integer(dummy_cycles), CNT_W);
pp_r <= '0';
end if;
if n_edges /= 0 then busy_r <= '1'; else busy_r <= '0'; end if;
elsif busy_r = '1' then
-- Byte assembly, identical in both modes -- the property
-- that makes SDR and DDR interchangeable to the consumer.
nxt := shift_left(asm_sr, to_integer(lanes)) or
resize(group_in and lane_mask, 8);
if grp_left = 1 then
bo_r <= nxt;
bv_r <= '1';
grp_left <= gpb_r;
asm_sr <= (others => '0');
else
asm_sr <= nxt;
grp_left <= grp_left - 1;
end if;
if edges_left = 1 then
busy_r <= '0';
done_r <= '1';
else
edges_left <= edges_left - 1;
end if;
end if;
end if;
end process;
byte_out <= bo_r;
byte_valid <= bv_r;
edges_used <= eu_r;
sclk_periods <= sp_r;
total_periods <= tp_r;
partial_period <= pp_r;
busy <= busy_r;
done <= done_r;
end architecture;-- spi_ddr_gear_tb.vhd
--
-- The same checks as the SystemVerilog and Verilog testbenches. The central
-- property is EQUIVALENCE: the same group sequence pushed through SDR and
-- DDR must produce the same bytes, because the byte assembly does not depend
-- on the data rate. Only the SCLK period count differs, by exactly two.
--
-- The second property is the chapter's point: the dummy phase does not
-- halve, so the end-to-end speedup is below two.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_ddr_gear_tb is
end entity;
architecture sim of spi_ddr_gear_tb is
constant MAX_LANES : positive := 8;
constant CNT_W : positive := 16;
signal clk : std_logic := '0';
signal rst_n : std_logic := '0';
signal halt : boolean := false;
signal lanes : unsigned(3 downto 0) := to_unsigned(4, 4);
signal ddr : std_logic := '0';
signal start : std_logic := '0';
signal n_bytes : unsigned(CNT_W - 1 downto 0) := to_unsigned(4, CNT_W);
signal dummy_cycles : unsigned(5 downto 0) := to_unsigned(8, 6);
signal group_in : unsigned(MAX_LANES - 1 downto 0) := (others => '0');
signal byte_out : unsigned(7 downto 0);
signal byte_valid : std_logic;
signal edges_used : unsigned(CNT_W - 1 downto 0);
signal sclk_periods : unsigned(CNT_W - 1 downto 0);
signal total_periods : unsigned(CNT_W - 1 downto 0);
signal partial_period : std_logic;
signal busy, done : std_logic;
type byte_arr is array (0 to 63) of unsigned(7 downto 0);
signal got : byte_arr;
signal n_got : natural := 0;
signal clr : std_logic := '0';
signal errors : natural := 0;
begin
clk <= not clk after 5 ns when not halt else '0';
dut : entity work.spi_ddr_gear
generic map (MAX_LANES => MAX_LANES, CNT_W => CNT_W)
port map (
clk => clk, rst_n => rst_n,
lanes => lanes, ddr => ddr,
start => start, n_bytes => n_bytes, dummy_cycles => dummy_cycles,
group_in => group_in,
byte_out => byte_out, byte_valid => byte_valid,
edges_used => edges_used, sclk_periods => sclk_periods,
total_periods => total_periods, partial_period => partial_period,
busy => busy, done => done
);
-- Collector. The clear arrives on its own signal because two processes
-- driving one signal is a multiple-driver error in VHDL.
collect : process (clk)
begin
if rising_edge(clk) then
if clr = '1' then
n_got <= 0;
elsif rst_n = '1' and byte_valid = '1' and n_got < 64 then
got(n_got) <= byte_out;
n_got <= n_got + 1;
end if;
end if;
end process;
stim : process
variable errs : natural := 0;
-- The source pattern, as groups. Built once so both modes see
-- exactly the same sequence.
type grp_arr is array (0 to 511) of unsigned(MAX_LANES - 1 downto 0);
variable src : grp_arr;
variable n_src : natural := 0;
variable sdr_bytes : byte_arr;
variable sdr_n : natural;
variable sdr_edges, sdr_periods, sdr_total : natural;
variable guard, idx, gpb, nl : natural;
variable v : unsigned(7 downto 0);
procedure build_source(nl_i : natural; nb : natural) is
variable g : natural;
variable w : unsigned(7 downto 0);
begin
g := 8 / nl_i;
n_src := 0;
for b in 0 to nb - 1 loop
w := to_unsigned((16#31# + b) mod 256, 8);
for k in 0 to g - 1 loop
-- Most significant group first, matching Chapter 12.2.
src(n_src) := resize(
shift_right(w, 8 - nl_i * (k + 1)) and
to_unsigned(2 ** nl_i - 1, 8), MAX_LANES);
n_src := n_src + 1;
end loop;
end loop;
end procedure;
procedure run(nl_i : natural; is_ddr : std_logic;
nb : natural; dum : natural) is
begin
wait until falling_edge(clk);
clr <= '1';
wait until falling_edge(clk);
clr <= '0';
lanes <= to_unsigned(nl_i, 4);
ddr <= is_ddr;
n_bytes <= to_unsigned(nb, CNT_W);
dummy_cycles <= to_unsigned(dum, 6);
start <= '1';
wait until falling_edge(clk);
start <= '0';
idx := 0; guard := 0;
while busy = '1' and guard < 1024 loop
if idx < n_src then group_in <= src(idx);
else group_in <= (others => '0'); end if;
idx := idx + 1;
wait until falling_edge(clk);
guard := guard + 1;
end loop;
group_in <= (others => '0');
wait until falling_edge(clk);
end procedure;
begin
for k in 0 to 2 loop
wait until falling_edge(clk);
end loop;
rst_n <= '1';
wait until falling_edge(clk);
-- 1. SDR at four lanes, four bytes. Establish the baseline.
build_source(4, 4);
run(4, '0', 4, 8);
if n_got /= 4 then
report " FAIL: SDR did not produce four bytes"; errs := errs + 1;
end if;
for i in 0 to n_got - 1 loop
sdr_bytes(i) := got(i);
end loop;
sdr_n := n_got;
sdr_edges := to_integer(edges_used);
sdr_periods := to_integer(sclk_periods);
sdr_total := to_integer(total_periods);
-- The bytes must be the source pattern, not merely self-consistent.
for i in 0 to 3 loop
if to_integer(got(i)) /= 16#31# + i then
report " FAIL: an SDR byte does not match the source pattern";
errs := errs + 1;
end if;
end loop;
report " 4 lanes SDR, 4 bytes: " &
integer'image(sdr_edges) & " edges, " &
integer'image(sdr_periods) & " SCLK periods, " &
integer'image(sdr_total) & " with dummy";
-- 2. DDR, THE SAME source sequence.
run(4, '1', 4, 8);
if n_got /= sdr_n then
report " FAIL: DDR produced a different number of bytes than SDR";
errs := errs + 1;
end if;
for i in 0 to n_got - 1 loop
if got(i) /= sdr_bytes(i) then
report " FAIL: a DDR byte differs from the SDR byte";
errs := errs + 1;
end if;
end loop;
if to_integer(edges_used) /= sdr_edges then
report " FAIL: the edge counts differ between the modes";
errs := errs + 1;
end if;
if to_integer(sclk_periods) /= sdr_periods / 2 then
report " FAIL: DDR did not halve the SCLK periods";
errs := errs + 1;
end if;
report " 4 lanes DDR, 4 bytes: " &
integer'image(to_integer(edges_used)) & " edges (same), " &
integer'image(to_integer(sclk_periods)) &
" SCLK periods (half), " &
integer'image(to_integer(total_periods)) & " with dummy";
-- 3. THE DUMMY PHASE DOES NOT HALVE.
if to_integer(total_periods) /= to_integer(sclk_periods) + 8 then
report " FAIL: the total is not the data periods plus the dummy count";
errs := errs + 1;
end if;
report " speedup with dummy: " & integer'image(sdr_total) & " -> " &
integer'image(to_integer(total_periods)) & " periods, x100 = " &
integer'image((sdr_total * 100) / to_integer(total_periods)) &
" -- not 200";
if (sdr_total * 100) / to_integer(total_periods) >= 200 then
report " FAIL: DDR achieved 2x or better"; errs := errs + 1;
end if;
-- 4. A SHORT transfer, where the dummy phase dominates.
build_source(4, 1);
run(4, '0', 1, 8);
sdr_total := to_integer(total_periods);
run(4, '1', 1, 8);
report " 1 byte, 4 lanes: " & integer'image(sdr_total) & " -> " &
integer'image(to_integer(total_periods)) &
" periods with dummy, x100 = " &
integer'image((sdr_total * 100) / to_integer(total_periods));
if (sdr_total * 100) / to_integer(total_periods) >= 130 then
report " FAIL: on one byte DDR gave 1.3x or better";
errs := errs + 1;
end if;
-- 5. WITHOUT a dummy phase, DDR does give exactly two.
build_source(4, 16);
run(4, '0', 16, 0);
sdr_total := to_integer(total_periods);
run(4, '1', 16, 0);
if to_integer(total_periods) /= sdr_total / 2 then
report " FAIL: with no dummy phase DDR did not give exactly 2x";
errs := errs + 1;
end if;
report " no dummy phase: " & integer'image(sdr_total) & " -> " &
integer'image(to_integer(total_periods)) &
" periods, exactly 2x";
-- 6. EQUIVALENCE ACROSS EVERY WIDTH.
for wi in 0 to 3 loop
case wi is
when 0 => nl := 1;
when 1 => nl := 2;
when 2 => nl := 4;
when others => nl := 8;
end case;
build_source(nl, 8);
run(nl, '0', 8, 8);
for i in 0 to n_got - 1 loop
sdr_bytes(i) := got(i);
end loop;
sdr_n := n_got;
sdr_edges := to_integer(edges_used);
sdr_periods := to_integer(sclk_periods);
run(nl, '1', 8, 8);
if n_got /= sdr_n then
report " FAIL: byte counts differ between the modes";
errs := errs + 1;
end if;
for i in 0 to n_got - 1 loop
if got(i) /= sdr_bytes(i) then
report " FAIL: a byte differs between the modes";
errs := errs + 1;
end if;
if to_integer(got(i)) /= 16#31# + i then
report " FAIL: a byte does not match the source pattern";
errs := errs + 1;
end if;
end loop;
if to_integer(edges_used) /= sdr_edges then
report " FAIL: the edge counts differ between the modes";
errs := errs + 1;
end if;
if to_integer(sclk_periods) /= sdr_periods / 2 then
report " FAIL: DDR did not halve the periods at this width";
errs := errs + 1;
end if;
end loop;
report " equivalence at 1, 2, 4 and 8 lanes: identical bytes, identical edges, half the periods";
-- 7. AN ODD GROUP COUNT in DDR.
build_source(8, 1);
run(8, '1', 1, 8);
if partial_period /= '1' then
report " FAIL: one group in DDR was not reported as partial";
errs := errs + 1;
end if;
report " 1 byte at 8 lanes in DDR: 1 edge, partial period reported";
-- 8. An even count must NOT be reported, and SDR never is.
build_source(8, 2);
run(8, '1', 2, 8);
if partial_period = '1' then
report " FAIL: two groups in DDR was reported as partial";
errs := errs + 1;
end if;
build_source(8, 1);
run(8, '0', 1, 8);
if partial_period = '1' then
report " FAIL: SDR reported a partial period"; errs := errs + 1;
end if;
report " pairing: odd reported in DDR only, even never, SDR never";
errors <= errs;
if errs = 0 then
report "PASS: the same group sequence produces identical bytes in SDR and DDR at every width and uses the same number of edges, DDR halves the SCLK periods the data occupies and nothing else, the dummy phase is unchanged so the end-to-end speedup is below two and far below it on a short transfer, removing the dummy phase restores exactly two, and an odd group count in DDR is reported rather than rounded";
else
report "FAIL: " & integer'image(errs) & " error(s)" severity error;
end if;
halt <= true;
wait;
end process;
end architecture;Parity
All three implement the same gearbox: identical ports and generics, an assembly path with no dependence on the mode, an edge count independent of the rate, a period count halved only in DDR, a dummy count added unchanged, and an odd group count reported in DDR only. All three testbenches build the same source sequence, replay it through both modes at all four widths, and report identical figures — 1.33× on four bytes, 1.11× on one, and exactly 2× with no dummy phase.
8. Why a Verification Engineer Cares
// 1. MODE EQUIVALENCE. The assembled byte does not depend on the rate.
// This is the property the design's structure should make
// unfalsifiable -- the assembly path has no mode input at all -- and
// stating it anyway documents that the structure is doing the work.
a_rate_irrelevant : assert property (
@(posedge clk) disable iff (!rst_n)
(byte_valid) |-> (byte_out == expected_from_groups))
else $error("the assembled byte depended on the data rate");
// 2. THE EDGE COUNT IS RATE-INDEPENDENT. The same bits cross the same
// lanes either way, so a design reporting fewer edges in DDR has
// confused edges with periods -- which is the central confusion this
// chapter exists to prevent.
a_edges_invariant : assert property (
@(posedge clk) disable iff (!rst_n)
(start) |=> (edges_used == ((n_bytes * 8) / lanes)))
else $error("the edge count changed with the data rate");
// 3. AND THE PERIOD COUNT IS EXACTLY HALVED -- no rounding, which is why
// the odd case needs its own report rather than a division.
a_periods_halved : assert property (
@(posedge clk) disable iff (!rst_n)
(start && ddr) |=> (sclk_periods == (edges_used / 2)))
else $error("DDR did not halve the SCLK period count");
// 4. THE DUMMY PHASE IS UNCHANGED. Total minus data equals the dummy
// count, in both modes -- the arithmetic that bounds the speedup.
a_dummy_unchanged : assert property (
@(posedge clk) disable iff (!rst_n)
(start) |=> ((total_periods - sclk_periods) == dummy_cycles))
else $error("the dummy phase was scaled by the data rate");
// 5. THE ODD CASE IS REPORTED. A transfer rounded up by one group has
// read a byte nobody asked for; rounded down, it has lost one.
a_odd_reported : assert property (
@(posedge clk) disable iff (!rst_n)
(start && ddr && edges_used[0]) |=> partial_period)
else $error("an odd group count in DDR was not reported");Property 1 deserves the same comment Chapter 12.2's enable property got, because it is the same situation. The assembly path has no mode input, so the property cannot fail by construction — and that is not a reason to omit it. It records that the structure is carrying the guarantee, and it will fail loudly if someone later adds a mode-dependent branch "for the odd case".
Property 2 is the one that catches the real confusion. Edges and periods are different quantities, and DDR changes only the relationship between them. A design that reports halved edges has conflated them, and every downstream latency calculation will then be wrong by two in a direction that looks like an improvement.
Coverage must cross the rate with the length, because the speedup is a function of both:
covergroup spi_ddr_cg @(posedge clk iff start);
cp_ddr : coverpoint ddr { bins sdr = {0}; bins ddr = {1}; }
cp_lanes : coverpoint lanes {
bins one = {1}; bins two = {2}; bins four = {4}; bins eight = {8};
}
// Length decides how much of the transfer the dummy phase is, and
// therefore what DDR can win. The short bins are where it wins least
// and are exactly what a throughput-focused suite omits.
cp_len : coverpoint n_bytes {
bins one = {1}; // 1.11x -- the worst case
bins tiny = {[2:8]};
bins medium = {[9:64]};
bins large = {[65:$]}; // approaches 2x
}
// ZERO dummy is the case that isolates the dummy phase as the cause,
// and it is the only one in which DDR reaches exactly 2x.
cp_dummy : coverpoint dummy_cycles {
bins none = {0}; // exactly 2x
bins sdr_typ = {8};
bins ddr_typ = {[9:16]}; // what DDR parts actually specify
bins many = {[17:32]};
}
// The pairing constraint: an odd group count exists only in DDR.
cp_parity : coverpoint edges_used[0] {
bins even = {0};
bins odd = {1};
}
x_ddr_len : cross cp_ddr, cp_len;
x_ddr_dummy : cross cp_ddr, cp_dummy;
x_ddr_parity : cross cp_ddr, cp_parity;
endgroupcp_dummy's ddr_typ bin is the one that reflects reality and is almost always absent. A suite comparing SDR-with-8 against DDR-with-8 is comparing something no datasheet offers — real DDR modes specify more dummy cycles, and §5 showed that can erase the benefit entirely on short transfers. Covering the dummy count without covering the realistic DDR value measures a configuration that does not exist.
9. Why an FPGA or ASIC Engineer Cares
Do not confuse edges with periods. Edges are how much data crossed; periods are how long it took. DDR changes only the ratio, and a design that halves the edge count will get every latency calculation wrong in the flattering direction.
Compare against the device's ACTUAL DDR dummy count. SDR-with-8 against DDR-with-8 is a comparison of one real configuration against one imaginary one. Take both numbers from the command table.
Enable DDR for long sequential transfers and question it for short ones. 1.94× at 256 bytes and 1.00× at four, on representative numbers. The access pattern decides, not the feature list.
Handle the odd group count explicitly. Rounding up reads a byte nobody asked for — on a flash, possibly across a boundary; rounding down loses one. The device's rule is in its datasheet and belongs in the profile.
Budget the timing before the throughput. Halving the time per bit halves every margin in Chapter 10.3's table simultaneously. The throughput arithmetic is the easy part; the reason DDR arrived late is the other half.
Capture on both edges with explicit registers, not a doubled clock, if you can. Rising- and falling-edge registers feeding a common synchroniser are easier to constrain than a 2× clock domain, and they keep the double-rate behaviour confined to the pad logic rather than spreading it through the controller.
10. Failure Signature — DDR That Works Cold and Fails Warm on Long Transfers Only
Symptom. A DDR quad read works during bring-up. In the final enclosure it returns corrupted bytes — but only on transfers longer than a few hundred bytes, and only once the unit has been running for some minutes. Short transfers are always correct. Reverting to SDR fixes it completely.
What "short transfers always correct" establishes. This is the observation that steers the whole case, and its implication is not obvious. A constant timing violation would corrupt short and long transfers alike, since every bit has the same budget. Corruption that depends on transfer length means something is degrading during a transfer — so the fault is cumulative rather than static.
Plausible mechanisms.
- Supply droop during a long burst. Four lanes switching at double rate draw current in bursts; decoupling that is adequate for a short transfer sags over a long one, and the device's output timing degrades with its supply. This fits length-dependence and temperature-dependence together.
- Self-heating of the flash during sustained access, so
t_Vgrows as the transfer proceeds — which fits both length and the warm-enclosure condition. - Insufficient dummy cycles for DDR, per §5. But that would be a constant error and would corrupt short transfers too.
- A duty-cycle problem on SCLK. DDR samples on both edges, so an asymmetric clock gives one edge less setup than the other — and the corruption would fall on alternate groups. This does not obviously depend on length.
- A PLL or clock source drifting with temperature, which fits warm but not length.
The discriminating observation. Length-dependence and temperature-dependence together point at a cumulative effect, and the two candidates that fit both are supply droop and self-heating. They are distinguishable:
Insert idle gaps mid-transfer. Split a 1 KB read into eight 128-byte reads with a short pause between them. If the corruption disappears while the total data is unchanged, something is recovering during the gaps — which is droop or heating, not a static timing error. That test costs nothing and eliminates half the list.
Then separate droop from heating. Droop responds to decoupling and to lowering the clock; heating responds to time-at-temperature and shows hysteresis — it stays broken briefly after the transfer stops. If the first read after a long idle is always correct and corruption builds within a single transfer, droop is indicated; if corruption persists into the next transfer after a pause, heating is.
Check the duty cycle regardless, because DDR makes it matter in a way SDR never did: an asymmetric SCLK gives the two edges unequal budgets, and a 45/55 clock that was invisible at single rate removes 10% of the margin from every falling-edge group. That is worth measuring once even if it is not the cause here.
The fix depends on the mechanism — decoupling, a lower clock, more dummy cycles, or splitting long transfers — but the diagnostic order is what matters: establish whether the fault is static or cumulative before measuring anything, because the two lead to entirely different instruments.
Why DDR surfaces this and SDR did not. Not because DDR is fragile in itself, but because it halves every margin at once (§1). A board with 10% of margin to spare at single rate has none at double, so effects that were always present — droop, heating, duty-cycle asymmetry — become visible together. DDR did not cause the problem; it removed the slack that was hiding it.
11. Common Misconceptions
12. Reason It Through
Work this before reading the answer.
A controller can be configured three ways and only one change is affordable:
(a) 1-1-4 SDR, 8 dummy → 1-1-4 DDR, 12 dummy (b) 1-1-4 SDR, 8 dummy → 4-4-4 SDR, 8 dummy (c) 1-1-4 SDR, 8 dummy → 1-1-8 SDR (octal data), 10 dummy
The workload is 32-byte cache-line fetches with three address bytes. Which wins, and what changes if the workload becomes 4 KB block reads?
Baseline: 1-1-4 SDR, 8 dummy, 32 bytes.
cmd 8 + addr 24 + dummy 8 + data (256/4 = 64) = 104 periodsOption (a): DDR on four lanes, 12 dummy. Data edges are unchanged at 64; DDR halves the periods they occupy to 32. The command and address are on one lane and also halve in periods — DDR applies to every phase, not only data:
cmd 8/2 = 4 + addr 24/2 = 12 + dummy 12 + data 32 = 60 periods → 1.73×Option (b): 4-4-4 SDR, 8 dummy. Command and address widen to four lanes:
cmd 2 + addr 6 + dummy 8 + data 64 = 80 periods → 1.30×Option (c): octal data only, 10 dummy. Data halves again; command and address stay at one lane:
cmd 8 + addr 24 + dummy 10 + data (256/8 = 32) = 74 periods → 1.41×Option (a) wins at 1.73×, and the reason is worth extracting: DDR is the only one of the three that shortens the command and address phases without needing the device to support a wider command mode. It halves every phase measured in cycles-of-data, including the one-lane ones — which on a 32-byte fetch is 36 periods of overhead reduced to 16.
Now 4 KB blocks. Data dominates and the overhead becomes noise:
baseline: 8 + 24 + 8 + 8192 = 8232
(a) DDR: 4 + 12 + 12 + 4096 = 4124 → 2.00×
(b) 4-4-4: 2 + 6 + 8 + 8192 = 8208 → 1.00×
(c) octal: 8 + 24 + 10 + 4096 = 4138 → 1.99×(a) and (c) both reach essentially 2× and (b) achieves nothing. At 4 KB, widening the command and address is worth 24 periods out of 8232 — completely irrelevant — while halving the data phase is worth everything.
The comparison worth carrying:
32 bytes 4 KB
(a) DDR 1.73× 2.00×
(b) 4-4-4 1.30× 1.00×
(c) octal 1.41× 1.99×Two conclusions. Option (a) wins at both lengths, which makes it the right single choice — but it wins for different reasons: at 32 bytes because it shrinks overhead, at 4 KB because it shrinks data. And option (b) is worth 1.30× on short fetches and nothing on long ones, which inverts the usual assumption that a "more complete" mode is a better one.
The general lesson. Each technique attacks specific terms: width shrinks the phases measured in bits, DDR shrinks every phase measured in periods, and overhead removal — Chapter 12.5's subject — removes phases entirely. Which wins depends on where the periods currently are, so the arithmetic must come before the decision, and the answer changes with the workload rather than with the device.
13. Understanding Check
14. Summary
Width and data rate are independent axes that multiply: bits per period is lanes times two in DDR, so an octal DDR part moves two bytes per SCLK period.
The byte assembly is identical in both modes. Same group size, same lane mapping, same order — so the same group sequence must produce the same bytes, and DDR changes exactly one thing:
edges = 8 × n_bytes / lanes ← identical
sclk_periods = ddr ? edges / 2 : edges ← the only differenceConfusing edges with periods is the error to avoid, because it makes every latency calculation wrong by two in the flattering direction.
The dummy phase is counted in periods, so DDR does not halve it — 1.33× on four bytes, 1.11× on one. Removing the dummy phase restores exactly 2×, which isolates it as the sole cause.
And DDR modes usually specify more dummy cycles, because halving the time per bit halves the sampling budget. On representative numbers that gives 1.00× on four bytes and 1.94× at 256 — making DDR a large-transfer feature whose value is decided by the access pattern rather than by the datasheet.
Groups arrive in pairs, so an odd count leaves a period with one unused edge — reported rather than rounded, because rounding up reads an unasked-for byte and rounding down loses one.
DDR's real cost is not pins but half of every timing margin at once, which is why it removes the slack that was hiding effects always present on a board — and why a DDR bring-up failure is usually a board measurement rather than a protocol one.
15. What Comes Next
Width shrinks the phases measured in bits. DDR shrinks the phases measured in periods. Neither can remove a phase entirely — and that is the one thing left.
Chapter 12.5 — XIP and Memory-Mapped Controllers closes the module by removing the command and address phases outright. A sequential access continues a read stream the device already has open, so it pays no command, no address and no dummy phase — just data. XIP removes the command phase from even a random access. The result inverts everything this module has found: where fixed overhead limited what width and rate could win on short transfers, removing that overhead helps short transfers most — which is exactly the instruction-fetch pattern a processor generates, and the reason a serial flash can be code memory at all.
Continue learning
Related tutorials
- Related topic
Dummy Phases and Read Latency
Why a device needs turnaround before it can answer, why dummy is counted in clock cycles rather than bytes, how its length grows with frequency, and the one-byte data offset a mismatch produces.
- Related topic
Back-to-Back Transactions and Inter-Frame Gap
How soon chip select may fall again after it rises: the minimum deselect time, why a device needs it, what the gap costs in throughput, and the hardware enforcement that keeps software from violating it.
- Related topic
Raw vs Effective Payload Throughput
The three throughput numbers a link has — raw bit rate, protocol rate and effective payload rate — why they differ by four times on an ordinary flash read, and the performance counter that measures the gap instead of estimating it.
- Related topic
Multi-Byte and Back-to-Back Transactions
Keeping the wire busy between frames with one register's worth of flops: why one deep is enough and what a deeper FIFO actually buys, why a stall must be reported rather than absorbed, and why an overrun is lost data.
