Skip to content
VLSI Mentor

SPI · Module 12

SDR, DDR Transfers, and Throughput

Why data on both edges is independent of lane width, why the byte assembly is identical in both modes, why the dummy phase does not halve so DDR gives under two times, and why DDR parts specify more dummy cycles than their SDR modes.

Chapter 12.2 multiplied the bits per edge. There is one more axis, and it multiplies the edges per period.

DDR moves data on both edges of every SCLK period, so it should double the throughput. It does not — it gives about 1.33× on a short transfer. What did it fail to halve?

The same thing width failed to shrink. The dummy phase is counted in clock periods, and DDR does not change how many periods a cycle count occupies.

1. Two Independent Axes

Width and data rate multiply, and it is worth being precise about why they are independent.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   bits per SCLK period  =  lanes  ×  (ddr ? 2 : 1)

Width decides how many bits cross simultaneously. Rate decides how many times per period that happens. Neither constrains the other, so all eight combinations exist:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
             SDR          DDR
   1 lane    1 bit/period  2 bits/period
   2 lanes   2             4
   4 lanes   4             8
   8 lanes   8            16

An octal DDR part moves sixteen bits per SCLK period — two bytes. At 100 MHz that is 200 MB/s from a serial interface, which is the point at which "serial" stops being a useful description of what is happening.

2. The Byte Assembly Does Not Change

Here is the property the whole chapter rests on, and it is easy to miss because it sounds like a technicality.

A group of bits arrives per edge. In SDR, one edge per period carries data; in DDR, both do. But the group itself is the same size, the bit-to-lane mapping is the same (Chapter 12.2), and the order of groups is the same.

So the byte assembly is identical. Given the same sequence of groups, SDR and DDR produce the same bytes. The only difference is how long the sequence takes:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   edges         = 8 × n_bytes / lanes        ← identical in both modes
   sclk_periods  = ddr ? edges / 2 : edges    ← the ONLY thing DDR changes

That framing is worth insisting on because it collapses a confusing feature into one line of arithmetic. DDR is not a different protocol, a different encoding or a different bit order. It is the same groups in the same order occupying half as many periods — and §6's testbench proves it by replaying one group sequence through both modes and requiring identical bytes.

3. Watching Both Edges

One group per edge, or two per period

8 cycles
A clock toggling over eight half-periods. The single-rate row carries a group only on the rising edges, so four groups in four periods. The double-rate row carries a group on every edge, so eight groups in the same four periods.rising onlyrising onlyboth edgesboth edgesSCLKSDR--G0--G1--G2--G3DDRG0G1G2G3G4G5G6G7t0t1t2t3t4t5t6t7
Figure 1 — the same groups, two rates. SDR carries one group per SCLK period, on the rising edge only; DDR carries one on each edge, so it delivers eight groups in the four periods that SDR needs for four.

The -- cells in the SDR row are the edges DDR reclaims. Note what the figure does not show: the groups are the same size in both rows. Widening them is a separate axis, and doing both fills the DDR row with wider pills rather than more of them.

4. What DDR Does Not Halve

The dummy phase is specified in SCLK cycles (Chapter 10.5), so DDR leaves it exactly as it was. Take four lanes, four bytes, eight dummy cycles:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
                    edges   data periods   + dummy   total
   SDR                 8          8            8      16
   DDR                 8          4            8      12

16 → 12 periods, which is 1.33× — not 2×. The data periods did halve; the total did not, because eight of the sixteen were never data.

And on a one-byte transfer the effect is starker:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   SDR:  2 data periods + 8 dummy = 10
   DDR:  1 data period  + 8 dummy =  9      → 1.11×

1.11× from a feature described as double data rate. This is Chapter 12.1's Amdahl bound arriving from the other direction: there, the fixed phases limited what widening could win; here they limit what doubling the rate can win, and by the same arithmetic.

Removing the dummy phase confirms the diagnosis exactly. With zero dummy cycles on a sixteen-byte transfer, DDR gives 32 → 16 periods — precisely 2×. So the dummy phase is not merely a reason the speedup falls short; on these numbers it is the only one.

5. And DDR Usually Needs MORE Dummy Cycles

This is the part that turns a disappointing result into an actively awkward one.

Halving the time per bit halves the sampling budget. The device's internal path from array to pin has not changed, so a part that needed eight dummy cycles at single rate may need ten or twelve at double rate to present data with adequate setup. Real DDR parts specify exactly that.

So the honest comparison is not SDR-with-8 against DDR-with-8. It is:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   SDR, 8 dummy,  4 bytes, 4 lanes:   8 + 8 = 16 periods
   DDR, 12 dummy, 4 bytes, 4 lanes:   4 + 12 = 16 periods      → 1.00×

No gain at all on a four-byte transfer. The extra dummy cycles consumed the entire benefit.

The gain returns with length, because the dummy phase is a fixed cost:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   64 bytes:  SDR  128 + 8  = 136      DDR  64 + 12 = 76    → 1.79×
  256 bytes:  SDR  512 + 8  = 520      DDR 256 + 12 = 268   → 1.94×

Which gives the rule. DDR is a large-transfer feature. On the long sequential reads that Chapter 12.5's execute-in-place generates it approaches its nominal 2×; on the short scattered accesses a register-polling driver generates it can be worth nothing whatever.

6. The Pairing Constraint

One structural detail with no SDR equivalent.

Groups arrive in pairs in DDR — one per edge of a period. So a transfer needing an odd number of groups leaves a period carrying only one, and what happens on the unused edge is a question the transaction has to answer.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   1 byte at 8 lanes  = 1 group   → odd:  one period, one edge used
   1 byte at 4 lanes  = 2 groups  → even: one period, both edges
   3 bytes at 8 lanes = 3 groups  → odd

Devices handle this by requiring even-length transfers in some modes, or by defining the trailing edge as ignored. Either way it is a constraint a controller must know about, and §6's gearbox reports it rather than silently rounding — because a transfer rounded up by one group has read a byte nobody asked for, and rounded down has lost one.

7. Building the DDR Gearbox — Three HDLs

The circuit

Circuit. A byte assembler clocked once per data-carrying edge.

State. An assembly register, a group counter, and an edge counter.

Datapath. Identical in both modes — §2's property expressed as the absence of a mode input anywhere in the assembly path. The mode affects only the reported period count.

Control. Counts groups into bytes and edges to completion.

Clock and reset. One tick per data-carrying edge; asynchronous active-low reset.

Enables. partial_period reports §6's odd group count in DDR.

Timing. One group per tick, byte_valid per completed byte.

Synthesis. One 8-bit register, two counters, a shifter.

Limitations. The edge-per-tick clocking is a modelling choice, not an implementation one. A real DDR receiver captures on both physical edges of one clock, which needs either a double-rate clock or explicit rising- and falling-edge registers — and that is where the timing difficulty of §1 actually lives. This block separates the counting from the electricals so the counting can be reasoned about; it is not a DDR input buffer.

That limitation is worth stating plainly rather than glossing: the block is honest about what DDR changes and says nothing about how to capture on a falling edge reliably, which is a physical-design problem rather than an arithmetic one.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_ddr_gear.sv — the same assembly, half the periods
// spi_ddr_gear.sv
//
// Chapter 12.4 -- single and double data rate on the same lanes.
//
// Widening the bus (Chapter 12.2) moves more bits per EDGE. Double data
// rate moves data on BOTH edges of each SCLK period, so it moves the same
// bits per edge and twice as many edges per period. The two are
// independent: an octal DDR part moves 16 bits per SCLK period.
//
// This block is clocked once per data-carrying EDGE, which makes the
// difference between the modes purely a counting relationship:
//
//   edges         = 8 * n_bytes / lanes        -- the same in both modes
//   sclk_periods  = ddr ? edges / 2 : edges    -- the only thing DDR changes
//
// Expressing it that way is the point. The byte assembly is IDENTICAL in
// both modes, so the same group sequence must produce the same bytes -- and
// the testbench proves that by replaying one sequence through both.
//
// THE DUMMY PHASE DOES NOT HALVE. It is specified in SCLK cycles (Chapter
// 10.5), so DDR leaves it untouched and the speedup is again below the
// nominal factor -- the same Amdahl argument as Chapter 12.1, arriving from
// the other direction. In practice DDR parts specify MORE dummy cycles than
// their SDR modes, because halving the time per bit tightens the sampling
// budget and the device needs the extra latency to meet it.
//
// A DDR transfer also needs an EVEN number of groups, because groups arrive
// in pairs -- one per edge of a period. An odd count leaves a period
// carrying a single group, which is reported rather than silently rounded.

module spi_ddr_gear #(
    parameter int MAX_LANES = 8,
    parameter int CNT_W     = 16
) (
    input  logic                 clk,      // one tick per data-carrying edge
    input  logic                 rst_n,

    input  logic [3:0]           lanes,    // 1, 2, 4 or 8
    input  logic                 ddr,      // 0 one edge per period, 1 both

    input  logic                 start,
    input  logic [CNT_W-1:0]     n_bytes,
    input  logic [5:0]           dummy_cycles,   // SCLK CYCLES -- not halved

    // One group arrives per tick while busy.
    input  logic [MAX_LANES-1:0] group_in,

    output logic [7:0]           byte_out,
    output logic                 byte_valid,

    output logic [CNT_W-1:0]     edges_used,
    output logic [CNT_W-1:0]     sclk_periods,   // what DDR actually reduces
    output logic [CNT_W-1:0]     total_periods,  // including the dummy phase
    output logic                 partial_period, // an odd group count in DDR
    output logic                 busy,
    output logic                 done
);

    logic [7:0]       asm_sr;      // assembling byte
    logic [3:0]       grp_left;    // groups remaining in this byte
    logic [CNT_W-1:0] edges_left;
    logic [3:0]       grp_per_byte;

    function automatic logic [3:0] group_count(input logic [3:0] n);
        case (n)
            4'd1:    group_count = 4'd8;
            4'd2:    group_count = 4'd4;
            4'd4:    group_count = 4'd2;
            4'd8:    group_count = 4'd1;
            default: group_count = 4'd8;
        endcase
    endfunction

    function automatic logic [2:0] lane_shift(input logic [3:0] n);
        case (n)
            4'd1:    lane_shift = 3'd0;
            4'd2:    lane_shift = 3'd1;
            4'd4:    lane_shift = 3'd2;
            4'd8:    lane_shift = 3'd3;
            default: lane_shift = 3'd0;
        endcase
    endfunction

    wire [MAX_LANES-1:0] lane_mask = MAX_LANES'((1 << lanes) - 1);

    // The total edge count, which does NOT depend on the data rate -- the
    // same bits must cross the same lanes either way.
    wire [CNT_W-1:0] n_edges = (n_bytes << 3) >> lane_shift(lanes);

    always_ff @(posedge clk or negedge rst_n) begin
        if (!rst_n) begin
            asm_sr         <= 8'h00;
            grp_left       <= 4'd0;
            edges_left     <= {CNT_W{1'b0}};
            grp_per_byte   <= 4'd8;
            byte_out       <= 8'h00;
            byte_valid     <= 1'b0;
            edges_used     <= {CNT_W{1'b0}};
            sclk_periods   <= {CNT_W{1'b0}};
            total_periods  <= {CNT_W{1'b0}};
            partial_period <= 1'b0;
            busy           <= 1'b0;
            done           <= 1'b0;
        end else begin
            byte_valid <= 1'b0;
            done       <= 1'b0;

            if (start && !busy) begin
                asm_sr       <= 8'h00;
                grp_per_byte <= group_count(lanes);
                grp_left     <= group_count(lanes);
                edges_left   <= n_edges;
                edges_used   <= n_edges;

                // The ONLY thing the data rate changes: how many SCLK
                // periods those edges occupy.
                sclk_periods <= ddr ? (n_edges >> 1) : n_edges;

                // ...and the dummy phase is added unchanged, because it is
                // specified in SCLK cycles rather than in bits.
                total_periods <= (ddr ? (n_edges >> 1) : n_edges)
                                 + CNT_W'(dummy_cycles);

                // A DDR transfer needs an even number of groups: they arrive
                // in pairs, one per edge of a period. An odd count leaves a
                // period carrying a single group.
                partial_period <= ddr && (n_edges[0] == 1'b1);

                busy <= (n_edges != {CNT_W{1'b0}});
            end else if (busy) begin
                // Byte assembly, identical in both modes -- which is the
                // property that makes SDR and DDR interchangeable from the
                // consumer's point of view.
                asm_sr <= (asm_sr << lanes) | 8'(group_in & lane_mask);

                if (grp_left == 4'd1) begin
                    byte_out   <= 8'((asm_sr << lanes) | 8'(group_in & lane_mask));
                    byte_valid <= 1'b1;
                    grp_left   <= grp_per_byte;
                    asm_sr     <= 8'h00;
                end else begin
                    grp_left <= grp_left - 4'd1;
                end

                if (edges_left == CNT_W'(1)) begin
                    busy <= 1'b0;
                    done <= 1'b1;
                end else begin
                    edges_left <= edges_left - 1'b1;
                end
            end
        end
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_ddr_gear_tb.sv — one group sequence, both modes
// spi_ddr_gear_tb.sv
//
// The central property is EQUIVALENCE: the same group sequence pushed
// through SDR and DDR must produce the same bytes, because the byte
// assembly does not depend on the data rate. Only the SCLK period count
// differs, and it must differ by exactly two.
//
// The second property is the one the chapter is about: the dummy phase does
// not halve, so the end-to-end speedup is below two -- and well below it on
// a short transfer.

`timescale 1ns/1ps

module spi_ddr_gear_tb;

    localparam int MAX_LANES = 8;
    localparam int CNT_W     = 16;

    logic clk = 1'b0;
    logic rst_n = 1'b0;
    always #5 clk = ~clk;

    logic [3:0]           lanes = 4'd4;
    logic                 ddr = 1'b0;
    logic                 start = 1'b0;
    logic [CNT_W-1:0]     n_bytes = 16'd4;
    logic [5:0]           dummy_cycles = 6'd8;
    logic [MAX_LANES-1:0] group_in = 8'h00;

    logic [7:0]           byte_out;
    logic                 byte_valid;
    logic [CNT_W-1:0]     edges_used, sclk_periods, total_periods;
    logic                 partial_period, busy, done;

    int errors = 0;

    spi_ddr_gear #(.MAX_LANES(MAX_LANES), .CNT_W(CNT_W)) dut (
        .clk(clk), .rst_n(rst_n),
        .lanes(lanes), .ddr(ddr),
        .start(start), .n_bytes(n_bytes), .dummy_cycles(dummy_cycles),
        .group_in(group_in),
        .byte_out(byte_out), .byte_valid(byte_valid),
        .edges_used(edges_used), .sclk_periods(sclk_periods),
        .total_periods(total_periods), .partial_period(partial_period),
        .busy(busy), .done(done)
    );

    // Received bytes, for the equivalence comparison.
    logic [7:0] got [0:63];
    int         n_got;

    always_ff @(posedge clk) begin
        if (rst_n && byte_valid && n_got < 64) begin
            got[n_got] <= byte_out;
            n_got      <= n_got + 1;
        end
    end

    // The source pattern, as groups. Built once so both modes see exactly
    // the same sequence -- which is what makes the comparison meaningful.
    logic [MAX_LANES-1:0] src [0:511];
    int                   n_src;

    task automatic build_source(input int nl, input int nb);
        int gpb, k;
        logic [7:0] v;
        begin
            gpb = 8 / nl;
            n_src = 0;
            for (int b = 0; b < nb; b++) begin
                v = 8'(8'h31 + b[7:0]);     // a pattern with no symmetry
                for (k = 0; k < gpb; k++) begin
                    // Most significant group first, matching Chapter 12.2.
                    src[n_src] = MAX_LANES'((v >> (8 - nl * (k + 1)))
                                            & ((1 << nl) - 1));
                    n_src++;
                end
            end
        end
    endtask

    task automatic run(input int nl, input bit is_ddr, input int nb,
                       input int dum);
        int guard, idx;
        begin
            @(negedge clk);
            lanes = 4'(nl); ddr = is_ddr; n_bytes = CNT_W'(nb);
            dummy_cycles = 6'(dum);
            n_got = 0;
            start = 1'b1;
            @(negedge clk);
            start = 1'b0;
            idx = 0; guard = 0;
            while (busy && guard < 1024) begin
                group_in = (idx < n_src) ? src[idx] : 8'h00;
                idx++;
                @(negedge clk);
                guard++;
            end
            group_in = 8'h00;
            @(negedge clk);
        end
    endtask

    // Captured SDR results, for comparison against DDR.
    logic [7:0]       sdr_bytes [0:63];
    int               sdr_n;
    logic [CNT_W-1:0] sdr_edges, sdr_periods, sdr_total;

    initial begin
        n_got = 0; n_src = 0;
        repeat (3) @(negedge clk);
        rst_n = 1'b1;
        @(negedge clk);

        // 1. SDR at four lanes, four bytes. Establish the baseline.
        build_source(4, 4);
        run(4, 1'b0, 4, 8);
        if (n_got != 4) begin
            $display("  FAIL: SDR produced %0d bytes, expected 4", n_got);
            errors++;
        end
        for (int i = 0; i < n_got; i++) sdr_bytes[i] = got[i];
        sdr_n = n_got;
        sdr_edges = edges_used; sdr_periods = sclk_periods;
        sdr_total = total_periods;
        // The bytes must be the source pattern, not merely self-consistent.
        for (int i = 0; i < 4; i++) begin
            if (got[i] !== 8'(8'h31 + i[7:0])) begin
                $display("  FAIL: SDR byte %0d is 0x%02h, expected 0x%02h",
                         i, got[i], 8'h31 + i[7:0]);
                errors++;
            end
        end
        $display("  4 lanes SDR, 4 bytes: %0d edges, %0d SCLK periods, %0d with dummy",
                 edges_used, sclk_periods, total_periods);

        // 2. DDR, THE SAME source sequence. Identical bytes, half the SCLK
        //    periods -- and the same edge count, because the same bits must
        //    still cross the same lanes.
        run(4, 1'b1, 4, 8);
        if (n_got != sdr_n) begin
            $display("  FAIL: DDR produced %0d bytes, SDR produced %0d",
                     n_got, sdr_n);
            errors++;
        end
        for (int i = 0; i < n_got; i++) begin
            if (got[i] !== sdr_bytes[i]) begin
                $display("  FAIL: DDR byte %0d is 0x%02h, SDR gave 0x%02h",
                         i, got[i], sdr_bytes[i]);
                errors++;
            end
        end
        if (edges_used !== sdr_edges) begin
            $display("  FAIL: DDR used %0d edges, SDR used %0d -- these must match",
                     edges_used, sdr_edges);
            errors++;
        end
        if (sclk_periods !== (sdr_periods >> 1)) begin
            $display("  FAIL: DDR took %0d SCLK periods, expected half of %0d",
                     sclk_periods, sdr_periods);
            errors++;
        end
        $display("  4 lanes DDR, 4 bytes: %0d edges (same), %0d SCLK periods (half), %0d with dummy",
                 edges_used, sclk_periods, total_periods);

        // 3. THE DUMMY PHASE DOES NOT HALVE. The data periods halved; the
        //    total did not, because the dummy count is in SCLK cycles.
        if (total_periods !== (sclk_periods + 16'd8)) begin
            $display("  FAIL: DDR total %0d is not data %0d plus dummy 8",
                     total_periods, sclk_periods);
            errors++;
        end
        $display("  speedup with dummy: %0d -> %0d periods, which is %0d.%02dx -- not 2x",
                 sdr_total, total_periods,
                 sdr_total / total_periods,
                 ((sdr_total * 100) / total_periods) % 100);
        if ((sdr_total * 100) / total_periods >= 200) begin
            $display("  FAIL: DDR achieved 2x or better -- the dummy phase must cost something");
            errors++;
        end

        // 4. A SHORT transfer, where the dummy phase dominates and DDR
        //    buys even less. The same Amdahl argument as Chapter 12.1.
        build_source(4, 1);
        run(4, 1'b0, 1, 8);
        sdr_total = total_periods;
        run(4, 1'b1, 1, 8);
        $display("  1 byte, 4 lanes: %0d -> %0d periods with dummy, %0d.%02dx",
                 sdr_total, total_periods,
                 sdr_total / total_periods,
                 ((sdr_total * 100) / total_periods) % 100);
        if ((sdr_total * 100) / total_periods >= 130) begin
            $display("  FAIL: on one byte DDR gave 1.3x or better");
            errors++;
        end

        // 5. WITHOUT a dummy phase, DDR does give exactly two -- which
        //    isolates the dummy phase as the whole reason it does not.
        build_source(4, 16);
        run(4, 1'b0, 16, 0);
        sdr_total = total_periods;
        run(4, 1'b1, 16, 0);
        if (total_periods !== (sdr_total >> 1)) begin
            $display("  FAIL: with no dummy phase DDR gave %0d, expected half of %0d",
                     total_periods, sdr_total);
            errors++;
        end
        $display("  no dummy phase: %0d -> %0d periods, exactly 2x",
                 sdr_total, total_periods);

        // 6. EQUIVALENCE ACROSS EVERY WIDTH. The byte assembly must not
        //    depend on the data rate at any width.
        for (int wi = 0; wi < 4; wi++) begin
            int nl;
            nl = (wi == 0) ? 1 : (wi == 1) ? 2 : (wi == 2) ? 4 : 8;
            build_source(nl, 8);
            run(nl, 1'b0, 8, 8);
            for (int i = 0; i < n_got; i++) sdr_bytes[i] = got[i];
            sdr_n = n_got; sdr_edges = edges_used; sdr_periods = sclk_periods;
            run(nl, 1'b1, 8, 8);
            if (n_got != sdr_n) begin
                $display("  FAIL: %0d lanes: DDR gave %0d bytes, SDR %0d",
                         nl, n_got, sdr_n);
                errors++;
            end
            for (int i = 0; i < n_got; i++) begin
                if (got[i] !== sdr_bytes[i]) begin
                    $display("  FAIL: %0d lanes byte %0d differs: DDR 0x%02h, SDR 0x%02h",
                             nl, i, got[i], sdr_bytes[i]);
                    errors++;
                end
                if (got[i] !== 8'(8'h31 + i[7:0])) begin
                    $display("  FAIL: %0d lanes byte %0d is 0x%02h, source was 0x%02h",
                             nl, i, got[i], 8'h31 + i[7:0]);
                    errors++;
                end
            end
            if (edges_used !== sdr_edges) begin
                $display("  FAIL: %0d lanes: edge counts differ between modes", nl);
                errors++;
            end
            if (sclk_periods !== (sdr_periods >> 1)) begin
                $display("  FAIL: %0d lanes: DDR periods %0d, expected half of %0d",
                         nl, sclk_periods, sdr_periods);
                errors++;
            end
        end
        $display("  equivalence at 1, 2, 4 and 8 lanes: identical bytes, identical edges, half the periods");

        // 7. AN ODD GROUP COUNT. Groups arrive in pairs in DDR, so an odd
        //    count leaves a period carrying one group -- reported rather
        //    than silently rounded. One byte at eight lanes is one group.
        build_source(8, 1);
        run(8, 1'b1, 1, 8);
        if (!partial_period) begin
            $display("  FAIL: one group in DDR was not reported as a partial period");
            errors++;
        end
        $display("  1 byte at 8 lanes in DDR: %0d edge, partial period reported",
                 edges_used);

        // 8. An even count must NOT be reported, and SDR never is -- a
        //    single-edge mode has no pairing requirement at all.
        build_source(8, 2);
        run(8, 1'b1, 2, 8);
        if (partial_period) begin
            $display("  FAIL: two groups in DDR was reported as partial");
            errors++;
        end
        build_source(8, 1);
        run(8, 1'b0, 1, 8);
        if (partial_period) begin
            $display("  FAIL: SDR reported a partial period -- it has no pairing requirement");
            errors++;
        end
        $display("  pairing: odd reported in DDR only, even never, SDR never");

        if (errors == 0)
            $display("PASS: the same group sequence produces identical bytes in SDR and DDR at every width and uses the same number of edges, DDR halves the SCLK periods the data occupies and nothing else, the dummy phase is unchanged so the end-to-end speedup is below two and far below it on a short transfer, removing the dummy phase restores exactly two, and an odd group count in DDR is reported rather than rounded");
        else
            $display("FAIL: %0d error(s)", errors);
        $finish;
    end

endmodule

The testbench is built around the equivalence claim, and the construction matters as much as the checks.

The source sequence is built once, into an array, before either mode runs. Both modes then see exactly the same groups in the same order. A testbench that generated the stimulus separately for each mode could not distinguish "the two modes agree" from "the two generators agree".

The bytes are checked against the source pattern, not merely against each other. That distinction is the same one Chapter 12.2 drew about round trips: two modes agreeing proves consistency, not correctness, so the absolute check is required alongside.

Then the three quantitative properties. The edge count must be identical between modes — the same bits cross the same lanes either way. The period count must be exactly half. And the total must be the data periods plus the dummy count, which is the arithmetic of §4 stated as a check.

Two bound checks encode the chapter's claims. DDR with a dummy phase must give strictly less than 2×, and on one byte less than 1.3×. Then removing the dummy phase must restore exactly 2× — which is what isolates the dummy phase as the sole cause rather than one of several.

The equivalence is finally swept across all four widths, because the assembly path's mode-independence must hold at every group size and not just at four lanes.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_ddr_gear.v — the same gearbox in Verilog-2001
// spi_ddr_gear.v
//
// Chapter 12.4 -- single and double data rate on the same lanes, in
// Verilog-2001.
//
// Widening the bus moves more bits per EDGE. Double data rate moves data on
// BOTH edges of each SCLK period, so it moves the same bits per edge and
// twice as many edges per period. The two are independent: an octal DDR part
// moves 16 bits per SCLK period.
//
// This block is clocked once per data-carrying EDGE, which makes the
// difference between the modes purely a counting relationship:
//
//   edges         = 8 * n_bytes / lanes        -- the same in both modes
//   sclk_periods  = ddr ? edges / 2 : edges    -- the only thing DDR changes
//
// The byte assembly is IDENTICAL in both modes, so the same group sequence
// must produce the same bytes.
//
// THE DUMMY PHASE DOES NOT HALVE. It is specified in SCLK cycles, so DDR
// leaves it untouched and the speedup is below the nominal factor. In
// practice DDR parts specify MORE dummy cycles than their SDR modes, because
// halving the time per bit tightens the sampling budget.
//
// A DDR transfer also needs an EVEN number of groups, because groups arrive
// in pairs -- one per edge of a period. An odd count is reported rather than
// silently rounded.

module spi_ddr_gear #(
    parameter MAX_LANES = 8,
    parameter CNT_W     = 16
) (
    input  wire                 clk,      // one tick per data-carrying edge
    input  wire                 rst_n,

    input  wire [3:0]           lanes,    // 1, 2, 4 or 8
    input  wire                 ddr,      // 0 one edge per period, 1 both

    input  wire                 start,
    input  wire [CNT_W-1:0]     n_bytes,
    input  wire [5:0]           dummy_cycles,   // SCLK CYCLES -- not halved

    // One group arrives per tick while busy.
    input  wire [MAX_LANES-1:0] group_in,

    output reg  [7:0]           byte_out,
    output reg                  byte_valid,

    output reg  [CNT_W-1:0]     edges_used,
    output reg  [CNT_W-1:0]     sclk_periods,   // what DDR actually reduces
    output reg  [CNT_W-1:0]     total_periods,  // including the dummy phase
    output reg                  partial_period, // an odd group count in DDR
    output reg                  busy,
    output reg                  done
);

    reg [7:0]       asm_sr;      // assembling byte
    reg [3:0]       grp_left;    // groups remaining in this byte
    reg [CNT_W-1:0] edges_left;
    reg [3:0]       grp_per_byte;

    function [3:0] group_count;
        input [3:0] n;
        begin
            case (n)
                4'd1:    group_count = 4'd8;
                4'd2:    group_count = 4'd4;
                4'd4:    group_count = 4'd2;
                4'd8:    group_count = 4'd1;
                default: group_count = 4'd8;
            endcase
        end
    endfunction

    function [2:0] lane_shift;
        input [3:0] n;
        begin
            case (n)
                4'd1:    lane_shift = 3'd0;
                4'd2:    lane_shift = 3'd1;
                4'd4:    lane_shift = 3'd2;
                4'd8:    lane_shift = 3'd3;
                default: lane_shift = 3'd0;
            endcase
        end
    endfunction

    wire [MAX_LANES-1:0] lane_mask = (1 << lanes) - 1;

    // The total edge count, which does NOT depend on the data rate -- the
    // same bits must cross the same lanes either way.
    wire [CNT_W-1:0] n_edges = (n_bytes << 3) >> lane_shift(lanes);

    always @(posedge clk or negedge rst_n) begin
        if (!rst_n) begin
            asm_sr         <= 8'h00;
            grp_left       <= 4'd0;
            edges_left     <= {CNT_W{1'b0}};
            grp_per_byte   <= 4'd8;
            byte_out       <= 8'h00;
            byte_valid     <= 1'b0;
            edges_used     <= {CNT_W{1'b0}};
            sclk_periods   <= {CNT_W{1'b0}};
            total_periods  <= {CNT_W{1'b0}};
            partial_period <= 1'b0;
            busy           <= 1'b0;
            done           <= 1'b0;
        end else begin
            byte_valid <= 1'b0;
            done       <= 1'b0;

            if (start && !busy) begin
                asm_sr       <= 8'h00;
                grp_per_byte <= group_count(lanes);
                grp_left     <= group_count(lanes);
                edges_left   <= n_edges;
                edges_used   <= n_edges;

                // The ONLY thing the data rate changes: how many SCLK
                // periods those edges occupy.
                sclk_periods <= ddr ? (n_edges >> 1) : n_edges;

                // ...and the dummy phase is added unchanged, because it is
                // specified in SCLK cycles rather than in bits.
                total_periods <= (ddr ? (n_edges >> 1) : n_edges)
                                 + dummy_cycles;

                // A DDR transfer needs an even number of groups: they arrive
                // in pairs, one per edge of a period.
                partial_period <= ddr && (n_edges[0] == 1'b1);

                busy <= (n_edges != {CNT_W{1'b0}});
            end else if (busy) begin
                // Byte assembly, identical in both modes -- the property
                // that makes SDR and DDR interchangeable to the consumer.
                asm_sr <= (asm_sr << lanes) | (group_in & lane_mask);

                if (grp_left == 4'd1) begin
                    byte_out   <= (asm_sr << lanes) | (group_in & lane_mask);
                    byte_valid <= 1'b1;
                    grp_left   <= grp_per_byte;
                    asm_sr     <= 8'h00;
                end else begin
                    grp_left <= grp_left - 4'd1;
                end

                if (edges_left == 1) begin
                    busy <= 1'b0;
                    done <= 1'b1;
                end else begin
                    edges_left <= edges_left - 1'b1;
                end
            end
        end
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_ddr_gear_tb.v — the same equivalence in Verilog-2001
// spi_ddr_gear_tb.v
//
// The same checks as the SystemVerilog testbench. The central property is
// EQUIVALENCE: the same group sequence pushed through SDR and DDR must
// produce the same bytes, because the byte assembly does not depend on the
// data rate. Only the SCLK period count differs, by exactly two.
//
// The second property is the chapter's point: the dummy phase does not
// halve, so the end-to-end speedup is below two.

`timescale 1ns/1ps

module spi_ddr_gear_tb;

    parameter MAX_LANES = 8;
    parameter CNT_W     = 16;

    reg clk;
    reg rst_n;

    reg  [3:0]           lanes;
    reg                  ddr;
    reg                  start;
    reg  [CNT_W-1:0]     n_bytes;
    reg  [5:0]           dummy_cycles;
    reg  [MAX_LANES-1:0] group_in;

    wire [7:0]           byte_out;
    wire                 byte_valid;
    wire [CNT_W-1:0]     edges_used, sclk_periods, total_periods;
    wire                 partial_period, busy, done;

    integer errors;
    integer i, wi, nl, b, k, gpb, idx, guard;
    reg [7:0] v;

    initial begin
        clk = 1'b0; rst_n = 1'b0;
        lanes = 4'd4; ddr = 1'b0; start = 1'b0;
        n_bytes = 16'd4; dummy_cycles = 6'd8; group_in = 8'h00;
        errors = 0;
    end
    always #5 clk = ~clk;

    spi_ddr_gear #(.MAX_LANES(MAX_LANES), .CNT_W(CNT_W)) dut (
        .clk(clk), .rst_n(rst_n),
        .lanes(lanes), .ddr(ddr),
        .start(start), .n_bytes(n_bytes), .dummy_cycles(dummy_cycles),
        .group_in(group_in),
        .byte_out(byte_out), .byte_valid(byte_valid),
        .edges_used(edges_used), .sclk_periods(sclk_periods),
        .total_periods(total_periods), .partial_period(partial_period),
        .busy(busy), .done(done)
    );

    // Received bytes, for the equivalence comparison.
    reg [7:0] got [0:63];
    integer   n_got;

    initial n_got = 0;

    always @(posedge clk) begin
        if (rst_n && byte_valid && n_got < 64) begin
            got[n_got] <= byte_out;
            n_got      <= n_got + 1;
        end
    end

    // The source pattern, as groups. Built once so both modes see exactly
    // the same sequence -- which is what makes the comparison meaningful.
    reg [MAX_LANES-1:0] src [0:511];
    integer             n_src;

    task build_source;
        input integer nl_i;
        input integer nb;
        begin
            gpb = 8 / nl_i;
            n_src = 0;
            for (b = 0; b < nb; b = b + 1) begin
                v = 8'h31 + b[7:0];       // a pattern with no symmetry
                for (k = 0; k < gpb; k = k + 1) begin
                    // Most significant group first, matching Chapter 12.2.
                    src[n_src] = (v >> (8 - nl_i * (k + 1))) & ((1 << nl_i) - 1);
                    n_src = n_src + 1;
                end
            end
        end
    endtask

    task run;
        input integer nl_i;
        input         is_ddr;
        input integer nb;
        input integer dum;
        begin
            @(negedge clk);
            lanes = nl_i[3:0]; ddr = is_ddr; n_bytes = nb[CNT_W-1:0];
            dummy_cycles = dum[5:0];
            n_got = 0;
            start = 1'b1;
            @(negedge clk);
            start = 1'b0;
            idx = 0; guard = 0;
            while (busy && guard < 1024) begin
                if (idx < n_src) group_in = src[idx];
                else             group_in = 8'h00;
                idx = idx + 1;
                @(negedge clk);
                guard = guard + 1;
            end
            group_in = 8'h00;
            @(negedge clk);
        end
    endtask

    // Captured SDR results, for comparison against DDR.
    reg [7:0]       sdr_bytes [0:63];
    integer         sdr_n;
    reg [CNT_W-1:0] sdr_edges, sdr_periods, sdr_total;

    initial begin
        n_src = 0;
        repeat (3) @(negedge clk);
        rst_n = 1'b1;
        @(negedge clk);

        // 1. SDR at four lanes, four bytes. Establish the baseline.
        build_source(4, 4);
        run(4, 1'b0, 4, 8);
        if (n_got != 4) begin
            $display("  FAIL: SDR produced %0d bytes, expected 4", n_got);
            errors = errors + 1;
        end
        for (i = 0; i < n_got; i = i + 1) sdr_bytes[i] = got[i];
        sdr_n = n_got;
        sdr_edges = edges_used; sdr_periods = sclk_periods;
        sdr_total = total_periods;
        // The bytes must be the source pattern, not merely self-consistent.
        for (i = 0; i < 4; i = i + 1) begin
            if (got[i] !== (8'h31 + i[7:0])) begin
                $display("  FAIL: SDR byte %0d is 0x%02h, expected 0x%02h",
                         i, got[i], 8'h31 + i[7:0]);
                errors = errors + 1;
            end
        end
        $display("  4 lanes SDR, 4 bytes: %0d edges, %0d SCLK periods, %0d with dummy",
                 edges_used, sclk_periods, total_periods);

        // 2. DDR, THE SAME source sequence.
        run(4, 1'b1, 4, 8);
        if (n_got != sdr_n) begin
            $display("  FAIL: DDR produced %0d bytes, SDR produced %0d",
                     n_got, sdr_n);
            errors = errors + 1;
        end
        for (i = 0; i < n_got; i = i + 1) begin
            if (got[i] !== sdr_bytes[i]) begin
                $display("  FAIL: DDR byte %0d is 0x%02h, SDR gave 0x%02h",
                         i, got[i], sdr_bytes[i]);
                errors = errors + 1;
            end
        end
        if (edges_used !== sdr_edges) begin
            $display("  FAIL: DDR used %0d edges, SDR used %0d -- these must match",
                     edges_used, sdr_edges);
            errors = errors + 1;
        end
        if (sclk_periods !== (sdr_periods >> 1)) begin
            $display("  FAIL: DDR took %0d SCLK periods, expected half of %0d",
                     sclk_periods, sdr_periods);
            errors = errors + 1;
        end
        $display("  4 lanes DDR, 4 bytes: %0d edges (same), %0d SCLK periods (half), %0d with dummy",
                 edges_used, sclk_periods, total_periods);

        // 3. THE DUMMY PHASE DOES NOT HALVE.
        if (total_periods !== (sclk_periods + 16'd8)) begin
            $display("  FAIL: DDR total %0d is not data %0d plus dummy 8",
                     total_periods, sclk_periods);
            errors = errors + 1;
        end
        $display("  speedup with dummy: %0d -> %0d periods, which is %0d.%02dx -- not 2x",
                 sdr_total, total_periods,
                 sdr_total / total_periods,
                 ((sdr_total * 100) / total_periods) % 100);
        if ((sdr_total * 100) / total_periods >= 200) begin
            $display("  FAIL: DDR achieved 2x or better -- the dummy phase must cost something");
            errors = errors + 1;
        end

        // 4. A SHORT transfer, where the dummy phase dominates.
        build_source(4, 1);
        run(4, 1'b0, 1, 8);
        sdr_total = total_periods;
        run(4, 1'b1, 1, 8);
        $display("  1 byte, 4 lanes: %0d -> %0d periods with dummy, %0d.%02dx",
                 sdr_total, total_periods,
                 sdr_total / total_periods,
                 ((sdr_total * 100) / total_periods) % 100);
        if ((sdr_total * 100) / total_periods >= 130) begin
            $display("  FAIL: on one byte DDR gave 1.3x or better");
            errors = errors + 1;
        end

        // 5. WITHOUT a dummy phase, DDR does give exactly two -- which
        //    isolates the dummy phase as the whole reason it does not.
        build_source(4, 16);
        run(4, 1'b0, 16, 0);
        sdr_total = total_periods;
        run(4, 1'b1, 16, 0);
        if (total_periods !== (sdr_total >> 1)) begin
            $display("  FAIL: with no dummy phase DDR gave %0d, expected half of %0d",
                     total_periods, sdr_total);
            errors = errors + 1;
        end
        $display("  no dummy phase: %0d -> %0d periods, exactly 2x",
                 sdr_total, total_periods);

        // 6. EQUIVALENCE ACROSS EVERY WIDTH.
        for (wi = 0; wi < 4; wi = wi + 1) begin
            if (wi == 0)      nl = 1;
            else if (wi == 1) nl = 2;
            else if (wi == 2) nl = 4;
            else              nl = 8;
            build_source(nl, 8);
            run(nl, 1'b0, 8, 8);
            for (i = 0; i < n_got; i = i + 1) sdr_bytes[i] = got[i];
            sdr_n = n_got; sdr_edges = edges_used; sdr_periods = sclk_periods;
            run(nl, 1'b1, 8, 8);
            if (n_got != sdr_n) begin
                $display("  FAIL: %0d lanes: DDR gave %0d bytes, SDR %0d",
                         nl, n_got, sdr_n);
                errors = errors + 1;
            end
            for (i = 0; i < n_got; i = i + 1) begin
                if (got[i] !== sdr_bytes[i]) begin
                    $display("  FAIL: %0d lanes byte %0d differs: DDR 0x%02h, SDR 0x%02h",
                             nl, i, got[i], sdr_bytes[i]);
                    errors = errors + 1;
                end
                if (got[i] !== (8'h31 + i[7:0])) begin
                    $display("  FAIL: %0d lanes byte %0d is 0x%02h, source was 0x%02h",
                             nl, i, got[i], 8'h31 + i[7:0]);
                    errors = errors + 1;
                end
            end
            if (edges_used !== sdr_edges) begin
                $display("  FAIL: %0d lanes: edge counts differ between modes", nl);
                errors = errors + 1;
            end
            if (sclk_periods !== (sdr_periods >> 1)) begin
                $display("  FAIL: %0d lanes: DDR periods %0d, expected half of %0d",
                         nl, sclk_periods, sdr_periods);
                errors = errors + 1;
            end
        end
        $display("  equivalence at 1, 2, 4 and 8 lanes: identical bytes, identical edges, half the periods");

        // 7. AN ODD GROUP COUNT in DDR.
        build_source(8, 1);
        run(8, 1'b1, 1, 8);
        if (!partial_period) begin
            $display("  FAIL: one group in DDR was not reported as a partial period");
            errors = errors + 1;
        end
        $display("  1 byte at 8 lanes in DDR: %0d edge, partial period reported",
                 edges_used);

        // 8. An even count must NOT be reported, and SDR never is.
        build_source(8, 2);
        run(8, 1'b1, 2, 8);
        if (partial_period) begin
            $display("  FAIL: two groups in DDR was reported as partial");
            errors = errors + 1;
        end
        build_source(8, 1);
        run(8, 1'b0, 1, 8);
        if (partial_period) begin
            $display("  FAIL: SDR reported a partial period");
            errors = errors + 1;
        end
        $display("  pairing: odd reported in DDR only, even never, SDR never");

        if (errors == 0)
            $display("PASS: the same group sequence produces identical bytes in SDR and DDR at every width and uses the same number of edges, DDR halves the SCLK periods the data occupies and nothing else, the dummy phase is unchanged so the end-to-end speedup is below two and far below it on a short transfer, removing the dummy phase restores exactly two, and an odd group count in DDR is reported rather than rounded");
        else
            $display("FAIL: %0d error(s)", errors);
        $finish;
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_ddr_gear.vhd — the same gearbox in VHDL
-- spi_ddr_gear.vhd
--
-- Chapter 12.4 -- single and double data rate on the same lanes, in VHDL.
--
-- Widening the bus moves more bits per EDGE. Double data rate moves data on
-- BOTH edges of each SCLK period, so it moves the same bits per edge and
-- twice as many edges per period. The two are independent: an octal DDR part
-- moves 16 bits per SCLK period.
--
-- This block is clocked once per data-carrying EDGE, which makes the
-- difference between the modes purely a counting relationship:
--
--   edges         = 8 * n_bytes / lanes        -- the same in both modes
--   sclk_periods  = ddr ? edges / 2 : edges    -- the only thing DDR changes
--
-- The byte assembly is IDENTICAL in both modes, so the same group sequence
-- must produce the same bytes.
--
-- THE DUMMY PHASE DOES NOT HALVE. It is specified in SCLK cycles, so DDR
-- leaves it untouched and the speedup is below the nominal factor. In
-- practice DDR parts specify MORE dummy cycles than their SDR modes, because
-- halving the time per bit tightens the sampling budget.
--
-- A DDR transfer also needs an EVEN number of groups, because groups arrive
-- in pairs -- one per edge of a period.

library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

entity spi_ddr_gear is
    generic (
        MAX_LANES : positive := 8;
        CNT_W     : positive := 16
    );
    port (
        clk            : in  std_logic;   -- one tick per data-carrying edge
        rst_n          : in  std_logic;

        lanes          : in  unsigned(3 downto 0);   -- 1, 2, 4 or 8
        ddr            : in  std_logic;   -- 0 one edge per period, 1 both

        start          : in  std_logic;
        n_bytes        : in  unsigned(CNT_W - 1 downto 0);
        dummy_cycles   : in  unsigned(5 downto 0);   -- SCLK CYCLES

        group_in       : in  unsigned(MAX_LANES - 1 downto 0);

        byte_out       : out unsigned(7 downto 0);
        byte_valid     : out std_logic;

        edges_used     : out unsigned(CNT_W - 1 downto 0);
        sclk_periods   : out unsigned(CNT_W - 1 downto 0);
        total_periods  : out unsigned(CNT_W - 1 downto 0);
        partial_period : out std_logic;
        busy           : out std_logic;
        done           : out std_logic
    );
end entity;

architecture rtl of spi_ddr_gear is

    function group_count(n : unsigned(3 downto 0)) return natural is
    begin
        case to_integer(n) is
            when 1      => return 8;
            when 2      => return 4;
            when 4      => return 2;
            when 8      => return 1;
            when others => return 8;
        end case;
    end function;

    function lane_div(n : unsigned(3 downto 0)) return natural is
    begin
        case to_integer(n) is
            when 1      => return 1;
            when 2      => return 2;
            when 4      => return 4;
            when 8      => return 8;
            when others => return 1;
        end case;
    end function;

    signal asm_sr     : unsigned(7 downto 0) := (others => '0');
    signal grp_left   : natural := 0;
    signal gpb_r      : natural := 8;
    signal edges_left : natural := 0;

    signal bo_r   : unsigned(7 downto 0) := (others => '0');
    signal bv_r   : std_logic := '0';
    signal eu_r   : unsigned(CNT_W - 1 downto 0) := (others => '0');
    signal sp_r   : unsigned(CNT_W - 1 downto 0) := (others => '0');
    signal tp_r   : unsigned(CNT_W - 1 downto 0) := (others => '0');
    signal pp_r   : std_logic := '0';
    signal busy_r : std_logic := '0';
    signal done_r : std_logic := '0';

    signal lane_mask : unsigned(MAX_LANES - 1 downto 0) := (others => '0');

begin

    lane_mask <= to_unsigned(2 ** to_integer(lanes) - 1, MAX_LANES)
                 when to_integer(lanes) <= MAX_LANES
                 else to_unsigned(0, MAX_LANES);

    gear : process (clk, rst_n)
        variable n_edges : natural;
        variable nxt     : unsigned(7 downto 0);
    begin
        if rst_n = '0' then
            asm_sr     <= (others => '0');
            grp_left   <= 0;
            gpb_r      <= 8;
            edges_left <= 0;
            bo_r       <= (others => '0');
            bv_r       <= '0';
            eu_r       <= (others => '0');
            sp_r       <= (others => '0');
            tp_r       <= (others => '0');
            pp_r       <= '0';
            busy_r     <= '0';
            done_r     <= '0';
        elsif rising_edge(clk) then
            bv_r   <= '0';
            done_r <= '0';

            if start = '1' and busy_r = '0' then
                -- The total edge count does NOT depend on the data rate --
                -- the same bits must cross the same lanes either way.
                n_edges := (to_integer(n_bytes) * 8) / lane_div(lanes);

                asm_sr     <= (others => '0');
                gpb_r      <= group_count(lanes);
                grp_left   <= group_count(lanes);
                edges_left <= n_edges;
                eu_r       <= to_unsigned(n_edges, CNT_W);

                -- The ONLY thing the data rate changes: how many SCLK
                -- periods those edges occupy.
                if ddr = '1' then
                    sp_r <= to_unsigned(n_edges / 2, CNT_W);
                    tp_r <= to_unsigned(n_edges / 2 +
                                        to_integer(dummy_cycles), CNT_W);
                    -- A DDR transfer needs an even number of groups.
                    if (n_edges mod 2) = 1 then pp_r <= '1';
                    else                        pp_r <= '0'; end if;
                else
                    sp_r <= to_unsigned(n_edges, CNT_W);
                    tp_r <= to_unsigned(n_edges +
                                        to_integer(dummy_cycles), CNT_W);
                    pp_r <= '0';
                end if;

                if n_edges /= 0 then busy_r <= '1'; else busy_r <= '0'; end if;
            elsif busy_r = '1' then
                -- Byte assembly, identical in both modes -- the property
                -- that makes SDR and DDR interchangeable to the consumer.
                nxt := shift_left(asm_sr, to_integer(lanes)) or
                       resize(group_in and lane_mask, 8);

                if grp_left = 1 then
                    bo_r     <= nxt;
                    bv_r     <= '1';
                    grp_left <= gpb_r;
                    asm_sr   <= (others => '0');
                else
                    asm_sr   <= nxt;
                    grp_left <= grp_left - 1;
                end if;

                if edges_left = 1 then
                    busy_r <= '0';
                    done_r <= '1';
                else
                    edges_left <= edges_left - 1;
                end if;
            end if;
        end if;
    end process;

    byte_out       <= bo_r;
    byte_valid     <= bv_r;
    edges_used     <= eu_r;
    sclk_periods   <= sp_r;
    total_periods  <= tp_r;
    partial_period <= pp_r;
    busy           <= busy_r;
    done           <= done_r;

end architecture;
Azvya Education Pvt. Ltd.VLSI Mentor
spi_ddr_gear_tb.vhd — the same equivalence in VHDL
-- spi_ddr_gear_tb.vhd
--
-- The same checks as the SystemVerilog and Verilog testbenches. The central
-- property is EQUIVALENCE: the same group sequence pushed through SDR and
-- DDR must produce the same bytes, because the byte assembly does not depend
-- on the data rate. Only the SCLK period count differs, by exactly two.
--
-- The second property is the chapter's point: the dummy phase does not
-- halve, so the end-to-end speedup is below two.

library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

entity spi_ddr_gear_tb is
end entity;

architecture sim of spi_ddr_gear_tb is

    constant MAX_LANES : positive := 8;
    constant CNT_W     : positive := 16;

    signal clk   : std_logic := '0';
    signal rst_n : std_logic := '0';
    signal halt  : boolean   := false;

    signal lanes        : unsigned(3 downto 0) := to_unsigned(4, 4);
    signal ddr          : std_logic := '0';
    signal start        : std_logic := '0';
    signal n_bytes      : unsigned(CNT_W - 1 downto 0) := to_unsigned(4, CNT_W);
    signal dummy_cycles : unsigned(5 downto 0) := to_unsigned(8, 6);
    signal group_in     : unsigned(MAX_LANES - 1 downto 0) := (others => '0');

    signal byte_out       : unsigned(7 downto 0);
    signal byte_valid     : std_logic;
    signal edges_used     : unsigned(CNT_W - 1 downto 0);
    signal sclk_periods   : unsigned(CNT_W - 1 downto 0);
    signal total_periods  : unsigned(CNT_W - 1 downto 0);
    signal partial_period : std_logic;
    signal busy, done     : std_logic;

    type byte_arr is array (0 to 63) of unsigned(7 downto 0);
    signal got   : byte_arr;
    signal n_got : natural := 0;
    signal clr   : std_logic := '0';

    signal errors : natural := 0;

begin

    clk <= not clk after 5 ns when not halt else '0';

    dut : entity work.spi_ddr_gear
        generic map (MAX_LANES => MAX_LANES, CNT_W => CNT_W)
        port map (
            clk => clk, rst_n => rst_n,
            lanes => lanes, ddr => ddr,
            start => start, n_bytes => n_bytes, dummy_cycles => dummy_cycles,
            group_in => group_in,
            byte_out => byte_out, byte_valid => byte_valid,
            edges_used => edges_used, sclk_periods => sclk_periods,
            total_periods => total_periods, partial_period => partial_period,
            busy => busy, done => done
        );

    -- Collector. The clear arrives on its own signal because two processes
    -- driving one signal is a multiple-driver error in VHDL.
    collect : process (clk)
    begin
        if rising_edge(clk) then
            if clr = '1' then
                n_got <= 0;
            elsif rst_n = '1' and byte_valid = '1' and n_got < 64 then
                got(n_got) <= byte_out;
                n_got      <= n_got + 1;
            end if;
        end if;
    end process;

    stim : process
        variable errs : natural := 0;

        -- The source pattern, as groups. Built once so both modes see
        -- exactly the same sequence.
        type grp_arr is array (0 to 511) of unsigned(MAX_LANES - 1 downto 0);
        variable src   : grp_arr;
        variable n_src : natural := 0;

        variable sdr_bytes : byte_arr;
        variable sdr_n     : natural;
        variable sdr_edges, sdr_periods, sdr_total : natural;
        variable guard, idx, gpb, nl : natural;
        variable v : unsigned(7 downto 0);

        procedure build_source(nl_i : natural; nb : natural) is
            variable g : natural;
            variable w : unsigned(7 downto 0);
        begin
            g := 8 / nl_i;
            n_src := 0;
            for b in 0 to nb - 1 loop
                w := to_unsigned((16#31# + b) mod 256, 8);
                for k in 0 to g - 1 loop
                    -- Most significant group first, matching Chapter 12.2.
                    src(n_src) := resize(
                        shift_right(w, 8 - nl_i * (k + 1)) and
                        to_unsigned(2 ** nl_i - 1, 8), MAX_LANES);
                    n_src := n_src + 1;
                end loop;
            end loop;
        end procedure;

        procedure run(nl_i : natural; is_ddr : std_logic;
                      nb : natural; dum : natural) is
        begin
            wait until falling_edge(clk);
            clr <= '1';
            wait until falling_edge(clk);
            clr <= '0';
            lanes        <= to_unsigned(nl_i, 4);
            ddr          <= is_ddr;
            n_bytes      <= to_unsigned(nb, CNT_W);
            dummy_cycles <= to_unsigned(dum, 6);
            start        <= '1';
            wait until falling_edge(clk);
            start        <= '0';
            idx := 0; guard := 0;
            while busy = '1' and guard < 1024 loop
                if idx < n_src then group_in <= src(idx);
                else                group_in <= (others => '0'); end if;
                idx := idx + 1;
                wait until falling_edge(clk);
                guard := guard + 1;
            end loop;
            group_in <= (others => '0');
            wait until falling_edge(clk);
        end procedure;
    begin
        for k in 0 to 2 loop
            wait until falling_edge(clk);
        end loop;
        rst_n <= '1';
        wait until falling_edge(clk);

        -- 1. SDR at four lanes, four bytes. Establish the baseline.
        build_source(4, 4);
        run(4, '0', 4, 8);
        if n_got /= 4 then
            report "  FAIL: SDR did not produce four bytes"; errs := errs + 1;
        end if;
        for i in 0 to n_got - 1 loop
            sdr_bytes(i) := got(i);
        end loop;
        sdr_n := n_got;
        sdr_edges := to_integer(edges_used);
        sdr_periods := to_integer(sclk_periods);
        sdr_total := to_integer(total_periods);
        -- The bytes must be the source pattern, not merely self-consistent.
        for i in 0 to 3 loop
            if to_integer(got(i)) /= 16#31# + i then
                report "  FAIL: an SDR byte does not match the source pattern";
                errs := errs + 1;
            end if;
        end loop;
        report "  4 lanes SDR, 4 bytes: " &
               integer'image(sdr_edges) & " edges, " &
               integer'image(sdr_periods) & " SCLK periods, " &
               integer'image(sdr_total) & " with dummy";

        -- 2. DDR, THE SAME source sequence.
        run(4, '1', 4, 8);
        if n_got /= sdr_n then
            report "  FAIL: DDR produced a different number of bytes than SDR";
            errs := errs + 1;
        end if;
        for i in 0 to n_got - 1 loop
            if got(i) /= sdr_bytes(i) then
                report "  FAIL: a DDR byte differs from the SDR byte";
                errs := errs + 1;
            end if;
        end loop;
        if to_integer(edges_used) /= sdr_edges then
            report "  FAIL: the edge counts differ between the modes";
            errs := errs + 1;
        end if;
        if to_integer(sclk_periods) /= sdr_periods / 2 then
            report "  FAIL: DDR did not halve the SCLK periods";
            errs := errs + 1;
        end if;
        report "  4 lanes DDR, 4 bytes: " &
               integer'image(to_integer(edges_used)) & " edges (same), " &
               integer'image(to_integer(sclk_periods)) &
               " SCLK periods (half), " &
               integer'image(to_integer(total_periods)) & " with dummy";

        -- 3. THE DUMMY PHASE DOES NOT HALVE.
        if to_integer(total_periods) /= to_integer(sclk_periods) + 8 then
            report "  FAIL: the total is not the data periods plus the dummy count";
            errs := errs + 1;
        end if;
        report "  speedup with dummy: " & integer'image(sdr_total) & " -> " &
               integer'image(to_integer(total_periods)) & " periods, x100 = " &
               integer'image((sdr_total * 100) / to_integer(total_periods)) &
               " -- not 200";
        if (sdr_total * 100) / to_integer(total_periods) >= 200 then
            report "  FAIL: DDR achieved 2x or better"; errs := errs + 1;
        end if;

        -- 4. A SHORT transfer, where the dummy phase dominates.
        build_source(4, 1);
        run(4, '0', 1, 8);
        sdr_total := to_integer(total_periods);
        run(4, '1', 1, 8);
        report "  1 byte, 4 lanes: " & integer'image(sdr_total) & " -> " &
               integer'image(to_integer(total_periods)) &
               " periods with dummy, x100 = " &
               integer'image((sdr_total * 100) / to_integer(total_periods));
        if (sdr_total * 100) / to_integer(total_periods) >= 130 then
            report "  FAIL: on one byte DDR gave 1.3x or better";
            errs := errs + 1;
        end if;

        -- 5. WITHOUT a dummy phase, DDR does give exactly two.
        build_source(4, 16);
        run(4, '0', 16, 0);
        sdr_total := to_integer(total_periods);
        run(4, '1', 16, 0);
        if to_integer(total_periods) /= sdr_total / 2 then
            report "  FAIL: with no dummy phase DDR did not give exactly 2x";
            errs := errs + 1;
        end if;
        report "  no dummy phase: " & integer'image(sdr_total) & " -> " &
               integer'image(to_integer(total_periods)) &
               " periods, exactly 2x";

        -- 6. EQUIVALENCE ACROSS EVERY WIDTH.
        for wi in 0 to 3 loop
            case wi is
                when 0      => nl := 1;
                when 1      => nl := 2;
                when 2      => nl := 4;
                when others => nl := 8;
            end case;
            build_source(nl, 8);
            run(nl, '0', 8, 8);
            for i in 0 to n_got - 1 loop
                sdr_bytes(i) := got(i);
            end loop;
            sdr_n := n_got;
            sdr_edges := to_integer(edges_used);
            sdr_periods := to_integer(sclk_periods);
            run(nl, '1', 8, 8);
            if n_got /= sdr_n then
                report "  FAIL: byte counts differ between the modes";
                errs := errs + 1;
            end if;
            for i in 0 to n_got - 1 loop
                if got(i) /= sdr_bytes(i) then
                    report "  FAIL: a byte differs between the modes";
                    errs := errs + 1;
                end if;
                if to_integer(got(i)) /= 16#31# + i then
                    report "  FAIL: a byte does not match the source pattern";
                    errs := errs + 1;
                end if;
            end loop;
            if to_integer(edges_used) /= sdr_edges then
                report "  FAIL: the edge counts differ between the modes";
                errs := errs + 1;
            end if;
            if to_integer(sclk_periods) /= sdr_periods / 2 then
                report "  FAIL: DDR did not halve the periods at this width";
                errs := errs + 1;
            end if;
        end loop;
        report "  equivalence at 1, 2, 4 and 8 lanes: identical bytes, identical edges, half the periods";

        -- 7. AN ODD GROUP COUNT in DDR.
        build_source(8, 1);
        run(8, '1', 1, 8);
        if partial_period /= '1' then
            report "  FAIL: one group in DDR was not reported as partial";
            errs := errs + 1;
        end if;
        report "  1 byte at 8 lanes in DDR: 1 edge, partial period reported";

        -- 8. An even count must NOT be reported, and SDR never is.
        build_source(8, 2);
        run(8, '1', 2, 8);
        if partial_period = '1' then
            report "  FAIL: two groups in DDR was reported as partial";
            errs := errs + 1;
        end if;
        build_source(8, 1);
        run(8, '0', 1, 8);
        if partial_period = '1' then
            report "  FAIL: SDR reported a partial period"; errs := errs + 1;
        end if;
        report "  pairing: odd reported in DDR only, even never, SDR never";

        errors <= errs;
        if errs = 0 then
            report "PASS: the same group sequence produces identical bytes in SDR and DDR at every width and uses the same number of edges, DDR halves the SCLK periods the data occupies and nothing else, the dummy phase is unchanged so the end-to-end speedup is below two and far below it on a short transfer, removing the dummy phase restores exactly two, and an odd group count in DDR is reported rather than rounded";
        else
            report "FAIL: " & integer'image(errs) & " error(s)" severity error;
        end if;
        halt <= true;
        wait;
    end process;

end architecture;

Parity

All three implement the same gearbox: identical ports and generics, an assembly path with no dependence on the mode, an edge count independent of the rate, a period count halved only in DDR, a dummy count added unchanged, and an odd group count reported in DDR only. All three testbenches build the same source sequence, replay it through both modes at all four widths, and report identical figures — 1.33× on four bytes, 1.11× on one, and exactly 2× with no dummy phase.

8. Why a Verification Engineer Cares

Azvya Education Pvt. Ltd.VLSI Mentor
spi_ddr_gear.sva — the equivalence, and the bound
   // 1. MODE EQUIVALENCE. The assembled byte does not depend on the rate.
   //    This is the property the design's structure should make
   //    unfalsifiable -- the assembly path has no mode input at all -- and
   //    stating it anyway documents that the structure is doing the work.
   a_rate_irrelevant : assert property (
       @(posedge clk) disable iff (!rst_n)
           (byte_valid) |-> (byte_out == expected_from_groups))
       else $error("the assembled byte depended on the data rate");

   // 2. THE EDGE COUNT IS RATE-INDEPENDENT. The same bits cross the same
   //    lanes either way, so a design reporting fewer edges in DDR has
   //    confused edges with periods -- which is the central confusion this
   //    chapter exists to prevent.
   a_edges_invariant : assert property (
       @(posedge clk) disable iff (!rst_n)
           (start) |=> (edges_used == ((n_bytes * 8) / lanes)))
       else $error("the edge count changed with the data rate");

   // 3. AND THE PERIOD COUNT IS EXACTLY HALVED -- no rounding, which is why
   //    the odd case needs its own report rather than a division.
   a_periods_halved : assert property (
       @(posedge clk) disable iff (!rst_n)
           (start && ddr) |=> (sclk_periods == (edges_used / 2)))
       else $error("DDR did not halve the SCLK period count");

   // 4. THE DUMMY PHASE IS UNCHANGED. Total minus data equals the dummy
   //    count, in both modes -- the arithmetic that bounds the speedup.
   a_dummy_unchanged : assert property (
       @(posedge clk) disable iff (!rst_n)
           (start) |=> ((total_periods - sclk_periods) == dummy_cycles))
       else $error("the dummy phase was scaled by the data rate");

   // 5. THE ODD CASE IS REPORTED. A transfer rounded up by one group has
   //    read a byte nobody asked for; rounded down, it has lost one.
   a_odd_reported : assert property (
       @(posedge clk) disable iff (!rst_n)
           (start && ddr && edges_used[0]) |=> partial_period)
       else $error("an odd group count in DDR was not reported");

Property 1 deserves the same comment Chapter 12.2's enable property got, because it is the same situation. The assembly path has no mode input, so the property cannot fail by construction — and that is not a reason to omit it. It records that the structure is carrying the guarantee, and it will fail loudly if someone later adds a mode-dependent branch "for the odd case".

Property 2 is the one that catches the real confusion. Edges and periods are different quantities, and DDR changes only the relationship between them. A design that reports halved edges has conflated them, and every downstream latency calculation will then be wrong by two in a direction that looks like an improvement.

Coverage must cross the rate with the length, because the speedup is a function of both:

Azvya Education Pvt. Ltd.VLSI Mentor
spi_ddr_cg.sv — the gain lives in the cross
   covergroup spi_ddr_cg @(posedge clk iff start);
       cp_ddr   : coverpoint ddr   { bins sdr = {0}; bins ddr = {1}; }

       cp_lanes : coverpoint lanes {
           bins one = {1}; bins two = {2}; bins four = {4}; bins eight = {8};
       }

       // Length decides how much of the transfer the dummy phase is, and
       // therefore what DDR can win. The short bins are where it wins least
       // and are exactly what a throughput-focused suite omits.
       cp_len : coverpoint n_bytes {
           bins one     = {1};            // 1.11x -- the worst case
           bins tiny    = {[2:8]};
           bins medium  = {[9:64]};
           bins large   = {[65:$]};       // approaches 2x
       }

       // ZERO dummy is the case that isolates the dummy phase as the cause,
       // and it is the only one in which DDR reaches exactly 2x.
       cp_dummy : coverpoint dummy_cycles {
           bins none    = {0};            // exactly 2x
           bins sdr_typ = {8};
           bins ddr_typ = {[9:16]};       // what DDR parts actually specify
           bins many    = {[17:32]};
       }

       // The pairing constraint: an odd group count exists only in DDR.
       cp_parity : coverpoint edges_used[0] {
           bins even = {0};
           bins odd  = {1};
       }

       x_ddr_len    : cross cp_ddr, cp_len;
       x_ddr_dummy  : cross cp_ddr, cp_dummy;
       x_ddr_parity : cross cp_ddr, cp_parity;
   endgroup

cp_dummy's ddr_typ bin is the one that reflects reality and is almost always absent. A suite comparing SDR-with-8 against DDR-with-8 is comparing something no datasheet offers — real DDR modes specify more dummy cycles, and §5 showed that can erase the benefit entirely on short transfers. Covering the dummy count without covering the realistic DDR value measures a configuration that does not exist.

9. Why an FPGA or ASIC Engineer Cares

Do not confuse edges with periods. Edges are how much data crossed; periods are how long it took. DDR changes only the ratio, and a design that halves the edge count will get every latency calculation wrong in the flattering direction.

Compare against the device's ACTUAL DDR dummy count. SDR-with-8 against DDR-with-8 is a comparison of one real configuration against one imaginary one. Take both numbers from the command table.

Enable DDR for long sequential transfers and question it for short ones. 1.94× at 256 bytes and 1.00× at four, on representative numbers. The access pattern decides, not the feature list.

Handle the odd group count explicitly. Rounding up reads a byte nobody asked for — on a flash, possibly across a boundary; rounding down loses one. The device's rule is in its datasheet and belongs in the profile.

Budget the timing before the throughput. Halving the time per bit halves every margin in Chapter 10.3's table simultaneously. The throughput arithmetic is the easy part; the reason DDR arrived late is the other half.

Capture on both edges with explicit registers, not a doubled clock, if you can. Rising- and falling-edge registers feeding a common synchroniser are easier to constrain than a 2× clock domain, and they keep the double-rate behaviour confined to the pad logic rather than spreading it through the controller.

10. Failure Signature — DDR That Works Cold and Fails Warm on Long Transfers Only

Symptom. A DDR quad read works during bring-up. In the final enclosure it returns corrupted bytes — but only on transfers longer than a few hundred bytes, and only once the unit has been running for some minutes. Short transfers are always correct. Reverting to SDR fixes it completely.

What "short transfers always correct" establishes. This is the observation that steers the whole case, and its implication is not obvious. A constant timing violation would corrupt short and long transfers alike, since every bit has the same budget. Corruption that depends on transfer length means something is degrading during a transfer — so the fault is cumulative rather than static.

Plausible mechanisms.

  • Supply droop during a long burst. Four lanes switching at double rate draw current in bursts; decoupling that is adequate for a short transfer sags over a long one, and the device's output timing degrades with its supply. This fits length-dependence and temperature-dependence together.
  • Self-heating of the flash during sustained access, so t_V grows as the transfer proceeds — which fits both length and the warm-enclosure condition.
  • Insufficient dummy cycles for DDR, per §5. But that would be a constant error and would corrupt short transfers too.
  • A duty-cycle problem on SCLK. DDR samples on both edges, so an asymmetric clock gives one edge less setup than the other — and the corruption would fall on alternate groups. This does not obviously depend on length.
  • A PLL or clock source drifting with temperature, which fits warm but not length.

The discriminating observation. Length-dependence and temperature-dependence together point at a cumulative effect, and the two candidates that fit both are supply droop and self-heating. They are distinguishable:

Insert idle gaps mid-transfer. Split a 1 KB read into eight 128-byte reads with a short pause between them. If the corruption disappears while the total data is unchanged, something is recovering during the gaps — which is droop or heating, not a static timing error. That test costs nothing and eliminates half the list.

Then separate droop from heating. Droop responds to decoupling and to lowering the clock; heating responds to time-at-temperature and shows hysteresis — it stays broken briefly after the transfer stops. If the first read after a long idle is always correct and corruption builds within a single transfer, droop is indicated; if corruption persists into the next transfer after a pause, heating is.

Check the duty cycle regardless, because DDR makes it matter in a way SDR never did: an asymmetric SCLK gives the two edges unequal budgets, and a 45/55 clock that was invisible at single rate removes 10% of the margin from every falling-edge group. That is worth measuring once even if it is not the cause here.

The fix depends on the mechanism — decoupling, a lower clock, more dummy cycles, or splitting long transfers — but the diagnostic order is what matters: establish whether the fault is static or cumulative before measuring anything, because the two lead to entirely different instruments.

Why DDR surfaces this and SDR did not. Not because DDR is fragile in itself, but because it halves every margin at once (§1). A board with 10% of margin to spare at single rate has none at double, so effects that were always present — droop, heating, duty-cycle asymmetry — become visible together. DDR did not cause the problem; it removed the slack that was hiding it.

11. Common Misconceptions

12. Reason It Through

Work this before reading the answer.

A controller can be configured three ways and only one change is affordable:

(a) 1-1-4 SDR, 8 dummy → 1-1-4 DDR, 12 dummy (b) 1-1-4 SDR, 8 dummy → 4-4-4 SDR, 8 dummy (c) 1-1-4 SDR, 8 dummy → 1-1-8 SDR (octal data), 10 dummy

The workload is 32-byte cache-line fetches with three address bytes. Which wins, and what changes if the workload becomes 4 KB block reads?

Baseline: 1-1-4 SDR, 8 dummy, 32 bytes.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   cmd 8 + addr 24 + dummy 8 + data (256/4 = 64) = 104 periods

Option (a): DDR on four lanes, 12 dummy. Data edges are unchanged at 64; DDR halves the periods they occupy to 32. The command and address are on one lane and also halve in periods — DDR applies to every phase, not only data:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   cmd 8/2 = 4 + addr 24/2 = 12 + dummy 12 + data 32 = 60 periods   → 1.73×

Option (b): 4-4-4 SDR, 8 dummy. Command and address widen to four lanes:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   cmd 2 + addr 6 + dummy 8 + data 64 = 80 periods                  → 1.30×

Option (c): octal data only, 10 dummy. Data halves again; command and address stay at one lane:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   cmd 8 + addr 24 + dummy 10 + data (256/8 = 32) = 74 periods       → 1.41×

Option (a) wins at 1.73×, and the reason is worth extracting: DDR is the only one of the three that shortens the command and address phases without needing the device to support a wider command mode. It halves every phase measured in cycles-of-data, including the one-lane ones — which on a 32-byte fetch is 36 periods of overhead reduced to 16.

Now 4 KB blocks. Data dominates and the overhead becomes noise:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   baseline:  8 + 24 + 8 + 8192 = 8232
   (a) DDR:   4 + 12 + 12 + 4096 = 4124          → 2.00×
   (b) 4-4-4: 2 + 6 + 8 + 8192  = 8208           → 1.00×
   (c) octal: 8 + 24 + 10 + 4096 = 4138          → 1.99×

(a) and (c) both reach essentially 2× and (b) achieves nothing. At 4 KB, widening the command and address is worth 24 periods out of 8232 — completely irrelevant — while halving the data phase is worth everything.

The comparison worth carrying:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
              32 bytes    4 KB
   (a) DDR      1.73×     2.00×
   (b) 4-4-4    1.30×     1.00×
   (c) octal    1.41×     1.99×

Two conclusions. Option (a) wins at both lengths, which makes it the right single choice — but it wins for different reasons: at 32 bytes because it shrinks overhead, at 4 KB because it shrinks data. And option (b) is worth 1.30× on short fetches and nothing on long ones, which inverts the usual assumption that a "more complete" mode is a better one.

The general lesson. Each technique attacks specific terms: width shrinks the phases measured in bits, DDR shrinks every phase measured in periods, and overhead removal — Chapter 12.5's subject — removes phases entirely. Which wins depends on where the periods currently are, so the arithmetic must come before the decision, and the answer changes with the workload rather than with the device.

13. Understanding Check

14. Summary

Width and data rate are independent axes that multiply: bits per period is lanes times two in DDR, so an octal DDR part moves two bytes per SCLK period.

The byte assembly is identical in both modes. Same group size, same lane mapping, same order — so the same group sequence must produce the same bytes, and DDR changes exactly one thing:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   edges        = 8 × n_bytes / lanes      ← identical
   sclk_periods = ddr ? edges / 2 : edges  ← the only difference

Confusing edges with periods is the error to avoid, because it makes every latency calculation wrong by two in the flattering direction.

The dummy phase is counted in periods, so DDR does not halve it — 1.33× on four bytes, 1.11× on one. Removing the dummy phase restores exactly 2×, which isolates it as the sole cause.

And DDR modes usually specify more dummy cycles, because halving the time per bit halves the sampling budget. On representative numbers that gives 1.00× on four bytes and 1.94× at 256 — making DDR a large-transfer feature whose value is decided by the access pattern rather than by the datasheet.

Groups arrive in pairs, so an odd count leaves a period with one unused edge — reported rather than rounded, because rounding up reads an unasked-for byte and rounding down loses one.

DDR's real cost is not pins but half of every timing margin at once, which is why it removes the slack that was hiding effects always present on a board — and why a DDR bring-up failure is usually a board measurement rather than a protocol one.

15. What Comes Next

Width shrinks the phases measured in bits. DDR shrinks the phases measured in periods. Neither can remove a phase entirely — and that is the one thing left.

Chapter 12.5 — XIP and Memory-Mapped Controllers closes the module by removing the command and address phases outright. A sequential access continues a read stream the device already has open, so it pays no command, no address and no dummy phase — just data. XIP removes the command phase from even a random access. The result inverts everything this module has found: where fixed overhead limited what width and rate could win on short transfers, removing that overhead helps short transfers most — which is exactly the instruction-fetch pattern a processor generates, and the reason a serial flash can be code memory at all.

Continue learning