Skip to content
VLSI Mentor

SPI · Module 13

TX and RX Shift Registers

The shift datapath, and why there are two registers rather than the one SPI's symmetry seems to permit: the three failures a circulating register cannot express, why transmit needs alignment and receive does not, and why MOSI must be a flop.

Chapter 13.5 produced three strobes — preload, launch, capture — and nothing that holds a bit. This chapter builds the part that does, and it opens with a genuinely good idea that this design does not use.

SPI is exactly full-duplex and exactly symmetric. Why not one shift register, MOSI from the top and MISO into the bottom?

It works. After N shifts the transmit word has walked out and the receive word has walked in, and one register has done both jobs. There are three reasons not to, and all three are about what happens when a transfer does not go as planned.

1. The Circulating Register

The idea is worth stating properly, because it is elegant and it appears in real designs:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   before:   [ t7 t6 t5 t4 t3 t2 t1 t0 ]
                ^                    ^
                MOSI taken here      MISO shifted in here

   after 8 shifts:
             [ r7 r6 r5 r4 r3 r2 r1 r0 ]

Every launched bit is paired with a captured one, so the register is never half empty. It is half the flops of the two-register version and it has a pleasing symmetry that mirrors the protocol's own.

2. Three Reasons It Is Not Enough

The transmit word is destroyed by the transfer. After the shifts, the word that was sent is gone — overwritten by the word that arrived. A retry has nothing left to re-send. That matters immediately in practice: a flash driver that gets a busy status back must repeat its command, and the pattern "issue command, read status, repeat if busy" is Chapter 11.4's entire subject. A circulating register means every retry needs the software to hold its own copy and rewrite it, which is fine until an interrupt handler retries and the copy is in another context.

The receive word is only correct at exactly N shifts. Stop early — an abort, Chapter 13.10 — and the register holds a mixture of both words with nothing marking the boundary. There is no way to recover the partial receive data, and no way to know how much of what is there is receive data at all. With two registers, a partial transfer leaves a partial receive word in a known position and an intact transmit word.

Transmit and receive widths are forced equal and equal to the register width. Chapter 13.8's configurable frame width has nowhere to go: a 13-bit frame in a 32-bit circulating register leaves the receive bits 13 places from where a reader expects them, and the transmit bits have to be pre-positioned to compensate, which couples the two directions together.

Two registers cost one more register's worth of flops. On any modern process that is not a trade, it is a rounding error — and it removes all three.

3. The Alignment Asymmetry

The two directions do not need the same treatment, which surprises people.

Shifting MSB-first out of the top of the register means the first bit sent must be sitting in the top bit at load time. A len-bit word arrives from software right-aligned, so it has to be moved:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   len = 4, MAX_W = 8, tx_data = 0000_1101

   loaded as:              1101_0000        shifted left by MAX_W - len
                           ^
                           first bit out

The receive side needs nothing. Bits arrive at the bottom and walk upward, so after exactly len captures the word is already right-aligned in the low len bits with zeros above:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   after 4 captures:       0000_r3r2r1r0
                                ^
                                first bit in ends up here

One side shifts at load, the other does not. A design that treats them symmetrically gets one of them wrong — and it is worth noticing which way the mistake usually goes: shifting the receive word as well produces a value that is correct for len == MAX_W and wrong for every narrower frame, so it passes the byte-wide smoke test and fails the first 12-bit ADC.

4. MOSI Is A Register

mosi is the output of a flop, not a mux off the top of the shift register. The distinction matters because of who samples it.

A slave samples MOSI on a clock edge that this master generates. What matters is that the pin is stable well before that edge and changes at exactly one known moment. A flop guarantees that. Combinational logic off a shifting register does not: during the cycle the register shifts, the top bit passes through whatever intermediate values the shift produces, and any of them can appear on the pin.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   mosi = tx_sr[MAX_W-1]        the pin follows the register's own settling
   mosi <= tx_sr[MAX_W-1]       the pin changes once, at a clock edge

The cost is one flop and one cycle of latency, and the latency is free because the launch strobe already arrives a cycle before the capture that consumes the bit.

The corresponding property is checkable and the testbench in §6 checks it continuously: MOSI may change only on the cycle after a drive strobe. Any other movement is a glitch a slave could sample. The monitor also asserts that MOSI moved at all, because a monitor that can only fail if something moves proves nothing when nothing does.

5. The Block

A shift datapath with two registers. A transmit word from software passes through a barrel shifter that left-aligns it by the datapath width minus the frame length, into a transmit shift register. The register's top bit feeds a MOSI output flop, driven on preload and launch strobes, which also shift the register. A MISO input feeds the bottom of a receive shift register, shifted on capture strobes. A capture counter compares against the frame length and produces a valid pulse one cycle later, when the received word is complete.tx_dataright-aligned, fromsoftwareLeft-alignshift by MAX_W - len, onceper frameTX registershifts left, top bit outpreload / launchfrom 13.5 — the sameoperationMOSI flopchanges once, at a clockedgecapturefrom 13.5RX registershifts left, MISO into thebottomCapture countthe frame's positionrx_validdelayed one cycle,deliberatelyrx_dataright-aligned,zero-extendedMOSI pinstable long before the edgeloadtop bitdrive and advanceshift incountlen reached12
Figure 1 — the shift datapath. Two registers, one per direction, and the asymmetry between them: the transmit side is left-aligned by a barrel shifter at load time, and the receive side needs no alignment at all because bits arriving at the bottom end up right-aligned. MOSI is a flop. The valid pulse is delayed by one cycle so that when it is high the received word is complete.

6. Building the Shift Datapath — Three HDLs

The circuit

Two registers, one counter, two flops for MOSI and the valid pulse. The interesting decisions:

Preload and launch do the same thing. Both drive the top bit and advance. They are separate strobes only because Chapter 13.5 knows they happen at different moments, and the datapath does not need to care. Writing them as one branch is not a shortcut — it is the correct statement of what the datapath does, and separating them is what produces the bit-zero-twice bug that Chapter 13.5 describes.

rx_valid_stb is delayed by one cycle, deliberately. The final bit is shifted into rx_sr by the same clock edge that sees the last capture_stb, so on that cycle rx_data still holds the word one bit short. Pulsing valid there would mean "the data is good next cycle" — a contract every consumer gets wrong once. Registering it costs one flop and makes valid mean what it says: while rx_valid_stb is high, rx_data is the complete received word.

The left-align is a barrel shifter, and it is off the critical path. A shift by a variable amount is real logic, but it happens once per frame on the load cycle, and len comes from the configuration latch of Chapter 13.2 so it is stable. Nothing on the per-bit path does arithmetic.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_shift_datapath.sv — two registers, one alignment, and MOSI in a flop
// spi_shift_datapath.sv
//
// Chapter 13.6 -- the shift registers, and why there are two of them.
//
// This block sits underneath the mode logic of 13.5 and sees only three
// strobes: preload, launch, capture. It never learns the mode, the divisor,
// or the polarity. Everything it does is one of:
//
//   load     take a parallel word from the CPU side
//   drive    put the next bit on MOSI          (preload / launch)
//   sample   take the next bit from MISO       (capture)
//   present  hand the assembled word back      (rx_valid_stb)
//
// ONE REGISTER OR TWO?
//
// SPI is exactly full-duplex and exactly symmetric: every launched bit is
// paired with a captured one. That makes a famous trick available -- a SINGLE
// circulating register, MOSI taken from the top bit, MISO shifted into the
// bottom. After N shifts the transmit word has walked out and the receive
// word has walked in, and one register has done both jobs.
//
// It is genuinely elegant and this design does not use it, for three reasons
// that matter in a real master:
//
//   1. The transmit word is DESTROYED by the transfer. A retry after an error
//      has nothing left to re-send, and a flash driver that must repeat a
//      command after a busy status hits this immediately.
//   2. The receive word is only correct at exactly N shifts. Stop early --
//      an abort, Chapter 13.10 -- and what is in the register is a mixture of
//      both words with no boundary marking where.
//   3. It forces the transmit and receive widths to be equal and to equal the
//      register width. Chapter 13.8's configurable width has nowhere to go.
//
// Two registers cost one more register's worth of flops and remove all three.
// On any modern process that is not a trade, it is a rounding error.
//
// THE ALIGNMENT ASYMMETRY.
//
// The two directions do NOT need the same treatment, which surprises people.
// Shifting MSB-first out of the top of the register means the first bit sent
// must be sitting in the TOP bit at load time, so a `len`-bit word has to be
// LEFT-ALIGNED as it is loaded:
//
//     tx_sr <= tx_data << (MAX_W - len)
//
// The receive side needs nothing. Bits arrive at the bottom and walk upward,
// so after exactly `len` captures the word is already right-aligned in the
// low `len` bits with zeros above. One side shifts at load, the other does
// not, and a design that treats them symmetrically gets one of them wrong.
//
// (Bit ORDER is deliberately not here. This block is MSB-first; Chapter 13.8
// generalises it, and keeping that out makes the alignment argument above
// possible to state in one line.)

module spi_shift_datapath #(
    parameter int MAX_W = 32,       // widest frame the datapath supports
    parameter int LEN_W = 6         // width of the bit-count field
) (
    input  wire               clk,
    input  wire               rst_n,

    input  wire [MAX_W-1:0]   tx_data,     // the word to send, right-aligned
    input  wire [LEN_W-1:0]   len,         // bits in this frame, 1 .. MAX_W
    input  wire               load_stb,    // capture tx_data and align it

    input  wire               preload_stb, // from 13.5
    input  wire               launch_stb,
    input  wire               capture_stb,

    input  wire               miso,        // already synchronised upstream

    output wire               mosi,
    output wire [MAX_W-1:0]   rx_data,     // right-aligned, zero-extended
    output wire               rx_valid_stb // one cycle, at the final capture
);

    reg [MAX_W-1:0] tx_sr;
    reg [MAX_W-1:0] rx_sr;
    reg [LEN_W-1:0] caps;
    reg             mosi_r;
    reg             rx_valid_r;

    // MOSI is a REGISTER, not a mux off the shift register. A slave samples
    // this pin on a clock edge the master also generates, so what matters is
    // that it is stable well before that edge and changes at exactly one
    // known moment -- which a flop guarantees and combinational logic off a
    // shifting register does not.
    assign mosi    = mosi_r;
    assign rx_data = rx_sr;

    // DELAYED BY ONE CYCLE, DELIBERATELY. The final bit is shifted into
    // `rx_sr` by the same clock edge that sees the last `capture_stb`, so on
    // that cycle `rx_data` still holds the word one bit short. Pulsing valid
    // there would mean "the data is good next cycle" -- a contract every
    // consumer gets wrong once. Registering it costs one flop and makes
    // valid mean what it says: while `rx_valid_stb` is high, `rx_data` is
    // the complete received word.
    assign rx_valid_stb = rx_valid_r;

    // The bit that goes out next is always the top of the transmit register.
    wire next_tx_bit = tx_sr[MAX_W-1];

    always_ff @(posedge clk or negedge rst_n) begin
        if (!rst_n) begin
            tx_sr      <= {MAX_W{1'b0}};
            rx_sr      <= {MAX_W{1'b0}};
            caps       <= {LEN_W{1'b0}};
            mosi_r     <= 1'b0;
            rx_valid_r <= 1'b0;
        end else begin
            rx_valid_r <= capture_stb & (caps == (len - 1'b1));

            if (load_stb) begin
                // Left-align so the first bit out is in the top position.
                // A shift by a variable amount is a barrel shifter; at these
                // widths it is a few levels of logic and it happens once per
                // frame, off the SCLK path entirely.
                tx_sr <= tx_data << (MAX_W - len);
                rx_sr <= {MAX_W{1'b0}};
                caps  <= {LEN_W{1'b0}};
            end

            // Preload and launch do the SAME thing to the datapath: drive
            // the top bit and advance. They are separate strobes only
            // because 13.5 knows they happen at different MOMENTS -- one
            // before the clock starts, one on an edge -- and the datapath
            // does not need to care.
            //
            // The tempting mistake is to make the preload drive without
            // advancing, on the reasoning that it has not consumed an edge.
            // It has consumed a BIT, and that is what the register counts.
            // Leaving it in place makes the first launch re-send bit zero,
            // so a CPHA=0 frame ships bit zero twice and never ships the
            // last bit at all -- and because the slave usually echoes, the
            // damage shows up as a receive word shifted right by one with a
            // stale bit on the front, which reads like a capture-timing bug
            // rather than a transmit one.
            if (preload_stb || launch_stb) begin
                mosi_r <= next_tx_bit;
                tx_sr  <= {tx_sr[MAX_W-2:0], 1'b0};
            end

            if (capture_stb) begin
                rx_sr <= {rx_sr[MAX_W-2:0], miso};
                caps  <= caps + 1'b1;
            end
        end
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_shift_datapath_tb.sv — loopback and an independent slave, 91 frames
// spi_shift_datapath_tb.sv
//
// By this chapter the testbench has quietly become a master: the divider of
// 13.4 drives the mode logic of 13.5, which drives the datapath under test.
// Only the sequencing (chip select, load, start) is still done by hand, and
// that is exactly what Chapters 13.7 and 13.11 take over.
//
// Two experiments run over every configuration:
//
//   LOOPBACK   MOSI tied to MISO. Whatever is sent must come back, which
//              tests the two registers against each other and catches any
//              off-by-one in the preload.
//   SLAVE      an independent pattern driven on MISO, advanced on the same
//              drive strobes the master uses. This catches the failure
//              loopback CANNOT see: a receive path that is really just
//              echoing the transmit register.
//
// Plus a continuous check that MOSI moves ONLY on a drive strobe -- the
// property that makes it safe for a slave to sample it on a clock edge.

`timescale 1ns/1ps

module spi_shift_datapath_tb;

    localparam int MAX_W = 32;
    localparam int LEN_W = 6;
    localparam int DIV_W = 8;

    logic clk = 1'b0;
    logic rst_n = 1'b0;
    always #5 clk = ~clk;

    // --- 13.4 -------------------------------------------------------------
    logic             en   = 1'b0;
    logic [DIV_W-1:0] div  = 8'd4;
    logic             cpol = 1'b0;
    wire              sclk, edge_a_stb, edge_b_stb, bit_done, div_err;
    wire [DIV_W-1:0]  half_a, half_b;

    spi_clkdiv_strobe #(.DIV_W(DIV_W)) u_div (
        .clk(clk), .rst_n(rst_n), .en(en), .div(div), .cpol(cpol),
        .sclk(sclk), .edge_a_stb(edge_a_stb), .edge_b_stb(edge_b_stb),
        .bit_done(bit_done), .half_a(half_a), .half_b(half_b),
        .div_err(div_err)
    );

    // --- 13.5 -------------------------------------------------------------
    logic             cpha      = 1'b0;
    logic [LEN_W-1:0] len       = 6'd8;
    logic             active    = 1'b0;
    logic             start_stb = 1'b0;
    wire              preload_stb, launch_stb, capture_stb, frame_done;
    wire [LEN_W-1:0]  bit_idx;

    spi_mode_edges #(.LEN_W(LEN_W)) u_mode (
        .clk(clk), .rst_n(rst_n), .cpha(cpha), .len(len),
        .active(active), .start_stb(start_stb),
        .edge_a_stb(edge_a_stb), .edge_b_stb(edge_b_stb),
        .preload_stb(preload_stb), .launch_stb(launch_stb),
        .capture_stb(capture_stb), .bit_idx(bit_idx), .frame_done(frame_done)
    );

    // --- 13.6, the block under test ---------------------------------------
    logic [MAX_W-1:0] tx_data  = 32'h0;
    logic             load_stb = 1'b0;
    wire              miso;
    wire              mosi;
    wire [MAX_W-1:0]  rx_data;
    wire              rx_valid_stb;

    spi_shift_datapath #(.MAX_W(MAX_W), .LEN_W(LEN_W)) dut (
        .clk(clk), .rst_n(rst_n),
        .tx_data(tx_data), .len(len), .load_stb(load_stb),
        .preload_stb(preload_stb), .launch_stb(launch_stb),
        .capture_stb(capture_stb),
        .miso(miso), .mosi(mosi),
        .rx_data(rx_data), .rx_valid_stb(rx_valid_stb)
    );

    // --- the slave model --------------------------------------------------
    // An ideal slave: it presents its next bit on exactly the strobes the
    // master drives on, which is what an edge-aligned slave does on the wire.
    // A real slave is the master's mirror: it puts its bit out on the same
    // strobe the master launches on and HOLDS it in a flop until the far end
    // samples. Modelling it combinationally off the shift register instead
    // makes CPHA=1 read one bit early, because the register has already moved
    // on by the time the trailing edge captures.
    logic             loopback = 1'b1;
    logic [MAX_W-1:0] slave_sr = 32'h0;
    logic             slave_bit_r = 1'b0;

    assign miso = loopback ? mosi : slave_bit_r;

    always_ff @(posedge clk) begin
        if (preload_stb || launch_stb) begin
            slave_bit_r <= slave_sr[MAX_W-1];
            slave_sr    <= {slave_sr[MAX_W-2:0], 1'b0};
        end
    end

    // --- MOSI stability monitor -------------------------------------------
    // MOSI may change only on a cycle carrying a drive strobe. Anything else
    // is a glitch a slave could sample.
    logic mosi_q;
    logic preload_q, launch_q;
    integer mosi_moves_illegally = 0;
    integer mosi_moves = 0;
    always_ff @(posedge clk) begin
        mosi_q <= mosi;
        if (rst_n) begin
            if (mosi !== mosi_q) begin
                mosi_moves <= mosi_moves + 1;
                // The strobes are sampled one cycle back, because MOSI is
                // registered and so moves the cycle AFTER the strobe.
                if (!(preload_q || launch_q))
                    mosi_moves_illegally <= mosi_moves_illegally + 1;
            end
        end
    end
    always_ff @(posedge clk) begin
        preload_q <= preload_stb;
        launch_q  <= launch_stb;
    end

    // --- rx_data stability ------------------------------------------------
    // Once the frame's last bit has been captured the received word must not
    // move again until the next load. A datapath that keeps shifting on
    // stray strobes fails here and nowhere else.
    logic [MAX_W-1:0] rx_at_valid;
    logic             watch_rx = 1'b0;
    integer           rx_disturbed = 0;
    always_ff @(posedge clk) begin
        if (rx_valid_stb) begin
            rx_at_valid <= rx_data;
            watch_rx    <= 1'b1;
        end else if (load_stb) begin
            watch_rx <= 1'b0;
        end else if (watch_rx && rx_data !== rx_at_valid) begin
            rx_disturbed <= rx_disturbed + 1;
        end
    end

    integer errors = 0;
    integer n_frames = 0;
    integer valid_pulses = 0;
    always_ff @(posedge clk) if (rx_valid_stb) valid_pulses <= valid_pulses + 1;

    function automatic [MAX_W-1:0] mask_to(input integer nbits);
        begin
            mask_to = (nbits >= MAX_W) ? {MAX_W{1'b1}}
                                       : ((32'h1 << nbits) - 32'h1);
        end
    endfunction

    task automatic run_frame(input integer dv, input bit pol, input bit pha,
                             input integer nbits, input [MAX_W-1:0] tx,
                             input [MAX_W-1:0] sl, input bit loop_en);
        integer guard;
        begin
            @(negedge clk);
            en = 1'b0; active = 1'b0; start_stb = 1'b0; load_stb = 1'b0;
            div = dv[DIV_W-1:0]; cpol = pol; cpha = pha;
            len = nbits[LEN_W-1:0];
            tx_data  = tx & mask_to(nbits);
            loopback = loop_en;
            // The slave presents its first bit immediately, left-aligned the
            // same way the master's transmit word is.
            slave_sr = (sl & mask_to(nbits)) << (MAX_W - nbits);
            repeat (3) @(negedge clk);

            @(negedge clk);
            load_stb = 1'b1;
            @(negedge clk);
            load_stb = 1'b0;
            @(negedge clk);

            start_stb = 1'b1; active = 1'b1;
            @(negedge clk);
            start_stb = 1'b0;
            repeat (3) @(negedge clk);

            en = 1'b1;
            guard = dv * (nbits + 4) + 60;
            while (!frame_done && guard > 0) begin
                @(negedge clk);
                guard = guard - 1;
            end
            if (guard == 0) begin
                $display("  FAIL: frame never completed (div=%0d len=%0d)",
                         dv, nbits);
                errors = errors + 1;
            end
            repeat (dv + 2) @(negedge clk);
            en = 1'b0; active = 1'b0;
            repeat (3) @(negedge clk);
            n_frames = n_frames + 1;
        end
    endtask

    task automatic expect_rx(input [MAX_W-1:0] want, input integer nbits,
                             input string tag);
        begin
            if (rx_data !== (want & mask_to(nbits))) begin
                $display("  FAIL: %0s len=%0d div=%0d cpol=%0d cpha=%0d expected rx %08h, got %08h",
                         tag, nbits, div, cpol, cpha, want & mask_to(nbits), rx_data);
                errors = errors + 1;
            end
        end
    endtask

    integer lens [0:6];
    integer dvs  [0:2];
    integer l, d, p, h, seed, k;
    logic [MAX_W-1:0] tv, sv;

    initial begin
        lens[0] = 1;  lens[1] = 2;  lens[2] = 4; lens[3] = 8;
        lens[4] = 12; lens[5] = 16; lens[6] = 32;
        dvs[0] = 2; dvs[1] = 3; dvs[2] = 4;
        seed = 32'h1234_5678;

        repeat (3) @(negedge clk);
        rst_n = 1'b1;
        @(negedge clk);

        // 1. LOOPBACK, one byte, mode 0. The simplest thing that can work.
        run_frame(4, 1'b0, 1'b0, 8, 32'hA5, 32'h00, 1'b1);
        expect_rx(32'hA5, 8, "loopback mode 0");
        $display("  loopback, mode 0, 8 bits: sent A5, received %02h",
                 rx_data[7:0]);

        // 2. THE SAME BYTE WITH AN INDEPENDENT SLAVE. If the receive path
        //    were secretly echoing the transmit register, test 1 would still
        //    pass and this one would not.
        run_frame(4, 1'b0, 1'b0, 8, 32'hA5, 32'h3C, 1'b0);
        expect_rx(32'h3C, 8, "independent slave mode 0");
        $display("  independent slave: sent A5, received %02h -- the receive path is its own register",
                 rx_data[7:0]);

        // 3. ALIGNMENT. A 4-bit frame must send the LOW four bits of the
        //    transmit word and return them in the LOW four bits of rx_data,
        //    with the upper bits zero rather than stale.
        run_frame(4, 1'b0, 1'b0, 4, 32'h0000_00F9, 32'h0, 1'b1);
        expect_rx(32'h9, 4, "4-bit alignment");
        if (rx_data[MAX_W-1:4] !== {(MAX_W-4){1'b0}}) begin
            $display("  FAIL: a 4-bit frame left rubbish above bit 3: %08h",
                     rx_data);
            errors = errors + 1;
        end
        $display("  4-bit frame of 0x9: rx_data = %08h -- right-aligned, zero above",
                 rx_data);

        // 4. ALL FOUR MODES, LOOPBACK. The datapath is supposed to be mode
        //    agnostic; this is where that claim is cashed.
        for (p = 0; p <= 1; p = p + 1)
            for (h = 0; h <= 1; h = h + 1) begin
                run_frame(4, p[0], h[0], 8, 32'hC3, 32'h00, 1'b1);
                expect_rx(32'hC3, 8, "all-modes loopback");
            end
        $display("  all four modes returned C3 unchanged -- the datapath never learns the mode");

        // 5. THE SWEEP. Widths 1..32, three divisors, all four modes, a
        //    fresh pattern each time, against an independent slave.
        for (l = 0; l < 7; l = l + 1)
            for (d = 0; d < 3; d = d + 1)
                for (p = 0; p <= 1; p = p + 1)
                    for (h = 0; h <= 1; h = h + 1) begin
                        seed = (seed * 32'h0019_660D) + 32'h3C6E_F35F;
                        tv   = seed;
                        seed = (seed * 32'h0019_660D) + 32'h3C6E_F35F;
                        sv   = seed;
                        run_frame(dvs[d], p[0], h[0], lens[l], tv, sv, 1'b0);
                        expect_rx(sv, lens[l], "sweep");
                    end
        $display("  %0d frames swept: widths 1..32, three divisors, all four modes",
                 n_frames);

        // 6. THE CONTINUOUS PROPERTIES, over everything above.
        if (mosi_moves_illegally != 0) begin
            $display("  FAIL: MOSI moved on %0d cycles with no drive strobe behind them",
                     mosi_moves_illegally);
            errors = errors + 1;
        end
        if (mosi_moves == 0) begin
            $display("  FAIL: MOSI never moved at all -- the monitor proves nothing");
            errors = errors + 1;
        end
        if (rx_disturbed != 0) begin
            $display("  FAIL: rx_data changed %0d times after the frame completed",
                     rx_disturbed);
            errors = errors + 1;
        end
        if (valid_pulses != n_frames) begin
            $display("  FAIL: %0d frames produced %0d rx_valid pulses",
                     n_frames, valid_pulses);
            errors = errors + 1;
        end
        $display("  MOSI moved %0d times, every one of them on a drive strobe; rx_data held still after all %0d completions",
                 mosi_moves, valid_pulses);

        if (errors == 0)
            $display("PASS: the two registers keep the transmit and receive words apart -- an independent slave pattern comes back intact where a secretly-echoing receive path would not -- a frame narrower than the datapath is left-aligned on load and arrives right-aligned with zeros above it, every one of %0d frames spanning widths from 1 to 32 bits, three divisors and all four modes returns exactly what the slave drove, MOSI moves only on the cycle after a preload or launch strobe and never otherwise, rx_data never moves again once the frame completes, and each frame raises rx_valid exactly once", n_frames);
        else
            $display("FAIL: %0d error(s)", errors);
        $finish;
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_shift_datapath.v — the same datapath in Verilog-2001
// spi_shift_datapath.v
//
// Chapter 13.6 -- the shift registers, and why there are two of them.
//
// This block sits underneath the mode logic of 13.5 and sees only three
// strobes: preload, launch, capture. It never learns the mode, the divisor,
// or the polarity. Everything it does is one of:
//
//   load     take a parallel word from the CPU side
//   drive    put the next bit on MOSI          (preload / launch)
//   sample   take the next bit from MISO       (capture)
//   present  hand the assembled word back      (rx_valid_stb)
//
// ONE REGISTER OR TWO?
//
// SPI is exactly full-duplex and exactly symmetric: every launched bit is
// paired with a captured one. That makes a famous trick available -- a SINGLE
// circulating register, MOSI taken from the top bit, MISO shifted into the
// bottom. After N shifts the transmit word has walked out and the receive
// word has walked in, and one register has done both jobs.
//
// It is genuinely elegant and this design does not use it, for three reasons
// that matter in a real master:
//
//   1. The transmit word is DESTROYED by the transfer. A retry after an error
//      has nothing left to re-send, and a flash driver that must repeat a
//      command after a busy status hits this immediately.
//   2. The receive word is only correct at exactly N shifts. Stop early --
//      an abort, Chapter 13.10 -- and what is in the register is a mixture of
//      both words with no boundary marking where.
//   3. It forces the transmit and receive widths to be equal and to equal the
//      register width. Chapter 13.8's configurable width has nowhere to go.
//
// Two registers cost one more register's worth of flops and remove all three.
// On any modern process that is not a trade, it is a rounding error.
//
// THE ALIGNMENT ASYMMETRY.
//
// The two directions do NOT need the same treatment, which surprises people.
// Shifting MSB-first out of the top of the register means the first bit sent
// must be sitting in the TOP bit at load time, so a `len`-bit word has to be
// LEFT-ALIGNED as it is loaded:
//
//     tx_sr <= tx_data << (MAX_W - len)
//
// The receive side needs nothing. Bits arrive at the bottom and walk upward,
// so after exactly `len` captures the word is already right-aligned in the
// low `len` bits with zeros above. One side shifts at load, the other does
// not, and a design that treats them symmetrically gets one of them wrong.
//
// (Bit ORDER is deliberately not here. This block is MSB-first; Chapter 13.8
// generalises it, and keeping that out makes the alignment argument above
// possible to state in one line.)

module spi_shift_datapath #(
    parameter MAX_W = 32,       // widest frame the datapath supports
    parameter LEN_W = 6         // width of the bit-count field
) (
    input  wire               clk,
    input  wire               rst_n,

    input  wire [MAX_W-1:0]   tx_data,     // the word to send, right-aligned
    input  wire [LEN_W-1:0]   len,         // bits in this frame, 1 .. MAX_W
    input  wire               load_stb,    // capture tx_data and align it

    input  wire               preload_stb, // from 13.5
    input  wire               launch_stb,
    input  wire               capture_stb,

    input  wire               miso,        // already synchronised upstream

    output wire               mosi,
    output wire [MAX_W-1:0]   rx_data,     // right-aligned, zero-extended
    output wire               rx_valid_stb // one cycle, at the final capture
);

    reg [MAX_W-1:0] tx_sr;
    reg [MAX_W-1:0] rx_sr;
    reg [LEN_W-1:0] caps;
    reg             mosi_r;
    reg             rx_valid_r;

    // MOSI is a REGISTER, not a mux off the shift register. A slave samples
    // this pin on a clock edge the master also generates, so what matters is
    // that it is stable well before that edge and changes at exactly one
    // known moment -- which a flop guarantees and combinational logic off a
    // shifting register does not.
    assign mosi    = mosi_r;
    assign rx_data = rx_sr;

    // DELAYED BY ONE CYCLE, DELIBERATELY. The final bit is shifted into
    // `rx_sr` by the same clock edge that sees the last `capture_stb`, so on
    // that cycle `rx_data` still holds the word one bit short. Pulsing valid
    // there would mean "the data is good next cycle" -- a contract every
    // consumer gets wrong once. Registering it costs one flop and makes
    // valid mean what it says: while `rx_valid_stb` is high, `rx_data` is
    // the complete received word.
    assign rx_valid_stb = rx_valid_r;

    // The bit that goes out next is always the top of the transmit register.
    wire next_tx_bit = tx_sr[MAX_W-1];

    always @(posedge clk or negedge rst_n) begin
        if (!rst_n) begin
            tx_sr      <= {MAX_W{1'b0}};
            rx_sr      <= {MAX_W{1'b0}};
            caps       <= {LEN_W{1'b0}};
            mosi_r     <= 1'b0;
            rx_valid_r <= 1'b0;
        end else begin
            rx_valid_r <= capture_stb & (caps == (len - 1'b1));

            if (load_stb) begin
                // Left-align so the first bit out is in the top position.
                // A shift by a variable amount is a barrel shifter; at these
                // widths it is a few levels of logic and it happens once per
                // frame, off the SCLK path entirely.
                tx_sr <= tx_data << (MAX_W - len);
                rx_sr <= {MAX_W{1'b0}};
                caps  <= {LEN_W{1'b0}};
            end

            // Preload and launch do the SAME thing to the datapath: drive
            // the top bit and advance. They are separate strobes only
            // because 13.5 knows they happen at different MOMENTS -- one
            // before the clock starts, one on an edge -- and the datapath
            // does not need to care.
            //
            // The tempting mistake is to make the preload drive without
            // advancing, on the reasoning that it has not consumed an edge.
            // It has consumed a BIT, and that is what the register counts.
            // Leaving it in place makes the first launch re-send bit zero,
            // so a CPHA=0 frame ships bit zero twice and never ships the
            // last bit at all -- and because the slave usually echoes, the
            // damage shows up as a receive word shifted right by one with a
            // stale bit on the front, which reads like a capture-timing bug
            // rather than a transmit one.
            if (preload_stb || launch_stb) begin
                mosi_r <= next_tx_bit;
                tx_sr  <= {tx_sr[MAX_W-2:0], 1'b0};
            end

            if (capture_stb) begin
                rx_sr <= {rx_sr[MAX_W-2:0], miso};
                caps  <= caps + 1'b1;
            end
        end
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_shift_datapath_tb.v — the same 91-frame sweep in Verilog-2001
// spi_shift_datapath_tb.v
//
// By this chapter the testbench has quietly become a master: the divider of
// 13.4 drives the mode logic of 13.5, which drives the datapath under test.
// Only the sequencing (chip select, load, start) is still done by hand, and
// that is exactly what Chapters 13.7 and 13.11 take over.
//
// Two experiments run over every configuration:
//
//   LOOPBACK   MOSI tied to MISO. Whatever is sent must come back, which
//              tests the two registers against each other and catches any
//              off-by-one in the preload.
//   SLAVE      an independent pattern driven on MISO, advanced on the same
//              drive strobes the master uses. This catches the failure
//              loopback CANNOT see: a receive path that is really just
//              echoing the transmit register.
//
// Plus a continuous check that MOSI moves ONLY on a drive strobe -- the
// property that makes it safe for a slave to sample it on a clock edge.

`timescale 1ns/1ps

module spi_shift_datapath_tb;

    localparam MAX_W = 32;
    localparam LEN_W = 6;
    localparam DIV_W = 8;

    reg clk;
    reg rst_n;
    always #5 clk = ~clk;

    // --- 13.4 -------------------------------------------------------------
    reg             en;
    reg [DIV_W-1:0] div;
    reg             cpol;
    wire              sclk, edge_a_stb, edge_b_stb, bit_done, div_err;
    wire [DIV_W-1:0]  half_a, half_b;

    spi_clkdiv_strobe #(.DIV_W(DIV_W)) u_div (
        .clk(clk), .rst_n(rst_n), .en(en), .div(div), .cpol(cpol),
        .sclk(sclk), .edge_a_stb(edge_a_stb), .edge_b_stb(edge_b_stb),
        .bit_done(bit_done), .half_a(half_a), .half_b(half_b),
        .div_err(div_err)
    );

    // --- 13.5 -------------------------------------------------------------
    reg             cpha;
    reg [LEN_W-1:0] len;
    reg             active;
    reg             start_stb;
    wire              preload_stb, launch_stb, capture_stb, frame_done;
    wire [LEN_W-1:0]  bit_idx;

    spi_mode_edges #(.LEN_W(LEN_W)) u_mode (
        .clk(clk), .rst_n(rst_n), .cpha(cpha), .len(len),
        .active(active), .start_stb(start_stb),
        .edge_a_stb(edge_a_stb), .edge_b_stb(edge_b_stb),
        .preload_stb(preload_stb), .launch_stb(launch_stb),
        .capture_stb(capture_stb), .bit_idx(bit_idx), .frame_done(frame_done)
    );

    // --- 13.6, the block under test ---------------------------------------
    reg [MAX_W-1:0] tx_data;
    reg             load_stb;
    wire              miso;
    wire              mosi;
    wire [MAX_W-1:0]  rx_data;
    wire              rx_valid_stb;

    spi_shift_datapath #(.MAX_W(MAX_W), .LEN_W(LEN_W)) dut (
        .clk(clk), .rst_n(rst_n),
        .tx_data(tx_data), .len(len), .load_stb(load_stb),
        .preload_stb(preload_stb), .launch_stb(launch_stb),
        .capture_stb(capture_stb),
        .miso(miso), .mosi(mosi),
        .rx_data(rx_data), .rx_valid_stb(rx_valid_stb)
    );

    // --- the slave model --------------------------------------------------
    // An ideal slave: it presents its next bit on exactly the strobes the
    // master drives on, which is what an edge-aligned slave does on the wire.
    // A real slave is the master's mirror: it puts its bit out on the same
    // strobe the master launches on and HOLDS it in a flop until the far end
    // samples. Modelling it combinationally off the shift register instead
    // makes CPHA=1 read one bit early, because the register has already moved
    // on by the time the trailing edge captures.
    reg             loopback;
    reg [MAX_W-1:0] slave_sr;
    reg             slave_bit_r;

    assign miso = loopback ? mosi : slave_bit_r;

    always @(posedge clk) begin
        if (preload_stb || launch_stb) begin
            slave_bit_r <= slave_sr[MAX_W-1];
            slave_sr    <= {slave_sr[MAX_W-2:0], 1'b0};
        end
    end

    // --- MOSI stability monitor -------------------------------------------
    // MOSI may change only on a cycle carrying a drive strobe. Anything else
    // is a glitch a slave could sample.
    reg mosi_q;
    reg preload_q, launch_q;
    integer mosi_moves_illegally;
    integer mosi_moves;
    always @(posedge clk) begin
        mosi_q <= mosi;
        if (rst_n) begin
            if (mosi !== mosi_q) begin
                mosi_moves <= mosi_moves + 1;
                // The strobes are sampled one cycle back, because MOSI is
                // registered and so moves the cycle AFTER the strobe.
                if (!(preload_q || launch_q))
                    mosi_moves_illegally <= mosi_moves_illegally + 1;
            end
        end
    end
    always @(posedge clk) begin
        preload_q <= preload_stb;
        launch_q  <= launch_stb;
    end

    // --- rx_data stability ------------------------------------------------
    // Once the frame's last bit has been captured the received word must not
    // move again until the next load. A datapath that keeps shifting on
    // stray strobes fails here and nowhere else.
    reg [MAX_W-1:0] rx_at_valid;
    reg             watch_rx;
    integer           rx_disturbed;
    always @(posedge clk) begin
        if (rx_valid_stb) begin
            rx_at_valid <= rx_data;
            watch_rx    <= 1'b1;
        end else if (load_stb) begin
            watch_rx <= 1'b0;
        end else if (watch_rx && rx_data !== rx_at_valid) begin
            rx_disturbed <= rx_disturbed + 1;
        end
    end

    integer errors;
    integer n_frames;
    integer valid_pulses;
    always @(posedge clk) if (rx_valid_stb) valid_pulses <= valid_pulses + 1;

        function [MAX_W-1:0] mask_to;
        input integer nbits;
        begin
            mask_to = (nbits >= MAX_W) ? {MAX_W{1'b1}}
                                       : ((32'h1 << nbits) - 32'h1);
        end
    endfunction

        task run_frame;
        input integer dv;
        input pol;
        input pha;
        input integer nbits;
        input [MAX_W-1:0] tx;
        input [MAX_W-1:0] sl;
        input loop_en;
        integer guard;
        begin
            @(negedge clk);
            en = 1'b0; active = 1'b0; start_stb = 1'b0; load_stb = 1'b0;
            div = dv[DIV_W-1:0]; cpol = pol; cpha = pha;
            len = nbits[LEN_W-1:0];
            tx_data  = tx & mask_to(nbits);
            loopback = loop_en;
            // The slave presents its first bit immediately, left-aligned the
            // same way the master's transmit word is.
            slave_sr = (sl & mask_to(nbits)) << (MAX_W - nbits);
            repeat (3) @(negedge clk);

            @(negedge clk);
            load_stb = 1'b1;
            @(negedge clk);
            load_stb = 1'b0;
            @(negedge clk);

            start_stb = 1'b1; active = 1'b1;
            @(negedge clk);
            start_stb = 1'b0;
            repeat (3) @(negedge clk);

            en = 1'b1;
            guard = dv * (nbits + 4) + 60;
            while (!frame_done && guard > 0) begin
                @(negedge clk);
                guard = guard - 1;
            end
            if (guard == 0) begin
                $display("  FAIL: frame never completed (div=%0d len=%0d)",
                         dv, nbits);
                errors = errors + 1;
            end
            repeat (dv + 2) @(negedge clk);
            en = 1'b0; active = 1'b0;
            repeat (3) @(negedge clk);
            n_frames = n_frames + 1;
        end
    endtask

        task expect_rx;
        input [MAX_W-1:0] want;
        input integer nbits;
        input [8*40:1] tag;
        begin
            if (rx_data !== (want & mask_to(nbits))) begin
                $display("  FAIL: %0s len=%0d div=%0d cpol=%0d cpha=%0d expected rx %08h, got %08h",
                         tag, nbits, div, cpol, cpha, want & mask_to(nbits), rx_data);
                errors = errors + 1;
            end
        end
    endtask

    integer lens [0:6];
    integer dvs  [0:2];
    integer l, d, p, h, seed, k;
    reg [MAX_W-1:0] tv, sv;

    initial begin
        lens[0] = 1;  lens[1] = 2;  lens[2] = 4; lens[3] = 8;
        lens[4] = 12; lens[5] = 16; lens[6] = 32;
        dvs[0] = 2; dvs[1] = 3; dvs[2] = 4;
        seed = 32'h1234_5678;

        repeat (3) @(negedge clk);
        rst_n = 1'b1;
        @(negedge clk);

        // 1. LOOPBACK, one byte, mode 0. The simplest thing that can work.
        run_frame(4, 1'b0, 1'b0, 8, 32'hA5, 32'h00, 1'b1);
        expect_rx(32'hA5, 8, "loopback mode 0");
        $display("  loopback, mode 0, 8 bits: sent A5, received %02h",
                 rx_data[7:0]);

        // 2. THE SAME BYTE WITH AN INDEPENDENT SLAVE. If the receive path
        //    were secretly echoing the transmit register, test 1 would still
        //    pass and this one would not.
        run_frame(4, 1'b0, 1'b0, 8, 32'hA5, 32'h3C, 1'b0);
        expect_rx(32'h3C, 8, "independent slave mode 0");
        $display("  independent slave: sent A5, received %02h -- the receive path is its own register",
                 rx_data[7:0]);

        // 3. ALIGNMENT. A 4-bit frame must send the LOW four bits of the
        //    transmit word and return them in the LOW four bits of rx_data,
        //    with the upper bits zero rather than stale.
        run_frame(4, 1'b0, 1'b0, 4, 32'h0000_00F9, 32'h0, 1'b1);
        expect_rx(32'h9, 4, "4-bit alignment");
        if (rx_data[MAX_W-1:4] !== {(MAX_W-4){1'b0}}) begin
            $display("  FAIL: a 4-bit frame left rubbish above bit 3: %08h",
                     rx_data);
            errors = errors + 1;
        end
        $display("  4-bit frame of 0x9: rx_data = %08h -- right-aligned, zero above",
                 rx_data);

        // 4. ALL FOUR MODES, LOOPBACK. The datapath is supposed to be mode
        //    agnostic; this is where that claim is cashed.
        for (p = 0; p <= 1; p = p + 1)
            for (h = 0; h <= 1; h = h + 1) begin
                run_frame(4, p[0], h[0], 8, 32'hC3, 32'h00, 1'b1);
                expect_rx(32'hC3, 8, "all-modes loopback");
            end
        $display("  all four modes returned C3 unchanged -- the datapath never learns the mode");

        // 5. THE SWEEP. Widths 1..32, three divisors, all four modes, a
        //    fresh pattern each time, against an independent slave.
        for (l = 0; l < 7; l = l + 1)
            for (d = 0; d < 3; d = d + 1)
                for (p = 0; p <= 1; p = p + 1)
                    for (h = 0; h <= 1; h = h + 1) begin
                        seed = (seed * 32'h0019_660D) + 32'h3C6E_F35F;
                        tv   = seed;
                        seed = (seed * 32'h0019_660D) + 32'h3C6E_F35F;
                        sv   = seed;
                        run_frame(dvs[d], p[0], h[0], lens[l], tv, sv, 1'b0);
                        expect_rx(sv, lens[l], "sweep");
                    end
        $display("  %0d frames swept: widths 1..32, three divisors, all four modes",
                 n_frames);

        // 6. THE CONTINUOUS PROPERTIES, over everything above.
        if (mosi_moves_illegally != 0) begin
            $display("  FAIL: MOSI moved on %0d cycles with no drive strobe behind them",
                     mosi_moves_illegally);
            errors = errors + 1;
        end
        if (mosi_moves == 0) begin
            $display("  FAIL: MOSI never moved at all -- the monitor proves nothing");
            errors = errors + 1;
        end
        if (rx_disturbed != 0) begin
            $display("  FAIL: rx_data changed %0d times after the frame completed",
                     rx_disturbed);
            errors = errors + 1;
        end
        if (valid_pulses != n_frames) begin
            $display("  FAIL: %0d frames produced %0d rx_valid pulses",
                     n_frames, valid_pulses);
            errors = errors + 1;
        end
        $display("  MOSI moved %0d times, every one of them on a drive strobe; rx_data held still after all %0d completions",
                 mosi_moves, valid_pulses);

        if (errors == 0)
            $display("PASS: the two registers keep the transmit and receive words apart -- an independent slave pattern comes back intact where a secretly-echoing receive path would not -- a frame narrower than the datapath is left-aligned on load and arrives right-aligned with zeros above it, every one of %0d frames spanning widths from 1 to 32 bits, three divisors and all four modes returns exactly what the slave drove, MOSI moves only on the cycle after a preload or launch strobe and never otherwise, rx_data never moves again once the frame completes, and each frame raises rx_valid exactly once", n_frames);
        else
            $display("FAIL: %0d error(s)", errors);
        $finish;
    end


    initial begin
        clk = 1'b0;
        rst_n = 1'b0;
        en = 1'b0;
        div = 8'd4;
        cpol = 1'b0;
        cpha = 1'b0;
        len = 6'd8;
        active = 1'b0;
        start_stb = 1'b0;
        tx_data = 32'h0;
        load_stb = 1'b0;
        loopback = 1'b1;
        slave_sr = 32'h0;
        slave_bit_r = 1'b0;
        mosi_moves_illegally = 0;
        mosi_moves = 0;
        watch_rx = 1'b0;
        rx_disturbed = 0;
        errors = 0;
        n_frames = 0;
        valid_pulses = 0;
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_shift_datapath.vhd — the same datapath in VHDL
-- spi_shift_datapath.vhd
--
-- Chapter 13.6 -- the shift registers, and why there are two of them.
--
-- This block sits underneath the mode logic of 13.5 and sees only three
-- strobes: preload, launch, capture. It never learns the mode, the divisor,
-- or the polarity.
--
-- ONE REGISTER OR TWO? SPI is exactly full-duplex and exactly symmetric, so a
-- SINGLE circulating register works -- MOSI from the top bit, MISO into the
-- bottom. It is elegant and this design does not use it, for three reasons:
--
--   1. The transmit word is DESTROYED by the transfer, so a retry after an
--      error has nothing left to re-send.
--   2. The receive word is only correct at exactly N shifts. Abort early
--      (Chapter 13.10) and the register holds a mixture of both words with
--      nothing marking the boundary.
--   3. Transmit and receive widths are forced equal and equal to the
--      register width, so Chapter 13.8's configurable width has nowhere to
--      go.
--
-- THE ALIGNMENT ASYMMETRY. Shifting MSB-first out of the TOP means the first
-- bit sent must be in the top position at load time, so a `len`-bit word is
-- LEFT-ALIGNED as it is loaded. The receive side needs nothing: bits arrive
-- at the bottom and walk upward, so after exactly `len` captures the word is
-- already right-aligned with zeros above. One side shifts at load, the other
-- does not.
--
-- (Bit ORDER is deliberately not here. This block is MSB-first; Chapter 13.8
-- generalises it.)

library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

entity spi_shift_datapath is
    generic (
        MAX_W : positive := 32;        -- widest frame the datapath supports
        LEN_W : positive := 6          -- width of the bit-count field
    );
    port (
        clk          : in  std_logic;
        rst_n        : in  std_logic;

        tx_data      : in  std_logic_vector(MAX_W - 1 downto 0);
        len          : in  unsigned(LEN_W - 1 downto 0);  -- 1 .. MAX_W
        load_stb     : in  std_logic;

        preload_stb  : in  std_logic;  -- from 13.5
        launch_stb   : in  std_logic;
        capture_stb  : in  std_logic;

        miso         : in  std_logic;  -- already synchronised upstream

        mosi         : out std_logic;
        rx_data      : out std_logic_vector(MAX_W - 1 downto 0);
        rx_valid_stb : out std_logic
    );
end entity;

architecture rtl of spi_shift_datapath is

    signal tx_sr      : std_logic_vector(MAX_W - 1 downto 0)
                        := (others => '0');
    signal rx_sr      : std_logic_vector(MAX_W - 1 downto 0)
                        := (others => '0');
    signal caps       : unsigned(LEN_W - 1 downto 0) := (others => '0');
    signal mosi_r     : std_logic := '0';
    signal rx_valid_r : std_logic := '0';

begin

    -- MOSI is a REGISTER, not a mux off the shift register. A slave samples
    -- this pin on a clock edge the master also generates, so what matters is
    -- that it is stable well before that edge and changes at exactly one
    -- known moment -- which a flop guarantees and combinational logic off a
    -- shifting register does not.
    mosi    <= mosi_r;
    rx_data <= rx_sr;

    -- DELAYED BY ONE CYCLE, DELIBERATELY. The final bit is shifted into
    -- `rx_sr` by the same clock edge that sees the last `capture_stb`, so on
    -- that cycle the output still holds the word one bit short. Pulsing valid
    -- there would mean "the data is good next cycle" -- a contract every
    -- consumer gets wrong once. Registering it costs one flop and makes valid
    -- mean what it says.
    rx_valid_stb <= rx_valid_r;

    shift : process (clk, rst_n)
    begin
        if rst_n = '0' then
            tx_sr      <= (others => '0');
            rx_sr      <= (others => '0');
            caps       <= (others => '0');
            mosi_r     <= '0';
            rx_valid_r <= '0';
        elsif rising_edge(clk) then
            if capture_stb = '1' and caps = (len - 1) then
                rx_valid_r <= '1';
            else
                rx_valid_r <= '0';
            end if;

            if load_stb = '1' then
                -- Left-align so the first bit out is in the top position. A
                -- shift by a variable amount is a barrel shifter; at these
                -- widths it is a few levels of logic and it happens once per
                -- frame, off the SCLK path entirely.
                tx_sr <= std_logic_vector(
                             shift_left(unsigned(tx_data),
                                        MAX_W - to_integer(len)));
                rx_sr <= (others => '0');
                caps  <= (others => '0');
            end if;

            -- Preload and launch do the SAME thing to the datapath: drive the
            -- top bit and advance. They are separate strobes only because
            -- 13.5 knows they happen at different MOMENTS.
            --
            -- The tempting mistake is to make the preload drive without
            -- advancing, on the reasoning that it has not consumed an edge.
            -- It has consumed a BIT. Leaving it in place makes the first
            -- launch re-send bit zero, so a CPHA=0 frame ships bit zero twice
            -- and never ships the last bit -- and because the slave usually
            -- echoes, that shows up as a receive word shifted right by one
            -- with a stale bit on the front, which reads like a capture-
            -- timing bug rather than a transmit one.
            if preload_stb = '1' or launch_stb = '1' then
                mosi_r <= tx_sr(MAX_W - 1);
                tx_sr  <= tx_sr(MAX_W - 2 downto 0) & '0';
            end if;

            if capture_stb = '1' then
                rx_sr <= rx_sr(MAX_W - 2 downto 0) & miso;
                caps  <= caps + 1;
            end if;
        end if;
    end process;

end architecture;
Azvya Education Pvt. Ltd.VLSI Mentor
spi_shift_datapath_tb.vhd — the same 91-frame sweep in VHDL
-- spi_shift_datapath_tb.vhd
--
-- By this chapter the testbench has quietly become a master: the divider of
-- 13.4 drives the mode logic of 13.5, which drives the datapath under test.
-- Only the sequencing is still done by hand, and that is exactly what
-- Chapters 13.7 and 13.11 take over.
--
--   LOOPBACK   MOSI tied to MISO. Whatever is sent must come back.
--   SLAVE      an independent pattern on MISO, advanced on the same drive
--              strobes the master uses -- this catches the failure loopback
--              CANNOT see: a receive path that is really echoing the
--              transmit register.
--
-- Plus a continuous check that MOSI moves ONLY on a drive strobe.

library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

entity spi_shift_datapath_tb is
end entity;

architecture sim of spi_shift_datapath_tb is

    constant MAX_W : positive := 32;
    constant LEN_W : positive := 6;
    constant DIV_W : positive := 8;

    signal clk   : std_logic := '0';
    signal rst_n : std_logic := '0';
    signal halt  : boolean   := false;

    -- 13.4
    signal en   : std_logic := '0';
    signal div  : unsigned(DIV_W - 1 downto 0) := to_unsigned(4, DIV_W);
    signal cpol : std_logic := '0';
    signal sclk, edge_a_stb, edge_b_stb, bit_done, div_err : std_logic;
    signal half_a, half_b : unsigned(DIV_W - 1 downto 0);

    -- 13.5
    signal cpha      : std_logic := '0';
    signal len       : unsigned(LEN_W - 1 downto 0) := to_unsigned(8, LEN_W);
    signal active    : std_logic := '0';
    signal start_stb : std_logic := '0';
    signal preload_stb, launch_stb, capture_stb, frame_done : std_logic;
    signal bit_idx   : unsigned(LEN_W - 1 downto 0);

    -- 13.6, the block under test
    signal tx_data  : std_logic_vector(MAX_W - 1 downto 0) := (others => '0');
    signal load_stb : std_logic := '0';
    signal miso     : std_logic;
    signal mosi     : std_logic;
    signal rx_data  : std_logic_vector(MAX_W - 1 downto 0);
    signal rx_valid_stb : std_logic;

    -- the slave model. `slave_sr` is driven by the slave process ALONE --
    -- VHDL resolves multiple drivers rather than letting the last writer win,
    -- so having the stimulus load it directly would make every bit the two
    -- disagreed on resolve to 'X'. The stimulus hands over a seed and a load
    -- pulse instead.
    signal loopback    : std_logic := '1';
    signal slave_seed  : std_logic_vector(MAX_W - 1 downto 0)
                         := (others => '0');
    signal slave_load  : std_logic := '0';
    signal slave_sr    : std_logic_vector(MAX_W - 1 downto 0)
                         := (others => '0');
    signal slave_bit_r : std_logic := '0';

    -- monitors
    signal mosi_q : std_logic := '0';
    signal preload_q, launch_q : std_logic := '0';
    signal mosi_moves            : natural := 0;
    signal mosi_moves_illegally  : natural := 0;
    signal rx_at_valid : std_logic_vector(MAX_W - 1 downto 0)
                         := (others => '0');
    signal watch_rx    : std_logic := '0';
    signal rx_disturbed  : natural := 0;
    signal valid_pulses  : natural := 0;
    signal n_frames      : natural := 0;
    signal errors        : natural := 0;

    function mask_to(nbits : natural) return std_logic_vector is
        variable m : std_logic_vector(MAX_W - 1 downto 0);
    begin
        m := (others => '0');
        for k in 0 to MAX_W - 1 loop
            if k < nbits then
                m(k) := '1';
            end if;
        end loop;
        return m;
    end function;

    function hex8(v : std_logic_vector) return string is
        constant DIGITS : string(1 to 16) := "0123456789abcdef";
        variable u : unsigned(MAX_W - 1 downto 0);
        variable r : string(1 to 8);
    begin
        u := unsigned(v);
        for k in 8 downto 1 loop
            r(k) := DIGITS(to_integer(u(3 downto 0)) + 1);
            u := shift_right(u, 4);
        end loop;
        return r;
    end function;

begin

    clk <= not clk after 5 ns when not halt else '0';

    u_div : entity work.spi_clkdiv_strobe
        generic map (DIV_W => DIV_W)
        port map (clk => clk, rst_n => rst_n, en => en, div => div,
                  cpol => cpol, sclk => sclk,
                  edge_a_stb => edge_a_stb, edge_b_stb => edge_b_stb,
                  bit_done => bit_done, half_a => half_a, half_b => half_b,
                  div_err => div_err);

    u_mode : entity work.spi_mode_edges
        generic map (LEN_W => LEN_W)
        port map (clk => clk, rst_n => rst_n, cpha => cpha, len => len,
                  active => active, start_stb => start_stb,
                  edge_a_stb => edge_a_stb, edge_b_stb => edge_b_stb,
                  preload_stb => preload_stb, launch_stb => launch_stb,
                  capture_stb => capture_stb, bit_idx => bit_idx,
                  frame_done => frame_done);

    dut : entity work.spi_shift_datapath
        generic map (MAX_W => MAX_W, LEN_W => LEN_W)
        port map (clk => clk, rst_n => rst_n,
                  tx_data => tx_data, len => len, load_stb => load_stb,
                  preload_stb => preload_stb, launch_stb => launch_stb,
                  capture_stb => capture_stb,
                  miso => miso, mosi => mosi,
                  rx_data => rx_data, rx_valid_stb => rx_valid_stb);

    -- A real slave is the master's mirror: it puts its bit out on the same
    -- strobe the master launches on and HOLDS it in a flop until the far end
    -- samples. Modelling it combinationally off the shift register instead
    -- makes CPHA=1 read one bit early.
    miso <= mosi when loopback = '1' else slave_bit_r;

    slave_p : process (clk)
    begin
        if rising_edge(clk) then
            if slave_load = '1' then
                slave_sr    <= slave_seed;
                slave_bit_r <= '0';
            elsif preload_stb = '1' or launch_stb = '1' then
                slave_bit_r <= slave_sr(MAX_W - 1);
                slave_sr    <= slave_sr(MAX_W - 2 downto 0) & '0';
            end if;
        end if;
    end process;

    -- MOSI may change only on a cycle carrying a drive strobe. Anything else
    -- is a glitch a slave could sample. The strobes are compared one cycle
    -- back, because MOSI is registered and so moves the cycle AFTER.
    mosi_mon : process (clk)
    begin
        if rising_edge(clk) then
            mosi_q    <= mosi;
            preload_q <= preload_stb;
            launch_q  <= launch_stb;
            if rst_n = '1' and mosi /= mosi_q then
                mosi_moves <= mosi_moves + 1;
                if not (preload_q = '1' or launch_q = '1') then
                    mosi_moves_illegally <= mosi_moves_illegally + 1;
                end if;
            end if;
        end if;
    end process;

    -- Once the frame's last bit has been captured the received word must not
    -- move again until the next load.
    rx_mon : process (clk)
    begin
        if rising_edge(clk) then
            if rx_valid_stb = '1' then
                rx_at_valid  <= rx_data;
                watch_rx     <= '1';
                valid_pulses <= valid_pulses + 1;
            elsif load_stb = '1' then
                watch_rx <= '0';
            elsif watch_rx = '1' and rx_data /= rx_at_valid then
                rx_disturbed <= rx_disturbed + 1;
            end if;
        end if;
    end process;

    stim : process
        variable errs : natural := 0;
        variable guard : natural;
        variable seed : unsigned(31 downto 0) := x"12345678";
        variable tv, sv : std_logic_vector(MAX_W - 1 downto 0);

        procedure run_frame(dv : natural; pol : std_logic; pha : std_logic;
                            nbits : natural;
                            tx : std_logic_vector(MAX_W - 1 downto 0);
                            sl : std_logic_vector(MAX_W - 1 downto 0);
                            loop_en : std_logic) is
        begin
            wait until falling_edge(clk);
            en <= '0'; active <= '0'; start_stb <= '0'; load_stb <= '0';
            div <= to_unsigned(dv, DIV_W); cpol <= pol; cpha <= pha;
            len <= to_unsigned(nbits, LEN_W);
            tx_data  <= tx and mask_to(nbits);
            loopback <= loop_en;
            -- The slave presents its first bit left-aligned, the same way the
            -- master's transmit word is.
            slave_seed <= std_logic_vector(
                              shift_left(unsigned(sl and mask_to(nbits)),
                                         MAX_W - nbits));
            for k in 1 to 3 loop wait until falling_edge(clk); end loop;

            wait until falling_edge(clk);
            load_stb <= '1'; slave_load <= '1';
            wait until falling_edge(clk);
            load_stb <= '0'; slave_load <= '0';
            wait until falling_edge(clk);

            start_stb <= '1'; active <= '1';
            wait until falling_edge(clk);
            start_stb <= '0';
            for k in 1 to 3 loop wait until falling_edge(clk); end loop;

            en <= '1';
            guard := dv * (nbits + 4) + 60;
            while frame_done = '0' and guard > 0 loop
                wait until falling_edge(clk);
                guard := guard - 1;
            end loop;
            if guard = 0 then
                report "  FAIL: frame never completed";
                errs := errs + 1;
            end if;
            for k in 1 to dv + 2 loop wait until falling_edge(clk); end loop;
            en <= '0'; active <= '0';
            for k in 1 to 3 loop wait until falling_edge(clk); end loop;
            n_frames <= n_frames + 1;
            wait until falling_edge(clk);
        end procedure;

        procedure expect_rx(want : std_logic_vector(MAX_W - 1 downto 0);
                            nbits : natural; tag : string) is
        begin
            if rx_data /= (want and mask_to(nbits)) then
                report "  FAIL: " & tag & " len=" & integer'image(nbits) &
                       " expected rx " & hex8(want and mask_to(nbits)) &
                       ", got " & hex8(rx_data);
                errs := errs + 1;
            end if;
        end procedure;

        procedure next_rand(variable v : out std_logic_vector(MAX_W - 1 downto 0)) is
        begin
            seed := resize(seed * x"0019660D", 32) + x"3C6EF35F";
            v := std_logic_vector(seed);
        end procedure;

        type int_vec is array (natural range <>) of natural;
        constant LENS : int_vec(0 to 6) := (1, 2, 4, 8, 12, 16, 32);
        constant DVS  : int_vec(0 to 2) := (2, 3, 4);
        variable pol_v, pha_v : std_logic;
        constant ZERO : std_logic_vector(MAX_W - 1 downto 0) := (others => '0');
    begin
        for k in 0 to 2 loop wait until falling_edge(clk); end loop;
        rst_n <= '1';
        wait until falling_edge(clk);

        -- 1. LOOPBACK, one byte, mode 0. The simplest thing that can work.
        run_frame(4, '0', '0', 8, x"000000A5", ZERO, '1');
        expect_rx(x"000000A5", 8, "loopback mode 0");
        report "  loopback, mode 0, 8 bits: sent A5, received " &
               hex8(rx_data);

        -- 2. THE SAME BYTE WITH AN INDEPENDENT SLAVE. If the receive path
        --    were secretly echoing the transmit register, test 1 would still
        --    pass and this one would not.
        run_frame(4, '0', '0', 8, x"000000A5", x"0000003C", '0');
        expect_rx(x"0000003C", 8, "independent slave mode 0");
        report "  independent slave: sent A5, received " & hex8(rx_data) &
               " -- the receive path is its own register";

        -- 3. ALIGNMENT. A 4-bit frame must send the LOW four bits and return
        --    them in the LOW four bits, with the upper bits zero not stale.
        run_frame(4, '0', '0', 4, x"000000F9", ZERO, '1');
        expect_rx(x"00000009", 4, "4-bit alignment");
        if rx_data(MAX_W - 1 downto 4) /= ZERO(MAX_W - 1 downto 4) then
            report "  FAIL: a 4-bit frame left rubbish above bit 3";
            errs := errs + 1;
        end if;
        report "  4-bit frame of 0x9: rx_data = " & hex8(rx_data) &
               " -- right-aligned, zero above";

        -- 4. ALL FOUR MODES, LOOPBACK. The datapath is supposed to be mode
        --    agnostic; this is where that claim is cashed.
        for p in 0 to 1 loop
            for h in 0 to 1 loop
                if p = 1 then pol_v := '1'; else pol_v := '0'; end if;
                if h = 1 then pha_v := '1'; else pha_v := '0'; end if;
                run_frame(4, pol_v, pha_v, 8, x"000000C3", ZERO, '1');
                expect_rx(x"000000C3", 8, "all-modes loopback");
            end loop;
        end loop;
        report "  all four modes returned C3 unchanged -- the datapath never learns the mode";

        -- 5. THE SWEEP. Widths 1..32, three divisors, all four modes, a fresh
        --    pattern each time, against an independent slave.
        for l in LENS'range loop
            for d in DVS'range loop
                for p in 0 to 1 loop
                    for h in 0 to 1 loop
                        if p = 1 then pol_v := '1'; else pol_v := '0'; end if;
                        if h = 1 then pha_v := '1'; else pha_v := '0'; end if;
                        next_rand(tv);
                        next_rand(sv);
                        run_frame(DVS(d), pol_v, pha_v, LENS(l), tv, sv, '0');
                        expect_rx(sv, LENS(l), "sweep");
                    end loop;
                end loop;
            end loop;
        end loop;
        report "  " & integer'image(n_frames) &
               " frames swept: widths 1..32, three divisors, all four modes";

        -- 6. THE CONTINUOUS PROPERTIES, over everything above.
        if mosi_moves_illegally /= 0 then
            report "  FAIL: MOSI moved with no drive strobe behind it";
            errs := errs + 1;
        end if;
        if mosi_moves = 0 then
            report "  FAIL: MOSI never moved at all -- the monitor proves nothing";
            errs := errs + 1;
        end if;
        if rx_disturbed /= 0 then
            report "  FAIL: rx_data changed after the frame completed";
            errs := errs + 1;
        end if;
        if valid_pulses /= n_frames then
            report "  FAIL: " & integer'image(n_frames) & " frames produced " &
                   integer'image(valid_pulses) & " rx_valid pulses";
            errs := errs + 1;
        end if;
        report "  MOSI moved " & integer'image(mosi_moves) &
               " times, every one of them on a drive strobe; rx_data held still after all " &
               integer'image(valid_pulses) & " completions";

        errors <= errs;
        if errs = 0 then
            report "PASS: the two registers keep the transmit and receive words apart -- an independent slave pattern comes back intact where a secretly-echoing receive path would not -- a frame narrower than the datapath is left-aligned on load and arrives right-aligned with zeros above it, every one of " & integer'image(n_frames) & " frames spanning widths from 1 to 32 bits, three divisors and all four modes returns exactly what the slave drove, MOSI moves only on the cycle after a preload or launch strobe and never otherwise, rx_data never moves again once the frame completes, and each frame raises rx_valid exactly once";
        else
            report "FAIL: " & integer'image(errs) & " error(s)" severity error;
        end if;
        halt <= true;
        wait;
    end process;

end architecture;

Parity

All three implementations run 91 frames across widths from 1 to 32 bits, three divisors and all four modes, against an independent slave rather than a loopback — and all three report 451 MOSI transitions, every one of them on the cycle after a drive strobe.

7. Why a Verification Engineer Cares

Azvya Education Pvt. Ltd.VLSI Mentor
spi_shift_datapath.sva — the pin's stability, and the valid contract
// Two properties matter here and neither is about data. The first is that MOSI
// moves only when it is allowed to; the second is that `rx_valid` means what it
// says. Data correctness is the scoreboard's business, not an assertion's.

module spi_shift_datapath_sva #(parameter int MAX_W = 32, LEN_W = 6) (
    input logic              clk,
    input logic              rst_n,
    input logic [LEN_W-1:0]  len,
    input logic              load_stb,
    input logic              preload_stb,
    input logic              launch_stb,
    input logic              capture_stb,
    input logic              mosi,
    input logic [MAX_W-1:0]  rx_data,
    input logic              rx_valid_stb
);

    default clocking cb @(posedge clk); endclocking
    default disable iff (!rst_n);

    // MOSI may change only on the cycle after a drive strobe. Anything else is a
    // glitch a slave could sample, and this is the property that makes the flop
    // worth its one cycle of latency.
    a_mosi_stable: assert property (
        !$stable(mosi) |-> $past(preload_stb || launch_stb)
    );

    // And it must move at some point, or the assertion above is vacuous. A
    // cover, not an assert: it is a statement about the stimulus.
    c_mosi_moves: cover property (!$stable(mosi));

    // The valid contract: when valid is high the data is COMPLETE, which means
    // it must not change on that cycle or the consumer reads a moving value.
    a_valid_stable: assert property (rx_valid_stb |-> $stable(rx_data));

    // Once complete, the received word holds until the next load. A datapath
    // that keeps shifting on a stray strobe fails here and nowhere else.
    a_rx_held: assert property (
        rx_valid_stb ##1 !load_stb[*1:$] ##0 !capture_stb |-> $stable(rx_data)
    );

    // Exactly one valid per frame, in both directions.
    int caps;
    always_ff @(posedge clk) begin
        if (!rst_n || load_stb) caps <= 0;
        else if (capture_stb)   caps <= caps + 1;
    end
    a_valid_at_len:   assert property (rx_valid_stb |-> $past(caps) == len);
    a_len_gives_valid: assert property (
        capture_stb && caps == len - 1 |=> rx_valid_stb
    );

    // A frame narrower than the datapath must leave zeros above it, not stale
    // bits from the previous frame -- the failure a byte-only suite never sees.
    a_zero_extended: assert property (
        rx_valid_stb && len < MAX_W |-> rx_data >> len == '0
    );

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_shift_datapath_cg.sv — widths, and where the alignment bites
// The axis that matters is the frame width relative to the datapath width,
// because that is what the alignment shift depends on. An all-byte suite
// exercises exactly one value of it.

covergroup cg_shift_datapath (int MAX_W) @(posedge clk);

    // Absolute width, with the corners named. One bit makes every shift
    // degenerate; MAX_W makes the alignment shift zero.
    width: coverpoint len iff (load_stb) {
        bins one        = {1};
        bins two        = {2};
        bins narrow     = {[3:7]};
        bins byte_      = {8};
        bins odd_middle = {[9:15]};
        bins wide       = {[16:31]};
        bins full       = {32};
    }

    // The alignment shift amount, which is the thing actually being exercised.
    // Zero means no shift at all and is a distinct case.
    shift_amt: coverpoint (MAX_W - len) iff (load_stb) {
        bins none   = {0};
        bins small  = {[1:8]};
        bins medium = {[9:24]};
        bins large  = {[25:31]};
    }

    // Whether the receive data's upper bits were non-zero BEFORE the frame.
    // A frame that follows an all-zero frame cannot detect a missing clear.
    prior_upper: coverpoint prior_rx_upper_nonzero iff (load_stb && len < MAX_W) {
        bins was_dirty = {1};
        bins was_clean = {0};
    }

    // And the independent-slave versus loopback axis, because a suite that is
    // all loopback cannot see an echoing receive path.
    source: coverpoint miso_source {
        bins loopback    = {0};
        bins independent = {1};
    }

    x_width_source: cross width, source;
    x_narrow_dirty: cross width, prior_upper;

endgroup

8. Why an FPGA or ASIC Engineer Cares

Two MAX_W registers instead of one is the whole cost. At MAX_W = 32 that is 32 extra flops against a master of a few hundred. In exchange the design can retry, abort and support configurable widths, all three of which the circulating version cannot express.

The barrel shifter is the only significant combinational block, and it is on the load path. A variable shift of a 32-bit word is five levels of 2:1 muxes — real logic, and irrelevant here because it is evaluated on one cycle per frame while the SCLK-rate logic idles. Moving it onto the per-bit path, which is what a bidirectional shift register does, is Chapter 13.8's subject.

The per-bit path is a shift and a mux, and nothing else. tx_sr <= {tx_sr[MAX_W-2:0], 1'b0} is a wire-only transformation; rx_sr <= {rx_sr[MAX_W-2:0], miso} is the same plus one input. Neither has arithmetic, neither has a comparator, and both are enabled by strobes computed a cycle earlier. There is nothing here to fail timing.

MISO is sampled into a flop and used nowhere else. That is worth stating as an implementation rule rather than a design detail: MISO is an asynchronous input from another device, so the only correct thing to do with it is register it. Any combinational path from MISO to an output pin — a loopback mode, a pass-through, a status mux — is a path from an asynchronous input to an output with no flop in it, and no timing tool will tell you what to constrain.

9. Failure Signature — An ADC That Reads Correctly Only In Its Top Bits

Symptom. A 12-bit ADC is read over SPI. The top eight bits track the input correctly. The bottom four are always the same value, and that value changes when an unrelated 8-bit sensor on the same bus is read.

What the pattern says. Bits from a previous transfer are appearing in the current one, in the bit positions the current frame does not occupy. That is a stale-data fault in the receive register, and the fact that the stale bits come from the other device identifies which register: the receive shift register was not cleared, so the frame's 12 bits were shifted in on top of the previous frame's 8.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   after the 8-bit read:      0000_0000_0000_0000_0000_0000_aaaa_aaaa
   after 12 more shifts:      0000_0000_0000_0000_aaaa_aaaa_bbbb_bbbb
                                                  ^^^^^^^^^ stale
   software reads the low 12: aaaa_bbbb_bbbb
                              ^^^^ the other device's bits

Why the top eight bits looked right. They are not the top eight bits of the ADC value — they are the ADC's top eight, shifted down four places and read as bits 11 to 4 of a value whose low four bits are wrong. A monotonic input produces a monotonic-looking reading, and the error is a constant offset plus a scale factor, which looks like a calibration problem.

Why an 8-bit-only suite never found it. At len == MAX_W there are no bits above the frame, so nothing stale can appear. At len == 8 with MAX_W == 8 likewise. The fault requires len < MAX_W and a preceding frame that left non-zero bits above the current frame's width — which is the x_narrow_dirty cross in the coverage model above, and it is the reason that cross exists.

The fix. Clear rx_sr on load. It is one line, and the assertion a_zero_extended is what stops it from being removed by someone optimising the load path.

The general lesson. A register that accumulates must be cleared by the same event that defines what it accumulates. Relying on "it will be fully overwritten anyway" is only true when the frame fills the register, and a configurable width is precisely the feature that makes it untrue.

10. Common Misconceptions

"One circulating register is strictly better, since SPI is symmetric." It is smaller and it cannot represent a retry, a partial transfer, or a frame narrower than the register. The symmetry is real; what breaks is that the design needs states in which the symmetry does not hold.

"Both directions need the alignment shift." Only the transmit side does. Bits arriving at the bottom of the receive register end up right-aligned by construction. Shifting the receive word as well gives a result that is correct at full width and wrong at every narrower frame.

"MOSI can be a mux off the shift register; the slave samples half a period later anyway." It samples half a period after the edge, and the register settles during the cycle the shift happens — which is the same cycle. A flop makes the pin change once at a known moment; a mux makes it follow the register's settling, and every intermediate value is on the wire.

"rx_valid should pulse on the last capture, since that is when the last bit arrives." The bit arrives at that clock edge, so during that cycle the register still holds the word one bit short. Pulsing there means "good next cycle", which every consumer gets wrong once. One flop makes valid mean what it says.

"A loopback test exercises both directions at once and is therefore sufficient." It exercises one path twice. A receive path that returns the transmit register passes it, and so does a design with both directions' bit order reversed. Loopback is a useful first test and never a complete one.

11. Reason It Through

Why is the receive register cleared on load rather than relying on the frame to overwrite it?

Because a frame narrower than the register does not overwrite the whole of it. The bits above the frame's width retain the previous frame's data, and software reading the low len bits gets a value whose upper part came from the previous transfer — which is the ADC failure of §9. "It will be overwritten anyway" is true only when the frame fills the register, and configurable width is exactly what makes it untrue.

What is lost by aborting a transfer in the circulating-register design?

Everything. The register holds a mixture of the transmit and receive words with no marker at the boundary, so neither can be recovered: the partial receive data cannot be separated from the un-sent transmit data, and the transmit word cannot be re-sent. With two registers, the transmit word is intact and the receive register holds the captured bits in a known position with a known count — which is what makes Chapter 13.10's partial-bit report possible.

Why does the launch strobe advance the register while the capture strobe advances a different one?

Because they are different registers and different directions. launch_stb shifts tx_sr and updates the MOSI flop; capture_stb shifts rx_sr and increments the count. They happen on opposite edges of SCLK and are never simultaneous, so the two shifts never contend — which is another thing the circulating version cannot say, since there the two operations are the same shift.

The MOSI monitor asserts both that MOSI never moved illegally and that it moved at all. Why the second?

Because the first is vacuous if nothing moves. A design that drove MOSI to a constant would satisfy "never moved illegally" perfectly, and a testbench reporting only that would be reporting nothing. Any assertion of the form "X never happens without Y" needs a companion cover that X happened, and the companion is cheap.

A design passes every loopback test and fails against a real device in one direction only. What is the most likely class of fault?

Something that is symmetric under loopback and asymmetric in reality. The candidates are: a receive path that returns the transmit register rather than the captured bits; both directions' bit order reversed, which cancels; or an alignment applied to both directions when it belongs to one. All three are invisible to loopback by construction, and all three are found by driving an independent pattern on MISO — which is why the independent-slave test exists rather than merely being a nicer version of loopback.

12. Understanding Check

13. Summary

SPI's symmetry permits a single circulating register, and this design does not use it for three reasons: the transmit word is destroyed so a retry has nothing to re-send; the receive word is correct only at exactly N shifts so an abort leaves an unrecoverable mixture; and the two widths are forced equal to the register width so configurable frames have nowhere to go. The cost of avoiding all three is one register's worth of flops.

The general form of that argument is worth keeping: a single-resource optimisation relying on two things being exactly matched fails when they stop being matched, and the decisive objection is not that it is hard to read but that it cannot represent states the design needs.

The two directions are asymmetric. Transmit needs a left-align at load because the first bit out must be in the top position; receive needs nothing, because bits arriving at the bottom end up right-aligned. Shifting both gives a result that is correct at full width and wrong at every narrower frame.

MOSI is a flop. A slave samples it on an edge this master generates, so the pin must be stable well before that edge and move at exactly one known moment — which a register guarantees and a mux off a shifting register does not. The companion property is that MOSI moves only on the cycle after a drive strobe, plus a cover that it moved at all.

rx_valid_stb is delayed one cycle so that when it is high the data is complete rather than one bit short. And rx_sr is cleared on load, because a frame narrower than the register does not overwrite the bits above it — the failure that makes a 12-bit ADC read with another device's bits in its low nibble.

Preload and launch are the same datapath operation: drive the top bit, advance the register. Separating them produces the bit-zero-twice bug of Chapter 13.5.

For verification, loopback is never sufficient — it exercises one path twice and passes a receive path that echoes the transmit register. An independent MISO pattern is required, and building the slave model surfaced that a peer device must register its output bit for the same reason the master does. Coverage's real axis is the frame width relative to the datapath width, crossed with whether the upper bits were dirty beforehand.

14. What Comes Next

Bits move, in both directions, in every mode. Nothing yet owns the select pin.

Chapter 13.7 — Chip-Select Generation builds the block that does, and it starts from the one-line implementation everybody writes first — cs_n = ~busy — which is wrong three times over. One of those three is the most common reason a flash driver reads back 0xFF.

Continue learning