Skip to content
VLSI Mentor

SPI · Module 6

Command-Then-Read Sequences

Why every SPI read is a write first, why request and response must share one CS frame, what the master drives once its half is done, what the request costs in bus time, and the master read sequencer in three HDLs.

Chapter 6.1 followed a read from the device's side: fetch, load, drive, release. This chapter takes the master's side, and it begins with an observation that sounds trivial and is not.

Before a device can return anything, it has to be told what to return. So a read begins as a transmission — and the master must change what it is doing part-way through a frame it is not allowed to interrupt.

That constraint shapes the master's hardware, its bus efficiency, and one of the more confusing failure modes in SPI.

1. Every Read Is a Write First

On a bus with separate command and response channels, a read is one operation. On SPI it is two halves of one frame:

The request half. The master transmits an opcode and usually an address on MOSI. During this, MISO carries nothing meaningful — the device does not yet know what is being asked (Chapter 6.1 §4).

The response half. The device drives MISO. During this, MOSI carries nothing meaningful — the master has already said everything it needs to.

So an SPI read is half-duplex traffic carried on a full-duplex bus. Both directions are physically active the whole time; only one of them means anything at any moment. Chapter 1.4 established that the bus cannot do otherwise — every edge shifts both registers — and this is where that fact starts costing something.

2. One Frame, Not Two

A natural question: why not send the command in one transaction, then read the data in another? It would simplify the driver considerably.

For most devices it does not work, and Chapter 5.1 §4 already supplies the reason. CS deassertion ends the transaction. A device that sees CS rise after the address concludes the transaction is over, discards its state, and resynchronises its sequencer (Chapter 4.4 §7). When CS falls again it expects a new command, so the master's read clocks are interpreted as an opcode followed by address bytes — decoding whatever the filler happens to be.

That is why the sequencer in §6 holds cs_n low from ST_CMD all the way through ST_DATA, and why ST_CLOSE exists as a separate state: the frame must not end until the last data byte has been collected.

There are real exceptions, and knowing which is which is a datasheet question:

  • Devices with a latched result. Some ADCs and sensors perform a conversion on one transaction and return it on the next. There the two-frame pattern is not a workaround but the specified protocol.
  • Status registers. A one-byte command whose response begins immediately often tolerates either arrangement.
  • Devices with an explicit continuous-read mode, which keep a pointer across frames deliberately.

The general rule stands: assume one frame unless the datasheet says otherwise, because the failure when you are wrong is silent and looks like data corruption rather than a framing error.

3. What the Master Drives During the Response

The request, the turnaround, and the response — one frame

8 cycles
Byte-time lanes for an SPI read. MOSI carries a read opcode and three address bytes, then two filler bytes of 0x00. MISO is high impedance for the first five byte times, then carries two data bytes. Chip select stays low throughout.request beginsrequest beginsresponse beginsresponse beginscs_nmosi--0x0B0x120x340x560x000x00--misoD0D1t0t1t2t3t4t5t6t7
Figure 1 — one frame spanning the direction change. MOSI carries the opcode and address, then filler bytes once the master's half is finished; MISO is high impedance until the device takes the line. Chip select stays low across the turnaround, because releasing it would end the transaction rather than pausing it.

The 0x00 bytes on MOSI during the response are the visible form of Chapter 5.2 §3: there is no "not sending". The master must drive every edge it clocks, so it supplies filler. Which value is a device-level agreement — 0x00 and 0xFF are both common, and a few devices specify one or interpret particular values as a subsequent command.

The sequencer in §6 makes that a parameter (FILLER) rather than hard-coding zero, because getting it wrong on a device that cares produces a fault that looks nothing like a filler problem.

4. What the Request Costs

The request is pure overhead from the payload's point of view, and quantifying it explains a great deal about how SPI devices are used.

For a read with a 1-byte opcode, a 3-byte address and 8 dummy cycles:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   opcode    8 cycles
   address  24 cycles
   dummy     8 cycles
   ──────────────────────
   request  40 cycles before the first payload bit

Efficiency therefore depends entirely on how much data follows:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   payload       total cycles         efficiency
      1 byte     40 +     8 =    48      16.7 %
      4 bytes    40 +    32 =    72      44.4 %
     32 bytes    40 +   256 =   296      86.5 %
    256 bytes    40 + 2048 =  2088       98.1 %

Two consequences worth holding.

Small reads are dominated by the request. Reading one byte spends five sixths of the bus restating what you want. Reading a single status register — the most common operation in a polling loop — is about as inefficient as SPI gets, which is why devices that expect to be polled often place the status byte behind a one-byte command with no address and no dummy phase, bringing the request down to 8 cycles.

Raising the clock does not help as much as it appears to. Doubling SCLK halves the wall-clock time of both the payload and the request, so the ratio is unchanged — and Chapter 4.5 §3 showed the dummy count may increase with frequency. For small reads, reducing the number of transactions beats raising the clock.

5. The Transaction

A sequence diagram of a master-side read sequence. The master asserts chip select, transmits an opcode and address, clocks a dummy phase while driving filler, then collects two data bytes from the device, and finally releases chip select.One frame, two halves, counted from CSmasterdeviceCS asserts — frameopensopcode, then addressbytesfiller during dummyphasedata byte 0data byte 1CS releases — framecloses
Figure 2 — the sequencer's view of a read. The master's states change at byte boundaries it counts internally; nothing on the bus marks them. The frame is opened once at the start and closed once at the end, with the direction of meaning reversing in the middle without any bus event to mark it.

6. Building the Master Read Sequencer — Three HDLs

The circuit

Circuit. A six-state machine above the byte engine, supplying a byte to transmit and deciding which received bytes are payload.

State. The phase, an address-byte index, a dummy-bit counter, and a payload-byte counter.

Datapath. tx_byte is presented one byte ahead of when it is needed — selected on the byte_done that completes the previous byte, so the engine always has the next value ready. Received bytes are forwarded only in ST_DATA.

Control. cs_n is asserted on start and released only in ST_CLOSE. That separation is the §2 requirement expressed structurally: there is no path from the data phase directly to idle that does not pass through a state whose only job is closing the frame.

Clock. The system clock. bit_stb and byte_done are strobes from the byte engine.

Reset. Asynchronous, active-low, to idle with CS deasserted — the safe state, since a slave must not see a frame opening out of reset.

Enables. Two different strobes advance the machine: byte_done in the byte-oriented phases and bit_stb in the dummy phase, because dummy is counted in cycles §3.

Timing. rx_valid is a single-cycle strobe accompanying each payload byte, and done a single pulse at completion. Neither is a level, so a consumer sees exactly one event per byte and one per transaction.

Synthesis. Three state bits, a small address index, a 5-bit dummy counter, an 8-bit payload counter, and a byte-wide multiplexer selecting the address byte. The multiplexer is the only non-trivial part and it scales with ADDR_BYTES.

Limitations. No FIFO, no clock generation, and a single outstanding transaction — all Module 13 concerns. n_data is a fixed count rather than a streaming interface.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_read_seq.sv — one frame that changes direction in the middle
// spi_read_seq.sv — the master-side read sequencer.
//
// A read is a write first. This FSM drives one CS frame that changes
// direction part-way through: it transmits an opcode and address, clocks a
// dummy phase, then collects data -- without ever releasing CS, because
// releasing it would end the transaction (Chapter 5.1 §4).
//
// Note what it supplies on MOSI during the dummy and data phases: filler.
// There is no "not sending" on SPI (Chapter 5.2 §3); the master drives every
// edge it clocks, and the filler value is a device-level agreement.
module spi_read_seq #(
    parameter int ADDR_BYTES = 3,
    parameter logic [7:0] FILLER = 8'h00
) (
    input  logic                     clk,
    input  logic                     rst_n,
    input  logic                     start,         // pulse: begin a read
    input  logic [7:0]               opcode,
    input  logic [8*ADDR_BYTES-1:0]  addr,
    input  logic [4:0]               dummy_cycles,
    input  logic [7:0]               n_data,        // bytes to collect

    // --- interface to the byte engine (Chapter 4.2) ---
    input  logic                     bit_stb,       // one pulse per SPI bit
    input  logic                     byte_done,     // one pulse per byte
    input  logic [7:0]               rx_byte,
    output logic [7:0]               tx_byte,       // what to shift out next

    // --- bus + result ---
    output logic                     cs_n,
    output logic                     rx_valid,      // pulse: rx_data is a payload byte
    output logic [7:0]               rx_data,
    output logic                     busy,
    output logic                     done           // pulse: transaction complete
);
    typedef enum logic [2:0] {
        ST_IDLE, ST_CMD, ST_ADDR, ST_DUMMY, ST_DATA, ST_CLOSE
    } state_t;

    state_t state;
    localparam int ACW = (ADDR_BYTES > 1) ? $clog2(ADDR_BYTES) : 1;
    logic [ACW-1:0] addr_cnt;
    logic [4:0]     dummy_cnt;
    logic [7:0]     data_cnt;

    // The address byte currently being presented, most-significant first.
    function automatic logic [7:0] addr_byte(input logic [ACW-1:0] idx);
        return addr[8*(ADDR_BYTES-1-idx) +: 8];
    endfunction

    always_ff @(posedge clk or negedge rst_n) begin
        if (!rst_n) begin
            state     <= ST_IDLE;
            cs_n      <= 1'b1;
            tx_byte   <= FILLER;
            addr_cnt  <= '0;
            dummy_cnt <= 5'd0;
            data_cnt  <= 8'd0;
            rx_valid  <= 1'b0;
            rx_data   <= 8'h00;
            done      <= 1'b0;
        end else begin
            rx_valid <= 1'b0;                  // strobes
            done     <= 1'b0;

            case (state)
                ST_IDLE: if (start) begin
                    cs_n      <= 1'b0;         // open the frame and keep it open
                    tx_byte   <= opcode;
                    addr_cnt  <= '0;
                    dummy_cnt <= 5'd0;
                    data_cnt  <= 8'd0;
                    state     <= ST_CMD;
                end

                ST_CMD: if (byte_done) begin
                    if (ADDR_BYTES > 0) begin
                        tx_byte  <= addr_byte('0);
                        addr_cnt <= '0;
                        state    <= ST_ADDR;
                    end else if (dummy_cycles != 5'd0) begin
                        tx_byte <= FILLER;
                        state   <= ST_DUMMY;
                    end else begin
                        tx_byte <= FILLER;
                        state   <= ST_DATA;
                    end
                end

                ST_ADDR: if (byte_done) begin
                    if (addr_cnt == ACW'(ADDR_BYTES - 1)) begin
                        // Address complete. Everything from here is filler on
                        // MOSI; the meaning is travelling the other way.
                        tx_byte <= FILLER;
                        if (dummy_cycles != 5'd0) state <= ST_DUMMY;
                        else                      state <= ST_DATA;
                    end else begin
                        addr_cnt <= addr_cnt + 1'b1;
                        tx_byte  <= addr_byte(addr_cnt + 1'b1);
                    end
                end

                // Dummy is counted in BITS, not bytes (Chapter 4.5 §3).
                ST_DUMMY: if (bit_stb) begin
                    if (dummy_cnt == dummy_cycles - 5'd1) state <= ST_DATA;
                    else dummy_cnt <= dummy_cnt + 5'd1;
                end

                ST_DATA: if (byte_done) begin
                    rx_data  <= rx_byte;
                    rx_valid <= 1'b1;
                    if (data_cnt == n_data - 8'd1) state <= ST_CLOSE;
                    else data_cnt <= data_cnt + 8'd1;
                end

                ST_CLOSE: begin
                    cs_n  <= 1'b1;             // only now is the frame ended
                    done  <= 1'b1;
                    state <= ST_IDLE;
                end

                default: state <= ST_IDLE;
            endcase
        end
    end

    assign busy = (state != ST_IDLE);
endmodule

The ST_CLOSE state deserves its own comment. It would be tempting to raise cs_n in ST_DATA on the final byte_done and return straight to idle. That works, and it couples the frame-closing decision to the payload-counting decision — so a later change to the counting logic can silently change when the frame ends. Keeping a state whose entire responsibility is close the frame makes the one-frame requirement of §2 a structural property rather than an emergent one.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_read_seq_tb.sv — a full read, a no-dummy read, and CS held across the turnaround
// spi_read_seq_tb.sv — a full read, a no-dummy read, and CS held across the
// direction change.
`timescale 1ns/1ps
module spi_read_seq_tb;
    logic clk = 0, rst_n = 0;
    always #5 clk = ~clk;

    localparam int AB = 3;
    logic start = 0, bit_stb = 0, byte_done = 0;
    logic [7:0] opcode = 8'h0B, rx_byte = 8'h00, n_data = 8'd2;
    logic [8*AB-1:0] addr = 24'h123456;
    logic [4:0] dummy_cycles = 5'd8;
    logic [7:0] tx_byte, rx_data;
    logic cs_n, rx_valid, busy, done;

    spi_read_seq #(.ADDR_BYTES(AB)) dut (
        .clk, .rst_n, .start, .opcode, .addr, .dummy_cycles, .n_data,
        .bit_stb, .byte_done, .rx_byte, .tx_byte,
        .cs_n, .rx_valid, .rx_data, .busy, .done);

    int errors = 0, cs_rises = 0;
    logic [7:0] tx_seen [$];
    logic [7:0] rx_got  [$];

    task automatic chk(input string what, input int got, input int exp);
        if (got !== exp) begin $display("FAIL %s: got 0x%0h exp 0x%0h", what, got, exp); errors++; end
    endtask

    logic cs_n_q;
    always @(posedge clk) if (rst_n) begin
        cs_n_q <= cs_n;
        if (cs_n && !cs_n_q) cs_rises++;
        if (rx_valid) rx_got.push_back(rx_data);
    end

    // Model one SPI bit; a byte is eight of them with byte_done on the last.
    task automatic one_bit(input logic last);
        @(negedge clk); bit_stb = 1; byte_done = last;
        @(negedge clk); bit_stb = 0; byte_done = 0;
        @(negedge clk);
    endtask

    task automatic one_byte(input logic [7:0] miso_val);
        tx_seen.push_back(tx_byte);       // what the master is presenting
        rx_byte = miso_val;
        for (int i = 0; i < 8; i++) one_bit(i == 7);
    endtask

    initial begin
        repeat (3) @(negedge clk); rst_n = 1; @(negedge clk);
        chk("idle: cs_n high", cs_n, 1);
        chk("idle: not busy",  busy, 0);

        // --- a full read: opcode, 3 address bytes, 8 dummy bits, 2 data ---
        start = 1; @(negedge clk); start = 0; @(negedge clk);
        chk("frame opened", cs_n, 0);
        chk("busy",         busy, 1);

        one_byte(8'h00);                  // opcode byte time
        one_byte(8'h00);                  // addr[23:16]
        one_byte(8'h00);                  // addr[15:8]
        one_byte(8'h00);                  // addr[7:0]
        chk("CS still low after address", cs_n, 0);   // the direction change
        repeat (8) one_bit(1'b0);         // dummy phase, counted in bits
        one_byte(8'hDE);                  // data byte 0
        one_byte(8'hAD);                  // data byte 1
        repeat (4) @(negedge clk);

        chk("frame closed",   cs_n,  1);
        chk("cs rose once",   cs_rises, 1);
        chk("two bytes back", rx_got.size(), 2);
        chk("data byte 0",    rx_got[0], 8'hDE);
        chk("data byte 1",    rx_got[1], 8'hAD);

        // The master must have presented opcode then address MSB-first, then
        // filler once the address was done.
        chk("tx opcode",   tx_seen[0], 8'h0B);
        chk("tx addr hi",  tx_seen[1], 8'h12);
        chk("tx addr mid", tx_seen[2], 8'h34);
        chk("tx addr lo",  tx_seen[3], 8'h56);
        chk("tx filler in data phase", tx_seen[4], 8'h00);

        // --- a read with NO dummy phase must go straight to data ---
        tx_seen.delete(); rx_got.delete();
        dummy_cycles = 5'd0; n_data = 8'd1; opcode = 8'h03;
        @(negedge clk);
        start = 1; @(negedge clk); start = 0; @(negedge clk);
        one_byte(8'h00); one_byte(8'h00); one_byte(8'h00); one_byte(8'h00);
        one_byte(8'hC3);                  // data arrives immediately
        repeat (4) @(negedge clk);
        chk("no-dummy: one byte back", rx_got.size(), 1);
        chk("no-dummy: value",         rx_got[0],     8'hC3);
        chk("no-dummy: frame closed",  cs_n,          1);
        chk("no-dummy: cs rose twice", cs_rises,      2);

        if (errors == 0)
            $display("PASS: one frame spans the direction change, the address is sent MSB-first, filler is driven once the request is complete, and the dummy phase is skipped when zero");
        else
            $display("FAILED with %0d error(s)", errors);
        $finish;
    end

    initial begin #500000; $display("FAIL: watchdog timeout"); $finish; end
endmodule

The check that matters most is chk("CS still low after address", cs_n, 0), performed immediately after the final address byte. That is §2 as an assertion: a sequencer that released CS at the end of its transmit half would still collect bytes afterwards and could still pass every data comparison in a testbench that did not look at CS.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_read_seq.v — the same sequencer in Verilog-2001
// spi_read_seq.v — the same master-side read sequencer in Verilog-2001.
module spi_read_seq #(
    parameter ADDR_BYTES = 3,
    parameter FILLER     = 8'h00
) (
    input  wire                     clk,
    input  wire                     rst_n,
    input  wire                     start,
    input  wire [7:0]               opcode,
    input  wire [8*ADDR_BYTES-1:0]  addr,
    input  wire [4:0]               dummy_cycles,
    input  wire [7:0]               n_data,

    input  wire                     bit_stb,
    input  wire                     byte_done,
    input  wire [7:0]               rx_byte,
    output reg  [7:0]               tx_byte,

    output reg                      cs_n,
    output reg                      rx_valid,
    output reg  [7:0]               rx_data,
    output wire                     busy,
    output reg                      done
);
    localparam ST_IDLE  = 3'd0,
               ST_CMD   = 3'd1,
               ST_ADDR  = 3'd2,
               ST_DUMMY = 3'd3,
               ST_DATA  = 3'd4,
               ST_CLOSE = 3'd5;

    function integer clogb2;
        input integer value;
        integer v;
        begin
            v = value - 1;
            for (clogb2 = 0; v > 0; clogb2 = clogb2 + 1) v = v >> 1;
        end
    endfunction

    localparam ACW = (ADDR_BYTES > 1) ? clogb2(ADDR_BYTES) : 1;

    reg [2:0]     state;
    reg [ACW-1:0] addr_cnt;
    reg [4:0]     dummy_cnt;
    reg [7:0]     data_cnt;

    // Address byte `idx`, most-significant first. Verilog-2001 has no
    // indexed part-select on a function result, so the shift is explicit.
    function [7:0] addr_byte;
        input integer idx;
        begin
            addr_byte = (addr >> (8 * (ADDR_BYTES - 1 - idx))) & 8'hFF;
        end
    endfunction

    always @(posedge clk or negedge rst_n) begin
        if (!rst_n) begin
            state     <= ST_IDLE;
            cs_n      <= 1'b1;
            tx_byte   <= FILLER;
            addr_cnt  <= {ACW{1'b0}};
            dummy_cnt <= 5'd0;
            data_cnt  <= 8'd0;
            rx_valid  <= 1'b0;
            rx_data   <= 8'h00;
            done      <= 1'b0;
        end else begin
            rx_valid <= 1'b0;
            done     <= 1'b0;

            case (state)
                ST_IDLE: if (start) begin
                    cs_n      <= 1'b0;
                    tx_byte   <= opcode;
                    addr_cnt  <= {ACW{1'b0}};
                    dummy_cnt <= 5'd0;
                    data_cnt  <= 8'd0;
                    state     <= ST_CMD;
                end

                ST_CMD: if (byte_done) begin
                    if (ADDR_BYTES > 0) begin
                        tx_byte  <= addr_byte(0);
                        addr_cnt <= {ACW{1'b0}};
                        state    <= ST_ADDR;
                    end else if (dummy_cycles != 5'd0) begin
                        tx_byte <= FILLER;
                        state   <= ST_DUMMY;
                    end else begin
                        tx_byte <= FILLER;
                        state   <= ST_DATA;
                    end
                end

                ST_ADDR: if (byte_done) begin
                    if (addr_cnt == (ADDR_BYTES - 1)) begin
                        tx_byte <= FILLER;
                        if (dummy_cycles != 5'd0) state <= ST_DUMMY;
                        else                      state <= ST_DATA;
                    end else begin
                        addr_cnt <= addr_cnt + 1'b1;
                        tx_byte  <= addr_byte(addr_cnt + 1'b1);
                    end
                end

                ST_DUMMY: if (bit_stb) begin
                    if (dummy_cnt == (dummy_cycles - 5'd1)) state <= ST_DATA;
                    else dummy_cnt <= dummy_cnt + 5'd1;
                end

                ST_DATA: if (byte_done) begin
                    rx_data  <= rx_byte;
                    rx_valid <= 1'b1;
                    if (data_cnt == (n_data - 8'd1)) state <= ST_CLOSE;
                    else data_cnt <= data_cnt + 8'd1;
                end

                ST_CLOSE: begin
                    cs_n  <= 1'b1;
                    done  <= 1'b1;
                    state <= ST_IDLE;
                end

                default: state <= ST_IDLE;
            endcase
        end
    end

    assign busy = (state != ST_IDLE);
endmodule

Note the address-byte selector. SystemVerilog's indexed part-select addr[8*i +: 8] has no Verilog-2001 equivalent, so the byte is extracted with an explicit shift and mask — the portable idiom, and worth recognising because it appears wherever parameterised field extraction is needed.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_read_seq_tb.v — the same checks in Verilog-2001
// spi_read_seq_tb.v — the same checks in Verilog-2001.
`timescale 1ns/1ps
module spi_read_seq_tb;
    reg clk = 0, rst_n = 0;
    always #5 clk = ~clk;

    parameter AB = 3;
    reg start = 0, bit_stb = 0, byte_done = 0;
    reg [7:0] opcode = 8'h0B, rx_byte = 8'h00, n_data = 8'd2;
    reg [8*AB-1:0] addr = 24'h123456;
    reg [4:0] dummy_cycles = 5'd8;
    wire [7:0] tx_byte, rx_data;
    wire cs_n, rx_valid, busy, done;

    spi_read_seq #(.ADDR_BYTES(AB)) dut (
        .clk(clk), .rst_n(rst_n), .start(start), .opcode(opcode), .addr(addr),
        .dummy_cycles(dummy_cycles), .n_data(n_data), .bit_stb(bit_stb),
        .byte_done(byte_done), .rx_byte(rx_byte), .tx_byte(tx_byte),
        .cs_n(cs_n), .rx_valid(rx_valid), .rx_data(rx_data), .busy(busy), .done(done));

    integer errors = 0, cs_rises = 0, tx_n = 0, rx_n = 0, i;
    reg [7:0] tx_seen [0:15];
    reg [7:0] rx_got  [0:15];
    reg cs_n_q;

    task chk;
        input [80*8-1:0] what;
        input [31:0] got, exp;
        begin
            if (got !== exp) begin
                $display("FAIL %0s: got 0x%0h exp 0x%0h", what, got, exp);
                errors = errors + 1;
            end
        end
    endtask

    always @(posedge clk) if (rst_n) begin
        cs_n_q <= cs_n;
        if (cs_n && !cs_n_q) cs_rises = cs_rises + 1;
        if (rx_valid) begin rx_got[rx_n] = rx_data; rx_n = rx_n + 1; end
    end

    task one_bit;
        input last;
        begin
            @(negedge clk); bit_stb = 1; byte_done = last;
            @(negedge clk); bit_stb = 0; byte_done = 0;
            @(negedge clk);
        end
    endtask

    task one_byte;
        input [7:0] miso_val;
        integer k;
        begin
            tx_seen[tx_n] = tx_byte; tx_n = tx_n + 1;
            rx_byte = miso_val;
            for (k = 0; k < 8; k = k + 1) one_bit(k == 7);
        end
    endtask

    initial begin
        repeat (3) @(negedge clk); rst_n = 1; @(negedge clk);
        chk("idle: cs_n high", cs_n, 1);
        chk("idle: not busy",  busy, 0);

        start = 1; @(negedge clk); start = 0; @(negedge clk);
        chk("frame opened", cs_n, 0);
        chk("busy",         busy, 1);

        one_byte(8'h00); one_byte(8'h00); one_byte(8'h00); one_byte(8'h00);
        chk("CS still low after address", cs_n, 0);
        repeat (8) one_bit(1'b0);
        one_byte(8'hDE);
        one_byte(8'hAD);
        repeat (4) @(negedge clk);

        chk("frame closed",   cs_n,     1);
        chk("cs rose once",   cs_rises, 1);
        chk("two bytes back", rx_n,     2);
        chk("data byte 0",    rx_got[0], 8'hDE);
        chk("data byte 1",    rx_got[1], 8'hAD);

        chk("tx opcode",   tx_seen[0], 8'h0B);
        chk("tx addr hi",  tx_seen[1], 8'h12);
        chk("tx addr mid", tx_seen[2], 8'h34);
        chk("tx addr lo",  tx_seen[3], 8'h56);
        chk("tx filler in data phase", tx_seen[4], 8'h00);

        tx_n = 0; rx_n = 0;
        dummy_cycles = 5'd0; n_data = 8'd1; opcode = 8'h03;
        @(negedge clk);
        start = 1; @(negedge clk); start = 0; @(negedge clk);
        one_byte(8'h00); one_byte(8'h00); one_byte(8'h00); one_byte(8'h00);
        one_byte(8'hC3);
        repeat (4) @(negedge clk);
        chk("no-dummy: one byte back", rx_n,      1);
        chk("no-dummy: value",         rx_got[0], 8'hC3);
        chk("no-dummy: frame closed",  cs_n,      1);
        chk("no-dummy: cs rose twice", cs_rises,  2);

        if (errors == 0)
            $display("PASS: one frame spans the direction change, the address is sent MSB-first, filler is driven once the request is complete, and the dummy phase is skipped when zero");
        else
            $display("FAILED with %0d error(s)", errors);
        $finish;
    end

    initial begin #500000; $display("FAIL: watchdog timeout"); $finish; end
endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_read_seq.vhd — the same sequencer in VHDL
-- spi_read_seq.vhd — the same master-side read sequencer in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

entity spi_read_seq is
    generic (
        ADDR_BYTES : positive               := 3;
        FILLER     : std_logic_vector(7 downto 0) := x"00"
    );
    port (
        clk          : in  std_logic;
        rst_n        : in  std_logic;
        start        : in  std_logic;
        opcode       : in  std_logic_vector(7 downto 0);
        addr         : in  std_logic_vector(8 * ADDR_BYTES - 1 downto 0);
        dummy_cycles : in  unsigned(4 downto 0);
        n_data       : in  unsigned(7 downto 0);

        bit_stb      : in  std_logic;
        byte_done    : in  std_logic;
        rx_byte      : in  std_logic_vector(7 downto 0);
        tx_byte      : out std_logic_vector(7 downto 0);

        cs_n         : out std_logic;
        rx_valid     : out std_logic;
        rx_data      : out std_logic_vector(7 downto 0);
        busy         : out std_logic;
        done         : out std_logic
    );
end entity spi_read_seq;

architecture rtl of spi_read_seq is
    type state_t is (ST_IDLE, ST_CMD, ST_ADDR, ST_DUMMY, ST_DATA, ST_CLOSE);
    signal state : state_t;

    signal addr_cnt  : integer range 0 to ADDR_BYTES - 1;
    signal dummy_cnt : unsigned(4 downto 0);
    signal data_cnt  : unsigned(7 downto 0);

    -- Address byte `idx`, most-significant first.
    function addr_byte (a : std_logic_vector; idx : integer) return std_logic_vector is
        variable hi : integer;
    begin
        hi := 8 * (ADDR_BYTES - 1 - idx) + 7;
        return a(hi downto hi - 7);
    end function;
begin

    process (clk, rst_n) is
    begin
        if rst_n = '0' then
            state     <= ST_IDLE;
            cs_n      <= '1';
            tx_byte   <= FILLER;
            addr_cnt  <= 0;
            dummy_cnt <= (others => '0');
            data_cnt  <= (others => '0');
            rx_valid  <= '0';
            rx_data   <= (others => '0');
            done      <= '0';
        elsif rising_edge(clk) then
            rx_valid <= '0';
            done     <= '0';

            case state is
                when ST_IDLE =>
                    if start = '1' then
                        cs_n      <= '0';       -- open the frame and keep it open
                        tx_byte   <= opcode;
                        addr_cnt  <= 0;
                        dummy_cnt <= (others => '0');
                        data_cnt  <= (others => '0');
                        state     <= ST_CMD;
                    end if;

                when ST_CMD =>
                    if byte_done = '1' then
                        tx_byte  <= addr_byte(addr, 0);
                        addr_cnt <= 0;
                        state    <= ST_ADDR;
                    end if;

                when ST_ADDR =>
                    if byte_done = '1' then
                        if addr_cnt = ADDR_BYTES - 1 then
                            -- Request complete: everything after this is
                            -- filler on MOSI.
                            tx_byte <= FILLER;
                            if dummy_cycles /= 0 then
                                state <= ST_DUMMY;
                            else
                                state <= ST_DATA;
                            end if;
                        else
                            addr_cnt <= addr_cnt + 1;
                            tx_byte  <= addr_byte(addr, addr_cnt + 1);
                        end if;
                    end if;

                -- Dummy is counted in BITS, not bytes.
                when ST_DUMMY =>
                    if bit_stb = '1' then
                        if dummy_cnt = dummy_cycles - 1 then
                            state <= ST_DATA;
                        else
                            dummy_cnt <= dummy_cnt + 1;
                        end if;
                    end if;

                when ST_DATA =>
                    if byte_done = '1' then
                        rx_data  <= rx_byte;
                        rx_valid <= '1';
                        if data_cnt = n_data - 1 then
                            state <= ST_CLOSE;
                        else
                            data_cnt <= data_cnt + 1;
                        end if;
                    end if;

                when ST_CLOSE =>
                    cs_n  <= '1';               -- only now is the frame ended
                    done  <= '1';
                    state <= ST_IDLE;
            end case;
        end if;
    end process;

    busy <= '0' when state = ST_IDLE else '1';

end architecture rtl;
Azvya Education Pvt. Ltd.VLSI Mentor
spi_read_seq_tb.vhd — the same checks in VHDL
-- spi_read_seq_tb.vhd — the same checks in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

entity spi_read_seq_tb is
end entity spi_read_seq_tb;

architecture tb of spi_read_seq_tb is
    constant AB : positive := 3;

    signal clk          : std_logic := '0';
    signal rst_n        : std_logic := '0';
    signal start        : std_logic := '0';
    signal bit_stb      : std_logic := '0';
    signal byte_done    : std_logic := '0';
    signal halt         : boolean   := false;

    signal opcode       : std_logic_vector(7 downto 0) := x"0B";
    signal addr         : std_logic_vector(8 * AB - 1 downto 0) := x"123456";
    signal dummy_cycles : unsigned(4 downto 0) := to_unsigned(8, 5);
    signal n_data       : unsigned(7 downto 0) := to_unsigned(2, 8);
    signal rx_byte      : std_logic_vector(7 downto 0) := (others => '0');

    signal tx_byte  : std_logic_vector(7 downto 0);
    signal cs_n     : std_logic;
    signal rx_valid : std_logic;
    signal rx_data  : std_logic_vector(7 downto 0);
    signal busy     : std_logic;
    signal done     : std_logic;

    signal errors   : natural := 0;
    signal cs_rises : natural := 0;

    type bytes_t is array (0 to 15) of std_logic_vector(7 downto 0);
    signal rx_got : bytes_t := (others => (others => '0'));
    signal rx_n   : natural := 0;
    signal rx_rst : std_logic := '0';
begin

    clk <= not clk after 5 ns when not halt else '0';

    dut : entity work.spi_read_seq
        generic map (ADDR_BYTES => AB)
        port map (clk => clk, rst_n => rst_n, start => start, opcode => opcode,
                  addr => addr, dummy_cycles => dummy_cycles, n_data => n_data,
                  bit_stb => bit_stb, byte_done => byte_done, rx_byte => rx_byte,
                  tx_byte => tx_byte, cs_n => cs_n, rx_valid => rx_valid,
                  rx_data => rx_data, busy => busy, done => done);

    -- One process owns rx_got/rx_n/cs_rises; stim only requests a rewind.
    collect : process (clk) is
        variable cs_n_q : std_logic := '1';
    begin
        if rising_edge(clk) then
            if rst_n = '1' then
                if cs_n = '1' and cs_n_q = '0' then
                    cs_rises <= cs_rises + 1;
                end if;
                if rx_rst = '1' then
                    rx_n <= 0;
                elsif rx_valid = '1' then
                    rx_got(rx_n) <= rx_data;
                    rx_n         <= rx_n + 1;
                end if;
            end if;
            cs_n_q := cs_n;
        end if;
    end process;

    stim : process is
        variable tx_seen : bytes_t := (others => (others => '0'));
        variable tx_n    : natural := 0;

        procedure chk_n (what : string; got, exp : natural) is
        begin
            if got /= exp then
                report "FAIL " & what & ": got " & integer'image(got)
                    & " exp " & integer'image(exp) severity error;
                errors <= errors + 1;
            end if;
        end procedure;

        procedure chk_v (what : string; got, exp : std_logic_vector) is
        begin
            if got /= exp then
                report "FAIL " & what & ": got 0x" & to_hstring(got)
                    & " exp 0x" & to_hstring(exp) severity error;
                errors <= errors + 1;
            end if;
        end procedure;

        procedure chk_b (what : string; got, exp : std_logic) is
        begin
            if got /= exp then
                report "FAIL " & what severity error;
                errors <= errors + 1;
            end if;
        end procedure;

        procedure one_bit (last : std_logic) is
        begin
            wait until falling_edge(clk); bit_stb <= '1'; byte_done <= last;
            wait until falling_edge(clk); bit_stb <= '0'; byte_done <= '0';
            wait until falling_edge(clk);
        end procedure;

        procedure one_byte (miso_val : std_logic_vector(7 downto 0)) is
        begin
            tx_seen(tx_n) := tx_byte; tx_n := tx_n + 1;
            rx_byte <= miso_val;
            for k in 0 to 7 loop
                if k = 7 then one_bit('1'); else one_bit('0'); end if;
            end loop;
        end procedure;
    begin
        for i in 0 to 2 loop
            wait until falling_edge(clk);
        end loop;
        rst_n <= '1';
        wait until falling_edge(clk);
        chk_b("idle: cs_n high", cs_n, '1');
        chk_b("idle: not busy",  busy, '0');

        start <= '1'; wait until falling_edge(clk); start <= '0';
        wait until falling_edge(clk);
        chk_b("frame opened", cs_n, '0');
        chk_b("busy",         busy, '1');

        one_byte(x"00"); one_byte(x"00"); one_byte(x"00"); one_byte(x"00");
        chk_b("CS still low after address", cs_n, '0');
        for i in 0 to 7 loop one_bit('0'); end loop;
        one_byte(x"DE");
        one_byte(x"AD");
        for i in 0 to 3 loop wait until falling_edge(clk); end loop;

        chk_b("frame closed",   cs_n,     '1');
        chk_n("cs rose once",   cs_rises, 1);
        chk_n("two bytes back", rx_n,     2);
        chk_v("data byte 0",    rx_got(0), x"DE");
        chk_v("data byte 1",    rx_got(1), x"AD");

        chk_v("tx opcode",   tx_seen(0), x"0B");
        chk_v("tx addr hi",  tx_seen(1), x"12");
        chk_v("tx addr mid", tx_seen(2), x"34");
        chk_v("tx addr lo",  tx_seen(3), x"56");
        chk_v("tx filler in data phase", tx_seen(4), x"00");

        -- A read with no dummy phase must go straight to data.
        tx_n := 0;
        rx_rst <= '1'; wait until falling_edge(clk); rx_rst <= '0';
        dummy_cycles <= to_unsigned(0, 5);
        n_data       <= to_unsigned(1, 8);
        opcode       <= x"03";
        wait until falling_edge(clk);
        start <= '1'; wait until falling_edge(clk); start <= '0';
        wait until falling_edge(clk);
        one_byte(x"00"); one_byte(x"00"); one_byte(x"00"); one_byte(x"00");
        one_byte(x"C3");
        for i in 0 to 3 loop wait until falling_edge(clk); end loop;
        chk_n("no-dummy: one byte back", rx_n,      1);
        chk_v("no-dummy: value",         rx_got(0), x"C3");
        chk_b("no-dummy: frame closed",  cs_n,      '1');
        chk_n("no-dummy: cs rose twice", cs_rises,  2);

        if errors = 0 then
            report "PASS: one frame spans the direction change, the address is sent "
                 & "MSB-first, filler is driven once the request is complete, and the "
                 & "dummy phase is skipped when zero" severity note;
        else
            report "FAILED with " & integer'image(errors) & " error(s)" severity error;
        end if;
        halt <= true;
        wait;
    end process;

    watchdog : process is
    begin
        wait for 500 us;
        if not halt then
            report "FAIL: watchdog timeout" severity failure;
        end if;
        wait;
    end process;

end architecture tb;

Parity

All three implement the same machine: identical ports, asynchronous active-low reset to idle with CS high, CS asserted on start and released only in the closing state, the address presented most-significant byte first, filler driven once the request completes, a dummy phase counted in bits and skipped when zero, payload bytes forwarded with a single-cycle valid strobe, and a single done pulse.

One deliberate difference: the Verilog dialects guard the address phase with if (ADDR_BYTES > 0) while the VHDL does not, because VHDL's positive generic type makes a zero value impossible at elaboration. The behaviour is identical for every legal parameter value; the type system is simply doing the check in one language and not the other.

7. Why a Verification Engineer Cares

The one-frame property is the one to assert, because violating it produces silent data corruption rather than an error:

Azvya Education Pvt. Ltd.VLSI Mentor
spi_read_seq.sva — the frame spans the turnaround
   // 1. THE property of this chapter. CS must not rise between the start of
   //    the transaction and its completion. A two-frame read looks like two
   //    unrelated transactions to the device (§2).
   property p_single_frame;
       @(posedge clk) disable iff (!rst_n)
           $fell(cs_n) |-> (!cs_n throughout done[->1]);
   endproperty
   a_single_frame : assert property (p_single_frame)
       else $error("CS rose before the transaction completed -- request and response split across frames");

   // 2. Payload bytes are forwarded ONLY from the data phase. Forwarding an
   //    address-phase byte would hand the scoreboard the device's pre-response
   //    garbage as though it were data.
   a_rx_only_in_data : assert property (
       @(posedge clk) disable iff (!rst_n) rx_valid |-> (state == ST_DATA))
       else $error("payload byte forwarded from outside the data phase");

   // 3. Exactly n_data payload bytes per transaction -- no more, no fewer.
   property p_payload_count;
       @(posedge clk) disable iff (!rst_n)
           $rose(busy) |-> (rx_valid[->n_data] ##1 done[->1]);
   endproperty

   // 4. The master must drive the agreed filler once its half is done.
   //    Some devices interpret MOSI during the response (Chapter 5.2 §3).
   a_filler_in_response : assert property (
       @(posedge clk) disable iff (!rst_n)
           (state inside {ST_DUMMY, ST_DATA}) |-> (tx_byte == FILLER))
       else $error("master drove a non-filler byte during the response");

What these prove. That the sequencer holds one frame across the turnaround, forwards only payload, and drives the agreed filler. What they do not prove is that the device expects a single frame — a device with a latched-result protocol wants two, and asserting p_single_frame against it would be asserting the wrong specification. That mapping is a datasheet fact, and it is Chapter 4.1 §7's boundary once more.

Coverage should target the sequencer's structural axes:

Azvya Education Pvt. Ltd.VLSI Mentor
spi_read_seq_cg.sv — the request shapes a master must produce
   covergroup spi_read_seq_cg @(posedge done);
       cp_addr_bytes : coverpoint cfg.addr_bytes {
           bins none  = {0};          // status-style command, no address
           bins one   = {1};
           bins three = {3};          // the common memory case
           bins four  = {4};          // large-capacity addressing
       }

       // §2 and Chapter 4.5: zero dummy takes a different path through the
       // FSM than any non-zero value, and it is a distinct branch.
       cp_dummy : coverpoint cfg.dummy_cycles {
           bins none      = {0};      // MUST be hit -- the skip path
           bins sub_byte  = {[1:7]};
           bins one_byte  = {8};
           bins many      = {[9:31]};
       }

       cp_payload : coverpoint cfg.n_data {
           bins one   = {1};          // never crosses a byte boundary
           bins two   = {2};          // the smallest burst
           bins burst = {[3:255]};
       }

       x_addr_dummy : cross cp_addr_bytes, cp_dummy;
   endgroup

The cp_dummy.none bin is the one to insist on, for the same reason as Chapter 4.5 §8: zero dummy cycles takes a different branch — straight from the address phase to the data phase — and a suite that always configures a non-zero count never executes it.

8. Why an FPGA or ASIC Engineer Cares

tx_byte must be ready before the engine needs it. The sequencer selects the next transmit byte on the byte_done that finishes the previous one, giving the byte engine a full byte time to consume it. Selecting it when the engine asks would put the address multiplexer in the critical path between a strobe and the first launch edge — which at high SCLK is exactly where you do not want a wide multiplexer.

The address multiplexer grows with address length. Four address bytes means a 4:1 byte-wide mux. That is small, but it is combinational logic selected by a counter, and on a design supporting both 3- and 4-byte addressing at runtime the select becomes a register output rather than a constant. Same trade as Chapter 4.2 §8.

Reset must leave CS deasserted. A sequencer resetting with cs_n low would open a frame the instant reset released, and any slave on the bus would begin decoding whatever the clock did next. This is the same reasoning as Chapter 5.2 §6's synchroniser resetting high, applied at the other end of the link.

Think about what happens if the fabric stalls. This sequencer assumes the byte engine keeps clocking. If the transmit data came from a FIFO that underran mid-frame, the master would have to either stall SCLK — legal, since SPI has no timeout, and the reason SPI tolerates a paused clock at all — or keep clocking with stale data. Stalling is correct and it is worth knowing the bus permits it; many first designs do the other thing by accident.

9. Failure Signature — Every Read Returns the Previous Read's Data

Symptom. Reads return plausible values that are consistently one transaction stale: read register A, get whatever register B returned last time; read A again immediately and now it is correct. Writes work. The pattern is perfectly reproducible and unaffected by clock rate.

What "one transaction stale" tells you. The data is real and correctly transferred — it is simply the wrong data, arriving intact. That rules out mode, bit order, framing and signal integrity in one step, since all of those corrupt values rather than displacing them in time (Chapter 4.3 §8's reasoning applied to a different axis).

Plausible mechanisms.

  • The device uses a latched-result protocol: the transaction that issues a command returns the previous command's result, and the driver is treating it as a one-frame read. This is the §2 exception, and it is the leading candidate.
  • The master is capturing one byte early — collecting a byte from the dummy phase as payload, so every byte shifts by one position and the last is stale.
  • The driver reads a buffer that the previous transfer filled, a software bug with no SPI cause at all.
  • The device requires a dummy read after a command, and the first result is discarded by design.

The discriminating observation. Issue the same read twice in succession and compare. If the second is correct while the first is stale, the device is one transaction behind and you have the latched-result protocol — the fix is to read twice and discard the first, or to restructure as command-then-separate-read, which for this device class is the specified arrangement rather than a workaround.

If both reads are stale by one, the pipeline is in the master or the driver, not the device.

That test costs one extra transaction and separates a device protocol from a master bug, which are otherwise indistinguishable from the symptom.

Why the investigation goes wrong. Because stale-but-valid data does not look like a protocol problem — it looks like a caching or buffering bug, and people search the driver. The SPI-level cause is a device whose read semantics are genuinely two-frame, which contradicts the reasonable default of §2. The default is right often enough that the exception is rarely suspected.

10. Common Misconceptions

11. Reason It Through

Work this before reading the answer.

A driver polls a sensor's status register in a tight loop, then reads a 4-byte sample when the ready bit sets. Each status poll is an 8-bit command plus an 8-bit response — 16 clock cycles — and each sample read is a command, a 2-byte address, 8 dummy cycles and 4 data bytes.

At 10 MHz the system reads 8,000 samples per second and the CPU is mostly idle waiting. The team proposes raising SCLK to 40 MHz to get 32,000 samples per second. Will they?

Work out where the time actually goes. A sample read is 8 + 16 + 8 + 32 = 64 cycles. At 10 MHz that is 6.4 µs. At 8,000 samples per second the sample reads occupy 8000 × 6.4 µs = 51 ms of each second — about 5% of the bus.

So what is the other 95%? The polling loop. Each poll is 16 cycles, 1.6 µs at 10 MHz, and the loop runs continuously between samples. In the ~119 µs between samples the driver issues roughly 74 status polls per sample.

Now apply the proposed change. Quadrupling SCLK makes every transaction four times faster — including the polls. The sample read drops to 1.6 µs and the polls to 400 ns. But the sample rate is set by the sensor, not by the bus: it produces 8,000 samples per second because that is its conversion rate. Going faster on the bus means the driver simply polls more times per sample, burning the same proportion of bus time on the same useless question.

So the answer is no, and the reason is that the bus was never the bottleneck. The team measured "the CPU is idle waiting" and concluded the transfer was slow, when the waiting was for the sensor.

What would actually help? Stop polling. If the device has a data-ready interrupt line — many do, and it is the single most valuable pin on a sensor — the polls disappear entirely and the bus goes idle between samples. Failing that, poll on a timer sized to the conversion time rather than in a tight loop, which converts 74 polls per sample into one or two.

And if the sample rate itself must rise? That is a sensor configuration question, and only then does the bus rate matter — because at 32,000 samples per second the sample reads alone would need 32000 × 6.4 µs = 205 ms, still only 20% of the bus at 10 MHz. The bus has ample headroom either way.

The general lesson. Section 4's efficiency arithmetic answers "how much of this transaction is overhead". It does not answer "is the bus the bottleneck" — and on a polled sensor the answer is usually no. Measure which transactions dominate before optimising the clock.

12. Understanding Check

13. Summary

Every SPI read is a write first: the master transmits an opcode and usually an address before anything comes back. A read is therefore half-duplex traffic on a full-duplex bus, and what changes at the turnaround is not the direction of transmission — both directions are always active — but which direction the master pays attention to.

The request and the response must normally occupy one CS frame, because CS deassertion ends the transaction and discards the request. The exceptions are real and are datasheet facts: latched-result converters, continuous-read modes, and some status registers.

Throughout the response the master drives filler on MOSI, since every clocked edge transmits something. The value is a device agreement and belongs in configuration.

The request is overhead, and for a 1-byte opcode, 3-byte address and 8 dummy cycles it costs 40 cycles before the first payload bit — making a single-byte read about 17% efficient and a 256-byte burst about 98%. Raising SCLK does not change that ratio, because it scales request and payload equally; fewer, larger transactions is the lever that does.

In RTL the master side is a six-state machine that asserts CS on start and releases it only in a dedicated closing state — which is what makes the one-frame property structural rather than emergent — presents the next transmit byte a full byte time ahead, counts the dummy phase in bits rather than bytes, and forwards received bytes only from the data phase.

For verification the property worth asserting is that CS does not rise before the transaction completes, remembering that a latched-result device genuinely wants the opposite. And when reads come back one transaction stale, the data is intact and merely displaced in time — issue the same read twice to separate a device protocol from a master pipeline bug.

14. What Comes Next

This chapter established when the device's data arrives. It said nothing about when that data may be believed — and on a real board those are different instants, separated by the device's output delay, the flight time down the trace, and the master's own sampling choice. Chapter 6.3 — MISO Valid Timing makes that window explicit, turns Module 2's round-trip budget into a transaction-level concern, and builds the master's configurable input sampler in all three HDLs.

Continue learning