SPI · Module 6
Command-Then-Read Sequences
Why every SPI read is a write first, why request and response must share one CS frame, what the master drives once its half is done, what the request costs in bus time, and the master read sequencer in three HDLs.
Chapter 6.1 followed a read from the device's side: fetch, load, drive, release. This chapter takes the master's side, and it begins with an observation that sounds trivial and is not.
Before a device can return anything, it has to be told what to return. So a read begins as a transmission — and the master must change what it is doing part-way through a frame it is not allowed to interrupt.
That constraint shapes the master's hardware, its bus efficiency, and one of the more confusing failure modes in SPI.
1. Every Read Is a Write First
On a bus with separate command and response channels, a read is one operation. On SPI it is two halves of one frame:
The request half. The master transmits an opcode and usually an address on MOSI. During this, MISO carries nothing meaningful — the device does not yet know what is being asked (Chapter 6.1 §4).
The response half. The device drives MISO. During this, MOSI carries nothing meaningful — the master has already said everything it needs to.
So an SPI read is half-duplex traffic carried on a full-duplex bus. Both directions are physically active the whole time; only one of them means anything at any moment. Chapter 1.4 established that the bus cannot do otherwise — every edge shifts both registers — and this is where that fact starts costing something.
2. One Frame, Not Two
A natural question: why not send the command in one transaction, then read the data in another? It would simplify the driver considerably.
For most devices it does not work, and Chapter 5.1 §4 already supplies the reason. CS deassertion ends the transaction. A device that sees CS rise after the address concludes the transaction is over, discards its state, and resynchronises its sequencer (Chapter 4.4 §7). When CS falls again it expects a new command, so the master's read clocks are interpreted as an opcode followed by address bytes — decoding whatever the filler happens to be.
That is why the sequencer in §6 holds cs_n low from ST_CMD all the way through ST_DATA, and why ST_CLOSE exists as a separate state: the frame must not end until the last data byte has been collected.
There are real exceptions, and knowing which is which is a datasheet question:
- Devices with a latched result. Some ADCs and sensors perform a conversion on one transaction and return it on the next. There the two-frame pattern is not a workaround but the specified protocol.
- Status registers. A one-byte command whose response begins immediately often tolerates either arrangement.
- Devices with an explicit continuous-read mode, which keep a pointer across frames deliberately.
The general rule stands: assume one frame unless the datasheet says otherwise, because the failure when you are wrong is silent and looks like data corruption rather than a framing error.
3. What the Master Drives During the Response
The request, the turnaround, and the response — one frame
8 cyclesThe 0x00 bytes on MOSI during the response are the visible form of Chapter 5.2 §3: there is no "not sending". The master must drive every edge it clocks, so it supplies filler. Which value is a device-level agreement — 0x00 and 0xFF are both common, and a few devices specify one or interpret particular values as a subsequent command.
The sequencer in §6 makes that a parameter (FILLER) rather than hard-coding zero, because getting it wrong on a device that cares produces a fault that looks nothing like a filler problem.
4. What the Request Costs
The request is pure overhead from the payload's point of view, and quantifying it explains a great deal about how SPI devices are used.
For a read with a 1-byte opcode, a 3-byte address and 8 dummy cycles:
opcode 8 cycles
address 24 cycles
dummy 8 cycles
──────────────────────
request 40 cycles before the first payload bitEfficiency therefore depends entirely on how much data follows:
payload total cycles efficiency
1 byte 40 + 8 = 48 16.7 %
4 bytes 40 + 32 = 72 44.4 %
32 bytes 40 + 256 = 296 86.5 %
256 bytes 40 + 2048 = 2088 98.1 %Two consequences worth holding.
Small reads are dominated by the request. Reading one byte spends five sixths of the bus restating what you want. Reading a single status register — the most common operation in a polling loop — is about as inefficient as SPI gets, which is why devices that expect to be polled often place the status byte behind a one-byte command with no address and no dummy phase, bringing the request down to 8 cycles.
Raising the clock does not help as much as it appears to. Doubling SCLK halves the wall-clock time of both the payload and the request, so the ratio is unchanged — and Chapter 4.5 §3 showed the dummy count may increase with frequency. For small reads, reducing the number of transactions beats raising the clock.
5. The Transaction
6. Building the Master Read Sequencer — Three HDLs
The circuit
Circuit. A six-state machine above the byte engine, supplying a byte to transmit and deciding which received bytes are payload.
State. The phase, an address-byte index, a dummy-bit counter, and a payload-byte counter.
Datapath. tx_byte is presented one byte ahead of when it is needed — selected on the byte_done that completes the previous byte, so the engine always has the next value ready. Received bytes are forwarded only in ST_DATA.
Control. cs_n is asserted on start and released only in ST_CLOSE. That separation is the §2 requirement expressed structurally: there is no path from the data phase directly to idle that does not pass through a state whose only job is closing the frame.
Clock. The system clock. bit_stb and byte_done are strobes from the byte engine.
Reset. Asynchronous, active-low, to idle with CS deasserted — the safe state, since a slave must not see a frame opening out of reset.
Enables. Two different strobes advance the machine: byte_done in the byte-oriented phases and bit_stb in the dummy phase, because dummy is counted in cycles §3.
Timing. rx_valid is a single-cycle strobe accompanying each payload byte, and done a single pulse at completion. Neither is a level, so a consumer sees exactly one event per byte and one per transaction.
Synthesis. Three state bits, a small address index, a 5-bit dummy counter, an 8-bit payload counter, and a byte-wide multiplexer selecting the address byte. The multiplexer is the only non-trivial part and it scales with ADDR_BYTES.
Limitations. No FIFO, no clock generation, and a single outstanding transaction — all Module 13 concerns. n_data is a fixed count rather than a streaming interface.
// spi_read_seq.sv — the master-side read sequencer.
//
// A read is a write first. This FSM drives one CS frame that changes
// direction part-way through: it transmits an opcode and address, clocks a
// dummy phase, then collects data -- without ever releasing CS, because
// releasing it would end the transaction (Chapter 5.1 §4).
//
// Note what it supplies on MOSI during the dummy and data phases: filler.
// There is no "not sending" on SPI (Chapter 5.2 §3); the master drives every
// edge it clocks, and the filler value is a device-level agreement.
module spi_read_seq #(
parameter int ADDR_BYTES = 3,
parameter logic [7:0] FILLER = 8'h00
) (
input logic clk,
input logic rst_n,
input logic start, // pulse: begin a read
input logic [7:0] opcode,
input logic [8*ADDR_BYTES-1:0] addr,
input logic [4:0] dummy_cycles,
input logic [7:0] n_data, // bytes to collect
// --- interface to the byte engine (Chapter 4.2) ---
input logic bit_stb, // one pulse per SPI bit
input logic byte_done, // one pulse per byte
input logic [7:0] rx_byte,
output logic [7:0] tx_byte, // what to shift out next
// --- bus + result ---
output logic cs_n,
output logic rx_valid, // pulse: rx_data is a payload byte
output logic [7:0] rx_data,
output logic busy,
output logic done // pulse: transaction complete
);
typedef enum logic [2:0] {
ST_IDLE, ST_CMD, ST_ADDR, ST_DUMMY, ST_DATA, ST_CLOSE
} state_t;
state_t state;
localparam int ACW = (ADDR_BYTES > 1) ? $clog2(ADDR_BYTES) : 1;
logic [ACW-1:0] addr_cnt;
logic [4:0] dummy_cnt;
logic [7:0] data_cnt;
// The address byte currently being presented, most-significant first.
function automatic logic [7:0] addr_byte(input logic [ACW-1:0] idx);
return addr[8*(ADDR_BYTES-1-idx) +: 8];
endfunction
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
state <= ST_IDLE;
cs_n <= 1'b1;
tx_byte <= FILLER;
addr_cnt <= '0;
dummy_cnt <= 5'd0;
data_cnt <= 8'd0;
rx_valid <= 1'b0;
rx_data <= 8'h00;
done <= 1'b0;
end else begin
rx_valid <= 1'b0; // strobes
done <= 1'b0;
case (state)
ST_IDLE: if (start) begin
cs_n <= 1'b0; // open the frame and keep it open
tx_byte <= opcode;
addr_cnt <= '0;
dummy_cnt <= 5'd0;
data_cnt <= 8'd0;
state <= ST_CMD;
end
ST_CMD: if (byte_done) begin
if (ADDR_BYTES > 0) begin
tx_byte <= addr_byte('0);
addr_cnt <= '0;
state <= ST_ADDR;
end else if (dummy_cycles != 5'd0) begin
tx_byte <= FILLER;
state <= ST_DUMMY;
end else begin
tx_byte <= FILLER;
state <= ST_DATA;
end
end
ST_ADDR: if (byte_done) begin
if (addr_cnt == ACW'(ADDR_BYTES - 1)) begin
// Address complete. Everything from here is filler on
// MOSI; the meaning is travelling the other way.
tx_byte <= FILLER;
if (dummy_cycles != 5'd0) state <= ST_DUMMY;
else state <= ST_DATA;
end else begin
addr_cnt <= addr_cnt + 1'b1;
tx_byte <= addr_byte(addr_cnt + 1'b1);
end
end
// Dummy is counted in BITS, not bytes (Chapter 4.5 §3).
ST_DUMMY: if (bit_stb) begin
if (dummy_cnt == dummy_cycles - 5'd1) state <= ST_DATA;
else dummy_cnt <= dummy_cnt + 5'd1;
end
ST_DATA: if (byte_done) begin
rx_data <= rx_byte;
rx_valid <= 1'b1;
if (data_cnt == n_data - 8'd1) state <= ST_CLOSE;
else data_cnt <= data_cnt + 8'd1;
end
ST_CLOSE: begin
cs_n <= 1'b1; // only now is the frame ended
done <= 1'b1;
state <= ST_IDLE;
end
default: state <= ST_IDLE;
endcase
end
end
assign busy = (state != ST_IDLE);
endmoduleThe ST_CLOSE state deserves its own comment. It would be tempting to raise cs_n in ST_DATA on the final byte_done and return straight to idle. That works, and it couples the frame-closing decision to the payload-counting decision — so a later change to the counting logic can silently change when the frame ends. Keeping a state whose entire responsibility is close the frame makes the one-frame requirement of §2 a structural property rather than an emergent one.
// spi_read_seq_tb.sv — a full read, a no-dummy read, and CS held across the
// direction change.
`timescale 1ns/1ps
module spi_read_seq_tb;
logic clk = 0, rst_n = 0;
always #5 clk = ~clk;
localparam int AB = 3;
logic start = 0, bit_stb = 0, byte_done = 0;
logic [7:0] opcode = 8'h0B, rx_byte = 8'h00, n_data = 8'd2;
logic [8*AB-1:0] addr = 24'h123456;
logic [4:0] dummy_cycles = 5'd8;
logic [7:0] tx_byte, rx_data;
logic cs_n, rx_valid, busy, done;
spi_read_seq #(.ADDR_BYTES(AB)) dut (
.clk, .rst_n, .start, .opcode, .addr, .dummy_cycles, .n_data,
.bit_stb, .byte_done, .rx_byte, .tx_byte,
.cs_n, .rx_valid, .rx_data, .busy, .done);
int errors = 0, cs_rises = 0;
logic [7:0] tx_seen [$];
logic [7:0] rx_got [$];
task automatic chk(input string what, input int got, input int exp);
if (got !== exp) begin $display("FAIL %s: got 0x%0h exp 0x%0h", what, got, exp); errors++; end
endtask
logic cs_n_q;
always @(posedge clk) if (rst_n) begin
cs_n_q <= cs_n;
if (cs_n && !cs_n_q) cs_rises++;
if (rx_valid) rx_got.push_back(rx_data);
end
// Model one SPI bit; a byte is eight of them with byte_done on the last.
task automatic one_bit(input logic last);
@(negedge clk); bit_stb = 1; byte_done = last;
@(negedge clk); bit_stb = 0; byte_done = 0;
@(negedge clk);
endtask
task automatic one_byte(input logic [7:0] miso_val);
tx_seen.push_back(tx_byte); // what the master is presenting
rx_byte = miso_val;
for (int i = 0; i < 8; i++) one_bit(i == 7);
endtask
initial begin
repeat (3) @(negedge clk); rst_n = 1; @(negedge clk);
chk("idle: cs_n high", cs_n, 1);
chk("idle: not busy", busy, 0);
// --- a full read: opcode, 3 address bytes, 8 dummy bits, 2 data ---
start = 1; @(negedge clk); start = 0; @(negedge clk);
chk("frame opened", cs_n, 0);
chk("busy", busy, 1);
one_byte(8'h00); // opcode byte time
one_byte(8'h00); // addr[23:16]
one_byte(8'h00); // addr[15:8]
one_byte(8'h00); // addr[7:0]
chk("CS still low after address", cs_n, 0); // the direction change
repeat (8) one_bit(1'b0); // dummy phase, counted in bits
one_byte(8'hDE); // data byte 0
one_byte(8'hAD); // data byte 1
repeat (4) @(negedge clk);
chk("frame closed", cs_n, 1);
chk("cs rose once", cs_rises, 1);
chk("two bytes back", rx_got.size(), 2);
chk("data byte 0", rx_got[0], 8'hDE);
chk("data byte 1", rx_got[1], 8'hAD);
// The master must have presented opcode then address MSB-first, then
// filler once the address was done.
chk("tx opcode", tx_seen[0], 8'h0B);
chk("tx addr hi", tx_seen[1], 8'h12);
chk("tx addr mid", tx_seen[2], 8'h34);
chk("tx addr lo", tx_seen[3], 8'h56);
chk("tx filler in data phase", tx_seen[4], 8'h00);
// --- a read with NO dummy phase must go straight to data ---
tx_seen.delete(); rx_got.delete();
dummy_cycles = 5'd0; n_data = 8'd1; opcode = 8'h03;
@(negedge clk);
start = 1; @(negedge clk); start = 0; @(negedge clk);
one_byte(8'h00); one_byte(8'h00); one_byte(8'h00); one_byte(8'h00);
one_byte(8'hC3); // data arrives immediately
repeat (4) @(negedge clk);
chk("no-dummy: one byte back", rx_got.size(), 1);
chk("no-dummy: value", rx_got[0], 8'hC3);
chk("no-dummy: frame closed", cs_n, 1);
chk("no-dummy: cs rose twice", cs_rises, 2);
if (errors == 0)
$display("PASS: one frame spans the direction change, the address is sent MSB-first, filler is driven once the request is complete, and the dummy phase is skipped when zero");
else
$display("FAILED with %0d error(s)", errors);
$finish;
end
initial begin #500000; $display("FAIL: watchdog timeout"); $finish; end
endmoduleThe check that matters most is chk("CS still low after address", cs_n, 0), performed immediately after the final address byte. That is §2 as an assertion: a sequencer that released CS at the end of its transmit half would still collect bytes afterwards and could still pass every data comparison in a testbench that did not look at CS.
// spi_read_seq.v — the same master-side read sequencer in Verilog-2001.
module spi_read_seq #(
parameter ADDR_BYTES = 3,
parameter FILLER = 8'h00
) (
input wire clk,
input wire rst_n,
input wire start,
input wire [7:0] opcode,
input wire [8*ADDR_BYTES-1:0] addr,
input wire [4:0] dummy_cycles,
input wire [7:0] n_data,
input wire bit_stb,
input wire byte_done,
input wire [7:0] rx_byte,
output reg [7:0] tx_byte,
output reg cs_n,
output reg rx_valid,
output reg [7:0] rx_data,
output wire busy,
output reg done
);
localparam ST_IDLE = 3'd0,
ST_CMD = 3'd1,
ST_ADDR = 3'd2,
ST_DUMMY = 3'd3,
ST_DATA = 3'd4,
ST_CLOSE = 3'd5;
function integer clogb2;
input integer value;
integer v;
begin
v = value - 1;
for (clogb2 = 0; v > 0; clogb2 = clogb2 + 1) v = v >> 1;
end
endfunction
localparam ACW = (ADDR_BYTES > 1) ? clogb2(ADDR_BYTES) : 1;
reg [2:0] state;
reg [ACW-1:0] addr_cnt;
reg [4:0] dummy_cnt;
reg [7:0] data_cnt;
// Address byte `idx`, most-significant first. Verilog-2001 has no
// indexed part-select on a function result, so the shift is explicit.
function [7:0] addr_byte;
input integer idx;
begin
addr_byte = (addr >> (8 * (ADDR_BYTES - 1 - idx))) & 8'hFF;
end
endfunction
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
state <= ST_IDLE;
cs_n <= 1'b1;
tx_byte <= FILLER;
addr_cnt <= {ACW{1'b0}};
dummy_cnt <= 5'd0;
data_cnt <= 8'd0;
rx_valid <= 1'b0;
rx_data <= 8'h00;
done <= 1'b0;
end else begin
rx_valid <= 1'b0;
done <= 1'b0;
case (state)
ST_IDLE: if (start) begin
cs_n <= 1'b0;
tx_byte <= opcode;
addr_cnt <= {ACW{1'b0}};
dummy_cnt <= 5'd0;
data_cnt <= 8'd0;
state <= ST_CMD;
end
ST_CMD: if (byte_done) begin
if (ADDR_BYTES > 0) begin
tx_byte <= addr_byte(0);
addr_cnt <= {ACW{1'b0}};
state <= ST_ADDR;
end else if (dummy_cycles != 5'd0) begin
tx_byte <= FILLER;
state <= ST_DUMMY;
end else begin
tx_byte <= FILLER;
state <= ST_DATA;
end
end
ST_ADDR: if (byte_done) begin
if (addr_cnt == (ADDR_BYTES - 1)) begin
tx_byte <= FILLER;
if (dummy_cycles != 5'd0) state <= ST_DUMMY;
else state <= ST_DATA;
end else begin
addr_cnt <= addr_cnt + 1'b1;
tx_byte <= addr_byte(addr_cnt + 1'b1);
end
end
ST_DUMMY: if (bit_stb) begin
if (dummy_cnt == (dummy_cycles - 5'd1)) state <= ST_DATA;
else dummy_cnt <= dummy_cnt + 5'd1;
end
ST_DATA: if (byte_done) begin
rx_data <= rx_byte;
rx_valid <= 1'b1;
if (data_cnt == (n_data - 8'd1)) state <= ST_CLOSE;
else data_cnt <= data_cnt + 8'd1;
end
ST_CLOSE: begin
cs_n <= 1'b1;
done <= 1'b1;
state <= ST_IDLE;
end
default: state <= ST_IDLE;
endcase
end
end
assign busy = (state != ST_IDLE);
endmoduleNote the address-byte selector. SystemVerilog's indexed part-select addr[8*i +: 8] has no Verilog-2001 equivalent, so the byte is extracted with an explicit shift and mask — the portable idiom, and worth recognising because it appears wherever parameterised field extraction is needed.
// spi_read_seq_tb.v — the same checks in Verilog-2001.
`timescale 1ns/1ps
module spi_read_seq_tb;
reg clk = 0, rst_n = 0;
always #5 clk = ~clk;
parameter AB = 3;
reg start = 0, bit_stb = 0, byte_done = 0;
reg [7:0] opcode = 8'h0B, rx_byte = 8'h00, n_data = 8'd2;
reg [8*AB-1:0] addr = 24'h123456;
reg [4:0] dummy_cycles = 5'd8;
wire [7:0] tx_byte, rx_data;
wire cs_n, rx_valid, busy, done;
spi_read_seq #(.ADDR_BYTES(AB)) dut (
.clk(clk), .rst_n(rst_n), .start(start), .opcode(opcode), .addr(addr),
.dummy_cycles(dummy_cycles), .n_data(n_data), .bit_stb(bit_stb),
.byte_done(byte_done), .rx_byte(rx_byte), .tx_byte(tx_byte),
.cs_n(cs_n), .rx_valid(rx_valid), .rx_data(rx_data), .busy(busy), .done(done));
integer errors = 0, cs_rises = 0, tx_n = 0, rx_n = 0, i;
reg [7:0] tx_seen [0:15];
reg [7:0] rx_got [0:15];
reg cs_n_q;
task chk;
input [80*8-1:0] what;
input [31:0] got, exp;
begin
if (got !== exp) begin
$display("FAIL %0s: got 0x%0h exp 0x%0h", what, got, exp);
errors = errors + 1;
end
end
endtask
always @(posedge clk) if (rst_n) begin
cs_n_q <= cs_n;
if (cs_n && !cs_n_q) cs_rises = cs_rises + 1;
if (rx_valid) begin rx_got[rx_n] = rx_data; rx_n = rx_n + 1; end
end
task one_bit;
input last;
begin
@(negedge clk); bit_stb = 1; byte_done = last;
@(negedge clk); bit_stb = 0; byte_done = 0;
@(negedge clk);
end
endtask
task one_byte;
input [7:0] miso_val;
integer k;
begin
tx_seen[tx_n] = tx_byte; tx_n = tx_n + 1;
rx_byte = miso_val;
for (k = 0; k < 8; k = k + 1) one_bit(k == 7);
end
endtask
initial begin
repeat (3) @(negedge clk); rst_n = 1; @(negedge clk);
chk("idle: cs_n high", cs_n, 1);
chk("idle: not busy", busy, 0);
start = 1; @(negedge clk); start = 0; @(negedge clk);
chk("frame opened", cs_n, 0);
chk("busy", busy, 1);
one_byte(8'h00); one_byte(8'h00); one_byte(8'h00); one_byte(8'h00);
chk("CS still low after address", cs_n, 0);
repeat (8) one_bit(1'b0);
one_byte(8'hDE);
one_byte(8'hAD);
repeat (4) @(negedge clk);
chk("frame closed", cs_n, 1);
chk("cs rose once", cs_rises, 1);
chk("two bytes back", rx_n, 2);
chk("data byte 0", rx_got[0], 8'hDE);
chk("data byte 1", rx_got[1], 8'hAD);
chk("tx opcode", tx_seen[0], 8'h0B);
chk("tx addr hi", tx_seen[1], 8'h12);
chk("tx addr mid", tx_seen[2], 8'h34);
chk("tx addr lo", tx_seen[3], 8'h56);
chk("tx filler in data phase", tx_seen[4], 8'h00);
tx_n = 0; rx_n = 0;
dummy_cycles = 5'd0; n_data = 8'd1; opcode = 8'h03;
@(negedge clk);
start = 1; @(negedge clk); start = 0; @(negedge clk);
one_byte(8'h00); one_byte(8'h00); one_byte(8'h00); one_byte(8'h00);
one_byte(8'hC3);
repeat (4) @(negedge clk);
chk("no-dummy: one byte back", rx_n, 1);
chk("no-dummy: value", rx_got[0], 8'hC3);
chk("no-dummy: frame closed", cs_n, 1);
chk("no-dummy: cs rose twice", cs_rises, 2);
if (errors == 0)
$display("PASS: one frame spans the direction change, the address is sent MSB-first, filler is driven once the request is complete, and the dummy phase is skipped when zero");
else
$display("FAILED with %0d error(s)", errors);
$finish;
end
initial begin #500000; $display("FAIL: watchdog timeout"); $finish; end
endmodule-- spi_read_seq.vhd — the same master-side read sequencer in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_read_seq is
generic (
ADDR_BYTES : positive := 3;
FILLER : std_logic_vector(7 downto 0) := x"00"
);
port (
clk : in std_logic;
rst_n : in std_logic;
start : in std_logic;
opcode : in std_logic_vector(7 downto 0);
addr : in std_logic_vector(8 * ADDR_BYTES - 1 downto 0);
dummy_cycles : in unsigned(4 downto 0);
n_data : in unsigned(7 downto 0);
bit_stb : in std_logic;
byte_done : in std_logic;
rx_byte : in std_logic_vector(7 downto 0);
tx_byte : out std_logic_vector(7 downto 0);
cs_n : out std_logic;
rx_valid : out std_logic;
rx_data : out std_logic_vector(7 downto 0);
busy : out std_logic;
done : out std_logic
);
end entity spi_read_seq;
architecture rtl of spi_read_seq is
type state_t is (ST_IDLE, ST_CMD, ST_ADDR, ST_DUMMY, ST_DATA, ST_CLOSE);
signal state : state_t;
signal addr_cnt : integer range 0 to ADDR_BYTES - 1;
signal dummy_cnt : unsigned(4 downto 0);
signal data_cnt : unsigned(7 downto 0);
-- Address byte `idx`, most-significant first.
function addr_byte (a : std_logic_vector; idx : integer) return std_logic_vector is
variable hi : integer;
begin
hi := 8 * (ADDR_BYTES - 1 - idx) + 7;
return a(hi downto hi - 7);
end function;
begin
process (clk, rst_n) is
begin
if rst_n = '0' then
state <= ST_IDLE;
cs_n <= '1';
tx_byte <= FILLER;
addr_cnt <= 0;
dummy_cnt <= (others => '0');
data_cnt <= (others => '0');
rx_valid <= '0';
rx_data <= (others => '0');
done <= '0';
elsif rising_edge(clk) then
rx_valid <= '0';
done <= '0';
case state is
when ST_IDLE =>
if start = '1' then
cs_n <= '0'; -- open the frame and keep it open
tx_byte <= opcode;
addr_cnt <= 0;
dummy_cnt <= (others => '0');
data_cnt <= (others => '0');
state <= ST_CMD;
end if;
when ST_CMD =>
if byte_done = '1' then
tx_byte <= addr_byte(addr, 0);
addr_cnt <= 0;
state <= ST_ADDR;
end if;
when ST_ADDR =>
if byte_done = '1' then
if addr_cnt = ADDR_BYTES - 1 then
-- Request complete: everything after this is
-- filler on MOSI.
tx_byte <= FILLER;
if dummy_cycles /= 0 then
state <= ST_DUMMY;
else
state <= ST_DATA;
end if;
else
addr_cnt <= addr_cnt + 1;
tx_byte <= addr_byte(addr, addr_cnt + 1);
end if;
end if;
-- Dummy is counted in BITS, not bytes.
when ST_DUMMY =>
if bit_stb = '1' then
if dummy_cnt = dummy_cycles - 1 then
state <= ST_DATA;
else
dummy_cnt <= dummy_cnt + 1;
end if;
end if;
when ST_DATA =>
if byte_done = '1' then
rx_data <= rx_byte;
rx_valid <= '1';
if data_cnt = n_data - 1 then
state <= ST_CLOSE;
else
data_cnt <= data_cnt + 1;
end if;
end if;
when ST_CLOSE =>
cs_n <= '1'; -- only now is the frame ended
done <= '1';
state <= ST_IDLE;
end case;
end if;
end process;
busy <= '0' when state = ST_IDLE else '1';
end architecture rtl;-- spi_read_seq_tb.vhd — the same checks in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_read_seq_tb is
end entity spi_read_seq_tb;
architecture tb of spi_read_seq_tb is
constant AB : positive := 3;
signal clk : std_logic := '0';
signal rst_n : std_logic := '0';
signal start : std_logic := '0';
signal bit_stb : std_logic := '0';
signal byte_done : std_logic := '0';
signal halt : boolean := false;
signal opcode : std_logic_vector(7 downto 0) := x"0B";
signal addr : std_logic_vector(8 * AB - 1 downto 0) := x"123456";
signal dummy_cycles : unsigned(4 downto 0) := to_unsigned(8, 5);
signal n_data : unsigned(7 downto 0) := to_unsigned(2, 8);
signal rx_byte : std_logic_vector(7 downto 0) := (others => '0');
signal tx_byte : std_logic_vector(7 downto 0);
signal cs_n : std_logic;
signal rx_valid : std_logic;
signal rx_data : std_logic_vector(7 downto 0);
signal busy : std_logic;
signal done : std_logic;
signal errors : natural := 0;
signal cs_rises : natural := 0;
type bytes_t is array (0 to 15) of std_logic_vector(7 downto 0);
signal rx_got : bytes_t := (others => (others => '0'));
signal rx_n : natural := 0;
signal rx_rst : std_logic := '0';
begin
clk <= not clk after 5 ns when not halt else '0';
dut : entity work.spi_read_seq
generic map (ADDR_BYTES => AB)
port map (clk => clk, rst_n => rst_n, start => start, opcode => opcode,
addr => addr, dummy_cycles => dummy_cycles, n_data => n_data,
bit_stb => bit_stb, byte_done => byte_done, rx_byte => rx_byte,
tx_byte => tx_byte, cs_n => cs_n, rx_valid => rx_valid,
rx_data => rx_data, busy => busy, done => done);
-- One process owns rx_got/rx_n/cs_rises; stim only requests a rewind.
collect : process (clk) is
variable cs_n_q : std_logic := '1';
begin
if rising_edge(clk) then
if rst_n = '1' then
if cs_n = '1' and cs_n_q = '0' then
cs_rises <= cs_rises + 1;
end if;
if rx_rst = '1' then
rx_n <= 0;
elsif rx_valid = '1' then
rx_got(rx_n) <= rx_data;
rx_n <= rx_n + 1;
end if;
end if;
cs_n_q := cs_n;
end if;
end process;
stim : process is
variable tx_seen : bytes_t := (others => (others => '0'));
variable tx_n : natural := 0;
procedure chk_n (what : string; got, exp : natural) is
begin
if got /= exp then
report "FAIL " & what & ": got " & integer'image(got)
& " exp " & integer'image(exp) severity error;
errors <= errors + 1;
end if;
end procedure;
procedure chk_v (what : string; got, exp : std_logic_vector) is
begin
if got /= exp then
report "FAIL " & what & ": got 0x" & to_hstring(got)
& " exp 0x" & to_hstring(exp) severity error;
errors <= errors + 1;
end if;
end procedure;
procedure chk_b (what : string; got, exp : std_logic) is
begin
if got /= exp then
report "FAIL " & what severity error;
errors <= errors + 1;
end if;
end procedure;
procedure one_bit (last : std_logic) is
begin
wait until falling_edge(clk); bit_stb <= '1'; byte_done <= last;
wait until falling_edge(clk); bit_stb <= '0'; byte_done <= '0';
wait until falling_edge(clk);
end procedure;
procedure one_byte (miso_val : std_logic_vector(7 downto 0)) is
begin
tx_seen(tx_n) := tx_byte; tx_n := tx_n + 1;
rx_byte <= miso_val;
for k in 0 to 7 loop
if k = 7 then one_bit('1'); else one_bit('0'); end if;
end loop;
end procedure;
begin
for i in 0 to 2 loop
wait until falling_edge(clk);
end loop;
rst_n <= '1';
wait until falling_edge(clk);
chk_b("idle: cs_n high", cs_n, '1');
chk_b("idle: not busy", busy, '0');
start <= '1'; wait until falling_edge(clk); start <= '0';
wait until falling_edge(clk);
chk_b("frame opened", cs_n, '0');
chk_b("busy", busy, '1');
one_byte(x"00"); one_byte(x"00"); one_byte(x"00"); one_byte(x"00");
chk_b("CS still low after address", cs_n, '0');
for i in 0 to 7 loop one_bit('0'); end loop;
one_byte(x"DE");
one_byte(x"AD");
for i in 0 to 3 loop wait until falling_edge(clk); end loop;
chk_b("frame closed", cs_n, '1');
chk_n("cs rose once", cs_rises, 1);
chk_n("two bytes back", rx_n, 2);
chk_v("data byte 0", rx_got(0), x"DE");
chk_v("data byte 1", rx_got(1), x"AD");
chk_v("tx opcode", tx_seen(0), x"0B");
chk_v("tx addr hi", tx_seen(1), x"12");
chk_v("tx addr mid", tx_seen(2), x"34");
chk_v("tx addr lo", tx_seen(3), x"56");
chk_v("tx filler in data phase", tx_seen(4), x"00");
-- A read with no dummy phase must go straight to data.
tx_n := 0;
rx_rst <= '1'; wait until falling_edge(clk); rx_rst <= '0';
dummy_cycles <= to_unsigned(0, 5);
n_data <= to_unsigned(1, 8);
opcode <= x"03";
wait until falling_edge(clk);
start <= '1'; wait until falling_edge(clk); start <= '0';
wait until falling_edge(clk);
one_byte(x"00"); one_byte(x"00"); one_byte(x"00"); one_byte(x"00");
one_byte(x"C3");
for i in 0 to 3 loop wait until falling_edge(clk); end loop;
chk_n("no-dummy: one byte back", rx_n, 1);
chk_v("no-dummy: value", rx_got(0), x"C3");
chk_b("no-dummy: frame closed", cs_n, '1');
chk_n("no-dummy: cs rose twice", cs_rises, 2);
if errors = 0 then
report "PASS: one frame spans the direction change, the address is sent "
& "MSB-first, filler is driven once the request is complete, and the "
& "dummy phase is skipped when zero" severity note;
else
report "FAILED with " & integer'image(errors) & " error(s)" severity error;
end if;
halt <= true;
wait;
end process;
watchdog : process is
begin
wait for 500 us;
if not halt then
report "FAIL: watchdog timeout" severity failure;
end if;
wait;
end process;
end architecture tb;Parity
All three implement the same machine: identical ports, asynchronous active-low reset to idle with CS high, CS asserted on start and released only in the closing state, the address presented most-significant byte first, filler driven once the request completes, a dummy phase counted in bits and skipped when zero, payload bytes forwarded with a single-cycle valid strobe, and a single done pulse.
One deliberate difference: the Verilog dialects guard the address phase with if (ADDR_BYTES > 0) while the VHDL does not, because VHDL's positive generic type makes a zero value impossible at elaboration. The behaviour is identical for every legal parameter value; the type system is simply doing the check in one language and not the other.
7. Why a Verification Engineer Cares
The one-frame property is the one to assert, because violating it produces silent data corruption rather than an error:
// 1. THE property of this chapter. CS must not rise between the start of
// the transaction and its completion. A two-frame read looks like two
// unrelated transactions to the device (§2).
property p_single_frame;
@(posedge clk) disable iff (!rst_n)
$fell(cs_n) |-> (!cs_n throughout done[->1]);
endproperty
a_single_frame : assert property (p_single_frame)
else $error("CS rose before the transaction completed -- request and response split across frames");
// 2. Payload bytes are forwarded ONLY from the data phase. Forwarding an
// address-phase byte would hand the scoreboard the device's pre-response
// garbage as though it were data.
a_rx_only_in_data : assert property (
@(posedge clk) disable iff (!rst_n) rx_valid |-> (state == ST_DATA))
else $error("payload byte forwarded from outside the data phase");
// 3. Exactly n_data payload bytes per transaction -- no more, no fewer.
property p_payload_count;
@(posedge clk) disable iff (!rst_n)
$rose(busy) |-> (rx_valid[->n_data] ##1 done[->1]);
endproperty
// 4. The master must drive the agreed filler once its half is done.
// Some devices interpret MOSI during the response (Chapter 5.2 §3).
a_filler_in_response : assert property (
@(posedge clk) disable iff (!rst_n)
(state inside {ST_DUMMY, ST_DATA}) |-> (tx_byte == FILLER))
else $error("master drove a non-filler byte during the response");What these prove. That the sequencer holds one frame across the turnaround, forwards only payload, and drives the agreed filler. What they do not prove is that the device expects a single frame — a device with a latched-result protocol wants two, and asserting p_single_frame against it would be asserting the wrong specification. That mapping is a datasheet fact, and it is Chapter 4.1 §7's boundary once more.
Coverage should target the sequencer's structural axes:
covergroup spi_read_seq_cg @(posedge done);
cp_addr_bytes : coverpoint cfg.addr_bytes {
bins none = {0}; // status-style command, no address
bins one = {1};
bins three = {3}; // the common memory case
bins four = {4}; // large-capacity addressing
}
// §2 and Chapter 4.5: zero dummy takes a different path through the
// FSM than any non-zero value, and it is a distinct branch.
cp_dummy : coverpoint cfg.dummy_cycles {
bins none = {0}; // MUST be hit -- the skip path
bins sub_byte = {[1:7]};
bins one_byte = {8};
bins many = {[9:31]};
}
cp_payload : coverpoint cfg.n_data {
bins one = {1}; // never crosses a byte boundary
bins two = {2}; // the smallest burst
bins burst = {[3:255]};
}
x_addr_dummy : cross cp_addr_bytes, cp_dummy;
endgroupThe cp_dummy.none bin is the one to insist on, for the same reason as Chapter 4.5 §8: zero dummy cycles takes a different branch — straight from the address phase to the data phase — and a suite that always configures a non-zero count never executes it.
8. Why an FPGA or ASIC Engineer Cares
tx_byte must be ready before the engine needs it. The sequencer selects the next transmit byte on the byte_done that finishes the previous one, giving the byte engine a full byte time to consume it. Selecting it when the engine asks would put the address multiplexer in the critical path between a strobe and the first launch edge — which at high SCLK is exactly where you do not want a wide multiplexer.
The address multiplexer grows with address length. Four address bytes means a 4:1 byte-wide mux. That is small, but it is combinational logic selected by a counter, and on a design supporting both 3- and 4-byte addressing at runtime the select becomes a register output rather than a constant. Same trade as Chapter 4.2 §8.
Reset must leave CS deasserted. A sequencer resetting with cs_n low would open a frame the instant reset released, and any slave on the bus would begin decoding whatever the clock did next. This is the same reasoning as Chapter 5.2 §6's synchroniser resetting high, applied at the other end of the link.
Think about what happens if the fabric stalls. This sequencer assumes the byte engine keeps clocking. If the transmit data came from a FIFO that underran mid-frame, the master would have to either stall SCLK — legal, since SPI has no timeout, and the reason SPI tolerates a paused clock at all — or keep clocking with stale data. Stalling is correct and it is worth knowing the bus permits it; many first designs do the other thing by accident.
9. Failure Signature — Every Read Returns the Previous Read's Data
Symptom. Reads return plausible values that are consistently one transaction stale: read register A, get whatever register B returned last time; read A again immediately and now it is correct. Writes work. The pattern is perfectly reproducible and unaffected by clock rate.
What "one transaction stale" tells you. The data is real and correctly transferred — it is simply the wrong data, arriving intact. That rules out mode, bit order, framing and signal integrity in one step, since all of those corrupt values rather than displacing them in time (Chapter 4.3 §8's reasoning applied to a different axis).
Plausible mechanisms.
- The device uses a latched-result protocol: the transaction that issues a command returns the previous command's result, and the driver is treating it as a one-frame read. This is the §2 exception, and it is the leading candidate.
- The master is capturing one byte early — collecting a byte from the dummy phase as payload, so every byte shifts by one position and the last is stale.
- The driver reads a buffer that the previous transfer filled, a software bug with no SPI cause at all.
- The device requires a dummy read after a command, and the first result is discarded by design.
The discriminating observation. Issue the same read twice in succession and compare. If the second is correct while the first is stale, the device is one transaction behind and you have the latched-result protocol — the fix is to read twice and discard the first, or to restructure as command-then-separate-read, which for this device class is the specified arrangement rather than a workaround.
If both reads are stale by one, the pipeline is in the master or the driver, not the device.
That test costs one extra transaction and separates a device protocol from a master bug, which are otherwise indistinguishable from the symptom.
Why the investigation goes wrong. Because stale-but-valid data does not look like a protocol problem — it looks like a caching or buffering bug, and people search the driver. The SPI-level cause is a device whose read semantics are genuinely two-frame, which contradicts the reasonable default of §2. The default is right often enough that the exception is rarely suspected.
10. Common Misconceptions
11. Reason It Through
Work this before reading the answer.
A driver polls a sensor's status register in a tight loop, then reads a 4-byte sample when the ready bit sets. Each status poll is an 8-bit command plus an 8-bit response — 16 clock cycles — and each sample read is a command, a 2-byte address, 8 dummy cycles and 4 data bytes.
At 10 MHz the system reads 8,000 samples per second and the CPU is mostly idle waiting. The team proposes raising SCLK to 40 MHz to get 32,000 samples per second. Will they?
Work out where the time actually goes. A sample read is 8 + 16 + 8 + 32 = 64 cycles. At 10 MHz that is 6.4 µs. At 8,000 samples per second the sample reads occupy 8000 × 6.4 µs = 51 ms of each second — about 5% of the bus.
So what is the other 95%? The polling loop. Each poll is 16 cycles, 1.6 µs at 10 MHz, and the loop runs continuously between samples. In the ~119 µs between samples the driver issues roughly 74 status polls per sample.
Now apply the proposed change. Quadrupling SCLK makes every transaction four times faster — including the polls. The sample read drops to 1.6 µs and the polls to 400 ns. But the sample rate is set by the sensor, not by the bus: it produces 8,000 samples per second because that is its conversion rate. Going faster on the bus means the driver simply polls more times per sample, burning the same proportion of bus time on the same useless question.
So the answer is no, and the reason is that the bus was never the bottleneck. The team measured "the CPU is idle waiting" and concluded the transfer was slow, when the waiting was for the sensor.
What would actually help? Stop polling. If the device has a data-ready interrupt line — many do, and it is the single most valuable pin on a sensor — the polls disappear entirely and the bus goes idle between samples. Failing that, poll on a timer sized to the conversion time rather than in a tight loop, which converts 74 polls per sample into one or two.
And if the sample rate itself must rise? That is a sensor configuration question, and only then does the bus rate matter — because at 32,000 samples per second the sample reads alone would need 32000 × 6.4 µs = 205 ms, still only 20% of the bus at 10 MHz. The bus has ample headroom either way.
The general lesson. Section 4's efficiency arithmetic answers "how much of this transaction is overhead". It does not answer "is the bus the bottleneck" — and on a polled sensor the answer is usually no. Measure which transactions dominate before optimising the clock.
12. Understanding Check
13. Summary
Every SPI read is a write first: the master transmits an opcode and usually an address before anything comes back. A read is therefore half-duplex traffic on a full-duplex bus, and what changes at the turnaround is not the direction of transmission — both directions are always active — but which direction the master pays attention to.
The request and the response must normally occupy one CS frame, because CS deassertion ends the transaction and discards the request. The exceptions are real and are datasheet facts: latched-result converters, continuous-read modes, and some status registers.
Throughout the response the master drives filler on MOSI, since every clocked edge transmits something. The value is a device agreement and belongs in configuration.
The request is overhead, and for a 1-byte opcode, 3-byte address and 8 dummy cycles it costs 40 cycles before the first payload bit — making a single-byte read about 17% efficient and a 256-byte burst about 98%. Raising SCLK does not change that ratio, because it scales request and payload equally; fewer, larger transactions is the lever that does.
In RTL the master side is a six-state machine that asserts CS on start and releases it only in a dedicated closing state — which is what makes the one-frame property structural rather than emergent — presents the next transmit byte a full byte time ahead, counts the dummy phase in bits rather than bytes, and forwards received bytes only from the data phase.
For verification the property worth asserting is that CS does not rise before the transaction completes, remembering that a latched-result device genuinely wants the opposite. And when reads come back one transaction stale, the data is intact and merely displaced in time — issue the same read twice to separate a device protocol from a master pipeline bug.
14. What Comes Next
This chapter established when the device's data arrives. It said nothing about when that data may be believed — and on a real board those are different instants, separated by the device's output delay, the flight time down the trace, and the master's own sampling choice. Chapter 6.3 — MISO Valid Timing makes that window explicit, turns Module 2's round-trip budget into a transaction-level concern, and builds the master's configurable input sampler in all three HDLs.
Continue learning
Related tutorials
- Related topic
Command, Address, and Data Phases
How a device layers a transaction onto a raw byte stream: why the opcode decides the shape of everything after it, how a slave tracks phases with no phase marker, and the sequencer that requires in three HDLs.
- Related topic
The Master Control FSM
The five states an SPI master transfer passes through, why the fifth differs from the first by exactly one behaviour, why outputs must be decoded from state alone, and the failure that disappears the moment you add a print statement.
- Related topic
CS Detection and Transaction Boundaries
Chip select is SPI's only framing and therefore its only resynchronisation point: why counters must reset on assert, how to classify a transaction with a running remainder instead of a divider, why a CPHA mismatch is invisible to the edge count, and a transaction detector verified in three HDLs.
- Related topic
Launch and Sample Edges
One edge of each bit time places a bit on the wire, the other captures it, and they must never be the same edge. Why the separation is forced, why it buys half a period, and how RTL maps physical edges onto those roles.
