SPI · Module 7
Continuous Transfers Under One CS
Holding chip select low across many bytes and what the device assumes: why the frame is the transaction, why the master may legally stop the clock when its data runs dry, and the streaming controller in three HDLs.
Modules 5 and 6 each followed a transaction of known length: a write of one register, a read of a few bytes. Both ended when the master decided it had finished. This module asks what happens when transfers stop being small.
Chip select goes low and stays low for a thousand byte times. What does the device believe is happening — and what is the master obliged to do for the entire duration?
The second half has an answer that surprises people, and it follows from something SPI does not have.
1. One Frame, Many Bytes, No Boundaries
Chapter 4.2 §2 established that CS delimits the frame and the master's bit counter is private bookkeeping. A continuous transfer takes that to its conclusion: CS falls once, a large number of bytes cross, and CS rises once.
From the device's side there is one transaction. Not a thousand transactions sharing a frame — one. That matters because everything the device tracks is scoped to the frame:
- Its phase sequencer decoded one command and has been in the data phase ever since (Chapter 4.4).
- Its address pointer has been advancing from the single address it was given (Chapter 5.3).
- Its output enable has been asserted since the data phase opened (Chapter 6.1).
None of that is re-evaluated at byte boundaries, because the device cannot see byte boundaries. It counts edges. Byte 700 is distinguished from byte 699 only by its position in a count that began at CS.
2. The Master Owns the Clock — and May Stop It
Here is the fact that shapes the rest of this chapter.
SPI has no timeout. There is no minimum data rate, no maximum gap between edges, and no mechanism by which a device can notice that time has passed. The device is a shift register with combinational control: it responds to edges, and between edges it does nothing at all.
So if the master's data source runs dry mid-frame, its options are:
Stop clocking. Hold CS low, stop producing SCLK edges, and resume when data is available. The device waits indefinitely and notices nothing. This is correct.
Keep clocking with whatever is in the register. This transmits a stale byte the device will accept as real data. It is silent, unreported, and indistinguishable from intended traffic (Chapter 5.1 §5). This is wrong and is the bug §10 describes.
Deassert CS. This ends the transaction. The device discards its state, and there is no way to resume a continuous transfer from the middle — the next frame starts from a fresh command. This is catastrophic for a long transfer and merely wasteful for a short one.
The asymmetry with I²C is worth naming. I²C gives the slave a way to stall the master — clock stretching, where the slave holds SCL low until ready. SPI gives the master a way to stall itself and gives the slave nothing. The device cannot ask for time; it can only be given it.
3. The Stall
The clock may stop; the selection may not
10 cyclesThe stall in the figure is three bit times. It could equally be three milliseconds. Nothing about the device's behaviour changes with the length of the pause, which is the practical consequence of having no timeout, and it is what makes stalling a genuine solution rather than a delaying tactic.
4. The Transaction
5. Building the Streaming Controller — Three HDLs
The circuit
Circuit. A five-state machine mediating between a data source with a valid/ready handshake and the byte engine, with authority over both CS and the clock enable.
State. Which of five conditions the frame is in: idle, primed-but-not-yet-clocking, running, stalled, or closing.
Datapath. One registered byte, taken from the source and handed to the byte engine.
Control. The want_byte term is expressed once and covers all three states that need data. Writing the handshake condition separately in each state is how a valid/ready interface drifts out of agreement with itself — a byte consumed in one state and not another, or consumed twice.
Clock. The system clock. sclk_en gates the SPI clock generator; it is a control output, not a gated clock inside this module.
Reset. Asynchronous, active-low, to idle with CS deasserted and the clock disabled — the safe state, since a slave must not see a frame open out of reset (Chapter 5.2 §6).
Enables. The ST_PRIME state exists specifically so that CS asserts before the clock starts. Opening the frame and immediately clocking would transmit whatever the register held; priming guarantees the first byte is real.
Timing. underrun is a single-cycle strobe, not a level, so software sees one event per stall rather than a condition to poll.
Synthesis. Three state bits, one WIDTH register, and a handful of gates. Trivially small — the value here is entirely in the control discipline.
Limitations. One byte of buffering. A real controller places a FIFO upstream, which is §8's subject, and the depth of that FIFO is what determines whether stalls happen at all.
// spi_stream_ctrl.sv — a continuous transfer under one chip select.
//
// The idea that makes this module interesting: SPI has no timeout and no
// minimum data rate. The master owns the clock, so if its data source runs
// dry mid-frame the correct response is to STOP CLOCKING and hold CS low.
// The device simply waits -- it has no way to know or care that time passed.
//
// The wrong response is to keep clocking with whatever is in the register,
// which transmits stale data the device will accept as real. That failure is
// silent, because nothing on an SPI bus reports anything (Chapter 5.1 §5).
module spi_stream_ctrl #(
parameter int WIDTH = 8
) (
input logic clk,
input logic rst_n,
// --- control ---
input logic start, // pulse: open a streaming frame
input logic stop, // level: close after the current byte
// --- upstream data source (FIFO-like handshake) ---
input logic [WIDTH-1:0] src_data,
input logic src_valid,
output logic src_ready, // pulse: byte consumed this cycle
// --- byte engine ---
input logic byte_done, // one pulse per transmitted byte
output logic [WIDTH-1:0] tx_byte,
// --- bus ---
output logic cs_n,
output logic sclk_en, // gates the clock generator
output logic underrun, // pulse: had to stall for data
output logic busy
);
typedef enum logic [2:0] {
ST_IDLE, // no frame
ST_PRIME, // CS low, waiting for the first byte -- not yet clocking
ST_RUN, // clocking
ST_STALL, // CS low, clock stopped, waiting for data
ST_CLOSE // finish the frame
} state_t;
state_t state;
// Consume a byte exactly when we are in a state that needs one and the
// source has it. Expressed once so the handshake cannot drift.
logic want_byte;
assign want_byte = (state == ST_PRIME)
|| (state == ST_STALL)
|| (state == ST_RUN && byte_done && !stop);
assign src_ready = want_byte && src_valid;
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
state <= ST_IDLE;
cs_n <= 1'b1;
sclk_en <= 1'b0;
tx_byte <= '0;
underrun <= 1'b0;
end else begin
underrun <= 1'b0; // strobe
case (state)
ST_IDLE: if (start) begin
cs_n <= 1'b0; // frame opens and stays open
sclk_en <= 1'b0; // but do not clock without data
state <= ST_PRIME;
end
ST_PRIME: if (src_valid) begin
tx_byte <= src_data;
sclk_en <= 1'b1;
state <= ST_RUN;
end
ST_RUN: if (byte_done) begin
if (stop) begin
state <= ST_CLOSE;
end else if (src_valid) begin
tx_byte <= src_data; // seamless: no gap in clocking
end else begin
// UNDERRUN. Stop the clock; CS stays low. The device
// is mid-transaction and must remain selected.
sclk_en <= 1'b0;
underrun <= 1'b1;
state <= ST_STALL;
end
end
ST_STALL: if (src_valid) begin
tx_byte <= src_data;
sclk_en <= 1'b1; // resume exactly where we stopped
state <= ST_RUN;
end
ST_CLOSE: begin
cs_n <= 1'b1;
sclk_en <= 1'b0;
state <= ST_IDLE;
end
default: state <= ST_IDLE;
endcase
end
end
assign busy = (state != ST_IDLE);
endmoduleThe ST_STALL state is the module's whole point, and note what it does not touch: cs_n is left alone. There is no path in this machine from ST_STALL to anything that raises CS. That is deliberate and structural — the one mistake this design exists to prevent is ending a transaction because data ran out.
// spi_stream_ctrl_tb.sv — seamless streaming, an underrun mid-frame, and the
// property that matters most: CS stays low while the clock is stopped.
`timescale 1ns/1ps
module spi_stream_ctrl_tb;
logic clk = 0, rst_n = 0;
always #5 clk = ~clk;
localparam int W = 8;
logic start = 0, stop = 0, src_valid = 0, byte_done = 0;
logic [W-1:0] src_data = '0;
logic src_ready, cs_n, sclk_en, underrun, busy;
logic [W-1:0] tx_byte;
spi_stream_ctrl #(.WIDTH(W)) dut (
.clk, .rst_n, .start, .stop, .src_data, .src_valid, .src_ready,
.byte_done, .tx_byte, .cs_n, .sclk_en, .underrun, .busy);
int errors = 0, underruns = 0, cs_rises = 0, taken = 0;
logic [W-1:0] sent [$];
logic cs_n_q;
task automatic chk(input string what, input int g, input int e);
if (g !== e) begin $display("FAIL %s: got 0x%0h exp 0x%0h", what, g, e); errors++; end
endtask
// Record every accepted byte at the edge that accepts it.
always @(posedge clk) if (rst_n) begin
cs_n_q <= cs_n;
if (cs_n && !cs_n_q) cs_rises++;
if (underrun) underruns++;
if (src_ready) begin sent.push_back(src_data); taken++; end
// CRITICAL INVARIANT: CS must never be high while a frame is in
// progress. The clock may stop; the selection may not.
if (busy && cs_n) begin
$display("FAIL: CS deasserted mid-frame at t=%0t", $time); errors++;
end
end
// Present a byte and wait until the controller actually takes it.
// src_ready is combinational, so settle before sampling it.
task automatic offer(input logic [W-1:0] b);
src_data = b; src_valid = 1;
#1;
while (!src_ready) begin @(negedge clk); #1; end
@(posedge clk); // the edge that consumes it
@(negedge clk); src_valid = 0;
endtask
// One transmitted byte time.
task automatic byte_time();
repeat (3) @(negedge clk);
byte_done = 1; @(negedge clk); byte_done = 0; @(negedge clk);
endtask
initial begin
repeat (3) @(negedge clk); rst_n = 1; @(negedge clk);
chk("idle: cs_n high", cs_n, 1);
chk("idle: no clock", sclk_en, 0);
// --- open a frame; the clock must NOT start before data exists ---
start = 1; @(negedge clk); start = 0; @(negedge clk);
chk("frame open", cs_n, 0);
chk("not clocking while dry", sclk_en, 0);
offer(8'hA1);
repeat (2) @(negedge clk);
chk("clocking once primed", sclk_en, 1);
chk("first byte loaded", tx_byte, 8'hA1);
// --- seamless: hold the next byte valid across the byte boundary ---
src_data = 8'hB2; src_valid = 1;
byte_time();
@(negedge clk); src_valid = 0; @(negedge clk);
chk("no underrun when fed in time", underruns, 0);
chk("still clocking", sclk_en, 1);
chk("next byte loaded", tx_byte, 8'hB2);
// --- UNDERRUN: let this byte finish with nothing available ---
byte_time();
repeat (2) @(negedge clk);
chk("underrun reported", underruns, 1);
chk("clock stopped", sclk_en, 0);
chk("CS STILL LOW", cs_n, 0); // the property that matters
chk("still busy", busy, 1);
// --- the stall may last arbitrarily long; SPI has no timeout ---
repeat (40) @(negedge clk);
chk("stall persists harmlessly", sclk_en, 0);
chk("CS still low after long stall", cs_n, 0);
chk("no spurious underruns", underruns, 1);
// --- resume exactly where it stopped ---
offer(8'hC3);
repeat (2) @(negedge clk);
chk("resumed clocking", sclk_en, 1);
chk("resumed byte", tx_byte, 8'hC3);
// --- close the frame ---
stop = 1;
byte_time();
repeat (3) @(negedge clk);
stop = 0; @(negedge clk);
chk("frame closed", cs_n, 1);
chk("clock off", sclk_en, 0);
chk("cs rose once", cs_rises, 1);
chk("not busy", busy, 0);
// --- exactly three bytes consumed, in order, none repeated ---
chk("three bytes consumed", taken, 3);
chk("byte 0", sent[0], 8'hA1);
chk("byte 1", sent[1], 8'hB2);
chk("byte 2", sent[2], 8'hC3);
if (errors == 0)
$display("PASS: the frame stays open across an arbitrarily long clock stall, the underrun is reported, and streaming resumes with no byte lost or repeated");
else
$display("FAILED with %0d error(s)", errors);
$finish;
end
initial begin #500000; $display("FAIL: watchdog timeout"); $finish; end
endmoduleTwo things in that testbench are doing real work.
The continuous invariant check — if (busy && cs_n) sampled every clock — asserts that CS is never high while a frame is in progress. It runs throughout the whole simulation rather than at chosen points, which is what makes it a proof about the design rather than a spot check.
The forty-cycle stall is not padding. It asserts that a long pause is harmless and produces no additional underrun strobes, catching a design that re-reports the condition every cycle and floods software with events.
// spi_stream_ctrl.v — the same continuous-transfer controller in Verilog-2001.
module spi_stream_ctrl #(
parameter WIDTH = 8
) (
input wire clk,
input wire rst_n,
input wire start,
input wire stop,
input wire [WIDTH-1:0] src_data,
input wire src_valid,
output wire src_ready,
input wire byte_done,
output reg [WIDTH-1:0] tx_byte,
output reg cs_n,
output reg sclk_en,
output reg underrun,
output wire busy
);
localparam ST_IDLE = 3'd0,
ST_PRIME = 3'd1,
ST_RUN = 3'd2,
ST_STALL = 3'd3,
ST_CLOSE = 3'd4;
reg [2:0] state;
wire want_byte;
// Expressed once so the handshake cannot drift between states.
assign want_byte = (state == ST_PRIME)
|| (state == ST_STALL)
|| (state == ST_RUN && byte_done && !stop);
assign src_ready = want_byte && src_valid;
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
state <= ST_IDLE;
cs_n <= 1'b1;
sclk_en <= 1'b0;
tx_byte <= {WIDTH{1'b0}};
underrun <= 1'b0;
end else begin
underrun <= 1'b0;
case (state)
ST_IDLE: if (start) begin
cs_n <= 1'b0; // frame opens and stays open
sclk_en <= 1'b0; // but do not clock without data
state <= ST_PRIME;
end
ST_PRIME: if (src_valid) begin
tx_byte <= src_data;
sclk_en <= 1'b1;
state <= ST_RUN;
end
ST_RUN: if (byte_done) begin
if (stop) begin
state <= ST_CLOSE;
end else if (src_valid) begin
tx_byte <= src_data;
end else begin
// UNDERRUN. Stop the clock; CS stays low.
sclk_en <= 1'b0;
underrun <= 1'b1;
state <= ST_STALL;
end
end
ST_STALL: if (src_valid) begin
tx_byte <= src_data;
sclk_en <= 1'b1;
state <= ST_RUN;
end
ST_CLOSE: begin
cs_n <= 1'b1;
sclk_en <= 1'b0;
state <= ST_IDLE;
end
default: state <= ST_IDLE;
endcase
end
end
assign busy = (state != ST_IDLE);
endmodule// spi_stream_ctrl_tb.v — the same checks in Verilog-2001.
`timescale 1ns/1ps
module spi_stream_ctrl_tb;
reg clk = 0, rst_n = 0;
always #5 clk = ~clk;
parameter W = 8;
reg start = 0, stop = 0, src_valid = 0, byte_done = 0;
reg [W-1:0] src_data = 0;
wire src_ready, cs_n, sclk_en, underrun, busy;
wire [W-1:0] tx_byte;
spi_stream_ctrl #(.WIDTH(W)) dut (
.clk(clk), .rst_n(rst_n), .start(start), .stop(stop), .src_data(src_data),
.src_valid(src_valid), .src_ready(src_ready), .byte_done(byte_done),
.tx_byte(tx_byte), .cs_n(cs_n), .sclk_en(sclk_en), .underrun(underrun),
.busy(busy));
integer errors = 0, underruns = 0, cs_rises = 0, taken = 0;
reg [W-1:0] sent [0:7];
reg cs_n_q;
task chk;
input [80*8-1:0] what;
input [31:0] g, e;
begin
if (g !== e) begin
$display("FAIL %0s: got 0x%0h exp 0x%0h", what, g, e);
errors = errors + 1;
end
end
endtask
always @(posedge clk) if (rst_n) begin
cs_n_q <= cs_n;
if (cs_n && !cs_n_q) cs_rises = cs_rises + 1;
if (underrun) underruns = underruns + 1;
if (src_ready) begin sent[taken] = src_data; taken = taken + 1; end
if (busy && cs_n) begin
$display("FAIL: CS deasserted mid-frame at t=%0t", $time);
errors = errors + 1;
end
end
task offer;
input [W-1:0] b;
begin
src_data = b; src_valid = 1;
#1;
while (!src_ready) begin @(negedge clk); #1; end
@(posedge clk);
@(negedge clk); src_valid = 0;
end
endtask
task byte_time;
begin
repeat (3) @(negedge clk);
byte_done = 1; @(negedge clk); byte_done = 0; @(negedge clk);
end
endtask
initial begin
repeat (3) @(negedge clk); rst_n = 1; @(negedge clk);
chk("idle: cs_n high", cs_n, 1);
chk("idle: no clock", sclk_en, 0);
start = 1; @(negedge clk); start = 0; @(negedge clk);
chk("frame open", cs_n, 0);
chk("not clocking while dry", sclk_en, 0);
offer(8'hA1);
repeat (2) @(negedge clk);
chk("clocking once primed", sclk_en, 1);
chk("first byte loaded", tx_byte, 8'hA1);
src_data = 8'hB2; src_valid = 1;
byte_time;
@(negedge clk); src_valid = 0; @(negedge clk);
chk("no underrun when fed in time", underruns, 0);
chk("still clocking", sclk_en, 1);
chk("next byte loaded", tx_byte, 8'hB2);
byte_time;
repeat (2) @(negedge clk);
chk("underrun reported", underruns, 1);
chk("clock stopped", sclk_en, 0);
chk("CS STILL LOW", cs_n, 0);
chk("still busy", busy, 1);
repeat (40) @(negedge clk);
chk("stall persists harmlessly", sclk_en, 0);
chk("CS still low after long stall", cs_n, 0);
chk("no spurious underruns", underruns, 1);
offer(8'hC3);
repeat (2) @(negedge clk);
chk("resumed clocking", sclk_en, 1);
chk("resumed byte", tx_byte, 8'hC3);
stop = 1;
byte_time;
repeat (3) @(negedge clk);
stop = 0; @(negedge clk);
chk("frame closed", cs_n, 1);
chk("clock off", sclk_en, 0);
chk("cs rose once", cs_rises, 1);
chk("not busy", busy, 0);
chk("three bytes consumed", taken, 3);
chk("byte 0", sent[0], 8'hA1);
chk("byte 1", sent[1], 8'hB2);
chk("byte 2", sent[2], 8'hC3);
if (errors == 0)
$display("PASS: the frame stays open across an arbitrarily long clock stall, the underrun is reported, and streaming resumes with no byte lost or repeated");
else
$display("FAILED with %0d error(s)", errors);
$finish;
end
initial begin #500000; $display("FAIL: watchdog timeout"); $finish; end
endmodule-- spi_stream_ctrl.vhd — the same continuous-transfer controller in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_stream_ctrl is
generic (
WIDTH : positive := 8
);
port (
clk : in std_logic;
rst_n : in std_logic;
start : in std_logic; -- pulse
stop : in std_logic; -- level
src_data : in std_logic_vector(WIDTH - 1 downto 0);
src_valid : in std_logic;
src_ready : out std_logic; -- pulse
byte_done : in std_logic;
tx_byte : out std_logic_vector(WIDTH - 1 downto 0);
cs_n : out std_logic;
sclk_en : out std_logic; -- gates the clock
underrun : out std_logic; -- pulse
busy : out std_logic
);
end entity spi_stream_ctrl;
architecture rtl of spi_stream_ctrl is
type state_t is (ST_IDLE, ST_PRIME, ST_RUN, ST_STALL, ST_CLOSE);
signal state : state_t;
signal want_byte : std_logic;
begin
-- Expressed once so the handshake cannot drift between states.
want_byte <= '1' when state = ST_PRIME
or state = ST_STALL
or (state = ST_RUN and byte_done = '1' and stop = '0')
else '0';
src_ready <= want_byte and src_valid;
busy <= '0' when state = ST_IDLE else '1';
process (clk, rst_n) is
begin
if rst_n = '0' then
state <= ST_IDLE;
cs_n <= '1';
sclk_en <= '0';
tx_byte <= (others => '0');
underrun <= '0';
elsif rising_edge(clk) then
underrun <= '0';
case state is
when ST_IDLE =>
if start = '1' then
cs_n <= '0'; -- frame opens and stays open
sclk_en <= '0'; -- but do not clock without data
state <= ST_PRIME;
end if;
when ST_PRIME =>
if src_valid = '1' then
tx_byte <= src_data;
sclk_en <= '1';
state <= ST_RUN;
end if;
when ST_RUN =>
if byte_done = '1' then
if stop = '1' then
state <= ST_CLOSE;
elsif src_valid = '1' then
tx_byte <= src_data; -- seamless
else
-- UNDERRUN. Stop the clock; CS stays low.
sclk_en <= '0';
underrun <= '1';
state <= ST_STALL;
end if;
end if;
when ST_STALL =>
if src_valid = '1' then
tx_byte <= src_data;
sclk_en <= '1';
state <= ST_RUN;
end if;
when ST_CLOSE =>
cs_n <= '1';
sclk_en <= '0';
state <= ST_IDLE;
end case;
end if;
end process;
end architecture rtl;-- spi_stream_ctrl_tb.vhd — the same checks in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_stream_ctrl_tb is
end entity spi_stream_ctrl_tb;
architecture tb of spi_stream_ctrl_tb is
constant W : positive := 8;
signal clk : std_logic := '0';
signal rst_n : std_logic := '0';
signal start : std_logic := '0';
signal stop : std_logic := '0';
signal src_valid : std_logic := '0';
signal byte_done : std_logic := '0';
signal src_data : std_logic_vector(W - 1 downto 0) := (others => '0');
signal halt : boolean := false;
signal src_ready : std_logic;
signal tx_byte : std_logic_vector(W - 1 downto 0);
signal cs_n : std_logic;
signal sclk_en : std_logic;
signal underrun : std_logic;
signal busy : std_logic;
signal errors : natural := 0; -- owned by stim
signal inv_err : natural := 0; -- owned by observe
signal underruns : natural := 0;
signal cs_rises : natural := 0;
signal taken : natural := 0;
type bytes_t is array (0 to 7) of std_logic_vector(W - 1 downto 0);
signal sent : bytes_t := (others => (others => '0'));
begin
clk <= not clk after 5 ns when not halt else '0';
dut : entity work.spi_stream_ctrl
generic map (WIDTH => W)
port map (clk => clk, rst_n => rst_n, start => start, stop => stop,
src_data => src_data, src_valid => src_valid, src_ready => src_ready,
byte_done => byte_done, tx_byte => tx_byte, cs_n => cs_n,
sclk_en => sclk_en, underrun => underrun, busy => busy);
-- One process owns the observation counters.
observe : process (clk) is
variable cs_n_q : std_logic := '1';
begin
if rising_edge(clk) then
if rst_n = '1' then
if cs_n = '1' and cs_n_q = '0' then
cs_rises <= cs_rises + 1;
end if;
if underrun = '1' then
underruns <= underruns + 1;
end if;
if src_ready = '1' then
sent(taken) <= src_data;
taken <= taken + 1;
end if;
-- CRITICAL INVARIANT: CS must never be high mid-frame.
if busy = '1' and cs_n = '1' then
report "FAIL: CS deasserted mid-frame" severity error;
inv_err <= inv_err + 1;
end if;
end if;
cs_n_q := cs_n;
end if;
end process;
stim : process is
procedure chk_b (what : string; g, e : std_logic) is
begin
if g /= e then
report "FAIL " & what severity error;
errors <= errors + 1;
end if;
end procedure;
procedure chk_v (what : string; g, e : std_logic_vector) is
begin
if g /= e then
report "FAIL " & what & ": got 0x" & to_hstring(g)
& " exp 0x" & to_hstring(e) severity error;
errors <= errors + 1;
end if;
end procedure;
procedure chk_n (what : string; g, e : natural) is
begin
if g /= e then
report "FAIL " & what & ": got " & integer'image(g)
& " exp " & integer'image(e) severity error;
errors <= errors + 1;
end if;
end procedure;
-- src_ready is combinational; settle before sampling it.
procedure offer (b : std_logic_vector(W - 1 downto 0)) is
begin
src_data <= b;
src_valid <= '1';
wait for 1 ns;
while src_ready /= '1' loop
wait until falling_edge(clk);
wait for 1 ns;
end loop;
wait until rising_edge(clk); -- the edge that consumes it
wait until falling_edge(clk);
src_valid <= '0';
end procedure;
procedure byte_time is
begin
for i in 0 to 2 loop
wait until falling_edge(clk);
end loop;
byte_done <= '1';
wait until falling_edge(clk);
byte_done <= '0';
wait until falling_edge(clk);
end procedure;
begin
for i in 0 to 2 loop
wait until falling_edge(clk);
end loop;
rst_n <= '1';
wait until falling_edge(clk);
chk_b("idle: cs_n high", cs_n, '1');
chk_b("idle: no clock", sclk_en, '0');
start <= '1'; wait until falling_edge(clk);
start <= '0'; wait until falling_edge(clk);
chk_b("frame open", cs_n, '0');
chk_b("not clocking while dry", sclk_en, '0');
offer(x"A1");
for i in 0 to 1 loop wait until falling_edge(clk); end loop;
chk_b("clocking once primed", sclk_en, '1');
chk_v("first byte loaded", tx_byte, x"A1");
src_data <= x"B2"; src_valid <= '1';
byte_time;
wait until falling_edge(clk); src_valid <= '0';
wait until falling_edge(clk);
chk_n("no underrun when fed in time", underruns, 0);
chk_b("still clocking", sclk_en, '1');
chk_v("next byte loaded", tx_byte, x"B2");
byte_time;
for i in 0 to 1 loop wait until falling_edge(clk); end loop;
chk_n("underrun reported", underruns, 1);
chk_b("clock stopped", sclk_en, '0');
chk_b("CS STILL LOW", cs_n, '0');
chk_b("still busy", busy, '1');
for i in 0 to 39 loop wait until falling_edge(clk); end loop;
chk_b("stall persists harmlessly", sclk_en, '0');
chk_b("CS still low after long stall", cs_n, '0');
chk_n("no spurious underruns", underruns, 1);
offer(x"C3");
for i in 0 to 1 loop wait until falling_edge(clk); end loop;
chk_b("resumed clocking", sclk_en, '1');
chk_v("resumed byte", tx_byte, x"C3");
stop <= '1';
byte_time;
for i in 0 to 2 loop wait until falling_edge(clk); end loop;
stop <= '0';
wait until falling_edge(clk);
chk_b("frame closed", cs_n, '1');
chk_b("clock off", sclk_en, '0');
chk_n("cs rose once", cs_rises, 1);
chk_b("not busy", busy, '0');
chk_n("three bytes consumed", taken, 3);
chk_v("byte 0", sent(0), x"A1");
chk_v("byte 1", sent(1), x"B2");
chk_v("byte 2", sent(2), x"C3");
chk_n("CS never deasserted mid-frame", inv_err, 0);
if errors = 0 then
report "PASS: the frame stays open across an arbitrarily long clock stall, "
& "the underrun is reported, and streaming resumes with no byte lost "
& "or repeated" severity note;
else
report "FAILED with " & integer'image(errors) & " error(s)" severity error;
end if;
halt <= true;
wait;
end process;
watchdog : process is
begin
wait for 500 us;
if not halt then
report "FAIL: watchdog timeout" severity failure;
end if;
wait;
end process;
end architecture tb;Parity
All three implement the same machine: identical ports, asynchronous active-low reset to idle with CS high and the clock disabled, a single want_byte term feeding src_ready, CS asserted before clocking begins, an underrun strobe on a dry byte boundary, cs_n untouched in the stall state, and resumption with no byte lost or repeated. All three testbenches run the same sequence — prime, seamless byte, underrun, forty-cycle stall, resume, close — and verify the same three bytes were consumed in order.
6. The UVM Streaming Sequence
Streaming is where a sequence stops being a list of transactions and starts modelling a rate.
// A conventional sequence sends items as fast as the driver accepts them,
// which means the DUT is never starved and the stall path is NEVER
// exercised. Streaming needs a sequence that models a source with gaps.
class spi_stream_seq extends uvm_sequence #(spi_stream_item);
`uvm_object_utils(spi_stream_seq)
rand int unsigned n_bytes;
rand int unsigned gap_probability; // percent chance of a gap per byte
rand int unsigned max_gap_cycles;
constraint c_shape {
n_bytes inside {[1:1024]};
gap_probability inside {[0:60]};
max_gap_cycles inside {[1:500]};
}
task body();
for (int i = 0; i < n_bytes; i++) begin
spi_stream_item it = spi_stream_item::type_id::create("it");
start_item(it);
// The gap BEFORE the byte is the stimulus. A zero gap is the
// seamless case; a long one forces the controller to stall.
assert (it.randomize() with {
pre_gap_cycles dist {
0 := 100 - gap_probability,
[1:max_gap_cycles] := gap_probability
};
});
finish_item(it);
end
endtask
endclassThe point is the pre_gap_cycles distribution. A sequence that always supplies data immediately produces a perfectly valid test in which the stall logic never executes — and since the stall logic is the entire reason this module exists, such a suite verifies everything except the thing that matters.
This generalises well beyond SPI: when a design's purpose is to handle a shortage, the stimulus must be able to create one, and the shortage must be a randomised dimension rather than a directed test appended at the end.
7. Why a Verification Engineer Cares
The properties are about what must not happen during a stall:
// 1. THE property. CS may not rise while the controller considers itself
// mid-frame. Everything else in this module is in service of this.
a_cs_held : assert property (
@(posedge clk) disable iff (!rst_n) busy |-> !cs_n)
else $error("CS deasserted while a frame was in progress");
// 2. No clocking without data. Clocking a stale register transmits bytes
// the device accepts as real -- the silent failure of §2.
a_no_blind_clocking : assert property (
@(posedge clk) disable iff (!rst_n)
(state == ST_PRIME || state == ST_STALL) |-> !sclk_en)
else $error("clock enabled while waiting for data");
// 3. A byte is consumed at most once. A valid/ready interface that
// double-consumes silently duplicates data.
a_single_consume : assert property (
@(posedge clk) disable iff (!rst_n) src_ready |=> !src_ready)
else $error("src_ready asserted on consecutive cycles");
// 4. One underrun strobe per stall, not one per cycle. Catches a level
// masquerading as an event.
a_underrun_is_pulse : assert property (
@(posedge clk) disable iff (!rst_n) underrun |=> !underrun)
else $error("underrun held for more than one cycle");What these prove. That the frame is held across stalls, that the clock never runs without data behind it, and that the handshake neither drops nor duplicates bytes.
What they do not prove is that a stall is acceptable to the device. Most devices genuinely do not care — they have no timebase. But some do: a device with an internal timeout, an analogue sample-and-hold that droops, or a self-refresh interval can be damaged by a long pause mid-frame. That is a datasheet fact, and it is Chapter 4.1 §7's boundary appearing in a new place — the bus permits the stall, and the device may not.
Coverage must target the rate, not the data:
covergroup spi_stream_cg @(posedge byte_done);
cp_gap : coverpoint pre_gap_cycles {
bins seamless = {0}; // no stall -- the easy path
bins tiny = {[1:3]}; // stall of a cycle or two
bins short = {[4:64]};
bins long = {[65:$]}; // stall longer than a byte time
}
cp_frame_len : coverpoint bytes_in_frame {
bins one = {1};
bins small = {[2:16]};
bins large = {[17:1024]};
bins huge = {[1025:$]}; // where counter widths get tested
}
// A stall in the FIRST byte exercises ST_PRIME; a stall later
// exercises ST_STALL. They are different states and different bugs.
cp_stall_position : coverpoint stall_byte_index {
bins before_first = {0};
bins mid_stream = {[1:$]};
}
x_gap_position : cross cp_gap, cp_stall_position;
endgroupcp_stall_position is the distinction worth drawing. A source that is dry when the frame opens exercises ST_PRIME; a source that runs dry at byte 500 exercises ST_STALL. They look like the same condition and are handled by different states, so a suite that only ever starts with data ready has tested one of them.
8. Why an FPGA or ASIC Engineer Cares
FIFO depth is the design parameter, and it follows from an arithmetic. A stall costs nothing functionally but costs throughput, so the question is how deep a buffer prevents them. If the upstream source can be unavailable for t_gap and a byte time is 8 × T_sclk, the FIFO needs ceil(t_gap / (8 × T_sclk)) entries to ride through. At 50 MHz a byte time is 160 ns, so tolerating a 10 µs DMA arbitration delay needs about 63 bytes of buffering. That number is computable at design time and is far more useful than choosing a depth by instinct.
sclk_en gates a clock generator, not a clock. The output here enables a divider (Chapter 2.1); it must not be ANDed with a clock net to produce SCLK. Gating a clock combinationally produces glitches and is unanalysable by static timing. The correct structure is a generator that stops on a synchronous enable — which is also why stopping cleanly leaves SCLK parked at its CPOL idle level rather than at an arbitrary point.
Stopping the clock must leave it at the idle level. If the generator halts mid-period, SCLK rests at the wrong polarity, and a device that cares (Chapter 3.1 §4) will see a spurious edge when clocking resumes. The stall must therefore complete the current period before halting — a detail that belongs in the generator and is easy to miss because simulation with an ideal model shows nothing wrong.
On an ASIC, think about what a long stall does to the device. The bus permits an arbitrary pause; the silicon may not. A device holding an analogue value, or one with a charge-pump that needs periodic activity, has a real maximum. That constraint does not appear in any SPI document and must come from the datasheet.
9. Failure Signature — Corruption That Appears Only Under System Load
Symptom. A design streams a display buffer or a firmware image over SPI. It works perfectly in isolation. When other subsystems are active — a network interface, a DMA-heavy task, a burst of interrupts — the transferred data develops corruption: occasional bytes are wrong, and the wrong bytes are usually repeats of the preceding byte.
What "repeats of the preceding byte" tells you. This is the decisive observation and it names the mechanism almost by itself. Duplicated bytes are not corruption in the usual sense — nothing was garbled, and no bit flipped. A byte was transmitted twice because the shift register was re-clocked with its previous contents. That is precisely what happens when a master keeps clocking with a dry source.
Plausible mechanisms.
- The controller keeps clocking on underrun instead of stalling — §2's wrong option, implemented by accident.
- The transmit FIFO is too shallow for the system's worst-case latency, so underruns occur at all.
- A DMA engine is losing arbitration for long enough to starve the interface.
- The software feeding the FIFO is being preempted by higher-priority work.
The discriminating observation. Correlate with load, then look at the rate. If corruption frequency rises with system activity and the corrupt bytes are duplicates, it is an underrun. If it is load-correlated but the bytes are garbled rather than duplicated, suspect electrical coupling from the other active subsystem instead — a genuinely different fault with a similar trigger.
Then check whether the controller even has a stall capability. Many simple SPI peripherals do not: they clock whatever is in the register, and the only defence is never letting the FIFO run dry. Knowing which kind of controller you have determines whether the fix is a deeper FIFO or a different controller.
Why the investigation goes wrong. Because load-correlated faults suggest electrical interference, and the search moves to grounding and layout. The duplicate-byte signature is the evidence that separates them, and it is visible in the received data without any instrument — but only if someone compares the corrupt bytes against their neighbours rather than against the expected values.
10. Common Misconceptions
11. Reason It Through
Work this before reading the answer.
An FPGA streams 4 KB frames to a display over SPI at 40 MHz, fed from DDR by a DMA engine through a 16-byte FIFO. The display shows occasional horizontal streaks. Measurements show the DMA engine's worst-case latency, when the video scaler is also active, is 12 µs.
Is the FIFO adequate, and what is the fix?
Compute the byte time. At 40 MHz, T_sclk = 25 ns, so one byte is 8 × 25 = 200 ns.
Compute what the FIFO buys. Sixteen bytes at 200 ns each is 3.2 µs of ride-through — the time the interface can keep streaming with no new data.
Compare against the worst case. The DMA can be absent for 12 µs. The FIFO covers just over a quarter of that, so underruns are not merely possible — they are guaranteed whenever the scaler contends, which is exactly the load correlation the symptom describes.
Why streaks specifically? Because the display consumes a raster. A duplicated or delayed byte shifts every subsequent pixel in that line, so a single underrun corrupts a horizontal run rather than a single pixel. The shape of the artifact tells you the failure is in the stream rather than in the pixel data — which is worth noticing, because a per-pixel fault would look like noise instead.
The required depth.
entries needed = ceil( t_gap / byte_time )
= ceil( 12 µs / 200 ns )
= 60 bytesWith margin for a worse-than-measured case, 64 or 128 bytes is the right choice — and 128 costs almost nothing on an FPGA, where a FIFO of this size is a fraction of one block RAM.
Is a deeper FIFO the only option? No, and the alternatives are worth weighing.
Stall on underrun — if the controller supports it, this converts corruption into a small throughput loss, which for a display refresh is very likely acceptable. It is the more robust fix because it degrades gracefully rather than failing at a threshold, and it protects against a latency worse than the one measured.
Raise the DMA's priority so it cannot be starved for 12 µs. This fixes the cause but trades against whatever was contending.
Lower SCLK to lengthen the byte time. At 20 MHz the same 16-byte FIFO buys 6.4 µs — still insufficient, and it halves throughput. A poor trade here.
The best answer combines two. Deepen the FIFO so stalls are rare, and ensure the controller stalls rather than clocking blind so that an unanticipated latency spike degrades throughput instead of corrupting the frame. Depth handles the expected case; stalling handles the case you did not measure.
The general lesson. A streaming interface has a ride-through time — depth × byte_time — that must exceed the source's worst-case latency. That is a two-line calculation, it is rarely done, and it converts "add a bigger FIFO and see" into a number with a justification.
12. Understanding Check
13. Summary
A continuous transfer is one transaction that happens to be long, not many short ones. The device's phase state, address pointer and output enable all persist for the whole frame, and none of it is re-evaluated at byte boundaries — which the device cannot see, because it counts edges from CS.
SPI has no timeout, so the master may legally stop the clock mid-frame and resume arbitrarily later. The device waits and notices nothing, because it does nothing between edges. Stalling is therefore a genuine solution rather than a delaying tactic.
The three responses to a dry data source are not equivalent. Stalling is correct. Clocking blind transmits stale bytes the device accepts as real, silently. Deasserting CS destroys the transaction's state, and a continuous transfer cannot be resumed from the middle.
The asymmetry with I²C is that a slave there can stall the master by stretching the clock; on SPI the device has no way to ask for time at all.
In RTL the controller is a small machine whose entire discipline is that no path from the stall state touches CS, plus a priming state so the clock never starts before real data exists, plus a single want_byte term so the handshake cannot drift between states.
For verification the stimulus must be able to create a shortage — a sequence that always has data ready never executes the stall path — and the coverage must distinguish a source dry at the frame's start from one dry mid-stream, since those are different states.
For implementation the number that matters is ride-through time, depth × byte_time, which must exceed the source's worst-case latency. And when a stream corrupts under load, duplicated bytes mean an underrun while garbled bytes mean interference.
14. What Comes Next
This chapter held the frame open and streamed bytes without asking where they were going. Chapter 7.2 — Burst Transfers and Auto-Increment supplies that: how a device advances its own address through a burst, why read bursts and write bursts advance under different rules — reads usually crossing page boundaries freely while writes wrap within a page — and the address generator that has to implement all of it, in three HDLs.
Continue learning
Related tutorials
- Related topic
CS-to-SCLK and SCLK-to-CS Timing
Chip select has timing requirements of its own: the lead before the first clock edge, the lag after the last, and the minimum deselect between transactions. Why violating them breaks a transfer whose every SCLK edge was correct.
- Related topic
Anatomy of a Read Transaction
The four phases of a device read, why a read cannot be a write reversed, when the slave takes and releases MISO, the obligation to have the first data bit valid before any edge can launch it, and the slave read datapath in three HDLs.
- Related topic
How to Read an SPI Device Datasheet
A repeatable four-pass route through any SPI peripheral datasheet — pins, timing, commands, registers — ending in the seven numbers a controller needs, and the register block that validates them atomically so no half-applied profile can ever be in force.
- Related topic
From Datasheet to Transaction Specification
A converter datasheet worked end to end into a mode, a divisor, a phase schedule and an executable element sequence — with the specification engine that generates it from a profile and the UVM environment that verifies a device described entirely by data.
