Skip to content
VLSI Mentor

SPI · Module 7

Continuous Transfers Under One CS

Holding chip select low across many bytes and what the device assumes: why the frame is the transaction, why the master may legally stop the clock when its data runs dry, and the streaming controller in three HDLs.

Modules 5 and 6 each followed a transaction of known length: a write of one register, a read of a few bytes. Both ended when the master decided it had finished. This module asks what happens when transfers stop being small.

Chip select goes low and stays low for a thousand byte times. What does the device believe is happening — and what is the master obliged to do for the entire duration?

The second half has an answer that surprises people, and it follows from something SPI does not have.

1. One Frame, Many Bytes, No Boundaries

Chapter 4.2 §2 established that CS delimits the frame and the master's bit counter is private bookkeeping. A continuous transfer takes that to its conclusion: CS falls once, a large number of bytes cross, and CS rises once.

From the device's side there is one transaction. Not a thousand transactions sharing a frame — one. That matters because everything the device tracks is scoped to the frame:

  • Its phase sequencer decoded one command and has been in the data phase ever since (Chapter 4.4).
  • Its address pointer has been advancing from the single address it was given (Chapter 5.3).
  • Its output enable has been asserted since the data phase opened (Chapter 6.1).

None of that is re-evaluated at byte boundaries, because the device cannot see byte boundaries. It counts edges. Byte 700 is distinguished from byte 699 only by its position in a count that began at CS.

2. The Master Owns the Clock — and May Stop It

Here is the fact that shapes the rest of this chapter.

SPI has no timeout. There is no minimum data rate, no maximum gap between edges, and no mechanism by which a device can notice that time has passed. The device is a shift register with combinational control: it responds to edges, and between edges it does nothing at all.

So if the master's data source runs dry mid-frame, its options are:

Stop clocking. Hold CS low, stop producing SCLK edges, and resume when data is available. The device waits indefinitely and notices nothing. This is correct.

Keep clocking with whatever is in the register. This transmits a stale byte the device will accept as real data. It is silent, unreported, and indistinguishable from intended traffic (Chapter 5.1 §5). This is wrong and is the bug §10 describes.

Deassert CS. This ends the transaction. The device discards its state, and there is no way to resume a continuous transfer from the middle — the next frame starts from a fresh command. This is catastrophic for a long transfer and merely wasteful for a short one.

The asymmetry with I²C is worth naming. I²C gives the slave a way to stall the master — clock stretching, where the slave holds SCL low until ready. SPI gives the master a way to stall itself and gives the slave nothing. The device cannot ask for time; it can only be given it.

3. The Stall

The clock may stop; the selection may not

10 cycles
A continuous SPI transfer. Chip select is low throughout. SCLK toggles, then holds low for three bit times while a source-ready signal is deasserted, then resumes toggling. Chip select never returns high during the pause.clock stopsclock stopsresumesresumescs_nsclksrc okt0t1t2t3t4t5t6t7t8t9
Figure 1 — a continuous transfer interrupted by an underrun. The clock stops for three bit times while the master's source is dry, then resumes. Chip select does not move: the transaction is still in progress and the device is still selected. From the device's side nothing happened at all, because nothing happens between edges.

The stall in the figure is three bit times. It could equally be three milliseconds. Nothing about the device's behaviour changes with the length of the pause, which is the practical consequence of having no timeout, and it is what makes stalling a genuine solution rather than a delaying tactic.

4. The Transaction

A sequence diagram of a continuous SPI transfer. The master asserts chip select, sends an opcode and address, streams several data bytes, pauses the clock during an underrun, resumes streaming, and finally releases chip select.One command, one address, an interrupted streammasterdeviceCS asserts — frameopensopcode + addressdata bytes streamclock stops — sourcedrystreaming resumesCS releases — framecloses
Figure 2 — a streaming write. One opcode and one address serve the entire burst, which is the economic reason to hold the frame open. The stall is drawn as a master-side event because the master is the only party that can cause or resolve it; the device is unaware of it.

5. Building the Streaming Controller — Three HDLs

The circuit

Circuit. A five-state machine mediating between a data source with a valid/ready handshake and the byte engine, with authority over both CS and the clock enable.

State. Which of five conditions the frame is in: idle, primed-but-not-yet-clocking, running, stalled, or closing.

Datapath. One registered byte, taken from the source and handed to the byte engine.

Control. The want_byte term is expressed once and covers all three states that need data. Writing the handshake condition separately in each state is how a valid/ready interface drifts out of agreement with itself — a byte consumed in one state and not another, or consumed twice.

Clock. The system clock. sclk_en gates the SPI clock generator; it is a control output, not a gated clock inside this module.

Reset. Asynchronous, active-low, to idle with CS deasserted and the clock disabled — the safe state, since a slave must not see a frame open out of reset (Chapter 5.2 §6).

Enables. The ST_PRIME state exists specifically so that CS asserts before the clock starts. Opening the frame and immediately clocking would transmit whatever the register held; priming guarantees the first byte is real.

Timing. underrun is a single-cycle strobe, not a level, so software sees one event per stall rather than a condition to poll.

Synthesis. Three state bits, one WIDTH register, and a handful of gates. Trivially small — the value here is entirely in the control discipline.

Limitations. One byte of buffering. A real controller places a FIFO upstream, which is §8's subject, and the depth of that FIFO is what determines whether stalls happen at all.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_stream_ctrl.sv — a frame that outlives its data source
// spi_stream_ctrl.sv — a continuous transfer under one chip select.
//
// The idea that makes this module interesting: SPI has no timeout and no
// minimum data rate. The master owns the clock, so if its data source runs
// dry mid-frame the correct response is to STOP CLOCKING and hold CS low.
// The device simply waits -- it has no way to know or care that time passed.
//
// The wrong response is to keep clocking with whatever is in the register,
// which transmits stale data the device will accept as real. That failure is
// silent, because nothing on an SPI bus reports anything (Chapter 5.1 §5).
module spi_stream_ctrl #(
    parameter int WIDTH = 8
) (
    input  logic             clk,
    input  logic             rst_n,

    // --- control ---
    input  logic             start,       // pulse: open a streaming frame
    input  logic             stop,        // level: close after the current byte

    // --- upstream data source (FIFO-like handshake) ---
    input  logic [WIDTH-1:0] src_data,
    input  logic             src_valid,
    output logic             src_ready,   // pulse: byte consumed this cycle

    // --- byte engine ---
    input  logic             byte_done,   // one pulse per transmitted byte
    output logic [WIDTH-1:0] tx_byte,

    // --- bus ---
    output logic             cs_n,
    output logic             sclk_en,     // gates the clock generator
    output logic             underrun,    // pulse: had to stall for data
    output logic             busy
);
    typedef enum logic [2:0] {
        ST_IDLE,   // no frame
        ST_PRIME,  // CS low, waiting for the first byte -- not yet clocking
        ST_RUN,    // clocking
        ST_STALL,  // CS low, clock stopped, waiting for data
        ST_CLOSE   // finish the frame
    } state_t;

    state_t state;

    // Consume a byte exactly when we are in a state that needs one and the
    // source has it. Expressed once so the handshake cannot drift.
    logic want_byte;
    assign want_byte = (state == ST_PRIME)
                    || (state == ST_STALL)
                    || (state == ST_RUN && byte_done && !stop);
    assign src_ready = want_byte && src_valid;

    always_ff @(posedge clk or negedge rst_n) begin
        if (!rst_n) begin
            state    <= ST_IDLE;
            cs_n     <= 1'b1;
            sclk_en  <= 1'b0;
            tx_byte  <= '0;
            underrun <= 1'b0;
        end else begin
            underrun <= 1'b0;                 // strobe

            case (state)
                ST_IDLE: if (start) begin
                    cs_n    <= 1'b0;          // frame opens and stays open
                    sclk_en <= 1'b0;          // but do not clock without data
                    state   <= ST_PRIME;
                end

                ST_PRIME: if (src_valid) begin
                    tx_byte <= src_data;
                    sclk_en <= 1'b1;
                    state   <= ST_RUN;
                end

                ST_RUN: if (byte_done) begin
                    if (stop) begin
                        state <= ST_CLOSE;
                    end else if (src_valid) begin
                        tx_byte <= src_data;  // seamless: no gap in clocking
                    end else begin
                        // UNDERRUN. Stop the clock; CS stays low. The device
                        // is mid-transaction and must remain selected.
                        sclk_en  <= 1'b0;
                        underrun <= 1'b1;
                        state    <= ST_STALL;
                    end
                end

                ST_STALL: if (src_valid) begin
                    tx_byte <= src_data;
                    sclk_en <= 1'b1;          // resume exactly where we stopped
                    state   <= ST_RUN;
                end

                ST_CLOSE: begin
                    cs_n    <= 1'b1;
                    sclk_en <= 1'b0;
                    state   <= ST_IDLE;
                end

                default: state <= ST_IDLE;
            endcase
        end
    end

    assign busy = (state != ST_IDLE);
endmodule

The ST_STALL state is the module's whole point, and note what it does not touch: cs_n is left alone. There is no path in this machine from ST_STALL to anything that raises CS. That is deliberate and structural — the one mistake this design exists to prevent is ending a transaction because data ran out.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_stream_ctrl_tb.sv — streaming, an underrun, and a very long stall
// spi_stream_ctrl_tb.sv — seamless streaming, an underrun mid-frame, and the
// property that matters most: CS stays low while the clock is stopped.
`timescale 1ns/1ps
module spi_stream_ctrl_tb;
    logic clk = 0, rst_n = 0;
    always #5 clk = ~clk;

    localparam int W = 8;
    logic start = 0, stop = 0, src_valid = 0, byte_done = 0;
    logic [W-1:0] src_data = '0;
    logic src_ready, cs_n, sclk_en, underrun, busy;
    logic [W-1:0] tx_byte;

    spi_stream_ctrl #(.WIDTH(W)) dut (
        .clk, .rst_n, .start, .stop, .src_data, .src_valid, .src_ready,
        .byte_done, .tx_byte, .cs_n, .sclk_en, .underrun, .busy);

    int errors = 0, underruns = 0, cs_rises = 0, taken = 0;
    logic [W-1:0] sent [$];
    logic cs_n_q;

    task automatic chk(input string what, input int g, input int e);
        if (g !== e) begin $display("FAIL %s: got 0x%0h exp 0x%0h", what, g, e); errors++; end
    endtask

    // Record every accepted byte at the edge that accepts it.
    always @(posedge clk) if (rst_n) begin
        cs_n_q <= cs_n;
        if (cs_n && !cs_n_q) cs_rises++;
        if (underrun) underruns++;
        if (src_ready) begin sent.push_back(src_data); taken++; end
        // CRITICAL INVARIANT: CS must never be high while a frame is in
        // progress. The clock may stop; the selection may not.
        if (busy && cs_n) begin
            $display("FAIL: CS deasserted mid-frame at t=%0t", $time); errors++;
        end
    end

    // Present a byte and wait until the controller actually takes it.
    // src_ready is combinational, so settle before sampling it.
    task automatic offer(input logic [W-1:0] b);
        src_data = b; src_valid = 1;
        #1;
        while (!src_ready) begin @(negedge clk); #1; end
        @(posedge clk);                    // the edge that consumes it
        @(negedge clk); src_valid = 0;
    endtask

    // One transmitted byte time.
    task automatic byte_time();
        repeat (3) @(negedge clk);
        byte_done = 1; @(negedge clk); byte_done = 0; @(negedge clk);
    endtask

    initial begin
        repeat (3) @(negedge clk); rst_n = 1; @(negedge clk);
        chk("idle: cs_n high", cs_n,    1);
        chk("idle: no clock",  sclk_en, 0);

        // --- open a frame; the clock must NOT start before data exists ---
        start = 1; @(negedge clk); start = 0; @(negedge clk);
        chk("frame open",             cs_n,    0);
        chk("not clocking while dry", sclk_en, 0);

        offer(8'hA1);
        repeat (2) @(negedge clk);
        chk("clocking once primed", sclk_en, 1);
        chk("first byte loaded",    tx_byte, 8'hA1);

        // --- seamless: hold the next byte valid across the byte boundary ---
        src_data = 8'hB2; src_valid = 1;
        byte_time();
        @(negedge clk); src_valid = 0; @(negedge clk);
        chk("no underrun when fed in time", underruns, 0);
        chk("still clocking",               sclk_en,   1);
        chk("next byte loaded",             tx_byte,   8'hB2);

        // --- UNDERRUN: let this byte finish with nothing available ---
        byte_time();
        repeat (2) @(negedge clk);
        chk("underrun reported", underruns, 1);
        chk("clock stopped",     sclk_en,   0);
        chk("CS STILL LOW",      cs_n,      0);   // the property that matters
        chk("still busy",        busy,      1);

        // --- the stall may last arbitrarily long; SPI has no timeout ---
        repeat (40) @(negedge clk);
        chk("stall persists harmlessly",     sclk_en, 0);
        chk("CS still low after long stall", cs_n,    0);
        chk("no spurious underruns",         underruns, 1);

        // --- resume exactly where it stopped ---
        offer(8'hC3);
        repeat (2) @(negedge clk);
        chk("resumed clocking", sclk_en, 1);
        chk("resumed byte",     tx_byte, 8'hC3);

        // --- close the frame ---
        stop = 1;
        byte_time();
        repeat (3) @(negedge clk);
        stop = 0; @(negedge clk);
        chk("frame closed", cs_n,     1);
        chk("clock off",    sclk_en,  0);
        chk("cs rose once", cs_rises, 1);
        chk("not busy",     busy,     0);

        // --- exactly three bytes consumed, in order, none repeated ---
        chk("three bytes consumed", taken, 3);
        chk("byte 0", sent[0], 8'hA1);
        chk("byte 1", sent[1], 8'hB2);
        chk("byte 2", sent[2], 8'hC3);

        if (errors == 0)
            $display("PASS: the frame stays open across an arbitrarily long clock stall, the underrun is reported, and streaming resumes with no byte lost or repeated");
        else
            $display("FAILED with %0d error(s)", errors);
        $finish;
    end

    initial begin #500000; $display("FAIL: watchdog timeout"); $finish; end
endmodule

Two things in that testbench are doing real work.

The continuous invariant check — if (busy && cs_n) sampled every clock — asserts that CS is never high while a frame is in progress. It runs throughout the whole simulation rather than at chosen points, which is what makes it a proof about the design rather than a spot check.

The forty-cycle stall is not padding. It asserts that a long pause is harmless and produces no additional underrun strobes, catching a design that re-reports the condition every cycle and floods software with events.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_stream_ctrl.v — the same controller in Verilog-2001
// spi_stream_ctrl.v — the same continuous-transfer controller in Verilog-2001.
module spi_stream_ctrl #(
    parameter WIDTH = 8
) (
    input  wire             clk,
    input  wire             rst_n,

    input  wire             start,
    input  wire             stop,

    input  wire [WIDTH-1:0] src_data,
    input  wire             src_valid,
    output wire             src_ready,

    input  wire             byte_done,
    output reg  [WIDTH-1:0] tx_byte,

    output reg              cs_n,
    output reg              sclk_en,
    output reg              underrun,
    output wire             busy
);
    localparam ST_IDLE  = 3'd0,
               ST_PRIME = 3'd1,
               ST_RUN   = 3'd2,
               ST_STALL = 3'd3,
               ST_CLOSE = 3'd4;

    reg [2:0] state;
    wire      want_byte;

    // Expressed once so the handshake cannot drift between states.
    assign want_byte = (state == ST_PRIME)
                    || (state == ST_STALL)
                    || (state == ST_RUN && byte_done && !stop);
    assign src_ready = want_byte && src_valid;

    always @(posedge clk or negedge rst_n) begin
        if (!rst_n) begin
            state    <= ST_IDLE;
            cs_n     <= 1'b1;
            sclk_en  <= 1'b0;
            tx_byte  <= {WIDTH{1'b0}};
            underrun <= 1'b0;
        end else begin
            underrun <= 1'b0;

            case (state)
                ST_IDLE: if (start) begin
                    cs_n    <= 1'b0;       // frame opens and stays open
                    sclk_en <= 1'b0;       // but do not clock without data
                    state   <= ST_PRIME;
                end

                ST_PRIME: if (src_valid) begin
                    tx_byte <= src_data;
                    sclk_en <= 1'b1;
                    state   <= ST_RUN;
                end

                ST_RUN: if (byte_done) begin
                    if (stop) begin
                        state <= ST_CLOSE;
                    end else if (src_valid) begin
                        tx_byte <= src_data;
                    end else begin
                        // UNDERRUN. Stop the clock; CS stays low.
                        sclk_en  <= 1'b0;
                        underrun <= 1'b1;
                        state    <= ST_STALL;
                    end
                end

                ST_STALL: if (src_valid) begin
                    tx_byte <= src_data;
                    sclk_en <= 1'b1;
                    state   <= ST_RUN;
                end

                ST_CLOSE: begin
                    cs_n    <= 1'b1;
                    sclk_en <= 1'b0;
                    state   <= ST_IDLE;
                end

                default: state <= ST_IDLE;
            endcase
        end
    end

    assign busy = (state != ST_IDLE);
endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_stream_ctrl_tb.v — the same checks in Verilog-2001
// spi_stream_ctrl_tb.v — the same checks in Verilog-2001.
`timescale 1ns/1ps
module spi_stream_ctrl_tb;
    reg clk = 0, rst_n = 0;
    always #5 clk = ~clk;

    parameter W = 8;
    reg start = 0, stop = 0, src_valid = 0, byte_done = 0;
    reg [W-1:0] src_data = 0;
    wire src_ready, cs_n, sclk_en, underrun, busy;
    wire [W-1:0] tx_byte;

    spi_stream_ctrl #(.WIDTH(W)) dut (
        .clk(clk), .rst_n(rst_n), .start(start), .stop(stop), .src_data(src_data),
        .src_valid(src_valid), .src_ready(src_ready), .byte_done(byte_done),
        .tx_byte(tx_byte), .cs_n(cs_n), .sclk_en(sclk_en), .underrun(underrun),
        .busy(busy));

    integer errors = 0, underruns = 0, cs_rises = 0, taken = 0;
    reg [W-1:0] sent [0:7];
    reg cs_n_q;

    task chk;
        input [80*8-1:0] what;
        input [31:0] g, e;
        begin
            if (g !== e) begin
                $display("FAIL %0s: got 0x%0h exp 0x%0h", what, g, e);
                errors = errors + 1;
            end
        end
    endtask

    always @(posedge clk) if (rst_n) begin
        cs_n_q <= cs_n;
        if (cs_n && !cs_n_q) cs_rises = cs_rises + 1;
        if (underrun) underruns = underruns + 1;
        if (src_ready) begin sent[taken] = src_data; taken = taken + 1; end
        if (busy && cs_n) begin
            $display("FAIL: CS deasserted mid-frame at t=%0t", $time);
            errors = errors + 1;
        end
    end

    task offer;
        input [W-1:0] b;
        begin
            src_data = b; src_valid = 1;
            #1;
            while (!src_ready) begin @(negedge clk); #1; end
            @(posedge clk);
            @(negedge clk); src_valid = 0;
        end
    endtask

    task byte_time;
        begin
            repeat (3) @(negedge clk);
            byte_done = 1; @(negedge clk); byte_done = 0; @(negedge clk);
        end
    endtask

    initial begin
        repeat (3) @(negedge clk); rst_n = 1; @(negedge clk);
        chk("idle: cs_n high", cs_n,    1);
        chk("idle: no clock",  sclk_en, 0);

        start = 1; @(negedge clk); start = 0; @(negedge clk);
        chk("frame open",             cs_n,    0);
        chk("not clocking while dry", sclk_en, 0);

        offer(8'hA1);
        repeat (2) @(negedge clk);
        chk("clocking once primed", sclk_en, 1);
        chk("first byte loaded",    tx_byte, 8'hA1);

        src_data = 8'hB2; src_valid = 1;
        byte_time;
        @(negedge clk); src_valid = 0; @(negedge clk);
        chk("no underrun when fed in time", underruns, 0);
        chk("still clocking",               sclk_en,   1);
        chk("next byte loaded",             tx_byte,   8'hB2);

        byte_time;
        repeat (2) @(negedge clk);
        chk("underrun reported", underruns, 1);
        chk("clock stopped",     sclk_en,   0);
        chk("CS STILL LOW",      cs_n,      0);
        chk("still busy",        busy,      1);

        repeat (40) @(negedge clk);
        chk("stall persists harmlessly",     sclk_en,   0);
        chk("CS still low after long stall", cs_n,      0);
        chk("no spurious underruns",         underruns, 1);

        offer(8'hC3);
        repeat (2) @(negedge clk);
        chk("resumed clocking", sclk_en, 1);
        chk("resumed byte",     tx_byte, 8'hC3);

        stop = 1;
        byte_time;
        repeat (3) @(negedge clk);
        stop = 0; @(negedge clk);
        chk("frame closed", cs_n,     1);
        chk("clock off",    sclk_en,  0);
        chk("cs rose once", cs_rises, 1);
        chk("not busy",     busy,     0);

        chk("three bytes consumed", taken, 3);
        chk("byte 0", sent[0], 8'hA1);
        chk("byte 1", sent[1], 8'hB2);
        chk("byte 2", sent[2], 8'hC3);

        if (errors == 0)
            $display("PASS: the frame stays open across an arbitrarily long clock stall, the underrun is reported, and streaming resumes with no byte lost or repeated");
        else
            $display("FAILED with %0d error(s)", errors);
        $finish;
    end

    initial begin #500000; $display("FAIL: watchdog timeout"); $finish; end
endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_stream_ctrl.vhd — the same controller in VHDL
-- spi_stream_ctrl.vhd — the same continuous-transfer controller in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

entity spi_stream_ctrl is
    generic (
        WIDTH : positive := 8
    );
    port (
        clk       : in  std_logic;
        rst_n     : in  std_logic;

        start     : in  std_logic;                                -- pulse
        stop      : in  std_logic;                                -- level

        src_data  : in  std_logic_vector(WIDTH - 1 downto 0);
        src_valid : in  std_logic;
        src_ready : out std_logic;                                -- pulse

        byte_done : in  std_logic;
        tx_byte   : out std_logic_vector(WIDTH - 1 downto 0);

        cs_n      : out std_logic;
        sclk_en   : out std_logic;                                -- gates the clock
        underrun  : out std_logic;                                -- pulse
        busy      : out std_logic
    );
end entity spi_stream_ctrl;

architecture rtl of spi_stream_ctrl is
    type state_t is (ST_IDLE, ST_PRIME, ST_RUN, ST_STALL, ST_CLOSE);
    signal state : state_t;

    signal want_byte : std_logic;
begin

    -- Expressed once so the handshake cannot drift between states.
    want_byte <= '1' when state = ST_PRIME
                       or state = ST_STALL
                       or (state = ST_RUN and byte_done = '1' and stop = '0')
                 else '0';

    src_ready <= want_byte and src_valid;
    busy      <= '0' when state = ST_IDLE else '1';

    process (clk, rst_n) is
    begin
        if rst_n = '0' then
            state    <= ST_IDLE;
            cs_n     <= '1';
            sclk_en  <= '0';
            tx_byte  <= (others => '0');
            underrun <= '0';
        elsif rising_edge(clk) then
            underrun <= '0';

            case state is
                when ST_IDLE =>
                    if start = '1' then
                        cs_n    <= '0';    -- frame opens and stays open
                        sclk_en <= '0';    -- but do not clock without data
                        state   <= ST_PRIME;
                    end if;

                when ST_PRIME =>
                    if src_valid = '1' then
                        tx_byte <= src_data;
                        sclk_en <= '1';
                        state   <= ST_RUN;
                    end if;

                when ST_RUN =>
                    if byte_done = '1' then
                        if stop = '1' then
                            state <= ST_CLOSE;
                        elsif src_valid = '1' then
                            tx_byte <= src_data;   -- seamless
                        else
                            -- UNDERRUN. Stop the clock; CS stays low.
                            sclk_en  <= '0';
                            underrun <= '1';
                            state    <= ST_STALL;
                        end if;
                    end if;

                when ST_STALL =>
                    if src_valid = '1' then
                        tx_byte <= src_data;
                        sclk_en <= '1';
                        state   <= ST_RUN;
                    end if;

                when ST_CLOSE =>
                    cs_n    <= '1';
                    sclk_en <= '0';
                    state   <= ST_IDLE;
            end case;
        end if;
    end process;

end architecture rtl;
Azvya Education Pvt. Ltd.VLSI Mentor
spi_stream_ctrl_tb.vhd — the same checks in VHDL
-- spi_stream_ctrl_tb.vhd — the same checks in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

entity spi_stream_ctrl_tb is
end entity spi_stream_ctrl_tb;

architecture tb of spi_stream_ctrl_tb is
    constant W : positive := 8;

    signal clk       : std_logic := '0';
    signal rst_n     : std_logic := '0';
    signal start     : std_logic := '0';
    signal stop      : std_logic := '0';
    signal src_valid : std_logic := '0';
    signal byte_done : std_logic := '0';
    signal src_data  : std_logic_vector(W - 1 downto 0) := (others => '0');
    signal halt      : boolean := false;

    signal src_ready : std_logic;
    signal tx_byte   : std_logic_vector(W - 1 downto 0);
    signal cs_n      : std_logic;
    signal sclk_en   : std_logic;
    signal underrun  : std_logic;
    signal busy      : std_logic;

    signal errors    : natural := 0;   -- owned by stim
    signal inv_err   : natural := 0;   -- owned by observe
    signal underruns : natural := 0;
    signal cs_rises  : natural := 0;
    signal taken     : natural := 0;

    type bytes_t is array (0 to 7) of std_logic_vector(W - 1 downto 0);
    signal sent : bytes_t := (others => (others => '0'));
begin

    clk <= not clk after 5 ns when not halt else '0';

    dut : entity work.spi_stream_ctrl
        generic map (WIDTH => W)
        port map (clk => clk, rst_n => rst_n, start => start, stop => stop,
                  src_data => src_data, src_valid => src_valid, src_ready => src_ready,
                  byte_done => byte_done, tx_byte => tx_byte, cs_n => cs_n,
                  sclk_en => sclk_en, underrun => underrun, busy => busy);

    -- One process owns the observation counters.
    observe : process (clk) is
        variable cs_n_q : std_logic := '1';
    begin
        if rising_edge(clk) then
            if rst_n = '1' then
                if cs_n = '1' and cs_n_q = '0' then
                    cs_rises <= cs_rises + 1;
                end if;
                if underrun = '1' then
                    underruns <= underruns + 1;
                end if;
                if src_ready = '1' then
                    sent(taken) <= src_data;
                    taken       <= taken + 1;
                end if;
                -- CRITICAL INVARIANT: CS must never be high mid-frame.
                if busy = '1' and cs_n = '1' then
                    report "FAIL: CS deasserted mid-frame" severity error;
                    inv_err <= inv_err + 1;
                end if;
            end if;
            cs_n_q := cs_n;
        end if;
    end process;

    stim : process is
        procedure chk_b (what : string; g, e : std_logic) is
        begin
            if g /= e then
                report "FAIL " & what severity error;
                errors <= errors + 1;
            end if;
        end procedure;

        procedure chk_v (what : string; g, e : std_logic_vector) is
        begin
            if g /= e then
                report "FAIL " & what & ": got 0x" & to_hstring(g)
                    & " exp 0x" & to_hstring(e) severity error;
                errors <= errors + 1;
            end if;
        end procedure;

        procedure chk_n (what : string; g, e : natural) is
        begin
            if g /= e then
                report "FAIL " & what & ": got " & integer'image(g)
                    & " exp " & integer'image(e) severity error;
                errors <= errors + 1;
            end if;
        end procedure;

        -- src_ready is combinational; settle before sampling it.
        procedure offer (b : std_logic_vector(W - 1 downto 0)) is
        begin
            src_data  <= b;
            src_valid <= '1';
            wait for 1 ns;
            while src_ready /= '1' loop
                wait until falling_edge(clk);
                wait for 1 ns;
            end loop;
            wait until rising_edge(clk);     -- the edge that consumes it
            wait until falling_edge(clk);
            src_valid <= '0';
        end procedure;

        procedure byte_time is
        begin
            for i in 0 to 2 loop
                wait until falling_edge(clk);
            end loop;
            byte_done <= '1';
            wait until falling_edge(clk);
            byte_done <= '0';
            wait until falling_edge(clk);
        end procedure;
    begin
        for i in 0 to 2 loop
            wait until falling_edge(clk);
        end loop;
        rst_n <= '1';
        wait until falling_edge(clk);
        chk_b("idle: cs_n high", cs_n,    '1');
        chk_b("idle: no clock",  sclk_en, '0');

        start <= '1'; wait until falling_edge(clk);
        start <= '0'; wait until falling_edge(clk);
        chk_b("frame open",             cs_n,    '0');
        chk_b("not clocking while dry", sclk_en, '0');

        offer(x"A1");
        for i in 0 to 1 loop wait until falling_edge(clk); end loop;
        chk_b("clocking once primed", sclk_en, '1');
        chk_v("first byte loaded",    tx_byte, x"A1");

        src_data <= x"B2"; src_valid <= '1';
        byte_time;
        wait until falling_edge(clk); src_valid <= '0';
        wait until falling_edge(clk);
        chk_n("no underrun when fed in time", underruns, 0);
        chk_b("still clocking",               sclk_en,   '1');
        chk_v("next byte loaded",             tx_byte,   x"B2");

        byte_time;
        for i in 0 to 1 loop wait until falling_edge(clk); end loop;
        chk_n("underrun reported", underruns, 1);
        chk_b("clock stopped",     sclk_en,   '0');
        chk_b("CS STILL LOW",      cs_n,      '0');
        chk_b("still busy",        busy,      '1');

        for i in 0 to 39 loop wait until falling_edge(clk); end loop;
        chk_b("stall persists harmlessly",     sclk_en,   '0');
        chk_b("CS still low after long stall", cs_n,      '0');
        chk_n("no spurious underruns",         underruns, 1);

        offer(x"C3");
        for i in 0 to 1 loop wait until falling_edge(clk); end loop;
        chk_b("resumed clocking", sclk_en, '1');
        chk_v("resumed byte",     tx_byte, x"C3");

        stop <= '1';
        byte_time;
        for i in 0 to 2 loop wait until falling_edge(clk); end loop;
        stop <= '0';
        wait until falling_edge(clk);
        chk_b("frame closed", cs_n,     '1');
        chk_b("clock off",    sclk_en,  '0');
        chk_n("cs rose once", cs_rises, 1);
        chk_b("not busy",     busy,     '0');

        chk_n("three bytes consumed", taken, 3);
        chk_v("byte 0", sent(0), x"A1");
        chk_v("byte 1", sent(1), x"B2");
        chk_v("byte 2", sent(2), x"C3");

        chk_n("CS never deasserted mid-frame", inv_err, 0);

        if errors = 0 then
            report "PASS: the frame stays open across an arbitrarily long clock stall, "
                 & "the underrun is reported, and streaming resumes with no byte lost "
                 & "or repeated" severity note;
        else
            report "FAILED with " & integer'image(errors) & " error(s)" severity error;
        end if;
        halt <= true;
        wait;
    end process;

    watchdog : process is
    begin
        wait for 500 us;
        if not halt then
            report "FAIL: watchdog timeout" severity failure;
        end if;
        wait;
    end process;

end architecture tb;

Parity

All three implement the same machine: identical ports, asynchronous active-low reset to idle with CS high and the clock disabled, a single want_byte term feeding src_ready, CS asserted before clocking begins, an underrun strobe on a dry byte boundary, cs_n untouched in the stall state, and resumption with no byte lost or repeated. All three testbenches run the same sequence — prime, seamless byte, underrun, forty-cycle stall, resume, close — and verify the same three bytes were consumed in order.

6. The UVM Streaming Sequence

Streaming is where a sequence stops being a list of transactions and starts modelling a rate.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_stream_seq.sv — a sequence that deliberately starves the DUT
   // A conventional sequence sends items as fast as the driver accepts them,
   // which means the DUT is never starved and the stall path is NEVER
   // exercised. Streaming needs a sequence that models a source with gaps.
   class spi_stream_seq extends uvm_sequence #(spi_stream_item);
       `uvm_object_utils(spi_stream_seq)

       rand int unsigned n_bytes;
       rand int unsigned gap_probability;   // percent chance of a gap per byte
       rand int unsigned max_gap_cycles;

       constraint c_shape {
           n_bytes         inside {[1:1024]};
           gap_probability inside {[0:60]};
           max_gap_cycles  inside {[1:500]};
       }

       task body();
           for (int i = 0; i < n_bytes; i++) begin
               spi_stream_item it = spi_stream_item::type_id::create("it");
               start_item(it);
               // The gap BEFORE the byte is the stimulus. A zero gap is the
               // seamless case; a long one forces the controller to stall.
               assert (it.randomize() with {
                   pre_gap_cycles dist {
                       0                        := 100 - gap_probability,
                       [1:max_gap_cycles]       := gap_probability
                   };
               });
               finish_item(it);
           end
       endtask
   endclass

The point is the pre_gap_cycles distribution. A sequence that always supplies data immediately produces a perfectly valid test in which the stall logic never executes — and since the stall logic is the entire reason this module exists, such a suite verifies everything except the thing that matters.

This generalises well beyond SPI: when a design's purpose is to handle a shortage, the stimulus must be able to create one, and the shortage must be a randomised dimension rather than a directed test appended at the end.

7. Why a Verification Engineer Cares

The properties are about what must not happen during a stall:

Azvya Education Pvt. Ltd.VLSI Mentor
spi_stream_ctrl.sva — the frame survives the drought
   // 1. THE property. CS may not rise while the controller considers itself
   //    mid-frame. Everything else in this module is in service of this.
   a_cs_held : assert property (
       @(posedge clk) disable iff (!rst_n) busy |-> !cs_n)
       else $error("CS deasserted while a frame was in progress");

   // 2. No clocking without data. Clocking a stale register transmits bytes
   //    the device accepts as real -- the silent failure of §2.
   a_no_blind_clocking : assert property (
       @(posedge clk) disable iff (!rst_n)
           (state == ST_PRIME || state == ST_STALL) |-> !sclk_en)
       else $error("clock enabled while waiting for data");

   // 3. A byte is consumed at most once. A valid/ready interface that
   //    double-consumes silently duplicates data.
   a_single_consume : assert property (
       @(posedge clk) disable iff (!rst_n) src_ready |=> !src_ready)
       else $error("src_ready asserted on consecutive cycles");

   // 4. One underrun strobe per stall, not one per cycle. Catches a level
   //    masquerading as an event.
   a_underrun_is_pulse : assert property (
       @(posedge clk) disable iff (!rst_n) underrun |=> !underrun)
       else $error("underrun held for more than one cycle");

What these prove. That the frame is held across stalls, that the clock never runs without data behind it, and that the handshake neither drops nor duplicates bytes.

What they do not prove is that a stall is acceptable to the device. Most devices genuinely do not care — they have no timebase. But some do: a device with an internal timeout, an analogue sample-and-hold that droops, or a self-refresh interval can be damaged by a long pause mid-frame. That is a datasheet fact, and it is Chapter 4.1 §7's boundary appearing in a new place — the bus permits the stall, and the device may not.

Coverage must target the rate, not the data:

Azvya Education Pvt. Ltd.VLSI Mentor
spi_stream_cg.sv — cover the gaps, because the gaps are the stimulus
   covergroup spi_stream_cg @(posedge byte_done);
       cp_gap : coverpoint pre_gap_cycles {
           bins seamless = {0};           // no stall -- the easy path
           bins tiny     = {[1:3]};       // stall of a cycle or two
           bins short    = {[4:64]};
           bins long     = {[65:$]};      // stall longer than a byte time
       }

       cp_frame_len : coverpoint bytes_in_frame {
           bins one     = {1};
           bins small   = {[2:16]};
           bins large   = {[17:1024]};
           bins huge    = {[1025:$]};     // where counter widths get tested
       }

       // A stall in the FIRST byte exercises ST_PRIME; a stall later
       // exercises ST_STALL. They are different states and different bugs.
       cp_stall_position : coverpoint stall_byte_index {
           bins before_first = {0};
           bins mid_stream   = {[1:$]};
       }

       x_gap_position : cross cp_gap, cp_stall_position;
   endgroup

cp_stall_position is the distinction worth drawing. A source that is dry when the frame opens exercises ST_PRIME; a source that runs dry at byte 500 exercises ST_STALL. They look like the same condition and are handled by different states, so a suite that only ever starts with data ready has tested one of them.

8. Why an FPGA or ASIC Engineer Cares

FIFO depth is the design parameter, and it follows from an arithmetic. A stall costs nothing functionally but costs throughput, so the question is how deep a buffer prevents them. If the upstream source can be unavailable for t_gap and a byte time is 8 × T_sclk, the FIFO needs ceil(t_gap / (8 × T_sclk)) entries to ride through. At 50 MHz a byte time is 160 ns, so tolerating a 10 µs DMA arbitration delay needs about 63 bytes of buffering. That number is computable at design time and is far more useful than choosing a depth by instinct.

sclk_en gates a clock generator, not a clock. The output here enables a divider (Chapter 2.1); it must not be ANDed with a clock net to produce SCLK. Gating a clock combinationally produces glitches and is unanalysable by static timing. The correct structure is a generator that stops on a synchronous enable — which is also why stopping cleanly leaves SCLK parked at its CPOL idle level rather than at an arbitrary point.

Stopping the clock must leave it at the idle level. If the generator halts mid-period, SCLK rests at the wrong polarity, and a device that cares (Chapter 3.1 §4) will see a spurious edge when clocking resumes. The stall must therefore complete the current period before halting — a detail that belongs in the generator and is easy to miss because simulation with an ideal model shows nothing wrong.

On an ASIC, think about what a long stall does to the device. The bus permits an arbitrary pause; the silicon may not. A device holding an analogue value, or one with a charge-pump that needs periodic activity, has a real maximum. That constraint does not appear in any SPI document and must come from the datasheet.

9. Failure Signature — Corruption That Appears Only Under System Load

Symptom. A design streams a display buffer or a firmware image over SPI. It works perfectly in isolation. When other subsystems are active — a network interface, a DMA-heavy task, a burst of interrupts — the transferred data develops corruption: occasional bytes are wrong, and the wrong bytes are usually repeats of the preceding byte.

What "repeats of the preceding byte" tells you. This is the decisive observation and it names the mechanism almost by itself. Duplicated bytes are not corruption in the usual sense — nothing was garbled, and no bit flipped. A byte was transmitted twice because the shift register was re-clocked with its previous contents. That is precisely what happens when a master keeps clocking with a dry source.

Plausible mechanisms.

  • The controller keeps clocking on underrun instead of stalling — §2's wrong option, implemented by accident.
  • The transmit FIFO is too shallow for the system's worst-case latency, so underruns occur at all.
  • A DMA engine is losing arbitration for long enough to starve the interface.
  • The software feeding the FIFO is being preempted by higher-priority work.

The discriminating observation. Correlate with load, then look at the rate. If corruption frequency rises with system activity and the corrupt bytes are duplicates, it is an underrun. If it is load-correlated but the bytes are garbled rather than duplicated, suspect electrical coupling from the other active subsystem instead — a genuinely different fault with a similar trigger.

Then check whether the controller even has a stall capability. Many simple SPI peripherals do not: they clock whatever is in the register, and the only defence is never letting the FIFO run dry. Knowing which kind of controller you have determines whether the fix is a deeper FIFO or a different controller.

Why the investigation goes wrong. Because load-correlated faults suggest electrical interference, and the search moves to grounding and layout. The duplicate-byte signature is the evidence that separates them, and it is visible in the received data without any instrument — but only if someone compares the corrupt bytes against their neighbours rather than against the expected values.

10. Common Misconceptions

11. Reason It Through

Work this before reading the answer.

An FPGA streams 4 KB frames to a display over SPI at 40 MHz, fed from DDR by a DMA engine through a 16-byte FIFO. The display shows occasional horizontal streaks. Measurements show the DMA engine's worst-case latency, when the video scaler is also active, is 12 µs.

Is the FIFO adequate, and what is the fix?

Compute the byte time. At 40 MHz, T_sclk = 25 ns, so one byte is 8 × 25 = 200 ns.

Compute what the FIFO buys. Sixteen bytes at 200 ns each is 3.2 µs of ride-through — the time the interface can keep streaming with no new data.

Compare against the worst case. The DMA can be absent for 12 µs. The FIFO covers just over a quarter of that, so underruns are not merely possible — they are guaranteed whenever the scaler contends, which is exactly the load correlation the symptom describes.

Why streaks specifically? Because the display consumes a raster. A duplicated or delayed byte shifts every subsequent pixel in that line, so a single underrun corrupts a horizontal run rather than a single pixel. The shape of the artifact tells you the failure is in the stream rather than in the pixel data — which is worth noticing, because a per-pixel fault would look like noise instead.

The required depth.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   entries needed = ceil( t_gap / byte_time )
                  = ceil( 12 µs / 200 ns )
                  = 60 bytes

With margin for a worse-than-measured case, 64 or 128 bytes is the right choice — and 128 costs almost nothing on an FPGA, where a FIFO of this size is a fraction of one block RAM.

Is a deeper FIFO the only option? No, and the alternatives are worth weighing.

Stall on underrun — if the controller supports it, this converts corruption into a small throughput loss, which for a display refresh is very likely acceptable. It is the more robust fix because it degrades gracefully rather than failing at a threshold, and it protects against a latency worse than the one measured.

Raise the DMA's priority so it cannot be starved for 12 µs. This fixes the cause but trades against whatever was contending.

Lower SCLK to lengthen the byte time. At 20 MHz the same 16-byte FIFO buys 6.4 µs — still insufficient, and it halves throughput. A poor trade here.

The best answer combines two. Deepen the FIFO so stalls are rare, and ensure the controller stalls rather than clocking blind so that an unanticipated latency spike degrades throughput instead of corrupting the frame. Depth handles the expected case; stalling handles the case you did not measure.

The general lesson. A streaming interface has a ride-through time — depth × byte_time — that must exceed the source's worst-case latency. That is a two-line calculation, it is rarely done, and it converts "add a bigger FIFO and see" into a number with a justification.

12. Understanding Check

13. Summary

A continuous transfer is one transaction that happens to be long, not many short ones. The device's phase state, address pointer and output enable all persist for the whole frame, and none of it is re-evaluated at byte boundaries — which the device cannot see, because it counts edges from CS.

SPI has no timeout, so the master may legally stop the clock mid-frame and resume arbitrarily later. The device waits and notices nothing, because it does nothing between edges. Stalling is therefore a genuine solution rather than a delaying tactic.

The three responses to a dry data source are not equivalent. Stalling is correct. Clocking blind transmits stale bytes the device accepts as real, silently. Deasserting CS destroys the transaction's state, and a continuous transfer cannot be resumed from the middle.

The asymmetry with I²C is that a slave there can stall the master by stretching the clock; on SPI the device has no way to ask for time at all.

In RTL the controller is a small machine whose entire discipline is that no path from the stall state touches CS, plus a priming state so the clock never starts before real data exists, plus a single want_byte term so the handshake cannot drift between states.

For verification the stimulus must be able to create a shortage — a sequence that always has data ready never executes the stall path — and the coverage must distinguish a source dry at the frame's start from one dry mid-stream, since those are different states.

For implementation the number that matters is ride-through time, depth × byte_time, which must exceed the source's worst-case latency. And when a stream corrupts under load, duplicated bytes mean an underrun while garbled bytes mean interference.

14. What Comes Next

This chapter held the frame open and streamed bytes without asking where they were going. Chapter 7.2 — Burst Transfers and Auto-Increment supplies that: how a device advances its own address through a burst, why read bursts and write bursts advance under different rules — reads usually crossing page boundaries freely while writes wrap within a page — and the address generator that has to implement all of it, in three HDLs.

Continue learning