SPI · Module 7
Back-to-Back Transactions and Inter-Frame Gap
How soon chip select may fall again after it rises: the minimum deselect time, why a device needs it, what the gap costs in throughput, and the hardware enforcement that keeps software from violating it.
Chapters 7.1 and 7.2 kept one frame open for as long as possible. This chapter asks the opposite question, and it is the one a system with many small transactions actually faces.
Chip select has just risen. How soon may it fall again?
There is a minimum. It is a device parameter that appears in no SPI document, violating it makes transactions vanish, and — as always — nothing reports the failure.
1. The Interval Nobody Specifies
Chapter 2.5 covered the CS lead and lag — the intervals inside a frame, before the first clock edge and after the last. This is a third interval, outside the frame entirely: the time CS must remain high between one frame and the next.
Datasheets call it various things — CS deselect time, CS high time, t_CSH, t_SHSL — and the naming inconsistency is itself informative. SPI does not define it. There is no specification to be consistent with, so each vendor names and specifies it independently, exactly as with width, bit order and dummy cycles.
What makes it easy to miss is that it constrains something most people do not think of as having timing at all. The lead and lag are visibly part of a transfer. The gap is the absence of a transfer, and an absence does not obviously need a duration.
2. Why a Device Needs the Time
The gap is not arbitrary padding. Chapter 5.1 §4 established that CS rising is a device's only atomic event, and a surprising amount happens on it:
- The staged write commits to the destination.
- The phase sequencer resets to expect a command (Chapter 4.4 §7).
- The output driver releases MISO to high impedance (Chapter 6.1 §4).
- On non-volatile memory, a programming or erase operation may begin.
None of that is instantaneous. The device needs CS to stay high long enough for its internal reset to propagate and for its interface logic to settle into the deselected state. If CS falls again before that completes, the device is still finishing the previous transaction and never registers the new frame at all.
There is a second, distinct reason on some devices: a slave whose CS input crosses into its clock domain through a synchroniser (Chapter 5.2 §5) needs CS to be stable for at least the synchroniser's latency in each state. A pulse shorter than that may be seen inconsistently or missed entirely — which makes the minimum deselect time, on an FPGA slave, a direct consequence of its own clock ratio.
3. Legal and Too Fast
The same two frames, spaced two ways
10 cyclesNote that the clock is identical in both cases, and the data would be too. Only the CS framing differs, which is why this failure is invisible in a byte-list view of a capture and visible immediately in a timing view — the same distinction Chapter 4.2 §3 drew for a different reason.
4. What the Gap Costs
For a system doing large transfers the gap is irrelevant. For one doing many small ones it can dominate.
Take a status-register poll: a 1-byte command and a 1-byte response, 16 clock cycles. At 50 MHz that is 320 ns of clocking. Add a hypothetical 100 ns deselect requirement and a 50 ns CS lead:
CS lead .......... 50 ns
clocking .......... 320 ns
CS lag .......... 20 ns
deselect gap .......... 100 ns
────────────────────────────────
per poll 490 ns of which 170 ns is framingThirty-five percent of the transaction is framing overhead, and none of it scales with payload — it is paid per frame, not per byte.
That arithmetic is the quantitative case for everything in this module. Merging transactions removes gaps entirely: a 32-byte auto-incrementing read pays the framing cost once instead of thirty-two times. It is the same conclusion Chapter 6.4 §4 reached about request overhead, arriving now from the framing side, and the two costs are additive — which is why small-transaction-heavy designs are so much slower than their clock rate suggests.
5. Building the Gap Enforcer — Three HDLs
The circuit
Circuit. A three-state machine with a saturating counter, sitting above the frame sequencer and owning CS.
State. Whether the interface is ready, mid-frame, or waiting out a gap; plus the elapsed gap count.
Datapath. None — this module carries no data. It converts a request into a permitted frame start.
Control. A request arriving during a gap is held, not dropped. That is the design's central decision: software issuing back-to-back transfers must be delayed, never silently denied, because a dropped transaction is indistinguishable from a completed one on SPI.
Clock. The system clock, so min_gap is expressed in system-clock cycles and is independent of SCLK.
Reset. Asynchronous, active-low, to ready with CS deasserted — and with the gap counter initialised to all ones, so the first request after reset is not delayed by a gap that never had a preceding frame.
Enables. The counter increments only while the requirement is unmet. Saturating rather than wrapping matters: a long idle period must not roll the counter back under the threshold and manufacture a spurious wait.
Timing. start_frame is a single-cycle strobe; gap_wait is a level so software can observe that it is being throttled.
Synthesis. Two state bits, a GAP_W-bit counter and a comparator. The comparator's second input is a register rather than a constant, since the requirement is per-device.
Limitations. One request depth — a queue of pending transactions belongs upstream. min_gap is in system-clock cycles, so a design that changes its system clock must recompute it.
// spi_frame_gap.sv — enforcing the minimum deselect time between frames.
//
// Chapter 2.5 covered the CS lead and lag INSIDE a frame. This is the
// interval BETWEEN frames -- the time CS must remain high before it may fall
// again. It is a device parameter (often written t_CSH or "CS deselect
// time"), it appears in no SPI document, and violating it means the device
// never registers the second transaction at all.
//
// Enforcing it in hardware rather than in software is the point: a driver
// issuing back-to-back transfers has no reliable way to insert a sub-
// microsecond delay, and the failure is silent (Chapter 5.1 §5).
module spi_frame_gap #(
parameter int GAP_W = 8
) (
input logic clk,
input logic rst_n,
input logic req, // level: a frame is wanted
input logic [GAP_W-1:0] min_gap, // minimum CS-high cycles
input logic frame_done, // pulse: frame content complete
output logic cs_n,
output logic start_frame, // pulse: begin clocking
output logic gap_wait, // level: request held off by the gap
output logic busy
);
typedef enum logic [1:0] { ST_READY, ST_ACTIVE, ST_GAP } state_t;
state_t state;
logic [GAP_W-1:0] gap_cnt;
// The gap is satisfied once the counter has reached the requirement.
// Saturating rather than wrapping means a long idle period cannot roll
// the counter back under the threshold.
logic gap_ok;
assign gap_ok = (gap_cnt >= min_gap);
// A request that arrives too early is HELD, never dropped.
assign gap_wait = req && (state == ST_GAP) && !gap_ok;
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
state <= ST_READY;
cs_n <= 1'b1;
start_frame <= 1'b0;
gap_cnt <= '1; // start fully "gapped" so the first
// request is not delayed after reset
end else begin
start_frame <= 1'b0; // strobe
case (state)
ST_READY: if (req) begin
cs_n <= 1'b0;
start_frame <= 1'b1;
state <= ST_ACTIVE;
end
ST_ACTIVE: if (frame_done) begin
cs_n <= 1'b1;
gap_cnt <= '0; // the gap starts now
state <= ST_GAP;
end
ST_GAP: begin
if (!gap_ok) begin
gap_cnt <= gap_cnt + 1'b1;
end else if (req) begin
// Gap satisfied and a request is waiting: go straight
// into the next frame with no further delay.
cs_n <= 1'b0;
start_frame <= 1'b1;
state <= ST_ACTIVE;
end else begin
state <= ST_READY;
end
end
default: state <= ST_READY;
endcase
end
end
assign busy = (state == ST_ACTIVE);
endmoduleThe ST_GAP state's ordering is the design. It counts first and only then considers the request, so there is no path by which a pending request can shorten the gap. Checking req before the counter — the more natural way to write it — would let a waiting request start a frame on the cycle the counter reached its target, which is off by one in the unsafe direction.
// spi_frame_gap_tb.sv — measure the ACTUAL enforced gap, prove a held
// request is not lost, and prove min_gap=0 permits back-to-back frames.
`timescale 1ns/1ps
module spi_frame_gap_tb;
logic clk = 0, rst_n = 0;
always #5 clk = ~clk;
localparam int GW = 8;
logic req = 0, frame_done = 0;
logic [GW-1:0] min_gap = 8'd5;
logic cs_n, start_frame, gap_wait, busy;
spi_frame_gap #(.GAP_W(GW)) dut (
.clk, .rst_n, .req, .min_gap, .frame_done,
.cs_n, .start_frame, .gap_wait, .busy);
int errors = 0, frames = 0;
task automatic chk(input string what, input int g, input int e);
if (g !== e) begin $display("FAIL %s: got %0d exp %0d", what, g, e); errors++; end
endtask
// Measure CS-high duration, in clock cycles, between consecutive frames.
int gap_len = 0, measured_gap = -1;
logic cs_n_q = 1'b1, measuring = 1'b0;
always @(posedge clk) if (rst_n) begin
cs_n_q <= cs_n;
if (start_frame) frames++;
if (cs_n && !cs_n_q) begin // CS just rose: start measuring
gap_len <= 1;
measuring <= 1'b1;
end else if (!cs_n && cs_n_q) begin // CS just fell: record
measured_gap <= gap_len;
measuring <= 1'b0;
end else if (measuring) begin
gap_len <= gap_len + 1;
end
end
// Run the content of a frame, then signal completion.
task automatic run_frame();
repeat (4) @(negedge clk);
frame_done = 1; @(negedge clk); frame_done = 0; @(negedge clk);
endtask
// Wait for CS to fall, then let the observers settle before sampling
// counters -- start_frame is registered and lands one edge later.
task automatic wait_frame_open();
while (cs_n) @(negedge clk);
repeat (2) @(negedge clk);
endtask
initial begin
repeat (3) @(negedge clk); rst_n = 1; @(negedge clk);
chk("idle: cs_n high", cs_n, 1);
// --- frame 1: no delay after reset ---
req = 1;
wait_frame_open();
chk("frame 1 started promptly", frames, 1);
chk("cs low", cs_n, 0);
// --- req stays high, so frame 2 must be HELD for the gap ---
run_frame();
repeat (2) @(negedge clk);
chk("request held by the gap", gap_wait, 1);
chk("CS high during the gap", cs_n, 1);
wait_frame_open();
chk("frame 2 started after the gap", frames, 2);
if (measured_gap < 5) begin
$display("FAIL: enforced gap was %0d cycles, minimum is 5", measured_gap);
errors++;
end
$display(" enforced gap with min_gap=5: %0d cycles", measured_gap);
// --- a request held through an entire frame must still produce the
// next frame -- held, never dropped ---
run_frame();
wait_frame_open();
chk("held request produced frame 3", frames, 3);
if (measured_gap < 5) begin
$display("FAIL: enforced gap was %0d cycles, minimum is 5", measured_gap);
errors++;
end
run_frame();
req = 0;
repeat (12) @(negedge clk);
// --- min_gap = 0 permits genuinely back-to-back frames. req is held
// high ACROSS the boundary so the measured interval is the
// enforced gap and not idle time. ---
min_gap = 8'd0;
repeat (2) @(negedge clk);
req = 1;
wait_frame_open();
chk("gapless: frame 4 open", frames, 4);
run_frame();
wait_frame_open();
chk("gapless: frame 5 open", frames, 5);
$display(" enforced gap with min_gap=0: %0d cycles", measured_gap);
if (measured_gap > 2) begin
$display("FAIL: min_gap=0 should allow a near-gapless frame, got %0d", measured_gap);
errors++;
end
run_frame();
req = 0;
repeat (8) @(negedge clk);
if (errors == 0)
$display("PASS: the enforced gap is never shorter than min_gap, a request arriving early is held rather than dropped, and min_gap=0 permits back-to-back frames");
else
$display("FAILED with %0d error(s)", errors);
$finish;
end
initial begin #500000; $display("FAIL: watchdog timeout"); $finish; end
endmoduleThat testbench measures the gap rather than assuming it. It counts CS-high cycles between consecutive frames and asserts the result is never shorter than the requirement — and it prints the measured value, so the margin is visible rather than merely adequate. All three implementations report 6 cycles for a minimum of 5, and 1 cycle when the minimum is 0.
The min_gap = 0 case is worth keeping. Some devices genuinely permit gapless framing, and a design that cannot express "no gap required" imposes a cost its device does not ask for.
// spi_frame_gap.v — the same inter-frame gap enforcer in Verilog-2001.
module spi_frame_gap #(
parameter GAP_W = 8
) (
input wire clk,
input wire rst_n,
input wire req,
input wire [GAP_W-1:0] min_gap,
input wire frame_done,
output reg cs_n,
output reg start_frame,
output wire gap_wait,
output wire busy
);
localparam ST_READY = 2'd0,
ST_ACTIVE = 2'd1,
ST_GAP = 2'd2;
reg [1:0] state;
reg [GAP_W-1:0] gap_cnt;
// Saturating rather than wrapping: a long idle period cannot roll the
// counter back under the threshold.
wire gap_ok = (gap_cnt >= min_gap);
// A request that arrives too early is HELD, never dropped.
assign gap_wait = req && (state == ST_GAP) && !gap_ok;
assign busy = (state == ST_ACTIVE);
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
state <= ST_READY;
cs_n <= 1'b1;
start_frame <= 1'b0;
gap_cnt <= {GAP_W{1'b1}}; // start fully "gapped"
end else begin
start_frame <= 1'b0;
case (state)
ST_READY: if (req) begin
cs_n <= 1'b0;
start_frame <= 1'b1;
state <= ST_ACTIVE;
end
ST_ACTIVE: if (frame_done) begin
cs_n <= 1'b1;
gap_cnt <= {GAP_W{1'b0}}; // the gap starts now
state <= ST_GAP;
end
ST_GAP: begin
if (!gap_ok) begin
gap_cnt <= gap_cnt + 1'b1;
end else if (req) begin
cs_n <= 1'b0;
start_frame <= 1'b1;
state <= ST_ACTIVE;
end else begin
state <= ST_READY;
end
end
default: state <= ST_READY;
endcase
end
end
endmodule// spi_frame_gap_tb.v — the same checks in Verilog-2001.
`timescale 1ns/1ps
module spi_frame_gap_tb;
reg clk = 0, rst_n = 0;
always #5 clk = ~clk;
parameter GW = 8;
reg req = 0, frame_done = 0;
reg [GW-1:0] min_gap = 8'd5;
wire cs_n, start_frame, gap_wait, busy;
spi_frame_gap #(.GAP_W(GW)) dut (
.clk(clk), .rst_n(rst_n), .req(req), .min_gap(min_gap),
.frame_done(frame_done), .cs_n(cs_n), .start_frame(start_frame),
.gap_wait(gap_wait), .busy(busy));
integer errors = 0, frames = 0, gap_len = 0, measured_gap = -1;
reg cs_n_q = 1'b1, measuring = 1'b0;
task chk;
input [80*8-1:0] what;
input [31:0] g, e;
begin
if (g !== e) begin
$display("FAIL %0s: got %0d exp %0d", what, g, e);
errors = errors + 1;
end
end
endtask
always @(posedge clk) if (rst_n) begin
cs_n_q <= cs_n;
if (start_frame) frames = frames + 1;
if (cs_n && !cs_n_q) begin
gap_len <= 1;
measuring <= 1'b1;
end else if (!cs_n && cs_n_q) begin
measured_gap <= gap_len;
measuring <= 1'b0;
end else if (measuring) begin
gap_len <= gap_len + 1;
end
end
task run_frame;
begin
repeat (4) @(negedge clk);
frame_done = 1; @(negedge clk); frame_done = 0; @(negedge clk);
end
endtask
task wait_frame_open;
begin
while (cs_n) @(negedge clk);
repeat (2) @(negedge clk);
end
endtask
initial begin
repeat (3) @(negedge clk); rst_n = 1; @(negedge clk);
chk("idle: cs_n high", cs_n, 1);
req = 1;
wait_frame_open;
chk("frame 1 started promptly", frames, 1);
chk("cs low", cs_n, 0);
run_frame;
repeat (2) @(negedge clk);
chk("request held by the gap", gap_wait, 1);
chk("CS high during the gap", cs_n, 1);
wait_frame_open;
chk("frame 2 started after the gap", frames, 2);
if (measured_gap < 5) begin
$display("FAIL: enforced gap was %0d cycles, minimum is 5", measured_gap);
errors = errors + 1;
end
$display(" enforced gap with min_gap=5: %0d cycles", measured_gap);
run_frame;
wait_frame_open;
chk("held request produced frame 3", frames, 3);
if (measured_gap < 5) begin
$display("FAIL: enforced gap was %0d cycles, minimum is 5", measured_gap);
errors = errors + 1;
end
run_frame;
req = 0;
repeat (12) @(negedge clk);
min_gap = 8'd0;
repeat (2) @(negedge clk);
req = 1;
wait_frame_open;
chk("gapless: frame 4 open", frames, 4);
run_frame;
wait_frame_open;
chk("gapless: frame 5 open", frames, 5);
$display(" enforced gap with min_gap=0: %0d cycles", measured_gap);
if (measured_gap > 2) begin
$display("FAIL: min_gap=0 should allow a near-gapless frame, got %0d", measured_gap);
errors = errors + 1;
end
run_frame;
req = 0;
repeat (8) @(negedge clk);
if (errors == 0)
$display("PASS: the enforced gap is never shorter than min_gap, a request arriving early is held rather than dropped, and min_gap=0 permits back-to-back frames");
else
$display("FAILED with %0d error(s)", errors);
$finish;
end
initial begin #500000; $display("FAIL: watchdog timeout"); $finish; end
endmodule-- spi_frame_gap.vhd — the same inter-frame gap enforcer in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_frame_gap is
generic (
GAP_W : positive := 8
);
port (
clk : in std_logic;
rst_n : in std_logic;
req : in std_logic; -- level
min_gap : in unsigned(GAP_W - 1 downto 0); -- minimum CS-high cycles
frame_done : in std_logic; -- pulse
cs_n : out std_logic;
start_frame : out std_logic; -- pulse
gap_wait : out std_logic; -- level
busy : out std_logic
);
end entity spi_frame_gap;
architecture rtl of spi_frame_gap is
type state_t is (ST_READY, ST_ACTIVE, ST_GAP);
signal state : state_t;
signal gap_cnt : unsigned(GAP_W - 1 downto 0);
signal gap_ok : std_logic;
begin
-- Saturating rather than wrapping: a long idle period cannot roll the
-- counter back under the threshold.
gap_ok <= '1' when gap_cnt >= min_gap else '0';
-- A request that arrives too early is HELD, never dropped.
gap_wait <= '1' when req = '1' and state = ST_GAP and gap_ok = '0' else '0';
busy <= '1' when state = ST_ACTIVE else '0';
process (clk, rst_n) is
begin
if rst_n = '0' then
state <= ST_READY;
cs_n <= '1';
start_frame <= '0';
gap_cnt <= (others => '1'); -- start fully "gapped"
elsif rising_edge(clk) then
start_frame <= '0';
case state is
when ST_READY =>
if req = '1' then
cs_n <= '0';
start_frame <= '1';
state <= ST_ACTIVE;
end if;
when ST_ACTIVE =>
if frame_done = '1' then
cs_n <= '1';
gap_cnt <= (others => '0'); -- the gap starts now
state <= ST_GAP;
end if;
when ST_GAP =>
if gap_ok = '0' then
gap_cnt <= gap_cnt + 1;
elsif req = '1' then
cs_n <= '0';
start_frame <= '1';
state <= ST_ACTIVE;
else
state <= ST_READY;
end if;
end case;
end if;
end process;
end architecture rtl;-- spi_frame_gap_tb.vhd — the same checks in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_frame_gap_tb is
end entity spi_frame_gap_tb;
architecture tb of spi_frame_gap_tb is
constant GW : positive := 8;
signal clk : std_logic := '0';
signal rst_n : std_logic := '0';
signal req : std_logic := '0';
signal frame_done : std_logic := '0';
signal min_gap : unsigned(GW - 1 downto 0) := to_unsigned(5, GW);
signal halt : boolean := false;
signal cs_n : std_logic;
signal start_frame : std_logic;
signal gap_wait : std_logic;
signal busy : std_logic;
signal errors : natural := 0;
signal frames : natural := 0;
signal measured_gap : natural := 0;
begin
clk <= not clk after 5 ns when not halt else '0';
dut : entity work.spi_frame_gap
generic map (GAP_W => GW)
port map (clk => clk, rst_n => rst_n, req => req, min_gap => min_gap,
frame_done => frame_done, cs_n => cs_n, start_frame => start_frame,
gap_wait => gap_wait, busy => busy);
-- One process owns the observation counters.
observe : process (clk) is
variable cs_n_q : std_logic := '1';
variable gap_len : natural := 0;
variable measuring : boolean := false;
begin
if rising_edge(clk) and rst_n = '1' then
if start_frame = '1' then
frames <= frames + 1;
end if;
if cs_n = '1' and cs_n_q = '0' then
gap_len := 1;
measuring := true;
elsif cs_n = '0' and cs_n_q = '1' then
measured_gap <= gap_len;
measuring := false;
elsif measuring then
gap_len := gap_len + 1;
end if;
cs_n_q := cs_n;
end if;
end process;
stim : process is
procedure chk_n (what : string; g, e : natural) is
begin
if g /= e then
report "FAIL " & what & ": got " & integer'image(g)
& " exp " & integer'image(e) severity error;
errors <= errors + 1;
end if;
end procedure;
procedure chk_b (what : string; g, e : std_logic) is
begin
if g /= e then
report "FAIL " & what severity error;
errors <= errors + 1;
end if;
end procedure;
procedure run_frame is
begin
for i in 0 to 3 loop
wait until falling_edge(clk);
end loop;
frame_done <= '1';
wait until falling_edge(clk);
frame_done <= '0';
wait until falling_edge(clk);
end procedure;
-- Wait for CS to fall, then let the observers settle: start_frame is
-- registered and lands one edge later.
procedure wait_frame_open is
begin
while cs_n = '1' loop
wait until falling_edge(clk);
end loop;
for i in 0 to 1 loop
wait until falling_edge(clk);
end loop;
end procedure;
begin
for i in 0 to 2 loop
wait until falling_edge(clk);
end loop;
rst_n <= '1';
wait until falling_edge(clk);
chk_b("idle: cs_n high", cs_n, '1');
req <= '1';
wait_frame_open;
chk_n("frame 1 started promptly", frames, 1);
chk_b("cs low", cs_n, '0');
run_frame;
for i in 0 to 1 loop wait until falling_edge(clk); end loop;
chk_b("request held by the gap", gap_wait, '1');
chk_b("CS high during the gap", cs_n, '1');
wait_frame_open;
chk_n("frame 2 started after the gap", frames, 2);
if measured_gap < 5 then
report "FAIL: enforced gap was " & integer'image(measured_gap)
& " cycles, minimum is 5" severity error;
errors <= errors + 1;
end if;
report " enforced gap with min_gap=5: " & integer'image(measured_gap)
& " cycles" severity note;
run_frame;
wait_frame_open;
chk_n("held request produced frame 3", frames, 3);
if measured_gap < 5 then
report "FAIL: enforced gap too short" severity error;
errors <= errors + 1;
end if;
run_frame;
req <= '0';
for i in 0 to 11 loop wait until falling_edge(clk); end loop;
-- min_gap = 0 permits genuinely back-to-back frames.
min_gap <= to_unsigned(0, GW);
for i in 0 to 1 loop wait until falling_edge(clk); end loop;
req <= '1';
wait_frame_open;
chk_n("gapless: frame 4 open", frames, 4);
run_frame;
wait_frame_open;
chk_n("gapless: frame 5 open", frames, 5);
report " enforced gap with min_gap=0: " & integer'image(measured_gap)
& " cycles" severity note;
if measured_gap > 2 then
report "FAIL: min_gap=0 should allow a near-gapless frame" severity error;
errors <= errors + 1;
end if;
run_frame;
req <= '0';
for i in 0 to 7 loop wait until falling_edge(clk); end loop;
if errors = 0 then
report "PASS: the enforced gap is never shorter than min_gap, a request "
& "arriving early is held rather than dropped, and min_gap=0 permits "
& "back-to-back frames" severity note;
else
report "FAILED with " & integer'image(errors) & " error(s)" severity error;
end if;
halt <= true;
wait;
end process;
watchdog : process is
begin
wait for 500 us;
if not halt then
report "FAIL: watchdog timeout" severity failure;
end if;
wait;
end process;
end architecture tb;Parity
All three implement the same machine: identical ports, asynchronous active-low reset to ready with CS high and the counter saturated, a held rather than dropped request, counting before considering the request, and a single-cycle start_frame with a level gap_wait. All three testbenches measure the enforced interval and report the same values — 6 cycles against a minimum of 5, and 1 cycle against a minimum of 0.
6. Why a Verification Engineer Cares
This is a property about an interval, which is what makes it worth asserting rather than inspecting:
// 1. THE property. Between CS rising and CS falling again, at least
// min_gap cycles must elapse. Everything else here serves this.
property p_min_deselect;
int unsigned g;
@(posedge clk) disable iff (!rst_n)
($rose(cs_n), g = min_gap) |-> ##[1:$] ($fell(cs_n) ##0 (gap_counter >= g));
endproperty
a_min_deselect : assert property (p_min_deselect)
else $error("CS fell before the minimum deselect time elapsed");
// 2. A request is never dropped. If req is held, a frame must eventually
// start -- a liveness property, and the reason software can trust the
// throttling rather than having to implement it.
property p_request_not_dropped;
@(posedge clk) disable iff (!rst_n)
(req && !busy) |-> ##[1:$] start_frame;
endproperty
a_request_lives : assert property (p_request_not_dropped)
else $error("a held request never produced a frame");
// 3. start_frame implies CS is being asserted, not merely that the
// sequencer felt like starting.
a_start_implies_select : assert property (
@(posedge clk) disable iff (!rst_n) start_frame |-> !cs_n)
else $error("frame started without asserting CS");Property 2 is the one most often left out, and it is the interesting kind: a liveness property rather than a safety one. The safety properties say the gap is never too short; property 2 says the throttle never becomes a block. A design that satisfied only the safety properties could deadlock and pass — and deadlock is exactly the failure mode a badly written gap enforcer produces.
What these prove. That the enforced interval meets the configured requirement and that requests are delayed rather than lost. What they cannot prove is that min_gap matches what the device needs — a datasheet fact, and Chapter 4.1 §7's boundary once more. The enforcer can be provably correct and configured with a number that is too small.
Coverage should target the timing relationship, not the frames:
covergroup spi_gap_cg @(posedge cs_fell);
// The measured interval RELATIVE to the requirement. Absolute cycle
// counts say nothing; the ratio is the whole question.
cp_margin : coverpoint (actual_gap - cfg.min_gap) {
bins exact = {0}; // landed exactly on the limit
bins tight = {[1:3]};
bins comfortable = {[4:63]};
bins idle = {[64:$]}; // long idle, not back-to-back at all
}
cp_min_gap : coverpoint cfg.min_gap {
bins none = {0}; // gapless permitted -- distinct path
bins small = {[1:15]};
bins large = {[16:$]};
}
// Whether the request was waiting when the gap expired. A request that
// arrives AFTER the gap has elapsed never exercises the hold path.
cp_pending : coverpoint req_was_pending_at_expiry {
bins arrived_late = {0};
bins was_waiting = {1}; // the back-to-back case
}
x_margin_pending : cross cp_margin, cp_pending;
endgroupcp_margin.exact is the bin to insist on. A design that is off by one in the gap comparison produces a gap one cycle short — correct-looking in every comfortable case and wrong only when a request is waiting at the moment the counter expires. That is precisely the exact × was_waiting cell, and random stimulus reaches it rarely.
7. Why an FPGA or ASIC Engineer Cares
Enforce it in hardware, not in software. A driver cannot reliably insert a 100 ns delay: the granularity of a software timer is orders of magnitude coarser, a busy-wait burns CPU and is not portable, and on a preemptive system the delay may become arbitrarily long rather than arbitrarily short — which is safe but destroys throughput. A counter in the controller costs a handful of flip-flops and removes the problem from software entirely.
min_gap is in system-clock cycles and must be recomputed if the clock changes. A design specifying 10 cycles at 100 MHz enforces 100 ns; the same constant at 50 MHz enforces 200 ns, which is merely slow, and at 200 MHz enforces 50 ns, which may violate the device. Deriving the constant from a ns value and the clock period at elaboration is the robust form.
On a slave, the requirement is partly your own doing. If your CS input crosses through a two-flop synchroniser and your system clock is slow relative to the master's framing rate, you are the reason the deselect time is what it is. Publishing a number in your own datasheet means computing (SYNC_STAGES + 1) × T_system and adding margin — the same arithmetic as Chapter 5.2 §12, applied to the other edge.
Watch out for a CS driven by a GPIO. Many designs drive CS from software rather than from the SPI controller, to hold it across multiple controller transfers (Chapter 5.1 §11). That works, and it moves the gap enforcement into software where it cannot be guaranteed. If CS is a GPIO, the gap is only as reliable as the code toggling it.
8. Failure Signature — Every Other Transaction Is Ignored
Symptom. A driver issues a rapid sequence of short transactions — a burst of register writes during initialisation, or a tight polling loop. Roughly half of them take effect. Reading back shows some registers written and others at their reset values, and which ones vary between runs. Inserting a delay in the loop makes the problem disappear entirely.
What "a delay fixes it" establishes. That the fault is temporal rather than logical. The bytes are correct — a slower version of the same code works with the same data — so the mode, framing, addressing and command encoding are all right. Something about the rate is the problem.
What "roughly half" establishes. This is the sharper clue. A device that misses a transaction because it is still processing the previous one will reliably accept transaction 1, miss 2, accept 3, miss 4 — because after each missed transaction the device has had a full transaction's worth of idle time to recover. The alternation is the signature of a per-transaction recovery requirement, and it is quite different from random loss.
Plausible mechanisms.
- The inter-frame gap is being violated, so the device misses every frame that arrives too soon after the previous one.
- The device is busy with an internal operation — a programming or erase cycle — and ignores commands until it completes (Chapter 5.1 §5).
- A write-enable interlock not reissued per transaction — but that produces a consistent pattern, not an alternating one.
The discriminating observation. Capture CS as a timing waveform and measure the high time between frames, then compare against the device's deselect specification. That is a direct measurement of the suspected quantity and settles it immediately.
To separate it from the busy case without an instrument: vary the delay and watch how the success rate changes. A deselect violation has a sharp threshold — below the required gap almost nothing gets through, above it everything does. A busy-device problem improves gradually as the delay approaches the operation's completion time, because the operation's duration varies.
Why the investigation goes wrong. Because "adding a delay fixes it" is treated as the solution rather than as evidence. The delay is usually inserted generously, the symptom disappears, and the actual requirement is never measured — leaving a design that works by an unknown margin and will fail on a faster CPU, a different compiler's optimisation, or a device from a different batch.
9. Common Misconceptions
10. Reason It Through
Work this before reading the answer.
An FPGA SPI master runs from a 100 MHz system clock. Its slave is another FPGA whose SPI interface runs from a 25 MHz clock and uses a two-flop synchroniser on CS. The master issues back-to-back single-byte transactions with 40 ns between CS rising and falling again.
Will the slave see every transaction? What is the correct minimum, and where should it be enforced?
Find the slave's reaction time. A 25 MHz clock is a 40 ns period. A two-flop synchroniser plus the edge-detect register means the slave needs three of its own clock periods — 120 ns — to recognise a CS edge (Chapter 5.2 §12).
Compare against the supplied gap. The master leaves 40 ns, which is one slave clock period. The slave has not finished recognising that CS went high before it goes low again.
What actually happens? The synchroniser samples a CS that has already returned low, so the rising edge may never appear in the synchronised signal at all. The slave's frame_end strobe never fires, its sequencer never resets, and it treats the second transaction as a continuation of the first — interpreting the new opcode as data in the previous frame's data phase.
That is worse than missing the transaction. The slave does not drop the second frame; it merges it into the first, so the data goes somewhere unintended rather than nowhere.
The correct minimum. The slave must observe CS high for at least its own recognition time, with margin:
minimum gap = (SYNC_STAGES + 1) × T_slave
= 3 × 40 ns
= 120 ns absolute floor
with margin → 160–200 nsWhere should it be enforced? In the master's hardware, because that is the only place it can be guaranteed. Expressed in the master's own 100 MHz clock, 200 ns is 20 cycles, so min_gap = 20 — and note that this constant is a property of the slave, computed from the slave's clock, but configured in the master.
And the deeper point. The slave's designer should publish this number. It is not a property the master can discover, it does not appear in any SPI document, and it arises entirely from an implementation choice — the synchroniser depth and the interface clock rate — that the master's designer has no visibility of. An FPGA slave that omits a deselect-time specification from its interface documentation has left a trap.
What if the master cannot be changed? Then the slave must be. Raising the slave's SPI-interface clock is the direct fix: at 100 MHz the same three stages cost 30 ns and the existing 40 ns gap is sufficient. That is the same conclusion as Chapter 5.2 §12 — the ratio between the two clocks is the real parameter, and it governs both the CS lead and the inter-frame gap.
11. Understanding Check
12. Summary
Between one frame and the next there is a minimum deselect time — how long CS must stay high before it may fall again. SPI does not define it, vendors name it inconsistently, and it must come from the datasheet.
It exists because CS rising is a device's only atomic event and a great deal happens on it: the write commits, the sequencer resets, the output releases. The gap is the time that response takes, which makes it the tail of the previous transaction rather than dead time before the next.
Violating it means the device never registers the new frame, and nothing reports the loss. On an FPGA slave with a synchronised CS the requirement is partly self-inflicted — (SYNC_STAGES + 1) × T_system — and can produce something worse than a dropped frame: a second transaction merged into the first.
The cost is real for small transactions. Framing overhead is paid per frame, not per byte, and can be a third of a short poll — which is the quantitative case for merging transactions that Chapters 7.1 and 7.2 made qualitatively.
In RTL the enforcer is a three-state machine with a saturating counter that counts before considering the request, so no pending request can shorten the gap, and that holds rather than drops an early request — because a dropped transaction is indistinguishable from a completed one.
For verification the safety property is that the interval is never too short, and the property most often missing is the liveness one: a held request must eventually produce a frame, since a gap enforcer's characteristic failure is deadlock. Coverage should target the margin relative to the requirement, with the exact-limit-plus-pending-request cell being where an off-by-one hides.
And when half of a rapid sequence succeeds, alternation is the signature of a per-transaction recovery requirement — and "a delay fixed it" is evidence, not a solution.
13. What Comes Next
That closes Module 7. Transfers are no longer single: one frame can stream indefinitely and stall without harm, the device advances its own address under three different policies, and separate frames must respect a gap the bus never mentions.
Every module so far has assumed one device on the bus. Module 8 removes that assumption and is the longest in the track for good reason: multiple slaves sharing MISO, chip-select decoding, the contention that Chapter 1.2 warned about becoming a real design problem, daisy-chaining, and the bus-ownership rules that keep several devices on three shared wires from destroying each other's data — and occasionally each other.
Browse the path on the SPI curriculum index, or revisit Continuous Transfers for the framing discipline this chapter bounds from the outside.
Continue learning
Related tutorials
- Related topic
Dummy Phases and Read Latency
Why a device needs turnaround before it can answer, why dummy is counted in clock cycles rather than bytes, how its length grows with frequency, and the one-byte data offset a mismatch produces.
- Related topic
Raw vs Effective Payload Throughput
The three throughput numbers a link has — raw bit rate, protocol rate and effective payload rate — why they differ by four times on an ordinary flash read, and the performance counter that measures the gap instead of estimating it.
- Related topic
SDR, DDR Transfers, and Throughput
Why data on both edges is independent of lane width, why the byte assembly is identical in both modes, why the dummy phase does not halve so DDR gives under two times, and why DDR parts specify more dummy cycles than their SDR modes.
- Related topic
Multi-Byte and Back-to-Back Transactions
Keeping the wire busy between frames with one register's worth of flops: why one deep is enough and what a deeper FIFO actually buys, why a stall must be reported rather than absorbed, and why an overrun is lost data.
