SPI · Module 2
CS-to-SCLK and SCLK-to-CS Timing
Chip select has timing requirements of its own: the lead before the first clock edge, the lag after the last, and the minimum deselect between transactions. Why violating them breaks a transfer whose every SCLK edge was correct.
Every chapter so far has treated chip select as a boundary marker: Chapter 1.2 established it as the line that grants a device the right to drive MISO, and Chapter 1.5 as the thing that distinguishes one peripheral from another. Both descriptions are about which device, not about when.
That was a simplification, and it is the one this chapter removes.
Why can a transaction fail when every SCLK edge is correct?
Because CS is not only a selector. It is a signal with its own setup and hold requirements relative to the clock burst, and a device that is selected too late — or released too early — will mis-handle a transfer whose clock and data are flawless. The resulting failure has a signature unlike anything in Chapter 2.4: it is perfectly repeatable, frequency-independent, and confined to the edges of the transaction rather than distributed through it.
1. What a Device Does When CS Falls
To see why CS needs timing, look at what its assertion actually triggers inside the peripheral.
The falling edge of CS is not a passive notification. It is the event that prepares the device to participate, and on a typical peripheral it sets several things in motion at once: an internal bit counter is cleared so the next clock edge counts as bit one; the transaction state machine leaves idle and begins interpreting incoming bits as an opcode; the MISO output driver is enabled (Chapter 1.2); and under one of the two edge-role configurations from Chapter 2.3, the first bit must be launched onto MISO by this event rather than by a clock edge, because there is no earlier edge available.
All of that is real logic with real propagation delay. None of it is instantaneous.
So if the first SCLK edge arrives too soon after CS falls, it arrives at a device that has not finished preparing. The edge may be missed entirely, counted against a stale bit position, or captured while the internal state machine is still in transition. The clock was correct; the device was not ready to receive it.
The mirror image applies at the end. The rising edge of CS ends the transaction, and the device needs the last clock edge's effects to have settled before that happens — the final bit captured, the internal state committed, the output released cleanly. Release CS too soon after the last edge and the final bit of the transfer can be lost.
2. The Four Requirements
Device datasheets name these differently, but the same four intervals appear across essentially all SPI peripherals. Fixing the vocabulary makes datasheet reading mechanical.
CS lead time — the minimum interval from CS asserting to the first SCLK edge. Often written t_CSS, t_SSC, or "CS setup time." This is the interval that lets the device finish the preparation described in §1.
CS lag time — the minimum interval from the last SCLK edge to CS deasserting. Often t_CSH, t_CHS, or "CS hold time." This lets the final edge's effects settle before the transaction is torn down.
CS deselect time — the minimum interval CS must remain deasserted between two transactions. Often t_CSD or t_SSH. A device needs time to conclude one transaction and return to idle; asserting CS again immediately can merge two transactions into one from the device's point of view, or find it mid-teardown.
CS-to-data-valid time — how long after CS asserts the device guarantees MISO carries meaningful data. This one belongs to the device's output behaviour rather than to what the master must provide, and Chapter 2.6 develops it.
The first three are obligations the master must satisfy. They are not negotiable and they are not implied by getting the clock right — which is exactly why a controller with a perfect divider can still fail.
CS lead, the clock burst, and CS lag
10 cyclesNote what the figure does not show: any irregularity in the clock. The burst is clean, the periods are equal, the data changes away from the edges. All the timing content is in the two shaded gaps, which is precisely why this class of fault survives a casual look at a capture.
3. Why the First Byte Is the Casualty
CS timing produces a distinctive failure pattern, and recognising it is most of the diagnosis.
A lead-time violation damages the beginning of the transaction. The device was still preparing when the first edge arrived, so its bit counter is offset, or it missed the edge entirely. Everything after that is processed correctly — the device is fully awake by bit three — but it is now interpreting the stream one position out, or it has taken the opcode to be something else.
The symptom is therefore: the first byte is wrong, and the rest is consistent with the device having misread it. Not random corruption; a structured misinterpretation. If the opcode was misread, the device may return data from the wrong register, or refuse to respond at all, and the subsequent bytes will be perfectly formed answers to the wrong question.
A lag-time violation damages the end. The last bit may not be committed, so a write loses its final bit, or a read returns a final byte that is short. The first bytes of the transfer are flawless.
Both are completely repeatable, because they are caused by a fixed relationship between two signals the master generates. Chapter 2.4's margin failures are intermittent and frequency-dependent; these are neither. That contrast is diagnostically decisive and costs nothing to observe.
And — the part that catches people — lowering the clock frequency often does not help. If a master asserts CS and starts clocking in the same system-clock cycle, the lead interval is near zero regardless of how slow SCLK is. The gap is set by the controller's own sequencing, not by the bus rate. An engineer who reflexively slows the clock and sees no improvement has, without realising it, gathered strong evidence that the problem is CS timing rather than margin.
4. Building the Guard — Three HDLs
Now the hardware. A master must sequence CS and the clock enable with defined gaps, and that is a small state machine worth building explicitly.
Circuit
A five-state machine with one counter. It asserts CS, holds off the clock for a programmed number of system-clock cycles, enables the divider from Chapter 2.1 for the duration of the burst, holds CS after the burst completes, releases it, and then enforces a minimum deselect interval before another transaction can begin.
Registers
Three pieces of state: the state variable itself, a cycle counter shared across the three timed intervals, and the registered cs_n output. Registering CS matters — a combinationally derived select can glitch while the state encoding changes, which is the hazard Chapter 1.5 raised for decoders and which applies just as directly here.
Combinational logic
Next-state logic, the terminal-count comparisons, and two derived outputs: clk_en (asserted only in the ACTIVE state) and busy.
Clock
The system clock. Note that this block deliberately knows nothing about SCLK — it gates the divider rather than counting SCLK edges, so the lead and lag intervals are expressed in system-clock cycles and are therefore independent of the SCLK divisor. That is the right choice: a device's t_CSS is a time in nanoseconds, not a number of bus clocks, so pinning it to the system clock keeps it constant as the bus rate changes.
Enables
start requests a transaction and is accepted only in IDLE, so the deselect interval cannot be skipped by an eager requester. burst_done is the transfer engine reporting that the final SCLK edge has occurred.
Reset
Asynchronous, active-low, returning to IDLE with CS deasserted. That direction is deliberate: an unknown or asserted select at power-up would let a device drive the shared MISO net before the master has decided anything, the hazard from Chapter 1.2.
Timing
LEAD_CYCLES system-clock cycles elapse between CS falling and clk_en rising; LAG_CYCLES elapse between burst_done and CS rising; IDLE_CYCLES elapse before the next start is accepted. Converting a device's nanosecond requirement into these parameters is a division by the system clock period, rounded up.
Synthesis
A three-bit state register, a small counter, comparators, and a handful of gates. Negligible.
Limitation
This is the CS sequencer only. It has no bit counter — burst_done comes from elsewhere — no shift register, no mode configuration, and no multi-slave select decoding (Chapter 1.5 built that separately). A real master fuses this sequencing with the transfer engine so that burst_done is derived from the bit count rather than supplied; Module 13 does that.
module spi_cs_guard #(
parameter int LEAD_CYCLES = 3, // clk cycles: CS low before the first SCLK edge
parameter int LAG_CYCLES = 3, // clk cycles: last SCLK edge before CS releases
parameter int IDLE_CYCLES = 4 // clk cycles: minimum CS-high between transactions
) (
input logic clk,
input logic rst_n, // asynchronous, active-low
input logic start, // request a transaction (accepted only when idle)
input logic burst_done, // transfer engine: the last SCLK edge has occurred
output logic cs_n, // registered — a combinational select can glitch
output logic clk_en, // enables the divider of Chapter 2.1
output logic busy
);
typedef enum logic [2:0] { IDLE, LEAD, ACTIVE, LAG, RECOVER } state_t;
state_t state;
// One counter serves all three intervals; size it for the largest.
localparam int MAXC =
(LEAD_CYCLES > LAG_CYCLES)
? ((LEAD_CYCLES > IDLE_CYCLES) ? LEAD_CYCLES : IDLE_CYCLES)
: ((LAG_CYCLES > IDLE_CYCLES) ? LAG_CYCLES : IDLE_CYCLES);
localparam int CW = (MAXC <= 1) ? 1 : $clog2(MAXC);
logic [CW-1:0] cnt;
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
state <= IDLE;
cnt <= '0;
cs_n <= 1'b1; // reset MUST deselect
end else begin
case (state)
IDLE: begin
cnt <= '0;
if (start) begin
cs_n <= 1'b0; // select, then wait
state <= LEAD;
end
end
LEAD: begin // CS is low, clock still parked
if (cnt == CW'(LEAD_CYCLES - 1)) begin
cnt <= '0;
state <= ACTIVE;
end else begin
cnt <= cnt + 1'b1;
end
end
ACTIVE: begin // clk_en is high; divider runs
cnt <= '0;
if (burst_done) state <= LAG;
end
LAG: begin // clock parked again, CS still low
if (cnt == CW'(LAG_CYCLES - 1)) begin
cnt <= '0;
cs_n <= 1'b1; // release only now
state <= RECOVER;
end else begin
cnt <= cnt + 1'b1;
end
end
RECOVER: begin // enforce minimum deselect time
if (cnt == CW'(IDLE_CYCLES - 1)) begin
cnt <= '0;
state <= IDLE;
end else begin
cnt <= cnt + 1'b1;
end
end
default: state <= IDLE;
endcase
end
end
assign clk_en = (state == ACTIVE);
assign busy = (state != IDLE);
endmoduleThe invariant worth asserting is the lead requirement itself, because it is the one a refactor is most likely to lose:
// After CS falls, clk_en must remain low for LEAD_CYCLES cycles. If a future
// change enables the divider from the same decision that asserts CS, this
// fails immediately — which is the point.
a_cs_lead : assert property (
@(posedge clk) disable iff (!rst_n)
$fell(cs_n) |-> (!clk_en)[*LEAD_CYCLES]
) else $error("SCLK enabled before the CS lead interval elapsed");
// And the mirror: CS must not release while the divider is still enabled.
a_cs_lag : assert property (
@(posedge clk) disable iff (!rst_n)
$rose(cs_n) |-> $past(!clk_en, 1)
) else $error("CS released while the clock was still enabled");What these prove and what they do not: they establish that this controller's own sequencing respects the programmed intervals, measured in system-clock cycles. They cannot establish that those intervals are long enough for the attached device — that is a comparison between the parameter you chose and a number in the datasheet, and no assertion on internal signals can see the datasheet. Nor do they say anything about the delay from these internal signals to the actual pins, which Chapter 2.7 and Module 15 address.
The testbench measures the intervals rather than trusting them, and checks that the deselect requirement cannot be bypassed.
module spi_cs_guard_tb;
localparam int LEAD = 3, LAG = 3, IDLEC = 4;
logic clk = 1'b0, rst_n, start, burst_done;
logic cs_n, clk_en, busy;
int errors = 0;
int measured;
spi_cs_guard #(.LEAD_CYCLES(LEAD), .LAG_CYCLES(LAG), .IDLE_CYCLES(IDLEC)) dut (
.clk(clk), .rst_n(rst_n), .start(start), .burst_done(burst_done),
.cs_n(cs_n), .clk_en(clk_en), .busy(busy));
always #5 clk = ~clk;
// A watchdog turns a hang into a FAILURE. Without it, a sequencer that
// never leaves a state produces an infinite run rather than a red test.
initial begin
#10_000;
$display("FAIL: timeout — the guard never completed its sequence");
$finish;
end
// Continuous invariant: the clock must never be enabled while CS is high.
always @(posedge clk) if (rst_n && clk_en && cs_n) begin
$error("clk_en asserted while CS deasserted"); errors++;
end
initial begin
rst_n = 1'b0; start = 1'b0; burst_done = 1'b0;
@(posedge clk); #1;
if (cs_n !== 1'b1) begin $error("reset must deselect"); errors++; end
if (clk_en !== 1'b0) begin $error("reset must park the clock"); errors++; end
rst_n = 1'b1;
@(posedge clk); #1;
// --- LEAD: count edges from the accepted start until the clock is enabled.
start = 1'b1; @(posedge clk); #1; start = 1'b0;
if (cs_n !== 1'b0) begin $error("CS did not assert on start"); errors++; end
measured = 0;
while (clk_en !== 1'b1) begin
@(posedge clk); #1;
measured++;
if (measured > LEAD) break; // watchdog handles a true hang
end
if (measured != LEAD) begin
$error("lead = %0d cycles, expected %0d", measured, LEAD); errors++;
end
// --- Nominal burst, then completion.
repeat (6) @(posedge clk); #1;
burst_done = 1'b1; @(posedge clk); #1; burst_done = 1'b0;
if (clk_en !== 1'b0) begin
$error("clk_en still asserted after burst_done"); errors++;
end
// --- LAG: count edges from burst_done until CS releases.
measured = 0;
while (cs_n !== 1'b1) begin
@(posedge clk); #1;
measured++;
if (measured > LAG) break;
end
if (measured != LAG) begin
$error("lag = %0d cycles, expected %0d", measured, LAG); errors++;
end
// --- DESELECT: request immediately and confirm it is held off.
start = 1'b1; // HOLD until it is accepted
repeat (IDLEC - 1) begin
@(posedge clk); #1;
if (cs_n !== 1'b1) begin
$error("CS re-asserted before the deselect interval elapsed");
errors++;
end
end
wait (cs_n === 1'b0); // accepted once RECOVER expires
start = 1'b0;
if (errors == 0)
$display("PASS: lead=%0d, lag=%0d, deselect honoured, clock gated by CS",
LEAD, LAG);
else
$display("FAIL: %0d errors", errors);
$finish;
end
endmodule module spi_cs_guard #(
parameter LEAD_CYCLES = 3,
parameter LAG_CYCLES = 3,
parameter IDLE_CYCLES = 4
) (
input clk,
input rst_n,
input start,
input burst_done,
output reg cs_n,
output clk_en,
output busy
);
// Verilog-2001 has no $clog2 and no typedef enum.
function integer clog2;
input integer value;
integer i;
begin
clog2 = 0;
for (i = value - 1; i > 0; i = i >> 1) clog2 = clog2 + 1;
end
endfunction
localparam IDLE = 3'd0, LEAD = 3'd1, ACTIVE = 3'd2,
LAG = 3'd3, RECOVER = 3'd4;
localparam MAXC = (LEAD_CYCLES > LAG_CYCLES)
? ((LEAD_CYCLES > IDLE_CYCLES) ? LEAD_CYCLES : IDLE_CYCLES)
: ((LAG_CYCLES > IDLE_CYCLES) ? LAG_CYCLES : IDLE_CYCLES);
localparam CW = (MAXC <= 1) ? 1 : clog2(MAXC);
reg [2:0] state;
reg [CW-1:0] cnt;
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
state <= IDLE;
cnt <= {CW{1'b0}};
cs_n <= 1'b1;
end else begin
case (state)
IDLE: begin
cnt <= {CW{1'b0}};
if (start) begin
cs_n <= 1'b0;
state <= LEAD;
end
end
LEAD: begin
if (cnt == (LEAD_CYCLES - 1)) begin
cnt <= {CW{1'b0}};
state <= ACTIVE;
end else begin
cnt <= cnt + 1'b1;
end
end
ACTIVE: begin
cnt <= {CW{1'b0}};
if (burst_done) state <= LAG;
end
LAG: begin
if (cnt == (LAG_CYCLES - 1)) begin
cnt <= {CW{1'b0}};
cs_n <= 1'b1;
state <= RECOVER;
end else begin
cnt <= cnt + 1'b1;
end
end
RECOVER: begin
if (cnt == (IDLE_CYCLES - 1)) begin
cnt <= {CW{1'b0}};
state <= IDLE;
end else begin
cnt <= cnt + 1'b1;
end
end
default: state <= IDLE;
endcase
end
end
assign clk_en = (state == ACTIVE);
assign busy = (state != IDLE);
endmodule module spi_cs_guard_tb;
parameter LEAD = 3, LAG = 3, IDLEC = 4;
reg clk = 1'b0;
reg rst_n, start, burst_done;
wire cs_n, clk_en, busy;
integer errors = 0, measured = 0, i;
spi_cs_guard #(.LEAD_CYCLES(LEAD), .LAG_CYCLES(LAG), .IDLE_CYCLES(IDLEC)) dut (
.clk(clk), .rst_n(rst_n), .start(start), .burst_done(burst_done),
.cs_n(cs_n), .clk_en(clk_en), .busy(busy));
always #5 clk = ~clk;
// A watchdog turns a hang into a FAILURE rather than an infinite run.
initial begin
#10000;
$display("FAIL: timeout - the guard never completed its sequence");
$finish;
end
always @(posedge clk) if (rst_n && clk_en && cs_n) begin
$display("ERROR clk_en asserted while CS deasserted");
errors = errors + 1;
end
initial begin
rst_n = 1'b0; start = 1'b0; burst_done = 1'b0;
@(posedge clk); #1;
if (cs_n !== 1'b1) begin
$display("ERROR reset must deselect"); errors = errors + 1;
end
if (clk_en !== 1'b0) begin
$display("ERROR reset must park the clock"); errors = errors + 1;
end
rst_n = 1'b1;
@(posedge clk); #1;
start = 1'b1; @(posedge clk); #1; start = 1'b0;
if (cs_n !== 1'b0) begin
$display("ERROR CS did not assert on start"); errors = errors + 1;
end
measured = 0;
while (clk_en !== 1'b1 && measured <= LEAD) begin
@(posedge clk); #1;
measured = measured + 1;
end
if (measured != LEAD) begin
$display("ERROR lead = %0d cycles, expected %0d", measured, LEAD);
errors = errors + 1;
end
for (i = 0; i < 6; i = i + 1) @(posedge clk);
#1;
burst_done = 1'b1; @(posedge clk); #1; burst_done = 1'b0;
if (clk_en !== 1'b0) begin
$display("ERROR clk_en still asserted after burst_done");
errors = errors + 1;
end
measured = 0;
while (cs_n !== 1'b1 && measured <= LAG) begin
@(posedge clk); #1;
measured = measured + 1;
end
if (measured != LAG) begin
$display("ERROR lag = %0d cycles, expected %0d", measured, LAG);
errors = errors + 1;
end
start = 1'b1; // hold until accepted
for (i = 0; i < IDLEC - 1; i = i + 1) begin
@(posedge clk); #1;
if (cs_n !== 1'b1) begin
$display("ERROR CS re-asserted before deselect elapsed");
errors = errors + 1;
end
end
wait (cs_n === 1'b0);
start = 1'b0;
if (errors == 0)
$display("PASS: lead=%0d, lag=%0d, deselect honoured, clock gated by CS",
LEAD, LAG);
else
$display("FAIL: %0d errors", errors);
$finish;
end
endmodule library ieee;
use ieee.std_logic_1164.all;
entity spi_cs_guard is
generic (
LEAD_CYCLES : positive := 3;
LAG_CYCLES : positive := 3;
IDLE_CYCLES : positive := 4
);
port (
clk : in std_logic;
rst_n : in std_logic; -- asynchronous, active-low
start : in std_logic;
burst_done : in std_logic;
cs_n : out std_logic;
clk_en : out std_logic;
busy : out std_logic
);
end entity spi_cs_guard;
architecture rtl of spi_cs_guard is
type state_t is (S_IDLE, S_LEAD, S_ACTIVE, S_LAG, S_RECOVER);
signal state : state_t := S_IDLE;
function max3 (a, b, c : positive) return positive is
variable m : positive := a;
begin
if b > m then m := b; end if;
if c > m then m := c; end if;
return m;
end function max3;
constant MAXC : positive := max3(LEAD_CYCLES, LAG_CYCLES, IDLE_CYCLES);
-- A range-constrained integer: no clog2, and simulation range-checks it.
signal cnt : integer range 0 to MAXC - 1 := 0;
signal cs_n_q : std_logic := '1';
begin
fsm : process (clk, rst_n)
begin
if rst_n = '0' then
state <= S_IDLE;
cnt <= 0;
cs_n_q <= '1'; -- reset MUST deselect
elsif rising_edge(clk) then
case state is
when S_IDLE =>
cnt <= 0;
if start = '1' then
cs_n_q <= '0';
state <= S_LEAD;
end if;
when S_LEAD =>
if cnt = LEAD_CYCLES - 1 then
cnt <= 0;
state <= S_ACTIVE;
else
cnt <= cnt + 1;
end if;
when S_ACTIVE =>
cnt <= 0;
if burst_done = '1' then
state <= S_LAG;
end if;
when S_LAG =>
if cnt = LAG_CYCLES - 1 then
cnt <= 0;
cs_n_q <= '1'; -- release only now
state <= S_RECOVER;
else
cnt <= cnt + 1;
end if;
when S_RECOVER =>
if cnt = IDLE_CYCLES - 1 then
cnt <= 0;
state <= S_IDLE;
else
cnt <= cnt + 1;
end if;
end case;
end if;
end process fsm;
cs_n <= cs_n_q;
clk_en <= '1' when state = S_ACTIVE else '0';
busy <= '0' when state = S_IDLE else '1';
end architecture rtl; library ieee;
use ieee.std_logic_1164.all;
entity spi_cs_guard_tb is
end entity spi_cs_guard_tb;
architecture sim of spi_cs_guard_tb is
constant LEAD : positive := 3;
constant LAG : positive := 3;
constant IDLEC : positive := 4;
constant TP : time := 10 ns;
signal clk : std_logic := '0';
signal rst_n : std_logic := '0';
signal start : std_logic := '0';
signal burst_done : std_logic := '0';
signal cs_n : std_logic;
signal clk_en : std_logic;
signal busy : std_logic;
signal done : boolean := false;
signal errors : natural := 0;
begin
clk <= not clk after TP/2 when not done else '0';
dut : entity work.spi_cs_guard
generic map (LEAD_CYCLES => LEAD, LAG_CYCLES => LAG, IDLE_CYCLES => IDLEC)
port map (clk => clk, rst_n => rst_n, start => start,
burst_done => burst_done, cs_n => cs_n,
clk_en => clk_en, busy => busy);
-- Continuous invariant: the clock must never run while CS is deasserted.
invariant : process (clk)
begin
if rising_edge(clk) and rst_n = '1' then
if clk_en = '1' and cs_n = '1' then
report "clk_en asserted while CS deasserted" severity error;
errors <= errors + 1;
end if;
end if;
end process invariant;
-- A watchdog turns a hang into a failure rather than an endless run.
watchdog : process
begin
wait for 10 us;
if not done then
report "timeout - the guard never completed its sequence" severity failure;
end if;
wait;
end process watchdog;
stim : process
variable errs : natural := 0;
variable measured : natural;
begin
wait until rising_edge(clk); wait for 1 ns;
if cs_n /= '1' then
report "reset must deselect" severity error; errs := errs + 1;
end if;
if clk_en /= '0' then
report "reset must park the clock" severity error; errs := errs + 1;
end if;
rst_n <= '1';
wait until rising_edge(clk); wait for 1 ns;
-- LEAD
start <= '1';
wait until rising_edge(clk); wait for 1 ns;
start <= '0';
if cs_n /= '0' then
report "CS did not assert on start" severity error; errs := errs + 1;
end if;
measured := 0;
while clk_en /= '1' and measured <= LEAD loop
wait until rising_edge(clk); wait for 1 ns;
measured := measured + 1;
end loop;
if measured /= LEAD then
report "lead interval incorrect" severity error; errs := errs + 1;
end if;
-- Burst, then completion.
for i in 0 to 5 loop
wait until rising_edge(clk);
end loop;
wait for 1 ns;
burst_done <= '1';
wait until rising_edge(clk); wait for 1 ns;
burst_done <= '0';
if clk_en /= '0' then
report "clk_en still asserted after burst_done" severity error;
errs := errs + 1;
end if;
-- LAG
measured := 0;
while cs_n /= '1' and measured <= LAG loop
wait until rising_edge(clk); wait for 1 ns;
measured := measured + 1;
end loop;
if measured /= LAG then
report "lag interval incorrect" severity error; errs := errs + 1;
end if;
-- DESELECT: request immediately, hold until accepted.
start <= '1';
for i in 0 to IDLEC - 2 loop
wait until rising_edge(clk); wait for 1 ns;
if cs_n /= '1' then
report "CS re-asserted before deselect elapsed" severity error;
errs := errs + 1;
end if;
end loop;
wait until cs_n = '0';
start <= '0';
if errs = 0 and errors = 0 then
report "PASS: lead, lag, deselect and clock gating all correct" severity note;
else
report "FAIL" severity error;
end if;
done <= true;
wait;
end process stim;
end architecture sim;What the three agree on, and where they differ
All three implement the same five-state sequence with the same counter semantics, the same registered CS, the same reset-to-deselected behaviour, and the same rule that clk_en is asserted in exactly one state. The parameters mean the same number of system-clock cycles in each.
The language differences are all about how the state and the counter are declared. SystemVerilog uses typedef enum logic [2:0], which gives the tool both an encoding hint and simulator-visible state names. Verilog-2001 has neither enums nor $clog2, so states are localparam constants and the counter width comes from a constant function — an idiom worth recognising because it appears throughout older production code. VHDL uses an enumerated type with no encoding specified at all, leaving the choice to synthesis, and a range-constrained integer counter that is both self-documenting and range-checked at simulation time. Of the three, the VHDL states the designer's intent most directly and the Verilog states the implementation most directly.
5. Converting a Datasheet Number into Parameters
The parameters above are cycle counts; datasheets specify nanoseconds. The conversion is short and the rounding direction matters.
Suppose a device requires a CS lead of at least 20 ns and the system clock is 100 MHz, a 10 ns period. Then LEAD_CYCLES = ceil(20 / 10) = 2. Round up, always — rounding down produces a gap shorter than the device demands, and the resulting bug is §3's first-byte corruption.
Two cautions that are easy to miss.
The parameter counts cycles of the system clock, not SCLK. That is the design choice §4 justified, and it means the lead interval stays constant in nanoseconds as the SCLK divisor changes. A controller that instead expressed the gap in SCLK periods would silently shrink it whenever someone raised the bus rate — which is exactly when margin is scarcest.
The requirement applies at the device's pins, not at your registers. The interval that matters is between CS arriving at the peripheral and the first clock edge arriving there. Your internal gap is reduced or extended by the difference in propagation between the two paths, and by the difference in your own output delays. At modest frequencies this is negligible; near the limit it is not, and it is Chapter 2.7's territory.
6. Why a Verification Engineer Cares
CS timing is the first thing in this module that a testbench can check thoroughly, because both signals are generated by the master and both are visible on the pins.
A timing monitor is a timestamp recorder. The useful structure is not a protocol decoder but something simpler: record the simulation time of six events per transaction — CS assertion, the first SCLK edge, each launch edge, each sample edge, the last SCLK edge, CS deassertion — and then compute the intervals between them. That single capability supports every check in this chapter and the next two.
// A timing monitor is a stopwatch, not a decoder. It records WHEN, and a
// checker compares those times against a configuration object holding the
// device's requirements. Fragment: no reset handling, no multi-slave
// identity, no transaction reconstruction — Module 16 builds the real agent.
class spi_timing_cfg extends uvm_object;
`uvm_object_utils(spi_timing_cfg)
time t_cs_lead_min; // from the DEVICE datasheet, not from the design
time t_cs_lag_min;
time t_cs_deselect_min;
function new(string name = "spi_timing_cfg"); super.new(name); endfunction
endclass
task automatic watch_transaction(spi_timing_cfg cfg);
time t_cs_fall, t_first_edge, t_last_edge, t_cs_rise;
static time t_prev_cs_rise = 0;
@(negedge vif.cs_n); t_cs_fall = $time;
if (t_prev_cs_rise != 0 && (t_cs_fall - t_prev_cs_rise) < cfg.t_cs_deselect_min)
`uvm_error("SPI_TIMING", $sformatf("deselect %0t < required %0t",
t_cs_fall - t_prev_cs_rise, cfg.t_cs_deselect_min))
@(posedge vif.sclk); t_first_edge = $time; // first edge, idle-low case
if ((t_first_edge - t_cs_fall) < cfg.t_cs_lead_min)
`uvm_error("SPI_TIMING", $sformatf("CS lead %0t < required %0t",
t_first_edge - t_cs_fall, cfg.t_cs_lead_min))
// ... the transfer runs; the transaction monitor of Chapter 1.4 collects data ...
@(posedge vif.cs_n); t_cs_rise = $time;
t_prev_cs_rise = t_cs_rise;
if ((t_cs_rise - t_last_edge) < cfg.t_cs_lag_min)
`uvm_error("SPI_TIMING", $sformatf("CS lag %0t < required %0t",
t_cs_rise - t_last_edge, cfg.t_cs_lag_min))
endtaskThree points about this structure are worth more than the code.
The requirements live in a configuration object, not in the monitor. They are properties of the attached device, so hard-coding them makes the agent single-use — the same argument Chapter 2.2 made about the resting level. A configuration object populated from the datasheet lets one agent verify against any peripheral.
Timing checks are separate from data checks. Chapter 1.4's transaction monitor reconstructs what was exchanged; this checks when the boundaries occurred. A transfer can carry perfect data and violate CS lead, and only a timing check catches it — which is precisely §3's failure.
What this can and cannot establish. In a timing-annotated or behavioural simulation with realistic delays, these intervals are meaningful. In a zero-delay RTL simulation they are measuring the testbench's own clock period, and a passing result says only that the design's sequencing is right — the distinction Chapter 2.4's layer table drew. Neither tells you anything about the intervals at the actual pins on a board.
7. Why an FPGA or ASIC Engineer Cares
CS and SCLK take different paths, and the difference eats your margin. The lead interval the device sees is your internal gap, plus the delay from your CS register to its pin, minus the delay from your SCLK register to its pin, plus the difference in board propagation. If CS leaves through a slow path and SCLK through a fast one, the device sees a shorter lead than you programmed. Registering both in I/O blocks makes the two paths similar and predictable — the same argument Chapter 2.1 made for SCLK alone, now with a reason that involves two signals.
Both are output-delay-constrained paths. CS is not a static configuration signal; it is a timed output whose position relative to SCLK matters. It needs an output delay constraint like any other interface signal, and an unconstrained CS is a path the tool was never asked about. Module 15 covers writing them.
A glitch on CS is a transaction. §4 registered cs_n deliberately. A select that briefly glitches — from combinational decoding, or from a state encoding passing through an intermediate value — can look to a device like a transaction boundary, resetting its bit counter mid-transfer. Chapter 1.5 raised this for multi-slave decoders; it applies to a single select too, and it is one of the reasons a registered output is not merely stylistic.
8. Failure Signature — First Byte Wrong, Rest Fine
Symptom. A transaction returns a first byte that is wrong, while the remaining bytes are well-formed and self-consistent — often exactly the data the device would return if it had received a different opcode. Completely repeatable. Unchanged by lowering the clock.
Candidate mechanisms. A CS lead violation is the leading hypothesis: the device was still preparing when the first edge arrived, so it mis-counted or missed it, and the opcode it assembled is not the one that was sent. The competing explanations are a first-bit launch problem under the edge-role configuration that requires the first bit before the first edge (Chapter 2.3), and a device that was never in a fit state because the previous transaction's deselect interval was violated.
The discriminating observations. First, does lowering SCLK change anything? If not, the fault is not margin — and that single negative result eliminates Chapter 2.4 and Chapter 2.7 entirely, which is a large fraction of the search space for free.
Second, measure the gap between CS falling and the first SCLK edge and compare it against the device's specified minimum. This is a direct test of the hypothesis and needs one capture with both signals visible.
Third, if the lead is adequate, look at the previous transaction's end: measure CS-high time between the two. A deselect violation leaves the device mid-teardown when the next transaction begins, producing an identical first-byte symptom from a completely different cause — and it appears only when transactions are issued back to back, which is why it survives single-shot testing.
Why a back-to-back-only failure is so revealing. If single transactions work perfectly and only bursts of them fail, the deselect interval is almost certainly the mechanism. Nothing else in this chapter distinguishes an isolated transaction from a repeated one.
9. Common Misconceptions
10. Reason It Through
Work this before reading the answers.
A controller talks to a flash device. Single reads issued from a debugger always succeed. The production driver, which issues a stream of reads back to back as fast as the controller allows, gets a corrupt first byte on roughly one read in five — and always on a read that immediately follows another. Lowering SCLK from 20 MHz to 2 MHz changes nothing.
What does the frequency insensitivity establish? That this is not a margin problem. Setup margin scales with the period, so a ten-fold reduction in clock rate would transform any Chapter 2.4 or Chapter 2.7 failure. A fault that ignores a ten-to-one change in frequency is being caused by something whose duration is not derived from the SCLK period — which points squarely at controller sequencing.
What does the back-to-back correlation establish? This is the decisive clue. A transaction that works in isolation and fails only when preceded by another differs in exactly one respect: how long CS was high beforehand. That is the deselect interval, and it is the only requirement in this chapter that a single-shot test cannot exercise. The debugger's transactions are separated by milliseconds of human and host latency; the driver's are separated by whatever the controller does between them, which may be almost nothing.
Why would a short deselect corrupt the first byte specifically? Because the device is still concluding the previous transaction when the next CS assertion arrives. Its state machine has not returned to idle, so the falling edge does not perform the full preparation of §1 — the bit counter may not clear, or the opcode decoder may still be busy. The first byte is therefore mis-assembled while everything afterwards, by which point the device has caught up, is processed normally.
What is the measurement? Capture CS across two consecutive transactions and measure the high time between them. Compare it against the device's specified minimum deselect time. One capture settles it.
Why one in five rather than every time? Because the interval is not constant. The driver's inter-transaction gap depends on what the controller and the software are doing — an interrupt, a bus arbitration delay, or a cache miss can lengthen it enough to satisfy the device. So the failures cluster on the fastest-issued pairs, which is exactly the data-dependent, load-dependent intermittency that makes this bug look mysterious until the mechanism is named.
What is the fix, and what is the trap? The fix is to enforce the minimum in hardware — the RECOVER state of §4's guard, parameterised from the datasheet — so the requirement cannot be bypassed by a driver issuing requests quickly. The trap is to add a software delay between transactions: it works, it is invisible to review, it costs throughput on every transfer, and it will be removed by someone optimising the driver two years later, at which point the bug returns without any hardware having changed.
11. Understanding Check
12. Summary
Chip select is a timed signal, not a static enable. Its falling edge starts real work inside the peripheral — clearing a bit counter, starting a state machine, enabling an output driver, and in one configuration launching the first bit — and that work takes time the master must allow.
Three intervals are the master's obligation. CS lead, from CS asserting to the first SCLK edge. CS lag, from the last SCLK edge to CS deasserting. CS deselect, the minimum CS-high time between transactions. Violating the first corrupts the beginning of a transfer; violating the second corrupts its end; violating the third corrupts a transfer only when it follows another closely, which is why it survives single-shot testing and appears in production.
The failures are perfectly repeatable and largely frequency-independent, because the intervals are set by the controller's sequencing rather than by the SCLK period. That is the strongest diagnostic in the chapter: a first-byte corruption that does not improve when the clock is slowed is pointing here, not at margin.
In hardware the answer is a small sequencer — assert CS, wait, enable the divider, wait after the burst, release, then enforce a recovery interval — with CS registered so it cannot glitch, reset deselecting, and the intervals counted in system-clock cycles so they stay constant in nanoseconds as the bus rate changes. Converting a datasheet requirement into those parameters is a division by the system clock period, rounded up.
In verification the right structure is a stopwatch: timestamp CS assertion, the first and last clock edges, and CS deassertion, then compare the intervals against a configuration object holding the device's published requirements — separate from, and complementary to, the data checking of Chapter 1.4.
13. What Comes Next
So far every requirement in this module has been an obligation the master owes. Chapter 2.6 — Slave Output Valid Timing turns the relationship around and asks what the peripheral owes: how long after a clock edge it may take before MISO carries trustworthy data, why that number depends on the load your board presents rather than only on the device, and why it is the single largest term in the budget Chapter 2.7 finally closes.
Browse the path on the SPI curriculum index, or revisit Setup, Hold, and Timing Margin for the window this chapter applies to a boundary signal.
Continue learning
Related tutorials
- Related topic
The Master Control FSM
The five states an SPI master transfer passes through, why the fifth differs from the first by exactly one behaviour, why outputs must be decoded from state alone, and the failure that disappears the moment you add a print statement.
- Related topic
Chip-Select Generation
Chip select is a state machine, not a wire: the three ways deriving it from a busy signal fails, why the between-frames pause and the between-transactions pause are opposites, and why a select for a slave that is not fitted must be refused.
- Related topic
USB vs SPI
SPI selects a peripheral with a wire routed at layout time and USB with an address the host assigned — so a chip-select contention is invisible to every slave (0 of 11) while a duplicate USB address is detected every time (274 of 274).
- Related topic
Bus Topologies and Daisy-Chain
What the four-wire model costs as devices are added: shared clock and data with one select per device, or a daisy chain that turns several peripherals into one long shift ring. Pin arithmetic, ownership consequences, and why chaining is a device property.
