SPI · Module 13
Chip-Select Generation
Chip select is a state machine, not a wire: the three ways deriving it from a busy signal fails, why the between-frames pause and the between-transactions pause are opposites, and why a select for a slave that is not fitted must be refused.
Everything so far moves bits. Nothing yet decides when a device is listening, and the first implementation everybody writes is one line.
assign cs_n = ~busy;— what is wrong with it?
Three things, and each one is a bug that reaches silicon. One of them is the most common reason a flash driver reads back 0xFF.
1. The Three Failures Of One Line
No lead or lag. Every slave datasheet specifies a setup time from chip select falling to the first SCLK edge (t_CSS, or t_SLCH in flash notation) and a hold from the last edge to chip select rising (t_CSH / t_CHSH). Deriving chip select from busy gives both of them zero. A part that needs 5 ns of lead works at 1 MHz and fails at 20 MHz — which is the worst possible failure mode, because it looks like a signal-integrity problem and sends the investigation to the layout.
It drops chip select between the bytes of one command. A flash read is one transaction of command, address, dummy and data bytes. Chip select rising anywhere inside it ends the transaction: the flash returns to idle and the address it was given is gone. The master must be able to run several frames under one continuous assertion, and a design that raises chip select per frame cannot talk to a flash at all. This is the 0xFF case.
No inter-transaction gap. Slaves specify a minimum chip-select-high time (t_CSD / t_SHSL) before the next assertion. Back-to-back transactions with no gap look legal on a scope and are ignored by the part — which is Chapter 13.3's stale-value failure.
cs_n = ~busy lead = 0 fails above some frequency
lag = 0 the final bit is at risk
per-frame CS a flash transaction cannot exist
gap = 0 the second transaction is ignoredSo chip select gets its own machine.
2. HOLD And GAP: The Same Pause, The Opposite Pin
The machine has six states, and the interesting one is the one that distinguishes it from Chapter 13.3's:
IDLE nothing selected; a request is accepted
LEAD selected, clock stopped — paying t_CSS
XFER the clock runs
HOLD selected, clock stopped — between FRAMES of one transaction
LAG selected, clock stopped — paying t_CSH at the end
GAP deselected — paying t_CSD before anything elseHOLD and GAP are both "the clock is stopped and we are waiting". They differ in exactly one thing, and it is the thing that matters:
HOLD chip select stays LOW the transaction continues
GAP chip select goes HIGH the transaction is overThat is the whole of multi-frame support. A transaction is a sequence of frames under one assertion, so between frames the machine pauses with the select still asserted; between transactions it pauses with the select released. Same counter, same waiting, opposite pin.
3. The Machine
4. Why The Abort Input Lives Here
Chapter 13.10 is the chapter about abandoning a transfer, and the abort input is in this block. That placement is a decision worth explaining here, because it was arrived at by getting it wrong first.
The first attempt kept this block untouched and aborted from outside: synthesise a "frame done" pulse and simultaneously pull hold low, so the machine would fall out of XFER into LAG and pay both intervals on its own. It does not work, and the reason is structural: this machine latches hold when the frame is requested. Mid-frame the live input is not consulted, so the machine went to HOLD — clock stopped, chip select still low, waiting for a request that the aborting logic was busy refusing. The bus hung with the slave still selected: the exact failure the abort exists to prevent, produced by the code meant to prevent it.
So the abort input belongs where the pins are owned. It does not release chip select — it redirects the machine into LAG, so the programmed hold and the CS-high gap are still paid in full. Releasing the select on the spot would violate the slave's hold time on the way out, and a slave that samples one last edge inside that violation may latch a bit that was never meant for it.
there is exactly one piece of logic that knows how to leave the bus,
and it is this one5. What The Intervals Look Like
Read the cs_n row across cycles 1 to 14. One continuous assertion covering two frames and the pause between them — which is what a flash transaction requires and what cs_n = ~busy cannot produce.
6. A Select For A Slave That Is Not There
The controller decodes sel into a one-hot-low select. With N_CS slaves on a SEL_W-bit select and N_CS < 2**SEL_W, some encodings name nothing:
N_CS = 3, SEL_W = 2 sel = 0, 1, 2 are fitted; sel = 3 is notA truncating decoder asserts slave 0 for sel = 3, which turns a driver bug into a write to the wrong device. On a bus with a flash and a sensor that is a corrupted flash image produced by a sensor driver, and it is very hard to attribute.
So the controller publishes sel_err and refuses the request. It is worth noting how the comparison must be written:
sel_err = (sel >= N_CS[SEL_W-1:0]) WRONG
sel_err = (sel > (N_CS - 1)) rightWith N_CS = 4 and SEL_W = 2, N_CS truncated to two bits is zero, so the first form rejects every request instead of none — a design that refuses to work at all, which at least fails loudly. With N_CS = 3 it truncates to 3 and rejects only sel = 3 by coincidence. Comparing as an integer is correct for every combination.
7. Building the Chip-Select Controller — Three HDLs
The circuit
Six states, one shared interval counter, a one-hot-low decoder and a latched select. Three details:
The counters load with the interval minus one. The cycle a state is entered on is already part of that interval. Loading the full value gives every timing parameter one cycle more than asked for — harmless on lead and lag, and a waste of real throughput on gap, paid on every transaction for the life of the product.
The select is latched for the whole transaction. sel is sampled on entry from IDLE and held, so a mid-transaction change of the requested slave cannot move the select. That is the same argument as Chapter 13.2's configuration snapshot, applied to the one field whose mid-transaction change would be catastrophic rather than merely wrong.
Reset releases every select line, asynchronously. A master coming out of reset while a slave still sees chip select low has that slave mid-transaction with a master that has forgotten about it. This is the one place in the design where the reset style is not a preference.
// spi_cs_ctrl.sv
//
// Chapter 13.7 -- chip select is a state machine, not a wire.
//
// The naive version of chip select is one line of code:
//
// assign cs_n = ~busy; // and this is wrong
//
// It is wrong three times over, and each one is a bug that reaches silicon.
//
// 1. NO LEAD OR LAG. Every slave datasheet specifies a setup time from CS
// falling to the first SCLK edge (t_CSS / t_SLCH) and a hold from the
// last edge to CS rising (t_CSH / t_CHSH). Deriving CS from `busy` gives
// both of them zero. A part that needs 5 ns of lead works at 1 MHz and
// fails at 20 MHz, which is the worst possible failure mode because it
// looks like a signal-integrity problem.
//
// 2. IT DROPS CS BETWEEN THE BYTES OF ONE COMMAND. A flash read is one
// transaction of command, address, dummy and data bytes. CS rising
// anywhere inside it ENDS the transaction -- the flash returns to idle
// and the address it was given is gone. The master must be able to run
// several frames under one continuous assertion. This is the single most
// common reason a flash driver reads back 0xFF.
//
// 3. NO INTER-TRANSACTION GAP. Slaves specify a minimum CS-high time
// (t_CSD / t_SHSL) before the next assertion. Back-to-back transactions
// with no gap are legal-looking on a scope and ignored by the part.
//
// So chip select gets its own small machine:
//
// IDLE -> LEAD -> XFER -> HOLD -> XFER ... (a burst)
// \-> LAG -> GAP -> IDLE (burst ends)
//
// HOLD is the state that distinguishes this block from the transfer FSM of
// Chapter 13.3. Its GAP means "between transactions, CS high". HOLD means
// "between FRAMES of one transaction, CS still low" -- the same pause on the
// clock, the opposite thing on the select pin.
module spi_cs_ctrl #(
parameter int N_CS = 4, // how many slaves hang off this master
parameter int SEL_W = 2, // ceil(log2(N_CS)), supplied not derived
parameter int CNT_W = 8 // width of the lead/lag/gap counters
) (
input wire clk,
input wire rst_n,
input wire req, // a frame is wanted
input wire hold, // ... and it is not the last one
input wire [SEL_W-1:0] sel, // which slave
input wire [CNT_W-1:0] lead_cyc, // CS low -> first SCLK edge
input wire [CNT_W-1:0] lag_cyc, // last SCLK edge -> CS high
input wire [CNT_W-1:0] gap_cyc, // CS high -> CS low again
input wire core_done, // the shift engine finished a frame
// ABORT: abandon the transaction NOW, but leave the bus legally. It does
// not release chip select -- it redirects the machine into LAG, so the
// programmed hold and the CS-high gap are still paid in full. Releasing CS
// on the spot would violate the slave's hold time on the way out, and a
// slave that samples one last edge inside that violation may latch a bit
// that was never meant for it. This input exists here rather than in the
// supervisor of Chapter 13.10 because this machine is the one that owns
// the pins, and there should be exactly one piece of logic that knows how
// to leave the bus.
input wire abort,
output reg [N_CS-1:0] cs_n, // active low, one-hot-low
output wire shift_en, // gates the divider of 13.4
output wire start_stb, // one cycle, at each frame's start
output wire busy,
output wire sel_err, // `sel` named a slave that is not there
output wire [2:0] state_id // for waveform capture
);
localparam [2:0] S_IDLE = 3'd0,
S_LEAD = 3'd1,
S_XFER = 3'd2,
S_HOLD = 3'd3,
S_LAG = 3'd4,
S_GAP = 3'd5;
reg [2:0] state;
reg [CNT_W-1:0] cnt;
reg [SEL_W-1:0] sel_q; // latched for the whole burst
reg hold_q;
reg start_r;
// A `sel` outside the installed range must not silently decode to slave
// zero -- which is what a truncating one-hot decoder does, and it is how
// a driver bug becomes a write to the wrong device.
// Compared as an integer, NOT against a truncated N_CS: with N_CS = 4
// and SEL_W = 2, `N_CS[SEL_W-1:0]` is zero and the comparison would
// reject every request instead of none.
assign sel_err = (sel > (N_CS - 1));
assign shift_en = (state == S_XFER);
assign start_stb = start_r;
assign busy = (state != S_IDLE);
assign state_id = state;
// The counters are loaded with the interval MINUS ONE, because the cycle
// the state is entered on is already part of the interval. Loading the
// full value gives every timing parameter one cycle more than asked for,
// which is harmless on lead and lag and wastes real throughput on gap.
wire [CNT_W-1:0] lead_m1 = (lead_cyc == {CNT_W{1'b0}})
? {CNT_W{1'b0}} : lead_cyc - 1'b1;
wire [CNT_W-1:0] lag_m1 = (lag_cyc == {CNT_W{1'b0}})
? {CNT_W{1'b0}} : lag_cyc - 1'b1;
wire [CNT_W-1:0] gap_m1 = (gap_cyc == {CNT_W{1'b0}})
? {CNT_W{1'b0}} : gap_cyc - 1'b1;
wire cnt_done = (cnt == {CNT_W{1'b0}});
integer k;
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
// Reset MUST release every select line, and asynchronously. A
// master coming out of reset while a slave still sees CS low has
// that slave mid-transaction with a master that has forgotten
// about it.
state <= S_IDLE;
cnt <= {CNT_W{1'b0}};
cs_n <= {N_CS{1'b1}};
sel_q <= {SEL_W{1'b0}};
hold_q <= 1'b0;
start_r <= 1'b0;
end else begin
start_r <= 1'b0;
case (state)
S_IDLE: begin
cs_n <= {N_CS{1'b1}};
if (req && !sel_err) begin
sel_q <= sel;
hold_q <= hold;
for (k = 0; k < N_CS; k = k + 1)
cs_n[k] <= ~(sel == k[SEL_W-1:0]);
cnt <= lead_m1;
state <= S_LEAD;
end
end
S_LEAD: begin
// CS is already low; the clock is not running yet. This
// is the whole of t_CSS.
if (abort) begin
cnt <= lag_m1;
state <= S_LAG;
end else if (cnt_done) begin
start_r <= 1'b1;
state <= S_XFER;
end else begin
cnt <= cnt - 1'b1;
end
end
S_XFER: begin
if (core_done || abort) begin
// An abort always ends the TRANSACTION, never merely
// the frame, so it overrides the burst's own hold.
if (hold_q && !abort) begin
// Another frame in the same transaction: pause
// the clock, keep CS asserted.
cnt <= gap_m1;
state <= S_HOLD;
end else begin
cnt <= lag_m1;
state <= S_LAG;
end
end
end
S_HOLD: begin
// CS STAYS LOW here. That is the entire point of the
// state, and the reason this machine is not the one in
// Chapter 13.3.
//
// Which also means an abort arriving HERE still has a
// chip select to release, and still owes the hold time for
// the edges of the frame that just finished.
if (abort) begin
cnt <= lag_m1;
state <= S_LAG;
end else if (!cnt_done) begin
cnt <= cnt - 1'b1;
end else if (req) begin
hold_q <= hold;
start_r <= 1'b1;
state <= S_XFER;
end
end
S_LAG: begin
if (cnt_done) begin
cs_n <= {N_CS{1'b1}};
cnt <= gap_m1;
state <= S_GAP;
end else begin
cnt <= cnt - 1'b1;
end
end
S_GAP: begin
// CS is high and must stay high. A request arriving now
// is not refused, it is simply not acted on until the
// gap has been paid.
if (cnt_done) begin
state <= S_IDLE;
end else begin
cnt <= cnt - 1'b1;
end
end
default: state <= S_IDLE;
endcase
end
end
endmodule// spi_cs_ctrl_tb.sv
//
// Every timing figure here is MEASURED off the pins -- cycles from CS falling
// to the first SCLK edge, from the last edge to CS rising, from CS rising to
// the next CS falling -- and compared against what was asked for. Nothing is
// inferred from the state machine's internals, because the state machine is
// what is on trial.
//
// The checks are two-sided on purpose. Too SHORT violates the slave's
// datasheet. Too LONG is a correctness-preserving bug that quietly costs
// throughput on every transaction for the life of the product, and it is the
// one nobody ever finds.
`timescale 1ns/1ps
module spi_cs_ctrl_tb;
// Three slaves on a two-bit select, so that `sel == 3` is a request for
// hardware that is not installed and `sel_err` has something to reject.
localparam int N_CS = 3;
localparam int SEL_W = 2;
localparam int CNT_W = 8;
localparam int DIV_W = 8;
logic clk = 1'b0;
logic rst_n = 1'b0;
always #5 clk = ~clk;
logic req = 1'b0;
logic hold = 1'b0;
logic [SEL_W-1:0] sel = 2'd0;
logic [CNT_W-1:0] lead_cyc = 8'd4;
logic [CNT_W-1:0] lag_cyc = 8'd4;
logic [CNT_W-1:0] gap_cyc = 8'd6;
logic core_done = 1'b0;
// Held low for the whole of this chapter's tests: the abort path is
// Chapter 13.10's subject, and the point here is that adding the input
// changes nothing about normal operation.
logic abort = 1'b0;
wire [N_CS-1:0] cs_n;
wire shift_en, start_stb, busy, sel_err;
wire [2:0] state_id;
spi_cs_ctrl #(.N_CS(N_CS), .SEL_W(SEL_W), .CNT_W(CNT_W)) dut (
.clk(clk), .rst_n(rst_n),
.req(req), .hold(hold), .sel(sel),
.lead_cyc(lead_cyc), .lag_cyc(lag_cyc), .gap_cyc(gap_cyc),
.core_done(core_done), .abort(abort),
.cs_n(cs_n), .shift_en(shift_en), .start_stb(start_stb),
.busy(busy), .sel_err(sel_err), .state_id(state_id)
);
// The real divider, gated by the controller, so the SCLK edges the
// measurements are taken against are the ones a slave would see.
logic [DIV_W-1:0] div = 8'd4;
logic cpol = 1'b0;
wire sclk, edge_a_stb, edge_b_stb, bit_done, div_err;
wire [DIV_W-1:0] half_a, half_b;
spi_clkdiv_strobe #(.DIV_W(DIV_W)) u_div (
.clk(clk), .rst_n(rst_n), .en(shift_en), .div(div), .cpol(cpol),
.sclk(sclk), .edge_a_stb(edge_a_stb), .edge_b_stb(edge_b_stb),
.bit_done(bit_done), .half_a(half_a), .half_b(half_b),
.div_err(div_err)
);
// --- the shift engine, standing in for Chapters 13.5 and 13.6 ---------
// It exists only to say "frame finished" after the right number of bit
// periods, which is all this block needs from it.
logic [5:0] frame_len = 6'd8;
logic [5:0] bits_left;
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
bits_left <= 6'd0;
core_done <= 1'b0;
end else begin
core_done <= 1'b0;
if (start_stb)
bits_left <= frame_len;
else if (shift_en && bit_done) begin
if (bits_left <= 6'd1) begin
bits_left <= 6'd0;
core_done <= 1'b1;
end else begin
bits_left <= bits_left - 6'd1;
end
end
end
end
// --- the pin monitor --------------------------------------------------
wire any_cs_low = ~(&cs_n);
logic any_cs_low_q;
integer c_since_fall, c_since_edge, c_since_rise;
integer f_now, e_now, r_now;
integer saw_first_edge;
integer min_lead, min_lag, min_gap;
integer max_lead, max_lag, max_gap;
integer n_lead, n_lag, n_gap;
integer multi_low; // two selects low at once -- must stay zero
integer low_while_idle; // a select low outside a transaction
integer edge_while_high; // an SCLK edge with no CS asserted
integer cs_rises; // how many times CS went back high
integer i, lowbits;
always_ff @(posedge clk) begin
if (!rst_n) begin
c_since_fall <= 0; c_since_edge <= 0; c_since_rise <= 0;
any_cs_low_q <= 1'b0;
saw_first_edge <= 0;
end else begin
any_cs_low_q <= any_cs_low;
// Snapshot before anything is reloaded, so a counter that is
// being restarted this cycle is still read at its old value.
f_now = c_since_fall;
e_now = c_since_edge;
r_now = c_since_rise;
if (any_cs_low && !any_cs_low_q) begin // CS just fell
c_since_fall <= 1;
saw_first_edge <= 0;
if (n_gap > 0 || cs_rises > 0) begin
if (r_now < min_gap) min_gap <= r_now;
if (r_now > max_gap) max_gap <= r_now;
n_gap <= n_gap + 1;
end
end else begin
c_since_fall <= f_now + 1;
end
if (!any_cs_low && any_cs_low_q) begin // CS just rose
c_since_rise <= 1;
cs_rises <= cs_rises + 1;
if (e_now < min_lag) min_lag <= e_now;
if (e_now > max_lag) max_lag <= e_now;
n_lag <= n_lag + 1;
end else begin
c_since_rise <= r_now + 1;
end
if (edge_a_stb || edge_b_stb) begin
c_since_edge <= 1;
if (!any_cs_low) edge_while_high <= edge_while_high + 1;
if (!saw_first_edge) begin
saw_first_edge <= 1;
if (f_now < min_lead) min_lead <= f_now;
if (f_now > max_lead) max_lead <= f_now;
n_lead <= n_lead + 1;
end
end else begin
c_since_edge <= e_now + 1;
end
// At most one select may be low, ever.
lowbits = 0;
for (i = 0; i < N_CS; i = i + 1)
if (!cs_n[i]) lowbits = lowbits + 1;
if (lowbits > 1) multi_low <= multi_low + 1;
if (lowbits > 0 && !busy) low_while_idle <= low_while_idle + 1;
end
end
task automatic clear_stats;
begin
min_lead = 9999; max_lead = 0; n_lead = 0;
min_lag = 9999; max_lag = 0; n_lag = 0;
min_gap = 9999; max_gap = 0; n_gap = 0;
end
endtask
integer errors = 0;
// A transaction: `nframes` frames under one continuous assertion.
task automatic transaction(input integer slave, input integer nframes,
input integer nbits);
integer guard, f;
begin
frame_len = nbits[5:0];
for (f = 0; f < nframes; f = f + 1) begin
@(negedge clk);
sel = slave[SEL_W-1:0];
hold = (f < nframes - 1);
req = 1'b1;
// Hold `req` until the controller acts on it, which is what a
// level request means.
guard = 4000;
while (!start_stb && guard > 0) begin
@(negedge clk);
guard = guard - 1;
end
req = 1'b0;
if (guard == 0) begin
$display(" FAIL: the controller never started a frame");
errors = errors + 1;
end
guard = 4000;
while (!core_done && guard > 0) begin
@(negedge clk);
guard = guard - 1;
end
if (guard == 0) begin
$display(" FAIL: the frame never finished");
errors = errors + 1;
end
end
// Deliberately does NOT wait for the machine to walk out through
// LAG and GAP. The next request is left pending while it does,
// which is how a real driver behaves and which makes the measured
// CS-high time the minimum the CONTROLLER enforces rather than
// the time the testbench took to ask again.
end
endtask
task automatic wait_idle;
integer guard;
begin
guard = 4000;
while (busy && guard > 0) begin
@(negedge clk);
guard = guard - 1;
end
repeat (2) @(negedge clk);
end
endtask
task automatic check_window(input integer got, input integer want,
input string what);
begin
if (got < want) begin
$display(" FAIL: %0s was %0d cycles, below the %0d asked for",
what, got, want);
errors = errors + 1;
end
if (got > want + 3) begin
$display(" FAIL: %0s was %0d cycles, wasting %0d beyond the %0d asked for",
what, got, got - want, want);
errors = errors + 1;
end
end
endtask
integer leads [0:3];
integer lags [0:3];
integer gaps [0:3];
integer a, b, c, s, nf;
initial begin
clear_stats();
multi_low = 0; low_while_idle = 0; edge_while_high = 0; cs_rises = 0;
saw_first_edge = 0;
c_since_fall = 0; c_since_edge = 0; c_since_rise = 0;
n_lead = 0; n_lag = 0; n_gap = 0;
leads[0] = 1; leads[1] = 2; leads[2] = 5; leads[3] = 12;
lags[0] = 1; lags[1] = 3; lags[2] = 6; lags[3] = 10;
gaps[0] = 1; gaps[1] = 4; gaps[2] = 8; gaps[3] = 16;
repeat (3) @(negedge clk);
// 1. RESET RELEASES EVERY SELECT, before anything else happens.
if (cs_n !== {N_CS{1'b1}}) begin
$display(" FAIL: reset did not release every select line");
errors = errors + 1;
end
$display(" reset: all %0d select lines released", N_CS);
rst_n = 1'b1;
@(negedge clk);
// 2. A SINGLE FRAME with generous windows.
lead_cyc = 8'd5; lag_cyc = 8'd6; gap_cyc = 8'd8;
clear_stats();
transaction(1, 1, 8);
wait_idle();
check_window(min_lead, 5, "the lead into the first edge");
check_window(min_lag, 6, "the lag out of the last edge");
$display(" one frame on slave 1: lead %0d (asked 5), lag %0d (asked 6)",
min_lead, min_lag);
// 3. A BURST. Four frames, ONE assertion. If CS drops between them a
// flash would abandon the transaction, so this counts rises.
clear_stats();
cs_rises = 0;
transaction(2, 4, 8);
wait_idle();
if (cs_rises != 1) begin
$display(" FAIL: a four-frame transaction raised CS %0d times",
cs_rises);
errors = errors + 1;
end
$display(" four frames, one assertion: CS rose %0d time -- the transaction was never broken",
cs_rises);
// 4. BACK-TO-BACK TRANSACTIONS must pay the gap.
transaction(0, 1, 8);
clear_stats();
transaction(0, 1, 8);
transaction(1, 1, 8);
wait_idle();
check_window(min_gap, 8, "the gap between transactions");
$display(" three transactions back to back: shortest CS-high gap %0d cycles (asked 8)",
min_gap);
// 5. A REQUEST FOR A SLAVE THAT IS NOT THERE. With three selects on
// two bits, `sel = 3` decodes to nothing; a truncating decoder
// would assert slave 0 instead.
@(negedge clk);
sel = 2'd3; hold = 1'b0; req = 1'b1;
// `sel_err` is combinational, so it settles a delta after `sel` is
// driven; reading it in the same statement sequence sees the old
// value and reports a flag that is in fact working.
@(negedge clk);
if (!sel_err) begin
$display(" FAIL: sel=3 was not flagged with only three slaves fitted");
errors = errors + 1;
end
repeat (30) @(negedge clk);
if (busy || cs_n !== {N_CS{1'b1}}) begin
$display(" FAIL: a request for a missing slave asserted something");
errors = errors + 1;
end
req = 1'b0; sel = 2'd0;
wait_idle();
$display(" sel=3 with three slaves fitted: flagged, and no select asserted");
// 6. THE SWEEP. Every combination of lead, lag and gap, on every
// slave, as single frames and as bursts.
for (a = 0; a < 4; a = a + 1)
for (b = 0; b < 4; b = b + 1)
for (c = 0; c < 4; c = c + 1) begin
lead_cyc = leads[a][CNT_W-1:0];
lag_cyc = lags[b][CNT_W-1:0];
gap_cyc = gaps[c][CNT_W-1:0];
s = (a + b + c) % N_CS;
nf = 1 + ((a + c) % 3);
// The first transaction runs under the NEW windows but
// its own CS fall closes a gap that was paid under the
// OLD ones, so the statistics start after it.
transaction(s, 1, 8);
clear_stats();
transaction(s, nf, 8);
wait_idle();
check_window(min_lead, leads[a], "a swept lead");
check_window(min_lag, lags[b], "a swept lag");
check_window(min_gap, gaps[c], "a swept gap");
end
$display(" 64 (lead, lag, gap) combinations swept across all three slaves, single frames and bursts");
// 7. THE CONTINUOUS PROPERTIES.
if (multi_low != 0) begin
$display(" FAIL: two selects were low together on %0d cycles",
multi_low);
errors = errors + 1;
end
if (low_while_idle != 0) begin
$display(" FAIL: a select was low on %0d cycles with no transaction",
low_while_idle);
errors = errors + 1;
end
if (edge_while_high != 0) begin
$display(" FAIL: %0d SCLK edges happened with no slave selected",
edge_while_high);
errors = errors + 1;
end
$display(" across the whole run: never two selects low, never a select low outside a transaction, never an SCLK edge with nothing selected");
if (errors == 0)
$display("PASS: chip select comes out of reset released, asserts one line and only one line, and holds it low across every frame of a multi-frame transaction so the transaction is never broken -- the measured lead from CS falling to the first SCLK edge, the lag from the last edge to CS rising, and the CS-high gap between transactions all meet what was programmed without overshooting it, across 64 combinations of the three windows on all three slaves as single frames and as bursts -- no SCLK edge ever occurs with nothing selected, and a request naming a slave that is not fitted is flagged and refused rather than decoding to slave zero");
else
$display("FAIL: %0d error(s)", errors);
$finish;
end
endmodule// spi_cs_ctrl.v
//
// Chapter 13.7 -- chip select is a state machine, not a wire.
//
// The naive version of chip select is one line of code:
//
// assign cs_n = ~busy; // and this is wrong
//
// It is wrong three times over, and each one is a bug that reaches silicon.
//
// 1. NO LEAD OR LAG. Every slave datasheet specifies a setup time from CS
// falling to the first SCLK edge (t_CSS / t_SLCH) and a hold from the
// last edge to CS rising (t_CSH / t_CHSH). Deriving CS from `busy` gives
// both of them zero. A part that needs 5 ns of lead works at 1 MHz and
// fails at 20 MHz, which is the worst possible failure mode because it
// looks like a signal-integrity problem.
//
// 2. IT DROPS CS BETWEEN THE BYTES OF ONE COMMAND. A flash read is one
// transaction of command, address, dummy and data bytes. CS rising
// anywhere inside it ENDS the transaction -- the flash returns to idle
// and the address it was given is gone. The master must be able to run
// several frames under one continuous assertion. This is the single most
// common reason a flash driver reads back 0xFF.
//
// 3. NO INTER-TRANSACTION GAP. Slaves specify a minimum CS-high time
// (t_CSD / t_SHSL) before the next assertion. Back-to-back transactions
// with no gap are legal-looking on a scope and ignored by the part.
//
// So chip select gets its own small machine:
//
// IDLE -> LEAD -> XFER -> HOLD -> XFER ... (a burst)
// \-> LAG -> GAP -> IDLE (burst ends)
//
// HOLD is the state that distinguishes this block from the transfer FSM of
// Chapter 13.3. Its GAP means "between transactions, CS high". HOLD means
// "between FRAMES of one transaction, CS still low" -- the same pause on the
// clock, the opposite thing on the select pin.
module spi_cs_ctrl #(
parameter N_CS = 4, // how many slaves hang off this master
parameter SEL_W = 2, // ceil(log2(N_CS)), supplied not derived
parameter CNT_W = 8 // width of the lead/lag/gap counters
) (
input wire clk,
input wire rst_n,
input wire req, // a frame is wanted
input wire hold, // ... and it is not the last one
input wire [SEL_W-1:0] sel, // which slave
input wire [CNT_W-1:0] lead_cyc, // CS low -> first SCLK edge
input wire [CNT_W-1:0] lag_cyc, // last SCLK edge -> CS high
input wire [CNT_W-1:0] gap_cyc, // CS high -> CS low again
input wire core_done, // the shift engine finished a frame
// ABORT: abandon the transaction NOW, but leave the bus legally. It does
// not release chip select -- it redirects the machine into LAG, so the
// programmed hold and the CS-high gap are still paid in full. Releasing CS
// on the spot would violate the slave's hold time on the way out, and a
// slave that samples one last edge inside that violation may latch a bit
// that was never meant for it. This input exists here rather than in the
// supervisor of Chapter 13.10 because this machine is the one that owns
// the pins, and there should be exactly one piece of logic that knows how
// to leave the bus.
input wire abort,
output reg [N_CS-1:0] cs_n, // active low, one-hot-low
output wire shift_en, // gates the divider of 13.4
output wire start_stb, // one cycle, at each frame's start
output wire busy,
output wire sel_err, // `sel` named a slave that is not there
output wire [2:0] state_id // for waveform capture
);
localparam [2:0] S_IDLE = 3'd0,
S_LEAD = 3'd1,
S_XFER = 3'd2,
S_HOLD = 3'd3,
S_LAG = 3'd4,
S_GAP = 3'd5;
reg [2:0] state;
reg [CNT_W-1:0] cnt;
reg [SEL_W-1:0] sel_q; // latched for the whole burst
reg hold_q;
reg start_r;
// A `sel` outside the installed range must not silently decode to slave
// zero -- which is what a truncating one-hot decoder does, and it is how
// a driver bug becomes a write to the wrong device.
// Compared as an integer, NOT against a truncated N_CS: with N_CS = 4
// and SEL_W = 2, `N_CS[SEL_W-1:0]` is zero and the comparison would
// reject every request instead of none.
assign sel_err = (sel > (N_CS - 1));
assign shift_en = (state == S_XFER);
assign start_stb = start_r;
assign busy = (state != S_IDLE);
assign state_id = state;
// The counters are loaded with the interval MINUS ONE, because the cycle
// the state is entered on is already part of the interval. Loading the
// full value gives every timing parameter one cycle more than asked for,
// which is harmless on lead and lag and wastes real throughput on gap.
wire [CNT_W-1:0] lead_m1 = (lead_cyc == {CNT_W{1'b0}})
? {CNT_W{1'b0}} : lead_cyc - 1'b1;
wire [CNT_W-1:0] lag_m1 = (lag_cyc == {CNT_W{1'b0}})
? {CNT_W{1'b0}} : lag_cyc - 1'b1;
wire [CNT_W-1:0] gap_m1 = (gap_cyc == {CNT_W{1'b0}})
? {CNT_W{1'b0}} : gap_cyc - 1'b1;
wire cnt_done = (cnt == {CNT_W{1'b0}});
integer k;
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
// Reset MUST release every select line, and asynchronously. A
// master coming out of reset while a slave still sees CS low has
// that slave mid-transaction with a master that has forgotten
// about it.
state <= S_IDLE;
cnt <= {CNT_W{1'b0}};
cs_n <= {N_CS{1'b1}};
sel_q <= {SEL_W{1'b0}};
hold_q <= 1'b0;
start_r <= 1'b0;
end else begin
start_r <= 1'b0;
case (state)
S_IDLE: begin
cs_n <= {N_CS{1'b1}};
if (req && !sel_err) begin
sel_q <= sel;
hold_q <= hold;
for (k = 0; k < N_CS; k = k + 1)
cs_n[k] <= ~(sel == k[SEL_W-1:0]);
cnt <= lead_m1;
state <= S_LEAD;
end
end
S_LEAD: begin
// CS is already low; the clock is not running yet. This
// is the whole of t_CSS.
if (abort) begin
cnt <= lag_m1;
state <= S_LAG;
end else if (cnt_done) begin
start_r <= 1'b1;
state <= S_XFER;
end else begin
cnt <= cnt - 1'b1;
end
end
S_XFER: begin
if (core_done || abort) begin
// An abort always ends the TRANSACTION, never merely
// the frame, so it overrides the burst's own hold.
if (hold_q && !abort) begin
// Another frame in the same transaction: pause
// the clock, keep CS asserted.
cnt <= gap_m1;
state <= S_HOLD;
end else begin
cnt <= lag_m1;
state <= S_LAG;
end
end
end
S_HOLD: begin
// CS STAYS LOW here. That is the entire point of the
// state, and the reason this machine is not the one in
// Chapter 13.3.
//
// Which also means an abort arriving HERE still has a
// chip select to release, and still owes the hold time for
// the edges of the frame that just finished.
if (abort) begin
cnt <= lag_m1;
state <= S_LAG;
end else if (!cnt_done) begin
cnt <= cnt - 1'b1;
end else if (req) begin
hold_q <= hold;
start_r <= 1'b1;
state <= S_XFER;
end
end
S_LAG: begin
if (cnt_done) begin
cs_n <= {N_CS{1'b1}};
cnt <= gap_m1;
state <= S_GAP;
end else begin
cnt <= cnt - 1'b1;
end
end
S_GAP: begin
// CS is high and must stay high. A request arriving now
// is not refused, it is simply not acted on until the
// gap has been paid.
if (cnt_done) begin
state <= S_IDLE;
end else begin
cnt <= cnt - 1'b1;
end
end
default: state <= S_IDLE;
endcase
end
end
endmodule// spi_cs_ctrl_tb.v
//
// Every timing figure here is MEASURED off the pins -- cycles from CS falling
// to the first SCLK edge, from the last edge to CS rising, from CS rising to
// the next CS falling -- and compared against what was asked for. Nothing is
// inferred from the state machine's internals, because the state machine is
// what is on trial.
//
// The checks are two-sided on purpose. Too SHORT violates the slave's
// datasheet. Too LONG is a correctness-preserving bug that quietly costs
// throughput on every transaction for the life of the product, and it is the
// one nobody ever finds.
`timescale 1ns/1ps
module spi_cs_ctrl_tb;
// Three slaves on a two-bit select, so that `sel == 3` is a request for
// hardware that is not installed and `sel_err` has something to reject.
localparam N_CS = 3;
localparam SEL_W = 2;
localparam CNT_W = 8;
localparam DIV_W = 8;
reg clk;
reg rst_n;
always #5 clk = ~clk;
reg req;
reg hold;
reg [SEL_W-1:0] sel;
reg [CNT_W-1:0] lead_cyc;
reg [CNT_W-1:0] lag_cyc;
reg [CNT_W-1:0] gap_cyc;
reg core_done;
// Held low for the whole of this chapter's tests: the abort path is
// Chapter 13.10's subject, and the point here is that adding the input
// changes nothing about normal operation.
reg abort;
wire [N_CS-1:0] cs_n;
wire shift_en, start_stb, busy, sel_err;
wire [2:0] state_id;
spi_cs_ctrl #(.N_CS(N_CS), .SEL_W(SEL_W), .CNT_W(CNT_W)) dut (
.clk(clk), .rst_n(rst_n),
.req(req), .hold(hold), .sel(sel),
.lead_cyc(lead_cyc), .lag_cyc(lag_cyc), .gap_cyc(gap_cyc),
.core_done(core_done), .abort(abort),
.cs_n(cs_n), .shift_en(shift_en), .start_stb(start_stb),
.busy(busy), .sel_err(sel_err), .state_id(state_id)
);
// The real divider, gated by the controller, so the SCLK edges the
// measurements are taken against are the ones a slave would see.
reg [DIV_W-1:0] div;
reg cpol;
wire sclk, edge_a_stb, edge_b_stb, bit_done, div_err;
wire [DIV_W-1:0] half_a, half_b;
spi_clkdiv_strobe #(.DIV_W(DIV_W)) u_div (
.clk(clk), .rst_n(rst_n), .en(shift_en), .div(div), .cpol(cpol),
.sclk(sclk), .edge_a_stb(edge_a_stb), .edge_b_stb(edge_b_stb),
.bit_done(bit_done), .half_a(half_a), .half_b(half_b),
.div_err(div_err)
);
// --- the shift engine, standing in for Chapters 13.5 and 13.6 ---------
// It exists only to say "frame finished" after the right number of bit
// periods, which is all this block needs from it.
reg [5:0] frame_len;
reg [5:0] bits_left;
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
bits_left <= 6'd0;
core_done <= 1'b0;
end else begin
core_done <= 1'b0;
if (start_stb)
bits_left <= frame_len;
else if (shift_en && bit_done) begin
if (bits_left <= 6'd1) begin
bits_left <= 6'd0;
core_done <= 1'b1;
end else begin
bits_left <= bits_left - 6'd1;
end
end
end
end
// --- the pin monitor --------------------------------------------------
wire any_cs_low = ~(&cs_n);
reg any_cs_low_q;
integer c_since_fall, c_since_edge, c_since_rise;
integer f_now, e_now, r_now;
integer saw_first_edge;
integer min_lead, min_lag, min_gap;
integer max_lead, max_lag, max_gap;
integer n_lead, n_lag, n_gap;
integer multi_low; // two selects low at once -- must stay zero
integer low_while_idle; // a select low outside a transaction
integer edge_while_high; // an SCLK edge with no CS asserted
integer cs_rises; // how many times CS went back high
integer i, lowbits;
always @(posedge clk) begin
if (!rst_n) begin
c_since_fall <= 0; c_since_edge <= 0; c_since_rise <= 0;
any_cs_low_q <= 1'b0;
saw_first_edge <= 0;
end else begin
any_cs_low_q <= any_cs_low;
// Snapshot before anything is reloaded, so a counter that is
// being restarted this cycle is still read at its old value.
f_now = c_since_fall;
e_now = c_since_edge;
r_now = c_since_rise;
if (any_cs_low && !any_cs_low_q) begin // CS just fell
c_since_fall <= 1;
saw_first_edge <= 0;
if (n_gap > 0 || cs_rises > 0) begin
if (r_now < min_gap) min_gap <= r_now;
if (r_now > max_gap) max_gap <= r_now;
n_gap <= n_gap + 1;
end
end else begin
c_since_fall <= f_now + 1;
end
if (!any_cs_low && any_cs_low_q) begin // CS just rose
c_since_rise <= 1;
cs_rises <= cs_rises + 1;
if (e_now < min_lag) min_lag <= e_now;
if (e_now > max_lag) max_lag <= e_now;
n_lag <= n_lag + 1;
end else begin
c_since_rise <= r_now + 1;
end
if (edge_a_stb || edge_b_stb) begin
c_since_edge <= 1;
if (!any_cs_low) edge_while_high <= edge_while_high + 1;
if (!saw_first_edge) begin
saw_first_edge <= 1;
if (f_now < min_lead) min_lead <= f_now;
if (f_now > max_lead) max_lead <= f_now;
n_lead <= n_lead + 1;
end
end else begin
c_since_edge <= e_now + 1;
end
// At most one select may be low, ever.
lowbits = 0;
for (i = 0; i < N_CS; i = i + 1)
if (!cs_n[i]) lowbits = lowbits + 1;
if (lowbits > 1) multi_low <= multi_low + 1;
if (lowbits > 0 && !busy) low_while_idle <= low_while_idle + 1;
end
end
task clear_stats;
begin
min_lead = 9999; max_lead = 0; n_lead = 0;
min_lag = 9999; max_lag = 0; n_lag = 0;
min_gap = 9999; max_gap = 0; n_gap = 0;
end
endtask
integer errors;
// A transaction: `nframes` frames under one continuous assertion.
task transaction;
input integer slave;
input integer nframes;
input integer nbits;
integer guard, f;
begin
frame_len = nbits[5:0];
for (f = 0; f < nframes; f = f + 1) begin
@(negedge clk);
sel = slave[SEL_W-1:0];
hold = (f < nframes - 1);
req = 1'b1;
// Hold `req` until the controller acts on it, which is what a
// level request means.
guard = 4000;
while (!start_stb && guard > 0) begin
@(negedge clk);
guard = guard - 1;
end
req = 1'b0;
if (guard == 0) begin
$display(" FAIL: the controller never started a frame");
errors = errors + 1;
end
guard = 4000;
while (!core_done && guard > 0) begin
@(negedge clk);
guard = guard - 1;
end
if (guard == 0) begin
$display(" FAIL: the frame never finished");
errors = errors + 1;
end
end
// Deliberately does NOT wait for the machine to walk out through
// LAG and GAP. The next request is left pending while it does,
// which is how a real driver behaves and which makes the measured
// CS-high time the minimum the CONTROLLER enforces rather than
// the time the testbench took to ask again.
end
endtask
task wait_idle;
integer guard;
begin
guard = 4000;
while (busy && guard > 0) begin
@(negedge clk);
guard = guard - 1;
end
repeat (2) @(negedge clk);
end
endtask
task check_window;
input integer got;
input integer want;
input [8*40:1] what;
begin
if (got < want) begin
$display(" FAIL: %0s was %0d cycles, below the %0d asked for",
what, got, want);
errors = errors + 1;
end
if (got > want + 3) begin
$display(" FAIL: %0s was %0d cycles, wasting %0d beyond the %0d asked for",
what, got, got - want, want);
errors = errors + 1;
end
end
endtask
integer leads [0:3];
integer lags [0:3];
integer gaps [0:3];
integer a, b, c, s, nf;
initial begin
clear_stats();
multi_low = 0; low_while_idle = 0; edge_while_high = 0; cs_rises = 0;
saw_first_edge = 0;
c_since_fall = 0; c_since_edge = 0; c_since_rise = 0;
n_lead = 0; n_lag = 0; n_gap = 0;
leads[0] = 1; leads[1] = 2; leads[2] = 5; leads[3] = 12;
lags[0] = 1; lags[1] = 3; lags[2] = 6; lags[3] = 10;
gaps[0] = 1; gaps[1] = 4; gaps[2] = 8; gaps[3] = 16;
repeat (3) @(negedge clk);
// 1. RESET RELEASES EVERY SELECT, before anything else happens.
if (cs_n !== {N_CS{1'b1}}) begin
$display(" FAIL: reset did not release every select line");
errors = errors + 1;
end
$display(" reset: all %0d select lines released", N_CS);
rst_n = 1'b1;
@(negedge clk);
// 2. A SINGLE FRAME with generous windows.
lead_cyc = 8'd5; lag_cyc = 8'd6; gap_cyc = 8'd8;
clear_stats();
transaction(1, 1, 8);
wait_idle();
check_window(min_lead, 5, "the lead into the first edge");
check_window(min_lag, 6, "the lag out of the last edge");
$display(" one frame on slave 1: lead %0d (asked 5), lag %0d (asked 6)",
min_lead, min_lag);
// 3. A BURST. Four frames, ONE assertion. If CS drops between them a
// flash would abandon the transaction, so this counts rises.
clear_stats();
cs_rises = 0;
transaction(2, 4, 8);
wait_idle();
if (cs_rises != 1) begin
$display(" FAIL: a four-frame transaction raised CS %0d times",
cs_rises);
errors = errors + 1;
end
$display(" four frames, one assertion: CS rose %0d time -- the transaction was never broken",
cs_rises);
// 4. BACK-TO-BACK TRANSACTIONS must pay the gap.
transaction(0, 1, 8);
clear_stats();
transaction(0, 1, 8);
transaction(1, 1, 8);
wait_idle();
check_window(min_gap, 8, "the gap between transactions");
$display(" three transactions back to back: shortest CS-high gap %0d cycles (asked 8)",
min_gap);
// 5. A REQUEST FOR A SLAVE THAT IS NOT THERE. With three selects on
// two bits, `sel = 3` decodes to nothing; a truncating decoder
// would assert slave 0 instead.
@(negedge clk);
sel = 2'd3; hold = 1'b0; req = 1'b1;
// `sel_err` is combinational, so it settles a delta after `sel` is
// driven; reading it in the same statement sequence sees the old
// value and reports a flag that is in fact working.
@(negedge clk);
if (!sel_err) begin
$display(" FAIL: sel=3 was not flagged with only three slaves fitted");
errors = errors + 1;
end
repeat (30) @(negedge clk);
if (busy || cs_n !== {N_CS{1'b1}}) begin
$display(" FAIL: a request for a missing slave asserted something");
errors = errors + 1;
end
req = 1'b0; sel = 2'd0;
wait_idle();
$display(" sel=3 with three slaves fitted: flagged, and no select asserted");
// 6. THE SWEEP. Every combination of lead, lag and gap, on every
// slave, as single frames and as bursts.
for (a = 0; a < 4; a = a + 1)
for (b = 0; b < 4; b = b + 1)
for (c = 0; c < 4; c = c + 1) begin
lead_cyc = leads[a][CNT_W-1:0];
lag_cyc = lags[b][CNT_W-1:0];
gap_cyc = gaps[c][CNT_W-1:0];
s = (a + b + c) % N_CS;
nf = 1 + ((a + c) % 3);
// The first transaction runs under the NEW windows but
// its own CS fall closes a gap that was paid under the
// OLD ones, so the statistics start after it.
transaction(s, 1, 8);
clear_stats();
transaction(s, nf, 8);
wait_idle();
check_window(min_lead, leads[a], "a swept lead");
check_window(min_lag, lags[b], "a swept lag");
check_window(min_gap, gaps[c], "a swept gap");
end
$display(" 64 (lead, lag, gap) combinations swept across all three slaves, single frames and bursts");
// 7. THE CONTINUOUS PROPERTIES.
if (multi_low != 0) begin
$display(" FAIL: two selects were low together on %0d cycles",
multi_low);
errors = errors + 1;
end
if (low_while_idle != 0) begin
$display(" FAIL: a select was low on %0d cycles with no transaction",
low_while_idle);
errors = errors + 1;
end
if (edge_while_high != 0) begin
$display(" FAIL: %0d SCLK edges happened with no slave selected",
edge_while_high);
errors = errors + 1;
end
$display(" across the whole run: never two selects low, never a select low outside a transaction, never an SCLK edge with nothing selected");
if (errors == 0)
$display("PASS: chip select comes out of reset released, asserts one line and only one line, and holds it low across every frame of a multi-frame transaction so the transaction is never broken -- the measured lead from CS falling to the first SCLK edge, the lag from the last edge to CS rising, and the CS-high gap between transactions all meet what was programmed without overshooting it, across 64 combinations of the three windows on all three slaves as single frames and as bursts -- no SCLK edge ever occurs with nothing selected, and a request naming a slave that is not fitted is flagged and refused rather than decoding to slave zero");
else
$display("FAIL: %0d error(s)", errors);
$finish;
end
initial begin
clk = 1'b0;
rst_n = 1'b0;
req = 1'b0;
hold = 1'b0;
sel = 2'd0;
lead_cyc = 8'd4;
lag_cyc = 8'd4;
gap_cyc = 8'd6;
core_done = 1'b0;
abort = 1'b0;
div = 8'd4;
cpol = 1'b0;
frame_len = 6'd8;
errors = 0;
end
endmodule-- spi_cs_ctrl.vhd
--
-- Chapter 13.7 -- chip select is a state machine, not a wire.
--
-- The naive version is one line, `cs_n <= not busy`, and it is wrong three
-- times over:
--
-- 1. NO LEAD OR LAG. Every slave specifies a setup from CS falling to the
-- first SCLK edge and a hold from the last edge to CS rising. Deriving
-- CS from `busy` gives both of them zero, so the part works at 1 MHz and
-- fails at 20 -- the worst failure mode, because it looks like signal
-- integrity.
-- 2. IT DROPS CS BETWEEN THE BYTES OF ONE COMMAND. A flash read is one
-- transaction of command, address, dummy and data. CS rising anywhere
-- inside it ENDS the transaction. This is the most common reason a
-- flash driver reads back 0xFF.
-- 3. NO INTER-TRANSACTION GAP. Slaves specify a minimum CS-high time before
-- the next assertion.
--
-- So chip select gets its own small machine:
--
-- IDLE -> LEAD -> XFER -> HOLD -> XFER ... (a burst)
-- \-> LAG -> GAP -> IDLE (burst ends)
--
-- HOLD is what distinguishes this block from the transfer FSM of Chapter
-- 13.3. That machine's GAP means "between transactions, CS high". HOLD means
-- "between FRAMES of one transaction, CS still low" -- the same pause on the
-- clock, the opposite thing on the select pin.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_cs_ctrl is
generic (
N_CS : positive := 4; -- how many slaves hang off this master
SEL_W : positive := 2; -- ceil(log2(N_CS)), supplied not derived
CNT_W : positive := 8 -- width of the lead/lag/gap counters
);
port (
clk : in std_logic;
rst_n : in std_logic;
req : in std_logic; -- a frame is wanted
hold : in std_logic; -- .. and not the last
sel : in unsigned(SEL_W - 1 downto 0); -- which slave
lead_cyc : in unsigned(CNT_W - 1 downto 0); -- CS low -> 1st edge
lag_cyc : in unsigned(CNT_W - 1 downto 0); -- last edge -> CS high
gap_cyc : in unsigned(CNT_W - 1 downto 0); -- CS high -> CS low
core_done : in std_logic; -- the shift engine finished a frame
-- ABORT: abandon the transaction NOW, but leave the bus legally. It
-- does not release chip select -- it redirects the machine into LAG, so
-- the programmed hold and the CS-high gap are still paid in full.
-- Releasing CS on the spot would violate the slave's hold time on the
-- way out. This input lives here rather than in the supervisor of
-- Chapter 13.10 because this machine owns the pins, and there should be
-- exactly one piece of logic that knows how to leave the bus.
abort : in std_logic;
cs_n : out std_logic_vector(N_CS - 1 downto 0); -- active low
shift_en : out std_logic; -- gates the divider of 13.4
start_stb : out std_logic; -- one cycle, at each frame's start
busy : out std_logic;
sel_err : out std_logic; -- `sel` named a slave that is not there
state_id : out unsigned(2 downto 0)
);
end entity;
architecture rtl of spi_cs_ctrl is
constant S_IDLE : unsigned(2 downto 0) := "000";
constant S_LEAD : unsigned(2 downto 0) := "001";
constant S_XFER : unsigned(2 downto 0) := "010";
constant S_HOLD : unsigned(2 downto 0) := "011";
constant S_LAG : unsigned(2 downto 0) := "100";
constant S_GAP : unsigned(2 downto 0) := "101";
signal state : unsigned(2 downto 0) := S_IDLE;
signal cnt : unsigned(CNT_W - 1 downto 0) := (others => '0');
signal sel_q : unsigned(SEL_W - 1 downto 0) := (others => '0');
signal hold_q : std_logic := '0';
signal start_r : std_logic := '0';
signal cs_r : std_logic_vector(N_CS - 1 downto 0) := (others => '1');
signal err_i : std_logic;
-- The counters are loaded with the interval MINUS ONE, because the cycle
-- the state is entered on is already part of the interval. Loading the
-- full value gives every timing parameter one cycle more than asked for,
-- which is harmless on lead and lag and wastes real throughput on gap.
function minus1(v : unsigned) return unsigned is
begin
if v = 0 then
return v;
else
return v - 1;
end if;
end function;
begin
-- A `sel` outside the installed range must not silently decode to slave
-- zero -- which is what a truncating one-hot decoder does, and it is how
-- a driver bug becomes a write to the wrong device. Compared as an
-- integer, not against a truncated N_CS.
err_i <= '1' when to_integer(sel) > (N_CS - 1) else '0';
sel_err <= err_i;
cs_n <= cs_r;
shift_en <= '1' when state = S_XFER else '0';
start_stb <= start_r;
busy <= '0' when state = S_IDLE else '1';
state_id <= state;
fsm : process (clk, rst_n)
begin
if rst_n = '0' then
-- Reset MUST release every select line, and asynchronously. A
-- master coming out of reset while a slave still sees CS low has
-- that slave mid-transaction with a master that has forgotten
-- about it.
state <= S_IDLE;
cnt <= (others => '0');
cs_r <= (others => '1');
sel_q <= (others => '0');
hold_q <= '0';
start_r <= '0';
elsif rising_edge(clk) then
start_r <= '0';
case to_integer(state) is
when 0 => -- S_IDLE
cs_r <= (others => '1');
if req = '1' and err_i = '0' then
sel_q <= sel;
hold_q <= hold;
for k in 0 to N_CS - 1 loop
if to_integer(sel) = k then
cs_r(k) <= '0';
else
cs_r(k) <= '1';
end if;
end loop;
cnt <= minus1(lead_cyc);
state <= S_LEAD;
end if;
when 1 => -- S_LEAD
-- CS is already low; the clock is not running yet. This
-- is the whole of t_CSS.
if abort = '1' then
cnt <= minus1(lag_cyc);
state <= S_LAG;
elsif cnt = 0 then
start_r <= '1';
state <= S_XFER;
else
cnt <= cnt - 1;
end if;
when 2 => -- S_XFER
if core_done = '1' or abort = '1' then
-- An abort always ends the TRANSACTION, never merely
-- the frame, so it overrides the burst's own hold.
if hold_q = '1' and abort = '0' then
-- Another frame in the same transaction: pause
-- the clock, keep CS asserted.
cnt <= minus1(gap_cyc);
state <= S_HOLD;
else
cnt <= minus1(lag_cyc);
state <= S_LAG;
end if;
end if;
when 3 => -- S_HOLD
-- CS STAYS LOW here. That is the entire point of the
-- state, and the reason this machine is not the one in
-- Chapter 13.3.
--
-- Which also means an abort arriving HERE still has a chip
-- select to release, and still owes the hold time for the
-- edges of the frame that just finished.
if abort = '1' then
cnt <= minus1(lag_cyc);
state <= S_LAG;
elsif cnt /= 0 then
cnt <= cnt - 1;
elsif req = '1' then
hold_q <= hold;
start_r <= '1';
state <= S_XFER;
end if;
when 4 => -- S_LAG
if cnt = 0 then
cs_r <= (others => '1');
cnt <= minus1(gap_cyc);
state <= S_GAP;
else
cnt <= cnt - 1;
end if;
when 5 => -- S_GAP
-- CS is high and must stay high. A request arriving now
-- is not refused, it is simply not acted on until the gap
-- has been paid.
if cnt = 0 then
state <= S_IDLE;
else
cnt <= cnt - 1;
end if;
when others =>
state <= S_IDLE;
end case;
end if;
end process;
end architecture;-- spi_cs_ctrl_tb.vhd
--
-- Every timing figure here is MEASURED off the pins -- cycles from CS falling
-- to the first SCLK edge, from the last edge to CS rising, from CS rising to
-- the next CS falling -- and compared against what was asked for. Nothing is
-- inferred from the state machine's internals, because the state machine is
-- what is on trial.
--
-- The checks are two-sided on purpose. Too SHORT violates the slave's
-- datasheet. Too LONG is a correctness-preserving bug that quietly costs
-- throughput on every transaction for the life of the product.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_cs_ctrl_tb is
end entity;
architecture sim of spi_cs_ctrl_tb is
-- Three slaves on a two-bit select, so that `sel = 3` is a request for
-- hardware that is not installed and `sel_err` has something to reject.
constant N_CS : positive := 3;
constant SEL_W : positive := 2;
constant CNT_W : positive := 8;
constant DIV_W : positive := 8;
signal clk : std_logic := '0';
signal rst_n : std_logic := '0';
signal halt : boolean := false;
signal req : std_logic := '0';
signal hold : std_logic := '0';
signal sel : unsigned(SEL_W - 1 downto 0) := (others => '0');
signal lead_cyc : unsigned(CNT_W - 1 downto 0) := to_unsigned(4, CNT_W);
signal lag_cyc : unsigned(CNT_W - 1 downto 0) := to_unsigned(4, CNT_W);
signal gap_cyc : unsigned(CNT_W - 1 downto 0) := to_unsigned(6, CNT_W);
signal core_done : std_logic := '0';
-- Held low for the whole of this chapter's tests: the abort path is
-- Chapter 13.10's subject, and the point here is that adding the input
-- changes nothing about normal operation.
signal abort : std_logic := '0';
signal cs_n : std_logic_vector(N_CS - 1 downto 0);
signal shift_en, start_stb, busy, sel_err : std_logic;
signal state_id : unsigned(2 downto 0);
signal div : unsigned(DIV_W - 1 downto 0) := to_unsigned(4, DIV_W);
signal cpol : std_logic := '0';
signal sclk, edge_a_stb, edge_b_stb, bit_done, div_err : std_logic;
signal half_a, half_b : unsigned(DIV_W - 1 downto 0);
signal frame_len : natural := 8;
signal frame_set : natural := 8;
signal any_cs_low, any_cs_low_q : std_logic := '0';
signal min_lead, min_lag, min_gap : natural := 9999;
signal max_lead, max_lag, max_gap : natural := 0;
signal n_lead, n_lag, n_gap : natural := 0;
signal clear_stb : std_logic := '0';
signal multi_low : natural := 0;
signal low_while_idle : natural := 0;
signal edge_while_high : natural := 0;
signal cs_rises : natural := 0;
signal clear_rises : std_logic := '0';
signal errors : natural := 0;
begin
clk <= not clk after 5 ns when not halt else '0';
dut : entity work.spi_cs_ctrl
generic map (N_CS => N_CS, SEL_W => SEL_W, CNT_W => CNT_W)
port map (clk => clk, rst_n => rst_n,
req => req, hold => hold, sel => sel,
lead_cyc => lead_cyc, lag_cyc => lag_cyc, gap_cyc => gap_cyc,
core_done => core_done, abort => abort,
cs_n => cs_n, shift_en => shift_en, start_stb => start_stb,
busy => busy, sel_err => sel_err, state_id => state_id);
-- The real divider, gated by the controller, so the SCLK edges the
-- measurements are taken against are the ones a slave would see.
u_div : entity work.spi_clkdiv_strobe
generic map (DIV_W => DIV_W)
port map (clk => clk, rst_n => rst_n, en => shift_en, div => div,
cpol => cpol, sclk => sclk,
edge_a_stb => edge_a_stb, edge_b_stb => edge_b_stb,
bit_done => bit_done, half_a => half_a, half_b => half_b,
div_err => div_err);
-- The shift engine, standing in for Chapters 13.5 and 13.6. It exists
-- only to say "frame finished" after the right number of bit periods.
core : process (clk, rst_n)
variable bits_left : natural := 0;
begin
if rst_n = '0' then
bits_left := 0;
core_done <= '0';
elsif rising_edge(clk) then
core_done <= '0';
if start_stb = '1' then
bits_left := frame_len;
elsif shift_en = '1' and bit_done = '1' then
if bits_left <= 1 then
bits_left := 0;
core_done <= '1';
else
bits_left := bits_left - 1;
end if;
end if;
end if;
end process;
any_cs_low <= '0' when cs_n = (cs_n'range => '1') else '1';
monitor : process (clk)
variable f_now, e_now, r_now : natural;
variable since_fall, since_edge, since_rise : natural := 0;
variable saw_first_edge : boolean := false;
variable lowbits : natural;
begin
if rising_edge(clk) then
if rst_n = '0' then
since_fall := 0; since_edge := 0; since_rise := 0;
any_cs_low_q <= '0';
saw_first_edge := false;
else
any_cs_low_q <= any_cs_low;
if clear_stb = '1' then
min_lead <= 9999; max_lead <= 0; n_lead <= 0;
min_lag <= 9999; max_lag <= 0; n_lag <= 0;
min_gap <= 9999; max_gap <= 0; n_gap <= 0;
end if;
if clear_rises = '1' then
cs_rises <= 0;
end if;
-- Snapshot before anything is reloaded, so a counter that is
-- being restarted this cycle is still read at its old value.
f_now := since_fall;
e_now := since_edge;
r_now := since_rise;
if any_cs_low = '1' and any_cs_low_q = '0' then -- CS fell
since_fall := 1;
saw_first_edge := false;
if cs_rises > 0 then
if r_now < min_gap then min_gap <= r_now; end if;
if r_now > max_gap then max_gap <= r_now; end if;
n_gap <= n_gap + 1;
end if;
else
since_fall := f_now + 1;
end if;
if any_cs_low = '0' and any_cs_low_q = '1' then -- CS rose
since_rise := 1;
if clear_rises = '0' then
cs_rises <= cs_rises + 1;
end if;
if e_now < min_lag then min_lag <= e_now; end if;
if e_now > max_lag then max_lag <= e_now; end if;
n_lag <= n_lag + 1;
else
since_rise := r_now + 1;
end if;
if edge_a_stb = '1' or edge_b_stb = '1' then
since_edge := 1;
if any_cs_low = '0' then
edge_while_high <= edge_while_high + 1;
end if;
if not saw_first_edge then
saw_first_edge := true;
if f_now < min_lead then min_lead <= f_now; end if;
if f_now > max_lead then max_lead <= f_now; end if;
n_lead <= n_lead + 1;
end if;
else
since_edge := e_now + 1;
end if;
-- At most one select may be low, ever.
lowbits := 0;
for k in 0 to N_CS - 1 loop
if cs_n(k) = '0' then lowbits := lowbits + 1; end if;
end loop;
if lowbits > 1 then
multi_low <= multi_low + 1;
end if;
if lowbits > 0 and busy = '0' then
low_while_idle <= low_while_idle + 1;
end if;
end if;
end if;
end process;
stim : process
variable errs : natural := 0;
variable guard : natural;
procedure clear_stats is
begin
clear_stb <= '1';
wait until falling_edge(clk);
clear_stb <= '0';
end procedure;
-- A transaction: `nframes` frames under one continuous assertion.
-- Deliberately does NOT wait for the machine to walk out through LAG
-- and GAP: the next request is left pending while it does, which is
-- how a real driver behaves and which makes the measured CS-high time
-- the minimum the CONTROLLER enforces rather than the time the
-- testbench took to ask again.
procedure transaction(slave : natural; nframes : natural;
nbits : natural) is
begin
frame_set <= nbits;
for f in 0 to nframes - 1 loop
wait until falling_edge(clk);
sel <= to_unsigned(slave, SEL_W);
if f < nframes - 1 then hold <= '1'; else hold <= '0'; end if;
req <= '1';
guard := 4000;
while start_stb = '0' and guard > 0 loop
wait until falling_edge(clk);
guard := guard - 1;
end loop;
req <= '0';
if guard = 0 then
report " FAIL: the controller never started a frame";
errs := errs + 1;
end if;
guard := 4000;
while core_done = '0' and guard > 0 loop
wait until falling_edge(clk);
guard := guard - 1;
end loop;
if guard = 0 then
report " FAIL: the frame never finished";
errs := errs + 1;
end if;
end loop;
end procedure;
procedure wait_idle is
begin
guard := 4000;
while busy = '1' and guard > 0 loop
wait until falling_edge(clk);
guard := guard - 1;
end loop;
for k in 1 to 2 loop wait until falling_edge(clk); end loop;
end procedure;
procedure check_window(got : natural; want : natural; what : string) is
begin
if got < want then
report " FAIL: " & what & " was " & integer'image(got) &
" cycles, below the " & integer'image(want) &
" asked for";
errs := errs + 1;
end if;
if got > want + 3 then
report " FAIL: " & what & " was " & integer'image(got) &
" cycles, wasting " & integer'image(got - want) &
" beyond the " & integer'image(want) & " asked for";
errs := errs + 1;
end if;
end procedure;
type int_vec is array (natural range <>) of natural;
constant LEADS : int_vec(0 to 3) := (1, 2, 5, 12);
constant LAGS : int_vec(0 to 3) := (1, 3, 6, 10);
constant GAPS : int_vec(0 to 3) := (1, 4, 8, 16);
variable s, nf : natural;
constant ALL_HIGH : std_logic_vector(N_CS - 1 downto 0)
:= (others => '1');
begin
for k in 1 to 3 loop wait until falling_edge(clk); end loop;
-- 1. RESET RELEASES EVERY SELECT, before anything else happens.
if cs_n /= ALL_HIGH then
report " FAIL: reset did not release every select line";
errs := errs + 1;
end if;
report " reset: all " & integer'image(N_CS) & " select lines released";
rst_n <= '1';
wait until falling_edge(clk);
-- 2. A SINGLE FRAME with generous windows.
lead_cyc <= to_unsigned(5, CNT_W);
lag_cyc <= to_unsigned(6, CNT_W);
gap_cyc <= to_unsigned(8, CNT_W);
clear_stats;
transaction(1, 1, 8);
wait_idle;
check_window(min_lead, 5, "the lead into the first edge");
check_window(min_lag, 6, "the lag out of the last edge");
report " one frame on slave 1: lead " & integer'image(min_lead) &
" (asked 5), lag " & integer'image(min_lag) & " (asked 6)";
-- 3. A BURST. Four frames, ONE assertion. If CS drops between them a
-- flash would abandon the transaction, so this counts rises.
clear_stats;
clear_rises <= '1';
wait until falling_edge(clk);
clear_rises <= '0';
transaction(2, 4, 8);
wait_idle;
if cs_rises /= 1 then
report " FAIL: a four-frame transaction raised CS " &
integer'image(cs_rises) & " times";
errs := errs + 1;
end if;
report " four frames, one assertion: CS rose " &
integer'image(cs_rises) &
" time -- the transaction was never broken";
-- 4. BACK-TO-BACK TRANSACTIONS must pay the gap.
transaction(0, 1, 8);
clear_stats;
transaction(0, 1, 8);
transaction(1, 1, 8);
wait_idle;
check_window(min_gap, 8, "the gap between transactions");
report " three transactions back to back: shortest CS-high gap " &
integer'image(min_gap) & " cycles (asked 8)";
-- 5. A REQUEST FOR A SLAVE THAT IS NOT THERE. With three selects on
-- two bits, `sel = 3` decodes to nothing; a truncating decoder
-- would assert slave 0 instead.
wait until falling_edge(clk);
sel <= to_unsigned(3, SEL_W); hold <= '0'; req <= '1';
wait until falling_edge(clk);
if sel_err /= '1' then
report " FAIL: sel=3 was not flagged with only three slaves fitted";
errs := errs + 1;
end if;
for k in 1 to 30 loop wait until falling_edge(clk); end loop;
if busy /= '0' or cs_n /= ALL_HIGH then
report " FAIL: a request for a missing slave asserted something";
errs := errs + 1;
end if;
req <= '0'; sel <= to_unsigned(0, SEL_W);
wait_idle;
report " sel=3 with three slaves fitted: flagged, and no select asserted";
-- 6. THE SWEEP. Every combination of lead, lag and gap, on every
-- slave, as single frames and as bursts.
for a in LEADS'range loop
for b in LAGS'range loop
for c in GAPS'range loop
lead_cyc <= to_unsigned(LEADS(a), CNT_W);
lag_cyc <= to_unsigned(LAGS(b), CNT_W);
gap_cyc <= to_unsigned(GAPS(c), CNT_W);
s := (a + b + c) mod N_CS;
nf := 1 + ((a + c) mod 3);
-- The first transaction runs under the NEW windows but
-- its own CS fall closes a gap that was paid under the
-- OLD ones, so the statistics start after it.
transaction(s, 1, 8);
clear_stats;
transaction(s, nf, 8);
wait_idle;
check_window(min_lead, LEADS(a), "a swept lead");
check_window(min_lag, LAGS(b), "a swept lag");
check_window(min_gap, GAPS(c), "a swept gap");
end loop;
end loop;
end loop;
report " 64 (lead, lag, gap) combinations swept across all three slaves, single frames and bursts";
-- 7. THE CONTINUOUS PROPERTIES.
if multi_low /= 0 then
report " FAIL: two selects were low together";
errs := errs + 1;
end if;
if low_while_idle /= 0 then
report " FAIL: a select was low with no transaction running";
errs := errs + 1;
end if;
if edge_while_high /= 0 then
report " FAIL: an SCLK edge happened with no slave selected";
errs := errs + 1;
end if;
report " across the whole run: never two selects low, never a select low outside a transaction, never an SCLK edge with nothing selected";
errors <= errs;
if errs = 0 then
report "PASS: chip select comes out of reset released, asserts one line and only one line, and holds it low across every frame of a multi-frame transaction so the transaction is never broken -- the measured lead from CS falling to the first SCLK edge, the lag from the last edge to CS rising, and the CS-high gap between transactions all meet what was programmed without overshooting it, across 64 combinations of the three windows on all three slaves as single frames and as bursts -- no SCLK edge ever occurs with nothing selected, and a request naming a slave that is not fitted is flagged and refused rather than decoding to slave zero";
else
report "FAIL: " & integer'image(errs) & " error(s)" severity error;
end if;
halt <= true;
wait;
end process;
frame_len <= frame_set;
end architecture;Parity
All three implementations sweep 64 combinations of lead, lag and gap across all three slaves, as single frames and as bursts, measuring every interval from the pins. All three report a four-frame transaction raising chip select exactly once, and all three confirm that no SCLK edge ever occurs with nothing selected.
8. Why a Verification Engineer Cares
// The timing properties are two-sided on purpose: too short violates the slave's
// datasheet, too long is a correctness-preserving bug that costs throughput on
// every transaction for the life of the product and that nobody ever finds.
module spi_cs_ctrl_sva #(
parameter int N_CS = 4,
parameter int CNT_W = 8
) (
input logic clk,
input logic rst_n,
input logic req,
input logic hold,
input logic abort,
input logic [CNT_W-1:0] lead_cyc,
input logic [CNT_W-1:0] lag_cyc,
input logic [CNT_W-1:0] gap_cyc,
input logic [N_CS-1:0] cs_n,
input logic shift_en,
input logic start_stb,
input logic busy,
input logic sel_err
);
default clocking cb @(posedge clk); endclocking
default disable iff (!rst_n);
wire any_low = ~(&cs_n);
// At most one select low, ever. Written with $countones so it fails on the
// cycle two go low rather than on a downstream data mismatch.
a_one_hot: assert property ($countones(~cs_n) <= 1);
// No select low outside a transaction, and no clock edge with nothing
// selected -- requirement R1 from Chapter 13.1, enforced where it is caused.
a_low_implies_busy: assert property (any_low |-> busy);
a_shift_implies_low: assert property (shift_en |-> any_low);
// The lead: once the select falls, the clock may not run for lead_cyc
// cycles. Stated about `shift_en` rather than about SCLK, because this block
// does not drive SCLK and should not be held to another block's output.
property p_lead;
$fell(any_low) |=> (!shift_en)[*lead_cyc-1];
endproperty
a_lead: assert property (p_lead);
// The lag: the clock must have been stopped for lag_cyc cycles before the
// select is allowed to rise.
property p_lag;
$rose(any_low) |-> $past(!shift_en, 1) && $past(!shift_en, lag_cyc-1);
endproperty
a_lag: assert property (p_lag);
// The gap: once released, the select stays released for gap_cyc cycles --
// regardless of how insistently a request is asserted.
property p_gap;
$rose(any_low) |=> (!any_low)[*gap_cyc-1];
endproperty
a_gap: assert property (p_gap);
// THE burst property. A frame that reports `hold` must not release the
// select, which is the whole of multi-frame support in one line.
a_hold_keeps_cs: assert property (
start_stb && hold |-> ##[1:$] (any_low throughout (!start_stb)[*1:$])
);
// Stated more usably as: the select may only rise out of LAG, never out of
// HOLD -- so a rise must be preceded by a frame that was NOT held.
a_rise_needs_last: assert property (
$rose(any_low) |-> !$past(hold_latched)
);
// A slave that is not fitted is refused, not truncated.
a_sel_err_refused: assert property (sel_err && req && !busy |=> !any_low);
a_sel_err_flagged: assert property ((sel > (N_CS - 1)) == sel_err);
// An abort must end the TRANSACTION, never merely the frame: the select
// must rise, and it must not rise early.
a_abort_ends: assert property (abort && busy |-> ##[1:$] $rose(any_low));
endmodule// The intervals' absolute values are uninteresting above about four cycles. What
// matters is the value ONE -- where "minus one" makes the counter degenerate --
// and the burst length, because HOLD is only exercised by a transaction with
// more than one frame.
covergroup cg_cs_ctrl @(posedge clk);
lead: coverpoint lead_cyc iff (start_txn) {
bins one = {1};
bins two = {2};
bins small = {[3:8]};
bins large = {[9:$]};
}
lag: coverpoint lag_cyc iff (start_txn) {
bins one = {1};
bins small = {[2:8]};
bins large = {[9:$]};
}
gap: coverpoint gap_cyc iff (start_txn) {
bins one = {1};
bins small = {[2:8]};
bins large = {[9:$]};
}
// Frames per transaction. One frame never enters HOLD at all, so a suite of
// single-frame transfers has not tested multi-frame support.
frames: coverpoint frames_this_txn iff (txn_end) {
bins one = {1};
bins two = {2};
bins few = {[3:8]};
bins many = {[9:$]};
}
// Which slave, including the one that is not fitted.
slave: coverpoint sel iff (req) {
bins fitted = {[0:2]};
bins not_there = {3};
}
// Whether a request was PENDING when the gap expired. That is the case where
// the controller must hold it off and then honour it, and a suite that always
// waits for idle never produces it.
pending: coverpoint req iff (gap_expiring) {
bins waiting = {1};
bins idle = {0};
}
// Where an abort landed. Each of the three states that accept one behaves
// differently and each must be seen.
abort_state: coverpoint state iff (abort) {
bins in_lead = {1};
bins in_xfer = {2};
bins in_hold = {3};
}
x_frames_gap: cross frames, gap;
x_pending_gap: cross pending, gap;
endgroup9. Why an FPGA or ASIC Engineer Cares
Six states and one shared counter. The three intervals are never simultaneously active, so one counter with a three-way mux on its load input serves all of them — logic on the load path, which is evaluated once per state entry, instead of two extra counters.
The select decoder is the only place fan-out matters. cs_n is N_CS output pins, each driven from a flop, and the decode happens on the load of those flops rather than combinationally on their outputs. That keeps the pins glitch-free, which is the point: a select pin that glitches during a transaction is a deselection, and a deselection ends the transaction.
Reset must be asynchronous on the select flops specifically. Elsewhere in the master a synchronous reset would be acceptable. Here it is not: a synchronous reset needs a clock edge, and a master held in reset with no clock running would drive whatever the flops powered up in. cs_n must be high from the instant reset asserts.
shift_en is the block's only timing-relevant output, and it is a single-state decode — one term in a one-hot encoding. It gates the divider, so it is the one signal in this block worth looking at if the master fails timing.
Cost. CNT_W flops for the counter, three for the state in binary or six in one-hot, SEL_W for the latched select, N_CS for the select outputs, one for the start pulse and one for the latched hold. At the defaults with four slaves that is about 20 flops.
10. Failure Signature — A Flash That Returns 0xFF To Every Read
Symptom. A flash driver issues a READ: command byte 0x03, three address bytes, then reads four data bytes. Every data byte comes back 0xFF. The JEDEC ID command — a single command byte followed by three reads — works correctly and returns the right manufacturer and device IDs.
What that pattern rules out. The ID command working rules out wiring, mode, bit order and the shift datapath: a device that returns its correct ID is being clocked correctly and is decoding commands. So the fault is specific to the longer transaction.
What distinguishes the two commands. The ID command is short — one command byte and three reads. The READ is longer: eight frames. If the master's chip select is per-frame, both transactions are broken, but the ID command happens to work anyway on many parts because its command is a single byte and some devices latch it before the deselection. The READ cannot survive it, because the address is spread across three frames and the deselection between them discards it.
correct: CS ___________________________________/‾‾‾
cmd a2 a1 a0 d0 d1 d2 d3
per-frame: CS ___/‾\___/‾\___/‾\___/‾\___/‾\___/‾\___
cmd a2 a1 a0 ...
^ transaction ends here; address never completedWhy 0xFF specifically. A flash with no command in progress drives its MISO pin either high or not at all, and a pulled-up or floating line reads as all ones. So 0xFF is not data — it is the absence of data, and it is the single most informative value a flash read can return: it means the device is not answering, which is a transaction-framing problem rather than a data problem.
The diagnostic that identifies it in one step. Count chip-select falling edges per transaction. The requirements monitor of Chapter 13.1 reports this directly, and the expected value for an eight-frame READ is one. Eight means per-frame chip select; one means look elsewhere.
The fix, and why HOLD is where it goes. A hold bit per frame, and a state that pauses the clock without releasing the select. Not a longer frame — a 64-bit frame would work for this particular READ and not for a 260-byte page program, and it would also require the shift datapath to be 64 bits wide. The framing and the frame width are separate concerns and the fix belongs to the framing.
11. Common Misconceptions
"cs_n = ~busy is fine for a simple master." It has zero lead and lag, cannot hold across frames, and has no inter-transaction gap. The first fails above some frequency, the second makes flash impossible, and the third makes back-to-back transactions unreliable. There is no frequency or device for which all three are safe.
"HOLD and GAP could share one state with a flag for the select level." They could, and then the state's outputs depend on a flag, which is two states written as one. Naming them separately makes the machine's own diagram say which intervals are paid with the select asserted, and a reviewer does not have to read the counter logic to find out whether multi-frame support exists.
"The controller should count the frames of a transaction." A status poll reads until a bit clears; a streaming read continues until software stops. The length is not always known at the start, so the decision is made one frame at a time by whoever knows.
"An abort should release chip select immediately — that is what abort means." It means abandon the transfer, not violate the slave's timing. Releasing the select on the spot skips the hold time, and a slave sampling one last edge inside the violation may latch a bit that was never meant for it. An abort redirects the machine into LAG; it does not bypass it.
"A select encoding that names no slave will just do nothing." A truncating decoder asserts slave 0, which turns a driver bug into a write to the wrong device. It has to be refused explicitly, and the comparison has to be done as an integer — N_CS truncated to SEL_W bits is zero whenever N_CS is a power of two.
12. Reason It Through
Why do the interval counters load with the interval minus one?
Because the cycle the state is entered on is already part of the interval being timed. Loading the full value gives one cycle more than programmed — invisible on lead and lag, and a permanent throughput cost on gap, paid on every transaction. It is the kind of error that never fails a test and never gets found.
Why is sel latched on entry rather than read continuously?
Because a mid-transaction change would move the select from one device to another while a transfer was in flight, which deselects the first device — ending its transaction — and selects the second in the middle of a frame it has no context for. It is the same snapshot argument as Chapter 13.2, applied to the field where the consequence is worst.
A request is asserted throughout GAP. What does the machine do, and why is that the right answer?
Nothing, until the gap has been paid; then it honours it on reaching IDLE. Refusing the request outright would push an error path into software for a condition that is purely a matter of timing, and accepting it early would violate the slave's minimum deselect time. Holding it is the only option that is both correct and requires nothing from the driver.
Why is the lead asserted about shift_en rather than about SCLK?
Because this block does not drive SCLK. Holding it to another block's output couples the two, so a change in the divider's registered-output latency would fail this block's assertion — which is a false failure and points at the wrong place. shift_en is the interface this block actually controls, and the divider's own assertions cover the step from shift_en to the pin.
Why did the first attempt at aborting from outside this block hang the bus, and what does that generalise to?
Because the machine latches hold at frame request and does not consult the live input mid-frame, so pulling hold low during a transfer changed nothing and the machine went to HOLD — clock stopped, select still low — waiting for a request the aborting logic was refusing. The bus hung with the slave selected.
The generalisation is that a block which latches a control input cannot be steered by that input afterwards, and any attempt to redirect it from outside must go through an input it does consult. That is why the abort is an input here rather than a manipulation of the existing ones: one piece of logic knows how to leave the bus, and it is the one that owns the pins.
13. Understanding Check
14. Summary
cs_n = ~busy is wrong three times over: zero lead and lag, so the slave's setup and hold are both violated and the part works at low frequency and fails at high; chip select dropped between the frames of one transaction, which ends it and is the most common reason a flash read returns 0xFF; and zero inter-transaction gap, so the next transaction is ignored.
The machine has six states, and HOLD and GAP are the same pause with opposite pin behaviour — HOLD keeps the select asserted because the transaction continues, GAP releases it because it is over. That single difference is the whole of multi-frame support.
hold is a per-frame input, not a frame count, because a status poll or a streaming read does not know its length when it starts. The cost is that the last frame must be marked, which is why Chapter 13.11 gives it its own register address rather than a flag bit.
The abort input lives here, because this machine latches hold and cannot be steered from outside afterwards — the first attempt to abort externally left the bus hung with the slave selected, which was the exact failure it was meant to prevent. An abort redirects into LAG; it does not bypass the hold time.
A select naming a slave that is not fitted is refused and flagged, not truncated to slave zero — and the comparison must be done as an integer, because N_CS truncated to the select width is zero whenever N_CS is a power of two.
Counters load with the interval minus one, because the entry cycle is part of the interval. Reset releases every select asynchronously, because a master in reset with no clock running must not hold a slave selected.
For verification the interval checks are two-sided — too short violates the datasheet, too long silently costs throughput forever — and two measurement traps produced false failures: the first interval of each configuration belongs to the previous one, and waiting for idle before re-requesting measures the testbench. Coverage crosses burst length against the intervals, because a suite of single-frame transfers never enters HOLD at all.
15. What Comes Next
The master can now run a correct transaction of any number of frames to any fitted slave. Every frame is the same width and the same bit order.
Chapter 13.8 — Configurable Transfer Width and Bit Order adds both, and the interesting part is where the logic goes. The obvious implementation puts a mux in the shift path — the one piece of the design clocked at the bit rate — and there is an alternative that puts a single transform at the boundaries instead, where nothing is clocked at all. It turns out to be the same function on both sides.
Continue learning
Related tutorials
- Related topic
Launch and Sample Edges
One edge of each bit time places a bit on the wire, the other captures it, and they must never be the same edge. Why the separation is forced, why it buys half a period, and how RTL maps physical edges onto those roles.
- Related topic
CS-to-SCLK and SCLK-to-CS Timing
Chip select has timing requirements of its own: the lead before the first clock edge, the lag after the last, and the minimum deselect between transactions. Why violating them breaks a transfer whose every SCLK edge was correct.
- Related topic
Chip Select Semantics and Device Selection
What selection means on an SPI bus: why active-low is a convention, what each CS edge commits the device to, why selection is physical rather than addressed, and the generator that makes multi-select structurally impossible.
- Related topic
Bus Turnaround and the Contention Window
The overlap between one driver releasing and the next asserting: how long the gap really is once pad turn-off counts, why re-selecting the same device needs none, and the guard that enforces dead time only on a handover.
