SPI · Module 14
MOSI Capture and Bit Counting
Why the bit counter takes its zero from chip select and nothing else, why a one-bit slip produces well-formed words at the wrong offsets, the complete taxonomy of which configuration mismatches a slave can detect, and a receive path verified in three HDLs across 210 words.
The capture is one line of RTL:
if (cap_stb) sr <= {sr[MAX_W-2:0], mosi_q};That is the whole of "capturing MOSI". Everything difficult in this chapter is about the counter — and about a question that has no good answer:
The slave has assembled a well-formed 8-bit word. How does it know the word boundary is in the right place?
It does not. It cannot. And working out precisely why leads to the taxonomy that the rest of Module 14 is organised around.
1. The Counter Takes Its Zero From The Pin
Chapter 14.2 established that chip select is the only framing SPI has. The bit position within a transaction is therefore (edges since assert) / 2, and this block counts captures rather than edges — which is the same number, divided by two, for free.
The zero comes from txn_start_stb, and the word boundary is derived from it by counting to len. Nothing else resets the counter.
The alternative — a counter that rolls over at len with no reference to the select pin — is worth examining because it works, right up until it does not:
counter reset by the assert counter that just rolls over
--------------------------------- ------------------------------------
the pin defines word boundaries the counter defines word boundaries
a lost edge shifts the rest of a lost edge shifts the rest of the
the transaction transaction, identically
the NEXT transaction is correct the next transaction is also correct,
because the assert happens to reset
it anyway in most designs
the authority is external and the authority is internal and cannot
cannot be wrong about framing notice that it is wrongThe rows are nearly identical, which is why this decision gets made carelessly. The difference is the last one, and it is the whole difference: an internal authority that is wrong stays wrong confidently, producing words of exactly the right length at exactly the right times, all of them cut in the wrong place.
2. What A Slip Actually Looks Like
Suppose an edge is lost — the failure Chapter 14.1's precondition exists to prevent, and the one no simulation can produce. From that point on:
- Every subsequent word in the transaction is shifted by one bit.
- The words are still exactly
lenbits long. - They still arrive at exactly the right times, one per
lencaptures. - The bit counter's own arithmetic is perfectly self-consistent.
Nothing in the received data says anything is wrong. A status byte of 0x00 is still 0x00 shifted by one. A data byte is a different data byte. The only observation that reveals it is Chapter 14.2's edge count, and only at the end of the transaction — by which time the slave has already published every word.
3. The Taxonomy Worth Carrying Out Of This Chapter
Three configuration mismatches are possible between a master and a slave, they look identical from the system side — plausible data of the right length at the right time — and they need completely different responses.
POLARITY mismatch DETECTABLE FROM A LEVEL
SCLK is not at the configured idle level when chip
select asserts. Caught before a single bit has moved.
PHASE mismatch DETECTABLE FROM A MOTION, and NOT from the edge count
A mismatched pair still exchanges a whole number of
frames. What gives it away is that MOSI is changing at
the instant the slave samples it.
BIT-ORDER mismatch NOT DETECTABLE AT ALL
Both orders produce a well-formed word of the right
length at the right time. There is no observation the
slave can make that distinguishes them.The third entry is the one worth sitting with. It is not a gap in this design; it is a property of the protocol. A byte sent MSB-first and read LSB-first is a different byte, and both are legal bytes. There is no redundancy anywhere in SPI that could distinguish them — no parity, no delimiter, no reserved value.
So the honest engineering response is: bit order must be configured correctly, and the design must not pretend to check it. A flag that guessed — "this looks like a reversed ASCII string" — would be worse than no flag, because it would occasionally be right and would therefore be trusted.
The other two are Chapter 14.6's subject, and that chapter had to be rewritten once because the plausible detection mechanism for a phase mismatch does not work. The correction is instructive enough that it is left visible there rather than tidied away.
4. Bit Order Is The Master's Transform, Applied Once
Reversing the low len bits of a word is its own inverse. That was established on the master side in Chapter 13.8, and it means the slave applies exactly the same function the master applies — not an inverse function that has to be kept in step with it.
Two consequences:
The reversal happens at the word boundary, not in the shift path. The shift register always shifts one way. When a word completes, the low len bits are reversed if lsb_first is set. A design that reverses the shift direction instead has two shift paths to verify and two ways for len to interact with alignment; a design that reverses at the boundary has one shift path and one function, and the function is shared with the master.
There is no second implementation to keep in step. If the reversal is ever wrong it is wrong on both sides in the same way, which is a far better failure mode than a master and slave that disagree about what "LSB first" means for a 5-bit frame.
5. What The Block Publishes
rx_data the completed word, right-aligned in MAX_W bits
rx_valid_stb one cycle, WITH valid data on rx_data
bit_idx the position within the word in progress
words_in_txn how many complete words this transaction has produced
rx_partial_sr the word IN PROGRESS, right-aligned as far as it has gotrx_partial_sr needs justifying, because publishing an internal shift register looks like a layering violation.
Chapter 14.7 has to be able to report a partial word when a transaction ends mid-frame. It has two ways to get one: this block publishes its in-progress shift register, or 14.7 keeps its own copy of the same state. The second means two copies of the same shift register, updated by the same strobe, which is one more copy than can be kept correct across a width change or a reset. Publishing the state is the smaller evil, and it is the reason the partial path in 14.7 has no shift logic at all.
rx_valid_stb has a one-cycle subtlety that is worth stating because getting it wrong produces a symptom that points somewhere else entirely — see §7.
6. Building the Receive Path — Three HDLs
The circuit
A shift register, a capture counter that takes its zero from the transaction start, a comparator against len, and a reversal at the word boundary. The only structural decision visible in the code is that rx_valid_stb is registered rather than combinational from the comparator.
// spi_slave_rx.sv
//
// Chapter 14.3 -- capturing MOSI, and counting bits that cannot be recounted.
//
// The capture itself is one line. Everything interesting is about the COUNTER,
// and about what the slave can and cannot detect when the counter is wrong.
//
// THE COUNTER IS CLEARED BY THE TRANSACTION, AND BY THE WORD BOUNDARY IT
// DERIVES FROM IT -- AND BY NOTHING ELSE.
//
// Chapter 14.2 established that chip select is the only framing SPI has, so the
// bit position within a transaction is (edges since assert) / 2. This block
// counts captures rather than edges, which is the same number divided by two,
// and it takes its zero from `txn_start_stb`. A counter that rolled over on
// `len` without any reference to the select pin would be asserting that its own
// arithmetic defines the word boundary -- and if it were ever wrong it would
// stay wrong for the rest of the assertion, because there is nothing in the data
// stream to resynchronise against.
//
// WHAT THE SLAVE CAN DETECT, AND WHAT IT CANNOT. This is the taxonomy worth
// carrying out of the chapter, because the three mismatches look identical from
// the system side and need completely different responses:
//
// POLARITY mismatch DETECTABLE FROM A LEVEL. SCLK is not at the configured idle
// level when chip select asserts (Chapter 14.6).
// PHASE mismatch DETECTABLE FROM A MOTION, and NOT from the edge count: a
// mismatched pair still exchanges a whole number of frames.
// What gives it away is that MOSI is changing at the instant
// the slave samples it (Chapter 14.6).
// BIT-ORDER mismatch NOT DETECTABLE AT ALL. Both orders produce a well-formed
// word of the right length at the right time. There is no
// observation the slave can make that distinguishes them, so
// it must be configured correctly and there is nothing to
// report.
//
// A SLIP IS ALSO NOT DETECTABLE FROM THE DATA. If an edge is lost -- the failure
// Chapter 14.1's precondition exists to prevent -- every subsequent word in the
// transaction is shifted by one bit. The words are still the right length and
// still arrive at the right times, because their boundaries come from a counter
// that is itself wrong. Only the edge count of Chapter 14.2 reveals it, and only
// at the END of the transaction. That asymmetry is why the edge count is
// published at all.
//
// BIT ORDER IS THE SAME TRANSFORM THE MASTER USES. Reversing the low `len` bits
// is its own inverse (Chapter 13.8), so the slave applies exactly the function
// the master applies, on the boundary rather than in the shift path. There is no
// second implementation to keep in step.
module spi_slave_rx #(
parameter int MAX_W = 32,
parameter int LEN_W = 6,
parameter int CNT_W = 12
) (
input wire clk,
input wire rst_n,
// --- from 14.2 ---------------------------------------------------------
input wire txn_active,
input wire txn_start_stb,
// --- the capture strobe, from the mode logic of 14.6 -------------------
input wire cap_stb,
input wire mosi_q, // delay-matched, from 14.1
// --- configuration -----------------------------------------------------
input wire [LEN_W-1:0] len,
input wire lsb_first,
// --- the received words ------------------------------------------------
output wire [MAX_W-1:0] rx_data,
output reg rx_valid_stb, // one cycle, WITH valid data
output wire [LEN_W-1:0] bit_idx, // captures into the current word
output wire [CNT_W-1:0] words_in_txn,
// The word IN PROGRESS, right-aligned as far as it has got. Published because
// Chapter 14.7 has to be able to report a partial, and the only alternative is
// for 14.7 to duplicate the shift register -- two copies of the same state,
// which is one more than can be kept correct.
output wire [MAX_W-1:0] rx_partial_sr
);
reg [MAX_W-1:0] rx_sr;
reg [MAX_W-1:0] rx_hold;
reg [LEN_W-1:0] bits;
reg [CNT_W-1:0] words;
assign bit_idx = bits;
assign words_in_txn = words;
assign rx_partial_sr = rx_sr;
// A fixed full-width reversal: pure wiring, no logic.
function automatic [MAX_W-1:0] rev_all(input [MAX_W-1:0] x);
integer b;
begin
rev_all = {MAX_W{1'b0}};
for (b = 0; b < MAX_W; b = b + 1)
rev_all[b] = x[MAX_W-1-b];
end
endfunction
// Reverse the low `len` bits, zero above -- its own inverse, and the same
// function the master applies (Chapter 13.8).
wire [MAX_W-1:0] rx_rev = rev_all(rx_hold) >> (MAX_W - len);
assign rx_data = lsb_first ? rx_rev : rx_hold;
// The word completes on the capture that brings the count to len-1, so this
// is the LAST capture of the word rather than the cycle after it.
wire last_cap = cap_stb & txn_active & (bits == (len - 1'b1));
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
rx_sr <= {MAX_W{1'b0}};
rx_hold <= {MAX_W{1'b0}};
bits <= {LEN_W{1'b0}};
words <= {CNT_W{1'b0}};
rx_valid_stb <= 1'b0;
end else begin
// Delayed by one cycle, deliberately: the final bit reaches `rx_sr`
// on the same edge that sees the last capture, so on that cycle the
// word is one bit short. Pulsing valid there means "good next
// cycle", which every consumer gets wrong once (Chapter 13.6).
rx_valid_stb <= last_cap;
if (txn_start_stb) begin
// The transaction's zero. Chapter 14.2's rule, applied here.
rx_sr <= {MAX_W{1'b0}};
bits <= {LEN_W{1'b0}};
words <= {CNT_W{1'b0}};
end else if (cap_stb && txn_active) begin
// Bits arrive at the bottom and walk up, so after `len` captures
// the word is already right-aligned with zeros above -- no
// alignment shift on this side, the same asymmetry as the
// master's datapath.
rx_sr <= {rx_sr[MAX_W-2:0], mosi_q};
if (bits == (len - 1'b1)) begin
// The completed word is copied out and the shift register is
// CLEARED, not merely overwritten. A frame narrower than the
// datapath does not overwrite the bits above it, and a
// transaction of several narrow words would otherwise carry
// the previous word's bits in the positions above the
// current one.
rx_hold <= {rx_sr[MAX_W-2:0], mosi_q};
rx_sr <= {MAX_W{1'b0}};
bits <= {LEN_W{1'b0}};
words <= words + 1'b1;
end else begin
bits <= bits + 1'b1;
end
end
end
end
`ifdef SPI_CHECKS
always_ff @(posedge clk) if (rst_n) begin
if (cap_stb && !txn_active && txn_active !== 1'bx)
$fatal(1, "a capture strobe arrived with no transaction open");
if (bits >= len && len != {LEN_W{1'b0}})
$fatal(1, "the bit index reached the frame width");
end
`endif
endmodule// spi_slave_rx.v
//
// Chapter 14.3 -- capturing MOSI, and counting bits that cannot be recounted.
//
// The capture itself is one line. Everything interesting is about the COUNTER,
// and about what the slave can and cannot detect when the counter is wrong.
//
// THE COUNTER IS CLEARED BY THE TRANSACTION, AND BY THE WORD BOUNDARY IT
// DERIVES FROM IT -- AND BY NOTHING ELSE.
//
// Chapter 14.2 established that chip select is the only framing SPI has, so the
// bit position within a transaction is (edges since assert) / 2. This block
// counts captures rather than edges, which is the same number divided by two,
// and it takes its zero from `txn_start_stb`. A counter that rolled over on
// `len` without any reference to the select pin would be asserting that its own
// arithmetic defines the word boundary -- and if it were ever wrong it would
// stay wrong for the rest of the assertion, because there is nothing in the data
// stream to resynchronise against.
//
// WHAT THE SLAVE CAN DETECT, AND WHAT IT CANNOT. This is the taxonomy worth
// carrying out of the chapter, because the three mismatches look identical from
// the system side and need completely different responses:
//
// POLARITY mismatch DETECTABLE FROM A LEVEL. SCLK is not at the configured idle
// level when chip select asserts (Chapter 14.6).
// PHASE mismatch DETECTABLE FROM A MOTION, and NOT from the edge count: a
// mismatched pair still exchanges a whole number of frames.
// What gives it away is that MOSI is changing at the instant
// the slave samples it (Chapter 14.6).
// BIT-ORDER mismatch NOT DETECTABLE AT ALL. Both orders produce a well-formed
// word of the right length at the right time. There is no
// observation the slave can make that distinguishes them, so
// it must be configured correctly and there is nothing to
// report.
//
// A SLIP IS ALSO NOT DETECTABLE FROM THE DATA. If an edge is lost -- the failure
// Chapter 14.1's precondition exists to prevent -- every subsequent word in the
// transaction is shifted by one bit. The words are still the right length and
// still arrive at the right times, because their boundaries come from a counter
// that is itself wrong. Only the edge count of Chapter 14.2 reveals it, and only
// at the END of the transaction. That asymmetry is why the edge count is
// published at all.
//
// BIT ORDER IS THE SAME TRANSFORM THE MASTER USES. Reversing the low `len` bits
// is its own inverse (Chapter 13.8), so the slave applies exactly the function
// the master applies, on the boundary rather than in the shift path. There is no
// second implementation to keep in step.
module spi_slave_rx #(
parameter MAX_W = 32,
parameter LEN_W = 6,
parameter CNT_W = 12
) (
input wire clk,
input wire rst_n,
// --- from 14.2 ---------------------------------------------------------
input wire txn_active,
input wire txn_start_stb,
// --- the capture strobe, from the mode logic of 14.6 -------------------
input wire cap_stb,
input wire mosi_q, // delay-matched, from 14.1
// --- configuration -----------------------------------------------------
input wire [LEN_W-1:0] len,
input wire lsb_first,
// --- the received words ------------------------------------------------
output wire [MAX_W-1:0] rx_data,
output reg rx_valid_stb, // one cycle, WITH valid data
output wire [LEN_W-1:0] bit_idx, // captures into the current word
output wire [CNT_W-1:0] words_in_txn,
// The word IN PROGRESS, right-aligned as far as it has got. Published because
// Chapter 14.7 has to be able to report a partial, and the only alternative is
// for 14.7 to duplicate the shift register -- two copies of the same state,
// which is one more than can be kept correct.
output wire [MAX_W-1:0] rx_partial_sr
);
reg [MAX_W-1:0] rx_sr;
reg [MAX_W-1:0] rx_hold;
reg [LEN_W-1:0] bits;
reg [CNT_W-1:0] words;
assign bit_idx = bits;
assign words_in_txn = words;
assign rx_partial_sr = rx_sr;
// A fixed full-width reversal: pure wiring, no logic.
function [MAX_W-1:0] rev_all;
input [MAX_W-1:0] x;
integer b;
begin
rev_all = {MAX_W{1'b0}};
for (b = 0; b < MAX_W; b = b + 1)
rev_all[b] = x[MAX_W-1-b];
end
endfunction
// Reverse the low `len` bits, zero above -- its own inverse, and the same
// function the master applies (Chapter 13.8).
wire [MAX_W-1:0] rx_rev = rev_all(rx_hold) >> (MAX_W - len);
assign rx_data = lsb_first ? rx_rev : rx_hold;
// The word completes on the capture that brings the count to len-1, so this
// is the LAST capture of the word rather than the cycle after it.
wire last_cap = cap_stb & txn_active & (bits == (len - 1'b1));
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
rx_sr <= {MAX_W{1'b0}};
rx_hold <= {MAX_W{1'b0}};
bits <= {LEN_W{1'b0}};
words <= {CNT_W{1'b0}};
rx_valid_stb <= 1'b0;
end else begin
// Delayed by one cycle, deliberately: the final bit reaches `rx_sr`
// on the same edge that sees the last capture, so on that cycle the
// word is one bit short. Pulsing valid there means "good next
// cycle", which every consumer gets wrong once (Chapter 13.6).
rx_valid_stb <= last_cap;
if (txn_start_stb) begin
// The transaction's zero. Chapter 14.2's rule, applied here.
rx_sr <= {MAX_W{1'b0}};
bits <= {LEN_W{1'b0}};
words <= {CNT_W{1'b0}};
end else if (cap_stb && txn_active) begin
// Bits arrive at the bottom and walk up, so after `len` captures
// the word is already right-aligned with zeros above -- no
// alignment shift on this side, the same asymmetry as the
// master's datapath.
rx_sr <= {rx_sr[MAX_W-2:0], mosi_q};
if (bits == (len - 1'b1)) begin
// The completed word is copied out and the shift register is
// CLEARED, not merely overwritten. A frame narrower than the
// datapath does not overwrite the bits above it, and a
// transaction of several narrow words would otherwise carry
// the previous word's bits in the positions above the
// current one.
rx_hold <= {rx_sr[MAX_W-2:0], mosi_q};
rx_sr <= {MAX_W{1'b0}};
bits <= {LEN_W{1'b0}};
words <= words + 1'b1;
end else begin
bits <= bits + 1'b1;
end
end
end
end
`ifdef SPI_CHECKS
always @(posedge clk) if (rst_n) begin
if (cap_stb && !txn_active && txn_active !== 1'bx)
$fatal(1, "a capture strobe arrived with no transaction open");
if (bits >= len && len != {LEN_W{1'b0}})
$fatal(1, "the bit index reached the frame width");
end
`endif
endmodule-- spi_slave_rx.vhd
--
-- Chapter 14.3 -- capturing MOSI, and counting bits that cannot be recounted.
--
-- The capture itself is one line. Everything interesting is about the COUNTER,
-- and about what the slave can and cannot detect when the counter is wrong.
--
-- THE COUNTER IS CLEARED BY THE TRANSACTION, AND BY THE WORD BOUNDARY IT DERIVES
-- FROM IT -- AND BY NOTHING ELSE.
--
-- Chapter 14.2 established that chip select is the only framing SPI has, so the
-- bit position within a transaction is (edges since assert) / 2. This block counts
-- captures, which is the same number divided by two, and takes its zero from
-- `txn_start_stb`. A counter that rolled over on `len` without any reference to
-- the select pin would be asserting that its own arithmetic defines the word
-- boundary -- and if it were ever wrong it would stay wrong for the rest of the
-- assertion, because there is nothing in the data stream to resynchronise against.
--
-- WHAT THE SLAVE CAN DETECT, AND WHAT IT CANNOT. The three mismatches look
-- identical from the system side and need completely different responses:
--
-- POLARITY mismatch DETECTABLE FROM A LEVEL. SCLK is not at the configured idle
-- level when chip select asserts (Chapter 14.6).
-- PHASE mismatch DETECTABLE FROM A MOTION, and NOT from the edge count: a
-- mismatched pair still exchanges a whole number of frames.
-- What gives it away is that MOSI is changing at the instant
-- the slave samples it (Chapter 14.6).
-- BIT-ORDER mismatch NOT DETECTABLE AT ALL. Both orders produce a well-formed
-- word of the right length at the right time. There is no
-- observation the slave can make that distinguishes them, so
-- it must be configured correctly and there is nothing to
-- report.
--
-- A SLIP IS ALSO NOT DETECTABLE FROM THE DATA. If an edge is lost -- the failure
-- Chapter 14.1's precondition exists to prevent -- every subsequent word in the
-- transaction is shifted by one bit. The words are still the right length and
-- still arrive at the right times, because their boundaries come from a counter
-- that is itself wrong. Only the edge count of Chapter 14.2 reveals it, and only
-- at the END of the transaction. That asymmetry is why the edge count is published.
--
-- BIT ORDER IS THE SAME TRANSFORM THE MASTER USES. Reversing the low `len` bits is
-- its own inverse (Chapter 13.8), so the slave applies exactly the function the
-- master applies, on the boundary rather than in the shift path.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_slave_rx is
generic (
MAX_W : positive := 32;
LEN_W : positive := 6;
CNT_W : positive := 12
);
port (
clk : in std_logic;
rst_n : in std_logic;
-- from 14.2
txn_active : in std_logic;
txn_start_stb : in std_logic;
-- the capture strobe, from the mode logic of 14.6
cap_stb : in std_logic;
mosi_q : in std_logic; -- delay-matched, from 14.1
-- configuration
len : in unsigned(LEN_W - 1 downto 0);
lsb_first : in std_logic;
-- the received words
rx_data : out std_logic_vector(MAX_W - 1 downto 0);
rx_valid_stb : out std_logic; -- one cycle, WITH valid data
bit_idx : out unsigned(LEN_W - 1 downto 0);
words_in_txn : out unsigned(CNT_W - 1 downto 0);
-- The word IN PROGRESS, right-aligned as far as it has got. Published
-- because Chapter 14.7 has to be able to report a partial, and the only
-- alternative is for 14.7 to duplicate the shift register -- two copies of
-- the same state, which is one more than can be kept correct.
rx_partial_sr : out std_logic_vector(MAX_W - 1 downto 0)
);
end entity;
architecture rtl of spi_slave_rx is
signal rx_sr : std_logic_vector(MAX_W - 1 downto 0) := (others => '0');
signal rx_hold : std_logic_vector(MAX_W - 1 downto 0) := (others => '0');
signal bits : unsigned(LEN_W - 1 downto 0) := (others => '0');
signal words : unsigned(CNT_W - 1 downto 0) := (others => '0');
signal valid_r : std_logic := '0';
signal last_cap : std_logic;
-- A fixed full-width reversal: pure wiring, no logic.
function rev_all(x : std_logic_vector(MAX_W - 1 downto 0))
return std_logic_vector is
variable r : std_logic_vector(MAX_W - 1 downto 0);
begin
for b in 0 to MAX_W - 1 loop
r(b) := x(MAX_W - 1 - b);
end loop;
return r;
end function;
signal n_bits : natural;
signal rx_rev : std_logic_vector(MAX_W - 1 downto 0);
begin
bit_idx <= bits;
words_in_txn <= words;
rx_partial_sr <= rx_sr;
rx_valid_stb <= valid_r;
-- Clamped so the shift amount stays in range while an illegal width is being
-- presented: an out-of-range shift is an undefined result, not an error, and
-- an undefined value propagating into a data path turns a configuration
-- mistake into an intermittent corruption.
n_bits <= 1 when len = 0
else MAX_W when to_integer(len) > MAX_W
else to_integer(len);
-- Reverse the low `len` bits, zero above -- its own inverse, and the same
-- function the master applies (Chapter 13.8).
rx_rev <= std_logic_vector(shift_right(unsigned(rev_all(rx_hold)),
MAX_W - n_bits));
rx_data <= rx_rev when lsb_first = '1' else rx_hold;
-- The word completes on the capture that brings the count to len-1, so this
-- is the LAST capture of the word rather than the cycle after it.
last_cap <= '1' when (cap_stb = '1' and txn_active = '1' and
bits = (len - 1)) else '0';
capture : process (clk, rst_n)
begin
if rst_n = '0' then
rx_sr <= (others => '0');
rx_hold <= (others => '0');
bits <= (others => '0');
words <= (others => '0');
valid_r <= '0';
elsif rising_edge(clk) then
-- Delayed by one cycle, deliberately: the final bit reaches `rx_sr` on
-- the same edge that sees the last capture, so on that cycle the word
-- is one bit short. Pulsing valid there means "good next cycle", which
-- every consumer gets wrong once (Chapter 13.6).
valid_r <= last_cap;
if txn_start_stb = '1' then
-- The transaction's zero. Chapter 14.2's rule, applied here.
rx_sr <= (others => '0');
bits <= (others => '0');
words <= (others => '0');
elsif cap_stb = '1' and txn_active = '1' then
-- Bits arrive at the bottom and walk up, so after `len` captures
-- the word is already right-aligned with zeros above -- no
-- alignment shift on this side, the same asymmetry as the
-- master's datapath.
if bits = (len - 1) then
-- The completed word is copied out and the shift register is
-- CLEARED, not merely overwritten. A frame narrower than the
-- datapath does not overwrite the bits above it, and a
-- transaction of several narrow words would otherwise carry
-- the previous word's bits in the positions above this one.
rx_hold <= rx_sr(MAX_W - 2 downto 0) & mosi_q;
rx_sr <= (others => '0');
bits <= (others => '0');
words <= words + 1;
else
rx_sr <= rx_sr(MAX_W - 2 downto 0) & mosi_q;
bits <= bits + 1;
end if;
end if;
end if;
end process;
check : process (clk)
begin
if rising_edge(clk) and rst_n = '1' then
assert not (cap_stb = '1' and txn_active = '0')
report "a capture strobe arrived with no transaction open"
severity failure;
assert not (bits >= len and len /= 0)
report "the bit index reached the frame width" severity failure;
end if;
end process;
end architecture;The testbench
Seven tests, and the fifth is the one that would catch a real bug.
- One byte, MSB-first, mode 0. The simplest thing that can work.
- Several words in one transaction. The word boundary comes from the counter, so multiple words in a single assertion exercise it repeatedly rather than once.
- Narrow words in sequence. Five-bit words back to back, which is where an off-by-one in the boundary comparison shows up as words drifting rather than as a single wrong word.
- Both bit orders. The wire carries a different sequence in each, so this is a test of the transform rather than of the plumbing.
- The sweep. Widths from 1 to 32, both orders, both phases, several words each — 210 words in total. The interesting widths are 1 (where the word boundary is every capture), and 32 (where
MAX_Wandlencoincide and the reversal touches every bit). - A transaction that ends mid-word. The completed words must be delivered and the partial must not be presented as a word.
- The continuous properties.
rx_valid_stbnever fires outside a transaction;bit_idxis always less thanlen; the counter is zero on the cycle after every start.
// spi_slave_rx_tb.sv
//
// The real front end of 14.1 and the real transaction detector of 14.2 sit in
// front of the block under test, so every capture is a recovered one and the word
// boundaries come from the select pin exactly as they will in the finished slave.
//
// Two properties get most of the attention because both are invisible to the
// obvious test:
//
// NARROW WORDS IN SEQUENCE. A five-bit word followed by another five-bit word
// must not carry the first word's bits in the positions above the second's. A
// suite whose words are all as wide as the datapath cannot fail this, and a
// suite whose narrow words are all identical cannot either.
//
// THE WORD COUNT MUST AGREE WITH THE EDGE COUNT. Chapter 14.2 counts edges and
// this block counts captures; they are the same number divided by two. If they
// ever disagree, one of them has invented a boundary -- which is the failure
// 14.2's rule exists to prevent, and the only place it can be observed is here,
// where both numbers exist at once.
`timescale 1ns/1ps
module spi_slave_rx_tb;
localparam int MAX_W = 32;
localparam int LEN_W = 6;
localparam int CNT_W = 12;
logic clk = 1'b0;
logic rst_n = 1'b0;
always #5 clk = ~clk;
// --- 14.1 ---------------------------------------------------------------
logic cpol = 1'b0;
logic sclk_pin = 1'b0;
logic cs_n_pin = 1'b1;
logic mosi_pin = 1'b0;
wire sclk_q, cs_active, mosi_q;
wire edge_a_stb, edge_b_stb;
wire cs_assert_stb, cs_deassert_stb;
wire [11:0] min_half;
wire ratio_err;
spi_slave_frontend #(.SYNC_N(2), .HALF_MIN(3), .CNT_W(12)) u_fe (
.clk(clk), .rst_n(rst_n), .cpol(cpol),
.sclk_pin(sclk_pin), .cs_n_pin(cs_n_pin), .mosi_pin(mosi_pin),
.sclk_q(sclk_q), .cs_active(cs_active), .mosi_q(mosi_q),
.edge_a_stb(edge_a_stb), .edge_b_stb(edge_b_stb),
.cs_assert_stb(cs_assert_stb), .cs_deassert_stb(cs_deassert_stb),
.min_half(min_half), .ratio_err(ratio_err), .clr_flags(1'b0)
);
// --- 14.2 ---------------------------------------------------------------
logic [LEN_W-1:0] len = 6'd8;
logic lsb_first = 1'b0;
logic cpha_now = 1'b0;
wire txn_active, txn_start_stb, txn_end_stb;
wire [CNT_W-1:0] edges_in_txn, frames_in_txn;
wire txn_clean, txn_trunc, txn_empty, txn_report_stb;
wire [2:0] cs_state;
spi_slave_cs #(.LEN_W(LEN_W), .CNT_W(CNT_W)) u_cs (
.clk(clk), .rst_n(rst_n),
.cs_assert_stb(cs_assert_stb), .cs_deassert_stb(cs_deassert_stb),
.edge_a_stb(edge_a_stb), .edge_b_stb(edge_b_stb), .len(len),
.txn_active(txn_active), .txn_start_stb(txn_start_stb),
.txn_end_stb(txn_end_stb),
.edges_in_txn(edges_in_txn), .frames_in_txn(frames_in_txn),
.txn_clean(txn_clean), .txn_trunc(txn_trunc), .txn_empty(txn_empty),
.txn_report_stb(txn_report_stb), .state_id(cs_state)
);
// The mode logic is Chapter 14.6; here the phase simply selects which
// recovered edge is the capture.
wire cap_stb = txn_active & (cpha_now ? edge_b_stb : edge_a_stb);
// --- 14.3, the block under test ----------------------------------------
wire [MAX_W-1:0] rx_data;
wire rx_valid_stb;
wire [LEN_W-1:0] bit_idx;
wire [CNT_W-1:0] words_in_txn;
wire [MAX_W-1:0] rx_partial_sr;
spi_slave_rx #(.MAX_W(MAX_W), .LEN_W(LEN_W), .CNT_W(CNT_W)) dut (
.clk(clk), .rst_n(rst_n),
.txn_active(txn_active), .txn_start_stb(txn_start_stb),
.cap_stb(cap_stb), .mosi_q(mosi_q),
.len(len), .lsb_first(lsb_first),
.rx_data(rx_data), .rx_valid_stb(rx_valid_stb),
.bit_idx(bit_idx), .words_in_txn(words_in_txn),
.rx_partial_sr(rx_partial_sr)
);
// --- the expectation ----------------------------------------------------
integer seed;
logic [MAX_W-1:0] exp_q [0:63];
integer exp_w, got_w;
integer word_bad, upper_bad, valid_at_wrong_time;
integer n_valid;
always_ff @(posedge clk) begin
if (rst_n) begin
if (txn_start_stb) begin
got_w <= 0;
end
if (rx_valid_stb) begin
n_valid <= n_valid + 1;
if (got_w < exp_w) begin
if (rx_data !== exp_q[got_w]) word_bad <= word_bad + 1;
// Nothing above the frame width, ever. This is the property
// a narrow word following another narrow word breaks.
if (len < MAX_W && (rx_data >> len) != {MAX_W{1'b0}})
upper_bad <= upper_bad + 1;
end else begin
valid_at_wrong_time <= valid_at_wrong_time + 1;
end
got_w <= got_w + 1;
end
end
end
integer errors = 0;
task automatic adv(input integer n);
begin repeat (n) @(negedge clk); end
endtask
task automatic clear_exp;
begin
exp_w = 0; got_w = 0;
end
endtask
// Drives a whole transaction of `nwords` words of `nbits` bits, in the given
// mode, with the bit order the SLAVE is configured for -- so what goes on the
// wire is the order that configuration implies, and the received word must
// come back equal to the value sent.
task automatic drive_txn(input integer half, input bit pol, input bit pha,
input integer nbits, input integer nwords,
input bit lsb);
integer w, i, seedl;
logic [MAX_W-1:0] word;
logic b;
begin
cpol = pol;
cpha_now = pha;
len = nbits[LEN_W-1:0];
lsb_first = lsb;
sclk_pin = pol;
cs_n_pin = 1'b1;
adv(8);
clear_exp();
cs_n_pin = 1'b0;
adv(4);
for (w = 0; w < nwords; w = w + 1) begin
seed = (seed * 32'h0019_660D) + 32'h3C6E_F35F;
word = seed & ((nbits >= 32) ? 32'hFFFF_FFFF
: ((32'h1 << nbits) - 1));
exp_q[exp_w] = word;
exp_w = exp_w + 1;
for (i = 0; i < nbits; i = i + 1) begin
// The wire carries the configured order: bit i of the word
// first when LSB-first, bit nbits-1-i first otherwise.
b = lsb ? word[i] : word[nbits-1-i];
if (!pha) begin
// Capture is on the leading edge, so the bit must be on
// the wire before it -- driven at CS assert for the first
// bit and on the trailing edge for the rest.
mosi_pin = b;
adv(2);
sclk_pin = ~pol; // leading: captured here
adv(half);
sclk_pin = pol; // trailing
adv(half - 2 < 1 ? 1 : half - 2);
end else begin
sclk_pin = ~pol; // leading: launch
mosi_pin = b;
adv(half);
sclk_pin = pol; // trailing: captured here
adv(half);
end
end
end
adv(4);
cs_n_pin = 1'b1;
adv(8);
sclk_pin = pol;
adv(8);
end
endtask
integer h, p, o, nb, nw, k;
initial begin
exp_w = 0; got_w = 0; n_valid = 0;
word_bad = 0; upper_bad = 0; valid_at_wrong_time = 0;
seed = 32'h0C4F_7E21;
adv(3);
rst_n = 1'b1;
adv(2);
// 1. ONE BYTE, MSB-first, mode 0. The simplest thing that can work.
drive_txn(4, 1'b0, 1'b0, 8, 1, 1'b0);
if (got_w != 1) begin
$display(" FAIL: one byte produced %0d received words", got_w);
errors = errors + 1;
end
$display(" one byte, mode 0, MSB-first: received %02h and expected %02h",
rx_data[7:0], exp_q[0][7:0]);
// 2. SEVERAL WORDS IN ONE TRANSACTION, in order. The word boundary comes
// from a counter whose zero is the select pin, so a transaction of
// four words must produce four words and the fourth must not be the
// first one's bits shifted.
drive_txn(4, 1'b0, 1'b0, 8, 4, 1'b0);
if (got_w != 4) begin
$display(" FAIL: four bytes produced %0d received words", got_w);
errors = errors + 1;
end
if (words_in_txn != 4 || frames_in_txn != 4) begin
$display(" FAIL: the capture counter says %0d words and the edge counter says %0d frames",
words_in_txn, frames_in_txn);
errors = errors + 1;
end
$display(" four bytes in one transaction: four words received in order, and the capture and edge counters agree");
// 3. NARROW WORDS IN SEQUENCE. Five-bit words, one after another, with
// different values -- the case where a receive register that is not
// cleared at the word boundary carries the previous word's bits above
// the current one.
drive_txn(4, 1'b0, 1'b0, 5, 6, 1'b0);
if (got_w != 6) begin
$display(" FAIL: six five-bit words produced %0d", got_w);
errors = errors + 1;
end
$display(" six five-bit words in one transaction: each right-aligned with nothing above bit 4");
// 4. BOTH BIT ORDERS. The wire carries a different sequence in each, so
// a slave that ignored the setting would return the bit-reverse.
drive_txn(4, 1'b0, 1'b0, 8, 3, 1'b0);
drive_txn(4, 1'b0, 1'b0, 8, 3, 1'b1);
$display(" both bit orders returned the value that was sent, which a slave ignoring the setting could not do");
// 5. THE SWEEP. Widths from 1 to 32, both orders, both phases, several
// ratios, multi-word transactions.
for (o = 0; o <= 1; o = o + 1)
for (p = 0; p <= 1; p = p + 1)
for (h = 2; h <= 6; h = h + 2) begin
drive_txn(h, 1'b0, p[0], 1, 3, o[0]);
drive_txn(h, 1'b0, p[0], 2, 3, o[0]);
drive_txn(h, 1'b0, p[0], 5, 3, o[0]);
drive_txn(h, 1'b0, p[0], 8, 3, o[0]);
drive_txn(h, 1'b0, p[0], 13, 2, o[0]);
drive_txn(h, 1'b0, p[0], 32, 2, o[0]);
end
$display(" %0d words swept: widths 1, 2, 5, 8, 13 and 32, both bit orders, both phases, three ratios",
n_valid);
// 6. A TRANSACTION THAT ENDS MID-WORD. The completed words must be
// unaffected and the partial count must be readable -- the policy for
// what to DO with it is Chapter 14.7's.
len = 6'd8; lsb_first = 1'b0; cpha_now = 1'b0;
adv(2);
clear_exp();
// The bits driven below are k[0] for k = 0..10, so the first eight --
// 0 1 0 1 0 1 0 1 read MSB-first -- are 0x55. Registering it matters:
// an unexpected word and a missing word are different failures, and a
// monitor that cannot tell them apart reports the wrong one.
exp_q[0] = 32'h55;
exp_w = 1;
cpol = 1'b0; sclk_pin = 1'b0; cs_n_pin = 1'b1; adv(8);
cs_n_pin = 1'b0; adv(4);
// one complete byte, then three bits of a second
for (k = 0; k < 11; k = k + 1) begin
mosi_pin = k[0];
adv(2);
sclk_pin = 1'b1; adv(4);
sclk_pin = 1'b0; adv(2);
end
adv(4);
if (bit_idx != 3) begin
$display(" FAIL: after 11 bits the partial count was %0d, expected 3",
bit_idx);
errors = errors + 1;
end
if (words_in_txn != 1) begin
$display(" FAIL: after 11 bits %0d words completed, expected 1",
words_in_txn);
errors = errors + 1;
end
cs_n_pin = 1'b1; adv(8); sclk_pin = 1'b0; adv(8);
if (!txn_trunc) begin
$display(" FAIL: a transaction ending mid-word was not reported truncated");
errors = errors + 1;
end
$display(" a transaction of 11 bits at eight bits per word: one complete word, a partial count of 3, and the transaction reported truncated");
// 7. THE CONTINUOUS PROPERTIES.
if (word_bad != 0) begin
$display(" FAIL: %0d of %0d received words were wrong",
word_bad, n_valid);
errors = errors + 1;
end
if (upper_bad != 0) begin
$display(" FAIL: %0d words carried bits above the frame width",
upper_bad);
errors = errors + 1;
end
if (valid_at_wrong_time != 0) begin
$display(" FAIL: %0d words arrived that nothing was sent for",
valid_at_wrong_time);
errors = errors + 1;
end
$display(" %0d received words, none wrong, none carrying bits above its width, and none unaccounted for",
n_valid);
if (errors == 0)
$display("PASS: MOSI is captured on the mode's capture edge and assembled into words whose boundaries come from a counter cleared by the select pin, so nothing the slave counts can redefine where a word ends -- all %0d words across widths of 1, 2, 5, 8, 13 and 32 bits, both bit orders, both phases and three frequency ratios came back exactly as sent, every one right-aligned with nothing above its width even when narrow words of differing values followed one another, the capture counter and the edge counter of Chapter 14.2 agree on every transaction, and a transaction ending mid-word leaves its completed words intact with a readable partial count and is reported as truncated", n_valid);
else
$display("FAIL: %0d error(s)", errors);
$finish;
end
endmodule// spi_slave_rx_tb.v
//
// The real front end of 14.1 and the real transaction detector of 14.2 sit in
// front of the block under test, so every capture is a recovered one and the word
// boundaries come from the select pin exactly as they will in the finished slave.
//
// Two properties get most of the attention because both are invisible to the
// obvious test:
//
// NARROW WORDS IN SEQUENCE. A five-bit word followed by another five-bit word
// must not carry the first word's bits in the positions above the second's. A
// suite whose words are all as wide as the datapath cannot fail this, and a
// suite whose narrow words are all identical cannot either.
//
// THE WORD COUNT MUST AGREE WITH THE EDGE COUNT. Chapter 14.2 counts edges and
// this block counts captures; they are the same number divided by two. If they
// ever disagree, one of them has invented a boundary -- which is the failure
// 14.2's rule exists to prevent, and the only place it can be observed is here,
// where both numbers exist at once.
`timescale 1ns/1ps
module spi_slave_rx_tb;
localparam MAX_W = 32;
localparam LEN_W = 6;
localparam CNT_W = 12;
reg clk;
reg rst_n;
always #5 clk = ~clk;
// --- 14.1 ---------------------------------------------------------------
reg cpol;
reg sclk_pin;
reg cs_n_pin;
reg mosi_pin;
wire sclk_q, cs_active, mosi_q;
wire edge_a_stb, edge_b_stb;
wire cs_assert_stb, cs_deassert_stb;
wire [11:0] min_half;
wire ratio_err;
spi_slave_frontend #(.SYNC_N(2), .HALF_MIN(3), .CNT_W(12)) u_fe (
.clk(clk), .rst_n(rst_n), .cpol(cpol),
.sclk_pin(sclk_pin), .cs_n_pin(cs_n_pin), .mosi_pin(mosi_pin),
.sclk_q(sclk_q), .cs_active(cs_active), .mosi_q(mosi_q),
.edge_a_stb(edge_a_stb), .edge_b_stb(edge_b_stb),
.cs_assert_stb(cs_assert_stb), .cs_deassert_stb(cs_deassert_stb),
.min_half(min_half), .ratio_err(ratio_err), .clr_flags(1'b0)
);
// --- 14.2 ---------------------------------------------------------------
reg [LEN_W-1:0] len;
reg lsb_first;
reg cpha_now;
wire txn_active, txn_start_stb, txn_end_stb;
wire [CNT_W-1:0] edges_in_txn, frames_in_txn;
wire txn_clean, txn_trunc, txn_empty, txn_report_stb;
wire [2:0] cs_state;
spi_slave_cs #(.LEN_W(LEN_W), .CNT_W(CNT_W)) u_cs (
.clk(clk), .rst_n(rst_n),
.cs_assert_stb(cs_assert_stb), .cs_deassert_stb(cs_deassert_stb),
.edge_a_stb(edge_a_stb), .edge_b_stb(edge_b_stb), .len(len),
.txn_active(txn_active), .txn_start_stb(txn_start_stb),
.txn_end_stb(txn_end_stb),
.edges_in_txn(edges_in_txn), .frames_in_txn(frames_in_txn),
.txn_clean(txn_clean), .txn_trunc(txn_trunc), .txn_empty(txn_empty),
.txn_report_stb(txn_report_stb), .state_id(cs_state)
);
// The mode logic is Chapter 14.6; here the phase simply selects which
// recovered edge is the capture.
wire cap_stb = txn_active & (cpha_now ? edge_b_stb : edge_a_stb);
// --- 14.3, the block under test ----------------------------------------
wire [MAX_W-1:0] rx_data;
wire rx_valid_stb;
wire [LEN_W-1:0] bit_idx;
wire [CNT_W-1:0] words_in_txn;
wire [MAX_W-1:0] rx_partial_sr;
spi_slave_rx #(.MAX_W(MAX_W), .LEN_W(LEN_W), .CNT_W(CNT_W)) dut (
.clk(clk), .rst_n(rst_n),
.txn_active(txn_active), .txn_start_stb(txn_start_stb),
.cap_stb(cap_stb), .mosi_q(mosi_q),
.len(len), .lsb_first(lsb_first),
.rx_data(rx_data), .rx_valid_stb(rx_valid_stb),
.bit_idx(bit_idx), .words_in_txn(words_in_txn),
.rx_partial_sr(rx_partial_sr)
);
// --- the expectation ----------------------------------------------------
integer seed;
reg [MAX_W-1:0] exp_q [0:63];
integer exp_w, got_w;
integer word_bad, upper_bad, valid_at_wrong_time;
integer n_valid;
always @(posedge clk) begin
if (rst_n) begin
if (txn_start_stb) begin
got_w <= 0;
end
if (rx_valid_stb) begin
n_valid <= n_valid + 1;
if (got_w < exp_w) begin
if (rx_data !== exp_q[got_w]) word_bad <= word_bad + 1;
// Nothing above the frame width, ever. This is the property
// a narrow word following another narrow word breaks.
if (len < MAX_W && (rx_data >> len) != {MAX_W{1'b0}})
upper_bad <= upper_bad + 1;
end else begin
valid_at_wrong_time <= valid_at_wrong_time + 1;
end
got_w <= got_w + 1;
end
end
end
integer errors;
task adv;
input integer n;
begin repeat (n) @(negedge clk); end
endtask
task clear_exp;
begin
exp_w = 0; got_w = 0;
end
endtask
// Drives a whole transaction of `nwords` words of `nbits` bits, in the given
// mode, with the bit order the SLAVE is configured for -- so what goes on the
// wire is the order that configuration implies, and the received word must
// come back equal to the value sent.
task drive_txn;
input integer half;
input pol;
input pha;
input integer nbits;
input integer nwords;
input lsb;
integer w, i, seedl;
reg [MAX_W-1:0] word;
reg b;
begin
cpol = pol;
cpha_now = pha;
len = nbits[LEN_W-1:0];
lsb_first = lsb;
sclk_pin = pol;
cs_n_pin = 1'b1;
adv(8);
clear_exp();
cs_n_pin = 1'b0;
adv(4);
for (w = 0; w < nwords; w = w + 1) begin
seed = (seed * 32'h0019_660D) + 32'h3C6E_F35F;
word = seed & ((nbits >= 32) ? 32'hFFFF_FFFF
: ((32'h1 << nbits) - 1));
exp_q[exp_w] = word;
exp_w = exp_w + 1;
for (i = 0; i < nbits; i = i + 1) begin
// The wire carries the configured order: bit i of the word
// first when LSB-first, bit nbits-1-i first otherwise.
b = lsb ? word[i] : word[nbits-1-i];
if (!pha) begin
// Capture is on the leading edge, so the bit must be on
// the wire before it -- driven at CS assert for the first
// bit and on the trailing edge for the rest.
mosi_pin = b;
adv(2);
sclk_pin = ~pol; // leading: captured here
adv(half);
sclk_pin = pol; // trailing
adv(half - 2 < 1 ? 1 : half - 2);
end else begin
sclk_pin = ~pol; // leading: launch
mosi_pin = b;
adv(half);
sclk_pin = pol; // trailing: captured here
adv(half);
end
end
end
adv(4);
cs_n_pin = 1'b1;
adv(8);
sclk_pin = pol;
adv(8);
end
endtask
integer h, p, o, nb, nw, k;
initial begin
exp_w = 0; got_w = 0; n_valid = 0;
word_bad = 0; upper_bad = 0; valid_at_wrong_time = 0;
seed = 32'h0C4F_7E21;
adv(3);
rst_n = 1'b1;
adv(2);
// 1. ONE BYTE, MSB-first, mode 0. The simplest thing that can work.
drive_txn(4, 1'b0, 1'b0, 8, 1, 1'b0);
if (got_w != 1) begin
$display(" FAIL: one byte produced %0d received words", got_w);
errors = errors + 1;
end
$display(" one byte, mode 0, MSB-first: received %02h and expected %02h",
rx_data[7:0], exp_q[0][7:0]);
// 2. SEVERAL WORDS IN ONE TRANSACTION, in order. The word boundary comes
// from a counter whose zero is the select pin, so a transaction of
// four words must produce four words and the fourth must not be the
// first one's bits shifted.
drive_txn(4, 1'b0, 1'b0, 8, 4, 1'b0);
if (got_w != 4) begin
$display(" FAIL: four bytes produced %0d received words", got_w);
errors = errors + 1;
end
if (words_in_txn != 4 || frames_in_txn != 4) begin
$display(" FAIL: the capture counter says %0d words and the edge counter says %0d frames",
words_in_txn, frames_in_txn);
errors = errors + 1;
end
$display(" four bytes in one transaction: four words received in order, and the capture and edge counters agree");
// 3. NARROW WORDS IN SEQUENCE. Five-bit words, one after another, with
// different values -- the case where a receive register that is not
// cleared at the word boundary carries the previous word's bits above
// the current one.
drive_txn(4, 1'b0, 1'b0, 5, 6, 1'b0);
if (got_w != 6) begin
$display(" FAIL: six five-bit words produced %0d", got_w);
errors = errors + 1;
end
$display(" six five-bit words in one transaction: each right-aligned with nothing above bit 4");
// 4. BOTH BIT ORDERS. The wire carries a different sequence in each, so
// a slave that ignored the setting would return the bit-reverse.
drive_txn(4, 1'b0, 1'b0, 8, 3, 1'b0);
drive_txn(4, 1'b0, 1'b0, 8, 3, 1'b1);
$display(" both bit orders returned the value that was sent, which a slave ignoring the setting could not do");
// 5. THE SWEEP. Widths from 1 to 32, both orders, both phases, several
// ratios, multi-word transactions.
for (o = 0; o <= 1; o = o + 1)
for (p = 0; p <= 1; p = p + 1)
for (h = 2; h <= 6; h = h + 2) begin
drive_txn(h, 1'b0, p[0], 1, 3, o[0]);
drive_txn(h, 1'b0, p[0], 2, 3, o[0]);
drive_txn(h, 1'b0, p[0], 5, 3, o[0]);
drive_txn(h, 1'b0, p[0], 8, 3, o[0]);
drive_txn(h, 1'b0, p[0], 13, 2, o[0]);
drive_txn(h, 1'b0, p[0], 32, 2, o[0]);
end
$display(" %0d words swept: widths 1, 2, 5, 8, 13 and 32, both bit orders, both phases, three ratios",
n_valid);
// 6. A TRANSACTION THAT ENDS MID-WORD. The completed words must be
// unaffected and the partial count must be readable -- the policy for
// what to DO with it is Chapter 14.7's.
len = 6'd8; lsb_first = 1'b0; cpha_now = 1'b0;
adv(2);
clear_exp();
// The bits driven below are k[0] for k = 0..10, so the first eight --
// 0 1 0 1 0 1 0 1 read MSB-first -- are 0x55. Registering it matters:
// an unexpected word and a missing word are different failures, and a
// monitor that cannot tell them apart reports the wrong one.
exp_q[0] = 32'h55;
exp_w = 1;
cpol = 1'b0; sclk_pin = 1'b0; cs_n_pin = 1'b1; adv(8);
cs_n_pin = 1'b0; adv(4);
// one complete byte, then three bits of a second
for (k = 0; k < 11; k = k + 1) begin
mosi_pin = k[0];
adv(2);
sclk_pin = 1'b1; adv(4);
sclk_pin = 1'b0; adv(2);
end
adv(4);
if (bit_idx != 3) begin
$display(" FAIL: after 11 bits the partial count was %0d, expected 3",
bit_idx);
errors = errors + 1;
end
if (words_in_txn != 1) begin
$display(" FAIL: after 11 bits %0d words completed, expected 1",
words_in_txn);
errors = errors + 1;
end
cs_n_pin = 1'b1; adv(8); sclk_pin = 1'b0; adv(8);
if (!txn_trunc) begin
$display(" FAIL: a transaction ending mid-word was not reported truncated");
errors = errors + 1;
end
$display(" a transaction of 11 bits at eight bits per word: one complete word, a partial count of 3, and the transaction reported truncated");
// 7. THE CONTINUOUS PROPERTIES.
if (word_bad != 0) begin
$display(" FAIL: %0d of %0d received words were wrong",
word_bad, n_valid);
errors = errors + 1;
end
if (upper_bad != 0) begin
$display(" FAIL: %0d words carried bits above the frame width",
upper_bad);
errors = errors + 1;
end
if (valid_at_wrong_time != 0) begin
$display(" FAIL: %0d words arrived that nothing was sent for",
valid_at_wrong_time);
errors = errors + 1;
end
$display(" %0d received words, none wrong, none carrying bits above its width, and none unaccounted for",
n_valid);
if (errors == 0)
$display("PASS: MOSI is captured on the mode's capture edge and assembled into words whose boundaries come from a counter cleared by the select pin, so nothing the slave counts can redefine where a word ends -- all %0d words across widths of 1, 2, 5, 8, 13 and 32 bits, both bit orders, both phases and three frequency ratios came back exactly as sent, every one right-aligned with nothing above its width even when narrow words of differing values followed one another, the capture counter and the edge counter of Chapter 14.2 agree on every transaction, and a transaction ending mid-word leaves its completed words intact with a readable partial count and is reported as truncated", n_valid);
else
$display("FAIL: %0d error(s)", errors);
$finish;
end
initial begin
clk = 1'b0;
rst_n = 1'b0;
cpol = 1'b0;
sclk_pin = 1'b0;
cs_n_pin = 1'b1;
mosi_pin = 1'b0;
len = 6'd8;
lsb_first = 1'b0;
cpha_now = 1'b0;
errors = 0;
end
endmodule-- spi_slave_rx_tb.vhd
--
-- The real front end of 14.1 and the real transaction detector of 14.2 sit in
-- front of the block under test, so every capture is a recovered one and the word
-- boundaries come from the select pin exactly as they will in the finished slave.
--
-- Two properties get most of the attention because both are invisible to the
-- obvious test:
--
-- NARROW WORDS IN SEQUENCE. A five-bit word followed by another five-bit word
-- must not carry the first word's bits in the positions above the second's. A
-- suite whose words are all as wide as the datapath cannot fail this, and one
-- whose narrow words are all identical cannot either.
--
-- THE WORD COUNT MUST AGREE WITH THE EDGE COUNT. Chapter 14.2 counts edges and
-- this block counts captures; they are the same number divided by two. If they
-- ever disagree, one of them has invented a boundary -- and the only place that
-- can be observed is here, where both numbers exist at once.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_slave_rx_tb is
end entity;
architecture sim of spi_slave_rx_tb is
constant MAX_W : positive := 32;
constant LEN_W : positive := 6;
constant CNT_W : positive := 12;
type word_array is array (0 to 63) of std_logic_vector(MAX_W - 1 downto 0);
signal clk : std_logic := '0';
signal rst_n : std_logic := '0';
signal halt : boolean := false;
signal cpol : std_logic := '0';
signal sclk_pin : std_logic := '0';
signal cs_n_pin : std_logic := '1';
signal mosi_pin : std_logic := '0';
signal sclk_q, cs_active, mosi_q : std_logic;
signal edge_a_stb, edge_b_stb : std_logic;
signal cs_assert_stb, cs_deassert_stb : std_logic;
signal min_half : unsigned(11 downto 0);
signal ratio_err : std_logic;
signal len : unsigned(LEN_W - 1 downto 0) := to_unsigned(8, LEN_W);
signal lsb_first : std_logic := '0';
signal cpha_now : std_logic := '0';
signal txn_active, txn_start_stb, txn_end_stb : std_logic;
signal edges_in_txn, frames_in_txn : unsigned(CNT_W - 1 downto 0);
signal txn_clean, txn_trunc, txn_empty, txn_report_stb : std_logic;
signal cs_state : unsigned(2 downto 0);
signal cap_stb : std_logic;
signal rx_data : std_logic_vector(MAX_W - 1 downto 0);
signal rx_valid_stb : std_logic;
signal bit_idx : unsigned(LEN_W - 1 downto 0);
signal words_in_txn : unsigned(CNT_W - 1 downto 0);
signal rx_partial_sr : std_logic_vector(MAX_W - 1 downto 0);
signal exp_q : word_array := (others => (others => '0'));
signal exp_w : natural := 0;
signal got_w : natural := 0;
signal word_bad, upper_bad, valid_at_wrong_time, n_valid : natural := 0;
signal errors : natural := 0;
function hex8(v : std_logic_vector) return string is
constant DIGITS : string(1 to 16) := "0123456789abcdef";
variable u : unsigned(MAX_W - 1 downto 0);
variable r : string(1 to 8);
begin
u := unsigned(v);
for k in 8 downto 1 loop
r(k) := DIGITS(to_integer(u(3 downto 0)) + 1);
u := shift_right(u, 4);
end loop;
return r;
end function;
begin
clk <= not clk after 5 ns when not halt else '0';
u_fe : entity work.spi_slave_frontend
generic map (SYNC_N => 2, HALF_MIN => 3, CNT_W => 12)
port map (clk => clk, rst_n => rst_n, cpol => cpol,
sclk_pin => sclk_pin, cs_n_pin => cs_n_pin,
mosi_pin => mosi_pin,
sclk_q => sclk_q, cs_active => cs_active, mosi_q => mosi_q,
edge_a_stb => edge_a_stb, edge_b_stb => edge_b_stb,
cs_assert_stb => cs_assert_stb,
cs_deassert_stb => cs_deassert_stb,
min_half => min_half, ratio_err => ratio_err,
clr_flags => '0');
u_cs : entity work.spi_slave_cs
generic map (LEN_W => LEN_W, CNT_W => CNT_W)
port map (clk => clk, rst_n => rst_n,
cs_assert_stb => cs_assert_stb,
cs_deassert_stb => cs_deassert_stb,
edge_a_stb => edge_a_stb, edge_b_stb => edge_b_stb,
len => len,
txn_active => txn_active, txn_start_stb => txn_start_stb,
txn_end_stb => txn_end_stb,
edges_in_txn => edges_in_txn, frames_in_txn => frames_in_txn,
txn_clean => txn_clean, txn_trunc => txn_trunc,
txn_empty => txn_empty,
txn_report_stb => txn_report_stb, state_id => cs_state);
-- The mode logic is Chapter 14.6; here the phase simply selects which
-- recovered edge is the capture.
cap_stb <= txn_active and edge_b_stb when cpha_now = '1'
else txn_active and edge_a_stb;
dut : entity work.spi_slave_rx
generic map (MAX_W => MAX_W, LEN_W => LEN_W, CNT_W => CNT_W)
port map (clk => clk, rst_n => rst_n,
txn_active => txn_active, txn_start_stb => txn_start_stb,
cap_stb => cap_stb, mosi_q => mosi_q,
len => len, lsb_first => lsb_first,
rx_data => rx_data, rx_valid_stb => rx_valid_stb,
bit_idx => bit_idx, words_in_txn => words_in_txn,
rx_partial_sr => rx_partial_sr);
monitor : process (clk)
variable above : std_logic_vector(MAX_W - 1 downto 0);
begin
if rising_edge(clk) then
if rst_n = '1' then
if txn_start_stb = '1' then
got_w <= 0;
end if;
if rx_valid_stb = '1' then
n_valid <= n_valid + 1;
if got_w < exp_w then
if rx_data /= exp_q(got_w) then
word_bad <= word_bad + 1;
end if;
-- Nothing above the frame width, ever. This is the
-- property a narrow word following another breaks.
if to_integer(len) < MAX_W then
above := std_logic_vector(
shift_right(unsigned(rx_data),
to_integer(len)));
if above /= (above'range => '0') then
upper_bad <= upper_bad + 1;
end if;
end if;
else
valid_at_wrong_time <= valid_at_wrong_time + 1;
end if;
got_w <= got_w + 1;
end if;
end if;
end if;
end process;
stim : process
variable errs : natural := 0;
variable seed : unsigned(31 downto 0) := x"0C4F7E21";
variable word : std_logic_vector(MAX_W - 1 downto 0);
variable b : std_logic;
variable mask : unsigned(MAX_W - 1 downto 0);
procedure adv(n : natural) is
begin
for k in 1 to n loop wait until falling_edge(clk); end loop;
end procedure;
-- Only `exp_w` is cleared here. `got_w` belongs to the monitor process
-- and is cleared there on the transaction start: VHDL resolves multiple
-- drivers rather than letting the last writer win, so a signal written
-- from both processes has no defined value at all.
procedure clear_exp is
begin
exp_w <= 0;
end procedure;
-- Drives a whole transaction of `nwords` words of `nbits` bits, in the
-- given mode, with the bit order the SLAVE is configured for -- so what
-- goes on the wire is the order that configuration implies.
procedure drive_txn(half : natural; pol : std_logic; pha : std_logic;
nbits : natural; nwords : natural;
lsb : std_logic) is
variable n_exp : natural := 0;
begin
cpol <= pol;
cpha_now <= pha;
len <= to_unsigned(nbits, LEN_W);
lsb_first <= lsb;
sclk_pin <= pol;
cs_n_pin <= '1';
adv(8);
clear_exp;
n_exp := 0;
cs_n_pin <= '0';
adv(4);
for w in 0 to nwords - 1 loop
seed := resize(seed * x"0019660D", 32) + x"3C6EF35F";
mask := (others => '0');
for k in 0 to MAX_W - 1 loop
if k < nbits then mask(k) := '1'; end if;
end loop;
word := std_logic_vector(seed and mask);
exp_q(n_exp) <= word;
n_exp := n_exp + 1;
exp_w <= n_exp;
for i in 0 to nbits - 1 loop
if lsb = '1' then
b := word(i);
else
b := word(nbits - 1 - i);
end if;
if pha = '0' then
-- Capture is on the leading edge, so the bit must be on
-- the wire before it.
mosi_pin <= b;
adv(2);
sclk_pin <= not pol; -- leading: captured here
adv(half);
sclk_pin <= pol; -- trailing
if half > 2 then adv(half - 2); else adv(1); end if;
else
sclk_pin <= not pol; -- leading: launch
mosi_pin <= b;
adv(half);
sclk_pin <= pol; -- trailing: captured here
adv(half);
end if;
end loop;
end loop;
adv(4);
cs_n_pin <= '1';
adv(8);
sclk_pin <= pol;
adv(8);
end procedure;
variable pha_v, lsb_v : std_logic;
begin
adv(3);
rst_n <= '1';
adv(2);
-- 1. ONE BYTE, MSB-first, mode 0.
drive_txn(4, '0', '0', 8, 1, '0');
if got_w /= 1 then
report " FAIL: one byte produced " & integer'image(got_w) &
" received words";
errs := errs + 1;
end if;
report " one byte, mode 0, MSB-first: received " & hex8(rx_data) &
" and expected " & hex8(exp_q(0));
-- 2. SEVERAL WORDS IN ONE TRANSACTION, in order.
drive_txn(4, '0', '0', 8, 4, '0');
if got_w /= 4 then
report " FAIL: four bytes produced " & integer'image(got_w) &
" received words";
errs := errs + 1;
end if;
if to_integer(words_in_txn) /= 4 or to_integer(frames_in_txn) /= 4 then
report " FAIL: the capture counter says " &
integer'image(to_integer(words_in_txn)) &
" words and the edge counter says " &
integer'image(to_integer(frames_in_txn)) & " frames";
errs := errs + 1;
end if;
report " four bytes in one transaction: four words received in order, and the capture and edge counters agree";
-- 3. NARROW WORDS IN SEQUENCE.
drive_txn(4, '0', '0', 5, 6, '0');
if got_w /= 6 then
report " FAIL: six five-bit words produced " &
integer'image(got_w);
errs := errs + 1;
end if;
report " six five-bit words in one transaction: each right-aligned with nothing above bit 4";
-- 4. BOTH BIT ORDERS.
drive_txn(4, '0', '0', 8, 3, '0');
drive_txn(4, '0', '0', 8, 3, '1');
report " both bit orders returned the value that was sent, which a slave ignoring the setting could not do";
-- 5. THE SWEEP.
for o in 0 to 1 loop
if o = 1 then lsb_v := '1'; else lsb_v := '0'; end if;
for p in 0 to 1 loop
if p = 1 then pha_v := '1'; else pha_v := '0'; end if;
for h in 1 to 3 loop
drive_txn(h * 2, '0', pha_v, 1, 3, lsb_v);
drive_txn(h * 2, '0', pha_v, 2, 3, lsb_v);
drive_txn(h * 2, '0', pha_v, 5, 3, lsb_v);
drive_txn(h * 2, '0', pha_v, 8, 3, lsb_v);
drive_txn(h * 2, '0', pha_v, 13, 2, lsb_v);
drive_txn(h * 2, '0', pha_v, 32, 2, lsb_v);
end loop;
end loop;
end loop;
report " " & integer'image(n_valid) &
" words swept: widths 1, 2, 5, 8, 13 and 32, both bit orders, both phases, three ratios";
-- 6. A TRANSACTION THAT ENDS MID-WORD.
len <= to_unsigned(8, LEN_W);
lsb_first <= '0';
cpha_now <= '0';
adv(2);
clear_exp;
-- The bits driven below are k mod 2 for k = 0..10, so the first eight --
-- 0 1 0 1 0 1 0 1 read MSB-first -- are 0x55. Registering it matters: an
-- unexpected word and a missing word are different failures.
exp_q(0) <= x"00000055";
exp_w <= 1;
cpol <= '0'; sclk_pin <= '0'; cs_n_pin <= '1'; adv(8);
cs_n_pin <= '0'; adv(4);
for k in 0 to 10 loop
if (k mod 2) = 0 then mosi_pin <= '0'; else mosi_pin <= '1'; end if;
adv(2);
sclk_pin <= '1'; adv(4);
sclk_pin <= '0'; adv(2);
end loop;
adv(4);
if to_integer(bit_idx) /= 3 then
report " FAIL: after 11 bits the partial count was " &
integer'image(to_integer(bit_idx)) & ", expected 3";
errs := errs + 1;
end if;
if to_integer(words_in_txn) /= 1 then
report " FAIL: after 11 bits " &
integer'image(to_integer(words_in_txn)) &
" words completed, expected 1";
errs := errs + 1;
end if;
cs_n_pin <= '1'; adv(8); sclk_pin <= '0'; adv(8);
if txn_trunc /= '1' then
report " FAIL: a transaction ending mid-word was not reported truncated";
errs := errs + 1;
end if;
report " a transaction of 11 bits at eight bits per word: one complete word, a partial count of 3, and the transaction reported truncated";
-- 7. THE CONTINUOUS PROPERTIES.
if word_bad /= 0 then
report " FAIL: " & integer'image(word_bad) & " of " &
integer'image(n_valid) & " received words were wrong";
errs := errs + 1;
end if;
if upper_bad /= 0 then
report " FAIL: " & integer'image(upper_bad) &
" words carried bits above the frame width";
errs := errs + 1;
end if;
if valid_at_wrong_time /= 0 then
report " FAIL: " & integer'image(valid_at_wrong_time) &
" words arrived that nothing was sent for";
errs := errs + 1;
end if;
report " " & integer'image(n_valid) &
" received words, none wrong, none carrying bits above its width, and none unaccounted for";
errors <= errs;
if errs = 0 then
report "PASS: MOSI is captured on the mode's capture edge and assembled into words whose boundaries come from a counter cleared by the select pin, so nothing the slave counts can redefine where a word ends -- all " & integer'image(n_valid) & " words across widths of 1, 2, 5, 8, 13 and 32 bits, both bit orders, both phases and three frequency ratios came back exactly as sent, every one right-aligned with nothing above its width even when narrow words of differing values followed one another, the capture counter and the edge counter of Chapter 14.2 agree on every transaction, and a transaction ending mid-word leaves its completed words intact with a readable partial count and is reported as truncated";
else
report "FAIL: " & integer'image(errs) & " error(s)" severity error;
end if;
halt <= true;
wait;
end process;
end architecture;7. Why a Verification Engineer Cares
rx_valid_stb is registered, and the reason is a bug from Module 13 that repeats here. If the strobe is combinational from "the counter has reached len", it fires on the same cycle as the final capture — and on that cycle the shift register has not yet absorbed the final bit, because both are clocked by the same edge. The published word is one bit short, shifted, with a zero in the low position.
The symptom is what makes it worth telling: it presents as a receive-side shift, which sends everyone looking at the capture edge, the synchroniser depth, and the master's output timing. The actual fault is one cycle of strobe timing. Chapter 13.6 had the same bug with the same misleading symptom, and the general rule is:
A "data valid" strobe must be delayed by one cycle from the condition that makes the data complete, whenever the data and the condition are clocked together.
The property to check is word conservation plus per-word content, not per-bit timing. Count the words the bench sent and the words the DUT published; then compare content by index. A checker that tries to predict when each word appears re-implements the counter, and then agrees with the design about its off-by-one.
Sweep the width, and include 1 and MAX_W. A width of 1 makes every capture a word boundary, which is the only case where the boundary logic runs on consecutive cycles. A width of MAX_W makes the reversal touch every bit and the right-alignment a no-op. Both are boundaries and both are where the arithmetic is most likely to be wrong.
// Properties for the receive path.
property p_valid_only_in_txn;
// A word can only be published inside a transaction. A strobe outside one
// means the counter is being advanced by something other than the capture
// strobe -- which is the failure mode of a counter that rolls over freely.
@(posedge clk) disable iff (!rst_n)
rx_valid_stb |-> txn_active;
endproperty
property p_bit_idx_in_range;
// The position within the word is always a legal position. An index that
// reaches len means the boundary comparison is off by one, and the symptom
// is words of len+1 bits that look like a shift.
@(posedge clk) disable iff (!rst_n)
txn_active |-> (bit_idx < len);
endproperty
property p_counter_zeroed_at_start;
// The zero comes from the pin. This is section 1's rule, checkable in one
// line, and it is the property that makes the next transaction correct
// after this one has slipped.
@(posedge clk) disable iff (!rst_n)
txn_start_stb |=> (bit_idx == 0);
endproperty
property p_valid_follows_final_capture;
// The strobe is exactly one cycle after the capture that completed the
// word -- not the same cycle. This is the off-by-one in section 7, stated
// as a property rather than left to a code review.
@(posedge clk) disable iff (!rst_n)
(cap_stb && bit_idx == len - 1) |=> rx_valid_stb;
endproperty
property p_no_valid_without_final_capture;
// And the converse, which is what catches a strobe generated from a level
// rather than from the capture.
@(posedge clk) disable iff (!rst_n)
rx_valid_stb |-> $past(cap_stb);
endproperty
property p_partial_tracks_sr;
// The published in-progress word changes only on a capture. A partial that
// moves on any other cycle means 14.7 is reading a register that something
// else writes.
@(posedge clk) disable iff (!rst_n)
!cap_stb |=> $stable(rx_partial_sr);
endproperty// Coverage. The two axes are WIDTH and ORDER, and they must be crossed, because
// the reversal is a function of the width and the alignment is a function of
// both.
covergroup cg_rx @(posedge clk iff rx_valid_stb);
option.per_instance = 1;
// Widths chosen at the boundaries rather than spread evenly: 1 is where a
// word boundary happens every capture, MAX_W is where the reversal touches
// every bit and the right-alignment is a no-op.
width: coverpoint len {
bins one = {1};
bins narrow = {[2:7]};
bins byte_w = {8};
bins odd = {5, 9, 13, 17}; // widths that are not powers of two
bins wide = {[18:31]};
bins full = {32};
}
order: coverpoint lsb_first { bins msb = {0}; bins lsb = {1}; }
// How many words this transaction has produced by the time this one
// arrives. The first word of a transaction and the fifth exercise
// different paths through the counter's reset.
position: coverpoint words_in_txn {
bins first = {1};
bins second = {2};
bins middle = {[3:7]};
bins many = {[8:$]};
}
x_width_order: cross width, order;
x_width_pos: cross width, position;
endgroup8. Why an FPGA or ASIC Engineer Cares
The shift register is MAX_W bits wide and fully utilised regardless of len. Words shift into the low end and are right-aligned, so a 5-bit frame uses 5 of 32 flops for data and the rest hold history that is masked off. That is the cheap choice: the alternative is a variable-length shift register, which means a barrel shifter, which is far more logic than 27 idle flops.
The reversal is a MAX_W-wide mux per bit, and it is the widest combinational path in the block. It sits between the shift register and the output register, with a full system clock to complete — never critical at any plausible SCLK. But it is MAX_W wide, so on a design with MAX_W = 32 it is 32 muxes, and that is worth knowing before someone parameterises MAX_W to 64 for one command that needs it.
The capture counter is LEN_W bits and compares against a run-time len. One comparator, one incrementer, no arithmetic on a variable. bit_idx is published so that Chapter 14.7 can size a partial without recomputing it.
Nothing here is clocked by SCLK. Worth restating in an implementation section because it is what makes the block ordinary: it is a shift register and a counter on the system clock, advanced by a strobe, and every timing question about it is a system-clock question.
9. Failure Signature — Every Byte Is Right Except The First
The symptom:
"Multi-byte reads work. The first byte of every transaction is wrong — it looks like the previous transaction's last byte shifted left by one."
What is happening: rx_valid_stb is combinational from the counter comparison, so it fires on the same cycle as the final capture. The shift register has not absorbed that last bit yet. Every published word is the previous len - 1 bits plus a zero — which, for the second and subsequent words of a transaction, is the previous word's tail, and for the first word of a transaction is whatever the register held.
Why it reads as a capture-edge problem rather than a strobe problem:
- The corruption is a shift, and a shift is what a wrong capture edge produces.
- It is consistent, reproducible and identical at every clock ratio, so it does not look like the synchroniser.
- Widening the synchroniser, changing
CPHA, and slowing the master all fail to fix it, which exhausts the usual suspects and sends people to the master.
How to find it in one waveform: look at rx_valid_stb and the final cap_stb of a word. If they are on the same cycle, that is the bug, and the fix is one register. The general form is in §7 and it is worth remembering because it recurs: a valid strobe and the capture that completes the data cannot share a cycle.
10. Common Misconceptions
"A bit counter that wraps at len is equivalent to one reset by chip select, because chip select resets it anyway." In most designs it does, which is why the decision gets made carelessly. The difference is what happens within a transaction after a slip: the wrapping counter continues confidently, and its confidence is the problem.
"A one-bit slip will be obvious in the data." It produces words of exactly the right length at exactly the right times. A status register full of zeros is still zeros. Only the edge count reveals it, and only at the deassert.
"The slave should be able to detect a bit-order mismatch — the data will look like nonsense." Sometimes it will, and a detector that fires on "looks like nonsense" is a detector that will be trusted the times it is wrong. There is no redundancy in SPI to distinguish the two orders. Configure it correctly and report nothing.
"Reversing the bit order means reversing the shift direction." It means reversing the word at the boundary, which is one function shared with the master and its own inverse. Reversing the shift direction gives two shift paths, two alignment rules, and two places for len to interact with them.
"Publishing the in-progress shift register breaks the layering." The alternative is two copies of the same state in two blocks, updated by the same strobe. One published register is a smaller cost than one duplicated one, and Chapter 14.7 has no shift logic at all as a result.
"rx_valid_stb should fire as early as possible to reduce latency." It should fire as early as the data is complete, which is one cycle after the final capture. Firing on the capture cycle does not reduce latency; it publishes a word that is one bit short.
11. Reason It Through
Q. len = 8 and the slave receives 8 captures, then chip select deasserts. rx_valid_stb fired once. Now len = 8 and the slave receives 9 captures before the deassert. How many times does rx_valid_stb fire, and what happens to the ninth bit?
Once, and the ninth bit stays in the shift register as the first bit of an incomplete word. rx_valid_stb fires only at a word boundary, and the ninth capture is one bit into the second word. The ninth bit is not lost — it is published on rx_partial_sr, and Chapter 14.7 reports it on the partial channel with partial_bits = 1. It must not appear as a word, because a one-bit word is not what the configuration asked for.
Q. Why does a width of 1 deserve its own coverage bin?
Because it is the only width at which the word boundary occurs on consecutive captures. Every other width gives the boundary logic at least one cycle between firings; at len = 1 the comparison, the reversal, the strobe and the counter reset all happen on every capture. Any design that accidentally depends on a gap between word boundaries fails only there.
Q. A reviewer proposes deleting rx_partial_sr and having Chapter 14.7 maintain its own shift register from cap_stb and mosi_q. What is the argument against, beyond duplication?
That the two copies would have to agree about the reset, not just the shift. This block zeroes its register on txn_start_stb; 14.7 would have to do the same, from the same strobe, and also agree about what happens when len changes mid-transaction, and also agree about the reversal. Every one of those is a place the copies can diverge, and a divergence shows up as a partial that disagrees with the words around it — which looks like a partial-path bug in 14.7 rather than a duplication problem.
Q. The slave's len is 8 and the master sends 16-bit frames. What does the receive path produce, and is anything wrong from this block's point of view?
Two 8-bit words per master frame, both well-formed, both published on time. Nothing is wrong from this block's point of view, and nothing is detectable: 32 edges is a whole number of 8-bit frames, so Chapter 14.2 reports txn_clean too. The fault is a width disagreement between two devices, it is invisible to every observation the slave can make, and the only evidence is that frames_in_txn is twice what software expected — which is a statement about software's expectation rather than about the bus.
Q. Why is it correct for rx_valid_stb to fire during the last word of a transaction that is about to be truncated mid-word, but not for the truncated fragment?
Because the last complete word is complete: len bits arrived and were assembled, and the transaction ending afterwards does not retroactively damage them. The fragment is different in kind — it is fewer than len bits — and publishing it as a word would mean the system could not distinguish "a 3-bit word arrived" from "a transaction was cut after 3 bits". Keeping them on separate channels is Chapter 14.7's central decision, and this block enables it by strobing only on whole words.
12. Understanding Check
13. Summary
The capture is one line. The counter is the design, and it takes its zero from the chip-select assert and nothing else — because a counter that rolls over on its own arithmetic has made itself the authority on word boundaries, and an authority that can be wrong without noticing produces well-formed words at the wrong offsets.
A one-bit slip is invisible in the data. Every word is still len bits, still on time, still self-consistent. Only Chapter 14.2's edge count reveals it, and only at the deassert — which is after the wrong words have been published. That ordering is unfixable, and it is why the verdict has to be readable alongside the words rather than separately.
The taxonomy: a polarity mismatch is detectable from a level, a phase mismatch from a motion and not from the edge count, and a bit-order mismatch not at all. The third is a property of the protocol rather than a gap in the design, and the right response is to configure it correctly and report nothing, because a flag that guessed would be trusted the times it was wrong.
Bit order is the master's transform, applied once. Reversing the low len bits is its own inverse, so there is one function rather than two to keep in step, and it is applied at the word boundary rather than in the shift path.
rx_valid_stb is registered, because the comparison that completes a word and the capture that completes it are the same clock edge. A combinational strobe publishes a word one bit short, and the symptom presents as a capture-edge problem — which is what makes it expensive.
rx_partial_sr is published so that Chapter 14.7 needs no shift register of its own. One published register beats one duplicated one.
For verification: check word conservation plus content by index, not per-bit timing; sweep the width including 1 and MAX_W, because those are the boundaries where the arithmetic breaks; and keep the reference model a pure function of the wire and the configuration, with no shift register in it.
For implementation: the shift register is MAX_W wide regardless of len, because masking idle flops is cheaper than a barrel shifter; the reversal is the widest combinational path and it is MAX_W muxes, which matters if someone widens MAX_W for one command.
14. What Comes Next
The slave can receive. It cannot yet say anything back.
Chapter 14.4 — MISO Generation and Launch Timing builds the transmit path, and it contains the hardest timing problem in the module: the first bit. Under CPHA=0 the master samples MOSI and MISO on the very first SCLK edge, which means the slave must have MISO on the wire before any edge exists — with nothing but the chip-select assert to trigger it, and that assert arriving SYNC_N cycles late. The chapter derives exactly how much lead time the slave needs, measures it, and reports when a master does not provide it.
Continue learning
Related tutorials
- Related topic
TX and RX Shift Registers
The shift datapath, and why there are two registers rather than the one SPI's symmetry seems to permit: the three failures a circulating register cannot express, why transmit needs alignment and receive does not, and why MOSI must be a flop.
- Related topic
Configurable Transfer Width and Bit Order
Frame width from 1 to 32 bits and a choice of bit order without touching the shift register: why a bidirectional shifter is the wrong trade, why reversing the low bits is its own inverse, and why loopback cannot test bit order at all.
- Related topic
Slave Microarchitecture and Clocking Assumptions
The two architectures available to a slave that does not own its clock, the one precondition the chosen architecture rests on and why no simulation can test it, why MOSI must pass through exactly as many flops as SCLK, and a delay-matched front end verified in three HDLs.
- Related topic
CS Detection and Transaction Boundaries
Chip select is SPI's only framing and therefore its only resynchronisation point: why counters must reset on assert, how to classify a transaction with a running remainder instead of a divider, why a CPHA mismatch is invisible to the edge count, and a transaction detector verified in three HDLs.
