SPI · Module 4
Transfer Width
SPI has no native word size. What a 12-bit ADC or 16-bit codec requires of a master, why chip select rather than a bit count delimits a frame, and a width-parameterized transfer engine in three HDLs.
Chapter 4.1 established that SPI defines signalling and leaves meaning to the device, and listed the fields it declines to specify. This chapter takes the first of them.
If SPI has no native word size, what decides how many bits a transfer contains — and what does the hardware at each end actually have to do about it?
The question sounds trivial and is not. The answer determines a counter in your RTL, a configuration field in your verification environment, and one of the most commonly misdiagnosed corruption patterns in SPI bring-up.
1. Eight Bits Is a Convention, Not a Rule
There is no eight anywhere in SPI. The bus clocks edges; it does not group them. A transfer is as long as the master keeps clocking, and it ends when chip select deasserts.
Eight bits dominates for reasons that are historical and practical rather than technical:
- Host software is byte-oriented. A driver hands the controller a buffer of bytes; a byte-sized transfer makes the mapping trivial.
- Controller hardware is byte-oriented. Most SPI peripherals expose an eight-bit data register and a FIFO of bytes, so eight bits is what the register file naturally holds.
- Memory devices are byte-addressed. Flash, EEPROM and their command sets are defined in bytes, and they represent a large share of SPI traffic.
None of that is a property of the bus. It is a property of the ecosystem on both sides of it — and the moment a device is not byte-shaped, the convention stops helping.
2. What the Device Actually Counts
Here is the fact that makes width a genuinely hard topic, and it follows directly from Chapter 1.3's model.
A slave does not know how many bits you intend to send. It has a shift register and a clock input. Every active edge shifts. It has no counter that agrees with yours, no way to see your counter, and no notification that you have finished — until CS deasserts.
So a device with a 12-bit register does not check that you sent twelve bits. It shifts on every edge it receives, and at the end of the frame it uses whatever happens to be in its register. Send eleven bits and it acts on eleven bits' worth of shifted state, misaligned by one position. Send thirteen and the first bit has already been shifted out the far end and lost.
This also explains why a truncated frame is not an error the bus can report. If CS deasserts early, the device simply has a partially-shifted register. Whether it discards that state or acts on it is a device design decision, documented — if you are fortunate — in the datasheet. Well-designed devices ignore a frame whose bit count they did not expect; many do not check at all.
3. One Frame or Two — the Distinction That Breaks Systems
This is where width stops being an abstraction. Consider a device with a 16-bit control register, and a master that wants to write 0x1234 to it.
Identical bytes; only chip select differs
6 cyclesA 16-bit device sees the upper case as one complete write and the lower case as two truncated 8-bit transfers, each half a command. It will usually act on neither.
This is why "send two bytes" is an ambiguous instruction on SPI, and why controller APIs distinguish between a single transfer of two bytes and two transfers of one byte. The bytes are the same; the framing is not. On a bus with an explicit frame format this ambiguity cannot arise — which is Chapter 4.1's absence, showing up as an integration bug.
4. Widths You Meet in Practice
A representative sample, and the reason each is shaped that way:
| Width | Typical device | Why |
|---|---|---|
| 8 | Flash, EEPROM, most sensors | Byte-oriented command sets |
| 9 | Some displays | Eight data bits plus a command/data flag |
| 10, 12 | ADCs | The converter's native resolution |
| 16 | Codecs, DACs, control registers | Register-oriented, often address + value packed |
| 24 | Flash addresses, high-resolution ADCs | Three-byte address space, or 24-bit samples |
| 32 | Some ADCs and DACs | Sample plus status flags |
The 9-bit case is worth noting because it defeats every byte-oriented assumption at once: it cannot be expressed as a whole number of bytes, so a controller with an 8-bit data register must either support a 9-bit mode natively or the design must work around it — commonly by bit-banging, or by using a 16-bit transfer and discarding seven bits.
5. Building a Width-Parameterized Engine — Three HDLs
The hardware consequence of everything above is small and specific: a bit counter and the done strobe it produces. The shift register itself is Chapter 1.3's, and the launch/sample strobes are Chapter 3.3's mode decoder output — this module adds only the counting.
The circuit
Circuit. Two shift registers — one transmit, one receive — plus a counter that tracks position within the word.
State. The transmit register's remaining bits, the receive register's accumulated bits, the bit position, and a registered copy of CS used to detect the start of a frame.
Datapath. On each launch strobe the transmit register shifts left, moving the next bit to the MSB tap. On each sample strobe the receive register shifts left with sdi entering at the LSB.
Control. cs_n reloads and resets the count; the counter increments on each sample strobe and wraps at WIDTH-1, where it also latches rx_data and pulses done.
Clock. Everything is clocked by the system clock, not by SCLK. The launch and sample strobes are single-cycle pulses marking the SPI edges — the architecture Chapter 2.3 established, which keeps the design in one clock domain.
Reset. Asynchronous, active-low, to a constant. Every register has a defined start value and the counter starts at zero.
Enables. No state changes except on a strobe or a CS transition, so the engine is idle-safe between transfers.
Timing. done is a single-cycle pulse coincident with the capture of the final bit, not a level — so software or an upstream FSM sees exactly one event per completed word.
Synthesis. A counter of ceil(log2(WIDTH)) flip-flops, two shift registers of WIDTH flip-flops, one comparator, and a small amount of control logic. No latches; no gated clocks.
Limitations. The engine assumes MSB-first, which Chapter 4.3 makes configurable. It does not generate SCLK or CS — those belong to a controller (Chapter 2.1, Chapter 2.5) — and it has no FIFO, which is Module 13's concern.
// spi_frame_width.sv — a transfer engine whose frame length is a parameter.
//
// The only thing this module adds to Chapter 1.3's shift core is a BIT
// COUNTER and the `done` strobe it produces. That counter is the master's
// private bookkeeping: nothing it does appears on the wire, and the slave
// never sees it. The slave counts its own edges.
module spi_frame_width #(
parameter int WIDTH = 8
) (
input logic clk,
input logic rst_n,
input logic cs_n, // frame: asserted low
input logic launch_stb, // one clk pulse on the launch edge
input logic sample_stb, // one clk pulse on the sample edge
input logic [WIDTH-1:0] tx_data,
input logic sdi,
output logic sdo,
output logic [WIDTH-1:0] rx_data,
output logic done // one pulse as bit WIDTH-1 is captured
);
// Counter must hold 0..WIDTH-1, so it needs $clog2(WIDTH) bits, with a
// floor of 1 so WIDTH=1 still elaborates.
localparam int CW = (WIDTH > 1) ? $clog2(WIDTH) : 1;
logic [WIDTH-1:0] tx_sh, rx_sh;
logic [CW-1:0] bit_cnt;
logic cs_n_q;
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
tx_sh <= '0;
rx_sh <= '0;
bit_cnt <= '0;
rx_data <= '0;
done <= 1'b0;
cs_n_q <= 1'b1;
end else begin
cs_n_q <= cs_n;
done <= 1'b0; // default: a strobe, not a level
if (cs_n) begin
// Deselected: reload and reset the count. CS -- not the
// counter -- is what truly delimits the frame.
tx_sh <= tx_data;
bit_cnt <= '0;
end else begin
if (cs_n_q) begin // CS just asserted: fresh frame
tx_sh <= tx_data;
bit_cnt <= '0;
end
if (sample_stb) begin
rx_sh <= {rx_sh[WIDTH-2:0], sdi};
if (bit_cnt == CW'(WIDTH - 1)) begin
bit_cnt <= '0;
rx_data <= {rx_sh[WIDTH-2:0], sdi};
done <= 1'b1;
end else begin
bit_cnt <= bit_cnt + 1'b1;
end
end
if (launch_stb) tx_sh <= {tx_sh[WIDTH-2:0], 1'b0};
end
end
end
assign sdo = tx_sh[WIDTH-1]; // MSB-first (Chapter 4.3 makes this configurable)
endmoduleA note on CW'(WIDTH - 1): the cast keeps the comparison at the counter's own width. Without it the literal is a 32-bit integer and the comparison still works, but tools differ in how loudly they warn about the width mismatch, and being explicit documents the intent.
// spi_frame_width_tb.sv — self-checking: two widths, back-to-back frames.
`timescale 1ns/1ps
module spi_frame_width_tb;
logic clk = 0, rst_n = 0;
always #5 clk = ~clk;
localparam int W8 = 8, W12 = 12;
logic cs_n = 1, launch = 0, sample = 0, sdi = 0;
logic [W8-1:0] tx8; logic sdo8; logic [W8-1:0] rx8; logic done8;
logic [W12-1:0] tx12; logic sdo12; logic [W12-1:0] rx12; logic done12;
spi_frame_width #(.WIDTH(W8)) u8 (.clk, .rst_n, .cs_n, .launch_stb(launch), .sample_stb(sample),
.tx_data(tx8), .sdi, .sdo(sdo8), .rx_data(rx8), .done(done8));
spi_frame_width #(.WIDTH(W12)) u12 (.clk, .rst_n, .cs_n, .launch_stb(launch), .sample_stb(sample),
.tx_data(tx12), .sdi, .sdo(sdo12), .rx_data(rx12), .done(done12));
int errors = 0, done8_pulses = 0, done12_pulses = 0;
task automatic chk(input string what, input int got, input int exp);
if (got !== exp) begin $display("FAIL %s: got %0d exp %0d", what, got, exp); errors++; end
endtask
// Count done pulses so "exactly one per frame" is actually checked.
always @(posedge clk) if (rst_n) begin
if (done8) done8_pulses++;
if (done12) done12_pulses++;
end
// Drive one bit time: sample edge then launch edge (mode-0 ordering).
logic [W12-1:0] rx_expect;
task automatic bit_time(input logic din);
sdi = din;
@(negedge clk); sample = 1; @(negedge clk); sample = 0;
@(negedge clk); launch = 1; @(negedge clk); launch = 0;
endtask
task automatic run_frame(input int nbits, input logic [W12-1:0] din_pat,
output logic [W12-1:0] captured_sdo8);
captured_sdo8 = '0;
cs_n = 0; @(negedge clk);
for (int i = 0; i < nbits; i++) begin
captured_sdo8 = {captured_sdo8[W12-2:0], sdo8}; // sdo BEFORE the edges
bit_time(din_pat[nbits-1-i]);
end
cs_n = 1; @(negedge clk);
endtask
logic [W12-1:0] got_sdo;
initial begin
tx8 = 8'hA5; tx12 = 12'h5A3;
repeat (3) @(negedge clk); rst_n = 1; @(negedge clk);
// --- frame 1: 8 bits. Check MSB-first output and captured input. ---
run_frame(W8, 12'h0C9, got_sdo);
chk("rx8 frame1", rx8, 8'hC9);
chk("sdo8 frame1", got_sdo[W8-1:0], 8'hA5);
chk("done8 pulses after frame1", done8_pulses, 1);
// --- frame 2: back-to-back, new data, must reload from tx_data ---
tx8 = 8'h3C;
run_frame(W8, 12'h017, got_sdo);
chk("rx8 frame2", rx8, 8'h17);
chk("sdo8 frame2", got_sdo[W8-1:0], 8'h3C);
chk("done8 pulses after frame2", done8_pulses, 2);
// --- 12-bit instance over a 12-bit frame ---
done12_pulses = 0;
run_frame(W12, 12'hE47, got_sdo);
chk("rx12", rx12, 12'hE47);
chk("done12 pulses", done12_pulses, 1);
// --- a SHORT frame: CS closes after 5 bits of a 12-bit word.
// done must NOT fire; the engine must recover on the next frame. ---
done12_pulses = 0;
run_frame(5, 12'h01F, got_sdo);
chk("done12 pulses on truncated frame", done12_pulses, 0);
run_frame(W12, 12'h123, got_sdo);
chk("rx12 after truncation", rx12, 12'h123);
chk("done12 pulses after recovery", done12_pulses, 1);
if (errors == 0) $display("PASS: width is a parameter; done fires once per complete frame; a truncated frame is discarded and the engine recovers");
else $display("FAILED with %0d error(s)", errors);
$finish;
end
initial begin #200000; $display("FAIL: watchdog timeout"); $finish; end
endmoduleThe Verilog-2001 version needs one genuine accommodation. $clog2 is a SystemVerilog system function, so the counter width must come from a constant function evaluated at elaboration time — the portable idiom, and worth knowing because it appears wherever Verilog-2001 code is parameterized.
// spi_frame_width.v — the same width-parameterized engine in Verilog-2001.
module spi_frame_width #(
parameter WIDTH = 8
) (
input wire clk,
input wire rst_n,
input wire cs_n,
input wire launch_stb,
input wire sample_stb,
input wire [WIDTH-1:0] tx_data,
input wire sdi,
output wire sdo,
output reg [WIDTH-1:0] rx_data,
output reg done
);
// Verilog-2001 has no $clog2, so the bit width is computed by a constant
// function. This is the portable idiom, evaluated at elaboration time.
function integer clogb2;
input integer value;
integer v;
begin
v = value - 1;
for (clogb2 = 0; v > 0; clogb2 = clogb2 + 1) v = v >> 1;
end
endfunction
localparam CW = (WIDTH > 1) ? clogb2(WIDTH) : 1;
reg [WIDTH-1:0] tx_sh, rx_sh;
reg [CW-1:0] bit_cnt;
reg cs_n_q;
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
tx_sh <= {WIDTH{1'b0}};
rx_sh <= {WIDTH{1'b0}};
bit_cnt <= {CW{1'b0}};
rx_data <= {WIDTH{1'b0}};
done <= 1'b0;
cs_n_q <= 1'b1;
end else begin
cs_n_q <= cs_n;
done <= 1'b0;
if (cs_n) begin
tx_sh <= tx_data;
bit_cnt <= {CW{1'b0}};
end else begin
if (cs_n_q) begin
tx_sh <= tx_data;
bit_cnt <= {CW{1'b0}};
end
if (sample_stb) begin
rx_sh <= {rx_sh[WIDTH-2:0], sdi};
if (bit_cnt == (WIDTH - 1)) begin
bit_cnt <= {CW{1'b0}};
rx_data <= {rx_sh[WIDTH-2:0], sdi};
done <= 1'b1;
end else begin
bit_cnt <= bit_cnt + 1'b1;
end
end
if (launch_stb) tx_sh <= {tx_sh[WIDTH-2:0], 1'b0};
end
end
end
assign sdo = tx_sh[WIDTH-1];
endmodule// spi_frame_width_tb.v — the same checks in Verilog-2001.
`timescale 1ns/1ps
module spi_frame_width_tb;
reg clk = 0, rst_n = 0;
always #5 clk = ~clk;
parameter W8 = 8, W12 = 12;
reg cs_n = 1, launch = 0, sample = 0, sdi = 0;
reg [W8-1:0] tx8; wire sdo8; wire [W8-1:0] rx8; wire done8;
reg [W12-1:0] tx12; wire sdo12; wire [W12-1:0] rx12; wire done12;
spi_frame_width #(.WIDTH(W8)) u8 (
.clk(clk), .rst_n(rst_n), .cs_n(cs_n), .launch_stb(launch), .sample_stb(sample),
.tx_data(tx8), .sdi(sdi), .sdo(sdo8), .rx_data(rx8), .done(done8));
spi_frame_width #(.WIDTH(W12)) u12 (
.clk(clk), .rst_n(rst_n), .cs_n(cs_n), .launch_stb(launch), .sample_stb(sample),
.tx_data(tx12), .sdi(sdi), .sdo(sdo12), .rx_data(rx12), .done(done12));
integer errors = 0, done8_pulses = 0, done12_pulses = 0, i;
reg [W12-1:0] got_sdo, din_pat;
task chk;
input [80*8-1:0] what;
input integer got, exp;
begin
if (got !== exp) begin
$display("FAIL %0s: got %0d exp %0d", what, got, exp);
errors = errors + 1;
end
end
endtask
always @(posedge clk) if (rst_n) begin
if (done8) done8_pulses = done8_pulses + 1;
if (done12) done12_pulses = done12_pulses + 1;
end
task bit_time;
input din;
begin
sdi = din;
@(negedge clk); sample = 1; @(negedge clk); sample = 0;
@(negedge clk); launch = 1; @(negedge clk); launch = 0;
end
endtask
task run_frame;
input integer nbits;
begin
got_sdo = {W12{1'b0}};
cs_n = 0; @(negedge clk);
for (i = 0; i < nbits; i = i + 1) begin
got_sdo = {got_sdo[W12-2:0], sdo8};
bit_time(din_pat[nbits-1-i]);
end
cs_n = 1; @(negedge clk);
end
endtask
initial begin
tx8 = 8'hA5; tx12 = 12'h5A3;
repeat (3) @(negedge clk); rst_n = 1; @(negedge clk);
din_pat = 12'h0C9; run_frame(W8);
chk("rx8 frame1", rx8, 8'hC9);
chk("sdo8 frame1", got_sdo[W8-1:0], 8'hA5);
chk("done8 pulses after frame1", done8_pulses, 1);
tx8 = 8'h3C; din_pat = 12'h017; run_frame(W8);
chk("rx8 frame2", rx8, 8'h17);
chk("sdo8 frame2", got_sdo[W8-1:0], 8'h3C);
chk("done8 pulses after frame2", done8_pulses, 2);
done12_pulses = 0;
din_pat = 12'hE47; run_frame(W12);
chk("rx12", rx12, 12'hE47);
chk("done12 pulses", done12_pulses, 1);
done12_pulses = 0;
din_pat = 12'h01F; run_frame(5);
chk("done12 pulses on truncated frame", done12_pulses, 0);
din_pat = 12'h123; run_frame(W12);
chk("rx12 after truncation", rx12, 12'h123);
chk("done12 pulses after recovery", done12_pulses, 1);
if (errors == 0)
$display("PASS: width is a parameter; done fires once per complete frame; a truncated frame is discarded and the engine recovers");
else
$display("FAILED with %0d error(s)", errors);
$finish;
end
initial begin #200000; $display("FAIL: watchdog timeout"); $finish; end
endmoduleVHDL needs no width calculation at all. A range-constrained integer states what the counter holds, and synthesis derives the storage from the range — arguably the cleanest of the three expressions of the same hardware.
-- spi_frame_width.vhd — the same width-parameterized engine in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_frame_width is
generic (
WIDTH : positive := 8
);
port (
clk : in std_logic;
rst_n : in std_logic;
cs_n : in std_logic;
launch_stb : in std_logic;
sample_stb : in std_logic;
tx_data : in std_logic_vector(WIDTH - 1 downto 0);
sdi : in std_logic;
sdo : out std_logic;
rx_data : out std_logic_vector(WIDTH - 1 downto 0);
done : out std_logic
);
end entity spi_frame_width;
architecture rtl of spi_frame_width is
-- A range-constrained integer says exactly what the counter holds, so no
-- clog2 is needed at all: synthesis derives the width from the range.
signal bit_cnt : integer range 0 to WIDTH - 1;
signal tx_sh : std_logic_vector(WIDTH - 1 downto 0);
signal rx_sh : std_logic_vector(WIDTH - 1 downto 0);
signal cs_n_q : std_logic;
begin
process (clk, rst_n) is
begin
if rst_n = '0' then
tx_sh <= (others => '0');
rx_sh <= (others => '0');
bit_cnt <= 0;
rx_data <= (others => '0');
done <= '0';
cs_n_q <= '1';
elsif rising_edge(clk) then
cs_n_q <= cs_n;
done <= '0';
if cs_n = '1' then
tx_sh <= tx_data;
bit_cnt <= 0;
else
if cs_n_q = '1' then
tx_sh <= tx_data;
bit_cnt <= 0;
end if;
if sample_stb = '1' then
rx_sh <= rx_sh(WIDTH - 2 downto 0) & sdi;
if bit_cnt = WIDTH - 1 then
bit_cnt <= 0;
rx_data <= rx_sh(WIDTH - 2 downto 0) & sdi;
done <= '1';
else
bit_cnt <= bit_cnt + 1;
end if;
end if;
if launch_stb = '1' then
tx_sh <= tx_sh(WIDTH - 2 downto 0) & '0';
end if;
end if;
end if;
end process;
sdo <= tx_sh(WIDTH - 1);
end architecture rtl;-- spi_frame_width_tb.vhd — the same checks in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_frame_width_tb is
end entity spi_frame_width_tb;
architecture tb of spi_frame_width_tb is
constant W8 : positive := 8;
constant W12 : positive := 12;
signal clk : std_logic := '0';
signal rst_n : std_logic := '0';
signal cs_n : std_logic := '1';
signal launch : std_logic := '0';
signal sample : std_logic := '0';
signal sdi : std_logic := '0';
signal halt : boolean := false;
signal tx8 : std_logic_vector(W8 - 1 downto 0) := x"A5";
signal sdo8 : std_logic;
signal rx8 : std_logic_vector(W8 - 1 downto 0);
signal done8 : std_logic;
signal tx12 : std_logic_vector(W12 - 1 downto 0) := x"5A3";
signal sdo12 : std_logic;
signal rx12 : std_logic_vector(W12 - 1 downto 0);
signal done12 : std_logic;
signal errors : natural := 0;
signal done8_pulses : natural := 0;
signal done12_pulses : natural := 0;
signal got_sdo : std_logic_vector(W12 - 1 downto 0) := (others => '0');
begin
clk <= not clk after 5 ns when not halt else '0';
dut8 : entity work.spi_frame_width
generic map (WIDTH => W8)
port map (clk => clk, rst_n => rst_n, cs_n => cs_n, launch_stb => launch,
sample_stb => sample, tx_data => tx8, sdi => sdi, sdo => sdo8,
rx_data => rx8, done => done8);
dut12 : entity work.spi_frame_width
generic map (WIDTH => W12)
port map (clk => clk, rst_n => rst_n, cs_n => cs_n, launch_stb => launch,
sample_stb => sample, tx_data => tx12, sdi => sdi, sdo => sdo12,
rx_data => rx12, done => done12);
-- Count done pulses so "exactly one per frame" is genuinely checked.
counter : process (clk) is
begin
if rising_edge(clk) and rst_n = '1' then
if done8 = '1' then
done8_pulses <= done8_pulses + 1;
end if;
if done12 = '1' then
done12_pulses <= done12_pulses + 1;
end if;
end if;
end process;
stim : process is
variable din_pat : std_logic_vector(W12 - 1 downto 0);
variable base12 : natural;
procedure chk (what : string; got : natural; exp : natural) is
begin
if got /= exp then
report "FAIL " & what & ": got " & integer'image(got)
& " exp " & integer'image(exp) severity error;
errors <= errors + 1;
end if;
end procedure;
procedure bit_time (din : std_logic) is
begin
sdi <= din;
wait until falling_edge(clk); sample <= '1';
wait until falling_edge(clk); sample <= '0';
wait until falling_edge(clk); launch <= '1';
wait until falling_edge(clk); launch <= '0';
end procedure;
procedure run_frame (nbits : positive; pat : std_logic_vector) is
begin
got_sdo <= (others => '0');
cs_n <= '0';
wait until falling_edge(clk);
for i in 0 to nbits - 1 loop
got_sdo <= got_sdo(W12 - 2 downto 0) & sdo8;
bit_time(pat(nbits - 1 - i));
end loop;
cs_n <= '1';
wait until falling_edge(clk);
end procedure;
begin
for i in 0 to 2 loop
wait until falling_edge(clk);
end loop;
rst_n <= '1';
wait until falling_edge(clk);
din_pat := x"0C9";
run_frame(W8, din_pat);
chk("rx8 frame1", to_integer(unsigned(rx8)), 16#C9#);
chk("sdo8 frame1", to_integer(unsigned(got_sdo(W8 - 1 downto 0))), 16#A5#);
chk("done8 pulses after frame1", done8_pulses, 1);
tx8 <= x"3C";
din_pat := x"017";
run_frame(W8, din_pat);
chk("rx8 frame2", to_integer(unsigned(rx8)), 16#17#);
chk("sdo8 frame2", to_integer(unsigned(got_sdo(W8 - 1 downto 0))), 16#3C#);
chk("done8 pulses after frame2", done8_pulses, 2);
base12 := done12_pulses;
din_pat := x"E47";
run_frame(W12, din_pat);
chk("rx12", to_integer(unsigned(rx12)), 16#E47#);
chk("done12 pulses", done12_pulses - base12, 1);
base12 := done12_pulses;
din_pat := x"01F";
run_frame(5, din_pat(4 downto 0));
chk("done12 pulses on truncated frame", done12_pulses - base12, 0);
din_pat := x"123";
run_frame(W12, din_pat);
chk("rx12 after truncation", to_integer(unsigned(rx12)), 16#123#);
chk("done12 pulses after recovery", done12_pulses - base12, 1);
wait until falling_edge(clk);
if errors = 0 then
report "PASS: width is a generic; done fires once per complete frame; "
& "a truncated frame is discarded and the engine recovers" severity note;
else
report "FAILED with " & integer'image(errors) & " error(s)" severity error;
end if;
halt <= true;
wait;
end process;
watchdog : process is
begin
wait for 200 us;
if not halt then
report "FAIL: watchdog timeout" severity failure;
end if;
wait;
end process;
end architecture tb;Parity
All three describe the same hardware: identical ports, an active-low asynchronous reset to constants, CS reload with priority over shifting, a counter advancing only on sample_stb, MSB-first output from the top bit, and done as a one-cycle pulse coincident with the final capture. The only differences are how each language expresses the counter's width — $clog2, a constant function, and a range constraint respectively — and all three produce the same number of flip-flops.
6. Why the Counter Is Not on the Wire
Worth stating precisely, because it has a direct verification consequence.
The done strobe and bit_cnt are internal. A device cannot observe either. Two masters — one counting 0 to 7, another counting 8 down to 1 — are indistinguishable to any device or analyzer, provided they clock the same number of edges within the same CS assertion.
What is observable is the edge count per frame, and that is the only width-related thing an assertion can check at the bus level:
// The ONLY bus-visible consequence of width. Counts active edges between
// CS assertion and deassertion and compares against the configured width.
//
// ASSUMPTION: `sample_stb` is one pulse per active SPI edge, in the same
// clock domain -- the Chapter 2.3 architecture. A monitor clocked by SCLK
// itself would need a different formulation.
int unsigned edges_seen;
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) edges_seen <= 0;
else if (cs_n) edges_seen <= 0; // frame closed
else if (sample_stb) edges_seen <= edges_seen + 1;
end
// Check at the CLOSING edge of the frame, when the count is final.
property p_frame_edge_count;
@(posedge clk) disable iff (!rst_n)
$rose(cs_n) |-> (edges_seen == cfg_width);
endproperty
a_frame_edge_count : assert property (p_frame_edge_count)
else $error("frame carried %0d edges, configuration says %0d", edges_seen, cfg_width);What it proves. That the master clocked exactly the configured number of edges inside each frame — which catches a truncated frame, an over-clocked frame, and a CS that dropped mid-word.
What it does not prove. That the configured width is the width the device expects. That is a datasheet fact, and no signal-level property can reach it — the same boundary Chapter 4.1 §7 drew. A master and a checker can agree perfectly with each other and both be wrong about the device.
7. Why a Verification Engineer Cares
Width is a small, enumerable configuration axis, which makes it cheap to cover and embarrassing to leave uncovered.
covergroup spi_width_cg @(posedge transfer_done);
// Not just "8 or 16" -- the interesting bins are the ones that break
// byte-oriented assumptions somewhere in the stack.
cp_width : coverpoint cfg.width {
bins byte_sized = {8, 16, 24, 32};
bins sub_byte = {[1:7]};
bins non_byte = {9, 10, 12, 18, 20}; // NOT a whole number of bytes
bins wide = {[33:64]};
}
// A width bug usually hides at the ends of the word, not the middle.
cp_first_bit : coverpoint first_bit_value;
cp_last_bit : coverpoint last_bit_value;
x_width_ends : cross cp_width, cp_first_bit, cp_last_bit;
// Chapter 4.1 §2: CS -- not the counter -- delimits. Cover the case
// where CS closes early, because that is what a width disagreement
// looks like from the device's side.
cp_truncated : coverpoint frame_truncated { bins no = {0}; bins yes = {1}; }
endgroupThe non_byte bin is the one that matters most. A design tested only at 8, 16 and 32 can have a counter that is wrong for every width that is not a multiple of eight and still pass its entire regression — because every test happened to use a width where the bug is invisible. The 9-bit and 12-bit cases are where byte-oriented assumptions buried in a FIFO, a packer or a driver surface.
The cp_truncated bin encodes §2's insight: from the device's side, a width disagreement is a truncated or over-long frame, so a suite that never truncates a frame has never exercised the failure this chapter is about.
8. Why an FPGA or ASIC Engineer Cares
The parameter is free; the data path around it is not. Widening the shift registers costs flip-flops linearly, which is nothing. What costs is everything the word touches upstream: a FIFO whose width must match, a bus interface that may be byte-oriented, and the packing logic between them. A 12-bit word on a 32-bit AXI interface needs an explicit packing decision — and that decision, not the shift register, is where width bugs actually live.
A runtime-configurable width is a different design. WIDTH as a parameter fixes the register sizes at elaboration. A controller that supports 4-to-32 bits at runtime must size everything for the maximum and compare the counter against a register rather than a constant. That is a modest change here and a significant one for timing: the comparator's input is now a configuration register instead of a constant, so the tools can no longer fold it away. It is a genuine trade, not an oversight, and Module 13 makes it deliberately.
The counter is not on the SCLK path. Because the engine is clocked by the system clock and advances on strobes, the counter never appears in SCLK's timing. Widening it does not affect the achievable SPI frequency — which is worth knowing because the instinct that "a wider counter slows the interface down" is wrong for this architecture, and right for one clocked directly by SCLK.
9. Failure Signature — The First Byte Is Right and Everything After It Is Wrong
Symptom. A multi-byte read from a device returns a correct first byte and nonsense afterwards. Re-reading gives the same wrong values, so it is reproducible. Slowing the clock changes nothing.
Why reproducibility matters immediately. Chapter 2.4 established that a margin problem improves as the clock slows and eventually becomes perfect. This does not, so it is a configuration or logic fault, not a timing one. That single observation eliminates the entire electrical hypothesis space before any probe is attached.
Plausible mechanisms.
- A width disagreement: the master sends 8-bit transfers where the device expects 16-bit, so after the first word every subsequent one is misaligned.
- CS toggling between bytes where the device requires a continuous frame — §3's exact case.
- A missing dummy phase, so data is read one phase too early (Chapter 4.5).
- An off-by-one bit counter in the master, clocking 7 or 9 edges per word.
The discriminating observations. The first byte being correct is itself strong evidence, and it is worth reading carefully: it means the mode is right — polarity, phase and bit order all produce a correct byte, so none of them is the fault. A mode mismatch corrupts the first byte as much as the rest. This symptom points specifically at something that goes wrong after the first word, which is alignment.
Then count edges per CS assertion on a timing view, not a byte view. That measurement separates all four candidates at once: it directly reveals a width disagreement, an off-by-one counter, and CS toggling, and its absence points at the dummy phase.
Why the investigation goes wrong. Because "the first byte is right" is read as reassurance rather than as evidence. It is the most informative observation available and it is routinely skipped past on the way to re-checking the mode — which the first correct byte has already exonerated.
10. Common Misconceptions
11. Reason It Through
Work this before reading the answer.
A 12-bit ADC is read by an SPI controller whose data register is 8 bits wide. The engineer configures two 8-bit transfers per sample and concatenates the results in software, discarding the low four bits of the second byte. Samples come back looking plausible — the right general magnitude, changing sensibly with the input — but consistently wrong by a varying amount.
What is happening, and why does the result look almost right?
Start with what "almost right" tells you. A sample that tracks the input but is consistently offset means the bits are arriving and are mostly in the correct positions — so the mode is right and the device is responding. This is an alignment problem, not a communication one. If the mode were wrong the values would not track the input at all.
Now count. Two 8-bit transfers clock 16 edges. The ADC delivers 12 bits. The extra four edges clock something — on most ADCs, either zeros, repeated data, or whatever the shift register holds after the sample is exhausted.
The decisive question is where those four extra bits land, and that is determined by whether CS stayed asserted. If CS dropped between the two 8-bit transfers, the ADC very likely restarted its output on the second frame, so the second byte contains the top bits of the sample again rather than its bottom bits. Concatenating them produces a number that is a distorted function of the real one — which is precisely "tracks the input, consistently wrong."
And if CS stayed asserted? Then the 16 edges form one frame and the 12 data bits arrive first, followed by four trailing bits. In that case the software is wrong in a different way: discarding the low four bits of the second byte is correct only if the sample is left-aligned in the 16-bit field. If the ADC right-aligns its 12 bits, the correct operation is to mask the top four bits of the first byte instead. Same symptom, opposite fix.
So what is the discriminating measurement? Look at CS across the whole sample read. It separates the two cases immediately and requires no knowledge of the ADC's internals. Then read the datasheet's timing diagram specifically for where the 12 bits sit within the clocked field — leading, trailing, or offset by a leading null bit, which is extremely common on ADCs and accounts for a great many "off by a factor of two" bugs.
The general lesson. When a width is not a multiple of eight, a byte-oriented controller forces you to clock more bits than the device produces, and the alignment of the real bits inside that padded field is a datasheet fact you must look up. It cannot be inferred from the data, because a shifted sample still looks like a plausible sample.
12. Understanding Check
13. Summary
SPI has no native word size. Eight bits is an ecosystem convention — byte-oriented software, byte-oriented controllers, byte-oriented command sets — and devices routinely use 9, 10, 12, 16, 24 or 32 bits instead.
Chip select delimits the frame; the bit counter is private bookkeeping. A slave has no counter that agrees with the master's, cannot see the master's, and is not notified when the master considers a word finished. It shifts on every edge and uses whatever its register holds when CS rises. A truncated frame therefore produces a misaligned register rather than an error, and whether the device discards it is a device decision the bus does not govern.
The practical consequence is that two bytes and one 16-bit word are different transfers, distinguished only by whether CS stays asserted between them — identical on MOSI, and invisible in any byte-list view.
In RTL the cost of width is one counter and a done strobe, expressed as $clog2 in SystemVerilog, a constant function in Verilog-2001, and a range-constrained integer in VHDL. At the bus level, edge count per frame is the only width-related thing an assertion can check, and it proves consistency with the configuration, never with the device.
For coverage, the bins that matter are the widths that are not multiples of eight, because byte-oriented bugs are invisible at 8, 16 and 32 — and a frame that is deliberately truncated, because that is what a width disagreement looks like from the device's side.
14. What Comes Next
Chapter 4.3 — Bit Ordering takes the second undefined field. Having settled how many bits a transfer carries, it asks which end of the register goes out first — why MSB-first became usual without becoming required, what changes in the shift engine when the order is configurable, and why a bit-order disagreement produces a corruption signature that is both perfectly deterministic and, on certain data, completely invisible.
Continue learning
Related tutorials
- Related topic
RTL Implementation and Mode Handling
The capstone controller in SystemVerilog, Verilog-2001 and VHDL with all four SPI modes derived from a parity, plus five defects found by running it — three of them because the three languages disagreed.
- Related topic
USB vs SPI
SPI selects a peripheral with a wire routed at layout time and USB with an address the host assigned — so a chip-select contention is invisible to every slave (0 of 11) while a duplicate USB address is detected every time (274 of 274).
- Related topic
Interrupt Latency Requirements
bInterval is a request, not a contract, and it bounds one term of seven. A latency monitor in three HDLs, and the measurement showing 3000 randomised steps could not reach two of its own boundaries.
- Related topic
Real-Time Data Streams
A sensor stream must know exactly how much data is missing, not just that something is. Sequence numbers, the modular arithmetic that survives a wrap, and the half-space limit beyond which a gap cannot be measured at all.
