SPI · Module 20
RTL Implementation and Mode Handling
The capstone controller in SystemVerilog, Verilog-2001 and VHDL with all four SPI modes derived from a parity, plus five defects found by running it — three of them because the three languages disagreed.
Chapter 20.2 chose an architecture. This chapter implements it three times, runs 284 checks against each implementation, and compares the three transcripts line for line.
All four SPI modes come out of one parity expression. The bug that took longest to find was correct in modes 0 and 2 and off by one bit in modes 1 and 3.
1. The Mode Decoder, In Three Languages
Start with the payoff. Everything Chapter 20.2 argued about mode handling reduces to this, and it is the same shape in all three languages:
wire leading = ~edge_i[0];
wire sample_event = cpha_q ? ~leading : leading;
wire launch_event = cpha_q ? leading : ~leading;leading <= not edge_i(0);
sample_event <= (not leading) when cpha_q = '1' else leading;
launch_event <= leading when cpha_q = '1' else (not leading);edge_i is the index of the SCLK transition about to be made. Even indices move the clock away from its parked level; odd indices bring it back. CPOL appears nowhere, because the index already counts from whatever level CPOL parked the clock at.
Two consequences worth naming before the code:
All four modes are verified by exercising two bits. There is no fourth case to forget, because there are no cases — there is a parity and a conditional.
An even transition count makes REQ-MODE-001 free. Each frame makes exactly 2N transitions, so the clock necessarily ends where it began. Nothing re-parks it; the arithmetic does.
2. The Control State Machine
spi_capstone_ctrl control FSM
fsm3. CPHA At The Pins
The same 4-bit word, the same device, the same divider — and the only difference is cfg_cpha. Both waveforms are the exact behaviour of the RTL below, at cfg_div = 0, cfg_lead = 1, cfg_lag = 1, sending 0xD and receiving 0xA.
Mode 0: sample on the rising edge, launch on the falling edge
14 cyclesMode 1: launch on the rising edge, sample on the falling edge
14 cycles4. The Implementation — SystemVerilog
// spi_capstone_ctrl.sv
//
// THE CAPSTONE CONTROLLER. One design, carried from Chapter 20.1's specification
// through to Chapter 20.7's design review. Every later chapter verifies THIS module;
// none of them redesigns it.
//
// WHAT IT IS
//
// A configurable SPI master. Four modes, either bit order, a transfer width chosen
// per request, a programmable clock divider, four chip selects, and programmable
// lead / lag / turnaround times.
//
// THE FOUR ARCHITECTURAL COMMITMENTS
//
// These are decided here, defended in 20.2, constrained in 20.4, and questioned in
// 20.7. They are stated at the top because every line below follows from them.
//
// (1) EVERY REGISTER IN THIS MODULE IS CLOCKED BY `clk`.
// SCLK is an OUTPUT WAVEFORM this module generates, not a clock it uses. No
// internal state is clocked by SCLK, so there is no internal clock-domain
// crossing to synchronise. That is a deliberate architectural choice with a
// cost (20.4 §"what this buys and what it costs") -- it is not an oversight,
// and inventing a synchroniser here would be inventing a problem.
//
// (2) CONFIGURATION IS SAMPLED WHEN A REQUEST IS ACCEPTED.
// The `cfg_*` inputs are live wires that software may change at any time. On
// acceptance they are copied into `*_q` registers, and the frame in flight uses
// ONLY the copies. A write during a frame therefore affects the NEXT frame.
// Every use of configuration below reads a `_q`; if any line read a live
// `cfg_*` while busy, that would be the bug 20.7 injects on purpose.
//
// (3) BIT ORDER IS AN ALIGNMENT PROBLEM, NOT A DATAPATH PROBLEM.
// There is exactly ONE shift direction in this module: left. MOSI is always
// `tx_sr[DATA_W-1]`; MISO always shifts into `rx_sr[0]`. LSB-first is produced
// by REVERSING the word as it is loaded and reversing it back as it is read
// out. A second shift direction would double the datapath state and every
// proof about it; one reversal function at each boundary does not.
//
// (4) TIMING GENERATION IS SEPARATE FROM DATAPATH CONTROL.
// A divider produces `tick`, one pulse per SCLK half-period. The FSM consumes
// ticks and knows nothing about system-clock counting. A mode decoder turns
// tick-driven SCLK transitions into `launch_event` / `sample_event`, and the
// datapath knows nothing about CPOL or CPHA. Three concerns, three places.
//
// THE ONE IDEA WORTH THE WHOLE CHAPTER
//
// CPOL and CPHA are not a four-entry lookup table. Every SCLK transition is either
// LEADING (away from the idle level) or TRAILING (back to it). CPOL decides which
// physical direction "away from idle" is; CPHA decides which of the two edge kinds
// launches and which samples. Given an edge index, `leading` is just its parity --
// and everything else follows:
//
// leading = ~edge_i[0] even transitions leave idle
// sample_event = cpha ? trailing : leading
// launch_event = cpha ? leading : trailing
//
// That is the entire mode-handling logic. CPOL never appears in it, because CPOL
// only decides the LEVEL SCLK idles at -- and the edge index already counts from
// that level.
`timescale 1ns/1ps
module spi_capstone_ctrl #(
// Datapath width. The widest transfer the hardware can carry; `cfg_width`
// selects any width from MIN_WIDTH to DATA_W per request.
parameter int DATA_W = 16,
// Narrowest legal transfer. Requests below this are rejected, not clamped:
// silently transferring a different number of bits than asked for is the
// failure mode REQ-ERR-001 exists to prevent.
parameter int MIN_WIDTH = 4,
parameter int NDEV = 4
) (
input wire clk,
// Asynchronous, active low. Release is assumed already synchronised to `clk`
// by the integrator -- see 20.4, and note that an RTL simulation cannot check
// that assumption.
input wire rst_n,
// ---- live configuration (sampled at request acceptance, commitment 2) ----
input wire cfg_cpol,
input wire cfg_cpha,
input wire cfg_lsb_first,
input wire [4:0] cfg_width, // MIN_WIDTH .. DATA_W
input wire [7:0] cfg_div, // SCLK half-period = cfg_div + 1 clk
input wire [1:0] cfg_dev, // which chip select
input wire [3:0] cfg_lead, // CS-low to first edge, in half-periods
input wire [3:0] cfg_lag, // last edge to CS-high, in half-periods
input wire [3:0] cfg_idle, // enforced turnaround, in half-periods
// ---- request ----
input wire start, // one-cycle pulse; ignored unless idle
input wire [DATA_W-1:0] tx_data,
input wire abort, // one-cycle pulse; abandons the frame
// ---- response ----
output wire busy,
output reg done, // exactly one cycle per COMPLETED frame
output reg cfg_err, // one cycle; request rejected, not run
output reg [DATA_W-1:0] rx_data, // right-aligned, valid from `done`
output reg [4:0] bits_done, // bits sampled; survives an abort
// ---- pins ----
output reg sclk,
output reg mosi,
output reg [NDEV-1:0] cs_n,
input wire miso
);
// ---------------------------------------------------------------------------
// Control states.
//
// S_GAP is part of `busy` ON PURPOSE. The turnaround of REQ-TIM-004 is not a
// soft timing hope the integrator has to honour -- it is enforced by refusing
// the next request until it has elapsed. A controller that reported itself idle
// during the turnaround would push that obligation onto software, which is
// exactly where such requirements get lost.
// ---------------------------------------------------------------------------
localparam [2:0] S_IDLE = 3'd0,
S_LEAD = 3'd1, // CS asserted, counting lead half-periods
S_XFER = 3'd2, // generating 2*N SCLK transitions
S_LAG = 3'd3, // SCLK back at idle, CS still asserted
S_GAP = 3'd4; // CS released, enforcing turnaround
reg [2:0] st;
assign busy = (st != S_IDLE);
// ---- captured configuration (commitment 2) --------------------------------
reg cpol_q, cpha_q, lsb_q;
reg [4:0] width_q;
reg [7:0] div_q;
reg [1:0] dev_q;
reg [3:0] lead_q, lag_q, idle_q;
// ---- timing generator (commitment 4) --------------------------------------
reg [7:0] div_cnt;
wire tick = (div_cnt == 8'd0);
// ---- transfer state -------------------------------------------------------
reg [5:0] edge_i; // 0 .. 2*width_q-1, counts SCLK transitions
reg [4:0] tx_idx; // how many bits have been PRESENTED on MOSI
reg [3:0] phase_cnt; // shared lead / lag / gap half-period counter
reg [DATA_W-1:0] tx_sr, rx_sr;
reg aborted;
// ---- mode decode (the one idea, commitment 4) -----------------------------
// `edge_i` is the index of the transition ABOUT to be made. Even indices move
// SCLK away from its idle level (leading); odd indices return it (trailing).
// 2*N transitions per frame, computed in SIX bits. `width_q << 1` would be a
// FIVE-bit expression -- the shift does not widen its operand -- so a 16-bit
// transfer would compute 32 truncated to 0 and the frame would end on its first
// edge. Widen first, then shift: the concatenation is the fix, not the cast.
wire [5:0] edge_total = {1'b0, width_q} << 1;
wire leading = ~edge_i[0];
wire sample_event = cpha_q ? ~leading : leading;
wire launch_event = cpha_q ? leading : ~leading;
// Reject a width the datapath cannot honestly carry. Checked on the LIVE inputs,
// because a request is accepted or refused before anything is captured.
wire cfg_bad = (cfg_width < MIN_WIDTH) || (cfg_width > DATA_W);
// ---- bit reversal: the whole of bit-order handling (commitment 3) ----------
function automatic [DATA_W-1:0] rev;
input [DATA_W-1:0] v;
integer b;
begin
rev = {DATA_W{1'b0}};
for (b = 0; b < DATA_W; b = b + 1)
rev[DATA_W-1-b] = v[b];
end
endfunction
// The word as the shift register wants it: MSB-first sends bit width-1 first, so
// left-align it; LSB-first sends bit 0 first, so reverse it (which puts bit 0 at
// the top and makes the SAME left shift produce the opposite order).
function automatic [DATA_W-1:0] load_align;
input [DATA_W-1:0] v;
input [4:0] w;
input lsb;
begin
load_align = lsb ? rev(v) : (v << (DATA_W - w));
end
endfunction
// The received stream, un-aligned. First bit received sits at rx_sr[w-1].
// MSB-first wants it at bit w-1 already; LSB-first wants it at bit 0.
function automatic [DATA_W-1:0] store_align;
input [DATA_W-1:0] v;
input [4:0] w;
input lsb;
reg [DATA_W-1:0] m;
begin
m = (w >= DATA_W) ? {DATA_W{1'b1}}
: ((({{(DATA_W-1){1'b0}}, 1'b1}) << w) - 1'b1);
store_align = lsb ? ((rev(v) >> (DATA_W - w)) & m) : (v & m);
end
endfunction
// Scratch value for the acceptance cycle. It is a VARIABLE assigned with a
// BLOCKING assignment inside the clocked block below, and both of those choices
// are load-bearing:
//
// * A part-select of a function CALL does not parse in Verilog or
// SystemVerilog -- `load_align(...)[DATA_W-1]` is a syntax error -- so the
// aligned word has to have a name before a bit of it can be taken.
//
// * The obvious name is a wire: `wire [15:0] ld = load_align(...)`. That
// compiles, reads correctly from a testbench, and IS WRONG HERE. A
// continuous assignment whose right-hand side is a function call settles a
// time step after its inputs change, so a clock edge in that same step
// samples the PREVIOUS value. Measured: the shift register loaded 0000 while
// the wire read d000 one cycle later, and every transmitted bit was zero
// while every other captured field was correct.
//
// Computing it in the clocked block removes the question: the function is called
// at the edge, with the values the edge itself sampled. Note that this is the
// opposite of Module 19's VHDL lesson -- a variable is wrong for state another
// process reads, and right for a value used and discarded inside one invocation.
reg [DATA_W-1:0] ld;
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
// REQ-RST-001. Pins first: every chip select released, SCLK at a DEFINED
// level. Note that this level is 0, not cfg_cpol -- reset does not read
// configuration. From the first clock after release SCLK follows CPOL
// (see the S_IDLE arm), and 20.7 asks what a slave sees in between.
st <= S_IDLE;
cs_n <= {NDEV{1'b1}};
sclk <= 1'b0;
mosi <= 1'b0;
done <= 1'b0;
cfg_err <= 1'b0;
rx_data <= {DATA_W{1'b0}};
bits_done <= 5'd0;
div_cnt <= 8'd0;
edge_i <= 6'd0;
tx_idx <= 5'd0;
phase_cnt <= 4'd0;
tx_sr <= {DATA_W{1'b0}};
rx_sr <= {DATA_W{1'b0}};
aborted <= 1'b0;
cpol_q <= 1'b0; cpha_q <= 1'b0; lsb_q <= 1'b0;
width_q <= 5'd0; div_q <= 8'd0; dev_q <= 2'd0;
lead_q <= 4'd0; lag_q <= 4'd0; idle_q <= 4'd0;
end else begin
// Single-cycle outputs default low; the arms below re-raise them.
done <= 1'b0;
cfg_err <= 1'b0;
// ---- divider: runs only while a frame is in progress ---------------
if (st == S_IDLE) div_cnt <= 8'd0;
else if (tick) div_cnt <= div_q;
else div_cnt <= div_cnt - 8'd1;
// ---- abort (REQ-ABT-001) ------------------------------------------
// Honoured from any active state. SCLK returns to the captured idle
// level and the frame proceeds to S_LAG so the selected device still
// gets its release time -- dropping CS instantly would violate the
// device's own hold requirement to punish the controller's caller.
if (abort && (st == S_LEAD || st == S_XFER)) begin
aborted <= 1'b1;
sclk <= cpol_q;
st <= S_LAG;
phase_cnt <= lag_q;
end else begin
case (st)
S_IDLE: begin
// Idle SCLK tracks LIVE cpol, so the pin is correct before any
// frame starts. This is the one place a live cfg_* is read, and
// it is safe precisely because no frame is in flight.
sclk <= cfg_cpol;
if (start) begin
if (cfg_bad) begin
// REQ-ERR-001: refuse, report, stay idle. No frame, no
// `done`, no chip select, nothing captured.
cfg_err <= 1'b1;
end else begin
// COMPUTED ONLY AFTER THE WIDTH IS KNOWN LEGAL. Here that
// ordering looks like tidiness; in VHDL the same code
// shifted by DATA_W - 17 = -1 and aborted the simulation,
// because `shift_left` takes a NATURAL. Verilog wrapped
// the shift amount, produced a garbage word, discarded it
// with the refused request, and said nothing. The bug was
// in both: a function evaluated outside its domain.
ld = load_align(tx_data, cfg_width, cfg_lsb_first);
cpol_q <= cfg_cpol; cpha_q <= cfg_cpha;
lsb_q <= cfg_lsb_first;
width_q <= cfg_width; div_q <= cfg_div;
dev_q <= cfg_dev;
lead_q <= cfg_lead; lag_q <= cfg_lag;
idle_q <= cfg_idle;
cs_n <= {NDEV{1'b1}};
cs_n[cfg_dev] <= 1'b0;
sclk <= cfg_cpol;
rx_sr <= {DATA_W{1'b0}};
// CPHA=0 needs bit 0 valid BEFORE the first leading edge,
// so present it now, with CS. CPHA=1 launches on that
// edge instead, so drive a defined 0 until it does --
// which is why MOSI visibly differs between the two modes
// during the lead time.
// PRESENT THEN SHIFT, uniformly. CPHA=0 owes the
// device a valid bit before its first sampling edge, so
// the pre-launch happens here and consumes bit 0; CPHA=1
// launches on that edge instead, so MOSI holds a defined
// 0 and bit 0 is still pending.
mosi <= cfg_cpha ? 1'b0 : ld[DATA_W-1];
tx_sr <= cfg_cpha ? ld : (ld << 1);
tx_idx <= cfg_cpha ? 5'd0 : 5'd1;
edge_i <= 6'd0;
bits_done <= 5'd0;
aborted <= 1'b0;
div_cnt <= cfg_div;
phase_cnt <= cfg_lead;
st <= S_LEAD;
end
end
end
S_LEAD: begin
if (tick) begin
if (phase_cnt == 4'd0) st <= S_XFER;
else phase_cnt <= phase_cnt - 4'd1;
end
end
S_XFER: begin
if (tick) begin
sclk <= ~sclk;
// Sample BEFORE the launch below, so that when a single
// transition both samples one bit and launches the next
// (it never does in a legal mode, but a mutation in 20.7
// makes it happen) the order is defined rather than lucky.
if (sample_event && (bits_done < width_q)) begin
rx_sr <= {rx_sr[DATA_W-2:0], miso};
bits_done <= bits_done + 5'd1;
end
// The SAME two lines as the pre-launch above: present the
// top bit, then shift it away. The alternative -- shift first
// and present the new top -- is correct for CPHA=0 and off by
// one bit for CPHA=1, because CPHA=1 has no pre-launch to
// have consumed bit 0. Measured: modes 0 and 2 passed at all
// four widths while modes 1 and 3 returned every word shifted
// up one position, in both directions at once.
if (launch_event && (tx_idx < width_q)) begin
mosi <= tx_sr[DATA_W-1];
tx_sr <= {tx_sr[DATA_W-2:0], 1'b0};
tx_idx <= tx_idx + 5'd1;
end
// 2*width_q transitions per frame: N leading + N trailing.
// An even count is why SCLK is guaranteed back at the idle
// level when the frame ends (REQ-MODE-001) -- it is a
// property of the count, not a separate assignment.
if (edge_i == edge_total - 6'd1) begin
st <= S_LAG;
phase_cnt <= lag_q;
end else begin
edge_i <= edge_i + 6'd1;
end
end
end
S_LAG: begin
if (tick) begin
if (phase_cnt == 4'd0) begin
cs_n <= {NDEV{1'b1}};
st <= S_GAP;
phase_cnt <= idle_q;
// REQ-FUNC-005 / 006: a completed frame publishes its
// data and pulses `done` exactly once. An aborted frame
// does NEITHER -- rx_data keeps its previous value, so
// a caller that ignores `done` reads stale data rather
// than a plausible-looking partial word.
if (!aborted) begin
rx_data <= store_align(rx_sr, width_q, lsb_q);
done <= 1'b1;
end
end else begin
phase_cnt <= phase_cnt - 4'd1;
end
end
end
S_GAP: begin
if (tick) begin
if (phase_cnt == 4'd0) st <= S_IDLE;
else phase_cnt <= phase_cnt - 4'd1;
end
end
default: st <= S_IDLE;
endcase
end
end
end
endmodule5. The Implementation — Verilog-2001
The conversion required six line changes, all of them keyword substitutions, and none in the testbench. It was not accepted on that basis: it was compiled with -g2001 and simulated independently, because a mechanical conversion that compiles is not evidence that it behaves the same.
parameter int → parameter integer 3 occurrences
function automatic → function 3 occurrences// spi_capstone_ctrl.sv
//
// THE CAPSTONE CONTROLLER. One design, carried from Chapter 20.1's specification
// through to Chapter 20.7's design review. Every later chapter verifies THIS module;
// none of them redesigns it.
//
// WHAT IT IS
//
// A configurable SPI master. Four modes, either bit order, a transfer width chosen
// per request, a programmable clock divider, four chip selects, and programmable
// lead / lag / turnaround times.
//
// THE FOUR ARCHITECTURAL COMMITMENTS
//
// These are decided here, defended in 20.2, constrained in 20.4, and questioned in
// 20.7. They are stated at the top because every line below follows from them.
//
// (1) EVERY REGISTER IN THIS MODULE IS CLOCKED BY `clk`.
// SCLK is an OUTPUT WAVEFORM this module generates, not a clock it uses. No
// internal state is clocked by SCLK, so there is no internal clock-domain
// crossing to synchronise. That is a deliberate architectural choice with a
// cost (20.4 §"what this buys and what it costs") -- it is not an oversight,
// and inventing a synchroniser here would be inventing a problem.
//
// (2) CONFIGURATION IS SAMPLED WHEN A REQUEST IS ACCEPTED.
// The `cfg_*` inputs are live wires that software may change at any time. On
// acceptance they are copied into `*_q` registers, and the frame in flight uses
// ONLY the copies. A write during a frame therefore affects the NEXT frame.
// Every use of configuration below reads a `_q`; if any line read a live
// `cfg_*` while busy, that would be the bug 20.7 injects on purpose.
//
// (3) BIT ORDER IS AN ALIGNMENT PROBLEM, NOT A DATAPATH PROBLEM.
// There is exactly ONE shift direction in this module: left. MOSI is always
// `tx_sr[DATA_W-1]`; MISO always shifts into `rx_sr[0]`. LSB-first is produced
// by REVERSING the word as it is loaded and reversing it back as it is read
// out. A second shift direction would double the datapath state and every
// proof about it; one reversal function at each boundary does not.
//
// (4) TIMING GENERATION IS SEPARATE FROM DATAPATH CONTROL.
// A divider produces `tick`, one pulse per SCLK half-period. The FSM consumes
// ticks and knows nothing about system-clock counting. A mode decoder turns
// tick-driven SCLK transitions into `launch_event` / `sample_event`, and the
// datapath knows nothing about CPOL or CPHA. Three concerns, three places.
//
// THE ONE IDEA WORTH THE WHOLE CHAPTER
//
// CPOL and CPHA are not a four-entry lookup table. Every SCLK transition is either
// LEADING (away from the idle level) or TRAILING (back to it). CPOL decides which
// physical direction "away from idle" is; CPHA decides which of the two edge kinds
// launches and which samples. Given an edge index, `leading` is just its parity --
// and everything else follows:
//
// leading = ~edge_i[0] even transitions leave idle
// sample_event = cpha ? trailing : leading
// launch_event = cpha ? leading : trailing
//
// That is the entire mode-handling logic. CPOL never appears in it, because CPOL
// only decides the LEVEL SCLK idles at -- and the edge index already counts from
// that level.
`timescale 1ns/1ps
module spi_capstone_ctrl #(
// Datapath width. The widest transfer the hardware can carry; `cfg_width`
// selects any width from MIN_WIDTH to DATA_W per request.
parameter integer DATA_W = 16,
// Narrowest legal transfer. Requests below this are rejected, not clamped:
// silently transferring a different number of bits than asked for is the
// failure mode REQ-ERR-001 exists to prevent.
parameter integer MIN_WIDTH = 4,
parameter integer NDEV = 4
) (
input wire clk,
// Asynchronous, active low. Release is assumed already synchronised to `clk`
// by the integrator -- see 20.4, and note that an RTL simulation cannot check
// that assumption.
input wire rst_n,
// ---- live configuration (sampled at request acceptance, commitment 2) ----
input wire cfg_cpol,
input wire cfg_cpha,
input wire cfg_lsb_first,
input wire [4:0] cfg_width, // MIN_WIDTH .. DATA_W
input wire [7:0] cfg_div, // SCLK half-period = cfg_div + 1 clk
input wire [1:0] cfg_dev, // which chip select
input wire [3:0] cfg_lead, // CS-low to first edge, in half-periods
input wire [3:0] cfg_lag, // last edge to CS-high, in half-periods
input wire [3:0] cfg_idle, // enforced turnaround, in half-periods
// ---- request ----
input wire start, // one-cycle pulse; ignored unless idle
input wire [DATA_W-1:0] tx_data,
input wire abort, // one-cycle pulse; abandons the frame
// ---- response ----
output wire busy,
output reg done, // exactly one cycle per COMPLETED frame
output reg cfg_err, // one cycle; request rejected, not run
output reg [DATA_W-1:0] rx_data, // right-aligned, valid from `done`
output reg [4:0] bits_done, // bits sampled; survives an abort
// ---- pins ----
output reg sclk,
output reg mosi,
output reg [NDEV-1:0] cs_n,
input wire miso
);
// ---------------------------------------------------------------------------
// Control states.
//
// S_GAP is part of `busy` ON PURPOSE. The turnaround of REQ-TIM-004 is not a
// soft timing hope the integrator has to honour -- it is enforced by refusing
// the next request until it has elapsed. A controller that reported itself idle
// during the turnaround would push that obligation onto software, which is
// exactly where such requirements get lost.
// ---------------------------------------------------------------------------
localparam [2:0] S_IDLE = 3'd0,
S_LEAD = 3'd1, // CS asserted, counting lead half-periods
S_XFER = 3'd2, // generating 2*N SCLK transitions
S_LAG = 3'd3, // SCLK back at idle, CS still asserted
S_GAP = 3'd4; // CS released, enforcing turnaround
reg [2:0] st;
assign busy = (st != S_IDLE);
// ---- captured configuration (commitment 2) --------------------------------
reg cpol_q, cpha_q, lsb_q;
reg [4:0] width_q;
reg [7:0] div_q;
reg [1:0] dev_q;
reg [3:0] lead_q, lag_q, idle_q;
// ---- timing generator (commitment 4) --------------------------------------
reg [7:0] div_cnt;
wire tick = (div_cnt == 8'd0);
// ---- transfer state -------------------------------------------------------
reg [5:0] edge_i; // 0 .. 2*width_q-1, counts SCLK transitions
reg [4:0] tx_idx; // how many bits have been PRESENTED on MOSI
reg [3:0] phase_cnt; // shared lead / lag / gap half-period counter
reg [DATA_W-1:0] tx_sr, rx_sr;
reg aborted;
// ---- mode decode (the one idea, commitment 4) -----------------------------
// `edge_i` is the index of the transition ABOUT to be made. Even indices move
// SCLK away from its idle level (leading); odd indices return it (trailing).
// 2*N transitions per frame, computed in SIX bits. `width_q << 1` would be a
// FIVE-bit expression -- the shift does not widen its operand -- so a 16-bit
// transfer would compute 32 truncated to 0 and the frame would end on its first
// edge. Widen first, then shift: the concatenation is the fix, not the cast.
wire [5:0] edge_total = {1'b0, width_q} << 1;
wire leading = ~edge_i[0];
wire sample_event = cpha_q ? ~leading : leading;
wire launch_event = cpha_q ? leading : ~leading;
// Reject a width the datapath cannot honestly carry. Checked on the LIVE inputs,
// because a request is accepted or refused before anything is captured.
wire cfg_bad = (cfg_width < MIN_WIDTH) || (cfg_width > DATA_W);
// ---- bit reversal: the whole of bit-order handling (commitment 3) ----------
function [DATA_W-1:0] rev;
input [DATA_W-1:0] v;
integer b;
begin
rev = {DATA_W{1'b0}};
for (b = 0; b < DATA_W; b = b + 1)
rev[DATA_W-1-b] = v[b];
end
endfunction
// The word as the shift register wants it: MSB-first sends bit width-1 first, so
// left-align it; LSB-first sends bit 0 first, so reverse it (which puts bit 0 at
// the top and makes the SAME left shift produce the opposite order).
function [DATA_W-1:0] load_align;
input [DATA_W-1:0] v;
input [4:0] w;
input lsb;
begin
load_align = lsb ? rev(v) : (v << (DATA_W - w));
end
endfunction
// The received stream, un-aligned. First bit received sits at rx_sr[w-1].
// MSB-first wants it at bit w-1 already; LSB-first wants it at bit 0.
function [DATA_W-1:0] store_align;
input [DATA_W-1:0] v;
input [4:0] w;
input lsb;
reg [DATA_W-1:0] m;
begin
m = (w >= DATA_W) ? {DATA_W{1'b1}}
: ((({{(DATA_W-1){1'b0}}, 1'b1}) << w) - 1'b1);
store_align = lsb ? ((rev(v) >> (DATA_W - w)) & m) : (v & m);
end
endfunction
// Scratch value for the acceptance cycle. It is a VARIABLE assigned with a
// BLOCKING assignment inside the clocked block below, and both of those choices
// are load-bearing:
//
// * A part-select of a function CALL does not parse in Verilog or
// SystemVerilog -- `load_align(...)[DATA_W-1]` is a syntax error -- so the
// aligned word has to have a name before a bit of it can be taken.
//
// * The obvious name is a wire: `wire [15:0] ld = load_align(...)`. That
// compiles, reads correctly from a testbench, and IS WRONG HERE. A
// continuous assignment whose right-hand side is a function call settles a
// time step after its inputs change, so a clock edge in that same step
// samples the PREVIOUS value. Measured: the shift register loaded 0000 while
// the wire read d000 one cycle later, and every transmitted bit was zero
// while every other captured field was correct.
//
// Computing it in the clocked block removes the question: the function is called
// at the edge, with the values the edge itself sampled. Note that this is the
// opposite of Module 19's VHDL lesson -- a variable is wrong for state another
// process reads, and right for a value used and discarded inside one invocation.
reg [DATA_W-1:0] ld;
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
// REQ-RST-001. Pins first: every chip select released, SCLK at a DEFINED
// level. Note that this level is 0, not cfg_cpol -- reset does not read
// configuration. From the first clock after release SCLK follows CPOL
// (see the S_IDLE arm), and 20.7 asks what a slave sees in between.
st <= S_IDLE;
cs_n <= {NDEV{1'b1}};
sclk <= 1'b0;
mosi <= 1'b0;
done <= 1'b0;
cfg_err <= 1'b0;
rx_data <= {DATA_W{1'b0}};
bits_done <= 5'd0;
div_cnt <= 8'd0;
edge_i <= 6'd0;
tx_idx <= 5'd0;
phase_cnt <= 4'd0;
tx_sr <= {DATA_W{1'b0}};
rx_sr <= {DATA_W{1'b0}};
aborted <= 1'b0;
cpol_q <= 1'b0; cpha_q <= 1'b0; lsb_q <= 1'b0;
width_q <= 5'd0; div_q <= 8'd0; dev_q <= 2'd0;
lead_q <= 4'd0; lag_q <= 4'd0; idle_q <= 4'd0;
end else begin
// Single-cycle outputs default low; the arms below re-raise them.
done <= 1'b0;
cfg_err <= 1'b0;
// ---- divider: runs only while a frame is in progress ---------------
if (st == S_IDLE) div_cnt <= 8'd0;
else if (tick) div_cnt <= div_q;
else div_cnt <= div_cnt - 8'd1;
// ---- abort (REQ-ABT-001) ------------------------------------------
// Honoured from any active state. SCLK returns to the captured idle
// level and the frame proceeds to S_LAG so the selected device still
// gets its release time -- dropping CS instantly would violate the
// device's own hold requirement to punish the controller's caller.
if (abort && (st == S_LEAD || st == S_XFER)) begin
aborted <= 1'b1;
sclk <= cpol_q;
st <= S_LAG;
phase_cnt <= lag_q;
end else begin
case (st)
S_IDLE: begin
// Idle SCLK tracks LIVE cpol, so the pin is correct before any
// frame starts. This is the one place a live cfg_* is read, and
// it is safe precisely because no frame is in flight.
sclk <= cfg_cpol;
if (start) begin
if (cfg_bad) begin
// REQ-ERR-001: refuse, report, stay idle. No frame, no
// `done`, no chip select, nothing captured.
cfg_err <= 1'b1;
end else begin
// COMPUTED ONLY AFTER THE WIDTH IS KNOWN LEGAL. Here that
// ordering looks like tidiness; in VHDL the same code
// shifted by DATA_W - 17 = -1 and aborted the simulation,
// because `shift_left` takes a NATURAL. Verilog wrapped
// the shift amount, produced a garbage word, discarded it
// with the refused request, and said nothing. The bug was
// in both: a function evaluated outside its domain.
ld = load_align(tx_data, cfg_width, cfg_lsb_first);
cpol_q <= cfg_cpol; cpha_q <= cfg_cpha;
lsb_q <= cfg_lsb_first;
width_q <= cfg_width; div_q <= cfg_div;
dev_q <= cfg_dev;
lead_q <= cfg_lead; lag_q <= cfg_lag;
idle_q <= cfg_idle;
cs_n <= {NDEV{1'b1}};
cs_n[cfg_dev] <= 1'b0;
sclk <= cfg_cpol;
rx_sr <= {DATA_W{1'b0}};
// CPHA=0 needs bit 0 valid BEFORE the first leading edge,
// so present it now, with CS. CPHA=1 launches on that
// edge instead, so drive a defined 0 until it does --
// which is why MOSI visibly differs between the two modes
// during the lead time.
// PRESENT THEN SHIFT, uniformly. CPHA=0 owes the
// device a valid bit before its first sampling edge, so
// the pre-launch happens here and consumes bit 0; CPHA=1
// launches on that edge instead, so MOSI holds a defined
// 0 and bit 0 is still pending.
mosi <= cfg_cpha ? 1'b0 : ld[DATA_W-1];
tx_sr <= cfg_cpha ? ld : (ld << 1);
tx_idx <= cfg_cpha ? 5'd0 : 5'd1;
edge_i <= 6'd0;
bits_done <= 5'd0;
aborted <= 1'b0;
div_cnt <= cfg_div;
phase_cnt <= cfg_lead;
st <= S_LEAD;
end
end
end
S_LEAD: begin
if (tick) begin
if (phase_cnt == 4'd0) st <= S_XFER;
else phase_cnt <= phase_cnt - 4'd1;
end
end
S_XFER: begin
if (tick) begin
sclk <= ~sclk;
// Sample BEFORE the launch below, so that when a single
// transition both samples one bit and launches the next
// (it never does in a legal mode, but a mutation in 20.7
// makes it happen) the order is defined rather than lucky.
if (sample_event && (bits_done < width_q)) begin
rx_sr <= {rx_sr[DATA_W-2:0], miso};
bits_done <= bits_done + 5'd1;
end
// The SAME two lines as the pre-launch above: present the
// top bit, then shift it away. The alternative -- shift first
// and present the new top -- is correct for CPHA=0 and off by
// one bit for CPHA=1, because CPHA=1 has no pre-launch to
// have consumed bit 0. Measured: modes 0 and 2 passed at all
// four widths while modes 1 and 3 returned every word shifted
// up one position, in both directions at once.
if (launch_event && (tx_idx < width_q)) begin
mosi <= tx_sr[DATA_W-1];
tx_sr <= {tx_sr[DATA_W-2:0], 1'b0};
tx_idx <= tx_idx + 5'd1;
end
// 2*width_q transitions per frame: N leading + N trailing.
// An even count is why SCLK is guaranteed back at the idle
// level when the frame ends (REQ-MODE-001) -- it is a
// property of the count, not a separate assignment.
if (edge_i == edge_total - 6'd1) begin
st <= S_LAG;
phase_cnt <= lag_q;
end else begin
edge_i <= edge_i + 6'd1;
end
end
end
S_LAG: begin
if (tick) begin
if (phase_cnt == 4'd0) begin
cs_n <= {NDEV{1'b1}};
st <= S_GAP;
phase_cnt <= idle_q;
// REQ-FUNC-005 / 006: a completed frame publishes its
// data and pulses `done` exactly once. An aborted frame
// does NEITHER -- rx_data keeps its previous value, so
// a caller that ignores `done` reads stale data rather
// than a plausible-looking partial word.
if (!aborted) begin
rx_data <= store_align(rx_sr, width_q, lsb_q);
done <= 1'b1;
end
end else begin
phase_cnt <= phase_cnt - 4'd1;
end
end
end
S_GAP: begin
if (tick) begin
if (phase_cnt == 4'd0) st <= S_IDLE;
else phase_cnt <= phase_cnt - 4'd1;
end
end
default: st <= S_IDLE;
endcase
end
end
end
endmoduleEvery .v file in this module was scanned for logic, always_ff, always_comb, typedef, enum, interface, class, assert property, covergroup, $urandom and $countones. Excluding prose inside comments, the count of such constructs in executable code is zero across all four files.
6. The Implementation — VHDL
-- spi_capstone_ctrl.vhd
--
-- The capstone controller in VHDL-2008. Same specification, same architecture, same
-- four commitments as the SystemVerilog and Verilog-2001 versions -- and the same
-- cycle-level behaviour, which is checked by comparing all three transcripts rather
-- than by reading the three files side by side.
--
-- THE VHDL-SPECIFIC DISCIPLINE THIS FILE FOLLOWS
--
-- Module 19 cost four separate defects to one root cause, so the rule is written
-- down here rather than rediscovered:
--
-- A SIGNAL is the VHDL counterpart of a non-blocking register assignment. Its new
-- value is visible only after the process suspends, which is exactly what a
-- flip-flop does, and exactly what `<=` means in the other two languages.
--
-- A VARIABLE is visible to the next statement. It is therefore correct for a value
-- computed and consumed inside ONE invocation, and wrong for anything that
-- represents state. Every register below is a signal. The only variable is `ld`,
-- which exists because the alignment function must be evaluated with the values
-- the edge itself sampled -- the same reason the other two languages compute it
-- inside the clocked block instead of on a wire.
--
-- Also observed and deliberately avoided here:
-- * every design unit carries its own context clause (they do not carry over)
-- * every function formal is CONSTRAINED, so no slice direction is inherited
-- * arithmetic goes through `resize`/`shift_*`, never a bare `&`, so no
-- concatenation is ambiguous
-- * no identifier is a reserved word, and none collides with another under case
-- folding -- VHDL is case-INSENSITIVE, so `S_LEAD` and `s_lead` are one name
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_capstone_ctrl is
generic (
DATA_W : integer := 16;
MIN_WIDTH : integer := 4;
NDEV : integer := 4
);
port (
clk : in std_logic;
rst_n : in std_logic;
cfg_cpol : in std_logic;
cfg_cpha : in std_logic;
cfg_lsb_first : in std_logic;
cfg_width : in std_logic_vector(4 downto 0);
cfg_div : in std_logic_vector(7 downto 0);
cfg_dev : in std_logic_vector(1 downto 0);
cfg_lead : in std_logic_vector(3 downto 0);
cfg_lag : in std_logic_vector(3 downto 0);
cfg_idle : in std_logic_vector(3 downto 0);
start : in std_logic;
tx_data : in std_logic_vector(DATA_W-1 downto 0);
abort : in std_logic;
busy : out std_logic;
done : out std_logic;
cfg_err : out std_logic;
rx_data : out std_logic_vector(DATA_W-1 downto 0);
bits_done : out std_logic_vector(4 downto 0);
sclk : out std_logic;
mosi : out std_logic;
cs_n : out std_logic_vector(NDEV-1 downto 0);
miso : in std_logic
);
end entity spi_capstone_ctrl;
architecture rtl of spi_capstone_ctrl is
type state_t is (S_IDLE, S_LEAD, S_XFER, S_LAG, S_GAP);
signal st : state_t;
-- captured configuration
signal cpol_q, cpha_q, lsb_q : std_logic := '0';
signal width_q : unsigned(4 downto 0) := (others => '0');
signal div_q : unsigned(7 downto 0) := (others => '0');
signal dev_q : unsigned(1 downto 0) := (others => '0');
signal lead_q, lag_q, idle_q : unsigned(3 downto 0) := (others => '0');
-- Timing generator. The declaration initialisers on the registers in this
-- architecture mirror their RESET values exactly. They exist so that the
-- concurrent decode below does not evaluate `div_cnt = 0` against an
-- uninitialised value at time zero, which NUMERIC_STD reports as
-- "metavalue detected, returning FALSE". They are simulation initial values:
-- synthesis ignores them, and the asynchronous reset -- not the initialiser --
-- is what defines this hardware's start-up state.
signal div_cnt : unsigned(7 downto 0) := (others => '0');
signal tick : std_logic;
-- transfer state
signal edge_i : unsigned(5 downto 0) := (others => '0');
signal edge_total : unsigned(5 downto 0);
signal tx_idx : unsigned(4 downto 0) := (others => '0');
signal bits_cnt : unsigned(4 downto 0) := (others => '0');
signal phase_cnt : unsigned(3 downto 0) := (others => '0');
signal tx_sr : std_logic_vector(DATA_W-1 downto 0);
signal rx_sr : std_logic_vector(DATA_W-1 downto 0);
signal aborted : std_logic;
-- mode decode
signal leading : std_logic;
signal sample_event : std_logic;
signal launch_event : std_logic;
signal cfg_bad : std_logic;
-- registered pin / status state. The ports are driven concurrently from these, so
-- every signal below has exactly ONE driver. VHDL's std_logic is a RESOLVED type:
-- a second driver would not be an error, it would quietly produce 'X'.
signal sclk_r : std_logic;
signal mosi_r : std_logic;
signal cs_r : std_logic_vector(NDEV-1 downto 0);
signal done_r : std_logic;
signal err_r : std_logic;
signal rx_data_r : std_logic_vector(DATA_W-1 downto 0);
-- Reverse the whole word. The formal is CONSTRAINED: an unconstrained formal takes
-- its range from the actual, and a concatenation or literal can hand it an
-- ASCENDING range, after which every index in here refers to the wrong end.
function rev (v : std_logic_vector(DATA_W-1 downto 0))
return std_logic_vector is
variable r : std_logic_vector(DATA_W-1 downto 0);
begin
for b in 0 to DATA_W-1 loop
r(DATA_W-1-b) := v(b);
end loop;
return r;
end function rev;
-- Low-w-bit mask, built by iteration rather than by `2**w - 1`. The arithmetic
-- form has to be evaluated in a type wide enough to hold 2**DATA_W, and getting
-- that wrong for w = DATA_W produces an all-zero mask that makes every masked
-- comparison succeed against zero.
function maskw (w : integer) return unsigned is
variable m : unsigned(DATA_W-1 downto 0);
begin
m := (others => '0');
for b in 0 to DATA_W-1 loop
if b < w then
m(b) := '1';
end if;
end loop;
return m;
end function maskw;
function load_align (v : std_logic_vector(DATA_W-1 downto 0);
w : integer; lsb : std_logic)
return std_logic_vector is
begin
if lsb = '1' then
return rev(v);
else
return std_logic_vector(shift_left(unsigned(v), DATA_W - w));
end if;
end function load_align;
function store_align (v : std_logic_vector(DATA_W-1 downto 0);
w : integer; lsb : std_logic)
return std_logic_vector is
begin
if lsb = '1' then
-- The shift already clears the vacated high bits, so no mask is needed.
return std_logic_vector(shift_right(unsigned(rev(v)), DATA_W - w));
else
return std_logic_vector(unsigned(v) and maskw(w));
end if;
end function store_align;
begin
-- Concurrent decode. Each of these reads a signal, so inside the clocked process
-- below they hold their PRE-EDGE values -- which is what makes them behave like
-- the `wire` declarations in the other two languages rather than like variables.
tick <= '1' when div_cnt = 0 else '0';
edge_total <= resize(width_q, 6) sll 1;
leading <= not edge_i(0);
sample_event <= (not leading) when cpha_q = '1' else leading;
launch_event <= leading when cpha_q = '1' else (not leading);
cfg_bad <= '1' when (to_integer(unsigned(cfg_width)) < MIN_WIDTH)
or (to_integer(unsigned(cfg_width)) > DATA_W) else '0';
busy <= '0' when st = S_IDLE else '1';
sclk <= sclk_r;
mosi <= mosi_r;
cs_n <= cs_r;
done <= done_r;
cfg_err <= err_r;
rx_data <= rx_data_r;
bits_done <= std_logic_vector(bits_cnt);
process (clk, rst_n)
-- The ONE variable in this design. It is computed and consumed within a single
-- invocation, which is precisely what a variable is for.
variable ld : std_logic_vector(DATA_W-1 downto 0);
begin
if rst_n = '0' then
st <= S_IDLE;
cs_r <= (others => '1');
sclk_r <= '0';
mosi_r <= '0';
done_r <= '0';
err_r <= '0';
rx_data_r <= (others => '0');
bits_cnt <= (others => '0');
div_cnt <= (others => '0');
edge_i <= (others => '0');
tx_idx <= (others => '0');
phase_cnt <= (others => '0');
tx_sr <= (others => '0');
rx_sr <= (others => '0');
aborted <= '0';
cpol_q <= '0'; cpha_q <= '0'; lsb_q <= '0';
width_q <= (others => '0');
div_q <= (others => '0');
dev_q <= (others => '0');
lead_q <= (others => '0');
lag_q <= (others => '0');
idle_q <= (others => '0');
elsif rising_edge(clk) then
done_r <= '0';
err_r <= '0';
if st = S_IDLE then
div_cnt <= (others => '0');
elsif tick = '1' then
div_cnt <= div_q;
else
div_cnt <= div_cnt - 1;
end if;
if abort = '1' and (st = S_LEAD or st = S_XFER) then
aborted <= '1';
sclk_r <= cpol_q;
st <= S_LAG;
phase_cnt <= lag_q;
else
case st is
when S_IDLE =>
sclk_r <= cfg_cpol;
if start = '1' then
if cfg_bad = '1' then
err_r <= '1';
else
-- COMPUTED ONLY AFTER THE WIDTH IS KNOWN LEGAL, and the
-- third language is why. `load_align` shifts by
-- DATA_W - w; for the rejected width 17 that is -1, and
-- `shift_left` takes a NATURAL:
--
-- ** Fatal: value -1 outside of NATURAL range
--
-- Verilog computes the same expression as an unsigned
-- wrap, produces a garbage word, discards it because the
-- request is refused, and says nothing. Both languages
-- were evaluating a function outside its domain; only one
-- of them admitted it. Validating first is not a VHDL
-- workaround -- it is the fix, in all three.
ld := load_align(tx_data,
to_integer(unsigned(cfg_width)),
cfg_lsb_first);
cpol_q <= cfg_cpol;
cpha_q <= cfg_cpha;
lsb_q <= cfg_lsb_first;
width_q <= unsigned(cfg_width);
div_q <= unsigned(cfg_div);
dev_q <= unsigned(cfg_dev);
lead_q <= unsigned(cfg_lead);
lag_q <= unsigned(cfg_lag);
idle_q <= unsigned(cfg_idle);
cs_r <= (others => '1');
cs_r(to_integer(unsigned(cfg_dev))) <= '0';
sclk_r <= cfg_cpol;
rx_sr <= (others => '0');
if cfg_cpha = '1' then
mosi_r <= '0';
tx_sr <= ld;
tx_idx <= (others => '0');
else
mosi_r <= ld(DATA_W-1);
tx_sr <= std_logic_vector(shift_left(unsigned(ld), 1));
tx_idx <= to_unsigned(1, 5);
end if;
bits_cnt <= (others => '0');
edge_i <= (others => '0');
aborted <= '0';
div_cnt <= unsigned(cfg_div);
phase_cnt <= unsigned(cfg_lead);
st <= S_LEAD;
end if;
end if;
when S_LEAD =>
if tick = '1' then
if phase_cnt = 0 then
st <= S_XFER;
else
phase_cnt <= phase_cnt - 1;
end if;
end if;
when S_XFER =>
if tick = '1' then
sclk_r <= not sclk_r;
if sample_event = '1' and bits_cnt < width_q then
rx_sr <= rx_sr(DATA_W-2 downto 0) & miso;
bits_cnt <= bits_cnt + 1;
end if;
if launch_event = '1' and tx_idx < width_q then
mosi_r <= tx_sr(DATA_W-1);
tx_sr <= tx_sr(DATA_W-2 downto 0) & '0';
tx_idx <= tx_idx + 1;
end if;
if edge_i = edge_total - 1 then
st <= S_LAG;
phase_cnt <= lag_q;
else
edge_i <= edge_i + 1;
end if;
end if;
when S_LAG =>
if tick = '1' then
if phase_cnt = 0 then
cs_r <= (others => '1');
st <= S_GAP;
phase_cnt <= idle_q;
if aborted = '0' then
rx_data_r <= store_align(rx_sr,
to_integer(width_q),
lsb_q);
done_r <= '1';
end if;
else
phase_cnt <= phase_cnt - 1;
end if;
end if;
when S_GAP =>
if tick = '1' then
if phase_cnt = 0 then
st <= S_IDLE;
else
phase_cnt <= phase_cnt - 1;
end if;
end if;
end case;
end if;
end if;
end process;
end architecture rtl;7. The Directed Testbench
284 checks per language, in nine groups. Two things in it are worth reading before the code.
The expected values come from specification arithmetic — a mask and a reversal — and never from asking the controller what it did. And the bench uses a pin-level device model rather than a loopback, for the reason Chapter 20.1 §10 gave: a loopback cannot see a bit-order fault, and Chapter 20.7 proves it by injecting one.
// spi_capstone_ctrl_tb.sv
//
// Chapter 20.3 -- the DIRECTED suite for the capstone controller.
//
// This bench answers one question per group, and every expected value in it is
// derived from Chapter 20.1's SPECIFICATION by arithmetic -- never by asking the
// controller what it did. The distinction matters most for bit order, and there is a
// trap in it worth stating before any code:
//
// A LOOPBACK TEST CANNOT DETECT A BIT-ORDER BUG.
//
// Tie MOSI to MISO and the received word equals the transmitted word in BOTH bit
// orders -- the reversal that goes out is undone coming back. A suite built on
// loopback reports a clean pass with the MSB/LSB logic wired backwards. So this
// bench uses a PIN-LEVEL SLAVE MODEL with an asymmetric return word, and checks two
// independent observations per transfer:
//
// * what the master received (depends on bit order)
// * what the slave received (depends on bit order, the other way round)
//
// Chapter 20.7 injects exactly that bug and shows the loopback check surviving it.
//
// WHAT THE SLAVE MODEL IS AND IS NOT
//
// It is the mirror-image DEVICE: it samples on the edge the master launches on, and
// launches on the edge the master samples on. It is NOT the source of the expected
// values -- those come from `mask` and `revw` below, which know nothing about either
// the controller or the model. If the model and the controller were both wrong in
// the same direction the arithmetic would still catch it.
//
// And per 20.4's rule about models: the model is configured for every mode the
// controller supports. A slave model that only works in mode 0 tests mode 0 and
// reports on four.
`timescale 1ns/1ps
module spi_capstone_ctrl_tb;
localparam integer DATA_W = 16;
// ---- DUT interface -------------------------------------------------------
reg clk, rst_n;
reg cfg_cpol, cfg_cpha, cfg_lsb;
reg [4:0] cfg_width;
reg [7:0] cfg_div;
reg [1:0] cfg_dev;
reg [3:0] cfg_lead, cfg_lag, cfg_idle;
reg start, abort;
reg [15:0] tx_data;
wire busy, done, cfg_err;
wire [15:0] rx_data;
wire [4:0] bits_done;
wire sclk, mosi;
wire [3:0] cs_n;
wire miso;
// ---- bookkeeping ---------------------------------------------------------
integer n_chk, n_err, n_neg_ok;
integer cyc;
reg [8*22-1:0] gname;
spi_capstone_ctrl #(.DATA_W(16), .MIN_WIDTH(4), .NDEV(4)) dut (
.clk(clk), .rst_n(rst_n),
.cfg_cpol(cfg_cpol), .cfg_cpha(cfg_cpha), .cfg_lsb_first(cfg_lsb),
.cfg_width(cfg_width), .cfg_div(cfg_div), .cfg_dev(cfg_dev),
.cfg_lead(cfg_lead), .cfg_lag(cfg_lag), .cfg_idle(cfg_idle),
.start(start), .tx_data(tx_data), .abort(abort),
.busy(busy), .done(done), .cfg_err(cfg_err),
.rx_data(rx_data), .bits_done(bits_done),
.sclk(sclk), .mosi(mosi), .cs_n(cs_n), .miso(miso)
);
always #5 clk = ~clk;
// =========================================================================
// Specification arithmetic. The independent oracle.
// =========================================================================
// Low-w-bit mask. Built by shifting a widened 1 -- `(1 << w) - 1` in a 16-bit
// expression gives 0 for w = 16, which would mask every checked value to zero and
// make the whole suite pass vacuously.
function [15:0] mask;
input [4:0] w;
reg [16:0] one;
begin
one = 17'd1;
mask = ((one << w) - 17'd1);
end
endfunction
// Reverse the low w bits of v, leaving the rest zero. This is the ONLY place the
// bench encodes what LSB-first means.
function [15:0] revw;
input [15:0] v;
input [4:0] w;
integer b;
begin
revw = 16'd0;
for (b = 0; b < 16; b = b + 1)
if (b < w) revw[w-1-b] = v[b];
end
endfunction
// What the SLAVE should end up holding, given what the master was asked to send.
function [15:0] exp_slave_rx;
input [15:0] tx; input [4:0] w; input lsb;
begin
exp_slave_rx = lsb ? revw(tx & mask(w), w) : (tx & mask(w));
end
endfunction
// What the MASTER should end up holding, given what the slave returns.
function [15:0] exp_master_rx;
input [15:0] sw; input [4:0] w; input lsb;
begin
exp_master_rx = lsb ? revw(sw & mask(w), w) : (sw & mask(w));
end
endfunction
// =========================================================================
// Pin-level slave model -- configurable for all four modes.
// =========================================================================
reg slv_cpol, slv_cpha;
reg [15:0] slv_word, slv_sr, slv_rx;
reg [4:0] slv_w, slv_idx, slv_nrx;
reg slv_miso, force_x;
reg lead_s;
wire cs_any = ~(&cs_n);
assign miso = force_x ? 1'bx : slv_miso;
// A select fell: load, and for CPHA=0 present the first bit immediately -- the
// master will sample it on the very first edge, so there is no edge left to
// launch it on.
always @(posedge cs_any) begin
slv_sr = slv_word << (16 - slv_w);
slv_rx = 16'd0;
slv_idx = 5'd0;
slv_nrx = 5'd0;
if (!slv_cpha) begin
slv_miso = slv_sr[15];
slv_sr = slv_sr << 1;
slv_idx = 5'd1;
end else begin
slv_miso = 1'b0;
end
end
always @(sclk) begin
if (cs_any === 1'b1) begin
// A transition AWAY from the idle level is leading; back to it, trailing.
// CPOL never appears anywhere else in this model.
lead_s = (sclk !== slv_cpol);
// NOT exchanged: a slave uses the SAME edges as its master. Both
// sample on one edge and both change their output on the other -- that is
// what makes the link work off a single clock. The model's two conditions
// are therefore identical to the controller's, not mirrored.
if (slv_cpha ? !lead_s : lead_s) begin
if (slv_nrx < slv_w) begin
slv_rx = {slv_rx[14:0], mosi};
slv_nrx = slv_nrx + 5'd1;
end
end
if (slv_cpha ? lead_s : !lead_s) begin
if (slv_idx < slv_w) begin
slv_miso = slv_sr[15];
slv_sr = slv_sr << 1;
slv_idx = slv_idx + 5'd1;
end
end
end
end
// =========================================================================
// Pin monitors. Every timing number reported by this bench is MEASURED at the
// pins, never computed from the configuration -- 20.4 depends on that, and
// Module 19.4 is why (hand-derived arithmetic was one cycle out).
// =========================================================================
reg cs_d, sclk_d;
integer t_cs_fall, t_cs_rise, t_first_edge, t_last_edge, t_prev_rise;
integer n_edge, m_half, m_lead, m_lag, m_gap, n_overlap;
integer n_low;
always @(posedge clk) begin
if (!rst_n) begin
cyc <= 0; cs_d <= 1'b0; sclk_d <= 1'b0; n_edge <= 0; n_overlap <= 0;
end else begin
cyc <= cyc + 1;
n_low = (cs_n[0] ? 0 : 1) + (cs_n[1] ? 0 : 1)
+ (cs_n[2] ? 0 : 1) + (cs_n[3] ? 0 : 1);
if (n_low > 1) n_overlap <= n_overlap + 1;
if (cs_any && !cs_d) begin
t_cs_fall <= cyc;
m_gap <= cyc - t_prev_rise;
n_edge <= 0;
end
if (!cs_any && cs_d) begin
t_cs_rise <= cyc;
t_prev_rise <= cyc;
m_lag <= cyc - t_last_edge;
end
// `cs_d` as well as `cs_any`: an edge is only a FRAME edge if a device
// was ALREADY selected last cycle. A transition in the very cycle the
// select falls is SCLK reaching its new idle level, not a clocking
// edge -- the controller parks SCLK and asserts CS together, so when
// the previous idle level differed the two coincide.
//
// Measured cost of omitting `cs_d`: the first transaction after reset
// with CPOL=1 counted 17 edges instead of 16 in VHDL and 16 in
// SystemVerilog, because a one-cycle difference in reset-release
// timing decided whether the re-park landed inside the window. The
// received data was correct in both. With the gate the measurement no
// longer depends on that phase at all.
if (cs_any && cs_d && (sclk !== sclk_d)) begin
if (n_edge == 0) begin
t_first_edge <= cyc;
m_lead <= cyc - t_cs_fall;
end else begin
m_half <= cyc - t_last_edge;
end
t_last_edge <= cyc;
n_edge <= n_edge + 1;
end
cs_d <= cs_any;
sclk_d <= sclk;
end
end
// =========================================================================
// Checkers
// =========================================================================
// `!==` not `!=`: an X anywhere in `got` must FAIL. With `!=` an X compares as
// "unknown", the `if` treats it as false, and the check silently passes -- which
// is the mechanism group 8 deliberately exercises.
task chk16;
input [8*22-1:0] nm; input [15:0] got; input [15:0] exp;
begin
n_chk = n_chk + 1;
if (got !== exp) begin
n_err = n_err + 1;
$display(" FAIL %0s: got %04h expected %04h", nm, got, exp);
end
end
endtask
task chki;
input [8*22-1:0] nm; input integer got; input integer exp;
begin
n_chk = n_chk + 1;
if (got !== exp) begin
n_err = n_err + 1;
$display(" FAIL %0s: got %0d expected %0d", nm, got, exp);
end
end
endtask
// =========================================================================
// Stimulus helpers
// =========================================================================
task set_cfg;
input cpol_i; input cpha_i; input lsb_i; input [4:0] w;
input [7:0] dv; input [1:0] dv_n;
input [3:0] ld; input [3:0] lg; input [3:0] id;
begin
cfg_cpol = cpol_i; cfg_cpha = cpha_i; cfg_lsb = lsb_i;
cfg_width = w; cfg_div = dv; cfg_dev = dv_n;
cfg_lead = ld; cfg_lag = lg; cfg_idle = id;
slv_cpol = cpol_i; slv_cpha = cpha_i; slv_w = w;
end
endtask
// EVERY input is driven on the NEGEDGE, and this is not a style preference.
//
// Driving `start` on the posedge the controller samples it on is a race: the
// assignment that clears it and the controller's clocked block are both in the
// same region of the same time step, and their order is not defined by the
// language. Measured cost of getting this wrong: the FIRST frame of the run was
// accepted and every later one was silently dropped, so 31 of 32 transfers
// compared a stale `rx_data` against a fresh expectation and the suite reported
// a wall of mismatches that had nothing to do with the controller.
//
// A negedge-driven `start` is high across exactly one sampling edge, and nothing
// the bench does can collide with the edge that matters. This is Chapter 16.3's
// clocking-block discipline written out by hand.
task fire;
input [15:0] d;
begin
@(negedge clk);
tx_data = d;
start = 1'b1;
@(negedge clk);
start = 1'b0;
end
endtask
// Bounded wait. Observation happens on the NEGEDGE so every non-blocking update
// from the clock edge has settled -- reading a one-cycle `done` right after
// `@(posedge clk)` reads its pre-edge value and misses it entirely.
task wait_idle;
input integer maxc; output got_done; output integer took;
integer g; reg seen;
begin
g = 0; seen = 1'b0;
while (g < maxc) begin
@(negedge clk);
g = g + 1;
if (done) seen = 1'b1;
if (!busy && seen) g = maxc;
else if (!busy && g > 4) g = maxc;
end
got_done = seen;
took = g;
end
endtask
// =========================================================================
// Test sequence
// =========================================================================
integer mi, oi, wi, k;
reg [4:0] WID [0:3];
reg [15:0] TXP [0:3];
reg [15:0] SWP [0:3];
reg gd;
integer tk, grp_err;
reg [15:0] rx_before;
initial begin
clk = 1'b0; rst_n = 1'b0;
start = 1'b0; abort = 1'b0; tx_data = 16'd0;
cfg_cpol = 1'b0; cfg_cpha = 1'b0; cfg_lsb = 1'b0;
cfg_width = 5'd8; cfg_div = 8'd1; cfg_dev = 2'd0;
cfg_lead = 4'd1; cfg_lag = 4'd1; cfg_idle = 4'd1;
slv_cpol = 1'b0; slv_cpha = 1'b0; slv_w = 5'd8;
slv_word = 16'd0; slv_miso = 1'b0; force_x = 1'b0;
slv_sr = 16'd0; slv_rx = 16'd0; slv_idx = 5'd0; slv_nrx = 5'd0;
n_chk = 0; n_err = 0; n_neg_ok = 0;
cyc = 0; n_edge = 0; n_overlap = 0;
t_cs_fall = 0; t_cs_rise = 0; t_first_edge = 0; t_last_edge = 0;
t_prev_rise = 0; m_half = 0; m_lead = 0; m_lag = 0; m_gap = 0;
cs_d = 1'b0; sclk_d = 1'b0;
WID[0] = 5'd4; WID[1] = 5'd8; WID[2] = 5'd13; WID[3] = 5'd16;
TXP[0] = 16'hB39D; TXP[1] = 16'hB39D; TXP[2] = 16'hB39D; TXP[3] = 16'hB39D;
SWP[0] = 16'h4E7A; SWP[1] = 16'h4E7A; SWP[2] = 16'h4E7A; SWP[3] = 16'h4E7A;
$display("=== Chapter 20.3 -- capstone controller, directed suite ===");
repeat (4) @(posedge clk);
rst_n = 1'b1;
repeat (2) @(posedge clk);
// -----------------------------------------------------------------
// GROUP 1 -- every mode, both bit orders, four widths.
// The question: does the controller implement the specification's mode
// and bit-order semantics, for every width the datapath claims?
// -----------------------------------------------------------------
$display(" G1 mode x bit-order x width");
for (mi = 0; mi < 4; mi = mi + 1) begin
for (oi = 0; oi < 2; oi = oi + 1) begin
grp_err = n_err;
for (wi = 0; wi < 4; wi = wi + 1) begin
set_cfg(mi[1], mi[0], oi[0], WID[wi], 8'd1, 2'd1, 4'd2, 4'd2, 4'd2);
slv_word = SWP[wi] & mask(WID[wi]);
fire(TXP[wi]);
wait_idle(4000, gd, tk);
chki ("done pulsed", gd ? 1 : 0, 1);
chk16("master rx", rx_data,
exp_master_rx(SWP[wi] & mask(WID[wi]), WID[wi], oi[0]));
chk16("slave rx", slv_rx & mask(WID[wi]),
exp_slave_rx(TXP[wi], WID[wi], oi[0]));
chki ("sclk edges", n_edge, 2 * WID[wi]);
chki ("bits sampled", bits_done, WID[wi]);
chki ("sclk parked", (sclk === mi[1]) ? 1 : 0, 1);
chki ("no cs overlap", n_overlap, 0);
end
$display(" mode %0d %0s : rx(w4,w8,w13,w16) checked, %0s",
mi, oi[0] ? "lsb" : "msb",
(n_err == grp_err) ? "all match" : "MISMATCH");
end
end
// -----------------------------------------------------------------
// GROUP 2 -- measured pin intervals.
// The question: are the lead / lag / turnaround requirements MET at the
// pins, and what is the actual margin? The requirement is ">=", the
// measurement is exact, and the difference is reported rather than assumed.
// -----------------------------------------------------------------
$display(" G2 measured pin intervals, in clk cycles");
$display(" div lead lag idle | t_half t_lead t_lag t_gap");
for (k = 0; k < 4; k = k + 1) begin
set_cfg(1'b0, 1'b0, 1'b0, 5'd8,
(k == 0) ? 8'd0 : (k == 1) ? 8'd1 : (k == 2) ? 8'd3 : 8'd1,
2'd2,
(k == 3) ? 4'd5 : 4'd2,
(k == 3) ? 4'd4 : 4'd2,
(k == 3) ? 4'd6 : 4'd3);
slv_word = 16'h5A;
// Two frames: the second one's gap is the interval between them.
fire(16'h33); wait_idle(4000, gd, tk);
fire(16'h33); wait_idle(4000, gd, tk);
$display(" %3d %4d %3d %4d | %6d %6d %5d %5d",
cfg_div, cfg_lead, cfg_lag, cfg_idle, m_half, m_lead, m_lag, m_gap);
// REQ-TIM-001 is an equality and is checked as one.
chki("t_half exact", m_half, cfg_div + 1);
// REQ-TIM-002/003/004 are ">=" requirements, and they are checked as
// stated -- but a ">=" check is WEAK: an implementation that loses one
// half-period of lead still satisfies it whenever cfg_lead > 0. So each
// interval is ALSO checked against the exact closed form the measurements
// above establish, which is what actually catches a one-tick regression.
if (m_lead < (cfg_lead * (cfg_div + 1))) begin
n_err = n_err + 1; $display(" FAIL lead below requirement");
end
n_chk = n_chk + 1;
if (m_lag < (cfg_lag * (cfg_div + 1))) begin
n_err = n_err + 1; $display(" FAIL lag below requirement");
end
n_chk = n_chk + 1;
if (m_gap < (cfg_idle * (cfg_div + 1))) begin
n_err = n_err + 1; $display(" FAIL turnaround below requirement");
end
n_chk = n_chk + 1;
// Exact implementation forms, in system clocks. The "+2" and "+1" terms
// are the ticks the state machine spends LEAVING a phase, and they are
// margin above the requirement, not part of it:
//
// t_lead = (cfg_lead + 2) * t_half one tick to exit S_LEAD, one more
// before the first transition
// t_lag = (cfg_lag + 1) * t_half one tick to exit S_LAG
// t_gap = (cfg_idle + 1) * t_half + 2
//
// The gap is the ONLY interval here with a term that is not a multiple of
// the half-period, and the reason is worth naming: those 2 cycles are
// S_IDLE plus ONE CYCLE OF THE REQUESTER'S OWN LATENCY. With no request
// queue, the bench cannot post the next transfer until `busy` falls, so
// the measured gap is the design's floor PLUS however long the requester
// took to react. 20.2 defends that trade-off and 20.7 questions it.
chki("t_lead exact", m_lead, (cfg_lead + 2) * (cfg_div + 1));
chki("t_lag exact", m_lag, (cfg_lag + 1) * (cfg_div + 1));
chki("t_gap exact", m_gap, (cfg_idle + 1) * (cfg_div + 1) + 2);
end
// -----------------------------------------------------------------
// GROUP 3 -- illegal configuration is REFUSED, at both boundaries.
// -----------------------------------------------------------------
$display(" G3 illegal width");
for (k = 0; k < 4; k = k + 1) begin
set_cfg(1'b0, 1'b0, 1'b0,
(k == 0) ? 5'd3 : (k == 1) ? 5'd17 : (k == 2) ? 5'd4 : 5'd16,
8'd1, 2'd0, 4'd2, 4'd2, 4'd2);
slv_word = 16'hFFFF & mask(cfg_width);
slv_w = (k < 2) ? 5'd8 : cfg_width;
fire(16'h5555);
// `fire` returns on the negedge immediately after the sampling edge, so a
// one-cycle output is readable RIGHT HERE. One more negedge and `cfg_err`
// has already gone low -- which is how a correct rejection reads as a
// missing one.
if (k < 2) begin
chki("cfg_err raised", cfg_err ? 1 : 0, 1);
chki("stayed idle", busy ? 1 : 0, 0);
chki("no cs asserted", cs_any ? 1 : 0, 0);
$display(" width %2d rejected: cfg_err=1 busy=0 cs=idle", cfg_width);
end else begin
chki("accepted", busy ? 1 : 0, 1);
chki("no cfg_err", cfg_err ? 1 : 0, 0);
wait_idle(4000, gd, tk);
chki("done pulsed", gd ? 1 : 0, 1);
$display(" width %2d accepted: cfg_err=0 done=1", cfg_width);
end
end
// -----------------------------------------------------------------
// GROUP 4 -- a request while busy has NO effect (REQ-FUNC-007).
// -----------------------------------------------------------------
$display(" G4 start while busy");
set_cfg(1'b0, 1'b0, 1'b0, 5'd8, 8'd2, 2'd0, 4'd2, 4'd2, 4'd2);
slv_word = 16'h96;
fire(16'hA5);
repeat (6) @(posedge clk);
fire(16'h3C); // ignored: the frame in flight owns the bus
wait_idle(4000, gd, tk);
chk16("rx from FIRST word", rx_data, exp_master_rx(16'h96, 5'd8, 1'b0));
chk16("slave got FIRST word", slv_rx & mask(5'd8), exp_slave_rx(16'hA5, 5'd8, 1'b0));
chki ("edges of one frame", n_edge, 16);
$display(" second request ignored: one frame, 16 edges, slave saw a5");
// -----------------------------------------------------------------
// GROUP 5 -- reset in the middle of a frame (REQ-RST-002).
// -----------------------------------------------------------------
$display(" G5 reset during a frame");
set_cfg(1'b1, 1'b0, 1'b0, 5'd16, 8'd3, 2'd3, 4'd2, 4'd2, 4'd2);
slv_word = 16'hBEEF;
fire(16'hDEAD);
repeat (20) @(posedge clk);
chki("frame really started", cs_any ? 1 : 0, 1);
@(negedge clk); rst_n = 1'b0;
@(negedge clk);
chki("all selects released", cs_any ? 1 : 0, 0);
chki("no done on reset", done ? 1 : 0, 0);
chki("not busy", busy ? 1 : 0, 0);
chki("sclk defined 0", (sclk === 1'b0) ? 1 : 0, 1);
$display(" reset mid-frame: cs released, no done, sclk=0 (not cpol)");
@(negedge clk); rst_n = 1'b1; repeat (3) @(posedge clk);
// -----------------------------------------------------------------
// GROUP 6 -- abort (REQ-ABT-001).
// -----------------------------------------------------------------
$display(" G6 abort mid-frame");
set_cfg(1'b0, 1'b0, 1'b0, 5'd16, 8'd1, 2'd0, 4'd2, 4'd2, 4'd2);
slv_word = 16'h1234;
fire(16'h4321);
wait_idle(4000, gd, tk); // a clean frame first, to set rx_data
rx_before = rx_data;
slv_word = 16'hFFFF;
fire(16'h0F0F);
repeat (24) @(posedge clk);
@(negedge clk); abort = 1'b1; @(negedge clk); abort = 1'b0;
wait_idle(4000, gd, tk);
chki ("no done on abort", gd ? 1 : 0, 0);
chk16("rx_data unchanged", rx_data, rx_before);
chki ("selects released", cs_any ? 1 : 0, 0);
chki ("sclk back at idle", (sclk === 1'b0) ? 1 : 0, 1);
if (bits_done == 5'd0 || bits_done >= 5'd16) begin
n_err = n_err + 1;
$display(" FAIL bits_done not partial: %0d", bits_done);
end
n_chk = n_chk + 1;
$display(" abort: done=0, rx held %04h, bits_done=%0d of 16",
rx_data, bits_done);
// -----------------------------------------------------------------
// GROUP 7 -- configuration captured at acceptance (commitment 2).
// The frame must ignore a mid-flight rewrite of EVERY field it uses.
// -----------------------------------------------------------------
$display(" G7 configuration captured at acceptance");
set_cfg(1'b0, 1'b0, 1'b0, 5'd8, 8'd2, 2'd0, 4'd2, 4'd2, 4'd2);
slv_word = 16'h7E;
fire(16'h81);
repeat (5) @(posedge clk);
// Rewrite everything to values that would produce a different frame.
cfg_cpol = 1'b1; cfg_cpha = 1'b1; cfg_lsb = 1'b1;
cfg_width = 5'd4; cfg_div = 8'd7; cfg_dev = 2'd3;
wait_idle(4000, gd, tk);
chk16("rx used captured cfg", rx_data, exp_master_rx(16'h7E, 5'd8, 1'b0));
chki ("edges of captured width", n_edge, 16);
chki ("captured device kept", cs_n[0] === 1'b1 ? 1 : 0, 1);
$display(" mid-frame rewrite ignored: 16 edges, msb-order rx=%04h", rx_data);
set_cfg(1'b0, 1'b0, 1'b0, 5'd8, 8'd1, 2'd0, 4'd2, 4'd2, 4'd2);
// -----------------------------------------------------------------
// GROUP 8 -- an X on MISO must NOT be able to produce a pass.
// This is a NEGATIVE test: the checker is EXPECTED to fail here, and the
// bench counts that as evidence rather than as a defect.
// -----------------------------------------------------------------
$display(" G8 X on MISO cannot slip through");
slv_word = 16'h5A;
force_x = 1'b1;
fire(16'hA5);
wait_idle(4000, gd, tk);
n_chk = n_chk + 1;
if (rx_data !== exp_master_rx(16'h5A, 5'd8, 1'b0)) begin
n_neg_ok = n_neg_ok + 1;
$display(" X reached rx_data and the !== comparison rejected it");
end else begin
n_err = n_err + 1;
$display(" FAIL an all-X rx_data compared EQUAL -- checker is blind");
end
force_x = 1'b0;
// -----------------------------------------------------------------
// GROUP 9 -- prove the checkers can fail at all.
// A suite that has never printed FAIL has not been shown to be able to.
// -----------------------------------------------------------------
$display(" G9 deliberately wrong expectations must FAIL");
// The pattern here is 9d, and the reason it is not a5 or 5a is worth more
// than the test it enables:
//
// a5 = 10100101 reversed = 10100101
// 5a = 01011010 reversed = 01011010
//
// The two most-used test bytes in digital design are both EIGHT-BIT
// PALINDROMES. Neither can distinguish MSB-first from LSB-first, so a suite
// built on them reports a clean pass with the bit-order logic inverted -- and
// the self-test of the oracle below reported a defect in the ORACLE when the
// only defect was the choice of stimulus. 9d reverses to b9.
k = 0;
slv_word = 16'h9D;
fire(16'h9D);
wait_idle(4000, gd, tk);
n_chk = n_chk + 1;
if (rx_data !== (exp_master_rx(16'h9D, 5'd8, 1'b0) ^ 16'h0001)) k = k + 1;
else begin n_err = n_err + 1; $display(" FAIL rx checker accepted a wrong value"); end
n_chk = n_chk + 1;
if (n_edge !== 15) k = k + 1;
else begin n_err = n_err + 1; $display(" FAIL edge checker accepted a wrong count"); end
n_chk = n_chk + 1;
if (exp_master_rx(16'h9D, 5'd8, 1'b1) !== exp_master_rx(16'h9D, 5'd8, 1'b0)) k = k + 1;
else begin n_err = n_err + 1; $display(" FAIL oracle is blind to bit order"); end
n_neg_ok = n_neg_ok + k;
$display(" %0d of 3 wrong expectations rejected; oracle separates 9d from b9", k);
$display("=== SUMMARY checks=%0d negatives=%0d failures=%0d : %0s ===",
n_chk, n_neg_ok, n_err, (n_err == 0) ? "PASS" : "FAIL");
$finish;
end
// Global watchdog. A hung frame must end the run with a FAILURE, not with a
// simulator that stops printing.
initial begin
#4000000;
$display(" FATAL global timeout -- a frame never completed");
$display("=== SUMMARY checks=%0d negatives=%0d failures=%0d : FAIL ===",
n_chk, n_neg_ok, n_err + 1);
$finish;
end
endmodule// spi_capstone_ctrl_tb.sv
//
// Chapter 20.3 -- the DIRECTED suite for the capstone controller.
//
// This bench answers one question per group, and every expected value in it is
// derived from Chapter 20.1's SPECIFICATION by arithmetic -- never by asking the
// controller what it did. The distinction matters most for bit order, and there is a
// trap in it worth stating before any code:
//
// A LOOPBACK TEST CANNOT DETECT A BIT-ORDER BUG.
//
// Tie MOSI to MISO and the received word equals the transmitted word in BOTH bit
// orders -- the reversal that goes out is undone coming back. A suite built on
// loopback reports a clean pass with the MSB/LSB logic wired backwards. So this
// bench uses a PIN-LEVEL SLAVE MODEL with an asymmetric return word, and checks two
// independent observations per transfer:
//
// * what the master received (depends on bit order)
// * what the slave received (depends on bit order, the other way round)
//
// Chapter 20.7 injects exactly that bug and shows the loopback check surviving it.
//
// WHAT THE SLAVE MODEL IS AND IS NOT
//
// It is the mirror-image DEVICE: it samples on the edge the master launches on, and
// launches on the edge the master samples on. It is NOT the source of the expected
// values -- those come from `mask` and `revw` below, which know nothing about either
// the controller or the model. If the model and the controller were both wrong in
// the same direction the arithmetic would still catch it.
//
// And per 20.4's rule about models: the model is configured for every mode the
// controller supports. A slave model that only works in mode 0 tests mode 0 and
// reports on four.
`timescale 1ns/1ps
module spi_capstone_ctrl_tb;
localparam integer DATA_W = 16;
// ---- DUT interface -------------------------------------------------------
reg clk, rst_n;
reg cfg_cpol, cfg_cpha, cfg_lsb;
reg [4:0] cfg_width;
reg [7:0] cfg_div;
reg [1:0] cfg_dev;
reg [3:0] cfg_lead, cfg_lag, cfg_idle;
reg start, abort;
reg [15:0] tx_data;
wire busy, done, cfg_err;
wire [15:0] rx_data;
wire [4:0] bits_done;
wire sclk, mosi;
wire [3:0] cs_n;
wire miso;
// ---- bookkeeping ---------------------------------------------------------
integer n_chk, n_err, n_neg_ok;
integer cyc;
reg [8*22-1:0] gname;
spi_capstone_ctrl #(.DATA_W(16), .MIN_WIDTH(4), .NDEV(4)) dut (
.clk(clk), .rst_n(rst_n),
.cfg_cpol(cfg_cpol), .cfg_cpha(cfg_cpha), .cfg_lsb_first(cfg_lsb),
.cfg_width(cfg_width), .cfg_div(cfg_div), .cfg_dev(cfg_dev),
.cfg_lead(cfg_lead), .cfg_lag(cfg_lag), .cfg_idle(cfg_idle),
.start(start), .tx_data(tx_data), .abort(abort),
.busy(busy), .done(done), .cfg_err(cfg_err),
.rx_data(rx_data), .bits_done(bits_done),
.sclk(sclk), .mosi(mosi), .cs_n(cs_n), .miso(miso)
);
always #5 clk = ~clk;
// =========================================================================
// Specification arithmetic. The independent oracle.
// =========================================================================
// Low-w-bit mask. Built by shifting a widened 1 -- `(1 << w) - 1` in a 16-bit
// expression gives 0 for w = 16, which would mask every checked value to zero and
// make the whole suite pass vacuously.
function [15:0] mask;
input [4:0] w;
reg [16:0] one;
begin
one = 17'd1;
mask = ((one << w) - 17'd1);
end
endfunction
// Reverse the low w bits of v, leaving the rest zero. This is the ONLY place the
// bench encodes what LSB-first means.
function [15:0] revw;
input [15:0] v;
input [4:0] w;
integer b;
begin
revw = 16'd0;
for (b = 0; b < 16; b = b + 1)
if (b < w) revw[w-1-b] = v[b];
end
endfunction
// What the SLAVE should end up holding, given what the master was asked to send.
function [15:0] exp_slave_rx;
input [15:0] tx; input [4:0] w; input lsb;
begin
exp_slave_rx = lsb ? revw(tx & mask(w), w) : (tx & mask(w));
end
endfunction
// What the MASTER should end up holding, given what the slave returns.
function [15:0] exp_master_rx;
input [15:0] sw; input [4:0] w; input lsb;
begin
exp_master_rx = lsb ? revw(sw & mask(w), w) : (sw & mask(w));
end
endfunction
// =========================================================================
// Pin-level slave model -- configurable for all four modes.
// =========================================================================
reg slv_cpol, slv_cpha;
reg [15:0] slv_word, slv_sr, slv_rx;
reg [4:0] slv_w, slv_idx, slv_nrx;
reg slv_miso, force_x;
reg lead_s;
wire cs_any = ~(&cs_n);
assign miso = force_x ? 1'bx : slv_miso;
// A select fell: load, and for CPHA=0 present the first bit immediately -- the
// master will sample it on the very first edge, so there is no edge left to
// launch it on.
always @(posedge cs_any) begin
slv_sr = slv_word << (16 - slv_w);
slv_rx = 16'd0;
slv_idx = 5'd0;
slv_nrx = 5'd0;
if (!slv_cpha) begin
slv_miso = slv_sr[15];
slv_sr = slv_sr << 1;
slv_idx = 5'd1;
end else begin
slv_miso = 1'b0;
end
end
always @(sclk) begin
if (cs_any === 1'b1) begin
// A transition AWAY from the idle level is leading; back to it, trailing.
// CPOL never appears anywhere else in this model.
lead_s = (sclk !== slv_cpol);
// NOT exchanged: a slave uses the SAME edges as its master. Both
// sample on one edge and both change their output on the other -- that is
// what makes the link work off a single clock. The model's two conditions
// are therefore identical to the controller's, not mirrored.
if (slv_cpha ? !lead_s : lead_s) begin
if (slv_nrx < slv_w) begin
slv_rx = {slv_rx[14:0], mosi};
slv_nrx = slv_nrx + 5'd1;
end
end
if (slv_cpha ? lead_s : !lead_s) begin
if (slv_idx < slv_w) begin
slv_miso = slv_sr[15];
slv_sr = slv_sr << 1;
slv_idx = slv_idx + 5'd1;
end
end
end
end
// =========================================================================
// Pin monitors. Every timing number reported by this bench is MEASURED at the
// pins, never computed from the configuration -- 20.4 depends on that, and
// Module 19.4 is why (hand-derived arithmetic was one cycle out).
// =========================================================================
reg cs_d, sclk_d;
integer t_cs_fall, t_cs_rise, t_first_edge, t_last_edge, t_prev_rise;
integer n_edge, m_half, m_lead, m_lag, m_gap, n_overlap;
integer n_low;
always @(posedge clk) begin
if (!rst_n) begin
cyc <= 0; cs_d <= 1'b0; sclk_d <= 1'b0; n_edge <= 0; n_overlap <= 0;
end else begin
cyc <= cyc + 1;
n_low = (cs_n[0] ? 0 : 1) + (cs_n[1] ? 0 : 1)
+ (cs_n[2] ? 0 : 1) + (cs_n[3] ? 0 : 1);
if (n_low > 1) n_overlap <= n_overlap + 1;
if (cs_any && !cs_d) begin
t_cs_fall <= cyc;
m_gap <= cyc - t_prev_rise;
n_edge <= 0;
end
if (!cs_any && cs_d) begin
t_cs_rise <= cyc;
t_prev_rise <= cyc;
m_lag <= cyc - t_last_edge;
end
// `cs_d` as well as `cs_any`: an edge is only a FRAME edge if a device
// was ALREADY selected last cycle. A transition in the very cycle the
// select falls is SCLK reaching its new idle level, not a clocking
// edge -- the controller parks SCLK and asserts CS together, so when
// the previous idle level differed the two coincide.
//
// Measured cost of omitting `cs_d`: the first transaction after reset
// with CPOL=1 counted 17 edges instead of 16 in VHDL and 16 in
// SystemVerilog, because a one-cycle difference in reset-release
// timing decided whether the re-park landed inside the window. The
// received data was correct in both. With the gate the measurement no
// longer depends on that phase at all.
if (cs_any && cs_d && (sclk !== sclk_d)) begin
if (n_edge == 0) begin
t_first_edge <= cyc;
m_lead <= cyc - t_cs_fall;
end else begin
m_half <= cyc - t_last_edge;
end
t_last_edge <= cyc;
n_edge <= n_edge + 1;
end
cs_d <= cs_any;
sclk_d <= sclk;
end
end
// =========================================================================
// Checkers
// =========================================================================
// `!==` not `!=`: an X anywhere in `got` must FAIL. With `!=` an X compares as
// "unknown", the `if` treats it as false, and the check silently passes -- which
// is the mechanism group 8 deliberately exercises.
task chk16;
input [8*22-1:0] nm; input [15:0] got; input [15:0] exp;
begin
n_chk = n_chk + 1;
if (got !== exp) begin
n_err = n_err + 1;
$display(" FAIL %0s: got %04h expected %04h", nm, got, exp);
end
end
endtask
task chki;
input [8*22-1:0] nm; input integer got; input integer exp;
begin
n_chk = n_chk + 1;
if (got !== exp) begin
n_err = n_err + 1;
$display(" FAIL %0s: got %0d expected %0d", nm, got, exp);
end
end
endtask
// =========================================================================
// Stimulus helpers
// =========================================================================
task set_cfg;
input cpol_i; input cpha_i; input lsb_i; input [4:0] w;
input [7:0] dv; input [1:0] dv_n;
input [3:0] ld; input [3:0] lg; input [3:0] id;
begin
cfg_cpol = cpol_i; cfg_cpha = cpha_i; cfg_lsb = lsb_i;
cfg_width = w; cfg_div = dv; cfg_dev = dv_n;
cfg_lead = ld; cfg_lag = lg; cfg_idle = id;
slv_cpol = cpol_i; slv_cpha = cpha_i; slv_w = w;
end
endtask
// EVERY input is driven on the NEGEDGE, and this is not a style preference.
//
// Driving `start` on the posedge the controller samples it on is a race: the
// assignment that clears it and the controller's clocked block are both in the
// same region of the same time step, and their order is not defined by the
// language. Measured cost of getting this wrong: the FIRST frame of the run was
// accepted and every later one was silently dropped, so 31 of 32 transfers
// compared a stale `rx_data` against a fresh expectation and the suite reported
// a wall of mismatches that had nothing to do with the controller.
//
// A negedge-driven `start` is high across exactly one sampling edge, and nothing
// the bench does can collide with the edge that matters. This is Chapter 16.3's
// clocking-block discipline written out by hand.
task fire;
input [15:0] d;
begin
@(negedge clk);
tx_data = d;
start = 1'b1;
@(negedge clk);
start = 1'b0;
end
endtask
// Bounded wait. Observation happens on the NEGEDGE so every non-blocking update
// from the clock edge has settled -- reading a one-cycle `done` right after
// `@(posedge clk)` reads its pre-edge value and misses it entirely.
task wait_idle;
input integer maxc; output got_done; output integer took;
integer g; reg seen;
begin
g = 0; seen = 1'b0;
while (g < maxc) begin
@(negedge clk);
g = g + 1;
if (done) seen = 1'b1;
if (!busy && seen) g = maxc;
else if (!busy && g > 4) g = maxc;
end
got_done = seen;
took = g;
end
endtask
// =========================================================================
// Test sequence
// =========================================================================
integer mi, oi, wi, k;
reg [4:0] WID [0:3];
reg [15:0] TXP [0:3];
reg [15:0] SWP [0:3];
reg gd;
integer tk, grp_err;
reg [15:0] rx_before;
initial begin
clk = 1'b0; rst_n = 1'b0;
start = 1'b0; abort = 1'b0; tx_data = 16'd0;
cfg_cpol = 1'b0; cfg_cpha = 1'b0; cfg_lsb = 1'b0;
cfg_width = 5'd8; cfg_div = 8'd1; cfg_dev = 2'd0;
cfg_lead = 4'd1; cfg_lag = 4'd1; cfg_idle = 4'd1;
slv_cpol = 1'b0; slv_cpha = 1'b0; slv_w = 5'd8;
slv_word = 16'd0; slv_miso = 1'b0; force_x = 1'b0;
slv_sr = 16'd0; slv_rx = 16'd0; slv_idx = 5'd0; slv_nrx = 5'd0;
n_chk = 0; n_err = 0; n_neg_ok = 0;
cyc = 0; n_edge = 0; n_overlap = 0;
t_cs_fall = 0; t_cs_rise = 0; t_first_edge = 0; t_last_edge = 0;
t_prev_rise = 0; m_half = 0; m_lead = 0; m_lag = 0; m_gap = 0;
cs_d = 1'b0; sclk_d = 1'b0;
WID[0] = 5'd4; WID[1] = 5'd8; WID[2] = 5'd13; WID[3] = 5'd16;
TXP[0] = 16'hB39D; TXP[1] = 16'hB39D; TXP[2] = 16'hB39D; TXP[3] = 16'hB39D;
SWP[0] = 16'h4E7A; SWP[1] = 16'h4E7A; SWP[2] = 16'h4E7A; SWP[3] = 16'h4E7A;
$display("=== Chapter 20.3 -- capstone controller, directed suite ===");
repeat (4) @(posedge clk);
rst_n = 1'b1;
repeat (2) @(posedge clk);
// -----------------------------------------------------------------
// GROUP 1 -- every mode, both bit orders, four widths.
// The question: does the controller implement the specification's mode
// and bit-order semantics, for every width the datapath claims?
// -----------------------------------------------------------------
$display(" G1 mode x bit-order x width");
for (mi = 0; mi < 4; mi = mi + 1) begin
for (oi = 0; oi < 2; oi = oi + 1) begin
grp_err = n_err;
for (wi = 0; wi < 4; wi = wi + 1) begin
set_cfg(mi[1], mi[0], oi[0], WID[wi], 8'd1, 2'd1, 4'd2, 4'd2, 4'd2);
slv_word = SWP[wi] & mask(WID[wi]);
fire(TXP[wi]);
wait_idle(4000, gd, tk);
chki ("done pulsed", gd ? 1 : 0, 1);
chk16("master rx", rx_data,
exp_master_rx(SWP[wi] & mask(WID[wi]), WID[wi], oi[0]));
chk16("slave rx", slv_rx & mask(WID[wi]),
exp_slave_rx(TXP[wi], WID[wi], oi[0]));
chki ("sclk edges", n_edge, 2 * WID[wi]);
chki ("bits sampled", bits_done, WID[wi]);
chki ("sclk parked", (sclk === mi[1]) ? 1 : 0, 1);
chki ("no cs overlap", n_overlap, 0);
end
$display(" mode %0d %0s : rx(w4,w8,w13,w16) checked, %0s",
mi, oi[0] ? "lsb" : "msb",
(n_err == grp_err) ? "all match" : "MISMATCH");
end
end
// -----------------------------------------------------------------
// GROUP 2 -- measured pin intervals.
// The question: are the lead / lag / turnaround requirements MET at the
// pins, and what is the actual margin? The requirement is ">=", the
// measurement is exact, and the difference is reported rather than assumed.
// -----------------------------------------------------------------
$display(" G2 measured pin intervals, in clk cycles");
$display(" div lead lag idle | t_half t_lead t_lag t_gap");
for (k = 0; k < 4; k = k + 1) begin
set_cfg(1'b0, 1'b0, 1'b0, 5'd8,
(k == 0) ? 8'd0 : (k == 1) ? 8'd1 : (k == 2) ? 8'd3 : 8'd1,
2'd2,
(k == 3) ? 4'd5 : 4'd2,
(k == 3) ? 4'd4 : 4'd2,
(k == 3) ? 4'd6 : 4'd3);
slv_word = 16'h5A;
// Two frames: the second one's gap is the interval between them.
fire(16'h33); wait_idle(4000, gd, tk);
fire(16'h33); wait_idle(4000, gd, tk);
$display(" %3d %4d %3d %4d | %6d %6d %5d %5d",
cfg_div, cfg_lead, cfg_lag, cfg_idle, m_half, m_lead, m_lag, m_gap);
// REQ-TIM-001 is an equality and is checked as one.
chki("t_half exact", m_half, cfg_div + 1);
// REQ-TIM-002/003/004 are ">=" requirements, and they are checked as
// stated -- but a ">=" check is WEAK: an implementation that loses one
// half-period of lead still satisfies it whenever cfg_lead > 0. So each
// interval is ALSO checked against the exact closed form the measurements
// above establish, which is what actually catches a one-tick regression.
if (m_lead < (cfg_lead * (cfg_div + 1))) begin
n_err = n_err + 1; $display(" FAIL lead below requirement");
end
n_chk = n_chk + 1;
if (m_lag < (cfg_lag * (cfg_div + 1))) begin
n_err = n_err + 1; $display(" FAIL lag below requirement");
end
n_chk = n_chk + 1;
if (m_gap < (cfg_idle * (cfg_div + 1))) begin
n_err = n_err + 1; $display(" FAIL turnaround below requirement");
end
n_chk = n_chk + 1;
// Exact implementation forms, in system clocks. The "+2" and "+1" terms
// are the ticks the state machine spends LEAVING a phase, and they are
// margin above the requirement, not part of it:
//
// t_lead = (cfg_lead + 2) * t_half one tick to exit S_LEAD, one more
// before the first transition
// t_lag = (cfg_lag + 1) * t_half one tick to exit S_LAG
// t_gap = (cfg_idle + 1) * t_half + 2
//
// The gap is the ONLY interval here with a term that is not a multiple of
// the half-period, and the reason is worth naming: those 2 cycles are
// S_IDLE plus ONE CYCLE OF THE REQUESTER'S OWN LATENCY. With no request
// queue, the bench cannot post the next transfer until `busy` falls, so
// the measured gap is the design's floor PLUS however long the requester
// took to react. 20.2 defends that trade-off and 20.7 questions it.
chki("t_lead exact", m_lead, (cfg_lead + 2) * (cfg_div + 1));
chki("t_lag exact", m_lag, (cfg_lag + 1) * (cfg_div + 1));
chki("t_gap exact", m_gap, (cfg_idle + 1) * (cfg_div + 1) + 2);
end
// -----------------------------------------------------------------
// GROUP 3 -- illegal configuration is REFUSED, at both boundaries.
// -----------------------------------------------------------------
$display(" G3 illegal width");
for (k = 0; k < 4; k = k + 1) begin
set_cfg(1'b0, 1'b0, 1'b0,
(k == 0) ? 5'd3 : (k == 1) ? 5'd17 : (k == 2) ? 5'd4 : 5'd16,
8'd1, 2'd0, 4'd2, 4'd2, 4'd2);
slv_word = 16'hFFFF & mask(cfg_width);
slv_w = (k < 2) ? 5'd8 : cfg_width;
fire(16'h5555);
// `fire` returns on the negedge immediately after the sampling edge, so a
// one-cycle output is readable RIGHT HERE. One more negedge and `cfg_err`
// has already gone low -- which is how a correct rejection reads as a
// missing one.
if (k < 2) begin
chki("cfg_err raised", cfg_err ? 1 : 0, 1);
chki("stayed idle", busy ? 1 : 0, 0);
chki("no cs asserted", cs_any ? 1 : 0, 0);
$display(" width %2d rejected: cfg_err=1 busy=0 cs=idle", cfg_width);
end else begin
chki("accepted", busy ? 1 : 0, 1);
chki("no cfg_err", cfg_err ? 1 : 0, 0);
wait_idle(4000, gd, tk);
chki("done pulsed", gd ? 1 : 0, 1);
$display(" width %2d accepted: cfg_err=0 done=1", cfg_width);
end
end
// -----------------------------------------------------------------
// GROUP 4 -- a request while busy has NO effect (REQ-FUNC-007).
// -----------------------------------------------------------------
$display(" G4 start while busy");
set_cfg(1'b0, 1'b0, 1'b0, 5'd8, 8'd2, 2'd0, 4'd2, 4'd2, 4'd2);
slv_word = 16'h96;
fire(16'hA5);
repeat (6) @(posedge clk);
fire(16'h3C); // ignored: the frame in flight owns the bus
wait_idle(4000, gd, tk);
chk16("rx from FIRST word", rx_data, exp_master_rx(16'h96, 5'd8, 1'b0));
chk16("slave got FIRST word", slv_rx & mask(5'd8), exp_slave_rx(16'hA5, 5'd8, 1'b0));
chki ("edges of one frame", n_edge, 16);
$display(" second request ignored: one frame, 16 edges, slave saw a5");
// -----------------------------------------------------------------
// GROUP 5 -- reset in the middle of a frame (REQ-RST-002).
// -----------------------------------------------------------------
$display(" G5 reset during a frame");
set_cfg(1'b1, 1'b0, 1'b0, 5'd16, 8'd3, 2'd3, 4'd2, 4'd2, 4'd2);
slv_word = 16'hBEEF;
fire(16'hDEAD);
repeat (20) @(posedge clk);
chki("frame really started", cs_any ? 1 : 0, 1);
@(negedge clk); rst_n = 1'b0;
@(negedge clk);
chki("all selects released", cs_any ? 1 : 0, 0);
chki("no done on reset", done ? 1 : 0, 0);
chki("not busy", busy ? 1 : 0, 0);
chki("sclk defined 0", (sclk === 1'b0) ? 1 : 0, 1);
$display(" reset mid-frame: cs released, no done, sclk=0 (not cpol)");
@(negedge clk); rst_n = 1'b1; repeat (3) @(posedge clk);
// -----------------------------------------------------------------
// GROUP 6 -- abort (REQ-ABT-001).
// -----------------------------------------------------------------
$display(" G6 abort mid-frame");
set_cfg(1'b0, 1'b0, 1'b0, 5'd16, 8'd1, 2'd0, 4'd2, 4'd2, 4'd2);
slv_word = 16'h1234;
fire(16'h4321);
wait_idle(4000, gd, tk); // a clean frame first, to set rx_data
rx_before = rx_data;
slv_word = 16'hFFFF;
fire(16'h0F0F);
repeat (24) @(posedge clk);
@(negedge clk); abort = 1'b1; @(negedge clk); abort = 1'b0;
wait_idle(4000, gd, tk);
chki ("no done on abort", gd ? 1 : 0, 0);
chk16("rx_data unchanged", rx_data, rx_before);
chki ("selects released", cs_any ? 1 : 0, 0);
chki ("sclk back at idle", (sclk === 1'b0) ? 1 : 0, 1);
if (bits_done == 5'd0 || bits_done >= 5'd16) begin
n_err = n_err + 1;
$display(" FAIL bits_done not partial: %0d", bits_done);
end
n_chk = n_chk + 1;
$display(" abort: done=0, rx held %04h, bits_done=%0d of 16",
rx_data, bits_done);
// -----------------------------------------------------------------
// GROUP 7 -- configuration captured at acceptance (commitment 2).
// The frame must ignore a mid-flight rewrite of EVERY field it uses.
// -----------------------------------------------------------------
$display(" G7 configuration captured at acceptance");
set_cfg(1'b0, 1'b0, 1'b0, 5'd8, 8'd2, 2'd0, 4'd2, 4'd2, 4'd2);
slv_word = 16'h7E;
fire(16'h81);
repeat (5) @(posedge clk);
// Rewrite everything to values that would produce a different frame.
cfg_cpol = 1'b1; cfg_cpha = 1'b1; cfg_lsb = 1'b1;
cfg_width = 5'd4; cfg_div = 8'd7; cfg_dev = 2'd3;
wait_idle(4000, gd, tk);
chk16("rx used captured cfg", rx_data, exp_master_rx(16'h7E, 5'd8, 1'b0));
chki ("edges of captured width", n_edge, 16);
chki ("captured device kept", cs_n[0] === 1'b1 ? 1 : 0, 1);
$display(" mid-frame rewrite ignored: 16 edges, msb-order rx=%04h", rx_data);
set_cfg(1'b0, 1'b0, 1'b0, 5'd8, 8'd1, 2'd0, 4'd2, 4'd2, 4'd2);
// -----------------------------------------------------------------
// GROUP 8 -- an X on MISO must NOT be able to produce a pass.
// This is a NEGATIVE test: the checker is EXPECTED to fail here, and the
// bench counts that as evidence rather than as a defect.
// -----------------------------------------------------------------
$display(" G8 X on MISO cannot slip through");
slv_word = 16'h5A;
force_x = 1'b1;
fire(16'hA5);
wait_idle(4000, gd, tk);
n_chk = n_chk + 1;
if (rx_data !== exp_master_rx(16'h5A, 5'd8, 1'b0)) begin
n_neg_ok = n_neg_ok + 1;
$display(" X reached rx_data and the !== comparison rejected it");
end else begin
n_err = n_err + 1;
$display(" FAIL an all-X rx_data compared EQUAL -- checker is blind");
end
force_x = 1'b0;
// -----------------------------------------------------------------
// GROUP 9 -- prove the checkers can fail at all.
// A suite that has never printed FAIL has not been shown to be able to.
// -----------------------------------------------------------------
$display(" G9 deliberately wrong expectations must FAIL");
// The pattern here is 9d, and the reason it is not a5 or 5a is worth more
// than the test it enables:
//
// a5 = 10100101 reversed = 10100101
// 5a = 01011010 reversed = 01011010
//
// The two most-used test bytes in digital design are both EIGHT-BIT
// PALINDROMES. Neither can distinguish MSB-first from LSB-first, so a suite
// built on them reports a clean pass with the bit-order logic inverted -- and
// the self-test of the oracle below reported a defect in the ORACLE when the
// only defect was the choice of stimulus. 9d reverses to b9.
k = 0;
slv_word = 16'h9D;
fire(16'h9D);
wait_idle(4000, gd, tk);
n_chk = n_chk + 1;
if (rx_data !== (exp_master_rx(16'h9D, 5'd8, 1'b0) ^ 16'h0001)) k = k + 1;
else begin n_err = n_err + 1; $display(" FAIL rx checker accepted a wrong value"); end
n_chk = n_chk + 1;
if (n_edge !== 15) k = k + 1;
else begin n_err = n_err + 1; $display(" FAIL edge checker accepted a wrong count"); end
n_chk = n_chk + 1;
if (exp_master_rx(16'h9D, 5'd8, 1'b1) !== exp_master_rx(16'h9D, 5'd8, 1'b0)) k = k + 1;
else begin n_err = n_err + 1; $display(" FAIL oracle is blind to bit order"); end
n_neg_ok = n_neg_ok + k;
$display(" %0d of 3 wrong expectations rejected; oracle separates 9d from b9", k);
$display("=== SUMMARY checks=%0d negatives=%0d failures=%0d : %0s ===",
n_chk, n_neg_ok, n_err, (n_err == 0) ? "PASS" : "FAIL");
$finish;
end
// Global watchdog. A hung frame must end the run with a FAILURE, not with a
// simulator that stops printing.
initial begin
#4000000;
$display(" FATAL global timeout -- a frame never completed");
$display("=== SUMMARY checks=%0d negatives=%0d failures=%0d : FAIL ===",
n_chk, n_neg_ok, n_err + 1);
$finish;
end
endmodule-- spi_capstone_ctrl_tb.vhd
--
-- Chapter 20.3 -- the directed suite in VHDL-2008. Same nine groups, same stimulus,
-- same expected values, and a transcript that must match the other two languages
-- LINE FOR LINE. That comparison is the real cross-language check: reading three files
-- side by side proves they look alike; comparing three transcripts proves they behave
-- alike, and only one of those is evidence.
--
-- THREE STRUCTURAL DECISIONS FORCED BY VHDL, ALL WORTH KNOWING
--
-- (1) THE SLAVE MODEL IS ONE PROCESS, NOT TWO.
-- The SystemVerilog model uses two always blocks -- one on the select falling,
-- one on every SCLK transition -- and both assign MISO. In VHDL that is TWO
-- DRIVERS on a resolved type, and the result is not an error: it is 'X' on the
-- pin, silently, for the whole run. The two events are merged into a single
-- process with an explicit priority between them.
--
-- (2) EVERY MONITOR TIMESTAMP IS A SIGNAL, NOT A VARIABLE.
-- The monitor reads `n_edge` and `t_last_edge` in the same invocation that
-- writes them, and the SystemVerilog original reads the PRE-EDGE values because
-- non-blocking assignment defers the update. Variables would be visible
-- immediately and every measured interval would come out one cycle short. This
-- is the Module 19 defect that made a device model under-report its own timing
-- requirement, and signals are the fix.
--
-- (3) THE CHECK COUNTERS ARE PROCESS VARIABLES.
-- Signals would need a single driving process, and every check happens inside
-- the stimulus process anyway. Variables are correct here for exactly the reason
-- they were wrong in (2): nothing outside this process reads them.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
use std.textio.all;
use std.env.all;
entity spi_capstone_ctrl_tb is
end entity spi_capstone_ctrl_tb;
architecture tb of spi_capstone_ctrl_tb is
signal clk : std_logic := '0';
signal rst_n : std_logic := '0';
signal cfg_cpol : std_logic := '0';
signal cfg_cpha : std_logic := '0';
signal cfg_lsb : std_logic := '0';
signal cfg_width : std_logic_vector(4 downto 0) := "01000";
signal cfg_div : std_logic_vector(7 downto 0) := x"01";
signal cfg_dev : std_logic_vector(1 downto 0) := "00";
signal cfg_lead : std_logic_vector(3 downto 0) := x"1";
signal cfg_lag : std_logic_vector(3 downto 0) := x"1";
signal cfg_idle : std_logic_vector(3 downto 0) := x"1";
signal start : std_logic := '0';
signal abort : std_logic := '0';
signal tx_data : std_logic_vector(15 downto 0) := (others => '0');
signal busy : std_logic;
signal done : std_logic;
signal cfg_err : std_logic;
signal rx_data : std_logic_vector(15 downto 0);
signal bits_done : std_logic_vector(4 downto 0);
signal sclk : std_logic;
signal mosi : std_logic;
signal cs_n : std_logic_vector(3 downto 0);
signal miso : std_logic;
-- slave model
signal slv_cpol : std_logic := '0';
signal slv_cpha : std_logic := '0';
signal slv_word : std_logic_vector(15 downto 0) := (others => '0');
signal slv_rx : std_logic_vector(15 downto 0) := (others => '0');
signal slv_w : unsigned(4 downto 0) := to_unsigned(8, 5);
signal slv_miso : std_logic := '0';
signal force_x : std_logic := '0';
signal cs_any : std_logic;
-- monitor state: signals, for the reason in note (2) above
signal cyc : integer := 0;
signal n_edge : integer := 0;
signal n_overlap : integer := 0;
signal m_half : integer := 0;
signal m_lead : integer := 0;
signal m_lag : integer := 0;
signal m_gap : integer := 0;
signal cs_d : std_logic := '0';
signal sclk_d : std_logic := '0';
signal t_cs_fall : integer := 0;
signal t_last_edge : integer := 0;
signal t_prev_rise : integer := 0;
-- ---------------- specification arithmetic: the independent oracle -------------
function maskw (w : integer) return unsigned is
variable m : unsigned(15 downto 0);
begin
m := (others => '0');
for b in 0 to 15 loop
if b < w then m(b) := '1'; end if;
end loop;
return m;
end function maskw;
function revw (v : std_logic_vector(15 downto 0); w : integer)
return std_logic_vector is
variable r : std_logic_vector(15 downto 0);
begin
r := (others => '0');
for b in 0 to 15 loop
if b < w then r(w-1-b) := v(b); end if;
end loop;
return r;
end function revw;
function exp_slave_rx (tx : std_logic_vector(15 downto 0);
w : integer; lsb : std_logic)
return std_logic_vector is
variable m : std_logic_vector(15 downto 0);
begin
m := std_logic_vector(unsigned(tx) and maskw(w));
if lsb = '1' then return revw(m, w); else return m; end if;
end function exp_slave_rx;
function exp_master_rx (sw : std_logic_vector(15 downto 0);
w : integer; lsb : std_logic)
return std_logic_vector is
variable m : std_logic_vector(15 downto 0);
begin
m := std_logic_vector(unsigned(sw) and maskw(w));
if lsb = '1' then return revw(m, w); else return m; end if;
end function exp_master_rx;
-- ---------------- output formatting -------------------------------------------
function hex4 (v : std_logic_vector(15 downto 0)) return string is
constant D : string(1 to 16) := "0123456789abcdef";
variable s : string(1 to 4);
variable n : integer;
begin
if is_x(v) then return "xxxx"; end if;
n := to_integer(unsigned(v));
for i in 4 downto 1 loop
s(i) := D((n mod 16) + 1);
n := n / 16;
end loop;
return s;
end function hex4;
-- Right-align an integer in w columns. The return is CONSTRAINED to w so no
-- caller inherits a surprising range, and the digits are built into a local
-- scratch string rather than concatenated, which would make the length depend on
-- the value -- a fatal length mismatch waiting for the first two-digit result.
function ipad (v : integer; w : integer) return string is
variable s : string(1 to w);
variable t : string(1 to 20);
variable n : integer;
variable len : integer;
begin
t := (others => ' ');
n := v;
len := 0;
if n = 0 then
len := 1; t(1) := '0';
else
while n > 0 loop
len := len + 1;
t(len) := character'val(character'pos('0') + (n mod 10));
n := n / 10;
end loop;
end if;
s := (others => ' ');
for i in 1 to len loop
s(w - i + 1) := t(i);
end loop;
return s;
end function ipad;
function i0 (v : integer) return string is
begin
return integer'image(v);
end function i0;
procedure pr (s : string) is
variable l : line;
begin
write(l, s);
writeline(output, l);
end procedure pr;
-- "msb" / "lsb" as a fixed THREE-character result. A conditional expression
-- (`if ... then ... else ...` inside an expression) is VHDL-2019, not 2008, so it
-- cannot be used here; and returning strings of different lengths from one
-- function is the length-mismatch fault Module 18 hit. Both are 3.
function ord_s (lsb : integer) return string is
begin
if lsb = 1 then return "lsb"; else return "msb"; end if;
end function ord_s;
function sl (b : boolean) return std_logic is
begin
if b then return '1'; else return '0'; end if;
end function sl;
type int4_t is array (0 to 3) of integer;
constant WID : int4_t := (4, 8, 13, 16);
constant TXP : std_logic_vector(15 downto 0) := x"B39D";
constant SWP : std_logic_vector(15 downto 0) := x"4E7A";
begin
cs_any <= not (cs_n(0) and cs_n(1) and cs_n(2) and cs_n(3));
miso <= 'X' when force_x = '1' else slv_miso;
dut : entity work.spi_capstone_ctrl
generic map (DATA_W => 16, MIN_WIDTH => 4, NDEV => 4)
port map (
clk => clk, rst_n => rst_n,
cfg_cpol => cfg_cpol, cfg_cpha => cfg_cpha, cfg_lsb_first => cfg_lsb,
cfg_width => cfg_width, cfg_div => cfg_div, cfg_dev => cfg_dev,
cfg_lead => cfg_lead, cfg_lag => cfg_lag, cfg_idle => cfg_idle,
start => start, tx_data => tx_data, abort => abort,
busy => busy, done => done, cfg_err => cfg_err,
rx_data => rx_data, bits_done => bits_done,
sclk => sclk, mosi => mosi, cs_n => cs_n, miso => miso
);
clkgen : process
begin
clk <= '0'; wait for 5 ns;
clk <= '1'; wait for 5 ns;
end process clkgen;
slave : process (cs_any, sclk)
variable sr : std_logic_vector(15 downto 0);
variable idx : unsigned(4 downto 0);
variable nrx : unsigned(4 downto 0);
variable lead_s : boolean;
begin
if rising_edge(cs_any) then
sr := std_logic_vector(shift_left(unsigned(slv_word),
16 - to_integer(slv_w)));
slv_rx <= (others => '0');
nrx := (others => '0');
if slv_cpha = '0' then
slv_miso <= sr(15);
sr := sr(14 downto 0) & '0';
idx := to_unsigned(1, 5);
else
slv_miso <= '0';
idx := (others => '0');
end if;
elsif sclk'event and cs_any = '1' then
-- A transition AWAY from the idle level is leading; back to it, trailing.
-- NOT mirrored from the controller: a slave uses the SAME edges as its
-- master. Both sample on one edge and both change their output on the
-- other -- that is what makes the link work off a single clock.
lead_s := (sclk /= slv_cpol);
if (slv_cpha = '1' and not lead_s) or (slv_cpha = '0' and lead_s) then
if nrx < slv_w then
slv_rx <= slv_rx(14 downto 0) & mosi;
nrx := nrx + 1;
end if;
end if;
if (slv_cpha = '1' and lead_s) or (slv_cpha = '0' and not lead_s) then
if idx < slv_w then
slv_miso <= sr(15);
sr := sr(14 downto 0) & '0';
idx := idx + 1;
end if;
end if;
end if;
end process slave;
mon : process (clk)
variable n_low : integer;
begin
if rising_edge(clk) then
if rst_n = '0' then
cyc <= 0;
n_edge <= 0;
n_overlap <= 0;
cs_d <= '0';
sclk_d <= '0';
else
cyc <= cyc + 1;
n_low := 0;
for i in 0 to 3 loop
if cs_n(i) = '0' then n_low := n_low + 1; end if;
end loop;
if n_low > 1 then n_overlap <= n_overlap + 1; end if;
if cs_any = '1' and cs_d = '0' then
t_cs_fall <= cyc;
m_gap <= cyc - t_prev_rise;
n_edge <= 0;
end if;
if cs_any = '0' and cs_d = '1' then
t_prev_rise <= cyc;
m_lag <= cyc - t_last_edge;
end if;
-- `cs_d` as well as `cs_any`: an edge is only a FRAME edge if a device
-- was ALREADY selected last cycle. A transition in the cycle the
-- select falls is SCLK reaching its new idle level, not a clocking
-- edge. Without this gate the first CPOL=1 transaction counted 17
-- edges here and 16 in SystemVerilog, on identical data.
if cs_any = '1' and cs_d = '1' and sclk /= sclk_d then
if n_edge = 0 then
m_lead <= cyc - t_cs_fall;
else
m_half <= cyc - t_last_edge;
end if;
t_last_edge <= cyc;
n_edge <= n_edge + 1;
end if;
cs_d <= cs_any;
sclk_d <= sclk;
end if;
end if;
end process mon;
wd : process
begin
wait for 4 ms;
pr(" FATAL global timeout -- a frame never completed");
pr("=== SUMMARY checks=0 negatives=0 failures=1 : FAIL ===");
finish;
end process wd;
main : process
variable n_chk : integer := 0;
variable n_err : integer := 0;
variable n_neg : integer := 0;
variable grp_err : integer := 0;
variable gd : std_logic;
variable tk : integer;
variable kk : integer;
variable w : integer;
variable rx_before : std_logic_vector(15 downto 0);
procedure chk16 (nm : string; got, exp : std_logic_vector(15 downto 0)) is
begin
n_chk := n_chk + 1;
-- Case comparison: an 'X' anywhere in `got` must FAIL. With `=` on
-- std_logic_vector an 'X' does not match '0' or '1' either, but printing
-- it needs the is_x guard in hex4 -- group 8 exercises exactly this path.
if got /= exp then
n_err := n_err + 1;
pr(" FAIL " & nm & ": got " & hex4(got) & " expected " & hex4(exp));
end if;
end procedure chk16;
procedure chki (nm : string; got, exp : integer) is
begin
n_chk := n_chk + 1;
if got /= exp then
n_err := n_err + 1;
pr(" FAIL " & nm & ": got " & i0(got) & " expected " & i0(exp));
end if;
end procedure chki;
procedure set_cfg (cpol_i, cpha_i, lsb_i : std_logic; wv : integer;
dv : integer; dv_n : integer;
ld, lg, idl : integer) is
begin
cfg_cpol <= cpol_i;
cfg_cpha <= cpha_i;
cfg_lsb <= lsb_i;
cfg_width <= std_logic_vector(to_unsigned(wv, 5));
cfg_div <= std_logic_vector(to_unsigned(dv, 8));
cfg_dev <= std_logic_vector(to_unsigned(dv_n, 2));
cfg_lead <= std_logic_vector(to_unsigned(ld, 4));
cfg_lag <= std_logic_vector(to_unsigned(lg, 4));
cfg_idle <= std_logic_vector(to_unsigned(idl, 4));
slv_cpol <= cpol_i;
slv_cpha <= cpha_i;
slv_w <= to_unsigned(wv, 5);
end procedure set_cfg;
-- Every input is driven on the FALLING edge. Driving `start` on the edge the
-- controller samples it on is a race whose measured cost was that the FIRST
-- frame of the run was accepted and every later one silently dropped, so 31
-- of 32 transfers compared a stale `rx_data` against a fresh expectation.
procedure fire (d : std_logic_vector(15 downto 0)) is
begin
wait until falling_edge(clk);
tx_data <= d;
start <= '1';
wait until falling_edge(clk);
start <= '0';
end procedure fire;
procedure wait_idle (maxc : integer; got_done : out std_logic;
took : out integer) is
variable g : integer;
variable seen : std_logic;
begin
g := 0;
seen := '0';
while g < maxc loop
wait until falling_edge(clk);
g := g + 1;
if done = '1' then seen := '1'; end if;
if busy = '0' and seen = '1' then
g := maxc;
elsif busy = '0' and g > 4 then
g := maxc;
end if;
end loop;
got_done := seen;
took := g;
end procedure wait_idle;
begin
pr("=== Chapter 20.3 -- capstone controller, directed suite ===");
-- Reset is RELEASED ON A FALLING EDGE, for the same reason `start` is driven on
-- one: released on a rising edge it races every clocked block that tests it.
-- The measured cost was a one-cycle difference in when the monitor left reset,
-- which counted SCLK's idle re-park as a frame edge in VHDL but not in
-- SystemVerilog -- 17 edges against 16, on the first transaction only.
for i in 1 to 4 loop wait until rising_edge(clk); end loop;
wait until falling_edge(clk);
rst_n <= '1';
for i in 1 to 2 loop wait until rising_edge(clk); end loop;
----------------------------------------------------------------------
pr(" G1 mode x bit-order x width");
for mi in 0 to 3 loop
for oi in 0 to 1 loop
grp_err := n_err;
for wi in 0 to 3 loop
w := WID(wi);
set_cfg(sl(mi / 2 = 1), sl(mi mod 2 = 1), sl(oi = 1),
w, 1, 1, 2, 2, 2);
slv_word <= std_logic_vector(unsigned(SWP) and maskw(w));
fire(TXP);
wait_idle(4000, gd, tk);
chki ("done pulsed", to_integer(unsigned'('0' & gd)), 1);
chk16("master rx", rx_data,
exp_master_rx(std_logic_vector(unsigned(SWP) and maskw(w)),
w, sl(oi = 1)));
chk16("slave rx",
std_logic_vector(unsigned(slv_rx) and maskw(w)),
exp_slave_rx(TXP, w, sl(oi = 1)));
chki ("sclk edges", n_edge, 2 * w);
chki ("bits sampled", to_integer(unsigned(bits_done)), w);
chki ("sclk parked",
to_integer(unsigned'('0' & sl(sclk = sl(mi / 2 = 1)))), 1);
chki ("no cs overlap", n_overlap, 0);
end loop;
if n_err = grp_err then
pr(" mode " & i0(mi) & " " & ord_s(oi) &
" : rx(w4,w8,w13,w16) checked, all match");
else
pr(" mode " & i0(mi) & " " & ord_s(oi) &
" : rx(w4,w8,w13,w16) checked, MISMATCH");
end if;
end loop;
end loop;
----------------------------------------------------------------------
pr(" G2 measured pin intervals, in clk cycles");
pr(" div lead lag idle | t_half t_lead t_lag t_gap");
for k in 0 to 3 loop
if k = 0 then
set_cfg('0', '0', '0', 8, 0, 2, 2, 2, 3);
elsif k = 1 then
set_cfg('0', '0', '0', 8, 1, 2, 2, 2, 3);
elsif k = 2 then
set_cfg('0', '0', '0', 8, 3, 2, 2, 2, 3);
else
set_cfg('0', '0', '0', 8, 1, 2, 5, 4, 6);
end if;
slv_word <= x"005A";
fire(x"0033"); wait_idle(4000, gd, tk);
fire(x"0033"); wait_idle(4000, gd, tk);
pr(" " & ipad(to_integer(unsigned(cfg_div)), 3) &
" " & ipad(to_integer(unsigned(cfg_lead)), 4) &
" " & ipad(to_integer(unsigned(cfg_lag)), 3) &
" " & ipad(to_integer(unsigned(cfg_idle)), 4) &
" | " & ipad(m_half, 6) &
" " & ipad(m_lead, 6) &
" " & ipad(m_lag, 5) &
" " & ipad(m_gap, 5));
chki("t_half exact", m_half, to_integer(unsigned(cfg_div)) + 1);
if m_lead < to_integer(unsigned(cfg_lead))
* (to_integer(unsigned(cfg_div)) + 1) then
n_err := n_err + 1; pr(" FAIL lead below requirement");
end if;
n_chk := n_chk + 1;
if m_lag < to_integer(unsigned(cfg_lag))
* (to_integer(unsigned(cfg_div)) + 1) then
n_err := n_err + 1; pr(" FAIL lag below requirement");
end if;
n_chk := n_chk + 1;
if m_gap < to_integer(unsigned(cfg_idle))
* (to_integer(unsigned(cfg_div)) + 1) then
n_err := n_err + 1; pr(" FAIL turnaround below requirement");
end if;
n_chk := n_chk + 1;
chki("t_lead exact", m_lead,
(to_integer(unsigned(cfg_lead)) + 2)
* (to_integer(unsigned(cfg_div)) + 1));
chki("t_lag exact", m_lag,
(to_integer(unsigned(cfg_lag)) + 1)
* (to_integer(unsigned(cfg_div)) + 1));
chki("t_gap exact", m_gap,
(to_integer(unsigned(cfg_idle)) + 1)
* (to_integer(unsigned(cfg_div)) + 1) + 2);
end loop;
----------------------------------------------------------------------
pr(" G3 illegal width");
for k in 0 to 3 loop
if k = 0 then w := 3;
elsif k = 1 then w := 17;
elsif k = 2 then w := 4;
else w := 16;
end if;
set_cfg('0', '0', '0', w, 1, 0, 2, 2, 2);
if k < 2 then
slv_w <= to_unsigned(8, 5);
slv_word <= x"00FF";
else
slv_word <= std_logic_vector(unsigned'(x"FFFF") and maskw(w));
end if;
fire(x"5555");
if k < 2 then
chki("cfg_err raised", to_integer(unsigned'('0' & cfg_err)), 1);
chki("stayed idle", to_integer(unsigned'('0' & busy)), 0);
chki("no cs asserted", to_integer(unsigned'('0' & cs_any)), 0);
pr(" width " & ipad(w, 2) &
" rejected: cfg_err=1 busy=0 cs=idle");
else
chki("accepted", to_integer(unsigned'('0' & busy)), 1);
chki("no cfg_err", to_integer(unsigned'('0' & cfg_err)), 0);
wait_idle(4000, gd, tk);
chki("done pulsed", to_integer(unsigned'('0' & gd)), 1);
pr(" width " & ipad(w, 2) & " accepted: cfg_err=0 done=1");
end if;
end loop;
----------------------------------------------------------------------
pr(" G4 start while busy");
set_cfg('0', '0', '0', 8, 2, 0, 2, 2, 2);
slv_word <= x"0096";
fire(x"00A5");
for i in 1 to 6 loop wait until rising_edge(clk); end loop;
fire(x"003C");
wait_idle(4000, gd, tk);
chk16("rx from FIRST word", rx_data, exp_master_rx(x"0096", 8, '0'));
chk16("slave got FIRST word",
std_logic_vector(unsigned(slv_rx) and maskw(8)),
exp_slave_rx(x"00A5", 8, '0'));
chki ("edges of one frame", n_edge, 16);
pr(" second request ignored: one frame, 16 edges, slave saw a5");
----------------------------------------------------------------------
pr(" G5 reset during a frame");
set_cfg('1', '0', '0', 16, 3, 3, 2, 2, 2);
slv_word <= x"BEEF";
fire(x"DEAD");
for i in 1 to 20 loop wait until rising_edge(clk); end loop;
chki("frame really started", to_integer(unsigned'('0' & cs_any)), 1);
wait until falling_edge(clk);
rst_n <= '0';
wait until falling_edge(clk);
chki("all selects released", to_integer(unsigned'('0' & cs_any)), 0);
chki("no done on reset", to_integer(unsigned'('0' & done)), 0);
chki("not busy", to_integer(unsigned'('0' & busy)), 0);
chki("sclk defined 0", to_integer(unsigned'('0' & sl(sclk = '0'))), 1);
pr(" reset mid-frame: cs released, no done, sclk=0 (not cpol)");
wait until falling_edge(clk);
rst_n <= '1';
for i in 1 to 3 loop wait until rising_edge(clk); end loop;
----------------------------------------------------------------------
pr(" G6 abort mid-frame");
set_cfg('0', '0', '0', 16, 1, 0, 2, 2, 2);
slv_word <= x"1234";
fire(x"4321");
wait_idle(4000, gd, tk);
rx_before := rx_data;
slv_word <= x"FFFF";
fire(x"0F0F");
for i in 1 to 24 loop wait until rising_edge(clk); end loop;
wait until falling_edge(clk);
abort <= '1';
wait until falling_edge(clk);
abort <= '0';
wait_idle(4000, gd, tk);
chki ("no done on abort", to_integer(unsigned'('0' & gd)), 0);
chk16("rx_data unchanged", rx_data, rx_before);
chki ("selects released", to_integer(unsigned'('0' & cs_any)), 0);
chki ("sclk back at idle", to_integer(unsigned'('0' & sl(sclk = '0'))), 1);
if to_integer(unsigned(bits_done)) = 0
or to_integer(unsigned(bits_done)) >= 16 then
n_err := n_err + 1;
pr(" FAIL bits_done not partial: " &
i0(to_integer(unsigned(bits_done))));
end if;
n_chk := n_chk + 1;
pr(" abort: done=0, rx held " & hex4(rx_data) & ", bits_done=" &
i0(to_integer(unsigned(bits_done))) & " of 16");
----------------------------------------------------------------------
pr(" G7 configuration captured at acceptance");
set_cfg('0', '0', '0', 8, 2, 0, 2, 2, 2);
slv_word <= x"007E";
fire(x"0081");
for i in 1 to 5 loop wait until rising_edge(clk); end loop;
cfg_cpol <= '1';
cfg_cpha <= '1';
cfg_lsb <= '1';
cfg_width <= std_logic_vector(to_unsigned(4, 5));
cfg_div <= std_logic_vector(to_unsigned(7, 8));
cfg_dev <= std_logic_vector(to_unsigned(3, 2));
wait_idle(4000, gd, tk);
chk16("rx used captured cfg", rx_data, exp_master_rx(x"007E", 8, '0'));
chki ("edges of captured width", n_edge, 16);
chki ("captured device kept",
to_integer(unsigned'('0' & sl(cs_n(0) = '1'))), 1);
pr(" mid-frame rewrite ignored: 16 edges, msb-order rx=" & hex4(rx_data));
set_cfg('0', '0', '0', 8, 1, 0, 2, 2, 2);
----------------------------------------------------------------------
pr(" G8 X on MISO cannot slip through");
slv_word <= x"005A";
force_x <= '1';
fire(x"00A5");
wait_idle(4000, gd, tk);
n_chk := n_chk + 1;
if rx_data /= exp_master_rx(x"005A", 8, '0') then
n_neg := n_neg + 1;
pr(" X reached rx_data and the !== comparison rejected it");
else
n_err := n_err + 1;
pr(" FAIL an all-X rx_data compared EQUAL -- checker is blind");
end if;
force_x <= '0';
----------------------------------------------------------------------
pr(" G9 deliberately wrong expectations must FAIL");
-- 9d, not a5 or 5a: those two are both EIGHT-BIT PALINDROMES
-- a5 = 10100101 reversed = 10100101
-- 5a = 01011010 reversed = 01011010
-- so neither can distinguish MSB-first from LSB-first, and the oracle
-- self-test below reported a defect in the ORACLE when the only defect was
-- the choice of stimulus. 9d reverses to b9.
kk := 0;
slv_word <= x"009D";
fire(x"009D");
wait_idle(4000, gd, tk);
n_chk := n_chk + 1;
if rx_data /= (exp_master_rx(x"009D", 8, '0') xor x"0001") then
kk := kk + 1;
else
n_err := n_err + 1;
pr(" FAIL rx checker accepted a wrong value");
end if;
n_chk := n_chk + 1;
if n_edge /= 15 then
kk := kk + 1;
else
n_err := n_err + 1;
pr(" FAIL edge checker accepted a wrong count");
end if;
n_chk := n_chk + 1;
if exp_master_rx(x"009D", 8, '1') /= exp_master_rx(x"009D", 8, '0') then
kk := kk + 1;
else
n_err := n_err + 1;
pr(" FAIL oracle is blind to bit order");
end if;
n_neg := n_neg + kk;
pr(" " & i0(kk) &
" of 3 wrong expectations rejected; oracle separates 9d from b9");
----------------------------------------------------------------------
if n_err = 0 then
pr("=== SUMMARY checks=" & i0(n_chk) & " negatives=" & i0(n_neg) &
" failures=" & i0(n_err) & " : PASS ===");
else
pr("=== SUMMARY checks=" & i0(n_chk) & " negatives=" & i0(n_neg) &
" failures=" & i0(n_err) & " : FAIL ===");
end if;
finish;
end process main;
end architecture tb;8. What It Measures
The transcript is identical in all three languages, byte for byte:
=== Chapter 20.3 -- capstone controller, directed suite ===
G1 mode x bit-order x width
mode 0 msb : rx(w4,w8,w13,w16) checked, all match
mode 0 lsb : rx(w4,w8,w13,w16) checked, all match
mode 1 msb : rx(w4,w8,w13,w16) checked, all match
mode 1 lsb : rx(w4,w8,w13,w16) checked, all match
mode 2 msb : rx(w4,w8,w13,w16) checked, all match
mode 2 lsb : rx(w4,w8,w13,w16) checked, all match
mode 3 msb : rx(w4,w8,w13,w16) checked, all match
mode 3 lsb : rx(w4,w8,w13,w16) checked, all match
G2 measured pin intervals, in clk cycles
div lead lag idle | t_half t_lead t_lag t_gap
0 2 2 3 | 1 4 3 6
1 2 2 3 | 2 8 6 10
3 2 2 3 | 4 16 12 18
1 5 4 6 | 2 14 10 16
G3 illegal width
width 3 rejected: cfg_err=1 busy=0 cs=idle
width 17 rejected: cfg_err=1 busy=0 cs=idle
width 4 accepted: cfg_err=0 done=1
width 16 accepted: cfg_err=0 done=1
G4 start while busy
second request ignored: one frame, 16 edges, slave saw a5
G5 reset during a frame
reset mid-frame: cs released, no done, sclk=0 (not cpol)
G6 abort mid-frame
abort: done=0, rx held 1234, bits_done=5 of 16
G7 configuration captured at acceptance
mid-frame rewrite ignored: 16 edges, msb-order rx=007e
G8 X on MISO cannot slip through
X reached rx_data and the !== comparison rejected it
G9 deliberately wrong expectations must FAIL
3 of 3 wrong expectations rejected; oracle separates 9d from b9
=== SUMMARY checks=284 negatives=4 failures=0 : PASS ===The timing table is the interesting part
Group 2 measures four intervals at the pins and finds every one of them an exact function of the configuration:
t_half = cfg_div + 1 system clocks
t_lead = (cfg_lead + 2) x t_half CS low to first edge
t_lag = (cfg_lag + 1) x t_half last edge to CS high
t_gap = (cfg_idle + 1) x t_half + 2 CS high to next CS lowCheck the last row against the third: at cfg_div = 1, cfg_lead = 5, cfg_lag = 4, cfg_idle = 6 the predicted values are (5+2)x2 = 14, (4+1)x2 = 10 and (6+1)x2+2 = 16, and the measured values are 14, 10 and 16.
The +1 and +2 terms are the ticks the state machine spends leaving a phase. They are margin above REQ-TIM-002 and REQ-TIM-003, not part of them, which is why the bench checks both the requirement's >= form and the exact form. The >= form is what the specification promises; the exact form is what catches a regression, because a controller that loses one half-period of lead still satisfies >= whenever cfg_lead is greater than zero.
The gap is the one interval with a term that is not a multiple of the half-period, and those 2 cycles are worth naming: one is the cycle spent in S_IDLE, and one is the requester's own reaction time. With no request queue, the bench cannot post the next transfer until busy falls. The measured gap is therefore the design's floor plus however long the caller took — which is Chapter 20.2 §6's cost, showing up as an arithmetic term.
9. Five Defects Found Writing This
Every one of these was found by running something, and three of them were found because the three languages disagreed.
D1 — A function-valued wire lags its inputs by a time step
The aligned word needed a name before a bit of it could be taken, because a part-select of a function call does not parse. The obvious name is a wire:
wire [15:0] ld = load_align(tx_data, cfg_width, cfg_lsb_first); // WRONG hereThat compiles, reads correctly from a testbench, and loaded the shift register with zero. A continuous assignment whose right-hand side is a function call settles a time step after its inputs change, so the clock edge in that same step sampled the previous value. Measured: tx_sr loaded 0000 while the wire read d000 one cycle later, and every transmitted bit was zero while every other captured field was correct.
The fix is to compute it inside the clocked block, where the function is called at the edge with the values the edge itself sampled. Note that this is the opposite of Module 19's VHDL lesson: a variable is wrong for state another process reads, and right for a value used and discarded inside one invocation.
D2 — Shift-then-present is correct in two modes and wrong in two
On a launch event, the choice is present the top bit and then shift, or shift and then present the new top bit. The second is what a first version naturally does, and it is correct for CPHA = 0 — because a bit was already presented when the chip select asserted, so the next launch genuinely wants the following bit.
For CPHA = 1 there is no pre-launch. Bit 0 is still pending at the first leading edge, so shifting first skips it.
modes 0 and 2 : all four widths passed
modes 1 and 3 : every word shifted up one position, in BOTH directions at onceBoth directions moved together because the same ordering is used by the controller and by the device model. The fix is one uniform rule — present, then shift — used by the pre-launch and by every launch, which is why the code has the same two lines in both places.
D3 — Driving start on the edge that samples it
The bench drove start high, waited for the sampling edge, and cleared it. The clearing assignment and the controller's clocked block are in the same region of the same time step, and their order is not defined by the language.
frame 1 accepted
frames 2..32 silently dropped31 of 32 transfers compared a stale rx_data against a fresh expectation, so the suite reported a wall of mismatches that had nothing to do with the controller. Every input is now driven on the negedge, so start is high across exactly one sampling edge and nothing the bench does can collide with the edge that matters. That is Chapter 16.3's clocking-block discipline written out by hand.
D4 — A function evaluated outside its domain, which only VHDL admitted
load_align shifts by DATA_W - w. For the rejected width of 17 that is -1, and it was being computed before the width was checked:
VHDL ** Fatal: value -1 outside of NATURAL range 0 to 2147483647
Verilog computed an unsigned wrap, produced a garbage word, discarded it
with the refused request, and said nothingBoth languages were evaluating a function outside its domain. Only one of them said so. The fix — validate the width before computing anything from it — went into all three, and this is the clearest case in the module of the third language acting as verification rather than as a translation.
D5 — The monitor counted the idle re-park as a frame edge
The controller parks SCLK at the configured idle level and asserts the chip select in the same cycle. When the previous frame's polarity differed, those two coincide, so an edge counter gated only on a device is selected counts the re-park.
VHDL 17 edges on the first CPOL=1 frame
SV 16 edges, identical received dataA one-cycle difference in when the monitor left reset decided whether the re-park landed inside the window. Gating the counter on a device was already selected last cycle is both the fix and the correct definition: a transition in the assert cycle is the clock reaching its idle level, not a clocking edge. The measurement no longer depends on that phase.
10. The Tri-HDL Result
| SystemVerilog | Verilog-2001 | VHDL-2008 | |
|---|---|---|---|
| Tool | iverilog -g2012 | iverilog -g2001 | nvc 1.23.0 |
| RTL lines | 401 | 401 | 361 |
| Testbench lines | 621 | 621 | 687 |
| Checks | 284 | 284 | 284 |
| Failures | 0 | 0 | 0 |
| Transcript | 34 lines | 34 lines | 34 lines |
The three transcripts are byte-identical, and the three simulations end at the same simulated time. Not one line differs by design in this chapter.
That identity is the actual cross-language evidence. Reading three files side by side proves they look alike; running them under one stimulus and comparing the output proves they behave alike, and only the second is evidence. Parity was checked across all thirteen axes — ports, widths, generics, clock edge, reset, priority, counter widths, output latency, bit order, handshake, abort, error semantics, and the shift alignment — but the byte-identical transcript is what makes the claim checkable rather than asserted.
11. Summary
The capstone controller exists in three languages, passes 284 checks in each, and produces byte-identical transcripts. All four SPI modes come out of ~edge_i[0] and two conditional expressions with no mention of CPOL, and an even transition count gives the parked-level requirement away for free.
The measured timing is an exact function of the configuration — t_half = cfg_div + 1, t_lead = (cfg_lead + 2) x t_half, t_lag = (cfg_lag + 1) x t_half, and a gap whose two extra cycles are S_IDLE plus the requester's own reaction time. That last term is the absence of a request queue showing up as arithmetic, and it means no back-to-back throughput figure from this interface describes the controller alone.
Five defects were found by running things. A function-valued wire settled a time step late and loaded zeros into the shift register while reading correctly from the bench. Shift-then-present was correct in modes 0 and 2 and off by one bit in modes 1 and 3, in both directions at once. Driving start on the sampling edge dropped 31 of 32 transfers. A width of 17 was shifted by −1, which VHDL called fatal and Verilog silently wrapped. And an edge counter gated only on selected counted the idle re-park, which differed between languages by one cycle of reset timing.
Three of those five were found because the three implementations disagreed, which is the argument for the third language stated as a measurement rather than a principle.
12. What Comes Next
Chapter 20.4 leaves the cycle counts behind and turns them into nanoseconds: the MISO round trip that decides the maximum SCLK rate, the constraints that tell a tool what this design's SCLK actually is, and an honest account of which crossings here are clock-domain problems and which are static timing problems wearing the wrong name.
Continue learning
Related tutorials
- Related topic
USB vs SPI
SPI selects a peripheral with a wire routed at layout time and USB with an address the host assigned — so a chip-select contention is invisible to every slave (0 of 11) while a duplicate USB address is detected every time (274 of 274).
- Related topic
Transfer Width
SPI has no native word size. What a 12-bit ADC or 16-bit codec requires of a master, why chip select rather than a bit count delimits a frame, and a width-parameterized transfer engine in three HDLs.
- Related topic
The Polling Model
The host runs a periodic schedule the device cannot see, so device hardware must count frames rather than trust a period — built in Verilog, SystemVerilog and VHDL with a schedule model that disagrees independently.
- Related topic
Keyboards on USB
A key matrix produces a set; the boot report has six slots. The mechanism that reconciles them has its own reserved code — and the chapter where the verification transaction stops being the pins.
