Skip to content
VLSI Mentor

SPI · Module 20

RTL Implementation and Mode Handling

The capstone controller in SystemVerilog, Verilog-2001 and VHDL with all four SPI modes derived from a parity, plus five defects found by running it — three of them because the three languages disagreed.

Chapter 20.2 chose an architecture. This chapter implements it three times, runs 284 checks against each implementation, and compares the three transcripts line for line.

All four SPI modes come out of one parity expression. The bug that took longest to find was correct in modes 0 and 2 and off by one bit in modes 1 and 3.

1. The Mode Decoder, In Three Languages

Start with the payoff. Everything Chapter 20.2 argued about mode handling reduces to this, and it is the same shape in all three languages:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
wire leading      = ~edge_i[0];
wire sample_event = cpha_q ? ~leading :  leading;
wire launch_event = cpha_q ?  leading : ~leading;
Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
leading      <= not edge_i(0);
sample_event <= (not leading) when cpha_q = '1' else leading;
launch_event <= leading       when cpha_q = '1' else (not leading);

edge_i is the index of the SCLK transition about to be made. Even indices move the clock away from its parked level; odd indices bring it back. CPOL appears nowhere, because the index already counts from whatever level CPOL parked the clock at.

Two consequences worth naming before the code:

All four modes are verified by exercising two bits. There is no fourth case to forget, because there are no cases — there is a parity and a conditional.

An even transition count makes REQ-MODE-001 free. Each frame makes exactly 2N transitions, so the clock necessarily ends where it began. Nothing re-parks it; the arithmetic does.

2. The Control State Machine

spi_capstone_ctrl control FSM

fsm
Five states. S_IDLE goes to S_LEAD on an accepted start, and has a self-loop for a refused request. S_LEAD goes to S_XFER when the lead count expires, and to S_LAG on abort. S_XFER goes to S_LAG after 2N edges or on abort. S_LAG goes to S_GAP when the lag expires, asserting done. S_GAP returns to S_IDLE when the turnaround expires.S_IDLES_LEADS_XFERS_LAGS_GAPstart acceptedstartacceptedlead elapsedlead elapsed2N edges, or abort2N edges, or abortlag elapsed, donelag elapsed, doneturnaround elapsedturnaround elapsedabortabortbad width: cfg_errbad width: cfg_err
busy is decoded as any state other than S_IDLE, which is what makes the turnaround in S_GAP structurally enforced rather than left to the caller.
Figure 1 — the five phases. Every transition out of a phase is driven by a tick from the divider, so each arrow is one SCLK half-period. The abort input reaches S_LAG from both active phases, which is what gives an abandoned frame its full lag time; and the refused-request self-loop on S_IDLE is REQ-ERR-001, where a bad width pulses cfg_err and nothing else happens.

3. CPHA At The Pins

The same 4-bit word, the same device, the same divider — and the only difference is cfg_cpha. Both waveforms are the exact behaviour of the RTL below, at cfg_div = 0, cfg_lead = 1, cfg_lag = 1, sending 0xD and receiving 0xA.

Mode 0: sample on the rising edge, launch on the falling edge

14 cycles
Fourteen cycles. Chip select is low from cycle zero to eleven. SCLK is low until cycle three, then alternates high and low each cycle until cycle ten. MOSI is high from the start, drops low at cycle six and returns high at cycle eight. MISO is high until cycle four, low until cycle six, high until cycle eight, then low. Phase bands mark the lead, the eight transitions, the lag, and the released gap.leadlead2N = 8 edges2N = 8 edgeslaglagidleidlefirst sample: MISO bit 3first sample: MISO bit 3first launch: MOSI bit 2first launch: MOSI bit 2done, rx_data = adone, rx_data = acs_nsclkmosimisobit00033221100000t0t1t2t3t4t5t6t7t8t9t10t11t12t13
Figure 2 — mode 0. The controller samples on every rising edge of SCLK and changes MOSI on every falling edge. Because the first sampling edge arrives before any falling edge exists, the first bit is already on MOSI when the chip select goes low — which is why MOSI is valid during the whole lead time.

Mode 1: launch on the rising edge, sample on the falling edge

14 cycles
Fourteen cycles, the same chip select and SCLK as the previous figure. MOSI is low until cycle three, then high until cycle seven, low until cycle nine, then high. MISO is low until cycle three, high until cycle five, low until cycle seven, high until cycle nine, then low.leadlead2N = 8 edges2N = 8 edgeslaglagidleidlefirst launch: MOSI bit 3first launch: MOSI bit 3first sample: MISO bit 3first sample: MISO bit 3done, rx_data = adone, rx_data = acs_nsclkmosimisobit00033221100000t0t1t2t3t4t5t6t7t8t9t10t11t12t13
Figure 3 — mode 1, everything else identical. Now the controller changes MOSI on the rising edge and samples on the falling one. MOSI holds a defined low through the whole lead time because no launch has happened yet, and the first data bit appears only at cycle three. Comparing the MOSI row here against Figure 2 is the entire content of CPHA.

4. The Implementation — SystemVerilog

Azvya Education Pvt. Ltd.VLSI Mentor
spi_capstone_ctrl.sv — the capstone controller, and the four commitments every line follows from
// spi_capstone_ctrl.sv
//
// THE CAPSTONE CONTROLLER. One design, carried from Chapter 20.1's specification
// through to Chapter 20.7's design review. Every later chapter verifies THIS module;
// none of them redesigns it.
//
// WHAT IT IS
//
//   A configurable SPI master. Four modes, either bit order, a transfer width chosen
//   per request, a programmable clock divider, four chip selects, and programmable
//   lead / lag / turnaround times.
//
// THE FOUR ARCHITECTURAL COMMITMENTS
//
//   These are decided here, defended in 20.2, constrained in 20.4, and questioned in
//   20.7. They are stated at the top because every line below follows from them.
//
//   (1) EVERY REGISTER IN THIS MODULE IS CLOCKED BY `clk`.
//       SCLK is an OUTPUT WAVEFORM this module generates, not a clock it uses. No
//       internal state is clocked by SCLK, so there is no internal clock-domain
//       crossing to synchronise. That is a deliberate architectural choice with a
//       cost (20.4 §"what this buys and what it costs") -- it is not an oversight,
//       and inventing a synchroniser here would be inventing a problem.
//
//   (2) CONFIGURATION IS SAMPLED WHEN A REQUEST IS ACCEPTED.
//       The `cfg_*` inputs are live wires that software may change at any time. On
//       acceptance they are copied into `*_q` registers, and the frame in flight uses
//       ONLY the copies. A write during a frame therefore affects the NEXT frame.
//       Every use of configuration below reads a `_q`; if any line read a live
//       `cfg_*` while busy, that would be the bug 20.7 injects on purpose.
//
//   (3) BIT ORDER IS AN ALIGNMENT PROBLEM, NOT A DATAPATH PROBLEM.
//       There is exactly ONE shift direction in this module: left. MOSI is always
//       `tx_sr[DATA_W-1]`; MISO always shifts into `rx_sr[0]`. LSB-first is produced
//       by REVERSING the word as it is loaded and reversing it back as it is read
//       out. A second shift direction would double the datapath state and every
//       proof about it; one reversal function at each boundary does not.
//
//   (4) TIMING GENERATION IS SEPARATE FROM DATAPATH CONTROL.
//       A divider produces `tick`, one pulse per SCLK half-period. The FSM consumes
//       ticks and knows nothing about system-clock counting. A mode decoder turns
//       tick-driven SCLK transitions into `launch_event` / `sample_event`, and the
//       datapath knows nothing about CPOL or CPHA. Three concerns, three places.
//
// THE ONE IDEA WORTH THE WHOLE CHAPTER
//
//   CPOL and CPHA are not a four-entry lookup table. Every SCLK transition is either
//   LEADING (away from the idle level) or TRAILING (back to it). CPOL decides which
//   physical direction "away from idle" is; CPHA decides which of the two edge kinds
//   launches and which samples. Given an edge index, `leading` is just its parity --
//   and everything else follows:
//
//       leading      = ~edge_i[0]                 even transitions leave idle
//       sample_event = cpha ? trailing : leading
//       launch_event = cpha ? leading  : trailing
//
//   That is the entire mode-handling logic. CPOL never appears in it, because CPOL
//   only decides the LEVEL SCLK idles at -- and the edge index already counts from
//   that level.

`timescale 1ns/1ps

module spi_capstone_ctrl #(
    // Datapath width. The widest transfer the hardware can carry; `cfg_width`
    // selects any width from MIN_WIDTH to DATA_W per request.
    parameter int DATA_W    = 16,
    // Narrowest legal transfer. Requests below this are rejected, not clamped:
    // silently transferring a different number of bits than asked for is the
    // failure mode REQ-ERR-001 exists to prevent.
    parameter int MIN_WIDTH = 4,
    parameter int NDEV      = 4
) (
    input  wire                clk,
    // Asynchronous, active low. Release is assumed already synchronised to `clk`
    // by the integrator -- see 20.4, and note that an RTL simulation cannot check
    // that assumption.
    input  wire                rst_n,

    // ---- live configuration (sampled at request acceptance, commitment 2) ----
    input  wire                cfg_cpol,
    input  wire                cfg_cpha,
    input  wire                cfg_lsb_first,
    input  wire [4:0]          cfg_width,     // MIN_WIDTH .. DATA_W
    input  wire [7:0]          cfg_div,       // SCLK half-period = cfg_div + 1 clk
    input  wire [1:0]          cfg_dev,       // which chip select
    input  wire [3:0]          cfg_lead,      // CS-low to first edge, in half-periods
    input  wire [3:0]          cfg_lag,       // last edge to CS-high, in half-periods
    input  wire [3:0]          cfg_idle,      // enforced turnaround, in half-periods

    // ---- request ----
    input  wire                start,         // one-cycle pulse; ignored unless idle
    input  wire [DATA_W-1:0]   tx_data,
    input  wire                abort,         // one-cycle pulse; abandons the frame

    // ---- response ----
    output wire                busy,
    output reg                 done,          // exactly one cycle per COMPLETED frame
    output reg                 cfg_err,       // one cycle; request rejected, not run
    output reg  [DATA_W-1:0]   rx_data,       // right-aligned, valid from `done`
    output reg  [4:0]          bits_done,     // bits sampled; survives an abort

    // ---- pins ----
    output reg                 sclk,
    output reg                 mosi,
    output reg  [NDEV-1:0]     cs_n,
    input  wire                miso
);

    // ---------------------------------------------------------------------------
    // Control states.
    //
    // S_GAP is part of `busy` ON PURPOSE. The turnaround of REQ-TIM-004 is not a
    // soft timing hope the integrator has to honour -- it is enforced by refusing
    // the next request until it has elapsed. A controller that reported itself idle
    // during the turnaround would push that obligation onto software, which is
    // exactly where such requirements get lost.
    // ---------------------------------------------------------------------------
    localparam [2:0] S_IDLE = 3'd0,
                     S_LEAD = 3'd1,   // CS asserted, counting lead half-periods
                     S_XFER = 3'd2,   // generating 2*N SCLK transitions
                     S_LAG  = 3'd3,   // SCLK back at idle, CS still asserted
                     S_GAP  = 3'd4;   // CS released, enforcing turnaround

    reg [2:0] st;
    assign busy = (st != S_IDLE);

    // ---- captured configuration (commitment 2) --------------------------------
    reg         cpol_q, cpha_q, lsb_q;
    reg [4:0]   width_q;
    reg [7:0]   div_q;
    reg [1:0]   dev_q;
    reg [3:0]   lead_q, lag_q, idle_q;

    // ---- timing generator (commitment 4) --------------------------------------
    reg  [7:0]  div_cnt;
    wire        tick = (div_cnt == 8'd0);

    // ---- transfer state -------------------------------------------------------
    reg  [5:0]  edge_i;      // 0 .. 2*width_q-1, counts SCLK transitions
    reg  [4:0]  tx_idx;      // how many bits have been PRESENTED on MOSI
    reg  [3:0]  phase_cnt;   // shared lead / lag / gap half-period counter
    reg  [DATA_W-1:0] tx_sr, rx_sr;
    reg         aborted;

    // ---- mode decode (the one idea, commitment 4) -----------------------------
    // `edge_i` is the index of the transition ABOUT to be made. Even indices move
    // SCLK away from its idle level (leading); odd indices return it (trailing).
    // 2*N transitions per frame, computed in SIX bits. `width_q << 1` would be a
    // FIVE-bit expression -- the shift does not widen its operand -- so a 16-bit
    // transfer would compute 32 truncated to 0 and the frame would end on its first
    // edge. Widen first, then shift: the concatenation is the fix, not the cast.
    wire [5:0] edge_total = {1'b0, width_q} << 1;

    wire leading      = ~edge_i[0];
    wire sample_event = cpha_q ? ~leading :  leading;
    wire launch_event = cpha_q ?  leading : ~leading;

    // Reject a width the datapath cannot honestly carry. Checked on the LIVE inputs,
    // because a request is accepted or refused before anything is captured.
    wire cfg_bad = (cfg_width < MIN_WIDTH) || (cfg_width > DATA_W);

    // ---- bit reversal: the whole of bit-order handling (commitment 3) ----------
    function automatic [DATA_W-1:0] rev;
        input [DATA_W-1:0] v;
        integer b;
        begin
            rev = {DATA_W{1'b0}};
            for (b = 0; b < DATA_W; b = b + 1)
                rev[DATA_W-1-b] = v[b];
        end
    endfunction

    // The word as the shift register wants it: MSB-first sends bit width-1 first, so
    // left-align it; LSB-first sends bit 0 first, so reverse it (which puts bit 0 at
    // the top and makes the SAME left shift produce the opposite order).
    function automatic [DATA_W-1:0] load_align;
        input [DATA_W-1:0] v;
        input [4:0]        w;
        input              lsb;
        begin
            load_align = lsb ? rev(v) : (v << (DATA_W - w));
        end
    endfunction

    // The received stream, un-aligned. First bit received sits at rx_sr[w-1].
    // MSB-first wants it at bit w-1 already; LSB-first wants it at bit 0.
    function automatic [DATA_W-1:0] store_align;
        input [DATA_W-1:0] v;
        input [4:0]        w;
        input              lsb;
        reg   [DATA_W-1:0] m;
        begin
            m = (w >= DATA_W) ? {DATA_W{1'b1}}
                              : ((({{(DATA_W-1){1'b0}}, 1'b1}) << w) - 1'b1);
            store_align = lsb ? ((rev(v) >> (DATA_W - w)) & m) : (v & m);
        end
    endfunction

    // Scratch value for the acceptance cycle. It is a VARIABLE assigned with a
    // BLOCKING assignment inside the clocked block below, and both of those choices
    // are load-bearing:
    //
    //   * A part-select of a function CALL does not parse in Verilog or
    //     SystemVerilog -- `load_align(...)[DATA_W-1]` is a syntax error -- so the
    //     aligned word has to have a name before a bit of it can be taken.
    //
    //   * The obvious name is a wire: `wire [15:0] ld = load_align(...)`. That
    //     compiles, reads correctly from a testbench, and IS WRONG HERE. A
    //     continuous assignment whose right-hand side is a function call settles a
    //     time step after its inputs change, so a clock edge in that same step
    //     samples the PREVIOUS value. Measured: the shift register loaded 0000 while
    //     the wire read d000 one cycle later, and every transmitted bit was zero
    //     while every other captured field was correct.
    //
    // Computing it in the clocked block removes the question: the function is called
    // at the edge, with the values the edge itself sampled. Note that this is the
    // opposite of Module 19's VHDL lesson -- a variable is wrong for state another
    // process reads, and right for a value used and discarded inside one invocation.
    reg [DATA_W-1:0] ld;

    always @(posedge clk or negedge rst_n) begin
        if (!rst_n) begin
            // REQ-RST-001. Pins first: every chip select released, SCLK at a DEFINED
            // level. Note that this level is 0, not cfg_cpol -- reset does not read
            // configuration. From the first clock after release SCLK follows CPOL
            // (see the S_IDLE arm), and 20.7 asks what a slave sees in between.
            st        <= S_IDLE;
            cs_n      <= {NDEV{1'b1}};
            sclk      <= 1'b0;
            mosi      <= 1'b0;
            done      <= 1'b0;
            cfg_err   <= 1'b0;
            rx_data   <= {DATA_W{1'b0}};
            bits_done <= 5'd0;
            div_cnt   <= 8'd0;
            edge_i    <= 6'd0;
            tx_idx    <= 5'd0;
            phase_cnt <= 4'd0;
            tx_sr     <= {DATA_W{1'b0}};
            rx_sr     <= {DATA_W{1'b0}};
            aborted   <= 1'b0;
            cpol_q    <= 1'b0; cpha_q <= 1'b0; lsb_q <= 1'b0;
            width_q   <= 5'd0; div_q  <= 8'd0; dev_q <= 2'd0;
            lead_q    <= 4'd0; lag_q  <= 4'd0; idle_q <= 4'd0;
        end else begin
            // Single-cycle outputs default low; the arms below re-raise them.
            done    <= 1'b0;
            cfg_err <= 1'b0;

            // ---- divider: runs only while a frame is in progress ---------------
            if (st == S_IDLE) div_cnt <= 8'd0;
            else if (tick)    div_cnt <= div_q;
            else              div_cnt <= div_cnt - 8'd1;

            // ---- abort (REQ-ABT-001) ------------------------------------------
            // Honoured from any active state. SCLK returns to the captured idle
            // level and the frame proceeds to S_LAG so the selected device still
            // gets its release time -- dropping CS instantly would violate the
            // device's own hold requirement to punish the controller's caller.
            if (abort && (st == S_LEAD || st == S_XFER)) begin
                aborted   <= 1'b1;
                sclk      <= cpol_q;
                st        <= S_LAG;
                phase_cnt <= lag_q;
            end else begin
                case (st)

                S_IDLE: begin
                    // Idle SCLK tracks LIVE cpol, so the pin is correct before any
                    // frame starts. This is the one place a live cfg_* is read, and
                    // it is safe precisely because no frame is in flight.
                    sclk <= cfg_cpol;
                    if (start) begin
                        if (cfg_bad) begin
                            // REQ-ERR-001: refuse, report, stay idle. No frame, no
                            // `done`, no chip select, nothing captured.
                            cfg_err <= 1'b1;
                        end else begin
                            // COMPUTED ONLY AFTER THE WIDTH IS KNOWN LEGAL. Here that
                            // ordering looks like tidiness; in VHDL the same code
                            // shifted by DATA_W - 17 = -1 and aborted the simulation,
                            // because `shift_left` takes a NATURAL. Verilog wrapped
                            // the shift amount, produced a garbage word, discarded it
                            // with the refused request, and said nothing. The bug was
                            // in both: a function evaluated outside its domain.
                            ld = load_align(tx_data, cfg_width, cfg_lsb_first);
                            cpol_q  <= cfg_cpol;  cpha_q <= cfg_cpha;
                            lsb_q   <= cfg_lsb_first;
                            width_q <= cfg_width; div_q  <= cfg_div;
                            dev_q   <= cfg_dev;
                            lead_q  <= cfg_lead;  lag_q  <= cfg_lag;
                            idle_q  <= cfg_idle;

                            cs_n              <= {NDEV{1'b1}};
                            cs_n[cfg_dev]     <= 1'b0;
                            sclk              <= cfg_cpol;
                            rx_sr             <= {DATA_W{1'b0}};
                            // CPHA=0 needs bit 0 valid BEFORE the first leading edge,
                            // so present it now, with CS. CPHA=1 launches on that
                            // edge instead, so drive a defined 0 until it does --
                            // which is why MOSI visibly differs between the two modes
                            // during the lead time.
                            // PRESENT THEN SHIFT, uniformly. CPHA=0 owes the
                            // device a valid bit before its first sampling edge, so
                            // the pre-launch happens here and consumes bit 0; CPHA=1
                            // launches on that edge instead, so MOSI holds a defined
                            // 0 and bit 0 is still pending.
                            mosi              <= cfg_cpha ? 1'b0 : ld[DATA_W-1];
                            tx_sr             <= cfg_cpha ? ld : (ld << 1);
                            tx_idx            <= cfg_cpha ? 5'd0 : 5'd1;
                            edge_i            <= 6'd0;
                            bits_done         <= 5'd0;
                            aborted           <= 1'b0;
                            div_cnt           <= cfg_div;
                            phase_cnt         <= cfg_lead;
                            st                <= S_LEAD;
                        end
                    end
                end

                S_LEAD: begin
                    if (tick) begin
                        if (phase_cnt == 4'd0) st <= S_XFER;
                        else phase_cnt <= phase_cnt - 4'd1;
                    end
                end

                S_XFER: begin
                    if (tick) begin
                        sclk <= ~sclk;

                        // Sample BEFORE the launch below, so that when a single
                        // transition both samples one bit and launches the next
                        // (it never does in a legal mode, but a mutation in 20.7
                        // makes it happen) the order is defined rather than lucky.
                        if (sample_event && (bits_done < width_q)) begin
                            rx_sr     <= {rx_sr[DATA_W-2:0], miso};
                            bits_done <= bits_done + 5'd1;
                        end

                        // The SAME two lines as the pre-launch above: present the
                        // top bit, then shift it away. The alternative -- shift first
                        // and present the new top -- is correct for CPHA=0 and off by
                        // one bit for CPHA=1, because CPHA=1 has no pre-launch to
                        // have consumed bit 0. Measured: modes 0 and 2 passed at all
                        // four widths while modes 1 and 3 returned every word shifted
                        // up one position, in both directions at once.
                        if (launch_event && (tx_idx < width_q)) begin
                            mosi   <= tx_sr[DATA_W-1];
                            tx_sr  <= {tx_sr[DATA_W-2:0], 1'b0};
                            tx_idx <= tx_idx + 5'd1;
                        end

                        // 2*width_q transitions per frame: N leading + N trailing.
                        // An even count is why SCLK is guaranteed back at the idle
                        // level when the frame ends (REQ-MODE-001) -- it is a
                        // property of the count, not a separate assignment.
                        if (edge_i == edge_total - 6'd1) begin
                            st        <= S_LAG;
                            phase_cnt <= lag_q;
                        end else begin
                            edge_i <= edge_i + 6'd1;
                        end
                    end
                end

                S_LAG: begin
                    if (tick) begin
                        if (phase_cnt == 4'd0) begin
                            cs_n      <= {NDEV{1'b1}};
                            st        <= S_GAP;
                            phase_cnt <= idle_q;
                            // REQ-FUNC-005 / 006: a completed frame publishes its
                            // data and pulses `done` exactly once. An aborted frame
                            // does NEITHER -- rx_data keeps its previous value, so
                            // a caller that ignores `done` reads stale data rather
                            // than a plausible-looking partial word.
                            if (!aborted) begin
                                rx_data <= store_align(rx_sr, width_q, lsb_q);
                                done    <= 1'b1;
                            end
                        end else begin
                            phase_cnt <= phase_cnt - 4'd1;
                        end
                    end
                end

                S_GAP: begin
                    if (tick) begin
                        if (phase_cnt == 4'd0) st <= S_IDLE;
                        else phase_cnt <= phase_cnt - 4'd1;
                    end
                end

                default: st <= S_IDLE;
                endcase
            end
        end
    end

endmodule

5. The Implementation — Verilog-2001

The conversion required six line changes, all of them keyword substitutions, and none in the testbench. It was not accepted on that basis: it was compiled with -g2001 and simulated independently, because a mechanical conversion that compiles is not evidence that it behaves the same.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   parameter int      →  parameter integer      3 occurrences
   function automatic →  function               3 occurrences
Azvya Education Pvt. Ltd.VLSI Mentor
spi_capstone_ctrl.v — the same hardware in Verilog-2001, compiled and simulated independently
// spi_capstone_ctrl.sv
//
// THE CAPSTONE CONTROLLER. One design, carried from Chapter 20.1's specification
// through to Chapter 20.7's design review. Every later chapter verifies THIS module;
// none of them redesigns it.
//
// WHAT IT IS
//
//   A configurable SPI master. Four modes, either bit order, a transfer width chosen
//   per request, a programmable clock divider, four chip selects, and programmable
//   lead / lag / turnaround times.
//
// THE FOUR ARCHITECTURAL COMMITMENTS
//
//   These are decided here, defended in 20.2, constrained in 20.4, and questioned in
//   20.7. They are stated at the top because every line below follows from them.
//
//   (1) EVERY REGISTER IN THIS MODULE IS CLOCKED BY `clk`.
//       SCLK is an OUTPUT WAVEFORM this module generates, not a clock it uses. No
//       internal state is clocked by SCLK, so there is no internal clock-domain
//       crossing to synchronise. That is a deliberate architectural choice with a
//       cost (20.4 §"what this buys and what it costs") -- it is not an oversight,
//       and inventing a synchroniser here would be inventing a problem.
//
//   (2) CONFIGURATION IS SAMPLED WHEN A REQUEST IS ACCEPTED.
//       The `cfg_*` inputs are live wires that software may change at any time. On
//       acceptance they are copied into `*_q` registers, and the frame in flight uses
//       ONLY the copies. A write during a frame therefore affects the NEXT frame.
//       Every use of configuration below reads a `_q`; if any line read a live
//       `cfg_*` while busy, that would be the bug 20.7 injects on purpose.
//
//   (3) BIT ORDER IS AN ALIGNMENT PROBLEM, NOT A DATAPATH PROBLEM.
//       There is exactly ONE shift direction in this module: left. MOSI is always
//       `tx_sr[DATA_W-1]`; MISO always shifts into `rx_sr[0]`. LSB-first is produced
//       by REVERSING the word as it is loaded and reversing it back as it is read
//       out. A second shift direction would double the datapath state and every
//       proof about it; one reversal function at each boundary does not.
//
//   (4) TIMING GENERATION IS SEPARATE FROM DATAPATH CONTROL.
//       A divider produces `tick`, one pulse per SCLK half-period. The FSM consumes
//       ticks and knows nothing about system-clock counting. A mode decoder turns
//       tick-driven SCLK transitions into `launch_event` / `sample_event`, and the
//       datapath knows nothing about CPOL or CPHA. Three concerns, three places.
//
// THE ONE IDEA WORTH THE WHOLE CHAPTER
//
//   CPOL and CPHA are not a four-entry lookup table. Every SCLK transition is either
//   LEADING (away from the idle level) or TRAILING (back to it). CPOL decides which
//   physical direction "away from idle" is; CPHA decides which of the two edge kinds
//   launches and which samples. Given an edge index, `leading` is just its parity --
//   and everything else follows:
//
//       leading      = ~edge_i[0]                 even transitions leave idle
//       sample_event = cpha ? trailing : leading
//       launch_event = cpha ? leading  : trailing
//
//   That is the entire mode-handling logic. CPOL never appears in it, because CPOL
//   only decides the LEVEL SCLK idles at -- and the edge index already counts from
//   that level.

`timescale 1ns/1ps

module spi_capstone_ctrl #(
    // Datapath width. The widest transfer the hardware can carry; `cfg_width`
    // selects any width from MIN_WIDTH to DATA_W per request.
    parameter integer DATA_W    = 16,
    // Narrowest legal transfer. Requests below this are rejected, not clamped:
    // silently transferring a different number of bits than asked for is the
    // failure mode REQ-ERR-001 exists to prevent.
    parameter integer MIN_WIDTH = 4,
    parameter integer NDEV      = 4
) (
    input  wire                clk,
    // Asynchronous, active low. Release is assumed already synchronised to `clk`
    // by the integrator -- see 20.4, and note that an RTL simulation cannot check
    // that assumption.
    input  wire                rst_n,

    // ---- live configuration (sampled at request acceptance, commitment 2) ----
    input  wire                cfg_cpol,
    input  wire                cfg_cpha,
    input  wire                cfg_lsb_first,
    input  wire [4:0]          cfg_width,     // MIN_WIDTH .. DATA_W
    input  wire [7:0]          cfg_div,       // SCLK half-period = cfg_div + 1 clk
    input  wire [1:0]          cfg_dev,       // which chip select
    input  wire [3:0]          cfg_lead,      // CS-low to first edge, in half-periods
    input  wire [3:0]          cfg_lag,       // last edge to CS-high, in half-periods
    input  wire [3:0]          cfg_idle,      // enforced turnaround, in half-periods

    // ---- request ----
    input  wire                start,         // one-cycle pulse; ignored unless idle
    input  wire [DATA_W-1:0]   tx_data,
    input  wire                abort,         // one-cycle pulse; abandons the frame

    // ---- response ----
    output wire                busy,
    output reg                 done,          // exactly one cycle per COMPLETED frame
    output reg                 cfg_err,       // one cycle; request rejected, not run
    output reg  [DATA_W-1:0]   rx_data,       // right-aligned, valid from `done`
    output reg  [4:0]          bits_done,     // bits sampled; survives an abort

    // ---- pins ----
    output reg                 sclk,
    output reg                 mosi,
    output reg  [NDEV-1:0]     cs_n,
    input  wire                miso
);

    // ---------------------------------------------------------------------------
    // Control states.
    //
    // S_GAP is part of `busy` ON PURPOSE. The turnaround of REQ-TIM-004 is not a
    // soft timing hope the integrator has to honour -- it is enforced by refusing
    // the next request until it has elapsed. A controller that reported itself idle
    // during the turnaround would push that obligation onto software, which is
    // exactly where such requirements get lost.
    // ---------------------------------------------------------------------------
    localparam [2:0] S_IDLE = 3'd0,
                     S_LEAD = 3'd1,   // CS asserted, counting lead half-periods
                     S_XFER = 3'd2,   // generating 2*N SCLK transitions
                     S_LAG  = 3'd3,   // SCLK back at idle, CS still asserted
                     S_GAP  = 3'd4;   // CS released, enforcing turnaround

    reg [2:0] st;
    assign busy = (st != S_IDLE);

    // ---- captured configuration (commitment 2) --------------------------------
    reg         cpol_q, cpha_q, lsb_q;
    reg [4:0]   width_q;
    reg [7:0]   div_q;
    reg [1:0]   dev_q;
    reg [3:0]   lead_q, lag_q, idle_q;

    // ---- timing generator (commitment 4) --------------------------------------
    reg  [7:0]  div_cnt;
    wire        tick = (div_cnt == 8'd0);

    // ---- transfer state -------------------------------------------------------
    reg  [5:0]  edge_i;      // 0 .. 2*width_q-1, counts SCLK transitions
    reg  [4:0]  tx_idx;      // how many bits have been PRESENTED on MOSI
    reg  [3:0]  phase_cnt;   // shared lead / lag / gap half-period counter
    reg  [DATA_W-1:0] tx_sr, rx_sr;
    reg         aborted;

    // ---- mode decode (the one idea, commitment 4) -----------------------------
    // `edge_i` is the index of the transition ABOUT to be made. Even indices move
    // SCLK away from its idle level (leading); odd indices return it (trailing).
    // 2*N transitions per frame, computed in SIX bits. `width_q << 1` would be a
    // FIVE-bit expression -- the shift does not widen its operand -- so a 16-bit
    // transfer would compute 32 truncated to 0 and the frame would end on its first
    // edge. Widen first, then shift: the concatenation is the fix, not the cast.
    wire [5:0] edge_total = {1'b0, width_q} << 1;

    wire leading      = ~edge_i[0];
    wire sample_event = cpha_q ? ~leading :  leading;
    wire launch_event = cpha_q ?  leading : ~leading;

    // Reject a width the datapath cannot honestly carry. Checked on the LIVE inputs,
    // because a request is accepted or refused before anything is captured.
    wire cfg_bad = (cfg_width < MIN_WIDTH) || (cfg_width > DATA_W);

    // ---- bit reversal: the whole of bit-order handling (commitment 3) ----------
    function [DATA_W-1:0] rev;
        input [DATA_W-1:0] v;
        integer b;
        begin
            rev = {DATA_W{1'b0}};
            for (b = 0; b < DATA_W; b = b + 1)
                rev[DATA_W-1-b] = v[b];
        end
    endfunction

    // The word as the shift register wants it: MSB-first sends bit width-1 first, so
    // left-align it; LSB-first sends bit 0 first, so reverse it (which puts bit 0 at
    // the top and makes the SAME left shift produce the opposite order).
    function [DATA_W-1:0] load_align;
        input [DATA_W-1:0] v;
        input [4:0]        w;
        input              lsb;
        begin
            load_align = lsb ? rev(v) : (v << (DATA_W - w));
        end
    endfunction

    // The received stream, un-aligned. First bit received sits at rx_sr[w-1].
    // MSB-first wants it at bit w-1 already; LSB-first wants it at bit 0.
    function [DATA_W-1:0] store_align;
        input [DATA_W-1:0] v;
        input [4:0]        w;
        input              lsb;
        reg   [DATA_W-1:0] m;
        begin
            m = (w >= DATA_W) ? {DATA_W{1'b1}}
                              : ((({{(DATA_W-1){1'b0}}, 1'b1}) << w) - 1'b1);
            store_align = lsb ? ((rev(v) >> (DATA_W - w)) & m) : (v & m);
        end
    endfunction

    // Scratch value for the acceptance cycle. It is a VARIABLE assigned with a
    // BLOCKING assignment inside the clocked block below, and both of those choices
    // are load-bearing:
    //
    //   * A part-select of a function CALL does not parse in Verilog or
    //     SystemVerilog -- `load_align(...)[DATA_W-1]` is a syntax error -- so the
    //     aligned word has to have a name before a bit of it can be taken.
    //
    //   * The obvious name is a wire: `wire [15:0] ld = load_align(...)`. That
    //     compiles, reads correctly from a testbench, and IS WRONG HERE. A
    //     continuous assignment whose right-hand side is a function call settles a
    //     time step after its inputs change, so a clock edge in that same step
    //     samples the PREVIOUS value. Measured: the shift register loaded 0000 while
    //     the wire read d000 one cycle later, and every transmitted bit was zero
    //     while every other captured field was correct.
    //
    // Computing it in the clocked block removes the question: the function is called
    // at the edge, with the values the edge itself sampled. Note that this is the
    // opposite of Module 19's VHDL lesson -- a variable is wrong for state another
    // process reads, and right for a value used and discarded inside one invocation.
    reg [DATA_W-1:0] ld;

    always @(posedge clk or negedge rst_n) begin
        if (!rst_n) begin
            // REQ-RST-001. Pins first: every chip select released, SCLK at a DEFINED
            // level. Note that this level is 0, not cfg_cpol -- reset does not read
            // configuration. From the first clock after release SCLK follows CPOL
            // (see the S_IDLE arm), and 20.7 asks what a slave sees in between.
            st        <= S_IDLE;
            cs_n      <= {NDEV{1'b1}};
            sclk      <= 1'b0;
            mosi      <= 1'b0;
            done      <= 1'b0;
            cfg_err   <= 1'b0;
            rx_data   <= {DATA_W{1'b0}};
            bits_done <= 5'd0;
            div_cnt   <= 8'd0;
            edge_i    <= 6'd0;
            tx_idx    <= 5'd0;
            phase_cnt <= 4'd0;
            tx_sr     <= {DATA_W{1'b0}};
            rx_sr     <= {DATA_W{1'b0}};
            aborted   <= 1'b0;
            cpol_q    <= 1'b0; cpha_q <= 1'b0; lsb_q <= 1'b0;
            width_q   <= 5'd0; div_q  <= 8'd0; dev_q <= 2'd0;
            lead_q    <= 4'd0; lag_q  <= 4'd0; idle_q <= 4'd0;
        end else begin
            // Single-cycle outputs default low; the arms below re-raise them.
            done    <= 1'b0;
            cfg_err <= 1'b0;

            // ---- divider: runs only while a frame is in progress ---------------
            if (st == S_IDLE) div_cnt <= 8'd0;
            else if (tick)    div_cnt <= div_q;
            else              div_cnt <= div_cnt - 8'd1;

            // ---- abort (REQ-ABT-001) ------------------------------------------
            // Honoured from any active state. SCLK returns to the captured idle
            // level and the frame proceeds to S_LAG so the selected device still
            // gets its release time -- dropping CS instantly would violate the
            // device's own hold requirement to punish the controller's caller.
            if (abort && (st == S_LEAD || st == S_XFER)) begin
                aborted   <= 1'b1;
                sclk      <= cpol_q;
                st        <= S_LAG;
                phase_cnt <= lag_q;
            end else begin
                case (st)

                S_IDLE: begin
                    // Idle SCLK tracks LIVE cpol, so the pin is correct before any
                    // frame starts. This is the one place a live cfg_* is read, and
                    // it is safe precisely because no frame is in flight.
                    sclk <= cfg_cpol;
                    if (start) begin
                        if (cfg_bad) begin
                            // REQ-ERR-001: refuse, report, stay idle. No frame, no
                            // `done`, no chip select, nothing captured.
                            cfg_err <= 1'b1;
                        end else begin
                            // COMPUTED ONLY AFTER THE WIDTH IS KNOWN LEGAL. Here that
                            // ordering looks like tidiness; in VHDL the same code
                            // shifted by DATA_W - 17 = -1 and aborted the simulation,
                            // because `shift_left` takes a NATURAL. Verilog wrapped
                            // the shift amount, produced a garbage word, discarded it
                            // with the refused request, and said nothing. The bug was
                            // in both: a function evaluated outside its domain.
                            ld = load_align(tx_data, cfg_width, cfg_lsb_first);
                            cpol_q  <= cfg_cpol;  cpha_q <= cfg_cpha;
                            lsb_q   <= cfg_lsb_first;
                            width_q <= cfg_width; div_q  <= cfg_div;
                            dev_q   <= cfg_dev;
                            lead_q  <= cfg_lead;  lag_q  <= cfg_lag;
                            idle_q  <= cfg_idle;

                            cs_n              <= {NDEV{1'b1}};
                            cs_n[cfg_dev]     <= 1'b0;
                            sclk              <= cfg_cpol;
                            rx_sr             <= {DATA_W{1'b0}};
                            // CPHA=0 needs bit 0 valid BEFORE the first leading edge,
                            // so present it now, with CS. CPHA=1 launches on that
                            // edge instead, so drive a defined 0 until it does --
                            // which is why MOSI visibly differs between the two modes
                            // during the lead time.
                            // PRESENT THEN SHIFT, uniformly. CPHA=0 owes the
                            // device a valid bit before its first sampling edge, so
                            // the pre-launch happens here and consumes bit 0; CPHA=1
                            // launches on that edge instead, so MOSI holds a defined
                            // 0 and bit 0 is still pending.
                            mosi              <= cfg_cpha ? 1'b0 : ld[DATA_W-1];
                            tx_sr             <= cfg_cpha ? ld : (ld << 1);
                            tx_idx            <= cfg_cpha ? 5'd0 : 5'd1;
                            edge_i            <= 6'd0;
                            bits_done         <= 5'd0;
                            aborted           <= 1'b0;
                            div_cnt           <= cfg_div;
                            phase_cnt         <= cfg_lead;
                            st                <= S_LEAD;
                        end
                    end
                end

                S_LEAD: begin
                    if (tick) begin
                        if (phase_cnt == 4'd0) st <= S_XFER;
                        else phase_cnt <= phase_cnt - 4'd1;
                    end
                end

                S_XFER: begin
                    if (tick) begin
                        sclk <= ~sclk;

                        // Sample BEFORE the launch below, so that when a single
                        // transition both samples one bit and launches the next
                        // (it never does in a legal mode, but a mutation in 20.7
                        // makes it happen) the order is defined rather than lucky.
                        if (sample_event && (bits_done < width_q)) begin
                            rx_sr     <= {rx_sr[DATA_W-2:0], miso};
                            bits_done <= bits_done + 5'd1;
                        end

                        // The SAME two lines as the pre-launch above: present the
                        // top bit, then shift it away. The alternative -- shift first
                        // and present the new top -- is correct for CPHA=0 and off by
                        // one bit for CPHA=1, because CPHA=1 has no pre-launch to
                        // have consumed bit 0. Measured: modes 0 and 2 passed at all
                        // four widths while modes 1 and 3 returned every word shifted
                        // up one position, in both directions at once.
                        if (launch_event && (tx_idx < width_q)) begin
                            mosi   <= tx_sr[DATA_W-1];
                            tx_sr  <= {tx_sr[DATA_W-2:0], 1'b0};
                            tx_idx <= tx_idx + 5'd1;
                        end

                        // 2*width_q transitions per frame: N leading + N trailing.
                        // An even count is why SCLK is guaranteed back at the idle
                        // level when the frame ends (REQ-MODE-001) -- it is a
                        // property of the count, not a separate assignment.
                        if (edge_i == edge_total - 6'd1) begin
                            st        <= S_LAG;
                            phase_cnt <= lag_q;
                        end else begin
                            edge_i <= edge_i + 6'd1;
                        end
                    end
                end

                S_LAG: begin
                    if (tick) begin
                        if (phase_cnt == 4'd0) begin
                            cs_n      <= {NDEV{1'b1}};
                            st        <= S_GAP;
                            phase_cnt <= idle_q;
                            // REQ-FUNC-005 / 006: a completed frame publishes its
                            // data and pulses `done` exactly once. An aborted frame
                            // does NEITHER -- rx_data keeps its previous value, so
                            // a caller that ignores `done` reads stale data rather
                            // than a plausible-looking partial word.
                            if (!aborted) begin
                                rx_data <= store_align(rx_sr, width_q, lsb_q);
                                done    <= 1'b1;
                            end
                        end else begin
                            phase_cnt <= phase_cnt - 4'd1;
                        end
                    end
                end

                S_GAP: begin
                    if (tick) begin
                        if (phase_cnt == 4'd0) st <= S_IDLE;
                        else phase_cnt <= phase_cnt - 4'd1;
                    end
                end

                default: st <= S_IDLE;
                endcase
            end
        end
    end

endmodule

Every .v file in this module was scanned for logic, always_ff, always_comb, typedef, enum, interface, class, assert property, covergroup, $urandom and $countones. Excluding prose inside comments, the count of such constructs in executable code is zero across all four files.

6. The Implementation — VHDL

Azvya Education Pvt. Ltd.VLSI Mentor
spi_capstone_ctrl.vhd — the same hardware in VHDL-2008, where a signal is a register and a variable is not
-- spi_capstone_ctrl.vhd
--
-- The capstone controller in VHDL-2008. Same specification, same architecture, same
-- four commitments as the SystemVerilog and Verilog-2001 versions -- and the same
-- cycle-level behaviour, which is checked by comparing all three transcripts rather
-- than by reading the three files side by side.
--
-- THE VHDL-SPECIFIC DISCIPLINE THIS FILE FOLLOWS
--
--   Module 19 cost four separate defects to one root cause, so the rule is written
--   down here rather than rediscovered:
--
--     A SIGNAL is the VHDL counterpart of a non-blocking register assignment. Its new
--     value is visible only after the process suspends, which is exactly what a
--     flip-flop does, and exactly what `<=` means in the other two languages.
--
--     A VARIABLE is visible to the next statement. It is therefore correct for a value
--     computed and consumed inside ONE invocation, and wrong for anything that
--     represents state. Every register below is a signal. The only variable is `ld`,
--     which exists because the alignment function must be evaluated with the values
--     the edge itself sampled -- the same reason the other two languages compute it
--     inside the clocked block instead of on a wire.
--
--   Also observed and deliberately avoided here:
--     * every design unit carries its own context clause (they do not carry over)
--     * every function formal is CONSTRAINED, so no slice direction is inherited
--     * arithmetic goes through `resize`/`shift_*`, never a bare `&`, so no
--       concatenation is ambiguous
--     * no identifier is a reserved word, and none collides with another under case
--       folding -- VHDL is case-INSENSITIVE, so `S_LEAD` and `s_lead` are one name

library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

entity spi_capstone_ctrl is
    generic (
        DATA_W    : integer := 16;
        MIN_WIDTH : integer := 4;
        NDEV      : integer := 4
    );
    port (
        clk           : in  std_logic;
        rst_n         : in  std_logic;

        cfg_cpol      : in  std_logic;
        cfg_cpha      : in  std_logic;
        cfg_lsb_first : in  std_logic;
        cfg_width     : in  std_logic_vector(4 downto 0);
        cfg_div       : in  std_logic_vector(7 downto 0);
        cfg_dev       : in  std_logic_vector(1 downto 0);
        cfg_lead      : in  std_logic_vector(3 downto 0);
        cfg_lag       : in  std_logic_vector(3 downto 0);
        cfg_idle      : in  std_logic_vector(3 downto 0);

        start         : in  std_logic;
        tx_data       : in  std_logic_vector(DATA_W-1 downto 0);
        abort         : in  std_logic;

        busy          : out std_logic;
        done          : out std_logic;
        cfg_err       : out std_logic;
        rx_data       : out std_logic_vector(DATA_W-1 downto 0);
        bits_done     : out std_logic_vector(4 downto 0);

        sclk          : out std_logic;
        mosi          : out std_logic;
        cs_n          : out std_logic_vector(NDEV-1 downto 0);
        miso          : in  std_logic
    );
end entity spi_capstone_ctrl;

architecture rtl of spi_capstone_ctrl is

    type state_t is (S_IDLE, S_LEAD, S_XFER, S_LAG, S_GAP);
    signal st : state_t;

    -- captured configuration
    signal cpol_q, cpha_q, lsb_q : std_logic := '0';
    signal width_q               : unsigned(4 downto 0) := (others => '0');
    signal div_q                 : unsigned(7 downto 0) := (others => '0');
    signal dev_q                 : unsigned(1 downto 0) := (others => '0');
    signal lead_q, lag_q, idle_q : unsigned(3 downto 0) := (others => '0');

    -- Timing generator. The declaration initialisers on the registers in this
    -- architecture mirror their RESET values exactly. They exist so that the
    -- concurrent decode below does not evaluate `div_cnt = 0` against an
    -- uninitialised value at time zero, which NUMERIC_STD reports as
    -- "metavalue detected, returning FALSE". They are simulation initial values:
    -- synthesis ignores them, and the asynchronous reset -- not the initialiser --
    -- is what defines this hardware's start-up state.
    signal div_cnt : unsigned(7 downto 0) := (others => '0');
    signal tick    : std_logic;

    -- transfer state
    signal edge_i     : unsigned(5 downto 0) := (others => '0');
    signal edge_total : unsigned(5 downto 0);
    signal tx_idx     : unsigned(4 downto 0) := (others => '0');
    signal bits_cnt   : unsigned(4 downto 0) := (others => '0');
    signal phase_cnt  : unsigned(3 downto 0) := (others => '0');
    signal tx_sr      : std_logic_vector(DATA_W-1 downto 0);
    signal rx_sr      : std_logic_vector(DATA_W-1 downto 0);
    signal aborted    : std_logic;

    -- mode decode
    signal leading      : std_logic;
    signal sample_event : std_logic;
    signal launch_event : std_logic;
    signal cfg_bad      : std_logic;

    -- registered pin / status state. The ports are driven concurrently from these, so
    -- every signal below has exactly ONE driver. VHDL's std_logic is a RESOLVED type:
    -- a second driver would not be an error, it would quietly produce 'X'.
    signal sclk_r    : std_logic;
    signal mosi_r    : std_logic;
    signal cs_r      : std_logic_vector(NDEV-1 downto 0);
    signal done_r    : std_logic;
    signal err_r     : std_logic;
    signal rx_data_r : std_logic_vector(DATA_W-1 downto 0);

    -- Reverse the whole word. The formal is CONSTRAINED: an unconstrained formal takes
    -- its range from the actual, and a concatenation or literal can hand it an
    -- ASCENDING range, after which every index in here refers to the wrong end.
    function rev (v : std_logic_vector(DATA_W-1 downto 0))
        return std_logic_vector is
        variable r : std_logic_vector(DATA_W-1 downto 0);
    begin
        for b in 0 to DATA_W-1 loop
            r(DATA_W-1-b) := v(b);
        end loop;
        return r;
    end function rev;

    -- Low-w-bit mask, built by iteration rather than by `2**w - 1`. The arithmetic
    -- form has to be evaluated in a type wide enough to hold 2**DATA_W, and getting
    -- that wrong for w = DATA_W produces an all-zero mask that makes every masked
    -- comparison succeed against zero.
    function maskw (w : integer) return unsigned is
        variable m : unsigned(DATA_W-1 downto 0);
    begin
        m := (others => '0');
        for b in 0 to DATA_W-1 loop
            if b < w then
                m(b) := '1';
            end if;
        end loop;
        return m;
    end function maskw;

    function load_align (v : std_logic_vector(DATA_W-1 downto 0);
                         w : integer; lsb : std_logic)
        return std_logic_vector is
    begin
        if lsb = '1' then
            return rev(v);
        else
            return std_logic_vector(shift_left(unsigned(v), DATA_W - w));
        end if;
    end function load_align;

    function store_align (v : std_logic_vector(DATA_W-1 downto 0);
                          w : integer; lsb : std_logic)
        return std_logic_vector is
    begin
        if lsb = '1' then
            -- The shift already clears the vacated high bits, so no mask is needed.
            return std_logic_vector(shift_right(unsigned(rev(v)), DATA_W - w));
        else
            return std_logic_vector(unsigned(v) and maskw(w));
        end if;
    end function store_align;

begin

    -- Concurrent decode. Each of these reads a signal, so inside the clocked process
    -- below they hold their PRE-EDGE values -- which is what makes them behave like
    -- the `wire` declarations in the other two languages rather than like variables.
    tick         <= '1' when div_cnt = 0 else '0';
    edge_total   <= resize(width_q, 6) sll 1;
    leading      <= not edge_i(0);
    sample_event <= (not leading) when cpha_q = '1' else leading;
    launch_event <= leading       when cpha_q = '1' else (not leading);
    cfg_bad      <= '1' when (to_integer(unsigned(cfg_width)) < MIN_WIDTH)
                          or (to_integer(unsigned(cfg_width)) > DATA_W) else '0';

    busy      <= '0' when st = S_IDLE else '1';
    sclk      <= sclk_r;
    mosi      <= mosi_r;
    cs_n      <= cs_r;
    done      <= done_r;
    cfg_err   <= err_r;
    rx_data   <= rx_data_r;
    bits_done <= std_logic_vector(bits_cnt);

    process (clk, rst_n)
        -- The ONE variable in this design. It is computed and consumed within a single
        -- invocation, which is precisely what a variable is for.
        variable ld : std_logic_vector(DATA_W-1 downto 0);
    begin
        if rst_n = '0' then
            st        <= S_IDLE;
            cs_r      <= (others => '1');
            sclk_r    <= '0';
            mosi_r    <= '0';
            done_r    <= '0';
            err_r     <= '0';
            rx_data_r <= (others => '0');
            bits_cnt  <= (others => '0');
            div_cnt   <= (others => '0');
            edge_i    <= (others => '0');
            tx_idx    <= (others => '0');
            phase_cnt <= (others => '0');
            tx_sr     <= (others => '0');
            rx_sr     <= (others => '0');
            aborted   <= '0';
            cpol_q    <= '0'; cpha_q <= '0'; lsb_q <= '0';
            width_q   <= (others => '0');
            div_q     <= (others => '0');
            dev_q     <= (others => '0');
            lead_q    <= (others => '0');
            lag_q     <= (others => '0');
            idle_q    <= (others => '0');

        elsif rising_edge(clk) then
            done_r <= '0';
            err_r  <= '0';

            if st = S_IDLE then
                div_cnt <= (others => '0');
            elsif tick = '1' then
                div_cnt <= div_q;
            else
                div_cnt <= div_cnt - 1;
            end if;

            if abort = '1' and (st = S_LEAD or st = S_XFER) then
                aborted   <= '1';
                sclk_r    <= cpol_q;
                st        <= S_LAG;
                phase_cnt <= lag_q;
            else
                case st is

                when S_IDLE =>
                    sclk_r <= cfg_cpol;
                    if start = '1' then
                        if cfg_bad = '1' then
                            err_r <= '1';
                        else
                            -- COMPUTED ONLY AFTER THE WIDTH IS KNOWN LEGAL, and the
                            -- third language is why. `load_align` shifts by
                            -- DATA_W - w; for the rejected width 17 that is -1, and
                            -- `shift_left` takes a NATURAL:
                            --
                            --   ** Fatal: value -1 outside of NATURAL range
                            --
                            -- Verilog computes the same expression as an unsigned
                            -- wrap, produces a garbage word, discards it because the
                            -- request is refused, and says nothing. Both languages
                            -- were evaluating a function outside its domain; only one
                            -- of them admitted it. Validating first is not a VHDL
                            -- workaround -- it is the fix, in all three.
                            ld := load_align(tx_data,
                                             to_integer(unsigned(cfg_width)),
                                             cfg_lsb_first);
                            cpol_q  <= cfg_cpol;
                            cpha_q  <= cfg_cpha;
                            lsb_q   <= cfg_lsb_first;
                            width_q <= unsigned(cfg_width);
                            div_q   <= unsigned(cfg_div);
                            dev_q   <= unsigned(cfg_dev);
                            lead_q  <= unsigned(cfg_lead);
                            lag_q   <= unsigned(cfg_lag);
                            idle_q  <= unsigned(cfg_idle);

                            cs_r    <= (others => '1');
                            cs_r(to_integer(unsigned(cfg_dev))) <= '0';
                            sclk_r  <= cfg_cpol;
                            rx_sr   <= (others => '0');
                            if cfg_cpha = '1' then
                                mosi_r <= '0';
                                tx_sr  <= ld;
                                tx_idx <= (others => '0');
                            else
                                mosi_r <= ld(DATA_W-1);
                                tx_sr  <= std_logic_vector(shift_left(unsigned(ld), 1));
                                tx_idx <= to_unsigned(1, 5);
                            end if;
                            bits_cnt  <= (others => '0');
                            edge_i    <= (others => '0');
                            aborted   <= '0';
                            div_cnt   <= unsigned(cfg_div);
                            phase_cnt <= unsigned(cfg_lead);
                            st        <= S_LEAD;
                        end if;
                    end if;

                when S_LEAD =>
                    if tick = '1' then
                        if phase_cnt = 0 then
                            st <= S_XFER;
                        else
                            phase_cnt <= phase_cnt - 1;
                        end if;
                    end if;

                when S_XFER =>
                    if tick = '1' then
                        sclk_r <= not sclk_r;

                        if sample_event = '1' and bits_cnt < width_q then
                            rx_sr    <= rx_sr(DATA_W-2 downto 0) & miso;
                            bits_cnt <= bits_cnt + 1;
                        end if;

                        if launch_event = '1' and tx_idx < width_q then
                            mosi_r <= tx_sr(DATA_W-1);
                            tx_sr  <= tx_sr(DATA_W-2 downto 0) & '0';
                            tx_idx <= tx_idx + 1;
                        end if;

                        if edge_i = edge_total - 1 then
                            st        <= S_LAG;
                            phase_cnt <= lag_q;
                        else
                            edge_i <= edge_i + 1;
                        end if;
                    end if;

                when S_LAG =>
                    if tick = '1' then
                        if phase_cnt = 0 then
                            cs_r      <= (others => '1');
                            st        <= S_GAP;
                            phase_cnt <= idle_q;
                            if aborted = '0' then
                                rx_data_r <= store_align(rx_sr,
                                                         to_integer(width_q),
                                                         lsb_q);
                                done_r    <= '1';
                            end if;
                        else
                            phase_cnt <= phase_cnt - 1;
                        end if;
                    end if;

                when S_GAP =>
                    if tick = '1' then
                        if phase_cnt = 0 then
                            st <= S_IDLE;
                        else
                            phase_cnt <= phase_cnt - 1;
                        end if;
                    end if;

                end case;
            end if;
        end if;
    end process;

end architecture rtl;

7. The Directed Testbench

284 checks per language, in nine groups. Two things in it are worth reading before the code.

The expected values come from specification arithmetic — a mask and a reversal — and never from asking the controller what it did. And the bench uses a pin-level device model rather than a loopback, for the reason Chapter 20.1 §10 gave: a loopback cannot see a bit-order fault, and Chapter 20.7 proves it by injecting one.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_capstone_ctrl_tb.sv — the directed suite: nine groups, 284 checks, and a device model rather than a loopback
// spi_capstone_ctrl_tb.sv
//
// Chapter 20.3 -- the DIRECTED suite for the capstone controller.
//
// This bench answers one question per group, and every expected value in it is
// derived from Chapter 20.1's SPECIFICATION by arithmetic -- never by asking the
// controller what it did. The distinction matters most for bit order, and there is a
// trap in it worth stating before any code:
//
//   A LOOPBACK TEST CANNOT DETECT A BIT-ORDER BUG.
//
//   Tie MOSI to MISO and the received word equals the transmitted word in BOTH bit
//   orders -- the reversal that goes out is undone coming back. A suite built on
//   loopback reports a clean pass with the MSB/LSB logic wired backwards. So this
//   bench uses a PIN-LEVEL SLAVE MODEL with an asymmetric return word, and checks two
//   independent observations per transfer:
//
//     * what the master received   (depends on bit order)
//     * what the slave received    (depends on bit order, the other way round)
//
//   Chapter 20.7 injects exactly that bug and shows the loopback check surviving it.
//
// WHAT THE SLAVE MODEL IS AND IS NOT
//
//   It is the mirror-image DEVICE: it samples on the edge the master launches on, and
//   launches on the edge the master samples on. It is NOT the source of the expected
//   values -- those come from `mask` and `revw` below, which know nothing about either
//   the controller or the model. If the model and the controller were both wrong in
//   the same direction the arithmetic would still catch it.
//
//   And per 20.4's rule about models: the model is configured for every mode the
//   controller supports. A slave model that only works in mode 0 tests mode 0 and
//   reports on four.

`timescale 1ns/1ps

module spi_capstone_ctrl_tb;

    localparam integer DATA_W = 16;

    // ---- DUT interface -------------------------------------------------------
    reg         clk, rst_n;
    reg         cfg_cpol, cfg_cpha, cfg_lsb;
    reg  [4:0]  cfg_width;
    reg  [7:0]  cfg_div;
    reg  [1:0]  cfg_dev;
    reg  [3:0]  cfg_lead, cfg_lag, cfg_idle;
    reg         start, abort;
    reg  [15:0] tx_data;
    wire        busy, done, cfg_err;
    wire [15:0] rx_data;
    wire [4:0]  bits_done;
    wire        sclk, mosi;
    wire [3:0]  cs_n;
    wire        miso;

    // ---- bookkeeping ---------------------------------------------------------
    integer n_chk, n_err, n_neg_ok;
    integer cyc;
    reg [8*22-1:0] gname;

    spi_capstone_ctrl #(.DATA_W(16), .MIN_WIDTH(4), .NDEV(4)) dut (
        .clk(clk), .rst_n(rst_n),
        .cfg_cpol(cfg_cpol), .cfg_cpha(cfg_cpha), .cfg_lsb_first(cfg_lsb),
        .cfg_width(cfg_width), .cfg_div(cfg_div), .cfg_dev(cfg_dev),
        .cfg_lead(cfg_lead), .cfg_lag(cfg_lag), .cfg_idle(cfg_idle),
        .start(start), .tx_data(tx_data), .abort(abort),
        .busy(busy), .done(done), .cfg_err(cfg_err),
        .rx_data(rx_data), .bits_done(bits_done),
        .sclk(sclk), .mosi(mosi), .cs_n(cs_n), .miso(miso)
    );

    always #5 clk = ~clk;

    // =========================================================================
    // Specification arithmetic. The independent oracle.
    // =========================================================================

    // Low-w-bit mask. Built by shifting a widened 1 -- `(1 << w) - 1` in a 16-bit
    // expression gives 0 for w = 16, which would mask every checked value to zero and
    // make the whole suite pass vacuously.
    function [15:0] mask;
        input [4:0] w;
        reg [16:0] one;
        begin
            one  = 17'd1;
            mask = ((one << w) - 17'd1);
        end
    endfunction

    // Reverse the low w bits of v, leaving the rest zero. This is the ONLY place the
    // bench encodes what LSB-first means.
    function [15:0] revw;
        input [15:0] v;
        input [4:0]  w;
        integer b;
        begin
            revw = 16'd0;
            for (b = 0; b < 16; b = b + 1)
                if (b < w) revw[w-1-b] = v[b];
        end
    endfunction

    // What the SLAVE should end up holding, given what the master was asked to send.
    function [15:0] exp_slave_rx;
        input [15:0] tx; input [4:0] w; input lsb;
        begin
            exp_slave_rx = lsb ? revw(tx & mask(w), w) : (tx & mask(w));
        end
    endfunction

    // What the MASTER should end up holding, given what the slave returns.
    function [15:0] exp_master_rx;
        input [15:0] sw; input [4:0] w; input lsb;
        begin
            exp_master_rx = lsb ? revw(sw & mask(w), w) : (sw & mask(w));
        end
    endfunction

    // =========================================================================
    // Pin-level slave model -- configurable for all four modes.
    // =========================================================================
    reg        slv_cpol, slv_cpha;
    reg [15:0] slv_word, slv_sr, slv_rx;
    reg [4:0]  slv_w, slv_idx, slv_nrx;
    reg        slv_miso, force_x;
    reg        lead_s;

    wire cs_any = ~(&cs_n);
    assign miso = force_x ? 1'bx : slv_miso;

    // A select fell: load, and for CPHA=0 present the first bit immediately -- the
    // master will sample it on the very first edge, so there is no edge left to
    // launch it on.
    always @(posedge cs_any) begin
        slv_sr   = slv_word << (16 - slv_w);
        slv_rx   = 16'd0;
        slv_idx  = 5'd0;
        slv_nrx  = 5'd0;
        if (!slv_cpha) begin
            slv_miso = slv_sr[15];
            slv_sr   = slv_sr << 1;
            slv_idx  = 5'd1;
        end else begin
            slv_miso = 1'b0;
        end
    end

    always @(sclk) begin
        if (cs_any === 1'b1) begin
            // A transition AWAY from the idle level is leading; back to it, trailing.
            // CPOL never appears anywhere else in this model.
            lead_s = (sclk !== slv_cpol);
            // NOT exchanged: a slave uses the SAME edges as its master. Both
            // sample on one edge and both change their output on the other -- that is
            // what makes the link work off a single clock. The model's two conditions
            // are therefore identical to the controller's, not mirrored.
            if (slv_cpha ? !lead_s : lead_s) begin
                if (slv_nrx < slv_w) begin
                    slv_rx  = {slv_rx[14:0], mosi};
                    slv_nrx = slv_nrx + 5'd1;
                end
            end
            if (slv_cpha ? lead_s : !lead_s) begin
                if (slv_idx < slv_w) begin
                    slv_miso = slv_sr[15];
                    slv_sr   = slv_sr << 1;
                    slv_idx  = slv_idx + 5'd1;
                end
            end
        end
    end

    // =========================================================================
    // Pin monitors. Every timing number reported by this bench is MEASURED at the
    // pins, never computed from the configuration -- 20.4 depends on that, and
    // Module 19.4 is why (hand-derived arithmetic was one cycle out).
    // =========================================================================
    reg        cs_d, sclk_d;
    integer    t_cs_fall, t_cs_rise, t_first_edge, t_last_edge, t_prev_rise;
    integer    n_edge, m_half, m_lead, m_lag, m_gap, n_overlap;
    integer    n_low;

    always @(posedge clk) begin
        if (!rst_n) begin
            cyc <= 0; cs_d <= 1'b0; sclk_d <= 1'b0; n_edge <= 0; n_overlap <= 0;
        end else begin
            cyc <= cyc + 1;

            n_low = (cs_n[0] ? 0 : 1) + (cs_n[1] ? 0 : 1)
                  + (cs_n[2] ? 0 : 1) + (cs_n[3] ? 0 : 1);
            if (n_low > 1) n_overlap <= n_overlap + 1;

            if (cs_any && !cs_d) begin
                t_cs_fall <= cyc;
                m_gap     <= cyc - t_prev_rise;
                n_edge    <= 0;
            end
            if (!cs_any && cs_d) begin
                t_cs_rise  <= cyc;
                t_prev_rise <= cyc;
                m_lag      <= cyc - t_last_edge;
            end
                // `cs_d` as well as `cs_any`: an edge is only a FRAME edge if a device
                // was ALREADY selected last cycle. A transition in the very cycle the
                // select falls is SCLK reaching its new idle level, not a clocking
                // edge -- the controller parks SCLK and asserts CS together, so when
                // the previous idle level differed the two coincide.
                //
                // Measured cost of omitting `cs_d`: the first transaction after reset
                // with CPOL=1 counted 17 edges instead of 16 in VHDL and 16 in
                // SystemVerilog, because a one-cycle difference in reset-release
                // timing decided whether the re-park landed inside the window. The
                // received data was correct in both. With the gate the measurement no
                // longer depends on that phase at all.
            if (cs_any && cs_d && (sclk !== sclk_d)) begin
                if (n_edge == 0) begin
                    t_first_edge <= cyc;
                    m_lead       <= cyc - t_cs_fall;
                end else begin
                    m_half <= cyc - t_last_edge;
                end
                t_last_edge <= cyc;
                n_edge      <= n_edge + 1;
            end
            cs_d   <= cs_any;
            sclk_d <= sclk;
        end
    end

    // =========================================================================
    // Checkers
    // =========================================================================

    // `!==` not `!=`: an X anywhere in `got` must FAIL. With `!=` an X compares as
    // "unknown", the `if` treats it as false, and the check silently passes -- which
    // is the mechanism group 8 deliberately exercises.
    task chk16;
        input [8*22-1:0] nm; input [15:0] got; input [15:0] exp;
        begin
            n_chk = n_chk + 1;
            if (got !== exp) begin
                n_err = n_err + 1;
                $display("    FAIL %0s: got %04h expected %04h", nm, got, exp);
            end
        end
    endtask

    task chki;
        input [8*22-1:0] nm; input integer got; input integer exp;
        begin
            n_chk = n_chk + 1;
            if (got !== exp) begin
                n_err = n_err + 1;
                $display("    FAIL %0s: got %0d expected %0d", nm, got, exp);
            end
        end
    endtask

    // =========================================================================
    // Stimulus helpers
    // =========================================================================
    task set_cfg;
        input cpol_i; input cpha_i; input lsb_i; input [4:0] w;
        input [7:0] dv; input [1:0] dv_n;
        input [3:0] ld; input [3:0] lg; input [3:0] id;
        begin
            cfg_cpol = cpol_i; cfg_cpha = cpha_i; cfg_lsb = lsb_i;
            cfg_width = w; cfg_div = dv; cfg_dev = dv_n;
            cfg_lead = ld; cfg_lag = lg; cfg_idle = id;
            slv_cpol = cpol_i; slv_cpha = cpha_i; slv_w = w;
        end
    endtask

    // EVERY input is driven on the NEGEDGE, and this is not a style preference.
    //
    // Driving `start` on the posedge the controller samples it on is a race: the
    // assignment that clears it and the controller's clocked block are both in the
    // same region of the same time step, and their order is not defined by the
    // language. Measured cost of getting this wrong: the FIRST frame of the run was
    // accepted and every later one was silently dropped, so 31 of 32 transfers
    // compared a stale `rx_data` against a fresh expectation and the suite reported
    // a wall of mismatches that had nothing to do with the controller.
    //
    // A negedge-driven `start` is high across exactly one sampling edge, and nothing
    // the bench does can collide with the edge that matters. This is Chapter 16.3's
    // clocking-block discipline written out by hand.
    task fire;
        input [15:0] d;
        begin
            @(negedge clk);
            tx_data = d;
            start   = 1'b1;
            @(negedge clk);
            start   = 1'b0;
        end
    endtask

    // Bounded wait. Observation happens on the NEGEDGE so every non-blocking update
    // from the clock edge has settled -- reading a one-cycle `done` right after
    // `@(posedge clk)` reads its pre-edge value and misses it entirely.
    task wait_idle;
        input integer maxc; output got_done; output integer took;
        integer g; reg seen;
        begin
            g = 0; seen = 1'b0;
            while (g < maxc) begin
                @(negedge clk);
                g = g + 1;
                if (done) seen = 1'b1;
                if (!busy && seen) g = maxc;
                else if (!busy && g > 4) g = maxc;
            end
            got_done = seen;
            took     = g;
        end
    endtask

    // =========================================================================
    // Test sequence
    // =========================================================================
    integer mi, oi, wi, k;
    reg [4:0]  WID [0:3];
    reg [15:0] TXP [0:3];
    reg [15:0] SWP [0:3];
    reg        gd;
    integer    tk, grp_err;
    reg [15:0] rx_before;

    initial begin
        clk = 1'b0; rst_n = 1'b0;
        start = 1'b0; abort = 1'b0; tx_data = 16'd0;
        cfg_cpol = 1'b0; cfg_cpha = 1'b0; cfg_lsb = 1'b0;
        cfg_width = 5'd8; cfg_div = 8'd1; cfg_dev = 2'd0;
        cfg_lead = 4'd1; cfg_lag = 4'd1; cfg_idle = 4'd1;
        slv_cpol = 1'b0; slv_cpha = 1'b0; slv_w = 5'd8;
        slv_word = 16'd0; slv_miso = 1'b0; force_x = 1'b0;
        slv_sr = 16'd0; slv_rx = 16'd0; slv_idx = 5'd0; slv_nrx = 5'd0;
        n_chk = 0; n_err = 0; n_neg_ok = 0;
        cyc = 0; n_edge = 0; n_overlap = 0;
        t_cs_fall = 0; t_cs_rise = 0; t_first_edge = 0; t_last_edge = 0;
        t_prev_rise = 0; m_half = 0; m_lead = 0; m_lag = 0; m_gap = 0;
        cs_d = 1'b0; sclk_d = 1'b0;

        WID[0] = 5'd4;  WID[1] = 5'd8;  WID[2] = 5'd13; WID[3] = 5'd16;
        TXP[0] = 16'hB39D; TXP[1] = 16'hB39D; TXP[2] = 16'hB39D; TXP[3] = 16'hB39D;
        SWP[0] = 16'h4E7A; SWP[1] = 16'h4E7A; SWP[2] = 16'h4E7A; SWP[3] = 16'h4E7A;

        $display("=== Chapter 20.3 -- capstone controller, directed suite ===");

        repeat (4) @(posedge clk);
        rst_n = 1'b1;
        repeat (2) @(posedge clk);

        // -----------------------------------------------------------------
        // GROUP 1 -- every mode, both bit orders, four widths.
        // The question: does the controller implement the specification's mode
        // and bit-order semantics, for every width the datapath claims?
        // -----------------------------------------------------------------
        $display("  G1 mode x bit-order x width");
        for (mi = 0; mi < 4; mi = mi + 1) begin
            for (oi = 0; oi < 2; oi = oi + 1) begin
                grp_err = n_err;
                for (wi = 0; wi < 4; wi = wi + 1) begin
                    set_cfg(mi[1], mi[0], oi[0], WID[wi], 8'd1, 2'd1, 4'd2, 4'd2, 4'd2);
                    slv_word = SWP[wi] & mask(WID[wi]);
                    fire(TXP[wi]);
                    wait_idle(4000, gd, tk);
                    chki ("done pulsed",   gd ? 1 : 0, 1);
                    chk16("master rx",     rx_data,
                          exp_master_rx(SWP[wi] & mask(WID[wi]), WID[wi], oi[0]));
                    chk16("slave rx",      slv_rx & mask(WID[wi]),
                          exp_slave_rx(TXP[wi], WID[wi], oi[0]));
                    chki ("sclk edges",    n_edge, 2 * WID[wi]);
                    chki ("bits sampled",  bits_done, WID[wi]);
                    chki ("sclk parked",   (sclk === mi[1]) ? 1 : 0, 1);
                    chki ("no cs overlap", n_overlap, 0);
                end
                $display("    mode %0d %0s : rx(w4,w8,w13,w16) checked, %0s",
                         mi, oi[0] ? "lsb" : "msb",
                         (n_err == grp_err) ? "all match" : "MISMATCH");
            end
        end

        // -----------------------------------------------------------------
        // GROUP 2 -- measured pin intervals.
        // The question: are the lead / lag / turnaround requirements MET at the
        // pins, and what is the actual margin? The requirement is ">=", the
        // measurement is exact, and the difference is reported rather than assumed.
        // -----------------------------------------------------------------
        $display("  G2 measured pin intervals, in clk cycles");
        $display("    div lead lag idle | t_half t_lead t_lag t_gap");
        for (k = 0; k < 4; k = k + 1) begin
            set_cfg(1'b0, 1'b0, 1'b0, 5'd8,
                    (k == 0) ? 8'd0 : (k == 1) ? 8'd1 : (k == 2) ? 8'd3 : 8'd1,
                    2'd2,
                    (k == 3) ? 4'd5 : 4'd2,
                    (k == 3) ? 4'd4 : 4'd2,
                    (k == 3) ? 4'd6 : 4'd3);
            slv_word = 16'h5A;
            // Two frames: the second one's gap is the interval between them.
            fire(16'h33); wait_idle(4000, gd, tk);
            fire(16'h33); wait_idle(4000, gd, tk);
            $display("    %3d %4d %3d %4d | %6d %6d %5d %5d",
                     cfg_div, cfg_lead, cfg_lag, cfg_idle, m_half, m_lead, m_lag, m_gap);
            // REQ-TIM-001 is an equality and is checked as one.
            chki("t_half exact", m_half, cfg_div + 1);

            // REQ-TIM-002/003/004 are ">=" requirements, and they are checked as
            // stated -- but a ">=" check is WEAK: an implementation that loses one
            // half-period of lead still satisfies it whenever cfg_lead > 0. So each
            // interval is ALSO checked against the exact closed form the measurements
            // above establish, which is what actually catches a one-tick regression.
            if (m_lead < (cfg_lead * (cfg_div + 1))) begin
                n_err = n_err + 1; $display("    FAIL lead below requirement");
            end
            n_chk = n_chk + 1;
            if (m_lag < (cfg_lag * (cfg_div + 1))) begin
                n_err = n_err + 1; $display("    FAIL lag below requirement");
            end
            n_chk = n_chk + 1;
            if (m_gap < (cfg_idle * (cfg_div + 1))) begin
                n_err = n_err + 1; $display("    FAIL turnaround below requirement");
            end
            n_chk = n_chk + 1;

            // Exact implementation forms, in system clocks. The "+2" and "+1" terms
            // are the ticks the state machine spends LEAVING a phase, and they are
            // margin above the requirement, not part of it:
            //
            //   t_lead = (cfg_lead + 2) * t_half   one tick to exit S_LEAD, one more
            //                                      before the first transition
            //   t_lag  = (cfg_lag  + 1) * t_half   one tick to exit S_LAG
            //   t_gap  = (cfg_idle + 1) * t_half + 2
            //
            // The gap is the ONLY interval here with a term that is not a multiple of
            // the half-period, and the reason is worth naming: those 2 cycles are
            // S_IDLE plus ONE CYCLE OF THE REQUESTER'S OWN LATENCY. With no request
            // queue, the bench cannot post the next transfer until `busy` falls, so
            // the measured gap is the design's floor PLUS however long the requester
            // took to react. 20.2 defends that trade-off and 20.7 questions it.
            chki("t_lead exact", m_lead, (cfg_lead + 2) * (cfg_div + 1));
            chki("t_lag exact",  m_lag,  (cfg_lag  + 1) * (cfg_div + 1));
            chki("t_gap exact",  m_gap,  (cfg_idle + 1) * (cfg_div + 1) + 2);
        end

        // -----------------------------------------------------------------
        // GROUP 3 -- illegal configuration is REFUSED, at both boundaries.
        // -----------------------------------------------------------------
        $display("  G3 illegal width");
        for (k = 0; k < 4; k = k + 1) begin
            set_cfg(1'b0, 1'b0, 1'b0,
                    (k == 0) ? 5'd3 : (k == 1) ? 5'd17 : (k == 2) ? 5'd4 : 5'd16,
                    8'd1, 2'd0, 4'd2, 4'd2, 4'd2);
            slv_word = 16'hFFFF & mask(cfg_width);
            slv_w    = (k < 2) ? 5'd8 : cfg_width;
            fire(16'h5555);
            // `fire` returns on the negedge immediately after the sampling edge, so a
            // one-cycle output is readable RIGHT HERE. One more negedge and `cfg_err`
            // has already gone low -- which is how a correct rejection reads as a
            // missing one.
            if (k < 2) begin
                chki("cfg_err raised", cfg_err ? 1 : 0, 1);
                chki("stayed idle",    busy ? 1 : 0,    0);
                chki("no cs asserted", cs_any ? 1 : 0,  0);
                $display("    width %2d rejected: cfg_err=1 busy=0 cs=idle", cfg_width);
            end else begin
                chki("accepted",       busy ? 1 : 0,    1);
                chki("no cfg_err",     cfg_err ? 1 : 0, 0);
                wait_idle(4000, gd, tk);
                chki("done pulsed",    gd ? 1 : 0, 1);
                $display("    width %2d accepted: cfg_err=0 done=1", cfg_width);
            end
        end

        // -----------------------------------------------------------------
        // GROUP 4 -- a request while busy has NO effect (REQ-FUNC-007).
        // -----------------------------------------------------------------
        $display("  G4 start while busy");
        set_cfg(1'b0, 1'b0, 1'b0, 5'd8, 8'd2, 2'd0, 4'd2, 4'd2, 4'd2);
        slv_word = 16'h96;
        fire(16'hA5);
        repeat (6) @(posedge clk);
        fire(16'h3C);                        // ignored: the frame in flight owns the bus
        wait_idle(4000, gd, tk);
        chk16("rx from FIRST word", rx_data, exp_master_rx(16'h96, 5'd8, 1'b0));
        chk16("slave got FIRST word", slv_rx & mask(5'd8), exp_slave_rx(16'hA5, 5'd8, 1'b0));
        chki ("edges of one frame", n_edge, 16);
        $display("    second request ignored: one frame, 16 edges, slave saw a5");

        // -----------------------------------------------------------------
        // GROUP 5 -- reset in the middle of a frame (REQ-RST-002).
        // -----------------------------------------------------------------
        $display("  G5 reset during a frame");
        set_cfg(1'b1, 1'b0, 1'b0, 5'd16, 8'd3, 2'd3, 4'd2, 4'd2, 4'd2);
        slv_word = 16'hBEEF;
        fire(16'hDEAD);
        repeat (20) @(posedge clk);
        chki("frame really started", cs_any ? 1 : 0, 1);
        @(negedge clk); rst_n = 1'b0;
        @(negedge clk);
        chki("all selects released", cs_any ? 1 : 0, 0);
        chki("no done on reset",     done ? 1 : 0,   0);
        chki("not busy",             busy ? 1 : 0,   0);
        chki("sclk defined 0",       (sclk === 1'b0) ? 1 : 0, 1);
        $display("    reset mid-frame: cs released, no done, sclk=0 (not cpol)");
        @(negedge clk); rst_n = 1'b1; repeat (3) @(posedge clk);

        // -----------------------------------------------------------------
        // GROUP 6 -- abort (REQ-ABT-001).
        // -----------------------------------------------------------------
        $display("  G6 abort mid-frame");
        set_cfg(1'b0, 1'b0, 1'b0, 5'd16, 8'd1, 2'd0, 4'd2, 4'd2, 4'd2);
        slv_word = 16'h1234;
        fire(16'h4321);
        wait_idle(4000, gd, tk);             // a clean frame first, to set rx_data
        rx_before = rx_data;
        slv_word  = 16'hFFFF;
        fire(16'h0F0F);
        repeat (24) @(posedge clk);
        @(negedge clk); abort = 1'b1; @(negedge clk); abort = 1'b0;
        wait_idle(4000, gd, tk);
        chki ("no done on abort",   gd ? 1 : 0, 0);
        chk16("rx_data unchanged",  rx_data, rx_before);
        chki ("selects released",   cs_any ? 1 : 0, 0);
        chki ("sclk back at idle",  (sclk === 1'b0) ? 1 : 0, 1);
        if (bits_done == 5'd0 || bits_done >= 5'd16) begin
            n_err = n_err + 1;
            $display("    FAIL bits_done not partial: %0d", bits_done);
        end
        n_chk = n_chk + 1;
        $display("    abort: done=0, rx held %04h, bits_done=%0d of 16",
                 rx_data, bits_done);

        // -----------------------------------------------------------------
        // GROUP 7 -- configuration captured at acceptance (commitment 2).
        // The frame must ignore a mid-flight rewrite of EVERY field it uses.
        // -----------------------------------------------------------------
        $display("  G7 configuration captured at acceptance");
        set_cfg(1'b0, 1'b0, 1'b0, 5'd8, 8'd2, 2'd0, 4'd2, 4'd2, 4'd2);
        slv_word = 16'h7E;
        fire(16'h81);
        repeat (5) @(posedge clk);
        // Rewrite everything to values that would produce a different frame.
        cfg_cpol = 1'b1; cfg_cpha = 1'b1; cfg_lsb = 1'b1;
        cfg_width = 5'd4; cfg_div = 8'd7; cfg_dev = 2'd3;
        wait_idle(4000, gd, tk);
        chk16("rx used captured cfg", rx_data, exp_master_rx(16'h7E, 5'd8, 1'b0));
        chki ("edges of captured width", n_edge, 16);
        chki ("captured device kept", cs_n[0] === 1'b1 ? 1 : 0, 1);
        $display("    mid-frame rewrite ignored: 16 edges, msb-order rx=%04h", rx_data);
        set_cfg(1'b0, 1'b0, 1'b0, 5'd8, 8'd1, 2'd0, 4'd2, 4'd2, 4'd2);

        // -----------------------------------------------------------------
        // GROUP 8 -- an X on MISO must NOT be able to produce a pass.
        // This is a NEGATIVE test: the checker is EXPECTED to fail here, and the
        // bench counts that as evidence rather than as a defect.
        // -----------------------------------------------------------------
        $display("  G8 X on MISO cannot slip through");
        slv_word = 16'h5A;
        force_x  = 1'b1;
        fire(16'hA5);
        wait_idle(4000, gd, tk);
        n_chk = n_chk + 1;
        if (rx_data !== exp_master_rx(16'h5A, 5'd8, 1'b0)) begin
            n_neg_ok = n_neg_ok + 1;
            $display("    X reached rx_data and the !== comparison rejected it");
        end else begin
            n_err = n_err + 1;
            $display("    FAIL an all-X rx_data compared EQUAL -- checker is blind");
        end
        force_x = 1'b0;

        // -----------------------------------------------------------------
        // GROUP 9 -- prove the checkers can fail at all.
        // A suite that has never printed FAIL has not been shown to be able to.
        // -----------------------------------------------------------------
        $display("  G9 deliberately wrong expectations must FAIL");
        // The pattern here is 9d, and the reason it is not a5 or 5a is worth more
        // than the test it enables:
        //
        //   a5 = 10100101   reversed = 10100101
        //   5a = 01011010   reversed = 01011010
        //
        // The two most-used test bytes in digital design are both EIGHT-BIT
        // PALINDROMES. Neither can distinguish MSB-first from LSB-first, so a suite
        // built on them reports a clean pass with the bit-order logic inverted -- and
        // the self-test of the oracle below reported a defect in the ORACLE when the
        // only defect was the choice of stimulus. 9d reverses to b9.
        k = 0;
        slv_word = 16'h9D;
        fire(16'h9D);
        wait_idle(4000, gd, tk);
        n_chk = n_chk + 1;
        if (rx_data !== (exp_master_rx(16'h9D, 5'd8, 1'b0) ^ 16'h0001)) k = k + 1;
        else begin n_err = n_err + 1; $display("    FAIL rx checker accepted a wrong value"); end
        n_chk = n_chk + 1;
        if (n_edge !== 15) k = k + 1;
        else begin n_err = n_err + 1; $display("    FAIL edge checker accepted a wrong count"); end
        n_chk = n_chk + 1;
        if (exp_master_rx(16'h9D, 5'd8, 1'b1) !== exp_master_rx(16'h9D, 5'd8, 1'b0)) k = k + 1;
        else begin n_err = n_err + 1; $display("    FAIL oracle is blind to bit order"); end
        n_neg_ok = n_neg_ok + k;
        $display("    %0d of 3 wrong expectations rejected; oracle separates 9d from b9", k);

        $display("=== SUMMARY checks=%0d negatives=%0d failures=%0d : %0s ===",
                 n_chk, n_neg_ok, n_err, (n_err == 0) ? "PASS" : "FAIL");
        $finish;
    end

    // Global watchdog. A hung frame must end the run with a FAILURE, not with a
    // simulator that stops printing.
    initial begin
        #4000000;
        $display("    FATAL global timeout -- a frame never completed");
        $display("=== SUMMARY checks=%0d negatives=%0d failures=%0d : FAIL ===",
                 n_chk, n_neg_ok, n_err + 1);
        $finish;
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_capstone_ctrl_tb.v — the same bench in Verilog-2001 — it needed no changes at all
// spi_capstone_ctrl_tb.sv
//
// Chapter 20.3 -- the DIRECTED suite for the capstone controller.
//
// This bench answers one question per group, and every expected value in it is
// derived from Chapter 20.1's SPECIFICATION by arithmetic -- never by asking the
// controller what it did. The distinction matters most for bit order, and there is a
// trap in it worth stating before any code:
//
//   A LOOPBACK TEST CANNOT DETECT A BIT-ORDER BUG.
//
//   Tie MOSI to MISO and the received word equals the transmitted word in BOTH bit
//   orders -- the reversal that goes out is undone coming back. A suite built on
//   loopback reports a clean pass with the MSB/LSB logic wired backwards. So this
//   bench uses a PIN-LEVEL SLAVE MODEL with an asymmetric return word, and checks two
//   independent observations per transfer:
//
//     * what the master received   (depends on bit order)
//     * what the slave received    (depends on bit order, the other way round)
//
//   Chapter 20.7 injects exactly that bug and shows the loopback check surviving it.
//
// WHAT THE SLAVE MODEL IS AND IS NOT
//
//   It is the mirror-image DEVICE: it samples on the edge the master launches on, and
//   launches on the edge the master samples on. It is NOT the source of the expected
//   values -- those come from `mask` and `revw` below, which know nothing about either
//   the controller or the model. If the model and the controller were both wrong in
//   the same direction the arithmetic would still catch it.
//
//   And per 20.4's rule about models: the model is configured for every mode the
//   controller supports. A slave model that only works in mode 0 tests mode 0 and
//   reports on four.

`timescale 1ns/1ps

module spi_capstone_ctrl_tb;

    localparam integer DATA_W = 16;

    // ---- DUT interface -------------------------------------------------------
    reg         clk, rst_n;
    reg         cfg_cpol, cfg_cpha, cfg_lsb;
    reg  [4:0]  cfg_width;
    reg  [7:0]  cfg_div;
    reg  [1:0]  cfg_dev;
    reg  [3:0]  cfg_lead, cfg_lag, cfg_idle;
    reg         start, abort;
    reg  [15:0] tx_data;
    wire        busy, done, cfg_err;
    wire [15:0] rx_data;
    wire [4:0]  bits_done;
    wire        sclk, mosi;
    wire [3:0]  cs_n;
    wire        miso;

    // ---- bookkeeping ---------------------------------------------------------
    integer n_chk, n_err, n_neg_ok;
    integer cyc;
    reg [8*22-1:0] gname;

    spi_capstone_ctrl #(.DATA_W(16), .MIN_WIDTH(4), .NDEV(4)) dut (
        .clk(clk), .rst_n(rst_n),
        .cfg_cpol(cfg_cpol), .cfg_cpha(cfg_cpha), .cfg_lsb_first(cfg_lsb),
        .cfg_width(cfg_width), .cfg_div(cfg_div), .cfg_dev(cfg_dev),
        .cfg_lead(cfg_lead), .cfg_lag(cfg_lag), .cfg_idle(cfg_idle),
        .start(start), .tx_data(tx_data), .abort(abort),
        .busy(busy), .done(done), .cfg_err(cfg_err),
        .rx_data(rx_data), .bits_done(bits_done),
        .sclk(sclk), .mosi(mosi), .cs_n(cs_n), .miso(miso)
    );

    always #5 clk = ~clk;

    // =========================================================================
    // Specification arithmetic. The independent oracle.
    // =========================================================================

    // Low-w-bit mask. Built by shifting a widened 1 -- `(1 << w) - 1` in a 16-bit
    // expression gives 0 for w = 16, which would mask every checked value to zero and
    // make the whole suite pass vacuously.
    function [15:0] mask;
        input [4:0] w;
        reg [16:0] one;
        begin
            one  = 17'd1;
            mask = ((one << w) - 17'd1);
        end
    endfunction

    // Reverse the low w bits of v, leaving the rest zero. This is the ONLY place the
    // bench encodes what LSB-first means.
    function [15:0] revw;
        input [15:0] v;
        input [4:0]  w;
        integer b;
        begin
            revw = 16'd0;
            for (b = 0; b < 16; b = b + 1)
                if (b < w) revw[w-1-b] = v[b];
        end
    endfunction

    // What the SLAVE should end up holding, given what the master was asked to send.
    function [15:0] exp_slave_rx;
        input [15:0] tx; input [4:0] w; input lsb;
        begin
            exp_slave_rx = lsb ? revw(tx & mask(w), w) : (tx & mask(w));
        end
    endfunction

    // What the MASTER should end up holding, given what the slave returns.
    function [15:0] exp_master_rx;
        input [15:0] sw; input [4:0] w; input lsb;
        begin
            exp_master_rx = lsb ? revw(sw & mask(w), w) : (sw & mask(w));
        end
    endfunction

    // =========================================================================
    // Pin-level slave model -- configurable for all four modes.
    // =========================================================================
    reg        slv_cpol, slv_cpha;
    reg [15:0] slv_word, slv_sr, slv_rx;
    reg [4:0]  slv_w, slv_idx, slv_nrx;
    reg        slv_miso, force_x;
    reg        lead_s;

    wire cs_any = ~(&cs_n);
    assign miso = force_x ? 1'bx : slv_miso;

    // A select fell: load, and for CPHA=0 present the first bit immediately -- the
    // master will sample it on the very first edge, so there is no edge left to
    // launch it on.
    always @(posedge cs_any) begin
        slv_sr   = slv_word << (16 - slv_w);
        slv_rx   = 16'd0;
        slv_idx  = 5'd0;
        slv_nrx  = 5'd0;
        if (!slv_cpha) begin
            slv_miso = slv_sr[15];
            slv_sr   = slv_sr << 1;
            slv_idx  = 5'd1;
        end else begin
            slv_miso = 1'b0;
        end
    end

    always @(sclk) begin
        if (cs_any === 1'b1) begin
            // A transition AWAY from the idle level is leading; back to it, trailing.
            // CPOL never appears anywhere else in this model.
            lead_s = (sclk !== slv_cpol);
            // NOT exchanged: a slave uses the SAME edges as its master. Both
            // sample on one edge and both change their output on the other -- that is
            // what makes the link work off a single clock. The model's two conditions
            // are therefore identical to the controller's, not mirrored.
            if (slv_cpha ? !lead_s : lead_s) begin
                if (slv_nrx < slv_w) begin
                    slv_rx  = {slv_rx[14:0], mosi};
                    slv_nrx = slv_nrx + 5'd1;
                end
            end
            if (slv_cpha ? lead_s : !lead_s) begin
                if (slv_idx < slv_w) begin
                    slv_miso = slv_sr[15];
                    slv_sr   = slv_sr << 1;
                    slv_idx  = slv_idx + 5'd1;
                end
            end
        end
    end

    // =========================================================================
    // Pin monitors. Every timing number reported by this bench is MEASURED at the
    // pins, never computed from the configuration -- 20.4 depends on that, and
    // Module 19.4 is why (hand-derived arithmetic was one cycle out).
    // =========================================================================
    reg        cs_d, sclk_d;
    integer    t_cs_fall, t_cs_rise, t_first_edge, t_last_edge, t_prev_rise;
    integer    n_edge, m_half, m_lead, m_lag, m_gap, n_overlap;
    integer    n_low;

    always @(posedge clk) begin
        if (!rst_n) begin
            cyc <= 0; cs_d <= 1'b0; sclk_d <= 1'b0; n_edge <= 0; n_overlap <= 0;
        end else begin
            cyc <= cyc + 1;

            n_low = (cs_n[0] ? 0 : 1) + (cs_n[1] ? 0 : 1)
                  + (cs_n[2] ? 0 : 1) + (cs_n[3] ? 0 : 1);
            if (n_low > 1) n_overlap <= n_overlap + 1;

            if (cs_any && !cs_d) begin
                t_cs_fall <= cyc;
                m_gap     <= cyc - t_prev_rise;
                n_edge    <= 0;
            end
            if (!cs_any && cs_d) begin
                t_cs_rise  <= cyc;
                t_prev_rise <= cyc;
                m_lag      <= cyc - t_last_edge;
            end
                // `cs_d` as well as `cs_any`: an edge is only a FRAME edge if a device
                // was ALREADY selected last cycle. A transition in the very cycle the
                // select falls is SCLK reaching its new idle level, not a clocking
                // edge -- the controller parks SCLK and asserts CS together, so when
                // the previous idle level differed the two coincide.
                //
                // Measured cost of omitting `cs_d`: the first transaction after reset
                // with CPOL=1 counted 17 edges instead of 16 in VHDL and 16 in
                // SystemVerilog, because a one-cycle difference in reset-release
                // timing decided whether the re-park landed inside the window. The
                // received data was correct in both. With the gate the measurement no
                // longer depends on that phase at all.
            if (cs_any && cs_d && (sclk !== sclk_d)) begin
                if (n_edge == 0) begin
                    t_first_edge <= cyc;
                    m_lead       <= cyc - t_cs_fall;
                end else begin
                    m_half <= cyc - t_last_edge;
                end
                t_last_edge <= cyc;
                n_edge      <= n_edge + 1;
            end
            cs_d   <= cs_any;
            sclk_d <= sclk;
        end
    end

    // =========================================================================
    // Checkers
    // =========================================================================

    // `!==` not `!=`: an X anywhere in `got` must FAIL. With `!=` an X compares as
    // "unknown", the `if` treats it as false, and the check silently passes -- which
    // is the mechanism group 8 deliberately exercises.
    task chk16;
        input [8*22-1:0] nm; input [15:0] got; input [15:0] exp;
        begin
            n_chk = n_chk + 1;
            if (got !== exp) begin
                n_err = n_err + 1;
                $display("    FAIL %0s: got %04h expected %04h", nm, got, exp);
            end
        end
    endtask

    task chki;
        input [8*22-1:0] nm; input integer got; input integer exp;
        begin
            n_chk = n_chk + 1;
            if (got !== exp) begin
                n_err = n_err + 1;
                $display("    FAIL %0s: got %0d expected %0d", nm, got, exp);
            end
        end
    endtask

    // =========================================================================
    // Stimulus helpers
    // =========================================================================
    task set_cfg;
        input cpol_i; input cpha_i; input lsb_i; input [4:0] w;
        input [7:0] dv; input [1:0] dv_n;
        input [3:0] ld; input [3:0] lg; input [3:0] id;
        begin
            cfg_cpol = cpol_i; cfg_cpha = cpha_i; cfg_lsb = lsb_i;
            cfg_width = w; cfg_div = dv; cfg_dev = dv_n;
            cfg_lead = ld; cfg_lag = lg; cfg_idle = id;
            slv_cpol = cpol_i; slv_cpha = cpha_i; slv_w = w;
        end
    endtask

    // EVERY input is driven on the NEGEDGE, and this is not a style preference.
    //
    // Driving `start` on the posedge the controller samples it on is a race: the
    // assignment that clears it and the controller's clocked block are both in the
    // same region of the same time step, and their order is not defined by the
    // language. Measured cost of getting this wrong: the FIRST frame of the run was
    // accepted and every later one was silently dropped, so 31 of 32 transfers
    // compared a stale `rx_data` against a fresh expectation and the suite reported
    // a wall of mismatches that had nothing to do with the controller.
    //
    // A negedge-driven `start` is high across exactly one sampling edge, and nothing
    // the bench does can collide with the edge that matters. This is Chapter 16.3's
    // clocking-block discipline written out by hand.
    task fire;
        input [15:0] d;
        begin
            @(negedge clk);
            tx_data = d;
            start   = 1'b1;
            @(negedge clk);
            start   = 1'b0;
        end
    endtask

    // Bounded wait. Observation happens on the NEGEDGE so every non-blocking update
    // from the clock edge has settled -- reading a one-cycle `done` right after
    // `@(posedge clk)` reads its pre-edge value and misses it entirely.
    task wait_idle;
        input integer maxc; output got_done; output integer took;
        integer g; reg seen;
        begin
            g = 0; seen = 1'b0;
            while (g < maxc) begin
                @(negedge clk);
                g = g + 1;
                if (done) seen = 1'b1;
                if (!busy && seen) g = maxc;
                else if (!busy && g > 4) g = maxc;
            end
            got_done = seen;
            took     = g;
        end
    endtask

    // =========================================================================
    // Test sequence
    // =========================================================================
    integer mi, oi, wi, k;
    reg [4:0]  WID [0:3];
    reg [15:0] TXP [0:3];
    reg [15:0] SWP [0:3];
    reg        gd;
    integer    tk, grp_err;
    reg [15:0] rx_before;

    initial begin
        clk = 1'b0; rst_n = 1'b0;
        start = 1'b0; abort = 1'b0; tx_data = 16'd0;
        cfg_cpol = 1'b0; cfg_cpha = 1'b0; cfg_lsb = 1'b0;
        cfg_width = 5'd8; cfg_div = 8'd1; cfg_dev = 2'd0;
        cfg_lead = 4'd1; cfg_lag = 4'd1; cfg_idle = 4'd1;
        slv_cpol = 1'b0; slv_cpha = 1'b0; slv_w = 5'd8;
        slv_word = 16'd0; slv_miso = 1'b0; force_x = 1'b0;
        slv_sr = 16'd0; slv_rx = 16'd0; slv_idx = 5'd0; slv_nrx = 5'd0;
        n_chk = 0; n_err = 0; n_neg_ok = 0;
        cyc = 0; n_edge = 0; n_overlap = 0;
        t_cs_fall = 0; t_cs_rise = 0; t_first_edge = 0; t_last_edge = 0;
        t_prev_rise = 0; m_half = 0; m_lead = 0; m_lag = 0; m_gap = 0;
        cs_d = 1'b0; sclk_d = 1'b0;

        WID[0] = 5'd4;  WID[1] = 5'd8;  WID[2] = 5'd13; WID[3] = 5'd16;
        TXP[0] = 16'hB39D; TXP[1] = 16'hB39D; TXP[2] = 16'hB39D; TXP[3] = 16'hB39D;
        SWP[0] = 16'h4E7A; SWP[1] = 16'h4E7A; SWP[2] = 16'h4E7A; SWP[3] = 16'h4E7A;

        $display("=== Chapter 20.3 -- capstone controller, directed suite ===");

        repeat (4) @(posedge clk);
        rst_n = 1'b1;
        repeat (2) @(posedge clk);

        // -----------------------------------------------------------------
        // GROUP 1 -- every mode, both bit orders, four widths.
        // The question: does the controller implement the specification's mode
        // and bit-order semantics, for every width the datapath claims?
        // -----------------------------------------------------------------
        $display("  G1 mode x bit-order x width");
        for (mi = 0; mi < 4; mi = mi + 1) begin
            for (oi = 0; oi < 2; oi = oi + 1) begin
                grp_err = n_err;
                for (wi = 0; wi < 4; wi = wi + 1) begin
                    set_cfg(mi[1], mi[0], oi[0], WID[wi], 8'd1, 2'd1, 4'd2, 4'd2, 4'd2);
                    slv_word = SWP[wi] & mask(WID[wi]);
                    fire(TXP[wi]);
                    wait_idle(4000, gd, tk);
                    chki ("done pulsed",   gd ? 1 : 0, 1);
                    chk16("master rx",     rx_data,
                          exp_master_rx(SWP[wi] & mask(WID[wi]), WID[wi], oi[0]));
                    chk16("slave rx",      slv_rx & mask(WID[wi]),
                          exp_slave_rx(TXP[wi], WID[wi], oi[0]));
                    chki ("sclk edges",    n_edge, 2 * WID[wi]);
                    chki ("bits sampled",  bits_done, WID[wi]);
                    chki ("sclk parked",   (sclk === mi[1]) ? 1 : 0, 1);
                    chki ("no cs overlap", n_overlap, 0);
                end
                $display("    mode %0d %0s : rx(w4,w8,w13,w16) checked, %0s",
                         mi, oi[0] ? "lsb" : "msb",
                         (n_err == grp_err) ? "all match" : "MISMATCH");
            end
        end

        // -----------------------------------------------------------------
        // GROUP 2 -- measured pin intervals.
        // The question: are the lead / lag / turnaround requirements MET at the
        // pins, and what is the actual margin? The requirement is ">=", the
        // measurement is exact, and the difference is reported rather than assumed.
        // -----------------------------------------------------------------
        $display("  G2 measured pin intervals, in clk cycles");
        $display("    div lead lag idle | t_half t_lead t_lag t_gap");
        for (k = 0; k < 4; k = k + 1) begin
            set_cfg(1'b0, 1'b0, 1'b0, 5'd8,
                    (k == 0) ? 8'd0 : (k == 1) ? 8'd1 : (k == 2) ? 8'd3 : 8'd1,
                    2'd2,
                    (k == 3) ? 4'd5 : 4'd2,
                    (k == 3) ? 4'd4 : 4'd2,
                    (k == 3) ? 4'd6 : 4'd3);
            slv_word = 16'h5A;
            // Two frames: the second one's gap is the interval between them.
            fire(16'h33); wait_idle(4000, gd, tk);
            fire(16'h33); wait_idle(4000, gd, tk);
            $display("    %3d %4d %3d %4d | %6d %6d %5d %5d",
                     cfg_div, cfg_lead, cfg_lag, cfg_idle, m_half, m_lead, m_lag, m_gap);
            // REQ-TIM-001 is an equality and is checked as one.
            chki("t_half exact", m_half, cfg_div + 1);

            // REQ-TIM-002/003/004 are ">=" requirements, and they are checked as
            // stated -- but a ">=" check is WEAK: an implementation that loses one
            // half-period of lead still satisfies it whenever cfg_lead > 0. So each
            // interval is ALSO checked against the exact closed form the measurements
            // above establish, which is what actually catches a one-tick regression.
            if (m_lead < (cfg_lead * (cfg_div + 1))) begin
                n_err = n_err + 1; $display("    FAIL lead below requirement");
            end
            n_chk = n_chk + 1;
            if (m_lag < (cfg_lag * (cfg_div + 1))) begin
                n_err = n_err + 1; $display("    FAIL lag below requirement");
            end
            n_chk = n_chk + 1;
            if (m_gap < (cfg_idle * (cfg_div + 1))) begin
                n_err = n_err + 1; $display("    FAIL turnaround below requirement");
            end
            n_chk = n_chk + 1;

            // Exact implementation forms, in system clocks. The "+2" and "+1" terms
            // are the ticks the state machine spends LEAVING a phase, and they are
            // margin above the requirement, not part of it:
            //
            //   t_lead = (cfg_lead + 2) * t_half   one tick to exit S_LEAD, one more
            //                                      before the first transition
            //   t_lag  = (cfg_lag  + 1) * t_half   one tick to exit S_LAG
            //   t_gap  = (cfg_idle + 1) * t_half + 2
            //
            // The gap is the ONLY interval here with a term that is not a multiple of
            // the half-period, and the reason is worth naming: those 2 cycles are
            // S_IDLE plus ONE CYCLE OF THE REQUESTER'S OWN LATENCY. With no request
            // queue, the bench cannot post the next transfer until `busy` falls, so
            // the measured gap is the design's floor PLUS however long the requester
            // took to react. 20.2 defends that trade-off and 20.7 questions it.
            chki("t_lead exact", m_lead, (cfg_lead + 2) * (cfg_div + 1));
            chki("t_lag exact",  m_lag,  (cfg_lag  + 1) * (cfg_div + 1));
            chki("t_gap exact",  m_gap,  (cfg_idle + 1) * (cfg_div + 1) + 2);
        end

        // -----------------------------------------------------------------
        // GROUP 3 -- illegal configuration is REFUSED, at both boundaries.
        // -----------------------------------------------------------------
        $display("  G3 illegal width");
        for (k = 0; k < 4; k = k + 1) begin
            set_cfg(1'b0, 1'b0, 1'b0,
                    (k == 0) ? 5'd3 : (k == 1) ? 5'd17 : (k == 2) ? 5'd4 : 5'd16,
                    8'd1, 2'd0, 4'd2, 4'd2, 4'd2);
            slv_word = 16'hFFFF & mask(cfg_width);
            slv_w    = (k < 2) ? 5'd8 : cfg_width;
            fire(16'h5555);
            // `fire` returns on the negedge immediately after the sampling edge, so a
            // one-cycle output is readable RIGHT HERE. One more negedge and `cfg_err`
            // has already gone low -- which is how a correct rejection reads as a
            // missing one.
            if (k < 2) begin
                chki("cfg_err raised", cfg_err ? 1 : 0, 1);
                chki("stayed idle",    busy ? 1 : 0,    0);
                chki("no cs asserted", cs_any ? 1 : 0,  0);
                $display("    width %2d rejected: cfg_err=1 busy=0 cs=idle", cfg_width);
            end else begin
                chki("accepted",       busy ? 1 : 0,    1);
                chki("no cfg_err",     cfg_err ? 1 : 0, 0);
                wait_idle(4000, gd, tk);
                chki("done pulsed",    gd ? 1 : 0, 1);
                $display("    width %2d accepted: cfg_err=0 done=1", cfg_width);
            end
        end

        // -----------------------------------------------------------------
        // GROUP 4 -- a request while busy has NO effect (REQ-FUNC-007).
        // -----------------------------------------------------------------
        $display("  G4 start while busy");
        set_cfg(1'b0, 1'b0, 1'b0, 5'd8, 8'd2, 2'd0, 4'd2, 4'd2, 4'd2);
        slv_word = 16'h96;
        fire(16'hA5);
        repeat (6) @(posedge clk);
        fire(16'h3C);                        // ignored: the frame in flight owns the bus
        wait_idle(4000, gd, tk);
        chk16("rx from FIRST word", rx_data, exp_master_rx(16'h96, 5'd8, 1'b0));
        chk16("slave got FIRST word", slv_rx & mask(5'd8), exp_slave_rx(16'hA5, 5'd8, 1'b0));
        chki ("edges of one frame", n_edge, 16);
        $display("    second request ignored: one frame, 16 edges, slave saw a5");

        // -----------------------------------------------------------------
        // GROUP 5 -- reset in the middle of a frame (REQ-RST-002).
        // -----------------------------------------------------------------
        $display("  G5 reset during a frame");
        set_cfg(1'b1, 1'b0, 1'b0, 5'd16, 8'd3, 2'd3, 4'd2, 4'd2, 4'd2);
        slv_word = 16'hBEEF;
        fire(16'hDEAD);
        repeat (20) @(posedge clk);
        chki("frame really started", cs_any ? 1 : 0, 1);
        @(negedge clk); rst_n = 1'b0;
        @(negedge clk);
        chki("all selects released", cs_any ? 1 : 0, 0);
        chki("no done on reset",     done ? 1 : 0,   0);
        chki("not busy",             busy ? 1 : 0,   0);
        chki("sclk defined 0",       (sclk === 1'b0) ? 1 : 0, 1);
        $display("    reset mid-frame: cs released, no done, sclk=0 (not cpol)");
        @(negedge clk); rst_n = 1'b1; repeat (3) @(posedge clk);

        // -----------------------------------------------------------------
        // GROUP 6 -- abort (REQ-ABT-001).
        // -----------------------------------------------------------------
        $display("  G6 abort mid-frame");
        set_cfg(1'b0, 1'b0, 1'b0, 5'd16, 8'd1, 2'd0, 4'd2, 4'd2, 4'd2);
        slv_word = 16'h1234;
        fire(16'h4321);
        wait_idle(4000, gd, tk);             // a clean frame first, to set rx_data
        rx_before = rx_data;
        slv_word  = 16'hFFFF;
        fire(16'h0F0F);
        repeat (24) @(posedge clk);
        @(negedge clk); abort = 1'b1; @(negedge clk); abort = 1'b0;
        wait_idle(4000, gd, tk);
        chki ("no done on abort",   gd ? 1 : 0, 0);
        chk16("rx_data unchanged",  rx_data, rx_before);
        chki ("selects released",   cs_any ? 1 : 0, 0);
        chki ("sclk back at idle",  (sclk === 1'b0) ? 1 : 0, 1);
        if (bits_done == 5'd0 || bits_done >= 5'd16) begin
            n_err = n_err + 1;
            $display("    FAIL bits_done not partial: %0d", bits_done);
        end
        n_chk = n_chk + 1;
        $display("    abort: done=0, rx held %04h, bits_done=%0d of 16",
                 rx_data, bits_done);

        // -----------------------------------------------------------------
        // GROUP 7 -- configuration captured at acceptance (commitment 2).
        // The frame must ignore a mid-flight rewrite of EVERY field it uses.
        // -----------------------------------------------------------------
        $display("  G7 configuration captured at acceptance");
        set_cfg(1'b0, 1'b0, 1'b0, 5'd8, 8'd2, 2'd0, 4'd2, 4'd2, 4'd2);
        slv_word = 16'h7E;
        fire(16'h81);
        repeat (5) @(posedge clk);
        // Rewrite everything to values that would produce a different frame.
        cfg_cpol = 1'b1; cfg_cpha = 1'b1; cfg_lsb = 1'b1;
        cfg_width = 5'd4; cfg_div = 8'd7; cfg_dev = 2'd3;
        wait_idle(4000, gd, tk);
        chk16("rx used captured cfg", rx_data, exp_master_rx(16'h7E, 5'd8, 1'b0));
        chki ("edges of captured width", n_edge, 16);
        chki ("captured device kept", cs_n[0] === 1'b1 ? 1 : 0, 1);
        $display("    mid-frame rewrite ignored: 16 edges, msb-order rx=%04h", rx_data);
        set_cfg(1'b0, 1'b0, 1'b0, 5'd8, 8'd1, 2'd0, 4'd2, 4'd2, 4'd2);

        // -----------------------------------------------------------------
        // GROUP 8 -- an X on MISO must NOT be able to produce a pass.
        // This is a NEGATIVE test: the checker is EXPECTED to fail here, and the
        // bench counts that as evidence rather than as a defect.
        // -----------------------------------------------------------------
        $display("  G8 X on MISO cannot slip through");
        slv_word = 16'h5A;
        force_x  = 1'b1;
        fire(16'hA5);
        wait_idle(4000, gd, tk);
        n_chk = n_chk + 1;
        if (rx_data !== exp_master_rx(16'h5A, 5'd8, 1'b0)) begin
            n_neg_ok = n_neg_ok + 1;
            $display("    X reached rx_data and the !== comparison rejected it");
        end else begin
            n_err = n_err + 1;
            $display("    FAIL an all-X rx_data compared EQUAL -- checker is blind");
        end
        force_x = 1'b0;

        // -----------------------------------------------------------------
        // GROUP 9 -- prove the checkers can fail at all.
        // A suite that has never printed FAIL has not been shown to be able to.
        // -----------------------------------------------------------------
        $display("  G9 deliberately wrong expectations must FAIL");
        // The pattern here is 9d, and the reason it is not a5 or 5a is worth more
        // than the test it enables:
        //
        //   a5 = 10100101   reversed = 10100101
        //   5a = 01011010   reversed = 01011010
        //
        // The two most-used test bytes in digital design are both EIGHT-BIT
        // PALINDROMES. Neither can distinguish MSB-first from LSB-first, so a suite
        // built on them reports a clean pass with the bit-order logic inverted -- and
        // the self-test of the oracle below reported a defect in the ORACLE when the
        // only defect was the choice of stimulus. 9d reverses to b9.
        k = 0;
        slv_word = 16'h9D;
        fire(16'h9D);
        wait_idle(4000, gd, tk);
        n_chk = n_chk + 1;
        if (rx_data !== (exp_master_rx(16'h9D, 5'd8, 1'b0) ^ 16'h0001)) k = k + 1;
        else begin n_err = n_err + 1; $display("    FAIL rx checker accepted a wrong value"); end
        n_chk = n_chk + 1;
        if (n_edge !== 15) k = k + 1;
        else begin n_err = n_err + 1; $display("    FAIL edge checker accepted a wrong count"); end
        n_chk = n_chk + 1;
        if (exp_master_rx(16'h9D, 5'd8, 1'b1) !== exp_master_rx(16'h9D, 5'd8, 1'b0)) k = k + 1;
        else begin n_err = n_err + 1; $display("    FAIL oracle is blind to bit order"); end
        n_neg_ok = n_neg_ok + k;
        $display("    %0d of 3 wrong expectations rejected; oracle separates 9d from b9", k);

        $display("=== SUMMARY checks=%0d negatives=%0d failures=%0d : %0s ===",
                 n_chk, n_neg_ok, n_err, (n_err == 0) ? "PASS" : "FAIL");
        $finish;
    end

    // Global watchdog. A hung frame must end the run with a FAILURE, not with a
    // simulator that stops printing.
    initial begin
        #4000000;
        $display("    FATAL global timeout -- a frame never completed");
        $display("=== SUMMARY checks=%0d negatives=%0d failures=%0d : FAIL ===",
                 n_chk, n_neg_ok, n_err + 1);
        $finish;
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_capstone_ctrl_tb.vhd — the same bench in VHDL, with one slave process because two would resolve to X
-- spi_capstone_ctrl_tb.vhd
--
-- Chapter 20.3 -- the directed suite in VHDL-2008. Same nine groups, same stimulus,
-- same expected values, and a transcript that must match the other two languages
-- LINE FOR LINE. That comparison is the real cross-language check: reading three files
-- side by side proves they look alike; comparing three transcripts proves they behave
-- alike, and only one of those is evidence.
--
-- THREE STRUCTURAL DECISIONS FORCED BY VHDL, ALL WORTH KNOWING
--
--   (1) THE SLAVE MODEL IS ONE PROCESS, NOT TWO.
--       The SystemVerilog model uses two always blocks -- one on the select falling,
--       one on every SCLK transition -- and both assign MISO. In VHDL that is TWO
--       DRIVERS on a resolved type, and the result is not an error: it is 'X' on the
--       pin, silently, for the whole run. The two events are merged into a single
--       process with an explicit priority between them.
--
--   (2) EVERY MONITOR TIMESTAMP IS A SIGNAL, NOT A VARIABLE.
--       The monitor reads `n_edge` and `t_last_edge` in the same invocation that
--       writes them, and the SystemVerilog original reads the PRE-EDGE values because
--       non-blocking assignment defers the update. Variables would be visible
--       immediately and every measured interval would come out one cycle short. This
--       is the Module 19 defect that made a device model under-report its own timing
--       requirement, and signals are the fix.
--
--   (3) THE CHECK COUNTERS ARE PROCESS VARIABLES.
--       Signals would need a single driving process, and every check happens inside
--       the stimulus process anyway. Variables are correct here for exactly the reason
--       they were wrong in (2): nothing outside this process reads them.

library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
use std.textio.all;
use std.env.all;

entity spi_capstone_ctrl_tb is
end entity spi_capstone_ctrl_tb;

architecture tb of spi_capstone_ctrl_tb is

    signal clk       : std_logic := '0';
    signal rst_n     : std_logic := '0';
    signal cfg_cpol  : std_logic := '0';
    signal cfg_cpha  : std_logic := '0';
    signal cfg_lsb   : std_logic := '0';
    signal cfg_width : std_logic_vector(4 downto 0) := "01000";
    signal cfg_div   : std_logic_vector(7 downto 0) := x"01";
    signal cfg_dev   : std_logic_vector(1 downto 0) := "00";
    signal cfg_lead  : std_logic_vector(3 downto 0) := x"1";
    signal cfg_lag   : std_logic_vector(3 downto 0) := x"1";
    signal cfg_idle  : std_logic_vector(3 downto 0) := x"1";
    signal start     : std_logic := '0';
    signal abort     : std_logic := '0';
    signal tx_data   : std_logic_vector(15 downto 0) := (others => '0');
    signal busy      : std_logic;
    signal done      : std_logic;
    signal cfg_err   : std_logic;
    signal rx_data   : std_logic_vector(15 downto 0);
    signal bits_done : std_logic_vector(4 downto 0);
    signal sclk      : std_logic;
    signal mosi      : std_logic;
    signal cs_n      : std_logic_vector(3 downto 0);
    signal miso      : std_logic;

    -- slave model
    signal slv_cpol  : std_logic := '0';
    signal slv_cpha  : std_logic := '0';
    signal slv_word  : std_logic_vector(15 downto 0) := (others => '0');
    signal slv_rx    : std_logic_vector(15 downto 0) := (others => '0');
    signal slv_w     : unsigned(4 downto 0) := to_unsigned(8, 5);
    signal slv_miso  : std_logic := '0';
    signal force_x   : std_logic := '0';
    signal cs_any    : std_logic;

    -- monitor state: signals, for the reason in note (2) above
    signal cyc         : integer := 0;
    signal n_edge      : integer := 0;
    signal n_overlap   : integer := 0;
    signal m_half      : integer := 0;
    signal m_lead      : integer := 0;
    signal m_lag       : integer := 0;
    signal m_gap       : integer := 0;
    signal cs_d        : std_logic := '0';
    signal sclk_d      : std_logic := '0';
    signal t_cs_fall   : integer := 0;
    signal t_last_edge : integer := 0;
    signal t_prev_rise : integer := 0;

    -- ---------------- specification arithmetic: the independent oracle -------------
    function maskw (w : integer) return unsigned is
        variable m : unsigned(15 downto 0);
    begin
        m := (others => '0');
        for b in 0 to 15 loop
            if b < w then m(b) := '1'; end if;
        end loop;
        return m;
    end function maskw;

    function revw (v : std_logic_vector(15 downto 0); w : integer)
        return std_logic_vector is
        variable r : std_logic_vector(15 downto 0);
    begin
        r := (others => '0');
        for b in 0 to 15 loop
            if b < w then r(w-1-b) := v(b); end if;
        end loop;
        return r;
    end function revw;

    function exp_slave_rx (tx : std_logic_vector(15 downto 0);
                           w : integer; lsb : std_logic)
        return std_logic_vector is
        variable m : std_logic_vector(15 downto 0);
    begin
        m := std_logic_vector(unsigned(tx) and maskw(w));
        if lsb = '1' then return revw(m, w); else return m; end if;
    end function exp_slave_rx;

    function exp_master_rx (sw : std_logic_vector(15 downto 0);
                            w : integer; lsb : std_logic)
        return std_logic_vector is
        variable m : std_logic_vector(15 downto 0);
    begin
        m := std_logic_vector(unsigned(sw) and maskw(w));
        if lsb = '1' then return revw(m, w); else return m; end if;
    end function exp_master_rx;

    -- ---------------- output formatting -------------------------------------------
    function hex4 (v : std_logic_vector(15 downto 0)) return string is
        constant D : string(1 to 16) := "0123456789abcdef";
        variable s : string(1 to 4);
        variable n : integer;
    begin
        if is_x(v) then return "xxxx"; end if;
        n := to_integer(unsigned(v));
        for i in 4 downto 1 loop
            s(i) := D((n mod 16) + 1);
            n := n / 16;
        end loop;
        return s;
    end function hex4;

    -- Right-align an integer in w columns. The return is CONSTRAINED to w so no
    -- caller inherits a surprising range, and the digits are built into a local
    -- scratch string rather than concatenated, which would make the length depend on
    -- the value -- a fatal length mismatch waiting for the first two-digit result.
    function ipad (v : integer; w : integer) return string is
        variable s   : string(1 to w);
        variable t   : string(1 to 20);
        variable n   : integer;
        variable len : integer;
    begin
        t   := (others => ' ');
        n   := v;
        len := 0;
        if n = 0 then
            len := 1; t(1) := '0';
        else
            while n > 0 loop
                len    := len + 1;
                t(len) := character'val(character'pos('0') + (n mod 10));
                n      := n / 10;
            end loop;
        end if;
        s := (others => ' ');
        for i in 1 to len loop
            s(w - i + 1) := t(i);
        end loop;
        return s;
    end function ipad;

    function i0 (v : integer) return string is
    begin
        return integer'image(v);
    end function i0;

    procedure pr (s : string) is
        variable l : line;
    begin
        write(l, s);
        writeline(output, l);
    end procedure pr;

    -- "msb" / "lsb" as a fixed THREE-character result. A conditional expression
    -- (`if ... then ... else ...` inside an expression) is VHDL-2019, not 2008, so it
    -- cannot be used here; and returning strings of different lengths from one
    -- function is the length-mismatch fault Module 18 hit. Both are 3.
    function ord_s (lsb : integer) return string is
    begin
        if lsb = 1 then return "lsb"; else return "msb"; end if;
    end function ord_s;

    function sl (b : boolean) return std_logic is
    begin
        if b then return '1'; else return '0'; end if;
    end function sl;

    type int4_t is array (0 to 3) of integer;
    constant WID : int4_t := (4, 8, 13, 16);
    constant TXP : std_logic_vector(15 downto 0) := x"B39D";
    constant SWP : std_logic_vector(15 downto 0) := x"4E7A";

begin

    cs_any <= not (cs_n(0) and cs_n(1) and cs_n(2) and cs_n(3));
    miso   <= 'X' when force_x = '1' else slv_miso;

    dut : entity work.spi_capstone_ctrl
        generic map (DATA_W => 16, MIN_WIDTH => 4, NDEV => 4)
        port map (
            clk => clk, rst_n => rst_n,
            cfg_cpol => cfg_cpol, cfg_cpha => cfg_cpha, cfg_lsb_first => cfg_lsb,
            cfg_width => cfg_width, cfg_div => cfg_div, cfg_dev => cfg_dev,
            cfg_lead => cfg_lead, cfg_lag => cfg_lag, cfg_idle => cfg_idle,
            start => start, tx_data => tx_data, abort => abort,
            busy => busy, done => done, cfg_err => cfg_err,
            rx_data => rx_data, bits_done => bits_done,
            sclk => sclk, mosi => mosi, cs_n => cs_n, miso => miso
        );

    clkgen : process
    begin
        clk <= '0'; wait for 5 ns;
        clk <= '1'; wait for 5 ns;
    end process clkgen;

    slave : process (cs_any, sclk)
        variable sr     : std_logic_vector(15 downto 0);
        variable idx    : unsigned(4 downto 0);
        variable nrx    : unsigned(4 downto 0);
        variable lead_s : boolean;
    begin
        if rising_edge(cs_any) then
            sr := std_logic_vector(shift_left(unsigned(slv_word),
                                              16 - to_integer(slv_w)));
            slv_rx <= (others => '0');
            nrx    := (others => '0');
            if slv_cpha = '0' then
                slv_miso <= sr(15);
                sr       := sr(14 downto 0) & '0';
                idx      := to_unsigned(1, 5);
            else
                slv_miso <= '0';
                idx      := (others => '0');
            end if;
        elsif sclk'event and cs_any = '1' then
            -- A transition AWAY from the idle level is leading; back to it, trailing.
            -- NOT mirrored from the controller: a slave uses the SAME edges as its
            -- master. Both sample on one edge and both change their output on the
            -- other -- that is what makes the link work off a single clock.
            lead_s := (sclk /= slv_cpol);
            if (slv_cpha = '1' and not lead_s) or (slv_cpha = '0' and lead_s) then
                if nrx < slv_w then
                    slv_rx <= slv_rx(14 downto 0) & mosi;
                    nrx    := nrx + 1;
                end if;
            end if;
            if (slv_cpha = '1' and lead_s) or (slv_cpha = '0' and not lead_s) then
                if idx < slv_w then
                    slv_miso <= sr(15);
                    sr       := sr(14 downto 0) & '0';
                    idx      := idx + 1;
                end if;
            end if;
        end if;
    end process slave;

    mon : process (clk)
        variable n_low : integer;
    begin
        if rising_edge(clk) then
            if rst_n = '0' then
                cyc       <= 0;
                n_edge    <= 0;
                n_overlap <= 0;
                cs_d      <= '0';
                sclk_d    <= '0';
            else
                cyc <= cyc + 1;

                n_low := 0;
                for i in 0 to 3 loop
                    if cs_n(i) = '0' then n_low := n_low + 1; end if;
                end loop;
                if n_low > 1 then n_overlap <= n_overlap + 1; end if;

                if cs_any = '1' and cs_d = '0' then
                    t_cs_fall <= cyc;
                    m_gap     <= cyc - t_prev_rise;
                    n_edge    <= 0;
                end if;
                if cs_any = '0' and cs_d = '1' then
                    t_prev_rise <= cyc;
                    m_lag       <= cyc - t_last_edge;
                end if;
                -- `cs_d` as well as `cs_any`: an edge is only a FRAME edge if a device
                -- was ALREADY selected last cycle. A transition in the cycle the
                -- select falls is SCLK reaching its new idle level, not a clocking
                -- edge. Without this gate the first CPOL=1 transaction counted 17
                -- edges here and 16 in SystemVerilog, on identical data.
                if cs_any = '1' and cs_d = '1' and sclk /= sclk_d then
                    if n_edge = 0 then
                        m_lead <= cyc - t_cs_fall;
                    else
                        m_half <= cyc - t_last_edge;
                    end if;
                    t_last_edge <= cyc;
                    n_edge      <= n_edge + 1;
                end if;
                cs_d   <= cs_any;
                sclk_d <= sclk;
            end if;
        end if;
    end process mon;

    wd : process
    begin
        wait for 4 ms;
        pr("    FATAL global timeout -- a frame never completed");
        pr("=== SUMMARY checks=0 negatives=0 failures=1 : FAIL ===");
        finish;
    end process wd;

    main : process
        variable n_chk   : integer := 0;
        variable n_err   : integer := 0;
        variable n_neg   : integer := 0;
        variable grp_err : integer := 0;
        variable gd      : std_logic;
        variable tk      : integer;
        variable kk      : integer;
        variable w       : integer;
        variable rx_before : std_logic_vector(15 downto 0);

        procedure chk16 (nm : string; got, exp : std_logic_vector(15 downto 0)) is
        begin
            n_chk := n_chk + 1;
            -- Case comparison: an 'X' anywhere in `got` must FAIL. With `=` on
            -- std_logic_vector an 'X' does not match '0' or '1' either, but printing
            -- it needs the is_x guard in hex4 -- group 8 exercises exactly this path.
            if got /= exp then
                n_err := n_err + 1;
                pr("    FAIL " & nm & ": got " & hex4(got) & " expected " & hex4(exp));
            end if;
        end procedure chk16;

        procedure chki (nm : string; got, exp : integer) is
        begin
            n_chk := n_chk + 1;
            if got /= exp then
                n_err := n_err + 1;
                pr("    FAIL " & nm & ": got " & i0(got) & " expected " & i0(exp));
            end if;
        end procedure chki;

        procedure set_cfg (cpol_i, cpha_i, lsb_i : std_logic; wv : integer;
                           dv : integer; dv_n : integer;
                           ld, lg, idl : integer) is
        begin
            cfg_cpol  <= cpol_i;
            cfg_cpha  <= cpha_i;
            cfg_lsb   <= lsb_i;
            cfg_width <= std_logic_vector(to_unsigned(wv, 5));
            cfg_div   <= std_logic_vector(to_unsigned(dv, 8));
            cfg_dev   <= std_logic_vector(to_unsigned(dv_n, 2));
            cfg_lead  <= std_logic_vector(to_unsigned(ld, 4));
            cfg_lag   <= std_logic_vector(to_unsigned(lg, 4));
            cfg_idle  <= std_logic_vector(to_unsigned(idl, 4));
            slv_cpol  <= cpol_i;
            slv_cpha  <= cpha_i;
            slv_w     <= to_unsigned(wv, 5);
        end procedure set_cfg;

        -- Every input is driven on the FALLING edge. Driving `start` on the edge the
        -- controller samples it on is a race whose measured cost was that the FIRST
        -- frame of the run was accepted and every later one silently dropped, so 31
        -- of 32 transfers compared a stale `rx_data` against a fresh expectation.
        procedure fire (d : std_logic_vector(15 downto 0)) is
        begin
            wait until falling_edge(clk);
            tx_data <= d;
            start   <= '1';
            wait until falling_edge(clk);
            start   <= '0';
        end procedure fire;

        procedure wait_idle (maxc : integer; got_done : out std_logic;
                             took : out integer) is
            variable g    : integer;
            variable seen : std_logic;
        begin
            g    := 0;
            seen := '0';
            while g < maxc loop
                wait until falling_edge(clk);
                g := g + 1;
                if done = '1' then seen := '1'; end if;
                if busy = '0' and seen = '1' then
                    g := maxc;
                elsif busy = '0' and g > 4 then
                    g := maxc;
                end if;
            end loop;
            got_done := seen;
            took     := g;
        end procedure wait_idle;

    begin
        pr("=== Chapter 20.3 -- capstone controller, directed suite ===");

        -- Reset is RELEASED ON A FALLING EDGE, for the same reason `start` is driven on
        -- one: released on a rising edge it races every clocked block that tests it.
        -- The measured cost was a one-cycle difference in when the monitor left reset,
        -- which counted SCLK's idle re-park as a frame edge in VHDL but not in
        -- SystemVerilog -- 17 edges against 16, on the first transaction only.
        for i in 1 to 4 loop wait until rising_edge(clk); end loop;
        wait until falling_edge(clk);
        rst_n <= '1';
        for i in 1 to 2 loop wait until rising_edge(clk); end loop;

        ----------------------------------------------------------------------
        pr("  G1 mode x bit-order x width");
        for mi in 0 to 3 loop
            for oi in 0 to 1 loop
                grp_err := n_err;
                for wi in 0 to 3 loop
                    w := WID(wi);
                    set_cfg(sl(mi / 2 = 1), sl(mi mod 2 = 1), sl(oi = 1),
                            w, 1, 1, 2, 2, 2);
                    slv_word <= std_logic_vector(unsigned(SWP) and maskw(w));
                    fire(TXP);
                    wait_idle(4000, gd, tk);
                    chki ("done pulsed", to_integer(unsigned'('0' & gd)), 1);
                    chk16("master rx", rx_data,
                          exp_master_rx(std_logic_vector(unsigned(SWP) and maskw(w)),
                                        w, sl(oi = 1)));
                    chk16("slave rx",
                          std_logic_vector(unsigned(slv_rx) and maskw(w)),
                          exp_slave_rx(TXP, w, sl(oi = 1)));
                    chki ("sclk edges", n_edge, 2 * w);
                    chki ("bits sampled", to_integer(unsigned(bits_done)), w);
                    chki ("sclk parked",
                          to_integer(unsigned'('0' & sl(sclk = sl(mi / 2 = 1)))), 1);
                    chki ("no cs overlap", n_overlap, 0);
                end loop;
                if n_err = grp_err then
                    pr("    mode " & i0(mi) & " " & ord_s(oi) &
                       " : rx(w4,w8,w13,w16) checked, all match");
                else
                    pr("    mode " & i0(mi) & " " & ord_s(oi) &
                       " : rx(w4,w8,w13,w16) checked, MISMATCH");
                end if;
            end loop;
        end loop;

        ----------------------------------------------------------------------
        pr("  G2 measured pin intervals, in clk cycles");
        pr("    div lead lag idle | t_half t_lead t_lag t_gap");
        for k in 0 to 3 loop
            if k = 0 then
                set_cfg('0', '0', '0', 8, 0, 2, 2, 2, 3);
            elsif k = 1 then
                set_cfg('0', '0', '0', 8, 1, 2, 2, 2, 3);
            elsif k = 2 then
                set_cfg('0', '0', '0', 8, 3, 2, 2, 2, 3);
            else
                set_cfg('0', '0', '0', 8, 1, 2, 5, 4, 6);
            end if;
            slv_word <= x"005A";
            fire(x"0033"); wait_idle(4000, gd, tk);
            fire(x"0033"); wait_idle(4000, gd, tk);
            pr("    " & ipad(to_integer(unsigned(cfg_div)), 3) &
               " "    & ipad(to_integer(unsigned(cfg_lead)), 4) &
               " "    & ipad(to_integer(unsigned(cfg_lag)), 3) &
               " "    & ipad(to_integer(unsigned(cfg_idle)), 4) &
               " | "  & ipad(m_half, 6) &
               " "    & ipad(m_lead, 6) &
               " "    & ipad(m_lag, 5) &
               " "    & ipad(m_gap, 5));
            chki("t_half exact", m_half, to_integer(unsigned(cfg_div)) + 1);
            if m_lead < to_integer(unsigned(cfg_lead))
                        * (to_integer(unsigned(cfg_div)) + 1) then
                n_err := n_err + 1; pr("    FAIL lead below requirement");
            end if;
            n_chk := n_chk + 1;
            if m_lag < to_integer(unsigned(cfg_lag))
                       * (to_integer(unsigned(cfg_div)) + 1) then
                n_err := n_err + 1; pr("    FAIL lag below requirement");
            end if;
            n_chk := n_chk + 1;
            if m_gap < to_integer(unsigned(cfg_idle))
                       * (to_integer(unsigned(cfg_div)) + 1) then
                n_err := n_err + 1; pr("    FAIL turnaround below requirement");
            end if;
            n_chk := n_chk + 1;
            chki("t_lead exact", m_lead,
                 (to_integer(unsigned(cfg_lead)) + 2)
                 * (to_integer(unsigned(cfg_div)) + 1));
            chki("t_lag exact", m_lag,
                 (to_integer(unsigned(cfg_lag)) + 1)
                 * (to_integer(unsigned(cfg_div)) + 1));
            chki("t_gap exact", m_gap,
                 (to_integer(unsigned(cfg_idle)) + 1)
                 * (to_integer(unsigned(cfg_div)) + 1) + 2);
        end loop;

        ----------------------------------------------------------------------
        pr("  G3 illegal width");
        for k in 0 to 3 loop
            if    k = 0 then w := 3;
            elsif k = 1 then w := 17;
            elsif k = 2 then w := 4;
            else                 w := 16;
            end if;
            set_cfg('0', '0', '0', w, 1, 0, 2, 2, 2);
            if k < 2 then
                slv_w    <= to_unsigned(8, 5);
                slv_word <= x"00FF";
            else
                slv_word <= std_logic_vector(unsigned'(x"FFFF") and maskw(w));
            end if;
            fire(x"5555");
            if k < 2 then
                chki("cfg_err raised", to_integer(unsigned'('0' & cfg_err)), 1);
                chki("stayed idle",    to_integer(unsigned'('0' & busy)), 0);
                chki("no cs asserted", to_integer(unsigned'('0' & cs_any)), 0);
                pr("    width " & ipad(w, 2) &
                   " rejected: cfg_err=1 busy=0 cs=idle");
            else
                chki("accepted",   to_integer(unsigned'('0' & busy)), 1);
                chki("no cfg_err", to_integer(unsigned'('0' & cfg_err)), 0);
                wait_idle(4000, gd, tk);
                chki("done pulsed", to_integer(unsigned'('0' & gd)), 1);
                pr("    width " & ipad(w, 2) & " accepted: cfg_err=0 done=1");
            end if;
        end loop;

        ----------------------------------------------------------------------
        pr("  G4 start while busy");
        set_cfg('0', '0', '0', 8, 2, 0, 2, 2, 2);
        slv_word <= x"0096";
        fire(x"00A5");
        for i in 1 to 6 loop wait until rising_edge(clk); end loop;
        fire(x"003C");
        wait_idle(4000, gd, tk);
        chk16("rx from FIRST word", rx_data, exp_master_rx(x"0096", 8, '0'));
        chk16("slave got FIRST word",
              std_logic_vector(unsigned(slv_rx) and maskw(8)),
              exp_slave_rx(x"00A5", 8, '0'));
        chki ("edges of one frame", n_edge, 16);
        pr("    second request ignored: one frame, 16 edges, slave saw a5");

        ----------------------------------------------------------------------
        pr("  G5 reset during a frame");
        set_cfg('1', '0', '0', 16, 3, 3, 2, 2, 2);
        slv_word <= x"BEEF";
        fire(x"DEAD");
        for i in 1 to 20 loop wait until rising_edge(clk); end loop;
        chki("frame really started", to_integer(unsigned'('0' & cs_any)), 1);
        wait until falling_edge(clk);
        rst_n <= '0';
        wait until falling_edge(clk);
        chki("all selects released", to_integer(unsigned'('0' & cs_any)), 0);
        chki("no done on reset", to_integer(unsigned'('0' & done)), 0);
        chki("not busy", to_integer(unsigned'('0' & busy)), 0);
        chki("sclk defined 0", to_integer(unsigned'('0' & sl(sclk = '0'))), 1);
        pr("    reset mid-frame: cs released, no done, sclk=0 (not cpol)");
        wait until falling_edge(clk);
        rst_n <= '1';
        for i in 1 to 3 loop wait until rising_edge(clk); end loop;

        ----------------------------------------------------------------------
        pr("  G6 abort mid-frame");
        set_cfg('0', '0', '0', 16, 1, 0, 2, 2, 2);
        slv_word <= x"1234";
        fire(x"4321");
        wait_idle(4000, gd, tk);
        rx_before := rx_data;
        slv_word  <= x"FFFF";
        fire(x"0F0F");
        for i in 1 to 24 loop wait until rising_edge(clk); end loop;
        wait until falling_edge(clk);
        abort <= '1';
        wait until falling_edge(clk);
        abort <= '0';
        wait_idle(4000, gd, tk);
        chki ("no done on abort", to_integer(unsigned'('0' & gd)), 0);
        chk16("rx_data unchanged", rx_data, rx_before);
        chki ("selects released", to_integer(unsigned'('0' & cs_any)), 0);
        chki ("sclk back at idle", to_integer(unsigned'('0' & sl(sclk = '0'))), 1);
        if to_integer(unsigned(bits_done)) = 0
           or to_integer(unsigned(bits_done)) >= 16 then
            n_err := n_err + 1;
            pr("    FAIL bits_done not partial: " &
               i0(to_integer(unsigned(bits_done))));
        end if;
        n_chk := n_chk + 1;
        pr("    abort: done=0, rx held " & hex4(rx_data) & ", bits_done=" &
           i0(to_integer(unsigned(bits_done))) & " of 16");

        ----------------------------------------------------------------------
        pr("  G7 configuration captured at acceptance");
        set_cfg('0', '0', '0', 8, 2, 0, 2, 2, 2);
        slv_word <= x"007E";
        fire(x"0081");
        for i in 1 to 5 loop wait until rising_edge(clk); end loop;
        cfg_cpol  <= '1';
        cfg_cpha  <= '1';
        cfg_lsb   <= '1';
        cfg_width <= std_logic_vector(to_unsigned(4, 5));
        cfg_div   <= std_logic_vector(to_unsigned(7, 8));
        cfg_dev   <= std_logic_vector(to_unsigned(3, 2));
        wait_idle(4000, gd, tk);
        chk16("rx used captured cfg", rx_data, exp_master_rx(x"007E", 8, '0'));
        chki ("edges of captured width", n_edge, 16);
        chki ("captured device kept",
              to_integer(unsigned'('0' & sl(cs_n(0) = '1'))), 1);
        pr("    mid-frame rewrite ignored: 16 edges, msb-order rx=" & hex4(rx_data));
        set_cfg('0', '0', '0', 8, 1, 0, 2, 2, 2);

        ----------------------------------------------------------------------
        pr("  G8 X on MISO cannot slip through");
        slv_word <= x"005A";
        force_x  <= '1';
        fire(x"00A5");
        wait_idle(4000, gd, tk);
        n_chk := n_chk + 1;
        if rx_data /= exp_master_rx(x"005A", 8, '0') then
            n_neg := n_neg + 1;
            pr("    X reached rx_data and the !== comparison rejected it");
        else
            n_err := n_err + 1;
            pr("    FAIL an all-X rx_data compared EQUAL -- checker is blind");
        end if;
        force_x <= '0';

        ----------------------------------------------------------------------
        pr("  G9 deliberately wrong expectations must FAIL");
        -- 9d, not a5 or 5a: those two are both EIGHT-BIT PALINDROMES
        --   a5 = 10100101   reversed = 10100101
        --   5a = 01011010   reversed = 01011010
        -- so neither can distinguish MSB-first from LSB-first, and the oracle
        -- self-test below reported a defect in the ORACLE when the only defect was
        -- the choice of stimulus. 9d reverses to b9.
        kk := 0;
        slv_word <= x"009D";
        fire(x"009D");
        wait_idle(4000, gd, tk);
        n_chk := n_chk + 1;
        if rx_data /= (exp_master_rx(x"009D", 8, '0') xor x"0001") then
            kk := kk + 1;
        else
            n_err := n_err + 1;
            pr("    FAIL rx checker accepted a wrong value");
        end if;
        n_chk := n_chk + 1;
        if n_edge /= 15 then
            kk := kk + 1;
        else
            n_err := n_err + 1;
            pr("    FAIL edge checker accepted a wrong count");
        end if;
        n_chk := n_chk + 1;
        if exp_master_rx(x"009D", 8, '1') /= exp_master_rx(x"009D", 8, '0') then
            kk := kk + 1;
        else
            n_err := n_err + 1;
            pr("    FAIL oracle is blind to bit order");
        end if;
        n_neg := n_neg + kk;
        pr("    " & i0(kk) &
           " of 3 wrong expectations rejected; oracle separates 9d from b9");

        ----------------------------------------------------------------------
        if n_err = 0 then
            pr("=== SUMMARY checks=" & i0(n_chk) & " negatives=" & i0(n_neg) &
               " failures=" & i0(n_err) & " : PASS ===");
        else
            pr("=== SUMMARY checks=" & i0(n_chk) & " negatives=" & i0(n_neg) &
               " failures=" & i0(n_err) & " : FAIL ===");
        end if;
        finish;
    end process main;

end architecture tb;

8. What It Measures

The transcript is identical in all three languages, byte for byte:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
=== Chapter 20.3 -- capstone controller, directed suite ===
  G1 mode x bit-order x width
    mode 0 msb : rx(w4,w8,w13,w16) checked, all match
    mode 0 lsb : rx(w4,w8,w13,w16) checked, all match
    mode 1 msb : rx(w4,w8,w13,w16) checked, all match
    mode 1 lsb : rx(w4,w8,w13,w16) checked, all match
    mode 2 msb : rx(w4,w8,w13,w16) checked, all match
    mode 2 lsb : rx(w4,w8,w13,w16) checked, all match
    mode 3 msb : rx(w4,w8,w13,w16) checked, all match
    mode 3 lsb : rx(w4,w8,w13,w16) checked, all match
  G2 measured pin intervals, in clk cycles
    div lead lag idle | t_half t_lead t_lag t_gap
      0    2   2    3 |      1      4     3     6
      1    2   2    3 |      2      8     6    10
      3    2   2    3 |      4     16    12    18
      1    5   4    6 |      2     14    10    16
  G3 illegal width
    width  3 rejected: cfg_err=1 busy=0 cs=idle
    width 17 rejected: cfg_err=1 busy=0 cs=idle
    width  4 accepted: cfg_err=0 done=1
    width 16 accepted: cfg_err=0 done=1
  G4 start while busy
    second request ignored: one frame, 16 edges, slave saw a5
  G5 reset during a frame
    reset mid-frame: cs released, no done, sclk=0 (not cpol)
  G6 abort mid-frame
    abort: done=0, rx held 1234, bits_done=5 of 16
  G7 configuration captured at acceptance
    mid-frame rewrite ignored: 16 edges, msb-order rx=007e
  G8 X on MISO cannot slip through
    X reached rx_data and the !== comparison rejected it
  G9 deliberately wrong expectations must FAIL
    3 of 3 wrong expectations rejected; oracle separates 9d from b9
=== SUMMARY checks=284 negatives=4 failures=0 : PASS ===

The timing table is the interesting part

Group 2 measures four intervals at the pins and finds every one of them an exact function of the configuration:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   t_half = cfg_div + 1                                  system clocks
   t_lead = (cfg_lead + 2) x t_half                      CS low to first edge
   t_lag  = (cfg_lag  + 1) x t_half                      last edge to CS high
   t_gap  = (cfg_idle + 1) x t_half + 2                  CS high to next CS low

Check the last row against the third: at cfg_div = 1, cfg_lead = 5, cfg_lag = 4, cfg_idle = 6 the predicted values are (5+2)x2 = 14, (4+1)x2 = 10 and (6+1)x2+2 = 16, and the measured values are 14, 10 and 16.

The +1 and +2 terms are the ticks the state machine spends leaving a phase. They are margin above REQ-TIM-002 and REQ-TIM-003, not part of them, which is why the bench checks both the requirement's >= form and the exact form. The >= form is what the specification promises; the exact form is what catches a regression, because a controller that loses one half-period of lead still satisfies >= whenever cfg_lead is greater than zero.

The gap is the one interval with a term that is not a multiple of the half-period, and those 2 cycles are worth naming: one is the cycle spent in S_IDLE, and one is the requester's own reaction time. With no request queue, the bench cannot post the next transfer until busy falls. The measured gap is therefore the design's floor plus however long the caller took — which is Chapter 20.2 §6's cost, showing up as an arithmetic term.

9. Five Defects Found Writing This

Every one of these was found by running something, and three of them were found because the three languages disagreed.

D1 — A function-valued wire lags its inputs by a time step

The aligned word needed a name before a bit of it could be taken, because a part-select of a function call does not parse. The obvious name is a wire:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
wire [15:0] ld = load_align(tx_data, cfg_width, cfg_lsb_first);   // WRONG here

That compiles, reads correctly from a testbench, and loaded the shift register with zero. A continuous assignment whose right-hand side is a function call settles a time step after its inputs change, so the clock edge in that same step sampled the previous value. Measured: tx_sr loaded 0000 while the wire read d000 one cycle later, and every transmitted bit was zero while every other captured field was correct.

The fix is to compute it inside the clocked block, where the function is called at the edge with the values the edge itself sampled. Note that this is the opposite of Module 19's VHDL lesson: a variable is wrong for state another process reads, and right for a value used and discarded inside one invocation.

D2 — Shift-then-present is correct in two modes and wrong in two

On a launch event, the choice is present the top bit and then shift, or shift and then present the new top bit. The second is what a first version naturally does, and it is correct for CPHA = 0 — because a bit was already presented when the chip select asserted, so the next launch genuinely wants the following bit.

For CPHA = 1 there is no pre-launch. Bit 0 is still pending at the first leading edge, so shifting first skips it.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   modes 0 and 2 : all four widths passed
   modes 1 and 3 : every word shifted up one position, in BOTH directions at once

Both directions moved together because the same ordering is used by the controller and by the device model. The fix is one uniform rule — present, then shift — used by the pre-launch and by every launch, which is why the code has the same two lines in both places.

D3 — Driving start on the edge that samples it

The bench drove start high, waited for the sampling edge, and cleared it. The clearing assignment and the controller's clocked block are in the same region of the same time step, and their order is not defined by the language.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   frame 1        accepted
   frames 2..32   silently dropped

31 of 32 transfers compared a stale rx_data against a fresh expectation, so the suite reported a wall of mismatches that had nothing to do with the controller. Every input is now driven on the negedge, so start is high across exactly one sampling edge and nothing the bench does can collide with the edge that matters. That is Chapter 16.3's clocking-block discipline written out by hand.

D4 — A function evaluated outside its domain, which only VHDL admitted

load_align shifts by DATA_W - w. For the rejected width of 17 that is -1, and it was being computed before the width was checked:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   VHDL      ** Fatal: value -1 outside of NATURAL range 0 to 2147483647
   Verilog   computed an unsigned wrap, produced a garbage word, discarded it
             with the refused request, and said nothing

Both languages were evaluating a function outside its domain. Only one of them said so. The fix — validate the width before computing anything from it — went into all three, and this is the clearest case in the module of the third language acting as verification rather than as a translation.

D5 — The monitor counted the idle re-park as a frame edge

The controller parks SCLK at the configured idle level and asserts the chip select in the same cycle. When the previous frame's polarity differed, those two coincide, so an edge counter gated only on a device is selected counts the re-park.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   VHDL      17 edges on the first CPOL=1 frame
   SV        16 edges, identical received data

A one-cycle difference in when the monitor left reset decided whether the re-park landed inside the window. Gating the counter on a device was already selected last cycle is both the fix and the correct definition: a transition in the assert cycle is the clock reaching its idle level, not a clocking edge. The measurement no longer depends on that phase.

10. The Tri-HDL Result

SystemVerilogVerilog-2001VHDL-2008
Tooliverilog -g2012iverilog -g2001nvc 1.23.0
RTL lines401401361
Testbench lines621621687
Checks284284284
Failures000
Transcript34 lines34 lines34 lines

The three transcripts are byte-identical, and the three simulations end at the same simulated time. Not one line differs by design in this chapter.

That identity is the actual cross-language evidence. Reading three files side by side proves they look alike; running them under one stimulus and comparing the output proves they behave alike, and only the second is evidence. Parity was checked across all thirteen axes — ports, widths, generics, clock edge, reset, priority, counter widths, output latency, bit order, handshake, abort, error semantics, and the shift alignment — but the byte-identical transcript is what makes the claim checkable rather than asserted.

11. Summary

The capstone controller exists in three languages, passes 284 checks in each, and produces byte-identical transcripts. All four SPI modes come out of ~edge_i[0] and two conditional expressions with no mention of CPOL, and an even transition count gives the parked-level requirement away for free.

The measured timing is an exact function of the configuration — t_half = cfg_div + 1, t_lead = (cfg_lead + 2) x t_half, t_lag = (cfg_lag + 1) x t_half, and a gap whose two extra cycles are S_IDLE plus the requester's own reaction time. That last term is the absence of a request queue showing up as arithmetic, and it means no back-to-back throughput figure from this interface describes the controller alone.

Five defects were found by running things. A function-valued wire settled a time step late and loaded zeros into the shift register while reading correctly from the bench. Shift-then-present was correct in modes 0 and 2 and off by one bit in modes 1 and 3, in both directions at once. Driving start on the sampling edge dropped 31 of 32 transfers. A width of 17 was shifted by −1, which VHDL called fatal and Verilog silently wrapped. And an edge counter gated only on selected counted the idle re-park, which differed between languages by one cycle of reset timing.

Three of those five were found because the three implementations disagreed, which is the argument for the third language stated as a measurement rather than a principle.

12. What Comes Next

Chapter 20.4 leaves the cycle counts behind and turns them into nanoseconds: the MISO round trip that decides the maximum SCLK rate, the constraints that tell a tool what this design's SCLK actually is, and an honest account of which crossings here are clock-domain problems and which are static timing problems wearing the wrong name.

Continue learning