Skip to content
VLSI Mentor

SPI · Module 13

Multi-Byte and Back-to-Back Transactions

Keeping the wire busy between frames with one register's worth of flops: why one deep is enough and what a deeper FIFO actually buys, why a stall must be reported rather than absorbed, and why an overrun is lost data.

Everything so far transfers one frame correctly. A real transaction is a stream, and the gap between its frames is where the throughput arithmetic of Module 9 is won or lost.

A flash page program is 4 command and address bytes followed by 256 data bytes. Between which two of those 260 frames is the wire idle?

With a single transmit register, between all 259 boundaries — and for however long the CPU takes to notice each one. The fix is one register deep, and it is worth being precise about why one is enough.

1. The Problem With A Single Transmit Register

The datapath of Chapter 13.6 holds the word being shifted. If that is the only place a word can live, the next word cannot be written until the current one has left — and "has left" means the frame is over:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   shift word N ... frame ends ... CPU notices ... CPU writes word N+1
                ... controller starts ... shift word N+1

Every one of those ellipses is dead time on SCLK, and none of it is under the hardware's control: it is however long the CPU takes to notice. At 50 MHz SCLK and a 100 MHz CPU, a byte takes 160 ns and an interrupt-driven driver can easily spend more time between bytes than inside them. The transfer is correct and the wire is half idle.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   8-bit frame at div=4, 100 MHz clock      320 ns
   interrupt latency + handler              1000 ns, typically
   wire utilisation                         24%

That number is not a pathological case. It is what an interrupt-per-byte driver achieves, and it is why the register below exists.

2. One Register Deep, And What Deeper Buys

Add a shadow register in front of the datapath. The CPU writes the shadow; the controller moves the shadow into the datapath between frames. The two are decoupled, so the CPU can write word N+1 at any point while word N is still shifting — an entire frame's worth of time instead of an inter-frame gap.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   without:   CPU has [inter-frame gap] to produce the next word
   with:      CPU has [a whole frame]  to produce the next word

One deep is enough to remove the dependency, and this is worth stating carefully because it is where intuition goes wrong:

A deeper FIFO buys latency TOLERANCE, not throughput.

Once the producer can stay one word ahead, the wire is already saturated and nothing deeper makes it faster. What depth buys is survival of a longer stall: a four-deep FIFO lets the CPU disappear for four frame times instead of one without the wire going idle. That is a real and useful property — it is what lets a driver batch its work — and it is not a throughput property.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   depth 1   producer must return within 1 frame time   wire saturated
   depth 4   producer must return within 4 frame times  wire saturated
   depth 16  producer must return within 16 frame times wire saturated

All three saturate the wire. They differ in how long the producer may be absent, which is a question about the system, not about the interface. So the right depth is decided by the worst-case interrupt latency divided by the frame time — and for a design where that ratio is below one, depth one is the correct answer and anything more is flops spent on nothing.

3. What The Block Reports

A design that silently absorbs a slow producer is worse than one that says so, because the lost throughput is invisible. So two flags:

stalled — the transaction is open, the clock is not running, and nothing is queued. The wire is idle because the CPU was late. Every cycle of this is throughput that was paid for and not used. It is not an error: it is the exact measure of how late the producer was, and a driver that is late by design — a low-priority background transfer, say — will assert it constantly and correctly.

rx_overrun — a received word was overwritten before it was read. Unlike a stall this is lost data, and the flag must be sticky: a receive path that drops a word and carries on is how a flash image ends up with one wrong word in it, and a driver that only checks at the end of a transaction must still find out.

The distinction between the two is worth keeping in mind as a general pattern:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   a performance flag reports a COST      -- not an error, clear on read is fine
   a data-loss flag reports a LOSS        -- an error, must be sticky

4. The Tail Is Not A Stall

stalled has a subtlety that produced a false report in the first version of this block, and it generalises.

After the final frame of a transaction, the chip-select controller of Chapter 13.7 walks out through LAG and GAP. Throughout that, the transaction is still open — busy is asserted — the clock is not running, and nothing is queued, because by design there is nothing more to send.

That is exactly the condition stalled tests for. Without a guard, every transaction ends with a burst of stall cycles equal to the lag plus the gap, and a driver reading the counter concludes it is always late. A flag that fires on correct behaviour is a flag that gets ignored, and then it is worse than absent.

So the block tracks tail: set when the final frame of a transaction starts, cleared on the next load, and ANDed out of stalled. Two flops, and the flag now means what it says.

5. The Two Cases

Twenty cycles across five rows. A prompt-producer SCLK row shows four frames separated only by the programmed two-cycle hold. A late-producer SCLK row shows the same four frames separated by much longer idle intervals. A shadow-full row shows the slot occupied for most of each frame in the prompt case. A stalled row is flat in the prompt case and asserted during the idle intervals in the late case.shadow freed at load, not at frame endshadow freed at load, notat frame endlate: four idle cycles, countedlate: four idle cycles,countedtail — not a stall, by designtail — not a stall, bydesignprompt sclksh_fulllate sclklate stalltailt0t1t2t3t4t5t6t7t8t9t10t11t12t13t14t15t16t17t18t19
Figure 1 — the same four-frame transaction with a prompt producer and a late one. With the shadow refilled during each frame, the only idle cycles are the programmed inter-frame hold. With the producer arriving late, the wire sits idle and the stall counter records exactly how long — the transfer is still correct and the cost is now visible.

Read the sh_full row against the prompt sclk row. The slot empties at cycle 3 — the load — and refills at cycle 5, giving the CPU a window that spans the rest of the frame. And read the tail row at the far right: the final frame has started, so the idle cycles that follow are the transaction closing, not the producer being late.

6. Building the Streaming Controller — Three HDLs

The circuit

A one-word shadow with a valid bit, an armed flag, a one-word receive holding register, and the two flags. The decisions:

The datapath may only be loaded while the clock is stopped. Loading it mid-frame would overwrite the word being shifted — the one failure a single-register design cannot even express, and that a shadow design has to be careful about. The guard is !shift_en.

armed stops a second load landing on the first. It means "the datapath is loaded and a frame has been requested, but the frame has not started yet". Without it, a shadow refilled during the lead time would overwrite the word the controller is about to send.

A push and a load on the same cycle both succeed. The shadow is freed by the load and filled by the push, and the two are ordered so neither is lost. That matters more than it looks: the CPU polling tx_ready and writing the instant it goes high will hit exactly this cycle, every time, on a tight loop.

flush throws queued work away. A word queued for a transaction that is being abandoned must not become the first word of the next one, and an armed request left standing would start that next transaction by itself the moment the bus came free. Chapter 13.10 drives it; here it is simply the one input that can discard work.

clr_flags clears the sticky overrun, and the clear is applied first. A clear arriving on the same cycle as a fresh overrun must not swallow it — losing a report of lost data is the same bug twice.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_multibyte.sv — one word of shadow, and the flags that make its cost visible
// spi_multibyte.sv
//
// Chapter 13.9 -- keeping the wire busy between frames.
//
// Everything built so far transfers ONE frame correctly. A real transaction is
// a stream: a flash page program is a command byte, three address bytes and
// then 256 data bytes, all under one chip select. Getting each of those bytes
// right individually is not the same as getting the stream right, and the gap
// between them is where the throughput of Chapter 9 is won or lost.
//
// THE PROBLEM WITH A SINGLE TRANSMIT REGISTER.
//
// The datapath of 13.6 holds the word being shifted. If that is the only place
// a word can live, then the next byte cannot be written until the current one
// has left -- and "has left" means the frame is over. So the sequence is:
//
//     shift byte N ... frame ends ... CPU notices ... CPU writes byte N+1
//     ... controller starts ... shift byte N+1
//
// Every one of those dots is dead time on SCLK, and none of it is under the
// hardware's control: it is however long the CPU takes to notice. At 50 MHz
// SCLK and a 100 MHz CPU, an interrupt-driven driver can easily spend more
// time between bytes than inside them. The transfer is correct and the wire is
// half idle.
//
// THE FIX IS ONE REGISTER DEEP.
//
// Add a SHADOW register in front of the datapath. The CPU writes into the
// shadow; the controller moves the shadow into the datapath between frames.
// The two are decoupled, so the CPU can write byte N+1 at any point while byte
// N is still shifting -- it has an entire frame's worth of time to do it,
// instead of the inter-frame gap.
//
// One deep is enough to remove the dependency, and that is worth being clear
// about: a deeper FIFO buys latency TOLERANCE (surviving a longer CPU stall),
// not extra throughput. Once the producer can stay one byte ahead, the wire is
// already saturated and nothing deeper makes it faster.
//
// WHAT THIS BLOCK REPORTS.
//
// A design that silently absorbs a slow producer is worse than one that says
// so, because the lost throughput is invisible. So:
//
//   stalled    the transaction is open, the clock is not running, and there is
//              nothing queued -- the wire is idle because the CPU was late.
//              Every cycle of this is throughput that was paid for and not
//              used.
//   rx_overrun a received word was overwritten before it was read. Unlike a
//              stall this is not a performance problem, it is LOST DATA, and
//              it must be sticky: a receive path that drops a byte and carries
//              on is how a flash image ends up with one wrong word in it.

module spi_multibyte #(
    parameter int MAX_W = 32,
    parameter int LEN_W = 6
) (
    input  wire               clk,
    input  wire               rst_n,

    // --- CPU side -------------------------------------------------------
    input  wire               tx_push,     // write a word into the shadow
    input  wire [MAX_W-1:0]   tx_wdata,
    input  wire               tx_last,     // ... and end the transaction after it
    output wire               tx_ready,    // the shadow is free

    input  wire               flush,       // drop whatever is queued
    input  wire               rx_pop,      // read the received word
    output wire [MAX_W-1:0]   rx_rdata,
    output wire               rx_ready,    // a received word is waiting
    output reg                rx_overrun,  // sticky: a word was lost
    // Clearing a sticky flag is the REGISTER INTERFACE's business, not this
    // block's -- but the flag lives here, because this is where the loss is
    // detected. So each flag has exactly one owner and exactly one clear path,
    // and Chapter 13.11 wires this to a write of the command register.
    input  wire               clr_flags,

    // --- transfer engine side (13.7 and 13.6/13.8) ----------------------
    input  wire               core_busy,   // a transaction is open
    input  wire               shift_en,    // the clock is running
    input  wire               start_stb,   // the engine took a frame
    input  wire               rx_valid_stb,
    input  wire [MAX_W-1:0]   rx_word,

    output wire               req,         // to the chip-select controller
    output wire               hold,
    output reg  [MAX_W-1:0]   tx_data,     // to the shift datapath
    output wire               load_stb,

    output wire               stalled      // the wire is idle for want of data
);

    // The shadow: one word, its end-of-transaction flag, and a valid bit.
    reg [MAX_W-1:0] sh_data;
    reg             sh_last;
    reg             sh_valid;

    // `armed` means the datapath has been loaded and the controller has been
    // asked for a frame, but the frame has not started yet. It stops a second
    // load from landing on top of the first.
    reg             armed;
    reg             armed_last;
    reg             load_r;

    // `tail` means the final frame of the transaction has started, so the
    // controller is walking out through its lag and gap with nothing more to
    // send. Without it the closing states look exactly like a producer that
    // went quiet, and `stalled` would accuse the CPU of being late at the end
    // of every single transaction.
    reg             tail;

    assign tx_ready = ~sh_valid;
    assign req      = armed;
    assign hold     = ~armed_last;
    assign load_stb = load_r;

    // The datapath may only be loaded while the clock is stopped. Loading it
    // mid-frame would overwrite the word being shifted -- the one failure
    // that a single-register design cannot even express, and that a shadow
    // design has to be careful about.
    wire can_load = sh_valid & ~armed & ~shift_en & ~flush;

    // Idle inside an open transaction with nothing to send. Note it is NOT an
    // error: it is the exact measure of how late the producer was.
    assign stalled = core_busy & ~shift_en & ~armed & ~sh_valid & ~tail;

    reg [MAX_W-1:0] rx_hold;
    reg             rx_full;

    assign rx_rdata = rx_hold;
    assign rx_ready = rx_full;

    always_ff @(posedge clk or negedge rst_n) begin
        if (!rst_n) begin
            sh_data    <= {MAX_W{1'b0}};
            sh_last    <= 1'b0;
            sh_valid   <= 1'b0;
            armed      <= 1'b0;
            armed_last <= 1'b0;
            load_r     <= 1'b0;
            tail       <= 1'b0;
            tx_data    <= {MAX_W{1'b0}};
            rx_hold    <= {MAX_W{1'b0}};
            rx_full    <= 1'b0;
            rx_overrun <= 1'b0;
        end else begin
            load_r <= 1'b0;

            // --- the CPU writing into the shadow ---
            // Accepted whenever the shadow is free, which includes the same
            // cycle it is being emptied into the datapath below: the two are
            // ordered so a push and a load on one cycle both succeed.
            // FLUSH takes priority over everything on the transmit side. A
            // word queued for a transaction that is being abandoned must not
            // become the first word of the next one, and an `armed` request
            // left standing would start that next transaction by itself the
            // moment the bus came free. Chapter 13.10 is where this is driven
            // from; here it is simply the one input that can throw work away.
            if (flush) begin
                sh_valid   <= 1'b0;
                armed      <= 1'b0;
                armed_last <= 1'b0;
                tail       <= 1'b0;
            end else if (tx_push && tx_ready) begin
                sh_data  <= tx_wdata;
                sh_last  <= tx_last;
                sh_valid <= 1'b1;
            end

            // --- the shadow moving into the datapath ---
            if (can_load && !flush) begin
                tx_data    <= sh_data;
                armed_last <= sh_last;
                armed      <= 1'b1;
                load_r     <= 1'b1;
                tail       <= 1'b0;
                // Freed immediately, so the CPU may refill it on the very
                // next cycle rather than waiting for the frame to finish.
                // This single line is what makes the stream back-to-back.
                if (!(tx_push && tx_ready))
                    sh_valid <= 1'b0;
            end

            // --- the engine taking the frame ---
            if (start_stb) begin
                armed <= 1'b0;
                if (armed_last)
                    tail <= 1'b1;
            end

            // --- the receive side ---
            // The clear is applied FIRST so that a clear arriving on the same
            // cycle as a fresh overrun does not swallow it: the new event wins,
            // because losing a report of lost data is the same bug twice.
            if (clr_flags)
                rx_overrun <= 1'b0;

            if (rx_valid_stb) begin
                rx_hold <= rx_word;
                rx_full <= 1'b1;
                // A word arriving on top of an unread one is DATA LOST, and
                // the flag is sticky so a driver that only checks at the end
                // of the transaction still finds out.
                if (rx_full && !rx_pop)
                    rx_overrun <= 1'b1;
            end else if (rx_pop) begin
                rx_full <= 1'b0;
            end
        end
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_multibyte_tb.sv — the whole engine, a prompt producer and a late one
// spi_multibyte_tb.sv
//
// This is the first testbench that runs the whole engine: the divider of 13.4,
// the mode logic of 13.5, the datapath of 13.6 wrapped by 13.8, the chip-select
// controller of 13.7, and the streaming shadow of 13.9. Only the register
// interface is missing, and that is Chapter 13.11.
//
// The point of the chapter is throughput, so the measurements are about TIME as
// well as data:
//
//   - a prompt producer must stream N bytes under ONE chip select with no
//     stall cycles at all, and with the inter-frame idle equal to the gap the
//     controller was told to insert -- not a cycle more;
//   - a deliberately late producer must still be CORRECT, and must be
//     reported: the stall counter is how the lost throughput becomes visible;
//   - a receiver that does not read fast enough must set a STICKY overrun,
//     because a dropped word is lost data, not a slow day.
//
// Data is checked in both directions against an independent slave, byte by
// byte and in order, so a stream that gets every byte right but in the wrong
// sequence fails.

`timescale 1ns/1ps

module spi_multibyte_tb;

    localparam int MAX_W = 32;
    localparam int LEN_W = 6;
    localparam int DIV_W = 8;
    localparam int N_CS  = 2;
    localparam int SEL_W = 1;
    localparam int CNT_W = 8;

    logic clk = 1'b0;
    logic rst_n = 1'b0;
    always #5 clk = ~clk;

    // --- 13.9, the block under test ---------------------------------------
    logic             tx_push  = 1'b0;
    logic [MAX_W-1:0] tx_wdata = 32'h0;
    logic             tx_last  = 1'b0;
    wire              tx_ready;
    logic             flush     = 1'b0;
    logic             clr_flags = 1'b0;
    logic             rx_pop   = 1'b0;
    wire [MAX_W-1:0]  rx_rdata;
    wire              rx_ready, rx_overrun;

    wire              req, hold, load_stb, stalled;
    wire [MAX_W-1:0]  tx_data;

    // --- 13.7 -------------------------------------------------------------
    logic [SEL_W-1:0] sel = 1'b0;
    logic [CNT_W-1:0] lead_cyc = 8'd3;
    logic [CNT_W-1:0] lag_cyc  = 8'd3;
    logic [CNT_W-1:0] gap_cyc  = 8'd2;
    wire  [N_CS-1:0]  cs_n;
    wire              shift_en, start_stb, busy, sel_err;
    wire [2:0]        state_id;

    // --- 13.4 -------------------------------------------------------------
    logic [DIV_W-1:0] div  = 8'd4;
    logic             cpol = 1'b0;
    wire              sclk, edge_a_stb, edge_b_stb, bit_done, div_err;
    wire [DIV_W-1:0]  half_a, half_b;

    // --- 13.5 / 13.8 ------------------------------------------------------
    logic             cpha      = 1'b0;
    logic [LEN_W-1:0] len       = 6'd8;
    logic             lsb_first = 1'b0;
    wire              preload_stb, launch_stb, capture_stb, frame_done;
    wire [LEN_W-1:0]  bit_idx;
    wire              miso, mosi;
    wire [MAX_W-1:0]  rx_word;
    wire              rx_valid_stb, len_err;

    spi_clkdiv_strobe #(.DIV_W(DIV_W)) u_div (
        .clk(clk), .rst_n(rst_n), .en(shift_en), .div(div), .cpol(cpol),
        .sclk(sclk), .edge_a_stb(edge_a_stb), .edge_b_stb(edge_b_stb),
        .bit_done(bit_done), .half_a(half_a), .half_b(half_b),
        .div_err(div_err)
    );

    spi_mode_edges #(.LEN_W(LEN_W)) u_mode (
        .clk(clk), .rst_n(rst_n), .cpha(cpha), .len(len),
        .active(shift_en), .start_stb(start_stb),
        .edge_a_stb(edge_a_stb), .edge_b_stb(edge_b_stb),
        .preload_stb(preload_stb), .launch_stb(launch_stb),
        .capture_stb(capture_stb), .bit_idx(bit_idx), .frame_done(frame_done)
    );

    spi_width_order #(.MAX_W(MAX_W), .LEN_W(LEN_W)) u_data (
        .clk(clk), .rst_n(rst_n),
        .tx_data(tx_data), .len(len), .lsb_first(lsb_first),
        .load_stb(load_stb),
        .preload_stb(preload_stb), .launch_stb(launch_stb),
        .capture_stb(capture_stb),
        .miso(miso), .mosi(mosi),
        .rx_data(rx_word), .rx_valid_stb(rx_valid_stb), .len_err(len_err)
    );

    spi_cs_ctrl #(.N_CS(N_CS), .SEL_W(SEL_W), .CNT_W(CNT_W)) u_cs (
        .clk(clk), .rst_n(rst_n),
        .req(req), .hold(hold), .sel(sel),
        .lead_cyc(lead_cyc), .lag_cyc(lag_cyc), .gap_cyc(gap_cyc),
        .core_done(frame_done), .abort(1'b0),
        .cs_n(cs_n), .shift_en(shift_en), .start_stb(start_stb),
        .busy(busy), .sel_err(sel_err), .state_id(state_id)
    );

    spi_multibyte #(.MAX_W(MAX_W), .LEN_W(LEN_W)) dut (
        .clk(clk), .rst_n(rst_n),
        .tx_push(tx_push), .tx_wdata(tx_wdata), .tx_last(tx_last),
        .tx_ready(tx_ready),
        .flush(flush), .clr_flags(clr_flags),
        .rx_pop(rx_pop), .rx_rdata(rx_rdata), .rx_ready(rx_ready),
        .rx_overrun(rx_overrun),
        .core_busy(busy), .shift_en(shift_en), .start_stb(start_stb),
        .rx_valid_stb(rx_valid_stb), .rx_word(rx_word),
        .req(req), .hold(hold), .tx_data(tx_data), .load_stb(load_stb),
        .stalled(stalled)
    );

    // --- the independent slave, word by word ------------------------------
    logic [MAX_W-1:0] slv_tx_q [0:127];
    logic [MAX_W-1:0] slv_rx_q [0:127];
    integer           slv_tx_i, slv_rx_i;
    logic [MAX_W-1:0] slv_sr, slv_rx_sr;
    logic             slave_bit_r;

    assign miso = slave_bit_r;

    always_ff @(posedge clk) begin
        if (load_stb) begin
            slv_sr   <= slv_tx_q[slv_tx_i] << (MAX_W - len);
            slv_tx_i <= slv_tx_i + 1;
        end else if (preload_stb || launch_stb) begin
            slave_bit_r <= slv_sr[MAX_W-1];
            slv_sr      <= {slv_sr[MAX_W-2:0], 1'b0};
        end

        // Cleared at each load, so the recorded word holds THIS frame's bits
        // and not a running history of every frame before it.
        if (load_stb)
            slv_rx_sr <= {MAX_W{1'b0}};
        else if (capture_stb)
            slv_rx_sr <= {slv_rx_sr[MAX_W-2:0], mosi};

        // `rx_valid_stb` is one cycle after the final capture, so the slave's
        // own shift register is already complete when it is recorded.
        if (rx_valid_stb) begin
            slv_rx_q[slv_rx_i] <= slv_rx_sr;
            slv_rx_i           <= slv_rx_i + 1;
        end
    end

    // --- the throughput monitor -------------------------------------------
    integer stall_cycles;
    integer idle_in_txn;      // cycles inside a transaction with no clock
    integer frames_started;
    integer cs_falls;
    integer slack_cycles;     // room in the shadow WHILE the wire is busy
    integer valid_pulses;
    integer n_loads, n_caps, n_pre, n_lau, n_race;
    logic   busy_q;

    always_ff @(posedge clk) begin
        if (!rst_n) begin
            busy_q <= 1'b0;
        end else begin
            busy_q <= busy;
            if (stalled)               stall_cycles   <= stall_cycles + 1;
            if (busy && !shift_en)     idle_in_txn    <= idle_in_txn + 1;
            if (start_stb)             frames_started <= frames_started + 1;
            if (busy && !busy_q)       cs_falls       <= cs_falls + 1;
            if (tx_ready && shift_en)  slack_cycles   <= slack_cycles + 1;
            if (rx_valid_stb)          valid_pulses   <= valid_pulses + 1;
            if (load_stb)              n_loads        <= n_loads + 1;
            if (capture_stb)           n_caps         <= n_caps + 1;
            if (preload_stb)           n_pre          <= n_pre + 1;
            if (launch_stb)            n_lau          <= n_lau + 1;
            // The datapath is loaded by one strobe and told to drive its first
            // bit by another. If they ever land on the same cycle the load is
            // overwritten by the shift in the same always block and the frame
            // sends the PREVIOUS word. The handover is built so they cannot,
            // and this counts the cycles on which that claim would be false.
            if (load_stb && start_stb) n_race         <= n_race + 1;
        end
    end

    integer errors = 0;

    task automatic clear_counts;
        begin
            stall_cycles = 0; idle_in_txn = 0; frames_started = 0;
            cs_falls = 0; slack_cycles = 0; valid_pulses = 0;
            n_loads = 0; n_caps = 0; n_pre = 0; n_lau = 0; n_race = 0;
        end
    endtask

    // The CPU writing a word. `late` cycles of dithering before the write is
    // how a slow driver is modelled.
    task automatic push(input [MAX_W-1:0] w, input bit last,
                        input integer late);
        integer guard;
        begin
            repeat (late) @(negedge clk);
            guard = 8000;
            while (!tx_ready && guard > 0) begin
                @(negedge clk);
                guard = guard - 1;
            end
            if (guard == 0) begin
                $display("  FAIL: the shadow never became free");
                errors = errors + 1;
            end
            tx_wdata = w;
            tx_last  = last;
            tx_push  = 1'b1;
            @(negedge clk);
            tx_push = 1'b0;
        end
    endtask

    // The CPU reading. `drain` = 0 leaves words unread on purpose.
    integer got_n;
    logic [MAX_W-1:0] got_q [0:127];

    task automatic drain_rx;
        begin
            while (rx_ready) begin
                got_q[got_n] = rx_rdata;
                got_n        = got_n + 1;
                rx_pop       = 1'b1;
                @(negedge clk);
                rx_pop       = 1'b0;
                @(negedge clk);
            end
        end
    endtask

    task automatic wait_idle;
        integer guard;
        begin
            guard = 20000;
            while (busy && guard > 0) begin
                drain_rx();
                @(negedge clk);
                guard = guard - 1;
            end
            repeat (4) @(negedge clk);
            drain_rx();
        end
    endtask

    logic [MAX_W-1:0] sent_q [0:127];
    integer i, k, nb, late, seed, bad;

    initial begin
        clear_counts();
        slv_tx_i = 0; slv_rx_i = 0; got_n = 0;
        slv_sr = 32'h0; slv_rx_sr = 32'h0; slave_bit_r = 1'b0;
        seed = 32'h5EED_1234;
        for (i = 0; i < 128; i = i + 1) begin
            seed        = (seed * 32'h0019_660D) + 32'h3C6E_F35F;
            slv_tx_q[i] = (seed >> 7) & 32'hFF;
        end

        repeat (3) @(negedge clk);
        rst_n = 1'b1;
        @(negedge clk);

        // 1. A PROMPT PRODUCER, sixteen bytes, one chip select. This is the
        //    case the shadow register exists for.
        clear_counts();
        nb = 16;
        for (i = 0; i < nb; i = i + 1) begin
            seed      = (seed * 32'h0019_660D) + 32'h3C6E_F35F;
            sent_q[i] = (seed >> 11) & 32'hFF;
        end
        got_n = 0;
        for (i = 0; i < nb; i = i + 1) begin
            push(sent_q[i], (i == nb - 1), 0);
            drain_rx();
        end
        wait_idle();

        if (cs_falls != 1) begin
            $display("  FAIL: a %0d-byte stream took %0d chip selects",
                     nb, cs_falls);
            errors = errors + 1;
        end
        if (frames_started != nb) begin
            $display("  FAIL: %0d bytes produced %0d frames", nb,
                     frames_started);
            errors = errors + 1;
        end
        if (stall_cycles != 0) begin
            $display("  FAIL: a prompt producer still stalled the wire for %0d cycles",
                     stall_cycles);
            errors = errors + 1;
        end
        // Everything the slave received, in order.
        bad = 0;
        for (i = 0; i < nb; i = i + 1)
            if (slv_rx_q[i] !== sent_q[i]) bad = bad + 1;
        if (bad != 0) begin
            $display("  FAIL: the slave received %0d of %0d bytes wrongly",
                     bad, nb);
            errors = errors + 1;
        end
        // Everything the master received, in order.
        bad = 0;
        if (got_n != nb) begin
            $display("  FAIL: the master read back %0d of %0d bytes",
                     got_n, nb);
            errors = errors + 1;
        end else begin
            for (i = 0; i < nb; i = i + 1)
                if (got_q[i] !== slv_tx_q[i]) bad = bad + 1;
            if (bad != 0) begin
                $display("  FAIL: the master received %0d of %0d bytes wrongly",
                         bad, nb);
                errors = errors + 1;
            end
        end
        if (rx_overrun) begin
            $display("  FAIL: a fully drained receiver reported an overrun");
            errors = errors + 1;
        end
        // The per-frame bookkeeping, counted over the whole stream rather
        // than trusted: one load and one preload per frame, len captures per
        // frame, and len-1 launches because CPHA=0's preload is the first
        // drive (Chapter 13.5).
        if (n_loads != nb || n_pre != nb) begin
            $display("  FAIL: %0d frames produced %0d loads and %0d preloads",
                     nb, n_loads, n_pre);
            errors = errors + 1;
        end
        if (n_caps != nb * 8 || n_lau != nb * 7) begin
            $display("  FAIL: %0d eight-bit frames produced %0d captures and %0d launches, expected %0d and %0d",
                     nb, n_caps, n_lau, nb * 8, nb * 7);
            errors = errors + 1;
        end
        if (n_race != 0) begin
            $display("  FAIL: a load landed on a frame start %0d times",
                     n_race);
            errors = errors + 1;
        end
        if (valid_pulses != nb) begin
            $display("  FAIL: %0d frames raised %0d received-word pulses",
                     nb, valid_pulses);
            errors = errors + 1;
        end
        $display("  per-frame bookkeeping over the stream: %0d loads, %0d preloads, %0d launches, %0d captures, %0d received words, and no load ever landed on a frame start",
                 n_loads, n_pre, n_lau, n_caps, valid_pulses);
        $display("  16 bytes, one chip select, %0d frames, %0d stall cycles, %0d idle cycles inside the transaction -- %0d per byte boundary",
                 frames_started, stall_cycles, idle_in_txn,
                 idle_in_txn / (nb - 1));

        // 2. THE IDLE BETWEEN FRAMES IS THE GAP THAT WAS ASKED FOR, and not a
        //    cycle more. This is the check that would catch a shadow register
        //    that works but hands over one cycle late.
        if (idle_in_txn > (nb - 1) * (gap_cyc + 3)) begin
            $display("  FAIL: %0d idle cycles across %0d boundaries, more than the %0d-cycle gap explains",
                     idle_in_txn, nb - 1, gap_cyc);
            errors = errors + 1;
        end
        $display("  the inter-frame idle is the programmed %0d-cycle hold and nothing else",
                 gap_cyc);

        // 3. A LATE PRODUCER. Still correct, and now visibly costly.
        clear_counts();
        nb = 8;
        for (i = 0; i < nb; i = i + 1) begin
            seed      = (seed * 32'h0019_660D) + 32'h3C6E_F35F;
            sent_q[i] = (seed >> 11) & 32'hFF;
        end
        k     = slv_tx_i;
        got_n = 0;
        for (i = 0; i < nb; i = i + 1) begin
            push(sent_q[i], (i == nb - 1), 40);
            drain_rx();
        end
        wait_idle();

        if (stall_cycles == 0) begin
            $display("  FAIL: a producer 40 cycles late stalled the wire not at all -- the counter cannot be working");
            errors = errors + 1;
        end
        bad = 0;
        for (i = 0; i < nb; i = i + 1)
            if (slv_rx_q[k + i] !== sent_q[i]) bad = bad + 1;
        if (bad != 0) begin
            $display("  FAIL: a late producer corrupted %0d of %0d bytes",
                     bad, nb);
            errors = errors + 1;
        end
        if (cs_falls != 1) begin
            $display("  FAIL: a late producer broke the transaction into %0d chip selects",
                     cs_falls);
            errors = errors + 1;
        end
        $display("  the same stream with the producer 40 cycles late: every byte still correct, one chip select, but %0d stall cycles the fast run did not have",
                 stall_cycles);

        // 4. AN OVERRUN IS STICKY. Stream four bytes and never read one.
        clear_counts();
        nb = 4;
        for (i = 0; i < nb; i = i + 1)
            push(32'h5A, (i == nb - 1), 0);
        // Deliberately no drain_rx here.
        i = 20000;
        while (busy && i > 0) begin
            @(negedge clk);
            i = i - 1;
        end
        repeat (4) @(negedge clk);
        if (!rx_overrun) begin
            $display("  FAIL: four unread received words did not raise an overrun");
            errors = errors + 1;
        end
        $display("  four received words left unread: overrun raised and latched");

        // It must STAY raised through a clean transfer -- a flag that clears
        // itself is a flag that hides the one byte that was lost.
        drain_rx();
        clear_counts();
        for (i = 0; i < 2; i = i + 1) begin
            push(32'h33, (i == 1), 0);
            drain_rx();
        end
        wait_idle();
        if (!rx_overrun) begin
            $display("  FAIL: the overrun flag cleared itself on the next clean transfer");
            errors = errors + 1;
        end
        $display("  a clean transfer afterwards does not clear it -- the lost byte stays reported");

        // 5. THE CPU CAN REFILL WHILE THE WIRE IS BUSY. Without that the
        //    shadow buys nothing, so it is checked directly: at least one push
        //    must be accepted with the clock running.
        clear_counts();
        // Reset the sticky flag by resetting the block, since nothing else
        // clears it -- which is itself the design decision under test.
        rst_n = 1'b0;
        repeat (3) @(negedge clk);
        rst_n = 1'b1;
        repeat (2) @(negedge clk);
        slv_tx_i = 0; slv_rx_i = 0; got_n = 0;

        // What matters is not the instant the testbench happens to write, but
        // HOW LONG the shadow stays free while the wire is busy -- that window
        // is the slack the producer is given, and without the shadow it would
        // be zero.
        for (i = 0; i < 8; i = i + 1) begin
            push(32'hC5, (i == 7), 0);
            drain_rx();
        end
        wait_idle();
        if (slack_cycles == 0) begin
            $display("  FAIL: the shadow was never free while the clock was running -- it is not decoupling anything");
            errors = errors + 1;
        end
        if (stall_cycles != 0) begin
            $display("  FAIL: %0d stall cycles while refilling ahead",
                     stall_cycles);
            errors = errors + 1;
        end
        $display("  the shadow stood free for %0d of the cycles the wire was busy -- that window is the slack the producer gets, and it is a whole frame wide",
                 slack_cycles);

        if (errors == 0)
            $display("PASS: a one-deep shadow register decouples the producer from the wire -- sixteen bytes stream under a single chip select with every byte correct in both directions and in order, zero stall cycles, and an inter-frame idle equal to the programmed hold and nothing more -- words are accepted while the previous frame is still shifting, which is the whole reason the shadow exists -- a producer forty cycles late still transfers every byte correctly under one chip select but now reports the stall cycles it cost, so the lost throughput is visible instead of silent -- and a received word overwritten before it was read raises a sticky overrun that a later clean transfer does not clear");
        else
            $display("FAIL: %0d error(s)", errors);
        $finish;
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_multibyte.v — the same controller in Verilog-2001
// spi_multibyte.v
//
// Chapter 13.9 -- keeping the wire busy between frames.
//
// Everything built so far transfers ONE frame correctly. A real transaction is
// a stream: a flash page program is a command byte, three address bytes and
// then 256 data bytes, all under one chip select. Getting each of those bytes
// right individually is not the same as getting the stream right, and the gap
// between them is where the throughput of Chapter 9 is won or lost.
//
// THE PROBLEM WITH A SINGLE TRANSMIT REGISTER.
//
// The datapath of 13.6 holds the word being shifted. If that is the only place
// a word can live, then the next byte cannot be written until the current one
// has left -- and "has left" means the frame is over. So the sequence is:
//
//     shift byte N ... frame ends ... CPU notices ... CPU writes byte N+1
//     ... controller starts ... shift byte N+1
//
// Every one of those dots is dead time on SCLK, and none of it is under the
// hardware's control: it is however long the CPU takes to notice. At 50 MHz
// SCLK and a 100 MHz CPU, an interrupt-driven driver can easily spend more
// time between bytes than inside them. The transfer is correct and the wire is
// half idle.
//
// THE FIX IS ONE REGISTER DEEP.
//
// Add a SHADOW register in front of the datapath. The CPU writes into the
// shadow; the controller moves the shadow into the datapath between frames.
// The two are decoupled, so the CPU can write byte N+1 at any point while byte
// N is still shifting -- it has an entire frame's worth of time to do it,
// instead of the inter-frame gap.
//
// One deep is enough to remove the dependency, and that is worth being clear
// about: a deeper FIFO buys latency TOLERANCE (surviving a longer CPU stall),
// not extra throughput. Once the producer can stay one byte ahead, the wire is
// already saturated and nothing deeper makes it faster.
//
// WHAT THIS BLOCK REPORTS.
//
// A design that silently absorbs a slow producer is worse than one that says
// so, because the lost throughput is invisible. So:
//
//   stalled    the transaction is open, the clock is not running, and there is
//              nothing queued -- the wire is idle because the CPU was late.
//              Every cycle of this is throughput that was paid for and not
//              used.
//   rx_overrun a received word was overwritten before it was read. Unlike a
//              stall this is not a performance problem, it is LOST DATA, and
//              it must be sticky: a receive path that drops a byte and carries
//              on is how a flash image ends up with one wrong word in it.

module spi_multibyte #(
    parameter MAX_W = 32,
    parameter LEN_W = 6
) (
    input  wire               clk,
    input  wire               rst_n,

    // --- CPU side -------------------------------------------------------
    input  wire               tx_push,     // write a word into the shadow
    input  wire [MAX_W-1:0]   tx_wdata,
    input  wire               tx_last,     // ... and end the transaction after it
    output wire               tx_ready,    // the shadow is free

    input  wire               flush,       // drop whatever is queued
    input  wire               rx_pop,      // read the received word
    output wire [MAX_W-1:0]   rx_rdata,
    output wire               rx_ready,    // a received word is waiting
    output reg                rx_overrun,  // sticky: a word was lost
    // Clearing a sticky flag is the REGISTER INTERFACE's business, not this
    // block's -- but the flag lives here, because this is where the loss is
    // detected. So each flag has exactly one owner and exactly one clear path,
    // and Chapter 13.11 wires this to a write of the command register.
    input  wire               clr_flags,

    // --- transfer engine side (13.7 and 13.6/13.8) ----------------------
    input  wire               core_busy,   // a transaction is open
    input  wire               shift_en,    // the clock is running
    input  wire               start_stb,   // the engine took a frame
    input  wire               rx_valid_stb,
    input  wire [MAX_W-1:0]   rx_word,

    output wire               req,         // to the chip-select controller
    output wire               hold,
    output reg  [MAX_W-1:0]   tx_data,     // to the shift datapath
    output wire               load_stb,

    output wire               stalled      // the wire is idle for want of data
);

    // The shadow: one word, its end-of-transaction flag, and a valid bit.
    reg [MAX_W-1:0] sh_data;
    reg             sh_last;
    reg             sh_valid;

    // `armed` means the datapath has been loaded and the controller has been
    // asked for a frame, but the frame has not started yet. It stops a second
    // load from landing on top of the first.
    reg             armed;
    reg             armed_last;
    reg             load_r;

    // `tail` means the final frame of the transaction has started, so the
    // controller is walking out through its lag and gap with nothing more to
    // send. Without it the closing states look exactly like a producer that
    // went quiet, and `stalled` would accuse the CPU of being late at the end
    // of every single transaction.
    reg             tail;

    assign tx_ready = ~sh_valid;
    assign req      = armed;
    assign hold     = ~armed_last;
    assign load_stb = load_r;

    // The datapath may only be loaded while the clock is stopped. Loading it
    // mid-frame would overwrite the word being shifted -- the one failure
    // that a single-register design cannot even express, and that a shadow
    // design has to be careful about.
    wire can_load = sh_valid & ~armed & ~shift_en & ~flush;

    // Idle inside an open transaction with nothing to send. Note it is NOT an
    // error: it is the exact measure of how late the producer was.
    assign stalled = core_busy & ~shift_en & ~armed & ~sh_valid & ~tail;

    reg [MAX_W-1:0] rx_hold;
    reg             rx_full;

    assign rx_rdata = rx_hold;
    assign rx_ready = rx_full;

    always @(posedge clk or negedge rst_n) begin
        if (!rst_n) begin
            sh_data    <= {MAX_W{1'b0}};
            sh_last    <= 1'b0;
            sh_valid   <= 1'b0;
            armed      <= 1'b0;
            armed_last <= 1'b0;
            load_r     <= 1'b0;
            tail       <= 1'b0;
            tx_data    <= {MAX_W{1'b0}};
            rx_hold    <= {MAX_W{1'b0}};
            rx_full    <= 1'b0;
            rx_overrun <= 1'b0;
        end else begin
            load_r <= 1'b0;

            // --- the CPU writing into the shadow ---
            // Accepted whenever the shadow is free, which includes the same
            // cycle it is being emptied into the datapath below: the two are
            // ordered so a push and a load on one cycle both succeed.
            // FLUSH takes priority over everything on the transmit side. A
            // word queued for a transaction that is being abandoned must not
            // become the first word of the next one, and an `armed` request
            // left standing would start that next transaction by itself the
            // moment the bus came free. Chapter 13.10 is where this is driven
            // from; here it is simply the one input that can throw work away.
            if (flush) begin
                sh_valid   <= 1'b0;
                armed      <= 1'b0;
                armed_last <= 1'b0;
                tail       <= 1'b0;
            end else if (tx_push && tx_ready) begin
                sh_data  <= tx_wdata;
                sh_last  <= tx_last;
                sh_valid <= 1'b1;
            end

            // --- the shadow moving into the datapath ---
            if (can_load && !flush) begin
                tx_data    <= sh_data;
                armed_last <= sh_last;
                armed      <= 1'b1;
                load_r     <= 1'b1;
                tail       <= 1'b0;
                // Freed immediately, so the CPU may refill it on the very
                // next cycle rather than waiting for the frame to finish.
                // This single line is what makes the stream back-to-back.
                if (!(tx_push && tx_ready))
                    sh_valid <= 1'b0;
            end

            // --- the engine taking the frame ---
            if (start_stb) begin
                armed <= 1'b0;
                if (armed_last)
                    tail <= 1'b1;
            end

            // --- the receive side ---
            // The clear is applied FIRST so that a clear arriving on the same
            // cycle as a fresh overrun does not swallow it: the new event wins,
            // because losing a report of lost data is the same bug twice.
            if (clr_flags)
                rx_overrun <= 1'b0;

            if (rx_valid_stb) begin
                rx_hold <= rx_word;
                rx_full <= 1'b1;
                // A word arriving on top of an unread one is DATA LOST, and
                // the flag is sticky so a driver that only checks at the end
                // of the transaction still finds out.
                if (rx_full && !rx_pop)
                    rx_overrun <= 1'b1;
            end else if (rx_pop) begin
                rx_full <= 1'b0;
            end
        end
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_multibyte_tb.v — the same throughput measurements in Verilog-2001
// spi_multibyte_tb.v
//
// This is the first testbench that runs the whole engine: the divider of 13.4,
// the mode logic of 13.5, the datapath of 13.6 wrapped by 13.8, the chip-select
// controller of 13.7, and the streaming shadow of 13.9. Only the register
// interface is missing, and that is Chapter 13.11.
//
// The point of the chapter is throughput, so the measurements are about TIME as
// well as data:
//
//   - a prompt producer must stream N bytes under ONE chip select with no
//     stall cycles at all, and with the inter-frame idle equal to the gap the
//     controller was told to insert -- not a cycle more;
//   - a deliberately late producer must still be CORRECT, and must be
//     reported: the stall counter is how the lost throughput becomes visible;
//   - a receiver that does not read fast enough must set a STICKY overrun,
//     because a dropped word is lost data, not a slow day.
//
// Data is checked in both directions against an independent slave, byte by
// byte and in order, so a stream that gets every byte right but in the wrong
// sequence fails.

`timescale 1ns/1ps

module spi_multibyte_tb;

    localparam MAX_W = 32;
    localparam LEN_W = 6;
    localparam DIV_W = 8;
    localparam N_CS  = 2;
    localparam SEL_W = 1;
    localparam CNT_W = 8;

    reg clk;
    reg rst_n;
    always #5 clk = ~clk;

    // --- 13.9, the block under test ---------------------------------------
    reg             tx_push;
    reg [MAX_W-1:0] tx_wdata;
    reg             tx_last;
    wire              tx_ready;
    reg             flush;
    reg             clr_flags;
    reg             rx_pop;
    wire [MAX_W-1:0]  rx_rdata;
    wire              rx_ready, rx_overrun;

    wire              req, hold, load_stb, stalled;
    wire [MAX_W-1:0]  tx_data;

    // --- 13.7 -------------------------------------------------------------
    reg [SEL_W-1:0] sel;
    reg [CNT_W-1:0] lead_cyc;
    reg [CNT_W-1:0] lag_cyc;
    reg [CNT_W-1:0] gap_cyc;
    wire  [N_CS-1:0]  cs_n;
    wire              shift_en, start_stb, busy, sel_err;
    wire [2:0]        state_id;

    // --- 13.4 -------------------------------------------------------------
    reg [DIV_W-1:0] div;
    reg             cpol;
    wire              sclk, edge_a_stb, edge_b_stb, bit_done, div_err;
    wire [DIV_W-1:0]  half_a, half_b;

    // --- 13.5 / 13.8 ------------------------------------------------------
    reg             cpha;
    reg [LEN_W-1:0] len;
    reg             lsb_first;
    wire              preload_stb, launch_stb, capture_stb, frame_done;
    wire [LEN_W-1:0]  bit_idx;
    wire              miso, mosi;
    wire [MAX_W-1:0]  rx_word;
    wire              rx_valid_stb, len_err;

    spi_clkdiv_strobe #(.DIV_W(DIV_W)) u_div (
        .clk(clk), .rst_n(rst_n), .en(shift_en), .div(div), .cpol(cpol),
        .sclk(sclk), .edge_a_stb(edge_a_stb), .edge_b_stb(edge_b_stb),
        .bit_done(bit_done), .half_a(half_a), .half_b(half_b),
        .div_err(div_err)
    );

    spi_mode_edges #(.LEN_W(LEN_W)) u_mode (
        .clk(clk), .rst_n(rst_n), .cpha(cpha), .len(len),
        .active(shift_en), .start_stb(start_stb),
        .edge_a_stb(edge_a_stb), .edge_b_stb(edge_b_stb),
        .preload_stb(preload_stb), .launch_stb(launch_stb),
        .capture_stb(capture_stb), .bit_idx(bit_idx), .frame_done(frame_done)
    );

    spi_width_order #(.MAX_W(MAX_W), .LEN_W(LEN_W)) u_data (
        .clk(clk), .rst_n(rst_n),
        .tx_data(tx_data), .len(len), .lsb_first(lsb_first),
        .load_stb(load_stb),
        .preload_stb(preload_stb), .launch_stb(launch_stb),
        .capture_stb(capture_stb),
        .miso(miso), .mosi(mosi),
        .rx_data(rx_word), .rx_valid_stb(rx_valid_stb), .len_err(len_err)
    );

    spi_cs_ctrl #(.N_CS(N_CS), .SEL_W(SEL_W), .CNT_W(CNT_W)) u_cs (
        .clk(clk), .rst_n(rst_n),
        .req(req), .hold(hold), .sel(sel),
        .lead_cyc(lead_cyc), .lag_cyc(lag_cyc), .gap_cyc(gap_cyc),
        .core_done(frame_done), .abort(1'b0),
        .cs_n(cs_n), .shift_en(shift_en), .start_stb(start_stb),
        .busy(busy), .sel_err(sel_err), .state_id(state_id)
    );

    spi_multibyte #(.MAX_W(MAX_W), .LEN_W(LEN_W)) dut (
        .clk(clk), .rst_n(rst_n),
        .tx_push(tx_push), .tx_wdata(tx_wdata), .tx_last(tx_last),
        .tx_ready(tx_ready),
        .flush(flush), .clr_flags(clr_flags),
        .rx_pop(rx_pop), .rx_rdata(rx_rdata), .rx_ready(rx_ready),
        .rx_overrun(rx_overrun),
        .core_busy(busy), .shift_en(shift_en), .start_stb(start_stb),
        .rx_valid_stb(rx_valid_stb), .rx_word(rx_word),
        .req(req), .hold(hold), .tx_data(tx_data), .load_stb(load_stb),
        .stalled(stalled)
    );

    // --- the independent slave, word by word ------------------------------
    reg [MAX_W-1:0] slv_tx_q [0:127];
    reg [MAX_W-1:0] slv_rx_q [0:127];
    integer           slv_tx_i, slv_rx_i;
    reg [MAX_W-1:0] slv_sr, slv_rx_sr;
    reg             slave_bit_r;

    assign miso = slave_bit_r;

    always @(posedge clk) begin
        if (load_stb) begin
            slv_sr   <= slv_tx_q[slv_tx_i] << (MAX_W - len);
            slv_tx_i <= slv_tx_i + 1;
        end else if (preload_stb || launch_stb) begin
            slave_bit_r <= slv_sr[MAX_W-1];
            slv_sr      <= {slv_sr[MAX_W-2:0], 1'b0};
        end

        // Cleared at each load, so the recorded word holds THIS frame's bits
        // and not a running history of every frame before it.
        if (load_stb)
            slv_rx_sr <= {MAX_W{1'b0}};
        else if (capture_stb)
            slv_rx_sr <= {slv_rx_sr[MAX_W-2:0], mosi};

        // `rx_valid_stb` is one cycle after the final capture, so the slave's
        // own shift register is already complete when it is recorded.
        if (rx_valid_stb) begin
            slv_rx_q[slv_rx_i] <= slv_rx_sr;
            slv_rx_i           <= slv_rx_i + 1;
        end
    end

    // --- the throughput monitor -------------------------------------------
    integer stall_cycles;
    integer idle_in_txn;      // cycles inside a transaction with no clock
    integer frames_started;
    integer cs_falls;
    integer slack_cycles;     // room in the shadow WHILE the wire is busy
    integer valid_pulses;
    integer n_loads, n_caps, n_pre, n_lau, n_race;
    reg   busy_q;

    always @(posedge clk) begin
        if (!rst_n) begin
            busy_q <= 1'b0;
        end else begin
            busy_q <= busy;
            if (stalled)               stall_cycles   <= stall_cycles + 1;
            if (busy && !shift_en)     idle_in_txn    <= idle_in_txn + 1;
            if (start_stb)             frames_started <= frames_started + 1;
            if (busy && !busy_q)       cs_falls       <= cs_falls + 1;
            if (tx_ready && shift_en)  slack_cycles   <= slack_cycles + 1;
            if (rx_valid_stb)          valid_pulses   <= valid_pulses + 1;
            if (load_stb)              n_loads        <= n_loads + 1;
            if (capture_stb)           n_caps         <= n_caps + 1;
            if (preload_stb)           n_pre          <= n_pre + 1;
            if (launch_stb)            n_lau          <= n_lau + 1;
            // The datapath is loaded by one strobe and told to drive its first
            // bit by another. If they ever land on the same cycle the load is
            // overwritten by the shift in the same always block and the frame
            // sends the PREVIOUS word. The handover is built so they cannot,
            // and this counts the cycles on which that claim would be false.
            if (load_stb && start_stb) n_race         <= n_race + 1;
        end
    end

    integer errors;

    task clear_counts;
        begin
            stall_cycles = 0; idle_in_txn = 0; frames_started = 0;
            cs_falls = 0; slack_cycles = 0; valid_pulses = 0;
            n_loads = 0; n_caps = 0; n_pre = 0; n_lau = 0; n_race = 0;
        end
    endtask

    // The CPU writing a word. `late` cycles of dithering before the write is
    // how a slow driver is modelled.
        task push;
        input [MAX_W-1:0] w;
        input last;
        input integer late;
        integer guard;
        begin
            repeat (late) @(negedge clk);
            guard = 8000;
            while (!tx_ready && guard > 0) begin
                @(negedge clk);
                guard = guard - 1;
            end
            if (guard == 0) begin
                $display("  FAIL: the shadow never became free");
                errors = errors + 1;
            end
            tx_wdata = w;
            tx_last  = last;
            tx_push  = 1'b1;
            @(negedge clk);
            tx_push = 1'b0;
        end
    endtask

    // The CPU reading. `drain` = 0 leaves words unread on purpose.
    integer got_n;
    reg [MAX_W-1:0] got_q [0:127];

    task drain_rx;
        begin
            while (rx_ready) begin
                got_q[got_n] = rx_rdata;
                got_n        = got_n + 1;
                rx_pop       = 1'b1;
                @(negedge clk);
                rx_pop       = 1'b0;
                @(negedge clk);
            end
        end
    endtask

    task wait_idle;
        integer guard;
        begin
            guard = 20000;
            while (busy && guard > 0) begin
                drain_rx();
                @(negedge clk);
                guard = guard - 1;
            end
            repeat (4) @(negedge clk);
            drain_rx();
        end
    endtask

    reg [MAX_W-1:0] sent_q [0:127];
    integer i, k, nb, late, seed, bad;

    initial begin
        clear_counts();
        slv_tx_i = 0; slv_rx_i = 0; got_n = 0;
        slv_sr = 32'h0; slv_rx_sr = 32'h0; slave_bit_r = 1'b0;
        seed = 32'h5EED_1234;
        for (i = 0; i < 128; i = i + 1) begin
            seed        = (seed * 32'h0019_660D) + 32'h3C6E_F35F;
            slv_tx_q[i] = (seed >> 7) & 32'hFF;
        end

        repeat (3) @(negedge clk);
        rst_n = 1'b1;
        @(negedge clk);

        // 1. A PROMPT PRODUCER, sixteen bytes, one chip select. This is the
        //    case the shadow register exists for.
        clear_counts();
        nb = 16;
        for (i = 0; i < nb; i = i + 1) begin
            seed      = (seed * 32'h0019_660D) + 32'h3C6E_F35F;
            sent_q[i] = (seed >> 11) & 32'hFF;
        end
        got_n = 0;
        for (i = 0; i < nb; i = i + 1) begin
            push(sent_q[i], (i == nb - 1), 0);
            drain_rx();
        end
        wait_idle();

        if (cs_falls != 1) begin
            $display("  FAIL: a %0d-byte stream took %0d chip selects",
                     nb, cs_falls);
            errors = errors + 1;
        end
        if (frames_started != nb) begin
            $display("  FAIL: %0d bytes produced %0d frames", nb,
                     frames_started);
            errors = errors + 1;
        end
        if (stall_cycles != 0) begin
            $display("  FAIL: a prompt producer still stalled the wire for %0d cycles",
                     stall_cycles);
            errors = errors + 1;
        end
        // Everything the slave received, in order.
        bad = 0;
        for (i = 0; i < nb; i = i + 1)
            if (slv_rx_q[i] !== sent_q[i]) bad = bad + 1;
        if (bad != 0) begin
            $display("  FAIL: the slave received %0d of %0d bytes wrongly",
                     bad, nb);
            errors = errors + 1;
        end
        // Everything the master received, in order.
        bad = 0;
        if (got_n != nb) begin
            $display("  FAIL: the master read back %0d of %0d bytes",
                     got_n, nb);
            errors = errors + 1;
        end else begin
            for (i = 0; i < nb; i = i + 1)
                if (got_q[i] !== slv_tx_q[i]) bad = bad + 1;
            if (bad != 0) begin
                $display("  FAIL: the master received %0d of %0d bytes wrongly",
                         bad, nb);
                errors = errors + 1;
            end
        end
        if (rx_overrun) begin
            $display("  FAIL: a fully drained receiver reported an overrun");
            errors = errors + 1;
        end
        // The per-frame bookkeeping, counted over the whole stream rather
        // than trusted: one load and one preload per frame, len captures per
        // frame, and len-1 launches because CPHA=0's preload is the first
        // drive (Chapter 13.5).
        if (n_loads != nb || n_pre != nb) begin
            $display("  FAIL: %0d frames produced %0d loads and %0d preloads",
                     nb, n_loads, n_pre);
            errors = errors + 1;
        end
        if (n_caps != nb * 8 || n_lau != nb * 7) begin
            $display("  FAIL: %0d eight-bit frames produced %0d captures and %0d launches, expected %0d and %0d",
                     nb, n_caps, n_lau, nb * 8, nb * 7);
            errors = errors + 1;
        end
        if (n_race != 0) begin
            $display("  FAIL: a load landed on a frame start %0d times",
                     n_race);
            errors = errors + 1;
        end
        if (valid_pulses != nb) begin
            $display("  FAIL: %0d frames raised %0d received-word pulses",
                     nb, valid_pulses);
            errors = errors + 1;
        end
        $display("  per-frame bookkeeping over the stream: %0d loads, %0d preloads, %0d launches, %0d captures, %0d received words, and no load ever landed on a frame start",
                 n_loads, n_pre, n_lau, n_caps, valid_pulses);
        $display("  16 bytes, one chip select, %0d frames, %0d stall cycles, %0d idle cycles inside the transaction -- %0d per byte boundary",
                 frames_started, stall_cycles, idle_in_txn,
                 idle_in_txn / (nb - 1));

        // 2. THE IDLE BETWEEN FRAMES IS THE GAP THAT WAS ASKED FOR, and not a
        //    cycle more. This is the check that would catch a shadow register
        //    that works but hands over one cycle late.
        if (idle_in_txn > (nb - 1) * (gap_cyc + 3)) begin
            $display("  FAIL: %0d idle cycles across %0d boundaries, more than the %0d-cycle gap explains",
                     idle_in_txn, nb - 1, gap_cyc);
            errors = errors + 1;
        end
        $display("  the inter-frame idle is the programmed %0d-cycle hold and nothing else",
                 gap_cyc);

        // 3. A LATE PRODUCER. Still correct, and now visibly costly.
        clear_counts();
        nb = 8;
        for (i = 0; i < nb; i = i + 1) begin
            seed      = (seed * 32'h0019_660D) + 32'h3C6E_F35F;
            sent_q[i] = (seed >> 11) & 32'hFF;
        end
        k     = slv_tx_i;
        got_n = 0;
        for (i = 0; i < nb; i = i + 1) begin
            push(sent_q[i], (i == nb - 1), 40);
            drain_rx();
        end
        wait_idle();

        if (stall_cycles == 0) begin
            $display("  FAIL: a producer 40 cycles late stalled the wire not at all -- the counter cannot be working");
            errors = errors + 1;
        end
        bad = 0;
        for (i = 0; i < nb; i = i + 1)
            if (slv_rx_q[k + i] !== sent_q[i]) bad = bad + 1;
        if (bad != 0) begin
            $display("  FAIL: a late producer corrupted %0d of %0d bytes",
                     bad, nb);
            errors = errors + 1;
        end
        if (cs_falls != 1) begin
            $display("  FAIL: a late producer broke the transaction into %0d chip selects",
                     cs_falls);
            errors = errors + 1;
        end
        $display("  the same stream with the producer 40 cycles late: every byte still correct, one chip select, but %0d stall cycles the fast run did not have",
                 stall_cycles);

        // 4. AN OVERRUN IS STICKY. Stream four bytes and never read one.
        clear_counts();
        nb = 4;
        for (i = 0; i < nb; i = i + 1)
            push(32'h5A, (i == nb - 1), 0);
        // Deliberately no drain_rx here.
        i = 20000;
        while (busy && i > 0) begin
            @(negedge clk);
            i = i - 1;
        end
        repeat (4) @(negedge clk);
        if (!rx_overrun) begin
            $display("  FAIL: four unread received words did not raise an overrun");
            errors = errors + 1;
        end
        $display("  four received words left unread: overrun raised and latched");

        // It must STAY raised through a clean transfer -- a flag that clears
        // itself is a flag that hides the one byte that was lost.
        drain_rx();
        clear_counts();
        for (i = 0; i < 2; i = i + 1) begin
            push(32'h33, (i == 1), 0);
            drain_rx();
        end
        wait_idle();
        if (!rx_overrun) begin
            $display("  FAIL: the overrun flag cleared itself on the next clean transfer");
            errors = errors + 1;
        end
        $display("  a clean transfer afterwards does not clear it -- the lost byte stays reported");

        // 5. THE CPU CAN REFILL WHILE THE WIRE IS BUSY. Without that the
        //    shadow buys nothing, so it is checked directly: at least one push
        //    must be accepted with the clock running.
        clear_counts();
        // Reset the sticky flag by resetting the block, since nothing else
        // clears it -- which is itself the design decision under test.
        rst_n = 1'b0;
        repeat (3) @(negedge clk);
        rst_n = 1'b1;
        repeat (2) @(negedge clk);
        slv_tx_i = 0; slv_rx_i = 0; got_n = 0;

        // What matters is not the instant the testbench happens to write, but
        // HOW LONG the shadow stays free while the wire is busy -- that window
        // is the slack the producer is given, and without the shadow it would
        // be zero.
        for (i = 0; i < 8; i = i + 1) begin
            push(32'hC5, (i == 7), 0);
            drain_rx();
        end
        wait_idle();
        if (slack_cycles == 0) begin
            $display("  FAIL: the shadow was never free while the clock was running -- it is not decoupling anything");
            errors = errors + 1;
        end
        if (stall_cycles != 0) begin
            $display("  FAIL: %0d stall cycles while refilling ahead",
                     stall_cycles);
            errors = errors + 1;
        end
        $display("  the shadow stood free for %0d of the cycles the wire was busy -- that window is the slack the producer gets, and it is a whole frame wide",
                 slack_cycles);

        if (errors == 0)
            $display("PASS: a one-deep shadow register decouples the producer from the wire -- sixteen bytes stream under a single chip select with every byte correct in both directions and in order, zero stall cycles, and an inter-frame idle equal to the programmed hold and nothing more -- words are accepted while the previous frame is still shifting, which is the whole reason the shadow exists -- a producer forty cycles late still transfers every byte correctly under one chip select but now reports the stall cycles it cost, so the lost throughput is visible instead of silent -- and a received word overwritten before it was read raises a sticky overrun that a later clean transfer does not clear");
        else
            $display("FAIL: %0d error(s)", errors);
        $finish;
    end


    initial begin
        clk = 1'b0;
        rst_n = 1'b0;
        tx_push = 1'b0;
        tx_wdata = 32'h0;
        tx_last = 1'b0;
        flush = 1'b0;
        clr_flags = 1'b0;
        rx_pop = 1'b0;
        sel = 1'b0;
        lead_cyc = 8'd3;
        lag_cyc = 8'd3;
        gap_cyc = 8'd2;
        div = 8'd4;
        cpol = 1'b0;
        cpha = 1'b0;
        len = 6'd8;
        lsb_first = 1'b0;
        errors = 0;
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_multibyte.vhd — the same controller in VHDL
-- spi_multibyte.vhd
--
-- Chapter 13.9 -- keeping the wire busy between frames.
--
-- Everything built so far transfers ONE frame correctly. A real transaction is
-- a stream: a flash page program is a command byte, three address bytes and
-- then 256 data bytes, all under one chip select. The gap between those bytes
-- is where the throughput of Chapter 9 is won or lost.
--
-- THE PROBLEM WITH A SINGLE TRANSMIT REGISTER. If the datapath's shift
-- register is the only place a word can live, the next byte cannot be written
-- until the current one has left -- and "has left" means the frame is over:
--
--     shift byte N ... frame ends ... CPU notices ... CPU writes byte N+1
--     ... controller starts ... shift byte N+1
--
-- Every one of those dots is dead time on SCLK, and none of it is under the
-- hardware's control. The transfer is correct and the wire is half idle.
--
-- THE FIX IS ONE REGISTER DEEP. A SHADOW register in front of the datapath:
-- the CPU writes the shadow, the controller moves the shadow into the datapath
-- between frames. The CPU now has an entire FRAME to produce the next word
-- instead of an inter-frame gap.
--
-- One deep is enough, and that is worth being clear about: a deeper FIFO buys
-- latency TOLERANCE, not throughput. Once the producer can stay one word
-- ahead, the wire is saturated and nothing deeper makes it faster.
--
-- WHAT THIS BLOCK REPORTS. A design that silently absorbs a slow producer is
-- worse than one that says so, because the lost throughput is invisible:
--
--   stalled     the transaction is open, the clock is not running, and nothing
--               is queued -- the wire is idle because the CPU was late.
--   rx_overrun  a received word was overwritten before it was read. Unlike a
--               stall this is LOST DATA, and the flag is sticky: a receive
--               path that drops a word and carries on is how a flash image
--               ends up with one wrong word in it.

library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

entity spi_multibyte is
    generic (
        MAX_W : positive := 32;
        LEN_W : positive := 6
    );
    port (
        clk          : in  std_logic;
        rst_n        : in  std_logic;

        -- CPU side
        tx_push      : in  std_logic;   -- write a word into the shadow
        tx_wdata     : in  std_logic_vector(MAX_W - 1 downto 0);
        tx_last      : in  std_logic;   -- .. and end the transaction after it
        tx_ready     : out std_logic;   -- the shadow is free

        flush        : in  std_logic;   -- drop whatever is queued
        rx_pop       : in  std_logic;   -- read the received word
        rx_rdata     : out std_logic_vector(MAX_W - 1 downto 0);
        rx_ready     : out std_logic;   -- a received word is waiting
        rx_overrun   : out std_logic;   -- sticky: a word was lost
        -- Clearing a sticky flag is the REGISTER INTERFACE's business, not this
        -- block's -- but the flag lives here, because this is where the loss is
        -- detected. So each flag has exactly one owner and one clear path, and
        -- Chapter 13.11 wires this to a write of the command register.
        clr_flags    : in  std_logic;

        -- transfer engine side (13.7 and 13.6/13.8)
        core_busy    : in  std_logic;   -- a transaction is open
        shift_en     : in  std_logic;   -- the clock is running
        start_stb    : in  std_logic;   -- the engine took a frame
        rx_valid_stb : in  std_logic;
        rx_word      : in  std_logic_vector(MAX_W - 1 downto 0);

        req          : out std_logic;   -- to the chip-select controller
        hold         : out std_logic;
        tx_data      : out std_logic_vector(MAX_W - 1 downto 0);
        load_stb     : out std_logic;

        stalled      : out std_logic    -- the wire is idle for want of data
    );
end entity;

architecture rtl of spi_multibyte is

    -- The shadow: one word, its end-of-transaction flag, and a valid bit.
    signal sh_data  : std_logic_vector(MAX_W - 1 downto 0) := (others => '0');
    signal sh_last  : std_logic := '0';
    signal sh_valid : std_logic := '0';

    -- `armed` means the datapath has been loaded and the controller has been
    -- asked for a frame, but the frame has not started yet. It stops a second
    -- load from landing on top of the first.
    signal armed      : std_logic := '0';
    signal armed_last : std_logic := '0';
    signal load_r     : std_logic := '0';

    -- `tail` means the final frame of the transaction has started, so the
    -- controller is walking out through its lag and gap with nothing more to
    -- send. Without it the closing states look exactly like a producer that
    -- went quiet, and `stalled` would accuse the CPU of being late at the end
    -- of every single transaction.
    signal tail : std_logic := '0';

    signal tx_data_r : std_logic_vector(MAX_W - 1 downto 0) := (others => '0');
    signal rx_hold   : std_logic_vector(MAX_W - 1 downto 0) := (others => '0');
    signal rx_full   : std_logic := '0';
    signal ovr_r     : std_logic := '0';

    signal can_load : std_logic;
    signal ready_i  : std_logic;

begin

    ready_i  <= not sh_valid;
    tx_ready <= ready_i;
    req      <= armed;
    hold     <= not armed_last;
    load_stb <= load_r;
    tx_data  <= tx_data_r;

    rx_rdata   <= rx_hold;
    rx_ready   <= rx_full;
    rx_overrun <= ovr_r;

    -- The datapath may only be loaded while the clock is stopped. Loading it
    -- mid-frame would overwrite the word being shifted -- the one failure a
    -- single-register design cannot even express, and that a shadow design has
    -- to be careful about.
    can_load <= sh_valid and (not armed) and (not shift_en) and (not flush);

    -- Idle inside an open transaction with nothing to send. Note it is NOT an
    -- error: it is the exact measure of how late the producer was.
    stalled <= core_busy and (not shift_en) and (not armed) and
               (not sh_valid) and (not tail);

    seq : process (clk, rst_n)
    begin
        if rst_n = '0' then
            sh_data    <= (others => '0');
            sh_last    <= '0';
            sh_valid   <= '0';
            armed      <= '0';
            armed_last <= '0';
            load_r     <= '0';
            tail       <= '0';
            tx_data_r  <= (others => '0');
            rx_hold    <= (others => '0');
            rx_full    <= '0';
            ovr_r      <= '0';
        elsif rising_edge(clk) then
            load_r <= '0';

            -- The CPU writing into the shadow. Accepted whenever the shadow is
            -- free, which includes the same cycle it is being emptied into the
            -- datapath below: the two are ordered so a push and a load on one
            -- cycle both succeed.
            -- FLUSH takes priority over everything on the transmit side. A
            -- word queued for a transaction that is being abandoned must not
            -- become the first word of the next one, and an `armed` request
            -- left standing would start that next transaction by itself the
            -- moment the bus came free. Chapter 13.10 drives this; here it is
            -- simply the one input that can throw work away.
            if flush = '1' then
                sh_valid   <= '0';
                armed      <= '0';
                armed_last <= '0';
                tail       <= '0';
            elsif tx_push = '1' and ready_i = '1' then
                sh_data  <= tx_wdata;
                sh_last  <= tx_last;
                sh_valid <= '1';
            end if;

            -- The shadow moving into the datapath.
            if can_load = '1' and flush = '0' then
                tx_data_r  <= sh_data;
                armed_last <= sh_last;
                armed      <= '1';
                load_r     <= '1';
                tail       <= '0';
                -- Freed immediately, so the CPU may refill it on the very next
                -- cycle rather than waiting for the frame to finish. This
                -- single line is what makes the stream back-to-back.
                if not (tx_push = '1' and ready_i = '1') then
                    sh_valid <= '0';
                end if;
            end if;

            -- The engine taking the frame.
            if start_stb = '1' then
                armed <= '0';
                if armed_last = '1' then
                    tail <= '1';
                end if;
            end if;

            -- The receive side. The clear is applied FIRST so that a clear
            -- arriving on the same cycle as a fresh overrun does not swallow
            -- it: the new event wins, because losing a report of lost data is
            -- the same bug twice.
            if clr_flags = '1' then
                ovr_r <= '0';
            end if;

            if rx_valid_stb = '1' then
                rx_hold <= rx_word;
                rx_full <= '1';
                -- A word arriving on top of an unread one is DATA LOST, and
                -- the flag is sticky so a driver that only checks at the end
                -- of the transaction still finds out.
                if rx_full = '1' and rx_pop = '0' then
                    ovr_r <= '1';
                end if;
            elsif rx_pop = '1' then
                rx_full <= '0';
            end if;
        end if;
    end process;

end architecture;
Azvya Education Pvt. Ltd.VLSI Mentor
spi_multibyte_tb.vhd — the same throughput measurements in VHDL
-- spi_multibyte_tb.vhd
--
-- This is the first testbench that runs the whole engine: the divider of 13.4,
-- the mode logic of 13.5, the datapath of 13.6 wrapped by 13.8, the chip-select
-- controller of 13.7, and the streaming shadow of 13.9. Only the register
-- interface is missing, and that is Chapter 13.11.
--
-- The point of the chapter is throughput, so the measurements are about TIME as
-- well as data:
--
--   - a prompt producer must stream N words under ONE chip select with no stall
--     cycles at all, and with the inter-frame idle equal to the gap the
--     controller was told to insert -- not a cycle more;
--   - a deliberately late producer must still be CORRECT, and must be reported;
--   - a receiver that does not read fast enough must set a STICKY overrun.
--
-- Data is checked in both directions against an independent slave, word by word
-- and in order, so a stream that gets every word right but in the wrong
-- sequence fails.

library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

entity spi_multibyte_tb is
end entity;

architecture sim of spi_multibyte_tb is

    constant MAX_W : positive := 32;
    constant LEN_W : positive := 6;
    constant DIV_W : positive := 8;
    constant N_CS  : positive := 2;
    constant SEL_W : positive := 1;
    constant CNT_W : positive := 8;

    type word_array is array (0 to 127) of std_logic_vector(MAX_W - 1 downto 0);

    signal clk   : std_logic := '0';
    signal rst_n : std_logic := '0';
    signal halt  : boolean   := false;

    -- 13.9, the block under test
    signal tx_push  : std_logic := '0';
    signal tx_wdata : std_logic_vector(MAX_W - 1 downto 0) := (others => '0');
    signal tx_last  : std_logic := '0';
    signal tx_ready : std_logic;
    signal flush     : std_logic := '0';
    signal clr_flags : std_logic := '0';
    signal rx_pop   : std_logic := '0';
    signal rx_rdata : std_logic_vector(MAX_W - 1 downto 0);
    signal rx_ready, rx_overrun : std_logic;
    signal req, hold, load_stb, stalled : std_logic;
    signal tx_data  : std_logic_vector(MAX_W - 1 downto 0);

    -- 13.7
    signal sel      : unsigned(SEL_W - 1 downto 0) := (others => '0');
    signal lead_cyc : unsigned(CNT_W - 1 downto 0) := to_unsigned(3, CNT_W);
    signal lag_cyc  : unsigned(CNT_W - 1 downto 0) := to_unsigned(3, CNT_W);
    signal gap_cyc  : unsigned(CNT_W - 1 downto 0) := to_unsigned(2, CNT_W);
    signal cs_n     : std_logic_vector(N_CS - 1 downto 0);
    signal shift_en, start_stb, busy, sel_err : std_logic;
    signal state_id : unsigned(2 downto 0);

    -- 13.4
    signal div  : unsigned(DIV_W - 1 downto 0) := to_unsigned(4, DIV_W);
    signal cpol : std_logic := '0';
    signal sclk, edge_a_stb, edge_b_stb, bit_done, div_err : std_logic;
    signal half_a, half_b : unsigned(DIV_W - 1 downto 0);

    -- 13.5 / 13.8
    signal cpha      : std_logic := '0';
    signal len       : unsigned(LEN_W - 1 downto 0) := to_unsigned(8, LEN_W);
    signal lsb_first : std_logic := '0';
    signal preload_stb, launch_stb, capture_stb, frame_done : std_logic;
    signal bit_idx   : unsigned(LEN_W - 1 downto 0);
    signal miso, mosi : std_logic;
    signal rx_word   : std_logic_vector(MAX_W - 1 downto 0);
    signal rx_valid_stb, len_err : std_logic;

    -- the independent slave, word by word
    signal slv_tx_q : word_array := (others => (others => '0'));
    signal slv_rx_q : word_array := (others => (others => '0'));
    signal slv_tx_i : natural := 0;
    signal slv_rx_i : natural := 0;
    signal slv_sr, slv_rx_sr : std_logic_vector(MAX_W - 1 downto 0)
                               := (others => '0');
    signal slave_bit_r : std_logic := '0';
    signal slv_rst     : std_logic := '0';

    -- the throughput monitor
    signal stall_cycles, idle_in_txn, frames_started, cs_falls : natural := 0;
    signal slack_cycles, valid_pulses : natural := 0;
    signal n_loads, n_caps, n_pre, n_lau, n_race : natural := 0;
    signal busy_q    : std_logic := '0';
    signal clear_stb : std_logic := '0';

    signal errors : natural := 0;

    function hex8(v : std_logic_vector) return string is
        constant DIGITS : string(1 to 16) := "0123456789abcdef";
        variable u : unsigned(MAX_W - 1 downto 0);
        variable r : string(1 to 8);
    begin
        u := unsigned(v);
        for k in 8 downto 1 loop
            r(k) := DIGITS(to_integer(u(3 downto 0)) + 1);
            u := shift_right(u, 4);
        end loop;
        return r;
    end function;

begin

    clk <= not clk after 5 ns when not halt else '0';

    u_div : entity work.spi_clkdiv_strobe
        generic map (DIV_W => DIV_W)
        port map (clk => clk, rst_n => rst_n, en => shift_en, div => div,
                  cpol => cpol, sclk => sclk,
                  edge_a_stb => edge_a_stb, edge_b_stb => edge_b_stb,
                  bit_done => bit_done, half_a => half_a, half_b => half_b,
                  div_err => div_err);

    u_mode : entity work.spi_mode_edges
        generic map (LEN_W => LEN_W)
        port map (clk => clk, rst_n => rst_n, cpha => cpha, len => len,
                  active => shift_en, start_stb => start_stb,
                  edge_a_stb => edge_a_stb, edge_b_stb => edge_b_stb,
                  preload_stb => preload_stb, launch_stb => launch_stb,
                  capture_stb => capture_stb, bit_idx => bit_idx,
                  frame_done => frame_done);

    u_data : entity work.spi_width_order
        generic map (MAX_W => MAX_W, LEN_W => LEN_W)
        port map (clk => clk, rst_n => rst_n,
                  tx_data => tx_data, len => len, lsb_first => lsb_first,
                  load_stb => load_stb,
                  preload_stb => preload_stb, launch_stb => launch_stb,
                  capture_stb => capture_stb,
                  miso => miso, mosi => mosi,
                  rx_data => rx_word, rx_valid_stb => rx_valid_stb,
                  len_err => len_err);

    u_cs : entity work.spi_cs_ctrl
        generic map (N_CS => N_CS, SEL_W => SEL_W, CNT_W => CNT_W)
        port map (clk => clk, rst_n => rst_n,
                  req => req, hold => hold, sel => sel,
                  lead_cyc => lead_cyc, lag_cyc => lag_cyc, gap_cyc => gap_cyc,
                  core_done => frame_done, abort => '0',
                  cs_n => cs_n, shift_en => shift_en, start_stb => start_stb,
                  busy => busy, sel_err => sel_err, state_id => state_id);

    dut : entity work.spi_multibyte
        generic map (MAX_W => MAX_W, LEN_W => LEN_W)
        port map (clk => clk, rst_n => rst_n,
                  tx_push => tx_push, tx_wdata => tx_wdata, tx_last => tx_last,
                  tx_ready => tx_ready,
                  flush => flush, clr_flags => clr_flags,
                  rx_pop => rx_pop, rx_rdata => rx_rdata, rx_ready => rx_ready,
                  rx_overrun => rx_overrun,
                  core_busy => busy, shift_en => shift_en,
                  start_stb => start_stb,
                  rx_valid_stb => rx_valid_stb, rx_word => rx_word,
                  req => req, hold => hold, tx_data => tx_data,
                  load_stb => load_stb,
                  stalled => stalled);

    miso <= slave_bit_r;

    slave_p : process (clk)
    begin
        if rising_edge(clk) then
            if slv_rst = '1' then
                slv_tx_i    <= 0;
                slv_rx_i    <= 0;
                slv_sr      <= (others => '0');
                slv_rx_sr   <= (others => '0');
                slave_bit_r <= '0';
            else
                if load_stb = '1' then
                    slv_sr <= std_logic_vector(
                                  shift_left(unsigned(slv_tx_q(slv_tx_i)),
                                             MAX_W - to_integer(len)));
                    slv_tx_i <= slv_tx_i + 1;
                elsif preload_stb = '1' or launch_stb = '1' then
                    slave_bit_r <= slv_sr(MAX_W - 1);
                    slv_sr      <= slv_sr(MAX_W - 2 downto 0) & '0';
                end if;

                -- Cleared at each load, so the recorded word holds THIS
                -- frame's bits and not a running history of every frame
                -- before it.
                if load_stb = '1' then
                    slv_rx_sr <= (others => '0');
                elsif capture_stb = '1' then
                    slv_rx_sr <= slv_rx_sr(MAX_W - 2 downto 0) & mosi;
                end if;

                -- `rx_valid_stb` is one cycle after the final capture, so the
                -- slave's own shift register is already complete here.
                if rx_valid_stb = '1' then
                    slv_rx_q(slv_rx_i) <= slv_rx_sr;
                    slv_rx_i           <= slv_rx_i + 1;
                end if;
            end if;
        end if;
    end process;

    monitor : process (clk)
    begin
        if rising_edge(clk) then
            if rst_n = '1' then
                busy_q <= busy;
                if clear_stb = '1' then
                    stall_cycles <= 0; idle_in_txn <= 0; frames_started <= 0;
                    cs_falls <= 0; slack_cycles <= 0; valid_pulses <= 0;
                    n_loads <= 0; n_caps <= 0; n_pre <= 0; n_lau <= 0;
                    n_race <= 0;
                else
                    if stalled = '1' then
                        stall_cycles <= stall_cycles + 1;
                    end if;
                    if busy = '1' and shift_en = '0' then
                        idle_in_txn <= idle_in_txn + 1;
                    end if;
                    if start_stb = '1' then
                        frames_started <= frames_started + 1;
                    end if;
                    if busy = '1' and busy_q = '0' then
                        cs_falls <= cs_falls + 1;
                    end if;
                    if tx_ready = '1' and shift_en = '1' then
                        slack_cycles <= slack_cycles + 1;
                    end if;
                    if rx_valid_stb = '1' then
                        valid_pulses <= valid_pulses + 1;
                    end if;
                    if load_stb = '1' then n_loads <= n_loads + 1; end if;
                    if capture_stb = '1' then n_caps <= n_caps + 1; end if;
                    if preload_stb = '1' then n_pre <= n_pre + 1; end if;
                    if launch_stb = '1' then n_lau <= n_lau + 1; end if;
                    -- The datapath is loaded by one strobe and told to drive
                    -- its first bit by another. If they ever land on the same
                    -- cycle the load is overwritten by the shift in the same
                    -- process and the frame sends the PREVIOUS word. The
                    -- handover is built so they cannot, and this counts the
                    -- cycles on which that claim would be false.
                    if load_stb = '1' and start_stb = '1' then
                        n_race <= n_race + 1;
                    end if;
                end if;
            end if;
        end if;
    end process;

    stim : process
        variable errs  : natural := 0;
        variable guard : natural;
        variable seed  : unsigned(31 downto 0) := x"5EED1234";
        variable sent_q : word_array;
        variable got_q  : word_array;
        variable got_n  : natural := 0;
        variable nb, k, bad : natural;
        variable accepted_free : natural;

        procedure clear_counts is
        begin
            clear_stb <= '1';
            wait until falling_edge(clk);
            clear_stb <= '0';
        end procedure;

        procedure next_rand(variable v : out unsigned(31 downto 0)) is
        begin
            seed := resize(seed * x"0019660D", 32) + x"3C6EF35F";
            v := seed;
        end procedure;

        -- The CPU reading. Leaving this uncalled is how an overrun is made.
        procedure drain_rx is
        begin
            while rx_ready = '1' loop
                got_q(got_n) := rx_rdata;
                got_n        := got_n + 1;
                rx_pop       <= '1';
                wait until falling_edge(clk);
                rx_pop       <= '0';
                wait until falling_edge(clk);
            end loop;
        end procedure;

        -- The CPU writing a word. `late` cycles of dithering before the write
        -- is how a slow driver is modelled.
        procedure push(w : std_logic_vector(MAX_W - 1 downto 0);
                       last : std_logic; late : natural) is
        begin
            for j in 1 to late loop wait until falling_edge(clk); end loop;
            guard := 8000;
            while tx_ready = '0' and guard > 0 loop
                wait until falling_edge(clk);
                guard := guard - 1;
            end loop;
            if guard = 0 then
                report "  FAIL: the shadow never became free";
                errs := errs + 1;
            end if;
            tx_wdata <= w;
            tx_last  <= last;
            tx_push  <= '1';
            wait until falling_edge(clk);
            tx_push  <= '0';
        end procedure;

        procedure wait_idle is
        begin
            guard := 20000;
            while busy = '1' and guard > 0 loop
                drain_rx;
                wait until falling_edge(clk);
                guard := guard - 1;
            end loop;
            for j in 1 to 4 loop wait until falling_edge(clk); end loop;
            drain_rx;
        end procedure;

        variable r : unsigned(31 downto 0);
    begin
        for i in 0 to 127 loop
            next_rand(r);
            slv_tx_q(i) <= std_logic_vector(resize(shift_right(r, 7) and x"000000FF",
                                                  MAX_W));
        end loop;

        for k2 in 1 to 3 loop wait until falling_edge(clk); end loop;
        rst_n <= '1';
        wait until falling_edge(clk);

        -- 1. A PROMPT PRODUCER, sixteen words, one chip select. This is the
        --    case the shadow register exists for.
        clear_counts;
        nb := 16;
        for i in 0 to nb - 1 loop
            next_rand(r);
            sent_q(i) := std_logic_vector(resize(shift_right(r, 11) and x"000000FF",
                                                 MAX_W));
        end loop;
        got_n := 0;
        for i in 0 to nb - 1 loop
            if i = nb - 1 then push(sent_q(i), '1', 0);
            else               push(sent_q(i), '0', 0); end if;
            drain_rx;
        end loop;
        wait_idle;

        if cs_falls /= 1 then
            report "  FAIL: a " & integer'image(nb) & "-word stream took " &
                   integer'image(cs_falls) & " chip selects";
            errs := errs + 1;
        end if;
        if frames_started /= nb then
            report "  FAIL: " & integer'image(nb) & " words produced " &
                   integer'image(frames_started) & " frames";
            errs := errs + 1;
        end if;
        if stall_cycles /= 0 then
            report "  FAIL: a prompt producer still stalled the wire for " &
                   integer'image(stall_cycles) & " cycles";
            errs := errs + 1;
        end if;
        bad := 0;
        for i in 0 to nb - 1 loop
            if slv_rx_q(i) /= sent_q(i) then bad := bad + 1; end if;
        end loop;
        if bad /= 0 then
            report "  FAIL: the slave received " & integer'image(bad) & " of " &
                   integer'image(nb) & " words wrongly";
            errs := errs + 1;
        end if;
        bad := 0;
        if got_n /= nb then
            report "  FAIL: the master read back " & integer'image(got_n) &
                   " of " & integer'image(nb) & " words";
            errs := errs + 1;
        else
            for i in 0 to nb - 1 loop
                if got_q(i) /= slv_tx_q(i) then bad := bad + 1; end if;
            end loop;
            if bad /= 0 then
                report "  FAIL: the master received " & integer'image(bad) &
                       " of " & integer'image(nb) & " words wrongly";
                errs := errs + 1;
            end if;
        end if;
        if rx_overrun = '1' then
            report "  FAIL: a fully drained receiver reported an overrun";
            errs := errs + 1;
        end if;

        -- The per-frame bookkeeping, counted over the whole stream rather than
        -- trusted: one load and one preload per frame, len captures per frame,
        -- and len-1 launches because CPHA=0's preload is the first drive.
        if n_loads /= nb or n_pre /= nb then
            report "  FAIL: " & integer'image(nb) & " frames produced " &
                   integer'image(n_loads) & " loads and " &
                   integer'image(n_pre) & " preloads";
            errs := errs + 1;
        end if;
        if n_caps /= nb * 8 or n_lau /= nb * 7 then
            report "  FAIL: " & integer'image(nb) &
                   " eight-bit frames produced " & integer'image(n_caps) &
                   " captures and " & integer'image(n_lau) & " launches";
            errs := errs + 1;
        end if;
        if n_race /= 0 then
            report "  FAIL: a load landed on a frame start";
            errs := errs + 1;
        end if;
        if valid_pulses /= nb then
            report "  FAIL: " & integer'image(nb) & " frames raised " &
                   integer'image(valid_pulses) & " received-word pulses";
            errs := errs + 1;
        end if;
        report "  per-frame bookkeeping over the stream: " &
               integer'image(n_loads) & " loads, " & integer'image(n_pre) &
               " preloads, " & integer'image(n_lau) & " launches, " &
               integer'image(n_caps) & " captures, " &
               integer'image(valid_pulses) &
               " received words, and no load ever landed on a frame start";
        report "  16 words, one chip select, " &
               integer'image(frames_started) & " frames, " &
               integer'image(stall_cycles) & " stall cycles, " &
               integer'image(idle_in_txn) &
               " idle cycles inside the transaction -- " &
               integer'image(idle_in_txn / (nb - 1)) & " per word boundary";

        -- 2. THE IDLE BETWEEN FRAMES IS THE GAP THAT WAS ASKED FOR, and not a
        --    cycle more. This would catch a shadow register that works but
        --    hands over one cycle late.
        if idle_in_txn > (nb - 1) * (to_integer(gap_cyc) + 3) then
            report "  FAIL: more idle cycles than the programmed gap explains";
            errs := errs + 1;
        end if;
        report "  the inter-frame idle is the programmed " &
               integer'image(to_integer(gap_cyc)) &
               "-cycle hold and nothing else";

        -- 3. A LATE PRODUCER. Still correct, and now visibly costly.
        clear_counts;
        nb := 8;
        for i in 0 to nb - 1 loop
            next_rand(r);
            sent_q(i) := std_logic_vector(resize(shift_right(r, 11) and x"000000FF",
                                                 MAX_W));
        end loop;
        k     := slv_rx_i;
        got_n := 0;
        for i in 0 to nb - 1 loop
            if i = nb - 1 then push(sent_q(i), '1', 40);
            else               push(sent_q(i), '0', 40); end if;
            drain_rx;
        end loop;
        wait_idle;

        if stall_cycles = 0 then
            report "  FAIL: a producer 40 cycles late stalled the wire not at all -- the counter cannot be working";
            errs := errs + 1;
        end if;
        bad := 0;
        for i in 0 to nb - 1 loop
            if slv_rx_q(k + i) /= sent_q(i) then bad := bad + 1; end if;
        end loop;
        if bad /= 0 then
            report "  FAIL: a late producer corrupted " & integer'image(bad) &
                   " of " & integer'image(nb) & " words";
            errs := errs + 1;
        end if;
        if cs_falls /= 1 then
            report "  FAIL: a late producer broke the transaction into " &
                   integer'image(cs_falls) & " chip selects";
            errs := errs + 1;
        end if;
        report "  the same stream with the producer 40 cycles late: every word still correct, one chip select, but " &
               integer'image(stall_cycles) &
               " stall cycles the fast run did not have";

        -- 4. AN OVERRUN IS STICKY. Stream four words and never read one.
        clear_counts;
        nb := 4;
        for i in 0 to nb - 1 loop
            if i = nb - 1 then push(x"0000005A", '1', 0);
            else               push(x"0000005A", '0', 0); end if;
        end loop;
        guard := 20000;
        while busy = '1' and guard > 0 loop
            wait until falling_edge(clk);
            guard := guard - 1;
        end loop;
        for j in 1 to 4 loop wait until falling_edge(clk); end loop;
        if rx_overrun /= '1' then
            report "  FAIL: four unread received words did not raise an overrun";
            errs := errs + 1;
        end if;
        report "  four received words left unread: overrun raised and latched";

        -- It must STAY raised through a clean transfer -- a flag that clears
        -- itself is a flag that hides the one word that was lost.
        drain_rx;
        clear_counts;
        for i in 0 to 1 loop
            if i = 1 then push(x"00000033", '1', 0);
            else          push(x"00000033", '0', 0); end if;
            drain_rx;
        end loop;
        wait_idle;
        if rx_overrun /= '1' then
            report "  FAIL: the overrun flag cleared itself on the next clean transfer";
            errs := errs + 1;
        end if;
        report "  a clean transfer afterwards does not clear it -- the lost word stays reported";

        -- 5. HOW LONG THE SHADOW STAYS FREE WHILE THE WIRE IS BUSY. That
        --    window is the slack the producer is given, and without the shadow
        --    it would be zero. Reset first, because nothing else clears the
        --    sticky overrun -- which is itself the design decision under test.
        rst_n   <= '0';
        slv_rst <= '1';
        for j in 1 to 3 loop wait until falling_edge(clk); end loop;
        rst_n   <= '1';
        slv_rst <= '0';
        for j in 1 to 2 loop wait until falling_edge(clk); end loop;
        clear_counts;
        got_n := 0;
        for i in 0 to 7 loop
            if i = 7 then push(x"000000C5", '1', 0);
            else          push(x"000000C5", '0', 0); end if;
            drain_rx;
        end loop;
        wait_idle;
        if slack_cycles = 0 then
            report "  FAIL: the shadow was never free while the clock was running -- it is not decoupling anything";
            errs := errs + 1;
        end if;
        if stall_cycles /= 0 then
            report "  FAIL: " & integer'image(stall_cycles) &
                   " stall cycles while refilling ahead";
            errs := errs + 1;
        end if;
        report "  the shadow stood free for " & integer'image(slack_cycles) &
               " of the cycles the wire was busy -- that window is the slack the producer gets, and it is a whole frame wide";

        errors <= errs;
        if errs = 0 then
            report "PASS: a one-deep shadow register decouples the producer from the wire -- sixteen words stream under a single chip select with every word correct in both directions and in order, zero stall cycles, and an inter-frame idle equal to the programmed hold and nothing more -- words are accepted while the previous frame is still shifting, which is the whole reason the shadow exists -- a producer forty cycles late still transfers every word correctly under one chip select but now reports the stall cycles it cost, so the lost throughput is visible instead of silent -- and a received word overwritten before it was read raises a sticky overrun that a later clean transfer does not clear";
        else
            report "FAIL: " & integer'image(errs) & " error(s)" severity error;
        end if;
        halt <= true;
        wait;
    end process;

end architecture;

Parity

All three implementations stream sixteen words under a single chip select with zero stall cycles, 38 idle cycles across fifteen boundaries — the programmed two-cycle hold and nothing more — and exactly 16 loads, 16 preloads, 112 launches and 128 captures. The same stream with the producer forty cycles late transfers every word correctly under one chip select and reports 65 stall cycles that the fast run did not have. And the shadow stands free for 31 of the cycles the wire is busy, which at a divisor of four and an eight-bit frame is a whole frame wide.

7. Why a Verification Engineer Cares

Azvya Education Pvt. Ltd.VLSI Mentor
spi_multibyte.sva — the handover, the slack, and the two flags
// The properties divide cleanly: correctness properties about the handover, and
// a performance property about the slack. The second is unusual in an assertion
// set and belongs there -- a design that is correct and slow fails its
// requirement just as surely as one that is fast and wrong.

module spi_multibyte_sva #(parameter int MAX_W = 32) (
    input logic              clk,
    input logic              rst_n,
    input logic              tx_push,
    input logic              tx_ready,
    input logic              flush,
    input logic              clr_flags,
    input logic              rx_pop,
    input logic              rx_ready,
    input logic              rx_overrun,
    input logic              core_busy,
    input logic              shift_en,
    input logic              start_stb,
    input logic              rx_valid_stb,
    input logic              req,
    input logic              load_stb,
    input logic              stalled
);

    default clocking cb @(posedge clk); endclocking
    default disable iff (!rst_n);

    // THE handover rule. The datapath may never be loaded while the clock runs,
    // because that would overwrite the word being shifted.
    a_no_load_while_shifting: assert property (load_stb |-> !shift_en);

    // And a load must never coincide with a frame start, because the datapath's
    // load and its first drive are both sampled on the same edge and the drive
    // wins -- so a coinciding load is silently discarded and the frame sends the
    // PREVIOUS word.
    a_no_load_at_start: assert property (!(load_stb && start_stb));

    // A push is accepted exactly when there is room, and acceptance must not
    // depend on anything else -- a design that also required the bus to be idle
    // would remove the whole benefit.
    a_push_accepted: assert property (tx_push && tx_ready |=> !tx_ready || flush);

    // THE slack property. The shadow must be free for some part of every frame,
    // or the producer has no more time than it had without the shadow. Stated as
    // a cover because it is about the stimulus AND the design together.
    c_slack_exists: cover property (tx_ready && shift_en);

    // A stall means the wire is idle inside a transaction with nothing queued.
    // It must NOT fire while the transaction is closing, which is the tail guard.
    a_stall_implies_idle: assert property (
        stalled |-> core_busy && !shift_en && !req
    );

    // The receive side. An overrun is data loss and must be sticky.
    a_overrun_sticky: assert property (
        rx_overrun && !clr_flags |=> rx_overrun
    );
    a_overrun_cause: assert property (
        rx_valid_stb && rx_ready && !rx_pop |=> rx_overrun
    );
    // And it must not fire when the consumer kept up, or it means nothing.
    a_no_false_overrun: assert property (
        !rx_overrun && rx_valid_stb && !rx_ready |=> !rx_overrun
    );

    // A pop on the same cycle as an arriving word keeps the new word: the pop
    // took the old one and the new one has not been read.
    a_pop_and_valid: assert property (
        rx_pop && rx_valid_stb |=> rx_ready
    );

    // Flush must leave nothing armed, or the abandoned transaction's word starts
    // the next transaction by itself.
    a_flush_disarms: assert property (flush |=> !req && tx_ready);

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_multibyte_cg.sv — producer latency, not data
// The axis that matters is WHEN the producer wrote, relative to the frame. The
// data is irrelevant here -- Chapter 13.6 covered data -- and the interesting
// bins are the boundaries: writing on the load cycle, and writing too late.

covergroup cg_multibyte @(posedge clk);

    // Where in the frame the push landed, as a fraction of the frame.
    push_when: coverpoint push_phase iff (tx_push && tx_ready) {
        bins on_load_cycle   = {0};    // the simultaneous push-and-load case
        bins early_in_frame  = {1};
        bins late_in_frame   = {2};
        bins in_the_gap      = {3};
        bins too_late        = {4};    // after the gap: a stall occurred
    }

    // Stall cycles per boundary. Zero is the design goal; the others must be
    // reachable or the stall counter is untested.
    stall_len: coverpoint stalls_this_boundary iff (frame_start) {
        bins none   = {0};
        bins one    = {1};
        bins few    = {[2:8]};
        bins many   = {[9:$]};
    }

    // Transaction length in frames. One frame has no boundary at all, so it
    // cannot exercise the shadow; two is the minimum that can.
    frames: coverpoint frames_this_txn iff (txn_end) {
        bins one  = {1};
        bins two  = {2};
        bins few  = {[3:16]};
        bins many = {[17:$]};
    }

    // The receive side's drain behaviour, which is what produces or avoids an
    // overrun. "Never" must be reachable, because that is the overrun test.
    drain: coverpoint drain_policy iff (txn_end) {
        bins every_word  = {0};
        bins every_other = {1};
        bins never       = {2};
    }

    // The tail: a transaction closing with nothing queued, which must NOT be
    // counted as a stall. This bin existing is how the guard stays tested.
    tail_seen: coverpoint tail iff (core_busy && !shift_en) {
        bins closing = {1};
        bins running = {0};
    }

    x_when_stall:  cross push_when, stall_len;
    x_frames_drain: cross frames, drain;

endgroup

8. Why an FPGA or ASIC Engineer Cares

The cost is one MAX_W register plus five flops. At MAX_W = 32 the shadow is 32 flops, plus sh_last, sh_valid, armed, armed_last and tail. The receive holding register is another 32 plus rx_full and rx_overrun. That is the whole feature: about 70 flops, no arithmetic, no wide combinational logic.

Nothing here is on a timing-critical path. can_load is a four-term AND of registered signals; the handover is a register-to-register transfer evaluated once per frame. The CPU-side interface is a write enable and a read enable.

The receive holding register is what makes the read interface safe. Without it, software would read the shift register directly — which is being shifted during the next frame, so a read landing mid-frame returns a partially-shifted word. The holding register is loaded once per frame at a known moment and held, which means the read path has no timing relationship to the SCLK rate at all.

If depth is increased, the shadow becomes a FIFO and the handover does not change. can_load becomes "the FIFO is not empty" and the free-on-load becomes a read pointer increment. That is worth noting because it means the depth decision can be deferred: build one deep, measure the stall counter on real traffic, and add depth only if the measurement says the producer cannot keep up. The interface to the rest of the design is identical either way.

Cost of getting the depth wrong in the other direction. A 16-deep FIFO on a design whose producer always keeps up is 512 flops doing nothing. On a small FPGA that is a block RAM that could have been something else, and it is very commonly spent because "deeper is faster" is an intuition that survives contact with almost no evidence.

9. Failure Signature — A Flash Page Program That Takes Four Times Too Long

Symptom. A flash driver programs a 256-byte page. The operation takes 2.1 ms where the arithmetic says 0.55 ms. The flash's own program time is 0.4 ms of that, and the SPI transfer should be 150 µs. SCLK is running at the configured rate and every byte arrives correctly; the page verifies.

What that rules out. Correct data and a correct SCLK frequency rule out the divider, the mode, the framing and the datapath. The transfer is taking four times as long as its byte count implies, which means the wire is idle most of the time — and the only thing that idles a correct wire is a producer that is not keeping up.

The measurement that identifies it. stalled, accumulated over the transaction, against the 259 frame boundaries:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   transfer time    150 us expected, 1700 us measured
   frames           260
   idle per frame   (1700 - 150) / 259 = 6.0 us

Six microseconds per byte, at a 100 MHz clock, is 600 cycles. That is not a hardware interval — it is an interrupt latency plus a handler, which puts the fault firmly in the driver's structure rather than in the master.

The mechanism. The driver takes an interrupt per byte. Each interrupt writes one byte and returns; the next byte is not queued until the next interrupt, which fires when the current frame completes. So the shadow is empty at every boundary and the wire waits for the interrupt every time — the shadow register is present and unused.

Why the shadow did not help. Because the driver never writes ahead. A one-deep shadow gives the producer a frame's worth of slack, and a producer that is only ever invoked at the boundary cannot use slack that exists before the boundary. The hardware feature is correct and the software does not exploit it.

The fix, in the driver. Write the next byte inside the same interrupt that writes the current one, while tx_ready is still asserted — which it is, because the shadow was freed at load. That single change takes the driver from one byte per interrupt to one byte per frame with the interrupt off the critical path, and the measured stall drops to zero.

Why a deeper FIFO would also have fixed it, and why that is the wrong fix. A 16-deep FIFO lets the interrupt-per-byte driver fall 16 frames behind without the wire idling, so the symptom disappears. It also spends 512 flops to paper over a driver structure that will cause the same problem on the next interface, and it does nothing for a driver that falls 17 behind. The stall counter pointed at the driver; the fix belongs there.

The general lesson. A throughput shortfall with correct data is always a gap problem, and the gap is measurable. stalled turns "it is slower than it should be" into "the producer is 600 cycles late per frame", which is a number that identifies the layer at fault.

10. Common Misconceptions

"A deeper FIFO is faster." It is not. Once the producer can stay one word ahead the wire is saturated, and no depth improves that. Depth buys tolerance of a longer producer absence, which is a different and also useful property — and the right depth is the worst-case producer latency divided by the frame time.

"A stall is an error." It is a cost. A low-priority background transfer that stalls constantly is behaving correctly. Treating it as an error produces a flag that gets masked, and then the one transfer where the stall matters is silent too.

"An overrun and a stall are the same kind of problem." A stall costs time; an overrun loses data. The first can clear on read; the second must be sticky, because a driver that checks once at the end of a transaction has to find out about a word dropped two hundred frames earlier.

"The shadow should be freed when the frame completes — that is when the word has actually gone." Freeing it at frame completion means it is full for the whole of every frame, so the producer gets only the inter-frame gap and the decoupling is worth nothing. Freeing it at load is the entire feature.

"Loading the datapath mid-frame is harmless if the load only changes the upper bits." It overwrites the shift register, which is mid-shift, so the bits still to be sent are replaced by bits from a word that was meant for the next frame. The guard is !shift_en and it is not negotiable.

11. Reason It Through

Why must a push and a load on the same cycle both succeed?

Because that cycle is the common case, not a corner. The CPU polling tx_ready and writing the instant it rises will hit exactly the cycle after the load, and a tight loop will hit the load cycle itself. A design that lost one of the two would drop a word on most fast streams, and the symptom would be a transaction one byte short with no flag set.

Why does stalled need the tail guard, and what would happen without it?

Because after the final frame the transaction is still open, the clock is stopped, and nothing is queued — by design. Without the guard every transaction ends with lag-plus-gap stall cycles, so a driver reading the counter concludes it is always late. A flag that fires on correct behaviour gets masked, and then it is worse than absent.

What is the right depth for the transmit buffer, and how would you determine it?

The worst-case producer latency divided by the frame time, rounded up. Determine it by measurement rather than by argument: build one deep, instrument stalled, and run real traffic. If the stall count is zero the depth is right; if it is not, the counter tells you how many frame times the producer is behind, which is the depth needed.

Why is there a separate receive holding register rather than letting software read the shift register?

Because the shift register is being shifted during the next frame, so a read landing mid-frame returns a partially-shifted word. The holding register is loaded once at a known moment and held, which removes any timing relationship between the read path and the SCLK rate — and it is what makes the overrun flag meaningful, since there is now a definite word that can be overwritten.

A design has a 16-deep FIFO, its wire never idles, and its throughput is still half what the arithmetic predicts. Where would you look?

Not at the buffer — a wire that never idles is saturated, so the shortfall is not a gap problem. The remaining candidates are the intervals: a lead, lag or gap programmed larger than the slave requires, which Chapter 13.7's two-sided interval checks exist to catch; or a frame width smaller than it needs to be, so the fixed per-frame overhead is amortised over fewer bits, which is Chapter 9.2's arithmetic. Measure the idle cycles inside a transaction: if they equal the programmed gap times the boundary count, the gap is the answer.

12. Understanding Check

13. Summary

With a single transmit register the wire is idle at every frame boundary for however long the CPU takes to notice — 24% utilisation for an interrupt-per-byte driver at a plausible clock ratio. A one-deep shadow removes the dependency by giving the producer a whole frame instead of a gap.

A deeper FIFO buys latency tolerance, not throughput. Once the producer stays one word ahead the wire is saturated; depth determines how long the producer may be absent. The right depth is the worst-case producer latency divided by the frame time, and it should be measured rather than argued — build one deep, instrument the stall counter, add depth if the numbers say so.

The shadow is freed at load, not at frame completion. One line, and it is the difference between a shadow register and a delay.

Two flags, and the distinction between them is a general pattern: stalled reports a cost and is not an error — a low-priority transfer stalls constantly and correctly — while rx_overrun reports a loss and must be sticky, because a driver may only check at the end of a two-hundred-frame transaction.

stalled needs the tail guard, because after the final frame the transaction is still open with nothing queued by design. Without it every transaction ends with a burst of false stalls, and a flag that fires on correct behaviour gets masked.

The datapath may only be loaded while the clock is stopped, a push and a load on the same cycle must both succeed — that cycle is the common case for a tight loop, not a corner — and flush must leave nothing armed, or an abandoned transaction's word starts the next one by itself.

Integrating this block found a stale-level bug in Chapter 13.5 that no unit test could have found: a level meaning "the previous frame finished" was still true on the next frame's first cycle. Any consumer that restarts and observes in the same cycle needs the start term.

For verification, the driver's pacing is the feature under test — a driver that waits for completion before writing the next word exercises none of it — and the scoreboard must check stalled against the delay the sequence asked for, or both a broken flag and a broken handover are silent.

14. What Comes Next

The master streams correctly and reports what it costs. Every path through it so far assumes the transfer finishes.

Chapter 13.10 — Busy/Done/Valid Interface, Reset, and Abort handles the case where it does not, and the two ways of stopping turn out to be opposites: one is orderly and costs time, the other is instant and costs protocol correctness — and knowing which is which decides whether the slave is still usable afterwards.

Continue learning