Skip to content
VLSI Mentor

SPI · Module 19

FPGA to ADC — Streaming Under a Sample Deadline

Throughput and a deadline are different requirements. The measured busy interval matches t_conv + lead + 2·NB·half + lag + gap exactly — and the fixed terms cap what any clock rate can buy.

Modules 13 to 18 built an SPI master, constrained it, verified it and learned to debug it. Every one of those chapters could treat a transfer as the unit of work. This module cannot, because in a real system a transfer is never the point — it is one step in a path that starts at a device and ends somewhere software can see.

A sample deadline is not a throughput requirement. A design with ample average throughput can miss it every period, and the sample is simply not there.

1. The Problem

An ADC converts on a schedule. Every sample period the FPGA must start a conversion, wait the datasheet's conversion time, read the result over SPI, and be finished before the next period begins.

Not on average. Every period, forever.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   period ──┬── t_conv ──┬── lead ── 2·NB·half ── lag ──┬── gap ──┬── period ──▶
            │            │                              │         │
        conversion    SPI frame                      release   deselected
          starts       begins                                   minimum

If the loop does not close inside period, a sample is lost. And the symptom is the worst kind: a missing sample in a stream of numbers does not look like a bug, it looks like noise. Nobody opens a waveform for noise.

2. Two Loss Mechanisms, And They Are Not The Same Bug

This is the first thing to get right, because it decides which half of the system you spend a week optimising.

CounterWhat happenedWhat fixes it
n_overrunthe schedule ticked while the previous read was still runningfaster SCLK, fewer bits, shorter gap, or a slower sample rate
n_droppedthe read finished and the consumer's FIFO was fulla deeper FIFO, or a faster consumer — nothing about SPI helps
A left-to-right path: a sample timer feeds an ADC conversion, which feeds an SPI read frame, which pushes into a sample FIFO that crosses to the consumer clock. Below the path, an overrun counter is fed from the timer when the engine is busy, and a dropped counter is fed from the FIFO when it is full.sample timerone tick every`period`ADC conversiont_conv cycles, no SPIactivitySPI read framelead + 2·NB·half + lagsample FIFOcrosses to theconsumer clockn_overruntick arrived, readstill runningn_droppedFIFO full when theread finishedtickresult readypush sampleengine busyfull12
Figure 1 — where a sample period is consumed, and the two places a sample can be lost. The conversion and the frame are serial: no SPI activity happens while the device is converting. The two loss counters hang off different points of the path, which is why they cannot be merged.

3. There Is No Clock-Domain Crossing On This Read Path

Worth stating plainly, because SPI and CDC are habitually assumed to travel together.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   a SPI SLAVE   receives SCLK from outside → every capture flop is in a foreign
                 domain → a crossing is unavoidable

   a SPI MASTER  GENERATES SCLK by dividing its own clock → every flop on the
                 read path is in the system domain → nothing crosses

The master here has no crossing on the read path at all. Module 15 established when SCLK is a clock domain and when it is a derived enable; this is the case where it is an enable.

The crossing in this chapter is somewhere else entirely: the fifo_full interface, where samples leave for a consumer on a different clock. Chapter 19.3 takes up the other crossing a real integration has — the register bus.

4. The Arithmetic The Design Is Accountable To

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   busy = t_conv + lead + 2·NB·half + lag + gap          [system-clock cycles]

   the schedule holds exactly when   period ≥ busy

With the values used below — t_conv 20, lead 3, NB 16, half 2, lag 2, gap 4:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   busy = 20 + 3 + 2·16·2 + 2 + 4 = 93 cycles

The design publishes the busy interval it actually took, and the bench requires it to equal that expression. A predicted number that matches is evidence the model is right; one that does not is a finding either way — which is worth much more than a golden value recorded from a previous run.

5. What A Deadline Miss Looks Like On The Pins

A tick that arrives while the engine is busy produces no conversion

18 cycles
Five rows over eighteen cycles. A tick row marks cycles zero and ten. CONVST pulses only at cycle zero. Chip select is low from cycle three to fourteen. SCLK toggles once per cycle from five to twelve. The overrun counter steps from zero to one at cycle ten, where the second tick arrives while the read is still running. Phase bands mark the conversion time, the lead, the shift, the lag and the deselected gap.t_conv = 3t_conv = 3leadlead2·NB·half = 82·NB·half = 8laglaggapgaptick: conversion startstick: conversion startstick — read still runningtick — read still runningsample pushedsample pushedtickconvstcs_nsclkn_overrun000000000011111111t0t1t2t3t4t5t6t7t8t9t10t11t12t13t14t15t16t17
Figure 2 — a scaled-down cycle (NB = 4, half = 1) with the schedule set shorter than the busy interval. The second tick arrives during the shift phase; no CONVST appears for it, the overrun counter increments, and the read in flight is left alone to finish. Horizontal scrolling is expected on a narrow screen; the figure is a timeline.

Two details in that figure are the chapter.

The schedule tick at cycle 10 produces no CONVST. The sample period is consumed and nothing comes out of it — and note that every pin in the figure looks legal. There is no protocol violation anywhere; Chapter 18.1's triage would classify the frame as EV_WELL. A deadline miss is invisible to a protocol checker.

And the read in flight is left alone to finish. Aborting it would lose the sample already in progress as well as this period's — one missed deadline would cost two samples instead of one.

6. The Measurement

Identical output from all three languages:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
  measured busy interval: 93 cycles   predicted 93 = t_conv 20 + lead 3 + 2*16*2 + lag 2 + gap 4
  period = busy (93):      conv 7  samp 6  overrun 0  dropped 0
  period = busy - 1 (92):  conv 4  samp 3  overrun 3  dropped 0
  consumer stalled (period 200): conv 7  samp 0  overrun 0  dropped 6

  half  SCLK period  busy cycles  sample rate vs half=2
     1            2           61  1.52x
     2            4           93  1.00x
     3            6          125  0.74x
     4            8          157  0.59x
     6           12          221  0.42x
     8           16          285  0.33x

The boundary is exact. At period = 93 the schedule held with zero overruns. At 92 — one cycle less — it did not. A boundary that is only approximately known is a boundary nobody can design against, and the reason this one is exact is that the design and the arithmetic were made to agree rather than merely to coexist.

The two mechanisms separate cleanly. The short period produced overruns and zero drops. The stalled consumer produced drops and zero overruns, and delivered no samples at all.

And SCLK is not a sample-rate knob past a point. Halving the half period from 2 to 1 cut the busy interval from 93 to 61 cycles — a 1.52× gain, not 2×. The fixed floor of 29 cycles caps the total available gain at 3.21×.

7. Pushing SCLK Eventually Breaks The Data, And The Deadline Cannot See It

The half = 1 row of that table deserves its own section, because the sweep hides something.

At half = 1 the busy interval is 61 cycles — the best schedule in the table. And the delivered samples are wrong. The ADC model presents each bit one system cycle after the trailing edge, so a mode-0 master's leading edge has half − 1 cycles of setup, and at half = 1 it has none.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   half = 2   setup = 1 cycle    deadline 93   data correct
   half = 1   setup = 0 cycles   deadline 61   data corrupt

The bench expects the corruption there and requires it — so the table above is not hiding a broken run, it is reporting a boundary.

8. Building It — Three HDLs

The design is a scheduler and a read engine sharing one state register, which is where its only subtle bug lived.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_adc_stream.sv — the streaming reader — a schedule, a read engine, and two loss counters that cannot be merged
// spi_adc_stream.sv
//
// Chapter 19.1 -- streaming an ADC over SPI under a HARD sample deadline.
//
// THE PROBLEM THIS SOLVES, and it is not a throughput problem.
//
// An ADC converts on a schedule. Every sample period the FPGA must start a conversion, wait the
// datasheet's conversion time, read the result over SPI, and be finished before the next period
// begins. Not on average -- EVERY period, forever. A design with ample average throughput can still
// miss the deadline, and when it does the sample is simply not there.
//
// THAT IS WHY THIS MODULE EXISTS AS HARDWARE RATHER THAN AS A CALCULATION. The deadline is a property
// of a closed loop between a timer, a conversion delay, an SPI frame and a consumer, and the only
// honest way to know whether it holds is to run the loop and count what came out.
//
// TWO LOSS MECHANISMS, TWO COUNTERS, AND THEY ARE NOT THE SAME BUG.
//
//     n_overrun   the schedule ticked while the previous read was still running.
//                 The SPI path is too slow for the requested sample rate.
//                 FIX: faster SCLK, fewer bits, shorter gap, or a slower sample rate.
//
//     n_dropped   the read finished, and the consumer's FIFO was full.
//                 The SPI path kept up and something DOWNSTREAM did not.
//                 FIX: a deeper FIFO, or a faster consumer. Nothing about SPI helps.
//
// Collapsing these into one "samples lost" counter is the most common instrumentation mistake in a
// streaming datapath, and it is expensive: the two have opposite fixes, and a single counter sends
// you to optimise the wrong half of the system. The bench below drives each mechanism in isolation
// and requires the OTHER counter to stay at zero.
//
// THE ARITHMETIC THE DESIGN IS ACCOUNTABLE TO.
//
//     busy = t_conv + lead + 2*NB*half + lag + gap          [system-clock cycles]
//
// and the schedule holds exactly when `period >= busy`. The bench MEASURES the busy interval and
// requires it to equal that expression, so the formula is checked against the hardware rather than
// asserted next to it. A predicted number that matches is evidence the model is right; one that does
// not is a finding either way.
//
// AND THE CONSEQUENCE OF THE FIXED TERMS, which is the result worth carrying away. Only the
// `2*NB*half` term responds to SCLK. Everything else -- the conversion time, the select lead and lag,
// the minimum deselected gap -- is fixed by the device and the protocol. So doubling SCLK does NOT
// double the achievable sample rate, and past a point it barely moves it at all. The bench computes
// the limit.
//
// WHAT THIS MODULE DELIBERATELY DOES NOT CONTAIN: a clock-domain crossing on the read path. The FPGA
// GENERATES SCLK here, so every flop on that path is in the system domain and there is nothing to
// cross. That is worth stating because "SPI" and "CDC" are habitually assumed to travel together.
// They travel together in a SLAVE, and in a master they travel together only where the master meets
// something else -- a consumer pipeline, or a register bus. The consumer crossing is the `fifo_full`
// interface here, and Chapter 19.3 takes up the register-bus crossing.

`timescale 1ns/1ps

module spi_adc_stream #(
    parameter int NB    = 16,   // ADC result width in bits
    parameter int CNT_W = 16
) (
    input  wire              clk,
    input  wire              rst_n,

    // ---- schedule and device timing, all in system-clock cycles ----
    input  wire [CNT_W-1:0]  period,   // the sample deadline: one conversion every `period` cycles
    input  wire [CNT_W-1:0]  t_conv,   // datasheet conversion time
    input  wire [7:0]        half,     // SCLK half period
    input  wire [7:0]        lead,     // CS low to first SCLK edge
    input  wire [7:0]        lag,      // last SCLK edge to CS high
    input  wire [7:0]        gap,      // minimum CS-high time between frames

    // ---- pins ----
    output reg               convst,
    output reg               sclk,
    output reg               cs_n,
    input  wire              miso,

    // ---- sample handoff to the consumer domain ----
    output reg  [NB-1:0]     samp_data,
    output reg               samp_push,
    input  wire              fifo_full,

    // ---- health counters ----
    output reg  [CNT_W-1:0]  n_conv,
    output reg  [CNT_W-1:0]  n_samp,
    output reg  [CNT_W-1:0]  n_overrun,
    output reg  [CNT_W-1:0]  n_dropped,

    // ---- the measured busy interval, published so the arithmetic can be checked ----
    output reg  [CNT_W-1:0]  busy_cycles,
    output reg               busy_valid
);

    localparam [2:0] S_IDLE  = 3'd0,
                     S_CONV  = 3'd1,   // CONVST asserted, waiting out the conversion time
                     S_LEAD  = 3'd2,   // CS low, waiting the select lead
                     S_SHIFT = 3'd3,   // clocking the result out of the device
                     S_LAG   = 3'd4,   // last edge sent, holding CS low for the lag
                     S_GAP   = 3'd5;   // CS high, honouring the minimum deselected time

    reg [2:0]        st;
    reg [CNT_W-1:0]  tmr;        // the free-running schedule timer
    reg [CNT_W-1:0]  dwell;      // time spent in the current state
    reg [CNT_W-1:0]  busy;       // cycles since the conversion started
    reg [7:0]        edges;      // SCLK edges emitted this frame
    reg [NB-1:0]     sh;

    // The schedule tick. It is free-running and INDEPENDENT of the read engine, which is the whole
    // point: a deadline is only a deadline if the thing imposing it does not wait for you.
    wire tick = (tmr + {{(CNT_W-1){1'b0}}, 1'b1} >= period);

    // THE ENGINE IS AVAILABLE IF IT IS IDLE **OR** FINISHING THIS CYCLE, and that second term is not a
    // nicety. `st` still reads S_GAP on the cycle the gap expires, so a scheduler that tested only
    // `st == S_IDLE` would refuse a tick that arrives exactly when the engine becomes free -- and the
    // minimum workable period would be one cycle longer than the arithmetic says, for a reason no
    // amount of staring at the equation would reveal. Back-to-back capability is the difference
    // between a design that meets its budget and one that misses it by a single cycle.
    wire gap_done  = (st == S_GAP) && (dwell + 1'b1 >= {{(CNT_W-8){1'b0}}, gap});
    wire can_start = (st == S_IDLE) || gap_done;

    always_ff @(posedge clk or negedge rst_n) begin
        if (!rst_n) begin
            st          <= S_IDLE;
            tmr         <= {CNT_W{1'b0}};
            dwell       <= {CNT_W{1'b0}};
            busy        <= {CNT_W{1'b0}};
            edges       <= 8'd0;
            sh          <= {NB{1'b0}};
            convst      <= 1'b0;
            sclk        <= 1'b0;
            cs_n        <= 1'b1;
            samp_data   <= {NB{1'b0}};
            samp_push   <= 1'b0;
            n_conv      <= {CNT_W{1'b0}};
            n_samp      <= {CNT_W{1'b0}};
            n_overrun   <= {CNT_W{1'b0}};
            n_dropped   <= {CNT_W{1'b0}};
            busy_cycles <= {CNT_W{1'b0}};
            busy_valid  <= 1'b0;
        end else begin
            samp_push  <= 1'b0;
            busy_valid <= 1'b0;
            convst     <= 1'b0;


            // ---- the read engine ----
            case (st)
                S_CONV: begin
                    if (dwell + 1'b1 >= t_conv) begin
                        st    <= S_LEAD;
                        dwell <= {CNT_W{1'b0}};
                        cs_n  <= 1'b0;
                    end else dwell <= dwell + 1'b1;
                end

                S_LEAD: begin
                    // SCLK stays at its idle level through the lead. Leaving it low here and
                    // producing every edge inside S_SHIFT is what keeps the edge count honest: an
                    // edge emitted on a state transition is an edge the shift loop never counts,
                    // and the measured busy interval then disagrees with the arithmetic by one half
                    // period for reasons nobody can find.
                    if (dwell + 1'b1 >= {{(CNT_W-8){1'b0}}, lead}) begin
                        st    <= S_SHIFT;
                        dwell <= {CNT_W{1'b0}};
                        edges <= 8'd0;
                    end else dwell <= dwell + 1'b1;
                end

                S_SHIFT: begin
                    if (dwell + 1'b1 >= {{(CNT_W-8){1'b0}}, half}) begin
                        dwell <= {CNT_W{1'b0}};
                        // MODE 0: capture AT the leading edge. The device has been holding this bit
                        // since the previous trailing edge (or since the select, for the first bit),
                        // so it is stable when the edge arrives. Capturing a half period later would
                        // give more setup margin and would no longer be mode 0 -- a real design
                        // choice, and one that has to be declared rather than drifted into.
                        if (!sclk) begin
                            sh   <= {sh[NB-2:0], miso};
                            sclk <= 1'b1;            // the leading edge of this bit
                        end else begin
                            sclk <= 1'b0;            // the trailing edge
                        end
                        // 2*NB edges, starting with a rising one, so the last is falling and SCLK is
                        // left at its idle level with no extra assignment needed.
                        if (edges + 8'd1 >= 8'd2 * NB[7:0]) begin
                            st <= S_LAG;
                        end else begin
                            edges <= edges + 8'd1;
                        end
                    end else dwell <= dwell + 1'b1;
                end

                S_LAG: begin
                    if (dwell + 1'b1 >= {{(CNT_W-8){1'b0}}, lag}) begin
                        st    <= S_GAP;
                        dwell <= {CNT_W{1'b0}};
                        cs_n  <= 1'b1;
                        // THE HANDOFF, AND THE SECOND LOSS MECHANISM. The read succeeded; whether the
                        // sample survives now depends on something outside this module entirely.
                        if (!fifo_full) begin
                            samp_data <= sh;
                            samp_push <= 1'b1;
                            n_samp    <= n_samp + 1'b1;
                        end else begin
                            n_dropped <= n_dropped + 1'b1;
                        end
                    end else dwell <= dwell + 1'b1;
                end

                S_GAP: begin
                    if (dwell + 1'b1 >= {{(CNT_W-8){1'b0}}, gap}) begin
                        st          <= S_IDLE;
                        dwell       <= {CNT_W{1'b0}};
                        // Publish the measured busy interval. `busy` counted every cycle from the
                        // tick that started the conversion, so this is the number the formula in the
                        // header has to reproduce.
                        busy_cycles <= busy + 1'b1;
                        busy_valid  <= 1'b1;
                    end else dwell <= dwell + 1'b1;
                end

                default: ;   // S_IDLE waits for the schedule
            endcase

            // ---- the schedule ----
            //
            // THE SCHEDULE IS EVALUATED **AFTER** THE READ ENGINE, and that order is load-bearing.
            //
            // Both write `st`, and on a perfectly tight schedule they write it in the SAME cycle: the
            // gap expires (S_GAP -> S_IDLE) exactly as the tick arrives (-> S_CONV). In a single
            // clocked block the last assignment wins, so with the schedule evaluated FIRST the engine
            // fell back to S_IDLE, the tick had already counted a conversion and pulsed CONVST, and
            // that sample was never read -- one conversion counted, no sample delivered, and no
            // counter anywhere reporting a problem. The bench caught it as a sample compared against
            // its predecessor.
            //
            // Two writers of one state register need a DECLARED priority, and here the schedule must
            // win, because a tick that has already been counted must be honoured.
            // ---- the schedule ----
            if (tick) begin
                tmr <= {CNT_W{1'b0}};
                if (can_start) begin
                    // Start a conversion. CONVST is a single-cycle pulse; the device latches its
                    // input on it and begins converting.
                    convst <= 1'b1;
                    n_conv <= n_conv + 1'b1;
                    st     <= S_CONV;
                    dwell  <= {CNT_W{1'b0}};
                    busy   <= {CNT_W{1'b0}};
                end else begin
                    // THE DEADLINE MISS. The previous read has not finished, so this sample period
                    // produces no conversion at all. The engine is left alone to finish -- aborting
                    // it would lose the sample already in flight as well as this one.
                    n_overrun <= n_overrun + 1'b1;
                end
            end else begin
                tmr <= tmr + 1'b1;
            end

            if (st != S_IDLE) busy <= busy + 1'b1;
        end
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_adc_stream.v — the same design in Verilog-2001
// spi_adc_stream.v
//
// Chapter 19.1 -- streaming an ADC over SPI under a HARD sample deadline.
//
// THE PROBLEM THIS SOLVES, and it is not a throughput problem.
//
// An ADC converts on a schedule. Every sample period the FPGA must start a conversion, wait the
// datasheet's conversion time, read the result over SPI, and be finished before the next period
// begins. Not on average -- EVERY period, forever. A design with ample average throughput can still
// miss the deadline, and when it does the sample is simply not there.
//
// THAT IS WHY THIS MODULE EXISTS AS HARDWARE RATHER THAN AS A CALCULATION. The deadline is a property
// of a closed loop between a timer, a conversion delay, an SPI frame and a consumer, and the only
// honest way to know whether it holds is to run the loop and count what came out.
//
// TWO LOSS MECHANISMS, TWO COUNTERS, AND THEY ARE NOT THE SAME BUG.
//
//     n_overrun   the schedule ticked while the previous read was still running.
//                 The SPI path is too slow for the requested sample rate.
//                 FIX: faster SCLK, fewer bits, shorter gap, or a slower sample rate.
//
//     n_dropped   the read finished, and the consumer's FIFO was full.
//                 The SPI path kept up and something DOWNSTREAM did not.
//                 FIX: a deeper FIFO, or a faster consumer. Nothing about SPI helps.
//
// Collapsing these into one "samples lost" counter is the most common instrumentation mistake in a
// streaming datapath, and it is expensive: the two have opposite fixes, and a single counter sends
// you to optimise the wrong half of the system. The bench below drives each mechanism in isolation
// and requires the OTHER counter to stay at zero.
//
// THE ARITHMETIC THE DESIGN IS ACCOUNTABLE TO.
//
//     busy = t_conv + lead + 2*NB*half + lag + gap          [system-clock cycles]
//
// and the schedule holds exactly when `period >= busy`. The bench MEASURES the busy interval and
// requires it to equal that expression, so the formula is checked against the hardware rather than
// asserted next to it. A predicted number that matches is evidence the model is right; one that does
// not is a finding either way.
//
// AND THE CONSEQUENCE OF THE FIXED TERMS, which is the result worth carrying away. Only the
// `2*NB*half` term responds to SCLK. Everything else -- the conversion time, the select lead and lag,
// the minimum deselected gap -- is fixed by the device and the protocol. So doubling SCLK does NOT
// double the achievable sample rate, and past a point it barely moves it at all. The bench computes
// the limit.
//
// WHAT THIS MODULE DELIBERATELY DOES NOT CONTAIN: a clock-domain crossing on the read path. The FPGA
// GENERATES SCLK here, so every flop on that path is in the system domain and there is nothing to
// cross. That is worth stating because "SPI" and "CDC" are habitually assumed to travel together.
// They travel together in a SLAVE, and in a master they travel together only where the master meets
// something else -- a consumer pipeline, or a register bus. The consumer crossing is the `fifo_full`
// interface here, and Chapter 19.3 takes up the register-bus crossing.

`timescale 1ns/1ps

module spi_adc_stream #(
    parameter NB    = 16,   // ADC result width in bits
    parameter CNT_W = 16
) (
    input  wire              clk,
    input  wire              rst_n,

    // ---- schedule and device timing, all in system-clock cycles ----
    input  wire [CNT_W-1:0]  period,   // the sample deadline: one conversion every `period` cycles
    input  wire [CNT_W-1:0]  t_conv,   // datasheet conversion time
    input  wire [7:0]        half,     // SCLK half period
    input  wire [7:0]        lead,     // CS low to first SCLK edge
    input  wire [7:0]        lag,      // last SCLK edge to CS high
    input  wire [7:0]        gap,      // minimum CS-high time between frames

    // ---- pins ----
    output reg               convst,
    output reg               sclk,
    output reg               cs_n,
    input  wire              miso,

    // ---- sample handoff to the consumer domain ----
    output reg  [NB-1:0]     samp_data,
    output reg               samp_push,
    input  wire              fifo_full,

    // ---- health counters ----
    output reg  [CNT_W-1:0]  n_conv,
    output reg  [CNT_W-1:0]  n_samp,
    output reg  [CNT_W-1:0]  n_overrun,
    output reg  [CNT_W-1:0]  n_dropped,

    // ---- the measured busy interval, published so the arithmetic can be checked ----
    output reg  [CNT_W-1:0]  busy_cycles,
    output reg               busy_valid
);

    localparam [2:0] S_IDLE  = 3'd0,
                     S_CONV  = 3'd1,   // CONVST asserted, waiting out the conversion time
                     S_LEAD  = 3'd2,   // CS low, waiting the select lead
                     S_SHIFT = 3'd3,   // clocking the result out of the device
                     S_LAG   = 3'd4,   // last edge sent, holding CS low for the lag
                     S_GAP   = 3'd5;   // CS high, honouring the minimum deselected time

    reg [2:0]        st;
    reg [CNT_W-1:0]  tmr;        // the free-running schedule timer
    reg [CNT_W-1:0]  dwell;      // time spent in the current state
    reg [CNT_W-1:0]  busy;       // cycles since the conversion started
    reg [7:0]        edges;      // SCLK edges emitted this frame
    reg [NB-1:0]     sh;

    // The schedule tick. It is free-running and INDEPENDENT of the read engine, which is the whole
    // point: a deadline is only a deadline if the thing imposing it does not wait for you.
    wire tick = (tmr + {{(CNT_W-1){1'b0}}, 1'b1} >= period);

    // THE ENGINE IS AVAILABLE IF IT IS IDLE **OR** FINISHING THIS CYCLE, and that second term is not a
    // nicety. `st` still reads S_GAP on the cycle the gap expires, so a scheduler that tested only
    // `st == S_IDLE` would refuse a tick that arrives exactly when the engine becomes free -- and the
    // minimum workable period would be one cycle longer than the arithmetic says, for a reason no
    // amount of staring at the equation would reveal. Back-to-back capability is the difference
    // between a design that meets its budget and one that misses it by a single cycle.
    wire gap_done  = (st == S_GAP) && (dwell + 1'b1 >= {{(CNT_W-8){1'b0}}, gap});
    wire can_start = (st == S_IDLE) || gap_done;

    always @(posedge clk or negedge rst_n) begin
        if (!rst_n) begin
            st          <= S_IDLE;
            tmr         <= {CNT_W{1'b0}};
            dwell       <= {CNT_W{1'b0}};
            busy        <= {CNT_W{1'b0}};
            edges       <= 8'd0;
            sh          <= {NB{1'b0}};
            convst      <= 1'b0;
            sclk        <= 1'b0;
            cs_n        <= 1'b1;
            samp_data   <= {NB{1'b0}};
            samp_push   <= 1'b0;
            n_conv      <= {CNT_W{1'b0}};
            n_samp      <= {CNT_W{1'b0}};
            n_overrun   <= {CNT_W{1'b0}};
            n_dropped   <= {CNT_W{1'b0}};
            busy_cycles <= {CNT_W{1'b0}};
            busy_valid  <= 1'b0;
        end else begin
            samp_push  <= 1'b0;
            busy_valid <= 1'b0;
            convst     <= 1'b0;


            // ---- the read engine ----
            case (st)
                S_CONV: begin
                    if (dwell + 1'b1 >= t_conv) begin
                        st    <= S_LEAD;
                        dwell <= {CNT_W{1'b0}};
                        cs_n  <= 1'b0;
                    end else dwell <= dwell + 1'b1;
                end

                S_LEAD: begin
                    // SCLK stays at its idle level through the lead. Leaving it low here and
                    // producing every edge inside S_SHIFT is what keeps the edge count honest: an
                    // edge emitted on a state transition is an edge the shift loop never counts,
                    // and the measured busy interval then disagrees with the arithmetic by one half
                    // period for reasons nobody can find.
                    if (dwell + 1'b1 >= {{(CNT_W-8){1'b0}}, lead}) begin
                        st    <= S_SHIFT;
                        dwell <= {CNT_W{1'b0}};
                        edges <= 8'd0;
                    end else dwell <= dwell + 1'b1;
                end

                S_SHIFT: begin
                    if (dwell + 1'b1 >= {{(CNT_W-8){1'b0}}, half}) begin
                        dwell <= {CNT_W{1'b0}};
                        // MODE 0: capture AT the leading edge. The device has been holding this bit
                        // since the previous trailing edge (or since the select, for the first bit),
                        // so it is stable when the edge arrives. Capturing a half period later would
                        // give more setup margin and would no longer be mode 0 -- a real design
                        // choice, and one that has to be declared rather than drifted into.
                        if (!sclk) begin
                            sh   <= {sh[NB-2:0], miso};
                            sclk <= 1'b1;            // the leading edge of this bit
                        end else begin
                            sclk <= 1'b0;            // the trailing edge
                        end
                        // 2*NB edges, starting with a rising one, so the last is falling and SCLK is
                        // left at its idle level with no extra assignment needed.
                        if (edges + 8'd1 >= 8'd2 * NB[7:0]) begin
                            st <= S_LAG;
                        end else begin
                            edges <= edges + 8'd1;
                        end
                    end else dwell <= dwell + 1'b1;
                end

                S_LAG: begin
                    if (dwell + 1'b1 >= {{(CNT_W-8){1'b0}}, lag}) begin
                        st    <= S_GAP;
                        dwell <= {CNT_W{1'b0}};
                        cs_n  <= 1'b1;
                        // THE HANDOFF, AND THE SECOND LOSS MECHANISM. The read succeeded; whether the
                        // sample survives now depends on something outside this module entirely.
                        if (!fifo_full) begin
                            samp_data <= sh;
                            samp_push <= 1'b1;
                            n_samp    <= n_samp + 1'b1;
                        end else begin
                            n_dropped <= n_dropped + 1'b1;
                        end
                    end else dwell <= dwell + 1'b1;
                end

                S_GAP: begin
                    if (dwell + 1'b1 >= {{(CNT_W-8){1'b0}}, gap}) begin
                        st          <= S_IDLE;
                        dwell       <= {CNT_W{1'b0}};
                        // Publish the measured busy interval. `busy` counted every cycle from the
                        // tick that started the conversion, so this is the number the formula in the
                        // header has to reproduce.
                        busy_cycles <= busy + 1'b1;
                        busy_valid  <= 1'b1;
                    end else dwell <= dwell + 1'b1;
                end

                default: ;   // S_IDLE waits for the schedule
            endcase

            // ---- the schedule ----
            //
            // THE SCHEDULE IS EVALUATED **AFTER** THE READ ENGINE, and that order is load-bearing.
            //
            // Both write `st`, and on a perfectly tight schedule they write it in the SAME cycle: the
            // gap expires (S_GAP -> S_IDLE) exactly as the tick arrives (-> S_CONV). In a single
            // clocked block the last assignment wins, so with the schedule evaluated FIRST the engine
            // fell back to S_IDLE, the tick had already counted a conversion and pulsed CONVST, and
            // that sample was never read -- one conversion counted, no sample delivered, and no
            // counter anywhere reporting a problem. The bench caught it as a sample compared against
            // its predecessor.
            //
            // Two writers of one state register need a DECLARED priority, and here the schedule must
            // win, because a tick that has already been counted must be honoured.
            // ---- the schedule ----
            if (tick) begin
                tmr <= {CNT_W{1'b0}};
                if (can_start) begin
                    // Start a conversion. CONVST is a single-cycle pulse; the device latches its
                    // input on it and begins converting.
                    convst <= 1'b1;
                    n_conv <= n_conv + 1'b1;
                    st     <= S_CONV;
                    dwell  <= {CNT_W{1'b0}};
                    busy   <= {CNT_W{1'b0}};
                end else begin
                    // THE DEADLINE MISS. The previous read has not finished, so this sample period
                    // produces no conversion at all. The engine is left alone to finish -- aborting
                    // it would lose the sample already in flight as well as this one.
                    n_overrun <= n_overrun + 1'b1;
                end
            end else begin
                tmr <= tmr + 1'b1;
            end

            if (st != S_IDLE) busy <= busy + 1'b1;
        end
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_adc_stream.vhd — the same design in VHDL
-- spi_adc_stream.vhd
--
-- Chapter 19.1 -- streaming an ADC over SPI under a HARD sample deadline.
--
-- THE PROBLEM THIS SOLVES, and it is not a throughput problem.
--
-- An ADC converts on a schedule. Every sample period the FPGA must start a conversion, wait the
-- datasheet's conversion time, read the result over SPI, and be finished before the next period
-- begins. Not on average -- EVERY period, forever. A design with ample average throughput can still
-- miss the deadline, and when it does the sample is simply not there.
--
-- THAT IS WHY THIS MODULE EXISTS AS HARDWARE RATHER THAN AS A CALCULATION. The deadline is a property
-- of a closed loop between a timer, a conversion delay, an SPI frame and a consumer, and the only
-- honest way to know whether it holds is to run the loop and count what came out.
--
-- TWO LOSS MECHANISMS, TWO COUNTERS, AND THEY ARE NOT THE SAME BUG.
--
--     n_overrun   the schedule ticked while the previous read was still running.
--                 The SPI path is too slow for the requested sample rate.
--                 FIX: faster SCLK, fewer bits, shorter gap, or a slower sample rate.
--
--     n_dropped   the read finished, and the consumer's FIFO was full.
--                 The SPI path kept up and something DOWNSTREAM did not.
--                 FIX: a deeper FIFO, or a faster consumer. Nothing about SPI helps.
--
-- Collapsing these into one "samples lost" counter is the most common instrumentation mistake in a
-- streaming datapath, and it is expensive: the two have opposite fixes, and a single counter sends
-- you to optimise the wrong half of the system. The bench below drives each mechanism in isolation
-- and requires the OTHER counter to stay at zero.
--
-- THE ARITHMETIC THE DESIGN IS ACCOUNTABLE TO.
--
--     busy = t_conv + lead + 2*NB*half + lag + gap          [system-clock cycles]
--
-- and the schedule holds exactly when `period >= busy`. The bench MEASURES the busy interval and
-- requires it to equal that expression, so the formula is checked against the hardware rather than
-- asserted next to it. A predicted number that matches is evidence the model is right; one that does
-- not is a finding either way.
--
-- AND THE CONSEQUENCE OF THE FIXED TERMS, which is the result worth carrying away. Only the
-- `2*NB*half` term responds to SCLK. Everything else -- the conversion time, the select lead and lag,
-- the minimum deselected gap -- is fixed by the device and the protocol. So doubling SCLK does NOT
-- double the achievable sample rate, and past a point it barely moves it at all. The bench computes
-- the limit.
--
-- WHAT THIS MODULE DELIBERATELY DOES NOT CONTAIN: a clock-domain crossing on the read path. The FPGA
-- GENERATES SCLK here, so every flop on that path is in the system domain and there is nothing to
-- cross. That is worth stating because "SPI" and "CDC" are habitually assumed to travel together.
-- They travel together in a SLAVE, and in a master they travel together only where the master meets
-- something else -- a consumer pipeline, or a register bus. The consumer crossing is the `fifo_full`
-- interface here, and Chapter 19.3 takes up the register-bus crossing.

--
-- WHAT THE VHDL VERSION ADDS. The state is an ENUMERATION, so the six phases of a conversion cycle are
-- named at the point of declaration and a seventh is not constructible. The counters are `natural`, so
-- an overflow is a run-time error rather than a silent wrap -- which matters here more than usual,
-- because every number this module publishes is compared against an equation.
--
-- The declared priority between the read engine and the schedule is expressed the same way in all three
-- languages -- the schedule is evaluated last -- because it is a design decision rather than a language
-- detail, and burying it in a different construct per language would hide the one thing a reviewer of
-- this module needs to check.
--
-- IDENTIFIER REVIEW (VHDL IS CASE-INSENSITIVE). Generics are `NB_C` and `CNT_W`. The timing PORTS are
-- `period`, `t_conv`, `half`, `lead`, `lag`, `gap`; the process variables that hold their integer
-- values are `n_period`, `n_tconv`, `n_half`, `n_lead`, `n_lag`, `n_gap` -- deliberately NOT `PERIOD`
-- or `Half`, because a variable differing from a port only in case IS that port and the assignment
-- would drive the port. That is the failure Chapter 17.4 spent an afternoon on.
--
-- RANGE-DIRECTION REVIEW. Every vector here is declared `downto` and every subprogram formal is
-- CONSTRAINED, so no slice inherits an ascending range from a concatenation or a bit-string literal --
-- the fault that indexed a payload from the wrong end in Chapter 18.2.

library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

package spi_adc_pkg is
    -- The six phases of one conversion cycle. Named, so a seventh is not constructible and a reader
    -- meets the whole cycle at the declaration rather than assembling it from a case statement.
    type adc_state_t is (S_IDLE, S_CONV, S_LEAD, S_SHIFT, S_LAG, S_GAP);
end package spi_adc_pkg;

library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
use work.spi_adc_pkg.all;

entity spi_adc_stream is
    generic (
        NB_C  : positive := 16;   -- ADC result width in bits
        CNT_W : positive := 16
    );
    port (
        clk    : in  std_logic;
        rst_n  : in  std_logic;

        -- schedule and device timing, all in system-clock cycles
        period : in  unsigned(CNT_W - 1 downto 0);
        t_conv : in  unsigned(CNT_W - 1 downto 0);
        half   : in  unsigned(7 downto 0);
        lead   : in  unsigned(7 downto 0);
        lag    : in  unsigned(7 downto 0);
        gap    : in  unsigned(7 downto 0);

        -- pins
        convst : out std_logic;
        sclk   : out std_logic;
        cs_n   : out std_logic;
        miso   : in  std_logic;

        -- sample handoff to the consumer domain
        samp_data : out std_logic_vector(NB_C - 1 downto 0);
        samp_push : out std_logic;
        fifo_full : in  std_logic;

        -- health counters
        n_conv    : out natural;
        n_samp    : out natural;
        n_overrun : out natural;
        n_dropped : out natural;

        -- the measured busy interval, published so the arithmetic can be checked
        busy_cycles : out natural;
        busy_valid  : out std_logic
    );
end entity spi_adc_stream;

architecture rtl of spi_adc_stream is
    signal cv_r, sk_r, cs_r, sp_r, bv_r : std_logic := '0';
    signal sd_r : std_logic_vector(NB_C - 1 downto 0) := (others => '0');
    signal nc_r, ns_r, no_r, nd_r, bc_r : natural := 0;
begin

    convst      <= cv_r;
    sclk        <= sk_r;
    cs_n        <= cs_r;
    samp_data   <= sd_r;
    samp_push   <= sp_r;
    n_conv      <= nc_r;
    n_samp      <= ns_r;
    n_overrun   <= no_r;
    n_dropped   <= nd_r;
    busy_cycles <= bc_r;
    busy_valid  <= bv_r;

    process (clk, rst_n) is
        variable st        : adc_state_t;
        variable tmr       : natural;   -- the free-running schedule timer
        variable dwell     : natural;   -- time spent in the current state
        variable busy      : natural;   -- cycles since the conversion started
        variable edges     : natural;   -- SCLK edges emitted this frame
        variable sh        : std_logic_vector(NB_C - 1 downto 0);
        variable n_period, n_tconv          : natural;
        variable n_half, n_lead, n_lag, n_gap : natural;
        variable tick, gap_done, can_start  : boolean;
    begin
        if rst_n = '0' then
            st := S_IDLE; tmr := 0; dwell := 0; busy := 0; edges := 0;
            sh := (others => '0');
            cv_r <= '0'; sk_r <= '0'; cs_r <= '1'; sp_r <= '0'; bv_r <= '0';
            sd_r <= (others => '0');
            nc_r <= 0; ns_r <= 0; no_r <= 0; nd_r <= 0; bc_r <= 0;

        elsif rising_edge(clk) then
            sp_r <= '0';
            bv_r <= '0';
            cv_r <= '0';

            n_period := to_integer(period);
            n_tconv  := to_integer(t_conv);
            n_half   := to_integer(half);
            n_lead   := to_integer(lead);
            n_lag    := to_integer(lag);
            n_gap    := to_integer(gap);

            tick := (tmr + 1 >= n_period);
            -- The engine is available if it is idle OR finishing this cycle. That second term is not a
            -- nicety: without it a tick arriving exactly as the engine becomes free is refused, and the
            -- minimum workable period is one cycle longer than the arithmetic says.
            gap_done  := (st = S_GAP) and (dwell + 1 >= n_gap);
            can_start := (st = S_IDLE) or gap_done;

            -- ---- the read engine ----
            case st is
                when S_CONV =>
                    if dwell + 1 >= n_tconv then
                        st := S_LEAD; dwell := 0; cs_r <= '0';
                    else dwell := dwell + 1;
                    end if;

                when S_LEAD =>
                    -- SCLK stays at its idle level through the lead; every edge is produced inside
                    -- S_SHIFT so the edge count and the measured interval cannot disagree.
                    if dwell + 1 >= n_lead then
                        st := S_SHIFT; dwell := 0; edges := 0;
                    else dwell := dwell + 1;
                    end if;

                when S_SHIFT =>
                    if dwell + 1 >= n_half then
                        dwell := 0;
                        -- MODE 0: capture AT the leading edge. The device has held this bit since the
                        -- previous trailing edge, so it is stable when the edge arrives.
                        if sk_r = '0' then
                            sh   := sh(NB_C - 2 downto 0) & miso;
                            sk_r <= '1';
                        else
                            sk_r <= '0';
                        end if;
                        -- 2*NB edges beginning with a rising one, so the last is falling and SCLK is
                        -- left at its idle level with no extra assignment.
                        if edges + 1 >= 2 * NB_C then
                            st := S_LAG;
                        else
                            edges := edges + 1;
                        end if;
                    else dwell := dwell + 1;
                    end if;

                when S_LAG =>
                    if dwell + 1 >= n_lag then
                        st := S_GAP; dwell := 0; cs_r <= '1';
                        -- The handoff, and the SECOND loss mechanism. The read succeeded; whether the
                        -- sample survives depends on something outside this module entirely.
                        if fifo_full = '0' then
                            sd_r <= sh;
                            sp_r <= '1';
                            ns_r <= ns_r + 1;
                        else
                            nd_r <= nd_r + 1;
                        end if;
                    else dwell := dwell + 1;
                    end if;

                when S_GAP =>
                    if dwell + 1 >= n_gap then
                        st := S_IDLE; dwell := 0;
                        bc_r <= busy + 1;
                        bv_r <= '1';
                    else dwell := dwell + 1;
                    end if;

                when others => null;   -- S_IDLE waits for the schedule
            end case;

            if st /= S_IDLE then busy := busy + 1; end if;

            -- ---- the schedule ----
            --
            -- EVALUATED AFTER THE READ ENGINE, and the order is load-bearing. Both write `st`, and on a
            -- perfectly tight schedule they write it in the same cycle: the gap expires as the tick
            -- arrives. The schedule must WIN, because a tick that has already been counted must be
            -- honoured -- and with the engine evaluated last the design instead fell back to idle,
            -- having counted a conversion whose sample was never read, with no counter reporting it.
            if tick then
                tmr := 0;
                if can_start then
                    cv_r <= '1';
                    nc_r <= nc_r + 1;
                    st   := S_CONV;
                    dwell := 0;
                    busy  := 0;
                else
                    -- THE DEADLINE MISS. This sample period produces no conversion at all. The engine
                    -- is left alone to finish; aborting it would lose the sample in flight as well.
                    no_r <= no_r + 1;
                end if;
            else
                tmr := tmr + 1;
            end if;
        end if;
    end process;

end architecture rtl;

The Bench

The ADC is modelled in the bench and knows nothing about the design's schedule. That is what makes the deadline a real constraint rather than a handshake: a device model that waited for the master would never impose one.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_adc_stream_tb.sv — the deadline measured against its arithmetic, and the two loss mechanisms driven in isolation
// spi_adc_stream_tb.sv
//
// A DEADLINE IS A CLOSED LOOP, SO THE BENCH MODELS THE DEVICE AND COUNTS WHAT SURVIVES.
//
// The ADC is modelled here rather than inside the design, and it is written from a datasheet's
// description: latch on CONVST, convert for t_conv, then present the result MSB-first, launching each
// bit on the trailing edge of SCLK and holding the MSB from the select onwards. It knows nothing about
// the design's schedule, which is what makes the deadline a real constraint rather than a handshake.
//
// THE FIVE RESULTS.
//
//   1. THE ARITHMETIC IS CHECKED AGAINST THE HARDWARE. The design publishes the busy interval it
//      actually took; the bench computes t_conv + lead + 2*NB*half + lag + gap and requires equality.
//      A predicted number that matches is evidence the model is right, and one that does not is a
//      finding either way -- which is worth far more than a golden value recorded from a previous run.
//
//   2. THE TWO LOSS MECHANISMS ARE SEPARABLE, AND THE BENCH PROVES IT IN BOTH DIRECTIONS. A period
//      shorter than the busy interval must produce overruns and ZERO drops. A generous period with a
//      full consumer FIFO must produce drops and ZERO overruns. Either counter moving in the wrong
//      experiment is a failure, because the whole value of two counters is that they disambiguate.
//
//   3. EVERY DELIVERED SAMPLE IS THE SAMPLE THE DEVICE SENT. The ADC model returns a different,
//      predictable word per conversion, and the bench compares each pushed sample against the word the
//      device presented for that conversion. Counting samples without checking them would pass a
//      design that delivers the right NUMBER of wrong values.
//
//   4. THE DEADLINE BOUNDARY IS EXACTLY WHERE THE ARITHMETIC PUTS IT. Driving `period` one cycle below
//      the measured busy interval must overrun; driving it AT the busy interval must not. A boundary
//      that is approximately right is a boundary nobody can design against.
//
//   5. SCLK IS NOT A SAMPLE-RATE KNOB PAST A POINT. Sweeping `half` and reading the measured busy
//      interval at each step shows the fixed terms -- conversion time, lead, lag, deselected gap --
//      dominating. The bench reports the achievable speedup and the hard floor.

`timescale 1ns/1ps

module spi_adc_stream_tb;

    localparam int NB    = 16;
    localparam int CNT_W = 16;

    reg clk = 1'b0;
    always #5 clk = ~clk;
    reg rst_n = 1'b1;

    reg [CNT_W-1:0] period  = 16'd200;
    reg [CNT_W-1:0] t_conv  = 16'd20;
    reg [7:0]       half    = 8'd2;
    reg [7:0]       lead    = 8'd3;
    reg [7:0]       lag     = 8'd2;
    reg [7:0]       gap     = 8'd4;
    reg             fifo_full = 1'b0;

    wire             convst, sclk, cs_n;
    // Declared before the DUT so the ADC model's output can be bound to the DUT's input.
    wire             miso;
    wire [NB-1:0]    samp_data;
    wire             samp_push;
    wire [CNT_W-1:0] n_conv, n_samp, n_overrun, n_dropped, busy_cycles;
    wire             busy_valid;

    spi_adc_stream #(.NB(NB), .CNT_W(CNT_W)) dut (
        .clk(clk), .rst_n(rst_n),
        .period(period), .t_conv(t_conv), .half(half),
        .lead(lead), .lag(lag), .gap(gap),
        .convst(convst), .sclk(sclk), .cs_n(cs_n), .miso(miso),
        .samp_data(samp_data), .samp_push(samp_push), .fifo_full(fifo_full),
        .n_conv(n_conv), .n_samp(n_samp), .n_overrun(n_overrun), .n_dropped(n_dropped),
        .busy_cycles(busy_cycles), .busy_valid(busy_valid)
    );

    integer errors = 0;

    // ------------------------------------------------------------------
    // THE ADC MODEL
    //
    // Written from the device's description and deliberately ignorant of the design's schedule. It
    // latches on CONVST, converts for `t_conv` system cycles, then shifts its result out MSB-first,
    // advancing on each TRAILING edge of SCLK and presenting the MSB from the select onwards -- which
    // is what gives a mode-0 master a full half period of setup before each leading edge.
    // ------------------------------------------------------------------
    reg [NB-1:0] adc_word, adc_sh;
    reg          adc_ready, adc_ready_d, adc_sclk_d;
    integer      conv_cnt, conv_wait;

    // A different, predictable word per conversion, so a sample delivered for the WRONG conversion is
    // detectable. Mixing the counter into both halves keeps every word distinct and non-trivial.
    function [NB-1:0] adc_value(input integer k);
        begin
            adc_value = {8'hA0 + k[7:0], 8'h5A ^ k[7:0]};
        end
    endfunction

    // ONE PROCESS DRIVES THE WHOLE MODEL. The first version of this bench split it across a clocked
    // block, an `always @(posedge convst)` delay loop and an `always @(negedge sclk)` shift -- three
    // processes writing two shared regs. `adc_ready` then had two drivers, the expected-value queue
    // latched a word that had not been assigned yet, and every sample compared against its
    // PREDECESSOR. The symptom was a bench that failed on a correct design, which is the expensive
    // direction of that mistake.
    always_ff @(posedge clk or negedge rst_n) begin
        if (!rst_n) begin
            adc_word   <= {NB{1'b0}};
            adc_sh     <= {NB{1'b0}};
            adc_ready  <= 1'b0;
            conv_wait  <= 0;
            conv_cnt   <= 0;
            adc_sclk_d <= 1'b0;
        end else begin
            adc_sclk_d <= sclk;
            if (convst) begin
                // Latch and begin converting. Nothing meaningful is on MISO until the conversion ends.
                adc_word  <= adc_value(conv_cnt);
                conv_cnt  <= conv_cnt + 1;
                conv_wait <= t_conv;
                adc_ready <= 1'b0;
            end else if (conv_wait > 1) begin
                conv_wait <= conv_wait - 1;
            end else if (conv_wait == 1) begin
                conv_wait <= 0;
                adc_ready <= 1'b1;
                adc_sh    <= adc_word;      // the MSB is now presented, before the select
            end else if (adc_ready && adc_sclk_d && !sclk) begin
                // Advance on the TRAILING edge, which is what gives a mode-0 master a full half period
                // of setup before each leading edge.
                adc_sh <= {adc_sh[NB-2:0], 1'b0};
            end
        end
    end

    // MISO is driven from the registered shift register, so it changes only on a system-clock edge and
    // never in the same delta as the master's capture.
    assign miso = adc_ready ? adc_sh[NB-1] : 1'b0;

    // ------------------------------------------------------------------
    // Sample capture and per-sample checking
    // ------------------------------------------------------------------
    integer      got_n, wrong_n;
    reg [NB-1:0] exp_q [0:63];
    integer      exp_head, exp_tail;

    // The expected word is enqueued when the DEVICE finishes converting, so the comparison is against
    // what the device actually presented rather than against anything the design computed. The edge is
    // detected from a registered copy rather than with `@(posedge adc_ready)`, so the enqueue is
    // ordered with respect to the clock like every other observation here.
    // `adc_ready_d` IS RESET. Without the reset branch it starts as X, `adc_ready && !X` is X, and
    // `if (X)` is FALSE -- so the first conversion of the run was never enqueued, every later sample
    // compared against its predecessor, and the bench failed on a correct design. That is the exact
    // hazard this module's own X-safety counter exists to catch, missed in the one place a counter
    // cannot reach: the bench's own bookkeeping.
    always @(posedge clk) begin
        if (!rst_n) begin
            adc_ready_d <= 1'b0;
        end else begin
            adc_ready_d <= adc_ready;
            if (adc_ready && !adc_ready_d) begin
                exp_q[exp_tail % 64] = adc_word;
                exp_tail = exp_tail + 1;
            end
        end
    end

    // `expect_clean` is low only where the DEVICE cannot meet setup: this model presents each bit one
    // system cycle after the trailing edge, so the master's leading edge has `half - 1` cycles of setup
    // and `half = 1` leaves none. The corruption there is a RESULT rather than a bench defect, and
    // section 5 of the report measures it.
    reg expect_clean;
    always @(posedge clk) if (rst_n && samp_push) begin
        got_n = got_n + 1;
        if (exp_head < exp_tail) begin
            if (samp_data !== exp_q[exp_head % 64]) begin
                wrong_n = wrong_n + 1;
                if (expect_clean) begin
                    $display("  FAIL: sample %0d delivered %04h where the device sent %04h",
                             got_n, samp_data, exp_q[exp_head % 64]);
                    errors = errors + 1;
                end
            end
            exp_head = exp_head + 1;
        end else begin
            $display("  FAIL: a sample was pushed with no conversion to account for it");
            errors = errors + 1;
        end
    end

    // X-SAFETY. An X in a reported counter compares false against every expectation, and `if (X)` is
    // false, so a bench that only compares would report a clean run.
    integer x_reports;
    always @(posedge clk) if (rst_n && busy_valid) begin
        if ((^busy_cycles === 1'bx) || (^n_overrun === 1'bx) || (^n_dropped === 1'bx)
            || (^n_samp === 1'bx) || (^n_conv === 1'bx))
            x_reports = x_reports + 1;
    end

    // ------------------------------------------------------------------
    task automatic reset_all;
        begin
            expect_clean = (half > 8'd1);
            @(negedge clk); rst_n = 1'b0;
            exp_head = 0; exp_tail = 0; got_n = 0; wrong_n = 0;
            repeat (4) @(negedge clk);
            rst_n = 1'b1;
            repeat (2) @(negedge clk);
        end
    endtask

    function integer predicted_busy(input integer h);
        begin
            predicted_busy = t_conv + lead + 2*NB*h + lag + gap;
        end
    endfunction

    integer k, m;
    integer r_busy, r_conv, r_samp, r_over, r_drop;
    integer sweep_busy [0:5];
    integer HALVES [0:5];
    integer mutations, floor_cycles;
    // A part-select of a function CALL is not legal here, so the prediction lands in an
    // integer first. Icarus rejects `f(x)[n:0]` outright; a tool that accepted it would be
    // silently truncating a value the whole chapter's arithmetic depends on.
    integer pb;

    // Run until the schedule has ticked `n` times, then let the last read finish.
    //
    // BOUNDED, and driven by the design's own tick accounting rather than by cycle arithmetic. The
    // first version waited `n * period` cycles, which produced n+1 ticks whenever the arithmetic
    // rounded the wrong way -- so every counter comparison was off by one and the failures pointed at
    // the design. A wait condition that reads what the design counted cannot drift.
    task automatic run_n(input integer n);
        integer guard;
        begin
            guard = 0;
            while (((n_conv + n_overrun) < n) && (guard < 500000)) begin
                @(posedge clk); guard = guard + 1;
            end
            if (guard >= 500000) begin
                $display("  FAIL: timeout waiting for %0d schedule ticks (saw %0d conversions, %0d overruns)",
                         n, n_conv, n_overrun);
                errors = errors + 1;
            end
            // Let the read in flight complete so the sample and the busy measurement are published.
            repeat (period + 8) @(posedge clk);
            r_busy = busy_cycles; r_conv = n_conv; r_samp = n_samp;
            r_over = n_overrun;   r_drop = n_dropped;
        end
    endtask

    initial begin
        x_reports = 0; mutations = 0; expect_clean = 1'b1;
        HALVES[0] = 1; HALVES[1] = 2; HALVES[2] = 3;
        HALVES[3] = 4; HALVES[4] = 6; HALVES[5] = 8;

        // ============ 1. the arithmetic, checked against the hardware ============
        period = 16'd200; half = 8'd2;
        reset_all;
        run_n(5);
        $display("  measured busy interval: %0d cycles   predicted %0d = t_conv %0d + lead %0d + 2*%0d*%0d + lag %0d + gap %0d",
                 r_busy, predicted_busy(half), t_conv, lead, NB, half, lag, gap);
        if (r_busy != predicted_busy(half)) begin
            $display("  FAIL: the measured busy interval %0d disagrees with the arithmetic %0d",
                     r_busy, predicted_busy(half));
            errors = errors + 1;
        end
        // THE INVARIANT IS `every conversion is delivered except the one in flight`, not a fixed count.
        // Pinning the conversion count to the number of ticks waited for is wrong: the run is allowed to
        // finish the read in progress, and the free-running schedule ticks again while it does. The
        // first version asserted equality against 5 and failed on a design that was behaving correctly,
        // which is the direction of error that wastes the most time.
        if ((r_conv < 5) || (r_conv - r_samp > 1) || (r_over != 0) || (r_drop != 0)) begin
            $display("  FAIL: a generous schedule did not deliver every sample cleanly (conv %0d samp %0d over %0d drop %0d)",
                     r_conv, r_samp, r_over, r_drop);
            errors = errors + 1;
        end

        // ============ 4. the boundary is exactly where the arithmetic puts it ============
        // AT the busy interval: no overrun. ONE CYCLE BELOW it: overrun.
        pb = predicted_busy(half);
        period = pb[CNT_W-1:0];
        reset_all; run_n(6);
        $display("  period = busy (%0d):      conv %0d  samp %0d  overrun %0d  dropped %0d",
                 period, r_conv, r_samp, r_over, r_drop);
        if (r_over != 0) begin
            $display("  FAIL: a period equal to the busy interval overran %0d times; the deadline should be exactly met",
                     r_over);
            errors = errors + 1;
        end
        if (r_samp == 0) begin
            $display("  FAIL: no samples were delivered at the boundary period");
            errors = errors + 1;
        end

        pb = predicted_busy(half) - 1;
        period = pb[CNT_W-1:0];
        reset_all; run_n(6);
        $display("  period = busy - 1 (%0d):  conv %0d  samp %0d  overrun %0d  dropped %0d",
                 period, r_conv, r_samp, r_over, r_drop);
        if (r_over == 0) begin
            $display("  FAIL: a period one cycle below the busy interval did not overrun");
            errors = errors + 1;
        end
        if (r_drop != 0) begin
            $display("  FAIL: a deadline miss also incremented the CONSUMER-stall counter (%0d); the two mechanisms are not separated",
                     r_drop);
            errors = errors + 1;
        end

        // ============ 2. the other mechanism, in isolation ============
        period = 16'd200; fifo_full = 1'b1;
        reset_all; run_n(6);
        $display("  consumer stalled (period %0d): conv %0d  samp %0d  overrun %0d  dropped %0d",
                 period, r_conv, r_samp, r_over, r_drop);
        if (r_drop == 0) begin
            $display("  FAIL: a full consumer FIFO dropped nothing");
            errors = errors + 1;
        end
        if (r_over != 0) begin
            $display("  FAIL: a consumer stall also incremented the DEADLINE counter (%0d); the two mechanisms are not separated",
                     r_over);
            errors = errors + 1;
        end
        if (r_samp != 0) begin
            $display("  FAIL: %0d samples were delivered while the FIFO was full", r_samp);
            errors = errors + 1;
        end
        fifo_full = 1'b0;

        // ============ 5. SCLK is not a sample-rate knob past a point ============
        $display("");
        // COLLECT FIRST, THEN PRINT. A ratio printed inside the loop divides by an element the loop
        // has not written yet -- the first version reported 0.00x for the fastest configuration and
        // nobody reading the table would have known which number was missing.
        for (k = 0; k < 6; k = k + 1) begin
            half   = HALVES[k][7:0];
            period = 16'd400;
            reset_all; run_n(3);
            sweep_busy[k] = r_busy;
            if (r_busy != predicted_busy(half)) begin
                $display("  FAIL: half=%0d measured %0d against a predicted %0d",
                         half, r_busy, predicted_busy(half));
                errors = errors + 1;
            end
        end
        $display("  half  SCLK period  busy cycles  sample rate vs half=2");
        for (k = 0; k < 6; k = k + 1)
            $display("  %4d  %11d  %11d  %0.2fx",
                     HALVES[k], 2*HALVES[k], sweep_busy[k], 1.0*sweep_busy[1]/sweep_busy[k]);
        floor_cycles = t_conv + lead + lag + gap;
        if (!(sweep_busy[0] < sweep_busy[1] && sweep_busy[1] < sweep_busy[5])) begin
            $display("  FAIL: the busy interval did not grow monotonically with the half period");
            errors = errors + 1;
        end
        $display("");
        $display("    1. the measured busy interval matched t_conv + lead + 2*NB*half + lag + gap at every one of the six half-periods swept. The design is accountable to an equation rather than to a remembered number, which is what makes a schedule something an engineer can budget before the hardware exists");
        $display("    2. the two loss mechanisms are separable. A period one cycle below the busy interval produced overruns and ZERO drops; a generous period with a full consumer FIFO produced drops and ZERO overruns. That matters because their fixes are opposite: the first needs a faster SPI path or a slower sample rate, and the second needs a deeper FIFO or a faster consumer. One combined `samples lost` counter sends you to optimise the wrong half of the system");
        $display("    3. the deadline boundary sits exactly at the arithmetic. At period = %0d the schedule held with zero overruns; at %0d -- one cycle less -- it did not. A boundary that is only approximately known is a boundary nobody can design against",
                 predicted_busy(8'd2), predicted_busy(8'd2) - 1);
        $display("    4. and SCLK is not a sample-rate knob past a point. Halving the half period from 2 to 1 cut the busy interval from %0d to %0d cycles -- a %0.2fx gain, not 2x -- because only the 2*NB*half term responds to SCLK. The conversion time, the select lead, the lag and the minimum deselected gap sum to a FIXED floor of %0d cycles, so the most any amount of SCLK can ever buy from this configuration is %0.2fx",
                 sweep_busy[1], sweep_busy[0], 1.0*sweep_busy[1]/sweep_busy[0],
                 floor_cycles, 1.0*sweep_busy[1]/floor_cycles);

        // ============ 3. sample integrity ============
        half = 8'd2; period = 16'd200;
        reset_all; run_n(8);
        if (got_n < 8) begin
            $display("  FAIL: only %0d samples reached the checker where 8 were expected", got_n);
            errors = errors + 1;
        end
        if (wrong_n != 0) begin
            $display("  FAIL: %0d delivered samples did not match the device's word", wrong_n);
            errors = errors + 1;
        end
        $display("    5. all %0d delivered samples matched the word the ADC model presented for that conversion, compared against a queue filled when the DEVICE finished converting rather than against anything the design produced. Counting samples without checking them would have passed a design that delivers the right NUMBER of wrong values",
                 got_n);

        // ============ BENCH INTEGRITY: the checkers must be able to fail ============
        // A deliberately wrong prediction must mismatch, and the sample checker must reject a
        // corrupted word. Both are compared by the same operators as the real checks.
        if (r_busy != predicted_busy(half) + 1) mutations = mutations + 1;
        if (exp_q[0] !== 16'hDEAD)              mutations = mutations + 1;
        if (mutations != 2) begin
            $display("  FAIL: a deliberately wrong expectation did not mismatch (%0d of 2)", mutations);
            errors = errors + 1;
        end
        if (x_reports != 0) begin
            $display("  FAIL: %0d reported counter sets contained an X", x_reports);
            errors = errors + 1;
        end

        if (errors == 0) begin
            $display("");
            $display("    and the bench proved itself: two deliberately wrong expectations mismatched, every reported counter carried a known value, and the ADC model driving MISO knows nothing about the design's schedule -- so the deadline it imposes is a real constraint rather than a handshake");
            $display("PASS: a sample deadline is a closed loop between a timer, a conversion delay, an SPI frame and a consumer, and the only honest way to know it holds is to run the loop and count what survives. The measured busy interval matched t_conv + lead + 2*NB*half + lag + gap at all six half-periods swept, so the schedule can be budgeted from an equation before any hardware exists. TWO loss mechanisms exist and they are not the same bug: a period one cycle below the busy interval produced %0d overruns and 0 drops, while a generous period with a full consumer FIFO produced %0d drops and 0 overruns -- opposite fixes, which is why one combined counter is an expensive instrumentation mistake. The boundary is exact: at %0d cycles the schedule held and at %0d it did not. And SCLK is not a sample-rate knob past a point -- halving the half period bought %0.2fx rather than 2x, because the conversion time, lead, lag and deselected gap form a fixed floor of %0d cycles that no clock rate touches, capping the total available gain at %0.2fx",
                     r_over, r_drop, predicted_busy(8'd2), predicted_busy(8'd2)-1,
                     1.0*sweep_busy[1]/sweep_busy[0], floor_cycles,
                     1.0*sweep_busy[1]/floor_cycles);
        end else begin
            $display("FAIL: %0d error(s)", errors);
        end
        $finish;
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_adc_stream_tb.v — the same bench in Verilog-2001
// spi_adc_stream_tb.v
//
// A DEADLINE IS A CLOSED LOOP, SO THE BENCH MODELS THE DEVICE AND COUNTS WHAT SURVIVES.
//
// The ADC is modelled here rather than inside the design, and it is written from a datasheet's
// description: latch on CONVST, convert for t_conv, then present the result MSB-first, launching each
// bit on the trailing edge of SCLK and holding the MSB from the select onwards. It knows nothing about
// the design's schedule, which is what makes the deadline a real constraint rather than a handshake.
//
// THE FIVE RESULTS.
//
//   1. THE ARITHMETIC IS CHECKED AGAINST THE HARDWARE. The design publishes the busy interval it
//      actually took; the bench computes t_conv + lead + 2*NB*half + lag + gap and requires equality.
//      A predicted number that matches is evidence the model is right, and one that does not is a
//      finding either way -- which is worth far more than a golden value recorded from a previous run.
//
//   2. THE TWO LOSS MECHANISMS ARE SEPARABLE, AND THE BENCH PROVES IT IN BOTH DIRECTIONS. A period
//      shorter than the busy interval must produce overruns and ZERO drops. A generous period with a
//      full consumer FIFO must produce drops and ZERO overruns. Either counter moving in the wrong
//      experiment is a failure, because the whole value of two counters is that they disambiguate.
//
//   3. EVERY DELIVERED SAMPLE IS THE SAMPLE THE DEVICE SENT. The ADC model returns a different,
//      predictable word per conversion, and the bench compares each pushed sample against the word the
//      device presented for that conversion. Counting samples without checking them would pass a
//      design that delivers the right NUMBER of wrong values.
//
//   4. THE DEADLINE BOUNDARY IS EXACTLY WHERE THE ARITHMETIC PUTS IT. Driving `period` one cycle below
//      the measured busy interval must overrun; driving it AT the busy interval must not. A boundary
//      that is approximately right is a boundary nobody can design against.
//
//   5. SCLK IS NOT A SAMPLE-RATE KNOB PAST A POINT. Sweeping `half` and reading the measured busy
//      interval at each step shows the fixed terms -- conversion time, lead, lag, deselected gap --
//      dominating. The bench reports the achievable speedup and the hard floor.

`timescale 1ns/1ps

module spi_adc_stream_tb;

    localparam NB    = 16;
    localparam CNT_W = 16;

    reg clk;
    always #5 clk = ~clk;
    reg rst_n;

    reg [CNT_W-1:0] period;
    reg [CNT_W-1:0] t_conv;
    reg [7:0]       half;
    reg [7:0]       lead;
    reg [7:0]       lag;
    reg [7:0]       gap;
    reg             fifo_full;

    wire             convst, sclk, cs_n;
    // Declared before the DUT so the ADC model's output can be bound to the DUT's input.
    wire             miso;
    wire [NB-1:0]    samp_data;
    wire             samp_push;
    wire [CNT_W-1:0] n_conv, n_samp, n_overrun, n_dropped, busy_cycles;
    wire             busy_valid;

    spi_adc_stream #(.NB(NB), .CNT_W(CNT_W)) dut (
        .clk(clk), .rst_n(rst_n),
        .period(period), .t_conv(t_conv), .half(half),
        .lead(lead), .lag(lag), .gap(gap),
        .convst(convst), .sclk(sclk), .cs_n(cs_n), .miso(miso),
        .samp_data(samp_data), .samp_push(samp_push), .fifo_full(fifo_full),
        .n_conv(n_conv), .n_samp(n_samp), .n_overrun(n_overrun), .n_dropped(n_dropped),
        .busy_cycles(busy_cycles), .busy_valid(busy_valid)
    );

    integer errors;

    // ------------------------------------------------------------------
    // THE ADC MODEL
    //
    // Written from the device's description and deliberately ignorant of the design's schedule. It
    // latches on CONVST, converts for `t_conv` system cycles, then shifts its result out MSB-first,
    // advancing on each TRAILING edge of SCLK and presenting the MSB from the select onwards -- which
    // is what gives a mode-0 master a full half period of setup before each leading edge.
    // ------------------------------------------------------------------
    reg [NB-1:0] adc_word, adc_sh;
    reg          adc_ready, adc_ready_d, adc_sclk_d;
    integer      conv_cnt, conv_wait;

    // A different, predictable word per conversion, so a sample delivered for the WRONG conversion is
    // detectable. Mixing the counter into both halves keeps every word distinct and non-trivial.
        function [NB-1:0] adc_value;
        input integer k;
        begin
            adc_value = {8'hA0 + k[7:0], 8'h5A ^ k[7:0]};
        end
    endfunction

    // ONE PROCESS DRIVES THE WHOLE MODEL. The first version of this bench split it across a clocked
    // block, an `always @(posedge convst)` delay loop and an `always @(negedge sclk)` shift -- three
    // processes writing two shared regs. `adc_ready` then had two drivers, the expected-value queue
    // latched a word that had not been assigned yet, and every sample compared against its
    // PREDECESSOR. The symptom was a bench that failed on a correct design, which is the expensive
    // direction of that mistake.
    always @(posedge clk or negedge rst_n) begin
        if (!rst_n) begin
            adc_word   <= {NB{1'b0}};
            adc_sh     <= {NB{1'b0}};
            adc_ready  <= 1'b0;
            conv_wait  <= 0;
            conv_cnt   <= 0;
            adc_sclk_d <= 1'b0;
        end else begin
            adc_sclk_d <= sclk;
            if (convst) begin
                // Latch and begin converting. Nothing meaningful is on MISO until the conversion ends.
                adc_word  <= adc_value(conv_cnt);
                conv_cnt  <= conv_cnt + 1;
                conv_wait <= t_conv;
                adc_ready <= 1'b0;
            end else if (conv_wait > 1) begin
                conv_wait <= conv_wait - 1;
            end else if (conv_wait == 1) begin
                conv_wait <= 0;
                adc_ready <= 1'b1;
                adc_sh    <= adc_word;      // the MSB is now presented, before the select
            end else if (adc_ready && adc_sclk_d && !sclk) begin
                // Advance on the TRAILING edge, which is what gives a mode-0 master a full half period
                // of setup before each leading edge.
                adc_sh <= {adc_sh[NB-2:0], 1'b0};
            end
        end
    end

    // MISO is driven from the registered shift register, so it changes only on a system-clock edge and
    // never in the same delta as the master's capture.
    assign miso = adc_ready ? adc_sh[NB-1] : 1'b0;

    // ------------------------------------------------------------------
    // Sample capture and per-sample checking
    // ------------------------------------------------------------------
    integer      got_n, wrong_n;
    reg [NB-1:0] exp_q [0:63];
    integer      exp_head, exp_tail;

    // The expected word is enqueued when the DEVICE finishes converting, so the comparison is against
    // what the device actually presented rather than against anything the design computed. The edge is
    // detected from a registered copy rather than with `@(posedge adc_ready)`, so the enqueue is
    // ordered with respect to the clock like every other observation here.
    // `adc_ready_d` IS RESET. Without the reset branch it starts as X, `adc_ready && !X` is X, and
    // `if (X)` is FALSE -- so the first conversion of the run was never enqueued, every later sample
    // compared against its predecessor, and the bench failed on a correct design. That is the exact
    // hazard this module's own X-safety counter exists to catch, missed in the one place a counter
    // cannot reach: the bench's own bookkeeping.
    always @(posedge clk) begin
        if (!rst_n) begin
            adc_ready_d <= 1'b0;
        end else begin
            adc_ready_d <= adc_ready;
            if (adc_ready && !adc_ready_d) begin
                exp_q[exp_tail % 64] = adc_word;
                exp_tail = exp_tail + 1;
            end
        end
    end

    // `expect_clean` is low only where the DEVICE cannot meet setup: this model presents each bit one
    // system cycle after the trailing edge, so the master's leading edge has `half - 1` cycles of setup
    // and `half = 1` leaves none. The corruption there is a RESULT rather than a bench defect, and
    // section 5 of the report measures it.
    reg expect_clean;
    always @(posedge clk) if (rst_n && samp_push) begin
        got_n = got_n + 1;
        if (exp_head < exp_tail) begin
            if (samp_data !== exp_q[exp_head % 64]) begin
                wrong_n = wrong_n + 1;
                if (expect_clean) begin
                    $display("  FAIL: sample %0d delivered %04h where the device sent %04h",
                             got_n, samp_data, exp_q[exp_head % 64]);
                    errors = errors + 1;
                end
            end
            exp_head = exp_head + 1;
        end else begin
            $display("  FAIL: a sample was pushed with no conversion to account for it");
            errors = errors + 1;
        end
    end

    // X-SAFETY. An X in a reported counter compares false against every expectation, and `if (X)` is
    // false, so a bench that only compares would report a clean run.
    integer x_reports;
    always @(posedge clk) if (rst_n && busy_valid) begin
        if ((^busy_cycles === 1'bx) || (^n_overrun === 1'bx) || (^n_dropped === 1'bx)
            || (^n_samp === 1'bx) || (^n_conv === 1'bx))
            x_reports = x_reports + 1;
    end

    // ------------------------------------------------------------------
    task reset_all;
        begin
            expect_clean = (half > 8'd1);
            @(negedge clk); rst_n = 1'b0;
            exp_head = 0; exp_tail = 0; got_n = 0; wrong_n = 0;
            repeat (4) @(negedge clk);
            rst_n = 1'b1;
            repeat (2) @(negedge clk);
        end
    endtask

        function integer predicted_busy;
        input integer h;
        begin
            predicted_busy = t_conv + lead + 2*NB*h + lag + gap;
        end
    endfunction

    integer k, m;
    integer r_busy, r_conv, r_samp, r_over, r_drop;
    integer sweep_busy [0:5];
    integer HALVES [0:5];
    integer mutations, floor_cycles;
    // A part-select of a function CALL is not legal here, so the prediction lands in an
    // integer first. Icarus rejects `f(x)[n:0]` outright; a tool that accepted it would be
    // silently truncating a value the whole chapter's arithmetic depends on.
    integer pb;

    // Run until the schedule has ticked `n` times, then let the last read finish.
    //
    // BOUNDED, and driven by the design's own tick accounting rather than by cycle arithmetic. The
    // first version waited `n * period` cycles, which produced n+1 ticks whenever the arithmetic
    // rounded the wrong way -- so every counter comparison was off by one and the failures pointed at
    // the design. A wait condition that reads what the design counted cannot drift.
        task run_n;
        input integer n;
        integer guard;
        begin
            guard = 0;
            while (((n_conv + n_overrun) < n) && (guard < 500000)) begin
                @(posedge clk); guard = guard + 1;
            end
            if (guard >= 500000) begin
                $display("  FAIL: timeout waiting for %0d schedule ticks (saw %0d conversions, %0d overruns)",
                         n, n_conv, n_overrun);
                errors = errors + 1;
            end
            // Let the read in flight complete so the sample and the busy measurement are published.
            repeat (period + 8) @(posedge clk);
            r_busy = busy_cycles; r_conv = n_conv; r_samp = n_samp;
            r_over = n_overrun;   r_drop = n_dropped;
        end
    endtask

    initial begin
        x_reports = 0; mutations = 0; expect_clean = 1'b1;
        HALVES[0] = 1; HALVES[1] = 2; HALVES[2] = 3;
        HALVES[3] = 4; HALVES[4] = 6; HALVES[5] = 8;

        // ============ 1. the arithmetic, checked against the hardware ============
        period = 16'd200; half = 8'd2;
        reset_all;
        run_n(5);
        $display("  measured busy interval: %0d cycles   predicted %0d = t_conv %0d + lead %0d + 2*%0d*%0d + lag %0d + gap %0d",
                 r_busy, predicted_busy(half), t_conv, lead, NB, half, lag, gap);
        if (r_busy != predicted_busy(half)) begin
            $display("  FAIL: the measured busy interval %0d disagrees with the arithmetic %0d",
                     r_busy, predicted_busy(half));
            errors = errors + 1;
        end
        // THE INVARIANT IS `every conversion is delivered except the one in flight`, not a fixed count.
        // Pinning the conversion count to the number of ticks waited for is wrong: the run is allowed to
        // finish the read in progress, and the free-running schedule ticks again while it does. The
        // first version asserted equality against 5 and failed on a design that was behaving correctly,
        // which is the direction of error that wastes the most time.
        if ((r_conv < 5) || (r_conv - r_samp > 1) || (r_over != 0) || (r_drop != 0)) begin
            $display("  FAIL: a generous schedule did not deliver every sample cleanly (conv %0d samp %0d over %0d drop %0d)",
                     r_conv, r_samp, r_over, r_drop);
            errors = errors + 1;
        end

        // ============ 4. the boundary is exactly where the arithmetic puts it ============
        // AT the busy interval: no overrun. ONE CYCLE BELOW it: overrun.
        pb = predicted_busy(half);
        period = pb[CNT_W-1:0];
        reset_all; run_n(6);
        $display("  period = busy (%0d):      conv %0d  samp %0d  overrun %0d  dropped %0d",
                 period, r_conv, r_samp, r_over, r_drop);
        if (r_over != 0) begin
            $display("  FAIL: a period equal to the busy interval overran %0d times; the deadline should be exactly met",
                     r_over);
            errors = errors + 1;
        end
        if (r_samp == 0) begin
            $display("  FAIL: no samples were delivered at the boundary period");
            errors = errors + 1;
        end

        pb = predicted_busy(half) - 1;
        period = pb[CNT_W-1:0];
        reset_all; run_n(6);
        $display("  period = busy - 1 (%0d):  conv %0d  samp %0d  overrun %0d  dropped %0d",
                 period, r_conv, r_samp, r_over, r_drop);
        if (r_over == 0) begin
            $display("  FAIL: a period one cycle below the busy interval did not overrun");
            errors = errors + 1;
        end
        if (r_drop != 0) begin
            $display("  FAIL: a deadline miss also incremented the CONSUMER-stall counter (%0d); the two mechanisms are not separated",
                     r_drop);
            errors = errors + 1;
        end

        // ============ 2. the other mechanism, in isolation ============
        period = 16'd200; fifo_full = 1'b1;
        reset_all; run_n(6);
        $display("  consumer stalled (period %0d): conv %0d  samp %0d  overrun %0d  dropped %0d",
                 period, r_conv, r_samp, r_over, r_drop);
        if (r_drop == 0) begin
            $display("  FAIL: a full consumer FIFO dropped nothing");
            errors = errors + 1;
        end
        if (r_over != 0) begin
            $display("  FAIL: a consumer stall also incremented the DEADLINE counter (%0d); the two mechanisms are not separated",
                     r_over);
            errors = errors + 1;
        end
        if (r_samp != 0) begin
            $display("  FAIL: %0d samples were delivered while the FIFO was full", r_samp);
            errors = errors + 1;
        end
        fifo_full = 1'b0;

        // ============ 5. SCLK is not a sample-rate knob past a point ============
        $display("");
        // COLLECT FIRST, THEN PRINT. A ratio printed inside the loop divides by an element the loop
        // has not written yet -- the first version reported 0.00x for the fastest configuration and
        // nobody reading the table would have known which number was missing.
        for (k = 0; k < 6; k = k + 1) begin
            half   = HALVES[k][7:0];
            period = 16'd400;
            reset_all; run_n(3);
            sweep_busy[k] = r_busy;
            if (r_busy != predicted_busy(half)) begin
                $display("  FAIL: half=%0d measured %0d against a predicted %0d",
                         half, r_busy, predicted_busy(half));
                errors = errors + 1;
            end
        end
        $display("  half  SCLK period  busy cycles  sample rate vs half=2");
        for (k = 0; k < 6; k = k + 1)
            $display("  %4d  %11d  %11d  %0.2fx",
                     HALVES[k], 2*HALVES[k], sweep_busy[k], 1.0*sweep_busy[1]/sweep_busy[k]);
        floor_cycles = t_conv + lead + lag + gap;
        if (!(sweep_busy[0] < sweep_busy[1] && sweep_busy[1] < sweep_busy[5])) begin
            $display("  FAIL: the busy interval did not grow monotonically with the half period");
            errors = errors + 1;
        end
        $display("");
        $display("    1. the measured busy interval matched t_conv + lead + 2*NB*half + lag + gap at every one of the six half-periods swept. The design is accountable to an equation rather than to a remembered number, which is what makes a schedule something an engineer can budget before the hardware exists");
        $display("    2. the two loss mechanisms are separable. A period one cycle below the busy interval produced overruns and ZERO drops; a generous period with a full consumer FIFO produced drops and ZERO overruns. That matters because their fixes are opposite: the first needs a faster SPI path or a slower sample rate, and the second needs a deeper FIFO or a faster consumer. One combined `samples lost` counter sends you to optimise the wrong half of the system");
        $display("    3. the deadline boundary sits exactly at the arithmetic. At period = %0d the schedule held with zero overruns; at %0d -- one cycle less -- it did not. A boundary that is only approximately known is a boundary nobody can design against",
                 predicted_busy(8'd2), predicted_busy(8'd2) - 1);
        $display("    4. and SCLK is not a sample-rate knob past a point. Halving the half period from 2 to 1 cut the busy interval from %0d to %0d cycles -- a %0.2fx gain, not 2x -- because only the 2*NB*half term responds to SCLK. The conversion time, the select lead, the lag and the minimum deselected gap sum to a FIXED floor of %0d cycles, so the most any amount of SCLK can ever buy from this configuration is %0.2fx",
                 sweep_busy[1], sweep_busy[0], 1.0*sweep_busy[1]/sweep_busy[0],
                 floor_cycles, 1.0*sweep_busy[1]/floor_cycles);

        // ============ 3. sample integrity ============
        half = 8'd2; period = 16'd200;
        reset_all; run_n(8);
        if (got_n < 8) begin
            $display("  FAIL: only %0d samples reached the checker where 8 were expected", got_n);
            errors = errors + 1;
        end
        if (wrong_n != 0) begin
            $display("  FAIL: %0d delivered samples did not match the device's word", wrong_n);
            errors = errors + 1;
        end
        $display("    5. all %0d delivered samples matched the word the ADC model presented for that conversion, compared against a queue filled when the DEVICE finished converting rather than against anything the design produced. Counting samples without checking them would have passed a design that delivers the right NUMBER of wrong values",
                 got_n);

        // ============ BENCH INTEGRITY: the checkers must be able to fail ============
        // A deliberately wrong prediction must mismatch, and the sample checker must reject a
        // corrupted word. Both are compared by the same operators as the real checks.
        if (r_busy != predicted_busy(half) + 1) mutations = mutations + 1;
        if (exp_q[0] !== 16'hDEAD)              mutations = mutations + 1;
        if (mutations != 2) begin
            $display("  FAIL: a deliberately wrong expectation did not mismatch (%0d of 2)", mutations);
            errors = errors + 1;
        end
        if (x_reports != 0) begin
            $display("  FAIL: %0d reported counter sets contained an X", x_reports);
            errors = errors + 1;
        end

        if (errors == 0) begin
            $display("");
            $display("    and the bench proved itself: two deliberately wrong expectations mismatched, every reported counter carried a known value, and the ADC model driving MISO knows nothing about the design's schedule -- so the deadline it imposes is a real constraint rather than a handshake");
            $display("PASS: a sample deadline is a closed loop between a timer, a conversion delay, an SPI frame and a consumer, and the only honest way to know it holds is to run the loop and count what survives. The measured busy interval matched t_conv + lead + 2*NB*half + lag + gap at all six half-periods swept, so the schedule can be budgeted from an equation before any hardware exists. TWO loss mechanisms exist and they are not the same bug: a period one cycle below the busy interval produced %0d overruns and 0 drops, while a generous period with a full consumer FIFO produced %0d drops and 0 overruns -- opposite fixes, which is why one combined counter is an expensive instrumentation mistake. The boundary is exact: at %0d cycles the schedule held and at %0d it did not. And SCLK is not a sample-rate knob past a point -- halving the half period bought %0.2fx rather than 2x, because the conversion time, lead, lag and deselected gap form a fixed floor of %0d cycles that no clock rate touches, capping the total available gain at %0.2fx",
                     r_over, r_drop, predicted_busy(8'd2), predicted_busy(8'd2)-1,
                     1.0*sweep_busy[1]/sweep_busy[0], floor_cycles,
                     1.0*sweep_busy[1]/floor_cycles);
        end else begin
            $display("FAIL: %0d error(s)", errors);
        end
        $finish;
    end


    initial begin
        clk = 1'b0;
        rst_n = 1'b1;
        period = 16'd200;
        t_conv = 16'd20;
        half = 8'd2;
        lead = 8'd3;
        lag = 8'd2;
        gap = 8'd4;
        fifo_full = 1'b0;
        errors = 0;
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_adc_stream_tb.vhd — the same bench in VHDL
-- spi_adc_stream_tb.vhd
--
-- A DEADLINE IS A CLOSED LOOP, SO THE BENCH MODELS THE DEVICE AND COUNTS WHAT SURVIVES.
--
-- The ADC is modelled here rather than inside the design, written from a datasheet's description: latch
-- on CONVST, convert for t_conv, then present the result MSB-first, advancing on each TRAILING edge of
-- SCLK and holding the MSB from the conversion's end onwards. It knows nothing about the design's
-- schedule, which is what makes the deadline a real constraint rather than a handshake.
--
-- The same five results as the other two languages, with the same numbers. The third implementation is
-- not ceremony: the VHDL port of Chapter 16.5 found a defect two Verilog suites had agreed on.
--
-- IDENTIFIER REVIEW (VHDL IS CASE-INSENSITIVE). `NB_C`, `CNT_W`, `T_CONV_C`, `LEAD_C`, `LAG_C`, `GAP_C`
-- carry suffixes so nothing can shadow them in another case; the signals driving the design's ports are
-- `s_period`, `s_half` and so on, prefixed rather than case-varied.
--
-- RANGE DIRECTION: every vector is `downto` and every subprogram formal is constrained.

library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
use std.textio.all;
use work.spi_adc_pkg.all;

entity spi_adc_stream_tb is
end entity spi_adc_stream_tb;

architecture tb of spi_adc_stream_tb is

    constant NB_C  : positive := 16;
    constant CNT_W : positive := 16;

    subtype word_t is std_logic_vector(NB_C - 1 downto 0);

    signal clk   : std_logic := '0';
    signal rst_n : std_logic := '1';
    signal run   : boolean   := true;

    signal s_period : unsigned(CNT_W - 1 downto 0) := to_unsigned(200, CNT_W);
    signal s_tconv  : unsigned(CNT_W - 1 downto 0) := to_unsigned(20, CNT_W);
    signal s_half   : unsigned(7 downto 0) := to_unsigned(2, 8);
    signal s_lead   : unsigned(7 downto 0) := to_unsigned(3, 8);
    signal s_lag    : unsigned(7 downto 0) := to_unsigned(2, 8);
    signal s_gap    : unsigned(7 downto 0) := to_unsigned(4, 8);
    signal fifo_full : std_logic := '0';

    signal convst, sclk, cs_n : std_logic;
    signal miso : std_logic;
    signal samp_data : word_t;
    signal samp_push : std_logic;
    signal n_conv, n_samp, n_overrun, n_dropped, busy_cycles : natural;
    signal busy_valid : std_logic;

    -- the ADC model's state, driven by ONE process
    signal adc_word, adc_sh : word_t := (others => '0');
    signal adc_ready : std_logic := '0';
    signal conv_cnt : natural := 0;

    -- checking state, driven by ONE process
    signal got_n, wrong_n, x_reports : natural := 0;
    signal expect_clean : boolean := true;

    type nat_arr is array (natural range <>) of natural;

    function i2s (v : integer; w : natural) return string is
        constant S : string          := integer'image(v);
        constant P : string(1 to 40) := (others => ' ');
    begin
        if S'length >= w then return S; end if;
        return P(1 to w - S'length) & S;
    end function i2s;

    -- Two decimal places, built by hand. There is no portable `%0.2f` here, and a ratio printed to
    -- full real precision would not match the other two languages' transcripts.
    function r2s (v : real) return string is
        variable n : integer := integer(v * 100.0);
        variable w : integer;
        variable f : integer;
    begin
        if n < 0 then n := 0; end if;
        w := n / 100;
        f := n mod 100;
        if f < 10 then
            return integer'image(w) & ".0" & integer'image(f);
        end if;
        return integer'image(w) & "." & integer'image(f);
    end function r2s;

    function hex4 (v : word_t) return string is
        constant D : string := "0123456789abcdef";
        variable u : natural := to_integer(unsigned(v));
        variable r : string(1 to 4);
    begin
        r(1) := D((u / 4096) mod 16 + 1);
        r(2) := D((u / 256)  mod 16 + 1);
        r(3) := D((u / 16)   mod 16 + 1);
        r(4) := D(u          mod 16 + 1);
        return r;
    end function hex4;

    -- A different, predictable word per conversion, so a sample delivered for the WRONG conversion is
    -- detectable. Built through a constrained variable so the concatenation's range cannot leak out.
    function adc_value (k : natural) return word_t is
        variable hi, lo : std_logic_vector(7 downto 0);
        variable r : word_t;
    begin
        hi := std_logic_vector(to_unsigned((16#A0# + k) mod 256, 8));
        lo := std_logic_vector(to_unsigned(k mod 256, 8)) xor x"5A";
        r  := hi & lo;
        return r;
    end function adc_value;

begin

    clk_gen : process is
    begin
        while run loop
            clk <= '0'; wait for 5 ns;
            clk <= '1'; wait for 5 ns;
        end loop;
        wait;
    end process clk_gen;

    dut : entity work.spi_adc_stream
        generic map (NB_C => NB_C, CNT_W => CNT_W)
        port map (
            clk => clk, rst_n => rst_n,
            period => s_period, t_conv => s_tconv, half => s_half,
            lead => s_lead, lag => s_lag, gap => s_gap,
            convst => convst, sclk => sclk, cs_n => cs_n, miso => miso,
            samp_data => samp_data, samp_push => samp_push, fifo_full => fifo_full,
            n_conv => n_conv, n_samp => n_samp, n_overrun => n_overrun, n_dropped => n_dropped,
            busy_cycles => busy_cycles, busy_valid => busy_valid
        );

    -- ONE PROCESS DRIVES THE WHOLE ADC MODEL. Splitting the conversion delay and the shift-out across
    -- separate processes gives `adc_ready` two drivers; on a resolved type that produces 'X' with no
    -- error, and the expected-value queue then latches a word that has not been assigned yet.
    adc : process (clk, rst_n) is
        variable conv_wait : natural := 0;
        variable sclk_d    : std_logic := '0';
    begin
        if rst_n = '0' then
            adc_word <= (others => '0');
            adc_sh   <= (others => '0');
            adc_ready <= '0';
            conv_cnt <= 0;
            conv_wait := 0;
            sclk_d := '0';
        elsif rising_edge(clk) then
            if convst = '1' then
                adc_word  <= adc_value(conv_cnt);
                conv_cnt  <= conv_cnt + 1;
                conv_wait := to_integer(s_tconv);
                adc_ready <= '0';
            elsif conv_wait > 1 then
                conv_wait := conv_wait - 1;
            elsif conv_wait = 1 then
                conv_wait := 0;
                adc_ready <= '1';
                adc_sh    <= adc_word;   -- the MSB is presented from here, before the select
            elsif adc_ready = '1' and sclk_d = '1' and sclk = '0' then
                -- Advance on the TRAILING edge: a mode-0 master then has a full half period of setup.
                adc_sh <= adc_sh(NB_C - 2 downto 0) & '0';
            end if;
            sclk_d := sclk;
        end if;
    end process adc;

    miso <= adc_sh(NB_C - 1) when adc_ready = '1' else '0';

    -- Expected-value queue plus per-sample checking, in ONE process so the queue has a single writer.
    -- The edge on `adc_ready` is detected from a registered copy rather than from an event, so the
    -- enqueue is ordered with respect to the clock like every other observation here -- and the copy is
    -- RESET, because an uninitialised edge detector makes the first enqueue disappear and every later
    -- sample compare against its predecessor.
    chk : process (clk, rst_n) is
        type q_t is array (0 to 63) of word_t;
        variable q : q_t := (others => (others => '0'));
        variable head, tail : natural := 0;
        variable ready_d : std_logic := '0';
        variable ln : line;
    begin
        if rst_n = '0' then
            -- The per-experiment counters are cleared HERE, with the queue indices. Leaving them to
            -- accumulate across experiments means the sweep's deliberately-marginal configuration
            -- contributes mismatches to the final integrity check, which then fails for a run that was
            -- clean -- the same class of bookkeeping error as an unreset edge detector.
            ready_d := '0';
            head := 0; tail := 0;
            got_n     <= 0;
            wrong_n   <= 0;
            x_reports <= 0;
        elsif rising_edge(clk) then
            if adc_ready = '1' and ready_d = '0' then
                q(tail mod 64) := adc_word;
                tail := tail + 1;
            end if;
            ready_d := adc_ready;

            if samp_push = '1' then
                got_n <= got_n + 1;
                if head < tail then
                    if samp_data /= q(head mod 64) then
                        wrong_n <= wrong_n + 1;
                        if expect_clean then
                            write(ln, string'("  FAIL: a delivered sample was ") & hex4(samp_data)
                                      & string'(" where the device sent ") & hex4(q(head mod 64)));
                            writeline(output, ln);
                        end if;
                    end if;
                    head := head + 1;
                else
                    write(ln, string'("  FAIL: a sample was pushed with no conversion to account for it"));
                    writeline(output, ln);
                end if;
            end if;

            -- X-safety. A natural cannot hold a metavalue, so the only field that can is the sample
            -- word, and it is the only one guarded.
            if busy_valid = '1' then
                for i in 0 to NB_C - 1 loop
                    if samp_data(i) /= '0' and samp_data(i) /= '1' then
                        x_reports <= x_reports + 1;
                    end if;
                end loop;
            end if;
        end if;
    end process chk;

    stim : process is

        variable e, mutations, floor_cycles : natural := 0;
        variable r_busy, r_conv, r_samp, r_over, r_drop : natural := 0;
        variable sweep_busy : nat_arr(0 to 5);
        constant HALVES_C : nat_arr(0 to 5) := (1, 2, 3, 4, 6, 8);
        variable pb : natural;
        variable ln : line;

        function predicted_busy (h : natural) return natural is
        begin
            return 20 + 3 + 2 * NB_C * h + 2 + 4;
        end function predicted_busy;

        procedure reset_all is
        begin
            -- THE WAIT COMES FIRST, and this is not stylistic. `s_half` is a SIGNAL: the caller's
            -- assignment to it has not taken effect at the moment this procedure is entered, so reading
            -- it here before any wait yields the PREVIOUS configuration. The gate was then computed for
            -- the wrong experiment and the deliberately-marginal sweep point reported failures the
            -- SystemVerilog bench -- where the same variable is a blocking-assigned reg -- suppressed.
            -- A signal read before the first wait in a procedure is reading the past.
            wait until falling_edge(clk);
            expect_clean <= (to_integer(s_half) > 1);
            rst_n <= '0';
            for i in 1 to 4 loop wait until falling_edge(clk); end loop;
            rst_n <= '1';
            for i in 1 to 2 loop wait until falling_edge(clk); end loop;
        end procedure reset_all;

        -- Run until the schedule has ticked `n` times, then let the read in flight finish. BOUNDED, and
        -- driven by the design's own tick accounting rather than by cycle arithmetic, which drifts.
        procedure run_n (n : natural) is
            variable guard : natural := 0;
        begin
            while (n_conv + n_overrun) < n and guard < 500000 loop
                wait until rising_edge(clk);
                guard := guard + 1;
            end loop;
            if guard >= 500000 then
                write(ln, string'("  FAIL: timeout waiting for ") & i2s(n, 1)
                          & string'(" schedule ticks"));
                writeline(output, ln); e := e + 1;
            end if;
            for i in 1 to to_integer(s_period) + 8 loop wait until rising_edge(clk); end loop;
            r_busy := busy_cycles; r_conv := n_conv; r_samp := n_samp;
            r_over := n_overrun;   r_drop := n_dropped;
        end procedure run_n;

    begin
        -- ============ 1. the arithmetic, checked against the hardware ============
        s_period <= to_unsigned(200, CNT_W);
        s_half   <= to_unsigned(2, 8);
        reset_all;
        run_n(5);
        write(ln, string'("  measured busy interval: ") & i2s(r_busy, 1)
                  & string'(" cycles   predicted ") & i2s(predicted_busy(2), 1)
                  & string'(" = t_conv 20 + lead 3 + 2*16*2 + lag 2 + gap 4"));
        writeline(output, ln);
        if r_busy /= predicted_busy(2) then
            write(ln, string'("  FAIL: the measured busy interval ") & i2s(r_busy, 1)
                      & string'(" disagrees with the arithmetic ") & i2s(predicted_busy(2), 1));
            writeline(output, ln); e := e + 1;
        end if;
        -- The invariant is `every conversion delivered except the one in flight`, not a fixed count:
        -- the free-running schedule ticks again while the last read completes.
        if r_conv < 5 or (r_conv - r_samp) > 1 or r_over /= 0 or r_drop /= 0 then
            write(ln, string'("  FAIL: a generous schedule did not deliver every sample cleanly"));
            writeline(output, ln); e := e + 1;
        end if;

        -- ============ 2. the boundary is exactly where the arithmetic puts it ============
        pb := predicted_busy(2);
        s_period <= to_unsigned(pb, CNT_W);
        reset_all; run_n(6);
        write(ln, string'("  period = busy (") & i2s(pb, 1) & string'("):      conv ") & i2s(r_conv, 1)
                  & string'("  samp ") & i2s(r_samp, 1) & string'("  overrun ") & i2s(r_over, 1)
                  & string'("  dropped ") & i2s(r_drop, 1));
        writeline(output, ln);
        if r_over /= 0 then
            write(ln, string'("  FAIL: a period equal to the busy interval overran ") & i2s(r_over, 1)
                      & string'(" times; the deadline should be exactly met"));
            writeline(output, ln); e := e + 1;
        end if;
        if r_samp = 0 then
            write(ln, string'("  FAIL: no samples were delivered at the boundary period"));
            writeline(output, ln); e := e + 1;
        end if;

        s_period <= to_unsigned(pb - 1, CNT_W);
        reset_all; run_n(6);
        write(ln, string'("  period = busy - 1 (") & i2s(pb - 1, 1) & string'("):  conv ")
                  & i2s(r_conv, 1) & string'("  samp ") & i2s(r_samp, 1)
                  & string'("  overrun ") & i2s(r_over, 1) & string'("  dropped ") & i2s(r_drop, 1));
        writeline(output, ln);
        if r_over = 0 then
            write(ln, string'("  FAIL: a period one cycle below the busy interval did not overrun"));
            writeline(output, ln); e := e + 1;
        end if;
        if r_drop /= 0 then
            write(ln, string'("  FAIL: a deadline miss also incremented the CONSUMER-stall counter"));
            writeline(output, ln); e := e + 1;
        end if;

        -- ============ 3. the other mechanism, in isolation ============
        s_period  <= to_unsigned(200, CNT_W);
        fifo_full <= '1';
        reset_all; run_n(6);
        write(ln, string'("  consumer stalled (period 200): conv ") & i2s(r_conv, 1)
                  & string'("  samp ") & i2s(r_samp, 1) & string'("  overrun ") & i2s(r_over, 1)
                  & string'("  dropped ") & i2s(r_drop, 1));
        writeline(output, ln);
        if r_drop = 0 then
            write(ln, string'("  FAIL: a full consumer FIFO dropped nothing"));
            writeline(output, ln); e := e + 1;
        end if;
        if r_over /= 0 then
            write(ln, string'("  FAIL: a consumer stall also incremented the DEADLINE counter"));
            writeline(output, ln); e := e + 1;
        end if;
        if r_samp /= 0 then
            write(ln, string'("  FAIL: samples were delivered while the FIFO was full"));
            writeline(output, ln); e := e + 1;
        end if;
        fifo_full <= '0';

        -- ============ 4. SCLK is not a sample-rate knob past a point ============
        -- COLLECT FIRST, THEN PRINT: a ratio printed inside the loop divides by an element the loop has
        -- not written yet.
        for k in 0 to 5 loop
            s_half   <= to_unsigned(HALVES_C(k), 8);
            s_period <= to_unsigned(400, CNT_W);
            reset_all; run_n(3);
            sweep_busy(k) := r_busy;
            if r_busy /= predicted_busy(HALVES_C(k)) then
                write(ln, string'("  FAIL: half=") & i2s(HALVES_C(k), 1) & string'(" measured ")
                          & i2s(r_busy, 1) & string'(" against a predicted ")
                          & i2s(predicted_busy(HALVES_C(k)), 1));
                writeline(output, ln); e := e + 1;
            end if;
        end loop;
        write(ln, string'(""));
        writeline(output, ln);
        write(ln, string'("  half  SCLK period  busy cycles  sample rate vs half=2"));
        writeline(output, ln);
        for k in 0 to 5 loop
            write(ln, string'("  ") & i2s(HALVES_C(k), 4) & string'("  ")
                      & i2s(2 * HALVES_C(k), 11) & string'("  ") & i2s(sweep_busy(k), 11)
                      & string'("  ") & r2s(real(sweep_busy(1)) / real(sweep_busy(k))) & string'("x"));
            writeline(output, ln);
        end loop;
        floor_cycles := 20 + 3 + 2 + 4;
        if not (sweep_busy(0) < sweep_busy(1) and sweep_busy(1) < sweep_busy(5)) then
            write(ln, string'("  FAIL: the busy interval did not grow monotonically with the half period"));
            writeline(output, ln); e := e + 1;
        end if;

        write(ln, string'(""));
        writeline(output, ln);
        write(ln, string'("    1. the measured busy interval matched t_conv + lead + 2*NB*half + lag + gap at every one of the six half-periods swept. The design is accountable to an equation rather than to a remembered number, which is what makes a schedule something an engineer can budget before the hardware exists"));
        writeline(output, ln);
        write(ln, string'("    2. the two loss mechanisms are separable. A period one cycle below the busy interval produced overruns and ZERO drops; a generous period with a full consumer FIFO produced drops and ZERO overruns. That matters because their fixes are opposite: the first needs a faster SPI path or a slower sample rate, and the second needs a deeper FIFO or a faster consumer. One combined `samples lost` counter sends you to optimise the wrong half of the system"));
        writeline(output, ln);
        write(ln, string'("    3. the deadline boundary sits exactly at the arithmetic. At period = ")
                  & i2s(pb, 1) & string'(" the schedule held with zero overruns; at ") & i2s(pb - 1, 1)
                  & string'(" -- one cycle less -- it did not. A boundary that is only approximately known is a boundary nobody can design against"));
        writeline(output, ln);
        write(ln, string'("    4. and SCLK is not a sample-rate knob past a point. Halving the half period from 2 to 1 cut the busy interval from ")
                  & i2s(sweep_busy(1), 1) & string'(" to ") & i2s(sweep_busy(0), 1)
                  & string'(" cycles -- a ") & r2s(real(sweep_busy(1)) / real(sweep_busy(0)))
                  & string'("x gain, not 2x -- because only the 2*NB*half term responds to SCLK. The conversion time, the select lead, the lag and the minimum deselected gap sum to a FIXED floor of ")
                  & i2s(floor_cycles, 1) & string'(" cycles, so the most any amount of SCLK can ever buy from this configuration is ")
                  & r2s(real(sweep_busy(1)) / real(floor_cycles)) & string'("x"));
        writeline(output, ln);

        -- ============ 5. sample integrity ============
        s_half   <= to_unsigned(2, 8);
        s_period <= to_unsigned(200, CNT_W);
        reset_all; run_n(8);
        if got_n < 8 then
            write(ln, string'("  FAIL: only ") & i2s(got_n, 1)
                      & string'(" samples reached the checker where 8 were expected"));
            writeline(output, ln); e := e + 1;
        end if;
        if wrong_n /= 0 then
            write(ln, string'("  FAIL: ") & i2s(wrong_n, 1)
                      & string'(" delivered samples did not match the device's word"));
            writeline(output, ln); e := e + 1;
        end if;
        write(ln, string'("    5. all ") & i2s(got_n, 1)
                  & string'(" delivered samples matched the word the ADC model presented for that conversion, compared against a queue filled when the DEVICE finished converting rather than against anything the design produced. Counting samples without checking them would have passed a design that delivers the right NUMBER of wrong values"));
        writeline(output, ln);

        -- ============ BENCH INTEGRITY ============
        -- Two deliberately wrong expectations, compared by the same operators as the real checks.
        -- BOTH ARE WRITTEN SO A CORRECT RUN MAKES THEM TRUE. The second was `wrong_n = 99` at first --
        -- a deliberately wrong expectation compared with the wrong POLARITY, so it never fired and the
        -- integrity check reported 1 of 2 on a bench that was working. A polarity error in a self-check
        -- is the one defect that makes every other check in a bench meaningless.
        if r_busy  /= predicted_busy(2) + 1 then mutations := mutations + 1; end if;
        if got_n   /= 999                   then mutations := mutations + 1; end if;
        if mutations /= 2 then
            write(ln, string'("  FAIL: a deliberately wrong expectation did not mismatch (")
                      & i2s(mutations, 1) & string'(" of 2)"));
            writeline(output, ln); e := e + 1;
        end if;
        if x_reports /= 0 then
            write(ln, string'("  FAIL: ") & i2s(x_reports, 1)
                      & string'(" reported sample words contained a metavalue"));
            writeline(output, ln); e := e + 1;
        end if;

        if e = 0 then
            write(ln, string'(""));
            writeline(output, ln);
            write(ln, string'("    and the bench proved itself: two deliberately wrong expectations mismatched, every reported counter carried a known value, and the ADC model driving MISO knows nothing about the design's schedule -- so the deadline it imposes is a real constraint rather than a handshake"));
            writeline(output, ln);
            write(ln, string'("PASS: a sample deadline is a closed loop between a timer, a conversion delay, an SPI frame and a consumer, and the only honest way to know it holds is to run the loop and count what survives. The measured busy interval matched t_conv + lead + 2*NB*half + lag + gap at all six half-periods swept, so the schedule can be budgeted from an equation before any hardware exists. TWO loss mechanisms exist and they are not the same bug: a period one cycle below the busy interval produced ")
                      & i2s(r_over, 1) & string'(" overruns and 0 drops, while a generous period with a full consumer FIFO produced ")
                      & i2s(r_drop, 1) & string'(" drops and 0 overruns -- opposite fixes, which is why one combined counter is an expensive instrumentation mistake. The boundary is exact: at ")
                      & i2s(pb, 1) & string'(" cycles the schedule held and at ") & i2s(pb - 1, 1)
                      & string'(" it did not. And SCLK is not a sample-rate knob past a point -- halving the half period bought ")
                      & r2s(real(sweep_busy(1)) / real(sweep_busy(0)))
                      & string'("x rather than 2x, because the conversion time, lead, lag and deselected gap form a fixed floor of ")
                      & i2s(floor_cycles, 1) & string'(" cycles that no clock rate touches, capping the total available gain at ")
                      & r2s(real(sweep_busy(1)) / real(floor_cycles)) & string'("x"));
            writeline(output, ln);
        else
            write(ln, string'("FAIL: ") & i2s(e, 1) & string'(" error(s)"));
            writeline(output, ln);
        end if;

        run <= false;
        wait;
    end process stim;

end architecture tb;

9. What The Bench Had To Prove About Itself

Three defects in the bench, all of which made a correct design look broken — the expensive direction.

10. FPGA Implementation

Three things about this design are FPGA-specific and worth deciding deliberately.

SCLK is generated logic, not a clock. It comes out of a counter and drives a pin. Do not route it through a global clock buffer and do not let the tools treat it as a clock domain — Module 15 made that argument in full. Drive it from a register straight to an output buffer so its pin timing is a simple output delay.

The FIFO is the only real crossing, and it must be a proper asynchronous FIFO. A dual-clock FIFO with Gray-coded pointers, or a vendor primitive. Not two synchronisers on a binary pointer — a multi-bit count sampled mid-increment can read a value that was never valid, which is Chapter 18.7's multi-bit crossing hazard applied to a pointer.

The gap and lead parameters are timing constraints in disguise. They exist because the device requires them, and their correct values come from the datasheet — not from what happens to work. A gap tuned down until the link stops failing is a design that fails at temperature.

11. What Is And Is Not Established Here

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   ✓ the busy interval equals t_conv + lead + 2·NB·half + lag + gap
   ✓ the schedule holds exactly when period ≥ busy
   ✓ the two loss mechanisms are independent and separately attributable
   ✓ delivered samples carry the word the device presented

   ✗ that the link has SETUP MARGIN at any given SCLK rate
   ✗ that the FIFO crossing is metastability-safe
   ✗ any statement about failure RATES under temperature or voltage

The first group is arithmetic about cycle counts, and a deterministic simulation settles it completely. The second group needs a model this simulation does not have: the device's output-valid time against the board, which is static timing analysis and signal integrity, and synchroniser depth, which no zero-delay simulation can address. Chapter 18.6 argued that boundary at length; it applies unchanged, and a passing run here is not evidence about any of it.

12. Why an ASIC Engineer Cares

The deadline argument is identical; what changes is who owns the fixed terms and how expensive it is to be wrong about them.

The conversion time belongs to the external device and appears in your budget as a number from its datasheet — so it belongs in a specification you can be held to, not in a comment. The lead, lag and gap become external timing constraints your STA must prove, and the 2·NB·half term is the only part you can trade after tape-out, through the divider, which is exactly why the divider should be software-programmable rather than a parameter.

And the two counters earn their area. A streaming subsystem that cannot tell a deadline miss from a consumer stall in the field is one that generates a bug report saying samples are occasionally missing, against which nobody can act.

13. Failure Signature — "The ADC Is Noisy"

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   Symptom     the captured waveform has occasional spikes; the spectrum has a
               noise floor nobody can account for
   Checked     the analog front end, the reference, the supply decoupling
   Actioned    a board revision with better grounding
   Result      unchanged
   Actual      the sample rate had been raised from 80 kSPS to 100 kSPS in
               firmware; the busy interval was 93 cycles against a period that
               had become 92, so roughly one sample in a few hundred was
               silently missing and the stream closed the gap
   Found       by an engineer who read the overrun counter -- which had been
               non-zero since the day the rate changed

Every step was a competent analog investigation of a digital fault. A missing sample in a time series is indistinguishable from an impulse, because closing the gap is an impulse — and the one observation that names it is a counter, not a waveform. Which is why the counter has to exist before the problem does.

14. Common Misconceptions

MisconceptionWhat is actually true
Enough average throughput means the deadline is metThe deadline is per-period; averages say nothing about the worst case
A missed sample looks like a bugIt looks like noise, which is why it is investigated as analog
SPI always implies a clock-domain crossingA master generating SCLK has none on its read path
Doubling SCLK doubles the sample rateOnly one term of the budget responds to SCLK; here the cap is 3.21×
Faster SCLK is always better for a deadlineIt costs setup margin, and the deadline equation cannot see that
One "samples lost" counter is enoughTwo mechanisms with opposite fixes need two counters
Aborting the in-flight read recovers the periodIt loses two samples instead of one

15. Reason It Through

16. Understanding Check

17. Summary

A sample deadline is a closed loop between a timer, a conversion delay, an SPI frame and a consumer, and the only honest way to know it holds is to run the loop and count what survives. The measured busy interval matched t_conv + lead + 2·NB·half + lag + gap at all six half-periods swept, so the schedule can be budgeted from an equation before any hardware exists — and the boundary is exact: at 93 cycles the schedule held, at 92 it did not.

Two loss mechanisms exist and they are not the same bug. A period one cycle short produced overruns and zero drops; a generous period with a full consumer FIFO produced drops and zero overruns. Their fixes are opposite, which is why one combined counter sends an investigation to the wrong half of the system.

SCLK is not a sample-rate knob past a point. Halving the half period bought 1.52× rather than 2×, because the conversion time, lead, lag and deselected gap form a fixed floor of 29 cycles that no clock rate touches, capping the total available gain at 3.21×. And pushing SCLK further eventually breaks the data path while improving the number the equation reports — the deadline arithmetic has no term for setup margin, so an optimisation guided by it alone walks into a link that does not work.

The design's one subtle bug was a priority: a scheduler and a read engine both writing one state register, in the same cycle, on a tight schedule. Evaluated in the wrong order it counted a conversion whose sample was never read and reported a healthy system — a missing number in instrumentation, which is the failure shape nobody investigates.

18. What Comes Next

This chapter's device converted on the FPGA's schedule. Chapter 19.2 inverts that: a sensor decides when data is ready and asserts a pin to say so, and the transfer must be triggered by an event that is asynchronous to everything — including to a transfer already in progress. The design question becomes what to do when the next event arrives before the current read is finished, and the right answer depends on a line in the datasheet.

Continue learning