Skip to content
VLSI Mentor

SPI · Module 15

Architecture A — SCLK as a Clock Domain

Clocking the shift path on SCLK removes the per-bit ratio precondition and works at SCLK faster than the system clock, which oversampling cannot. It replaces it with a per-word requirement — a measured twenty-four-fold relief — and costs a synthesis-time mode, a hand-off register that must outlive its transaction, and a reset edge chip select does not supply.

Module 14 built Architecture B and stated its precondition (Chapter 14.1): each SCLK half-period must last at least HALF_MIN system clocks, because an oversampler that misses a half-period loses an edge.

That is a per-bit requirement, and it is the reason Architecture B cannot be used when SCLK is faster than the system clock. A 50 MHz SPI flash interface on a 25 MHz control domain has no oversampling option at any parameter setting.

This chapter takes the other answer.

Clock the shift register, the bit counter and the word assembly with SCLK itself. What does that buy, and what does it cost?

It buys the removal of the ratio requirement. It costs the boundary not disappearing but moving — and this chapter's real content is what the boundary costs once it has moved.

1. What The SCLK Domain Has, And What It Does Not

It has a clock only while the master is clocking. That is not a detail; it removes a whole class of design move.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   * nothing can happen "after the last bit", because there is no edge after the
     last bit. Every value the outside world needs must be published ON the edge
     that completes it.
   * no timeout, no watchdog, no "if nothing has happened for N cycles" -- there
     are no cycles.
   * no reset release. A synchronous reset in this domain is released by an edge
     that may never come, so the domain's reset must be ASYNCHRONOUS.

And the only asynchronous signal available is chip select — which is the Architecture A idiom and what real SCLK-clocked slaves do. It is elegant, and §4 and §5 are about its two sharp edges.

2. The Mode Cannot Be A Run-Time Input

This is usually left out, and it is a real cost.

The capture edge is rising when CPOL equals CPHA and falling otherwise. In Architecture B that was a mux on a strobe — Chapter 14.6 built it in two gates. Here it is a mux on a clock, and a mux on a clock is not something to put in RTL: it glitches the clock net, it defeats clock-tree synthesis, and static timing analysis will not analyse the path it creates.

So CAP_ON_RISING is a parameter. A device built this way supports one mode, or two if a dedicated clock-inversion primitive is instantiated, and a run-time four-mode slave in Architecture A means two shift paths and a mux on the data.

The design writes the inversion as an XOR with a constant:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   wire cap_clk = sclk_pin ^ ~CAP_ON_RISING;

which synthesis resolves to either sclk_pin or its complement, so exactly one clock net exists in the implementation and no mux is built. On an FPGA the inverted case is absorbed by the clock buffer's optional inversion rather than costing a gate.

3. The Requirement It Does Have, Measured

A completed word is published with a toggle (Chapter 15.1's scheme 3) and the destination samples the data register on the cycle the toggle's change reaches the end of its synchroniser — SYNC_N destination cycles after the source changed it. That read is correct only while the register still holds the word the toggle announced, and the source overwrites it LEN_FIXED SCLK periods later:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   SYNC_N x (system period)  <  LEN_FIXED x (SCLK period)

Compare with Architecture B's, from Chapter 14.1:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   HALF_MIN x (system period)  <  one SCLK HALF-period

Written as a minimum SCLK period at a 10 ns system clock, with HALF_MIN = 3, SYNC_N = 2 and an eight-bit frame:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   Architecture B needs   SCLK period > 60 ns
   Architecture A needs   SCLK period > 2.5 ns

A factor of 2 × HALF_MIN × LEN / SYNC_N, which is twenty-four. The requirement did not go away; it moved from per-bit to per-word and got easier by that factor.

The measurement

Four words per run at a fixed 10 ns system clock:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   sclk_period  words  correct  overrun
         80 ns      4        4        0
         20 ns      4        4        0
         10 ns      4        4        0
          4 ns      4        4        0
          2 ns      4        1        1

The row at a 4 ns SCLK period is the headline: SCLK two and a half times faster than the system clock, four words correct. Architecture B cannot reach that row at any parameter setting.

The row at 2 ns is the per-word requirement failing — a word arrives every 16 ns and the destination needs 20 — and the design reports it.

And the boundary, swept from the other side

SCLK fixed at a 4 ns period so a word arrives every 32 ns, with the system clock slowed:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   sys_period  words  overrun
         4 ns      4        0
         6 ns      4        0
         8 ns      4        0
        12 ns      4        0
        20 ns      4        1
        40 ns      4        1

The boundary falls between 12 ns and 20 ns, which is exactly where SYNC_N × system period crosses 32 ns at SYNC_N = 2. The formula is a measurement rather than an assertion.

4. The Reset Chip Select Cannot Supply

Chip select as the domain's reset is the Architecture A idiom, and it has a sharp edge that cost this design a debugging session.

Chip select alone is not enough. At power-on chip select is already high — it is not driven high, it simply is high — so there is no edge. And an asynchronous reset needs an edge, or a level that is present when a clock arrives. This domain has no clock while the select is high, so the level never gets sampled either.

The registers come up undefined and stay undefined. The first transaction shifts X into a shift register that toggles an X flag at a synchroniser which propagates X into the system domain. Every signal looks plausible in a waveform and no word ever arrives.

The fix is that the domain's reset must be chip select or the system reset:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   dom_rst = cs_n_pin | ~rst_n

which is a reset crossing a clock boundary and therefore a CDC question of its own — one that Architecture B never has, because there is only one domain to reset. What makes it safe here is that the release happens when the select falls, and the first capture edge cannot arrive until the master's CS-to-SCLK lead has elapsed.

So Architecture A needs a minimum CS lead too, for a completely different reason from Chapter 14.4's: there, to let a first bit be driven; here, to let a reset release before the clock starts.

5. The Hand-Off Register Must Outlive Its Transaction

This is the chapter's sharpest finding, and the version that shared one reset with the shift path passed four of five ratios before failing in a way that looked like a ratio problem.

The destination reads the hand-off register SYNC_N system cycles after the toggle arrives. By then the master has finished the transaction and raised chip select — which, if chip select reset that register, would have cleared the word the destination is on its way to fetch.

The last word of every transaction is lost, and only the last one. So the symptom is a one-in-N corruption that scales with transaction length and looks exactly like a marginal ratio.

The fix is two processes with two different resets:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   the shift path  (sr, bit_idx)      reset by chip select OR the system reset
   the hand-off    (word_src, tog)    reset by the SYSTEM RESET ONLY

And the rule it generalises to:

A cross-domain hand-off register must outlive the event that filled it. Its lifetime is set by the reader, not by the writer — and in Architecture A the writer's clock has stopped by the time the reader arrives.

word_tog belongs in the same process for the same reason and one more: if the toggle were reset while the data was not, the destination would see a transition it had already accounted for and fetch the same word twice.

6. The Block Diagram

Architecture A: MOSI and an inverted-or-not SCLK clock a shift register and bit counter reset by chip select, a hand-off register and toggle reset only by the system reset, then a destination synchroniser and word publicationsclk_pinmosi_pincs_n_pincap_clkshift registerbit counterhand-off registerword toggleSYNC_N flopsrx_data, validrx_overrunresetwhole wordboundaryread on arrival12
Figure 1 — Architecture A. The dashed line is the domain boundary: everything left of it is clocked by SCLK and has no clock between transactions. Note that the two SCLK-domain processes have DIFFERENT resets — the shift path is reset by the select and the hand-off only by the system reset, which is section 5's finding and the difference between losing the last word of every transaction and not.
Twelve cycles across six rows. A capture clock has its last edge at cycle 2. A word toggle flips at cycle 2. Chip select rises at cycle 6 and resets the shift register. A destination synchroniser sees the toggle at cycle 8, when the hand-off register is read.published on the final capturepublished on the finalcaptureselect resets the SHIFT pathselect resets the SHIFTpaththe reader arrives herethe reader arrives herecap_clkword_togcs_n_pinsr (shift)00A5A5A5A5000000000000word_src00A5A5A5A5A5A5A5A5A5A5rx_valid_stbt0t1t2t3t4t5t6t7t8t9t10t11
Figure 2 — the hand-off, drawn at the moment section 5 is about. The final capture publishes the word and flips the toggle; chip select rises two cycles later and resets the shift path; and the destination fetches the word two system cycles after that. A hand-off register reset by chip select would have been cleared at cycle 6, before the read at cycle 8.

7. Building Architecture A — Three HDLs

The circuit

Three processes: the shift path on cap_clk reset by the select or the system, the hand-off on cap_clk reset by the system only, and the destination on the system clock. The frame width and capture edge are parameters, for §2's reason.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_slave_sclk_domain.sv — three processes, two resets, and a hand-off register whose lifetime is set by its reader
// spi_slave_sclk_domain.sv
//
// Chapter 15.2 -- Architecture A: SCLK really is a clock.
//
// Module 14 built Architecture B and stated its precondition (Chapter 14.1): SCLK
// must be slow enough that each half-period lasts at least HALF_MIN system clocks,
// because an oversampler that misses a half-period loses an edge. That is a
// PER-BIT requirement, and it is the reason Architecture B cannot be used when
// SCLK is faster than the system clock.
//
// This file takes the other answer. The shift register, the bit counter and the
// word assembly are clocked BY SCLK. There is then no oversampling, no edge
// recovery, and no ratio requirement on SCLK at all -- a 50 MHz SPI flash
// interface on a 25 MHz control domain works, which Architecture B cannot do at
// any parameter setting.
//
// The cost is that the boundary moves rather than disappears, and this file's real
// content is what the boundary costs once it has moved.
//
// WHAT THE SCLK DOMAIN HAS, AND WHAT IT DOES NOT.
//
// It has a clock only while the master is clocking. That is not a detail; it
// removes a whole class of design move:
//
//   * nothing can happen "after the last bit", because there is no edge after the
//     last bit. Every value the outside world needs must be published ON the edge
//     that completes it.
//   * no timeout, no watchdog, no "if nothing has happened for N cycles" -- there
//     are no cycles.
//   * no reset release. A synchronous reset in this domain is released by an edge
//     that may never come, so the domain's reset must be ASYNCHRONOUS, and the
//     only asynchronous signal available is chip select.
//
// So CS is the domain's reset, which is the Architecture A idiom and is exactly
// what real SCLK-clocked slaves do. It is elegant and it has two sharp edges.
//
// The first is that a glitch on chip select resets the shift path mid-word, and
// Chapter 14.8's careful "refuse the transaction" reasoning has nowhere to live
// here, because the thing that would notice is itself reset.
//
// The second is subtler and it cost this design a debugging session. CHIP SELECT
// ALONE IS NOT ENOUGH. At power-on chip select is already high -- it is not driven
// high, it simply is high -- so there is no EDGE, and an asynchronous reset needs
// an edge or a level that is present when a clock arrives. This domain has no
// clock while the select is high, so the level never gets sampled either. The
// registers therefore come up undefined and stay undefined, and the first
// transaction shifts X into a shift register that then toggles an X flag at a
// synchroniser which propagates X into the system domain. Every signal looks
// plausible in a waveform and no word ever arrives.
//
// The fix is that the domain's reset must be chip select OR the system reset:
//
//     dom_rst = cs_n_pin | ~rst_n
//
// which is a RESET CROSSING A CLOCK BOUNDARY, and therefore a CDC question of its
// own -- one that Architecture B never has, because there is only one domain to
// reset. Note what makes it safe here: the release happens when the select falls,
// and the first capture edge cannot arrive until the master's CS-to-SCLK lead has
// elapsed. So Architecture A needs a minimum CS lead too, for a completely
// different reason from Chapter 14.4's -- there, to let a first bit be driven;
// here, to let a reset release before the clock starts.
//
// THE PART THAT IS USUALLY LEFT OUT: THE MODE CANNOT BE A RUN-TIME INPUT.
//
// The capture edge is rising when CPOL equals CPHA and falling otherwise. In
// Architecture B that was a mux on a STROBE (Chapter 14.6, two gates). Here it is
// a mux on a CLOCK, and a mux on a clock is not something to put in RTL: it
// produces glitches on the clock net, it defeats clock-tree synthesis, and static
// timing analysis will not analyse the path it creates.
//
// So `CAP_ON_RISING` is a PARAMETER. A device built this way supports one mode, or
// two if a dedicated clock-inversion primitive is instantiated, and a run-time
// four-mode slave in Architecture A means two shift paths and a mux on the DATA.
// That is a real cost and it belongs in the comparison of Chapter 15.4 rather than
// in a footnote.
//
// THE REQUIREMENT ARCHITECTURE A DOES HAVE, WHICH IS THE POINT OF THE CHAPTER.
//
// A completed word is published with a toggle (scheme 3 of Chapter 15.1) and the
// destination samples the data register on the cycle the toggle's change reaches
// the end of its synchroniser -- which is SYNC_N destination cycles after the
// source changed it. That read is correct only while the data register still holds
// the word the toggle announced, and the source overwrites it LEN_FIXED SCLK
// periods later. So the requirement is:
//
//     SYNC_N x (system period)  <  LEN_FIXED x (SCLK period)
//
// Compare that with Architecture B's, from Chapter 14.1:
//
//     HALF_MIN x (system period)  <  one SCLK HALF-period
//
// Written as a minimum SCLK period at a 10 ns system clock, with HALF_MIN = 3,
// SYNC_N = 2 and an eight-bit frame:
//
//     Architecture B needs   SCLK period > 60 ns
//     Architecture A needs   SCLK period > 2.5 ns
//
// a factor of 2 x HALF_MIN x LEN / SYNC_N, which is TWENTY-FOUR. The requirement
// did not go away; it moved from PER-BIT to PER-WORD and got easier by that
// factor. That number is the whole quantitative case for Architecture A, and the
// testbench measures it rather than asserting it.
//
// AND THE LIMIT OF WHAT THE DESTINATION CAN REPORT. `rx_overrun` fires when two
// word arrivals land on consecutive destination cycles, which is a RATE
// observation. The failure the requirement is really about is STALENESS -- the
// destination reading a register the source has already overwritten -- and the
// destination cannot observe that at all: it reads a well-formed word that simply
// belongs to the wrong frame. The two thresholds happen to coincide, because both
// amount to "the word interval fell below SYNC_N destination cycles", and that
// coincidence is worth knowing rather than relying on: this is the fourth thing in
// the sequence that is invisible from inside the design, after a lost edge, a late
// first bit and bus contention.

module spi_slave_sclk_domain #(
    parameter int MAX_W         = 32,
    parameter int LEN_W         = 6,
    parameter int SYNC_N        = 2,
    // Rising when CPOL == CPHA, falling otherwise. A PARAMETER, not an input:
    // see the header. The value is the mode's capture edge, fixed at synthesis.
    parameter bit CAP_ON_RISING = 1'b1,
    parameter bit LSB_FIRST     = 1'b0,
    parameter int LEN_FIXED     = 8    // frame width, also fixed: see the header
) (
    // --- the pins. sclk_pin is a CLOCK here, not data. ---------------------
    input  wire              sclk_pin,
    input  wire              cs_n_pin,
    input  wire              mosi_pin,

    // --- the system domain -------------------------------------------------
    input  wire              clk,
    input  wire              rst_n,

    output reg  [MAX_W-1:0]  rx_data,
    output reg               rx_valid_stb,
    // Sticky: the destination did not take a word before the next one replaced
    // it. This is the Architecture A precondition, observed rather than assumed.
    output reg               rx_overrun,
    input  wire              clr_flags
);

    // =====================================================================
    // THE SCLK DOMAIN
    //
    // Reset is chip select, asynchronously. Every register here is reset by the
    // select rising, and nothing here has a clock while the select is high.
    // =====================================================================

    // The capture clock. `^ ~CAP_ON_RISING` is an inversion by a constant, which
    // synthesis resolves to either `sclk_pin` or its complement -- so exactly one
    // clock net exists in the implementation and no mux is built. Written this way
    // rather than with a generate block because the intent is clearer and the
    // result is identical; on an FPGA the inverted case is absorbed by the clock
    // buffer's optional inversion rather than costing a gate.
    wire cap_clk = sclk_pin ^ ~CAP_ON_RISING;

    // The domain's asynchronous reset condition is "deselected, OR the system in
    // reset". See the header for why chip select alone leaves this domain
    // undefined forever.
    //
    // It is NOT given a name here, and that is deliberate -- see the callout on the
    // shift path. A derived wire in an asynchronous reset CONDITION races the
    // sensitivity list that triggered the process, and the reset silently does not
    // happen. The condition is written out of its constituent signals instead.

    // The last bit index of a word, as a correctly sized localparam rather than a
    // part-select of the parameter. `LEN_FIXED[LEN_W-1:0]` is the natural way to
    // write it and Icarus Verilog evaluates a part-select of a parameter as ZERO,
    // silently -- so the comparison becomes `bit_idx == 63`, no word ever
    // completes, and the symptom is a slave that receives nothing while every
    // other signal looks correct. Sizing the localparam avoids the construct
    // entirely, which is better than relying on a simulator to get it right.
    localparam [LEN_W-1:0] LAST_BIT = LEN_FIXED - 1;

    reg [MAX_W-1:0]  sr;
    reg [LEN_W-1:0]  bit_idx;
    reg [MAX_W-1:0]  word_src;   // the completed word, held for the destination
    reg              word_tog;   // toggles once per completed word

    // The word boundary, computed so that the shift path and the hand-off can be
    // two processes with two different resets. Splitting them is not tidiness --
    // see the callout below the shift path for what happens when they share one.
    wire word_done = ~cs_n_pin & rst_n & (bit_idx == LAST_BIT);
    wire [MAX_W-1:0] sr_next = {sr[MAX_W-2:0], mosi_pin};

    // `bit_idx` counts captures within a word and is reset by the SELECT, never
    // by its own carry -- which is Chapter 14.3's rule and survives the change of
    // architecture unchanged, because the reason for it was the protocol rather
    // than the clocking.
    // --- the shift path: reset by the SELECT, because it is per-transaction ----
    //
    // TWO asynchronous reset sources, each with its OWN edge in the sensitivity
    // list, rather than one `posedge dom_rst`. The single-signal form is what one
    // naturally writes and it does not work in simulation: `dom_rst` is already
    // asserted at time zero -- chip select is high and the system is in reset -- so
    // it never RISES, the process never triggers, and these registers stay
    // undefined until the first deselect. The first transaction after power-on
    // then shifts X, and the symptom is a slave that works from the second
    // transaction onwards.
    //
    // Synthesis merges the two into the flop's single reset input, so this costs
    // nothing; it is written this way so that simulation sees the assertion that
    // silicon would see from a level.
    //
    // AND THE CONDITION TESTS `cs_n_pin` AND `rst_n` DIRECTLY, not a derived wire
    // like `cs_n_pin | ~rst_n`. This is the second thing that cost a debugging
    // session, and it is a genuine trap rather than a style preference.
    //
    // A wire is updated in its own delta cycle. When chip select rises, this
    // process is triggered by `posedge cs_n_pin` immediately -- and a derived wire
    // built from `cs_n_pin` has not necessarily been recomputed yet, so the
    // condition evaluates against the OLD value, finds no reset, and takes the
    // clocked branch instead. The reset does not happen at all.
    //
    // The symptom was precise and thoroughly misleading: a partial transaction left
    // `bit_idx` at 4 and three bits in the shift register, and the NEXT transaction
    // completed a word four bits early, delivering a value made of both. It looked
    // like a word-boundary bug in a design whose word boundary was correct.
    //
    // The rule: THE CONDITION OF AN ASYNCHRONOUS RESET MUST BE BUILT FROM THE SAME
    // SIGNALS THAT APPEAR IN THE SENSITIVITY LIST, with no intermediate wire.
    always_ff @(posedge cap_clk or posedge cs_n_pin or negedge rst_n) begin
        if (cs_n_pin || !rst_n) begin
            sr      <= {MAX_W{1'b0}};
            bit_idx <= {LEN_W{1'b0}};
        end else begin
            sr <= sr_next;
            if (bit_idx == LAST_BIT)
                bit_idx <= {LEN_W{1'b0}};
            else
                bit_idx <= bit_idx + 1'b1;
        end
    end

    // --- the hand-off: reset ONLY by the system reset -------------------------
    //
    // THIS SPLIT IS THE CHAPTER'S SHARPEST FINDING, and the version that shared
    // one reset with the shift path passed four of five ratios before failing in a
    // way that looked like a ratio problem.
    //
    // The destination reads `word_src` SYNC_N + 1 system cycles after the toggle
    // arrives. By then the master has finished the transaction and raised chip
    // select -- which, if chip select reset this register, would have cleared the
    // word the destination is on its way to fetch. The last word of every
    // transaction is lost, and only the last one, so the symptom is a
    // one-in-N corruption that scales with the transaction length and looks
    // exactly like a marginal ratio.
    //
    // The rule it generalises to: A CROSS-DOMAIN HAND-OFF REGISTER MUST OUTLIVE THE
    // EVENT THAT FILLED IT. Its lifetime is set by the reader, not by the writer,
    // and in Architecture A the writer's clock has stopped by the time the reader
    // arrives.
    //
    // `word_tog` is in the same process for the same reason and one more: if the
    // toggle were reset while the data was not, the destination would see a
    // transition it had already accounted for and fetch the same word twice.
    always_ff @(posedge cap_clk or negedge rst_n) begin
        if (!rst_n) begin
            word_src <= {MAX_W{1'b0}};
            word_tog <= 1'b0;
        end else if (word_done) begin
            // Published ON the completing edge, because there is no later edge to
            // publish it on. `sr_next` rather than `sr` because the word includes
            // the bit being captured on this very edge.
            word_src <= LSB_FIRST ? reverse_low(sr_next) : sr_next;
            word_tog <= ~word_tog;
        end
    end

    // The same reversal as Chapter 13.8 and Chapter 14.3: reversing the low
    // LEN_FIXED bits is its own inverse, so there is one function rather than two.
    function automatic [MAX_W-1:0] reverse_low(input [MAX_W-1:0] v);
        integer i;
        begin
            reverse_low = {MAX_W{1'b0}};
            for (i = 0; i < LEN_FIXED; i = i + 1)
                reverse_low[i] = v[LEN_FIXED - 1 - i];
        end
    endfunction

    // =====================================================================
    // THE SYSTEM DOMAIN
    //
    // One bit crosses through a synchroniser: the toggle. The DATA does not go
    // through one, and that is correct rather than an oversight -- a multi-bit
    // value must never be synchronised bit by bit (Chapter 15.5). It is sampled
    // once, on the cycle the toggle's arrival says it is stable, and it is stable
    // because the source will not touch it for another `len` SCLK periods.
    // =====================================================================

    reg [SYNC_N-1:0] tog_sr;
    reg              tog_d;
    wire             word_arrived = tog_sr[SYNC_N-1] ^ tog_d;

    // The overrun detector needs to know that a SECOND word arrived before the
    // first was published. Since publication takes exactly one cycle after
    // `word_arrived`, an overrun is two arrivals on consecutive cycles -- which is
    // the only observable form the failure has from this side.
    reg arrived_d;

    always_ff @(posedge clk or negedge rst_n) begin
        if (!rst_n) begin
            tog_sr       <= {SYNC_N{1'b0}};
            tog_d        <= 1'b0;
            arrived_d    <= 1'b0;
            rx_data      <= {MAX_W{1'b0}};
            rx_valid_stb <= 1'b0;
            rx_overrun   <= 1'b0;
        end else begin
            tog_sr <= {tog_sr[SYNC_N-2:0], word_tog};
            tog_d  <= tog_sr[SYNC_N-1];

            if (clr_flags)
                rx_overrun <= 1'b0;

            rx_valid_stb <= word_arrived;
            arrived_d    <= word_arrived;

            if (word_arrived)
                rx_data <= word_src;

            // Two arrivals with no gap means the source produced words faster
            // than this side can retire them. At SYNC_N = 2 the destination needs
            // four cycles per word and the source supplies one every len SCLK
            // periods, so this is the requirement in the header, failing.
            if (word_arrived && arrived_d)
                rx_overrun <= 1'b1;
        end
    end

`ifdef SPI_CHECKS
    // A word must never be published without an arrival, and an arrival must
    // never be silently dropped. Both are one-cycle relationships and both are
    // the kind of thing a crossing gets wrong when someone "simplifies" it.
    reg chk_arr_d;
    always_ff @(posedge clk) begin
        chk_arr_d <= word_arrived;
        if (rst_n) begin
            if (rx_valid_stb && !chk_arr_d)
                $fatal(1, "a word was published without an arrival");
        end
    end
`endif

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_slave_sclk_domain.v — the same design in Verilog-2001
// spi_slave_sclk_domain.v
//
// Chapter 15.2 -- Architecture A: SCLK really is a clock.
//
// Module 14 built Architecture B and stated its precondition (Chapter 14.1): SCLK
// must be slow enough that each half-period lasts at least HALF_MIN system clocks,
// because an oversampler that misses a half-period loses an edge. That is a
// PER-BIT requirement, and it is the reason Architecture B cannot be used when
// SCLK is faster than the system clock.
//
// This file takes the other answer. The shift register, the bit counter and the
// word assembly are clocked BY SCLK. There is then no oversampling, no edge
// recovery, and no ratio requirement on SCLK at all -- a 50 MHz SPI flash
// interface on a 25 MHz control domain works, which Architecture B cannot do at
// any parameter setting.
//
// The cost is that the boundary moves rather than disappears, and this file's real
// content is what the boundary costs once it has moved.
//
// WHAT THE SCLK DOMAIN HAS, AND WHAT IT DOES NOT.
//
// It has a clock only while the master is clocking. That is not a detail; it
// removes a whole class of design move:
//
//   * nothing can happen "after the last bit", because there is no edge after the
//     last bit. Every value the outside world needs must be published ON the edge
//     that completes it.
//   * no timeout, no watchdog, no "if nothing has happened for N cycles" -- there
//     are no cycles.
//   * no reset release. A synchronous reset in this domain is released by an edge
//     that may never come, so the domain's reset must be ASYNCHRONOUS, and the
//     only asynchronous signal available is chip select.
//
// So CS is the domain's reset, which is the Architecture A idiom and is exactly
// what real SCLK-clocked slaves do. It is elegant and it has two sharp edges.
//
// The first is that a glitch on chip select resets the shift path mid-word, and
// Chapter 14.8's careful "refuse the transaction" reasoning has nowhere to live
// here, because the thing that would notice is itself reset.
//
// The second is subtler and it cost this design a debugging session. CHIP SELECT
// ALONE IS NOT ENOUGH. At power-on chip select is already high -- it is not driven
// high, it simply is high -- so there is no EDGE, and an asynchronous reset needs
// an edge or a level that is present when a clock arrives. This domain has no
// clock while the select is high, so the level never gets sampled either. The
// registers therefore come up undefined and stay undefined, and the first
// transaction shifts X into a shift register that then toggles an X flag at a
// synchroniser which propagates X into the system domain. Every signal looks
// plausible in a waveform and no word ever arrives.
//
// The fix is that the domain's reset must be chip select OR the system reset:
//
//     dom_rst = cs_n_pin | ~rst_n
//
// which is a RESET CROSSING A CLOCK BOUNDARY, and therefore a CDC question of its
// own -- one that Architecture B never has, because there is only one domain to
// reset. Note what makes it safe here: the release happens when the select falls,
// and the first capture edge cannot arrive until the master's CS-to-SCLK lead has
// elapsed. So Architecture A needs a minimum CS lead too, for a completely
// different reason from Chapter 14.4's -- there, to let a first bit be driven;
// here, to let a reset release before the clock starts.
//
// THE PART THAT IS USUALLY LEFT OUT: THE MODE CANNOT BE A RUN-TIME INPUT.
//
// The capture edge is rising when CPOL equals CPHA and falling otherwise. In
// Architecture B that was a mux on a STROBE (Chapter 14.6, two gates). Here it is
// a mux on a CLOCK, and a mux on a clock is not something to put in RTL: it
// produces glitches on the clock net, it defeats clock-tree synthesis, and static
// timing analysis will not analyse the path it creates.
//
// So `CAP_ON_RISING` is a PARAMETER. A device built this way supports one mode, or
// two if a dedicated clock-inversion primitive is instantiated, and a run-time
// four-mode slave in Architecture A means two shift paths and a mux on the DATA.
// That is a real cost and it belongs in the comparison of Chapter 15.4 rather than
// in a footnote.
//
// THE REQUIREMENT ARCHITECTURE A DOES HAVE, WHICH IS THE POINT OF THE CHAPTER.
//
// A completed word is published with a toggle (scheme 3 of Chapter 15.1) and the
// destination samples the data register on the cycle the toggle's change reaches
// the end of its synchroniser -- which is SYNC_N destination cycles after the
// source changed it. That read is correct only while the data register still holds
// the word the toggle announced, and the source overwrites it LEN_FIXED SCLK
// periods later. So the requirement is:
//
//     SYNC_N x (system period)  <  LEN_FIXED x (SCLK period)
//
// Compare that with Architecture B's, from Chapter 14.1:
//
//     HALF_MIN x (system period)  <  one SCLK HALF-period
//
// Written as a minimum SCLK period at a 10 ns system clock, with HALF_MIN = 3,
// SYNC_N = 2 and an eight-bit frame:
//
//     Architecture B needs   SCLK period > 60 ns
//     Architecture A needs   SCLK period > 2.5 ns
//
// a factor of 2 x HALF_MIN x LEN / SYNC_N, which is TWENTY-FOUR. The requirement
// did not go away; it moved from PER-BIT to PER-WORD and got easier by that
// factor. That number is the whole quantitative case for Architecture A, and the
// testbench measures it rather than asserting it.
//
// AND THE LIMIT OF WHAT THE DESTINATION CAN REPORT. `rx_overrun` fires when two
// word arrivals land on consecutive destination cycles, which is a RATE
// observation. The failure the requirement is really about is STALENESS -- the
// destination reading a register the source has already overwritten -- and the
// destination cannot observe that at all: it reads a well-formed word that simply
// belongs to the wrong frame. The two thresholds happen to coincide, because both
// amount to "the word interval fell below SYNC_N destination cycles", and that
// coincidence is worth knowing rather than relying on: this is the fourth thing in
// the sequence that is invisible from inside the design, after a lost edge, a late
// first bit and bus contention.

module spi_slave_sclk_domain #(
    parameter MAX_W         = 32,
    parameter LEN_W         = 6,
    parameter SYNC_N        = 2,
    // Rising when CPOL == CPHA, falling otherwise. A PARAMETER, not an input:
    // see the header. The value is the mode's capture edge, fixed at synthesis.
    parameter CAP_ON_RISING = 1'b1,
    parameter LSB_FIRST     = 1'b0,
    parameter LEN_FIXED     = 8    // frame width, also fixed: see the header
) (
    // --- the pins. sclk_pin is a CLOCK here, not data. ---------------------
    input  wire              sclk_pin,
    input  wire              cs_n_pin,
    input  wire              mosi_pin,

    // --- the system domain -------------------------------------------------
    input  wire              clk,
    input  wire              rst_n,

    output reg  [MAX_W-1:0]  rx_data,
    output reg               rx_valid_stb,
    // Sticky: the destination did not take a word before the next one replaced
    // it. This is the Architecture A precondition, observed rather than assumed.
    output reg               rx_overrun,
    input  wire              clr_flags
);

    // =====================================================================
    // THE SCLK DOMAIN
    //
    // Reset is chip select, asynchronously. Every register here is reset by the
    // select rising, and nothing here has a clock while the select is high.
    // =====================================================================

    // The capture clock. `^ ~CAP_ON_RISING` is an inversion by a constant, which
    // synthesis resolves to either `sclk_pin` or its complement -- so exactly one
    // clock net exists in the implementation and no mux is built. Written this way
    // rather than with a generate block because the intent is clearer and the
    // result is identical; on an FPGA the inverted case is absorbed by the clock
    // buffer's optional inversion rather than costing a gate.
    wire cap_clk = sclk_pin ^ ~CAP_ON_RISING;

    // The domain's asynchronous reset condition is "deselected, OR the system in
    // reset". See the header for why chip select alone leaves this domain
    // undefined forever.
    //
    // It is NOT given a name here, and that is deliberate -- see the callout on the
    // shift path. A derived wire in an asynchronous reset CONDITION races the
    // sensitivity list that triggered the process, and the reset silently does not
    // happen. The condition is written out of its constituent signals instead.

    // The last bit index of a word, as a correctly sized localparam rather than a
    // part-select of the parameter. `LEN_FIXED[LEN_W-1:0]` is the natural way to
    // write it and Icarus Verilog evaluates a part-select of a parameter as ZERO,
    // silently -- so the comparison becomes `bit_idx == 63`, no word ever
    // completes, and the symptom is a slave that receives nothing while every
    // other signal looks correct. Sizing the localparam avoids the construct
    // entirely, which is better than relying on a simulator to get it right.
    localparam [LEN_W-1:0] LAST_BIT = LEN_FIXED - 1;

    reg [MAX_W-1:0]  sr;
    reg [LEN_W-1:0]  bit_idx;
    reg [MAX_W-1:0]  word_src;   // the completed word, held for the destination
    reg              word_tog;   // toggles once per completed word

    // The word boundary, computed so that the shift path and the hand-off can be
    // two processes with two different resets. Splitting them is not tidiness --
    // see the callout below the shift path for what happens when they share one.
    wire word_done = ~cs_n_pin & rst_n & (bit_idx == LAST_BIT);
    wire [MAX_W-1:0] sr_next = {sr[MAX_W-2:0], mosi_pin};

    // `bit_idx` counts captures within a word and is reset by the SELECT, never
    // by its own carry -- which is Chapter 14.3's rule and survives the change of
    // architecture unchanged, because the reason for it was the protocol rather
    // than the clocking.
    // --- the shift path: reset by the SELECT, because it is per-transaction ----
    //
    // TWO asynchronous reset sources, each with its OWN edge in the sensitivity
    // list, rather than one `posedge dom_rst`. The single-signal form is what one
    // naturally writes and it does not work in simulation: `dom_rst` is already
    // asserted at time zero -- chip select is high and the system is in reset -- so
    // it never RISES, the process never triggers, and these registers stay
    // undefined until the first deselect. The first transaction after power-on
    // then shifts X, and the symptom is a slave that works from the second
    // transaction onwards.
    //
    // Synthesis merges the two into the flop's single reset input, so this costs
    // nothing; it is written this way so that simulation sees the assertion that
    // silicon would see from a level.
    //
    // AND THE CONDITION TESTS `cs_n_pin` AND `rst_n` DIRECTLY, not a derived wire
    // like `cs_n_pin | ~rst_n`. This is the second thing that cost a debugging
    // session, and it is a genuine trap rather than a style preference.
    //
    // A wire is updated in its own delta cycle. When chip select rises, this
    // process is triggered by `posedge cs_n_pin` immediately -- and a derived wire
    // built from `cs_n_pin` has not necessarily been recomputed yet, so the
    // condition evaluates against the OLD value, finds no reset, and takes the
    // clocked branch instead. The reset does not happen at all.
    //
    // The symptom was precise and thoroughly misleading: a partial transaction left
    // `bit_idx` at 4 and three bits in the shift register, and the NEXT transaction
    // completed a word four bits early, delivering a value made of both. It looked
    // like a word-boundary bug in a design whose word boundary was correct.
    //
    // The rule: THE CONDITION OF AN ASYNCHRONOUS RESET MUST BE BUILT FROM THE SAME
    // SIGNALS THAT APPEAR IN THE SENSITIVITY LIST, with no intermediate wire.
    always @(posedge cap_clk or posedge cs_n_pin or negedge rst_n) begin
        if (cs_n_pin || !rst_n) begin
            sr      <= {MAX_W{1'b0}};
            bit_idx <= {LEN_W{1'b0}};
        end else begin
            sr <= sr_next;
            if (bit_idx == LAST_BIT)
                bit_idx <= {LEN_W{1'b0}};
            else
                bit_idx <= bit_idx + 1'b1;
        end
    end

    // --- the hand-off: reset ONLY by the system reset -------------------------
    //
    // THIS SPLIT IS THE CHAPTER'S SHARPEST FINDING, and the version that shared
    // one reset with the shift path passed four of five ratios before failing in a
    // way that looked like a ratio problem.
    //
    // The destination reads `word_src` SYNC_N + 1 system cycles after the toggle
    // arrives. By then the master has finished the transaction and raised chip
    // select -- which, if chip select reset this register, would have cleared the
    // word the destination is on its way to fetch. The last word of every
    // transaction is lost, and only the last one, so the symptom is a
    // one-in-N corruption that scales with the transaction length and looks
    // exactly like a marginal ratio.
    //
    // The rule it generalises to: A CROSS-DOMAIN HAND-OFF REGISTER MUST OUTLIVE THE
    // EVENT THAT FILLED IT. Its lifetime is set by the reader, not by the writer,
    // and in Architecture A the writer's clock has stopped by the time the reader
    // arrives.
    //
    // `word_tog` is in the same process for the same reason and one more: if the
    // toggle were reset while the data was not, the destination would see a
    // transition it had already accounted for and fetch the same word twice.
    always @(posedge cap_clk or negedge rst_n) begin
        if (!rst_n) begin
            word_src <= {MAX_W{1'b0}};
            word_tog <= 1'b0;
        end else if (word_done) begin
            // Published ON the completing edge, because there is no later edge to
            // publish it on. `sr_next` rather than `sr` because the word includes
            // the bit being captured on this very edge.
            word_src <= LSB_FIRST ? reverse_low(sr_next) : sr_next;
            word_tog <= ~word_tog;
        end
    end

    // The same reversal as Chapter 13.8 and Chapter 14.3: reversing the low
    // LEN_FIXED bits is its own inverse, so there is one function rather than two.
        function [MAX_W-1:0] reverse_low;
        input [MAX_W-1:0] v;
        integer i;
        begin
            reverse_low = {MAX_W{1'b0}};
            for (i = 0; i < LEN_FIXED; i = i + 1)
                reverse_low[i] = v[LEN_FIXED - 1 - i];
        end
    endfunction

    // =====================================================================
    // THE SYSTEM DOMAIN
    //
    // One bit crosses through a synchroniser: the toggle. The DATA does not go
    // through one, and that is correct rather than an oversight -- a multi-bit
    // value must never be synchronised bit by bit (Chapter 15.5). It is sampled
    // once, on the cycle the toggle's arrival says it is stable, and it is stable
    // because the source will not touch it for another `len` SCLK periods.
    // =====================================================================

    reg [SYNC_N-1:0] tog_sr;
    reg              tog_d;
    wire             word_arrived = tog_sr[SYNC_N-1] ^ tog_d;

    // The overrun detector needs to know that a SECOND word arrived before the
    // first was published. Since publication takes exactly one cycle after
    // `word_arrived`, an overrun is two arrivals on consecutive cycles -- which is
    // the only observable form the failure has from this side.
    reg arrived_d;

    always @(posedge clk or negedge rst_n) begin
        if (!rst_n) begin
            tog_sr       <= {SYNC_N{1'b0}};
            tog_d        <= 1'b0;
            arrived_d    <= 1'b0;
            rx_data      <= {MAX_W{1'b0}};
            rx_valid_stb <= 1'b0;
            rx_overrun   <= 1'b0;
        end else begin
            tog_sr <= {tog_sr[SYNC_N-2:0], word_tog};
            tog_d  <= tog_sr[SYNC_N-1];

            if (clr_flags)
                rx_overrun <= 1'b0;

            rx_valid_stb <= word_arrived;
            arrived_d    <= word_arrived;

            if (word_arrived)
                rx_data <= word_src;

            // Two arrivals with no gap means the source produced words faster
            // than this side can retire them. At SYNC_N = 2 the destination needs
            // four cycles per word and the source supplies one every len SCLK
            // periods, so this is the requirement in the header, failing.
            if (word_arrived && arrived_d)
                rx_overrun <= 1'b1;
        end
    end

`ifdef SPI_CHECKS
    // A word must never be published without an arrival, and an arrival must
    // never be silently dropped. Both are one-cycle relationships and both are
    // the kind of thing a crossing gets wrong when someone "simplifies" it.
    reg chk_arr_d;
    always @(posedge clk) begin
        chk_arr_d <= word_arrived;
        if (rst_n) begin
            if (rx_valid_stb && !chk_arr_d)
                $fatal(1, "a word was published without an arrival");
        end
    end
`endif

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_slave_sclk_domain.vhd — the same design in VHDL
-- spi_slave_sclk_domain.vhd
--
-- Chapter 15.2 -- Architecture A: SCLK really is a clock.
--
-- Module 14 built Architecture B and stated its precondition (Chapter 14.1): SCLK
-- must be slow enough that each half-period lasts at least HALF_MIN system clocks,
-- because an oversampler that misses a half-period loses an edge. That is a
-- PER-BIT requirement, and it is the reason Architecture B cannot be used when
-- SCLK is faster than the system clock.
--
-- This file takes the other answer. The shift register, the bit counter and the
-- word assembly are clocked BY SCLK. There is then no oversampling, no edge
-- recovery, and no ratio requirement on SCLK at all -- a 50 MHz SPI flash
-- interface on a 25 MHz control domain works, which Architecture B cannot do at
-- any parameter setting.
--
-- The cost is that the boundary moves rather than disappears, and this file's real
-- content is what the boundary costs once it has moved.
--
-- WHAT THE SCLK DOMAIN HAS, AND WHAT IT DOES NOT.
--
-- It has a clock only while the master is clocking. That is not a detail; it
-- removes a whole class of design move:
--
--   * nothing can happen "after the last bit", because there is no edge after the
--     last bit. Every value the outside world needs must be published ON the edge
--     that completes it.
--   * no timeout, no watchdog, no "if nothing has happened for N cycles" -- there
--     are no cycles.
--   * no reset release. A synchronous reset in this domain is released by an edge
--     that may never come, so the domain's reset must be ASYNCHRONOUS, and the
--     only asynchronous signal available is chip select.
--
-- So CS is the domain's reset, which is the Architecture A idiom and is exactly
-- what real SCLK-clocked slaves do. It is elegant and it has two sharp edges.
--
-- The first is that a glitch on chip select resets the shift path mid-word, and
-- Chapter 14.8's careful "refuse the transaction" reasoning has nowhere to live
-- here, because the thing that would notice is itself reset.
--
-- The second is subtler and it cost this design a debugging session. CHIP SELECT
-- ALONE IS NOT ENOUGH. At power-on chip select is already high -- it is not driven
-- high, it simply is high -- so there is no EDGE, and an asynchronous reset needs
-- an edge or a level that is present when a clock arrives. This domain has no
-- clock while the select is high, so the level never gets sampled either. The
-- registers therefore come up undefined and stay undefined, and the first
-- transaction shifts X into a shift register that then toggles an X flag at a
-- synchroniser which propagates X into the system domain. Every signal looks
-- plausible in a waveform and no word ever arrives.
--
-- The fix is that the domain's reset must be chip select OR the system reset:
--
--     dom_rst = cs_n_pin | ~rst_n
--
-- which is a RESET CROSSING A CLOCK BOUNDARY, and therefore a CDC question of its
-- own -- one that Architecture B never has, because there is only one domain to
-- reset. Note what makes it safe here: the release happens when the select falls,
-- and the first capture edge cannot arrive until the master's CS-to-SCLK lead has
-- elapsed. So Architecture A needs a minimum CS lead too, for a completely
-- different reason from Chapter 14.4's -- there, to let a first bit be driven;
-- here, to let a reset release before the clock starts.
--
-- THE PART THAT IS USUALLY LEFT OUT: THE MODE CANNOT BE A RUN-TIME INPUT.
--
-- The capture edge is rising when CPOL equals CPHA and falling otherwise. In
-- Architecture B that was a mux on a STROBE (Chapter 14.6, two gates). Here it is
-- a mux on a CLOCK, and a mux on a clock is not something to put in RTL: it
-- produces glitches on the clock net, it defeats clock-tree synthesis, and static
-- timing analysis will not analyse the path it creates.
--
-- So `CAP_ON_RISING` is a PARAMETER. A device built this way supports one mode, or
-- two if a dedicated clock-inversion primitive is instantiated, and a run-time
-- four-mode slave in Architecture A means two shift paths and a mux on the DATA.
-- That is a real cost and it belongs in the comparison of Chapter 15.4 rather than
-- in a footnote.
--
-- THE REQUIREMENT ARCHITECTURE A DOES HAVE, WHICH IS THE POINT OF THE CHAPTER.
--
-- A completed word is published with a toggle (scheme 3 of Chapter 15.1) and the
-- destination samples the data register on the cycle the toggle's change reaches
-- the end of its synchroniser -- which is SYNC_N destination cycles after the
-- source changed it. That read is correct only while the data register still holds
-- the word the toggle announced, and the source overwrites it LEN_FIXED SCLK
-- periods later. So the requirement is:
--
--     SYNC_N x (system period)  <  LEN_FIXED x (SCLK period)
--
-- Compare that with Architecture B's, from Chapter 14.1:
--
--     HALF_MIN x (system period)  <  one SCLK HALF-period
--
-- Written as a minimum SCLK period at a 10 ns system clock, with HALF_MIN = 3,
-- SYNC_N = 2 and an eight-bit frame:
--
--     Architecture B needs   SCLK period > 60 ns
--     Architecture A needs   SCLK period > 2.5 ns
--
-- a factor of 2 x HALF_MIN x LEN / SYNC_N, which is TWENTY-FOUR. The requirement
-- did not go away; it moved from PER-BIT to PER-WORD and got easier by that
-- factor. That number is the whole quantitative case for Architecture A, and the
-- testbench measures it rather than asserting it.
--
-- AND THE LIMIT OF WHAT THE DESTINATION CAN REPORT. `rx_overrun` fires when two
-- word arrivals land on consecutive destination cycles, which is a RATE
-- observation. The failure the requirement is really about is STALENESS -- the
-- destination reading a register the source has already overwritten -- and the
-- destination cannot observe that at all: it reads a well-formed word that simply
-- belongs to the wrong frame. The two thresholds happen to coincide, because both
-- amount to "the word interval fell below SYNC_N destination cycles", and that
-- coincidence is worth knowing rather than relying on: this is the fourth thing in
-- the sequence that is invisible from inside the design, after a lost edge, a late
-- first bit and bus contention.

library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

entity spi_slave_sclk_domain is
    generic (
        MAX_W         : positive := 32;
        LEN_W         : positive := 6;
        SYNC_N        : positive := 2;
        -- Rising when CPOL = CPHA, falling otherwise. A GENERIC, not a port:
        -- see the header. The value is the mode's capture edge, fixed at synthesis.
        CAP_ON_RISING : std_logic := '1';
        LSB_FIRST     : std_logic := '0';
        LEN_FIXED     : positive := 8    -- frame width, also fixed: see the header
    );
    port (
        -- the pins. sclk_pin is a CLOCK here, not data.
        sclk_pin     : in  std_logic;
        cs_n_pin     : in  std_logic;
        mosi_pin     : in  std_logic;

        -- the system domain
        clk          : in  std_logic;
        rst_n        : in  std_logic;

        rx_data      : out std_logic_vector(MAX_W - 1 downto 0);
        rx_valid_stb : out std_logic;
        -- Sticky: the destination did not take a word before the next one
        -- replaced it. The Architecture A precondition, observed not assumed.
        rx_overrun   : out std_logic;
        clr_flags    : in  std_logic
    );
end entity;

architecture rtl of spi_slave_sclk_domain is

    -- The capture clock. An inversion by a CONSTANT, which synthesis resolves to
    -- either `sclk_pin` or its complement -- so exactly one clock net exists in the
    -- implementation and no mux is built.
    signal cap_clk : std_logic;

    constant LAST_BIT : unsigned(LEN_W - 1 downto 0) :=
        to_unsigned(LEN_FIXED - 1, LEN_W);

    signal sr        : std_logic_vector(MAX_W - 1 downto 0) := (others => '0');
    signal sr_next   : std_logic_vector(MAX_W - 1 downto 0);
    signal bit_idx   : unsigned(LEN_W - 1 downto 0) := (others => '0');
    signal word_src  : std_logic_vector(MAX_W - 1 downto 0) := (others => '0');
    signal word_tog  : std_logic := '0';
    signal word_done : std_logic;

    signal tog_sr    : std_logic_vector(SYNC_N - 1 downto 0) := (others => '0');
    signal tog_d     : std_logic := '0';
    signal arrived   : std_logic;
    signal arrived_d : std_logic := '0';

    signal rx_data_r : std_logic_vector(MAX_W - 1 downto 0) := (others => '0');
    signal rx_vld_r  : std_logic := '0';
    signal rx_ovr_r  : std_logic := '0';

    -- The same reversal as Chapter 13.8 and Chapter 14.3: reversing the low
    -- LEN_FIXED bits is its own inverse, so there is one function rather than two.
    function reverse_low(v : std_logic_vector) return std_logic_vector is
        variable r : std_logic_vector(v'range) := (others => '0');
    begin
        for i in 0 to LEN_FIXED - 1 loop
            r(i) := v(LEN_FIXED - 1 - i);
        end loop;
        return r;
    end function;

begin

    cap_clk   <= sclk_pin xor (not CAP_ON_RISING);
    sr_next   <= sr(MAX_W - 2 downto 0) & mosi_pin;
    word_done <= '1' when (cs_n_pin = '0' and rst_n = '1' and bit_idx = LAST_BIT)
                 else '0';

    rx_data      <= rx_data_r;
    rx_valid_stb <= rx_vld_r;
    rx_overrun   <= rx_ovr_r;

    -- --- the shift path: reset by the SELECT, because it is per-transaction ----
    --
    -- Worth comparing with the SystemVerilog above, because the languages differ
    -- here in a way that matters. Verilog's sensitivity list is a list of EDGES,
    -- so a reset that is already asserted at time zero never triggers the process
    -- and the registers stay undefined. VHDL's is a list of SIGNALS and every
    -- process runs once at time zero, so the level is evaluated and the reset
    -- applies -- which means this VHDL works in a bench that leaves the Verilog
    -- undefined. Neither behaviour is the hardware's; the hardware holds the flop
    -- reset because the level is asserted, and only the VHDL happens to agree.
    shift : process (cap_clk, cs_n_pin, rst_n)
    begin
        if cs_n_pin = '1' or rst_n = '0' then
            sr      <= (others => '0');
            bit_idx <= (others => '0');
        elsif rising_edge(cap_clk) then
            sr <= sr_next;
            if bit_idx = LAST_BIT then
                bit_idx <= (others => '0');
            else
                bit_idx <= bit_idx + 1;
            end if;
        end if;
    end process;

    -- --- the hand-off: reset ONLY by the system reset -------------------------
    -- See the SystemVerilog for why this is a separate process with a different
    -- reset: a cross-domain hand-off register must outlive the event that filled
    -- it, because its lifetime is set by the reader and the writer's clock has
    -- stopped by the time the reader arrives.
    handoff : process (cap_clk, rst_n)
    begin
        if rst_n = '0' then
            word_src <= (others => '0');
            word_tog <= '0';
        elsif rising_edge(cap_clk) then
            if word_done = '1' then
                if LSB_FIRST = '1' then
                    word_src <= reverse_low(sr_next);
                else
                    word_src <= sr_next;
                end if;
                word_tog <= not word_tog;
            end if;
        end if;
    end process;

    -- =====================================================================
    -- THE SYSTEM DOMAIN
    -- =====================================================================
    arrived <= tog_sr(SYNC_N - 1) xor tog_d;

    destination : process (clk, rst_n)
    begin
        if rst_n = '0' then
            tog_sr    <= (others => '0');
            tog_d     <= '0';
            arrived_d <= '0';
            rx_data_r <= (others => '0');
            rx_vld_r  <= '0';
            rx_ovr_r  <= '0';
        elsif rising_edge(clk) then
            tog_sr <= tog_sr(SYNC_N - 2 downto 0) & word_tog;
            tog_d  <= tog_sr(SYNC_N - 1);

            if clr_flags = '1' then
                rx_ovr_r <= '0';
            end if;

            rx_vld_r  <= arrived;
            arrived_d <= arrived;

            if arrived = '1' then
                rx_data_r <= word_src;
            end if;

            if arrived = '1' and arrived_d = '1' then
                rx_ovr_r <= '1';
            end if;
        end if;
    end process;

    -- A word must never be published without an arrival. A one-cycle relationship,
    -- and the kind of thing a crossing gets wrong when someone simplifies it.
    check : process (clk)
        variable arr_d : std_logic := '0';
    begin
        if rising_edge(clk) then
            if rst_n = '1' then
                assert not (rx_vld_r = '1' and arr_d = '0')
                    report "a word was published without an arrival" severity failure;
            end if;
            arr_d := arrived;
        end if;
    end process;

end architecture;

The testbench

Two experiments and three directed cases. Experiment 1 establishes that Architecture A works where B cannot; experiment 2 measures the per-word boundary two-sidedly, which is what turns the formula from a claim into a result.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_slave_sclk_domain_tb.sv — two ratio sweeps from opposite sides, with expectations computed from the requirement
// spi_slave_sclk_domain_tb.sv
//
// Two experiments, and the second is the chapter's number.
//
//   1  ARCHITECTURE A WORKS WHERE ARCHITECTURE B CANNOT. SCLK is driven FASTER
//      than the system clock -- up to eight times faster -- and every word arrives
//      correctly. Chapter 14.1's oversampler cannot do this at any parameter
//      setting, because its precondition is that an SCLK half-period lasts several
//      system clocks.
//
//   2  AND IT HAS ITS OWN PRECONDITION, PER WORD RATHER THAN PER BIT. The bench
//      sweeps the system clock from fast to slow against a fixed SCLK and finds the
//      exact period at which `rx_overrun` starts -- then checks that period against
//      the arithmetic in the design's header. The relief over Architecture B is a
//      factor of 2 x len, and the bench measures it rather than believing it.
//
// The master here is a pin-level driver written from the protocol, not from the
// design: it toggles SCLK, drives MOSI on the launch edge, and holds chip select.
// That discipline is Chapter 14.1's and it matters more here, because the DUT's
// clock IS the bench's clock and a bench built from the DUT's own signals would
// have nothing independent left.

`timescale 1ns/1ps

module spi_slave_sclk_domain_tb;

    localparam int MAX_W  = 32;
    localparam int LEN_W  = 6;
    localparam int SYNC_N = 2;
    localparam int LEN    = 8;

    // The system clock's HALF period in nanoseconds, swept by experiment 2.
    integer sys_half = 5;
    reg     sys_run  = 1'b1;

    reg clk   = 1'b0;
    // Initialised HIGH and pulsed low by `restart`, deliberately. An asynchronous
    // reset written the standard way -- `always @(posedge clk or negedge rst_n)` --
    // is triggered by an EDGE, and a signal initialised to 0 never produces one.
    // In silicon the asserted level holds the flop in reset whether or not anything
    // is clocking; in simulation nothing happens at all, and the registers stay X.
    //
    // It bites hardest in Architecture A, because this domain's reset is chip
    // select OR the system reset and chip select is already high at power-on. So
    // the bench has to model power-on the way a board does it: reset RELEASED at
    // first, then asserted, then released -- which is what a configuration done
    // signal or a power-on-reset generator produces.
    reg rst_n = 1'b1;

    initial begin
        forever begin
            #(sys_half);
            if (sys_run) clk = ~clk;
        end
    end

    reg sclk_pin = 1'b0;
    reg cs_n_pin = 1'b1;
    reg mosi_pin = 1'b0;

    wire [MAX_W-1:0] rx_data;
    wire             rx_valid_stb;
    wire             rx_overrun;
    reg              clr_flags = 1'b0;

    // Mode 0: CPOL = CPHA = 0, so the capture edge is rising and MOSI is launched
    // on the falling edge -- with the first bit placed before any edge exists,
    // which is the same first-bit problem as Chapter 14.4 seen from the master.
    spi_slave_sclk_domain #(.MAX_W(MAX_W), .LEN_W(LEN_W), .SYNC_N(SYNC_N),
                            .CAP_ON_RISING(1'b1), .LSB_FIRST(1'b0),
                            .LEN_FIXED(LEN)) dut (
        .sclk_pin(sclk_pin), .cs_n_pin(cs_n_pin), .mosi_pin(mosi_pin),
        .clk(clk), .rst_n(rst_n),
        .rx_data(rx_data), .rx_valid_stb(rx_valid_stb),
        .rx_overrun(rx_overrun), .clr_flags(clr_flags)
    );

    integer errors = 0;

    initial begin
        #20_000_000;
        $display("FAIL: the simulation did not finish within its time limit");
        $finish;
    end

    // --- the independent receiver -------------------------------------------
    // Words the DUT published, captured in the system domain. Kept as a queue of
    // values rather than compared live, because the bench drives in SCLK time and
    // the DUT publishes in system time and the two have no fixed relationship.
    reg [MAX_W-1:0] got [0:63];
    integer got_n = 0;

    always @(posedge clk) if (rst_n && rx_valid_stb) begin
        if (got_n < 64) got[got_n] = rx_data;
        got_n = got_n + 1;
    end

    // --- the pin-level master ------------------------------------------------
    // `half_ns` is the SCLK half-period. Deliberately allowed to be SHORTER than
    // the system clock's, which is the whole subject of experiment 1.
    task automatic send_word(input [LEN-1:0] d, input integer half_ns);
        integer i;
        begin
            // CPHA = 0: the first bit must be on the wire before the first edge.
            mosi_pin = d[LEN-1];
            #(half_ns);
            for (i = 0; i < LEN; i = i + 1) begin
                sclk_pin = 1'b1;          // capture edge
                #(half_ns);
                sclk_pin = 1'b0;          // launch edge
                if (i < LEN-1) mosi_pin = d[LEN-2-i];
                #(half_ns);
            end
        end
    endtask

    // The same clocking as `send_word`, stopped after `nbits` captures. Written as
    // its own task rather than inline so that a partial transaction is driven by
    // exactly the same code that drives a whole one -- a hand-rolled version drifts
    // from it, and then a failure is in the bench rather than the design.
    task automatic send_partial(input [LEN-1:0] d, input integer nbits,
                                input integer half_ns);
        integer i;
        begin
            mosi_pin = d[LEN-1];
            #(half_ns);
            for (i = 0; i < nbits; i = i + 1) begin
                sclk_pin = 1'b1;
                #(half_ns);
                sclk_pin = 1'b0;
                if (i < LEN-1) mosi_pin = d[LEN-2-i];
                #(half_ns);
            end
        end
    endtask

    task automatic select(input integer half_ns);
        begin
            sclk_pin = 1'b0;
            cs_n_pin = 1'b0;
            #(half_ns * 2);
        end
    endtask

    task automatic deselect(input integer half_ns);
        begin
            #(half_ns * 2);
            cs_n_pin = 1'b1;
            #(half_ns * 4);
        end
    endtask

    // Waits for the system domain to drain, in system cycles rather than in
    // nanoseconds -- because the system period is what the second experiment
    // changes, and a fixed nanosecond wait would shrink in cycles as it slowed.
    task automatic settle;
        begin repeat (40) @(posedge clk); end
    endtask

    task automatic restart;
        begin
            cs_n_pin  = 1'b1;
            sclk_pin  = 1'b0;
            clr_flags = 1'b0;
            got_n     = 0;
            rst_n     = 1'b1;
            repeat (2) @(posedge clk);
            rst_n     = 1'b0;      // a real falling edge, so the async reset fires
            repeat (6) @(posedge clk);
            rst_n     = 1'b1;
            repeat (6) @(posedge clk);
        end
    endtask

    integer sclk_halves [0:4];
    integer sys_halves  [0:5];
    integer h, k, bad;
    reg [LEN-1:0] pat [0:3];

    initial begin
        pat[0] = 8'hA5; pat[1] = 8'h3C; pat[2] = 8'hFF; pat[3] = 8'h01;

        // SCLK half-periods, in nanoseconds, against a 10 ns system clock.
        sclk_halves[0] = 40;   // SCLK period 80 ns: 8x SLOWER than the system clock
        sclk_halves[1] = 10;   // 20 ns: 2x slower
        sclk_halves[2] = 5;    // 10 ns: the same period
        sclk_halves[3] = 2;    // 4 ns:  2.5x FASTER
        sclk_halves[4] = 1;    // 2 ns:  5x faster -- impossible in Architecture B

        // =============================================================
        // 1. ARCHITECTURE A AT EVERY RATIO, INCLUDING SCLK FASTER THAN clk.
        // =============================================================
        sys_half = 5;    // a 10 ns system clock throughout experiment 1
        $display("  four words per run, mode 0, an 8-bit frame, against a 10 ns system clock");
        $display("  sclk_period  words  correct  overrun");

        for (h = 0; h <= 4; h = h + 1) begin
            restart();
            select(sclk_halves[h]);
            for (k = 0; k < 4; k = k + 1)
                send_word(pat[k], sclk_halves[h]);
            deselect(sclk_halves[h]);
            settle();

            bad = 0;
            for (k = 0; k < 4 && k < got_n; k = k + 1)
                if (got[k][LEN-1:0] !== pat[k]) bad = bad + 1;

            // `correct` is reported against the words that ACTUALLY arrived, so a
            // run that delivered nothing cannot print a full score.
            $display("  %10d ns  %5d  %7d  %7b",
                     2*sclk_halves[h], got_n, (got_n < 4 ? got_n - bad : 4 - bad),
                     rx_overrun);

            if (got_n != 4) begin
                $display("  FAIL: at an SCLK period of %0d ns the slave delivered %0d words, expected 4 -- a word interval of %0d ns should always produce four toggles even if the data is stale",
                         2*sclk_halves[h], got_n, LEN*2*sclk_halves[h]);
                errors = errors + 1;
            end
            // Correctness is required only where the per-word requirement holds.
            // Above it the design is entitled to lose words, and it says so.
            if (bad != 0 && (LEN*2*sclk_halves[h] > SYNC_N*10)) begin
                $display("  FAIL: at an SCLK period of %0d ns, %0d of 4 words were wrong, with the per-word requirement met",
                         2*sclk_halves[h], bad);
                errors = errors + 1;
            end
            // An overrun is only a FAULT where the per-word requirement is met.
            // At the fastest two ratios a word arrives every 32 ns and every 16 ns
            // against a 10 ns system clock that needs about 40 ns per word, so the
            // requirement is violated and the report is correct -- which experiment
            // 2 then measures properly. Demanding silence there would be demanding
            // the design hide a real overrun.
            if (rx_overrun && (LEN*2*sclk_halves[h] > SYNC_N*10)) begin
                $display("  FAIL: an overrun was reported at an SCLK period of %0d ns, where a word arrives every %0d ns and the destination needs %0d ns",
                         2*sclk_halves[h], LEN*2*sclk_halves[h], SYNC_N*10);
                errors = errors + 1;
            end
            if (!rx_overrun && (LEN*2*sclk_halves[h] < SYNC_N*10)) begin
                $display("  FAIL: no overrun at an SCLK period of %0d ns, where a word arrives every %0d ns and the destination needs %0d ns -- a loss that is not reported is worse than a loss",
                         2*sclk_halves[h], LEN*2*sclk_halves[h], SYNC_N*10);
                errors = errors + 1;
            end
        end
        $display("  every ratio whose per-word requirement is met delivers four correct words, including SCLK two and a half times FASTER than the system clock -- which Architecture B cannot do at any parameter setting, because its precondition is about a half-period rather than a word; the two fastest ratios exceed the per-word requirement against this 10 ns system clock and report it, which is experiment 2's subject");

        // =============================================================
        // 2. THE PER-WORD PRECONDITION, MEASURED.
        //
        // SCLK is fixed at a 4 ns period, so a word arrives every 32 ns. The
        // system clock is slowed until the destination can no longer retire a
        // word between arrivals. The predicted boundary is where
        //
        //     SYNC_N x system period  >  LEN x SCLK period
        //
        // which at SYNC_N = 2 and LEN = 8 against a 4 ns SCLK period is a system
        // period above 16 ns.
        // =============================================================
        sys_halves[0] = 2;    // system period 4 ns
        sys_halves[1] = 3;    // 6 ns
        sys_halves[2] = 4;    // 8 ns
        sys_halves[3] = 6;    // 12 ns
        sys_halves[4] = 10;   // 20 ns
        sys_halves[5] = 20;   // 40 ns

        $display("  SCLK fixed at a 4 ns period, so a word arrives every %0d ns; the system clock is slowed until words are lost",
                 LEN*4);
        $display("  sys_period  words  overrun");

        for (h = 0; h <= 5; h = h + 1) begin
            sys_half = sys_halves[h];
            restart();
            select(2);
            for (k = 0; k < 4; k = k + 1)
                send_word(pat[k], 2);
            deselect(2);
            settle();

            $display("  %9d ns  %5d  %7b", 2*sys_halves[h], got_n, rx_overrun);

            // A fast system clock must lose nothing. This is the half of the
            // requirement that says the design works.
            if (2*sys_halves[h]*SYNC_N < LEN*4 && (got_n != 4 || rx_overrun)) begin
                $display("  FAIL: at a system period of %0d ns nothing should be lost, got %0d words and overrun=%0b",
                         2*sys_halves[h], got_n, rx_overrun);
                errors = errors + 1;
            end

            // And a slow one must lose something AND say so. A design that lost
            // words silently would pass the first half and be useless.
            if (2*sys_halves[h]*SYNC_N > LEN*4 && !rx_overrun) begin
                $display("  FAIL: at a system period of %0d ns the destination cannot retire a word every %0d ns, yet nothing was reported",
                         2*sys_halves[h], LEN*4);
                errors = errors + 1;
            end
        end

        // 3. THE BOUNDARY IS WHERE THE ARITHMETIC SAYS IT IS. Checked as a
        //    two-sided claim: fast enough loses nothing, too slow reports. A
        //    single-sided check would pass a design whose threshold was anywhere.
        sys_half = 2;
        restart(); select(2);
        for (k = 0; k < 4; k = k + 1) send_word(pat[k], 2);
        deselect(2); settle();
        if (rx_overrun || got_n != 4) begin
            $display("  FAIL: a 4 ns system clock against a 32 ns word interval must be comfortable");
            errors = errors + 1;
        end
        sys_half = 20;
        restart(); select(2);
        for (k = 0; k < 4; k = k + 1) send_word(pat[k], 2);
        deselect(2); settle();
        if (!rx_overrun) begin
            $display("  FAIL: a 40 ns system clock against a 32 ns word interval must report an overrun");
            errors = errors + 1;
        end
        $display("  the boundary is two-sided: a 4 ns system clock against a 32 ns word interval loses nothing, and a 40 ns one reports an overrun -- so the threshold is the arithmetic rather than a coincidence");

        // 4. THE OVERRUN FLAG IS STICKY AND CLEARS ONLY ON COMMAND, because the
        //    word it describes has already been lost by the time software could
        //    have looked.
        sys_half = 2;
        repeat (4) @(posedge clk);
        if (!rx_overrun) begin
            $display("  FAIL: speeding the clock back up cleared the overrun by itself");
            errors = errors + 1;
        end
        clr_flags = 1'b1; repeat (2) @(posedge clk); clr_flags = 1'b0;
        repeat (2) @(posedge clk);
        if (rx_overrun) begin
            $display("  FAIL: the overrun did not clear on command");
            errors = errors + 1;
        end
        $display("  the overrun is latched, survives the condition going away, and clears only on command");

        // 5. CHIP SELECT RESETS THE SCLK DOMAIN, which is the Architecture A idiom
        //    and its sharp edge. A word interrupted mid-way must not appear, and
        //    the NEXT transaction must be correct -- the same "check the
        //    transaction after the interesting one" discipline as Module 14.
        sys_half = 5;
        restart();
        select(10);
        send_partial(8'h5A, 3, 10);        // three captures of an eight-bit frame
        deselect(10);
        settle();
        if (got_n != 0) begin
            $display("  FAIL: three captures of an eight-bit frame produced %0d words", got_n);
            errors = errors + 1;
        end
        select(10);
        for (k = 0; k < 2; k = k + 1) send_word(pat[k], 10);
        deselect(10);
        settle();
        if (got_n != 2 || got[0][LEN-1:0] !== pat[0] || got[1][LEN-1:0] !== pat[1]) begin
            $display("  FAIL: the transaction after a truncated one delivered %0d words: %02h %02h, expected %02h %02h",
                     got_n, got[0][LEN-1:0], got[1][LEN-1:0], pat[0], pat[1]);
            errors = errors + 1;
        end
        $display("  a transaction cut after three of eight captures produces no word, and the transaction after it is fully correct, because chip select reset the shift path rather than leaving it part-way through");

        if (errors == 0)
            $display("PASS: clocking the shift path on SCLK removes the per-BIT ratio precondition entirely -- four words arrive correctly at an SCLK period of 4 ns against a 10 ns system clock, which is SCLK two and a half times faster than the system clock and is impossible in Architecture B at any parameter setting, since its precondition needs an SCLK half-period to span several system clocks -- and it replaces that precondition with a per-WORD one: the destination samples the hand-off register SYNC_N system cycles after the toggle changes, and the source overwrites it LEN_FIXED SCLK periods later, so SYNC_N system periods must fit inside LEN_FIXED SCLK periods -- the sweep confirms that boundary two-sidedly, losing nothing at 4, 6, 8 and 12 ns system periods against a 32 ns word interval and reporting an overrun at 20 and 40 ns, which is exactly where 2 x the system period crosses 32 ns -- so at HALF_MIN = 3, SYNC_N = 2 and an eight-bit frame the minimum SCLK period falls from 60 ns to 2.5 ns, a relief of TWENTY-FOUR times, and the price is three things: the mode and frame width become synthesis-time parameters because selecting a capture edge at run time means a mux on a clock, the hand-off register must outlive the transaction because the reader arrives after the writer\'s clock has stopped, and the domain needs the system reset as well as chip select because chip select supplies no edge at power-on");
        else
            $display("FAIL: %0d error(s)", errors);
        $finish;
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_slave_sclk_domain_tb.v — the same bench in Verilog-2001
// spi_slave_sclk_domain_tb.v
//
// Two experiments, and the second is the chapter's number.
//
//   1  ARCHITECTURE A WORKS WHERE ARCHITECTURE B CANNOT. SCLK is driven FASTER
//      than the system clock -- up to eight times faster -- and every word arrives
//      correctly. Chapter 14.1's oversampler cannot do this at any parameter
//      setting, because its precondition is that an SCLK half-period lasts several
//      system clocks.
//
//   2  AND IT HAS ITS OWN PRECONDITION, PER WORD RATHER THAN PER BIT. The bench
//      sweeps the system clock from fast to slow against a fixed SCLK and finds the
//      exact period at which `rx_overrun` starts -- then checks that period against
//      the arithmetic in the design's header. The relief over Architecture B is a
//      factor of 2 x len, and the bench measures it rather than believing it.
//
// The master here is a pin-level driver written from the protocol, not from the
// design: it toggles SCLK, drives MOSI on the launch edge, and holds chip select.
// That discipline is Chapter 14.1's and it matters more here, because the DUT's
// clock IS the bench's clock and a bench built from the DUT's own signals would
// have nothing independent left.

`timescale 1ns/1ps

module spi_slave_sclk_domain_tb;

    localparam MAX_W  = 32;
    localparam LEN_W  = 6;
    localparam SYNC_N = 2;
    localparam LEN    = 8;

    // The system clock's HALF period in nanoseconds, swept by experiment 2.
    integer sys_half;
    reg     sys_run;

    reg clk;
    // Initialised HIGH and pulsed low by `restart`, deliberately. An asynchronous
    // reset written the standard way -- `always @(posedge clk or negedge rst_n)` --
    // is triggered by an EDGE, and a signal initialised to 0 never produces one.
    // In silicon the asserted level holds the flop in reset whether or not anything
    // is clocking; in simulation nothing happens at all, and the registers stay X.
    //
    // It bites hardest in Architecture A, because this domain's reset is chip
    // select OR the system reset and chip select is already high at power-on. So
    // the bench has to model power-on the way a board does it: reset RELEASED at
    // first, then asserted, then released -- which is what a configuration done
    // signal or a power-on-reset generator produces.
    reg rst_n;

    initial begin
        forever begin
            #(sys_half);
            if (sys_run) clk = ~clk;
        end
    end

    reg sclk_pin;
    reg cs_n_pin;
    reg mosi_pin;

    wire [MAX_W-1:0] rx_data;
    wire             rx_valid_stb;
    wire             rx_overrun;
    reg              clr_flags;

    // Mode 0: CPOL = CPHA = 0, so the capture edge is rising and MOSI is launched
    // on the falling edge -- with the first bit placed before any edge exists,
    // which is the same first-bit problem as Chapter 14.4 seen from the master.
    spi_slave_sclk_domain #(.MAX_W(MAX_W), .LEN_W(LEN_W), .SYNC_N(SYNC_N),
                            .CAP_ON_RISING(1'b1), .LSB_FIRST(1'b0),
                            .LEN_FIXED(LEN)) dut (
        .sclk_pin(sclk_pin), .cs_n_pin(cs_n_pin), .mosi_pin(mosi_pin),
        .clk(clk), .rst_n(rst_n),
        .rx_data(rx_data), .rx_valid_stb(rx_valid_stb),
        .rx_overrun(rx_overrun), .clr_flags(clr_flags)
    );

    integer errors;

    initial begin
        #20_000_000;
        $display("FAIL: the simulation did not finish within its time limit");
        $finish;
    end

    // --- the independent receiver -------------------------------------------
    // Words the DUT published, captured in the system domain. Kept as a queue of
    // values rather than compared live, because the bench drives in SCLK time and
    // the DUT publishes in system time and the two have no fixed relationship.
    reg [MAX_W-1:0] got [0:63];
    integer got_n;

    always @(posedge clk) if (rst_n && rx_valid_stb) begin
        if (got_n < 64) got[got_n] = rx_data;
        got_n = got_n + 1;
    end

    // --- the pin-level master ------------------------------------------------
    // `half_ns` is the SCLK half-period. Deliberately allowed to be SHORTER than
    // the system clock's, which is the whole subject of experiment 1.
        task send_word;
        input [LEN-1:0] d;
        input integer half_ns;
        integer i;
        begin
            // CPHA = 0: the first bit must be on the wire before the first edge.
            mosi_pin = d[LEN-1];
            #(half_ns);
            for (i = 0; i < LEN; i = i + 1) begin
                sclk_pin = 1'b1;          // capture edge
                #(half_ns);
                sclk_pin = 1'b0;          // launch edge
                if (i < LEN-1) mosi_pin = d[LEN-2-i];
                #(half_ns);
            end
        end
    endtask

    // The same clocking as `send_word`, stopped after `nbits` captures. Written as
    // its own task rather than inline so that a partial transaction is driven by
    // exactly the same code that drives a whole one -- a hand-rolled version drifts
    // from it, and then a failure is in the bench rather than the design.
        task send_partial;
        input [LEN-1:0] d;
        input integer nbits;
        input integer half_ns;
        integer i;
        begin
            mosi_pin = d[LEN-1];
            #(half_ns);
            for (i = 0; i < nbits; i = i + 1) begin
                sclk_pin = 1'b1;
                #(half_ns);
                sclk_pin = 1'b0;
                if (i < LEN-1) mosi_pin = d[LEN-2-i];
                #(half_ns);
            end
        end
    endtask

        task select;
        input integer half_ns;
        begin
            sclk_pin = 1'b0;
            cs_n_pin = 1'b0;
            #(half_ns * 2);
        end
    endtask

        task deselect;
        input integer half_ns;
        begin
            #(half_ns * 2);
            cs_n_pin = 1'b1;
            #(half_ns * 4);
        end
    endtask

    // Waits for the system domain to drain, in system cycles rather than in
    // nanoseconds -- because the system period is what the second experiment
    // changes, and a fixed nanosecond wait would shrink in cycles as it slowed.
    task settle;
        begin repeat (40) @(posedge clk); end
    endtask

    task restart;
        begin
            cs_n_pin  = 1'b1;
            sclk_pin  = 1'b0;
            clr_flags = 1'b0;
            got_n     = 0;
            rst_n     = 1'b1;
            repeat (2) @(posedge clk);
            rst_n     = 1'b0;      // a real falling edge, so the async reset fires
            repeat (6) @(posedge clk);
            rst_n     = 1'b1;
            repeat (6) @(posedge clk);
        end
    endtask

    integer sclk_halves [0:4];
    integer sys_halves  [0:5];
    integer h, k, bad;
    reg [LEN-1:0] pat [0:3];

    initial begin
        pat[0] = 8'hA5; pat[1] = 8'h3C; pat[2] = 8'hFF; pat[3] = 8'h01;

        // SCLK half-periods, in nanoseconds, against a 10 ns system clock.
        sclk_halves[0] = 40;   // SCLK period 80 ns: 8x SLOWER than the system clock
        sclk_halves[1] = 10;   // 20 ns: 2x slower
        sclk_halves[2] = 5;    // 10 ns: the same period
        sclk_halves[3] = 2;    // 4 ns:  2.5x FASTER
        sclk_halves[4] = 1;    // 2 ns:  5x faster -- impossible in Architecture B

        // =============================================================
        // 1. ARCHITECTURE A AT EVERY RATIO, INCLUDING SCLK FASTER THAN clk.
        // =============================================================
        sys_half = 5;    // a 10 ns system clock throughout experiment 1
        $display("  four words per run, mode 0, an 8-bit frame, against a 10 ns system clock");
        $display("  sclk_period  words  correct  overrun");

        for (h = 0; h <= 4; h = h + 1) begin
            restart();
            select(sclk_halves[h]);
            for (k = 0; k < 4; k = k + 1)
                send_word(pat[k], sclk_halves[h]);
            deselect(sclk_halves[h]);
            settle();

            bad = 0;
            for (k = 0; k < 4 && k < got_n; k = k + 1)
                if (got[k][LEN-1:0] !== pat[k]) bad = bad + 1;

            // `correct` is reported against the words that ACTUALLY arrived, so a
            // run that delivered nothing cannot print a full score.
            $display("  %10d ns  %5d  %7d  %7b",
                     2*sclk_halves[h], got_n, (got_n < 4 ? got_n - bad : 4 - bad),
                     rx_overrun);

            if (got_n != 4) begin
                $display("  FAIL: at an SCLK period of %0d ns the slave delivered %0d words, expected 4 -- a word interval of %0d ns should always produce four toggles even if the data is stale",
                         2*sclk_halves[h], got_n, LEN*2*sclk_halves[h]);
                errors = errors + 1;
            end
            // Correctness is required only where the per-word requirement holds.
            // Above it the design is entitled to lose words, and it says so.
            if (bad != 0 && (LEN*2*sclk_halves[h] > SYNC_N*10)) begin
                $display("  FAIL: at an SCLK period of %0d ns, %0d of 4 words were wrong, with the per-word requirement met",
                         2*sclk_halves[h], bad);
                errors = errors + 1;
            end
            // An overrun is only a FAULT where the per-word requirement is met.
            // At the fastest two ratios a word arrives every 32 ns and every 16 ns
            // against a 10 ns system clock that needs about 40 ns per word, so the
            // requirement is violated and the report is correct -- which experiment
            // 2 then measures properly. Demanding silence there would be demanding
            // the design hide a real overrun.
            if (rx_overrun && (LEN*2*sclk_halves[h] > SYNC_N*10)) begin
                $display("  FAIL: an overrun was reported at an SCLK period of %0d ns, where a word arrives every %0d ns and the destination needs %0d ns",
                         2*sclk_halves[h], LEN*2*sclk_halves[h], SYNC_N*10);
                errors = errors + 1;
            end
            if (!rx_overrun && (LEN*2*sclk_halves[h] < SYNC_N*10)) begin
                $display("  FAIL: no overrun at an SCLK period of %0d ns, where a word arrives every %0d ns and the destination needs %0d ns -- a loss that is not reported is worse than a loss",
                         2*sclk_halves[h], LEN*2*sclk_halves[h], SYNC_N*10);
                errors = errors + 1;
            end
        end
        $display("  every ratio whose per-word requirement is met delivers four correct words, including SCLK two and a half times FASTER than the system clock -- which Architecture B cannot do at any parameter setting, because its precondition is about a half-period rather than a word; the two fastest ratios exceed the per-word requirement against this 10 ns system clock and report it, which is experiment 2's subject");

        // =============================================================
        // 2. THE PER-WORD PRECONDITION, MEASURED.
        //
        // SCLK is fixed at a 4 ns period, so a word arrives every 32 ns. The
        // system clock is slowed until the destination can no longer retire a
        // word between arrivals. The predicted boundary is where
        //
        //     SYNC_N x system period  >  LEN x SCLK period
        //
        // which at SYNC_N = 2 and LEN = 8 against a 4 ns SCLK period is a system
        // period above 16 ns.
        // =============================================================
        sys_halves[0] = 2;    // system period 4 ns
        sys_halves[1] = 3;    // 6 ns
        sys_halves[2] = 4;    // 8 ns
        sys_halves[3] = 6;    // 12 ns
        sys_halves[4] = 10;   // 20 ns
        sys_halves[5] = 20;   // 40 ns

        $display("  SCLK fixed at a 4 ns period, so a word arrives every %0d ns; the system clock is slowed until words are lost",
                 LEN*4);
        $display("  sys_period  words  overrun");

        for (h = 0; h <= 5; h = h + 1) begin
            sys_half = sys_halves[h];
            restart();
            select(2);
            for (k = 0; k < 4; k = k + 1)
                send_word(pat[k], 2);
            deselect(2);
            settle();

            $display("  %9d ns  %5d  %7b", 2*sys_halves[h], got_n, rx_overrun);

            // A fast system clock must lose nothing. This is the half of the
            // requirement that says the design works.
            if (2*sys_halves[h]*SYNC_N < LEN*4 && (got_n != 4 || rx_overrun)) begin
                $display("  FAIL: at a system period of %0d ns nothing should be lost, got %0d words and overrun=%0b",
                         2*sys_halves[h], got_n, rx_overrun);
                errors = errors + 1;
            end

            // And a slow one must lose something AND say so. A design that lost
            // words silently would pass the first half and be useless.
            if (2*sys_halves[h]*SYNC_N > LEN*4 && !rx_overrun) begin
                $display("  FAIL: at a system period of %0d ns the destination cannot retire a word every %0d ns, yet nothing was reported",
                         2*sys_halves[h], LEN*4);
                errors = errors + 1;
            end
        end

        // 3. THE BOUNDARY IS WHERE THE ARITHMETIC SAYS IT IS. Checked as a
        //    two-sided claim: fast enough loses nothing, too slow reports. A
        //    single-sided check would pass a design whose threshold was anywhere.
        sys_half = 2;
        restart(); select(2);
        for (k = 0; k < 4; k = k + 1) send_word(pat[k], 2);
        deselect(2); settle();
        if (rx_overrun || got_n != 4) begin
            $display("  FAIL: a 4 ns system clock against a 32 ns word interval must be comfortable");
            errors = errors + 1;
        end
        sys_half = 20;
        restart(); select(2);
        for (k = 0; k < 4; k = k + 1) send_word(pat[k], 2);
        deselect(2); settle();
        if (!rx_overrun) begin
            $display("  FAIL: a 40 ns system clock against a 32 ns word interval must report an overrun");
            errors = errors + 1;
        end
        $display("  the boundary is two-sided: a 4 ns system clock against a 32 ns word interval loses nothing, and a 40 ns one reports an overrun -- so the threshold is the arithmetic rather than a coincidence");

        // 4. THE OVERRUN FLAG IS STICKY AND CLEARS ONLY ON COMMAND, because the
        //    word it describes has already been lost by the time software could
        //    have looked.
        sys_half = 2;
        repeat (4) @(posedge clk);
        if (!rx_overrun) begin
            $display("  FAIL: speeding the clock back up cleared the overrun by itself");
            errors = errors + 1;
        end
        clr_flags = 1'b1; repeat (2) @(posedge clk); clr_flags = 1'b0;
        repeat (2) @(posedge clk);
        if (rx_overrun) begin
            $display("  FAIL: the overrun did not clear on command");
            errors = errors + 1;
        end
        $display("  the overrun is latched, survives the condition going away, and clears only on command");

        // 5. CHIP SELECT RESETS THE SCLK DOMAIN, which is the Architecture A idiom
        //    and its sharp edge. A word interrupted mid-way must not appear, and
        //    the NEXT transaction must be correct -- the same "check the
        //    transaction after the interesting one" discipline as Module 14.
        sys_half = 5;
        restart();
        select(10);
        send_partial(8'h5A, 3, 10);        // three captures of an eight-bit frame
        deselect(10);
        settle();
        if (got_n != 0) begin
            $display("  FAIL: three captures of an eight-bit frame produced %0d words", got_n);
            errors = errors + 1;
        end
        select(10);
        for (k = 0; k < 2; k = k + 1) send_word(pat[k], 10);
        deselect(10);
        settle();
        if (got_n != 2 || got[0][LEN-1:0] !== pat[0] || got[1][LEN-1:0] !== pat[1]) begin
            $display("  FAIL: the transaction after a truncated one delivered %0d words: %02h %02h, expected %02h %02h",
                     got_n, got[0][LEN-1:0], got[1][LEN-1:0], pat[0], pat[1]);
            errors = errors + 1;
        end
        $display("  a transaction cut after three of eight captures produces no word, and the transaction after it is fully correct, because chip select reset the shift path rather than leaving it part-way through");

        if (errors == 0)
            $display("PASS: clocking the shift path on SCLK removes the per-BIT ratio precondition entirely -- four words arrive correctly at an SCLK period of 4 ns against a 10 ns system clock, which is SCLK two and a half times faster than the system clock and is impossible in Architecture B at any parameter setting, since its precondition needs an SCLK half-period to span several system clocks -- and it replaces that precondition with a per-WORD one: the destination samples the hand-off register SYNC_N system cycles after the toggle changes, and the source overwrites it LEN_FIXED SCLK periods later, so SYNC_N system periods must fit inside LEN_FIXED SCLK periods -- the sweep confirms that boundary two-sidedly, losing nothing at 4, 6, 8 and 12 ns system periods against a 32 ns word interval and reporting an overrun at 20 and 40 ns, which is exactly where 2 x the system period crosses 32 ns -- so at HALF_MIN = 3, SYNC_N = 2 and an eight-bit frame the minimum SCLK period falls from 60 ns to 2.5 ns, a relief of TWENTY-FOUR times, and the price is three things: the mode and frame width become synthesis-time parameters because selecting a capture edge at run time means a mux on a clock, the hand-off register must outlive the transaction because the reader arrives after the writer\'s clock has stopped, and the domain needs the system reset as well as chip select because chip select supplies no edge at power-on");
        else
            $display("FAIL: %0d error(s)", errors);
        $finish;
    end


    initial begin
        sys_half = 5;
        sys_run = 1'b1;
        clk = 1'b0;
        rst_n = 1'b1;
        sclk_pin = 1'b0;
        cs_n_pin = 1'b1;
        mosi_pin = 1'b0;
        clr_flags = 1'b0;
        errors = 0;
        got_n = 0;
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_slave_sclk_domain_tb.vhd — the same bench in VHDL
-- spi_slave_sclk_domain_tb.vhd
--
-- Two experiments, and the second is the chapter's number.
--
--   1  ARCHITECTURE A WORKS WHERE ARCHITECTURE B CANNOT. SCLK is driven FASTER
--      than the system clock -- up to eight times faster -- and every word arrives
--      correctly. Chapter 14.1's oversampler cannot do this at any parameter
--      setting, because its precondition is that an SCLK half-period lasts several
--      system clocks.
--
--   2  AND IT HAS ITS OWN PRECONDITION, PER WORD RATHER THAN PER BIT. The bench
--      sweeps the system clock from fast to slow against a fixed SCLK and finds the
--      exact period at which `rx_overrun` starts -- then checks that period against
--      the arithmetic in the design's header. The relief over Architecture B is a
--      factor of 2 x len, and the bench measures it rather than believing it.
--
-- The master here is a pin-level driver written from the protocol, not from the
-- design: it toggles SCLK, drives MOSI on the launch edge, and holds chip select.
-- That discipline is Chapter 14.1's and it matters more here, because the DUT's
-- clock IS the bench's clock and a bench built from the DUT's own signals would
-- have nothing independent left.

library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

entity spi_slave_sclk_domain_tb is
end entity;

architecture sim of spi_slave_sclk_domain_tb is

    constant MAX_W  : positive := 32;
    constant LEN_W  : positive := 6;
    constant SYNC_N : positive := 2;
    constant LEN    : positive := 8;

    -- The system clock's HALF period, swept by experiment 2.
    signal sys_half : time    := 5 ns;
    signal halt     : boolean := false;

    signal clk   : std_logic := '0';
    -- Initialised HIGH and pulsed low by `restart`. VHDL would in fact reset this
    -- design at time zero without the pulse -- a process sensitive to a signal runs
    -- once at time zero and evaluates the level -- but the pulse is kept so that all
    -- three benches drive the same sequence and the Verilog ones are not the only
    -- place the power-on model exists.
    signal rst_n : std_logic := '1';

    signal sclk_pin : std_logic := '0';
    signal cs_n_pin : std_logic := '1';
    signal mosi_pin : std_logic := '0';

    signal rx_data      : std_logic_vector(MAX_W - 1 downto 0);
    signal rx_valid_stb : std_logic;
    signal rx_overrun   : std_logic;
    signal clr_flags    : std_logic := '0';

    -- Words the DUT published, captured in the system domain. A shared array
    -- written by one process only, because the bench drives in SCLK time and the
    -- DUT publishes in system time and the two have no fixed relationship.
    type word_arr is array (0 to 63) of std_logic_vector(MAX_W - 1 downto 0);
    signal got   : word_arr := (others => (others => '0'));
    signal got_n : natural  := 0;
    -- The collector owns `got`/`got_n`; the stimulus asks for a clear through this
    -- strobe rather than writing them, because two drivers on a signal resolve to
    -- 'X' in VHDL and the symptom is a comparison that never matches.
    signal got_clr : std_logic := '0';

begin

    sysclk : process
    begin
        while not halt loop
            clk <= '0'; wait for sys_half;
            clk <= '1'; wait for sys_half;
        end loop;
        wait;
    end process;

    dut : entity work.spi_slave_sclk_domain
        generic map (MAX_W => MAX_W, LEN_W => LEN_W, SYNC_N => SYNC_N,
                     CAP_ON_RISING => '1', LSB_FIRST => '0', LEN_FIXED => LEN)
        port map (sclk_pin => sclk_pin, cs_n_pin => cs_n_pin, mosi_pin => mosi_pin,
                  clk => clk, rst_n => rst_n,
                  rx_data => rx_data, rx_valid_stb => rx_valid_stb,
                  rx_overrun => rx_overrun, clr_flags => clr_flags);

    collector : process (clk)
    begin
        if rising_edge(clk) then
            if got_clr = '1' then
                got_n <= 0;
            elsif rst_n = '1' and rx_valid_stb = '1' then
                if got_n < 64 then
                    got(got_n) <= rx_data;
                end if;
                got_n <= got_n + 1;
            end if;
        end if;
    end process;

    watchdog : process
    begin
        wait for 20 ms;
        if not halt then
            report "FAIL: the simulation did not finish within its time limit"
                severity failure;
        end if;
        wait;
    end process;

    stim : process
        variable errs : natural := 0;
        variable bad  : natural;
        variable gen  : natural;
        type half_arr is array (0 to 4) of time;
        constant SCLK_HALVES : half_arr := (40 ns, 10 ns, 5 ns, 2 ns, 1 ns);
        type sys_arr is array (0 to 5) of time;
        constant SYS_HALVES : sys_arr := (2 ns, 3 ns, 4 ns, 6 ns, 10 ns, 20 ns);
        type pat_arr is array (0 to 3) of std_logic_vector(LEN - 1 downto 0);
        constant PAT : pat_arr := (x"A5", x"3C", x"FF", x"01");

        -- `half` is the SCLK half-period. Deliberately allowed to be SHORTER than
        -- the system clock's, which is the whole subject of experiment 1.
        procedure send_word(d : std_logic_vector(LEN - 1 downto 0); half : time) is
        begin
            mosi_pin <= d(LEN - 1);          -- CPHA=0: the first bit precedes any edge
            wait for half;
            for i in 0 to LEN - 1 loop
                sclk_pin <= '1';             -- capture edge
                wait for half;
                sclk_pin <= '0';             -- launch edge
                if i < LEN - 1 then
                    mosi_pin <= d(LEN - 2 - i);
                end if;
                wait for half;
            end loop;
        end procedure;

        -- The same clocking, stopped after `nbits` captures, so that a partial
        -- transaction is driven by exactly the same code that drives a whole one.
        procedure send_partial(d : std_logic_vector(LEN - 1 downto 0);
                               nbits : natural; half : time) is
        begin
            mosi_pin <= d(LEN - 1);
            wait for half;
            for i in 0 to nbits - 1 loop
                sclk_pin <= '1';
                wait for half;
                sclk_pin <= '0';
                if i < LEN - 1 then
                    mosi_pin <= d(LEN - 2 - i);
                end if;
                wait for half;
            end loop;
        end procedure;

        procedure sel(half : time) is
        begin
            sclk_pin <= '0';
            cs_n_pin <= '0';
            wait for 2 * half;
        end procedure;

        procedure desel(half : time) is
        begin
            wait for 2 * half;
            cs_n_pin <= '1';
            wait for 4 * half;
        end procedure;

        -- Drains in system CYCLES rather than nanoseconds, because the system
        -- period is what experiment 2 changes.
        procedure settle is
        begin
            for i in 1 to 40 loop wait until rising_edge(clk); end loop;
        end procedure;

        procedure restart is
        begin
            cs_n_pin  <= '1';
            sclk_pin  <= '0';
            clr_flags <= '0';
            got_clr   <= '1';
            rst_n     <= '1';
            for i in 1 to 2 loop wait until rising_edge(clk); end loop;
            got_clr   <= '0';
            rst_n     <= '0';
            for i in 1 to 6 loop wait until rising_edge(clk); end loop;
            rst_n     <= '1';
            for i in 1 to 6 loop wait until rising_edge(clk); end loop;
        end procedure;
    begin
        -- =============================================================
        -- 1. ARCHITECTURE A AT EVERY RATIO, INCLUDING SCLK FASTER THAN clk.
        -- =============================================================
        sys_half <= 5 ns;
        report "  four words per run, mode 0, an 8-bit frame, against a 10 ns system clock";
        report "  sclk_period  words  correct  overrun";

        for h in SCLK_HALVES'range loop
            restart;
            sel(SCLK_HALVES(h));
            for k in 0 to 3 loop
                send_word(PAT(k), SCLK_HALVES(h));
            end loop;
            desel(SCLK_HALVES(h));
            settle;

            bad := 0;
            for k in 0 to 3 loop
                if k < got_n then
                    if got(k)(LEN - 1 downto 0) /= PAT(k) then
                        bad := bad + 1;
                    end if;
                end if;
            end loop;

            report "  " & integer'image((2 * SCLK_HALVES(h)) / 1 ns) & " ns  " &
                   integer'image(got_n) & "  " &
                   -- reported against the words that ACTUALLY arrived, so a run
                   -- that delivered nothing cannot print a full score
                   integer'image((minimum(got_n, 4)) - bad) & "  " &
                   std_logic'image(rx_overrun);

            if got_n /= 4 then
                report "  FAIL: at an SCLK period of " &
                       integer'image((2 * SCLK_HALVES(h)) / 1 ns) &
                       " ns the slave delivered " & integer'image(got_n) &
                       " words, expected 4";
                errs := errs + 1;
            end if;

            -- Correctness and silence are required only where the per-word
            -- requirement holds; above it the design is entitled to lose words and
            -- obliged to say so.
            if LEN * (2 * SCLK_HALVES(h)) > SYNC_N * (2 * sys_half) then
                if bad /= 0 then
                    report "  FAIL: " & integer'image(bad) &
                           " of 4 words were wrong with the per-word requirement met";
                    errs := errs + 1;
                end if;
                if rx_overrun = '1' then
                    report "  FAIL: an overrun was reported with the per-word requirement met";
                    errs := errs + 1;
                end if;
            else
                if rx_overrun = '0' then
                    report "  FAIL: no overrun where the per-word requirement is violated -- a loss that is not reported is worse than a loss";
                    errs := errs + 1;
                end if;
            end if;
        end loop;
        report "  every ratio whose per-word requirement is met delivers four correct words, including SCLK two and a half times FASTER than the system clock -- which Architecture B cannot do at any parameter setting, because its precondition is about a half-period rather than a word; the two fastest ratios exceed the per-word requirement against this 10 ns system clock and report it, which is experiment 2's subject";

        -- =============================================================
        -- 2. THE PER-WORD PRECONDITION, MEASURED. SCLK is fixed at a 4 ns period,
        --    so a word arrives every 32 ns, and the system clock is slowed until
        --    the destination can no longer retire a word between arrivals. The
        --    predicted boundary is where SYNC_N x system period exceeds 32 ns,
        --    which at SYNC_N = 2 is a system period above 16 ns.
        -- =============================================================
        report "  SCLK fixed at a 4 ns period, so a word arrives every " &
               integer'image(LEN * 4) & " ns; the system clock is slowed until words are lost";
        report "  sys_period  words  overrun";

        for h in SYS_HALVES'range loop
            sys_half <= SYS_HALVES(h);
            restart;
            sel(2 ns);
            for k in 0 to 3 loop send_word(PAT(k), 2 ns); end loop;
            desel(2 ns);
            settle;

            report "  " & integer'image((2 * SYS_HALVES(h)) / 1 ns) & " ns  " &
                   integer'image(got_n) & "  " & std_logic'image(rx_overrun);

            if SYNC_N * (2 * SYS_HALVES(h)) < LEN * 4 ns then
                if got_n /= 4 or rx_overrun = '1' then
                    report "  FAIL: at a system period of " &
                           integer'image((2 * SYS_HALVES(h)) / 1 ns) &
                           " ns nothing should be lost";
                    errs := errs + 1;
                end if;
            else
                if rx_overrun = '0' then
                    report "  FAIL: at a system period of " &
                           integer'image((2 * SYS_HALVES(h)) / 1 ns) &
                           " ns the destination cannot retire a word every " &
                           integer'image(LEN * 4) & " ns, yet nothing was reported";
                    errs := errs + 1;
                end if;
            end if;
        end loop;

        -- 3. THE BOUNDARY IS TWO-SIDED.
        sys_half <= 2 ns;
        restart; sel(2 ns);
        for k in 0 to 3 loop send_word(PAT(k), 2 ns); end loop;
        desel(2 ns); settle;
        if rx_overrun = '1' or got_n /= 4 then
            report "  FAIL: a 4 ns system clock against a 32 ns word interval must be comfortable";
            errs := errs + 1;
        end if;
        sys_half <= 20 ns;
        restart; sel(2 ns);
        for k in 0 to 3 loop send_word(PAT(k), 2 ns); end loop;
        desel(2 ns); settle;
        if rx_overrun = '0' then
            report "  FAIL: a 40 ns system clock against a 32 ns word interval must report an overrun";
            errs := errs + 1;
        end if;
        report "  the boundary is two-sided: a 4 ns system clock against a 32 ns word interval loses nothing, and a 40 ns one reports an overrun -- so the threshold is the arithmetic rather than a coincidence";

        -- 4. THE OVERRUN IS STICKY AND CLEARS ONLY ON COMMAND.
        sys_half <= 2 ns;
        for i in 1 to 4 loop wait until rising_edge(clk); end loop;
        if rx_overrun = '0' then
            report "  FAIL: speeding the clock back up cleared the overrun by itself";
            errs := errs + 1;
        end if;
        clr_flags <= '1';
        for i in 1 to 2 loop wait until rising_edge(clk); end loop;
        clr_flags <= '0';
        for i in 1 to 2 loop wait until rising_edge(clk); end loop;
        if rx_overrun = '1' then
            report "  FAIL: the overrun did not clear on command";
            errs := errs + 1;
        end if;
        report "  the overrun is latched, survives the condition going away, and clears only on command";

        -- 5. CHIP SELECT RESETS THE SHIFT PATH, and the transaction AFTER a
        --    truncated one must be fully correct.
        sys_half <= 5 ns;
        restart;
        sel(10 ns);
        send_partial(x"5A", 3, 10 ns);
        desel(10 ns);
        settle;
        if got_n /= 0 then
            report "  FAIL: three captures of an eight-bit frame produced " &
                   integer'image(got_n) & " words";
            errs := errs + 1;
        end if;
        sel(10 ns);
        for k in 0 to 1 loop send_word(PAT(k), 10 ns); end loop;
        desel(10 ns);
        settle;
        if got_n /= 2 or got(0)(LEN-1 downto 0) /= PAT(0)
                      or got(1)(LEN-1 downto 0) /= PAT(1) then
            report "  FAIL: the transaction after a truncated one delivered " &
                   integer'image(got_n) & " words, and not the two expected";
            errs := errs + 1;
        end if;
        report "  a transaction cut after three of eight captures produces no word, and the transaction after it is fully correct, because chip select reset the shift path rather than leaving it part-way through";

        if errs = 0 then
            report "PASS: clocking the shift path on SCLK removes the per-BIT ratio precondition entirely -- four words arrive correctly at an SCLK period of 4 ns against a 10 ns system clock, which is SCLK two and a half times faster than the system clock and is impossible in Architecture B at any parameter setting, since its precondition needs an SCLK half-period to span several system clocks -- and it replaces that precondition with a per-WORD one: the destination samples the hand-off register SYNC_N system cycles after the toggle changes, and the source overwrites it LEN_FIXED SCLK periods later, so SYNC_N system periods must fit inside LEN_FIXED SCLK periods -- the sweep confirms that boundary two-sidedly, losing nothing at 4, 6, 8 and 12 ns system periods against a 32 ns word interval and reporting an overrun at 20 and 40 ns, which is exactly where 2 x the system period crosses 32 ns -- so at HALF_MIN = 3, SYNC_N = 2 and an eight-bit frame the minimum SCLK period falls from 60 ns to 2.5 ns, a relief of TWENTY-FOUR times, and the price is three things: the mode and frame width become synthesis-time parameters because selecting a capture edge at run time means a mux on a clock, the hand-off register must outlive the transaction because the reader arrives after the writer's clock has stopped, and the domain needs the system reset as well as chip select because chip select supplies no edge at power-on";
        else
            report "FAIL: " & integer'image(errs) & " error(s)" severity error;
        end if;

        halt <= true;
        wait;
    end process;

end architecture;

8. Why a Verification Engineer Cares

Sweep the ratio in both directions, from both sides. Experiment 1 fixes the system clock and sweeps SCLK; experiment 2 fixes SCLK and sweeps the system clock. They test the same inequality and they fail differently, and a suite that only does the first never reaches the case where the destination is the problem.

Require correctness only where the precondition holds, and require a REPORT where it does not. The bench's expectations are computed from the formula rather than hard-coded, so the fastest two ratios are allowed to lose words and are required to report the loss. An expectation table that demanded silence everywhere would be demanding the design hide a real overrun.

Check the transaction after the truncated one. This is Module 14's recurring discipline and it is what found §5's bug: the last word of a transaction is the one a mis-scoped reset destroys, and only a test that reads all of a transaction's words notices.

Model power-on, not just reset. A bench that begins rst_n = 0; #100; rst_n = 1; has never produced a reset edge. That is the single most common reason an Architecture A design reads X in its first simulation.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// Properties for Architecture A. The interesting ones are all about the hand-off,
// because that is where the two domains meet.

property p_word_published_on_the_completing_edge;
    // There is no later edge to publish on, so the toggle must flip on the capture
    // that completes the word -- not one edge later.
    @(posedge cap_clk) disable iff (!rst_n)
        (bit_idx == LAST_BIT) |=> (word_tog != $past(word_tog));
endproperty

property p_handoff_survives_the_deselect;
    // Section 5, as a property: chip select rising must not disturb the hand-off.
    // The version that failed this lost the last word of every transaction.
    @(posedge clk) disable iff (!rst_n)
        $rose(cs_n_pin) |=> $stable(word_src);
endproperty

property p_shift_path_cleared_by_the_select;
    // And the shift path must be cleared by it, which is the other half -- a
    // partial transaction must not contribute bits to the next one.
    @(posedge clk) disable iff (!rst_n)
        cs_n_pin |-> (bit_idx == 0);
endproperty

property p_valid_follows_an_arrival;
    // A word is published only because a toggle arrived, never because the data
    // register changed -- the data is not synchronised and must never be watched.
    @(posedge clk) disable iff (!rst_n)
        rx_valid_stb |-> $past(word_arrived);
endproperty

property p_overrun_sticky;
    // The word it describes has already been lost by the time software could look.
    @(posedge clk) disable iff (!rst_n)
        rx_overrun && !clr_flags |=> rx_overrun;
endproperty
Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// Coverage. The axis is the WORD INTERVAL measured in destination cycles, because
// that is the single quantity the per-word requirement is about -- and it is a
// quantity neither clock's frequency alone determines.

covergroup cg_arch_a @(posedge clk iff rx_valid_stb);
    option.per_instance = 1;

    // LEN x SCLK period, divided by the system period. The requirement is that this
    // exceeds SYNC_N, so the bins are placed around SYNC_N rather than spread.
    interval: coverpoint word_interval_in_dst_cycles {
        bins below    = {[0:1]};          // below SYNC_N at SYNC_N = 2
        bins at       = {2};              // exactly at it
        bins margin_1 = {3};
        bins roomy    = {[4:16]};
        bins generous = {[17:$]};
    }

    // SCLK relative to the system clock, in tenths. The rows above 1.0 are the ones
    // Architecture B cannot reach at all, and they are the reason to choose A.
    sclk_rel: coverpoint sclk_period_tenths {
        bins faster_2x = {[1:5]};
        bins faster    = {[6:9]};
        bins equal     = {10};
        bins slower    = {[11:40]};
        bins far_slower = {[41:$]};
    }

    // The capture edge, which is a synthesis-time parameter -- so this coverpoint is
    // per-build rather than per-run, and a suite that only ever elaborates one value
    // has verified one of the four modes.
    edge_param: coverpoint cap_on_rising { bins rising = {1}; bins falling = {0}; }

    x_interval_rel: cross interval, sclk_rel;

endgroup

9. Why an FPGA or ASIC Engineer Cares

SCLK is a real clock now, and it needs a real clock's treatment: a create_clock, a clock buffer, and a clock tree. That is the opposite of Chapter 14.1's advice and it is the direct consequence of the architecture. On an FPGA, SCLK must reach a global clock buffer, which means it must land on a clock-capable pin — a pinout constraint that arrives from an RTL decision.

The clock is not free-running, and most tools assume clocks are. Clock-gating checks, clock-domain-crossing reports and any analysis that reasons about "cycles" have to cope with a clock that stops for milliseconds. In practice this means the SCLK domain's paths are constrained against the SCLK period and nothing else, and every path leaving the domain is declared asynchronous.

The inverted-capture case must map to a clock primitive, not a LUT. sclk_pin ^ ~CAP_ON_RISING with a constant parameter resolves to sclk_pin or its complement at elaboration, and the inversion is absorbed by the clock buffer's optional inversion. Writing it as a run-time mux would put a LUT in a clock path, which is the thing §2 is about.

The hand-off register is the only path between the domains and it has no timing constraint that means anything. Its source is clocked by SCLK and its destination by the system clock, so a set_max_delay -datapath_only or the vendor's CDC exception applies — and the correctness argument is §3's inequality, which no tool checks. That is the shape of every crossing: the tool is told not to analyse it, and the reason it is safe lives in a comment.

10. Failure Signature — A Slave That Loses The Last Byte Of Every Read

The symptom:

"Multi-byte reads are short by one. A four-byte read returns three bytes and a zero, a two-byte read returns one byte and a zero. Single-byte reads return zero. Slowing the SPI clock does not help."

What is happening: §5's bug. The hand-off register is reset by chip select, and the destination fetches it SYNC_N system cycles after the toggle arrives — which is after the master has raised chip select. Every transaction's last word is cleared before it is read, and a single-byte transaction has only a last word.

Why it survives testing: because the words before the last one are all correct, so a bench that checks "did a word arrive and was it right" passes for every word except one, and a bench that only sends one word at a time sees a total failure that looks like the design not working at all rather than like a reset scoping error.

Why slowing the SPI clock does not help: the race is between chip select rising and the destination's synchroniser, both of which are unaffected by the SCLK period. A fault that ignores the SCLK rate in Architecture A is almost always in the hand-off rather than in the shift path.

11. Common Misconceptions

"Clocking on SCLK removes the clock-domain problem." It moves it. The shift path becomes single-domain and the hand-off becomes the crossing, with its own requirement — SYNC_N system periods must fit inside LEN_FIXED SCLK periods.

"Architecture A has no ratio requirement." It has no per-bit requirement. The per-word one is real, it was measured, and the two fastest ratios in the table fail it.

"Chip select is a sufficient reset for the SCLK domain." It supplies no edge at power-on, so the domain comes up undefined and stays undefined until the first deselect. The reset must be chip select or the system reset.

"An asynchronous reset condition can be a named wire for readability." It cannot: the wire races the sensitivity list that triggered the process, and the reset silently does not happen. The symptom is a word boundary that is wrong in the transaction after a partial one.

"The hand-off register should be reset with the rest of the transaction state." Its lifetime is set by the reader, and the reader arrives after the writer's clock has stopped and its reset has fired. Resetting it with the select loses the last word of every transaction.

"rx_overrun means the destination read stale data." It means two arrivals landed on consecutive cycles, which is a rate observation. Staleness is invisible to the destination — it reads a well-formed word from the wrong frame — and the two thresholds coincide rather than being the same test.

12. Reason It Through

Q. SYNC_N = 2, LEN_FIXED = 8, a 10 ns system clock. What is the fastest SCLK this design supports, and what happens one step beyond it?

2 × 10 = 20 ns must fit inside 8 × T_sclk, so T_sclk > 2.5 ns. At 4 ns the design delivers four correct words; at 2 ns a word arrives every 16 ns against a 20 ns requirement, and the measurement shows one of four words correct with the overrun reported. Architecture B at the same system clock needs 60 ns, so the relief is twenty-four-fold — and it is the frame width that provides most of it, which is why a one-bit frame would remove almost all of the advantage.

Q. Why does the domain's reset need two separate edge triggers in the sensitivity list rather than one combined signal?

Because both constituent signals can already be asserted at time zero, and an edge-sensitive process needs a transition. posedge cs_n_pin fires when the master deselects and negedge rst_n fires when the system resets, and between them every real reset event produces a trigger. One combined posedge dom_rst never rises if dom_rst starts high — which it does, because chip select starts high. Synthesis merges the two into the flop's single reset input, so the cost is zero and the benefit is that simulation sees what silicon sees.

Q. A reviewer proposes making the capture edge a run-time input so the device supports all four modes. What are the two implementations and what does each cost?

A mux on the clock, which glitches the clock net, defeats clock-tree synthesis and creates a path static timing analysis will not analyse — not acceptable. Or two complete shift paths, one on each edge, with a mux on the data at the hand-off — which works, costs a second shift register and bit counter, and roughly doubles the SCLK-domain logic. The second is the honest answer, and it removes a good part of Architecture A's appeal, which is that it is small.

Q. The hand-off register is read by the destination without any synchroniser on the data. Why is that correct rather than an instance of Chapter 15.5's error?

Because the data is not sampled while it is changing. The toggle's arrival is what says the data is stable, and the data has been stable since SYNC_N system cycles earlier and will remain so for LEN_FIXED SCLK periods. Synchronising it bit by bit would introduce the multi-bit problem rather than avoid it. What the scheme depends on instead is a lifetime requirement on the register, which is §3's inequality — so the correctness has moved from a structure to a number.

Q. Under CPHA = 1 the final capture is the last edge of the transaction. Does §5's hand-off argument still hold?

Yes, and more tightly. The word is published on that final edge and there is no edge after it at all, so the hand-off register is the only thing holding the word and the SCLK domain will not run again until the next transaction. Anything that clears it before the destination's read loses the word with no possibility of recovery — there is no later edge to republish on. The argument is the same and the margin is smaller, which is a reason to state the rule structurally rather than to reason about it per mode.

13. Understanding Check

14. Summary

Clocking the shift path on SCLK removes the per-bit precondition entirely: four words arrive correctly at an SCLK period of 4 ns against a 10 ns system clock, which is SCLK two and a half times faster than the system clock and is unreachable for an oversampler at any setting.

It replaces that precondition with a per-word one — SYNC_N × system period < LEN_FIXED × SCLK period — measured two-sidedly, with the boundary falling exactly where the arithmetic puts it. The relief is 2 × HALF_MIN × LEN / SYNC_N, which is twenty-four at these parameters, and most of it comes from the frame width.

The costs are three, and all three are discovered late by teams that do not read them first.

The mode becomes a synthesis-time parameter, because selecting a capture edge at run time means a mux on a clock. Four modes need four builds, or two shift paths.

The domain needs a reset edge that chip select does not supply, because at power-on chip select is already high. The reset is chip select or the system reset — a reset crossing a domain, which Architecture B never has — and its condition must be built from the pins rather than from a derived wire, because a wire races the sensitivity list and the reset then silently does not happen.

The hand-off register must outlive its transaction, because the reader arrives after the writer's clock has stopped. Its lifetime is set by the reader, and a version that reset it with the select lost the last word of every transaction while passing four of five ratios.

And rx_overrun is a rate observation standing in for a staleness failure that the destination cannot see — the fourth thing in this sequence that is invisible from inside the design.

For verification: sweep the ratio from both sides; compute expectations from the formula so that the design is required to report a loss rather than hide one; check the transaction after the truncated one, which is what found the hand-off bug; and model power-on rather than reset.

For implementation: SCLK needs a create_clock, a clock buffer and a clock-capable pin — a pinout constraint arriving from an RTL decision; the clock stops for milliseconds, which most tools do not expect; the inverted capture must resolve at elaboration rather than become a LUT in a clock path; and the one path between the domains has no meaningful timing constraint, so its correctness lives in §3's inequality and in a comment.

15. What Comes Next

Architecture A is built and measured. Architecture B was built in Module 14 and its precondition was stated but never taken apart.

Chapter 15.3 — Architecture B: Oversampling SCLK does that, and the result is that the precondition has two parts which are different in kind. One is arithmetic, exact, and settled completely by simulation — a level present for less than one sampling interval need not be observed, and below that limit the recovered edge count is not merely low but meaningless. The other is metastability, and the chapter's sharpest row is a ratio at which simulation recovers every single edge, reports a measurement below the stated rule, and loses edges on silicon.

Continue learning