Skip to content
VLSI Mentor

SPI · Module 19

SoC to SPI Peripheral Through a Register Bus

Software becomes part of the design. A configuration shadow, two sticky status bits and a toggle handshake — the three races that ship, measured at four unrelated clock ratios.

Chapters 19.1 and 19.2 had hardware decide everything — a timer, or a pin. This chapter puts a processor in the loop. It writes configuration and data, reads status, and does so on its own clock, asynchronously to the SPI engine.

Every interesting failure here is in the seam between software and hardware, and each one is fixed by a contract rather than by a circuit.

1. The Three Races

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   1. software writes CTRL while a transfer is running
   2. software polls DONE and misses a completion
   3. a multi-bit value crosses between the bus clock and the SPI clock

None of them is exotic. All three ship regularly, and all three are cheap to prevent once named.

Top row, left to right: an APB register file feeds a configuration and transmit shadow, which feeds a request toggle synchronised into the SPI domain, which starts the SPI engine, which drives the pins. Bottom row, right to left: the engine's result is latched at the frame end, an acknowledge toggle carries it back across two flops, it lands in RXDATA, and it sets the sticky DONE bit in the status register.APB register fileCTRL, TXDATA, STATUS,IEconfig + txshadowcaptured at transferstartreq toggletwo flops into the SPIdomainSPI engineshift, divide, drivethe pinspinssclk, cs_n, mosi, misoSTATUS bitsBUSY level, DONE andOVERRUN stickyRXDATAreadable once DONE issetack toggletwo flops back to thebusrx capturelatched at the frameendstartheld stablereq toggleframeresultdonebytesets DONE12
Figure 1 — the two domains and what crosses between them. Configuration and transmit data go one way through a shadow register and a request toggle; the received byte and the completion come back through an acknowledge toggle. Nothing multi-bit is synchronised bit-by-bit: the handshake holds the data still while it is sampled.

2. Race One — The Configuration Shadow

Software may write CTRL at any time. It has no reason not to, and nothing stops it.

If the engine reads the live CTRL register, then a write landing mid-transfer changes the mode or the divider during the frame — and the transfer on the wire is neither the old configuration nor the new one. Changing CPOL mid-frame is the worst of them: Chapter 18.2 showed that leading edge is defined relative to the idle level, so inverting CPOL redefines which edges are captures.

The fix is one register:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   at transfer start:  cfg_shadow <= ctrl;  tx_shadow <= pwdata
   for the whole frame: the engine reads ONLY the shadow

The live register moves; the shadowed configuration does not

18 cycles
Five rows over eighteen SPI reference cycles, mid-frame. Chip select stays low. SCLK toggles every three cycles. A live-CTRL row changes from 0x21 to 0x37 at cycle seven, where an APB write marker appears. The shadow-CTRL row stays at 0x21 throughout, and SCLK's spacing does not change.before the writebefore the writeafter it — the wire is unchangedafter it — the wire is unchangedsoftware writes CTRL = 0x37software writes CTRL = 0x37SCLK spacing unchangedSCLK spacing unchangedcs_nsclkAPB writelive ctrl212121212121213737373737373737373737shadow ctrl212121212121212121212121212121212121t0t1t2t3t4t5t6t7t8t9t10t11t12t13t14t15t16t17
Figure 2 — a mid-frame CTRL write, viewed in the SPI reference domain. Software changes the live register from 0x21 (mode 0, divider 2) to 0x37 (mode 3, divider 3). The shadowed copy does not move, so SCLK keeps its original spacing and its original idle level for the rest of the frame.

3. Race Two — Status Bit Semantics

Three bits, three different disciplines, and the reasoning for each is worth stating because getting one wrong loses data silently.

BitKindWhy
BUSYlevelit means what it says at the instant it is read; there is nothing to remember
DONEsticky, write-one-to-cleara level DONE exists only between one transfer ending and the next beginning, so software polling slower than that sees nothing
OVERRUNsticky, write-one-to-cleara sticky DONE has nowhere to record a second completion; without OVERRUN that completion is silently lost

There is one more ordering decision, and it is invisible until it bites:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   the write decode runs FIRST, then the completion
   → a completion landing in the same cycle as a write-one-to-clear WINS

The other order silently drops a completion whenever software happens to clear in that exact cycle — a race software can neither avoid nor observe.

4. Race Three — The Crossing

The register bus runs on pclk; the engine runs on sclk_ref, an unrelated clock. What crosses is multi-bit: the configuration and the transmit byte one way, the received byte the other.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   WRONG   a synchroniser per bit. Each bit resolves independently, so a value
           that was never valid can appear on the other side.

   RIGHT   a request/acknowledge TOGGLE handshake, with the data held stable
           across it. One synchroniser is then enough -- because the HANDSHAKE
           guarantees stability, not the synchroniser.

A toggle rather than a pulse, for the reason Chapter 18.7 measured: a pulse narrower than the receiving clock period is caught only by luck, and a toggle persists until the next one.

5. The Measurement

Identical output from all three languages:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
  shadow test: rxdata=a5 (slave sent a5)   slave received 3c (master sent 3c)   shadow_ctrl=21 (live ctrl=37)   edges=16
  sticky bits after two transfers with no clear: busy=0 done=1 overrun=1
  after writing 0x02:                          busy=0 done=0 overrun=1
  a second start while busy:                    done=1 overrun=1  slave received aa

  sclk_ref half  ratio to pclk  rxdata  frames  edges  status
              7      unrelated      96       1     16  done=1 overrun=0
              3      unrelated      96       1     16  done=1 overrun=0
             11      unrelated      96       1     16  done=1 overrun=0
              4      unrelated      96       1     16  done=1 overrun=0

The shadow test checks four things, not one. The received byte is correct, the byte the slave received is correct, the shadowed configuration still reads 0x21 while the live register reads 0x37, and the wire carried exactly 16 edges. That last check matters because a divider that changed mid-frame would still produce 16 edges — at the wrong spacing — while a mode change would corrupt data with the count looking fine. Checking only the data would miss half the fault space.

A second start while busy is refused and recorded, and the transfer in flight is untouched — the slave still received 0xaa, not some mixture.

And the results do not depend on the clock ratio. Four unrelated sclk_ref half periods against a fixed 5 ns bus half period, and every received byte, frame count and edge count is identical.

6. Building It — Three HDLs

Azvya Education Pvt. Ltd.VLSI Mentor
spi_regbus_ctrl.sv — the register-mapped controller — a shadow, two sticky bits, and a toggle handshake across two clocks
// spi_regbus_ctrl.sv
//
// Chapter 19.3 -- an SPI master behind a register bus, across two clock domains.
//
// THIS IS THE FIRST CHAPTER IN WHICH SOFTWARE IS PART OF THE DESIGN. Chapters 19.1 and 19.2 had hardware
// decide everything: a timer, or a pin. Here a processor writes configuration and data, reads status, and
// does so on ITS OWN clock, asynchronously to the SPI engine -- and the interesting failures are all in
// the seam between the two.
//
// THREE THINGS THIS MODULE EXISTS TO GET RIGHT.
//
//   1. THE SHADOW REGISTER. Software may write CTRL at any time, including while a transfer is running.
//      If the engine reads the live CTRL register, the mode or the divider changes MID-FRAME and the
//      transfer on the wire is neither the old configuration nor the new one. The fix is to latch the
//      configuration at transfer start and use only the latched copy for the duration.
//
//      The alternative is to put the burden on software -- "poll BUSY before writing CTRL" -- which is a
//      real design choice and a worse one, because it is a rule that has to be obeyed by every driver
//      ever written against the part, including the ones written by people who never read this sentence.
//      A shadow register costs one register's worth of flops and removes the rule.
//
//   2. THE STATUS BIT SEMANTICS. `BUSY` is a LEVEL: it means what it says at the instant it is read.
//      `DONE` is STICKY and cleared by writing a one to it (write-one-to-clear), because a level DONE can
//      be missed entirely -- it would be asserted between one transfer ending and the next beginning, and
//      software that polls slower than that sees nothing.
//
//      Sticky introduces its own failure, and it needs its own bit: if a SECOND transfer completes before
//      software has cleared DONE, that completion has nowhere to be recorded. `OVERRUN` exists precisely
//      for that, and a status register without it silently loses completions.
//
//   3. THE CROSSING. The register bus runs on `pclk` and the SPI engine on `sclk_ref`, an unrelated
//      clock. What crosses is multi-bit -- the configuration and the transmit byte one way, the received
//      byte the other -- so the mechanism is a REQUEST/ACKNOWLEDGE TOGGLE HANDSHAKE with the data held
//      stable across it, not a synchroniser per bit.
//
//      Synchronising a multi-bit bus bit-by-bit is the classic error: each bit resolves independently and
//      a value that was never valid can appear on the other side. A handshake makes the data static while
//      it is sampled, which is what makes one synchroniser enough.
//
// WHAT THE SIMULATION CANNOT ESTABLISH, said here rather than implied: nothing about metastability or
// synchroniser depth. Both are two flops deep because that is correct design practice, and a single flop
// would behave identically in zero-delay RTL. What simulation DOES establish is that the handshake holds
// the data stable, that no transfer is lost or duplicated across the crossing, and that the results do
// not depend on the clock ratio -- which is the property a crossing defect would break.

`timescale 1ns/1ps

module spi_regbus_ctrl #(
    parameter int AW = 8
) (
    // ---- the register bus domain ----
    input  wire            pclk,
    input  wire            prst_n,
    input  wire            psel,
    input  wire            penable,
    input  wire            pwrite,
    input  wire [AW-1:0]   paddr,
    input  wire [31:0]     pwdata,
    output reg  [31:0]     prdata,
    output wire            pready,
    output reg             irq,

    // ---- the SPI domain ----
    input  wire            sclk_ref,
    input  wire            srst_n,
    output reg             sclk,
    output reg             cs_n,
    output reg             mosi,
    input  wire            miso,

    // ---- observation, for the bench only ----
    output wire [7:0]      dbg_shadow_ctrl,
    output wire [7:0]      dbg_txshadow
);

    // ---- the register map ----
    localparam [AW-1:0] A_CTRL   = 'h0,   // [0] enable  [1] cpol  [2] cpha  [7:4] divider
                        A_TXDATA = 'h4,   // write starts a transfer
                        A_RXDATA = 'h8,   // read-only
                        A_STATUS = 'hC,   // [0] BUSY (level)  [1] DONE (W1C)  [2] OVERRUN (W1C)
                        A_IE     = 'h10;  // [1] DONE enable   [2] OVERRUN enable

    assign pready = 1'b1;                 // a single-cycle slave; no wait states

    // An APB access completes in the ENABLE phase. Decoding the write in the setup phase would perform
    // it twice, which for a write-one-to-clear bit means clearing something software never wrote.
    wire acc_wr = psel & penable & pwrite;
    wire acc_rd = psel & penable & ~pwrite;

    // ---- the SPI domain's state, declared here so the bus domain can read what crosses back ----
    //
    // Declared above both processes rather than beside the one that writes them. Verilog resolves names
    // in file order, so a cross-domain signal read before its declaration is an elaboration error -- and
    // hoisting them makes the crossing's two directions visible in one place, which is what a reviewer of
    // a clock-domain crossing actually needs.
    reg        req_s1, req_s2, req_s3;
    wire       req_edge = req_s2 ^ req_s3;
    reg        ack_tog;
    reg [7:0]  rx_capture;

    reg [2:0]  est;
    localparam [2:0] E_IDLE = 3'd0, E_LEAD = 3'd1, E_SHIFT = 3'd2, E_LAG = 3'd3;
    reg [7:0]  ediv;
    reg [3:0]  cfg_div;      // the shadowed divider, kept so every edge reloads from the same value
    reg [3:0]  eedge;
    reg [7:0]  esh_tx, esh_rx;
    reg        ecpol, ecpha;

    // ================= the register-bus domain =================
    reg [7:0]  ctrl;        // live configuration, writable at any time
    reg [7:0]  txdata;
    reg [7:0]  rxdata;
    reg [7:0]  ie;
    reg        st_done, st_over;

    // The request side of the crossing: a TOGGLE, because a toggle cannot be missed at any clock ratio.
    reg        req_tog;
    reg [7:0]  cfg_shadow;  // the configuration captured at transfer start
    reg [7:0]  tx_shadow;

    assign dbg_shadow_ctrl = cfg_shadow;
    assign dbg_txshadow    = tx_shadow;

    // The acknowledge coming back, synchronised into this domain.
    reg        ack_s1, ack_s2, ack_s3;
    wire       ack_edge = ack_s2 ^ ack_s3;

    // BUSY is a level derived from the two toggles disagreeing: a request has gone out and its
    // acknowledge has not come back. It needs no separate register and cannot get out of step with the
    // engine, which a hand-maintained busy flag can.
    wire       busy = req_tog ^ ack_s2;

    always_ff @(posedge pclk or negedge prst_n) begin
        if (!prst_n) begin
            ctrl       <= 8'h00;
            txdata     <= 8'h00;
            rxdata     <= 8'h00;
            ie         <= 8'h00;
            st_done    <= 1'b0;
            st_over    <= 1'b0;
            req_tog    <= 1'b0;
            cfg_shadow <= 8'h00;
            tx_shadow  <= 8'h00;
            ack_s1     <= 1'b0;
            ack_s2     <= 1'b0;
            ack_s3     <= 1'b0;
            prdata     <= 32'h0;
            irq        <= 1'b0;
        end else begin
            ack_s1 <= ack_tog;
            ack_s2 <= ack_s1;
            ack_s3 <= ack_s2;

            // ---- writes ----
            if (acc_wr) begin
                case (paddr)
                    A_CTRL: ctrl <= pwdata[7:0];
                    A_IE:   ie   <= pwdata[7:0];
                    A_TXDATA: begin
                        // A write to TXDATA starts a transfer -- but only if one is not already running.
                        // Accepting a second start while busy would either corrupt the frame in flight or
                        // silently discard the byte; refusing it and recording an OVERRUN says what
                        // happened.
                        if (!busy && ctrl[0]) begin
                            // THE SHADOW. The configuration and the data are captured HERE, once, and the
                            // engine sees nothing else for the whole transfer. A later CTRL write changes
                            // `ctrl` and cannot change `cfg_shadow`.
                            cfg_shadow <= ctrl;
                            tx_shadow  <= pwdata[7:0];
                            txdata     <= pwdata[7:0];
                            req_tog    <= ~req_tog;
                        end else begin
                            st_over <= 1'b1;
                        end
                    end
                    A_STATUS: begin
                        // WRITE-ONE-TO-CLEAR. Writing a zero must leave a bit alone, so that a driver
                        // clearing DONE cannot accidentally clear OVERRUN it has not read yet.
                        if (pwdata[1]) st_done <= 1'b0;
                        if (pwdata[2]) st_over <= 1'b0;
                    end
                    default: ;
                endcase
            end

            // ---- the transfer completing ----
            //
            // Ordered AFTER the write decode so that a completion landing in the same cycle as a
            // write-one-to-clear is NOT lost: the completion wins, and software's next poll sees it. The
            // other order silently drops a completion whenever software happens to clear in that cycle,
            // which is a race software cannot avoid and cannot see.
            if (ack_edge) begin
                rxdata <= rx_capture;
                if (st_done) st_over <= 1'b1;   // a second completion with the first unread
                st_done <= 1'b1;
            end

            // ---- reads ----
            if (acc_rd) begin
                case (paddr)
                    A_CTRL:   prdata <= {24'h0, ctrl};
                    A_TXDATA: prdata <= {24'h0, txdata};
                    A_RXDATA: prdata <= {24'h0, rxdata};
                    A_STATUS: prdata <= {29'h0, st_over, st_done, busy};
                    A_IE:     prdata <= {24'h0, ie};
                    default:  prdata <= 32'h0;
                endcase
            end

            irq <= (st_done & ie[1]) | (st_over & ie[2]);
        end
    end

    // ================= the SPI domain =================
    always_ff @(posedge sclk_ref or negedge srst_n) begin
        if (!srst_n) begin
            req_s1 <= 1'b0; req_s2 <= 1'b0; req_s3 <= 1'b0;
            ack_tog    <= 1'b0;
            rx_capture <= 8'h00;
            est        <= E_IDLE;
            ediv       <= 8'h00;
            cfg_div    <= 4'h0;
            eedge      <= 4'h0;
            esh_tx     <= 8'h00;
            esh_rx     <= 8'h00;
            ecpol      <= 1'b0;
            ecpha      <= 1'b0;
            sclk       <= 1'b0;
            cs_n       <= 1'b1;
            mosi       <= 1'b0;
        end else begin
            req_s1 <= req_tog;
            req_s2 <= req_s1;
            req_s3 <= req_s2;

            case (est)
                E_IDLE: begin
                    if (req_edge) begin
                        // THE DATA IS SAMPLED HERE, from registers the other domain has held stable since
                        // before the toggle flipped. That is what makes one synchroniser enough for a
                        // multi-bit value: the handshake, not the synchroniser, guarantees stability.
                        ecpol   <= dbg_shadow_ctrl[1];
                        ecpha   <= dbg_shadow_ctrl[2];
                        cfg_div <= dbg_shadow_ctrl[7:4];
                        ediv    <= {4'h0, dbg_shadow_ctrl[7:4]};
                        esh_tx  <= dbg_txshadow;
                        esh_rx  <= 8'h00;
                        sclk    <= dbg_shadow_ctrl[1];
                        cs_n    <= 1'b0;
                        mosi    <= dbg_txshadow[7];
                        eedge   <= 4'h0;
                        est     <= E_LEAD;
                    end
                end

                E_LEAD: begin
                    est <= E_SHIFT;
                end

                E_SHIFT: begin
                    // One SCLK edge per reference cycle when the divider is zero; otherwise one per
                    // (divider + 1). The divider is deliberately part of the shadowed configuration, so
                    // changing it mid-transfer is impossible by construction rather than by convention.
                    if (ediv == 8'h00) begin
                        // THE DIVIDER RELOADS AFTER EVERY EDGE. The first version loaded it once at
                        // transfer start and decremented to zero, after which edges came out one per
                        // reference cycle forever -- so the divider delayed the FIRST edge and did
                        // nothing else. The frame still carried 16 edges and the received byte was
                        // still correct at low rates, so nothing failed until the slave ran out of
                        // setup margin. A divider that only works at one setting is worse than none.
                        ediv <= {4'h0, cfg_div};
                        if (sclk == ecpol) begin
                            // the leading edge: capture for CPHA = 0
                            if (!ecpha) esh_rx <= {esh_rx[6:0], miso};
                            sclk <= ~ecpol;
                        end else begin
                            if (ecpha) esh_rx <= {esh_rx[6:0], miso};
                            sclk   <= ecpol;
                            esh_tx <= {esh_tx[6:0], 1'b0};
                            mosi   <= esh_tx[6];
                        end
                        if (eedge == 4'hF) begin
                            est <= E_LAG;
                        end else begin
                            eedge <= eedge + 4'h1;
                        end
                    end else begin
                        ediv <= ediv - 8'h01;
                    end
                end

                E_LAG: begin
                    cs_n       <= 1'b1;
                    rx_capture <= esh_rx;
                    ack_tog    <= ~ack_tog;
                    est        <= E_IDLE;
                end

                default: est <= E_IDLE;
            endcase
        end
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_regbus_ctrl.v — the same design in Verilog-2001
// spi_regbus_ctrl.v
//
// Chapter 19.3 -- an SPI master behind a register bus, across two clock domains.
//
// THIS IS THE FIRST CHAPTER IN WHICH SOFTWARE IS PART OF THE DESIGN. Chapters 19.1 and 19.2 had hardware
// decide everything: a timer, or a pin. Here a processor writes configuration and data, reads status, and
// does so on ITS OWN clock, asynchronously to the SPI engine -- and the interesting failures are all in
// the seam between the two.
//
// THREE THINGS THIS MODULE EXISTS TO GET RIGHT.
//
//   1. THE SHADOW REGISTER. Software may write CTRL at any time, including while a transfer is running.
//      If the engine reads the live CTRL register, the mode or the divider changes MID-FRAME and the
//      transfer on the wire is neither the old configuration nor the new one. The fix is to latch the
//      configuration at transfer start and use only the latched copy for the duration.
//
//      The alternative is to put the burden on software -- "poll BUSY before writing CTRL" -- which is a
//      real design choice and a worse one, because it is a rule that has to be obeyed by every driver
//      ever written against the part, including the ones written by people who never read this sentence.
//      A shadow register costs one register's worth of flops and removes the rule.
//
//   2. THE STATUS BIT SEMANTICS. `BUSY` is a LEVEL: it means what it says at the instant it is read.
//      `DONE` is STICKY and cleared by writing a one to it (write-one-to-clear), because a level DONE can
//      be missed entirely -- it would be asserted between one transfer ending and the next beginning, and
//      software that polls slower than that sees nothing.
//
//      Sticky introduces its own failure, and it needs its own bit: if a SECOND transfer completes before
//      software has cleared DONE, that completion has nowhere to be recorded. `OVERRUN` exists precisely
//      for that, and a status register without it silently loses completions.
//
//   3. THE CROSSING. The register bus runs on `pclk` and the SPI engine on `sclk_ref`, an unrelated
//      clock. What crosses is multi-bit -- the configuration and the transmit byte one way, the received
//      byte the other -- so the mechanism is a REQUEST/ACKNOWLEDGE TOGGLE HANDSHAKE with the data held
//      stable across it, not a synchroniser per bit.
//
//      Synchronising a multi-bit bus bit-by-bit is the classic error: each bit resolves independently and
//      a value that was never valid can appear on the other side. A handshake makes the data static while
//      it is sampled, which is what makes one synchroniser enough.
//
// WHAT THE SIMULATION CANNOT ESTABLISH, said here rather than implied: nothing about metastability or
// synchroniser depth. Both are two flops deep because that is correct design practice, and a single flop
// would behave identically in zero-delay RTL. What simulation DOES establish is that the handshake holds
// the data stable, that no transfer is lost or duplicated across the crossing, and that the results do
// not depend on the clock ratio -- which is the property a crossing defect would break.

`timescale 1ns/1ps

module spi_regbus_ctrl #(
    parameter AW = 8
) (
    // ---- the register bus domain ----
    input  wire            pclk,
    input  wire            prst_n,
    input  wire            psel,
    input  wire            penable,
    input  wire            pwrite,
    input  wire [AW-1:0]   paddr,
    input  wire [31:0]     pwdata,
    output reg  [31:0]     prdata,
    output wire            pready,
    output reg             irq,

    // ---- the SPI domain ----
    input  wire            sclk_ref,
    input  wire            srst_n,
    output reg             sclk,
    output reg             cs_n,
    output reg             mosi,
    input  wire            miso,

    // ---- observation, for the bench only ----
    output wire [7:0]      dbg_shadow_ctrl,
    output wire [7:0]      dbg_txshadow
);

    // ---- the register map ----
    localparam [AW-1:0] A_CTRL   = 'h0,   // [0] enable  [1] cpol  [2] cpha  [7:4] divider
                        A_TXDATA = 'h4,   // write starts a transfer
                        A_RXDATA = 'h8,   // read-only
                        A_STATUS = 'hC,   // [0] BUSY (level)  [1] DONE (W1C)  [2] OVERRUN (W1C)
                        A_IE     = 'h10;  // [1] DONE enable   [2] OVERRUN enable

    assign pready = 1'b1;                 // a single-cycle slave; no wait states

    // An APB access completes in the ENABLE phase. Decoding the write in the setup phase would perform
    // it twice, which for a write-one-to-clear bit means clearing something software never wrote.
    wire acc_wr = psel & penable & pwrite;
    wire acc_rd = psel & penable & ~pwrite;

    // ---- the SPI domain's state, declared here so the bus domain can read what crosses back ----
    //
    // Declared above both processes rather than beside the one that writes them. Verilog resolves names
    // in file order, so a cross-domain signal read before its declaration is an elaboration error -- and
    // hoisting them makes the crossing's two directions visible in one place, which is what a reviewer of
    // a clock-domain crossing actually needs.
    reg        req_s1, req_s2, req_s3;
    wire       req_edge = req_s2 ^ req_s3;
    reg        ack_tog;
    reg [7:0]  rx_capture;

    reg [2:0]  est;
    localparam [2:0] E_IDLE = 3'd0, E_LEAD = 3'd1, E_SHIFT = 3'd2, E_LAG = 3'd3;
    reg [7:0]  ediv;
    reg [3:0]  cfg_div;      // the shadowed divider, kept so every edge reloads from the same value
    reg [3:0]  eedge;
    reg [7:0]  esh_tx, esh_rx;
    reg        ecpol, ecpha;

    // ================= the register-bus domain =================
    reg [7:0]  ctrl;        // live configuration, writable at any time
    reg [7:0]  txdata;
    reg [7:0]  rxdata;
    reg [7:0]  ie;
    reg        st_done, st_over;

    // The request side of the crossing: a TOGGLE, because a toggle cannot be missed at any clock ratio.
    reg        req_tog;
    reg [7:0]  cfg_shadow;  // the configuration captured at transfer start
    reg [7:0]  tx_shadow;

    assign dbg_shadow_ctrl = cfg_shadow;
    assign dbg_txshadow    = tx_shadow;

    // The acknowledge coming back, synchronised into this domain.
    reg        ack_s1, ack_s2, ack_s3;
    wire       ack_edge = ack_s2 ^ ack_s3;

    // BUSY is a level derived from the two toggles disagreeing: a request has gone out and its
    // acknowledge has not come back. It needs no separate register and cannot get out of step with the
    // engine, which a hand-maintained busy flag can.
    wire       busy = req_tog ^ ack_s2;

    always @(posedge pclk or negedge prst_n) begin
        if (!prst_n) begin
            ctrl       <= 8'h00;
            txdata     <= 8'h00;
            rxdata     <= 8'h00;
            ie         <= 8'h00;
            st_done    <= 1'b0;
            st_over    <= 1'b0;
            req_tog    <= 1'b0;
            cfg_shadow <= 8'h00;
            tx_shadow  <= 8'h00;
            ack_s1     <= 1'b0;
            ack_s2     <= 1'b0;
            ack_s3     <= 1'b0;
            prdata     <= 32'h0;
            irq        <= 1'b0;
        end else begin
            ack_s1 <= ack_tog;
            ack_s2 <= ack_s1;
            ack_s3 <= ack_s2;

            // ---- writes ----
            if (acc_wr) begin
                case (paddr)
                    A_CTRL: ctrl <= pwdata[7:0];
                    A_IE:   ie   <= pwdata[7:0];
                    A_TXDATA: begin
                        // A write to TXDATA starts a transfer -- but only if one is not already running.
                        // Accepting a second start while busy would either corrupt the frame in flight or
                        // silently discard the byte; refusing it and recording an OVERRUN says what
                        // happened.
                        if (!busy && ctrl[0]) begin
                            // THE SHADOW. The configuration and the data are captured HERE, once, and the
                            // engine sees nothing else for the whole transfer. A later CTRL write changes
                            // `ctrl` and cannot change `cfg_shadow`.
                            cfg_shadow <= ctrl;
                            tx_shadow  <= pwdata[7:0];
                            txdata     <= pwdata[7:0];
                            req_tog    <= ~req_tog;
                        end else begin
                            st_over <= 1'b1;
                        end
                    end
                    A_STATUS: begin
                        // WRITE-ONE-TO-CLEAR. Writing a zero must leave a bit alone, so that a driver
                        // clearing DONE cannot accidentally clear OVERRUN it has not read yet.
                        if (pwdata[1]) st_done <= 1'b0;
                        if (pwdata[2]) st_over <= 1'b0;
                    end
                    default: ;
                endcase
            end

            // ---- the transfer completing ----
            //
            // Ordered AFTER the write decode so that a completion landing in the same cycle as a
            // write-one-to-clear is NOT lost: the completion wins, and software's next poll sees it. The
            // other order silently drops a completion whenever software happens to clear in that cycle,
            // which is a race software cannot avoid and cannot see.
            if (ack_edge) begin
                rxdata <= rx_capture;
                if (st_done) st_over <= 1'b1;   // a second completion with the first unread
                st_done <= 1'b1;
            end

            // ---- reads ----
            if (acc_rd) begin
                case (paddr)
                    A_CTRL:   prdata <= {24'h0, ctrl};
                    A_TXDATA: prdata <= {24'h0, txdata};
                    A_RXDATA: prdata <= {24'h0, rxdata};
                    A_STATUS: prdata <= {29'h0, st_over, st_done, busy};
                    A_IE:     prdata <= {24'h0, ie};
                    default:  prdata <= 32'h0;
                endcase
            end

            irq <= (st_done & ie[1]) | (st_over & ie[2]);
        end
    end

    // ================= the SPI domain =================
    always @(posedge sclk_ref or negedge srst_n) begin
        if (!srst_n) begin
            req_s1 <= 1'b0; req_s2 <= 1'b0; req_s3 <= 1'b0;
            ack_tog    <= 1'b0;
            rx_capture <= 8'h00;
            est        <= E_IDLE;
            ediv       <= 8'h00;
            cfg_div    <= 4'h0;
            eedge      <= 4'h0;
            esh_tx     <= 8'h00;
            esh_rx     <= 8'h00;
            ecpol      <= 1'b0;
            ecpha      <= 1'b0;
            sclk       <= 1'b0;
            cs_n       <= 1'b1;
            mosi       <= 1'b0;
        end else begin
            req_s1 <= req_tog;
            req_s2 <= req_s1;
            req_s3 <= req_s2;

            case (est)
                E_IDLE: begin
                    if (req_edge) begin
                        // THE DATA IS SAMPLED HERE, from registers the other domain has held stable since
                        // before the toggle flipped. That is what makes one synchroniser enough for a
                        // multi-bit value: the handshake, not the synchroniser, guarantees stability.
                        ecpol   <= dbg_shadow_ctrl[1];
                        ecpha   <= dbg_shadow_ctrl[2];
                        cfg_div <= dbg_shadow_ctrl[7:4];
                        ediv    <= {4'h0, dbg_shadow_ctrl[7:4]};
                        esh_tx  <= dbg_txshadow;
                        esh_rx  <= 8'h00;
                        sclk    <= dbg_shadow_ctrl[1];
                        cs_n    <= 1'b0;
                        mosi    <= dbg_txshadow[7];
                        eedge   <= 4'h0;
                        est     <= E_LEAD;
                    end
                end

                E_LEAD: begin
                    est <= E_SHIFT;
                end

                E_SHIFT: begin
                    // One SCLK edge per reference cycle when the divider is zero; otherwise one per
                    // (divider + 1). The divider is deliberately part of the shadowed configuration, so
                    // changing it mid-transfer is impossible by construction rather than by convention.
                    if (ediv == 8'h00) begin
                        // THE DIVIDER RELOADS AFTER EVERY EDGE. The first version loaded it once at
                        // transfer start and decremented to zero, after which edges came out one per
                        // reference cycle forever -- so the divider delayed the FIRST edge and did
                        // nothing else. The frame still carried 16 edges and the received byte was
                        // still correct at low rates, so nothing failed until the slave ran out of
                        // setup margin. A divider that only works at one setting is worse than none.
                        ediv <= {4'h0, cfg_div};
                        if (sclk == ecpol) begin
                            // the leading edge: capture for CPHA = 0
                            if (!ecpha) esh_rx <= {esh_rx[6:0], miso};
                            sclk <= ~ecpol;
                        end else begin
                            if (ecpha) esh_rx <= {esh_rx[6:0], miso};
                            sclk   <= ecpol;
                            esh_tx <= {esh_tx[6:0], 1'b0};
                            mosi   <= esh_tx[6];
                        end
                        if (eedge == 4'hF) begin
                            est <= E_LAG;
                        end else begin
                            eedge <= eedge + 4'h1;
                        end
                    end else begin
                        ediv <= ediv - 8'h01;
                    end
                end

                E_LAG: begin
                    cs_n       <= 1'b1;
                    rx_capture <= esh_rx;
                    ack_tog    <= ~ack_tog;
                    est        <= E_IDLE;
                end

                default: est <= E_IDLE;
            endcase
        end
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_regbus_ctrl.vhd — the same design in VHDL
-- spi_regbus_ctrl.vhd
--
-- Chapter 19.3 -- an SPI master behind a register bus, across two clock domains.
--
-- THIS IS THE FIRST CHAPTER IN WHICH SOFTWARE IS PART OF THE DESIGN. Chapters 19.1 and 19.2 had hardware
-- decide everything: a timer, or a pin. Here a processor writes configuration and data, reads status, and
-- does so on ITS OWN clock, asynchronously to the SPI engine -- and the interesting failures are all in
-- the seam between the two.
--
-- THREE THINGS THIS MODULE EXISTS TO GET RIGHT.
--
--   1. THE SHADOW REGISTER. Software may write CTRL at any time, including while a transfer is running.
--      If the engine reads the live CTRL register, the mode or the divider changes MID-FRAME and the
--      transfer on the wire is neither the old configuration nor the new one. The fix is to latch the
--      configuration at transfer start and use only the latched copy for the duration.
--
--      The alternative is to put the burden on software -- "poll BUSY before writing CTRL" -- which is a
--      real design choice and a worse one, because it is a rule that has to be obeyed by every driver
--      ever written against the part, including the ones written by people who never read this sentence.
--      A shadow register costs one register's worth of flops and removes the rule.
--
--   2. THE STATUS BIT SEMANTICS. `BUSY` is a LEVEL: it means what it says at the instant it is read.
--      `DONE` is STICKY and cleared by writing a one to it (write-one-to-clear), because a level DONE can
--      be missed entirely -- it would be asserted between one transfer ending and the next beginning, and
--      software that polls slower than that sees nothing.
--
--      Sticky introduces its own failure, and it needs its own bit: if a SECOND transfer completes before
--      software has cleared DONE, that completion has nowhere to be recorded. `OVERRUN` exists precisely
--      for that, and a status register without it silently loses completions.
--
--   3. THE CROSSING. The register bus runs on `pclk` and the SPI engine on `sclk_ref`, an unrelated
--      clock. What crosses is multi-bit -- the configuration and the transmit byte one way, the received
--      byte the other -- so the mechanism is a REQUEST/ACKNOWLEDGE TOGGLE HANDSHAKE with the data held
--      stable across it, not a synchroniser per bit.
--
--      Synchronising a multi-bit bus bit-by-bit is the classic error: each bit resolves independently and
--      a value that was never valid can appear on the other side. A handshake makes the data static while
--      it is sampled, which is what makes one synchroniser enough.
--
-- WHAT THE SIMULATION CANNOT ESTABLISH, said here rather than implied: nothing about metastability or
-- synchroniser depth. Both are two flops deep because that is correct design practice, and a single flop
-- would behave identically in zero-delay RTL. What simulation DOES establish is that the handshake holds
-- the data stable, that no transfer is lost or duplicated across the crossing, and that the results do
-- not depend on the clock ratio -- which is the property a crossing defect would break.

--
-- WHAT THE VHDL VERSION ADDS, and it is structural rather than cosmetic.
--
-- EVERY REGISTER HERE IS A SIGNAL, not a process variable. Chapter 19.2 found four separate defects that
-- all came from the same fact -- a process variable is updated immediately and a non-blocking reg is not --
-- and each one produced a design that behaved plausibly. Using signals throughout makes the semantics
-- match the SystemVerilog and Verilog versions by construction rather than by care, which is the right
-- trade for a module whose whole subject is two clock domains disagreeing about time.
--
-- The register addresses are named constants of a constrained subtype, so an access to an address outside
-- the map is a decode miss rather than an aliased hit.
--
-- IDENTIFIER REVIEW (VHDL IS CASE-INSENSITIVE). The bus ports are `psel`, `penable`, `pwrite`, `paddr`,
-- `pwdata`, `prdata`; the internal registers are `ctrl_r`, `tx_r`, `rx_r`, `ie_r` -- suffixed rather than
-- case-varied, so nothing is distinguished from a port by case alone. `A_CTRL` and friends are constants
-- and no signal shares their spelling.
--
-- RESERVED-WORD REVIEW. Nothing is named `label`, `range`, `next`, `access`, `body`, `bus`, `register`,
-- `open`, `guarded`, `block`, `new`, `abs`, `rem`, `mod`, `exit`, `file`, `group` or `literal`. The last
-- one is worth checking in a register-file design, where `register` and `literal` are tempting names.
--
-- RANGE-DIRECTION REVIEW. Every vector is `downto`; the shift registers are constrained subtypes; and the
-- bus read data is assembled with explicit zero padding rather than by concatenating a literal, so no
-- slice inherits an ascending range.

library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

package spi_regbus_pkg is
    subtype byte_t is std_logic_vector(7 downto 0);
    type eng_state_t is (E_IDLE, E_LEAD, E_SHIFT, E_LAG);
end package spi_regbus_pkg;

library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
use work.spi_regbus_pkg.all;

entity spi_regbus_ctrl is
    generic (
        AW : positive := 8
    );
    port (
        -- the register bus domain
        pclk    : in  std_logic;
        prst_n  : in  std_logic;
        psel    : in  std_logic;
        penable : in  std_logic;
        pwrite  : in  std_logic;
        paddr   : in  std_logic_vector(AW - 1 downto 0);
        pwdata  : in  std_logic_vector(31 downto 0);
        prdata  : out std_logic_vector(31 downto 0);
        pready  : out std_logic;
        irq     : out std_logic;

        -- the SPI domain
        sclk_ref : in  std_logic;
        srst_n   : in  std_logic;
        sclk     : out std_logic;
        cs_n     : out std_logic;
        mosi     : out std_logic;
        miso     : in  std_logic;

        -- observation, for the bench only
        dbg_shadow_ctrl : out byte_t;
        dbg_txshadow    : out byte_t
    );
end entity spi_regbus_ctrl;

architecture rtl of spi_regbus_ctrl is

    constant A_CTRL   : std_logic_vector(AW - 1 downto 0) := std_logic_vector(to_unsigned(16#00#, AW));
    constant A_TXDATA : std_logic_vector(AW - 1 downto 0) := std_logic_vector(to_unsigned(16#04#, AW));
    constant A_RXDATA : std_logic_vector(AW - 1 downto 0) := std_logic_vector(to_unsigned(16#08#, AW));
    constant A_STATUS : std_logic_vector(AW - 1 downto 0) := std_logic_vector(to_unsigned(16#0C#, AW));
    constant A_IE     : std_logic_vector(AW - 1 downto 0) := std_logic_vector(to_unsigned(16#10#, AW));

    -- ---- the register-bus domain ----
    signal ctrl_r, tx_r, rx_r, ie_r : byte_t := (others => '0');
    signal st_done, st_over : std_logic := '0';
    signal req_tog : std_logic := '0';
    signal cfg_shadow, tx_shadow : byte_t := (others => '0');
    signal ack_sr : std_logic_vector(2 downto 0) := (others => '0');
    signal prdata_r, irq_r : std_logic := '0';
    signal prd_r : std_logic_vector(31 downto 0) := (others => '0');

    -- ---- the SPI domain, declared here so the bus domain can read what crosses back ----
    signal req_sr     : std_logic_vector(2 downto 0) := (others => '0');
    signal ack_tog    : std_logic := '0';
    signal rx_capture : byte_t := (others => '0');
    signal est        : eng_state_t := E_IDLE;
    signal ediv       : unsigned(7 downto 0) := (others => '0');
    signal cfg_div    : std_logic_vector(3 downto 0) := (others => '0');
    signal eedge      : unsigned(3 downto 0) := (others => '0');
    signal esh_tx, esh_rx : byte_t := (others => '0');
    signal ecpol, ecpha : std_logic := '0';
    signal sclk_r, cs_r, mosi_r : std_logic := '0';

    -- An APB access completes in the ENABLE phase. Decoding a write in the setup phase performs it twice,
    -- which for a write-one-to-clear bit means clearing something software never wrote.
    signal acc_wr, acc_rd : std_logic;

    -- BUSY is a level derived from the two toggles disagreeing: a request has gone out and its acknowledge
    -- has not come back. It needs no separate register and cannot get out of step with the engine, which a
    -- hand-maintained busy flag can.
    signal busy_w : std_logic;
    signal ack_edge_w, req_edge_w : std_logic;

begin

    pready <= '1';                      -- a single-cycle slave; no wait states
    prdata <= prd_r;
    irq    <= irq_r;
    sclk   <= sclk_r;
    cs_n   <= cs_r;
    mosi   <= mosi_r;
    dbg_shadow_ctrl <= cfg_shadow;
    dbg_txshadow    <= tx_shadow;

    acc_wr <= psel and penable and pwrite;
    acc_rd <= psel and penable and (not pwrite);

    busy_w     <= req_tog xor ack_sr(1);
    ack_edge_w <= ack_sr(1) xor ack_sr(2);
    req_edge_w <= req_sr(1) xor req_sr(2);

    -- ================= the register-bus domain =================
    bus_dom : process (pclk, prst_n) is
    begin
        if prst_n = '0' then
            ctrl_r <= (others => '0'); tx_r <= (others => '0');
            rx_r   <= (others => '0'); ie_r <= (others => '0');
            st_done <= '0'; st_over <= '0';
            req_tog <= '0';
            cfg_shadow <= (others => '0'); tx_shadow <= (others => '0');
            ack_sr <= (others => '0');
            prd_r  <= (others => '0');
            irq_r  <= '0';

        elsif rising_edge(pclk) then
            ack_sr <= ack_sr(1 downto 0) & ack_tog;

            -- ---- writes ----
            if acc_wr = '1' then
                if paddr = A_CTRL then
                    ctrl_r <= pwdata(7 downto 0);
                elsif paddr = A_IE then
                    ie_r <= pwdata(7 downto 0);
                elsif paddr = A_TXDATA then
                    -- A write to TXDATA starts a transfer, but only if one is not already running.
                    -- Accepting a second start while busy would either corrupt the frame in flight or
                    -- silently discard the byte; refusing it and recording an OVERRUN says what happened.
                    if busy_w = '0' and ctrl_r(0) = '1' then
                        -- THE SHADOW. Configuration and data are captured HERE, once, and the engine sees
                        -- nothing else for the whole transfer. A later CTRL write changes `ctrl_r` and
                        -- cannot change `cfg_shadow`.
                        cfg_shadow <= ctrl_r;
                        tx_shadow  <= pwdata(7 downto 0);
                        tx_r       <= pwdata(7 downto 0);
                        req_tog    <= not req_tog;
                    else
                        st_over <= '1';
                    end if;
                elsif paddr = A_STATUS then
                    -- WRITE-ONE-TO-CLEAR. Writing a zero must leave a bit alone, so a driver clearing
                    -- DONE cannot accidentally clear an OVERRUN it has not read yet.
                    if pwdata(1) = '1' then st_done <= '0'; end if;
                    if pwdata(2) = '1' then st_over <= '0'; end if;
                end if;
            end if;

            -- ---- the transfer completing ----
            --
            -- Ordered AFTER the write decode so that a completion landing in the same cycle as a
            -- write-one-to-clear is NOT lost: the completion wins and software's next poll sees it. The
            -- other order silently drops a completion whenever software happens to clear in that cycle,
            -- which is a race software can neither avoid nor see.
            if ack_edge_w = '1' then
                rx_r <= rx_capture;
                if st_done = '1' then
                    st_over <= '1';            -- a second completion with the first unread
                end if;
                st_done <= '1';
            end if;

            -- ---- reads ----
            if acc_rd = '1' then
                if paddr = A_CTRL then
                    prd_r <= x"000000" & ctrl_r;
                elsif paddr = A_TXDATA then
                    prd_r <= x"000000" & tx_r;
                elsif paddr = A_RXDATA then
                    prd_r <= x"000000" & rx_r;
                elsif paddr = A_STATUS then
                    prd_r <= x"000000" & "00000" & st_over & st_done & busy_w;
                elsif paddr = A_IE then
                    prd_r <= x"000000" & ie_r;
                else
                    prd_r <= (others => '0');
                end if;
            end if;

            irq_r <= (st_done and ie_r(1)) or (st_over and ie_r(2));
        end if;
    end process bus_dom;

    -- ================= the SPI domain =================
    spi_dom : process (sclk_ref, srst_n) is
    begin
        if srst_n = '0' then
            req_sr <= (others => '0');
            ack_tog <= '0';
            rx_capture <= (others => '0');
            est   <= E_IDLE;
            ediv  <= (others => '0');
            cfg_div <= (others => '0');
            eedge <= (others => '0');
            esh_tx <= (others => '0'); esh_rx <= (others => '0');
            ecpol <= '0'; ecpha <= '0';
            sclk_r <= '0'; cs_r <= '1'; mosi_r <= '0';

        elsif rising_edge(sclk_ref) then
            req_sr <= req_sr(1 downto 0) & req_tog;

            case est is
                when E_IDLE =>
                    if req_edge_w = '1' then
                        -- THE DATA IS SAMPLED HERE, from registers the other domain has held stable since
                        -- before the toggle flipped. That is what makes one synchroniser enough for a
                        -- multi-bit value: the handshake, not the synchroniser, guarantees stability.
                        ecpol   <= cfg_shadow(1);
                        ecpha   <= cfg_shadow(2);
                        cfg_div <= cfg_shadow(7 downto 4);
                        ediv    <= unsigned(x"0" & cfg_shadow(7 downto 4));
                        esh_tx  <= tx_shadow;
                        esh_rx  <= (others => '0');
                        sclk_r  <= cfg_shadow(1);
                        cs_r    <= '0';
                        mosi_r  <= tx_shadow(7);
                        eedge   <= (others => '0');
                        est     <= E_LEAD;
                    end if;

                when E_LEAD =>
                    est <= E_SHIFT;

                when E_SHIFT =>
                    if ediv = 0 then
                        -- The divider RELOADS after every edge. Loading it once and decrementing to zero
                        -- makes it delay the first edge and nothing else, which still produces 16 edges
                        -- and a correct byte until the slave runs out of setup margin.
                        ediv <= unsigned(x"0" & cfg_div);
                        if sclk_r = ecpol then
                            if ecpha = '0' then
                                esh_rx <= esh_rx(6 downto 0) & miso;
                            end if;
                            sclk_r <= not ecpol;
                        else
                            if ecpha = '1' then
                                esh_rx <= esh_rx(6 downto 0) & miso;
                            end if;
                            sclk_r <= ecpol;
                            -- The read comes before the shift in the OTHER languages because a reg's
                            -- non-blocking shift is not visible in the same cycle. Here both are signal
                            -- assignments, so `esh_tx` still reads its pre-edge value and the order does
                            -- not matter -- which is exactly why signals are used throughout this file.
                            mosi_r <= esh_tx(6);
                            esh_tx <= esh_tx(6 downto 0) & '0';
                        end if;
                        if eedge = 15 then
                            est <= E_LAG;
                        else
                            eedge <= eedge + 1;
                        end if;
                    else
                        ediv <= ediv - 1;
                    end if;

                when E_LAG =>
                    cs_r       <= '1';
                    rx_capture <= esh_rx;
                    ack_tog    <= not ack_tog;
                    est        <= E_IDLE;
            end case;
        end if;
    end process spi_dom;

end architecture rtl;

The Bench

The bench drives the register bus the way a driver would — through APB accesses, with no visibility into the engine — and models the slave on the pins. The pin monitor counts SCLK edges independently of the design, so the frame had 16 edges is an observation rather than a claim.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_regbus_ctrl_tb.sv — the three races driven deliberately, and rate invariance measured at four unrelated clock ratios
// spi_regbus_ctrl_tb.sv
//
// SOFTWARE ON ONE CLOCK, AN SPI ENGINE ON ANOTHER, AND THE THREE RACES THAT SHIP.
//
// The bench drives the register bus the way a driver would -- writes and reads through APB accesses, with
// no visibility into the engine -- and models the slave on the pins. The two clocks are unrelated and
// their ratio is swept, because a crossing defect is a phase effect and a single ratio cannot find one.
//
// THE FOUR RESULTS.
//
//   1. THE SHADOW REGISTER MAKES A MID-TRANSFER CONFIGURATION WRITE HARMLESS. Software writes CTRL with a
//      different mode AND a different divider while a transfer is in flight. The bench requires the
//      received byte to be correct, the shadowed configuration to be UNCHANGED, and the number of SCLK
//      edges on the wire to be exactly 16 -- because a divider that changed mid-frame would still produce
//      16 edges at the wrong spacing, and a mode that changed would corrupt the data while the count
//      looked fine. Both are checked.
//
//   2. WRITE-ONE-TO-CLEAR MEANS WRITING A ZERO LEAVES A BIT ALONE. The bench sets both sticky bits, writes
//      a word with only one bit set, and requires the other to survive. A status register that clears on
//      any write destroys a bit the driver had not read yet, and the symptom is a lost completion.
//
//   3. A SECOND COMPLETION WITH THE FIRST UNREAD MUST BE RECORDED, NOT DISCARDED. Two transfers are run
//      without clearing DONE between them, and OVERRUN must be set. A sticky DONE with no OVERRUN silently
//      loses completions, which is the failure a level DONE was replaced to avoid -- traded for a
//      different one rather than fixed.
//
//   4. THE RESULTS DO NOT DEPEND ON THE CLOCK RATIO. The same traffic is run at four unrelated
//      pclk-to-sclk_ref ratios and every received byte, every status word and every transfer count must be
//      IDENTICAL. That is rate invariance, and it is the property a multi-bit crossing defect breaks --
//      the one thing a simulation can genuinely establish about a crossing.

`timescale 1ns/1ps

module spi_regbus_ctrl_tb;

    localparam int AW = 8;

    localparam [AW-1:0] A_CTRL = 'h0, A_TXDATA = 'h4, A_RXDATA = 'h8,
                        A_STATUS = 'hC, A_IE = 'h10;

    // ---- the register-bus clock: fixed ----
    reg pclk = 1'b0;
    always #5 pclk = ~pclk;
    reg prst_n = 1'b1;

    // ---- the SPI reference clock: its period is the swept variable ----
    reg  sclk_ref = 1'b0;
    reg  srst_n   = 1'b1;
    integer sref_half = 7;                 // half period in ns; deliberately not a divisor of 10
    always begin
        #(sref_half) sclk_ref = ~sclk_ref;
    end

    reg         psel = 1'b0, penable = 1'b0, pwrite = 1'b0;
    // `'0` is a SystemVerilog literal and does not survive the Verilog-2001 conversion; an
    // explicit width is the only spelling that means the same thing in both.
    reg [AW-1:0] paddr = {AW{1'b0}};
    reg [31:0]  pwdata = 32'h0;
    wire [31:0] prdata;
    wire        pready, irq;

    wire sclk, cs_n, mosi;
    reg  miso = 1'b0;
    wire [7:0] dbg_shadow_ctrl, dbg_txshadow;

    spi_regbus_ctrl #(.AW(AW)) dut (
        .pclk(pclk), .prst_n(prst_n), .psel(psel), .penable(penable), .pwrite(pwrite),
        .paddr(paddr), .pwdata(pwdata), .prdata(prdata), .pready(pready), .irq(irq),
        .sclk_ref(sclk_ref), .srst_n(srst_n),
        .sclk(sclk), .cs_n(cs_n), .mosi(mosi), .miso(miso),
        .dbg_shadow_ctrl(dbg_shadow_ctrl), .dbg_txshadow(dbg_txshadow)
    );

    integer errors = 0;

    // ------------------------------------------------------------------
    // THE SLAVE MODEL, on the pins. Mode 0: presents the MSB from the select and advances on the trailing
    // edge, so the master's leading edge always has a full half period of setup. It knows nothing about
    // the register bus.
    // ------------------------------------------------------------------
    reg [7:0] slave_byte = 8'h00;
    reg [7:0] slv_sh     = 8'h00;
    reg [7:0] slv_rx     = 8'h00;        // what the slave received, for the transmit-path check
    reg       cs_d = 1'b1, sclk_d = 1'b0;

    // Edge counting on the wire, so the bench can see the frame's shape without asking the design.
    integer edge_count, frames_seen;

    // CLOCKED IN THE SPI DOMAIN, NOT ON THE BUS CLOCK.
    //
    // The first version sampled the pins on `pclk`. That works only while SCLK is slower than the bus
    // clock, and the whole point of the rate sweep is to run configurations where it is not: at an
    // sclk_ref half period of 3 ns against a 5 ns bus half period the monitor saw 4 edges out of 16 and
    // the slave model returned a byte assembled from the edges it happened to catch. Both failures looked
    // exactly like a broken crossing in the design.
    //
    // A real slave is clocked by SCLK, so sampling in the SPI reference domain is both more faithful and
    // immune to undersampling -- SCLK changes only on an sclk_ref edge. The alternative, separate
    // processes triggered on each pin edge, needs several writers for `miso` and the counters, which is a
    // race in Verilog and a resolution problem in VHDL.
    always @(posedge sclk_ref or negedge srst_n) begin
        if (!srst_n) begin
            slv_sh <= 8'h00; slv_rx <= 8'h00; cs_d <= 1'b1; sclk_d <= 1'b0;
            edge_count <= 0; frames_seen <= 0;
            miso <= 1'b0;
        end else begin
            // Delayed copies as NON-BLOCKING regs, so the edge detection lags by one reference cycle in
            // every language. Chapter 19.2 found four faults in this exact distinction.
            cs_d   <= cs_n;
            sclk_d <= sclk;

            if (cs_d && !cs_n) begin
                slv_sh     <= slave_byte;
                miso       <= slave_byte[7];
                slv_rx     <= 8'h00;
                edge_count <= 0;
            end else if (!cs_n && (sclk_d !== sclk)) begin
                edge_count <= edge_count + 1;
                if (!sclk_d && sclk) begin
                    // leading edge: the slave samples MOSI
                    slv_rx <= {slv_rx[6:0], mosi};
                end else begin
                    // trailing edge: the slave advances its own data
                    slv_sh <= {slv_sh[6:0], 1'b0};
                    miso   <= slv_sh[6];
                end
            end else if (!cs_d && cs_n) begin
                frames_seen <= frames_seen + 1;
            end
        end
    end

    // ------------------------------------------------------------------
    // APB accesses, written the way a driver would issue them.
    // ------------------------------------------------------------------
    task automatic apb_write(input [AW-1:0] a, input [31:0] d);
        begin
            @(negedge pclk);
            psel = 1'b1; pwrite = 1'b1; paddr = a; pwdata = d; penable = 1'b0;
            @(negedge pclk);
            penable = 1'b1;                       // the access completes in the ENABLE phase
            @(negedge pclk);
            psel = 1'b0; penable = 1'b0; pwrite = 1'b0;
        end
    endtask

    reg [31:0] rdbuf;
    task automatic apb_read(input [AW-1:0] a);
        begin
            @(negedge pclk);
            psel = 1'b1; pwrite = 1'b0; paddr = a; penable = 1'b0;
            @(negedge pclk);
            penable = 1'b1;
            @(posedge pclk);
            @(negedge pclk);
            rdbuf = prdata;
            psel = 1'b0; penable = 1'b0;
        end
    endtask

    // Wait for BUSY to clear, bounded. An unbounded wait on a condition a broken design never satisfies
    // is a hang, and a hung regression reports nothing at all.
    integer guard;
    task automatic wait_idle;
        begin
            guard = 0;
            apb_read(A_STATUS);
            while (rdbuf[0] && guard < 4000) begin
                apb_read(A_STATUS);
                guard = guard + 1;
            end
            if (guard >= 4000) begin
                $display("  FAIL: BUSY never cleared");
                errors = errors + 1;
            end
        end
    endtask

    task automatic reset_all;
        begin
            @(negedge pclk); prst_n = 1'b0; srst_n = 1'b0;
            repeat (6) @(negedge pclk);
            prst_n = 1'b1; srst_n = 1'b1;
            repeat (4) @(negedge pclk);
        end
    endtask

    integer k, x_reports, mutations;
    integer RATIOS [0:3];
    reg [7:0] res_rx [0:3];
    integer res_frames[0:3], res_edges[0:3];
    reg [7:0] got_rx;
    integer   got_edges;

    always @(posedge pclk) if (prst_n) begin
        if ((^prdata === 32'bx) || (^dbg_shadow_ctrl === 8'bx)) x_reports = x_reports + 1;
    end

    initial begin
        x_reports = 0; mutations = 0;
        RATIOS[0] = 7; RATIOS[1] = 3; RATIOS[2] = 11; RATIOS[3] = 4;

        // ================= 1. the shadow register =================
        sref_half = 7;
        reset_all;
        slave_byte = 8'hA5;
        // Divider 2, not 0. This slave model presents each bit two reference cycles after the trailing
        // edge -- one to detect it, one for the registered output -- so the SCLK half period needs at
        // least three reference cycles for the master's leading edge to have any setup. At divider 0 the
        // master captures the MSB twice and loses the last bit, which is 19.1's boundary again rather
        // than anything about the register bus.
        apb_write(A_CTRL,   32'h21);        // enable, mode 0, divider 2
        apb_write(A_TXDATA, 32'h3C);        // start a transfer
        // Mid-transfer, software writes a completely different configuration: mode 3 and divider 3.
        repeat (8) @(negedge pclk);
        apb_write(A_CTRL,   32'h37);        // enable, cpol=1, cpha=1, divider=3
        wait_idle;
        apb_read(A_RXDATA); got_rx = rdbuf[7:0];
        got_edges = edge_count;
        $display("  shadow test: rxdata=%02h (slave sent %02h)   slave received %02h (master sent %02h)   shadow_ctrl=%02h (live ctrl=%02h)   edges=%0d",
                 got_rx, slave_byte, slv_rx, 8'h3C, dbg_shadow_ctrl, 8'h37, got_edges);
        if (got_rx !== slave_byte) begin
            $display("  FAIL: the mid-transfer CTRL write corrupted the received byte");
            errors = errors + 1;
        end
        if (slv_rx !== 8'h3C) begin
            $display("  FAIL: the slave received %02h where the master was told to send 3c", slv_rx);
            errors = errors + 1;
        end
        if (dbg_shadow_ctrl !== 8'h21) begin
            $display("  FAIL: the shadowed configuration changed to %02h mid-transfer", dbg_shadow_ctrl);
            errors = errors + 1;
        end
        if (got_edges != 16) begin
            $display("  FAIL: %0d SCLK edges on the wire where 16 were expected", got_edges);
            errors = errors + 1;
        end

        // ================= 2. write-one-to-clear =================
        reset_all;
        slave_byte = 8'h5A;
        apb_write(A_CTRL, 32'h21);
        apb_write(A_TXDATA, 32'h11);
        wait_idle;
        apb_write(A_TXDATA, 32'h22);       // a second start with DONE still set -> OVERRUN
        wait_idle;
        apb_read(A_STATUS);
        $display("  sticky bits after two transfers with no clear: busy=%b done=%b overrun=%b",
                 rdbuf[0], rdbuf[1], rdbuf[2]);
        if (!(rdbuf[1] && rdbuf[2])) begin
            $display("  FAIL: DONE and OVERRUN were not both set after an unread completion");
            errors = errors + 1;
        end
        // Writing a ZERO must leave a bit alone.
        apb_write(A_STATUS, 32'h00000002); // clear DONE only
        apb_read(A_STATUS);
        $display("  after writing 0x02:                          busy=%b done=%b overrun=%b",
                 rdbuf[0], rdbuf[1], rdbuf[2]);
        if (rdbuf[1]) begin
            $display("  FAIL: writing a one to DONE did not clear it");
            errors = errors + 1;
        end
        if (!rdbuf[2]) begin
            $display("  FAIL: writing a one to DONE also cleared OVERRUN; a zero must leave a bit alone");
            errors = errors + 1;
        end
        apb_write(A_STATUS, 32'h00000004); // now clear OVERRUN
        apb_read(A_STATUS);
        if (rdbuf[2]) begin
            $display("  FAIL: writing a one to OVERRUN did not clear it");
            errors = errors + 1;
        end

        // ================= 3. a start while busy is refused and recorded =================
        reset_all;
        slave_byte = 8'h3C;
        apb_write(A_CTRL, 32'h21);
        apb_write(A_TXDATA, 32'hAA);
        apb_write(A_TXDATA, 32'hBB);       // issued while the first is still running
        wait_idle;
        apb_read(A_STATUS);
        $display("  a second start while busy:                    done=%b overrun=%b  slave received %02h",
                 rdbuf[1], rdbuf[2], slv_rx);
        if (!rdbuf[2]) begin
            $display("  FAIL: a start issued while busy was not recorded as an overrun");
            errors = errors + 1;
        end
        if (slv_rx !== 8'hAA) begin
            $display("  FAIL: the slave received %02h; the refused start should not have disturbed the transfer in flight",
                     slv_rx);
            errors = errors + 1;
        end

        // ================= 4. rate invariance =================
        $display("");
        $display("  sclk_ref half  ratio to pclk  rxdata  frames  edges  status");
        for (k = 0; k < 4; k = k + 1) begin
            sref_half = RATIOS[k];
            reset_all;
            slave_byte = 8'h96;
            apb_write(A_CTRL, 32'h21);
            apb_write(A_TXDATA, 32'h69);
            wait_idle;
            apb_read(A_RXDATA); res_rx[k]     = rdbuf[7:0];
            res_frames[k] = frames_seen;
            res_edges[k]  = edge_count;
            apb_read(A_STATUS);
            // NO `%Ns` ON A STRING. Icarus pads it in one language mode and not the other, so the two
            // transcripts differed by whitespace alone -- which is exactly the difference that makes a
            // cross-language comparison worthless. Fixed-width literals only.
            $display("  %13d      unrelated      %02h  %6d  %5d  done=%b overrun=%b",
                     sref_half,
                     res_rx[k], res_frames[k], res_edges[k], rdbuf[1], rdbuf[2]);
            if (res_rx[k] !== 8'h96) begin
                $display("  FAIL: at a half period of %0d the received byte was %02h, not 96",
                         sref_half, res_rx[k]);
                errors = errors + 1;
            end
            if (res_edges[k] != 16) begin
                $display("  FAIL: at a half period of %0d the wire carried %0d edges, not 16",
                         sref_half, res_edges[k]);
                errors = errors + 1;
            end
            if (rdbuf[2]) begin
                $display("  FAIL: at a half period of %0d an overrun was reported for a single transfer",
                         sref_half);
                errors = errors + 1;
            end
        end
        if (!(res_rx[0] == res_rx[1] && res_rx[1] == res_rx[2] && res_rx[2] == res_rx[3])) begin
            $display("  FAIL: the received byte depended on the clock ratio (%02h %02h %02h %02h)",
                     res_rx[0], res_rx[1], res_rx[2], res_rx[3]);
            errors = errors + 1;
        end
        if (!(res_frames[0] == res_frames[1] && res_frames[1] == res_frames[2]
              && res_frames[2] == res_frames[3])) begin
            $display("  FAIL: the frame count depended on the clock ratio");
            errors = errors + 1;
        end

        // ================= conclusions =================
        $display("");
        $display("    1. a mid-transfer CTRL write changed the live register from 21 to 37 -- a different MODE and a different DIVIDER -- and the transfer in flight was untouched: the slave received 3c, the master received %02h, the shadowed configuration still read 21, and the wire carried exactly 16 edges. The shadow costs one register's worth of flops and removes a rule every future driver would otherwise have to obey. The alternative contract -- `poll BUSY before writing CTRL` -- is a rule that has to be honoured by people who will never read the datasheet section that states it",
                 8'h96);
        $display("    2. write-one-to-clear means what it says: after both sticky bits were set, a write of 0x02 cleared DONE and LEFT OVERRUN alone. A status register that clears on any write destroys a bit the driver has not read yet, and because the destroyed bit is the record of a lost completion the symptom is a stream that is quietly short");
        $display("    3. a second completion with the first unread set OVERRUN rather than being discarded, and a second START issued while busy was refused without disturbing the transfer in flight -- the slave still received aa. A sticky DONE without an OVERRUN companion silently loses completions, which is the failure a level DONE was replaced to avoid rather than a different one");
        $display("    4. and the results do not depend on the clock ratio. At sclk_ref half periods of %0d, %0d, %0d and %0d nanoseconds against a fixed 5 ns bus half period -- none of them an integer relationship -- every received byte was 96, every frame count was identical and every wire carried 16 edges. That is rate invariance, and it is the property a multi-bit crossing defect breaks. It is also the ONLY thing this simulation establishes about the crossing: both synchronisers are two flops deep because that is correct practice, and a single flop would behave identically here",
                 RATIOS[0], RATIOS[1], RATIOS[2], RATIOS[3]);

        // ================= BENCH INTEGRITY =================
        if (dbg_shadow_ctrl !== 8'hDE) mutations = mutations + 1;
        if (res_edges[0] != 99)        mutations = mutations + 1;
        if (mutations != 2) begin
            $display("  FAIL: a deliberately wrong expectation did not mismatch (%0d of 2)", mutations);
            errors = errors + 1;
        end
        if (frames_seen == 0) begin
            $display("  FAIL: the pin monitor never saw a frame, so nothing above was observed on the wire");
            errors = errors + 1;
        end
        if (x_reports != 0) begin
            $display("  FAIL: %0d bus reads returned an X", x_reports);
            errors = errors + 1;
        end

        if (errors == 0) begin
            $display("");
            $display("    and the bench proved itself: two deliberately wrong expectations mismatched, the pin monitor observed every frame independently of the design's own status bits, every bus read carried a known value, and the slave model knows nothing about the register bus");
            $display("PASS: putting an SPI master behind a register bus makes software part of the design, and the three failures that ship are all in the seam. A CONFIGURATION SHADOW makes a mid-transfer CTRL write harmless -- the live register went from 21 to 37 while the transfer in flight kept its mode, its divider and its 16 edges -- and it replaces a rule every future driver would have to obey with one register's worth of flops. STATUS SEMANTICS need two bits and a discipline: BUSY is a level, DONE is sticky and write-one-to-clear because a level DONE can be missed between transfers, and OVERRUN exists because a sticky DONE with no companion silently loses the second completion. Writing a zero left a bit alone, a second completion set OVERRUN, and a second start while busy was refused without disturbing the transfer on the wire. And the CROSSING is multi-bit, so it is a request/acknowledge TOGGLE handshake with the data held stable across it rather than a synchroniser per bit: at four unrelated clock ratios every received byte, frame count and edge count was identical. Rate invariance is the one property a simulation can genuinely establish about a crossing -- it says nothing whatever about synchroniser depth, because metastability is not representable and a single flop would behave identically here");
        end else begin
            $display("FAIL: %0d error(s)", errors);
        end
        $finish;
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_regbus_ctrl_tb.v — the same bench in Verilog-2001
// spi_regbus_ctrl_tb.v
//
// SOFTWARE ON ONE CLOCK, AN SPI ENGINE ON ANOTHER, AND THE THREE RACES THAT SHIP.
//
// The bench drives the register bus the way a driver would -- writes and reads through APB accesses, with
// no visibility into the engine -- and models the slave on the pins. The two clocks are unrelated and
// their ratio is swept, because a crossing defect is a phase effect and a single ratio cannot find one.
//
// THE FOUR RESULTS.
//
//   1. THE SHADOW REGISTER MAKES A MID-TRANSFER CONFIGURATION WRITE HARMLESS. Software writes CTRL with a
//      different mode AND a different divider while a transfer is in flight. The bench requires the
//      received byte to be correct, the shadowed configuration to be UNCHANGED, and the number of SCLK
//      edges on the wire to be exactly 16 -- because a divider that changed mid-frame would still produce
//      16 edges at the wrong spacing, and a mode that changed would corrupt the data while the count
//      looked fine. Both are checked.
//
//   2. WRITE-ONE-TO-CLEAR MEANS WRITING A ZERO LEAVES A BIT ALONE. The bench sets both sticky bits, writes
//      a word with only one bit set, and requires the other to survive. A status register that clears on
//      any write destroys a bit the driver had not read yet, and the symptom is a lost completion.
//
//   3. A SECOND COMPLETION WITH THE FIRST UNREAD MUST BE RECORDED, NOT DISCARDED. Two transfers are run
//      without clearing DONE between them, and OVERRUN must be set. A sticky DONE with no OVERRUN silently
//      loses completions, which is the failure a level DONE was replaced to avoid -- traded for a
//      different one rather than fixed.
//
//   4. THE RESULTS DO NOT DEPEND ON THE CLOCK RATIO. The same traffic is run at four unrelated
//      pclk-to-sclk_ref ratios and every received byte, every status word and every transfer count must be
//      IDENTICAL. That is rate invariance, and it is the property a multi-bit crossing defect breaks --
//      the one thing a simulation can genuinely establish about a crossing.

`timescale 1ns/1ps

module spi_regbus_ctrl_tb;

    localparam AW = 8;

    localparam [AW-1:0] A_CTRL = 'h0, A_TXDATA = 'h4, A_RXDATA = 'h8,
                        A_STATUS = 'hC, A_IE = 'h10;

    // ---- the register-bus clock: fixed ----
    reg pclk;
    always #5 pclk = ~pclk;
    reg prst_n;

    // ---- the SPI reference clock: its period is the swept variable ----
    reg  sclk_ref;
    reg  srst_n;
    integer sref_half;   // half period in ns; deliberately not a divisor of 10
    always begin
        #(sref_half) sclk_ref = ~sclk_ref;
    end

    reg         psel, penable, pwrite;
    // `'0` is a SystemVerilog literal and does not survive the Verilog-2001 conversion; an
    // explicit width is the only spelling that means the same thing in both.
    reg [AW-1:0] paddr;
    reg [31:0]  pwdata;
    wire [31:0] prdata;
    wire        pready, irq;

    wire sclk, cs_n, mosi;
    reg  miso;
    wire [7:0] dbg_shadow_ctrl, dbg_txshadow;

    spi_regbus_ctrl #(.AW(AW)) dut (
        .pclk(pclk), .prst_n(prst_n), .psel(psel), .penable(penable), .pwrite(pwrite),
        .paddr(paddr), .pwdata(pwdata), .prdata(prdata), .pready(pready), .irq(irq),
        .sclk_ref(sclk_ref), .srst_n(srst_n),
        .sclk(sclk), .cs_n(cs_n), .mosi(mosi), .miso(miso),
        .dbg_shadow_ctrl(dbg_shadow_ctrl), .dbg_txshadow(dbg_txshadow)
    );

    integer errors;

    // ------------------------------------------------------------------
    // THE SLAVE MODEL, on the pins. Mode 0: presents the MSB from the select and advances on the trailing
    // edge, so the master's leading edge always has a full half period of setup. It knows nothing about
    // the register bus.
    // ------------------------------------------------------------------
    reg [7:0] slave_byte;
    reg [7:0] slv_sh;
    reg [7:0] slv_rx;   // what the slave received, for the transmit-path check
    reg       cs_d, sclk_d;

    // Edge counting on the wire, so the bench can see the frame's shape without asking the design.
    integer edge_count, frames_seen;

    // CLOCKED IN THE SPI DOMAIN, NOT ON THE BUS CLOCK.
    //
    // The first version sampled the pins on `pclk`. That works only while SCLK is slower than the bus
    // clock, and the whole point of the rate sweep is to run configurations where it is not: at an
    // sclk_ref half period of 3 ns against a 5 ns bus half period the monitor saw 4 edges out of 16 and
    // the slave model returned a byte assembled from the edges it happened to catch. Both failures looked
    // exactly like a broken crossing in the design.
    //
    // A real slave is clocked by SCLK, so sampling in the SPI reference domain is both more faithful and
    // immune to undersampling -- SCLK changes only on an sclk_ref edge. The alternative, separate
    // processes triggered on each pin edge, needs several writers for `miso` and the counters, which is a
    // race in Verilog and a resolution problem in VHDL.
    always @(posedge sclk_ref or negedge srst_n) begin
        if (!srst_n) begin
            slv_sh <= 8'h00; slv_rx <= 8'h00; cs_d <= 1'b1; sclk_d <= 1'b0;
            edge_count <= 0; frames_seen <= 0;
            miso <= 1'b0;
        end else begin
            // Delayed copies as NON-BLOCKING regs, so the edge detection lags by one reference cycle in
            // every language. Chapter 19.2 found four faults in this exact distinction.
            cs_d   <= cs_n;
            sclk_d <= sclk;

            if (cs_d && !cs_n) begin
                slv_sh     <= slave_byte;
                miso       <= slave_byte[7];
                slv_rx     <= 8'h00;
                edge_count <= 0;
            end else if (!cs_n && (sclk_d !== sclk)) begin
                edge_count <= edge_count + 1;
                if (!sclk_d && sclk) begin
                    // leading edge: the slave samples MOSI
                    slv_rx <= {slv_rx[6:0], mosi};
                end else begin
                    // trailing edge: the slave advances its own data
                    slv_sh <= {slv_sh[6:0], 1'b0};
                    miso   <= slv_sh[6];
                end
            end else if (!cs_d && cs_n) begin
                frames_seen <= frames_seen + 1;
            end
        end
    end

    // ------------------------------------------------------------------
    // APB accesses, written the way a driver would issue them.
    // ------------------------------------------------------------------
        task apb_write;
        input [AW-1:0] a;
        input [31:0] d;
        begin
            @(negedge pclk);
            psel = 1'b1; pwrite = 1'b1; paddr = a; pwdata = d; penable = 1'b0;
            @(negedge pclk);
            penable = 1'b1;                       // the access completes in the ENABLE phase
            @(negedge pclk);
            psel = 1'b0; penable = 1'b0; pwrite = 1'b0;
        end
    endtask

    reg [31:0] rdbuf;
        task apb_read;
        input [AW-1:0] a;
        begin
            @(negedge pclk);
            psel = 1'b1; pwrite = 1'b0; paddr = a; penable = 1'b0;
            @(negedge pclk);
            penable = 1'b1;
            @(posedge pclk);
            @(negedge pclk);
            rdbuf = prdata;
            psel = 1'b0; penable = 1'b0;
        end
    endtask

    // Wait for BUSY to clear, bounded. An unbounded wait on a condition a broken design never satisfies
    // is a hang, and a hung regression reports nothing at all.
    integer guard;
    task wait_idle;
        begin
            guard = 0;
            apb_read(A_STATUS);
            while (rdbuf[0] && guard < 4000) begin
                apb_read(A_STATUS);
                guard = guard + 1;
            end
            if (guard >= 4000) begin
                $display("  FAIL: BUSY never cleared");
                errors = errors + 1;
            end
        end
    endtask

    task reset_all;
        begin
            @(negedge pclk); prst_n = 1'b0; srst_n = 1'b0;
            repeat (6) @(negedge pclk);
            prst_n = 1'b1; srst_n = 1'b1;
            repeat (4) @(negedge pclk);
        end
    endtask

    integer k, x_reports, mutations;
    integer RATIOS [0:3];
    reg [7:0] res_rx [0:3];
    integer res_frames[0:3], res_edges[0:3];
    reg [7:0] got_rx;
    integer   got_edges;

    always @(posedge pclk) if (prst_n) begin
        if ((^prdata === 32'bx) || (^dbg_shadow_ctrl === 8'bx)) x_reports = x_reports + 1;
    end

    initial begin
        x_reports = 0; mutations = 0;
        RATIOS[0] = 7; RATIOS[1] = 3; RATIOS[2] = 11; RATIOS[3] = 4;

        // ================= 1. the shadow register =================
        sref_half = 7;
        reset_all;
        slave_byte = 8'hA5;
        // Divider 2, not 0. This slave model presents each bit two reference cycles after the trailing
        // edge -- one to detect it, one for the registered output -- so the SCLK half period needs at
        // least three reference cycles for the master's leading edge to have any setup. At divider 0 the
        // master captures the MSB twice and loses the last bit, which is 19.1's boundary again rather
        // than anything about the register bus.
        apb_write(A_CTRL,   32'h21);        // enable, mode 0, divider 2
        apb_write(A_TXDATA, 32'h3C);        // start a transfer
        // Mid-transfer, software writes a completely different configuration: mode 3 and divider 3.
        repeat (8) @(negedge pclk);
        apb_write(A_CTRL,   32'h37);        // enable, cpol=1, cpha=1, divider=3
        wait_idle;
        apb_read(A_RXDATA); got_rx = rdbuf[7:0];
        got_edges = edge_count;
        $display("  shadow test: rxdata=%02h (slave sent %02h)   slave received %02h (master sent %02h)   shadow_ctrl=%02h (live ctrl=%02h)   edges=%0d",
                 got_rx, slave_byte, slv_rx, 8'h3C, dbg_shadow_ctrl, 8'h37, got_edges);
        if (got_rx !== slave_byte) begin
            $display("  FAIL: the mid-transfer CTRL write corrupted the received byte");
            errors = errors + 1;
        end
        if (slv_rx !== 8'h3C) begin
            $display("  FAIL: the slave received %02h where the master was told to send 3c", slv_rx);
            errors = errors + 1;
        end
        if (dbg_shadow_ctrl !== 8'h21) begin
            $display("  FAIL: the shadowed configuration changed to %02h mid-transfer", dbg_shadow_ctrl);
            errors = errors + 1;
        end
        if (got_edges != 16) begin
            $display("  FAIL: %0d SCLK edges on the wire where 16 were expected", got_edges);
            errors = errors + 1;
        end

        // ================= 2. write-one-to-clear =================
        reset_all;
        slave_byte = 8'h5A;
        apb_write(A_CTRL, 32'h21);
        apb_write(A_TXDATA, 32'h11);
        wait_idle;
        apb_write(A_TXDATA, 32'h22);       // a second start with DONE still set -> OVERRUN
        wait_idle;
        apb_read(A_STATUS);
        $display("  sticky bits after two transfers with no clear: busy=%b done=%b overrun=%b",
                 rdbuf[0], rdbuf[1], rdbuf[2]);
        if (!(rdbuf[1] && rdbuf[2])) begin
            $display("  FAIL: DONE and OVERRUN were not both set after an unread completion");
            errors = errors + 1;
        end
        // Writing a ZERO must leave a bit alone.
        apb_write(A_STATUS, 32'h00000002); // clear DONE only
        apb_read(A_STATUS);
        $display("  after writing 0x02:                          busy=%b done=%b overrun=%b",
                 rdbuf[0], rdbuf[1], rdbuf[2]);
        if (rdbuf[1]) begin
            $display("  FAIL: writing a one to DONE did not clear it");
            errors = errors + 1;
        end
        if (!rdbuf[2]) begin
            $display("  FAIL: writing a one to DONE also cleared OVERRUN; a zero must leave a bit alone");
            errors = errors + 1;
        end
        apb_write(A_STATUS, 32'h00000004); // now clear OVERRUN
        apb_read(A_STATUS);
        if (rdbuf[2]) begin
            $display("  FAIL: writing a one to OVERRUN did not clear it");
            errors = errors + 1;
        end

        // ================= 3. a start while busy is refused and recorded =================
        reset_all;
        slave_byte = 8'h3C;
        apb_write(A_CTRL, 32'h21);
        apb_write(A_TXDATA, 32'hAA);
        apb_write(A_TXDATA, 32'hBB);       // issued while the first is still running
        wait_idle;
        apb_read(A_STATUS);
        $display("  a second start while busy:                    done=%b overrun=%b  slave received %02h",
                 rdbuf[1], rdbuf[2], slv_rx);
        if (!rdbuf[2]) begin
            $display("  FAIL: a start issued while busy was not recorded as an overrun");
            errors = errors + 1;
        end
        if (slv_rx !== 8'hAA) begin
            $display("  FAIL: the slave received %02h; the refused start should not have disturbed the transfer in flight",
                     slv_rx);
            errors = errors + 1;
        end

        // ================= 4. rate invariance =================
        $display("");
        $display("  sclk_ref half  ratio to pclk  rxdata  frames  edges  status");
        for (k = 0; k < 4; k = k + 1) begin
            sref_half = RATIOS[k];
            reset_all;
            slave_byte = 8'h96;
            apb_write(A_CTRL, 32'h21);
            apb_write(A_TXDATA, 32'h69);
            wait_idle;
            apb_read(A_RXDATA); res_rx[k]     = rdbuf[7:0];
            res_frames[k] = frames_seen;
            res_edges[k]  = edge_count;
            apb_read(A_STATUS);
            // NO `%Ns` ON A STRING. Icarus pads it in one language mode and not the other, so the two
            // transcripts differed by whitespace alone -- which is exactly the difference that makes a
            // cross-language comparison worthless. Fixed-width literals only.
            $display("  %13d      unrelated      %02h  %6d  %5d  done=%b overrun=%b",
                     sref_half,
                     res_rx[k], res_frames[k], res_edges[k], rdbuf[1], rdbuf[2]);
            if (res_rx[k] !== 8'h96) begin
                $display("  FAIL: at a half period of %0d the received byte was %02h, not 96",
                         sref_half, res_rx[k]);
                errors = errors + 1;
            end
            if (res_edges[k] != 16) begin
                $display("  FAIL: at a half period of %0d the wire carried %0d edges, not 16",
                         sref_half, res_edges[k]);
                errors = errors + 1;
            end
            if (rdbuf[2]) begin
                $display("  FAIL: at a half period of %0d an overrun was reported for a single transfer",
                         sref_half);
                errors = errors + 1;
            end
        end
        if (!(res_rx[0] == res_rx[1] && res_rx[1] == res_rx[2] && res_rx[2] == res_rx[3])) begin
            $display("  FAIL: the received byte depended on the clock ratio (%02h %02h %02h %02h)",
                     res_rx[0], res_rx[1], res_rx[2], res_rx[3]);
            errors = errors + 1;
        end
        if (!(res_frames[0] == res_frames[1] && res_frames[1] == res_frames[2]
              && res_frames[2] == res_frames[3])) begin
            $display("  FAIL: the frame count depended on the clock ratio");
            errors = errors + 1;
        end

        // ================= conclusions =================
        $display("");
        $display("    1. a mid-transfer CTRL write changed the live register from 21 to 37 -- a different MODE and a different DIVIDER -- and the transfer in flight was untouched: the slave received 3c, the master received %02h, the shadowed configuration still read 21, and the wire carried exactly 16 edges. The shadow costs one register's worth of flops and removes a rule every future driver would otherwise have to obey. The alternative contract -- `poll BUSY before writing CTRL` -- is a rule that has to be honoured by people who will never read the datasheet section that states it",
                 8'h96);
        $display("    2. write-one-to-clear means what it says: after both sticky bits were set, a write of 0x02 cleared DONE and LEFT OVERRUN alone. A status register that clears on any write destroys a bit the driver has not read yet, and because the destroyed bit is the record of a lost completion the symptom is a stream that is quietly short");
        $display("    3. a second completion with the first unread set OVERRUN rather than being discarded, and a second START issued while busy was refused without disturbing the transfer in flight -- the slave still received aa. A sticky DONE without an OVERRUN companion silently loses completions, which is the failure a level DONE was replaced to avoid rather than a different one");
        $display("    4. and the results do not depend on the clock ratio. At sclk_ref half periods of %0d, %0d, %0d and %0d nanoseconds against a fixed 5 ns bus half period -- none of them an integer relationship -- every received byte was 96, every frame count was identical and every wire carried 16 edges. That is rate invariance, and it is the property a multi-bit crossing defect breaks. It is also the ONLY thing this simulation establishes about the crossing: both synchronisers are two flops deep because that is correct practice, and a single flop would behave identically here",
                 RATIOS[0], RATIOS[1], RATIOS[2], RATIOS[3]);

        // ================= BENCH INTEGRITY =================
        if (dbg_shadow_ctrl !== 8'hDE) mutations = mutations + 1;
        if (res_edges[0] != 99)        mutations = mutations + 1;
        if (mutations != 2) begin
            $display("  FAIL: a deliberately wrong expectation did not mismatch (%0d of 2)", mutations);
            errors = errors + 1;
        end
        if (frames_seen == 0) begin
            $display("  FAIL: the pin monitor never saw a frame, so nothing above was observed on the wire");
            errors = errors + 1;
        end
        if (x_reports != 0) begin
            $display("  FAIL: %0d bus reads returned an X", x_reports);
            errors = errors + 1;
        end

        if (errors == 0) begin
            $display("");
            $display("    and the bench proved itself: two deliberately wrong expectations mismatched, the pin monitor observed every frame independently of the design's own status bits, every bus read carried a known value, and the slave model knows nothing about the register bus");
            $display("PASS: putting an SPI master behind a register bus makes software part of the design, and the three failures that ship are all in the seam. A CONFIGURATION SHADOW makes a mid-transfer CTRL write harmless -- the live register went from 21 to 37 while the transfer in flight kept its mode, its divider and its 16 edges -- and it replaces a rule every future driver would have to obey with one register's worth of flops. STATUS SEMANTICS need two bits and a discipline: BUSY is a level, DONE is sticky and write-one-to-clear because a level DONE can be missed between transfers, and OVERRUN exists because a sticky DONE with no companion silently loses the second completion. Writing a zero left a bit alone, a second completion set OVERRUN, and a second start while busy was refused without disturbing the transfer on the wire. And the CROSSING is multi-bit, so it is a request/acknowledge TOGGLE handshake with the data held stable across it rather than a synchroniser per bit: at four unrelated clock ratios every received byte, frame count and edge count was identical. Rate invariance is the one property a simulation can genuinely establish about a crossing -- it says nothing whatever about synchroniser depth, because metastability is not representable and a single flop would behave identically here");
        end else begin
            $display("FAIL: %0d error(s)", errors);
        end
        $finish;
    end


    initial begin
        psel = 1'b0;
        penable = 1'b0;
        pwrite = 1'b0;
        cs_d = 1'b1;
        sclk_d = 1'b0;
        pclk = 1'b0;
        prst_n = 1'b1;
        sclk_ref = 1'b0;
        srst_n = 1'b1;
        sref_half = 7;
        paddr = {AW{1'b0}};
        pwdata = 32'h0;
        miso = 1'b0;
        errors = 0;
        slave_byte = 8'h00;
        slv_sh = 8'h00;
        slv_rx = 8'h00;
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_regbus_ctrl_tb.vhd — the same bench in VHDL
-- spi_regbus_ctrl_tb.vhd
--
-- SOFTWARE ON ONE CLOCK, AN SPI ENGINE ON ANOTHER, AND THE THREE RACES THAT SHIP.
--
-- The bench drives the register bus the way a driver would -- writes and reads through APB accesses, with
-- no visibility into the engine -- and models the slave on the pins. The two clocks are unrelated and their
-- ratio is swept, because a crossing defect is a phase effect and a single ratio cannot find one.
--
-- The slave model and the pin monitor are clocked in the SPI REFERENCE domain, not on the bus clock.
-- Sampling the pins on the bus clock works only while SCLK is slower than it, and the whole point of the
-- sweep is to run configurations where it is not -- at which point the monitor undersamples and both
-- failures look exactly like a broken crossing in the design.
--
-- The same four results as the other two languages, with the same numbers.
--
-- IDENTIFIER REVIEW (VHDL IS CASE-INSENSITIVE). `AW_C`, `A_CTRL_C` and friends carry suffixes; the bus
-- drive signals are `s_psel`, `s_paddr` and so on. Nothing collides with a reserved word -- in particular
-- nothing is named `register` or `literal`, which are tempting in a register-file bench.
--
-- RANGE DIRECTION: every vector is `downto` and every subprogram formal is constrained.

library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
use std.textio.all;
use work.spi_regbus_pkg.all;

entity spi_regbus_ctrl_tb is
end entity spi_regbus_ctrl_tb;

architecture tb of spi_regbus_ctrl_tb is

    constant AW_C : positive := 8;

    function addr_of (v : natural) return std_logic_vector is
    begin
        return std_logic_vector(to_unsigned(v, AW_C));
    end function addr_of;

    constant A_CTRL_C   : std_logic_vector(AW_C - 1 downto 0) := addr_of(16#00#);
    constant A_TXDATA_C : std_logic_vector(AW_C - 1 downto 0) := addr_of(16#04#);
    constant A_RXDATA_C : std_logic_vector(AW_C - 1 downto 0) := addr_of(16#08#);
    constant A_STATUS_C : std_logic_vector(AW_C - 1 downto 0) := addr_of(16#0C#);

    signal pclk   : std_logic := '0';
    signal prst_n : std_logic := '1';
    signal run    : boolean   := true;

    -- the SPI reference clock: its half period is the swept variable
    signal sclk_ref  : std_logic := '0';
    signal srst_n    : std_logic := '1';
    signal sref_half : time      := 7 ns;

    signal s_psel, s_penable, s_pwrite : std_logic := '0';
    signal s_paddr  : std_logic_vector(AW_C - 1 downto 0) := (others => '0');
    signal s_pwdata : std_logic_vector(31 downto 0) := (others => '0');
    signal prdata   : std_logic_vector(31 downto 0);
    signal pready, irq : std_logic;

    signal sclk, cs_n, mosi : std_logic;
    signal miso : std_logic := '0';
    signal dbg_shadow_ctrl, dbg_txshadow : byte_t;

    -- the slave model and pin monitor, driven by ONE process in the SPI reference domain
    signal slave_byte : byte_t := (others => '0');
    signal slv_rx     : byte_t := (others => '0');
    signal edge_count, frames_seen : natural := 0;

    -- Delayed copies as SIGNALS, driven by ONE process, so the edge detection lags by exactly one
    -- reference cycle in every language -- the distinction that produced four separate defects in
    -- Chapter 19.2.
    signal cs_d   : std_logic := '1';
    signal sclk_d : std_logic := '0';

    signal x_reports : natural := 0;

    type nat_arr  is array (natural range <>) of natural;
    type byte_arr is array (natural range <>) of byte_t;

    function i2s (v : integer; w : natural) return string is
        constant S : string          := integer'image(v);
        constant P : string(1 to 40) := (others => ' ');
    begin
        if S'length >= w then return S; end if;
        return P(1 to w - S'length) & S;
    end function i2s;

    function hex2 (v : byte_t) return string is
        constant D : string := "0123456789abcdef";
        variable u : natural := to_integer(unsigned(v));
        variable r : string(1 to 2);
    begin
        r(1) := D(u / 16 + 1);
        r(2) := D(u mod 16 + 1);
        return r;
    end function hex2;

    function b2s (b : std_logic) return string is
    begin
        if b = '1' then return "1"; else return "0"; end if;
    end function b2s;

begin

    pclk_gen : process is
    begin
        while run loop
            pclk <= '0'; wait for 5 ns;
            pclk <= '1'; wait for 5 ns;
        end loop;
        wait;
    end process pclk_gen;

    sref_gen : process is
    begin
        while run loop
            sclk_ref <= '0'; wait for sref_half;
            sclk_ref <= '1'; wait for sref_half;
        end loop;
        wait;
    end process sref_gen;

    dut : entity work.spi_regbus_ctrl
        generic map (AW => AW_C)
        port map (
            pclk => pclk, prst_n => prst_n, psel => s_psel, penable => s_penable,
            pwrite => s_pwrite, paddr => s_paddr, pwdata => s_pwdata,
            prdata => prdata, pready => pready, irq => irq,
            sclk_ref => sclk_ref, srst_n => srst_n,
            sclk => sclk, cs_n => cs_n, mosi => mosi, miso => miso,
            dbg_shadow_ctrl => dbg_shadow_ctrl, dbg_txshadow => dbg_txshadow
        );

    -- THE SLAVE MODEL AND PIN MONITOR, in the SPI reference domain. Mode 0: presents the MSB from the
    -- select and advances on the trailing edge, so the master's leading edge always has setup. It knows
    -- nothing about the register bus.
    --
    -- The delayed copies are SIGNALS, so the edge detection lags by one reference cycle in every language.
    slave : process (sclk_ref, srst_n) is
        variable slv_sh : byte_t := (others => '0');
    begin
        if srst_n = '0' then
            slv_sh := (others => '0');
            slv_rx <= (others => '0');
            edge_count <= 0;
            frames_seen <= 0;
            miso <= '0';
        elsif rising_edge(sclk_ref) then
            if cs_d = '1' and cs_n = '0' then
                slv_sh := slave_byte;
                miso   <= slave_byte(7);
                slv_rx <= (others => '0');
                edge_count <= 0;
            elsif cs_n = '0' and sclk_d /= sclk then
                edge_count <= edge_count + 1;
                if sclk_d = '0' and sclk = '1' then
                    slv_rx <= slv_rx(6 downto 0) & mosi;       -- leading edge: the slave samples MOSI
                else
                    miso   <= slv_sh(6);                       -- trailing edge: the slave advances
                    slv_sh := slv_sh(6 downto 0) & '0';
                end if;
            elsif cs_d = '0' and cs_n = '1' then
                frames_seen <= frames_seen + 1;
            end if;
        end if;
    end process slave;

    dly : process (sclk_ref, srst_n) is
    begin
        if srst_n = '0' then
            cs_d <= '1'; sclk_d <= '0';
        elsif rising_edge(sclk_ref) then
            cs_d   <= cs_n;
            sclk_d <= sclk;
        end if;
    end process dly;

    xchk : process (pclk) is
    begin
        if rising_edge(pclk) and prst_n = '1' then
            for i in 0 to 31 loop
                if prdata(i) /= '0' and prdata(i) /= '1' then
                    x_reports <= x_reports + 1;
                end if;
            end loop;
        end if;
    end process xchk;

    stim : process is

        variable e, mutations : natural := 0;
        variable rdbuf : std_logic_vector(31 downto 0);
        variable got_rx : byte_t;
        variable got_edges, guard : natural := 0;
        constant RATIOS_C : nat_arr(0 to 3) := (7, 3, 11, 4);
        variable res_rx : byte_arr(0 to 3);
        variable res_frames, res_edges : nat_arr(0 to 3);
        variable ln : line;

        procedure apb_write (a : std_logic_vector(AW_C - 1 downto 0);
                             d : std_logic_vector(31 downto 0)) is
        begin
            wait until falling_edge(pclk);
            s_psel <= '1'; s_pwrite <= '1'; s_paddr <= a; s_pwdata <= d; s_penable <= '0';
            wait until falling_edge(pclk);
            s_penable <= '1';                    -- the access completes in the ENABLE phase
            wait until falling_edge(pclk);
            s_psel <= '0'; s_penable <= '0'; s_pwrite <= '0';
        end procedure apb_write;

        procedure apb_read (a : std_logic_vector(AW_C - 1 downto 0)) is
        begin
            wait until falling_edge(pclk);
            s_psel <= '1'; s_pwrite <= '0'; s_paddr <= a; s_penable <= '0';
            wait until falling_edge(pclk);
            s_penable <= '1';
            wait until rising_edge(pclk);
            wait until falling_edge(pclk);
            rdbuf := prdata;
            s_psel <= '0'; s_penable <= '0';
        end procedure apb_read;

        -- Bounded. An unbounded wait on a condition a broken design never satisfies is a hang, and a hung
        -- regression reports nothing at all.
        procedure wait_idle is
        begin
            guard := 0;
            apb_read(A_STATUS_C);
            while rdbuf(0) = '1' and guard < 4000 loop
                apb_read(A_STATUS_C);
                guard := guard + 1;
            end loop;
            if guard >= 4000 then
                write(ln, string'("  FAIL: BUSY never cleared")); writeline(output, ln); e := e + 1;
            end if;
        end procedure wait_idle;

        procedure reset_all is
        begin
            wait until falling_edge(pclk);
            prst_n <= '0'; srst_n <= '0';
            for i in 1 to 6 loop wait until falling_edge(pclk); end loop;
            prst_n <= '1'; srst_n <= '1';
            for i in 1 to 4 loop wait until falling_edge(pclk); end loop;
        end procedure reset_all;

    begin
        -- ============ 1. the shadow register ============
        sref_half <= 7 ns;
        reset_all;
        slave_byte <= x"A5";
        -- Divider 2, not 0. This slave model presents each bit two reference cycles after the trailing
        -- edge, so the SCLK half period needs at least three reference cycles for the master's leading edge
        -- to have any setup. That boundary belongs to Chapter 19.1, not to the register bus.
        apb_write(A_CTRL_C,   x"00000021");
        apb_write(A_TXDATA_C, x"0000003C");
        -- Mid-transfer, software writes a completely different configuration: mode 3 and divider 3.
        for i in 1 to 8 loop wait until falling_edge(pclk); end loop;
        apb_write(A_CTRL_C,   x"00000037");
        wait_idle;
        apb_read(A_RXDATA_C); got_rx := rdbuf(7 downto 0);
        got_edges := edge_count;
        write(ln, string'("  shadow test: rxdata=") & hex2(got_rx) & string'(" (slave sent ")
                  & hex2(slave_byte) & string'(")   slave received ") & hex2(slv_rx)
                  & string'(" (master sent 3c)   shadow_ctrl=") & hex2(dbg_shadow_ctrl)
                  & string'(" (live ctrl=37)   edges=") & i2s(got_edges, 1));
        writeline(output, ln);
        if got_rx /= slave_byte then
            write(ln, string'("  FAIL: the mid-transfer CTRL write corrupted the received byte"));
            writeline(output, ln); e := e + 1;
        end if;
        if slv_rx /= x"3C" then
            write(ln, string'("  FAIL: the slave received the wrong byte"));
            writeline(output, ln); e := e + 1;
        end if;
        if dbg_shadow_ctrl /= x"21" then
            write(ln, string'("  FAIL: the shadowed configuration changed mid-transfer"));
            writeline(output, ln); e := e + 1;
        end if;
        if got_edges /= 16 then
            write(ln, string'("  FAIL: ") & i2s(got_edges, 1)
                      & string'(" SCLK edges on the wire where 16 were expected"));
            writeline(output, ln); e := e + 1;
        end if;

        -- ============ 2. write-one-to-clear ============
        reset_all;
        slave_byte <= x"5A";
        apb_write(A_CTRL_C, x"00000021");
        apb_write(A_TXDATA_C, x"00000011");
        wait_idle;
        apb_write(A_TXDATA_C, x"00000022");     -- a second start with DONE still set -> OVERRUN
        wait_idle;
        apb_read(A_STATUS_C);
        write(ln, string'("  sticky bits after two transfers with no clear: busy=") & b2s(rdbuf(0))
                  & string'(" done=") & b2s(rdbuf(1)) & string'(" overrun=") & b2s(rdbuf(2)));
        writeline(output, ln);
        if not (rdbuf(1) = '1' and rdbuf(2) = '1') then
            write(ln, string'("  FAIL: DONE and OVERRUN were not both set after an unread completion"));
            writeline(output, ln); e := e + 1;
        end if;
        apb_write(A_STATUS_C, x"00000002");     -- clear DONE only
        apb_read(A_STATUS_C);
        write(ln, string'("  after writing 0x02:                          busy=") & b2s(rdbuf(0))
                  & string'(" done=") & b2s(rdbuf(1)) & string'(" overrun=") & b2s(rdbuf(2)));
        writeline(output, ln);
        if rdbuf(1) = '1' then
            write(ln, string'("  FAIL: writing a one to DONE did not clear it"));
            writeline(output, ln); e := e + 1;
        end if;
        if rdbuf(2) = '0' then
            write(ln, string'("  FAIL: writing a one to DONE also cleared OVERRUN; a zero must leave a bit alone"));
            writeline(output, ln); e := e + 1;
        end if;
        apb_write(A_STATUS_C, x"00000004");
        apb_read(A_STATUS_C);
        if rdbuf(2) = '1' then
            write(ln, string'("  FAIL: writing a one to OVERRUN did not clear it"));
            writeline(output, ln); e := e + 1;
        end if;

        -- ============ 3. a start while busy is refused and recorded ============
        reset_all;
        slave_byte <= x"3C";
        apb_write(A_CTRL_C, x"00000021");
        apb_write(A_TXDATA_C, x"000000AA");
        apb_write(A_TXDATA_C, x"000000BB");     -- issued while the first is still running
        wait_idle;
        apb_read(A_STATUS_C);
        write(ln, string'("  a second start while busy:                    done=") & b2s(rdbuf(1))
                  & string'(" overrun=") & b2s(rdbuf(2)) & string'("  slave received ") & hex2(slv_rx));
        writeline(output, ln);
        if rdbuf(2) = '0' then
            write(ln, string'("  FAIL: a start issued while busy was not recorded as an overrun"));
            writeline(output, ln); e := e + 1;
        end if;
        if slv_rx /= x"AA" then
            write(ln, string'("  FAIL: the refused start disturbed the transfer in flight"));
            writeline(output, ln); e := e + 1;
        end if;

        -- ============ 4. rate invariance ============
        write(ln, string'(""));
        writeline(output, ln);
        write(ln, string'("  sclk_ref half  ratio to pclk  rxdata  frames  edges  status"));
        writeline(output, ln);
        for k in 0 to 3 loop
            sref_half <= RATIOS_C(k) * 1 ns;
            reset_all;
            slave_byte <= x"96";
            apb_write(A_CTRL_C, x"00000021");
            apb_write(A_TXDATA_C, x"00000069");
            wait_idle;
            apb_read(A_RXDATA_C); res_rx(k) := rdbuf(7 downto 0);
            res_frames(k) := frames_seen;
            res_edges(k)  := edge_count;
            apb_read(A_STATUS_C);
            write(ln, string'("  ") & i2s(RATIOS_C(k), 13) & string'("      unrelated      ")
                      & hex2(res_rx(k)) & string'("  ") & i2s(res_frames(k), 6)
                      & string'("  ") & i2s(res_edges(k), 5) & string'("  done=") & b2s(rdbuf(1))
                      & string'(" overrun=") & b2s(rdbuf(2)));
            writeline(output, ln);
            if res_rx(k) /= x"96" then
                write(ln, string'("  FAIL: at a half period of ") & i2s(RATIOS_C(k), 1)
                          & string'(" the received byte was wrong"));
                writeline(output, ln); e := e + 1;
            end if;
            if res_edges(k) /= 16 then
                write(ln, string'("  FAIL: at a half period of ") & i2s(RATIOS_C(k), 1)
                          & string'(" the wire carried ") & i2s(res_edges(k), 1)
                          & string'(" edges, not 16"));
                writeline(output, ln); e := e + 1;
            end if;
            if rdbuf(2) = '1' then
                write(ln, string'("  FAIL: an overrun was reported for a single transfer"));
                writeline(output, ln); e := e + 1;
            end if;
        end loop;
        if not (res_rx(0) = res_rx(1) and res_rx(1) = res_rx(2) and res_rx(2) = res_rx(3)) then
            write(ln, string'("  FAIL: the received byte depended on the clock ratio"));
            writeline(output, ln); e := e + 1;
        end if;
        if not (res_frames(0) = res_frames(1) and res_frames(1) = res_frames(2)
                and res_frames(2) = res_frames(3)) then
            write(ln, string'("  FAIL: the frame count depended on the clock ratio"));
            writeline(output, ln); e := e + 1;
        end if;

        -- ============ conclusions ============
        write(ln, string'(""));
        writeline(output, ln);
        write(ln, string'("    1. a mid-transfer CTRL write changed the live register from 21 to 37 -- a different MODE and a different DIVIDER -- and the transfer in flight was untouched: the slave received 3c, the master received 96, the shadowed configuration still read 21, and the wire carried exactly 16 edges. The shadow costs one register's worth of flops and removes a rule every future driver would otherwise have to obey. The alternative contract -- `poll BUSY before writing CTRL` -- is a rule that has to be honoured by people who will never read the datasheet section that states it"));
        writeline(output, ln);
        write(ln, string'("    2. write-one-to-clear means what it says: after both sticky bits were set, a write of 0x02 cleared DONE and LEFT OVERRUN alone. A status register that clears on any write destroys a bit the driver has not read yet, and because the destroyed bit is the record of a lost completion the symptom is a stream that is quietly short"));
        writeline(output, ln);
        write(ln, string'("    3. a second completion with the first unread set OVERRUN rather than being discarded, and a second START issued while busy was refused without disturbing the transfer in flight -- the slave still received aa. A sticky DONE without an OVERRUN companion silently loses completions, which is the failure a level DONE was replaced to avoid rather than a different one"));
        writeline(output, ln);
        write(ln, string'("    4. and the results do not depend on the clock ratio. At sclk_ref half periods of ")
                  & i2s(RATIOS_C(0), 1) & string'(", ") & i2s(RATIOS_C(1), 1) & string'(", ")
                  & i2s(RATIOS_C(2), 1) & string'(" and ") & i2s(RATIOS_C(3), 1)
                  & string'(" nanoseconds against a fixed 5 ns bus half period -- none of them an integer relationship -- every received byte was 96, every frame count was identical and every wire carried 16 edges. That is rate invariance, and it is the property a multi-bit crossing defect breaks. It is also the ONLY thing this simulation establishes about the crossing: both synchronisers are two flops deep because that is correct practice, and a single flop would behave identically here"));
        writeline(output, ln);

        -- ============ BENCH INTEGRITY ============
        if dbg_shadow_ctrl /= x"DE" then mutations := mutations + 1; end if;
        if res_edges(0) /= 99         then mutations := mutations + 1; end if;
        if mutations /= 2 then
            write(ln, string'("  FAIL: a deliberately wrong expectation did not mismatch ("
                      )) ; write(ln, i2s(mutations, 1) & string'(" of 2)"));
            writeline(output, ln); e := e + 1;
        end if;
        if frames_seen = 0 then
            write(ln, string'("  FAIL: the pin monitor never saw a frame, so nothing above was observed on the wire"));
            writeline(output, ln); e := e + 1;
        end if;
        if x_reports /= 0 then
            write(ln, string'("  FAIL: bus reads returned a metavalue"));
            writeline(output, ln); e := e + 1;
        end if;

        if e = 0 then
            write(ln, string'(""));
            writeline(output, ln);
            write(ln, string'("    and the bench proved itself: two deliberately wrong expectations mismatched, the pin monitor observed every frame independently of the design's own status bits, every bus read carried a known value, and the slave model knows nothing about the register bus"));
            writeline(output, ln);
            write(ln, string'("PASS: putting an SPI master behind a register bus makes software part of the design, and the three failures that ship are all in the seam. A CONFIGURATION SHADOW makes a mid-transfer CTRL write harmless -- the live register went from 21 to 37 while the transfer in flight kept its mode, its divider and its 16 edges -- and it replaces a rule every future driver would have to obey with one register's worth of flops. STATUS SEMANTICS need two bits and a discipline: BUSY is a level, DONE is sticky and write-one-to-clear because a level DONE can be missed between transfers, and OVERRUN exists because a sticky DONE with no companion silently loses the second completion. Writing a zero left a bit alone, a second completion set OVERRUN, and a second start while busy was refused without disturbing the transfer on the wire. And the CROSSING is multi-bit, so it is a request/acknowledge TOGGLE handshake with the data held stable across it rather than a synchroniser per bit: at four unrelated clock ratios every received byte, frame count and edge count was identical. Rate invariance is the one property a simulation can genuinely establish about a crossing -- it says nothing whatever about synchroniser depth, because metastability is not representable and a single flop would behave identically here"));
            writeline(output, ln);
        else
            write(ln, string'("FAIL: ") & i2s(e, 1) & string'(" error(s)"));
            writeline(output, ln);
        end if;

        run <= false;
        wait;
    end process stim;

end architecture tb;

7. Where UVM RAL Genuinely Belongs

This is the first chapter in the module with a register map, and a register abstraction layer is the right tool here rather than a decoration.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   a register model mirrors CTRL, TXDATA, RXDATA, STATUS and IE
   frontdoor access drives real APB transactions
   the model PREDICTS the mirror after each access

8. The Assertions Worth Writing

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   BUSY is high if and only if the request and acknowledge toggles disagree
   no frame starts while BUSY is high
   the shadowed configuration does not change between the start and the end
     of a frame
   every frame carries exactly 16 SCLK edges

9. FPGA And ASIC Implementation

On an FPGA, the two clocks need to be declared asynchronous to each other — set_clock_groups -asynchronous or the vendor equivalent — or static timing analysis will try to close a path between them and report a violation on a crossing that is deliberately unconstrained. The synchroniser flops should carry whatever the vendor's false path or max delay convention is, so the tool does not optimise them into one.

The register file is small enough to live in logic. What is worth thinking about is prdata: a wide read mux across five registers is a combinational cone that can become the bus's critical path. Registering it — as this design does — costs one cycle of read latency that APB is happy to absorb.

On an ASIC, the same crossing needs a set_clock_groups and a CDC tool run, and the synchroniser cells usually come from a library with a documented mean-time-between-failure rather than from inferred flops. That MTBF number is the thing simulation cannot give you and the reason the tool exists.

And the register map is a specification deliverable. Field widths, reset values, access types and volatility go in a machine-readable description that generates both the RTL and the register model — because a map maintained in two places diverges, and the divergence appears as a driver that works against the documentation and not against the silicon.

10. Failure Signature — "The Driver Loses Every Fourth Byte"

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   Symptom     a streaming driver loses roughly one byte in four at high rates
               and none at low rates
   Checked     the SPI pins on a scope -- every frame present and correct
   Checked     the driver's own accounting -- one write per byte, as expected
   Concluded   the hardware drops transfers under load
   Actual      the driver polled DONE, saw it set, read RXDATA, and cleared
               STATUS with a write of 0xFF -- which also cleared an OVERRUN
               that had been set by a completion arriving during the read
   Found       by an engineer who noticed OVERRUN was never once observed as
               set, in a system that was demonstrably overrunning

The scope was right, the driver's accounting was right, and the conclusion followed. What made it invisible is that the evidence was being destroyed by the act of reading: a blanket 0xFF write to a status register clears bits the driver has not looked at, and the bit it destroyed was the only record that anything had gone wrong.

A counter that is never observed to be non-zero in a system that is demonstrably failing is not evidence of correctness — it is evidence that something is clearing it.

11. Common Misconceptions

MisconceptionWhat is actually true
Software should just poll BUSY before writing CTRLThat is a rule every future driver must obey; a shadow register removes it
A sticky DONE fixes missed completionsIt trades them for lost second completions, which is why OVERRUN exists
Writing 0xFF to a status register is a safe way to clear itIt destroys bits the driver has not read, including the record of a loss
A multi-bit value can be synchronised bit by bitEach bit resolves independently; a value that was never valid can appear
A busy flag is simpler than deriving BUSY from the togglesA separate flag can disagree with the engine; a derived one cannot
A divider that produces the right number of edges is workingIt may be delaying only the first edge and running free afterwards
A passing rate sweep proves the crossing is safeIt proves the handshake's shape; it says nothing about synchroniser depth

12. Reason It Through

13. Understanding Check

14. Summary

Putting an SPI master behind a register bus makes software part of the design, and the three failures that ship are all in the seam.

A configuration shadow makes a mid-transfer CTRL write harmless: the live register went from 0x21 to 0x37 while the transfer in flight kept its mode, its divider and its 16 edges. It replaces a rule every future driver would have to obey with one register's worth of flops — prefer a hardware invariant to a software obligation, because the invariant is enforced once and the obligation is re-tested by every integration.

Status semantics need two bits and a discipline. BUSY is a level, derived from the crossing's own toggles rather than maintained separately, so it cannot disagree with the engine. DONE is sticky and write-one-to-clear, because a level DONE exists only between transfers and is missed by any driver doing work. And OVERRUN exists because a sticky DONE has nowhere to record a second completion — it is the other half of the design, not a diagnostic. Writing a zero left a bit alone, a second completion set OVERRUN, and a second start while busy was refused without disturbing the transfer on the wire.

The crossing is multi-bit, so it is a request/acknowledge toggle handshake with the data held stable across it rather than a synchroniser per bit: at four unrelated clock ratios every received byte, frame count and edge count was identical. Rate invariance is the one property a simulation can genuinely establish about a crossing; it says nothing whatever about synchroniser depth.

Two defects found along the way are worth carrying. A divider that never reloads delays the first edge and nothing else, while the edge count and the data both stay correct — it surfaced only when a slave ran out of setup margin, looking like a crossing fault. And the bench's slave, clocked on the bus clock, undersampled a faster SCLK and produced two failures that looked exactly like a broken design: a monitor must be at least as fast as the thing it monitors, and a rate sweep is the experiment that finds out whether it is.

15. What Comes Next

One master, one slave, one configuration. Chapter 19.4 adds devices: several slaves on one bus, each with its own select, its own mode and its own maximum clock rate — and the configuration switch between devices becomes the bug factory, because changing CPOL changes SCLK's idle level while another device is still watching the bus.

Continue learning