Skip to content
VLSI Mentor

SPI · Module 15

Why an SPI Slave Is a Clock-Domain Problem

An externally generated SCLK turns a shift register into a CDC question, and there are only four ways to carry an event across: a level, a synchronised level, a toggle, and a handshake. Each is limited by something different, no scheme creates bandwidth, and the difference that matters most is the one no simulation can show.

Module 13 built a master and never had this problem: the master generates SCLK, so there is one clock and one domain.

Module 14 built a slave and assumed its way past the problem. Chapter 14.1 recovered SCLK's edges on the system clock and stated a ratio precondition — which is one answer to a question that was never properly asked.

This module asks it.

An event happens in the SCLK domain. It has to reach the system domain. What are the options, and what does each one cost?

There are four, they are not equivalent, and the differences between them are the whole subject of Module 15.

1. The Four Schemes

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   1 LEVEL      one destination flop, edge-detected.
                Limited by PULSE WIDTH: a source pulse narrower than a destination
                clock period may fall between two sampling instants and never be seen.

   2 SYNC       two destination flops, edge-detected.
                Limited by PULSE WIDTH in exactly the same way -- the second flop
                buys metastability settling, not bandwidth.

   3 TOGGLE     the source TOGGLES a bit per event; the destination synchronises it
                and detects BOTH edges. A toggle is a level that persists until the
                next event, so pulse width stops mattering.
                Limited instead by EVENT RATE: two toggles inside one destination
                sampling period leave the level unchanged and both events vanish.

   4 HANDSHAKE  the source toggles and then WAITS for an acknowledgement.
                Limited by nothing -- and the price is that the source stalls.

The conclusion is worth stating before the evidence, because the evidence then reads as confirmation rather than as a surprise:

No crossing scheme creates bandwidth.

A toggle moves the limit from pulse width to event rate. A handshake removes the limit by removing the source's freedom to keep going. Nothing makes a destination clock able to observe more events than it has cycles.

2. The Block Diagram

Four crossing schemes: a one-cycle pulse into one flop and into two flops, a toggle into two flops with both-edge detection, and the same toggle with an acknowledgement returning through its own synchronisersrc_eventpulse registertoggle registertoggle + busy1 flopSYNC_N flopsSYNC_N flopsSYNC_N flopsrising-edge detectboth-edge detectack, mirroreddelivered events12
Figure 1 — the four schemes on one source event. Schemes 1 and 2 differ only in the number of destination flops, which is why their counters can never disagree in simulation. Scheme 4's dashed path is the acknowledgement crossing back — a handshake is two crossings rather than one, and that return trip is half of what it costs.

3. The Measurement

The same twenty events, one every two source cycles, at six source periods against a fixed 10 ns destination clock:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   src_period  generated  scheme1  scheme2  scheme3  scheme4
        100 ns         20       20       20       20       10
         50 ns         20       20       20       20       10
         20 ns         20       20       20       20       10
         10 ns         20       20       20       20       10
          6 ns         20       12       12       20        7
          4 ns         20        4        4       20        6

Read it column by column, because each column fails for its own reason.

Twelve source-clock cycles across six rows. A source clock runs throughout. A pulse row is high for one cycle at column 3 and again at column 7. A toggle row rises at column 3 and falls at column 7. A destination sampling row marks columns 0, 3, 6 and 9. The one-flop sampled level rises after the first pulse and never sees the second; the sampled toggle follows both transitions.caught: a sample lands on itcaught: a sample lands onitmissed: no sample lands on itmissed: no sample lands onitthe toggle is still seenthe toggle is still seensrc_clkpulse (pin)toggle (pin)dst sampleslvl_q (sch 1,2)tog_q (sch 3)t0t1t2t3t4t5t6t7t8t9t10t11
Figure 2 — twelve SOURCE cycles, with the destination sampling every third one. The first pulse happens to coincide with a sampling instant and is delivered; the second falls between two and is gone. The toggle survives both, because a toggle is a level that persists until the next event rather than a pulse that has to be caught.

Schemes 1 and 2 degrade together — 20, 20, 20, 20, 12, 4. Once the source period is shorter than the destination's, a one-source-cycle pulse can fall between two sampling instants, and it does so more often as the source speeds up.

Scheme 3 never loses one — 20 at every ratio. A toggle has no width to lose.

Scheme 4 loses events at every ratio, including the slowest, and that is not a defect. The source in this table ignores the busy signal, which is what a slave does — and a handshake's round trip is SYNC_N destination cycles out and SYNC_N source cycles back, so a source producing events every two source cycles outruns it however slow it is.

4. The Toggle's Limit, And The Handshake's Price

Two further measurements finish the picture, and both are about what happens when you push each scheme past its own boundary.

Push the event rate up and the toggle fails too. At a 4 ns source period with an event every source cycle, the toggle delivered 12 of 20 — because two toggles inside one destination sampling period leave the level unchanged and both events vanish. The toggle moved the limit from pulse width to event rate; it did not remove it.

Give the handshake a source that waits and it loses nothing. The same 4 ns source period, honouring src_busy, delivered 20 of 20. Exact, at the fastest ratio in the table.

Give it a source that cannot wait and the guarantee evaporates. The same handshake against a source ignoring src_busy delivered 4 of 20 — no better than the pulse.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   scheme      limited by        an SPI slave can honour it?
   ---------   ---------------   ----------------------------------------
   1, 2        pulse width       yes, by making SCLK slow enough
   3           event rate        yes, and the limit is far away
   4           nothing           NO -- SCLK does not stop for the slave

That last row is why Module 15 does not end here. A handshake is the only scheme with no bandwidth limit, and its precondition is a freedom an SPI slave structurally lacks. Chapter 15.6 has to build a FIFO instead, and the reason is in this table.

5. Building the Four Schemes — Three HDLs

The circuit

One module, four clearly separated crossings, sharing one source event so that any difference between their counters comes from the crossing rather than from the stimulus. The event counts live in their own domains — a count of what was generated cannot itself be delivered across the boundary without raising the question the block exists to answer.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_cdc_probe.sv — four crossings on one event, so that any difference between their counters comes from the crossing
// spi_cdc_probe.sv
//
// Chapter 15.1 -- why an SPI slave is a clock-domain problem, made measurable.
//
// Modules 13 and 14 both assumed their way past this. The master owned SCLK, so
// there was no second domain. The slave (Chapter 14.1) recovered SCLK's edges on
// the system clock and stated a ratio precondition -- which is one answer to the
// question, chosen before the question was properly asked.
//
// This block asks it. An event happens in the SCLK domain. It has to be delivered
// to the system domain. There are four ways to do that, they are not equivalent,
// and the differences between them are the whole subject of Module 15.
//
// THE FOUR SCHEMES, AND WHAT EACH ONE IS LIMITED BY.
//
//   1 LEVEL      one destination flop, edge-detected. Limited by PULSE WIDTH: a
//                source pulse narrower than a destination clock period may fall
//                between two sampling instants and never be seen.
//
//   2 SYNC       two destination flops, edge-detected. Limited by PULSE WIDTH in
//                exactly the same way -- the second flop buys metastability
//                settling, not bandwidth.
//
//   3 TOGGLE     the source TOGGLES a bit per event; the destination
//                synchronises it and detects BOTH edges. A toggle is a level
//                that persists until the next event, so pulse width stops
//                mattering. Limited instead by EVENT RATE: two toggles inside
//                one destination sampling period leave the level unchanged and
//                both events vanish.
//
//   4 HANDSHAKE  the source toggles and then WAITS for an acknowledgement before
//                producing another event. Limited by nothing -- and the price is
//                that the source stalls, which is a price a slave often cannot
//                pay, because SCLK does not stop for it.
//
// WHAT THIS MEANS, STATED BEFORE THE CODE SO THE CODE READS AS EVIDENCE:
//
//     no crossing scheme creates bandwidth.
//
// A toggle moves the limit from pulse width to event rate. A handshake removes
// the limit by removing the source's freedom to keep going. Nothing makes a
// destination clock able to observe more events than it has cycles.
//
// AND THE PART THAT MATTERS MOST FOR HOW MODULE 15 IS VERIFIED.
//
// Schemes 1 and 2 behave IDENTICALLY in simulation, apart from one cycle of
// latency. Every event scheme 1 delivers, scheme 2 delivers. Every event scheme 1
// loses, scheme 2 loses. A regression that compares their counters finds no
// difference, at any ratio, with any stimulus, forever.
//
// The difference between them is that scheme 1's flop can be sampled while its
// input is changing and can then drive a metastable value into logic that is
// about to make a decision on it -- and a simulator's flop always captures a
// defined value. So the single most important rule in this module is the one
// rule no testbench in this module can check.
//
// That is why the counters here count what IS observable: events generated
// against events delivered, per scheme. Simulation settles the bandwidth
// question completely and says nothing at all about the settling question, and
// keeping those two apart is the difference between a CDC review that works and
// one that consists of running the regression again.

module spi_cdc_probe #(
    parameter int SYNC_N = 2,    // destination synchroniser depth for schemes 2-4
    parameter int CNT_W  = 16    // width of the event counters
) (
    // --- the source domain: SCLK, or anything else somebody else owns --------
    input  wire              sclk,
    input  wire              src_rst_n,
    input  wire              src_event,     // one sclk cycle per event
    output wire              src_busy,      // scheme 4 only: an event is in flight

    // --- the destination domain: the system clock ---------------------------
    input  wire              clk,
    input  wire              dst_rst_n,

    output reg               dst_level_stb,   // scheme 1
    output reg               dst_sync_stb,    // scheme 2
    output reg               dst_toggle_stb,  // scheme 3
    output reg               dst_hs_stb,      // scheme 4

    // --- what was generated, and what arrived -------------------------------
    // The source count is in the SOURCE domain deliberately: a count of what was
    // generated cannot itself be delivered across the boundary without raising
    // the same question the block exists to answer.
    output reg  [CNT_W-1:0]  src_events,
    output reg  [CNT_W-1:0]  dst_level_n,
    output reg  [CNT_W-1:0]  dst_sync_n,
    output reg  [CNT_W-1:0]  dst_toggle_n,
    output reg  [CNT_W-1:0]  dst_hs_n
);

    // The acknowledgement, declared here because it is written by the DESTINATION
    // and read by the SOURCE. A signal crossing back is easy to miss when reading
    // a handshake, and it is half of what a handshake costs.
    reg dst_hs_ack;

    // =====================================================================
    // SOURCE DOMAIN
    // =====================================================================

    // Scheme 1 and 2 share one source signal: a pulse one source cycle wide.
    // Sharing it is the point -- the two schemes differ only in the destination,
    // so any difference in their counters would have to come from the flop count.
    reg pulse_src;

    // Scheme 3: a toggle. Its whole virtue is that it has no width; it is a
    // level that stays wherever the last event left it.
    reg tog_src;

    // Scheme 4: the same toggle, plus the source refusing to move again until
    // the destination has answered. `hs_ack_q` is the acknowledgement having
    // come BACK across the boundary, which needs its own synchroniser -- a
    // handshake is two crossings, not one, and that is half of what it costs.
    reg hs_tog_src;
    reg [SYNC_N-1:0] hs_ack_sr;
    wire hs_ack_q = hs_ack_sr[SYNC_N-1];
    reg  hs_ack_seen;          // the ack for the event in flight has returned

    assign src_busy = hs_tog_src ^ hs_ack_seen;

    always_ff @(posedge sclk or negedge src_rst_n) begin
        if (!src_rst_n) begin
            pulse_src    <= 1'b0;
            tog_src      <= 1'b0;
            hs_tog_src   <= 1'b0;
            hs_ack_sr    <= {SYNC_N{1'b0}};
            hs_ack_seen  <= 1'b0;
            src_events   <= {CNT_W{1'b0}};
        end else begin
            hs_ack_sr <= {hs_ack_sr[SYNC_N-2:0], dst_hs_ack};

            // The pulse is one source cycle wide and then gone. That is what
            // makes it a pulse, and it is what makes it losable.
            pulse_src <= src_event;

            if (src_event) begin
                src_events <= src_events + 1'b1;
                tog_src    <= ~tog_src;

                // Scheme 4 only accepts an event when the previous one has been
                // acknowledged. A source that ignores `src_busy` and toggles
                // anyway has built scheme 3 with extra logic.
                if (!src_busy)
                    hs_tog_src <= ~hs_tog_src;
            end

            // The ack has come back; the handshake is free for the next event.
            if (hs_ack_q != hs_ack_seen)
                hs_ack_seen <= hs_ack_q;
        end
    end

    // =====================================================================
    // DESTINATION DOMAIN
    // =====================================================================

    // Scheme 1: ONE flop, then an edge detect. The second flop of the edge
    // detector is not a synchroniser stage -- it holds the previous value so the
    // rise can be found, and it is fed by a flop whose input crossed the
    // boundary. This is the shape a CDC review rejects, and it is here so that
    // the counters can show it is not the bandwidth that the review is about.
    reg lvl_q, lvl_q_d;

    // Scheme 2: SYNC_N flops, then the same edge detect.
    reg [SYNC_N-1:0] sync_sr;
    reg              sync_q_d;

    // Scheme 3 and 4: the toggles, synchronised, then BOTH edges detected --
    // because a toggle carries one event per transition in either direction.
    reg [SYNC_N-1:0] tog_sr;
    reg              tog_q_d;
    reg [SYNC_N-1:0] hs_tog_sr;
    reg              hs_tog_q_d;

    // The acknowledgement the source is waiting for: the destination mirrors the
    // toggle it observed. Mirroring rather than pulsing matters -- an ack pulse
    // crossing back would be scheme 1 in the reverse direction, with the same
    // pulse-width problem, in the path that exists to avoid it. Declared above,
    // with the source-domain signals, because it crosses in the other direction.

    always_ff @(posedge clk or negedge dst_rst_n) begin
        if (!dst_rst_n) begin
            lvl_q          <= 1'b0;
            lvl_q_d        <= 1'b0;
            sync_sr        <= {SYNC_N{1'b0}};
            sync_q_d       <= 1'b0;
            tog_sr         <= {SYNC_N{1'b0}};
            tog_q_d        <= 1'b0;
            hs_tog_sr      <= {SYNC_N{1'b0}};
            hs_tog_q_d     <= 1'b0;
            dst_hs_ack     <= 1'b0;
            dst_level_stb  <= 1'b0;
            dst_sync_stb   <= 1'b0;
            dst_toggle_stb <= 1'b0;
            dst_hs_stb     <= 1'b0;
            dst_level_n    <= {CNT_W{1'b0}};
            dst_sync_n     <= {CNT_W{1'b0}};
            dst_toggle_n   <= {CNT_W{1'b0}};
            dst_hs_n       <= {CNT_W{1'b0}};
        end else begin
            lvl_q     <= pulse_src;
            lvl_q_d   <= lvl_q;
            sync_sr   <= {sync_sr[SYNC_N-2:0], pulse_src};
            sync_q_d  <= sync_sr[SYNC_N-1];
            tog_sr    <= {tog_sr[SYNC_N-2:0], tog_src};
            tog_q_d   <= tog_sr[SYNC_N-1];
            hs_tog_sr <= {hs_tog_sr[SYNC_N-2:0], hs_tog_src};
            hs_tog_q_d<= hs_tog_sr[SYNC_N-1];

            // --- the four deliveries -------------------------------------
            dst_level_stb  <= lvl_q & ~lvl_q_d;
            dst_sync_stb   <= sync_sr[SYNC_N-1] & ~sync_q_d;
            dst_toggle_stb <= tog_sr[SYNC_N-1] ^ tog_q_d;
            dst_hs_stb     <= hs_tog_sr[SYNC_N-1] ^ hs_tog_q_d;

            // The ack mirrors the observed toggle, one cycle behind the strobe
            // so it cannot be mistaken for the event itself.
            if (hs_tog_sr[SYNC_N-1] != hs_tog_q_d)
                dst_hs_ack <= hs_tog_sr[SYNC_N-1];

            if (dst_level_stb)  dst_level_n  <= dst_level_n  + 1'b1;
            if (dst_sync_stb)   dst_sync_n   <= dst_sync_n   + 1'b1;
            if (dst_toggle_stb) dst_toggle_n <= dst_toggle_n + 1'b1;
            if (dst_hs_stb)     dst_hs_n     <= dst_hs_n     + 1'b1;
        end
    end

`ifdef SPI_CHECKS
    // Every strobe is one cycle wide. A delivery that lasts two cycles would be
    // counted twice, and a counter that disagreed with reality would make this
    // whole block useless as evidence.
    //
    // Written with explicit delayed copies rather than `$past`, because Icarus
    // Verilog does not implement it -- and a check that only exists in a
    // simulator nobody runs is not a check.
    // The check applies to the PULSE schemes only, and the reason is worth a note
    // because the obvious generalisation is wrong. Schemes 1 and 2 detect a RISING
    // edge, so two deliveries need a fall between them and cannot be adjacent.
    // Schemes 3 and 4 detect BOTH edges, so two toggles on consecutive
    // destination cycles produce two adjacent strobes -- which is two events
    // correctly delivered, not one event delivered twice. A width check there
    // would fail on the design working.
    reg chk_lvl_d, chk_syn_d;
    always_ff @(posedge clk) begin
        chk_lvl_d <= dst_level_stb;
        chk_syn_d <= dst_sync_stb;
        if (dst_rst_n) begin
            if (dst_level_stb && chk_lvl_d)
                $fatal(1, "level strobe wider than one cycle");
            if (dst_sync_stb && chk_syn_d)
                $fatal(1, "sync strobe wider than one cycle");
        end
    end
`endif

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_cdc_probe.v — the same design in Verilog-2001
// spi_cdc_probe.v
//
// Chapter 15.1 -- why an SPI slave is a clock-domain problem, made measurable.
//
// Modules 13 and 14 both assumed their way past this. The master owned SCLK, so
// there was no second domain. The slave (Chapter 14.1) recovered SCLK's edges on
// the system clock and stated a ratio precondition -- which is one answer to the
// question, chosen before the question was properly asked.
//
// This block asks it. An event happens in the SCLK domain. It has to be delivered
// to the system domain. There are four ways to do that, they are not equivalent,
// and the differences between them are the whole subject of Module 15.
//
// THE FOUR SCHEMES, AND WHAT EACH ONE IS LIMITED BY.
//
//   1 LEVEL      one destination flop, edge-detected. Limited by PULSE WIDTH: a
//                source pulse narrower than a destination clock period may fall
//                between two sampling instants and never be seen.
//
//   2 SYNC       two destination flops, edge-detected. Limited by PULSE WIDTH in
//                exactly the same way -- the second flop buys metastability
//                settling, not bandwidth.
//
//   3 TOGGLE     the source TOGGLES a bit per event; the destination
//                synchronises it and detects BOTH edges. A toggle is a level
//                that persists until the next event, so pulse width stops
//                mattering. Limited instead by EVENT RATE: two toggles inside
//                one destination sampling period leave the level unchanged and
//                both events vanish.
//
//   4 HANDSHAKE  the source toggles and then WAITS for an acknowledgement before
//                producing another event. Limited by nothing -- and the price is
//                that the source stalls, which is a price a slave often cannot
//                pay, because SCLK does not stop for it.
//
// WHAT THIS MEANS, STATED BEFORE THE CODE SO THE CODE READS AS EVIDENCE:
//
//     no crossing scheme creates bandwidth.
//
// A toggle moves the limit from pulse width to event rate. A handshake removes
// the limit by removing the source's freedom to keep going. Nothing makes a
// destination clock able to observe more events than it has cycles.
//
// AND THE PART THAT MATTERS MOST FOR HOW MODULE 15 IS VERIFIED.
//
// Schemes 1 and 2 behave IDENTICALLY in simulation, apart from one cycle of
// latency. Every event scheme 1 delivers, scheme 2 delivers. Every event scheme 1
// loses, scheme 2 loses. A regression that compares their counters finds no
// difference, at any ratio, with any stimulus, forever.
//
// The difference between them is that scheme 1's flop can be sampled while its
// input is changing and can then drive a metastable value into logic that is
// about to make a decision on it -- and a simulator's flop always captures a
// defined value. So the single most important rule in this module is the one
// rule no testbench in this module can check.
//
// That is why the counters here count what IS observable: events generated
// against events delivered, per scheme. Simulation settles the bandwidth
// question completely and says nothing at all about the settling question, and
// keeping those two apart is the difference between a CDC review that works and
// one that consists of running the regression again.

module spi_cdc_probe #(
    parameter SYNC_N = 2,    // destination synchroniser depth for schemes 2-4
    parameter CNT_W  = 16    // width of the event counters
) (
    // --- the source domain: SCLK, or anything else somebody else owns --------
    input  wire              sclk,
    input  wire              src_rst_n,
    input  wire              src_event,     // one sclk cycle per event
    output wire              src_busy,      // scheme 4 only: an event is in flight

    // --- the destination domain: the system clock ---------------------------
    input  wire              clk,
    input  wire              dst_rst_n,

    output reg               dst_level_stb,   // scheme 1
    output reg               dst_sync_stb,    // scheme 2
    output reg               dst_toggle_stb,  // scheme 3
    output reg               dst_hs_stb,      // scheme 4

    // --- what was generated, and what arrived -------------------------------
    // The source count is in the SOURCE domain deliberately: a count of what was
    // generated cannot itself be delivered across the boundary without raising
    // the same question the block exists to answer.
    output reg  [CNT_W-1:0]  src_events,
    output reg  [CNT_W-1:0]  dst_level_n,
    output reg  [CNT_W-1:0]  dst_sync_n,
    output reg  [CNT_W-1:0]  dst_toggle_n,
    output reg  [CNT_W-1:0]  dst_hs_n
);

    // The acknowledgement, declared here because it is written by the DESTINATION
    // and read by the SOURCE. A signal crossing back is easy to miss when reading
    // a handshake, and it is half of what a handshake costs.
    reg dst_hs_ack;

    // =====================================================================
    // SOURCE DOMAIN
    // =====================================================================

    // Scheme 1 and 2 share one source signal: a pulse one source cycle wide.
    // Sharing it is the point -- the two schemes differ only in the destination,
    // so any difference in their counters would have to come from the flop count.
    reg pulse_src;

    // Scheme 3: a toggle. Its whole virtue is that it has no width; it is a
    // level that stays wherever the last event left it.
    reg tog_src;

    // Scheme 4: the same toggle, plus the source refusing to move again until
    // the destination has answered. `hs_ack_q` is the acknowledgement having
    // come BACK across the boundary, which needs its own synchroniser -- a
    // handshake is two crossings, not one, and that is half of what it costs.
    reg hs_tog_src;
    reg [SYNC_N-1:0] hs_ack_sr;
    wire hs_ack_q = hs_ack_sr[SYNC_N-1];
    reg  hs_ack_seen;          // the ack for the event in flight has returned

    assign src_busy = hs_tog_src ^ hs_ack_seen;

    always @(posedge sclk or negedge src_rst_n) begin
        if (!src_rst_n) begin
            pulse_src    <= 1'b0;
            tog_src      <= 1'b0;
            hs_tog_src   <= 1'b0;
            hs_ack_sr    <= {SYNC_N{1'b0}};
            hs_ack_seen  <= 1'b0;
            src_events   <= {CNT_W{1'b0}};
        end else begin
            hs_ack_sr <= {hs_ack_sr[SYNC_N-2:0], dst_hs_ack};

            // The pulse is one source cycle wide and then gone. That is what
            // makes it a pulse, and it is what makes it losable.
            pulse_src <= src_event;

            if (src_event) begin
                src_events <= src_events + 1'b1;
                tog_src    <= ~tog_src;

                // Scheme 4 only accepts an event when the previous one has been
                // acknowledged. A source that ignores `src_busy` and toggles
                // anyway has built scheme 3 with extra logic.
                if (!src_busy)
                    hs_tog_src <= ~hs_tog_src;
            end

            // The ack has come back; the handshake is free for the next event.
            if (hs_ack_q != hs_ack_seen)
                hs_ack_seen <= hs_ack_q;
        end
    end

    // =====================================================================
    // DESTINATION DOMAIN
    // =====================================================================

    // Scheme 1: ONE flop, then an edge detect. The second flop of the edge
    // detector is not a synchroniser stage -- it holds the previous value so the
    // rise can be found, and it is fed by a flop whose input crossed the
    // boundary. This is the shape a CDC review rejects, and it is here so that
    // the counters can show it is not the bandwidth that the review is about.
    reg lvl_q, lvl_q_d;

    // Scheme 2: SYNC_N flops, then the same edge detect.
    reg [SYNC_N-1:0] sync_sr;
    reg              sync_q_d;

    // Scheme 3 and 4: the toggles, synchronised, then BOTH edges detected --
    // because a toggle carries one event per transition in either direction.
    reg [SYNC_N-1:0] tog_sr;
    reg              tog_q_d;
    reg [SYNC_N-1:0] hs_tog_sr;
    reg              hs_tog_q_d;

    // The acknowledgement the source is waiting for: the destination mirrors the
    // toggle it observed. Mirroring rather than pulsing matters -- an ack pulse
    // crossing back would be scheme 1 in the reverse direction, with the same
    // pulse-width problem, in the path that exists to avoid it. Declared above,
    // with the source-domain signals, because it crosses in the other direction.

    always @(posedge clk or negedge dst_rst_n) begin
        if (!dst_rst_n) begin
            lvl_q          <= 1'b0;
            lvl_q_d        <= 1'b0;
            sync_sr        <= {SYNC_N{1'b0}};
            sync_q_d       <= 1'b0;
            tog_sr         <= {SYNC_N{1'b0}};
            tog_q_d        <= 1'b0;
            hs_tog_sr      <= {SYNC_N{1'b0}};
            hs_tog_q_d     <= 1'b0;
            dst_hs_ack     <= 1'b0;
            dst_level_stb  <= 1'b0;
            dst_sync_stb   <= 1'b0;
            dst_toggle_stb <= 1'b0;
            dst_hs_stb     <= 1'b0;
            dst_level_n    <= {CNT_W{1'b0}};
            dst_sync_n     <= {CNT_W{1'b0}};
            dst_toggle_n   <= {CNT_W{1'b0}};
            dst_hs_n       <= {CNT_W{1'b0}};
        end else begin
            lvl_q     <= pulse_src;
            lvl_q_d   <= lvl_q;
            sync_sr   <= {sync_sr[SYNC_N-2:0], pulse_src};
            sync_q_d  <= sync_sr[SYNC_N-1];
            tog_sr    <= {tog_sr[SYNC_N-2:0], tog_src};
            tog_q_d   <= tog_sr[SYNC_N-1];
            hs_tog_sr <= {hs_tog_sr[SYNC_N-2:0], hs_tog_src};
            hs_tog_q_d<= hs_tog_sr[SYNC_N-1];

            // --- the four deliveries -------------------------------------
            dst_level_stb  <= lvl_q & ~lvl_q_d;
            dst_sync_stb   <= sync_sr[SYNC_N-1] & ~sync_q_d;
            dst_toggle_stb <= tog_sr[SYNC_N-1] ^ tog_q_d;
            dst_hs_stb     <= hs_tog_sr[SYNC_N-1] ^ hs_tog_q_d;

            // The ack mirrors the observed toggle, one cycle behind the strobe
            // so it cannot be mistaken for the event itself.
            if (hs_tog_sr[SYNC_N-1] != hs_tog_q_d)
                dst_hs_ack <= hs_tog_sr[SYNC_N-1];

            if (dst_level_stb)  dst_level_n  <= dst_level_n  + 1'b1;
            if (dst_sync_stb)   dst_sync_n   <= dst_sync_n   + 1'b1;
            if (dst_toggle_stb) dst_toggle_n <= dst_toggle_n + 1'b1;
            if (dst_hs_stb)     dst_hs_n     <= dst_hs_n     + 1'b1;
        end
    end

`ifdef SPI_CHECKS
    // Every strobe is one cycle wide. A delivery that lasts two cycles would be
    // counted twice, and a counter that disagreed with reality would make this
    // whole block useless as evidence.
    //
    // Written with explicit delayed copies rather than `$past`, because Icarus
    // Verilog does not implement it -- and a check that only exists in a
    // simulator nobody runs is not a check.
    // The check applies to the PULSE schemes only, and the reason is worth a note
    // because the obvious generalisation is wrong. Schemes 1 and 2 detect a RISING
    // edge, so two deliveries need a fall between them and cannot be adjacent.
    // Schemes 3 and 4 detect BOTH edges, so two toggles on consecutive
    // destination cycles produce two adjacent strobes -- which is two events
    // correctly delivered, not one event delivered twice. A width check there
    // would fail on the design working.
    reg chk_lvl_d, chk_syn_d;
    always @(posedge clk) begin
        chk_lvl_d <= dst_level_stb;
        chk_syn_d <= dst_sync_stb;
        if (dst_rst_n) begin
            if (dst_level_stb && chk_lvl_d)
                $fatal(1, "level strobe wider than one cycle");
            if (dst_sync_stb && chk_syn_d)
                $fatal(1, "sync strobe wider than one cycle");
        end
    end
`endif

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_cdc_probe.vhd — the same design in VHDL
-- spi_cdc_probe.vhd
--
-- Chapter 15.1 -- why an SPI slave is a clock-domain problem, made measurable.
--
-- Modules 13 and 14 both assumed their way past this. The master owned SCLK, so
-- there was no second domain. The slave (Chapter 14.1) recovered SCLK's edges on
-- the system clock and stated a ratio precondition -- which is one answer to the
-- question, chosen before the question was properly asked.
--
-- This block asks it. An event happens in the SCLK domain. It has to be delivered
-- to the system domain. There are four ways to do that, they are not equivalent,
-- and the differences between them are the whole subject of Module 15.
--
-- THE FOUR SCHEMES, AND WHAT EACH ONE IS LIMITED BY.
--
--   1 LEVEL      one destination flop, edge-detected. Limited by PULSE WIDTH: a
--                source pulse narrower than a destination clock period may fall
--                between two sampling instants and never be seen.
--
--   2 SYNC       two destination flops, edge-detected. Limited by PULSE WIDTH in
--                exactly the same way -- the second flop buys metastability
--                settling, not bandwidth.
--
--   3 TOGGLE     the source TOGGLES a bit per event; the destination
--                synchronises it and detects BOTH edges. A toggle is a level
--                that persists until the next event, so pulse width stops
--                mattering. Limited instead by EVENT RATE: two toggles inside
--                one destination sampling period leave the level unchanged and
--                both events vanish.
--
--   4 HANDSHAKE  the source toggles and then WAITS for an acknowledgement before
--                producing another event. Limited by nothing -- and the price is
--                that the source stalls, which is a price a slave often cannot
--                pay, because SCLK does not stop for it.
--
-- WHAT THIS MEANS, STATED BEFORE THE CODE SO THE CODE READS AS EVIDENCE:
--
--     no crossing scheme creates bandwidth.
--
-- A toggle moves the limit from pulse width to event rate. A handshake removes
-- the limit by removing the source's freedom to keep going. Nothing makes a
-- destination clock able to observe more events than it has cycles.
--
-- AND THE PART THAT MATTERS MOST FOR HOW MODULE 15 IS VERIFIED.
--
-- Schemes 1 and 2 behave IDENTICALLY in simulation, apart from one cycle of
-- latency. Every event scheme 1 delivers, scheme 2 delivers. Every event scheme 1
-- loses, scheme 2 loses. A regression that compares their counters finds no
-- difference, at any ratio, with any stimulus, forever.
--
-- The difference between them is that scheme 1's flop can be sampled while its
-- input is changing and can then drive a metastable value into logic that is
-- about to make a decision on it -- and a simulator's flop always captures a
-- defined value. So the single most important rule in this module is the one
-- rule no testbench in this module can check.
--
-- That is why the counters here count what IS observable: events generated
-- against events delivered, per scheme. Simulation settles the bandwidth
-- question completely and says nothing at all about the settling question, and
-- keeping those two apart is the difference between a CDC review that works and
-- one that consists of running the regression again.

library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

entity spi_cdc_probe is
    generic (
        SYNC_N : positive := 2;   -- destination synchroniser depth for schemes 2-4
        CNT_W  : positive := 16   -- width of the event counters
    );
    port (
        -- the source domain: SCLK, or anything else somebody else owns
        sclk           : in  std_logic;
        src_rst_n      : in  std_logic;
        src_event      : in  std_logic;   -- one sclk cycle per event
        src_busy       : out std_logic;   -- scheme 4 only: an event is in flight

        -- the destination domain: the system clock
        clk            : in  std_logic;
        dst_rst_n      : in  std_logic;

        dst_level_stb  : out std_logic;   -- scheme 1
        dst_sync_stb   : out std_logic;   -- scheme 2
        dst_toggle_stb : out std_logic;   -- scheme 3
        dst_hs_stb     : out std_logic;   -- scheme 4

        -- what was generated, and what arrived
        src_events     : out unsigned(CNT_W - 1 downto 0);
        dst_level_n    : out unsigned(CNT_W - 1 downto 0);
        dst_sync_n     : out unsigned(CNT_W - 1 downto 0);
        dst_toggle_n   : out unsigned(CNT_W - 1 downto 0);
        dst_hs_n       : out unsigned(CNT_W - 1 downto 0)
    );
end entity;

architecture rtl of spi_cdc_probe is

    -- The acknowledgement, declared here because it is written by the DESTINATION
    -- and read by the SOURCE. A signal crossing back is easy to miss when reading
    -- a handshake, and it is half of what a handshake costs.
    signal dst_hs_ack : std_logic := '0';

    -- source domain
    signal pulse_src   : std_logic := '0';
    signal tog_src     : std_logic := '0';
    signal hs_tog_src  : std_logic := '0';
    signal hs_ack_sr   : std_logic_vector(SYNC_N - 1 downto 0) := (others => '0');
    signal hs_ack_seen : std_logic := '0';
    signal src_ev_r    : unsigned(CNT_W - 1 downto 0) := (others => '0');
    signal busy_i      : std_logic;

    -- destination domain
    signal lvl_q, lvl_q_d       : std_logic := '0';
    signal sync_sr              : std_logic_vector(SYNC_N - 1 downto 0) := (others => '0');
    signal sync_q_d             : std_logic := '0';
    signal tog_sr               : std_logic_vector(SYNC_N - 1 downto 0) := (others => '0');
    signal tog_q_d              : std_logic := '0';
    signal hs_tog_sr            : std_logic_vector(SYNC_N - 1 downto 0) := (others => '0');
    signal hs_tog_q_d           : std_logic := '0';

    signal lvl_stb_r, syn_stb_r, tog_stb_r, hs_stb_r : std_logic := '0';
    signal lvl_n_r, syn_n_r, tog_n_r, hs_n_r
        : unsigned(CNT_W - 1 downto 0) := (others => '0');

begin

    busy_i   <= hs_tog_src xor hs_ack_seen;
    src_busy <= busy_i;

    src_events     <= src_ev_r;
    dst_level_stb  <= lvl_stb_r;
    dst_sync_stb   <= syn_stb_r;
    dst_toggle_stb <= tog_stb_r;
    dst_hs_stb     <= hs_stb_r;
    dst_level_n    <= lvl_n_r;
    dst_sync_n     <= syn_n_r;
    dst_toggle_n   <= tog_n_r;
    dst_hs_n       <= hs_n_r;

    -- =====================================================================
    -- SOURCE DOMAIN
    -- =====================================================================
    source : process (sclk, src_rst_n)
    begin
        if src_rst_n = '0' then
            pulse_src   <= '0';
            tog_src     <= '0';
            hs_tog_src  <= '0';
            hs_ack_sr   <= (others => '0');
            hs_ack_seen <= '0';
            src_ev_r    <= (others => '0');
        elsif rising_edge(sclk) then
            hs_ack_sr <= hs_ack_sr(SYNC_N - 2 downto 0) & dst_hs_ack;

            -- The pulse is one source cycle wide and then gone. That is what
            -- makes it a pulse, and it is what makes it losable.
            pulse_src <= src_event;

            if src_event = '1' then
                src_ev_r <= src_ev_r + 1;
                tog_src  <= not tog_src;

                -- Scheme 4 only accepts an event when the previous one has been
                -- acknowledged. A source that ignores `src_busy` and toggles
                -- anyway has built scheme 3 with extra logic.
                if busy_i = '0' then
                    hs_tog_src <= not hs_tog_src;
                end if;
            end if;

            -- The ack has come back; the handshake is free for the next event.
            if hs_ack_sr(SYNC_N - 1) /= hs_ack_seen then
                hs_ack_seen <= hs_ack_sr(SYNC_N - 1);
            end if;
        end if;
    end process;

    -- =====================================================================
    -- DESTINATION DOMAIN
    -- =====================================================================
    destination : process (clk, dst_rst_n)
    begin
        if dst_rst_n = '0' then
            lvl_q      <= '0';
            lvl_q_d    <= '0';
            sync_sr    <= (others => '0');
            sync_q_d   <= '0';
            tog_sr     <= (others => '0');
            tog_q_d    <= '0';
            hs_tog_sr  <= (others => '0');
            hs_tog_q_d <= '0';
            dst_hs_ack <= '0';
            lvl_stb_r  <= '0';
            syn_stb_r  <= '0';
            tog_stb_r  <= '0';
            hs_stb_r   <= '0';
            lvl_n_r    <= (others => '0');
            syn_n_r    <= (others => '0');
            tog_n_r    <= (others => '0');
            hs_n_r     <= (others => '0');
        elsif rising_edge(clk) then
            lvl_q      <= pulse_src;
            lvl_q_d    <= lvl_q;
            sync_sr    <= sync_sr(SYNC_N - 2 downto 0) & pulse_src;
            sync_q_d   <= sync_sr(SYNC_N - 1);
            tog_sr     <= tog_sr(SYNC_N - 2 downto 0) & tog_src;
            tog_q_d    <= tog_sr(SYNC_N - 1);
            hs_tog_sr  <= hs_tog_sr(SYNC_N - 2 downto 0) & hs_tog_src;
            hs_tog_q_d <= hs_tog_sr(SYNC_N - 1);

            -- the four deliveries
            lvl_stb_r <= lvl_q and (not lvl_q_d);
            syn_stb_r <= sync_sr(SYNC_N - 1) and (not sync_q_d);
            tog_stb_r <= tog_sr(SYNC_N - 1) xor tog_q_d;
            hs_stb_r  <= hs_tog_sr(SYNC_N - 1) xor hs_tog_q_d;

            -- The ack mirrors the observed toggle, one cycle behind the strobe
            -- so it cannot be mistaken for the event itself.
            if hs_tog_sr(SYNC_N - 1) /= hs_tog_q_d then
                dst_hs_ack <= hs_tog_sr(SYNC_N - 1);
            end if;

            if lvl_stb_r = '1' then lvl_n_r <= lvl_n_r + 1; end if;
            if syn_stb_r = '1' then syn_n_r <= syn_n_r + 1; end if;
            if tog_stb_r = '1' then tog_n_r <= tog_n_r + 1; end if;
            if hs_stb_r  = '1' then hs_n_r  <= hs_n_r  + 1; end if;
        end if;
    end process;

    -- The check applies to the PULSE schemes only, and the reason is worth a note
    -- because the obvious generalisation is wrong. Schemes 1 and 2 detect a RISING
    -- edge, so two deliveries need a fall between them and cannot be adjacent.
    -- Schemes 3 and 4 detect BOTH edges, so two toggles on consecutive
    -- destination cycles produce two adjacent strobes -- which is two events
    -- correctly delivered, not one event delivered twice. A width check there
    -- would fail on the design working.
    check : process (clk)
        variable lvl_d : std_logic := '0';
        variable syn_d : std_logic := '0';
    begin
        if rising_edge(clk) then
            if dst_rst_n = '1' then
                assert not (lvl_stb_r = '1' and lvl_d = '1')
                    report "level strobe wider than one cycle" severity failure;
                assert not (syn_stb_r = '1' and syn_d = '1')
                    report "sync strobe wider than one cycle" severity failure;
            end if;
            lvl_d := lvl_stb_r;
            syn_d := syn_stb_r;
        end if;
    end process;

end architecture;

The testbench

One sweep and five directed cases. The sweep produces the table in §3; the directed cases push each scheme past its own limit.

The bench's most unusual assertion is the one in §3's callout: schemes 1 and 2 must agree at every ratio. A bench that found a difference there would have found a bug in an edge detector, not evidence about metastability, and stating that in the bench keeps a reader from mistaking a green run for a proof.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_cdc_probe_tb.sv — one ratio sweep and five directed cases, including an assertion that two schemes must AGREE
// spi_cdc_probe_tb.sv
//
// One experiment, run at many clock ratios, and the result is a table rather than
// a pass/fail: how many events each of the four crossing schemes delivered out of
// how many were generated.
//
// The ratios are chosen around the two limits the schemes have:
//
//   PULSE WIDTH   a source pulse is one source period wide, so schemes 1 and 2
//                 start losing events as soon as the source period is shorter
//                 than a destination period.
//   EVENT RATE    a toggle survives any pulse width, so scheme 3 only loses
//                 events when two of them fall inside one destination period.
//
// And the experiment that is NOT here, because it cannot be: nothing in this
// bench distinguishes scheme 1 from scheme 2. That is checked explicitly -- the
// bench asserts their counters are EQUAL at every ratio -- so the file states the
// limit of its own evidence rather than leaving a reader to assume the regression
// covered it.

`timescale 1ns/1ps

module spi_cdc_probe_tb;

    localparam int SYNC_N = 2;
    localparam int CNT_W  = 16;

    // The destination clock is fixed at 10 ns. The source clock is swept.
    localparam time DST_PERIOD = 10;

    reg clk   = 1'b0;
    reg sclk  = 1'b0;
    reg dst_rst_n = 1'b0;
    reg src_rst_n = 1'b0;
    reg src_event = 1'b0;

    // The source half-period in nanoseconds, driven by the sweep.
    integer src_half = 50;
    reg     src_run  = 1'b0;

    always #(DST_PERIOD/2) clk = ~clk;

    // A source clock whose period the test changes between runs. Generated from a
    // variable delay rather than a parameter, because the whole point is to sweep
    // the ratio within one simulation.
    initial begin
        forever begin
            #(src_half);
            if (src_run) sclk = ~sclk;
        end
    end

    wire              src_busy;
    wire              dst_level_stb, dst_sync_stb, dst_toggle_stb, dst_hs_stb;
    wire [CNT_W-1:0]  src_events, dst_level_n, dst_sync_n, dst_toggle_n, dst_hs_n;

    spi_cdc_probe #(.SYNC_N(SYNC_N), .CNT_W(CNT_W)) dut (
        .sclk(sclk), .src_rst_n(src_rst_n), .src_event(src_event),
        .src_busy(src_busy),
        .clk(clk), .dst_rst_n(dst_rst_n),
        .dst_level_stb(dst_level_stb), .dst_sync_stb(dst_sync_stb),
        .dst_toggle_stb(dst_toggle_stb), .dst_hs_stb(dst_hs_stb),
        .src_events(src_events),
        .dst_level_n(dst_level_n), .dst_sync_n(dst_sync_n),
        .dst_toggle_n(dst_toggle_n), .dst_hs_n(dst_hs_n)
    );

    integer errors = 0;

    // A watchdog, because a bench with two independent clocks and a handshake has
    // a real way to wait forever.
    initial begin
        #4_000_000;
        $display("FAIL: the simulation did not finish within its time limit");
        $finish;
    end

    // The source clock must be RUNNING while the source reset is asserted. Its
    // reset is asynchronous, which means it takes effect on a reset edge or on a
    // clock edge -- and at time zero there is no reset edge, because the signal
    // was initialised low rather than driven low. A source clock that only starts
    // after reset releases therefore never resets the source registers at all,
    // and they stay X for the whole run.
    //
    // This is not a subtlety of this bench. It is the ordinary reason a design
    // with an externally supplied clock reads X in its first simulation, and the
    // fix is the same: clock the domain while it is held in reset.
    task automatic restart(input integer half_ns);
        begin
            src_run   = 1'b0;
            src_event = 1'b0;
            dst_rst_n = 1'b0;
            src_rst_n = 1'b0;
            sclk      = 1'b0;
            src_half  = half_ns;
            repeat (4) @(posedge clk);
            src_run   = 1'b1;               // clock it WHILE reset is asserted
            repeat (4) @(posedge sclk);
            repeat (4) @(posedge clk);
            dst_rst_n = 1'b1;
            src_rst_n = 1'b1;
            repeat (4) @(posedge clk);
            repeat (2) @(posedge sclk);
        end
    endtask

    // Generates `n` events, one every `gap` source cycles, IGNORING src_busy --
    // which is what a slave does, because SCLK does not stop to be convenient.
    // Scheme 4 therefore drops events here too, and that is the honest result:
    // a handshake only conserves events for a source that can be stalled.
    task automatic burst_free(input integer n, input integer gap);
        integer i, g;
        begin
            for (i = 0; i < n; i = i + 1) begin
                @(negedge sclk);
                src_event = 1'b1;
                @(negedge sclk);
                src_event = 1'b0;
                for (g = 1; g < gap; g = g + 1) @(negedge sclk);
            end
        end
    endtask

    // The same burst, but the source WAITS when the handshake is busy. This is the
    // stimulus scheme 4 was designed for, and the difference between the two tasks
    // is the price of scheme 4 rather than a property of it.
    task automatic burst_polite(input integer n, input integer guard);
        integer i, w;
        begin
            for (i = 0; i < n; i = i + 1) begin
                w = 0;
                while (src_busy && w < guard) begin
                    @(negedge sclk);
                    w = w + 1;
                end
                if (w == guard) begin
                    $display("  FAIL: the handshake never freed within %0d source cycles", guard);
                    errors = errors + 1;
                end
                @(negedge sclk);
                src_event = 1'b1;
                @(negedge sclk);
                src_event = 1'b0;
            end
        end
    endtask

    // Lets every scheme's pipeline drain before the counters are read. Each
    // scheme is at most SYNC_N + 2 destination cycles deep, and the handshake
    // needs a round trip, so the settle is generous rather than tight.
    task automatic settle;
        begin
            repeat (40) @(posedge clk);
            repeat (8)  @(posedge sclk);
            repeat (40) @(posedge clk);
        end
    endtask

    integer halves [0:5];
    integer h, k;
    integer gen, lvl, syn, tog, hs;

    initial begin
        halves[0] = 50;   // source period 100 ns: 10x slower than the destination
        halves[1] = 25;   // 50 ns:  5x slower
        halves[2] = 10;   // 20 ns:  2x slower
        halves[3] = 5;    // 10 ns:  the same period
        halves[4] = 3;    // 6 ns:   faster than the destination
        halves[5] = 2;    // 4 ns:   2.5x faster

        $display("  ratio sweep: 20 events per run, one event every 2 source cycles");
        $display("  src_period  generated  scheme1  scheme2  scheme3  scheme4");

        for (h = 0; h <= 5; h = h + 1) begin
            restart(halves[h]);
            burst_free(20, 2);
            settle();

            gen = src_events; lvl = dst_level_n; syn = dst_sync_n;
            tog = dst_toggle_n; hs = dst_hs_n;

            $display("  %8d ns  %9d  %7d  %7d  %7d  %7d",
                     2*halves[h], gen, lvl, syn, tog, hs);

            // 1. THE SOURCE GENERATED WHAT THE BENCH ASKED FOR. Checked first,
            //    because every other number is relative to this one.
            if (gen != 20) begin
                $display("  FAIL: the source generated %0d events, expected 20", gen);
                errors = errors + 1;
            end

            // 2. SCHEMES 1 AND 2 ARE INDISTINGUISHABLE HERE. This is the
            //    assertion that states the limit of the whole bench: the extra
            //    flop buys settling, and settling is not modelled, so the two
            //    counters must agree at every ratio. A bench that found a
            //    difference would have found a bug in the edge detectors, not
            //    evidence about metastability.
            if (lvl != syn) begin
                $display("  FAIL: schemes 1 and 2 delivered %0d and %0d -- they must agree in simulation, because the only difference between them is unsimulatable",
                         lvl, syn);
                errors = errors + 1;
            end

            // 3. NO SCHEME INVENTS EVENTS. Losing them is a bandwidth limit;
            //    gaining them is a design error, and the two must never be
            //    confused by a reader of the table.
            if (lvl > gen || tog > gen || hs > gen) begin
                $display("  FAIL: a scheme delivered more events than were generated");
                errors = errors + 1;
            end

            // 4. A COMFORTABLE RATIO LOSES NOTHING, in any scheme. Two source
            //    cycles between events at a source period of 100 ns is twenty
            //    destination cycles, so there is no excuse for a loss.
            if (halves[h] >= 25 && (lvl != gen || tog != gen)) begin
                $display("  FAIL: at a source period of %0d ns nothing should be lost, got %0d and %0d of %0d",
                         2*halves[h], lvl, tog, gen);
                errors = errors + 1;
            end

            // 5. THE TOGGLE IS NEVER WORSE THAN THE PULSE. It cannot be: a
            //    toggle is a level that persists and a pulse is one that does
            //    not, so any event the pulse survived the toggle survived too.
            if (tog < lvl) begin
                $display("  FAIL: the toggle delivered %0d and the pulse %0d -- the toggle cannot be worse",
                         tog, lvl);
                errors = errors + 1;
            end
        end

        // 5b. THE HANDSHAKE'S ROUND TRIP IS A LIMIT AT EVERY RATIO, not only at a
        //     fast one, and the table above shows it: an impolite source loses
        //     events even at a source period ten times the destination's. The
        //     round trip is SYNC_N destination cycles out and SYNC_N SOURCE
        //     cycles back, so a source producing events every two source cycles
        //     outruns it however slow it is. This is the number Chapter 15.6 has
        //     to size a FIFO against.
        restart(50);
        burst_free(20, 2);
        settle();
        if (dst_hs_n >= src_events) begin
            $display("  FAIL: an impolite source two source cycles apart should outrun the handshake's round trip even at a slow ratio, got %0d of %0d",
                     dst_hs_n, src_events);
            errors = errors + 1;
        end
        $display("  at a source period of 100 ns -- ten times the destination's -- an impolite source still lost events through the handshake, delivering %0d of %0d, because the round trip is SYNC_N destination cycles out and SYNC_N SOURCE cycles back and no ratio makes that free",
                 dst_hs_n, src_events);

        // 6. THE PULSE SCHEMES ACTUALLY DO LOSE EVENTS AT A FAST SOURCE. Without
        //    this the table above could be produced by four identical schemes,
        //    and the chapter's claim would rest on nothing.
        restart(2);
        burst_free(20, 2);
        settle();
        if (dst_level_n >= src_events) begin
            $display("  FAIL: at a source period of 4 ns the one-flop scheme delivered %0d of %0d -- a pulse one source cycle wide cannot survive a destination period 2.5x longer",
                     dst_level_n, src_events);
            errors = errors + 1;
        end
        // The exact count here is phase-dependent: which pulses happen to straddle a
        // destination edge depends on the relationship between two clocks that have
        // none. So the assertion is that events are LOST, never that a particular
        // number survives -- the three language benches report 4 and 5 for the same
        // reason two boards would.
        $display("  at a source period of 4 ns against a 10 ns destination, the pulse schemes delivered %0d of %0d and the toggle delivered %0d -- the exact survivor count is phase-dependent, which is why the check is that events were lost rather than how many",
                 dst_level_n, src_events, dst_toggle_n);

        // 7. AND THE TOGGLE LOSES THEM TOO, ONCE THE EVENT RATE IS HIGH ENOUGH.
        //    This is the point that stops "use a toggle" from being the answer:
        //    it moves the limit, it does not remove it. Events every source cycle
        //    at a source period of 4 ns is one event per 4 ns into a 10 ns clock.
        restart(2);
        burst_free(20, 1);
        settle();
        if (dst_toggle_n >= src_events) begin
            $display("  FAIL: at one event per source cycle the toggle delivered %0d of %0d -- two toggles inside one destination period must cancel",
                     dst_toggle_n, src_events);
            errors = errors + 1;
        end
        $display("  at one event per source cycle the toggle delivered %0d of %0d, because two toggles inside one destination period leave the level unchanged",
                 dst_toggle_n, src_events);

        // 8. THE HANDSHAKE CONSERVES EVERYTHING -- FOR A SOURCE THAT WAITS. The
        //    same fast source, now honouring src_busy, loses nothing at all. That
        //    is the whole trade: correctness bought with the source's freedom.
        restart(2);
        burst_polite(20, 2000);
        settle();
        if (dst_hs_n != src_events) begin
            $display("  FAIL: with a source that waits, the handshake delivered %0d of %0d",
                     dst_hs_n, src_events);
            errors = errors + 1;
        end
        if (src_events != 20) begin
            $display("  FAIL: the polite source generated %0d events, expected 20", src_events);
            errors = errors + 1;
        end
        $display("  with a source that honours src_busy, the handshake delivered %0d of %0d at the same 4 ns source period -- nothing lost, and the source stalled to achieve it",
                 dst_hs_n, src_events);

        // 9. AND A SOURCE THAT CANNOT WAIT GETS NO BENEFIT FROM THE HANDSHAKE.
        //    A slave is that source: SCLK does not stop. So the handshake's
        //    guarantee is conditional on something an SPI slave cannot offer,
        //    which is why Chapter 15.6 has to size a FIFO instead.
        restart(2);
        burst_free(20, 1);
        settle();
        if (dst_hs_n >= src_events) begin
            $display("  FAIL: an impolite fast source should lose events even with the handshake, got %0d of %0d",
                     dst_hs_n, src_events);
            errors = errors + 1;
        end
        $display("  the same handshake against a source that ignores src_busy delivered %0d of %0d -- the guarantee was never about the scheme, it was about the source being stallable",
                 dst_hs_n, src_events);

        if (errors == 0)
            $display("PASS: an event crossing a clock boundary is limited by something in every scheme, and simulation settles which -- a one-source-cycle pulse is lost as soon as the source period is shorter than the destination's, and adding a second synchroniser flop changes that not at all, because the two schemes' counters agree at every ratio and the only difference between them is a settling time no simulator models -- a toggle survives any pulse width and moves the limit to the EVENT RATE, losing events only when two of them fall inside one destination period, which they do at one event per source cycle -- and a four-phase handshake loses nothing at all, but only for a source that honours its busy signal, so its guarantee is conditional on a freedom an SPI slave does not have, because SCLK does not stop: no crossing scheme creates bandwidth, and the choice between them is a choice about which limit to accept");
        else
            $display("FAIL: %0d error(s)", errors);
        $finish;
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_cdc_probe_tb.v — the same bench in Verilog-2001
// spi_cdc_probe_tb.v
//
// One experiment, run at many clock ratios, and the result is a table rather than
// a pass/fail: how many events each of the four crossing schemes delivered out of
// how many were generated.
//
// The ratios are chosen around the two limits the schemes have:
//
//   PULSE WIDTH   a source pulse is one source period wide, so schemes 1 and 2
//                 start losing events as soon as the source period is shorter
//                 than a destination period.
//   EVENT RATE    a toggle survives any pulse width, so scheme 3 only loses
//                 events when two of them fall inside one destination period.
//
// And the experiment that is NOT here, because it cannot be: nothing in this
// bench distinguishes scheme 1 from scheme 2. That is checked explicitly -- the
// bench asserts their counters are EQUAL at every ratio -- so the file states the
// limit of its own evidence rather than leaving a reader to assume the regression
// covered it.

`timescale 1ns/1ps

module spi_cdc_probe_tb;

    localparam SYNC_N = 2;
    localparam CNT_W  = 16;

    // The destination clock is fixed at 10 ns. The source clock is swept.
    localparam time DST_PERIOD = 10;

    reg clk;
    reg sclk;
    reg dst_rst_n;
    reg src_rst_n;
    reg src_event;

    // The source half-period in nanoseconds, driven by the sweep.
    integer src_half;
    reg     src_run;

    always #(DST_PERIOD/2) clk = ~clk;

    // A source clock whose period the test changes between runs. Generated from a
    // variable delay rather than a parameter, because the whole point is to sweep
    // the ratio within one simulation.
    initial begin
        forever begin
            #(src_half);
            if (src_run) sclk = ~sclk;
        end
    end

    wire              src_busy;
    wire              dst_level_stb, dst_sync_stb, dst_toggle_stb, dst_hs_stb;
    wire [CNT_W-1:0]  src_events, dst_level_n, dst_sync_n, dst_toggle_n, dst_hs_n;

    spi_cdc_probe #(.SYNC_N(SYNC_N), .CNT_W(CNT_W)) dut (
        .sclk(sclk), .src_rst_n(src_rst_n), .src_event(src_event),
        .src_busy(src_busy),
        .clk(clk), .dst_rst_n(dst_rst_n),
        .dst_level_stb(dst_level_stb), .dst_sync_stb(dst_sync_stb),
        .dst_toggle_stb(dst_toggle_stb), .dst_hs_stb(dst_hs_stb),
        .src_events(src_events),
        .dst_level_n(dst_level_n), .dst_sync_n(dst_sync_n),
        .dst_toggle_n(dst_toggle_n), .dst_hs_n(dst_hs_n)
    );

    integer errors;

    // A watchdog, because a bench with two independent clocks and a handshake has
    // a real way to wait forever.
    initial begin
        #4_000_000;
        $display("FAIL: the simulation did not finish within its time limit");
        $finish;
    end

    // The source clock must be RUNNING while the source reset is asserted. Its
    // reset is asynchronous, which means it takes effect on a reset edge or on a
    // clock edge -- and at time zero there is no reset edge, because the signal
    // was initialised low rather than driven low. A source clock that only starts
    // after reset releases therefore never resets the source registers at all,
    // and they stay X for the whole run.
    //
    // This is not a subtlety of this bench. It is the ordinary reason a design
    // with an externally supplied clock reads X in its first simulation, and the
    // fix is the same: clock the domain while it is held in reset.
        task restart;
        input integer half_ns;
        begin
            src_run   = 1'b0;
            src_event = 1'b0;
            dst_rst_n = 1'b0;
            src_rst_n = 1'b0;
            sclk      = 1'b0;
            src_half  = half_ns;
            repeat (4) @(posedge clk);
            src_run   = 1'b1;               // clock it WHILE reset is asserted
            repeat (4) @(posedge sclk);
            repeat (4) @(posedge clk);
            dst_rst_n = 1'b1;
            src_rst_n = 1'b1;
            repeat (4) @(posedge clk);
            repeat (2) @(posedge sclk);
        end
    endtask

    // Generates `n` events, one every `gap` source cycles, IGNORING src_busy --
    // which is what a slave does, because SCLK does not stop to be convenient.
    // Scheme 4 therefore drops events here too, and that is the honest result:
    // a handshake only conserves events for a source that can be stalled.
        task burst_free;
        input integer n;
        input integer gap;
        integer i, g;
        begin
            for (i = 0; i < n; i = i + 1) begin
                @(negedge sclk);
                src_event = 1'b1;
                @(negedge sclk);
                src_event = 1'b0;
                for (g = 1; g < gap; g = g + 1) @(negedge sclk);
            end
        end
    endtask

    // The same burst, but the source WAITS when the handshake is busy. This is the
    // stimulus scheme 4 was designed for, and the difference between the two tasks
    // is the price of scheme 4 rather than a property of it.
        task burst_polite;
        input integer n;
        input integer guard;
        integer i, w;
        begin
            for (i = 0; i < n; i = i + 1) begin
                w = 0;
                while (src_busy && w < guard) begin
                    @(negedge sclk);
                    w = w + 1;
                end
                if (w == guard) begin
                    $display("  FAIL: the handshake never freed within %0d source cycles", guard);
                    errors = errors + 1;
                end
                @(negedge sclk);
                src_event = 1'b1;
                @(negedge sclk);
                src_event = 1'b0;
            end
        end
    endtask

    // Lets every scheme's pipeline drain before the counters are read. Each
    // scheme is at most SYNC_N + 2 destination cycles deep, and the handshake
    // needs a round trip, so the settle is generous rather than tight.
    task settle;
        begin
            repeat (40) @(posedge clk);
            repeat (8)  @(posedge sclk);
            repeat (40) @(posedge clk);
        end
    endtask

    integer halves [0:5];
    integer h, k;
    integer gen, lvl, syn, tog, hs;

    initial begin
        halves[0] = 50;   // source period 100 ns: 10x slower than the destination
        halves[1] = 25;   // 50 ns:  5x slower
        halves[2] = 10;   // 20 ns:  2x slower
        halves[3] = 5;    // 10 ns:  the same period
        halves[4] = 3;    // 6 ns:   faster than the destination
        halves[5] = 2;    // 4 ns:   2.5x faster

        $display("  ratio sweep: 20 events per run, one event every 2 source cycles");
        $display("  src_period  generated  scheme1  scheme2  scheme3  scheme4");

        for (h = 0; h <= 5; h = h + 1) begin
            restart(halves[h]);
            burst_free(20, 2);
            settle();

            gen = src_events; lvl = dst_level_n; syn = dst_sync_n;
            tog = dst_toggle_n; hs = dst_hs_n;

            $display("  %8d ns  %9d  %7d  %7d  %7d  %7d",
                     2*halves[h], gen, lvl, syn, tog, hs);

            // 1. THE SOURCE GENERATED WHAT THE BENCH ASKED FOR. Checked first,
            //    because every other number is relative to this one.
            if (gen != 20) begin
                $display("  FAIL: the source generated %0d events, expected 20", gen);
                errors = errors + 1;
            end

            // 2. SCHEMES 1 AND 2 ARE INDISTINGUISHABLE HERE. This is the
            //    assertion that states the limit of the whole bench: the extra
            //    flop buys settling, and settling is not modelled, so the two
            //    counters must agree at every ratio. A bench that found a
            //    difference would have found a bug in the edge detectors, not
            //    evidence about metastability.
            if (lvl != syn) begin
                $display("  FAIL: schemes 1 and 2 delivered %0d and %0d -- they must agree in simulation, because the only difference between them is unsimulatable",
                         lvl, syn);
                errors = errors + 1;
            end

            // 3. NO SCHEME INVENTS EVENTS. Losing them is a bandwidth limit;
            //    gaining them is a design error, and the two must never be
            //    confused by a reader of the table.
            if (lvl > gen || tog > gen || hs > gen) begin
                $display("  FAIL: a scheme delivered more events than were generated");
                errors = errors + 1;
            end

            // 4. A COMFORTABLE RATIO LOSES NOTHING, in any scheme. Two source
            //    cycles between events at a source period of 100 ns is twenty
            //    destination cycles, so there is no excuse for a loss.
            if (halves[h] >= 25 && (lvl != gen || tog != gen)) begin
                $display("  FAIL: at a source period of %0d ns nothing should be lost, got %0d and %0d of %0d",
                         2*halves[h], lvl, tog, gen);
                errors = errors + 1;
            end

            // 5. THE TOGGLE IS NEVER WORSE THAN THE PULSE. It cannot be: a
            //    toggle is a level that persists and a pulse is one that does
            //    not, so any event the pulse survived the toggle survived too.
            if (tog < lvl) begin
                $display("  FAIL: the toggle delivered %0d and the pulse %0d -- the toggle cannot be worse",
                         tog, lvl);
                errors = errors + 1;
            end
        end

        // 5b. THE HANDSHAKE'S ROUND TRIP IS A LIMIT AT EVERY RATIO, not only at a
        //     fast one, and the table above shows it: an impolite source loses
        //     events even at a source period ten times the destination's. The
        //     round trip is SYNC_N destination cycles out and SYNC_N SOURCE
        //     cycles back, so a source producing events every two source cycles
        //     outruns it however slow it is. This is the number Chapter 15.6 has
        //     to size a FIFO against.
        restart(50);
        burst_free(20, 2);
        settle();
        if (dst_hs_n >= src_events) begin
            $display("  FAIL: an impolite source two source cycles apart should outrun the handshake's round trip even at a slow ratio, got %0d of %0d",
                     dst_hs_n, src_events);
            errors = errors + 1;
        end
        $display("  at a source period of 100 ns -- ten times the destination's -- an impolite source still lost events through the handshake, delivering %0d of %0d, because the round trip is SYNC_N destination cycles out and SYNC_N SOURCE cycles back and no ratio makes that free",
                 dst_hs_n, src_events);

        // 6. THE PULSE SCHEMES ACTUALLY DO LOSE EVENTS AT A FAST SOURCE. Without
        //    this the table above could be produced by four identical schemes,
        //    and the chapter's claim would rest on nothing.
        restart(2);
        burst_free(20, 2);
        settle();
        if (dst_level_n >= src_events) begin
            $display("  FAIL: at a source period of 4 ns the one-flop scheme delivered %0d of %0d -- a pulse one source cycle wide cannot survive a destination period 2.5x longer",
                     dst_level_n, src_events);
            errors = errors + 1;
        end
        // The exact count here is phase-dependent: which pulses happen to straddle a
        // destination edge depends on the relationship between two clocks that have
        // none. So the assertion is that events are LOST, never that a particular
        // number survives -- the three language benches report 4 and 5 for the same
        // reason two boards would.
        $display("  at a source period of 4 ns against a 10 ns destination, the pulse schemes delivered %0d of %0d and the toggle delivered %0d -- the exact survivor count is phase-dependent, which is why the check is that events were lost rather than how many",
                 dst_level_n, src_events, dst_toggle_n);

        // 7. AND THE TOGGLE LOSES THEM TOO, ONCE THE EVENT RATE IS HIGH ENOUGH.
        //    This is the point that stops "use a toggle" from being the answer:
        //    it moves the limit, it does not remove it. Events every source cycle
        //    at a source period of 4 ns is one event per 4 ns into a 10 ns clock.
        restart(2);
        burst_free(20, 1);
        settle();
        if (dst_toggle_n >= src_events) begin
            $display("  FAIL: at one event per source cycle the toggle delivered %0d of %0d -- two toggles inside one destination period must cancel",
                     dst_toggle_n, src_events);
            errors = errors + 1;
        end
        $display("  at one event per source cycle the toggle delivered %0d of %0d, because two toggles inside one destination period leave the level unchanged",
                 dst_toggle_n, src_events);

        // 8. THE HANDSHAKE CONSERVES EVERYTHING -- FOR A SOURCE THAT WAITS. The
        //    same fast source, now honouring src_busy, loses nothing at all. That
        //    is the whole trade: correctness bought with the source's freedom.
        restart(2);
        burst_polite(20, 2000);
        settle();
        if (dst_hs_n != src_events) begin
            $display("  FAIL: with a source that waits, the handshake delivered %0d of %0d",
                     dst_hs_n, src_events);
            errors = errors + 1;
        end
        if (src_events != 20) begin
            $display("  FAIL: the polite source generated %0d events, expected 20", src_events);
            errors = errors + 1;
        end
        $display("  with a source that honours src_busy, the handshake delivered %0d of %0d at the same 4 ns source period -- nothing lost, and the source stalled to achieve it",
                 dst_hs_n, src_events);

        // 9. AND A SOURCE THAT CANNOT WAIT GETS NO BENEFIT FROM THE HANDSHAKE.
        //    A slave is that source: SCLK does not stop. So the handshake's
        //    guarantee is conditional on something an SPI slave cannot offer,
        //    which is why Chapter 15.6 has to size a FIFO instead.
        restart(2);
        burst_free(20, 1);
        settle();
        if (dst_hs_n >= src_events) begin
            $display("  FAIL: an impolite fast source should lose events even with the handshake, got %0d of %0d",
                     dst_hs_n, src_events);
            errors = errors + 1;
        end
        $display("  the same handshake against a source that ignores src_busy delivered %0d of %0d -- the guarantee was never about the scheme, it was about the source being stallable",
                 dst_hs_n, src_events);

        if (errors == 0)
            $display("PASS: an event crossing a clock boundary is limited by something in every scheme, and simulation settles which -- a one-source-cycle pulse is lost as soon as the source period is shorter than the destination's, and adding a second synchroniser flop changes that not at all, because the two schemes' counters agree at every ratio and the only difference between them is a settling time no simulator models -- a toggle survives any pulse width and moves the limit to the EVENT RATE, losing events only when two of them fall inside one destination period, which they do at one event per source cycle -- and a four-phase handshake loses nothing at all, but only for a source that honours its busy signal, so its guarantee is conditional on a freedom an SPI slave does not have, because SCLK does not stop: no crossing scheme creates bandwidth, and the choice between them is a choice about which limit to accept");
        else
            $display("FAIL: %0d error(s)", errors);
        $finish;
    end


    initial begin
        clk = 1'b0;
        sclk = 1'b0;
        dst_rst_n = 1'b0;
        src_rst_n = 1'b0;
        src_event = 1'b0;
        src_half = 50;
        src_run = 1'b0;
        errors = 0;
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_cdc_probe_tb.vhd — the same bench in VHDL
-- spi_cdc_probe_tb.vhd
--
-- One experiment, run at many clock ratios, and the result is a table rather than
-- a pass/fail: how many events each of the four crossing schemes delivered out of
-- how many were generated.
--
-- The ratios are chosen around the two limits the schemes have:
--
--   PULSE WIDTH   a source pulse is one source period wide, so schemes 1 and 2
--                 start losing events as soon as the source period is shorter
--                 than a destination period.
--   EVENT RATE    a toggle survives any pulse width, so scheme 3 only loses
--                 events when two of them fall inside one destination period.
--
-- And the experiment that is NOT here, because it cannot be: nothing in this
-- bench distinguishes scheme 1 from scheme 2. That is checked explicitly -- the
-- bench asserts their counters are EQUAL at every ratio -- so the file states the
-- limit of its own evidence rather than leaving a reader to assume the regression
-- covered it.

library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

entity spi_cdc_probe_tb is
end entity;

architecture sim of spi_cdc_probe_tb is

    constant SYNC_N : positive := 2;
    constant CNT_W  : positive := 16;

    -- The destination clock is fixed at 10 ns. The source clock is swept.
    constant DST_PERIOD : time := 10 ns;

    signal clk       : std_logic := '0';
    signal sclk      : std_logic := '0';
    signal dst_rst_n : std_logic := '0';
    signal src_rst_n : std_logic := '0';
    signal src_event : std_logic := '0';
    signal halt      : boolean   := false;

    -- The source half-period, driven by the sweep. A shared variable would be the
    -- obvious shape and is not portable; a signal read by the clock process is.
    signal src_half : time      := 50 ns;
    signal src_run  : boolean   := false;

    signal src_busy : std_logic;
    signal dst_level_stb, dst_sync_stb, dst_toggle_stb, dst_hs_stb : std_logic;
    signal src_events, dst_level_n, dst_sync_n, dst_toggle_n, dst_hs_n
        : unsigned(CNT_W - 1 downto 0);

begin

    dstclk : process
    begin
        while not halt loop
            clk <= '0'; wait for DST_PERIOD / 2;
            clk <= '1'; wait for DST_PERIOD / 2;
        end loop;
        wait;
    end process;

    -- A source clock whose period the test changes between runs, generated from a
    -- signal rather than a generic because the whole point is to sweep the ratio
    -- within one simulation.
    --
    -- This process is the ONLY driver of `sclk`, and that is not a stylistic
    -- choice. The SystemVerilog bench above also parks the clock low inside
    -- `restart`, which is harmless there because the last procedural write wins.
    -- In VHDL two processes driving one `std_logic` resolve to 'X', so
    -- `rising_edge(sclk)` never fires again and the bench waits forever -- a hang
    -- rather than a mismatch, which is the worst way for a bench to fail. So the
    -- clock is parked here, by its owner, rather than by the stimulus.
    srcclk : process
    begin
        while not halt loop
            if src_run then
                wait for src_half;
                sclk <= not sclk;
            else
                sclk <= '0';
                wait for 1 ns;
            end if;
        end loop;
        wait;
    end process;

    dut : entity work.spi_cdc_probe
        generic map (SYNC_N => SYNC_N, CNT_W => CNT_W)
        port map (sclk => sclk, src_rst_n => src_rst_n, src_event => src_event,
                  src_busy => src_busy,
                  clk => clk, dst_rst_n => dst_rst_n,
                  dst_level_stb => dst_level_stb, dst_sync_stb => dst_sync_stb,
                  dst_toggle_stb => dst_toggle_stb, dst_hs_stb => dst_hs_stb,
                  src_events => src_events,
                  dst_level_n => dst_level_n, dst_sync_n => dst_sync_n,
                  dst_toggle_n => dst_toggle_n, dst_hs_n => dst_hs_n);

    -- A watchdog, because a bench with two independent clocks and a handshake has
    -- a real way to wait forever. The `halt` guard matters: stopping the clocks
    -- does not stop simulation TIME, so a watchdog without it fires after the test
    -- has already passed.
    watchdog : process
    begin
        wait for 4 ms;
        if not halt then
            report "FAIL: the simulation did not finish within its time limit"
                severity failure;
        end if;
        wait;
    end process;

    stim : process
        variable errs : natural := 0;
        variable gen, lvl, syn, tog, hs : natural;
        type half_arr is array (0 to 5) of time;
        constant HALVES : half_arr := (50 ns, 25 ns, 10 ns, 5 ns, 3 ns, 2 ns);

        -- The source clock must be RUNNING while the source reset is asserted. Its
        -- reset is asynchronous, which means it takes effect on a reset edge or on
        -- a clock edge -- and at time zero there is no reset edge, because the
        -- signal was initialised low rather than driven low. A source clock that
        -- only starts after reset releases therefore never resets the source
        -- registers at all, and they stay undefined for the whole run.
        --
        -- This is not a subtlety of this bench. It is the ordinary reason a design
        -- with an externally supplied clock reads X in its first simulation, and
        -- the fix is the same: clock the domain while it is held in reset.
        procedure restart(half : time) is
        begin
            src_run   <= false;
            src_event <= '0';
            dst_rst_n <= '0';
            src_rst_n <= '0';
            src_half  <= half;   -- `sclk` is parked by its owning process
            for i in 1 to 4 loop wait until rising_edge(clk); end loop;
            src_run   <= true;               -- clock it WHILE reset is asserted
            for i in 1 to 4 loop wait until rising_edge(sclk); end loop;
            for i in 1 to 4 loop wait until rising_edge(clk); end loop;
            dst_rst_n <= '1';
            src_rst_n <= '1';
            for i in 1 to 4 loop wait until rising_edge(clk); end loop;
            for i in 1 to 2 loop wait until rising_edge(sclk); end loop;
        end procedure;

        -- Generates `n` events, one every `gap` source cycles, IGNORING src_busy --
        -- which is what a slave does, because SCLK does not stop to be convenient.
        procedure burst_free(n : natural; gap : natural) is
        begin
            for i in 1 to n loop
                wait until falling_edge(sclk);
                src_event <= '1';
                wait until falling_edge(sclk);
                src_event <= '0';
                for g in 2 to gap loop wait until falling_edge(sclk); end loop;
            end loop;
        end procedure;

        -- The same burst, but the source WAITS when the handshake is busy. The
        -- difference between the two procedures is the price of scheme 4 rather
        -- than a property of it.
        procedure burst_polite(n : natural; guard : natural) is
            variable w : natural;
        begin
            for i in 1 to n loop
                w := 0;
                while src_busy = '1' and w < guard loop
                    wait until falling_edge(sclk);
                    w := w + 1;
                end loop;
                if w = guard then
                    report "  FAIL: the handshake never freed within " &
                           integer'image(guard) & " source cycles";
                    errs := errs + 1;
                end if;
                wait until falling_edge(sclk);
                src_event <= '1';
                wait until falling_edge(sclk);
                src_event <= '0';
            end loop;
        end procedure;

        -- Lets every scheme's pipeline drain before the counters are read.
        procedure settle is
        begin
            for i in 1 to 40 loop wait until rising_edge(clk);  end loop;
            for i in 1 to 8  loop wait until rising_edge(sclk); end loop;
            for i in 1 to 40 loop wait until rising_edge(clk);  end loop;
        end procedure;
    begin
        report "  ratio sweep: 20 events per run, one event every 2 source cycles";
        report "  src_period  generated  scheme1  scheme2  scheme3  scheme4";

        for h in HALVES'range loop
            restart(HALVES(h));
            burst_free(20, 2);
            settle;

            gen := to_integer(src_events);   lvl := to_integer(dst_level_n);
            syn := to_integer(dst_sync_n);   tog := to_integer(dst_toggle_n);
            hs  := to_integer(dst_hs_n);

            -- `time'image` prints femtoseconds, which is unreadable in a table,
            -- so the period is divided down to an integer number of nanoseconds.
            report "  " & integer'image((2 * HALVES(h)) / 1 ns) &
                   " ns  generated " & integer'image(gen) & "  " &
                   integer'image(lvl) & "  " & integer'image(syn) & "  " &
                   integer'image(tog) & "  " & integer'image(hs);

            -- 1. THE SOURCE GENERATED WHAT THE BENCH ASKED FOR.
            if gen /= 20 then
                report "  FAIL: the source generated " & integer'image(gen) &
                       " events, expected 20";
                errs := errs + 1;
            end if;

            -- 2. SCHEMES 1 AND 2 ARE INDISTINGUISHABLE HERE. This is the
            --    assertion that states the limit of the whole bench.
            if lvl /= syn then
                report "  FAIL: schemes 1 and 2 delivered " & integer'image(lvl) &
                       " and " & integer'image(syn) &
                       " -- they must agree in simulation, because the only difference between them is unsimulatable";
                errs := errs + 1;
            end if;

            -- 3. NO SCHEME INVENTS EVENTS.
            if lvl > gen or tog > gen or hs > gen then
                report "  FAIL: a scheme delivered more events than were generated";
                errs := errs + 1;
            end if;

            -- 4. A COMFORTABLE RATIO LOSES NOTHING, in the pulse and toggle
            --    schemes. Scheme 4 is excluded deliberately: its round trip is a
            --    limit at every ratio, which test 5b establishes.
            if HALVES(h) >= 25 ns and (lvl /= gen or tog /= gen) then
                report "  FAIL: at a slow source nothing should be lost, got " &
                       integer'image(lvl) & " and " & integer'image(tog) &
                       " of " & integer'image(gen);
                errs := errs + 1;
            end if;

            -- 5. THE TOGGLE IS NEVER WORSE THAN THE PULSE.
            if tog < lvl then
                report "  FAIL: the toggle delivered " & integer'image(tog) &
                       " and the pulse " & integer'image(lvl) &
                       " -- the toggle cannot be worse";
                errs := errs + 1;
            end if;
        end loop;

        -- 5b. THE HANDSHAKE'S ROUND TRIP IS A LIMIT AT EVERY RATIO.
        restart(50 ns);
        burst_free(20, 2);
        settle;
        if to_integer(dst_hs_n) >= to_integer(src_events) then
            report "  FAIL: an impolite source two source cycles apart should outrun the handshake's round trip even at a slow ratio, got " &
                   integer'image(to_integer(dst_hs_n)) & " of " &
                   integer'image(to_integer(src_events));
            errs := errs + 1;
        end if;
        report "  at a source period of 100 ns -- ten times the destination's -- an impolite source still lost events through the handshake, delivering " &
               integer'image(to_integer(dst_hs_n)) & " of " &
               integer'image(to_integer(src_events)) &
               ", because the round trip is SYNC_N destination cycles out and SYNC_N SOURCE cycles back and no ratio makes that free";

        -- 6. THE PULSE SCHEMES ACTUALLY DO LOSE EVENTS AT A FAST SOURCE.
        restart(2 ns);
        burst_free(20, 2);
        settle;
        if to_integer(dst_level_n) >= to_integer(src_events) then
            report "  FAIL: at a source period of 4 ns the one-flop scheme delivered " &
                   integer'image(to_integer(dst_level_n)) & " of " &
                   integer'image(to_integer(src_events)) &
                   " -- a pulse one source cycle wide cannot survive a destination period 2.5x longer";
            errs := errs + 1;
        end if;
        -- The exact count here is phase-dependent: which pulses happen to straddle
        -- a destination edge depends on the relationship between two clocks that
        -- have none. So the assertion is that events are LOST, never that a
        -- particular number survives -- the three language benches report 4 and 5
        -- for the same reason two boards would.
        report "  at a source period of 4 ns against a 10 ns destination, the pulse schemes delivered " &
               integer'image(to_integer(dst_level_n)) & " of " &
               integer'image(to_integer(src_events)) & " and the toggle delivered " &
               integer'image(to_integer(dst_toggle_n)) &
               " -- the exact survivor count is phase-dependent, which is why the check is that events were lost rather than how many";

        -- 7. AND THE TOGGLE LOSES THEM TOO, ONCE THE EVENT RATE IS HIGH ENOUGH.
        restart(2 ns);
        burst_free(20, 1);
        settle;
        if to_integer(dst_toggle_n) >= to_integer(src_events) then
            report "  FAIL: at one event per source cycle the toggle delivered " &
                   integer'image(to_integer(dst_toggle_n)) & " of " &
                   integer'image(to_integer(src_events)) &
                   " -- two toggles inside one destination period must cancel";
            errs := errs + 1;
        end if;
        report "  at one event per source cycle the toggle delivered " &
               integer'image(to_integer(dst_toggle_n)) & " of " &
               integer'image(to_integer(src_events)) &
               ", because two toggles inside one destination period leave the level unchanged";

        -- 8. THE HANDSHAKE CONSERVES EVERYTHING -- FOR A SOURCE THAT WAITS.
        restart(2 ns);
        burst_polite(20, 2000);
        settle;
        if to_integer(dst_hs_n) /= to_integer(src_events) then
            report "  FAIL: with a source that waits, the handshake delivered " &
                   integer'image(to_integer(dst_hs_n)) & " of " &
                   integer'image(to_integer(src_events));
            errs := errs + 1;
        end if;
        if to_integer(src_events) /= 20 then
            report "  FAIL: the polite source generated " &
                   integer'image(to_integer(src_events)) & " events, expected 20";
            errs := errs + 1;
        end if;
        report "  with a source that honours src_busy, the handshake delivered " &
               integer'image(to_integer(dst_hs_n)) & " of " &
               integer'image(to_integer(src_events)) &
               " at the same 4 ns source period -- nothing lost, and the source stalled to achieve it";

        -- 9. AND A SOURCE THAT CANNOT WAIT GETS NO BENEFIT FROM THE HANDSHAKE.
        restart(2 ns);
        burst_free(20, 1);
        settle;
        if to_integer(dst_hs_n) >= to_integer(src_events) then
            report "  FAIL: an impolite fast source should lose events even with the handshake, got " &
                   integer'image(to_integer(dst_hs_n)) & " of " &
                   integer'image(to_integer(src_events));
            errs := errs + 1;
        end if;
        report "  the same handshake against a source that ignores src_busy delivered " &
               integer'image(to_integer(dst_hs_n)) & " of " &
               integer'image(to_integer(src_events)) &
               " -- the guarantee was never about the scheme, it was about the source being stallable";

        if errs = 0 then
            report "PASS: an event crossing a clock boundary is limited by something in every scheme, and simulation settles which -- a one-source-cycle pulse is lost as soon as the source period is shorter than the destination's, and adding a second synchroniser flop changes that not at all, because the two schemes' counters agree at every ratio and the only difference between them is a settling time no simulator models -- a toggle survives any pulse width and moves the limit to the EVENT RATE, losing events only when two of them fall inside one destination period, which they do at one event per source cycle -- and a four-phase handshake loses nothing at all, but only for a source that honours its busy signal, so its guarantee is conditional on a freedom an SPI slave does not have, because SCLK does not stop: no crossing scheme creates bandwidth, and the choice between them is a choice about which limit to accept";
        else
            report "FAIL: " & integer'image(errs) & " error(s)" severity error;
        end if;

        halt <= true;
        wait;
    end process;

end architecture;

6. Why a Verification Engineer Cares

Count what was generated and what arrived, separately, and in their own domains. The whole table rests on that: a bench that read the design's delivery counter to decide what was sent could not measure a loss. The generated count lives in the source domain and is read only after the run.

Sweep the ratio, and put the boundary inside the sweep. The interesting rows are the ones either side of "source period equals destination period", and a suite that runs at one ratio has measured nothing about a crossing.

Assert that two schemes AGREE where simulation cannot distinguish them. This is the technique worth stealing. Stating "schemes 1 and 2 must have equal counters" documents the limit of the evidence inside the regression, so the limit is visible to whoever reads the log rather than living in someone's memory.

Do not assert an exact survivor count below the limit. Which pulses happen to straddle a destination edge depends on the relationship between two clocks that have none, so the three language benches report 4 and 5 for the same design — for the same reason two boards would. The assertion is that events were lost, never how many.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// Properties for an event crossing. Written as SVA; the simulated benches use
// counters, because Icarus Verilog rejects SVA outright.

property p_strobe_one_cycle_pulse_schemes;
    // The rising-edge detectors cannot produce two adjacent deliveries: two rises
    // need a fall between them. Note this does NOT hold for the both-edge
    // detectors, where two adjacent strobes are two events correctly delivered --
    // a width check there fails on the design working.
    @(posedge clk) disable iff (!dst_rst_n)
        dst_level_stb |=> !dst_level_stb;
endproperty

property p_no_delivery_without_a_source_event;
    // No scheme may invent an event. Losing them is a bandwidth limit; gaining one
    // is a design fault, and a reader of the table must never have to wonder which
    // a number represents.
    @(posedge clk) disable iff (!dst_rst_n)
        dst_toggle_stb |-> (src_events > dst_toggle_n);
endproperty

property p_counters_monotone;
    // Delivery counts only increase. Weak, and true -- and about the limit of what
    // can be asserted about a crossing from inside it.
    @(posedge clk) disable iff (!dst_rst_n)
        1 |=> (dst_toggle_n >= $past(dst_toggle_n));
endproperty

property p_handshake_busy_blocks_a_second_event;
    // The handshake's contract, and the only one of the four schemes that HAS a
    // contract with its source: while busy, a new event must not advance the
    // request. A source that ignores this has built scheme 3 with extra logic.
    @(posedge sclk) disable iff (!src_rst_n)
        (src_event && src_busy) |=> $stable(hs_tog_src);
endproperty
Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// Coverage. The two axes are the RATIO and the EVENT SPACING, and they must be
// crossed -- because the pulse schemes are limited by the first and the toggle by
// the second, so a suite that varies only one has tested half the table.

covergroup cg_crossing @(posedge clk);
    option.per_instance = 1;

    // Source period as a multiple of the destination's, in tenths, so that the
    // fractional rows either side of 1.0 are distinguishable.
    ratio: coverpoint src_period_tenths {
        bins far_slower = {[40:$]};
        bins slower     = {[15:39]};
        bins equal      = {[9:14]};       // around 1.0
        bins faster     = {[4:8]};
        bins far_faster = {[1:3]};
    }

    // Source cycles between events. One is the hardest case for the toggle and the
    // normal case for a FIFO pointer.
    spacing: coverpoint event_gap_src_cycles {
        bins back_to_back = {1};
        bins two          = {2};
        bins few          = {[3:8]};
        bins sparse       = {[9:$]};
    }

    // Whether the source honoured the handshake's busy signal. The handshake's
    // guarantee is conditional on this, so a suite that never sets it to 0 has
    // verified a guarantee nobody can rely on.
    polite: coverpoint src_honours_busy { bins yes = {1}; bins no = {0}; }

    x_ratio_spacing: cross ratio, spacing;
    x_spacing_polite: cross spacing, polite;

endgroup

7. Why an FPGA or ASIC Engineer Cares

The two flops of a synchroniser must be kept as two flops, and the tool would rather not. A chain of flops with no logic between them is exactly what retiming and shift-register inference look for, and either will change the depth. On Xilinx that means ASYNC_REG = "TRUE" on both flops and SHREG_EXTRACT = "no"; on Intel, an ALTERA_ATTRIBUTE with -name SYNCHRONIZER_IDENTIFICATION. The attribute does two jobs — it stops the optimisation, and it tells timing analysis that the first flop's input is an asynchronous arrival rather than a path to report as failing forever.

The toggle scheme costs one flop in each domain and nothing else. That is worth knowing because it is often assumed to be expensive: the source's toggle register, the destination's synchroniser chain, and an XOR. Against the pulse scheme it costs one flop and buys immunity to pulse width.

The handshake costs two crossings, and the second one is easy to miss when reading a design. The acknowledgement travels back through its own synchroniser in the source's domain, so the round trip is SYNC_N cycles of each clock. A review that counts crossings by looking at the forward direction will count half of them.

The counters are the only wide logic here and they run at their own clock with a full cycle available. If a crossing fails timing, it is not the counter.

8. Failure Signature — A Slave That Misses Interrupts Under Load

The symptom, as it arrives:

"The slave's word_ready interrupt is occasionally missed. It happens more at high SPI clock rates and never on the bench setup. The word is in the buffer — it is the notification that vanishes."

What is happening: word_ready is a one-cycle pulse in the SCLK domain, crossed as a level (scheme 1 or 2). At the bench's SPI rate a source cycle is many destination cycles and the pulse is always seen. At the production rate the source cycle is shorter than the destination's sampling interval, and the pulse falls between two samples.

Why the data being present is the diagnostic clue: the word crossed correctly because it crossed as data held stable, and only the notification crossed as a pulse. A fault that loses the announcement and not the payload is almost always a pulse crossed as a level.

How to confirm in one change: convert the notification to a toggle. If the misses stop, that was it — and if they only reduce, the event rate is also at its limit and §4's second measurement is the one to read.

9. Common Misconceptions

"A two-flop synchroniser makes a signal safe to cross." It makes the value settled. It does not choose which of two values you get, it does not convey a pulse, and on a multi-bit bus it is wrong — which is Chapter 15.5's subject.

"Adding a second flop reduces the chance of losing a pulse." It changes nothing about pulse width. The measurement is in §3: schemes 1 and 2 are identical in every row.

"A toggle solves event crossing." It solves pulse width. Two toggles inside one destination sampling period cancel, and the measurement is 12 of 20.

"A handshake is always correct." It is correct for a source that honours its busy signal. An SPI slave cannot, because SCLK does not stop — and the same handshake against an impolite source delivered 4 of 20.

"If the regression passes at several clock ratios, the crossing is verified." Two of the five faults in Chapter 15.9 pass at every ratio, with every stimulus, in every simulator. Ratio sweeping is necessary and it is not sufficient.

"Simulation will show the difference between one flop and two if the stimulus is aggressive enough." There is no aggressive stimulus for metastability, because there is no metastability. This is the one place in the module where a green regression carries no information at all.

10. Reason It Through

Q. A source with a 6 ns period sends twenty events, two source cycles apart, to a destination with a 10 ns period. Schemes 1 and 2 deliver 12. Why 12 rather than 0 or 20?

Because a 6 ns pulse straddles a 10 ns sampling interval sometimes and not others. Whether a given pulse is seen depends on where it falls relative to the destination's edge, and with two clocks that have no fixed relationship that phase walks through the period — so some are caught and some are not. The fraction is not a property of the design; it is a property of the two periods, which is why the bench asserts that events were lost rather than how many.

Q. Why can the source's event count not simply be crossed to the destination so that a single counter compares them?

Because crossing a multi-bit count is the problem the block exists to study. A count carried across by N independent synchronisers would be Chapter 15.5's scheme 1 and would report values the source never held; carried by a handshake it would be limited by the handshake. Either way the instrument would share the fault it is measuring. The count stays in its own domain and is read once, at the end.

Q. The handshake loses half its events at a source period ten times the destination's. Derive why.

The round trip is SYNC_N destination cycles for the request to arrive plus SYNC_N source cycles for the acknowledgement to return — the ack crosses back and needs its own synchroniser, clocked by the source. At SYNC_N = 2 that is at least two source cycles, and the stimulus sends an event every two source cycles. So the ack arrives exactly as the next event does, src_busy is still asserted when the source looks, and every second event is refused. Slowing the source does not help, because the source-side half of the round trip scales with the source period.

Q. A design crosses a one-cycle pulse and its author argues that SCLK will always be at least ten times slower than the system clock, so the pulse is always seen. Is the argument sound, and is the design acceptable?

The argument is sound and the design is fragile. At a ten-times ratio the pulse spans ten destination cycles and cannot be missed. What makes it fragile is that the correctness depends on a ratio that nothing enforces: a later product variant with a faster SPI interface or a slower system clock breaks it silently, and the failure is a missed notification with the data intact — §8's signature, which is expensive to diagnose. A toggle costs one flop and removes the dependency, which is a better trade than a comment.

Q. Why is it correct for the both-edge detectors to produce two strobes on consecutive destination cycles, when the rising-edge detectors cannot?

Because a toggle carries one event per transition in either direction, so two events in quick succession genuinely produce two adjacent changes in the synchronised level, and two adjacent strobes are two events correctly delivered. A rising-edge detector needs a fall between two rises, so two adjacent deliveries would be impossible and a width check there is a valid invariant. The design's assertions check the width only for the pulse schemes, and the first version checked all four and fired on the design working.

11. Understanding Check

12. Summary

A master owns its clock and has no crossing. A slave is handed one, and Chapter 14.1 answered the resulting question before asking it.

There are four ways to carry an event across the boundary. A level through one flop and a level through several are both limited by pulse width, and they are identical in simulation. A toggle has no width and is limited by event rate. A handshake has no limit and requires the source to stall.

The measurements: the pulse schemes go 20, 20, 20, 20, 12, 4 as the source speeds up; the toggle holds 20 throughout and then fails at 12 of 20 when events come every source cycle; the handshake delivers 20 of 20 to a polite source at the fastest ratio and 4 of 20 to an impolite one.

So no crossing scheme creates bandwidth, and the choice between them is a choice about which limit to accept.

The most important row in the table is the one where nothing happens: schemes 1 and 2 never disagree, because their only difference is a settling time and no simulator models one. The module's central rule is the one rule its benches cannot check — which is why Chapter 15.3 separates the simulatable limit from the unsimulatable one and Chapter 15.9 shows five broken crossings of which the two worst pass everything.

And the handshake's precondition is a freedom an SPI slave lacks, because SCLK does not stop — which is why Chapter 15.6 has to size a FIFO.

For verification: keep the generated count in the source's own domain; put the ratio boundary inside the sweep; assert that two schemes agree where simulation cannot tell them apart, so the limit of the evidence is in the log; and never assert an exact survivor count below the limit.

For implementation: ASYNC_REG and an anti-shift-register attribute on every synchroniser, because the tool would rather make it one flop; a toggle costs one flop per domain; and a handshake is two crossings, the second of which a forward-looking review will miss.

13. What Comes Next

The crossing is understood as a problem. Neither of the two architectures that solve it has been built.

Chapter 15.2 — Architecture A: SCLK as a Clock Domain clocks the shift path on SCLK itself, which removes the ratio precondition entirely and works where oversampling cannot — at SCLK faster than the system clock. It replaces that precondition with a different one, per word rather than per bit, and it costs three things that are easy to discover late: the mode becomes a synthesis-time parameter, the hand-off register has to outlive the transaction that filled it, and the domain needs a reset edge that chip select alone does not supply.

Continue learning