Skip to content
VLSI Mentor

SPI · Module 15

Synchronizers and Their Limits

A two-flop synchroniser gives a settled value and nothing else — not which value, not a pulse, and not a bus. Measured with the per-bit skew routing produces, independent synchronisers on an 8-bit counter observe values the source never held, and are still wrong when the source is slowed. Gray coding observes zero; data-plus-a-flag fails too unless the source holds still.

A two-flop synchroniser solves exactly one problem: it takes a signal that may be sampled while it is changing and produces a value that has finished settling before anything downstream looks at it.

That is all. In particular:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   IT DOES NOT  guarantee WHICH of the two values you get. The first flop may
                resolve to the old one, which is why Chapter 15.3's soft limit exists.

   IT DOES NOT  convey a pulse. A pulse narrower than a destination clock period is
                lost with two flops exactly as with one -- Chapter 15.1 measured it.

   IT DOES NOT  work on a BUS.

The third is the one that gets built anyway.

An N-bit value has to cross a clock boundary. Why is one synchroniser per bit wrong, and what is right instead?

1. Why N Independent Synchronisers Is Wrong

Each bit settles independently. The destination therefore samples a mixture: some bits from before the source's update and some from after. The result is a value the source never held.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   a counter going 0111 -> 1000 changes four bits at once
   catch two of them and the destination reads 1111 or 0000
   both plausible, neither ever true

And the damage is unbounded: an address that briefly reads 1111 addresses the wrong memory, and no downstream check can tell, because the value is well formed. There is no bit pattern a consumer could reject.

2. What Makes This Simulatable, Which Is The Interesting Part

Bit-level settling differences are metastability and are not modelled. But the same incoherence is produced by skew — the source bits arriving at the destination's flops at slightly different times, which is what unequal routing does on any real device and is entirely ordinary.

So the measurement drives the crossing with per-bit skew, and that reproduces the failure exactly, for a reason a reviewer can point at on a floorplan.

A design that is safe against skew is safe against the settling differences too, because both amount to "the bits do not arrive together".

A bench that drives the vector with all bits changing at the same instant proves nothing at all — and is what almost every bench does.

3. The Three Correct Answers, And When Each Applies

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   GRAY CODE         for a value that changes by ONE STEP at a time -- a counter or a
                     FIFO pointer. Consecutive Gray codes differ in exactly one bit,
                     so a destination catching an update part-way reads either the old
                     value or the new one, never a third.
                     Does NOT work for arbitrary data, which does not change by one step.

   DATA PLUS A FLAG  for arbitrary data. The data crosses with NO synchroniser at all
                     and is simply held stable; one flag bit crosses through a
                     synchroniser and says when the data is safe to sample.
                     This is what Chapter 15.2 does, and the requirement it creates is
                     a LIFETIME on the data register.

   A FIFO            for arbitrary data that keeps coming. Chapter 15.6.

4. The Block Diagram

A source counter feeding three crossings: a binary vector through independent synchronisers, the same vector Gray coded through identical synchronisers then decoded, and the raw vector sampled once when a synchronised flag arrivessource counterbinary, as isGray encodetoggle a flagW synchronisersW synchronisersONE synchroniservalues never heldGray decodesample the datawhen stable12
Figure 1 — three crossings of one source value. Schemes 1 and 2 use the identical synchroniser structure and differ only in what the source encoded, which is the point: the synchronisers are no better in scheme 2, but consecutive Gray codes differ in one bit so a partial observation is the old value or the new one. Scheme 3's data path has no synchroniser at all, deliberately — synchronising it would reintroduce scheme 1's problem.
Ten cycles across five rows. A destination clock runs throughout. The source value changes from 7 to 8 at cycle 4, with the bits arriving over cycles 4 to 6. A sampling instant at cycle 5 reads the mixture 15. The Gray-coded row changes by a single bit and the sample reads either 7 or 8.four bits change at oncefour bits change at oncebinary reads 15: never heldbinary reads 15: never heldGray reads 7 or 8, never 15Gray reads 7 or 8, never 15dst_clksrc (binary)7777??8888bits arrivingsampled binary777771515888sampled Gray7777778888t0t1t2t3t4t5t6t7t8t9
Figure 2 — one update of a four-bit counter, with the top bits arriving late. The source moves from 7 to 8, which changes every bit; the destination's sampling instant falls inside the skew window and reads a mixture. Gray coding removes the possibility because consecutive codes differ in one bit, so a partial observation is 7 or 8 and nothing else.

5. The Measurement

An 8-bit counter crossing two unrelated clocks — 14 ns source, 10 ns destination — with up to 3 ns of per-bit skew on the seam. An observation is impossible if the source did not hold it within the last six source cycles.

Phase 1 — the source never holds still

One update per source cycle, which is the hardest case for every scheme and the normal case for a FIFO pointer:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   scheme                        observations  impossible values
   1  N synchronisers on a bus           1138                 13
   2  Gray code, then decoded            1138                  0
   3  data plus a flag                    400                 13

Scheme 1 produced 12 distinct impossible values out of the 256 the counter can hold — so it is not one unlucky alignment. And every one was a legal 8-bit pattern, so no downstream check could reject one.

Scheme 3 failing here was not the expected result and is the more useful one. Its precondition is that the source holds the data still until the flag has been observed — and a source updating every cycle violates that, so the destination samples a register that is mid-change and reads a mixture, exactly as scheme 1 does.

The flag is not the weak part. The holding is.

Phase 2 — the source holds still

One update every eight source cycles:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   scheme                        observations  impossible values
   1  N synchronisers on a bus            775                  9
   2  Gray code, then decoded             775                  0
   3  data plus a flag                     60                  0

Gray is zero again. Data-plus-a-flag is now exact, because its precondition is met.

And scheme 1 is still wrong — 9 times in 775 observations against 13 in 1138. Slowing the source reduced the rate and not the fault.

The exact counts are phase- and simulator-dependent and are deliberately not asserted: which update a destination edge lands inside depends on two clocks with no relationship, so the three language benches report different totals from the same design. What is asserted is the shape — scheme 1 non-zero in both regimes, Gray zero in both, data-plus-a-flag non-zero only when the source never holds still — and that is identical in all three, because it follows from the encoding rather than from the alignment.

6. Building the Three Crossings — Three HDLs

The circuit

One module, three crossings of the same source value, with the skew seam brought out as a port pair.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_cdc_vector.sv — three crossings of one value, with the skew seam brought out as a port pair
// spi_cdc_vector.sv
//
// Chapter 15.5 -- what a two-flop synchroniser is for, and the thing it is
// routinely used for that it cannot do.
//
// A two-flop synchroniser solves exactly one problem: it takes a signal that may
// be sampled while it is changing and produces a value that has finished settling
// before anything downstream looks at it. That is all. In particular:
//
//   IT DOES NOT guarantee WHICH of the two values you get. The first flop may
//   resolve to the old one, which is why Chapter 15.3's soft limit exists.
//
//   IT DOES NOT convey a pulse. A pulse narrower than a destination clock period
//   is lost with two flops exactly as it is with one (Chapter 15.1 measured this).
//
//   IT DOES NOT work on a BUS. This is the one that gets built anyway, and this
//   file exists to make the failure visible rather than arguable.
//
// WHY N INDEPENDENT SYNCHRONISERS ON AN N-BIT VALUE IS WRONG.
//
// Each bit settles independently. The destination therefore samples a MIXTURE: some
// bits from before the source's update and some from after. The result is a value
// the source NEVER HELD.
//
// A counter going from 0111 to 1000 changes four bits at once. If the destination
// catches two of them, it reads 1111 or 0000 -- both plausible, neither ever true.
// And the damage is unbounded: an address that briefly reads 1111 addresses the
// wrong memory, and no downstream check can tell because the value is well formed.
//
// THE THREE CORRECT ANSWERS, AND WHEN EACH APPLIES.
//
//   GRAY CODE          for a value that changes by ONE STEP at a time, such as a
//                      counter or a FIFO pointer. Consecutive Gray codes differ in
//                      exactly one bit, so a destination that catches the update
//                      part-way reads either the old value or the new one -- never
//                      a third. It does NOT work for arbitrary data, because
//                      arbitrary data does not change by one step.
//
//   DATA PLUS A FLAG   for arbitrary data. The data crosses with NO synchroniser
//                      at all and is simply held stable; one flag bit crosses
//                      through a synchroniser and says when the data is safe to
//                      sample. This is what Chapter 15.2 does, and the requirement
//                      it creates is a LIFETIME on the data register.
//
//   A FIFO             for arbitrary data that keeps coming. Chapter 15.6.
//
// This module builds the wrong scheme and two of the right ones side by side, on
// the same source value, so that the testbench can count how many times each
// destination observed a value the source never held.
//
// WHAT MAKES THE FAILURE SIMULATABLE, WHICH IS THE INTERESTING PART.
//
// Bit-level settling differences are metastability and are not modelled. But the
// same incoherence is produced by SKEW -- the source bits arriving at the
// destination's flops at slightly different times, which is what unequal routing
// does on any real device and is entirely ordinary. So the testbench drives the
// source vector with a small per-bit skew, and that reproduces the failure exactly,
// for a reason a reviewer can point at on a floorplan.
//
// A design that is safe against skew is safe against the settling differences too,
// because both amount to "the bits do not arrive together". A bench that drives the
// vector with all bits changing at the same instant proves nothing at all, and is
// what almost every bench does.

module spi_cdc_vector #(
    parameter int W      = 4,    // the width of the value being crossed
    parameter int SYNC_N = 2
) (
    input  wire              src_clk,
    input  wire              src_rst_n,
    input  wire [W-1:0]      src_bin,     // the value, binary, as the source holds it
    input  wire [W-1:0]      src_gray,    // the same value, Gray-coded
    input  wire              src_update,  // one source cycle: the value just changed

    input  wire              dst_clk,
    input  wire              dst_rst_n,

    // --- scheme 1: N INDEPENDENT SYNCHRONISERS ON A BUS. Wrong, and built here
    //     so that the count of impossible values is a measurement.
    output wire [W-1:0]      dst_bin_wrong,

    // --- scheme 2: GRAY CODE through the same N synchronisers, then decoded.
    output wire [W-1:0]      dst_gray_ok,

    // --- scheme 3: DATA PLUS A FLAG. The data is not synchronised at all.
    output reg  [W-1:0]      dst_flagged,
    output reg               dst_flagged_stb
);

    // ---------------------------------------------------------------- scheme 1
    // One synchroniser per bit, which is the natural thing to write and is the
    // bug. Nothing here is wrong on its own; the error is applying it to a value
    // whose bits have to agree with each other.
    reg [W-1:0] w_sr [0:SYNC_N-1];
    integer i1;
    always_ff @(posedge dst_clk or negedge dst_rst_n) begin
        if (!dst_rst_n) begin
            for (i1 = 0; i1 < SYNC_N; i1 = i1 + 1)
                w_sr[i1] <= {W{1'b0}};
        end else begin
            w_sr[0] <= src_bin;
            for (i1 = 1; i1 < SYNC_N; i1 = i1 + 1)
                w_sr[i1] <= w_sr[i1-1];
        end
    end
    assign dst_bin_wrong = w_sr[SYNC_N-1];

    // ---------------------------------------------------------------- scheme 2
    // The identical structure, on a Gray-coded value. The synchronisers are no
    // better; what changed is that consecutive values differ in ONE bit, so
    // catching an update part-way yields the old value or the new one.
    reg [W-1:0] g_sr [0:SYNC_N-1];
    integer i2;
    always_ff @(posedge dst_clk or negedge dst_rst_n) begin
        if (!dst_rst_n) begin
            for (i2 = 0; i2 < SYNC_N; i2 = i2 + 1)
                g_sr[i2] <= {W{1'b0}};
        end else begin
            g_sr[0] <= src_gray;
            for (i2 = 1; i2 < SYNC_N; i2 = i2 + 1)
                g_sr[i2] <= g_sr[i2-1];
        end
    end

    // Gray to binary: each output bit is the XOR of all Gray bits at or above it.
    // Combinational and W-1 XORs deep, which is the whole cost of the scheme.
    function automatic [W-1:0] gray_to_bin(input [W-1:0] g);
        integer k;
        begin
            gray_to_bin[W-1] = g[W-1];
            for (k = W-2; k >= 0; k = k - 1)
                gray_to_bin[k] = gray_to_bin[k+1] ^ g[k];
        end
    endfunction

    assign dst_gray_ok = gray_to_bin(g_sr[SYNC_N-1]);

    // ---------------------------------------------------------------- scheme 3
    // Data plus a flag. The data path has NO synchroniser -- deliberately, and it
    // is the point: synchronising it would reintroduce scheme 1's problem. What
    // makes it safe is that the source holds it still, and the flag says when.
    reg              flag_src;
    always_ff @(posedge src_clk or negedge src_rst_n) begin
        if (!src_rst_n) flag_src <= 1'b0;
        else if (src_update) flag_src <= ~flag_src;
    end

    reg [SYNC_N-1:0] f_sr;
    reg              f_d;
    wire             f_arrived = f_sr[SYNC_N-1] ^ f_d;

    always_ff @(posedge dst_clk or negedge dst_rst_n) begin
        if (!dst_rst_n) begin
            f_sr            <= {SYNC_N{1'b0}};
            f_d             <= 1'b0;
            dst_flagged     <= {W{1'b0}};
            dst_flagged_stb <= 1'b0;
        end else begin
            f_sr <= {f_sr[SYNC_N-2:0], flag_src};
            f_d  <= f_sr[SYNC_N-1];
            dst_flagged_stb <= f_arrived;
            if (f_arrived)
                dst_flagged <= src_bin;   // sampled once, unsynchronised, on trust
        end
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_cdc_vector.v — the same design in Verilog-2001
// spi_cdc_vector.v
//
// Chapter 15.5 -- what a two-flop synchroniser is for, and the thing it is
// routinely used for that it cannot do.
//
// A two-flop synchroniser solves exactly one problem: it takes a signal that may
// be sampled while it is changing and produces a value that has finished settling
// before anything downstream looks at it. That is all. In particular:
//
//   IT DOES NOT guarantee WHICH of the two values you get. The first flop may
//   resolve to the old one, which is why Chapter 15.3's soft limit exists.
//
//   IT DOES NOT convey a pulse. A pulse narrower than a destination clock period
//   is lost with two flops exactly as it is with one (Chapter 15.1 measured this).
//
//   IT DOES NOT work on a BUS. This is the one that gets built anyway, and this
//   file exists to make the failure visible rather than arguable.
//
// WHY N INDEPENDENT SYNCHRONISERS ON AN N-BIT VALUE IS WRONG.
//
// Each bit settles independently. The destination therefore samples a MIXTURE: some
// bits from before the source's update and some from after. The result is a value
// the source NEVER HELD.
//
// A counter going from 0111 to 1000 changes four bits at once. If the destination
// catches two of them, it reads 1111 or 0000 -- both plausible, neither ever true.
// And the damage is unbounded: an address that briefly reads 1111 addresses the
// wrong memory, and no downstream check can tell because the value is well formed.
//
// THE THREE CORRECT ANSWERS, AND WHEN EACH APPLIES.
//
//   GRAY CODE          for a value that changes by ONE STEP at a time, such as a
//                      counter or a FIFO pointer. Consecutive Gray codes differ in
//                      exactly one bit, so a destination that catches the update
//                      part-way reads either the old value or the new one -- never
//                      a third. It does NOT work for arbitrary data, because
//                      arbitrary data does not change by one step.
//
//   DATA PLUS A FLAG   for arbitrary data. The data crosses with NO synchroniser
//                      at all and is simply held stable; one flag bit crosses
//                      through a synchroniser and says when the data is safe to
//                      sample. This is what Chapter 15.2 does, and the requirement
//                      it creates is a LIFETIME on the data register.
//
//   A FIFO             for arbitrary data that keeps coming. Chapter 15.6.
//
// This module builds the wrong scheme and two of the right ones side by side, on
// the same source value, so that the testbench can count how many times each
// destination observed a value the source never held.
//
// WHAT MAKES THE FAILURE SIMULATABLE, WHICH IS THE INTERESTING PART.
//
// Bit-level settling differences are metastability and are not modelled. But the
// same incoherence is produced by SKEW -- the source bits arriving at the
// destination's flops at slightly different times, which is what unequal routing
// does on any real device and is entirely ordinary. So the testbench drives the
// source vector with a small per-bit skew, and that reproduces the failure exactly,
// for a reason a reviewer can point at on a floorplan.
//
// A design that is safe against skew is safe against the settling differences too,
// because both amount to "the bits do not arrive together". A bench that drives the
// vector with all bits changing at the same instant proves nothing at all, and is
// what almost every bench does.

module spi_cdc_vector #(
    parameter W      = 4,    // the width of the value being crossed
    parameter SYNC_N = 2
) (
    input  wire              src_clk,
    input  wire              src_rst_n,
    input  wire [W-1:0]      src_bin,     // the value, binary, as the source holds it
    input  wire [W-1:0]      src_gray,    // the same value, Gray-coded
    input  wire              src_update,  // one source cycle: the value just changed

    input  wire              dst_clk,
    input  wire              dst_rst_n,

    // --- scheme 1: N INDEPENDENT SYNCHRONISERS ON A BUS. Wrong, and built here
    //     so that the count of impossible values is a measurement.
    output wire [W-1:0]      dst_bin_wrong,

    // --- scheme 2: GRAY CODE through the same N synchronisers, then decoded.
    output wire [W-1:0]      dst_gray_ok,

    // --- scheme 3: DATA PLUS A FLAG. The data is not synchronised at all.
    output reg  [W-1:0]      dst_flagged,
    output reg               dst_flagged_stb
);

    // ---------------------------------------------------------------- scheme 1
    // One synchroniser per bit, which is the natural thing to write and is the
    // bug. Nothing here is wrong on its own; the error is applying it to a value
    // whose bits have to agree with each other.
    reg [W-1:0] w_sr [0:SYNC_N-1];
    integer i1;
    always @(posedge dst_clk or negedge dst_rst_n) begin
        if (!dst_rst_n) begin
            for (i1 = 0; i1 < SYNC_N; i1 = i1 + 1)
                w_sr[i1] <= {W{1'b0}};
        end else begin
            w_sr[0] <= src_bin;
            for (i1 = 1; i1 < SYNC_N; i1 = i1 + 1)
                w_sr[i1] <= w_sr[i1-1];
        end
    end
    assign dst_bin_wrong = w_sr[SYNC_N-1];

    // ---------------------------------------------------------------- scheme 2
    // The identical structure, on a Gray-coded value. The synchronisers are no
    // better; what changed is that consecutive values differ in ONE bit, so
    // catching an update part-way yields the old value or the new one.
    reg [W-1:0] g_sr [0:SYNC_N-1];
    integer i2;
    always @(posedge dst_clk or negedge dst_rst_n) begin
        if (!dst_rst_n) begin
            for (i2 = 0; i2 < SYNC_N; i2 = i2 + 1)
                g_sr[i2] <= {W{1'b0}};
        end else begin
            g_sr[0] <= src_gray;
            for (i2 = 1; i2 < SYNC_N; i2 = i2 + 1)
                g_sr[i2] <= g_sr[i2-1];
        end
    end

    // Gray to binary: each output bit is the XOR of all Gray bits at or above it.
    // Combinational and W-1 XORs deep, which is the whole cost of the scheme.
        function [W-1:0] gray_to_bin;
        input [W-1:0] g;
        integer k;
        begin
            gray_to_bin[W-1] = g[W-1];
            for (k = W-2; k >= 0; k = k - 1)
                gray_to_bin[k] = gray_to_bin[k+1] ^ g[k];
        end
    endfunction

    assign dst_gray_ok = gray_to_bin(g_sr[SYNC_N-1]);

    // ---------------------------------------------------------------- scheme 3
    // Data plus a flag. The data path has NO synchroniser -- deliberately, and it
    // is the point: synchronising it would reintroduce scheme 1's problem. What
    // makes it safe is that the source holds it still, and the flag says when.
    reg              flag_src;
    always @(posedge src_clk or negedge src_rst_n) begin
        if (!src_rst_n) flag_src <= 1'b0;
        else if (src_update) flag_src <= ~flag_src;
    end

    reg [SYNC_N-1:0] f_sr;
    reg              f_d;
    wire             f_arrived = f_sr[SYNC_N-1] ^ f_d;

    always @(posedge dst_clk or negedge dst_rst_n) begin
        if (!dst_rst_n) begin
            f_sr            <= {SYNC_N{1'b0}};
            f_d             <= 1'b0;
            dst_flagged     <= {W{1'b0}};
            dst_flagged_stb <= 1'b0;
        end else begin
            f_sr <= {f_sr[SYNC_N-2:0], flag_src};
            f_d  <= f_sr[SYNC_N-1];
            dst_flagged_stb <= f_arrived;
            if (f_arrived)
                dst_flagged <= src_bin;   // sampled once, unsynchronised, on trust
        end
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_cdc_vector.vhd — the same design in VHDL
-- spi_cdc_vector.vhd
--
-- Chapter 15.5 -- what a two-flop synchroniser is for, and the thing it is
-- routinely used for that it cannot do.
--
-- A two-flop synchroniser solves exactly one problem: it takes a signal that may
-- be sampled while it is changing and produces a value that has finished settling
-- before anything downstream looks at it. That is all. In particular:
--
--   IT DOES NOT guarantee WHICH of the two values you get. The first flop may
--   resolve to the old one, which is why Chapter 15.3's soft limit exists.
--
--   IT DOES NOT convey a pulse. A pulse narrower than a destination clock period
--   is lost with two flops exactly as it is with one (Chapter 15.1 measured this).
--
--   IT DOES NOT work on a BUS. This is the one that gets built anyway, and this
--   file exists to make the failure visible rather than arguable.
--
-- WHY N INDEPENDENT SYNCHRONISERS ON AN N-BIT VALUE IS WRONG.
--
-- Each bit settles independently. The destination therefore samples a MIXTURE: some
-- bits from before the source's update and some from after. The result is a value
-- the source NEVER HELD.
--
-- A counter going from 0111 to 1000 changes four bits at once. If the destination
-- catches two of them, it reads 1111 or 0000 -- both plausible, neither ever true.
-- And the damage is unbounded: an address that briefly reads 1111 addresses the
-- wrong memory, and no downstream check can tell because the value is well formed.
--
-- THE THREE CORRECT ANSWERS, AND WHEN EACH APPLIES.
--
--   GRAY CODE          for a value that changes by ONE STEP at a time, such as a
--                      counter or a FIFO pointer. Consecutive Gray codes differ in
--                      exactly one bit, so a destination that catches the update
--                      part-way reads either the old value or the new one -- never
--                      a third. It does NOT work for arbitrary data, because
--                      arbitrary data does not change by one step.
--
--   DATA PLUS A FLAG   for arbitrary data. The data crosses with NO synchroniser
--                      at all and is simply held stable; one flag bit crosses
--                      through a synchroniser and says when the data is safe to
--                      sample. This is what Chapter 15.2 does, and the requirement
--                      it creates is a LIFETIME on the data register.
--
--   A FIFO             for arbitrary data that keeps coming. Chapter 15.6.
--
-- This module builds the wrong scheme and two of the right ones side by side, on
-- the same source value, so that the testbench can count how many times each
-- destination observed a value the source never held.
--
-- WHAT MAKES THE FAILURE SIMULATABLE, WHICH IS THE INTERESTING PART.
--
-- Bit-level settling differences are metastability and are not modelled. But the
-- same incoherence is produced by SKEW -- the source bits arriving at the
-- destination's flops at slightly different times, which is what unequal routing
-- does on any real device and is entirely ordinary. So the testbench drives the
-- source vector with a small per-bit skew, and that reproduces the failure exactly,
-- for a reason a reviewer can point at on a floorplan.
--
-- A design that is safe against skew is safe against the settling differences too,
-- because both amount to "the bits do not arrive together". A bench that drives the
-- vector with all bits changing at the same instant proves nothing at all, and is
-- what almost every bench does.

library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

entity spi_cdc_vector is
    generic (
        W      : positive := 8;   -- the width of the value being crossed
        SYNC_N : positive := 2
    );
    port (
        src_clk         : in  std_logic;
        src_rst_n       : in  std_logic;
        src_bin         : in  std_logic_vector(W - 1 downto 0);
        src_gray        : in  std_logic_vector(W - 1 downto 0);
        src_update      : in  std_logic;

        dst_clk         : in  std_logic;
        dst_rst_n       : in  std_logic;

        -- scheme 1: N INDEPENDENT SYNCHRONISERS ON A BUS. Wrong, and built here
        -- so that the count of impossible values is a measurement.
        dst_bin_wrong   : out std_logic_vector(W - 1 downto 0);
        -- scheme 2: GRAY CODE through the same N synchronisers, then decoded.
        dst_gray_ok     : out std_logic_vector(W - 1 downto 0);
        -- scheme 3: DATA PLUS A FLAG. The data is not synchronised at all.
        dst_flagged     : out std_logic_vector(W - 1 downto 0);
        dst_flagged_stb : out std_logic
    );
end entity;

architecture rtl of spi_cdc_vector is

    type vec_arr is array (0 to SYNC_N - 1) of std_logic_vector(W - 1 downto 0);

    signal w_sr : vec_arr := (others => (others => '0'));
    signal g_sr : vec_arr := (others => (others => '0'));

    signal flag_src  : std_logic := '0';
    signal f_sr      : std_logic_vector(SYNC_N - 1 downto 0) := (others => '0');
    signal f_d       : std_logic := '0';
    signal f_arrived : std_logic;

    signal flagged_r : std_logic_vector(W - 1 downto 0) := (others => '0');
    signal flag_stb_r : std_logic := '0';

    -- Gray to binary: each output bit is the XOR of all Gray bits at or above it.
    -- Combinational and W-1 XORs deep, which is the whole cost of the scheme.
    function gray_to_bin(g : std_logic_vector) return std_logic_vector is
        variable r : std_logic_vector(g'range);
    begin
        r(r'high) := g(g'high);
        for k in g'high - 1 downto 0 loop
            r(k) := r(k + 1) xor g(k);
        end loop;
        return r;
    end function;

begin

    -- ---------------------------------------------------------------- scheme 1
    -- One synchroniser per bit, which is the natural thing to write and is the bug.
    -- Nothing here is wrong on its own; the error is applying it to a value whose
    -- bits have to agree with each other.
    wrong : process (dst_clk, dst_rst_n)
    begin
        if dst_rst_n = '0' then
            w_sr <= (others => (others => '0'));
        elsif rising_edge(dst_clk) then
            w_sr(0) <= src_bin;
            for i in 1 to SYNC_N - 1 loop
                w_sr(i) <= w_sr(i - 1);
            end loop;
        end if;
    end process;
    dst_bin_wrong <= w_sr(SYNC_N - 1);

    -- ---------------------------------------------------------------- scheme 2
    -- The identical structure, on a Gray-coded value. The synchronisers are no
    -- better; what changed is that consecutive values differ in ONE bit.
    grayx : process (dst_clk, dst_rst_n)
    begin
        if dst_rst_n = '0' then
            g_sr <= (others => (others => '0'));
        elsif rising_edge(dst_clk) then
            g_sr(0) <= src_gray;
            for i in 1 to SYNC_N - 1 loop
                g_sr(i) <= g_sr(i - 1);
            end loop;
        end if;
    end process;
    dst_gray_ok <= gray_to_bin(g_sr(SYNC_N - 1));

    -- ---------------------------------------------------------------- scheme 3
    -- Data plus a flag. The data path has NO synchroniser -- deliberately, and it
    -- is the point: synchronising it would reintroduce scheme 1's problem. What
    -- makes it safe is that the source holds it still, and the flag says when.
    srcflag : process (src_clk, src_rst_n)
    begin
        if src_rst_n = '0' then
            flag_src <= '0';
        elsif rising_edge(src_clk) then
            if src_update = '1' then
                flag_src <= not flag_src;
            end if;
        end if;
    end process;

    f_arrived       <= f_sr(SYNC_N - 1) xor f_d;
    dst_flagged     <= flagged_r;
    dst_flagged_stb <= flag_stb_r;

    dstflag : process (dst_clk, dst_rst_n)
    begin
        if dst_rst_n = '0' then
            f_sr       <= (others => '0');
            f_d        <= '0';
            flagged_r  <= (others => '0');
            flag_stb_r <= '0';
        elsif rising_edge(dst_clk) then
            f_sr       <= f_sr(SYNC_N - 2 downto 0) & flag_src;
            f_d        <= f_sr(SYNC_N - 1);
            flag_stb_r <= f_arrived;
            if f_arrived = '1' then
                flagged_r <= src_bin;   -- sampled once, unsynchronised, on trust
            end if;
        end if;
    end process;

end architecture;

The testbench

Two phases, one per regime, with the recency window as the impossibility test and per-bit skew on the seam.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_cdc_vector_tb.sv — two regimes, a recency window, and an assertion about the shape rather than the counts
// spi_cdc_vector_tb.sv
//
// The measurement is "how many times did the destination observe a value the
// source never held?", and it is made three times on the same stimulus.
//
// HOW THE BENCH KNOWS WHAT THE SOURCE HELD, and why the obvious version of this
// does not work. The first attempt kept a bitmap of every value the source had EVER
// taken -- and a counter passes through every value, so within sixteen cycles the
// bitmap was all ones and no observation could ever be impossible. The bench
// reported zero faults on a scheme that is definitely broken, which is a good
// reminder that a checker returning the answer you hoped for is not evidence.
//
// The correct question is about RECENCY, not existence: the destination is reading a
// value through a synchroniser, so it must see one the source held WITHIN THE LAST
// FEW SOURCE CYCLES. The bench keeps a short history of the source's recent values
// and any observation outside that window is impossible -- exactly, not
// probabilistically, because the source cannot have held it.
//
// The counter is eight bits wide so that the 256-value space is far larger than the
// six-deep window; a mixture of two adjacent values then lands outside the window
// almost always rather than sometimes.
//
// HOW THE FAILURE IS MADE TO HAPPEN. The source bits are driven to the DUT through
// per-bit skew -- bit k is delayed by k times a small delta -- which is what
// unequal routing does on any real device. That is the honest way to exhibit
// incoherence in a simulator: bit-level settling differences are metastability and
// are not modelled, but skew is ordinary, it produces the identical failure, and a
// design that survives skew survives the settling case for the same reason.
//
// A bench that changes all bits at the same instant finds nothing wrong with scheme
// 1, which is why scheme 1 keeps being built.

`timescale 1ns/1ps

module spi_cdc_vector_tb;

    localparam int W      = 8;
    localparam int SYNC_N = 2;
    localparam int NVAL   = 1 << W;
    // How far back a legitimate observation may come from: the synchroniser is
    // SYNC_N deep and the two clocks are close in period, so a value three or four
    // source cycles old is plausible. Six is generous, which matters -- a window
    // that is too tight reports correct designs as broken.
    localparam int HIST   = 6;

    // Two unrelated clocks. The ratio is deliberately not an integer, so that the
    // destination's sampling instant walks through the source's period rather than
    // sitting at one point in it.
    reg src_clk = 1'b0;
    reg dst_clk = 1'b0;
    always #7  src_clk = ~src_clk;    // 14 ns
    always #5  dst_clk = ~dst_clk;    // 10 ns

    reg src_rst_n = 1'b1;
    reg dst_rst_n = 1'b1;

    // The source: a plain counter, and its Gray encoding.
    reg [W-1:0] cnt = {W{1'b0}};
    reg         update = 1'b0;

    wire [W-1:0] cnt_gray = cnt ^ (cnt >> 1);

    // --- the skew injection -------------------------------------------------
    // Each bit of the value reaches the DUT through its own delay. This is the
    // whole reason the bench can see the bug, and it is a model of routing rather
    // than of metastability -- which is what makes it defensible.
    localparam time SKEW = 1;
    wire [W-1:0] bin_skewed;
    wire [W-1:0] gray_skewed;
    genvar gi;
    generate
        for (gi = 0; gi < W; gi = gi + 1) begin : g_skew
            // Bit 0 arrives on time, bit W-1 arrives (W-1)*SKEW late. A real
            // device's spread is smaller and less orderly; an orderly spread makes
            // the result reproducible, which a chapter needs.
            assign #(gi * SKEW) bin_skewed[gi]  = cnt[gi];
            assign #(gi * SKEW) gray_skewed[gi] = cnt_gray[gi];
        end
    endgenerate

    wire [W-1:0] dst_bin_wrong, dst_gray_ok, dst_flagged;
    wire         dst_flagged_stb;

    spi_cdc_vector #(.W(W), .SYNC_N(SYNC_N)) dut (
        .src_clk(src_clk), .src_rst_n(src_rst_n),
        .src_bin(bin_skewed), .src_gray(gray_skewed), .src_update(update),
        .dst_clk(dst_clk), .dst_rst_n(dst_rst_n),
        .dst_bin_wrong(dst_bin_wrong), .dst_gray_ok(dst_gray_ok),
        .dst_flagged(dst_flagged), .dst_flagged_stb(dst_flagged_stb)
    );

    integer errors = 0;

    initial begin
        #2_000_000;
        $display("FAIL: the simulation did not finish within its time limit");
        $finish;
    end

    // The source's recent values, newest first. Maintained in SOURCE time, which is
    // the only place the source's value means anything.
    reg [W-1:0] hist [0:HIST-1];

    function automatic is_recent(input [W-1:0] v);
        integer j;
        begin
            is_recent = 1'b0;
            for (j = 0; j < HIST; j = j + 1)
                if (hist[j] == v) is_recent = 1'b1;
        end
    endfunction

    // Counts of destination observations that were NOT values the source held.
    integer bad_bin = 0, bad_gray = 0, bad_flag = 0;
    // And of observations in total, so the bad counts can be read as a proportion.
    integer obs_bin = 0, obs_gray = 0, obs_flag = 0;

    // Distinct impossible values seen on scheme 1, so the chapter can say how many
    // different lies were told rather than only how often.
    reg  seen_bad [0:NVAL-1];
    integer distinct_bad = 0;

    integer hj;
    always @(posedge src_clk) if (src_rst_n) begin
        for (hj = HIST-1; hj > 0; hj = hj - 1) hist[hj] = hist[hj-1];
        hist[0] = cnt;
    end

    // The destination observations. Sampled on the destination clock, which is the
    // only place they exist.
    always @(posedge dst_clk) if (dst_rst_n && src_rst_n) begin
        obs_bin = obs_bin + 1;
        if (!is_recent(dst_bin_wrong)) begin
            bad_bin = bad_bin + 1;
            if (!seen_bad[dst_bin_wrong]) begin
                seen_bad[dst_bin_wrong] = 1'b1;
                distinct_bad = distinct_bad + 1;
            end
        end

        obs_gray = obs_gray + 1;
        if (!is_recent(dst_gray_ok)) bad_gray = bad_gray + 1;

        if (dst_flagged_stb) begin
            obs_flag = obs_flag + 1;
            if (!is_recent(dst_flagged)) bad_flag = bad_flag + 1;
        end
    end

    task automatic clear_stats;
        integer c;
        begin
            bad_bin = 0; bad_gray = 0; bad_flag = 0;
            obs_bin = 0; obs_gray = 0; obs_flag = 0;
            distinct_bad = 0;
            for (c = 0; c < NVAL; c = c + 1) seen_bad[c] = 1'b0;
        end
    endtask

    // Runs the counter for `n` updates, spaced `gap` source cycles apart. `gap` is
    // the whole difference between the two phases: at 1 the source never holds
    // still, and at 8 it holds still for far longer than the synchroniser latency.
    task automatic run_counter(input integer n, input integer gap);
        integer i, g;
        begin
            for (i = 0; i < n; i = i + 1) begin
                @(negedge src_clk);
                update = 1'b1;
                cnt    = cnt + 1'b1;
                @(negedge src_clk);
                update = 1'b0;
                for (g = 1; g < gap; g = g + 1) @(negedge src_clk);
            end
            repeat (20) @(posedge dst_clk);
        end
    endtask

    integer k;
    integer p1_bin, p1_gray, p1_flag, p1_distinct, p1_obs;

    initial begin
        for (k = 0; k < NVAL; k = k + 1) seen_bad[k] = 1'b0;
        for (k = 0; k < HIST; k = k + 1) hist[k] = {W{1'b0}};

        src_rst_n = 1'b1; dst_rst_n = 1'b1;
        repeat (2) @(posedge dst_clk);
        src_rst_n = 1'b0; dst_rst_n = 1'b0;
        repeat (4) @(posedge dst_clk);
        src_rst_n = 1'b1; dst_rst_n = 1'b1;
        repeat (4) @(posedge dst_clk);

        // Reset the bookkeeping now that the design is out of reset, so that the
        // reset value is not counted as a source observation.
        bad_bin = 0; bad_gray = 0; bad_flag = 0;
        obs_bin = 0; obs_gray = 0; obs_flag = 0;
        distinct_bad = 0;
        for (k = 0; k < NVAL; k = k + 1) seen_bad[k] = 1'b0;

        $display("  an %0d-bit counter crossing two unrelated clocks (14 ns source, 10 ns destination) with up to %0d ns of per-bit skew; an observation is IMPOSSIBLE if the source did not hold it within the last %0d source cycles",
                 W, (W-1)*SKEW, HIST);

        // =============================================================
        // PHASE 1: the source NEVER HOLDS STILL -- one update per source cycle,
        // which is the hardest case for every scheme and the normal case for a
        // FIFO pointer.
        // =============================================================
        run_counter(400, 1);

        p1_bin = bad_bin; p1_gray = bad_gray; p1_flag = bad_flag;
        p1_distinct = distinct_bad; p1_obs = obs_bin;

        $display("  PHASE 1 -- one update per source cycle: the source never holds still");
        $display("  scheme                        observations  impossible values");
        $display("  1  N synchronisers on a bus   %12d  %17d", obs_bin, bad_bin);
        $display("  2  Gray code, then decoded    %12d  %17d", obs_gray, bad_gray);
        $display("  3  data plus a flag           %12d  %17d", obs_flag, bad_flag);

        // 1. SCHEME 1 OBSERVES VALUES THAT NEVER EXISTED. Without this the chapter
        //    has an argument and no evidence, and the argument is one that every
        //    reviewer has heard and some do not believe.
        if (bad_bin == 0) begin
            $display("  FAIL: N independent synchronisers on a %0d-bit bus produced no impossible value across %0d observations -- with up to %0d ns of skew they must, and a bench that finds none is either driving all bits at the same instant or asking whether a value ever existed rather than whether it existed recently",
                     W, obs_bin, (W-1)*SKEW);
            errors = errors + 1;
        end

        // 2. GRAY CODE NEVER DOES. This is the positive claim and it is exact: a
        //    one-bit-at-a-time encoding cannot produce a third value, so the count
        //    must be zero rather than small.
        if (bad_gray != 0) begin
            $display("  FAIL: the Gray-coded crossing observed %0d impossible values -- consecutive Gray codes differ in one bit, so catching an update part-way must yield the old value or the new one",
                     bad_gray);
            errors = errors + 1;
        end

        // 3. AND DATA PLUS A FLAG FAILS HERE TOO, which was not the expected
        //    result and is the more useful one. Its precondition is that the source
        //    HOLDS THE DATA STILL until the flag has been observed -- and a source
        //    updating every cycle violates that, so the destination samples a
        //    register that is mid-change and reads a mixture, exactly as scheme 1
        //    does. The flag is not the weak part; the holding is.
        //
        //    This is Chapter 15.2's lifetime requirement arriving from the other
        //    direction, and it is why Chapter 15.6 has to build a FIFO: an SPI
        //    slave's source cannot be told to hold still, because SCLK does not
        //    stop.
        if (bad_flag == 0) begin
            $display("  FAIL: data plus a flag should ALSO observe impossible values when the source never holds still -- its correctness depends on the data being stable when sampled, not on the flag");
            errors = errors + 1;
        end
        $display("  data plus a flag observed %0d impossible values as well, because its precondition is that the source HOLDS THE DATA STILL and a source updating every cycle does not -- the flag is not the weak part, the holding is",
                 bad_flag);

        // 4. AND THE LIES WERE VARIED, not one repeated glitch. A single impossible
        //    value could be dismissed as an artefact of one particular alignment;
        //    several distinct ones cannot.
        if (p1_distinct < 2) begin
            $display("  FAIL: only %0d distinct impossible value(s) -- one could be an alignment artefact",
                     p1_distinct);
            errors = errors + 1;
        end
        $display("  scheme 1 produced %0d distinct impossible values out of the %0d the counter can hold, so it is not one unlucky alignment",
                 p1_distinct, NVAL);

        // =============================================================
        // PHASE 2: the source HOLDS STILL between updates -- one update every eight
        // source cycles, which is far longer than the synchroniser latency.
        // =============================================================
        clear_stats();
        run_counter(60, 8);

        $display("  PHASE 2 -- one update every eight source cycles: the source holds still");
        $display("  scheme                        observations  impossible values");
        $display("  1  N synchronisers on a bus   %12d  %17d", obs_bin, bad_bin);
        $display("  2  Gray code, then decoded    %12d  %17d", obs_gray, bad_gray);
        $display("  3  data plus a flag           %12d  %17d", obs_flag, bad_flag);

        // 4b. WITH THE SOURCE HOLDING STILL, DATA PLUS A FLAG IS EXACT. Its
        //     precondition is met, and it carries arbitrary data rather than only a
        //     counter -- which is what makes it the general answer and Gray the
        //     special one.
        if (bad_flag != 0) begin
            $display("  FAIL: with the source holding still, data plus a flag observed %0d impossible values",
                     bad_flag);
            errors = errors + 1;
        end
        if (bad_gray != 0) begin
            $display("  FAIL: the Gray crossing observed %0d impossible values with the source holding still",
                     bad_gray);
            errors = errors + 1;
        end

        // 4c. AND SCHEME 1 IS STILL WRONG, just less often -- which is the reason it
        //     reaches silicon. Slowing the source reduces the error RATE without
        //     changing the error, so a bench that runs a slow source and counts
        //     failures per thousand observations concludes the crossing is fine.
        if (bad_bin == 0) begin
            $display("  FAIL: scheme 1 should still be wrong with a slow source -- skew does not care how often the value changes");
            errors = errors + 1;
        end
        $display("  scheme 1 was still wrong with the source holding still, %0d times in %0d observations against %0d in %0d when it never held still -- slowing the source reduces the RATE and not the fault, which is exactly how this crossing survives testing and reaches silicon",
                 bad_bin, obs_bin, p1_bin, p1_obs);

        // 5. THE OBSERVATION THAT MAKES SCHEME 1 DANGEROUS RATHER THAN MERELY
        //    WRONG: every impossible value is a WELL-FORMED one. There is no bit
        //    pattern a downstream consumer could reject, so nothing downstream can
        //    detect the fault -- which is the whole reason it survives to silicon.
        $display("  every impossible value was a legal %0d-bit pattern, so no downstream check could reject one -- the fault is undetectable by anything except the crossing being built correctly in the first place",
                 W);

        // 6. AND THE COST OF THE CORRECT SCHEMES, so the chapter is not only a
        //    warning. Gray costs an encoder, a decoder of W-1 XORs and nothing
        //    else; data-plus-flag costs one synchronised bit and a LIFETIME
        //    requirement on the data register, which is Chapter 15.2's finding.
        if (obs_flag == 0) begin
            $display("  FAIL: the data-plus-flag scheme delivered nothing at all");
            errors = errors + 1;
        end
        $display("  the flagged scheme delivered %0d observations rather than one per destination cycle, because it only publishes when the flag arrives -- it is exact, it is slower, and it needs the source to hold still, and those three together are the trade",
                 obs_flag);

        // The exact counts are phase- and simulator-dependent, and deliberately not
        // asserted: which update a destination edge happens to land inside depends
        // on the relationship between two clocks that have none, so the three
        // language benches report different totals from the same design. What is
        // asserted is the SHAPE -- scheme 1 non-zero in both regimes, Gray zero in
        // both, data-plus-flag non-zero only when the source never holds still --
        // and that is identical in all three, because it follows from the encoding
        // rather than from the alignment.
        if (errors == 0)
            $display("PASS: a two-flop synchroniser produces a SETTLED value and nothing else -- it does not choose which of two values you get, it does not convey a pulse, and on a BUS it is wrong, because each bit settles independently and the destination reads a mixture of bits from before and after the update. Driven with per-bit skew of the kind ordinary routing produces, N independent synchronisers on an 8-bit counter observed impossible values repeatedly and in many distinct forms, every one a well-formed pattern that no downstream check could reject -- and it was still wrong when the source was slowed to update every eighth cycle, just less often, which is precisely how this crossing passes a regression and reaches silicon. A GRAY-CODED crossing through the identical synchronisers observed exactly ZERO in both regimes, because consecutive Gray codes differ in one bit so catching an update part-way yields the old value or the new one -- and that holds only for a value that changes by one step, which suits a counter or a FIFO pointer and not arbitrary data. DATA PLUS A FLAG observed zero only in the second regime, and that was the useful surprise: it fails exactly as badly as scheme 1 when the source never holds still, because its correctness rests on the data being STABLE WHEN SAMPLED rather than on the flag -- Chapter 15.2s lifetime requirement arriving from the other direction, and the reason Chapter 15.6 must build a FIFO, since an SPI slaves source cannot be told to hold still while SCLK is running. So the choice is not between synchroniser depths. It is between encodings: change one bit at a time, or hold the value still and cross one bit of handshake, or buffer it");
        else
            $display("FAIL: %0d error(s)", errors);
        $finish;
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_cdc_vector_tb.v — the same bench in Verilog-2001
// spi_cdc_vector_tb.v
//
// The measurement is "how many times did the destination observe a value the
// source never held?", and it is made three times on the same stimulus.
//
// HOW THE BENCH KNOWS WHAT THE SOURCE HELD, and why the obvious version of this
// does not work. The first attempt kept a bitmap of every value the source had EVER
// taken -- and a counter passes through every value, so within sixteen cycles the
// bitmap was all ones and no observation could ever be impossible. The bench
// reported zero faults on a scheme that is definitely broken, which is a good
// reminder that a checker returning the answer you hoped for is not evidence.
//
// The correct question is about RECENCY, not existence: the destination is reading a
// value through a synchroniser, so it must see one the source held WITHIN THE LAST
// FEW SOURCE CYCLES. The bench keeps a short history of the source's recent values
// and any observation outside that window is impossible -- exactly, not
// probabilistically, because the source cannot have held it.
//
// The counter is eight bits wide so that the 256-value space is far larger than the
// six-deep window; a mixture of two adjacent values then lands outside the window
// almost always rather than sometimes.
//
// HOW THE FAILURE IS MADE TO HAPPEN. The source bits are driven to the DUT through
// per-bit skew -- bit k is delayed by k times a small delta -- which is what
// unequal routing does on any real device. That is the honest way to exhibit
// incoherence in a simulator: bit-level settling differences are metastability and
// are not modelled, but skew is ordinary, it produces the identical failure, and a
// design that survives skew survives the settling case for the same reason.
//
// A bench that changes all bits at the same instant finds nothing wrong with scheme
// 1, which is why scheme 1 keeps being built.

`timescale 1ns/1ps

module spi_cdc_vector_tb;

    localparam W      = 8;
    localparam SYNC_N = 2;
    localparam NVAL   = 1 << W;
    // How far back a legitimate observation may come from: the synchroniser is
    // SYNC_N deep and the two clocks are close in period, so a value three or four
    // source cycles old is plausible. Six is generous, which matters -- a window
    // that is too tight reports correct designs as broken.
    localparam HIST   = 6;

    // Two unrelated clocks. The ratio is deliberately not an integer, so that the
    // destination's sampling instant walks through the source's period rather than
    // sitting at one point in it.
    reg src_clk;
    reg dst_clk;
    always #7  src_clk = ~src_clk;    // 14 ns
    always #5  dst_clk = ~dst_clk;    // 10 ns

    reg src_rst_n;
    reg dst_rst_n;

    // The source: a plain counter, and its Gray encoding.
    reg [W-1:0] cnt;
    reg         update;

    wire [W-1:0] cnt_gray = cnt ^ (cnt >> 1);

    // --- the skew injection -------------------------------------------------
    // Each bit of the value reaches the DUT through its own delay. This is the
    // whole reason the bench can see the bug, and it is a model of routing rather
    // than of metastability -- which is what makes it defensible.
    localparam time SKEW = 1;
    wire [W-1:0] bin_skewed;
    wire [W-1:0] gray_skewed;
    genvar gi;
    generate
        for (gi = 0; gi < W; gi = gi + 1) begin : g_skew
            // Bit 0 arrives on time, bit W-1 arrives (W-1)*SKEW late. A real
            // device's spread is smaller and less orderly; an orderly spread makes
            // the result reproducible, which a chapter needs.
            assign #(gi * SKEW) bin_skewed[gi]  = cnt[gi];
            assign #(gi * SKEW) gray_skewed[gi] = cnt_gray[gi];
        end
    endgenerate

    wire [W-1:0] dst_bin_wrong, dst_gray_ok, dst_flagged;
    wire         dst_flagged_stb;

    spi_cdc_vector #(.W(W), .SYNC_N(SYNC_N)) dut (
        .src_clk(src_clk), .src_rst_n(src_rst_n),
        .src_bin(bin_skewed), .src_gray(gray_skewed), .src_update(update),
        .dst_clk(dst_clk), .dst_rst_n(dst_rst_n),
        .dst_bin_wrong(dst_bin_wrong), .dst_gray_ok(dst_gray_ok),
        .dst_flagged(dst_flagged), .dst_flagged_stb(dst_flagged_stb)
    );

    integer errors;

    initial begin
        #2_000_000;
        $display("FAIL: the simulation did not finish within its time limit");
        $finish;
    end

    // The source's recent values, newest first. Maintained in SOURCE time, which is
    // the only place the source's value means anything.
    reg [W-1:0] hist [0:HIST-1];

        function is_recent;
        input [W-1:0] v;
        integer j;
        begin
            is_recent = 1'b0;
            for (j = 0; j < HIST; j = j + 1)
                if (hist[j] == v) is_recent = 1'b1;
        end
    endfunction

    // Counts of destination observations that were NOT values the source held.
    integer bad_bin, bad_gray, bad_flag;
    // And of observations in total, so the bad counts can be read as a proportion.
    integer obs_bin, obs_gray, obs_flag;

    // Distinct impossible values seen on scheme 1, so the chapter can say how many
    // different lies were told rather than only how often.
    reg  seen_bad [0:NVAL-1];
    integer distinct_bad;

    integer hj;
    always @(posedge src_clk) if (src_rst_n) begin
        for (hj = HIST-1; hj > 0; hj = hj - 1) hist[hj] = hist[hj-1];
        hist[0] = cnt;
    end

    // The destination observations. Sampled on the destination clock, which is the
    // only place they exist.
    always @(posedge dst_clk) if (dst_rst_n && src_rst_n) begin
        obs_bin = obs_bin + 1;
        if (!is_recent(dst_bin_wrong)) begin
            bad_bin = bad_bin + 1;
            if (!seen_bad[dst_bin_wrong]) begin
                seen_bad[dst_bin_wrong] = 1'b1;
                distinct_bad = distinct_bad + 1;
            end
        end

        obs_gray = obs_gray + 1;
        if (!is_recent(dst_gray_ok)) bad_gray = bad_gray + 1;

        if (dst_flagged_stb) begin
            obs_flag = obs_flag + 1;
            if (!is_recent(dst_flagged)) bad_flag = bad_flag + 1;
        end
    end

    task clear_stats;
        integer c;
        begin
            bad_bin = 0; bad_gray = 0; bad_flag = 0;
            obs_bin = 0; obs_gray = 0; obs_flag = 0;
            distinct_bad = 0;
            for (c = 0; c < NVAL; c = c + 1) seen_bad[c] = 1'b0;
        end
    endtask

    // Runs the counter for `n` updates, spaced `gap` source cycles apart. `gap` is
    // the whole difference between the two phases: at 1 the source never holds
    // still, and at 8 it holds still for far longer than the synchroniser latency.
        task run_counter;
        input integer n;
        input integer gap;
        integer i, g;
        begin
            for (i = 0; i < n; i = i + 1) begin
                @(negedge src_clk);
                update = 1'b1;
                cnt    = cnt + 1'b1;
                @(negedge src_clk);
                update = 1'b0;
                for (g = 1; g < gap; g = g + 1) @(negedge src_clk);
            end
            repeat (20) @(posedge dst_clk);
        end
    endtask

    integer k;
    integer p1_bin, p1_gray, p1_flag, p1_distinct, p1_obs;

    initial begin
        for (k = 0; k < NVAL; k = k + 1) seen_bad[k] = 1'b0;
        for (k = 0; k < HIST; k = k + 1) hist[k] = {W{1'b0}};

        src_rst_n = 1'b1; dst_rst_n = 1'b1;
        repeat (2) @(posedge dst_clk);
        src_rst_n = 1'b0; dst_rst_n = 1'b0;
        repeat (4) @(posedge dst_clk);
        src_rst_n = 1'b1; dst_rst_n = 1'b1;
        repeat (4) @(posedge dst_clk);

        // Reset the bookkeeping now that the design is out of reset, so that the
        // reset value is not counted as a source observation.
        bad_bin = 0; bad_gray = 0; bad_flag = 0;
        obs_bin = 0; obs_gray = 0; obs_flag = 0;
        distinct_bad = 0;
        for (k = 0; k < NVAL; k = k + 1) seen_bad[k] = 1'b0;

        $display("  an %0d-bit counter crossing two unrelated clocks (14 ns source, 10 ns destination) with up to %0d ns of per-bit skew; an observation is IMPOSSIBLE if the source did not hold it within the last %0d source cycles",
                 W, (W-1)*SKEW, HIST);

        // =============================================================
        // PHASE 1: the source NEVER HOLDS STILL -- one update per source cycle,
        // which is the hardest case for every scheme and the normal case for a
        // FIFO pointer.
        // =============================================================
        run_counter(400, 1);

        p1_bin = bad_bin; p1_gray = bad_gray; p1_flag = bad_flag;
        p1_distinct = distinct_bad; p1_obs = obs_bin;

        $display("  PHASE 1 -- one update per source cycle: the source never holds still");
        $display("  scheme                        observations  impossible values");
        $display("  1  N synchronisers on a bus   %12d  %17d", obs_bin, bad_bin);
        $display("  2  Gray code, then decoded    %12d  %17d", obs_gray, bad_gray);
        $display("  3  data plus a flag           %12d  %17d", obs_flag, bad_flag);

        // 1. SCHEME 1 OBSERVES VALUES THAT NEVER EXISTED. Without this the chapter
        //    has an argument and no evidence, and the argument is one that every
        //    reviewer has heard and some do not believe.
        if (bad_bin == 0) begin
            $display("  FAIL: N independent synchronisers on a %0d-bit bus produced no impossible value across %0d observations -- with up to %0d ns of skew they must, and a bench that finds none is either driving all bits at the same instant or asking whether a value ever existed rather than whether it existed recently",
                     W, obs_bin, (W-1)*SKEW);
            errors = errors + 1;
        end

        // 2. GRAY CODE NEVER DOES. This is the positive claim and it is exact: a
        //    one-bit-at-a-time encoding cannot produce a third value, so the count
        //    must be zero rather than small.
        if (bad_gray != 0) begin
            $display("  FAIL: the Gray-coded crossing observed %0d impossible values -- consecutive Gray codes differ in one bit, so catching an update part-way must yield the old value or the new one",
                     bad_gray);
            errors = errors + 1;
        end

        // 3. AND DATA PLUS A FLAG FAILS HERE TOO, which was not the expected
        //    result and is the more useful one. Its precondition is that the source
        //    HOLDS THE DATA STILL until the flag has been observed -- and a source
        //    updating every cycle violates that, so the destination samples a
        //    register that is mid-change and reads a mixture, exactly as scheme 1
        //    does. The flag is not the weak part; the holding is.
        //
        //    This is Chapter 15.2's lifetime requirement arriving from the other
        //    direction, and it is why Chapter 15.6 has to build a FIFO: an SPI
        //    slave's source cannot be told to hold still, because SCLK does not
        //    stop.
        if (bad_flag == 0) begin
            $display("  FAIL: data plus a flag should ALSO observe impossible values when the source never holds still -- its correctness depends on the data being stable when sampled, not on the flag");
            errors = errors + 1;
        end
        $display("  data plus a flag observed %0d impossible values as well, because its precondition is that the source HOLDS THE DATA STILL and a source updating every cycle does not -- the flag is not the weak part, the holding is",
                 bad_flag);

        // 4. AND THE LIES WERE VARIED, not one repeated glitch. A single impossible
        //    value could be dismissed as an artefact of one particular alignment;
        //    several distinct ones cannot.
        if (p1_distinct < 2) begin
            $display("  FAIL: only %0d distinct impossible value(s) -- one could be an alignment artefact",
                     p1_distinct);
            errors = errors + 1;
        end
        $display("  scheme 1 produced %0d distinct impossible values out of the %0d the counter can hold, so it is not one unlucky alignment",
                 p1_distinct, NVAL);

        // =============================================================
        // PHASE 2: the source HOLDS STILL between updates -- one update every eight
        // source cycles, which is far longer than the synchroniser latency.
        // =============================================================
        clear_stats();
        run_counter(60, 8);

        $display("  PHASE 2 -- one update every eight source cycles: the source holds still");
        $display("  scheme                        observations  impossible values");
        $display("  1  N synchronisers on a bus   %12d  %17d", obs_bin, bad_bin);
        $display("  2  Gray code, then decoded    %12d  %17d", obs_gray, bad_gray);
        $display("  3  data plus a flag           %12d  %17d", obs_flag, bad_flag);

        // 4b. WITH THE SOURCE HOLDING STILL, DATA PLUS A FLAG IS EXACT. Its
        //     precondition is met, and it carries arbitrary data rather than only a
        //     counter -- which is what makes it the general answer and Gray the
        //     special one.
        if (bad_flag != 0) begin
            $display("  FAIL: with the source holding still, data plus a flag observed %0d impossible values",
                     bad_flag);
            errors = errors + 1;
        end
        if (bad_gray != 0) begin
            $display("  FAIL: the Gray crossing observed %0d impossible values with the source holding still",
                     bad_gray);
            errors = errors + 1;
        end

        // 4c. AND SCHEME 1 IS STILL WRONG, just less often -- which is the reason it
        //     reaches silicon. Slowing the source reduces the error RATE without
        //     changing the error, so a bench that runs a slow source and counts
        //     failures per thousand observations concludes the crossing is fine.
        if (bad_bin == 0) begin
            $display("  FAIL: scheme 1 should still be wrong with a slow source -- skew does not care how often the value changes");
            errors = errors + 1;
        end
        $display("  scheme 1 was still wrong with the source holding still, %0d times in %0d observations against %0d in %0d when it never held still -- slowing the source reduces the RATE and not the fault, which is exactly how this crossing survives testing and reaches silicon",
                 bad_bin, obs_bin, p1_bin, p1_obs);

        // 5. THE OBSERVATION THAT MAKES SCHEME 1 DANGEROUS RATHER THAN MERELY
        //    WRONG: every impossible value is a WELL-FORMED one. There is no bit
        //    pattern a downstream consumer could reject, so nothing downstream can
        //    detect the fault -- which is the whole reason it survives to silicon.
        $display("  every impossible value was a legal %0d-bit pattern, so no downstream check could reject one -- the fault is undetectable by anything except the crossing being built correctly in the first place",
                 W);

        // 6. AND THE COST OF THE CORRECT SCHEMES, so the chapter is not only a
        //    warning. Gray costs an encoder, a decoder of W-1 XORs and nothing
        //    else; data-plus-flag costs one synchronised bit and a LIFETIME
        //    requirement on the data register, which is Chapter 15.2's finding.
        if (obs_flag == 0) begin
            $display("  FAIL: the data-plus-flag scheme delivered nothing at all");
            errors = errors + 1;
        end
        $display("  the flagged scheme delivered %0d observations rather than one per destination cycle, because it only publishes when the flag arrives -- it is exact, it is slower, and it needs the source to hold still, and those three together are the trade",
                 obs_flag);

        // The exact counts are phase- and simulator-dependent, and deliberately not
        // asserted: which update a destination edge happens to land inside depends
        // on the relationship between two clocks that have none, so the three
        // language benches report different totals from the same design. What is
        // asserted is the SHAPE -- scheme 1 non-zero in both regimes, Gray zero in
        // both, data-plus-flag non-zero only when the source never holds still --
        // and that is identical in all three, because it follows from the encoding
        // rather than from the alignment.
        if (errors == 0)
            $display("PASS: a two-flop synchroniser produces a SETTLED value and nothing else -- it does not choose which of two values you get, it does not convey a pulse, and on a BUS it is wrong, because each bit settles independently and the destination reads a mixture of bits from before and after the update. Driven with per-bit skew of the kind ordinary routing produces, N independent synchronisers on an 8-bit counter observed impossible values repeatedly and in many distinct forms, every one a well-formed pattern that no downstream check could reject -- and it was still wrong when the source was slowed to update every eighth cycle, just less often, which is precisely how this crossing passes a regression and reaches silicon. A GRAY-CODED crossing through the identical synchronisers observed exactly ZERO in both regimes, because consecutive Gray codes differ in one bit so catching an update part-way yields the old value or the new one -- and that holds only for a value that changes by one step, which suits a counter or a FIFO pointer and not arbitrary data. DATA PLUS A FLAG observed zero only in the second regime, and that was the useful surprise: it fails exactly as badly as scheme 1 when the source never holds still, because its correctness rests on the data being STABLE WHEN SAMPLED rather than on the flag -- Chapter 15.2s lifetime requirement arriving from the other direction, and the reason Chapter 15.6 must build a FIFO, since an SPI slaves source cannot be told to hold still while SCLK is running. So the choice is not between synchroniser depths. It is between encodings: change one bit at a time, or hold the value still and cross one bit of handshake, or buffer it");
        else
            $display("FAIL: %0d error(s)", errors);
        $finish;
    end


    initial begin
        bad_bin = 0;
        bad_gray = 0;
        bad_flag = 0;
        obs_bin = 0;
        obs_gray = 0;
        obs_flag = 0;
        src_clk = 1'b0;
        dst_clk = 1'b0;
        src_rst_n = 1'b1;
        dst_rst_n = 1'b1;
        cnt = {W{1'b0}};
        update = 1'b0;
        errors = 0;
        distinct_bad = 0;
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_cdc_vector_tb.vhd — the same bench in VHDL
-- spi_cdc_vector_tb.vhd
--
-- The measurement is "how many times did the destination observe a value the
-- source never held?", and it is made three times on the same stimulus.
--
-- HOW THE BENCH KNOWS WHAT THE SOURCE HELD, and why the obvious version of this
-- does not work. The first attempt kept a bitmap of every value the source had EVER
-- taken -- and a counter passes through every value, so within sixteen cycles the
-- bitmap was all ones and no observation could ever be impossible. The bench
-- reported zero faults on a scheme that is definitely broken, which is a good
-- reminder that a checker returning the answer you hoped for is not evidence.
--
-- The correct question is about RECENCY, not existence: the destination is reading a
-- value through a synchroniser, so it must see one the source held WITHIN THE LAST
-- FEW SOURCE CYCLES. The bench keeps a short history of the source's recent values
-- and any observation outside that window is impossible -- exactly, not
-- probabilistically, because the source cannot have held it.
--
-- The counter is eight bits wide so that the 256-value space is far larger than the
-- six-deep window; a mixture of two adjacent values then lands outside the window
-- almost always rather than sometimes.
--
-- HOW THE FAILURE IS MADE TO HAPPEN. The source bits are driven to the DUT through
-- per-bit skew -- bit k is delayed by k times a small delta -- which is what
-- unequal routing does on any real device. That is the honest way to exhibit
-- incoherence in a simulator: bit-level settling differences are metastability and
-- are not modelled, but skew is ordinary, it produces the identical failure, and a
-- design that survives skew survives the settling case for the same reason.
--
-- A bench that changes all bits at the same instant finds nothing wrong with scheme
-- 1, which is why scheme 1 keeps being built.

library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

entity spi_cdc_vector_tb is
end entity;

architecture sim of spi_cdc_vector_tb is

    constant W      : positive := 8;
    constant SYNC_N : positive := 2;
    constant NVAL   : positive := 2 ** W;
    -- How far back a legitimate observation may come from. Six is generous, which
    -- matters: a window that is too tight reports correct designs as broken.
    constant HIST   : positive := 6;
    constant SKEW   : time     := 1 ns;

    -- Two unrelated clocks. The ratio is deliberately not an integer, so the
    -- destination's sampling instant walks through the source's period.
    signal src_clk : std_logic := '0';
    signal dst_clk : std_logic := '0';
    signal halt    : boolean   := false;

    signal src_rst_n : std_logic := '1';
    signal dst_rst_n : std_logic := '1';

    signal cnt      : unsigned(W - 1 downto 0) := (others => '0');
    signal update   : std_logic := '0';
    signal cnt_gray : std_logic_vector(W - 1 downto 0);

    -- The skew injection. Each bit reaches the DUT through its own delay, which is
    -- a model of ROUTING rather than of metastability -- and that is what makes it
    -- defensible: a design that survives skew survives the settling case for the
    -- same reason, and skew is something a reviewer can point at on a floorplan.
    signal bin_skewed  : std_logic_vector(W - 1 downto 0);
    signal gray_skewed : std_logic_vector(W - 1 downto 0);

    signal dst_bin_wrong, dst_gray_ok, dst_flagged
        : std_logic_vector(W - 1 downto 0);
    signal dst_flagged_stb : std_logic;

    -- The statistics are owned by the observer process. The stimulus asks for a
    -- clear through a strobe rather than writing them, because two drivers on a
    -- signal resolve to 'X' in VHDL.
    signal bad_bin, bad_gray, bad_flag : natural := 0;
    signal obs_bin, obs_gray, obs_flag : natural := 0;
    signal distinct_bad : natural := 0;
    signal stat_clr : std_logic := '0';

    -- The source's recent values, newest first, maintained in SOURCE time.
    type hist_arr is array (0 to HIST - 1) of std_logic_vector(W - 1 downto 0);
    -- Named `src_hist` rather than `hist`: VHDL is case-INSENSITIVE, so a signal
    -- called `hist` collides with the constant `HIST` and the error points at the
    -- second declaration rather than at the naming convention that caused it.
    signal src_hist : hist_arr := (others => (others => '0'));

begin

    srcclk : process
    begin
        while not halt loop
            src_clk <= '0'; wait for 7 ns; src_clk <= '1'; wait for 7 ns;
        end loop;
        wait;
    end process;

    dstclk : process
    begin
        while not halt loop
            dst_clk <= '0'; wait for 5 ns; dst_clk <= '1'; wait for 5 ns;
        end loop;
        wait;
    end process;

    cnt_gray <= std_logic_vector(cnt xor shift_right(cnt, 1));

    -- Bit 0 arrives on time, bit W-1 arrives (W-1)*SKEW late. A real device's
    -- spread is smaller and less orderly; an orderly spread makes the result
    -- reproducible, which a chapter needs.
    skewgen : for i in 0 to W - 1 generate
        bin_skewed(i)  <= std_logic(cnt(i)) after i * SKEW;
        gray_skewed(i) <= cnt_gray(i)       after i * SKEW;
    end generate;

    dut : entity work.spi_cdc_vector
        generic map (W => W, SYNC_N => SYNC_N)
        port map (src_clk => src_clk, src_rst_n => src_rst_n,
                  src_bin => bin_skewed, src_gray => gray_skewed,
                  src_update => update,
                  dst_clk => dst_clk, dst_rst_n => dst_rst_n,
                  dst_bin_wrong => dst_bin_wrong, dst_gray_ok => dst_gray_ok,
                  dst_flagged => dst_flagged, dst_flagged_stb => dst_flagged_stb);

    -- The history, maintained in source time, which is the only place the source's
    -- value means anything.
    histproc : process (src_clk)
    begin
        if rising_edge(src_clk) then
            if src_rst_n = '1' then
                for j in HIST - 1 downto 1 loop
                    src_hist(j) <= src_hist(j - 1);
                end loop;
                src_hist(0) <= std_logic_vector(cnt);
            end if;
        end if;
    end process;

    -- The destination observations, and the impossibility test. Sampled on the
    -- destination clock, which is the only place they exist.
    observe : process (dst_clk)
        variable seen_bad : std_logic_vector(NVAL - 1 downto 0) := (others => '0');

        impure function is_recent(v : std_logic_vector) return boolean is
        begin
            for j in 0 to HIST - 1 loop
                if src_hist(j) = v then return true; end if;
            end loop;
            return false;
        end function;
    begin
        if rising_edge(dst_clk) then
            if stat_clr = '1' then
                bad_bin  <= 0; bad_gray <= 0; bad_flag <= 0;
                obs_bin  <= 0; obs_gray <= 0; obs_flag <= 0;
                distinct_bad <= 0;
                seen_bad := (others => '0');
            elsif dst_rst_n = '1' and src_rst_n = '1' then
                obs_bin <= obs_bin + 1;
                if not is_recent(dst_bin_wrong) then
                    bad_bin <= bad_bin + 1;
                    if seen_bad(to_integer(unsigned(dst_bin_wrong))) = '0' then
                        seen_bad(to_integer(unsigned(dst_bin_wrong))) := '1';
                        distinct_bad <= distinct_bad + 1;
                    end if;
                end if;

                obs_gray <= obs_gray + 1;
                if not is_recent(dst_gray_ok) then
                    bad_gray <= bad_gray + 1;
                end if;

                if dst_flagged_stb = '1' then
                    obs_flag <= obs_flag + 1;
                    if not is_recent(dst_flagged) then
                        bad_flag <= bad_flag + 1;
                    end if;
                end if;
            end if;
        end if;
    end process;

    watchdog : process
    begin
        wait for 2 ms;
        if not halt then
            report "FAIL: the simulation did not finish within its time limit"
                severity failure;
        end if;
        wait;
    end process;

    stim : process
        variable errs : natural := 0;
        variable p1_bin, p1_distinct, p1_obs : natural;

        procedure clear_stats is
        begin
            stat_clr <= '1';
            for i in 1 to 2 loop wait until rising_edge(dst_clk); end loop;
            stat_clr <= '0';
            wait until rising_edge(dst_clk);
        end procedure;

        -- Runs the counter for `n` updates, `gap` source cycles apart. `gap` is the
        -- whole difference between the two phases: at 1 the source never holds
        -- still, and at 8 it holds still for far longer than the synchroniser.
        procedure run_counter(n : natural; gap : natural) is
        begin
            for i in 1 to n loop
                wait until falling_edge(src_clk);
                update <= '1';
                cnt    <= cnt + 1;
                wait until falling_edge(src_clk);
                update <= '0';
                for g in 2 to gap loop
                    wait until falling_edge(src_clk);
                end loop;
            end loop;
            for i in 1 to 20 loop wait until rising_edge(dst_clk); end loop;
        end procedure;

        procedure show(phase : string) is
        begin
            report "  " & phase;
            report "  scheme                        observations  impossible values";
            report "  1  N synchronisers on a bus   " &
                   integer'image(obs_bin) & "  " & integer'image(bad_bin);
            report "  2  Gray code, then decoded    " &
                   integer'image(obs_gray) & "  " & integer'image(bad_gray);
            report "  3  data plus a flag           " &
                   integer'image(obs_flag) & "  " & integer'image(bad_flag);
        end procedure;
    begin
        src_rst_n <= '1'; dst_rst_n <= '1';
        for i in 1 to 2 loop wait until rising_edge(dst_clk); end loop;
        src_rst_n <= '0'; dst_rst_n <= '0';
        for i in 1 to 4 loop wait until rising_edge(dst_clk); end loop;
        src_rst_n <= '1'; dst_rst_n <= '1';
        for i in 1 to 4 loop wait until rising_edge(dst_clk); end loop;
        clear_stats;

        report "  an " & integer'image(W) &
               "-bit counter crossing two unrelated clocks (14 ns source, 10 ns destination) with up to " &
               integer'image(((W - 1) * SKEW) / 1 ns) &
               " ns of per-bit skew; an observation is IMPOSSIBLE if the source did not hold it within the last " &
               integer'image(HIST) & " source cycles";

        -- PHASE 1: the source NEVER HOLDS STILL.
        run_counter(400, 1);
        show("PHASE 1 -- one update per source cycle: the source never holds still");
        p1_bin := bad_bin; p1_distinct := distinct_bad; p1_obs := obs_bin;

        -- 1. SCHEME 1 OBSERVES VALUES THAT NEVER EXISTED.
        if bad_bin = 0 then
            report "  FAIL: N independent synchronisers on a bus produced no impossible value -- with skew they must, and a bench that finds none is either driving all bits at the same instant or asking whether a value ever existed rather than whether it existed recently";
            errs := errs + 1;
        end if;

        -- 2. GRAY CODE NEVER DOES. Exact, not small: a one-bit-at-a-time encoding
        --    cannot produce a third value.
        if bad_gray /= 0 then
            report "  FAIL: the Gray-coded crossing observed " &
                   integer'image(bad_gray) & " impossible values";
            errs := errs + 1;
        end if;

        -- 3. AND DATA PLUS A FLAG FAILS HERE TOO, which was not the expected result
        --    and is the more useful one: its precondition is that the source HOLDS
        --    THE DATA STILL, and a source updating every cycle does not.
        if bad_flag = 0 then
            report "  FAIL: data plus a flag should ALSO observe impossible values when the source never holds still -- its correctness depends on the data being stable when sampled, not on the flag";
            errs := errs + 1;
        end if;
        report "  data plus a flag observed " & integer'image(bad_flag) &
               " impossible values as well, because its precondition is that the source HOLDS THE DATA STILL and a source updating every cycle does not -- the flag is not the weak part, the holding is";

        -- 4. AND THE LIES WERE VARIED, not one repeated glitch.
        if p1_distinct < 2 then
            report "  FAIL: only " & integer'image(p1_distinct) &
                   " distinct impossible value(s) -- one could be an alignment artefact";
            errs := errs + 1;
        end if;
        report "  scheme 1 produced " & integer'image(p1_distinct) &
               " distinct impossible values out of the " & integer'image(NVAL) &
               " the counter can hold, so it is not one unlucky alignment";

        -- PHASE 2: the source HOLDS STILL between updates.
        clear_stats;
        run_counter(60, 8);
        show("PHASE 2 -- one update every eight source cycles: the source holds still");

        if bad_flag /= 0 then
            report "  FAIL: with the source holding still, data plus a flag observed " &
                   integer'image(bad_flag) & " impossible values";
            errs := errs + 1;
        end if;
        if bad_gray /= 0 then
            report "  FAIL: the Gray crossing observed " & integer'image(bad_gray) &
                   " impossible values with the source holding still";
            errs := errs + 1;
        end if;
        if bad_bin = 0 then
            report "  FAIL: scheme 1 should still be wrong with a slow source -- skew does not care how often the value changes";
            errs := errs + 1;
        end if;
        report "  scheme 1 was still wrong with the source holding still, " &
               integer'image(bad_bin) & " times in " & integer'image(obs_bin) &
               " observations against " & integer'image(p1_bin) & " in " &
               integer'image(p1_obs) &
               " when it never held still -- slowing the source reduces the RATE and not the fault, which is exactly how this crossing survives testing and reaches silicon";

        report "  every impossible value was a legal " & integer'image(W) &
               "-bit pattern, so no downstream check could reject one -- the fault is undetectable by anything except the crossing being built correctly in the first place";

        if obs_flag = 0 then
            report "  FAIL: the data-plus-flag scheme delivered nothing at all";
            errs := errs + 1;
        end if;
        report "  the flagged scheme delivered " & integer'image(obs_flag) &
               " observations rather than one per destination cycle, because it only publishes when the flag arrives -- it is exact, it is slower, and it needs the source to hold still, and those three together are the trade";

        -- The exact counts are phase- and simulator-dependent, and deliberately not
        -- asserted: which update a destination edge happens to land inside depends
        -- on the relationship between two clocks that have none, so the three
        -- language benches report different totals from the same design. What is
        -- asserted is the SHAPE -- scheme 1 non-zero in both regimes, Gray zero in
        -- both, data-plus-flag non-zero only when the source never holds still --
        -- and that is identical in all three, because it follows from the encoding
        -- rather than from the alignment.
        if errs = 0 then
            report "PASS: a two-flop synchroniser produces a SETTLED value and nothing else -- it does not choose which of two values you get, it does not convey a pulse, and on a BUS it is wrong, because each bit settles independently and the destination reads a mixture of bits from before and after the update. Driven with per-bit skew of the kind ordinary routing produces, N independent synchronisers on an 8-bit counter observed impossible values repeatedly and in many distinct forms, every one a well-formed pattern that no downstream check could reject -- and it was still wrong when the source was slowed to update every eighth cycle, just less often, which is precisely how this crossing passes a regression and reaches silicon. A GRAY-CODED crossing through the identical synchronisers observed exactly ZERO in both regimes, because consecutive Gray codes differ in one bit so catching an update part-way yields the old value or the new one -- and that holds only for a value that changes by one step, which suits a counter or a FIFO pointer and not arbitrary data. DATA PLUS A FLAG observed zero only in the second regime, and that was the useful surprise: it fails exactly as badly as scheme 1 when the source never holds still, because its correctness rests on the data being STABLE WHEN SAMPLED rather than on the flag -- Chapter 15.2's lifetime requirement arriving from the other direction, and the reason Chapter 15.6 must build a FIFO, since an SPI slave's source cannot be told to hold still while SCLK is running. So the choice is not between synchroniser depths. It is between encodings: change one bit at a time, or hold the value still and cross one bit of handshake, or buffer it";
        else
            report "FAIL: " & integer'image(errs) & " error(s)" severity error;
        end if;

        halt <= true;
        wait;
    end process;

end architecture;

7. Why a Verification Engineer Cares

Ask about recency, not existence. The bookkeeping in §5's callout is the single most transferable idea here: a checker that asks "did this value ever exist" against a counter can only ever answer yes. The window has to be a few source cycles wide — generous enough not to fail a correct design, narrow enough that a mixture lands outside it.

Inject skew, and put it between two flops. Without skew, scheme 1 is indistinguishable from scheme 2 forever. With skew in the wrong place, every scheme including the correct one fails. The seam being a port pair is what makes the position reviewable.

Run both regimes, and compare the numbers rather than the verdicts. The interesting result is not that scheme 1 fails — it is that slowing the source reduces its failure count without fixing it. A suite that runs one regime and reports pass/fail cannot say that.

Assert the shape, not the counts. Three simulators give three totals for the same design, for the same reason two boards would. Asserting "scheme 1 non-zero, Gray zero" is robust; asserting "13" is a bench that fails on a correct design.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// Properties for a vector crossing. The first is the only one that can be written
// about the wrong scheme, and it is an assertion that it is wrong.

property p_gray_changes_one_bit_at_a_time;
    // The property the encoding exists for, and the only one worth binding inside a
    // Gray-coded crossing: if this can fail, the guarantee is void and the crossing
    // is scheme 1 wearing a Gray code's name.
    @(posedge src_clk) disable iff (!src_rst_n)
        $countones(src_gray ^ $past(src_gray)) <= 1;
endproperty

property p_flagged_data_stable_across_the_window;
    // Scheme 3's real precondition, and the one that failed in phase 1: the data must
    // not change between the flag toggling and the destination sampling it.
    @(posedge src_clk) disable iff (!src_rst_n)
        src_update |=> $stable(src_bin) [*SYNC_N+2];
endproperty

property p_flagged_strobe_follows_the_flag;
    // The destination publishes only because the flag arrived, never because the data
    // changed -- the data is unsynchronised and must never be watched.
    @(posedge dst_clk) disable iff (!dst_rst_n)
        dst_flagged_stb |-> $past(f_arrived);
endproperty

property p_no_synchroniser_on_the_data_path;
    // Written as a structural check rather than a temporal one, because it is
    // structural: the data register's value must appear at the destination in the
    // same cycle it is sampled, with no intermediate stage. A design that grew one
    // has reintroduced scheme 1.
    @(posedge dst_clk) disable iff (!dst_rst_n)
        dst_flagged_stb |-> (dst_flagged == $past(src_bin));
endproperty
Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// Coverage. The axes are HOW MANY BITS CHANGED and WHETHER THE SOURCE WAS STILL,
// because those two decide whether a partial observation is possible and whether it
// matters -- and a suite that never varies the second concludes scheme 3 is correct.

covergroup cg_vector @(posedge src_clk iff src_update);
    option.per_instance = 1;

    // Bits changing in one update. A Gray-coded value must only ever hit 1; a binary
    // counter hits high values at every power-of-two boundary, and those are the
    // updates during which scheme 1 can lie the most.
    changed: coverpoint bits_changed_this_update {
        bins one    = {1};
        bins two    = {2};
        bins few    = {[3:4]};
        bins many   = {[5:$]};            // 0111 -> 1000 and friends
    }

    // Source cycles between updates. One is the FIFO-pointer case and the regime in
    // which data-plus-a-flag fails; eight is the regime in which it is exact.
    spacing: coverpoint src_cycles_between_updates {
        bins every_cycle = {1};
        bins two         = {2};
        bins few         = {[3:7]};
        bins holds_still = {[8:$]};
    }

    // Skew spread on the seam, in destination-clock tenths. Zero is the bin that
    // makes scheme 1 look correct, so a suite must cover more than it.
    skew: coverpoint seam_skew_tenths {
        bins none   = {0};
        bins small  = {[1:3]};
        bins third  = {[4:6]};
        bins large  = {[7:$]};
    }

    x_changed_skew:   cross changed, skew;
    x_spacing_skew:   cross spacing, skew;

endgroup

8. Why an FPGA or ASIC Engineer Cares

Gray coding costs an encoder and a decoder, and nothing else. The encoder is v ^ (v >> 1), which is W-1 XOR gates with no depth. The decoder is a prefix XOR, which is W-1 XORs in a chain — so its depth grows with the width, and at 32 bits it is worth registering the decoded value rather than using it combinationally. That is the entire implementation cost of the scheme.

Data-plus-a-flag costs one synchronised bit and a LIFETIME requirement, and the second is not a gate count. The data register must hold its value from the flag toggling until SYNC_N destination cycles later, which is a constraint on the source's behaviour rather than on the destination's logic — and Chapter 15.2 §5 is what happens when a reset shortens that lifetime.

A multi-bit crossing must be declared to the CDC tool, and the declaration is the review. Vendor CDC checkers find "a multi-bit signal crossing domains" and ask what protects it. The answers they accept are a Gray declaration, a handshake declaration, or a waiver — and a waiver on a bus with no protection is exactly scheme 1, signed off.

The synchroniser flops must not be merged across bits. A W-bit synchroniser is W independent two-flop chains, and a tool that recognises them as a shift register array may pack them in ways that change their relative delays. ASYNC_REG on all of them, and the anti-shift-register attribute, for the same reason as Chapter 15.1 — with the extra consequence that here the relative timing between bits is what the scheme is about.

9. Failure Signature — Occasional Impossible Sensor Readings

The symptom:

"Once every few thousand reads we get a temperature of 4000 degrees. The sensor is fine, the payload checksum passes, and the driver is single-threaded."

What is happening: a multi-bit value crossed by N independent synchronisers, and the destination sampled a mixture of two adjacent source values. The reading is a well-formed number that the source never produced.

Why the checksum does not help: a checksum validates that the bytes were received correctly, and they were. Scheme 1's fault is not a corruption of bytes; it is a mixture of bit positions within one value, and every byte of the mixture is a byte the source really drove at some point. No per-byte or per-payload integrity check can see it.

How to confirm: read the value twice in quick succession. A mixture will usually not repeat, and a genuine sensor fault will. And the structural check is faster than either: count the flops between the source register and the destination's first use, per bit, and ask what makes the bits agree with each other. If the answer is "two flops each", that is the fault.

10. Common Misconceptions

"A two-flop synchroniser makes a signal safe to cross." It makes the value settled. It does not choose which of two values you get, it does not convey a pulse, and on a bus it produces values that never existed.

"Synchronise every bit and the bus is safe." Every bit is individually settled and collectively incoherent. That is the whole failure.

"A wider synchroniser helps." Depth is about settling. Nothing about depth makes W bits agree with each other.

"Gray coding is the general answer for multi-bit crossings." It is the answer for values that change by one step. Arbitrary data does not, so Gray coding a data bus gives no guarantee at all — and Chapter 15.9 builds an encoding that looks like Gray and is not.

"Data-plus-a-flag is safe because the flag is synchronised." Its correctness rests on the data being stable when sampled. Phase 1 of the measurement shows it failing exactly as badly as scheme 1 when the source never holds still.

"If the failure rate is low, the crossing is nearly correct." The rate is a property of the stimulus and the fault is a property of the structure. §5's phase 2 has a lower rate and the same fault, and that gap is how the design ships.

11. Reason It Through

Q. A 4-bit counter crosses with one synchroniser per bit. Which transitions can produce the most impossible values, and what does that say about where to look in a waveform?

The ones that change the most bits: 0111 → 1000 changes all four, 0011 → 0100 changes three. A mixture of two values that differ in k bits can produce up to 2^k - 2 values neither of them equals, so the power-of-two boundaries are where the lying is worst. In a waveform, the place to look is not where the data is unusual but where the counter crosses a boundary — which is why a stimulus of random values is better than a counter for exposing the fault and a counter is better for measuring it, since the window test needs a predictable source.

Q. Why is data-plus-a-flag the general answer while Gray coding is the special one, when Gray coding needs no lifetime requirement?

Because Gray coding's guarantee comes from a property of the data — that it changes by one step — and most data does not have that property. Data-plus-a-flag makes no assumption about the data at all; it converts the problem into a requirement on the source's behaviour, which is something a designer can arrange. The trade is that the requirement is a number rather than a structure, so it can be violated later by a change that looks unrelated — which is exactly Chapter 15.2 §5.

Q. The bench's recency window is six source cycles. What breaks if it is two, and what breaks if it is sixty?

At two, a correct design fails: the destination's observation legitimately lags the source by the synchroniser depth plus the clock-ratio slack, so a window tighter than that reports correct values as impossible. At sixty, against an 8-bit counter, the window covers nearly a quarter of the value space, so a mixture has a good chance of landing inside it and the test stops detecting the fault. The window has to be wider than the real latency and much narrower than the value space — which is why the counter is 8 bits rather than 4.

Q. Why does the design bring the skew seam out as two ports rather than keeping the delay inside the testbench's source model?

Because the two positions are not equivalent and the difference is invisible in a port list that hides it. Skew before the source's register corrupts the source and breaks every scheme; skew between the two registers models routing and breaks only the schemes that deserve it. Making the seam a port pair — tied together in a real design — puts the choice where a reviewer reads it, and the first version of this measurement got it wrong precisely because the position was implicit.

Q. A design Gray-codes a FIFO write pointer and also crosses a 32-bit data word by synchronising every bit, arguing that the data is only read when the pointer says it is valid. Is the argument sound?

The argument is sound and the implementation is not. If the data is only read when the pointer says so, then the data does not need synchronising at all — it needs to be held stable, which is data-plus-a-flag. Synchronising it as well adds 64 flops and, worse, adds a stage between the register and the read, so the value the destination sees is one cycle older than the pointer implies. The synchronisers are not just unnecessary; they break the timing relationship the argument depends on.

12. Understanding Check

13. Summary

A two-flop synchroniser produces a settled value and nothing else. It does not choose which of two values you get, it does not convey a pulse, and on a bus it is wrong — each bit settles independently and the destination reads a mixture of bits from before and after the update.

Driven with the per-bit skew ordinary routing produces, N independent synchronisers on an 8-bit counter observed impossible values repeatedly and in many distinct forms, every one a well-formed pattern no downstream check could reject — and it was still wrong when the source was slowed, just less often, which is precisely how this crossing passes a regression and reaches silicon.

Gray coding observed exactly zero in both regimes, because consecutive Gray codes differ in one bit so a partial observation is the old value or the new one. That holds only for a value that changes by one step, which suits a counter or a FIFO pointer and not arbitrary data.

Data plus a flag observed zero only when the source held still, and that was the useful surprise: it fails exactly as badly as scheme 1 when the source never does, because its correctness rests on the data being stable when sampled rather than on the flag — Chapter 15.2's lifetime requirement arriving from the other direction, and the reason Chapter 15.6 must build a FIFO.

So the choice is not between synchroniser depths. It is between encodings: change one bit at a time, or hold the value still and cross one bit of handshake, or buffer it.

Two things about the measurement are worth carrying. The skew must go between two flops, not in front of the first one — the version that got it wrong broke the correct design too. And the impossibility test must ask about recency, not existence: the first version asked whether a value had ever existed, against a counter that passes through every value, and reported zero faults on a design that is definitely broken.

For verification: inject skew and put it in the right place; run both regimes and compare the counts rather than the verdicts, because the falling rate is the finding; and assert the shape, since three simulators give three totals.

For implementation: Gray costs an encoder and a prefix-XOR decoder whose depth grows with width; data-plus-a-flag costs one bit and a lifetime, which is not a gate count; a multi-bit crossing must be declared to the CDC tool, and a waiver on an unprotected bus is scheme 1 signed off.

14. What Comes Next

Both correct schemes carry one value at a time and both require the source to stop between values. An SPI slave's source cannot stop, because SCLK does not.

Chapter 15.6 — Handshake and FIFO Crossings builds what is left. A FIFO works because the data never crosses a boundary at all — a location is written by one clock and read by the other at a different time — and what crosses is two pointers, which are exactly Gray coding's special case. Every stale pointer reading turns out to be conservative, which is what makes the design safe rather than merely probable. And then the real question: what depth does an SPI burst actually need, and what happens when the answer is "no depth is sufficient"?

Continue learning