SPI · Module 15
Synchronizers and Their Limits
A two-flop synchroniser gives a settled value and nothing else — not which value, not a pulse, and not a bus. Measured with the per-bit skew routing produces, independent synchronisers on an 8-bit counter observe values the source never held, and are still wrong when the source is slowed. Gray coding observes zero; data-plus-a-flag fails too unless the source holds still.
A two-flop synchroniser solves exactly one problem: it takes a signal that may be sampled while it is changing and produces a value that has finished settling before anything downstream looks at it.
That is all. In particular:
IT DOES NOT guarantee WHICH of the two values you get. The first flop may
resolve to the old one, which is why Chapter 15.3's soft limit exists.
IT DOES NOT convey a pulse. A pulse narrower than a destination clock period is
lost with two flops exactly as with one -- Chapter 15.1 measured it.
IT DOES NOT work on a BUS.The third is the one that gets built anyway.
An N-bit value has to cross a clock boundary. Why is one synchroniser per bit wrong, and what is right instead?
1. Why N Independent Synchronisers Is Wrong
Each bit settles independently. The destination therefore samples a mixture: some bits from before the source's update and some from after. The result is a value the source never held.
a counter going 0111 -> 1000 changes four bits at once
catch two of them and the destination reads 1111 or 0000
both plausible, neither ever trueAnd the damage is unbounded: an address that briefly reads 1111 addresses the wrong memory, and no downstream check can tell, because the value is well formed. There is no bit pattern a consumer could reject.
2. What Makes This Simulatable, Which Is The Interesting Part
Bit-level settling differences are metastability and are not modelled. But the same incoherence is produced by skew — the source bits arriving at the destination's flops at slightly different times, which is what unequal routing does on any real device and is entirely ordinary.
So the measurement drives the crossing with per-bit skew, and that reproduces the failure exactly, for a reason a reviewer can point at on a floorplan.
A design that is safe against skew is safe against the settling differences too, because both amount to "the bits do not arrive together".
A bench that drives the vector with all bits changing at the same instant proves nothing at all — and is what almost every bench does.
3. The Three Correct Answers, And When Each Applies
GRAY CODE for a value that changes by ONE STEP at a time -- a counter or a
FIFO pointer. Consecutive Gray codes differ in exactly one bit,
so a destination catching an update part-way reads either the old
value or the new one, never a third.
Does NOT work for arbitrary data, which does not change by one step.
DATA PLUS A FLAG for arbitrary data. The data crosses with NO synchroniser at all
and is simply held stable; one flag bit crosses through a
synchroniser and says when the data is safe to sample.
This is what Chapter 15.2 does, and the requirement it creates is
a LIFETIME on the data register.
A FIFO for arbitrary data that keeps coming. Chapter 15.6.4. The Block Diagram
5. The Measurement
An 8-bit counter crossing two unrelated clocks — 14 ns source, 10 ns destination — with up to 3 ns of per-bit skew on the seam. An observation is impossible if the source did not hold it within the last six source cycles.
Phase 1 — the source never holds still
One update per source cycle, which is the hardest case for every scheme and the normal case for a FIFO pointer:
scheme observations impossible values
1 N synchronisers on a bus 1138 13
2 Gray code, then decoded 1138 0
3 data plus a flag 400 13Scheme 1 produced 12 distinct impossible values out of the 256 the counter can hold — so it is not one unlucky alignment. And every one was a legal 8-bit pattern, so no downstream check could reject one.
Scheme 3 failing here was not the expected result and is the more useful one. Its precondition is that the source holds the data still until the flag has been observed — and a source updating every cycle violates that, so the destination samples a register that is mid-change and reads a mixture, exactly as scheme 1 does.
The flag is not the weak part. The holding is.
Phase 2 — the source holds still
One update every eight source cycles:
scheme observations impossible values
1 N synchronisers on a bus 775 9
2 Gray code, then decoded 775 0
3 data plus a flag 60 0Gray is zero again. Data-plus-a-flag is now exact, because its precondition is met.
And scheme 1 is still wrong — 9 times in 775 observations against 13 in 1138. Slowing the source reduced the rate and not the fault.
The exact counts are phase- and simulator-dependent and are deliberately not asserted: which update a destination edge lands inside depends on two clocks with no relationship, so the three language benches report different totals from the same design. What is asserted is the shape — scheme 1 non-zero in both regimes, Gray zero in both, data-plus-a-flag non-zero only when the source never holds still — and that is identical in all three, because it follows from the encoding rather than from the alignment.
6. Building the Three Crossings — Three HDLs
The circuit
One module, three crossings of the same source value, with the skew seam brought out as a port pair.
// spi_cdc_vector.sv
//
// Chapter 15.5 -- what a two-flop synchroniser is for, and the thing it is
// routinely used for that it cannot do.
//
// A two-flop synchroniser solves exactly one problem: it takes a signal that may
// be sampled while it is changing and produces a value that has finished settling
// before anything downstream looks at it. That is all. In particular:
//
// IT DOES NOT guarantee WHICH of the two values you get. The first flop may
// resolve to the old one, which is why Chapter 15.3's soft limit exists.
//
// IT DOES NOT convey a pulse. A pulse narrower than a destination clock period
// is lost with two flops exactly as it is with one (Chapter 15.1 measured this).
//
// IT DOES NOT work on a BUS. This is the one that gets built anyway, and this
// file exists to make the failure visible rather than arguable.
//
// WHY N INDEPENDENT SYNCHRONISERS ON AN N-BIT VALUE IS WRONG.
//
// Each bit settles independently. The destination therefore samples a MIXTURE: some
// bits from before the source's update and some from after. The result is a value
// the source NEVER HELD.
//
// A counter going from 0111 to 1000 changes four bits at once. If the destination
// catches two of them, it reads 1111 or 0000 -- both plausible, neither ever true.
// And the damage is unbounded: an address that briefly reads 1111 addresses the
// wrong memory, and no downstream check can tell because the value is well formed.
//
// THE THREE CORRECT ANSWERS, AND WHEN EACH APPLIES.
//
// GRAY CODE for a value that changes by ONE STEP at a time, such as a
// counter or a FIFO pointer. Consecutive Gray codes differ in
// exactly one bit, so a destination that catches the update
// part-way reads either the old value or the new one -- never
// a third. It does NOT work for arbitrary data, because
// arbitrary data does not change by one step.
//
// DATA PLUS A FLAG for arbitrary data. The data crosses with NO synchroniser
// at all and is simply held stable; one flag bit crosses
// through a synchroniser and says when the data is safe to
// sample. This is what Chapter 15.2 does, and the requirement
// it creates is a LIFETIME on the data register.
//
// A FIFO for arbitrary data that keeps coming. Chapter 15.6.
//
// This module builds the wrong scheme and two of the right ones side by side, on
// the same source value, so that the testbench can count how many times each
// destination observed a value the source never held.
//
// WHAT MAKES THE FAILURE SIMULATABLE, WHICH IS THE INTERESTING PART.
//
// Bit-level settling differences are metastability and are not modelled. But the
// same incoherence is produced by SKEW -- the source bits arriving at the
// destination's flops at slightly different times, which is what unequal routing
// does on any real device and is entirely ordinary. So the testbench drives the
// source vector with a small per-bit skew, and that reproduces the failure exactly,
// for a reason a reviewer can point at on a floorplan.
//
// A design that is safe against skew is safe against the settling differences too,
// because both amount to "the bits do not arrive together". A bench that drives the
// vector with all bits changing at the same instant proves nothing at all, and is
// what almost every bench does.
module spi_cdc_vector #(
parameter int W = 4, // the width of the value being crossed
parameter int SYNC_N = 2
) (
input wire src_clk,
input wire src_rst_n,
input wire [W-1:0] src_bin, // the value, binary, as the source holds it
input wire [W-1:0] src_gray, // the same value, Gray-coded
input wire src_update, // one source cycle: the value just changed
input wire dst_clk,
input wire dst_rst_n,
// --- scheme 1: N INDEPENDENT SYNCHRONISERS ON A BUS. Wrong, and built here
// so that the count of impossible values is a measurement.
output wire [W-1:0] dst_bin_wrong,
// --- scheme 2: GRAY CODE through the same N synchronisers, then decoded.
output wire [W-1:0] dst_gray_ok,
// --- scheme 3: DATA PLUS A FLAG. The data is not synchronised at all.
output reg [W-1:0] dst_flagged,
output reg dst_flagged_stb
);
// ---------------------------------------------------------------- scheme 1
// One synchroniser per bit, which is the natural thing to write and is the
// bug. Nothing here is wrong on its own; the error is applying it to a value
// whose bits have to agree with each other.
reg [W-1:0] w_sr [0:SYNC_N-1];
integer i1;
always_ff @(posedge dst_clk or negedge dst_rst_n) begin
if (!dst_rst_n) begin
for (i1 = 0; i1 < SYNC_N; i1 = i1 + 1)
w_sr[i1] <= {W{1'b0}};
end else begin
w_sr[0] <= src_bin;
for (i1 = 1; i1 < SYNC_N; i1 = i1 + 1)
w_sr[i1] <= w_sr[i1-1];
end
end
assign dst_bin_wrong = w_sr[SYNC_N-1];
// ---------------------------------------------------------------- scheme 2
// The identical structure, on a Gray-coded value. The synchronisers are no
// better; what changed is that consecutive values differ in ONE bit, so
// catching an update part-way yields the old value or the new one.
reg [W-1:0] g_sr [0:SYNC_N-1];
integer i2;
always_ff @(posedge dst_clk or negedge dst_rst_n) begin
if (!dst_rst_n) begin
for (i2 = 0; i2 < SYNC_N; i2 = i2 + 1)
g_sr[i2] <= {W{1'b0}};
end else begin
g_sr[0] <= src_gray;
for (i2 = 1; i2 < SYNC_N; i2 = i2 + 1)
g_sr[i2] <= g_sr[i2-1];
end
end
// Gray to binary: each output bit is the XOR of all Gray bits at or above it.
// Combinational and W-1 XORs deep, which is the whole cost of the scheme.
function automatic [W-1:0] gray_to_bin(input [W-1:0] g);
integer k;
begin
gray_to_bin[W-1] = g[W-1];
for (k = W-2; k >= 0; k = k - 1)
gray_to_bin[k] = gray_to_bin[k+1] ^ g[k];
end
endfunction
assign dst_gray_ok = gray_to_bin(g_sr[SYNC_N-1]);
// ---------------------------------------------------------------- scheme 3
// Data plus a flag. The data path has NO synchroniser -- deliberately, and it
// is the point: synchronising it would reintroduce scheme 1's problem. What
// makes it safe is that the source holds it still, and the flag says when.
reg flag_src;
always_ff @(posedge src_clk or negedge src_rst_n) begin
if (!src_rst_n) flag_src <= 1'b0;
else if (src_update) flag_src <= ~flag_src;
end
reg [SYNC_N-1:0] f_sr;
reg f_d;
wire f_arrived = f_sr[SYNC_N-1] ^ f_d;
always_ff @(posedge dst_clk or negedge dst_rst_n) begin
if (!dst_rst_n) begin
f_sr <= {SYNC_N{1'b0}};
f_d <= 1'b0;
dst_flagged <= {W{1'b0}};
dst_flagged_stb <= 1'b0;
end else begin
f_sr <= {f_sr[SYNC_N-2:0], flag_src};
f_d <= f_sr[SYNC_N-1];
dst_flagged_stb <= f_arrived;
if (f_arrived)
dst_flagged <= src_bin; // sampled once, unsynchronised, on trust
end
end
endmodule// spi_cdc_vector.v
//
// Chapter 15.5 -- what a two-flop synchroniser is for, and the thing it is
// routinely used for that it cannot do.
//
// A two-flop synchroniser solves exactly one problem: it takes a signal that may
// be sampled while it is changing and produces a value that has finished settling
// before anything downstream looks at it. That is all. In particular:
//
// IT DOES NOT guarantee WHICH of the two values you get. The first flop may
// resolve to the old one, which is why Chapter 15.3's soft limit exists.
//
// IT DOES NOT convey a pulse. A pulse narrower than a destination clock period
// is lost with two flops exactly as it is with one (Chapter 15.1 measured this).
//
// IT DOES NOT work on a BUS. This is the one that gets built anyway, and this
// file exists to make the failure visible rather than arguable.
//
// WHY N INDEPENDENT SYNCHRONISERS ON AN N-BIT VALUE IS WRONG.
//
// Each bit settles independently. The destination therefore samples a MIXTURE: some
// bits from before the source's update and some from after. The result is a value
// the source NEVER HELD.
//
// A counter going from 0111 to 1000 changes four bits at once. If the destination
// catches two of them, it reads 1111 or 0000 -- both plausible, neither ever true.
// And the damage is unbounded: an address that briefly reads 1111 addresses the
// wrong memory, and no downstream check can tell because the value is well formed.
//
// THE THREE CORRECT ANSWERS, AND WHEN EACH APPLIES.
//
// GRAY CODE for a value that changes by ONE STEP at a time, such as a
// counter or a FIFO pointer. Consecutive Gray codes differ in
// exactly one bit, so a destination that catches the update
// part-way reads either the old value or the new one -- never
// a third. It does NOT work for arbitrary data, because
// arbitrary data does not change by one step.
//
// DATA PLUS A FLAG for arbitrary data. The data crosses with NO synchroniser
// at all and is simply held stable; one flag bit crosses
// through a synchroniser and says when the data is safe to
// sample. This is what Chapter 15.2 does, and the requirement
// it creates is a LIFETIME on the data register.
//
// A FIFO for arbitrary data that keeps coming. Chapter 15.6.
//
// This module builds the wrong scheme and two of the right ones side by side, on
// the same source value, so that the testbench can count how many times each
// destination observed a value the source never held.
//
// WHAT MAKES THE FAILURE SIMULATABLE, WHICH IS THE INTERESTING PART.
//
// Bit-level settling differences are metastability and are not modelled. But the
// same incoherence is produced by SKEW -- the source bits arriving at the
// destination's flops at slightly different times, which is what unequal routing
// does on any real device and is entirely ordinary. So the testbench drives the
// source vector with a small per-bit skew, and that reproduces the failure exactly,
// for a reason a reviewer can point at on a floorplan.
//
// A design that is safe against skew is safe against the settling differences too,
// because both amount to "the bits do not arrive together". A bench that drives the
// vector with all bits changing at the same instant proves nothing at all, and is
// what almost every bench does.
module spi_cdc_vector #(
parameter W = 4, // the width of the value being crossed
parameter SYNC_N = 2
) (
input wire src_clk,
input wire src_rst_n,
input wire [W-1:0] src_bin, // the value, binary, as the source holds it
input wire [W-1:0] src_gray, // the same value, Gray-coded
input wire src_update, // one source cycle: the value just changed
input wire dst_clk,
input wire dst_rst_n,
// --- scheme 1: N INDEPENDENT SYNCHRONISERS ON A BUS. Wrong, and built here
// so that the count of impossible values is a measurement.
output wire [W-1:0] dst_bin_wrong,
// --- scheme 2: GRAY CODE through the same N synchronisers, then decoded.
output wire [W-1:0] dst_gray_ok,
// --- scheme 3: DATA PLUS A FLAG. The data is not synchronised at all.
output reg [W-1:0] dst_flagged,
output reg dst_flagged_stb
);
// ---------------------------------------------------------------- scheme 1
// One synchroniser per bit, which is the natural thing to write and is the
// bug. Nothing here is wrong on its own; the error is applying it to a value
// whose bits have to agree with each other.
reg [W-1:0] w_sr [0:SYNC_N-1];
integer i1;
always @(posedge dst_clk or negedge dst_rst_n) begin
if (!dst_rst_n) begin
for (i1 = 0; i1 < SYNC_N; i1 = i1 + 1)
w_sr[i1] <= {W{1'b0}};
end else begin
w_sr[0] <= src_bin;
for (i1 = 1; i1 < SYNC_N; i1 = i1 + 1)
w_sr[i1] <= w_sr[i1-1];
end
end
assign dst_bin_wrong = w_sr[SYNC_N-1];
// ---------------------------------------------------------------- scheme 2
// The identical structure, on a Gray-coded value. The synchronisers are no
// better; what changed is that consecutive values differ in ONE bit, so
// catching an update part-way yields the old value or the new one.
reg [W-1:0] g_sr [0:SYNC_N-1];
integer i2;
always @(posedge dst_clk or negedge dst_rst_n) begin
if (!dst_rst_n) begin
for (i2 = 0; i2 < SYNC_N; i2 = i2 + 1)
g_sr[i2] <= {W{1'b0}};
end else begin
g_sr[0] <= src_gray;
for (i2 = 1; i2 < SYNC_N; i2 = i2 + 1)
g_sr[i2] <= g_sr[i2-1];
end
end
// Gray to binary: each output bit is the XOR of all Gray bits at or above it.
// Combinational and W-1 XORs deep, which is the whole cost of the scheme.
function [W-1:0] gray_to_bin;
input [W-1:0] g;
integer k;
begin
gray_to_bin[W-1] = g[W-1];
for (k = W-2; k >= 0; k = k - 1)
gray_to_bin[k] = gray_to_bin[k+1] ^ g[k];
end
endfunction
assign dst_gray_ok = gray_to_bin(g_sr[SYNC_N-1]);
// ---------------------------------------------------------------- scheme 3
// Data plus a flag. The data path has NO synchroniser -- deliberately, and it
// is the point: synchronising it would reintroduce scheme 1's problem. What
// makes it safe is that the source holds it still, and the flag says when.
reg flag_src;
always @(posedge src_clk or negedge src_rst_n) begin
if (!src_rst_n) flag_src <= 1'b0;
else if (src_update) flag_src <= ~flag_src;
end
reg [SYNC_N-1:0] f_sr;
reg f_d;
wire f_arrived = f_sr[SYNC_N-1] ^ f_d;
always @(posedge dst_clk or negedge dst_rst_n) begin
if (!dst_rst_n) begin
f_sr <= {SYNC_N{1'b0}};
f_d <= 1'b0;
dst_flagged <= {W{1'b0}};
dst_flagged_stb <= 1'b0;
end else begin
f_sr <= {f_sr[SYNC_N-2:0], flag_src};
f_d <= f_sr[SYNC_N-1];
dst_flagged_stb <= f_arrived;
if (f_arrived)
dst_flagged <= src_bin; // sampled once, unsynchronised, on trust
end
end
endmodule-- spi_cdc_vector.vhd
--
-- Chapter 15.5 -- what a two-flop synchroniser is for, and the thing it is
-- routinely used for that it cannot do.
--
-- A two-flop synchroniser solves exactly one problem: it takes a signal that may
-- be sampled while it is changing and produces a value that has finished settling
-- before anything downstream looks at it. That is all. In particular:
--
-- IT DOES NOT guarantee WHICH of the two values you get. The first flop may
-- resolve to the old one, which is why Chapter 15.3's soft limit exists.
--
-- IT DOES NOT convey a pulse. A pulse narrower than a destination clock period
-- is lost with two flops exactly as it is with one (Chapter 15.1 measured this).
--
-- IT DOES NOT work on a BUS. This is the one that gets built anyway, and this
-- file exists to make the failure visible rather than arguable.
--
-- WHY N INDEPENDENT SYNCHRONISERS ON AN N-BIT VALUE IS WRONG.
--
-- Each bit settles independently. The destination therefore samples a MIXTURE: some
-- bits from before the source's update and some from after. The result is a value
-- the source NEVER HELD.
--
-- A counter going from 0111 to 1000 changes four bits at once. If the destination
-- catches two of them, it reads 1111 or 0000 -- both plausible, neither ever true.
-- And the damage is unbounded: an address that briefly reads 1111 addresses the
-- wrong memory, and no downstream check can tell because the value is well formed.
--
-- THE THREE CORRECT ANSWERS, AND WHEN EACH APPLIES.
--
-- GRAY CODE for a value that changes by ONE STEP at a time, such as a
-- counter or a FIFO pointer. Consecutive Gray codes differ in
-- exactly one bit, so a destination that catches the update
-- part-way reads either the old value or the new one -- never
-- a third. It does NOT work for arbitrary data, because
-- arbitrary data does not change by one step.
--
-- DATA PLUS A FLAG for arbitrary data. The data crosses with NO synchroniser
-- at all and is simply held stable; one flag bit crosses
-- through a synchroniser and says when the data is safe to
-- sample. This is what Chapter 15.2 does, and the requirement
-- it creates is a LIFETIME on the data register.
--
-- A FIFO for arbitrary data that keeps coming. Chapter 15.6.
--
-- This module builds the wrong scheme and two of the right ones side by side, on
-- the same source value, so that the testbench can count how many times each
-- destination observed a value the source never held.
--
-- WHAT MAKES THE FAILURE SIMULATABLE, WHICH IS THE INTERESTING PART.
--
-- Bit-level settling differences are metastability and are not modelled. But the
-- same incoherence is produced by SKEW -- the source bits arriving at the
-- destination's flops at slightly different times, which is what unequal routing
-- does on any real device and is entirely ordinary. So the testbench drives the
-- source vector with a small per-bit skew, and that reproduces the failure exactly,
-- for a reason a reviewer can point at on a floorplan.
--
-- A design that is safe against skew is safe against the settling differences too,
-- because both amount to "the bits do not arrive together". A bench that drives the
-- vector with all bits changing at the same instant proves nothing at all, and is
-- what almost every bench does.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_cdc_vector is
generic (
W : positive := 8; -- the width of the value being crossed
SYNC_N : positive := 2
);
port (
src_clk : in std_logic;
src_rst_n : in std_logic;
src_bin : in std_logic_vector(W - 1 downto 0);
src_gray : in std_logic_vector(W - 1 downto 0);
src_update : in std_logic;
dst_clk : in std_logic;
dst_rst_n : in std_logic;
-- scheme 1: N INDEPENDENT SYNCHRONISERS ON A BUS. Wrong, and built here
-- so that the count of impossible values is a measurement.
dst_bin_wrong : out std_logic_vector(W - 1 downto 0);
-- scheme 2: GRAY CODE through the same N synchronisers, then decoded.
dst_gray_ok : out std_logic_vector(W - 1 downto 0);
-- scheme 3: DATA PLUS A FLAG. The data is not synchronised at all.
dst_flagged : out std_logic_vector(W - 1 downto 0);
dst_flagged_stb : out std_logic
);
end entity;
architecture rtl of spi_cdc_vector is
type vec_arr is array (0 to SYNC_N - 1) of std_logic_vector(W - 1 downto 0);
signal w_sr : vec_arr := (others => (others => '0'));
signal g_sr : vec_arr := (others => (others => '0'));
signal flag_src : std_logic := '0';
signal f_sr : std_logic_vector(SYNC_N - 1 downto 0) := (others => '0');
signal f_d : std_logic := '0';
signal f_arrived : std_logic;
signal flagged_r : std_logic_vector(W - 1 downto 0) := (others => '0');
signal flag_stb_r : std_logic := '0';
-- Gray to binary: each output bit is the XOR of all Gray bits at or above it.
-- Combinational and W-1 XORs deep, which is the whole cost of the scheme.
function gray_to_bin(g : std_logic_vector) return std_logic_vector is
variable r : std_logic_vector(g'range);
begin
r(r'high) := g(g'high);
for k in g'high - 1 downto 0 loop
r(k) := r(k + 1) xor g(k);
end loop;
return r;
end function;
begin
-- ---------------------------------------------------------------- scheme 1
-- One synchroniser per bit, which is the natural thing to write and is the bug.
-- Nothing here is wrong on its own; the error is applying it to a value whose
-- bits have to agree with each other.
wrong : process (dst_clk, dst_rst_n)
begin
if dst_rst_n = '0' then
w_sr <= (others => (others => '0'));
elsif rising_edge(dst_clk) then
w_sr(0) <= src_bin;
for i in 1 to SYNC_N - 1 loop
w_sr(i) <= w_sr(i - 1);
end loop;
end if;
end process;
dst_bin_wrong <= w_sr(SYNC_N - 1);
-- ---------------------------------------------------------------- scheme 2
-- The identical structure, on a Gray-coded value. The synchronisers are no
-- better; what changed is that consecutive values differ in ONE bit.
grayx : process (dst_clk, dst_rst_n)
begin
if dst_rst_n = '0' then
g_sr <= (others => (others => '0'));
elsif rising_edge(dst_clk) then
g_sr(0) <= src_gray;
for i in 1 to SYNC_N - 1 loop
g_sr(i) <= g_sr(i - 1);
end loop;
end if;
end process;
dst_gray_ok <= gray_to_bin(g_sr(SYNC_N - 1));
-- ---------------------------------------------------------------- scheme 3
-- Data plus a flag. The data path has NO synchroniser -- deliberately, and it
-- is the point: synchronising it would reintroduce scheme 1's problem. What
-- makes it safe is that the source holds it still, and the flag says when.
srcflag : process (src_clk, src_rst_n)
begin
if src_rst_n = '0' then
flag_src <= '0';
elsif rising_edge(src_clk) then
if src_update = '1' then
flag_src <= not flag_src;
end if;
end if;
end process;
f_arrived <= f_sr(SYNC_N - 1) xor f_d;
dst_flagged <= flagged_r;
dst_flagged_stb <= flag_stb_r;
dstflag : process (dst_clk, dst_rst_n)
begin
if dst_rst_n = '0' then
f_sr <= (others => '0');
f_d <= '0';
flagged_r <= (others => '0');
flag_stb_r <= '0';
elsif rising_edge(dst_clk) then
f_sr <= f_sr(SYNC_N - 2 downto 0) & flag_src;
f_d <= f_sr(SYNC_N - 1);
flag_stb_r <= f_arrived;
if f_arrived = '1' then
flagged_r <= src_bin; -- sampled once, unsynchronised, on trust
end if;
end if;
end process;
end architecture;The testbench
Two phases, one per regime, with the recency window as the impossibility test and per-bit skew on the seam.
// spi_cdc_vector_tb.sv
//
// The measurement is "how many times did the destination observe a value the
// source never held?", and it is made three times on the same stimulus.
//
// HOW THE BENCH KNOWS WHAT THE SOURCE HELD, and why the obvious version of this
// does not work. The first attempt kept a bitmap of every value the source had EVER
// taken -- and a counter passes through every value, so within sixteen cycles the
// bitmap was all ones and no observation could ever be impossible. The bench
// reported zero faults on a scheme that is definitely broken, which is a good
// reminder that a checker returning the answer you hoped for is not evidence.
//
// The correct question is about RECENCY, not existence: the destination is reading a
// value through a synchroniser, so it must see one the source held WITHIN THE LAST
// FEW SOURCE CYCLES. The bench keeps a short history of the source's recent values
// and any observation outside that window is impossible -- exactly, not
// probabilistically, because the source cannot have held it.
//
// The counter is eight bits wide so that the 256-value space is far larger than the
// six-deep window; a mixture of two adjacent values then lands outside the window
// almost always rather than sometimes.
//
// HOW THE FAILURE IS MADE TO HAPPEN. The source bits are driven to the DUT through
// per-bit skew -- bit k is delayed by k times a small delta -- which is what
// unequal routing does on any real device. That is the honest way to exhibit
// incoherence in a simulator: bit-level settling differences are metastability and
// are not modelled, but skew is ordinary, it produces the identical failure, and a
// design that survives skew survives the settling case for the same reason.
//
// A bench that changes all bits at the same instant finds nothing wrong with scheme
// 1, which is why scheme 1 keeps being built.
`timescale 1ns/1ps
module spi_cdc_vector_tb;
localparam int W = 8;
localparam int SYNC_N = 2;
localparam int NVAL = 1 << W;
// How far back a legitimate observation may come from: the synchroniser is
// SYNC_N deep and the two clocks are close in period, so a value three or four
// source cycles old is plausible. Six is generous, which matters -- a window
// that is too tight reports correct designs as broken.
localparam int HIST = 6;
// Two unrelated clocks. The ratio is deliberately not an integer, so that the
// destination's sampling instant walks through the source's period rather than
// sitting at one point in it.
reg src_clk = 1'b0;
reg dst_clk = 1'b0;
always #7 src_clk = ~src_clk; // 14 ns
always #5 dst_clk = ~dst_clk; // 10 ns
reg src_rst_n = 1'b1;
reg dst_rst_n = 1'b1;
// The source: a plain counter, and its Gray encoding.
reg [W-1:0] cnt = {W{1'b0}};
reg update = 1'b0;
wire [W-1:0] cnt_gray = cnt ^ (cnt >> 1);
// --- the skew injection -------------------------------------------------
// Each bit of the value reaches the DUT through its own delay. This is the
// whole reason the bench can see the bug, and it is a model of routing rather
// than of metastability -- which is what makes it defensible.
localparam time SKEW = 1;
wire [W-1:0] bin_skewed;
wire [W-1:0] gray_skewed;
genvar gi;
generate
for (gi = 0; gi < W; gi = gi + 1) begin : g_skew
// Bit 0 arrives on time, bit W-1 arrives (W-1)*SKEW late. A real
// device's spread is smaller and less orderly; an orderly spread makes
// the result reproducible, which a chapter needs.
assign #(gi * SKEW) bin_skewed[gi] = cnt[gi];
assign #(gi * SKEW) gray_skewed[gi] = cnt_gray[gi];
end
endgenerate
wire [W-1:0] dst_bin_wrong, dst_gray_ok, dst_flagged;
wire dst_flagged_stb;
spi_cdc_vector #(.W(W), .SYNC_N(SYNC_N)) dut (
.src_clk(src_clk), .src_rst_n(src_rst_n),
.src_bin(bin_skewed), .src_gray(gray_skewed), .src_update(update),
.dst_clk(dst_clk), .dst_rst_n(dst_rst_n),
.dst_bin_wrong(dst_bin_wrong), .dst_gray_ok(dst_gray_ok),
.dst_flagged(dst_flagged), .dst_flagged_stb(dst_flagged_stb)
);
integer errors = 0;
initial begin
#2_000_000;
$display("FAIL: the simulation did not finish within its time limit");
$finish;
end
// The source's recent values, newest first. Maintained in SOURCE time, which is
// the only place the source's value means anything.
reg [W-1:0] hist [0:HIST-1];
function automatic is_recent(input [W-1:0] v);
integer j;
begin
is_recent = 1'b0;
for (j = 0; j < HIST; j = j + 1)
if (hist[j] == v) is_recent = 1'b1;
end
endfunction
// Counts of destination observations that were NOT values the source held.
integer bad_bin = 0, bad_gray = 0, bad_flag = 0;
// And of observations in total, so the bad counts can be read as a proportion.
integer obs_bin = 0, obs_gray = 0, obs_flag = 0;
// Distinct impossible values seen on scheme 1, so the chapter can say how many
// different lies were told rather than only how often.
reg seen_bad [0:NVAL-1];
integer distinct_bad = 0;
integer hj;
always @(posedge src_clk) if (src_rst_n) begin
for (hj = HIST-1; hj > 0; hj = hj - 1) hist[hj] = hist[hj-1];
hist[0] = cnt;
end
// The destination observations. Sampled on the destination clock, which is the
// only place they exist.
always @(posedge dst_clk) if (dst_rst_n && src_rst_n) begin
obs_bin = obs_bin + 1;
if (!is_recent(dst_bin_wrong)) begin
bad_bin = bad_bin + 1;
if (!seen_bad[dst_bin_wrong]) begin
seen_bad[dst_bin_wrong] = 1'b1;
distinct_bad = distinct_bad + 1;
end
end
obs_gray = obs_gray + 1;
if (!is_recent(dst_gray_ok)) bad_gray = bad_gray + 1;
if (dst_flagged_stb) begin
obs_flag = obs_flag + 1;
if (!is_recent(dst_flagged)) bad_flag = bad_flag + 1;
end
end
task automatic clear_stats;
integer c;
begin
bad_bin = 0; bad_gray = 0; bad_flag = 0;
obs_bin = 0; obs_gray = 0; obs_flag = 0;
distinct_bad = 0;
for (c = 0; c < NVAL; c = c + 1) seen_bad[c] = 1'b0;
end
endtask
// Runs the counter for `n` updates, spaced `gap` source cycles apart. `gap` is
// the whole difference between the two phases: at 1 the source never holds
// still, and at 8 it holds still for far longer than the synchroniser latency.
task automatic run_counter(input integer n, input integer gap);
integer i, g;
begin
for (i = 0; i < n; i = i + 1) begin
@(negedge src_clk);
update = 1'b1;
cnt = cnt + 1'b1;
@(negedge src_clk);
update = 1'b0;
for (g = 1; g < gap; g = g + 1) @(negedge src_clk);
end
repeat (20) @(posedge dst_clk);
end
endtask
integer k;
integer p1_bin, p1_gray, p1_flag, p1_distinct, p1_obs;
initial begin
for (k = 0; k < NVAL; k = k + 1) seen_bad[k] = 1'b0;
for (k = 0; k < HIST; k = k + 1) hist[k] = {W{1'b0}};
src_rst_n = 1'b1; dst_rst_n = 1'b1;
repeat (2) @(posedge dst_clk);
src_rst_n = 1'b0; dst_rst_n = 1'b0;
repeat (4) @(posedge dst_clk);
src_rst_n = 1'b1; dst_rst_n = 1'b1;
repeat (4) @(posedge dst_clk);
// Reset the bookkeeping now that the design is out of reset, so that the
// reset value is not counted as a source observation.
bad_bin = 0; bad_gray = 0; bad_flag = 0;
obs_bin = 0; obs_gray = 0; obs_flag = 0;
distinct_bad = 0;
for (k = 0; k < NVAL; k = k + 1) seen_bad[k] = 1'b0;
$display(" an %0d-bit counter crossing two unrelated clocks (14 ns source, 10 ns destination) with up to %0d ns of per-bit skew; an observation is IMPOSSIBLE if the source did not hold it within the last %0d source cycles",
W, (W-1)*SKEW, HIST);
// =============================================================
// PHASE 1: the source NEVER HOLDS STILL -- one update per source cycle,
// which is the hardest case for every scheme and the normal case for a
// FIFO pointer.
// =============================================================
run_counter(400, 1);
p1_bin = bad_bin; p1_gray = bad_gray; p1_flag = bad_flag;
p1_distinct = distinct_bad; p1_obs = obs_bin;
$display(" PHASE 1 -- one update per source cycle: the source never holds still");
$display(" scheme observations impossible values");
$display(" 1 N synchronisers on a bus %12d %17d", obs_bin, bad_bin);
$display(" 2 Gray code, then decoded %12d %17d", obs_gray, bad_gray);
$display(" 3 data plus a flag %12d %17d", obs_flag, bad_flag);
// 1. SCHEME 1 OBSERVES VALUES THAT NEVER EXISTED. Without this the chapter
// has an argument and no evidence, and the argument is one that every
// reviewer has heard and some do not believe.
if (bad_bin == 0) begin
$display(" FAIL: N independent synchronisers on a %0d-bit bus produced no impossible value across %0d observations -- with up to %0d ns of skew they must, and a bench that finds none is either driving all bits at the same instant or asking whether a value ever existed rather than whether it existed recently",
W, obs_bin, (W-1)*SKEW);
errors = errors + 1;
end
// 2. GRAY CODE NEVER DOES. This is the positive claim and it is exact: a
// one-bit-at-a-time encoding cannot produce a third value, so the count
// must be zero rather than small.
if (bad_gray != 0) begin
$display(" FAIL: the Gray-coded crossing observed %0d impossible values -- consecutive Gray codes differ in one bit, so catching an update part-way must yield the old value or the new one",
bad_gray);
errors = errors + 1;
end
// 3. AND DATA PLUS A FLAG FAILS HERE TOO, which was not the expected
// result and is the more useful one. Its precondition is that the source
// HOLDS THE DATA STILL until the flag has been observed -- and a source
// updating every cycle violates that, so the destination samples a
// register that is mid-change and reads a mixture, exactly as scheme 1
// does. The flag is not the weak part; the holding is.
//
// This is Chapter 15.2's lifetime requirement arriving from the other
// direction, and it is why Chapter 15.6 has to build a FIFO: an SPI
// slave's source cannot be told to hold still, because SCLK does not
// stop.
if (bad_flag == 0) begin
$display(" FAIL: data plus a flag should ALSO observe impossible values when the source never holds still -- its correctness depends on the data being stable when sampled, not on the flag");
errors = errors + 1;
end
$display(" data plus a flag observed %0d impossible values as well, because its precondition is that the source HOLDS THE DATA STILL and a source updating every cycle does not -- the flag is not the weak part, the holding is",
bad_flag);
// 4. AND THE LIES WERE VARIED, not one repeated glitch. A single impossible
// value could be dismissed as an artefact of one particular alignment;
// several distinct ones cannot.
if (p1_distinct < 2) begin
$display(" FAIL: only %0d distinct impossible value(s) -- one could be an alignment artefact",
p1_distinct);
errors = errors + 1;
end
$display(" scheme 1 produced %0d distinct impossible values out of the %0d the counter can hold, so it is not one unlucky alignment",
p1_distinct, NVAL);
// =============================================================
// PHASE 2: the source HOLDS STILL between updates -- one update every eight
// source cycles, which is far longer than the synchroniser latency.
// =============================================================
clear_stats();
run_counter(60, 8);
$display(" PHASE 2 -- one update every eight source cycles: the source holds still");
$display(" scheme observations impossible values");
$display(" 1 N synchronisers on a bus %12d %17d", obs_bin, bad_bin);
$display(" 2 Gray code, then decoded %12d %17d", obs_gray, bad_gray);
$display(" 3 data plus a flag %12d %17d", obs_flag, bad_flag);
// 4b. WITH THE SOURCE HOLDING STILL, DATA PLUS A FLAG IS EXACT. Its
// precondition is met, and it carries arbitrary data rather than only a
// counter -- which is what makes it the general answer and Gray the
// special one.
if (bad_flag != 0) begin
$display(" FAIL: with the source holding still, data plus a flag observed %0d impossible values",
bad_flag);
errors = errors + 1;
end
if (bad_gray != 0) begin
$display(" FAIL: the Gray crossing observed %0d impossible values with the source holding still",
bad_gray);
errors = errors + 1;
end
// 4c. AND SCHEME 1 IS STILL WRONG, just less often -- which is the reason it
// reaches silicon. Slowing the source reduces the error RATE without
// changing the error, so a bench that runs a slow source and counts
// failures per thousand observations concludes the crossing is fine.
if (bad_bin == 0) begin
$display(" FAIL: scheme 1 should still be wrong with a slow source -- skew does not care how often the value changes");
errors = errors + 1;
end
$display(" scheme 1 was still wrong with the source holding still, %0d times in %0d observations against %0d in %0d when it never held still -- slowing the source reduces the RATE and not the fault, which is exactly how this crossing survives testing and reaches silicon",
bad_bin, obs_bin, p1_bin, p1_obs);
// 5. THE OBSERVATION THAT MAKES SCHEME 1 DANGEROUS RATHER THAN MERELY
// WRONG: every impossible value is a WELL-FORMED one. There is no bit
// pattern a downstream consumer could reject, so nothing downstream can
// detect the fault -- which is the whole reason it survives to silicon.
$display(" every impossible value was a legal %0d-bit pattern, so no downstream check could reject one -- the fault is undetectable by anything except the crossing being built correctly in the first place",
W);
// 6. AND THE COST OF THE CORRECT SCHEMES, so the chapter is not only a
// warning. Gray costs an encoder, a decoder of W-1 XORs and nothing
// else; data-plus-flag costs one synchronised bit and a LIFETIME
// requirement on the data register, which is Chapter 15.2's finding.
if (obs_flag == 0) begin
$display(" FAIL: the data-plus-flag scheme delivered nothing at all");
errors = errors + 1;
end
$display(" the flagged scheme delivered %0d observations rather than one per destination cycle, because it only publishes when the flag arrives -- it is exact, it is slower, and it needs the source to hold still, and those three together are the trade",
obs_flag);
// The exact counts are phase- and simulator-dependent, and deliberately not
// asserted: which update a destination edge happens to land inside depends
// on the relationship between two clocks that have none, so the three
// language benches report different totals from the same design. What is
// asserted is the SHAPE -- scheme 1 non-zero in both regimes, Gray zero in
// both, data-plus-flag non-zero only when the source never holds still --
// and that is identical in all three, because it follows from the encoding
// rather than from the alignment.
if (errors == 0)
$display("PASS: a two-flop synchroniser produces a SETTLED value and nothing else -- it does not choose which of two values you get, it does not convey a pulse, and on a BUS it is wrong, because each bit settles independently and the destination reads a mixture of bits from before and after the update. Driven with per-bit skew of the kind ordinary routing produces, N independent synchronisers on an 8-bit counter observed impossible values repeatedly and in many distinct forms, every one a well-formed pattern that no downstream check could reject -- and it was still wrong when the source was slowed to update every eighth cycle, just less often, which is precisely how this crossing passes a regression and reaches silicon. A GRAY-CODED crossing through the identical synchronisers observed exactly ZERO in both regimes, because consecutive Gray codes differ in one bit so catching an update part-way yields the old value or the new one -- and that holds only for a value that changes by one step, which suits a counter or a FIFO pointer and not arbitrary data. DATA PLUS A FLAG observed zero only in the second regime, and that was the useful surprise: it fails exactly as badly as scheme 1 when the source never holds still, because its correctness rests on the data being STABLE WHEN SAMPLED rather than on the flag -- Chapter 15.2s lifetime requirement arriving from the other direction, and the reason Chapter 15.6 must build a FIFO, since an SPI slaves source cannot be told to hold still while SCLK is running. So the choice is not between synchroniser depths. It is between encodings: change one bit at a time, or hold the value still and cross one bit of handshake, or buffer it");
else
$display("FAIL: %0d error(s)", errors);
$finish;
end
endmodule// spi_cdc_vector_tb.v
//
// The measurement is "how many times did the destination observe a value the
// source never held?", and it is made three times on the same stimulus.
//
// HOW THE BENCH KNOWS WHAT THE SOURCE HELD, and why the obvious version of this
// does not work. The first attempt kept a bitmap of every value the source had EVER
// taken -- and a counter passes through every value, so within sixteen cycles the
// bitmap was all ones and no observation could ever be impossible. The bench
// reported zero faults on a scheme that is definitely broken, which is a good
// reminder that a checker returning the answer you hoped for is not evidence.
//
// The correct question is about RECENCY, not existence: the destination is reading a
// value through a synchroniser, so it must see one the source held WITHIN THE LAST
// FEW SOURCE CYCLES. The bench keeps a short history of the source's recent values
// and any observation outside that window is impossible -- exactly, not
// probabilistically, because the source cannot have held it.
//
// The counter is eight bits wide so that the 256-value space is far larger than the
// six-deep window; a mixture of two adjacent values then lands outside the window
// almost always rather than sometimes.
//
// HOW THE FAILURE IS MADE TO HAPPEN. The source bits are driven to the DUT through
// per-bit skew -- bit k is delayed by k times a small delta -- which is what
// unequal routing does on any real device. That is the honest way to exhibit
// incoherence in a simulator: bit-level settling differences are metastability and
// are not modelled, but skew is ordinary, it produces the identical failure, and a
// design that survives skew survives the settling case for the same reason.
//
// A bench that changes all bits at the same instant finds nothing wrong with scheme
// 1, which is why scheme 1 keeps being built.
`timescale 1ns/1ps
module spi_cdc_vector_tb;
localparam W = 8;
localparam SYNC_N = 2;
localparam NVAL = 1 << W;
// How far back a legitimate observation may come from: the synchroniser is
// SYNC_N deep and the two clocks are close in period, so a value three or four
// source cycles old is plausible. Six is generous, which matters -- a window
// that is too tight reports correct designs as broken.
localparam HIST = 6;
// Two unrelated clocks. The ratio is deliberately not an integer, so that the
// destination's sampling instant walks through the source's period rather than
// sitting at one point in it.
reg src_clk;
reg dst_clk;
always #7 src_clk = ~src_clk; // 14 ns
always #5 dst_clk = ~dst_clk; // 10 ns
reg src_rst_n;
reg dst_rst_n;
// The source: a plain counter, and its Gray encoding.
reg [W-1:0] cnt;
reg update;
wire [W-1:0] cnt_gray = cnt ^ (cnt >> 1);
// --- the skew injection -------------------------------------------------
// Each bit of the value reaches the DUT through its own delay. This is the
// whole reason the bench can see the bug, and it is a model of routing rather
// than of metastability -- which is what makes it defensible.
localparam time SKEW = 1;
wire [W-1:0] bin_skewed;
wire [W-1:0] gray_skewed;
genvar gi;
generate
for (gi = 0; gi < W; gi = gi + 1) begin : g_skew
// Bit 0 arrives on time, bit W-1 arrives (W-1)*SKEW late. A real
// device's spread is smaller and less orderly; an orderly spread makes
// the result reproducible, which a chapter needs.
assign #(gi * SKEW) bin_skewed[gi] = cnt[gi];
assign #(gi * SKEW) gray_skewed[gi] = cnt_gray[gi];
end
endgenerate
wire [W-1:0] dst_bin_wrong, dst_gray_ok, dst_flagged;
wire dst_flagged_stb;
spi_cdc_vector #(.W(W), .SYNC_N(SYNC_N)) dut (
.src_clk(src_clk), .src_rst_n(src_rst_n),
.src_bin(bin_skewed), .src_gray(gray_skewed), .src_update(update),
.dst_clk(dst_clk), .dst_rst_n(dst_rst_n),
.dst_bin_wrong(dst_bin_wrong), .dst_gray_ok(dst_gray_ok),
.dst_flagged(dst_flagged), .dst_flagged_stb(dst_flagged_stb)
);
integer errors;
initial begin
#2_000_000;
$display("FAIL: the simulation did not finish within its time limit");
$finish;
end
// The source's recent values, newest first. Maintained in SOURCE time, which is
// the only place the source's value means anything.
reg [W-1:0] hist [0:HIST-1];
function is_recent;
input [W-1:0] v;
integer j;
begin
is_recent = 1'b0;
for (j = 0; j < HIST; j = j + 1)
if (hist[j] == v) is_recent = 1'b1;
end
endfunction
// Counts of destination observations that were NOT values the source held.
integer bad_bin, bad_gray, bad_flag;
// And of observations in total, so the bad counts can be read as a proportion.
integer obs_bin, obs_gray, obs_flag;
// Distinct impossible values seen on scheme 1, so the chapter can say how many
// different lies were told rather than only how often.
reg seen_bad [0:NVAL-1];
integer distinct_bad;
integer hj;
always @(posedge src_clk) if (src_rst_n) begin
for (hj = HIST-1; hj > 0; hj = hj - 1) hist[hj] = hist[hj-1];
hist[0] = cnt;
end
// The destination observations. Sampled on the destination clock, which is the
// only place they exist.
always @(posedge dst_clk) if (dst_rst_n && src_rst_n) begin
obs_bin = obs_bin + 1;
if (!is_recent(dst_bin_wrong)) begin
bad_bin = bad_bin + 1;
if (!seen_bad[dst_bin_wrong]) begin
seen_bad[dst_bin_wrong] = 1'b1;
distinct_bad = distinct_bad + 1;
end
end
obs_gray = obs_gray + 1;
if (!is_recent(dst_gray_ok)) bad_gray = bad_gray + 1;
if (dst_flagged_stb) begin
obs_flag = obs_flag + 1;
if (!is_recent(dst_flagged)) bad_flag = bad_flag + 1;
end
end
task clear_stats;
integer c;
begin
bad_bin = 0; bad_gray = 0; bad_flag = 0;
obs_bin = 0; obs_gray = 0; obs_flag = 0;
distinct_bad = 0;
for (c = 0; c < NVAL; c = c + 1) seen_bad[c] = 1'b0;
end
endtask
// Runs the counter for `n` updates, spaced `gap` source cycles apart. `gap` is
// the whole difference between the two phases: at 1 the source never holds
// still, and at 8 it holds still for far longer than the synchroniser latency.
task run_counter;
input integer n;
input integer gap;
integer i, g;
begin
for (i = 0; i < n; i = i + 1) begin
@(negedge src_clk);
update = 1'b1;
cnt = cnt + 1'b1;
@(negedge src_clk);
update = 1'b0;
for (g = 1; g < gap; g = g + 1) @(negedge src_clk);
end
repeat (20) @(posedge dst_clk);
end
endtask
integer k;
integer p1_bin, p1_gray, p1_flag, p1_distinct, p1_obs;
initial begin
for (k = 0; k < NVAL; k = k + 1) seen_bad[k] = 1'b0;
for (k = 0; k < HIST; k = k + 1) hist[k] = {W{1'b0}};
src_rst_n = 1'b1; dst_rst_n = 1'b1;
repeat (2) @(posedge dst_clk);
src_rst_n = 1'b0; dst_rst_n = 1'b0;
repeat (4) @(posedge dst_clk);
src_rst_n = 1'b1; dst_rst_n = 1'b1;
repeat (4) @(posedge dst_clk);
// Reset the bookkeeping now that the design is out of reset, so that the
// reset value is not counted as a source observation.
bad_bin = 0; bad_gray = 0; bad_flag = 0;
obs_bin = 0; obs_gray = 0; obs_flag = 0;
distinct_bad = 0;
for (k = 0; k < NVAL; k = k + 1) seen_bad[k] = 1'b0;
$display(" an %0d-bit counter crossing two unrelated clocks (14 ns source, 10 ns destination) with up to %0d ns of per-bit skew; an observation is IMPOSSIBLE if the source did not hold it within the last %0d source cycles",
W, (W-1)*SKEW, HIST);
// =============================================================
// PHASE 1: the source NEVER HOLDS STILL -- one update per source cycle,
// which is the hardest case for every scheme and the normal case for a
// FIFO pointer.
// =============================================================
run_counter(400, 1);
p1_bin = bad_bin; p1_gray = bad_gray; p1_flag = bad_flag;
p1_distinct = distinct_bad; p1_obs = obs_bin;
$display(" PHASE 1 -- one update per source cycle: the source never holds still");
$display(" scheme observations impossible values");
$display(" 1 N synchronisers on a bus %12d %17d", obs_bin, bad_bin);
$display(" 2 Gray code, then decoded %12d %17d", obs_gray, bad_gray);
$display(" 3 data plus a flag %12d %17d", obs_flag, bad_flag);
// 1. SCHEME 1 OBSERVES VALUES THAT NEVER EXISTED. Without this the chapter
// has an argument and no evidence, and the argument is one that every
// reviewer has heard and some do not believe.
if (bad_bin == 0) begin
$display(" FAIL: N independent synchronisers on a %0d-bit bus produced no impossible value across %0d observations -- with up to %0d ns of skew they must, and a bench that finds none is either driving all bits at the same instant or asking whether a value ever existed rather than whether it existed recently",
W, obs_bin, (W-1)*SKEW);
errors = errors + 1;
end
// 2. GRAY CODE NEVER DOES. This is the positive claim and it is exact: a
// one-bit-at-a-time encoding cannot produce a third value, so the count
// must be zero rather than small.
if (bad_gray != 0) begin
$display(" FAIL: the Gray-coded crossing observed %0d impossible values -- consecutive Gray codes differ in one bit, so catching an update part-way must yield the old value or the new one",
bad_gray);
errors = errors + 1;
end
// 3. AND DATA PLUS A FLAG FAILS HERE TOO, which was not the expected
// result and is the more useful one. Its precondition is that the source
// HOLDS THE DATA STILL until the flag has been observed -- and a source
// updating every cycle violates that, so the destination samples a
// register that is mid-change and reads a mixture, exactly as scheme 1
// does. The flag is not the weak part; the holding is.
//
// This is Chapter 15.2's lifetime requirement arriving from the other
// direction, and it is why Chapter 15.6 has to build a FIFO: an SPI
// slave's source cannot be told to hold still, because SCLK does not
// stop.
if (bad_flag == 0) begin
$display(" FAIL: data plus a flag should ALSO observe impossible values when the source never holds still -- its correctness depends on the data being stable when sampled, not on the flag");
errors = errors + 1;
end
$display(" data plus a flag observed %0d impossible values as well, because its precondition is that the source HOLDS THE DATA STILL and a source updating every cycle does not -- the flag is not the weak part, the holding is",
bad_flag);
// 4. AND THE LIES WERE VARIED, not one repeated glitch. A single impossible
// value could be dismissed as an artefact of one particular alignment;
// several distinct ones cannot.
if (p1_distinct < 2) begin
$display(" FAIL: only %0d distinct impossible value(s) -- one could be an alignment artefact",
p1_distinct);
errors = errors + 1;
end
$display(" scheme 1 produced %0d distinct impossible values out of the %0d the counter can hold, so it is not one unlucky alignment",
p1_distinct, NVAL);
// =============================================================
// PHASE 2: the source HOLDS STILL between updates -- one update every eight
// source cycles, which is far longer than the synchroniser latency.
// =============================================================
clear_stats();
run_counter(60, 8);
$display(" PHASE 2 -- one update every eight source cycles: the source holds still");
$display(" scheme observations impossible values");
$display(" 1 N synchronisers on a bus %12d %17d", obs_bin, bad_bin);
$display(" 2 Gray code, then decoded %12d %17d", obs_gray, bad_gray);
$display(" 3 data plus a flag %12d %17d", obs_flag, bad_flag);
// 4b. WITH THE SOURCE HOLDING STILL, DATA PLUS A FLAG IS EXACT. Its
// precondition is met, and it carries arbitrary data rather than only a
// counter -- which is what makes it the general answer and Gray the
// special one.
if (bad_flag != 0) begin
$display(" FAIL: with the source holding still, data plus a flag observed %0d impossible values",
bad_flag);
errors = errors + 1;
end
if (bad_gray != 0) begin
$display(" FAIL: the Gray crossing observed %0d impossible values with the source holding still",
bad_gray);
errors = errors + 1;
end
// 4c. AND SCHEME 1 IS STILL WRONG, just less often -- which is the reason it
// reaches silicon. Slowing the source reduces the error RATE without
// changing the error, so a bench that runs a slow source and counts
// failures per thousand observations concludes the crossing is fine.
if (bad_bin == 0) begin
$display(" FAIL: scheme 1 should still be wrong with a slow source -- skew does not care how often the value changes");
errors = errors + 1;
end
$display(" scheme 1 was still wrong with the source holding still, %0d times in %0d observations against %0d in %0d when it never held still -- slowing the source reduces the RATE and not the fault, which is exactly how this crossing survives testing and reaches silicon",
bad_bin, obs_bin, p1_bin, p1_obs);
// 5. THE OBSERVATION THAT MAKES SCHEME 1 DANGEROUS RATHER THAN MERELY
// WRONG: every impossible value is a WELL-FORMED one. There is no bit
// pattern a downstream consumer could reject, so nothing downstream can
// detect the fault -- which is the whole reason it survives to silicon.
$display(" every impossible value was a legal %0d-bit pattern, so no downstream check could reject one -- the fault is undetectable by anything except the crossing being built correctly in the first place",
W);
// 6. AND THE COST OF THE CORRECT SCHEMES, so the chapter is not only a
// warning. Gray costs an encoder, a decoder of W-1 XORs and nothing
// else; data-plus-flag costs one synchronised bit and a LIFETIME
// requirement on the data register, which is Chapter 15.2's finding.
if (obs_flag == 0) begin
$display(" FAIL: the data-plus-flag scheme delivered nothing at all");
errors = errors + 1;
end
$display(" the flagged scheme delivered %0d observations rather than one per destination cycle, because it only publishes when the flag arrives -- it is exact, it is slower, and it needs the source to hold still, and those three together are the trade",
obs_flag);
// The exact counts are phase- and simulator-dependent, and deliberately not
// asserted: which update a destination edge happens to land inside depends
// on the relationship between two clocks that have none, so the three
// language benches report different totals from the same design. What is
// asserted is the SHAPE -- scheme 1 non-zero in both regimes, Gray zero in
// both, data-plus-flag non-zero only when the source never holds still --
// and that is identical in all three, because it follows from the encoding
// rather than from the alignment.
if (errors == 0)
$display("PASS: a two-flop synchroniser produces a SETTLED value and nothing else -- it does not choose which of two values you get, it does not convey a pulse, and on a BUS it is wrong, because each bit settles independently and the destination reads a mixture of bits from before and after the update. Driven with per-bit skew of the kind ordinary routing produces, N independent synchronisers on an 8-bit counter observed impossible values repeatedly and in many distinct forms, every one a well-formed pattern that no downstream check could reject -- and it was still wrong when the source was slowed to update every eighth cycle, just less often, which is precisely how this crossing passes a regression and reaches silicon. A GRAY-CODED crossing through the identical synchronisers observed exactly ZERO in both regimes, because consecutive Gray codes differ in one bit so catching an update part-way yields the old value or the new one -- and that holds only for a value that changes by one step, which suits a counter or a FIFO pointer and not arbitrary data. DATA PLUS A FLAG observed zero only in the second regime, and that was the useful surprise: it fails exactly as badly as scheme 1 when the source never holds still, because its correctness rests on the data being STABLE WHEN SAMPLED rather than on the flag -- Chapter 15.2s lifetime requirement arriving from the other direction, and the reason Chapter 15.6 must build a FIFO, since an SPI slaves source cannot be told to hold still while SCLK is running. So the choice is not between synchroniser depths. It is between encodings: change one bit at a time, or hold the value still and cross one bit of handshake, or buffer it");
else
$display("FAIL: %0d error(s)", errors);
$finish;
end
initial begin
bad_bin = 0;
bad_gray = 0;
bad_flag = 0;
obs_bin = 0;
obs_gray = 0;
obs_flag = 0;
src_clk = 1'b0;
dst_clk = 1'b0;
src_rst_n = 1'b1;
dst_rst_n = 1'b1;
cnt = {W{1'b0}};
update = 1'b0;
errors = 0;
distinct_bad = 0;
end
endmodule-- spi_cdc_vector_tb.vhd
--
-- The measurement is "how many times did the destination observe a value the
-- source never held?", and it is made three times on the same stimulus.
--
-- HOW THE BENCH KNOWS WHAT THE SOURCE HELD, and why the obvious version of this
-- does not work. The first attempt kept a bitmap of every value the source had EVER
-- taken -- and a counter passes through every value, so within sixteen cycles the
-- bitmap was all ones and no observation could ever be impossible. The bench
-- reported zero faults on a scheme that is definitely broken, which is a good
-- reminder that a checker returning the answer you hoped for is not evidence.
--
-- The correct question is about RECENCY, not existence: the destination is reading a
-- value through a synchroniser, so it must see one the source held WITHIN THE LAST
-- FEW SOURCE CYCLES. The bench keeps a short history of the source's recent values
-- and any observation outside that window is impossible -- exactly, not
-- probabilistically, because the source cannot have held it.
--
-- The counter is eight bits wide so that the 256-value space is far larger than the
-- six-deep window; a mixture of two adjacent values then lands outside the window
-- almost always rather than sometimes.
--
-- HOW THE FAILURE IS MADE TO HAPPEN. The source bits are driven to the DUT through
-- per-bit skew -- bit k is delayed by k times a small delta -- which is what
-- unequal routing does on any real device. That is the honest way to exhibit
-- incoherence in a simulator: bit-level settling differences are metastability and
-- are not modelled, but skew is ordinary, it produces the identical failure, and a
-- design that survives skew survives the settling case for the same reason.
--
-- A bench that changes all bits at the same instant finds nothing wrong with scheme
-- 1, which is why scheme 1 keeps being built.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_cdc_vector_tb is
end entity;
architecture sim of spi_cdc_vector_tb is
constant W : positive := 8;
constant SYNC_N : positive := 2;
constant NVAL : positive := 2 ** W;
-- How far back a legitimate observation may come from. Six is generous, which
-- matters: a window that is too tight reports correct designs as broken.
constant HIST : positive := 6;
constant SKEW : time := 1 ns;
-- Two unrelated clocks. The ratio is deliberately not an integer, so the
-- destination's sampling instant walks through the source's period.
signal src_clk : std_logic := '0';
signal dst_clk : std_logic := '0';
signal halt : boolean := false;
signal src_rst_n : std_logic := '1';
signal dst_rst_n : std_logic := '1';
signal cnt : unsigned(W - 1 downto 0) := (others => '0');
signal update : std_logic := '0';
signal cnt_gray : std_logic_vector(W - 1 downto 0);
-- The skew injection. Each bit reaches the DUT through its own delay, which is
-- a model of ROUTING rather than of metastability -- and that is what makes it
-- defensible: a design that survives skew survives the settling case for the
-- same reason, and skew is something a reviewer can point at on a floorplan.
signal bin_skewed : std_logic_vector(W - 1 downto 0);
signal gray_skewed : std_logic_vector(W - 1 downto 0);
signal dst_bin_wrong, dst_gray_ok, dst_flagged
: std_logic_vector(W - 1 downto 0);
signal dst_flagged_stb : std_logic;
-- The statistics are owned by the observer process. The stimulus asks for a
-- clear through a strobe rather than writing them, because two drivers on a
-- signal resolve to 'X' in VHDL.
signal bad_bin, bad_gray, bad_flag : natural := 0;
signal obs_bin, obs_gray, obs_flag : natural := 0;
signal distinct_bad : natural := 0;
signal stat_clr : std_logic := '0';
-- The source's recent values, newest first, maintained in SOURCE time.
type hist_arr is array (0 to HIST - 1) of std_logic_vector(W - 1 downto 0);
-- Named `src_hist` rather than `hist`: VHDL is case-INSENSITIVE, so a signal
-- called `hist` collides with the constant `HIST` and the error points at the
-- second declaration rather than at the naming convention that caused it.
signal src_hist : hist_arr := (others => (others => '0'));
begin
srcclk : process
begin
while not halt loop
src_clk <= '0'; wait for 7 ns; src_clk <= '1'; wait for 7 ns;
end loop;
wait;
end process;
dstclk : process
begin
while not halt loop
dst_clk <= '0'; wait for 5 ns; dst_clk <= '1'; wait for 5 ns;
end loop;
wait;
end process;
cnt_gray <= std_logic_vector(cnt xor shift_right(cnt, 1));
-- Bit 0 arrives on time, bit W-1 arrives (W-1)*SKEW late. A real device's
-- spread is smaller and less orderly; an orderly spread makes the result
-- reproducible, which a chapter needs.
skewgen : for i in 0 to W - 1 generate
bin_skewed(i) <= std_logic(cnt(i)) after i * SKEW;
gray_skewed(i) <= cnt_gray(i) after i * SKEW;
end generate;
dut : entity work.spi_cdc_vector
generic map (W => W, SYNC_N => SYNC_N)
port map (src_clk => src_clk, src_rst_n => src_rst_n,
src_bin => bin_skewed, src_gray => gray_skewed,
src_update => update,
dst_clk => dst_clk, dst_rst_n => dst_rst_n,
dst_bin_wrong => dst_bin_wrong, dst_gray_ok => dst_gray_ok,
dst_flagged => dst_flagged, dst_flagged_stb => dst_flagged_stb);
-- The history, maintained in source time, which is the only place the source's
-- value means anything.
histproc : process (src_clk)
begin
if rising_edge(src_clk) then
if src_rst_n = '1' then
for j in HIST - 1 downto 1 loop
src_hist(j) <= src_hist(j - 1);
end loop;
src_hist(0) <= std_logic_vector(cnt);
end if;
end if;
end process;
-- The destination observations, and the impossibility test. Sampled on the
-- destination clock, which is the only place they exist.
observe : process (dst_clk)
variable seen_bad : std_logic_vector(NVAL - 1 downto 0) := (others => '0');
impure function is_recent(v : std_logic_vector) return boolean is
begin
for j in 0 to HIST - 1 loop
if src_hist(j) = v then return true; end if;
end loop;
return false;
end function;
begin
if rising_edge(dst_clk) then
if stat_clr = '1' then
bad_bin <= 0; bad_gray <= 0; bad_flag <= 0;
obs_bin <= 0; obs_gray <= 0; obs_flag <= 0;
distinct_bad <= 0;
seen_bad := (others => '0');
elsif dst_rst_n = '1' and src_rst_n = '1' then
obs_bin <= obs_bin + 1;
if not is_recent(dst_bin_wrong) then
bad_bin <= bad_bin + 1;
if seen_bad(to_integer(unsigned(dst_bin_wrong))) = '0' then
seen_bad(to_integer(unsigned(dst_bin_wrong))) := '1';
distinct_bad <= distinct_bad + 1;
end if;
end if;
obs_gray <= obs_gray + 1;
if not is_recent(dst_gray_ok) then
bad_gray <= bad_gray + 1;
end if;
if dst_flagged_stb = '1' then
obs_flag <= obs_flag + 1;
if not is_recent(dst_flagged) then
bad_flag <= bad_flag + 1;
end if;
end if;
end if;
end if;
end process;
watchdog : process
begin
wait for 2 ms;
if not halt then
report "FAIL: the simulation did not finish within its time limit"
severity failure;
end if;
wait;
end process;
stim : process
variable errs : natural := 0;
variable p1_bin, p1_distinct, p1_obs : natural;
procedure clear_stats is
begin
stat_clr <= '1';
for i in 1 to 2 loop wait until rising_edge(dst_clk); end loop;
stat_clr <= '0';
wait until rising_edge(dst_clk);
end procedure;
-- Runs the counter for `n` updates, `gap` source cycles apart. `gap` is the
-- whole difference between the two phases: at 1 the source never holds
-- still, and at 8 it holds still for far longer than the synchroniser.
procedure run_counter(n : natural; gap : natural) is
begin
for i in 1 to n loop
wait until falling_edge(src_clk);
update <= '1';
cnt <= cnt + 1;
wait until falling_edge(src_clk);
update <= '0';
for g in 2 to gap loop
wait until falling_edge(src_clk);
end loop;
end loop;
for i in 1 to 20 loop wait until rising_edge(dst_clk); end loop;
end procedure;
procedure show(phase : string) is
begin
report " " & phase;
report " scheme observations impossible values";
report " 1 N synchronisers on a bus " &
integer'image(obs_bin) & " " & integer'image(bad_bin);
report " 2 Gray code, then decoded " &
integer'image(obs_gray) & " " & integer'image(bad_gray);
report " 3 data plus a flag " &
integer'image(obs_flag) & " " & integer'image(bad_flag);
end procedure;
begin
src_rst_n <= '1'; dst_rst_n <= '1';
for i in 1 to 2 loop wait until rising_edge(dst_clk); end loop;
src_rst_n <= '0'; dst_rst_n <= '0';
for i in 1 to 4 loop wait until rising_edge(dst_clk); end loop;
src_rst_n <= '1'; dst_rst_n <= '1';
for i in 1 to 4 loop wait until rising_edge(dst_clk); end loop;
clear_stats;
report " an " & integer'image(W) &
"-bit counter crossing two unrelated clocks (14 ns source, 10 ns destination) with up to " &
integer'image(((W - 1) * SKEW) / 1 ns) &
" ns of per-bit skew; an observation is IMPOSSIBLE if the source did not hold it within the last " &
integer'image(HIST) & " source cycles";
-- PHASE 1: the source NEVER HOLDS STILL.
run_counter(400, 1);
show("PHASE 1 -- one update per source cycle: the source never holds still");
p1_bin := bad_bin; p1_distinct := distinct_bad; p1_obs := obs_bin;
-- 1. SCHEME 1 OBSERVES VALUES THAT NEVER EXISTED.
if bad_bin = 0 then
report " FAIL: N independent synchronisers on a bus produced no impossible value -- with skew they must, and a bench that finds none is either driving all bits at the same instant or asking whether a value ever existed rather than whether it existed recently";
errs := errs + 1;
end if;
-- 2. GRAY CODE NEVER DOES. Exact, not small: a one-bit-at-a-time encoding
-- cannot produce a third value.
if bad_gray /= 0 then
report " FAIL: the Gray-coded crossing observed " &
integer'image(bad_gray) & " impossible values";
errs := errs + 1;
end if;
-- 3. AND DATA PLUS A FLAG FAILS HERE TOO, which was not the expected result
-- and is the more useful one: its precondition is that the source HOLDS
-- THE DATA STILL, and a source updating every cycle does not.
if bad_flag = 0 then
report " FAIL: data plus a flag should ALSO observe impossible values when the source never holds still -- its correctness depends on the data being stable when sampled, not on the flag";
errs := errs + 1;
end if;
report " data plus a flag observed " & integer'image(bad_flag) &
" impossible values as well, because its precondition is that the source HOLDS THE DATA STILL and a source updating every cycle does not -- the flag is not the weak part, the holding is";
-- 4. AND THE LIES WERE VARIED, not one repeated glitch.
if p1_distinct < 2 then
report " FAIL: only " & integer'image(p1_distinct) &
" distinct impossible value(s) -- one could be an alignment artefact";
errs := errs + 1;
end if;
report " scheme 1 produced " & integer'image(p1_distinct) &
" distinct impossible values out of the " & integer'image(NVAL) &
" the counter can hold, so it is not one unlucky alignment";
-- PHASE 2: the source HOLDS STILL between updates.
clear_stats;
run_counter(60, 8);
show("PHASE 2 -- one update every eight source cycles: the source holds still");
if bad_flag /= 0 then
report " FAIL: with the source holding still, data plus a flag observed " &
integer'image(bad_flag) & " impossible values";
errs := errs + 1;
end if;
if bad_gray /= 0 then
report " FAIL: the Gray crossing observed " & integer'image(bad_gray) &
" impossible values with the source holding still";
errs := errs + 1;
end if;
if bad_bin = 0 then
report " FAIL: scheme 1 should still be wrong with a slow source -- skew does not care how often the value changes";
errs := errs + 1;
end if;
report " scheme 1 was still wrong with the source holding still, " &
integer'image(bad_bin) & " times in " & integer'image(obs_bin) &
" observations against " & integer'image(p1_bin) & " in " &
integer'image(p1_obs) &
" when it never held still -- slowing the source reduces the RATE and not the fault, which is exactly how this crossing survives testing and reaches silicon";
report " every impossible value was a legal " & integer'image(W) &
"-bit pattern, so no downstream check could reject one -- the fault is undetectable by anything except the crossing being built correctly in the first place";
if obs_flag = 0 then
report " FAIL: the data-plus-flag scheme delivered nothing at all";
errs := errs + 1;
end if;
report " the flagged scheme delivered " & integer'image(obs_flag) &
" observations rather than one per destination cycle, because it only publishes when the flag arrives -- it is exact, it is slower, and it needs the source to hold still, and those three together are the trade";
-- The exact counts are phase- and simulator-dependent, and deliberately not
-- asserted: which update a destination edge happens to land inside depends
-- on the relationship between two clocks that have none, so the three
-- language benches report different totals from the same design. What is
-- asserted is the SHAPE -- scheme 1 non-zero in both regimes, Gray zero in
-- both, data-plus-flag non-zero only when the source never holds still --
-- and that is identical in all three, because it follows from the encoding
-- rather than from the alignment.
if errs = 0 then
report "PASS: a two-flop synchroniser produces a SETTLED value and nothing else -- it does not choose which of two values you get, it does not convey a pulse, and on a BUS it is wrong, because each bit settles independently and the destination reads a mixture of bits from before and after the update. Driven with per-bit skew of the kind ordinary routing produces, N independent synchronisers on an 8-bit counter observed impossible values repeatedly and in many distinct forms, every one a well-formed pattern that no downstream check could reject -- and it was still wrong when the source was slowed to update every eighth cycle, just less often, which is precisely how this crossing passes a regression and reaches silicon. A GRAY-CODED crossing through the identical synchronisers observed exactly ZERO in both regimes, because consecutive Gray codes differ in one bit so catching an update part-way yields the old value or the new one -- and that holds only for a value that changes by one step, which suits a counter or a FIFO pointer and not arbitrary data. DATA PLUS A FLAG observed zero only in the second regime, and that was the useful surprise: it fails exactly as badly as scheme 1 when the source never holds still, because its correctness rests on the data being STABLE WHEN SAMPLED rather than on the flag -- Chapter 15.2's lifetime requirement arriving from the other direction, and the reason Chapter 15.6 must build a FIFO, since an SPI slave's source cannot be told to hold still while SCLK is running. So the choice is not between synchroniser depths. It is between encodings: change one bit at a time, or hold the value still and cross one bit of handshake, or buffer it";
else
report "FAIL: " & integer'image(errs) & " error(s)" severity error;
end if;
halt <= true;
wait;
end process;
end architecture;7. Why a Verification Engineer Cares
Ask about recency, not existence. The bookkeeping in §5's callout is the single most transferable idea here: a checker that asks "did this value ever exist" against a counter can only ever answer yes. The window has to be a few source cycles wide — generous enough not to fail a correct design, narrow enough that a mixture lands outside it.
Inject skew, and put it between two flops. Without skew, scheme 1 is indistinguishable from scheme 2 forever. With skew in the wrong place, every scheme including the correct one fails. The seam being a port pair is what makes the position reviewable.
Run both regimes, and compare the numbers rather than the verdicts. The interesting result is not that scheme 1 fails — it is that slowing the source reduces its failure count without fixing it. A suite that runs one regime and reports pass/fail cannot say that.
Assert the shape, not the counts. Three simulators give three totals for the same design, for the same reason two boards would. Asserting "scheme 1 non-zero, Gray zero" is robust; asserting "13" is a bench that fails on a correct design.
// Properties for a vector crossing. The first is the only one that can be written
// about the wrong scheme, and it is an assertion that it is wrong.
property p_gray_changes_one_bit_at_a_time;
// The property the encoding exists for, and the only one worth binding inside a
// Gray-coded crossing: if this can fail, the guarantee is void and the crossing
// is scheme 1 wearing a Gray code's name.
@(posedge src_clk) disable iff (!src_rst_n)
$countones(src_gray ^ $past(src_gray)) <= 1;
endproperty
property p_flagged_data_stable_across_the_window;
// Scheme 3's real precondition, and the one that failed in phase 1: the data must
// not change between the flag toggling and the destination sampling it.
@(posedge src_clk) disable iff (!src_rst_n)
src_update |=> $stable(src_bin) [*SYNC_N+2];
endproperty
property p_flagged_strobe_follows_the_flag;
// The destination publishes only because the flag arrived, never because the data
// changed -- the data is unsynchronised and must never be watched.
@(posedge dst_clk) disable iff (!dst_rst_n)
dst_flagged_stb |-> $past(f_arrived);
endproperty
property p_no_synchroniser_on_the_data_path;
// Written as a structural check rather than a temporal one, because it is
// structural: the data register's value must appear at the destination in the
// same cycle it is sampled, with no intermediate stage. A design that grew one
// has reintroduced scheme 1.
@(posedge dst_clk) disable iff (!dst_rst_n)
dst_flagged_stb |-> (dst_flagged == $past(src_bin));
endproperty// Coverage. The axes are HOW MANY BITS CHANGED and WHETHER THE SOURCE WAS STILL,
// because those two decide whether a partial observation is possible and whether it
// matters -- and a suite that never varies the second concludes scheme 3 is correct.
covergroup cg_vector @(posedge src_clk iff src_update);
option.per_instance = 1;
// Bits changing in one update. A Gray-coded value must only ever hit 1; a binary
// counter hits high values at every power-of-two boundary, and those are the
// updates during which scheme 1 can lie the most.
changed: coverpoint bits_changed_this_update {
bins one = {1};
bins two = {2};
bins few = {[3:4]};
bins many = {[5:$]}; // 0111 -> 1000 and friends
}
// Source cycles between updates. One is the FIFO-pointer case and the regime in
// which data-plus-a-flag fails; eight is the regime in which it is exact.
spacing: coverpoint src_cycles_between_updates {
bins every_cycle = {1};
bins two = {2};
bins few = {[3:7]};
bins holds_still = {[8:$]};
}
// Skew spread on the seam, in destination-clock tenths. Zero is the bin that
// makes scheme 1 look correct, so a suite must cover more than it.
skew: coverpoint seam_skew_tenths {
bins none = {0};
bins small = {[1:3]};
bins third = {[4:6]};
bins large = {[7:$]};
}
x_changed_skew: cross changed, skew;
x_spacing_skew: cross spacing, skew;
endgroup8. Why an FPGA or ASIC Engineer Cares
Gray coding costs an encoder and a decoder, and nothing else. The encoder is v ^ (v >> 1), which is W-1 XOR gates with no depth. The decoder is a prefix XOR, which is W-1 XORs in a chain — so its depth grows with the width, and at 32 bits it is worth registering the decoded value rather than using it combinationally. That is the entire implementation cost of the scheme.
Data-plus-a-flag costs one synchronised bit and a LIFETIME requirement, and the second is not a gate count. The data register must hold its value from the flag toggling until SYNC_N destination cycles later, which is a constraint on the source's behaviour rather than on the destination's logic — and Chapter 15.2 §5 is what happens when a reset shortens that lifetime.
A multi-bit crossing must be declared to the CDC tool, and the declaration is the review. Vendor CDC checkers find "a multi-bit signal crossing domains" and ask what protects it. The answers they accept are a Gray declaration, a handshake declaration, or a waiver — and a waiver on a bus with no protection is exactly scheme 1, signed off.
The synchroniser flops must not be merged across bits. A W-bit synchroniser is W independent two-flop chains, and a tool that recognises them as a shift register array may pack them in ways that change their relative delays. ASYNC_REG on all of them, and the anti-shift-register attribute, for the same reason as Chapter 15.1 — with the extra consequence that here the relative timing between bits is what the scheme is about.
9. Failure Signature — Occasional Impossible Sensor Readings
The symptom:
"Once every few thousand reads we get a temperature of 4000 degrees. The sensor is fine, the payload checksum passes, and the driver is single-threaded."
What is happening: a multi-bit value crossed by N independent synchronisers, and the destination sampled a mixture of two adjacent source values. The reading is a well-formed number that the source never produced.
Why the checksum does not help: a checksum validates that the bytes were received correctly, and they were. Scheme 1's fault is not a corruption of bytes; it is a mixture of bit positions within one value, and every byte of the mixture is a byte the source really drove at some point. No per-byte or per-payload integrity check can see it.
How to confirm: read the value twice in quick succession. A mixture will usually not repeat, and a genuine sensor fault will. And the structural check is faster than either: count the flops between the source register and the destination's first use, per bit, and ask what makes the bits agree with each other. If the answer is "two flops each", that is the fault.
10. Common Misconceptions
"A two-flop synchroniser makes a signal safe to cross." It makes the value settled. It does not choose which of two values you get, it does not convey a pulse, and on a bus it produces values that never existed.
"Synchronise every bit and the bus is safe." Every bit is individually settled and collectively incoherent. That is the whole failure.
"A wider synchroniser helps." Depth is about settling. Nothing about depth makes W bits agree with each other.
"Gray coding is the general answer for multi-bit crossings." It is the answer for values that change by one step. Arbitrary data does not, so Gray coding a data bus gives no guarantee at all — and Chapter 15.9 builds an encoding that looks like Gray and is not.
"Data-plus-a-flag is safe because the flag is synchronised." Its correctness rests on the data being stable when sampled. Phase 1 of the measurement shows it failing exactly as badly as scheme 1 when the source never holds still.
"If the failure rate is low, the crossing is nearly correct." The rate is a property of the stimulus and the fault is a property of the structure. §5's phase 2 has a lower rate and the same fault, and that gap is how the design ships.
11. Reason It Through
Q. A 4-bit counter crosses with one synchroniser per bit. Which transitions can produce the most impossible values, and what does that say about where to look in a waveform?
The ones that change the most bits: 0111 → 1000 changes all four, 0011 → 0100 changes three. A mixture of two values that differ in k bits can produce up to 2^k - 2 values neither of them equals, so the power-of-two boundaries are where the lying is worst. In a waveform, the place to look is not where the data is unusual but where the counter crosses a boundary — which is why a stimulus of random values is better than a counter for exposing the fault and a counter is better for measuring it, since the window test needs a predictable source.
Q. Why is data-plus-a-flag the general answer while Gray coding is the special one, when Gray coding needs no lifetime requirement?
Because Gray coding's guarantee comes from a property of the data — that it changes by one step — and most data does not have that property. Data-plus-a-flag makes no assumption about the data at all; it converts the problem into a requirement on the source's behaviour, which is something a designer can arrange. The trade is that the requirement is a number rather than a structure, so it can be violated later by a change that looks unrelated — which is exactly Chapter 15.2 §5.
Q. The bench's recency window is six source cycles. What breaks if it is two, and what breaks if it is sixty?
At two, a correct design fails: the destination's observation legitimately lags the source by the synchroniser depth plus the clock-ratio slack, so a window tighter than that reports correct values as impossible. At sixty, against an 8-bit counter, the window covers nearly a quarter of the value space, so a mixture has a good chance of landing inside it and the test stops detecting the fault. The window has to be wider than the real latency and much narrower than the value space — which is why the counter is 8 bits rather than 4.
Q. Why does the design bring the skew seam out as two ports rather than keeping the delay inside the testbench's source model?
Because the two positions are not equivalent and the difference is invisible in a port list that hides it. Skew before the source's register corrupts the source and breaks every scheme; skew between the two registers models routing and breaks only the schemes that deserve it. Making the seam a port pair — tied together in a real design — puts the choice where a reviewer reads it, and the first version of this measurement got it wrong precisely because the position was implicit.
Q. A design Gray-codes a FIFO write pointer and also crosses a 32-bit data word by synchronising every bit, arguing that the data is only read when the pointer says it is valid. Is the argument sound?
The argument is sound and the implementation is not. If the data is only read when the pointer says so, then the data does not need synchronising at all — it needs to be held stable, which is data-plus-a-flag. Synchronising it as well adds 64 flops and, worse, adds a stage between the register and the read, so the value the destination sees is one cycle older than the pointer implies. The synchronisers are not just unnecessary; they break the timing relationship the argument depends on.
12. Understanding Check
13. Summary
A two-flop synchroniser produces a settled value and nothing else. It does not choose which of two values you get, it does not convey a pulse, and on a bus it is wrong — each bit settles independently and the destination reads a mixture of bits from before and after the update.
Driven with the per-bit skew ordinary routing produces, N independent synchronisers on an 8-bit counter observed impossible values repeatedly and in many distinct forms, every one a well-formed pattern no downstream check could reject — and it was still wrong when the source was slowed, just less often, which is precisely how this crossing passes a regression and reaches silicon.
Gray coding observed exactly zero in both regimes, because consecutive Gray codes differ in one bit so a partial observation is the old value or the new one. That holds only for a value that changes by one step, which suits a counter or a FIFO pointer and not arbitrary data.
Data plus a flag observed zero only when the source held still, and that was the useful surprise: it fails exactly as badly as scheme 1 when the source never does, because its correctness rests on the data being stable when sampled rather than on the flag — Chapter 15.2's lifetime requirement arriving from the other direction, and the reason Chapter 15.6 must build a FIFO.
So the choice is not between synchroniser depths. It is between encodings: change one bit at a time, or hold the value still and cross one bit of handshake, or buffer it.
Two things about the measurement are worth carrying. The skew must go between two flops, not in front of the first one — the version that got it wrong broke the correct design too. And the impossibility test must ask about recency, not existence: the first version asked whether a value had ever existed, against a counter that passes through every value, and reported zero faults on a design that is definitely broken.
For verification: inject skew and put it in the right place; run both regimes and compare the counts rather than the verdicts, because the falling rate is the finding; and assert the shape, since three simulators give three totals.
For implementation: Gray costs an encoder and a prefix-XOR decoder whose depth grows with width; data-plus-a-flag costs one bit and a lifetime, which is not a gate count; a multi-bit crossing must be declared to the CDC tool, and a waiver on an unprotected bus is scheme 1 signed off.
14. What Comes Next
Both correct schemes carry one value at a time and both require the source to stop between values. An SPI slave's source cannot stop, because SCLK does not.
Chapter 15.6 — Handshake and FIFO Crossings builds what is left. A FIFO works because the data never crosses a boundary at all — a location is written by one clock and read by the other at a different time — and what crosses is two pointers, which are exactly Gray coding's special case. Every stale pointer reading turns out to be conservative, which is what makes the design safe rather than merely probable. And then the real question: what depth does an SPI burst actually need, and what happens when the answer is "no depth is sufficient"?
Continue learning
Related tutorials
- Related topic
Slave Microarchitecture and Clocking Assumptions
The two architectures available to a slave that does not own its clock, the one precondition the chosen architecture rests on and why no simulation can test it, why MOSI must pass through exactly as many flops as SCLK, and a delay-matched front end verified in three HDLs.
- Related topic
Why an SPI Slave Is a Clock-Domain Problem
An externally generated SCLK turns a shift register into a CDC question, and there are only four ways to carry an event across: a level, a synchronised level, a toggle, and a handshake. Each is limited by something different, no scheme creates bandwidth, and the difference that matters most is the one no simulation can show.
- Related topic
Handshake and FIFO Crossings
A FIFO works because the data never crosses a boundary at all — what crosses is two pointers, which are Gray coding's special case, and every stale reading is conservative rather than optimistic. Then the sizing question a conservation test never asks: what depth does an SPI burst need, and what to do when no depth is sufficient.
- Related topic
Back-to-Back Transactions and Inter-Frame Gap
How soon chip select may fall again after it rises: the minimum deselect time, why a device needs it, what the gap costs in throughput, and the hardware enforcement that keeps software from violating it.
