UART · Module 12
UART Clock Domains: What Is and Is Not CDC
Asynchronous serial timing is not clock-domain crossing. Where a UART genuinely needs CDC structures, why pointers cross as gray code — measured — and the asynchronous FIFO that follows.
"Asynchronous" is doing two completely different jobs in this subject, and conflating them produces both of the available mistakes: building CDC structures where none are needed, and missing the crossing that actually exists.
Asynchronous serial means the link carries no clock and the receiver recovers timing from the data — Modules 2 through 5. Clock-domain crossing means a signal generated by one clock is captured by another. A UART is asynchronous in the first sense by definition and, as built through Module 11, contains exactly one crossing in the second sense: the receive pin.
1. Two Different Meanings of One Word
| Asynchronous serial | Clock-domain crossing | |
|---|---|---|
| What is asynchronous | the link — no clock is transmitted | two clocks inside one chip |
| The problem | recovering bit timing from data | capturing a signal safely |
| The solution | oversampling and start-edge alignment | synchronisers, gray code, handshakes |
| Where it lives | Modules 2–5 | this module |
| How many in the UART | the whole protocol | one signal, until a second clock arrives |
Oversampling is not a CDC technique. It solves a frequency and phase mismatch between two devices that each have a stable clock — the far end's bit rate versus this design's sampling grid — and it solves it statistically, by sampling near bit centres. It does nothing whatsoever about metastability, which is why the receiver has a synchroniser in addition.
And a synchroniser is not timing recovery. It gives a settled value quantised to a clock edge, and knows nothing about where bit boundaries are.
2. The UART As Built: One Crossing
Taking the inventory honestly for the IP of Chapter 11.2, which runs entirely on clk:
| Signal | Crosses? | Why |
|---|---|---|
rx_i | YES | driven by a device with its own oscillator |
tx_o | no | driven by clk; the far end's problem, not ours |
cts_n_i | YES | a pin, driven by the far end |
rts_n_o | no | an output of this domain |
| FIFO pointers and levels | no | one clock on both ports |
os_tick, baud_tick | no | enables in this domain, not clocks — Chapter 8.1 |
| configuration, status | no | same clock as everything else |
Two input pins cross, and both are single-bit levels, which is the easy case. rx_i is synchronised inside uart_rx (Chapter 12.2); cts_n_i is synchronised inside uart_tx_gate (Chapter 10.5 §7). Nothing else in the design crosses anything.
That is the whole CDC story for a single-clock UART, and it is worth stating plainly because a great deal of anxious complexity gets added to designs that are in exactly this position.
3. Where a Real Crossing Appears
It appears when a second clock does — and for a UART that is essentially always the register interface. The serial side wants a clock with a convenient relationship to the baud rate; the bus wants the system clock; these are rarely the same, and forcing them to be is a real constraint on the rest of the SoC.
With two clocks, three separate kinds of crossing appear, and they need three different mechanisms:
| What crosses | Mechanism | Why not the others |
|---|---|---|
| a single-bit level (an enable, a status flag) | two-flop synchroniser | nothing else is needed |
| a single-bit event (a one-cycle pulse) | toggle, then synchronise, then edge-detect | a pulse shorter than the destination period is lost |
| a multi-bit value (data, pointers) | gray code, or a handshake, or an async FIFO | §4 — independent bits arrive at independent times |
4. Why a Multi-Bit Value Crosses as Gray Code
The standard explanation is that gray code changes one bit per increment, so a sample taken during a transition yields one of the two adjacent values. That is correct and it understates how badly the alternative behaves.
Measured. A 5-bit counter in one domain, sampled by an unrelated clock, with a different routing delay on each of the five wires — which is what a real bus does. Both a binary and a gray-coded version of the same counter, sampled at the same instants, over 3,999 samples:
BINARY sampled 0 (00000) — source held 9/8/7 : IMPOSSIBLE
BINARY sampled 7 (00111) — source held 4/3/2 : IMPOSSIBLE
BINARY sampled 19 (10011) — source held 24/23/22 : IMPOSSIBLE
BINARY sampled 24 (11000) — source held 17/16/15 : IMPOSSIBLE
3999 samples taken across two unrelated clocks, with per-bit skew
BINARY pointer : 215 samples were values the counter NEVER held (5.38%)
GRAY pointer : 0 samples were values the counter NEVER held (0.00%)One sample in nineteen of the binary pointer was a value that never existed. Not a stale value — stale would be harmless — a fabricated one. Reading 0 while the counter held 8 is the case that destroys a FIFO: the reader concludes the queue is empty and stops, while eight entries sit in it.
The gray-coded counter produced none, and can produce none. With one bit changing per increment, a sample during the transition resolves that single bit to either its old or its new value, and both of those are real counter values. The property is structural rather than probabilistic.
5. The Asynchronous FIFO
Put together: arbitrary data crossing a clock boundary, with both sides free-running. The write side needs to know whether the FIFO is full and the read side whether it is empty, and neither may read the other's pointer directly.
// ===========================================================================
// A genuine asynchronous FIFO.
//
// This is NOT the FIFO of Module 10. That one is synchronous: one clock, one
// level counter, and a comparison any part of the design may read. Here the
// write and read sides are in unrelated clock domains, and NOTHING about the
// other side's state may be read directly.
//
// Three structural decisions carry the whole design:
// 1. Pointers are one bit WIDER than the address, so full and empty are
// distinguishable rather than both meaning "pointers equal".
// 2. Pointers cross as GRAY code, so a sample taken during a transition is
// always one of the two adjacent values and never a third.
// 3. Each side computes its own flag PESSIMISTICALLY from a stale view of
// the other, so the error is always in the safe direction.
// 4. Both flags are REGISTERED. That is not a style choice: the next
// pointer value depends on the flag and the flag depends on the next
// pointer value, so a combinational flag closes a genuine loop.
// ===========================================================================
module uart_async_fifo #(
parameter int unsigned WIDTH = 8,
parameter int unsigned DEPTH_L2 = 4 // DEPTH = 2**DEPTH_L2
) (
// write domain
input logic wclk,
input logic wrst_n,
input logic [WIDTH-1:0] wdata_i,
input logic wpush_i,
output logic wfull_o,
// read domain
input logic rclk,
input logic rrst_n,
output logic [WIDTH-1:0] rdata_o,
input logic rpop_i,
output logic rempty_o
);
localparam int unsigned DEPTH = 1 << DEPTH_L2;
localparam int unsigned PW = DEPTH_L2 + 1; // one extra bit — see (1)
localparam logic [PW-1:0] LAP = 3 << (PW-2); // the top two bits set
if (DEPTH_L2 < 1) begin : g_bad_depth
$error("uart_async_fifo: DEPTH_L2 must be at least 1");
end
logic [WIDTH-1:0] mem [0:DEPTH-1];
// Both pointers declared up front: each domain's synchroniser reads the
// other's gray register, so neither can be declared after its reader.
logic [PW-1:0] wgray_q, rgray_q;
// ---- write domain ------------------------------------------------------
logic [PW-1:0] wbin_q, wbin_nxt, wgray_nxt;
logic [PW-1:0] wq1_rgray, wq2_rgray;
assign wbin_nxt = wbin_q + (wpush_i && !wfull_o);
assign wgray_nxt = (wbin_nxt >> 1) ^ wbin_nxt; // binary -> gray
always_ff @(posedge wclk or negedge wrst_n)
if (!wrst_n) begin wbin_q <= '0; wgray_q <= '0; end
else begin wbin_q <= wbin_nxt; wgray_q <= wgray_nxt; end
always_ff @(posedge wclk)
if (wpush_i && !wfull_o) mem[wbin_q[DEPTH_L2-1:0]] <= wdata_i;
// The read pointer, synchronised into the WRITE domain. Two flops, and
// the ONLY thing this domain is allowed to know about the other one.
always_ff @(posedge wclk or negedge wrst_n)
if (!wrst_n) begin wq1_rgray <= '0; wq2_rgray <= '0; end
else begin wq1_rgray <= rgray_q; wq2_rgray <= wq1_rgray; end
// Full: the write pointer is one lap ahead of the read pointer. In binary
// that is "low bits equal, MSB different". The gray image of that is "the
// two vectors differ in exactly the top TWO bits and nowhere else", which
// is what LAP tests. Written this way rather than as
// wgray_nxt == {~wq2_rgray[PW-1:PW-2], wq2_rgray[PW-3:0]}
// because that spelling contains the part-select [PW-3:0], which goes
// negative at PW=2 and makes the smallest legal FIFO (DEPTH_L2=1) refuse
// to elaborate. The XOR form has no part-select and is correct for every
// PW >= 2.
// REGISTERED, see (4): wbin_nxt consumes the registered wfull_o, so the
// path flag -> next-pointer -> flag is broken by a flop.
logic wfull_d;
assign wfull_d = ((wgray_nxt ^ wq2_rgray) == LAP);
always_ff @(posedge wclk or negedge wrst_n)
if (!wrst_n) wfull_o <= 1'b0; else wfull_o <= wfull_d;
// ---- read domain -------------------------------------------------------
logic [PW-1:0] rbin_q, rbin_nxt, rgray_nxt;
logic [PW-1:0] rq1_wgray, rq2_wgray;
assign rbin_nxt = rbin_q + (rpop_i && !rempty_o);
assign rgray_nxt = (rbin_nxt >> 1) ^ rbin_nxt;
always_ff @(posedge rclk or negedge rrst_n)
if (!rrst_n) begin rbin_q <= '0; rgray_q <= '0; end
else begin rbin_q <= rbin_nxt; rgray_q <= rgray_nxt; end
assign rdata_o = mem[rbin_q[DEPTH_L2-1:0]];
always_ff @(posedge rclk or negedge rrst_n)
if (!rrst_n) begin rq1_wgray <= '0; rq2_wgray <= '0; end
else begin rq1_wgray <= wgray_q; rq2_wgray <= rq1_wgray; end
// Empty: the read pointer has caught the (stale) write pointer exactly.
// Registered for the same reason, and reset ASSERTED — an empty FIFO is
// the safe claim to make before anything has been written.
logic rempty_d;
assign rempty_d = (rgray_nxt == rq2_wgray);
always_ff @(posedge rclk or negedge rrst_n)
if (!rrst_n) rempty_o <= 1'b1; else rempty_o <= rempty_d;
endmoduleDecision 3 is the one worth dwelling on. Each side sees the other's pointer two clocks late, so each side's view is stale — and the staleness is deliberately arranged to be conservative. The write side's full is computed against an out-of-date read pointer, so it can claim full when the reader has actually freed a slot: it refuses a write it could have accepted. The read side's empty is computed against an out-of-date write pointer, so it can claim empty when a word has actually been written: it delays a read it could have made.
Both errors cost throughput and neither can corrupt data. That asymmetry is the entire safety argument, and it is why the flags are never "corrected" by exchanging more information.
Verified
Two deliberately incommensurate clocks — 7 ns and 13 ns, with no repeating edge pattern — and randomised push and pop pressure on both sides:
wrote 1499, read 1485, in flight 14
pass data comes out in order, unmodified
pass throughput achieved (over a thousand transfers)
pass never over-filled: in flight <= DEPTH
pass never under-ran: read count never exceeds write count
pass drains to empty
== 5 checks, 0 failures ==6. What the Module 10 FIFO Is Not
The synchronous FIFO of Chapter 10.2 must not be used across clock domains, and the reason is precisely §4:
It maintains a binary occupancy counter that both ports read. Drive its two ports from different clocks and that counter is a multi-bit binary value crossing a domain boundary — the exact case that produced 215 fabricated samples out of 3,999. A FIFO whose level register can read as 0 while it holds 8 entries will report empty and lose them.
Adding synchronisers to it does not repair this. Synchronising each bit of a binary counter independently is what the experiment did; the bits are still independent and the fabricated combinations still occur. The structure has to change — separate pointers per domain, gray encoding, no shared counter — which is a different module, not a modification.
Use the right one for the situation. Single clock: the Module 10 FIFO, which is smaller, has an exact occupancy for watermarks, and reports full and empty with no latency. Two clocks: this one, which costs two synchroniser stages of latency in each direction and cannot give either side an exact occupancy.
7. The Same FIFO in Verilog-2001 and VHDL-2008
The four decisions above are structural, not linguistic. They survive translation intact, and translating them is a useful exercise precisely because it forces each one to be stated in a language that will not let you get away with hand-waving.
Verilog-2001 carries the design with no difficulty. There is no logic, no
int unsigned, and no always_ff, and none of that matters here: every signal
is a vector of bits and every process is a flop.
//===========================================================================
// uart_async_fifo_v — Synthesizable Verilog-2001
//
// A genuine asynchronous FIFO. This is NOT the FIFO of Module 10: that one
// is synchronous, with one clock and one level counter any part of the
// design may read. Here the write and read sides are in unrelated clock
// domains and NOTHING about the other side's state may be read directly.
//
// Four structural decisions carry the whole design:
// 1. Pointers are one bit WIDER than the address, so full and empty are
// distinguishable rather than both meaning "pointers equal".
// 2. Pointers cross as GRAY code, so a sample taken during a transition
// is always one of the two adjacent values and never a third.
// 3. Each side computes its own flag PESSIMISTICALLY from a stale view of
// the other, so the error is always in the safe direction.
// 4. Both flags are REGISTERED. Not a style choice: the next pointer
// depends on the flag and the flag depends on the next pointer, so a
// combinational flag closes a genuine loop.
//===========================================================================
module uart_async_fifo_v #(
parameter WIDTH = 8,
parameter DEPTH_L2 = 4 // DEPTH = 2**DEPTH_L2; MUST be >= 1
) (
// write domain
input wire wclk,
input wire wrst_n,
input wire [WIDTH-1:0] wdata_i,
input wire wpush_i,
output reg wfull_o,
// read domain
input wire rclk,
input wire rrst_n,
output wire [WIDTH-1:0] rdata_o,
input wire rpop_i,
output reg rempty_o
);
localparam DEPTH = 1 << DEPTH_L2;
localparam PW = DEPTH_L2 + 1; // one extra bit -- see (1)
localparam [PW-1:0] LAP = 3 << (PW-2); // the top two bits set
reg [WIDTH-1:0] mem [0:DEPTH-1];
// Both gray pointers declared up front: each domain's synchroniser reads
// the other's gray register, so neither can be declared after its reader.
reg [PW-1:0] wgray_q, rgray_q;
//---- write domain -----------------------------------------------------
reg [PW-1:0] wbin_q, wq1_rgray, wq2_rgray;
wire [PW-1:0] wbin_nxt = wbin_q + (wpush_i && !wfull_o);
wire [PW-1:0] wgray_nxt = (wbin_nxt >> 1) ^ wbin_nxt; // binary -> gray
always @(posedge wclk or negedge wrst_n) begin
if (!wrst_n) begin
wbin_q <= {PW{1'b0}};
wgray_q <= {PW{1'b0}};
end else begin
wbin_q <= wbin_nxt;
wgray_q <= wgray_nxt;
end
end
always @(posedge wclk)
if (wpush_i && !wfull_o) mem[wbin_q[DEPTH_L2-1:0]] <= wdata_i;
// The read pointer, synchronised into the WRITE domain. Two flops, and
// the ONLY thing this domain is allowed to know about the other one.
always @(posedge wclk or negedge wrst_n) begin
if (!wrst_n) begin
wq1_rgray <= {PW{1'b0}};
wq2_rgray <= {PW{1'b0}};
end else begin
wq1_rgray <= rgray_q;
wq2_rgray <= wq1_rgray;
end
end
// Full: the write pointer is one lap ahead of the read pointer. In binary
// that is "low bits equal, MSB different". The gray image of that is "the
// two vectors differ in exactly the top TWO bits and nowhere else", which
// is what LAP tests. Written this way rather than as
// wgray_nxt == {~wq2_rgray[PW-1:PW-2], wq2_rgray[PW-3:0]}
// because that spelling contains the part-select [PW-3:0], which goes
// negative at PW=2 and makes the smallest legal FIFO (DEPTH_L2=1) refuse
// to elaborate. The XOR form has no part-select and is correct for every
// PW >= 2.
// REGISTERED, see (4): wbin_nxt consumes the registered wfull_o, so the
// path flag -> next-pointer -> flag is broken by a flop.
wire wfull_d = ((wgray_nxt ^ wq2_rgray) == LAP);
always @(posedge wclk or negedge wrst_n)
if (!wrst_n) wfull_o <= 1'b0; else wfull_o <= wfull_d;
//---- read domain ------------------------------------------------------
reg [PW-1:0] rbin_q, rq1_wgray, rq2_wgray;
wire [PW-1:0] rbin_nxt = rbin_q + (rpop_i && !rempty_o);
wire [PW-1:0] rgray_nxt = (rbin_nxt >> 1) ^ rbin_nxt;
always @(posedge rclk or negedge rrst_n) begin
if (!rrst_n) begin
rbin_q <= {PW{1'b0}};
rgray_q <= {PW{1'b0}};
end else begin
rbin_q <= rbin_nxt;
rgray_q <= rgray_nxt;
end
end
assign rdata_o = mem[rbin_q[DEPTH_L2-1:0]];
always @(posedge rclk or negedge rrst_n) begin
if (!rrst_n) begin
rq1_wgray <= {PW{1'b0}};
rq2_wgray <= {PW{1'b0}};
end else begin
rq1_wgray <= wgray_q;
rq2_wgray <= rq1_wgray;
end
end
// Empty: the read pointer has caught the (stale) write pointer exactly.
// Registered for the same reason, and reset ASSERTED -- an empty FIFO is
// the safe claim to make before anything has been written.
wire rempty_d = (rgray_nxt == rq2_wgray);
always @(posedge rclk or negedge rrst_n)
if (!rrst_n) rempty_o <= 1'b1; else rempty_o <= rempty_d;
endmoduleVHDL is the one that argues back, in three places worth knowing about.
It will not let the architecture read its own outputs. wfull_o is an
out port, and wbin_nxt needs its value. So the design keeps an internal
wfull_s, drives the port from it concurrently, and reads the internal signal
everywhere. This is pure ceremony in VHDL-2008 — the language relaxed the rule
in VHDL-2019 — but it makes the data flow explicit in a way the Verilog does
not.
Gray conversion is a shift and an XOR, spelled with functions. wbin_nxt srl 1 is not available on unsigned in the way you might expect, so the code
says shift_right(wbin_nxt, 1) xor wbin_nxt, which is exactly the
(wbin_nxt >> 1) ^ wbin_nxt of the other two.
Its assert is the best parameter guard of the three languages. A
concurrent assertion with severity failure is checked at elaboration and
needs no process, no clock and no simulation time — it simply refuses to build
a FIFO whose depth is illegal.
--==========================================================================
-- uart_async_fifo -- Synthesizable VHDL-2008
--
-- A genuine asynchronous FIFO. The write and read sides are in unrelated
-- clock domains and NOTHING about the other side's state may be read
-- directly.
--
-- Four structural decisions carry the whole design:
-- 1. Pointers are one bit WIDER than the address, so full and empty are
-- distinguishable rather than both meaning "pointers equal".
-- 2. Pointers cross as GRAY code, so a sample taken during a transition
-- is always one of the two adjacent values and never a third.
-- 3. Each side computes its own flag PESSIMISTICALLY from a stale view of
-- the other, so the error is always in the safe direction.
-- 4. Both flags are REGISTERED -- a combinational flag closes a genuine
-- loop through the next-pointer arithmetic.
--
-- wfull and rempty are read by the architecture (the next pointer depends
-- on them), and VHDL forbids reading an out port, so each is an internal
-- signal with a concurrent driver.
--==========================================================================
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity uart_async_fifo is
generic (
WIDTH : positive := 8;
DEPTH_L2 : positive := 4 -- DEPTH = 2**DEPTH_L2
);
port (
-- write domain
wclk : in std_logic;
wrst_n : in std_logic;
wdata_i : in std_logic_vector(WIDTH-1 downto 0);
wpush_i : in std_logic;
wfull_o : out std_logic;
-- read domain
rclk : in std_logic;
rrst_n : in std_logic;
rdata_o : out std_logic_vector(WIDTH-1 downto 0);
rpop_i : in std_logic;
rempty_o : out std_logic
);
end entity uart_async_fifo;
architecture rtl of uart_async_fifo is
constant DEPTH : positive := 2 ** DEPTH_L2;
constant PW : positive := DEPTH_L2 + 1; -- one extra bit -- see (1)
-- "one lap ahead" in gray code: the two vectors differ in exactly the
-- top two bits. Expressed as a mask so that no slice of the pointer has
-- to be taken -- a (PW-3 downto 0) slice goes null at PW=2 and breaks
-- the smallest legal FIFO.
constant LAP : unsigned(PW-1 downto 0)
:= shift_left(to_unsigned(3, PW), PW-2);
type mem_t is array (0 to DEPTH-1) of std_logic_vector(WIDTH-1 downto 0);
signal mem : mem_t;
-- Both gray pointers declared up front: each domain's synchroniser reads
-- the other's gray register.
signal wgray_q, rgray_q : unsigned(PW-1 downto 0) := (others => '0');
signal wbin_q, wq1_rgray, wq2_rgray : unsigned(PW-1 downto 0) := (others => '0');
signal wbin_nxt, wgray_nxt : unsigned(PW-1 downto 0);
signal wfull_s, wfull_d : std_logic;
signal rbin_q, rq1_wgray, rq2_wgray : unsigned(PW-1 downto 0) := (others => '0');
signal rbin_nxt, rgray_nxt : unsigned(PW-1 downto 0);
signal rempty_s, rempty_d : std_logic;
begin
assert DEPTH_L2 >= 1
report "uart_async_fifo: DEPTH_L2 must be at least 1" severity failure;
wfull_o <= wfull_s;
rempty_o <= rempty_s;
------------------------------------------------------------------ write
wbin_nxt <= wbin_q + 1 when (wpush_i = '1' and wfull_s = '0') else wbin_q;
wgray_nxt <= shift_right(wbin_nxt, 1) xor wbin_nxt; -- binary -> gray
process (wclk, wrst_n)
begin
if wrst_n = '0' then
wbin_q <= (others => '0');
wgray_q <= (others => '0');
elsif rising_edge(wclk) then
wbin_q <= wbin_nxt;
wgray_q <= wgray_nxt;
end if;
end process;
process (wclk)
begin
if rising_edge(wclk) then
if wpush_i = '1' and wfull_s = '0' then
mem(to_integer(wbin_q(DEPTH_L2-1 downto 0))) <= wdata_i;
end if;
end if;
end process;
-- The read pointer, synchronised into the WRITE domain. Two flops, and
-- the ONLY thing this domain is allowed to know about the other one.
process (wclk, wrst_n)
begin
if wrst_n = '0' then
wq1_rgray <= (others => '0');
wq2_rgray <= (others => '0');
elsif rising_edge(wclk) then
wq1_rgray <= rgray_q;
wq2_rgray <= wq1_rgray;
end if;
end process;
-- Full: pointers equal in the low bits, MSBs differ. In gray code that
-- is the top TWO bits inverted -- the gray equivalent of "one lap ahead".
wfull_d <= '1' when (wgray_nxt xor wq2_rgray) = LAP else '0';
process (wclk, wrst_n)
begin
if wrst_n = '0' then
wfull_s <= '0';
elsif rising_edge(wclk) then
wfull_s <= wfull_d;
end if;
end process;
------------------------------------------------------------------- read
rbin_nxt <= rbin_q + 1 when (rpop_i = '1' and rempty_s = '0') else rbin_q;
rgray_nxt <= shift_right(rbin_nxt, 1) xor rbin_nxt;
process (rclk, rrst_n)
begin
if rrst_n = '0' then
rbin_q <= (others => '0');
rgray_q <= (others => '0');
elsif rising_edge(rclk) then
rbin_q <= rbin_nxt;
rgray_q <= rgray_nxt;
end if;
end process;
rdata_o <= mem(to_integer(rbin_q(DEPTH_L2-1 downto 0)));
process (rclk, rrst_n)
begin
if rrst_n = '0' then
rq1_wgray <= (others => '0');
rq2_wgray <= (others => '0');
elsif rising_edge(rclk) then
rq1_wgray <= wgray_q;
rq2_wgray <= rq1_wgray;
end if;
end process;
-- Empty: the read pointer has caught the (stale) write pointer exactly.
-- Reset ASSERTED -- an empty FIFO is the safe claim before anything has
-- been written.
rempty_d <= '1' when rgray_nxt = rq2_wgray else '0';
process (rclk, rrst_n)
begin
if rrst_n = '0' then
rempty_s <= '1';
elsif rising_edge(rclk) then
rempty_s <= rempty_d;
end if;
end process;
end architecture rtl;8. Testing a FIFO Whose Two Halves Never Agree
The testbench has a problem the Module 10 FIFO's did not: there is no moment at which a correct answer exists. Ask "how many entries are in the FIFO" and the write side and the read side will give different numbers, both correct, because each is describing a different instant. A scoreboard that compares against occupancy is therefore comparing against a fiction.
So the oracle is built on the only thing both sides can agree about — the sequence. The writer emits a counting sequence. The reader keeps its own counter and demands that every word it pops be the next value of that sequence. No shared model, no peeking at pointers, and every failure mode lands on the same check:
- a lost word → the next compare mismatches
- a duplicated word → the next compare mismatches
- a reordered word → the next compare mismatches
and at the end of each phase, drained, the two counters must be equal.
On top of that the suite makes the clocks genuinely unrelated — periods are held in variables and changed at run time to 3/11, 11/3, 7/5 and 5/7 — so the phase relationship drifts continuously and the crossing is exercised at every alignment rather than at one.
//===========================================================================
// tb_uart_async_fifo — self-checking SystemVerilog testbench
//
// The two clocks here are UNRELATED. Their periods are set at run time and
// deliberately chosen to be incommensurate, so the phase relationship keeps
// drifting and the design is exercised at every alignment rather than at one.
//
// THE ORACLE IS INDEPENDENT. The writer emits a counting sequence; the
// reader keeps its OWN counter and demands that every popped word be the
// next value of that sequence. No shared model, no peeking at pointers:
// a lost word → the reader's next compare mismatches
// a duplicated → the reader's next compare mismatches
// a reordered → the reader's next compare mismatches
// and at the end of each phase, drained, the two counters must be equal.
//
// Two properties are checked STRUCTURALLY, through hierarchical references,
// because they are the reason the design is shaped the way it is:
// - every gray pointer transition changes exactly ONE bit (decision 2)
// - outstanding words never exceed DEPTH (decision 1, the extra bit)
//
// Same 36 counted checks as the Verilog twin.
//===========================================================================
`timescale 1ns/1ps
module tb_uart_async_fifo;
localparam int WIDTH = 8;
localparam int DL2 = 3;
localparam int DEPTH = 1 << DL2; // 8
localparam int PW = DL2 + 1; // 4
//---- two independent clocks, periods settable at run time --------------
int whalf = 5, rhalf = 7;
logic wclk = 1'b0, rclk = 1'b0;
always #(whalf) wclk = ~wclk;
always #(rhalf) rclk = ~rclk;
logic wrst_n = 1'b0, rrst_n = 1'b0;
logic w_en = 1'b0, r_en = 1'b0, w_force = 1'b0;
logic [WIDTH-1:0] wdata, rdata;
logic wfull, rempty;
wire wpush = w_force ? 1'b1 : (w_en && !wfull);
wire rpop = r_en && !rempty;
int wnext = 0, rexp = 0, mism = 0;
assign wdata = wnext[WIDTH-1:0];
uart_async_fifo #(.WIDTH(WIDTH), .DEPTH_L2(DL2)) u8 (
.wclk(wclk), .wrst_n(wrst_n), .wdata_i(wdata), .wpush_i(wpush), .wfull_o(wfull),
.rclk(rclk), .rrst_n(rrst_n), .rdata_o(rdata), .rpop_i(rpop), .rempty_o(rempty));
// the SMALLEST legal FIFO — DEPTH_L2=1, the configuration whose part-select
// the original spelling of the full comparison could not even elaborate.
logic s_wrst_n = 1'b0, s_rrst_n = 1'b0, s_push = 1'b0, s_pop = 1'b0;
logic [WIDTH-1:0] s_wdata = 8'h00;
logic [WIDTH-1:0] s_rdata;
logic s_full, s_empty;
uart_async_fifo #(.WIDTH(WIDTH), .DEPTH_L2(1)) u2 (
.wclk(wclk), .wrst_n(s_wrst_n), .wdata_i(s_wdata), .wpush_i(s_push), .wfull_o(s_full),
.rclk(rclk), .rrst_n(s_rrst_n), .rdata_o(s_rdata), .rpop_i(s_pop), .rempty_o(s_empty));
//---- the writer: emit a counting sequence ------------------------------
always_ff @(posedge wclk or negedge wrst_n)
if (!wrst_n) wnext <= 0;
else if (wpush && !wfull) wnext <= wnext + 1;
//---- the reader: demand the next value of that sequence ----------------
always @(posedge rclk or negedge rrst_n)
if (!rrst_n) rexp <= 0;
else if (rpop) begin
if (rdata !== rexp[WIDTH-1:0]) mism <= mism + 1;
a_seq: assert (rdata === rexp[WIDTH-1:0])
else $error("popped %02h, expected %02h", rdata, rexp[WIDTH-1:0]);
rexp <= rexp + 1;
end
//---- structural observer 1: outstanding words --------------------------
logic obs_en = 1'b0;
int maxout = 0, minout = 0, outst = 0;
initial begin
#0.37;
forever begin
if (obs_en) begin
outst = wnext - rexp;
if (outst > maxout) maxout = outst;
if (outst < minout) minout = outst;
a_cap: assert (outst <= DEPTH && outst >= 0)
else $error("outstanding=%0d outside 0..%0d", outst, DEPTH);
end
#1;
end
end
//---- structural observer 2: the gray property --------------------------
wire [PW-1:0] wg = u8.wgray_q; // hierarchical — observation only
wire [PW-1:0] rg = u8.rgray_q;
logic [PW-1:0] wg_prev, rg_prev;
int wg_bad = 0, rg_bad = 0, wg_moves = 0, rg_moves = 0;
// $countones is the idiomatic SystemVerilog spelling of a population count,
// and it is what this observer WANTS to say. It is not what it says, because
// the Icarus Verilog 13.0 build these listings are executed on miscomputes
// it: called repeatedly inside a procedural block it returns the vector
// WIDTH rather than the number of set bits from the second call onward
// ($countones(4'b0010) == 4). Published code has to run on the tool it is
// published with, so the count is written out. On a simulator without the
// defect, `pc(v)` may be replaced by `$countones(v)` unchanged.
function automatic int pc(input logic [PW-1:0] v);
pc = 0;
for (int b = 0; b < PW; b++) pc += v[b];
endfunction
always @(wg) begin
if (obs_en) begin
wg_moves++;
if (pc(wg ^ wg_prev) != 1) wg_bad++;
a_wgray: assert (pc(wg ^ wg_prev) == 1)
else $error("write gray pointer %b -> %b is not a single-bit step",
wg_prev, wg);
end
wg_prev = wg;
end
always @(rg) begin
if (obs_en) begin
rg_moves++;
if (pc(rg ^ rg_prev) != 1) rg_bad++;
a_rgray: assert (pc(rg ^ rg_prev) == 1)
else $error("read gray pointer %b -> %b is not a single-bit step",
rg_prev, rg);
end
rg_prev = rg;
end
//---- structural observer 3: the SYNCHRONISER DEPTH ---------------------
// This is the observer that matters most in Module 12, and the reason it
// exists is worth stating plainly.
//
// Cutting the pointer synchronisers from two flops to one does not change
// ANY functional result. The FIFO still loses nothing, duplicates nothing
// and reorders nothing; every other check in this file still passes. The
// property a second flop buys — that a metastable first stage is given a
// whole clock to resolve before anything reads it — has no representation
// in an event-driven simulator, so no amount of stimulus can expose its
// absence.
//
// What IS observable is the STRUCTURE: with two flops in the chain, the
// value of wq2_rgray after any edge must equal the value wq1_rgray held
// BEFORE that edge. Cut the chain to one flop and the two move together.
// So the testbench checks the shape of the logic rather than the behaviour
// of the data — which is exactly what CDC lint does on a netlist, done
// here with the only instrument a simulation has.
logic [PW-1:0] wq1_dly, rq1_dly;
int sdepth_bad_w = 0, sdepth_bad_r = 0, sdepth_obs = 0;
always @(posedge wclk) begin
#1; // read POST-edge values
if (obs_en && wrst_n) begin
sdepth_obs++;
if (u8.wq2_rgray !== wq1_dly) sdepth_bad_w++;
a_wdepth: assert (u8.wq2_rgray === wq1_dly)
else $error("write pointer synchroniser is not two flops deep");
end
wq1_dly = u8.wq1_rgray;
end
always @(posedge rclk) begin
#1;
if (obs_en && rrst_n) begin
if (u8.rq2_wgray !== rq1_dly) sdepth_bad_r++;
a_rdepth: assert (u8.rq2_wgray === rq1_dly)
else $error("read pointer synchroniser is not two flops deep");
end
rq1_dly = u8.rq1_wgray;
end
task automatic obs_start; // arm the observers from a known state
wg_prev = wg; rg_prev = rg;
wq1_dly = u8.wq1_rgray; rq1_dly = u8.rq1_wgray;
obs_en = 1'b1;
endtask
//---- check plumbing ----------------------------------------------------
int checks = 0, failures = 0;
task automatic check(input logic cond, input string name);
checks++;
if (cond) $display(" PASS %0s", name);
else begin failures++; $display(" FAIL %0s", name); end
endtask
int base_w, base_r, base_m, lat, i, guard;
// Run one traffic phase at a given clock ratio and drain it.
task automatic run_phase(input int wh, input int rh, input int nwords);
whalf = wh; rhalf = rh;
base_w = wnext; base_r = rexp;
w_en = 1'b1; r_en = 1'b1;
// The #1 after each edge matters twice over: it lets the loop read
// SETTLED counters rather than their pre-NBA values, and it moves the
// deassertion of w_en off the very edge the DUT samples wpush on.
// Without it an extra word slips in at the boundary.
guard = 0;
while (wnext - base_w < nwords && guard < 200000) begin
@(posedge wclk); #1; guard++;
end
w_en = 1'b0; // stop writing, then drain
guard = 0;
while (rexp != wnext && guard < 200000) begin
@(posedge rclk); #1; guard++;
end
repeat (6) @(posedge rclk);
r_en = 1'b0;
endtask
initial begin
#500_000;
$display(" FAIL watchdog: simulation did not finish");
$display("== %0d checks, %0d failures ==", checks+1, failures+1);
$display(" RESULT: SYSTEMVERILOG ASYNC-FIFO TESTS FAILED (timeout)");
$finish;
end
initial begin
$display("== uart_async_fifo : self-checking SystemVerilog testbench ==");
wrst_n = 1'b0; rrst_n = 1'b0; s_wrst_n = 1'b0; s_rrst_n = 1'b0;
repeat (8) @(posedge wclk);
check(rempty === 1'b1 && wfull === 1'b0,
"out of reset: empty asserted, full deasserted");
check(s_empty === 1'b1 && s_full === 1'b0,
"the DEPTH=2 FIFO resets the same way");
wrst_n = 1'b1; rrst_n = 1'b1; s_wrst_n = 1'b1; s_rrst_n = 1'b1;
repeat (8) @(posedge rclk);
obs_start;
//=== phase 1: writer much faster than reader ========================
run_phase(3, 11, 40);
check(mism == 0, "fast-write / slow-read: sequence intact");
check(wnext == 40, "fast-write / slow-read: exactly 40 written");
check(rexp == wnext, "fast-write / slow-read: nothing left behind");
check(rempty === 1'b1, "fast-write / slow-read: drained to empty");
check(wfull === 1'b0, "and full has released");
//=== phase 2: reader much faster than writer ========================
base_r = rexp;
run_phase(11, 3, 40);
check(mism == 0, "slow-write / fast-read: sequence intact");
check(rexp - base_r == 40, "slow-write / fast-read: all 40 words arrived");
//=== phase 3: near-equal but incommensurate =========================
run_phase(7, 5, 60);
check(mism == 0, "incommensurate ratio: sequence intact");
check(rexp == 140, "incommensurate ratio: 40+40+60 words total");
check(rexp == wnext, "incommensurate ratio: counts agree");
//=== the extra pointer bit: capacity is exactly DEPTH ===============
whalf = 5; rhalf = 7;
base_w = wnext;
r_en = 1'b0; // reader stalled: nothing can drain
w_en = 1'b1;
guard = 0;
while (wfull !== 1'b1 && guard < 200) begin @(posedge wclk); #1; guard++; end
check(wfull === 1'b1, "reader stalled: full eventually asserts");
repeat (4) @(posedge wclk);
check(wnext - base_w == DEPTH,
"exactly DEPTH words were accepted -- not DEPTH-1, not DEPTH+1");
//=== a full FIFO refuses further writes =============================
base_w = wnext;
w_force = 1'b1; // push HARD, ignoring wfull
repeat (20) @(posedge wclk);
w_force = 1'b0; w_en = 1'b0;
check(wnext - base_w == 0, "20 forced pushes into a full FIFO: all ignored");
//=== and the DEPTH words are still the right DEPTH words ============
base_r = rexp;
r_en = 1'b1;
guard = 0;
while (rexp != wnext && guard < 2000) begin @(posedge rclk); #1; guard++; end
repeat (6) @(posedge rclk);
r_en = 1'b0;
check(rexp - base_r == DEPTH, "draining recovers exactly DEPTH words");
check(mism == 0, "and they are the right words, in order");
check(rempty === 1'b1, "empty again");
//=== visibility latency across the boundary =========================
// One word into an empty FIFO. The datum is in memory on the writing
// edge; the READ side may not believe it until the gray pointer has
// crossed two synchroniser flops. That cost is the price of safety.
w_en = 1'b1;
@(posedge wclk);
@(negedge wclk) w_en = 1'b0;
lat = 0;
while (rempty !== 1'b0 && lat < 12) begin @(posedge rclk); #0.1; lat++; end
check(lat >= 2, "a new word costs at least TWO rclk edges to become visible");
check(lat <= 5, "and the crossing is bounded, not open-ended");
r_en = 1'b1;
guard = 0;
while (rexp != wnext && guard < 200) begin @(posedge rclk); #1; guard++; end
r_en = 1'b0;
check(mism == 0, "the single word read back correctly");
//=== the DEPTH=2 FIFO: capacity two, and it really is two ===========
for (i = 0; i < 6; i++) begin
@(negedge wclk);
s_wdata = 8'hA0 + i[7:0];
s_push = !s_full;
end
@(negedge wclk) s_push = 1'b0;
repeat (4) @(posedge wclk);
check(s_full === 1'b1, "DEPTH=2 FIFO fills");
repeat (6) @(posedge rclk);
check(s_empty === 1'b0, "and the read side sees its contents");
@(negedge rclk) s_pop = 1'b1;
repeat (2) @(posedge rclk);
@(negedge rclk) s_pop = 1'b0;
repeat (8) @(posedge rclk);
check(s_empty === 1'b1, "two pops empty it");
//=== the structural properties ======================================
check(wg_moves > 50 && rg_moves > 50,
"both gray pointers moved enough to be worth judging");
check(wg_bad == 0, "every write-pointer transition changed exactly ONE bit");
check(rg_bad == 0, "every read-pointer transition changed exactly ONE bit");
check(maxout <= DEPTH,
"outstanding words never exceeded DEPTH -- no silent overwrite");
check(minout >= 0,
"outstanding words never went negative -- nothing read before written");
check(sdepth_obs > 300,
"the synchroniser-depth observer ran on enough edges to judge");
check(sdepth_bad_w == 0,
"write-side pointer synchroniser is TWO flops deep, not one");
check(sdepth_bad_r == 0,
"read-side pointer synchroniser is TWO flops deep, not one");
//=== reset during traffic ===========================================
obs_en = 1'b0;
w_en = 1'b1; r_en = 1'b1;
repeat (20) @(posedge wclk);
wrst_n = 1'b0; rrst_n = 1'b0; // both domains reset together
w_en = 1'b0; r_en = 1'b0;
repeat (6) @(posedge wclk);
check(rempty === 1'b1 && wfull === 1'b0,
"reset during traffic returns the FIFO to empty");
wrst_n = 1'b1; rrst_n = 1'b1;
repeat (8) @(posedge rclk);
obs_start;
// A baseline rather than a reset of the counter: this asks "were there
// NEW mismatches after the reset", which is the actual question,
// instead of destroying the earlier evidence. The VHDL twin has no
// choice -- mism is driven by the reader process there, and a second
// driver on an unresolved type is illegal -- and it is better practice
// here too.
base_m = mism;
run_phase(5, 7, 24);
check(mism == base_m, "traffic after reset is clean from the first word");
check(rexp == wnext, "and balanced");
check(wg_bad == 0 && rg_bad == 0,
"the gray property survived the reset too");
$display("== %0d checks, %0d failures ==", checks, failures);
if (failures == 0) $display(" RESULT: ALL SYSTEMVERILOG ASYNC-FIFO TESTS PASSED");
else $display(" RESULT: SYSTEMVERILOG ASYNC-FIFO TESTS FAILED");
$finish;
end
endmodule//===========================================================================
// tb_uart_async_fifo_v — self-checking Verilog-2001 testbench
//
// The two clocks here are UNRELATED. Their periods are set at run time and
// deliberately chosen to be incommensurate, so the phase relationship keeps
// drifting and the design is exercised at every alignment rather than at one.
//
// THE ORACLE IS INDEPENDENT. The writer emits a counting sequence; the
// reader keeps its OWN counter and demands that every popped word be the
// next value of that sequence. No shared model, no peeking at pointers:
// a lost word -> the reader's next compare mismatches
// a duplicated -> the reader's next compare mismatches
// a reordered -> the reader's next compare mismatches
// and at the end of each phase, drained, the two counters must be equal.
//
// Two properties are checked STRUCTURALLY, through hierarchical references,
// because they are the reason the design is shaped the way it is:
// - every gray pointer transition changes exactly ONE bit (decision 2)
// - outstanding words never exceed DEPTH (decision 1, the extra bit)
//===========================================================================
`timescale 1ns/1ps
module tb_uart_async_fifo_v;
localparam WIDTH = 8;
localparam DL2 = 3;
localparam DEPTH = 1 << DL2; // 8
localparam PW = DL2 + 1; // 4
//---- two independent clocks, periods settable at run time --------------
integer whalf = 5, rhalf = 7;
reg wclk = 1'b0, rclk = 1'b0;
always #(whalf) wclk = ~wclk;
always #(rhalf) rclk = ~rclk;
reg wrst_n = 1'b0, rrst_n = 1'b0;
reg w_en = 1'b0, r_en = 1'b0, w_force = 1'b0;
wire [WIDTH-1:0] wdata, rdata;
wire wfull, rempty;
wire wpush = w_force ? 1'b1 : (w_en && !wfull);
wire rpop = r_en && !rempty;
integer wnext = 0, rexp = 0, mism = 0;
assign wdata = wnext[WIDTH-1:0];
uart_async_fifo_v #(.WIDTH(WIDTH), .DEPTH_L2(DL2)) u8 (
.wclk(wclk), .wrst_n(wrst_n), .wdata_i(wdata), .wpush_i(wpush), .wfull_o(wfull),
.rclk(rclk), .rrst_n(rrst_n), .rdata_o(rdata), .rpop_i(rpop), .rempty_o(rempty));
// the SMALLEST legal FIFO -- DEPTH_L2=1, the configuration whose part-select
// the original spelling of the full comparison could not even elaborate.
reg s_wrst_n = 1'b0, s_rrst_n = 1'b0, s_push = 1'b0, s_pop = 1'b0;
reg [WIDTH-1:0] s_wdata = 8'h00;
wire [WIDTH-1:0] s_rdata;
wire s_full, s_empty;
uart_async_fifo_v #(.WIDTH(WIDTH), .DEPTH_L2(1)) u2 (
.wclk(wclk), .wrst_n(s_wrst_n), .wdata_i(s_wdata), .wpush_i(s_push), .wfull_o(s_full),
.rclk(rclk), .rrst_n(s_rrst_n), .rdata_o(s_rdata), .rpop_i(s_pop), .rempty_o(s_empty));
//---- the writer: emit a counting sequence ------------------------------
always @(posedge wclk or negedge wrst_n)
if (!wrst_n) wnext <= 0;
else if (wpush && !wfull) wnext <= wnext + 1;
//---- the reader: demand the next value of that sequence ----------------
always @(posedge rclk or negedge rrst_n)
if (!rrst_n) rexp <= 0;
else if (rpop) begin
if (rdata !== rexp[WIDTH-1:0]) mism <= mism + 1;
rexp <= rexp + 1;
end
//---- structural observer 1: outstanding words --------------------------
// Sampled off the integer-ns stimulus grid so that every read is of
// settled values (see the reset-sync testbench for why this matters).
reg obs_en = 1'b0;
integer maxout = 0, minout = 0, outst = 0;
initial begin
#0.37;
forever begin
if (obs_en) begin
outst = wnext - rexp;
if (outst > maxout) maxout = outst;
if (outst < minout) minout = outst;
end
#1;
end
end
//---- structural observer 2: the gray property --------------------------
wire [PW-1:0] wg = u8.wgray_q; // hierarchical -- observation only
wire [PW-1:0] rg = u8.rgray_q;
reg [PW-1:0] wg_prev, rg_prev;
integer wg_bad = 0, rg_bad = 0, wg_moves = 0, rg_moves = 0;
function [2:0] pc; // population count, 4 bits
input [PW-1:0] v;
integer i;
begin pc = 0; for (i = 0; i < PW; i = i + 1) pc = pc + v[i]; end
endfunction
always @(wg) begin
if (obs_en) begin
wg_moves = wg_moves + 1;
if (pc(wg ^ wg_prev) !== 3'd1) wg_bad = wg_bad + 1;
end
wg_prev = wg;
end
always @(rg) begin
if (obs_en) begin
rg_moves = rg_moves + 1;
if (pc(rg ^ rg_prev) !== 3'd1) rg_bad = rg_bad + 1;
end
rg_prev = rg;
end
//---- structural observer 3: the SYNCHRONISER DEPTH ---------------------
// This is the observer that matters most in Module 12, and the reason it
// exists is worth stating plainly.
//
// Cutting the pointer synchronisers from two flops to one does not change
// ANY functional result. The FIFO still loses nothing, duplicates nothing
// and reorders nothing; every other check in this file still passes. The
// property a second flop buys -- that a metastable first stage is given a
// whole clock to resolve before anything reads it -- has no representation
// in an event-driven simulator, so no amount of stimulus can expose its
// absence.
//
// What IS observable is the STRUCTURE: with two flops in the chain, the
// value of wq2_rgray after any edge must equal the value wq1_rgray held
// BEFORE that edge. Cut the chain to one flop and the two move together.
// So the testbench checks the shape of the logic rather than the behaviour
// of the data -- which is exactly what CDC lint does on a netlist, done
// here with the only instrument a simulation has.
reg [PW-1:0] wq1_dly, rq1_dly;
integer sdepth_bad_w = 0, sdepth_bad_r = 0, sdepth_obs = 0;
always @(posedge wclk) begin
#1; // read POST-edge values
if (obs_en && wrst_n) begin
sdepth_obs = sdepth_obs + 1;
if (u8.wq2_rgray !== wq1_dly) sdepth_bad_w = sdepth_bad_w + 1;
end
wq1_dly = u8.wq1_rgray;
end
always @(posedge rclk) begin
#1;
if (obs_en && rrst_n) begin
if (u8.rq2_wgray !== rq1_dly) sdepth_bad_r = sdepth_bad_r + 1;
end
rq1_dly = u8.rq1_wgray;
end
task obs_start; // arm the observers from a known state
begin wg_prev = wg; rg_prev = rg;
wq1_dly = u8.wq1_rgray; rq1_dly = u8.rq1_wgray;
obs_en = 1'b1; end
endtask
//---- check plumbing ----------------------------------------------------
integer checks = 0, failures = 0;
task check;
input cond;
input [8*80-1:0] name;
begin
checks = checks + 1;
if (cond) $display(" PASS %0s", name);
else begin failures = failures + 1; $display(" FAIL %0s", name); end
end
endtask
integer base_w, base_r, base_m, lat, i, guard;
// Run one traffic phase at a given clock ratio and drain it.
task run_phase;
input integer wh;
input integer rh;
input integer nwords;
begin
whalf = wh; rhalf = rh;
base_w = wnext; base_r = rexp;
w_en = 1'b1; r_en = 1'b1;
// The #1 after each edge matters twice over: it lets the loop read
// SETTLED counters rather than their pre-NBA values, and it moves the
// deassertion of w_en off the very edge the DUT samples wpush on.
// Without it an extra word slips in at the boundary and the phase
// writes nwords+1.
guard = 0;
while (wnext - base_w < nwords && guard < 200000) begin
@(posedge wclk); #1; guard = guard + 1;
end
w_en = 1'b0; // stop writing, then drain
guard = 0;
while (rexp != wnext && guard < 200000) begin
@(posedge rclk); #1; guard = guard + 1;
end
repeat (6) @(posedge rclk);
r_en = 1'b0;
end
endtask
initial begin
#500_000;
$display(" FAIL watchdog: simulation did not finish");
$display("== %0d checks, %0d failures ==", checks+1, failures+1);
$display(" RESULT: VERILOG ASYNC-FIFO TESTS FAILED (timeout)");
$finish;
end
initial begin
$display("== uart_async_fifo_v : self-checking Verilog testbench ==");
wrst_n = 1'b0; rrst_n = 1'b0; s_wrst_n = 1'b0; s_rrst_n = 1'b0;
repeat (8) @(posedge wclk);
check(rempty === 1'b1 && wfull === 1'b0,
"out of reset: empty asserted, full deasserted");
check(s_empty === 1'b1 && s_full === 1'b0,
"the DEPTH=2 FIFO resets the same way");
wrst_n = 1'b1; rrst_n = 1'b1; s_wrst_n = 1'b1; s_rrst_n = 1'b1;
repeat (8) @(posedge rclk);
obs_start;
//=== phase 1: writer much faster than reader ========================
run_phase(3, 11, 40);
check(mism == 0, "fast-write / slow-read: sequence intact");
check(wnext == 40, "fast-write / slow-read: exactly 40 written");
check(rexp == wnext, "fast-write / slow-read: nothing left behind");
check(rempty === 1'b1, "fast-write / slow-read: drained to empty");
check(wfull === 1'b0, "and full has released");
//=== phase 2: reader much faster than writer ========================
base_r = rexp;
run_phase(11, 3, 40);
check(mism == 0, "slow-write / fast-read: sequence intact");
check(rexp - base_r == 40, "slow-write / fast-read: all 40 words arrived");
//=== phase 3: near-equal but incommensurate =========================
run_phase(7, 5, 60);
check(mism == 0, "incommensurate ratio: sequence intact");
check(rexp == 140, "incommensurate ratio: 40+40+60 words total");
check(rexp == wnext, "incommensurate ratio: counts agree");
//=== the extra pointer bit: capacity is exactly DEPTH ===============
whalf = 5; rhalf = 7;
base_w = wnext;
r_en = 1'b0; // reader stalled: nothing can drain
w_en = 1'b1;
guard = 0;
while (wfull !== 1'b1 && guard < 200) begin @(posedge wclk); guard = guard + 1; end
check(wfull === 1'b1, "reader stalled: full eventually asserts");
repeat (4) @(posedge wclk);
check(wnext - base_w == DEPTH,
"exactly DEPTH words were accepted -- not DEPTH-1, not DEPTH+1");
//=== a full FIFO refuses further writes =============================
base_w = wnext;
w_force = 1'b1; // push HARD, ignoring wfull
repeat (20) @(posedge wclk);
w_force = 1'b0; w_en = 1'b0;
check(wnext - base_w == 0, "20 forced pushes into a full FIFO: all ignored");
//=== and the DEPTH words are still the right DEPTH words ============
base_r = rexp;
r_en = 1'b1;
guard = 0;
while (rexp != wnext && guard < 2000) begin @(posedge rclk); guard = guard + 1; end
repeat (6) @(posedge rclk);
r_en = 1'b0;
check(rexp - base_r == DEPTH, "draining recovers exactly DEPTH words");
check(mism == 0, "and they are the right words, in order");
check(rempty === 1'b1, "empty again");
//=== visibility latency across the boundary =========================
// One word into an empty FIFO. The datum is in memory on the writing
// edge; the READ side may not believe it until the gray pointer has
// crossed two synchroniser flops. That cost is the price of safety.
w_en = 1'b1;
@(posedge wclk);
@(negedge wclk) w_en = 1'b0;
lat = 0;
while (rempty !== 1'b0 && lat < 12) begin @(posedge rclk); #0.1; lat = lat + 1; end
check(lat >= 2, "a new word costs at least TWO rclk edges to become visible");
check(lat <= 5, "and the crossing is bounded, not open-ended");
r_en = 1'b1;
guard = 0;
while (rexp != wnext && guard < 200) begin @(posedge rclk); guard = guard + 1; end
r_en = 1'b0;
check(mism == 0, "the single word read back correctly");
//=== the DEPTH=2 FIFO: capacity two, and it really is two ===========
for (i = 0; i < 6; i = i + 1) begin
@(negedge wclk);
s_wdata = 8'hA0 + i[7:0];
s_push = !s_full;
end
@(negedge wclk) s_push = 1'b0;
repeat (4) @(posedge wclk);
check(s_full === 1'b1, "DEPTH=2 FIFO fills");
repeat (6) @(posedge rclk);
check(s_empty === 1'b0, "and the read side sees its contents");
@(negedge rclk) s_pop = 1'b1;
repeat (2) @(posedge rclk);
@(negedge rclk) s_pop = 1'b0;
repeat (8) @(posedge rclk);
check(s_empty === 1'b1, "two pops empty it");
//=== the structural properties ======================================
check(wg_moves > 50 && rg_moves > 50,
"both gray pointers moved enough to be worth judging");
check(wg_bad == 0, "every write-pointer transition changed exactly ONE bit");
check(rg_bad == 0, "every read-pointer transition changed exactly ONE bit");
check(maxout <= DEPTH,
"outstanding words never exceeded DEPTH -- no silent overwrite");
check(minout >= 0,
"outstanding words never went negative -- nothing read before written");
check(sdepth_obs > 300,
"the synchroniser-depth observer ran on enough edges to judge");
check(sdepth_bad_w == 0,
"write-side pointer synchroniser is TWO flops deep, not one");
check(sdepth_bad_r == 0,
"read-side pointer synchroniser is TWO flops deep, not one");
//=== reset during traffic ===========================================
obs_en = 1'b0;
w_en = 1'b1; r_en = 1'b1;
repeat (20) @(posedge wclk);
wrst_n = 1'b0; rrst_n = 1'b0; // both domains reset together
w_en = 1'b0; r_en = 1'b0;
repeat (6) @(posedge wclk);
check(rempty === 1'b1 && wfull === 1'b0,
"reset during traffic returns the FIFO to empty");
wrst_n = 1'b1; rrst_n = 1'b1;
repeat (8) @(posedge rclk);
obs_start;
// A baseline rather than a reset of the counter: this asks "were there
// NEW mismatches after the reset", which is the actual question,
// instead of destroying the earlier evidence. The VHDL twin has no
// choice -- mism is driven by the reader process there, and a second
// driver on an unresolved type is illegal -- and it is better practice
// here too.
base_m = mism;
run_phase(5, 7, 24);
check(mism == base_m, "traffic after reset is clean from the first word");
check(rexp == wnext, "and balanced");
check(wg_bad == 0 && rg_bad == 0,
"the gray property survived the reset too");
$display("== %0d checks, %0d failures ==", checks, failures);
if (failures == 0) $display(" RESULT: ALL VERILOG ASYNC-FIFO TESTS PASSED");
else $display(" RESULT: VERILOG ASYNC-FIFO TESTS FAILED");
$finish;
end
endmodule--===========================================================================
-- tb_uart_async_fifo — self-checking VHDL-2008 testbench
--
-- The two clocks here are UNRELATED. Their periods are held in signals of
-- type TIME and changed at run time, deliberately to incommensurate values,
-- so the phase relationship keeps drifting and the design is exercised at
-- every alignment rather than at one.
--
-- THE ORACLE IS INDEPENDENT. The writer emits a counting sequence; the
-- reader keeps its OWN counter and demands that every popped word be the
-- next value of that sequence. No shared model, no peeking at pointers:
-- a lost word → the reader's next compare mismatches
-- a duplicated → the reader's next compare mismatches
-- a reordered → the reader's next compare mismatches
-- and at the end of each phase, drained, the two counters must be equal.
--
-- Two properties are checked STRUCTURALLY, through VHDL-2008 EXTERNAL NAMES,
-- because they are the reason the design is shaped the way it is:
-- - every gray pointer transition changes exactly ONE bit (decision 2)
-- - outstanding words never exceed DEPTH (decision 1, the extra bit)
-- External names are the VHDL equivalent of Verilog's hierarchical reference,
-- and like it they are for OBSERVATION only — nothing here drives the DUT's
-- internals or derives an expectation from them.
--
-- Same 36 counted checks as the Verilog and SystemVerilog twins.
--===========================================================================
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity tb_uart_async_fifo is
end entity tb_uart_async_fifo;
architecture sim of tb_uart_async_fifo is
constant WIDTH : positive := 8;
constant DL2 : positive := 3;
constant DEPTH : positive := 2 ** DL2; -- 8
constant PW : positive := DL2 + 1; -- 4
signal whalf : time := 5 ns;
signal rhalf : time := 7 ns;
signal wclk : std_logic := '0';
signal rclk : std_logic := '0';
signal done : boolean := false;
signal wrst_n, rrst_n : std_logic := '0';
signal w_en, r_en, w_force : std_logic := '0';
signal wdata, rdata : std_logic_vector(WIDTH-1 downto 0);
signal wfull, rempty : std_logic;
signal wpush, rpop : std_logic;
signal wnext, rexp : natural := 0;
signal mism : natural := 0;
-- the SMALLEST legal FIFO — DEPTH_L2=1, the configuration whose part-select
-- the original spelling of the full comparison could not even elaborate.
signal s_wrst_n, s_rrst_n, s_push, s_pop : std_logic := '0';
signal s_wdata, s_rdata : std_logic_vector(WIDTH-1 downto 0) := (others => '0');
signal s_full, s_empty : std_logic;
signal obs_en : boolean := false;
signal maxout, minout : integer := 0;
signal wg_bad, rg_bad, wg_moves, rg_moves : natural := 0;
signal sdepth_bad_w, sdepth_bad_r, sdepth_obs : natural := 0;
begin
---- two independent clocks, periods settable at run time ------------------
wgen : process
begin
while not done loop
wait for whalf;
wclk <= not wclk;
end loop;
wait;
end process wgen;
rgen : process
begin
while not done loop
wait for rhalf;
rclk <= not rclk;
end loop;
wait;
end process rgen;
wpush <= '1' when w_force = '1' else (w_en and not wfull);
rpop <= r_en and not rempty;
wdata <= std_logic_vector(to_unsigned(wnext mod 2**WIDTH, WIDTH));
u8 : entity work.uart_async_fifo
generic map (WIDTH => WIDTH, DEPTH_L2 => DL2)
port map (wclk => wclk, wrst_n => wrst_n, wdata_i => wdata,
wpush_i => wpush, wfull_o => wfull,
rclk => rclk, rrst_n => rrst_n, rdata_o => rdata,
rpop_i => rpop, rempty_o => rempty);
u2 : entity work.uart_async_fifo
generic map (WIDTH => WIDTH, DEPTH_L2 => 1)
port map (wclk => wclk, wrst_n => s_wrst_n, wdata_i => s_wdata,
wpush_i => s_push, wfull_o => s_full,
rclk => rclk, rrst_n => s_rrst_n, rdata_o => s_rdata,
rpop_i => s_pop, rempty_o => s_empty);
---- the writer: emit a counting sequence ----------------------------------
wproc : process (wclk, wrst_n)
begin
if wrst_n = '0' then
wnext <= 0;
elsif rising_edge(wclk) then
if wpush = '1' and wfull = '0' then
wnext <= wnext + 1;
end if;
end if;
end process wproc;
---- the reader: demand the next value of that sequence --------------------
rproc : process (rclk, rrst_n)
begin
if rrst_n = '0' then
rexp <= 0;
elsif rising_edge(rclk) then
if rpop = '1' then
if rdata /= std_logic_vector(to_unsigned(rexp mod 2**WIDTH, WIDTH)) then
mism <= mism + 1;
report "popped word does not match the expected sequence"
severity error;
end if;
rexp <= rexp + 1;
end if;
end if;
end process rproc;
---- structural observer 1: outstanding words ------------------------------
-- Sampled off the whole-nanosecond stimulus grid so that every read is of
-- settled values, matching the Verilog and SystemVerilog twins.
obs_cap : process
variable outst : integer;
begin
wait for 0.37 ns;
loop
if obs_en then
outst := wnext - rexp;
if outst > maxout then maxout <= outst; end if;
if outst < minout then minout <= outst; end if;
assert outst <= DEPTH and outst >= 0
report "outstanding=" & integer'image(outst)
& " outside 0.." & integer'image(DEPTH)
severity error;
end if;
exit when done;
wait for 1 ns;
end loop;
wait;
end process obs_cap;
---- structural observer 2: the gray property ------------------------------
-- VHDL-2008 external names. Observation only.
obs_gray : process
alias wg is << signal .tb_uart_async_fifo.u8.wgray_q : unsigned(PW-1 downto 0) >>;
alias rg is << signal .tb_uart_async_fifo.u8.rgray_q : unsigned(PW-1 downto 0) >>;
variable wg_prev, rg_prev : unsigned(PW-1 downto 0) := (others => '0');
-- VHDL-2008 has a unary XOR reduction but no population count, so the
-- count is written out. It is the same count the Verilog twin computes
-- in a function and the SystemVerilog twin would spell $countones.
function pc(v : unsigned) return natural is
variable n : natural := 0;
begin
for b in v'range loop
if v(b) = '1' then n := n + 1; end if;
end loop;
return n;
end function pc;
begin
wait on wg, rg;
loop
if wg /= wg_prev then
if obs_en then
wg_moves <= wg_moves + 1;
if pc(wg xor wg_prev) /= 1 then
wg_bad <= wg_bad + 1;
report "write gray pointer step is not a single bit"
severity error;
end if;
end if;
wg_prev := wg;
end if;
if rg /= rg_prev then
if obs_en then
rg_moves <= rg_moves + 1;
if pc(rg xor rg_prev) /= 1 then
rg_bad <= rg_bad + 1;
report "read gray pointer step is not a single bit"
severity error;
end if;
end if;
rg_prev := rg;
end if;
exit when done;
wait on wg, rg, done;
end loop;
wait;
end process obs_gray;
---- structural observer 3: the SYNCHRONISER DEPTH -------------------------
-- This is the observer that matters most in Module 12, and the reason it
-- exists is worth stating plainly.
--
-- Cutting the pointer synchronisers from two flops to one does not change
-- ANY functional result. The FIFO still loses nothing, duplicates nothing
-- and reorders nothing; every other check in this file still passes. The
-- property a second flop buys — that a metastable first stage is given a
-- whole clock to resolve before anything reads it — has no representation
-- in an event-driven simulator, so no amount of stimulus can expose its
-- absence.
--
-- What IS observable is the STRUCTURE: with two flops in the chain, the
-- value of wq2_rgray after any edge must equal the value wq1_rgray held
-- BEFORE that edge. Cut the chain to one flop and the two move together.
-- So the testbench checks the shape of the logic rather than the behaviour
-- of the data — which is exactly what CDC lint does on a netlist, done
-- here with the only instrument a simulation has.
obs_depth_w : process
alias wq1 is << signal .tb_uart_async_fifo.u8.wq1_rgray : unsigned(PW-1 downto 0) >>;
alias wq2 is << signal .tb_uart_async_fifo.u8.wq2_rgray : unsigned(PW-1 downto 0) >>;
variable wq1_dly : unsigned(PW-1 downto 0) := (others => '0');
begin
loop
wait until rising_edge(wclk);
wait for 1 ns; -- read POST-edge values
if obs_en and wrst_n = '1' then
sdepth_obs <= sdepth_obs + 1;
if wq2 /= wq1_dly then
sdepth_bad_w <= sdepth_bad_w + 1;
report "write pointer synchroniser is not two flops deep"
severity error;
end if;
end if;
wq1_dly := wq1;
exit when done;
end loop;
wait;
end process obs_depth_w;
obs_depth_r : process
alias rq1 is << signal .tb_uart_async_fifo.u8.rq1_wgray : unsigned(PW-1 downto 0) >>;
alias rq2 is << signal .tb_uart_async_fifo.u8.rq2_wgray : unsigned(PW-1 downto 0) >>;
variable rq1_dly : unsigned(PW-1 downto 0) := (others => '0');
begin
loop
wait until rising_edge(rclk);
wait for 1 ns;
if obs_en and rrst_n = '1' then
if rq2 /= rq1_dly then
sdepth_bad_r <= sdepth_bad_r + 1;
report "read pointer synchroniser is not two flops deep"
severity error;
end if;
end if;
rq1_dly := rq1;
exit when done;
end loop;
wait;
end process obs_depth_r;
watchdog : process
begin
wait for 500 us;
report "watchdog: simulation did not finish" severity failure;
end process watchdog;
stim : process
variable checks, failures : natural := 0;
variable lat, guard : natural := 0;
variable base_w, base_r : natural := 0;
variable base_m : natural := 0;
procedure check(cond : boolean; name : string) is
begin
checks := checks + 1;
if cond then
report " PASS " & name severity note;
else
failures := failures + 1;
report " FAIL " & name severity error;
end if;
end procedure check;
-- Run one traffic phase at a given clock ratio and drain it.
-- The 1 ns after each edge matters twice over: it lets the loop read
-- SETTLED counters, and it moves the deassertion of w_en off the very
-- edge the DUT samples wpush on. Without it an extra word slips in at
-- the boundary and the phase writes nwords+1.
procedure run_phase(wh : time; rh : time; nwords : natural) is
variable bw, g : natural := 0;
begin
whalf <= wh; rhalf <= rh;
bw := wnext;
w_en <= '1'; r_en <= '1';
g := 0;
while (wnext - bw) < nwords and g < 200000 loop
wait until rising_edge(wclk); wait for 1 ns; g := g + 1;
end loop;
w_en <= '0';
g := 0;
while rexp /= wnext and g < 200000 loop
wait until rising_edge(rclk); wait for 1 ns; g := g + 1;
end loop;
for i in 1 to 6 loop wait until rising_edge(rclk); end loop;
r_en <= '0';
end procedure run_phase;
begin
report "== uart_async_fifo : self-checking VHDL testbench ==" severity note;
wrst_n <= '0'; rrst_n <= '0'; s_wrst_n <= '0'; s_rrst_n <= '0';
for i in 1 to 8 loop wait until rising_edge(wclk); end loop;
check(rempty = '1' and wfull = '0',
"out of reset: empty asserted, full deasserted");
check(s_empty = '1' and s_full = '0',
"the DEPTH=2 FIFO resets the same way");
wrst_n <= '1'; rrst_n <= '1'; s_wrst_n <= '1'; s_rrst_n <= '1';
for i in 1 to 8 loop wait until rising_edge(rclk); end loop;
obs_en <= true;
wait for 1 ns;
--=== phase 1: writer much faster than reader ========================
run_phase(3 ns, 11 ns, 40);
check(mism = 0, "fast-write / slow-read: sequence intact");
check(wnext = 40, "fast-write / slow-read: exactly 40 written");
check(rexp = wnext, "fast-write / slow-read: nothing left behind");
check(rempty = '1', "fast-write / slow-read: drained to empty");
check(wfull = '0', "and full has released");
--=== phase 2: reader much faster than writer ========================
base_r := rexp;
run_phase(11 ns, 3 ns, 40);
check(mism = 0, "slow-write / fast-read: sequence intact");
check(rexp - base_r = 40, "slow-write / fast-read: all 40 words arrived");
--=== phase 3: near-equal but incommensurate =========================
run_phase(7 ns, 5 ns, 60);
check(mism = 0, "incommensurate ratio: sequence intact");
check(rexp = 140, "incommensurate ratio: 40+40+60 words total");
check(rexp = wnext, "incommensurate ratio: counts agree");
--=== the extra pointer bit: capacity is exactly DEPTH ===============
whalf <= 5 ns; rhalf <= 7 ns;
base_w := wnext;
r_en <= '0'; -- reader stalled: nothing can drain
w_en <= '1';
guard := 0;
while wfull /= '1' and guard < 200 loop
wait until rising_edge(wclk); wait for 1 ns; guard := guard + 1;
end loop;
check(wfull = '1', "reader stalled: full eventually asserts");
for i in 1 to 4 loop wait until rising_edge(wclk); end loop;
check(wnext - base_w = DEPTH,
"exactly DEPTH words were accepted -- not DEPTH-1, not DEPTH+1");
--=== a full FIFO refuses further writes =============================
base_w := wnext;
w_force <= '1'; -- push HARD, ignoring wfull
for i in 1 to 20 loop wait until rising_edge(wclk); end loop;
w_force <= '0'; w_en <= '0';
wait for 1 ns;
check(wnext - base_w = 0, "20 forced pushes into a full FIFO: all ignored");
--=== and the DEPTH words are still the right DEPTH words ============
base_r := rexp;
r_en <= '1';
guard := 0;
while rexp /= wnext and guard < 2000 loop
wait until rising_edge(rclk); wait for 1 ns; guard := guard + 1;
end loop;
for i in 1 to 6 loop wait until rising_edge(rclk); end loop;
r_en <= '0';
check(rexp - base_r = DEPTH, "draining recovers exactly DEPTH words");
check(mism = 0, "and they are the right words, in order");
check(rempty = '1', "empty again");
--=== visibility latency across the boundary =========================
-- One word into an empty FIFO. The datum is in memory on the writing
-- edge; the READ side may not believe it until the gray pointer has
-- crossed two synchroniser flops. That cost is the price of safety.
w_en <= '1';
wait until rising_edge(wclk);
wait until falling_edge(wclk); w_en <= '0';
lat := 0;
while rempty /= '0' and lat < 12 loop
wait until rising_edge(rclk); wait for 0.1 ns; lat := lat + 1;
end loop;
check(lat >= 2, "a new word costs at least TWO rclk edges to become visible");
check(lat <= 5, "and the crossing is bounded, not open-ended");
r_en <= '1';
guard := 0;
while rexp /= wnext and guard < 200 loop
wait until rising_edge(rclk); wait for 1 ns; guard := guard + 1;
end loop;
r_en <= '0';
check(mism = 0, "the single word read back correctly");
--=== the DEPTH=2 FIFO: capacity two, and it really is two ===========
for i in 0 to 5 loop
wait until falling_edge(wclk);
s_wdata <= std_logic_vector(to_unsigned(16#A0# + i, WIDTH));
if s_full = '0' then s_push <= '1'; else s_push <= '0'; end if;
end loop;
wait until falling_edge(wclk); s_push <= '0';
for i in 1 to 4 loop wait until rising_edge(wclk); end loop;
check(s_full = '1', "DEPTH=2 FIFO fills");
for i in 1 to 6 loop wait until rising_edge(rclk); end loop;
check(s_empty = '0', "and the read side sees its contents");
wait until falling_edge(rclk); s_pop <= '1';
for i in 1 to 2 loop wait until rising_edge(rclk); end loop;
wait until falling_edge(rclk); s_pop <= '0';
for i in 1 to 8 loop wait until rising_edge(rclk); end loop;
check(s_empty = '1', "two pops empty it");
--=== the structural properties ======================================
check(wg_moves > 50 and rg_moves > 50,
"both gray pointers moved enough to be worth judging");
check(wg_bad = 0, "every write-pointer transition changed exactly ONE bit");
check(rg_bad = 0, "every read-pointer transition changed exactly ONE bit");
check(maxout <= DEPTH,
"outstanding words never exceeded DEPTH -- no silent overwrite");
check(minout >= 0,
"outstanding words never went negative -- nothing read before written");
check(sdepth_obs > 300,
"the synchroniser-depth observer ran on enough edges to judge");
check(sdepth_bad_w = 0,
"write-side pointer synchroniser is TWO flops deep, not one");
check(sdepth_bad_r = 0,
"read-side pointer synchroniser is TWO flops deep, not one");
--=== reset during traffic ===========================================
obs_en <= false;
wait for 1 ns;
w_en <= '1'; r_en <= '1';
for i in 1 to 20 loop wait until rising_edge(wclk); end loop;
wrst_n <= '0'; rrst_n <= '0'; -- both domains reset together
w_en <= '0'; r_en <= '0';
for i in 1 to 6 loop wait until rising_edge(wclk); end loop;
check(rempty = '1' and wfull = '0',
"reset during traffic returns the FIFO to empty");
wrst_n <= '1'; rrst_n <= '1';
for i in 1 to 8 loop wait until rising_edge(rclk); end loop;
obs_en <= true;
-- mism cannot simply be zeroed here: it is driven by rproc, and a
-- second driver on an unresolved type (natural) is illegal in VHDL.
-- A baseline is better practice in every language anyway — it asks
-- "were there NEW mismatches after the reset", which is the actual
-- question, instead of destroying the earlier evidence.
base_m := mism;
wait for 1 ns;
run_phase(5 ns, 7 ns, 24);
check(mism = base_m, "traffic after reset is clean from the first word");
check(rexp = wnext, "and balanced");
check(wg_bad = 0 and rg_bad = 0,
"the gray property survived the reset too");
report "== " & integer'image(checks) & " checks, "
& integer'image(failures) & " failures ==" severity note;
if failures = 0 then
report " RESULT: ALL VHDL ASYNC-FIFO TESTS PASSED" severity note;
else
report " RESULT: VHDL ASYNC-FIFO TESTS FAILED" severity error;
end if;
done <= true;
wait;
end process stim;
end architecture sim;Thirty-six checks, and all three languages agree:
PASS out of reset: empty asserted, full deasserted
PASS the DEPTH=2 FIFO resets the same way
PASS fast-write / slow-read: sequence intact
PASS slow-write / fast-read: all 40 words arrived
PASS incommensurate ratio: 40+40+60 words total
PASS reader stalled: full eventually asserts
PASS exactly DEPTH words were accepted -- not DEPTH-1, not DEPTH+1
PASS 20 forced pushes into a full FIFO: all ignored
PASS draining recovers exactly DEPTH words
PASS a new word costs at least TWO rclk edges to become visible
PASS and the crossing is bounded, not open-ended
PASS DEPTH=2 FIFO fills
PASS every write-pointer transition changed exactly ONE bit
PASS every read-pointer transition changed exactly ONE bit
PASS outstanding words never exceeded DEPTH -- no silent overwrite
PASS write-side pointer synchroniser is TWO flops deep, not one
PASS read-side pointer synchroniser is TWO flops deep, not one
PASS reset during traffic returns the FIFO to empty
== 36 checks, 0 failures ==
Verilog-2001 : 36 checks, 0 failures
SystemVerilog : 36 checks, 0 failures
VHDL-2008 : 36 checks, 0 failures9. Verification
An asynchronous FIFO must be tested with clocks that have no rational relationship, or the test is not exercising the crossing. Two clocks at a 2:1 ratio produce a repeating edge pattern that visits a small set of phase relationships; 7 ns against 13 ns does not.
Sweep the ratio in both directions. Fast writer with slow reader exercises full and the write-side pointer crossing; the reverse exercises empty. A test where both sides run at similar rates and the FIFO hovers half-full exercises neither flag.
// Assertion — in the write domain. The pessimism argument, stated: full may
// be wrong, but only by claiming full when the FIFO is not.
property p_full_is_pessimistic;
@(posedge wclk) disable iff (!wrst_n)
wfull_o |-> (wbin_q - rbin_q_observed) >= 0;
endproperty
// Assertion — data integrity, which is what the crossing must not break.
// Every word read is the word written that many pushes ago, in order.
// Checked procedurally in the suite above: 0 mismatches over ~1500 transfers.
// Assertion — the depth is never exceeded. This is the one that catches a
// broken `full`, and it must be computed by the TESTBENCH, which is allowed
// to know both pointers at once. The DESIGN is not.
property p_never_overfilled;
@(posedge wclk) (tb_pushes - tb_pops) <= DEPTH;
endpropertyThe testbench may look at both domains; the design may not. That asymmetry is the whole point of a scoreboard here, and it is also the trap: it is easy to write a check that quietly assumes an ordering between the two clocks that does not exist. A scoreboard for an asynchronous FIFO should count events, not compare states at an instant.
Run CDC lint on the result (Chapter 12.2 §8). It verifies structurally what the simulation cannot: that every crossing has a synchroniser, that the gray encoding is actually in the crossing path, and that no combinational logic slipped between the stages.
Mutation testing, and the one defect that got away
Four defects were installed in uart_async_fifo, one at a time, each verified
to have actually changed the source before the result was scored.
| # | Defect installed | Result |
|---|---|---|
| M6 | pointers cross as binary, not gray | killed — the gray invariant fires, then the FIFO deadlocks |
| M7 | full compares plain equality, without the one-lap inversion | killed — the design hangs; caught by the watchdog |
| M8 | the write pointer advances even when full | killed, 5 checks |
| M9 | the pointer synchronisers cut to one flop | survived every behavioural check |
M6 and M7 are killed by deadlock rather than by a mismatch, which is why every loop in the testbench carries a guard and the run carries a watchdog. A test that hangs reports nothing; a test that hangs and times out reports a failure. The difference is a few lines and it is the difference between a result and a hung CI job.
M9 is the one that matters. Cutting each two-flop pointer synchroniser to a single flop is not a subtle defect — it is the exact hazard this entire module exists to prevent, and in silicon it is the difference between a FIFO that runs for years and one that corrupts a pointer when two edges land too close together.
It changed nothing a simulation could see:
--- M9: one-flop pointer synchronisers, behavioural checks only ---
PASS all 34 behavioural checks
the FIFO lost nothing, duplicated nothing, reordered nothing,
full and empty were correct, capacity was exactly DEPTHThe FIFO is still functionally perfect, because the second flop's job is to give the first one a clock in which to resolve — and in an event simulator there is nothing to resolve. Every data check, every flag check, every capacity check passes. The design is broken and the testbench is happy.
What finally caught it was the third structural observer: after any write-clock
edge, wq2_rgray must hold the value wq1_rgray had before that edge. With
one flop the two move together and the invariant breaks on the first pointer
movement.
--- M9 against the structural observer ---
FAIL write-side pointer synchroniser is TWO flops deep, not one
FAIL read-side pointer synchroniser is TWO flops deep, not one
== 36 checks, 2 failures ==Two checks out of thirty-six, and they are the only two in the file that could possibly have fired. Note what they do: they do not test behaviour at all. They test the shape of the logic, which is what a CDC linter does against a netlist — performed here with the only instrument a simulation has, a hierarchical reference and a one-deep delay.
10. What This Means on an FPGA
Ask first whether you need two clocks at all. A UART running on the system clock with a fractional baud generator (Chapter 8.3) removes this entire chapter from the design. That is frequently the right answer and it is frequently not considered, because "the UART should have its own clock" sounds tidy.
Vendor FIFO generators exist and are usually better than a hand-written one. Xilinx FIFO Generator and Intel's DCFIFO implement this structure, are verified, and map onto block RAM with the correct primitives. Write your own to understand it; instantiate theirs to ship it — unless you need something they do not offer.
The extra pointer bit is not optional. It is what distinguishes full from empty when the pointers are otherwise equal. A FIFO built with DEPTH_L2 pointer bits instead of DEPTH_L2 + 1 reports empty when it is full, and the symptom is catastrophic and intermittent.
Constrain the crossings. The gray pointer paths need set_max_delay or equivalent, not a default setup check between unrelated clocks — Chapter 12.5. Left alone, the tool will either fail timing on a path that has no real requirement, or meet an imaginary one while allowing skew that breaks the one-bit-at-a-time guarantee.
11. Understanding Check
12. Summary
Asynchronous serial and clock-domain crossing are different subjects wearing the same word. The first is solved by oversampling and is the whole protocol; the second is solved by synchronisers and, in the UART as built, applies to two input pins and nothing else.
Both confusions are expensive: CDC structures on the serial path add latency where margin is scarce, and a missed crossing produces intermittent failures that no test reproduces.
A real crossing appears when a second clock does — in practice, a register interface on the bus clock. Three kinds then need three mechanisms: levels take a synchroniser, pulses must become levels before crossing, and multi-bit values take gray code, a handshake, or a FIFO.
Gray code does not make a bad sample less likely — it makes it unrepresentable. Measured with per-bit skew: a binary pointer produced 215 fabricated values in 3,999 samples (5.38%), including reading 0 while the counter held 8. The gray pointer produced none, structurally.
The asynchronous FIFO keeps a pointer per domain, crosses them as gray, compares locally against a deliberately stale view, and registers both flags — which is required, not stylistic, because a combinational flag closes a real loop. Verified across 7 ns and 13 ns clocks with zero data mismatches.
The Module 10 FIFO cannot be adapted: its shared binary level counter is the exact failure the measurement demonstrates, and synchronising it repairs nothing.
And the question that resolves every case: is this signal generated by one clock and captured by another? Not whether it is slow, external, or has "asynchronous" in its name.
13. What Comes Next
Two clocks raise a second question this chapter has quietly set aside: what does reset mean when there are two of them?
Chapter 12.4 takes reset seriously — asynchronous assertion with synchronous release, why the release edge is the dangerous one, how reset is distributed across domains, and the question the whole module has been building toward: what a reset asserted mid-frame leaves behind, on both ends of a link where only one end was reset.
Browse the full path on the UART tutorials index. For the synchronous FIFO this chapter contrasts against, read back to Chapter 10.2.
Continue learning
Related tutorials
- Related topic
Why Receiving Is Harder Than Transmitting
A transmitter executes a schedule it wrote itself. A receiver must decide whether something is happening, whether it was real, where the positions are, and what value was there — four judgements from one edge on an input it does not control.
- Related topic
Metastability and the Asynchronous RX Input
Why the RX pin is a true asynchronous input, what metastability actually is, and — with the MTBF arithmetic worked out — what a synchroniser does and does not guarantee.
- Related topic
Synchroniser Placement and Edge Detection After Synchronisation
Where the synchroniser belongs, what its latency costs the start-edge measurement, and a real placement defect found in the IP built in Module 11 — which every functional test passed.
- Related topic
Reset Strategy and Reset During Active Traffic
Reset style and release, and what a reset asserted mid-frame actually leaves behind — measured, and worse than a framing error: well-formed frames carrying bytes nobody sent.
Where this fits
Part of the UART curriculum.
