Skip to content
VLSI Mentor

UART · Module 12

UART Clock Domains: What Is and Is Not CDC

Asynchronous serial timing is not clock-domain crossing. Where a UART genuinely needs CDC structures, why pointers cross as gray code — measured — and the asynchronous FIFO that follows.

"Asynchronous" is doing two completely different jobs in this subject, and conflating them produces both of the available mistakes: building CDC structures where none are needed, and missing the crossing that actually exists.

Asynchronous serial means the link carries no clock and the receiver recovers timing from the data — Modules 2 through 5. Clock-domain crossing means a signal generated by one clock is captured by another. A UART is asynchronous in the first sense by definition and, as built through Module 11, contains exactly one crossing in the second sense: the receive pin.

1. Two Different Meanings of One Word

Asynchronous serialClock-domain crossing
What is asynchronousthe link — no clock is transmittedtwo clocks inside one chip
The problemrecovering bit timing from datacapturing a signal safely
The solutionoversampling and start-edge alignmentsynchronisers, gray code, handshakes
Where it livesModules 2–5this module
How many in the UARTthe whole protocolone signal, until a second clock arrives

Oversampling is not a CDC technique. It solves a frequency and phase mismatch between two devices that each have a stable clock — the far end's bit rate versus this design's sampling grid — and it solves it statistically, by sampling near bit centres. It does nothing whatsoever about metastability, which is why the receiver has a synchroniser in addition.

And a synchroniser is not timing recovery. It gives a settled value quantised to a clock edge, and knows nothing about where bit boundaries are.

2. The UART As Built: One Crossing

Taking the inventory honestly for the IP of Chapter 11.2, which runs entirely on clk:

SignalCrosses?Why
rx_iYESdriven by a device with its own oscillator
tx_onodriven by clk; the far end's problem, not ours
cts_n_iYESa pin, driven by the far end
rts_n_onoan output of this domain
FIFO pointers and levelsnoone clock on both ports
os_tick, baud_ticknoenables in this domain, not clocks — Chapter 8.1
configuration, statusnosame clock as everything else

Two input pins cross, and both are single-bit levels, which is the easy case. rx_i is synchronised inside uart_rx (Chapter 12.2); cts_n_i is synchronised inside uart_tx_gate (Chapter 10.5 §7). Nothing else in the design crosses anything.

That is the whole CDC story for a single-clock UART, and it is worth stating plainly because a great deal of anxious complexity gets added to designs that are in exactly this position.

3. Where a Real Crossing Appears

It appears when a second clock does — and for a UART that is essentially always the register interface. The serial side wants a clock with a convenient relationship to the baud rate; the bus wants the system clock; these are rarely the same, and forcing them to be is a real constraint on the rest of the SoC.

With two clocks, three separate kinds of crossing appear, and they need three different mechanisms:

What crossesMechanismWhy not the others
a single-bit level (an enable, a status flag)two-flop synchronisernothing else is needed
a single-bit event (a one-cycle pulse)toggle, then synchronise, then edge-detecta pulse shorter than the destination period is lost
a multi-bit value (data, pointers)gray code, or a handshake, or an async FIFO§4 — independent bits arrive at independent times

4. Why a Multi-Bit Value Crosses as Gray Code

The standard explanation is that gray code changes one bit per increment, so a sample taken during a transition yields one of the two adjacent values. That is correct and it understates how badly the alternative behaves.

Measured. A 5-bit counter in one domain, sampled by an unrelated clock, with a different routing delay on each of the five wires — which is what a real bus does. Both a binary and a gray-coded version of the same counter, sampled at the same instants, over 3,999 samples:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
  BINARY  sampled 0 (00000) — source held 9/8/7 : IMPOSSIBLE
  BINARY  sampled 7 (00111) — source held 4/3/2 : IMPOSSIBLE
  BINARY  sampled 19 (10011) — source held 24/23/22 : IMPOSSIBLE
  BINARY  sampled 24 (11000) — source held 17/16/15 : IMPOSSIBLE

  3999 samples taken across two unrelated clocks, with per-bit skew
  BINARY pointer : 215 samples were values the counter NEVER held  (5.38%)
  GRAY   pointer : 0 samples were values the counter NEVER held  (0.00%)

One sample in nineteen of the binary pointer was a value that never existed. Not a stale value — stale would be harmless — a fabricated one. Reading 0 while the counter held 8 is the case that destroys a FIFO: the reader concludes the queue is empty and stops, while eight entries sit in it.

The gray-coded counter produced none, and can produce none. With one bit changing per increment, a sample during the transition resolves that single bit to either its old or its new value, and both of those are real counter values. The property is structural rather than probabilistic.

5. The Asynchronous FIFO

Put together: arbitrary data crossing a clock boundary, with both sides free-running. The write side needs to know whether the FIFO is full and the read side whether it is empty, and neither may read the other's pointer directly.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// ===========================================================================
//  A genuine asynchronous FIFO.
//
//  This is NOT the FIFO of Module 10. That one is synchronous: one clock, one
//  level counter, and a comparison any part of the design may read. Here the
//  write and read sides are in unrelated clock domains, and NOTHING about the
//  other side's state may be read directly.
//
//  Three structural decisions carry the whole design:
//    1. Pointers are one bit WIDER than the address, so full and empty are
//       distinguishable rather than both meaning "pointers equal".
//    2. Pointers cross as GRAY code, so a sample taken during a transition is
//       always one of the two adjacent values and never a third.
//    3. Each side computes its own flag PESSIMISTICALLY from a stale view of
//       the other, so the error is always in the safe direction.
//    4. Both flags are REGISTERED. That is not a style choice: the next
//       pointer value depends on the flag and the flag depends on the next
//       pointer value, so a combinational flag closes a genuine loop.
// ===========================================================================
module uart_async_fifo #(
    parameter int unsigned WIDTH    = 8,
    parameter int unsigned DEPTH_L2 = 4          // DEPTH = 2**DEPTH_L2
) (
    // write domain
    input  logic             wclk,
    input  logic             wrst_n,
    input  logic [WIDTH-1:0] wdata_i,
    input  logic             wpush_i,
    output logic             wfull_o,
    // read domain
    input  logic             rclk,
    input  logic             rrst_n,
    output logic [WIDTH-1:0] rdata_o,
    input  logic             rpop_i,
    output logic             rempty_o
);
    localparam int unsigned DEPTH = 1 << DEPTH_L2;
    localparam int unsigned   PW  = DEPTH_L2 + 1;  // one extra bit — see (1)
    localparam logic [PW-1:0] LAP = 3 << (PW-2);   // the top two bits set

    if (DEPTH_L2 < 1) begin : g_bad_depth
        $error("uart_async_fifo: DEPTH_L2 must be at least 1");
    end

    logic [WIDTH-1:0] mem [0:DEPTH-1];

    // Both pointers declared up front: each domain's synchroniser reads the
    // other's gray register, so neither can be declared after its reader.
    logic [PW-1:0] wgray_q, rgray_q;

    // ---- write domain ------------------------------------------------------
    logic [PW-1:0] wbin_q, wbin_nxt, wgray_nxt;
    logic [PW-1:0] wq1_rgray, wq2_rgray;

    assign wbin_nxt  = wbin_q + (wpush_i && !wfull_o);
    assign wgray_nxt = (wbin_nxt >> 1) ^ wbin_nxt;          // binary -> gray

    always_ff @(posedge wclk or negedge wrst_n)
        if (!wrst_n) begin wbin_q <= '0; wgray_q <= '0; end
        else         begin wbin_q <= wbin_nxt; wgray_q <= wgray_nxt; end

    always_ff @(posedge wclk)
        if (wpush_i && !wfull_o) mem[wbin_q[DEPTH_L2-1:0]] <= wdata_i;

    // The read pointer, synchronised into the WRITE domain. Two flops, and
    // the ONLY thing this domain is allowed to know about the other one.
    always_ff @(posedge wclk or negedge wrst_n)
        if (!wrst_n) begin wq1_rgray <= '0; wq2_rgray <= '0; end
        else         begin wq1_rgray <= rgray_q; wq2_rgray <= wq1_rgray; end

    // Full: the write pointer is one lap ahead of the read pointer. In binary
    // that is "low bits equal, MSB different". The gray image of that is "the
    // two vectors differ in exactly the top TWO bits and nowhere else", which
    // is what LAP tests. Written this way rather than as
    //     wgray_nxt == {~wq2_rgray[PW-1:PW-2], wq2_rgray[PW-3:0]}
    // because that spelling contains the part-select [PW-3:0], which goes
    // negative at PW=2 and makes the smallest legal FIFO (DEPTH_L2=1) refuse
    // to elaborate. The XOR form has no part-select and is correct for every
    // PW >= 2.
    // REGISTERED, see (4): wbin_nxt consumes the registered wfull_o, so the
    // path flag -> next-pointer -> flag is broken by a flop.
    logic wfull_d;
    assign wfull_d = ((wgray_nxt ^ wq2_rgray) == LAP);
    always_ff @(posedge wclk or negedge wrst_n)
        if (!wrst_n) wfull_o <= 1'b0; else wfull_o <= wfull_d;

    // ---- read domain -------------------------------------------------------
    logic [PW-1:0] rbin_q, rbin_nxt, rgray_nxt;
    logic [PW-1:0] rq1_wgray, rq2_wgray;

    assign rbin_nxt  = rbin_q + (rpop_i && !rempty_o);
    assign rgray_nxt = (rbin_nxt >> 1) ^ rbin_nxt;

    always_ff @(posedge rclk or negedge rrst_n)
        if (!rrst_n) begin rbin_q <= '0; rgray_q <= '0; end
        else         begin rbin_q <= rbin_nxt; rgray_q <= rgray_nxt; end

    assign rdata_o = mem[rbin_q[DEPTH_L2-1:0]];

    always_ff @(posedge rclk or negedge rrst_n)
        if (!rrst_n) begin rq1_wgray <= '0; rq2_wgray <= '0; end
        else         begin rq1_wgray <= wgray_q; rq2_wgray <= rq1_wgray; end

    // Empty: the read pointer has caught the (stale) write pointer exactly.
    // Registered for the same reason, and reset ASSERTED — an empty FIFO is
    // the safe claim to make before anything has been written.
    logic rempty_d;
    assign rempty_d = (rgray_nxt == rq2_wgray);
    always_ff @(posedge rclk or negedge rrst_n)
        if (!rrst_n) rempty_o <= 1'b1; else rempty_o <= rempty_d;
endmodule

Decision 3 is the one worth dwelling on. Each side sees the other's pointer two clocks late, so each side's view is stale — and the staleness is deliberately arranged to be conservative. The write side's full is computed against an out-of-date read pointer, so it can claim full when the reader has actually freed a slot: it refuses a write it could have accepted. The read side's empty is computed against an out-of-date write pointer, so it can claim empty when a word has actually been written: it delays a read it could have made.

Both errors cost throughput and neither can corrupt data. That asymmetry is the entire safety argument, and it is why the flags are never "corrected" by exchanging more information.

A block diagram of an asynchronous FIFO. On the left is the write domain, clocked by the write clock: a write pointer counts in binary, is converted to gray code, and addresses the shared memory for writing. On the right is the read domain, clocked by the read clock: a read pointer counts in binary, is converted to gray code, and addresses the same memory for reading. The write domain's gray pointer crosses through a two-flop synchroniser into the read domain, where it is compared against the read pointer to produce the empty flag. The read domain's gray pointer crosses through a two-flop synchroniser into the write domain, where it is compared against the write pointer to produce the full flag. Both flags are registered. The memory sits between the two domains and is the only shared storage.write ptrbinary, +1 per pushto grayone bit changesmemorywritten by wclkto grayone bit changesread ptrbinary, +1 per pop2FF syncwgray into rclk2FF syncrgray into wclkrempty_oregistered,pessimisticwfull_oregistered,pessimisticencodecrossescompareencodecrossescompareaddrdata12
Figure 1 — the asynchronous FIFO. The memory is written by one clock and read by the other. Neither domain reads the other's pointer directly: each pointer is converted to gray code, crosses through a two-flop synchroniser, and is compared locally. Both flags are computed from a deliberately stale view of the far pointer, so each errs only toward refusing work.

Verified

Two deliberately incommensurate clocks — 7 ns and 13 ns, with no repeating edge pattern — and randomised push and pop pressure on both sides:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
  wrote 1499, read 1485, in flight 14
  pass data comes out in order, unmodified
  pass throughput achieved (over a thousand transfers)
  pass never over-filled: in flight <= DEPTH
  pass never under-ran: read count never exceeds write count
  pass drains to empty

== 5 checks, 0 failures ==

6. What the Module 10 FIFO Is Not

The synchronous FIFO of Chapter 10.2 must not be used across clock domains, and the reason is precisely §4:

It maintains a binary occupancy counter that both ports read. Drive its two ports from different clocks and that counter is a multi-bit binary value crossing a domain boundary — the exact case that produced 215 fabricated samples out of 3,999. A FIFO whose level register can read as 0 while it holds 8 entries will report empty and lose them.

Adding synchronisers to it does not repair this. Synchronising each bit of a binary counter independently is what the experiment did; the bits are still independent and the fabricated combinations still occur. The structure has to change — separate pointers per domain, gray encoding, no shared counter — which is a different module, not a modification.

Use the right one for the situation. Single clock: the Module 10 FIFO, which is smaller, has an exact occupancy for watermarks, and reports full and empty with no latency. Two clocks: this one, which costs two synchroniser stages of latency in each direction and cannot give either side an exact occupancy.

7. The Same FIFO in Verilog-2001 and VHDL-2008

The four decisions above are structural, not linguistic. They survive translation intact, and translating them is a useful exercise precisely because it forces each one to be stated in a language that will not let you get away with hand-waving.

Verilog-2001 carries the design with no difficulty. There is no logic, no int unsigned, and no always_ff, and none of that matters here: every signal is a vector of bits and every process is a flop.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
//===========================================================================
//  uart_async_fifo_v — Synthesizable Verilog-2001
//
//  A genuine asynchronous FIFO. This is NOT the FIFO of Module 10: that one
//  is synchronous, with one clock and one level counter any part of the
//  design may read. Here the write and read sides are in unrelated clock
//  domains and NOTHING about the other side's state may be read directly.
//
//  Four structural decisions carry the whole design:
//    1. Pointers are one bit WIDER than the address, so full and empty are
//       distinguishable rather than both meaning "pointers equal".
//    2. Pointers cross as GRAY code, so a sample taken during a transition
//       is always one of the two adjacent values and never a third.
//    3. Each side computes its own flag PESSIMISTICALLY from a stale view of
//       the other, so the error is always in the safe direction.
//    4. Both flags are REGISTERED. Not a style choice: the next pointer
//       depends on the flag and the flag depends on the next pointer, so a
//       combinational flag closes a genuine loop.
//===========================================================================
module uart_async_fifo_v #(
    parameter WIDTH    = 8,
    parameter DEPTH_L2 = 4          // DEPTH = 2**DEPTH_L2; MUST be >= 1
) (
    // write domain
    input  wire             wclk,
    input  wire             wrst_n,
    input  wire [WIDTH-1:0] wdata_i,
    input  wire             wpush_i,
    output reg              wfull_o,
    // read domain
    input  wire             rclk,
    input  wire             rrst_n,
    output wire [WIDTH-1:0] rdata_o,
    input  wire             rpop_i,
    output reg              rempty_o
);
    localparam DEPTH = 1 << DEPTH_L2;
    localparam PW    = DEPTH_L2 + 1;      // one extra bit -- see (1)
    localparam [PW-1:0] LAP = 3 << (PW-2);   // the top two bits set

    reg [WIDTH-1:0] mem [0:DEPTH-1];

    // Both gray pointers declared up front: each domain's synchroniser reads
    // the other's gray register, so neither can be declared after its reader.
    reg [PW-1:0] wgray_q, rgray_q;

    //---- write domain -----------------------------------------------------
    reg  [PW-1:0] wbin_q, wq1_rgray, wq2_rgray;
    wire [PW-1:0] wbin_nxt  = wbin_q + (wpush_i && !wfull_o);
    wire [PW-1:0] wgray_nxt = (wbin_nxt >> 1) ^ wbin_nxt;        // binary -> gray

    always @(posedge wclk or negedge wrst_n) begin
        if (!wrst_n) begin
            wbin_q  <= {PW{1'b0}};
            wgray_q <= {PW{1'b0}};
        end else begin
            wbin_q  <= wbin_nxt;
            wgray_q <= wgray_nxt;
        end
    end

    always @(posedge wclk)
        if (wpush_i && !wfull_o) mem[wbin_q[DEPTH_L2-1:0]] <= wdata_i;

    // The read pointer, synchronised into the WRITE domain. Two flops, and
    // the ONLY thing this domain is allowed to know about the other one.
    always @(posedge wclk or negedge wrst_n) begin
        if (!wrst_n) begin
            wq1_rgray <= {PW{1'b0}};
            wq2_rgray <= {PW{1'b0}};
        end else begin
            wq1_rgray <= rgray_q;
            wq2_rgray <= wq1_rgray;
        end
    end

    // Full: the write pointer is one lap ahead of the read pointer. In binary
    // that is "low bits equal, MSB different". The gray image of that is "the
    // two vectors differ in exactly the top TWO bits and nowhere else", which
    // is what LAP tests. Written this way rather than as
    //     wgray_nxt == {~wq2_rgray[PW-1:PW-2], wq2_rgray[PW-3:0]}
    // because that spelling contains the part-select [PW-3:0], which goes
    // negative at PW=2 and makes the smallest legal FIFO (DEPTH_L2=1) refuse
    // to elaborate. The XOR form has no part-select and is correct for every
    // PW >= 2.
    // REGISTERED, see (4): wbin_nxt consumes the registered wfull_o, so the
    // path flag -> next-pointer -> flag is broken by a flop.
    wire wfull_d = ((wgray_nxt ^ wq2_rgray) == LAP);
    always @(posedge wclk or negedge wrst_n)
        if (!wrst_n) wfull_o <= 1'b0; else wfull_o <= wfull_d;

    //---- read domain ------------------------------------------------------
    reg  [PW-1:0] rbin_q, rq1_wgray, rq2_wgray;
    wire [PW-1:0] rbin_nxt  = rbin_q + (rpop_i && !rempty_o);
    wire [PW-1:0] rgray_nxt = (rbin_nxt >> 1) ^ rbin_nxt;

    always @(posedge rclk or negedge rrst_n) begin
        if (!rrst_n) begin
            rbin_q  <= {PW{1'b0}};
            rgray_q <= {PW{1'b0}};
        end else begin
            rbin_q  <= rbin_nxt;
            rgray_q <= rgray_nxt;
        end
    end

    assign rdata_o = mem[rbin_q[DEPTH_L2-1:0]];

    always @(posedge rclk or negedge rrst_n) begin
        if (!rrst_n) begin
            rq1_wgray <= {PW{1'b0}};
            rq2_wgray <= {PW{1'b0}};
        end else begin
            rq1_wgray <= wgray_q;
            rq2_wgray <= rq1_wgray;
        end
    end

    // Empty: the read pointer has caught the (stale) write pointer exactly.
    // Registered for the same reason, and reset ASSERTED -- an empty FIFO is
    // the safe claim to make before anything has been written.
    wire rempty_d = (rgray_nxt == rq2_wgray);
    always @(posedge rclk or negedge rrst_n)
        if (!rrst_n) rempty_o <= 1'b1; else rempty_o <= rempty_d;
endmodule

VHDL is the one that argues back, in three places worth knowing about.

It will not let the architecture read its own outputs. wfull_o is an out port, and wbin_nxt needs its value. So the design keeps an internal wfull_s, drives the port from it concurrently, and reads the internal signal everywhere. This is pure ceremony in VHDL-2008 — the language relaxed the rule in VHDL-2019 — but it makes the data flow explicit in a way the Verilog does not.

Gray conversion is a shift and an XOR, spelled with functions. wbin_nxt srl 1 is not available on unsigned in the way you might expect, so the code says shift_right(wbin_nxt, 1) xor wbin_nxt, which is exactly the (wbin_nxt >> 1) ^ wbin_nxt of the other two.

Its assert is the best parameter guard of the three languages. A concurrent assertion with severity failure is checked at elaboration and needs no process, no clock and no simulation time — it simply refuses to build a FIFO whose depth is illegal.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
--==========================================================================
--  uart_async_fifo -- Synthesizable VHDL-2008
--
--  A genuine asynchronous FIFO. The write and read sides are in unrelated
--  clock domains and NOTHING about the other side's state may be read
--  directly.
--
--  Four structural decisions carry the whole design:
--    1. Pointers are one bit WIDER than the address, so full and empty are
--       distinguishable rather than both meaning "pointers equal".
--    2. Pointers cross as GRAY code, so a sample taken during a transition
--       is always one of the two adjacent values and never a third.
--    3. Each side computes its own flag PESSIMISTICALLY from a stale view of
--       the other, so the error is always in the safe direction.
--    4. Both flags are REGISTERED -- a combinational flag closes a genuine
--       loop through the next-pointer arithmetic.
--
--  wfull and rempty are read by the architecture (the next pointer depends
--  on them), and VHDL forbids reading an out port, so each is an internal
--  signal with a concurrent driver.
--==========================================================================
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

entity uart_async_fifo is
    generic (
        WIDTH    : positive := 8;
        DEPTH_L2 : positive := 4        -- DEPTH = 2**DEPTH_L2
    );
    port (
        -- write domain
        wclk     : in  std_logic;
        wrst_n   : in  std_logic;
        wdata_i  : in  std_logic_vector(WIDTH-1 downto 0);
        wpush_i  : in  std_logic;
        wfull_o  : out std_logic;
        -- read domain
        rclk     : in  std_logic;
        rrst_n   : in  std_logic;
        rdata_o  : out std_logic_vector(WIDTH-1 downto 0);
        rpop_i   : in  std_logic;
        rempty_o : out std_logic
    );
end entity uart_async_fifo;

architecture rtl of uart_async_fifo is

    constant DEPTH : positive := 2 ** DEPTH_L2;
    constant PW    : positive := DEPTH_L2 + 1;   -- one extra bit -- see (1)
    -- "one lap ahead" in gray code: the two vectors differ in exactly the
    -- top two bits. Expressed as a mask so that no slice of the pointer has
    -- to be taken -- a (PW-3 downto 0) slice goes null at PW=2 and breaks
    -- the smallest legal FIFO.
    constant LAP   : unsigned(PW-1 downto 0)
                   := shift_left(to_unsigned(3, PW), PW-2);

    type mem_t is array (0 to DEPTH-1) of std_logic_vector(WIDTH-1 downto 0);
    signal mem : mem_t;

    -- Both gray pointers declared up front: each domain's synchroniser reads
    -- the other's gray register.
    signal wgray_q, rgray_q : unsigned(PW-1 downto 0) := (others => '0');

    signal wbin_q, wq1_rgray, wq2_rgray : unsigned(PW-1 downto 0) := (others => '0');
    signal wbin_nxt, wgray_nxt          : unsigned(PW-1 downto 0);
    signal wfull_s, wfull_d             : std_logic;

    signal rbin_q, rq1_wgray, rq2_wgray : unsigned(PW-1 downto 0) := (others => '0');
    signal rbin_nxt, rgray_nxt          : unsigned(PW-1 downto 0);
    signal rempty_s, rempty_d           : std_logic;


begin

    assert DEPTH_L2 >= 1
        report "uart_async_fifo: DEPTH_L2 must be at least 1" severity failure;

    wfull_o  <= wfull_s;
    rempty_o <= rempty_s;

    ------------------------------------------------------------------ write
    wbin_nxt  <= wbin_q + 1 when (wpush_i = '1' and wfull_s = '0') else wbin_q;
    wgray_nxt <= shift_right(wbin_nxt, 1) xor wbin_nxt;         -- binary -> gray

    process (wclk, wrst_n)
    begin
        if wrst_n = '0' then
            wbin_q  <= (others => '0');
            wgray_q <= (others => '0');
        elsif rising_edge(wclk) then
            wbin_q  <= wbin_nxt;
            wgray_q <= wgray_nxt;
        end if;
    end process;

    process (wclk)
    begin
        if rising_edge(wclk) then
            if wpush_i = '1' and wfull_s = '0' then
                mem(to_integer(wbin_q(DEPTH_L2-1 downto 0))) <= wdata_i;
            end if;
        end if;
    end process;

    -- The read pointer, synchronised into the WRITE domain. Two flops, and
    -- the ONLY thing this domain is allowed to know about the other one.
    process (wclk, wrst_n)
    begin
        if wrst_n = '0' then
            wq1_rgray <= (others => '0');
            wq2_rgray <= (others => '0');
        elsif rising_edge(wclk) then
            wq1_rgray <= rgray_q;
            wq2_rgray <= wq1_rgray;
        end if;
    end process;

    -- Full: pointers equal in the low bits, MSBs differ. In gray code that
    -- is the top TWO bits inverted -- the gray equivalent of "one lap ahead".
    wfull_d <= '1' when (wgray_nxt xor wq2_rgray) = LAP else '0';

    process (wclk, wrst_n)
    begin
        if wrst_n = '0' then
            wfull_s <= '0';
        elsif rising_edge(wclk) then
            wfull_s <= wfull_d;
        end if;
    end process;

    ------------------------------------------------------------------- read
    rbin_nxt  <= rbin_q + 1 when (rpop_i = '1' and rempty_s = '0') else rbin_q;
    rgray_nxt <= shift_right(rbin_nxt, 1) xor rbin_nxt;

    process (rclk, rrst_n)
    begin
        if rrst_n = '0' then
            rbin_q  <= (others => '0');
            rgray_q <= (others => '0');
        elsif rising_edge(rclk) then
            rbin_q  <= rbin_nxt;
            rgray_q <= rgray_nxt;
        end if;
    end process;

    rdata_o <= mem(to_integer(rbin_q(DEPTH_L2-1 downto 0)));

    process (rclk, rrst_n)
    begin
        if rrst_n = '0' then
            rq1_wgray <= (others => '0');
            rq2_wgray <= (others => '0');
        elsif rising_edge(rclk) then
            rq1_wgray <= wgray_q;
            rq2_wgray <= rq1_wgray;
        end if;
    end process;

    -- Empty: the read pointer has caught the (stale) write pointer exactly.
    -- Reset ASSERTED -- an empty FIFO is the safe claim before anything has
    -- been written.
    rempty_d <= '1' when rgray_nxt = rq2_wgray else '0';

    process (rclk, rrst_n)
    begin
        if rrst_n = '0' then
            rempty_s <= '1';
        elsif rising_edge(rclk) then
            rempty_s <= rempty_d;
        end if;
    end process;

end architecture rtl;

8. Testing a FIFO Whose Two Halves Never Agree

The testbench has a problem the Module 10 FIFO's did not: there is no moment at which a correct answer exists. Ask "how many entries are in the FIFO" and the write side and the read side will give different numbers, both correct, because each is describing a different instant. A scoreboard that compares against occupancy is therefore comparing against a fiction.

So the oracle is built on the only thing both sides can agree about — the sequence. The writer emits a counting sequence. The reader keeps its own counter and demands that every word it pops be the next value of that sequence. No shared model, no peeking at pointers, and every failure mode lands on the same check:

  • a lost word → the next compare mismatches
  • a duplicated word → the next compare mismatches
  • a reordered word → the next compare mismatches

and at the end of each phase, drained, the two counters must be equal.

On top of that the suite makes the clocks genuinely unrelated — periods are held in variables and changed at run time to 3/11, 11/3, 7/5 and 5/7 — so the phase relationship drifts continuously and the crossing is exercised at every alignment rather than at one.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
//===========================================================================
//  tb_uart_async_fifo — self-checking SystemVerilog testbench
//
//  The two clocks here are UNRELATED. Their periods are set at run time and
//  deliberately chosen to be incommensurate, so the phase relationship keeps
//  drifting and the design is exercised at every alignment rather than at one.
//
//  THE ORACLE IS INDEPENDENT. The writer emits a counting sequence; the
//  reader keeps its OWN counter and demands that every popped word be the
//  next value of that sequence. No shared model, no peeking at pointers:
//     a lost word   → the reader's next compare mismatches
//     a duplicated  → the reader's next compare mismatches
//     a reordered   → the reader's next compare mismatches
//  and at the end of each phase, drained, the two counters must be equal.
//
//  Two properties are checked STRUCTURALLY, through hierarchical references,
//  because they are the reason the design is shaped the way it is:
//     - every gray pointer transition changes exactly ONE bit (decision 2)
//     - outstanding words never exceed DEPTH        (decision 1, the extra bit)
//
//  Same 36 counted checks as the Verilog twin.
//===========================================================================
`timescale 1ns/1ps

module tb_uart_async_fifo;

    localparam int WIDTH = 8;
    localparam int DL2   = 3;
    localparam int DEPTH = 1 << DL2;        // 8
    localparam int PW    = DL2 + 1;         // 4

    //---- two independent clocks, periods settable at run time --------------
    int   whalf = 5, rhalf = 7;
    logic wclk = 1'b0, rclk = 1'b0;
    always #(whalf) wclk = ~wclk;
    always #(rhalf) rclk = ~rclk;

    logic wrst_n = 1'b0, rrst_n = 1'b0;
    logic w_en = 1'b0, r_en = 1'b0, w_force = 1'b0;

    logic [WIDTH-1:0] wdata, rdata;
    logic             wfull, rempty;
    wire              wpush = w_force ? 1'b1 : (w_en && !wfull);
    wire              rpop  = r_en && !rempty;

    int   wnext = 0, rexp = 0, mism = 0;
    assign wdata = wnext[WIDTH-1:0];

    uart_async_fifo #(.WIDTH(WIDTH), .DEPTH_L2(DL2)) u8 (
        .wclk(wclk), .wrst_n(wrst_n), .wdata_i(wdata), .wpush_i(wpush), .wfull_o(wfull),
        .rclk(rclk), .rrst_n(rrst_n), .rdata_o(rdata), .rpop_i(rpop),  .rempty_o(rempty));

    // the SMALLEST legal FIFO — DEPTH_L2=1, the configuration whose part-select
    // the original spelling of the full comparison could not even elaborate.
    logic s_wrst_n = 1'b0, s_rrst_n = 1'b0, s_push = 1'b0, s_pop = 1'b0;
    logic [WIDTH-1:0] s_wdata = 8'h00;
    logic [WIDTH-1:0] s_rdata;
    logic s_full, s_empty;
    uart_async_fifo #(.WIDTH(WIDTH), .DEPTH_L2(1)) u2 (
        .wclk(wclk), .wrst_n(s_wrst_n), .wdata_i(s_wdata), .wpush_i(s_push), .wfull_o(s_full),
        .rclk(rclk), .rrst_n(s_rrst_n), .rdata_o(s_rdata), .rpop_i(s_pop),  .rempty_o(s_empty));

    //---- the writer: emit a counting sequence ------------------------------
    always_ff @(posedge wclk or negedge wrst_n)
        if (!wrst_n)              wnext <= 0;
        else if (wpush && !wfull) wnext <= wnext + 1;

    //---- the reader: demand the next value of that sequence ----------------
    always @(posedge rclk or negedge rrst_n)
        if (!rrst_n) rexp <= 0;
        else if (rpop) begin
            if (rdata !== rexp[WIDTH-1:0]) mism <= mism + 1;
            a_seq: assert (rdata === rexp[WIDTH-1:0])
                else $error("popped %02h, expected %02h", rdata, rexp[WIDTH-1:0]);
            rexp <= rexp + 1;
        end

    //---- structural observer 1: outstanding words --------------------------
    logic obs_en = 1'b0;
    int   maxout = 0, minout = 0, outst = 0;
    initial begin
        #0.37;
        forever begin
            if (obs_en) begin
                outst = wnext - rexp;
                if (outst > maxout) maxout = outst;
                if (outst < minout) minout = outst;
                a_cap: assert (outst <= DEPTH && outst >= 0)
                    else $error("outstanding=%0d outside 0..%0d", outst, DEPTH);
            end
            #1;
        end
    end

    //---- structural observer 2: the gray property --------------------------
    wire [PW-1:0] wg = u8.wgray_q;      // hierarchical — observation only
    wire [PW-1:0] rg = u8.rgray_q;
    logic [PW-1:0] wg_prev, rg_prev;
    int   wg_bad = 0, rg_bad = 0, wg_moves = 0, rg_moves = 0;

    // $countones is the idiomatic SystemVerilog spelling of a population count,
    // and it is what this observer WANTS to say. It is not what it says, because
    // the Icarus Verilog 13.0 build these listings are executed on miscomputes
    // it: called repeatedly inside a procedural block it returns the vector
    // WIDTH rather than the number of set bits from the second call onward
    // ($countones(4'b0010) == 4). Published code has to run on the tool it is
    // published with, so the count is written out. On a simulator without the
    // defect, `pc(v)` may be replaced by `$countones(v)` unchanged.
    function automatic int pc(input logic [PW-1:0] v);
        pc = 0;
        for (int b = 0; b < PW; b++) pc += v[b];
    endfunction

    always @(wg) begin
        if (obs_en) begin
            wg_moves++;
            if (pc(wg ^ wg_prev) != 1) wg_bad++;
            a_wgray: assert (pc(wg ^ wg_prev) == 1)
                else $error("write gray pointer %b -> %b is not a single-bit step",
                            wg_prev, wg);
        end
        wg_prev = wg;
    end
    always @(rg) begin
        if (obs_en) begin
            rg_moves++;
            if (pc(rg ^ rg_prev) != 1) rg_bad++;
            a_rgray: assert (pc(rg ^ rg_prev) == 1)
                else $error("read gray pointer %b -> %b is not a single-bit step",
                            rg_prev, rg);
        end
        rg_prev = rg;
    end

    //---- structural observer 3: the SYNCHRONISER DEPTH ---------------------
    // This is the observer that matters most in Module 12, and the reason it
    // exists is worth stating plainly.
    //
    // Cutting the pointer synchronisers from two flops to one does not change
    // ANY functional result. The FIFO still loses nothing, duplicates nothing
    // and reorders nothing; every other check in this file still passes. The
    // property a second flop buys — that a metastable first stage is given a
    // whole clock to resolve before anything reads it — has no representation
    // in an event-driven simulator, so no amount of stimulus can expose its
    // absence.
    //
    // What IS observable is the STRUCTURE: with two flops in the chain, the
    // value of wq2_rgray after any edge must equal the value wq1_rgray held
    // BEFORE that edge. Cut the chain to one flop and the two move together.
    // So the testbench checks the shape of the logic rather than the behaviour
    // of the data — which is exactly what CDC lint does on a netlist, done
    // here with the only instrument a simulation has.
    logic [PW-1:0] wq1_dly, rq1_dly;
    int   sdepth_bad_w = 0, sdepth_bad_r = 0, sdepth_obs = 0;

    always @(posedge wclk) begin
        #1;                              // read POST-edge values
        if (obs_en && wrst_n) begin
            sdepth_obs++;
            if (u8.wq2_rgray !== wq1_dly) sdepth_bad_w++;
            a_wdepth: assert (u8.wq2_rgray === wq1_dly)
                else $error("write pointer synchroniser is not two flops deep");
        end
        wq1_dly = u8.wq1_rgray;
    end

    always @(posedge rclk) begin
        #1;
        if (obs_en && rrst_n) begin
            if (u8.rq2_wgray !== rq1_dly) sdepth_bad_r++;
            a_rdepth: assert (u8.rq2_wgray === rq1_dly)
                else $error("read pointer synchroniser is not two flops deep");
        end
        rq1_dly = u8.rq1_wgray;
    end

    task automatic obs_start;           // arm the observers from a known state
        wg_prev = wg; rg_prev = rg;
        wq1_dly = u8.wq1_rgray; rq1_dly = u8.rq1_wgray;
        obs_en = 1'b1;
    endtask

    //---- check plumbing ----------------------------------------------------
    int checks = 0, failures = 0;
    task automatic check(input logic cond, input string name);
        checks++;
        if (cond) $display("  PASS %0s", name);
        else begin failures++; $display("  FAIL %0s", name); end
    endtask

    int base_w, base_r, base_m, lat, i, guard;

    // Run one traffic phase at a given clock ratio and drain it.
    task automatic run_phase(input int wh, input int rh, input int nwords);
        whalf = wh; rhalf = rh;
        base_w = wnext; base_r = rexp;
        w_en = 1'b1; r_en = 1'b1;
        // The #1 after each edge matters twice over: it lets the loop read
        // SETTLED counters rather than their pre-NBA values, and it moves the
        // deassertion of w_en off the very edge the DUT samples wpush on.
        // Without it an extra word slips in at the boundary.
        guard = 0;
        while (wnext - base_w < nwords && guard < 200000) begin
            @(posedge wclk); #1; guard++;
        end
        w_en = 1'b0;                     // stop writing, then drain
        guard = 0;
        while (rexp != wnext && guard < 200000) begin
            @(posedge rclk); #1; guard++;
        end
        repeat (6) @(posedge rclk);
        r_en = 1'b0;
    endtask

    initial begin
        #500_000;
        $display("  FAIL watchdog: simulation did not finish");
        $display("== %0d checks, %0d failures ==", checks+1, failures+1);
        $display("   RESULT: SYSTEMVERILOG ASYNC-FIFO TESTS FAILED (timeout)");
        $finish;
    end

    initial begin
        $display("== uart_async_fifo : self-checking SystemVerilog testbench ==");

        wrst_n = 1'b0; rrst_n = 1'b0; s_wrst_n = 1'b0; s_rrst_n = 1'b0;
        repeat (8) @(posedge wclk);
        check(rempty === 1'b1 && wfull === 1'b0,
              "out of reset: empty asserted, full deasserted");
        check(s_empty === 1'b1 && s_full === 1'b0,
              "the DEPTH=2 FIFO resets the same way");
        wrst_n = 1'b1; rrst_n = 1'b1; s_wrst_n = 1'b1; s_rrst_n = 1'b1;
        repeat (8) @(posedge rclk);
        obs_start;

        //=== phase 1: writer much faster than reader ========================
        run_phase(3, 11, 40);
        check(mism == 0,             "fast-write / slow-read: sequence intact");
        check(wnext == 40,           "fast-write / slow-read: exactly 40 written");
        check(rexp == wnext,         "fast-write / slow-read: nothing left behind");
        check(rempty === 1'b1,       "fast-write / slow-read: drained to empty");
        check(wfull === 1'b0,        "and full has released");

        //=== phase 2: reader much faster than writer ========================
        base_r = rexp;
        run_phase(11, 3, 40);
        check(mism == 0,             "slow-write / fast-read: sequence intact");
        check(rexp - base_r == 40,   "slow-write / fast-read: all 40 words arrived");

        //=== phase 3: near-equal but incommensurate =========================
        run_phase(7, 5, 60);
        check(mism == 0,             "incommensurate ratio: sequence intact");
        check(rexp == 140,           "incommensurate ratio: 40+40+60 words total");
        check(rexp == wnext,         "incommensurate ratio: counts agree");

        //=== the extra pointer bit: capacity is exactly DEPTH ===============
        whalf = 5; rhalf = 7;
        base_w = wnext;
        r_en = 1'b0;                     // reader stalled: nothing can drain
        w_en = 1'b1;
        guard = 0;
        while (wfull !== 1'b1 && guard < 200) begin @(posedge wclk); #1; guard++; end
        check(wfull === 1'b1,        "reader stalled: full eventually asserts");
        repeat (4) @(posedge wclk);
        check(wnext - base_w == DEPTH,
              "exactly DEPTH words were accepted -- not DEPTH-1, not DEPTH+1");

        //=== a full FIFO refuses further writes =============================
        base_w = wnext;
        w_force = 1'b1;                  // push HARD, ignoring wfull
        repeat (20) @(posedge wclk);
        w_force = 1'b0; w_en = 1'b0;
        check(wnext - base_w == 0,   "20 forced pushes into a full FIFO: all ignored");

        //=== and the DEPTH words are still the right DEPTH words ============
        base_r = rexp;
        r_en = 1'b1;
        guard = 0;
        while (rexp != wnext && guard < 2000) begin @(posedge rclk); #1; guard++; end
        repeat (6) @(posedge rclk);
        r_en = 1'b0;
        check(rexp - base_r == DEPTH, "draining recovers exactly DEPTH words");
        check(mism == 0,             "and they are the right words, in order");
        check(rempty === 1'b1,       "empty again");

        //=== visibility latency across the boundary =========================
        // One word into an empty FIFO. The datum is in memory on the writing
        // edge; the READ side may not believe it until the gray pointer has
        // crossed two synchroniser flops. That cost is the price of safety.
        w_en = 1'b1;
        @(posedge wclk);
        @(negedge wclk) w_en = 1'b0;
        lat = 0;
        while (rempty !== 1'b0 && lat < 12) begin @(posedge rclk); #0.1; lat++; end
        check(lat >= 2,  "a new word costs at least TWO rclk edges to become visible");
        check(lat <= 5,  "and the crossing is bounded, not open-ended");
        r_en = 1'b1;
        guard = 0;
        while (rexp != wnext && guard < 200) begin @(posedge rclk); #1; guard++; end
        r_en = 1'b0;
        check(mism == 0, "the single word read back correctly");

        //=== the DEPTH=2 FIFO: capacity two, and it really is two ===========
        for (i = 0; i < 6; i++) begin
            @(negedge wclk);
            s_wdata = 8'hA0 + i[7:0];
            s_push  = !s_full;
        end
        @(negedge wclk) s_push = 1'b0;
        repeat (4) @(posedge wclk);
        check(s_full === 1'b1, "DEPTH=2 FIFO fills");
        repeat (6) @(posedge rclk);
        check(s_empty === 1'b0, "and the read side sees its contents");
        @(negedge rclk) s_pop = 1'b1;
        repeat (2) @(posedge rclk);
        @(negedge rclk) s_pop = 1'b0;
        repeat (8) @(posedge rclk);
        check(s_empty === 1'b1, "two pops empty it");

        //=== the structural properties ======================================
        check(wg_moves > 50 && rg_moves > 50,
              "both gray pointers moved enough to be worth judging");
        check(wg_bad == 0,  "every write-pointer transition changed exactly ONE bit");
        check(rg_bad == 0,  "every read-pointer transition changed exactly ONE bit");
        check(maxout <= DEPTH,
              "outstanding words never exceeded DEPTH -- no silent overwrite");
        check(minout >= 0,
              "outstanding words never went negative -- nothing read before written");

        check(sdepth_obs > 300,
              "the synchroniser-depth observer ran on enough edges to judge");
        check(sdepth_bad_w == 0,
              "write-side pointer synchroniser is TWO flops deep, not one");
        check(sdepth_bad_r == 0,
              "read-side pointer synchroniser is TWO flops deep, not one");

        //=== reset during traffic ===========================================
        obs_en = 1'b0;
        w_en = 1'b1; r_en = 1'b1;
        repeat (20) @(posedge wclk);
        wrst_n = 1'b0; rrst_n = 1'b0;    // both domains reset together
        w_en = 1'b0; r_en = 1'b0;
        repeat (6) @(posedge wclk);
        check(rempty === 1'b1 && wfull === 1'b0,
              "reset during traffic returns the FIFO to empty");
        wrst_n = 1'b1; rrst_n = 1'b1;
        repeat (8) @(posedge rclk);
        obs_start;
        // A baseline rather than a reset of the counter: this asks "were there
        // NEW mismatches after the reset", which is the actual question,
        // instead of destroying the earlier evidence. The VHDL twin has no
        // choice -- mism is driven by the reader process there, and a second
        // driver on an unresolved type is illegal -- and it is better practice
        // here too.
        base_m = mism;
        run_phase(5, 7, 24);
        check(mism == base_m, "traffic after reset is clean from the first word");
        check(rexp == wnext, "and balanced");
        check(wg_bad == 0 && rg_bad == 0,
              "the gray property survived the reset too");

        $display("== %0d checks, %0d failures ==", checks, failures);
        if (failures == 0) $display("   RESULT: ALL SYSTEMVERILOG ASYNC-FIFO TESTS PASSED");
        else               $display("   RESULT: SYSTEMVERILOG ASYNC-FIFO TESTS FAILED");
        $finish;
    end
endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
//===========================================================================
//  tb_uart_async_fifo_v — self-checking Verilog-2001 testbench
//
//  The two clocks here are UNRELATED. Their periods are set at run time and
//  deliberately chosen to be incommensurate, so the phase relationship keeps
//  drifting and the design is exercised at every alignment rather than at one.
//
//  THE ORACLE IS INDEPENDENT. The writer emits a counting sequence; the
//  reader keeps its OWN counter and demands that every popped word be the
//  next value of that sequence. No shared model, no peeking at pointers:
//     a lost word   -> the reader's next compare mismatches
//     a duplicated  -> the reader's next compare mismatches
//     a reordered   -> the reader's next compare mismatches
//  and at the end of each phase, drained, the two counters must be equal.
//
//  Two properties are checked STRUCTURALLY, through hierarchical references,
//  because they are the reason the design is shaped the way it is:
//     - every gray pointer transition changes exactly ONE bit (decision 2)
//     - outstanding words never exceed DEPTH        (decision 1, the extra bit)
//===========================================================================
`timescale 1ns/1ps

module tb_uart_async_fifo_v;

    localparam WIDTH = 8;
    localparam DL2   = 3;
    localparam DEPTH = 1 << DL2;        // 8
    localparam PW    = DL2 + 1;         // 4

    //---- two independent clocks, periods settable at run time --------------
    integer whalf = 5, rhalf = 7;
    reg wclk = 1'b0, rclk = 1'b0;
    always #(whalf) wclk = ~wclk;
    always #(rhalf) rclk = ~rclk;

    reg  wrst_n = 1'b0, rrst_n = 1'b0;
    reg  w_en = 1'b0, r_en = 1'b0, w_force = 1'b0;

    wire [WIDTH-1:0] wdata, rdata;
    wire             wfull, rempty;
    wire             wpush = w_force ? 1'b1 : (w_en && !wfull);
    wire             rpop  = r_en && !rempty;

    integer wnext = 0, rexp = 0, mism = 0;
    assign wdata = wnext[WIDTH-1:0];

    uart_async_fifo_v #(.WIDTH(WIDTH), .DEPTH_L2(DL2)) u8 (
        .wclk(wclk), .wrst_n(wrst_n), .wdata_i(wdata), .wpush_i(wpush), .wfull_o(wfull),
        .rclk(rclk), .rrst_n(rrst_n), .rdata_o(rdata), .rpop_i(rpop),  .rempty_o(rempty));

    // the SMALLEST legal FIFO -- DEPTH_L2=1, the configuration whose part-select
    // the original spelling of the full comparison could not even elaborate.
    reg  s_wrst_n = 1'b0, s_rrst_n = 1'b0, s_push = 1'b0, s_pop = 1'b0;
    reg  [WIDTH-1:0] s_wdata = 8'h00;
    wire [WIDTH-1:0] s_rdata;
    wire s_full, s_empty;
    uart_async_fifo_v #(.WIDTH(WIDTH), .DEPTH_L2(1)) u2 (
        .wclk(wclk), .wrst_n(s_wrst_n), .wdata_i(s_wdata), .wpush_i(s_push), .wfull_o(s_full),
        .rclk(rclk), .rrst_n(s_rrst_n), .rdata_o(s_rdata), .rpop_i(s_pop),  .rempty_o(s_empty));

    //---- the writer: emit a counting sequence ------------------------------
    always @(posedge wclk or negedge wrst_n)
        if (!wrst_n)            wnext <= 0;
        else if (wpush && !wfull) wnext <= wnext + 1;

    //---- the reader: demand the next value of that sequence ----------------
    always @(posedge rclk or negedge rrst_n)
        if (!rrst_n) rexp <= 0;
        else if (rpop) begin
            if (rdata !== rexp[WIDTH-1:0]) mism <= mism + 1;
            rexp <= rexp + 1;
        end

    //---- structural observer 1: outstanding words --------------------------
    // Sampled off the integer-ns stimulus grid so that every read is of
    // settled values (see the reset-sync testbench for why this matters).
    reg obs_en = 1'b0;
    integer maxout = 0, minout = 0, outst = 0;
    initial begin
        #0.37;
        forever begin
            if (obs_en) begin
                outst = wnext - rexp;
                if (outst > maxout) maxout = outst;
                if (outst < minout) minout = outst;
            end
            #1;
        end
    end

    //---- structural observer 2: the gray property --------------------------
    wire [PW-1:0] wg = u8.wgray_q;      // hierarchical -- observation only
    wire [PW-1:0] rg = u8.rgray_q;
    reg  [PW-1:0] wg_prev, rg_prev;
    integer wg_bad = 0, rg_bad = 0, wg_moves = 0, rg_moves = 0;

    function [2:0] pc;                  // population count, 4 bits
        input [PW-1:0] v;
        integer i;
        begin pc = 0; for (i = 0; i < PW; i = i + 1) pc = pc + v[i]; end
    endfunction

    always @(wg) begin
        if (obs_en) begin
            wg_moves = wg_moves + 1;
            if (pc(wg ^ wg_prev) !== 3'd1) wg_bad = wg_bad + 1;
        end
        wg_prev = wg;
    end
    always @(rg) begin
        if (obs_en) begin
            rg_moves = rg_moves + 1;
            if (pc(rg ^ rg_prev) !== 3'd1) rg_bad = rg_bad + 1;
        end
        rg_prev = rg;
    end

    //---- structural observer 3: the SYNCHRONISER DEPTH ---------------------
    // This is the observer that matters most in Module 12, and the reason it
    // exists is worth stating plainly.
    //
    // Cutting the pointer synchronisers from two flops to one does not change
    // ANY functional result. The FIFO still loses nothing, duplicates nothing
    // and reorders nothing; every other check in this file still passes. The
    // property a second flop buys -- that a metastable first stage is given a
    // whole clock to resolve before anything reads it -- has no representation
    // in an event-driven simulator, so no amount of stimulus can expose its
    // absence.
    //
    // What IS observable is the STRUCTURE: with two flops in the chain, the
    // value of wq2_rgray after any edge must equal the value wq1_rgray held
    // BEFORE that edge. Cut the chain to one flop and the two move together.
    // So the testbench checks the shape of the logic rather than the behaviour
    // of the data -- which is exactly what CDC lint does on a netlist, done
    // here with the only instrument a simulation has.
    reg [PW-1:0] wq1_dly, rq1_dly;
    integer sdepth_bad_w = 0, sdepth_bad_r = 0, sdepth_obs = 0;

    always @(posedge wclk) begin
        #1;                              // read POST-edge values
        if (obs_en && wrst_n) begin
            sdepth_obs = sdepth_obs + 1;
            if (u8.wq2_rgray !== wq1_dly) sdepth_bad_w = sdepth_bad_w + 1;
        end
        wq1_dly = u8.wq1_rgray;
    end

    always @(posedge rclk) begin
        #1;
        if (obs_en && rrst_n) begin
            if (u8.rq2_wgray !== rq1_dly) sdepth_bad_r = sdepth_bad_r + 1;
        end
        rq1_dly = u8.rq1_wgray;
    end

    task obs_start;                     // arm the observers from a known state
        begin wg_prev = wg; rg_prev = rg;
              wq1_dly = u8.wq1_rgray; rq1_dly = u8.rq1_wgray;
              obs_en = 1'b1; end
    endtask


    //---- check plumbing ----------------------------------------------------
    integer checks = 0, failures = 0;
    task check;
        input cond;
        input [8*80-1:0] name;
        begin
            checks = checks + 1;
            if (cond) $display("  PASS %0s", name);
            else begin failures = failures + 1; $display("  FAIL %0s", name); end
        end
    endtask

    integer base_w, base_r, base_m, lat, i, guard;

    // Run one traffic phase at a given clock ratio and drain it.
    task run_phase;
        input integer wh;
        input integer rh;
        input integer nwords;
        begin
            whalf = wh; rhalf = rh;
            base_w = wnext; base_r = rexp;
            w_en = 1'b1; r_en = 1'b1;
            // The #1 after each edge matters twice over: it lets the loop read
            // SETTLED counters rather than their pre-NBA values, and it moves the
            // deassertion of w_en off the very edge the DUT samples wpush on.
            // Without it an extra word slips in at the boundary and the phase
            // writes nwords+1.
            guard = 0;
            while (wnext - base_w < nwords && guard < 200000) begin
                @(posedge wclk); #1; guard = guard + 1;
            end
            w_en = 1'b0;                 // stop writing, then drain
            guard = 0;
            while (rexp != wnext && guard < 200000) begin
                @(posedge rclk); #1; guard = guard + 1;
            end
            repeat (6) @(posedge rclk);
            r_en = 1'b0;
        end
    endtask

    initial begin
        #500_000;
        $display("  FAIL watchdog: simulation did not finish");
        $display("== %0d checks, %0d failures ==", checks+1, failures+1);
        $display("   RESULT: VERILOG ASYNC-FIFO TESTS FAILED (timeout)");
        $finish;
    end

    initial begin
        $display("== uart_async_fifo_v : self-checking Verilog testbench ==");

        wrst_n = 1'b0; rrst_n = 1'b0; s_wrst_n = 1'b0; s_rrst_n = 1'b0;
        repeat (8) @(posedge wclk);
        check(rempty === 1'b1 && wfull === 1'b0,
              "out of reset: empty asserted, full deasserted");
        check(s_empty === 1'b1 && s_full === 1'b0,
              "the DEPTH=2 FIFO resets the same way");
        wrst_n = 1'b1; rrst_n = 1'b1; s_wrst_n = 1'b1; s_rrst_n = 1'b1;
        repeat (8) @(posedge rclk);
        obs_start;

        //=== phase 1: writer much faster than reader ========================
        run_phase(3, 11, 40);
        check(mism == 0,             "fast-write / slow-read: sequence intact");
        check(wnext == 40,           "fast-write / slow-read: exactly 40 written");
        check(rexp == wnext,         "fast-write / slow-read: nothing left behind");
        check(rempty === 1'b1,       "fast-write / slow-read: drained to empty");
        check(wfull === 1'b0,        "and full has released");

        //=== phase 2: reader much faster than writer ========================
        base_r = rexp;
        run_phase(11, 3, 40);
        check(mism == 0,             "slow-write / fast-read: sequence intact");
        check(rexp - base_r == 40,   "slow-write / fast-read: all 40 words arrived");

        //=== phase 3: near-equal but incommensurate =========================
        run_phase(7, 5, 60);
        check(mism == 0,             "incommensurate ratio: sequence intact");
        check(rexp == 140,           "incommensurate ratio: 40+40+60 words total");
        check(rexp == wnext,         "incommensurate ratio: counts agree");

        //=== the extra pointer bit: capacity is exactly DEPTH ===============
        whalf = 5; rhalf = 7;
        base_w = wnext;
        r_en = 1'b0;                     // reader stalled: nothing can drain
        w_en = 1'b1;
        guard = 0;
        while (wfull !== 1'b1 && guard < 200) begin @(posedge wclk); guard = guard + 1; end
        check(wfull === 1'b1,        "reader stalled: full eventually asserts");
        repeat (4) @(posedge wclk);
        check(wnext - base_w == DEPTH,
              "exactly DEPTH words were accepted -- not DEPTH-1, not DEPTH+1");

        //=== a full FIFO refuses further writes =============================
        base_w = wnext;
        w_force = 1'b1;                  // push HARD, ignoring wfull
        repeat (20) @(posedge wclk);
        w_force = 1'b0; w_en = 1'b0;
        check(wnext - base_w == 0,   "20 forced pushes into a full FIFO: all ignored");

        //=== and the DEPTH words are still the right DEPTH words ============
        base_r = rexp;
        r_en = 1'b1;
        guard = 0;
        while (rexp != wnext && guard < 2000) begin @(posedge rclk); guard = guard + 1; end
        repeat (6) @(posedge rclk);
        r_en = 1'b0;
        check(rexp - base_r == DEPTH, "draining recovers exactly DEPTH words");
        check(mism == 0,             "and they are the right words, in order");
        check(rempty === 1'b1,       "empty again");

        //=== visibility latency across the boundary =========================
        // One word into an empty FIFO. The datum is in memory on the writing
        // edge; the READ side may not believe it until the gray pointer has
        // crossed two synchroniser flops. That cost is the price of safety.
        w_en = 1'b1;
        @(posedge wclk);
        @(negedge wclk) w_en = 1'b0;
        lat = 0;
        while (rempty !== 1'b0 && lat < 12) begin @(posedge rclk); #0.1; lat = lat + 1; end
        check(lat >= 2,  "a new word costs at least TWO rclk edges to become visible");
        check(lat <= 5,  "and the crossing is bounded, not open-ended");
        r_en = 1'b1;
        guard = 0;
        while (rexp != wnext && guard < 200) begin @(posedge rclk); guard = guard + 1; end
        r_en = 1'b0;
        check(mism == 0, "the single word read back correctly");

        //=== the DEPTH=2 FIFO: capacity two, and it really is two ===========
        for (i = 0; i < 6; i = i + 1) begin
            @(negedge wclk);
            s_wdata = 8'hA0 + i[7:0];
            s_push  = !s_full;
        end
        @(negedge wclk) s_push = 1'b0;
        repeat (4) @(posedge wclk);
        check(s_full === 1'b1, "DEPTH=2 FIFO fills");
        repeat (6) @(posedge rclk);
        check(s_empty === 1'b0, "and the read side sees its contents");
        @(negedge rclk) s_pop = 1'b1;
        repeat (2) @(posedge rclk);
        @(negedge rclk) s_pop = 1'b0;
        repeat (8) @(posedge rclk);
        check(s_empty === 1'b1, "two pops empty it");

        //=== the structural properties ======================================
        check(wg_moves > 50 && rg_moves > 50,
              "both gray pointers moved enough to be worth judging");
        check(wg_bad == 0,  "every write-pointer transition changed exactly ONE bit");
        check(rg_bad == 0,  "every read-pointer transition changed exactly ONE bit");
        check(maxout <= DEPTH,
              "outstanding words never exceeded DEPTH -- no silent overwrite");
        check(minout >= 0,
              "outstanding words never went negative -- nothing read before written");

        check(sdepth_obs > 300,
              "the synchroniser-depth observer ran on enough edges to judge");
        check(sdepth_bad_w == 0,
              "write-side pointer synchroniser is TWO flops deep, not one");
        check(sdepth_bad_r == 0,
              "read-side pointer synchroniser is TWO flops deep, not one");

        //=== reset during traffic ===========================================
        obs_en = 1'b0;
        w_en = 1'b1; r_en = 1'b1;
        repeat (20) @(posedge wclk);
        wrst_n = 1'b0; rrst_n = 1'b0;    // both domains reset together
        w_en = 1'b0; r_en = 1'b0;
        repeat (6) @(posedge wclk);
        check(rempty === 1'b1 && wfull === 1'b0,
              "reset during traffic returns the FIFO to empty");
        wrst_n = 1'b1; rrst_n = 1'b1;
        repeat (8) @(posedge rclk);
        obs_start;
        // A baseline rather than a reset of the counter: this asks "were there
        // NEW mismatches after the reset", which is the actual question,
        // instead of destroying the earlier evidence. The VHDL twin has no
        // choice -- mism is driven by the reader process there, and a second
        // driver on an unresolved type is illegal -- and it is better practice
        // here too.
        base_m = mism;
        run_phase(5, 7, 24);
        check(mism == base_m, "traffic after reset is clean from the first word");
        check(rexp == wnext, "and balanced");
        check(wg_bad == 0 && rg_bad == 0,
              "the gray property survived the reset too");

        $display("== %0d checks, %0d failures ==", checks, failures);
        if (failures == 0) $display("   RESULT: ALL VERILOG ASYNC-FIFO TESTS PASSED");
        else               $display("   RESULT: VERILOG ASYNC-FIFO TESTS FAILED");
        $finish;
    end
endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
--===========================================================================
--  tb_uart_async_fifo — self-checking VHDL-2008 testbench
--
--  The two clocks here are UNRELATED. Their periods are held in signals of
--  type TIME and changed at run time, deliberately to incommensurate values,
--  so the phase relationship keeps drifting and the design is exercised at
--  every alignment rather than at one.
--
--  THE ORACLE IS INDEPENDENT. The writer emits a counting sequence; the
--  reader keeps its OWN counter and demands that every popped word be the
--  next value of that sequence. No shared model, no peeking at pointers:
--     a lost word   → the reader's next compare mismatches
--     a duplicated  → the reader's next compare mismatches
--     a reordered   → the reader's next compare mismatches
--  and at the end of each phase, drained, the two counters must be equal.
--
--  Two properties are checked STRUCTURALLY, through VHDL-2008 EXTERNAL NAMES,
--  because they are the reason the design is shaped the way it is:
--     - every gray pointer transition changes exactly ONE bit (decision 2)
--     - outstanding words never exceed DEPTH        (decision 1, the extra bit)
--  External names are the VHDL equivalent of Verilog's hierarchical reference,
--  and like it they are for OBSERVATION only — nothing here drives the DUT's
--  internals or derives an expectation from them.
--
--  Same 36 counted checks as the Verilog and SystemVerilog twins.
--===========================================================================
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

entity tb_uart_async_fifo is
end entity tb_uart_async_fifo;

architecture sim of tb_uart_async_fifo is

    constant WIDTH : positive := 8;
    constant DL2   : positive := 3;
    constant DEPTH : positive := 2 ** DL2;      -- 8
    constant PW    : positive := DL2 + 1;       -- 4

    signal whalf : time := 5 ns;
    signal rhalf : time := 7 ns;
    signal wclk  : std_logic := '0';
    signal rclk  : std_logic := '0';
    signal done  : boolean   := false;

    signal wrst_n, rrst_n : std_logic := '0';
    signal w_en, r_en, w_force : std_logic := '0';

    signal wdata, rdata : std_logic_vector(WIDTH-1 downto 0);
    signal wfull, rempty : std_logic;
    signal wpush, rpop   : std_logic;

    signal wnext, rexp : natural := 0;
    signal mism        : natural := 0;

    -- the SMALLEST legal FIFO — DEPTH_L2=1, the configuration whose part-select
    -- the original spelling of the full comparison could not even elaborate.
    signal s_wrst_n, s_rrst_n, s_push, s_pop : std_logic := '0';
    signal s_wdata, s_rdata : std_logic_vector(WIDTH-1 downto 0) := (others => '0');
    signal s_full, s_empty  : std_logic;

    signal obs_en : boolean := false;
    signal maxout, minout : integer := 0;
    signal wg_bad, rg_bad, wg_moves, rg_moves : natural := 0;
    signal sdepth_bad_w, sdepth_bad_r, sdepth_obs : natural := 0;

begin

    ---- two independent clocks, periods settable at run time ------------------
    wgen : process
    begin
        while not done loop
            wait for whalf;
            wclk <= not wclk;
        end loop;
        wait;
    end process wgen;

    rgen : process
    begin
        while not done loop
            wait for rhalf;
            rclk <= not rclk;
        end loop;
        wait;
    end process rgen;

    wpush <= '1' when w_force = '1' else (w_en and not wfull);
    rpop  <= r_en and not rempty;
    wdata <= std_logic_vector(to_unsigned(wnext mod 2**WIDTH, WIDTH));

    u8 : entity work.uart_async_fifo
        generic map (WIDTH => WIDTH, DEPTH_L2 => DL2)
        port map (wclk => wclk, wrst_n => wrst_n, wdata_i => wdata,
                  wpush_i => wpush, wfull_o => wfull,
                  rclk => rclk, rrst_n => rrst_n, rdata_o => rdata,
                  rpop_i => rpop, rempty_o => rempty);

    u2 : entity work.uart_async_fifo
        generic map (WIDTH => WIDTH, DEPTH_L2 => 1)
        port map (wclk => wclk, wrst_n => s_wrst_n, wdata_i => s_wdata,
                  wpush_i => s_push, wfull_o => s_full,
                  rclk => rclk, rrst_n => s_rrst_n, rdata_o => s_rdata,
                  rpop_i => s_pop, rempty_o => s_empty);

    ---- the writer: emit a counting sequence ----------------------------------
    wproc : process (wclk, wrst_n)
    begin
        if wrst_n = '0' then
            wnext <= 0;
        elsif rising_edge(wclk) then
            if wpush = '1' and wfull = '0' then
                wnext <= wnext + 1;
            end if;
        end if;
    end process wproc;

    ---- the reader: demand the next value of that sequence --------------------
    rproc : process (rclk, rrst_n)
    begin
        if rrst_n = '0' then
            rexp <= 0;
        elsif rising_edge(rclk) then
            if rpop = '1' then
                if rdata /= std_logic_vector(to_unsigned(rexp mod 2**WIDTH, WIDTH)) then
                    mism <= mism + 1;
                    report "popped word does not match the expected sequence"
                        severity error;
                end if;
                rexp <= rexp + 1;
            end if;
        end if;
    end process rproc;

    ---- structural observer 1: outstanding words ------------------------------
    -- Sampled off the whole-nanosecond stimulus grid so that every read is of
    -- settled values, matching the Verilog and SystemVerilog twins.
    obs_cap : process
        variable outst : integer;
    begin
        wait for 0.37 ns;
        loop
            if obs_en then
                outst := wnext - rexp;
                if outst > maxout then maxout <= outst; end if;
                if outst < minout then minout <= outst; end if;
                assert outst <= DEPTH and outst >= 0
                    report "outstanding=" & integer'image(outst)
                         & " outside 0.." & integer'image(DEPTH)
                    severity error;
            end if;
            exit when done;
            wait for 1 ns;
        end loop;
        wait;
    end process obs_cap;

    ---- structural observer 2: the gray property ------------------------------
    -- VHDL-2008 external names. Observation only.
    obs_gray : process
        alias wg is << signal .tb_uart_async_fifo.u8.wgray_q : unsigned(PW-1 downto 0) >>;
        alias rg is << signal .tb_uart_async_fifo.u8.rgray_q : unsigned(PW-1 downto 0) >>;
        variable wg_prev, rg_prev : unsigned(PW-1 downto 0) := (others => '0');

        -- VHDL-2008 has a unary XOR reduction but no population count, so the
        -- count is written out. It is the same count the Verilog twin computes
        -- in a function and the SystemVerilog twin would spell $countones.
        function pc(v : unsigned) return natural is
            variable n : natural := 0;
        begin
            for b in v'range loop
                if v(b) = '1' then n := n + 1; end if;
            end loop;
            return n;
        end function pc;
    begin
        wait on wg, rg;
        loop
            if wg /= wg_prev then
                if obs_en then
                    wg_moves <= wg_moves + 1;
                    if pc(wg xor wg_prev) /= 1 then
                        wg_bad <= wg_bad + 1;
                        report "write gray pointer step is not a single bit"
                            severity error;
                    end if;
                end if;
                wg_prev := wg;
            end if;
            if rg /= rg_prev then
                if obs_en then
                    rg_moves <= rg_moves + 1;
                    if pc(rg xor rg_prev) /= 1 then
                        rg_bad <= rg_bad + 1;
                        report "read gray pointer step is not a single bit"
                            severity error;
                    end if;
                end if;
                rg_prev := rg;
            end if;
            exit when done;
            wait on wg, rg, done;
        end loop;
        wait;
    end process obs_gray;

    ---- structural observer 3: the SYNCHRONISER DEPTH -------------------------
    -- This is the observer that matters most in Module 12, and the reason it
    -- exists is worth stating plainly.
    --
    -- Cutting the pointer synchronisers from two flops to one does not change
    -- ANY functional result. The FIFO still loses nothing, duplicates nothing
    -- and reorders nothing; every other check in this file still passes. The
    -- property a second flop buys — that a metastable first stage is given a
    -- whole clock to resolve before anything reads it — has no representation
    -- in an event-driven simulator, so no amount of stimulus can expose its
    -- absence.
    --
    -- What IS observable is the STRUCTURE: with two flops in the chain, the
    -- value of wq2_rgray after any edge must equal the value wq1_rgray held
    -- BEFORE that edge. Cut the chain to one flop and the two move together.
    -- So the testbench checks the shape of the logic rather than the behaviour
    -- of the data — which is exactly what CDC lint does on a netlist, done
    -- here with the only instrument a simulation has.
    obs_depth_w : process
        alias wq1 is << signal .tb_uart_async_fifo.u8.wq1_rgray : unsigned(PW-1 downto 0) >>;
        alias wq2 is << signal .tb_uart_async_fifo.u8.wq2_rgray : unsigned(PW-1 downto 0) >>;
        variable wq1_dly : unsigned(PW-1 downto 0) := (others => '0');
    begin
        loop
            wait until rising_edge(wclk);
            wait for 1 ns;                       -- read POST-edge values
            if obs_en and wrst_n = '1' then
                sdepth_obs <= sdepth_obs + 1;
                if wq2 /= wq1_dly then
                    sdepth_bad_w <= sdepth_bad_w + 1;
                    report "write pointer synchroniser is not two flops deep"
                        severity error;
                end if;
            end if;
            wq1_dly := wq1;
            exit when done;
        end loop;
        wait;
    end process obs_depth_w;

    obs_depth_r : process
        alias rq1 is << signal .tb_uart_async_fifo.u8.rq1_wgray : unsigned(PW-1 downto 0) >>;
        alias rq2 is << signal .tb_uart_async_fifo.u8.rq2_wgray : unsigned(PW-1 downto 0) >>;
        variable rq1_dly : unsigned(PW-1 downto 0) := (others => '0');
    begin
        loop
            wait until rising_edge(rclk);
            wait for 1 ns;
            if obs_en and rrst_n = '1' then
                if rq2 /= rq1_dly then
                    sdepth_bad_r <= sdepth_bad_r + 1;
                    report "read pointer synchroniser is not two flops deep"
                        severity error;
                end if;
            end if;
            rq1_dly := rq1;
            exit when done;
        end loop;
        wait;
    end process obs_depth_r;

    watchdog : process
    begin
        wait for 500 us;
        report "watchdog: simulation did not finish" severity failure;
    end process watchdog;

    stim : process
        variable checks, failures : natural := 0;
        variable lat, guard       : natural := 0;
        variable base_w, base_r   : natural := 0;
        variable base_m           : natural := 0;

        procedure check(cond : boolean; name : string) is
        begin
            checks := checks + 1;
            if cond then
                report "  PASS " & name severity note;
            else
                failures := failures + 1;
                report "  FAIL " & name severity error;
            end if;
        end procedure check;

        -- Run one traffic phase at a given clock ratio and drain it.
        -- The 1 ns after each edge matters twice over: it lets the loop read
        -- SETTLED counters, and it moves the deassertion of w_en off the very
        -- edge the DUT samples wpush on. Without it an extra word slips in at
        -- the boundary and the phase writes nwords+1.
        procedure run_phase(wh : time; rh : time; nwords : natural) is
            variable bw, g : natural := 0;
        begin
            whalf <= wh; rhalf <= rh;
            bw := wnext;
            w_en <= '1'; r_en <= '1';
            g := 0;
            while (wnext - bw) < nwords and g < 200000 loop
                wait until rising_edge(wclk); wait for 1 ns; g := g + 1;
            end loop;
            w_en <= '0';
            g := 0;
            while rexp /= wnext and g < 200000 loop
                wait until rising_edge(rclk); wait for 1 ns; g := g + 1;
            end loop;
            for i in 1 to 6 loop wait until rising_edge(rclk); end loop;
            r_en <= '0';
        end procedure run_phase;
    begin
        report "== uart_async_fifo : self-checking VHDL testbench ==" severity note;

        wrst_n <= '0'; rrst_n <= '0'; s_wrst_n <= '0'; s_rrst_n <= '0';
        for i in 1 to 8 loop wait until rising_edge(wclk); end loop;
        check(rempty = '1' and wfull = '0',
              "out of reset: empty asserted, full deasserted");
        check(s_empty = '1' and s_full = '0',
              "the DEPTH=2 FIFO resets the same way");
        wrst_n <= '1'; rrst_n <= '1'; s_wrst_n <= '1'; s_rrst_n <= '1';
        for i in 1 to 8 loop wait until rising_edge(rclk); end loop;
        obs_en <= true;
        wait for 1 ns;

        --=== phase 1: writer much faster than reader ========================
        run_phase(3 ns, 11 ns, 40);
        check(mism = 0,       "fast-write / slow-read: sequence intact");
        check(wnext = 40,     "fast-write / slow-read: exactly 40 written");
        check(rexp = wnext,   "fast-write / slow-read: nothing left behind");
        check(rempty = '1',   "fast-write / slow-read: drained to empty");
        check(wfull = '0',    "and full has released");

        --=== phase 2: reader much faster than writer ========================
        base_r := rexp;
        run_phase(11 ns, 3 ns, 40);
        check(mism = 0,              "slow-write / fast-read: sequence intact");
        check(rexp - base_r = 40,    "slow-write / fast-read: all 40 words arrived");

        --=== phase 3: near-equal but incommensurate =========================
        run_phase(7 ns, 5 ns, 60);
        check(mism = 0,       "incommensurate ratio: sequence intact");
        check(rexp = 140,     "incommensurate ratio: 40+40+60 words total");
        check(rexp = wnext,   "incommensurate ratio: counts agree");

        --=== the extra pointer bit: capacity is exactly DEPTH ===============
        whalf <= 5 ns; rhalf <= 7 ns;
        base_w := wnext;
        r_en <= '0';                     -- reader stalled: nothing can drain
        w_en <= '1';
        guard := 0;
        while wfull /= '1' and guard < 200 loop
            wait until rising_edge(wclk); wait for 1 ns; guard := guard + 1;
        end loop;
        check(wfull = '1',    "reader stalled: full eventually asserts");
        for i in 1 to 4 loop wait until rising_edge(wclk); end loop;
        check(wnext - base_w = DEPTH,
              "exactly DEPTH words were accepted -- not DEPTH-1, not DEPTH+1");

        --=== a full FIFO refuses further writes =============================
        base_w := wnext;
        w_force <= '1';                  -- push HARD, ignoring wfull
        for i in 1 to 20 loop wait until rising_edge(wclk); end loop;
        w_force <= '0'; w_en <= '0';
        wait for 1 ns;
        check(wnext - base_w = 0, "20 forced pushes into a full FIFO: all ignored");

        --=== and the DEPTH words are still the right DEPTH words ============
        base_r := rexp;
        r_en <= '1';
        guard := 0;
        while rexp /= wnext and guard < 2000 loop
            wait until rising_edge(rclk); wait for 1 ns; guard := guard + 1;
        end loop;
        for i in 1 to 6 loop wait until rising_edge(rclk); end loop;
        r_en <= '0';
        check(rexp - base_r = DEPTH, "draining recovers exactly DEPTH words");
        check(mism = 0,              "and they are the right words, in order");
        check(rempty = '1',          "empty again");

        --=== visibility latency across the boundary =========================
        -- One word into an empty FIFO. The datum is in memory on the writing
        -- edge; the READ side may not believe it until the gray pointer has
        -- crossed two synchroniser flops. That cost is the price of safety.
        w_en <= '1';
        wait until rising_edge(wclk);
        wait until falling_edge(wclk); w_en <= '0';
        lat := 0;
        while rempty /= '0' and lat < 12 loop
            wait until rising_edge(rclk); wait for 0.1 ns; lat := lat + 1;
        end loop;
        check(lat >= 2, "a new word costs at least TWO rclk edges to become visible");
        check(lat <= 5, "and the crossing is bounded, not open-ended");
        r_en <= '1';
        guard := 0;
        while rexp /= wnext and guard < 200 loop
            wait until rising_edge(rclk); wait for 1 ns; guard := guard + 1;
        end loop;
        r_en <= '0';
        check(mism = 0, "the single word read back correctly");

        --=== the DEPTH=2 FIFO: capacity two, and it really is two ===========
        for i in 0 to 5 loop
            wait until falling_edge(wclk);
            s_wdata <= std_logic_vector(to_unsigned(16#A0# + i, WIDTH));
            if s_full = '0' then s_push <= '1'; else s_push <= '0'; end if;
        end loop;
        wait until falling_edge(wclk); s_push <= '0';
        for i in 1 to 4 loop wait until rising_edge(wclk); end loop;
        check(s_full = '1', "DEPTH=2 FIFO fills");
        for i in 1 to 6 loop wait until rising_edge(rclk); end loop;
        check(s_empty = '0', "and the read side sees its contents");
        wait until falling_edge(rclk); s_pop <= '1';
        for i in 1 to 2 loop wait until rising_edge(rclk); end loop;
        wait until falling_edge(rclk); s_pop <= '0';
        for i in 1 to 8 loop wait until rising_edge(rclk); end loop;
        check(s_empty = '1', "two pops empty it");

        --=== the structural properties ======================================
        check(wg_moves > 50 and rg_moves > 50,
              "both gray pointers moved enough to be worth judging");
        check(wg_bad = 0, "every write-pointer transition changed exactly ONE bit");
        check(rg_bad = 0, "every read-pointer transition changed exactly ONE bit");
        check(maxout <= DEPTH,
              "outstanding words never exceeded DEPTH -- no silent overwrite");
        check(minout >= 0,
              "outstanding words never went negative -- nothing read before written");

        check(sdepth_obs > 300,
              "the synchroniser-depth observer ran on enough edges to judge");
        check(sdepth_bad_w = 0,
              "write-side pointer synchroniser is TWO flops deep, not one");
        check(sdepth_bad_r = 0,
              "read-side pointer synchroniser is TWO flops deep, not one");

        --=== reset during traffic ===========================================
        obs_en <= false;
        wait for 1 ns;
        w_en <= '1'; r_en <= '1';
        for i in 1 to 20 loop wait until rising_edge(wclk); end loop;
        wrst_n <= '0'; rrst_n <= '0';    -- both domains reset together
        w_en <= '0'; r_en <= '0';
        for i in 1 to 6 loop wait until rising_edge(wclk); end loop;
        check(rempty = '1' and wfull = '0',
              "reset during traffic returns the FIFO to empty");
        wrst_n <= '1'; rrst_n <= '1';
        for i in 1 to 8 loop wait until rising_edge(rclk); end loop;
        obs_en <= true;
        -- mism cannot simply be zeroed here: it is driven by rproc, and a
        -- second driver on an unresolved type (natural) is illegal in VHDL.
        -- A baseline is better practice in every language anyway — it asks
        -- "were there NEW mismatches after the reset", which is the actual
        -- question, instead of destroying the earlier evidence.
        base_m := mism;
        wait for 1 ns;
        run_phase(5 ns, 7 ns, 24);
        check(mism = base_m,  "traffic after reset is clean from the first word");
        check(rexp = wnext,   "and balanced");
        check(wg_bad = 0 and rg_bad = 0,
              "the gray property survived the reset too");

        report "== " & integer'image(checks) & " checks, "
                     & integer'image(failures) & " failures ==" severity note;
        if failures = 0 then
            report "   RESULT: ALL VHDL ASYNC-FIFO TESTS PASSED" severity note;
        else
            report "   RESULT: VHDL ASYNC-FIFO TESTS FAILED" severity error;
        end if;
        done <= true;
        wait;
    end process stim;

end architecture sim;

Thirty-six checks, and all three languages agree:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
  PASS out of reset: empty asserted, full deasserted
  PASS the DEPTH=2 FIFO resets the same way
  PASS fast-write / slow-read: sequence intact
  PASS slow-write / fast-read: all 40 words arrived
  PASS incommensurate ratio: 40+40+60 words total
  PASS reader stalled: full eventually asserts
  PASS exactly DEPTH words were accepted -- not DEPTH-1, not DEPTH+1
  PASS 20 forced pushes into a full FIFO: all ignored
  PASS draining recovers exactly DEPTH words
  PASS a new word costs at least TWO rclk edges to become visible
  PASS and the crossing is bounded, not open-ended
  PASS DEPTH=2 FIFO fills
  PASS every write-pointer transition changed exactly ONE bit
  PASS every read-pointer transition changed exactly ONE bit
  PASS outstanding words never exceeded DEPTH -- no silent overwrite
  PASS write-side pointer synchroniser is TWO flops deep, not one
  PASS read-side pointer synchroniser is TWO flops deep, not one
  PASS reset during traffic returns the FIFO to empty
== 36 checks, 0 failures ==

Verilog-2001    : 36 checks, 0 failures
SystemVerilog   : 36 checks, 0 failures
VHDL-2008       : 36 checks, 0 failures

9. Verification

An asynchronous FIFO must be tested with clocks that have no rational relationship, or the test is not exercising the crossing. Two clocks at a 2:1 ratio produce a repeating edge pattern that visits a small set of phase relationships; 7 ns against 13 ns does not.

Sweep the ratio in both directions. Fast writer with slow reader exercises full and the write-side pointer crossing; the reverse exercises empty. A test where both sides run at similar rates and the FIFO hovers half-full exercises neither flag.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// Assertion — in the write domain. The pessimism argument, stated: full may
// be wrong, but only by claiming full when the FIFO is not.
property p_full_is_pessimistic;
    @(posedge wclk) disable iff (!wrst_n)
        wfull_o |-> (wbin_q - rbin_q_observed) >= 0;
endproperty

// Assertion — data integrity, which is what the crossing must not break.
// Every word read is the word written that many pushes ago, in order.
// Checked procedurally in the suite above: 0 mismatches over ~1500 transfers.

// Assertion — the depth is never exceeded. This is the one that catches a
// broken `full`, and it must be computed by the TESTBENCH, which is allowed
// to know both pointers at once. The DESIGN is not.
property p_never_overfilled;
    @(posedge wclk) (tb_pushes - tb_pops) <= DEPTH;
endproperty

The testbench may look at both domains; the design may not. That asymmetry is the whole point of a scoreboard here, and it is also the trap: it is easy to write a check that quietly assumes an ordering between the two clocks that does not exist. A scoreboard for an asynchronous FIFO should count events, not compare states at an instant.

Run CDC lint on the result (Chapter 12.2 §8). It verifies structurally what the simulation cannot: that every crossing has a synchroniser, that the gray encoding is actually in the crossing path, and that no combinational logic slipped between the stages.

Mutation testing, and the one defect that got away

Four defects were installed in uart_async_fifo, one at a time, each verified to have actually changed the source before the result was scored.

#Defect installedResult
M6pointers cross as binary, not graykilled — the gray invariant fires, then the FIFO deadlocks
M7full compares plain equality, without the one-lap inversionkilled — the design hangs; caught by the watchdog
M8the write pointer advances even when fullkilled, 5 checks
M9the pointer synchronisers cut to one flopsurvived every behavioural check

M6 and M7 are killed by deadlock rather than by a mismatch, which is why every loop in the testbench carries a guard and the run carries a watchdog. A test that hangs reports nothing; a test that hangs and times out reports a failure. The difference is a few lines and it is the difference between a result and a hung CI job.

M9 is the one that matters. Cutting each two-flop pointer synchroniser to a single flop is not a subtle defect — it is the exact hazard this entire module exists to prevent, and in silicon it is the difference between a FIFO that runs for years and one that corrupts a pointer when two edges land too close together.

It changed nothing a simulation could see:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
--- M9: one-flop pointer synchronisers, behavioural checks only ---
  PASS  all 34 behavioural checks
  the FIFO lost nothing, duplicated nothing, reordered nothing,
  full and empty were correct, capacity was exactly DEPTH

The FIFO is still functionally perfect, because the second flop's job is to give the first one a clock in which to resolve — and in an event simulator there is nothing to resolve. Every data check, every flag check, every capacity check passes. The design is broken and the testbench is happy.

What finally caught it was the third structural observer: after any write-clock edge, wq2_rgray must hold the value wq1_rgray had before that edge. With one flop the two move together and the invariant breaks on the first pointer movement.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
--- M9 against the structural observer ---
  FAIL write-side pointer synchroniser is TWO flops deep, not one
  FAIL read-side pointer synchroniser is TWO flops deep, not one
== 36 checks, 2 failures ==

Two checks out of thirty-six, and they are the only two in the file that could possibly have fired. Note what they do: they do not test behaviour at all. They test the shape of the logic, which is what a CDC linter does against a netlist — performed here with the only instrument a simulation has, a hierarchical reference and a one-deep delay.

10. What This Means on an FPGA

Ask first whether you need two clocks at all. A UART running on the system clock with a fractional baud generator (Chapter 8.3) removes this entire chapter from the design. That is frequently the right answer and it is frequently not considered, because "the UART should have its own clock" sounds tidy.

Vendor FIFO generators exist and are usually better than a hand-written one. Xilinx FIFO Generator and Intel's DCFIFO implement this structure, are verified, and map onto block RAM with the correct primitives. Write your own to understand it; instantiate theirs to ship it — unless you need something they do not offer.

The extra pointer bit is not optional. It is what distinguishes full from empty when the pointers are otherwise equal. A FIFO built with DEPTH_L2 pointer bits instead of DEPTH_L2 + 1 reports empty when it is full, and the symptom is catastrophic and intermittent.

Constrain the crossings. The gray pointer paths need set_max_delay or equivalent, not a default setup check between unrelated clocks — Chapter 12.5. Left alone, the tool will either fail timing on a path that has no real requirement, or meet an imaginary one while allowing skew that breaks the one-bit-at-a-time guarantee.

11. Understanding Check

12. Summary

Asynchronous serial and clock-domain crossing are different subjects wearing the same word. The first is solved by oversampling and is the whole protocol; the second is solved by synchronisers and, in the UART as built, applies to two input pins and nothing else.

Both confusions are expensive: CDC structures on the serial path add latency where margin is scarce, and a missed crossing produces intermittent failures that no test reproduces.

A real crossing appears when a second clock does — in practice, a register interface on the bus clock. Three kinds then need three mechanisms: levels take a synchroniser, pulses must become levels before crossing, and multi-bit values take gray code, a handshake, or a FIFO.

Gray code does not make a bad sample less likely — it makes it unrepresentable. Measured with per-bit skew: a binary pointer produced 215 fabricated values in 3,999 samples (5.38%), including reading 0 while the counter held 8. The gray pointer produced none, structurally.

The asynchronous FIFO keeps a pointer per domain, crosses them as gray, compares locally against a deliberately stale view, and registers both flags — which is required, not stylistic, because a combinational flag closes a real loop. Verified across 7 ns and 13 ns clocks with zero data mismatches.

The Module 10 FIFO cannot be adapted: its shared binary level counter is the exact failure the measurement demonstrates, and synchronising it repairs nothing.

And the question that resolves every case: is this signal generated by one clock and captured by another? Not whether it is slow, external, or has "asynchronous" in its name.

13. What Comes Next

Two clocks raise a second question this chapter has quietly set aside: what does reset mean when there are two of them?

Chapter 12.4 takes reset seriously — asynchronous assertion with synchronous release, why the release edge is the dangerous one, how reset is distributed across domains, and the question the whole module has been building toward: what a reset asserted mid-frame leaves behind, on both ends of a link where only one end was reset.

Browse the full path on the UART tutorials index. For the synchronous FIFO this chapter contrasts against, read back to Chapter 10.2.

Continue learning

Where this fits

Part of the UART curriculum.