Skip to content
VLSI Mentor

UART · Module 12

Synchroniser Placement and Edge Detection After Synchronisation

Where the synchroniser belongs, what its latency costs the start-edge measurement, and a real placement defect found in the IP built in Module 11 — which every functional test passed.

Chapter 12.1 established what a synchroniser guarantees. This chapter is about where it goes, which sounds like a detail and is the part that actually gets built wrong.

The rule is one sentence: synchronise first, then do everything else. Edge detection, glitch filtering, counting, comparison — all of it operates on the synchronised signal, never on the pin.

The rest of the chapter is about why that rule has to be stated as a rule. It ends with a placement defect in the IP assembled in Chapter 11.2, which passed 86 functional checks across four testbenches and three parameter sets, because — per Chapter 12.1 §7 — it was always going to.

1. The Rule, and the Tempting Violation

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// CORRECT — this is what uart_rx has done since Chapter 6.1.
always_ff @(posedge clk or negedge rst_n) begin
    if (!rst_n) begin
        rx_meta_q   <= 1'b1;
        rx_sync_q   <= 1'b1;
        rx_sync_d_q <= 1'b1;
    end else begin
        rx_meta_q   <= rx_i;        // may go metastable — read by NOTHING else
        rx_sync_q   <= rx_meta_q;   // settled
        rx_sync_d_q <= rx_sync_q;   // history, for edge detection
    end
end

// Edge detection on SYNCHRONISED history only.
assign start_cand = rx_sync_d_q && !rx_sync_q;

The tempting alternative detects the edge as early as possible, on the pin itself:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// WRONG — the start edge is taken from the RAW pin while every data sample
// still comes from the synchronised one.
logic rx_raw_d_q;
always_ff @(posedge clk or negedge rst_n)
    if (!rst_n) rx_raw_d_q <= 1'b1; else rx_raw_d_q <= rx_i;

assign start_cand = rx_raw_d_q && !rx_i;    // combinational on the pin

It is tempting for a real reason: it detects the start edge two clocks sooner, and start-edge latency eats directly into the receiver's timing budget. The instinct is sound and the implementation is not.

2. What the Synchroniser Costs, Measured

Dropping the line 1 ns after a clock edge at 100 MHz and timestamping each stage:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
=== 100 MHz clock, 10.00 ns period. Falling edge on rx at t=286.00 ns
  rx_meta_q changed at      295.00 ns  (+0.90 clocks)
  rx_sync_q changed at      305.00 ns  (+1.90 clocks)
  FSM consumes start_cand   315.00 ns  (+2.90 clocks)
  raw-edge FSM consumes     295.00 ns  (+0.90 clocks)
  raw-edge detects 2.00 clocks EARLIER

Three clock edges, correct placement; one, raw. The difference is exactly two clocks — the two synchroniser stages — and it is the entire benefit the wrong version buys.

What those two clocks are worth depends only on the clocks-per-bit ratio:

Clock / baudclocks per bitclocks per os tick2 clocks as a fraction of a bit
100 MHz / 115,200868.0654.250.230%
50 MHz / 460,800108.516.781.84%
24 MHz / 460,80052.083.263.84%
24 MHz / 921,60026.041.637.68%
20 MHz / 1,000,00020.001.2510.0%

The right-hand column is what the argument is actually about. At 868 clocks per bit the synchroniser is free. At 20 clocks per bit it costs a tenth of a bit time — and the receiver's budget, built in Chapter 4.5, has to absorb it alongside everything else.

A block diagram comparing two synchroniser placements. In the correct arrangement along the top, the receive pin feeds the first synchroniser flop, which feeds the second, which feeds both a history register and an edge detector; the edge detector reads only synchronised values and its output goes to the state machine, and the same synchronised signal also supplies the data sampler. In the wrong arrangement along the bottom, the receive pin feeds a single raw history register and also feeds an edge detector combinationally, so the edge detector output is driven directly by the unsynchronised pin and fans out into the state machine's next-state logic, while the data sampler still reads the synchronised signal from a separate path.rx_ithe pinrx_meta_qisolatedrx_sync_qsettlededge detecton sync historyFSM + samplerONE view of the linerx_ithe pinrx_raw_d_qraw historyedge detectCOMBINATIONAL on pinFSM next-stateunresolved fan-outsamplerreads sync — 2 clkslaterasyncresolvesafestartasyncstartdisagree12
Figure 1 — the two placements. Correct: the pin is synchronised, and the edge detector reads only synchronised history, so no combinational logic ever touches the pin. Wrong: the edge detector's output is combinational on the pin, so an unresolved value fans out into the state machine's next-state logic — while the data path still reads the synchronised signal, leaving the two halves reasoning about different versions of the same line.

Sync-then-edge versus raw edge

8 cycles
A timing trace of eight clock cycles comparing two synchroniser placements against the same line transition. A clock runs throughout. The asynchronous receive line falls during the second cycle. In the correct placement, the first synchroniser flop captures the change at the next clock edge, the second flop a clock later, and the state machine consumes the resulting start candidate at the third clock edge after the transition. In the wrong placement, the start candidate is produced combinationally from the pin and is consumed by the state machine at the first clock edge after the transition, two clocks earlier. Both are shown decoding the same frame successfully, because an event simulator cannot represent the unresolved value that makes the second arrangement unsafe.raw acts here — 2 clocks earlyraw acts here — 2 clocksearlycorrect acts herecorrect acts hereclkrx_i (async)rx_meta_qrx_sync_qstart (sync)start (raw)t0t1t2t3t4t5t6t7
Figure 2 — the same line transition through both placements. The correct path reaches the state machine at the third clock edge; the raw path reaches it at the first, two clocks sooner. Both decode the byte correctly in simulation. The difference is not what they decode — it is that the raw path's start candidate is driven combinationally by the pin.

3. One Synchroniser Per Signal

A second rule, less often stated and just as binding: an asynchronous signal is synchronised once, and every consumer reads that same synchronised copy.

Two independent synchronisers on one pin each resolve independently. When a transition lands near an edge, one may resolve to 0 and the other to 1 — on the same clock, from the same wire. The two consumers then disagree about when the line changed, and the design contains a contradiction that no single block is wrong about.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// WRONG — two consumers, two synchronisers, two opinions.
uart_rx        u_rx   (.rx_i(rx_line),      ...);   // synchronises internally
uart_break_det u_break(.rx_sync_i(rx_line), ...);   // ALSO sees the raw pin

// RIGHT — one synchroniser, published, shared.
uart_rx        u_rx   (.rx_i(rx_line), .rx_sync_o(rx_synced), ...);
uart_break_det u_break(.rx_sync_i(rx_synced), ...);

4. The Defect in the Module 11 IP

The wrong version above is not hypothetical. It is what Chapter 11.2 shipped.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// uart_break_detect's port, as declared in Module 9:
input  logic rx_sync_i,          // synchronised line — Chapter 5.1

// what uart_ip connected to it:
uart_break_detect #(...) u_break (
    .clk(clk), .rst_n(rst_n), .os_tick_i(os_tick),
    .rx_sync_i(rx_line),       // <-- the RAW PIN
    ...);

The port documents its own requirement in a comment, and the integration violated it. The break detector has no synchroniser of its own — it registers rx_sync_i straight into its low-tick counter — so the assembled IP put an asynchronous input directly onto a flop, in a block whose interface says in words that it must not.

The fix gives the IP one synchroniser and shares it:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// in uart_rx — publish the line it has already synchronised
output logic rx_sync_o
...
assign rx_sync_o = rx_sync_q;

// in uart_ip — the break detector now sees the SAME resolved value the
// receiver does, two clocks after the pin, exactly like every other consumer.
uart_rx #(...) u_rx (..., .rx_active_o(rx_active), .rx_sync_o(rx_synced));

uart_break_detect #(...) u_break (
    .clk(clk), .rst_n(rst_n), .os_tick_i(os_tick),
    .rx_sync_i(rx_synced), ...);

The cost is two clocks of extra break-detection latency against a threshold of eleven bit times — 20 ns against 95.5 µs at 100 MHz and 115,200 baud, or 0.02% of the detection window. There was never a trade-off to weigh.

Regression after the fix: 86 checks, 0 failures across all four Module 11 suites — unchanged, as expected.

5. The Experiment That Did Not Work

Wanting a measurable consequence rather than only a structural argument, I built the raw-edge receiver of §1 and ran it beside the correct one at five clocks-per-bit ratios, from 868 down to 20, on eight byte patterns each.

Both received all eight bytes correctly at every ratio.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
-- 100 MHz / 115200 baud : 868.06 clk/bit, 54.253 clk per os tick
     sync-then-edge (correct) : 8/8 delivered, 8 correct
     raw-edge       (wrong)   : 8/8 delivered, 8 correct
-- 20 MHz / 1000000 baud : 20.00 clk/bit, 1.250 clk per os tick
     sync-then-edge (correct) : 8/8 delivered, 8 correct
     raw-edge       (wrong)   : 8/8 delivered, 8 correct

The hypothesis was wrong. I expected the two-clock head start to shift the sampling grid enough to erode margin at low clocks-per-tick ratios. It does shift the grid — §2 measured exactly two clocks — but the receiver re-centres on the start edge it detected, so the whole grid moves together and the samples still land near bit centres. A 1.6-tick shift against roughly ±8 ticks of margin is simply absorbed.

6. The Synchroniser As a Named Block

Everything above argues that one structure is correct and a very similar one is not. Section 4 then showed an integration that got it wrong — and got it wrong precisely because the correct structure existed only as lines typed inside uart_rx, not as a thing with a name. A block that is re-typed is a block that can be re-typed incorrectly, and no reviewer can diff a pattern against a pattern that was never written down.

So write it down. uart_sync_edge is the §1 rule with a port list: STAGES flops of synchroniser, edge detection taken from the settled history, and a reset value chosen so that leaving reset does not look like a start bit.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// ===========================================================================
//  uart_sync_edge — Synthesizable SystemVerilog
//
//  The block Module 12 is about, published as a module rather than as a
//  fragment inside the receiver. Two jobs, and the order is the whole point:
//
//    1. Bring an ASYNCHRONOUS input into this clock domain through STAGES
//       flops. The first flop may go metastable; nothing but the second flop
//       is allowed to read it.
//    2. Detect edges on the SYNCHRONISED history only — never on the pin.
//
//  Chapter 12.2 shows what edge-detecting the raw pin costs. This module
//  exists so that the correct structure has one name, one owner and one
//  testbench instead of being re-typed inside every consumer.
// ===========================================================================
module uart_sync_edge #(
    parameter int unsigned STAGES      = 2,
    // The value the chain holds during reset. For a UART receive line this
    // is MARK: coming out of reset believing the line is low would look
    // exactly like a start bit.
    parameter logic        RESET_VALUE = 1'b1
) (
    input  logic clk,
    input  logic rst_n,
    input  logic async_i,     // genuinely asynchronous — the pin
    output logic sync_o,      // settled, safe for the whole domain to read
    output logic rise_o,      // ONE cycle, on a synchronised 0 -> 1
    output logic fall_o       // ONE cycle, on a synchronised 1 -> 0
);
    if (STAGES < 2) begin : g_bad_stages
        $error("uart_sync_edge: STAGES must be at least 2");
    end

    // chain_q[0] is the metastability-absorbing flop. It is READ BY NOTHING
    // except chain_q[1]; that restriction is the synchroniser.
    logic [STAGES-1:0] chain_q;
    logic              sync_d_q;      // one-cycle history, for the edges

    always_ff @(posedge clk or negedge rst_n) begin
        if (!rst_n) begin
            chain_q  <= {STAGES{RESET_VALUE}};
            sync_d_q <= RESET_VALUE;
        end else begin
            chain_q  <= {chain_q[STAGES-2:0], async_i};
            sync_d_q <= chain_q[STAGES-1];
        end
    end

    assign sync_o = chain_q[STAGES-1];

    // Edges from SYNCHRONISED history only. Both are one clock wide because
    // they compare two adjacent samples of the same settled signal.
    assign rise_o = ~sync_d_q &  sync_o;
    assign fall_o =  sync_d_q & ~sync_o;
endmodule

The Verilog-2001 version is the same design with the two facilities that version of the language does not have written out: no logic, and no elaboration-time $error, so the parameter check becomes a runtime one in an initial block. It is a weaker guard — it fires when the simulation starts rather than when the design elaborates — and it is the strongest guard Verilog-2001 offers.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
//===========================================================================
//  uart_sync_edge_v — Synthesizable Verilog-2001
//
//  Two jobs, and the order is the whole point:
//    1. Bring an ASYNCHRONOUS input into this clock domain through STAGES
//       flops. The first flop may go metastable; nothing but the second is
//       allowed to read it.
//    2. Detect edges on the SYNCHRONISED history only -- never on the pin.
//
//  Verilog-2001 has no elaboration-time assertion, so the STAGES >= 2 rule
//  is a documented contract here rather than an enforced one. Chapter 11.4
//  section 4 covers that gap.
//===========================================================================
module uart_sync_edge_v #(
    parameter STAGES      = 2,     // MUST be >= 2 -- not enforceable in 2001
    // The value the chain holds during reset. For a UART receive line this
    // is MARK: leaving reset believing the line is low looks exactly like a
    // start bit.
    parameter RESET_VALUE = 1'b1
) (
    input  wire clk,
    input  wire rst_n,
    input  wire async_i,       // genuinely asynchronous -- the pin
    output wire sync_o,        // settled, safe for the whole domain to read
    output wire rise_o,        // ONE cycle, on a synchronised 0 -> 1
    output wire fall_o         // ONE cycle, on a synchronised 1 -> 0
);
    // chain_q[0] is the metastability-absorbing flop. It is READ BY NOTHING
    // except chain_q[1]; that restriction IS the synchroniser.
    reg [STAGES-1:0] chain_q;
    reg              sync_d_q;

    always @(posedge clk or negedge rst_n) begin
        if (!rst_n) begin
            chain_q  <= {STAGES{RESET_VALUE}};
            sync_d_q <= RESET_VALUE;
        end else begin
            chain_q  <= {chain_q[STAGES-2:0], async_i};
            sync_d_q <= chain_q[STAGES-1];
        end
    end

    assign sync_o = chain_q[STAGES-1];

    // Edges from SYNCHRONISED history only. Both are one clock wide because
    // they compare two adjacent samples of the same settled signal.
    assign rise_o = ~sync_d_q &  sync_o;
    assign fall_o =  sync_d_q & ~sync_o;
endmodule

VHDL expresses the parameter check best of the three: a concurrent assert with severity failure inside the architecture, checked at elaboration, needing no process and no simulation time.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
--==========================================================================
--  uart_sync_edge -- Synthesizable VHDL-2008
--
--  Two jobs, and the order is the whole point:
--    1. Bring an ASYNCHRONOUS input into this clock domain through STAGES
--       flops. The first flop may go metastable; nothing but the second is
--       allowed to read it.
--    2. Detect edges on the SYNCHRONISED history only -- never on the pin.
--==========================================================================
library ieee;
use ieee.std_logic_1164.all;

entity uart_sync_edge is
    generic (
        STAGES      : positive  := 2;
        -- The value the chain holds during reset. For a UART receive line
        -- this is MARK: leaving reset believing the line is low looks
        -- exactly like a start bit.
        RESET_VALUE : std_logic := '1'
    );
    port (
        clk     : in  std_logic;
        rst_n   : in  std_logic;
        async_i : in  std_logic;    -- genuinely asynchronous -- the pin
        sync_o  : out std_logic;    -- settled, safe for the domain to read
        rise_o  : out std_logic;    -- ONE cycle, on a synchronised 0 -> 1
        fall_o  : out std_logic     -- ONE cycle, on a synchronised 1 -> 0
    );
end entity uart_sync_edge;

architecture rtl of uart_sync_edge is
    -- chain(0) is the metastability-absorbing flop. It is READ BY NOTHING
    -- except chain(1); that restriction IS the synchroniser.
    signal chain_q  : std_logic_vector(STAGES-1 downto 0);
    signal sync_d_q : std_logic;
    signal sync_s   : std_logic;
begin

    assert STAGES >= 2
        report "uart_sync_edge: STAGES must be at least 2" severity failure;

    sync_s <= chain_q(STAGES-1);
    sync_o <= sync_s;

    -- Edges from SYNCHRONISED history only. Both are one clock wide because
    -- they compare two adjacent samples of the same settled signal.
    rise_o <= (not sync_d_q) and sync_s;
    fall_o <= sync_d_q and (not sync_s);

    process (clk, rst_n)
    begin
        if rst_n = '0' then
            chain_q  <= (others => RESET_VALUE);
            sync_d_q <= RESET_VALUE;
        elsif rising_edge(clk) then
            chain_q  <= chain_q(STAGES-2 downto 0) & async_i;
            sync_d_q <= chain_q(STAGES-1);
        end if;
    end process;

end architecture rtl;

7. Three Testbenches That Say What They Cannot Prove

A testbench for a synchroniser has an unusual obligation: it has to be explicit about the fact that it cannot test the thing the block exists for. There is no metastability in an RTL simulation, so there is nothing for the first flop to absorb, and a design with the flop and a design without it produce the same waveform. Section 5 demonstrated exactly that over forty frames.

What the testbench can pin down is everything else, and everything else is worth pinning down:

PropertyHow it is checkedKills
a pin change appears after exactly STAGES clocksmeasured with a counting loop, not assumedM1
the same change at STAGES=3 takes threea second instance, same stimulusM1
RESET_VALUE holds the chain during resettwo instances, reset MARK and reset SPACEM3
edge pulses are exactly one clock widea whole-run observer counting wide pulses—
rise_o and fall_o never assert togetherwhole-run observer—
an edge only ever follows a real sync_o movewhole-run observer comparing against historyM2
a one-clock pulse still produces its edgesexplicit test — a synchroniser is not a filter—

That last row is the one that surprises people. A synchroniser aligns; it does not debounce. A one-clock glitch on the pin arrives as a one-clock glitch on sync_o, and rejecting it is the start-bit validation logic's job from Chapter 5.2, not this block's. A testbench that did not check this would leave room for someone to "improve" the synchroniser into a filter and break start detection in a way that only shows up on a noisy line.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
//===========================================================================
//  tb_uart_sync_edge — self-checking SystemVerilog testbench
//
//  WHAT A SIMULATION CAN AND CANNOT CHECK HERE, stated up front because it
//  is the whole lesson of Module 12:
//
//    CAN    — the LATENCY contract (STAGES clocks), the pulse widths, the
//             reset value, and that edges are derived from synchronised
//             history rather than from the pin.
//    CANNOT — that the first flop actually absorbs metastability. There is
//             no metastability in an RTL simulation to absorb. That property
//             is checked by CDC lint against the netlist, and a clean
//             simulation is not evidence for it.
//
//  This file carries the same 19 counted checks as its Verilog twin, plus a
//  layer of SystemVerilog immediate assertions that fire the instant an
//  invariant breaks rather than at the end of the run. The assertions do not
//  replace the counted checks — an assertion that never fires proves nothing
//  unless the stimulus reached the situation it guards.
//===========================================================================
`timescale 1ns/1ps

module tb_uart_sync_edge;

    logic clk = 1'b0, rst_n = 1'b0;
    always #5 clk = ~clk;

    logic a2 = 1'b1, a3 = 1'b1, a0 = 1'b0;
    logic s2, r2, f2;      // STAGES=2, reset MARK
    logic s3, r3, f3;      // STAGES=3, reset MARK
    logic s0, r0, f0;      // STAGES=2, reset SPACE

    uart_sync_edge #(.STAGES(2), .RESET_VALUE(1'b1)) d2 (
        .clk(clk), .rst_n(rst_n), .async_i(a2),
        .sync_o(s2), .rise_o(r2), .fall_o(f2));
    uart_sync_edge #(.STAGES(3), .RESET_VALUE(1'b1)) d3 (
        .clk(clk), .rst_n(rst_n), .async_i(a3),
        .sync_o(s3), .rise_o(r3), .fall_o(f3));
    uart_sync_edge #(.STAGES(2), .RESET_VALUE(1'b0)) d0 (
        .clk(clk), .rst_n(rst_n), .async_i(a0),
        .sync_o(s0), .rise_o(r0), .fall_o(f0));

    int checks = 0, failures = 0;

    task automatic check(input logic cond, input string name);
        checks++;
        if (cond) $display("  PASS %0s", name);
        else begin failures++; $display("  FAIL %0s", name); end
    endtask

    // Whole-run observers, computed from the OUTPUTS alone.
    int   n_rise = 0, n_fall = 0, wide = 0, both = 0;
    logic r_prev = 1'b0, f_prev = 1'b0;
    // Plain `always`, not `always_ff`: this is an observer, not hardware, and
    // it deliberately contains a $error call that no synthesiser would accept.
    always @(posedge clk) if (rst_n) begin
        if (r2) begin n_rise++; if (r_prev) wide++; end
        if (f2) begin n_fall++; if (f_prev) wide++; end
        if (r2 && f2) both++;
        // SystemVerilog immediate assertion: rise and fall are mutually
        // exclusive by construction, so a failure here is a design defect,
        // not a stimulus one.
        a_excl: assert (!(r2 && f2))
            else $error("rise_o and fall_o asserted in the same clock");
        r_prev <= r2; f_prev <= f2;
    end

    // An edge may only appear one clock after sync_o actually moved.
    int   edge_without_move = 0;
    logic s_prev = 1'b1;
    always @(posedge clk) if (rst_n) begin
        if ((r2 || f2) && (s2 === s_prev)) edge_without_move++;
        a_move: assert (!((r2 || f2) && (s2 === s_prev)))
            else $error("an edge pulse without a sync_o transition");
        s_prev <= s2;
    end

    int lat, i, base;

    initial begin
        #400_000;
        $display("  FAIL watchdog: simulation did not finish");
        $display("== %0d checks, %0d failures ==", checks+1, failures+1);
        $display("   RESULT: SYSTEMVERILOG SYNC-EDGE TESTS FAILED (timeout)");
        $finish;
    end

    initial begin
        $display("== uart_sync_edge : self-checking SystemVerilog testbench ==");

        // ---- reset contract ------------------------------------------------
        rst_n = 1'b0; a2 = 1'b1; a3 = 1'b1; a0 = 1'b0;
        repeat (6) @(negedge clk);
        check(s2 === 1'b1, "reset: RESET_VALUE=1 holds sync at MARK");
        check(s0 === 1'b0, "reset: RESET_VALUE=0 holds sync at SPACE");
        check(r2 === 1'b0 && f2 === 1'b0, "reset: no edge pulses");
        @(negedge clk) rst_n = 1'b1;
        repeat (4) @(negedge clk);
        check(r2 === 1'b0 && f2 === 1'b0, "leaving reset with a steady input: no edge");

        // ---- THE latency contract, measured not assumed --------------------
        @(negedge clk) a2 = 1'b0;
        lat = 0;
        while (s2 !== 1'b0 && lat < 10) begin @(posedge clk); #1; lat++; end
        check(lat == 2, "STAGES=2: a pin change takes exactly TWO clocks");

        @(negedge clk) a3 = 1'b0;
        lat = 0;
        while (s3 !== 1'b0 && lat < 10) begin @(posedge clk); #1; lat++; end
        check(lat == 3, "STAGES=3: the same change takes exactly THREE clocks");
        check(s2 === 1'b0, "and the 2-stage instance already settled");

        // ---- the edge pulses ------------------------------------------------
        base = n_fall;
        @(negedge clk) a2 = 1'b1;
        repeat (4) @(negedge clk);
        check(n_fall == base, "returning to mark raises no FALL");
        base = n_rise;
        @(negedge clk) a2 = 1'b0;
        repeat (4) @(negedge clk);
        check(n_rise == base, "going to space raises no RISE");

        // ---- a start-bit edge, which is what this block exists for ---------
        @(negedge clk) a2 = 1'b1;
        repeat (4) @(negedge clk);
        base = n_fall;
        @(negedge clk) a2 = 1'b0;
        repeat (4) @(negedge clk);
        check(n_fall == base + 1, "a 1->0 pin transition yields exactly one FALL");
        check(f2 === 1'b0,        "and the pulse has already returned low");

        // ---- a one-clock pulse is PASSED, not filtered ---------------------
        @(negedge clk) a2 = 1'b1;
        repeat (4) @(negedge clk);
        base = n_fall;
        @(negedge clk) a2 = 1'b0;
        @(negedge clk) a2 = 1'b1;            // one clock wide
        repeat (5) @(negedge clk);
        check(n_fall == base + 1, "a one-clock low pulse still produces its FALL");
        check(n_rise > 0,         "and its matching RISE -- alignment, not filtering");

        // ---- back-to-back transitions ---------------------------------------
        base = n_fall;
        for (i = 0; i < 8; i++) begin
            @(negedge clk) a2 = 1'b0;
            repeat (3) @(negedge clk);
            @(negedge clk) a2 = 1'b1;
            repeat (3) @(negedge clk);
        end
        check(n_fall == base + 8, "eight transitions produce exactly eight FALLs");

        // ---- reset during activity ------------------------------------------
        @(negedge clk) a2 = 1'b0;
        @(negedge clk) rst_n = 1'b0;
        repeat (3) @(negedge clk);
        check(s2 === 1'b1, "reset mid-activity forces sync back to RESET_VALUE");
        check(r2 === 1'b0 && f2 === 1'b0, "reset suppresses the edge outputs");
        @(negedge clk) rst_n = 1'b1;
        repeat (4) @(negedge clk);

        // ---- whole-run invariants -------------------------------------------
        check(wide == 0,              "no edge pulse ever wider than one clock");
        check(both == 0,              "rise and fall never assert together");
        check(edge_without_move == 0, "an edge only ever follows a real sync_o move");

        $display("== %0d checks, %0d failures ==", checks, failures);
        if (failures == 0) $display("   RESULT: ALL SYSTEMVERILOG SYNC-EDGE TESTS PASSED");
        else               $display("   RESULT: SYSTEMVERILOG SYNC-EDGE TESTS FAILED");
        $finish;
    end
endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
//===========================================================================
//  tb_uart_sync_edge_v — self-checking Verilog-2001 testbench
//
//  WHAT A SIMULATION CAN AND CANNOT CHECK HERE, stated up front because it
//  is the whole lesson of Module 12:
//
//    CAN   -- the LATENCY contract (STAGES clocks), the pulse widths, the
//             reset value, and that edges are derived from synchronised
//             history rather than from the pin.
//    CANNOT -- that the first flop actually absorbs metastability. There is
//             no metastability in an RTL simulation to absorb. That property
//             is checked by CDC lint against the netlist, and a clean
//             simulation is not evidence for it.
//
//  INDEPENDENCE: every expectation comes from the stimulus the testbench
//  drove, never from reading chain_q.
//===========================================================================
`timescale 1ns/1ps

module tb_uart_sync_edge_v;

    reg clk = 1'b0, rst_n = 1'b0;
    always #5 clk = ~clk;

    reg  a2 = 1'b1, a3 = 1'b1, a0 = 1'b0;
    wire s2, r2, f2;      // STAGES=2, reset MARK
    wire s3, r3, f3;      // STAGES=3, reset MARK
    wire s0, r0, f0;      // STAGES=2, reset SPACE

    uart_sync_edge_v #(.STAGES(2), .RESET_VALUE(1'b1)) d2 (
        .clk(clk), .rst_n(rst_n), .async_i(a2),
        .sync_o(s2), .rise_o(r2), .fall_o(f2));
    uart_sync_edge_v #(.STAGES(3), .RESET_VALUE(1'b1)) d3 (
        .clk(clk), .rst_n(rst_n), .async_i(a3),
        .sync_o(s3), .rise_o(r3), .fall_o(f3));
    uart_sync_edge_v #(.STAGES(2), .RESET_VALUE(1'b0)) d0 (
        .clk(clk), .rst_n(rst_n), .async_i(a0),
        .sync_o(s0), .rise_o(r0), .fall_o(f0));

    integer checks = 0, failures = 0;

    task check;
        input cond;
        input [8*80-1:0] name;
        begin
            checks = checks + 1;
            if (cond) $display("  PASS %0s", name);
            else begin failures = failures + 1; $display("  FAIL %0s", name); end
        end
    endtask

    // Whole-run observers, computed from the OUTPUTS alone.
    integer n_rise = 0, n_fall = 0, wide = 0, both = 0;
    reg r_prev = 1'b0, f_prev = 1'b0;
    always @(posedge clk) if (rst_n) begin
        if (r2) begin n_rise = n_rise + 1; if (r_prev) wide = wide + 1; end
        if (f2) begin n_fall = n_fall + 1; if (f_prev) wide = wide + 1; end
        if (r2 && f2) both = both + 1;          // must never both fire
        r_prev = r2; f_prev = f2;
    end

    // An edge may only appear one clock after sync_o actually moved.
    integer edge_without_move = 0;
    reg s_prev = 1'b1;
    always @(posedge clk) if (rst_n) begin
        if ((r2 || f2) && (s2 === s_prev)) edge_without_move = edge_without_move + 1;
        s_prev = s2;
    end

    integer lat, i, base;

    initial begin
        #400_000;
        $display("  FAIL watchdog: simulation did not finish");
        $display("== %0d checks, %0d failures ==", checks+1, failures+1);
        $display("   RESULT: VERILOG SYNC-EDGE TESTS FAILED (timeout)");
        $finish;
    end

    initial begin
        $display("== uart_sync_edge_v : self-checking Verilog testbench ==");

        // ---- reset contract ------------------------------------------------
        rst_n = 1'b0; a2 = 1'b1; a3 = 1'b1; a0 = 1'b0;
        repeat (6) @(negedge clk);
        check(s2 === 1'b1, "reset: RESET_VALUE=1 holds sync at MARK");
        check(s0 === 1'b0, "reset: RESET_VALUE=0 holds sync at SPACE");
        check(r2 === 1'b0 && f2 === 1'b0, "reset: no edge pulses");
        @(negedge clk) rst_n = 1'b1;
        repeat (4) @(negedge clk);
        check(r2 === 1'b0 && f2 === 1'b0, "leaving reset with a steady input: no edge");

        // ---- THE latency contract, measured not assumed --------------------
        // A change on the pin must take exactly STAGES clocks to appear on
        // sync_o. The rest of the receiver's timing budget depends on this
        // number, which is why it is measured rather than trusted.
        @(negedge clk) a2 = 1'b0;
        lat = 0;
        while (s2 !== 1'b0 && lat < 10) begin
            @(posedge clk); #1; lat = lat + 1;
        end
        check(lat == 2, "STAGES=2: a pin change takes exactly TWO clocks");

        @(negedge clk) a3 = 1'b0;
        lat = 0;
        while (s3 !== 1'b0 && lat < 10) begin
            @(posedge clk); #1; lat = lat + 1;
        end
        check(lat == 3, "STAGES=3: the same change takes exactly THREE clocks");
        check(s2 === 1'b0, "and the 2-stage instance already settled");

        // ---- the edge pulses ------------------------------------------------
        base = n_fall;
        @(negedge clk) a2 = 1'b1;
        repeat (4) @(negedge clk);
        check(n_fall == base, "returning to mark raises no FALL");
        base = n_rise;
        @(negedge clk) a2 = 1'b0;
        repeat (4) @(negedge clk);
        check(n_rise == base, "going to space raises no RISE");

        // ---- a start-bit edge, which is what this block exists for ---------
        @(negedge clk) a2 = 1'b1;
        repeat (4) @(negedge clk);
        base = n_fall;
        @(negedge clk) a2 = 1'b0;            // the falling edge = start candidate
        repeat (4) @(negedge clk);
        check(n_fall == base + 1, "a 1->0 pin transition yields exactly one FALL");
        check(f2 === 1'b0,        "and the pulse has already returned low");

        // ---- a one-clock pulse is PASSED, not filtered ---------------------
        // A synchroniser aligns; it does not debounce. A pulse one clock wide
        // arrives one clock wide, and rejecting it is the start-validation
        // logic's job (Chapter 5.2), not this block's.
        @(negedge clk) a2 = 1'b1;
        repeat (4) @(negedge clk);
        base = n_fall;
        @(negedge clk) a2 = 1'b0;
        @(negedge clk) a2 = 1'b1;            // one clock wide
        repeat (5) @(negedge clk);
        check(n_fall == base + 1, "a one-clock low pulse still produces its FALL");
        check(n_rise > 0,         "and its matching RISE -- alignment, not filtering");

        // ---- back-to-back transitions ---------------------------------------
        base = n_fall;
        for (i = 0; i < 8; i = i + 1) begin
            @(negedge clk) a2 = 1'b0;
            repeat (3) @(negedge clk);
            @(negedge clk) a2 = 1'b1;
            repeat (3) @(negedge clk);
        end
        check(n_fall == base + 8, "eight transitions produce exactly eight FALLs");

        // ---- reset during activity ------------------------------------------
        @(negedge clk) a2 = 1'b0;
        @(negedge clk) rst_n = 1'b0;
        repeat (3) @(negedge clk);
        check(s2 === 1'b1, "reset mid-activity forces sync back to RESET_VALUE");
        check(r2 === 1'b0 && f2 === 1'b0, "reset suppresses the edge outputs");
        @(negedge clk) rst_n = 1'b1;
        repeat (4) @(negedge clk);

        // ---- whole-run invariants -------------------------------------------
        check(wide == 0,              "no edge pulse ever wider than one clock");
        check(both == 0,              "rise and fall never assert together");
        check(edge_without_move == 0, "an edge only ever follows a real sync_o move");

        $display("== %0d checks, %0d failures ==", checks, failures);
        if (failures == 0) $display("   RESULT: ALL VERILOG SYNC-EDGE TESTS PASSED");
        else               $display("   RESULT: VERILOG SYNC-EDGE TESTS FAILED");
        $finish;
    end
endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
--===========================================================================
--  tb_uart_sync_edge — self-checking VHDL-2008 testbench
--
--  WHAT A SIMULATION CAN AND CANNOT CHECK HERE, stated up front because it
--  is the whole lesson of Module 12:
--
--    CAN    — the LATENCY contract (STAGES clocks), the pulse widths, the
--             reset value, and that edges are derived from synchronised
--             history rather than from the pin.
--    CANNOT — that the first flop actually absorbs metastability. There is
--             no metastability in an RTL simulation to absorb. That property
--             is checked by CDC lint against the netlist, and a clean
--             simulation is not evidence for it.
--
--  INDEPENDENCE: every expectation comes from the stimulus the testbench
--  drove, never from reading the synchroniser chain.
--
--  Same 19 counted checks as the Verilog and SystemVerilog twins.
--===========================================================================
library ieee;
use ieee.std_logic_1164.all;

entity tb_uart_sync_edge is
end entity tb_uart_sync_edge;

architecture sim of tb_uart_sync_edge is

    constant TCLK : time := 10 ns;

    signal clk   : std_logic := '0';
    signal rst_n : std_logic := '0';
    signal done  : boolean   := false;

    signal a2 : std_logic := '1';
    signal a3 : std_logic := '1';
    signal a0 : std_logic := '0';

    signal s2, r2, f2 : std_logic;      -- STAGES=2, reset MARK
    signal s3, r3, f3 : std_logic;      -- STAGES=3, reset MARK
    signal s0, r0, f0 : std_logic;      -- STAGES=2, reset SPACE

    -- whole-run observers, computed from the OUTPUTS alone
    signal n_rise, n_fall, wide, both : natural := 0;
    signal edge_without_move          : natural := 0;

begin

    clk <= not clk after TCLK/2 when not done else '0';

    d2 : entity work.uart_sync_edge
        generic map (STAGES => 2, RESET_VALUE => '1')
        port map (clk => clk, rst_n => rst_n, async_i => a2,
                  sync_o => s2, rise_o => r2, fall_o => f2);
    d3 : entity work.uart_sync_edge
        generic map (STAGES => 3, RESET_VALUE => '1')
        port map (clk => clk, rst_n => rst_n, async_i => a3,
                  sync_o => s3, rise_o => r3, fall_o => f3);
    d0 : entity work.uart_sync_edge
        generic map (STAGES => 2, RESET_VALUE => '0')
        port map (clk => clk, rst_n => rst_n, async_i => a0,
                  sync_o => s0, rise_o => r0, fall_o => f0);

    -- One process per observed signal group: VHDL forbids two drivers on an
    -- unresolved type, so the counters cannot be shared between processes.
    obs_pulses : process (clk)
        variable r_prev, f_prev : std_logic := '0';
    begin
        if rising_edge(clk) and rst_n = '1' then
            if r2 = '1' then
                n_rise <= n_rise + 1;
                if r_prev = '1' then wide <= wide + 1; end if;
            end if;
            if f2 = '1' then
                n_fall <= n_fall + 1;
                if f_prev = '1' then wide <= wide + 1; end if;
            end if;
            if r2 = '1' and f2 = '1' then both <= both + 1; end if;
            r_prev := r2; f_prev := f2;
        end if;
    end process obs_pulses;

    -- An edge may only appear one clock after sync_o actually moved.
    obs_move : process (clk)
        variable s_prev : std_logic := '1';
    begin
        if rising_edge(clk) and rst_n = '1' then
            if (r2 = '1' or f2 = '1') and s2 = s_prev then
                edge_without_move <= edge_without_move + 1;
            end if;
            s_prev := s2;
        end if;
    end process obs_move;

    watchdog : process
    begin
        wait for 400 us;
        report "watchdog: simulation did not finish" severity failure;
    end process watchdog;

    stim : process
        variable checks, failures : natural := 0;
        variable lat, base        : natural := 0;

        procedure check(cond : boolean; name : string) is
        begin
            checks := checks + 1;
            if cond then
                report "  PASS " & name severity note;
            else
                failures := failures + 1;
                report "  FAIL " & name severity error;
            end if;
        end procedure check;
    begin
        report "== uart_sync_edge : self-checking VHDL testbench ==" severity note;

        ---- reset contract --------------------------------------------------
        rst_n <= '0'; a2 <= '1'; a3 <= '1'; a0 <= '0';
        for i in 1 to 6 loop wait until falling_edge(clk); end loop;
        check(s2 = '1', "reset: RESET_VALUE=1 holds sync at MARK");
        check(s0 = '0', "reset: RESET_VALUE=0 holds sync at SPACE");
        check(r2 = '0' and f2 = '0', "reset: no edge pulses");
        wait until falling_edge(clk); rst_n <= '1';
        for i in 1 to 4 loop wait until falling_edge(clk); end loop;
        check(r2 = '0' and f2 = '0', "leaving reset with a steady input: no edge");

        ---- THE latency contract, measured not assumed -----------------------
        -- A change on the pin must take exactly STAGES clocks to appear on
        -- sync_o. The rest of the receiver's timing budget depends on this
        -- number, which is why it is measured rather than trusted.
        wait until falling_edge(clk); a2 <= '0';
        lat := 0;
        while s2 /= '0' and lat < 10 loop
            wait until rising_edge(clk); wait for 1 ns; lat := lat + 1;
        end loop;
        check(lat = 2, "STAGES=2: a pin change takes exactly TWO clocks");

        wait until falling_edge(clk); a3 <= '0';
        lat := 0;
        while s3 /= '0' and lat < 10 loop
            wait until rising_edge(clk); wait for 1 ns; lat := lat + 1;
        end loop;
        check(lat = 3, "STAGES=3: the same change takes exactly THREE clocks");
        check(s2 = '0', "and the 2-stage instance already settled");

        ---- the edge pulses --------------------------------------------------
        base := n_fall;
        wait until falling_edge(clk); a2 <= '1';
        for i in 1 to 4 loop wait until falling_edge(clk); end loop;
        check(n_fall = base, "returning to mark raises no FALL");
        base := n_rise;
        wait until falling_edge(clk); a2 <= '0';
        for i in 1 to 4 loop wait until falling_edge(clk); end loop;
        check(n_rise = base, "going to space raises no RISE");

        ---- a start-bit edge, which is what this block exists for ------------
        wait until falling_edge(clk); a2 <= '1';
        for i in 1 to 4 loop wait until falling_edge(clk); end loop;
        base := n_fall;
        wait until falling_edge(clk); a2 <= '0';   -- falling edge = start cand.
        for i in 1 to 4 loop wait until falling_edge(clk); end loop;
        check(n_fall = base + 1, "a 1->0 pin transition yields exactly one FALL");
        check(f2 = '0',          "and the pulse has already returned low");

        ---- a one-clock pulse is PASSED, not filtered ------------------------
        -- A synchroniser aligns; it does not debounce. A pulse one clock wide
        -- arrives one clock wide, and rejecting it is the start-validation
        -- logic's job (Chapter 5.2), not this block's.
        wait until falling_edge(clk); a2 <= '1';
        for i in 1 to 4 loop wait until falling_edge(clk); end loop;
        base := n_fall;
        wait until falling_edge(clk); a2 <= '0';
        wait until falling_edge(clk); a2 <= '1';   -- one clock wide
        for i in 1 to 5 loop wait until falling_edge(clk); end loop;
        check(n_fall = base + 1, "a one-clock low pulse still produces its FALL");
        check(n_rise > 0,        "and its matching RISE -- alignment, not filtering");

        ---- back-to-back transitions -----------------------------------------
        base := n_fall;
        for k in 0 to 7 loop
            wait until falling_edge(clk); a2 <= '0';
            for i in 1 to 3 loop wait until falling_edge(clk); end loop;
            wait until falling_edge(clk); a2 <= '1';
            for i in 1 to 3 loop wait until falling_edge(clk); end loop;
        end loop;
        check(n_fall = base + 8, "eight transitions produce exactly eight FALLs");

        ---- reset during activity ---------------------------------------------
        wait until falling_edge(clk); a2 <= '0';
        wait until falling_edge(clk); rst_n <= '0';
        for i in 1 to 3 loop wait until falling_edge(clk); end loop;
        check(s2 = '1', "reset mid-activity forces sync back to RESET_VALUE");
        check(r2 = '0' and f2 = '0', "reset suppresses the edge outputs");
        wait until falling_edge(clk); rst_n <= '1';
        for i in 1 to 4 loop wait until falling_edge(clk); end loop;

        ---- whole-run invariants -----------------------------------------------
        check(wide = 0,              "no edge pulse ever wider than one clock");
        check(both = 0,              "rise and fall never assert together");
        check(edge_without_move = 0, "an edge only ever follows a real sync_o move");

        report "== " & integer'image(checks) & " checks, "
                     & integer'image(failures) & " failures ==" severity note;
        if failures = 0 then
            report "   RESULT: ALL VHDL SYNC-EDGE TESTS PASSED" severity note;
        else
            report "   RESULT: VHDL SYNC-EDGE TESTS FAILED" severity error;
        end if;
        done <= true;
        wait;
    end process stim;

end architecture sim;

All three run the same 19 checks and agree:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
== uart_sync_edge : self-checking SystemVerilog testbench ==
  PASS reset: RESET_VALUE=1 holds sync at MARK
  PASS reset: RESET_VALUE=0 holds sync at SPACE
  PASS STAGES=2: a pin change takes exactly TWO clocks
  PASS STAGES=3: the same change takes exactly THREE clocks
  PASS a 1->0 pin transition yields exactly one FALL
  PASS a one-clock low pulse still produces its FALL
  PASS and its matching RISE -- alignment, not filtering
  PASS eight transitions produce exactly eight FALLs
  PASS no edge pulse ever wider than one clock
  PASS rise and fall never assert together
  PASS an edge only ever follows a real sync_o move
== 19 checks, 0 failures ==

Verilog-2001    : 19 checks, 0 failures
SystemVerilog   : 19 checks, 0 failures
VHDL-2008       : 19 checks, 0 failures

8. Verification

Structural checks, because functional ones cannot help. The list is short and every item is mechanically checkable on the netlist:

RuleWhy
every asynchronous input has a synchroniser§1
nothing but stage 2 reads stage 1unresolved fan-out
no combinational logic between stagesit spends the settling time
exactly one synchroniser per async signal§3 — two can disagree
the synchronised signal is what feeds edge detection and filtering§1
the synchroniser source is registered, not combinationala glitch would be faithfully resolved
Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// What simulation CAN check is the consequence of the fix: every consumer
// sees the same version of the line. This is the §3 rule as an assertion, and
// it fails on the pre-fix integration.
property p_one_view_of_the_line;
    @(posedge clk) disable iff (!rst_n)
        u_break.rx_sync_i == u_rx.rx_sync_o;
endproperty
assert property (p_one_view_of_the_line);

Measured: 40 violations before the fix, 0 after — the assertion discriminates, and it is the only check in this module's arsenal that does.

Run CDC lint and treat it as a gate. It found nothing in this design because it was not run; reading the connection by hand found the defect, which is not a method that scales past one reviewer with time on their hands. The assertion above is the cheap partial substitute — it catches divergent views, not missing synchronisers.

Mutation testing: does the suite actually discriminate?

Nineteen passing checks are worth nothing until something has made them fail. Three defects were installed in uart_sync_edge, one at a time, and the suite re-run against each. The mutation was verified to have actually changed the source before the result was scored — a patch that silently matches nothing produces a "survivor" that is really just the original design, which is the easiest way to fool yourself in a mutation run.

#Defect installedResult
M1sync_o read from the first flop — a one-flop synchroniserkilled, 7 checks
M2fall_o derived from the raw pin, not the settled historykilled, 5 checks
M3RESET_VALUE ignored; the chain resets to 0killed, 4 checks

M1 is worth dwelling on. The one-flop synchroniser is the defect this whole module is about, and here it is caught — but read carefully why. It is not caught because the simulation saw metastability. It is caught because reading chain_q[0] instead of chain_q[STAGES-1] changes the latency from two clocks to one, and the testbench measures latency. The hazard stayed invisible; a side effect of the hazard was visible, and the test was written to look at it.

That is the pattern to take away, and Chapter 12.3 shows the case where it is not available: a one-flop synchroniser buried inside a FIFO's pointer crossing changes no latency any output can see, and needs a structural check instead.

9. Debugging

10. What This Means on an FPGA

Mark the synchroniser. ASYNC_REG on Xilinx and the equivalent elsewhere tells the placer these flops belong together, preserving the settling time Chapter 12.1 §3 assumed. Without it the stages can be routed apart and the margin quietly spent.

Do not let retiming or optimisation touch the pair. Synthesis will happily absorb a "redundant" flop if nothing tells it otherwise, and the attribute is what tells it.

Publish the synchronised signal, once, from one place. §3's fix is one output port. It also makes the structural rule visible at the integration level rather than buried inside a block.

The two-clock latency is real and is almost never the problem. Below roughly 50 clocks per bit it is worth putting in the budget; above that it is noise against the oversampling margin.

11. Understanding Check

12. Summary

Synchronise first, then do everything else. The pin feeds one flop; that flop feeds one flop; everything else reads the second.

The correct placement costs exactly two clocks — three clock edges from line transition to the state machine acting, against one for the raw version. That is 0.23% of a bit at 868 clocks per bit and 10% at 20, and only the ratio matters.

An asynchronous signal is synchronised once and shared. Two synchronisers can resolve the same transition differently on the same clock, leaving two consumers with contradictory views and neither of them wrong.

The IP built in Module 11 violated both rules: the break detector's port is documented as requiring a synchronised line and the top level handed it the raw pin. 86 functional checks passed, and none of them could have failed. Fixed by publishing the receiver's synchronised line and sharing it, for two clocks of extra break latency against a window of eleven bit times — 0.02%.

A deliberately wrong placement was functionally indistinguishable from the correct one across five clock ratios and forty frames. The negative result is the finding: no test of outputs distinguishes correct CDC from incorrect, so the case for these rules is physical and the gate that enforces them is a CDC linter.

But a structural rule can have an observable shadow, and this one does: asserting that all consumers see one synchronised line gave 40 violations before the fix and 0 after. Where such a shadow exists, assert on it — it is three lines and it runs in the tests you already have.

13. What Comes Next

The receive pin is handled. Chapter 12.3 asks the question this module has been circling: which parts of a UART are actually clock-domain crossings?

The answer is narrower than most people assume, and getting it wrong in either direction is expensive — treating asynchronous serial timing as CDC produces structures that solve nothing, while missing a genuine crossing produces the failures this chapter has been describing. It also confronts the case this IP has so far avoided entirely: a register interface running on a different clock from the UART itself, and the asynchronous FIFO that then becomes unavoidable.

Browse the full path on the UART tutorials index. For the guarantees this chapter's placement rules protect, read back to Chapter 12.1.

Continue learning

Where this fits

Part of the UART curriculum.