UART · Module 12
Synchroniser Placement and Edge Detection After Synchronisation
Where the synchroniser belongs, what its latency costs the start-edge measurement, and a real placement defect found in the IP built in Module 11 — which every functional test passed.
Chapter 12.1 established what a synchroniser guarantees. This chapter is about where it goes, which sounds like a detail and is the part that actually gets built wrong.
The rule is one sentence: synchronise first, then do everything else. Edge detection, glitch filtering, counting, comparison — all of it operates on the synchronised signal, never on the pin.
The rest of the chapter is about why that rule has to be stated as a rule. It ends with a placement defect in the IP assembled in Chapter 11.2, which passed 86 functional checks across four testbenches and three parameter sets, because — per Chapter 12.1 §7 — it was always going to.
1. The Rule, and the Tempting Violation
// CORRECT — this is what uart_rx has done since Chapter 6.1.
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
rx_meta_q <= 1'b1;
rx_sync_q <= 1'b1;
rx_sync_d_q <= 1'b1;
end else begin
rx_meta_q <= rx_i; // may go metastable — read by NOTHING else
rx_sync_q <= rx_meta_q; // settled
rx_sync_d_q <= rx_sync_q; // history, for edge detection
end
end
// Edge detection on SYNCHRONISED history only.
assign start_cand = rx_sync_d_q && !rx_sync_q;The tempting alternative detects the edge as early as possible, on the pin itself:
// WRONG — the start edge is taken from the RAW pin while every data sample
// still comes from the synchronised one.
logic rx_raw_d_q;
always_ff @(posedge clk or negedge rst_n)
if (!rst_n) rx_raw_d_q <= 1'b1; else rx_raw_d_q <= rx_i;
assign start_cand = rx_raw_d_q && !rx_i; // combinational on the pinIt is tempting for a real reason: it detects the start edge two clocks sooner, and start-edge latency eats directly into the receiver's timing budget. The instinct is sound and the implementation is not.
2. What the Synchroniser Costs, Measured
Dropping the line 1 ns after a clock edge at 100 MHz and timestamping each stage:
=== 100 MHz clock, 10.00 ns period. Falling edge on rx at t=286.00 ns
rx_meta_q changed at 295.00 ns (+0.90 clocks)
rx_sync_q changed at 305.00 ns (+1.90 clocks)
FSM consumes start_cand 315.00 ns (+2.90 clocks)
raw-edge FSM consumes 295.00 ns (+0.90 clocks)
raw-edge detects 2.00 clocks EARLIERThree clock edges, correct placement; one, raw. The difference is exactly two clocks — the two synchroniser stages — and it is the entire benefit the wrong version buys.
What those two clocks are worth depends only on the clocks-per-bit ratio:
| Clock / baud | clocks per bit | clocks per os tick | 2 clocks as a fraction of a bit |
|---|---|---|---|
| 100 MHz / 115,200 | 868.06 | 54.25 | 0.230% |
| 50 MHz / 460,800 | 108.51 | 6.78 | 1.84% |
| 24 MHz / 460,800 | 52.08 | 3.26 | 3.84% |
| 24 MHz / 921,600 | 26.04 | 1.63 | 7.68% |
| 20 MHz / 1,000,000 | 20.00 | 1.25 | 10.0% |
The right-hand column is what the argument is actually about. At 868 clocks per bit the synchroniser is free. At 20 clocks per bit it costs a tenth of a bit time — and the receiver's budget, built in Chapter 4.5, has to absorb it alongside everything else.
Sync-then-edge versus raw edge
8 cycles3. One Synchroniser Per Signal
A second rule, less often stated and just as binding: an asynchronous signal is synchronised once, and every consumer reads that same synchronised copy.
Two independent synchronisers on one pin each resolve independently. When a transition lands near an edge, one may resolve to 0 and the other to 1 — on the same clock, from the same wire. The two consumers then disagree about when the line changed, and the design contains a contradiction that no single block is wrong about.
// WRONG — two consumers, two synchronisers, two opinions.
uart_rx u_rx (.rx_i(rx_line), ...); // synchronises internally
uart_break_det u_break(.rx_sync_i(rx_line), ...); // ALSO sees the raw pin
// RIGHT — one synchroniser, published, shared.
uart_rx u_rx (.rx_i(rx_line), .rx_sync_o(rx_synced), ...);
uart_break_det u_break(.rx_sync_i(rx_synced), ...);4. The Defect in the Module 11 IP
The wrong version above is not hypothetical. It is what Chapter 11.2 shipped.
// uart_break_detect's port, as declared in Module 9:
input logic rx_sync_i, // synchronised line — Chapter 5.1
// what uart_ip connected to it:
uart_break_detect #(...) u_break (
.clk(clk), .rst_n(rst_n), .os_tick_i(os_tick),
.rx_sync_i(rx_line), // <-- the RAW PIN
...);The port documents its own requirement in a comment, and the integration violated it. The break detector has no synchroniser of its own — it registers rx_sync_i straight into its low-tick counter — so the assembled IP put an asynchronous input directly onto a flop, in a block whose interface says in words that it must not.
The fix gives the IP one synchroniser and shares it:
// in uart_rx — publish the line it has already synchronised
output logic rx_sync_o
...
assign rx_sync_o = rx_sync_q;
// in uart_ip — the break detector now sees the SAME resolved value the
// receiver does, two clocks after the pin, exactly like every other consumer.
uart_rx #(...) u_rx (..., .rx_active_o(rx_active), .rx_sync_o(rx_synced));
uart_break_detect #(...) u_break (
.clk(clk), .rst_n(rst_n), .os_tick_i(os_tick),
.rx_sync_i(rx_synced), ...);The cost is two clocks of extra break-detection latency against a threshold of eleven bit times — 20 ns against 95.5 µs at 100 MHz and 115,200 baud, or 0.02% of the detection window. There was never a trade-off to weigh.
Regression after the fix: 86 checks, 0 failures across all four Module 11 suites — unchanged, as expected.
5. The Experiment That Did Not Work
Wanting a measurable consequence rather than only a structural argument, I built the raw-edge receiver of §1 and ran it beside the correct one at five clocks-per-bit ratios, from 868 down to 20, on eight byte patterns each.
Both received all eight bytes correctly at every ratio.
-- 100 MHz / 115200 baud : 868.06 clk/bit, 54.253 clk per os tick
sync-then-edge (correct) : 8/8 delivered, 8 correct
raw-edge (wrong) : 8/8 delivered, 8 correct
-- 20 MHz / 1000000 baud : 20.00 clk/bit, 1.250 clk per os tick
sync-then-edge (correct) : 8/8 delivered, 8 correct
raw-edge (wrong) : 8/8 delivered, 8 correctThe hypothesis was wrong. I expected the two-clock head start to shift the sampling grid enough to erode margin at low clocks-per-tick ratios. It does shift the grid — §2 measured exactly two clocks — but the receiver re-centres on the start edge it detected, so the whole grid moves together and the samples still land near bit centres. A 1.6-tick shift against roughly ±8 ticks of margin is simply absorbed.
6. The Synchroniser As a Named Block
Everything above argues that one structure is correct and a very similar one is
not. Section 4 then showed an integration that got it wrong — and got it wrong
precisely because the correct structure existed only as lines typed inside
uart_rx, not as a thing with a name. A block that is re-typed is a block that
can be re-typed incorrectly, and no reviewer can diff a pattern against a
pattern that was never written down.
So write it down. uart_sync_edge is the §1 rule with a port list: STAGES
flops of synchroniser, edge detection taken from the settled history, and a
reset value chosen so that leaving reset does not look like a start bit.
// ===========================================================================
// uart_sync_edge — Synthesizable SystemVerilog
//
// The block Module 12 is about, published as a module rather than as a
// fragment inside the receiver. Two jobs, and the order is the whole point:
//
// 1. Bring an ASYNCHRONOUS input into this clock domain through STAGES
// flops. The first flop may go metastable; nothing but the second flop
// is allowed to read it.
// 2. Detect edges on the SYNCHRONISED history only — never on the pin.
//
// Chapter 12.2 shows what edge-detecting the raw pin costs. This module
// exists so that the correct structure has one name, one owner and one
// testbench instead of being re-typed inside every consumer.
// ===========================================================================
module uart_sync_edge #(
parameter int unsigned STAGES = 2,
// The value the chain holds during reset. For a UART receive line this
// is MARK: coming out of reset believing the line is low would look
// exactly like a start bit.
parameter logic RESET_VALUE = 1'b1
) (
input logic clk,
input logic rst_n,
input logic async_i, // genuinely asynchronous — the pin
output logic sync_o, // settled, safe for the whole domain to read
output logic rise_o, // ONE cycle, on a synchronised 0 -> 1
output logic fall_o // ONE cycle, on a synchronised 1 -> 0
);
if (STAGES < 2) begin : g_bad_stages
$error("uart_sync_edge: STAGES must be at least 2");
end
// chain_q[0] is the metastability-absorbing flop. It is READ BY NOTHING
// except chain_q[1]; that restriction is the synchroniser.
logic [STAGES-1:0] chain_q;
logic sync_d_q; // one-cycle history, for the edges
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
chain_q <= {STAGES{RESET_VALUE}};
sync_d_q <= RESET_VALUE;
end else begin
chain_q <= {chain_q[STAGES-2:0], async_i};
sync_d_q <= chain_q[STAGES-1];
end
end
assign sync_o = chain_q[STAGES-1];
// Edges from SYNCHRONISED history only. Both are one clock wide because
// they compare two adjacent samples of the same settled signal.
assign rise_o = ~sync_d_q & sync_o;
assign fall_o = sync_d_q & ~sync_o;
endmoduleThe Verilog-2001 version is the same design with the two facilities that
version of the language does not have written out: no logic, and no
elaboration-time $error, so the parameter check becomes a runtime one in an
initial block. It is a weaker guard — it fires when the simulation starts
rather than when the design elaborates — and it is the strongest guard
Verilog-2001 offers.
//===========================================================================
// uart_sync_edge_v — Synthesizable Verilog-2001
//
// Two jobs, and the order is the whole point:
// 1. Bring an ASYNCHRONOUS input into this clock domain through STAGES
// flops. The first flop may go metastable; nothing but the second is
// allowed to read it.
// 2. Detect edges on the SYNCHRONISED history only -- never on the pin.
//
// Verilog-2001 has no elaboration-time assertion, so the STAGES >= 2 rule
// is a documented contract here rather than an enforced one. Chapter 11.4
// section 4 covers that gap.
//===========================================================================
module uart_sync_edge_v #(
parameter STAGES = 2, // MUST be >= 2 -- not enforceable in 2001
// The value the chain holds during reset. For a UART receive line this
// is MARK: leaving reset believing the line is low looks exactly like a
// start bit.
parameter RESET_VALUE = 1'b1
) (
input wire clk,
input wire rst_n,
input wire async_i, // genuinely asynchronous -- the pin
output wire sync_o, // settled, safe for the whole domain to read
output wire rise_o, // ONE cycle, on a synchronised 0 -> 1
output wire fall_o // ONE cycle, on a synchronised 1 -> 0
);
// chain_q[0] is the metastability-absorbing flop. It is READ BY NOTHING
// except chain_q[1]; that restriction IS the synchroniser.
reg [STAGES-1:0] chain_q;
reg sync_d_q;
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
chain_q <= {STAGES{RESET_VALUE}};
sync_d_q <= RESET_VALUE;
end else begin
chain_q <= {chain_q[STAGES-2:0], async_i};
sync_d_q <= chain_q[STAGES-1];
end
end
assign sync_o = chain_q[STAGES-1];
// Edges from SYNCHRONISED history only. Both are one clock wide because
// they compare two adjacent samples of the same settled signal.
assign rise_o = ~sync_d_q & sync_o;
assign fall_o = sync_d_q & ~sync_o;
endmoduleVHDL expresses the parameter check best of the three: a concurrent assert
with severity failure inside the architecture, checked at elaboration,
needing no process and no simulation time.
--==========================================================================
-- uart_sync_edge -- Synthesizable VHDL-2008
--
-- Two jobs, and the order is the whole point:
-- 1. Bring an ASYNCHRONOUS input into this clock domain through STAGES
-- flops. The first flop may go metastable; nothing but the second is
-- allowed to read it.
-- 2. Detect edges on the SYNCHRONISED history only -- never on the pin.
--==========================================================================
library ieee;
use ieee.std_logic_1164.all;
entity uart_sync_edge is
generic (
STAGES : positive := 2;
-- The value the chain holds during reset. For a UART receive line
-- this is MARK: leaving reset believing the line is low looks
-- exactly like a start bit.
RESET_VALUE : std_logic := '1'
);
port (
clk : in std_logic;
rst_n : in std_logic;
async_i : in std_logic; -- genuinely asynchronous -- the pin
sync_o : out std_logic; -- settled, safe for the domain to read
rise_o : out std_logic; -- ONE cycle, on a synchronised 0 -> 1
fall_o : out std_logic -- ONE cycle, on a synchronised 1 -> 0
);
end entity uart_sync_edge;
architecture rtl of uart_sync_edge is
-- chain(0) is the metastability-absorbing flop. It is READ BY NOTHING
-- except chain(1); that restriction IS the synchroniser.
signal chain_q : std_logic_vector(STAGES-1 downto 0);
signal sync_d_q : std_logic;
signal sync_s : std_logic;
begin
assert STAGES >= 2
report "uart_sync_edge: STAGES must be at least 2" severity failure;
sync_s <= chain_q(STAGES-1);
sync_o <= sync_s;
-- Edges from SYNCHRONISED history only. Both are one clock wide because
-- they compare two adjacent samples of the same settled signal.
rise_o <= (not sync_d_q) and sync_s;
fall_o <= sync_d_q and (not sync_s);
process (clk, rst_n)
begin
if rst_n = '0' then
chain_q <= (others => RESET_VALUE);
sync_d_q <= RESET_VALUE;
elsif rising_edge(clk) then
chain_q <= chain_q(STAGES-2 downto 0) & async_i;
sync_d_q <= chain_q(STAGES-1);
end if;
end process;
end architecture rtl;7. Three Testbenches That Say What They Cannot Prove
A testbench for a synchroniser has an unusual obligation: it has to be explicit about the fact that it cannot test the thing the block exists for. There is no metastability in an RTL simulation, so there is nothing for the first flop to absorb, and a design with the flop and a design without it produce the same waveform. Section 5 demonstrated exactly that over forty frames.
What the testbench can pin down is everything else, and everything else is worth pinning down:
| Property | How it is checked | Kills |
|---|---|---|
| a pin change appears after exactly STAGES clocks | measured with a counting loop, not assumed | M1 |
| the same change at STAGES=3 takes three | a second instance, same stimulus | M1 |
RESET_VALUE holds the chain during reset | two instances, reset MARK and reset SPACE | M3 |
| edge pulses are exactly one clock wide | a whole-run observer counting wide pulses | — |
rise_o and fall_o never assert together | whole-run observer | — |
an edge only ever follows a real sync_o move | whole-run observer comparing against history | M2 |
| a one-clock pulse still produces its edges | explicit test — a synchroniser is not a filter | — |
That last row is the one that surprises people. A synchroniser aligns; it
does not debounce. A one-clock glitch on the pin arrives as a one-clock
glitch on sync_o, and rejecting it is the start-bit validation logic's job
from Chapter 5.2, not this block's.
A testbench that did not check this would leave room for someone to "improve"
the synchroniser into a filter and break start detection in a way that only
shows up on a noisy line.
//===========================================================================
// tb_uart_sync_edge — self-checking SystemVerilog testbench
//
// WHAT A SIMULATION CAN AND CANNOT CHECK HERE, stated up front because it
// is the whole lesson of Module 12:
//
// CAN — the LATENCY contract (STAGES clocks), the pulse widths, the
// reset value, and that edges are derived from synchronised
// history rather than from the pin.
// CANNOT — that the first flop actually absorbs metastability. There is
// no metastability in an RTL simulation to absorb. That property
// is checked by CDC lint against the netlist, and a clean
// simulation is not evidence for it.
//
// This file carries the same 19 counted checks as its Verilog twin, plus a
// layer of SystemVerilog immediate assertions that fire the instant an
// invariant breaks rather than at the end of the run. The assertions do not
// replace the counted checks — an assertion that never fires proves nothing
// unless the stimulus reached the situation it guards.
//===========================================================================
`timescale 1ns/1ps
module tb_uart_sync_edge;
logic clk = 1'b0, rst_n = 1'b0;
always #5 clk = ~clk;
logic a2 = 1'b1, a3 = 1'b1, a0 = 1'b0;
logic s2, r2, f2; // STAGES=2, reset MARK
logic s3, r3, f3; // STAGES=3, reset MARK
logic s0, r0, f0; // STAGES=2, reset SPACE
uart_sync_edge #(.STAGES(2), .RESET_VALUE(1'b1)) d2 (
.clk(clk), .rst_n(rst_n), .async_i(a2),
.sync_o(s2), .rise_o(r2), .fall_o(f2));
uart_sync_edge #(.STAGES(3), .RESET_VALUE(1'b1)) d3 (
.clk(clk), .rst_n(rst_n), .async_i(a3),
.sync_o(s3), .rise_o(r3), .fall_o(f3));
uart_sync_edge #(.STAGES(2), .RESET_VALUE(1'b0)) d0 (
.clk(clk), .rst_n(rst_n), .async_i(a0),
.sync_o(s0), .rise_o(r0), .fall_o(f0));
int checks = 0, failures = 0;
task automatic check(input logic cond, input string name);
checks++;
if (cond) $display(" PASS %0s", name);
else begin failures++; $display(" FAIL %0s", name); end
endtask
// Whole-run observers, computed from the OUTPUTS alone.
int n_rise = 0, n_fall = 0, wide = 0, both = 0;
logic r_prev = 1'b0, f_prev = 1'b0;
// Plain `always`, not `always_ff`: this is an observer, not hardware, and
// it deliberately contains a $error call that no synthesiser would accept.
always @(posedge clk) if (rst_n) begin
if (r2) begin n_rise++; if (r_prev) wide++; end
if (f2) begin n_fall++; if (f_prev) wide++; end
if (r2 && f2) both++;
// SystemVerilog immediate assertion: rise and fall are mutually
// exclusive by construction, so a failure here is a design defect,
// not a stimulus one.
a_excl: assert (!(r2 && f2))
else $error("rise_o and fall_o asserted in the same clock");
r_prev <= r2; f_prev <= f2;
end
// An edge may only appear one clock after sync_o actually moved.
int edge_without_move = 0;
logic s_prev = 1'b1;
always @(posedge clk) if (rst_n) begin
if ((r2 || f2) && (s2 === s_prev)) edge_without_move++;
a_move: assert (!((r2 || f2) && (s2 === s_prev)))
else $error("an edge pulse without a sync_o transition");
s_prev <= s2;
end
int lat, i, base;
initial begin
#400_000;
$display(" FAIL watchdog: simulation did not finish");
$display("== %0d checks, %0d failures ==", checks+1, failures+1);
$display(" RESULT: SYSTEMVERILOG SYNC-EDGE TESTS FAILED (timeout)");
$finish;
end
initial begin
$display("== uart_sync_edge : self-checking SystemVerilog testbench ==");
// ---- reset contract ------------------------------------------------
rst_n = 1'b0; a2 = 1'b1; a3 = 1'b1; a0 = 1'b0;
repeat (6) @(negedge clk);
check(s2 === 1'b1, "reset: RESET_VALUE=1 holds sync at MARK");
check(s0 === 1'b0, "reset: RESET_VALUE=0 holds sync at SPACE");
check(r2 === 1'b0 && f2 === 1'b0, "reset: no edge pulses");
@(negedge clk) rst_n = 1'b1;
repeat (4) @(negedge clk);
check(r2 === 1'b0 && f2 === 1'b0, "leaving reset with a steady input: no edge");
// ---- THE latency contract, measured not assumed --------------------
@(negedge clk) a2 = 1'b0;
lat = 0;
while (s2 !== 1'b0 && lat < 10) begin @(posedge clk); #1; lat++; end
check(lat == 2, "STAGES=2: a pin change takes exactly TWO clocks");
@(negedge clk) a3 = 1'b0;
lat = 0;
while (s3 !== 1'b0 && lat < 10) begin @(posedge clk); #1; lat++; end
check(lat == 3, "STAGES=3: the same change takes exactly THREE clocks");
check(s2 === 1'b0, "and the 2-stage instance already settled");
// ---- the edge pulses ------------------------------------------------
base = n_fall;
@(negedge clk) a2 = 1'b1;
repeat (4) @(negedge clk);
check(n_fall == base, "returning to mark raises no FALL");
base = n_rise;
@(negedge clk) a2 = 1'b0;
repeat (4) @(negedge clk);
check(n_rise == base, "going to space raises no RISE");
// ---- a start-bit edge, which is what this block exists for ---------
@(negedge clk) a2 = 1'b1;
repeat (4) @(negedge clk);
base = n_fall;
@(negedge clk) a2 = 1'b0;
repeat (4) @(negedge clk);
check(n_fall == base + 1, "a 1->0 pin transition yields exactly one FALL");
check(f2 === 1'b0, "and the pulse has already returned low");
// ---- a one-clock pulse is PASSED, not filtered ---------------------
@(negedge clk) a2 = 1'b1;
repeat (4) @(negedge clk);
base = n_fall;
@(negedge clk) a2 = 1'b0;
@(negedge clk) a2 = 1'b1; // one clock wide
repeat (5) @(negedge clk);
check(n_fall == base + 1, "a one-clock low pulse still produces its FALL");
check(n_rise > 0, "and its matching RISE -- alignment, not filtering");
// ---- back-to-back transitions ---------------------------------------
base = n_fall;
for (i = 0; i < 8; i++) begin
@(negedge clk) a2 = 1'b0;
repeat (3) @(negedge clk);
@(negedge clk) a2 = 1'b1;
repeat (3) @(negedge clk);
end
check(n_fall == base + 8, "eight transitions produce exactly eight FALLs");
// ---- reset during activity ------------------------------------------
@(negedge clk) a2 = 1'b0;
@(negedge clk) rst_n = 1'b0;
repeat (3) @(negedge clk);
check(s2 === 1'b1, "reset mid-activity forces sync back to RESET_VALUE");
check(r2 === 1'b0 && f2 === 1'b0, "reset suppresses the edge outputs");
@(negedge clk) rst_n = 1'b1;
repeat (4) @(negedge clk);
// ---- whole-run invariants -------------------------------------------
check(wide == 0, "no edge pulse ever wider than one clock");
check(both == 0, "rise and fall never assert together");
check(edge_without_move == 0, "an edge only ever follows a real sync_o move");
$display("== %0d checks, %0d failures ==", checks, failures);
if (failures == 0) $display(" RESULT: ALL SYSTEMVERILOG SYNC-EDGE TESTS PASSED");
else $display(" RESULT: SYSTEMVERILOG SYNC-EDGE TESTS FAILED");
$finish;
end
endmodule//===========================================================================
// tb_uart_sync_edge_v — self-checking Verilog-2001 testbench
//
// WHAT A SIMULATION CAN AND CANNOT CHECK HERE, stated up front because it
// is the whole lesson of Module 12:
//
// CAN -- the LATENCY contract (STAGES clocks), the pulse widths, the
// reset value, and that edges are derived from synchronised
// history rather than from the pin.
// CANNOT -- that the first flop actually absorbs metastability. There is
// no metastability in an RTL simulation to absorb. That property
// is checked by CDC lint against the netlist, and a clean
// simulation is not evidence for it.
//
// INDEPENDENCE: every expectation comes from the stimulus the testbench
// drove, never from reading chain_q.
//===========================================================================
`timescale 1ns/1ps
module tb_uart_sync_edge_v;
reg clk = 1'b0, rst_n = 1'b0;
always #5 clk = ~clk;
reg a2 = 1'b1, a3 = 1'b1, a0 = 1'b0;
wire s2, r2, f2; // STAGES=2, reset MARK
wire s3, r3, f3; // STAGES=3, reset MARK
wire s0, r0, f0; // STAGES=2, reset SPACE
uart_sync_edge_v #(.STAGES(2), .RESET_VALUE(1'b1)) d2 (
.clk(clk), .rst_n(rst_n), .async_i(a2),
.sync_o(s2), .rise_o(r2), .fall_o(f2));
uart_sync_edge_v #(.STAGES(3), .RESET_VALUE(1'b1)) d3 (
.clk(clk), .rst_n(rst_n), .async_i(a3),
.sync_o(s3), .rise_o(r3), .fall_o(f3));
uart_sync_edge_v #(.STAGES(2), .RESET_VALUE(1'b0)) d0 (
.clk(clk), .rst_n(rst_n), .async_i(a0),
.sync_o(s0), .rise_o(r0), .fall_o(f0));
integer checks = 0, failures = 0;
task check;
input cond;
input [8*80-1:0] name;
begin
checks = checks + 1;
if (cond) $display(" PASS %0s", name);
else begin failures = failures + 1; $display(" FAIL %0s", name); end
end
endtask
// Whole-run observers, computed from the OUTPUTS alone.
integer n_rise = 0, n_fall = 0, wide = 0, both = 0;
reg r_prev = 1'b0, f_prev = 1'b0;
always @(posedge clk) if (rst_n) begin
if (r2) begin n_rise = n_rise + 1; if (r_prev) wide = wide + 1; end
if (f2) begin n_fall = n_fall + 1; if (f_prev) wide = wide + 1; end
if (r2 && f2) both = both + 1; // must never both fire
r_prev = r2; f_prev = f2;
end
// An edge may only appear one clock after sync_o actually moved.
integer edge_without_move = 0;
reg s_prev = 1'b1;
always @(posedge clk) if (rst_n) begin
if ((r2 || f2) && (s2 === s_prev)) edge_without_move = edge_without_move + 1;
s_prev = s2;
end
integer lat, i, base;
initial begin
#400_000;
$display(" FAIL watchdog: simulation did not finish");
$display("== %0d checks, %0d failures ==", checks+1, failures+1);
$display(" RESULT: VERILOG SYNC-EDGE TESTS FAILED (timeout)");
$finish;
end
initial begin
$display("== uart_sync_edge_v : self-checking Verilog testbench ==");
// ---- reset contract ------------------------------------------------
rst_n = 1'b0; a2 = 1'b1; a3 = 1'b1; a0 = 1'b0;
repeat (6) @(negedge clk);
check(s2 === 1'b1, "reset: RESET_VALUE=1 holds sync at MARK");
check(s0 === 1'b0, "reset: RESET_VALUE=0 holds sync at SPACE");
check(r2 === 1'b0 && f2 === 1'b0, "reset: no edge pulses");
@(negedge clk) rst_n = 1'b1;
repeat (4) @(negedge clk);
check(r2 === 1'b0 && f2 === 1'b0, "leaving reset with a steady input: no edge");
// ---- THE latency contract, measured not assumed --------------------
// A change on the pin must take exactly STAGES clocks to appear on
// sync_o. The rest of the receiver's timing budget depends on this
// number, which is why it is measured rather than trusted.
@(negedge clk) a2 = 1'b0;
lat = 0;
while (s2 !== 1'b0 && lat < 10) begin
@(posedge clk); #1; lat = lat + 1;
end
check(lat == 2, "STAGES=2: a pin change takes exactly TWO clocks");
@(negedge clk) a3 = 1'b0;
lat = 0;
while (s3 !== 1'b0 && lat < 10) begin
@(posedge clk); #1; lat = lat + 1;
end
check(lat == 3, "STAGES=3: the same change takes exactly THREE clocks");
check(s2 === 1'b0, "and the 2-stage instance already settled");
// ---- the edge pulses ------------------------------------------------
base = n_fall;
@(negedge clk) a2 = 1'b1;
repeat (4) @(negedge clk);
check(n_fall == base, "returning to mark raises no FALL");
base = n_rise;
@(negedge clk) a2 = 1'b0;
repeat (4) @(negedge clk);
check(n_rise == base, "going to space raises no RISE");
// ---- a start-bit edge, which is what this block exists for ---------
@(negedge clk) a2 = 1'b1;
repeat (4) @(negedge clk);
base = n_fall;
@(negedge clk) a2 = 1'b0; // the falling edge = start candidate
repeat (4) @(negedge clk);
check(n_fall == base + 1, "a 1->0 pin transition yields exactly one FALL");
check(f2 === 1'b0, "and the pulse has already returned low");
// ---- a one-clock pulse is PASSED, not filtered ---------------------
// A synchroniser aligns; it does not debounce. A pulse one clock wide
// arrives one clock wide, and rejecting it is the start-validation
// logic's job (Chapter 5.2), not this block's.
@(negedge clk) a2 = 1'b1;
repeat (4) @(negedge clk);
base = n_fall;
@(negedge clk) a2 = 1'b0;
@(negedge clk) a2 = 1'b1; // one clock wide
repeat (5) @(negedge clk);
check(n_fall == base + 1, "a one-clock low pulse still produces its FALL");
check(n_rise > 0, "and its matching RISE -- alignment, not filtering");
// ---- back-to-back transitions ---------------------------------------
base = n_fall;
for (i = 0; i < 8; i = i + 1) begin
@(negedge clk) a2 = 1'b0;
repeat (3) @(negedge clk);
@(negedge clk) a2 = 1'b1;
repeat (3) @(negedge clk);
end
check(n_fall == base + 8, "eight transitions produce exactly eight FALLs");
// ---- reset during activity ------------------------------------------
@(negedge clk) a2 = 1'b0;
@(negedge clk) rst_n = 1'b0;
repeat (3) @(negedge clk);
check(s2 === 1'b1, "reset mid-activity forces sync back to RESET_VALUE");
check(r2 === 1'b0 && f2 === 1'b0, "reset suppresses the edge outputs");
@(negedge clk) rst_n = 1'b1;
repeat (4) @(negedge clk);
// ---- whole-run invariants -------------------------------------------
check(wide == 0, "no edge pulse ever wider than one clock");
check(both == 0, "rise and fall never assert together");
check(edge_without_move == 0, "an edge only ever follows a real sync_o move");
$display("== %0d checks, %0d failures ==", checks, failures);
if (failures == 0) $display(" RESULT: ALL VERILOG SYNC-EDGE TESTS PASSED");
else $display(" RESULT: VERILOG SYNC-EDGE TESTS FAILED");
$finish;
end
endmodule--===========================================================================
-- tb_uart_sync_edge — self-checking VHDL-2008 testbench
--
-- WHAT A SIMULATION CAN AND CANNOT CHECK HERE, stated up front because it
-- is the whole lesson of Module 12:
--
-- CAN — the LATENCY contract (STAGES clocks), the pulse widths, the
-- reset value, and that edges are derived from synchronised
-- history rather than from the pin.
-- CANNOT — that the first flop actually absorbs metastability. There is
-- no metastability in an RTL simulation to absorb. That property
-- is checked by CDC lint against the netlist, and a clean
-- simulation is not evidence for it.
--
-- INDEPENDENCE: every expectation comes from the stimulus the testbench
-- drove, never from reading the synchroniser chain.
--
-- Same 19 counted checks as the Verilog and SystemVerilog twins.
--===========================================================================
library ieee;
use ieee.std_logic_1164.all;
entity tb_uart_sync_edge is
end entity tb_uart_sync_edge;
architecture sim of tb_uart_sync_edge is
constant TCLK : time := 10 ns;
signal clk : std_logic := '0';
signal rst_n : std_logic := '0';
signal done : boolean := false;
signal a2 : std_logic := '1';
signal a3 : std_logic := '1';
signal a0 : std_logic := '0';
signal s2, r2, f2 : std_logic; -- STAGES=2, reset MARK
signal s3, r3, f3 : std_logic; -- STAGES=3, reset MARK
signal s0, r0, f0 : std_logic; -- STAGES=2, reset SPACE
-- whole-run observers, computed from the OUTPUTS alone
signal n_rise, n_fall, wide, both : natural := 0;
signal edge_without_move : natural := 0;
begin
clk <= not clk after TCLK/2 when not done else '0';
d2 : entity work.uart_sync_edge
generic map (STAGES => 2, RESET_VALUE => '1')
port map (clk => clk, rst_n => rst_n, async_i => a2,
sync_o => s2, rise_o => r2, fall_o => f2);
d3 : entity work.uart_sync_edge
generic map (STAGES => 3, RESET_VALUE => '1')
port map (clk => clk, rst_n => rst_n, async_i => a3,
sync_o => s3, rise_o => r3, fall_o => f3);
d0 : entity work.uart_sync_edge
generic map (STAGES => 2, RESET_VALUE => '0')
port map (clk => clk, rst_n => rst_n, async_i => a0,
sync_o => s0, rise_o => r0, fall_o => f0);
-- One process per observed signal group: VHDL forbids two drivers on an
-- unresolved type, so the counters cannot be shared between processes.
obs_pulses : process (clk)
variable r_prev, f_prev : std_logic := '0';
begin
if rising_edge(clk) and rst_n = '1' then
if r2 = '1' then
n_rise <= n_rise + 1;
if r_prev = '1' then wide <= wide + 1; end if;
end if;
if f2 = '1' then
n_fall <= n_fall + 1;
if f_prev = '1' then wide <= wide + 1; end if;
end if;
if r2 = '1' and f2 = '1' then both <= both + 1; end if;
r_prev := r2; f_prev := f2;
end if;
end process obs_pulses;
-- An edge may only appear one clock after sync_o actually moved.
obs_move : process (clk)
variable s_prev : std_logic := '1';
begin
if rising_edge(clk) and rst_n = '1' then
if (r2 = '1' or f2 = '1') and s2 = s_prev then
edge_without_move <= edge_without_move + 1;
end if;
s_prev := s2;
end if;
end process obs_move;
watchdog : process
begin
wait for 400 us;
report "watchdog: simulation did not finish" severity failure;
end process watchdog;
stim : process
variable checks, failures : natural := 0;
variable lat, base : natural := 0;
procedure check(cond : boolean; name : string) is
begin
checks := checks + 1;
if cond then
report " PASS " & name severity note;
else
failures := failures + 1;
report " FAIL " & name severity error;
end if;
end procedure check;
begin
report "== uart_sync_edge : self-checking VHDL testbench ==" severity note;
---- reset contract --------------------------------------------------
rst_n <= '0'; a2 <= '1'; a3 <= '1'; a0 <= '0';
for i in 1 to 6 loop wait until falling_edge(clk); end loop;
check(s2 = '1', "reset: RESET_VALUE=1 holds sync at MARK");
check(s0 = '0', "reset: RESET_VALUE=0 holds sync at SPACE");
check(r2 = '0' and f2 = '0', "reset: no edge pulses");
wait until falling_edge(clk); rst_n <= '1';
for i in 1 to 4 loop wait until falling_edge(clk); end loop;
check(r2 = '0' and f2 = '0', "leaving reset with a steady input: no edge");
---- THE latency contract, measured not assumed -----------------------
-- A change on the pin must take exactly STAGES clocks to appear on
-- sync_o. The rest of the receiver's timing budget depends on this
-- number, which is why it is measured rather than trusted.
wait until falling_edge(clk); a2 <= '0';
lat := 0;
while s2 /= '0' and lat < 10 loop
wait until rising_edge(clk); wait for 1 ns; lat := lat + 1;
end loop;
check(lat = 2, "STAGES=2: a pin change takes exactly TWO clocks");
wait until falling_edge(clk); a3 <= '0';
lat := 0;
while s3 /= '0' and lat < 10 loop
wait until rising_edge(clk); wait for 1 ns; lat := lat + 1;
end loop;
check(lat = 3, "STAGES=3: the same change takes exactly THREE clocks");
check(s2 = '0', "and the 2-stage instance already settled");
---- the edge pulses --------------------------------------------------
base := n_fall;
wait until falling_edge(clk); a2 <= '1';
for i in 1 to 4 loop wait until falling_edge(clk); end loop;
check(n_fall = base, "returning to mark raises no FALL");
base := n_rise;
wait until falling_edge(clk); a2 <= '0';
for i in 1 to 4 loop wait until falling_edge(clk); end loop;
check(n_rise = base, "going to space raises no RISE");
---- a start-bit edge, which is what this block exists for ------------
wait until falling_edge(clk); a2 <= '1';
for i in 1 to 4 loop wait until falling_edge(clk); end loop;
base := n_fall;
wait until falling_edge(clk); a2 <= '0'; -- falling edge = start cand.
for i in 1 to 4 loop wait until falling_edge(clk); end loop;
check(n_fall = base + 1, "a 1->0 pin transition yields exactly one FALL");
check(f2 = '0', "and the pulse has already returned low");
---- a one-clock pulse is PASSED, not filtered ------------------------
-- A synchroniser aligns; it does not debounce. A pulse one clock wide
-- arrives one clock wide, and rejecting it is the start-validation
-- logic's job (Chapter 5.2), not this block's.
wait until falling_edge(clk); a2 <= '1';
for i in 1 to 4 loop wait until falling_edge(clk); end loop;
base := n_fall;
wait until falling_edge(clk); a2 <= '0';
wait until falling_edge(clk); a2 <= '1'; -- one clock wide
for i in 1 to 5 loop wait until falling_edge(clk); end loop;
check(n_fall = base + 1, "a one-clock low pulse still produces its FALL");
check(n_rise > 0, "and its matching RISE -- alignment, not filtering");
---- back-to-back transitions -----------------------------------------
base := n_fall;
for k in 0 to 7 loop
wait until falling_edge(clk); a2 <= '0';
for i in 1 to 3 loop wait until falling_edge(clk); end loop;
wait until falling_edge(clk); a2 <= '1';
for i in 1 to 3 loop wait until falling_edge(clk); end loop;
end loop;
check(n_fall = base + 8, "eight transitions produce exactly eight FALLs");
---- reset during activity ---------------------------------------------
wait until falling_edge(clk); a2 <= '0';
wait until falling_edge(clk); rst_n <= '0';
for i in 1 to 3 loop wait until falling_edge(clk); end loop;
check(s2 = '1', "reset mid-activity forces sync back to RESET_VALUE");
check(r2 = '0' and f2 = '0', "reset suppresses the edge outputs");
wait until falling_edge(clk); rst_n <= '1';
for i in 1 to 4 loop wait until falling_edge(clk); end loop;
---- whole-run invariants -----------------------------------------------
check(wide = 0, "no edge pulse ever wider than one clock");
check(both = 0, "rise and fall never assert together");
check(edge_without_move = 0, "an edge only ever follows a real sync_o move");
report "== " & integer'image(checks) & " checks, "
& integer'image(failures) & " failures ==" severity note;
if failures = 0 then
report " RESULT: ALL VHDL SYNC-EDGE TESTS PASSED" severity note;
else
report " RESULT: VHDL SYNC-EDGE TESTS FAILED" severity error;
end if;
done <= true;
wait;
end process stim;
end architecture sim;All three run the same 19 checks and agree:
== uart_sync_edge : self-checking SystemVerilog testbench ==
PASS reset: RESET_VALUE=1 holds sync at MARK
PASS reset: RESET_VALUE=0 holds sync at SPACE
PASS STAGES=2: a pin change takes exactly TWO clocks
PASS STAGES=3: the same change takes exactly THREE clocks
PASS a 1->0 pin transition yields exactly one FALL
PASS a one-clock low pulse still produces its FALL
PASS and its matching RISE -- alignment, not filtering
PASS eight transitions produce exactly eight FALLs
PASS no edge pulse ever wider than one clock
PASS rise and fall never assert together
PASS an edge only ever follows a real sync_o move
== 19 checks, 0 failures ==
Verilog-2001 : 19 checks, 0 failures
SystemVerilog : 19 checks, 0 failures
VHDL-2008 : 19 checks, 0 failures8. Verification
Structural checks, because functional ones cannot help. The list is short and every item is mechanically checkable on the netlist:
| Rule | Why |
|---|---|
| every asynchronous input has a synchroniser | §1 |
| nothing but stage 2 reads stage 1 | unresolved fan-out |
| no combinational logic between stages | it spends the settling time |
| exactly one synchroniser per async signal | §3 — two can disagree |
| the synchronised signal is what feeds edge detection and filtering | §1 |
| the synchroniser source is registered, not combinational | a glitch would be faithfully resolved |
// What simulation CAN check is the consequence of the fix: every consumer
// sees the same version of the line. This is the §3 rule as an assertion, and
// it fails on the pre-fix integration.
property p_one_view_of_the_line;
@(posedge clk) disable iff (!rst_n)
u_break.rx_sync_i == u_rx.rx_sync_o;
endproperty
assert property (p_one_view_of_the_line);Measured: 40 violations before the fix, 0 after — the assertion discriminates, and it is the only check in this module's arsenal that does.
Run CDC lint and treat it as a gate. It found nothing in this design because it was not run; reading the connection by hand found the defect, which is not a method that scales past one reviewer with time on their hands. The assertion above is the cheap partial substitute — it catches divergent views, not missing synchronisers.
Mutation testing: does the suite actually discriminate?
Nineteen passing checks are worth nothing until something has made them fail.
Three defects were installed in uart_sync_edge, one at a time, and the suite
re-run against each. The mutation was verified to have actually changed the
source before the result was scored — a patch that silently matches nothing
produces a "survivor" that is really just the original design, which is the
easiest way to fool yourself in a mutation run.
| # | Defect installed | Result |
|---|---|---|
| M1 | sync_o read from the first flop — a one-flop synchroniser | killed, 7 checks |
| M2 | fall_o derived from the raw pin, not the settled history | killed, 5 checks |
| M3 | RESET_VALUE ignored; the chain resets to 0 | killed, 4 checks |
M1 is worth dwelling on. The one-flop synchroniser is the defect this whole
module is about, and here it is caught — but read carefully why. It is
not caught because the simulation saw metastability. It is caught because
reading chain_q[0] instead of chain_q[STAGES-1] changes the latency from
two clocks to one, and the testbench measures latency. The hazard stayed
invisible; a side effect of the hazard was visible, and the test was written
to look at it.
That is the pattern to take away, and Chapter 12.3 shows the case where it is not available: a one-flop synchroniser buried inside a FIFO's pointer crossing changes no latency any output can see, and needs a structural check instead.
9. Debugging
10. What This Means on an FPGA
Mark the synchroniser. ASYNC_REG on Xilinx and the equivalent elsewhere tells the placer these flops belong together, preserving the settling time Chapter 12.1 §3 assumed. Without it the stages can be routed apart and the margin quietly spent.
Do not let retiming or optimisation touch the pair. Synthesis will happily absorb a "redundant" flop if nothing tells it otherwise, and the attribute is what tells it.
Publish the synchronised signal, once, from one place. §3's fix is one output port. It also makes the structural rule visible at the integration level rather than buried inside a block.
The two-clock latency is real and is almost never the problem. Below roughly 50 clocks per bit it is worth putting in the budget; above that it is noise against the oversampling margin.
11. Understanding Check
12. Summary
Synchronise first, then do everything else. The pin feeds one flop; that flop feeds one flop; everything else reads the second.
The correct placement costs exactly two clocks — three clock edges from line transition to the state machine acting, against one for the raw version. That is 0.23% of a bit at 868 clocks per bit and 10% at 20, and only the ratio matters.
An asynchronous signal is synchronised once and shared. Two synchronisers can resolve the same transition differently on the same clock, leaving two consumers with contradictory views and neither of them wrong.
The IP built in Module 11 violated both rules: the break detector's port is documented as requiring a synchronised line and the top level handed it the raw pin. 86 functional checks passed, and none of them could have failed. Fixed by publishing the receiver's synchronised line and sharing it, for two clocks of extra break latency against a window of eleven bit times — 0.02%.
A deliberately wrong placement was functionally indistinguishable from the correct one across five clock ratios and forty frames. The negative result is the finding: no test of outputs distinguishes correct CDC from incorrect, so the case for these rules is physical and the gate that enforces them is a CDC linter.
But a structural rule can have an observable shadow, and this one does: asserting that all consumers see one synchronised line gave 40 violations before the fix and 0 after. Where such a shadow exists, assert on it — it is three lines and it runs in the tests you already have.
13. What Comes Next
The receive pin is handled. Chapter 12.3 asks the question this module has been circling: which parts of a UART are actually clock-domain crossings?
The answer is narrower than most people assume, and getting it wrong in either direction is expensive — treating asynchronous serial timing as CDC produces structures that solve nothing, while missing a genuine crossing produces the failures this chapter has been describing. It also confronts the case this IP has so far avoided entirely: a register interface running on a different clock from the UART itself, and the asynchronous FIFO that then becomes unavoidable.
Browse the full path on the UART tutorials index. For the guarantees this chapter's placement rules protect, read back to Chapter 12.1.
Continue learning
Related tutorials
- Related topic
Why Receiving Is Harder Than Transmitting
A transmitter executes a schedule it wrote itself. A receiver must decide whether something is happening, whether it was real, where the positions are, and what value was there — four judgements from one edge on an input it does not control.
- Related topic
Metastability and the Asynchronous RX Input
Why the RX pin is a true asynchronous input, what metastability actually is, and — with the MTBF arithmetic worked out — what a synchroniser does and does not guarantee.
- Related topic
Sampling Centres and the Timing Margin Budget
Half a bit period separates an interval's centre from its boundary. That half-bit is a budget spent by origin uncertainty, interval construction, accumulated drift and the decision mechanism — and the last interval of a frame is where it runs out first.
- Related topic
Parity Generation, Checking and Error Detection
One interval, one XOR reduction, and a detection guarantee with a sharp edge: parity catches every corruption that flips an odd number of protected bits and provably misses every even-numbered one — demonstrated, not asserted.
Where this fits
Part of the UART curriculum.
