Skip to content
VLSI Mentor

UART · Module 12

Metastability and the Asynchronous RX Input

Why the RX pin is a true asynchronous input, what metastability actually is, and — with the MTBF arithmetic worked out — what a synchroniser does and does not guarantee.

Everything built so far assumed the receive line could simply be read. Modules 5 through 11 sampled it, filtered it, counted its low intervals and decoded frames from it, and every one of those chapters was correct given a signal in this clock domain.

The receive pin is not in this clock domain. It is driven by a device with its own crystal, and there is no relationship whatsoever between its transitions and this design's clock edges. It is the one genuinely asynchronous input in the entire UART, and this module is where that is taken seriously.

1. What Makes an Input Asynchronous

An input is asynchronous when nothing constrains when it changes relative to the sampling clock. Not "it changes rarely", not "it changes slowly" — those are properties of the data rate, and they do not help.

The UART receive line qualifies in the strictest sense:

Driven bya different device, with its own oscillator
Frequency relationship to clknone — two free-running crystals
Phase relationshipnone, and drifting continuously
Can it change during this design's setup/hold window?yes, and eventually it will

The last row is the whole chapter. Over a long enough run, every phase relationship between the far end's bit transitions and this design's clock edges occurs, including the ones that land inside the sampling flop's aperture.

2. What Metastability Actually Is

A flip-flop is a bistable circuit: two stable states, and a balance point between them. Data arriving safely before the clock edge drives it firmly toward one of the stable states. Data arriving inside the aperture can leave it near the balance point, and from there it resolves — but not in a bounded time.

The output during that interval is not a valid logic level. It may sit near mid-supply, it may oscillate, and downstream gates reading it may interpret it differently from one another. It always resolves; what is unbounded is when.

A block diagram of the asynchronous input boundary. On the left, a far-end device driven by its own independent oscillator drives the receive line across a cable. That line arrives at the receive pin, which is the asynchronous boundary. The pin feeds the first synchroniser flop, whose output is permitted to be metastable and must not be read by anything else. That flop feeds a second synchroniser flop, which is given a full clock period to allow any metastable state to resolve. The second flop's output is a settled value in the local clock domain, and only that output is consumed by the rest of the design: the edge detector, the receive state machine, and the break detector all read the same synchronised signal.far-end deviceits own crystalrx_i (the pin)ASYNCHRONOUSrx_meta_qmay be metastablerx_sync_qsettled — safe to useedge detectstart candidateRX FSMsamples databreak detectcounts low tickscableasync1 clocksettledsettledsame12
Figure 1 — the receive pin crossing into the clock domain. The pin is driven by a device with an unrelated oscillator, so its transitions can land inside the first flop's aperture. That flop's output is allowed to be metastable; the second flop gives it a full clock period to resolve before anything else in the design is permitted to look at it.

The second flop does not prevent metastability — it contains it. The first flop is allowed to go metastable; what the structure guarantees is that nothing except the second flop ever sees it, and that the second flop gets a full clock period of settling time before it samples.

Containment, not prevention

8 cycles
A timing trace of eight clock cycles showing metastability being contained. A clock runs throughout. The asynchronous receive line falls in the third cycle, at a moment that lands inside the first synchroniser flop's setup and hold aperture. The first flop's output becomes invalid during the fourth cycle, shown as an unknown value, because it was left near its balance point. By the fifth clock edge it has resolved to a settled level. The second synchroniser flop, which samples one clock later, therefore captures a clean logic level and its output is valid throughout. The rest of the design reads only the second flop, so it never observes the invalid interval.metametatransition inside the aperturetransition inside theapertureresolved — either value is legalresolved — either value islegalclkrx_i (async)rx_meta_qXrx_sync_qsafe to readt0t1t2t3t4t5t6t7
Figure 2 — a line transition landing inside the first flop's aperture. The first synchroniser output is invalid for part of a clock period and must be treated as unreadable. By the next edge it has resolved — to either value, both of which are legitimate — and the second flop captures a clean level. The design never sees the invalid interval.

3. The MTBF Arithmetic, Done Properly

The standard model is

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
                    exp(t_r / tau)
    MTBF  =  ──────────────────────────────
              T_w  x  f_clk  x  f_data

where t_r is the time allowed for resolution, tau the flop's resolution time constant, T_w its metastability window, f_clk the sampling clock and f_data the rate of asynchronous transitions.

t_r is what a synchroniser buys. One flop leaves roughly one clock period minus setup; each extra stage adds a full period. And because t_r appears in an exponent, adding a stage does not improve MTBF by a factor — it improves it by orders of magnitude.

Worked exactly, for a 115,200 baud line (f_data = 115,200 transitions/s worst case) sampled at 100 MHz, with T_w = 50 ps and t_setup = 0.5 ns:

tau1 flop2 flops3 flops
10 ps10⁴⁰² years10⁸³⁷ years10¹²⁷¹ years
20 ps10¹⁹⁶ years10⁴¹³ years10⁶³⁰ years
50 ps10⁷² years10¹⁵⁹ years10²⁴⁶ years
100 ps10³¹ years10⁷⁴ years10¹¹⁸ years
200 ps10¹⁰ years10³² years10⁵⁴ years

And a UART is the gentle case. Raising f_data from a serial line to a 100 MHz parallel interface — the more typical CDC situation — moves two-flop MTBF at 500 MHz from 3.9 hours to 16 seconds. Same structure, same arithmetic, an unusable answer. That is why wider and faster crossings need more than flops, which is Chapter 12.3's subject.

4. What a Synchroniser Does Not Give You

Three guarantees people assume and do not get:

It does not tell you which value you got. When a transition lands near an edge, the synchroniser resolves to 0 or 1 — and both are correct. The line genuinely was changing; either answer is a truthful report of an ambiguous instant. What you are guaranteed is a settled value, not a particular one.

It does not preserve the timing of the transition. The edge is quantised to a clock boundary, which is exactly the δ_origin term Chapter 5.2 put into the timing budget — up to one clock period of uncertainty in when the start edge is deemed to have arrived. The synchroniser is where that uncertainty is created, and the oversampling design is what absorbs it.

It does not filter anything. A glitch wider than a clock period passes straight through, resolved and clean-looking. Noise rejection is Chapter 5.4's majority vote, and it operates after synchronisation on a signal that is already settled — never instead of it.

5. What It Costs

Measured on the receiver as built, with the line dropped 1 ns after a clock edge at 100 MHz:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
=== 100 MHz clock, 10.00 ns period. Falling edge on rx at t=286.00 ns
  rx_meta_q changed at      295.00 ns  (+0.90 clocks)
  rx_sync_q changed at      305.00 ns  (+1.90 clocks)
  FSM consumes start_cand   315.00 ns  (+2.90 clocks)

Three clock edges from line transition to the state machine acting on it: one to capture, one to resolve, one for the registered edge detection that follows. At 100 MHz that is 30 ns against an 8.68 µs bit period — 0.35% of a bit, absorbed without comment by a design with ±8 oversample ticks of margin.

The ratio that matters is clocks per bit, not nanoseconds. The same three clocks at 20 MHz with a 1 Mbaud link — 20 clocks per bit — is 15% of a bit time. Chapter 12.2 develops what that does to the sampling point and where the limit actually is.

6. Proving the Claim: a Broken Design That Passes Everything

Section 6 will state that simulation cannot find metastability. That is a strong claim and it deserves more than an assertion, so here is the experiment.

Take the asynchronous FIFO built in Chapter 12.3 — a design whose entire correctness rests on crossing pointers safely between two unrelated clocks. Cut each of its two-flop pointer synchronisers down to one flop. Nothing else changes:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// correct — the second flop gives the first a full clock to resolve
wq1_rgray <= rgray_q;
wq2_rgray <= wq1_rgray;

// the defect — one flop, no resolution time, read immediately
wq2_rgray <= rgray_q;

In silicon this is the difference between a FIFO that runs for years and one that corrupts a pointer the first time two edges land inside the aperture window. Now run it against a suite of thirty-four behavioural checks — data integrity across four clock ratios, capacity, full and empty behaviour, overflow rejection, reset during traffic, 164 words through the crossing:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
--- one-flop pointer synchronisers, behavioural checks only ---
  PASS fast-write / slow-read: sequence intact
  PASS slow-write / fast-read: all 40 words arrived
  PASS incommensurate ratio: 40+40+60 words total
  PASS exactly DEPTH words were accepted -- not DEPTH-1, not DEPTH+1
  PASS draining recovers exactly DEPTH words
  PASS outstanding words never exceeded DEPTH -- no silent overwrite
  PASS reset during traffic returns the FIFO to empty
  ... 34 of 34 behavioural checks PASS

Nothing lost, nothing duplicated, nothing reordered. The design is broken in the most fundamental way a CDC design can be broken, and the regression is green.

This is not a weakness of that particular testbench. It is a property of the simulator. There is no aperture window in an event-driven model, so there is no violation to commit; there is no unresolved value, so there is nothing for the second flop to resolve. The second flop's entire job is to do nothing, observably, and a simulator agrees that nothing is what it does.

7. Verification

Simulation cannot find metastability. An event simulator has no aperture, no balance point and no resolution time; a flop sampling a changing signal simply takes whichever value the scheduler says is current. A design with no synchroniser at all passes every functional test, at every parameter set, forever.

That is the single most important verification fact in this module, and it has a direct consequence: synchroniser correctness is enforced structurally, not by testing.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// Assertion — the raw pin must never reach anything except the first
// synchroniser flop. This is a STRUCTURAL rule; no stimulus can violate it
// and no stimulus can confirm it, which is exactly why it is written down.
//
// In practice this is enforced by CDC lint (Spyglass CDC, Questa CDC,
// Conformal CDC), which reads the netlist rather than running it. A simulator
// cannot help here and reporting a clean simulation as evidence is a mistake.

// What CAN be checked in simulation is the LATENCY contract, which the rest
// of the design's timing budget depends on:
property p_sync_is_two_stages;
    @(posedge clk) disable iff (!rst_n)
        1 |=> (rx_sync_q == $past(rx_meta_q));
endproperty
assert property (p_sync_is_two_stages);

Icarus Verilog 13.0 supports neither named properties nor $past, so what was actually executed is the procedural equivalent — and it was mutation-checked: collapsing the pair to a single flop (rx_sync_q <= rx_i) makes it fire.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
logic meta_d; int lat_fail = 0;
always @(posedge clk) if (rst_n_dly) begin
    if (rx_sync_q !== meta_d) lat_fail++;   // fires on the 1-flop mutant
    meta_d <= rx_meta_q;
end
Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
  pass 2FF contract: rx_sync_q is rx_meta_q delayed exactly one clock
== 5 checks, 0 failures ==

Run a CDC linter, and treat its report as a gate. The rules it applies — every asynchronous input synchronised, exactly one synchroniser per signal, no combinational logic between the stages, no fan-out from the first stage — are all structural, and all invisible to the testbench that passed.

8. What This Means on an FPGA

Two flops, and they are free. The receiver already contains them — rx_meta_q and rx_sync_q in Chapter 6.1's datapath, which is why every functional chapter since was entitled to treat the line as readable.

Keep the pair close and let the tool know. FPGA flows recognise synchronisers and place the stages adjacently — Xilinx ASYNC_REG, Intel's equivalent — which maximises the settling time actually available. Without the attribute the placer is free to put a long route between the stages and quietly spend the margin this chapter just computed.

Do not add a third flop reflexively. At 100 MHz on a serial line the table says the second flop already bought twenty-two orders of magnitude of margin. A third is warranted for genuinely fast clocks or high-rate crossings, and elsewhere it is one more flop and one more clock of latency for nothing.

The pin needs a constraint, not just a synchroniser. Its path has no meaningful setup requirement, and leaving it unconstrained lets the tool try to meet an imaginary one — Chapter 12.5.

9. Understanding Check

10. Summary

The receive pin is the one genuinely asynchronous input in the UART — driven by an unrelated oscillator, with a phase relationship that drifts through every possible value including the ones inside a flop's aperture.

A slow line is not a safe line. Rarity sets the rate of the problem, not its possibility, and the difference is settled by arithmetic rather than intuition.

The second flop contains metastability rather than preventing it, by giving the first a full clock to resolve and ensuring nothing else ever reads it.

Resolution time is an exponent, so margin collapses rather than degrading: at 100 MHz one flop gives 10¹⁰ years and at 250 MHz the same flop gives 7.8 hours. The RTL is identical and records nothing about which case it is. One extra flop restores twenty-two orders of magnitude.

A synchroniser guarantees a settled value — not a particular value, not the transition's timing, and no filtering at all. The timing uncertainty it creates is the start-edge term the oversampling design was built to absorb.

It costs three clock edges from line transition to the state machine acting — 0.35% of a bit at 100 MHz and 115,200 baud, and 15% at 20 MHz and 1 Mbaud.

And the fact that governs this entire module: simulation cannot find a missing synchroniser. A design with none passes everything. Correctness here is structural, and a CDC linter is the tool that checks it.

11. What Comes Next

The synchroniser exists and its guarantees are now precise. Chapter 12.2 asks where it belongs — and the answer turns out to have been violated in the IP assembled in Module 11.

The break detector's input port is named rx_sync_i and documents itself as expecting a synchronised line. The top level hands it the raw pin. Every functional test passed, because §6's fact guarantees they would.

Browse the full path on the UART tutorials index. For the start-edge uncertainty this chapter's quantisation creates, read back to Chapter 5.2.

Continue learning

Where this fits

Part of the UART curriculum.