UART · Module 12
FPGA and ASIC Implementation Differences
Reset resources, memory choice for FIFOs, I/O and clocking — the places where identical UART RTL becomes two different pieces of hardware, with the numbers worked out.
The RTL built across Modules 6 through 12 is technology-independent, and that is a real achievement rather than a given — it took the enable-not-clock decision of Chapter 8.1, the parameterisation of Chapter 11.4 and the structural CDC of Chapter 12.2 to keep it that way.
It is still not the same hardware on both targets. Four things differ enough to change decisions: reset, memory, I/O and clocking. This chapter is about those four, and it closes the module.
1. Reset: Free on One Target, Mandatory on the Other
An FPGA loads a bitstream, and that bitstream sets every flip-flop's initial value. The device comes up in a defined state without a reset ever being asserted.
An ASIC powers on with every flop at an unknown value. Anything read before it is written must be reset, and there is no alternative.
This changes a real design decision:
| FPGA | ASIC | |
|---|---|---|
| state at power-on | defined by the bitstream | unknown |
| reset needed for correctness? | only for re-initialisation | yes, on everything read before written |
| cost of resetting everything | routing and fanout for no benefit | the same cost, and it is necessary |
| what to reset | control state: FSMs, pointers, flags | control and datapath as required |
In the UART, reset what carries meaning. The state machines, the FIFO pointers, the sticky error flags, the configuration shadow — all must reset on both targets, because all are read before they are written. The FIFO memory is different: an entry is written before it is read, by construction, so resetting the array buys nothing and on an FPGA it can prevent the array from mapping into a memory primitive at all.
2. Memory: The FIFO Is Far Too Small for a Block RAM
At its default parameters the IP's queues hold:
TX FIFO 16 x 8 = 128 bits
RX FIFO 16 x 10 = 160 bits (data + framing + parity — Chapter 10.2)
total = 288 bitsAgainst the memory primitives available:
| Primitive | Capacity | A 16×10 RX FIFO uses | Wasted |
|---|---|---|---|
| Xilinx RAMB18E1 | 18,432 bits | 0.87% | 99.13% |
| Xilinx RAMB36E1 | 36,864 bits | 0.43% | 99.57% |
| Intel M9K | 9,216 bits | 1.74% | 98.26% |
| Intel M20K | 20,480 bits | 0.78% | 99.22% |
Using a block RAM for a UART FIFO wastes over 99% of a scarce resource. Distributed RAM — the LUTs themselves used as small memories — or plain flip-flops are the right choice, and any sensible tool picks one automatically at this size.
Where the crossover sits, exactly: an 18 Kb block holds 1,843 entries of 10 bits. And in link terms, at 115,200 baud 8N1:
| Buffered | Link time it represents |
|---|---|
| 16 bytes | 1.39 ms |
| 256 bytes | 22.22 ms |
| 1,024 bytes | 88.89 ms |
| 2,048 bytes | 177.78 ms |
A UART FIFO deep enough to fill a block RAM buffers about a sixth of a second of link. That is occasionally what a design wants — a logger that must survive a long interrupt-disabled window — and it is far more than Chapter 10.3's service-latency argument usually justifies. Size the FIFO from the service-latency requirement, and then notice which primitive that lands in; sizing it to fill a memory block is choosing a number for the wrong reason.
On an ASIC the same arithmetic points the other way. A compiled SRAM macro has a fixed overhead — address decode, sense amplifiers, a periphery that does not shrink — and at 288 bits the overhead dominates so completely that flip-flops win outright. The crossover is typically in the low thousands of bits, so a UART FIFO is always flops.
3. I/O: Four Pins, Two Very Different Boundaries
The UART presents four pins — rx_i, tx_o, cts_n_i, rts_n_o — and what sits behind them differs completely.
On an FPGA the boundary is already built. The I/O block provides buffering, optional registers on the pin, programmable pull-ups and selectable standards, all configured by constraints rather than RTL. Two settings are worth naming:
A pull-up on rx_i matters more than it looks. An unconnected receive pin floats, and a floating input drifts across the threshold, producing start edges from noise — Chapter 9.5's false starts, from nothing but an empty connector. A pull-up holds it at mark, which is idle, so a disconnected link is correctly silent.
Register tx_o in the I/O block. It is already registered in the RTL; placing that register in the IOB makes the pin-to-pad delay fixed and short instead of dependent on placement.
On an ASIC the boundary is a design task. Pad cells must be instantiated, with ESD structures, chosen drive strength and slew, possibly level shifters if the I/O supply differs from the core. Drive strength is a real decision: too little and the edge is slow enough to matter at high baud rates over a long cable; too much and the design radiates and draws current spikes.
Both targets need the pull-up decision made deliberately, and on an ASIC it is a pad-cell option rather than a constraint.
4. Clocking and Synchronisers
FPGAs have dedicated global clock networks — a small, countable number of them. Put clk on one. This is also the strongest practical argument for Chapter 8.1's rule: a generated clock either consumes one of these scarce resources or, worse, travels on ordinary routing with ordinary skew.
ASICs synthesise a clock tree balanced to the design's skew budget. A UART is small and undemanding, and its clock will simply be part of the larger tree.
Synchroniser primitives differ, and both need to be told. On an FPGA, ASYNC_REG (Xilinx) or the equivalent keeps the two stages adjacent so the full clock period is available for settling — the assumption Chapter 12.1's MTBF arithmetic rests on. On an ASIC, standard-cell libraries usually provide characterised synchroniser cells with a published τ, and using them turns that arithmetic from an estimate into a number with a datasheet behind it.
Neither is automatic. A synchroniser that is not marked is a synchroniser the tool may optimise, retime or separate.
5. A Porting Checklist
Moving this IP between targets, in order:
| # | Check | Why |
|---|---|---|
| 1 | is the clock frequency the same? | CLK_HZ is a parameter and the baud divisor follows from it — Chapter 11.4 |
| 2 | re-run the elaboration legality checks | they catch a re-target that violates a parameter relationship |
| 3 | reset strategy still appropriate? | §1 — an FPGA design that relied on bitstream init needs explicit reset on an ASIC |
| 4 | FIFO storage still mapping sensibly? | §2 — and confirm nothing reset the array |
| 5 | synchronisers still marked? | §4 — attributes are target-specific and get lost |
| 6 | pull-up on rx_i still present? | §3 — a constraint on one target, a pad option on the other |
| 7 | constraints ported, not copied | Chapter 12.5 — cell names change |
| 8 | re-run CDC lint | Chapter 12.2 — it is structural, and the structure just moved |
Item 1 is the one that actually bites. Changing CLK_HZ changes the baud divisor, which changes the baud error, which changes the timing budget. Chapter 4.3 computes it; the elaboration checks of Chapter 11.4 catch the illegal cases; nothing catches a legal configuration with an error large enough to be marginal. Re-run the budget arithmetic on every re-target, because it is the one thing in this list that produces a link which works on the bench and fails at temperature.
6. The UVM Shape of a Two-Clock Testbench
Everything this module verified was verified with directed, self-checking testbenches, because that is the right instrument for a block with three ports and one structural property. A UART that has reached this point, though, is about to be attached to a bus in Module 13 and dropped into a constrained-random environment — so it is worth being precise about what changes in UVM and what does not.
The structural difference from every previous module is that there is no longer
one clock, so there can no longer be one driver, one monitor, or one
uvm_event that both halves agree on. A CDC environment is two agents that
never share a timebase, joined only by a scoreboard that compares order
rather than time.
// ---------------------------------------------------------------------------
// uart_cdc_env — the shape of a UVM environment for an asynchronous FIFO
//
// The essential point: NOTHING in this environment is clocked by both
// clocks. The write agent lives entirely in wclk, the read agent entirely
// in rclk, and they meet only in the scoreboard, which compares ORDER and
// never time.
// ---------------------------------------------------------------------------
class uart_cdc_env extends uvm_env;
`uvm_component_utils(uart_cdc_env)
uart_wr_agent m_wr; // driver + monitor, wclk domain only
uart_rd_agent m_rd; // driver + monitor, rclk domain only
uart_cdc_sb m_sb;
function new(string name, uvm_component parent);
super.new(name, parent);
endfunction
function void build_phase(uvm_phase phase);
super.build_phase(phase);
m_wr = uart_wr_agent::type_id::create("m_wr", this);
m_rd = uart_rd_agent::type_id::create("m_rd", this);
m_sb = uart_cdc_sb ::type_id::create("m_sb", this);
endfunction
function void connect_phase(uvm_phase phase);
// Two analysis ports, two independent timebases, one scoreboard.
m_wr.mon.ap.connect(m_sb.wr_export);
m_rd.mon.ap.connect(m_sb.rd_export);
endfunction
endclassThe scoreboard is where the CDC-specific discipline lives. A conventional scoreboard often compares against an expected value at a time; that is not available here, because there is no instant at which both sides agree on the FIFO's contents. What is available — and what the directed testbenches of Chapter 12.3 already relied on — is that order is preserved.
class uart_cdc_sb extends uvm_scoreboard;
`uvm_component_utils(uart_cdc_sb)
uvm_analysis_imp_wr #(uart_word_item, uart_cdc_sb) wr_export;
uvm_analysis_imp_rd #(uart_word_item, uart_cdc_sb) rd_export;
// The model is a QUEUE, not a count and not an expected value. It is the
// only structure that encodes "order is preserved" without also claiming
// to know how many entries are resident at a given instant.
protected byte unsigned m_q[$];
protected int m_pushed, m_popped, m_mismatch;
function void write_wr(uart_word_item t);
m_q.push_back(t.data); // accepted pushes only -- see below
m_pushed++;
endfunction
function void write_rd(uart_word_item t);
byte unsigned expected;
if (m_q.size() == 0) begin
`uvm_error("CDC_SB", "a word was popped that was never pushed")
return;
end
expected = m_q.pop_front();
m_popped++;
if (t.data !== expected) begin
m_mismatch++;
`uvm_error("CDC_SB", $sformatf(
"popped %02h, expected %02h (word %0d)", t.data, expected, m_popped))
end
endfunction
function void check_phase(uvm_phase phase);
// Drained at end of test: everything pushed came out, in order.
if (m_q.size() != 0)
`uvm_error("CDC_SB", $sformatf("%0d words never arrived", m_q.size()))
endfunction
endclassConstrained-random gains something real here, which is not true of every block. The variable worth randomising is the clock ratio itself:
class uart_cdc_cfg extends uvm_object;
rand int unsigned wclk_ps;
rand int unsigned rclk_ps;
// Unrelated periods, spanning both directions of the ratio. The
// constraint that matters is the LAST one: two periods with a small
// common factor produce a repeating edge pattern that visits only a
// handful of phase relationships, which is precisely the case a CDC
// test must avoid.
constraint c_range { wclk_ps inside {[2000:40000]};
rclk_ps inside {[2000:40000]}; }
constraint c_spread { (wclk_ps * 4 < rclk_ps) || (rclk_ps * 4 < wclk_ps)
|| (wclk_ps != rclk_ps); }
constraint c_coprime { wclk_ps % 1000 != 0 || rclk_ps % 1000 != 0; }
endclassAnd a functional coverage model that is about the boundaries, since the middle of the range is where nothing happens:
covergroup cg_cdc @(posedge wclk);
cp_full : coverpoint wfull;
cp_empty : coverpoint rempty iff (rrst_n);
cp_ratio : coverpoint ratio_bucket {
bins wr_much_faster = {0}; // both directions of the ratio must
bins similar = {1}; // be reached, or one flag is never
bins rd_much_faster = {2}; // exercised at all
}
cp_occup : coverpoint occupancy_estimate {
bins empty = {0};
bins low = {[1:2]};
bins mid = {[3:DEPTH-2]};
bins high = {DEPTH-1};
bins full = {DEPTH}; // must be hit, or capacity is untested
}
x_full_ratio : cross cp_full, cp_ratio;
x_empty_ratio: cross cp_empty, cp_ratio;
endcovergroup7. Module 12 Verification Evidence
Every listing published in this module was extracted from the page you are reading, compiled with a real tool, and simulated. The numbers below are the output of those runs, not a summary of intent.
Three blocks, three languages, nine designs:
| Block | SystemVerilog | Verilog-2001 | VHDL-2008 | Chapter |
|---|---|---|---|---|
uart_sync_edge | ✅ | ✅ | ✅ | 12.2 |
uart_reset_sync | ✅ | ✅ | ✅ | 12.4 |
uart_async_fifo | ✅ | ✅ | ✅ | 12.3 |
Nine testbenches, identical check counts across all three languages:
| Suite | Checks | Verilog-2001 | SystemVerilog | VHDL-2008 |
|---|---|---|---|---|
uart_sync_edge | 19 | 19 / 0 | 19 / 0 | 19 / 0 |
uart_reset_sync | 16 | 16 / 0 | 16 / 0 | 16 / 0 |
uart_async_fifo | 36 | 36 / 0 | 36 / 0 | 36 / 0 |
| Total per language | 71 | 71 / 0 | 71 / 0 | 71 / 0 |
213 checks across the three languages, 0 failures. Tooling: Icarus Verilog
13.0 (-g2001 and -g2012) and NVC 1.23.0 for VHDL-2008.
Mutation campaign — nine defects, nine killed:
| # | Block | Defect installed | Killed by |
|---|---|---|---|
| M1 | sync_edge | sync_o read from the first flop | latency measurement (7 checks) |
| M2 | sync_edge | edge taken from the raw pin | edge-vs-history observer (5) |
| M3 | sync_edge | RESET_VALUE ignored | reset-value checks (4) |
| M4 | reset_sync | assertion made synchronous | the clock-stopped block (7) |
| M5 | reset_sync | output from the first flop | release latency (6) |
| M6 | async_fifo | binary pointer crossing | gray invariant, then deadlock |
| M7 | async_fifo | full compares plain equality | deadlock, caught by watchdog |
| M8 | async_fifo | pointer advances when full | capacity + data (5) |
| M9 | async_fifo | synchronisers cut to one flop | structural check only (2) |
Each mutation was verified to have actually modified the source before its result was scored. This matters more than it sounds: a patch that silently matches nothing produces a "surviving mutant" that is really the original design, and it is the easiest way to award yourself a passing grade you did not earn. Two candidate mutations in this campaign failed to apply on the first attempt and were caught by that guard.
Two design defects were found and fixed during this module, both by activities other than running the tests:
-
The asynchronous FIFO would not elaborate at its smallest documented size. The textbook spelling of the full comparison contains a part-select that goes negative at
DEPTH_L2 = 1. Found by translating to a second language and by testing the boundary of the documented parameter range rather than its middle. Fixed in all three languages by replacing the part-select with a mask comparison. (12.3) -
A break detector wired to the raw pin instead of the synchronised line, in the IP assembled in Module 11. Found by reading the connection. It survived 86 passing checks, none of which could have failed. (12.2)
Neither was found by simulation, and that is the single sentence this module would keep if it had to discard the rest.
8. Verification
The functional suites do not change, and that is the payoff for keeping the RTL technology-independent. The same 86 checks from Module 11, plus this module's, run unchanged on either target.
What changes is the gates around them:
| Gate | FPGA | ASIC |
|---|---|---|
| synthesis warnings as errors | latch, combinational loop, inferred clock | the same |
| CDC lint | required | required |
| utilisation review | LUT/FF/BRAM report | area report |
| timing | post-route, worst corner | multi-corner, multi-mode |
| gate-level simulation | rare | usually required |
Gate-level simulation is where an ASIC finds what RTL simulation cannot — X-propagation from uninitialised flops in particular, which is §1's problem made visible. An FPGA design rarely runs it because the bitstream removed the question.
9. Debugging
10. What This Means in Practice
Prototype on an FPGA, and do not trust it about reset. An FPGA prototype validates function and finds protocol bugs cheaply. It cannot validate a reset strategy, because it initialises state the ASIC will not.
Watch the IP's size in perspective. 288 bits of queue and a few dozen registers is small enough that on either target the UART is a rounding error — which is precisely why it is worth building well rather than optimising.
Check what the tool actually inferred. The utilisation report tells you whether the FIFO went into LUT RAM or flops, and whether an unexpected block RAM appeared. Both are cheap to read and both occasionally surprise.
11. Understanding Check
12. Summary
The RTL is technology-independent and four things around it are not: reset, memory, I/O and clocking.
An FPGA initialises state from its bitstream; an ASIC does not. Reset control state on both; reset the FIFO pointers and never the array, which on an FPGA would cost the memory mapping and on an ASIC a pointless reset tree.
A UART FIFO is far too small for a block RAM — 160 bits against 18,432, using 0.87% and wasting over 99%. One block would hold 1,843 entries, about 160 ms of link time. Size the queue from service latency and let the primitive follow.
On an ASIC the same size argues for flip-flops, because a compiled memory's periphery overhead dominates at a few hundred bits.
A pull-up on rx_i turns a disconnected link from a noise source into silence — a constraint on one target, a pad option on the other, and a deliberate decision on both.
Synchronisers must be marked on either target, or the tool may separate, retime or optimise the stages and quietly spend the settling margin the MTBF arithmetic assumed.
And the checklist item that bites: changing CLK_HZ changes the baud error, and a legal-but-marginal divisor produces a link that passes on the bench and fails at temperature.
13. What Comes Next
Module 12 is complete, and so is the hardware. The UART has been designed, integrated, parameterised, made testable, synchronised, reset properly, constrained and targeted. It is a piece of IP that would survive review.
It still has no way to talk to software. Every configuration input is a wire and every status output is a wire, and something has to turn those into addresses a driver can write.
Module 13 attaches it to a bus: the register map this module has repeatedly deferred, the interrupt controller that turns rx_trigger_o into something a CPU notices, and the DMA hooks that stop a processor spending its life servicing a byte at a time. The second clock Chapter 12.3 built the asynchronous FIFO for arrives there, and the structures from this module are what make that attachment safe.
Browse the full path on the UART tutorials index. For the crossings this chapter's targets have to implement, read back to Chapter 12.3.
Continue learning
Related tutorials
- Related topic
What a UART Actually Is
Two digital systems need to exchange a small amount of data over very few wires, and no clock travels with it. A UART is the logic that answers that problem — it converts between locally meaningful parallel data and timed activity on a single line, and the timing agreement it depends on is what the rest of the curriculum builds.
- Related topic
Synchronous vs Asynchronous Serial Links
A forwarded clock is a sampling reference generated by the same source as the data. Remove it and the receiver must assemble one from a configured rate, an observable event in the signal, and its own local clock — the responsibility shift that turns a receiver into a state machine and shapes every UART design decision that follows.
- Related topic
The UART Link: TX, RX, Idle and Full Duplex
A UART link is two independent one-way conductors, not one bidirectional bus — which removes arbitration, turnaround and direction control from the design, makes the naming endpoint-relative, and means full duplex guarantees simultaneity and nothing else.
- Related topic
UART vs RS-232, RS-485 and the Physical Layer
A UART controller decides the bit pattern; a transceiver decides what those bits physically are, how far they reach and how many devices can share the medium. Same controller, different transceiver, different network — and why topology is never a property of the framing.
Where this fits
Part of the UART curriculum.
