Skip to content
VLSI Mentor

UART · Module 13

DMA Interaction and Bulk Transfer

Request and acknowledge handshaking with a DMA engine, where the residual bytes at the end of a transfer go, and a real defect that only a DMA-style reader could expose.

Interrupts make a UART usable by software. They also mean a processor takes an exception, saves context, runs a handler and returns every few bytes — and at high baud rates that stops being free.

DMA moves the data path out of the processor entirely. The UART's side of that is small: two request lines and an enable bit. What is not small is the question every DMA-driven UART has to answer — what happens to the last few bytes, the ones below any threshold that a request-driven engine will never come to collect.

1. What the UART Has To Provide

Very little, and that is the point:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// Chapter 13.4 — the entire DMA interface.
assign dma_tx_req_o = dma_tx_en_q && tx_trigger;
assign dma_rx_req_o = dma_rx_en_q && (rx_trigger || rx_timeout_q);

Two levels, each an enable ANDed with a condition the design already computed. No new state, no new logic, no protocol.

SignalMeansDeasserts when
dma_tx_req_othe transmit queue has roomthe engine fills it above the trigger
dma_rx_req_othe receive queue has data worth collectingthe engine drains it below the trigger

The acknowledge is the transfer itself. There is no separate ack line, because a DMA engine servicing a request performs bus accesses to DATA, and those accesses are what change the condition. A request line that needed an explicit acknowledgement would be adding a handshake on top of a handshake.

This is the same level-source reasoning as Chapter 13.3. A request is a condition, not an event; it is cleared by servicing, not by acknowledging; and it must never be latched.

2. Trigger Levels Mean Something Different Now

Under interrupts, the trigger level trades interrupt frequency against latency and overrun headroom — Chapter 10.3. Under DMA the processor is not involved, so one side of that trade disappears and a different one takes its place:

Interrupt-drivenDMA-driven
Cost per service eventcontext switch, handler, returna burst of bus transfers
What a low trigger costsmany interruptsmany short bursts — poor bus efficiency
What a high trigger costslatency, less overrun headroomlatency, less overrun headroom
Natural settingas high as latency allowsmatched to the engine's burst size

The right DMA trigger is the burst size the engine is configured for. If the engine moves eight bytes per request and the trigger is four, every other burst runs the queue dry halfway through and the engine issues a bus transaction for nothing. If the trigger is twelve and the burst is eight, four bytes are left behind on every service — which is a small version of §4's problem occurring continuously.

This is why Chapter 13.2 made the trigger a register. The correct value depends on how the DMA engine is programmed, which is a software decision made long after the UART was synthesised.

Measured, with a trigger of 8 and an engine that moves 8 per request:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
-- 1. 16 bytes with an RX trigger of 8
      DMA moved 16 bytes in 2 bursts, CPU interrupts: 0
  pass all 16 bytes moved by DMA
  pass no CPU interrupt was needed

Sixteen bytes, two bus events, zero processor involvement. The interrupt-driven equivalent at a trigger of one would be sixteen exceptions.

3. Why Not Just Set the Trigger to One

Because then DMA buys almost nothing. A request per byte means a bus transaction per byte, and while that is cheaper than an interrupt per byte, it wastes the thing DMA exists to exploit: amortising the cost of getting to the peripheral across many bytes.

The measurement, with 11 bytes and a trigger of 8:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
-- 3. what DMA actually saves
      11 bytes, trigger 8: DMA needed 2 service events
      interrupt-per-byte would have needed 11

Two service events instead of eleven. The saving grows with the trigger level, and so does the latency — the same trade Chapter 10.3 analysed, with a different cost on one axis.

4. The Residual Bytes

Here is the problem that makes this chapter necessary.

A DMA engine moves data when requested. The request is a threshold condition. A message whose length is not a multiple of the threshold leaves a remainder that never reaches the threshold, and the engine never comes back for it.

Eleven bytes arriving with a trigger of eight:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
-- 2. 11 bytes with a trigger of 8: one whole burst, 3 left over
      after the link went quiet: DMA moved 8 of 11 bytes
      RX empty = 0   (bytes still sitting in the queue)
  pass the residual bytes are STRANDED without a timeout
  pass and they really are still in the queue

Three bytes sat in the queue indefinitely. Not lost, not corrupted — correctly received, correctly stored, and unreachable. The link is idle, the engine is idle, the processor believes the transfer is in progress, and nothing will ever change.

The receive-idle timeout is what solves it

Chapter 13.3 §6 built the timeout so that a short message would still raise an interrupt. It turns out to be the same mechanism, and this is why it is wired into the request:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
assign dma_rx_req_o = dma_rx_en_q && (rx_trigger || rx_timeout_q);

The request has two causes: enough data, or no more data coming. The second is what collects the remainder.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   -> waiting for the receive-idle timeout
  pass idle timeout asserts
  pass and it raises the DMA request for the residual
      after the timeout: DMA moved 11 of 11 bytes
  pass all 11 bytes recovered
  pass queue now empty
A sequence showing eleven received bytes being collected by a DMA engine with a trigger level of eight. Eleven bytes arrive from the line and are pushed into the receive queue. When the queue level reaches eight, the receive trigger condition asserts and raises the DMA request. The engine performs a burst of eight reads, which drains the queue to three and withdraws the trigger condition, so the request deasserts. At this point three bytes remain in the queue below the threshold and the engine has no reason to return. The line then goes quiet, and after the configured idle period the receive timeout asserts. The timeout raises the same DMA request by a different cause, the engine performs a short burst, and the final three bytes are collected, after which the queue is empty and both causes are gone.The lineRX queuedma_rx_reqDMA engine11 bytes arrivelevel reaches 8 —triggerrequest assertedburst of 8 readslevel 3 — triggerwithdrawnrequest deasserted:3 STRANDEDline goes quietidle timeout — SAMErequest, new causeshort burst collectsthe residual
Figure 1 — how eleven bytes reach a DMA engine whose trigger is eight. The first eight cross on the threshold condition; the remaining three cross only because the line going quiet raises the same request by a different cause.

One request line, two causes

7 cycles
A trace of seven events showing the receive DMA request line across a complete transfer of eleven bytes with a trigger level of eight. The queue level rises from zero as bytes arrive, reaching eight. The receive trigger condition asserts at that point and the DMA request rises with it. The engine performs a burst of eight reads, taking the queue level down to three, at which point the trigger condition falls away and the request deasserts, leaving three bytes below the threshold with no reason for the engine to return. The line then goes quiet and the receive-idle timeout asserts after the configured interval, raising the same request line again from a different cause. The engine performs a short burst, the queue level reaches zero, and both the timeout and the request fall for good.no reason for the engine to returnno reason for theengine to return3 bytes STRANDED here3 bytes STRANDED heresame request, new causesame request, new causeeventidle8 inburst3 leftquiettimeoutdonerx queue level0883330rx_triggerrx_timeoutdma_rx_req_oengine activet0t1t2t3t4t5t6
Figure 2 — the request line across the whole transfer. Columns are events, not clock cycles. The request rises on the threshold, falls when the burst drains the queue below it, and — crucially — rises a second time from a completely different cause once the line goes quiet. Without that second cause the trace would end with three bytes still in the queue and the request low forever.

5. The Defect the DMA Engine Found

Modelling the engine exposed a real bug in the receive path — one that eleven testbenches across three modules had never triggered.

The symptom. With the receive queue full and the engine reading back to back, the same byte was returned forever and the queue level never fell:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
        [dbg] read 17: data=50 rxlevel=16 rd_ready=1
        [dbg] read 18: data=50 rxlevel=16 rd_ready=1
        [dbg] read 19: data=50 rxlevel=16 rd_ready=1
      after the link went quiet: DMA moved 400 of 19 bytes

Four hundred bytes out of nineteen sent, all of them the same value. The queue was pinned at full.

The cause was in Chapter 11.2's integration, one line:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
assign rx_push       = rx_char_valid && !cfg_rx_flush_i;   // WRONG
assign rx_char_ready = !rx_fifo_full;

The push was qualified on the receiver's offer and the handshake on the queue's room, and those are not the same condition. When the queue is full and a pop happens on the same edge, the FIFO has room for the push and accepts it — but rx_fifo_full still reads high, so rx_char_ready is low, so the receiver's handshake does not complete and it continues to offer the same character. Every subsequent pop re-pushes it.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// Chapter 13.4 — the push must be the HANDSHAKE, not merely the offer.
// Gating only on rx_char_valid meant that when the queue was full and a
// pop happened on the same edge, the FIFO accepted a push while the
// receiver's handshake did NOT complete — so the same character was
// written again on every subsequent pop and the queue never drained.
assign rx_push = rx_char_valid && rx_char_ready && !cfg_rx_flush_i;

Verified after the fix, across every suite in Modules 11, 12 and 13:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
  tb_ip: 53   tb_param: 9    tb_break_param: 11   tb_loop: 13
  tb_lat: 5   tb_sync: 5     tb_view: 0 violations
  tb_afifo: 5 tb_rstsync: 5  tb_reset: 7          tb_synth: 6
  tb_regs: 40 tb_irq: 15     tb_dma: 9            tb_pop: 2

  186 checks, 0 failures

6. What a Driver Still Has To Do

DMA removes the processor from the data path, not from the peripheral:

Errors still need attention. A framing or parity error sets a sticky flag and the byte still goes into the queue with its per-byte status — and DMA moves the byte without looking at either. The error interrupt remains enabled under DMA, and it is the only thing that will tell software something went wrong. A DMA-driven UART with errors masked reports nothing and delivers corrupt data silently.

Per-byte status is lost to DMA. The engine reads DATA and stores the low byte; bits 8 and 9 go nowhere. A design that must attribute an error to a specific byte cannot use DMA for that stream, or must store all 32 bits per byte — which is four times the memory and is occasionally the right answer.

The end of a transfer is a software decision. The timeout collects the residual; it does not tell software the message is complete. That is the protocol's job above.

Flow control still matters. DMA makes the UART faster to service, which reduces overruns — it does not eliminate them, because the engine's latency under bus contention is not bounded by anything the UART knows.

7. The DMA Interface in Three Languages

Section 1 gave the whole interface as two assignments. Here it is as a module in each language, which adds nothing functional and one thing worth having: a name, so that the reasoning about it lives next to it instead of in a comment inside a larger file.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// ===========================================================================
//  uart_dma_if — Synthesizable SystemVerilog
//
//  Chapter 13.4. The entire DMA interface, and its smallness is the point.
//
//  Two levels, each an enable ANDed with a condition the design had already
//  computed. No new state, no new logic, no protocol.
//
//  THERE IS NO ACKNOWLEDGE LINE, and that is deliberate. A DMA engine
//  servicing a request performs bus accesses to DATA, and those accesses are
//  what change the condition. A request line that needed an explicit
//  acknowledgement would be adding a handshake on top of a handshake.
//
//  This is the same reasoning as Chapter 13.3: a request is a CONDITION, not
//  an event. It is cleared by servicing, not by acknowledging, and it must
//  never be latched. A latched request survives the transfer that satisfied
//  it and the engine reads a queue that has nothing in it.
//
//  The receive request includes the idle timeout as well as the trigger,
//  which is what rescues the RESIDUAL BYTES at the end of a transfer
//  (Chapter 13.4 section 4). Without it a message whose length is not a
//  multiple of the trigger level leaves its tail stranded in the queue.
// ===========================================================================
module uart_dma_if (
    input  logic clk,             // present for the assertions below only
    input  logic rst_n,

    input  logic dma_tx_en_i,     // DMA_CTRL bit 0
    input  logic dma_rx_en_i,     // DMA_CTRL bit 1

    input  logic tx_trig_i,       // transmit queue has room
    input  logic rx_trig_i,       // receive queue has data worth collecting
    input  logic rx_timeout_i,    // ...or has had some for a while

    output logic dma_tx_req_o,
    output logic dma_rx_req_o
);
    assign dma_tx_req_o = dma_tx_en_i && tx_trig_i;
    assign dma_rx_req_o = dma_rx_en_i && (rx_trig_i || rx_timeout_i);

`ifndef SYNTHESIS
    // A request must never outlive its cause. This is the property that a
    // latched implementation violates, and it is checkable in simulation
    // because the cause is an input to this module.
    always @(posedge clk) if (rst_n) begin
        a_tx_no_latch: assert (dma_tx_req_o == (dma_tx_en_i && tx_trig_i))
            else $error("dma_tx_req_o does not track its cause");
        a_rx_no_latch: assert (dma_rx_req_o ==
                               (dma_rx_en_i && (rx_trig_i || rx_timeout_i)))
            else $error("dma_rx_req_o does not track its cause");
    end
`endif
endmodule

The SystemVerilog listing carries two immediate assertions stating that each request equals its cause. They look redundant — the assignment immediately above says the same thing — and they are not, because what they defend against is a future edit. The line someone adds in six months to "fix" a marginal timing path by registering the request breaks the assertion on the first simulation.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
//===========================================================================
//  uart_dma_if_v — Synthesizable Verilog-2001
//
//  Chapter 13.4. The entire DMA interface, and its smallness is the point.
//
//  Two levels, each an enable ANDed with a condition the design had already
//  computed. No new state, no new logic, no protocol.
//
//  THERE IS NO ACKNOWLEDGE LINE, and that is deliberate. A DMA engine
//  servicing a request performs bus accesses to DATA, and those accesses
//  are what change the condition. A request line needing an explicit
//  acknowledgement would be a handshake on top of a handshake.
//
//  Same reasoning as Chapter 13.3: a request is a CONDITION, not an event.
//  Cleared by servicing, never by acknowledging, and never latched -- a
//  latched request survives the transfer that satisfied it, and the engine
//  then reads a queue with nothing in it.
//
//  The receive request includes the idle timeout as well as the trigger,
//  which is what rescues the RESIDUAL BYTES at the end of a transfer
//  (Chapter 13.4 section 4).
//
//  Verilog-2001 note: the SystemVerilog listing carries two immediate
//  assertions stating that each request tracks its cause. Verilog-2001 has
//  no assertion construct at all, so that property lives only in the
//  testbench here. It is the same property, checked in the same place the
//  simulator can see it -- but the design file no longer states it, which
//  is a real loss of self-documentation rather than a stylistic one.
//===========================================================================
module uart_dma_if_v (
    input  wire dma_tx_en_i,      // DMA_CTRL bit 0
    input  wire dma_rx_en_i,      // DMA_CTRL bit 1

    input  wire tx_trig_i,        // transmit queue has room
    input  wire rx_trig_i,        // receive queue has data worth collecting
    input  wire rx_timeout_i,     // ...or has had some for a while

    output wire dma_tx_req_o,
    output wire dma_rx_req_o
);
    assign dma_tx_req_o = dma_tx_en_i && tx_trig_i;
    assign dma_rx_req_o = dma_rx_en_i && (rx_trig_i || rx_timeout_i);
endmodule

Verilog-2001 has no assertion construct at all, so that property can only live in the testbench. The design file no longer states its own invariant, which is a real loss rather than a stylistic one: the next person to edit the module reads the module, not the testbench.

VHDL gets the better of the two, and for a reason specific to the language. Its assertions here are concurrent, not clocked — so the claim is that the outputs track their cause at every instant, which is exactly right for combinational logic. A clocked assertion only checks at sample points, and a request that glitched between two edges would pass it.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
--===========================================================================
--  uart_dma_if — Synthesizable VHDL-2008
--
--  Chapter 13.4. The entire DMA interface, and its smallness is the point.
--
--  Two levels, each an enable ANDed with a condition the design had already
--  computed. No new state, no new logic, no protocol.
--
--  THERE IS NO ACKNOWLEDGE LINE, and that is deliberate. A DMA engine
--  servicing a request performs bus accesses to DATA, and those accesses
--  are what change the condition. A request line needing an explicit
--  acknowledgement would be a handshake on top of a handshake.
--
--  Same reasoning as Chapter 13.3: a request is a CONDITION, not an event.
--  Cleared by servicing, never by acknowledging, and never latched -- a
--  latched request survives the transfer that satisfied it, and the engine
--  then reads a queue with nothing in it.
--
--  The receive request includes the idle timeout as well as the trigger,
--  which is what rescues the RESIDUAL BYTES at the end of a transfer
--  (Chapter 13.4 section 4). Without it a message whose length is not a
--  multiple of the trigger level leaves its tail stranded in the queue.
--
--  VHDL note: the property the SystemVerilog listing states with two
--  immediate assertions is stated here as two CONCURRENT assertions. They
--  are checked continuously rather than on a clock edge, which is arguably
--  the better fit -- these are combinational outputs and the claim is that
--  they track their cause at every instant, not merely at sample points.
--===========================================================================
library ieee;
use ieee.std_logic_1164.all;

entity uart_dma_if is
    port (
        dma_tx_en_i : in std_logic;     -- DMA_CTRL bit 0
        dma_rx_en_i : in std_logic;     -- DMA_CTRL bit 1

        tx_trig_i    : in std_logic;    -- transmit queue has room
        rx_trig_i    : in std_logic;    -- receive queue has data worth collecting
        rx_timeout_i : in std_logic;    -- ...or has had some for a while

        dma_tx_req_o : out std_logic;
        dma_rx_req_o : out std_logic
    );
end entity uart_dma_if;

architecture rtl of uart_dma_if is
    signal tx_req_s, rx_req_s : std_logic;
begin
    tx_req_s <= dma_tx_en_i and tx_trig_i;
    rx_req_s <= dma_rx_en_i and (rx_trig_i or rx_timeout_i);

    dma_tx_req_o <= tx_req_s;
    dma_rx_req_o <= rx_req_s;

    -- A request must never outlive its cause.
    assert not (tx_req_s /= (dma_tx_en_i and tx_trig_i))
        report "dma_tx_req_o does not track its cause" severity error;
    assert not (rx_req_s /= (dma_rx_en_i and (rx_trig_i or rx_timeout_i)))
        report "dma_rx_req_o does not track its cause" severity error;
end architecture rtl;

8. Testing a Module That Must Not Remember

Four lines of logic, and a testbench worth eighteen checks — because what makes this module correct is not what it does but what it refuses to do.

A latched request is the classic DMA integration bug. It survives the transfer that satisfied it, the engine comes back to a queue with nothing in it, and the symptom is either a spurious read or a transfer that never terminates. Section 5 of this chapter has the matching story from the RTL side: a pushed byte that was never handshaken, four hundred bytes moved out of nineteen.

The defect is invisible to any test that only ever asserts conditions. It appears the instant one is removed — so the suite is built around removals:

SequenceWhat it proves
enable, raise the trigger, drain below itthe request clears with no acknowledge
enable, raise room, fill above the triggerthe same, on the transmit side
a few bytes below the trigger, then the timeoutthe residual tail is collected
the timeout goes awayand the request goes with it
600 pseudo-random steps over all five inputsno latch of any depth, on either request
Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
//===========================================================================
//  tb_uart_dma_if — self-checking SystemVerilog testbench
//
//  The module under test is four lines long, and a testbench for it is
//  worth writing anyway -- because the thing that makes it correct is not
//  what it does but what it REFUSES to do: it never remembers a request.
//
//  A latched request is the classic DMA bug. It survives the transfer that
//  satisfied it, and the engine comes back to read a queue with nothing in
//  it. The defect is invisible in any test that only ever asserts
//  conditions; it appears the instant one is REMOVED.
//
//  So the central check here is a whole-run invariant -- each request is
//  exactly its cause, on every clock of a pseudo-random soak -- and the
//  narrative tests are about the two cases that motivated the design: an
//  engine servicing a request, and the residual bytes at the end of a
//  transfer.
//===========================================================================
`timescale 1ns/1ps

module tb_uart_dma_if;

    logic clk = 1'b0;                 // the DUT is combinational; this clock
    always #5 clk = ~clk;           // exists so the observer can sample
    logic rst_n = 1'b0;

    logic dma_tx_en = 1'b0, dma_rx_en = 1'b0;
    logic tx_trig = 1'b0, rx_trig = 1'b0, rx_tmo = 1'b0;
    wire dma_tx_req, dma_rx_req;

    uart_dma_if dut (
        // clk/rst_n exist on the SystemVerilog listing only: they clock the two
        // immediate assertions inside it. Leaving them unconnected compiles
        // cleanly and silently disables those assertions, which is exactly the
        // kind of hole that makes an assertion-based claim worthless.
        .clk(clk), .rst_n(rst_n),
        .dma_tx_en_i(dma_tx_en), .dma_rx_en_i(dma_rx_en),
        .tx_trig_i(tx_trig), .rx_trig_i(rx_trig), .rx_timeout_i(rx_tmo),
        .dma_tx_req_o(dma_tx_req), .dma_rx_req_o(dma_rx_req));

    // ---- THE invariant: a request never outlives its cause ---------------
    int tx_bad = 0, rx_bad = 0, obs = 0, tx_high = 0, rx_high = 0;
    always @(posedge clk) if (rst_n) begin
        obs++;
        if (dma_tx_req !== (dma_tx_en && tx_trig))              tx_bad++;
        if (dma_rx_req !== (dma_rx_en && (rx_trig || rx_tmo)))  rx_bad++;
        if (dma_tx_req) tx_high++;
        if (dma_rx_req) rx_high++;
    end

    int checks = 0, failures = 0;
    task automatic check(input logic cond, input string name);
        checks++;
        if (cond) $display("  PASS %0s", name);
        else begin failures++; $display("  FAIL %0s", name); end
    endtask

    task automatic settle; repeat (2) @(negedge clk); endtask

    int i;
    logic [15:0] lfsr = 16'hBEEF;

    initial begin
        #500_000;
        $display("  FAIL watchdog: simulation did not finish");
        $display("== %0d checks, %0d failures ==", checks+1, failures+1);
        $display("   RESULT: SYSTEMVERILOG DMA TESTS FAILED (timeout)");
        $finish;
    end

    initial begin
        $display("== uart_dma_if : self-checking SystemVerilog testbench ==");
        rst_n = 1'b0; settle; rst_n = 1'b1; settle;

        //=== disabled means silent ==========================================
        @(negedge clk) tx_trig = 1'b1; rx_trig = 1'b1; rx_tmo = 1'b1; settle;
        check(dma_tx_req === 1'b0 && dma_rx_req === 1'b0,
              "every condition true but DMA disabled: no request at all");

        //=== the enables are independent =====================================
        @(negedge clk) dma_tx_en = 1'b1; settle;
        check(dma_tx_req === 1'b1 && dma_rx_req === 1'b0,
              "enabling TX raises only the TX request");
        @(negedge clk) dma_rx_en = 1'b1; settle;
        check(dma_rx_req === 1'b1, "enabling RX raises the RX request too");
        @(negedge clk) dma_tx_en = 1'b0; settle;
        check(dma_tx_req === 1'b0 && dma_rx_req === 1'b1,
              "and disabling one does not touch the other");
        @(negedge clk) dma_tx_en = 1'b1; settle;

        //=== an engine servicing a request ===================================
        // The acknowledge IS the transfer: the engine's reads change the
        // condition, and the request follows. There is no ack line because
        // there is nothing an ack would add.
        @(negedge clk) rx_tmo = 1'b0; rx_trig = 1'b1; settle;
        check(dma_rx_req === 1'b1, "queue above trigger: the engine is asked to collect");
        @(negedge clk) rx_trig = 1'b0; settle;       // the engine drained it
        check(dma_rx_req === 1'b0,
              "draining below the trigger clears the request -- no acknowledge needed");

        @(negedge clk) tx_trig = 1'b1; settle;
        check(dma_tx_req === 1'b1, "queue has room: the engine is asked to fill it");
        @(negedge clk) tx_trig = 1'b0; settle;
        check(dma_tx_req === 1'b0, "filling it above the trigger clears the request");

        //=== the residual bytes ==============================================
        // A message whose length is not a multiple of the trigger level
        // leaves a tail below the threshold. Without the timeout in the
        // request expression, that tail is stranded: the queue holds real
        // data and the engine is never asked for it.
        @(negedge clk) rx_trig = 1'b0; rx_tmo = 1'b0; settle;
        check(dma_rx_req === 1'b0, "a few bytes below the trigger raise nothing yet");
        @(negedge clk) rx_tmo = 1'b1; settle;
        check(dma_rx_req === 1'b1,
              "the idle timeout rescues them -- the residual tail is collected");
        @(negedge clk) rx_tmo = 1'b0; settle;
        check(dma_rx_req === 1'b0, "and once collected the request goes away");

        //=== trigger and timeout are an OR, not an AND ========================
        @(negedge clk) rx_trig = 1'b1; rx_tmo = 1'b0; settle;
        check(dma_rx_req === 1'b1, "trigger alone is enough");
        @(negedge clk) rx_trig = 1'b0; rx_tmo = 1'b1; settle;
        check(dma_rx_req === 1'b1, "timeout alone is enough");
        @(negedge clk) rx_trig = 1'b1; rx_tmo = 1'b1; settle;
        check(dma_rx_req === 1'b1, "and both together is still just a request");

        //=== a pseudo-random soak ============================================
        for (i = 0; i < 600; i++) begin
            @(negedge clk);
            lfsr = {lfsr[14:0], lfsr[15]^lfsr[13]^lfsr[12]^lfsr[10]};
            dma_tx_en = lfsr[0]; dma_rx_en = lfsr[1];
            tx_trig   = lfsr[2]; rx_trig   = lfsr[3]; rx_tmo = lfsr[4];
        end
        settle;

        //=== whole-run invariants =============================================
        check(obs > 600,     "the invariant observer ran on enough clocks to judge");
        check(tx_high > 50 && rx_high > 50,
              "both requests were asserted often enough for their removal to matter");
        check(tx_bad == 0,
              "the TX request was ALWAYS exactly its cause -- never once latched");
        check(rx_bad == 0, "and so was the RX request");

        $display("== %0d checks, %0d failures ==", checks, failures);
        if (failures == 0) $display("   RESULT: ALL SYSTEMVERILOG DMA TESTS PASSED");
        else               $display("   RESULT: SYSTEMVERILOG DMA TESTS FAILED");
        $finish;
    end
endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
//===========================================================================
//  tb_uart_dma_if_v — self-checking Verilog-2001 testbench
//
//  The module under test is four lines long, and a testbench for it is
//  worth writing anyway -- because the thing that makes it correct is not
//  what it does but what it REFUSES to do: it never remembers a request.
//
//  A latched request is the classic DMA bug. It survives the transfer that
//  satisfied it, and the engine comes back to read a queue with nothing in
//  it. The defect is invisible in any test that only ever asserts
//  conditions; it appears the instant one is REMOVED.
//
//  So the central check here is a whole-run invariant -- each request is
//  exactly its cause, on every clock of a pseudo-random soak -- and the
//  narrative tests are about the two cases that motivated the design: an
//  engine servicing a request, and the residual bytes at the end of a
//  transfer.
//===========================================================================
`timescale 1ns/1ps

module tb_uart_dma_if_v;

    reg clk = 1'b0;                 // the DUT is combinational; this clock
    always #5 clk = ~clk;           // exists so the observer can sample
    reg rst_n = 1'b0;

    reg dma_tx_en = 1'b0, dma_rx_en = 1'b0;
    reg tx_trig = 1'b0, rx_trig = 1'b0, rx_tmo = 1'b0;
    wire dma_tx_req, dma_rx_req;

    uart_dma_if_v dut (
        .dma_tx_en_i(dma_tx_en), .dma_rx_en_i(dma_rx_en),
        .tx_trig_i(tx_trig), .rx_trig_i(rx_trig), .rx_timeout_i(rx_tmo),
        .dma_tx_req_o(dma_tx_req), .dma_rx_req_o(dma_rx_req));

    // ---- THE invariant: a request never outlives its cause ---------------
    integer tx_bad = 0, rx_bad = 0, obs = 0, tx_high = 0, rx_high = 0;
    always @(posedge clk) if (rst_n) begin
        obs = obs + 1;
        if (dma_tx_req !== (dma_tx_en && tx_trig))              tx_bad = tx_bad + 1;
        if (dma_rx_req !== (dma_rx_en && (rx_trig || rx_tmo)))  rx_bad = rx_bad + 1;
        if (dma_tx_req) tx_high = tx_high + 1;
        if (dma_rx_req) rx_high = rx_high + 1;
    end

    integer checks = 0, failures = 0;
    task check;
        input cond;
        input [8*80-1:0] name;
        begin
            checks = checks + 1;
            if (cond) $display("  PASS %0s", name);
            else begin failures = failures + 1; $display("  FAIL %0s", name); end
        end
    endtask

    task settle; begin repeat (2) @(negedge clk); end endtask

    integer i;
    reg [15:0] lfsr = 16'hBEEF;

    initial begin
        #500_000;
        $display("  FAIL watchdog: simulation did not finish");
        $display("== %0d checks, %0d failures ==", checks+1, failures+1);
        $display("   RESULT: VERILOG DMA TESTS FAILED (timeout)");
        $finish;
    end

    initial begin
        $display("== uart_dma_if_v : self-checking Verilog testbench ==");
        rst_n = 1'b0; settle; rst_n = 1'b1; settle;

        //=== disabled means silent ==========================================
        @(negedge clk) tx_trig = 1'b1; rx_trig = 1'b1; rx_tmo = 1'b1; settle;
        check(dma_tx_req === 1'b0 && dma_rx_req === 1'b0,
              "every condition true but DMA disabled: no request at all");

        //=== the enables are independent =====================================
        @(negedge clk) dma_tx_en = 1'b1; settle;
        check(dma_tx_req === 1'b1 && dma_rx_req === 1'b0,
              "enabling TX raises only the TX request");
        @(negedge clk) dma_rx_en = 1'b1; settle;
        check(dma_rx_req === 1'b1, "enabling RX raises the RX request too");
        @(negedge clk) dma_tx_en = 1'b0; settle;
        check(dma_tx_req === 1'b0 && dma_rx_req === 1'b1,
              "and disabling one does not touch the other");
        @(negedge clk) dma_tx_en = 1'b1; settle;

        //=== an engine servicing a request ===================================
        // The acknowledge IS the transfer: the engine's reads change the
        // condition, and the request follows. There is no ack line because
        // there is nothing an ack would add.
        @(negedge clk) rx_tmo = 1'b0; rx_trig = 1'b1; settle;
        check(dma_rx_req === 1'b1, "queue above trigger: the engine is asked to collect");
        @(negedge clk) rx_trig = 1'b0; settle;       // the engine drained it
        check(dma_rx_req === 1'b0,
              "draining below the trigger clears the request -- no acknowledge needed");

        @(negedge clk) tx_trig = 1'b1; settle;
        check(dma_tx_req === 1'b1, "queue has room: the engine is asked to fill it");
        @(negedge clk) tx_trig = 1'b0; settle;
        check(dma_tx_req === 1'b0, "filling it above the trigger clears the request");

        //=== the residual bytes ==============================================
        // A message whose length is not a multiple of the trigger level
        // leaves a tail below the threshold. Without the timeout in the
        // request expression, that tail is stranded: the queue holds real
        // data and the engine is never asked for it.
        @(negedge clk) rx_trig = 1'b0; rx_tmo = 1'b0; settle;
        check(dma_rx_req === 1'b0, "a few bytes below the trigger raise nothing yet");
        @(negedge clk) rx_tmo = 1'b1; settle;
        check(dma_rx_req === 1'b1,
              "the idle timeout rescues them -- the residual tail is collected");
        @(negedge clk) rx_tmo = 1'b0; settle;
        check(dma_rx_req === 1'b0, "and once collected the request goes away");

        //=== trigger and timeout are an OR, not an AND ========================
        @(negedge clk) rx_trig = 1'b1; rx_tmo = 1'b0; settle;
        check(dma_rx_req === 1'b1, "trigger alone is enough");
        @(negedge clk) rx_trig = 1'b0; rx_tmo = 1'b1; settle;
        check(dma_rx_req === 1'b1, "timeout alone is enough");
        @(negedge clk) rx_trig = 1'b1; rx_tmo = 1'b1; settle;
        check(dma_rx_req === 1'b1, "and both together is still just a request");

        //=== a pseudo-random soak ============================================
        for (i = 0; i < 600; i = i + 1) begin
            @(negedge clk);
            lfsr = {lfsr[14:0], lfsr[15]^lfsr[13]^lfsr[12]^lfsr[10]};
            dma_tx_en = lfsr[0]; dma_rx_en = lfsr[1];
            tx_trig   = lfsr[2]; rx_trig   = lfsr[3]; rx_tmo = lfsr[4];
        end
        settle;

        //=== whole-run invariants =============================================
        check(obs > 600,     "the invariant observer ran on enough clocks to judge");
        check(tx_high > 50 && rx_high > 50,
              "both requests were asserted often enough for their removal to matter");
        check(tx_bad == 0,
              "the TX request was ALWAYS exactly its cause -- never once latched");
        check(rx_bad == 0, "and so was the RX request");

        $display("== %0d checks, %0d failures ==", checks, failures);
        if (failures == 0) $display("   RESULT: ALL VERILOG DMA TESTS PASSED");
        else               $display("   RESULT: VERILOG DMA TESTS FAILED");
        $finish;
    end
endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
--===========================================================================
--  tb_uart_dma_if — self-checking VHDL-2008 testbench
--
--  The module under test is four lines long, and a testbench for it is
--  worth writing anyway -- because the thing that makes it correct is not
--  what it does but what it REFUSES to do: it never remembers a request.
--
--  A latched request is the classic DMA bug. It survives the transfer that
--  satisfied it, and the engine comes back to read a queue with nothing in
--  it. The defect is invisible in any test that only ever asserts
--  conditions; it appears the instant one is REMOVED.
--
--  So the central check here is a whole-run invariant -- each request is
--  exactly its cause, on every clock of a pseudo-random soak -- and the
--  narrative tests are about the two cases that motivated the design: an
--  engine servicing a request, and the residual bytes at the end of a
--  transfer.
--
--  Same 18 counted checks as the Verilog and SystemVerilog twins.
--===========================================================================
library ieee;
use ieee.std_logic_1164.all;

entity tb_uart_dma_if is
end entity tb_uart_dma_if;

architecture sim of tb_uart_dma_if is

    constant TCLK : time := 10 ns;

    -- The DUT is purely combinational; this clock exists so the invariant
    -- observer has something to sample on.
    signal clk   : std_logic := '0';
    signal rst_n : std_logic := '0';
    signal done  : boolean   := false;

    signal dma_tx_en, dma_rx_en : std_logic := '0';
    signal tx_trig, rx_trig, rx_tmo : std_logic := '0';
    signal dma_tx_req, dma_rx_req : std_logic;

    signal tx_bad, rx_bad, obs, tx_high, rx_high : natural := 0;

begin

    clk <= not clk after TCLK/2 when not done else '0';

    dut : entity work.uart_dma_if
        port map (dma_tx_en_i => dma_tx_en, dma_rx_en_i => dma_rx_en,
                  tx_trig_i => tx_trig, rx_trig_i => rx_trig,
                  rx_timeout_i => rx_tmo,
                  dma_tx_req_o => dma_tx_req, dma_rx_req_o => dma_rx_req);

    -- THE invariant: a request never outlives its cause.
    invariant_obs : process (clk)
    begin
        if rising_edge(clk) and rst_n = '1' then
            obs <= obs + 1;
            if dma_tx_req /= (dma_tx_en and tx_trig) then
                tx_bad <= tx_bad + 1;
                report "dma_tx_req outlived its cause" severity error;
            end if;
            if dma_rx_req /= (dma_rx_en and (rx_trig or rx_tmo)) then
                rx_bad <= rx_bad + 1;
                report "dma_rx_req outlived its cause" severity error;
            end if;
            if dma_tx_req = '1' then tx_high <= tx_high + 1; end if;
            if dma_rx_req = '1' then rx_high <= rx_high + 1; end if;
        end if;
    end process invariant_obs;

    watchdog : process
    begin
        wait for 500 us;
        report "watchdog: simulation did not finish" severity failure;
    end process watchdog;

    stim : process
        variable checks, failures : natural := 0;
        variable lfsr : std_logic_vector(15 downto 0) := x"BEEF";

        procedure check(cond : boolean; name : string) is
        begin
            checks := checks + 1;
            if cond then
                report "  PASS " & name severity note;
            else
                failures := failures + 1;
                report "  FAIL " & name severity error;
            end if;
        end procedure check;

        procedure settle is
        begin
            for i in 1 to 2 loop wait until falling_edge(clk); end loop;
        end procedure settle;
    begin
        report "== uart_dma_if : self-checking VHDL testbench ==" severity note;
        rst_n <= '0'; settle; rst_n <= '1'; settle;

        --=== disabled means silent ==========================================
        wait until falling_edge(clk);
        tx_trig <= '1'; rx_trig <= '1'; rx_tmo <= '1';
        settle;
        check(dma_tx_req = '0' and dma_rx_req = '0',
              "every condition true but DMA disabled: no request at all");

        --=== the enables are independent =====================================
        wait until falling_edge(clk); dma_tx_en <= '1'; settle;
        check(dma_tx_req = '1' and dma_rx_req = '0',
              "enabling TX raises only the TX request");
        wait until falling_edge(clk); dma_rx_en <= '1'; settle;
        check(dma_rx_req = '1', "enabling RX raises the RX request too");
        wait until falling_edge(clk); dma_tx_en <= '0'; settle;
        check(dma_tx_req = '0' and dma_rx_req = '1',
              "and disabling one does not touch the other");
        wait until falling_edge(clk); dma_tx_en <= '1'; settle;

        --=== an engine servicing a request ===================================
        wait until falling_edge(clk); rx_tmo <= '0'; rx_trig <= '1'; settle;
        check(dma_rx_req = '1',
              "queue above trigger: the engine is asked to collect");
        wait until falling_edge(clk); rx_trig <= '0'; settle;
        check(dma_rx_req = '0',
              "draining below the trigger clears the request -- no acknowledge needed");

        wait until falling_edge(clk); tx_trig <= '1'; settle;
        check(dma_tx_req = '1', "queue has room: the engine is asked to fill it");
        wait until falling_edge(clk); tx_trig <= '0'; settle;
        check(dma_tx_req = '0', "filling it above the trigger clears the request");

        --=== the residual bytes ==============================================
        wait until falling_edge(clk); rx_trig <= '0'; rx_tmo <= '0'; settle;
        check(dma_rx_req = '0', "a few bytes below the trigger raise nothing yet");
        wait until falling_edge(clk); rx_tmo <= '1'; settle;
        check(dma_rx_req = '1',
              "the idle timeout rescues them -- the residual tail is collected");
        wait until falling_edge(clk); rx_tmo <= '0'; settle;
        check(dma_rx_req = '0', "and once collected the request goes away");

        --=== trigger and timeout are an OR, not an AND ========================
        wait until falling_edge(clk); rx_trig <= '1'; rx_tmo <= '0'; settle;
        check(dma_rx_req = '1', "trigger alone is enough");
        wait until falling_edge(clk); rx_trig <= '0'; rx_tmo <= '1'; settle;
        check(dma_rx_req = '1', "timeout alone is enough");
        wait until falling_edge(clk); rx_trig <= '1'; rx_tmo <= '1'; settle;
        check(dma_rx_req = '1', "and both together is still just a request");

        --=== a pseudo-random soak ============================================
        for i in 1 to 600 loop
            wait until falling_edge(clk);
            lfsr := lfsr(14 downto 0)
                  & (lfsr(15) xor lfsr(13) xor lfsr(12) xor lfsr(10));
            dma_tx_en <= lfsr(0); dma_rx_en <= lfsr(1);
            tx_trig   <= lfsr(2); rx_trig   <= lfsr(3); rx_tmo <= lfsr(4);
        end loop;
        settle;

        --=== whole-run invariants =============================================
        check(obs > 600, "the invariant observer ran on enough clocks to judge");
        check(tx_high > 50 and rx_high > 50,
              "both requests were asserted often enough for their removal to matter");
        check(tx_bad = 0,
              "the TX request was ALWAYS exactly its cause -- never once latched");
        check(rx_bad = 0, "and so was the RX request");

        report "== " & integer'image(checks) & " checks, "
                     & integer'image(failures) & " failures ==" severity note;
        if failures = 0 then
            report "   RESULT: ALL VHDL DMA TESTS PASSED" severity note;
        else
            report "   RESULT: VHDL DMA TESTS FAILED" severity error;
        end if;
        done <= true;
        wait;
    end process stim;

end architecture sim;

Eighteen checks, and all three languages agree:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
  PASS every condition true but DMA disabled: no request at all
  PASS enabling TX raises only the TX request
  PASS and disabling one does not touch the other
  PASS draining below the trigger clears the request -- no acknowledge needed
  PASS filling it above the trigger clears the request
  PASS a few bytes below the trigger raise nothing yet
  PASS the idle timeout rescues them -- the residual tail is collected
  PASS and once collected the request goes away
  PASS trigger alone is enough
  PASS timeout alone is enough
  PASS both requests were asserted often enough for their removal to matter
  PASS the TX request was ALWAYS exactly its cause -- never once latched
== 18 checks, 0 failures ==

Verilog-2001    : 18 checks, 0 failures
SystemVerilog   : 18 checks, 0 failures
VHDL-2008       : 18 checks, 0 failures

9. Verification

Model the engine, do not stub it. The defect in §5 was found because the model read at bus speed from a full queue. A stub that pops one byte when the request is high would not have found it.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// Assertion — a request is a level derived from live conditions, never
// latched. Same rule as the interrupt status, for the same reason.
property p_dma_req_is_transparent;
    @(posedge clk) disable iff (!rst_n)
        dma_rx_req_o == (dma_rx_en_q && (rx_trigger || rx_timeout_q));
endproperty

// Assertion — a push and its handshake are the SAME condition. This is the
// section 5 defect stated directly, and it is checkable by inspection too.
property p_push_is_the_handshake;
    @(posedge clk) disable iff (!rst_n)
        rx_push |-> (rx_char_valid && rx_char_ready);
endproperty

// Assertion — nothing is ever stranded. If the queue is non-empty and the
// line has been idle past the timeout, a request must be pending.
property p_no_stranded_bytes;
    @(posedge clk) disable iff (!rst_n)
        (!rx_empty && rx_timeout_q && dma_rx_en_q) |-> dma_rx_req_o;
endproperty

Test a message length that is not a multiple of the trigger. This is the single highest-value DMA test and it is trivially easy to omit, because round numbers are what people reach for. §4 used 11 bytes with a trigger of 8 precisely so there would be a remainder.

Test with the timeout disabled, and confirm the bytes really do strand. A test that only ever runs with the timeout enabled cannot tell whether the timeout is doing anything.

Count bus events, not just bytes. DMA's purpose is fewer, larger accesses; a test that confirms the data arrives says nothing about whether the mechanism is working. Two service events for eleven bytes is the result that matters.

10. Debugging

11. What This Means on an FPGA

The DMA interface costs two AND gates. Everything it needs already existed; only the enables are new.

The request must reach the engine as a level. If the interconnect converts it to a pulse, the residual mechanism breaks — the timeout asserts once and a missed pulse strands the data permanently.

Match the trigger to the burst size and say so in the driver. The two numbers live in different places — one in a UART register, one in the DMA engine's descriptor — and nothing enforces the relationship. It is worth a comment in both.

At 115,200 baud none of this is about throughput. A byte every 86.8 µs is trivial for any processor to service by interrupt. DMA on a slow UART is about not taking an exception every 86.8 µs in a system that has real-time work to do, which is a latency and jitter argument rather than a bandwidth one. At 3 Mbaud the bandwidth argument appears as well.

12. Understanding Check

13. Summary

The UART's DMA interface is two request levels and two enable bits, each an enable ANDed with a condition that already existed. The transfer itself is the acknowledgement.

Requests are levels, never latched — the same rule and the same reasoning as Chapter 13.3's interrupt status.

Under DMA the trigger should match the engine's burst size. Measured: sixteen bytes moved in two bus events with zero processor involvement, against sixteen exceptions for an interrupt-per-byte driver.

The residual-byte problem is the defining hazard. Eleven bytes with a trigger of eight left three stranded indefinitely — received correctly and unreachable. The receive-idle timeout gives the request a second cause, no more data coming, and recovers them.

The transmit side has no equivalent, because the engine knows the length. It does need busy rather than queue-empty to know the last character has left the wire.

A modelled DMA engine found a real defect that 186 checks across eleven testbenches had not: a receive push qualified on the offer rather than the handshake, which duplicated a held character indefinitely whenever a full queue was drained at bus speed — 400 bytes delivered from 19 sent.

And the rule that generalises: a push and its handshake must be the same condition, and a new master's access pattern must be modelled rather than assumed covered.

14. What Comes Next

Everything so far has assumed APB, which was chosen because it is the smallest bus with a real protocol and it makes the side-effect timing of Chapter 13.1 easy to state exactly.

Chapter 13.5 asks what changes on an AXI-class bus — outstanding transactions, separate address and data phases, bursts, and the uncomfortable question of what a peripheral with read side effects does when a master is allowed to speculate. It also draws the line this module has been observing throughout: what the UART must provide, and what belongs to the interconnect.

Browse the full path on the UART tutorials index. For the timeout this chapter depends on, read back to Chapter 13.3.

Continue learning

Where this fits

Part of the UART curriculum.