Skip to content
VLSI Mentor

SPI · Module 15

Handshake and FIFO Crossings

A FIFO works because the data never crosses a boundary at all — what crosses is two pointers, which are Gray coding's special case, and every stale reading is conservative rather than optimistic. Then the sizing question a conservation test never asks: what depth does an SPI burst need, and what to do when no depth is sufficient.

Chapter 15.1 measured what a four-phase handshake costs: it conserves every event, and only for a source that honours its busy signal. An SPI slave cannot — SCLK does not stop because the system side is busy.

Chapter 15.5 measured why the data cannot simply be synchronised, and showed that both working schemes carry one value at a time and both require the source to stop between values.

The source cannot be stopped and the data keeps coming. What is left?

A FIFO. Its job is not to be faster; it is to decouple, so that the source never waits for the destination and the only remaining question is one of depth.

1. How A FIFO Crosses The Boundary Without Violating Chapter 15.5

The data does not cross a boundary at all.

It is written into a dual-port memory by the source and read out of it by the destination, and no data bit is ever sampled by the other clock while it is changing — because a location is written once and read later, never both at once.

What does cross is the two pointers, and they are exactly the case Gray coding was made for: each increments by one, so consecutive values differ in a single bit, so a destination that catches an update part-way reads either the old pointer or the new one.

And here is the part that makes the design safe rather than merely probable:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   the WRITE pointer read late by the reader  makes the FIFO look EMPTIER
   the READ pointer read late by the writer   makes the FIFO look FULLER

Every stale reading is conservative. A reader that has not yet seen a write waits; a writer that has not yet seen a read holds off. The FIFO can transiently report less space than it has and never more — which is why full and empty are computed in different domains from different pointer pairs, rather than from one shared occupancy count.

2. The Block Diagram

A write side with a binary and Gray pointer writing a dual-port memory, a read side with its own pointers reading it, and each pointer crossing to the other domain through a two-flop synchroniser to compute full and emptywr_en, wr_datawrite pointerdual-port memoryfullSYNC_N flopsSYNC_N flopsread pointeremptyrd_en, rd_dataoverflowwritewrite refusedreadgates12
Figure 1 — the crossing. The memory is the only shared object and it is never accessed by both clocks at one location at one time, which is what keeps the data out of Chapter 15.5's problem. The two dashed paths are the pointers crossing, each Gray coded and each read STALE in the conservative direction — the write pointer late makes the FIFO look emptier and the read pointer late makes it look fuller, so neither flag is ever optimistic.

3. Data Conservation, Across Ratios

Thirty-two words per run, W = 8, an eight-deep-or-more FIFO with the reader draining continuously:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   wclk  rclk  words_read  wrong  overflow  underflow
     40    10          32      0         0          0
     14    10          32      0         0          0
     10    10          32      0         0          0
     10    14          32      0         0          0
     10    40          28      3         1          0

The first four rows are the Gray-coded pointer crossing working: every word arrives exactly once and in order, at ratios from 8:1 through 1:1 to 1:1.4.

The last row is not a crossing fault. The writer sends one word every two write cycles and the reader drains one per read cycle, so the reader keeps up while one read period is at most two write periods. At a 10 ns write clock against a 40 ns read clock it is not, the writer outruns the reader indefinitely, and no depth is enough.

A FIFO decouples. It does not create bandwidth.

Which is Chapter 15.1's conclusion in a different shape — and the shortfall is reported rather than hidden, which is the only honest thing available.

Ten cycles across five rows. A write pointer advances at cycle 2. Two synchroniser stages carry it into the read domain, arriving at cycle 4. The empty flag stays asserted until then, so the reader waits for two cycles during which the FIFO is not actually empty.the word is committedthe word is committedstill looks empty: reader waitsstill looks empty: readerwaitsempty clears -- late, never earlyempty clears -- late, neverearlyrclkwptr_gray0011111111sync stage 10000111111sync stage 20000001111rd_emptyt0t1t2t3t4t5t6t7t8t9
Figure 2 — why a stale pointer is safe. The writer commits a word at cycle 2, but the reader's synchronised copy of the write pointer does not show it until cycle 4. Between those the FIFO looks emptier than it is, so the reader waits — which costs latency and never correctness. A stale reading in the other direction makes it look fuller, so the writer holds off, and neither flag is ever optimistic.

4. The Flags Are Conservative, Not Exact

Two directed checks establish the direction of the error, which is the property that makes the design safe.

Sixty read attempts at an empty FIFO took nothing and set no underflow, because rd_en is gated on empty and empty errs towards staying asserted. A stale write pointer makes the FIFO look emptier, never fuller — so the reader can never take a word that was not written.

Forty words into a depth-2 FIFO with no reader overflowed and said so, and the word was dropped rather than corrupting one already held. An overflow is a reported loss, which is the only honest thing a FIFO can do when the source cannot be stalled.

5. The Sizing Question

This is the one a data-conservation test never asks, and the one a real design has to answer.

Eight words arrive while the reader is blocked — an interrupt, a lost arbitration, a software loop that took longer than usual:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   depth 2   overflow = 1
   depth 4   overflow = 1
   depth 16  overflow = 0

And once the reader resumed, the depth-16 FIFO returned all eight words in order. So a blocked reader costs latency and nothing else, provided the depth covers the block.

The rule the experiment confirms:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   DEPTH >= (words the source produces while the destination is not draining)

For an SPI slave that is (blocked time) / (LEN × SCLK period). A 4 KB burst at 8-bit frames and a 10 MHz SPI clock produces a word every 800 ns, so a 100 µs interrupt latency needs 125 entries.

6. Building the FIFO — Three HDLs

The circuit

A dual-port memory, two binary pointers, two Gray pointers, two synchroniser pairs, and two sticky flags. The full condition decodes the synchronised Gray read pointer rather than comparing Gray codes directly, for a reason worth reading in the code.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_cdc_fifo.sv — a dual-port memory, two Gray pointers, and flags whose errors are deliberately one-directional
// spi_cdc_fifo.sv
//
// Chapter 15.6 -- moving whole frames across the boundary, and sizing the crossing.
//
// Chapter 15.1 measured what a four-phase handshake costs: it conserves every event,
// and only for a source that honours its busy signal. An SPI slave cannot. SCLK does
// not stop because the system side is busy, so a handshake's guarantee is
// conditional on a freedom the slave does not have.
//
// Chapter 15.5 measured why the data cannot simply be synchronised: N independent
// synchronisers on a bus observe values the source never held, and the two schemes
// that work -- Gray code and data-plus-a-flag -- each carry exactly one value at a
// time and each require the source to stop between values.
//
// What is left, and what every real SPI slave uses, is a FIFO. Its job is not to be
// faster; it is to DECOUPLE, so that the source never has to wait for the
// destination and the only remaining question is one of DEPTH.
//
// HOW A FIFO CROSSES THE BOUNDARY WITHOUT VIOLATING CHAPTER 15.5.
//
// The data does not cross a boundary at all. It is written into a dual-port memory
// by the source and read out of it by the destination, and no data bit is ever
// sampled by the other clock while it is changing -- because a location is written
// once and read later, never both at once.
//
// What DOES cross is the two POINTERS, and they are exactly the case Gray code was
// made for: each one increments by one, so consecutive values differ in a single
// bit, so a destination that catches an update part-way reads either the old pointer
// or the new one. Both are SAFE, in a specific and slightly surprising way:
//
//   the WRITE pointer read late by the reader makes the FIFO look EMPTIER
//   the READ pointer read late by the writer makes the FIFO look FULLER
//
// So every stale reading is CONSERVATIVE. A reader that has not yet seen a write
// waits; a writer that has not yet seen a read holds off. The FIFO can transiently
// report less space than it has and never more, which is what makes the design safe
// rather than merely probable -- and it is the reason `full` and `empty` are computed
// in different domains from different pointer pairs.
//
// THE POINTERS ARE ONE BIT WIDER THAN THE ADDRESS.
//
// With DEPTH = 2^AW entries and AW-bit pointers, full and empty are the same
// condition. The extra bit distinguishes them: equal pointers including the top bit
// means EMPTY, and equal low bits with a differing top bit means FULL. That is the
// standard construction and the reason is worth stating because the alternative --
// a separate occupancy counter -- would be a value that BOTH domains increment, and
// there is no correct way to cross that.
//
// AND THE SIZING QUESTION, WHICH IS THE CHAPTER'S NUMBER.
//
// The source writes one word every LEN SCLK periods for the length of a burst. The
// destination drains one word per read, whenever it gets round to it. So:
//
//     DEPTH >= (words the source can produce) - (words the destination can drain
//              in the same time)
//
// For an SPI slave the worst case is a burst arriving while the destination is
// blocked -- an interrupt, a bus arbitration loss, a software loop that took longer
// than usual. If the destination can be blocked for T and the source produces a word
// every T_word, the FIFO needs T/T_word entries, and the honest answer for a system
// that cannot bound T is that no depth is sufficient and the slave must be able to
// report the overflow. Which is why `overflow` is an output rather than an assertion.

module spi_cdc_fifo #(
    parameter int W  = 8,    // word width
    parameter int AW = 2     // address width; DEPTH = 2**AW
) (
    // --- the write side: the source's clock ---------------------------------
    input  wire          wclk,
    input  wire          wrst_n,
    input  wire          wr_en,
    input  wire [W-1:0]  wr_data,
    output wire          wr_full,
    // Sticky: a write arrived with no room. The word is LOST, and a FIFO that
    // dropped it silently would make an unbounded destination stall look like
    // corrupt data.
    output reg           wr_overflow,

    // --- the read side: the destination's clock -----------------------------
    input  wire          rclk,
    input  wire          rrst_n,
    input  wire          rd_en,
    output wire [W-1:0]  rd_data,
    output wire          rd_empty,
    // Sticky: a read arrived with nothing in the FIFO.
    output reg           rd_underflow
);

    localparam int DEPTH = 1 << AW;

    // The memory. Written by one clock, read by the other, and never both at the
    // same location at the same time -- which is what keeps the DATA out of
    // Chapter 15.5's problem entirely.
    reg [W-1:0] mem [0:DEPTH-1];

    // Pointers are AW+1 bits: the extra bit is what distinguishes full from empty.
    reg [AW:0] wptr_bin, wptr_gray;
    reg [AW:0] rptr_bin, rptr_gray;

    // Each pointer, synchronised into the OTHER domain. Two flops each, and Gray
    // coded, which is the only reason a multi-bit value may be synchronised here.
    reg [AW:0] rptr_gray_w1, rptr_gray_w2;   // read pointer, seen by the writer
    reg [AW:0] wptr_gray_r1, wptr_gray_r2;   // write pointer, seen by the reader

    function automatic [AW:0] bin_to_gray(input [AW:0] b);
        bin_to_gray = b ^ (b >> 1);
    endfunction

    // Gray back to binary, used in the WRITE domain on the synchronised read
    // pointer. The textbook full condition compares Gray codes directly:
    //
    //     full = (wgray_next == {~rgray_sync[AW:AW-1], rgray_sync[AW-2:0]})
    //
    // which is two gates cheaper and does not generalise: at AW = 1 the low slice
    // `rgray_sync[-1:0]` does not exist, and a depth-2 FIFO is a perfectly
    // reasonable thing to build. Decoding is AW XORs on a path with a full write
    // cycle available, and it is correct at every depth -- which is the better
    // trade for a parameterised block.
    function automatic [AW:0] gray_to_bin(input [AW:0] g);
        integer k;
        begin
            gray_to_bin[AW] = g[AW];
            for (k = AW-1; k >= 0; k = k - 1)
                gray_to_bin[k] = gray_to_bin[k+1] ^ g[k];
        end
    endfunction

    // --- the write side -----------------------------------------------------
    wire [AW:0] wptr_next = wptr_bin + 1'b1;

    // FULL: the next write pointer would have the same address as the read pointer
    // but the opposite top bit -- which is the whole reason the pointers are one bit
    // wider than the address. A stale read pointer makes this fire EARLY, which is
    // the conservative direction: the FIFO reports less room than it has, never
    // more.
    wire [AW:0] rptr_bin_w = gray_to_bin(rptr_gray_w2);
    assign wr_full = (wptr_next[AW-1:0] == rptr_bin_w[AW-1:0]) &&
                     (wptr_next[AW]     != rptr_bin_w[AW]);

    always_ff @(posedge wclk or negedge wrst_n) begin
        if (!wrst_n) begin
            wptr_bin     <= {(AW+1){1'b0}};
            wptr_gray    <= {(AW+1){1'b0}};
            rptr_gray_w1 <= {(AW+1){1'b0}};
            rptr_gray_w2 <= {(AW+1){1'b0}};
            wr_overflow  <= 1'b0;
        end else begin
            rptr_gray_w1 <= rptr_gray;
            rptr_gray_w2 <= rptr_gray_w1;

            if (wr_en) begin
                if (!wr_full) begin
                    mem[wptr_bin[AW-1:0]] <= wr_data;
                    wptr_bin  <= wptr_next;
                    wptr_gray <= bin_to_gray(wptr_next);
                end else begin
                    wr_overflow <= 1'b1;
                end
            end
        end
    end

    // --- the read side ------------------------------------------------------
    wire [AW:0] rptr_next = rptr_bin + 1'b1;

    // EMPTY: the pointers are equal including the top bit. Compared in GRAY,
    // against the synchronised write pointer -- and a stale write pointer makes
    // this stay asserted LONGER, which is again conservative.
    assign rd_empty = (rptr_gray == wptr_gray_r2);
    assign rd_data  = mem[rptr_bin[AW-1:0]];

    always_ff @(posedge rclk or negedge rrst_n) begin
        if (!rrst_n) begin
            rptr_bin     <= {(AW+1){1'b0}};
            rptr_gray    <= {(AW+1){1'b0}};
            wptr_gray_r1 <= {(AW+1){1'b0}};
            wptr_gray_r2 <= {(AW+1){1'b0}};
            rd_underflow <= 1'b0;
        end else begin
            wptr_gray_r1 <= wptr_gray;
            wptr_gray_r2 <= wptr_gray_r1;

            if (rd_en) begin
                if (!rd_empty) begin
                    rptr_bin  <= rptr_next;
                    rptr_gray <= bin_to_gray(rptr_next);
                end else begin
                    rd_underflow <= 1'b1;
                end
            end
        end
    end

`ifdef SPI_CHECKS
    // The one structural property worth asserting, and it is about the POINTERS
    // rather than the data: a Gray pointer must never change by more than one bit
    // at a time. If it does, the encoding's whole guarantee is void and the
    // crossing is Chapter 15.5's scheme 1 wearing a Gray code's name.
    function automatic integer popcount(input [AW:0] v);
        integer b;
        begin
            popcount = 0;
            for (b = 0; b <= AW; b = b + 1) if (v[b]) popcount = popcount + 1;
        end
    endfunction

    // The history registers are RESET, and that is not a detail. Written as
    // `always_ff @(posedge clk) if (rst_n)`, the history keeps its pre-reset value
    // while the pointer it is compared against is cleared to zero -- so the first
    // check after reset release compares a stale pointer with a fresh one and fires
    // on a design that is behaving perfectly. A checker that reports a fault at
    // every reset gets disabled, and then it reports nothing at all.
    reg [AW:0] wg_d, rg_d;
    always_ff @(posedge wclk or negedge wrst_n) begin
        if (!wrst_n) begin
            wg_d <= {(AW+1){1'b0}};
        end else begin
            if (popcount(wptr_gray ^ wg_d) > 1)
                $fatal(1, "the write Gray pointer changed by more than one bit");
            wg_d <= wptr_gray;
        end
    end
    always_ff @(posedge rclk or negedge rrst_n) begin
        if (!rrst_n) begin
            rg_d <= {(AW+1){1'b0}};
        end else begin
            if (popcount(rptr_gray ^ rg_d) > 1)
                $fatal(1, "the read Gray pointer changed by more than one bit");
            rg_d <= rptr_gray;
        end
    end
`endif

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_cdc_fifo.v — the same design in Verilog-2001
// spi_cdc_fifo.v
//
// Chapter 15.6 -- moving whole frames across the boundary, and sizing the crossing.
//
// Chapter 15.1 measured what a four-phase handshake costs: it conserves every event,
// and only for a source that honours its busy signal. An SPI slave cannot. SCLK does
// not stop because the system side is busy, so a handshake's guarantee is
// conditional on a freedom the slave does not have.
//
// Chapter 15.5 measured why the data cannot simply be synchronised: N independent
// synchronisers on a bus observe values the source never held, and the two schemes
// that work -- Gray code and data-plus-a-flag -- each carry exactly one value at a
// time and each require the source to stop between values.
//
// What is left, and what every real SPI slave uses, is a FIFO. Its job is not to be
// faster; it is to DECOUPLE, so that the source never has to wait for the
// destination and the only remaining question is one of DEPTH.
//
// HOW A FIFO CROSSES THE BOUNDARY WITHOUT VIOLATING CHAPTER 15.5.
//
// The data does not cross a boundary at all. It is written into a dual-port memory
// by the source and read out of it by the destination, and no data bit is ever
// sampled by the other clock while it is changing -- because a location is written
// once and read later, never both at once.
//
// What DOES cross is the two POINTERS, and they are exactly the case Gray code was
// made for: each one increments by one, so consecutive values differ in a single
// bit, so a destination that catches an update part-way reads either the old pointer
// or the new one. Both are SAFE, in a specific and slightly surprising way:
//
//   the WRITE pointer read late by the reader makes the FIFO look EMPTIER
//   the READ pointer read late by the writer makes the FIFO look FULLER
//
// So every stale reading is CONSERVATIVE. A reader that has not yet seen a write
// waits; a writer that has not yet seen a read holds off. The FIFO can transiently
// report less space than it has and never more, which is what makes the design safe
// rather than merely probable -- and it is the reason `full` and `empty` are computed
// in different domains from different pointer pairs.
//
// THE POINTERS ARE ONE BIT WIDER THAN THE ADDRESS.
//
// With DEPTH = 2^AW entries and AW-bit pointers, full and empty are the same
// condition. The extra bit distinguishes them: equal pointers including the top bit
// means EMPTY, and equal low bits with a differing top bit means FULL. That is the
// standard construction and the reason is worth stating because the alternative --
// a separate occupancy counter -- would be a value that BOTH domains increment, and
// there is no correct way to cross that.
//
// AND THE SIZING QUESTION, WHICH IS THE CHAPTER'S NUMBER.
//
// The source writes one word every LEN SCLK periods for the length of a burst. The
// destination drains one word per read, whenever it gets round to it. So:
//
//     DEPTH >= (words the source can produce) - (words the destination can drain
//              in the same time)
//
// For an SPI slave the worst case is a burst arriving while the destination is
// blocked -- an interrupt, a bus arbitration loss, a software loop that took longer
// than usual. If the destination can be blocked for T and the source produces a word
// every T_word, the FIFO needs T/T_word entries, and the honest answer for a system
// that cannot bound T is that no depth is sufficient and the slave must be able to
// report the overflow. Which is why `overflow` is an output rather than an assertion.

module spi_cdc_fifo #(
    parameter W  = 8,    // word width
    parameter AW = 2     // address width; DEPTH = 2**AW
) (
    // --- the write side: the source's clock ---------------------------------
    input  wire          wclk,
    input  wire          wrst_n,
    input  wire          wr_en,
    input  wire [W-1:0]  wr_data,
    output wire          wr_full,
    // Sticky: a write arrived with no room. The word is LOST, and a FIFO that
    // dropped it silently would make an unbounded destination stall look like
    // corrupt data.
    output reg           wr_overflow,

    // --- the read side: the destination's clock -----------------------------
    input  wire          rclk,
    input  wire          rrst_n,
    input  wire          rd_en,
    output wire [W-1:0]  rd_data,
    output wire          rd_empty,
    // Sticky: a read arrived with nothing in the FIFO.
    output reg           rd_underflow
);

    localparam DEPTH = 1 << AW;

    // The memory. Written by one clock, read by the other, and never both at the
    // same location at the same time -- which is what keeps the DATA out of
    // Chapter 15.5's problem entirely.
    reg [W-1:0] mem [0:DEPTH-1];

    // Pointers are AW+1 bits: the extra bit is what distinguishes full from empty.
    reg [AW:0] wptr_bin, wptr_gray;
    reg [AW:0] rptr_bin, rptr_gray;

    // Each pointer, synchronised into the OTHER domain. Two flops each, and Gray
    // coded, which is the only reason a multi-bit value may be synchronised here.
    reg [AW:0] rptr_gray_w1, rptr_gray_w2;   // read pointer, seen by the writer
    reg [AW:0] wptr_gray_r1, wptr_gray_r2;   // write pointer, seen by the reader

        function [AW:0] bin_to_gray;
        input [AW:0] b;
        bin_to_gray = b ^ (b >> 1);
    endfunction

    // Gray back to binary, used in the WRITE domain on the synchronised read
    // pointer. The textbook full condition compares Gray codes directly:
    //
    //     full = (wgray_next == {~rgray_sync[AW:AW-1], rgray_sync[AW-2:0]})
    //
    // which is two gates cheaper and does not generalise: at AW = 1 the low slice
    // `rgray_sync[-1:0]` does not exist, and a depth-2 FIFO is a perfectly
    // reasonable thing to build. Decoding is AW XORs on a path with a full write
    // cycle available, and it is correct at every depth -- which is the better
    // trade for a parameterised block.
        function [AW:0] gray_to_bin;
        input [AW:0] g;
        integer k;
        begin
            gray_to_bin[AW] = g[AW];
            for (k = AW-1; k >= 0; k = k - 1)
                gray_to_bin[k] = gray_to_bin[k+1] ^ g[k];
        end
    endfunction

    // --- the write side -----------------------------------------------------
    wire [AW:0] wptr_next = wptr_bin + 1'b1;

    // FULL: the next write pointer would have the same address as the read pointer
    // but the opposite top bit -- which is the whole reason the pointers are one bit
    // wider than the address. A stale read pointer makes this fire EARLY, which is
    // the conservative direction: the FIFO reports less room than it has, never
    // more.
    wire [AW:0] rptr_bin_w = gray_to_bin(rptr_gray_w2);
    assign wr_full = (wptr_next[AW-1:0] == rptr_bin_w[AW-1:0]) &&
                     (wptr_next[AW]     != rptr_bin_w[AW]);

    always @(posedge wclk or negedge wrst_n) begin
        if (!wrst_n) begin
            wptr_bin     <= {(AW+1){1'b0}};
            wptr_gray    <= {(AW+1){1'b0}};
            rptr_gray_w1 <= {(AW+1){1'b0}};
            rptr_gray_w2 <= {(AW+1){1'b0}};
            wr_overflow  <= 1'b0;
        end else begin
            rptr_gray_w1 <= rptr_gray;
            rptr_gray_w2 <= rptr_gray_w1;

            if (wr_en) begin
                if (!wr_full) begin
                    mem[wptr_bin[AW-1:0]] <= wr_data;
                    wptr_bin  <= wptr_next;
                    wptr_gray <= bin_to_gray(wptr_next);
                end else begin
                    wr_overflow <= 1'b1;
                end
            end
        end
    end

    // --- the read side ------------------------------------------------------
    wire [AW:0] rptr_next = rptr_bin + 1'b1;

    // EMPTY: the pointers are equal including the top bit. Compared in GRAY,
    // against the synchronised write pointer -- and a stale write pointer makes
    // this stay asserted LONGER, which is again conservative.
    assign rd_empty = (rptr_gray == wptr_gray_r2);
    assign rd_data  = mem[rptr_bin[AW-1:0]];

    always @(posedge rclk or negedge rrst_n) begin
        if (!rrst_n) begin
            rptr_bin     <= {(AW+1){1'b0}};
            rptr_gray    <= {(AW+1){1'b0}};
            wptr_gray_r1 <= {(AW+1){1'b0}};
            wptr_gray_r2 <= {(AW+1){1'b0}};
            rd_underflow <= 1'b0;
        end else begin
            wptr_gray_r1 <= wptr_gray;
            wptr_gray_r2 <= wptr_gray_r1;

            if (rd_en) begin
                if (!rd_empty) begin
                    rptr_bin  <= rptr_next;
                    rptr_gray <= bin_to_gray(rptr_next);
                end else begin
                    rd_underflow <= 1'b1;
                end
            end
        end
    end

`ifdef SPI_CHECKS
    // The one structural property worth asserting, and it is about the POINTERS
    // rather than the data: a Gray pointer must never change by more than one bit
    // at a time. If it does, the encoding's whole guarantee is void and the
    // crossing is Chapter 15.5's scheme 1 wearing a Gray code's name.
        function integer popcount;
        input [AW:0] v;
        integer b;
        begin
            popcount = 0;
            for (b = 0; b <= AW; b = b + 1) if (v[b]) popcount = popcount + 1;
        end
    endfunction

    // The history registers are RESET, and that is not a detail. Written as
    // `always @(posedge clk) if (rst_n)`, the history keeps its pre-reset value
    // while the pointer it is compared against is cleared to zero -- so the first
    // check after reset release compares a stale pointer with a fresh one and fires
    // on a design that is behaving perfectly. A checker that reports a fault at
    // every reset gets disabled, and then it reports nothing at all.
    reg [AW:0] wg_d, rg_d;
    always @(posedge wclk or negedge wrst_n) begin
        if (!wrst_n) begin
            wg_d <= {(AW+1){1'b0}};
        end else begin
            if (popcount(wptr_gray ^ wg_d) > 1)
                $fatal(1, "the write Gray pointer changed by more than one bit");
            wg_d <= wptr_gray;
        end
    end
    always @(posedge rclk or negedge rrst_n) begin
        if (!rrst_n) begin
            rg_d <= {(AW+1){1'b0}};
        end else begin
            if (popcount(rptr_gray ^ rg_d) > 1)
                $fatal(1, "the read Gray pointer changed by more than one bit");
            rg_d <= rptr_gray;
        end
    end
`endif

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_cdc_fifo.vhd — the same design in VHDL
-- spi_cdc_fifo.vhd
--
-- Chapter 15.6 -- moving whole frames across the boundary, and sizing the crossing.
--
-- Chapter 15.1 measured what a four-phase handshake costs: it conserves every event,
-- and only for a source that honours its busy signal. An SPI slave cannot. SCLK does
-- not stop because the system side is busy, so a handshake's guarantee is
-- conditional on a freedom the slave does not have.
--
-- Chapter 15.5 measured why the data cannot simply be synchronised: N independent
-- synchronisers on a bus observe values the source never held, and the two schemes
-- that work -- Gray code and data-plus-a-flag -- each carry exactly one value at a
-- time and each require the source to stop between values.
--
-- What is left, and what every real SPI slave uses, is a FIFO. Its job is not to be
-- faster; it is to DECOUPLE, so that the source never has to wait for the
-- destination and the only remaining question is one of DEPTH.
--
-- HOW A FIFO CROSSES THE BOUNDARY WITHOUT VIOLATING CHAPTER 15.5.
--
-- The data does not cross a boundary at all. It is written into a dual-port memory
-- by the source and read out of it by the destination, and no data bit is ever
-- sampled by the other clock while it is changing -- because a location is written
-- once and read later, never both at once.
--
-- What DOES cross is the two POINTERS, and they are exactly the case Gray code was
-- made for: each one increments by one, so consecutive values differ in a single
-- bit, so a destination that catches an update part-way reads either the old pointer
-- or the new one. Both are SAFE, in a specific and slightly surprising way:
--
--   the WRITE pointer read late by the reader makes the FIFO look EMPTIER
--   the READ pointer read late by the writer makes the FIFO look FULLER
--
-- So every stale reading is CONSERVATIVE. A reader that has not yet seen a write
-- waits; a writer that has not yet seen a read holds off. The FIFO can transiently
-- report less space than it has and never more, which is what makes the design safe
-- rather than merely probable -- and it is the reason `full` and `empty` are computed
-- in different domains from different pointer pairs.
--
-- THE POINTERS ARE ONE BIT WIDER THAN THE ADDRESS.
--
-- With DEPTH = 2^AW entries and AW-bit pointers, full and empty are the same
-- condition. The extra bit distinguishes them: equal pointers including the top bit
-- means EMPTY, and equal low bits with a differing top bit means FULL. That is the
-- standard construction and the reason is worth stating because the alternative --
-- a separate occupancy counter -- would be a value that BOTH domains increment, and
-- there is no correct way to cross that.
--
-- AND THE SIZING QUESTION, WHICH IS THE CHAPTER'S NUMBER.
--
-- The source writes one word every LEN SCLK periods for the length of a burst. The
-- destination drains one word per read, whenever it gets round to it. So:
--
--     DEPTH >= (words the source can produce) - (words the destination can drain
--              in the same time)
--
-- For an SPI slave the worst case is a burst arriving while the destination is
-- blocked -- an interrupt, a bus arbitration loss, a software loop that took longer
-- than usual. If the destination can be blocked for T and the source produces a word
-- every T_word, the FIFO needs T/T_word entries, and the honest answer for a system
-- that cannot bound T is that no depth is sufficient and the slave must be able to
-- report the overflow. Which is why `overflow` is an output rather than an assertion.

library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

entity spi_cdc_fifo is
    generic (
        W  : positive := 8;   -- word width
        AW : positive := 2    -- address width; DEPTH = 2**AW
    );
    port (
        -- the write side: the source's clock
        wclk         : in  std_logic;
        wrst_n       : in  std_logic;
        wr_en        : in  std_logic;
        wr_data      : in  std_logic_vector(W - 1 downto 0);
        wr_full      : out std_logic;
        -- Sticky: a write arrived with no room. The word is LOST, and a FIFO that
        -- dropped it silently would make an unbounded destination stall look like
        -- corrupt data.
        wr_overflow  : out std_logic;

        -- the read side: the destination's clock
        rclk         : in  std_logic;
        rrst_n       : in  std_logic;
        rd_en        : in  std_logic;
        rd_data      : out std_logic_vector(W - 1 downto 0);
        rd_empty     : out std_logic;
        rd_underflow : out std_logic
    );
end entity;

architecture rtl of spi_cdc_fifo is

    constant DEPTH : positive := 2 ** AW;

    -- The memory. Written by one clock, read by the other, and never both at the
    -- same location at the same time -- which is what keeps the DATA out of
    -- Chapter 15.5's problem entirely.
    type mem_t is array (0 to DEPTH - 1) of std_logic_vector(W - 1 downto 0);
    signal mem : mem_t := (others => (others => '0'));

    -- Pointers are AW+1 bits: the extra bit distinguishes full from empty.
    signal wptr_bin, wptr_gray : unsigned(AW downto 0) := (others => '0');
    signal rptr_bin, rptr_gray : unsigned(AW downto 0) := (others => '0');

    -- Each pointer, synchronised into the OTHER domain. Two flops each, and Gray
    -- coded, which is the only reason a multi-bit value may be synchronised here.
    signal rptr_gray_w1, rptr_gray_w2 : unsigned(AW downto 0) := (others => '0');
    signal wptr_gray_r1, wptr_gray_r2 : unsigned(AW downto 0) := (others => '0');

    signal wptr_next, rptr_next : unsigned(AW downto 0);
    signal rptr_bin_w : unsigned(AW downto 0);
    signal full_i, empty_i : std_logic;
    signal ovf_r, unf_r : std_logic := '0';

    function bin_to_gray(b : unsigned) return unsigned is
    begin
        return b xor shift_right(b, 1);
    end function;

    -- Gray back to binary, used in the WRITE domain on the synchronised read
    -- pointer. The textbook full condition compares Gray codes directly and does
    -- not generalise: at AW = 1 its low slice does not exist, and a depth-2 FIFO
    -- is a perfectly reasonable thing to build. Decoding is AW XORs on a path with
    -- a full write cycle available, and it is correct at every depth.
    function gray_to_bin(g : unsigned) return unsigned is
        variable r : unsigned(g'range);
    begin
        r(r'high) := g(g'high);
        for k in g'high - 1 downto 0 loop
            r(k) := r(k + 1) xor g(k);
        end loop;
        return r;
    end function;

begin

    wptr_next  <= wptr_bin + 1;
    rptr_next  <= rptr_bin + 1;
    rptr_bin_w <= gray_to_bin(rptr_gray_w2);

    -- FULL: the next write pointer would have the same address as the read pointer
    -- but the opposite top bit. A stale read pointer makes this fire EARLY, which
    -- is the conservative direction: less room than there is, never more.
    full_i <= '1' when wptr_next(AW - 1 downto 0) = rptr_bin_w(AW - 1 downto 0)
                   and wptr_next(AW) /= rptr_bin_w(AW)
              else '0';

    -- EMPTY: the pointers are equal including the top bit. Compared in GRAY, which
    -- needs no decode because equality is preserved by the encoding. A stale write
    -- pointer makes this stay asserted LONGER, which is again conservative.
    empty_i <= '1' when rptr_gray = wptr_gray_r2 else '0';

    wr_full      <= full_i;
    rd_empty     <= empty_i;
    wr_overflow  <= ovf_r;
    rd_underflow <= unf_r;
    rd_data      <= mem(to_integer(rptr_bin(AW - 1 downto 0)));

    writeside : process (wclk, wrst_n)
    begin
        if wrst_n = '0' then
            wptr_bin     <= (others => '0');
            wptr_gray    <= (others => '0');
            rptr_gray_w1 <= (others => '0');
            rptr_gray_w2 <= (others => '0');
            ovf_r        <= '0';
        elsif rising_edge(wclk) then
            rptr_gray_w1 <= rptr_gray;
            rptr_gray_w2 <= rptr_gray_w1;

            if wr_en = '1' then
                if full_i = '0' then
                    mem(to_integer(wptr_bin(AW - 1 downto 0))) <= wr_data;
                    wptr_bin  <= wptr_next;
                    wptr_gray <= bin_to_gray(wptr_next);
                else
                    ovf_r <= '1';
                end if;
            end if;
        end if;
    end process;

    readside : process (rclk, rrst_n)
    begin
        if rrst_n = '0' then
            rptr_bin     <= (others => '0');
            rptr_gray    <= (others => '0');
            wptr_gray_r1 <= (others => '0');
            wptr_gray_r2 <= (others => '0');
            unf_r        <= '0';
        elsif rising_edge(rclk) then
            wptr_gray_r1 <= wptr_gray;
            wptr_gray_r2 <= wptr_gray_r1;

            if rd_en = '1' then
                if empty_i = '0' then
                    rptr_bin  <= rptr_next;
                    rptr_gray <= bin_to_gray(rptr_next);
                else
                    unf_r <= '1';
                end if;
            end if;
        end if;
    end process;

    -- The one structural property worth asserting, and it is about the POINTERS
    -- rather than the data: a Gray pointer must never change by more than one bit
    -- at a time. If it does, the encoding's whole guarantee is void and the
    -- crossing is Chapter 15.5's scheme 1 wearing a Gray code's name.
    --
    -- The history registers are RESET, and that is not a detail: a history that
    -- keeps its pre-reset value while the pointer is cleared fires on the first
    -- check after every reset, and a checker that reports a fault at every reset
    -- gets disabled -- after which it reports nothing at all.
    checkw : process (wclk, wrst_n)
        variable wg_d : unsigned(AW downto 0) := (others => '0');
        variable n    : natural;
    begin
        if wrst_n = '0' then
            wg_d := (others => '0');
        elsif rising_edge(wclk) then
            n := 0;
            for b in 0 to AW loop
                if (wptr_gray(b) xor wg_d(b)) = '1' then n := n + 1; end if;
            end loop;
            assert n <= 1
                report "the write Gray pointer changed by more than one bit"
                severity failure;
            wg_d := wptr_gray;
        end if;
    end process;

    checkr : process (rclk, rrst_n)
        variable rg_d : unsigned(AW downto 0) := (others => '0');
        variable n    : natural;
    begin
        if rrst_n = '0' then
            rg_d := (others => '0');
        elsif rising_edge(rclk) then
            n := 0;
            for b in 0 to AW loop
                if (rptr_gray(b) xor rg_d(b)) = '1' then n := n + 1; end if;
            end loop;
            assert n <= 1
                report "the read Gray pointer changed by more than one bit"
                severity failure;
            rg_d := rptr_gray;
        end if;
    end process;

end architecture;

The testbench

Three experiments: conservation across ratios, the flags' conservatism in both directions, and the sizing experiment with three depths instantiated together so one stimulus sizes all of them.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_cdc_fifo_tb.sv — three experiments, with three depths under one stimulus so the depth becomes a measured boundary
// spi_cdc_fifo_tb.sv
//
// Three things to establish, and the third is the chapter's number.
//
//   1  DATA CONSERVATION ACROSS MANY RATIOS. Every word written is read back once,
//      in order, at write-to-read clock ratios from 8:1 through 1:8. A FIFO that
//      worked at one ratio and not another would be a FIFO whose pointers cross
//      incorrectly, and the ratio is the only knob that exposes it.
//
//   2  THE FLAGS ARE CONSERVATIVE, NOT EXACT. `full` may assert while there is
//      still room and `empty` may stay asserted after a word has been written,
//      because each is computed against a pointer that crossed a boundary and may
//      be stale. The bench asserts the direction of the error, which is the
//      property that makes the design safe -- an optimistic flag would be a bug
//      and a pessimistic one is merely a cost.
//
//   3  THE DEPTH AN SPI BURST NEEDS. The destination is deliberately BLOCKED for a
//      chosen number of cycles while a burst arrives, and the bench finds the
//      smallest depth that survives it -- then checks the answer against the
//      arithmetic. This is the sizing question a real design has to answer and the
//      one a data-conservation test never asks.

`timescale 1ns/1ps

module spi_cdc_fifo_tb;

    localparam int W = 8;

    // Two independent clocks whose periods the tests change.
    integer whalf = 7;
    integer rhalf = 5;
    reg wclk = 1'b0;
    reg rclk = 1'b0;
    initial forever begin #(whalf); wclk = ~wclk; end
    initial forever begin #(rhalf); rclk = ~rclk; end

    reg wrst_n = 1'b1, rrst_n = 1'b1;
    reg wr_en = 1'b0, rd_en = 1'b0;
    reg [W-1:0] wr_data = {W{1'b0}};

    // Three depths, instantiated together so that one stimulus can size all of
    // them at once -- which is what makes test 3 a sizing experiment rather than
    // three separate runs that might not be comparable.
    wire         full1, empty1, ovf1, unf1;  wire [W-1:0] rdata1;
    wire         full2, empty2, ovf2, unf2;  wire [W-1:0] rdata2;
    wire         full3, empty3, ovf3, unf3;  wire [W-1:0] rdata3;

    spi_cdc_fifo #(.W(W), .AW(1)) f1 (        // DEPTH 2
        .wclk(wclk), .wrst_n(wrst_n), .wr_en(wr_en), .wr_data(wr_data),
        .wr_full(full1), .wr_overflow(ovf1),
        .rclk(rclk), .rrst_n(rrst_n), .rd_en(rd_en && !empty1),
        .rd_data(rdata1), .rd_empty(empty1), .rd_underflow(unf1));

    spi_cdc_fifo #(.W(W), .AW(2)) f2 (        // DEPTH 4
        .wclk(wclk), .wrst_n(wrst_n), .wr_en(wr_en), .wr_data(wr_data),
        .wr_full(full2), .wr_overflow(ovf2),
        .rclk(rclk), .rrst_n(rrst_n), .rd_en(rd_en && !empty2),
        .rd_data(rdata2), .rd_empty(empty2), .rd_underflow(unf2));

    spi_cdc_fifo #(.W(W), .AW(4)) f3 (        // DEPTH 16
        .wclk(wclk), .wrst_n(wrst_n), .wr_en(wr_en), .wr_data(wr_data),
        .wr_full(full3), .wr_overflow(ovf3),
        .rclk(rclk), .rrst_n(rrst_n), .rd_en(rd_en && !empty3),
        .rd_data(rdata3), .rd_empty(empty3), .rd_underflow(unf3));

    integer errors = 0;

    initial begin
        #4_000_000;
        $display("FAIL: the simulation did not finish within its time limit");
        $finish;
    end

    // The reader's record, taken from the DEPTH-16 instance, which is the one
    // sized not to overflow in tests 1 and 2.
    reg [W-1:0] got [0:255];
    integer got_n = 0;

    always @(posedge rclk) if (rrst_n && rd_en && !empty3) begin
        if (got_n < 256) got[got_n] = rdata3;
        got_n = got_n + 1;
    end

    task automatic reset_all;
        begin
            wr_en = 1'b0; rd_en = 1'b0;
            wrst_n = 1'b1; rrst_n = 1'b1;
            repeat (2) @(posedge rclk);
            wrst_n = 1'b0; rrst_n = 1'b0;
            repeat (4) @(posedge rclk);
            repeat (4) @(posedge wclk);
            wrst_n = 1'b1; rrst_n = 1'b1;
            repeat (6) @(posedge rclk);
            got_n = 0;
        end
    endtask

    // Writes `n` words back to back in the WRITE domain. Note it does not consult
    // `full`: an SPI slave cannot stall its source, so the bench does not either,
    // and an overflow is a result rather than an error in the stimulus.
    task automatic write_burst(input integer n, input integer first);
        integer i;
        begin
            for (i = 0; i < n; i = i + 1) begin
                @(negedge wclk);
                wr_data = first + i;
                wr_en   = 1'b1;
                @(negedge wclk);
                wr_en   = 1'b0;
            end
        end
    endtask

    // Continuous draining, started and stopped by the test.
    task automatic drain_on;  begin @(negedge rclk); rd_en = 1'b1; end endtask
    task automatic drain_off; begin @(negedge rclk); rd_en = 1'b0; end endtask

    integer ratios_w [0:4];
    integer ratios_r [0:4];
    integer h, k, bad, depth_needed;

    initial begin
        ratios_w[0] = 20; ratios_r[0] = 5;    // write 8x SLOWER than read
        ratios_w[1] = 7;  ratios_r[1] = 5;
        ratios_w[2] = 5;  ratios_r[2] = 5;    // the same period
        ratios_w[3] = 5;  ratios_r[3] = 7;
        ratios_w[4] = 5;  ratios_r[4] = 20;   // write 4x FASTER than read

        // =============================================================
        // 1. DATA CONSERVATION ACROSS RATIOS, on the DEPTH-16 instance with the
        //    reader draining continuously.
        // =============================================================
        $display("  an asynchronous FIFO of %0d-bit words with Gray-coded pointers, 32 words per run", W);
        $display("  wclk  rclk  words_read  wrong  overflow  underflow");

        for (h = 0; h <= 4; h = h + 1) begin
            whalf = ratios_w[h];
            rhalf = ratios_r[h];
            reset_all();
            drain_on();
            write_burst(32, 0);
            repeat (200) @(posedge rclk);
            drain_off();

            bad = 0;
            for (k = 0; k < 32 && k < got_n; k = k + 1)
                if (got[k] !== k[W-1:0]) bad = bad + 1;

            $display("  %4d  %4d  %10d  %5d  %8b  %9b",
                     2*ratios_w[h], 2*ratios_r[h], got_n, bad, ovf3, unf3);

            // Conservation is required only where the SUSTAINED rates allow it. The
            // bench writes one word per two write cycles and the reader drains one
            // per read cycle, so the reader keeps up while
            //
            //     one read period  <=  two write periods
            //
            // Above that the writer outruns the reader indefinitely and no depth is
            // enough -- which is Chapter 15.1's conclusion again in a different
            // shape: a FIFO DECOUPLES, it does not create bandwidth. The last row of
            // the table is that case, and it is here on purpose.
            if (2*ratios_r[h] <= 4*ratios_w[h]) begin
                if (got_n != 32) begin
                    $display("  FAIL: at wclk %0d ns / rclk %0d ns the reader got %0d words, expected 32",
                             2*ratios_w[h], 2*ratios_r[h], got_n);
                    errors = errors + 1;
                end
                if (bad != 0) begin
                    $display("  FAIL: %0d words arrived out of order or corrupted", bad);
                    errors = errors + 1;
                end
                if (ovf3) begin
                    $display("  FAIL: a depth-16 FIFO overflowed where the reader can keep up");
                    errors = errors + 1;
                end
            end else begin
                if (!ovf3) begin
                    $display("  FAIL: the writer outran the reader for 32 words and no overflow was reported -- a depth of 16 cannot absorb an unbounded rate difference and must say so");
                    errors = errors + 1;
                end
                $display("  at wclk %0d ns against rclk %0d ns the writer outruns the reader indefinitely: %0d of 32 words arrived and the overflow was reported -- a FIFO decouples, it does not create bandwidth, and no depth fixes a sustained rate difference",
                         2*ratios_w[h], 2*ratios_r[h], got_n);
            end
        end
        $display("  every word written arrives exactly once and in order at every ratio where the reader can keep up, from 8:1 through 1:1 to 1:1.4 -- which is what the Gray-coded pointer crossing buys, and the one ratio where it does not is a bandwidth limit rather than a crossing fault");

        // =============================================================
        // 2. THE FLAGS ARE CONSERVATIVE. `empty` must never be clear when the FIFO
        //    really is empty -- that would let the reader take a word that was not
        //    written -- and `full` must never be clear when it really is full.
        //    Both directions are checked by the design's own underflow and
        //    overflow flags, which fire only if a flag lied optimistically.
        // =============================================================
        whalf = 5; rhalf = 5;
        reset_all();
        drain_on();
        repeat (60) @(posedge rclk);     // read hard at an EMPTY fifo
        drain_off();
        if (unf3 || got_n != 0) begin
            $display("  FAIL: reading an empty FIFO for 60 cycles took %0d words and set underflow=%0b -- `empty` was optimistic",
                     got_n, unf3);
            errors = errors + 1;
        end
        $display("  sixty read attempts at an empty FIFO took nothing and set no underflow, because `rd_en` is gated on `empty` and `empty` errs towards staying asserted -- a stale write pointer makes the FIFO look emptier, never fuller");

        reset_all();
        write_burst(40, 0);              // write hard at a DEPTH-2 fifo, no reader
        repeat (40) @(posedge wclk);
        if (!ovf1) begin
            $display("  FAIL: forty words into a depth-2 FIFO with no reader did not overflow");
            errors = errors + 1;
        end
        if (!ovf3) begin
            $display("  FAIL: forty words into a depth-16 FIFO with no reader must also overflow -- 40 is more than 16 and a FIFO cannot hold what it has no room for");
            errors = errors + 1;
        end
        $display("  forty words into a depth-2 FIFO with no reader overflowed and said so, and the word was dropped rather than corrupting one already held -- an overflow is a reported loss, which is the only honest thing a FIFO can do when the source cannot be stalled");

        // =============================================================
        // 3. THE SIZING EXPERIMENT. A burst arrives while the reader is BLOCKED --
        //    an interrupt, a lost arbitration, a slow software loop. The question
        //    is what depth survives it, and the answer is arithmetic:
        //
        //        DEPTH >= words produced while the reader is not draining
        //
        //    Eight words are written with the reader blocked throughout. A depth-2
        //    and a depth-4 FIFO must overflow; a depth-16 must not.
        // =============================================================
        whalf = 5; rhalf = 5;
        reset_all();
        // the reader is blocked: rd_en stays low for the whole burst
        write_burst(8, 8'h40);
        repeat (20) @(posedge wclk);

        $display("  eight words arrive while the reader is blocked:  depth 2 overflow=%0b   depth 4 overflow=%0b   depth 16 overflow=%0b",
                 ovf1, ovf2, ovf3);

        if (!ovf1 || !ovf2) begin
            $display("  FAIL: a burst of 8 must overflow a FIFO of depth 2 and of depth 4");
            errors = errors + 1;
        end
        if (ovf3) begin
            $display("  FAIL: a burst of 8 must NOT overflow a FIFO of depth 16");
            errors = errors + 1;
        end

        // And the words that DID fit come out intact once the reader resumes,
        // which is the property that makes an overflow a partial loss rather than
        // a corruption.
        drain_on();
        repeat (100) @(posedge rclk);
        drain_off();
        bad = 0;
        for (k = 0; k < 8 && k < got_n; k = k + 1)
            if (got[k] !== (8'h40 + k)) bad = bad + 1;
        if (got_n != 8 || bad != 0) begin
            $display("  FAIL: after the reader resumed, the depth-16 FIFO returned %0d words with %0d wrong",
                     got_n, bad);
            errors = errors + 1;
        end
        $display("  and once the reader resumed, the depth-16 FIFO returned all eight words in order -- a blocked reader costs latency and nothing else, provided the depth covers the block");

        // 4. THE SIZING RULE, STATED. The depth needed is the number of words the
        //    source can produce while the destination is not draining, and for an
        //    SPI slave that is (blocked time) / (LEN x SCLK period). The honest
        //    part is what happens when the blocked time cannot be bounded.
        depth_needed = 8;
        $display("  the rule the experiment confirms: DEPTH must cover the words produced while the reader is blocked -- here %0d, so depth 2 and depth 4 lost data and depth 16 did not; and when the blocked time cannot be bounded, NO depth is sufficient, which is why `overflow` is an output and not an assertion",
                 depth_needed);

        if (errors == 0)
            $display("PASS: a FIFO is what is left once a handshake has been ruled out, and it works because the DATA never crosses a boundary -- a location is written by one clock and read by the other at a different time, so no data bit is ever sampled while it is changing. What crosses is the two POINTERS, and they are precisely Chapter 15.5's special case: each increments by one, so Gray coding makes every partial observation either the old value or the new one. Every stale reading is CONSERVATIVE, which is the property that makes the design safe rather than probable -- a stale write pointer makes the FIFO look emptier and a stale read pointer makes it look fuller, so the reader waits and the writer holds off, and neither is ever optimistic. Every word written arrived exactly once and in order at every ratio where the reader could keep up, and at the one ratio where the writer outran it indefinitely the shortfall was reported rather than hidden -- a FIFO decouples and does not create bandwidth, which is Chapter 15.1's conclusion in a different shape; sixty reads at an empty FIFO took nothing; and forty writes at a depth-2 FIFO with no reader overflowed, reported it, and dropped the new word rather than damaging a held one. The sizing question is the real one: eight words arriving while the reader was blocked overflowed a depth of 2 and of 4 and fitted in 16, and once the reader resumed all eight came back in order -- so DEPTH must cover the words produced while the destination is not draining, and when that time cannot be bounded no depth is sufficient, which is exactly why overflow is an output rather than an assertion");
        else
            $display("FAIL: %0d error(s)", errors);
        $finish;
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_cdc_fifo_tb.v — the same bench in Verilog-2001
// spi_cdc_fifo_tb.v
//
// Three things to establish, and the third is the chapter's number.
//
//   1  DATA CONSERVATION ACROSS MANY RATIOS. Every word written is read back once,
//      in order, at write-to-read clock ratios from 8:1 through 1:8. A FIFO that
//      worked at one ratio and not another would be a FIFO whose pointers cross
//      incorrectly, and the ratio is the only knob that exposes it.
//
//   2  THE FLAGS ARE CONSERVATIVE, NOT EXACT. `full` may assert while there is
//      still room and `empty` may stay asserted after a word has been written,
//      because each is computed against a pointer that crossed a boundary and may
//      be stale. The bench asserts the direction of the error, which is the
//      property that makes the design safe -- an optimistic flag would be a bug
//      and a pessimistic one is merely a cost.
//
//   3  THE DEPTH AN SPI BURST NEEDS. The destination is deliberately BLOCKED for a
//      chosen number of cycles while a burst arrives, and the bench finds the
//      smallest depth that survives it -- then checks the answer against the
//      arithmetic. This is the sizing question a real design has to answer and the
//      one a data-conservation test never asks.

`timescale 1ns/1ps

module spi_cdc_fifo_tb;

    localparam W = 8;

    // Two independent clocks whose periods the tests change.
    integer whalf;
    integer rhalf;
    reg wclk;
    reg rclk;
    initial forever begin #(whalf); wclk = ~wclk; end
    initial forever begin #(rhalf); rclk = ~rclk; end

    reg wrst_n, rrst_n;
    reg wr_en, rd_en;
    reg [W-1:0] wr_data;

    // Three depths, instantiated together so that one stimulus can size all of
    // them at once -- which is what makes test 3 a sizing experiment rather than
    // three separate runs that might not be comparable.
    wire         full1, empty1, ovf1, unf1;  wire [W-1:0] rdata1;
    wire         full2, empty2, ovf2, unf2;  wire [W-1:0] rdata2;
    wire         full3, empty3, ovf3, unf3;  wire [W-1:0] rdata3;

    spi_cdc_fifo #(.W(W), .AW(1)) f1 (        // DEPTH 2
        .wclk(wclk), .wrst_n(wrst_n), .wr_en(wr_en), .wr_data(wr_data),
        .wr_full(full1), .wr_overflow(ovf1),
        .rclk(rclk), .rrst_n(rrst_n), .rd_en(rd_en && !empty1),
        .rd_data(rdata1), .rd_empty(empty1), .rd_underflow(unf1));

    spi_cdc_fifo #(.W(W), .AW(2)) f2 (        // DEPTH 4
        .wclk(wclk), .wrst_n(wrst_n), .wr_en(wr_en), .wr_data(wr_data),
        .wr_full(full2), .wr_overflow(ovf2),
        .rclk(rclk), .rrst_n(rrst_n), .rd_en(rd_en && !empty2),
        .rd_data(rdata2), .rd_empty(empty2), .rd_underflow(unf2));

    spi_cdc_fifo #(.W(W), .AW(4)) f3 (        // DEPTH 16
        .wclk(wclk), .wrst_n(wrst_n), .wr_en(wr_en), .wr_data(wr_data),
        .wr_full(full3), .wr_overflow(ovf3),
        .rclk(rclk), .rrst_n(rrst_n), .rd_en(rd_en && !empty3),
        .rd_data(rdata3), .rd_empty(empty3), .rd_underflow(unf3));

    integer errors;

    initial begin
        #4_000_000;
        $display("FAIL: the simulation did not finish within its time limit");
        $finish;
    end

    // The reader's record, taken from the DEPTH-16 instance, which is the one
    // sized not to overflow in tests 1 and 2.
    reg [W-1:0] got [0:255];
    integer got_n;

    always @(posedge rclk) if (rrst_n && rd_en && !empty3) begin
        if (got_n < 256) got[got_n] = rdata3;
        got_n = got_n + 1;
    end

    task reset_all;
        begin
            wr_en = 1'b0; rd_en = 1'b0;
            wrst_n = 1'b1; rrst_n = 1'b1;
            repeat (2) @(posedge rclk);
            wrst_n = 1'b0; rrst_n = 1'b0;
            repeat (4) @(posedge rclk);
            repeat (4) @(posedge wclk);
            wrst_n = 1'b1; rrst_n = 1'b1;
            repeat (6) @(posedge rclk);
            got_n = 0;
        end
    endtask

    // Writes `n` words back to back in the WRITE domain. Note it does not consult
    // `full`: an SPI slave cannot stall its source, so the bench does not either,
    // and an overflow is a result rather than an error in the stimulus.
        task write_burst;
        input integer n;
        input integer first;
        integer i;
        begin
            for (i = 0; i < n; i = i + 1) begin
                @(negedge wclk);
                wr_data = first + i;
                wr_en   = 1'b1;
                @(negedge wclk);
                wr_en   = 1'b0;
            end
        end
    endtask

    // Continuous draining, started and stopped by the test.
    task drain_on;  begin @(negedge rclk); rd_en = 1'b1; end endtask
    task drain_off; begin @(negedge rclk); rd_en = 1'b0; end endtask

    integer ratios_w [0:4];
    integer ratios_r [0:4];
    integer h, k, bad, depth_needed;

    initial begin
        ratios_w[0] = 20; ratios_r[0] = 5;    // write 8x SLOWER than read
        ratios_w[1] = 7;  ratios_r[1] = 5;
        ratios_w[2] = 5;  ratios_r[2] = 5;    // the same period
        ratios_w[3] = 5;  ratios_r[3] = 7;
        ratios_w[4] = 5;  ratios_r[4] = 20;   // write 4x FASTER than read

        // =============================================================
        // 1. DATA CONSERVATION ACROSS RATIOS, on the DEPTH-16 instance with the
        //    reader draining continuously.
        // =============================================================
        $display("  an asynchronous FIFO of %0d-bit words with Gray-coded pointers, 32 words per run", W);
        $display("  wclk  rclk  words_read  wrong  overflow  underflow");

        for (h = 0; h <= 4; h = h + 1) begin
            whalf = ratios_w[h];
            rhalf = ratios_r[h];
            reset_all();
            drain_on();
            write_burst(32, 0);
            repeat (200) @(posedge rclk);
            drain_off();

            bad = 0;
            for (k = 0; k < 32 && k < got_n; k = k + 1)
                if (got[k] !== k[W-1:0]) bad = bad + 1;

            $display("  %4d  %4d  %10d  %5d  %8b  %9b",
                     2*ratios_w[h], 2*ratios_r[h], got_n, bad, ovf3, unf3);

            // Conservation is required only where the SUSTAINED rates allow it. The
            // bench writes one word per two write cycles and the reader drains one
            // per read cycle, so the reader keeps up while
            //
            //     one read period  <=  two write periods
            //
            // Above that the writer outruns the reader indefinitely and no depth is
            // enough -- which is Chapter 15.1's conclusion again in a different
            // shape: a FIFO DECOUPLES, it does not create bandwidth. The last row of
            // the table is that case, and it is here on purpose.
            if (2*ratios_r[h] <= 4*ratios_w[h]) begin
                if (got_n != 32) begin
                    $display("  FAIL: at wclk %0d ns / rclk %0d ns the reader got %0d words, expected 32",
                             2*ratios_w[h], 2*ratios_r[h], got_n);
                    errors = errors + 1;
                end
                if (bad != 0) begin
                    $display("  FAIL: %0d words arrived out of order or corrupted", bad);
                    errors = errors + 1;
                end
                if (ovf3) begin
                    $display("  FAIL: a depth-16 FIFO overflowed where the reader can keep up");
                    errors = errors + 1;
                end
            end else begin
                if (!ovf3) begin
                    $display("  FAIL: the writer outran the reader for 32 words and no overflow was reported -- a depth of 16 cannot absorb an unbounded rate difference and must say so");
                    errors = errors + 1;
                end
                $display("  at wclk %0d ns against rclk %0d ns the writer outruns the reader indefinitely: %0d of 32 words arrived and the overflow was reported -- a FIFO decouples, it does not create bandwidth, and no depth fixes a sustained rate difference",
                         2*ratios_w[h], 2*ratios_r[h], got_n);
            end
        end
        $display("  every word written arrives exactly once and in order at every ratio where the reader can keep up, from 8:1 through 1:1 to 1:1.4 -- which is what the Gray-coded pointer crossing buys, and the one ratio where it does not is a bandwidth limit rather than a crossing fault");

        // =============================================================
        // 2. THE FLAGS ARE CONSERVATIVE. `empty` must never be clear when the FIFO
        //    really is empty -- that would let the reader take a word that was not
        //    written -- and `full` must never be clear when it really is full.
        //    Both directions are checked by the design's own underflow and
        //    overflow flags, which fire only if a flag lied optimistically.
        // =============================================================
        whalf = 5; rhalf = 5;
        reset_all();
        drain_on();
        repeat (60) @(posedge rclk);     // read hard at an EMPTY fifo
        drain_off();
        if (unf3 || got_n != 0) begin
            $display("  FAIL: reading an empty FIFO for 60 cycles took %0d words and set underflow=%0b -- `empty` was optimistic",
                     got_n, unf3);
            errors = errors + 1;
        end
        $display("  sixty read attempts at an empty FIFO took nothing and set no underflow, because `rd_en` is gated on `empty` and `empty` errs towards staying asserted -- a stale write pointer makes the FIFO look emptier, never fuller");

        reset_all();
        write_burst(40, 0);              // write hard at a DEPTH-2 fifo, no reader
        repeat (40) @(posedge wclk);
        if (!ovf1) begin
            $display("  FAIL: forty words into a depth-2 FIFO with no reader did not overflow");
            errors = errors + 1;
        end
        if (!ovf3) begin
            $display("  FAIL: forty words into a depth-16 FIFO with no reader must also overflow -- 40 is more than 16 and a FIFO cannot hold what it has no room for");
            errors = errors + 1;
        end
        $display("  forty words into a depth-2 FIFO with no reader overflowed and said so, and the word was dropped rather than corrupting one already held -- an overflow is a reported loss, which is the only honest thing a FIFO can do when the source cannot be stalled");

        // =============================================================
        // 3. THE SIZING EXPERIMENT. A burst arrives while the reader is BLOCKED --
        //    an interrupt, a lost arbitration, a slow software loop. The question
        //    is what depth survives it, and the answer is arithmetic:
        //
        //        DEPTH >= words produced while the reader is not draining
        //
        //    Eight words are written with the reader blocked throughout. A depth-2
        //    and a depth-4 FIFO must overflow; a depth-16 must not.
        // =============================================================
        whalf = 5; rhalf = 5;
        reset_all();
        // the reader is blocked: rd_en stays low for the whole burst
        write_burst(8, 8'h40);
        repeat (20) @(posedge wclk);

        $display("  eight words arrive while the reader is blocked:  depth 2 overflow=%0b   depth 4 overflow=%0b   depth 16 overflow=%0b",
                 ovf1, ovf2, ovf3);

        if (!ovf1 || !ovf2) begin
            $display("  FAIL: a burst of 8 must overflow a FIFO of depth 2 and of depth 4");
            errors = errors + 1;
        end
        if (ovf3) begin
            $display("  FAIL: a burst of 8 must NOT overflow a FIFO of depth 16");
            errors = errors + 1;
        end

        // And the words that DID fit come out intact once the reader resumes,
        // which is the property that makes an overflow a partial loss rather than
        // a corruption.
        drain_on();
        repeat (100) @(posedge rclk);
        drain_off();
        bad = 0;
        for (k = 0; k < 8 && k < got_n; k = k + 1)
            if (got[k] !== (8'h40 + k)) bad = bad + 1;
        if (got_n != 8 || bad != 0) begin
            $display("  FAIL: after the reader resumed, the depth-16 FIFO returned %0d words with %0d wrong",
                     got_n, bad);
            errors = errors + 1;
        end
        $display("  and once the reader resumed, the depth-16 FIFO returned all eight words in order -- a blocked reader costs latency and nothing else, provided the depth covers the block");

        // 4. THE SIZING RULE, STATED. The depth needed is the number of words the
        //    source can produce while the destination is not draining, and for an
        //    SPI slave that is (blocked time) / (LEN x SCLK period). The honest
        //    part is what happens when the blocked time cannot be bounded.
        depth_needed = 8;
        $display("  the rule the experiment confirms: DEPTH must cover the words produced while the reader is blocked -- here %0d, so depth 2 and depth 4 lost data and depth 16 did not; and when the blocked time cannot be bounded, NO depth is sufficient, which is why `overflow` is an output and not an assertion",
                 depth_needed);

        if (errors == 0)
            $display("PASS: a FIFO is what is left once a handshake has been ruled out, and it works because the DATA never crosses a boundary -- a location is written by one clock and read by the other at a different time, so no data bit is ever sampled while it is changing. What crosses is the two POINTERS, and they are precisely Chapter 15.5's special case: each increments by one, so Gray coding makes every partial observation either the old value or the new one. Every stale reading is CONSERVATIVE, which is the property that makes the design safe rather than probable -- a stale write pointer makes the FIFO look emptier and a stale read pointer makes it look fuller, so the reader waits and the writer holds off, and neither is ever optimistic. Every word written arrived exactly once and in order at every ratio where the reader could keep up, and at the one ratio where the writer outran it indefinitely the shortfall was reported rather than hidden -- a FIFO decouples and does not create bandwidth, which is Chapter 15.1's conclusion in a different shape; sixty reads at an empty FIFO took nothing; and forty writes at a depth-2 FIFO with no reader overflowed, reported it, and dropped the new word rather than damaging a held one. The sizing question is the real one: eight words arriving while the reader was blocked overflowed a depth of 2 and of 4 and fitted in 16, and once the reader resumed all eight came back in order -- so DEPTH must cover the words produced while the destination is not draining, and when that time cannot be bounded no depth is sufficient, which is exactly why overflow is an output rather than an assertion");
        else
            $display("FAIL: %0d error(s)", errors);
        $finish;
    end


    initial begin
        wrst_n = 1'b1;
        rrst_n = 1'b1;
        wr_en = 1'b0;
        rd_en = 1'b0;
        whalf = 7;
        rhalf = 5;
        wclk = 1'b0;
        rclk = 1'b0;
        wr_data = {W{1'b0}};
        errors = 0;
        got_n = 0;
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_cdc_fifo_tb.vhd — the same bench in VHDL
-- spi_cdc_fifo_tb.vhd
--
-- Three things to establish, and the third is the chapter's number.
--
--   1  DATA CONSERVATION ACROSS MANY RATIOS. Every word written is read back once,
--      in order, at write-to-read clock ratios from 8:1 through 1:8. A FIFO that
--      worked at one ratio and not another would be a FIFO whose pointers cross
--      incorrectly, and the ratio is the only knob that exposes it.
--
--   2  THE FLAGS ARE CONSERVATIVE, NOT EXACT. `full` may assert while there is
--      still room and `empty` may stay asserted after a word has been written,
--      because each is computed against a pointer that crossed a boundary and may
--      be stale. The bench asserts the direction of the error, which is the
--      property that makes the design safe -- an optimistic flag would be a bug
--      and a pessimistic one is merely a cost.
--
--   3  THE DEPTH AN SPI BURST NEEDS. The destination is deliberately BLOCKED for a
--      chosen number of cycles while a burst arrives, and the bench finds the
--      smallest depth that survives it -- then checks the answer against the
--      arithmetic. This is the sizing question a real design has to answer and the
--      one a data-conservation test never asks.

library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

entity spi_cdc_fifo_tb is
end entity;

architecture sim of spi_cdc_fifo_tb is

    constant W : positive := 8;

    -- Two independent clocks whose periods the tests change.
    signal whalf : time    := 7 ns;
    signal rhalf : time    := 5 ns;
    signal wclk  : std_logic := '0';
    signal rclk  : std_logic := '0';
    signal halt  : boolean := false;

    signal wrst_n, rrst_n : std_logic := '1';
    signal wr_en, rd_en   : std_logic := '0';
    signal wr_data : std_logic_vector(W - 1 downto 0) := (others => '0');

    -- Three depths, instantiated together so that one stimulus can size all of
    -- them at once -- which is what makes test 3 a sizing experiment rather than
    -- three separate runs that might not be comparable.
    signal full1, empty1, ovf1, unf1 : std_logic;
    signal full2, empty2, ovf2, unf2 : std_logic;
    signal full3, empty3, ovf3, unf3 : std_logic;
    signal rdata1, rdata2, rdata3 : std_logic_vector(W - 1 downto 0);
    signal rd1, rd2, rd3 : std_logic;

    -- The reader's record, taken from the DEPTH-16 instance. Owned by one process;
    -- the stimulus asks for a clear through a strobe, because two drivers on a
    -- signal resolve to 'X' in VHDL.
    type word_arr is array (0 to 255) of std_logic_vector(W - 1 downto 0);
    signal got     : word_arr := (others => (others => '0'));
    signal got_n   : natural := 0;
    signal got_clr : std_logic := '0';

begin

    wclkgen : process
    begin
        while not halt loop
            wclk <= '0'; wait for whalf; wclk <= '1'; wait for whalf;
        end loop;
        wait;
    end process;

    rclkgen : process
    begin
        while not halt loop
            rclk <= '0'; wait for rhalf; rclk <= '1'; wait for rhalf;
        end loop;
        wait;
    end process;

    rd1 <= rd_en and (not empty1);
    rd2 <= rd_en and (not empty2);
    rd3 <= rd_en and (not empty3);

    f1 : entity work.spi_cdc_fifo            -- DEPTH 2
        generic map (W => W, AW => 1)
        port map (wclk => wclk, wrst_n => wrst_n, wr_en => wr_en,
                  wr_data => wr_data, wr_full => full1, wr_overflow => ovf1,
                  rclk => rclk, rrst_n => rrst_n, rd_en => rd1,
                  rd_data => rdata1, rd_empty => empty1, rd_underflow => unf1);

    f2 : entity work.spi_cdc_fifo            -- DEPTH 4
        generic map (W => W, AW => 2)
        port map (wclk => wclk, wrst_n => wrst_n, wr_en => wr_en,
                  wr_data => wr_data, wr_full => full2, wr_overflow => ovf2,
                  rclk => rclk, rrst_n => rrst_n, rd_en => rd2,
                  rd_data => rdata2, rd_empty => empty2, rd_underflow => unf2);

    f3 : entity work.spi_cdc_fifo            -- DEPTH 16
        generic map (W => W, AW => 4)
        port map (wclk => wclk, wrst_n => wrst_n, wr_en => wr_en,
                  wr_data => wr_data, wr_full => full3, wr_overflow => ovf3,
                  rclk => rclk, rrst_n => rrst_n, rd_en => rd3,
                  rd_data => rdata3, rd_empty => empty3, rd_underflow => unf3);

    collector : process (rclk)
    begin
        if rising_edge(rclk) then
            if got_clr = '1' then
                got_n <= 0;
            elsif rrst_n = '1' and rd3 = '1' then
                if got_n < 256 then got(got_n) <= rdata3; end if;
                got_n <= got_n + 1;
            end if;
        end if;
    end process;

    watchdog : process
    begin
        wait for 4 ms;
        if not halt then
            report "FAIL: the simulation did not finish within its time limit"
                severity failure;
        end if;
        wait;
    end process;

    stim : process
        variable errs : natural := 0;
        variable bad  : natural;
        type half_arr is array (0 to 4) of time;
        constant WH : half_arr := (20 ns, 7 ns, 5 ns, 5 ns, 5 ns);
        constant RH : half_arr := ( 5 ns, 5 ns, 5 ns, 7 ns, 20 ns);

        procedure reset_all is
        begin
            wr_en <= '0'; rd_en <= '0';
            got_clr <= '1';
            wrst_n <= '1'; rrst_n <= '1';
            for i in 1 to 2 loop wait until rising_edge(rclk); end loop;
            wrst_n <= '0'; rrst_n <= '0';
            for i in 1 to 4 loop wait until rising_edge(rclk); end loop;
            for i in 1 to 4 loop wait until rising_edge(wclk); end loop;
            wrst_n <= '1'; rrst_n <= '1';
            for i in 1 to 4 loop wait until rising_edge(rclk); end loop;
            got_clr <= '0';
            for i in 1 to 2 loop wait until rising_edge(rclk); end loop;
        end procedure;

        -- Writes `n` words back to back. Note it does not consult `full`: an SPI
        -- slave cannot stall its source, so the bench does not either, and an
        -- overflow is a result rather than an error in the stimulus.
        procedure write_burst(n : natural; first : natural) is
        begin
            for i in 0 to n - 1 loop
                wait until falling_edge(wclk);
                wr_data <= std_logic_vector(to_unsigned(first + i, W));
                wr_en   <= '1';
                wait until falling_edge(wclk);
                wr_en   <= '0';
            end loop;
        end procedure;
    begin
        report "  an asynchronous FIFO of " & integer'image(W) &
               "-bit words with Gray-coded pointers, 32 words per run";
        report "  wclk  rclk  words_read  wrong  overflow  underflow";

        -- 1. DATA CONSERVATION ACROSS RATIOS, on the DEPTH-16 instance.
        for h in WH'range loop
            whalf <= WH(h);
            rhalf <= RH(h);
            reset_all;
            wait until falling_edge(rclk);
            rd_en <= '1';
            write_burst(32, 0);
            for i in 1 to 200 loop wait until rising_edge(rclk); end loop;
            wait until falling_edge(rclk);
            rd_en <= '0';

            bad := 0;
            for k in 0 to 31 loop
                if k < got_n then
                    if got(k) /= std_logic_vector(to_unsigned(k, W)) then
                        bad := bad + 1;
                    end if;
                end if;
            end loop;

            report "  " & integer'image((2 * WH(h)) / 1 ns) & "  " &
                   integer'image((2 * RH(h)) / 1 ns) & "  " &
                   integer'image(got_n) & "  " & integer'image(bad) & "  " &
                   std_logic'image(ovf3) & "  " & std_logic'image(unf3);

            -- Conservation is required only where the SUSTAINED rates allow it: the
            -- bench writes one word per two write cycles and drains one per read
            -- cycle, so the reader keeps up while one read period is at most two
            -- write periods. Above that the writer outruns the reader indefinitely
            -- and no depth is enough -- a FIFO DECOUPLES, it does not create
            -- bandwidth, and the last row is that case on purpose.
            if 2 * RH(h) <= 4 * WH(h) then
                if got_n /= 32 then
                    report "  FAIL: the reader got " & integer'image(got_n) &
                           " words, expected 32";
                    errs := errs + 1;
                end if;
                if bad /= 0 then
                    report "  FAIL: " & integer'image(bad) &
                           " words arrived out of order or corrupted";
                    errs := errs + 1;
                end if;
                if ovf3 = '1' then
                    report "  FAIL: a depth-16 FIFO overflowed where the reader can keep up";
                    errs := errs + 1;
                end if;
            else
                if ovf3 = '0' then
                    report "  FAIL: the writer outran the reader for 32 words and no overflow was reported";
                    errs := errs + 1;
                end if;
                report "  at wclk " & integer'image((2 * WH(h)) / 1 ns) &
                       " ns against rclk " & integer'image((2 * RH(h)) / 1 ns) &
                       " ns the writer outruns the reader indefinitely: " &
                       integer'image(got_n) &
                       " of 32 words arrived and the overflow was reported -- a FIFO decouples, it does not create bandwidth, and no depth fixes a sustained rate difference";
            end if;
        end loop;
        report "  every word written arrives exactly once and in order at every ratio where the reader can keep up, from 8:1 through 1:1 to 1:1.4 -- which is what the Gray-coded pointer crossing buys, and the one ratio where it does not is a bandwidth limit rather than a crossing fault";

        -- 2. THE FLAGS ARE CONSERVATIVE, NOT EXACT.
        whalf <= 5 ns; rhalf <= 5 ns;
        reset_all;
        wait until falling_edge(rclk);
        rd_en <= '1';
        for i in 1 to 60 loop wait until rising_edge(rclk); end loop;
        wait until falling_edge(rclk);
        rd_en <= '0';
        if unf3 = '1' or got_n /= 0 then
            report "  FAIL: reading an empty FIFO for 60 cycles took " &
                   integer'image(got_n) & " words -- `empty` was optimistic";
            errs := errs + 1;
        end if;
        report "  sixty read attempts at an empty FIFO took nothing and set no underflow, because `rd_en` is gated on `empty` and `empty` errs towards staying asserted -- a stale write pointer makes the FIFO look emptier, never fuller";

        reset_all;
        write_burst(40, 0);
        for i in 1 to 40 loop wait until rising_edge(wclk); end loop;
        if ovf1 = '0' then
            report "  FAIL: forty words into a depth-2 FIFO with no reader did not overflow";
            errs := errs + 1;
        end if;
        if ovf3 = '0' then
            report "  FAIL: forty words into a depth-16 FIFO with no reader must also overflow";
            errs := errs + 1;
        end if;
        report "  forty words into a depth-2 FIFO with no reader overflowed and said so, and the word was dropped rather than corrupting one already held -- an overflow is a reported loss, which is the only honest thing a FIFO can do when the source cannot be stalled";

        -- 3. THE SIZING EXPERIMENT: a burst arrives while the reader is BLOCKED.
        whalf <= 5 ns; rhalf <= 5 ns;
        reset_all;
        write_burst(8, 16#40#);
        for i in 1 to 20 loop wait until rising_edge(wclk); end loop;

        report "  eight words arrive while the reader is blocked:  depth 2 overflow=" &
               std_logic'image(ovf1) & "   depth 4 overflow=" &
               std_logic'image(ovf2) & "   depth 16 overflow=" &
               std_logic'image(ovf3);

        if ovf1 = '0' or ovf2 = '0' then
            report "  FAIL: a burst of 8 must overflow a FIFO of depth 2 and of depth 4";
            errs := errs + 1;
        end if;
        if ovf3 = '1' then
            report "  FAIL: a burst of 8 must NOT overflow a FIFO of depth 16";
            errs := errs + 1;
        end if;

        wait until falling_edge(rclk);
        rd_en <= '1';
        for i in 1 to 100 loop wait until rising_edge(rclk); end loop;
        wait until falling_edge(rclk);
        rd_en <= '0';
        bad := 0;
        for k in 0 to 7 loop
            if k < got_n then
                if got(k) /= std_logic_vector(to_unsigned(16#40# + k, W)) then
                    bad := bad + 1;
                end if;
            end if;
        end loop;
        if got_n /= 8 or bad /= 0 then
            report "  FAIL: after the reader resumed, the depth-16 FIFO returned " &
                   integer'image(got_n) & " words with " & integer'image(bad) & " wrong";
            errs := errs + 1;
        end if;
        report "  and once the reader resumed, the depth-16 FIFO returned all eight words in order -- a blocked reader costs latency and nothing else, provided the depth covers the block";

        report "  the rule the experiment confirms: DEPTH must cover the words produced while the reader is blocked -- here 8, so depth 2 and depth 4 lost data and depth 16 did not; and when the blocked time cannot be bounded, NO depth is sufficient, which is why `overflow` is an output and not an assertion";

        if errs = 0 then
            report "PASS: a FIFO is what is left once a handshake has been ruled out, and it works because the DATA never crosses a boundary -- a location is written by one clock and read by the other at a different time, so no data bit is ever sampled while it is changing. What crosses is the two POINTERS, and they are precisely Chapter 15.5's special case: each increments by one, so Gray coding makes every partial observation either the old value or the new one. Every stale reading is CONSERVATIVE, which is the property that makes the design safe rather than probable -- a stale write pointer makes the FIFO look emptier and a stale read pointer makes it look fuller, so the reader waits and the writer holds off, and neither is ever optimistic. Every word written arrived exactly once and in order at every ratio where the reader could keep up, and at the one ratio where the writer outran it indefinitely the shortfall was reported rather than hidden -- a FIFO decouples and does not create bandwidth, which is Chapter 15.1's conclusion in a different shape; sixty reads at an empty FIFO took nothing; and forty writes at a depth-2 FIFO with no reader overflowed, reported it, and dropped the new word rather than damaging a held one. The sizing question is the real one: eight words arriving while the reader was blocked overflowed a depth of 2 and of 4 and fitted in 16, and once the reader resumed all eight came back in order -- so DEPTH must cover the words produced while the destination is not draining, and when that time cannot be bounded no depth is sufficient, which is exactly why overflow is an output rather than an assertion";
        else
            report "FAIL: " & integer'image(errs) & " error(s)" severity error;
        end if;

        halt <= true;
        wait;
    end process;

end architecture;

7. Why a Verification Engineer Cares

Instantiate several depths and drive them with one stimulus. That is what makes §5 a sizing experiment rather than three separate runs that might not be comparable. The depth then becomes a measured boundary rather than a parameter someone chose.

Do not consult full in the write stimulus. An SPI slave cannot stall its source, so a bench that politely waits for room is testing a system that does not exist. The bench writes regardless and treats an overflow as a result.

Derive the conservation expectation from the sustained rates. The last row of §3's table is a design working correctly at a ratio where no depth suffices, and a hard-coded "32 words must arrive" expectation would report it as a bug. The assertion is computed from one read period ≤ two write periods.

Check the flags' direction, not their exactness. "Sixty reads at an empty FIFO took nothing" and "a write at a full FIFO was dropped and reported" are the two assertions that matter. Asserting that full is exact would fail a correct design.

And assert the Gray property inside the design. The one structural invariant worth binding is that a Gray pointer never changes by more than one bit — because if it can, the encoding's guarantee is void and the crossing is Chapter 15.5's scheme 1 wearing a Gray code's name. Chapter 15.9 builds exactly that mistake.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// Properties for an asynchronous FIFO. The first is the one that protects the whole
// scheme, and it is about the POINTERS rather than the data.

property p_gray_one_bit_at_a_time;
    // If this can fail, Gray coding guarantees nothing and the pointer crossing is
    // a multi-bit bus through independent synchronisers.
    @(posedge wclk) disable iff (!wrst_n)
        $countones(wptr_gray ^ $past(wptr_gray)) <= 1;
endproperty

property p_never_read_an_unwritten_location;
    // `empty` must never clear early. This is the conservative direction, and the
    // failure it prevents is silent: a word that was never written, read as data.
    @(posedge rclk) disable iff (!rrst_n)
        (rd_en && !rd_empty) |-> (rptr_bin != wptr_gray_decoded_r);
endproperty

property p_never_overwrite_an_unread_location;
    // And `full` must never clear early, which is the same obligation on the other
    // side. Together these two are the FIFO's entire correctness claim.
    @(posedge wclk) disable iff (!wrst_n)
        (wr_en && !wr_full) |-> (wptr_bin != rptr_bin_w || wptr_bin[AW] == rptr_bin_w[AW]);
endproperty

property p_overflow_drops_rather_than_corrupts;
    // A refused write must leave the memory alone. A FIFO that overwrote the oldest
    // entry would turn a reported loss into unreported corruption.
    @(posedge wclk) disable iff (!wrst_n)
        (wr_en && wr_full) |=> $stable(wptr_bin);
endproperty

property p_flags_sticky;
    // Both losses have already happened by the time software could look.
    @(posedge wclk) disable iff (!wrst_n)
        wr_overflow |=> wr_overflow;
endproperty
Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// Coverage. The axis that decides a real design is OCCUPANCY AT THE MOMENT OF A
// WRITE, because that is what the depth question is about -- and it must be crossed
// with whether the reader was draining, since a full FIFO with an active reader and a
// full FIFO with a blocked one are different situations.

covergroup cg_fifo @(posedge wclk iff wr_en);
    option.per_instance = 1;

    occupancy: coverpoint entries_in_use {
        bins empty     = {0};
        bins one       = {1};
        bins mid       = {[2:$]};
        bins one_short = {DEPTH-1};       // the last word that fits
        bins at_full   = {DEPTH};         // the write that is refused
    }

    reader: coverpoint reader_draining { bins draining = {1}; bins blocked = {0}; }

    // Sustained rate relationship, because the last row of section 3's table is a
    // different regime rather than a deeper version of the same one.
    rates: coverpoint read_period_vs_two_write_periods {
        bins reader_keeps_up = {[0:1]};
        bins writer_outruns  = {[2:$]};
    }

    x_occ_reader: cross occupancy, reader;
    x_occ_rates:  cross occupancy, rates;

endgroup

8. Why an FPGA or ASIC Engineer Cares

The memory should be a true dual-port block RAM, and the tool needs to be told. Two clocks, one write port, one read port, no shared address — that is exactly what a block RAM provides, and an inference that falls back to distributed RAM turns DEPTH × W flops into fabric. On Xilinx that means a ram_style attribute; on Intel, ramstyle. At DEPTH = 128 and W = 32 the difference is 4096 flops against one block.

The read path is combinational from the memory in this design, which is what makes rd_data valid in the same cycle as !rd_empty. A registered block-RAM output adds a cycle and changes the read protocol — so a design that switches to a registered output must switch its consumers too, and the register map does not announce the change (Chapter 14.9 §7 has the same trap).

The two Gray pointers and their synchronisers are the only paths between the domains, so they get the CDC exception and everything else is constrained normally. That is the strongest practical argument for a FIFO over an ad-hoc crossing: the number of exceptions is two, and a reviewer can check two.

And AW, not DEPTH, is the parameter. A non-power-of-two depth breaks the one-extra-bit construction, because the pointer wrap no longer coincides with the address wrap. A design that needs 100 entries uses AW = 7 and accepts 128, or builds a different structure — and discovering that late is a redesign rather than a parameter change.

9. Failure Signature — A Slave That Loses Bytes Only During Interrupts

The symptom:

"Large SPI transfers occasionally lose a run of bytes in the middle. It correlates with network activity on the same CPU. Short transfers are always fine."

What is happening: §5, exactly. The receive FIFO's depth covers the normal reader latency and not the latency during an interrupt storm, so a burst arriving while the reader is blocked overflows — and the words that fit arrive correctly, which is why the loss is a run in the middle rather than corruption.

Why short transfers are fine: they finish inside the reader's normal service interval, so the FIFO never fills. The fault is a function of burst length and reader latency, and either one alone looks innocent.

How to size the fix rather than guess it: read overflow and the words actually received, and compute backwards. Words lost × LEN × SCLK period is the blocked time that caused it — which is a measurement of the real interrupt latency, obtained from the SPI slave. That is the most useful thing the flag does, and a design with an assertion instead of a flag cannot do it.

10. Common Misconceptions

"A FIFO fixes a clock-domain crossing." It relocates it. The data stops crossing and the pointers start, and the pointers are safe because they are Gray coded and read conservatively — not because they are in a FIFO.

"A deeper FIFO fixes an overflow." It fixes an overflow caused by a bounded reader stall. It does nothing for a sustained rate difference, which §3's last row is, and nothing for an unbounded stall.

"An occupancy counter is simpler than two pointers." It is simpler and unbuildable: both domains would write it, and there is no correct way to cross a value with two writers.

"The full and empty flags should be exact." They cannot be — each is computed from a pointer that crossed a boundary. What matters is the direction of the error, and it is arranged so that every stale reading reports less capacity rather than more.

"An overflow should be an assertion, because it means the design is broken." It means the environment exceeded what the depth covers, which is a system property the slave cannot control. An assertion converts a reportable condition into a simulation failure that never fires on hardware — and throws away §9's measurement.

"The Gray-code comparison for full is standard, so it is safe to copy." The textbook form has a lower bound at AW = 2 that the textbook does not mention, and a depth-2 FIFO built from it does not elaborate.

11. Reason It Through

Q. An 8-bit-frame slave at a 10 MHz SPI clock feeds a reader that can be blocked for up to 200 µs. What depth is needed, and what should the design do if that depth is unaffordable?

A word arrives every 8 × 100 ns = 800 ns, so 200 µs of blocking produces 250 words. AW = 8 gives 256 entries, which covers it. If 256 entries of the word width is unaffordable, the design cannot be made correct by a smaller number — so it should size for the typical latency, report overflow, and expose the count so the system can either bound its interrupt latency or shorten its bursts. Choosing a smaller depth silently is the only option that is definitely wrong.

Q. Why does a stale write pointer make the FIFO look emptier rather than fuller, and why is that the safe direction?

Because empty is read pointer equals synchronised write pointer, and a stale write pointer is an older, smaller value — so it is equal to the read pointer for longer, and empty stays asserted after a word has actually been written. The reader therefore waits for a word that is already there, which costs latency. The opposite — empty clearing before the word is committed — would let the reader take a location that has not been written, which is silent corruption. Latency is recoverable and corruption is not.

Q. A design changes DEPTH from 16 to 20. What breaks?

The one-extra-bit full/empty construction. With AW = 5 the pointers wrap at 32 while the addresses are used only up to 19, so "equal low bits with a differing top bit" no longer coincides with "the FIFO is full" — the pointers can differ by 20 without the top bits disagreeing. The construction requires DEPTH = 2^AW. A design needing 20 entries takes 32, or uses an occupancy scheme with a single-writer counter in each domain, which is a different and larger design.

Q. Why is it correct for the bench to write without consulting full, when a real system's writer might well check it?

Because this FIFO's writer is an SPI slave's receive path, and its source is SCLK. A slave cannot refuse a word the master is clocking in. So the honest model of the writer is one that writes regardless, and full exists to tell it that the word was dropped rather than to ask permission. A bench that waited for room would verify a FIFO in a system where the source can be stalled — which is Chapter 15.1's handshake, already ruled out.

Q. The read data path is combinational from the memory. What does that buy, and what would registering it cost beyond a cycle of latency?

It buys rd_data being valid in the same cycle rd_empty is low, so a reader can test and take in one access. Registering it would add a cycle and change the protocol: the consumer would have to assert rd_en, wait, then sample — and nothing in the register map announces that, so a driver written against the combinational version reads stale data on the registered one. That is the same compatibility break as Chapter 14.9's registered buffer read, and it is worth deciding once rather than discovering after silicon.

12. Understanding Check

13. Summary

A FIFO is what is left once a handshake has been ruled out, and it works because the data never crosses a boundary — a location is written by one clock and read by the other at a different time, so no data bit is ever sampled while it is changing.

What crosses is the two pointers, and they are exactly Chapter 15.5's special case: each increments by one, so Gray coding makes every partial observation the old value or the new one.

Every stale reading is conservative — a stale write pointer makes the FIFO look emptier and a stale read pointer makes it look fuller, so the reader waits and the writer holds off, and neither flag is ever optimistic. That is the property that makes the design safe rather than probable, and it is why full and empty come from different pointer pairs in different domains rather than from one occupancy count, which would have two writers and no correct way to cross.

The measurements: every word written arrived exactly once and in order at every ratio where the reader could keep up; sixty reads at an empty FIFO took nothing; forty writes at a depth-2 FIFO with no reader overflowed, reported it, and dropped the new word rather than damaging a held one. And at the one ratio where the writer outran the reader indefinitely, the shortfall was reported — because a FIFO decouples and does not create bandwidth.

The sizing question is the real one. Eight words arriving while the reader was blocked overflowed depths of 2 and 4 and fitted in 16, and once the reader resumed all eight came back in order — so a blocked reader costs latency and nothing else, provided the depth covers the block. The rule is DEPTH ≥ words produced while the destination is not draining, and when that time cannot be bounded no depth is sufficient, which is exactly why overflow is an output rather than an assertion.

For verification: instantiate several depths under one stimulus so the depth is a measured boundary; do not consult full in the write stimulus, because an SPI slave cannot stall SCLK; derive the conservation expectation from the sustained rates, or the last table row reports a correct design as broken; check the flags' direction rather than their exactness; and bind the one-bit-at-a-time Gray invariant, because without it the pointer crossing is scheme 1 under another name.

For implementation: a true dual-port block RAM and the attribute to get one — 4096 flops against one block at realistic sizes; a combinational read path that a registered output would silently make incompatible; two CDC exceptions in total, which is the strongest practical argument for a FIFO; and AW rather than DEPTH as the parameter, because a non-power-of-two depth breaks the construction entirely.

14. What Comes Next

Every crossing in this module has been built and measured, and every one of them rests on a constraints file that has been deferred four times.

Chapter 15.7 — Input and Output Delay Constraints writes it. The chapter's premise is that a constraint cannot be simulated — set_input_delay describes a board, and no board is present — so the RTL side is one register per pin and no logic, and the bench checks only what the constraints assume: that every input follows its pin after exactly one cycle, and the same one cycle. That shell also changes three numbers Module 14 published, which makes adding a pad register a datasheet change rather than an implementation detail.

Continue learning