Skip to content
VLSI Mentor

USB · Module 23

USB Timing Considerations

A multi-bit value cannot be bit-synchronised — each bit resolves independently, so the destination can sample a value that was never sent. The four-phase handshake crosses one bit and lets the data through untouched.

Chapters 23.1 to 23.4 all assumed one clock. A USB device does not have one clock.

The PHY runs at the bus rate — 60 MHz for a UTMI high-speed interface, derived from the recovered bit clock. The controller, the FIFOs and the CPU do not. Every value that moves between them crosses a boundary where the two clocks have no fixed relationship at all, and the rules on the other side of that boundary are different.

1. The Two-Flop Synchroniser, and Where It Stops Working

For one bit, the standard answer is correct and well understood: two flops in the destination domain, so the first flop's metastability has a full destination cycle to resolve before anything downstream looks at it.

Put one on each bit of a bus, and the design is broken.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   source changes    0111  ->  1000        (7 -> 8)

   Four bits change at once. Each bit's synchroniser resolves
   INDEPENDENTLY -- there is no mechanism that makes them agree.
   So the destination can sample:

       0111   1111   1000   0000   1011   0100   ...

   ANY of the sixteen values, for one destination cycle,
   before the bus settles.

Those intermediate values are not noise to be filtered out. They are indices into a descriptor table, byte counts, endpoint numbers. A transient 0000 on an endpoint number is a packet delivered to endpoint 0. A transient 1111 on a byte count is a DMA of 15 bytes that should have been 8.

Why a per-bit synchroniser fails on a bus

A four-bit bus crossing a clock boundary through four independent two-flop synchronisers. The source changes from 0111 to 1000, changing all four bits at once. Because each synchroniser resolves independently, the destination can observe any of sixteen intermediate values, including ones the source never drove.source bus0111 → 1000bit 3 syncresolves when it resolvesbit 2 syncindependentlybit 1 syncindependentlybit 0 syncindependentlydestinationany of 16 values0000never sent, acted on12
Each bit's pair of flops is correct in isolation and resolves on its own schedule. Nothing couples them, so a four-bit change produces up to sixteen possible sampled values — including values the source never held.

2. The Answer: Do Not Synchronise the Data At All

Synchronise one bit — a request — and use it to say "the data has been stable for a long time and will stay stable until I hear back."

The four-phase handshake

A sequence diagram of a four-phase handshake between a source clock domain and a destination clock domain. The source latches the data and raises req. Two synchroniser stages later the destination sees req, captures the data and raises ack. Two synchroniser stages later the source sees ack and drops req. The destination sees req fall and drops ack. The source sees ack fall and only then may it change the data.usb_cdc_handshake — one transfer, five stepsSource domain (clk_s)Destination domain (clk_d)1. latch the datainto holdreq rises — crossesas ONE bit2 flops later:capture hold, it hasnot movedack rises — also ONEbit3. two flops later:drop reqreq falls4. ack falls5. and only NOW maythe data change
One bit crosses each way. The data bus crosses with no synchroniser on it whatsoever, which is safe precisely because it is never sampled while it is changing.

The data path crosses the boundary with no synchroniser on it whatsoever. That looks alarming on a schematic and is the entire point: it is never sampled while it is changing, so it never needs one.

3. Why the Fourth Phase Cannot Be Dropped

The tempting simplification is three phases: the source drops req when it sees ack, and declares itself ready.

But ack is still high — the destination has not been told to drop it yet.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   three-phase:

     transfer N     req up ... ack up ... req down ... READY
     transfer N+1   req up -- and ack is STILL HIGH from N

                    the source sees "ack" immediately,
                    drops req in one cycle, completes in
                    zero time,

                    and the destination never captured
                    anything at all.

   The value is lost. No error. No stall. Nothing.

That is mutation C5 in §11, and it is the single most common way a hand-written handshake is wrong.

4. A Handshake Is Slow, and That Is the Trade

Four phases, each costing SYNC_N cycles in one domain or the other, plus edge detection. Roughly five to six cycles of the slower clock per transfer.

Use it forNot for
a device address, a configuration valuea byte stream
a status word, an interrupt flagpacket payload
a completion codeanything at line rate

For streaming data the answer is an asynchronous FIFO — which is this same idea applied to pointers rather than to payload: the pointers are Gray-coded (a counter, so single-bit transitions are guaranteed) and the memory is dual-ported, so the data never crosses a boundary at all. That is a different block and the reason this one is not it.

5. What We Are Building

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
  usb_cdc_handshake  #(DW = 8, SYNC_N = 2)

  source domain (clk_s)          destination domain (clk_d)
  ---------------------          --------------------------
  s_data / s_valid               d_data
  s_ready                        d_valid    ONE cycle per transfer
  s_req    the bit that crosses  d_ack      the bit that crosses back
  s_hold   the bus that crosses
           with NO synchroniser
  n_sent / n_stalls              n_received

s_hold is exposed deliberately. "It never moves while req is high" is a property, and a property you cannot observe is a property you are assuming.

6. Verilog-2005 Implementation

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// usb_cdc_handshake -- crossing a multi-bit value between two clocks, and
// the reason you cannot do it the obvious way.
//
// THE ASSUMPTION THAT JUST BROKE
//
// Chapters 23.1 to 23.4 all had one clock. A USB device does not. The PHY
// runs at the bus rate -- 60 MHz for a UTMI high-speed interface -- and the
// controller, the FIFOs and the CPU do not. Every value that moves between
// them crosses a boundary where the two clocks have no fixed relationship
// at all.
//
// A MULTI-BIT VALUE CANNOT BE BIT-SYNCHRONISED
//
// The two-flop synchroniser is the standard answer for a single bit, and it
// is correct: it gives the first flop's metastability time to resolve before
// anything downstream looks at it.
//
// Put one on each bit of a bus and the design is broken.
//
//     source changes   0111 -> 1000     (7 -> 8, four bits at once)
//
//     Each bit's synchroniser resolves INDEPENDENTLY. There is no
//     mechanism that makes them agree. The destination can sample
//
//         0111  1111  1000  0000  1011  0100  ...
//
//     any of the sixteen values, for one cycle, before settling.
//
// Those intermediate values are not noise to be filtered. They are indices
// into a descriptor table, byte counts, endpoint numbers. A transient 0000
// on an endpoint number is a packet delivered to endpoint 0.
//
// And a Gray code does not rescue the general case. It works for a counter,
// where only one bit changes per increment BY CONSTRUCTION. Arbitrary data
// changes arbitrarily many bits at once, and no encoding fixes that.
//
// THE ANSWER: DO NOT SYNCHRONISE THE DATA AT ALL
//
// Synchronise ONE bit -- a request -- and use it to say "the data has been
// stable for a long time and will stay stable until I hear back".
//
//     1. source latches the data, raises req
//     2. destination sees req (2 flops later), captures the data, raises ack
//     3. source sees ack (2 flops later), drops req
//     4. destination sees req drop, drops ack
//     5. source sees ack drop -- and only NOW may it change the data
//
// That is the FOUR-PHASE handshake, and every one of the four phases is
// load-bearing. The data path crosses the boundary with no synchroniser on
// it whatsoever, which looks alarming and is the entire point: it is not
// sampled while it is changing, so it never needs one.
//
// WHY PHASE 4 CANNOT BE DROPPED
//
// The tempting simplification is three phases: the source drops req on ack
// and declares itself ready. But ack is still HIGH -- it has not been told
// to drop yet -- so the next transfer starts, raises req, and immediately
// sees the PREVIOUS ack. It completes in zero time, the destination never
// captures, and the value is silently lost.
//
// A HANDSHAKE IS SLOW, AND THAT IS THE TRADE
//
// Four phases, each costing SYNC_N destination or source cycles plus edge
// detection. Roughly 5 to 6 cycles of the slower clock per transfer. For
// control values -- an address, a configuration, a status word -- that is
// free. For streaming data it is not, and the answer there is an
// asynchronous FIFO, which is this handshake's idea applied to pointers
// instead of to payload.
module usb_cdc_handshake #(
  parameter integer DW     = 8,
  parameter integer SYNC_N = 2   // synchroniser depth (>= 2; see below)
) (
  // ---- source domain ----
  input  wire          clk_s,
  input  wire          rst_s_n,
  input  wire [DW-1:0] s_data,
  input  wire          s_valid,   // the source has a value to send
  output wire          s_ready,   // ...and this says it may send it
  output wire          s_req,     // THE wire that crosses, and it is 1 bit
  output wire [DW-1:0] s_hold,    // ...and this is the bus that crosses with
                                  // no synchroniser on it at all. Exposed
                                  // because "it never moves while req is
                                  // high" is a property, and a property you
                                  // cannot observe is a property you are
                                  // assuming.
  output wire [1:0]    s_state,
  output wire [31:0]   n_sent,
  output wire [31:0]   n_stalls,

  // ---- destination domain ----
  input  wire          clk_d,
  input  wire          rst_d_n,
  output wire [DW-1:0] d_data,
  output wire          d_valid,   // ONE destination cycle per transfer
  output wire          d_ack,     // the wire that crosses back
  output wire [31:0]   n_received
);

  localparam [1:0] SS_IDLE = 2'd0,  // ready; the data may be changed
                   SS_REQ  = 2'd1,  // req raised, waiting for ack
                   SS_DROP = 2'd2;  // req dropped, waiting for ack to fall

  // ==================================================================
  // SOURCE DOMAIN
  // ==================================================================
  reg [1:0]    ss_r;
  reg [DW-1:0] hold_r;     // the CDC data path. NOT synchronised. Held.
  reg          req_r;
  reg [31:0]   sent_r, stalls_r;

  // The ack, brought into the source domain through SYNC_N flops. This is
  // a single bit, which is the only kind of signal a synchroniser is
  // allowed to be used on.
  reg [SYNC_N-1:0] ack_sync_r;

  wire ack_seen = ack_sync_r[SYNC_N-1];

  assign s_req   = req_r;
  assign s_hold  = hold_r;
  assign s_state = ss_r;
  assign s_ready = (ss_r == SS_IDLE);
  assign n_sent  = sent_r;
  assign n_stalls = stalls_r;

  always @(posedge clk_s or negedge rst_s_n) begin
    if (!rst_s_n) begin
      ss_r       <= SS_IDLE;
      hold_r     <= {DW{1'b0}};
      req_r      <= 1'b0;
      ack_sync_r <= {SYNC_N{1'b0}};
      sent_r     <= 32'd0;
      stalls_r   <= 32'd0;
    end else begin
      ack_sync_r <= {ack_sync_r[SYNC_N-2:0], d_ack};

      case (ss_r)
        SS_IDLE: begin
          if (s_valid) begin
            // ---- PHASE 1. Latch the data, THEN raise req. ----
            //
            // The latch is what makes the data path safe: from here until
            // phase 5 the destination is looking at a value that does not
            // move. Passing s_data through combinationally instead would
            // put a live, changing bus across the boundary -- which is the
            // bit-synchroniser bug with the synchronisers removed.
            hold_r <= s_data;
            req_r  <= 1'b1;
            ss_r   <= SS_REQ;
          end
        end

        SS_REQ: begin
          // ---- PHASE 3. The destination has captured it. Drop req. ----
          if (ack_seen) begin
            req_r <= 1'b0;
            ss_r  <= SS_DROP;
          end
        end

        SS_DROP: begin
          // ---- PHASE 5. And WAIT for ack to fall before saying ready.
          //
          // Returning to IDLE here without waiting is the three-phase
          // shortcut: the next req would be raised while ack is still high
          // from the previous transfer, the source would see it instantly,
          // and the destination would never capture anything.
          if (!ack_seen) begin
            ss_r   <= SS_IDLE;
            sent_r <= sent_r + 32'd1;
          end
        end

        default: ss_r <= SS_IDLE;
      endcase

      // A source that wants to send while the handshake is busy. Counted,
      // because the cost of a handshake is exactly this and it should be
      // measured rather than assumed.
      if (s_valid && (ss_r != SS_IDLE)) stalls_r <= stalls_r + 32'd1;
    end
  end

  // ==================================================================
  // DESTINATION DOMAIN
  // ==================================================================
  reg [SYNC_N-1:0] req_sync_r;
  reg              req_d_r;      // one more flop, for edge detection
  reg [DW-1:0]     d_data_r;
  reg              d_valid_r, ack_r;
  reg [31:0]       recv_r;

  wire req_seen = req_sync_r[SYNC_N-1];

  // The RISING EDGE, not the level. A level would re-capture and re-pulse
  // on every destination cycle for as long as req is high, which at a fast
  // destination clock is a single transfer delivered dozens of times.
  wire req_rise = req_seen && !req_d_r;

  assign d_data     = d_data_r;
  assign d_valid    = d_valid_r;
  assign d_ack      = ack_r;
  assign n_received = recv_r;

  always @(posedge clk_d or negedge rst_d_n) begin
    if (!rst_d_n) begin
      req_sync_r <= {SYNC_N{1'b0}};
      req_d_r    <= 1'b0;
      d_data_r   <= {DW{1'b0}};
      d_valid_r  <= 1'b0;
      ack_r      <= 1'b0;
      recv_r     <= 32'd0;
    end else begin
      req_sync_r <= {req_sync_r[SYNC_N-2:0], s_req};
      req_d_r    <= req_seen;

      d_valid_r <= 1'b0;

      if (req_rise) begin
        // ---- PHASE 2. Capture the data and acknowledge. ----
        //
        // hold_r has been stable for at least SYNC_N destination cycles by
        // now -- that is precisely what the synchroniser bought -- so this
        // read is of a static value and needs no synchroniser of its own.
        d_data_r  <= hold_r;
        d_valid_r <= 1'b1;
        ack_r     <= 1'b1;
        recv_r    <= recv_r + 32'd1;
      end else if (!req_seen) begin
        // ---- PHASE 4. req has fallen; release ack. ----
        ack_r <= 1'b0;
      end
    end
  end
endmodule

7. SystemVerilog Implementation

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// usb_cdc_handshake -- crossing a multi-bit value between two clocks, and
// the reason you cannot do it the obvious way.
//
// THE ASSUMPTION THAT JUST BROKE
//
// Chapters 23.1 to 23.4 all had one clock. A USB device does not. The PHY
// runs at the bus rate -- 60 MHz for a UTMI high-speed interface -- and the
// controller, the FIFOs and the CPU do not. Every value that moves between
// them crosses a boundary where the two clocks have no fixed relationship
// at all.
//
// A MULTI-BIT VALUE CANNOT BE BIT-SYNCHRONISED
//
// The two-flop synchroniser is the standard answer for a single bit, and it
// is correct: it gives the first flop's metastability time to resolve before
// anything downstream looks at it.
//
// Put one on each bit of a bus and the design is broken.
//
//     source changes   0111 -> 1000     (7 -> 8, four bits at once)
//
//     Each bit's synchroniser resolves INDEPENDENTLY. There is no
//     mechanism that makes them agree. The destination can sample
//
//         0111  1111  1000  0000  1011  0100  ...
//
//     any of the sixteen values, for one cycle, before settling.
//
// Those intermediate values are not noise to be filtered. They are indices
// into a descriptor table, byte counts, endpoint numbers. A transient 0000
// on an endpoint number is a packet delivered to endpoint 0.
//
// And a Gray code does not rescue the general case. It works for a counter,
// where only one bit changes per increment BY CONSTRUCTION. Arbitrary data
// changes arbitrarily many bits at once, and no encoding fixes that.
//
// THE ANSWER: DO NOT SYNCHRONISE THE DATA AT ALL
//
// Synchronise ONE bit -- a request -- and use it to say "the data has been
// stable for a long time and will stay stable until I hear back".
//
//     1. source latches the data, raises req
//     2. destination sees req (2 flops later), captures the data, raises ack
//     3. source sees ack (2 flops later), drops req
//     4. destination sees req drop, drops ack
//     5. source sees ack drop -- and only NOW may it change the data
//
// That is the FOUR-PHASE handshake, and every one of the four phases is
// load-bearing. The data path crosses the boundary with no synchroniser on
// it whatsoever, which looks alarming and is the entire point: it is not
// sampled while it is changing, so it never needs one.
//
// WHY PHASE 4 CANNOT BE DROPPED
//
// The tempting simplification is three phases: the source drops req on ack
// and declares itself ready. But ack is still HIGH -- it has not been told
// to drop yet -- so the next transfer starts, raises req, and immediately
// sees the PREVIOUS ack. It completes in zero time, the destination never
// captures, and the value is silently lost.
//
// A HANDSHAKE IS SLOW, AND THAT IS THE TRADE
//
// Four phases, each costing SYNC_N destination or source cycles plus edge
// detection. Roughly 5 to 6 cycles of the slower clock per transfer. For
// control values -- an address, a configuration, a status word -- that is
// free. For streaming data it is not, and the answer there is an
// asynchronous FIFO, which is this handshake's idea applied to pointers
// instead of to payload.
package usb_cdc_pkg;
  // The three source-side phases, named. They are an enumeration rather
  // than an encoding because SS_DROP exists ONLY to wait for the fourth
  // phase of the handshake, and a name is the cheapest way to stop somebody
  // deleting it.
  typedef enum logic [1:0] {
    SS_IDLE = 2'd0,   // ready; the data may be changed
    SS_REQ  = 2'd1,   // req raised, waiting for ack
    SS_DROP = 2'd2    // req dropped, waiting for ack to fall
  } src_state_e;
endpackage

module usb_cdc_handshake
  import usb_cdc_pkg::*;
#(
  parameter int DW     = 8,
  parameter int SYNC_N = 2   // synchroniser depth (>= 2; see below)
) (
  // ---- source domain ----
  input  logic          clk_s,
  input  logic          rst_s_n,
  input  logic [DW-1:0] s_data,
  input  logic          s_valid,   // the source has a value to send
  output logic          s_ready,   // ...and this says it may send it
  output logic          s_req,     // THE wire that crosses, and it is 1 bit
  output logic [DW-1:0] s_hold,    // ...and this is the bus that crosses with
                                   // no synchroniser on it at all. Exposed
                                   // because "it never moves while req is
                                   // high" is a property, and a property you
                                   // cannot observe is a property you are
                                   // assuming.
  output src_state_e    s_state,
  output logic [31:0]   n_sent,
  output logic [31:0]   n_stalls,

  // ---- destination domain ----
  input  logic          clk_d,
  input  logic          rst_d_n,
  output logic [DW-1:0] d_data,
  output logic          d_valid,   // ONE destination cycle per transfer
  output logic          d_ack,     // the wire that crosses back
  output logic [31:0]   n_received
);

  // ==================================================================
  // SOURCE DOMAIN
  // ==================================================================
  src_state_e    ss_r;
  logic [DW-1:0] hold_r;    // the CDC data path. NOT synchronised. Held.
  logic          req_r;
  logic [31:0]   sent_r, stalls_r;

  // The ack, brought into the source domain through SYNC_N flops. This is
  // a single bit, which is the only kind of signal a synchroniser is
  // allowed to be used on.
  logic [SYNC_N-1:0] ack_sync_r;

  logic ack_seen;
  assign ack_seen = ack_sync_r[SYNC_N-1];

  assign s_req    = req_r;
  assign s_hold   = hold_r;
  assign s_state  = ss_r;
  assign s_ready  = (ss_r == SS_IDLE);
  assign n_sent   = sent_r;
  assign n_stalls = stalls_r;

  always_ff @(posedge clk_s or negedge rst_s_n) begin
    if (!rst_s_n) begin
      ss_r       <= SS_IDLE;
      hold_r     <= '0;
      req_r      <= 1'b0;
      ack_sync_r <= '0;
      sent_r     <= '0;
      stalls_r   <= '0;
    end else begin
      ack_sync_r <= {ack_sync_r[SYNC_N-2:0], d_ack};

      unique case (ss_r)
        SS_IDLE: begin
          if (s_valid) begin
            // ---- PHASE 1. Latch the data, THEN raise req. ----
            //
            // The latch is what makes the data path safe: from here until
            // phase 5 the destination is looking at a value that does not
            // move. Passing s_data through combinationally instead would
            // put a live, changing bus across the boundary -- which is the
            // bit-synchroniser bug with the synchronisers removed.
            hold_r <= s_data;
            req_r  <= 1'b1;
            ss_r   <= SS_REQ;
          end
        end

        SS_REQ: begin
          // ---- PHASE 3. The destination has captured it. Drop req. ----
          if (ack_seen) begin
            req_r <= 1'b0;
            ss_r  <= SS_DROP;
          end
        end

        SS_DROP: begin
          // ---- PHASE 5. And WAIT for ack to fall before saying ready.
          //
          // Returning to IDLE here without waiting is the three-phase
          // shortcut: the next req would be raised while ack is still high
          // from the previous transfer, the source would see it instantly,
          // and the destination would never capture anything.
          if (!ack_seen) begin
            ss_r   <= SS_IDLE;
            sent_r <= sent_r + 32'd1;
          end
        end

        default: ss_r <= SS_IDLE;
      endcase

      // A source that wants to send while the handshake is busy. Counted,
      // because the cost of a handshake is exactly this and it should be
      // measured rather than assumed.
      if (s_valid && (ss_r != SS_IDLE)) stalls_r <= stalls_r + 32'd1;
    end
  end

  // ==================================================================
  // DESTINATION DOMAIN
  // ==================================================================
  logic [SYNC_N-1:0] req_sync_r;
  logic              req_d_r;      // one more flop, for edge detection
  logic [DW-1:0]     d_data_r;
  logic              d_valid_r, ack_r;
  logic [31:0]       recv_r;

  logic req_seen, req_rise;
  assign req_seen = req_sync_r[SYNC_N-1];

  // The RISING EDGE, not the level. A level would re-capture and re-pulse
  // on every destination cycle for as long as req is high, which at a fast
  // destination clock is a single transfer delivered dozens of times.
  assign req_rise = req_seen && !req_d_r;

  assign d_data     = d_data_r;
  assign d_valid    = d_valid_r;
  assign d_ack      = ack_r;
  assign n_received = recv_r;

  always_ff @(posedge clk_d or negedge rst_d_n) begin
    if (!rst_d_n) begin
      req_sync_r <= '0;
      req_d_r    <= 1'b0;
      d_data_r   <= '0;
      d_valid_r  <= 1'b0;
      ack_r      <= 1'b0;
      recv_r     <= '0;
    end else begin
      req_sync_r <= {req_sync_r[SYNC_N-2:0], s_req};
      req_d_r    <= req_seen;

      d_valid_r <= 1'b0;

      if (req_rise) begin
        // ---- PHASE 2. Capture the data and acknowledge. ----
        //
        // hold_r has been stable for at least SYNC_N destination cycles by
        // now -- that is precisely what the synchroniser bought -- so this
        // read is of a static value and needs no synchroniser of its own.
        d_data_r  <= hold_r;
        d_valid_r <= 1'b1;
        ack_r     <= 1'b1;
        recv_r    <= recv_r + 32'd1;
      end else if (!req_seen) begin
        // ---- PHASE 4. req has fallen; release ack. ----
        ack_r <= 1'b0;
      end
    end
  end
endmodule

8. VHDL-2008 Implementation

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
-- usb_cdc_handshake -- crossing a multi-bit value between two clocks, and
-- the reason you cannot do it the obvious way.
--
-- THE ASSUMPTION THAT JUST BROKE
--
-- Chapters 23.1 to 23.4 all had one clock. A USB device does not. The PHY
-- runs at the bus rate -- 60 MHz for a UTMI high-speed interface -- and the
-- controller, the FIFOs and the CPU do not. Every value that moves between
-- them crosses a boundary where the two clocks have no fixed relationship
-- at all.
--
-- A MULTI-BIT VALUE CANNOT BE BIT-SYNCHRONISED
--
-- The two-flop synchroniser is the standard answer for a single bit, and it
-- is correct: it gives the first flop's metastability time to resolve before
-- anything downstream looks at it.
--
-- Put one on each bit of a bus and the design is broken.
--
--     source changes   0111 -> 1000     (7 -> 8, four bits at once)
--
--     Each bit's synchroniser resolves INDEPENDENTLY. There is no
--     mechanism that makes them agree. The destination can sample
--
--         0111  1111  1000  0000  1011  0100  ...
--
--     any of the sixteen values, for one cycle, before settling.
--
-- Those intermediate values are not noise to be filtered. They are indices
-- into a descriptor table, byte counts, endpoint numbers. A transient 0000
-- on an endpoint number is a packet delivered to endpoint 0.
--
-- And a Gray code does not rescue the general case. It works for a counter,
-- where only one bit changes per increment BY CONSTRUCTION. Arbitrary data
-- changes arbitrarily many bits at once, and no encoding fixes that.
--
-- THE ANSWER: DO NOT SYNCHRONISE THE DATA AT ALL
--
-- Synchronise ONE bit -- a request -- and use it to say "the data has been
-- stable for a long time and will stay stable until I hear back".
--
--     1. source latches the data, raises req
--     2. destination sees req (2 flops later), captures the data, raises ack
--     3. source sees ack (2 flops later), drops req
--     4. destination sees req drop, drops ack
--     5. source sees ack drop -- and only NOW may it change the data
--
-- That is the FOUR-PHASE handshake, and every one of the four phases is
-- load-bearing. The data path crosses the boundary with no synchroniser on
-- it whatsoever, which looks alarming and is the entire point: it is not
-- sampled while it is changing, so it never needs one.
--
-- WHY PHASE 4 CANNOT BE DROPPED
--
-- The tempting simplification is three phases: the source drops req on ack
-- and declares itself ready. But ack is still HIGH -- it has not been told
-- to drop yet -- so the next transfer starts, raises req, and immediately
-- sees the PREVIOUS ack. It completes in zero time, the destination never
-- captures, and the value is silently lost.
--
-- A HANDSHAKE IS SLOW, AND THAT IS THE TRADE
--
-- Four phases, each costing SYNC_N destination or source cycles plus edge
-- detection. Roughly 5 to 6 cycles of the slower clock per transfer. For
-- control values -- an address, a configuration, a status word -- that is
-- free. For streaming data it is not, and the answer there is an
-- asynchronous FIFO, which is this handshake's idea applied to pointers
-- instead of to payload.
library ieee;
use ieee.std_logic_1164.all;

package usb_cdc_pkg is
  -- The three source-side phases, named. They are an enumeration rather
  -- than an encoding because SS_DROP exists ONLY to wait for the fourth
  -- phase of the handshake, and a name is the cheapest way to stop somebody
  -- deleting it.
  type src_state_t is (SS_IDLE, SS_REQ, SS_DROP);
  function ss_code (s : src_state_t) return std_logic_vector;
end package usb_cdc_pkg;

package body usb_cdc_pkg is
  function ss_code (s : src_state_t) return std_logic_vector is
  begin
    case s is
      when SS_IDLE => return "00";
      when SS_REQ  => return "01";
      when SS_DROP => return "10";
    end case;
  end function;
end package body usb_cdc_pkg;

library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
use work.usb_cdc_pkg.all;

entity usb_cdc_handshake is
  generic (
    DW     : integer := 8;
    SYNC_N : integer := 2    -- synchroniser depth (>= 2; see below)
  );
  port (
    -- ---- source domain ----
    clk_s      : in  std_logic;
    rst_s_n    : in  std_logic;
    s_data     : in  std_logic_vector(DW-1 downto 0);
    s_valid    : in  std_logic;                        -- a value to send
    s_ready    : out std_logic;                        -- ...and it may go
    s_req      : out std_logic;                        -- THE wire: 1 bit
    s_hold     : out std_logic_vector(DW-1 downto 0);  -- the bus that
                                                       -- crosses with no
                                                       -- synchroniser at all
    s_state    : out std_logic_vector(1 downto 0);
    n_sent     : out std_logic_vector(31 downto 0);
    n_stalls   : out std_logic_vector(31 downto 0);

    -- ---- destination domain ----
    clk_d      : in  std_logic;
    rst_d_n    : in  std_logic;
    d_data     : out std_logic_vector(DW-1 downto 0);
    d_valid    : out std_logic;   -- ONE destination cycle per transfer
    d_ack      : out std_logic;   -- the wire that crosses back
    n_received : out std_logic_vector(31 downto 0)
  );
end entity usb_cdc_handshake;

architecture rtl of usb_cdc_handshake is

  -- ================================================================
  -- SOURCE DOMAIN
  -- ================================================================
  signal ss_r     : src_state_t := SS_IDLE;
  signal hold_r   : std_logic_vector(DW-1 downto 0) := (others => '0');
  signal req_r    : std_logic := '0';
  signal sent_r   : unsigned(31 downto 0) := (others => '0');
  signal stalls_r : unsigned(31 downto 0) := (others => '0');

  -- The ack, brought into the source domain through SYNC_N flops. This is
  -- a single bit, which is the only kind of signal a synchroniser is
  -- allowed to be used on.
  signal ack_sync_r : std_logic_vector(SYNC_N-1 downto 0) := (others => '0');
  signal ack_seen   : std_logic;

  -- ================================================================
  -- DESTINATION DOMAIN
  -- ================================================================
  signal req_sync_r : std_logic_vector(SYNC_N-1 downto 0) := (others => '0');
  signal req_d_r    : std_logic := '0';
  signal d_data_r   : std_logic_vector(DW-1 downto 0) := (others => '0');
  signal d_valid_r  : std_logic := '0';
  signal ack_r      : std_logic := '0';
  signal recv_r     : unsigned(31 downto 0) := (others => '0');
  signal req_seen, req_rise : std_logic;

begin

  ack_seen <= ack_sync_r(SYNC_N-1);

  s_req    <= req_r;
  s_hold   <= hold_r;
  s_state  <= ss_code(ss_r);
  s_ready  <= '1' when ss_r = SS_IDLE else '0';
  n_sent   <= std_logic_vector(sent_r);
  n_stalls <= std_logic_vector(stalls_r);

  src : process (clk_s, rst_s_n)
  begin
    if rst_s_n = '0' then
      ss_r       <= SS_IDLE;
      hold_r     <= (others => '0');
      req_r      <= '0';
      ack_sync_r <= (others => '0');
      sent_r     <= (others => '0');
      stalls_r   <= (others => '0');
    elsif rising_edge(clk_s) then
      ack_sync_r <= ack_sync_r(SYNC_N-2 downto 0) & d_ack;

      case ss_r is
        when SS_IDLE =>
          if s_valid = '1' then
            -- ---- PHASE 1. Latch the data, THEN raise req. ----
            --
            -- The latch is what makes the data path safe: from here until
            -- phase 5 the destination is looking at a value that does not
            -- move. Passing s_data through combinationally instead would
            -- put a live, changing bus across the boundary -- which is the
            -- bit-synchroniser bug with the synchronisers removed.
            hold_r <= s_data;
            req_r  <= '1';
            ss_r   <= SS_REQ;
          end if;

        when SS_REQ =>
          -- ---- PHASE 3. The destination has captured it. Drop req. ----
          if ack_seen = '1' then
            req_r <= '0';
            ss_r  <= SS_DROP;
          end if;

        when SS_DROP =>
          -- ---- PHASE 5. And WAIT for ack to fall before saying ready.
          --
          -- Returning to IDLE here without waiting is the three-phase
          -- shortcut: the next req would be raised while ack is still high
          -- from the previous transfer, the source would see it instantly,
          -- and the destination would never capture anything.
          if ack_seen = '0' then
            ss_r   <= SS_IDLE;
            sent_r <= sent_r + 1;
          end if;
      end case;

      -- A source that wants to send while the handshake is busy. Counted,
      -- because the cost of a handshake is exactly this and it should be
      -- measured rather than assumed.
      if s_valid = '1' and ss_r /= SS_IDLE then
        stalls_r <= stalls_r + 1;
      end if;
    end if;
  end process;

  req_seen <= req_sync_r(SYNC_N-1);

  -- The RISING EDGE, not the level. A level would re-capture and re-pulse
  -- on every destination cycle for as long as req is high, which at a fast
  -- destination clock is a single transfer delivered dozens of times.
  req_rise <= req_seen and (not req_d_r);

  d_data     <= d_data_r;
  d_valid    <= d_valid_r;
  d_ack      <= ack_r;
  n_received <= std_logic_vector(recv_r);

  dst : process (clk_d, rst_d_n)
  begin
    if rst_d_n = '0' then
      req_sync_r <= (others => '0');
      req_d_r    <= '0';
      d_data_r   <= (others => '0');
      d_valid_r  <= '0';
      ack_r      <= '0';
      recv_r     <= (others => '0');
    elsif rising_edge(clk_d) then
      -- req_r rather than the s_req port: VHDL will not let an entity read
      -- its own output, and this is the identical wire. It is the one bit
      -- that crosses the boundary in this direction.
      req_sync_r <= req_sync_r(SYNC_N-2 downto 0) & req_r;
      req_d_r    <= req_seen;

      d_valid_r <= '0';

      if req_rise = '1' then
        -- ---- PHASE 2. Capture the data and acknowledge. ----
        --
        -- hold_r has been stable for at least SYNC_N destination cycles by
        -- now -- that is precisely what the synchroniser bought -- so this
        -- read is of a static value and needs no synchroniser of its own.
        d_data_r  <= hold_r;
        d_valid_r <= '1';
        ack_r     <= '1';
        recv_r    <= recv_r + 1;
      elsif req_seen = '0' then
        -- ---- PHASE 4. req has fallen; release ack. ----
        ack_r <= '0';
      end if;
    end if;
  end process;

end architecture rtl;

9. Seeing One Transfer Cross

One value crossing, with the destination clock slower than the source

usb_cdc_handshake — the four phases, in destination cycles

10 cycles
A ten-cycle waveform drawn in destination cycles. The source raises req with the held data already stable. Two synchroniser stages later req_sync is high, and on the following destination cycle d_valid pulses for exactly one cycle and ack rises. The source then drops req, the destination drops ack, and the held data is released only after that.2 sync flops: req is safe to use2 sync flops: req is safeto usecapture; one cycle of d_validcapture; one cycle ofd_validphase 5: only now may data movephase 5: only now may datamoveclk_ds_reqs_hold--5A5A5A5A5A5A5A5AC3req_syncd_validd_data--------5A5A5A5A5A5Ad_acks_readyt0t1t2t3t4t5t6t7t8t9
req rises at source cycle 1. The destination's two synchroniser flops fill over its next two cycles and the capture happens on the third — always three destination cycles, whatever the ratio. ack then takes the same route back, and only after it falls is the source free to change the data.

s_hold holds 5A for eight destination cycles and changes only at cycle 9, after s_ready returns. Every one of those cycles is a cycle in which the destination could have been reading it, and none of them is a cycle in which it was moving.

10. The Testbenches

The block claims to work at any clock ratio, so the suite sweeps the ratio instead of picking one:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   5 : 7     destination a little slower
   7 : 5     source a little slower
   5 : 31    destination MUCH slower
   31 : 5    source MUCH slower
   3 : 11    co-prime, close
   11 : 3
   5 : 6     nearly equal, never aligned
   6 : 5
   13 : 13   equal periods, offset phase
   2 : 29    extreme
   29 : 2

   ...and then a phase where BOTH periods change on every
   transfer, while transfers are in flight.

The two clocks are generated with a deliberate phase offset so that no edge in one domain ever coincides exactly with an edge in the other — which is the one alignment a real chip is guaranteed not to have, and the one a lazy testbench accidentally arranges.

Four things are checked:

  1. Every value arrives exactly once, in order, unchanged, compared element by element against an independent queue.
  2. The bus that crosses never moves while req is high. The driver deliberately puts a different value on s_data the instant a transfer is taken.
  3. d_valid is exactly one destination cycle per transfer.
  4. The minimum observed latency is at least SYNC_N + 1 destination cycles.

And every wait is bounded:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
      // Every wait is BOUNDED. A handshake that has lost a transfer does
      // not produce a wrong answer -- it produces no answer at all, and an
      // unbounded wait turns that into a suite that hangs instead of a
      // suite that fails. The bound is far larger than the slowest legal
      // handshake at the most extreme clock ratio in the sweep.
      guard = 0;
      while (!s_ready && (guard < 600)) begin @(posedge clk_s); #1; guard = guard + 1; end
      check(s_ready === 1'b1,
            "the source never became ready -- the handshake is deadlocked, which is what a lost ack looks like from this side");

10.1 Verilog testbench

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// Testbench for usb_cdc_handshake (Verilog-2005).
//
// TWO CLOCKS, AND NO RELATIONSHIP BETWEEN THEM
//
// The whole point of the block is that it works for ANY ratio, so the suite
// sweeps the ratio rather than picking one: source faster, destination
// faster, nearly equal, wildly unequal, and a phase where the destination
// period CHANGES while transfers are in flight.
//
// WHAT IS CHECKED
//
//   1. Every value sent arrives EXACTLY ONCE, IN ORDER, UNCHANGED, at
//      every ratio. Compared element by element against an independent
//      queue, not by counting.
//
//   2. The bus that crosses NEVER MOVES while req is high. That is the
//      property that makes an unsynchronised data path safe, and the
//      source deliberately drives a different value onto s_data the
//      instant the transfer is taken, so a design that passed s_data
//      through combinationally would be caught immediately.
//
//   3. d_valid is exactly ONE destination cycle per transfer. A level-
//      triggered capture delivers the same value dozens of times at a fast
//      destination clock, and a counter alone would not notice the
//      difference between that and a fast source.
//
//   4. The minimum observed latency is at least SYNC_N+1 destination
//      cycles -- the signature of a synchroniser that really is SYNC_N
//      deep. See the note in the chapter about what this check can and
//      cannot prove.
`timescale 1ns/1ps
module tb_ch_v;

  localparam integer DW     = 8;
  localparam integer SYNC_N = 2;

  localparam [1:0] SS_IDLE = 2'd0, SS_REQ = 2'd1, SS_DROP = 2'd2;

  integer sp = 5;     // source half period
  integer dp = 7;     // destination half period

  reg clk_s = 1'b0, rst_s_n = 1'b0;
  reg clk_d = 1'b0, rst_d_n = 1'b0;
  reg [DW-1:0] s_data = 8'd0;
  reg s_valid = 1'b0;

  wire s_ready, s_req, d_valid, d_ack;
  wire [DW-1:0] s_hold, d_data;
  wire [1:0] s_state;
  wire [31:0] n_sent, n_stalls, n_received;

  usb_cdc_handshake #(.DW(DW), .SYNC_N(SYNC_N)) dut (
    .clk_s(clk_s), .rst_s_n(rst_s_n),
    .s_data(s_data), .s_valid(s_valid), .s_ready(s_ready),
    .s_req(s_req), .s_hold(s_hold), .s_state(s_state),
    .n_sent(n_sent), .n_stalls(n_stalls),
    .clk_d(clk_d), .rst_d_n(rst_d_n),
    .d_data(d_data), .d_valid(d_valid), .d_ack(d_ack),
    .n_received(n_received)
  );

  // Two free-running clocks with a deliberate phase offset, so no edge in
  // one domain ever coincides exactly with an edge in the other -- which
  // would be the one alignment a real chip is guaranteed NOT to have.
  initial forever #sp clk_s = ~clk_s;
  initial begin #3; forever #dp clk_d = ~clk_d; end

  integer errors = 0, checks = 0;
  task check(input cond, input [1023:0] msg);
    begin
      checks = checks + 1;
      if (!cond) begin
        errors = errors + 1;
        if (errors <= 25)
          $display("FAIL @%0t: %0s | st=%0d req=%b ack=%b hold=%h d=%h",
                   $time, msg, s_state, s_req, d_ack, s_hold, d_data);
      end
    end
  endtask

  // ---- the two independent queues ----
  reg [DW-1:0] txq [0:20000];
  reg [DW-1:0] rxq [0:20000];
  integer tx_n = 0, rx_n = 0;
  integer min_lat = 1000000, max_lat = 0;
  integer n_ratios = 0;

  // ---- destination-side observation ----
  integer d_age = 0;
  reg d_valid_d = 1'b0;

  always @(posedge clk_d or negedge rst_d_n) begin
    if (!rst_d_n) d_age <= 0;
    else if (!s_req) d_age <= 0;
    else d_age <= d_age + 1;
  end

  always @(posedge clk_d) begin
    if (rst_d_n) begin
      // ---- d_valid is ONE cycle. Two in a row is a level-triggered
      // ---- capture delivering the same value repeatedly.
      check(!(d_valid && d_valid_d),
            "d_valid was asserted for two consecutive destination cycles -- the capture is level-triggered, not edge-triggered");
      d_valid_d <= d_valid;

      if (d_valid) begin
        // The index is bounded. A level-triggered capture delivers orders
        // of magnitude more values than a correct design does, and a queue
        // that runs off its end turns a measured kill into a crash.
        if (rx_n <= 20000) rxq[rx_n] = d_data;
        rx_n = rx_n + 1;
        if (d_age < min_lat) min_lat = d_age;
        if (d_age > max_lat) max_lat = d_age;
        // ---- the synchroniser really is SYNC_N deep ----
        check(d_age >= SYNC_N + 1,
              "a value arrived sooner than SYNC_N synchroniser stages can possibly allow -- a stage has been bypassed");
        // Note what is NOT checked here: that req is still high. The
        // monitor observes d_valid one destination cycle after the capture,
        // and a fast source can already have seen ack and completed phases
        // 3 to 5 by then. "One delivery per request" is a statement about
        // totals, so it is checked as one -- see n_req_rises below.
      end
    end
  end

  // ---- source-side observation: the bus that crosses must NOT MOVE ----
  reg [DW-1:0] hold_prev = 8'd0;
  reg          req_prev  = 1'b0;
  integer      n_req_rises = 0;
  always @(posedge clk_s) begin
    if (rst_s_n) begin
      // From the SECOND cycle of req onward. The edge that raises req is
      // the same edge that loads the hold register, so comparing across it
      // would flag the load itself.
      if (s_req && req_prev)
        check(s_hold === hold_prev,
              "the bus that crosses the boundary changed while req was asserted -- the destination may be sampling it mid-change, which is the bit-synchroniser bug in its purest form");
      // ---- and the source must not declare itself ready mid-transfer ----
      check(!(s_ready && s_req),
            "s_ready was asserted while a transfer was still in flight -- the next value can overwrite the one being captured");
      check(!(s_ready && (s_state != SS_IDLE)),
            "s_ready disagrees with the source state");
      if (s_req && !req_prev) n_req_rises = n_req_rises + 1;
    end
    hold_prev <= s_hold;
    req_prev  <= s_req;
  end

  // Hand one value to the source, wait until the handshake has taken it,
  // and then IMMEDIATELY drive something else onto s_data -- so that a
  // design which forwarded s_data combinationally is caught at once.
  integer guard;
  task send(input [DW-1:0] v);
    begin
      @(posedge clk_s); #1;
      // Every wait is BOUNDED. A handshake that has lost a transfer does
      // not produce a wrong answer -- it produces no answer at all, and an
      // unbounded wait turns that into a suite that hangs instead of a
      // suite that fails. The bound is far larger than the slowest legal
      // handshake at the most extreme clock ratio in the sweep.
      guard = 0;
      while (!s_ready && (guard < 600)) begin @(posedge clk_s); #1; guard = guard + 1; end
      check(s_ready === 1'b1,
            "the source never became ready -- the handshake is deadlocked, which is what a lost ack looks like from this side");
      s_data  = v;
      s_valid = 1'b1;
      @(posedge clk_s); #1;
      guard = 0;
      while ((s_state == SS_IDLE) && (guard < 600)) begin @(posedge clk_s); #1; guard = guard + 1; end
      check(s_state !== SS_IDLE,
            "the source never accepted the value offered to it");
      s_valid = 1'b0;
      txq[tx_n] = v;
      tx_n = tx_n + 1;
      s_data = ~v;              // the value is now HELD, or it is not
    end
  endtask

  task drain;
    integer g;
    begin
      for (g = 0; g < 400; g = g + 1) @(posedge clk_d);
      for (g = 0; g < 400; g = g + 1) @(posedge clk_s);
    end
  endtask

  // One ratio: reset nothing, just send a run of values and confirm the
  // queues still agree. State is deliberately NOT cleared between ratios,
  // so a transfer in flight when the period changes is part of the test.
  task run_ratio(input integer new_sp, input integer new_dp, input integer count);
    integer g;
    begin
      sp = new_sp;
      dp = new_dp;
      n_ratios = n_ratios + 1;
      for (g = 0; g < count; g = g + 1)
        send(((tx_n * 8'd37) ^ (tx_n >> 3)) & 8'hff);
      drain;
      check(rx_n == tx_n,
            "the number of values received does not equal the number sent at this clock ratio");
      check(n_received === n_sent,
            "the design's own source and destination counters disagree");
    end
  endtask

  integer i, cmp_n;
  initial begin
    repeat (6) @(posedge clk_s);
    rst_s_n = 1'b1;
    repeat (6) @(posedge clk_d);
    rst_d_n = 1'b1;
    repeat (4) @(posedge clk_s);

    // ---- Phase A: the state after reset ----
    #1;
    check(s_ready === 1'b1, "the source is not ready after reset");
    check(s_req   === 1'b0, "req is asserted after reset");
    check(d_ack   === 1'b0, "ack is asserted after reset");
    check(d_valid === 1'b0, "d_valid is asserted after reset");

    // ---- Phase B: the ratio sweep. Every shape of relationship the two
    // ---- clocks can have, including one that changes mid-flight.
    run_ratio( 5,  7, 300);   // destination a little slower
    run_ratio( 7,  5, 300);   // source a little slower
    run_ratio( 5, 31, 300);   // destination MUCH slower
    run_ratio(31,  5, 300);   // source MUCH slower
    run_ratio( 3, 11, 300);   // co-prime, close
    run_ratio(11,  3, 300);
    run_ratio( 5,  6, 300);   // nearly equal, never aligned
    run_ratio( 6,  5, 300);
    run_ratio(13, 13, 300);   // equal periods, offset phase
    run_ratio( 2, 29, 200);   // extreme
    run_ratio(29,  2, 200);

    // ---- Phase C: the destination period CHANGES while transfers are in
    // ---- flight. A handshake that depended on a fixed ratio anywhere --
    // ---- a fixed-delay req, a timed ack -- comes apart here.
    for (i = 0; i < 600; i = i + 1) begin
      dp = 2 + (i % 17);
      sp = 3 + ((i * 7) % 13);
      send(((tx_n * 8'd91) ^ 8'h5a) & 8'hff);
    end
    drain;
    check(rx_n == tx_n, "a value was lost while the clock ratio was changing");

    // Everything up to here went through send(), so the two queues are
    // element-for-element comparable. Phase D drives the source directly
    // and is checked by value instead.
    cmp_n = tx_n;

    // ---- Phase D: the source offers values back to back with no gap, so
    // ---- s_valid is high through entire handshakes and the stall path is
    // ---- exercised rather than merely present.
    sp = 5; dp = 9;
    s_data  = 8'hA5;
    s_valid = 1'b1;
    for (i = 0; i < 2000; i = i + 1) @(posedge clk_s);
    s_valid = 1'b0;
    drain;
    check(n_stalls > 32'd0, "the source never had to wait, so the stall path was never exercised");

    // ---- Final: every value, in order, unchanged. ----
    check(rx_n >= cmp_n, "fewer values arrived than were sent through send()");
    // EVERY mismatch is counted, not just the first. A design that puts a
    // live bus across the boundary gets one wrong value per transfer, and
    // stopping at the first would report that as a single error -- which
    // says nothing about how pervasive it is.
    for (i = 0; i < cmp_n; i = i + 1)
      check(rxq[i] === txq[i], "a value arrived corrupted or out of order");
    // Phase D held a single value on s_data throughout, so every delivery
    // from there on must be exactly that value -- a design that forwarded
    // a live bus would deliver the inverted copy the driver leaves behind.
    for (i = cmp_n; (i < rx_n) && (i <= 20000); i = i + 1)
      check(rxq[i] === 8'hA5,
            "a back-to-back transfer delivered a value that was never offered");

    check(cmp_n > 3000, "the suite did not send enough values to mean anything");
    check(n_received === n_sent,
          "the design's own source and destination counters disagree at the end of the run");

    // The latency is DETERMINISTIC in destination cycles -- SYNC_N stages
    // plus one for the edge detector -- no matter what the clock ratio is.
    // That is the signature of a synchroniser that is really SYNC_N deep,
    // and it is the same at every ratio because it is counted in the
    // destination's own cycles.
    check(min_lat >= SYNC_N + 1,
          "the fastest observed delivery was quicker than SYNC_N synchroniser stages allow -- a stage has been bypassed");
    check(max_lat <= SYNC_N + 2,
          "a delivery took longer than the synchroniser plus edge detection can account for");
    check(n_ratios == 11, "not every clock ratio was exercised");

    // ---- EXACTLY ONE DELIVERY PER REQUEST. Not "at least one", which a
    // ---- level-triggered capture also satisfies, and not "at most one",
    // ---- which a handshake that loses transfers also satisfies.
    check(rx_n == n_req_rises,
          "the number of values delivered does not equal the number of requests raised");

    $display("REACH ratios=%0d transfers=%0d received=%0d latency=%0d..%0d dest cycles (bound >= %0d)",
             n_ratios, cmp_n, rx_n, min_lat, max_lat, SYNC_N + 1);
    $display("COUNTERS sent=%0d received=%0d requests=%0d stalls=%0d",
             n_sent, n_received, n_req_rises, n_stalls);
    $display("%0s: %0d errors in %0d checks", (errors==0)?"PASS":"FAIL", errors, checks);
    $finish;
  end
endmodule

10.2 SystemVerilog testbench

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// Testbench for usb_cdc_handshake (SystemVerilog).
//
// TWO CLOCKS, AND NO RELATIONSHIP BETWEEN THEM
//
// The whole point of the block is that it works for ANY ratio, so the suite
// sweeps the ratio rather than picking one: source faster, destination
// faster, nearly equal, wildly unequal, and a phase where the destination
// period CHANGES while transfers are in flight.
//
// WHAT IS CHECKED
//
//   1. Every value sent arrives EXACTLY ONCE, IN ORDER, UNCHANGED, at
//      every ratio. Compared element by element against an independent
//      queue, not by counting.
//
//   2. The bus that crosses NEVER MOVES while req is high. That is the
//      property that makes an unsynchronised data path safe, and the
//      source deliberately drives a different value onto s_data the
//      instant the transfer is taken, so a design that passed s_data
//      through combinationally would be caught immediately.
//
//   3. d_valid is exactly ONE destination cycle per transfer. A level-
//      triggered capture delivers the same value dozens of times at a fast
//      destination clock, and a counter alone would not notice the
//      difference between that and a fast source.
//
//   4. The minimum observed latency is at least SYNC_N+1 destination
//      cycles -- the signature of a synchroniser that really is SYNC_N
//      deep. See the note in the chapter about what this check can and
//      cannot prove.
`timescale 1ns/1ps
module tb_ch_sv;
  import usb_cdc_pkg::*;

  localparam int DW     = 8;
  localparam int SYNC_N = 2;

  int sp = 5;     // source half period
  int dp = 7;     // destination half period

  logic clk_s = 1'b0, rst_s_n = 1'b0;
  logic clk_d = 1'b0, rst_d_n = 1'b0;
  logic [DW-1:0] s_data = '0;
  logic s_valid = 1'b0;

  logic s_ready, s_req, d_valid, d_ack;
  logic [DW-1:0] s_hold, d_data;
  src_state_e s_state;
  logic [31:0] n_sent, n_stalls, n_received;

  usb_cdc_handshake #(.DW(DW), .SYNC_N(SYNC_N)) dut (.*);

  // Two free-running clocks with a deliberate phase offset, so no edge in
  // one domain ever coincides exactly with an edge in the other -- which
  // would be the one alignment a real chip is guaranteed NOT to have.
  initial forever #sp clk_s = ~clk_s;
  initial begin #3; forever #dp clk_d = ~clk_d; end

  int errors = 0, checks = 0;
  task automatic check(input logic cond, input string msg);
    checks++;
    if (!cond) begin
      errors++;
      if (errors <= 25)
        $display("FAIL @%0t: %0s | st=%0d req=%b ack=%b hold=%h d=%h",
                 $time, msg, s_state, s_req, d_ack, s_hold, d_data);
    end
  endtask

  // ---- the two independent queues ----
  logic [DW-1:0] txq [20001];
  logic [DW-1:0] rxq [20001];
  int tx_n = 0, rx_n = 0;
  int min_lat = 1000000, max_lat = 0;
  int n_ratios = 0;

  // ---- destination-side observation ----
  int d_age = 0;
  logic d_valid_d = 1'b0;

  always @(posedge clk_d or negedge rst_d_n) begin
    if (!rst_d_n) d_age <= 0;
    else if (!s_req) d_age <= 0;
    else d_age <= d_age + 1;
  end

  always @(posedge clk_d) begin
    if (rst_d_n) begin
      // ---- d_valid is ONE cycle. Two in a row is a level-triggered
      // ---- capture delivering the same value repeatedly.
      check(!(d_valid && d_valid_d),
            "d_valid was asserted for two consecutive destination cycles -- the capture is level-triggered, not edge-triggered");
      d_valid_d <= d_valid;

      if (d_valid) begin
        // The index is bounded. A level-triggered capture delivers orders
        // of magnitude more values than a correct design does, and a queue
        // that runs off its end turns a measured kill into a crash.
        if (rx_n <= 20000) rxq[rx_n] = d_data;
        rx_n++;
        if (d_age < min_lat) min_lat = d_age;
        if (d_age > max_lat) max_lat = d_age;
        // ---- the synchroniser really is SYNC_N deep ----
        check(d_age >= SYNC_N + 1,
              "a value arrived sooner than SYNC_N synchroniser stages can possibly allow -- a stage has been bypassed");
        // Note what is NOT checked here: that req is still high. The
        // monitor observes d_valid one destination cycle after the capture,
        // and a fast source can already have seen ack and completed phases
        // 3 to 5 by then. "One delivery per request" is a statement about
        // totals, so it is checked as one -- see n_req_rises below.
      end
    end
  end

  // ---- source-side observation: the bus that crosses must NOT MOVE ----
  logic [DW-1:0] hold_prev = '0;
  logic          req_prev  = 1'b0;
  int            n_req_rises = 0;
  always @(posedge clk_s) begin
    if (rst_s_n) begin
      // From the SECOND cycle of req onward. The edge that raises req is
      // the same edge that loads the hold register, so comparing across it
      // would flag the load itself.
      if (s_req && req_prev)
        check(s_hold === hold_prev,
              "the bus that crosses the boundary changed while req was asserted -- the destination may be sampling it mid-change, which is the bit-synchroniser bug in its purest form");
      // ---- and the source must not declare itself ready mid-transfer ----
      check(!(s_ready && s_req),
            "s_ready was asserted while a transfer was still in flight -- the next value can overwrite the one being captured");
      check(!(s_ready && (s_state != SS_IDLE)),
            "s_ready disagrees with the source state");
      if (s_req && !req_prev) n_req_rises++;
    end
    hold_prev <= s_hold;
    req_prev  <= s_req;
  end

  // Hand one value to the source, wait until the handshake has taken it,
  // and then IMMEDIATELY drive something else onto s_data -- so that a
  // design which forwarded s_data combinationally is caught at once.
  int guard;
  task automatic send(input logic [DW-1:0] v);
    @(posedge clk_s); #1;
    // Every wait is BOUNDED. A handshake that has lost a transfer does
    // not produce a wrong answer -- it produces no answer at all, and an
    // unbounded wait turns that into a suite that hangs instead of a
    // suite that fails. The bound is far larger than the slowest legal
    // handshake at the most extreme clock ratio in the sweep.
    guard = 0;
    while (!s_ready && (guard < 600)) begin @(posedge clk_s); #1; guard++; end
    check(s_ready === 1'b1,
          "the source never became ready -- the handshake is deadlocked, which is what a lost ack looks like from this side");
    s_data  = v;
    s_valid = 1'b1;
    @(posedge clk_s); #1;
    guard = 0;
    while ((s_state == SS_IDLE) && (guard < 600)) begin @(posedge clk_s); #1; guard++; end
    check(s_state !== SS_IDLE,
          "the source never accepted the value offered to it");
    s_valid = 1'b0;
    txq[tx_n] = v;
    tx_n++;
    s_data = ~v;                // the value is now HELD, or it is not
  endtask

  task automatic drain();
    repeat (400) @(posedge clk_d);
    repeat (400) @(posedge clk_s);
  endtask

  // One ratio: reset nothing, just send a run of values and confirm the
  // queues still agree. State is deliberately NOT cleared between ratios,
  // so a transfer in flight when the period changes is part of the test.
  task automatic run_ratio(input int new_sp, input int new_dp, input int count);
    sp = new_sp;
    dp = new_dp;
    n_ratios++;
    repeat (count)
      send(8'((tx_n * 37) ^ (tx_n >> 3)));
    drain();
      check(rx_n == tx_n,
            "the number of values received does not equal the number sent at this clock ratio");
    check(n_received === n_sent,
          "the design's own source and destination counters disagree");
  endtask

  int i, cmp_n;
  initial begin
    repeat (6) @(posedge clk_s);
    rst_s_n = 1'b1;
    repeat (6) @(posedge clk_d);
    rst_d_n = 1'b1;
    repeat (4) @(posedge clk_s);

    // ---- Phase A: the state after reset ----
    #1;
    check(s_ready === 1'b1, "the source is not ready after reset");
    check(s_req   === 1'b0, "req is asserted after reset");
    check(d_ack   === 1'b0, "ack is asserted after reset");
    check(d_valid === 1'b0, "d_valid is asserted after reset");

    // ---- Phase B: the ratio sweep. Every shape of relationship the two
    // ---- clocks can have, including one that changes mid-flight.
    run_ratio( 5,  7, 300);   // destination a little slower
    run_ratio( 7,  5, 300);   // source a little slower
    run_ratio( 5, 31, 300);   // destination MUCH slower
    run_ratio(31,  5, 300);   // source MUCH slower
    run_ratio( 3, 11, 300);   // co-prime, close
    run_ratio(11,  3, 300);
    run_ratio( 5,  6, 300);   // nearly equal, never aligned
    run_ratio( 6,  5, 300);
    run_ratio(13, 13, 300);   // equal periods, offset phase
    run_ratio( 2, 29, 200);   // extreme
    run_ratio(29,  2, 200);

    // ---- Phase C: the destination period CHANGES while transfers are in
    // ---- flight. A handshake that depended on a fixed ratio anywhere --
    // ---- a fixed-delay req, a timed ack -- comes apart here.
    for (i = 0; i < 600; i++) begin
      dp = 2 + (i % 17);
      sp = 3 + ((i * 7) % 13);
      send(8'((tx_n * 91) ^ 8'h5a));
    end
    drain();
    check(rx_n == tx_n, "a value was lost while the clock ratio was changing");

    // Everything up to here went through send(), so the two queues are
    // element-for-element comparable. Phase D drives the source directly
    // and is checked by value instead.
    cmp_n = tx_n;

    // ---- Phase D: the source offers values back to back with no gap, so
    // ---- s_valid is high through entire handshakes and the stall path is
    // ---- exercised rather than merely present.
    sp = 5; dp = 9;
    s_data  = 8'hA5;
    s_valid = 1'b1;
    repeat (2000) @(posedge clk_s);
    s_valid = 1'b0;
    drain();
    check(n_stalls > 32'd0, "the source never had to wait, so the stall path was never exercised");

    // ---- Final: every value, in order, unchanged. ----
    check(rx_n >= cmp_n, "fewer values arrived than were sent through send()");
    // EVERY mismatch is counted, not just the first. A design that puts a
    // live bus across the boundary gets one wrong value per transfer, and
    // stopping at the first would report that as a single error -- which
    // says nothing about how pervasive it is.
    for (i = 0; i < cmp_n; i++)
      check(rxq[i] === txq[i], "a value arrived corrupted or out of order");
    // Phase D held a single value on s_data throughout, so every delivery
    // from there on must be exactly that value -- a design that forwarded
    // a live bus would deliver the inverted copy the driver leaves behind.
    for (i = cmp_n; (i < rx_n) && (i <= 20000); i++)
      check(rxq[i] === 8'hA5,
            "a back-to-back transfer delivered a value that was never offered");

    check(cmp_n > 3000, "the suite did not send enough values to mean anything");
    check(n_received === n_sent,
          "the design's own source and destination counters disagree at the end of the run");

    // The latency is DETERMINISTIC in destination cycles -- SYNC_N stages
    // plus one for the edge detector -- no matter what the clock ratio is.
    // That is the signature of a synchroniser that is really SYNC_N deep,
    // and it is the same at every ratio because it is counted in the
    // destination's own cycles.
    check(min_lat >= SYNC_N + 1,
          "the fastest observed delivery was quicker than SYNC_N synchroniser stages allow -- a stage has been bypassed");
    check(max_lat <= SYNC_N + 2,
          "a delivery took longer than the synchroniser plus edge detection can account for");
    check(n_ratios == 11, "not every clock ratio was exercised");

    // ---- EXACTLY ONE DELIVERY PER REQUEST. Not "at least one", which a
    // ---- level-triggered capture also satisfies, and not "at most one",
    // ---- which a handshake that loses transfers also satisfies.
    check(rx_n == n_req_rises,
          "the number of values delivered does not equal the number of requests raised");

    $display("REACH ratios=%0d transfers=%0d received=%0d latency=%0d..%0d dest cycles (bound >= %0d)",
             n_ratios, cmp_n, rx_n, min_lat, max_lat, SYNC_N + 1);
    $display("COUNTERS sent=%0d received=%0d requests=%0d stalls=%0d",
             n_sent, n_received, n_req_rises, n_stalls);
    $display("%0s: %0d errors in %0d checks", (errors==0)?"PASS":"FAIL", errors, checks);
    $finish;
  end
endmodule

10.3 VHDL testbench

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
-- Testbench for usb_cdc_handshake (VHDL-2008).
--
-- TWO CLOCKS, AND NO RELATIONSHIP BETWEEN THEM
--
-- The whole point of the block is that it works for ANY ratio, so the suite
-- sweeps the ratio rather than picking one: source faster, destination
-- faster, nearly equal, wildly unequal, and a phase where the destination
-- period CHANGES while transfers are in flight.
--
-- WHAT IS CHECKED
--
--   1. Every value sent arrives EXACTLY ONCE, IN ORDER, UNCHANGED, at
--      every ratio. Compared element by element against an independent
--      queue, not by counting.
--
--   2. The bus that crosses NEVER MOVES while req is high. That is the
--      property that makes an unsynchronised data path safe, and the
--      source deliberately drives a different value onto s_data the
--      instant the transfer is taken, so a design that passed s_data
--      through combinationally would be caught immediately.
--
--   3. d_valid is exactly ONE destination cycle per transfer. A level-
--      triggered capture delivers the same value dozens of times at a fast
--      destination clock, and a counter alone would not notice the
--      difference between that and a fast source.
--
--   4. The minimum observed latency is at least SYNC_N+1 destination
--      cycles -- the signature of a synchroniser that really is SYNC_N
--      deep. See the note in the chapter about what this check can and
--      cannot prove.
--
-- The checks are spread over three processes because they live in two
-- different clock domains, so each process keeps its own error and check
-- tallies and the totals are summed at the end. There is no shared mutable
-- state between them, which is the same discipline the design itself
-- follows.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
use std.textio.all;
use work.usb_cdc_pkg.all;

entity tb_ch_vhdl is
end entity tb_ch_vhdl;

architecture sim of tb_ch_vhdl is

  constant DW     : integer := 8;
  constant SYNC_N : integer := 2;

  type byte_arr is array (0 to 20000) of std_logic_vector(DW-1 downto 0);

  signal sp : time := 5 ns;   -- source half period
  signal dp : time := 7 ns;   -- destination half period

  signal clk_s   : std_logic := '0';
  signal rst_s_n : std_logic := '0';
  signal clk_d   : std_logic := '0';
  signal rst_d_n : std_logic := '0';
  signal done    : boolean   := false;

  signal s_data  : std_logic_vector(DW-1 downto 0) := (others => '0');
  signal s_valid : std_logic := '0';

  signal s_ready, s_req, d_valid, d_ack : std_logic;
  signal s_hold, d_data : std_logic_vector(DW-1 downto 0);
  signal s_state : std_logic_vector(1 downto 0);
  signal n_sent, n_stalls, n_received : std_logic_vector(31 downto 0);

  -- destination-side observation, owned by the dmon process
  signal rxq     : byte_arr := (others => (others => '0'));
  signal rx_n    : integer := 0;
  signal d_age   : integer := 0;
  signal min_lat : integer := 1000000;
  signal max_lat : integer := 0;
  signal e_dst, c_dst : integer := 0;

  -- source-side observation, owned by the smon process
  signal n_req_rises : integer := 0;
  signal e_src, c_src : integer := 0;

  -- the stimulus process's own tallies
  signal e_top, c_top : integer := 0;

begin

  dut : entity work.usb_cdc_handshake
    generic map (DW => DW, SYNC_N => SYNC_N)
    port map (
      clk_s => clk_s, rst_s_n => rst_s_n,
      s_data => s_data, s_valid => s_valid, s_ready => s_ready,
      s_req => s_req, s_hold => s_hold, s_state => s_state,
      n_sent => n_sent, n_stalls => n_stalls,
      clk_d => clk_d, rst_d_n => rst_d_n,
      d_data => d_data, d_valid => d_valid, d_ack => d_ack,
      n_received => n_received
    );

  -- Two free-running clocks with a deliberate phase offset, so no edge in
  -- one domain ever coincides exactly with an edge in the other -- which
  -- would be the one alignment a real chip is guaranteed NOT to have.
  clk_s_gen : process
  begin
    if done then wait; end if;
    wait for sp;
    clk_s <= not clk_s;
  end process;

  clk_d_gen : process
    variable started : boolean := false;
  begin
    if not started then
      wait for 3 ns;
      started := true;
    end if;
    if done then wait; end if;
    wait for dp;
    clk_d <= not clk_d;
  end process;

  -- ==================================================================
  -- Destination-domain monitor
  -- ==================================================================
  dmon : process (clk_d, rst_d_n)
    variable e, c : integer := 0;
    variable dv_d : std_logic := '0';
    variable n    : integer := 0;
    variable q    : byte_arr;

    procedure chk (cond : boolean; msg : string) is
    begin
      c := c + 1;
      if not cond then
        e := e + 1;
        if e <= 25 then
          report "FAIL(dst): " & msg severity note;
        end if;
      end if;
    end procedure;
  begin
    if rst_d_n = '0' then
      d_age <= 0;
      dv_d  := '0';
    elsif rising_edge(clk_d) then
      if s_req = '0' then d_age <= 0; else d_age <= d_age + 1; end if;

      -- ---- d_valid is ONE cycle. Two in a row is a level-triggered
      -- ---- capture delivering the same value repeatedly.
      chk(not (d_valid = '1' and dv_d = '1'),
          "d_valid was asserted for two consecutive destination cycles -- the capture is level-triggered, not edge-triggered");
      dv_d := d_valid;

      if d_valid = '1' then
        -- The index is bounded. A level-triggered capture delivers orders
        -- of magnitude more values than a correct design does, and a queue
        -- that runs off its end turns a measured kill into a crash.
        if n <= 20000 then
          q(n) := d_data;
          rxq <= q;
        end if;
        n := n + 1;
        rx_n <= n;
        if d_age < min_lat then min_lat <= d_age; end if;
        if d_age > max_lat then max_lat <= d_age; end if;
        -- ---- the synchroniser really is SYNC_N deep ----
        chk(d_age >= SYNC_N + 1,
            "a value arrived sooner than SYNC_N synchroniser stages can possibly allow -- a stage has been bypassed");
        -- Note what is NOT checked here: that req is still high. The
        -- monitor observes d_valid one destination cycle after the capture,
        -- and a fast source can already have seen ack and completed phases
        -- 3 to 5 by then. "One delivery per request" is a statement about
        -- totals, so it is checked as one -- see n_req_rises below.
      end if;

      e_dst <= e;
      c_dst <= c;
    end if;
  end process;

  -- ==================================================================
  -- Source-domain monitor: the bus that crosses must NOT MOVE
  -- ==================================================================
  smon : process (clk_s, rst_s_n)
    variable e, c : integer := 0;
    variable hold_prev : std_logic_vector(DW-1 downto 0) := (others => '0');
    variable req_prev  : std_logic := '0';
    variable rises     : integer := 0;

    procedure chk (cond : boolean; msg : string) is
    begin
      c := c + 1;
      if not cond then
        e := e + 1;
        if e <= 25 then
          report "FAIL(src): " & msg severity note;
        end if;
      end if;
    end procedure;
  begin
    if rst_s_n = '0' then
      hold_prev := (others => '0');
      req_prev  := '0';
    elsif rising_edge(clk_s) then
      -- From the SECOND cycle of req onward. The edge that raises req is
      -- the same edge that loads the hold register, so comparing across it
      -- would flag the load itself.
      if s_req = '1' and req_prev = '1' then
        chk(s_hold = hold_prev,
            "the bus that crosses the boundary changed while req was asserted -- the destination may be sampling it mid-change, which is the bit-synchroniser bug in its purest form");
      end if;
      -- ---- and the source must not declare itself ready mid-transfer ----
      chk(not (s_ready = '1' and s_req = '1'),
          "s_ready was asserted while a transfer was still in flight -- the next value can overwrite the one being captured");
      chk((s_ready = '1') = (s_state = ss_code(SS_IDLE)),
          "s_ready disagrees with the source state");

      if s_req = '1' and req_prev = '0' then
        rises := rises + 1;
        n_req_rises <= rises;
      end if;

      hold_prev := s_hold;
      req_prev  := s_req;
      e_src <= e;
      c_src <= c;
    end if;
  end process;

  -- ==================================================================
  -- Stimulus
  -- ==================================================================
  stim : process
    variable e, c : integer := 0;
    variable txq  : byte_arr := (others => (others => '0'));
    variable tx_n : integer := 0;
    variable n_ratios : integer := 0;
    variable cmp_n : integer := 0;
    variable ln : line;

    procedure chk (cond : boolean; msg : string) is
    begin
      c := c + 1;
      if not cond then
        e := e + 1;
        if e <= 25 then
          report "FAIL(top): " & msg severity note;
        end if;
      end if;
    end procedure;

    -- VHDL has no xor on INTEGER, so the pattern is built in unsigned --
    -- which is the same arithmetic the other two suites do implicitly and
    -- produces the identical byte sequence.
    function pat (n, mul, xr : integer) return std_logic_vector is
      variable a, b : unsigned(DW-1 downto 0);
    begin
      a := to_unsigned((n * mul) mod 256, DW);
      b := to_unsigned(xr mod 256, DW);
      return std_logic_vector(a xor b);
    end function;

    -- Hand one value to the source, wait until the handshake has taken it,
    -- and then IMMEDIATELY drive something else onto s_data -- so that a
    -- design which forwarded s_data combinationally is caught at once.
    procedure send (v : std_logic_vector(DW-1 downto 0)) is
      variable guard : integer := 0;
    begin
      wait until rising_edge(clk_s);
      wait for 1 ns;
      -- Every wait is BOUNDED. A handshake that has lost a transfer does
      -- not produce a wrong answer -- it produces no answer at all, and an
      -- unbounded wait turns that into a suite that hangs instead of a
      -- suite that fails. The bound is far larger than the slowest legal
      -- handshake at the most extreme clock ratio in the sweep.
      guard := 0;
      while s_ready = '0' and guard < 600 loop
        wait until rising_edge(clk_s);
        wait for 1 ns;
        guard := guard + 1;
      end loop;
      chk(s_ready = '1',
          "the source never became ready -- the handshake is deadlocked, which is what a lost ack looks like from this side");
      s_data  <= v;
      s_valid <= '1';
      wait until rising_edge(clk_s);
      wait for 1 ns;
      guard := 0;
      while s_state = ss_code(SS_IDLE) and guard < 600 loop
        wait until rising_edge(clk_s);
        wait for 1 ns;
        guard := guard + 1;
      end loop;
      chk(s_state /= ss_code(SS_IDLE),
          "the source never accepted the value offered to it");
      s_valid <= '0';
      txq(tx_n) := v;
      tx_n := tx_n + 1;
      s_data <= not v;          -- the value is now HELD, or it is not
    end procedure;

    procedure drain is
    begin
      for g in 1 to 400 loop wait until rising_edge(clk_d); end loop;
      for g in 1 to 400 loop wait until rising_edge(clk_s); end loop;
    end procedure;

    -- One ratio: reset nothing, just send a run of values and confirm the
    -- queues still agree. State is deliberately NOT cleared between ratios,
    -- so a transfer in flight when the period changes is part of the test.
    procedure run_ratio (new_sp, new_dp : time; count : integer) is
    begin
      sp <= new_sp;
      dp <= new_dp;
      n_ratios := n_ratios + 1;
      for g in 1 to count loop
        send(pat(tx_n, 37, tx_n / 8));
      end loop;
      drain;
      chk(rx_n = tx_n,
          "the number of values received does not equal the number sent at this clock ratio");
      chk(n_received = n_sent,
          "the design's own source and destination counters disagree");
    end procedure;

  begin
    for g in 1 to 6 loop wait until rising_edge(clk_s); end loop;
    rst_s_n <= '1';
    for g in 1 to 6 loop wait until rising_edge(clk_d); end loop;
    rst_d_n <= '1';
    for g in 1 to 4 loop wait until rising_edge(clk_s); end loop;
    wait for 1 ns;

    -- ---- Phase A: the state after reset ----
    chk(s_ready = '1', "the source is not ready after reset");
    chk(s_req   = '0', "req is asserted after reset");
    chk(d_ack   = '0', "ack is asserted after reset");
    chk(d_valid = '0', "d_valid is asserted after reset");

    -- ---- Phase B: the ratio sweep. Every shape of relationship the two
    -- ---- clocks can have, including one that changes mid-flight.
    run_ratio( 5 ns,  7 ns, 300);   -- destination a little slower
    run_ratio( 7 ns,  5 ns, 300);   -- source a little slower
    run_ratio( 5 ns, 31 ns, 300);   -- destination MUCH slower
    run_ratio(31 ns,  5 ns, 300);   -- source MUCH slower
    run_ratio( 3 ns, 11 ns, 300);   -- co-prime, close
    run_ratio(11 ns,  3 ns, 300);
    run_ratio( 5 ns,  6 ns, 300);   -- nearly equal, never aligned
    run_ratio( 6 ns,  5 ns, 300);
    run_ratio(13 ns, 13 ns, 300);   -- equal periods, offset phase
    run_ratio( 2 ns, 29 ns, 200);   -- extreme
    run_ratio(29 ns,  2 ns, 200);

    -- ---- Phase C: the destination period CHANGES while transfers are in
    -- ---- flight. A handshake that depended on a fixed ratio anywhere --
    -- ---- a fixed-delay req, a timed ack -- comes apart here.
    for i in 0 to 599 loop
      dp <= (2 + (i mod 17)) * 1 ns;
      sp <= (3 + ((i * 7) mod 13)) * 1 ns;
      send(pat(tx_n, 91, 16#5a#));
    end loop;
    drain;
    chk(rx_n = tx_n, "a value was lost while the clock ratio was changing");

    -- Everything up to here went through send(), so the two queues are
    -- element-for-element comparable. Phase D drives the source directly
    -- and is checked by value instead.
    cmp_n := tx_n;

    -- ---- Phase D: the source offers values back to back with no gap, so
    -- ---- s_valid is high through entire handshakes and the stall path is
    -- ---- exercised rather than merely present.
    sp <= 5 ns; dp <= 9 ns;
    s_data  <= x"A5";
    s_valid <= '1';
    for g in 1 to 2000 loop wait until rising_edge(clk_s); end loop;
    s_valid <= '0';
    drain;
    chk(unsigned(n_stalls) > 0,
        "the source never had to wait, so the stall path was never exercised");

    -- ---- Final: every value, in order, unchanged. ----
    chk(rx_n >= cmp_n, "fewer values arrived than were sent through send()");
    -- EVERY mismatch is counted, not just the first. A design that puts a
    -- live bus across the boundary gets one wrong value per transfer, and
    -- stopping at the first would report that as a single error -- which
    -- says nothing about how pervasive it is.
    for i in 0 to cmp_n - 1 loop
      chk(rxq(i) = txq(i), "a value arrived corrupted or out of order");
    end loop;
    -- Phase D held a single value on s_data throughout, so every delivery
    -- from there on must be exactly that value -- a design that forwarded
    -- a live bus would deliver the inverted copy the driver leaves behind.
    for i in cmp_n to minimum(rx_n - 1, 20000) loop
      chk(rxq(i) = x"A5",
          "a back-to-back transfer delivered a value that was never offered");
    end loop;

    chk(cmp_n > 3000, "the suite did not send enough values to mean anything");
    chk(n_received = n_sent,
        "the design's own source and destination counters disagree at the end of the run");

    -- The latency is DETERMINISTIC in destination cycles -- SYNC_N stages
    -- plus one for the edge detector -- no matter what the clock ratio is.
    -- That is the signature of a synchroniser that is really SYNC_N deep,
    -- and it is the same at every ratio because it is counted in the
    -- destination's own cycles.
    chk(min_lat >= SYNC_N + 1,
        "the fastest observed delivery was quicker than SYNC_N synchroniser stages allow -- a stage has been bypassed");
    chk(max_lat <= SYNC_N + 2,
        "a delivery took longer than the synchroniser plus edge detection can account for");
    chk(n_ratios = 11, "not every clock ratio was exercised");

    -- ---- EXACTLY ONE DELIVERY PER REQUEST. Not "at least one", which a
    -- ---- level-triggered capture also satisfies, and not "at most one",
    -- ---- which a handshake that loses transfers also satisfies.
    chk(rx_n = n_req_rises,
        "the number of values delivered does not equal the number of requests raised");

    e_top <= e;
    c_top <= c;
    wait for 1 ns;

    write(ln, string'("REACH ratios=") & integer'image(n_ratios) &
              " transfers=" & integer'image(cmp_n) &
              " received=" & integer'image(rx_n) &
              " latency=" & integer'image(min_lat) & ".." & integer'image(max_lat) &
              " dest cycles (bound >= " & integer'image(SYNC_N + 1) & ")");
    writeline(output, ln);
    write(ln, string'("COUNTERS sent=") & integer'image(to_integer(unsigned(n_sent))) &
              " received=" & integer'image(to_integer(unsigned(n_received))) &
              " requests=" & integer'image(n_req_rises) &
              " stalls=" & integer'image(to_integer(unsigned(n_stalls))));
    writeline(output, ln);
    if (e + e_src + e_dst) = 0 then
      write(ln, string'("PASS: 0 errors in ") &
                integer'image(c + c_src + c_dst) & " checks");
    else
      write(ln, string'("FAIL: ") & integer'image(e + e_src + e_dst) &
                " errors in " & integer'image(c + c_src + c_dst) & " checks");
    end if;
    writeline(output, ln);

    done <= true;
    wait;
  end process;

end architecture sim;

11. Exhaustive Verification

MeasureVerilogSystemVerilogVHDL
Clock ratios swept111111
Values sent through the checker370037003700
Total transfers (incl. back-to-back phase)382838283827
Checks executed302907302907302873
design's n_sent382838283827
design's n_received382838283827
requests raised382838283827
source stall cycles187218721873
delivery latency (destination cycles)3 .. 33 .. 33 .. 3
values lost, duplicated or corrupted000
ResultPASSPASSPASS

The latency row is the interesting one. Three destination cycles, minimum and maximum, at every ratio from 2:29 to 29:2 and through a phase where both periods changed on every transfer.

That is not a coincidence, and it is the signature of a correct synchroniser: two flops plus one for edge detection, counted in the destination's own cycles, which is the only frame in which the number is meaningful at all. Measured in source cycles the same transfer takes anywhere from 2 to 90.

n_sent, n_received and the count of requests raised agree exactly — three counters in two different clock domains, which is the property the whole block exists to provide.

12. Mutation Testing

#MutationVerilogSysVerVHDL
C3the capture follows req's LEVEL, not its rising edge555815556955563
C7s_ready is asserted one phase early107691076810770
C6ack is released on a timer, not when req falls682668156815
C4req is dropped on a timer, not when ack arrives582058085849
C5three phases instead of four519851865515
C1the synchroniser is tapped one stage early392439243924
C2the data is not held — a live bus crosses the boundary370037003700
—unmutated baseline000

All seven die in all three languages, all counts distinct, and the columns agree to within 6% — the suite is fully deterministic, so where they differ it is only because the VHDL scheduler resolves one back-to-back transfer differently in phase D.

C3 is by far the largest. A level-triggered capture re-delivers the same value on every destination cycle for as long as req is high, so one transfer becomes dozens. It is caught three separate ways — consecutive d_valid, the delivery count, and the end-to-end queue — which is what you want for the failure that produces the most wrong data.

C2 scores exactly 3700: one wrong value per transfer, every transfer, no exceptions. That exactness is worth pausing on. It is not a flaky, load-dependent, ratio-dependent bug. Putting a live bus across a clock boundary is wrong every single time, and the only reason it ever appears to work in a lab is that the two clocks happened to be related that afternoon.

C1 is the one to be careful about, and it is discussed below.

13. Debugging Walkthrough: The Device That Fails Once a Week

The report. A USB device fails roughly once a week per unit in the field, and never in the lab. When it fails it enumerates with the wrong number of endpoints and has to be unplugged. No error is logged by anything.

Step 1 — reproduce it. Impossible. Thousands of hours of soak testing, zero failures. This is the first and most important datum: a failure that will not reproduce in a stable environment is usually a failure that depends on an unstable one.

Step 2 — what is unstable in the field and stable in the lab? Temperature, voltage, and — crucially — the relationship between two clock sources. In the lab both were derived from one bench oscillator. In the field the PHY's clock comes from the recovered bus clock and the controller's from a local crystal, and the phase between them drifts continuously.

Step 3 — so look for clock crossings. Run a CDC lint. It reports 41 crossings. Forty of them are single bits with two-flop synchronisers. One is a 4-bit endpoint-count value with a two-flop synchroniser on each bit.

Step 4 — is that actually wrong? The value is written once at configuration time and never changes afterwards. The reasoning in the original review was "it is static, so there is nothing to synchronise."

Step 5 — but it changes once. From 0000 to 0100 at SET_CONFIGURATION. Two bits... no, one bit. From 0011 to 0100 on a reconfiguration: three bits at once. If the destination samples during that transition it can see 0111 — seven endpoints where there are four.

Step 6 — why once a week. The window is one destination clock period per reconfiguration, and the phase has to land inside it. That is a probability, not a condition, and a probability is exactly what produces "once a week per unit, never in the lab".

Step 7 — the fix. A four-phase handshake. The value is static, so the handshake costs nothing that matters.

14. UVM and Assertions

14.1 The sequences

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
class usb_cdc_item extends uvm_sequence_item;
  `uvm_object_utils(usb_cdc_item)

  rand bit [7:0] data;
  rand int unsigned gap;      // source cycles of idle before offering it

  constraint c_gap { gap dist {0 := 5, [1:3] := 3, [4:20] := 1}; }

  function new(string name = "usb_cdc_item"); super.new(name); endfunction
endclass

// THE sequence for this chapter, and it is not a data sequence at all --
// it is a CLOCK sequence. The block's claim is that it works at any ratio,
// so the ratio is the thing to randomise. A suite that fixes the ratio and
// randomises only the payload is testing one point of the space it cares
// about.
class ratio_sweep_seq extends uvm_sequence #(usb_cdc_item);
  `uvm_object_utils(ratio_sweep_seq)

  // Set by the test; the driver applies them to the clock generator.
  rand int unsigned src_period;
  rand int unsigned dst_period;
  constraint c_periods {
    src_period inside {[2:31]};
    dst_period inside {[2:31]};
  }

  function new(string name = "ratio_sweep_seq"); super.new(name); endfunction

  task body();
    repeat (40) begin
      if (!this.randomize())
        `uvm_error("RAND", "period randomize failed")
      uvm_config_db #(int unsigned)::set(null, "*", "src_period", src_period);
      uvm_config_db #(int unsigned)::set(null, "*", "dst_period", dst_period);

      repeat (100) begin
        usb_cdc_item it = usb_cdc_item::type_id::create("it");
        start_item(it);
        if (!it.randomize())
          `uvm_error("RAND", "item randomize failed")
        finish_item(it);
      end
    end
  endtask
endclass

// Back to back with no gap at all, so s_valid is high through entire
// handshakes and the source spends most of its time stalled. This is the
// sequence that exercises the cost of the handshake rather than its
// correctness, and it is the one that catches a source declaring itself
// ready a phase early.
class back_to_back_seq extends uvm_sequence #(usb_cdc_item);
  `uvm_object_utils(back_to_back_seq)
  function new(string name = "back_to_back_seq"); super.new(name); endfunction

  task body();
    repeat (2000) begin
      usb_cdc_item it = usb_cdc_item::type_id::create("it");
      start_item(it);
      if (!it.randomize() with { gap == 0; })
        `uvm_error("RAND", "back-to-back randomize failed")
      finish_item(it);
    end
  endtask
endclass

14.2 The scoreboard: two domains, two monitors, one queue

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// The source and destination monitors run in DIFFERENT clock domains and
// write into the same scoreboard. That is the one place in the whole
// environment where two domains meet, and it is deliberately the only one.
class usb_cdc_scoreboard extends uvm_scoreboard;
  `uvm_component_utils(usb_cdc_scoreboard)

  uvm_analysis_imp_src #(usb_cdc_src_item, usb_cdc_scoreboard) src_ap;
  uvm_analysis_imp_dst #(usb_cdc_dst_item, usb_cdc_scoreboard) dst_ap;

  localparam int SYNC_N = 2;

  byte unsigned expected[$];

  int unsigned n_sent, n_received;
  int unsigned min_lat = 32'hFFFF_FFFF, max_lat;
  bit [7:0]    hold_at_req;
  bit          req_active;

  function new(string name, uvm_component parent);
    super.new(name, parent);
    src_ap = new("src_ap", this);
    dst_ap = new("dst_ap", this);
  endfunction

  // ---- source side ----
  function void write_src(usb_cdc_src_item t);
    // THE property that makes an unsynchronised data path safe.
    if (t.s_req && req_active && (t.s_hold !== hold_at_req))
      `uvm_error("UNSTABLE",
        $sformatf("the bus that crosses the boundary changed from 0x%02h to 0x%02h while req was asserted -- the destination may be sampling it mid-change",
                  hold_at_req, t.s_hold))

    if (t.s_req && !req_active) begin
      req_active  = 1;
      hold_at_req = t.s_hold;
      expected.push_back(t.s_hold);
      n_sent++;
    end
    if (!t.s_req) req_active = 0;

    // A source that declares itself ready mid-transfer can overwrite the
    // value the destination is still reading.
    if (t.s_ready && t.s_req)
      `uvm_error("EARLY_READY",
        "s_ready was asserted while a transfer was still in flight")
  endfunction

  // ---- destination side ----
  function void write_dst(usb_cdc_dst_item t);
    if (!t.d_valid) return;

    // ---- EXACTLY ONE delivery, of EXACTLY the value that was sent. ----
    if (expected.size() == 0)
      `uvm_error("PHANTOM",
        $sformatf("0x%02h was delivered but nothing was sent -- the capture is following req's level, not its edge",
                  t.d_data))
    else begin
      byte unsigned want = expected.pop_front();
      if (t.d_data !== want)
        `uvm_error("DATA",
          $sformatf("received 0x%02h, expected 0x%02h -- a live bus crossed the boundary instead of a held one",
                    t.d_data, want))
    end
    n_received++;

    // ---- The synchroniser really is SYNC_N deep. See the warning in the
    // ---- chapter about exactly what this does and does not prove.
    if (t.latency_in_dst_cycles < SYNC_N + 1)
      `uvm_error("TOO_FAST",
        $sformatf("delivered after %0d destination cycles; %0d synchroniser stages plus edge detection cannot produce fewer than %0d",
                  t.latency_in_dst_cycles, SYNC_N, SYNC_N + 1))
    if (t.latency_in_dst_cycles < min_lat) min_lat = t.latency_in_dst_cycles;
    if (t.latency_in_dst_cycles > max_lat) max_lat = t.latency_in_dst_cycles;
  endfunction

  function void report_phase(uvm_phase phase);
    `uvm_info("SB", $sformatf("sent=%0d received=%0d latency=%0d..%0d dest cycles",
                              n_sent, n_received, min_lat, max_lat), UVM_LOW)

    // Nothing may be left in flight at the end of the run.
    if (expected.size() != 0)
      `uvm_error("LOST",
        $sformatf("%0d values were sent and never delivered", expected.size()))
    if (n_sent != n_received)
      `uvm_error("COUNT",
        $sformatf("sent %0d, received %0d", n_sent, n_received))

    // The latency must be CONSTANT in destination cycles, whatever the
    // ratio. A spread here means the handshake's timing depends on the
    // clock relationship, which is the thing it exists not to do.
    if (min_lat != max_lat)
      `uvm_error("JITTER",
        $sformatf("delivery latency varied from %0d to %0d destination cycles -- the handshake is sensitive to the clock ratio",
                  min_lat, max_lat))
  endfunction
endclass

14.3 Assertions, and which domain each one belongs to

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// Two assertion modules, not one. An assertion has a clock, and an
// assertion clocked in the wrong domain is a race dressed up as a check.
module usb_cdc_src_sva
  import usb_cdc_pkg::*;
#(
  parameter int DW = 8
) (
  input logic          clk_s,
  input logic          rst_s_n,
  input logic          s_valid,
  input logic          s_ready,
  input logic          s_req,
  input logic [DW-1:0] s_hold,
  input src_state_e    s_state
);
  default clocking cb @(posedge clk_s); endclocking
  default disable iff (!rst_s_n);

  // ---- 1. THE property. The bus does not move while req is high. ----
  //
  //      Written from the SECOND cycle of req, because the edge that
  //      raises req is the same edge that legitimately loads the hold
  //      register. See the chapter on why this cannot catch a purely
  //      combinational data path.
  property p_hold_is_stable;
    (s_req && $past(s_req)) |-> $stable(s_hold);
  endproperty
  a_hold_is_stable : assert property (p_hold_is_stable)
    else $error("the bus crossing the boundary moved while req was asserted");

  // ---- 2. The source is never ready mid-transfer. ----
  property p_not_ready_in_flight;
    s_req |-> !s_ready;
  endproperty
  a_not_ready_in_flight : assert property (p_not_ready_in_flight)
    else $error("s_ready asserted while a transfer was in flight -- the next value can overwrite the one being captured");

  // ---- 3. FOUR phases, in order. req goes up, and comes down only from
  // ----    SS_REQ; the machine then waits in SS_DROP.
  property p_req_falls_only_from_req_state;
    $fell(s_req) |-> $past(s_state == SS_REQ);
  endproperty
  a_req_falls_only_from_req_state :
    assert property (p_req_falls_only_from_req_state);

  property p_drop_precedes_idle;
    (s_state == SS_IDLE) && $past(s_state != SS_IDLE)
      |-> $past(s_state == SS_DROP);
  endproperty
  a_drop_precedes_idle : assert property (p_drop_precedes_idle)
    else $error("the source returned to IDLE without passing through SS_DROP -- this is the three-phase handshake, and it loses the next transfer");

  // ---- 4. req rises only from IDLE, and only when a value was offered. ----
  property p_req_rises_from_idle;
    $rose(s_req) |-> $past(s_state == SS_IDLE) && $past(s_valid);
  endproperty
  a_req_rises_from_idle : assert property (p_req_rises_from_idle);

  c_stalled : cover property ((s_valid && !s_ready));
  c_drop    : cover property ((s_state == SS_DROP));
endmodule

module usb_cdc_dst_sva #(
  parameter int DW = 8
) (
  input logic          clk_d,
  input logic          rst_d_n,
  input logic          s_req,      // asynchronous here, and only ever
                                   // referenced through its synchroniser
  input logic          req_seen,
  input logic          d_valid,
  input logic [DW-1:0] d_data,
  input logic          d_ack
);
  default clocking cb @(posedge clk_d); endclocking
  default disable iff (!rst_d_n);

  // ---- 5. d_valid is ONE cycle. Never two. ----
  property p_valid_is_a_pulse;
    d_valid |=> !d_valid;
  endproperty
  a_valid_is_a_pulse : assert property (p_valid_is_a_pulse)
    else $error("d_valid was asserted for two consecutive cycles -- the capture is level-triggered and is re-delivering the same value");

  // ---- 6. A delivery happens only on the synchronised RISING edge. ----
  property p_valid_needs_a_rise;
    d_valid |-> $past(req_seen) && !$past(req_seen, 2);
  endproperty
  a_valid_needs_a_rise : assert property (p_valid_needs_a_rise);

  // ---- 7. ack rises with the delivery and falls only after req does. ----
  property p_ack_rises_with_delivery;
    $rose(d_ack) |-> d_valid;
  endproperty
  a_ack_rises_with_delivery : assert property (p_ack_rises_with_delivery);

  property p_ack_falls_after_req;
    $fell(d_ack) |-> !req_seen;
  endproperty
  a_ack_falls_after_req : assert property (p_ack_falls_after_req)
    else $error("ack was released before req fell -- the source may miss it entirely and wait for ever");

  // ---- 8. The data is stable for the cycle it is delivered in. ----
  property p_data_stable_at_valid;
    d_valid |=> $stable(d_data);
  endproperty
  a_data_stable_at_valid : assert property (p_data_stable_at_valid);

  c_delivered : cover property (d_valid);
endmodule

bind usb_cdc_handshake usb_cdc_src_sva #(.DW(DW)) u_src_sva (.*);
bind usb_cdc_handshake usb_cdc_dst_sva #(.DW(DW)) u_dst_sva (.*);

15. Common Misconceptions

"Put a two-flop synchroniser on each bit." Each bit resolves independently. The destination can sample a value the source never held.

"Gray-code the bus." Gray codes work for counters, where single-bit transitions are guaranteed by construction. Arbitrary data changes arbitrarily many bits.

"The value is static, so it does not need synchronising." It changes once, and the failure window is the transition. Once a week per unit is a probability, not a condition.

"Three phases is enough — drop req on ack and go." ack is still high. The next transfer sees the previous ack, completes in zero time, and the destination never captures.

"Ack can be released on a timer." Then the source's synchroniser can miss the pulse entirely and wait for ever.

"A handshake is too slow." For a control value it costs nothing. For streaming data you wanted an asynchronous FIFO, which is this idea applied to Gray-coded pointers.

"The regression is green, so the CDC is fine." RTL simulation does not model metastability. A missing synchroniser stage is invisible to every simulation at every seed. Lint and STA are the tools; simulation is not.

16. Exercises

1. Work out how many distinct values a destination can sample when a 6-bit bus changes from 011111 to 100000 through per-bit synchronisers, and name two of them that would be actively harmful as an endpoint number.

2. Apply C5 (three phases) and trace, cycle by cycle, what happens to the second transfer. Then say why n_sent still increments.

3. C2 scores exactly 3700 — one per transfer. Explain why the source-side stability monitor does not catch it, and design a source-side check that would.

4. The delivery latency is 3 destination cycles at every ratio. Derive that number from the design, then work out the latency in source cycles at 2:29 and at 29:2.

5. C1 is killed only by its latency signature. Write down what a CDC lint would report for it, and explain why that report is stronger evidence than the mutation score.

6. Replace the handshake with an asynchronous FIFO for the same data. What changes, what stays, and why do the pointers get a Gray code when the data does not?

17. Summary

IdeaWhy it matters
A multi-bit value cannot be bit-synchronisedeach bit resolves independently; 16 possible samples
A Gray code fixes a counter, not a bussingle-bit transitions must be by construction
Synchronise one bit, not the datathe data crosses untouched, and is never sampled moving
Four phases, not threeack is still high when three-phase declares ready
The data is latched before req risesa live bus across a boundary is the original bug
The capture is on the edge, not the levela level re-delivers the same value dozens of times
Latency is constant in destination cyclesthe only frame in which the number means anything
"It is static" is the reasoning to distrustit changes once, and the window is the transition
A handshake is slowcontrol values, not streams; use an async FIFO for those
Simulation cannot see metastabilitya missing flop is invisible; lint and STA are the tools
11 clock ratios, 3828 transfers, 0 lost7 mutations, all killed in 3 languages

Tooling

StepCommand
Verilog-2005iverilog -g2005 -o ch_v.out ch_v.v ch_v_tb.v && ./ch_v.out
SystemVerilogiverilog -g2012 -o ch_sv.out ch_sv.sv ch_sv_tb.sv && ./ch_sv.out
VHDL-2008 analysenvc --std=2008 -a ch_vhdl.vhd ch_vhdl_tb.vhd
VHDL-2008 elaboratenvc --std=2008 -e tb_ch_vhdl
VHDL-2008 runnvc --std=2008 -r tb_ch_vhdl
One mutationiverilog -g2005 -DMUT_C2 -o mm ch_v_mut.v ch_v_tb.v && ./mm
And the one that mattersa CDC lint, plus STA with the crossing constrained

All three implementations pass with 0 errors: eleven clock ratios plus a phase in which both periods change on every transfer, ~3828 values delivered exactly once each in order and unchanged, and a delivery latency of exactly three destination cycles throughout.


That completes Module 23 — USB RTL Design. Five blocks: the address decode that decides whether a packet is ours, the arbiter that shares one port with bounded waiting, the FIFO whose watermark has to pay for its own backpressure latency, the packet recogniser that resynchronises after every error, and the handshake that carries a value between two clocks that have never met.

Each was written three times, verified exhaustively over a named domain, and mutated seven ways with every mutation killed in every language. What is worth carrying forward is not the five designs but the four habits underneath them: make the property observable, reach every state for real, measure the bound rather than assume it, and be explicit about what your verification cannot see.

Continue learning

Standards & specifications

Governing standard
USB-IF (Universal Serial Bus Specification)(opens USB Implementers Forum (USB-IF) in a new tab)

Defines the USB bus — its electrical signalling, connectors, packet and transaction model, device framework and the descriptors a device must expose — together with the device-class specifications layered on it. It does not define host-controller register interfaces (xHCI and EHCI are separate documents) nor any operating system's driver architecture.

This page also covers RTL structure, verification approach and debugging technique. Those are engineering practice built on the standard, not requirements the standard itself imposes.

Where this fits

Part of the USB curriculum.