USB · Module 23
USB Timing Considerations
A multi-bit value cannot be bit-synchronised — each bit resolves independently, so the destination can sample a value that was never sent. The four-phase handshake crosses one bit and lets the data through untouched.
Chapters 23.1 to 23.4 all assumed one clock. A USB device does not have one clock.
The PHY runs at the bus rate — 60 MHz for a UTMI high-speed interface, derived from the recovered bit clock. The controller, the FIFOs and the CPU do not. Every value that moves between them crosses a boundary where the two clocks have no fixed relationship at all, and the rules on the other side of that boundary are different.
1. The Two-Flop Synchroniser, and Where It Stops Working
For one bit, the standard answer is correct and well understood: two flops in the destination domain, so the first flop's metastability has a full destination cycle to resolve before anything downstream looks at it.
Put one on each bit of a bus, and the design is broken.
source changes 0111 -> 1000 (7 -> 8)
Four bits change at once. Each bit's synchroniser resolves
INDEPENDENTLY -- there is no mechanism that makes them agree.
So the destination can sample:
0111 1111 1000 0000 1011 0100 ...
ANY of the sixteen values, for one destination cycle,
before the bus settles.Those intermediate values are not noise to be filtered out. They are indices into a descriptor table, byte counts, endpoint numbers. A transient 0000 on an endpoint number is a packet delivered to endpoint 0. A transient 1111 on a byte count is a DMA of 15 bytes that should have been 8.
Why a per-bit synchroniser fails on a bus
2. The Answer: Do Not Synchronise the Data At All
Synchronise one bit — a request — and use it to say "the data has been stable for a long time and will stay stable until I hear back."
The four-phase handshake
The data path crosses the boundary with no synchroniser on it whatsoever. That looks alarming on a schematic and is the entire point: it is never sampled while it is changing, so it never needs one.
3. Why the Fourth Phase Cannot Be Dropped
The tempting simplification is three phases: the source drops req when it sees ack, and declares itself ready.
But ack is still high — the destination has not been told to drop it yet.
three-phase:
transfer N req up ... ack up ... req down ... READY
transfer N+1 req up -- and ack is STILL HIGH from N
the source sees "ack" immediately,
drops req in one cycle, completes in
zero time,
and the destination never captured
anything at all.
The value is lost. No error. No stall. Nothing.That is mutation C5 in §11, and it is the single most common way a hand-written handshake is wrong.
4. A Handshake Is Slow, and That Is the Trade
Four phases, each costing SYNC_N cycles in one domain or the other, plus edge detection. Roughly five to six cycles of the slower clock per transfer.
| Use it for | Not for |
|---|---|
| a device address, a configuration value | a byte stream |
| a status word, an interrupt flag | packet payload |
| a completion code | anything at line rate |
For streaming data the answer is an asynchronous FIFO — which is this same idea applied to pointers rather than to payload: the pointers are Gray-coded (a counter, so single-bit transitions are guaranteed) and the memory is dual-ported, so the data never crosses a boundary at all. That is a different block and the reason this one is not it.
5. What We Are Building
usb_cdc_handshake #(DW = 8, SYNC_N = 2)
source domain (clk_s) destination domain (clk_d)
--------------------- --------------------------
s_data / s_valid d_data
s_ready d_valid ONE cycle per transfer
s_req the bit that crosses d_ack the bit that crosses back
s_hold the bus that crosses
with NO synchroniser
n_sent / n_stalls n_receiveds_hold is exposed deliberately. "It never moves while req is high" is a property, and a property you cannot observe is a property you are assuming.
6. Verilog-2005 Implementation
// usb_cdc_handshake -- crossing a multi-bit value between two clocks, and
// the reason you cannot do it the obvious way.
//
// THE ASSUMPTION THAT JUST BROKE
//
// Chapters 23.1 to 23.4 all had one clock. A USB device does not. The PHY
// runs at the bus rate -- 60 MHz for a UTMI high-speed interface -- and the
// controller, the FIFOs and the CPU do not. Every value that moves between
// them crosses a boundary where the two clocks have no fixed relationship
// at all.
//
// A MULTI-BIT VALUE CANNOT BE BIT-SYNCHRONISED
//
// The two-flop synchroniser is the standard answer for a single bit, and it
// is correct: it gives the first flop's metastability time to resolve before
// anything downstream looks at it.
//
// Put one on each bit of a bus and the design is broken.
//
// source changes 0111 -> 1000 (7 -> 8, four bits at once)
//
// Each bit's synchroniser resolves INDEPENDENTLY. There is no
// mechanism that makes them agree. The destination can sample
//
// 0111 1111 1000 0000 1011 0100 ...
//
// any of the sixteen values, for one cycle, before settling.
//
// Those intermediate values are not noise to be filtered. They are indices
// into a descriptor table, byte counts, endpoint numbers. A transient 0000
// on an endpoint number is a packet delivered to endpoint 0.
//
// And a Gray code does not rescue the general case. It works for a counter,
// where only one bit changes per increment BY CONSTRUCTION. Arbitrary data
// changes arbitrarily many bits at once, and no encoding fixes that.
//
// THE ANSWER: DO NOT SYNCHRONISE THE DATA AT ALL
//
// Synchronise ONE bit -- a request -- and use it to say "the data has been
// stable for a long time and will stay stable until I hear back".
//
// 1. source latches the data, raises req
// 2. destination sees req (2 flops later), captures the data, raises ack
// 3. source sees ack (2 flops later), drops req
// 4. destination sees req drop, drops ack
// 5. source sees ack drop -- and only NOW may it change the data
//
// That is the FOUR-PHASE handshake, and every one of the four phases is
// load-bearing. The data path crosses the boundary with no synchroniser on
// it whatsoever, which looks alarming and is the entire point: it is not
// sampled while it is changing, so it never needs one.
//
// WHY PHASE 4 CANNOT BE DROPPED
//
// The tempting simplification is three phases: the source drops req on ack
// and declares itself ready. But ack is still HIGH -- it has not been told
// to drop yet -- so the next transfer starts, raises req, and immediately
// sees the PREVIOUS ack. It completes in zero time, the destination never
// captures, and the value is silently lost.
//
// A HANDSHAKE IS SLOW, AND THAT IS THE TRADE
//
// Four phases, each costing SYNC_N destination or source cycles plus edge
// detection. Roughly 5 to 6 cycles of the slower clock per transfer. For
// control values -- an address, a configuration, a status word -- that is
// free. For streaming data it is not, and the answer there is an
// asynchronous FIFO, which is this handshake's idea applied to pointers
// instead of to payload.
module usb_cdc_handshake #(
parameter integer DW = 8,
parameter integer SYNC_N = 2 // synchroniser depth (>= 2; see below)
) (
// ---- source domain ----
input wire clk_s,
input wire rst_s_n,
input wire [DW-1:0] s_data,
input wire s_valid, // the source has a value to send
output wire s_ready, // ...and this says it may send it
output wire s_req, // THE wire that crosses, and it is 1 bit
output wire [DW-1:0] s_hold, // ...and this is the bus that crosses with
// no synchroniser on it at all. Exposed
// because "it never moves while req is
// high" is a property, and a property you
// cannot observe is a property you are
// assuming.
output wire [1:0] s_state,
output wire [31:0] n_sent,
output wire [31:0] n_stalls,
// ---- destination domain ----
input wire clk_d,
input wire rst_d_n,
output wire [DW-1:0] d_data,
output wire d_valid, // ONE destination cycle per transfer
output wire d_ack, // the wire that crosses back
output wire [31:0] n_received
);
localparam [1:0] SS_IDLE = 2'd0, // ready; the data may be changed
SS_REQ = 2'd1, // req raised, waiting for ack
SS_DROP = 2'd2; // req dropped, waiting for ack to fall
// ==================================================================
// SOURCE DOMAIN
// ==================================================================
reg [1:0] ss_r;
reg [DW-1:0] hold_r; // the CDC data path. NOT synchronised. Held.
reg req_r;
reg [31:0] sent_r, stalls_r;
// The ack, brought into the source domain through SYNC_N flops. This is
// a single bit, which is the only kind of signal a synchroniser is
// allowed to be used on.
reg [SYNC_N-1:0] ack_sync_r;
wire ack_seen = ack_sync_r[SYNC_N-1];
assign s_req = req_r;
assign s_hold = hold_r;
assign s_state = ss_r;
assign s_ready = (ss_r == SS_IDLE);
assign n_sent = sent_r;
assign n_stalls = stalls_r;
always @(posedge clk_s or negedge rst_s_n) begin
if (!rst_s_n) begin
ss_r <= SS_IDLE;
hold_r <= {DW{1'b0}};
req_r <= 1'b0;
ack_sync_r <= {SYNC_N{1'b0}};
sent_r <= 32'd0;
stalls_r <= 32'd0;
end else begin
ack_sync_r <= {ack_sync_r[SYNC_N-2:0], d_ack};
case (ss_r)
SS_IDLE: begin
if (s_valid) begin
// ---- PHASE 1. Latch the data, THEN raise req. ----
//
// The latch is what makes the data path safe: from here until
// phase 5 the destination is looking at a value that does not
// move. Passing s_data through combinationally instead would
// put a live, changing bus across the boundary -- which is the
// bit-synchroniser bug with the synchronisers removed.
hold_r <= s_data;
req_r <= 1'b1;
ss_r <= SS_REQ;
end
end
SS_REQ: begin
// ---- PHASE 3. The destination has captured it. Drop req. ----
if (ack_seen) begin
req_r <= 1'b0;
ss_r <= SS_DROP;
end
end
SS_DROP: begin
// ---- PHASE 5. And WAIT for ack to fall before saying ready.
//
// Returning to IDLE here without waiting is the three-phase
// shortcut: the next req would be raised while ack is still high
// from the previous transfer, the source would see it instantly,
// and the destination would never capture anything.
if (!ack_seen) begin
ss_r <= SS_IDLE;
sent_r <= sent_r + 32'd1;
end
end
default: ss_r <= SS_IDLE;
endcase
// A source that wants to send while the handshake is busy. Counted,
// because the cost of a handshake is exactly this and it should be
// measured rather than assumed.
if (s_valid && (ss_r != SS_IDLE)) stalls_r <= stalls_r + 32'd1;
end
end
// ==================================================================
// DESTINATION DOMAIN
// ==================================================================
reg [SYNC_N-1:0] req_sync_r;
reg req_d_r; // one more flop, for edge detection
reg [DW-1:0] d_data_r;
reg d_valid_r, ack_r;
reg [31:0] recv_r;
wire req_seen = req_sync_r[SYNC_N-1];
// The RISING EDGE, not the level. A level would re-capture and re-pulse
// on every destination cycle for as long as req is high, which at a fast
// destination clock is a single transfer delivered dozens of times.
wire req_rise = req_seen && !req_d_r;
assign d_data = d_data_r;
assign d_valid = d_valid_r;
assign d_ack = ack_r;
assign n_received = recv_r;
always @(posedge clk_d or negedge rst_d_n) begin
if (!rst_d_n) begin
req_sync_r <= {SYNC_N{1'b0}};
req_d_r <= 1'b0;
d_data_r <= {DW{1'b0}};
d_valid_r <= 1'b0;
ack_r <= 1'b0;
recv_r <= 32'd0;
end else begin
req_sync_r <= {req_sync_r[SYNC_N-2:0], s_req};
req_d_r <= req_seen;
d_valid_r <= 1'b0;
if (req_rise) begin
// ---- PHASE 2. Capture the data and acknowledge. ----
//
// hold_r has been stable for at least SYNC_N destination cycles by
// now -- that is precisely what the synchroniser bought -- so this
// read is of a static value and needs no synchroniser of its own.
d_data_r <= hold_r;
d_valid_r <= 1'b1;
ack_r <= 1'b1;
recv_r <= recv_r + 32'd1;
end else if (!req_seen) begin
// ---- PHASE 4. req has fallen; release ack. ----
ack_r <= 1'b0;
end
end
end
endmodule7. SystemVerilog Implementation
// usb_cdc_handshake -- crossing a multi-bit value between two clocks, and
// the reason you cannot do it the obvious way.
//
// THE ASSUMPTION THAT JUST BROKE
//
// Chapters 23.1 to 23.4 all had one clock. A USB device does not. The PHY
// runs at the bus rate -- 60 MHz for a UTMI high-speed interface -- and the
// controller, the FIFOs and the CPU do not. Every value that moves between
// them crosses a boundary where the two clocks have no fixed relationship
// at all.
//
// A MULTI-BIT VALUE CANNOT BE BIT-SYNCHRONISED
//
// The two-flop synchroniser is the standard answer for a single bit, and it
// is correct: it gives the first flop's metastability time to resolve before
// anything downstream looks at it.
//
// Put one on each bit of a bus and the design is broken.
//
// source changes 0111 -> 1000 (7 -> 8, four bits at once)
//
// Each bit's synchroniser resolves INDEPENDENTLY. There is no
// mechanism that makes them agree. The destination can sample
//
// 0111 1111 1000 0000 1011 0100 ...
//
// any of the sixteen values, for one cycle, before settling.
//
// Those intermediate values are not noise to be filtered. They are indices
// into a descriptor table, byte counts, endpoint numbers. A transient 0000
// on an endpoint number is a packet delivered to endpoint 0.
//
// And a Gray code does not rescue the general case. It works for a counter,
// where only one bit changes per increment BY CONSTRUCTION. Arbitrary data
// changes arbitrarily many bits at once, and no encoding fixes that.
//
// THE ANSWER: DO NOT SYNCHRONISE THE DATA AT ALL
//
// Synchronise ONE bit -- a request -- and use it to say "the data has been
// stable for a long time and will stay stable until I hear back".
//
// 1. source latches the data, raises req
// 2. destination sees req (2 flops later), captures the data, raises ack
// 3. source sees ack (2 flops later), drops req
// 4. destination sees req drop, drops ack
// 5. source sees ack drop -- and only NOW may it change the data
//
// That is the FOUR-PHASE handshake, and every one of the four phases is
// load-bearing. The data path crosses the boundary with no synchroniser on
// it whatsoever, which looks alarming and is the entire point: it is not
// sampled while it is changing, so it never needs one.
//
// WHY PHASE 4 CANNOT BE DROPPED
//
// The tempting simplification is three phases: the source drops req on ack
// and declares itself ready. But ack is still HIGH -- it has not been told
// to drop yet -- so the next transfer starts, raises req, and immediately
// sees the PREVIOUS ack. It completes in zero time, the destination never
// captures, and the value is silently lost.
//
// A HANDSHAKE IS SLOW, AND THAT IS THE TRADE
//
// Four phases, each costing SYNC_N destination or source cycles plus edge
// detection. Roughly 5 to 6 cycles of the slower clock per transfer. For
// control values -- an address, a configuration, a status word -- that is
// free. For streaming data it is not, and the answer there is an
// asynchronous FIFO, which is this handshake's idea applied to pointers
// instead of to payload.
package usb_cdc_pkg;
// The three source-side phases, named. They are an enumeration rather
// than an encoding because SS_DROP exists ONLY to wait for the fourth
// phase of the handshake, and a name is the cheapest way to stop somebody
// deleting it.
typedef enum logic [1:0] {
SS_IDLE = 2'd0, // ready; the data may be changed
SS_REQ = 2'd1, // req raised, waiting for ack
SS_DROP = 2'd2 // req dropped, waiting for ack to fall
} src_state_e;
endpackage
module usb_cdc_handshake
import usb_cdc_pkg::*;
#(
parameter int DW = 8,
parameter int SYNC_N = 2 // synchroniser depth (>= 2; see below)
) (
// ---- source domain ----
input logic clk_s,
input logic rst_s_n,
input logic [DW-1:0] s_data,
input logic s_valid, // the source has a value to send
output logic s_ready, // ...and this says it may send it
output logic s_req, // THE wire that crosses, and it is 1 bit
output logic [DW-1:0] s_hold, // ...and this is the bus that crosses with
// no synchroniser on it at all. Exposed
// because "it never moves while req is
// high" is a property, and a property you
// cannot observe is a property you are
// assuming.
output src_state_e s_state,
output logic [31:0] n_sent,
output logic [31:0] n_stalls,
// ---- destination domain ----
input logic clk_d,
input logic rst_d_n,
output logic [DW-1:0] d_data,
output logic d_valid, // ONE destination cycle per transfer
output logic d_ack, // the wire that crosses back
output logic [31:0] n_received
);
// ==================================================================
// SOURCE DOMAIN
// ==================================================================
src_state_e ss_r;
logic [DW-1:0] hold_r; // the CDC data path. NOT synchronised. Held.
logic req_r;
logic [31:0] sent_r, stalls_r;
// The ack, brought into the source domain through SYNC_N flops. This is
// a single bit, which is the only kind of signal a synchroniser is
// allowed to be used on.
logic [SYNC_N-1:0] ack_sync_r;
logic ack_seen;
assign ack_seen = ack_sync_r[SYNC_N-1];
assign s_req = req_r;
assign s_hold = hold_r;
assign s_state = ss_r;
assign s_ready = (ss_r == SS_IDLE);
assign n_sent = sent_r;
assign n_stalls = stalls_r;
always_ff @(posedge clk_s or negedge rst_s_n) begin
if (!rst_s_n) begin
ss_r <= SS_IDLE;
hold_r <= '0;
req_r <= 1'b0;
ack_sync_r <= '0;
sent_r <= '0;
stalls_r <= '0;
end else begin
ack_sync_r <= {ack_sync_r[SYNC_N-2:0], d_ack};
unique case (ss_r)
SS_IDLE: begin
if (s_valid) begin
// ---- PHASE 1. Latch the data, THEN raise req. ----
//
// The latch is what makes the data path safe: from here until
// phase 5 the destination is looking at a value that does not
// move. Passing s_data through combinationally instead would
// put a live, changing bus across the boundary -- which is the
// bit-synchroniser bug with the synchronisers removed.
hold_r <= s_data;
req_r <= 1'b1;
ss_r <= SS_REQ;
end
end
SS_REQ: begin
// ---- PHASE 3. The destination has captured it. Drop req. ----
if (ack_seen) begin
req_r <= 1'b0;
ss_r <= SS_DROP;
end
end
SS_DROP: begin
// ---- PHASE 5. And WAIT for ack to fall before saying ready.
//
// Returning to IDLE here without waiting is the three-phase
// shortcut: the next req would be raised while ack is still high
// from the previous transfer, the source would see it instantly,
// and the destination would never capture anything.
if (!ack_seen) begin
ss_r <= SS_IDLE;
sent_r <= sent_r + 32'd1;
end
end
default: ss_r <= SS_IDLE;
endcase
// A source that wants to send while the handshake is busy. Counted,
// because the cost of a handshake is exactly this and it should be
// measured rather than assumed.
if (s_valid && (ss_r != SS_IDLE)) stalls_r <= stalls_r + 32'd1;
end
end
// ==================================================================
// DESTINATION DOMAIN
// ==================================================================
logic [SYNC_N-1:0] req_sync_r;
logic req_d_r; // one more flop, for edge detection
logic [DW-1:0] d_data_r;
logic d_valid_r, ack_r;
logic [31:0] recv_r;
logic req_seen, req_rise;
assign req_seen = req_sync_r[SYNC_N-1];
// The RISING EDGE, not the level. A level would re-capture and re-pulse
// on every destination cycle for as long as req is high, which at a fast
// destination clock is a single transfer delivered dozens of times.
assign req_rise = req_seen && !req_d_r;
assign d_data = d_data_r;
assign d_valid = d_valid_r;
assign d_ack = ack_r;
assign n_received = recv_r;
always_ff @(posedge clk_d or negedge rst_d_n) begin
if (!rst_d_n) begin
req_sync_r <= '0;
req_d_r <= 1'b0;
d_data_r <= '0;
d_valid_r <= 1'b0;
ack_r <= 1'b0;
recv_r <= '0;
end else begin
req_sync_r <= {req_sync_r[SYNC_N-2:0], s_req};
req_d_r <= req_seen;
d_valid_r <= 1'b0;
if (req_rise) begin
// ---- PHASE 2. Capture the data and acknowledge. ----
//
// hold_r has been stable for at least SYNC_N destination cycles by
// now -- that is precisely what the synchroniser bought -- so this
// read is of a static value and needs no synchroniser of its own.
d_data_r <= hold_r;
d_valid_r <= 1'b1;
ack_r <= 1'b1;
recv_r <= recv_r + 32'd1;
end else if (!req_seen) begin
// ---- PHASE 4. req has fallen; release ack. ----
ack_r <= 1'b0;
end
end
end
endmodule8. VHDL-2008 Implementation
-- usb_cdc_handshake -- crossing a multi-bit value between two clocks, and
-- the reason you cannot do it the obvious way.
--
-- THE ASSUMPTION THAT JUST BROKE
--
-- Chapters 23.1 to 23.4 all had one clock. A USB device does not. The PHY
-- runs at the bus rate -- 60 MHz for a UTMI high-speed interface -- and the
-- controller, the FIFOs and the CPU do not. Every value that moves between
-- them crosses a boundary where the two clocks have no fixed relationship
-- at all.
--
-- A MULTI-BIT VALUE CANNOT BE BIT-SYNCHRONISED
--
-- The two-flop synchroniser is the standard answer for a single bit, and it
-- is correct: it gives the first flop's metastability time to resolve before
-- anything downstream looks at it.
--
-- Put one on each bit of a bus and the design is broken.
--
-- source changes 0111 -> 1000 (7 -> 8, four bits at once)
--
-- Each bit's synchroniser resolves INDEPENDENTLY. There is no
-- mechanism that makes them agree. The destination can sample
--
-- 0111 1111 1000 0000 1011 0100 ...
--
-- any of the sixteen values, for one cycle, before settling.
--
-- Those intermediate values are not noise to be filtered. They are indices
-- into a descriptor table, byte counts, endpoint numbers. A transient 0000
-- on an endpoint number is a packet delivered to endpoint 0.
--
-- And a Gray code does not rescue the general case. It works for a counter,
-- where only one bit changes per increment BY CONSTRUCTION. Arbitrary data
-- changes arbitrarily many bits at once, and no encoding fixes that.
--
-- THE ANSWER: DO NOT SYNCHRONISE THE DATA AT ALL
--
-- Synchronise ONE bit -- a request -- and use it to say "the data has been
-- stable for a long time and will stay stable until I hear back".
--
-- 1. source latches the data, raises req
-- 2. destination sees req (2 flops later), captures the data, raises ack
-- 3. source sees ack (2 flops later), drops req
-- 4. destination sees req drop, drops ack
-- 5. source sees ack drop -- and only NOW may it change the data
--
-- That is the FOUR-PHASE handshake, and every one of the four phases is
-- load-bearing. The data path crosses the boundary with no synchroniser on
-- it whatsoever, which looks alarming and is the entire point: it is not
-- sampled while it is changing, so it never needs one.
--
-- WHY PHASE 4 CANNOT BE DROPPED
--
-- The tempting simplification is three phases: the source drops req on ack
-- and declares itself ready. But ack is still HIGH -- it has not been told
-- to drop yet -- so the next transfer starts, raises req, and immediately
-- sees the PREVIOUS ack. It completes in zero time, the destination never
-- captures, and the value is silently lost.
--
-- A HANDSHAKE IS SLOW, AND THAT IS THE TRADE
--
-- Four phases, each costing SYNC_N destination or source cycles plus edge
-- detection. Roughly 5 to 6 cycles of the slower clock per transfer. For
-- control values -- an address, a configuration, a status word -- that is
-- free. For streaming data it is not, and the answer there is an
-- asynchronous FIFO, which is this handshake's idea applied to pointers
-- instead of to payload.
library ieee;
use ieee.std_logic_1164.all;
package usb_cdc_pkg is
-- The three source-side phases, named. They are an enumeration rather
-- than an encoding because SS_DROP exists ONLY to wait for the fourth
-- phase of the handshake, and a name is the cheapest way to stop somebody
-- deleting it.
type src_state_t is (SS_IDLE, SS_REQ, SS_DROP);
function ss_code (s : src_state_t) return std_logic_vector;
end package usb_cdc_pkg;
package body usb_cdc_pkg is
function ss_code (s : src_state_t) return std_logic_vector is
begin
case s is
when SS_IDLE => return "00";
when SS_REQ => return "01";
when SS_DROP => return "10";
end case;
end function;
end package body usb_cdc_pkg;
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
use work.usb_cdc_pkg.all;
entity usb_cdc_handshake is
generic (
DW : integer := 8;
SYNC_N : integer := 2 -- synchroniser depth (>= 2; see below)
);
port (
-- ---- source domain ----
clk_s : in std_logic;
rst_s_n : in std_logic;
s_data : in std_logic_vector(DW-1 downto 0);
s_valid : in std_logic; -- a value to send
s_ready : out std_logic; -- ...and it may go
s_req : out std_logic; -- THE wire: 1 bit
s_hold : out std_logic_vector(DW-1 downto 0); -- the bus that
-- crosses with no
-- synchroniser at all
s_state : out std_logic_vector(1 downto 0);
n_sent : out std_logic_vector(31 downto 0);
n_stalls : out std_logic_vector(31 downto 0);
-- ---- destination domain ----
clk_d : in std_logic;
rst_d_n : in std_logic;
d_data : out std_logic_vector(DW-1 downto 0);
d_valid : out std_logic; -- ONE destination cycle per transfer
d_ack : out std_logic; -- the wire that crosses back
n_received : out std_logic_vector(31 downto 0)
);
end entity usb_cdc_handshake;
architecture rtl of usb_cdc_handshake is
-- ================================================================
-- SOURCE DOMAIN
-- ================================================================
signal ss_r : src_state_t := SS_IDLE;
signal hold_r : std_logic_vector(DW-1 downto 0) := (others => '0');
signal req_r : std_logic := '0';
signal sent_r : unsigned(31 downto 0) := (others => '0');
signal stalls_r : unsigned(31 downto 0) := (others => '0');
-- The ack, brought into the source domain through SYNC_N flops. This is
-- a single bit, which is the only kind of signal a synchroniser is
-- allowed to be used on.
signal ack_sync_r : std_logic_vector(SYNC_N-1 downto 0) := (others => '0');
signal ack_seen : std_logic;
-- ================================================================
-- DESTINATION DOMAIN
-- ================================================================
signal req_sync_r : std_logic_vector(SYNC_N-1 downto 0) := (others => '0');
signal req_d_r : std_logic := '0';
signal d_data_r : std_logic_vector(DW-1 downto 0) := (others => '0');
signal d_valid_r : std_logic := '0';
signal ack_r : std_logic := '0';
signal recv_r : unsigned(31 downto 0) := (others => '0');
signal req_seen, req_rise : std_logic;
begin
ack_seen <= ack_sync_r(SYNC_N-1);
s_req <= req_r;
s_hold <= hold_r;
s_state <= ss_code(ss_r);
s_ready <= '1' when ss_r = SS_IDLE else '0';
n_sent <= std_logic_vector(sent_r);
n_stalls <= std_logic_vector(stalls_r);
src : process (clk_s, rst_s_n)
begin
if rst_s_n = '0' then
ss_r <= SS_IDLE;
hold_r <= (others => '0');
req_r <= '0';
ack_sync_r <= (others => '0');
sent_r <= (others => '0');
stalls_r <= (others => '0');
elsif rising_edge(clk_s) then
ack_sync_r <= ack_sync_r(SYNC_N-2 downto 0) & d_ack;
case ss_r is
when SS_IDLE =>
if s_valid = '1' then
-- ---- PHASE 1. Latch the data, THEN raise req. ----
--
-- The latch is what makes the data path safe: from here until
-- phase 5 the destination is looking at a value that does not
-- move. Passing s_data through combinationally instead would
-- put a live, changing bus across the boundary -- which is the
-- bit-synchroniser bug with the synchronisers removed.
hold_r <= s_data;
req_r <= '1';
ss_r <= SS_REQ;
end if;
when SS_REQ =>
-- ---- PHASE 3. The destination has captured it. Drop req. ----
if ack_seen = '1' then
req_r <= '0';
ss_r <= SS_DROP;
end if;
when SS_DROP =>
-- ---- PHASE 5. And WAIT for ack to fall before saying ready.
--
-- Returning to IDLE here without waiting is the three-phase
-- shortcut: the next req would be raised while ack is still high
-- from the previous transfer, the source would see it instantly,
-- and the destination would never capture anything.
if ack_seen = '0' then
ss_r <= SS_IDLE;
sent_r <= sent_r + 1;
end if;
end case;
-- A source that wants to send while the handshake is busy. Counted,
-- because the cost of a handshake is exactly this and it should be
-- measured rather than assumed.
if s_valid = '1' and ss_r /= SS_IDLE then
stalls_r <= stalls_r + 1;
end if;
end if;
end process;
req_seen <= req_sync_r(SYNC_N-1);
-- The RISING EDGE, not the level. A level would re-capture and re-pulse
-- on every destination cycle for as long as req is high, which at a fast
-- destination clock is a single transfer delivered dozens of times.
req_rise <= req_seen and (not req_d_r);
d_data <= d_data_r;
d_valid <= d_valid_r;
d_ack <= ack_r;
n_received <= std_logic_vector(recv_r);
dst : process (clk_d, rst_d_n)
begin
if rst_d_n = '0' then
req_sync_r <= (others => '0');
req_d_r <= '0';
d_data_r <= (others => '0');
d_valid_r <= '0';
ack_r <= '0';
recv_r <= (others => '0');
elsif rising_edge(clk_d) then
-- req_r rather than the s_req port: VHDL will not let an entity read
-- its own output, and this is the identical wire. It is the one bit
-- that crosses the boundary in this direction.
req_sync_r <= req_sync_r(SYNC_N-2 downto 0) & req_r;
req_d_r <= req_seen;
d_valid_r <= '0';
if req_rise = '1' then
-- ---- PHASE 2. Capture the data and acknowledge. ----
--
-- hold_r has been stable for at least SYNC_N destination cycles by
-- now -- that is precisely what the synchroniser bought -- so this
-- read is of a static value and needs no synchroniser of its own.
d_data_r <= hold_r;
d_valid_r <= '1';
ack_r <= '1';
recv_r <= recv_r + 1;
elsif req_seen = '0' then
-- ---- PHASE 4. req has fallen; release ack. ----
ack_r <= '0';
end if;
end if;
end process;
end architecture rtl;9. Seeing One Transfer Cross
One value crossing, with the destination clock slower than the source
usb_cdc_handshake — the four phases, in destination cycles
10 cycless_hold holds 5A for eight destination cycles and changes only at cycle 9, after s_ready returns. Every one of those cycles is a cycle in which the destination could have been reading it, and none of them is a cycle in which it was moving.
10. The Testbenches
The block claims to work at any clock ratio, so the suite sweeps the ratio instead of picking one:
5 : 7 destination a little slower
7 : 5 source a little slower
5 : 31 destination MUCH slower
31 : 5 source MUCH slower
3 : 11 co-prime, close
11 : 3
5 : 6 nearly equal, never aligned
6 : 5
13 : 13 equal periods, offset phase
2 : 29 extreme
29 : 2
...and then a phase where BOTH periods change on every
transfer, while transfers are in flight.The two clocks are generated with a deliberate phase offset so that no edge in one domain ever coincides exactly with an edge in the other — which is the one alignment a real chip is guaranteed not to have, and the one a lazy testbench accidentally arranges.
Four things are checked:
- Every value arrives exactly once, in order, unchanged, compared element by element against an independent queue.
- The bus that crosses never moves while
reqis high. The driver deliberately puts a different value ons_datathe instant a transfer is taken. d_validis exactly one destination cycle per transfer.- The minimum observed latency is at least
SYNC_N + 1destination cycles.
And every wait is bounded:
// Every wait is BOUNDED. A handshake that has lost a transfer does
// not produce a wrong answer -- it produces no answer at all, and an
// unbounded wait turns that into a suite that hangs instead of a
// suite that fails. The bound is far larger than the slowest legal
// handshake at the most extreme clock ratio in the sweep.
guard = 0;
while (!s_ready && (guard < 600)) begin @(posedge clk_s); #1; guard = guard + 1; end
check(s_ready === 1'b1,
"the source never became ready -- the handshake is deadlocked, which is what a lost ack looks like from this side");10.1 Verilog testbench
// Testbench for usb_cdc_handshake (Verilog-2005).
//
// TWO CLOCKS, AND NO RELATIONSHIP BETWEEN THEM
//
// The whole point of the block is that it works for ANY ratio, so the suite
// sweeps the ratio rather than picking one: source faster, destination
// faster, nearly equal, wildly unequal, and a phase where the destination
// period CHANGES while transfers are in flight.
//
// WHAT IS CHECKED
//
// 1. Every value sent arrives EXACTLY ONCE, IN ORDER, UNCHANGED, at
// every ratio. Compared element by element against an independent
// queue, not by counting.
//
// 2. The bus that crosses NEVER MOVES while req is high. That is the
// property that makes an unsynchronised data path safe, and the
// source deliberately drives a different value onto s_data the
// instant the transfer is taken, so a design that passed s_data
// through combinationally would be caught immediately.
//
// 3. d_valid is exactly ONE destination cycle per transfer. A level-
// triggered capture delivers the same value dozens of times at a fast
// destination clock, and a counter alone would not notice the
// difference between that and a fast source.
//
// 4. The minimum observed latency is at least SYNC_N+1 destination
// cycles -- the signature of a synchroniser that really is SYNC_N
// deep. See the note in the chapter about what this check can and
// cannot prove.
`timescale 1ns/1ps
module tb_ch_v;
localparam integer DW = 8;
localparam integer SYNC_N = 2;
localparam [1:0] SS_IDLE = 2'd0, SS_REQ = 2'd1, SS_DROP = 2'd2;
integer sp = 5; // source half period
integer dp = 7; // destination half period
reg clk_s = 1'b0, rst_s_n = 1'b0;
reg clk_d = 1'b0, rst_d_n = 1'b0;
reg [DW-1:0] s_data = 8'd0;
reg s_valid = 1'b0;
wire s_ready, s_req, d_valid, d_ack;
wire [DW-1:0] s_hold, d_data;
wire [1:0] s_state;
wire [31:0] n_sent, n_stalls, n_received;
usb_cdc_handshake #(.DW(DW), .SYNC_N(SYNC_N)) dut (
.clk_s(clk_s), .rst_s_n(rst_s_n),
.s_data(s_data), .s_valid(s_valid), .s_ready(s_ready),
.s_req(s_req), .s_hold(s_hold), .s_state(s_state),
.n_sent(n_sent), .n_stalls(n_stalls),
.clk_d(clk_d), .rst_d_n(rst_d_n),
.d_data(d_data), .d_valid(d_valid), .d_ack(d_ack),
.n_received(n_received)
);
// Two free-running clocks with a deliberate phase offset, so no edge in
// one domain ever coincides exactly with an edge in the other -- which
// would be the one alignment a real chip is guaranteed NOT to have.
initial forever #sp clk_s = ~clk_s;
initial begin #3; forever #dp clk_d = ~clk_d; end
integer errors = 0, checks = 0;
task check(input cond, input [1023:0] msg);
begin
checks = checks + 1;
if (!cond) begin
errors = errors + 1;
if (errors <= 25)
$display("FAIL @%0t: %0s | st=%0d req=%b ack=%b hold=%h d=%h",
$time, msg, s_state, s_req, d_ack, s_hold, d_data);
end
end
endtask
// ---- the two independent queues ----
reg [DW-1:0] txq [0:20000];
reg [DW-1:0] rxq [0:20000];
integer tx_n = 0, rx_n = 0;
integer min_lat = 1000000, max_lat = 0;
integer n_ratios = 0;
// ---- destination-side observation ----
integer d_age = 0;
reg d_valid_d = 1'b0;
always @(posedge clk_d or negedge rst_d_n) begin
if (!rst_d_n) d_age <= 0;
else if (!s_req) d_age <= 0;
else d_age <= d_age + 1;
end
always @(posedge clk_d) begin
if (rst_d_n) begin
// ---- d_valid is ONE cycle. Two in a row is a level-triggered
// ---- capture delivering the same value repeatedly.
check(!(d_valid && d_valid_d),
"d_valid was asserted for two consecutive destination cycles -- the capture is level-triggered, not edge-triggered");
d_valid_d <= d_valid;
if (d_valid) begin
// The index is bounded. A level-triggered capture delivers orders
// of magnitude more values than a correct design does, and a queue
// that runs off its end turns a measured kill into a crash.
if (rx_n <= 20000) rxq[rx_n] = d_data;
rx_n = rx_n + 1;
if (d_age < min_lat) min_lat = d_age;
if (d_age > max_lat) max_lat = d_age;
// ---- the synchroniser really is SYNC_N deep ----
check(d_age >= SYNC_N + 1,
"a value arrived sooner than SYNC_N synchroniser stages can possibly allow -- a stage has been bypassed");
// Note what is NOT checked here: that req is still high. The
// monitor observes d_valid one destination cycle after the capture,
// and a fast source can already have seen ack and completed phases
// 3 to 5 by then. "One delivery per request" is a statement about
// totals, so it is checked as one -- see n_req_rises below.
end
end
end
// ---- source-side observation: the bus that crosses must NOT MOVE ----
reg [DW-1:0] hold_prev = 8'd0;
reg req_prev = 1'b0;
integer n_req_rises = 0;
always @(posedge clk_s) begin
if (rst_s_n) begin
// From the SECOND cycle of req onward. The edge that raises req is
// the same edge that loads the hold register, so comparing across it
// would flag the load itself.
if (s_req && req_prev)
check(s_hold === hold_prev,
"the bus that crosses the boundary changed while req was asserted -- the destination may be sampling it mid-change, which is the bit-synchroniser bug in its purest form");
// ---- and the source must not declare itself ready mid-transfer ----
check(!(s_ready && s_req),
"s_ready was asserted while a transfer was still in flight -- the next value can overwrite the one being captured");
check(!(s_ready && (s_state != SS_IDLE)),
"s_ready disagrees with the source state");
if (s_req && !req_prev) n_req_rises = n_req_rises + 1;
end
hold_prev <= s_hold;
req_prev <= s_req;
end
// Hand one value to the source, wait until the handshake has taken it,
// and then IMMEDIATELY drive something else onto s_data -- so that a
// design which forwarded s_data combinationally is caught at once.
integer guard;
task send(input [DW-1:0] v);
begin
@(posedge clk_s); #1;
// Every wait is BOUNDED. A handshake that has lost a transfer does
// not produce a wrong answer -- it produces no answer at all, and an
// unbounded wait turns that into a suite that hangs instead of a
// suite that fails. The bound is far larger than the slowest legal
// handshake at the most extreme clock ratio in the sweep.
guard = 0;
while (!s_ready && (guard < 600)) begin @(posedge clk_s); #1; guard = guard + 1; end
check(s_ready === 1'b1,
"the source never became ready -- the handshake is deadlocked, which is what a lost ack looks like from this side");
s_data = v;
s_valid = 1'b1;
@(posedge clk_s); #1;
guard = 0;
while ((s_state == SS_IDLE) && (guard < 600)) begin @(posedge clk_s); #1; guard = guard + 1; end
check(s_state !== SS_IDLE,
"the source never accepted the value offered to it");
s_valid = 1'b0;
txq[tx_n] = v;
tx_n = tx_n + 1;
s_data = ~v; // the value is now HELD, or it is not
end
endtask
task drain;
integer g;
begin
for (g = 0; g < 400; g = g + 1) @(posedge clk_d);
for (g = 0; g < 400; g = g + 1) @(posedge clk_s);
end
endtask
// One ratio: reset nothing, just send a run of values and confirm the
// queues still agree. State is deliberately NOT cleared between ratios,
// so a transfer in flight when the period changes is part of the test.
task run_ratio(input integer new_sp, input integer new_dp, input integer count);
integer g;
begin
sp = new_sp;
dp = new_dp;
n_ratios = n_ratios + 1;
for (g = 0; g < count; g = g + 1)
send(((tx_n * 8'd37) ^ (tx_n >> 3)) & 8'hff);
drain;
check(rx_n == tx_n,
"the number of values received does not equal the number sent at this clock ratio");
check(n_received === n_sent,
"the design's own source and destination counters disagree");
end
endtask
integer i, cmp_n;
initial begin
repeat (6) @(posedge clk_s);
rst_s_n = 1'b1;
repeat (6) @(posedge clk_d);
rst_d_n = 1'b1;
repeat (4) @(posedge clk_s);
// ---- Phase A: the state after reset ----
#1;
check(s_ready === 1'b1, "the source is not ready after reset");
check(s_req === 1'b0, "req is asserted after reset");
check(d_ack === 1'b0, "ack is asserted after reset");
check(d_valid === 1'b0, "d_valid is asserted after reset");
// ---- Phase B: the ratio sweep. Every shape of relationship the two
// ---- clocks can have, including one that changes mid-flight.
run_ratio( 5, 7, 300); // destination a little slower
run_ratio( 7, 5, 300); // source a little slower
run_ratio( 5, 31, 300); // destination MUCH slower
run_ratio(31, 5, 300); // source MUCH slower
run_ratio( 3, 11, 300); // co-prime, close
run_ratio(11, 3, 300);
run_ratio( 5, 6, 300); // nearly equal, never aligned
run_ratio( 6, 5, 300);
run_ratio(13, 13, 300); // equal periods, offset phase
run_ratio( 2, 29, 200); // extreme
run_ratio(29, 2, 200);
// ---- Phase C: the destination period CHANGES while transfers are in
// ---- flight. A handshake that depended on a fixed ratio anywhere --
// ---- a fixed-delay req, a timed ack -- comes apart here.
for (i = 0; i < 600; i = i + 1) begin
dp = 2 + (i % 17);
sp = 3 + ((i * 7) % 13);
send(((tx_n * 8'd91) ^ 8'h5a) & 8'hff);
end
drain;
check(rx_n == tx_n, "a value was lost while the clock ratio was changing");
// Everything up to here went through send(), so the two queues are
// element-for-element comparable. Phase D drives the source directly
// and is checked by value instead.
cmp_n = tx_n;
// ---- Phase D: the source offers values back to back with no gap, so
// ---- s_valid is high through entire handshakes and the stall path is
// ---- exercised rather than merely present.
sp = 5; dp = 9;
s_data = 8'hA5;
s_valid = 1'b1;
for (i = 0; i < 2000; i = i + 1) @(posedge clk_s);
s_valid = 1'b0;
drain;
check(n_stalls > 32'd0, "the source never had to wait, so the stall path was never exercised");
// ---- Final: every value, in order, unchanged. ----
check(rx_n >= cmp_n, "fewer values arrived than were sent through send()");
// EVERY mismatch is counted, not just the first. A design that puts a
// live bus across the boundary gets one wrong value per transfer, and
// stopping at the first would report that as a single error -- which
// says nothing about how pervasive it is.
for (i = 0; i < cmp_n; i = i + 1)
check(rxq[i] === txq[i], "a value arrived corrupted or out of order");
// Phase D held a single value on s_data throughout, so every delivery
// from there on must be exactly that value -- a design that forwarded
// a live bus would deliver the inverted copy the driver leaves behind.
for (i = cmp_n; (i < rx_n) && (i <= 20000); i = i + 1)
check(rxq[i] === 8'hA5,
"a back-to-back transfer delivered a value that was never offered");
check(cmp_n > 3000, "the suite did not send enough values to mean anything");
check(n_received === n_sent,
"the design's own source and destination counters disagree at the end of the run");
// The latency is DETERMINISTIC in destination cycles -- SYNC_N stages
// plus one for the edge detector -- no matter what the clock ratio is.
// That is the signature of a synchroniser that is really SYNC_N deep,
// and it is the same at every ratio because it is counted in the
// destination's own cycles.
check(min_lat >= SYNC_N + 1,
"the fastest observed delivery was quicker than SYNC_N synchroniser stages allow -- a stage has been bypassed");
check(max_lat <= SYNC_N + 2,
"a delivery took longer than the synchroniser plus edge detection can account for");
check(n_ratios == 11, "not every clock ratio was exercised");
// ---- EXACTLY ONE DELIVERY PER REQUEST. Not "at least one", which a
// ---- level-triggered capture also satisfies, and not "at most one",
// ---- which a handshake that loses transfers also satisfies.
check(rx_n == n_req_rises,
"the number of values delivered does not equal the number of requests raised");
$display("REACH ratios=%0d transfers=%0d received=%0d latency=%0d..%0d dest cycles (bound >= %0d)",
n_ratios, cmp_n, rx_n, min_lat, max_lat, SYNC_N + 1);
$display("COUNTERS sent=%0d received=%0d requests=%0d stalls=%0d",
n_sent, n_received, n_req_rises, n_stalls);
$display("%0s: %0d errors in %0d checks", (errors==0)?"PASS":"FAIL", errors, checks);
$finish;
end
endmodule10.2 SystemVerilog testbench
// Testbench for usb_cdc_handshake (SystemVerilog).
//
// TWO CLOCKS, AND NO RELATIONSHIP BETWEEN THEM
//
// The whole point of the block is that it works for ANY ratio, so the suite
// sweeps the ratio rather than picking one: source faster, destination
// faster, nearly equal, wildly unequal, and a phase where the destination
// period CHANGES while transfers are in flight.
//
// WHAT IS CHECKED
//
// 1. Every value sent arrives EXACTLY ONCE, IN ORDER, UNCHANGED, at
// every ratio. Compared element by element against an independent
// queue, not by counting.
//
// 2. The bus that crosses NEVER MOVES while req is high. That is the
// property that makes an unsynchronised data path safe, and the
// source deliberately drives a different value onto s_data the
// instant the transfer is taken, so a design that passed s_data
// through combinationally would be caught immediately.
//
// 3. d_valid is exactly ONE destination cycle per transfer. A level-
// triggered capture delivers the same value dozens of times at a fast
// destination clock, and a counter alone would not notice the
// difference between that and a fast source.
//
// 4. The minimum observed latency is at least SYNC_N+1 destination
// cycles -- the signature of a synchroniser that really is SYNC_N
// deep. See the note in the chapter about what this check can and
// cannot prove.
`timescale 1ns/1ps
module tb_ch_sv;
import usb_cdc_pkg::*;
localparam int DW = 8;
localparam int SYNC_N = 2;
int sp = 5; // source half period
int dp = 7; // destination half period
logic clk_s = 1'b0, rst_s_n = 1'b0;
logic clk_d = 1'b0, rst_d_n = 1'b0;
logic [DW-1:0] s_data = '0;
logic s_valid = 1'b0;
logic s_ready, s_req, d_valid, d_ack;
logic [DW-1:0] s_hold, d_data;
src_state_e s_state;
logic [31:0] n_sent, n_stalls, n_received;
usb_cdc_handshake #(.DW(DW), .SYNC_N(SYNC_N)) dut (.*);
// Two free-running clocks with a deliberate phase offset, so no edge in
// one domain ever coincides exactly with an edge in the other -- which
// would be the one alignment a real chip is guaranteed NOT to have.
initial forever #sp clk_s = ~clk_s;
initial begin #3; forever #dp clk_d = ~clk_d; end
int errors = 0, checks = 0;
task automatic check(input logic cond, input string msg);
checks++;
if (!cond) begin
errors++;
if (errors <= 25)
$display("FAIL @%0t: %0s | st=%0d req=%b ack=%b hold=%h d=%h",
$time, msg, s_state, s_req, d_ack, s_hold, d_data);
end
endtask
// ---- the two independent queues ----
logic [DW-1:0] txq [20001];
logic [DW-1:0] rxq [20001];
int tx_n = 0, rx_n = 0;
int min_lat = 1000000, max_lat = 0;
int n_ratios = 0;
// ---- destination-side observation ----
int d_age = 0;
logic d_valid_d = 1'b0;
always @(posedge clk_d or negedge rst_d_n) begin
if (!rst_d_n) d_age <= 0;
else if (!s_req) d_age <= 0;
else d_age <= d_age + 1;
end
always @(posedge clk_d) begin
if (rst_d_n) begin
// ---- d_valid is ONE cycle. Two in a row is a level-triggered
// ---- capture delivering the same value repeatedly.
check(!(d_valid && d_valid_d),
"d_valid was asserted for two consecutive destination cycles -- the capture is level-triggered, not edge-triggered");
d_valid_d <= d_valid;
if (d_valid) begin
// The index is bounded. A level-triggered capture delivers orders
// of magnitude more values than a correct design does, and a queue
// that runs off its end turns a measured kill into a crash.
if (rx_n <= 20000) rxq[rx_n] = d_data;
rx_n++;
if (d_age < min_lat) min_lat = d_age;
if (d_age > max_lat) max_lat = d_age;
// ---- the synchroniser really is SYNC_N deep ----
check(d_age >= SYNC_N + 1,
"a value arrived sooner than SYNC_N synchroniser stages can possibly allow -- a stage has been bypassed");
// Note what is NOT checked here: that req is still high. The
// monitor observes d_valid one destination cycle after the capture,
// and a fast source can already have seen ack and completed phases
// 3 to 5 by then. "One delivery per request" is a statement about
// totals, so it is checked as one -- see n_req_rises below.
end
end
end
// ---- source-side observation: the bus that crosses must NOT MOVE ----
logic [DW-1:0] hold_prev = '0;
logic req_prev = 1'b0;
int n_req_rises = 0;
always @(posedge clk_s) begin
if (rst_s_n) begin
// From the SECOND cycle of req onward. The edge that raises req is
// the same edge that loads the hold register, so comparing across it
// would flag the load itself.
if (s_req && req_prev)
check(s_hold === hold_prev,
"the bus that crosses the boundary changed while req was asserted -- the destination may be sampling it mid-change, which is the bit-synchroniser bug in its purest form");
// ---- and the source must not declare itself ready mid-transfer ----
check(!(s_ready && s_req),
"s_ready was asserted while a transfer was still in flight -- the next value can overwrite the one being captured");
check(!(s_ready && (s_state != SS_IDLE)),
"s_ready disagrees with the source state");
if (s_req && !req_prev) n_req_rises++;
end
hold_prev <= s_hold;
req_prev <= s_req;
end
// Hand one value to the source, wait until the handshake has taken it,
// and then IMMEDIATELY drive something else onto s_data -- so that a
// design which forwarded s_data combinationally is caught at once.
int guard;
task automatic send(input logic [DW-1:0] v);
@(posedge clk_s); #1;
// Every wait is BOUNDED. A handshake that has lost a transfer does
// not produce a wrong answer -- it produces no answer at all, and an
// unbounded wait turns that into a suite that hangs instead of a
// suite that fails. The bound is far larger than the slowest legal
// handshake at the most extreme clock ratio in the sweep.
guard = 0;
while (!s_ready && (guard < 600)) begin @(posedge clk_s); #1; guard++; end
check(s_ready === 1'b1,
"the source never became ready -- the handshake is deadlocked, which is what a lost ack looks like from this side");
s_data = v;
s_valid = 1'b1;
@(posedge clk_s); #1;
guard = 0;
while ((s_state == SS_IDLE) && (guard < 600)) begin @(posedge clk_s); #1; guard++; end
check(s_state !== SS_IDLE,
"the source never accepted the value offered to it");
s_valid = 1'b0;
txq[tx_n] = v;
tx_n++;
s_data = ~v; // the value is now HELD, or it is not
endtask
task automatic drain();
repeat (400) @(posedge clk_d);
repeat (400) @(posedge clk_s);
endtask
// One ratio: reset nothing, just send a run of values and confirm the
// queues still agree. State is deliberately NOT cleared between ratios,
// so a transfer in flight when the period changes is part of the test.
task automatic run_ratio(input int new_sp, input int new_dp, input int count);
sp = new_sp;
dp = new_dp;
n_ratios++;
repeat (count)
send(8'((tx_n * 37) ^ (tx_n >> 3)));
drain();
check(rx_n == tx_n,
"the number of values received does not equal the number sent at this clock ratio");
check(n_received === n_sent,
"the design's own source and destination counters disagree");
endtask
int i, cmp_n;
initial begin
repeat (6) @(posedge clk_s);
rst_s_n = 1'b1;
repeat (6) @(posedge clk_d);
rst_d_n = 1'b1;
repeat (4) @(posedge clk_s);
// ---- Phase A: the state after reset ----
#1;
check(s_ready === 1'b1, "the source is not ready after reset");
check(s_req === 1'b0, "req is asserted after reset");
check(d_ack === 1'b0, "ack is asserted after reset");
check(d_valid === 1'b0, "d_valid is asserted after reset");
// ---- Phase B: the ratio sweep. Every shape of relationship the two
// ---- clocks can have, including one that changes mid-flight.
run_ratio( 5, 7, 300); // destination a little slower
run_ratio( 7, 5, 300); // source a little slower
run_ratio( 5, 31, 300); // destination MUCH slower
run_ratio(31, 5, 300); // source MUCH slower
run_ratio( 3, 11, 300); // co-prime, close
run_ratio(11, 3, 300);
run_ratio( 5, 6, 300); // nearly equal, never aligned
run_ratio( 6, 5, 300);
run_ratio(13, 13, 300); // equal periods, offset phase
run_ratio( 2, 29, 200); // extreme
run_ratio(29, 2, 200);
// ---- Phase C: the destination period CHANGES while transfers are in
// ---- flight. A handshake that depended on a fixed ratio anywhere --
// ---- a fixed-delay req, a timed ack -- comes apart here.
for (i = 0; i < 600; i++) begin
dp = 2 + (i % 17);
sp = 3 + ((i * 7) % 13);
send(8'((tx_n * 91) ^ 8'h5a));
end
drain();
check(rx_n == tx_n, "a value was lost while the clock ratio was changing");
// Everything up to here went through send(), so the two queues are
// element-for-element comparable. Phase D drives the source directly
// and is checked by value instead.
cmp_n = tx_n;
// ---- Phase D: the source offers values back to back with no gap, so
// ---- s_valid is high through entire handshakes and the stall path is
// ---- exercised rather than merely present.
sp = 5; dp = 9;
s_data = 8'hA5;
s_valid = 1'b1;
repeat (2000) @(posedge clk_s);
s_valid = 1'b0;
drain();
check(n_stalls > 32'd0, "the source never had to wait, so the stall path was never exercised");
// ---- Final: every value, in order, unchanged. ----
check(rx_n >= cmp_n, "fewer values arrived than were sent through send()");
// EVERY mismatch is counted, not just the first. A design that puts a
// live bus across the boundary gets one wrong value per transfer, and
// stopping at the first would report that as a single error -- which
// says nothing about how pervasive it is.
for (i = 0; i < cmp_n; i++)
check(rxq[i] === txq[i], "a value arrived corrupted or out of order");
// Phase D held a single value on s_data throughout, so every delivery
// from there on must be exactly that value -- a design that forwarded
// a live bus would deliver the inverted copy the driver leaves behind.
for (i = cmp_n; (i < rx_n) && (i <= 20000); i++)
check(rxq[i] === 8'hA5,
"a back-to-back transfer delivered a value that was never offered");
check(cmp_n > 3000, "the suite did not send enough values to mean anything");
check(n_received === n_sent,
"the design's own source and destination counters disagree at the end of the run");
// The latency is DETERMINISTIC in destination cycles -- SYNC_N stages
// plus one for the edge detector -- no matter what the clock ratio is.
// That is the signature of a synchroniser that is really SYNC_N deep,
// and it is the same at every ratio because it is counted in the
// destination's own cycles.
check(min_lat >= SYNC_N + 1,
"the fastest observed delivery was quicker than SYNC_N synchroniser stages allow -- a stage has been bypassed");
check(max_lat <= SYNC_N + 2,
"a delivery took longer than the synchroniser plus edge detection can account for");
check(n_ratios == 11, "not every clock ratio was exercised");
// ---- EXACTLY ONE DELIVERY PER REQUEST. Not "at least one", which a
// ---- level-triggered capture also satisfies, and not "at most one",
// ---- which a handshake that loses transfers also satisfies.
check(rx_n == n_req_rises,
"the number of values delivered does not equal the number of requests raised");
$display("REACH ratios=%0d transfers=%0d received=%0d latency=%0d..%0d dest cycles (bound >= %0d)",
n_ratios, cmp_n, rx_n, min_lat, max_lat, SYNC_N + 1);
$display("COUNTERS sent=%0d received=%0d requests=%0d stalls=%0d",
n_sent, n_received, n_req_rises, n_stalls);
$display("%0s: %0d errors in %0d checks", (errors==0)?"PASS":"FAIL", errors, checks);
$finish;
end
endmodule10.3 VHDL testbench
-- Testbench for usb_cdc_handshake (VHDL-2008).
--
-- TWO CLOCKS, AND NO RELATIONSHIP BETWEEN THEM
--
-- The whole point of the block is that it works for ANY ratio, so the suite
-- sweeps the ratio rather than picking one: source faster, destination
-- faster, nearly equal, wildly unequal, and a phase where the destination
-- period CHANGES while transfers are in flight.
--
-- WHAT IS CHECKED
--
-- 1. Every value sent arrives EXACTLY ONCE, IN ORDER, UNCHANGED, at
-- every ratio. Compared element by element against an independent
-- queue, not by counting.
--
-- 2. The bus that crosses NEVER MOVES while req is high. That is the
-- property that makes an unsynchronised data path safe, and the
-- source deliberately drives a different value onto s_data the
-- instant the transfer is taken, so a design that passed s_data
-- through combinationally would be caught immediately.
--
-- 3. d_valid is exactly ONE destination cycle per transfer. A level-
-- triggered capture delivers the same value dozens of times at a fast
-- destination clock, and a counter alone would not notice the
-- difference between that and a fast source.
--
-- 4. The minimum observed latency is at least SYNC_N+1 destination
-- cycles -- the signature of a synchroniser that really is SYNC_N
-- deep. See the note in the chapter about what this check can and
-- cannot prove.
--
-- The checks are spread over three processes because they live in two
-- different clock domains, so each process keeps its own error and check
-- tallies and the totals are summed at the end. There is no shared mutable
-- state between them, which is the same discipline the design itself
-- follows.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
use std.textio.all;
use work.usb_cdc_pkg.all;
entity tb_ch_vhdl is
end entity tb_ch_vhdl;
architecture sim of tb_ch_vhdl is
constant DW : integer := 8;
constant SYNC_N : integer := 2;
type byte_arr is array (0 to 20000) of std_logic_vector(DW-1 downto 0);
signal sp : time := 5 ns; -- source half period
signal dp : time := 7 ns; -- destination half period
signal clk_s : std_logic := '0';
signal rst_s_n : std_logic := '0';
signal clk_d : std_logic := '0';
signal rst_d_n : std_logic := '0';
signal done : boolean := false;
signal s_data : std_logic_vector(DW-1 downto 0) := (others => '0');
signal s_valid : std_logic := '0';
signal s_ready, s_req, d_valid, d_ack : std_logic;
signal s_hold, d_data : std_logic_vector(DW-1 downto 0);
signal s_state : std_logic_vector(1 downto 0);
signal n_sent, n_stalls, n_received : std_logic_vector(31 downto 0);
-- destination-side observation, owned by the dmon process
signal rxq : byte_arr := (others => (others => '0'));
signal rx_n : integer := 0;
signal d_age : integer := 0;
signal min_lat : integer := 1000000;
signal max_lat : integer := 0;
signal e_dst, c_dst : integer := 0;
-- source-side observation, owned by the smon process
signal n_req_rises : integer := 0;
signal e_src, c_src : integer := 0;
-- the stimulus process's own tallies
signal e_top, c_top : integer := 0;
begin
dut : entity work.usb_cdc_handshake
generic map (DW => DW, SYNC_N => SYNC_N)
port map (
clk_s => clk_s, rst_s_n => rst_s_n,
s_data => s_data, s_valid => s_valid, s_ready => s_ready,
s_req => s_req, s_hold => s_hold, s_state => s_state,
n_sent => n_sent, n_stalls => n_stalls,
clk_d => clk_d, rst_d_n => rst_d_n,
d_data => d_data, d_valid => d_valid, d_ack => d_ack,
n_received => n_received
);
-- Two free-running clocks with a deliberate phase offset, so no edge in
-- one domain ever coincides exactly with an edge in the other -- which
-- would be the one alignment a real chip is guaranteed NOT to have.
clk_s_gen : process
begin
if done then wait; end if;
wait for sp;
clk_s <= not clk_s;
end process;
clk_d_gen : process
variable started : boolean := false;
begin
if not started then
wait for 3 ns;
started := true;
end if;
if done then wait; end if;
wait for dp;
clk_d <= not clk_d;
end process;
-- ==================================================================
-- Destination-domain monitor
-- ==================================================================
dmon : process (clk_d, rst_d_n)
variable e, c : integer := 0;
variable dv_d : std_logic := '0';
variable n : integer := 0;
variable q : byte_arr;
procedure chk (cond : boolean; msg : string) is
begin
c := c + 1;
if not cond then
e := e + 1;
if e <= 25 then
report "FAIL(dst): " & msg severity note;
end if;
end if;
end procedure;
begin
if rst_d_n = '0' then
d_age <= 0;
dv_d := '0';
elsif rising_edge(clk_d) then
if s_req = '0' then d_age <= 0; else d_age <= d_age + 1; end if;
-- ---- d_valid is ONE cycle. Two in a row is a level-triggered
-- ---- capture delivering the same value repeatedly.
chk(not (d_valid = '1' and dv_d = '1'),
"d_valid was asserted for two consecutive destination cycles -- the capture is level-triggered, not edge-triggered");
dv_d := d_valid;
if d_valid = '1' then
-- The index is bounded. A level-triggered capture delivers orders
-- of magnitude more values than a correct design does, and a queue
-- that runs off its end turns a measured kill into a crash.
if n <= 20000 then
q(n) := d_data;
rxq <= q;
end if;
n := n + 1;
rx_n <= n;
if d_age < min_lat then min_lat <= d_age; end if;
if d_age > max_lat then max_lat <= d_age; end if;
-- ---- the synchroniser really is SYNC_N deep ----
chk(d_age >= SYNC_N + 1,
"a value arrived sooner than SYNC_N synchroniser stages can possibly allow -- a stage has been bypassed");
-- Note what is NOT checked here: that req is still high. The
-- monitor observes d_valid one destination cycle after the capture,
-- and a fast source can already have seen ack and completed phases
-- 3 to 5 by then. "One delivery per request" is a statement about
-- totals, so it is checked as one -- see n_req_rises below.
end if;
e_dst <= e;
c_dst <= c;
end if;
end process;
-- ==================================================================
-- Source-domain monitor: the bus that crosses must NOT MOVE
-- ==================================================================
smon : process (clk_s, rst_s_n)
variable e, c : integer := 0;
variable hold_prev : std_logic_vector(DW-1 downto 0) := (others => '0');
variable req_prev : std_logic := '0';
variable rises : integer := 0;
procedure chk (cond : boolean; msg : string) is
begin
c := c + 1;
if not cond then
e := e + 1;
if e <= 25 then
report "FAIL(src): " & msg severity note;
end if;
end if;
end procedure;
begin
if rst_s_n = '0' then
hold_prev := (others => '0');
req_prev := '0';
elsif rising_edge(clk_s) then
-- From the SECOND cycle of req onward. The edge that raises req is
-- the same edge that loads the hold register, so comparing across it
-- would flag the load itself.
if s_req = '1' and req_prev = '1' then
chk(s_hold = hold_prev,
"the bus that crosses the boundary changed while req was asserted -- the destination may be sampling it mid-change, which is the bit-synchroniser bug in its purest form");
end if;
-- ---- and the source must not declare itself ready mid-transfer ----
chk(not (s_ready = '1' and s_req = '1'),
"s_ready was asserted while a transfer was still in flight -- the next value can overwrite the one being captured");
chk((s_ready = '1') = (s_state = ss_code(SS_IDLE)),
"s_ready disagrees with the source state");
if s_req = '1' and req_prev = '0' then
rises := rises + 1;
n_req_rises <= rises;
end if;
hold_prev := s_hold;
req_prev := s_req;
e_src <= e;
c_src <= c;
end if;
end process;
-- ==================================================================
-- Stimulus
-- ==================================================================
stim : process
variable e, c : integer := 0;
variable txq : byte_arr := (others => (others => '0'));
variable tx_n : integer := 0;
variable n_ratios : integer := 0;
variable cmp_n : integer := 0;
variable ln : line;
procedure chk (cond : boolean; msg : string) is
begin
c := c + 1;
if not cond then
e := e + 1;
if e <= 25 then
report "FAIL(top): " & msg severity note;
end if;
end if;
end procedure;
-- VHDL has no xor on INTEGER, so the pattern is built in unsigned --
-- which is the same arithmetic the other two suites do implicitly and
-- produces the identical byte sequence.
function pat (n, mul, xr : integer) return std_logic_vector is
variable a, b : unsigned(DW-1 downto 0);
begin
a := to_unsigned((n * mul) mod 256, DW);
b := to_unsigned(xr mod 256, DW);
return std_logic_vector(a xor b);
end function;
-- Hand one value to the source, wait until the handshake has taken it,
-- and then IMMEDIATELY drive something else onto s_data -- so that a
-- design which forwarded s_data combinationally is caught at once.
procedure send (v : std_logic_vector(DW-1 downto 0)) is
variable guard : integer := 0;
begin
wait until rising_edge(clk_s);
wait for 1 ns;
-- Every wait is BOUNDED. A handshake that has lost a transfer does
-- not produce a wrong answer -- it produces no answer at all, and an
-- unbounded wait turns that into a suite that hangs instead of a
-- suite that fails. The bound is far larger than the slowest legal
-- handshake at the most extreme clock ratio in the sweep.
guard := 0;
while s_ready = '0' and guard < 600 loop
wait until rising_edge(clk_s);
wait for 1 ns;
guard := guard + 1;
end loop;
chk(s_ready = '1',
"the source never became ready -- the handshake is deadlocked, which is what a lost ack looks like from this side");
s_data <= v;
s_valid <= '1';
wait until rising_edge(clk_s);
wait for 1 ns;
guard := 0;
while s_state = ss_code(SS_IDLE) and guard < 600 loop
wait until rising_edge(clk_s);
wait for 1 ns;
guard := guard + 1;
end loop;
chk(s_state /= ss_code(SS_IDLE),
"the source never accepted the value offered to it");
s_valid <= '0';
txq(tx_n) := v;
tx_n := tx_n + 1;
s_data <= not v; -- the value is now HELD, or it is not
end procedure;
procedure drain is
begin
for g in 1 to 400 loop wait until rising_edge(clk_d); end loop;
for g in 1 to 400 loop wait until rising_edge(clk_s); end loop;
end procedure;
-- One ratio: reset nothing, just send a run of values and confirm the
-- queues still agree. State is deliberately NOT cleared between ratios,
-- so a transfer in flight when the period changes is part of the test.
procedure run_ratio (new_sp, new_dp : time; count : integer) is
begin
sp <= new_sp;
dp <= new_dp;
n_ratios := n_ratios + 1;
for g in 1 to count loop
send(pat(tx_n, 37, tx_n / 8));
end loop;
drain;
chk(rx_n = tx_n,
"the number of values received does not equal the number sent at this clock ratio");
chk(n_received = n_sent,
"the design's own source and destination counters disagree");
end procedure;
begin
for g in 1 to 6 loop wait until rising_edge(clk_s); end loop;
rst_s_n <= '1';
for g in 1 to 6 loop wait until rising_edge(clk_d); end loop;
rst_d_n <= '1';
for g in 1 to 4 loop wait until rising_edge(clk_s); end loop;
wait for 1 ns;
-- ---- Phase A: the state after reset ----
chk(s_ready = '1', "the source is not ready after reset");
chk(s_req = '0', "req is asserted after reset");
chk(d_ack = '0', "ack is asserted after reset");
chk(d_valid = '0', "d_valid is asserted after reset");
-- ---- Phase B: the ratio sweep. Every shape of relationship the two
-- ---- clocks can have, including one that changes mid-flight.
run_ratio( 5 ns, 7 ns, 300); -- destination a little slower
run_ratio( 7 ns, 5 ns, 300); -- source a little slower
run_ratio( 5 ns, 31 ns, 300); -- destination MUCH slower
run_ratio(31 ns, 5 ns, 300); -- source MUCH slower
run_ratio( 3 ns, 11 ns, 300); -- co-prime, close
run_ratio(11 ns, 3 ns, 300);
run_ratio( 5 ns, 6 ns, 300); -- nearly equal, never aligned
run_ratio( 6 ns, 5 ns, 300);
run_ratio(13 ns, 13 ns, 300); -- equal periods, offset phase
run_ratio( 2 ns, 29 ns, 200); -- extreme
run_ratio(29 ns, 2 ns, 200);
-- ---- Phase C: the destination period CHANGES while transfers are in
-- ---- flight. A handshake that depended on a fixed ratio anywhere --
-- ---- a fixed-delay req, a timed ack -- comes apart here.
for i in 0 to 599 loop
dp <= (2 + (i mod 17)) * 1 ns;
sp <= (3 + ((i * 7) mod 13)) * 1 ns;
send(pat(tx_n, 91, 16#5a#));
end loop;
drain;
chk(rx_n = tx_n, "a value was lost while the clock ratio was changing");
-- Everything up to here went through send(), so the two queues are
-- element-for-element comparable. Phase D drives the source directly
-- and is checked by value instead.
cmp_n := tx_n;
-- ---- Phase D: the source offers values back to back with no gap, so
-- ---- s_valid is high through entire handshakes and the stall path is
-- ---- exercised rather than merely present.
sp <= 5 ns; dp <= 9 ns;
s_data <= x"A5";
s_valid <= '1';
for g in 1 to 2000 loop wait until rising_edge(clk_s); end loop;
s_valid <= '0';
drain;
chk(unsigned(n_stalls) > 0,
"the source never had to wait, so the stall path was never exercised");
-- ---- Final: every value, in order, unchanged. ----
chk(rx_n >= cmp_n, "fewer values arrived than were sent through send()");
-- EVERY mismatch is counted, not just the first. A design that puts a
-- live bus across the boundary gets one wrong value per transfer, and
-- stopping at the first would report that as a single error -- which
-- says nothing about how pervasive it is.
for i in 0 to cmp_n - 1 loop
chk(rxq(i) = txq(i), "a value arrived corrupted or out of order");
end loop;
-- Phase D held a single value on s_data throughout, so every delivery
-- from there on must be exactly that value -- a design that forwarded
-- a live bus would deliver the inverted copy the driver leaves behind.
for i in cmp_n to minimum(rx_n - 1, 20000) loop
chk(rxq(i) = x"A5",
"a back-to-back transfer delivered a value that was never offered");
end loop;
chk(cmp_n > 3000, "the suite did not send enough values to mean anything");
chk(n_received = n_sent,
"the design's own source and destination counters disagree at the end of the run");
-- The latency is DETERMINISTIC in destination cycles -- SYNC_N stages
-- plus one for the edge detector -- no matter what the clock ratio is.
-- That is the signature of a synchroniser that is really SYNC_N deep,
-- and it is the same at every ratio because it is counted in the
-- destination's own cycles.
chk(min_lat >= SYNC_N + 1,
"the fastest observed delivery was quicker than SYNC_N synchroniser stages allow -- a stage has been bypassed");
chk(max_lat <= SYNC_N + 2,
"a delivery took longer than the synchroniser plus edge detection can account for");
chk(n_ratios = 11, "not every clock ratio was exercised");
-- ---- EXACTLY ONE DELIVERY PER REQUEST. Not "at least one", which a
-- ---- level-triggered capture also satisfies, and not "at most one",
-- ---- which a handshake that loses transfers also satisfies.
chk(rx_n = n_req_rises,
"the number of values delivered does not equal the number of requests raised");
e_top <= e;
c_top <= c;
wait for 1 ns;
write(ln, string'("REACH ratios=") & integer'image(n_ratios) &
" transfers=" & integer'image(cmp_n) &
" received=" & integer'image(rx_n) &
" latency=" & integer'image(min_lat) & ".." & integer'image(max_lat) &
" dest cycles (bound >= " & integer'image(SYNC_N + 1) & ")");
writeline(output, ln);
write(ln, string'("COUNTERS sent=") & integer'image(to_integer(unsigned(n_sent))) &
" received=" & integer'image(to_integer(unsigned(n_received))) &
" requests=" & integer'image(n_req_rises) &
" stalls=" & integer'image(to_integer(unsigned(n_stalls))));
writeline(output, ln);
if (e + e_src + e_dst) = 0 then
write(ln, string'("PASS: 0 errors in ") &
integer'image(c + c_src + c_dst) & " checks");
else
write(ln, string'("FAIL: ") & integer'image(e + e_src + e_dst) &
" errors in " & integer'image(c + c_src + c_dst) & " checks");
end if;
writeline(output, ln);
done <= true;
wait;
end process;
end architecture sim;11. Exhaustive Verification
| Measure | Verilog | SystemVerilog | VHDL |
|---|---|---|---|
| Clock ratios swept | 11 | 11 | 11 |
| Values sent through the checker | 3700 | 3700 | 3700 |
| Total transfers (incl. back-to-back phase) | 3828 | 3828 | 3827 |
| Checks executed | 302907 | 302907 | 302873 |
design's n_sent | 3828 | 3828 | 3827 |
design's n_received | 3828 | 3828 | 3827 |
| requests raised | 3828 | 3828 | 3827 |
| source stall cycles | 1872 | 1872 | 1873 |
| delivery latency (destination cycles) | 3 .. 3 | 3 .. 3 | 3 .. 3 |
| values lost, duplicated or corrupted | 0 | 0 | 0 |
| Result | PASS | PASS | PASS |
The latency row is the interesting one. Three destination cycles, minimum and maximum, at every ratio from 2:29 to 29:2 and through a phase where both periods changed on every transfer.
That is not a coincidence, and it is the signature of a correct synchroniser: two flops plus one for edge detection, counted in the destination's own cycles, which is the only frame in which the number is meaningful at all. Measured in source cycles the same transfer takes anywhere from 2 to 90.
n_sent, n_received and the count of requests raised agree exactly — three counters in two different clock domains, which is the property the whole block exists to provide.
12. Mutation Testing
| # | Mutation | Verilog | SysVer | VHDL |
|---|---|---|---|---|
| C3 | the capture follows req's LEVEL, not its rising edge | 55581 | 55569 | 55563 |
| C7 | s_ready is asserted one phase early | 10769 | 10768 | 10770 |
| C6 | ack is released on a timer, not when req falls | 6826 | 6815 | 6815 |
| C4 | req is dropped on a timer, not when ack arrives | 5820 | 5808 | 5849 |
| C5 | three phases instead of four | 5198 | 5186 | 5515 |
| C1 | the synchroniser is tapped one stage early | 3924 | 3924 | 3924 |
| C2 | the data is not held — a live bus crosses the boundary | 3700 | 3700 | 3700 |
| — | unmutated baseline | 0 | 0 | 0 |
All seven die in all three languages, all counts distinct, and the columns agree to within 6% — the suite is fully deterministic, so where they differ it is only because the VHDL scheduler resolves one back-to-back transfer differently in phase D.
C3 is by far the largest. A level-triggered capture re-delivers the same value on every destination cycle for as long as req is high, so one transfer becomes dozens. It is caught three separate ways — consecutive d_valid, the delivery count, and the end-to-end queue — which is what you want for the failure that produces the most wrong data.
C2 scores exactly 3700: one wrong value per transfer, every transfer, no exceptions. That exactness is worth pausing on. It is not a flaky, load-dependent, ratio-dependent bug. Putting a live bus across a clock boundary is wrong every single time, and the only reason it ever appears to work in a lab is that the two clocks happened to be related that afternoon.
C1 is the one to be careful about, and it is discussed below.
13. Debugging Walkthrough: The Device That Fails Once a Week
The report. A USB device fails roughly once a week per unit in the field, and never in the lab. When it fails it enumerates with the wrong number of endpoints and has to be unplugged. No error is logged by anything.
Step 1 — reproduce it. Impossible. Thousands of hours of soak testing, zero failures. This is the first and most important datum: a failure that will not reproduce in a stable environment is usually a failure that depends on an unstable one.
Step 2 — what is unstable in the field and stable in the lab? Temperature, voltage, and — crucially — the relationship between two clock sources. In the lab both were derived from one bench oscillator. In the field the PHY's clock comes from the recovered bus clock and the controller's from a local crystal, and the phase between them drifts continuously.
Step 3 — so look for clock crossings. Run a CDC lint. It reports 41 crossings. Forty of them are single bits with two-flop synchronisers. One is a 4-bit endpoint-count value with a two-flop synchroniser on each bit.
Step 4 — is that actually wrong? The value is written once at configuration time and never changes afterwards. The reasoning in the original review was "it is static, so there is nothing to synchronise."
Step 5 — but it changes once. From 0000 to 0100 at SET_CONFIGURATION. Two bits... no, one bit. From 0011 to 0100 on a reconfiguration: three bits at once. If the destination samples during that transition it can see 0111 — seven endpoints where there are four.
Step 6 — why once a week. The window is one destination clock period per reconfiguration, and the phase has to land inside it. That is a probability, not a condition, and a probability is exactly what produces "once a week per unit, never in the lab".
Step 7 — the fix. A four-phase handshake. The value is static, so the handshake costs nothing that matters.
14. UVM and Assertions
14.1 The sequences
class usb_cdc_item extends uvm_sequence_item;
`uvm_object_utils(usb_cdc_item)
rand bit [7:0] data;
rand int unsigned gap; // source cycles of idle before offering it
constraint c_gap { gap dist {0 := 5, [1:3] := 3, [4:20] := 1}; }
function new(string name = "usb_cdc_item"); super.new(name); endfunction
endclass
// THE sequence for this chapter, and it is not a data sequence at all --
// it is a CLOCK sequence. The block's claim is that it works at any ratio,
// so the ratio is the thing to randomise. A suite that fixes the ratio and
// randomises only the payload is testing one point of the space it cares
// about.
class ratio_sweep_seq extends uvm_sequence #(usb_cdc_item);
`uvm_object_utils(ratio_sweep_seq)
// Set by the test; the driver applies them to the clock generator.
rand int unsigned src_period;
rand int unsigned dst_period;
constraint c_periods {
src_period inside {[2:31]};
dst_period inside {[2:31]};
}
function new(string name = "ratio_sweep_seq"); super.new(name); endfunction
task body();
repeat (40) begin
if (!this.randomize())
`uvm_error("RAND", "period randomize failed")
uvm_config_db #(int unsigned)::set(null, "*", "src_period", src_period);
uvm_config_db #(int unsigned)::set(null, "*", "dst_period", dst_period);
repeat (100) begin
usb_cdc_item it = usb_cdc_item::type_id::create("it");
start_item(it);
if (!it.randomize())
`uvm_error("RAND", "item randomize failed")
finish_item(it);
end
end
endtask
endclass
// Back to back with no gap at all, so s_valid is high through entire
// handshakes and the source spends most of its time stalled. This is the
// sequence that exercises the cost of the handshake rather than its
// correctness, and it is the one that catches a source declaring itself
// ready a phase early.
class back_to_back_seq extends uvm_sequence #(usb_cdc_item);
`uvm_object_utils(back_to_back_seq)
function new(string name = "back_to_back_seq"); super.new(name); endfunction
task body();
repeat (2000) begin
usb_cdc_item it = usb_cdc_item::type_id::create("it");
start_item(it);
if (!it.randomize() with { gap == 0; })
`uvm_error("RAND", "back-to-back randomize failed")
finish_item(it);
end
endtask
endclass14.2 The scoreboard: two domains, two monitors, one queue
// The source and destination monitors run in DIFFERENT clock domains and
// write into the same scoreboard. That is the one place in the whole
// environment where two domains meet, and it is deliberately the only one.
class usb_cdc_scoreboard extends uvm_scoreboard;
`uvm_component_utils(usb_cdc_scoreboard)
uvm_analysis_imp_src #(usb_cdc_src_item, usb_cdc_scoreboard) src_ap;
uvm_analysis_imp_dst #(usb_cdc_dst_item, usb_cdc_scoreboard) dst_ap;
localparam int SYNC_N = 2;
byte unsigned expected[$];
int unsigned n_sent, n_received;
int unsigned min_lat = 32'hFFFF_FFFF, max_lat;
bit [7:0] hold_at_req;
bit req_active;
function new(string name, uvm_component parent);
super.new(name, parent);
src_ap = new("src_ap", this);
dst_ap = new("dst_ap", this);
endfunction
// ---- source side ----
function void write_src(usb_cdc_src_item t);
// THE property that makes an unsynchronised data path safe.
if (t.s_req && req_active && (t.s_hold !== hold_at_req))
`uvm_error("UNSTABLE",
$sformatf("the bus that crosses the boundary changed from 0x%02h to 0x%02h while req was asserted -- the destination may be sampling it mid-change",
hold_at_req, t.s_hold))
if (t.s_req && !req_active) begin
req_active = 1;
hold_at_req = t.s_hold;
expected.push_back(t.s_hold);
n_sent++;
end
if (!t.s_req) req_active = 0;
// A source that declares itself ready mid-transfer can overwrite the
// value the destination is still reading.
if (t.s_ready && t.s_req)
`uvm_error("EARLY_READY",
"s_ready was asserted while a transfer was still in flight")
endfunction
// ---- destination side ----
function void write_dst(usb_cdc_dst_item t);
if (!t.d_valid) return;
// ---- EXACTLY ONE delivery, of EXACTLY the value that was sent. ----
if (expected.size() == 0)
`uvm_error("PHANTOM",
$sformatf("0x%02h was delivered but nothing was sent -- the capture is following req's level, not its edge",
t.d_data))
else begin
byte unsigned want = expected.pop_front();
if (t.d_data !== want)
`uvm_error("DATA",
$sformatf("received 0x%02h, expected 0x%02h -- a live bus crossed the boundary instead of a held one",
t.d_data, want))
end
n_received++;
// ---- The synchroniser really is SYNC_N deep. See the warning in the
// ---- chapter about exactly what this does and does not prove.
if (t.latency_in_dst_cycles < SYNC_N + 1)
`uvm_error("TOO_FAST",
$sformatf("delivered after %0d destination cycles; %0d synchroniser stages plus edge detection cannot produce fewer than %0d",
t.latency_in_dst_cycles, SYNC_N, SYNC_N + 1))
if (t.latency_in_dst_cycles < min_lat) min_lat = t.latency_in_dst_cycles;
if (t.latency_in_dst_cycles > max_lat) max_lat = t.latency_in_dst_cycles;
endfunction
function void report_phase(uvm_phase phase);
`uvm_info("SB", $sformatf("sent=%0d received=%0d latency=%0d..%0d dest cycles",
n_sent, n_received, min_lat, max_lat), UVM_LOW)
// Nothing may be left in flight at the end of the run.
if (expected.size() != 0)
`uvm_error("LOST",
$sformatf("%0d values were sent and never delivered", expected.size()))
if (n_sent != n_received)
`uvm_error("COUNT",
$sformatf("sent %0d, received %0d", n_sent, n_received))
// The latency must be CONSTANT in destination cycles, whatever the
// ratio. A spread here means the handshake's timing depends on the
// clock relationship, which is the thing it exists not to do.
if (min_lat != max_lat)
`uvm_error("JITTER",
$sformatf("delivery latency varied from %0d to %0d destination cycles -- the handshake is sensitive to the clock ratio",
min_lat, max_lat))
endfunction
endclass14.3 Assertions, and which domain each one belongs to
// Two assertion modules, not one. An assertion has a clock, and an
// assertion clocked in the wrong domain is a race dressed up as a check.
module usb_cdc_src_sva
import usb_cdc_pkg::*;
#(
parameter int DW = 8
) (
input logic clk_s,
input logic rst_s_n,
input logic s_valid,
input logic s_ready,
input logic s_req,
input logic [DW-1:0] s_hold,
input src_state_e s_state
);
default clocking cb @(posedge clk_s); endclocking
default disable iff (!rst_s_n);
// ---- 1. THE property. The bus does not move while req is high. ----
//
// Written from the SECOND cycle of req, because the edge that
// raises req is the same edge that legitimately loads the hold
// register. See the chapter on why this cannot catch a purely
// combinational data path.
property p_hold_is_stable;
(s_req && $past(s_req)) |-> $stable(s_hold);
endproperty
a_hold_is_stable : assert property (p_hold_is_stable)
else $error("the bus crossing the boundary moved while req was asserted");
// ---- 2. The source is never ready mid-transfer. ----
property p_not_ready_in_flight;
s_req |-> !s_ready;
endproperty
a_not_ready_in_flight : assert property (p_not_ready_in_flight)
else $error("s_ready asserted while a transfer was in flight -- the next value can overwrite the one being captured");
// ---- 3. FOUR phases, in order. req goes up, and comes down only from
// ---- SS_REQ; the machine then waits in SS_DROP.
property p_req_falls_only_from_req_state;
$fell(s_req) |-> $past(s_state == SS_REQ);
endproperty
a_req_falls_only_from_req_state :
assert property (p_req_falls_only_from_req_state);
property p_drop_precedes_idle;
(s_state == SS_IDLE) && $past(s_state != SS_IDLE)
|-> $past(s_state == SS_DROP);
endproperty
a_drop_precedes_idle : assert property (p_drop_precedes_idle)
else $error("the source returned to IDLE without passing through SS_DROP -- this is the three-phase handshake, and it loses the next transfer");
// ---- 4. req rises only from IDLE, and only when a value was offered. ----
property p_req_rises_from_idle;
$rose(s_req) |-> $past(s_state == SS_IDLE) && $past(s_valid);
endproperty
a_req_rises_from_idle : assert property (p_req_rises_from_idle);
c_stalled : cover property ((s_valid && !s_ready));
c_drop : cover property ((s_state == SS_DROP));
endmodule
module usb_cdc_dst_sva #(
parameter int DW = 8
) (
input logic clk_d,
input logic rst_d_n,
input logic s_req, // asynchronous here, and only ever
// referenced through its synchroniser
input logic req_seen,
input logic d_valid,
input logic [DW-1:0] d_data,
input logic d_ack
);
default clocking cb @(posedge clk_d); endclocking
default disable iff (!rst_d_n);
// ---- 5. d_valid is ONE cycle. Never two. ----
property p_valid_is_a_pulse;
d_valid |=> !d_valid;
endproperty
a_valid_is_a_pulse : assert property (p_valid_is_a_pulse)
else $error("d_valid was asserted for two consecutive cycles -- the capture is level-triggered and is re-delivering the same value");
// ---- 6. A delivery happens only on the synchronised RISING edge. ----
property p_valid_needs_a_rise;
d_valid |-> $past(req_seen) && !$past(req_seen, 2);
endproperty
a_valid_needs_a_rise : assert property (p_valid_needs_a_rise);
// ---- 7. ack rises with the delivery and falls only after req does. ----
property p_ack_rises_with_delivery;
$rose(d_ack) |-> d_valid;
endproperty
a_ack_rises_with_delivery : assert property (p_ack_rises_with_delivery);
property p_ack_falls_after_req;
$fell(d_ack) |-> !req_seen;
endproperty
a_ack_falls_after_req : assert property (p_ack_falls_after_req)
else $error("ack was released before req fell -- the source may miss it entirely and wait for ever");
// ---- 8. The data is stable for the cycle it is delivered in. ----
property p_data_stable_at_valid;
d_valid |=> $stable(d_data);
endproperty
a_data_stable_at_valid : assert property (p_data_stable_at_valid);
c_delivered : cover property (d_valid);
endmodule
bind usb_cdc_handshake usb_cdc_src_sva #(.DW(DW)) u_src_sva (.*);
bind usb_cdc_handshake usb_cdc_dst_sva #(.DW(DW)) u_dst_sva (.*);15. Common Misconceptions
"Put a two-flop synchroniser on each bit." Each bit resolves independently. The destination can sample a value the source never held.
"Gray-code the bus." Gray codes work for counters, where single-bit transitions are guaranteed by construction. Arbitrary data changes arbitrarily many bits.
"The value is static, so it does not need synchronising." It changes once, and the failure window is the transition. Once a week per unit is a probability, not a condition.
"Three phases is enough — drop req on ack and go." ack is still high. The next transfer sees the previous ack, completes in zero time, and the destination never captures.
"Ack can be released on a timer." Then the source's synchroniser can miss the pulse entirely and wait for ever.
"A handshake is too slow." For a control value it costs nothing. For streaming data you wanted an asynchronous FIFO, which is this idea applied to Gray-coded pointers.
"The regression is green, so the CDC is fine." RTL simulation does not model metastability. A missing synchroniser stage is invisible to every simulation at every seed. Lint and STA are the tools; simulation is not.
16. Exercises
1. Work out how many distinct values a destination can sample when a 6-bit bus changes from 011111 to 100000 through per-bit synchronisers, and name two of them that would be actively harmful as an endpoint number.
2. Apply C5 (three phases) and trace, cycle by cycle, what happens to the second transfer. Then say why n_sent still increments.
3. C2 scores exactly 3700 — one per transfer. Explain why the source-side stability monitor does not catch it, and design a source-side check that would.
4. The delivery latency is 3 destination cycles at every ratio. Derive that number from the design, then work out the latency in source cycles at 2:29 and at 29:2.
5. C1 is killed only by its latency signature. Write down what a CDC lint would report for it, and explain why that report is stronger evidence than the mutation score.
6. Replace the handshake with an asynchronous FIFO for the same data. What changes, what stays, and why do the pointers get a Gray code when the data does not?
17. Summary
| Idea | Why it matters |
|---|---|
| A multi-bit value cannot be bit-synchronised | each bit resolves independently; 16 possible samples |
| A Gray code fixes a counter, not a bus | single-bit transitions must be by construction |
| Synchronise one bit, not the data | the data crosses untouched, and is never sampled moving |
| Four phases, not three | ack is still high when three-phase declares ready |
| The data is latched before req rises | a live bus across a boundary is the original bug |
| The capture is on the edge, not the level | a level re-delivers the same value dozens of times |
| Latency is constant in destination cycles | the only frame in which the number means anything |
| "It is static" is the reasoning to distrust | it changes once, and the window is the transition |
| A handshake is slow | control values, not streams; use an async FIFO for those |
| Simulation cannot see metastability | a missing flop is invisible; lint and STA are the tools |
| 11 clock ratios, 3828 transfers, 0 lost | 7 mutations, all killed in 3 languages |
Tooling
| Step | Command |
|---|---|
| Verilog-2005 | iverilog -g2005 -o ch_v.out ch_v.v ch_v_tb.v && ./ch_v.out |
| SystemVerilog | iverilog -g2012 -o ch_sv.out ch_sv.sv ch_sv_tb.sv && ./ch_sv.out |
| VHDL-2008 analyse | nvc --std=2008 -a ch_vhdl.vhd ch_vhdl_tb.vhd |
| VHDL-2008 elaborate | nvc --std=2008 -e tb_ch_vhdl |
| VHDL-2008 run | nvc --std=2008 -r tb_ch_vhdl |
| One mutation | iverilog -g2005 -DMUT_C2 -o mm ch_v_mut.v ch_v_tb.v && ./mm |
| And the one that matters | a CDC lint, plus STA with the crossing constrained |
All three implementations pass with 0 errors: eleven clock ratios plus a phase in which both periods change on every transfer, ~3828 values delivered exactly once each in order and unchanged, and a delivery latency of exactly three destination cycles throughout.
That completes Module 23 — USB RTL Design. Five blocks: the address decode that decides whether a packet is ours, the arbiter that shares one port with bounded waiting, the FIFO whose watermark has to pay for its own backpressure latency, the packet recogniser that resynchronises after every error, and the handshake that carries a value between two clocks that have never met.
Each was written three times, verified exhaustively over a named domain, and mutated seven ways with every mutation killed in every language. What is worth carrying forward is not the five designs but the four habits underneath them: make the property observable, reach every state for real, measure the bound rather than assume it, and be explicit about what your verification cannot see.
Continue learning
Related tutorials
- Related topic
Hub Architecture
A hub is three devices in one package, and its repeater is deliberately asymmetric: downstream is a broadcast, upstream is a select of exactly one — and two talkers connects neither.
- Related topic
Port Management
Three uncoordinated sources move a hub port and collide in the same cycle — and the host, whose information is stale by construction, is the one that loses.
- Related topic
Downstream Device Discovery
A hub cannot interrupt the host, so every port event waits to be asked for — and the window between the poll and the acknowledgement is where devices are silently lost.
- Related topic
Cascading Hubs
The seven-tier rule is a timing constraint wearing a topology costume — and the limit applies to hubs, not to devices, which is one comparison operator almost everyone gets wrong.
Standards & specifications
- Governing standard
- USB-IF (Universal Serial Bus Specification)(opens USB Implementers Forum (USB-IF) in a new tab)
Defines the USB bus — its electrical signalling, connectors, packet and transaction model, device framework and the descriptors a device must expose — together with the device-class specifications layered on it. It does not define host-controller register interfaces (xHCI and EHCI are separate documents) nor any operating system's driver architecture.
This page also covers RTL structure, verification approach and debugging technique. Those are engineering practice built on the standard, not requirements the standard itself imposes.
Where this fits
Part of the USB curriculum.
