USB · Module 15
Interrupt Latency Requirements
bInterval is a request, not a contract, and it bounds one term of seven. A latency monitor in three HDLs, and the measurement showing 3000 randomised steps could not reach two of its own boundaries.
Chapter 15.1 built the schedule, and 15.2 and 15.3 built devices that took it as given: a poll arrives, eventually, within the interval the endpoint asked for.
This chapter asks what that interval actually buys, and the honest answer is much less than the number suggests. An endpoint that declares a 1 ms interval does not thereby give the user a 1 ms response. It bounds one link in a chain that starts at a sensor and ends at a pixel, and several of the other links are larger, less predictable, and outside the device's control entirely.
The interval bounds one link. The user feels the chain.
1. What bInterval Actually Encodes
The endpoint descriptor's bInterval field is one byte, and what that byte means depends on the speed of the device — not as a footnote, but as two genuinely different encodings.
| Speed | Unit | Encoding | Range | Resulting interval |
|---|---|---|---|---|
| Low speed | frame (1 ms) | linear — the value is the count | 10 – 255 | 10 ms – 255 ms |
| Full speed | frame (1 ms) | linear | 1 – 255 | 1 ms – 255 ms |
| High speed | microframe (125 µs) | logarithmic — 2^(bInterval−1) | 1 – 16 | 125 µs – 4.096 s |
| SuperSpeed | microframe (125 µs) | logarithmic | 1 – 16 | 125 µs – 4.096 s |
The kernel states the split plainly, in the documentation of the URB field that carries it:
* @interval: Specifies the polling interval for interrupt or isochronous
* transfers. The units are frames (milliseconds) for full and low
* speed devices, and microframes (1/8 millisecond) for highspeed
* and SuperSpeed devices.and is explicit that the descriptor's encoding is the driver's problem:
* (Note that for isochronous
* endpoints, as well as high speed interrupt endpoints, the encoding of
* the transfer interval in the endpoint descriptor is logarithmic.
* Device drivers must convert that value to linear units themselves.)The root hub's own descriptors are a clean worked example of both encodings, because the kernel builds one for each speed and annotates each with the interval it intends:
/* full-speed root hub, interrupt endpoint */
0xff /* __u8 ep_bInterval; (255ms -- usb 2.0 spec) */
/* high-speed root hub, interrupt endpoint */
0x0c /* __u8 ep_bInterval; (256ms -- usb 2.0 spec) */Check the arithmetic against the comments, because it is the fastest way to internalise the difference:
- Full speed, linear:
0xff= 255 frames × 1 ms = 255 ms. ✅ matches. - High speed, logarithmic:
0x0c= 12, so 2^(12−1) = 2048 microframes × 125 µs = 256 000 µs = 256 ms. ✅ matches.
The same intended interval — roughly a quarter of a second — is written
255at full speed and12at high speed. A device that copies a full-speed descriptor'sbIntervalinto a high-speed one does not get a slightly different interval.0xffat high speed is out of range; a value of 12 at full speed asks for 12 ms instead of 256 ms, a factor of twenty-one the wrong way.
2. The Interval Is a Request, Not a Contract
This is the chapter's central correction, and the kernel says it in one sentence:
* The polling interval may be more frequent than requested.
* For example, some controllers have a maximum interval of 32 milliseconds,
* while others support intervals of up to 1024 milliseconds.and then, about what happens after submission:
* After the URB has been submitted, the interval
* field reflects how the transfer was actually scheduled.Three separate facts are packed into that.
The host may poll more often than asked. bInterval is an upper bound on the gap, not a period. A device that asks for 8 ms may be polled every 8 ms, or every 4, or every 1. Nothing in the protocol promises the device a minimum gap, which means device hardware must not assume one — Chapter 15.1 §4 built the scheduler on exactly this assumption and it is the reason it counts frames rather than trusting a period.
Host controllers have hard limits that silently clamp. A controller with a 32 ms maximum cannot honour a request for 255 ms. It does not fail; it schedules 32 ms. A device asking for a long interval to save power may be polled eight times more often than it planned for.
The granted interval is readable, and it is a different number from the requested one. The interval field is annotated (modify) in the URB structure — the host writes back what it actually scheduled. A driver that wants to know the real polling rate must read it after submission, not compute it from the descriptor.
3. The Latency Budget, Term by Term
Here is the chain from a physical event to a visible response, with the polling interval in its actual place — one term among seven.
| # | Term | Typical magnitude | Bounded by | Who controls it |
|---|---|---|---|---|
| 1 | Sensor sampling period | 1 ms at 1 kHz; 8 ms at 125 Hz | the sensor's own clock | device |
| 2 | Device processing and arming the endpoint | µs | device firmware and RTL | device |
| 3 | Wait for the next poll | 0 to bInterval | bInterval | host schedule |
| 4 | The transaction on the wire | a few µs at high speed | packet size and speed | bus |
| 5 | Controller → driver completion | µs to ms | interrupt coalescing, IRQ latency | host OS |
| 6 | Driver → application | µs to tens of ms | scheduler, wakeup, contention | host OS |
| 7 | Application → visible output | one display frame: 16.7 ms at 60 Hz | the display pipeline | application |
Read the magnitudes, not the row count. Term 3 — the one bInterval bounds, and the only one most engineers think about — is often not the largest term. At 60 Hz, term 7 alone is 16.7 ms, more than sixteen times a 1 ms polling interval, and it is a hard floor that no amount of polling can move.
Two structural observations follow, and they are what make this a budget rather than a list.
Terms 1 and 3 do not add naively — they beat against each other. A 1 kHz sensor and a 1 ms poll are not synchronised, so the delay from a sample to the poll that carries it varies across the full interval. The worst case is close to the sum, the average is closer to the sum of the halves, and the variation itself is the problem for anything that cares about consistency rather than mean latency. This is the same phase relationship Chapter 15.1 §5 spread deliberately across endpoints.
Only terms 1–3 are bounded at all. Terms 5 and 6 are scheduling latencies on a general-purpose operating system, which has no upper bound under load. A device can guarantee its half of the chain and still be part of a system that misses a deadline, because the unbounded terms are on the other side.
Reducing
bIntervalfrom 8 ms to 1 ms improves one bounded term by 7 ms. If term 6 occasionally costs 20 ms under load, the user's experience is dominated by a term the device cannot see, and the eightfold increase in reserved bandwidth bought very little.
4. Why 1000 Hz Mice Are Not Eight Times Better
The budget explains a claim the peripheral market makes constantly and mostly wrongly.
A "1000 Hz" mouse declares a 125 µs high-speed interval (bInterval = 1) against a "125 Hz" mouse's 8 ms. The marketing difference is 7.875 ms in term 3. Against a 16.7 ms display frame in term 7 and a variable term 6, that is real but small — and it is not the eightfold improvement the numbers imply, because the other terms did not change.
What the higher rate genuinely does buy is worth stating precisely, because it is not nothing:
Lower variance, not just lower mean. The poll-wait term varies between 0 and bInterval; shrinking the interval shrinks the spread. For a human closing a control loop by hand, consistency is more perceptible than latency.
Less accumulation per report. Chapter 15.3 §4 showed the accumulator saturates at ±127. At 8 ms intervals a fast movement can genuinely exceed that and be clipped; at 125 µs it essentially cannot. The faster poll removes a source of non-linearity, which is a different benefit from latency and arguably the more important one.
And what it costs is term 3's bandwidth reservation, multiplied by eight, taken from the same capped budget every other periodic endpoint draws on.
5. Measuring It in Hardware: the Oldest-Item Rule
A device that must prove it meets a latency requirement needs to measure it, and the measurement has one trap that matters more than all the rest.
When data is already waiting and more data arrives, which item's age is the latency?
The answer is the oldest. The item that has been waiting longest is the one whose deadline is closest, and it is the one the user is actually waiting on. A monitor that restarts its timer whenever new data arrives reports the age of the newest item — and under load, when data arrives constantly, the newest item is always young.
6. The Hardware, Before Any Language
State retained: a pending flag, an age counter, the last measured latency, a running maximum, and a sticky deadline-miss flag.
On reset or bus reset: everything clears. A bus reset ends the relationship with the host, and statistics gathered for one host are not statistics for the next.
On a delivery (report_taken while pending): the age before this cycle's update is the measurement. It updates last_latency, updates max_latency if it is greater, and sets the sticky miss flag if it exceeds the deadline.
On an arrival with no item pending: pending sets and the age starts at zero.
On an arrival while an item is already pending: nothing happens. This is §5's rule, and it is expressed as the absence of an action, which is why it is easy to get wrong — the correct code has no line in it.
On a delivery and an arrival in the same cycle: the delivery measures the item it delivered, and the new arrival starts its own measurement at zero. Structurally the same collision as Chapter 15.3 §5.
On a frame tick while pending and none of the above: the age increments, saturating rather than wrapping.
Saturation matters more here than it did in 15.3. A wrapped mouse delta produces a visible glitch; a wrapped age reports a small latency for a very late item — the monitor actively asserting that a badly missed deadline was met.
7. Verilog
The RTL contract
- What it models: an in-hardware latency monitor for one interrupt endpoint.
- Why it exists: because §3's term 3 is the device's to prove, and proving it requires measuring the age of the oldest undelivered item (§5).
- Inputs:
frame_tick,event_valid(data became ready),report_taken(the host actually took it),bus_reset. - State retained:
pend_r,age_r,last_r,max_r,miss_r. - Outputs:
pending,age,last_latency,max_latency,deadline_miss. - Hardware implied: one flag flop, one saturating counter, two comparators, two holding registers, one sticky flop.
- Reset: asynchronous active-low
rst_n;bus_resetsynchronous and equivalent. - Assumptions:
report_takenmeans delivered, not attempted — a NAKed poll must not assert it, or the monitor measures the wrong interval and reports latencies that never happened. - Omissions: no CDC to a register block, no clear-on-read of the statistics, one endpoint only.
- What DV should verify: that the measurement is the age of the oldest item under every arrival pattern; that the deadline comparison is exclusive at the boundary; that the age saturates; that
maxis a maximum anddeadline_missis sticky.
// usb_latency_monitor -- measures the age of the OLDEST undelivered event.
//
// The subtlety this module exists for: when data is already waiting and more
// data arrives, the latency that matters is still the age of the FIRST item,
// because that is the one the user has been waiting on. A monitor that
// restarts its timer on each new event reports the age of the NEWEST item and
// systematically under-reports -- exactly the direction that hides a problem.
module usb_latency_monitor #(
parameter integer AGE_W = 8,
parameter integer DEADLINE = 10 // frames; a miss is age > DEADLINE
) (
input wire clk,
input wire rst_n,
input wire bus_reset,
input wire frame_tick,
input wire event_valid, // data became ready in the endpoint
input wire report_taken, // the host actually took it
output wire pending,
output wire [AGE_W-1:0] age,
output wire [AGE_W-1:0] last_latency,
output wire [AGE_W-1:0] max_latency,
output wire deadline_miss // sticky
);
localparam [AGE_W-1:0] AGE_MAX = {AGE_W{1'b1}};
reg pend_r;
reg [AGE_W-1:0] age_r, last_r, max_r;
reg miss_r;
assign pending = pend_r;
assign age = age_r;
assign last_latency = last_r;
assign max_latency = max_r;
assign deadline_miss= miss_r;
// The measured latency of the delivery happening this cycle: the age
// BEFORE this cycle's update, i.e. frames fully elapsed since arrival.
wire [AGE_W-1:0] measured = age_r;
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
pend_r <= 1'b0; age_r <= {AGE_W{1'b0}};
last_r <= {AGE_W{1'b0}}; max_r <= {AGE_W{1'b0}}; miss_r <= 1'b0;
end else if (bus_reset) begin
// A bus reset ends the relationship: the pending item and every
// statistic about the previous host are discarded.
pend_r <= 1'b0; age_r <= {AGE_W{1'b0}};
last_r <= {AGE_W{1'b0}}; max_r <= {AGE_W{1'b0}}; miss_r <= 1'b0;
end else begin
// 1. DELIVERY. Record the measurement before anything else changes.
if (report_taken && pend_r) begin
last_r <= measured;
if (measured > max_r) max_r <= measured;
if (measured > DEADLINE[AGE_W-1:0]) miss_r <= 1'b1;
end
// 2. PENDING and AGE.
if (report_taken && event_valid) begin
// The collision: the delivery empties the endpoint and the new event
// immediately refills it. The new item's age starts at zero NOW.
pend_r <= 1'b1;
age_r <= {AGE_W{1'b0}};
end else if (report_taken) begin
pend_r <= 1'b0;
age_r <= {AGE_W{1'b0}};
end else if (event_valid && !pend_r) begin
// A new oldest item. Only the FIRST event starts the clock.
pend_r <= 1'b1;
age_r <= {AGE_W{1'b0}};
end else if (frame_tick && pend_r) begin
// Saturate. An age that wraps reports a small latency for a very
// late item -- the worst possible direction for a monitor to fail.
if (age_r != AGE_MAX) age_r <= age_r + {{(AGE_W-1){1'b0}}, 1'b1};
end
end
end
endmoduleThe else if chain's order is the specification. Collision first, then delivery, then arrival, then tick. Reordering arrival ahead of delivery would make a simultaneous pair measure the wrong item; reordering tick ahead of arrival would charge a newly arrived item for a frame it did not wait. Exercise 6 asks for the other orderings to be worked out, because the chain reads as arbitrary until they are.
And note the branch that is not there. There is no arm for arrival while already pending — that case falls through every condition and changes nothing, which is precisely §5's rule. The correct code expresses it by omission, and that is exactly why mutation L1, which adds a plausible-looking arm, is the one to measure.
8. SystemVerilog
Same hardware, with two additions that pay for themselves: an elaboration-time check, and a named event type instead of an inferred chain.
module usb_latency_monitor_sv #(
parameter int unsigned AGE_W = 8,
parameter int unsigned DEADLINE = 10
) (
input logic clk,
input logic rst_n,
input logic bus_reset,
input logic frame_tick,
input logic event_valid,
input logic report_taken,
output logic pending,
output logic [AGE_W-1:0] age,
output logic [AGE_W-1:0] last_latency,
output logic [AGE_W-1:0] max_latency,
output logic deadline_miss
);
localparam logic [AGE_W-1:0] AGE_MAX = '1;
// An elaboration-time check: a deadline the age counter cannot represent
// would make deadline_miss permanently false, which is the worst way for a
// monitor to be wrong -- silently reporting that nothing is ever late.
initial begin
if (DEADLINE >= (1 << AGE_W))
$fatal(1, "DEADLINE=%0d is unrepresentable in AGE_W=%0d bits",
DEADLINE, AGE_W);
end
// What this cycle does, named rather than inferred from an if/else chain.
typedef enum logic [1:0] { EV_IDLE, EV_ARRIVE, EV_DELIVER, EV_BOTH } ev_e;
ev_e ev;
always_comb begin
if (report_taken && event_valid) ev = EV_BOTH;
else if (report_taken) ev = EV_DELIVER;
else if (event_valid && !pending) ev = EV_ARRIVE;
else ev = EV_IDLE;
end
wire [AGE_W-1:0] measured = age;
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n || bus_reset) begin
pending <= 1'b0; age <= '0;
last_latency <= '0; max_latency <= '0; deadline_miss <= 1'b0;
end else begin
if (report_taken && pending) begin
last_latency <= measured;
if (measured > max_latency) max_latency <= measured;
if (measured > AGE_W'(DEADLINE)) deadline_miss <= 1'b1;
end
unique case (ev)
EV_BOTH: begin pending <= 1'b1; age <= '0; end
EV_DELIVER: begin pending <= 1'b0; age <= '0; end
EV_ARRIVE: begin pending <= 1'b1; age <= '0; end
EV_IDLE: if (frame_tick && pending && age != AGE_MAX)
age <= age + 1'b1;
endcase
end
end
endmoduleThe ev_e enumeration is the real improvement, and not for readability. The Verilog chain encodes four mutually exclusive cases in a structure that does not say they are exclusive; naming them makes the exclusivity checkable, and it puts §5's "arrival while pending is not an arrival" into the definition of EV_ARRIVE rather than into the middle of a sequential chain.
$fatal catches a parameterisation that would silently disable the monitor. With AGE_W = 8 and DEADLINE = 300, the comparison measured > 300 can never be true, the miss flag never sets, and the module reports perfect compliance forever. That is not a subtle degradation — it is a monitor that lies — and it is a configuration error a designer can make in one keystroke. Catching it at elaboration costs four lines.
9. VHDL
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity usb_latency_monitor_vhdl is
generic (
AGE_W : positive := 8;
DEADLINE : natural := 10
);
port (
clk : in std_logic;
rst_n : in std_logic;
bus_reset : in std_logic;
frame_tick : in std_logic;
event_valid : in std_logic;
report_taken : in std_logic;
pending : out std_logic;
age : out unsigned(AGE_W-1 downto 0);
last_latency : out unsigned(AGE_W-1 downto 0);
max_latency : out unsigned(AGE_W-1 downto 0);
deadline_miss : out std_logic
);
end entity;
architecture rtl of usb_latency_monitor_vhdl is
constant AGE_MAX : unsigned(AGE_W-1 downto 0) := (others => '1');
constant DEADLINE_U : unsigned(AGE_W-1 downto 0) :=
to_unsigned(DEADLINE, AGE_W);
signal pend_r : std_logic := '0';
signal age_r, last_r, max_r : unsigned(AGE_W-1 downto 0)
:= (others => '0');
signal miss_r : std_logic := '0';
begin
-- A deadline the counter cannot represent would make deadline_miss
-- permanently false: a monitor silently reporting that nothing is late.
assert DEADLINE < 2**AGE_W
report "DEADLINE is unrepresentable in AGE_W bits"
severity failure;
pending <= pend_r;
age <= age_r;
last_latency <= last_r;
max_latency <= max_r;
deadline_miss <= miss_r;
process (clk, rst_n)
begin
if rst_n = '0' then
pend_r <= '0';
age_r <= (others => '0');
last_r <= (others => '0');
max_r <= (others => '0');
miss_r <= '0';
elsif rising_edge(clk) then
if bus_reset = '1' then
pend_r <= '0';
age_r <= (others => '0');
last_r <= (others => '0');
max_r <= (others => '0');
miss_r <= '0';
else
-- 1. DELIVERY: record the measurement of the item leaving now.
-- age_r is the age BEFORE this cycle's update, which is the
-- number of frames fully elapsed since the item arrived.
if report_taken = '1' and pend_r = '1' then
last_r <= age_r;
if age_r > max_r then
max_r <= age_r;
end if;
if age_r > DEADLINE_U then
miss_r <= '1';
end if;
end if;
-- 2. PENDING and AGE.
if report_taken = '1' and event_valid = '1' then
-- the collision: the delivery empties the endpoint and the new
-- event refills it, starting its own measurement at zero.
pend_r <= '1';
age_r <= (others => '0');
elsif report_taken = '1' then
pend_r <= '0';
age_r <= (others => '0');
elsif event_valid = '1' and pend_r = '0' then
-- only the FIRST event starts the clock; a later event arriving
-- while one is already pending must NOT restart it.
pend_r <= '1';
age_r <= (others => '0');
elsif frame_tick = '1' and pend_r = '1' then
-- saturate: an age that wraps reports a small latency for a very
-- late item, the worst direction for a monitor to fail in.
if age_r /= AGE_MAX then
age_r <= age_r + 1;
end if;
end if;
end if;
end if;
end process;
end architecture;assert ... severity failure is VHDL's elaboration check, and it is a concurrent statement rather than an initial block — it is evaluated as part of the architecture, not as a process that runs at time zero. The intent matches §8's $fatal exactly.
unsigned rather than signed here, deliberately, because an age is a count and counts are never negative. Chapter 15.3 §10 made the case that VHDL's distinction between the two types is load-bearing; this module is the other half of that case, where the other type is the right one and the compiler will enforce it.
age_r + 1 needs no conversion. numeric_std defines addition of an unsigned and a natural, so the literal is unambiguous — unlike the Verilog, where the increment is written {{(AGE_W-1){1'b0}}, 1'b1} to match the counter's width exactly.
10. Comparing the Three
| Concern | Verilog | SystemVerilog | VHDL |
|---|---|---|---|
| Illegal parameterisation | undetected | $fatal at elaboration | assert ... severity failure |
| The four cases | an else if chain | typedef enum plus unique case | an elsif chain |
| Incrementing the counter | width-matched concatenation | age + 1'b1 | age_r + 1, operator defined for the type |
| Counter type | reg [N-1:0], unsignedness by convention | logic [N-1:0], same | unsigned, enforced |
| Exclusivity of the cases | implied by position only | stated — though see §8's tooling note | implied by position only |
The honest summary is that all three describe identical hardware and the differences are about what the author's intent is written down. Verilog records the design; SystemVerilog records the design plus two claims about it; VHDL records the design plus the type of every quantity in it. None of the three prevents the design being wrong about which item to measure, which is why §13's L1 matters more than any of these rows.
11. The Testbenches
All three benches share the decision that made Chapter 15.3's bug findable: the model uses a different method from the design.
The design counts incrementally — an age register that increments on frame ticks. Every bench stores an absolute arrival timestamp and subtracts at delivery. A counter that increments at the wrong moment, or fails to saturate, or is restarted by the wrong event, cannot be reproduced by a subtraction of two timestamps.
// The model: absolute timestamps, applied in a defined order after the edge.
task automatic step(input bit ft, input bit ev, input bit rt);
frame_tick=ft; event_valid=ev; report_taken=rt;
@(posedge clk); #1;
frame_tick=0; event_valid=0; report_taken=0;
t_before = ticks; // the tick count BEFORE this cycle's
if (ft) ticks++; // tick, which is what a delivery sees
if (rt && m_pending) begin
m_lat = t_before - arrival; // SUBTRACTION, not counting
if (m_lat > AGE_MAX) m_lat = AGE_MAX;
m_last = m_lat;
if (m_lat > m_max) m_max = m_lat;
if (m_lat > DEADLINE) m_miss = 1;
m_pending = 0;
end
if (ev && (rt || !m_pending)) begin m_pending = 1; arrival = ticks; end
#1;
endtaskThe directed sequence covers accumulation, the oldest-item rule, both sides of the deadline boundary, max remaining a maximum after a short delivery, the collision, bus reset, and saturation. Then 3000 randomised steps draw the three stimulus bits and check all four outputs after every one.
12. Mutation Testing
Six mutations, each applied to the Verilog and the SystemVerilog design and run against that language's own bench, from a clean baseline.
| ID | Mutation | Verilog | SystemVerilog | Killed |
|---|---|---|---|---|
| — | baseline, no mutation | 0 | 0 | — |
| L1 | restart the timer on every arrival (newest, not oldest) | 7320 | 7320 | ✅ both |
| L2 | max_latency takes the last value, not the maximum | 2891 | 2891 | ✅ both |
| L3 | deadline comparison >= instead of > | 2 | 1 | ✅ both |
| L4 | deadline_miss not sticky | 2740 | 2739 | ✅ both |
| L5 | the age wraps instead of saturating | 2 | 2 | ✅ both |
| L6 | the collision arm removed — a simultaneous arrival is dropped | 3497 | 3497 | ✅ both |
L1 is the module's reason for existing and it dies hardest — 7320 failures, more than twice any other mutation, because restarting the timer corrupts every subsequent measurement rather than one.
The Verilog and SystemVerilog counts are nearly identical, which is itself worth reading correctly. It demonstrates that the two designs are equivalent; it does not demonstrate that the two benches are independent. They are transcriptions of one another driving the same seeded $random sequence, so they reach the same states in the same order. Cross-HDL equivalence is shown here; cross-HDL stimulus diversity is not, and claiming the second from these numbers would be wrong.
13. What 3000 Randomised Steps Could Not Reach
L3 and L5 were killed by one or two failures each. That is a pass, and treating it as one would miss the finding. Both were killed entirely by directed tests, and the randomised phase contributed nothing at all — which is measurable rather than a matter of opinion, so it was measured:
RANDOM-PHASE REACH over 3000 steps:
deliveries observed ......... 273
collisions (deliver+arrive) . 84
greatest age reached ........ 31 (saturation needs 255)
deliveries at exactly 10 .... 3The saturation boundary is not merely unexercised — it is unreachable. Reaching an age of 255 requires 255 consecutive steps with no delivery, and with a delivery probability of one in seven that has a probability of about 10^−17. No amount of running this bench longer would find L5. The stimulus distribution, not the check list, is what excludes it.
The collision, by contrast, occurred 84 times, which is why L6 dies by 3497 — the same bench covers one boundary richly and another not at all, and nothing in a pass/fail result distinguishes the two cases.
The sticky flag that masks its own detection
L3 — the off-by-one on the deadline comparison — has a second reason for its low count, and it is a genuinely surprising one. Here is where its two failures actually occurred:
FAIL: exactly the deadline is NOT a miss (age=0 last=10 max=10 miss=1, t=277000)
FAIL: at the deadline (age=0 last=10 max=10 miss=1, t=277000)Both at the same instant — the directed at-the-deadline test, and nothing afterwards.
The reason is the flag's own design. deadline_miss is sticky: once any genuine violation sets it, it stays set. The random phase reaches ages up to 31 against a deadline of 10, so real misses happen early and the flag is legitimately high for the rest of the run. After that point the mutation's false positive is indistinguishable from the true positive already present, and every subsequent opportunity to detect it is masked.
A sticky flag is saturating evidence. It answers did this ever happen and is therefore structurally incapable of detecting an error that makes it happen again. Verifying the boundary of a sticky condition requires a scenario in which the condition has not yet occurred — which means it must be checked early, or after a reset, and can never be checked by accumulating more stimulus.
This generalises well beyond USB. Error flags, overflow bits, fault latches and "has ever been out of spec" status registers are all saturating evidence, and all share the property that their most testable moment is their first.
14. The Waveform
Two deliveries: one at the deadline, one past it
14 cyclesFrame 6 against frame 13 is the boundary that mutation L3 attacks. A monitor written with >= would set the flag at frame 6 as well, reporting a violation for a delivery that met its deadline exactly — and §13 showed that once the flag is set, nothing afterwards can reveal the error.
15. Assertions
// A1. The oldest-item rule, stated directly: an arrival while something is
// already pending must not disturb the age. This is the property the
// RTL expresses by OMISSION, so it is the one most worth asserting.
property p_arrival_does_not_restart;
@(posedge clk) disable iff (!rst_n || bus_reset)
(event_valid && pending && !report_taken && !frame_tick)
|=> (age == $past(age));
endproperty
a_arrival_does_not_restart: assert property (p_arrival_does_not_restart);
// A2. max_latency is monotonically non-decreasing between resets.
property p_max_monotonic;
@(posedge clk) disable iff (!rst_n || bus_reset)
1'b1 |=> (max_latency >= $past(max_latency));
endproperty
a_max_monotonic: assert property (p_max_monotonic);
// A3. deadline_miss is sticky: it never falls except through a reset.
property p_miss_sticky;
@(posedge clk) disable iff (!rst_n || bus_reset)
$fell(deadline_miss) |-> 1'b0;
endproperty
a_miss_sticky: assert property (p_miss_sticky);
// A4. The boundary, in the direction section 13 showed is hard to test:
// a delivery at EXACTLY the deadline must not raise the flag. The
// antecedent requires the flag to be low BEFORE the delivery, which
// is the only window in which a sticky flag can be falsified.
property p_boundary_exclusive;
@(posedge clk) disable iff (!rst_n || bus_reset)
(report_taken && pending && !deadline_miss && age == AGE_W'(DEADLINE))
|=> !deadline_miss;
endproperty
a_boundary_exclusive: assert property (p_boundary_exclusive);
// A5. The age saturates: it never decreases on a tick, and never wraps.
property p_age_saturates;
@(posedge clk) disable iff (!rst_n || bus_reset)
(frame_tick && pending && !event_valid && !report_taken)
|=> (age >= $past(age));
endproperty
a_age_saturates: assert property (p_age_saturates);Assertion contracts
| What it claims | What would make it vacuous | How non-vacuity is established | |
|---|---|---|---|
| A1 | an arrival while pending leaves the age alone | arrivals never coincide with pending | the random phase produced 84 delivery/arrival collisions and many more arrive-while-pending cases |
| A2 | max_latency never decreases | trivially true if max never changes | the directed sequence raises max four times |
| A3 | the miss flag never falls | trivially true if the flag never rises | the directed sequence sets it, and the random phase sets it early |
| A4 | a delivery at exactly the deadline does not raise the flag | the antecedent requires deadline_miss low — after the first real miss it is never satisfied again | only the directed test satisfies it, which is §13's finding restated as an assertion |
| A5 | the age never decreases on a tick | ticks never occur while pending | ticks dominate the random phase |
A4's vacuity row is the whole point of writing the contracts down. The property is correct, it is not trivially true, and it is nevertheless satisfiable only in a narrow window that ordinary stimulus closes permanently. A vacuity report that merely said "A4 was exercised" would be true and useless; what matters is that it was exercised twice, both in the directed phase, and that no quantity of extra random stimulus would add a third.
16. Verification: What a UVM Environment Adds
The argument that justified UVM in 15.3 §15 — the property spans an unbounded stream — applies here in a sharper form, because the property being verified is itself a statistic.
| Component | Why it is justified |
|---|---|
| Arrival sequence | the distribution is the test (§13). Constrained-random on the gap between arrivals, with constraints that deliberately produce long droughts, is the only way to reach the saturation region |
| Poll agent | polls must be schedulable independently of arrivals, including deliberately delayed and deliberately absent, to produce the ages the distribution otherwise excludes |
| Monitor | observes event_valid and report_taken on the interface and timestamps them itself — it must not be told the ages by the sequence |
| Scoreboard | keeps its own record of every outstanding item and its arrival time, and predicts last, max and miss from that record |
| Coverage | must include a reach cover group: maximum age attained, deliveries at exactly the deadline, arrivals-while-pending, and collisions — the four quantities §13 had to measure by hand |
The coverage row is the one this chapter argues for hardest. §13's finding was not produced by a check and could not have been: it came from instrumenting what the stimulus reached. A functional coverage model that includes the boundaries the checks care about turns that from an ad-hoc measurement into a standing report — and would have shown a permanent hole at the saturation bin from the first regression.
And the scoreboard must keep a queue, not a counter. A scoreboard that tracks only "the current pending item's arrival time" has made the same simplification the design makes, and would therefore agree with mutation L1 — the newest-item bug — for exactly the same reason. It must model the endpoint as an ordered collection whose oldest member is the one being measured, which is a different structure from the design's single register and is what makes the comparison meaningful.
17. Debugging: the Device That Meets Its Deadline and Feels Slow
A device declares a 1 ms interval, its in-hardware latency monitor reports a maximum of 2 frames over hours of operation, and users report that it feels laggy and inconsistent. The monitor is not lying — the RTL was verified — and the complaint is real.
Both facts are true, and §3 is why. The monitor measures term 3 and part of term 2. The user feels terms 1 through 7. A device can be perfectly compliant on the only term it can see while the chain it sits in is dominated by terms it cannot.
Work the chain from the ends inward:
Term 1, the sensor period. A 1 ms poll on a 125 Hz sensor gives 8 ms of latency that the monitor never sees, because the monitor's clock starts when data becomes ready, not when the physical event occurred. This is the most common cause and the easiest to overlook, because the device's own instrumentation is structurally blind to it — the event does not exist, as far as the endpoint is concerned, until the sensor has already delayed it.
Term 6, the host wakeup. Measurable from the host: timestamp in the driver's completion handler and again in the application, and look at the distribution's tail, not its mean. "Inconsistent" is a statement about the tail.
Term 3's granted value. §2: the interval the device requested is not necessarily the interval the host scheduled. Read back the URB's interval field after submission. A controller clamp in the other direction — granting a shorter interval — would not cause lag, but a driver that re-submits slowly can leave the endpoint unpolled regardless of what was granted.
The report_taken definition. §7's stated assumption. If the monitor's report_taken is driven by a poll being attempted rather than data being delivered, then NAKed polls reset the measurement and the monitor reports the age since the last NAK instead of the age since arrival. It would report small numbers under exactly the congested conditions that produce large real latencies — the same failure direction as §5's newest-item bug, arriving by a different route.
18. Common Misconceptions
"bInterval is the polling period." It is an upper bound on the gap. The kernel states the host "may be more frequent than requested" (§2), and a device must not assume a minimum gap.
"bInterval means the same thing at every speed." It is linear in frames at low and full speed and logarithmic in microframes at high speed and above (§1). The root hub descriptors show the same 256 ms interval written as 0xff and as 0x0c.
"A 1 ms interval gives a 1 ms response." It bounds one of seven terms, and at 60 Hz the display term alone is 16.7 ms (§3).
"A 1000 Hz mouse is eight times better than a 125 Hz one." It improves one term by 7.875 ms and reduces variance and delta clipping (§4). The other terms are unchanged.
"The latency is the age of the data being sent." It is the age of the oldest undelivered item (§5). Measuring the newest under-reports precisely when the system is most loaded.
"A monitor that passes 3000 randomised steps has been thoroughly tested." Two of this module's six mutations were killed only by directed tests, and one of the two was in a region the random distribution reaches with probability about 10^−17 (§13).
"A sticky error flag is the safe way to record a violation." It is safe for recording and hostile to verifying, because after the first true assertion it can no longer be falsified (§13).
19. Exercises
1. Compute bInterval for a 4 ms interval at full speed and at high speed, then compute what each value would mean if used at the other speed. Confirm your high-speed answer against the root hub example in §1.
2. Take the §13 measurement further: instrument the randomised phase to report the full distribution of measured latencies, not just the maximum. Determine the smallest delivery probability for which an age of 255 becomes reachable within 3000 steps, and say whether that stimulus is still representative.
3. Implement A4 from §15 as a procedural check and confirm it fires under L3 — then show that it stops firing once a genuine miss has occurred, and propose a bench structure that keeps it checkable.
4. Add a second endpoint with a different interval and a shared frame counter. Determine whether one monitor per endpoint is required or whether the state can be shared, and justify the answer from §5 rather than from area.
5. Implement the report_taken-means-attempted defect from §17 and measure it. Note that the present bench cannot produce a NAK at all, and specify precisely what must be added to the stimulus before the defect is detectable.
6. Work out what the else if chain in §7 produces under the two incorrect orderings named there — arrival ahead of delivery, and tick ahead of arrival — and give a concrete stimulus that distinguishes each from the correct design.
20. Summary
bInterval is one byte with two encodings — linear frames at low and full speed, logarithmic microframes at high speed and above (§1) — and the kernel's own root hub descriptors verify the arithmetic: 0xff is 255 ms at full speed and 0x0c is 256 ms at high speed.
It is a request, not a contract (§2). The host may poll more frequently than asked, controllers clamp it to their own limits, and the granted interval is written back as a different number from the requested one. The only guarantee is a bound on the gap.
That bound covers one term of seven (§3). Two of the remaining terms are on the host side and unbounded under load, and one — the display frame — is larger at 60 Hz than the entire polling interval of a fast endpoint. This is why a 1000 Hz mouse is a real but modest improvement (§4), and why its most valuable effect is arguably the reduction in delta clipping rather than the latency itself.
Measuring latency in hardware has one rule that dominates the rest (§5): the measurement is the age of the oldest undelivered item. The newest-item version fails in the direction that hides congestion, and mutation L1 costs 7320 failures — more than twice any other mutation here.
Six mutations, all killed in both languages (§12) — but two of them were killed entirely by directed tests, and instrumenting the randomised phase showed why (§13): in 3000 steps the greatest age reached was 31 against the 255 saturation needs, a region the stimulus distribution reaches with probability around 10^−17. No amount of additional running would have found it.
And a sticky flag cannot be falsified twice (§13). Once a genuine deadline miss set deadline_miss, the off-by-one mutation on the boundary comparison became undetectable — its two failures both occurred at the single directed moment before the first real violation. Saturating evidence is most testable at the beginning.
Verilog and SystemVerilog compiled and simulated clean; VHDL was written and reviewed but not executed (§21), and no claim to the contrary is made anywhere above.
21. Tooling, Honestly
| Language | Design | Testbench | Compiled | Simulated | Mutations |
|---|---|---|---|---|---|
| Verilog-2005 | usb_latency_monitor | lat_v_tb.v | ✅ Icarus -g2005 | ✅ 0 errors | ✅ all six |
| SystemVerilog | usb_latency_monitor_sv | lat_sv_tb.sv | ✅ Icarus -g2012 | ✅ 0 errors | ✅ all six |
| VHDL | usb_latency_monitor_vhdl | lat_vhdl_tb.vhd | ❌ tool unavailable | ❌ tool unavailable | ❌ not run |
| SVA (§15) | — | — | ❌ unsupported by Icarus | ❌ | — |
unique case | — | — | ✅ accepted | ⚠️ quality ignored | — |
No VHDL analyser or simulator exists in this environment — ghdl, nvc, vcom, xvhdl and vsim were each checked — and installing one was out of scope. The VHDL was instead reviewed structurally against the properties a compiler would check: entity and architecture pairing, positional port association matching the declaration in count and order, every input driven and every output checked, numeric_std with no deprecated arithmetic package, an asynchronous-reset process sensitive to both clk and rst_n, a reset value for every registered signal, and the elsif chain ordered collision-before-delivery-before-arrival.
The unique case row is the subtler entry. Icarus accepts the keyword and reports sorry: Case unique/unique0 qualities are ignored — so on this tool the claim is documentation, not a check. A keyword that is accepted and not enforced is worse than one that is rejected, because it reads in review as a verified property.
22. What Comes Next
Module 15 built the transfer type that buys a bound on the gap and pays for it with a small, capped share of bandwidth — and this chapter showed how narrow that bound really is.
Module 16 — Isochronous Transfers — is the other periodic transfer type, and it makes the opposite trade. Isochronous transfers reserve bandwidth rather than buying a latency bound, and they pay for it by giving up something interrupt transfers never surrender: the retry. An isochronous packet that arrives corrupted is not resent, because by the time a retry could complete, the moment the data belonged to has passed.
That single difference reorganises everything — error handling, buffering, the meaning of a lost packet, and what the hardware must do when data does not arrive on time.
Browse the full path on the USB tutorials index.
Continue learning
Related tutorials
- Related topic
The Polling Model
The host runs a periodic schedule the device cannot see, so device hardware must count frames rather than trust a period — built in Verilog, SystemVerilog and VHDL with a schedule model that disagrees independently.
- Related topic
Keyboards on USB
A key matrix produces a set; the boot report has six slots. The mechanism that reconciles them has its own reserved code — and the chapter where the verification transaction stops being the pins.
- Related topic
Mice on USB
A relative report cannot be resent, so the accumulator must saturate rather than wrap and must be cleared by the act of being read — and a real signedness bug the testbench caught on its first check.
- Related topic
Real-Time Data Streams
A sensor stream must know exactly how much data is missing, not just that something is. Sequence numbers, the modular arithmetic that survives a wrap, and the half-space limit beyond which a gap cannot be measured at all.
Standards & specifications
- Governing standard
- USB-IF (Universal Serial Bus Specification)(opens USB Implementers Forum (USB-IF) in a new tab)
Defines the USB bus — its electrical signalling, connectors, packet and transaction model, device framework and the descriptors a device must expose — together with the device-class specifications layered on it. It does not define host-controller register interfaces (xHCI and EHCI are separate documents) nor any operating system's driver architecture.
This page also covers RTL structure, verification approach and debugging technique. Those are engineering practice built on the standard, not requirements the standard itself imposes.
Where this fits
Part of the USB curriculum.
