Skip to content
VLSI Mentor

USB · Module 15

Interrupt Latency Requirements

bInterval is a request, not a contract, and it bounds one term of seven. A latency monitor in three HDLs, and the measurement showing 3000 randomised steps could not reach two of its own boundaries.

Chapter 15.1 built the schedule, and 15.2 and 15.3 built devices that took it as given: a poll arrives, eventually, within the interval the endpoint asked for.

This chapter asks what that interval actually buys, and the honest answer is much less than the number suggests. An endpoint that declares a 1 ms interval does not thereby give the user a 1 ms response. It bounds one link in a chain that starts at a sensor and ends at a pixel, and several of the other links are larger, less predictable, and outside the device's control entirely.

The interval bounds one link. The user feels the chain.

1. What bInterval Actually Encodes

The endpoint descriptor's bInterval field is one byte, and what that byte means depends on the speed of the device — not as a footnote, but as two genuinely different encodings.

SpeedUnitEncodingRangeResulting interval
Low speedframe (1 ms)linear — the value is the count10 – 25510 ms – 255 ms
Full speedframe (1 ms)linear1 – 2551 ms – 255 ms
High speedmicroframe (125 µs)logarithmic — 2^(bInterval−1)1 – 16125 µs – 4.096 s
SuperSpeedmicroframe (125 µs)logarithmic1 – 16125 µs – 4.096 s

The kernel states the split plainly, in the documentation of the URB field that carries it:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
 * @interval: Specifies the polling interval for interrupt or isochronous
 *	transfers.  The units are frames (milliseconds) for full and low
 *	speed devices, and microframes (1/8 millisecond) for highspeed
 *	and SuperSpeed devices.

and is explicit that the descriptor's encoding is the driver's problem:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
 * (Note that for isochronous
 * endpoints, as well as high speed interrupt endpoints, the encoding of
 * the transfer interval in the endpoint descriptor is logarithmic.
 * Device drivers must convert that value to linear units themselves.)

The root hub's own descriptors are a clean worked example of both encodings, because the kernel builds one for each speed and annotates each with the interval it intends:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
/* full-speed root hub, interrupt endpoint */
0xff        /*  __u8  ep_bInterval; (255ms -- usb 2.0 spec) */

/* high-speed root hub, interrupt endpoint */
0x0c        /*  __u8  ep_bInterval; (256ms -- usb 2.0 spec) */

Check the arithmetic against the comments, because it is the fastest way to internalise the difference:

  • Full speed, linear: 0xff = 255 frames × 1 ms = 255 ms. ✅ matches.
  • High speed, logarithmic: 0x0c = 12, so 2^(12−1) = 2048 microframes × 125 µs = 256 000 µs = 256 ms. ✅ matches.

The same intended interval — roughly a quarter of a second — is written 255 at full speed and 12 at high speed. A device that copies a full-speed descriptor's bInterval into a high-speed one does not get a slightly different interval. 0xff at high speed is out of range; a value of 12 at full speed asks for 12 ms instead of 256 ms, a factor of twenty-one the wrong way.

2. The Interval Is a Request, Not a Contract

This is the chapter's central correction, and the kernel says it in one sentence:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
 * The polling interval may be more frequent than requested.
 * For example, some controllers have a maximum interval of 32 milliseconds,
 * while others support intervals of up to 1024 milliseconds.

and then, about what happens after submission:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
 * After the URB has been submitted, the interval
 * field reflects how the transfer was actually scheduled.

Three separate facts are packed into that.

The host may poll more often than asked. bInterval is an upper bound on the gap, not a period. A device that asks for 8 ms may be polled every 8 ms, or every 4, or every 1. Nothing in the protocol promises the device a minimum gap, which means device hardware must not assume one — Chapter 15.1 §4 built the scheduler on exactly this assumption and it is the reason it counts frames rather than trusting a period.

Host controllers have hard limits that silently clamp. A controller with a 32 ms maximum cannot honour a request for 255 ms. It does not fail; it schedules 32 ms. A device asking for a long interval to save power may be polled eight times more often than it planned for.

The granted interval is readable, and it is a different number from the requested one. The interval field is annotated (modify) in the URB structure — the host writes back what it actually scheduled. A driver that wants to know the real polling rate must read it after submission, not compute it from the descriptor.

3. The Latency Budget, Term by Term

Here is the chain from a physical event to a visible response, with the polling interval in its actual place — one term among seven.

#TermTypical magnitudeBounded byWho controls it
1Sensor sampling period1 ms at 1 kHz; 8 ms at 125 Hzthe sensor's own clockdevice
2Device processing and arming the endpointµsdevice firmware and RTLdevice
3Wait for the next poll0 to bIntervalbIntervalhost schedule
4The transaction on the wirea few µs at high speedpacket size and speedbus
5Controller → driver completionµs to msinterrupt coalescing, IRQ latencyhost OS
6Driver → applicationµs to tens of msscheduler, wakeup, contentionhost OS
7Application → visible outputone display frame: 16.7 ms at 60 Hzthe display pipelineapplication

Read the magnitudes, not the row count. Term 3 — the one bInterval bounds, and the only one most engineers think about — is often not the largest term. At 60 Hz, term 7 alone is 16.7 ms, more than sixteen times a 1 ms polling interval, and it is a hard floor that no amount of polling can move.

Two structural observations follow, and they are what make this a budget rather than a list.

Terms 1 and 3 do not add naively — they beat against each other. A 1 kHz sensor and a 1 ms poll are not synchronised, so the delay from a sample to the poll that carries it varies across the full interval. The worst case is close to the sum, the average is closer to the sum of the halves, and the variation itself is the problem for anything that cares about consistency rather than mean latency. This is the same phase relationship Chapter 15.1 §5 spread deliberately across endpoints.

Only terms 1–3 are bounded at all. Terms 5 and 6 are scheduling latencies on a general-purpose operating system, which has no upper bound under load. A device can guarantee its half of the chain and still be part of a system that misses a deadline, because the unbounded terms are on the other side.

Reducing bInterval from 8 ms to 1 ms improves one bounded term by 7 ms. If term 6 occasionally costs 20 ms under load, the user's experience is dominated by a term the device cannot see, and the eightfold increase in reserved bandwidth bought very little.

4. Why 1000 Hz Mice Are Not Eight Times Better

The budget explains a claim the peripheral market makes constantly and mostly wrongly.

A "1000 Hz" mouse declares a 125 µs high-speed interval (bInterval = 1) against a "125 Hz" mouse's 8 ms. The marketing difference is 7.875 ms in term 3. Against a 16.7 ms display frame in term 7 and a variable term 6, that is real but small — and it is not the eightfold improvement the numbers imply, because the other terms did not change.

What the higher rate genuinely does buy is worth stating precisely, because it is not nothing:

Lower variance, not just lower mean. The poll-wait term varies between 0 and bInterval; shrinking the interval shrinks the spread. For a human closing a control loop by hand, consistency is more perceptible than latency.

Less accumulation per report. Chapter 15.3 §4 showed the accumulator saturates at ±127. At 8 ms intervals a fast movement can genuinely exceed that and be clipped; at 125 µs it essentially cannot. The faster poll removes a source of non-linearity, which is a different benefit from latency and arguably the more important one.

And what it costs is term 3's bandwidth reservation, multiplied by eight, taken from the same capped budget every other periodic endpoint draws on.

5. Measuring It in Hardware: the Oldest-Item Rule

A device that must prove it meets a latency requirement needs to measure it, and the measurement has one trap that matters more than all the rest.

When data is already waiting and more data arrives, which item's age is the latency?

The answer is the oldest. The item that has been waiting longest is the one whose deadline is closest, and it is the one the user is actually waiting on. A monitor that restarts its timer whenever new data arrives reports the age of the newest item — and under load, when data arrives constantly, the newest item is always young.

6. The Hardware, Before Any Language

State retained: a pending flag, an age counter, the last measured latency, a running maximum, and a sticky deadline-miss flag.

On reset or bus reset: everything clears. A bus reset ends the relationship with the host, and statistics gathered for one host are not statistics for the next.

On a delivery (report_taken while pending): the age before this cycle's update is the measurement. It updates last_latency, updates max_latency if it is greater, and sets the sticky miss flag if it exceeds the deadline.

On an arrival with no item pending: pending sets and the age starts at zero.

On an arrival while an item is already pending: nothing happens. This is §5's rule, and it is expressed as the absence of an action, which is why it is easy to get wrong — the correct code has no line in it.

On a delivery and an arrival in the same cycle: the delivery measures the item it delivered, and the new arrival starts its own measurement at zero. Structurally the same collision as Chapter 15.3 §5.

On a frame tick while pending and none of the above: the age increments, saturating rather than wrapping.

Saturation matters more here than it did in 15.3. A wrapped mouse delta produces a visible glitch; a wrapped age reports a small latency for a very late item — the monitor actively asserting that a badly missed deadline was met.

7. Verilog

The RTL contract

  • What it models: an in-hardware latency monitor for one interrupt endpoint.
  • Why it exists: because §3's term 3 is the device's to prove, and proving it requires measuring the age of the oldest undelivered item (§5).
  • Inputs: frame_tick, event_valid (data became ready), report_taken (the host actually took it), bus_reset.
  • State retained: pend_r, age_r, last_r, max_r, miss_r.
  • Outputs: pending, age, last_latency, max_latency, deadline_miss.
  • Hardware implied: one flag flop, one saturating counter, two comparators, two holding registers, one sticky flop.
  • Reset: asynchronous active-low rst_n; bus_reset synchronous and equivalent.
  • Assumptions: report_taken means delivered, not attempted — a NAKed poll must not assert it, or the monitor measures the wrong interval and reports latencies that never happened.
  • Omissions: no CDC to a register block, no clear-on-read of the statistics, one endpoint only.
  • What DV should verify: that the measurement is the age of the oldest item under every arrival pattern; that the deadline comparison is exclusive at the boundary; that the age saturates; that max is a maximum and deadline_miss is sticky.
Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// usb_latency_monitor -- measures the age of the OLDEST undelivered event.
//
// The subtlety this module exists for: when data is already waiting and more
// data arrives, the latency that matters is still the age of the FIRST item,
// because that is the one the user has been waiting on. A monitor that
// restarts its timer on each new event reports the age of the NEWEST item and
// systematically under-reports -- exactly the direction that hides a problem.
module usb_latency_monitor #(
  parameter integer AGE_W    = 8,
  parameter integer DEADLINE = 10          // frames; a miss is age > DEADLINE
) (
  input  wire              clk,
  input  wire              rst_n,
  input  wire              bus_reset,
  input  wire              frame_tick,
  input  wire              event_valid,    // data became ready in the endpoint
  input  wire              report_taken,   // the host actually took it
  output wire              pending,
  output wire [AGE_W-1:0]  age,
  output wire [AGE_W-1:0]  last_latency,
  output wire [AGE_W-1:0]  max_latency,
  output wire              deadline_miss   // sticky
);
  localparam [AGE_W-1:0] AGE_MAX = {AGE_W{1'b1}};

  reg              pend_r;
  reg [AGE_W-1:0]  age_r, last_r, max_r;
  reg              miss_r;

  assign pending      = pend_r;
  assign age          = age_r;
  assign last_latency = last_r;
  assign max_latency  = max_r;
  assign deadline_miss= miss_r;

  // The measured latency of the delivery happening this cycle: the age
  // BEFORE this cycle's update, i.e. frames fully elapsed since arrival.
  wire [AGE_W-1:0] measured = age_r;

  always @(posedge clk or negedge rst_n) begin
    if (!rst_n) begin
      pend_r <= 1'b0; age_r <= {AGE_W{1'b0}};
      last_r <= {AGE_W{1'b0}}; max_r <= {AGE_W{1'b0}}; miss_r <= 1'b0;
    end else if (bus_reset) begin
      // A bus reset ends the relationship: the pending item and every
      // statistic about the previous host are discarded.
      pend_r <= 1'b0; age_r <= {AGE_W{1'b0}};
      last_r <= {AGE_W{1'b0}}; max_r <= {AGE_W{1'b0}}; miss_r <= 1'b0;
    end else begin
      // 1. DELIVERY. Record the measurement before anything else changes.
      if (report_taken && pend_r) begin
        last_r <= measured;
        if (measured > max_r) max_r <= measured;
        if (measured > DEADLINE[AGE_W-1:0]) miss_r <= 1'b1;
      end

      // 2. PENDING and AGE.
      if (report_taken && event_valid) begin
        // The collision: the delivery empties the endpoint and the new event
        // immediately refills it. The new item's age starts at zero NOW.
        pend_r <= 1'b1;
        age_r  <= {AGE_W{1'b0}};
      end else if (report_taken) begin
        pend_r <= 1'b0;
        age_r  <= {AGE_W{1'b0}};
      end else if (event_valid && !pend_r) begin
        // A new oldest item. Only the FIRST event starts the clock.
        pend_r <= 1'b1;
        age_r  <= {AGE_W{1'b0}};
      end else if (frame_tick && pend_r) begin
        // Saturate. An age that wraps reports a small latency for a very
        // late item -- the worst possible direction for a monitor to fail.
        if (age_r != AGE_MAX) age_r <= age_r + {{(AGE_W-1){1'b0}}, 1'b1};
      end
    end
  end
endmodule

The else if chain's order is the specification. Collision first, then delivery, then arrival, then tick. Reordering arrival ahead of delivery would make a simultaneous pair measure the wrong item; reordering tick ahead of arrival would charge a newly arrived item for a frame it did not wait. Exercise 6 asks for the other orderings to be worked out, because the chain reads as arbitrary until they are.

And note the branch that is not there. There is no arm for arrival while already pending — that case falls through every condition and changes nothing, which is precisely §5's rule. The correct code expresses it by omission, and that is exactly why mutation L1, which adds a plausible-looking arm, is the one to measure.

8. SystemVerilog

Same hardware, with two additions that pay for themselves: an elaboration-time check, and a named event type instead of an inferred chain.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
module usb_latency_monitor_sv #(
  parameter int unsigned AGE_W    = 8,
  parameter int unsigned DEADLINE = 10
) (
  input  logic             clk,
  input  logic             rst_n,
  input  logic             bus_reset,
  input  logic             frame_tick,
  input  logic             event_valid,
  input  logic             report_taken,
  output logic             pending,
  output logic [AGE_W-1:0] age,
  output logic [AGE_W-1:0] last_latency,
  output logic [AGE_W-1:0] max_latency,
  output logic             deadline_miss
);
  localparam logic [AGE_W-1:0] AGE_MAX = '1;

  // An elaboration-time check: a deadline the age counter cannot represent
  // would make deadline_miss permanently false, which is the worst way for a
  // monitor to be wrong -- silently reporting that nothing is ever late.
  initial begin
    if (DEADLINE >= (1 << AGE_W))
      $fatal(1, "DEADLINE=%0d is unrepresentable in AGE_W=%0d bits",
             DEADLINE, AGE_W);
  end

  // What this cycle does, named rather than inferred from an if/else chain.
  typedef enum logic [1:0] { EV_IDLE, EV_ARRIVE, EV_DELIVER, EV_BOTH } ev_e;
  ev_e ev;
  always_comb begin
    if      (report_taken && event_valid) ev = EV_BOTH;
    else if (report_taken)                ev = EV_DELIVER;
    else if (event_valid && !pending)     ev = EV_ARRIVE;
    else                                  ev = EV_IDLE;
  end

  wire [AGE_W-1:0] measured = age;

  always_ff @(posedge clk or negedge rst_n) begin
    if (!rst_n || bus_reset) begin
      pending <= 1'b0; age <= '0;
      last_latency <= '0; max_latency <= '0; deadline_miss <= 1'b0;
    end else begin
      if (report_taken && pending) begin
        last_latency <= measured;
        if (measured > max_latency)                 max_latency   <= measured;
        if (measured > AGE_W'(DEADLINE))            deadline_miss <= 1'b1;
      end

      unique case (ev)
        EV_BOTH:    begin pending <= 1'b1; age <= '0; end
        EV_DELIVER: begin pending <= 1'b0; age <= '0; end
        EV_ARRIVE:  begin pending <= 1'b1; age <= '0; end
        EV_IDLE:    if (frame_tick && pending && age != AGE_MAX)
                      age <= age + 1'b1;
      endcase
    end
  end
endmodule

The ev_e enumeration is the real improvement, and not for readability. The Verilog chain encodes four mutually exclusive cases in a structure that does not say they are exclusive; naming them makes the exclusivity checkable, and it puts §5's "arrival while pending is not an arrival" into the definition of EV_ARRIVE rather than into the middle of a sequential chain.

$fatal catches a parameterisation that would silently disable the monitor. With AGE_W = 8 and DEADLINE = 300, the comparison measured > 300 can never be true, the miss flag never sets, and the module reports perfect compliance forever. That is not a subtle degradation — it is a monitor that lies — and it is a configuration error a designer can make in one keystroke. Catching it at elaboration costs four lines.

9. VHDL

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

entity usb_latency_monitor_vhdl is
  generic (
    AGE_W    : positive := 8;
    DEADLINE : natural  := 10
  );
  port (
    clk           : in  std_logic;
    rst_n         : in  std_logic;
    bus_reset     : in  std_logic;
    frame_tick    : in  std_logic;
    event_valid   : in  std_logic;
    report_taken  : in  std_logic;
    pending       : out std_logic;
    age           : out unsigned(AGE_W-1 downto 0);
    last_latency  : out unsigned(AGE_W-1 downto 0);
    max_latency   : out unsigned(AGE_W-1 downto 0);
    deadline_miss : out std_logic
  );
end entity;

architecture rtl of usb_latency_monitor_vhdl is
  constant AGE_MAX  : unsigned(AGE_W-1 downto 0) := (others => '1');
  constant DEADLINE_U : unsigned(AGE_W-1 downto 0) :=
    to_unsigned(DEADLINE, AGE_W);

  signal pend_r : std_logic := '0';
  signal age_r, last_r, max_r : unsigned(AGE_W-1 downto 0)
    := (others => '0');
  signal miss_r : std_logic := '0';
begin
  -- A deadline the counter cannot represent would make deadline_miss
  -- permanently false: a monitor silently reporting that nothing is late.
  assert DEADLINE < 2**AGE_W
    report "DEADLINE is unrepresentable in AGE_W bits"
    severity failure;

  pending       <= pend_r;
  age           <= age_r;
  last_latency  <= last_r;
  max_latency   <= max_r;
  deadline_miss <= miss_r;

  process (clk, rst_n)
  begin
    if rst_n = '0' then
      pend_r <= '0';
      age_r  <= (others => '0');
      last_r <= (others => '0');
      max_r  <= (others => '0');
      miss_r <= '0';
    elsif rising_edge(clk) then
      if bus_reset = '1' then
        pend_r <= '0';
        age_r  <= (others => '0');
        last_r <= (others => '0');
        max_r  <= (others => '0');
        miss_r <= '0';
      else
        -- 1. DELIVERY: record the measurement of the item leaving now.
        --    age_r is the age BEFORE this cycle's update, which is the
        --    number of frames fully elapsed since the item arrived.
        if report_taken = '1' and pend_r = '1' then
          last_r <= age_r;
          if age_r > max_r then
            max_r <= age_r;
          end if;
          if age_r > DEADLINE_U then
            miss_r <= '1';
          end if;
        end if;

        -- 2. PENDING and AGE.
        if report_taken = '1' and event_valid = '1' then
          -- the collision: the delivery empties the endpoint and the new
          -- event refills it, starting its own measurement at zero.
          pend_r <= '1';
          age_r  <= (others => '0');
        elsif report_taken = '1' then
          pend_r <= '0';
          age_r  <= (others => '0');
        elsif event_valid = '1' and pend_r = '0' then
          -- only the FIRST event starts the clock; a later event arriving
          -- while one is already pending must NOT restart it.
          pend_r <= '1';
          age_r  <= (others => '0');
        elsif frame_tick = '1' and pend_r = '1' then
          -- saturate: an age that wraps reports a small latency for a very
          -- late item, the worst direction for a monitor to fail in.
          if age_r /= AGE_MAX then
            age_r <= age_r + 1;
          end if;
        end if;
      end if;
    end if;
  end process;
end architecture;

assert ... severity failure is VHDL's elaboration check, and it is a concurrent statement rather than an initial block — it is evaluated as part of the architecture, not as a process that runs at time zero. The intent matches §8's $fatal exactly.

unsigned rather than signed here, deliberately, because an age is a count and counts are never negative. Chapter 15.3 §10 made the case that VHDL's distinction between the two types is load-bearing; this module is the other half of that case, where the other type is the right one and the compiler will enforce it.

age_r + 1 needs no conversion. numeric_std defines addition of an unsigned and a natural, so the literal is unambiguous — unlike the Verilog, where the increment is written {{(AGE_W-1){1'b0}}, 1'b1} to match the counter's width exactly.

10. Comparing the Three

ConcernVerilogSystemVerilogVHDL
Illegal parameterisationundetected$fatal at elaborationassert ... severity failure
The four casesan else if chaintypedef enum plus unique casean elsif chain
Incrementing the counterwidth-matched concatenationage + 1'b1age_r + 1, operator defined for the type
Counter typereg [N-1:0], unsignedness by conventionlogic [N-1:0], sameunsigned, enforced
Exclusivity of the casesimplied by position onlystated — though see §8's tooling noteimplied by position only

The honest summary is that all three describe identical hardware and the differences are about what the author's intent is written down. Verilog records the design; SystemVerilog records the design plus two claims about it; VHDL records the design plus the type of every quantity in it. None of the three prevents the design being wrong about which item to measure, which is why §13's L1 matters more than any of these rows.

11. The Testbenches

All three benches share the decision that made Chapter 15.3's bug findable: the model uses a different method from the design.

The design counts incrementally — an age register that increments on frame ticks. Every bench stores an absolute arrival timestamp and subtracts at delivery. A counter that increments at the wrong moment, or fails to saturate, or is restarted by the wrong event, cannot be reproduced by a subtraction of two timestamps.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
  // The model: absolute timestamps, applied in a defined order after the edge.
  task automatic step(input bit ft, input bit ev, input bit rt);
    frame_tick=ft; event_valid=ev; report_taken=rt;
    @(posedge clk); #1;
    frame_tick=0; event_valid=0; report_taken=0;
    t_before = ticks;                    // the tick count BEFORE this cycle's
    if (ft) ticks++;                     // tick, which is what a delivery sees
    if (rt && m_pending) begin
      m_lat = t_before - arrival;        // SUBTRACTION, not counting
      if (m_lat > AGE_MAX) m_lat = AGE_MAX;
      m_last = m_lat;
      if (m_lat > m_max) m_max = m_lat;
      if (m_lat > DEADLINE) m_miss = 1;
      m_pending = 0;
    end
    if (ev && (rt || !m_pending)) begin m_pending = 1; arrival = ticks; end
    #1;
  endtask

The directed sequence covers accumulation, the oldest-item rule, both sides of the deadline boundary, max remaining a maximum after a short delivery, the collision, bus reset, and saturation. Then 3000 randomised steps draw the three stimulus bits and check all four outputs after every one.

12. Mutation Testing

Six mutations, each applied to the Verilog and the SystemVerilog design and run against that language's own bench, from a clean baseline.

IDMutationVerilogSystemVerilogKilled
—baseline, no mutation00—
L1restart the timer on every arrival (newest, not oldest)73207320✅ both
L2max_latency takes the last value, not the maximum28912891✅ both
L3deadline comparison >= instead of >21✅ both
L4deadline_miss not sticky27402739✅ both
L5the age wraps instead of saturating22✅ both
L6the collision arm removed — a simultaneous arrival is dropped34973497✅ both

L1 is the module's reason for existing and it dies hardest — 7320 failures, more than twice any other mutation, because restarting the timer corrupts every subsequent measurement rather than one.

The Verilog and SystemVerilog counts are nearly identical, which is itself worth reading correctly. It demonstrates that the two designs are equivalent; it does not demonstrate that the two benches are independent. They are transcriptions of one another driving the same seeded $random sequence, so they reach the same states in the same order. Cross-HDL equivalence is shown here; cross-HDL stimulus diversity is not, and claiming the second from these numbers would be wrong.

13. What 3000 Randomised Steps Could Not Reach

L3 and L5 were killed by one or two failures each. That is a pass, and treating it as one would miss the finding. Both were killed entirely by directed tests, and the randomised phase contributed nothing at all — which is measurable rather than a matter of opinion, so it was measured:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
  RANDOM-PHASE REACH over 3000 steps:
    deliveries observed ......... 273
    collisions (deliver+arrive) . 84
    greatest age reached ........ 31  (saturation needs 255)
    deliveries at exactly 10 .... 3

The saturation boundary is not merely unexercised — it is unreachable. Reaching an age of 255 requires 255 consecutive steps with no delivery, and with a delivery probability of one in seven that has a probability of about 10^−17. No amount of running this bench longer would find L5. The stimulus distribution, not the check list, is what excludes it.

The collision, by contrast, occurred 84 times, which is why L6 dies by 3497 — the same bench covers one boundary richly and another not at all, and nothing in a pass/fail result distinguishes the two cases.

The sticky flag that masks its own detection

L3 — the off-by-one on the deadline comparison — has a second reason for its low count, and it is a genuinely surprising one. Here is where its two failures actually occurred:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
  FAIL: exactly the deadline is NOT a miss  (age=0 last=10 max=10 miss=1, t=277000)
  FAIL: at the deadline                     (age=0 last=10 max=10 miss=1, t=277000)

Both at the same instant — the directed at-the-deadline test, and nothing afterwards.

The reason is the flag's own design. deadline_miss is sticky: once any genuine violation sets it, it stays set. The random phase reaches ages up to 31 against a deadline of 10, so real misses happen early and the flag is legitimately high for the rest of the run. After that point the mutation's false positive is indistinguishable from the true positive already present, and every subsequent opportunity to detect it is masked.

A sticky flag is saturating evidence. It answers did this ever happen and is therefore structurally incapable of detecting an error that makes it happen again. Verifying the boundary of a sticky condition requires a scenario in which the condition has not yet occurred — which means it must be checked early, or after a reset, and can never be checked by accumulating more stimulus.

This generalises well beyond USB. Error flags, overflow bits, fault latches and "has ever been out of spec" status registers are all saturating evidence, and all share the property that their most testable moment is their first.

14. The Waveform

Two deliveries: one at the deadline, one past it

14 cycles
A waveform of the latency monitor over fourteen frames with a deadline of four frames. In frame one an event arrives and the pending flag sets with the age at zero. Frames two to five each carry a frame tick and the age counts one, two, three, four. In frame six the host takes the report; the pending flag clears, the age returns to zero, and the last-latency output becomes four. Because the deadline is four and a miss requires the age to exceed it, the deadline-miss flag stays low: exactly the deadline is not a miss. In frame seven a second event arrives and the pending flag sets again with the age at zero. Frames eight to twelve each carry a frame tick and the age counts one, two, three, four, five. In frame thirteen the host takes the report; last-latency becomes five, which exceeds the deadline of four, and the deadline-miss flag sets and remains set.arrival — the clock startsarrival — the clock startslatency 4 = exactly the deadline: NOT a misslatency 4 = exactly thedeadline: NOT a misssecond arrivalsecond arrivallatency 5 exceeds 4 — miss, and it is stickylatency 5 exceeds 4 — miss,and it is stickyframe012345678910111213frame_tickevent_validreport_takenpendingage00123400123450last_latency00000044444445deadline_misst0t1t2t3t4t5t6t7t8t9t10t11t12t13
Figure 1 — two measurements against a deadline of 4 frames, taken from the simulator rather than drawn by hand. The first delivery lands at exactly the deadline and is not a miss; the second is one frame later and is.

Frame 6 against frame 13 is the boundary that mutation L3 attacks. A monitor written with >= would set the flag at frame 6 as well, reporting a violation for a delivery that met its deadline exactly — and §13 showed that once the flag is set, nothing afterwards can reveal the error.

15. Assertions

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
  // A1. The oldest-item rule, stated directly: an arrival while something is
  //     already pending must not disturb the age. This is the property the
  //     RTL expresses by OMISSION, so it is the one most worth asserting.
  property p_arrival_does_not_restart;
    @(posedge clk) disable iff (!rst_n || bus_reset)
      (event_valid && pending && !report_taken && !frame_tick)
        |=> (age == $past(age));
  endproperty
  a_arrival_does_not_restart: assert property (p_arrival_does_not_restart);

  // A2. max_latency is monotonically non-decreasing between resets.
  property p_max_monotonic;
    @(posedge clk) disable iff (!rst_n || bus_reset)
      1'b1 |=> (max_latency >= $past(max_latency));
  endproperty
  a_max_monotonic: assert property (p_max_monotonic);

  // A3. deadline_miss is sticky: it never falls except through a reset.
  property p_miss_sticky;
    @(posedge clk) disable iff (!rst_n || bus_reset)
      $fell(deadline_miss) |-> 1'b0;
  endproperty
  a_miss_sticky: assert property (p_miss_sticky);

  // A4. The boundary, in the direction section 13 showed is hard to test:
  //     a delivery at EXACTLY the deadline must not raise the flag. The
  //     antecedent requires the flag to be low BEFORE the delivery, which
  //     is the only window in which a sticky flag can be falsified.
  property p_boundary_exclusive;
    @(posedge clk) disable iff (!rst_n || bus_reset)
      (report_taken && pending && !deadline_miss && age == AGE_W'(DEADLINE))
        |=> !deadline_miss;
  endproperty
  a_boundary_exclusive: assert property (p_boundary_exclusive);

  // A5. The age saturates: it never decreases on a tick, and never wraps.
  property p_age_saturates;
    @(posedge clk) disable iff (!rst_n || bus_reset)
      (frame_tick && pending && !event_valid && !report_taken)
        |=> (age >= $past(age));
  endproperty
  a_age_saturates: assert property (p_age_saturates);

Assertion contracts

What it claimsWhat would make it vacuousHow non-vacuity is established
A1an arrival while pending leaves the age alonearrivals never coincide with pendingthe random phase produced 84 delivery/arrival collisions and many more arrive-while-pending cases
A2max_latency never decreasestrivially true if max never changesthe directed sequence raises max four times
A3the miss flag never fallstrivially true if the flag never risesthe directed sequence sets it, and the random phase sets it early
A4a delivery at exactly the deadline does not raise the flagthe antecedent requires deadline_miss low — after the first real miss it is never satisfied againonly the directed test satisfies it, which is §13's finding restated as an assertion
A5the age never decreases on a tickticks never occur while pendingticks dominate the random phase

A4's vacuity row is the whole point of writing the contracts down. The property is correct, it is not trivially true, and it is nevertheless satisfiable only in a narrow window that ordinary stimulus closes permanently. A vacuity report that merely said "A4 was exercised" would be true and useless; what matters is that it was exercised twice, both in the directed phase, and that no quantity of extra random stimulus would add a third.

16. Verification: What a UVM Environment Adds

The argument that justified UVM in 15.3 §15 — the property spans an unbounded stream — applies here in a sharper form, because the property being verified is itself a statistic.

ComponentWhy it is justified
Arrival sequencethe distribution is the test (§13). Constrained-random on the gap between arrivals, with constraints that deliberately produce long droughts, is the only way to reach the saturation region
Poll agentpolls must be schedulable independently of arrivals, including deliberately delayed and deliberately absent, to produce the ages the distribution otherwise excludes
Monitorobserves event_valid and report_taken on the interface and timestamps them itself — it must not be told the ages by the sequence
Scoreboardkeeps its own record of every outstanding item and its arrival time, and predicts last, max and miss from that record
Coveragemust include a reach cover group: maximum age attained, deliveries at exactly the deadline, arrivals-while-pending, and collisions — the four quantities §13 had to measure by hand

The coverage row is the one this chapter argues for hardest. §13's finding was not produced by a check and could not have been: it came from instrumenting what the stimulus reached. A functional coverage model that includes the boundaries the checks care about turns that from an ad-hoc measurement into a standing report — and would have shown a permanent hole at the saturation bin from the first regression.

And the scoreboard must keep a queue, not a counter. A scoreboard that tracks only "the current pending item's arrival time" has made the same simplification the design makes, and would therefore agree with mutation L1 — the newest-item bug — for exactly the same reason. It must model the endpoint as an ordered collection whose oldest member is the one being measured, which is a different structure from the design's single register and is what makes the comparison meaningful.

A sequence diagram of the end-to-end latency chain for a USB interrupt device. A physical event occurs and the sensor samples it, which takes up to one sampling period. The device processes the sample and arms the interrupt endpoint, taking microseconds. The endpoint then waits for the host's next poll, which takes between zero and the full bInterval; this is the only term that bInterval bounds. The host polls and the transaction completes on the wire in a few microseconds. The host controller then signals the driver, which takes microseconds to milliseconds depending on interrupt coalescing. The driver wakes the application, which takes microseconds to tens of milliseconds and has no upper bound under load. Finally the application renders to the display, which costs one display frame, about sixteen point seven milliseconds at sixty hertz. The diagram marks the poll wait as the only bounded term the device influences, and marks the driver wakeup and the display frame as the two largest contributors, both outside the device's control.From physical event to visible responseEventDeviceHostApplication1. sensor samples —up to one sampleperiod2. process and armthe endpoint —microseconds3. WAIT FOR POLL — 0to bInterval (theonly bounded term)IN — the pollarrives4. DATA on the wire— a few microseconds5. controller todriver — coalescing,IRQ latency6. driver wakes theapplication —UNBOUNDED under load7. render — onedisplay frame, 16.7ms at 60 Hz
Figure 2 — the end-to-end chain from §3, with the term bInterval bounds marked. Note that the two largest contributors sit on the host side of the bus, where the device has no visibility and no control.

17. Debugging: the Device That Meets Its Deadline and Feels Slow

A device declares a 1 ms interval, its in-hardware latency monitor reports a maximum of 2 frames over hours of operation, and users report that it feels laggy and inconsistent. The monitor is not lying — the RTL was verified — and the complaint is real.

Both facts are true, and §3 is why. The monitor measures term 3 and part of term 2. The user feels terms 1 through 7. A device can be perfectly compliant on the only term it can see while the chain it sits in is dominated by terms it cannot.

Work the chain from the ends inward:

Term 1, the sensor period. A 1 ms poll on a 125 Hz sensor gives 8 ms of latency that the monitor never sees, because the monitor's clock starts when data becomes ready, not when the physical event occurred. This is the most common cause and the easiest to overlook, because the device's own instrumentation is structurally blind to it — the event does not exist, as far as the endpoint is concerned, until the sensor has already delayed it.

Term 6, the host wakeup. Measurable from the host: timestamp in the driver's completion handler and again in the application, and look at the distribution's tail, not its mean. "Inconsistent" is a statement about the tail.

Term 3's granted value. §2: the interval the device requested is not necessarily the interval the host scheduled. Read back the URB's interval field after submission. A controller clamp in the other direction — granting a shorter interval — would not cause lag, but a driver that re-submits slowly can leave the endpoint unpolled regardless of what was granted.

The report_taken definition. §7's stated assumption. If the monitor's report_taken is driven by a poll being attempted rather than data being delivered, then NAKed polls reset the measurement and the monitor reports the age since the last NAK instead of the age since arrival. It would report small numbers under exactly the congested conditions that produce large real latencies — the same failure direction as §5's newest-item bug, arriving by a different route.

18. Common Misconceptions

"bInterval is the polling period." It is an upper bound on the gap. The kernel states the host "may be more frequent than requested" (§2), and a device must not assume a minimum gap.

"bInterval means the same thing at every speed." It is linear in frames at low and full speed and logarithmic in microframes at high speed and above (§1). The root hub descriptors show the same 256 ms interval written as 0xff and as 0x0c.

"A 1 ms interval gives a 1 ms response." It bounds one of seven terms, and at 60 Hz the display term alone is 16.7 ms (§3).

"A 1000 Hz mouse is eight times better than a 125 Hz one." It improves one term by 7.875 ms and reduces variance and delta clipping (§4). The other terms are unchanged.

"The latency is the age of the data being sent." It is the age of the oldest undelivered item (§5). Measuring the newest under-reports precisely when the system is most loaded.

"A monitor that passes 3000 randomised steps has been thoroughly tested." Two of this module's six mutations were killed only by directed tests, and one of the two was in a region the random distribution reaches with probability about 10^−17 (§13).

"A sticky error flag is the safe way to record a violation." It is safe for recording and hostile to verifying, because after the first true assertion it can no longer be falsified (§13).

19. Exercises

1. Compute bInterval for a 4 ms interval at full speed and at high speed, then compute what each value would mean if used at the other speed. Confirm your high-speed answer against the root hub example in §1.

2. Take the §13 measurement further: instrument the randomised phase to report the full distribution of measured latencies, not just the maximum. Determine the smallest delivery probability for which an age of 255 becomes reachable within 3000 steps, and say whether that stimulus is still representative.

3. Implement A4 from §15 as a procedural check and confirm it fires under L3 — then show that it stops firing once a genuine miss has occurred, and propose a bench structure that keeps it checkable.

4. Add a second endpoint with a different interval and a shared frame counter. Determine whether one monitor per endpoint is required or whether the state can be shared, and justify the answer from §5 rather than from area.

5. Implement the report_taken-means-attempted defect from §17 and measure it. Note that the present bench cannot produce a NAK at all, and specify precisely what must be added to the stimulus before the defect is detectable.

6. Work out what the else if chain in §7 produces under the two incorrect orderings named there — arrival ahead of delivery, and tick ahead of arrival — and give a concrete stimulus that distinguishes each from the correct design.

20. Summary

bInterval is one byte with two encodings — linear frames at low and full speed, logarithmic microframes at high speed and above (§1) — and the kernel's own root hub descriptors verify the arithmetic: 0xff is 255 ms at full speed and 0x0c is 256 ms at high speed.

It is a request, not a contract (§2). The host may poll more frequently than asked, controllers clamp it to their own limits, and the granted interval is written back as a different number from the requested one. The only guarantee is a bound on the gap.

That bound covers one term of seven (§3). Two of the remaining terms are on the host side and unbounded under load, and one — the display frame — is larger at 60 Hz than the entire polling interval of a fast endpoint. This is why a 1000 Hz mouse is a real but modest improvement (§4), and why its most valuable effect is arguably the reduction in delta clipping rather than the latency itself.

Measuring latency in hardware has one rule that dominates the rest (§5): the measurement is the age of the oldest undelivered item. The newest-item version fails in the direction that hides congestion, and mutation L1 costs 7320 failures — more than twice any other mutation here.

Six mutations, all killed in both languages (§12) — but two of them were killed entirely by directed tests, and instrumenting the randomised phase showed why (§13): in 3000 steps the greatest age reached was 31 against the 255 saturation needs, a region the stimulus distribution reaches with probability around 10^−17. No amount of additional running would have found it.

And a sticky flag cannot be falsified twice (§13). Once a genuine deadline miss set deadline_miss, the off-by-one mutation on the boundary comparison became undetectable — its two failures both occurred at the single directed moment before the first real violation. Saturating evidence is most testable at the beginning.

Verilog and SystemVerilog compiled and simulated clean; VHDL was written and reviewed but not executed (§21), and no claim to the contrary is made anywhere above.

21. Tooling, Honestly

LanguageDesignTestbenchCompiledSimulatedMutations
Verilog-2005usb_latency_monitorlat_v_tb.v✅ Icarus -g2005✅ 0 errors✅ all six
SystemVerilogusb_latency_monitor_svlat_sv_tb.sv✅ Icarus -g2012✅ 0 errors✅ all six
VHDLusb_latency_monitor_vhdllat_vhdl_tb.vhd❌ tool unavailable❌ tool unavailable❌ not run
SVA (§15)——❌ unsupported by Icarus❌—
unique case——✅ accepted⚠️ quality ignored—

No VHDL analyser or simulator exists in this environment — ghdl, nvc, vcom, xvhdl and vsim were each checked — and installing one was out of scope. The VHDL was instead reviewed structurally against the properties a compiler would check: entity and architecture pairing, positional port association matching the declaration in count and order, every input driven and every output checked, numeric_std with no deprecated arithmetic package, an asynchronous-reset process sensitive to both clk and rst_n, a reset value for every registered signal, and the elsif chain ordered collision-before-delivery-before-arrival.

The unique case row is the subtler entry. Icarus accepts the keyword and reports sorry: Case unique/unique0 qualities are ignored — so on this tool the claim is documentation, not a check. A keyword that is accepted and not enforced is worse than one that is rejected, because it reads in review as a verified property.

22. What Comes Next

Module 15 built the transfer type that buys a bound on the gap and pays for it with a small, capped share of bandwidth — and this chapter showed how narrow that bound really is.

Module 16 — Isochronous Transfers — is the other periodic transfer type, and it makes the opposite trade. Isochronous transfers reserve bandwidth rather than buying a latency bound, and they pay for it by giving up something interrupt transfers never surrender: the retry. An isochronous packet that arrives corrupted is not resent, because by the time a retry could complete, the moment the data belonged to has passed.

That single difference reorganises everything — error handling, buffering, the meaning of a lost packet, and what the hardware must do when data does not arrive on time.

Browse the full path on the USB tutorials index.

Continue learning

Standards & specifications

Governing standard
USB-IF (Universal Serial Bus Specification)(opens USB Implementers Forum (USB-IF) in a new tab)

Defines the USB bus — its electrical signalling, connectors, packet and transaction model, device framework and the descriptors a device must expose — together with the device-class specifications layered on it. It does not define host-controller register interfaces (xHCI and EHCI are separate documents) nor any operating system's driver architecture.

This page also covers RTL structure, verification approach and debugging technique. Those are engineering practice built on the standard, not requirements the standard itself imposes.

Where this fits

Part of the USB curriculum.