Skip to content
VLSI Mentor

USB · Module 16

Bandwidth Reservation

An isochronous endpoint reserves time on the wire, not bytes — and worst-case bit stuffing inflates every payload by a sixth. The host's admission arithmetic reproduced exactly, verified over its entire input domain.

The last three chapters all began after a decision had already been made. 16.1's feedback loop, 16.2's frame assembly and 16.3's sequence arithmetic each assume the endpoint gets its slot, every service interval, for as long as the configuration lasts.

This chapter is where that assumption is granted or refused.

And the refusal is absolute. An isochronous endpoint does not negotiate at run time and does not degrade gracefully: the host computes what the endpoint would cost, compares it against what remains, and either admits it or fails the configuration. There is no reduced rate, no best-effort fallback, no retry later. A device whose reservation cannot be met does not run slowly. It does not run.

1. What Is Actually Reserved

The first correction is the one that makes every later number surprising.

The reservation is time on the wire, not bytes in a buffer. A 1024-byte payload does not consume "1024 bytes" of some budget — it consumes the microseconds its transaction occupies, and a transaction is a great deal more than its payload.

ContributorWhat it is
token packetthe host's IN or OUT, with its own sync, PID, address, endpoint and CRC5
data packet overheadsync, PID, and a 16-bit CRC around the payload
inter-packet gapsthe bus turnaround the transaction requires
host delaythe controller's own fixed cost per transaction
bit stuffingthe one people forget — see §2

The Linux host stack computes this explicitly, because periodic transfers are the only ones that must be scheduled in software, and it cannot schedule what it has not costed:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
 * usb_calc_bus_time - approximate periodic transaction time in nanoseconds
 ...
 * See USB 2.0 spec section 5.11.3; only periodic transfers need to be
 * scheduled in software, this function is only used for such scheduling.

The answer is in nanoseconds. Not bytes, not packets — nanoseconds of bus occupancy, which is the resource actually being divided up.

2. Bit Stuffing Inflates Every Payload by a Sixth

USB encodes data as NRZI and guarantees transitions by inserting a stuffed bit after six consecutive ones. The receiver removes it, so it never appears in the data — but it absolutely appears in the time.

The kernel names the worst case directly:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
#define BitTime(bytecount) (7 * 8 * bytecount / 6) /* with integer truncation */
		/* Trying not to use worst-case bit-stuffing
		 * of (7/6 * 8 * bytecount) = 9.33 * bytecount */
		/* bytecount = data payload byte count */

Seven bit times for every six data bits. A payload of all-ones is the pathological case, and scheduling must assume it, because the schedule is built before the data is known.

PayloadData bitsBitTimeInflation
64 bytes512597+16.60%
512 bytes40964778+16.65%
1024 bytes81929557+16.66%

Every isochronous reservation is a sixth larger than its payload suggests, before a single byte of protocol overhead is added. A design that budgets in bytes and multiplies by the bit rate will under-reserve by about 17% and will disagree with the host about what fits.

Note the comment's own caveat. The kernel writes 7 * 8 * b / 6 with integer truncation rather than the true worst case of 9.33 bytes per byte — it is approximating, slightly optimistically, and saying so. §5 is about why a device controller must reproduce that approximation exactly rather than improve on it.

3. The Budget, and Where 80% Comes From

Periodic transfers may not consume the whole frame. Control transfers must always be able to get through — a host that cannot issue a control transfer cannot recover a device — so the specification caps periodic allocation and the kernel encodes the cap:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
#define FRAME_TIME_BITS			12000L	/* frame = 1 millisecond */
#define FRAME_TIME_MAX_BITS_ALLOC	(90L * FRAME_TIME_BITS / 100L)
#define FRAME_TIME_MAX_USECS_ALLOC	(90L * FRAME_TIME_USECS / 100L)

At full speed: 12 000 bit times in a 1 ms frame, of which 90% — 10 800 bits, or 900 µs — may be allocated to periodic transfers. The remaining 10% is reserved for control and whatever bulk can be fitted in.

At high speed the service interval is the 125 µs microframe and the cap is 80%, which is 100 µs of every 125. The tighter fraction reflects how much more traffic a microframe can hold and how much more control traffic a high-speed bus carries.

These caps are the entire reason admission control exists. Without a cap the arithmetic would be trivial — accept everything until the frame is full — and a single greedy device could make the bus unrecoverable.

4. The Numbers, Computed Rather Than Recalled

Applying the kernel's own formulas gives figures that are worth internalising, because they are far less generous than "480 Mb/s" suggests:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
  HS_NSECS_ISO(b) = ((38*8*2083) + 2083*(3 + BitTime(b))) / 1000 + 5
PayloadBus timeFraction of a 125 µs microframe
64 bytes1 888 ns1.51%
256 bytes5 620 ns4.50%
512 bytes10 597 ns8.48%
1024 bytes20 551 ns16.44%

And with the high-bandwidth multiplier from Chapter 16.2 §1:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
  one 1024-byte transaction  =  20.551 us
  mult = 3 (high bandwidth)  =  61.653 us
  the 80% budget             = 100.000 us

A single high-bandwidth isochronous endpoint consumes 61.7% of the entire periodic budget of every microframe. Two of them do not fit. That is the real constraint behind "why can't I run two 1080p webcams on one controller", and it is arithmetic, not a driver limitation.

Full speed is starker still. A 1023-byte full-speed isochronous IN transaction costs 806 µs of the 900 µs budget:

PayloadBus timeFraction of the 900 µs budget
8 bytes14.71 µs1.6%
64 bytes58.40 µs6.5%
512 bytes407.68 µs45.3%
1023 bytes806.17 µs89.6%

Exactly one maximum-size full-speed isochronous endpoint fits in a frame, with 94 µs left over — not enough for a second of any consequence.

5. Why the Device Must Reproduce the Host's Arithmetic Exactly

A device controller that tracks its own bandwidth — to decide which alternate settings it can offer, or to validate its own descriptors before shipping — faces a requirement that is easy to underestimate.

It must compute the same number the host computes, including the host's rounding.

Look again at what §2 quoted: 7 * 8 * bytecount / 6, with integer truncation, and the acknowledgement that this is deliberately not the true worst case. And §4's formula divides by 1000 with truncation again. Neither rounding is the mathematically natural one, and both are load-bearing.

If the device…Then
rounds up where the host truncatesit computes a larger cost, refuses configurations the host would accept
uses the true 9.33 stuffing factorsame — it is more correct and less compatible
rounds down anywhere the host does notit offers configurations the host rejects, and the failure appears at SET_CONFIGURATION

Being more accurate than the host is a bug. The specification's arithmetic is a protocol, not an estimate to improve on: the two sides must agree, and only one of them gets to define what agreement means.

And the disagreement is silent until it is fatal. A device that computes a slightly smaller cost enumerates, advertises an alternate setting, and fails only when a host with a nearly-full schedule refuses the configuration — which is to say, on someone else's machine, with another device plugged in.

6. The Hardware, Before Any Language

Two divisions, neither by a power of two. BitTime divides by 6; the nanosecond conversion divides by 1000. A hardware divider for either would be large and slow, and this block runs once per configuration rather than once per packet.

The standard answer is reciprocal multiplication: replace x / d with (x × M) >> S, choosing M = ceil(2^S / d) and S large enough that the result is exact for every input in the legal range.

Exact, not approximately right. §5 says the device must match the host bit for bit, so "close enough" is a defect. The constants used here are:

DivisionMSExact over
x / 643 69118x in 0 … 57 344
y / 10008 589 93533y in 0 … 20 546 712

Those ranges are not chosen for comfort — they are the full reachable input domain. The payload is at most 1024 bytes, so x = 56 × bytes is at most 57 344, and y follows. §10 verifies the whole pipeline over every point in that domain rather than sampling it.

The rest of the design is accounting:

State retained: a running total of granted reservations, and the registered decision outputs.

On reset or bus reset: the total clears. A bus reset ends the configuration, so every reservation made for it is void — carrying them forward would refuse a re-enumerating device its own bandwidth.

On clear: the total clears without a decision. This is the host beginning a new configuration.

On a request: compute the cost, compare used + cost against the budget, and admit or reject. The comparison is inclusive — a reservation that exactly fills the budget is legal, because the budget is what may be used, not what must be left over.

A rejected request reserves nothing. There is no partial admission (§0), so the running total is untouched and the next request sees the same budget.

A vertical pipeline of six stages showing how an isochronous bandwidth reservation is decided. The first stage is the endpoint descriptor's wMaxPacketSize field, which carries the payload size in bits ten down to zero and the transactions-per-microframe multiplier in bits twelve and eleven. The second stage computes BitTime, multiplying the payload by seven eighths over six to account for worst-case bit stuffing, which inflates every payload by one sixth; this stage divides by six and must truncate exactly as the host does. The third stage computes the per-transaction bus time in nanoseconds, adding the token packet, the data packet identifier and cyclic redundancy check, the inter-packet gaps and the host controller's own delay, then dividing by one thousand with the same truncation requirement. The fourth stage multiplies by the number of transactions per microframe, which is the multiplier field plus one. The fifth stage compares the running total of already-granted reservations plus this cost against eighty percent of the one hundred and twenty five microsecond microframe. The sixth and final stage issues the decision, which is either admit, adding the cost to the running total, or reject, which reserves nothing at all because there is no partial admission.wMaxPacketSizebits 10:0 payload bytes · bits 12:11 transactions per microframeBitTime = 7 × 8 × bytes ÷ 6worst-case bit stuffing — every payload grows by a sixthper-transaction bus time (ns)token · PID · CRC · gaps · host delay, then ÷ 1000× (mult + 1)up to three transactions per microframeused + cost ≤ 80% of 125 µs100 000 ns — the rest is kept for control trafficADMIT or REJECTgranted whole, or not at allpayload bytesbit timesns per transactionns per microframefits / does not fit12
Figure 1 — the admission pipeline, top to bottom. Each stage answers one question, and the two divisions in the middle are the stages that must reproduce the host's truncation exactly rather than round naturally.

The two orange stages are the ones §5 is about. Both divide, both truncate, and both must truncate the way the host truncates rather than the way the arithmetic would prefer.

7. Verilog

The RTL contract

  • What it models: admission control for high-speed isochronous bandwidth reservations.
  • Why it exists: because the reservation is time, not bytes (§1), bit stuffing inflates it by a sixth (§2), and the device must agree with the host exactly (§5).
  • Inputs: req_valid with req_bytes and req_mult, clear, bus_reset.
  • State retained: used_r (granted nanoseconds), and the registered decision outputs.
  • Outputs: cost_ns, ack, admit, reject, used_ns.
  • Hardware implied: one 16×16 multiplier, one 48-bit multiplier, three adders, one comparator, one 18-bit accumulator. The 48-bit multiplier is the area cost — acceptable here because this block runs once per configuration, not once per packet.
  • Reset: asynchronous active-low rst_n; bus_reset synchronous and equivalent, and both void every reservation.
  • Priority: clear outranks req_valid — a host beginning a new configuration is not also making a request.
  • Latency: one cycle; ack and the decision are registered.
  • Boundaries: the comparison is inclusive at the budget; a rejected request accumulates nothing.
  • Simultaneous events: none — one request per cycle.
  • Assumptions: req_bytes is 0…1024 and req_mult is 0…2, as wMaxPacketSize encodes them; the reciprocal constants are exact only over that domain, which §10 verifies exhaustively.
  • Omissions: high speed only, no full/low-speed formula, no split transactions, no cross-frame scheduling.
  • What DV should verify: that the cost matches the host's arithmetic for every legal input; that the budget boundary is inclusive; that a rejected request reserves nothing; that reset and clear both release everything.
Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// iso_bw_admission -- decides whether a high-speed isochronous endpoint's
// bandwidth reservation can be granted.
//
// What is reserved is TIME ON THE WIRE, not bytes in a buffer. A 1024-byte
// payload does not cost 1024 bytes of budget: it costs the token, the data
// PID, the CRC, the inter-packet gaps, the host delay, and -- the term people
// forget -- WORST-CASE BIT STUFFING, which inflates every payload by 7/6.
//
// The arithmetic below reproduces the host's own calculation exactly,
// including its integer truncation. That exactness is the point: a device
// controller that rounds differently from the host will disagree about
// whether a configuration fits, and the two will disagree in silence.
//
//   BitTime(b)      = 7 * 8 * b / 6            (truncating)
//   HS_NSECS_ISO(b) = ((38*8*2083) + 2083*(3 + BitTime(b))) / 1000 + 5
//
// Neither division is by a power of two, so each is done by multiplying by a
// reciprocal and shifting. The constants below are not approximations: they
// are EXACT for every input in the legal range, which the testbench proves
// exhaustively rather than by sampling.
module iso_bw_admission #(
  // 80% of a 125 us microframe, in nanoseconds, is what USB 2.0 allows
  // periodic transfers to reserve at high speed.
  parameter integer BUDGET_NS = 100000
) (
  input  wire        clk,
  input  wire        rst_n,
  input  wire        bus_reset,
  input  wire        clear,          // drop all reservations (reconfigure)
  input  wire        req_valid,      // a reservation request
  input  wire [10:0] req_bytes,      // wMaxPacketSize payload bytes, 0..1024
  input  wire [1:0]  req_mult,       // transactions per microframe minus one
  output wire [15:0] cost_ns,        // this request's cost
  output wire        ack,            // one cycle, the decision is valid
  output wire        admit,
  output wire        reject,
  output wire [17:0] used_ns         // running total of granted reservations
);
  // Reciprocal constants. EXACT over the legal input range -- see the
  // exhaustive check in the testbench, not a spot check.
  localparam [31:0] M6    = 32'd43691;      // 2**18 / 6,    shift 18
  localparam [47:0] M1000 = 48'd8589935;    // 2**33 / 1000, shift 33
  localparam [24:0] K38   = 25'd633232;     // 38 * 8 * 2083

  // ---- the cost pipeline, combinational ----
  // 56 * bytes, which is 7 * 8 * bytes before the divide by 6.
  wire [15:0] x   = {5'd0, req_bytes} * 16'd56;
  wire [31:0] p6  = {16'd0, x} * M6;
  wire [13:0] bt  = p6[31:18];              // BitTime(bytes), truncating

  // 2083 * (3 + BitTime) + 38*8*2083
  wire [24:0] y   = K38 + ({11'd0, bt} + 25'd3) * 25'd2083;
  wire [47:0] p1k = {23'd0, y} * M1000;
  wire [15:0] one_ns = p1k[47:33] + 16'd5;  // HS_NSECS_ISO(bytes)

  // mult is transactions-per-microframe MINUS ONE, exactly as the
  // wMaxPacketSize field encodes it. Forgetting the +1 under-reserves by a
  // third at mult=2, which is mutation B3.
  wire [17:0] cost = ({16'd0, req_mult} + 18'd1) * {2'd0, one_ns};

  reg [17:0] used_r;
  reg        ack_r, admit_r, reject_r;
  reg [15:0] cost_r;

  assign cost_ns = cost_r;
  assign ack     = ack_r;
  assign admit   = admit_r;
  assign reject  = reject_r;
  assign used_ns = used_r;

  // The decision. `<=` because a reservation that exactly fills the budget
  // is legal -- the budget is what MAY be used, not what must be left over.
  wire [18:0] would_use = {1'b0, used_r} + {1'b0, cost};
  wire        fits      = (would_use <= {1'b0, BUDGET_NS[17:0]});

  always @(posedge clk or negedge rst_n) begin
    if (!rst_n) begin
      used_r <= 18'd0; ack_r <= 1'b0; admit_r <= 1'b0; reject_r <= 1'b0;
      cost_r <= 16'd0;
    end else if (bus_reset) begin
      // A bus reset ends the configuration, so every reservation made for
      // it is void. Carrying them forward would refuse a device its own
      // bandwidth after a re-enumeration.
      used_r <= 18'd0; ack_r <= 1'b0; admit_r <= 1'b0; reject_r <= 1'b0;
      cost_r <= 16'd0;
    end else begin
      ack_r    <= 1'b0;
      admit_r  <= 1'b0;
      reject_r <= 1'b0;

      if (clear) begin
        used_r <= 18'd0;
      end else if (req_valid) begin
        ack_r  <= 1'b1;
        cost_r <= cost[15:0];
        if (fits) begin
          admit_r <= 1'b1;
          used_r  <= would_use[17:0];
        end else begin
          // A rejected reservation consumes NOTHING. There is no partial
          // admission: an isochronous endpoint that cannot have its full
          // reservation does not run at a reduced rate, it does not run.
          reject_r <= 1'b1;
        end
      end
    end
  end
endmodule

Three details are worth naming.

bt = p6[31:18] is the division. There is no divider. The shift is a bit-select, which is wiring — the entire cost of dividing by six is the 16×16 multiply that precedes it, and a synthesis tool will reduce that further because one operand is constant.

The widths are stated at every step and they are not obvious. x is 16 bits, p6 is 32, y is 25, p1k is 48, one_ns is 16. The intermediate is three times the width of the result, and letting any of those be inferred is precisely how a reciprocal multiplication stops being exact — the truncation moves, and the answer changes by one in a way no test that samples inputs is likely to find.

req_mult is transactions-per-microframe minus one. That is how wMaxPacketSize encodes it, so the design adds one rather than inventing a friendlier convention. §11's mutation B3 drops the multiplier entirely and under-reserves by up to two thirds.

8. SystemVerilog

Same arithmetic, with the request and the decision given types.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
package iso_bw_pkg;
  // A reservation request as one object. The two fields come from a single
  // descriptor field -- wMaxPacketSize packs the payload size in bits 10:0
  // and the transactions-per-microframe multiplier in bits 12:11 -- so
  // keeping them together in the type mirrors where they came from.
  typedef struct packed {
    logic [1:0]  mult;    // transactions per microframe MINUS ONE
    logic [10:0] bytes;   // payload bytes, 0..1024
  } iso_req_t;

  // The decision. Naming it removes any possibility of a state in which
  // both admit and reject are asserted, which two independent flags allow.
  typedef enum logic [1:0] {
    D_IDLE,
    D_ADMIT,
    D_REJECT
  } decision_e;
endpackage

module iso_bw_admission_sv
  import iso_bw_pkg::*;
#(
  parameter int unsigned BUDGET_NS = 100000
) (
  input  logic      clk,
  input  logic      rst_n,
  input  logic      bus_reset,
  input  logic      clear,
  input  logic      req_valid,
  input  iso_req_t  req,
  output logic [15:0] cost_ns,
  output logic        ack,
  output decision_e   decision,
  output logic [17:0] used_ns
);
  // Reciprocal constants, EXACT over the legal input range 0..1024 bytes --
  // proven exhaustively by the testbench rather than argued.
  localparam logic [31:0] M6    = 32'd43691;    // 2**18 / 6,    shift 18
  localparam logic [47:0] M1000 = 48'd8589935;  // 2**33 / 1000, shift 33
  localparam logic [24:0] K38   = 25'd633232;   // 38 * 8 * 2083

  initial begin
    if (BUDGET_NS > 125000)
      $fatal(1, "BUDGET_NS=%0d exceeds a 125us microframe", BUDGET_NS);
    if (BUDGET_NS == 0)
      $fatal(1, "BUDGET_NS=0 admits nothing at all");
  end

  // ---- the cost pipeline ----
  // Every width is stated. The intermediate products are much wider than the
  // result, and letting any of them be inferred is how a reciprocal
  // multiplication silently stops being exact.
  logic [15:0] x;
  logic [31:0] p6;
  logic [13:0] bt;
  logic [24:0] y;
  logic [47:0] p1k;
  logic [15:0] one_ns;
  logic [17:0] cost;

  always_comb begin
    x      = 16'(req.bytes) * 16'd56;          // 7 * 8 * bytes, before /6
    p6     = 32'(x) * M6;
    bt     = p6[31:18];                        // BitTime(bytes), truncating
    y      = K38 + (25'(bt) + 25'd3) * 25'd2083;
    p1k    = 48'(y) * M1000;
    one_ns = 16'(p1k[47:33]) + 16'd5;          // HS_NSECS_ISO(bytes)
    // mult is transactions-per-microframe MINUS ONE, exactly as the
    // descriptor encodes it. Dropping the +1 under-reserves by a third.
    cost   = (18'(req.mult) + 18'd1) * 18'(one_ns);
  end

  wire [18:0] would_use = 19'(used_ns) + 19'(cost);
  // `<=` because a reservation that exactly fills the budget is legal: the
  // budget is what MAY be used, not what must be left over.
  wire        fits      = (would_use <= 19'(BUDGET_NS));

  always_ff @(posedge clk or negedge rst_n) begin
    if (!rst_n || bus_reset) begin
      // A bus reset ends the configuration, so every reservation made for it
      // is void -- otherwise a re-enumerating device is refused its own
      // bandwidth because the previous configuration still holds it.
      used_ns <= '0; ack <= 1'b0; decision <= D_IDLE; cost_ns <= '0;
    end else begin
      ack      <= 1'b0;
      decision <= D_IDLE;

      if (clear) begin
        used_ns <= '0;
      end else if (req_valid) begin
        ack     <= 1'b1;
        cost_ns <= cost[15:0];
        if (fits) begin
          decision <= D_ADMIT;
          used_ns  <= would_use[17:0];
        end else begin
          // A rejected reservation consumes NOTHING. There is no partial
          // admission: an isochronous endpoint that cannot have its full
          // reservation does not run slowly, it does not run.
          decision <= D_REJECT;
        end
      end
    end
  end
endmodule

decision_e removes a state the Verilog permits. Two independent flags, admit and reject, can in principle both be high — nothing in the Verilog's structure forbids it, only the code's discipline does. A single enumerated output makes the illegal state unrepresentable, which is a stronger guarantee than an assertion that it never occurs.

iso_req_t mirrors where the fields came from. Payload size and multiplier are not two separate parameters a driver happens to pass together — they are two fields of one descriptor word, wMaxPacketSize, and keeping them in one type records that.

Every width is an explicit cast. 48'(y) * M1000, 19'(used_ns) + 19'(cost). §7 explains why: the exactness of the reciprocal multiplication depends on the product being computed at full width, and an inferred context is not a guarantee.

9. VHDL

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

package iso_bw_pkg is
  -- A reservation request as one object. Both fields come from a single
  -- descriptor field: wMaxPacketSize packs the payload size in bits 10:0 and
  -- the transactions-per-microframe multiplier in bits 12:11.
  type iso_req_t is record
    mult  : unsigned(1 downto 0);    -- transactions per microframe MINUS ONE
    bytes : unsigned(10 downto 0);   -- payload bytes, 0..1024
  end record;

  -- The decision as a distinct type. Two independent flags would permit a
  -- state in which both admit and reject are asserted; this cannot.
  type decision_t is (D_IDLE, D_ADMIT, D_REJECT);
end package;

library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
use work.iso_bw_pkg.all;

entity iso_bw_admission_vhdl is
  generic (
    BUDGET_NS : positive := 100000   -- 80% of a 125 us microframe
  );
  port (
    clk       : in  std_logic;
    rst_n     : in  std_logic;
    bus_reset : in  std_logic;
    clear     : in  std_logic;
    req_valid : in  std_logic;
    req       : in  iso_req_t;
    cost_ns   : out unsigned(15 downto 0);
    ack       : out std_logic;
    decision  : out decision_t;
    used_ns   : out unsigned(17 downto 0)
  );
end entity;

architecture rtl of iso_bw_admission_vhdl is
  -- Reciprocal constants, EXACT over the legal input range 0..1024 bytes.
  -- numeric_std does define "/" for unsigned, and it would simulate
  -- correctly -- but the reciprocal form is what the hardware actually is,
  -- and writing it explicitly means its exactness is VERIFIED (see the
  -- exhaustive check in the testbench) rather than trusted to a synthesis
  -- tool's own choice of constant.
  constant M6    : unsigned(15 downto 0) := to_unsigned(43691, 16);
  constant M1000 : unsigned(23 downto 0) := to_unsigned(8589935, 24);
  constant K38   : unsigned(24 downto 0) := to_unsigned(633232, 25);

  signal x      : unsigned(15 downto 0);
  signal p6     : unsigned(31 downto 0);
  signal bt     : unsigned(13 downto 0);
  signal y      : unsigned(24 downto 0);
  signal p1k    : unsigned(48 downto 0);
  signal one_ns : unsigned(15 downto 0);
  signal cost   : unsigned(17 downto 0);

  signal used_r : unsigned(17 downto 0) := (others => '0');
  signal cost_r : unsigned(15 downto 0) := (others => '0');
  signal ack_r  : std_logic := '0';
  signal dec_r  : decision_t := D_IDLE;

  signal would_use : unsigned(18 downto 0);
  signal fits      : boolean;
begin
  assert BUDGET_NS <= 125000
    report "BUDGET_NS exceeds a 125 us microframe" severity failure;

  cost_ns  <= cost_r;
  ack      <= ack_r;
  decision <= dec_r;
  used_ns  <= used_r;

  -- The cost pipeline. Every width is stated: resize() before each multiply
  -- so the product width is what the arithmetic needs, not what an inferred
  -- context happens to give.
  x   <= resize(resize(req.bytes, 16) * to_unsigned(56, 16), 16);
  p6  <= x * M6;
  bt  <= p6(31 downto 18);                        -- BitTime, truncating
  y   <= K38 + resize((resize(bt, 25) + to_unsigned(3, 25))
                      * to_unsigned(2083, 25), 25);
  p1k <= resize(y * M1000, 49);
  one_ns <= resize(p1k(48 downto 33), 16) + to_unsigned(5, 16);

  -- mult is transactions-per-microframe MINUS ONE, exactly as the descriptor
  -- encodes it. Dropping the +1 under-reserves by a third at mult = 2.
  cost <= resize((resize(req.mult, 18) + to_unsigned(1, 18))
                 * resize(one_ns, 18), 18);

  would_use <= resize(used_r, 19) + resize(cost, 19);
  -- "<=" because a reservation that exactly fills the budget is legal: the
  -- budget is what MAY be used, not what must be left over.
  fits <= would_use <= to_unsigned(BUDGET_NS, 19);

  process (clk, rst_n)
  begin
    if rst_n = '0' then
      used_r <= (others => '0'); cost_r <= (others => '0');
      ack_r <= '0'; dec_r <= D_IDLE;
    elsif rising_edge(clk) then
      if bus_reset = '1' then
        -- A bus reset ends the configuration, so every reservation made for
        -- it is void. Carrying them forward would refuse a re-enumerating
        -- device its own bandwidth.
        used_r <= (others => '0'); cost_r <= (others => '0');
        ack_r <= '0'; dec_r <= D_IDLE;
      else
        ack_r <= '0';
        dec_r <= D_IDLE;

        if clear = '1' then
          used_r <= (others => '0');
        elsif req_valid = '1' then
          ack_r  <= '1';
          cost_r <= resize(cost, 16);
          if fits then
            dec_r  <= D_ADMIT;
            used_r <= resize(would_use, 18);
          else
            -- A rejected reservation consumes NOTHING. There is no partial
            -- admission: an endpoint that cannot have its full reservation
            -- does not run slowly, it does not run.
            dec_r <= D_REJECT;
          end if;
        end if;
      end if;
    end if;
  end process;
end architecture;

numeric_std defines / for unsigned, and this design does not use it. That is a deliberate choice and worth stating plainly: y / 1000 would simulate perfectly in all three languages. The reciprocal form is written out because it is what the hardware will actually be — and because writing it explicitly means its exactness is something the testbench proves (§10) rather than something a synthesis tool is trusted to get right when it chooses its own constant.

resize before every multiply, for the reason §7 gives. In VHDL the product of an unsigned(24 downto 0) and an unsigned(23 downto 0) is 49 bits by definition of the operator — the width is a property of the types, not of the assignment context. That is a genuinely stronger position than the other two languages, where the author must assert the width and the tool must be persuaded to honour it.

decision_t is a distinct type with no encoding, so — as in 16.3 — there is no possibility of comparing it against an integer, and the case that consumes it must be exhaustive.

10. Comparing the Three

ConcernVerilogSystemVerilogVHDL
The two divisionsreciprocal multiply + bit-selectsame, with explicit castssame; / exists and is deliberately unused
Product widthauthor asserts itexplicit 48'() castsa property of the operand types
The decisiontwo independent flagsdecision_e — illegal state unrepresentabledecision_t, no encoding
The requesttwo portsstruct packedrecord
Illegal parameterisationundetectedtwo $fatal guardsone assert ... severity failure

All three compute identical numbers — §10's exhaustive check confirms it at every one of 3075 input points in each language — and the differences are entirely about how much the author has to assert versus how much the language guarantees.

11. The Testbenches, and an Exhaustive Proof

The independence principle here is unusually clean. The design divides by multiplying by a reciprocal and shifting. Every bench divides using the language's own integer division operator.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
  function automatic int f_bittime(input int b);  return (7*8*b)/6; endfunction
  function automatic int f_one_ns(input int b);
    return ((38*8*2083) + 2083*(3 + f_bittime(b)))/1000 + 5;
  endfunction

The two agree only if the reciprocal constants are exact. They are not approximately equivalent methods that should roughly match — they are a hardware shortcut and the arithmetic it is shortcutting, and any inexactness anywhere in the domain shows up as a mismatch.

And the domain is small enough to check completely. wMaxPacketSize allows 0…1024 payload bytes and the multiplier is 0…2, so the entire legal input space is 3075 points:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
    for (int m=0; m<3; m++)
      for (int i=0; i<=1024; i++) begin
        do_clear();
        do_req(i, m);
        n_exhaustive++;
      end
Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
  exhaustive cost check: 3075 of 3075 input points verified

The directed sequence then covers the accounting, which the exhaustive check deliberately does not — every exhaustive point clears the budget first, so the accumulator is never under test there:

ScenarioWhat it pins down
the three headline figures20 551 ns, 61 653 ns, 1 888 ns against §4
four 1024-byte endpointsaccumulate to 82 204 ns and all fit
a fifthrejected, and the total is unchanged
the largest payload that fits in the remainderadmitted
one byte morerejected
a total landing exactly on 100 000 nsadmitted — see §13
anything after that, however smallrejected
clear, and bus_resetboth release every reservation
4000 randomised requestswith periodic clears, so the budget spends time both nearly empty and nearly full
Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
  REACH: admitted=4621 rejected=2476 exhaustive points=3075

Both outcomes are well represented. A randomiser that never filled the budget would exercise only the admit path, and mutation B5 — which corrupts the accumulator — would be far harder to see.

12. Mutation Testing — Across All Three Languages

IDMutationVerilogSystemVerilogVHDLKilled
—baseline, no mutation000—
B1no worst-case bit stuffing (×8 instead of ×7×8÷6)150361460914610✅ all three
B2budget is 100% of a microframe, not 80%369031973117✅ all three
B3the mult field is ignored122861145011882✅ all three
B4budget boundary exclusive instead of inclusive977✅ all three — but see §13
B5an admitted reservation is not accumulated1205595779479✅ all three
B6reciprocal shift off by one158301500614961✅ all three

B1 and B6 are the largest, and both are §5's failure mode. Each makes the device compute a different number from the host — B1 by omitting bit stuffing, B6 by moving the truncation point by one bit — and each is caught at essentially every one of the 3075 exhaustive points. That is what an exhaustive check buys: a defect in the arithmetic cannot hide in an untested corner, because there are no untested corners.

B2 is smaller at ~3200 because it does not change any cost — only the threshold. It is detectable exactly when a request lands between the 80% and 100% marks, which the randomiser produces often but not always.

The VHDL counts differ by a few percent because uniform and $random produce different random sequences after the identical directed and exhaustive phases. The exhaustive portion contributes identically in all three, which is why B1 and B6 are so close across languages while B2 and B5 — whose detection depends on the random accumulation pattern — vary more.

13. The Mutation That Survived

B4 initially survived in all three languages. Zero failures.

It replaces the inclusive budget comparison with an exclusive one — used + cost < BUDGET instead of <= — so it differs from the correct design only when a reservation total lands exactly on 100 000 ns.

The first question is whether the mutation is equivalent, and it is not: the behaviours genuinely differ, on an input the stimulus had simply never produced. §10's directed tests included the largest payload that fits and one byte more, which bracket the boundary without landing on it, and the random phase had no reason to hit an exact value out of a 100 000 ns range.

The second question is whether the boundary is reachable at all. The costs are quantised — every reservation is one of 2907 distinct achievable values between 644 and 61 653 ns — so it is entirely possible for a target to be unreachable. Searching the achievable sums:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
  distinct achievable single-request costs: 2907   (range 644 .. 61653)
  no single request equals the budget exactly
  two requests summing to exactly 100000: None
  three requests: (0 bytes, mult=0) + (947 bytes, mult=1) + (1017 bytes, mult=2)

Verifying that combination:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
  (bytes=0,    mult=0) ->    644 ns
  (bytes=947,  mult=1) ->  38108 ns
  (bytes=1017, mult=2) ->  61248 ns
  total = 100000 ns   (budget = 100000)   exact

Adding those three requests as a directed test killed B4 in all three languages — 9, 7 and 7 failures — and the baseline stayed clean.

14. Assertions

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
  // A1. SAFETY, and the property §5 is about: the cost equals the host's own
  //     arithmetic, computed INDEPENDENTLY in the bench with real division.
  property p_cost_matches_host;
    @(posedge clk) disable iff (!rst_n || bus_reset)
      ack |-> (cost_ns == $past(host_cost_ns));
  endproperty
  a_cost_matches_host: assert property (p_cost_matches_host);

  // A2. SAFETY: the running total never exceeds the budget. This is the
  //     invariant the whole block exists to maintain.
  property p_never_over_budget;
    @(posedge clk) disable iff (!rst_n)
      (used_ns <= BUDGET_NS);
  endproperty
  a_never_over_budget: assert property (p_never_over_budget);

  // A3. SAFETY: a rejected request reserves NOTHING. There is no partial
  //     admission -- this is mutation B5's inverse, and it must hold on the
  //     reject path as strictly as the accumulation holds on the admit path.
  property p_reject_reserves_nothing;
    @(posedge clk) disable iff (!rst_n || bus_reset)
      reject |=> $stable(used_ns);
  endproperty
  a_reject_reserves_nothing: assert property (p_reject_reserves_nothing);

  // A4. SAFETY: admit and reject are mutually exclusive, and neither occurs
  //     without an ack. In the SystemVerilog this is structural -- the
  //     decision is one enumerated value -- so this property is really a
  //     check on the VERILOG's two-flag encoding.
  property p_decision_consistent;
    @(posedge clk) disable iff (!rst_n)
      $onehot0({admit, reject}) && ((admit || reject) |-> ack);
  endproperty
  a_decision_consistent: assert property (p_decision_consistent);

  // A5. SAFETY: the total moves ONLY by the cost just published, and only on
  //     an admit. Stated as an equality so it catches both a missing update
  //     (B5) and an update by the wrong amount.
  property p_total_moves_by_cost;
    @(posedge clk) disable iff (!rst_n || bus_reset)
      admit |=> (used_ns == $past(used_ns) + $past(cost));
  endproperty
  a_total_moves_by_cost: assert property (p_total_moves_by_cost);

  // A6. PROGRESS: a request always produces a decision. A block that simply
  //     never acknowledged would satisfy every safety property above.
  property p_request_is_answered;
    @(posedge clk) disable iff (!rst_n || bus_reset)
      (req_valid && !clear) |=> ack;
  endproperty
  a_request_is_answered: assert property (p_request_is_answered);

Assertion contracts

ClaimSafety / progressVacuity riskHow non-vacuity is established
A1the cost matches the host's arithmeticsafetylow3075 exhaustive points plus 4000 random
A2the total never exceeds the budgetsafetynone — no antecedentholds every cycle
A3a rejected request reserves nothingsafetymoderate2476 rejections measured
A4the decision is consistentsafetynone for the $onehot0 halfholds every cycle
A5the total moves by exactly the costsafetylow4621 admissions measured
A6every request is answeredprogresslowevery directed and random request

A6 is the one that a purely safety-driven review omits, and §30's argument applies exactly: a design that never asserted ack, never admitted anything, and left used_ns at zero for ever would satisfy A1 through A5 perfectly. It would also make every isochronous device on the bus fail to configure.

A5 is deliberately an equality rather than an inequality. The total never decreases would be true of a design that added the wrong amount; the total increased by the published cost catches both B5 (no update) and an update by a wrong value, and it is no harder to write.

A4 is interesting for what it says about §9's comparison table. In the SystemVerilog and VHDL the decision is a single enumerated value, so $onehot0({admit, reject}) is structurally guaranteed and the assertion is redundant. In the Verilog it is two independent flags and the property is doing real work. The same assertion is load-bearing in one implementation and decorative in another — which is a reasonable argument for the typed encoding, and a reminder that an assertion's value depends on the implementation it guards.

15. Verification: Why This Chapter Does Not Use UVM Either

Chapter 16.2 made the strong case for UVM — a structure spanning hundreds of packets, natural loss injection, genuinely multi-dimensional coverage. This block is the opposite of that, and the reason is §11.

The cost function's input domain is 3075 points and they were all checked. There is nothing for constrained-random to explore, no distribution to tune, and no coverage model that could report more than "all of it". A UVM environment here would randomise over a space that has already been enumerated, which is strictly less informative than the enumeration.

The accumulator is the only part with a sequence to it, and its interesting behaviours are: fill it, overfill it, land exactly on the boundary, clear it. Three of those four are directed tests by nature — §13's exact-boundary case in particular was constructed by search, not found by randomisation, and no constraint solver would have been pointed at it without first knowing the boundary mattered.

The honest rule this module has applied three times now: UVM's cost is justified by scenario complexity. 16.1 declined it for a scalar accumulator, 16.2 adopted it for frame reassembly with loss injection, 16.3 found a narrow case, and this chapter declines it for a block whose function is finite and whose sequencing is four scenarios long. Importance and complexity are different axes, and this block is important and simple.

What a UVM environment would add, if this block lived inside a larger host-controller verification effort, is worth naming precisely: a scheduler-level environment in which many endpoints of many speeds are configured and removed over time, and this admission block is one component. The interesting scenarios there are about ordering and fragmentation — does a bus that admitted A, B and C, then dropped B, still admit D? — which is genuinely a sequencing problem and genuinely beyond directed tests. That environment belongs to Module 17, not here.

16. Debugging: the Webcam That Works Alone and Fails in a Pair

A UVC camera streams 1080p correctly when it is the only device on its controller. Plugged in alongside a second identical camera, the second one fails to configure: the host reports that it cannot allocate bandwidth, and the device never reaches its streaming alternate setting. Both cameras work individually on the same port. Nothing is defective.

Nothing is defective. §4 already contains the answer, and the arithmetic takes thirty seconds.

A 1080p camera almost certainly uses a high-bandwidth isochronous endpoint — mult of 2 or 3 — at or near 1024 bytes. From §4:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
  1024 bytes, mult = 3  ->  61 653 ns
  the 80% budget        -> 100 000 ns
  two of them           -> 123 306 ns

Two do not fit, and it is not close. The second camera is not being refused because of a driver bug or a controller limitation; it is being refused because the arithmetic says no, and the arithmetic is the specification.

The chain, from the outside in:

Descriptors — what is the endpoint actually asking for? Read wMaxPacketSize and decode both fields: bits 10:0 for the payload, bits 12:11 for the multiplier. A device asking for mult = 3 at 1024 bytes has reserved 61.7% of every microframe, and a device that does that unnecessarily is the fault.

Alternate settings — is there a smaller one? UVC devices normally offer several alternate settings with different wMaxPacketSize values precisely so a host can pick one that fits. A device offering only its largest is a firmware defect, and it is the single most common cause of this symptom.

Topology — are both cameras on the same host controller? Bandwidth is reserved per controller, not per port. Moving one camera to a port on a different controller resolves it immediately, and whether that is possible is a property of the machine, not of USB.

The device's own calculation — does it agree with the host's? §5. If the device validates its own descriptors and believes the configuration fits when the host does not, compare the two numbers directly. A device that omitted bit stuffing will under-estimate by about 17% — enough to believe an endpoint fits when it does not, and mutation B1 measures exactly that error.

17. Common Misconceptions

"Bandwidth is reserved in bytes." It is reserved in time (§1) — nanoseconds of bus occupancy, including token, CRC, gaps and host delay.

"A 480 Mb/s bus can carry 60 MB/s of isochronous data." The periodic cap is 80% of each microframe (§3), and per-transaction overhead plus bit stuffing consumes a large part of what remains.

"Bit stuffing is a physical-layer detail that does not affect scheduling." It inflates every reservation by a sixth (§2), and the schedule must assume the worst case because it is built before the data exists.

"A bigger wMaxPacketSize is free if the device usually sends less." The reservation is taken every interval regardless (§4). An isochronous wMaxPacketSize is a commitment, not a ceiling.

"The device should compute bandwidth as accurately as possible." It should compute it identically to the host (§5), including the host's deliberate truncation. Being more accurate is a compatibility bug.

"mult is an optimisation hint." It multiplies the reservation by up to three (§4). Mutation B3 ignores it and under-reserves by two thirds.

"Two 1080p cameras on one controller is a driver limitation." It is arithmetic: 2 × 61 653 ns against a 100 000 ns budget (§16).

"A mutation that survives with zero failures is probably equivalent." §13: B4 was merely unreached, and the boundary it guards decides whether devices configure at all.

18. Exercises

1. Compute the bus time for a 512-byte high-speed isochronous endpoint with mult = 2, and determine how many such endpoints fit in the 80% budget. Then do the same for mult = 1 at 1024 bytes and explain why the answers differ despite equal payload per microframe.

2. Derive the full-speed figure in §4 for 1023 bytes from usb_calc_bus_time's full-speed isochronous branch, and confirm 806 µs. Then determine the largest full-speed isochronous payload for which two endpoints fit in the 900 µs budget.

3. §6 gives reciprocal constants exact over 0…57 344 and 0…20 546 712. Determine what happens to M = 43691, S = 18 at x = 57 345 and beyond, and state what would have to change if wMaxPacketSize allowed 2048 bytes.

4. §13 found an exact-budget combination by searching achievable costs. Write that search, then determine whether a two-request combination exists for a budget of 90 000 ns, and what that implies about how boundary stimulus should be generated in general.

5. Implement the full-speed cost formula alongside the high-speed one and add a speed input. Decide whether the budget comparison needs to change, and justify the answer from §3.

6. Mutation B1 removes bit stuffing and is caught at nearly every exhaustive point. Construct a smaller version of the same defect — one that is wrong only for some payload sizes — and determine whether the exhaustive check still catches it. Say what that implies about approximating the formula.

7. §14's A4 is structurally guaranteed in two of the three implementations. Find another assertion in this module that is load-bearing in one language and redundant in another, and say what that suggests about writing one assertion set for a tri-HDL design.

19. Summary

What is reserved is time on the wire (§1) — the token, the CRC, the gaps, the host delay, and the payload — computed in nanoseconds because nanoseconds are the resource being divided.

Worst-case bit stuffing inflates every payload by a sixth (§2), and the schedule must assume it because the schedule is built before the data is known. A design that budgets in bytes under-reserves by about 17%.

Periodic transfers are capped (§3): 90% of a full-speed frame, 80% of a high-speed microframe, with the remainder kept so that control transfers can always get through.

The resulting numbers are far less generous than the raw bit rate suggests (§4). A 1024-byte high-speed isochronous transaction costs 20 551 ns; with mult = 3 it costs 61 653 ns, which is 61.7% of the entire periodic budget of every microframe. Two such endpoints do not fit. At full speed, one 1023-byte endpoint consumes 806 of the 900 µs available.

The device must reproduce the host's arithmetic exactly, truncation included (§5). Being more accurate than the specification is a compatibility bug, and the resulting disagreement is silent until a configuration is refused on someone else's machine.

Neither division is by a power of two, so both are reciprocal multiplications (§6) — and the constants must be exact over the whole legal domain, not approximately right.

That domain is 3075 points, and all 3075 were verified in all three languages (§11). Where the input space can be enumerated, enumeration is the only claim worth making, and it leaves no coverage question about the cost function at all.

Six mutations, all killed in all three languages (§12) — but B4 survived the first run entirely (§13). It differed from the correct design only when a reservation total landed exactly on the budget, which no stimulus had produced. A search over the 2907 achievable cost values found a three-request combination summing to exactly 100 000 ns, and adding it killed the mutation. The mutation was not equivalent and the boundary was not unreachable — it was merely unreached, which is the answer that obliges you to do something.

20. Tooling, Honestly

LanguageDesignTestbenchAnalysed / compiledSimulatedMutations
Verilog-2005iso_bw_admissionbw_v_tb.v✅ Icarus -g2005✅ 0 errors✅ all six
SystemVerilogiso_bw_admission_svbw_sv_tb.sv✅ Icarus -g2012✅ 0 errors✅ all six
VHDL-2008iso_bw_admission_vhdlbw_vhdl_tb.vhd✅ nvc 1.23.0✅ 0 errors✅ all six
SVA (§14)——❌ unsupported by Icarus❌—

One tool limitation shaped the SystemVerilog testbench. Icarus rejects the enumeration .name() method applied to a net — The second argument to enum method $ivl_enum_method$name() must be an enumeration variable, found a vpiNet — so the decision is printed as its encoding with a comment giving the mapping. The check itself compares the enumerated values directly and is unaffected; only the diagnostic text changed. Module 15 found a bench whose diagnostics were silently truncated, so the distinction between the check and the message about the check is one this curriculum now treats as load-bearing.

21. What Comes Next

This chapter decided whether an isochronous endpoint may exist. The previous three built devices that assume it does. One question is left, and it is the one the whole module has been circling.

Every chapter so far has said the same thing in passing: there is no retry. 16.1 built a feedback loop because the device cannot say wait. 16.2 built a resynchronising frame marker because a lost packet is not resent. 16.3 built gap arithmetic because missing data stays missing.

Chapter 16.5 asks why. Not what to do about it — the last three chapters were that — but why a bus perfectly capable of retrying deliberately refuses to, what the alternative would cost, and what a device must do at the moment data does not arrive. The answer turns out to be a statement about time rather than about reliability, and it is the trade that defines the transfer type.

Browse the full path on the USB tutorials index.

Continue learning

Standards & specifications

Governing standard
USB-IF (Universal Serial Bus Specification)(opens USB Implementers Forum (USB-IF) in a new tab)

Defines the USB bus — its electrical signalling, connectors, packet and transaction model, device framework and the descriptors a device must expose — together with the device-class specifications layered on it. It does not define host-controller register interfaces (xHCI and EHCI are separate documents) nor any operating system's driver architecture.

This page also covers RTL structure, verification approach and debugging technique. Those are engineering practice built on the standard, not requirements the standard itself imposes.

Where this fits

Part of the USB curriculum.