USB · Module 16
Audio over Isochronous
Isochronous gives up the retry and the NAK, so a device cannot slow the host down — it can only report the rate it needs. The feedback accumulator in three HDLs, and the clamp that hid a defect from eight of nine chances to catch it.
Chapter 15.4 closed by measuring how little bInterval actually promises: a bound on the gap, and nothing else. Interrupt transfers buy that bound with a small, capped share of bandwidth, and every device in Module 15 had one property in common — when something went wrong, there was always another poll.
Isochronous transfers give that up. They reserve bandwidth rather than buying a latency bound, and they pay for the reservation by surrendering the two mechanisms every previous transfer type relied on: the handshake and the retry.
The consequence is not a smaller safety net. It is a different engineering problem, and this chapter is the first half of it: when a device cannot say wait, and cannot say send that again, what control does it actually have?
1. Best Effort Is the Whole Contract
The Linux USB core states the quality-of-service guarantee in one clause, and it is worth reading before any mechanism:
* Isochronous URBs have a different data transfer model, in part because
* the quality of service is only "best effort"."Best effort" here is precise, not vague. It means the host guarantees a reservation — a slot of bus time in every service interval — and guarantees nothing whatsoever about the contents of that slot arriving intact.
Compare the four transfer types on the one axis that matters now:
| Handshake | Retry on error | Flow control (NAK) | Guaranteed | |
|---|---|---|---|---|
| Control | yes | yes | yes | correctness |
| Bulk | yes | yes | yes | correctness, eventually |
| Interrupt | yes | yes | yes | a bound on the gap (15.4) |
| Isochronous | none | none | none | a reserved slot, best effort |
Read the last row as three separate losses, because each removes a different capability:
No handshake means the sender never learns whether the packet arrived. There is no ACK to wait for and no status to inspect.
No retry follows from that, but it is also deliberate: by the time a retransmission could be scheduled, the service interval the data belonged to has passed. A late sample is not a correct sample.
No NAK is the one engineers underestimate. In Chapter 14, a device that was not ready said NAK and the host tried again. An isochronous endpoint has no way to say not ready and no way to say slow down. The data arrives whether the device can take it or not.
The error model changes shape too. Because there is no retry, errors are not corrected — they are counted. The Linux URB carries a per-packet status array and a total:
* ISO transfer status is reported in the status and actual_length fields
* of the iso_frame_desc array, and the number of errors is reported in
* error_count.Per packet, not per transfer. One packet of a stream can fail while the rest succeed, and the stream continues regardless. Chapter 16.5 is entirely about what a device should do with that information; this chapter needs only the fact that nothing retries.
2. Two Clocks That Will Never Agree
Here is the physical problem, stated without protocol vocabulary.
A USB audio sink — a headset, a speaker, a DAC — converts samples to sound at a rate set by its own crystal. The host produces samples at a rate derived from its own clock. Both are nominally 48 000 Hz. Neither is exactly 48 000 Hz.
A crystal specified at ±100 ppm may be off by 4.8 Hz at 48 kHz. That sounds negligible, and over one millisecond it is: 0.0048 of a sample. The problem is that the error integrates.
| Elapsed time | Accumulated drift at 100 ppm |
|---|---|
| 1 ms (one frame) | 0.0048 samples |
| 1 second | 4.8 samples |
| 1 minute | 288 samples |
| 1 hour | 17 280 samples — about 0.36 s of audio |
The device buffers samples between arrival and playback, and that buffer is finite. If the host sends faster than the device consumes, the buffer fills and eventually overflows. If the host sends slower, the buffer empties and underruns. Either way the result is audible: a click, a dropout, or a gap.
The mismatch is not a fault. It is the guaranteed steady-state behaviour of two independent oscillators, and a design that does not address it has not failed to handle an edge case — it has failed to handle the normal case, slowly.
A bulk endpoint would never notice this, because its NAK absorbs it. The producer is throttled by the consumer automatically, and the two rates are reconciled by a mechanism nobody had to design. Isochronous removes that mechanism and hands the problem to the device.
3. The Four Synchronisation Types
USB's answer is to make the device declare how it relates to the bus clock, in two bits of the endpoint descriptor's bmAttributes. The Linux header names them exactly:
#define USB_ENDPOINT_SYNCTYPE 0x0c
#define USB_ENDPOINT_SYNC_NONE (0 << 2)
#define USB_ENDPOINT_SYNC_ASYNC (1 << 2)
#define USB_ENDPOINT_SYNC_ADAPTIVE (2 << 2)
#define USB_ENDPOINT_SYNC_SYNC (3 << 2)Bits 3:2 of bmAttributes — the same byte whose bits 1:0 select the transfer type. Each value describes a different answer to whose clock wins:
| Value | Name | The device's clock is… | Rate reconciliation |
|---|---|---|---|
0 << 2 | NONE | not synchronised and not declared | none — for endpoints where it is meaningless |
1 << 2 | ASYNCHRONOUS | free-running, its own | the device tells the host the rate it needs |
2 << 2 | ADAPTIVE | slaved to the incoming data rate | the device follows whatever arrives |
3 << 2 | SYNCHRONOUS | locked to the bus SOF | the device derives its clock from USB |
These are three genuinely different hardware architectures, not three configuration settings.
Synchronous puts a PLL in the device locked to the 1 kHz SOF. The device has no independent clock, so there is no drift to correct — but the audio clock now inherits whatever jitter the host's frame timing has, and SOF jitter is not a good audio reference.
Adaptive puts a rate estimator in the device: it measures how much data actually arrives and tunes its own sample clock to match. The device tracks the host. This works and is common, but the device's clock quality is now limited by how well it can estimate the host's rate.
Asynchronous lets the device run from a clean local crystal — the best possible audio clock — and accepts that it will never match the host. It then needs a channel to tell the host what rate to send at. That channel is a second endpoint, and USB gives it its own encoding:
#define USB_ENDPOINT_USAGE_MASK 0x30
#define USB_ENDPOINT_USAGE_DATA 0x00
#define USB_ENDPOINT_USAGE_FEEDBACK 0x10
#define USB_ENDPOINT_USAGE_IMPLICIT_FB 0x20 /* Implicit feedback Data endpoint */Bits 5:4 of the same byte. An endpoint declares itself as carrying data, as carrying feedback, or as an implicit-feedback data endpoint — one whose own data rate implicitly tells the host the device's rate, so no separate feedback endpoint is needed.
This chapter builds the asynchronous case, because it is the one with real hardware in it. Synchronous needs a PLL, which is a mixed-signal problem outside this curriculum's scope. Adaptive is a rate estimator — a close cousin of what we are about to build. Asynchronous is the one where the device must produce a number and put it on the wire, and that number is the interesting part.
4. The Feedback Value Is a Rate, in Fixed Point
The device reports, on its feedback endpoint, how many samples per (micro)frame it actually wants. Not a correction, not an error term — an absolute rate.
The encoding is unsigned fixed point, and the format differs by speed:
| Speed | Format | Width | Units |
|---|---|---|---|
| Full speed | 10.14 | 3 bytes | samples per frame (1 ms) |
| High speed | 16.16 | 4 bytes | samples per microframe (125 µs) |
Take full speed and 48 kHz. Nominal is 48 samples per frame, so the nominal feedback value is 48 << 14 = 786 432. A device running 0.5 samples per frame fast reports 48.5 << 14 = 794 624.
The fractional part is the entire point. A device that could only report whole samples per frame would be limited to 48 or 49 — a 2% rate step, wildly coarser than the ±100 ppm it is trying to correct. Fourteen fractional bits resolve 1/16384 of a sample per frame, which at 48 kHz is about 1.3 ppm: comfortably finer than the drift being corrected.
The format is not arbitrary precision for its own sake. 10 integer bits cover every audio rate USB full speed can carry, and 14 fractional bits are what make the value finer than the error it exists to correct.
5. The Window Length Is the Resolution
Now the hardware question, and it has an unusually elegant answer.
The device knows its true rate by counting how many samples its DAC actually consumed. Counting over one frame gives an integer — 48, or 47, or 49 — with no fractional information at all. The fraction has to come from somewhere, and it comes from counting over more than one frame.
Count over 2^K frames. The total is 2^K times the average per-frame rate, so the average is the total divided by 2^K — and dividing by a power of two is not a division. Reading the same count with K binary fractional digits is the average.
So the Q10.14 feedback value is:
fb = sample_count << (Q − K) where Q = 14A shift. There is no divider in this design, and there never needs to be.
The consequence is the design's central trade, and it is not a tuning knob added afterwards — it falls out of the arithmetic:
Window 2^K frames | Resolution (samples/frame) | In ppm at 48 kHz | Update period |
|---|---|---|---|
| K = 4 (16 frames) | 1/16 = 0.0625 | 1302 ppm | 16 ms |
| K = 6 (64 frames) | 1/64 = 0.0156 | 326 ppm | 64 ms |
| K = 8 (256 frames) | 1/256 = 0.0039 | 81 ppm | 256 ms |
| K = 10 (1024 frames) | 1/1024 = 0.00098 | 20 ppm | 1.02 s |
Longer window, finer measurement, slower answer. A device that needs to correct 100 ppm of drift needs at least K = 8 to see it at all — and accepts that it responds to a change a quarter of a second later.
6. The Hardware, Before Any Language
State retained: a sample counter scnt, a frame counter fcnt, the current feedback value, and a locked flag.
On reset or bus reset: the counters clear, and the feedback value initialises to nominal — not to zero. A feedback value of zero tells the host to send no samples at all, and a device that reports it during its first window has stopped its own stream before it started.
On a sample tick (the DAC consumed a sample): scnt increments, saturating rather than wrapping.
On a frame boundary (SOF) that is not a window boundary: fcnt increments.
On the window boundary — the 2^K-th SOF: the feedback value takes scnt << (Q−K), clamped to a legal band around nominal; fcnt and scnt reset; locked sets; a one-cycle fb_valid pulse fires.
On a sample tick in the very cycle the window closes: the closing window's value uses the count before this cycle's tick, and the new window starts with that tick already counted. Every tick lands in exactly one window — the same conservation rule the interrupt-transfer accumulators had to obey in Chapter 15.3, here protecting a rate rather than a distance.
Two of those decisions deserve their reasons stated, because both are the difference between a working device and a plausible one.
Why saturate the counter. A wrapped sample count does not produce an obviously wrong value — it produces a plausible small one. A device running away at twice the expected rate would wrap and report a low rate, and the host would send even more data. The failure would be amplified by the mechanism designed to correct it.
Why clamp the output. The same argument, one level up. A measurement corrupted by a startup transient, a glitching sample clock, or a partially elapsed first window produces a value the host will act on immediately. Clamping to nominal ± a small band means the worst a bad measurement can do is push the rate slightly the wrong way for one window. A stale feedback value is survivable; a wild one is not.
7. Verilog
The RTL contract
- What it models: the rate-measurement accumulator behind an asynchronous isochronous audio sink's feedback endpoint.
- Why it exists: because isochronous has no NAK (§1), so a device whose crystal differs from the host's (§2) can only influence the stream by reporting a rate (§4).
- Inputs:
sof(one pulse per frame),sample_tick(one pulse per sample the DAC consumed),bus_reset. - State retained:
scnt,fcnt,fb_r,valid_r,locked_r. - Outputs:
fb_value(Q10.14),fb_valid(one cycle),locked. - Hardware implied: one
CNT_W-bit saturating counter, oneK-bit counter, a left shift (wiring, not logic), two comparators for the clamp, one 24-bit register, two flags. - Reset: asynchronous active-low
rst_n;bus_resetsynchronous and equivalent. Both initialisefb_rto nominal, not zero. - Priority: the window boundary outranks the frame boundary, which outranks the plain tick.
- Latency: the feedback value is
2^Kframes old the instant it is published — that is §5's trade, not a defect. - Boundaries:
scntsaturates;fb_valueclamps toNOM ± BAND. - Collision semantics: a tick coinciding with the window close belongs to the opening window.
- Assumptions:
sample_tickis already synchronised into this clock domain, and asserts exactly once per consumed sample. - Omissions: no CDC, no packet formatting, no rate-change handling, one endpoint.
- What DV should verify: that no tick is lost or double-counted across every alignment of
sofandsample_tick; that the scaling is exact; that both clamp limits are reachable and correct; thatlockedis false until a full window has genuinely elapsed.
// usb_audio_feedback -- the rate-feedback accumulator of an ASYNCHRONOUS
// isochronous audio sink.
//
// The device's DAC runs from its own crystal. The host sends samples at a
// rate derived from ITS clock. Isochronous has no NAK and no retry, so the
// device cannot slow the host down by refusing data -- the only control it
// has is to TELL the host what rate to send at, on a feedback endpoint.
//
// The measurement is deliberately simple: count the samples the DAC actually
// consumed over 2**K frames. Interpreted with K fractional bits that count IS
// the average samples-per-frame, so the Q10.14 feedback value is a SHIFT and
// never a division:
//
// fb = sample_count << (Q - K)
//
// The window length is therefore not a tuning knob bolted on afterwards --
// it IS the fractional resolution. K=6 measures to 1/64 of a sample per
// frame and answers every 64 ms; K=10 measures 16x finer and answers 16x
// more slowly. That trade is the whole design.
module usb_audio_feedback #(
parameter integer K = 6, // window = 2**K frames
parameter integer Q = 14, // fractional bits (10.14 for full speed)
parameter integer CNT_W = 16, // sample counter width
parameter integer NOM = 48, // nominal samples per frame (48 kHz)
parameter integer BAND = 2 // legal band: NOM +/- BAND samples/frame
) (
input wire clk,
input wire rst_n,
input wire bus_reset,
input wire sof, // 1-cycle pulse at each frame start
input wire sample_tick, // 1-cycle pulse per sample consumed
output wire [Q+9:0] fb_value, // Q10.14 samples per frame (24 bits)
output wire fb_valid, // 1-cycle pulse when fb_value updates
output wire locked // a full window has completed
);
localparam integer SHIFT = Q - K;
// WIDTH AUDIT. The RESULT is 10.14 -- 24 bits. The INTERMEDIATE is not:
// a saturated counter shifted left needs CNT_W+SHIFT bits, which here is
// 26. Computing `raw` in 24 bits would silently discard exactly the
// overflow the clamp below exists to catch, and the clamp would then be
// comparing a number that had already wrapped.
localparam integer RAW_W = CNT_W + SHIFT; // 26
localparam integer FB_W = Q + 10; // 24 -- the wire format
localparam [RAW_W-1:0] FB_MIN = (NOM - BAND) << Q;
localparam [RAW_W-1:0] FB_MAX = (NOM + BAND) << Q;
localparam [CNT_W-1:0] CNT_MAX = {CNT_W{1'b1}};
reg [CNT_W-1:0] scnt; // samples in the window currently open
reg [K-1:0] fcnt; // frames elapsed in that window
reg [FB_W-1:0] fb_r;
reg valid_r, locked_r;
assign fb_value = fb_r;
assign fb_valid = valid_r;
assign locked = locked_r;
// The raw measurement, widened BEFORE the shift. Widening after the shift
// would discard exactly the high bits the shift just created.
wire [RAW_W-1:0] raw = {{SHIFT{1'b0}}, scnt} << SHIFT;
// Clamp. A wild feedback value is worse than a stale one: the host would
// act on it immediately and the buffer it is supposed to protect is the
// thing that overflows.
wire [RAW_W-1:0] clamped = (raw < FB_MIN) ? FB_MIN :
(raw > FB_MAX) ? FB_MAX : raw;
// Safe to narrow ONLY because the clamp above proves clamped <= FB_MAX,
// and FB_MAX = (NOM+BAND)<<Q = 50<<14 = 819200 < 2**24.
wire [FB_W-1:0] clamped_fb = clamped[FB_W-1:0];
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
scnt <= {CNT_W{1'b0}}; fcnt <= {K{1'b0}};
fb_r <= NOM << Q; // start at nominal, not at zero
valid_r <= 1'b0; locked_r <= 1'b0;
end else if (bus_reset) begin
scnt <= {CNT_W{1'b0}}; fcnt <= {K{1'b0}};
fb_r <= NOM << Q;
valid_r <= 1'b0; locked_r <= 1'b0;
end else begin
valid_r <= 1'b0;
if (sof && (fcnt == {K{1'b1}})) begin
// The window closes. fb takes the count as it stood BEFORE this
// cycle's sample tick; a tick arriving now belongs to the window
// that is opening. Every tick lands in exactly one window --
// conservation, the same rule the interrupt-transfer accumulators
// had to obey.
fb_r <= clamped_fb;
valid_r <= 1'b1;
locked_r<= 1'b1;
fcnt <= {K{1'b0}};
scnt <= sample_tick ? {{(CNT_W-1){1'b0}}, 1'b1} : {CNT_W{1'b0}};
end else begin
if (sof) fcnt <= fcnt + {{(K-1){1'b0}}, 1'b1};
// Saturate rather than wrap: a wrapped count reports a PLAUSIBLE
// small rate for a device that is actually running away.
if (sample_tick && (scnt != CNT_MAX))
scnt <= scnt + {{(CNT_W-1){1'b0}}, 1'b1};
end
end
end
endmoduleThree details are worth naming.
The shift is free. raw is scnt moved left by SHIFT bits — no adder, no divider, no multiplier. §5's arithmetic is the reason: the averaging is a power of two by construction, so the hardware cost of the fractional resolution is wiring.
The widening happens before the shift. {{SHIFT{1'b0}}, scnt} << SHIFT extends first and then shifts. Shifting a CNT_W-wide value and widening afterwards would discard exactly the high bits the shift just created — the ones that tell the clamp an overflow occurred.
raw is 26 bits and fb_value is 24. That is not sloppiness; it is §51's width audit made explicit. The result is a 10.14 wire format, but the intermediate of a saturated counter shifted left needs CNT_W + SHIFT bits. Computing raw at the result's width would let it wrap before the clamp ever inspected it, and the clamp would then be comparing a number that had already lost the information it was checking for.
8. SystemVerilog
Same hardware. What changes is that the fixed-point format stops being a comment.
package usb_audio_pkg;
// The full-speed feedback value is an unsigned 10.14 fixed-point number of
// samples per frame. Declaring it as a STRUCT rather than a 24-bit vector
// puts the binary point in the type system: whole.frac is checkable, and a
// reader cannot mistake which end the fraction lives at.
typedef struct packed {
logic [9:0] whole; // integer samples per frame
logic [13:0] frac; // fractional samples per frame
} fb_q10_14_t;
// What a window boundary is doing this cycle, named rather than inferred.
typedef enum logic [1:0] {
W_IDLE, // neither a frame boundary nor a window boundary
W_FRAME, // a frame boundary inside the window
W_CLOSE // the window boundary itself
} win_e;
endpackage
module usb_audio_feedback_sv
import usb_audio_pkg::*;
#(
parameter int unsigned K = 6,
parameter int unsigned Q = 14,
parameter int unsigned CNT_W = 16,
parameter int unsigned NOM = 48,
parameter int unsigned BAND = 2
) (
input logic clk,
input logic rst_n,
input logic bus_reset,
input logic sof,
input logic sample_tick,
output fb_q10_14_t fb_value,
output logic fb_valid,
output logic locked
);
localparam int unsigned SHIFT = Q - K;
localparam int unsigned RAW_W = CNT_W + SHIFT;
// Elaboration guards. Each of these silently disables a different part of
// the design if violated, and none of them would fail a simulation loudly.
initial begin
if (Q <= K)
$fatal(1, "Q=%0d must exceed K=%0d or the shift is negative", Q, K);
if (NOM + BAND >= (1 << 10))
$fatal(1, "NOM+BAND=%0d overflows the 10-bit whole part", NOM + BAND);
if ((NOM * (1 << K)) >= (1 << CNT_W))
$fatal(1, "a nominal window needs %0d counts, CNT_W=%0d cannot hold it",
NOM * (1 << K), CNT_W);
end
localparam logic [RAW_W-1:0] FB_MIN = (NOM - BAND) << Q;
localparam logic [RAW_W-1:0] FB_MAX = (NOM + BAND) << Q;
localparam logic [CNT_W-1:0] CNT_MAX = '1;
logic [CNT_W-1:0] scnt;
logic [K-1:0] fcnt;
win_e win;
always_comb begin
if (sof && (fcnt == '1)) win = W_CLOSE;
else if (sof) win = W_FRAME;
else win = W_IDLE;
end
// Width audit, stated in the code: the intermediate is RAW_W bits and the
// result is 24. The cast makes the widening explicit rather than relying
// on an inferred concatenation.
logic [RAW_W-1:0] raw, clamped;
always_comb begin
raw = RAW_W'(scnt) << SHIFT;
clamped = (raw < FB_MIN) ? FB_MIN :
(raw > FB_MAX) ? FB_MAX : raw;
end
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n || bus_reset) begin
scnt <= '0;
fcnt <= '0;
fb_value.whole<= 10'(NOM); // nominal, never zero
fb_value.frac <= '0;
fb_valid <= 1'b0;
locked <= 1'b0;
end else begin
fb_valid <= 1'b0;
unique case (win)
W_CLOSE: begin
// The window's value is the count as it stood BEFORE this cycle's
// tick; a tick arriving now opens the next window. Conservation.
// Exactly 10 bits, not clamped[RAW_W-1:Q] -- that slice is 12
// bits wide and would rely on an IMPLICIT truncation being safe.
// The clamp proves bits above Q+9 are zero; slice what is meant.
fb_value.whole <= clamped[Q+9:Q];
fb_value.frac <= clamped[Q-1:0];
fb_valid <= 1'b1;
locked <= 1'b1;
fcnt <= '0;
scnt <= sample_tick ? CNT_W'(1) : '0;
end
W_FRAME: begin
fcnt <= fcnt + 1'b1;
if (sample_tick && (scnt != CNT_MAX)) scnt <= scnt + 1'b1;
end
W_IDLE: begin
if (sample_tick && (scnt != CNT_MAX)) scnt <= scnt + 1'b1;
end
endcase
end
end
endmoduleThe struct packed puts the binary point in the type system. fb_q10_14_t has a whole field and a frac field, so a reader cannot mistake which end the fraction lives at, and the testbench can check fb_value.frac === 14'd1024 — one sixteenth of a sample per frame — rather than comparing a 24-bit magic number. The Verilog version encodes the same layout in a comment and a set of bit indices.
clamped[Q+9:Q], not clamped[RAW_W-1:Q]. Both name the whole part, and the second is a 12-bit slice assigned to a 10-bit field — an implicit truncation that happens to be safe because the clamp proves the top two bits are zero. Slicing exactly ten bits states the intent; relying on the truncation states nothing and would survive a later change to BAND that made it false.
The three $fatal guards catch parameterisations that silently disable the design. Q <= K makes the shift negative. NOM + BAND >= 1024 overflows the whole part. A CNT_W too small to hold a nominal window makes the counter saturate every time, pinning the feedback at maximum forever. None of these fails loudly at run time — each produces a design that runs and is wrong, which is precisely the category worth catching at elaboration.
9. VHDL
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
package usb_audio_pkg is
-- The full-speed feedback value is an unsigned 10.14 fixed-point number of
-- samples per frame. It is a RECORD here for the same reason it is a struct
-- in the SystemVerilog: the binary point belongs in the type, not in a
-- comment. Note that the record is fixed at 10.14 and is NOT parameterised
-- -- the wire format is fixed by the specification, so making it a generic
-- would offer a freedom the protocol does not actually grant.
type fb_q10_14_t is record
whole : unsigned(9 downto 0);
frac : unsigned(13 downto 0);
end record;
end package;
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
use work.usb_audio_pkg.all;
entity usb_audio_feedback_vhdl is
generic (
K : positive := 6;
Q : positive := 14; -- must be 14: the wire format is 10.14
CNT_W : positive := 16;
NOM : positive := 48;
BAND : natural := 2
);
port (
clk : in std_logic;
rst_n : in std_logic;
bus_reset : in std_logic;
sof : in std_logic;
sample_tick : in std_logic;
fb_value : out fb_q10_14_t;
fb_valid : out std_logic;
locked : out std_logic
);
end entity;
architecture rtl of usb_audio_feedback_vhdl is
constant SHIFT : natural := Q - K;
-- WIDTH AUDIT: the intermediate is wider than the result. A saturated
-- counter shifted left needs CNT_W+SHIFT bits; the wire format is 24.
constant RAW_W : natural := CNT_W + SHIFT;
constant FB_MIN : unsigned(RAW_W-1 downto 0) :=
to_unsigned((NOM - BAND) * 2**Q, RAW_W);
constant FB_MAX : unsigned(RAW_W-1 downto 0) :=
to_unsigned((NOM + BAND) * 2**Q, RAW_W);
constant CNT_MAX : unsigned(CNT_W-1 downto 0) := (others => '1');
signal scnt : unsigned(CNT_W-1 downto 0) := (others => '0');
signal fcnt : unsigned(K-1 downto 0) := (others => '0');
signal fb_r : fb_q10_14_t;
signal vld_r : std_logic := '0';
signal lock_r : std_logic := '0';
-- Initialised. Without this, `clamped` is evaluated once at delta 0 from
-- an uninitialised `raw`, and numeric_std's comparison reports a metavalue.
-- The transient is harmless -- it resolves before the first clock edge --
-- but VHDL SAYS SO and Verilog does not: see the cross-HDL note in the
-- chapter, where the identical expression returns a silent 'x'.
signal raw, clamped : unsigned(RAW_W-1 downto 0) := (others => '0');
begin
-- Elaboration guards. Each of these silently disables part of the design.
assert Q = 14
report "Q must be 14: the full-speed feedback wire format is 10.14"
severity failure;
assert Q > K
report "Q must exceed K or the scaling shift is negative"
severity failure;
assert NOM + BAND < 2**10
report "NOM+BAND overflows the 10-bit whole part of the wire format"
severity failure;
assert NOM * 2**K < 2**CNT_W
report "CNT_W cannot hold a nominal window's sample count"
severity failure;
fb_value <= fb_r;
fb_valid <= vld_r;
locked <= lock_r;
-- resize() before shift_left(): both are defined for UNSIGNED, so the
-- widening cannot accidentally become a signed extension and the shift
-- cannot silently drop the bits it just created.
raw <= shift_left(resize(scnt, RAW_W), SHIFT);
clamped <= FB_MIN when raw < FB_MIN else
FB_MAX when raw > FB_MAX else
raw;
process (clk, rst_n)
begin
if rst_n = '0' then
scnt <= (others => '0');
fcnt <= (others => '0');
fb_r.whole <= to_unsigned(NOM, 10); -- nominal, never zero
fb_r.frac <= (others => '0');
vld_r <= '0';
lock_r <= '0';
elsif rising_edge(clk) then
if bus_reset = '1' then
scnt <= (others => '0');
fcnt <= (others => '0');
fb_r.whole <= to_unsigned(NOM, 10);
fb_r.frac <= (others => '0');
vld_r <= '0';
lock_r <= '0';
else
vld_r <= '0';
if sof = '1' and fcnt = CNT_MAX(K-1 downto 0) then
-- The window closes on the count as it stood BEFORE this cycle's
-- tick; a tick arriving now opens the next window. Conservation:
-- every tick is counted in exactly one window.
fb_r.whole <= clamped(Q+9 downto Q);
fb_r.frac <= clamped(Q-1 downto 0);
vld_r <= '1';
lock_r <= '1';
fcnt <= (others => '0');
if sample_tick = '1' then
scnt <= to_unsigned(1, CNT_W);
else
scnt <= (others => '0');
end if;
else
if sof = '1' then
fcnt <= fcnt + 1;
end if;
-- Saturate: a wrapped count reports a plausible small rate for a
-- device that is in fact running away.
if sample_tick = '1' and scnt /= CNT_MAX then
scnt <= scnt + 1;
end if;
end if;
end if;
end if;
end process;
end architecture;The record does what the struct does, and the entity declares what the struct could not. fb_q10_14_t is deliberately not parameterised on Q: the wire format is fixed at 10.14 by the specification, so making it a generic would advertise a freedom the protocol does not grant. The assert Q = 14 immediately below says the same thing to anyone who tries.
resize then shift_left, both defined for unsigned. The widening cannot accidentally become a sign extension, because scnt is an unsigned and numeric_std resolves the operator from the type rather than from the author's intent. §12 returns to what that does and does not buy.
scnt + 1 needs no conversion, because numeric_std defines addition of an unsigned and a natural. The Verilog writes {{(CNT_W-1){1'b0}}, 1'b1} to match the counter's width exactly; the VHDL lets the type carry that.
10. Comparing the Three
| Concern | Verilog | SystemVerilog | VHDL |
|---|---|---|---|
| The 10.14 wire format | comment plus bit indices | struct packed — checkable | record — checkable, and fixed at 10.14 by design |
| Widening before the shift | explicit concatenation | RAW_W'(scnt) cast | resize, defined for the type |
| Narrowing to the wire | slice, author checks the bound | explicit 10-bit slice | explicit 10-bit slice |
| Illegal parameterisation | undetected | three $fatal guards | four assert ... severity failure |
| The four cycle cases | if/else if chain | typedef enum + unique case | if/elsif chain |
| Counter type | reg, unsigned by convention | logic, same | unsigned, enforced |
All three describe the same hardware. The honest summary is that the differences are about what the author's intent is written down, and §14's mutation results confirm it: the same six defects are caught by all three benches, at almost exactly the same cost.
11. The Testbenches
Each language has its own bench, and all three share the decision that made Chapter 15.3's real bug findable: the model does not use the design's method.
The design counts incrementally and shifts. Every bench instead keeps an absolute running total of ticks and derives each window by subtraction, then scales by multiplication:
task automatic step(input bit s, input bit t);
sof = s; sample_tick = t;
@(posedge clk); #1;
sof = 0; sample_tick = 0;
boundary_before = total_ticks; // the total BEFORE this cycle's tick
if (t) total_ticks++; // -- which is what a closing window sees
m_valid = 0;
if (s) begin
sof_count++;
if ((sof_count % (1<<K)) == 0) begin
win = boundary_before - ticks_at_close; // SUBTRACTION, not counting
raw = win * (1<<SHIFT); // MULTIPLICATION, not a shift
m_fb = clampf(raw);
ticks_at_close = boundary_before;
m_locked = 1; m_valid = 1;
end
end
#1;
check(fb_as_int(fb_value) === m_fb, "feedback value");
check(fb_valid === m_valid, "fb_valid pulse");
check(locked === m_locked, "locked");
endtaskWhy subtraction of running totals is the right independent method. It makes conservation structural rather than asserted: the sum of all windows is, by construction, the total number of ticks. If the design loses a tick anywhere, the model cannot lose the same one, because the model never decrements and never resets its total. That is the property under test, and the model is incapable of agreeing with a design that violates it.
Directed boundaries come before any randomisation — §26 of this curriculum's standard, and §13 below shows why it is not a formality:
| Scenario | What it pins down |
|---|---|
| a full window at exactly nominal | the scaling is exact, not approximately right |
| one extra sample in a 16-frame window | 1/16 of a sample resolves to frac === 14'd1024 exactly |
| back to nominal | the window genuinely reset |
| a window far too fast | the high clamp |
| a window far too slow | the low clamp |
| a tick in the closing cycle | the colliding tick is excluded from the closing window… |
| the following window | …and present in the next one |
bus_reset | feedback returns to nominal and locked clears |
| 4000 randomised steps | every alignment of sof and sample_tick |
The randomised phase draws its two stimulus bits independently. Chapter 15.4 found a bench that derived several supposedly independent bits from one random draw, which made the very collision it was meant to exercise unreachable. The VHDL bench here calls uniform twice, and the Verilog and SystemVerilog benches call $random twice:
-- TWO independent draws. Deriving both bits from one sample would
-- correlate them and the SOF/tick collision would become unreachable.
for i in 0 to 3999 loop
uniform(seed1, seed2, r_sof);
uniform(seed1, seed2, r_tick);
step(r_sof < 0.0588, r_tick < 0.667);
end loop;And the VHDL bench's counters are variables, not signals. A signal assignment does not take effect until the process waits, so two checks failing before the next wait would each compute errors + 1 from the same stale value and one increment would vanish. This is Chapter 15.3's finding applied deliberately rather than rediscovered.
Reporting was checked, not assumed. The Verilog benches declare message arguments at 80 characters (input [639:0] msg); Module 15 shipped a bench whose [255:0] argument silently truncated every message longer than 31 characters from the left, so the failure text named the wrong check while the pass/fail verdict stayed correct.
12. Mutation Testing — Across All Three Languages
Six mutations, each applied to the Verilog, the SystemVerilog and the VHDL, each run against that language's own bench from a clean baseline.
| ID | Mutation | Verilog | SystemVerilog | VHDL | Killed |
|---|---|---|---|---|---|
| — | baseline, no mutation | 0 | 0 | 0 | — |
| A1 | scale by Q instead of Q−K | 7728 | 7728 | 7636 | ✅ all three |
| A2 | the colliding sample tick is lost | 2 | 2 | 2 | ✅ all three |
| A3 | no clamp on the feedback value | 5115 | 5115 | 5023 | ✅ all three |
| A4 | the window closes one frame early | 3632 | 3632 | 3721 | ✅ all three |
| A5 | locked asserted from reset | 784 | 1112 | 784 | ✅ all three |
| A6 | the sample count is never cleared | 5975 | 5975 | 6055 | ✅ all three |
The VHDL numbers are real. Unlike Module 15, where no VHDL analyser existed and every VHDL claim had to be marked reviewed, not run, this module had a working simulator and the VHDL was analysed, elaborated, simulated and mutated like the other two. §21 records the tooling exactly.
The VHDL counts differ slightly from the other two — 7636 against 7728, 3721 against 3632 — and the reason is not a semantic difference. The randomisers are different. VHDL's uniform and Verilog's $random produce different sequences, so the three benches drive different random stimulus after the identical directed phase. That the three agree on every verdict while disagreeing on every count is exactly the right outcome: it shows the designs are equivalent and that the benches are not transcriptions of one another.
A5 is not a language difference, and checking mattered
A5 costs 1112 failures in SystemVerilog against 784 in the other two, which looks like a semantic divergence and is not. Reading the code settles it:
if (!rst_n) begin // Verilog: power-on reset only
...
end else if (bus_reset) begin // a SEPARATE branch if (!rst_n || bus_reset) begin // SystemVerilog: ONE merged branchThe SystemVerilog design merges the two reset paths, so the same textual mutation lands on both of them there and on only the power-on path in the other two. The SV mutant therefore also corrupts locked after every bus_reset, and the extra 328 failures are the bench's post-bus-reset checks.
§44 of this curriculum's standard warns not to assume a syntactically similar mutation has the same semantics in each HDL. This is a measured instance: the mutation is textually identical, the designs are behaviourally identical, and the mutants are not — because the designs differ in how they group a condition, not in what they do.
13. The Clamp That Hid the Collision
A2 — the colliding tick being lost — was killed by two failures. In a run with 4000 randomised steps. That is a pass, and stopping there would miss the most useful result in the chapter.
Here is where those two failures occurred:
FAIL: feedback value (fb=785408 exp=786432 valid=1 locked=1, t=54897000)
FAIL: the colliding tick was carried into the next window, not lost
(fb=785408 exp=786432 valid=1 locked=1, t=54897000)Both at the same instant — the single directed collision test. The randomised phase contributed nothing. And the reason is not the one Chapter 15.4 found; the collision here was reached repeatedly:
REACH: closes=21 collisions=8 clamp_lo=15 clamp_hi=1
REACH: UNCLAMPED closes=5 collisions at an unclamped close=1Eight collisions occurred, and seven of them were undetectable.
The clamp is why. FB_MIN corresponds to 736 samples in a 16-frame window and FB_MAX to 800. The randomised stimulus produced windows from 134 to 896 samples, so 16 of the 21 window closes were clamped — and a value pinned to FB_MIN or FB_MAX is identical whether or not one sample was lost on the way in. Only 5 closes were in the linear region at all, and only 1 collision landed on one of them.
The fix is not more randomisation. It is a directed band: hold the rate inside NOM ± BAND and sweep the collision across it, so every collision lands where a one-sample error is visible. Exercise 3 asks for that measurement, and the prediction is that A2's failure count rises by roughly the number of collisions generated.
14. The Waveform
Two windows: 4 samples reads 2.00, 5 samples reads 2.50
13 cyclesThis waveform is in the controller clock domain, not on the USB wire. sof here is the controller's internal one-cycle pulse marking a frame boundary, not the SOF packet's shape on D+/D−; sample_tick is an audio-clock event already synchronised into this domain. Nothing in this picture depicts bus signalling, and reading it as bus timing would suggest a frame is four clocks long.
Cycle 5 against cycle 2 is the whole window mechanism. Both carry sof. Cycle 2 merely advances fcnt; cycle 5 finds fcnt already at its maximum and closes the window. Mutation A4 changes exactly which of those two a given SOF is — and costs over 3600 failures.
Cycle 12 is why 14 fractional bits exist. Five samples over two frames is 2.5 per frame, a rate no integer encoding can express. The device is running 25% fast in this toy parameterisation, and the format reports it exactly rather than rounding it to 2 or 3.
15. Assertions
// A1. CONSERVATION, stated independently of the design's counters: between
// two window closes, the number of ticks observed on the interface must
// equal the number of samples the published value accounts for. This is
// the property the whole design exists to satisfy.
// tick_obs is the BENCH's own counter, not the DUT's scnt.
property p_no_tick_lost;
@(posedge clk) disable iff (!rst_n || bus_reset)
fb_valid |-> (window_ticks_observed == $past(scnt_expected));
endproperty
a_no_tick_lost: assert property (p_no_tick_lost);
// A2. SAFETY: the published value never leaves the legal band, whatever
// the measurement was. Note this is checked on the OUTPUT, so it holds
// even if the clamp logic itself is wrong.
property p_in_band;
@(posedge clk) disable iff (!rst_n)
locked |-> (fb_as_int(fb_value) >= FB_MIN) &&
(fb_as_int(fb_value) <= FB_MAX);
endproperty
a_in_band: assert property (p_in_band);
// A3. PROGRESS: a device that is being clocked and consuming samples must
// eventually publish. A design that never closes a window satisfies
// every safety property above and is useless.
property p_eventually_publishes;
@(posedge clk) disable iff (!rst_n || bus_reset)
$rose(locked) |-> ##[1:$] fb_valid;
endproperty
a_eventually_publishes: assert property (p_eventually_publishes);
// A4. STABILITY: fb_value changes ONLY on a valid pulse. A value that
// drifts between publications would be read differently depending on
// when the host happened to fetch it.
property p_stable_between;
@(posedge clk) disable iff (!rst_n || bus_reset)
!fb_valid |=> $stable(fb_value);
endproperty
a_stable_between: assert property (p_stable_between);
// A5. fb_valid is a ONE-CYCLE pulse, never a level.
property p_valid_is_pulse;
@(posedge clk) disable iff (!rst_n || bus_reset)
fb_valid |=> !fb_valid;
endproperty
a_valid_is_pulse: assert property (p_valid_is_pulse);
// A6. The window is exactly 2**K frames -- no more, no fewer. Counting
// SOFs on the interface is independent of the DUT's fcnt.
property p_window_length;
@(posedge clk) disable iff (!rst_n || bus_reset)
fb_valid |-> (sofs_since_last_valid == (1 << K));
endproperty
a_window_length: assert property (p_window_length);Assertion contracts
| Claim | Safety / progress | Vacuity risk | How non-vacuity is established | |
|---|---|---|---|---|
| A1 | no tick is lost or double-counted across a window | safety | high — a design that never publishes never triggers it | paired with A3, which requires publication |
| A2 | the published value is always in band | safety | low | locked is reached in every directed scenario |
| A3 | a running device eventually publishes | progress | moderate — locked never rising makes it vacuous | the directed sequence reaches locked before any random stimulus |
| A4 | the value is stable between publications | safety | low | the majority of cycles have fb_valid low |
| A5 | fb_valid is a pulse | safety | low | fb_valid asserts 21 times in the measured run |
| A6 | the window is exactly 2^K frames | safety | high — same as A1 | paired with A3 |
A1 and A6 are the two that must not be read alone, and the reason is §30's distinction. Both are implications antecedent on fb_valid. A design that never publishes anything satisfies both perfectly — and satisfies A2, A4 and A5 as well, because all five are safety properties and a design that does nothing violates no safety property. A3 is the only one in the list that a do-nothing design fails, which is precisely why it is there.
A1's independence is the other point worth stating. It compares the published value against ticks counted by the bench on the interface, not against the DUT's own scnt. §29 of this standard warns against properties phrased in terms of the decision under test; a property written as scnt_correct |-> fb_correct would be satisfied by a design that mis-counted consistently, which is exactly mutation A6.
16. Verification: Why This Chapter Does Not Use UVM
Chapter 15.3 argued that UVM earns its place when the property spans an unbounded stream and the transaction is not the pins. Both are arguably true here, and UVM is still the wrong tool for this block.
What the environment would have to reason about is one scalar published every 2^K frames. The transaction abstraction is a single number; the sequence is a rate profile; the scoreboard is one comparison. Wrapping that in a uvm_sequence_item, an agent, a driver and a monitor produces several hundred lines whose entire content is the dozen lines of step() in §11.
§33 of this curriculum's standard is explicit about this — do not wrap a 10-line decoder in 400 lines of UVM — and the honest reading is that the interesting complexity in this chapter is arithmetic and boundary behaviour, which directed tests and assertions address far more directly than a constrained-random environment.
Where UVM does start to earn its place in this module:
| Chapter | Why UVM becomes justified |
|---|---|
| 16.2 | a video frame spans many transactions with a header protocol, ordering requirements, and mid-frame loss — multiple transaction types and stateful sequencing |
| 16.5 | error injection across a stream, with classification and recovery policy — the canonical UVM case |
The rule this chapter is applying: UVM's cost is justified by scenario complexity, not by the importance of the block. This accumulator is important and simple. A block can be both.
17. Debugging: the Headset That Clicks Every Few Minutes
A USB headset plays cleanly for several minutes, then produces a brief click, then plays cleanly again. The interval between clicks is roughly constant. Audio quality is otherwise perfect, the device never disconnects, and nothing appears in any host log.
Read the symptom before opening anything. Periodic, brief, and otherwise perfect is the signature of a buffer boundary being reached, not of corruption. Corruption is aperiodic and correlates with bus activity; a drift problem is periodic and correlates with time.
And the period tells you the rate error directly. If the buffer holds N samples of slack and the click recurs every T seconds, the rate mismatch is about N / (48000 × T). A 480-sample buffer clicking every 100 seconds is roughly 100 ppm — a number you can compare against the crystal's specification before touching the hardware.
The chain, from the outside in:
Protocol analyser — is feedback being requested at all? If the host never issues IN transactions on the feedback endpoint, the descriptor is the fault, not the RTL. §3's warning applies: bmAttributes bits 3:2 must say asynchronous, and the feedback endpoint's bits 5:4 must say feedback. A device that declares synchronisation type NONE gets no feedback requests and will drift exactly like this.
Protocol analyser — is the reported value sane? Decode it as 10.14 and divide by 16384. A value that reads as 48.0000 forever means the measurement is not running; a value that reads 0 means §6's reset-to-nominal was not implemented; a value pinned at the band edge means the clamp is holding a measurement that is wildly wrong.
Controller state — is locked ever asserting? If the window never closes, the device is publishing its reset value forever and the host is dutifully sending nominal into a device that needs something else. Mutation A4 and A6 both produce variants of this.
RTL — the first divergence. Compare the bench's independent tick total against the design's scnt at each window close. The first close where they differ is the bug, and it is almost always a conservation failure at a boundary: a tick lost to the collision (A2), a window closing early (A4), or a counter that never cleared (A6).
18. Common Misconceptions
"Isochronous means real-time or fast." It means reserved and best-effort (§1). The reservation is about guaranteed bus time, not low latency, and the best-effort part means data can be lost with no retry.
"A device can NAK an isochronous transfer if it isn't ready." There is no handshake phase at all (§1). That is the whole reason this chapter exists.
"Crystal drift is a tolerance problem that good components solve." Two independent oscillators always diverge; ±100 ppm accumulates 17 280 samples an hour (§2). It is the steady-state behaviour, not a defect.
"The synchronisation type is a configuration detail." It selects between three different hardware architectures — a PLL, a rate estimator, or a feedback endpoint (§3).
"The feedback value is a correction or an error term." It is an absolute rate in samples per frame (§4).
"You can average over more frames for a better answer without cost." The window length is the update period (§5); finer resolution is slower response, exactly and unavoidably.
"A longer window is always more accurate." It is more precise. If the device's true rate changes — a sample-rate switch, a clock source change — a long window reports the old rate for 2^K frames afterwards.
"Clamping the output is a belt-and-braces safety measure with no downside." §13 measured the downside: the clamp erased seven of eight chances to detect a real defect.
19. Exercises
1. Compute the feedback value for a device that needs 44.1 kHz on a full-speed interface. Note that 44.1 samples per frame is not an integer and explain what the device sends in a frame where the host has been told 44.1 — then say which field of §4's format carries the 0.1.
2. Derive the minimum K needed to resolve a 50 ppm rate error at 48 kHz, and state the resulting update period. Compare it against the click period your answer implies for a 1024-sample buffer.
3. Implement §13's fix: constrain the randomised rate to stay inside NOM ± BAND so every window close lands in the linear region, then re-run mutation A2. Predict the new failure count before running it, using the measured collision count as your estimate.
4. Implement the adaptive sink instead of the asynchronous one: the device has no feedback endpoint and must tune its own sample clock from the observed arrival rate. Identify which parts of this chapter's RTL survive unchanged and which disappear entirely.
5. Add a rate-change input (a SET_INTERFACE switching 48 kHz to 96 kHz). Decide whether locked should clear, whether the window should restart, and what the device should report during the first window at the new rate. Justify each choice from §6 rather than from convenience.
6. Write an SVA property that would catch mutation A6 (scnt never cleared) without referring to the design's scnt. Then explain why A1 as written in §15 does catch it.
7. The VHDL reported a metavalue at time zero where the Verilog silently produced x (§10). Construct a case where the VHDL's substitution of FALSE would produce a worse outcome than Verilog's x propagation, and say which tool would have found it.
20. Summary
Isochronous surrenders the handshake, the retry and the NAK (§1), and the last of those is the one that creates this chapter's problem: losing the NAK is losing flow control, which is what every other transfer type used to reconcile a producer and a consumer running at different rates.
Two independent crystals always diverge (§2). At ±100 ppm the drift is 0.0048 samples per frame and 17 280 samples per hour, so the buffer between arrival and playback will eventually overflow or underrun. This is the normal case, slowly.
USB's answer is to make the device declare its clock relationship in two bits of bmAttributes (§3) — synchronous, adaptive, or asynchronous — and to give the asynchronous case a second endpoint, declared through two further bits, on which the device reports the rate it needs.
That rate is an absolute value in unsigned fixed point — 10.14 samples per frame at full speed (§4) — and the fourteen fractional bits exist because whole samples per frame would be a 2% step where the error being corrected is 0.01%.
The measurement's resolution and its response time are one parameter (§5). Counting over 2^K frames makes the Q10.14 value a shift and never a division, and the same K sets both the fractional resolution and the update period. That is not a design choice about how to measure; it is the information content of the observation.
All three HDL implementations model the same hardware and all three were simulated (§21), including the VHDL — which Module 15 could not do. Six mutations were killed in all three languages (§12), and one apparent language difference turned out to be a mutation-placement artefact of the SystemVerilog design merging two reset branches, not a semantic divergence.
And the chapter's most useful result is a verification finding, not a design one (§13). Mutation A2 died by two failures despite the collision occurring eight times in the randomised phase — because 16 of 21 window closes were clamped, and a one-sample error inside a clamped value is arithmetically erased. A numeric clamp is saturating evidence, the same structure as 15.4's sticky flag, and the bench's own wide random spread was what suppressed detection.
21. Tooling, Honestly
| Language | Design | Testbench | Analysed / compiled | Simulated | Mutations |
|---|---|---|---|---|---|
| Verilog-2005 | usb_audio_feedback | fb_v_tb.v | ✅ Icarus -g2005 | ✅ 0 errors | ✅ all six |
| SystemVerilog | usb_audio_feedback_sv | fb_sv_tb.sv | ✅ Icarus -g2012 | ✅ 0 errors | ✅ all six |
| VHDL-2008 | usb_audio_feedback_vhdl | fb_vhdl_tb.vhd | ✅ nvc 1.23.0 | ✅ 0 errors | ✅ all six |
| SVA (§15) | — | — | ❌ unsupported by Icarus | ❌ | — |
unique case | — | — | ✅ accepted | ⚠️ quality ignored | — |
The VHDL row is the one that changed. Module 15 had no VHDL analyser available and every VHDL claim there was marked reviewed, not run. A working simulator — nvc 1.23.0 — was available for this module, so the VHDL here was analysed, elaborated, simulated and mutation-tested exactly like the other two, and §12's VHDL column is measured rather than reasoned.
The unique case row is the subtler caveat. Icarus accepts the keyword and reports sorry: Case unique/unique0 qualities are ignored. The exclusivity is guaranteed by the always_comb that assigns win, so the design is correct — but on this tool the keyword documents intent rather than verifying it, and a keyword that is accepted and unenforced reads in review as a checked property.
22. What Comes Next
This chapter's device streamed a continuous, uniform thing. Every audio sample is the same size, every frame carries the same number of them, and a lost sample is one sample.
Chapter 16.2 streams something with structure. A video frame is many isochronous packets that must be reassembled, carrying a payload header that marks where each frame begins and ends — and the packets are not uniform, because the last packet of a frame is whatever is left over.
That changes the failure mode completely. A lost audio sample is a click. A lost video packet is a corrupted frame, and unless the device marks its frame boundaries in a way the host can resynchronise from, it is a corrupted frame and every frame after it. The mechanism that prevents that is a single toggling bit, and building it is the next chapter.
Browse the full path on the USB tutorials index.
Continue learning
Related tutorials
- Related topic
USB Audio Devices
An audio device runs on its own crystal, so a few parts per million empty the buffer every minute — with nothing lost and nothing to retry. The fix is one 10.14 number per frame, and the steady-state offset is a closed form: crystal error × 2^gain.
- Related topic
Bandwidth Reservation
An isochronous endpoint reserves time on the wire, not bytes — and worst-case bit stuffing inflates every payload by a sixth. The host's admission arithmetic reproduced exactly, verified over its entire input domain.
- Related topic
USB Webcams
UVC streams video over isochronous transfers, which have no retries. With only a frame-ID bit and an end-of-frame flag for framing, exactly 1 packet loss in P is detectable — 6% for a 16-packet frame, and under 1% for a real one.
- Related topic
Mice on USB
A relative report cannot be resent, so the accumulator must saturate rather than wrap and must be cleared by the act of being read — and a real signedness bug the testbench caught on its first check.
Standards & specifications
- Governing standard
- USB-IF (Universal Serial Bus Specification)(opens USB Implementers Forum (USB-IF) in a new tab)
Defines the USB bus — its electrical signalling, connectors, packet and transaction model, device framework and the descriptors a device must expose — together with the device-class specifications layered on it. It does not define host-controller register interfaces (xHCI and EHCI are separate documents) nor any operating system's driver architecture.
This page also covers RTL structure, verification approach and debugging technique. Those are engineering practice built on the standard, not requirements the standard itself imposes.
Where this fits
Part of the USB curriculum.
