USB · Module 16
Isochronous Error-Handling Tradeoffs
A bus capable of retrying deliberately refuses, because a retry arrives after the moment it was for. The concealment stage that must never stall — and the mutation VHDL's type system caught that the other languages could not distinguish.
Every chapter in this module has said the same thing in passing. 16.1 built a feedback loop because the device cannot say wait. 16.2 built a resynchronising frame marker because a lost packet is not resent. 16.3 built gap arithmetic because missing data stays missing. 16.4 reserved the bus time in advance because there is none spare to recover with.
All four took "there is no retry" as given. This chapter asks why.
Not what to do about it — the previous four chapters were that — but why a bus that retries control transfers, bulk transfers and interrupt transfers, using mechanisms it already has, deliberately refuses to retry these. The answer is not about reliability. It is about time, and once it is stated the entire transfer type follows from it.
1. A Retry Arrives After the Moment It Was For
Here is the argument in one line: a retransmission is only useful if it arrives before the data is needed, and for an isochronous stream it cannot.
Consider a 48 kHz audio sink. Samples are consumed every 20.8 µs; a packet carrying one service interval's worth of samples is due every 125 µs, and the consumer will want those samples during the next 125 µs whether or not they arrived.
Now price the retry. The host must notice the failure, reschedule the transaction, and issue it — and it cannot issue it in the current service interval, because that interval's bandwidth is already committed to every other periodic endpoint (Chapter 16.4 reserved it). The earliest a retry could land is the next service interval, which is exactly when the next packet is due.
| What arrives | When the consumer needed it | |
|---|---|---|
| original packet | interval n | interval n |
| retry | interval n+1, at the earliest | interval n |
| the packet the retry displaced | never, or interval n+2 | interval n+1 |
The third row is the part that makes retry not merely useless but harmful. The bus time in interval n+1 was reserved for interval n+1's data. A retry occupying it does not conjure extra bandwidth — it steals the slot from the next packet, converting one lost sample into two, and potentially cascading.
Retry works when the consumer can wait. An isochronous consumer cannot wait — it is a DAC, a display, a control loop — and a mechanism that delivers late data at the cost of the next data is worse than no mechanism at all.
2. The Error Is Counted, Not Corrected
If errors are not fixed, they must at least be reported — and the Linux URB does exactly that, per packet:
* ISO transfer status is reported in the status and actual_length fields
* of the iso_frame_desc array, and the number of errors is reported in
* error_count.Three design decisions are visible in that sentence.
Status is per packet, not per transfer. An isochronous URB carries many packets and each gets its own verdict. A single failure does not fail the submission, because the submission is a stream and the stream continues.
actual_length is per packet too, so a short or empty packet is distinguishable from a failed one. Chapter 16.2 §6 needed exactly that distinction: a zero-length packet is legal and inert; a failed packet is not.
error_count is a total. The interface offers no way to ask for a packet again, and it offers a counter instead — which is the whole philosophy in one field. The consumer's job is to decide what to do about the loss, and the protocol's job is to tell it how much there was.
3. The Property That Cannot Be Given Up
If nothing is retried and nothing waits, what must the hardware guarantee?
Exactly one thing: the output never stalls.
A consumer running at a fixed rate — a DAC clocking out samples, a display scanning lines, a servo loop closing — must receive something on every tick. Not the right thing, necessarily. Something.
Because the alternative is worse than wrong data. Chapter 16.3 §1 established that a sensor stream's sample positions are part of its measurement; the same is true of every isochronous sink. A stalled output does not delay the stream — it shifts every subsequent sample by one position, permanently, and the time base never recovers.
| On a missing sample | Immediate effect | Lasting effect |
|---|---|---|
| emit something | one sample is wrong | none — position preserved |
| stall | one sample is late | every later sample is misplaced |
A wrong sample is a local error. A missing sample is a time-base error. That is why the concealment stage in §6 asserts its valid output unconditionally, in every branch, and why §12's mutation E1 — which stalls on loss — is measured rather than argued about.
4. Conceal, Then Mute
Having established that something must be emitted, the question becomes what.
The cheapest answer is to repeat the last good sample. For a single missing sample at 48 kHz this is genuinely excellent: the waveform holds for 20.8 µs, which is inaudible. For two or three it is still the best cheap option.
For a hundred it is a disaster — and not because it sounds "more wrong". A held sample repeated indefinitely is no longer a concealment; it is a constant, which is a DC offset in the signal path or, if the last sample happened to be mid-waveform, a sustained tone. Both are far more objectionable than the gap they were covering, and a DC offset can damage a speaker.
So the policy has two stages and a limit between them:
| Consecutive missing samples | Emit | Why |
|---|---|---|
1 … CONCEAL_LIMIT | the last good sample | cheapest, least audible for short gaps |
beyond CONCEAL_LIMIT | silence | a repeated sample has become a constant |
Better concealments exist — linear interpolation, waveform-similarity overlap-add, spectral estimation — and a real audio controller may implement one. They are all refinements of the same two-stage structure: substitute something plausible, and give up gracefully when the gap is too long to be plausible about. The limit is the part that must exist regardless of how sophisticated the substitution is.
5. The Hardware, Before Any Language
State retained: the last good sample, the current output, a consecutive-concealment counter, and a saturating error total.
On reset or bus reset: everything clears, including the held sample. A sample from a session that has ended is not a sample.
On a corrupt payload arriving: increment the error total, whether or not a sample is due this cycle. The error happened on the wire; tying the count to the consumer's clock would under-report it.
On a sample tick — and this is the whole design — emit something, always:
with usable data (arrived and passed its check): emit it, store it as the last good sample, clear the concealment run.
with no usable data and the run within budget: emit the last good sample, increment the run.
with no usable data and the run at the limit: emit silence, and stay there until real data returns.
Note what "usable" means. An isochronous payload carries one CRC over the whole payload, so a failure says nothing about which bytes are wrong. There is no partially usable packet — it is all good or all suspect — and a design that emits a corrupt payload because "most of it is probably fine" has no basis for that belief. §12's mutation E6 measures it.
The error total saturates. A wrapped count reports a healthy stream for one that is failing constantly — the direction that keeps a fault hidden, which is the third time this module has made that argument (16.1 §6, 16.3 §5) and the third time it is the right one.
6. Verilog
The RTL contract
- What it models: the output stage of an isochronous sink, deciding what to emit when data did not arrive.
- Why it exists: because nothing is retried (§1) and a stalled output is a time-base error rather than a local one (§3).
- Inputs:
sample_tick(the consumer wants a sample now),data_valid,data_err,data_in,bus_reset. - State retained:
last_good,out_r,run_r,err_r, and the registered status flags. - Outputs:
data_out,out_valid,concealed,muted,err_count. - Hardware implied: two
DATA_Wregisters, one small counter, one saturatingERR_Wcounter, one comparator, three flags. - Reset: asynchronous active-low
rst_n;bus_resetsynchronous and equivalent, and both clear the held sample. - Priority: usable data outranks the limit check, which outranks the hold.
- Latency: one cycle; every output is registered.
- Boundaries: the run counter stops at
CONCEAL_LIMIT;err_rsaturates atERR_MAX. - Simultaneous events: an error arriving with a sample tick is both counted and concealed; an error arriving without one is counted only.
- Assumptions:
data_erris the payload's CRC verdict and covers the whole payload;sample_tickis the consumer's rate, already in this clock domain. - Omissions: no FIFO, no CDC, no interpolation, no mute ramp.
- What DV should verify: that
out_validasserts on every tick under every input combination; that a corrupt payload is never emitted; that the hold budget is exact at its boundary; that recovery from mute is immediate; that the error count saturates.
// iso_conceal -- the output stage of an isochronous sink.
//
// Every other transfer type answers "the data did not arrive" with "ask
// again". Isochronous cannot: by the time a retry could be scheduled the
// service interval the data belonged to has passed, and a late sample is not
// a correct sample. So the question this block answers is not how to recover
// the data -- it is what to EMIT when the data is not there.
//
// The one inviolable property is that the output NEVER STALLS. A consumer
// running at a fixed rate must receive something on every sample tick, or
// the stream's time base breaks and every later sample is misplaced. Whether
// that something is good data, a concealment, or silence is a quality
// decision; whether it arrives at all is not negotiable.
//
// The concealment policy has two stages, because one is not enough:
// * hold the last good sample -- excellent for one or two samples, and
// * after CONCEAL_LIMIT consecutive concealments, MUTE instead, because a
// held sample repeated indefinitely is not a gap, it is a constant --
// a DC offset or a tone, which is far more objectionable than silence.
module iso_conceal #(
parameter integer DATA_W = 16,
parameter integer ERR_W = 16,
parameter integer CONCEAL_LIMIT = 8 // consecutive conceals before mute
) (
input wire clk,
input wire rst_n,
input wire bus_reset,
input wire sample_tick, // the consumer wants a sample NOW
input wire data_valid, // a payload arrived for this tick
input wire data_err, // ...but the CRC failed
input wire [DATA_W-1:0] data_in,
output wire [DATA_W-1:0] data_out,
output wire out_valid, // asserted on EVERY sample tick
output wire concealed, // this output was not real data
output wire muted, // ...and was silence, not a hold
output wire [ERR_W-1:0] err_count // saturating
);
localparam [ERR_W-1:0] ERR_MAX = {ERR_W{1'b1}};
localparam integer RUN_W = $clog2(CONCEAL_LIMIT+1);
reg [DATA_W-1:0] last_good;
reg [DATA_W-1:0] out_r;
reg valid_r, conceal_r, mute_r;
reg [RUN_W-1:0] run_r;
reg [ERR_W-1:0] err_r;
assign data_out = out_r;
assign out_valid = valid_r;
assign concealed = conceal_r;
assign muted = mute_r;
assign err_count = err_r;
// Usable data is data that arrived AND passed its check. A corrupt payload
// is not partially usable: an isochronous packet has one CRC over the whole
// payload, so a failure says nothing about which bytes are wrong.
wire usable = data_valid && !data_err;
// This tick will exhaust the hold budget.
wire at_limit = (run_r >= CONCEAL_LIMIT[RUN_W-1:0]);
// Saturating error total. A wrapped count reports a healthy stream for one
// that is failing constantly -- the direction that keeps a fault hidden.
wire [ERR_W:0] err_sum = {1'b0, err_r} + {{ERR_W{1'b0}}, 1'b1};
wire [ERR_W-1:0] err_next = err_sum[ERR_W] ? ERR_MAX : err_sum[ERR_W-1:0];
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
last_good <= {DATA_W{1'b0}}; out_r <= {DATA_W{1'b0}};
valid_r <= 1'b0; conceal_r <= 1'b0; mute_r <= 1'b0;
run_r <= {RUN_W{1'b0}}; err_r <= {ERR_W{1'b0}};
end else if (bus_reset) begin
last_good <= {DATA_W{1'b0}}; out_r <= {DATA_W{1'b0}};
valid_r <= 1'b0; conceal_r <= 1'b0; mute_r <= 1'b0;
run_r <= {RUN_W{1'b0}}; err_r <= {ERR_W{1'b0}};
end else begin
valid_r <= 1'b0;
conceal_r <= 1'b0;
mute_r <= 1'b0;
// A corrupt payload is counted whenever it arrives, whether or not a
// sample was due this cycle -- the error happened on the wire, and
// tying the count to the consumer's clock would under-report it.
if (data_valid && data_err) err_r <= err_next;
if (sample_tick) begin
// out_valid is asserted here UNCONDITIONALLY. Every branch below
// chooses WHAT to emit; none of them chooses whether to emit.
valid_r <= 1'b1;
if (usable) begin
out_r <= data_in;
last_good <= data_in;
run_r <= {RUN_W{1'b0}};
end else if (at_limit) begin
// The hold budget is spent. Emit silence rather than a constant.
out_r <= {DATA_W{1'b0}};
conceal_r <= 1'b1;
mute_r <= 1'b1;
end else begin
// Hold the last good sample: the cheapest concealment there is,
// and for one or two samples the least audible.
out_r <= last_good;
conceal_r <= 1'b1;
run_r <= run_r + {{(RUN_W-1){1'b0}}, 1'b1};
end
end
end
end
endmoduleThree details are worth naming.
valid_r <= 1'b1; sits above the branch, not inside it. Every arm below chooses what to emit; none of them chooses whether to emit. That placement is §3's guarantee expressed structurally rather than by three separate assignments that a later edit could make inconsistent — and mutation E1 is precisely the edit that moves it inside.
The error count is updated outside the sample_tick block. An error is an event on the wire, and a payload can arrive corrupt in a cycle when no sample is due. Counting only on ticks would under-report by whatever fraction of intervals the consumer happens not to be asking in.
usable is one named condition. Chapter 16.2 §13 found a defect in this module's own Verilog caused by writing the same condition inline six times; this design names it once from the start.
7. SystemVerilog
Same behaviour, with one structural change to the interface that is worth arguing for.
package iso_conceal_pkg;
// What this output sample actually is. Three flags -- valid, concealed,
// muted -- can encode combinations that are meaningless (muted without
// concealed). One enumerated value cannot.
typedef enum logic [1:0] {
S_NONE, // no sample was due this cycle
S_GOOD, // real data, received and checked
S_HELD, // concealed by repeating the last good sample
S_MUTED // concealed by silence: the hold budget is spent
} sample_e;
endpackage
module iso_conceal_sv
import iso_conceal_pkg::*;
#(
parameter int unsigned DATA_W = 16,
parameter int unsigned ERR_W = 16,
parameter int unsigned CONCEAL_LIMIT = 8
) (
input logic clk,
input logic rst_n,
input logic bus_reset,
input logic sample_tick,
input logic data_valid,
input logic data_err,
input logic [DATA_W-1:0] data_in,
output logic [DATA_W-1:0] data_out,
output sample_e sample_kind,
output logic [ERR_W-1:0] err_count
);
initial begin
if (CONCEAL_LIMIT == 0)
$fatal(1, "CONCEAL_LIMIT=0 mutes on the first missing sample");
if (ERR_W < 1)
$fatal(1, "ERR_W=%0d cannot count anything", ERR_W);
end
localparam logic [ERR_W-1:0] ERR_MAX = '1;
localparam int unsigned RUN_W = $clog2(CONCEAL_LIMIT+1);
logic [DATA_W-1:0] last_good;
logic [RUN_W-1:0] run_r;
// Usable data is data that arrived AND passed its check. An isochronous
// payload carries one CRC over the whole thing, so a failure says nothing
// about which bytes are wrong -- none of it is partially usable.
wire usable = data_valid && !data_err;
wire at_limit = (run_r >= RUN_W'(CONCEAL_LIMIT));
// What this cycle emits, named rather than inferred from an if/else chain.
sample_e kind_next;
always_comb begin
if (!sample_tick) kind_next = S_NONE;
else if (usable) kind_next = S_GOOD;
else if (at_limit) kind_next = S_MUTED;
else kind_next = S_HELD;
end
// Saturating error total. A wrapped count reports a healthy stream for one
// that is failing constantly -- the direction that keeps a fault hidden.
wire [ERR_W:0] err_sum = {1'b0, err_count} + (ERR_W+1)'(1);
wire [ERR_W-1:0] err_next = err_sum[ERR_W] ? ERR_MAX : err_sum[ERR_W-1:0];
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n || bus_reset) begin
last_good <= '0; data_out <= '0; sample_kind <= S_NONE;
run_r <= '0; err_count <= '0;
end else begin
sample_kind <= kind_next;
// A corrupt payload is counted whenever it arrives, whether or not a
// sample was due -- the error happened on the wire, and tying the count
// to the consumer's clock would under-report it.
if (data_valid && data_err) err_count <= err_next;
unique case (kind_next)
S_NONE: ; // hold everything
S_GOOD: begin
data_out <= data_in;
last_good <= data_in;
run_r <= '0;
end
S_HELD: begin
// The cheapest concealment there is, and for one or two samples
// the least audible.
data_out <= last_good;
run_r <= run_r + 1'b1;
end
S_MUTED: begin
// The hold budget is spent. A held sample repeated indefinitely is
// not a gap, it is a constant -- a DC offset or a tone, which is
// far more objectionable than silence.
data_out <= '0;
end
endcase
end
end
endmoduleThree flags become one enumerated value, and the reason is not brevity.
out_valid, concealed and muted are three independent bits, so they encode eight states — of which exactly four are meaningful. muted without concealed is nonsense. concealed without out_valid is nonsense. Nothing in the Verilog prevents those combinations except the author's discipline, and an assertion that they never occur is a weaker guarantee than an encoding in which they cannot be written.
sample_e | Verilog equivalent |
|---|---|
S_NONE | out_valid = 0 |
S_GOOD | valid = 1, concealed = 0, muted = 0 |
S_HELD | valid = 1, concealed = 1, muted = 0 |
S_MUTED | valid = 1, concealed = 1, muted = 1 |
The mapping is total and the behaviour identical — §11's benches check the same scenarios against both — but the SystemVerilog interface is narrower, and a narrower interface is one in which fewer things need testing because fewer things are expressible.
kind_next also moves §5's decision into one always_comb. The Verilog interleaves what is this sample with what do I do about it; naming the four kinds separates the classification from the consequence, which is the same benefit 16.2's prole_e and 16.3's cls_e provide.
8. VHDL
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
package iso_conceal_pkg is
-- What this output sample actually is. Three independent flags -- valid,
-- concealed, muted -- admit eight combinations of which only four are
-- meaningful. One enumerated value admits exactly the four.
type sample_t is (S_NONE, S_GOOD, S_HELD, S_MUTED);
end package;
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
use work.iso_conceal_pkg.all;
entity iso_conceal_vhdl is
generic (
DATA_W : positive := 16;
ERR_W : positive := 16;
CONCEAL_LIMIT : positive := 8
);
port (
clk : in std_logic;
rst_n : in std_logic;
bus_reset : in std_logic;
sample_tick : in std_logic;
data_valid : in std_logic;
data_err : in std_logic;
data_in : in unsigned(DATA_W-1 downto 0);
data_out : out unsigned(DATA_W-1 downto 0);
sample_kind : out sample_t;
err_count : out unsigned(ERR_W-1 downto 0)
);
end entity;
architecture rtl of iso_conceal_vhdl is
constant ERR_MAX : unsigned(ERR_W-1 downto 0) := (others => '1');
signal last_good : unsigned(DATA_W-1 downto 0) := (others => '0');
signal out_r : unsigned(DATA_W-1 downto 0) := (others => '0');
signal kind_r : sample_t := S_NONE;
signal err_r : unsigned(ERR_W-1 downto 0) := (others => '0');
-- The run counter's range is part of its type, so a value outside it is an
-- error the simulator reports rather than a silent wrap.
signal run_r : integer range 0 to CONCEAL_LIMIT := 0;
signal usable : boolean;
signal at_limit : boolean;
signal kind_next : sample_t;
signal err_sum : unsigned(ERR_W downto 0) := (others => '0');
signal err_next : unsigned(ERR_W-1 downto 0) := (others => '0');
begin
assert CONCEAL_LIMIT >= 1
report "CONCEAL_LIMIT of 0 would mute on the first missing sample"
severity failure;
data_out <= out_r;
sample_kind <= kind_r;
err_count <= err_r;
-- Usable data is data that arrived AND passed its check. An isochronous
-- payload carries one CRC over the whole thing, so a failure says nothing
-- about which bytes are wrong: none of it is partially usable.
usable <= (data_valid = '1') and (data_err = '0');
at_limit <= run_r >= CONCEAL_LIMIT;
kind_next <= S_NONE when sample_tick = '0' else
S_GOOD when usable else
S_MUTED when at_limit else
S_HELD;
-- Saturating error total. A wrapped count reports a healthy stream for one
-- that is failing constantly -- the direction that keeps a fault hidden.
err_sum <= resize(err_r, ERR_W+1) + 1;
err_next <= ERR_MAX when err_sum(ERR_W) = '1'
else err_sum(ERR_W-1 downto 0);
process (clk, rst_n)
begin
if rst_n = '0' then
last_good <= (others => '0'); out_r <= (others => '0');
kind_r <= S_NONE; run_r <= 0; err_r <= (others => '0');
elsif rising_edge(clk) then
if bus_reset = '1' then
last_good <= (others => '0'); out_r <= (others => '0');
kind_r <= S_NONE; run_r <= 0; err_r <= (others => '0');
else
kind_r <= kind_next;
-- A corrupt payload is counted whenever it arrives, whether or not a
-- sample was due -- the error happened on the wire, and tying the
-- count to the consumer's clock would under-report it.
if data_valid = '1' and data_err = '1' then
err_r <= err_next;
end if;
case kind_next is
when S_NONE =>
null; -- hold everything
when S_GOOD =>
out_r <= data_in;
last_good <= data_in;
run_r <= 0;
when S_HELD =>
-- The cheapest concealment there is, and for one or two samples
-- the least audible.
out_r <= last_good;
run_r <= run_r + 1;
when S_MUTED =>
-- The hold budget is spent. A held sample repeated indefinitely
-- is not a gap, it is a constant -- a DC offset or a tone, which
-- is far more objectionable than silence.
out_r <= (others => '0');
end case;
end if;
end if;
end process;
end architecture;run_r is declared integer range 0 to CONCEAL_LIMIT, and that is not decoration. §13 is the account of what happens when a mutation makes the counter exceed it: the simulation stops with a range violation naming the signal and the line, at a point where the other two languages simply wrap and carry on. It is the single most consequential line in this chapter's VHDL.
sample_t has no numeric encoding at all, so — as in 16.2 and 16.3 — the four kinds cannot be compared against integers and the case must be exhaustive.
null in S_NONE is required and useful. VHDL forces the do-nothing branch to be written; the other two express it by the absence of a branch, where an omission and an oversight look identical.
9. Comparing the Three
| Concern | Verilog | SystemVerilog | VHDL |
|---|---|---|---|
| The output status | three flags — 8 encodings, 4 legal | sample_e — 4 encodings, 4 legal | sample_t, no numeric encoding |
| The never-stall guarantee | one assignment above the branch | the S_NONE case is the only non-emitting one | same |
| The run counter | $clog2-sized vector, wraps silently | same | integer range 0 to N — range-checked at run time |
| Illegal parameterisation | undetected | two $fatal guards | one assert ... severity failure |
| Case exhaustiveness | not checked | unique — but see §20 | enforced by the type |
The third row is where this chapter's headline result lives, and §13 measures it rather than asserting it.
10. The Testbenches
The model keeps its own consecutive-concealment tally and its own last-good register, and decides hold-versus-mute from those — never from the design's run_r, which is the signal mutations E3 and E4 corrupt.
if (!tk) e_kind = S_NONE;
else begin
if (v && !er) begin
e_kind = S_GOOD; e_out = d; m_last = d; m_run = 0; // model's own
end else if (m_run >= CONCEAL_LIMIT) begin // model's own
e_kind = S_MUTED; e_out = '0;
end else begin
e_kind = S_HELD; e_out = m_last; m_run++;
end
endThe directed sequence walks the policy's edges:
| Scenario | What it pins down |
|---|---|
| good data | passes through unchanged |
| a tick with no data | out_valid still asserts — §3's guarantee |
| a tick with corrupt data | concealed, the corrupt value is not emitted, error counted |
| good data after a conceal | resumes immediately |
exactly CONCEAL_LIMIT holds | all still S_HELD, still holding the last good value |
| one beyond | S_MUTED, silence, and still valid |
| a further tick | stays muted |
| good data after mute | recovers immediately, and the budget is reset |
| an error with no tick | counted, but produces no sample |
| bus reset | clears the count and the held sample |
| 65 585 corrupt payloads | the error count saturates |
| 6000 randomised cycles | four independent draws, with bursty dropouts |
11. Mutation Testing — Across All Three Languages
| ID | Mutation | Verilog | SystemVerilog | VHDL | Killed |
|---|---|---|---|---|---|
| — | baseline, no mutation | 0 | 0 | 0 | — |
| E1 | the output stalls when data is missing | 1648 | 1742 | 1805 | ✅ all three |
| E2 | conceal with zero instead of holding | 1561 | 1561 | 1643 | ✅ all three |
| E3 | no mute limit — holds the last sample for ever | 169 | 169 | FATAL | ✅ all three — see §12 |
| E4 | the conceal run counter never advances | 169 | 169 | 129 | ✅ all three |
| E5 | the error count wraps instead of saturating | 51 | 51 | 51 | ✅ all three |
| E6 | corrupt data is emitted instead of concealed | 1022 | 1018 | 1120 | ✅ all three |
E1 is the transfer type's defining property, measured. Stalling on loss costs 1648–1805 failures — the largest count here — because it does not corrupt one sample, it shifts every sample that follows.
E5's 51 is entirely from the directed saturation test, and honestly so: the randomised phase deliberately clears the error counter every 500 cycles to keep it in its linear region (§10), which means it never approaches 65 535 there. The saturation boundary is directed-only by construction, exactly as Chapter 16.3 §13 reported for its own loss total. "6000 randomised cycles passed" says nothing about E5.
E2 is worth noting for what it is not. Emitting zero instead of holding is still a concealment, still asserts valid, and still preserves the time base — it is merely a worse-sounding one. It is caught 1561 times not because anything structural breaks but because the bench checks the emitted value, which a bench that only checked out_valid would not.
12. The Mutation VHDL Caught That the Others Could Not Distinguish
E3 and E4 produced identical counts in Verilog and SystemVerilog — 169 each. That is suspicious in the way Chapter 16.3 §12 describes, so it was checked rather than recorded:
E3 failures: 169 E4 failures: 169
IDENTICAL failure sets -> observationally EQUIVALENT in Verilog
FAIL: muted (out=3000 v=1 c=1 m=0 err=1 | exp 0 1 1 1, t=157000) <- E3
FAIL: muted (out=3000 v=1 c=1 m=0 err=1 | exp 0 1 1 1, t=157000) <- E4Byte-identical failure sets, same first divergence, same timestamp. The two mutations are different code changes with the same observable behaviour:
- E3 disables the limit check, so
at_limitis never true. The counter keeps incrementing and wraps silently in its$clog2-sized vector. - E4 removes the increment, so the counter never grows and
at_limitis never true.
Either way the design never mutes, and run_r is not an output. In Verilog and SystemVerilog these are observationally equivalent mutations — a perfectly legitimate outcome, and one §43 of this curriculum's standard requires be identified rather than counted twice.
In VHDL they are not equivalent, and the difference is the type system:
** Fatal: 155ns+0: value 9 outside of INTEGER range 0 to 8 for signal RUN_R
> cn_vhdl_E3.vhd:111
|
111 | run_r <= run_r + 1;
| ^^^^^^^^^E3 does not produce 169 checker failures in VHDL. It stops the simulation at 155 nanoseconds, two nanoseconds before the testbench's first behavioural failure, naming the signal, the line and the illegal value. E4 — which never increments — violates no range and runs to completion with 129 failures.
13. The Stream Under Loss
Every arrow into the consumer is a tick that was answered. That is §3's guarantee drawn rather than argued: the sink's content changes from real, to held, to silent, to real again — and at no point does the consumer wait.
The transition at tick 10 is the one §4 exists for. Everything before it is a concealment; everything from it is an admission that the gap is too long to conceal. Mutations E3 and E4 both delete that transition, and §12 is the account of how differently the three languages reacted.
14. Assertions
// A1. THE property of this transfer type, and the one assertion that would
// be worth writing if only one were allowed: every consumer tick is
// answered, unconditionally, whatever the input looked like.
property p_never_stalls;
@(posedge clk) disable iff (!rst_n || bus_reset)
sample_tick |=> out_valid;
endproperty
a_never_stalls: assert property (p_never_stalls);
// A2. SAFETY: a corrupt payload is never emitted. Phrased on the OBSERVED
// inputs, not on the design's `usable` -- which is what E6 corrupts.
property p_no_corrupt_data;
@(posedge clk) disable iff (!rst_n || bus_reset)
(sample_tick && data_valid && data_err) |=> (data_out != $past(data_in));
endproperty
a_no_corrupt_data: assert property (p_no_corrupt_data);
// A3. BOUNDEDNESS, and the property VHDL's subtype enforces for free:
// the concealment run never exceeds its limit.
property p_run_bounded;
@(posedge clk) disable iff (!rst_n)
(run_r <= CONCEAL_LIMIT);
endproperty
a_run_bounded: assert property (p_run_bounded);
// A4. PROGRESS: a sustained dropout MUST eventually mute. A design that
// conceals for ever satisfies every safety property here -- including
// A1 -- and is the defect of section 4.
property p_eventually_mutes;
@(posedge clk) disable iff (!rst_n || bus_reset)
concealed [*CONCEAL_LIMIT+1] |-> muted;
endproperty
a_eventually_mutes: assert property (p_eventually_mutes);
// A5. RECOVERY: good data always ends concealment immediately. The stream
// must not need to "settle" before real samples resume.
property p_recovery_immediate;
@(posedge clk) disable iff (!rst_n || bus_reset)
(sample_tick && data_valid && !data_err) |=> (!concealed && !muted);
endproperty
a_recovery_immediate: assert property (p_recovery_immediate);
// A6. The error count never decreases except through a reset.
property p_err_monotonic;
@(posedge clk) disable iff (!rst_n || bus_reset)
1'b1 |=> (err_count >= $past(err_count));
endproperty
a_err_monotonic: assert property (p_err_monotonic);Assertion contracts
| Claim | Safety / progress | Vacuity risk | How non-vacuity is established | |
|---|---|---|---|---|
| A1 | every tick is answered | safety | low | 3997 ticks measured |
| A2 | a corrupt payload is never emitted | safety | moderate | corrupt payloads arrive constantly in the random phase |
| A3 | the run never exceeds its limit | safety | none — no antecedent | holds every cycle |
| A4 | a sustained dropout eventually mutes | progress | high — needs a run of CONCEAL_LIMIT+1 | 83 mutes after §10's bursty-loss fix; 2 before it |
| A5 | recovery is immediate | safety | low | 2351 good samples measured |
| A6 | the error count never decreases | safety | low | 66 236 corrupt payloads |
A4's vacuity row is §10's finding stated as a contract, and it is the clearest example this module has produced of why the vacuity question must be quantitative. Under the original stimulus the antecedent held twice in 6000 cycles. A vacuity report saying "A4 was exercised" would have been true and would have concealed that the module's central safety valve was resting on two samples.
A1 is the assertion to keep if all the others were deleted. It is short, it has no interesting antecedent, and it states the one property that distinguishes this transfer type from every other one in USB.
A4 is also the only progress property here, and §30's argument applies with unusual force: a design that concealed for ever would satisfy A1, A2, A3, A5 and A6 perfectly. It would also, given a long enough dropout, put a DC offset through a speaker.
15. Verification: the Chapter Where Error Injection Is the Test
Chapter 16.1 declined UVM, 16.2 adopted it, 16.3 found a narrow case and 16.4 declined it again. This block is the strongest case in the module, and the reason is precisely §10.
The test is the error model. Every interesting behaviour of this design is a response to a loss pattern, and §10 demonstrated that the shape of that pattern — independent versus bursty — changed the coverage of the central path by a factor of forty. That is a stimulus-modelling problem, and constrained-random with a well-chosen distribution is the right tool for it.
| Component | Why it is justified |
|---|---|
| Error-injection sequence | the whole test. Burst length, burst spacing, corrupt-versus-absent, and the correlation between them are the knobs that matter |
| Sequence item | one service interval: arrived / absent / corrupt, plus the payload |
| Driver | thin — drive the interval |
| Monitor | observes data_out and the status on the interface and reconstructs what should have been emitted itself |
| Reference model | keeps its own last-good register and conceal tally — never the design's run_r, which E3 and E4 corrupt |
| Coverage | where this environment earns most — see below |
The coverage model follows directly from §10 and §11:
bin: burst_length { 1, 2, .. CONCEAL_LIMIT-1, CONCEAL_LIMIT,
CONCEAL_LIMIT+1, long }
cross: burst_length × loss_kind (absent, corrupt, mixed)
bin: samples_emitted_by_kind (GOOD, HELD, MUTED)
bin: recovery_from { hold, mute }
bin: err_count_region { linear, saturated }
cross: tick_present × data_present -- the four alignmentsburst_length binned around CONCEAL_LIMIT is the bin that would have caught §10's problem on the first regression. The original stimulus would have shown CONCEAL_LIMIT+1 at a count of two while every shorter bin was full — a visible, standing report of exactly the hole that took a manual reach measurement to find.
And err_count_region is Chapter 16.1's clamp-masking lesson made permanent. This module has now hit that problem three times — 16.1's feedback clamp, 16.3's loss total, 16.5's error count — and each time the fix was noticed by hand. A coverage bin that distinguishes a counter's linear region from its saturated one turns a recurring manual discovery into a line in every regression report.
The reference model must not share the design's run counter, and the reason is §12. Mutations E3 and E4 both corrupt exactly that counter, and a model that consulted it would agree with both. It must maintain its own tally from the loss pattern it observed — which is the same principle as 15.3's wide accumulator, 16.2's packet list and 16.3's skip count. Four chapters, four different mechanisms, one rule.
16. Debugging: the Headset That Buzzes When the Wi-Fi Is Busy
A USB headset plays correctly most of the time. When the machine's wireless link is heavily loaded, the audio develops a loud, sustained buzz at a constant pitch — not a click, not a dropout, a tone. It stops when the transfer finishes. The device is not overheating, the cable is fine, and the same headset behaves identically on another machine.
The pitch is the diagnosis. A click is a discontinuity; a dropout is silence; a sustained tone at a constant pitch is a repeating waveform, and the only thing in an isochronous sink that repeats a waveform is §4's hold.
So the fault is not that data is being lost — the correlation with wireless load says data is being lost, and that is expected on a shared bus with an aggressive neighbour. The fault is that the concealment never gives up. The device is holding its last good sample through a dropout long enough that the hold has become a constant, which is exactly mutations E3 and E4.
And the pitch tells you the sample rate relationship. If the device holds a single sample, the output is DC — a thump and then silence-with-offset. A buzz at a pitch implies it is repeating a short block of samples, which means the concealment is a block repeat rather than a single hold, and the pitch is the block rate. That is a more sophisticated concealment than §6's, with the same missing limit.
The chain, from the outside in:
Host — is the device actually losing intervals? The isochronous frame descriptors carry per-packet status and error_count (§2). Rising error_count under wireless load confirms the premise and locates the contention, which is a scheduling problem rather than a device problem.
Device statistics — does the device report concealments, and does it report mutes? A device that counts concealments but has no mute state is E3. A device whose mute counter is always zero while its conceal counter climbs is the same bug with instrumentation attached.
RTL — the first divergence. Compare the bench's independent conceal tally against the design's run_r under a long dropout. The first cycle where the design should have muted and did not is the bug — and §12 is the record of how hard that is to see when the run counter is not an output, and how much easier VHDL's range check made it.
17. Common Misconceptions
"Isochronous has no error detection." It has a CRC like every other packet type (§2). What it has no mechanism for is correction.
"Retry was left out to keep the protocol simple." It was left out because a retry cannot arrive in time and would consume the next interval's reserved slot (§1) — turning one lost sample into two.
"A retry could use spare bandwidth." There is no spare bandwidth by construction: 16.4 reserved it all in advance, which is what makes the guarantee a guarantee (§1).
"Some of a corrupt packet is usable." One CRC covers the whole payload (§5), so a failure says nothing about which bytes are wrong. Mutation E6 emits it anyway.
"A stalled output just delays the stream slightly." It shifts every subsequent sample permanently (§3). Mutation E1 costs the largest failure count in the chapter.
"Holding the last sample is a complete concealment policy." It is half of one. Without a limit it becomes a DC offset or a tone (§4), which is worse than the gap — and is §16's buzzing headset.
"Muting is a failure mode." Muting is the correct behaviour once a gap exceeds what can be plausibly concealed. The failure mode is concealing for ever.
"Two mutations with the same failure count are probably the same mutation." §12: E3 and E4 were observationally equivalent in two languages and distinguishable in the third.
18. Exercises
1. At 48 kHz with a 125 µs service interval, compute how many samples a single lost packet costs and how long CONCEAL_LIMIT = 8 holds before muting. Then choose a limit for a 192 kHz stream and justify it in milliseconds rather than samples.
2. Replace the hold with linear interpolation between the last good sample and the first sample after the gap. Determine what that requires of the datapath that a hold does not, and state why it cannot be done in the architecture of §5.
3. §12 showed E3 and E4 are observationally equivalent in Verilog. Make them distinguishable without changing the language: add one output, and say what it costs.
4. Add a soft ramp into and out of mute. Decide whether the ramp belongs in this block or downstream, and justify the answer from the never-stall property of §3.
5. Implement A4 from §14 as a standing procedural check, then run it against the original stimulus from §10 that produced two mutes. Report how many times its antecedent held, and say what a vacuity report should have shown.
6. The error counter saturates and §10 clears it periodically to keep it testable. Design a hardware mechanism that preserves the information a saturated counter discards, and compare its cost against simply widening the counter.
7. §16's table maps audio artefacts to RTL defects. Extend it with the artefact produced by Chapter 16.1's feedback loop failing — a device whose reported rate is stuck at nominal — and say how you would distinguish it from a concealment problem by ear.
19. Summary
A retry cannot arrive before the data is needed (§1), and the slot it would use is already reserved for the next interval's data — so retrying converts one lost sample into two. The absence of retry is a consequence of the bandwidth guarantee, not an independent simplification: you can have reserved periodic bandwidth or unscheduled retries, and not both.
Errors are therefore counted rather than corrected (§2), per packet, with a running total — and the interface offers no way to ask for a packet again.
The one property that cannot be given up is that the output never stalls (§3). A wrong sample is a local error; a missing sample is a time-base error that displaces everything after it permanently. Mutation E1 measures the difference at 1648–1805 failures.
Concealment has two stages and the limit between them is the part that must exist (§4). Holding the last good sample is excellent for one or two samples and becomes a DC offset or a tone for a hundred — which is §16's buzzing headset, heard rather than measured.
All three HDL implementations were simulated (§20), and six mutations died in all three languages (§11) — but two of the six were observationally equivalent in Verilog and SystemVerilog and distinguishable in VHDL (§12). A run_r declared integer range 0 to CONCEAL_LIMIT stopped the simulation at 155 ns naming the signal and the line, where the other two languages wrapped silently and reported a symptom eight cycles later. The extra observability was in the type declaration, not in the design — and it is a simulation-time property that does not synthesise, so its entire value is in verification.
And the stimulus was wrong in a way no check could report (§10). Independent per-packet loss is statistically respectable and physically wrong for a shared serial bus, and it reached the mute path twice in 6000 cycles. Bursty dropouts raised that to 83. The central safety valve of the design was resting on two samples, and only a reach measurement said so.
20. Tooling, Honestly
| Language | Design | Testbench | Analysed / compiled | Simulated | Mutations |
|---|---|---|---|---|---|
| Verilog-2005 | iso_conceal | cn_v_tb.v | ✅ Icarus -g2005 | ✅ 0 errors | ✅ all six |
| SystemVerilog | iso_conceal_sv | cn_sv_tb.sv | ✅ Icarus -g2012 | ✅ 0 errors | ✅ all six |
| VHDL-2008 | iso_conceal_vhdl | cn_vhdl_tb.vhd | ✅ nvc 1.23.0 | ✅ 0 errors | ✅ all six |
| SVA (§14) | — | — | ❌ unsupported by Icarus | ❌ | — |
One mutation's VHDL result is not a number and is not recorded as one. E3 does not produce a failure count in VHDL — it produces a fatal range violation that ends the run. Reporting it as "0 errors" would be false and reporting it as some count would be invented, so the table says FATAL and §12 gives the transcript.
21. What Comes Next
Module 16 is complete. Isochronous transfers reserve bandwidth and surrender the retry, and the five chapters have built what a device needs on each side of that trade: a feedback loop for rate, a resynchronising marker for structure, sequence arithmetic for gaps, admission control for the reservation itself, and a concealment stage for the moment data does not arrive.
Every one of them took the host's schedule as given. The poll arrives, the slot exists, the interval is 125 µs and something happens in it. Module 17 — USB Scheduling — is where that schedule is actually built.
It is the question this module and Module 15 have both been deferring: given a bus with isochronous endpoints holding reservations, interrupt endpoints holding latency bounds, control transfers that must never be starved, and bulk transfers taking whatever is left — in what order does the host actually issue transactions, and how does it decide? The answer is a priority scheme and a per-frame budget, and it explains a great deal about why the numbers in 16.4 are what they are.
Browse the full path on the USB tutorials index.
Continue learning
Related tutorials
- Related topic
USB Webcams
UVC streams video over isochronous transfers, which have no retries. With only a frame-ID bit and an end-of-frame flag for framing, exactly 1 packet loss in P is detectable — 6% for a 16-packet frame, and under 1% for a real one.
- Related topic
USB Audio Devices
An audio device runs on its own crystal, so a few parts per million empty the buffer every minute — with nothing lost and nothing to retry. The fix is one 10.14 number per frame, and the steady-state offset is a closed form: crystal error × 2^gain.
- Related topic
The Polling Model
The host runs a periodic schedule the device cannot see, so device hardware must count frames rather than trust a period — built in Verilog, SystemVerilog and VHDL with a schedule model that disagrees independently.
- Related topic
Audio over Isochronous
Isochronous gives up the retry and the NAK, so a device cannot slow the host down — it can only report the rate it needs. The feedback accumulator in three HDLs, and the clamp that hid a defect from eight of nine chances to catch it.
Standards & specifications
- Governing standard
- USB-IF (Universal Serial Bus Specification)(opens USB Implementers Forum (USB-IF) in a new tab)
Defines the USB bus — its electrical signalling, connectors, packet and transaction model, device framework and the descriptors a device must expose — together with the device-class specifications layered on it. It does not define host-controller register interfaces (xHCI and EHCI are separate documents) nor any operating system's driver architecture.
This page also covers RTL structure, verification approach and debugging technique. Those are engineering practice built on the standard, not requirements the standard itself imposes.
Where this fits
Part of the USB curriculum.
