USB · Module 14
Bulk Reliability
Guarantee-correct and bandwidth-best-effort are two promises, not one — and the five-rung recovery ladder that makes the first true has rungs that are not interchangeable.
Chapter 14.1 assumed every transaction succeeds first time. None of them do, always.
And Bulk's promise about that is precise and routinely misquoted:
Bulk is guarantee-correct and bandwidth-best-effort.
Those are two separate promises, and collapsing them into best-effort — which is how Bulk is usually described — throws away the half that matters.
1. Two Promises, Not One
| Bulk promises | Bulk does not promise | |
|---|---|---|
| Correctness | data arrives intact or the error is visible | — |
| Bandwidth | — | any particular rate |
| Latency | — | any bound at all |
| Ordering | within an endpoint, yes | — |
The first row is absolute. A bulk transfer does not deliver corrupted data silently. Either the bytes are right or somebody is told they are not, and Chapter 10.1 §2's grid placed Bulk in the retry is useful column precisely because that guarantee is what retry buys.
The other rows are empty, which Chapter 10.3 covered as a service contract.
2. The Recovery Ladder
Correctness is not delivered by one mechanism. It is delivered by five, each recovering from a strictly larger class of failure at a strictly higher cost.
| Rung | Mechanism | Recovers from | Cost | Chapter |
|---|---|---|---|---|
| 1 | Discard — CRC failed, stay silent | a corrupted packet | one timeout | 11.6 |
| 2 | Toggle — absorb a duplicate | a lost handshake | one transaction | 11.2 |
| 3 | Retry — the host tries again | a lost packet, a busy device | one transaction | 12.5 |
| 4 | Halt — STALL, then ClearFeature | a condition the device cannot clear alone | a control transfer | 11.3 |
| 5 | Bus reset | the two ends disagree irrecoverably | re-enumeration | 8.3 |
Two properties of the ladder matter more than any individual rung.
It is ordered by cost, and the costs are wildly different. A discard costs a timeout; a bus reset costs the device's address, its configuration, every endpoint's state and a full re-enumeration. Five orders of magnitude separate the ends.
And the rungs are not interchangeable. Each recovers from a class of failure, and reaching for the wrong one fails in one of two specific ways:
Too low a rung recovers from nothing — the failure recurs immediately, forever.
Too high a rung recovers, and destroys state that was fine.
§6 measures both directions, and the asymmetry between them is §7's subject.
3. Where the Guarantee Actually Comes From
Worth being concrete, because bulk is reliable sounds like a property of the wire and is not.
Nothing about the wire is reliable. Chapter 11.6 §2 measured the CRC's strength — every single-bit error, every burst shorter than the register — and beyond that it is probabilistic.
The guarantee comes from the composition:
- The CRC turns a corrupted packet into a missing one — a packet that fails it is discarded, so the receiver never acts on damaged data.
- The timeout turns a missing packet into a retried one — Chapter 12.5: the host waits, hears nothing, and tries again.
- The toggle turns a duplicated packet into an absorbed one — Chapter 11.2: the retry that was not needed is recognised and discarded.
Each mechanism converts one failure mode into another that the next mechanism handles, and the last one converts it into nothing.
Which is why removing any single rung does not degrade the guarantee — it destroys it. §6's mutations are each the removal of one link in that chain, and every one of them produces silent data loss or silent duplication rather than a reduced error rate.
4. The Two Failure Directions
§2's claim, made precise, because the two errors have opposite symptoms.
Reaching too low — answering a permanent condition with a retry:
- the condition does not clear;
- the host retries, forever;
- every packet is well formed and no error is reported anywhere.
Chapter 11.3 §2 measured this as the NAK-for-STALL inversion, and it produces a transfer that never completes and never fails.
Reaching too high — answering a transient condition with a reset:
- the condition would have cleared on its own;
- the device loses its address, configuration and every endpoint's toggle;
- the transfer that was about to succeed is abandoned, along with every other transfer in flight.
A device that over-escalates is not being cautious. It is converting a local, recoverable failure into a global, unrecoverable one.
And the asymmetry is in who notices. Too low is silent and permanent; too high is loud and recovers. Which makes too-high the one that gets fixed and too-low the one that ships — the same ranking Chapter 11.2 §6 found between severity and difficulty.
5. The Recovery Classifier, as RTL
// ─────────────────────────────────────────────────────────────────────────
// usb_rel_pkg + usb_bulk_recovery
//
// Classification: SIMPLIFIED SYNTHESIZABLE TEACHING RTL plus a CONCEPTUAL
// package. It maps a failure event onto the LOWEST rung of section 2's
// ladder that can recover from it.
//
// WHAT IT MODELS. Section 2's ladder and section 4's two directions -- and
// nothing about how any rung is implemented, because each of them is a whole
// chapter elsewhere.
//
// WHAT IT DOES NOT MODEL. The CRC (Chapter 11.6), the toggle (Chapter 11.2),
// the timeout (Chapter 12.5), the STALL packet (Chapter 11.3) or the bus
// reset (Chapter 8.3) -- all five arrive here as classified EVENTS. Nor the
// data itself, nor the retry's cost, which is Chapter 14.3.
//
// ── WHY THE ENUM IS ORDERED ─────────────────────────────────────────────
// RC_NONE < RC_DISCARD < ... < RC_BUS_RESET is deliberate: "too high a rung"
// (section 4) is then a COMPARISON rather than a table, and the property
// that a recoverable failure never over-escalates is one line. An unordered
// encoding would make that property a case statement, which is a second
// place for the ladder's order to be written down wrongly.
// ─────────────────────────────────────────────────────────────────────────
package usb_rel_pkg;
// Section 2, weakest first. Each rung recovers from a strictly larger
// class of failure and costs strictly more.
typedef enum logic [2:0] {
RC_NONE = 3'd0,
RC_DISCARD = 3'd1, // CRC failed: drop it, say nothing
RC_TOGGLE = 3'd2, // a duplicate: absorb it
RC_RETRY = 3'd3, // timeout or NAK: the host tries again
RC_HALT = 3'd4, // STALL: the endpoint needs explicit clearing
RC_BUS_RESET = 3'd5 // the relationship is rebuilt
} recovery_e;
typedef enum logic [2:0] {
EV_OK = 3'd0,
EV_CRC_BAD = 3'd1,
EV_DUPLICATE = 3'd2,
EV_NO_REPLY = 3'd3,
EV_NAK = 3'd4,
EV_STALL = 3'd5,
EV_LOST_SYNC = 3'd6
} event_e;
endpackage
module usb_bulk_recovery
import usb_rel_pkg::*;
(
input logic clk,
input logic rst_n,
input logic ev_valid,
input event_e ev,
input logic halt_cleared, // ClearFeature(ENDPOINT_HALT)
input logic bus_reset,
output recovery_e action,
output logic data_delivered, // this event moved payload upward
output logic data_lost, // payload was lost, AND that is reported
output logic endpoint_halted,
output logic needs_host_action
);
logic halted_q;
assign endpoint_halted = halted_q;
always_comb begin
action = RC_NONE;
data_delivered = 1'b0;
data_lost = 1'b0;
if (ev_valid) begin
if (halted_q) begin
// Chapter 11.3: a halted endpoint answers nothing else until it is
// cleared. Classifying events normally while halted (section 6's R1)
// means the device services traffic it has told the host to stop.
action = RC_HALT;
end else begin
case (ev)
EV_OK: begin action = RC_NONE; data_delivered = 1'b1; end
// Chapter 11.3 section 3: SILENCE. The host's timeout is what
// turns this into a retry -- answering it (section 6's R2) asserts
// something the device does not know.
EV_CRC_BAD: begin action = RC_DISCARD; end
// Chapter 11.2: absorbed and acknowledged, NOT handed upward.
// Delivering it (section 6's R3) duplicates data silently.
EV_DUPLICATE: begin action = RC_TOGGLE; end
EV_NO_REPLY: begin action = RC_RETRY; end
EV_NAK: begin action = RC_RETRY; end
// Section 4: a permanent condition retried is retried forever.
EV_STALL: begin action = RC_HALT; end
// The only rung that admits loss -- and it REPORTS it, which is
// the whole of section 1's guarantee: data may be lost, but never
// lost silently.
EV_LOST_SYNC: begin action = RC_BUS_RESET; data_lost = 1'b1; end
default: begin action = RC_NONE; end
endcase
end
end
end
// Anything at or above RC_HALT is outside what the device can do alone --
// Chapter 12.4 section 2: every recovery mechanism is host-initiated.
assign needs_host_action = (action >= RC_HALT);
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) halted_q <= 1'b0;
else if (bus_reset) halted_q <= 1'b0;
else if (halt_cleared) halted_q <= 1'b0;
else if (ev_valid && (ev == EV_STALL)) halted_q <= 1'b1;
end
endmoduleWhat it models. The mapping from a failure event to the lowest rung that recovers from it.
Engineering reason. Because the rungs are not interchangeable, and both directions of getting it wrong have specific, opposite consequences.
Inputs. A classified failure event, the two ways a halt is cleared.
State retained. One flip-flop — the halt, which is the only condition here that persists.
Outputs. The chosen rung, whether payload moved, whether it was lost-and-reported, the halt, and whether the host must act.
Hardware implied. A case and one comparator.
Reset behaviour. The halt clears on hard reset and bus reset — the latter because a bus reset is rung 5, which subsumes rung 4.
Assumptions. That events arrive already classified — this block does not detect failures, it responds to them; that EV_LOST_SYNC means a disagreement the toggle cannot absorb; and that halt_cleared comes from an actual ClearFeature.
Omissions. Every mechanism, the data, and the retry's cost — in the header.
What DV should verify. That a corrupted packet is never answered; that a duplicate is never delivered twice; that a permanent condition is never retried; that a halted endpoint does nothing but report its halt; that any failure the device cannot fix escalates; and — the direction that is easy to omit — that a recoverable failure never reaches for a rung above retry.
6. Mutation Test
Five mutations over 623 events — every event type exhaustively, in both halted and unhalted states, a halt cleared by ClearFeature and separately by a bus reset, and 600 randomised events with halts cleared at random.
The unmutated block: 49 deliveries, 57 reported losses, 312 halts, 109 retries, and 0 violations of every obligation.
| corrupt answered | duplicate delivered | permanent retried | halted acted normally | loss unreported | failed to escalate | |
|---|---|---|---|---|---|---|
| golden | 0 | 0 | 0 | 0 | 0 | 0 |
| R1 halt ignored | 0 | 0 | 0 | 218 | 0 | 0 |
| R2 answer a bad CRC | 45 | 0 | 0 | 0 | 0 | 0 |
| R3 deliver a duplicate | 0 | 51 | 0 | 0 | 0 | 0 |
| R4 retry a STALL | 0 | 0 | 53 | 0 | 0 | 53 |
| R5 lost sync → retry | 0 | 0 | 0 | 0 | 57 | 57 |
R1 — classify events normally while halted
Measured: 218 events serviced by an endpoint that had told the host to stop.
The device STALLs, the host is supposed to stop and issue ClearFeature — and the device keeps working. Which sounds harmless and is not: the host and device now disagree about whether the endpoint is usable, and Chapter 12.4 §2's divergence follows.
R2 — answer a corrupted packet
Measured: 45 corrupted packets answered.
§3's first conversion broken. A corrupted packet is no longer turned into a missing one — it is answered, which asserts that something was received that the device cannot vouch for. Chapter 11.3 §3's argument in full: the device does not even know the packet was addressed to it.
R3 — deliver a duplicate upward
Measured: 51 duplicates delivered.
§3's third conversion broken. The retry that the toggle exists to absorb is handed to the application, so a lost handshake becomes duplicated data — silently, because every packet involved was valid.
This is the mutation that breaks §1's guarantee most directly. Not lost data — duplicated data, which no checksum above catches because each copy is individually correct.
R4 — retry a STALL
Measured: 53 permanent conditions retried, all 53 also failing to escalate.
§4's too low direction. The condition cannot clear, so the host retries forever and the endpoint spins — Chapter 11.3 §2's measured failure, arriving here as a ladder error rather than a packet error.
R5 — answer an irrecoverable disagreement with a retry
Measured: 57 losses unreported, and 57 failures to escalate.
§4's too low direction again, one rung further out. The two ends disagree in a way no retry can fix, and retrying produces the same disagreement each time.
The important column is the first one. With the bus-reset rung removed, data_lost is never asserted — so data is lost and nobody is told, which is the precise violation of §1's guarantee.
7. The Assertions
// ─────────────────────────────────────────────────────────────────────────
// Classification: TEACHING ASSERTIONS about the recovery ladder.
// N-properties are the NEVER rules of sections 2 to 4; E the escalation
// contract. Every one is stated against the EVENT and the RUNG, never
// against the design's internal case -- Chapter 11.3 section 8's rule.
// ─────────────────────────────────────────────────────────────────────────
// N1 -- A CORRUPTED PACKET IS NEVER ANSWERED. Section 6's R2.
property p_corrupt_is_silent;
@(posedge clk) disable iff (!rst_n)
(ev_valid && !endpoint_halted && (ev == EV_CRC_BAD)) |-> (action == RC_DISCARD);
endproperty
assert property (p_corrupt_is_silent);
// N2 -- A DUPLICATE IS NEVER DELIVERED. Section 6's R3, and the mutation
// that breaks section 1's guarantee most directly.
property p_duplicate_not_delivered;
@(posedge clk) disable iff (!rst_n)
(ev_valid && !endpoint_halted && (ev == EV_DUPLICATE)) |-> !data_delivered;
endproperty
assert property (p_duplicate_not_delivered);
// N3 -- A PERMANENT CONDITION IS NEVER RETRIED. Section 4's "too low"
// direction, and section 6's R4.
property p_permanent_not_retried;
@(posedge clk) disable iff (!rst_n)
(ev_valid && !endpoint_halted && (ev == EV_STALL)) |-> (action == RC_HALT);
endproperty
assert property (p_permanent_not_retried);
// N4 -- A RECOVERABLE CONDITION NEVER OVER-ESCALATES. Section 4's "too high"
// direction -- the one that is easy to omit, because over-escalation looks
// like caution. The ORDERED enum is what makes this one comparison.
property p_no_over_escalation;
@(posedge clk) disable iff (!rst_n)
(ev_valid && !endpoint_halted && ((ev == EV_NAK) || (ev == EV_NO_REPLY)))
|-> (action <= RC_RETRY);
endproperty
assert property (p_no_over_escalation);
// N5 -- A HALTED ENDPOINT DOES NOTHING ELSE. Section 6's R1.
property p_halted_only_halts;
@(posedge clk) disable iff (!rst_n)
(ev_valid && endpoint_halted) |-> (action == RC_HALT);
endproperty
assert property (p_halted_only_halts);
// E1 -- WHAT THE DEVICE CANNOT FIX, IT ESCALATES. Chapter 12.4 section 2:
// every recovery mechanism above the device's own reach is host-initiated,
// so the device's only contribution is to SAY SO.
property p_escalates;
@(posedge clk) disable iff (!rst_n)
(ev_valid && (action >= RC_HALT)) |-> needs_host_action;
endproperty
assert property (p_escalates);
// E2 -- AND DATA IS NEVER LOST WITHOUT A REPORT. Section 1's guarantee,
// stated as the single property that expresses it. Section 6's R5 violates
// exactly this and nothing else.
property p_no_silent_loss;
@(posedge clk) disable iff (!rst_n)
(ev_valid && !endpoint_halted && (ev == EV_LOST_SYNC)) |-> data_lost;
endproperty
assert property (p_no_silent_loss);N4 is the property that is usually missing, and the reason is psychological rather than technical: over-escalation reads as caution. A device that resets the bus when it sees a NAK is doing something safe-sounding, and nothing about it looks like a defect until you notice it has discarded a working configuration to solve a problem that would have cleared itself.
And N4 is one comparison only because the enum is ordered. With an unordered encoding it would be a case statement enumerating the acceptable rungs — a second place for the ladder's order to be written down, and therefore a second place for it to be wrong.
E2 is §1's guarantee as a single line. Data may be lost; it may never be lost silently. Everything else in this chapter exists to make that true.
8. Verification
This chapter's commit point is every failure got the cheapest rung that could actually fix it, and nothing was lost without a report.
Stimulus. Every event type, exhaustively, in both the halted and unhalted states — 14 combinations; a halt cleared by ClearFeature; a halt cleared by a bus reset; and 600 randomised events with halts cleared at random intervals.
Observation. The chosen rung, both data flags, the halt, and the escalation indication — against seven obligations stated from §§2–4 rather than from the design, which is what makes the bench able to disagree with the model.
Reference model. None, deliberately — the seven obligations do the work. A reference model here would be a second copy of the ladder, and Chapter 11.3 §8 measured what a model sharing the design's understanding is worth. The obligations are each one sentence from §§2–4, phrased in the protocol's terms.
Coverage — crosses:
- every event × halted and unhalted — all 14
- halt cleared by
ClearFeature× by bus reset EV_STALLarriving while already halted- consecutive failures of the same class
- a failure immediately after a halt is cleared
Negative cases with defined outcomes: a corrupted packet is never answered; a duplicate is never delivered; a permanent condition is never retried; a transient one never over-escalates; a halted endpoint does nothing else; and EV_LOST_SYNC always reports the loss.
9. Debugging: the Transfer That Completes Twice
An application reading from a bulk endpoint occasionally receives the same block of data twice in a row. The duplicate is byte-identical. No error is reported by any layer. The rate rises on a busy bus and falls to zero on a quiet one.
What does byte-identical tell you? That it is not corruption. It is the same data delivered twice, which is a completely different failure from data being wrong.
What mechanism exists specifically to prevent this? Chapter 11.2's toggle — §2's rung 2. Its entire job is absorbing a retransmission the sender made because it did not hear an acknowledgement.
Why would the rate rise with bus activity? Because retransmissions require lost handshakes, and lost handshakes scale with traffic. Load does not cause the bug; it causes the precondition — the same shape as Chapter 11.2 §9's.
So what is broken? §6's R3: the device recognises the duplicate and delivers it anyway. The toggle logic works; what it feeds does not.
How do you distinguish that from a toggle that does not work at all? By whether the duplicate is acknowledged. A device with no toggle logic would mis-sequence and eventually wedge — Chapter 11.2 §6's D1. A device that classifies correctly and delivers anyway keeps working perfectly, and only the application sees anything wrong.
Is there an alternative explanation? Yes, and it must be ruled out first: the application itself may be re-reading. Cheap to check, and being wrong about it is expensive.
And what makes this class of bug expensive? No layer reports an error, because none occurred. Every packet was valid, every handshake was correct, the transfer completed successfully — twice. The only evidence is in the data, and the only party that can see it is the one furthest from the cause.
The signature to keep: byte-identical duplicates at a rate that scales with bus load, with no errors anywhere, is a duplicate that was detected and delivered — and the detection working is what makes it invisible.
10. Common Misconceptions
11. Reason It Through
A device's firmware author proposes that on any bulk transfer error — CRC failure, timeout, NAK storm, anything — the device should stall the endpoint. The argument: a stall is unambiguous, the host knows exactly what to do, and it avoids the device having to classify failures it may not understand.
Is the premise true? Mostly. A STALL is unambiguous, and ClearFeature(ENDPOINT_HALT) is a well-defined recovery that also resynchronises the toggle.
So what is wrong with it? §4's too high direction. Most of the events being stalled would have cleared themselves at rung 1, 2 or 3 — a CRC failure needs a discard, a lost handshake needs the toggle, a busy moment needs a NAK.
What does stalling them cost? A control transfer per event, plus the transfer in progress, plus every queued transfer on that endpoint — Chapter 8.5. On a bus with a 1-in-10⁶ bit error rate and 45 MB/s of traffic, that is several halts per second.
Is the device still correct? Yes, which is what makes the proposal seductive. No data is lost, no duplicates are delivered, and every halt is honestly reported. It satisfies §1's guarantee completely.
So the objection is purely about throughput? No, and this is the part worth getting right. A halt is also visible to the driver, so the operating system logs an endpoint error several times a second. A device that reports constant errors while working perfectly trains everyone who sees it to ignore its error reports — and the one halt that matters arrives among thousands that do not.
What is the correct version of the author's concern? That classification is hard, and a device unsure what happened should escalate rather than guess. That is right, and the answer is to classify what can be classified and escalate the residue — not to escalate everything.
And the transferable point: an error-reporting mechanism has a signal-to-noise ratio, and a design that reports recoverable events as errors destroys it. The cost is not the reports; it is that the reports stop being read — which is a failure mode with no measurement and no alarm.
12. Understanding Check
13. Summary
Bulk makes two promises, not one: it is guarantee-correct and bandwidth-best-effort. The usual one-line summary keeps only the second, and applying best-effort to the data produces a specific design error — a second retry layer above a channel that already retries. Bulk is reliable but unpunctual, which needs buffering and timeouts rather than checksums and sequence numbers.
The guarantee comes from a chain of conversions, not from any single mechanism: the CRC turns corruption into absence, the timeout turns absence into a retry, the toggle turns the retry into nothing. Above them sit halt and bus reset, for what the device cannot fix alone.
The rungs are ordered by cost and are not interchangeable, and both directions of error are specific:
Too low a rung recovers from nothing — silent, permanent, every packet well formed.
Too high a rung recovers and destroys state that was fine — loud, global, and it reads as caution.
§6 measured five mutations over 623 events, and the finding is what they have in common:
Not one produced a higher error rate. Each let the original failure through unchanged and unreported — a device asserting what it cannot know, byte-identical data delivered twice, a transfer that never completes and never fails, or data lost with no report.
A correctness guarantee assembled from converting mechanisms is all-or-nothing, so a designer dropping a rung to save logic is not trading reliability for area — they are trading a guarantee for none, and no measurement will warn them.
§8's bench finding is Chapter 14.1's inverted. There the bench reported a value it should have checked; here it asserted a state it should have observed — passing a constant not halted through a loop in which EV_STALL halts the endpoint halfway. Both substitute the bench's belief for an observation, and this one presents as a correct design reporting violations, which invites weakening a check that was right.
A bench that tells the design what state it is in has stopped testing and started narrating.
14. What Comes Next
Rung 3 is where almost all real recovery happens, and this chapter treated it as free.
Chapter 14.3 prices it. A NAK is not an error — it is flow control, and the most common answer a busy bulk endpoint gives. But it costs a full transaction's bus time for zero bytes, which Chapter 14.1 §4's arithmetic already priced at about 927 ns at high speed.
The chapter is about what happens when that is the usual answer: a device that NAKs half the time has halved its throughput without anything being wrong, and the host's polling strategy — not the device's — decides how much of the bus that wastes.
Browse the full path on the USB tutorials index.
Continue learning
Related tutorials
- Related topic
Four Transfer Types Overview
The four transfer types are not a list to memorise — they fall out of two orthogonal questions plus one bootstrap question. Deriving them, and the policy block whose outputs are all derived and none stored.
- Related topic
Bulk Transfers
Promised nothing and permitted everything: why throughput and guarantee are different axes, and why a fixed-priority arbiter starves an endpoint forever while every safety property passes.
- Related topic
Bulk Throughput
Why 480 Mbit/s never means 60 MB/s — overhead is per transaction and fixed, so efficiency is a function of packet size alone, and a zero-length terminator costs a full transaction for no bytes.
- Related topic
Bulk Retries
A NAK costs a full transaction for zero bytes — and is still eleven times cheaper than a data transaction, so a device ready one time in ten keeps 56% of its throughput rather than 10%.
Standards & specifications
- Governing standard
- USB-IF (Universal Serial Bus Specification)(opens USB Implementers Forum (USB-IF) in a new tab)
Defines the USB bus — its electrical signalling, connectors, packet and transaction model, device framework and the descriptors a device must expose — together with the device-class specifications layered on it. It does not define host-controller register interfaces (xHCI and EHCI are separate documents) nor any operating system's driver architecture.
This page also covers RTL structure, verification approach and debugging technique. Those are engineering practice built on the standard, not requirements the standard itself imposes.
Where this fits
Part of the USB curriculum.
