USB · Module 14
Bulk Retries
A NAK costs a full transaction for zero bytes — and is still eleven times cheaper than a data transaction, so a device ready one time in ten keeps 56% of its throughput rather than 10%.
Chapter 14.2 treated rung 3 of the recovery ladder as free. This chapter prices it — and the price is not what the folklore says.
A NAK is not an error. Chapter 11.3 §2 established that it is flow control, and that on a typical bulk endpoint NAKs outnumber ACKs. What it did not do is say what they cost.
1. A NAK Costs a Full Transaction for Zero Bytes
A NAKed transaction is still a transaction: a token, an attempt at data, a handshake, two turnarounds. Only the payload is missing.
At high speed a NAKed bulk transaction costs about 927 ns and moves zero bytes.
Which sounds ruinous and is not, because Chapter 14.1 §2 priced the alternative:
| Bus time | Payload | |
|---|---|---|
| A full 512-byte transaction | 10 880 ns | 512 B |
| A NAKed transaction | 927 ns | 0 B |
| Ratio | 11.7 : 1 | — |
A NAK is about a twelfth of a data transaction. It has no payload and no per-byte cost — only the fixed overhead, and even that is smaller because there is no data packet to send.
So the intuition that a NAK-heavy endpoint is catastrophically slow is quantitatively wrong, and §2 measures by how much.
2. The Measurement
A device polled 100 times, ready a varying fraction of the time, with every transaction's bus time accumulated:
| Device ready | Good | NAK | Bus time | Wasted | Bus carried payload | MB/s | vs ideal |
|---|---|---|---|---|---|---|---|
| 1 in 10 | 10 | 90 | 192 250 ns | 83 430 | 56% | 26.63 | 56% |
| 2 in 10 | 20 | 80 | 291 800 | 74 160 | 74% | 35.09 | 74% |
| 3 in 10 | 30 | 70 | 391 350 | 64 890 | 83% | 39.24 | 83% |
| 5 in 10 | 50 | 50 | 590 450 | 46 350 | 92% | 43.35 | 92% |
| 8 in 10 | 80 | 20 | 889 100 | 18 540 | 97% | 46.06 | 97% |
| 10 in 10 | 100 | 0 | 1 088 200 | 0 | 100% | 47.05 | 100% |
3. The Cost Is Paid by Somebody Else
§2 measured the endpoint's own throughput. The bus time is a different quantity with a different owner.
At ready-1-in-10 the endpoint consumed 192 250 ns to deliver 5 120 bytes — and 83 430 ns of that moved nothing. That time was not idle. It was occupied, and Chapter 10.3 §1 established that bulk runs in the residue left by everything else.
A NAKing endpoint does not slow itself down as much as it slows down everything sharing the bus with it.
And the asymmetry is the point:
| Loses | |
|---|---|
| The NAKing endpoint | 44% of its own throughput |
| Every other bulk endpoint | the full 83 430 ns of residue that was consumed for nothing |
On a single-device bus this costs nothing that matters — the bus was idle anyway. On a shared bus the device is spending a common resource on its own unreadiness, which is Chapter 10.3 §4's fairness question arriving as a throughput one.
4. Who Decides How Often to Poll
The device does not. Chapter 10.1 §1: the host initiates every transaction, so the NAK rate is set by how often the host asks, not by how often the device is ready.
Which means the waste of §3 is under the host's control and not the device's. A device that NAKs has said not yet; whether that costs 927 ns once or a hundred times is the host's decision.
Hosts do back off, and the strategy is theirs — the specification does not mandate one. A host that polls a NAKing endpoint as fast as the bus allows is compliant and wasteful; one that backs off is compliant and slower to notice readiness.
5. The Retry-Cost Model, as RTL
// ─────────────────────────────────────────────────────────────────────────
// usb_bulk_retry_cost
//
// Classification: SIMPLIFIED SYNTHESIZABLE TEACHING RTL -- and in practice
// an instrumentation block, because its outputs are measurements rather
// than control. A real controller that wanted section 3's number would look
// very like this.
//
// WHAT IT MODELS. Sections 1 to 3: what each poll costs depending on whether
// the device was ready, and how much of the total moved nothing.
//
// WHAT IT DOES NOT MODEL. The packets (Module 11) or transactions (Module
// 12); WHY the device is unready, which is Chapter 9.5's buffering; the
// host's polling strategy (section 4) -- `attempt` arrives from outside; and
// errors, which are Chapter 14.2's ladder. Every NAK here is honest flow
// control, never a failure.
//
// ── ON SEPARATING bus_ns FROM wasted_ns ─────────────────────────────────
// They are the same measurement at different granularities, and a design
// that reports only the first cannot answer section 3's question -- which is
// not "how fast am I" but "how much of the shared bus did I spend on
// nothing". Section 6's K3 removes the second and nothing else, and the
// endpoint's own behaviour is unchanged.
// ─────────────────────────────────────────────────────────────────────────
module usb_bulk_retry_cost #(
parameter int unsigned NAK_NS = 927, // a NAKed transaction
parameter int unsigned OVERHEAD_NS = 2347, // a data transaction's fixed part
parameter int unsigned NS_PER_B_X100 = 1667,
// Section 4: how long the host waits before re-polling. Zero means "as
// fast as the bus allows", which is what a host does by default and what
// makes an unready device expensive to everything else.
parameter int unsigned BACKOFF_NS = 0
)(
input logic clk,
input logic rst_n,
input logic start,
input logic attempt, // the host polls now
input logic dev_ready, // we have data, or space for it
input logic [15:0] pkt_len,
output logic got_data,
output logic [31:0] bytes_moved,
output logic [31:0] bus_ns,
output logic [31:0] wasted_ns, // of bus_ns, the part that moved nothing
output logic [15:0] naks,
output logic [15:0] goods
);
logic [31:0] bytes_q, ns_q, waste_q;
logic [15:0] nak_q, good_q;
assign bytes_moved = bytes_q;
assign bus_ns = ns_q;
assign wasted_ns = waste_q;
assign naks = nak_q;
assign goods = good_q;
// A poll succeeds exactly when the device is ready. There is no third
// outcome here -- errors are Chapter 14.2's subject, and conflating an
// error with a NAK is that chapter's whole warning.
assign got_data = attempt && dev_ready;
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
bytes_q <= '0; ns_q <= '0; waste_q <= '0; nak_q <= '0; good_q <= '0;
end else if (start) begin
bytes_q <= '0; ns_q <= '0; waste_q <= '0; nak_q <= '0; good_q <= '0;
end else if (attempt) begin
if (got_data) begin
good_q <= good_q + 16'd1;
bytes_q <= bytes_q + pkt_len;
ns_q <= ns_q + OVERHEAD_NS + ((pkt_len * NS_PER_B_X100) / 100);
end else begin
nak_q <= nak_q + 16'd1;
// SECTION 1. A refusal is a whole transaction: token, attempt,
// handshake, two turnarounds. It is CHEAP relative to a data
// transaction and it is not free, and section 6's K2 is the
// difference between those two statements.
ns_q <= ns_q + NAK_NS + BACKOFF_NS;
waste_q <= waste_q + NAK_NS;
end
end
end
endmoduleWhat it models. The bus time a sequence of polls consumes, split by whether each one moved anything.
Engineering reason. Because the endpoint's throughput and the bus time it spends are different quantities with different owners, and only the second is a shared resource.
Inputs. A poll, the device's readiness, and the packet size.
State retained. Two 32-bit accumulators, a waste accumulator, and two counters.
Outputs. Whether this poll succeeded, and four running measurements.
Hardware implied. Three accumulators and a constant multiply.
Reset behaviour. All counters cleared.
Assumptions. That every poll is answered — a timeout is Chapter 14.2's territory and costs more than a NAK; that NAK_NS and OVERHEAD_NS match the link speed (Chapter 12.5 §1); and that BACKOFF_NS models the host's strategy, which the device does not control and cannot influence.
Omissions. Packets, transactions, the cause of unreadiness, the polling strategy and all errors — in the header.
What DV should verify. That a poll succeeds exactly when the device is ready; that every poll is charged, a NAKed one included; that the wasted time counts only NAKs; that the counts and the byte total agree; and that the two accumulators stay consistent with each other.
What each poll cost
10 cycles6. Mutation Test
Three mutations over 2 313 polls — a sweep of readiness from 1-in-10 to 10-in-10 at 100 polls each, directed corners at always-ready, 1-in-10 and 1-in-50, and 40 randomised runs.
The unmutated block reproduced §2's table exactly, and reported 0 errors of every kind.
got_data wrong | bytes wrong | bus time wrong | wasted wrong | counts wrong | |
|---|---|---|---|---|---|
| golden | 0 | 0 | 0 | 0 | 0 |
| K1 readiness ignored | 1 046 | 49 | 49 | 49 | 49 |
| K2 a NAK costs nothing | 0 | 0 | 49 | 49 | 0 |
| K3 waste not attributed | 0 | 0 | 0 | 49 | 0 |
K1 — every poll succeeds
Measured: 1 046 wrong outcomes, and every accumulator wrong with them.
The device delivers data it does not have. Catastrophic, immediate, and included as the contrast for K3.
K2 — charge nothing for a NAK
Measured: 49 runs with wrong bus time and wrong waste. The byte counts and the good/NAK counts are perfect.
§1's cheap misread as free. The endpoint's behaviour is entirely correct — it NAKs when it should, delivers when it should, and the packet stream is indistinguishable from the golden design's.
What is wrong is the number it reports about itself, and §3 is why that matters: a host or a driver using this figure to decide whether an endpoint is worth polling is reading a figure that omits the cost of polling it.
K3 — do not attribute the waste
// waste_q <= waste_q + NAK_NS; // MUTANT K3: removedMeasured: 49 runs with wrong wasted_ns. Bus time correct. Bytes correct. Counts correct. Behaviour identical.
The total is right and the breakdown is gone.
7. Verification
This chapter's commit point is every poll was charged what it actually cost, and the part that bought nothing was identified.
Stimulus. A readiness sweep from 1-in-10 to 10-in-10 at 100 polls each — which is what produces §2's table rather than a pass/fail; directed corners at always-ready, 1-in-10 and 1-in-50; and 40 randomised runs of varying length and readiness.
Observation. The success of each poll, both accumulators, and both counters — against expectations computed poll by poll in the bench, not from the totals.
Reference model. A per-poll accumulation in the bench: ready means overhead plus per-byte cost, unready means the NAK cost. Four lines, and it is independent because it is driven from the readiness pattern rather than from the design's outputs.
Coverage — crosses:
- readiness from 1-in-10 to 10-in-10, every step
- readiness 1-in-50 — the sparse extreme
- run lengths from 1 to 60 polls
- always-ready and never-ready endpoints
- packet sizes at the endpoint maximum
Negative cases with defined outcomes: no poll succeeds while the device is unready; no poll is uncharged; the wasted total never exceeds the bus total; and the wasted total counts only NAKs.
8. Debugging: the Device That Slows Down Its Neighbours
A bulk device achieves acceptable throughput. When it is attached, an unrelated device on the same hub loses about a third of its throughput. Removing the first device restores the second. The first device reports no errors and its own throughput is within specification.
What does no errors and acceptable throughput rule out? A malfunction. The device is working, and the effect on its neighbour is a consequence of how it works rather than of anything going wrong.
What can one bulk endpoint take from another? Chapter 10.3: the residue. Bulk endpoints share what periodic traffic leaves, so one consuming more leaves less.
But the device's throughput is normal — so it is not consuming more bus time for data. What else consumes bus time? §1: NAKs.
Does the arithmetic work? §2's table: at ready-1-in-10, an endpoint spends 192 250 ns to deliver 5 120 bytes, of which 83 430 ns bought nothing — 43% of its bus consumption is waste. A neighbour sharing the residue loses that.
How do you confirm it from outside? Count NAKs in a bus trace. A device NAKing heavily is unmistakable in a capture and invisible in any throughput figure — because §2's whole finding is that heavy NAKing barely dents the endpoint's own rate.
What is the fix? §4: the device's only lever is being ready more often, which means buffering — Chapter 9.5. And the host's lever is backing off, which the device cannot request.
Is the device at fault? Not in any way that a compliance test would catch, and that is the honest answer. It is compliant, functional, and a bad neighbour. The specification does not bound how often a device may NAK, so this is a quality-of-implementation issue rather than a conformance one.
The signature to keep: a device that harms its neighbours while measuring fine itself is spending shared bus time on refusals — and the only instrument that shows it is a NAK count, which no throughput measurement contains.
9. Common Misconceptions
10. Reason It Through
A device's firmware author adds a 50 µs delay before answering when the endpoint has no data, reasoning that it reduces the number of NAKs the host receives and therefore the bus time consumed.
Does it reduce the NAK count? Yes — the host cannot poll again until the current transaction resolves, so a slower answer means fewer polls per second.
Does it reduce bus time? No. It increases it dramatically. Chapter 12.5 §2: the device must respond within 7.5 bit times — about 15 ns at high speed. A 50 µs delay is three thousand times that bound.
So what actually happens? The device misses the turnaround window entirely, the host times out, and — Chapter 14.2 §2's rung 3 — retries. The device does this again. A transaction that would have cost 927 ns now costs a full timeout plus a retry, repeatedly.
Is the reasoning salvageable at a smaller delay? No, and this is the important part. Any delay in answering is wrong, because the answer is due within the turnaround window and there is no legal value above it. The lever the author is reaching for does not exist at the transaction level.
Where does the lever actually exist? §4: in the host, which chooses the polling rate, and in the device's readiness, which is buffering. Neither is reachable from the transaction response path.
Is there anything the device can do that resembles the intent? Yes, and it is the useful reframing: be ready more often. Buffering does not reduce the NAK cost — it reduces the NAK count, by converting refusals into deliveries. That is the same goal, achieved at the only layer that can achieve it.
And the transferable point: a protocol's timing bounds are not a budget to spend but a window to hit, and an optimisation that treats a deadline as a resource converts a cheap correct behaviour into an expensive incorrect one. The author's instinct — fewer NAKs is better — was right; the mechanism was the one layer that cannot deliver it.
11. Understanding Check
12. Summary
A NAK costs a full transaction for zero bytes — about 927 ns at high speed — and that is about a twelfth of a 512-byte data transaction's 10 880 ns.
Which makes NAK-based flow control efficient rather than ruinous. Measured across a readiness sweep: a device ready 1 in 10 keeps 56% of ideal throughput, not 10%. Nine refusals together cost less than one delivery.
And the curve is steep at the wrong end — 1-in-10 to 2-in-10 buys eighteen percentage points; 9-in-10 to always-ready buys one. Effort spent eliminating the last few NAKs is nearly worthless.
The endpoint's throughput and the bus time it spends are different quantities with different owners. At ready-1-in-10 the endpoint lost 44% of its own rate and consumed 83 430 ns of shared residue that moved nothing — which is what a neighbour on the same bus pays.
And the device cannot influence any of it. There is no poll me later packet; a NAK carries no timing information; the polling rate is the host's alone. The device's only lever is being ready, and its worst alternative is stalling — Chapter 14.2 §4's over-escalation as a throughput decision.
§6's three mutations, and K3 is the one worth keeping: removing the waste attribution changes no behaviour at all — same packets, same timing, same throughput, same total.
bus_nssays how expensive this endpoint is.wasted_nssays how much of that expense bought nothing. The first is a measurement; the second is a diagnosis.
§7's practice generalises past this chapter:
A parameter sweep is a measurement. It becomes a test when every point in it has an expected value — and the display statement and the comparison are different lines, so deleting the comparison leaves output that looks identical.
13. What Comes Next
Three chapters of Bulk as a mechanism. The next two are Bulk as a foundation — what real classes build on top of it.
Chapter 14.4 is USB Mass Storage, which is where nearly all bulk traffic in the world actually goes. Its protocol is two endpoints and three phases — a 31-byte command wrapper, a data phase, and a 13-byte status wrapper — and it is a case study in building a request/response protocol on a channel that has neither.
Its interesting failure is the one it had to invent a name for: a phase error, which is what happens when the two ends disagree about which of the three phases they are in — Chapter 12.4 §2's two beliefs, reappearing two layers up with the same structure and a new recovery mechanism.
Browse the full path on the USB tutorials index.
Continue learning
Related tutorials
- Related topic
Four Transfer Types Overview
The four transfer types are not a list to memorise — they fall out of two orthogonal questions plus one bootstrap question. Deriving them, and the policy block whose outputs are all derived and none stored.
- Related topic
Bulk Transfers
Promised nothing and permitted everything: why throughput and guarantee are different axes, and why a fixed-priority arbiter starves an endpoint forever while every safety property passes.
- Related topic
Handshake Packets
Four one-byte packets and a silence: the difference between not now and not ever, and why a reference model reported zero divergence on a design that was wrong.
- Related topic
Status Stage
A zero-length packet running the opposite way to the data stage — the only place a device can report that a request it already accepted has failed.
Standards & specifications
- Governing standard
- USB-IF (Universal Serial Bus Specification)(opens USB Implementers Forum (USB-IF) in a new tab)
Defines the USB bus — its electrical signalling, connectors, packet and transaction model, device framework and the descriptors a device must expose — together with the device-class specifications layered on it. It does not define host-controller register interfaces (xHCI and EHCI are separate documents) nor any operating system's driver architecture.
This page also covers RTL structure, verification approach and debugging technique. Those are engineering practice built on the standard, not requirements the standard itself imposes.
Where this fits
Part of the USB curriculum.
