Skip to content
VLSI Mentor

USB · Module 14

Bulk Retries

A NAK costs a full transaction for zero bytes — and is still eleven times cheaper than a data transaction, so a device ready one time in ten keeps 56% of its throughput rather than 10%.

Chapter 14.2 treated rung 3 of the recovery ladder as free. This chapter prices it — and the price is not what the folklore says.

A NAK is not an error. Chapter 11.3 §2 established that it is flow control, and that on a typical bulk endpoint NAKs outnumber ACKs. What it did not do is say what they cost.

1. A NAK Costs a Full Transaction for Zero Bytes

A NAKed transaction is still a transaction: a token, an attempt at data, a handshake, two turnarounds. Only the payload is missing.

At high speed a NAKed bulk transaction costs about 927 ns and moves zero bytes.

Which sounds ruinous and is not, because Chapter 14.1 §2 priced the alternative:

Bus timePayload
A full 512-byte transaction10 880 ns512 B
A NAKed transaction927 ns0 B
Ratio11.7 : 1—

A NAK is about a twelfth of a data transaction. It has no payload and no per-byte cost — only the fixed overhead, and even that is smaller because there is no data packet to send.

So the intuition that a NAK-heavy endpoint is catastrophically slow is quantitatively wrong, and §2 measures by how much.

2. The Measurement

A device polled 100 times, ready a varying fraction of the time, with every transaction's bus time accumulated:

Device readyGoodNAKBus timeWastedBus carried payloadMB/svs ideal
1 in 101090192 250 ns83 43056%26.6356%
2 in 102080291 80074 16074%35.0974%
3 in 103070391 35064 89083%39.2483%
5 in 105050590 45046 35092%43.3592%
8 in 108020889 10018 54097%46.0697%
10 in 1010001 088 2000100%47.05100%

3. The Cost Is Paid by Somebody Else

§2 measured the endpoint's own throughput. The bus time is a different quantity with a different owner.

At ready-1-in-10 the endpoint consumed 192 250 ns to deliver 5 120 bytes — and 83 430 ns of that moved nothing. That time was not idle. It was occupied, and Chapter 10.3 §1 established that bulk runs in the residue left by everything else.

A NAKing endpoint does not slow itself down as much as it slows down everything sharing the bus with it.

And the asymmetry is the point:

Loses
The NAKing endpoint44% of its own throughput
Every other bulk endpointthe full 83 430 ns of residue that was consumed for nothing

On a single-device bus this costs nothing that matters — the bus was idle anyway. On a shared bus the device is spending a common resource on its own unreadiness, which is Chapter 10.3 §4's fairness question arriving as a throughput one.

4. Who Decides How Often to Poll

The device does not. Chapter 10.1 §1: the host initiates every transaction, so the NAK rate is set by how often the host asks, not by how often the device is ready.

Which means the waste of §3 is under the host's control and not the device's. A device that NAKs has said not yet; whether that costs 927 ns once or a hundred times is the host's decision.

Hosts do back off, and the strategy is theirs — the specification does not mandate one. A host that polls a NAKing endpoint as fast as the bus allows is compliant and wasteful; one that backs off is compliant and slower to notice readiness.

A sequence diagram of ten bulk IN polls at an endpoint that is ready four times. The host issues an IN token and the device returns a five hundred and twelve byte data packet, costing about ten thousand eight hundred and eighty nanoseconds. The host polls three more times and the device answers each with a negative acknowledgement, costing about nine hundred and twenty-seven nanoseconds each and moving nothing. The host polls again and receives data. Two more negative acknowledgements follow, then data, then one more negative acknowledgement, then data. Across the ten polls the device delivered four packets and refused six times, and the six refusals together cost about half of what a single data transaction costs. An annotation notes that the host chooses the polling rate, not the device, and that a device has no way to ask to be polled less often.Ten polls, four answersHostBulk IN endpointINDATA 512 B — 10 880nsINNAK — 927 ns, 0bytesINNAK — 927 nsINNAK — 927 nsINDATA 512 B3 NAKs = 2 781 ns;one DATA = 10 880 nsthe HOST chose topoll — the devicecannot ask it not to
Figure 1 — ten polls at a device ready four times. The NAKs are visibly shorter than the data transactions, which is the whole of why the endpoint keeps 56% of its throughput at a 9:1 NAK ratio rather than 10%. The wasted time accumulates in small increments; the productive time in large ones.

5. The Retry-Cost Model, as RTL

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// ─────────────────────────────────────────────────────────────────────────
// usb_bulk_retry_cost
//
// Classification: SIMPLIFIED SYNTHESIZABLE TEACHING RTL -- and in practice
// an instrumentation block, because its outputs are measurements rather
// than control. A real controller that wanted section 3's number would look
// very like this.
//
// WHAT IT MODELS. Sections 1 to 3: what each poll costs depending on whether
// the device was ready, and how much of the total moved nothing.
//
// WHAT IT DOES NOT MODEL. The packets (Module 11) or transactions (Module
// 12); WHY the device is unready, which is Chapter 9.5's buffering; the
// host's polling strategy (section 4) -- `attempt` arrives from outside; and
// errors, which are Chapter 14.2's ladder. Every NAK here is honest flow
// control, never a failure.
//
// ── ON SEPARATING bus_ns FROM wasted_ns ─────────────────────────────────
// They are the same measurement at different granularities, and a design
// that reports only the first cannot answer section 3's question -- which is
// not "how fast am I" but "how much of the shared bus did I spend on
// nothing". Section 6's K3 removes the second and nothing else, and the
// endpoint's own behaviour is unchanged.
// ─────────────────────────────────────────────────────────────────────────
module usb_bulk_retry_cost #(
  parameter int unsigned NAK_NS        = 927,    // a NAKed transaction
  parameter int unsigned OVERHEAD_NS   = 2347,   // a data transaction's fixed part
  parameter int unsigned NS_PER_B_X100 = 1667,
  // Section 4: how long the host waits before re-polling. Zero means "as
  // fast as the bus allows", which is what a host does by default and what
  // makes an unready device expensive to everything else.
  parameter int unsigned BACKOFF_NS    = 0
)(
  input  logic        clk,
  input  logic        rst_n,

  input  logic        start,
  input  logic        attempt,          // the host polls now
  input  logic        dev_ready,        // we have data, or space for it
  input  logic [15:0] pkt_len,

  output logic        got_data,
  output logic [31:0] bytes_moved,
  output logic [31:0] bus_ns,
  output logic [31:0] wasted_ns,        // of bus_ns, the part that moved nothing
  output logic [15:0] naks,
  output logic [15:0] goods
);

  logic [31:0] bytes_q, ns_q, waste_q;
  logic [15:0] nak_q, good_q;

  assign bytes_moved = bytes_q;
  assign bus_ns      = ns_q;
  assign wasted_ns   = waste_q;
  assign naks        = nak_q;
  assign goods       = good_q;

  // A poll succeeds exactly when the device is ready. There is no third
  // outcome here -- errors are Chapter 14.2's subject, and conflating an
  // error with a NAK is that chapter's whole warning.
  assign got_data = attempt && dev_ready;

  always_ff @(posedge clk or negedge rst_n) begin
    if (!rst_n) begin
      bytes_q <= '0; ns_q <= '0; waste_q <= '0; nak_q <= '0; good_q <= '0;
    end else if (start) begin
      bytes_q <= '0; ns_q <= '0; waste_q <= '0; nak_q <= '0; good_q <= '0;
    end else if (attempt) begin
      if (got_data) begin
        good_q  <= good_q + 16'd1;
        bytes_q <= bytes_q + pkt_len;
        ns_q    <= ns_q + OVERHEAD_NS + ((pkt_len * NS_PER_B_X100) / 100);
      end else begin
        nak_q   <= nak_q + 16'd1;
        // SECTION 1. A refusal is a whole transaction: token, attempt,
        // handshake, two turnarounds. It is CHEAP relative to a data
        // transaction and it is not free, and section 6's K2 is the
        // difference between those two statements.
        ns_q    <= ns_q + NAK_NS + BACKOFF_NS;
        waste_q <= waste_q + NAK_NS;
      end
    end
  end

endmodule

What it models. The bus time a sequence of polls consumes, split by whether each one moved anything.

Engineering reason. Because the endpoint's throughput and the bus time it spends are different quantities with different owners, and only the second is a shared resource.

Inputs. A poll, the device's readiness, and the packet size.

State retained. Two 32-bit accumulators, a waste accumulator, and two counters.

Outputs. Whether this poll succeeded, and four running measurements.

Hardware implied. Three accumulators and a constant multiply.

Reset behaviour. All counters cleared.

Assumptions. That every poll is answered — a timeout is Chapter 14.2's territory and costs more than a NAK; that NAK_NS and OVERHEAD_NS match the link speed (Chapter 12.5 §1); and that BACKOFF_NS models the host's strategy, which the device does not control and cannot influence.

Omissions. Packets, transactions, the cause of unreadiness, the polling strategy and all errors — in the header.

What DV should verify. That a poll succeeds exactly when the device is ready; that every poll is charged, a NAKed one included; that the wasted time counts only NAKs; that the counts and the byte total agree; and that the two accumulators stay consistent with each other.

What each poll cost

10 cycles
A waveform of the retry cost model over ten polls. The first poll finds the device ready and delivers five hundred and twelve bytes, raising the bus time to ten thousand eight hundred and eighty-two nanoseconds. The second, third and fourth polls are refused, each adding nine hundred and twenty-seven nanoseconds to both the bus time and the wasted time while moving nothing. The fifth poll delivers another five hundred and twelve bytes. The sixth and seventh are refused, the eighth delivers, the ninth is refused and the tenth delivers. Across the ten polls four deliveries moved one thousand five hundred and thirty-six bytes at a total cost of thirty-eight thousand two hundred and eight nanoseconds, of which five thousand five hundred and sixty-two were spent on refusals.one delivery: +10 882 nsone delivery: +10 882 nsthree refusals: +2 781 ns totalthree refusals: +2 781 nstotal6 NAKs cost 5 562 of 38 208 ns — 15%6 NAKs cost 5 562 of 38 208ns — 15%poll12345678910answerDATANAKNAKNAKDATANAKNAKDATANAKDATAbytes051251251251210241024102415361536bus_ns0108821180912736136632454525472263993728138208wasted_ns009271854278127813708463546355562t0t1t2t3t4t5t6t7t8t9
Figure 2 — ten polls at a device ready four times, every column taken from a simulation of the block above. Watch the two accumulators: bus time jumps by 10 882 on a data poll and by 927 on a NAK, so six refusals add less than two thirds of one delivery. The wasted total ends at 5 562 ns against 38 208 consumed — 15%.

6. Mutation Test

Three mutations over 2 313 polls — a sweep of readiness from 1-in-10 to 10-in-10 at 100 polls each, directed corners at always-ready, 1-in-10 and 1-in-50, and 40 randomised runs.

The unmutated block reproduced §2's table exactly, and reported 0 errors of every kind.

got_data wrongbytes wrongbus time wrongwasted wrongcounts wrong
golden00000
K1 readiness ignored1 04649494949
K2 a NAK costs nothing0049490
K3 waste not attributed000490

K1 — every poll succeeds

Measured: 1 046 wrong outcomes, and every accumulator wrong with them.

The device delivers data it does not have. Catastrophic, immediate, and included as the contrast for K3.

K2 — charge nothing for a NAK

Measured: 49 runs with wrong bus time and wrong waste. The byte counts and the good/NAK counts are perfect.

§1's cheap misread as free. The endpoint's behaviour is entirely correct — it NAKs when it should, delivers when it should, and the packet stream is indistinguishable from the golden design's.

What is wrong is the number it reports about itself, and §3 is why that matters: a host or a driver using this figure to decide whether an endpoint is worth polling is reading a figure that omits the cost of polling it.

K3 — do not attribute the waste

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// waste_q <= waste_q + NAK_NS;     // MUTANT K3: removed

Measured: 49 runs with wrong wasted_ns. Bus time correct. Bytes correct. Counts correct. Behaviour identical.

The total is right and the breakdown is gone.

7. Verification

This chapter's commit point is every poll was charged what it actually cost, and the part that bought nothing was identified.

Stimulus. A readiness sweep from 1-in-10 to 10-in-10 at 100 polls each — which is what produces §2's table rather than a pass/fail; directed corners at always-ready, 1-in-10 and 1-in-50; and 40 randomised runs of varying length and readiness.

Observation. The success of each poll, both accumulators, and both counters — against expectations computed poll by poll in the bench, not from the totals.

Reference model. A per-poll accumulation in the bench: ready means overhead plus per-byte cost, unready means the NAK cost. Four lines, and it is independent because it is driven from the readiness pattern rather than from the design's outputs.

Coverage — crosses:

  • readiness from 1-in-10 to 10-in-10, every step
  • readiness 1-in-50 — the sparse extreme
  • run lengths from 1 to 60 polls
  • always-ready and never-ready endpoints
  • packet sizes at the endpoint maximum

Negative cases with defined outcomes: no poll succeeds while the device is unready; no poll is uncharged; the wasted total never exceeds the bus total; and the wasted total counts only NAKs.

8. Debugging: the Device That Slows Down Its Neighbours

A bulk device achieves acceptable throughput. When it is attached, an unrelated device on the same hub loses about a third of its throughput. Removing the first device restores the second. The first device reports no errors and its own throughput is within specification.

What does no errors and acceptable throughput rule out? A malfunction. The device is working, and the effect on its neighbour is a consequence of how it works rather than of anything going wrong.

What can one bulk endpoint take from another? Chapter 10.3: the residue. Bulk endpoints share what periodic traffic leaves, so one consuming more leaves less.

But the device's throughput is normal — so it is not consuming more bus time for data. What else consumes bus time? §1: NAKs.

Does the arithmetic work? §2's table: at ready-1-in-10, an endpoint spends 192 250 ns to deliver 5 120 bytes, of which 83 430 ns bought nothing — 43% of its bus consumption is waste. A neighbour sharing the residue loses that.

How do you confirm it from outside? Count NAKs in a bus trace. A device NAKing heavily is unmistakable in a capture and invisible in any throughput figure — because §2's whole finding is that heavy NAKing barely dents the endpoint's own rate.

What is the fix? §4: the device's only lever is being ready more often, which means buffering — Chapter 9.5. And the host's lever is backing off, which the device cannot request.

Is the device at fault? Not in any way that a compliance test would catch, and that is the honest answer. It is compliant, functional, and a bad neighbour. The specification does not bound how often a device may NAK, so this is a quality-of-implementation issue rather than a conformance one.

The signature to keep: a device that harms its neighbours while measuring fine itself is spending shared bus time on refusals — and the only instrument that shows it is a NAK count, which no throughput measurement contains.

9. Common Misconceptions

10. Reason It Through

A device's firmware author adds a 50 µs delay before answering when the endpoint has no data, reasoning that it reduces the number of NAKs the host receives and therefore the bus time consumed.

Does it reduce the NAK count? Yes — the host cannot poll again until the current transaction resolves, so a slower answer means fewer polls per second.

Does it reduce bus time? No. It increases it dramatically. Chapter 12.5 §2: the device must respond within 7.5 bit times — about 15 ns at high speed. A 50 µs delay is three thousand times that bound.

So what actually happens? The device misses the turnaround window entirely, the host times out, and — Chapter 14.2 §2's rung 3 — retries. The device does this again. A transaction that would have cost 927 ns now costs a full timeout plus a retry, repeatedly.

Is the reasoning salvageable at a smaller delay? No, and this is the important part. Any delay in answering is wrong, because the answer is due within the turnaround window and there is no legal value above it. The lever the author is reaching for does not exist at the transaction level.

Where does the lever actually exist? §4: in the host, which chooses the polling rate, and in the device's readiness, which is buffering. Neither is reachable from the transaction response path.

Is there anything the device can do that resembles the intent? Yes, and it is the useful reframing: be ready more often. Buffering does not reduce the NAK cost — it reduces the NAK count, by converting refusals into deliveries. That is the same goal, achieved at the only layer that can achieve it.

And the transferable point: a protocol's timing bounds are not a budget to spend but a window to hit, and an optimisation that treats a deadline as a resource converts a cheap correct behaviour into an expensive incorrect one. The author's instinct — fewer NAKs is better — was right; the mechanism was the one layer that cannot deliver it.

11. Understanding Check

12. Summary

A NAK costs a full transaction for zero bytes — about 927 ns at high speed — and that is about a twelfth of a 512-byte data transaction's 10 880 ns.

Which makes NAK-based flow control efficient rather than ruinous. Measured across a readiness sweep: a device ready 1 in 10 keeps 56% of ideal throughput, not 10%. Nine refusals together cost less than one delivery.

And the curve is steep at the wrong end — 1-in-10 to 2-in-10 buys eighteen percentage points; 9-in-10 to always-ready buys one. Effort spent eliminating the last few NAKs is nearly worthless.

The endpoint's throughput and the bus time it spends are different quantities with different owners. At ready-1-in-10 the endpoint lost 44% of its own rate and consumed 83 430 ns of shared residue that moved nothing — which is what a neighbour on the same bus pays.

And the device cannot influence any of it. There is no poll me later packet; a NAK carries no timing information; the polling rate is the host's alone. The device's only lever is being ready, and its worst alternative is stalling — Chapter 14.2 §4's over-escalation as a throughput decision.

§6's three mutations, and K3 is the one worth keeping: removing the waste attribution changes no behaviour at all — same packets, same timing, same throughput, same total.

bus_ns says how expensive this endpoint is. wasted_ns says how much of that expense bought nothing. The first is a measurement; the second is a diagnosis.

§7's practice generalises past this chapter:

A parameter sweep is a measurement. It becomes a test when every point in it has an expected value — and the display statement and the comparison are different lines, so deleting the comparison leaves output that looks identical.

13. What Comes Next

Three chapters of Bulk as a mechanism. The next two are Bulk as a foundation — what real classes build on top of it.

Chapter 14.4 is USB Mass Storage, which is where nearly all bulk traffic in the world actually goes. Its protocol is two endpoints and three phases — a 31-byte command wrapper, a data phase, and a 13-byte status wrapper — and it is a case study in building a request/response protocol on a channel that has neither.

Its interesting failure is the one it had to invent a name for: a phase error, which is what happens when the two ends disagree about which of the three phases they are in — Chapter 12.4 §2's two beliefs, reappearing two layers up with the same structure and a new recovery mechanism.

Browse the full path on the USB tutorials index.

Continue learning

Standards & specifications

Governing standard
USB-IF (Universal Serial Bus Specification)(opens USB Implementers Forum (USB-IF) in a new tab)

Defines the USB bus — its electrical signalling, connectors, packet and transaction model, device framework and the descriptors a device must expose — together with the device-class specifications layered on it. It does not define host-controller register interfaces (xHCI and EHCI are separate documents) nor any operating system's driver architecture.

This page also covers RTL structure, verification approach and debugging technique. Those are engineering practice built on the standard, not requirements the standard itself imposes.

Where this fits

Part of the USB curriculum.