USB · Module 14
Bulk for Data Streaming
Streaming avoids the framing mass storage had to build — and its throughput is decided by a number in the host's driver: queue depth, which takes a wire from 27% utilisation to 100%.
Chapter 14.4 built framing on a channel that has none. This chapter is the other end of the range — protocols that avoid needing it, and whose performance is decided somewhere else entirely.
1. Streaming Is the Absence of a Protocol
Printer class and most vendor-specific bulk usage have no wrappers, no commands and no status. A printer's bulk OUT endpoint takes a byte stream; a data-acquisition device's bulk IN endpoint produces one.
| Mass Storage | Streaming | |
|---|---|---|
| Framing | 31/13-byte wrappers | none |
| Request/response | invented — Chapter 14.4 §2 | not needed |
| Transfer boundary means | one command's data | nothing in particular |
| Recovery from desync | phase error + class reset | there is nothing to desynchronise |
That last row is the whole design. Chapter 14.4 §3's phase error exists because two ends track a phase neither transmits. A stream has no phase, so it has no phase error — and needs no mechanism to recover from one.
The simplest way to survive a class of failure is to have no state that can exhibit it.
Class code 7 is Printer and 0xFF is vendor-specific, against Mass Storage's 8 — and the difference in complexity between them is almost entirely this.
2. The Throughput Is Not Decided by the Device
Chapter 14.1 priced the wire and Chapter 14.3 priced the refusals. Neither explains why most real bulk streams run far below both figures.
The reason is that a transfer only occupies the bus while the host has one queued — and between a transfer completing and the next being submitted, the host's software has to run:
- the controller raises an interrupt;
- the driver's completion handler runs;
- it allocates and submits the next transfer;
- the controller schedules it.
That path is tens of microseconds. A 512-byte transaction is 10.9.
So a device with one transfer outstanding at a time spends most of its life waiting for software, and the wire is idle throughout — not busy with somebody else, idle.
3. Queue Depth
The mechanism, and it is entirely in the host's driver.
If the host keeps several transfers outstanding, the controller moves straight from one to the next without waiting for software. Software's latency is then overlapped with transmission instead of alternating with it.
Measured, with a 10 880 ns transaction and 40 000 ns of software latency:
| Depth | Drain time | Gap | Wire utilisation |
|---|---|---|---|
| 1 | 10 880 ns | 29 120 ns | 27% |
| 2 | 21 760 | 18 240 | 54% |
| 3 | 32 640 | 7 360 | 81% |
| 4 | 43 520 | 0 | 100% |
| 8 | 87 040 | 0 | 100% |
And the threshold is exact:
depth ≥ ⌈ software latency ÷ transaction time ⌉ = ⌈ 40 000 ÷ 10 880 ⌉ = 4
4. Who Owns the Number
Not the device, which is the uncomfortable part.
Queue depth is a property of the host's driver: how many transfers it submits before waiting. A device cannot request a depth, cannot observe one, and cannot compensate for a shallow one.
What a device can do is make each transfer bigger. The threshold is a ratio, so doubling the transfer size halves the depth needed — and unlike depth, transfer size is something a device's protocol can influence by not imposing small natural boundaries on its data.
| Lever | Owned by | Effect |
|---|---|---|
| Queue depth | the host driver | the direct mechanism |
| Transfer size | the protocol above | lowers the depth threshold |
| Packet size | the endpoint descriptor | Chapter 14.1 §2's efficiency |
| Readiness | the device | Chapter 14.3's NAK rate |
Three of the four are the device's and none of them is the mechanism — which is why a device vendor investigating poor streaming throughput often finds nothing wrong with the device.
5. The Queue-Depth Model, as RTL
// ─────────────────────────────────────────────────────────────────────────
// usb_bulk_queue
//
// Classification: SIMPLIFIED SYNTHESIZABLE TEACHING RTL -- in practice an
// accounting model, because the quantity it computes lives in a host driver
// rather than in a device.
//
// WHAT IT MODELS. Section 3's arithmetic: how long the wire is busy and how
// long it is idle, per round of DEPTH transfers, given a software latency
// the queue is trying to hide.
//
// WHAT IT DOES NOT MODEL. The transfers themselves (Chapter 14.1); NAKs
// (Chapter 14.3 -- every transaction here succeeds); errors (Chapter 14.2);
// other traffic; and the host's actual scheduling, which is more complex
// than "drain then refill" and does not change the threshold.
//
// ── WHY A ROUND AND NOT A TRANSACTION ───────────────────────────────────
// Driving this a ROUND at a time -- fill, drain, wait -- makes the result
// analytic: busy is DEPTH*TXN_NS and idle is whatever of SW_LAT_NS the
// draining did not cover. A transaction-level model of the same thing is
// larger, needs a scheduler, and produces the same number. Section 7 is
// about preferring the model whose expected value can be written down.
// ─────────────────────────────────────────────────────────────────────────
module usb_bulk_queue #(
parameter int unsigned DEPTH = 4,
parameter int unsigned TXN_NS = 10880, // Chapter 14.1: one 512 B txn
// Section 2: interrupt, completion handler, submission, scheduling. Tens
// of microseconds, against a transaction's ten.
parameter int unsigned SW_LAT_NS = 40000
)(
input logic clk,
input logic rst_n,
input logic start,
input logic round,
output logic [31:0] busy_ns,
output logic [31:0] idle_ns,
output logic [15:0] txns_done
);
logic [31:0] busy_q, idle_q;
logic [15:0] done_q;
assign busy_ns = busy_q;
assign idle_ns = idle_q;
assign txns_done = done_q;
// SECTION 3'S ARITHMETIC, as two constants.
localparam int unsigned DRAIN_NS = DEPTH * TXN_NS;
// The gap is what the draining did not cover -- and it SATURATES at zero,
// which is why depth is a step function rather than a curve. Past the
// threshold there is no gap left to remove.
localparam int unsigned GAP_NS = (DRAIN_NS >= SW_LAT_NS)
? 0 : (SW_LAT_NS - DRAIN_NS);
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
busy_q <= '0; idle_q <= '0; done_q <= '0;
end else if (start) begin
busy_q <= '0; idle_q <= '0; done_q <= '0;
end else if (round) begin
done_q <= done_q + DEPTH[15:0];
busy_q <= busy_q + DRAIN_NS;
// Section 2: the wire is IDLE here, not contended. Nothing else is
// using it -- which is why this cost is invisible in every
// device-side measurement.
idle_q <= idle_q + GAP_NS;
end
end
endmoduleWhat it models. The wire time a queue of a given depth wastes waiting for software.
Engineering reason. Because the dominant limit on real bulk streaming is neither the wire nor the device, and the threshold that removes it is computable rather than empirical.
Inputs. A round — fill, drain, wait.
State retained. Two 32-bit accumulators and a counter.
Outputs. Busy time, idle time, and transactions completed.
Hardware implied. Two adders with constant operands. GAP_NS folds away entirely at elaboration.
Reset behaviour. Counters cleared.
Assumptions. That every transaction succeeds — Chapter 14.3 adds NAKs on top of this, and they lengthen the drain, which helps; that SW_LAT_NS is constant, which it is not; and that the host refills only after a completion, which is the worst case and the one worth designing against.
Omissions. Transfers, NAKs, errors, other traffic and real scheduling — in the header.
What DV should verify. That busy time is DEPTH × TXN_NS per round; that the gap is max(0, SW_LAT_NS − DRAIN_NS); that the gap saturates at zero rather than going negative; that the transaction count is DEPTH per round; and that the utilisation curve has the step shape §3 predicts across a sweep of depths.
6. Mutation Test
Four mutations, checked across five depths — 1, 2, 3, 4 and 8 — over 100 rounds each, against expectations computed from §3's formula.
The unmutated block reproduced §3's table exactly, with 0 errors of every kind.
| busy wrong | idle wrong | count wrong | |
|---|---|---|---|
| golden | 0 | 0 | 0 |
| Q1 depth does not reduce the gap | 0 | 5 | 0 |
| Q2 no gap at any depth | 0 | 3 | 0 |
| Q3 one transaction per round | 0 | 0 | 4 |
| Q4 idle time not charged | 0 | 3 | 0 |
Q1 — the gap is always the full software latency
Measured: all 5 depths wrong.
Queue depth buys nothing — the model reports 27% utilisation at every depth including 8. A driver author reading this figure would conclude that depth does not help and stop at 1, which is precisely the belief §3 exists to correct.
Q2 — there is never a gap
Measured: 3 depths wrong — 1, 2 and 3.
Depths 4 and 8 are correct, because their real gap is already zero. So the mutation is invisible at exactly the depths a well-tuned driver uses, and visible only below the threshold.
Which is the wrong way round for a measurement tool: it reports 100% utilisation for a depth-1 queue that is actually achieving 27%, so the number is right when you do not need it and wrong when you do.
Q3 — count one transaction per round
Measured: 4 depths wrong — every depth except 1.
Busy time and idle time stay correct, so the utilisation figure is right and the throughput figure is wrong by a factor of DEPTH. A driver author would see a wire at 100% carrying a quarter of the expected data and look for the loss on the wire, where it is not.
Q4 — do not charge idle time
Measured: 3 depths wrong — 1, 2 and 3 again.
The busy time is right and the total is missing. Utilisation computed as busy / (busy + idle) becomes busy / busy = 100% at every depth.
7. Verification
This chapter's commit point is the utilisation curve had the shape the arithmetic predicts, at every depth.
Stimulus. A sweep across five depths — 1, 2, 3, 4 and 8 — at 100 rounds each, chosen to straddle the threshold: three below, one at, one well above.
Observation. Busy time, idle time and the transaction count at every depth, against §3's closed-form expectation — not against a second simulation.
Reference model. A three-line function. Its independence is total, because it is the algebra rather than a re-implementation — which is the property §7's callout is about.
Coverage — crosses:
- depth below, at and above the threshold
- the threshold boundary itself — depth 3 versus 4
- a depth where the gap is exactly zero
- transaction times spanning full speed and high speed — which moves the threshold
Negative cases with defined outcomes: the gap never goes negative; utilisation never exceeds 100%; and no depth above the threshold differs from the threshold.
8. Debugging: the Capture Device That Is Slow on One Machine
A data-acquisition device streams on a bulk IN endpoint. On a developer's workstation it sustains 40 MB/s. On the customer's embedded system it manages 11 MB/s with the same firmware, the same cable and the same host controller silicon. The device reports no NAKs and no errors on either.
What does no NAKs and no errors rule out? The device. It was ready every time it was asked, so Chapter 14.3's mechanism is not involved, and Chapter 14.2's ladder is not either.
So what differs between the machines? Not the wire and not the device — the software above the driver, and specifically how many transfers it keeps outstanding.
Does the arithmetic fit? §3's table: 11 MB/s against 40 is 27%, which is the depth-1 figure exactly. On the embedded system the driver is submitting one transfer at a time.
How do you confirm it? Look at the wire, not the device. A capture will show transactions in bursts separated by idle gaps — not NAKs, not other traffic, nothing at all. Measure a gap: it is the software latency, and it is the number that sets the threshold.
What is the fix and who applies it? §4: the host's driver submits more transfers before waiting. The device vendor cannot do it, and the depth required is ⌈gap ÷ transaction time⌉, which the capture just measured.
Is there anything the device can do? Make transfers larger, which lowers the threshold proportionally. It does not fix the driver; it makes the driver's shallowness matter less.
And why did the workstation hide it? Because its software latency is smaller. The same depth-1 driver on a faster machine may sit above the threshold by accident — which is why this is found at the customer and not in development.
The signature to keep: throughput that differs by a large factor between hosts, with no NAKs and no errors and visible idle gaps on the wire, is queue depth — and the gap on the trace is the whole diagnosis.
9. Common Misconceptions
10. Reason It Through
A vendor protocol sends 200-byte records over a bulk pipe, one record per transfer, and the receiving application reads one transfer at a time. It works. The vendor then increases the record size to 512 bytes to improve efficiency, and the application starts seeing records split across reads.
Why would larger records split? Because 512 is the endpoint's wMaxPacketSize — Chapter 14.1 §2. A 200-byte transfer is one short packet and self-delimiting; a 512-byte transfer is one full packet, and a full packet does not end a transfer.
What was the application actually relying on? Chapter 12.2 §2's short-packet rule, without knowing it. Every 200-byte record ended with a short packet, so every read returned exactly one record — by accident of the size.
What happens at 512? The transfer has no short packet, so it does not terminate — and the next record's data continues the same read. The application receives 1024 bytes and calls it a split record.
Is the device wrong? No, and the vendor's efficiency reasoning was sound: Chapter 14.1 §2 puts 512-byte packets at 78% efficiency against a 200-byte transfer's much lower figure.
Is the application wrong? Yes, and it was always wrong — it inferred message boundaries from transfer boundaries, which §1 says USB does not guarantee. The 200-byte case worked by coincidence.
What are the two fixes, and which is right? Send a zero-length terminator after each 512-byte record — Chapter 14.1 §4 prices it at a full transaction, which at one per record is 50% overhead. Or frame the records in the byte stream, which costs a few bytes per record and makes the protocol independent of packet size forever.
The second, decisively — and the first is what gets implemented, because it is a one-line change on the device and the other is a protocol revision.
And the transferable point: a protocol that works because of a size coincidence has an undocumented dependency on wMaxPacketSize, which is a configuration-time value that changes with link speed. It will break on a port, a speed change, or a record-size tweak — and the failure will look like whatever changed rather than like the assumption that was always there.
11. Understanding Check
12. Summary
Streaming avoids the framing Chapter 14.4 had to build — no wrappers, no commands, no status — and therefore has no phase and no phase error. The simplest way to survive a class of failure is to have no state that can exhibit it.
And its throughput is decided by neither the device nor the wire. Between a transfer completing and the next being submitted, the host's software runs — tens of microseconds against a transaction's ten — and the wire is idle, not contended, throughout.
Queue depth is the mechanism, measured:
| Depth | 1 | 2 | 3 | 4 | 8 |
|---|---|---|---|---|---|
| Wire utilisation | 27% | 54% | 81% | 100% | 100% |
with the threshold exactly ⌈ software latency ÷ transaction time ⌉ = 4. A depth-1 queue costs 73% of the wire; depth 5 buys nothing. Queue depth is a step function with one correct value, and that value is a division rather than an experiment.
It also moves with the link — a ratio whose denominator changes with speed — so a driver tuned at full speed is 3.7× too shallow at high speed without a line changing.
And the device cannot fix it. Depth belongs to the host's driver; the device's only lever is larger transfers, which lowers the threshold rather than meeting it.
§6's four mutations produced Chapter 14.4's pair again, and completing a pattern that runs through the whole module:
Every measurement has two independent failure modes — measuring the wrong thing, and recording the right thing wrongly — and they are indistinguishable in the output.
§7 is the chapter's own correction. The first model was transaction-level, produced an obviously wrong result — depth 2 worse than depth 1 — and disagreed with its own bench on 602 of 834 steps, because both had to replicate the same interacting subtleties and did not.
A model whose expected value you cannot write down is a model you cannot check — only compare against another model with the same subtleties.
Rewriting it a round at a time made the expectation one line of algebra, made the threshold fall out in closed form, and lost nothing: the transaction-level version computed the same number with more state and more ways to be wrong.
13. What Comes Next
Module 14 is complete: Bulk's throughput, its reliability, its retries, and the two ends of the range it carries — a class that built a full request/response protocol on it, and one that deliberately built nothing.
What every chapter so far has assumed is that the bits arrive. That a receiver can find a packet's boundaries on a wire carrying no clock, tell a one from a zero, and work out what speed to talk at before it can talk at all.
The remaining modules are where that assumption is discharged — beginning with how a bit is put on a wire such that the other end can find it again.
Browse the full path on the USB tutorials index.
Continue learning
Related tutorials
- Related topic
Four Transfer Types Overview
The four transfer types are not a list to memorise — they fall out of two orthogonal questions plus one bootstrap question. Deriving them, and the policy block whose outputs are all derived and none stored.
- Related topic
Bulk Transfers
Promised nothing and permitted everything: why throughput and guarantee are different axes, and why a fixed-priority arbiter starves an endpoint forever while every safety property passes.
- Related topic
Bulk Throughput
Why 480 Mbit/s never means 60 MB/s — overhead is per transaction and fixed, so efficiency is a function of packet size alone, and a zero-length terminator costs a full transaction for no bytes.
- Related topic
Bulk Reliability
Guarantee-correct and bandwidth-best-effort are two promises, not one — and the five-rung recovery ladder that makes the first true has rungs that are not interchangeable.
Standards & specifications
- Governing standard
- USB-IF (Universal Serial Bus Specification)(opens USB Implementers Forum (USB-IF) in a new tab)
Defines the USB bus — its electrical signalling, connectors, packet and transaction model, device framework and the descriptors a device must expose — together with the device-class specifications layered on it. It does not define host-controller register interfaces (xHCI and EHCI are separate documents) nor any operating system's driver architecture.
This page also covers RTL structure, verification approach and debugging technique. Those are engineering practice built on the standard, not requirements the standard itself imposes.
Where this fits
Part of the USB curriculum.
