USB · Module 14
Bulk Throughput
Why 480 Mbit/s never means 60 MB/s — overhead is per transaction and fixed, so efficiency is a function of packet size alone, and a zero-length terminator costs a full transaction for no bytes.
Chapter 10.3 established what Bulk is promised: nothing, and permitted everything. This module is about what it actually delivers — and this chapter is the arithmetic.
The headline number is a lie in a specific and predictable way. A high-speed bus signals at 480 Mbit/s, which is 60 MB/s, and no bulk endpoint has ever moved 60 MB/s of payload. The gap is not loss or contention. It is structure, and it is calculable.
1. Overhead Is Per Transaction, and Fixed
Module 12 built a transaction as three packets: token, data, handshake. Two of those three carry no payload at all, and the third carries a PID and a CRC alongside the payload.
The cost of a transaction is a fixed overhead plus a per-byte cost. The overhead does not shrink when the payload does.
Which has one consequence that determines everything else:
Efficiency is a function of packet size and nothing else.
Not of the transfer's length, not of the bus's load, not of the device's speed — of how many payload bytes ride on each fixed overhead.
The kernel supplies the numbers. HS_NSECS in linux/usb/hcd.h gives the bus time a handshaked high-speed transaction costs, and its structure is exactly §1's claim — a constant term plus a term proportional to the byte count:
#define HS_NSECS(bytes) (((55 * 8 * 2083) \
+ (2083UL * (3 + BitTime(bytes))))/1000 \
+ USB2_HOST_DELAY)2083 is 2.083 ns in picoseconds — one bit at 480 Mbit/s (Chapter 12.5 §1), and BitTime(bytes) is 7 * 8 * bytes / 6, the kernel's allowance for bit stuffing.
2. The Numbers
Working HS_NSECS through at each legal high-speed bulk packet size:
| Payload | Transaction | Payload bits alone | Efficiency | Per 125 µs microframe |
|---|---|---|---|---|
| 64 B | 2 171 ns | 1 067 ns | 49.1% | 57 |
| 128 B | 3 414 ns | 2 133 ns | 62.5% | 36 |
| 256 B | 5 904 ns | 4 267 ns | 72.3% | 21 |
| 512 B | 10 880 ns | 8 533 ns | 78.4% | 11 |
At the maximum bulk packet size, 78% of the bus time carries payload. At a quarter of it, half.
And the same shape at full speed, computed from the packet structure directly — a token of 3 bytes, a data packet of PID + payload + CRC16, a handshake of 1 byte, each with its own SYNC and EOP:
| Payload | Bytes on the wire | Efficiency | Per 1 ms frame |
|---|---|---|---|
| 8 B | ~18 B | 44.4% | 83 |
| 16 B | ~26 B | 61.5% | 57 |
| 32 B | ~42 B | 76.2% | 35 |
| 64 B | ~74 B | 86.5% | 20 |
3. Where the Rest of It Goes
The 22% that is not payload at 512 bytes, accounted for:
| Roughly | |
|---|---|
| Token packet — SYNC, PID, address, endpoint, CRC5, EOP | ~5 bytes on the wire |
| Data packet's own PID, CRC16, SYNC, EOP | ~8 bytes |
| Handshake packet | ~3 bytes |
| Inter-packet gaps and turnaround — Chapter 12.5 | two per transaction |
| Bit stuffing | up to one extra bit per six |
Bit stuffing is the one that surprises people. The encoding inserts a zero after six consecutive ones, so a payload of all-ones costs more bus time than a payload of alternating bits — and the kernel's BitTime allows for it with a flat 7/6 factor rather than computing it, because the transmitter cannot know the data in advance when it is reserving bandwidth.
The bus time a transfer consumes depends on its contents, which is a fact with no analogue in most buses and which makes exact throughput prediction impossible in principle.
4. A Zero-Length Packet Costs a Full Transaction
Chapter 13.2 §3 established that a transfer whose length is an exact multiple of wMaxPacketSize needs an explicit zero-length terminator. This chapter supplies its price.
A ZLP has no payload and pays the full fixed overhead — token, data packet, handshake, two turnarounds. At high speed that is:
~2 347 ns of bus time for zero bytes, which is 21.6% of a full 512-byte transaction.
And the transfers that need one are exactly the ones real software produces. A filesystem reads in 4 096-byte blocks; 4096 ÷ 512 = 8 exactly, so every block read ends with a ninth transaction carrying nothing.
| Transfer | Payload transactions | ZLP? | Total | Overhead |
|---|---|---|---|---|
| 4 096 B | 8 | yes | 9 | 1 transaction wasted in 9 — 11.1% |
| 4 000 B | 8 (last is short) | no | 8 | none |
| 65 536 B | 128 | yes | 129 | 0.8% |
The cost falls as the transfer grows, which is the reassuring half. The alarming half is that a 512-byte transfer with a terminator is 50% overhead — one transaction of payload and one of nothing.
5. The Packetiser, as RTL
// ─────────────────────────────────────────────────────────────────────────
// usb_bulk_packetiser
//
// Classification: SIMPLIFIED SYNTHESIZABLE TEACHING RTL. It answers the
// question throughput actually depends on: how many TRANSACTIONS does this
// transfer cost, and how much bus time is that?
//
// WHAT IT MODELS. Sections 1 to 4: splitting a transfer into packets, the
// terminator rule, and the bus time each transaction consumes.
//
// WHAT IT DOES NOT MODEL. The packets themselves (Module 11); the
// transactions (Module 12); NAKs and retries (Chapter 14.3 -- every
// transaction here succeeds first time); other traffic on the bus (Chapter
// 10.3's arbitration); and bit stuffing's DATA DEPENDENCE (section 3), which
// is folded into a constant here exactly as the kernel folds it into
// BitTime.
//
// ── ON THE TIMING PARAMETERS ────────────────────────────────────────────
// OVERHEAD_NS and NS_PER_B_X100 are the two terms of section 1's model, and
// their defaults are derived from the kernel's HS_NSECS at 480 Mbit/s. They
// are PARAMETERS because they change with speed -- Chapter 12.5 section 1's
// rule that a quantity varying with the link is not a property of the design.
// ─────────────────────────────────────────────────────────────────────────
module usb_bulk_packetiser #(
parameter int unsigned MAXP = 512,
parameter int unsigned OVERHEAD_NS = 2347, // fixed, per transaction
parameter int unsigned NS_PER_B_X100 = 1667 // 16.67 ns/byte at 480 Mbit/s
)(
input logic clk,
input logic rst_n,
input logic start,
input logic [31:0] xfer_len,
input logic want_zlp, // Chapter 13.2's terminator request
output logic [15:0] pkt_len,
output logic pkt_valid,
output logic is_zlp,
output logic xfer_done,
output logic [15:0] txn_count,
output logic [31:0] bytes_moved,
output logic [31:0] bus_ns
);
logic [31:0] rem_q, bytes_q, ns_q;
logic [15:0] txn_q;
logic active_q, zlp_owed_q;
assign txn_count = txn_q;
assign bytes_moved = bytes_q;
assign bus_ns = ns_q;
logic [15:0] this_len;
assign this_len = (rem_q >= MAXP) ? MAXP[15:0] : rem_q[15:0];
// A zero-length transfer IS a single zero-length transaction: that is what
// "transfer nothing" looks like on a wire with no other way to say it.
assign is_zlp = active_q && (rem_q == 32'd0);
assign pkt_valid = active_q;
assign pkt_len = is_zlp ? 16'd0 : this_len;
// Done on THIS transaction, not after it. Written against the packet being
// emitted now -- an output describing a state the block has already left
// can never assert, and section 8 is about how that was found.
assign xfer_done = active_q &&
( is_zlp || ((rem_q == {16'd0, this_len}) && !zlp_owed_q) );
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
rem_q <= '0; bytes_q <= '0; ns_q <= '0; txn_q <= '0;
active_q <= 1'b0; zlp_owed_q <= 1'b0;
end else if (start) begin
rem_q <= xfer_len;
bytes_q <= '0; ns_q <= '0; txn_q <= '0;
active_q <= 1'b1;
// SECTION 4. A transfer that is an exact multiple of MAXP contains no
// short packet, so it owes a terminator -- and that terminator is a
// whole transaction at full price.
zlp_owed_q <= want_zlp && ((xfer_len % MAXP) == 0);
end else if (active_q) begin
txn_q <= txn_q + 16'd1;
// EVERY transaction pays the fixed overhead, a zero-length one too.
// Section 4 is the whole reason this line has no exception in it.
ns_q <= ns_q + OVERHEAD_NS + ((pkt_len * NS_PER_B_X100) / 100);
bytes_q <= bytes_q + pkt_len;
if (is_zlp) begin
zlp_owed_q <= 1'b0;
active_q <= 1'b0;
end else begin
rem_q <= rem_q - pkt_len;
if ((rem_q - pkt_len) == 32'd0) begin
if (!zlp_owed_q) active_q <= 1'b0;
end
end
end
end
endmoduleWhat it models. The transaction count and bus time a bulk transfer costs.
Engineering reason. Because throughput is a question about transactions, not about bytes, and the transaction count is not the byte count divided by the packet size.
Inputs. A transfer length and whether its protocol wants a terminator.
State retained. A 32-bit remainder, byte and time accumulators, a 16-bit counter, and two flags.
Outputs. Each packet's length, the terminator indication, completion, and the three accumulated figures.
Hardware implied. A subtractor, a comparator, and three accumulators. The multiply is by a constant and folds away.
Reset behaviour. Everything cleared; a transfer does not survive a reset.
Assumptions. That every transaction succeeds first time — Chapter 14.3 removes this; that OVERHEAD_NS matches the link speed; and that bit stuffing is adequately modelled by a constant per byte, which §3 explains is a approximation the kernel also makes.
Omissions. Packets, transactions, retries, other traffic and data-dependent stuffing — in the header.
What DV should verify. That the transaction count matches the packet split including the terminator; that no packet exceeds MAXP; that a zero-length transfer is one transaction; that the bus-time accounting charges every transaction including a ZLP; that xfer_done asserts on the last transaction; and that a transfer always terminates.
Transactions, not bytes
9 cycles6. Mutation Test
Five mutations over 212 transfers and 4 159 packets, with directed cases for every relationship to MAXP — exact multiples, one-short, one-over, zero-length, and a 4 096-byte filesystem block — plus 200 randomised transfers.
The unmutated block: 6 transfers ending with a terminator, and 0 errors of every kind. It reports a 64 KiB transfer as 128 transactions, 1 392 896 ns, 47.05 MB/s — 78% of the nominal 60 MB/s, matching §2's derivation exactly.
| bytes wrong | transaction count wrong | ZLP wrong | bus time wrong | xfer_done missing | never ended | |
|---|---|---|---|---|---|---|
| golden | 0 | 0 | 0 | 0 | 0 | 0 |
| T1 last packet always full | 204 | 204 | 0 | 0 | 204 | 204 |
| T2 no terminator | 0 | 4 | 4 | 0 | 0 | 0 |
| T3 short transfers never report done | 0 | 0 | 0 | 0 | 206 | 0 |
| T4 a ZLP costs nothing | 0 | 0 | 0 | 6 | 0 | 0 |
| T5 terminator skipped at the end | 0 | 4 | 4 | 0 | 4 | 0 |
T1 — make every packet full-size
Measured: 204 transfers with the wrong byte count, wrong transaction count, and no termination.
The remainder never reaches zero because the last packet claims more bytes than remain. Catastrophic and immediate — included as the contrast for T4.
T2 — never owe a terminator
Measured: 4 of 212 transfers wrong.
§4's rule removed. Four is 1.9% of this bench and close to universal in production, for the reason Chapter 12.2 §6 measured: random lengths are rarely exact multiples of 512, and real software's lengths are almost always powers of two.
T3 — report completion only on a terminator
Measured: 206 transfers never reported done — every transfer that ended on a short packet rather than a ZLP.
And it is the mutation that revealed a bug in the golden design. §8.
T4 — charge nothing for a zero-length transaction
if (!is_zlp) ns_q <= ns_q + OVERHEAD_NS + ...; // MUTANT T4Measured: 6 transfers with mis-accounted bus time — exactly the six that ended with a terminator. Every other output identical.
The packet stream is correct; only the cost is wrong. A design with this defect packetises perfectly and under-reports its own bus consumption by a full transaction on every exact-multiple transfer.
Which matters because the number is used for scheduling. A host that believes a transfer is cheaper than it is over-commits the bus.
T5 — skip the terminator at the end of the transfer
Measured: 4 wrong transaction counts, 4 missing terminators, 4 missing completions.
The terminator is owed and the transfer ends without it. The same 4 transfers as T2 — and the same production-versus-random asymmetry.
7. Verification
This chapter's commit point is the transfer cost what the design says it cost.
Stimulus. Exact multiples of MAXP (512, 1024, 4096, 65536); one-under and one-over (511, 513); a zero-length transfer, with and without a terminator request; a 4 096-byte block with and without one; and 200 randomised lengths up to 20 000 bytes.
Observation. Every packet's length, the running counts, the accumulated bus time against an independently computed expectation, and xfer_done on every transaction.
Reference model. Three lines: ceil(len / MAXP) transactions, plus one if a terminator is owed, and overhead + len × rate of bus time per transaction. Its independence is real — it computes the totals from the transfer's parameters without reference to the packet stream, so it disagrees whenever the stream is wrong.
Coverage — crosses:
- transfer length modulo
MAXP: 0, 1,MAXP−1, mid-range want_zlp× exact multiple — the cell that produces a terminator- zero-length transfer × each
want_zlp - transfer sizes spanning 1 to 65 536 bytes
- 4 096 bytes specifically — the filesystem block
Negative cases with defined outcomes: no packet exceeds MAXP; no transfer fails to terminate; no transaction is un-charged; and xfer_done never asserts on a non-final transaction.
8. Debugging: the Device That Is Fast Except on Round Numbers
A bulk device sustains close to its expected throughput on arbitrary transfer sizes. When the application switches to reading in 4 096-byte blocks, throughput drops by about 11% and stays there. Larger blocks recover most of it; 512-byte blocks halve it.
What does drops on round numbers tell you? That the defect is triggered by a property of the length, not of the data or the rate. Round numbers on a USB bulk endpoint means exact multiples of 512.
What happens on an exact multiple? §4: no short packet, so a terminator is owed — and the terminator is a whole transaction.
Does the arithmetic match? 4 096 ÷ 512 = 8 payload transactions plus one terminator = 9. One wasted transaction in nine is 11.1%, which is the reported drop.
And the 512-byte case? One payload transaction plus one terminator = 50% overhead, which halves throughput. Both numbers follow from the same rule.
So is this a bug? No — and that is the useful part. The device and host are both correct; the terminator is required, and the cost is inherent. The fault, if there is one, is in the application's block size.
What can actually be done? Three things, in decreasing order of usefulness:
- Read in sizes that are not exact multiples — 4 095 or 4 097 bytes eliminates the terminator entirely;
- Read in larger blocks — 64 KiB amortises one wasted transaction over 128, which is 0.8%;
- Drop the terminator request if the protocol above does not need it — Chapter 13.2 §3: whether a transfer needs terminating is a property of the protocol running on top, not of USB.
What if none of those is available? Then the number is the number, and the right outcome of the investigation is a documented explanation rather than a fix.
The signature to keep: a throughput loss that appears at exact multiples of the packet size and scales as 1/(n+1) is a zero-length terminator — and it is arithmetic, not a defect.
9. Common Misconceptions
10. Reason It Through
Two devices are compared. Device A uses 512-byte bulk packets and sustains 40 MB/s. Device B declares 64-byte bulk packets and sustains 22 MB/s. The vendor of B argues that its smaller packets reduce latency and the throughput difference is a fair trade.
Is the throughput difference explained by the packet size? §2: 78.4% efficiency at 512 bytes against 49.1% at 64. The ratio is 0.63, and 22 ÷ 40 is 0.55 — close, and the remainder is accounted for by B needing 8× as many transactions, each with its own Chapter 12.5 turnaround.
So the arithmetic supports the vendor's premise. The loss is structural, not a defect.
Is the latency claim also true? Partly, and it is the weaker half. A smaller packet does complete sooner — but a bulk endpoint has no latency guarantee at all (Chapter 10.3 §1), so the latency being improved is not one anybody was promised.
What is the real effect of smaller packets on latency? It reduces the granularity of bus occupancy, which helps other endpoints — a 64-byte transaction occupies the bus for 2.2 µs where a 512-byte one occupies it for 10.9 µs. So B is not improving its own latency; it is being a better citizen to everything else on the bus.
Is that a good trade? It depends entirely on what else is attached, and that is the point the vendor's argument skips. On a bus with a latency-sensitive isochronous endpoint, B's smaller transactions genuinely help. On a bus where B is alone, it has given away 45% of its throughput for nothing.
And what should the evaluation actually measure? Not either device in isolation. Both devices alongside the traffic they will share a bus with — because the quantity the vendor is trading is not its own latency but somebody else's, and that quantity is invisible in a single-device test.
The transferable point: a trade-off argument names what is given up and often not who receives it. Here the beneficiary is a third party who may not exist, and the cost is borne by the device making the argument.
11. Understanding Check
12. Summary
Overhead is a fixed cost per transaction, so efficiency is a function of packet size alone — and the kernel's HS_NSECS has exactly that shape: a constant term plus one proportional to the bytes.
The numbers, derived rather than quoted: at high speed, 49.1% efficiency at 64 bytes rising to 78.4% at 512; at full speed, 44.4% at 8 bytes rising to 86.5% at 64. The faster bus is the less efficient one, because its fixed overhead is larger in bit times — raising the bit rate raises the rate at which overhead is paid.
At the maximum bulk packet size a high-speed endpoint gets 11 transactions per microframe — 45.1 MB/s against a nominal 60, which is 75%, before contention, retries or any other device.
And a zero-length terminator costs a full transaction — about 2 347 ns for no bytes, 21.6% of a 512-byte transaction. The transfers that need one are the ones real software produces: a 4 096-byte filesystem block is eight packets plus a ninth carrying nothing, which is the 11.1% of §8's debugging scenario, arithmetic rather than a defect.
§6 measured five mutations over 212 transfers, and the pair worth keeping is T2 and T4 — both about the terminator, and only one visible on the wire. T2 changes what is sent; T4 changes only what the design says it cost, and produces no protocol error at all.
A design that lies about its own cost is harder to catch than one that misbehaves, because nothing downstream contradicts it.
§7 is the chapter's finding, and it is about benches:
A quantity a bench reports is not a quantity a bench verifies — and a printed number looks identical whether the design is right or wrong.
Adding the missing check found more than a mutation. The completion output could never assert in the correct design: it was written against the state after the final transaction, by which time the block had already gone inactive. An output no test reads is an output no test constrains, and it is as likely to be dead as to be wrong.
13. What Comes Next
This chapter assumed every transaction succeeds first time. None of them do, always.
Chapter 14.2 is what Bulk actually guarantees — and the answer separates two things that sound like one: Bulk is guarantee-correct and bandwidth-best-effort. Data is delivered intact or the error is visible; when it is delivered is promised by nobody.
The chapter is about the ladder of mechanisms that make the first half true — CRC, the data toggle, retry, STALL, reset — and about what each one can and cannot recover from, because they are not interchangeable and a device that reaches for the wrong rung recovers from nothing.
Browse the full path on the USB tutorials index.
Continue learning
Related tutorials
- Related topic
Four Transfer Types Overview
The four transfer types are not a list to memorise — they fall out of two orthogonal questions plus one bootstrap question. Deriving them, and the policy block whose outputs are all derived and none stored.
- Related topic
Bulk Transfers
Promised nothing and permitted everything: why throughput and guarantee are different axes, and why a fixed-priority arbiter starves an endpoint forever while every safety property passes.
- Related topic
Bulk Reliability
Guarantee-correct and bandwidth-best-effort are two promises, not one — and the five-rung recovery ladder that makes the first true has rungs that are not interchangeable.
- Related topic
Bulk Retries
A NAK costs a full transaction for zero bytes — and is still eleven times cheaper than a data transaction, so a device ready one time in ten keeps 56% of its throughput rather than 10%.
Standards & specifications
- Governing standard
- USB-IF (Universal Serial Bus Specification)(opens USB Implementers Forum (USB-IF) in a new tab)
Defines the USB bus — its electrical signalling, connectors, packet and transaction model, device framework and the descriptors a device must expose — together with the device-class specifications layered on it. It does not define host-controller register interfaces (xHCI and EHCI are separate documents) nor any operating system's driver architecture.
This page also covers RTL structure, verification approach and debugging technique. Those are engineering practice built on the standard, not requirements the standard itself imposes.
Where this fits
Part of the USB curriculum.
