Skip to content
VLSI Mentor

USB · Module 14

Bulk Throughput

Why 480 Mbit/s never means 60 MB/s — overhead is per transaction and fixed, so efficiency is a function of packet size alone, and a zero-length terminator costs a full transaction for no bytes.

Chapter 10.3 established what Bulk is promised: nothing, and permitted everything. This module is about what it actually delivers — and this chapter is the arithmetic.

The headline number is a lie in a specific and predictable way. A high-speed bus signals at 480 Mbit/s, which is 60 MB/s, and no bulk endpoint has ever moved 60 MB/s of payload. The gap is not loss or contention. It is structure, and it is calculable.

1. Overhead Is Per Transaction, and Fixed

Module 12 built a transaction as three packets: token, data, handshake. Two of those three carry no payload at all, and the third carries a PID and a CRC alongside the payload.

The cost of a transaction is a fixed overhead plus a per-byte cost. The overhead does not shrink when the payload does.

Which has one consequence that determines everything else:

Efficiency is a function of packet size and nothing else.

Not of the transfer's length, not of the bus's load, not of the device's speed — of how many payload bytes ride on each fixed overhead.

The kernel supplies the numbers. HS_NSECS in linux/usb/hcd.h gives the bus time a handshaked high-speed transaction costs, and its structure is exactly §1's claim — a constant term plus a term proportional to the byte count:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
#define HS_NSECS(bytes) (((55 * 8 * 2083) \
	+ (2083UL * (3 + BitTime(bytes))))/1000 \
	+ USB2_HOST_DELAY)

2083 is 2.083 ns in picoseconds — one bit at 480 Mbit/s (Chapter 12.5 §1), and BitTime(bytes) is 7 * 8 * bytes / 6, the kernel's allowance for bit stuffing.

2. The Numbers

Working HS_NSECS through at each legal high-speed bulk packet size:

PayloadTransactionPayload bits aloneEfficiencyPer 125 µs microframe
64 B2 171 ns1 067 ns49.1%57
128 B3 414 ns2 133 ns62.5%36
256 B5 904 ns4 267 ns72.3%21
512 B10 880 ns8 533 ns78.4%11

At the maximum bulk packet size, 78% of the bus time carries payload. At a quarter of it, half.

And the same shape at full speed, computed from the packet structure directly — a token of 3 bytes, a data packet of PID + payload + CRC16, a handshake of 1 byte, each with its own SYNC and EOP:

PayloadBytes on the wireEfficiencyPer 1 ms frame
8 B~18 B44.4%83
16 B~26 B61.5%57
32 B~42 B76.2%35
64 B~74 B86.5%20

3. Where the Rest of It Goes

The 22% that is not payload at 512 bytes, accounted for:

Roughly
Token packet — SYNC, PID, address, endpoint, CRC5, EOP~5 bytes on the wire
Data packet's own PID, CRC16, SYNC, EOP~8 bytes
Handshake packet~3 bytes
Inter-packet gaps and turnaround — Chapter 12.5two per transaction
Bit stuffingup to one extra bit per six

Bit stuffing is the one that surprises people. The encoding inserts a zero after six consecutive ones, so a payload of all-ones costs more bus time than a payload of alternating bits — and the kernel's BitTime allows for it with a flat 7/6 factor rather than computing it, because the transmitter cannot know the data in advance when it is reserving bandwidth.

The bus time a transfer consumes depends on its contents, which is a fact with no analogue in most buses and which makes exact throughput prediction impossible in principle.

4. A Zero-Length Packet Costs a Full Transaction

Chapter 13.2 §3 established that a transfer whose length is an exact multiple of wMaxPacketSize needs an explicit zero-length terminator. This chapter supplies its price.

A ZLP has no payload and pays the full fixed overhead — token, data packet, handshake, two turnarounds. At high speed that is:

~2 347 ns of bus time for zero bytes, which is 21.6% of a full 512-byte transaction.

And the transfers that need one are exactly the ones real software produces. A filesystem reads in 4 096-byte blocks; 4096 ÷ 512 = 8 exactly, so every block read ends with a ninth transaction carrying nothing.

TransferPayload transactionsZLP?TotalOverhead
4 096 B8yes91 transaction wasted in 9 — 11.1%
4 000 B8 (last is short)no8none
65 536 B128yes1290.8%

The cost falls as the transfer grows, which is the reassuring half. The alarming half is that a 512-byte transfer with a terminator is 50% overhead — one transaction of payload and one of nothing.

A sequence diagram comparing two bulk transfers on an endpoint whose maximum packet size is five hundred and twelve bytes. The first transfer is one thousand and twenty-four bytes, an exact multiple, so it is sent as two full packets; because neither is short, the receiver cannot tell the transfer has ended, and a third transaction carrying a zero-length packet is required as a terminator. The second transfer is six hundred bytes, sent as one full packet of five hundred and twelve and one short packet of eighty-eight; the short packet terminates the transfer naturally, so only two transactions are needed. The smaller transfer therefore costs fewer transactions than the larger one.Two transfers, three transactions and twoHostBulk endpointtransfer 1024 B — anexact multipletxn 1 · 512 B —full, more followstxn 2 · 512 B —full, more followsneither was short.is there more?txn 3 · 0 B —terminator, fulloverheadtransfer 600 Btxn 1 · 512 B — fulltxn 2 · 88 B —SHORT, ends it1024 B cost 3transactions; 600 Bcost 2
Figure 1 — two transfers of similar size costing different numbers of transactions. 1024 bytes is an exact multiple and needs a terminator, so it costs three transactions; 600 bytes ends naturally short and costs two. The 424-byte difference in payload is a one-transaction difference in the wrong direction.

5. The Packetiser, as RTL

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// ─────────────────────────────────────────────────────────────────────────
// usb_bulk_packetiser
//
// Classification: SIMPLIFIED SYNTHESIZABLE TEACHING RTL. It answers the
// question throughput actually depends on: how many TRANSACTIONS does this
// transfer cost, and how much bus time is that?
//
// WHAT IT MODELS. Sections 1 to 4: splitting a transfer into packets, the
// terminator rule, and the bus time each transaction consumes.
//
// WHAT IT DOES NOT MODEL. The packets themselves (Module 11); the
// transactions (Module 12); NAKs and retries (Chapter 14.3 -- every
// transaction here succeeds first time); other traffic on the bus (Chapter
// 10.3's arbitration); and bit stuffing's DATA DEPENDENCE (section 3), which
// is folded into a constant here exactly as the kernel folds it into
// BitTime.
//
// ── ON THE TIMING PARAMETERS ────────────────────────────────────────────
// OVERHEAD_NS and NS_PER_B_X100 are the two terms of section 1's model, and
// their defaults are derived from the kernel's HS_NSECS at 480 Mbit/s. They
// are PARAMETERS because they change with speed -- Chapter 12.5 section 1's
// rule that a quantity varying with the link is not a property of the design.
// ─────────────────────────────────────────────────────────────────────────
module usb_bulk_packetiser #(
  parameter int unsigned MAXP          = 512,
  parameter int unsigned OVERHEAD_NS   = 2347,   // fixed, per transaction
  parameter int unsigned NS_PER_B_X100 = 1667    // 16.67 ns/byte at 480 Mbit/s
)(
  input  logic        clk,
  input  logic        rst_n,

  input  logic        start,
  input  logic [31:0] xfer_len,
  input  logic        want_zlp,        // Chapter 13.2's terminator request

  output logic [15:0] pkt_len,
  output logic        pkt_valid,
  output logic        is_zlp,
  output logic        xfer_done,
  output logic [15:0] txn_count,
  output logic [31:0] bytes_moved,
  output logic [31:0] bus_ns
);

  logic [31:0] rem_q, bytes_q, ns_q;
  logic [15:0] txn_q;
  logic        active_q, zlp_owed_q;

  assign txn_count   = txn_q;
  assign bytes_moved = bytes_q;
  assign bus_ns      = ns_q;

  logic [15:0] this_len;
  assign this_len = (rem_q >= MAXP) ? MAXP[15:0] : rem_q[15:0];

  // A zero-length transfer IS a single zero-length transaction: that is what
  // "transfer nothing" looks like on a wire with no other way to say it.
  assign is_zlp    = active_q && (rem_q == 32'd0);
  assign pkt_valid = active_q;
  assign pkt_len   = is_zlp ? 16'd0 : this_len;

  // Done on THIS transaction, not after it. Written against the packet being
  // emitted now -- an output describing a state the block has already left
  // can never assert, and section 8 is about how that was found.
  assign xfer_done = active_q &&
                     ( is_zlp || ((rem_q == {16'd0, this_len}) && !zlp_owed_q) );

  always_ff @(posedge clk or negedge rst_n) begin
    if (!rst_n) begin
      rem_q <= '0; bytes_q <= '0; ns_q <= '0; txn_q <= '0;
      active_q <= 1'b0; zlp_owed_q <= 1'b0;
    end else if (start) begin
      rem_q      <= xfer_len;
      bytes_q    <= '0; ns_q <= '0; txn_q <= '0;
      active_q   <= 1'b1;
      // SECTION 4. A transfer that is an exact multiple of MAXP contains no
      // short packet, so it owes a terminator -- and that terminator is a
      // whole transaction at full price.
      zlp_owed_q <= want_zlp && ((xfer_len % MAXP) == 0);
    end else if (active_q) begin
      txn_q   <= txn_q + 16'd1;
      // EVERY transaction pays the fixed overhead, a zero-length one too.
      // Section 4 is the whole reason this line has no exception in it.
      ns_q    <= ns_q + OVERHEAD_NS + ((pkt_len * NS_PER_B_X100) / 100);
      bytes_q <= bytes_q + pkt_len;

      if (is_zlp) begin
        zlp_owed_q <= 1'b0;
        active_q   <= 1'b0;
      end else begin
        rem_q <= rem_q - pkt_len;
        if ((rem_q - pkt_len) == 32'd0) begin
          if (!zlp_owed_q) active_q <= 1'b0;
        end
      end
    end
  end

endmodule

What it models. The transaction count and bus time a bulk transfer costs.

Engineering reason. Because throughput is a question about transactions, not about bytes, and the transaction count is not the byte count divided by the packet size.

Inputs. A transfer length and whether its protocol wants a terminator.

State retained. A 32-bit remainder, byte and time accumulators, a 16-bit counter, and two flags.

Outputs. Each packet's length, the terminator indication, completion, and the three accumulated figures.

Hardware implied. A subtractor, a comparator, and three accumulators. The multiply is by a constant and folds away.

Reset behaviour. Everything cleared; a transfer does not survive a reset.

Assumptions. That every transaction succeeds first time — Chapter 14.3 removes this; that OVERHEAD_NS matches the link speed; and that bit stuffing is adequately modelled by a constant per byte, which §3 explains is a approximation the kernel also makes.

Omissions. Packets, transactions, retries, other traffic and data-dependent stuffing — in the header.

What DV should verify. That the transaction count matches the packet split including the terminator; that no packet exceeds MAXP; that a zero-length transfer is one transaction; that the bus-time accounting charges every transaction including a ZLP; that xfer_done asserts on the last transaction; and that a transfer always terminates.

Transactions, not bytes

9 cycles
A waveform of the bulk packetiser over nine cycles covering two transfers. The first transfer of one thousand and twenty-four bytes is started, then emits a five hundred and twelve byte packet, a second five hundred and twelve byte packet which completes the byte count, and then a zero-length terminator packet on which the transfer-done output asserts, bringing the transaction count to three. The second transfer of six hundred bytes is started, emits a five hundred and twelve byte packet and then an eighty-eight byte short packet on which transfer-done asserts, bringing its transaction count to two. The smaller transfer costs fewer transactions than the larger one.a whole transaction for zero bytesa whole transaction forzero bytes1024 B cost 3 transactions1024 B cost 3 transactions600 B cost 2 — short ends it600 B cost 2 — short endsitphasestart 1024pkt 1pkt 2ZLPidlestart 600pkt 1pkt 2idlepkt_len—5125120——51288—txn_count001233012bytes005121024102410240512600xfer_donet0t1t2t3t4t5t6t7t8
Figure 2 — the two transfers of Figure 1, every column taken from a simulation of the block above. The 1024-byte transfer reaches three transactions with its byte count already complete after two; the 600-byte transfer finishes in two. Note that the transaction counter, not the byte counter, is what bus time is proportional to.

6. Mutation Test

Five mutations over 212 transfers and 4 159 packets, with directed cases for every relationship to MAXP — exact multiples, one-short, one-over, zero-length, and a 4 096-byte filesystem block — plus 200 randomised transfers.

The unmutated block: 6 transfers ending with a terminator, and 0 errors of every kind. It reports a 64 KiB transfer as 128 transactions, 1 392 896 ns, 47.05 MB/s — 78% of the nominal 60 MB/s, matching §2's derivation exactly.

bytes wrongtransaction count wrongZLP wrongbus time wrongxfer_done missingnever ended
golden000000
T1 last packet always full20420400204204
T2 no terminator044000
T3 short transfers never report done00002060
T4 a ZLP costs nothing000600
T5 terminator skipped at the end044040

T1 — make every packet full-size

Measured: 204 transfers with the wrong byte count, wrong transaction count, and no termination.

The remainder never reaches zero because the last packet claims more bytes than remain. Catastrophic and immediate — included as the contrast for T4.

T2 — never owe a terminator

Measured: 4 of 212 transfers wrong.

§4's rule removed. Four is 1.9% of this bench and close to universal in production, for the reason Chapter 12.2 §6 measured: random lengths are rarely exact multiples of 512, and real software's lengths are almost always powers of two.

T3 — report completion only on a terminator

Measured: 206 transfers never reported done — every transfer that ended on a short packet rather than a ZLP.

And it is the mutation that revealed a bug in the golden design. §8.

T4 — charge nothing for a zero-length transaction

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
if (!is_zlp) ns_q <= ns_q + OVERHEAD_NS + ...;   // MUTANT T4

Measured: 6 transfers with mis-accounted bus time — exactly the six that ended with a terminator. Every other output identical.

The packet stream is correct; only the cost is wrong. A design with this defect packetises perfectly and under-reports its own bus consumption by a full transaction on every exact-multiple transfer.

Which matters because the number is used for scheduling. A host that believes a transfer is cheaper than it is over-commits the bus.

T5 — skip the terminator at the end of the transfer

Measured: 4 wrong transaction counts, 4 missing terminators, 4 missing completions.

The terminator is owed and the transfer ends without it. The same 4 transfers as T2 — and the same production-versus-random asymmetry.

7. Verification

This chapter's commit point is the transfer cost what the design says it cost.

Stimulus. Exact multiples of MAXP (512, 1024, 4096, 65536); one-under and one-over (511, 513); a zero-length transfer, with and without a terminator request; a 4 096-byte block with and without one; and 200 randomised lengths up to 20 000 bytes.

Observation. Every packet's length, the running counts, the accumulated bus time against an independently computed expectation, and xfer_done on every transaction.

Reference model. Three lines: ceil(len / MAXP) transactions, plus one if a terminator is owed, and overhead + len × rate of bus time per transaction. Its independence is real — it computes the totals from the transfer's parameters without reference to the packet stream, so it disagrees whenever the stream is wrong.

Coverage — crosses:

  • transfer length modulo MAXP: 0, 1, MAXP−1, mid-range
  • want_zlp × exact multiple — the cell that produces a terminator
  • zero-length transfer × each want_zlp
  • transfer sizes spanning 1 to 65 536 bytes
  • 4 096 bytes specifically — the filesystem block

Negative cases with defined outcomes: no packet exceeds MAXP; no transfer fails to terminate; no transaction is un-charged; and xfer_done never asserts on a non-final transaction.

8. Debugging: the Device That Is Fast Except on Round Numbers

A bulk device sustains close to its expected throughput on arbitrary transfer sizes. When the application switches to reading in 4 096-byte blocks, throughput drops by about 11% and stays there. Larger blocks recover most of it; 512-byte blocks halve it.

What does drops on round numbers tell you? That the defect is triggered by a property of the length, not of the data or the rate. Round numbers on a USB bulk endpoint means exact multiples of 512.

What happens on an exact multiple? §4: no short packet, so a terminator is owed — and the terminator is a whole transaction.

Does the arithmetic match? 4 096 ÷ 512 = 8 payload transactions plus one terminator = 9. One wasted transaction in nine is 11.1%, which is the reported drop.

And the 512-byte case? One payload transaction plus one terminator = 50% overhead, which halves throughput. Both numbers follow from the same rule.

So is this a bug? No — and that is the useful part. The device and host are both correct; the terminator is required, and the cost is inherent. The fault, if there is one, is in the application's block size.

What can actually be done? Three things, in decreasing order of usefulness:

  • Read in sizes that are not exact multiples — 4 095 or 4 097 bytes eliminates the terminator entirely;
  • Read in larger blocks — 64 KiB amortises one wasted transaction over 128, which is 0.8%;
  • Drop the terminator request if the protocol above does not need it — Chapter 13.2 §3: whether a transfer needs terminating is a property of the protocol running on top, not of USB.

What if none of those is available? Then the number is the number, and the right outcome of the investigation is a documented explanation rather than a fix.

The signature to keep: a throughput loss that appears at exact multiples of the packet size and scales as 1/(n+1) is a zero-length terminator — and it is arithmetic, not a defect.

9. Common Misconceptions

10. Reason It Through

Two devices are compared. Device A uses 512-byte bulk packets and sustains 40 MB/s. Device B declares 64-byte bulk packets and sustains 22 MB/s. The vendor of B argues that its smaller packets reduce latency and the throughput difference is a fair trade.

Is the throughput difference explained by the packet size? §2: 78.4% efficiency at 512 bytes against 49.1% at 64. The ratio is 0.63, and 22 ÷ 40 is 0.55 — close, and the remainder is accounted for by B needing 8× as many transactions, each with its own Chapter 12.5 turnaround.

So the arithmetic supports the vendor's premise. The loss is structural, not a defect.

Is the latency claim also true? Partly, and it is the weaker half. A smaller packet does complete sooner — but a bulk endpoint has no latency guarantee at all (Chapter 10.3 §1), so the latency being improved is not one anybody was promised.

What is the real effect of smaller packets on latency? It reduces the granularity of bus occupancy, which helps other endpoints — a 64-byte transaction occupies the bus for 2.2 µs where a 512-byte one occupies it for 10.9 µs. So B is not improving its own latency; it is being a better citizen to everything else on the bus.

Is that a good trade? It depends entirely on what else is attached, and that is the point the vendor's argument skips. On a bus with a latency-sensitive isochronous endpoint, B's smaller transactions genuinely help. On a bus where B is alone, it has given away 45% of its throughput for nothing.

And what should the evaluation actually measure? Not either device in isolation. Both devices alongside the traffic they will share a bus with — because the quantity the vendor is trading is not its own latency but somebody else's, and that quantity is invisible in a single-device test.

The transferable point: a trade-off argument names what is given up and often not who receives it. Here the beneficiary is a third party who may not exist, and the cost is borne by the device making the argument.

11. Understanding Check

12. Summary

Overhead is a fixed cost per transaction, so efficiency is a function of packet size alone — and the kernel's HS_NSECS has exactly that shape: a constant term plus one proportional to the bytes.

The numbers, derived rather than quoted: at high speed, 49.1% efficiency at 64 bytes rising to 78.4% at 512; at full speed, 44.4% at 8 bytes rising to 86.5% at 64. The faster bus is the less efficient one, because its fixed overhead is larger in bit times — raising the bit rate raises the rate at which overhead is paid.

At the maximum bulk packet size a high-speed endpoint gets 11 transactions per microframe — 45.1 MB/s against a nominal 60, which is 75%, before contention, retries or any other device.

And a zero-length terminator costs a full transaction — about 2 347 ns for no bytes, 21.6% of a 512-byte transaction. The transfers that need one are the ones real software produces: a 4 096-byte filesystem block is eight packets plus a ninth carrying nothing, which is the 11.1% of §8's debugging scenario, arithmetic rather than a defect.

§6 measured five mutations over 212 transfers, and the pair worth keeping is T2 and T4 — both about the terminator, and only one visible on the wire. T2 changes what is sent; T4 changes only what the design says it cost, and produces no protocol error at all.

A design that lies about its own cost is harder to catch than one that misbehaves, because nothing downstream contradicts it.

§7 is the chapter's finding, and it is about benches:

A quantity a bench reports is not a quantity a bench verifies — and a printed number looks identical whether the design is right or wrong.

Adding the missing check found more than a mutation. The completion output could never assert in the correct design: it was written against the state after the final transaction, by which time the block had already gone inactive. An output no test reads is an output no test constrains, and it is as likely to be dead as to be wrong.

13. What Comes Next

This chapter assumed every transaction succeeds first time. None of them do, always.

Chapter 14.2 is what Bulk actually guarantees — and the answer separates two things that sound like one: Bulk is guarantee-correct and bandwidth-best-effort. Data is delivered intact or the error is visible; when it is delivered is promised by nobody.

The chapter is about the ladder of mechanisms that make the first half true — CRC, the data toggle, retry, STALL, reset — and about what each one can and cannot recover from, because they are not interchangeable and a device that reaches for the wrong rung recovers from nothing.

Browse the full path on the USB tutorials index.

Continue learning

Standards & specifications

Governing standard
USB-IF (Universal Serial Bus Specification)(opens USB Implementers Forum (USB-IF) in a new tab)

Defines the USB bus — its electrical signalling, connectors, packet and transaction model, device framework and the descriptors a device must expose — together with the device-class specifications layered on it. It does not define host-controller register interfaces (xHCI and EHCI are separate documents) nor any operating system's driver architecture.

This page also covers RTL structure, verification approach and debugging technique. Those are engineering practice built on the standard, not requirements the standard itself imposes.

Where this fits

Part of the USB curriculum.