USB · Module 30
Debug Review Checklist
A triage algorithm, not a list of tools. Classify the symptom, pick the cheapest observation that discriminates, know what each evidence source cannot prove, and find the first event after which expected and observed architectural state disagree.
Everything in the three previous chapters has been done, and the device is failing anyway. This chapter is the procedure for the hour after that sentence is first said.
1. Why This Chapter Has No New RTL
The three chapters before this one each built a specimen because each was about a decision that lives in code. Debugging is not a decision that lives in code; it is a decision about which observation to make next, and the thing being debugged should be something the reader already understands.
So this chapter uses the specimens already built — 30.1's transfer controller and 30.4's interrupt register — and the designs from 29.5 and 29.6. Writing a fourth one would teach nothing about triage and would cost the reader the advantage of already knowing how the first three behave.
That is itself a review standard worth stating: a worked debug example is only worth reading if the reader can predict the correct behaviour. An example on unfamiliar hardware degenerates into a story.
2. Classify The Symptom First
The single most expensive mistake in USB debugging is starting from a tool. "Put a protocol analyser on it" is not a plan, because the analyser answers one narrow question and six of the eight common symptoms are not that question.
SYMPTOM FIRST QUESTION IT RAISES
----------------------------- ---------------------------------------
1 not detected at all is anything electrically present?
2 detected, enumeration which request failed, and in which
fails device state?
3 enumerates, then one is that endpoint's configuration what
endpoint fails firmware thinks it is?
4 one DIRECTION fails arbitration or ownership -- 29.6
5 works briefly, then a resource that is acquired and not
stalls released. Liveness, not correctness.
6 throughput far below packetisation and latency, not
expectation bandwidth -- 29.6 section 3
7 intermittent, load- a collision between two agents
dependent
8 fails on ONE host only conformance, or interoperability --
30.3 section 23. The Evidence Ladder
Once classified, work outside in, cheapest first — and "cheapest" means fastest to obtain, not simplest to understand.
The evidence ladder, cheapest discriminating observation first
Rung 4 is the one that is skipped, and it is nearly free. A debugger reading six registers while the device is stuck answers the ownership question directly, and ownership is what most of these symptoms are about. The reason it gets skipped is cultural: the analyser is the USB tool, so the USB problem must be on the wire.
4. What Each Source Cannot Prove
This is the part that separates a senior debugger from a thorough one. Every instrument has a blind spot, and using an instrument outside its competence is how a team spends a week confidently looking in the wrong place.
SOURCE PROVES CANNOT PROVE
---------------------- ------------------------- --------------------
protocol analyser what was on the wire, who anything about WHY.
answered, with which PID It cannot see a
and which handshake disagreement between
two blocks that are
both inside the
device.
controller registers the controller's own view what FIRMWARE
of ownership, state and believes. The whole
sticky errors class of bug is the
two disagreeing.
firmware log what software believed that the hardware
and when event actually
occurred -- only that
firmware thinks it
did
integrated logic exact RTL behaviour in anything the host
analyser the window it captured stack believed, and
nothing outside the
window
simulation replay what the RTL does for a that the stimulus
given stimulus matches what the
silicon actually saw
a scope electrical reality any protocol meaning
"it works on my that this host, today, conformance, and
machine" exercised a path that nothing about any
worked other host5. First Divergence
THE PRINCIPLE
Find the first event after which EXPECTED and OBSERVED architectural
state disagree. Everything after that point is a consequence, including
every symptom that is easier to see.Debugging backwards from the visible symptom is the default and it is usually wrong, because the symptom is many steps downstream and each step has several possible causes. Working forward from the last known-good state has one cause at each step.
The method is the same every time, and it is cheap when the expected model is simple:
1 BUILD THE EXPECTED SEQUENCE from evidence you already have -- the
protocol trace is usually enough. For an endpoint that means: the
packets the host sent, with lengths and PIDs.
2 BUILD THE OBSERVED SEQUENCE from the controller's registers or an
ILA capture.
3 FIND THE FIRST INDEX AT WHICH THEY DIFFER.
4 THE DEFECT IS AT THAT EVENT, not at the symptom.6. Worked Triage — Class 3
"It enumerates fine, then the bulk OUT endpoint stops accepting data after a few transfers."
1 CLASSIFY
Works, then stops -> class 5 as much as class 3. Something was
acquired and not released. Do NOT open the analyser yet.
2 WHAT IS THE DEVICE ANSWERING?
This is the one analyser question worth asking first, because the
two answers lead to completely different halves of the tree.
NAKing forever -> a resource is held. Go to 3.
ACKing, data lost -> a toggle mismatch or a checker in the
device discarding what it accepted. That is
29.5's territory, not this one.
silent -> not this layer. PHY, clock, or reset.
3 READ THE CONTROLLER'S REGISTERS (rung 4, nearly free)
ep_busy set, xfer_done set, firmware idle
-> firmware is not acknowledging. Did the ISR run? Is the
interrupt masked? Did a mask write clear the status --
30.4's C-M2?
ep_busy set, xfer_done set, firmware says it acknowledged
-> the two sides disagree. This is the interesting case and
it is where rung 4 earns its place: the disagreement is
invisible on the wire.
n_lost non-zero
-> the layer BELOW violated the contract. The bug is not
here; it is in whatever offered a packet while busy.
4 COMPARE THE TWO SIDES
Firmware's completion count against the controller's n_xfer. If
they differ, the difference is the number of events one side did
not see, and n_lost usually equals it.
5 FIRST DIVERGENCE
From the packet log build the expected byte totals. From n_xfer
and the completion log build the observed ones. The first
transfer whose total differs is the event.The shape to notice: four of the five steps are free, and the expensive one (an ILA capture) is last and is often unnecessary. A triage tree whose early steps are expensive is a triage tree nobody follows under pressure.
7. Worked Triage — Class 5, And The Liveness Case
"It works for about thirty seconds, then the transmit direction stops completely. Receive keeps working."
One direction stopping while the other continues is nearly diagnostic on its own. Data is not being corrupted — it is not being scheduled.
1 CLASSIFY class 4 and class 5 together. Arbitration.
2 THE QUESTION
Is the transmit path BLOCKED (waiting for something that never
comes) or STARVED (never selected)? They look identical from
outside and have different fixes.
3 THE DISCRIMINATOR
Stop the receive traffic. If transmit immediately resumes, it was
STARVED. If it stays stopped, it was BLOCKED.
That single observation costs nothing and splits the tree in half.
It is the best example in this chapter of choosing an observation
by what it DISCRIMINATES rather than by what it shows.
4 IF STARVED
Read the arbiter's state: which direction went last, is the burst
budget refilling, is the alternation bit stuck. 29.6's M1 is
exactly this and its signature is precise -- the transmit path
never runs while receive is saturated.
5 IF BLOCKED
Find what it is waiting for and who was supposed to provide it.
Usually a credit, a completion or an ownership bit that was
consumed twice.8. Worked Triage — Class 7, The Intermittent
"About once an hour, under load, a transfer completes with the wrong length. It never happens with the debugger attached."
1 CLASSIFY class 7. The debugger observation is not a nuisance; it
is evidence. Attaching it changed firmware's timing,
which means firmware's timing is a variable, which means
two agents are colliding.
2 ENUMERATE THE COLLISIONS
From the design's own contract: which pairs of events can occur in
the same cycle? 30.1 lists them for the transfer controller;
30.4 lists them for the interrupt register. There are never many.
3 FOR EACH, ASK WHAT THE WRONG RESOLUTION WOULD LOOK LIKE
completion overwritten -> a wrong LENGTH. Matches.
halt set/clear collision -> a stuck halt. Does not match.
status set/clear collision -> a missed interrupt. Does not match
directly, but produces a LATE
acknowledgement, which produces the
first one. Keep it.
4 READ THE STICKY COUNTERS
n_lost non-zero is the answer to the whole question, and it is one
register read. If the design has no such counter, this step does
not exist and the next one costs a week.
5 IF THERE IS NO COUNTER
Reproduce under an ILA with a trigger on the collision condition
itself. You now need to know which collision to trigger on, which
is why step 2 came first.The lesson is the one from 30.1's observability item, arriving with a bill attached: the difference between step 4 and step 5 is one flip-flop, decided eighteen months earlier by somebody who did not have to debug this.
9. The Checklist
1 BEFORE TOUCHING A TOOL
[ ] classify the symptom -- eight classes, and 5 and 7 are the
ones that get misclassified
[ ] state the competing hypotheses, out loud, before observing
[ ] pick the observation that DISCRIMINATES between them, not the
one that shows the most
2 THE LADDER
[ ] host/software evidence first -- it is free
[ ] then the wire, and ask only what the wire can answer
[ ] then the controller's registers -- nearly free, usually skipped
[ ] then RTL state, and only with a trigger you chose deliberately
3 EVIDENCE LIMITS
[ ] for every observation, say what it CANNOT prove
[ ] a source on one side of a seam cannot settle a disagreement
about the seam
[ ] a firmware log proves belief, not occurrence
[ ] an analyser cannot see two device-internal blocks disagreeing
4 FIRST DIVERGENCE
[ ] build the expected sequence from evidence already in hand
[ ] build the observed sequence
[ ] find the FIRST index where they differ
[ ] the expected model must be structurally SIMPLER than the
design, or it is a second design and not an oracle
5 THE INTERMITTENT
[ ] "it disappears with the debugger attached" is evidence, not an
obstacle -- it means timing is a variable
[ ] enumerate the same-cycle collisions from the contract
[ ] read the sticky counters before capturing anything
6 CLOSING IT
[ ] the root cause explains EVERY observation, including the ones
that seemed irrelevant
[ ] if one observation is unexplained, the investigation is not
finished, whatever the fix did
[ ] the finding goes back into 30.1's or 30.4's checklist10. Exercises
1 CLASSIFY
For each, give the class and the first observation:
(a) the device works on Linux and not on one Windows laptop
(b) throughput is 400 kB/s where 900 was expected
(c) the device enumerates, then disappears after ~40 seconds
(d) every fourth file transfer hangs, always at the same size
2 DISCRIMINATION
For (c) above, list three hypotheses and, for each, one observation
that would rule it OUT. Then pick the single observation that rules
out the most.
3 EVIDENCE LIMITS
A firmware log shows 4,213 interrupts serviced. The controller's
n_xfer reads 4,300. Name three situations consistent with both, and
the one register read that separates them.
4 FIRST DIVERGENCE
Given a packet log of lengths 64, 64, 64, 8 against an endpoint
with maxpkt 64 and limit 1024, write the expected completion
sequence. Then given observed completions of 200 and 8, say at
which event they first diverge and what that implicates.
5 THE ABSENT COUNTER
Take 30.1's submitted design, which has no n_lost. Write the debug
procedure for the class 7 symptom in section 8 without it, and
estimate the cost in hours against the version that has it.
6 LIVENESS
Design the single measurement you would add to 29.6's bus master
so that a starvation failure is diagnosable from one register read
after the fact. State its width and its reset behaviour.
7 THE UNEXPLAINED OBSERVATION
An investigation closes with a fix that resolves the symptom, but
one early observation was never explained. Argue for reopening it,
in three sentences, to a manager who wants to ship.11. The Interview Answer
"A device enumerates and then permanently NAKs an OUT endpoint after a few transfers. Give me your plan."
First I classify, because "works then stops" is a different problem from "never worked" and they share almost no branches. Working-then-stopping means the data path is fine and something was acquired and not released — a completion slot, a buffer, a credit. That is a resource-accounting problem, so I am not opening a protocol analyser first.
There is one analyser question worth asking up front, because it splits the tree in half: is the device NAKing, or is it ACKing? NAKing forever means a resource is held. ACKing while the data never arrives means the device accepted and discarded it, which on a bulk endpoint means the data toggle is wrong, and the commonest cause of that is a bus reset that did not clear it.
Given NAKing, I go to the controller's registers, which is the rung people skip and which is nearly free. I want the ownership state: is a completion pending, does firmware think it acknowledged it, and is there a sticky counter saying the contract was violated. There are three outcomes and they are different bugs. Pending with firmware idle means the interrupt did not arrive or was masked away — and a mask write that also clears status will do exactly that, destroying the events a critical section was protecting. Pending with firmware claiming it acknowledged means the two sides disagree, and that is the case the analyser cannot see at all, because a NAKing device is a perfectly legal thing on the wire. A non-zero loss counter means the bug is in the layer below and I am looking at the victim.
Then first divergence. From the packet log I can build the expected sequence of transfer completions on paper — it is a running total and a completion rule. From the controller's counters I get the observed one. The first transfer where they differ is the event, and everything after it is consequence. That works only because the expected model is simpler than the design; if I had to model the pointers to predict the answer, I would have two designs and no oracle.
The thing I would say unprompted is about the instruments. Every one of them has a blind spot and using one outside its competence costs days. A firmware log proves what software believed, not what happened. An analyser proves what was on the wire and can tell you nothing about two blocks inside the device disagreeing. A source on one side of a seam cannot settle an argument about the seam — you need a reading from each side, and whether one exists was decided in the RTL review eighteen months ago, by somebody who did not have to debug this.
12. What Carries Forward
THE PROCEDURE
o classify the symptom before touching a tool; eight classes, and 5
(works then stops) and 7 (intermittent, load-dependent) are the two
that get misclassified
o choose an observation by what it DISCRIMINATES, not by what it
shows -- an observation every hypothesis predicts is free and
worthless
o the ladder is outside in, and rung 4 -- the controller's own
registers -- is nearly free and routinely skipped
EVIDENCE LIMITS
o a protocol analyser cannot see two device-internal blocks
disagreeing; a NAKing device is legal and unremarkable on the wire
o a firmware log proves BELIEF, not occurrence
o an ILA proves nothing outside its window, and the window needs a
trigger you chose for a reason
o a source on one side of a seam cannot settle a disagreement about
the seam
FIRST DIVERGENCE
o build expected from evidence in hand, observed from the device, and
find the first index where they differ
o the expected model must be structurally SIMPLER than the design --
otherwise it is a second design and not an oracle
THE TWO HARD CLASSES
o a liveness failure leaves no bad data behind, so nothing that looks
at data will find it: the evidence must be a MEASUREMENT against a
bound, and the bound has to have been designed in
o "it disappears with the debugger attached" is evidence: firmware
timing is a variable, which means two agents are colliding
o enumerate the same-cycle collisions from the contract -- there are
never many -- and read the sticky counters before capturing anything
THE BILL
o the difference between a one-register-read diagnosis and a week is
one flip-flop, decided in a review long before the failureThe last chapter turns all six reviews around and points them at the reader.
Continue learning
Related tutorials
- Related topic
Senior Silicon Debug
Walk a bring-up debug session from an analyser trace — narrowing by what the evidence excludes, never by what it suggests, because every USB symptom is consistent with five different faults.
- Related topic
Differential Signalling
Why the receiver's decision depends on the difference between two conductors rather than either one's level — what common-mode rejection actually buys, the routing and matching conditions it is conditional on, and the precise sense in which differential signalling does not remove noise.
- Related topic
Pull-Up Resistors
How a device announces itself with no protocol available: a resistor working against the host's opposing pull-down, why the ratio between them makes the announcement unambiguous, why the choice of conductor carries the speed, and what it means that a device can choose when to present it.
- Related topic
USB Protocol Checkers
Every ordering rule says what must happen next, and none of them fires when nothing happens at all — the timeout is a checker's only liveness tool, and a checker with false positives gets switched off.
Standards & specifications
- Governing standard
- USB-IF (Universal Serial Bus Specification)(opens USB Implementers Forum (USB-IF) in a new tab)
Defines the USB bus — its electrical signalling, connectors, packet and transaction model, device framework and the descriptors a device must expose — together with the device-class specifications layered on it. It does not define host-controller register interfaces (xHCI and EHCI are separate documents) nor any operating system's driver architecture.
This page also covers RTL structure, verification approach and debugging technique. Those are engineering practice built on the standard, not requirements the standard itself imposes.
Where this fits
Part of the USB curriculum.
