USB · Module 25
Endpoint Problems
A NAK is not an error and a STALL is not a NAK — one is flow control working, one is firmware refusing permanently, and a monitor that treats them alike either floods the log or misses the endpoint that has stopped.
Chapter 25.2 was about a device describing itself wrongly. This chapter is about a device that describes itself perfectly and then stops answering on one endpoint out of four.
1. Three Words That Look Alike on a Trace
Open any analyser capture of a working bulk transfer and the overwhelming majority of the lines say NAK. Open a capture of a broken one and the overwhelming majority of the lines also say NAK. The three handshakes an endpoint can return are:
ACK the transfer happened
NAK not ready YET -- try again.
Flow control. The endpoint is WORKING.
STALL no, and the answer will stay no until the host
sends CLEAR_FEATURE(ENDPOINT_HALT).
NYET (high speed) accepted, but do not send the next
one immediately.They occupy one line each in the log and they mean completely different things about the health of the device. The whole of this chapter follows from two sentences:
The same token, four answers
2. A NAK Is Not an Error — It Is Flow Control Working
A bulk endpoint has no guaranteed bandwidth. It gets whatever is left after the isochronous and interrupt traffic has been scheduled, and the way it refuses work it cannot take is by NAKing. A bulk IN endpoint with nothing to send NAKs every time it is polled, and the host polls it as often as the schedule allows.
So on a busy bus, a device doing exactly the right thing emits thousands of NAKs per second. The count is not a defect rate. It is not even loosely correlated with one.
This is chapter 24.1's false-positive argument in its most concrete form, and the consequence is the one that argument predicts:
3. What Matters Is the Run, Not the Count
If a NAK is not an error, what is? Not a bigger number of NAKs — a busy endpoint and a dead one both produce large numbers. The distinguishing property is consecutiveness:
healthy and busy NAK NAK NAK NAK ACK NAK NAK ACK NAK ...
wedged NAK NAK NAK NAK NAK NAK NAK NAK NAK ...
Same NAK rate over a long window. Same total.
Completely different devices.So the measurement is a run length — consecutive NAKs since this endpoint last answered — and the thing that resets it is a successful transfer, not the passage of time:
ACK -> run = 0 the endpoint answered
NYET -> run = 0 the endpoint answered (and asked to pause)
NAK -> run = run + 1
STALL -> run = 0 a different fault entirely; see section 5
run >= NAK_RUN_MAX -> WEDGEDAn endpoint that NAKs a hundred times and then ACKs once is healthy and busy. One that NAKs a hundred times and never ACKs has stopped. A total NAK count cannot separate those two; a run length separates them on the first sample past the threshold.
4. ...And the Run Must Be Per Endpoint
This is the part that is most often got wrong, and it is worth stating on its own because the wrong version passes a casual test.
A single run counter for the whole device is reset whenever any endpoint answers. On a device with four endpoints where one is wedged and three are healthy, the three healthy ones answer constantly — and each of their answers clears the shared counter. The wedged endpoint never reaches the threshold.
5. A STALL Is Sticky, and Polling a Halted Endpoint Is a Different Bug
A STALL puts the endpoint into the halted condition. It stays halted. Nothing on the wire clears it except CLEAR_FEATURE(ENDPOINT_HALT) addressed to that endpoint — not a successful transfer, not a SOF, not time.
A monitor that clears its halt flag when the endpoint next completes a transfer hides the single most useful fact about a stalled endpoint: that firmware set it and nobody ever cleared it.
And once an endpoint is halted, a host that keeps issuing tokens to it has not noticed. That is worth counting separately from the stall:
STALL seen -> the DEVICE took the endpoint out of service
token to halted ep -> the HOST has not noticed
Same two lines on a trace. Different team gets the bug.The reverse disagreement is worth reporting too. CLEAR_FEATURE(ENDPOINT_HALT) sent to an endpoint that is not halted is harmless on the wire and means the driver believes something the device does not — which is a disagreement worth knowing about before it turns into a lockup.
One endpoint's health, as a state machine
6. The Toggle Catches Two Opposite Faults and Cannot Tell Them Apart
Every successful transfer on an endpoint alternates the DATA toggle: DATA0, DATA1, DATA0, and so on. A transfer arriving with the wrong toggle means the sequence broke. It is the cheapest loss-and-duplication detector on the bus, and it is worth knowing exactly what it can and cannot do:
expected DATA0, got DATA1 -> the sequence broke
expected DATA1, got DATA0 -> the sequence broke
A one-bit field has exactly ONE wrong value.
"the device re-sent a packet the host already took"
"a packet was lost"
... produce the SAME symptom.One more thing the toggle needs: CLEAR_FEATURE(ENDPOINT_HALT) resets it to DATA0. An implementation that clears the halt and leaves the toggle where it was is out of sync with the host by exactly one packet, forever — which is chapter 21.1's point about a bus reset applied to a smaller piece of state.
7. What We Are Building
usb_endpoint_health watches completed transactions on four endpoints and reports exactly four things, each with its own counter:
| Code | Fires when | Points at |
|---|---|---|
H_WEDGED | NAK_RUN_MAX consecutive NAKs, once per wedge | the device |
H_POLLED_HALTED | a token arrives for a halted endpoint | the host / driver |
H_TOGGLE | the DATA toggle sequence broke | either — see section 6 |
H_CLEAR_UNHALTED | CLEAR_FEATURE on an endpoint that is not halted | the driver |
A NAK is not on that list, and neither is a STALL. Both are events, both are counted, and neither is an error.
usb_endpoint_health — three pieces of per-endpoint state
8. Verilog-2005 Implementation
// usb_endpoint_health -- three things that look alike on a trace and mean
// completely different things.
//
// ACK the transfer happened
// NAK the endpoint is not ready YET. Flow control, working.
// STALL firmware is saying NO, and will keep saying no until
// somebody explicitly clears it.
//
// A NAK IS NOT AN ERROR
//
// On a busy bus a bulk endpoint NAKs constantly -- that is what a bulk
// endpoint is FOR. Counting each NAK as an error produces thousands of
// entries per second and a log nobody reads, which is chapter 24.1's
// false-positive argument in its most concrete form.
//
// WHAT MATTERS IS THE RUN, NOT THE COUNT
//
// One NAK is nothing. A thousand CONSECUTIVE NAKs from one endpoint is an
// endpoint that has stopped working, and no individual NAK in that run is
// an error. So the measurement is a RUN LENGTH:
//
// nak_run[ep] consecutive NAKs since this endpoint last answered
//
// and it is reset by a successful transfer, not by time. An endpoint that
// NAKs a hundred times and then ACKs once is healthy and busy; one that
// NAKs a hundred times and never ACKs is wedged, and the two are
// indistinguishable by total count.
//
// ...AND THE RUN IS PER ENDPOINT
//
// A single global run counter is reset by any endpoint answering, so a
// wedged endpoint on a device with three healthy ones is invisible: the
// others keep clearing the counter. The wedged endpoint is exactly the one
// that never resets it, which is the whole signal, and a shared counter
// deletes it.
//
// A STALL IS STICKY
//
// A halted endpoint stays halted until the host sends
// CLEAR_FEATURE(ENDPOINT_HALT). It does not clear itself, it does not clear
// on the next successful transfer, and a monitor that clears it on an ACK
// hides the single most useful fact about a stalled endpoint: that firmware
// set it and nobody ever cleared it.
//
// AND POLLING A HALTED ENDPOINT IS THE HOST'S BUG
//
// A host that keeps issuing tokens to an endpoint that has STALLed is a
// host that has not noticed. That is worth counting SEPARATELY from the
// stall itself, because it points at the driver rather than at the device
// -- and on a trace the two look identical.
//
// THE DATA TOGGLE CATCHES LOSS AND DUPLICATION, AND CANNOT TELL THEM APART
//
// The toggle alternates on every successful transfer. A transfer arriving
// with the wrong toggle means the sequence broke -- but a one-bit field has
// only one wrong value, so "the packet was re-sent" and "a packet was lost"
// produce the SAME symptom.
//
// A one-bit toggle detects that something went wrong.
// It cannot tell you which of two opposite things it was.
//
// That is not a defect in this block; it is a property of a one-bit
// sequence number, and it is why chapter 24.3's scoreboard needs a real
// identity rather than a toggle.
module usb_endpoint_health #(
parameter integer N_EP = 4, // endpoints watched
parameter integer NAK_RUN_MAX = 16 // consecutive NAKs before "wedged"
) (
input wire clk,
input wire rst_n,
input wire xact_valid, // a transaction completed on this endpoint
input wire [1:0] ep,
input wire [1:0] resp, // 0 ACK, 1 NAK, 2 STALL, 3 NYET
input wire toggle, // the DATA toggle this transfer carried
input wire clear_halt, // CLEAR_FEATURE(ENDPOINT_HALT) for `ep`
input wire eot,
output wire [N_EP-1:0] halted,
output wire [7:0] nak_run, // the run on the endpoint named by `ep`
output wire [7:0] worst_run,
output wire exp_toggle, // what `ep` should send next
output wire err_pulse,
output wire [2:0] err_code,
output wire stall_pulse, // a NEW halt, not an error in itself
output reg [31:0] n_ack,
output reg [31:0] n_nak,
output reg [31:0] n_stall,
output reg [31:0] n_nyet,
output reg [31:0] n_wedged,
output reg [31:0] n_polled_halted,
output reg [31:0] n_toggle,
output reg [31:0] n_clear_unhalted,
output reg [31:0] n_clear
);
localparam [1:0] R_ACK = 2'd0, R_NAK = 2'd1, R_STALL = 2'd2, R_NYET = 2'd3;
localparam [2:0] H_NONE = 3'd0,
H_WEDGED = 3'd1, // NAK_RUN_MAX consecutive NAKs
H_POLLED_HALTED = 3'd2, // a token to a halted endpoint
H_TOGGLE = 3'd3, // the toggle sequence broke
H_CLEAR_UNHALTED = 3'd4; // CLEAR_FEATURE on a live one
reg [N_EP-1:0] halt_r;
reg [7:0] run_r [0:N_EP-1];
reg tog_r [0:N_EP-1];
reg wed_r [0:N_EP-1]; // already reported wedged: report ONCE
reg [7:0] worst_r;
reg [2:0] ec_r;
reg er_r, sp_r;
assign halted = halt_r;
assign nak_run = run_r[ep];
assign worst_run = worst_r;
assign exp_toggle = tog_r[ep];
assign err_pulse = er_r;
assign err_code = ec_r;
assign stall_pulse = sp_r;
integer i;
reg [N_EP-1:0] halt_n;
reg [7:0] run_n [0:N_EP-1];
reg tog_n [0:N_EP-1];
reg wed_n [0:N_EP-1];
reg [7:0] worst_n;
reg [2:0] ec_n;
reg er_n, sp_n;
reg ack_n, nak_n, stl_n, nyt_n, clr_n;
always @* begin
halt_n = halt_r;
worst_n = worst_r;
ec_n = H_NONE;
er_n = 1'b0;
sp_n = 1'b0;
ack_n = 1'b0; nak_n = 1'b0; stl_n = 1'b0; nyt_n = 1'b0; clr_n = 1'b0;
for (i = 0; i < N_EP; i = i + 1) begin
run_n[i] = run_r[i];
tog_n[i] = tog_r[i];
wed_n[i] = wed_r[i];
end
if (eot) begin
// Nothing: the counters are the report.
end else if (clear_halt) begin
// ---- CLEAR_FEATURE(ENDPOINT_HALT). ----
if (!halt_r[ep]) begin
// Clearing a halt that was never set. Harmless on the wire, and a
// useful signal: it means the driver believes the endpoint is
// halted and the device does not, which is a disagreement worth
// knowing about before it becomes a lockup.
er_n = 1'b1; ec_n = H_CLEAR_UNHALTED;
end
halt_n[ep] = 1'b0;
// A cleared endpoint restarts at DATA0. Chapter 21.1 makes the same
// point about a bus reset: the toggle is part of the state that has
// to be reset with everything else, and forgetting it costs exactly
// one silently-dropped packet.
tog_n[ep] = 1'b0;
run_n[ep] = 8'd0;
wed_n[ep] = 1'b0;
clr_n = 1'b1;
end else if (xact_valid) begin
if (halt_r[ep]) begin
// ---- The host is still polling a HALTED endpoint. ----
//
// Counted separately from the stall, because it points at the
// driver rather than at the device -- and on a trace the two look
// identical.
er_n = 1'b1; ec_n = H_POLLED_HALTED;
end else begin
case (resp)
R_ACK: begin
ack_n = 1'b1;
if (toggle != tog_r[ep]) begin
// ---- The toggle sequence broke. ----
//
// Which direction it broke in is NOT recoverable from one
// bit: a re-sent packet and a lost one produce the same
// symptom. See the note in the header.
er_n = 1'b1; ec_n = H_TOGGLE;
end
tog_n[ep] = ~tog_r[ep];
// ---- A successful transfer resets the run. ----
//
// By the transfer, not by time. An endpoint that NAKs a hundred
// times and then answers is healthy and busy.
run_n[ep] = 8'd0;
wed_n[ep] = 1'b0;
end
R_NAK: begin
// NOT an error. This is flow control working, and counting each
// one produces a log nobody reads.
nak_n = 1'b1;
if (run_r[ep] != 8'hFF) run_n[ep] = run_r[ep] + 8'd1;
if (run_n[ep] > worst_r) worst_n = run_n[ep];
if ((run_n[ep] >= NAK_RUN_MAX[7:0]) && !wed_r[ep]) begin
// ---- The RUN is the error, not the NAK. ----
//
// Reported ONCE per wedge: an endpoint that is stuck produces
// a NAK every microframe, and one report per NAK is the same
// flood the individual-NAK counting produced.
er_n = 1'b1; ec_n = H_WEDGED;
wed_n[ep] = 1'b1;
end
end
R_STALL: begin
// ---- STICKY. Until CLEAR_FEATURE, and nothing else. ----
stl_n = 1'b1;
sp_n = 1'b1;
halt_n[ep] = 1'b1;
run_n[ep] = 8'd0;
end
default: begin
// NYET: accepted, but not ready for the next one. Like a NAK it
// is not an error, and unlike a NAK it does not mean the
// transfer failed -- so it resets the run.
nyt_n = 1'b1;
run_n[ep] = 8'd0;
wed_n[ep] = 1'b0;
end
endcase
end
end
end
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
halt_r <= {N_EP{1'b0}};
worst_r <= 8'd0;
ec_r <= H_NONE;
er_r <= 1'b0;
sp_r <= 1'b0;
for (i = 0; i < N_EP; i = i + 1) begin
run_r[i] <= 8'd0;
tog_r[i] <= 1'b0;
wed_r[i] <= 1'b0;
end
n_ack <= 32'd0;
n_nak <= 32'd0;
n_stall <= 32'd0;
n_nyet <= 32'd0;
n_wedged <= 32'd0;
n_polled_halted <= 32'd0;
n_toggle <= 32'd0;
n_clear_unhalted <= 32'd0;
n_clear <= 32'd0;
end else begin
halt_r <= halt_n;
worst_r <= worst_n;
ec_r <= ec_n;
er_r <= er_n;
sp_r <= sp_n;
for (i = 0; i < N_EP; i = i + 1) begin
run_r[i] <= run_n[i];
tog_r[i] <= tog_n[i];
wed_r[i] <= wed_n[i];
end
if (ack_n) n_ack <= n_ack + 32'd1;
if (nak_n) n_nak <= n_nak + 32'd1;
if (stl_n) n_stall <= n_stall + 32'd1;
if (nyt_n) n_nyet <= n_nyet + 32'd1;
if (clr_n) n_clear <= n_clear + 32'd1;
// The per-cause counters are driven by the SAME pulse as the error,
// so they sum to it by construction (chapter 23.4).
if (er_n) begin
case (ec_n)
H_WEDGED: n_wedged <= n_wedged + 32'd1;
H_POLLED_HALTED: n_polled_halted <= n_polled_halted + 32'd1;
H_TOGGLE: n_toggle <= n_toggle + 32'd1;
H_CLEAR_UNHALTED: n_clear_unhalted <= n_clear_unhalted + 32'd1;
default: ;
endcase
end
end
end
endmodule9. SystemVerilog Implementation
Same block, with the two encodings promoted to enumerated types so the waveform viewer prints H_POLLED_HALTED instead of 3'd2.
// usb_endpoint_health -- three things that look alike on a trace and mean
// completely different things.
//
// ACK the transfer happened
// NAK the endpoint is not ready YET. Flow control, working.
// STALL firmware is saying NO, and will keep saying no until
// somebody explicitly clears it.
//
// A NAK IS NOT AN ERROR
//
// On a busy bus a bulk endpoint NAKs constantly -- that is what a bulk
// endpoint is FOR. Counting each NAK as an error produces thousands of
// entries per second and a log nobody reads, which is chapter 24.1's
// false-positive argument in its most concrete form.
//
// WHAT MATTERS IS THE RUN, NOT THE COUNT
//
// One NAK is nothing. A thousand CONSECUTIVE NAKs from one endpoint is an
// endpoint that has stopped working, and no individual NAK in that run is
// an error. So the measurement is a RUN LENGTH, reset by a successful
// transfer rather than by time. An endpoint that NAKs a hundred times and
// then ACKs once is healthy and busy; one that NAKs a hundred times and
// never ACKs is wedged, and the two are indistinguishable by total count.
//
// ...AND THE RUN IS PER ENDPOINT
//
// A single global run counter is reset by ANY endpoint answering, so a
// wedged endpoint on a device with three healthy ones is invisible: the
// others keep clearing the counter. The wedged endpoint is exactly the one
// that never resets it, which is the whole signal, and a shared counter
// deletes it.
//
// A STALL IS STICKY, AND POLLING A HALTED ENDPOINT IS THE HOST'S BUG
//
// A halted endpoint stays halted until CLEAR_FEATURE(ENDPOINT_HALT). It does
// not clear on the next successful transfer, and a host that keeps issuing
// tokens to it has not noticed -- which is worth counting SEPARATELY,
// because it points at the driver rather than at the device.
//
// THE TOGGLE CATCHES LOSS AND DUPLICATION AND CANNOT TELL THEM APART
//
// A one-bit sequence number has only one wrong value, so "re-sent" and
// "lost" produce the same symptom. That is a property of one bit, not a
// defect here, and it is why chapter 24.3's scoreboard needs a real
// identity.
package usb_eph_pkg;
typedef enum logic [1:0] {
R_ACK = 2'd0, // the transfer happened
R_NAK = 2'd1, // not ready YET -- flow control, not an error
R_STALL = 2'd2, // firmware says no, until explicitly cleared
R_NYET = 2'd3 // accepted, but not ready for the next one
} resp_e;
typedef enum logic [2:0] {
H_NONE = 3'd0,
H_WEDGED = 3'd1, // NAK_RUN_MAX consecutive NAKs
H_POLLED_HALTED = 3'd2, // a token to a halted endpoint
H_TOGGLE = 3'd3, // the toggle sequence broke
H_CLEAR_UNHALTED = 3'd4 // CLEAR_FEATURE on a live one
} health_e;
endpackage
module usb_endpoint_health
import usb_eph_pkg::*;
#(
parameter int N_EP = 4,
parameter int NAK_RUN_MAX = 16
) (
input logic clk,
input logic rst_n,
input logic xact_valid,
input logic [1:0] ep,
input resp_e resp,
input logic toggle,
input logic clear_halt,
input logic eot,
output logic [N_EP-1:0] halted,
output logic [7:0] nak_run,
output logic [7:0] worst_run,
output logic exp_toggle,
output logic err_pulse,
output health_e err_code,
output logic stall_pulse,
output logic [31:0] n_ack,
output logic [31:0] n_nak,
output logic [31:0] n_stall,
output logic [31:0] n_nyet,
output logic [31:0] n_wedged,
output logic [31:0] n_polled_halted,
output logic [31:0] n_toggle,
output logic [31:0] n_clear_unhalted,
output logic [31:0] n_clear
);
logic [N_EP-1:0] halt_r;
logic [7:0] run_r [N_EP];
logic tog_r [N_EP];
logic wed_r [N_EP]; // already reported wedged: report ONCE
logic [7:0] worst_r;
health_e ec_r;
logic er_r, sp_r;
assign halted = halt_r;
assign nak_run = run_r[ep];
assign worst_run = worst_r;
assign exp_toggle = tog_r[ep];
assign err_pulse = er_r;
assign err_code = ec_r;
assign stall_pulse = sp_r;
logic [N_EP-1:0] halt_n;
logic [7:0] run_n [N_EP];
logic tog_n [N_EP];
logic wed_n [N_EP];
logic [7:0] worst_n;
health_e ec_n;
logic er_n, sp_n;
logic ack_n, nak_n, stl_n, nyt_n, clr_n;
int i;
always_comb begin
halt_n = halt_r;
worst_n = worst_r;
ec_n = H_NONE;
er_n = 1'b0;
sp_n = 1'b0;
ack_n = 1'b0; nak_n = 1'b0; stl_n = 1'b0; nyt_n = 1'b0; clr_n = 1'b0;
for (i = 0; i < N_EP; i++) begin
run_n[i] = run_r[i];
tog_n[i] = tog_r[i];
wed_n[i] = wed_r[i];
end
if (eot) begin
// Nothing: the counters are the report.
end else if (clear_halt) begin
// ---- CLEAR_FEATURE(ENDPOINT_HALT). ----
if (!halt_r[ep]) begin
// Clearing a halt that was never set: the driver believes the
// endpoint is halted and the device does not. Harmless on the wire,
// and a disagreement worth knowing about before it becomes a lockup.
er_n = 1'b1; ec_n = H_CLEAR_UNHALTED;
end
halt_n[ep] = 1'b0;
// A cleared endpoint restarts at DATA0. Chapter 21.1 makes the same
// point about a bus reset: the toggle is part of the state that has
// to be reset with everything else, and forgetting it costs exactly
// one silently-dropped packet.
tog_n[ep] = 1'b0;
run_n[ep] = 8'd0;
wed_n[ep] = 1'b0;
clr_n = 1'b1;
end else if (xact_valid) begin
if (halt_r[ep]) begin
// ---- The host is still polling a HALTED endpoint. ----
er_n = 1'b1; ec_n = H_POLLED_HALTED;
end else begin
case (resp)
R_ACK: begin
ack_n = 1'b1;
if (toggle != tog_r[ep]) begin
// The sequence broke. WHICH WAY is not recoverable from one
// bit -- see the note in the header.
er_n = 1'b1; ec_n = H_TOGGLE;
end
tog_n[ep] = ~tog_r[ep];
// ---- A successful transfer resets the run. By the transfer,
// ---- not by time.
run_n[ep] = 8'd0;
wed_n[ep] = 1'b0;
end
R_NAK: begin
// NOT an error. Flow control working.
nak_n = 1'b1;
if (run_r[ep] != 8'hFF) run_n[ep] = run_r[ep] + 8'd1;
if (run_n[ep] > worst_r) worst_n = run_n[ep];
if ((run_n[ep] >= 8'(NAK_RUN_MAX)) && !wed_r[ep]) begin
// ---- The RUN is the error, not the NAK -- and it is
// ---- reported ONCE. A stuck endpoint NAKs every
// ---- microframe, and one report per NAK is the same flood
// ---- that counting individual NAKs produced.
er_n = 1'b1; ec_n = H_WEDGED;
wed_n[ep] = 1'b1;
end
end
R_STALL: begin
// ---- STICKY. Until CLEAR_FEATURE, and nothing else. ----
stl_n = 1'b1;
sp_n = 1'b1;
halt_n[ep] = 1'b1;
run_n[ep] = 8'd0;
end
R_NYET: begin
// Accepted, but not ready for the next one. Like a NAK it is
// not an error; unlike a NAK the transfer did happen, so it
// resets the run.
nyt_n = 1'b1;
run_n[ep] = 8'd0;
wed_n[ep] = 1'b0;
end
endcase
end
end
end
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
halt_r <= '0;
worst_r <= 8'd0;
ec_r <= H_NONE;
er_r <= 1'b0;
sp_r <= 1'b0;
for (int j = 0; j < N_EP; j++) begin
run_r[j] <= 8'd0;
tog_r[j] <= 1'b0;
wed_r[j] <= 1'b0;
end
n_ack <= 32'd0;
n_nak <= 32'd0;
n_stall <= 32'd0;
n_nyet <= 32'd0;
n_wedged <= 32'd0;
n_polled_halted <= 32'd0;
n_toggle <= 32'd0;
n_clear_unhalted <= 32'd0;
n_clear <= 32'd0;
end else begin
halt_r <= halt_n;
worst_r <= worst_n;
ec_r <= ec_n;
er_r <= er_n;
sp_r <= sp_n;
for (int j = 0; j < N_EP; j++) begin
run_r[j] <= run_n[j];
tog_r[j] <= tog_n[j];
wed_r[j] <= wed_n[j];
end
if (ack_n) n_ack <= n_ack + 32'd1;
if (nak_n) n_nak <= n_nak + 32'd1;
if (stl_n) n_stall <= n_stall + 32'd1;
if (nyt_n) n_nyet <= n_nyet + 32'd1;
if (clr_n) n_clear <= n_clear + 32'd1;
// The per-cause counters are driven by the SAME pulse as the error,
// so they sum to it by construction (chapter 23.4).
if (er_n) begin
case (ec_n)
H_WEDGED: n_wedged <= n_wedged + 32'd1;
H_POLLED_HALTED: n_polled_halted <= n_polled_halted + 32'd1;
H_TOGGLE: n_toggle <= n_toggle + 32'd1;
H_CLEAR_UNHALTED: n_clear_unhalted <= n_clear_unhalted + 32'd1;
default: ;
endcase
end
end
end
endmodule10. VHDL-2008 Implementation
-- usb_endpoint_health -- three things that look alike on a trace and mean
-- completely different things.
--
-- ACK the transfer happened
-- NAK the endpoint is not ready YET. Flow control, working.
-- STALL firmware is saying NO, and will keep saying no until
-- somebody explicitly clears it.
--
-- A NAK IS NOT AN ERROR
--
-- On a busy bus a bulk endpoint NAKs constantly -- that is what a bulk
-- endpoint is FOR. Counting each NAK as an error produces thousands of
-- entries per second and a log nobody reads, which is chapter 24.1's
-- false-positive argument in its most concrete form.
--
-- WHAT MATTERS IS THE RUN, NOT THE COUNT
--
-- One NAK is nothing. A thousand CONSECUTIVE NAKs from one endpoint is an
-- endpoint that has stopped working, and no individual NAK in that run is
-- an error. So the measurement is a RUN LENGTH, reset by a successful
-- transfer rather than by time. An endpoint that NAKs a hundred times and
-- then ACKs once is healthy and busy; one that NAKs a hundred times and
-- never ACKs is wedged, and the two are indistinguishable by total count.
--
-- ...AND THE RUN IS PER ENDPOINT
--
-- A single global run counter is reset by ANY endpoint answering, so a
-- wedged endpoint on a device with three healthy ones is invisible: the
-- others keep clearing the counter. The wedged endpoint is exactly the one
-- that never resets it, which is the whole signal, and a shared counter
-- deletes it.
--
-- A STALL IS STICKY, AND POLLING A HALTED ENDPOINT IS THE HOST'S BUG
--
-- A halted endpoint stays halted until CLEAR_FEATURE(ENDPOINT_HALT). It does
-- not clear on the next successful transfer, and a host that keeps issuing
-- tokens to it has not noticed -- which is worth counting SEPARATELY,
-- because it points at the driver rather than at the device.
--
-- THE TOGGLE CATCHES LOSS AND DUPLICATION AND CANNOT TELL THEM APART
--
-- A one-bit sequence number has only one wrong value, so "re-sent" and
-- "lost" produce the same symptom. That is a property of one bit, not a
-- defect here, and it is why chapter 24.3's scoreboard needs a real
-- identity.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
package usb_eph_pkg is
-- Response encodings. Named constants rather than an enumerated type,
-- because the port carries the two bits that appear on the analyser and
-- an enumeration would hide which value is which.
constant R_ACK : std_logic_vector(1 downto 0) := "00";
constant R_NAK : std_logic_vector(1 downto 0) := "01";
constant R_STALL : std_logic_vector(1 downto 0) := "10";
constant R_NYET : std_logic_vector(1 downto 0) := "11";
constant H_NONE : std_logic_vector(2 downto 0) := "000";
constant H_WEDGED : std_logic_vector(2 downto 0) := "001";
constant H_POLLED_HALTED : std_logic_vector(2 downto 0) := "010";
constant H_TOGGLE : std_logic_vector(2 downto 0) := "011";
constant H_CLEAR_UNHALTED : std_logic_vector(2 downto 0) := "100";
type run_array is array (natural range <>) of unsigned(7 downto 0);
end package;
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
use work.usb_eph_pkg.all;
entity usb_endpoint_health is
generic (
N_EP : integer := 4;
NAK_RUN_MAX : integer := 16
);
port (
clk : in std_logic;
rst_n : in std_logic;
xact_valid : in std_logic;
ep : in std_logic_vector(1 downto 0);
resp : in std_logic_vector(1 downto 0);
toggle : in std_logic;
clear_halt : in std_logic;
eot : in std_logic;
halted : out std_logic_vector(N_EP-1 downto 0);
nak_run : out std_logic_vector(7 downto 0);
worst_run : out std_logic_vector(7 downto 0);
exp_toggle : out std_logic;
err_pulse : out std_logic;
err_code : out std_logic_vector(2 downto 0);
stall_pulse : out std_logic;
n_ack : out unsigned(31 downto 0);
n_nak : out unsigned(31 downto 0);
n_stall : out unsigned(31 downto 0);
n_nyet : out unsigned(31 downto 0);
n_wedged : out unsigned(31 downto 0);
n_polled_halted : out unsigned(31 downto 0);
n_toggle : out unsigned(31 downto 0);
n_clear_unhalted : out unsigned(31 downto 0);
n_clear : out unsigned(31 downto 0)
);
end entity;
architecture rtl of usb_endpoint_health is
type bit_array is array (natural range <>) of std_logic;
signal halt_r : std_logic_vector(N_EP-1 downto 0);
signal run_r : run_array(0 to N_EP-1);
signal tog_r : bit_array(0 to N_EP-1);
signal wed_r : bit_array(0 to N_EP-1); -- reported wedged: report ONCE
signal worst_r : unsigned(7 downto 0);
signal ec_r : std_logic_vector(2 downto 0);
signal er_r : std_logic;
signal sp_r : std_logic;
signal ack_c, nak_c, stl_c, nyt_c, clr_c : unsigned(31 downto 0);
signal wedg_c, poll_c, tg_c, clru_c : unsigned(31 downto 0);
-- The endpoint index. Only two bits wide, so to_integer cannot overflow
-- INTEGER here -- but the habit is worth keeping: converting a full
-- 32-bit unsigned aborts the run rather than wrapping.
signal epi : integer range 0 to N_EP-1;
begin
epi <= to_integer(unsigned(ep));
halted <= halt_r;
nak_run <= std_logic_vector(run_r(epi));
worst_run <= std_logic_vector(worst_r);
exp_toggle <= tog_r(epi);
err_pulse <= er_r;
err_code <= ec_r;
stall_pulse <= sp_r;
n_ack <= ack_c;
n_nak <= nak_c;
n_stall <= stl_c;
n_nyet <= nyt_c;
n_wedged <= wedg_c;
n_polled_halted <= poll_c;
n_toggle <= tg_c;
n_clear_unhalted <= clru_c;
n_clear <= clr_c;
process (clk, rst_n)
variable e : integer range 0 to N_EP-1;
variable halt_v : std_logic_vector(N_EP-1 downto 0);
variable run_v : run_array(0 to N_EP-1);
variable tog_v : bit_array(0 to N_EP-1);
variable wed_v : bit_array(0 to N_EP-1);
variable wrst_v : unsigned(7 downto 0);
variable ec_v : std_logic_vector(2 downto 0);
variable er_v : std_logic;
variable sp_v : std_logic;
begin
if rst_n = '0' then
halt_r <= (others => '0');
run_r <= (others => (others => '0'));
tog_r <= (others => '0');
wed_r <= (others => '0');
worst_r <= (others => '0');
ec_r <= H_NONE;
er_r <= '0';
sp_r <= '0';
ack_c <= (others => '0');
nak_c <= (others => '0');
stl_c <= (others => '0');
nyt_c <= (others => '0');
clr_c <= (others => '0');
wedg_c <= (others => '0');
poll_c <= (others => '0');
tg_c <= (others => '0');
clru_c <= (others => '0');
elsif rising_edge(clk) then
e := to_integer(unsigned(ep));
halt_v := halt_r;
run_v := run_r;
tog_v := tog_r;
wed_v := wed_r;
wrst_v := worst_r;
ec_v := H_NONE;
er_v := '0';
sp_v := '0';
if eot = '1' then
-- Nothing: the counters are the report.
null;
elsif clear_halt = '1' then
-- ---- CLEAR_FEATURE(ENDPOINT_HALT). ----
if halt_v(e) = '0' then
-- Clearing a halt that was never set: the driver believes the
-- endpoint is halted and the device does not. Harmless on the
-- wire, and a disagreement worth knowing about before it becomes
-- a lockup.
er_v := '1';
ec_v := H_CLEAR_UNHALTED;
clru_c <= clru_c + 1;
end if;
halt_v(e) := '0';
-- A cleared endpoint restarts at DATA0. Chapter 21.1 makes the same
-- point about a bus reset: the toggle is part of the state that has
-- to be reset with everything else, and forgetting it costs exactly
-- one silently-dropped packet.
tog_v(e) := '0';
run_v(e) := (others => '0');
wed_v(e) := '0';
clr_c <= clr_c + 1;
elsif xact_valid = '1' then
if halt_v(e) = '1' then
-- ---- The host is still polling a HALTED endpoint. ----
er_v := '1';
ec_v := H_POLLED_HALTED;
poll_c <= poll_c + 1;
else
case resp is
when R_ACK =>
ack_c <= ack_c + 1;
if toggle /= tog_v(e) then
-- The sequence broke. WHICH WAY is not recoverable from one
-- bit -- see the note in the header.
er_v := '1';
ec_v := H_TOGGLE;
tg_c <= tg_c + 1;
end if;
tog_v(e) := not tog_v(e);
-- ---- A successful transfer resets the run. By the transfer,
-- ---- not by time.
run_v(e) := (others => '0');
wed_v(e) := '0';
when R_NAK =>
-- NOT an error. Flow control working.
nak_c <= nak_c + 1;
if run_v(e) /= x"FF" then
run_v(e) := run_v(e) + 1;
end if;
if run_v(e) > wrst_v then
wrst_v := run_v(e);
end if;
if (run_v(e) >= to_unsigned(NAK_RUN_MAX, 8)) and wed_v(e) = '0'
then
-- ---- The RUN is the error, not the NAK -- and it is
-- ---- reported ONCE. A stuck endpoint NAKs every
-- ---- microframe, and one report per NAK is the same flood
-- ---- that counting individual NAKs produced.
er_v := '1';
ec_v := H_WEDGED;
wedg_c <= wedg_c + 1;
wed_v(e) := '1';
end if;
when R_STALL =>
-- ---- STICKY. Until CLEAR_FEATURE, and nothing else. ----
stl_c <= stl_c + 1;
sp_v := '1';
halt_v(e) := '1';
run_v(e) := (others => '0');
when others =>
-- NYET: accepted, but not ready for the next one. Like a NAK
-- it is not an error; unlike a NAK the transfer did happen,
-- so it resets the run.
nyt_c <= nyt_c + 1;
run_v(e) := (others => '0');
wed_v(e) := '0';
end case;
end if;
end if;
halt_r <= halt_v;
run_r <= run_v;
tog_r <= tog_v;
wed_r <= wed_v;
worst_r <= wrst_v;
ec_r <= ec_v;
er_r <= er_v;
sp_r <= sp_v;
end if;
end process;
end architecture;11. Seeing the Wedge Form
The threshold crossing, the single report, and the re-arm:
Sixteen consecutive NAKs, one report
usb_endpoint_health — the run crosses the threshold once
10 cyclesAnd the halt, which behaves nothing like it:
A STALL, then four transactions that do not clear it
usb_endpoint_health — the halt is sticky
10 cycles12. The Testbenches
The oracle is a shadow model written from section 3 to section 6 rather than from the RTL, re-derived every cycle and compared against every output — seventeen checks per cycle, including the structural one that the four per-cause counters sum to the error total.
The exhaustive claim is over the situation:
endpoint 4 (0..3)
halted 2 (before the transaction)
response 4 (ACK / NAK / STALL / NYET)
toggle matches 2 (yes / no)
---
64 transaction situations
CLEAR_FEATURE x endpoint x halted
8 clear situations
---
72 and all 72 are REQUIREDThe halted axis is set up on the wire, by sending a STALL or a CLEAR_FEATURE, never by forcing a register. A state forced into the design is a state the design never proved it can reach, which is chapter 22.4's rule.
Eight phases, and the last two are the interesting ones:
| Phase | What it establishes |
|---|---|
| 1 | all 64 transaction situations, halt set up on the wire |
| 2 | CLEAR_FEATURE against both halt states, all four endpoints |
| 3 | the wedge boundary: silent at 15, fires at 16, silent at 56, re-arms after an ACK |
| 4 | one wedged endpoint among three healthy ones — per-endpoint independence |
| 5 | 64 correct toggles silent, one break reported, 32 more silent (resync) |
| 6 | a halt survives 50 transactions including ACKs; only CLEAR_FEATURE ends it |
| 7 | 40000 random transactions, mixed responses, mostly-correct toggles |
| 8 | 20000 random transactions with one sick endpoint at a time |
Verilog-2005 testbench
`timescale 1ns/1ps
// Testbench for usb_endpoint_health.
//
// The oracle is a SHADOW MODEL written from the chapter's rules rather than
// from the design's code: per-endpoint run lengths, a sticky halt, a
// report-once wedge latch, and an expected toggle. It is re-derived every
// cycle from the inputs and compared against every output, which is the only
// arrangement that catches a design that is self-consistently wrong.
//
// The exhaustive part is the SITUATION: every (endpoint, halted, response,
// toggle) combination, plus CLEAR_FEATURE against both a halted and an
// unhalted endpoint. 4x2x4x2 + 4x2 = 72, and all 72 are required to have
// been visited before the run is allowed to pass.
module tb_eh_v;
localparam integer N_EP = 4;
localparam integer NAK_RUN_MAX = 16;
localparam [1:0] R_ACK = 2'd0, R_NAK = 2'd1, R_STALL = 2'd2, R_NYET = 2'd3;
localparam [2:0] H_NONE = 3'd0,
H_WEDGED = 3'd1,
H_POLLED_HALTED = 3'd2,
H_TOGGLE = 3'd3,
H_CLEAR_UNHALTED = 3'd4;
reg clk = 1'b0, rst_n = 1'b0;
reg xact_valid = 1'b0, toggle = 1'b0, clear_halt = 1'b0, eot = 1'b0;
reg [1:0] ep = 2'd0, resp = 2'd0;
wire [N_EP-1:0] halted;
wire [7:0] nak_run, worst_run;
wire exp_toggle, err_pulse, stall_pulse;
wire [2:0] err_code;
wire [31:0] n_ack, n_nak, n_stall, n_nyet, n_wedged,
n_polled_halted, n_toggle, n_clear_unhalted, n_clear;
usb_endpoint_health #(.N_EP(N_EP), .NAK_RUN_MAX(NAK_RUN_MAX)) dut (
.clk(clk), .rst_n(rst_n),
.xact_valid(xact_valid), .ep(ep), .resp(resp), .toggle(toggle),
.clear_halt(clear_halt), .eot(eot),
.halted(halted), .nak_run(nak_run), .worst_run(worst_run),
.exp_toggle(exp_toggle), .err_pulse(err_pulse), .err_code(err_code),
.stall_pulse(stall_pulse),
.n_ack(n_ack), .n_nak(n_nak), .n_stall(n_stall), .n_nyet(n_nyet),
.n_wedged(n_wedged), .n_polled_halted(n_polled_halted),
.n_toggle(n_toggle), .n_clear_unhalted(n_clear_unhalted),
.n_clear(n_clear)
);
always #5 clk = ~clk;
// ---------------- the shadow model ----------------
reg [N_EP-1:0] m_halt;
reg [7:0] m_run [0:N_EP-1];
reg m_tog [0:N_EP-1];
reg m_wed [0:N_EP-1];
reg [7:0] m_worst;
reg m_err, m_stall_p;
reg [2:0] m_code;
reg [31:0] m_ack, m_nak, m_stl, m_nyt, m_wedg, m_poll, m_tgerr, m_clru, m_clr;
integer errors = 0, checks = 0, steps = 0;
integer k;
reg [0:0] reach [0:71];
integer n_reach;
task mark; input integer idx; begin reach[idx] = 1'b1; end endtask
task ck;
input [255:0] nm;
input [31:0] got, exp;
begin
checks = checks + 1;
if (got !== exp) begin
errors = errors + 1;
if (errors < 25)
$display("FAIL t=%0t step=%0d %0s got=%0d exp=%0d",
$time, steps, nm, got, exp);
end
end
endtask
reg m_halt_pre, m_tog_pre;
// One modelled cycle. Mirrors the rules, not the RTL.
task model_step;
integer j;
begin
m_err = 1'b0; m_code = H_NONE; m_stall_p = 1'b0;
if (eot) begin
// no state change
end else if (clear_halt) begin
if (!m_halt[ep]) begin
m_err = 1'b1; m_code = H_CLEAR_UNHALTED;
m_clru = m_clru + 1;
end
m_halt[ep] = 1'b0;
m_tog[ep] = 1'b0;
m_run[ep] = 8'd0;
m_wed[ep] = 1'b0;
m_clr = m_clr + 1;
end else if (xact_valid) begin
if (m_halt[ep]) begin
m_err = 1'b1; m_code = H_POLLED_HALTED;
m_poll = m_poll + 1;
end else if (resp == R_ACK) begin
m_ack = m_ack + 1;
if (toggle != m_tog[ep]) begin
m_err = 1'b1; m_code = H_TOGGLE;
m_tgerr = m_tgerr + 1;
end
m_tog[ep] = ~m_tog[ep];
m_run[ep] = 8'd0;
m_wed[ep] = 1'b0;
end else if (resp == R_NAK) begin
m_nak = m_nak + 1;
if (m_run[ep] != 8'hFF) m_run[ep] = m_run[ep] + 8'd1;
if (m_run[ep] > m_worst) m_worst = m_run[ep];
if ((m_run[ep] >= NAK_RUN_MAX) && !m_wed[ep]) begin
m_err = 1'b1; m_code = H_WEDGED;
m_wedg = m_wedg + 1;
m_wed[ep] = 1'b1;
end
end else if (resp == R_STALL) begin
m_stl = m_stl + 1;
m_stall_p = 1'b1;
m_halt[ep] = 1'b1;
m_run[ep] = 8'd0;
end else begin
m_nyt = m_nyt + 1;
m_run[ep] = 8'd0;
m_wed[ep] = 1'b0;
end
end
// --- record the situation reached, for the exhaustiveness proof ---
if (clear_halt)
mark(64 + ep*2 + (m_halt_pre ? 1 : 0));
else if (xact_valid)
mark(ep*16 + (m_halt_pre ? 8 : 0) + resp*2 + (m_tog_pre ? 1 : 0));
for (j = 0; j < N_EP; j = j + 1) ;
end
endtask
task check_out;
begin
ck("halted", {28'd0, halted}, {28'd0, m_halt});
ck("nak_run", {24'd0, nak_run}, {24'd0, m_run[ep]});
ck("worst_run", {24'd0, worst_run}, {24'd0, m_worst});
ck("exp_toggle", {31'd0, exp_toggle}, {31'd0, m_tog[ep]});
ck("err_pulse", {31'd0, err_pulse}, {31'd0, m_err});
ck("err_code", {29'd0, err_code}, {29'd0, m_code});
ck("stall_pulse", {31'd0, stall_pulse},{31'd0, m_stall_p});
ck("n_ack", n_ack, m_ack);
ck("n_nak", n_nak, m_nak);
ck("n_stall", n_stall, m_stl);
ck("n_nyet", n_nyet, m_nyt);
ck("n_wedged", n_wedged, m_wedg);
ck("n_polled_halted", n_polled_halted, m_poll);
ck("n_toggle", n_toggle, m_tgerr);
ck("n_clear_unhalted", n_clear_unhalted, m_clru);
ck("n_clear", n_clear, m_clr);
// ---- the structural invariant ----
// Every error is exactly one of the four causes, so the per-cause
// counters sum to the total. Chapter 23.4's rule: a report that
// cannot be decomposed cannot be trusted.
ck("sum", n_wedged + n_polled_halted + n_toggle + n_clear_unhalted,
m_wedg + m_poll + m_tgerr + m_clru);
end
endtask
task step;
begin
m_halt_pre = m_halt[ep];
m_tog_pre = (toggle == m_tog[ep]);
// The caller has already driven the inputs, so the model is stepped
// and the DUT clocked on the SAME edge. An extra @(negedge) here lets
// the intervening posedge sample the inputs a second time -- every
// transaction counted twice, and the shadow model is what says so.
model_step;
@(posedge clk);
#1;
steps = steps + 1;
check_out;
end
endtask
task xact; input [1:0] e; input [1:0] r; input tg;
begin
xact_valid = 1'b1; clear_halt = 1'b0; eot = 1'b0;
ep = e; resp = r; toggle = tg;
step;
end
endtask
task clr; input [1:0] e;
begin
xact_valid = 1'b0; clear_halt = 1'b1; eot = 1'b0;
ep = e;
step;
end
endtask
task nop;
begin
xact_valid = 1'b0; clear_halt = 1'b0; eot = 1'b0;
step;
end
endtask
// ---- OBSERVE an endpoint without driving it. ----
//
// `nak_run` and `exp_toggle` are indexed by the `ep` input, so pointing
// `ep` at an endpoint during an idle cycle reads that endpoint's state and
// changes nothing. A per-endpoint counter can only be proved per-endpoint
// by looking at the endpoints that are NOT being driven -- see phase 4.
task peek; input [1:0] e;
begin
xact_valid = 1'b0; clear_halt = 1'b0; eot = 1'b0;
ep = e;
step;
end
endtask
integer i, e, r, t, h, w;
integer base_wedged, base_poll, base_tg, base_clru;
initial begin
for (k = 0; k < 72; k = k + 1) reach[k] = 1'b0;
m_halt = {N_EP{1'b0}};
for (k = 0; k < N_EP; k = k + 1) begin
m_run[k] = 8'd0; m_tog[k] = 1'b0; m_wed[k] = 1'b0;
end
m_worst = 8'd0; m_err = 1'b0; m_code = H_NONE; m_stall_p = 1'b0;
m_ack=0; m_nak=0; m_stl=0; m_nyt=0;
m_wedg=0; m_poll=0; m_tgerr=0; m_clru=0; m_clr=0;
repeat (3) @(posedge clk);
rst_n = 1'b1;
@(negedge clk);
// ================= PHASE 1 -- the exhaustive situation sweep =========
//
// Every (endpoint, halted, response, toggle-matches) combination. The
// halted axis is set up on the wire, by STALLing or clearing, because a
// state forced into the design is a state the design never proved it can
// reach (chapter 22.4).
for (e = 0; e < N_EP; e = e + 1)
for (h = 0; h < 2; h = h + 1)
for (r = 0; r < 4; r = r + 1)
for (t = 0; t < 2; t = t + 1) begin
// bring endpoint e to the wanted halt state
if (h == 1 && !m_halt[e[1:0]]) xact(e[1:0], R_STALL, 1'b0);
if (h == 0 && m_halt[e[1:0]]) clr(e[1:0]);
// t==1 means "the toggle matches what the model expects"
xact(e[1:0], r[1:0], (t == 1) ? m_tog[e[1:0]] : ~m_tog[e[1:0]]);
// leave the endpoint clean for the next combination
if (m_halt[e[1:0]]) clr(e[1:0]);
end
// ================= PHASE 2 -- CLEAR against both halt states =========
for (e = 0; e < N_EP; e = e + 1) begin
clr(e[1:0]); // unhalted: must report
xact(e[1:0], R_STALL, 1'b0);
clr(e[1:0]); // halted: must NOT report
end
// ================= PHASE 3 -- the wedge boundary =====================
//
// NAK_RUN_MAX-1 NAKs must be silent; the NAK_RUN_MAX'th must fire; every
// NAK after it must be silent again (report ONCE); and an ACK must
// re-arm the latch so a second wedge is reported.
for (e = 0; e < N_EP; e = e + 1) begin
clr(e[1:0]); // known state, toggle at 0
base_wedged = m_wedg;
for (i = 0; i < NAK_RUN_MAX - 1; i = i + 1) xact(e[1:0], R_NAK, 1'b0);
if (m_wedg != base_wedged) begin
errors = errors + 1;
$display("FAIL early wedge at ep=%0d", e);
end
xact(e[1:0], R_NAK, 1'b0); // the threshold
if (m_wedg != base_wedged + 1) begin
errors = errors + 1;
$display("FAIL threshold wedge ep=%0d", e);
end
repeat (40) xact(e[1:0], R_NAK, 1'b0); // still ONE report
if (m_wedg != base_wedged + 1) begin
errors = errors + 1;
$display("FAIL repeated wedge reports ep=%0d", e);
end
xact(e[1:0], R_ACK, m_tog[e[1:0]]); // answered: run resets, latch arms
if (m_run[e[1:0]] != 0) begin
errors = errors + 1;
$display("FAIL run not reset by ACK ep=%0d", e);
end
repeat (NAK_RUN_MAX) xact(e[1:0], R_NAK, 1'b0);
if (m_wedg != base_wedged + 2) begin
errors = errors + 1;
$display("FAIL wedge not re-armed ep=%0d", e);
end
// A NYET also clears the run -- accepted, just not ready for more.
xact(e[1:0], R_NYET, 1'b0);
if (m_run[e[1:0]] != 0) begin
errors = errors + 1;
$display("FAIL run not reset by NYET ep=%0d", e);
end
end
// ================= PHASE 4 -- per-endpoint independence ==============
//
// The point of a PER-ENDPOINT run. Endpoint 0 NAKs forever while 1, 2
// and 3 answer normally. A single shared counter is reset by every one
// of those answers, so the wedged endpoint never reaches the threshold
// and the one real fault on the bus is the one thing not reported.
for (e = 0; e < N_EP; e = e + 1) clr(e[1:0]);
base_wedged = m_wedg;
for (i = 0; i < 3 * NAK_RUN_MAX; i = i + 1) begin
xact(2'd0, R_NAK, 1'b0);
// Read the three healthy endpoints BEFORE they answer. Without this,
// a run counter shared across the whole device is invisible here: the
// corruption lands on endpoints 1..3, and each one's own ACK wipes it
// before anything looks. The wedge COUNT on endpoint 0 is unchanged,
// so the phase written to catch a shared counter passes.
peek(2'd1); peek(2'd2); peek(2'd3);
xact(2'd1, R_ACK, m_tog[1]);
xact(2'd2, R_ACK, m_tog[2]);
xact(2'd3, R_NYET, 1'b0);
end
if (m_wedg != base_wedged + 1) begin
errors = errors + 1;
$display("FAIL interleaved wedge: got %0d expected %0d",
m_wedg - base_wedged, 1);
end
// ================= PHASE 5 -- the toggle sequence ====================
//
// A long correct alternation must be silent; one wrong toggle must fire
// exactly once; and the design must RESYNC afterwards rather than
// reporting every subsequent transfer.
for (e = 0; e < N_EP; e = e + 1) clr(e[1:0]);
base_tg = m_tgerr;
for (i = 0; i < 64; i = i + 1) xact(2'd1, R_ACK, m_tog[1]);
if (m_tgerr != base_tg) begin
errors = errors + 1;
$display("FAIL toggle false positive on a correct sequence");
end
xact(2'd1, R_ACK, ~m_tog[1]); // one break
if (m_tgerr != base_tg + 1) begin
errors = errors + 1;
$display("FAIL toggle break not reported");
end
for (i = 0; i < 32; i = i + 1) xact(2'd1, R_ACK, m_tog[1]);
if (m_tgerr != base_tg + 1) begin
errors = errors + 1;
$display("FAIL toggle did not resync (%0d extra)",
m_tgerr - base_tg - 1);
end
// ================= PHASE 6 -- the halt is STICKY =====================
//
// A STALL, then a long run of successful-looking traffic. Nothing about
// an ACK, a NYET or the passage of time clears a halt: only
// CLEAR_FEATURE does, and until it arrives every poll is the host's bug.
for (e = 0; e < N_EP; e = e + 1) clr(e[1:0]);
base_poll = m_poll;
xact(2'd2, R_STALL, 1'b0);
for (i = 0; i < 50; i = i + 1) xact(2'd2, (i[1:0]), m_tog[2]);
if (!m_halt[2]) begin
errors = errors + 1;
$display("FAIL halt did not stick");
end
if (m_poll != base_poll + 50) begin
errors = errors + 1;
$display("FAIL polled-halted count %0d expected 50",
m_poll - base_poll);
end
clr(2'd2);
if (m_halt[2]) begin
errors = errors + 1;
$display("FAIL CLEAR_FEATURE did not unhalt");
end
// ...and a cleared endpoint restarts at DATA0.
if (m_tog[2] !== 1'b0) begin
errors = errors + 1;
$display("FAIL toggle not reset by CLEAR_FEATURE");
end
// The two random phases are switchable, because a mutation score is
// only interesting once it is DECOMPOSED. A score that is identical
// with and without them was earned entirely by directed stimulus, and
// the random cycles -- however many there are -- bought nothing.
// Phases 1 and 2 already visit all 72 situations, so the exhaustiveness
// proof still holds with the random phases compiled out.
`ifndef DIRECTED_ONLY
// ================= PHASE 7 -- random =================================
//
// The response is drawn ONCE into a variable. Written as a chain of
// ternaries -- `($random%100 < 55) ? R_NAK : ($random%100 < 60) ? ...`
// -- each arm draws a NEW number, so the branches are independent 55%
// and 60% events rather than the nested distribution the code appears
// to describe. It still runs, still passes, and quietly delivers a mix
// nobody chose.
//
// The toggle is mostly CORRECT. A 50/50 toggle makes half of all ACKs a
// sequence error, which is not a bus and, more to the point, makes a
// dropped toggle check impossible to distinguish from a working one:
// when the error is the common case, the check that finds it is no
// longer doing any work.
for (i = 0; i < 40000; i = i + 1) begin
w = $unsigned($random) % 100;
e = $unsigned($random) % N_EP;
if (w < 14) begin
clr(e[1:0]);
end else if (w < 17) begin
nop;
end else begin
r = $unsigned($random) % 100;
t = ($unsigned($random) % 100) < 88; // 12% break the sequence
xact(e[1:0],
(r < 50) ? R_NAK : (r < 53) ? R_STALL :
(r < 63) ? R_NYET : R_ACK,
(t == 1) ? m_tog[e[1:0]] : ~m_tog[e[1:0]]);
end
end
// ================= PHASE 8 -- random, with a SICK endpoint ===========
//
// Phase 7 produced nine wedges and every one of them came from the
// directed phases. With four endpoints and a 50% NAK rate, sixteen
// consecutive NAKs on the SAME endpoint is a ~1-in-10^12 event: uniform
// random stimulus does not produce a wedge, because a wedge is
// SUSTAINED SILENCE and random traffic is the opposite of sustained.
//
// That is the same result as chapter 24.1's -- a liveness failure has
// to be constructed -- and it is why the fix is not "more random
// cycles". It is a stimulus that MODELS the fault: one endpoint at a
// time is sick, answering only occasionally, while the others behave.
// The sick one rotates, so every endpoint is both the fault and the
// background noise.
for (e = 0; e < N_EP; e = e + 1) clr(e[1:0]);
for (i = 0; i < 20000; i = i + 1) begin
if ((i % 201) == 0) h = $unsigned($random) % N_EP; // rotate the sick one
w = $unsigned($random) % 100;
e = $unsigned($random) % N_EP;
if (w < 3) begin
clr(e[1:0]);
end else if (e == h) begin
// The sick endpoint: NAKs almost always, answers once in a while.
r = $unsigned($random) % 100;
xact(e[1:0], (r < 94) ? R_NAK : R_ACK, m_tog[e[1:0]]);
end else begin
r = $unsigned($random) % 100;
t = ($unsigned($random) % 100) < 92;
xact(e[1:0],
(r < 20) ? R_NAK : (r < 22) ? R_STALL :
(r < 32) ? R_NYET : R_ACK,
(t == 1) ? m_tog[e[1:0]] : ~m_tog[e[1:0]]);
end
end
`endif
// ================= the exhaustiveness proof ==========================
n_reach = 0;
for (k = 0; k < 72; k = k + 1) n_reach = n_reach + reach[k];
if (n_reach != 72) begin
errors = errors + 1;
$display("FAIL situation reach %0d/72", n_reach);
for (k = 0; k < 72; k = k + 1)
if (!reach[k]) $display(" unreached situation %0d", k);
end
$display("steps=%0d checks=%0d reach=%0d/72 errors=%0d",
steps, checks, n_reach, errors);
$display("ack=%0d nak=%0d stall=%0d nyet=%0d clear=%0d",
n_ack, n_nak, n_stall, n_nyet, n_clear);
$display("wedged=%0d polled_halted=%0d toggle=%0d clear_unhalted=%0d",
n_wedged, n_polled_halted, n_toggle, n_clear_unhalted);
$display("%0s: %0d errors in %0d checks",
(errors == 0) ? "PASS" : "FAIL", errors, checks);
$finish;
end
endmoduleSystemVerilog testbench
`timescale 1ns/1ps
// Testbench for usb_endpoint_health (SystemVerilog).
//
// The oracle is a SHADOW MODEL written from the chapter's rules rather than
// from the design's code, and re-derived every cycle from the inputs. Every
// output is compared against it, which is the only arrangement that catches
// a design that is self-consistently wrong.
//
// The exhaustive part is the SITUATION: every (endpoint, halted, response,
// toggle-matches) combination plus CLEAR_FEATURE against both halt states.
// 4x2x4x2 + 4x2 = 72, and all 72 must have been visited before the run is
// allowed to pass.
module tb_eh_sv;
import usb_eph_pkg::*;
localparam int N_EP = 4;
localparam int NAK_RUN_MAX = 16;
logic clk = 1'b0, rst_n = 1'b0;
logic xact_valid = 1'b0, toggle = 1'b0, clear_halt = 1'b0, eot = 1'b0;
logic [1:0] ep = 2'd0;
resp_e resp = R_ACK;
logic [N_EP-1:0] halted;
logic [7:0] nak_run, worst_run;
logic exp_toggle, err_pulse, stall_pulse;
health_e err_code;
logic [31:0] n_ack, n_nak, n_stall, n_nyet, n_wedged,
n_polled_halted, n_toggle, n_clear_unhalted, n_clear;
usb_endpoint_health #(.N_EP(N_EP), .NAK_RUN_MAX(NAK_RUN_MAX)) dut (
.clk(clk), .rst_n(rst_n),
.xact_valid(xact_valid), .ep(ep), .resp(resp), .toggle(toggle),
.clear_halt(clear_halt), .eot(eot),
.halted(halted), .nak_run(nak_run), .worst_run(worst_run),
.exp_toggle(exp_toggle), .err_pulse(err_pulse), .err_code(err_code),
.stall_pulse(stall_pulse),
.n_ack(n_ack), .n_nak(n_nak), .n_stall(n_stall), .n_nyet(n_nyet),
.n_wedged(n_wedged), .n_polled_halted(n_polled_halted),
.n_toggle(n_toggle), .n_clear_unhalted(n_clear_unhalted),
.n_clear(n_clear)
);
always #5 clk = ~clk;
// ---------------- the shadow model ----------------
//
// Module scope, not a class. A class oracle is the UVM shape and the right
// one on a commercial simulator; Icarus reaches its limits on class
// properties quickly, and two of the ways it does so are worth knowing
// because neither produces a usable message:
//
// * an unpacked class array declared with a SIZE (`bit [7:0] run [N_EP]`)
// aborts the code generator in draw_class.c with no line number, while
// the identical descending RANGE (`run [N_EP-1:0]`) compiles;
// * a PACKED class property cannot be indexed inside a condition, though
// it can be indexed as an assignment target -- so the failure shows up
// on two lines out of six and reads like a typo.
//
// The UVM formulation of this same oracle is listed later in the chapter.
// What matters for correctness is that the model is written from the
// rules rather than from the RTL, and that is true either way.
logic [N_EP-1:0] m_halt;
logic [7:0] m_run [N_EP];
logic m_tog [N_EP];
logic m_wed [N_EP];
logic [7:0] m_worst;
logic m_err, m_stall_p;
health_e m_code;
int unsigned m_ack, m_nak, m_stl, m_nyt;
int unsigned m_wedg, m_poll, m_tgerr, m_clru, m_clr;
// One modelled cycle, derived from the chapter's rules rather than from
// the RTL. Written independently is the whole point: an oracle copied
// from the design agrees with the design's bugs.
task automatic model_step;
m_err = 1'b0; m_code = H_NONE; m_stall_p = 1'b0;
if (eot) begin
// no state change
end else if (clear_halt) begin
if (!m_halt[ep]) begin
m_err = 1'b1; m_code = H_CLEAR_UNHALTED; m_clru++;
end
m_halt[ep] = 1'b0;
m_tog[ep] = 1'b0;
m_run[ep] = 8'd0;
m_wed[ep] = 1'b0;
m_clr++;
end else if (xact_valid) begin
if (m_halt[ep]) begin
m_err = 1'b1; m_code = H_POLLED_HALTED; m_poll++;
end else if (resp == R_ACK) begin
m_ack++;
if (toggle != m_tog[ep]) begin
m_err = 1'b1; m_code = H_TOGGLE; m_tgerr++;
end
m_tog[ep] = ~m_tog[ep];
m_run[ep] = 8'd0;
m_wed[ep] = 1'b0;
end else if (resp == R_NAK) begin
m_nak++;
if (m_run[ep] != 8'hFF) m_run[ep] = m_run[ep] + 8'd1;
if (m_run[ep] > m_worst) m_worst = m_run[ep];
if ((m_run[ep] >= 8'(NAK_RUN_MAX)) && !m_wed[ep]) begin
m_err = 1'b1; m_code = H_WEDGED; m_wedg++; m_wed[ep] = 1'b1;
end
end else if (resp == R_STALL) begin
m_stl++; m_stall_p = 1'b1; m_halt[ep] = 1'b1; m_run[ep] = 8'd0;
end else begin
m_nyt++; m_run[ep] = 8'd0; m_wed[ep] = 1'b0;
end
end
endtask
int errors = 0, checks = 0, steps = 0;
bit reach [72];
int n_reach;
bit m_halt_pre, m_tog_pre;
task automatic ck(string nm, int unsigned got, int unsigned exp);
checks++;
if (got !== exp) begin
errors++;
if (errors < 25)
$display("FAIL t=%0t step=%0d %0s got=%0d exp=%0d",
$time, steps, nm, got, exp);
end
endtask
task automatic check_out;
ck("halted", halted, m_halt);
ck("nak_run", nak_run, m_run[ep]);
ck("worst_run", worst_run, m_worst);
ck("exp_toggle", exp_toggle, m_tog[ep]);
ck("err_pulse", err_pulse, m_err);
ck("err_code", err_code, m_code);
ck("stall_pulse", stall_pulse, m_stall_p);
ck("n_ack", n_ack, m_ack);
ck("n_nak", n_nak, m_nak);
ck("n_stall", n_stall, m_stl);
ck("n_nyet", n_nyet, m_nyt);
ck("n_wedged", n_wedged, m_wedg);
ck("n_polled_halted", n_polled_halted, m_poll);
ck("n_toggle", n_toggle, m_tgerr);
ck("n_clear_unhalted", n_clear_unhalted, m_clru);
ck("n_clear", n_clear, m_clr);
// Every error is exactly one of four causes, so the per-cause counters
// sum to the total. A report that cannot be decomposed cannot be
// trusted (chapter 23.4).
ck("sum", n_wedged + n_polled_halted + n_toggle + n_clear_unhalted,
m_wedg + m_poll + m_tgerr + m_clru);
endtask
task automatic step_one;
m_halt_pre = m_halt[ep];
m_tog_pre = (toggle == m_tog[ep]);
// The caller has already driven the inputs, so the model is stepped and
// the DUT clocked on the SAME edge. An extra @(negedge) here lets the
// intervening posedge sample the inputs a second time -- every
// transaction counted twice, and the shadow model is what says so.
model_step;
if (clear_halt)
reach[64 + ep*2 + (m_halt_pre ? 1 : 0)] = 1'b1;
else if (xact_valid)
reach[ep*16 + (m_halt_pre ? 8 : 0) + int'(resp)*2 + (m_tog_pre ? 1 : 0)]
= 1'b1;
@(posedge clk);
#1;
steps++;
check_out;
endtask
task automatic xact(bit [1:0] e, resp_e r, bit tg);
xact_valid = 1'b1; clear_halt = 1'b0; eot = 1'b0;
ep = e; resp = r; toggle = tg;
step_one;
endtask
task automatic clr(bit [1:0] e);
xact_valid = 1'b0; clear_halt = 1'b1; eot = 1'b0;
ep = e;
step_one;
endtask
task automatic nop;
xact_valid = 1'b0; clear_halt = 1'b0; eot = 1'b0;
step_one;
endtask
// ---- OBSERVE an endpoint without driving it. ----
//
// `nak_run` and `exp_toggle` are indexed by the `ep` input, so pointing
// `ep` at an endpoint during an idle cycle reads that endpoint's state and
// changes nothing. A per-endpoint counter can only be proved per-endpoint
// by looking at the endpoints that are NOT being driven -- see phase 4.
task automatic peek(bit [1:0] e);
xact_valid = 1'b0; clear_halt = 1'b0; eot = 1'b0;
ep = e;
step_one;
endtask
int i, e, r, t, h, w;
int base_wedged, base_poll, base_tg;
initial begin
foreach (reach[k]) reach[k] = 1'b0;
m_halt = '0; m_worst = 8'd0; m_err = 1'b0; m_code = H_NONE;
m_stall_p = 1'b0;
for (int j = 0; j < N_EP; j++) begin
m_run[j] = 8'd0; m_tog[j] = 1'b0; m_wed[j] = 1'b0;
end
m_ack = 0; m_nak = 0; m_stl = 0; m_nyt = 0;
m_wedg = 0; m_poll = 0; m_tgerr = 0; m_clru = 0; m_clr = 0;
repeat (3) @(posedge clk);
rst_n = 1'b1;
@(negedge clk);
// ================= PHASE 1 -- the exhaustive situation sweep =========
//
// The halted axis is set up ON THE WIRE, by STALLing or clearing,
// because a state forced into the design is a state the design never
// proved it can reach (chapter 22.4).
for (e = 0; e < N_EP; e++)
for (h = 0; h < 2; h++)
for (r = 0; r < 4; r++)
for (t = 0; t < 2; t++) begin
if (h == 1 && !m_halt[e[1:0]]) xact(e[1:0], R_STALL, 1'b0);
if (h == 0 && m_halt[e[1:0]]) clr(e[1:0]);
xact(e[1:0], resp_e'(r[1:0]),
(t == 1) ? m_tog[e[1:0]] : ~m_tog[e[1:0]]);
if (m_halt[e[1:0]]) clr(e[1:0]);
end
// ================= PHASE 2 -- CLEAR against both halt states =========
for (e = 0; e < N_EP; e++) begin
clr(e[1:0]); // unhalted: must report
xact(e[1:0], R_STALL, 1'b0);
clr(e[1:0]); // halted: must NOT report
end
// ================= PHASE 3 -- the wedge boundary =====================
//
// NAK_RUN_MAX-1 NAKs must be silent; the NAK_RUN_MAX'th must fire; every
// NAK after it must be silent again (report ONCE); and an ACK must
// re-arm the latch so a second wedge is reported.
for (e = 0; e < N_EP; e++) begin
clr(e[1:0]);
base_wedged = m_wedg;
for (i = 0; i < NAK_RUN_MAX - 1; i++) xact(e[1:0], R_NAK, 1'b0);
if (m_wedg != base_wedged) begin
errors++; $display("FAIL early wedge at ep=%0d", e);
end
xact(e[1:0], R_NAK, 1'b0);
if (m_wedg != base_wedged + 1) begin
errors++; $display("FAIL threshold wedge ep=%0d", e);
end
repeat (40) xact(e[1:0], R_NAK, 1'b0);
if (m_wedg != base_wedged + 1) begin
errors++; $display("FAIL repeated wedge reports ep=%0d", e);
end
xact(e[1:0], R_ACK, m_tog[e[1:0]]);
if (m_run[e[1:0]] != 0) begin
errors++; $display("FAIL run not reset by ACK ep=%0d", e);
end
repeat (NAK_RUN_MAX) xact(e[1:0], R_NAK, 1'b0);
if (m_wedg != base_wedged + 2) begin
errors++; $display("FAIL wedge not re-armed ep=%0d", e);
end
xact(e[1:0], R_NYET, 1'b0);
if (m_run[e[1:0]] != 0) begin
errors++; $display("FAIL run not reset by NYET ep=%0d", e);
end
end
// ================= PHASE 4 -- per-endpoint independence ==============
//
// Endpoint 0 NAKs forever while 1, 2 and 3 answer normally. A single
// shared counter is reset by every one of those answers, so the wedged
// endpoint never reaches the threshold and the one real fault on the
// bus is the one thing not reported.
for (e = 0; e < N_EP; e++) clr(e[1:0]);
base_wedged = m_wedg;
for (i = 0; i < 3 * NAK_RUN_MAX; i++) begin
xact(2'd0, R_NAK, 1'b0);
// Read the three healthy endpoints BEFORE they answer. Without this,
// a run counter shared across the whole device is invisible here: the
// corruption lands on endpoints 1..3, and each one's own ACK wipes it
// before anything looks. The wedge COUNT on endpoint 0 is unchanged,
// so the phase written to catch a shared counter passes.
peek(2'd1); peek(2'd2); peek(2'd3);
xact(2'd1, R_ACK, m_tog[1]);
xact(2'd2, R_ACK, m_tog[2]);
xact(2'd3, R_NYET, 1'b0);
end
if (m_wedg != base_wedged + 1) begin
errors++;
$display("FAIL interleaved wedge: got %0d expected 1",
m_wedg - base_wedged);
end
// ================= PHASE 5 -- the toggle sequence ====================
for (e = 0; e < N_EP; e++) clr(e[1:0]);
base_tg = m_tgerr;
for (i = 0; i < 64; i++) xact(2'd1, R_ACK, m_tog[1]);
if (m_tgerr != base_tg) begin
errors++; $display("FAIL toggle false positive on a correct sequence");
end
xact(2'd1, R_ACK, ~m_tog[1]);
if (m_tgerr != base_tg + 1) begin
errors++; $display("FAIL toggle break not reported");
end
for (i = 0; i < 32; i++) xact(2'd1, R_ACK, m_tog[1]);
if (m_tgerr != base_tg + 1) begin
errors++;
$display("FAIL toggle did not resync (%0d extra)",
m_tgerr - base_tg - 1);
end
// ================= PHASE 6 -- the halt is STICKY =====================
for (e = 0; e < N_EP; e++) clr(e[1:0]);
base_poll = m_poll;
xact(2'd2, R_STALL, 1'b0);
for (i = 0; i < 50; i++) xact(2'd2, resp_e'(i[1:0]), m_tog[2]);
if (!m_halt[2]) begin errors++; $display("FAIL halt did not stick"); end
if (m_poll != base_poll + 50) begin
errors++;
$display("FAIL polled-halted count %0d expected 50",
m_poll - base_poll);
end
clr(2'd2);
if (m_halt[2]) begin
errors++; $display("FAIL CLEAR_FEATURE did not unhalt");
end
if (m_tog[2] !== 1'b0) begin
errors++; $display("FAIL toggle not reset by CLEAR_FEATURE");
end
// ================= PHASE 7 -- random =================================
//
// The response is drawn ONCE into a variable. Written as a chain of
// ternaries each arm draws a NEW number, so the branches become
// independent events rather than the nested distribution the code
// appears to describe -- it still runs, still passes, and quietly
// delivers a mix nobody chose.
//
// The toggle is mostly CORRECT: a 50/50 toggle makes half of all ACKs a
// sequence error, and when the error is the common case the check that
// finds it is no longer doing any work.
for (i = 0; i < 40000; i++) begin
w = $unsigned($random) % 100;
e = $unsigned($random) % N_EP;
if (w < 14) begin
clr(e[1:0]);
end else if (w < 17) begin
nop;
end else begin
r = $unsigned($random) % 100;
t = (($unsigned($random) % 100) < 88) ? 1 : 0;
if (r < 50) xact(e[1:0], R_NAK,
(t == 1) ? m_tog[e[1:0]] : ~m_tog[e[1:0]]);
else if (r < 53) xact(e[1:0], R_STALL,
(t == 1) ? m_tog[e[1:0]] : ~m_tog[e[1:0]]);
else if (r < 63) xact(e[1:0], R_NYET,
(t == 1) ? m_tog[e[1:0]] : ~m_tog[e[1:0]]);
else xact(e[1:0], R_ACK,
(t == 1) ? m_tog[e[1:0]] : ~m_tog[e[1:0]]);
end
end
// ================= PHASE 8 -- random, with a SICK endpoint ===========
//
// Phase 7 produces wedges only where the directed phases put them. With
// four endpoints and a 50% NAK rate, sixteen consecutive NAKs on the
// SAME endpoint is a ~1-in-10^12 event: uniform random stimulus does
// not produce a wedge, because a wedge is SUSTAINED SILENCE and random
// traffic is the opposite of sustained. The fix is not more cycles; it
// is a stimulus that MODELS the fault -- one endpoint at a time is
// sick, answering only occasionally, and the sick one rotates.
for (e = 0; e < N_EP; e++) clr(e[1:0]);
for (i = 0; i < 20000; i++) begin
if ((i % 201) == 0) h = $unsigned($random) % N_EP;
w = $unsigned($random) % 100;
e = $unsigned($random) % N_EP;
if (w < 3) begin
clr(e[1:0]);
end else if (e == h) begin
r = $unsigned($random) % 100;
if (r < 94) xact(e[1:0], R_NAK, m_tog[e[1:0]]);
else xact(e[1:0], R_ACK, m_tog[e[1:0]]);
end else begin
r = $unsigned($random) % 100;
t = (($unsigned($random) % 100) < 92) ? 1 : 0;
if (r < 20) xact(e[1:0], R_NAK,
(t == 1) ? m_tog[e[1:0]] : ~m_tog[e[1:0]]);
else if (r < 22) xact(e[1:0], R_STALL,
(t == 1) ? m_tog[e[1:0]] : ~m_tog[e[1:0]]);
else if (r < 32) xact(e[1:0], R_NYET,
(t == 1) ? m_tog[e[1:0]] : ~m_tog[e[1:0]]);
else xact(e[1:0], R_ACK,
(t == 1) ? m_tog[e[1:0]] : ~m_tog[e[1:0]]);
end
end
// ================= the exhaustiveness proof ==========================
n_reach = 0;
foreach (reach[k]) n_reach += reach[k];
if (n_reach != 72) begin
errors++;
$display("FAIL situation reach %0d/72", n_reach);
foreach (reach[k]) if (!reach[k]) $display(" unreached situation %0d", k);
end
$display("steps=%0d checks=%0d reach=%0d/72 errors=%0d",
steps, checks, n_reach, errors);
$display("ack=%0d nak=%0d stall=%0d nyet=%0d clear=%0d",
n_ack, n_nak, n_stall, n_nyet, n_clear);
$display("wedged=%0d polled_halted=%0d toggle=%0d clear_unhalted=%0d",
n_wedged, n_polled_halted, n_toggle, n_clear_unhalted);
$display("%0s: %0d errors in %0d checks",
(errors == 0) ? "PASS" : "FAIL", errors, checks);
$finish;
end
endmoduleVHDL-2008 testbench
-- Testbench for usb_endpoint_health (VHDL-2008).
--
-- The oracle is a SHADOW MODEL held in process variables and written from
-- the chapter's rules rather than from the design's code. It is re-derived
-- every cycle from the inputs and compared against every output, which is
-- the only arrangement that catches a design that is self-consistently
-- wrong.
--
-- The exhaustive part is the SITUATION: every (endpoint, halted, response,
-- toggle-matches) combination plus CLEAR_FEATURE against both halt states.
-- 4x2x4x2 + 4x2 = 72, and all 72 must have been visited before the run is
-- allowed to pass.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
use std.textio.all;
use work.usb_eph_pkg.all;
entity tb_eh_vhdl is
end entity;
architecture sim of tb_eh_vhdl is
constant N_EP : integer := 4;
constant NAK_RUN_MAX : integer := 16;
signal clk : std_logic := '0';
signal rst_n : std_logic := '0';
signal xact_valid : std_logic := '0';
signal ep : std_logic_vector(1 downto 0) := "00";
signal resp : std_logic_vector(1 downto 0) := "00";
signal toggle : std_logic := '0';
signal clear_halt : std_logic := '0';
signal eot : std_logic := '0';
signal halted : std_logic_vector(N_EP-1 downto 0);
signal nak_run : std_logic_vector(7 downto 0);
signal worst_run : std_logic_vector(7 downto 0);
signal exp_toggle : std_logic;
signal err_pulse : std_logic;
signal err_code : std_logic_vector(2 downto 0);
signal stall_pulse : std_logic;
signal n_ack, n_nak, n_stall, n_nyet : unsigned(31 downto 0);
signal n_wedged, n_polled_halted : unsigned(31 downto 0);
signal n_toggle, n_clear_unhalted, n_clear : unsigned(31 downto 0);
signal done : boolean := false;
type bit_array is array (natural range <>) of std_logic;
type int_array is array (natural range <>) of integer;
begin
clk <= not clk after 5 ns when not done else '0';
dut : entity work.usb_endpoint_health
generic map (N_EP => N_EP, NAK_RUN_MAX => NAK_RUN_MAX)
port map (
clk => clk, rst_n => rst_n,
xact_valid => xact_valid, ep => ep, resp => resp, toggle => toggle,
clear_halt => clear_halt, eot => eot,
halted => halted, nak_run => nak_run, worst_run => worst_run,
exp_toggle => exp_toggle, err_pulse => err_pulse,
err_code => err_code, stall_pulse => stall_pulse,
n_ack => n_ack, n_nak => n_nak, n_stall => n_stall, n_nyet => n_nyet,
n_wedged => n_wedged, n_polled_halted => n_polled_halted,
n_toggle => n_toggle, n_clear_unhalted => n_clear_unhalted,
n_clear => n_clear
);
stim : process
-- ---------------- the shadow model ----------------
variable m_halt : std_logic_vector(N_EP-1 downto 0) := (others => '0');
variable m_run : run_array(0 to N_EP-1) := (others => (others => '0'));
variable m_tog : bit_array(0 to N_EP-1) := (others => '0');
variable m_wed : bit_array(0 to N_EP-1) := (others => '0');
variable m_worst : unsigned(7 downto 0) := (others => '0');
variable m_err : std_logic := '0';
variable m_sp : std_logic := '0';
variable m_code : std_logic_vector(2 downto 0) := H_NONE;
variable c_ack, c_nak, c_stl, c_nyt : integer := 0;
variable c_wedg, c_poll, c_tg, c_clru : integer := 0;
variable c_clr : integer := 0;
variable errors, checks, steps : integer := 0;
variable reach : int_array(0 to 71) := (others => 0);
variable n_reach : integer := 0;
variable halt_pre, tog_pre : std_logic := '0';
variable e_v, i_v, r_v, t_v, h_v, w_v, k_v : integer := 0;
variable base_w, base_p, base_t : integer := 0;
variable ln : line;
-- A deterministic LFSR, so a rerun reproduces exactly the same traffic.
variable lfsr : unsigned(31 downto 0) := x"D35C0FFE";
impure function rnd32 return unsigned is
begin
lfsr := lfsr(30 downto 0) &
(lfsr(31) xor lfsr(21) xor lfsr(1) xor lfsr(0));
return lfsr;
end function;
-- Only the low 30 bits are converted: a full 32-bit unsigned does not
-- fit in VHDL's INTEGER, and to_integer aborts the run rather than
-- wrapping.
impure function rnd_nat return integer is
variable u : unsigned(31 downto 0);
begin
u := rnd32;
return to_integer(u(29 downto 0));
end function;
procedure ck (nm : string; got, exp : integer) is
begin
checks := checks + 1;
if got /= exp then
errors := errors + 1;
if errors < 25 then
write(ln, string'("FAIL step=") & integer'image(steps) &
" " & nm & " got=" & integer'image(got) &
" exp=" & integer'image(exp));
writeline(output, ln);
end if;
end if;
end procedure;
function sl2i (s : std_logic) return integer is
begin
if s = '1' then return 1; else return 0; end if;
end function;
-- One modelled cycle, derived from the chapter's rules rather than
-- from the RTL. Written independently is the whole point: an oracle
-- copied from the design agrees with the design's bugs.
procedure model_step is
variable e : integer;
begin
e := to_integer(unsigned(ep));
m_err := '0'; m_code := H_NONE; m_sp := '0';
if eot = '1' then
null;
elsif clear_halt = '1' then
if m_halt(e) = '0' then
m_err := '1'; m_code := H_CLEAR_UNHALTED; c_clru := c_clru + 1;
end if;
m_halt(e) := '0';
m_tog(e) := '0';
m_run(e) := (others => '0');
m_wed(e) := '0';
c_clr := c_clr + 1;
elsif xact_valid = '1' then
if m_halt(e) = '1' then
m_err := '1'; m_code := H_POLLED_HALTED; c_poll := c_poll + 1;
elsif resp = R_ACK then
c_ack := c_ack + 1;
if toggle /= m_tog(e) then
m_err := '1'; m_code := H_TOGGLE; c_tg := c_tg + 1;
end if;
m_tog(e) := not m_tog(e);
m_run(e) := (others => '0');
m_wed(e) := '0';
elsif resp = R_NAK then
c_nak := c_nak + 1;
if m_run(e) /= x"FF" then m_run(e) := m_run(e) + 1; end if;
if m_run(e) > m_worst then m_worst := m_run(e); end if;
if (m_run(e) >= to_unsigned(NAK_RUN_MAX, 8)) and m_wed(e) = '0' then
m_err := '1'; m_code := H_WEDGED;
c_wedg := c_wedg + 1; m_wed(e) := '1';
end if;
elsif resp = R_STALL then
c_stl := c_stl + 1; m_sp := '1';
m_halt(e) := '1'; m_run(e) := (others => '0');
else
c_nyt := c_nyt + 1;
m_run(e) := (others => '0'); m_wed(e) := '0';
end if;
end if;
end procedure;
procedure check_out is
variable e : integer;
begin
e := to_integer(unsigned(ep));
ck("halted", to_integer(unsigned(halted)),
to_integer(unsigned(m_halt)));
ck("nak_run", to_integer(unsigned(nak_run)), to_integer(m_run(e)));
ck("worst_run", to_integer(unsigned(worst_run)), to_integer(m_worst));
ck("exp_toggle", sl2i(exp_toggle), sl2i(m_tog(e)));
ck("err_pulse", sl2i(err_pulse), sl2i(m_err));
ck("err_code", to_integer(unsigned(err_code)),
to_integer(unsigned(m_code)));
ck("stall_pulse", sl2i(stall_pulse), sl2i(m_sp));
ck("n_ack", to_integer(n_ack), c_ack);
ck("n_nak", to_integer(n_nak), c_nak);
ck("n_stall", to_integer(n_stall), c_stl);
ck("n_nyet", to_integer(n_nyet), c_nyt);
ck("n_wedged", to_integer(n_wedged), c_wedg);
ck("n_polled_halted", to_integer(n_polled_halted), c_poll);
ck("n_toggle", to_integer(n_toggle), c_tg);
ck("n_clear_unhalted", to_integer(n_clear_unhalted), c_clru);
ck("n_clear", to_integer(n_clear), c_clr);
-- Every error is exactly one of four causes, so the per-cause
-- counters sum to the total. A report that cannot be decomposed
-- cannot be trusted (chapter 23.4).
ck("sum", to_integer(n_wedged) + to_integer(n_polled_halted) +
to_integer(n_toggle) + to_integer(n_clear_unhalted),
c_wedg + c_poll + c_tg + c_clru);
end procedure;
procedure step_one is
variable e, idx : integer;
begin
e := to_integer(unsigned(ep));
halt_pre := m_halt(e);
if toggle = m_tog(e) then tog_pre := '1'; else tog_pre := '0'; end if;
model_step;
if clear_halt = '1' then
idx := 64 + e*2 + sl2i(halt_pre);
reach(idx) := 1;
elsif xact_valid = '1' then
idx := e*16 + sl2i(halt_pre)*8
+ to_integer(unsigned(resp))*2 + sl2i(tog_pre);
reach(idx) := 1;
end if;
wait until rising_edge(clk);
wait for 1 ns;
steps := steps + 1;
check_out;
end procedure;
procedure xact (e : integer; r : std_logic_vector(1 downto 0);
tg : std_logic) is
begin
xact_valid <= '1'; clear_halt <= '0'; eot <= '0';
ep <= std_logic_vector(to_unsigned(e, 2));
resp <= r; toggle <= tg;
wait for 0 ns;
step_one;
end procedure;
procedure clr (e : integer) is
begin
xact_valid <= '0'; clear_halt <= '1'; eot <= '0';
ep <= std_logic_vector(to_unsigned(e, 2));
wait for 0 ns;
step_one;
end procedure;
procedure nop is
begin
xact_valid <= '0'; clear_halt <= '0'; eot <= '0';
wait for 0 ns;
step_one;
end procedure;
-- ---- OBSERVE an endpoint without driving it. ----
--
-- nak_run and exp_toggle are indexed by the `ep` input, so pointing ep
-- at an endpoint during an idle cycle reads that endpoint's state and
-- changes nothing. A per-endpoint counter can only be proved
-- per-endpoint by looking at the endpoints that are NOT being driven --
-- see phase 4.
procedure peek (e : integer) is
begin
xact_valid <= '0'; clear_halt <= '0'; eot <= '0';
ep <= std_logic_vector(to_unsigned(e, 2));
wait for 0 ns;
step_one;
end procedure;
begin
wait until rising_edge(clk);
wait until rising_edge(clk);
wait until rising_edge(clk);
rst_n <= '1';
wait for 1 ns;
-- ================= PHASE 1 -- the exhaustive situation sweep =========
--
-- The halted axis is set up ON THE WIRE, by STALLing or clearing,
-- because a state forced into the design is a state the design never
-- proved it can reach (chapter 22.4).
for e in 0 to N_EP-1 loop
for h in 0 to 1 loop
for r in 0 to 3 loop
for t in 0 to 1 loop
if h = 1 and m_halt(e) = '0' then xact(e, R_STALL, '0'); end if;
if h = 0 and m_halt(e) = '1' then clr(e); end if;
if t = 1 then
xact(e, std_logic_vector(to_unsigned(r, 2)), m_tog(e));
else
xact(e, std_logic_vector(to_unsigned(r, 2)), not m_tog(e));
end if;
if m_halt(e) = '1' then clr(e); end if;
end loop;
end loop;
end loop;
end loop;
-- ================= PHASE 2 -- CLEAR against both halt states =========
for e in 0 to N_EP-1 loop
clr(e); -- unhalted: must report
xact(e, R_STALL, '0');
clr(e); -- halted: must NOT report
end loop;
-- ================= PHASE 3 -- the wedge boundary =====================
--
-- NAK_RUN_MAX-1 NAKs must be silent; the NAK_RUN_MAX'th must fire; every
-- NAK after it must be silent again (report ONCE); and an ACK must
-- re-arm the latch so a second wedge is reported.
for e in 0 to N_EP-1 loop
clr(e);
base_w := c_wedg;
for i in 0 to NAK_RUN_MAX-2 loop xact(e, R_NAK, '0'); end loop;
if c_wedg /= base_w then
errors := errors + 1;
write(ln, string'("FAIL early wedge ep=") & integer'image(e));
writeline(output, ln);
end if;
xact(e, R_NAK, '0');
if c_wedg /= base_w + 1 then
errors := errors + 1;
write(ln, string'("FAIL threshold wedge ep=") & integer'image(e));
writeline(output, ln);
end if;
for i in 0 to 39 loop xact(e, R_NAK, '0'); end loop;
if c_wedg /= base_w + 1 then
errors := errors + 1;
write(ln, string'("FAIL repeated wedge reports ep=")
& integer'image(e));
writeline(output, ln);
end if;
xact(e, R_ACK, m_tog(e));
if m_run(e) /= x"00" then
errors := errors + 1;
write(ln, string'("FAIL run not reset by ACK ep=") & integer'image(e));
writeline(output, ln);
end if;
for i in 0 to NAK_RUN_MAX-1 loop xact(e, R_NAK, '0'); end loop;
if c_wedg /= base_w + 2 then
errors := errors + 1;
write(ln, string'("FAIL wedge not re-armed ep=") & integer'image(e));
writeline(output, ln);
end if;
xact(e, R_NYET, '0');
if m_run(e) /= x"00" then
errors := errors + 1;
write(ln, string'("FAIL run not reset by NYET ep=")
& integer'image(e));
writeline(output, ln);
end if;
end loop;
-- ================= PHASE 4 -- per-endpoint independence ==============
--
-- Endpoint 0 NAKs forever while 1, 2 and 3 answer normally. A single
-- shared counter is reset by every one of those answers, so the wedged
-- endpoint never reaches the threshold and the one real fault on the
-- bus is the one thing not reported.
for e in 0 to N_EP-1 loop clr(e); end loop;
base_w := c_wedg;
for i in 0 to 3*NAK_RUN_MAX-1 loop
xact(0, R_NAK, '0');
-- Read the three healthy endpoints BEFORE they answer. Without this,
-- a run counter shared across the whole device is invisible here: the
-- corruption lands on endpoints 1..3, and each one's own ACK wipes it
-- before anything looks. The wedge COUNT on endpoint 0 is unchanged,
-- so the phase written to catch a shared counter passes.
peek(1); peek(2); peek(3);
xact(1, R_ACK, m_tog(1));
xact(2, R_ACK, m_tog(2));
xact(3, R_NYET, '0');
end loop;
if c_wedg /= base_w + 1 then
errors := errors + 1;
write(ln, string'("FAIL interleaved wedge: got ")
& integer'image(c_wedg - base_w) & " expected 1");
writeline(output, ln);
end if;
-- ================= PHASE 5 -- the toggle sequence ====================
for e in 0 to N_EP-1 loop clr(e); end loop;
base_t := c_tg;
for i in 0 to 63 loop xact(1, R_ACK, m_tog(1)); end loop;
if c_tg /= base_t then
errors := errors + 1;
write(ln, string'("FAIL toggle false positive on a correct sequence"));
writeline(output, ln);
end if;
xact(1, R_ACK, not m_tog(1));
if c_tg /= base_t + 1 then
errors := errors + 1;
write(ln, string'("FAIL toggle break not reported"));
writeline(output, ln);
end if;
for i in 0 to 31 loop xact(1, R_ACK, m_tog(1)); end loop;
if c_tg /= base_t + 1 then
errors := errors + 1;
write(ln, string'("FAIL toggle did not resync"));
writeline(output, ln);
end if;
-- ================= PHASE 6 -- the halt is STICKY =====================
for e in 0 to N_EP-1 loop clr(e); end loop;
base_p := c_poll;
xact(2, R_STALL, '0');
for i in 0 to 49 loop
xact(2, std_logic_vector(to_unsigned(i mod 4, 2)), m_tog(2));
end loop;
if m_halt(2) /= '1' then
errors := errors + 1;
write(ln, string'("FAIL halt did not stick")); writeline(output, ln);
end if;
if c_poll /= base_p + 50 then
errors := errors + 1;
write(ln, string'("FAIL polled-halted count ")
& integer'image(c_poll - base_p) & " expected 50");
writeline(output, ln);
end if;
clr(2);
if m_halt(2) /= '0' then
errors := errors + 1;
write(ln, string'("FAIL CLEAR_FEATURE did not unhalt"));
writeline(output, ln);
end if;
if m_tog(2) /= '0' then
errors := errors + 1;
write(ln, string'("FAIL toggle not reset by CLEAR_FEATURE"));
writeline(output, ln);
end if;
-- ================= PHASE 7 -- random =================================
--
-- The toggle is mostly CORRECT. A 50/50 toggle makes half of all ACKs a
-- sequence error, which is not a bus and, more to the point, makes a
-- dropped toggle check impossible to distinguish from a working one:
-- when the error is the common case, the check that finds it is no
-- longer doing any work.
for i in 0 to 39999 loop
w_v := rnd_nat mod 100;
e_v := rnd_nat mod N_EP;
if w_v < 14 then
clr(e_v);
elsif w_v < 17 then
nop;
else
r_v := rnd_nat mod 100;
if (rnd_nat mod 100) < 88 then t_v := 1; else t_v := 0; end if;
if r_v < 50 then
if t_v = 1 then xact(e_v, R_NAK, m_tog(e_v));
else xact(e_v, R_NAK, not m_tog(e_v)); end if;
elsif r_v < 53 then
if t_v = 1 then xact(e_v, R_STALL, m_tog(e_v));
else xact(e_v, R_STALL, not m_tog(e_v)); end if;
elsif r_v < 63 then
if t_v = 1 then xact(e_v, R_NYET, m_tog(e_v));
else xact(e_v, R_NYET, not m_tog(e_v)); end if;
else
if t_v = 1 then xact(e_v, R_ACK, m_tog(e_v));
else xact(e_v, R_ACK, not m_tog(e_v)); end if;
end if;
end if;
end loop;
-- ================= PHASE 8 -- random, with a SICK endpoint ===========
--
-- Phase 7 produces wedges only where the directed phases put them. With
-- four endpoints and a 50% NAK rate, sixteen consecutive NAKs on the
-- SAME endpoint is a ~1-in-10^12 event: uniform random stimulus does
-- not produce a wedge, because a wedge is SUSTAINED SILENCE and random
-- traffic is the opposite of sustained. The fix is not more cycles; it
-- is a stimulus that MODELS the fault -- one endpoint at a time is
-- sick, answering only occasionally, and the sick one rotates.
for e in 0 to N_EP-1 loop clr(e); end loop;
for i in 0 to 19999 loop
if (i mod 201) = 0 then h_v := rnd_nat mod N_EP; end if;
w_v := rnd_nat mod 100;
e_v := rnd_nat mod N_EP;
if w_v < 3 then
clr(e_v);
elsif e_v = h_v then
r_v := rnd_nat mod 100;
if r_v < 94 then xact(e_v, R_NAK, m_tog(e_v));
else xact(e_v, R_ACK, m_tog(e_v)); end if;
else
r_v := rnd_nat mod 100;
if (rnd_nat mod 100) < 92 then t_v := 1; else t_v := 0; end if;
if r_v < 20 then
if t_v = 1 then xact(e_v, R_NAK, m_tog(e_v));
else xact(e_v, R_NAK, not m_tog(e_v)); end if;
elsif r_v < 22 then
if t_v = 1 then xact(e_v, R_STALL, m_tog(e_v));
else xact(e_v, R_STALL, not m_tog(e_v)); end if;
elsif r_v < 32 then
if t_v = 1 then xact(e_v, R_NYET, m_tog(e_v));
else xact(e_v, R_NYET, not m_tog(e_v)); end if;
else
if t_v = 1 then xact(e_v, R_ACK, m_tog(e_v));
else xact(e_v, R_ACK, not m_tog(e_v)); end if;
end if;
end if;
end loop;
-- ================= the exhaustiveness proof ==========================
n_reach := 0;
for k in 0 to 71 loop n_reach := n_reach + reach(k); end loop;
if n_reach /= 72 then
errors := errors + 1;
write(ln, string'("FAIL situation reach ") & integer'image(n_reach)
& "/72");
writeline(output, ln);
for k in 0 to 71 loop
if reach(k) = 0 then
write(ln, string'(" unreached situation ") & integer'image(k));
writeline(output, ln);
end if;
end loop;
end if;
write(ln, string'("steps=") & integer'image(steps)
& " checks=" & integer'image(checks)
& " reach=" & integer'image(n_reach) & "/72"
& " errors=" & integer'image(errors));
writeline(output, ln);
write(ln, string'("ack=") & integer'image(to_integer(n_ack))
& " nak=" & integer'image(to_integer(n_nak))
& " stall=" & integer'image(to_integer(n_stall))
& " nyet=" & integer'image(to_integer(n_nyet))
& " clear=" & integer'image(to_integer(n_clear)));
writeline(output, ln);
write(ln, string'("wedged=") & integer'image(to_integer(n_wedged))
& " polled_halted="
& integer'image(to_integer(n_polled_halted))
& " toggle=" & integer'image(to_integer(n_toggle))
& " clear_unhalted="
& integer'image(to_integer(n_clear_unhalted)));
writeline(output, ln);
if errors = 0 then
write(ln, string'("PASS: 0 errors in ") & integer'image(checks)
& " checks");
else
write(ln, string'("FAIL: ") & integer'image(errors) & " errors in " &
integer'image(checks) & " checks");
end if;
writeline(output, ln);
done <= true;
wait;
end process;
end architecture;13. Why a Random Testbench Cannot Find a Wedge
Phase 7 runs 40000 random transactions across four endpoints with a 50% NAK rate. It produces, reliably, zero wedges.
The arithmetic is not close. A wedge needs sixteen consecutive NAKs on the same endpoint; each transaction picks one endpoint in four and NAKs with probability one half, so the chance of any given run of sixteen is about (1/8)^16, or roughly one in 10^14. Forty thousand transactions do not make a dent in that, and neither would forty million.
Phase 8 produces 80 wedges in the Verilog run against phase 3 and 4's 9. It is also the only phase in which a wedge ever coincides with ordinary traffic on the other three endpoints, which is what section 4 is about.
14. Exhaustive Verification
| Measure | Verilog | SystemVerilog | VHDL |
|---|---|---|---|
| situations reached | 72 / 72 | 72 / 72 | 72 / 72 |
| Steps | 60949 | 60949 | 60949 |
| Checks executed | 1036133 | 1036133 | 1036133 |
| ACKs | 17252 | 17252 | 18215 |
| NAKs | 20152 | 20152 | 21135 |
| STALLs | 1032 | 1032 | 1027 |
| NYETs | 3802 | 3802 | 3981 |
| CLEAR_FEATUREs | 6291 | 6291 | 6303 |
| wedged | 89 | 89 | 98 |
| polled while halted | 11045 | 11045 | 8925 |
| toggle breaks | 1754 | 1754 | 1481 |
| cleared while unhalted | 5261 | 5261 | 5278 |
| Result | PASS | PASS | PASS |
The number that carries the most weight is 72 / 72. It is not "we tried a stalled endpoint and a busy one"; it is every response, on every endpoint, in both halt states, with the toggle both right and wrong, plus CLEAR_FEATURE against an endpoint that is halted and one that is not.
15. Mutation Testing
| # | Mutation | Verilog | SysVer | VHDL |
|---|---|---|---|---|
| Q3 | a successful transfer clears the halt | 552279 | 552279 | 544031 |
| Q7 | an ACK does not reset the run | 197580 | 197580 | 200638 |
| Q1 | every NAK is reported as an error | 162020 | 162020 | 163968 |
| Q4 | a token to a halted endpoint is not reported | 143966 | 143966 | 139726 |
| Q5 | one run counter for the device, not per endpoint | 141225 | 141225 | 165463 |
| Q6 | the toggle check is dropped | 125406 | 125406 | 124860 |
| Q2 | the report-once latch is dropped | 123382 | 123382 | 123538 |
| — | unmutated baseline | 0 | 0 | 0 |
All seven die in all three languages.
Q3 is an order of magnitude above the rest because a halt that clears itself does not produce a wrong report — it produces a permanently divergent state. From the first ACK-while-halted onwards, the design and the model disagree about the halt vector on that endpoint, and every subsequent transaction to it takes the wrong branch. Nothing resynchronises them.
Q1 and Q2 are both floods, and Q2 is the subtler one. Q1 reports every NAK; Q2 reports only NAKs on an endpoint that is already wedged. Q1 is obvious the moment anybody looks at the log. Q2 looks correct until a device stays broken for a while, and then emits one line per microframe about a fault that was already reported.
Directed against random
Every score decomposes. The Verilog testbench compiled with -DDIRECTED_ONLY drops phases 7 and 8 and keeps the 72/72 proof, because phases 1 and 2 reach every situation on their own:
| # | All phases | Directed only | Random |
|---|---|---|---|
| Q1 | 162020 | 2556 | 159464 |
| Q2 | 123382 | 1944 | 121438 |
| Q3 | 552279 | 2139 | 550140 |
| Q4 | 143966 | 2032 | 141934 |
| Q5 | 141225 | 144 | 141081 |
| Q6 | 125406 | 1900 | 123506 |
| Q7 | 197580 | 943 | 196637 |
Every mutation is killed by directed stimulus alone. Nothing in this suite depends on getting lucky — which is the property worth checking, because a mutation killed only by the random phases is a mutation that a different seed may let through.
16. The Phase That Was Written to Catch Q5 and Did Not
Q5's directed column was 0 in the first version of this testbench, and that number is the most useful thing in the chapter.
Phase 4 exists specifically to catch a device-wide run counter. It NAKs endpoint 0 forever while endpoints 1, 2 and 3 answer normally, and then checks that endpoint 0 still reports exactly one wedge. It passed against the mutant.
Here is why. Q5 makes a NAK on any endpoint write the new run value to all four counters. Endpoint 0's own run is unaffected — it is still counting its own NAKs — so endpoint 0 still wedges exactly once and the check still passes. The corruption lands on endpoints 1, 2 and 3. And the very next thing the phase does is send each of them an ACK, which resets that endpoint's counter to zero:
xact(0, NAK) -> run[0..3] all corrupted
xact(1, ACK) -> run[1] reset. Corruption gone, unobserved.
xact(2, ACK) -> run[2] reset. Corruption gone, unobserved.
xact(3, NYET) -> run[3] reset. Corruption gone, unobserved.The testbench wrote the wrong value into three registers and then cleared all three before reading any of them.
And Q3 was dead code before it was a mutation
Q3's first formulation cleared the halt inside the R_ACK arm. It scored 0 in all three languages, which looks exactly like an escape.
It was not. The R_ACK arm is guarded by !halt_r[ep] — the design only classifies a response when the endpoint is not halted — so clearing a halt there clears a halt that is already clear. The mutation could not change behaviour under any stimulus.
Moved into the halted branch — where a successful-looking response really does clear a halt that is really set — the same idea scores 552279.
And the VHDL column of Q1 was a different mutation
Q1 first read 42074 in VHDL against 161732 in the other two. A 4x gap looks like a VHDL testbench gap.
It was a fidelity failure in the mutant. Verilog and SystemVerilog drive n_wedged from ec_n in the sequential block, so an injection that sets the code corrupts the counter for free. VHDL increments its counters at the point of decision, so the same injection set only the pulse — a strictly weaker mutation, scoring only the two pulse checks per NAK instead of those plus a permanently wrong counter.
Verilog 2 x 20152 NAKs + 2 x ~60000 cycles of a wrong counter
= ~160000 measured 161732
VHDL 2 x 21135 NAKs + 0
= ~42270 measured 42074Adding wedg_c <= wedg_c + 1; to the VHDL injection brought it to 163968. The lesson is chapter 24.2's in a new costume: divide the column by the count of the events it depends on before believing it says anything about the testbench.
17. Debugging Walkthrough: The Endpoint That Works Until You Plug In a Camera
The report. A USB audio device works perfectly on its own. Plug a webcam into the same hub and the audio device's bulk configuration endpoint stops responding — but only sometimes, and the vendor's own monitoring tool reports the device as healthy throughout.
Step 1 — what does the tool actually measure? It reports a NAK rate. Under load the rate rises from 40% to 61%. Both are "normal for bulk", and both are true.
Step 2 — so ask for consecutiveness instead. Re-run with a run-length monitor. Endpoint 2 reaches a consecutive-NAK run of 4,900 and never answers again. The rate barely moved because the other three endpoints kept the aggregate looking ordinary.
Step 3 — which is exactly section 4. The vendor tool keeps one counter per device. On a single-function device that is indistinguishable from per-endpoint. This is a four-endpoint device, so three healthy endpoints were resetting the wedged one's counter thousands of times a second.
Step 4 — why the camera matters. It does not cause the bug. It causes the scheduling pressure that makes endpoint 2's firmware take longer than its watchdog, wedge, and stop. The camera is the trigger and the shared counter is the reason nobody saw it for six months.
Step 5 — why it is intermittent. Endpoint 2 recovers on the next SET_INTERFACE, which the audio stack issues on every stream start. So the endpoint wedges, is invisible, and is cleared by a routine operation before anybody captures it.
The fix. One run counter per endpoint, reported on the transition, with the endpoint number in the message. The firmware bug took an afternoon once it could be seen.
18. UVM: Endpoint Health as a Subscriber
The health monitor is not a driver and not a scoreboard. It observes transactions that another component is already producing and maintains a longitudinal judgement about each endpoint — which is precisely what uvm_subscriber is for.
// A transaction carries what the analyser saw. Nothing more: the health
// monitor's whole value is that it derives its judgement from the wire,
// not from a status register the device also owns.
class usb_xact_item extends uvm_sequence_item;
`uvm_object_utils(usb_xact_item)
rand bit [3:0] ep;
rand resp_e resp; // ACK / NAK / STALL / NYET
rand bit toggle;
rand bit clear_halt;
function new(string name = "usb_xact_item"); super.new(name); endfunction
// A NAK is not an error, so it is not constrained AWAY. A generator that
// emits mostly ACKs produces a bus nobody has -- and, worse, one in which
// the toggle check and the wedge detector are never under load.
constraint c_mix {
resp dist { R_ACK := 35, R_NAK := 50, R_NYET := 12, R_STALL := 3 };
clear_halt dist { 0 := 97, 1 := 3 };
}
endclass
// ---------------------------------------------------------------------
// The health monitor. A SUBSCRIBER, because it consumes an analysis
// stream and produces a judgement rather than a response.
//
// The state is an ASSOCIATIVE ARRAY keyed by endpoint, not four fields
// and not a single counter. That is the UVM-shaped version of section 4:
// the data structure itself makes "one run counter for the device"
// impossible to write by accident.
// ---------------------------------------------------------------------
class usb_endpoint_health_sub extends uvm_subscriber #(usb_xact_item);
`uvm_component_utils(usb_endpoint_health_sub)
int unsigned NAK_RUN_MAX = 16;
typedef struct {
int unsigned run; // CONSECUTIVE NAKs, not total
bit halted; // sticky until CLEAR_FEATURE
bit wed; // reported: the event is the TRANSITION
bit tog; // expected DATA0/DATA1
} ep_health_t;
ep_health_t h [int];
int unsigned n_wedged, n_polled_halted, n_toggle, n_clear_unhalted;
function new(string name, uvm_component parent); super.new(name, parent);
endfunction
function ep_health_t get_ep(int ep);
if (!h.exists(ep)) h[ep] = '{run: 0, halted: 0, wed: 0, tog: 0};
return h[ep];
endfunction
function void write(usb_xact_item t);
ep_health_t e = get_ep(t.ep);
if (t.clear_halt) begin
if (!e.halted) begin
// The driver believes the endpoint is halted and the device does
// not. Harmless on the wire; a disagreement worth a message.
`uvm_warning("EP/CLEAR_UNHALTED",
$sformatf("ep %0d: CLEAR_FEATURE on an endpoint that is not halted",
t.ep))
n_clear_unhalted++;
end
// A cleared endpoint restarts at DATA0. Forgetting this costs exactly
// one silently dropped packet, for ever.
e = '{run: 0, halted: 0, wed: 0, tog: 0};
h[t.ep] = e;
return;
end
if (e.halted) begin
// Counted against the HOST, not the device. On a trace the two look
// identical and they belong to different teams.
`uvm_error("EP/POLLED_HALTED",
$sformatf("ep %0d: token issued to a halted endpoint", t.ep))
n_polled_halted++;
h[t.ep] = e;
return;
end
case (t.resp)
R_ACK: begin
if (t.toggle != e.tog) begin
// Loss and duplication are the SAME symptom in one bit. The
// message says so rather than guessing, because the two want
// opposite fixes.
`uvm_error("EP/TOGGLE",
$sformatf("ep %0d: toggle sequence broke (loss or duplication — one bit cannot say which)",
t.ep))
n_toggle++;
end
e.tog = ~e.tog;
e.run = 0; // a successful transfer, not the clock
e.wed = 0; // re-arm
end
R_NAK: begin
// NOT an error, NOT reported, and NOT counted as one. Only the run
// matters.
e.run++;
if (e.run >= NAK_RUN_MAX && !e.wed) begin
`uvm_error("EP/WEDGED",
$sformatf("ep %0d: %0d consecutive NAKs with no successful transfer",
t.ep, e.run))
n_wedged++;
e.wed = 1; // report the TRANSITION, once
end
end
R_STALL: begin
// An event, not an error. What would be an error is this halt
// still being set at the end of the test.
`uvm_info("EP/STALL",
$sformatf("ep %0d: halted by the device", t.ep), UVM_MEDIUM)
e.halted = 1;
e.run = 0;
end
default: begin // NYET: accepted, just not ready for more
e.run = 0;
e.wed = 0;
end
endcase
h[t.ep] = e;
endfunction
// ---- THE CHECK THAT ONLY A LONGITUDINAL COMPONENT CAN MAKE ----
//
// Every other check above fires on a transaction. This one fires on the
// ABSENCE of one: an endpoint still halted when the test ends was never
// cleared, and nothing in the transaction stream says so. A checker with
// no end-of-test phase cannot express it at all.
function void check_phase(uvm_phase phase);
super.check_phase(phase);
foreach (h[ep])
if (h[ep].halted)
`uvm_error("EP/LEFT_HALTED",
$sformatf("ep %0d: still halted at end of test — firmware stalled it and nobody cleared it",
ep))
endfunction
function void report_phase(uvm_phase phase);
super.report_phase(phase);
`uvm_info("EP/HEALTH",
$sformatf("wedged=%0d polled_halted=%0d toggle=%0d clear_unhalted=%0d",
n_wedged, n_polled_halted, n_toggle, n_clear_unhalted),
UVM_LOW)
endfunction
endclass19. Common Misconceptions
"A high NAK rate means the device is struggling." It means the device has a bulk endpoint and the host is polling it. Nothing more.
"A wedged endpoint shows up as a spike in the NAK count." It shows up as a run. The count may not move at all.
"NAKs and STALLs are both failures, just different severities." A NAK is not a failure at any severity. A STALL is a state change.
"A successful transfer proves the endpoint is not halted." In the design, yes — which is why the halted branch returns before classifying the response. On a trace, an ACK after a STALL means your capture is missing the CLEAR_FEATURE.
"CLEAR_FEATURE only needs to clear the halt." It must reset the toggle to DATA0 as well, or the next packet is silently dropped.
"A toggle error tells you a packet was lost." It tells you the sequence broke. Duplication produces the identical symptom and wants the opposite fix.
"One counter is fine; I will add the endpoint index later." The single-counter version passes on every simple device and fails only on the complex ones, which is the worst possible order to discover it in.
"Reporting a persistent fault repeatedly is harmless." It is how a channel gets muted, after which the checker is off and nobody knows.
"A mutation that scores zero means the testbench missed it." It can equally mean the mutation was dead code. Q3 scored zero in three languages for that reason.
20. Exercises
1. A device has one endpoint and a per-device NAK run counter. Show that its behaviour is indistinguishable from the correct design, and say what that implies about where this bug is found.
2. Q5's directed score is exactly 144. Derive that number from the phase 4 loop structure, and say which single check in check_out contributes it.
3. Phase 3 sends 40 further NAKs after the threshold and asserts the wedge count does not change. Remove the wed latch and work out how many errors those 40 NAKs produce — then check your answer against Q2's directed column.
4. Construct a transaction sequence in which a one-bit toggle cannot distinguish a lost packet from a duplicated one, and write the smallest change to the protocol that would make them distinguishable.
5. The design resets the NAK run on NYET. Argue both sides, then decide what a NYET-heavy high-speed endpoint would do to a wedge detector that treated NYET like a NAK.
6. Add a SET_INTERFACE-style input that clears every endpoint's state at once. Write the check that would have caught the intermittent bug in section 17 despite it.
7. The UVM subscriber's check_phase reports endpoints still halted at end of test. Write the stimulus that makes that check fire, and explain why no amount of constrained-random transaction generation produces it reliably.
21. Summary
| Idea | Why it matters |
|---|---|
| A NAK is not an error | it is flow control working, thousands of times a second |
| A STALL is not a NAK | one is not yet, the other is no, until cleared |
| The measurement is a run, not a count | a busy endpoint and a dead one have the same total |
| The run resets on a transfer, not on time | an endpoint that answers is healthy by definition |
| The run must be per endpoint | healthy endpoints reset a shared counter for the broken one |
| Report the transition, not the state | or a wedge produces one line per microframe |
| A halt is sticky | only CLEAR_FEATURE ends it, and it resets the toggle too |
| Polling a halted endpoint is the host's bug | same two trace lines, different team |
| One bit detects two opposite faults | and cannot say which; duplication and loss want opposite fixes |
| A wedge is sustained silence | random stimulus does not produce it, at any length |
| Observe the endpoints you are not driving | Q5's directed score was 0 until phase 4 peeked |
| A zero can mean the mutation was dead code | Q3 scored 0 in three languages and was unreachable |
| A weak mutant looks like a weak testbench | Q1's VHDL column was 42074 because the injection was smaller |
| 72/72 situations, 1036133 checks | 7 mutations, all killed in 3 languages, all by directed stimulus |
Tooling
| Step | Command |
|---|---|
| Verilog-2005 | iverilog -g2005 -o eh_v.out eh_v.v eh_v_tb.v && ./eh_v.out |
| SystemVerilog | iverilog -g2012 -o eh_sv.out eh_sv.sv eh_sv_tb.sv && ./eh_sv.out |
| VHDL-2008 analyse | nvc --std=2008 -a eh_vhdl.vhd eh_vhdl_tb.vhd |
| VHDL-2008 elaborate | nvc --std=2008 -e tb_eh_vhdl |
| VHDL-2008 run | nvc --std=2008 -r tb_eh_vhdl |
| One mutation | iverilog -g2005 -DMUT_Q5 -o mm eh_v_mut.v eh_v_tb.v && ./mm |
| Directed only | iverilog -g2005 -DDIRECTED_ONLY -o mm eh_v_mut.v eh_v_tb.v && ./mm |
All three implementations pass with 0 errors: every one of the 72 situations visited, 1036133 checks executed against an independently written shadow model, and every one of the seven mutations killed by directed stimulus alone.
Chapter 25.4 — Transfer Errors moves up one level, from the endpoint to the transfer built out of its transactions. The problem it opens with is the mirror image of this one: retries mask the error rate. A link losing one packet in fifty looks perfect at the application layer, because the host quietly re-sent them all — and the number that matters is not the error count but the retry-adjusted rate, attributed to a layer.
Continue learning
Related tutorials
- Related topic
Endpoint Logic
A lost ACK and a lost data packet look identical to the host, so it resends the same bytes — and the data toggle is the only thing that tells a device a retransmission from new data.
- Related topic
What Is USB?
The opening interview question answered with one load-bearing idea instead of a list — USB is host-scheduled, and polling, NAK, the frame and the missing interrupt line are all consequences of it.
- Related topic
The Endpoints Question
Endpoint 1 IN and endpoint 1 OUT are two different endpoints — separate buffers, toggles, halt states and packet sizes — and a table indexed by number alone is a device where halting a read kills its writes.
- Related topic
USB in Embedded Devices
On a microcontroller the controller is a peripheral, and the protocol can be perfectly correct while the device goes deaf. Everything turns on one question — who owns this buffer right now — answered by one bit per buffer and exhausted over 132 transitions in three languages.
Standards & specifications
- Governing standard
- USB-IF (Universal Serial Bus Specification)(opens USB Implementers Forum (USB-IF) in a new tab)
Defines the USB bus — its electrical signalling, connectors, packet and transaction model, device framework and the descriptors a device must expose — together with the device-class specifications layered on it. It does not define host-controller register interfaces (xHCI and EHCI are separate documents) nor any operating system's driver architecture.
This page also covers RTL structure, verification approach and debugging technique. Those are engineering practice built on the standard, not requirements the standard itself imposes.
Where this fits
Part of the USB curriculum.
