USB · Module 26
Firmware Interaction
A register interface looks like memory and is not — reading can change it, writing 1 can clear it, one register can apply another, and none of that is anything a compiler knows.
Chapter 26.4 got the driver's attention. This chapter is about what it does next, and about the gap between what a register interface looks like and what it is.
1. A Register Interface Is Not Memory
It looks like memory. It is addressed like memory, the compiler emits loads and stores for it, and every abstraction between the driver and the wire treats it as a place to put numbers.
It is not memory, because:
reading can CHANGE it read-to-clear
writing 1 can CLEAR it write-1-to-clear
writing one register can APPLY another shadow / commit
writing some registers is REJECTED locked while busyNone of which a compiler knows, and none of which appears in a header file. Every one of them is a source of driver bugs that look like hardware bugs.
2. Read-to-Clear and Write-1-to-Clear Are Not the Same Thing
Both hand the driver a set of events and then empty the register. The difference is when, and what happens to an event that arrives in between.
| How it clears | What it can destroy | |
|---|---|---|
| read-to-clear | the read returns the bits and clears them, indivisibly | an event arriving on the read cycle — it was not returned |
| write-1-to-clear | the driver clears exactly the bits it handled, later | nothing it did not see, because it never writes those bits |
3. A Multi-Field Configuration Must Be Applied Atomically
A configuration that spans two registers cannot be written live. Between the two stores the hardware is running on half the old value and half the new one, and for one bus cycle it is configured as something nobody asked for.
write CFG_LO -> a shadow
write CFG_HI -> a shadow
write COMMIT -> BOTH become live, in the same cycleThe shadow is not an optimisation. It is the only way to make a 32-bit change atomic across a 16-bit interface — and one bus cycle of a wrong configuration is thousands of bit times on a USB link.
4. And Some Writes Are Rejected
A register that configures a running engine cannot be changed while it runs.
5. Reserved Bits Are a Two-Sided Rule
They read as zero and ignore writes, and the two halves have to agree.
A driver doing read-modify-write writes back exactly what it read. If reserved bits read as zero but accept writes, the cycle is still a fixed point. If they read as something else and accept writes, the register walks somewhere new every time round — which is the kind of bug that appears only after the tenth iteration of a polling loop.
6. What We Are Building
usb_reg_file implements one of each behaviour, in eight addresses:
| Address | Register | Behaviour |
|---|---|---|
| 0 | CTRL | RW, upper byte reserved |
| 1 | STATUS | read-to-clear, and not writable |
| 2 | EVENT | write-1-to-clear |
| 3 / 4 | CFG_LO / CFG_HI | shadow — staged, not live |
| 5 | COMMIT | write-only; applies both halves at once |
| 6 | LOCKED | RW, rejected while busy |
| 7 | — | a hole in the map |
7. Verilog-2005 Implementation
// usb_reg_file -- why driver code that looks like memory access is not, and
// the four register behaviours that make it a protocol.
//
// A REGISTER INTERFACE IS NOT MEMORY
//
// It looks like memory. It is addressed like memory, the compiler emits
// loads and stores for it, and every abstraction between the driver and the
// wire treats it as a place to put numbers.
//
// It is not memory, because:
//
// reading can CHANGE it (read-to-clear)
// writing 1 can CLEAR it (write-1-to-clear)
// writing one register can APPLY another (shadow / commit)
// writing some registers is REJECTED at the wrong moment
//
// None of which a compiler knows, and none of which shows up in a header
// file. Every one of them is a source of driver bugs that look like hardware
// bugs.
//
// READ-TO-CLEAR AND WRITE-1-TO-CLEAR ARE NOT THE SAME THING
//
// Both hand the driver a set of events and then empty the register. The
// difference is WHEN, and what happens to an event that arrives in between.
//
// read-to-clear the read returns the bits AND clears them, in one
// indivisible step. Simple, and it has a trap: the
// register must clear only THE BITS IT RETURNED. An
// event arriving on the read cycle was not returned,
// so clearing it destroys an interrupt nobody ever saw.
//
// write-1-to-clear the driver clears exactly the bits it handled,
// explicitly, later. Nothing it did not see can be
// destroyed, because it never writes those bits.
//
// The same event source feeds both registers in this block, deliberately, so
// the two policies can be compared on identical stimulus.
//
// A MULTI-FIELD CONFIGURATION MUST BE APPLIED ATOMICALLY
//
// A configuration that spans two registers cannot be written live. Between
// the two stores the hardware is running on half the old value and half the
// new one, and for one bus cycle it is configured as something nobody asked
// for.
//
// write CFG_LO -> a shadow
// write CFG_HI -> a shadow
// write COMMIT -> BOTH become live, in the same cycle
//
// The shadow is not an optimisation. It is the only way to make a 32-bit
// change atomic across a 16-bit interface.
//
// AND SOME WRITES ARE REJECTED
//
// A register that configures a running engine cannot be changed while it
// runs. Silently ignoring the write is the worst option: the driver believes
// the setting took. This block reports it, because a rejected write is a
// driver bug and the driver is the only thing that can fix it.
module usb_reg_file #(
parameter integer N_EVT = 4 // event bits, in both status registers
) (
input wire clk,
input wire rst_n,
input wire [2:0] addr,
input wire rd,
input wire wr,
input wire [15:0] wdata,
input wire [N_EVT-1:0] evt, // one pulse per hardware event
input wire busy, // the engine is running: config is locked
input wire eot,
output wire [15:0] rdata,
output wire ack,
output wire err_pulse,
output wire [1:0] err_code,
output wire [15:0] ctrl,
output wire [N_EVT-1:0] status, // read-to-clear
output wire [N_EVT-1:0] event_w1c,
output wire [31:0] cfg_live, // what the engine is actually using
output wire [31:0] cfg_shadow, // what has been staged but not applied
output reg [31:0] n_read,
output reg [31:0] n_write,
output reg [31:0] n_commit,
output reg [31:0] n_rc_lost, // an event destroyed by a read-to-clear
output reg [31:0] n_locked, // a write rejected because busy
output reg [31:0] n_badaddr
);
// The map. Addresses are named once, here, because a driver header and an
// RTL case statement that disagree is a bug no simulation can find.
localparam [2:0] A_CTRL = 3'd0, // RW, with reserved bits
A_STATUS = 3'd1, // read-to-clear
A_EVENT = 3'd2, // write-1-to-clear
A_CFG_LO = 3'd3, // shadow
A_CFG_HI = 3'd4, // shadow
A_COMMIT = 3'd5, // write-only: applies the shadow
A_LOCKED = 3'd6, // RW, rejected while busy
A_RSVD = 3'd7; // not a register
localparam [1:0] E_NONE = 2'd0,
E_LOCKED = 2'd1, // a write while the engine is running
E_BADADDR = 2'd2; // an access to a hole in the map
// Reserved bits read as ZERO and ignore writes. Stated as a mask rather
// than assumed, because a driver doing read-modify-write on this register
// will write back whatever it read -- and if the reserved bits read back
// as something other than the value it must write, every such cycle
// corrupts them.
localparam [15:0] CTRL_MASK = 16'h00FF;
reg [15:0] ctrl_r;
reg [N_EVT-1:0] stat_r, evw_r;
reg [31:0] shadow_r, live_r;
reg [15:0] lock_r;
reg ack_r, er_r;
reg [1:0] ec_r;
// The RAW register, not the masked read value. The read path masks
// separately; exposing the stored value here is what makes "the reserved
// bits are never even STORED" checkable, rather than only "the read path
// hides them".
assign ctrl = ctrl_r;
assign status = stat_r;
assign event_w1c = evw_r;
assign cfg_live = live_r;
assign cfg_shadow = shadow_r;
assign ack = ack_r;
assign err_pulse = er_r;
assign err_code = ec_r;
// ---- The read data is COMBINATIONAL and has no side effect of its own.
//
// The side effect lives in the sequential block. Mixing the two -- reading
// a register that the same expression also clears -- is how a read-to-clear
// ends up returning the value it just destroyed.
reg [15:0] rd_mux;
always @* begin
case (addr)
A_CTRL: rd_mux = ctrl_r & CTRL_MASK;
A_STATUS: rd_mux = {{(16-N_EVT){1'b0}}, stat_r};
A_EVENT: rd_mux = {{(16-N_EVT){1'b0}}, evw_r};
A_CFG_LO: rd_mux = shadow_r[15:0];
A_CFG_HI: rd_mux = shadow_r[31:16];
// COMMIT is write-only. It reads as zero rather than as the last
// thing written, so a driver that reads it back does not conclude the
// commit is still pending.
A_COMMIT: rd_mux = 16'h0000;
A_LOCKED: rd_mux = lock_r;
default: rd_mux = 16'h0000;
endcase
end
assign rdata = rd_mux;
integer i;
reg [15:0] ctrl_n, lock_n;
reg [N_EVT-1:0] stat_n, evw_n;
reg [31:0] shadow_n, live_n;
reg ack_n, er_n;
reg [1:0] ec_n;
reg rd_n, wr_n, cm_n, lk_n, ba_n;
reg [31:0] rcl_n;
always @* begin
ctrl_n = ctrl_r;
lock_n = lock_r;
stat_n = stat_r;
evw_n = evw_r;
shadow_n = shadow_r;
live_n = live_r;
ack_n = 1'b0; er_n = 1'b0; ec_n = E_NONE;
rd_n = 1'b0; wr_n = 1'b0; cm_n = 1'b0; lk_n = 1'b0; ba_n = 1'b0;
rcl_n = 32'd0;
if (eot) begin
// Nothing: the counters are the report.
end else begin
// ---- The event sources come FIRST, and they SET. ----
//
// Before any clear, so that an event arriving on the same cycle as a
// read or a write survives it. This is chapter 26.4's rule applied to
// two different clear policies at once.
for (i = 0; i < N_EVT; i = i + 1) begin
if (evt[i]) begin
stat_n[i] = 1'b1;
evw_n[i] = 1'b1;
end
end
if (rd && wr) begin
// A simultaneous read and write of the same address is not something
// a single-port interface can do, and accepting it quietly gives the
// driver a value that belongs to neither.
er_n = 1'b1; ec_n = E_BADADDR;
ba_n = 1'b1;
end else if (rd) begin
ack_n = 1'b1;
rd_n = 1'b1;
if (addr == A_RSVD) begin
er_n = 1'b1; ec_n = E_BADADDR;
ba_n = 1'b1;
end else if (addr == A_STATUS) begin
// ---- READ-TO-CLEAR, and the whole trap in one line. ----
//
// Clear only the bits that were RETURNED. An event that arrived on
// this very cycle is in stat_n and was NOT in rd_mux, so clearing
// the whole register would destroy an interrupt nobody ever saw --
// and nothing else in the system will ever mention it again.
for (i = 0; i < N_EVT; i = i + 1)
if (stat_r[i] && !evt[i]) stat_n[i] = 1'b0;
// The count of events that WOULD have been lost by a naive clear.
// Reported so that the hazard is visible in a run rather than
// argued about in a review.
for (i = 0; i < N_EVT; i = i + 1)
if (evt[i] && !stat_r[i]) rcl_n = rcl_n + 32'd1;
end
end else if (wr) begin
ack_n = 1'b1;
wr_n = 1'b1;
case (addr)
A_CTRL: begin
// Reserved bits ignore writes as well as reading zero. The two
// halves of that rule have to agree or a read-modify-write walks
// the register somewhere new every time round.
ctrl_n = wdata & CTRL_MASK;
end
A_STATUS: begin
// A read-to-clear register is not writable. Accepting the write
// silently lets a driver that treats it like the W1C register
// clear events it never read.
er_n = 1'b1; ec_n = E_BADADDR;
ba_n = 1'b1;
end
A_EVENT: begin
// ---- WRITE-1-TO-CLEAR. Set still wins. ----
for (i = 0; i < N_EVT; i = i + 1)
if (wdata[i] && !evt[i]) evw_n[i] = 1'b0;
end
A_CFG_LO: shadow_n[15:0] = wdata;
A_CFG_HI: shadow_n[31:16] = wdata;
A_COMMIT: begin
// ---- The whole 32 bits become live in ONE cycle. ----
//
// Which is the only way to make the change atomic across a
// 16-bit interface. Writing the halves live means the engine
// runs on a configuration nobody asked for for one bus cycle,
// and one bus cycle is thousands of bit times on a USB link.
live_n = shadow_r;
cm_n = 1'b1;
end
A_LOCKED: begin
if (busy) begin
// Reported, not ignored. A silently dropped write leaves the
// driver believing a setting took effect, and the next thing
// it does is built on that belief.
er_n = 1'b1; ec_n = E_LOCKED;
lk_n = 1'b1;
end else begin
lock_n = wdata;
end
end
default: begin
er_n = 1'b1; ec_n = E_BADADDR;
ba_n = 1'b1;
end
endcase
end
end
end
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
ctrl_r <= 16'd0;
lock_r <= 16'd0;
stat_r <= {N_EVT{1'b0}};
evw_r <= {N_EVT{1'b0}};
shadow_r <= 32'd0;
live_r <= 32'd0;
ack_r <= 1'b0; er_r <= 1'b0; ec_r <= E_NONE;
n_read <= 32'd0;
n_write <= 32'd0;
n_commit <= 32'd0;
n_rc_lost <= 32'd0;
n_locked <= 32'd0;
n_badaddr <= 32'd0;
end else begin
ctrl_r <= ctrl_n;
lock_r <= lock_n;
stat_r <= stat_n;
evw_r <= evw_n;
shadow_r <= shadow_n;
live_r <= live_n;
ack_r <= ack_n; er_r <= er_n; ec_r <= ec_n;
if (rd_n) n_read <= n_read + 32'd1;
if (wr_n) n_write <= n_write + 32'd1;
if (cm_n) n_commit <= n_commit + 32'd1;
if (lk_n) n_locked <= n_locked + 32'd1;
if (ba_n) n_badaddr <= n_badaddr + 32'd1;
n_rc_lost <= n_rc_lost + rcl_n;
end
end
endmodule8. SystemVerilog Implementation
// usb_reg_file -- why driver code that looks like memory access is not, and
// the four register behaviours that make it a protocol.
//
// A REGISTER INTERFACE IS NOT MEMORY
//
// It looks like memory. It is addressed like memory, the compiler emits
// loads and stores for it, and every abstraction between the driver and the
// wire treats it as a place to put numbers.
//
// It is not memory, because:
//
// reading can CHANGE it (read-to-clear)
// writing 1 can CLEAR it (write-1-to-clear)
// writing one register can APPLY another (shadow / commit)
// writing some registers is REJECTED at the wrong moment
//
// None of which a compiler knows, and none of which shows up in a header
// file. Every one of them is a source of driver bugs that look like hardware
// bugs.
//
// READ-TO-CLEAR AND WRITE-1-TO-CLEAR ARE NOT THE SAME THING
//
// Both hand the driver a set of events and then empty the register. The
// difference is WHEN, and what happens to an event that arrives in between.
//
// read-to-clear the read returns the bits AND clears them, in one
// indivisible step. Simple, and it has a trap: the
// register must clear only THE BITS IT RETURNED. An
// event arriving on the read cycle was not returned,
// so clearing it destroys an interrupt nobody ever saw.
//
// write-1-to-clear the driver clears exactly the bits it handled,
// explicitly, later. Nothing it did not see can be
// destroyed, because it never writes those bits.
//
// The same event source feeds both registers in this block, deliberately, so
// the two policies can be compared on identical stimulus.
//
// A MULTI-FIELD CONFIGURATION MUST BE APPLIED ATOMICALLY
//
// A configuration that spans two registers cannot be written live. Between
// the two stores the hardware is running on half the old value and half the
// new one, and for one bus cycle it is configured as something nobody asked
// for.
//
// write CFG_LO -> a shadow
// write CFG_HI -> a shadow
// write COMMIT -> BOTH become live, in the same cycle
//
// The shadow is not an optimisation. It is the only way to make a 32-bit
// change atomic across a 16-bit interface.
//
// AND SOME WRITES ARE REJECTED
//
// A register that configures a running engine cannot be changed while it
// runs. Silently ignoring the write is the worst option: the driver believes
// the setting took. This block reports it, because a rejected write is a
// driver bug and the driver is the only thing that can fix it.
package usb_reg_pkg;
// The map. Named once, in a package, because a driver header and an RTL
// case statement that disagree is a bug no simulation can find.
typedef enum logic [2:0] {
A_CTRL = 3'd0, // RW, with reserved bits
A_STATUS = 3'd1, // read-to-clear
A_EVENT = 3'd2, // write-1-to-clear
A_CFG_LO = 3'd3, // shadow
A_CFG_HI = 3'd4, // shadow
A_COMMIT = 3'd5, // write-only: applies the shadow
A_LOCKED = 3'd6, // RW, rejected while busy
A_RSVD = 3'd7 // not a register
} reg_addr_e;
typedef enum logic [1:0] {
E_NONE = 2'd0,
E_LOCKED = 2'd1, // a write while the engine is running
E_BADADDR = 2'd2 // an access to a hole in the map
} reg_err_e;
endpackage
module usb_reg_file
import usb_reg_pkg::*;
#(
parameter int N_EVT = 4 // event bits, in both status registers
) (
input logic clk,
input logic rst_n,
input reg_addr_e addr,
input logic rd,
input logic wr,
input logic [15:0] wdata,
input logic [N_EVT-1:0] evt, // one pulse per hardware event
input logic busy, // the engine is running: config is locked
input logic eot,
output logic [15:0] rdata,
output logic ack,
output logic err_pulse,
output reg_err_e err_code,
output logic [15:0] ctrl,
output logic [N_EVT-1:0] status, // read-to-clear
output logic [N_EVT-1:0] event_w1c,
output logic [31:0] cfg_live, // what the engine is actually using
output logic [31:0] cfg_shadow, // what has been staged but not applied
output logic [31:0] n_read,
output logic [31:0] n_write,
output logic [31:0] n_commit,
output logic [31:0] n_rc_lost, // an event destroyed by a read-to-clear
output logic [31:0] n_locked, // a write rejected because busy
output logic [31:0] n_badaddr
);
// Reserved bits read as ZERO and ignore writes. Stated as a mask rather
// than assumed, because a driver doing read-modify-write on this register
// will write back whatever it read -- and if the reserved bits read back
// as something other than the value it must write, every such cycle
// corrupts them.
localparam logic [15:0] CTRL_MASK = 16'h00FF;
logic [15:0] ctrl_r;
logic [N_EVT-1:0] stat_r, evw_r;
logic [31:0] shadow_r, live_r;
logic [15:0] lock_r;
logic ack_r, er_r;
reg_err_e ec_r;
// The RAW register, not the masked read value. The read path masks
// separately; exposing the stored value here is what makes "the reserved
// bits are never even STORED" checkable, rather than only "the read path
// hides them".
assign ctrl = ctrl_r;
assign status = stat_r;
assign event_w1c = evw_r;
assign cfg_live = live_r;
assign cfg_shadow = shadow_r;
assign ack = ack_r;
assign err_pulse = er_r;
assign err_code = ec_r;
// ---- The read data is COMBINATIONAL and has no side effect of its own.
//
// The side effect lives in the sequential block. Mixing the two -- reading
// a register that the same expression also clears -- is how a read-to-clear
// ends up returning the value it just destroyed.
logic [15:0] rd_mux;
always_comb begin
case (addr)
A_CTRL: rd_mux = ctrl_r & CTRL_MASK;
A_STATUS: rd_mux = {{(16-N_EVT){1'b0}}, stat_r};
A_EVENT: rd_mux = {{(16-N_EVT){1'b0}}, evw_r};
A_CFG_LO: rd_mux = shadow_r[15:0];
A_CFG_HI: rd_mux = shadow_r[31:16];
// COMMIT is write-only. It reads as zero rather than as the last
// thing written, so a driver that reads it back does not conclude the
// commit is still pending.
A_COMMIT: rd_mux = 16'h0000;
A_LOCKED: rd_mux = lock_r;
default: rd_mux = 16'h0000;
endcase
end
assign rdata = rd_mux;
int i;
logic [15:0] ctrl_n, lock_n;
logic [N_EVT-1:0] stat_n, evw_n;
logic [31:0] shadow_n, live_n;
logic ack_n, er_n;
reg_err_e ec_n;
logic rd_n, wr_n, cm_n, lk_n, ba_n;
logic [31:0] rcl_n;
always_comb begin
ctrl_n = ctrl_r;
lock_n = lock_r;
stat_n = stat_r;
evw_n = evw_r;
shadow_n = shadow_r;
live_n = live_r;
ack_n = 1'b0; er_n = 1'b0; ec_n = E_NONE;
rd_n = 1'b0; wr_n = 1'b0; cm_n = 1'b0; lk_n = 1'b0; ba_n = 1'b0;
rcl_n = 32'd0;
if (eot) begin
// Nothing: the counters are the report.
end else begin
// ---- The event sources come FIRST, and they SET. ----
//
// Before any clear, so that an event arriving on the same cycle as a
// read or a write survives it. This is chapter 26.4's rule applied to
// two different clear policies at once.
for (i = 0; i < N_EVT; i = i + 1) begin
if (evt[i]) begin
stat_n[i] = 1'b1;
evw_n[i] = 1'b1;
end
end
if (rd && wr) begin
// A simultaneous read and write of the same address is not something
// a single-port interface can do, and accepting it quietly gives the
// driver a value that belongs to neither.
er_n = 1'b1; ec_n = E_BADADDR;
ba_n = 1'b1;
end else if (rd) begin
ack_n = 1'b1;
rd_n = 1'b1;
if (addr == A_RSVD) begin
er_n = 1'b1; ec_n = E_BADADDR;
ba_n = 1'b1;
end else if (addr == A_STATUS) begin
// ---- READ-TO-CLEAR, and the whole trap in one line. ----
//
// Clear only the bits that were RETURNED. An event that arrived on
// this very cycle is in stat_n and was NOT in rd_mux, so clearing
// the whole register would destroy an interrupt nobody ever saw --
// and nothing else in the system will ever mention it again.
for (i = 0; i < N_EVT; i = i + 1)
if (stat_r[i] && !evt[i]) stat_n[i] = 1'b0;
// The count of events that WOULD have been lost by a naive clear.
// Reported so that the hazard is visible in a run rather than
// argued about in a review.
for (i = 0; i < N_EVT; i = i + 1)
if (evt[i] && !stat_r[i]) rcl_n = rcl_n + 32'd1;
end
end else if (wr) begin
ack_n = 1'b1;
wr_n = 1'b1;
case (addr)
A_CTRL: begin
// Reserved bits ignore writes as well as reading zero. The two
// halves of that rule have to agree or a read-modify-write walks
// the register somewhere new every time round.
ctrl_n = wdata & CTRL_MASK;
end
A_STATUS: begin
// A read-to-clear register is not writable. Accepting the write
// silently lets a driver that treats it like the W1C register
// clear events it never read.
er_n = 1'b1; ec_n = E_BADADDR;
ba_n = 1'b1;
end
A_EVENT: begin
// ---- WRITE-1-TO-CLEAR. Set still wins. ----
for (i = 0; i < N_EVT; i = i + 1)
if (wdata[i] && !evt[i]) evw_n[i] = 1'b0;
end
A_CFG_LO: shadow_n[15:0] = wdata;
A_CFG_HI: shadow_n[31:16] = wdata;
A_COMMIT: begin
// ---- The whole 32 bits become live in ONE cycle. ----
//
// Which is the only way to make the change atomic across a
// 16-bit interface. Writing the halves live means the engine
// runs on a configuration nobody asked for for one bus cycle,
// and one bus cycle is thousands of bit times on a USB link.
live_n = shadow_r;
cm_n = 1'b1;
end
A_LOCKED: begin
if (busy) begin
// Reported, not ignored. A silently dropped write leaves the
// driver believing a setting took effect, and the next thing
// it does is built on that belief.
er_n = 1'b1; ec_n = E_LOCKED;
lk_n = 1'b1;
end else begin
lock_n = wdata;
end
end
default: begin
er_n = 1'b1; ec_n = E_BADADDR;
ba_n = 1'b1;
end
endcase
end
end
end
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
ctrl_r <= 16'd0;
lock_r <= 16'd0;
stat_r <= '0;
evw_r <= '0;
shadow_r <= 32'd0;
live_r <= 32'd0;
ack_r <= 1'b0; er_r <= 1'b0; ec_r <= E_NONE;
n_read <= 32'd0;
n_write <= 32'd0;
n_commit <= 32'd0;
n_rc_lost <= 32'd0;
n_locked <= 32'd0;
n_badaddr <= 32'd0;
end else begin
ctrl_r <= ctrl_n;
lock_r <= lock_n;
stat_r <= stat_n;
evw_r <= evw_n;
shadow_r <= shadow_n;
live_r <= live_n;
ack_r <= ack_n; er_r <= er_n; ec_r <= ec_n;
if (rd_n) n_read <= n_read + 32'd1;
if (wr_n) n_write <= n_write + 32'd1;
if (cm_n) n_commit <= n_commit + 32'd1;
if (lk_n) n_locked <= n_locked + 32'd1;
if (ba_n) n_badaddr <= n_badaddr + 32'd1;
n_rc_lost <= n_rc_lost + rcl_n;
end
end
endmodule9. VHDL-2008 Implementation
-- usb_reg_file -- why driver code that looks like memory access is not, and
-- the four register behaviours that make it a protocol.
--
-- A REGISTER INTERFACE IS NOT MEMORY
--
-- It looks like memory. It is addressed like memory, the compiler emits loads
-- and stores for it, and every abstraction between the driver and the wire
-- treats it as a place to put numbers.
--
-- It is not memory, because:
--
-- reading can CHANGE it (read-to-clear)
-- writing 1 can CLEAR it (write-1-to-clear)
-- writing one register can APPLY another (shadow / commit)
-- writing some registers is REJECTED at the wrong moment
--
-- None of which a compiler knows, and none of which shows up in a header
-- file. Every one of them is a source of driver bugs that look like hardware
-- bugs.
--
-- READ-TO-CLEAR AND WRITE-1-TO-CLEAR ARE NOT THE SAME THING
--
-- Both hand the driver a set of events and then empty the register. The
-- difference is WHEN, and what happens to an event that arrives in between.
--
-- read-to-clear the read returns the bits AND clears them, in one
-- indivisible step. Simple, and it has a trap: the
-- register must clear only THE BITS IT RETURNED. An event
-- arriving on the read cycle was not returned, so
-- clearing it destroys an interrupt nobody ever saw.
--
-- write-1-to-clear the driver clears exactly the bits it handled,
-- explicitly, later. Nothing it did not see can be
-- destroyed, because it never writes those bits.
--
-- The same event source feeds both registers in this block, deliberately, so
-- the two policies can be compared on identical stimulus.
--
-- A MULTI-FIELD CONFIGURATION MUST BE APPLIED ATOMICALLY
--
-- A configuration that spans two registers cannot be written live. Between
-- the two stores the hardware is running on half the old value and half the
-- new one, and for one bus cycle it is configured as something nobody asked
-- for.
--
-- write CFG_LO -> a shadow
-- write CFG_HI -> a shadow
-- write COMMIT -> BOTH become live, in the same cycle
--
-- The shadow is not an optimisation. It is the only way to make a 32-bit
-- change atomic across a 16-bit interface.
--
-- AND SOME WRITES ARE REJECTED
--
-- A register that configures a running engine cannot be changed while it
-- runs. Silently ignoring the write is the worst option: the driver believes
-- the setting took. This block reports it, because a rejected write is a
-- driver bug and the driver is the only thing that can fix it.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
package usb_reg_pkg is
-- The map. Named once, in a package, because a driver header and an RTL
-- case statement that disagree is a bug no simulation can find.
constant A_CTRL : std_logic_vector(2 downto 0) := "000";
constant A_STATUS : std_logic_vector(2 downto 0) := "001";
constant A_EVENT : std_logic_vector(2 downto 0) := "010";
constant A_CFG_LO : std_logic_vector(2 downto 0) := "011";
constant A_CFG_HI : std_logic_vector(2 downto 0) := "100";
constant A_COMMIT : std_logic_vector(2 downto 0) := "101";
constant A_LOCKED : std_logic_vector(2 downto 0) := "110";
constant A_RSVD : std_logic_vector(2 downto 0) := "111";
constant E_NONE : std_logic_vector(1 downto 0) := "00";
constant E_LOCKED : std_logic_vector(1 downto 0) := "01";
constant E_BADADDR : std_logic_vector(1 downto 0) := "10";
-- Reserved bits read as ZERO and ignore writes. Stated as a mask rather
-- than assumed, because a driver doing read-modify-write on this register
-- will write back whatever it read.
constant CTRL_MASK : std_logic_vector(15 downto 0) := x"00FF";
end package;
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
use work.usb_reg_pkg.all;
entity usb_reg_file is
generic (
N_EVT : integer := 4 -- event bits, in both status registers
);
port (
clk : in std_logic;
rst_n : in std_logic;
addr : in std_logic_vector(2 downto 0);
rd : in std_logic;
wr : in std_logic;
wdata : in std_logic_vector(15 downto 0);
evt : in std_logic_vector(N_EVT-1 downto 0);
busy : in std_logic;
eot : in std_logic;
rdata : out std_logic_vector(15 downto 0);
ack : out std_logic;
err_pulse : out std_logic;
err_code : out std_logic_vector(1 downto 0);
ctrl : out std_logic_vector(15 downto 0);
status : out std_logic_vector(N_EVT-1 downto 0); -- read-to-clear
event_w1c : out std_logic_vector(N_EVT-1 downto 0);
cfg_live : out std_logic_vector(31 downto 0);
cfg_shadow : out std_logic_vector(31 downto 0);
n_read : out unsigned(31 downto 0);
n_write : out unsigned(31 downto 0);
n_commit : out unsigned(31 downto 0);
n_rc_lost : out unsigned(31 downto 0);
n_locked : out unsigned(31 downto 0);
n_badaddr : out unsigned(31 downto 0)
);
end entity;
architecture rtl of usb_reg_file is
signal ctrl_r, lock_r : std_logic_vector(15 downto 0) := (others => '0');
signal stat_r, evw_r : std_logic_vector(N_EVT-1 downto 0)
:= (others => '0');
signal shadow_r, live_r : std_logic_vector(31 downto 0) := (others => '0');
signal ack_r, er_r : std_logic := '0';
signal ec_r : std_logic_vector(1 downto 0) := E_NONE;
signal rd_c, wr_c, cm_c : unsigned(31 downto 0) := (others => '0');
signal rcl_c, lk_c, ba_c : unsigned(31 downto 0) := (others => '0');
signal rd_mux : std_logic_vector(15 downto 0) := (others => '0');
begin
-- The RAW register, not the masked read value. The read path masks
-- separately; exposing the stored value here is what makes "the reserved
-- bits are never even STORED" checkable, rather than only "the read path
-- hides them". Masking here too would make a design that stores reserved
-- bits indistinguishable from one that does not.
ctrl <= ctrl_r;
status <= stat_r;
event_w1c <= evw_r;
cfg_live <= live_r;
cfg_shadow <= shadow_r;
ack <= ack_r;
err_pulse <= er_r;
err_code <= ec_r;
n_read <= rd_c;
n_write <= wr_c;
n_commit <= cm_c;
n_rc_lost <= rcl_c;
n_locked <= lk_c;
n_badaddr <= ba_c;
-- ---- The read data is COMBINATIONAL and has no side effect of its own.
--
-- The side effect lives in the clocked process. Mixing the two -- reading a
-- register that the same expression also clears -- is how a read-to-clear
-- ends up returning the value it just destroyed.
process (addr, ctrl_r, stat_r, evw_r, shadow_r, lock_r)
begin
case addr is
when A_CTRL => rd_mux <= ctrl_r and CTRL_MASK;
when A_STATUS =>
rd_mux <= (15 downto N_EVT => '0') & stat_r;
when A_EVENT =>
rd_mux <= (15 downto N_EVT => '0') & evw_r;
when A_CFG_LO => rd_mux <= shadow_r(15 downto 0);
when A_CFG_HI => rd_mux <= shadow_r(31 downto 16);
-- COMMIT is write-only. It reads as zero rather than as the last thing
-- written, so a driver that reads it back does not conclude the commit
-- is still pending.
when A_COMMIT => rd_mux <= (others => '0');
when A_LOCKED => rd_mux <= lock_r;
when others => rd_mux <= (others => '0');
end case;
end process;
rdata <= rd_mux;
process (clk, rst_n)
variable ctrl_v, lock_v : std_logic_vector(15 downto 0);
variable stat_v, evw_v : std_logic_vector(N_EVT-1 downto 0);
variable stat_pre : std_logic_vector(N_EVT-1 downto 0);
variable shadow_v, live_v : std_logic_vector(31 downto 0);
variable shadow_pre : std_logic_vector(31 downto 0);
variable ack_v, er_v : std_logic;
variable ec_v : std_logic_vector(1 downto 0);
variable d_rcl : integer;
begin
if rst_n = '0' then
ctrl_r <= (others => '0');
lock_r <= (others => '0');
stat_r <= (others => '0');
evw_r <= (others => '0');
shadow_r <= (others => '0');
live_r <= (others => '0');
ack_r <= '0'; er_r <= '0'; ec_r <= E_NONE;
rd_c <= (others => '0'); wr_c <= (others => '0');
cm_c <= (others => '0'); rcl_c <= (others => '0');
lk_c <= (others => '0'); ba_c <= (others => '0');
elsif rising_edge(clk) then
ctrl_v := ctrl_r;
lock_v := lock_r;
stat_v := stat_r;
evw_v := evw_r;
stat_pre := stat_r;
shadow_v := shadow_r;
shadow_pre := shadow_r;
live_v := live_r;
ack_v := '0'; er_v := '0'; ec_v := E_NONE;
d_rcl := 0;
if eot = '1' then
-- Nothing: the counters are the report.
null;
else
-- ---- The event sources come FIRST, and they SET. ----
--
-- Before any clear, so that an event arriving on the same cycle as a
-- read or a write survives it. This is chapter 26.4's rule applied to
-- two different clear policies at once.
for i in 0 to N_EVT-1 loop
if evt(i) = '1' then
stat_v(i) := '1';
evw_v(i) := '1';
end if;
end loop;
if rd = '1' and wr = '1' then
-- A simultaneous read and write of the same address is not
-- something a single-port interface can do, and accepting it
-- quietly gives the driver a value that belongs to neither.
er_v := '1'; ec_v := E_BADADDR;
ba_c <= ba_c + 1;
elsif rd = '1' then
ack_v := '1';
rd_c <= rd_c + 1;
if addr = A_RSVD then
er_v := '1'; ec_v := E_BADADDR;
ba_c <= ba_c + 1;
elsif addr = A_STATUS then
-- ---- READ-TO-CLEAR, and the whole trap in one line. ----
--
-- Clear only the bits that were RETURNED. An event that arrived
-- on this very cycle is in stat_v and was NOT in rd_mux, so
-- clearing the whole register would destroy an interrupt nobody
-- ever saw.
for i in 0 to N_EVT-1 loop
if stat_pre(i) = '1' and evt(i) = '0' then stat_v(i) := '0'; end if;
end loop;
-- The count of events that WOULD have been lost by a naive
-- clear, so the hazard is visible in a run rather than argued
-- about in a review.
for i in 0 to N_EVT-1 loop
if evt(i) = '1' and stat_pre(i) = '0' then d_rcl := d_rcl + 1; end if;
end loop;
end if;
elsif wr = '1' then
ack_v := '1';
wr_c <= wr_c + 1;
case addr is
when A_CTRL =>
-- Reserved bits ignore writes as well as reading zero. The two
-- halves of that rule have to agree or a read-modify-write
-- walks the register somewhere new every time round.
ctrl_v := wdata and CTRL_MASK;
when A_STATUS =>
-- A read-to-clear register is not writable. Accepting the
-- write silently lets a driver that treats it like the W1C
-- register clear events it never read.
er_v := '1'; ec_v := E_BADADDR;
ba_c <= ba_c + 1;
when A_EVENT =>
-- ---- WRITE-1-TO-CLEAR. Set still wins. ----
for i in 0 to N_EVT-1 loop
if wdata(i) = '1' and evt(i) = '0' then evw_v(i) := '0'; end if;
end loop;
when A_CFG_LO => shadow_v(15 downto 0) := wdata;
when A_CFG_HI => shadow_v(31 downto 16) := wdata;
when A_COMMIT =>
-- ---- The whole 32 bits become live in ONE cycle. ----
--
-- Which is the only way to make the change atomic across a
-- 16-bit interface. Writing the halves live means the engine
-- runs on a configuration nobody asked for for one bus cycle,
-- and one bus cycle is thousands of bit times on a USB link.
live_v := shadow_pre;
cm_c <= cm_c + 1;
when A_LOCKED =>
if busy = '1' then
-- Reported, not ignored. A silently dropped write leaves the
-- driver believing a setting took effect, and the next thing
-- it does is built on that belief.
er_v := '1'; ec_v := E_LOCKED;
lk_c <= lk_c + 1;
else
lock_v := wdata;
end if;
when others =>
er_v := '1'; ec_v := E_BADADDR;
ba_c <= ba_c + 1;
end case;
end if;
end if;
ctrl_r <= ctrl_v;
lock_r <= lock_v;
stat_r <= stat_v;
evw_r <= evw_v;
shadow_r <= shadow_v;
live_r <= live_v;
ack_r <= ack_v; er_r <= er_v; ec_r <= ec_v;
rcl_c <= rcl_c + to_unsigned(d_rcl, 32);
end if;
end process;
end architecture;10. Seeing the Two Clear Policies Diverge
The same event, two registers, one read
usb_reg_file — read-to-clear beside write-1-to-clear
10 cyclesAnd the trap, which only the read-to-clear register has:
An event arriving on the read cycle
usb_reg_file — the read-to-clear hazard
10 cycles11. The Testbenches
The oracle is a shadow model written from sections 2 to 5 rather than from the RTL, re-derived every cycle and compared against every output.
The exhaustive claim is over three things at once:
8 addresses x 3 accesses (read / write / idle) x 2 engine states
= 48, all requiredSweeping addresses alone misses the locked register; sweeping accesses alone misses that reading one register and writing it are different operations with different side effects; sweeping the engine state alone misses everything.
Verilog-2005 testbench
`timescale 1ns/1ps
// Testbench for usb_reg_file.
//
// The oracle is a shadow model written from the chapter's rules rather than
// from the RTL, re-derived every cycle and compared against every output.
//
// THE EXHAUSTIVE CLAIM IS OVER (ADDRESS, ACCESS, ENGINE STATE)
//
// A register interface's behaviour is a function of three things at once,
// and every one of them changes the answer:
//
// 8 addresses x 3 accesses (read / write / idle) x 2 engine states
// = 48, all required
//
// Sweeping addresses alone misses the locked register; sweeping accesses
// alone misses that reading one register and writing it are different
// operations with different side effects; sweeping the engine state alone
// misses everything.
module tb_rf_v;
localparam integer N_EVT = 4;
localparam [2:0] A_CTRL = 3'd0, A_STATUS = 3'd1, A_EVENT = 3'd2,
A_CFG_LO = 3'd3, A_CFG_HI = 3'd4, A_COMMIT = 3'd5,
A_LOCKED = 3'd6, A_RSVD = 3'd7;
localparam [1:0] E_NONE = 2'd0, E_LOCKED = 2'd1, E_BADADDR = 2'd2;
localparam [15:0] CTRL_MASK = 16'h00FF;
reg clk = 1'b0, rst_n = 1'b0;
reg [2:0] addr = 3'd0;
reg rd = 1'b0, wr = 1'b0, busy = 1'b0, eot = 1'b0;
reg [15:0] wdata = 16'd0;
reg [N_EVT-1:0] evt = {N_EVT{1'b0}};
wire [15:0] rdata, ctrl;
wire ack, err_pulse;
wire [1:0] err_code;
wire [N_EVT-1:0] status, event_w1c;
wire [31:0] cfg_live, cfg_shadow;
wire [31:0] n_read, n_write, n_commit, n_rc_lost, n_locked, n_badaddr;
usb_reg_file #(.N_EVT(N_EVT)) dut (
.clk(clk), .rst_n(rst_n),
.addr(addr), .rd(rd), .wr(wr), .wdata(wdata), .evt(evt), .busy(busy),
.eot(eot),
.rdata(rdata), .ack(ack), .err_pulse(err_pulse), .err_code(err_code),
.ctrl(ctrl), .status(status), .event_w1c(event_w1c),
.cfg_live(cfg_live), .cfg_shadow(cfg_shadow),
.n_read(n_read), .n_write(n_write), .n_commit(n_commit),
.n_rc_lost(n_rc_lost), .n_locked(n_locked), .n_badaddr(n_badaddr)
);
always #5 clk = ~clk;
// ---------------- the shadow model ----------------
reg [15:0] m_ctrl, m_lock;
reg [N_EVT-1:0] m_stat, m_evw;
reg [31:0] m_shadow, m_live;
reg m_ack, m_er;
reg [1:0] m_ec;
reg [31:0] c_rd, c_wr, c_cm, c_rcl, c_lk, c_ba;
integer errors = 0, checks = 0, steps = 0;
integer k;
// reach: address (8) x access (3) x busy (2)
reg [0:0] reach [0:47];
integer n_reach;
task ck;
input [255:0] nm;
input [31:0] got, exp;
begin
checks = checks + 1;
if (got !== exp) begin
errors = errors + 1;
if (errors < 25)
$display("FAIL t=%0t step=%0d %0s got=%0h exp=%0h",
$time, steps, nm, got, exp);
end
end
endtask
// The expected read data, from the map rather than from the design.
function [15:0] exp_rdata;
input [2:0] a;
begin
case (a)
A_CTRL: exp_rdata = m_ctrl & CTRL_MASK;
A_STATUS: exp_rdata = {{(16-N_EVT){1'b0}}, m_stat};
A_EVENT: exp_rdata = {{(16-N_EVT){1'b0}}, m_evw};
A_CFG_LO: exp_rdata = m_shadow[15:0];
A_CFG_HI: exp_rdata = m_shadow[31:16];
A_COMMIT: exp_rdata = 16'h0000;
A_LOCKED: exp_rdata = m_lock;
default: exp_rdata = 16'h0000;
endcase
end
endfunction
reg [N_EVT-1:0] m_stat_pre;
reg [31:0] m_shadow_pre;
task model_step;
integer i;
begin
m_ack = 1'b0; m_er = 1'b0; m_ec = E_NONE;
if (eot) begin
// nothing
end else begin
// The event sources come FIRST and they SET, before any clear, so an
// event arriving on the same cycle as a read or a write survives it.
for (i = 0; i < N_EVT; i = i + 1)
if (evt[i]) begin m_stat[i] = 1'b1; m_evw[i] = 1'b1; end
if (rd && wr) begin
m_er = 1'b1; m_ec = E_BADADDR; c_ba = c_ba + 1;
end else if (rd) begin
m_ack = 1'b1; c_rd = c_rd + 1;
if (addr == A_RSVD) begin
m_er = 1'b1; m_ec = E_BADADDR; c_ba = c_ba + 1;
end else if (addr == A_STATUS) begin
// READ-TO-CLEAR: clear only the bits that were RETURNED.
for (i = 0; i < N_EVT; i = i + 1)
if (m_stat_pre[i] && !evt[i]) m_stat[i] = 1'b0;
for (i = 0; i < N_EVT; i = i + 1)
if (evt[i] && !m_stat_pre[i]) c_rcl = c_rcl + 1;
end
end else if (wr) begin
m_ack = 1'b1; c_wr = c_wr + 1;
case (addr)
A_CTRL: m_ctrl = wdata & CTRL_MASK;
A_STATUS: begin
m_er = 1'b1; m_ec = E_BADADDR; c_ba = c_ba + 1;
end
A_EVENT: begin
for (i = 0; i < N_EVT; i = i + 1)
if (wdata[i] && !evt[i]) m_evw[i] = 1'b0;
end
A_CFG_LO: m_shadow[15:0] = wdata;
A_CFG_HI: m_shadow[31:16] = wdata;
A_COMMIT: begin
m_live = m_shadow_pre;
c_cm = c_cm + 1;
end
A_LOCKED: begin
if (busy) begin
m_er = 1'b1; m_ec = E_LOCKED; c_lk = c_lk + 1;
end else m_lock = wdata;
end
default: begin
m_er = 1'b1; m_ec = E_BADADDR; c_ba = c_ba + 1;
end
endcase
end
end
end
endtask
task check_out;
begin
ck("ctrl", {16'd0, ctrl}, {16'd0, m_ctrl & CTRL_MASK});
ck("status", {28'd0, status}, {28'd0, m_stat});
ck("event_w1c", {28'd0, event_w1c}, {28'd0, m_evw});
// Compared in halves, to match the VHDL port: to_integer on a full
// 32-bit unsigned overflows VHDL's INTEGER and ABORTS the run, so the
// three benches split it and the check counts line up.
ck("cfg_live lo", {16'd0, cfg_live[15:0]}, {16'd0, m_live[15:0]});
ck("cfg_live hi", {16'd0, cfg_live[31:16]}, {16'd0, m_live[31:16]});
ck("cfg_shadow lo", {16'd0, cfg_shadow[15:0]}, {16'd0, m_shadow[15:0]});
ck("cfg_shadow hi", {16'd0, cfg_shadow[31:16]}, {16'd0, m_shadow[31:16]});
ck("ack", {31'd0, ack}, {31'd0, m_ack});
ck("err_pulse", {31'd0, err_pulse}, {31'd0, m_er});
ck("err_code", {30'd0, err_code}, {30'd0, m_ec});
ck("n_read", n_read, c_rd);
ck("n_write", n_write, c_wr);
ck("n_commit", n_commit, c_cm);
ck("n_rc_lost", n_rc_lost, c_rcl);
ck("n_locked", n_locked, c_lk);
ck("n_badaddr", n_badaddr, c_ba);
// ---- the structural invariant ----
//
// Reserved bits read as zero, always. A driver doing read-modify-write
// on this register writes back exactly what it read, so the two halves
// of the reserved-bit rule have to agree or the register walks
// somewhere new every time round.
ck("ctrl reserved", {16'd0, ctrl & ~CTRL_MASK}, 32'd0);
end
endtask
task step;
begin
// Let the combinational read path settle before sampling it. `addr`
// was driven by a blocking assignment in the same time step, so
// without this the testbench reads the mux output from BEFORE the
// address changed -- and the mismatch looks like a decode bug in the
// design rather than a race in the bench.
#1;
m_stat_pre = m_stat;
m_shadow_pre = m_shadow;
// rdata is combinational and must be sampled BEFORE the edge: it is
// the value the bus returns for this access, not the value the
// register will hold afterwards. Checking it after the edge is
// checking the result of the side effect.
if (rd && !wr && !eot) ck("rdata", {16'd0, rdata}, {16'd0, exp_rdata(addr)});
model_step;
@(posedge clk);
#1;
steps = steps + 1;
check_out;
end
endtask
task acc; input [2:0] a; input r; input w; input [15:0] d;
input [N_EVT-1:0] e; input b;
begin
addr = a; rd = r; wr = w; wdata = d; evt = e; busy = b; eot = 1'b0;
step;
rd = 1'b0; wr = 1'b0; evt = {N_EVT{1'b0}};
end
endtask
task reset_all;
begin
rd = 1'b0; wr = 1'b0; evt = {N_EVT{1'b0}}; busy = 1'b0; eot = 1'b0;
rst_n = 1'b0;
@(posedge clk); #1;
rst_n = 1'b1;
m_ctrl = 16'd0; m_lock = 16'd0;
m_stat = {N_EVT{1'b0}}; m_evw = {N_EVT{1'b0}};
m_stat_pre = {N_EVT{1'b0}};
m_shadow = 32'd0; m_live = 32'd0; m_shadow_pre = 32'd0;
m_ack = 1'b0; m_er = 1'b0; m_ec = E_NONE;
c_rd=0; c_wr=0; c_cm=0; c_rcl=0; c_lk=0; c_ba=0;
@(negedge clk);
end
endtask
integer a, ac, bs, i, w, base, n_race_cases;
initial begin
for (k = 0; k < 48; k = k + 1) reach[k] = 1'b0;
n_race_cases = 0;
repeat (3) @(posedge clk);
rst_n = 1'b1;
@(negedge clk);
reset_all;
// ================= PHASE 1 -- address x access x engine state ========
for (a = 0; a < 8; a = a + 1)
for (ac = 0; ac < 3; ac = ac + 1)
for (bs = 0; bs < 2; bs = bs + 1) begin
acc(a[2:0], (ac == 0), (ac == 1), 16'hA5A5, {N_EVT{1'b0}}, bs[0]);
reach[a * 6 + ac * 2 + bs] = 1'b1;
end
// ================= PHASE 2 -- read-to-clear versus write-1-to-clear ===
//
// The same event feeds both registers. The driver reads STATUS (which
// empties it) and reads EVENT (which does not), and the difference is
// the whole argument.
reset_all;
acc(A_CTRL, 1'b0, 1'b0, 16'd0, 4'b0101, 1'b0); // events 0 and 2 fire
if (status !== 4'b0101 || event_w1c !== 4'b0101) begin
errors = errors + 1;
$display("FAIL the event did not reach both registers");
end
acc(A_STATUS, 1'b1, 1'b0, 16'd0, 4'b0000, 1'b0); // read STATUS
if (status !== 4'b0000) begin
errors = errors + 1;
$display("FAIL read-to-clear did not clear: %b", status);
end
if (event_w1c !== 4'b0101) begin
errors = errors + 1;
$display("FAIL reading STATUS disturbed the W1C register");
end
// ...and the W1C register clears only what the driver writes back.
acc(A_EVENT, 1'b0, 1'b1, 16'h0001, 4'b0000, 1'b0);
if (event_w1c !== 4'b0100) begin
errors = errors + 1;
$display("FAIL W1C cleared the wrong bits: %b", event_w1c);
end
// ================= PHASE 3 -- the two races, over EVERY bit ===========
//
// An event arrives on the very cycle the driver clears. It was not in the
// value returned (read-to-clear) or in the value written back (W1C), so
// destroying it loses an interrupt nobody ever saw -- and nothing else in
// the system will ever mention it again.
//
// The first version of this phase drove ONE bit with ONE pre-state, and
// the mutation that clears the whole register died on TWO checks. Both
// races are now driven for every event bit and both pre-states: whether
// the bit was already set before the racing event or not.
//
// 4 bits x 2 pre-states x 2 clear policies = 16 scenarios
n_race_cases = 0;
for (a = 0; a < N_EVT; a = a + 1)
for (bs = 0; bs < 2; bs = bs + 1)
for (ac = 0; ac < 2; ac = ac + 1) begin
reset_all;
// optionally pre-set the bit that is about to race
if (bs == 1)
acc(A_CTRL, 1'b0, 1'b0, 16'd0, (4'd1 << a), 1'b0);
// ...and always set a DIFFERENT bit, so there is something for the
// clear to legitimately remove. Without it a correct design and a
// design that clears everything are indistinguishable.
acc(A_CTRL, 1'b0, 1'b0, 16'd0, (4'd1 << ((a + 1) % N_EVT)), 1'b0);
base = c_rcl;
if (ac == 0) begin
// read-to-clear, with the event arriving on the read cycle
acc(A_STATUS, 1'b1, 1'b0, 16'd0, (4'd1 << a), 1'b0);
if (!status[a]) begin
errors = errors + 1;
$display("FAIL RC race bit %0d pre=%0d: the event was destroyed",
a, bs);
end
// the OTHER bit was returned, so it must be gone
if (status[(a + 1) % N_EVT] && ((a + 1) % N_EVT != a)) begin
errors = errors + 1;
$display("FAIL RC race bit %0d: the returned bit survived", a);
end
// and the hazard is counted exactly when the racing bit was NOT
// already set -- because only then was it absent from the value
// the driver got
if (c_rcl != base + ((bs == 0) ? 1 : 0)) begin
errors = errors + 1;
$display("FAIL RC hazard count bit %0d pre=%0d: %0d",
a, bs, c_rcl - base);
end
end else begin
// write-1-to-clear, with the event arriving on the write cycle
acc(A_EVENT, 1'b0, 1'b1, (16'd1 << a), (4'd1 << a), 1'b0);
if (!event_w1c[a]) begin
errors = errors + 1;
$display("FAIL W1C race bit %0d pre=%0d: set lost the race",
a, bs);
end
// ...and the bit nobody wrote back is untouched
if (!event_w1c[(a + 1) % N_EVT]) begin
errors = errors + 1;
$display("FAIL W1C race bit %0d: an unwritten bit was cleared",
a);
end
// ...and with no concurrent event the same write does clear it
acc(A_EVENT, 1'b0, 1'b1, (16'd1 << a), 4'd0, 1'b0);
if (event_w1c[a]) begin
errors = errors + 1;
$display("FAIL W1C bit %0d: a plain write did not clear", a);
end
end
n_race_cases = n_race_cases + 1;
end
if (n_race_cases != 16) begin
errors = errors + 1;
$display("FAIL race scenarios %0d, expected 16", n_race_cases);
end
// ================= PHASE 4 -- the shadow is ATOMIC ====================
//
// Two halves staged, then applied in one cycle. Between the two stores
// the live configuration must not change at all -- for one bus cycle the
// engine would otherwise be running on a value nobody asked for, and one
// bus cycle is thousands of bit times on a USB link.
reset_all;
acc(A_CFG_LO, 1'b0, 1'b1, 16'hBEEF, 4'b0000, 1'b0);
if (cfg_live !== 32'd0) begin
errors = errors + 1;
$display("FAIL writing CFG_LO changed the live configuration");
end
acc(A_CFG_HI, 1'b0, 1'b1, 16'hDEAD, 4'b0000, 1'b0);
if (cfg_live !== 32'd0) begin
errors = errors + 1;
$display("FAIL writing CFG_HI changed the live configuration");
end
if (cfg_shadow !== 32'hDEAD_BEEF) begin
errors = errors + 1;
$display("FAIL the shadow is 0x%0h, expected 0xDEADBEEF", cfg_shadow);
end
base = n_commit;
acc(A_COMMIT, 1'b0, 1'b1, 16'h0001, 4'b0000, 1'b0);
if (cfg_live !== 32'hDEAD_BEEF) begin
errors = errors + 1;
$display("FAIL the commit did not apply the shadow: 0x%0h", cfg_live);
end
if (n_commit != base + 1) begin
errors = errors + 1;
$display("FAIL the commit was not counted");
end
// ================= PHASE 5 -- a rejected write is REPORTED ============
//
// Silently dropping it is the worst option: the driver believes the
// setting took effect and the next thing it does is built on that
// belief.
reset_all;
acc(A_LOCKED, 1'b0, 1'b1, 16'h1234, 4'b0000, 1'b0); // idle: accepted
if (rdata !== 16'h0000) ; // rdata is for the access being driven
base = n_locked;
acc(A_LOCKED, 1'b0, 1'b1, 16'h5678, 4'b0000, 1'b1); // busy: rejected
if (n_locked != base + 1) begin
errors = errors + 1;
$display("FAIL a write while busy was not reported");
end
acc(A_LOCKED, 1'b1, 1'b0, 16'd0, 4'b0000, 1'b0);
if (rdata !== 16'h1234) begin
errors = errors + 1;
$display("FAIL the rejected write took effect anyway: 0x%0h", rdata);
end
// ================= PHASE 6 -- reserved bits ==========================
//
// They read as zero AND ignore writes. The two halves of that rule have
// to agree, because a driver doing read-modify-write writes back exactly
// what it read.
reset_all;
acc(A_CTRL, 1'b0, 1'b1, 16'hFFFF, 4'b0000, 1'b0);
if (ctrl !== 16'h00FF) begin
errors = errors + 1;
$display("FAIL reserved bits took a write: 0x%0h", ctrl);
end
acc(A_CTRL, 1'b1, 1'b0, 16'd0, 4'b0000, 1'b0);
if (rdata !== 16'h00FF) begin
errors = errors + 1;
$display("FAIL reserved bits did not read as zero: 0x%0h", rdata);
end
// ...and a read-modify-write cycle is a fixed point: write back what was
// read and nothing changes.
acc(A_CTRL, 1'b0, 1'b1, 16'h00FF, 4'b0000, 1'b0);
if (ctrl !== 16'h00FF) begin
errors = errors + 1;
$display("FAIL a read-modify-write cycle moved the register");
end
// ================= PHASE 7 -- the two rejections ======================
reset_all;
base = n_badaddr;
acc(A_RSVD, 1'b1, 1'b0, 16'd0, 4'b0000, 1'b0); // read a hole
if (n_badaddr != base + 1) begin
errors = errors + 1;
$display("FAIL reading a hole in the map was not reported");
end
base = n_badaddr;
acc(A_STATUS, 1'b0, 1'b1, 16'hFFFF, 4'b0000, 1'b0); // write an RC reg
if (n_badaddr != base + 1) begin
errors = errors + 1;
$display("FAIL writing a read-to-clear register was not reported");
end
base = n_badaddr;
acc(A_CTRL, 1'b1, 1'b1, 16'd0, 4'b0000, 1'b0); // read AND write
if (n_badaddr != base + 1) begin
errors = errors + 1;
$display("FAIL a simultaneous read and write was not reported");
end
// The random phase is switchable, because a mutation score is only
// interesting once it is DECOMPOSED. The directed phases already reach
// every situation the exhaustiveness proof requires, so nothing in that
// claim depends on it.
`ifndef DIRECTED_ONLY
// ================= PHASE 8 -- random =================================
//
// The event pattern is deliberately correlated with the accesses: an
// event drawn independently of the read almost never lands on the same
// cycle, and that coincidence is the read-to-clear hazard.
reset_all;
for (i = 0; i < 30000; i = i + 1) begin
w = $unsigned($random) % 100;
acc($unsigned($random) % 8,
(w < 40), (w >= 40) && (w < 85),
$unsigned($random) % 65536,
(w < 55) ? ($unsigned($random) % 16) : 4'd0,
($unsigned($random) % 100) < 25);
end
`endif
// ================= the exhaustiveness proof ==========================
n_reach = 0;
for (k = 0; k < 48; k = k + 1) n_reach = n_reach + reach[k];
if (n_reach != 48) begin
errors = errors + 1;
$display("FAIL address x access x busy reach %0d/48", n_reach);
for (k = 0; k < 48; k = k + 1)
if (!reach[k])
$display(" unreached addr=%0d access=%0d busy=%0d",
k / 6, (k % 6) / 2, k % 2);
end
$display("steps=%0d checks=%0d reach=%0d/48 errors=%0d",
steps, checks, n_reach, errors);
$display("read=%0d write=%0d commit=%0d rc_lost=%0d locked=%0d badaddr=%0d",
n_read, n_write, n_commit, n_rc_lost, n_locked, n_badaddr);
$display("%0s: %0d errors in %0d checks",
(errors == 0) ? "PASS" : "FAIL", errors, checks);
$finish;
end
endmoduleSystemVerilog testbench
`timescale 1ns/1ps
// Testbench for usb_reg_file.
//
// The oracle is a shadow model written from the chapter's rules rather than
// from the RTL, re-derived every cycle and compared against every output.
//
// THE EXHAUSTIVE CLAIM IS OVER (ADDRESS, ACCESS, ENGINE STATE)
//
// A register interface's behaviour is a function of three things at once,
// and every one of them changes the answer:
//
// 8 addresses x 3 accesses (read / write / idle) x 2 engine states
// = 48, all required
//
// Sweeping addresses alone misses the locked register; sweeping accesses
// alone misses that reading one register and writing it are different
// operations with different side effects; sweeping the engine state alone
// misses everything.
module tb_rf_sv;
import usb_reg_pkg::*;
localparam int N_EVT = 4;
localparam logic [15:0] CTRL_MASK = 16'h00FF;
logic clk = 1'b0, rst_n = 1'b0;
reg_addr_e addr = A_CTRL;
logic rd = 1'b0, wr = 1'b0, busy = 1'b0, eot = 1'b0;
logic [15:0] wdata = 16'd0;
logic [N_EVT-1:0] evt = '0;
logic [15:0] rdata, ctrl;
logic ack, err_pulse;
reg_err_e err_code;
logic [N_EVT-1:0] status, event_w1c;
logic [31:0] cfg_live, cfg_shadow;
logic [31:0] n_read, n_write, n_commit, n_rc_lost, n_locked, n_badaddr;
usb_reg_file #(.N_EVT(N_EVT)) dut (
.clk(clk), .rst_n(rst_n),
.addr(addr), .rd(rd), .wr(wr), .wdata(wdata), .evt(evt), .busy(busy),
.eot(eot),
.rdata(rdata), .ack(ack), .err_pulse(err_pulse), .err_code(err_code),
.ctrl(ctrl), .status(status), .event_w1c(event_w1c),
.cfg_live(cfg_live), .cfg_shadow(cfg_shadow),
.n_read(n_read), .n_write(n_write), .n_commit(n_commit),
.n_rc_lost(n_rc_lost), .n_locked(n_locked), .n_badaddr(n_badaddr)
);
always #5 clk = ~clk;
// ---------------- the shadow model ----------------
logic [15:0] m_ctrl, m_lock;
logic [N_EVT-1:0] m_stat, m_evw;
logic [31:0] m_shadow, m_live;
logic m_ack, m_er;
reg_err_e m_ec;
int unsigned c_rd, c_wr, c_cm, c_rcl, c_lk, c_ba;
int errors = 0, checks = 0, steps = 0;
int k;
// reach: address (8) x access (3) x busy (2)
bit reach [48];
int n_reach;
task automatic ck(string nm, int unsigned got, int unsigned exp);
checks++;
if (got !== exp) begin
errors++;
if (errors < 25)
$display("FAIL t=%0t step=%0d %0s got=%0h exp=%0h",
$time, steps, nm, got, exp);
end
endtask
// The expected read data, from the map rather than from the design.
function automatic logic [15:0] exp_rdata(input reg_addr_e a);
begin
case (a)
A_CTRL: exp_rdata = m_ctrl & CTRL_MASK;
A_STATUS: exp_rdata = {{(16-N_EVT){1'b0}}, m_stat};
A_EVENT: exp_rdata = {{(16-N_EVT){1'b0}}, m_evw};
A_CFG_LO: exp_rdata = m_shadow[15:0];
A_CFG_HI: exp_rdata = m_shadow[31:16];
A_COMMIT: exp_rdata = 16'h0000;
A_LOCKED: exp_rdata = m_lock;
default: exp_rdata = 16'h0000;
endcase
end
endfunction
logic [N_EVT-1:0] m_stat_pre;
logic [31:0] m_shadow_pre;
task automatic model_step;
int i;
begin
m_ack = 1'b0; m_er = 1'b0; m_ec = E_NONE;
if (eot) begin
// nothing
end else begin
// The event sources come FIRST and they SET, before any clear, so an
// event arriving on the same cycle as a read or a write survives it.
for (i = 0; i < N_EVT; i = i + 1)
if (evt[i]) begin m_stat[i] = 1'b1; m_evw[i] = 1'b1; end
if (rd && wr) begin
m_er = 1'b1; m_ec = E_BADADDR; c_ba = c_ba + 1;
end else if (rd) begin
m_ack = 1'b1; c_rd = c_rd + 1;
if (addr == A_RSVD) begin
m_er = 1'b1; m_ec = E_BADADDR; c_ba = c_ba + 1;
end else if (addr == A_STATUS) begin
// READ-TO-CLEAR: clear only the bits that were RETURNED.
for (i = 0; i < N_EVT; i = i + 1)
if (m_stat_pre[i] && !evt[i]) m_stat[i] = 1'b0;
for (i = 0; i < N_EVT; i = i + 1)
if (evt[i] && !m_stat_pre[i]) c_rcl = c_rcl + 1;
end
end else if (wr) begin
m_ack = 1'b1; c_wr = c_wr + 1;
case (addr)
A_CTRL: m_ctrl = wdata & CTRL_MASK;
A_STATUS: begin
m_er = 1'b1; m_ec = E_BADADDR; c_ba = c_ba + 1;
end
A_EVENT: begin
for (i = 0; i < N_EVT; i = i + 1)
if (wdata[i] && !evt[i]) m_evw[i] = 1'b0;
end
A_CFG_LO: m_shadow[15:0] = wdata;
A_CFG_HI: m_shadow[31:16] = wdata;
A_COMMIT: begin
m_live = m_shadow_pre;
c_cm = c_cm + 1;
end
A_LOCKED: begin
if (busy) begin
m_er = 1'b1; m_ec = E_LOCKED; c_lk = c_lk + 1;
end else m_lock = wdata;
end
default: begin
m_er = 1'b1; m_ec = E_BADADDR; c_ba = c_ba + 1;
end
endcase
end
end
end
endtask
task automatic check_out;
begin
ck("ctrl", {16'd0, ctrl}, {16'd0, m_ctrl & CTRL_MASK});
ck("status", {28'd0, status}, {28'd0, m_stat});
ck("event_w1c", {28'd0, event_w1c}, {28'd0, m_evw});
// Compared in halves, to match the VHDL port: to_integer on a full
// 32-bit unsigned overflows VHDL's INTEGER and ABORTS the run, so the
// three benches split it and the check counts line up.
ck("cfg_live lo", {16'd0, cfg_live[15:0]}, {16'd0, m_live[15:0]});
ck("cfg_live hi", {16'd0, cfg_live[31:16]}, {16'd0, m_live[31:16]});
ck("cfg_shadow lo", {16'd0, cfg_shadow[15:0]}, {16'd0, m_shadow[15:0]});
ck("cfg_shadow hi", {16'd0, cfg_shadow[31:16]}, {16'd0, m_shadow[31:16]});
ck("ack", {31'd0, ack}, {31'd0, m_ack});
ck("err_pulse", {31'd0, err_pulse}, {31'd0, m_er});
ck("err_code", err_code, m_ec);
ck("n_read", n_read, c_rd);
ck("n_write", n_write, c_wr);
ck("n_commit", n_commit, c_cm);
ck("n_rc_lost", n_rc_lost, c_rcl);
ck("n_locked", n_locked, c_lk);
ck("n_badaddr", n_badaddr, c_ba);
// ---- the structural invariant ----
//
// Reserved bits read as zero, always. A driver doing read-modify-write
// on this register writes back exactly what it read, so the two halves
// of the reserved-bit rule have to agree or the register walks
// somewhere new every time round.
ck("ctrl reserved", {16'd0, ctrl & ~CTRL_MASK}, 32'd0);
end
endtask
task automatic step;
begin
// Let the combinational read path settle before sampling it. `addr`
// was driven by a blocking assignment in the same time step, so
// without this the testbench reads the mux output from BEFORE the
// address changed -- and the mismatch looks like a decode bug in the
// design rather than a race in the bench.
#1;
m_stat_pre = m_stat;
m_shadow_pre = m_shadow;
// rdata is combinational and must be sampled BEFORE the edge: it is
// the value the bus returns for this access, not the value the
// register will hold afterwards. Checking it after the edge is
// checking the result of the side effect.
if (rd && !wr && !eot) ck("rdata", {16'd0, rdata}, {16'd0, exp_rdata(addr)});
model_step;
@(posedge clk);
#1;
steps = steps + 1;
check_out;
end
endtask
task automatic acc(reg_addr_e a, logic r, logic w, logic [15:0] d,
logic [N_EVT-1:0] e, logic b);
begin
addr = a; rd = r; wr = w; wdata = d; evt = e; busy = b; eot = 1'b0;
step;
rd = 1'b0; wr = 1'b0; evt = '0;
end
endtask
task automatic reset_all;
begin
rd = 1'b0; wr = 1'b0; evt = '0; busy = 1'b0; eot = 1'b0;
rst_n = 1'b0;
@(posedge clk); #1;
rst_n = 1'b1;
m_ctrl = 16'd0; m_lock = 16'd0;
m_stat = '0; m_evw = '0;
m_stat_pre = '0;
m_shadow = 32'd0; m_live = 32'd0; m_shadow_pre = 32'd0;
m_ack = 1'b0; m_er = 1'b0; m_ec = E_NONE;
c_rd=0; c_wr=0; c_cm=0; c_rcl=0; c_lk=0; c_ba=0;
@(negedge clk);
end
endtask
integer a, ac, bs, i, w, base, n_race_cases;
initial begin
foreach (reach[q]) reach[q] = 1'b0;
n_race_cases = 0;
repeat (3) @(posedge clk);
rst_n = 1'b1;
@(negedge clk);
reset_all;
// ================= PHASE 1 -- address x access x engine state ========
for (a = 0; a < 8; a = a + 1)
for (ac = 0; ac < 3; ac = ac + 1)
for (bs = 0; bs < 2; bs = bs + 1) begin
acc(reg_addr_e'(a[2:0]), (ac == 0), (ac == 1), 16'hA5A5, '0, bs[0]);
reach[a * 6 + ac * 2 + bs] = 1'b1;
end
// ================= PHASE 2 -- read-to-clear versus write-1-to-clear ===
//
// The same event feeds both registers. The driver reads STATUS (which
// empties it) and reads EVENT (which does not), and the difference is
// the whole argument.
reset_all;
acc(A_CTRL, 1'b0, 1'b0, 16'd0, 4'b0101, 1'b0); // events 0 and 2 fire
if (status !== 4'b0101 || event_w1c !== 4'b0101) begin
errors = errors + 1;
$display("FAIL the event did not reach both registers");
end
acc(A_STATUS, 1'b1, 1'b0, 16'd0, 4'b0000, 1'b0); // read STATUS
if (status !== 4'b0000) begin
errors = errors + 1;
$display("FAIL read-to-clear did not clear: %b", status);
end
if (event_w1c !== 4'b0101) begin
errors = errors + 1;
$display("FAIL reading STATUS disturbed the W1C register");
end
// ...and the W1C register clears only what the driver writes back.
acc(A_EVENT, 1'b0, 1'b1, 16'h0001, 4'b0000, 1'b0);
if (event_w1c !== 4'b0100) begin
errors = errors + 1;
$display("FAIL W1C cleared the wrong bits: %b", event_w1c);
end
// ================= PHASE 3 -- the two races, over EVERY bit ===========
//
// An event arrives on the very cycle the driver clears. It was not in the
// value returned (read-to-clear) or in the value written back (W1C), so
// destroying it loses an interrupt nobody ever saw -- and nothing else in
// the system will ever mention it again.
//
// The first version of this phase drove ONE bit with ONE pre-state, and
// the mutation that clears the whole register died on TWO checks. Both
// races are now driven for every event bit and both pre-states: whether
// the bit was already set before the racing event or not.
//
// 4 bits x 2 pre-states x 2 clear policies = 16 scenarios
n_race_cases = 0;
for (a = 0; a < N_EVT; a = a + 1)
for (bs = 0; bs < 2; bs = bs + 1)
for (ac = 0; ac < 2; ac = ac + 1) begin
reset_all;
// optionally pre-set the bit that is about to race
if (bs == 1)
acc(A_CTRL, 1'b0, 1'b0, 16'd0, 4'(4'd1 << a), 1'b0);
// ...and always set a DIFFERENT bit, so there is something for the
// clear to legitimately remove. Without it a correct design and a
// design that clears everything are indistinguishable.
acc(A_CTRL, 1'b0, 1'b0, 16'd0, 4'(4'd1 << ((a + 1) % N_EVT)), 1'b0);
base = int'(c_rcl);
if (ac == 0) begin
// read-to-clear, with the event arriving on the read cycle
acc(A_STATUS, 1'b1, 1'b0, 16'd0, 4'(4'd1 << a), 1'b0);
if (!status[a]) begin
errors = errors + 1;
$display("FAIL RC race bit %0d pre=%0d: the event was destroyed",
a, bs);
end
// the OTHER bit was returned, so it must be gone
if (status[(a + 1) % N_EVT] && ((a + 1) % N_EVT != a)) begin
errors = errors + 1;
$display("FAIL RC race bit %0d: the returned bit survived", a);
end
// and the hazard is counted exactly when the racing bit was NOT
// already set -- because only then was it absent from the value
// the driver got
if (int'(c_rcl) != base + ((bs == 0) ? 1 : 0)) begin
errors = errors + 1;
$display("FAIL RC hazard count bit %0d pre=%0d: %0d",
a, bs, int'(c_rcl) - base);
end
end else begin
// write-1-to-clear, with the event arriving on the write cycle
acc(A_EVENT, 1'b0, 1'b1, 16'(16'd1 << a), 4'(4'd1 << a), 1'b0);
if (!event_w1c[a]) begin
errors = errors + 1;
$display("FAIL W1C race bit %0d pre=%0d: set lost the race",
a, bs);
end
// ...and the bit nobody wrote back is untouched
if (!event_w1c[(a + 1) % N_EVT]) begin
errors = errors + 1;
$display("FAIL W1C race bit %0d: an unwritten bit was cleared",
a);
end
// ...and with no concurrent event the same write does clear it
acc(A_EVENT, 1'b0, 1'b1, 16'(16'd1 << a), 4'd0, 1'b0);
if (event_w1c[a]) begin
errors = errors + 1;
$display("FAIL W1C bit %0d: a plain write did not clear", a);
end
end
n_race_cases = n_race_cases + 1;
end
if (n_race_cases != 16) begin
errors = errors + 1;
$display("FAIL race scenarios %0d, expected 16", n_race_cases);
end
// ================= PHASE 4 -- the shadow is ATOMIC ====================
//
// Two halves staged, then applied in one cycle. Between the two stores
// the live configuration must not change at all -- for one bus cycle the
// engine would otherwise be running on a value nobody asked for, and one
// bus cycle is thousands of bit times on a USB link.
reset_all;
acc(A_CFG_LO, 1'b0, 1'b1, 16'hBEEF, 4'b0000, 1'b0);
if (cfg_live !== 32'd0) begin
errors = errors + 1;
$display("FAIL writing CFG_LO changed the live configuration");
end
acc(A_CFG_HI, 1'b0, 1'b1, 16'hDEAD, 4'b0000, 1'b0);
if (cfg_live !== 32'd0) begin
errors = errors + 1;
$display("FAIL writing CFG_HI changed the live configuration");
end
if (cfg_shadow !== 32'hDEAD_BEEF) begin
errors = errors + 1;
$display("FAIL the shadow is 0x%0h, expected 0xDEADBEEF", cfg_shadow);
end
base = int'(n_commit);
acc(A_COMMIT, 1'b0, 1'b1, 16'h0001, 4'b0000, 1'b0);
if (cfg_live !== 32'hDEAD_BEEF) begin
errors = errors + 1;
$display("FAIL the commit did not apply the shadow: 0x%0h", cfg_live);
end
if (int'(n_commit) != base + 1) begin
errors = errors + 1;
$display("FAIL the commit was not counted");
end
// ================= PHASE 5 -- a rejected write is REPORTED ============
//
// Silently dropping it is the worst option: the driver believes the
// setting took effect and the next thing it does is built on that
// belief.
reset_all;
acc(A_LOCKED, 1'b0, 1'b1, 16'h1234, 4'b0000, 1'b0); // idle: accepted
if (rdata !== 16'h0000) ; // rdata is for the access being driven
base = int'(n_locked);
acc(A_LOCKED, 1'b0, 1'b1, 16'h5678, 4'b0000, 1'b1); // busy: rejected
if (int'(n_locked) != base + 1) begin
errors = errors + 1;
$display("FAIL a write while busy was not reported");
end
acc(A_LOCKED, 1'b1, 1'b0, 16'd0, 4'b0000, 1'b0);
if (rdata !== 16'h1234) begin
errors = errors + 1;
$display("FAIL the rejected write took effect anyway: 0x%0h", rdata);
end
// ================= PHASE 6 -- reserved bits ==========================
//
// They read as zero AND ignore writes. The two halves of that rule have
// to agree, because a driver doing read-modify-write writes back exactly
// what it read.
reset_all;
acc(A_CTRL, 1'b0, 1'b1, 16'hFFFF, 4'b0000, 1'b0);
if (ctrl !== 16'h00FF) begin
errors = errors + 1;
$display("FAIL reserved bits took a write: 0x%0h", ctrl);
end
acc(A_CTRL, 1'b1, 1'b0, 16'd0, 4'b0000, 1'b0);
if (rdata !== 16'h00FF) begin
errors = errors + 1;
$display("FAIL reserved bits did not read as zero: 0x%0h", rdata);
end
// ...and a read-modify-write cycle is a fixed point: write back what was
// read and nothing changes.
acc(A_CTRL, 1'b0, 1'b1, 16'h00FF, 4'b0000, 1'b0);
if (ctrl !== 16'h00FF) begin
errors = errors + 1;
$display("FAIL a read-modify-write cycle moved the register");
end
// ================= PHASE 7 -- the two rejections ======================
reset_all;
base = int'(n_badaddr);
acc(A_RSVD, 1'b1, 1'b0, 16'd0, 4'b0000, 1'b0); // read a hole
if (int'(n_badaddr) != base + 1) begin
errors = errors + 1;
$display("FAIL reading a hole in the map was not reported");
end
base = int'(n_badaddr);
acc(A_STATUS, 1'b0, 1'b1, 16'hFFFF, 4'b0000, 1'b0); // write an RC reg
if (int'(n_badaddr) != base + 1) begin
errors = errors + 1;
$display("FAIL writing a read-to-clear register was not reported");
end
base = int'(n_badaddr);
acc(A_CTRL, 1'b1, 1'b1, 16'd0, 4'b0000, 1'b0); // read AND write
if (int'(n_badaddr) != base + 1) begin
errors = errors + 1;
$display("FAIL a simultaneous read and write was not reported");
end
// The random phase is switchable, because a mutation score is only
// interesting once it is DECOMPOSED. The directed phases already reach
// every situation the exhaustiveness proof requires, so nothing in that
// claim depends on it.
`ifndef DIRECTED_ONLY
// ================= PHASE 8 -- random =================================
//
// The event pattern is deliberately correlated with the accesses: an
// event drawn independently of the read almost never lands on the same
// cycle, and that coincidence is the read-to-clear hazard.
reset_all;
for (i = 0; i < 30000; i = i + 1) begin
w = $unsigned($random) % 100;
acc(reg_addr_e'($unsigned($random) % 8),
(w < 40), (w >= 40) && (w < 85),
16'($unsigned($random) % 65536),
(w < 55) ? 4'($unsigned($random) % 16) : 4'd0,
($unsigned($random) % 100) < 25);
end
`endif
// ================= the exhaustiveness proof ==========================
n_reach = 0;
for (k = 0; k < 48; k = k + 1) n_reach = n_reach + reach[k];
if (n_reach != 48) begin
errors = errors + 1;
$display("FAIL address x access x busy reach %0d/48", n_reach);
for (k = 0; k < 48; k = k + 1)
if (!reach[k])
$display(" unreached addr=%0d access=%0d busy=%0d",
k / 6, (k % 6) / 2, k % 2);
end
$display("steps=%0d checks=%0d reach=%0d/48 errors=%0d",
steps, checks, n_reach, errors);
$display("read=%0d write=%0d commit=%0d rc_lost=%0d locked=%0d badaddr=%0d",
n_read, n_write, n_commit, n_rc_lost, n_locked, n_badaddr);
$display("%0s: %0d errors in %0d checks",
(errors == 0) ? "PASS" : "FAIL", errors, checks);
$finish;
end
endmoduleVHDL-2008 testbench
-- Testbench for usb_reg_file (VHDL-2008).
--
-- The oracle is a shadow model held in process variables and written from the
-- chapter's rules rather than from the RTL, re-derived every cycle and
-- compared against every output.
--
-- THE EXHAUSTIVE CLAIM IS OVER (ADDRESS, ACCESS, ENGINE STATE)
--
-- A register interface's behaviour is a function of three things at once, and
-- every one of them changes the answer:
--
-- 8 addresses x 3 accesses (read / write / idle) x 2 engine states
-- = 48, all required
--
-- Sweeping addresses alone misses the locked register; sweeping accesses
-- alone misses that reading one register and writing it are different
-- operations with different side effects.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
use std.textio.all;
use work.usb_reg_pkg.all;
entity tb_rf_vhdl is
end entity;
architecture sim of tb_rf_vhdl is
constant N_EVT : integer := 4;
signal clk : std_logic := '0';
signal rst_n : std_logic := '0';
signal addr : std_logic_vector(2 downto 0) := "000";
signal rd, wr, busy, eot : std_logic := '0';
signal wdata : std_logic_vector(15 downto 0) := (others => '0');
signal evt : std_logic_vector(N_EVT-1 downto 0) := (others => '0');
signal rdata_s, ctrl_s : std_logic_vector(15 downto 0);
signal ack_s, err_s : std_logic;
signal ec_s : std_logic_vector(1 downto 0);
signal status_s, evw_s : std_logic_vector(N_EVT-1 downto 0);
signal live_s, shadow_s : std_logic_vector(31 downto 0);
signal n_rd_s, n_wr_s, n_cm_s : unsigned(31 downto 0);
signal n_rcl_s, n_lk_s, n_ba_s : unsigned(31 downto 0);
signal done : boolean := false;
type int_array is array (natural range <>) of integer;
begin
clk <= not clk after 5 ns when not done else '0';
dut : entity work.usb_reg_file
generic map (N_EVT => N_EVT)
port map (
clk => clk, rst_n => rst_n,
addr => addr, rd => rd, wr => wr, wdata => wdata, evt => evt,
busy => busy, eot => eot,
rdata => rdata_s, ack => ack_s, err_pulse => err_s, err_code => ec_s,
ctrl => ctrl_s, status => status_s, event_w1c => evw_s,
cfg_live => live_s, cfg_shadow => shadow_s,
n_read => n_rd_s, n_write => n_wr_s, n_commit => n_cm_s,
n_rc_lost => n_rcl_s, n_locked => n_lk_s, n_badaddr => n_ba_s
);
stim : process
variable m_ctrl, m_lock : std_logic_vector(15 downto 0) := (others => '0');
variable m_stat, m_evw : std_logic_vector(N_EVT-1 downto 0)
:= (others => '0');
variable m_stat_pre : std_logic_vector(N_EVT-1 downto 0)
:= (others => '0');
variable m_shadow, m_live : std_logic_vector(31 downto 0)
:= (others => '0');
variable m_shadow_pre : std_logic_vector(31 downto 0) := (others => '0');
variable m_ack, m_er : std_logic := '0';
variable m_ec : std_logic_vector(1 downto 0) := E_NONE;
variable c_rd, c_wr, c_cm, c_rcl, c_lk, c_ba : integer := 0;
variable errors, checks, steps : integer := 0;
variable reach : int_array(0 to 47) := (others => 0);
variable n_reach, base, w_v, n_race_cases : integer := 0;
variable o_v : std_logic_vector(N_EVT-1 downto 0) := (others => '0');
variable a_v : std_logic_vector(2 downto 0) := "000";
variable d_v : std_logic_vector(15 downto 0) := (others => '0');
variable e_v : std_logic_vector(N_EVT-1 downto 0) := (others => '0');
variable r_v, wv_v, b_v : std_logic := '0';
variable ln : line;
-- A deterministic LFSR, so a rerun reproduces exactly the same traffic.
variable lfsr : unsigned(31 downto 0) := x"5EA51DE5";
impure function rnd_nat return integer is
begin
lfsr := lfsr(30 downto 0) &
(lfsr(31) xor lfsr(21) xor lfsr(1) xor lfsr(0));
return to_integer(lfsr(29 downto 0));
end function;
procedure ck (nm : string; got, exp : integer) is
begin
checks := checks + 1;
if got /= exp then
errors := errors + 1;
if errors < 25 then
write(ln, string'("FAIL step=") & integer'image(steps) & " " & nm
& " got=" & integer'image(got)
& " exp=" & integer'image(exp));
writeline(output, ln);
end if;
end if;
end procedure;
procedure fail (nm : string) is
begin
errors := errors + 1;
if errors < 25 then
write(ln, string'("FAIL ") & nm); writeline(output, ln);
end if;
end procedure;
function sl2i (s : std_logic) return integer is
begin
if s = '1' then return 1; else return 0; end if;
end function;
-- The expected read data, from the map rather than from the design.
impure function exp_rdata (a : std_logic_vector(2 downto 0))
return std_logic_vector is
begin
if a = A_CTRL then return m_ctrl and CTRL_MASK;
elsif a = A_STATUS then return (15 downto N_EVT => '0') & m_stat;
elsif a = A_EVENT then return (15 downto N_EVT => '0') & m_evw;
elsif a = A_CFG_LO then return m_shadow(15 downto 0);
elsif a = A_CFG_HI then return m_shadow(31 downto 16);
elsif a = A_LOCKED then return m_lock;
else return x"0000";
end if;
end function;
procedure model_step is
begin
m_ack := '0'; m_er := '0'; m_ec := E_NONE;
if eot = '1' then
null;
else
-- The event sources come FIRST and they SET, before any clear.
for i in 0 to N_EVT-1 loop
if evt(i) = '1' then m_stat(i) := '1'; m_evw(i) := '1'; end if;
end loop;
if rd = '1' and wr = '1' then
m_er := '1'; m_ec := E_BADADDR; c_ba := c_ba + 1;
elsif rd = '1' then
m_ack := '1'; c_rd := c_rd + 1;
if addr = A_RSVD then
m_er := '1'; m_ec := E_BADADDR; c_ba := c_ba + 1;
elsif addr = A_STATUS then
-- READ-TO-CLEAR: clear only the bits that were RETURNED.
for i in 0 to N_EVT-1 loop
if m_stat_pre(i) = '1' and evt(i) = '0' then m_stat(i) := '0'; end if;
end loop;
for i in 0 to N_EVT-1 loop
if evt(i) = '1' and m_stat_pre(i) = '0' then c_rcl := c_rcl + 1; end if;
end loop;
end if;
elsif wr = '1' then
m_ack := '1'; c_wr := c_wr + 1;
if addr = A_CTRL then
m_ctrl := wdata and CTRL_MASK;
elsif addr = A_STATUS then
m_er := '1'; m_ec := E_BADADDR; c_ba := c_ba + 1;
elsif addr = A_EVENT then
for i in 0 to N_EVT-1 loop
if wdata(i) = '1' and evt(i) = '0' then m_evw(i) := '0'; end if;
end loop;
elsif addr = A_CFG_LO then
m_shadow(15 downto 0) := wdata;
elsif addr = A_CFG_HI then
m_shadow(31 downto 16) := wdata;
elsif addr = A_COMMIT then
m_live := m_shadow_pre;
c_cm := c_cm + 1;
elsif addr = A_LOCKED then
if busy = '1' then
m_er := '1'; m_ec := E_LOCKED; c_lk := c_lk + 1;
else
m_lock := wdata;
end if;
else
m_er := '1'; m_ec := E_BADADDR; c_ba := c_ba + 1;
end if;
end if;
end if;
end procedure;
procedure check_out is
begin
ck("ctrl", to_integer(unsigned(ctrl_s)),
to_integer(unsigned(m_ctrl)));
ck("status", to_integer(unsigned(status_s)),
to_integer(unsigned(m_stat)));
ck("event_w1c", to_integer(unsigned(evw_s)),
to_integer(unsigned(m_evw)));
-- Compared in halves. to_integer on a full 32-bit unsigned overflows
-- VHDL's INTEGER and ABORTS the run rather than returning a wrong
-- answer -- which is better than wrapping and much worse than a
-- failing check, because the run stops before it has told you
-- anything.
ck("cfg_live lo", to_integer(unsigned(live_s(15 downto 0))),
to_integer(unsigned(m_live(15 downto 0))));
ck("cfg_live hi", to_integer(unsigned(live_s(31 downto 16))),
to_integer(unsigned(m_live(31 downto 16))));
ck("cfg_shadow lo", to_integer(unsigned(shadow_s(15 downto 0))),
to_integer(unsigned(m_shadow(15 downto 0))));
ck("cfg_shadow hi", to_integer(unsigned(shadow_s(31 downto 16))),
to_integer(unsigned(m_shadow(31 downto 16))));
ck("ack", sl2i(ack_s), sl2i(m_ack));
ck("err_pulse", sl2i(err_s), sl2i(m_er));
ck("err_code", to_integer(unsigned(ec_s)),
to_integer(unsigned(m_ec)));
ck("n_read", to_integer(n_rd_s), c_rd);
ck("n_write", to_integer(n_wr_s), c_wr);
ck("n_commit", to_integer(n_cm_s), c_cm);
ck("n_rc_lost", to_integer(n_rcl_s), c_rcl);
ck("n_locked", to_integer(n_lk_s), c_lk);
ck("n_badaddr", to_integer(n_ba_s), c_ba);
-- ---- the structural invariant ----
--
-- Reserved bits read as zero, always.
ck("ctrl reserved",
to_integer(unsigned(ctrl_s and not CTRL_MASK)), 0);
end procedure;
procedure step is
begin
-- Let the combinational read path settle before sampling it. `addr` was
-- driven one delta ago, and the read mux plus the output assignment
-- need deltas of their own -- without this the testbench reads the mux
-- output from BEFORE the address changed, and the mismatch looks like
-- a decode bug in the design rather than a race in the bench.
wait for 1 ns;
m_stat_pre := m_stat;
m_shadow_pre := m_shadow;
-- rdata is combinational and must be sampled BEFORE the edge: it is the
-- value the bus returns for this access, not the value the register
-- will hold afterwards.
if rd = '1' and wr = '0' and eot = '0' then
ck("rdata", to_integer(unsigned(rdata_s)),
to_integer(unsigned(exp_rdata(addr))));
end if;
model_step;
wait until rising_edge(clk);
wait for 1 ns;
steps := steps + 1;
check_out;
end procedure;
procedure acc (a : std_logic_vector(2 downto 0); r, w : std_logic;
d : std_logic_vector(15 downto 0);
e : std_logic_vector(N_EVT-1 downto 0); b : std_logic) is
begin
addr <= a; rd <= r; wr <= w; wdata <= d; evt <= e; busy <= b; eot <= '0';
wait for 0 ns;
step;
rd <= '0'; wr <= '0'; evt <= (others => '0');
end procedure;
procedure reset_all is
begin
rd <= '0'; wr <= '0'; evt <= (others => '0'); busy <= '0'; eot <= '0';
rst_n <= '0';
wait until rising_edge(clk);
wait for 1 ns;
rst_n <= '1';
m_ctrl := (others => '0'); m_lock := (others => '0');
m_stat := (others => '0'); m_evw := (others => '0');
m_stat_pre := (others => '0');
m_shadow := (others => '0'); m_live := (others => '0');
m_shadow_pre := (others => '0');
m_ack := '0'; m_er := '0'; m_ec := E_NONE;
c_rd := 0; c_wr := 0; c_cm := 0; c_rcl := 0; c_lk := 0; c_ba := 0;
wait for 1 ns;
end procedure;
begin
wait until rising_edge(clk);
wait until rising_edge(clk);
wait until rising_edge(clk);
rst_n <= '1';
wait for 1 ns;
reset_all;
-- ================= PHASE 1 -- address x access x engine state ========
for a in 0 to 7 loop
for ac in 0 to 2 loop
for bs in 0 to 1 loop
if ac = 0 then
if bs = 1 then acc(std_logic_vector(to_unsigned(a,3)), '1','0', x"A5A5", (others=>'0'), '1');
else acc(std_logic_vector(to_unsigned(a,3)), '1','0', x"A5A5", (others=>'0'), '0'); end if;
elsif ac = 1 then
if bs = 1 then acc(std_logic_vector(to_unsigned(a,3)), '0','1', x"A5A5", (others=>'0'), '1');
else acc(std_logic_vector(to_unsigned(a,3)), '0','1', x"A5A5", (others=>'0'), '0'); end if;
else
if bs = 1 then acc(std_logic_vector(to_unsigned(a,3)), '0','0', x"A5A5", (others=>'0'), '1');
else acc(std_logic_vector(to_unsigned(a,3)), '0','0', x"A5A5", (others=>'0'), '0'); end if;
end if;
reach(a * 6 + ac * 2 + bs) := 1;
end loop;
end loop;
end loop;
-- ================= PHASE 2 -- read-to-clear versus write-1-to-clear ===
reset_all;
acc(A_CTRL, '0','0', x"0000", "0101", '0');
if status_s /= "0101" or evw_s /= "0101" then
fail("the event did not reach both registers");
end if;
acc(A_STATUS, '1','0', x"0000", "0000", '0');
if status_s /= "0000" then fail("read-to-clear did not clear"); end if;
if evw_s /= "0101" then fail("reading STATUS disturbed the W1C register"); end if;
acc(A_EVENT, '0','1', x"0001", "0000", '0');
if evw_s /= "0100" then fail("W1C cleared the wrong bits"); end if;
-- ================= PHASE 3 -- the two races, over EVERY bit ===========
--
-- An event arrives on the very cycle the driver clears. It was not in the
-- value returned (read-to-clear) or in the value written back (W1C), so
-- destroying it loses an interrupt nobody ever saw.
--
-- The first version of this phase drove ONE bit with ONE pre-state, and
-- the mutation that clears the whole register died on TWO checks. Both
-- races are now driven for every event bit and both pre-states.
--
-- 4 bits x 2 pre-states x 2 clear policies = 16 scenarios
n_race_cases := 0;
for av in 0 to N_EVT-1 loop
for bs in 0 to 1 loop
for ac in 0 to 1 loop
reset_all;
e_v := (others => '0'); e_v(av) := '1';
if bs = 1 then acc(A_CTRL, '0','0', x"0000", e_v, '0'); end if;
-- ...and always set a DIFFERENT bit, so there is something for the
-- clear to legitimately remove. Without it a correct design and a
-- design that clears everything are indistinguishable.
o_v := (others => '0'); o_v((av + 1) mod N_EVT) := '1';
acc(A_CTRL, '0','0', x"0000", o_v, '0');
base := c_rcl;
if ac = 0 then
acc(A_STATUS, '1','0', x"0000", e_v, '0');
if status_s(av) /= '1' then
fail("RC race: the same-cycle event was destroyed");
end if;
if ((av + 1) mod N_EVT) /= av
and status_s((av + 1) mod N_EVT) /= '0' then
fail("RC race: the returned bit survived");
end if;
if bs = 0 then
if c_rcl /= base + 1 then fail("RC hazard not counted"); end if;
else
if c_rcl /= base then fail("RC hazard counted wrongly"); end if;
end if;
else
d_v := (others => '0'); d_v(av) := '1';
acc(A_EVENT, '0','1', d_v, e_v, '0');
if evw_s(av) /= '1' then
fail("W1C race: set lost the race");
end if;
if evw_s((av + 1) mod N_EVT) /= '1' then
fail("W1C race: an unwritten bit was cleared");
end if;
acc(A_EVENT, '0','1', d_v, (others => '0'), '0');
if evw_s(av) /= '0' then
fail("W1C: a plain write did not clear");
end if;
end if;
n_race_cases := n_race_cases + 1;
end loop;
end loop;
end loop;
if n_race_cases /= 16 then
errors := errors + 1;
write(ln, string'("FAIL race scenarios ") & integer'image(n_race_cases));
writeline(output, ln);
end if;
-- ================= PHASE 4 -- the shadow is ATOMIC ====================
reset_all;
acc(A_CFG_LO, '0','1', x"BEEF", "0000", '0');
if live_s /= x"00000000" then
fail("writing CFG_LO changed the live configuration");
end if;
acc(A_CFG_HI, '0','1', x"DEAD", "0000", '0');
if live_s /= x"00000000" then
fail("writing CFG_HI changed the live configuration");
end if;
if shadow_s /= x"DEADBEEF" then fail("the shadow is wrong"); end if;
base := c_cm;
acc(A_COMMIT, '0','1', x"0001", "0000", '0');
if live_s /= x"DEADBEEF" then fail("the commit did not apply the shadow"); end if;
if c_cm /= base + 1 then fail("the commit was not counted"); end if;
-- ================= PHASE 5 -- a rejected write is REPORTED ============
reset_all;
acc(A_LOCKED, '0','1', x"1234", "0000", '0');
base := c_lk;
acc(A_LOCKED, '0','1', x"5678", "0000", '1');
if c_lk /= base + 1 then fail("a write while busy was not reported"); end if;
acc(A_LOCKED, '1','0', x"0000", "0000", '0');
if rdata_s /= x"1234" then fail("the rejected write took effect anyway"); end if;
-- ================= PHASE 6 -- reserved bits ==========================
reset_all;
acc(A_CTRL, '0','1', x"FFFF", "0000", '0');
if ctrl_s /= x"00FF" then fail("reserved bits took a write"); end if;
acc(A_CTRL, '1','0', x"0000", "0000", '0');
if rdata_s /= x"00FF" then fail("reserved bits did not read as zero"); end if;
acc(A_CTRL, '0','1', x"00FF", "0000", '0');
if ctrl_s /= x"00FF" then fail("a read-modify-write cycle moved the register"); end if;
-- ================= PHASE 7 -- the two rejections ======================
reset_all;
base := c_ba;
acc(A_RSVD, '1','0', x"0000", "0000", '0');
if c_ba /= base + 1 then fail("reading a hole in the map was not reported"); end if;
base := c_ba;
acc(A_STATUS, '0','1', x"FFFF", "0000", '0');
if c_ba /= base + 1 then
fail("writing a read-to-clear register was not reported");
end if;
base := c_ba;
acc(A_CTRL, '1','1', x"0000", "0000", '0');
if c_ba /= base + 1 then
fail("a simultaneous read and write was not reported");
end if;
-- ================= PHASE 8 -- random =================================
--
-- The event pattern is deliberately correlated with the accesses: an
-- event drawn independently of the read almost never lands on the same
-- cycle, and that coincidence is the read-to-clear hazard.
reset_all;
for i in 0 to 29999 loop
-- Built into variables rather than written as conditional expressions
-- in the argument list: those are VHDL-2019 and this file is 2008.
w_v := rnd_nat mod 100;
a_v := std_logic_vector(to_unsigned(rnd_nat mod 8, 3));
d_v := std_logic_vector(to_unsigned(rnd_nat mod 65536, 16));
if w_v < 40 then r_v := '1'; else r_v := '0'; end if;
if w_v >= 40 and w_v < 85 then wv_v := '1'; else wv_v := '0'; end if;
if w_v < 55 then
e_v := std_logic_vector(to_unsigned(rnd_nat mod 16, N_EVT));
else
e_v := (others => '0');
end if;
if (rnd_nat mod 100) < 25 then b_v := '1'; else b_v := '0'; end if;
acc(a_v, r_v, wv_v, d_v, e_v, b_v);
end loop;
-- ================= the exhaustiveness proof ==========================
n_reach := 0;
for k in 0 to 47 loop n_reach := n_reach + reach(k); end loop;
if n_reach /= 48 then
errors := errors + 1;
write(ln, string'("FAIL address x access x busy reach ")
& integer'image(n_reach) & "/48");
writeline(output, ln);
end if;
write(ln, string'("steps=") & integer'image(steps)
& " checks=" & integer'image(checks)
& " reach=" & integer'image(n_reach) & "/48"
& " errors=" & integer'image(errors));
writeline(output, ln);
write(ln, string'("read=") & integer'image(to_integer(n_rd_s))
& " write=" & integer'image(to_integer(n_wr_s))
& " commit=" & integer'image(to_integer(n_cm_s))
& " rc_lost=" & integer'image(to_integer(n_rcl_s))
& " locked=" & integer'image(to_integer(n_lk_s))
& " badaddr=" & integer'image(to_integer(n_ba_s)));
writeline(output, ln);
if errors = 0 then
write(ln, string'("PASS: 0 errors in ") & integer'image(checks)
& " checks");
else
write(ln, string'("FAIL: ") & integer'image(errors) & " errors in " &
integer'image(checks) & " checks");
end if;
writeline(output, ln);
done <= true;
wait;
end process;
end architecture;11. Exhaustive Verification
| Measure | Verilog | SystemVerilog | VHDL |
|---|---|---|---|
| (address × access) reached | 48 / 48 | 48 / 48 | 48 / 48 |
| read-to-clear race scenarios swept | 16 / 16 | 16 / 16 | 16 / 16 |
| W1C-against-a-live-source scenarios swept | 16 / 16 | 16 / 16 | 16 / 16 |
| Steps | 30111 | 30111 | 30111 |
| Checks executed | 523901 | 523901 | 523861 |
| reads | 11986 | 11986 | 11946 |
| writes | 13552 | 13552 | 13549 |
| commits | 1689 | 1689 | 1695 |
| events arriving on the read cycle | 260 | 260 | 123 |
| writes refused as locked | 429 | 429 | 386 |
| writes refused as bad address | 4858 | 4858 | 4937 |
| Result | PASS | PASS | PASS |
12. Mutation Testing
| # | Mutation | Verilog | SysVer | VHDL |
|---|---|---|---|---|
| Z7 | the read-to-clear register accepts a write | 69612 | 69612 | 69107 |
| Z4 | the control register accepts reserved bits | 59785 | 59785 | 59935 |
| Z1 | read-to-clear clears the whole register, not the bits returned | 36381 | 36381 | 37209 |
| Z5 | a write to a locked register is accepted silently | 30836 | 30836 | 30675 |
| Z2 | a config write reaches the live register, bypassing the shadow | 30066 | 30066 | 29409 |
| Z3 | commit copies only the low half of the shadow | 29997 | 29997 | 30015 |
| Z6 | W1C clears a bit whose source is still asserting | 1551 | 1551 | 591 |
| — | unmutated baseline | 0 | 0 | 0 |
All seven die in all three languages.
The one column that is out of line, and why it is not a finding
Z6 scores 1551 in Verilog and SystemVerilog and 591 in VHDL — a 2.6× spread, by far the largest in this module. Every other mutation agrees across languages to within 3%.
The rule for this situation is to divide the suspicious score by the baseline event count it depends on before concluding anything. Z6 can only be detected on a cycle where a source is still asserting while the driver writes 1 to clear it — the same enabling situation counted in the table above as "events arriving on the read cycle": 260 in Verilog and SystemVerilog, 123 in VHDL.
Z6 score / enabling events:
Verilog / SystemVerilog 1551 / 260 = 5.97
VHDL 591 / 123 = 4.80
A 2.6x spread in the SCORE is a 1.24x spread in the RATE.The design behaves the same way in all three languages. The stimulus produced the enabling situation half as often in VHDL, because the pseudo-random generators are different — Icarus seeds $random and $urandom identically, which is why the Verilog and SystemVerilog columns match exactly everywhere in this module, and NVC's is unrelated to either.
Directed against random
| # | All phases | Directed only | Random |
|---|---|---|---|
| Z1 | 36381 | 16 | 36365 |
| Z2 | 30066 | 23 | 30043 |
| Z3 | 29997 | 18 | 29979 |
| Z4 | 59785 | 97 | 59688 |
| Z5 | 30836 | 16 | 30820 |
| Z6 | 1551 | 16 | 1535 |
| Z7 | 69612 | 49 | 69563 |
Every one is killed by directed stimulus alone. Z1 and Z6 both read exactly 16, and that is not a coincidence — it is the size of the phase written for them: 4 status bits × 2 pre-states × 2 clear policies, driven deliberately rather than waited for.
13. Debugging Walkthrough: The Counter That Reads Zero
The report. A driver's error statistics are always zero. The device definitely has errors — the link renegotiates, the analyser shows retries — but every counter in the register file reads 0.
Step 1 — is the hardware counting? Probe internally: yes, the counters increment.
Step 2 — so the read path is broken? No. Read the register twice in quick succession from a debugger and the first read returns a non-zero value.
Step 3 — what is between the hardware and the driver? The driver's own diagnostic thread, polling the same register once a second. The register is read-to-clear, and two readers of a read-to-clear register destroy each other's data. Whoever reads first gets everything; whoever reads second gets zero.
Step 4 — and the driver's real reader is second. So it always reads zero.
Step 5 — the deeper problem. Even with one reader, the design as originally written cleared the whole register on a read rather than only the bits it returned. An event arriving on the read cycle was returned to nobody and erased. At a 2% arrival rate that is one lost event in fifty — invisible in testing, and exactly the class of loss that shows up as an unexplained hang.
14. UVM: A Register Model That Knows Its Own Access Policies
// A RAL model usually mirrors ADDRESSES. The bugs in this chapter are all in
// ACCESS POLICY -- read-to-clear, write-1-to-clear, shadow/commit, locked,
// reserved bits -- so the model carries the policy and checks it.
typedef enum { RW, RO, W1C, RC, SHADOW, LOCKED_RW } policy_e;
class reg_desc;
string name;
bit [7:0] addr;
policy_e pol;
bit [31:0] wmask; // reserved bits read back as zero
function new(string n, bit [7:0] a, policy_e p, bit [31:0] m = '1);
name = n; addr = a; pol = p; wmask = m;
endfunction
endclass
class rf_predictor extends uvm_component;
`uvm_component_utils(rf_predictor)
reg_desc map [bit [7:0]];
bit [31:0] model [bit [7:0]];
bit [31:0] shadow, live;
bit locked;
int unsigned n_rc_race, n_w1c_live, n_locked_refused, n_reserved_masked;
function new(string name, uvm_component parent); super.new(name, parent);
endfunction
// ---- A READ that also writes. ----
//
// `hw_set` is the set of bits hardware asserted DURING this read. The
// entire point of the design is that they survive it: the read returns the
// value the mux sampled, and the clear applies to THAT value -- never to
// the register as it stands after the event landed.
function bit [31:0] predict_read(bit [7:0] a, bit [31:0] hw_set);
bit [31:0] returned;
reg_desc d = map[a];
if (d == null) begin
`uvm_error("RF/ADDR", $sformatf("read of unmapped address 0x%02h", a))
return 32'h0;
end
returned = model[a];
if (d.pol == RC) begin
// Clear ONLY what was returned. A bit set by hardware this cycle is in
// `hw_set` and not in `returned`, so it stays.
model[a] = (model[a] & ~returned) | hw_set;
if ((hw_set & ~returned) != 0) n_rc_race++;
end else begin
model[a] = model[a] | hw_set;
end
return returned;
endfunction
function void predict_write(bit [7:0] a, bit [31:0] wdata, bit [31:0] hw_live);
reg_desc d = map[a];
if (d == null) begin
`uvm_error("RF/ADDR", $sformatf("write to unmapped address 0x%02h", a))
return;
end
// A read-to-clear register is NOT writable. Accepting the write lets a
// driver that treats it like the W1C register clear events it never read
// -- the single most destructive thing software can do to this block.
if (d.pol == RC || d.pol == RO) begin
`uvm_error("RF/POL",
$sformatf("write to %s, which is %s and must refuse it", d.name, d.pol.name()))
return;
end
if (locked && d.pol == LOCKED_RW) begin
// Refused, and REPORTED. A lock that silently drops writes is worse
// than no lock: software believes it reconfigured the device.
n_locked_refused++;
return;
end
case (d.pol)
W1C: begin
// Clear only bits whose source is NOT still asserting. Clearing a bit
// that hardware is holding produces a register that disagrees with
// the hardware until the next event -- and a driver that reads it
// concludes the condition went away.
foreach (wdata[i])
if (wdata[i] && !hw_live[i]) model[a][i] = 1'b0;
if ((wdata & hw_live) != 0) n_w1c_live++;
end
SHADOW: begin
// The write lands in the shadow and NOWHERE else. Nothing the device
// is currently doing changes until commit -- which is the entire
// reason a shadow exists: a 32-bit configuration written through a
// 16-bit bus is briefly half-old, and a device that acts on the
// half-written value acts on a configuration nobody asked for.
shadow = wdata;
end
default: begin
if ((wdata & ~d.wmask) != 0) n_reserved_masked++;
model[a] = wdata & d.wmask;
end
endcase
endfunction
function void predict_commit();
live = shadow; // the WHOLE shadow, both halves
endfunction
function void check_phase(uvm_phase phase);
super.check_phase(phase);
// Stimulus checks, as errors: a run that never produced the race has not
// tested the design's central property, however many cycles it ran for.
if (n_rc_race == 0)
`uvm_error("RF/COV",
"no event ever arrived during a read-to-clear read: the property this design exists for was never exercised")
if (n_w1c_live == 0)
`uvm_error("RF/COV",
"no W1C write ever collided with a live source: the W1C policy was never exercised")
`uvm_info("RF",
$sformatf("%0d RC races | %0d W1C-vs-live | %0d locked refusals | %0d reserved-bit maskings",
n_rc_race, n_w1c_live, n_locked_refused, n_reserved_masked), UVM_LOW)
endfunction
endclass15. Common Misconceptions
"Read-to-clear is convenient." It is a destructive read with exactly one legal owner. A second reader — a diagnostic thread, a second driver, a debugger — destroys the first one's data.
"Clearing on read means clearing the register." It means clearing what was returned. An event that arrived on the read cycle was returned to nobody.
"W1C should clear the bit unconditionally." Not if the source is still asserting. The register would disagree with the hardware, and the driver would conclude the condition went away.
"A shadow register is over-engineering." A 32-bit configuration written through a 16-bit bus is briefly half-old. Without a shadow the device acts on a value nobody asked for.
"Commit can copy the half that changed." Z3 copies only the low half and is invisible until someone writes the high one. Commit copies the whole shadow.
"Reserved bits can be ignored." They read back as whatever was written, software stores and restores the word, and a later silicon revision gives those bits meaning.
"A locked register can drop writes quietly." Then software believes it reconfigured the device. A refused write must be reported.
"An unmapped address can return zeros." Then a typo in a driver is indistinguishable from a register that is genuinely zero.
"Same-cycle read and write is a corner case." It happened 260 times in this run and it is the reason for every design decision in the block.
16. Exercises
1. Write the read-to-clear path as stat_n = 0 on any read and give the exact cycle on which it differs from the published design. Estimate how often that cycle occurs given the 260-in-11986 rate measured here.
2. Z6 scores 1551 and 591. Reproduce the division by enabling events, and propose a measurement that would account for the residual 1.24×.
3. Z1 and Z6 both have a directed score of exactly 16. Explain why, and construct a phase that would raise it without adding a single random cycle.
4. Give the shadow/commit design a partial commit — only the halves that were written since the last commit. Which mutations does this defeat, which new one does it need, and is it worth it?
5. Two drivers share the read-to-clear counter register. Design a non-destructive mirror for the diagnostic path and say what it costs in area and in coherency.
6. The control register masks reserved bits on write, so they read back as zero. Argue the opposite policy — store them verbatim — and say which one survives a silicon revision that gives the bits meaning.
7. Across this module the directed share of the mutation score ranges from 40% (chapter 26.3) to under 0.2% (chapter 26.4) to about 0.05% here. Explain the range from the shape of the three state spaces, and say what a high directed share would tell you about a suite.
17. Summary
| Idea | Why it matters |
|---|---|
| Read-to-clear clears what was returned | an event arriving on the read cycle was seen by nobody |
| A read-to-clear register has one reader | two readers destroy each other's data |
| A read-to-clear register is not writable | Z7 lets a driver clear events it never read |
| W1C respects a live source | or the register disagrees with the hardware |
| Configuration lands in a shadow | a 32-bit value through a 16-bit bus is briefly half-old |
| Commit copies the whole shadow | Z3 copies half and hides until someone writes the other |
| Reserved bits are masked, not stored | software saves and restores the word across revisions |
| A locked write is refused and reported | a silent drop makes software believe it succeeded |
| An unmapped access is an error | or a driver typo looks like a zero register |
| Expose the raw register, not a masked view | a masked port hid Z4 entirely in one language |
| Divide a spread by its enabling event count | 2.6× in score is 1.24× in rate |
| A directed score of 2 is luck | enumerate the domain: 4 bits × 2 states × 2 policies = 16 |
| 48 address × access states, 7 mutations | all killed by directed stimulus alone, in 3 languages |
Tooling
| Step | Command |
|---|---|
| Verilog-2005 | iverilog -g2005 -o rf_v.out rf_v.v rf_v_tb.v && ./rf_v.out |
| SystemVerilog | iverilog -g2012 -o rf_sv.out rf_sv.sv rf_sv_tb.sv && ./rf_sv.out |
| VHDL-2008 analyse | nvc --std=2008 -a rf_vhdl.vhd rf_vhdl_tb.vhd |
| VHDL-2008 elaborate | nvc --std=2008 -e tb_rf_vhdl |
| VHDL-2008 run | nvc --std=2008 -r tb_rf_vhdl |
| One mutation | iverilog -g2005 -DMUT_Z1 -o mm rf_v_mut.v rf_v_tb.v && ./mm |
| Directed only | iverilog -g2005 -DDIRECTED_ONLY -o mm rf_v_mut.v rf_v_tb.v && ./mm |
All three implementations pass with 0 errors: 48 address × access states reached, 260 same-cycle read/set races and 429 locked refusals exercised, every access policy checked against an independently written model, and every one of the seven mutations killed by directed stimulus alone.
18. Module 26 in one page
Five blocks, five defects, and one idea underneath all of them.
| Chapter | The block | The defect |
|---|---|---|
| 26.1 | endpoint buffer pool | a shared pool head-of-line blocks; a per-endpoint pool cannot |
| 26.2 | DMA descriptor engine | ceil(len / maxp) omits the ZLP and hangs on the sizes everybody uses |
| 26.3 | AXI burst splitter | a burst crossing a 4 KB boundary is legal on your board and corrupts on theirs |
| 26.4 | interrupt aggregator | hardware sets a bit as software clears it, and the clear wins |
| 26.5 | register file | read-to-clear clears the whole register, including what nobody saw |
Every one of the five is a race between two agents who each behave correctly. Nothing here is a typo or a missing case. The DMA engine's arithmetic is right for every length but the multiples. The AXI splitter tiles the address range exactly. The aggregator sets and clears exactly what it was told. The register file returns exactly what was asked for. They fail because two things happen on the same cycle and the design silently picks one.
That is what makes integration bugs different from RTL bugs, and it is why every chapter in this module counts how often the situation arose alongside whether the design survived it. A verification report that says "0 errors" and cannot say "and the race occurred 17920 times" has proved that a testbench ran, not that a design works.
Continue learning
Related tutorials
- Related topic
USB Interrupt Integration
Hardware sets a status bit on the same cycle the driver writes 1 to clear it, and if the clear wins that interrupt is not delayed — it is gone, along with every event behind it.
- Related topic
USB Controllers on SoC
DWC2, MUSB and xHCI differ in a hundred mechanical ways that do not matter and one architectural way that does — whether endpoints own their packet buffers or share a pool.
- Related topic
DMA Integration
A descriptor has a byte count and the wire has packets, and the rule that converts one to the other is not ceil(length / packet size) — the version that is hangs on exactly the buffer sizes everybody uses.
- Related topic
AXI Interfaces to USB
AXI forbids a burst from crossing a 4 KB address boundary, and the reason this bug reaches production is that many interconnects quietly split the burst for you — so the engine that breaks the rule works on the board you develop on and corrupts memory on the board you ship.
Standards & specifications
- Governing standard
- USB-IF (Universal Serial Bus Specification)(opens USB Implementers Forum (USB-IF) in a new tab)
Defines the USB bus — its electrical signalling, connectors, packet and transaction model, device framework and the descriptors a device must expose — together with the device-class specifications layered on it. It does not define host-controller register interfaces (xHCI and EHCI are separate documents) nor any operating system's driver architecture.
This page also covers RTL structure, verification approach and debugging technique. Those are engineering practice built on the standard, not requirements the standard itself imposes.
Where this fits
Part of the USB curriculum.
