Skip to content
VLSI Mentor

USB · Module 26

Firmware Interaction

A register interface looks like memory and is not — reading can change it, writing 1 can clear it, one register can apply another, and none of that is anything a compiler knows.

Chapter 26.4 got the driver's attention. This chapter is about what it does next, and about the gap between what a register interface looks like and what it is.

1. A Register Interface Is Not Memory

It looks like memory. It is addressed like memory, the compiler emits loads and stores for it, and every abstraction between the driver and the wire treats it as a place to put numbers.

It is not memory, because:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   reading can CHANGE it                  read-to-clear
   writing 1 can CLEAR it                 write-1-to-clear
   writing one register can APPLY another shadow / commit
   writing some registers is REJECTED     locked while busy

None of which a compiler knows, and none of which appears in a header file. Every one of them is a source of driver bugs that look like hardware bugs.

2. Read-to-Clear and Write-1-to-Clear Are Not the Same Thing

Both hand the driver a set of events and then empty the register. The difference is when, and what happens to an event that arrives in between.

How it clearsWhat it can destroy
read-to-clearthe read returns the bits and clears them, indivisiblyan event arriving on the read cycle — it was not returned
write-1-to-clearthe driver clears exactly the bits it handled, laternothing it did not see, because it never writes those bits

3. A Multi-Field Configuration Must Be Applied Atomically

A configuration that spans two registers cannot be written live. Between the two stores the hardware is running on half the old value and half the new one, and for one bus cycle it is configured as something nobody asked for.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   write CFG_LO  ->  a shadow
   write CFG_HI  ->  a shadow
   write COMMIT  ->  BOTH become live, in the same cycle

The shadow is not an optimisation. It is the only way to make a 32-bit change atomic across a 16-bit interface — and one bus cycle of a wrong configuration is thousands of bit times on a USB link.

4. And Some Writes Are Rejected

A register that configures a running engine cannot be changed while it runs.

5. Reserved Bits Are a Two-Sided Rule

They read as zero and ignore writes, and the two halves have to agree.

A driver doing read-modify-write writes back exactly what it read. If reserved bits read as zero but accept writes, the cycle is still a fixed point. If they read as something else and accept writes, the register walks somewhere new every time round — which is the kind of bug that appears only after the tenth iteration of a polling loop.

6. What We Are Building

usb_reg_file implements one of each behaviour, in eight addresses:

AddressRegisterBehaviour
0CTRLRW, upper byte reserved
1STATUSread-to-clear, and not writable
2EVENTwrite-1-to-clear
3 / 4CFG_LO / CFG_HIshadow — staged, not live
5COMMITwrite-only; applies both halves at once
6LOCKEDRW, rejected while busy
7—a hole in the map

7. Verilog-2005 Implementation

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// usb_reg_file -- why driver code that looks like memory access is not, and
// the four register behaviours that make it a protocol.
//
// A REGISTER INTERFACE IS NOT MEMORY
//
// It looks like memory. It is addressed like memory, the compiler emits
// loads and stores for it, and every abstraction between the driver and the
// wire treats it as a place to put numbers.
//
// It is not memory, because:
//
//     reading can CHANGE it          (read-to-clear)
//     writing 1 can CLEAR it         (write-1-to-clear)
//     writing one register can APPLY another   (shadow / commit)
//     writing some registers is REJECTED at the wrong moment
//
// None of which a compiler knows, and none of which shows up in a header
// file. Every one of them is a source of driver bugs that look like hardware
// bugs.
//
// READ-TO-CLEAR AND WRITE-1-TO-CLEAR ARE NOT THE SAME THING
//
// Both hand the driver a set of events and then empty the register. The
// difference is WHEN, and what happens to an event that arrives in between.
//
//     read-to-clear   the read returns the bits AND clears them, in one
//                     indivisible step. Simple, and it has a trap: the
//                     register must clear only THE BITS IT RETURNED. An
//                     event arriving on the read cycle was not returned,
//                     so clearing it destroys an interrupt nobody ever saw.
//
//     write-1-to-clear   the driver clears exactly the bits it handled,
//                        explicitly, later. Nothing it did not see can be
//                        destroyed, because it never writes those bits.
//
// The same event source feeds both registers in this block, deliberately, so
// the two policies can be compared on identical stimulus.
//
// A MULTI-FIELD CONFIGURATION MUST BE APPLIED ATOMICALLY
//
// A configuration that spans two registers cannot be written live. Between
// the two stores the hardware is running on half the old value and half the
// new one, and for one bus cycle it is configured as something nobody asked
// for.
//
//     write CFG_LO -> a shadow
//     write CFG_HI -> a shadow
//     write COMMIT -> BOTH become live, in the same cycle
//
// The shadow is not an optimisation. It is the only way to make a 32-bit
// change atomic across a 16-bit interface.
//
// AND SOME WRITES ARE REJECTED
//
// A register that configures a running engine cannot be changed while it
// runs. Silently ignoring the write is the worst option: the driver believes
// the setting took. This block reports it, because a rejected write is a
// driver bug and the driver is the only thing that can fix it.
module usb_reg_file #(
  parameter integer N_EVT = 4      // event bits, in both status registers
) (
  input  wire        clk,
  input  wire        rst_n,

  input  wire [2:0]  addr,
  input  wire        rd,
  input  wire        wr,
  input  wire [15:0] wdata,
  input  wire [N_EVT-1:0] evt,     // one pulse per hardware event
  input  wire        busy,         // the engine is running: config is locked
  input  wire        eot,

  output wire [15:0] rdata,
  output wire        ack,
  output wire        err_pulse,
  output wire [1:0]  err_code,

  output wire [15:0] ctrl,
  output wire [N_EVT-1:0] status,  // read-to-clear
  output wire [N_EVT-1:0] event_w1c,
  output wire [31:0] cfg_live,     // what the engine is actually using
  output wire [31:0] cfg_shadow,   // what has been staged but not applied

  output reg [31:0] n_read,
  output reg [31:0] n_write,
  output reg [31:0] n_commit,
  output reg [31:0] n_rc_lost,     // an event destroyed by a read-to-clear
  output reg [31:0] n_locked,      // a write rejected because busy
  output reg [31:0] n_badaddr
);

  // The map. Addresses are named once, here, because a driver header and an
  // RTL case statement that disagree is a bug no simulation can find.
  localparam [2:0] A_CTRL   = 3'd0,   // RW, with reserved bits
                   A_STATUS = 3'd1,   // read-to-clear
                   A_EVENT  = 3'd2,   // write-1-to-clear
                   A_CFG_LO = 3'd3,   // shadow
                   A_CFG_HI = 3'd4,   // shadow
                   A_COMMIT = 3'd5,   // write-only: applies the shadow
                   A_LOCKED = 3'd6,   // RW, rejected while busy
                   A_RSVD   = 3'd7;   // not a register

  localparam [1:0] E_NONE    = 2'd0,
                   E_LOCKED  = 2'd1,  // a write while the engine is running
                   E_BADADDR = 2'd2;  // an access to a hole in the map

  // Reserved bits read as ZERO and ignore writes. Stated as a mask rather
  // than assumed, because a driver doing read-modify-write on this register
  // will write back whatever it read -- and if the reserved bits read back
  // as something other than the value it must write, every such cycle
  // corrupts them.
  localparam [15:0] CTRL_MASK = 16'h00FF;

  reg [15:0]       ctrl_r;
  reg [N_EVT-1:0]  stat_r, evw_r;
  reg [31:0]       shadow_r, live_r;
  reg [15:0]       lock_r;
  reg              ack_r, er_r;
  reg [1:0]        ec_r;

  // The RAW register, not the masked read value. The read path masks
  // separately; exposing the stored value here is what makes "the reserved
  // bits are never even STORED" checkable, rather than only "the read path
  // hides them".
  assign ctrl       = ctrl_r;
  assign status     = stat_r;
  assign event_w1c  = evw_r;
  assign cfg_live   = live_r;
  assign cfg_shadow = shadow_r;
  assign ack        = ack_r;
  assign err_pulse  = er_r;
  assign err_code   = ec_r;

  // ---- The read data is COMBINATIONAL and has no side effect of its own.
  //
  // The side effect lives in the sequential block. Mixing the two -- reading
  // a register that the same expression also clears -- is how a read-to-clear
  // ends up returning the value it just destroyed.
  reg [15:0] rd_mux;
  always @* begin
    case (addr)
      A_CTRL:   rd_mux = ctrl_r & CTRL_MASK;
      A_STATUS: rd_mux = {{(16-N_EVT){1'b0}}, stat_r};
      A_EVENT:  rd_mux = {{(16-N_EVT){1'b0}}, evw_r};
      A_CFG_LO: rd_mux = shadow_r[15:0];
      A_CFG_HI: rd_mux = shadow_r[31:16];
      // COMMIT is write-only. It reads as zero rather than as the last
      // thing written, so a driver that reads it back does not conclude the
      // commit is still pending.
      A_COMMIT: rd_mux = 16'h0000;
      A_LOCKED: rd_mux = lock_r;
      default:  rd_mux = 16'h0000;
    endcase
  end
  assign rdata = rd_mux;

  integer i;
  reg [15:0]      ctrl_n, lock_n;
  reg [N_EVT-1:0] stat_n, evw_n;
  reg [31:0]      shadow_n, live_n;
  reg             ack_n, er_n;
  reg [1:0]       ec_n;
  reg             rd_n, wr_n, cm_n, lk_n, ba_n;
  reg [31:0]      rcl_n;

  always @* begin
    ctrl_n   = ctrl_r;
    lock_n   = lock_r;
    stat_n   = stat_r;
    evw_n    = evw_r;
    shadow_n = shadow_r;
    live_n   = live_r;
    ack_n    = 1'b0; er_n = 1'b0; ec_n = E_NONE;
    rd_n = 1'b0; wr_n = 1'b0; cm_n = 1'b0; lk_n = 1'b0; ba_n = 1'b0;
    rcl_n = 32'd0;

    if (eot) begin
      // Nothing: the counters are the report.
    end else begin
      // ---- The event sources come FIRST, and they SET. ----
      //
      // Before any clear, so that an event arriving on the same cycle as a
      // read or a write survives it. This is chapter 26.4's rule applied to
      // two different clear policies at once.
      for (i = 0; i < N_EVT; i = i + 1) begin
        if (evt[i]) begin
          stat_n[i] = 1'b1;
          evw_n[i]  = 1'b1;
        end
      end

      if (rd && wr) begin
        // A simultaneous read and write of the same address is not something
        // a single-port interface can do, and accepting it quietly gives the
        // driver a value that belongs to neither.
        er_n = 1'b1; ec_n = E_BADADDR;
        ba_n = 1'b1;
      end else if (rd) begin
        ack_n = 1'b1;
        rd_n  = 1'b1;
        if (addr == A_RSVD) begin
          er_n = 1'b1; ec_n = E_BADADDR;
          ba_n = 1'b1;
        end else if (addr == A_STATUS) begin
          // ---- READ-TO-CLEAR, and the whole trap in one line. ----
          //
          // Clear only the bits that were RETURNED. An event that arrived on
          // this very cycle is in stat_n and was NOT in rd_mux, so clearing
          // the whole register would destroy an interrupt nobody ever saw --
          // and nothing else in the system will ever mention it again.
          for (i = 0; i < N_EVT; i = i + 1)
            if (stat_r[i] && !evt[i]) stat_n[i] = 1'b0;
          // The count of events that WOULD have been lost by a naive clear.
          // Reported so that the hazard is visible in a run rather than
          // argued about in a review.
          for (i = 0; i < N_EVT; i = i + 1)
            if (evt[i] && !stat_r[i]) rcl_n = rcl_n + 32'd1;
        end
      end else if (wr) begin
        ack_n = 1'b1;
        wr_n  = 1'b1;
        case (addr)
          A_CTRL: begin
            // Reserved bits ignore writes as well as reading zero. The two
            // halves of that rule have to agree or a read-modify-write walks
            // the register somewhere new every time round.
            ctrl_n = wdata & CTRL_MASK;
          end
          A_STATUS: begin
            // A read-to-clear register is not writable. Accepting the write
            // silently lets a driver that treats it like the W1C register
            // clear events it never read.
            er_n = 1'b1; ec_n = E_BADADDR;
            ba_n = 1'b1;
          end
          A_EVENT: begin
            // ---- WRITE-1-TO-CLEAR. Set still wins. ----
            for (i = 0; i < N_EVT; i = i + 1)
              if (wdata[i] && !evt[i]) evw_n[i] = 1'b0;
          end
          A_CFG_LO: shadow_n[15:0]  = wdata;
          A_CFG_HI: shadow_n[31:16] = wdata;
          A_COMMIT: begin
            // ---- The whole 32 bits become live in ONE cycle. ----
            //
            // Which is the only way to make the change atomic across a
            // 16-bit interface. Writing the halves live means the engine
            // runs on a configuration nobody asked for for one bus cycle,
            // and one bus cycle is thousands of bit times on a USB link.
            live_n = shadow_r;
            cm_n   = 1'b1;
          end
          A_LOCKED: begin
            if (busy) begin
              // Reported, not ignored. A silently dropped write leaves the
              // driver believing a setting took effect, and the next thing
              // it does is built on that belief.
              er_n = 1'b1; ec_n = E_LOCKED;
              lk_n = 1'b1;
            end else begin
              lock_n = wdata;
            end
          end
          default: begin
            er_n = 1'b1; ec_n = E_BADADDR;
            ba_n = 1'b1;
          end
        endcase
      end
    end
  end

  always @(posedge clk or negedge rst_n) begin
    if (!rst_n) begin
      ctrl_r   <= 16'd0;
      lock_r   <= 16'd0;
      stat_r   <= {N_EVT{1'b0}};
      evw_r    <= {N_EVT{1'b0}};
      shadow_r <= 32'd0;
      live_r   <= 32'd0;
      ack_r    <= 1'b0; er_r <= 1'b0; ec_r <= E_NONE;
      n_read    <= 32'd0;
      n_write   <= 32'd0;
      n_commit  <= 32'd0;
      n_rc_lost <= 32'd0;
      n_locked  <= 32'd0;
      n_badaddr <= 32'd0;
    end else begin
      ctrl_r   <= ctrl_n;
      lock_r   <= lock_n;
      stat_r   <= stat_n;
      evw_r    <= evw_n;
      shadow_r <= shadow_n;
      live_r   <= live_n;
      ack_r    <= ack_n; er_r <= er_n; ec_r <= ec_n;

      if (rd_n) n_read    <= n_read    + 32'd1;
      if (wr_n) n_write   <= n_write   + 32'd1;
      if (cm_n) n_commit  <= n_commit  + 32'd1;
      if (lk_n) n_locked  <= n_locked  + 32'd1;
      if (ba_n) n_badaddr <= n_badaddr + 32'd1;
      n_rc_lost <= n_rc_lost + rcl_n;
    end
  end
endmodule

8. SystemVerilog Implementation

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// usb_reg_file -- why driver code that looks like memory access is not, and
// the four register behaviours that make it a protocol.
//
// A REGISTER INTERFACE IS NOT MEMORY
//
// It looks like memory. It is addressed like memory, the compiler emits
// loads and stores for it, and every abstraction between the driver and the
// wire treats it as a place to put numbers.
//
// It is not memory, because:
//
//     reading can CHANGE it          (read-to-clear)
//     writing 1 can CLEAR it         (write-1-to-clear)
//     writing one register can APPLY another   (shadow / commit)
//     writing some registers is REJECTED at the wrong moment
//
// None of which a compiler knows, and none of which shows up in a header
// file. Every one of them is a source of driver bugs that look like hardware
// bugs.
//
// READ-TO-CLEAR AND WRITE-1-TO-CLEAR ARE NOT THE SAME THING
//
// Both hand the driver a set of events and then empty the register. The
// difference is WHEN, and what happens to an event that arrives in between.
//
//     read-to-clear   the read returns the bits AND clears them, in one
//                     indivisible step. Simple, and it has a trap: the
//                     register must clear only THE BITS IT RETURNED. An
//                     event arriving on the read cycle was not returned,
//                     so clearing it destroys an interrupt nobody ever saw.
//
//     write-1-to-clear   the driver clears exactly the bits it handled,
//                        explicitly, later. Nothing it did not see can be
//                        destroyed, because it never writes those bits.
//
// The same event source feeds both registers in this block, deliberately, so
// the two policies can be compared on identical stimulus.
//
// A MULTI-FIELD CONFIGURATION MUST BE APPLIED ATOMICALLY
//
// A configuration that spans two registers cannot be written live. Between
// the two stores the hardware is running on half the old value and half the
// new one, and for one bus cycle it is configured as something nobody asked
// for.
//
//     write CFG_LO -> a shadow
//     write CFG_HI -> a shadow
//     write COMMIT -> BOTH become live, in the same cycle
//
// The shadow is not an optimisation. It is the only way to make a 32-bit
// change atomic across a 16-bit interface.
//
// AND SOME WRITES ARE REJECTED
//
// A register that configures a running engine cannot be changed while it
// runs. Silently ignoring the write is the worst option: the driver believes
// the setting took. This block reports it, because a rejected write is a
// driver bug and the driver is the only thing that can fix it.
package usb_reg_pkg;
  // The map. Named once, in a package, because a driver header and an RTL
  // case statement that disagree is a bug no simulation can find.
  typedef enum logic [2:0] {
    A_CTRL   = 3'd0,   // RW, with reserved bits
    A_STATUS = 3'd1,   // read-to-clear
    A_EVENT  = 3'd2,   // write-1-to-clear
    A_CFG_LO = 3'd3,   // shadow
    A_CFG_HI = 3'd4,   // shadow
    A_COMMIT = 3'd5,   // write-only: applies the shadow
    A_LOCKED = 3'd6,   // RW, rejected while busy
    A_RSVD   = 3'd7    // not a register
  } reg_addr_e;

  typedef enum logic [1:0] {
    E_NONE    = 2'd0,
    E_LOCKED  = 2'd1,  // a write while the engine is running
    E_BADADDR = 2'd2   // an access to a hole in the map
  } reg_err_e;
endpackage

module usb_reg_file
  import usb_reg_pkg::*;
 #(
  parameter int N_EVT = 4          // event bits, in both status registers
) (
  input  logic        clk,
  input  logic        rst_n,

  input  reg_addr_e   addr,
  input  logic        rd,
  input  logic        wr,
  input  logic [15:0] wdata,
  input  logic [N_EVT-1:0] evt,    // one pulse per hardware event
  input  logic        busy,        // the engine is running: config is locked
  input  logic        eot,

  output logic [15:0] rdata,
  output logic        ack,
  output logic        err_pulse,
  output reg_err_e    err_code,

  output logic [15:0] ctrl,
  output logic [N_EVT-1:0] status, // read-to-clear
  output logic [N_EVT-1:0] event_w1c,
  output logic [31:0] cfg_live,    // what the engine is actually using
  output logic [31:0] cfg_shadow,  // what has been staged but not applied

  output logic [31:0] n_read,
  output logic [31:0] n_write,
  output logic [31:0] n_commit,
  output logic [31:0] n_rc_lost,     // an event destroyed by a read-to-clear
  output logic [31:0] n_locked,      // a write rejected because busy
  output logic [31:0] n_badaddr
);

  // Reserved bits read as ZERO and ignore writes. Stated as a mask rather
  // than assumed, because a driver doing read-modify-write on this register
  // will write back whatever it read -- and if the reserved bits read back
  // as something other than the value it must write, every such cycle
  // corrupts them.
  localparam logic [15:0] CTRL_MASK = 16'h00FF;

  logic [15:0]      ctrl_r;
  logic [N_EVT-1:0] stat_r, evw_r;
  logic [31:0]      shadow_r, live_r;
  logic [15:0]      lock_r;
  logic             ack_r, er_r;
  reg_err_e         ec_r;

  // The RAW register, not the masked read value. The read path masks
  // separately; exposing the stored value here is what makes "the reserved
  // bits are never even STORED" checkable, rather than only "the read path
  // hides them".
  assign ctrl       = ctrl_r;
  assign status     = stat_r;
  assign event_w1c  = evw_r;
  assign cfg_live   = live_r;
  assign cfg_shadow = shadow_r;
  assign ack        = ack_r;
  assign err_pulse  = er_r;
  assign err_code   = ec_r;

  // ---- The read data is COMBINATIONAL and has no side effect of its own.
  //
  // The side effect lives in the sequential block. Mixing the two -- reading
  // a register that the same expression also clears -- is how a read-to-clear
  // ends up returning the value it just destroyed.
  logic [15:0] rd_mux;
  always_comb begin
    case (addr)
      A_CTRL:   rd_mux = ctrl_r & CTRL_MASK;
      A_STATUS: rd_mux = {{(16-N_EVT){1'b0}}, stat_r};
      A_EVENT:  rd_mux = {{(16-N_EVT){1'b0}}, evw_r};
      A_CFG_LO: rd_mux = shadow_r[15:0];
      A_CFG_HI: rd_mux = shadow_r[31:16];
      // COMMIT is write-only. It reads as zero rather than as the last
      // thing written, so a driver that reads it back does not conclude the
      // commit is still pending.
      A_COMMIT: rd_mux = 16'h0000;
      A_LOCKED: rd_mux = lock_r;
      default:  rd_mux = 16'h0000;
    endcase
  end
  assign rdata = rd_mux;

  int i;
  logic [15:0]      ctrl_n, lock_n;
  logic [N_EVT-1:0] stat_n, evw_n;
  logic [31:0]      shadow_n, live_n;
  logic             ack_n, er_n;
  reg_err_e         ec_n;
  logic             rd_n, wr_n, cm_n, lk_n, ba_n;
  logic [31:0]      rcl_n;

  always_comb begin
    ctrl_n   = ctrl_r;
    lock_n   = lock_r;
    stat_n   = stat_r;
    evw_n    = evw_r;
    shadow_n = shadow_r;
    live_n   = live_r;
    ack_n    = 1'b0; er_n = 1'b0; ec_n = E_NONE;
    rd_n = 1'b0; wr_n = 1'b0; cm_n = 1'b0; lk_n = 1'b0; ba_n = 1'b0;
    rcl_n = 32'd0;

    if (eot) begin
      // Nothing: the counters are the report.
    end else begin
      // ---- The event sources come FIRST, and they SET. ----
      //
      // Before any clear, so that an event arriving on the same cycle as a
      // read or a write survives it. This is chapter 26.4's rule applied to
      // two different clear policies at once.
      for (i = 0; i < N_EVT; i = i + 1) begin
        if (evt[i]) begin
          stat_n[i] = 1'b1;
          evw_n[i]  = 1'b1;
        end
      end

      if (rd && wr) begin
        // A simultaneous read and write of the same address is not something
        // a single-port interface can do, and accepting it quietly gives the
        // driver a value that belongs to neither.
        er_n = 1'b1; ec_n = E_BADADDR;
        ba_n = 1'b1;
      end else if (rd) begin
        ack_n = 1'b1;
        rd_n  = 1'b1;
        if (addr == A_RSVD) begin
          er_n = 1'b1; ec_n = E_BADADDR;
          ba_n = 1'b1;
        end else if (addr == A_STATUS) begin
          // ---- READ-TO-CLEAR, and the whole trap in one line. ----
          //
          // Clear only the bits that were RETURNED. An event that arrived on
          // this very cycle is in stat_n and was NOT in rd_mux, so clearing
          // the whole register would destroy an interrupt nobody ever saw --
          // and nothing else in the system will ever mention it again.
          for (i = 0; i < N_EVT; i = i + 1)
            if (stat_r[i] && !evt[i]) stat_n[i] = 1'b0;
          // The count of events that WOULD have been lost by a naive clear.
          // Reported so that the hazard is visible in a run rather than
          // argued about in a review.
          for (i = 0; i < N_EVT; i = i + 1)
            if (evt[i] && !stat_r[i]) rcl_n = rcl_n + 32'd1;
        end
      end else if (wr) begin
        ack_n = 1'b1;
        wr_n  = 1'b1;
        case (addr)
          A_CTRL: begin
            // Reserved bits ignore writes as well as reading zero. The two
            // halves of that rule have to agree or a read-modify-write walks
            // the register somewhere new every time round.
            ctrl_n = wdata & CTRL_MASK;
          end
          A_STATUS: begin
            // A read-to-clear register is not writable. Accepting the write
            // silently lets a driver that treats it like the W1C register
            // clear events it never read.
            er_n = 1'b1; ec_n = E_BADADDR;
            ba_n = 1'b1;
          end
          A_EVENT: begin
            // ---- WRITE-1-TO-CLEAR. Set still wins. ----
            for (i = 0; i < N_EVT; i = i + 1)
              if (wdata[i] && !evt[i]) evw_n[i] = 1'b0;
          end
          A_CFG_LO: shadow_n[15:0]  = wdata;
          A_CFG_HI: shadow_n[31:16] = wdata;
          A_COMMIT: begin
            // ---- The whole 32 bits become live in ONE cycle. ----
            //
            // Which is the only way to make the change atomic across a
            // 16-bit interface. Writing the halves live means the engine
            // runs on a configuration nobody asked for for one bus cycle,
            // and one bus cycle is thousands of bit times on a USB link.
            live_n = shadow_r;
            cm_n   = 1'b1;
          end
          A_LOCKED: begin
            if (busy) begin
              // Reported, not ignored. A silently dropped write leaves the
              // driver believing a setting took effect, and the next thing
              // it does is built on that belief.
              er_n = 1'b1; ec_n = E_LOCKED;
              lk_n = 1'b1;
            end else begin
              lock_n = wdata;
            end
          end
          default: begin
            er_n = 1'b1; ec_n = E_BADADDR;
            ba_n = 1'b1;
          end
        endcase
      end
    end
  end

  always_ff @(posedge clk or negedge rst_n) begin
    if (!rst_n) begin
      ctrl_r   <= 16'd0;
      lock_r   <= 16'd0;
      stat_r   <= '0;
      evw_r    <= '0;
      shadow_r <= 32'd0;
      live_r   <= 32'd0;
      ack_r    <= 1'b0; er_r <= 1'b0; ec_r <= E_NONE;
      n_read    <= 32'd0;
      n_write   <= 32'd0;
      n_commit  <= 32'd0;
      n_rc_lost <= 32'd0;
      n_locked  <= 32'd0;
      n_badaddr <= 32'd0;
    end else begin
      ctrl_r   <= ctrl_n;
      lock_r   <= lock_n;
      stat_r   <= stat_n;
      evw_r    <= evw_n;
      shadow_r <= shadow_n;
      live_r   <= live_n;
      ack_r    <= ack_n; er_r <= er_n; ec_r <= ec_n;

      if (rd_n) n_read    <= n_read    + 32'd1;
      if (wr_n) n_write   <= n_write   + 32'd1;
      if (cm_n) n_commit  <= n_commit  + 32'd1;
      if (lk_n) n_locked  <= n_locked  + 32'd1;
      if (ba_n) n_badaddr <= n_badaddr + 32'd1;
      n_rc_lost <= n_rc_lost + rcl_n;
    end
  end
endmodule

9. VHDL-2008 Implementation

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
-- usb_reg_file -- why driver code that looks like memory access is not, and
-- the four register behaviours that make it a protocol.
--
-- A REGISTER INTERFACE IS NOT MEMORY
--
-- It looks like memory. It is addressed like memory, the compiler emits loads
-- and stores for it, and every abstraction between the driver and the wire
-- treats it as a place to put numbers.
--
-- It is not memory, because:
--
--     reading can CHANGE it          (read-to-clear)
--     writing 1 can CLEAR it         (write-1-to-clear)
--     writing one register can APPLY another   (shadow / commit)
--     writing some registers is REJECTED at the wrong moment
--
-- None of which a compiler knows, and none of which shows up in a header
-- file. Every one of them is a source of driver bugs that look like hardware
-- bugs.
--
-- READ-TO-CLEAR AND WRITE-1-TO-CLEAR ARE NOT THE SAME THING
--
-- Both hand the driver a set of events and then empty the register. The
-- difference is WHEN, and what happens to an event that arrives in between.
--
--     read-to-clear   the read returns the bits AND clears them, in one
--                     indivisible step. Simple, and it has a trap: the
--                     register must clear only THE BITS IT RETURNED. An event
--                     arriving on the read cycle was not returned, so
--                     clearing it destroys an interrupt nobody ever saw.
--
--     write-1-to-clear   the driver clears exactly the bits it handled,
--                        explicitly, later. Nothing it did not see can be
--                        destroyed, because it never writes those bits.
--
-- The same event source feeds both registers in this block, deliberately, so
-- the two policies can be compared on identical stimulus.
--
-- A MULTI-FIELD CONFIGURATION MUST BE APPLIED ATOMICALLY
--
-- A configuration that spans two registers cannot be written live. Between
-- the two stores the hardware is running on half the old value and half the
-- new one, and for one bus cycle it is configured as something nobody asked
-- for.
--
--     write CFG_LO -> a shadow
--     write CFG_HI -> a shadow
--     write COMMIT -> BOTH become live, in the same cycle
--
-- The shadow is not an optimisation. It is the only way to make a 32-bit
-- change atomic across a 16-bit interface.
--
-- AND SOME WRITES ARE REJECTED
--
-- A register that configures a running engine cannot be changed while it
-- runs. Silently ignoring the write is the worst option: the driver believes
-- the setting took. This block reports it, because a rejected write is a
-- driver bug and the driver is the only thing that can fix it.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

package usb_reg_pkg is
  -- The map. Named once, in a package, because a driver header and an RTL
  -- case statement that disagree is a bug no simulation can find.
  constant A_CTRL   : std_logic_vector(2 downto 0) := "000";
  constant A_STATUS : std_logic_vector(2 downto 0) := "001";
  constant A_EVENT  : std_logic_vector(2 downto 0) := "010";
  constant A_CFG_LO : std_logic_vector(2 downto 0) := "011";
  constant A_CFG_HI : std_logic_vector(2 downto 0) := "100";
  constant A_COMMIT : std_logic_vector(2 downto 0) := "101";
  constant A_LOCKED : std_logic_vector(2 downto 0) := "110";
  constant A_RSVD   : std_logic_vector(2 downto 0) := "111";

  constant E_NONE    : std_logic_vector(1 downto 0) := "00";
  constant E_LOCKED  : std_logic_vector(1 downto 0) := "01";
  constant E_BADADDR : std_logic_vector(1 downto 0) := "10";

  -- Reserved bits read as ZERO and ignore writes. Stated as a mask rather
  -- than assumed, because a driver doing read-modify-write on this register
  -- will write back whatever it read.
  constant CTRL_MASK : std_logic_vector(15 downto 0) := x"00FF";
end package;

library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
use work.usb_reg_pkg.all;

entity usb_reg_file is
  generic (
    N_EVT : integer := 4          -- event bits, in both status registers
  );
  port (
    clk   : in std_logic;
    rst_n : in std_logic;

    addr  : in std_logic_vector(2 downto 0);
    rd    : in std_logic;
    wr    : in std_logic;
    wdata : in std_logic_vector(15 downto 0);
    evt   : in std_logic_vector(N_EVT-1 downto 0);
    busy  : in std_logic;
    eot   : in std_logic;

    rdata     : out std_logic_vector(15 downto 0);
    ack       : out std_logic;
    err_pulse : out std_logic;
    err_code  : out std_logic_vector(1 downto 0);

    ctrl       : out std_logic_vector(15 downto 0);
    status     : out std_logic_vector(N_EVT-1 downto 0);  -- read-to-clear
    event_w1c  : out std_logic_vector(N_EVT-1 downto 0);
    cfg_live   : out std_logic_vector(31 downto 0);
    cfg_shadow : out std_logic_vector(31 downto 0);

    n_read    : out unsigned(31 downto 0);
    n_write   : out unsigned(31 downto 0);
    n_commit  : out unsigned(31 downto 0);
    n_rc_lost : out unsigned(31 downto 0);
    n_locked  : out unsigned(31 downto 0);
    n_badaddr : out unsigned(31 downto 0)
  );
end entity;

architecture rtl of usb_reg_file is
  signal ctrl_r, lock_r : std_logic_vector(15 downto 0) := (others => '0');
  signal stat_r, evw_r  : std_logic_vector(N_EVT-1 downto 0)
    := (others => '0');
  signal shadow_r, live_r : std_logic_vector(31 downto 0) := (others => '0');
  signal ack_r, er_r : std_logic := '0';
  signal ec_r : std_logic_vector(1 downto 0) := E_NONE;

  signal rd_c, wr_c, cm_c : unsigned(31 downto 0) := (others => '0');
  signal rcl_c, lk_c, ba_c : unsigned(31 downto 0) := (others => '0');

  signal rd_mux : std_logic_vector(15 downto 0) := (others => '0');
begin
  -- The RAW register, not the masked read value. The read path masks
  -- separately; exposing the stored value here is what makes "the reserved
  -- bits are never even STORED" checkable, rather than only "the read path
  -- hides them". Masking here too would make a design that stores reserved
  -- bits indistinguishable from one that does not.
  ctrl       <= ctrl_r;
  status     <= stat_r;
  event_w1c  <= evw_r;
  cfg_live   <= live_r;
  cfg_shadow <= shadow_r;
  ack        <= ack_r;
  err_pulse  <= er_r;
  err_code   <= ec_r;

  n_read    <= rd_c;
  n_write   <= wr_c;
  n_commit  <= cm_c;
  n_rc_lost <= rcl_c;
  n_locked  <= lk_c;
  n_badaddr <= ba_c;

  -- ---- The read data is COMBINATIONAL and has no side effect of its own.
  --
  -- The side effect lives in the clocked process. Mixing the two -- reading a
  -- register that the same expression also clears -- is how a read-to-clear
  -- ends up returning the value it just destroyed.
  process (addr, ctrl_r, stat_r, evw_r, shadow_r, lock_r)
  begin
    case addr is
      when A_CTRL   => rd_mux <= ctrl_r and CTRL_MASK;
      when A_STATUS =>
        rd_mux <= (15 downto N_EVT => '0') & stat_r;
      when A_EVENT  =>
        rd_mux <= (15 downto N_EVT => '0') & evw_r;
      when A_CFG_LO => rd_mux <= shadow_r(15 downto 0);
      when A_CFG_HI => rd_mux <= shadow_r(31 downto 16);
      -- COMMIT is write-only. It reads as zero rather than as the last thing
      -- written, so a driver that reads it back does not conclude the commit
      -- is still pending.
      when A_COMMIT => rd_mux <= (others => '0');
      when A_LOCKED => rd_mux <= lock_r;
      when others   => rd_mux <= (others => '0');
    end case;
  end process;
  rdata <= rd_mux;

  process (clk, rst_n)
    variable ctrl_v, lock_v : std_logic_vector(15 downto 0);
    variable stat_v, evw_v  : std_logic_vector(N_EVT-1 downto 0);
    variable stat_pre       : std_logic_vector(N_EVT-1 downto 0);
    variable shadow_v, live_v : std_logic_vector(31 downto 0);
    variable shadow_pre     : std_logic_vector(31 downto 0);
    variable ack_v, er_v    : std_logic;
    variable ec_v           : std_logic_vector(1 downto 0);
    variable d_rcl          : integer;
  begin
    if rst_n = '0' then
      ctrl_r   <= (others => '0');
      lock_r   <= (others => '0');
      stat_r   <= (others => '0');
      evw_r    <= (others => '0');
      shadow_r <= (others => '0');
      live_r   <= (others => '0');
      ack_r <= '0'; er_r <= '0'; ec_r <= E_NONE;
      rd_c  <= (others => '0'); wr_c <= (others => '0');
      cm_c  <= (others => '0'); rcl_c <= (others => '0');
      lk_c  <= (others => '0'); ba_c <= (others => '0');
    elsif rising_edge(clk) then
      ctrl_v     := ctrl_r;
      lock_v     := lock_r;
      stat_v     := stat_r;
      evw_v      := evw_r;
      stat_pre   := stat_r;
      shadow_v   := shadow_r;
      shadow_pre := shadow_r;
      live_v     := live_r;
      ack_v := '0'; er_v := '0'; ec_v := E_NONE;
      d_rcl := 0;

      if eot = '1' then
        -- Nothing: the counters are the report.
        null;
      else
        -- ---- The event sources come FIRST, and they SET. ----
        --
        -- Before any clear, so that an event arriving on the same cycle as a
        -- read or a write survives it. This is chapter 26.4's rule applied to
        -- two different clear policies at once.
        for i in 0 to N_EVT-1 loop
          if evt(i) = '1' then
            stat_v(i) := '1';
            evw_v(i)  := '1';
          end if;
        end loop;

        if rd = '1' and wr = '1' then
          -- A simultaneous read and write of the same address is not
          -- something a single-port interface can do, and accepting it
          -- quietly gives the driver a value that belongs to neither.
          er_v := '1'; ec_v := E_BADADDR;
          ba_c <= ba_c + 1;
        elsif rd = '1' then
          ack_v := '1';
          rd_c  <= rd_c + 1;
          if addr = A_RSVD then
            er_v := '1'; ec_v := E_BADADDR;
            ba_c <= ba_c + 1;
          elsif addr = A_STATUS then
            -- ---- READ-TO-CLEAR, and the whole trap in one line. ----
            --
            -- Clear only the bits that were RETURNED. An event that arrived
            -- on this very cycle is in stat_v and was NOT in rd_mux, so
            -- clearing the whole register would destroy an interrupt nobody
            -- ever saw.
            for i in 0 to N_EVT-1 loop
              if stat_pre(i) = '1' and evt(i) = '0' then stat_v(i) := '0'; end if;
            end loop;
            -- The count of events that WOULD have been lost by a naive
            -- clear, so the hazard is visible in a run rather than argued
            -- about in a review.
            for i in 0 to N_EVT-1 loop
              if evt(i) = '1' and stat_pre(i) = '0' then d_rcl := d_rcl + 1; end if;
            end loop;
          end if;
        elsif wr = '1' then
          ack_v := '1';
          wr_c  <= wr_c + 1;
          case addr is
            when A_CTRL =>
              -- Reserved bits ignore writes as well as reading zero. The two
              -- halves of that rule have to agree or a read-modify-write
              -- walks the register somewhere new every time round.
              ctrl_v := wdata and CTRL_MASK;
            when A_STATUS =>
              -- A read-to-clear register is not writable. Accepting the
              -- write silently lets a driver that treats it like the W1C
              -- register clear events it never read.
              er_v := '1'; ec_v := E_BADADDR;
              ba_c <= ba_c + 1;
            when A_EVENT =>
              -- ---- WRITE-1-TO-CLEAR. Set still wins. ----
              for i in 0 to N_EVT-1 loop
                if wdata(i) = '1' and evt(i) = '0' then evw_v(i) := '0'; end if;
              end loop;
            when A_CFG_LO => shadow_v(15 downto 0)  := wdata;
            when A_CFG_HI => shadow_v(31 downto 16) := wdata;
            when A_COMMIT =>
              -- ---- The whole 32 bits become live in ONE cycle. ----
              --
              -- Which is the only way to make the change atomic across a
              -- 16-bit interface. Writing the halves live means the engine
              -- runs on a configuration nobody asked for for one bus cycle,
              -- and one bus cycle is thousands of bit times on a USB link.
              live_v := shadow_pre;
              cm_c   <= cm_c + 1;
            when A_LOCKED =>
              if busy = '1' then
                -- Reported, not ignored. A silently dropped write leaves the
                -- driver believing a setting took effect, and the next thing
                -- it does is built on that belief.
                er_v := '1'; ec_v := E_LOCKED;
                lk_c <= lk_c + 1;
              else
                lock_v := wdata;
              end if;
            when others =>
              er_v := '1'; ec_v := E_BADADDR;
              ba_c <= ba_c + 1;
          end case;
        end if;
      end if;

      ctrl_r   <= ctrl_v;
      lock_r   <= lock_v;
      stat_r   <= stat_v;
      evw_r    <= evw_v;
      shadow_r <= shadow_v;
      live_r   <= live_v;
      ack_r <= ack_v; er_r <= er_v; ec_r <= ec_v;
      rcl_c <= rcl_c + to_unsigned(d_rcl, 32);
    end if;
  end process;
end architecture;

10. Seeing the Two Clear Policies Diverge

The same event, two registers, one read

usb_reg_file — read-to-clear beside write-1-to-clear

10 cycles
A ten-cycle waveform. Two events set bits 0 and 2 in both the status and event registers. A read of the status register returns 0x5 and clears it, leaving the event register unchanged at 0x5. A write of 1 to bit 0 of the event register clears only that bit, leaving bit 2 set.the read returns 0x5 and empties itthe read returns 0x5 andempties itW1C untouched: nobody wrote to itW1C untouched: nobody wroteto itwrite 1 to bit 0: bit 2 surviveswrite 1 to bit 0: bit 2survivesclkevt5000000000addr00STATUSSTATUSEVENTEVENTEVENTEVENTEVENTEVENTrdwrrdata0055555555status0550000000event_w1c0555544444t0t1t2t3t4t5t6t7t8t9
Events 0 and 2 set both registers. Reading STATUS empties it; the W1C register is untouched, because nobody wrote to it. The driver then clears exactly bit 0 of EVENT and bit 2 survives — which is the property read-to-clear cannot offer.

And the trap, which only the read-to-clear register has:

An event arriving on the read cycle

usb_reg_file — the read-to-clear hazard

10 cycles
A ten-cycle waveform. Bit 0 of the status register is set by an event. On a later cycle the driver reads the status register while a second event sets bit 1. The read returns 0x1, bit 0 is cleared and bit 1 remains set, the read-to-clear hazard counter advances to one, and the write-1-to-clear register still holds both bits.read returns 0x1; bit 1 fires NOWread returns 0x1; bit 1fires NOWbit 1 survives: it was never returnedbit 1 survives: it wasnever returnedW1C kept both, having been asked for neitherW1C kept both, having beenasked for neitherclkevt1020000000rdrdata0011111111status0112222222event_w1c0113333333n_rc_lost0001111111t0t1t2t3t4t5t6t7t8t9
Bit 0 is already set when the driver reads; bit 1 fires on the same cycle. The read returns bit 0 only, so bit 0 is cleared and bit 1 must survive — it was never handed to anybody. The W1C register keeps both, because the driver has not written to it at all.

11. The Testbenches

The oracle is a shadow model written from sections 2 to 5 rather than from the RTL, re-derived every cycle and compared against every output.

The exhaustive claim is over three things at once:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   8 addresses x 3 accesses (read / write / idle) x 2 engine states
   = 48, all required

Sweeping addresses alone misses the locked register; sweeping accesses alone misses that reading one register and writing it are different operations with different side effects; sweeping the engine state alone misses everything.

Verilog-2005 testbench

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
`timescale 1ns/1ps
// Testbench for usb_reg_file.
//
// The oracle is a shadow model written from the chapter's rules rather than
// from the RTL, re-derived every cycle and compared against every output.
//
// THE EXHAUSTIVE CLAIM IS OVER (ADDRESS, ACCESS, ENGINE STATE)
//
// A register interface's behaviour is a function of three things at once,
// and every one of them changes the answer:
//
//     8 addresses  x  3 accesses (read / write / idle)  x  2 engine states
//     = 48, all required
//
// Sweeping addresses alone misses the locked register; sweeping accesses
// alone misses that reading one register and writing it are different
// operations with different side effects; sweeping the engine state alone
// misses everything.
module tb_rf_v;

  localparam integer N_EVT = 4;

  localparam [2:0] A_CTRL = 3'd0, A_STATUS = 3'd1, A_EVENT = 3'd2,
                   A_CFG_LO = 3'd3, A_CFG_HI = 3'd4, A_COMMIT = 3'd5,
                   A_LOCKED = 3'd6, A_RSVD = 3'd7;

  localparam [1:0] E_NONE = 2'd0, E_LOCKED = 2'd1, E_BADADDR = 2'd2;
  localparam [15:0] CTRL_MASK = 16'h00FF;

  reg         clk = 1'b0, rst_n = 1'b0;
  reg  [2:0]  addr = 3'd0;
  reg         rd = 1'b0, wr = 1'b0, busy = 1'b0, eot = 1'b0;
  reg  [15:0] wdata = 16'd0;
  reg  [N_EVT-1:0] evt = {N_EVT{1'b0}};

  wire [15:0] rdata, ctrl;
  wire        ack, err_pulse;
  wire [1:0]  err_code;
  wire [N_EVT-1:0] status, event_w1c;
  wire [31:0] cfg_live, cfg_shadow;
  wire [31:0] n_read, n_write, n_commit, n_rc_lost, n_locked, n_badaddr;

  usb_reg_file #(.N_EVT(N_EVT)) dut (
    .clk(clk), .rst_n(rst_n),
    .addr(addr), .rd(rd), .wr(wr), .wdata(wdata), .evt(evt), .busy(busy),
    .eot(eot),
    .rdata(rdata), .ack(ack), .err_pulse(err_pulse), .err_code(err_code),
    .ctrl(ctrl), .status(status), .event_w1c(event_w1c),
    .cfg_live(cfg_live), .cfg_shadow(cfg_shadow),
    .n_read(n_read), .n_write(n_write), .n_commit(n_commit),
    .n_rc_lost(n_rc_lost), .n_locked(n_locked), .n_badaddr(n_badaddr)
  );

  always #5 clk = ~clk;

  // ---------------- the shadow model ----------------
  reg [15:0] m_ctrl, m_lock;
  reg [N_EVT-1:0] m_stat, m_evw;
  reg [31:0] m_shadow, m_live;
  reg        m_ack, m_er;
  reg [1:0]  m_ec;
  reg [31:0] c_rd, c_wr, c_cm, c_rcl, c_lk, c_ba;

  integer errors = 0, checks = 0, steps = 0;
  integer k;

  // reach: address (8) x access (3) x busy (2)
  reg [0:0] reach [0:47];
  integer   n_reach;

  task ck;
    input [255:0] nm;
    input [31:0]  got, exp;
    begin
      checks = checks + 1;
      if (got !== exp) begin
        errors = errors + 1;
        if (errors < 25)
          $display("FAIL t=%0t step=%0d %0s got=%0h exp=%0h",
                   $time, steps, nm, got, exp);
      end
    end
  endtask

  // The expected read data, from the map rather than from the design.
  function [15:0] exp_rdata;
    input [2:0] a;
    begin
      case (a)
        A_CTRL:   exp_rdata = m_ctrl & CTRL_MASK;
        A_STATUS: exp_rdata = {{(16-N_EVT){1'b0}}, m_stat};
        A_EVENT:  exp_rdata = {{(16-N_EVT){1'b0}}, m_evw};
        A_CFG_LO: exp_rdata = m_shadow[15:0];
        A_CFG_HI: exp_rdata = m_shadow[31:16];
        A_COMMIT: exp_rdata = 16'h0000;
        A_LOCKED: exp_rdata = m_lock;
        default:  exp_rdata = 16'h0000;
      endcase
    end
  endfunction

  reg [N_EVT-1:0] m_stat_pre;
  reg [31:0]      m_shadow_pre;

  task model_step;
    integer i;
    begin
      m_ack = 1'b0; m_er = 1'b0; m_ec = E_NONE;

      if (eot) begin
        // nothing
      end else begin
        // The event sources come FIRST and they SET, before any clear, so an
        // event arriving on the same cycle as a read or a write survives it.
        for (i = 0; i < N_EVT; i = i + 1)
          if (evt[i]) begin m_stat[i] = 1'b1; m_evw[i] = 1'b1; end

        if (rd && wr) begin
          m_er = 1'b1; m_ec = E_BADADDR; c_ba = c_ba + 1;
        end else if (rd) begin
          m_ack = 1'b1; c_rd = c_rd + 1;
          if (addr == A_RSVD) begin
            m_er = 1'b1; m_ec = E_BADADDR; c_ba = c_ba + 1;
          end else if (addr == A_STATUS) begin
            // READ-TO-CLEAR: clear only the bits that were RETURNED.
            for (i = 0; i < N_EVT; i = i + 1)
              if (m_stat_pre[i] && !evt[i]) m_stat[i] = 1'b0;
            for (i = 0; i < N_EVT; i = i + 1)
              if (evt[i] && !m_stat_pre[i]) c_rcl = c_rcl + 1;
          end
        end else if (wr) begin
          m_ack = 1'b1; c_wr = c_wr + 1;
          case (addr)
            A_CTRL:   m_ctrl = wdata & CTRL_MASK;
            A_STATUS: begin
              m_er = 1'b1; m_ec = E_BADADDR; c_ba = c_ba + 1;
            end
            A_EVENT: begin
              for (i = 0; i < N_EVT; i = i + 1)
                if (wdata[i] && !evt[i]) m_evw[i] = 1'b0;
            end
            A_CFG_LO: m_shadow[15:0]  = wdata;
            A_CFG_HI: m_shadow[31:16] = wdata;
            A_COMMIT: begin
              m_live = m_shadow_pre;
              c_cm   = c_cm + 1;
            end
            A_LOCKED: begin
              if (busy) begin
                m_er = 1'b1; m_ec = E_LOCKED; c_lk = c_lk + 1;
              end else m_lock = wdata;
            end
            default: begin
              m_er = 1'b1; m_ec = E_BADADDR; c_ba = c_ba + 1;
            end
          endcase
        end
      end
    end
  endtask

  task check_out;
    begin
      ck("ctrl",       {16'd0, ctrl},       {16'd0, m_ctrl & CTRL_MASK});
      ck("status",     {28'd0, status},     {28'd0, m_stat});
      ck("event_w1c",  {28'd0, event_w1c},  {28'd0, m_evw});
      // Compared in halves, to match the VHDL port: to_integer on a full
      // 32-bit unsigned overflows VHDL's INTEGER and ABORTS the run, so the
      // three benches split it and the check counts line up.
      ck("cfg_live lo",   {16'd0, cfg_live[15:0]},    {16'd0, m_live[15:0]});
      ck("cfg_live hi",   {16'd0, cfg_live[31:16]},   {16'd0, m_live[31:16]});
      ck("cfg_shadow lo", {16'd0, cfg_shadow[15:0]},  {16'd0, m_shadow[15:0]});
      ck("cfg_shadow hi", {16'd0, cfg_shadow[31:16]}, {16'd0, m_shadow[31:16]});
      ck("ack",        {31'd0, ack},        {31'd0, m_ack});
      ck("err_pulse",  {31'd0, err_pulse},  {31'd0, m_er});
      ck("err_code",   {30'd0, err_code},   {30'd0, m_ec});
      ck("n_read",    n_read,    c_rd);
      ck("n_write",   n_write,   c_wr);
      ck("n_commit",  n_commit,  c_cm);
      ck("n_rc_lost", n_rc_lost, c_rcl);
      ck("n_locked",  n_locked,  c_lk);
      ck("n_badaddr", n_badaddr, c_ba);
      // ---- the structural invariant ----
      //
      // Reserved bits read as zero, always. A driver doing read-modify-write
      // on this register writes back exactly what it read, so the two halves
      // of the reserved-bit rule have to agree or the register walks
      // somewhere new every time round.
      ck("ctrl reserved", {16'd0, ctrl & ~CTRL_MASK}, 32'd0);
    end
  endtask

  task step;
    begin
      // Let the combinational read path settle before sampling it. `addr`
      // was driven by a blocking assignment in the same time step, so
      // without this the testbench reads the mux output from BEFORE the
      // address changed -- and the mismatch looks like a decode bug in the
      // design rather than a race in the bench.
      #1;
      m_stat_pre   = m_stat;
      m_shadow_pre = m_shadow;
      // rdata is combinational and must be sampled BEFORE the edge: it is
      // the value the bus returns for this access, not the value the
      // register will hold afterwards. Checking it after the edge is
      // checking the result of the side effect.
      if (rd && !wr && !eot) ck("rdata", {16'd0, rdata}, {16'd0, exp_rdata(addr)});
      model_step;
      @(posedge clk);
      #1;
      steps = steps + 1;
      check_out;
    end
  endtask

  task acc; input [2:0] a; input r; input w; input [15:0] d;
            input [N_EVT-1:0] e; input b;
    begin
      addr = a; rd = r; wr = w; wdata = d; evt = e; busy = b; eot = 1'b0;
      step;
      rd = 1'b0; wr = 1'b0; evt = {N_EVT{1'b0}};
    end
  endtask

  task reset_all;
    begin
      rd = 1'b0; wr = 1'b0; evt = {N_EVT{1'b0}}; busy = 1'b0; eot = 1'b0;
      rst_n = 1'b0;
      @(posedge clk); #1;
      rst_n = 1'b1;
      m_ctrl = 16'd0; m_lock = 16'd0;
      m_stat = {N_EVT{1'b0}}; m_evw = {N_EVT{1'b0}};
      m_stat_pre = {N_EVT{1'b0}};
      m_shadow = 32'd0; m_live = 32'd0; m_shadow_pre = 32'd0;
      m_ack = 1'b0; m_er = 1'b0; m_ec = E_NONE;
      c_rd=0; c_wr=0; c_cm=0; c_rcl=0; c_lk=0; c_ba=0;
      @(negedge clk);
    end
  endtask

  integer a, ac, bs, i, w, base, n_race_cases;

  initial begin
    for (k = 0; k < 48; k = k + 1) reach[k] = 1'b0;
    n_race_cases = 0;

    repeat (3) @(posedge clk);
    rst_n = 1'b1;
    @(negedge clk);
    reset_all;

    // ================= PHASE 1 -- address x access x engine state ========
    for (a = 0; a < 8; a = a + 1)
      for (ac = 0; ac < 3; ac = ac + 1)
        for (bs = 0; bs < 2; bs = bs + 1) begin
          acc(a[2:0], (ac == 0), (ac == 1), 16'hA5A5, {N_EVT{1'b0}}, bs[0]);
          reach[a * 6 + ac * 2 + bs] = 1'b1;
        end

    // ================= PHASE 2 -- read-to-clear versus write-1-to-clear ===
    //
    // The same event feeds both registers. The driver reads STATUS (which
    // empties it) and reads EVENT (which does not), and the difference is
    // the whole argument.
    reset_all;
    acc(A_CTRL, 1'b0, 1'b0, 16'd0, 4'b0101, 1'b0);     // events 0 and 2 fire
    if (status !== 4'b0101 || event_w1c !== 4'b0101) begin
      errors = errors + 1;
      $display("FAIL the event did not reach both registers");
    end
    acc(A_STATUS, 1'b1, 1'b0, 16'd0, 4'b0000, 1'b0);   // read STATUS
    if (status !== 4'b0000) begin
      errors = errors + 1;
      $display("FAIL read-to-clear did not clear: %b", status);
    end
    if (event_w1c !== 4'b0101) begin
      errors = errors + 1;
      $display("FAIL reading STATUS disturbed the W1C register");
    end
    // ...and the W1C register clears only what the driver writes back.
    acc(A_EVENT, 1'b0, 1'b1, 16'h0001, 4'b0000, 1'b0);
    if (event_w1c !== 4'b0100) begin
      errors = errors + 1;
      $display("FAIL W1C cleared the wrong bits: %b", event_w1c);
    end

    // ================= PHASE 3 -- the two races, over EVERY bit ===========
    //
    // An event arrives on the very cycle the driver clears. It was not in the
    // value returned (read-to-clear) or in the value written back (W1C), so
    // destroying it loses an interrupt nobody ever saw -- and nothing else in
    // the system will ever mention it again.
    //
    // The first version of this phase drove ONE bit with ONE pre-state, and
    // the mutation that clears the whole register died on TWO checks. Both
    // races are now driven for every event bit and both pre-states: whether
    // the bit was already set before the racing event or not.
    //
    //     4 bits x 2 pre-states x 2 clear policies = 16 scenarios
    n_race_cases = 0;
    for (a = 0; a < N_EVT; a = a + 1)
      for (bs = 0; bs < 2; bs = bs + 1)
        for (ac = 0; ac < 2; ac = ac + 1) begin
          reset_all;
          // optionally pre-set the bit that is about to race
          if (bs == 1)
            acc(A_CTRL, 1'b0, 1'b0, 16'd0, (4'd1 << a), 1'b0);
          // ...and always set a DIFFERENT bit, so there is something for the
          // clear to legitimately remove. Without it a correct design and a
          // design that clears everything are indistinguishable.
          acc(A_CTRL, 1'b0, 1'b0, 16'd0, (4'd1 << ((a + 1) % N_EVT)), 1'b0);
          base = c_rcl;
          if (ac == 0) begin
            // read-to-clear, with the event arriving on the read cycle
            acc(A_STATUS, 1'b1, 1'b0, 16'd0, (4'd1 << a), 1'b0);
            if (!status[a]) begin
              errors = errors + 1;
              $display("FAIL RC race bit %0d pre=%0d: the event was destroyed",
                       a, bs);
            end
            // the OTHER bit was returned, so it must be gone
            if (status[(a + 1) % N_EVT] && ((a + 1) % N_EVT != a)) begin
              errors = errors + 1;
              $display("FAIL RC race bit %0d: the returned bit survived", a);
            end
            // and the hazard is counted exactly when the racing bit was NOT
            // already set -- because only then was it absent from the value
            // the driver got
            if (c_rcl != base + ((bs == 0) ? 1 : 0)) begin
              errors = errors + 1;
              $display("FAIL RC hazard count bit %0d pre=%0d: %0d",
                       a, bs, c_rcl - base);
            end
          end else begin
            // write-1-to-clear, with the event arriving on the write cycle
            acc(A_EVENT, 1'b0, 1'b1, (16'd1 << a), (4'd1 << a), 1'b0);
            if (!event_w1c[a]) begin
              errors = errors + 1;
              $display("FAIL W1C race bit %0d pre=%0d: set lost the race",
                       a, bs);
            end
            // ...and the bit nobody wrote back is untouched
            if (!event_w1c[(a + 1) % N_EVT]) begin
              errors = errors + 1;
              $display("FAIL W1C race bit %0d: an unwritten bit was cleared",
                       a);
            end
            // ...and with no concurrent event the same write does clear it
            acc(A_EVENT, 1'b0, 1'b1, (16'd1 << a), 4'd0, 1'b0);
            if (event_w1c[a]) begin
              errors = errors + 1;
              $display("FAIL W1C bit %0d: a plain write did not clear", a);
            end
          end
          n_race_cases = n_race_cases + 1;
        end
    if (n_race_cases != 16) begin
      errors = errors + 1;
      $display("FAIL race scenarios %0d, expected 16", n_race_cases);
    end

    // ================= PHASE 4 -- the shadow is ATOMIC ====================
    //
    // Two halves staged, then applied in one cycle. Between the two stores
    // the live configuration must not change at all -- for one bus cycle the
    // engine would otherwise be running on a value nobody asked for, and one
    // bus cycle is thousands of bit times on a USB link.
    reset_all;
    acc(A_CFG_LO, 1'b0, 1'b1, 16'hBEEF, 4'b0000, 1'b0);
    if (cfg_live !== 32'd0) begin
      errors = errors + 1;
      $display("FAIL writing CFG_LO changed the live configuration");
    end
    acc(A_CFG_HI, 1'b0, 1'b1, 16'hDEAD, 4'b0000, 1'b0);
    if (cfg_live !== 32'd0) begin
      errors = errors + 1;
      $display("FAIL writing CFG_HI changed the live configuration");
    end
    if (cfg_shadow !== 32'hDEAD_BEEF) begin
      errors = errors + 1;
      $display("FAIL the shadow is 0x%0h, expected 0xDEADBEEF", cfg_shadow);
    end
    base = n_commit;
    acc(A_COMMIT, 1'b0, 1'b1, 16'h0001, 4'b0000, 1'b0);
    if (cfg_live !== 32'hDEAD_BEEF) begin
      errors = errors + 1;
      $display("FAIL the commit did not apply the shadow: 0x%0h", cfg_live);
    end
    if (n_commit != base + 1) begin
      errors = errors + 1;
      $display("FAIL the commit was not counted");
    end

    // ================= PHASE 5 -- a rejected write is REPORTED ============
    //
    // Silently dropping it is the worst option: the driver believes the
    // setting took effect and the next thing it does is built on that
    // belief.
    reset_all;
    acc(A_LOCKED, 1'b0, 1'b1, 16'h1234, 4'b0000, 1'b0);   // idle: accepted
    if (rdata !== 16'h0000) ;   // rdata is for the access being driven
    base = n_locked;
    acc(A_LOCKED, 1'b0, 1'b1, 16'h5678, 4'b0000, 1'b1);   // busy: rejected
    if (n_locked != base + 1) begin
      errors = errors + 1;
      $display("FAIL a write while busy was not reported");
    end
    acc(A_LOCKED, 1'b1, 1'b0, 16'd0, 4'b0000, 1'b0);
    if (rdata !== 16'h1234) begin
      errors = errors + 1;
      $display("FAIL the rejected write took effect anyway: 0x%0h", rdata);
    end

    // ================= PHASE 6 -- reserved bits ==========================
    //
    // They read as zero AND ignore writes. The two halves of that rule have
    // to agree, because a driver doing read-modify-write writes back exactly
    // what it read.
    reset_all;
    acc(A_CTRL, 1'b0, 1'b1, 16'hFFFF, 4'b0000, 1'b0);
    if (ctrl !== 16'h00FF) begin
      errors = errors + 1;
      $display("FAIL reserved bits took a write: 0x%0h", ctrl);
    end
    acc(A_CTRL, 1'b1, 1'b0, 16'd0, 4'b0000, 1'b0);
    if (rdata !== 16'h00FF) begin
      errors = errors + 1;
      $display("FAIL reserved bits did not read as zero: 0x%0h", rdata);
    end
    // ...and a read-modify-write cycle is a fixed point: write back what was
    // read and nothing changes.
    acc(A_CTRL, 1'b0, 1'b1, 16'h00FF, 4'b0000, 1'b0);
    if (ctrl !== 16'h00FF) begin
      errors = errors + 1;
      $display("FAIL a read-modify-write cycle moved the register");
    end

    // ================= PHASE 7 -- the two rejections ======================
    reset_all;
    base = n_badaddr;
    acc(A_RSVD, 1'b1, 1'b0, 16'd0, 4'b0000, 1'b0);       // read a hole
    if (n_badaddr != base + 1) begin
      errors = errors + 1;
      $display("FAIL reading a hole in the map was not reported");
    end
    base = n_badaddr;
    acc(A_STATUS, 1'b0, 1'b1, 16'hFFFF, 4'b0000, 1'b0);  // write an RC reg
    if (n_badaddr != base + 1) begin
      errors = errors + 1;
      $display("FAIL writing a read-to-clear register was not reported");
    end
    base = n_badaddr;
    acc(A_CTRL, 1'b1, 1'b1, 16'd0, 4'b0000, 1'b0);       // read AND write
    if (n_badaddr != base + 1) begin
      errors = errors + 1;
      $display("FAIL a simultaneous read and write was not reported");
    end

    // The random phase is switchable, because a mutation score is only
    // interesting once it is DECOMPOSED. The directed phases already reach
    // every situation the exhaustiveness proof requires, so nothing in that
    // claim depends on it.
`ifndef DIRECTED_ONLY
    // ================= PHASE 8 -- random =================================
    //
    // The event pattern is deliberately correlated with the accesses: an
    // event drawn independently of the read almost never lands on the same
    // cycle, and that coincidence is the read-to-clear hazard.
    reset_all;
    for (i = 0; i < 30000; i = i + 1) begin
      w = $unsigned($random) % 100;
      acc($unsigned($random) % 8,
          (w < 40), (w >= 40) && (w < 85),
          $unsigned($random) % 65536,
          (w < 55) ? ($unsigned($random) % 16) : 4'd0,
          ($unsigned($random) % 100) < 25);
    end

`endif

    // ================= the exhaustiveness proof ==========================
    n_reach = 0;
    for (k = 0; k < 48; k = k + 1) n_reach = n_reach + reach[k];
    if (n_reach != 48) begin
      errors = errors + 1;
      $display("FAIL address x access x busy reach %0d/48", n_reach);
      for (k = 0; k < 48; k = k + 1)
        if (!reach[k])
          $display("  unreached addr=%0d access=%0d busy=%0d",
                   k / 6, (k % 6) / 2, k % 2);
    end

    $display("steps=%0d checks=%0d reach=%0d/48 errors=%0d",
             steps, checks, n_reach, errors);
    $display("read=%0d write=%0d commit=%0d rc_lost=%0d locked=%0d badaddr=%0d",
             n_read, n_write, n_commit, n_rc_lost, n_locked, n_badaddr);
    $display("%0s: %0d errors in %0d checks",
             (errors == 0) ? "PASS" : "FAIL", errors, checks);
    $finish;
  end
endmodule

SystemVerilog testbench

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
`timescale 1ns/1ps
// Testbench for usb_reg_file.
//
// The oracle is a shadow model written from the chapter's rules rather than
// from the RTL, re-derived every cycle and compared against every output.
//
// THE EXHAUSTIVE CLAIM IS OVER (ADDRESS, ACCESS, ENGINE STATE)
//
// A register interface's behaviour is a function of three things at once,
// and every one of them changes the answer:
//
//     8 addresses  x  3 accesses (read / write / idle)  x  2 engine states
//     = 48, all required
//
// Sweeping addresses alone misses the locked register; sweeping accesses
// alone misses that reading one register and writing it are different
// operations with different side effects; sweeping the engine state alone
// misses everything.
module tb_rf_sv;
  import usb_reg_pkg::*;

  localparam int N_EVT = 4;

  localparam logic [15:0] CTRL_MASK = 16'h00FF;

  logic        clk = 1'b0, rst_n = 1'b0;
  reg_addr_e   addr = A_CTRL;
  logic        rd = 1'b0, wr = 1'b0, busy = 1'b0, eot = 1'b0;
  logic [15:0] wdata = 16'd0;
  logic [N_EVT-1:0] evt = '0;

  logic [15:0] rdata, ctrl;
  logic        ack, err_pulse;
  reg_err_e    err_code;
  logic [N_EVT-1:0] status, event_w1c;
  logic [31:0] cfg_live, cfg_shadow;
  logic [31:0] n_read, n_write, n_commit, n_rc_lost, n_locked, n_badaddr;

  usb_reg_file #(.N_EVT(N_EVT)) dut (
    .clk(clk), .rst_n(rst_n),
    .addr(addr), .rd(rd), .wr(wr), .wdata(wdata), .evt(evt), .busy(busy),
    .eot(eot),
    .rdata(rdata), .ack(ack), .err_pulse(err_pulse), .err_code(err_code),
    .ctrl(ctrl), .status(status), .event_w1c(event_w1c),
    .cfg_live(cfg_live), .cfg_shadow(cfg_shadow),
    .n_read(n_read), .n_write(n_write), .n_commit(n_commit),
    .n_rc_lost(n_rc_lost), .n_locked(n_locked), .n_badaddr(n_badaddr)
  );

  always #5 clk = ~clk;

  // ---------------- the shadow model ----------------
  logic [15:0] m_ctrl, m_lock;
  logic [N_EVT-1:0] m_stat, m_evw;
  logic [31:0] m_shadow, m_live;
  logic        m_ack, m_er;
  reg_err_e    m_ec;
  int unsigned c_rd, c_wr, c_cm, c_rcl, c_lk, c_ba;

  int errors = 0, checks = 0, steps = 0;
  int k;

  // reach: address (8) x access (3) x busy (2)
  bit reach [48];
  int n_reach;

  task automatic ck(string nm, int unsigned got, int unsigned exp);
    checks++;
    if (got !== exp) begin
      errors++;
      if (errors < 25)
        $display("FAIL t=%0t step=%0d %0s got=%0h exp=%0h",
                 $time, steps, nm, got, exp);
    end
  endtask

  // The expected read data, from the map rather than from the design.
  function automatic logic [15:0] exp_rdata(input reg_addr_e a);
    begin
      case (a)
        A_CTRL:   exp_rdata = m_ctrl & CTRL_MASK;
        A_STATUS: exp_rdata = {{(16-N_EVT){1'b0}}, m_stat};
        A_EVENT:  exp_rdata = {{(16-N_EVT){1'b0}}, m_evw};
        A_CFG_LO: exp_rdata = m_shadow[15:0];
        A_CFG_HI: exp_rdata = m_shadow[31:16];
        A_COMMIT: exp_rdata = 16'h0000;
        A_LOCKED: exp_rdata = m_lock;
        default:  exp_rdata = 16'h0000;
      endcase
    end
  endfunction

  logic [N_EVT-1:0] m_stat_pre;
  logic [31:0]      m_shadow_pre;

  task automatic model_step;
    int i;
    begin
      m_ack = 1'b0; m_er = 1'b0; m_ec = E_NONE;

      if (eot) begin
        // nothing
      end else begin
        // The event sources come FIRST and they SET, before any clear, so an
        // event arriving on the same cycle as a read or a write survives it.
        for (i = 0; i < N_EVT; i = i + 1)
          if (evt[i]) begin m_stat[i] = 1'b1; m_evw[i] = 1'b1; end

        if (rd && wr) begin
          m_er = 1'b1; m_ec = E_BADADDR; c_ba = c_ba + 1;
        end else if (rd) begin
          m_ack = 1'b1; c_rd = c_rd + 1;
          if (addr == A_RSVD) begin
            m_er = 1'b1; m_ec = E_BADADDR; c_ba = c_ba + 1;
          end else if (addr == A_STATUS) begin
            // READ-TO-CLEAR: clear only the bits that were RETURNED.
            for (i = 0; i < N_EVT; i = i + 1)
              if (m_stat_pre[i] && !evt[i]) m_stat[i] = 1'b0;
            for (i = 0; i < N_EVT; i = i + 1)
              if (evt[i] && !m_stat_pre[i]) c_rcl = c_rcl + 1;
          end
        end else if (wr) begin
          m_ack = 1'b1; c_wr = c_wr + 1;
          case (addr)
            A_CTRL:   m_ctrl = wdata & CTRL_MASK;
            A_STATUS: begin
              m_er = 1'b1; m_ec = E_BADADDR; c_ba = c_ba + 1;
            end
            A_EVENT: begin
              for (i = 0; i < N_EVT; i = i + 1)
                if (wdata[i] && !evt[i]) m_evw[i] = 1'b0;
            end
            A_CFG_LO: m_shadow[15:0]  = wdata;
            A_CFG_HI: m_shadow[31:16] = wdata;
            A_COMMIT: begin
              m_live = m_shadow_pre;
              c_cm   = c_cm + 1;
            end
            A_LOCKED: begin
              if (busy) begin
                m_er = 1'b1; m_ec = E_LOCKED; c_lk = c_lk + 1;
              end else m_lock = wdata;
            end
            default: begin
              m_er = 1'b1; m_ec = E_BADADDR; c_ba = c_ba + 1;
            end
          endcase
        end
      end
    end
  endtask

  task automatic check_out;
    begin
      ck("ctrl",       {16'd0, ctrl},       {16'd0, m_ctrl & CTRL_MASK});
      ck("status",     {28'd0, status},     {28'd0, m_stat});
      ck("event_w1c",  {28'd0, event_w1c},  {28'd0, m_evw});
      // Compared in halves, to match the VHDL port: to_integer on a full
      // 32-bit unsigned overflows VHDL's INTEGER and ABORTS the run, so the
      // three benches split it and the check counts line up.
      ck("cfg_live lo",   {16'd0, cfg_live[15:0]},    {16'd0, m_live[15:0]});
      ck("cfg_live hi",   {16'd0, cfg_live[31:16]},   {16'd0, m_live[31:16]});
      ck("cfg_shadow lo", {16'd0, cfg_shadow[15:0]},  {16'd0, m_shadow[15:0]});
      ck("cfg_shadow hi", {16'd0, cfg_shadow[31:16]}, {16'd0, m_shadow[31:16]});
      ck("ack",        {31'd0, ack},        {31'd0, m_ack});
      ck("err_pulse",  {31'd0, err_pulse},  {31'd0, m_er});
      ck("err_code",   err_code, m_ec);
      ck("n_read",    n_read,    c_rd);
      ck("n_write",   n_write,   c_wr);
      ck("n_commit",  n_commit,  c_cm);
      ck("n_rc_lost", n_rc_lost, c_rcl);
      ck("n_locked",  n_locked,  c_lk);
      ck("n_badaddr", n_badaddr, c_ba);
      // ---- the structural invariant ----
      //
      // Reserved bits read as zero, always. A driver doing read-modify-write
      // on this register writes back exactly what it read, so the two halves
      // of the reserved-bit rule have to agree or the register walks
      // somewhere new every time round.
      ck("ctrl reserved", {16'd0, ctrl & ~CTRL_MASK}, 32'd0);
    end
  endtask

  task automatic step;
    begin
      // Let the combinational read path settle before sampling it. `addr`
      // was driven by a blocking assignment in the same time step, so
      // without this the testbench reads the mux output from BEFORE the
      // address changed -- and the mismatch looks like a decode bug in the
      // design rather than a race in the bench.
      #1;
      m_stat_pre   = m_stat;
      m_shadow_pre = m_shadow;
      // rdata is combinational and must be sampled BEFORE the edge: it is
      // the value the bus returns for this access, not the value the
      // register will hold afterwards. Checking it after the edge is
      // checking the result of the side effect.
      if (rd && !wr && !eot) ck("rdata", {16'd0, rdata}, {16'd0, exp_rdata(addr)});
      model_step;
      @(posedge clk);
      #1;
      steps = steps + 1;
      check_out;
    end
  endtask

  task automatic acc(reg_addr_e a, logic r, logic w, logic [15:0] d,
                     logic [N_EVT-1:0] e, logic b);
    begin
      addr = a; rd = r; wr = w; wdata = d; evt = e; busy = b; eot = 1'b0;
      step;
      rd = 1'b0; wr = 1'b0; evt = '0;
    end
  endtask

  task automatic reset_all;
    begin
      rd = 1'b0; wr = 1'b0; evt = '0; busy = 1'b0; eot = 1'b0;
      rst_n = 1'b0;
      @(posedge clk); #1;
      rst_n = 1'b1;
      m_ctrl = 16'd0; m_lock = 16'd0;
      m_stat = '0; m_evw = '0;
      m_stat_pre = '0;
      m_shadow = 32'd0; m_live = 32'd0; m_shadow_pre = 32'd0;
      m_ack = 1'b0; m_er = 1'b0; m_ec = E_NONE;
      c_rd=0; c_wr=0; c_cm=0; c_rcl=0; c_lk=0; c_ba=0;
      @(negedge clk);
    end
  endtask

  integer a, ac, bs, i, w, base, n_race_cases;

  initial begin
    foreach (reach[q]) reach[q] = 1'b0;
    n_race_cases = 0;

    repeat (3) @(posedge clk);
    rst_n = 1'b1;
    @(negedge clk);
    reset_all;

    // ================= PHASE 1 -- address x access x engine state ========
    for (a = 0; a < 8; a = a + 1)
      for (ac = 0; ac < 3; ac = ac + 1)
        for (bs = 0; bs < 2; bs = bs + 1) begin
          acc(reg_addr_e'(a[2:0]), (ac == 0), (ac == 1), 16'hA5A5, '0, bs[0]);
          reach[a * 6 + ac * 2 + bs] = 1'b1;
        end

    // ================= PHASE 2 -- read-to-clear versus write-1-to-clear ===
    //
    // The same event feeds both registers. The driver reads STATUS (which
    // empties it) and reads EVENT (which does not), and the difference is
    // the whole argument.
    reset_all;
    acc(A_CTRL, 1'b0, 1'b0, 16'd0, 4'b0101, 1'b0);     // events 0 and 2 fire
    if (status !== 4'b0101 || event_w1c !== 4'b0101) begin
      errors = errors + 1;
      $display("FAIL the event did not reach both registers");
    end
    acc(A_STATUS, 1'b1, 1'b0, 16'd0, 4'b0000, 1'b0);   // read STATUS
    if (status !== 4'b0000) begin
      errors = errors + 1;
      $display("FAIL read-to-clear did not clear: %b", status);
    end
    if (event_w1c !== 4'b0101) begin
      errors = errors + 1;
      $display("FAIL reading STATUS disturbed the W1C register");
    end
    // ...and the W1C register clears only what the driver writes back.
    acc(A_EVENT, 1'b0, 1'b1, 16'h0001, 4'b0000, 1'b0);
    if (event_w1c !== 4'b0100) begin
      errors = errors + 1;
      $display("FAIL W1C cleared the wrong bits: %b", event_w1c);
    end

    // ================= PHASE 3 -- the two races, over EVERY bit ===========
    //
    // An event arrives on the very cycle the driver clears. It was not in the
    // value returned (read-to-clear) or in the value written back (W1C), so
    // destroying it loses an interrupt nobody ever saw -- and nothing else in
    // the system will ever mention it again.
    //
    // The first version of this phase drove ONE bit with ONE pre-state, and
    // the mutation that clears the whole register died on TWO checks. Both
    // races are now driven for every event bit and both pre-states: whether
    // the bit was already set before the racing event or not.
    //
    //     4 bits x 2 pre-states x 2 clear policies = 16 scenarios
    n_race_cases = 0;
    for (a = 0; a < N_EVT; a = a + 1)
      for (bs = 0; bs < 2; bs = bs + 1)
        for (ac = 0; ac < 2; ac = ac + 1) begin
          reset_all;
          // optionally pre-set the bit that is about to race
          if (bs == 1)
            acc(A_CTRL, 1'b0, 1'b0, 16'd0, 4'(4'd1 << a), 1'b0);
          // ...and always set a DIFFERENT bit, so there is something for the
          // clear to legitimately remove. Without it a correct design and a
          // design that clears everything are indistinguishable.
          acc(A_CTRL, 1'b0, 1'b0, 16'd0, 4'(4'd1 << ((a + 1) % N_EVT)), 1'b0);
          base = int'(c_rcl);
          if (ac == 0) begin
            // read-to-clear, with the event arriving on the read cycle
            acc(A_STATUS, 1'b1, 1'b0, 16'd0, 4'(4'd1 << a), 1'b0);
            if (!status[a]) begin
              errors = errors + 1;
              $display("FAIL RC race bit %0d pre=%0d: the event was destroyed",
                       a, bs);
            end
            // the OTHER bit was returned, so it must be gone
            if (status[(a + 1) % N_EVT] && ((a + 1) % N_EVT != a)) begin
              errors = errors + 1;
              $display("FAIL RC race bit %0d: the returned bit survived", a);
            end
            // and the hazard is counted exactly when the racing bit was NOT
            // already set -- because only then was it absent from the value
            // the driver got
            if (int'(c_rcl) != base + ((bs == 0) ? 1 : 0)) begin
              errors = errors + 1;
              $display("FAIL RC hazard count bit %0d pre=%0d: %0d",
                       a, bs, int'(c_rcl) - base);
            end
          end else begin
            // write-1-to-clear, with the event arriving on the write cycle
            acc(A_EVENT, 1'b0, 1'b1, 16'(16'd1 << a), 4'(4'd1 << a), 1'b0);
            if (!event_w1c[a]) begin
              errors = errors + 1;
              $display("FAIL W1C race bit %0d pre=%0d: set lost the race",
                       a, bs);
            end
            // ...and the bit nobody wrote back is untouched
            if (!event_w1c[(a + 1) % N_EVT]) begin
              errors = errors + 1;
              $display("FAIL W1C race bit %0d: an unwritten bit was cleared",
                       a);
            end
            // ...and with no concurrent event the same write does clear it
            acc(A_EVENT, 1'b0, 1'b1, 16'(16'd1 << a), 4'd0, 1'b0);
            if (event_w1c[a]) begin
              errors = errors + 1;
              $display("FAIL W1C bit %0d: a plain write did not clear", a);
            end
          end
          n_race_cases = n_race_cases + 1;
        end
    if (n_race_cases != 16) begin
      errors = errors + 1;
      $display("FAIL race scenarios %0d, expected 16", n_race_cases);
    end

    // ================= PHASE 4 -- the shadow is ATOMIC ====================
    //
    // Two halves staged, then applied in one cycle. Between the two stores
    // the live configuration must not change at all -- for one bus cycle the
    // engine would otherwise be running on a value nobody asked for, and one
    // bus cycle is thousands of bit times on a USB link.
    reset_all;
    acc(A_CFG_LO, 1'b0, 1'b1, 16'hBEEF, 4'b0000, 1'b0);
    if (cfg_live !== 32'd0) begin
      errors = errors + 1;
      $display("FAIL writing CFG_LO changed the live configuration");
    end
    acc(A_CFG_HI, 1'b0, 1'b1, 16'hDEAD, 4'b0000, 1'b0);
    if (cfg_live !== 32'd0) begin
      errors = errors + 1;
      $display("FAIL writing CFG_HI changed the live configuration");
    end
    if (cfg_shadow !== 32'hDEAD_BEEF) begin
      errors = errors + 1;
      $display("FAIL the shadow is 0x%0h, expected 0xDEADBEEF", cfg_shadow);
    end
    base = int'(n_commit);
    acc(A_COMMIT, 1'b0, 1'b1, 16'h0001, 4'b0000, 1'b0);
    if (cfg_live !== 32'hDEAD_BEEF) begin
      errors = errors + 1;
      $display("FAIL the commit did not apply the shadow: 0x%0h", cfg_live);
    end
    if (int'(n_commit) != base + 1) begin
      errors = errors + 1;
      $display("FAIL the commit was not counted");
    end

    // ================= PHASE 5 -- a rejected write is REPORTED ============
    //
    // Silently dropping it is the worst option: the driver believes the
    // setting took effect and the next thing it does is built on that
    // belief.
    reset_all;
    acc(A_LOCKED, 1'b0, 1'b1, 16'h1234, 4'b0000, 1'b0);   // idle: accepted
    if (rdata !== 16'h0000) ;   // rdata is for the access being driven
    base = int'(n_locked);
    acc(A_LOCKED, 1'b0, 1'b1, 16'h5678, 4'b0000, 1'b1);   // busy: rejected
    if (int'(n_locked) != base + 1) begin
      errors = errors + 1;
      $display("FAIL a write while busy was not reported");
    end
    acc(A_LOCKED, 1'b1, 1'b0, 16'd0, 4'b0000, 1'b0);
    if (rdata !== 16'h1234) begin
      errors = errors + 1;
      $display("FAIL the rejected write took effect anyway: 0x%0h", rdata);
    end

    // ================= PHASE 6 -- reserved bits ==========================
    //
    // They read as zero AND ignore writes. The two halves of that rule have
    // to agree, because a driver doing read-modify-write writes back exactly
    // what it read.
    reset_all;
    acc(A_CTRL, 1'b0, 1'b1, 16'hFFFF, 4'b0000, 1'b0);
    if (ctrl !== 16'h00FF) begin
      errors = errors + 1;
      $display("FAIL reserved bits took a write: 0x%0h", ctrl);
    end
    acc(A_CTRL, 1'b1, 1'b0, 16'd0, 4'b0000, 1'b0);
    if (rdata !== 16'h00FF) begin
      errors = errors + 1;
      $display("FAIL reserved bits did not read as zero: 0x%0h", rdata);
    end
    // ...and a read-modify-write cycle is a fixed point: write back what was
    // read and nothing changes.
    acc(A_CTRL, 1'b0, 1'b1, 16'h00FF, 4'b0000, 1'b0);
    if (ctrl !== 16'h00FF) begin
      errors = errors + 1;
      $display("FAIL a read-modify-write cycle moved the register");
    end

    // ================= PHASE 7 -- the two rejections ======================
    reset_all;
    base = int'(n_badaddr);
    acc(A_RSVD, 1'b1, 1'b0, 16'd0, 4'b0000, 1'b0);       // read a hole
    if (int'(n_badaddr) != base + 1) begin
      errors = errors + 1;
      $display("FAIL reading a hole in the map was not reported");
    end
    base = int'(n_badaddr);
    acc(A_STATUS, 1'b0, 1'b1, 16'hFFFF, 4'b0000, 1'b0);  // write an RC reg
    if (int'(n_badaddr) != base + 1) begin
      errors = errors + 1;
      $display("FAIL writing a read-to-clear register was not reported");
    end
    base = int'(n_badaddr);
    acc(A_CTRL, 1'b1, 1'b1, 16'd0, 4'b0000, 1'b0);       // read AND write
    if (int'(n_badaddr) != base + 1) begin
      errors = errors + 1;
      $display("FAIL a simultaneous read and write was not reported");
    end

    // The random phase is switchable, because a mutation score is only
    // interesting once it is DECOMPOSED. The directed phases already reach
    // every situation the exhaustiveness proof requires, so nothing in that
    // claim depends on it.
`ifndef DIRECTED_ONLY
    // ================= PHASE 8 -- random =================================
    //
    // The event pattern is deliberately correlated with the accesses: an
    // event drawn independently of the read almost never lands on the same
    // cycle, and that coincidence is the read-to-clear hazard.
    reset_all;
    for (i = 0; i < 30000; i = i + 1) begin
      w = $unsigned($random) % 100;
      acc(reg_addr_e'($unsigned($random) % 8),
          (w < 40), (w >= 40) && (w < 85),
          16'($unsigned($random) % 65536),
          (w < 55) ? 4'($unsigned($random) % 16) : 4'd0,
          ($unsigned($random) % 100) < 25);
    end

`endif

    // ================= the exhaustiveness proof ==========================
    n_reach = 0;
    for (k = 0; k < 48; k = k + 1) n_reach = n_reach + reach[k];
    if (n_reach != 48) begin
      errors = errors + 1;
      $display("FAIL address x access x busy reach %0d/48", n_reach);
      for (k = 0; k < 48; k = k + 1)
        if (!reach[k])
          $display("  unreached addr=%0d access=%0d busy=%0d",
                   k / 6, (k % 6) / 2, k % 2);
    end

    $display("steps=%0d checks=%0d reach=%0d/48 errors=%0d",
             steps, checks, n_reach, errors);
    $display("read=%0d write=%0d commit=%0d rc_lost=%0d locked=%0d badaddr=%0d",
             n_read, n_write, n_commit, n_rc_lost, n_locked, n_badaddr);
    $display("%0s: %0d errors in %0d checks",
             (errors == 0) ? "PASS" : "FAIL", errors, checks);
    $finish;
  end
endmodule

VHDL-2008 testbench

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
-- Testbench for usb_reg_file (VHDL-2008).
--
-- The oracle is a shadow model held in process variables and written from the
-- chapter's rules rather than from the RTL, re-derived every cycle and
-- compared against every output.
--
-- THE EXHAUSTIVE CLAIM IS OVER (ADDRESS, ACCESS, ENGINE STATE)
--
-- A register interface's behaviour is a function of three things at once, and
-- every one of them changes the answer:
--
--     8 addresses  x  3 accesses (read / write / idle)  x  2 engine states
--     = 48, all required
--
-- Sweeping addresses alone misses the locked register; sweeping accesses
-- alone misses that reading one register and writing it are different
-- operations with different side effects.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
use std.textio.all;
use work.usb_reg_pkg.all;

entity tb_rf_vhdl is
end entity;

architecture sim of tb_rf_vhdl is
  constant N_EVT : integer := 4;

  signal clk   : std_logic := '0';
  signal rst_n : std_logic := '0';
  signal addr  : std_logic_vector(2 downto 0) := "000";
  signal rd, wr, busy, eot : std_logic := '0';
  signal wdata : std_logic_vector(15 downto 0) := (others => '0');
  signal evt   : std_logic_vector(N_EVT-1 downto 0) := (others => '0');

  signal rdata_s, ctrl_s : std_logic_vector(15 downto 0);
  signal ack_s, err_s    : std_logic;
  signal ec_s            : std_logic_vector(1 downto 0);
  signal status_s, evw_s : std_logic_vector(N_EVT-1 downto 0);
  signal live_s, shadow_s : std_logic_vector(31 downto 0);
  signal n_rd_s, n_wr_s, n_cm_s : unsigned(31 downto 0);
  signal n_rcl_s, n_lk_s, n_ba_s : unsigned(31 downto 0);

  signal done : boolean := false;

  type int_array is array (natural range <>) of integer;
begin
  clk <= not clk after 5 ns when not done else '0';

  dut : entity work.usb_reg_file
    generic map (N_EVT => N_EVT)
    port map (
      clk => clk, rst_n => rst_n,
      addr => addr, rd => rd, wr => wr, wdata => wdata, evt => evt,
      busy => busy, eot => eot,
      rdata => rdata_s, ack => ack_s, err_pulse => err_s, err_code => ec_s,
      ctrl => ctrl_s, status => status_s, event_w1c => evw_s,
      cfg_live => live_s, cfg_shadow => shadow_s,
      n_read => n_rd_s, n_write => n_wr_s, n_commit => n_cm_s,
      n_rc_lost => n_rcl_s, n_locked => n_lk_s, n_badaddr => n_ba_s
    );

  stim : process
    variable m_ctrl, m_lock : std_logic_vector(15 downto 0) := (others => '0');
    variable m_stat, m_evw  : std_logic_vector(N_EVT-1 downto 0)
      := (others => '0');
    variable m_stat_pre     : std_logic_vector(N_EVT-1 downto 0)
      := (others => '0');
    variable m_shadow, m_live : std_logic_vector(31 downto 0)
      := (others => '0');
    variable m_shadow_pre   : std_logic_vector(31 downto 0) := (others => '0');
    variable m_ack, m_er    : std_logic := '0';
    variable m_ec           : std_logic_vector(1 downto 0) := E_NONE;
    variable c_rd, c_wr, c_cm, c_rcl, c_lk, c_ba : integer := 0;

    variable errors, checks, steps : integer := 0;
    variable reach : int_array(0 to 47) := (others => 0);
    variable n_reach, base, w_v, n_race_cases : integer := 0;
    variable o_v : std_logic_vector(N_EVT-1 downto 0) := (others => '0');
    variable a_v  : std_logic_vector(2 downto 0) := "000";
    variable d_v  : std_logic_vector(15 downto 0) := (others => '0');
    variable e_v  : std_logic_vector(N_EVT-1 downto 0) := (others => '0');
    variable r_v, wv_v, b_v : std_logic := '0';
    variable ln : line;

    -- A deterministic LFSR, so a rerun reproduces exactly the same traffic.
    variable lfsr : unsigned(31 downto 0) := x"5EA51DE5";

    impure function rnd_nat return integer is
    begin
      lfsr := lfsr(30 downto 0) &
              (lfsr(31) xor lfsr(21) xor lfsr(1) xor lfsr(0));
      return to_integer(lfsr(29 downto 0));
    end function;

    procedure ck (nm : string; got, exp : integer) is
    begin
      checks := checks + 1;
      if got /= exp then
        errors := errors + 1;
        if errors < 25 then
          write(ln, string'("FAIL step=") & integer'image(steps) & " " & nm
                    & " got=" & integer'image(got)
                    & " exp=" & integer'image(exp));
          writeline(output, ln);
        end if;
      end if;
    end procedure;

    procedure fail (nm : string) is
    begin
      errors := errors + 1;
      if errors < 25 then
        write(ln, string'("FAIL ") & nm); writeline(output, ln);
      end if;
    end procedure;

    function sl2i (s : std_logic) return integer is
    begin
      if s = '1' then return 1; else return 0; end if;
    end function;

    -- The expected read data, from the map rather than from the design.
    impure function exp_rdata (a : std_logic_vector(2 downto 0))
      return std_logic_vector is
    begin
      if    a = A_CTRL   then return m_ctrl and CTRL_MASK;
      elsif a = A_STATUS then return (15 downto N_EVT => '0') & m_stat;
      elsif a = A_EVENT  then return (15 downto N_EVT => '0') & m_evw;
      elsif a = A_CFG_LO then return m_shadow(15 downto 0);
      elsif a = A_CFG_HI then return m_shadow(31 downto 16);
      elsif a = A_LOCKED then return m_lock;
      else                    return x"0000";
      end if;
    end function;

    procedure model_step is
    begin
      m_ack := '0'; m_er := '0'; m_ec := E_NONE;

      if eot = '1' then
        null;
      else
        -- The event sources come FIRST and they SET, before any clear.
        for i in 0 to N_EVT-1 loop
          if evt(i) = '1' then m_stat(i) := '1'; m_evw(i) := '1'; end if;
        end loop;

        if rd = '1' and wr = '1' then
          m_er := '1'; m_ec := E_BADADDR; c_ba := c_ba + 1;
        elsif rd = '1' then
          m_ack := '1'; c_rd := c_rd + 1;
          if addr = A_RSVD then
            m_er := '1'; m_ec := E_BADADDR; c_ba := c_ba + 1;
          elsif addr = A_STATUS then
            -- READ-TO-CLEAR: clear only the bits that were RETURNED.
            for i in 0 to N_EVT-1 loop
              if m_stat_pre(i) = '1' and evt(i) = '0' then m_stat(i) := '0'; end if;
            end loop;
            for i in 0 to N_EVT-1 loop
              if evt(i) = '1' and m_stat_pre(i) = '0' then c_rcl := c_rcl + 1; end if;
            end loop;
          end if;
        elsif wr = '1' then
          m_ack := '1'; c_wr := c_wr + 1;
          if addr = A_CTRL then
            m_ctrl := wdata and CTRL_MASK;
          elsif addr = A_STATUS then
            m_er := '1'; m_ec := E_BADADDR; c_ba := c_ba + 1;
          elsif addr = A_EVENT then
            for i in 0 to N_EVT-1 loop
              if wdata(i) = '1' and evt(i) = '0' then m_evw(i) := '0'; end if;
            end loop;
          elsif addr = A_CFG_LO then
            m_shadow(15 downto 0) := wdata;
          elsif addr = A_CFG_HI then
            m_shadow(31 downto 16) := wdata;
          elsif addr = A_COMMIT then
            m_live := m_shadow_pre;
            c_cm := c_cm + 1;
          elsif addr = A_LOCKED then
            if busy = '1' then
              m_er := '1'; m_ec := E_LOCKED; c_lk := c_lk + 1;
            else
              m_lock := wdata;
            end if;
          else
            m_er := '1'; m_ec := E_BADADDR; c_ba := c_ba + 1;
          end if;
        end if;
      end if;
    end procedure;

    procedure check_out is
    begin
      ck("ctrl",       to_integer(unsigned(ctrl_s)),
                       to_integer(unsigned(m_ctrl)));
      ck("status",     to_integer(unsigned(status_s)),
                       to_integer(unsigned(m_stat)));
      ck("event_w1c",  to_integer(unsigned(evw_s)),
                       to_integer(unsigned(m_evw)));
      -- Compared in halves. to_integer on a full 32-bit unsigned overflows
      -- VHDL's INTEGER and ABORTS the run rather than returning a wrong
      -- answer -- which is better than wrapping and much worse than a
      -- failing check, because the run stops before it has told you
      -- anything.
      ck("cfg_live lo",   to_integer(unsigned(live_s(15 downto 0))),
                          to_integer(unsigned(m_live(15 downto 0))));
      ck("cfg_live hi",   to_integer(unsigned(live_s(31 downto 16))),
                          to_integer(unsigned(m_live(31 downto 16))));
      ck("cfg_shadow lo", to_integer(unsigned(shadow_s(15 downto 0))),
                          to_integer(unsigned(m_shadow(15 downto 0))));
      ck("cfg_shadow hi", to_integer(unsigned(shadow_s(31 downto 16))),
                          to_integer(unsigned(m_shadow(31 downto 16))));
      ck("ack",        sl2i(ack_s),  sl2i(m_ack));
      ck("err_pulse",  sl2i(err_s),  sl2i(m_er));
      ck("err_code",   to_integer(unsigned(ec_s)),
                       to_integer(unsigned(m_ec)));
      ck("n_read",    to_integer(n_rd_s),  c_rd);
      ck("n_write",   to_integer(n_wr_s),  c_wr);
      ck("n_commit",  to_integer(n_cm_s),  c_cm);
      ck("n_rc_lost", to_integer(n_rcl_s), c_rcl);
      ck("n_locked",  to_integer(n_lk_s),  c_lk);
      ck("n_badaddr", to_integer(n_ba_s),  c_ba);
      -- ---- the structural invariant ----
      --
      -- Reserved bits read as zero, always.
      ck("ctrl reserved",
         to_integer(unsigned(ctrl_s and not CTRL_MASK)), 0);
    end procedure;

    procedure step is
    begin
      -- Let the combinational read path settle before sampling it. `addr` was
      -- driven one delta ago, and the read mux plus the output assignment
      -- need deltas of their own -- without this the testbench reads the mux
      -- output from BEFORE the address changed, and the mismatch looks like
      -- a decode bug in the design rather than a race in the bench.
      wait for 1 ns;
      m_stat_pre   := m_stat;
      m_shadow_pre := m_shadow;
      -- rdata is combinational and must be sampled BEFORE the edge: it is the
      -- value the bus returns for this access, not the value the register
      -- will hold afterwards.
      if rd = '1' and wr = '0' and eot = '0' then
        ck("rdata", to_integer(unsigned(rdata_s)),
                    to_integer(unsigned(exp_rdata(addr))));
      end if;
      model_step;
      wait until rising_edge(clk);
      wait for 1 ns;
      steps := steps + 1;
      check_out;
    end procedure;

    procedure acc (a : std_logic_vector(2 downto 0); r, w : std_logic;
                   d : std_logic_vector(15 downto 0);
                   e : std_logic_vector(N_EVT-1 downto 0); b : std_logic) is
    begin
      addr <= a; rd <= r; wr <= w; wdata <= d; evt <= e; busy <= b; eot <= '0';
      wait for 0 ns;
      step;
      rd <= '0'; wr <= '0'; evt <= (others => '0');
    end procedure;

    procedure reset_all is
    begin
      rd <= '0'; wr <= '0'; evt <= (others => '0'); busy <= '0'; eot <= '0';
      rst_n <= '0';
      wait until rising_edge(clk);
      wait for 1 ns;
      rst_n <= '1';
      m_ctrl := (others => '0'); m_lock := (others => '0');
      m_stat := (others => '0'); m_evw := (others => '0');
      m_stat_pre := (others => '0');
      m_shadow := (others => '0'); m_live := (others => '0');
      m_shadow_pre := (others => '0');
      m_ack := '0'; m_er := '0'; m_ec := E_NONE;
      c_rd := 0; c_wr := 0; c_cm := 0; c_rcl := 0; c_lk := 0; c_ba := 0;
      wait for 1 ns;
    end procedure;
  begin
    wait until rising_edge(clk);
    wait until rising_edge(clk);
    wait until rising_edge(clk);
    rst_n <= '1';
    wait for 1 ns;
    reset_all;

    -- ================= PHASE 1 -- address x access x engine state ========
    for a in 0 to 7 loop
      for ac in 0 to 2 loop
        for bs in 0 to 1 loop
          if ac = 0 then
            if bs = 1 then acc(std_logic_vector(to_unsigned(a,3)), '1','0', x"A5A5", (others=>'0'), '1');
            else           acc(std_logic_vector(to_unsigned(a,3)), '1','0', x"A5A5", (others=>'0'), '0'); end if;
          elsif ac = 1 then
            if bs = 1 then acc(std_logic_vector(to_unsigned(a,3)), '0','1', x"A5A5", (others=>'0'), '1');
            else           acc(std_logic_vector(to_unsigned(a,3)), '0','1', x"A5A5", (others=>'0'), '0'); end if;
          else
            if bs = 1 then acc(std_logic_vector(to_unsigned(a,3)), '0','0', x"A5A5", (others=>'0'), '1');
            else           acc(std_logic_vector(to_unsigned(a,3)), '0','0', x"A5A5", (others=>'0'), '0'); end if;
          end if;
          reach(a * 6 + ac * 2 + bs) := 1;
        end loop;
      end loop;
    end loop;

    -- ================= PHASE 2 -- read-to-clear versus write-1-to-clear ===
    reset_all;
    acc(A_CTRL, '0','0', x"0000", "0101", '0');
    if status_s /= "0101" or evw_s /= "0101" then
      fail("the event did not reach both registers");
    end if;
    acc(A_STATUS, '1','0', x"0000", "0000", '0');
    if status_s /= "0000" then fail("read-to-clear did not clear"); end if;
    if evw_s /= "0101" then fail("reading STATUS disturbed the W1C register"); end if;
    acc(A_EVENT, '0','1', x"0001", "0000", '0');
    if evw_s /= "0100" then fail("W1C cleared the wrong bits"); end if;

    -- ================= PHASE 3 -- the two races, over EVERY bit ===========
    --
    -- An event arrives on the very cycle the driver clears. It was not in the
    -- value returned (read-to-clear) or in the value written back (W1C), so
    -- destroying it loses an interrupt nobody ever saw.
    --
    -- The first version of this phase drove ONE bit with ONE pre-state, and
    -- the mutation that clears the whole register died on TWO checks. Both
    -- races are now driven for every event bit and both pre-states.
    --
    --     4 bits x 2 pre-states x 2 clear policies = 16 scenarios
    n_race_cases := 0;
    for av in 0 to N_EVT-1 loop
      for bs in 0 to 1 loop
        for ac in 0 to 1 loop
          reset_all;
          e_v := (others => '0'); e_v(av) := '1';
          if bs = 1 then acc(A_CTRL, '0','0', x"0000", e_v, '0'); end if;
          -- ...and always set a DIFFERENT bit, so there is something for the
          -- clear to legitimately remove. Without it a correct design and a
          -- design that clears everything are indistinguishable.
          o_v := (others => '0'); o_v((av + 1) mod N_EVT) := '1';
          acc(A_CTRL, '0','0', x"0000", o_v, '0');
          base := c_rcl;
          if ac = 0 then
            acc(A_STATUS, '1','0', x"0000", e_v, '0');
            if status_s(av) /= '1' then
              fail("RC race: the same-cycle event was destroyed");
            end if;
            if ((av + 1) mod N_EVT) /= av
               and status_s((av + 1) mod N_EVT) /= '0' then
              fail("RC race: the returned bit survived");
            end if;
            if bs = 0 then
              if c_rcl /= base + 1 then fail("RC hazard not counted"); end if;
            else
              if c_rcl /= base then fail("RC hazard counted wrongly"); end if;
            end if;
          else
            d_v := (others => '0'); d_v(av) := '1';
            acc(A_EVENT, '0','1', d_v, e_v, '0');
            if evw_s(av) /= '1' then
              fail("W1C race: set lost the race");
            end if;
            if evw_s((av + 1) mod N_EVT) /= '1' then
              fail("W1C race: an unwritten bit was cleared");
            end if;
            acc(A_EVENT, '0','1', d_v, (others => '0'), '0');
            if evw_s(av) /= '0' then
              fail("W1C: a plain write did not clear");
            end if;
          end if;
          n_race_cases := n_race_cases + 1;
        end loop;
      end loop;
    end loop;
    if n_race_cases /= 16 then
      errors := errors + 1;
      write(ln, string'("FAIL race scenarios ") & integer'image(n_race_cases));
      writeline(output, ln);
    end if;

    -- ================= PHASE 4 -- the shadow is ATOMIC ====================
    reset_all;
    acc(A_CFG_LO, '0','1', x"BEEF", "0000", '0');
    if live_s /= x"00000000" then
      fail("writing CFG_LO changed the live configuration");
    end if;
    acc(A_CFG_HI, '0','1', x"DEAD", "0000", '0');
    if live_s /= x"00000000" then
      fail("writing CFG_HI changed the live configuration");
    end if;
    if shadow_s /= x"DEADBEEF" then fail("the shadow is wrong"); end if;
    base := c_cm;
    acc(A_COMMIT, '0','1', x"0001", "0000", '0');
    if live_s /= x"DEADBEEF" then fail("the commit did not apply the shadow"); end if;
    if c_cm /= base + 1 then fail("the commit was not counted"); end if;

    -- ================= PHASE 5 -- a rejected write is REPORTED ============
    reset_all;
    acc(A_LOCKED, '0','1', x"1234", "0000", '0');
    base := c_lk;
    acc(A_LOCKED, '0','1', x"5678", "0000", '1');
    if c_lk /= base + 1 then fail("a write while busy was not reported"); end if;
    acc(A_LOCKED, '1','0', x"0000", "0000", '0');
    if rdata_s /= x"1234" then fail("the rejected write took effect anyway"); end if;

    -- ================= PHASE 6 -- reserved bits ==========================
    reset_all;
    acc(A_CTRL, '0','1', x"FFFF", "0000", '0');
    if ctrl_s /= x"00FF" then fail("reserved bits took a write"); end if;
    acc(A_CTRL, '1','0', x"0000", "0000", '0');
    if rdata_s /= x"00FF" then fail("reserved bits did not read as zero"); end if;
    acc(A_CTRL, '0','1', x"00FF", "0000", '0');
    if ctrl_s /= x"00FF" then fail("a read-modify-write cycle moved the register"); end if;

    -- ================= PHASE 7 -- the two rejections ======================
    reset_all;
    base := c_ba;
    acc(A_RSVD, '1','0', x"0000", "0000", '0');
    if c_ba /= base + 1 then fail("reading a hole in the map was not reported"); end if;
    base := c_ba;
    acc(A_STATUS, '0','1', x"FFFF", "0000", '0');
    if c_ba /= base + 1 then
      fail("writing a read-to-clear register was not reported");
    end if;
    base := c_ba;
    acc(A_CTRL, '1','1', x"0000", "0000", '0');
    if c_ba /= base + 1 then
      fail("a simultaneous read and write was not reported");
    end if;

    -- ================= PHASE 8 -- random =================================
    --
    -- The event pattern is deliberately correlated with the accesses: an
    -- event drawn independently of the read almost never lands on the same
    -- cycle, and that coincidence is the read-to-clear hazard.
    reset_all;
    for i in 0 to 29999 loop
      -- Built into variables rather than written as conditional expressions
      -- in the argument list: those are VHDL-2019 and this file is 2008.
      w_v   := rnd_nat mod 100;
      a_v   := std_logic_vector(to_unsigned(rnd_nat mod 8, 3));
      d_v   := std_logic_vector(to_unsigned(rnd_nat mod 65536, 16));
      if w_v < 40 then r_v := '1'; else r_v := '0'; end if;
      if w_v >= 40 and w_v < 85 then wv_v := '1'; else wv_v := '0'; end if;
      if w_v < 55 then
        e_v := std_logic_vector(to_unsigned(rnd_nat mod 16, N_EVT));
      else
        e_v := (others => '0');
      end if;
      if (rnd_nat mod 100) < 25 then b_v := '1'; else b_v := '0'; end if;
      acc(a_v, r_v, wv_v, d_v, e_v, b_v);
    end loop;

    -- ================= the exhaustiveness proof ==========================
    n_reach := 0;
    for k in 0 to 47 loop n_reach := n_reach + reach(k); end loop;
    if n_reach /= 48 then
      errors := errors + 1;
      write(ln, string'("FAIL address x access x busy reach ")
                & integer'image(n_reach) & "/48");
      writeline(output, ln);
    end if;

    write(ln, string'("steps=") & integer'image(steps)
              & " checks=" & integer'image(checks)
              & " reach=" & integer'image(n_reach) & "/48"
              & " errors=" & integer'image(errors));
    writeline(output, ln);
    write(ln, string'("read=") & integer'image(to_integer(n_rd_s))
              & " write=" & integer'image(to_integer(n_wr_s))
              & " commit=" & integer'image(to_integer(n_cm_s))
              & " rc_lost=" & integer'image(to_integer(n_rcl_s))
              & " locked=" & integer'image(to_integer(n_lk_s))
              & " badaddr=" & integer'image(to_integer(n_ba_s)));
    writeline(output, ln);
    if errors = 0 then
      write(ln, string'("PASS: 0 errors in ") & integer'image(checks)
                & " checks");
    else
      write(ln, string'("FAIL: ") & integer'image(errors) & " errors in " &
                integer'image(checks) & " checks");
    end if;
    writeline(output, ln);
    done <= true;
    wait;
  end process;
end architecture;

11. Exhaustive Verification

MeasureVerilogSystemVerilogVHDL
(address × access) reached48 / 4848 / 4848 / 48
read-to-clear race scenarios swept16 / 1616 / 1616 / 16
W1C-against-a-live-source scenarios swept16 / 1616 / 1616 / 16
Steps301113011130111
Checks executed523901523901523861
reads119861198611946
writes135521355213549
commits168916891695
events arriving on the read cycle260260123
writes refused as locked429429386
writes refused as bad address485848584937
ResultPASSPASSPASS

12. Mutation Testing

#MutationVerilogSysVerVHDL
Z7the read-to-clear register accepts a write696126961269107
Z4the control register accepts reserved bits597855978559935
Z1read-to-clear clears the whole register, not the bits returned363813638137209
Z5a write to a locked register is accepted silently308363083630675
Z2a config write reaches the live register, bypassing the shadow300663006629409
Z3commit copies only the low half of the shadow299972999730015
Z6W1C clears a bit whose source is still asserting15511551591
—unmutated baseline000

All seven die in all three languages.

The one column that is out of line, and why it is not a finding

Z6 scores 1551 in Verilog and SystemVerilog and 591 in VHDL — a 2.6× spread, by far the largest in this module. Every other mutation agrees across languages to within 3%.

The rule for this situation is to divide the suspicious score by the baseline event count it depends on before concluding anything. Z6 can only be detected on a cycle where a source is still asserting while the driver writes 1 to clear it — the same enabling situation counted in the table above as "events arriving on the read cycle": 260 in Verilog and SystemVerilog, 123 in VHDL.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   Z6 score / enabling events:

     Verilog / SystemVerilog   1551 / 260  =  5.97
     VHDL                       591 / 123  =  4.80

   A 2.6x spread in the SCORE is a 1.24x spread in the RATE.

The design behaves the same way in all three languages. The stimulus produced the enabling situation half as often in VHDL, because the pseudo-random generators are different — Icarus seeds $random and $urandom identically, which is why the Verilog and SystemVerilog columns match exactly everywhere in this module, and NVC's is unrelated to either.

Directed against random

#All phasesDirected onlyRandom
Z1363811636365
Z2300662330043
Z3299971829979
Z4597859759688
Z5308361630820
Z61551161535
Z7696124969563

Every one is killed by directed stimulus alone. Z1 and Z6 both read exactly 16, and that is not a coincidence — it is the size of the phase written for them: 4 status bits × 2 pre-states × 2 clear policies, driven deliberately rather than waited for.

13. Debugging Walkthrough: The Counter That Reads Zero

The report. A driver's error statistics are always zero. The device definitely has errors — the link renegotiates, the analyser shows retries — but every counter in the register file reads 0.

Step 1 — is the hardware counting? Probe internally: yes, the counters increment.

Step 2 — so the read path is broken? No. Read the register twice in quick succession from a debugger and the first read returns a non-zero value.

Step 3 — what is between the hardware and the driver? The driver's own diagnostic thread, polling the same register once a second. The register is read-to-clear, and two readers of a read-to-clear register destroy each other's data. Whoever reads first gets everything; whoever reads second gets zero.

Step 4 — and the driver's real reader is second. So it always reads zero.

Step 5 — the deeper problem. Even with one reader, the design as originally written cleared the whole register on a read rather than only the bits it returned. An event arriving on the read cycle was returned to nobody and erased. At a 2% arrival rate that is one lost event in fifty — invisible in testing, and exactly the class of loss that shows up as an unexplained hang.

14. UVM: A Register Model That Knows Its Own Access Policies

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// A RAL model usually mirrors ADDRESSES. The bugs in this chapter are all in
// ACCESS POLICY -- read-to-clear, write-1-to-clear, shadow/commit, locked,
// reserved bits -- so the model carries the policy and checks it.
typedef enum { RW, RO, W1C, RC, SHADOW, LOCKED_RW } policy_e;

class reg_desc;
  string   name;
  bit [7:0] addr;
  policy_e pol;
  bit [31:0] wmask;        // reserved bits read back as zero

  function new(string n, bit [7:0] a, policy_e p, bit [31:0] m = '1);
    name = n; addr = a; pol = p; wmask = m;
  endfunction
endclass


class rf_predictor extends uvm_component;
  `uvm_component_utils(rf_predictor)

  reg_desc map [bit [7:0]];
  bit [31:0] model [bit [7:0]];
  bit [31:0] shadow, live;
  bit        locked;

  int unsigned n_rc_race, n_w1c_live, n_locked_refused, n_reserved_masked;

  function new(string name, uvm_component parent); super.new(name, parent);
  endfunction

  // ---- A READ that also writes. ----
  //
  // `hw_set` is the set of bits hardware asserted DURING this read. The
  // entire point of the design is that they survive it: the read returns the
  // value the mux sampled, and the clear applies to THAT value -- never to
  // the register as it stands after the event landed.
  function bit [31:0] predict_read(bit [7:0] a, bit [31:0] hw_set);
    bit [31:0] returned;
    reg_desc d = map[a];

    if (d == null) begin
      `uvm_error("RF/ADDR", $sformatf("read of unmapped address 0x%02h", a))
      return 32'h0;
    end

    returned = model[a];

    if (d.pol == RC) begin
      // Clear ONLY what was returned. A bit set by hardware this cycle is in
      // `hw_set` and not in `returned`, so it stays.
      model[a] = (model[a] & ~returned) | hw_set;
      if ((hw_set & ~returned) != 0) n_rc_race++;
    end else begin
      model[a] = model[a] | hw_set;
    end

    return returned;
  endfunction

  function void predict_write(bit [7:0] a, bit [31:0] wdata, bit [31:0] hw_live);
    reg_desc d = map[a];

    if (d == null) begin
      `uvm_error("RF/ADDR", $sformatf("write to unmapped address 0x%02h", a))
      return;
    end

    // A read-to-clear register is NOT writable. Accepting the write lets a
    // driver that treats it like the W1C register clear events it never read
    // -- the single most destructive thing software can do to this block.
    if (d.pol == RC || d.pol == RO) begin
      `uvm_error("RF/POL",
        $sformatf("write to %s, which is %s and must refuse it", d.name, d.pol.name()))
      return;
    end

    if (locked && d.pol == LOCKED_RW) begin
      // Refused, and REPORTED. A lock that silently drops writes is worse
      // than no lock: software believes it reconfigured the device.
      n_locked_refused++;
      return;
    end

    case (d.pol)
      W1C: begin
        // Clear only bits whose source is NOT still asserting. Clearing a bit
        // that hardware is holding produces a register that disagrees with
        // the hardware until the next event -- and a driver that reads it
        // concludes the condition went away.
        foreach (wdata[i])
          if (wdata[i] && !hw_live[i]) model[a][i] = 1'b0;
        if ((wdata & hw_live) != 0) n_w1c_live++;
      end

      SHADOW: begin
        // The write lands in the shadow and NOWHERE else. Nothing the device
        // is currently doing changes until commit -- which is the entire
        // reason a shadow exists: a 32-bit configuration written through a
        // 16-bit bus is briefly half-old, and a device that acts on the
        // half-written value acts on a configuration nobody asked for.
        shadow = wdata;
      end

      default: begin
        if ((wdata & ~d.wmask) != 0) n_reserved_masked++;
        model[a] = wdata & d.wmask;
      end
    endcase
  endfunction

  function void predict_commit();
    live = shadow;            // the WHOLE shadow, both halves
  endfunction

  function void check_phase(uvm_phase phase);
    super.check_phase(phase);

    // Stimulus checks, as errors: a run that never produced the race has not
    // tested the design's central property, however many cycles it ran for.
    if (n_rc_race == 0)
      `uvm_error("RF/COV",
        "no event ever arrived during a read-to-clear read: the property this design exists for was never exercised")
    if (n_w1c_live == 0)
      `uvm_error("RF/COV",
        "no W1C write ever collided with a live source: the W1C policy was never exercised")

    `uvm_info("RF",
      $sformatf("%0d RC races | %0d W1C-vs-live | %0d locked refusals | %0d reserved-bit maskings",
                n_rc_race, n_w1c_live, n_locked_refused, n_reserved_masked), UVM_LOW)
  endfunction
endclass

15. Common Misconceptions

"Read-to-clear is convenient." It is a destructive read with exactly one legal owner. A second reader — a diagnostic thread, a second driver, a debugger — destroys the first one's data.

"Clearing on read means clearing the register." It means clearing what was returned. An event that arrived on the read cycle was returned to nobody.

"W1C should clear the bit unconditionally." Not if the source is still asserting. The register would disagree with the hardware, and the driver would conclude the condition went away.

"A shadow register is over-engineering." A 32-bit configuration written through a 16-bit bus is briefly half-old. Without a shadow the device acts on a value nobody asked for.

"Commit can copy the half that changed." Z3 copies only the low half and is invisible until someone writes the high one. Commit copies the whole shadow.

"Reserved bits can be ignored." They read back as whatever was written, software stores and restores the word, and a later silicon revision gives those bits meaning.

"A locked register can drop writes quietly." Then software believes it reconfigured the device. A refused write must be reported.

"An unmapped address can return zeros." Then a typo in a driver is indistinguishable from a register that is genuinely zero.

"Same-cycle read and write is a corner case." It happened 260 times in this run and it is the reason for every design decision in the block.

16. Exercises

1. Write the read-to-clear path as stat_n = 0 on any read and give the exact cycle on which it differs from the published design. Estimate how often that cycle occurs given the 260-in-11986 rate measured here.

2. Z6 scores 1551 and 591. Reproduce the division by enabling events, and propose a measurement that would account for the residual 1.24×.

3. Z1 and Z6 both have a directed score of exactly 16. Explain why, and construct a phase that would raise it without adding a single random cycle.

4. Give the shadow/commit design a partial commit — only the halves that were written since the last commit. Which mutations does this defeat, which new one does it need, and is it worth it?

5. Two drivers share the read-to-clear counter register. Design a non-destructive mirror for the diagnostic path and say what it costs in area and in coherency.

6. The control register masks reserved bits on write, so they read back as zero. Argue the opposite policy — store them verbatim — and say which one survives a silicon revision that gives the bits meaning.

7. Across this module the directed share of the mutation score ranges from 40% (chapter 26.3) to under 0.2% (chapter 26.4) to about 0.05% here. Explain the range from the shape of the three state spaces, and say what a high directed share would tell you about a suite.

17. Summary

IdeaWhy it matters
Read-to-clear clears what was returnedan event arriving on the read cycle was seen by nobody
A read-to-clear register has one readertwo readers destroy each other's data
A read-to-clear register is not writableZ7 lets a driver clear events it never read
W1C respects a live sourceor the register disagrees with the hardware
Configuration lands in a shadowa 32-bit value through a 16-bit bus is briefly half-old
Commit copies the whole shadowZ3 copies half and hides until someone writes the other
Reserved bits are masked, not storedsoftware saves and restores the word across revisions
A locked write is refused and reporteda silent drop makes software believe it succeeded
An unmapped access is an erroror a driver typo looks like a zero register
Expose the raw register, not a masked viewa masked port hid Z4 entirely in one language
Divide a spread by its enabling event count2.6× in score is 1.24× in rate
A directed score of 2 is luckenumerate the domain: 4 bits × 2 states × 2 policies = 16
48 address × access states, 7 mutationsall killed by directed stimulus alone, in 3 languages

Tooling

StepCommand
Verilog-2005iverilog -g2005 -o rf_v.out rf_v.v rf_v_tb.v && ./rf_v.out
SystemVerilogiverilog -g2012 -o rf_sv.out rf_sv.sv rf_sv_tb.sv && ./rf_sv.out
VHDL-2008 analysenvc --std=2008 -a rf_vhdl.vhd rf_vhdl_tb.vhd
VHDL-2008 elaboratenvc --std=2008 -e tb_rf_vhdl
VHDL-2008 runnvc --std=2008 -r tb_rf_vhdl
One mutationiverilog -g2005 -DMUT_Z1 -o mm rf_v_mut.v rf_v_tb.v && ./mm
Directed onlyiverilog -g2005 -DDIRECTED_ONLY -o mm rf_v_mut.v rf_v_tb.v && ./mm

All three implementations pass with 0 errors: 48 address × access states reached, 260 same-cycle read/set races and 429 locked refusals exercised, every access policy checked against an independently written model, and every one of the seven mutations killed by directed stimulus alone.


18. Module 26 in one page

Five blocks, five defects, and one idea underneath all of them.

ChapterThe blockThe defect
26.1endpoint buffer poola shared pool head-of-line blocks; a per-endpoint pool cannot
26.2DMA descriptor engineceil(len / maxp) omits the ZLP and hangs on the sizes everybody uses
26.3AXI burst splittera burst crossing a 4 KB boundary is legal on your board and corrupts on theirs
26.4interrupt aggregatorhardware sets a bit as software clears it, and the clear wins
26.5register fileread-to-clear clears the whole register, including what nobody saw

Every one of the five is a race between two agents who each behave correctly. Nothing here is a typo or a missing case. The DMA engine's arithmetic is right for every length but the multiples. The AXI splitter tiles the address range exactly. The aggregator sets and clears exactly what it was told. The register file returns exactly what was asked for. They fail because two things happen on the same cycle and the design silently picks one.

That is what makes integration bugs different from RTL bugs, and it is why every chapter in this module counts how often the situation arose alongside whether the design survived it. A verification report that says "0 errors" and cannot say "and the race occurred 17920 times" has proved that a testbench ran, not that a design works.

Continue learning

Standards & specifications

Governing standard
USB-IF (Universal Serial Bus Specification)(opens USB Implementers Forum (USB-IF) in a new tab)

Defines the USB bus — its electrical signalling, connectors, packet and transaction model, device framework and the descriptors a device must expose — together with the device-class specifications layered on it. It does not define host-controller register interfaces (xHCI and EHCI are separate documents) nor any operating system's driver architecture.

This page also covers RTL structure, verification approach and debugging technique. Those are engineering practice built on the standard, not requirements the standard itself imposes.

Where this fits

Part of the USB curriculum.