Skip to content
VLSI Mentor

USB · Module 25

Using a Protocol Analyser

A trigger fires on the symptom, and the symptom is the end of the story — so a capture that starts at the trigger records the aftermath and never contains the cause.

This module has spent six chapters building blocks that watch a bus. This one is about the instrument that does it for you, and about the two settings that decide whether its output is evidence or decoration.

1. The Capture Does Not Contain the Bug

You set the analyser to trigger on a STALL. It triggers. You get a beautiful, complete, perfectly decoded capture of everything that happened after the STALL: the host giving up, the driver resetting the endpoint, the retry storm, the re-enumeration.

None of which is the bug.

The bug is whatever the firmware did three hundred transactions earlier that made it decide to stall. That is in the part of the trace you did not record, because you had not triggered yet.

Where the cause lives relative to where the trigger fires

A sequence diagram between a host and a device firmware. The firmware receives a control request it does not recognise and sets an internal error flag, which is the cause. Several hundred ordinary bulk transfers follow, all successful. The firmware's buffer then overruns as a consequence of the earlier flag, and it returns a STALL. The analyser triggers on the STALL, so its capture begins there and contains only the host's recovery attempts.A STALL, and the three hundred transactions that caused itHostDevice firmwarea control requestthe firmware doesnot knowACK — and aninternal flag is set~300 ordinary bulktransfers, all finebuffer overruns, asa consequenceSTALL — the analysertriggers HERECLEAR_FEATURE,reset, re-enumerate
The analyser's default is to begin recording when the trigger condition is met, which puts the whole capture to the right of the line. Everything that explains the symptom is to the left of it.

2. So Record Always, and Let the Trigger Decide What to Keep

The fix is a circular buffer that runs the entire time. The trigger does not start the recording; it stops it, some number of events later.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   arm ----> [ always writing, oldest overwritten ] ----> trigger
                          |<--- kept --->|<-- POST_N -->|

   The trigger chooses a WINDOW, and the window straddles it.

Every analyser worth using supports this, usually under a name like pre-trigger buffer or trigger position. Most of them ship with it set to zero.

3. And the Post-Trigger Window Eats the Pre-Trigger History

This is the part that surprises people, and it is pure arithmetic. The buffer is a fixed depth. Every event captured after the trigger overwrites the oldest pre-trigger entry:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   pre_kept = min(events before the trigger, DEPTH - post_kept)

   DEPTH = 64, POST_N = 8    ->  at most 56 pre-trigger events
   DEPTH = 64, POST_N = 32   ->  at most 32
   DEPTH = 64, POST_N = 64   ->  none at all

So asking for more aftermath costs you exactly that much of the cause, and the cause is the part you needed.

4. The Second Trap: Filtering at Capture Time

"Just capture the errors." It halves the file size and destroys the investigation.

The error is the symptom. The surrounding ordinary traffic is the evidence. A capture filtered down to errors cannot answer what was this endpoint doing immediately before, which is the only question worth asking of it.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   Filter on the way OUT, never on the way IN.

So the view mask in this block is applied at read time. Every event is in the buffer; the mask only decides what is shown, and the hidden ones are counted — so the number on screen is never mistaken for the number captured.

5. What We Are Building

usb_trace_filter — always writing, the trigger chooses the window

A block diagram of the trace filter. An event stream feeds a circular buffer unconditionally while the capture is armed. A trigger input moves a state machine from armed into a post-trigger countdown and then to frozen. The frozen state computes how many pre-trigger and post-trigger entries survived and whether any were lost, and a read port drains them oldest first through a view mask that counts shown and hidden entries separately.Event streamtype · endpoint · codeCircular bufferwritten UNCONDITIONALLYTriggerstops, does not startWindowpre_kept · post_kept ·pre_lostDrainoldest firstView maskhides, never deletes12
The event stream is written unconditionally while armed; no filter of any kind exists on the write path. The trigger moves the machine into a post-trigger countdown, and the freeze fixes the window. The read side computes the oldest retained entry from the write pointer and applies the view mask, counting what it hides.

The capture state machine

A four-state capture machine. IDLE is the start state and records nothing. An arm input moves it to ARMED, where every event is written to the circular buffer. A trigger moves ARMED to CAPT, where writing continues for POST_N more events. After those the machine moves to FROZEN, where events are ignored and the capture can be drained. An arm input returns any state to ARMED.IDLEARMEDCAPTFROZENarmarmevent: write, oldest lostevent: write, oldest lostevent:write,…triggertriggerPOST_N eventsPOST_N eventsevents ignoredeventsignoredarm: capture discardedarm: capture discardedarm:capture…
ARMED is the state the block spends almost all its time in, writing continuously and discarding the oldest entries. The trigger only moves it to CAPT; the POST_N countdown moves it to FROZEN. Re-arming from any state throws the capture away, including the pre-trigger count.

6. Verilog-2005 Implementation

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// usb_trace_filter -- why the capture you took does not contain the bug,
// and the two decisions that fix it.
//
// A TRIGGER FIRES ON THE SYMPTOM, AND THE SYMPTOM IS THE END OF THE STORY
//
// You set the analyser to trigger on a STALL. It triggers. You get a
// beautiful capture of everything that happened AFTER the STALL: the host
// giving up, the driver resetting the endpoint, the retry storm. None of
// which is the bug.
//
// The bug is whatever the firmware did three hundred transactions earlier
// that made it decide to stall. That is in the part of the trace you did
// not record, because you had not triggered yet.
//
//     A post-trigger-only capture records the AFTERMATH.
//     The cause is always before the trigger. Always.
//
// THE FIX IS TO RECORD ALWAYS AND LET THE TRIGGER DECIDE WHAT TO KEEP
//
// A circular buffer that is running the entire time. The trigger does not
// start the recording; it STOPS it, some number of events later. What you
// keep is a window that straddles the trigger.
//
//     arm ----> [ always writing, oldest overwritten ] ----> trigger
//                            |<--- kept --->|<-- POST_N -->|
//
// AND THE POST-TRIGGER WINDOW EATS THE PRE-TRIGGER HISTORY
//
// This is the part that surprises people. The buffer is a fixed DEPTH. Every
// event captured after the trigger overwrites the OLDEST pre-trigger entry.
// So asking for more post-trigger context costs you exactly that much
// pre-trigger context, and the pre-trigger context is where the bug is:
//
//     pre_kept = min(events before the trigger, DEPTH - post_kept)
//
// THE SECOND TRAP: FILTERING AT CAPTURE TIME
//
// "Just capture the errors" halves the file size and destroys the
// investigation. The error is the symptom; the surrounding ordinary traffic
// is the evidence. A capture filtered down to errors cannot answer "what was
// this endpoint doing immediately before", which is the only question.
//
//     Filter on the way OUT, never on the way IN.
//
// So the view mask here is applied at READ time. Every event is in the
// buffer; the mask only decides what is shown, and the hidden ones are
// counted so the number on screen is never mistaken for the number captured.
//
// AND THE CAPTURE HAS TO SAY WHEN IT IS NOT LONG ENOUGH
//
// If more events occurred before the trigger than the buffer could hold, the
// cause may be off the front -- and a full buffer looks exactly like a
// sufficient one. `pre_lost` is the difference between "the cause is not in
// this capture" and "there is no cause", which are not the same finding.
module usb_trace_filter #(
  parameter integer DEPTH  = 64,   // circular buffer entries
  parameter integer POST_N = 8     // events captured after the trigger
) (
  input  wire        clk,
  input  wire        rst_n,

  input  wire        arm,          // begin a fresh capture
  input  wire        ev_valid,
  input  wire [1:0]  ev_type,      // 0 token, 1 data, 2 handshake, 3 error
  input  wire [3:0]  ev_ep,
  input  wire [7:0]  ev_code,
  input  wire        trigger,      // the symptom fired

  input  wire        rd_en,
  input  wire [3:0]  view_mask,    // one bit per event type, applied ON READ

  output wire [1:0]  state,
  output wire [7:0]  n_valid,      // entries retained
  output wire [7:0]  pre_kept,
  output wire [7:0]  post_kept,
  output wire        pre_lost,     // the cause may be off the front

  output wire        rd_valid,
  output wire        rd_show,      // passes the view mask
  output wire [1:0]  rd_type,
  output wire [3:0]  rd_ep,
  output wire [7:0]  rd_code,
  output wire [7:0]  rd_index,     // 0 = oldest retained

  output reg [31:0] n_written,     // every event offered. NEVER filtered.
  output reg [31:0] n_overwritten, // lost to the circular buffer
  output reg [31:0] n_shown,
  output reg [31:0] n_hidden       // in the capture, not on the screen
);

  localparam [1:0] S_IDLE   = 2'd0,
                   S_ARMED  = 2'd1,   // recording, waiting for the symptom
                   S_CAPT   = 2'd2,   // triggered, taking POST_N more
                   S_FROZEN = 2'd3;   // done; drain it

  // The index width. Written as a localparam rather than $clog2 inline so
  // that the two places it is used cannot drift apart.
  localparam integer AW = (DEPTH <= 16)  ? 4 :
                          (DEPTH <= 32)  ? 5 :
                          (DEPTH <= 64)  ? 6 :
                          (DEPTH <= 128) ? 7 : 8;

  reg [1:0]  ty_m   [0:DEPTH-1];
  reg [3:0]  ep_m   [0:DEPTH-1];
  reg [7:0]  code_m [0:DEPTH-1];

  reg [1:0]      st_r;
  reg [AW-1:0]   wr_r;
  reg [15:0]     total_r;     // events written since arm, saturating
  reg [15:0]     pre_r;       // events written BEFORE the trigger
  reg [7:0]      post_r;      // events written after it
  reg [7:0]      rdi_r;       // drain index
  reg            rdv_r;

  // ---- What survived. ----
  //
  // post_kept is however many post-trigger events were actually taken --
  // fewer than POST_N if the capture was frozen early. pre_kept is what the
  // buffer had room for AFTER those, which is the subtraction that surprises
  // people.
  wire [7:0] post_k = post_r;
  // ---- DEPTH is bounded by the port widths, and that is stated here. ----
  //
  // pre_kept, post_kept and n_valid are eight-bit outputs, so a buffer
  // deeper than 255 cannot be reported through them whatever this expression
  // does. Rather than widen one number and leave the others, the bound is
  // made explicit: this block is correct for DEPTH <= 255. Chapter 25.4's
  // retry budget is the same question answered the other way, and the rule
  // is the same -- a width follows from what the value means, and a value
  // that cannot be expressed is a documented limit rather than a surprise.
  wire [7:0] depth8 = DEPTH;
  wire [7:0] room   = (depth8 > post_k) ? (depth8 - post_k) : 8'd0;
  wire [7:0] pre_k  = (pre_r > {8'd0, room}) ? room : pre_r[7:0];

  assign state     = st_r;
  assign pre_kept  = pre_k;
  assign post_kept = post_k;
  assign n_valid   = pre_k + post_k;
  // The cause may be off the front, and a full buffer looks exactly like a
  // sufficient one. This bit is the difference between "not in this capture"
  // and "no such event".
  assign pre_lost  = (pre_r > {8'd0, room});

  // ---- The read window. ----
  //
  // Chronological, oldest first. The oldest retained entry is n_valid slots
  // behind the write pointer, modulo the buffer -- which is the only place
  // the circularity is visible, and the only place it can be got wrong.
  wire [AW-1:0] rd_addr = wr_r - n_valid[AW-1:0] + rdi_r[AW-1:0];

  // ---- The read port is REGISTERED, all of it, together. ----
  //
  // A combinational read off `rdi_r` beside a registered `rd_valid` is off
  // by one: the pointer advances on the same edge that raises valid, so the
  // data presented alongside the valid pulse belongs to the NEXT entry. The
  // symptom is a capture that is correct in every respect except that it
  // starts one entry late and ends one entry early -- which is exactly the
  // kind of defect a length check cannot see, and exactly why this
  // testbench compares the contents rather than the count.
  reg [1:0] rdt_r;
  reg [3:0] rde_r;
  reg [7:0] rdc_r;
  reg [7:0] rdx_r;
  reg       rds_r;

  assign rd_valid = rdv_r;
  assign rd_type  = rdt_r;
  assign rd_ep    = rde_r;
  assign rd_code  = rdc_r;
  assign rd_index = rdx_r;
  // The mask decides what is SHOWN. The entry is in the buffer either way.
  assign rd_show  = rds_r;

  integer i;
  always @(posedge clk or negedge rst_n) begin
    if (!rst_n) begin
      st_r    <= S_IDLE;
      wr_r    <= {AW{1'b0}};
      total_r <= 16'd0;
      pre_r   <= 16'd0;
      post_r  <= 8'd0;
      rdi_r   <= 8'd0;
      rdv_r   <= 1'b0;
      rdt_r   <= 2'd0;
      rde_r   <= 4'd0;
      rdc_r   <= 8'd0;
      rdx_r   <= 8'd0;
      rds_r   <= 1'b0;
      n_written     <= 32'd0;
      n_overwritten <= 32'd0;
      n_shown       <= 32'd0;
      n_hidden      <= 32'd0;
      for (i = 0; i < DEPTH; i = i + 1) begin
        ty_m[i]   <= 2'd0;
        ep_m[i]   <= 4'd0;
        code_m[i] <= 8'd0;
      end
    end else begin
      rdv_r <= 1'b0;

      if (arm) begin
        st_r    <= S_ARMED;
        wr_r    <= {AW{1'b0}};
        total_r <= 16'd0;
        pre_r   <= 16'd0;
        post_r  <= 8'd0;
        rdi_r   <= 8'd0;
      end else begin
        // ---- WRITE FIRST, unconditionally, with no filter of any kind. ----
        //
        // Not "write if it is an error", not "write if it matches the mask".
        // The mask lives on the read side and nowhere else; the moment a
        // filter appears here, the context that explains the error is gone
        // and no amount of later processing brings it back.
        if (ev_valid && ((st_r == S_ARMED) || (st_r == S_CAPT))) begin
          ty_m[wr_r]   <= ev_type;
          ep_m[wr_r]   <= ev_ep;
          code_m[wr_r] <= ev_code;
          wr_r         <= wr_r + {{(AW-1){1'b0}}, 1'b1};
          n_written    <= n_written + 32'd1;
          if (total_r >= DEPTH) n_overwritten <= n_overwritten + 32'd1;
          if (total_r != 16'hFFFF) total_r <= total_r + 16'd1;

          if (st_r == S_ARMED) begin
            if (pre_r != 16'hFFFF) pre_r <= pre_r + 16'd1;
          end else begin
            post_r <= post_r + 8'd1;
            // ---- The post-trigger window is the last thing that happens.
            //
            // POST_N events after the symptom and the capture stops. Every
            // one of them has already overwritten an oldest pre-trigger
            // entry, which is the cost nobody budgets for.
            if (post_r + 8'd1 >= POST_N[7:0]) st_r <= S_FROZEN;
          end
        end

        // The trigger is sampled AFTER the write, so the event that caused
        // the symptom is itself the last pre-trigger entry rather than the
        // first post-trigger one. Off by one here moves the boundary of
        // every capture by one event, which is invisible until somebody
        // counts.
        if (trigger && (st_r == S_ARMED)) st_r <= S_CAPT;

        if (rd_en && (st_r == S_FROZEN)) begin
          if (rdi_r < n_valid) begin
            rdv_r <= 1'b1;
            rdi_r <= rdi_r + 8'd1;
            rdt_r <= ty_m[rd_addr];
            rde_r <= ep_m[rd_addr];
            rdc_r <= code_m[rd_addr];
            rdx_r <= rdi_r;
            rds_r <= view_mask[ty_m[rd_addr]];
            if (view_mask[ty_m[rd_addr]]) n_shown  <= n_shown  + 32'd1;
            else                          n_hidden <= n_hidden + 32'd1;
          end
        end
      end
    end
  end
endmodule

7. SystemVerilog Implementation

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// usb_trace_filter -- why the capture you took does not contain the bug,
// and the two decisions that fix it.
//
// A TRIGGER FIRES ON THE SYMPTOM, AND THE SYMPTOM IS THE END OF THE STORY
//
// You set the analyser to trigger on a STALL. It triggers. You get a
// beautiful capture of everything that happened AFTER the STALL: the host
// giving up, the driver resetting the endpoint, the retry storm. None of
// which is the bug.
//
// The bug is whatever the firmware did three hundred transactions earlier
// that made it decide to stall. That is in the part of the trace you did
// not record, because you had not triggered yet.
//
//     A post-trigger-only capture records the AFTERMATH.
//     The cause is always before the trigger. Always.
//
// THE FIX IS TO RECORD ALWAYS AND LET THE TRIGGER DECIDE WHAT TO KEEP
//
// A circular buffer that is running the entire time. The trigger does not
// start the recording; it STOPS it, some number of events later. What you
// keep is a window that straddles the trigger.
//
//     arm ----> [ always writing, oldest overwritten ] ----> trigger
//                            |<--- kept --->|<-- POST_N -->|
//
// AND THE POST-TRIGGER WINDOW EATS THE PRE-TRIGGER HISTORY
//
// This is the part that surprises people. The buffer is a fixed DEPTH. Every
// event captured after the trigger overwrites the OLDEST pre-trigger entry.
// So asking for more post-trigger context costs you exactly that much
// pre-trigger context, and the pre-trigger context is where the bug is:
//
//     pre_kept = min(events before the trigger, DEPTH - post_kept)
//
// THE SECOND TRAP: FILTERING AT CAPTURE TIME
//
// "Just capture the errors" halves the file size and destroys the
// investigation. The error is the symptom; the surrounding ordinary traffic
// is the evidence. A capture filtered down to errors cannot answer "what was
// this endpoint doing immediately before", which is the only question.
//
//     Filter on the way OUT, never on the way IN.
//
// So the view mask here is applied at READ time. Every event is in the
// buffer; the mask only decides what is shown, and the hidden ones are
// counted so the number on screen is never mistaken for the number captured.
//
// AND THE CAPTURE HAS TO SAY WHEN IT IS NOT LONG ENOUGH
//
// If more events occurred before the trigger than the buffer could hold, the
// cause may be off the front -- and a full buffer looks exactly like a
// sufficient one. `pre_lost` is the difference between "the cause is not in
// this capture" and "there is no cause", which are not the same finding.
package usb_trace_pkg;
  typedef enum logic [1:0] {
    S_IDLE   = 2'd0,
    S_ARMED  = 2'd1,   // recording, waiting for the symptom
    S_CAPT   = 2'd2,   // triggered, taking POST_N more
    S_FROZEN = 2'd3    // done; drain it
  } cap_state_e;

  typedef enum logic [1:0] {
    EV_TOKEN = 2'd0,
    EV_DATA  = 2'd1,
    EV_HSHK  = 2'd2,
    EV_ERROR = 2'd3
  } ev_type_e;
endpackage

module usb_trace_filter
  import usb_trace_pkg::*;
 #(
  parameter int DEPTH  = 64,       // circular buffer entries
  parameter int POST_N = 8         // events captured after the trigger
) (
  input  logic       clk,
  input  logic       rst_n,

  input  logic       arm,          // begin a fresh capture
  input  logic       ev_valid,
  input  ev_type_e   ev_type,
  input  logic [3:0] ev_ep,
  input  logic [7:0] ev_code,
  input  logic       trigger,      // the symptom fired

  input  logic       rd_en,
  input  logic [3:0] view_mask,    // one bit per event type, applied ON READ

  output cap_state_e state,
  output logic [7:0] n_valid,      // entries retained
  output logic [7:0] pre_kept,
  output logic [7:0] post_kept,
  output logic       pre_lost,     // the cause may be off the front

  output logic       rd_valid,
  output logic       rd_show,      // passes the view mask
  output ev_type_e   rd_type,
  output logic [3:0] rd_ep,
  output logic [7:0] rd_code,
  output logic [7:0] rd_index,     // 0 = oldest retained

  output logic [31:0] n_written,     // every event offered. NEVER filtered.
  output logic [31:0] n_overwritten, // lost to the circular buffer
  output logic [31:0] n_shown,
  output logic [31:0] n_hidden       // in the capture, not on the screen
);

  // The index width. Written as a localparam rather than $clog2 inline so
  // that the two places it is used cannot drift apart.
  localparam int AW = (DEPTH <= 16)  ? 4 :
                          (DEPTH <= 32)  ? 5 :
                          (DEPTH <= 64)  ? 6 :
                          (DEPTH <= 128) ? 7 : 8;

  ev_type_e    ty_m   [DEPTH-1:0];
  logic [3:0]  ep_m   [DEPTH-1:0];
  logic [7:0]  code_m [DEPTH-1:0];

  cap_state_e    st_r;
  logic [AW-1:0] wr_r;
  logic [15:0]   total_r;   // events written since arm, saturating
  logic [15:0]   pre_r;     // events written BEFORE the trigger
  logic [7:0]    post_r;    // events written after it
  logic [7:0]    rdi_r;     // drain index
  logic          rdv_r;

  // ---- What survived. ----
  //
  // post_kept is however many post-trigger events were actually taken --
  // fewer than POST_N if the capture was frozen early. pre_kept is what the
  // buffer had room for AFTER those, which is the subtraction that surprises
  // people.
  logic [7:0] post_k, room, pre_k;
  assign post_k = post_r;
  // ---- DEPTH is bounded by the port widths, and that is stated here. ----
  //
  // pre_kept, post_kept and n_valid are eight-bit outputs, so a buffer
  // deeper than 255 cannot be reported through them whatever this expression
  // does. Rather than widen one number and leave the others, the bound is
  // made explicit: this block is correct for DEPTH <= 255. Chapter 25.4's
  // retry budget is the same question answered the other way, and the rule
  // is the same -- a width follows from what the value means, and a value
  // that cannot be expressed is a documented limit rather than a surprise.
  localparam logic [7:0] DEPTH8 = 8'(DEPTH);
  assign room   = (DEPTH8 > post_k) ? (DEPTH8 - post_k) : 8'd0;
  assign pre_k  = (pre_r > {8'd0, room}) ? room : pre_r[7:0];

  assign state     = st_r;
  assign pre_kept  = pre_k;
  assign post_kept = post_k;
  assign n_valid   = pre_k + post_k;
  // The cause may be off the front, and a full buffer looks exactly like a
  // sufficient one. This bit is the difference between "not in this capture"
  // and "no such event".
  assign pre_lost  = (pre_r > {8'd0, room});

  // ---- The read window. ----
  //
  // Chronological, oldest first. The oldest retained entry is n_valid slots
  // behind the write pointer, modulo the buffer -- which is the only place
  // the circularity is visible, and the only place it can be got wrong.
  logic [AW-1:0] rd_addr; assign rd_addr = wr_r - n_valid[AW-1:0] + rdi_r[AW-1:0];

  // ---- The read port is REGISTERED, all of it, together. ----
  //
  // A combinational read off `rdi_r` beside a registered `rd_valid` is off
  // by one: the pointer advances on the same edge that raises valid, so the
  // data presented alongside the valid pulse belongs to the NEXT entry. The
  // symptom is a capture that is correct in every respect except that it
  // starts one entry late and ends one entry early -- which is exactly the
  // kind of defect a length check cannot see, and exactly why this
  // testbench compares the contents rather than the count.
  ev_type_e   rdt_r;
  logic [3:0] rde_r;
  logic [7:0] rdc_r;
  logic [7:0] rdx_r;
  logic       rds_r;

  assign rd_valid = rdv_r;
  assign rd_type  = rdt_r;
  assign rd_ep    = rde_r;
  assign rd_code  = rdc_r;
  assign rd_index = rdx_r;
  // The mask decides what is SHOWN. The entry is in the buffer either way.
  assign rd_show  = rds_r;

  int i;
  always_ff @(posedge clk or negedge rst_n) begin
    if (!rst_n) begin
      st_r    <= S_IDLE;
      wr_r    <= {AW{1'b0}};
      total_r <= 16'd0;
      pre_r   <= 16'd0;
      post_r  <= 8'd0;
      rdi_r   <= 8'd0;
      rdv_r   <= 1'b0;
      rdt_r   <= EV_TOKEN;
      rde_r   <= 4'd0;
      rdc_r   <= 8'd0;
      rdx_r   <= 8'd0;
      rds_r   <= 1'b0;
      n_written     <= 32'd0;
      n_overwritten <= 32'd0;
      n_shown       <= 32'd0;
      n_hidden      <= 32'd0;
      for (i = 0; i < DEPTH; i = i + 1) begin
        ty_m[i]   <= EV_TOKEN;
        ep_m[i]   <= 4'd0;
        code_m[i] <= 8'd0;
      end
    end else begin
      rdv_r <= 1'b0;

      if (arm) begin
        st_r    <= S_ARMED;
        wr_r    <= {AW{1'b0}};
        total_r <= 16'd0;
        pre_r   <= 16'd0;
        post_r  <= 8'd0;
        rdi_r   <= 8'd0;
      end else begin
        // ---- WRITE FIRST, unconditionally, with no filter of any kind. ----
        //
        // Not "write if it is an error", not "write if it matches the mask".
        // The mask lives on the read side and nowhere else; the moment a
        // filter appears here, the context that explains the error is gone
        // and no amount of later processing brings it back.
        if (ev_valid && ((st_r == S_ARMED) || (st_r == S_CAPT))) begin
          ty_m[wr_r]   <= ev_type;
          ep_m[wr_r]   <= ev_ep;
          code_m[wr_r] <= ev_code;
          wr_r         <= wr_r + {{(AW-1){1'b0}}, 1'b1};
          n_written    <= n_written + 32'd1;
          if (total_r >= 16'(DEPTH)) n_overwritten <= n_overwritten + 32'd1;
          if (total_r != 16'hFFFF) total_r <= total_r + 16'd1;

          if (st_r == S_ARMED) begin
            if (pre_r != 16'hFFFF) pre_r <= pre_r + 16'd1;
          end else begin
            post_r <= post_r + 8'd1;
            // ---- The post-trigger window is the last thing that happens.
            //
            // POST_N events after the symptom and the capture stops. Every
            // one of them has already overwritten an oldest pre-trigger
            // entry, which is the cost nobody budgets for.
            if ((post_r + 8'd1) >= 8'(POST_N)) st_r <= S_FROZEN;
          end
        end

        // The trigger is sampled AFTER the write, so the event that caused
        // the symptom is itself the last pre-trigger entry rather than the
        // first post-trigger one. Off by one here moves the boundary of
        // every capture by one event, which is invisible until somebody
        // counts.
        if (trigger && (st_r == S_ARMED)) st_r <= S_CAPT;

        if (rd_en && (st_r == S_FROZEN)) begin
          if (rdi_r < n_valid) begin
            rdv_r <= 1'b1;
            rdi_r <= rdi_r + 8'd1;
            rdt_r <= ty_m[rd_addr];
            rde_r <= ep_m[rd_addr];
            rdc_r <= code_m[rd_addr];
            rdx_r <= rdi_r;
            rds_r <= view_mask[ty_m[rd_addr]];
            if (view_mask[ty_m[rd_addr]]) n_shown  <= n_shown  + 32'd1;
            else                          n_hidden <= n_hidden + 32'd1;
          end
        end
      end
    end
  end
endmodule

8. VHDL-2008 Implementation

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
-- usb_trace_filter -- why the capture you took does not contain the bug, and
-- the two decisions that fix it.
--
-- A TRIGGER FIRES ON THE SYMPTOM, AND THE SYMPTOM IS THE END OF THE STORY
--
-- You set the analyser to trigger on a STALL. It triggers. You get a
-- beautiful capture of everything that happened AFTER the STALL: the host
-- giving up, the driver resetting the endpoint, the retry storm. None of
-- which is the bug.
--
-- The bug is whatever the firmware did three hundred transactions earlier
-- that made it decide to stall. That is in the part of the trace you did not
-- record, because you had not triggered yet.
--
--     A post-trigger-only capture records the AFTERMATH.
--     The cause is always before the trigger. Always.
--
-- THE FIX IS TO RECORD ALWAYS AND LET THE TRIGGER DECIDE WHAT TO KEEP
--
-- A circular buffer running the entire time. The trigger does not start the
-- recording; it STOPS it, some number of events later.
--
--     arm ----> [ always writing, oldest overwritten ] ----> trigger
--                            |<--- kept --->|<-- POST_N -->|
--
-- AND THE POST-TRIGGER WINDOW EATS THE PRE-TRIGGER HISTORY
--
-- The buffer is a fixed DEPTH. Every event captured after the trigger
-- overwrites the OLDEST pre-trigger entry, so asking for more post-trigger
-- context costs exactly that much pre-trigger context -- and the pre-trigger
-- context is where the bug is:
--
--     pre_kept = min(events before the trigger, DEPTH - post_kept)
--
-- THE SECOND TRAP: FILTERING AT CAPTURE TIME
--
-- "Just capture the errors" halves the file size and destroys the
-- investigation. The error is the symptom; the surrounding ordinary traffic
-- is the evidence.
--
--     Filter on the way OUT, never on the way IN.
--
-- So the view mask is applied at READ time. Every event is in the buffer;
-- the mask only decides what is shown, and the hidden ones are counted so
-- the number on screen is never mistaken for the number captured.
--
-- AND THE CAPTURE HAS TO SAY WHEN IT IS NOT LONG ENOUGH
--
-- If more events occurred before the trigger than the buffer could hold, the
-- cause may be off the front -- and a full buffer looks exactly like a
-- sufficient one. pre_lost is the difference between "the cause is not in
-- this capture" and "there is no cause", which are not the same finding.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

package usb_trace_pkg is
  constant S_IDLE   : std_logic_vector(1 downto 0) := "00";
  constant S_ARMED  : std_logic_vector(1 downto 0) := "01";
  constant S_CAPT   : std_logic_vector(1 downto 0) := "10";
  constant S_FROZEN : std_logic_vector(1 downto 0) := "11";
end package;

library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
use work.usb_trace_pkg.all;

entity usb_trace_filter is
  generic (
    DEPTH  : integer := 64;    -- circular buffer entries
    POST_N : integer := 8      -- events captured after the trigger
  );
  port (
    clk       : in  std_logic;
    rst_n     : in  std_logic;

    arm       : in  std_logic;                     -- begin a fresh capture
    ev_valid  : in  std_logic;
    ev_type   : in  std_logic_vector(1 downto 0);  -- token/data/hshk/error
    ev_ep     : in  std_logic_vector(3 downto 0);
    ev_code   : in  std_logic_vector(7 downto 0);
    trigger   : in  std_logic;                     -- the symptom fired

    rd_en     : in  std_logic;
    view_mask : in  std_logic_vector(3 downto 0);  -- applied ON READ

    state     : out std_logic_vector(1 downto 0);
    n_valid   : out unsigned(7 downto 0);          -- entries retained
    pre_kept  : out unsigned(7 downto 0);
    post_kept : out unsigned(7 downto 0);
    pre_lost  : out std_logic;                     -- cause may be off front

    rd_valid  : out std_logic;
    rd_show   : out std_logic;                     -- passes the view mask
    rd_type   : out std_logic_vector(1 downto 0);
    rd_ep     : out std_logic_vector(3 downto 0);
    rd_code   : out std_logic_vector(7 downto 0);
    rd_index  : out unsigned(7 downto 0);          -- 0 = oldest retained

    n_written     : out unsigned(31 downto 0);     -- every event. NEVER filtered.
    n_overwritten : out unsigned(31 downto 0);     -- lost to the circular buffer
    n_shown       : out unsigned(31 downto 0);
    n_hidden      : out unsigned(31 downto 0)      -- captured, not on screen
  );
end entity;

architecture rtl of usb_trace_filter is
  type ty_array   is array (0 to DEPTH-1) of std_logic_vector(1 downto 0);
  type ep_array   is array (0 to DEPTH-1) of std_logic_vector(3 downto 0);
  type code_array is array (0 to DEPTH-1) of std_logic_vector(7 downto 0);

  signal ty_m   : ty_array   := (others => (others => '0'));
  signal ep_m   : ep_array   := (others => (others => '0'));
  signal code_m : code_array := (others => (others => '0'));

  signal st_r    : std_logic_vector(1 downto 0) := S_IDLE;
  signal wr_r    : integer range 0 to DEPTH-1 := 0;
  signal total_r : unsigned(15 downto 0) := (others => '0');
  signal pre_r   : unsigned(15 downto 0) := (others => '0');
  signal post_r  : unsigned(7 downto 0)  := (others => '0');
  signal rdi_r   : unsigned(7 downto 0)  := (others => '0');
  signal rdv_r   : std_logic := '0';

  -- The read port is REGISTERED, all of it, together. A combinational read
  -- beside a registered rd_valid is off by one: the pointer advances on the
  -- same edge that raises valid, so the data presented alongside the valid
  -- pulse belongs to the NEXT entry. The symptom is a capture correct in
  -- every respect except that it starts one entry late and ends one early --
  -- exactly the defect a length check cannot see.
  signal rdt_r : std_logic_vector(1 downto 0) := (others => '0');
  signal rde_r : std_logic_vector(3 downto 0) := (others => '0');
  signal rdc_r : std_logic_vector(7 downto 0) := (others => '0');
  signal rdx_r : unsigned(7 downto 0) := (others => '0');
  signal rds_r : std_logic := '0';

  signal wrote_c, over_c, shown_c, hid_c : unsigned(31 downto 0)
    := (others => '0');

  -- Initialised at declaration. rd_addr is a concurrent expression over
  -- nval, so it evaluates at time 0 before any driver has resolved; left
  -- uninitialised these are all-U and numeric_std's to_integer emits a
  -- metavalue warning on every run. A warning that always fires is a
  -- warning nobody reads.
  signal post_k, room, pre_k, nval : unsigned(7 downto 0) := (others => '0');
  signal rd_addr : integer range 0 to DEPTH-1 := 0;
begin
  -- ---- What survived. ----
  --
  -- post_kept is however many post-trigger events were actually taken.
  -- pre_kept is what the buffer had room for AFTER those, which is the
  -- subtraction that surprises people.
  post_k <= post_r;
  room   <= to_unsigned(DEPTH, 8) - post_k when to_unsigned(DEPTH, 8) > post_k
            else (others => '0');
  pre_k  <= room when pre_r > resize(room, 16) else pre_r(7 downto 0);
  nval   <= pre_k + post_k;

  state     <= st_r;
  pre_kept  <= pre_k;
  post_kept <= post_k;
  n_valid   <= nval;
  -- The cause may be off the front, and a full buffer looks exactly like a
  -- sufficient one. This bit is the difference between "not in this capture"
  -- and "no such event".
  pre_lost  <= '1' when pre_r > resize(room, 16) else '0';

  -- ---- The read window. ----
  --
  -- Chronological, oldest first. The oldest retained entry is n_valid slots
  -- behind the write pointer, modulo the buffer -- the only place the
  -- circularity is visible, and the only place it can be got wrong.
  rd_addr <= (wr_r - to_integer(nval) + to_integer(rdi_r) + 2*DEPTH) mod DEPTH;

  rd_valid <= rdv_r;
  rd_type  <= rdt_r;
  rd_ep    <= rde_r;
  rd_code  <= rdc_r;
  rd_index <= rdx_r;
  rd_show  <= rds_r;

  n_written     <= wrote_c;
  n_overwritten <= over_c;
  n_shown       <= shown_c;
  n_hidden      <= hid_c;

  process (clk, rst_n)
  begin
    if rst_n = '0' then
      st_r    <= S_IDLE;
      wr_r    <= 0;
      total_r <= (others => '0');
      pre_r   <= (others => '0');
      post_r  <= (others => '0');
      rdi_r   <= (others => '0');
      rdv_r   <= '0';
      rdt_r   <= (others => '0');
      rde_r   <= (others => '0');
      rdc_r   <= (others => '0');
      rdx_r   <= (others => '0');
      rds_r   <= '0';
      wrote_c <= (others => '0');
      over_c  <= (others => '0');
      shown_c <= (others => '0');
      hid_c   <= (others => '0');
      ty_m    <= (others => (others => '0'));
      ep_m    <= (others => (others => '0'));
      code_m  <= (others => (others => '0'));
    elsif rising_edge(clk) then
      rdv_r <= '0';

      if arm = '1' then
        st_r    <= S_ARMED;
        wr_r    <= 0;
        total_r <= (others => '0');
        pre_r   <= (others => '0');
        post_r  <= (others => '0');
        rdi_r   <= (others => '0');
      else
        -- ---- WRITE FIRST, unconditionally, with no filter of any kind. ----
        --
        -- Not "write if it is an error", not "write if it matches the mask".
        -- The mask lives on the read side and nowhere else; the moment a
        -- filter appears here, the context that explains the error is gone
        -- and no amount of later processing brings it back.
        if ev_valid = '1' and (st_r = S_ARMED or st_r = S_CAPT) then
          ty_m(wr_r)   <= ev_type;
          ep_m(wr_r)   <= ev_ep;
          code_m(wr_r) <= ev_code;
          if wr_r = DEPTH-1 then wr_r <= 0; else wr_r <= wr_r + 1; end if;
          wrote_c <= wrote_c + 1;
          if total_r >= to_unsigned(DEPTH, 16) then over_c <= over_c + 1; end if;
          if total_r /= x"FFFF" then total_r <= total_r + 1; end if;

          if st_r = S_ARMED then
            if pre_r /= x"FFFF" then pre_r <= pre_r + 1; end if;
          else
            post_r <= post_r + 1;
            -- ---- The post-trigger window is the last thing that happens.
            --
            -- POST_N events after the symptom and the capture stops. Every
            -- one of them has already overwritten an oldest pre-trigger
            -- entry, which is the cost nobody budgets for.
            if (post_r + 1) >= to_unsigned(POST_N, 8) then
              st_r <= S_FROZEN;
            end if;
          end if;
        end if;

        -- The trigger is sampled AFTER the write, so the event that caused
        -- the symptom is itself the last pre-trigger entry rather than the
        -- first post-trigger one. Off by one here moves the boundary of
        -- every capture by one event, which is invisible until somebody
        -- counts.
        if trigger = '1' and st_r = S_ARMED then
          st_r <= S_CAPT;
        end if;

        if rd_en = '1' and st_r = S_FROZEN then
          if rdi_r < nval then
            rdv_r <= '1';
            rdi_r <= rdi_r + 1;
            rdt_r <= ty_m(rd_addr);
            rde_r <= ep_m(rd_addr);
            rdc_r <= code_m(rd_addr);
            rdx_r <= rdi_r;
            rds_r <= view_mask(to_integer(unsigned(ty_m(rd_addr))));
            if view_mask(to_integer(unsigned(ty_m(rd_addr)))) = '1' then
              shown_c <= shown_c + 1;
            else
              hid_c <= hid_c + 1;
            end if;
          end if;
        end if;
      end if;
    end if;
  end process;
end architecture;

9. Seeing the Window Close

Two events, the trigger, then the post-trigger countdown

usb_trace_filter — the trigger stops the recording, it does not start it

10 cycles
A ten-cycle waveform with a buffer depth of eight and a post-trigger count of two, for legibility. Three events are written while the state is ARMED and the pre-trigger count rises to three. A trigger pulse moves the state to CAPT without changing the counts. Two further events are written, the post-trigger count reaches two, and the state becomes FROZEN. The retained pre-trigger count is three and the retained post-trigger count is two.recording already, before any triggerrecording already, beforeany triggertrigger: stops, does not starttrigger: stops, does notstartfrozen: 3 before, 2 afterfrozen: 3 before, 2 afterclkev_validtriggerstateARMEDARMEDARMEDARMEDCAPTCAPTFROZENFROZENFROZENFROZENpre_kept0123333333post_kept0000012222n_valid0123345555pre_lostt0t1t2t3t4t5t6t7t8t9
The block is already recording when the window opens. The trigger moves it to CAPT without discarding anything; the POST_N countdown runs on events rather than cycles, and the freeze happens on the last of them. Note that pre_kept stops at DEPTH minus the post-trigger events — the aftermath has already overwritten that much of the cause.

And the case the block exists to report:

More history than the buffer can hold

usb_trace_filter — pre_lost is the honest part of the report

10 cycles
A ten-cycle waveform with a buffer depth of eight and a post-trigger count of two. Many events have already been written so the pre-trigger total is twenty, well past what the buffer can hold. The retained pre-trigger count saturates at six, the pre-lost flag is high, and after the trigger and two more events the state freezes with six pre-trigger and two post-trigger entries retained.20 events happened, 8 fit20 events happened, 8 fiteach aftermath event costs one causeeach aftermath event costsone causepre_lost: the cause may be off the frontpre_lost: the cause may beoff the frontclkev_validtriggerstateARMEDARMEDARMEDARMEDCAPTCAPTFROZENFROZENFROZENFROZENpre total18192020202020202020pre_kept8888876666post_kept0000012222pre_lostt0t1t2t3t4t5t6t7t8t9
The same capture with far more pre-trigger events than the buffer has room for. The retained pre-trigger count saturates at DEPTH minus the post-trigger events, and pre_lost asserts — which is the difference between a capture that does not contain the cause and one that proves there was none. Without that bit the two are indistinguishable.

10. The Testbenches

The claim this chapter makes is about which events survive, and a count cannot check that. The failure being described is a capture containing the wrong events — the aftermath instead of the cause — and a capture of the right length full of the wrong entries has the right count.

So the testbench keeps its own list of everything that was offered, computes the expected window from the chapter's rules, and drains the whole capture and compares it entry by entry, in order, on every one of its 400-odd captures.

The exhaustive part is the retention edge:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   The arithmetic changes behaviour at exactly one place:

       DEPTH - POST_N

   below it, everything before the trigger survives
   above it, the oldest entries are gone and pre_lost must say so

   Phase 1 sweeps the pre-trigger count 0 .. DEPTH + 8 --
   every value on both sides and well past -- and drains and
   compares the entire capture on each of the 73 runs.

   result: 57 intact, 16 lost

Six phases:

PhaseWhat it establishes
0traffic before the capture is armed is not recorded
1the retention edge, swept at every pre-trigger count from 0 to DEPTH + 8
2an event on the trigger cycle belongs to the pre side
3the same capture under all 16 view masks has identical contents
4draining before the freeze yields nothing
5re-arming discards the capture, and traffic after the freeze does not touch it
6300 random captures, every one drained and compared in full

Verilog-2005 testbench

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
`timescale 1ns/1ps
// Testbench for usb_trace_filter.
//
// THE CLAIM IS ABOUT WHAT SURVIVES, SO THE WHOLE CAPTURE IS COMPARED
//
// It is not enough to check that the block retained the right NUMBER of
// events. The failure this chapter is about is a capture that contains the
// wrong events -- the aftermath instead of the cause -- and a count cannot
// tell those apart. So every retained entry is drained and compared, in
// order, against an independently kept list of everything that was offered.
//
// THE EXHAUSTIVE CLAIM
//
// The retention arithmetic depends on how many events preceded the trigger,
// and it changes behaviour at exactly one place: DEPTH - POST_N, where the
// buffer stops being able to hold all of them. So phase 1 sweeps the
// pre-trigger event count from 0 to DEPTH + 8 -- every value on both sides
// of that edge and well past it -- and for each one drains the entire
// capture and compares it entry by entry.
module tb_tf_v;

  localparam integer DEPTH  = 64;
  localparam integer POST_N = 8;
  localparam integer MAXEV  = 400;

  localparam [1:0] S_IDLE = 2'd0, S_ARMED = 2'd1,
                   S_CAPT = 2'd2, S_FROZEN = 2'd3;

  reg        clk = 1'b0, rst_n = 1'b0;
  reg        arm = 1'b0, ev_valid = 1'b0, trigger = 1'b0, rd_en = 1'b0;
  reg [1:0]  ev_type = 2'd0;
  reg [3:0]  ev_ep = 4'd0;
  reg [7:0]  ev_code = 8'd0;
  reg [3:0]  view_mask = 4'hF;

  wire [1:0] state, rd_type;
  wire [7:0] n_valid, pre_kept, post_kept, rd_code, rd_index;
  wire [3:0] rd_ep;
  wire       pre_lost, rd_valid, rd_show;
  wire [31:0] n_written, n_overwritten, n_shown, n_hidden;

  usb_trace_filter #(.DEPTH(DEPTH), .POST_N(POST_N)) dut (
    .clk(clk), .rst_n(rst_n),
    .arm(arm), .ev_valid(ev_valid), .ev_type(ev_type), .ev_ep(ev_ep),
    .ev_code(ev_code), .trigger(trigger),
    .rd_en(rd_en), .view_mask(view_mask),
    .state(state), .n_valid(n_valid), .pre_kept(pre_kept),
    .post_kept(post_kept), .pre_lost(pre_lost),
    .rd_valid(rd_valid), .rd_show(rd_show), .rd_type(rd_type),
    .rd_ep(rd_ep), .rd_code(rd_code), .rd_index(rd_index),
    .n_written(n_written), .n_overwritten(n_overwritten),
    .n_shown(n_shown), .n_hidden(n_hidden)
  );

  always #5 clk = ~clk;

  // ---------------- the reference list ----------------
  //
  // Everything that was offered while the block was recording, in order.
  // The expected capture is a WINDOW into this list, and computing that
  // window from the chapter's rules is the whole oracle.
  reg [1:0] a_ty   [0:MAXEV-1];
  reg [3:0] a_ep   [0:MAXEV-1];
  reg [7:0] a_code [0:MAXEV-1];
  integer   n_all, n_pre;

  // ---------------- the shadow model of the recorder ----------------
  //
  // The drain comparison proves the CONTENTS. This proves the machine that
  // produced them: what state it is in, how much it wrote, and how much it
  // threw away -- every cycle, including the cycles where nothing happens,
  // because "nothing happens" is a claim too.
  reg [1:0]  m_state;
  integer    m_written, m_overwritten, m_total, m_post;

  integer errors = 0, checks = 0, steps = 0;
  integer k;

  reg [0:0] reach [0:15];    // state (4) x event type (4)
  integer   n_reach;

  task ck;
    input [255:0] nm;
    input [31:0]  got, exp;
    begin
      checks = checks + 1;
      if (got !== exp) begin
        errors = errors + 1;
        if (errors < 25)
          $display("FAIL t=%0t step=%0d %0s got=%0d exp=%0d",
                   $time, steps, nm, got, exp);
      end
    end
  endtask

  task model_cycle;
    begin
      if (arm) begin
        m_state = S_ARMED;
        m_total = 0;
        m_post  = 0;
      end else begin
        if (ev_valid && ((m_state == S_ARMED) || (m_state == S_CAPT))) begin
          m_written = m_written + 1;
          if (m_total >= DEPTH) m_overwritten = m_overwritten + 1;
          m_total = m_total + 1;
          if (m_state == S_CAPT) begin
            m_post = m_post + 1;
            if (m_post >= POST_N) m_state = S_FROZEN;
          end
        end
        // Sampled AFTER the write, so an event on the trigger cycle belongs
        // to the PRE side.
        if (trigger && (m_state == S_ARMED)) m_state = S_CAPT;
      end
    end
  endtask

  task tickc;
    begin
      model_cycle;
      @(posedge clk); #1;
      steps = steps + 1;
      ck("state",         {30'd0, state}, {30'd0, m_state});
      ck("n_written",     n_written,      m_written[31:0]);
      ck("n_overwritten", n_overwritten,  m_overwritten[31:0]);
    end
  endtask

  task do_arm;
    begin
      arm = 1'b1; ev_valid = 1'b0; trigger = 1'b0; rd_en = 1'b0;
      tickc;
      arm = 1'b0;
      n_all = 0; n_pre = 0;
    end
  endtask

  task ev; input [1:0] t; input [3:0] e; input [7:0] c;
    begin
      arm = 1'b0; ev_valid = 1'b1; trigger = 1'b0; rd_en = 1'b0;
      ev_type = t; ev_ep = e; ev_code = c;
      reach[{state, t}] = 1'b1;
      if ((state == S_ARMED) || (state == S_CAPT)) begin
        a_ty[n_all] = t; a_ep[n_all] = e; a_code[n_all] = c;
        n_all = n_all + 1;
        if (state == S_ARMED) n_pre = n_pre + 1;
      end
      tickc;
      ev_valid = 1'b0;
    end
  endtask

  task trig;
    begin
      arm = 1'b0; ev_valid = 1'b0; trigger = 1'b1; rd_en = 1'b0;
      tickc;
      trigger = 1'b0;
    end
  endtask

  task nothing;
    begin
      arm = 1'b0; ev_valid = 1'b0; trigger = 1'b0; rd_en = 1'b0;
      tickc;
    end
  endtask

  // ---- The expected retained window, from the chapter's rules. ----
  //
  //   post_kept = however many post-trigger events were taken
  //   pre_kept  = min(pre_total, DEPTH - post_kept)
  //
  // The subtraction is the point: every post-trigger event overwrote an
  // oldest pre-trigger one, so asking for more aftermath costs exactly that
  // much of the cause.
  function integer exp_post; input integer dummy;
    begin
      exp_post = n_all - n_pre;
      if (exp_post > POST_N) exp_post = POST_N;
    end
  endfunction

  function integer exp_pre; input integer dummy;
    integer room;
    begin
      room = DEPTH - exp_post(0);
      exp_pre = (n_pre > room) ? room : n_pre;
    end
  endfunction

  // Drain the whole capture and compare every entry, in order, against the
  // reference list. Returns the number of entries the view mask showed.
  task drain_and_check;
    integer want, first, idx, shown, hidden;
    begin
      want  = exp_pre(0) + exp_post(0);
      first = n_all - want;
      ck("n_valid",   {24'd0, n_valid},   want[31:0]);
      ck("pre_kept",  {24'd0, pre_kept},  exp_pre(0));
      ck("post_kept", {24'd0, post_kept}, exp_post(0));
      // The cause may be off the front, and a full buffer looks exactly like
      // a sufficient one.
      ck("pre_lost",  {31'd0, pre_lost},
         (n_pre > (DEPTH - exp_post(0))) ? 32'd1 : 32'd0);

      shown = 0; hidden = 0;
      for (idx = 0; idx < want; idx = idx + 1) begin
        rd_en = 1'b1;
        model_cycle;
        @(posedge clk); #1;
        steps = steps + 1;
        ck("state",         {30'd0, state}, {30'd0, m_state});
        ck("n_written",     n_written,      m_written[31:0]);
        ck("n_overwritten", n_overwritten,  m_overwritten[31:0]);
        ck("rd_valid", {31'd0, rd_valid}, 32'd1);
        ck("rd_index", {24'd0, rd_index}, idx[31:0]);
        // ---- The summary outputs are re-checked on every drain cycle. ----
        //
        // They are stable while frozen, so checking them once per capture is
        // "enough" in the sense that it catches a wrong value. It is not
        // enough in the sense that matters: pre_lost is the difference
        // between "the cause is not in this capture" and "there is no
        // cause", and a mutation that hard-wires it low was dying on ONE
        // check per capture. Re-checking it per entry costs nothing and
        // turns a cornered mutation into a killed one.
        ck("pre_lost stable",  {31'd0, pre_lost},
           (n_pre > (DEPTH - exp_post(0))) ? 32'd1 : 32'd0);
        ck("n_valid stable",   {24'd0, n_valid},   want[31:0]);
        ck("pre_kept stable",  {24'd0, pre_kept},  exp_pre(0));
        ck("post_kept stable", {24'd0, post_kept}, exp_post(0));
        // THE check. Not the count -- the contents, in order.
        ck("rd_type",  {30'd0, rd_type},  {30'd0, a_ty[first + idx]});
        ck("rd_ep",    {28'd0, rd_ep},    {28'd0, a_ep[first + idx]});
        ck("rd_code",  {24'd0, rd_code},  {24'd0, a_code[first + idx]});
        ck("rd_show",  {31'd0, rd_show},
           {31'd0, view_mask[a_ty[first + idx]]});
        if (view_mask[a_ty[first + idx]]) shown = shown + 1;
        else                              hidden = hidden + 1;
      end
      rd_en = 1'b0;
      // Reading past the end yields nothing at all rather than wrapping
      // round to the oldest entry again.
      rd_en = 1'b1;
      model_cycle;
      @(posedge clk); #1;
      steps = steps + 1;
      ck("rd_valid past end", {31'd0, rd_valid}, 32'd0);
      rd_en = 1'b0;
    end
  endtask

  integer p, i, m, w, base_shown, base_hidden;
  integer n_lost_cases, n_intact_cases;
  reg [1:0] save_ty [0:DEPTH-1];

  initial begin
    for (k = 0; k < 16; k = k + 1) reach[k] = 1'b0;
    n_all = 0; n_pre = 0;
    n_lost_cases = 0; n_intact_cases = 0;
    m_state = S_IDLE; m_written = 0; m_overwritten = 0;
    m_total = 0; m_post = 0;

    repeat (3) @(posedge clk);
    rst_n = 1'b1;
    @(negedge clk);

    // ================= PHASE 0 -- traffic before the capture is armed =====
    //
    // Every event type offered while the block is IDLE. None of it may be
    // recorded: a capture that quietly began before it was armed contains
    // events from whatever the bus was doing last time, in the same buffer,
    // with no marker between them.
    for (i = 0; i < 4; i = i + 1) ev(i[1:0], 4'd9, 8'h70 + i[7:0]);
    ck("idle state",       {30'd0, state}, {30'd0, S_IDLE});
    ck("nothing written",  n_written,      32'd0);

    // ================= PHASE 1 -- THE RETENTION EDGE, SWEPT ==============
    //
    // Every pre-trigger event count from 0 to DEPTH + 8. Below
    // DEPTH - POST_N everything before the trigger survives; above it the
    // oldest entries are gone and pre_lost must say so. The whole capture is
    // drained and compared entry by entry on every one of the 73 runs.
    for (p = 0; p <= DEPTH + 8; p = p + 1) begin
      do_arm;
      for (i = 0; i < p; i = i + 1)
        ev(i[1:0], i[3:0], i[7:0] + 8'd1);
      trig;
      for (i = 0; i < POST_N; i = i + 1)
        ev(2'd3, 4'hF, 8'hA0 + i[7:0]);
      ck("frozen", {30'd0, state}, {30'd0, S_FROZEN});
      drain_and_check;
      if (p > DEPTH - POST_N) n_lost_cases   = n_lost_cases + 1;
      else                    n_intact_cases = n_intact_cases + 1;
    end
    if (n_lost_cases != 16 || n_intact_cases != 57) begin
      errors = errors + 1;
      $display("FAIL edge split: lost=%0d intact=%0d expected 16 and 57",
               n_lost_cases, n_intact_cases);
    end

    // ================= PHASE 2 -- the trigger is sampled AFTER the write ==
    //
    // An event and the trigger on the same cycle: the event belongs to the
    // PRE side, because it is part of what led up to the symptom rather than
    // part of the aftermath. Off by one here moves the boundary of every
    // capture by one event, which is invisible until somebody counts.
    do_arm;
    for (i = 0; i < 10; i = i + 1) ev(2'd0, 4'd1, i[7:0] + 8'd1);
    arm = 1'b0; ev_valid = 1'b1; trigger = 1'b1; rd_en = 1'b0;
    ev_type = 2'd1; ev_ep = 4'd2; ev_code = 8'h5A;
    a_ty[n_all] = 2'd1; a_ep[n_all] = 4'd2; a_code[n_all] = 8'h5A;
    n_all = n_all + 1; n_pre = n_pre + 1;
    tickc;
    ev_valid = 1'b0; trigger = 1'b0;
    ck("state after simultaneous trigger", {30'd0, state}, {30'd0, S_CAPT});
    for (i = 0; i < POST_N; i = i + 1) ev(2'd3, 4'hF, 8'hB0 + i[7:0]);
    ck("pre_kept with simultaneous trigger", {24'd0, pre_kept}, 32'd11);
    drain_and_check;

    // ================= PHASE 3 -- the mask hides, it does not delete =====
    //
    // The same capture drained under all 16 view masks. The CONTENTS are
    // identical every time; only rd_show changes. A design that filtered on
    // the way in would return a different capture for each mask, and the
    // context that explains the error would be gone for fifteen of them.
    do_arm;
    for (i = 0; i < 40; i = i + 1) ev(i[1:0], i[3:0], i[7:0] + 8'd1);
    trig;
    for (i = 0; i < POST_N; i = i + 1) ev(2'd2, 4'd7, 8'hC0 + i[7:0]);
    for (m = 0; m < 16; m = m + 1) begin
      // Re-drain from the start. rd_index is only reset by a fresh arm, so
      // the capture is re-created identically instead -- which also proves
      // that an identical stimulus produces an identical capture.
      do_arm;
      for (i = 0; i < 40; i = i + 1) ev(i[1:0], i[3:0], i[7:0] + 8'd1);
      trig;
      for (i = 0; i < POST_N; i = i + 1) ev(2'd2, 4'd7, 8'hC0 + i[7:0]);
      view_mask = m[3:0];
      base_shown = n_shown; base_hidden = n_hidden;
      drain_and_check;
      // Everything drained is either shown or hidden, and nothing is lost.
      ck("shown + hidden",
         (n_shown - base_shown) + (n_hidden - base_hidden),
         {24'd0, n_valid});
    end
    view_mask = 4'hF;

    // ================= PHASE 4 -- draining before the freeze yields none ==
    do_arm;
    for (i = 0; i < 20; i = i + 1) ev(2'd0, 4'd3, i[7:0] + 8'd1);
    trig;
    ev(2'd3, 4'd3, 8'hD0);
    ck("still capturing", {30'd0, state}, {30'd0, S_CAPT});
    rd_en = 1'b1;
    model_cycle;
    @(posedge clk); #1; steps = steps + 1;
    ck("no drain while capturing", {31'd0, rd_valid}, 32'd0);
    rd_en = 1'b0;

    // ================= PHASE 5 -- re-arming throws the capture away ======
    //
    // And it must throw away the pre-trigger count with it. A stale pre_r
    // makes the NEXT capture report a pre_lost it did not earn.
    do_arm;
    for (i = 0; i < 5; i = i + 1) ev(2'd1, 4'd4, i[7:0] + 8'd1);
    trig;
    for (i = 0; i < POST_N; i = i + 1) ev(2'd3, 4'hE, 8'hE0 + i[7:0]);
    ck("re-armed pre_kept", {24'd0, pre_kept}, 32'd5);
    ck("re-armed pre_lost", {31'd0, pre_lost}, 32'd0);

    // ---- Traffic AFTER the freeze must not touch the capture. ----
    //
    // The bus does not stop because you triggered. If a frozen capture keeps
    // absorbing events, the trace changes while you are reading it -- and
    // the entries you already looked at are no longer the entries that are
    // there. Every type is offered, and the drain below is compared against
    // the model as though none of them happened, because none of them may.
    base_shown = n_written;
    for (i = 0; i < 4; i = i + 1) ev(i[1:0], 4'hA, 8'hF0 + i[7:0]);
    ck("frozen: nothing written", n_written, base_shown[31:0]);
    ck("frozen: still frozen",    {30'd0, state}, {30'd0, S_FROZEN});
    drain_and_check;

    // ================= PHASE 6 -- random ================================
    //
    // Random capture lengths and random event mixes, every one of them
    // drained and compared in full.
    for (i = 0; i < 300; i = i + 1) begin
      do_arm;
      w = $unsigned($random) % 90;
      for (p = 0; p < w; p = p + 1)
        ev($unsigned($random) % 4, $unsigned($random) % 16,
           $unsigned($random) % 256);
      trig;
      for (p = 0; p < POST_N; p = p + 1)
        ev($unsigned($random) % 4, $unsigned($random) % 16,
           $unsigned($random) % 256);
      view_mask = $unsigned($random) % 16;
      drain_and_check;
    end
    view_mask = 4'hF;

    // ================= the exhaustiveness proof ==========================
    n_reach = 0;
    for (k = 0; k < 16; k = k + 1) n_reach = n_reach + reach[k];
    if (n_reach != 16) begin
      errors = errors + 1;
      $display("FAIL state x type reach %0d/16", n_reach);
      for (k = 0; k < 16; k = k + 1)
        if (!reach[k]) $display("  unreached state=%0d type=%0d", k >> 2, k & 3);
    end

    $display("steps=%0d checks=%0d reach=%0d/16 errors=%0d",
             steps, checks, n_reach, errors);
    $display("retention edge: intact=%0d lost=%0d",
             n_intact_cases, n_lost_cases);
    $display("written=%0d overwritten=%0d shown=%0d hidden=%0d",
             n_written, n_overwritten, n_shown, n_hidden);
    $display("%0s: %0d errors in %0d checks",
             (errors == 0) ? "PASS" : "FAIL", errors, checks);
    $finish;
  end
endmodule

SystemVerilog testbench

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
`timescale 1ns/1ps
// Testbench for usb_trace_filter.
//
// THE CLAIM IS ABOUT WHAT SURVIVES, SO THE WHOLE CAPTURE IS COMPARED
//
// It is not enough to check that the block retained the right NUMBER of
// events. The failure this chapter is about is a capture that contains the
// wrong events -- the aftermath instead of the cause -- and a count cannot
// tell those apart. So every retained entry is drained and compared, in
// order, against an independently kept list of everything that was offered.
//
// THE EXHAUSTIVE CLAIM
//
// The retention arithmetic depends on how many events preceded the trigger,
// and it changes behaviour at exactly one place: DEPTH - POST_N, where the
// buffer stops being able to hold all of them. So phase 1 sweeps the
// pre-trigger event count from 0 to DEPTH + 8 -- every value on both sides
// of that edge and well past it -- and for each one drains the entire
// capture and compares it entry by entry.
module tb_tf_sv;
  import usb_trace_pkg::*;

  localparam int DEPTH  = 64;
  localparam int POST_N = 8;
  localparam int MAXEV  = 400;

  logic       clk = 1'b0, rst_n = 1'b0;
  logic       arm = 1'b0, ev_valid = 1'b0, trigger = 1'b0, rd_en = 1'b0;
  ev_type_e   ev_type = EV_TOKEN;
  logic [3:0] ev_ep = 4'd0;
  logic [7:0] ev_code = 8'd0;
  logic [3:0] view_mask = 4'hF;

  cap_state_e state;
  ev_type_e   rd_type;
  logic [7:0] n_valid, pre_kept, post_kept, rd_code, rd_index;
  logic [3:0] rd_ep;
  logic       pre_lost, rd_valid, rd_show;
  logic [31:0] n_written, n_overwritten, n_shown, n_hidden;

  usb_trace_filter #(.DEPTH(DEPTH), .POST_N(POST_N)) dut (
    .clk(clk), .rst_n(rst_n),
    .arm(arm), .ev_valid(ev_valid), .ev_type(ev_type), .ev_ep(ev_ep),
    .ev_code(ev_code), .trigger(trigger),
    .rd_en(rd_en), .view_mask(view_mask),
    .state(state), .n_valid(n_valid), .pre_kept(pre_kept),
    .post_kept(post_kept), .pre_lost(pre_lost),
    .rd_valid(rd_valid), .rd_show(rd_show), .rd_type(rd_type),
    .rd_ep(rd_ep), .rd_code(rd_code), .rd_index(rd_index),
    .n_written(n_written), .n_overwritten(n_overwritten),
    .n_shown(n_shown), .n_hidden(n_hidden)
  );

  always #5 clk = ~clk;

  // ---------------- the reference list ----------------
  //
  // Everything that was offered while the block was recording, in order.
  // The expected capture is a WINDOW into this list, and computing that
  // window from the chapter's rules is the whole oracle.
  ev_type_e   a_ty   [MAXEV-1:0];
  logic [3:0] a_ep   [MAXEV-1:0];
  logic [7:0] a_code [MAXEV-1:0];
  int         n_all, n_pre;

  // ---------------- the shadow model of the recorder ----------------
  //
  // The drain comparison proves the CONTENTS. This proves the machine that
  // produced them: what state it is in, how much it wrote, and how much it
  // threw away -- every cycle, including the cycles where nothing happens,
  // because "nothing happens" is a claim too.
  cap_state_e m_state;
  int         m_written, m_overwritten, m_total, m_post;

  int errors = 0, checks = 0, steps = 0;
  int k;

  bit reach [16];            // state (4) x event type (4)
  int n_reach;

  task automatic ck(string nm, int unsigned got, int unsigned exp);
    checks++;
    if (got !== exp) begin
      errors++;
      if (errors < 25)
        $display("FAIL t=%0t step=%0d %0s got=%0d exp=%0d",
                 $time, steps, nm, got, exp);
    end
  endtask

  task automatic model_cycle;
    begin
      if (arm) begin
        m_state = S_ARMED;
        m_total = 0;
        m_post  = 0;
      end else begin
        if (ev_valid && ((m_state == S_ARMED) || (m_state == S_CAPT))) begin
          m_written = m_written + 1;
          if (m_total >= DEPTH) m_overwritten = m_overwritten + 1;
          m_total = m_total + 1;
          if (m_state == S_CAPT) begin
            m_post = m_post + 1;
            if (m_post >= POST_N) m_state = S_FROZEN;
          end
        end
        // Sampled AFTER the write, so an event on the trigger cycle belongs
        // to the PRE side.
        if (trigger && (m_state == S_ARMED)) m_state = S_CAPT;
      end
    end
  endtask

  task automatic tickc;
    begin
      model_cycle;
      @(posedge clk); #1;
      steps = steps + 1;
      ck("state",         state, m_state);
      ck("n_written",     n_written,      m_written[31:0]);
      ck("n_overwritten", n_overwritten,  m_overwritten[31:0]);
    end
  endtask

  task automatic do_arm;
    begin
      arm = 1'b1; ev_valid = 1'b0; trigger = 1'b0; rd_en = 1'b0;
      tickc;
      arm = 1'b0;
      n_all = 0; n_pre = 0;
    end
  endtask

  task automatic ev(ev_type_e t, logic [3:0] e, logic [7:0] c);
    begin
      arm = 1'b0; ev_valid = 1'b1; trigger = 1'b0; rd_en = 1'b0;
      ev_type = t; ev_ep = e; ev_code = c;
      reach[int'(state) * 4 + int'(t)] = 1'b1;
      if ((state == S_ARMED) || (state == S_CAPT)) begin
        a_ty[n_all] = t; a_ep[n_all] = e; a_code[n_all] = c;
        n_all = n_all + 1;
        if (state == S_ARMED) n_pre = n_pre + 1;
      end
      tickc;
      ev_valid = 1'b0;
    end
  endtask

  task automatic trig;
    begin
      arm = 1'b0; ev_valid = 1'b0; trigger = 1'b1; rd_en = 1'b0;
      tickc;
      trigger = 1'b0;
    end
  endtask

  task automatic nothing;
    begin
      arm = 1'b0; ev_valid = 1'b0; trigger = 1'b0; rd_en = 1'b0;
      tickc;
    end
  endtask

  // ---- The expected retained window, from the chapter's rules. ----
  //
  //   post_kept = however many post-trigger events were taken
  //   pre_kept  = min(pre_total, DEPTH - post_kept)
  //
  // The subtraction is the point: every post-trigger event overwrote an
  // oldest pre-trigger one, so asking for more aftermath costs exactly that
  // much of the cause.
  function automatic int exp_post(input int dummy);
    begin
      exp_post = n_all - n_pre;
      if (exp_post > POST_N) exp_post = POST_N;
    end
  endfunction

  function automatic int exp_pre(input int dummy);
    int room;
    begin
      room = DEPTH - exp_post(0);
      exp_pre = (n_pre > room) ? room : n_pre;
    end
  endfunction

  // Drain the whole capture and compare every entry, in order, against the
  // reference list. Returns the number of entries the view mask showed.
  task automatic drain_and_check;
    int want, first, idx, shown, hidden;
    begin
      want  = exp_pre(0) + exp_post(0);
      first = n_all - want;
      ck("n_valid",   {24'd0, n_valid},   want[31:0]);
      ck("pre_kept",  {24'd0, pre_kept},  exp_pre(0));
      ck("post_kept", {24'd0, post_kept}, exp_post(0));
      // The cause may be off the front, and a full buffer looks exactly like
      // a sufficient one.
      ck("pre_lost",  {31'd0, pre_lost},
         (n_pre > (DEPTH - exp_post(0))) ? 32'd1 : 32'd0);

      shown = 0; hidden = 0;
      for (idx = 0; idx < want; idx = idx + 1) begin
        rd_en = 1'b1;
        model_cycle;
        @(posedge clk); #1;
        steps = steps + 1;
        ck("state",         state, m_state);
        ck("n_written",     n_written,      m_written[31:0]);
        ck("n_overwritten", n_overwritten,  m_overwritten[31:0]);
        ck("rd_valid", {31'd0, rd_valid}, 32'd1);
        ck("rd_index", {24'd0, rd_index}, idx[31:0]);
        // ---- The summary outputs are re-checked on every drain cycle. ----
        //
        // They are stable while frozen, so checking them once per capture is
        // "enough" in the sense that it catches a wrong value. It is not
        // enough in the sense that matters: pre_lost is the difference
        // between "the cause is not in this capture" and "there is no
        // cause", and a mutation that hard-wires it low was dying on ONE
        // check per capture. Re-checking it per entry costs nothing and
        // turns a cornered mutation into a killed one.
        ck("pre_lost stable",  {31'd0, pre_lost},
           (n_pre > (DEPTH - exp_post(0))) ? 32'd1 : 32'd0);
        ck("n_valid stable",   {24'd0, n_valid},   want[31:0]);
        ck("pre_kept stable",  {24'd0, pre_kept},  exp_pre(0));
        ck("post_kept stable", {24'd0, post_kept}, exp_post(0));
        // THE check. Not the count -- the contents, in order.
        ck("rd_type",  rd_type, a_ty[first + idx]);
        ck("rd_ep",    {28'd0, rd_ep},    {28'd0, a_ep[first + idx]});
        ck("rd_code",  {24'd0, rd_code},  {24'd0, a_code[first + idx]});
        ck("rd_show",  {31'd0, rd_show},
           {31'd0, view_mask[int'(a_ty[first + idx])]});
        if (view_mask[int'(a_ty[first + idx])]) shown = shown + 1;
        else                              hidden = hidden + 1;
      end
      rd_en = 1'b0;
      // Reading past the end yields nothing at all rather than wrapping
      // round to the oldest entry again.
      rd_en = 1'b1;
      model_cycle;
      @(posedge clk); #1;
      steps = steps + 1;
      ck("rd_valid past end", {31'd0, rd_valid}, 32'd0);
      rd_en = 1'b0;
    end
  endtask

  int p, i, m, w, base_shown, base_hidden;
  int n_lost_cases, n_intact_cases;

  initial begin
    foreach (reach[q]) reach[q] = 1'b0;
    n_all = 0; n_pre = 0;
    n_lost_cases = 0; n_intact_cases = 0;
    m_state = S_IDLE; m_written = 0; m_overwritten = 0;
    m_total = 0; m_post = 0;

    repeat (3) @(posedge clk);
    rst_n = 1'b1;
    @(negedge clk);

    // ================= PHASE 0 -- traffic before the capture is armed =====
    //
    // Every event type offered while the block is IDLE. None of it may be
    // recorded: a capture that quietly began before it was armed contains
    // events from whatever the bus was doing last time, in the same buffer,
    // with no marker between them.
    for (i = 0; i < 4; i = i + 1) ev(ev_type_e'(i[1:0]), 4'd9, 8'h70 + i[7:0]);
    ck("idle state",       state, S_IDLE);
    ck("nothing written",  n_written,      32'd0);

    // ================= PHASE 1 -- THE RETENTION EDGE, SWEPT ==============
    //
    // Every pre-trigger event count from 0 to DEPTH + 8. Below
    // DEPTH - POST_N everything before the trigger survives; above it the
    // oldest entries are gone and pre_lost must say so. The whole capture is
    // drained and compared entry by entry on every one of the 73 runs.
    for (p = 0; p <= DEPTH + 8; p = p + 1) begin
      do_arm;
      for (i = 0; i < p; i = i + 1)
        ev(ev_type_e'(i[1:0]), i[3:0], i[7:0] + 8'd1);
      trig;
      for (i = 0; i < POST_N; i = i + 1)
        ev(EV_ERROR, 4'hF, 8'hA0 + i[7:0]);
      ck("frozen", state, S_FROZEN);
      drain_and_check;
      if (p > DEPTH - POST_N) n_lost_cases   = n_lost_cases + 1;
      else                    n_intact_cases = n_intact_cases + 1;
    end
    if (n_lost_cases != 16 || n_intact_cases != 57) begin
      errors = errors + 1;
      $display("FAIL edge split: lost=%0d intact=%0d expected 16 and 57",
               n_lost_cases, n_intact_cases);
    end

    // ================= PHASE 2 -- the trigger is sampled AFTER the write ==
    //
    // An event and the trigger on the same cycle: the event belongs to the
    // PRE side, because it is part of what led up to the symptom rather than
    // part of the aftermath. Off by one here moves the boundary of every
    // capture by one event, which is invisible until somebody counts.
    do_arm;
    for (i = 0; i < 10; i = i + 1) ev(EV_TOKEN, 4'd1, i[7:0] + 8'd1);
    arm = 1'b0; ev_valid = 1'b1; trigger = 1'b1; rd_en = 1'b0;
    ev_type = EV_DATA; ev_ep = 4'd2; ev_code = 8'h5A;
    a_ty[n_all] = EV_DATA; a_ep[n_all] = 4'd2; a_code[n_all] = 8'h5A;
    n_all = n_all + 1; n_pre = n_pre + 1;
    tickc;
    ev_valid = 1'b0; trigger = 1'b0;
    ck("state after simultaneous trigger", state, S_CAPT);
    for (i = 0; i < POST_N; i = i + 1) ev(EV_ERROR, 4'hF, 8'hB0 + i[7:0]);
    ck("pre_kept with simultaneous trigger", {24'd0, pre_kept}, 32'd11);
    drain_and_check;

    // ================= PHASE 3 -- the mask hides, it does not delete =====
    //
    // The same capture drained under all 16 view masks. The CONTENTS are
    // identical every time; only rd_show changes. A design that filtered on
    // the way in would return a different capture for each mask, and the
    // context that explains the error would be gone for fifteen of them.
    do_arm;
    for (i = 0; i < 40; i = i + 1) ev(ev_type_e'(i[1:0]), i[3:0], i[7:0] + 8'd1);
    trig;
    for (i = 0; i < POST_N; i = i + 1) ev(EV_HSHK, 4'd7, 8'hC0 + i[7:0]);
    for (m = 0; m < 16; m = m + 1) begin
      // Re-drain from the start. rd_index is only reset by a fresh arm, so
      // the capture is re-created identically instead -- which also proves
      // that an identical stimulus produces an identical capture.
      do_arm;
      for (i = 0; i < 40; i = i + 1) ev(ev_type_e'(i[1:0]), i[3:0], i[7:0] + 8'd1);
      trig;
      for (i = 0; i < POST_N; i = i + 1) ev(EV_HSHK, 4'd7, 8'hC0 + i[7:0]);
      view_mask = m[3:0];
      base_shown = n_shown; base_hidden = n_hidden;
      drain_and_check;
      // Everything drained is either shown or hidden, and nothing is lost.
      ck("shown + hidden",
         (n_shown - base_shown) + (n_hidden - base_hidden),
         {24'd0, n_valid});
    end
    view_mask = 4'hF;

    // ================= PHASE 4 -- draining before the freeze yields none ==
    do_arm;
    for (i = 0; i < 20; i = i + 1) ev(EV_TOKEN, 4'd3, i[7:0] + 8'd1);
    trig;
    ev(EV_ERROR, 4'd3, 8'hD0);
    ck("still capturing", state, S_CAPT);
    rd_en = 1'b1;
    model_cycle;
    @(posedge clk); #1; steps = steps + 1;
    ck("no drain while capturing", {31'd0, rd_valid}, 32'd0);
    rd_en = 1'b0;

    // ================= PHASE 5 -- re-arming throws the capture away ======
    //
    // And it must throw away the pre-trigger count with it. A stale pre_r
    // makes the NEXT capture report a pre_lost it did not earn.
    do_arm;
    for (i = 0; i < 5; i = i + 1) ev(EV_DATA, 4'd4, i[7:0] + 8'd1);
    trig;
    for (i = 0; i < POST_N; i = i + 1) ev(EV_ERROR, 4'hE, 8'hE0 + i[7:0]);
    ck("re-armed pre_kept", {24'd0, pre_kept}, 32'd5);
    ck("re-armed pre_lost", {31'd0, pre_lost}, 32'd0);

    // ---- Traffic AFTER the freeze must not touch the capture. ----
    //
    // The bus does not stop because you triggered. If a frozen capture keeps
    // absorbing events, the trace changes while you are reading it -- and
    // the entries you already looked at are no longer the entries that are
    // there. Every type is offered, and the drain below is compared against
    // the model as though none of them happened, because none of them may.
    base_shown = n_written;
    for (i = 0; i < 4; i = i + 1) ev(ev_type_e'(i[1:0]), 4'hA, 8'hF0 + i[7:0]);
    ck("frozen: nothing written", n_written, base_shown[31:0]);
    ck("frozen: still frozen",    state, S_FROZEN);
    drain_and_check;

    // ================= PHASE 6 -- random ================================
    //
    // Random capture lengths and random event mixes, every one of them
    // drained and compared in full.
    for (i = 0; i < 300; i = i + 1) begin
      do_arm;
      w = $unsigned($random) % 90;
      for (p = 0; p < w; p = p + 1)
        ev(ev_type_e'($unsigned($random) % 4),
           4'($unsigned($random) % 16), 8'($unsigned($random) % 256));
      trig;
      for (p = 0; p < POST_N; p = p + 1)
        ev(ev_type_e'($unsigned($random) % 4),
           4'($unsigned($random) % 16), 8'($unsigned($random) % 256));
      view_mask = 4'($unsigned($random) % 16);
      drain_and_check;
    end
    view_mask = 4'hF;

    // ================= the exhaustiveness proof ==========================
    n_reach = 0;
    for (k = 0; k < 16; k = k + 1) n_reach = n_reach + reach[k];
    if (n_reach != 16) begin
      errors = errors + 1;
      $display("FAIL state x type reach %0d/16", n_reach);
      for (k = 0; k < 16; k = k + 1)
        if (!reach[k]) $display("  unreached state=%0d type=%0d", k / 4, k % 4);
    end

    $display("steps=%0d checks=%0d reach=%0d/16 errors=%0d",
             steps, checks, n_reach, errors);
    $display("retention edge: intact=%0d lost=%0d",
             n_intact_cases, n_lost_cases);
    $display("written=%0d overwritten=%0d shown=%0d hidden=%0d",
             n_written, n_overwritten, n_shown, n_hidden);
    $display("%0s: %0d errors in %0d checks",
             (errors == 0) ? "PASS" : "FAIL", errors, checks);
    $finish;
  end
endmodule

VHDL-2008 testbench

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
-- Testbench for usb_trace_filter (VHDL-2008).
--
-- THE CLAIM IS ABOUT WHAT SURVIVES, SO THE WHOLE CAPTURE IS COMPARED
--
-- It is not enough to check that the block retained the right NUMBER of
-- events. The failure this chapter is about is a capture that contains the
-- wrong events -- the aftermath instead of the cause -- and a count cannot
-- tell those apart. So every retained entry is drained and compared, in
-- order, against an independently kept list of everything that was offered.
--
-- THE EXHAUSTIVE CLAIM
--
-- The retention arithmetic changes behaviour at exactly one place:
-- DEPTH - POST_N, where the buffer stops being able to hold all the
-- pre-trigger events. Phase 1 sweeps the pre-trigger count from 0 to
-- DEPTH + 8 -- every value on both sides of that edge and well past it --
-- and for each one drains the entire capture and compares it entry by entry.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
use std.textio.all;
use work.usb_trace_pkg.all;

entity tb_tf_vhdl is
end entity;

architecture sim of tb_tf_vhdl is
  constant DEPTH  : integer := 64;
  constant POST_N : integer := 8;
  constant MAXEV  : integer := 400;

  signal clk       : std_logic := '0';
  signal rst_n     : std_logic := '0';
  signal arm       : std_logic := '0';
  signal ev_valid  : std_logic := '0';
  signal trigger   : std_logic := '0';
  signal rd_en     : std_logic := '0';
  signal ev_type   : std_logic_vector(1 downto 0) := "00";
  signal ev_ep     : std_logic_vector(3 downto 0) := "0000";
  signal ev_code   : std_logic_vector(7 downto 0) := (others => '0');
  signal view_mask : std_logic_vector(3 downto 0) := "1111";

  signal state_s     : std_logic_vector(1 downto 0);
  signal n_valid_s   : unsigned(7 downto 0);
  signal pre_kept_s  : unsigned(7 downto 0);
  signal post_kept_s : unsigned(7 downto 0);
  signal pre_lost_s  : std_logic;
  signal rd_valid_s  : std_logic;
  signal rd_show_s   : std_logic;
  signal rd_type_s   : std_logic_vector(1 downto 0);
  signal rd_ep_s     : std_logic_vector(3 downto 0);
  signal rd_code_s   : std_logic_vector(7 downto 0);
  signal rd_index_s  : unsigned(7 downto 0);
  signal n_written_s, n_over_s, n_shown_s, n_hidden_s : unsigned(31 downto 0);

  signal done : boolean := false;

  type ty_arr   is array (0 to MAXEV-1) of std_logic_vector(1 downto 0);
  type ep_arr   is array (0 to MAXEV-1) of std_logic_vector(3 downto 0);
  type code_arr is array (0 to MAXEV-1) of std_logic_vector(7 downto 0);
  type int_arr  is array (natural range <>) of integer;
begin
  clk <= not clk after 5 ns when not done else '0';

  dut : entity work.usb_trace_filter
    generic map (DEPTH => DEPTH, POST_N => POST_N)
    port map (
      clk => clk, rst_n => rst_n,
      arm => arm, ev_valid => ev_valid, ev_type => ev_type, ev_ep => ev_ep,
      ev_code => ev_code, trigger => trigger,
      rd_en => rd_en, view_mask => view_mask,
      state => state_s, n_valid => n_valid_s, pre_kept => pre_kept_s,
      post_kept => post_kept_s, pre_lost => pre_lost_s,
      rd_valid => rd_valid_s, rd_show => rd_show_s, rd_type => rd_type_s,
      rd_ep => rd_ep_s, rd_code => rd_code_s, rd_index => rd_index_s,
      n_written => n_written_s, n_overwritten => n_over_s,
      n_shown => n_shown_s, n_hidden => n_hidden_s
    );

  stim : process
    -- ---------------- the reference list ----------------
    --
    -- Everything offered while the block was recording, in order. The
    -- expected capture is a WINDOW into this list, and computing that window
    -- from the chapter's rules is the whole oracle.
    variable a_ty   : ty_arr;
    variable a_ep   : ep_arr;
    variable a_code : code_arr;
    variable n_all, n_pre : integer := 0;

    -- ---------------- the shadow model of the recorder ----------------
    variable m_state : std_logic_vector(1 downto 0) := S_IDLE;
    variable m_written, m_over, m_total, m_post : integer := 0;

    variable errors, checks, steps : integer := 0;
    variable reach   : int_arr(0 to 15) := (others => 0);
    variable n_reach : integer := 0;
    variable n_lost_cases, n_intact_cases : integer := 0;
    variable base_written : integer := 0;
    variable w_v : integer := 0;
    variable ln  : line;

    -- A deterministic LFSR, so a rerun reproduces exactly the same traffic.
    variable lfsr : unsigned(31 downto 0) := x"C0FFEE11";

    impure function rnd_nat return integer is
    begin
      lfsr := lfsr(30 downto 0) &
              (lfsr(31) xor lfsr(21) xor lfsr(1) xor lfsr(0));
      -- Only the low 30 bits: a full 32-bit unsigned does not fit in VHDL's
      -- INTEGER, and to_integer aborts the run rather than wrapping.
      return to_integer(lfsr(29 downto 0));
    end function;

    procedure ck (nm : string; got, exp : integer) is
    begin
      checks := checks + 1;
      if got /= exp then
        errors := errors + 1;
        if errors < 25 then
          write(ln, string'("FAIL step=") & integer'image(steps) & " " & nm
                    & " got=" & integer'image(got)
                    & " exp=" & integer'image(exp));
          writeline(output, ln);
        end if;
      end if;
    end procedure;

    function sl2i (s : std_logic) return integer is
    begin
      if s = '1' then return 1; else return 0; end if;
    end function;

    procedure model_cycle is
    begin
      if arm = '1' then
        m_state := S_ARMED;
        m_total := 0;
        m_post  := 0;
      else
        if ev_valid = '1' and (m_state = S_ARMED or m_state = S_CAPT) then
          m_written := m_written + 1;
          if m_total >= DEPTH then m_over := m_over + 1; end if;
          m_total := m_total + 1;
          if m_state = S_CAPT then
            m_post := m_post + 1;
            if m_post >= POST_N then m_state := S_FROZEN; end if;
          end if;
        end if;
        -- Sampled AFTER the write, so an event on the trigger cycle belongs
        -- to the PRE side.
        if trigger = '1' and m_state = S_ARMED then m_state := S_CAPT; end if;
      end if;
    end procedure;

    procedure tickc is
    begin
      model_cycle;
      wait until rising_edge(clk);
      wait for 1 ns;
      steps := steps + 1;
      ck("state",         to_integer(unsigned(state_s)),
                          to_integer(unsigned(m_state)));
      ck("n_written",     to_integer(n_written_s), m_written);
      ck("n_overwritten", to_integer(n_over_s),    m_over);
    end procedure;

    procedure do_arm is
    begin
      arm <= '1'; ev_valid <= '0'; trigger <= '0'; rd_en <= '0';
      wait for 0 ns;
      tickc;
      arm <= '0';
      n_all := 0; n_pre := 0;
    end procedure;

    procedure ev (t : std_logic_vector(1 downto 0);
                  e : std_logic_vector(3 downto 0);
                  c : std_logic_vector(7 downto 0)) is
    begin
      arm <= '0'; ev_valid <= '1'; trigger <= '0'; rd_en <= '0';
      ev_type <= t; ev_ep <= e; ev_code <= c;
      wait for 0 ns;
      reach(to_integer(unsigned(state_s)) * 4 + to_integer(unsigned(t))) := 1;
      if state_s = S_ARMED or state_s = S_CAPT then
        a_ty(n_all) := t; a_ep(n_all) := e; a_code(n_all) := c;
        n_all := n_all + 1;
        if state_s = S_ARMED then n_pre := n_pre + 1; end if;
      end if;
      tickc;
      ev_valid <= '0';
    end procedure;

    procedure trig is
    begin
      arm <= '0'; ev_valid <= '0'; trigger <= '1'; rd_en <= '0';
      wait for 0 ns;
      tickc;
      trigger <= '0';
    end procedure;

    -- ---- The expected retained window, from the chapter's rules. ----
    --
    --   post_kept = however many post-trigger events were taken
    --   pre_kept  = min(pre_total, DEPTH - post_kept)
    --
    -- The subtraction is the point: every post-trigger event overwrote an
    -- oldest pre-trigger one, so asking for more aftermath costs exactly
    -- that much of the cause.
    impure function exp_post return integer is
      variable p : integer;
    begin
      p := n_all - n_pre;
      if p > POST_N then p := POST_N; end if;
      return p;
    end function;

    impure function exp_pre return integer is
      variable room : integer;
    begin
      room := DEPTH - exp_post;
      if n_pre > room then return room; else return n_pre; end if;
    end function;

    -- Drain the whole capture and compare every entry, in order, against the
    -- reference list.
    procedure drain_and_check is
      variable want, first : integer;
    begin
      want  := exp_pre + exp_post;
      first := n_all - want;
      ck("n_valid",   to_integer(n_valid_s),   want);
      ck("pre_kept",  to_integer(pre_kept_s),  exp_pre);
      ck("post_kept", to_integer(post_kept_s), exp_post);
      if n_pre > (DEPTH - exp_post) then
        ck("pre_lost", sl2i(pre_lost_s), 1);
      else
        ck("pre_lost", sl2i(pre_lost_s), 0);
      end if;

      for idx in 0 to want-1 loop
        rd_en <= '1';
        wait for 0 ns;
        tickc;
        ck("rd_valid", sl2i(rd_valid_s), 1);
        ck("rd_index", to_integer(rd_index_s), idx);
        -- ---- The summary outputs are re-checked on every drain cycle. ----
        --
        -- They are stable while frozen, so checking them once per capture is
        -- "enough" in the sense that it catches a wrong value. It is not
        -- enough in the sense that matters: pre_lost is the difference
        -- between "the cause is not in this capture" and "there is no
        -- cause", and a mutation that hard-wires it low was dying on ONE
        -- check per capture. Re-checking it per entry costs nothing and
        -- turns a cornered mutation into a killed one.
        if n_pre > (DEPTH - exp_post) then
          ck("pre_lost stable", sl2i(pre_lost_s), 1);
        else
          ck("pre_lost stable", sl2i(pre_lost_s), 0);
        end if;
        ck("n_valid stable",   to_integer(n_valid_s),   want);
        ck("pre_kept stable",  to_integer(pre_kept_s),  exp_pre);
        ck("post_kept stable", to_integer(post_kept_s), exp_post);
        -- THE check. Not the count -- the contents, in order.
        ck("rd_type",  to_integer(unsigned(rd_type_s)),
                       to_integer(unsigned(a_ty(first + idx))));
        ck("rd_ep",    to_integer(unsigned(rd_ep_s)),
                       to_integer(unsigned(a_ep(first + idx))));
        ck("rd_code",  to_integer(unsigned(rd_code_s)),
                       to_integer(unsigned(a_code(first + idx))));
        ck("rd_show",  sl2i(rd_show_s),
           sl2i(view_mask(to_integer(unsigned(a_ty(first + idx))))));
      end loop;
      rd_en <= '0';
      -- Reading past the end yields nothing at all rather than wrapping
      -- round to the oldest entry again.
      rd_en <= '1';
      wait for 0 ns;
      tickc;
      ck("rd_valid past end", sl2i(rd_valid_s), 0);
      rd_en <= '0';
    end procedure;
  begin
    wait until rising_edge(clk);
    wait until rising_edge(clk);
    wait until rising_edge(clk);
    rst_n <= '1';
    wait for 1 ns;

    -- ================= PHASE 0 -- traffic before the capture is armed =====
    --
    -- Every event type offered while the block is IDLE. None of it may be
    -- recorded: a capture that quietly began before it was armed contains
    -- events from whatever the bus was doing last time, in the same buffer,
    -- with no marker between them.
    for i in 0 to 3 loop
      ev(std_logic_vector(to_unsigned(i, 2)), x"9",
         std_logic_vector(to_unsigned(16#70# + i, 8)));
    end loop;
    ck("idle state",      to_integer(unsigned(state_s)),
                          to_integer(unsigned(S_IDLE)));
    ck("nothing written", to_integer(n_written_s), 0);

    -- ================= PHASE 1 -- THE RETENTION EDGE, SWEPT ==============
    --
    -- Every pre-trigger event count from 0 to DEPTH + 8. Below
    -- DEPTH - POST_N everything before the trigger survives; above it the
    -- oldest entries are gone and pre_lost must say so. The whole capture is
    -- drained and compared entry by entry on every one of the 73 runs.
    for p in 0 to DEPTH + 8 loop
      do_arm;
      for i in 0 to p-1 loop
        ev(std_logic_vector(to_unsigned(i mod 4, 2)),
           std_logic_vector(to_unsigned(i mod 16, 4)),
           std_logic_vector(to_unsigned((i + 1) mod 256, 8)));
      end loop;
      trig;
      for i in 0 to POST_N-1 loop
        ev("11", x"F", std_logic_vector(to_unsigned(16#A0# + i, 8)));
      end loop;
      ck("frozen", to_integer(unsigned(state_s)),
                   to_integer(unsigned(S_FROZEN)));
      drain_and_check;
      if p > DEPTH - POST_N then n_lost_cases := n_lost_cases + 1;
      else                       n_intact_cases := n_intact_cases + 1;
      end if;
    end loop;
    if n_lost_cases /= 16 or n_intact_cases /= 57 then
      errors := errors + 1;
      write(ln, string'("FAIL edge split: lost=") & integer'image(n_lost_cases)
                & " intact=" & integer'image(n_intact_cases));
      writeline(output, ln);
    end if;

    -- ================= PHASE 2 -- the trigger is sampled AFTER the write ==
    --
    -- An event and the trigger on the same cycle: the event belongs to the
    -- PRE side, because it is part of what led up to the symptom rather than
    -- part of the aftermath.
    do_arm;
    for i in 0 to 9 loop
      ev("00", x"1", std_logic_vector(to_unsigned(i + 1, 8)));
    end loop;
    arm <= '0'; ev_valid <= '1'; trigger <= '1'; rd_en <= '0';
    ev_type <= "01"; ev_ep <= x"2"; ev_code <= x"5A";
    wait for 0 ns;
    a_ty(n_all) := "01"; a_ep(n_all) := x"2"; a_code(n_all) := x"5A";
    n_all := n_all + 1; n_pre := n_pre + 1;
    tickc;
    ev_valid <= '0'; trigger <= '0';
    ck("state after simultaneous trigger", to_integer(unsigned(state_s)),
       to_integer(unsigned(S_CAPT)));
    for i in 0 to POST_N-1 loop
      ev("11", x"F", std_logic_vector(to_unsigned(16#B0# + i, 8)));
    end loop;
    ck("pre_kept with simultaneous trigger", to_integer(pre_kept_s), 11);
    drain_and_check;

    -- ================= PHASE 3 -- the mask hides, it does not delete =====
    --
    -- The same capture re-created and drained under all 16 view masks. The
    -- CONTENTS are identical every time; only rd_show changes. A design that
    -- filtered on the way in would return a different capture for each mask.
    for m in 0 to 15 loop
      do_arm;
      for i in 0 to 39 loop
        ev(std_logic_vector(to_unsigned(i mod 4, 2)),
           std_logic_vector(to_unsigned(i mod 16, 4)),
           std_logic_vector(to_unsigned((i + 1) mod 256, 8)));
      end loop;
      trig;
      for i in 0 to POST_N-1 loop
        ev("10", x"7", std_logic_vector(to_unsigned(16#C0# + i, 8)));
      end loop;
      view_mask <= std_logic_vector(to_unsigned(m, 4));
      wait for 0 ns;
      drain_and_check;
    end loop;
    view_mask <= "1111";
    wait for 0 ns;

    -- ================= PHASE 4 -- draining before the freeze yields none ==
    do_arm;
    for i in 0 to 19 loop
      ev("00", x"3", std_logic_vector(to_unsigned(i + 1, 8)));
    end loop;
    trig;
    ev("11", x"3", x"D0");
    ck("still capturing", to_integer(unsigned(state_s)),
       to_integer(unsigned(S_CAPT)));
    rd_en <= '1';
    wait for 0 ns;
    tickc;
    ck("no drain while capturing", sl2i(rd_valid_s), 0);
    rd_en <= '0';

    -- ================= PHASE 5 -- re-arming throws the capture away ======
    --
    -- And it must throw away the pre-trigger count with it. A stale pre_r
    -- makes the NEXT capture report a pre_lost it did not earn.
    do_arm;
    for i in 0 to 4 loop
      ev("01", x"4", std_logic_vector(to_unsigned(i + 1, 8)));
    end loop;
    trig;
    for i in 0 to POST_N-1 loop
      ev("11", x"E", std_logic_vector(to_unsigned(16#E0# + i, 8)));
    end loop;
    ck("re-armed pre_kept", to_integer(pre_kept_s), 5);
    ck("re-armed pre_lost", sl2i(pre_lost_s), 0);

    -- ---- Traffic AFTER the freeze must not touch the capture. ----
    --
    -- The bus does not stop because you triggered. If a frozen capture keeps
    -- absorbing events, the trace changes while you are reading it -- and
    -- the entries you already looked at are no longer the entries that are
    -- there. Every type is offered, and the drain below is compared against
    -- the model as though none of them happened, because none of them may.
    base_written := to_integer(n_written_s);
    for i in 0 to 3 loop
      ev(std_logic_vector(to_unsigned(i, 2)), x"A",
         std_logic_vector(to_unsigned(16#F0# + i, 8)));
    end loop;
    ck("frozen: nothing written", to_integer(n_written_s), base_written);
    ck("frozen: still frozen", to_integer(unsigned(state_s)),
       to_integer(unsigned(S_FROZEN)));
    drain_and_check;

    -- ================= PHASE 6 -- random ================================
    --
    -- Random capture lengths and random event mixes, every one of them
    -- drained and compared in full.
    for i in 0 to 299 loop
      do_arm;
      w_v := rnd_nat mod 90;
      for p in 0 to w_v-1 loop
        ev(std_logic_vector(to_unsigned(rnd_nat mod 4, 2)),
           std_logic_vector(to_unsigned(rnd_nat mod 16, 4)),
           std_logic_vector(to_unsigned(rnd_nat mod 256, 8)));
      end loop;
      trig;
      for p in 0 to POST_N-1 loop
        ev(std_logic_vector(to_unsigned(rnd_nat mod 4, 2)),
           std_logic_vector(to_unsigned(rnd_nat mod 16, 4)),
           std_logic_vector(to_unsigned(rnd_nat mod 256, 8)));
      end loop;
      view_mask <= std_logic_vector(to_unsigned(rnd_nat mod 16, 4));
      wait for 0 ns;
      drain_and_check;
    end loop;
    view_mask <= "1111";
    wait for 0 ns;

    -- ================= the exhaustiveness proof ==========================
    n_reach := 0;
    for k in 0 to 15 loop n_reach := n_reach + reach(k); end loop;
    if n_reach /= 16 then
      errors := errors + 1;
      write(ln, string'("FAIL state x type reach ") & integer'image(n_reach)
                & "/16");
      writeline(output, ln);
      for k in 0 to 15 loop
        if reach(k) = 0 then
          write(ln, string'("  unreached state=") & integer'image(k / 4)
                    & " type=" & integer'image(k mod 4));
          writeline(output, ln);
        end if;
      end loop;
    end if;

    write(ln, string'("steps=") & integer'image(steps)
              & " checks=" & integer'image(checks)
              & " reach=" & integer'image(n_reach) & "/16"
              & " errors=" & integer'image(errors));
    writeline(output, ln);
    write(ln, string'("retention edge: intact=")
              & integer'image(n_intact_cases) & " lost="
              & integer'image(n_lost_cases));
    writeline(output, ln);
    write(ln, string'("written=") & integer'image(to_integer(n_written_s))
              & " overwritten=" & integer'image(to_integer(n_over_s))
              & " shown=" & integer'image(to_integer(n_shown_s))
              & " hidden=" & integer'image(to_integer(n_hidden_s)));
    writeline(output, ln);
    if errors = 0 then
      write(ln, string'("PASS: 0 errors in ") & integer'image(checks)
                & " checks");
    else
      write(ln, string'("FAIL: ") & integer'image(errors) & " errors in " &
                integer'image(checks) & " checks");
    end if;
    writeline(output, ln);
    done <= true;
    wait;
  end process;
end architecture;

11. Exhaustive Verification

MeasureVerilogSystemVerilogVHDL
pre-trigger counts swept73 / 7373 / 7373 / 73
intact / lost split57 / 1657 / 1657 / 16
capture state × event type16 / 1616 / 1616 / 16
Steps393553935539567
Checks executed299223299223301619
events written201422014220296
events overwritten204520452187
entries shown105321053210410
entries hidden749674967678
ResultPASSPASSPASS

57 / 16 is the row that pins the edge. DEPTH - POST_N is 56, so the counts from 0 to 56 inclusive must retain everything — 57 of them — and the 16 above must not. Both halves are required, and every one of the 73 runs is compared entry by entry rather than by length.

16 / 16 on state × event type includes the two states where events are supposed to be discarded. Offering traffic in IDLE and in FROZEN and requiring that nothing is written is not a formality: a capture that quietly began before it was armed contains events from whatever the bus was doing last time, in the same buffer, with no marker between them.

12. Mutation Testing

#MutationVerilogSysVerVHDL
U1recording starts at the trigger200369200369201324
U5the trigger is never taken; nothing ever freezes147226147226148503
U6re-arming does not clear the pre-trigger count555445554457097
U4the post-trigger window is not subtracted from the room445274452742973
U7the drain starts at the write pointer: newest first264802648027281
U2the view mask is applied on the way IN221922219222263
U3pre_lost is never asserted851585158320
—unmutated baseline000

All seven die in all three languages.

U1 is the chapter, and it is the largest score in the table. Recording only after the trigger is not an exotic defect — it is the factory default on most analysers, and it is what you get by not changing a setting.

U7 is worth reading carefully. It does not lose a single entry. The count is right, pre_kept and post_kept are right, pre_lost is right, and every event that was captured is in the output — starting from the newest and wrapping round to the oldest. A testbench that checked lengths would pass it completely, which is why this one compares contents in order.

U3 is the most dangerous mutation here despite being the smallest. Everything still works: the capture is taken, the window is right, the entries are correct. The only thing missing is the flag that says the cause may be off the front, and without it a capture that does not contain the bug is indistinguishable from one that proves there was none. A negative result becomes evidence.

13. The Mutation That Was Nearly Cornered, and a Column That Lied

U3 scored 131.

Not because the check was wrong — pre_lost was checked, correctly, once per capture. That is one check against roughly 400 captures, of which only the ~130 that actually overflowed could fail. Correct, and thin: one reordered phase from zero.

And U2's VHDL column read 22263 against 8313 for the other two.

A 2.7× gap looks like a VHDL testbench gap. It was a mutation-fidelity failure: the Verilog injection guarded only the type write, leaving the endpoint and code fields stored — a mislabelled entry rather than a missing one — while the VHDL injection guarded all three. Two different mutations in the same row.

Guarding all three in every language brought it to 22192 / 22192 / 22263.

14. Debugging Walkthrough: Reading a Capture That Is Telling the Truth

The setup. A composite device intermittently stops responding on its bulk endpoint. You have an analyser and one afternoon.

Step 1 — configure it to record, not to start recording. Pre-trigger buffer to maximum, trigger position as early as the tool allows. If the tool expresses it as a percentage, 5–10% post-trigger is usually right for a protocol bug: the host's recovery sequence is well specified and predictable, and it is not where the defect is.

Step 2 — do not filter at capture. Turn every class of traffic on, including SOFs if the tool will take them. The file will be large. The file being large is the point; you can filter the view afterwards, and you cannot un-filter a capture.

Step 3 — trigger on the symptom anyway. You do not have a better condition. What has changed is that the symptom now sits near the end of the window rather than at the start of it.

Step 4 — check whether the capture is long enough before reading it. If the tool reports that the pre-trigger buffer wrapped, the cause may be off the front and a capture that looks full tells you nothing about that. Increase the depth, reduce the post-trigger share, or narrow what you are recording — in that order, because the third one costs context.

Step 5 — read backwards from the symptom. Not forwards from the start. The question is what is the last thing that was different, and the answer is usually within a few dozen transactions of the trigger.

Step 6 — filter the view, repeatedly, and keep the capture. Show only errors to find the symptom. Show only that endpoint to see its history. Show everything around the interesting moment. Every one of those is a different question against the same bytes, which is only possible because nothing was discarded at capture time.

15. UVM: A Trace Recorder as an Environment Service

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// A trace recorder is not a checker. It produces no pass/fail of its own --
// it produces the CONTEXT that makes somebody else's failure readable, and
// it has to already be running when that failure happens.
//
// So it subscribes to the same analysis port everything else does, and it
// records unconditionally.
class usb_trace_recorder extends uvm_subscriber #(usb_bus_item);
  `uvm_component_utils(usb_trace_recorder)

  int DEPTH  = 4096;
  int POST_N = 64;

  usb_bus_item ring[$];      // the circular buffer
  int  post_left;
  bit  triggered, frozen;
  int  pre_total, post_total;
  bit  pre_lost;

  function new(string name, uvm_component parent);
    super.new(name, parent);
  endfunction

  // ---- WRITE FIRST, always, with no filter of any kind. ----
  //
  // Not "if it is an error", not "if UVM_HIGH is enabled". The moment a
  // condition appears here, the context that explains the error is gone and
  // no amount of later processing brings it back.
  function void write(usb_bus_item t);
    if (frozen) return;      // the bus does not stop; the capture does

    ring.push_back(t);
    if (ring.size() > DEPTH) begin
      void'(ring.pop_front());
      // A full buffer looks exactly like a sufficient one. This is the only
      // place that difference is observable, so it is recorded here.
      if (!triggered) pre_lost = 1;
    end

    if (triggered) begin
      post_total++;
      if (--post_left <= 0) frozen = 1;
    end else begin
      pre_total++;
    end
  endfunction

  // Called by whichever component detected the symptom -- a scoreboard, a
  // protocol checker, the endpoint health monitor from chapter 25.3. The
  // recorder does not decide what a symptom is; it only decides what to
  // keep when somebody else says one happened.
  function void trigger(string why);
    if (triggered) return;   // the FIRST symptom is the one with context
    triggered = 1;
    post_left = POST_N;
    `uvm_info("TRACE", $sformatf("triggered: %s", why), UVM_LOW)
  endfunction

  // ---- The report says what it does NOT have. ----
  function void report_phase(uvm_phase phase);
    super.report_phase(phase);
    if (!triggered) begin
      `uvm_info("TRACE", "no trigger: nothing to dump", UVM_LOW)
      return;
    end

    `uvm_info("TRACE",
      $sformatf("capture: %0d entries (%0d before the trigger, %0d after)",
                ring.size(), ring.size() - post_total, post_total),
      UVM_LOW)

    // The honest part. A capture missing its cause and a capture proving
    // there was none look identical from the inside.
    if (pre_lost)
      `uvm_warning("TRACE/PRE_LOST",
        $sformatf("the pre-trigger buffer wrapped: %0d events occurred before the trigger and %0d were kept -- the cause may be off the front of this capture",
                  pre_total, ring.size() - post_total))

    // And the post-trigger window's real cost, stated rather than implied.
    if (post_total * 4 > DEPTH)
      `uvm_warning("TRACE/POST_HEAVY",
        $sformatf("%0d of %0d entries are AFTER the symptom -- that is %0d%% of the buffer spent on the host's recovery sequence",
                  post_total, DEPTH, (100 * post_total) / DEPTH))

    dump(.mask('1));         // everything. Filter the VIEW, not the file.
  endfunction

  // The mask is a VIEW. The queue is untouched, so the same capture can be
  // dumped again with a different mask and answer a different question.
  function void dump(bit [3:0] mask);
    int shown, hidden;
    foreach (ring[i])
      if (mask[ring[i].kind]) begin
        `uvm_info("TRACE", $sformatf("[%0d] %s", i, ring[i].convert2string()),
                  UVM_HIGH)
        shown++;
      end else hidden++;
    // The number on screen is never allowed to be mistaken for the number
    // captured.
    `uvm_info("TRACE",
      $sformatf("%0d shown, %0d hidden by the view mask (all %0d still held)",
                shown, hidden, ring.size()), UVM_LOW)
  endfunction
endclass

16. Common Misconceptions

"The trigger starts the capture." On a default configuration, yes — and that is the setting this chapter is about.

"The capture is complete." It is complete about an interval. Which interval is the question.

"A bigger buffer fixes it." Only if the trigger position is also right. A 1 GB buffer used entirely for post-trigger data contains no causes.

"50% pre/post is a sensible default." For a protocol bug it spends half the buffer on the host's recovery sequence, which is specified behaviour.

"Filtering at capture saves space." It destroys the evidence and keeps the symptom, which is the opposite of useful.

"A full buffer means I have enough." A full buffer looks identical to a sufficient one. Only an explicit wrapped indicator distinguishes them.

"Nothing in the capture explains it, so the cause is elsewhere." Or it is off the front. Those are different conclusions and one bit tells them apart.

"The capture is frozen, so the bus has stopped." The bus has not stopped. If the capture keeps absorbing events, what you read is not what you looked at.

"A capture with the right number of events is the right capture." U7 keeps every event, gets every count right, and returns them newest-first.

17. Exercises

1. A buffer holds 4096 events and the trigger position is 50%. A bug's cause is 3000 transactions before its symptom. Say whether the capture contains it, and what the smallest change is that makes it certain.

2. The retention sweep reports 57 intact and 16 lost for DEPTH = 64, POST_N = 8. Derive both numbers, then give them for POST_N = 32.

3. U7 returns every captured event, in the wrong order, with every count correct. Write the smallest check that catches it and explain why checking n_valid cannot.

4. The design samples the trigger after the write. Construct the capture in which that decision changes which events survive, and say which event moves.

5. pre_lost is set when the buffer wraps before the trigger. Show that it cannot be recovered afterwards from the capture's contents alone.

6. The UVM recorder takes the first trigger and ignores later ones. Construct the failure where that is the wrong choice, and say what you would change.

7. Combine this block with 25.3's endpoint monitor: which of its four reports should call trigger(), and what POST_N does each one want?

18. Summary

IdeaWhy it matters
A trigger fires on the symptomand the symptom is the end of the story
The cause is always before the triggerso a capture that starts there records the aftermath
Record always; the trigger stops ita circular buffer chooses a window that straddles the trigger
The post-trigger window eats the pre-trigger historymore aftermath costs exactly that much cause
Filter on the way outthe error is the symptom; the ordinary traffic is the evidence
A full buffer looks like a sufficient oneonly an explicit flag distinguishes them
pre_lost separates two different findings"not in this capture" is not "there is no cause"
A frozen capture must stay frozenor the trace changes while you read it
Traffic before arm is not part of the captureor it contains events from the last one, unmarked
Right count, wrong order is still wrongU7 keeps everything and returns it newest-first
A stable output checked once is checked onceU3 went 131 → 8515 by moving four checks into the loop
An out-of-line column is usually the mutantthe third time in this module: Q1, R1, U2
73 counts swept, 16/16 state × type7 mutations, all killed in 3 languages

Tooling

StepCommand
Verilog-2005iverilog -g2005 -o tf_v.out tf_v.v tf_v_tb.v && ./tf_v.out
SystemVerilogiverilog -g2012 -o tf_sv.out tf_sv.sv tf_sv_tb.sv && ./tf_sv.out
VHDL-2008 analysenvc --std=2008 -a tf_vhdl.vhd tf_vhdl_tb.vhd
VHDL-2008 elaboratenvc --std=2008 -e tb_tf_vhdl
VHDL-2008 runnvc --std=2008 -r tb_tf_vhdl
One mutationiverilog -g2005 -DMUT_U7 -o mm tf_v_mut.v tf_v_tb.v && ./mm

All three implementations pass with 0 errors: the retention edge swept at all 73 pre-trigger counts with every capture drained and compared entry by entry, all 16 state × event-type situations reached including the two where events must be discarded, and every one of the seven mutations killed.


Module 25 Complete

Seven chapters, seven synthesisable blocks, and one argument running through all of them:

ChapterThe blockThe claim
25.1usb_enum_trackera retry restarts the sequence, not the diagnosis
25.2usb_descriptor_validatorthe error names a field that is perfectly correct
25.3usb_endpoint_healtha NAK is not an error and a STALL is not a NAK
25.4usb_error_triageretries mask the error rate; the denominator is attempts
25.5usb_crc_analyzerthe distribution names the layer; the count says nothing
25.6usb_power_monitorin-rush is legal and sag is not, and the difference is a window
25.7usb_trace_filterthe cause is always before the trigger

Every one of them is a variation on the same theme: the raw count is not the measurement. A NAK count, a transfer error rate, a CRC total, a voltage reading, a capture length — each is a number that looks like an answer and is not one until it has been divided by something, split by something, or placed in time.

Continue learning

Standards & specifications

Governing standard
USB-IF (Universal Serial Bus Specification)(opens USB Implementers Forum (USB-IF) in a new tab)

Defines the USB bus — its electrical signalling, connectors, packet and transaction model, device framework and the descriptors a device must expose — together with the device-class specifications layered on it. It does not define host-controller register interfaces (xHCI and EHCI are separate documents) nor any operating system's driver architecture.

This page also covers RTL structure, verification approach and debugging technique. Those are engineering practice built on the standard, not requirements the standard itself imposes.

Where this fits

Part of the USB curriculum.