Skip to content
VLSI Mentor

I²C · Module 20

Protocol Checking and Error-Injection Strategy

Where each protocol rule belongs, why forcing a flag inside the design proves nothing about detection, and why a fault's position is as much stimulus as its size. Ends with a simulation in which 0xFF arrives as 0x7F with every byte acknowledged — silent corruption under a completely clean protocol verdict.

Every chapter so far has verified correct traffic. The environment drives legal transfers, the target answers legally, and the checks confirm that what happened matches what the contract obliged.

That leaves out the conditions a real bus actually produces. A line held down by a device that crashed. A disturbance narrow enough to be invisible and wide enough not to be. Framing appearing in the middle of a byte. None of these can occur in an environment where every participant behaves, which means the design's handling of them is — up to this point — entirely untested.

1. Where a Rule Belongs

Before injecting anything, the rules have to live somewhere. There are three places, and the choice determines how many times each rule is checked.

rulewhere it liveswhy there
SDA changes only while SCL is low, except framingmonitoruniversal, so it should check every trace ever produced
START and STOP are SDA edges while SCL is highmonitorsame, and it is the monitor's own framing basis
a byte is eight bits and an acknowledgemonitorstructural; every consumer depends on it
the ninth slot has exactly one driverbus modelundetectable from the resolved line — see Section 2
no participant drives a line highport structuremake it impossible rather than checked
a write to a read-only register is refusedreference modela device decision, not a protocol rule
a stretch is boundeddriver policya choice, not a protocol fact

The interesting rows are the last three. "No participant drives high" is enforced by every component having only drive_low outputs, which makes the rule unbreakable rather than monitored. "A stretch is bounded" cannot be a protocol check at all, because the protocol places no bound — it is the environment stating how patient it intends to be. And the read-only refusal is not a protocol rule in any sense, which is why it lives with the contract.

2. The Injection Antipattern

Here is the shape that almost every first error-injection attempt takes.

Azvya Education Pvt. Ltd.VLSI Mentor
What this proves, and what it does not
   // "test that the target handles a framing error"
   dut.force_framing_error = 1'b1;
   @(posedge clk);
   ck("the target flagged an error", dut.error_flag, 1);

That test passes, and it establishes one fact: the target reacts to its own error input.

It proves nothing about whether the target would notice the corresponding real condition, because the real condition never occurred. No illegal edge appeared on the bus, nothing on the observation path was exercised, and the entire mechanism between "a bad thing happened on the wire" and "the design noticed" was bypassed. That mechanism is the part most likely to be wrong.

So the injector here is a third participant on the shared bus, not a hook inside the design. It pulls lines low exactly as every other device does, and the design has no idea it exists.

3. The Injector

Azvya Education Pvt. Ltd.VLSI Mentor
i2c_fault_inj.sv — faults as bus events, not as flags
   // -----------------------------------------------------------------------------
   // i2c_fault_inj.sv
   // A fault injector that is a DEVICE ON THE BUS, not a hook inside the DUT.
   //
   // WHY THAT DISTINCTION IS THE WHOLE CHAPTER. Setting `dut.force_error = 1` proves the
   // DUT reacts to its own error input. It proves nothing about whether the DUT would
   // notice the corresponding real condition, because the real condition never occurred
   // and nothing on the observation path was exercised.
   //
   // This injector is a third participant on the wired-AND bus. Everything it does, it
   // does by pulling a line LOW -- which is the only thing any device on this bus can do
   // -- so its faults reach the DUT through exactly the path a real disturbance would:
   // the resolved bus, the pads, the synchroniser, the filter. The monitor sees them
   // because they are genuinely there.
   //
   // THE FAULTS, and what makes each one a fault rather than just activity:
   //
   //   F_NONE        idle. The injector releases both lines and is invisible.
   //
   //   F_SDA_STUCK   hold SDA LOW indefinitely. The bus cannot be framed: a STOP needs
   //                 SDA to RISE while SCL is high, and it cannot. This is the classic
   //                 wedged bus, and the correct DUT response is a timeout.
   //
   //   F_SCL_STUCK   hold SCL LOW indefinitely. Every participant that waits for SCL to
   //                 rise waits forever, which is why Chapter 20.5's driver has a
   //                 stretch timeout rather than an unbounded wait.
   //
   //   F_SDA_GLITCH  pull SDA LOW for GLITCH_CLKS and release. Narrow enough to be
   //                 rejected by Chapter 19.5's filter, wide enough to be seen by an
   //                 unfiltered observer -- so it discriminates between the two.
   //
   //   F_EXTRA_START pull SDA LOW while SCL is HIGH, mid-transfer. Every conforming
   //                 device must read that as a START and abandon what it was doing.
   //                 This is the fault that tests framing recovery, and it is
   //                 indistinguishable from a real START because it IS one.
   //
   // WHAT IT NEVER DOES: drive a line HIGH. It has no way to. An injector that could
   // would be testing a condition the bus cannot physically produce, and the DUT's
   // response to it would not be worth knowing. If a push-pull contender is the thing
   // under test, that is a different component and it must be labelled as modelling
   // non-conforming hardware -- Chapter 20.9 §7 makes that argument.
   //
   // BOUNDED: `arm` is a level owned by the environment, and every timed fault counts
   // down. The injector never waits for anything.
   // -----------------------------------------------------------------------------

   module i2c_fault_inj #(
      // Clocks a glitch lasts. Chapter 19.5's filter rejects runs shorter than its
      // threshold, so a glitch below that must be invisible to a filtered observer and
      // visible to an unfiltered one -- which is how F_SDA_GLITCH discriminates.
      parameter int GLITCH_CLKS = 2
   ) (
      input  logic clk,
      input  logic rst_n,

      // ---- the fault to inject -------------------------------------------------
      input  logic [2:0] fault,      // see the F_* encoding below
      input  logic       arm,        // level: inject while asserted

      // ---- the bus: LOW or RELEASE, nothing else -------------------------------
      output logic scl_drive_low,
      output logic sda_drive_low,
      input  logic scl_in,           // RESOLVED -- needed to time an SDA-while-SCL-high
      input  logic sda_in,

      // ---- evidence that the fault actually happened ---------------------------
      // A fault the environment believes it injected but which never reached the bus is
      // the failure mode Chapter 20.9 §6 is about. These counters are how a bench proves
      // the injection was real rather than merely configured.
      output logic [15:0] n_injected,
      output logic        injecting
   );

      localparam [2:0] F_NONE        = 3'd0,
                       F_SDA_STUCK   = 3'd1,
                       F_SCL_STUCK   = 3'd2,
                       F_SDA_GLITCH  = 3'd3,
                       F_EXTRA_START = 3'd4;

      logic [15:0] cnt;
      logic        fired;       // this arming has already produced its one-shot fault

      always @(posedge clk or negedge rst_n) begin
         if (!rst_n) begin
            scl_drive_low <= 1'b0;
            sda_drive_low <= 1'b0;
            injecting     <= 1'b0;
            n_injected    <= 16'd0;
            cnt           <= 16'd0;
            fired         <= 1'b0;
         end else if (!arm) begin
            // Disarmed: release both lines and forget. An injector that kept holding a
            // line after being disarmed would corrupt every later test in the run.
            scl_drive_low <= 1'b0;
            sda_drive_low <= 1'b0;
            injecting     <= 1'b0;
            cnt           <= 16'd0;
            fired         <= 1'b0;
         end else begin
            case (fault)
               F_SDA_STUCK: begin
                  // A level fault: held for as long as the environment arms it.
                  if (!sda_drive_low) n_injected <= n_injected + 16'd1;
                  sda_drive_low <= 1'b1;
                  injecting     <= 1'b1;
               end

               F_SCL_STUCK: begin
                  if (!scl_drive_low) n_injected <= n_injected + 16'd1;
                  scl_drive_low <= 1'b1;
                  injecting     <= 1'b1;
               end

               F_SDA_GLITCH: begin
                  // A one-shot fault, GLITCH_CLKS wide. `fired` stops it repeating while
                  // still armed, so the environment gets exactly one glitch per arming
                  // and can count them.
                  if (!fired) begin
                     if (cnt == 16'd0) begin
                        n_injected    <= n_injected + 16'd1;
                        sda_drive_low <= 1'b1;
                        injecting     <= 1'b1;
                        cnt           <= 16'd1;
                     end else if (cnt < GLITCH_CLKS[15:0]) begin
                        cnt <= cnt + 16'd1;
                     end else begin
                        sda_drive_low <= 1'b0;
                        injecting     <= 1'b0;
                        fired         <= 1'b1;
                     end
                  end
               end

               F_EXTRA_START: begin
                  // Timed against the RESOLVED bus, not against a clock count: pulling
                  // SDA low is only a START if SCL is HIGH at that moment. Waiting for
                  // the real condition is what makes this fault indistinguishable from a
                  // genuine START -- because it is one.
                  if (!fired && scl_in && sda_in) begin
                     n_injected    <= n_injected + 16'd1;
                     sda_drive_low <= 1'b1;
                     injecting     <= 1'b1;
                     cnt           <= 16'd1;
                  end else if (injecting) begin
                     if (cnt < GLITCH_CLKS[15:0]) cnt <= cnt + 16'd1;
                     else begin
                        // Hold SDA low past the SCL fall so the framing is unambiguous,
                        // then release.
                        sda_drive_low <= 1'b0;
                        injecting     <= 1'b0;
                        fired         <= 1'b1;
                     end
                  end
               end

               default: begin       // F_NONE
                  scl_drive_low <= 1'b0;
                  sda_drive_low <= 1'b0;
                  injecting     <= 1'b0;
               end
            endcase
         end
      end

   endmodule

Five faults, and the choice of which are worth having is the design decision:

faultwhat it modelswhy it is interesting
F_SDA_STUCKa device that crashed holding SDAthe bus cannot be framed at all: a STOP needs SDA to rise
F_SCL_STUCKa device that crashed holding SCLtests the environment's timeout, not the design's
F_SDA_GLITCHa narrow disturbancethe sampling-window question, and it has two answers
F_EXTRA_STARTframing where none belongsthe only fault that is indistinguishable from legal traffic
F_NONEarmed, doing nothingthe baseline that makes the others attributable

F_EXTRA_START is the one worth dwelling on. It pulls SDA low while SCL is high — and by the specification that is a START. It is not a corrupted START or an approximation of one; every conforming device must read it as one, and the target reframing is correct behaviour rather than a bug. Timing it against the resolved bus rather than against a clock count is what makes it genuine:

Azvya Education Pvt. Ltd.VLSI Mentor
i2c_fault_inj.sv — timed against the bus, not against a counter
               F_EXTRA_START: begin
                  // Timed against the RESOLVED bus, not against a clock count: pulling
                  // SDA low is only a START if SCL is HIGH at that moment. Waiting for
                  // the real condition is what makes this fault indistinguishable from a
                  // genuine START -- because it is one.
                  if (!fired && scl_in && sda_in) begin
                     n_injected    <= n_injected + 16'd1;
                     sda_drive_low <= 1'b1;
                     injecting     <= 1'b1;
                     cnt           <= 16'd1;

4. Position Is Part of the Stimulus

The first version of this bench armed the injector before the transfer. F_EXTRA_START fires as soon as it sees SCL high with SDA high — and an idle bus already satisfies that. So it fired into the gap before the controller's own START, produced a complete one-byte transaction of its own, and left the real transfer entirely untouched.

The injection check passed. Nothing about mid-transfer recovery had been tested.

The second version armed it at the first observed byte, and the glitch landed in the acknowledge slot — where the target is pulling SDA low anyway and is not sampling it. Pulling an already-low line changes nothing, so the glitch was harmless for a reason that has nothing to do with its width. Widening it from 2 clocks to 60 in a mutation run changed no result at all, which is how the vacuity was found.

5. The Result: Silent Corruption Under a Clean Verdict

With position controlled, the glitch width becomes a measurable variable. The bench places a disturbance inside a data byte of 0xFF — every bit high, so pulling SDA low genuinely changes the line — sixty clocks after the second observed byte, and changes only the width.

glitch widthreaches the bus?every byte acknowledged?byte received
2 clocksyes, resolved SDA pulled lowyes0xFF — intact
40 clocksyes, same positionyes0x7F — corrupted

Two clocks does not span the rising edge at which the target samples SDA, so it is invisible. Forty clocks does, so bit 7 is captured low and 0xFF arrives as 0x7F.

And the acknowledges are unaffected in both cases.

There is a second observation in that test which is subtler and matters more for the module's central claim. The bench checks that the monitor's observed byte and the target's stored byte agree — and that both differ from what the driver was asked to send.

Azvya Education Pvt. Ltd.VLSI Mentor
i2c_fault_tb.sv — which two things are compared, and why
         ck("T2b but the wire did NOT carry the byte that was requested",
            (mon_last_data !== 8'hFF), 1);
         ck("T2b the observed byte is the one bit-7 corruption predicts",
            mon_last_data, 8'h7F);
         ck("T2b and the target stored exactly what the WIRE carried, not what was sent",
            dut_regs[1*8 +: 8], mon_last_data);

A scoreboard fed the driver's intent as its expected value would have called this a failure of the target. The target stored exactly what arrived; what arrived was not what was sent. The bus is the authority, and this is the concrete case that makes that principle worth the architecture it costs.

A disturbance matters if and only if it spans a sampling edge

10 cycles
A ten-cycle waveform. SCL alternates, with rising edges marked as sampling instants. An undisturbed SDA trace shows a high data bit. A narrow-glitch trace (gl 2) shows a brief low pulse falling between two rising edges. A wide-glitch trace (gl 40) shows a longer low period that covers a rising edge. Markers indicate the sampling edge for bit 7 and the acknowledge slot.wide: covers the edgewide: covers the edgeack slot: clean either wayack slot: cleaneither waysampling edge for this bitsampling edge for this bitthe narrow glitch lands here, between edgesthe narrow glitch landshere, between edgesSCLSDA cleanSDA gl 2SDA gl 40t0t1t2t3t4t5t6t7t8t9
The three SDA rows are the same bit with no glitch, a 2-clock glitch and a 40-clock glitch. The figure is schematic — the real widths are 2 and 40 clocks at a 32-clock bit period, and the glitch is placed 60 clocks after the second observed byte. What it shows faithfully is the mechanism: the sampling instant is a point, and a disturbance is visible if and only if it covers one.
Figure 1 — the same fault at the same position, twice. The narrow disturbance falls between sampling edges and is invisible to the target; the wide one spans the rising edge on which bit 7 is captured, so a high bit is read low. Note that the acknowledge slot is unaffected in both cases: the protocol layer reports a clean transfer either way, and only a check on the data content can tell them apart.

6. Proving Each Fault Reached the Wire

Every test in the bench establishes the consequence before the response. The order is what makes the second half meaningful.

Azvya Education Pvt. Ltd.VLSI Mentor
i2c_fault_tb.sv — faults on the wire, and proof they got there
   // -----------------------------------------------------------------------------
   // i2c_fault_tb.sv
   // Error injection against the REAL target, and proof each fault reached the bus.
   //
   //   controller BFM ─┐
   //   fault injector ─┼──▶ resolved bus ──▶ i2c_slave (Module 18, verified)
   //                   │         │
   //                   │         └──▶ i2c_mon (20.7): the independent witness
   //
   // EVERY TEST HERE HAS TWO HALVES, and the first is the one usually skipped:
   //
   //   1. DID THE FAULT REACH THE BUS? Not "was the injector armed" -- was the resolved
   //      line actually held, did the monitor actually see the framing event. An injector
   //      whose fault never reached the wire produces a test that passes for the wrong
   //      reason, and Chapter 20.9 §6 argues that configuration is not evidence.
   //
   //   2. DID THE DUT RESPOND CORRECTLY? Only meaningful once (1) is established.
   //
   // Every wait is bounded and a watchdog ends the run, which matters more here than
   // anywhere else in the module: two of these faults deliberately wedge the bus, and an
   // environment without timeouts would hang on its own stimulus.
   // -----------------------------------------------------------------------------
   `timescale 1ns/1ps

   module i2c_fault_tb;

      localparam int   HALF = 16;
      localparam [6:0] ADDR = 7'h50;
      localparam [2:0] F_NONE = 3'd0, F_SDA_STUCK = 3'd1, F_SCL_STUCK = 3'd2,
                       F_SDA_GLITCH = 3'd3, F_EXTRA_START = 3'd4;

      logic clk = 1'b0, rst_n = 1'b0;

      // three participants: controller BFM, the DUT, and the injector
      logic c_scl_low, c_sda_low, d_scl_low, d_sda_low;
      logic f_scl_low, f_sda_low, w_scl_low, w_sda_low;
      wire  scl, sda;
      wire [3:0] scl_in3, sda_in3, scl_rbl3, sda_rbl3;
      wire [7:0] scl_holders, sda_holders;

      // FOUR participants. The second injector is identical except for its glitch width,
      // and it exists because a glitch's consequence depends entirely on whether it spans
      // a sampling edge -- a fact no single width can demonstrate.
      i2c_line_model #(.N_DEV(4)) bus (
         .scl_drive_low({w_scl_low, f_scl_low, d_scl_low, c_scl_low}),
         .sda_drive_low({w_sda_low, f_sda_low, d_sda_low, c_sda_low}),
         .scl(scl), .sda(sda), .scl_in(scl_in3), .sda_in(sda_in3),
         .scl_released_but_low(scl_rbl3), .sda_released_but_low(sda_rbl3),
         .scl_holders(scl_holders), .sda_holders(sda_holders));

      logic       req = 1'b0, req_read = 1'b0, req_restart = 1'b0;
      logic [6:0] req_addr = ADDR;
      logic [2:0] req_len = 3'd1;
      logic [7:0] req_d0 = 8'h00, req_d1 = 8'h00, req_d2 = 8'h00;
      wire [7:0]  c_donecnt;
      wire        c_aacked, c_stretched, c_timeout, c_busy;
      wire [2:0]  c_nacked;
      wire [7:0]  c_r0, c_r1, c_r2;
      integer     dc_before;

      i2c_ctrl_bfm #(.HALF(HALF), .STRETCH_TIMEOUT(4000)) ctrl (
         .clk(clk), .rst_n(rst_n),
         .scl_drive_low(c_scl_low), .sda_drive_low(c_sda_low),
         .scl_in(scl), .sda_in(sda),
         .req(req), .req_addr(req_addr), .req_read(req_read), .req_len(req_len),
         .req_d0(req_d0), .req_d1(req_d1), .req_d2(req_d2), .req_restart(req_restart), .req_hold(1'b0),
         .done_count(c_donecnt), .obs_addr_acked(c_aacked), .obs_n_acked(c_nacked),
         .obs_r0(c_r0), .obs_r1(c_r1), .obs_r2(c_r2),
         .obs_stretched(c_stretched), .obs_timeout(c_timeout), .busy(c_busy));

      // ---- ASKING THE REAL TARGET TO STRETCH, FOR A BOUNDED TIME ---------------
      // `stall_req` is a LEVEL, not a pulse: the target stays busy for as long as it is
      // asserted. Holding it high for a whole transfer therefore does not mean "stretch
      // once", it means "never become ready" -- and the first version of T5b did exactly
      // that and produced a driver timeout, which is the T5 result arriving under the T5b
      // heading. A legal stretch has to be bounded to be legal, so the request is bounded
      // here, and it is timed off the monitor's observed address byte for the same reason
      // the injector is: so the stall lands inside a transfer rather than before it.
      logic       stall_arm = 1'b0;
      integer     stall_left;
      logic       stall_fired;        // evidence the request reached the target, latched
      wire        stall_req = (stall_left > 0);

      wire [8*8-1:0] dut_regs;
      wire [7:0] dut_pointer;
      wire dut_selected, dut_stretching;
      wire [15:0] dut_writes, dut_aborts, dut_conflict;

      i2c_slave #(.MY_ADDR(ADDR), .N_REG(8), .RO_MASK(8'h04),
                  .IDLE_CYCLES(900), .SYNC_DEPTH(2), .CNT_W(16)) dut (
         .clk(clk), .rst_n(rst_n),
         .scl_pin(scl), .sda_pin(sda),
         .scl_drive_low(d_scl_low), .sda_drive_low(d_sda_low),
         .stall_req(stall_req),
         .reg_flat(dut_regs), .pointer(dut_pointer),
         .selected(dut_selected), .stretching(dut_stretching),
         .n_phases(), .n_restarts(), .n_writes(dut_writes), .n_refused(),
         .n_reads(), .n_aborts(dut_aborts), .n_sda_conflict(dut_conflict));

      logic [2:0] fault = F_NONE;

      // ---- PLACING THE FAULT, TIMED BY THE MONITOR -----------------------------
      //
      // Two lessons are built into this block, both learned from injections that looked
      // successful and proved nothing.
      //
      // FIRST: `F_EXTRA_START` fires as soon as it sees SCL high with SDA high, and an
      // IDLE bus already satisfies that. Armed before a transfer it fired into the gap
      // BEFORE the controller's own START, produced a one-byte transaction of its own, and
      // left the real transfer untouched. The injection check passed; nothing about
      // mid-transfer recovery had been tested.
      //
      // SECOND: armed at the first observed byte, the glitch landed in the ACKNOWLEDGE
      // slot -- where the target is pulling SDA low anyway and is not sampling it. Pulling
      // an already-low line changes nothing, so the glitch was harmless for a reason that
      // has nothing to do with its width. Widening it from 2 clocks to 60 in a mutation
      // run made no difference at all, which is how the vacuity was found.
      //
      // The conclusion is not a smarter injector -- that would mean giving it its own
      // framing detector, duplicating Chapter 20.7's monitor inside a component meant to
      // be thin. The ENVIRONMENT decides WHERE the fault goes and the injector decides
      // WHAT it is, and the environment places it using the MONITOR's observed byte count
      // plus an explicit offset in clocks. Position is part of the stimulus, and a fault
      // whose position is not controlled is not a controlled experiment.
      integer arm_after_bytes = 0;    // arm once this many bytes have been OBSERVED
      integer arm_delay       = 0;    // ... plus this many clocks, to place it within a bit
      logic   arm_en          = 1'b0; // the narrow injector
      logic   armw_en         = 1'b0; // the wide one
      integer bytes_seen, delay_cnt;
      logic   farm_r, farmw_r;
      wire    placed = (bytes_seen >= arm_after_bytes) && (delay_cnt >= arm_delay);
      wire    farm   = farm_r;
      wire    farmw  = farmw_r;

      wire [15:0] f_ninj;
      wire        f_injecting;

      i2c_fault_inj #(.GLITCH_CLKS(2)) inj (
         .clk(clk), .rst_n(rst_n), .fault(fault), .arm(farm),
         .scl_drive_low(f_scl_low), .sda_drive_low(f_sda_low),
         .scl_in(scl), .sda_in(sda),
         .n_injected(f_ninj), .injecting(f_injecting));

      wire [15:0] w_ninj;
      wire        w_injecting;

      // The same module, a wider glitch. Nothing else differs, which is what makes the
      // comparison between T2 and T2b a measurement of width rather than of luck.
      i2c_fault_inj #(.GLITCH_CLKS(40)) injw (
         .clk(clk), .rst_n(rst_n), .fault(fault), .arm(farmw),
         .scl_drive_low(w_scl_low), .sda_drive_low(w_sda_low),
         .scl_in(scl), .sda_in(sda),
         .n_injected(w_ninj), .injecting(w_injecting));

      // ---- THE INDEPENDENT WITNESS ---------------------------------------------
      // The monitor has no connection to either injector. When it reports a repeated
      // START, that report can only have come from the resolved wire -- which is what
      // makes it evidence that a fault was injected rather than merely requested.
      wire m_start, m_restart, m_stop, m_bvalid, m_backed, m_bisaddr;
      wire [7:0] m_bdata;
      wire m_active, m_read, m_aacked, m_done, m_rs;
      wire [6:0] m_addr;
      wire [3:0] m_ndata;
      wire [15:0] m_ntxn, m_nbytes, m_nnacks;

      i2c_mon #(.MAX_BYTES(8)) mon (
         .clk(clk), .rst_n(rst_n), .scl(scl), .sda(sda),
         .saw_start(m_start), .saw_restart(m_restart), .saw_stop(m_stop),
         .byte_valid(m_bvalid), .byte_data(m_bdata), .byte_acked(m_backed),
         .byte_is_addr(m_bisaddr),
         .txn_active(m_active), .txn_addr(m_addr), .txn_read(m_read),
         .txn_addr_acked(m_aacked), .txn_n_data(m_ndata),
         .txn_done(m_done), .txn_ended_by_restart(m_rs),
         .n_txns(m_ntxn), .n_bytes(m_nbytes), .n_nacks(m_nnacks));

      always @(posedge clk or negedge rst_n) begin
         if (!rst_n) begin
            bytes_seen <= 0;
            delay_cnt  <= 0;
            farm_r     <= 1'b0;
            farmw_r    <= 1'b0;
         end else begin
            if (m_bvalid) bytes_seen <= bytes_seen + 1;
            if (bytes_seen >= arm_after_bytes && delay_cnt < arm_delay)
               delay_cnt <= delay_cnt + 1;
            farm_r  <= arm_en  && placed;
            farmw_r <= armw_en && placed;
         end
      end

      integer errors = 0, n;
      integer nrs = 0, sda_low_cycles = 0;

      // WHAT THE WIRE ACTUALLY CARRIED. Latched from the monitor, not from the driver's
      // request -- the distinction the whole module rests on, and the one that makes T2b
      // below a real measurement instead of a restatement of the stimulus.
      logic [7:0] mon_last_data;

      always @(posedge clk or negedge rst_n) begin
         if (!rst_n) begin
            stall_left  <= 0;
            stall_fired <= 1'b0;
         end else begin
            if (stall_arm && m_bvalid && m_bisaddr)  stall_left <= 300;
            else if (stall_left > 0)                 stall_left <= stall_left - 1;
            // Latched, because by the time the checks run the window has closed. Checking a
            // live signal after the event it describes is how T5b first reported that no
            // stall had been requested during a transfer that had visibly been stalled.
            if (stall_left > 0 && dut_stretching)    stall_fired <= 1'b1;
         end
      end

      always @(posedge clk) if (rst_n) begin
         if (m_restart) nrs <= nrs + 1;
         if (!sda) sda_low_cycles <= sda_low_cycles + 1;
         if (m_bvalid && !m_bisaddr) mon_last_data <= m_bdata;
      end

      always #5 clk = ~clk;
      task step; begin @(posedge clk); @(negedge clk); end endtask

      task do_reset;
         begin
            @(negedge clk); rst_n = 1'b0; req = 1'b0; stall_arm = 1'b0;
            arm_en = 1'b0; armw_en = 1'b0; fault = F_NONE;
            arm_after_bytes = 0; arm_delay = 0;
            step; step; step;
            @(negedge clk); rst_n = 1'b1;
            for (n = 0; n < 40; n = n + 1) step;
            nrs = 0; sda_low_cycles = 0; mon_last_data = 8'h00;
         end
      endtask

      task run_txn;
         begin
            dc_before = c_donecnt;
            @(negedge clk); req = 1'b1;
            @(negedge clk); req = 1'b0;
            n = 0;
            while (c_donecnt == dc_before && n < 400000) begin @(posedge clk); n = n + 1; end
            if (c_donecnt == dc_before) begin
               $display("  FAIL the controller BFM never completed a transaction");
               errors = errors + 1;
            end
            for (n = 0; n < 8; n = n + 1) step;
         end
      endtask

      task ck (input [200*8:1] what, input integer g, input integer e);
         begin
            if (g !== e) begin
               $display("  FAIL %0s: got %0d expected %0d", what, g, e);
               errors = errors + 1;
            end
         end
      endtask

      initial begin
         #8000000;
         $display("  FAIL watchdog: the fault environment did not finish");
         $display("=== i2c_fault: 1 CHECK(S) FAILED ===");
         $finish;
      end

      initial begin
         $display("=== i2c_fault: faults on the wire, not flags in the DUT ===");

         // ----------------------------------------------------------------
         // T1. THE BASELINE. Disarmed, the injector must be invisible: a clean transfer
         //     succeeds. Without this, a later "the fault broke it" result could equally
         //     mean the injector breaks everything all the time.
         // ----------------------------------------------------------------
         do_reset;
         fault = F_NONE; arm_en = 1'b0; armw_en = 1'b0;
         req_addr = ADDR; req_read = 1'b0; req_len = 3'd2;
         req_d0 = 8'h00; req_d1 = 8'h5A;
         run_txn;
         $display("T1  disarmed, the injector is invisible and the transfer succeeds");
         ck("T1 the address was acked",      c_aacked, 1);
         ck("T1 both data bytes acked",      c_nacked, 2);
         ck("T1 neither injector injected anything", f_ninj + w_ninj, 0);
         ck("T1 and neither holds a line",
            f_scl_low | f_sda_low | w_scl_low | w_sda_low, 0);
         ck("T1 the byte landed",            dut_regs[0*8 +: 8], 8'h5A);

         // ----------------------------------------------------------------
         // T1b. ARMED, BUT WITH NO FAULT SELECTED. Distinct from T1: there the injector
         //      was disarmed and its arm-low branch released the lines. Here it is fully
         //      armed and running its fault decode with `F_NONE` selected, so this is the
         //      only test that exercises that path -- and an injector that held a line in
         //      its no-fault case would corrupt every transfer in a real environment while
         //      looking, from its own configuration, like it was doing nothing.
         // ----------------------------------------------------------------
         do_reset;
         fault = F_NONE;
         @(negedge clk); arm_en = 1'b1;
         for (n = 0; n < 20; n = n + 1) step;
         $display("T1b armed with no fault selected changes nothing on the bus");
         ck("T1b the bus is still idle high",  scl & sda, 1);
         ck("T1b nothing was injected",        f_ninj, 0);
         ck("T1b and no hold is claimed",      f_injecting, 0);
         req_addr = ADDR; req_read = 1'b0; req_len = 3'd2;
         req_d0 = 8'h07; req_d1 = 8'h2D;
         run_txn;
         ck("T1b a transfer still succeeds",   c_aacked, 1);
         ck("T1b and its byte landed",         dut_regs[7*8 +: 8], 8'h2D);
         @(negedge clk); arm_en = 1'b0;

         // ----------------------------------------------------------------
         // T2. A NARROW GLITCH ON A LINE THAT IS ACTUALLY HIGH, AT A POSITION THAT
         //     ACTUALLY MATTERS -- and it is survived.
         //
         //     Everything about the placement is deliberate. The data byte is 0xFF, so
         //     every bit is HIGH and pulling SDA low genuinely changes the line. The fault
         //     is armed after two OBSERVED bytes (address, then pointer) and offset 60
         //     clocks, which lands it inside the data byte rather than in an acknowledge
         //     slot where the target drives SDA low itself and samples nothing.
         //
         //     Only then does "the glitch was survived" mean anything: 2 clocks at HALF=16
         //     does not span the rising edge at which Module 18 samples SDA, so it is
         //     invisible to the target. T2b holds the position fixed and changes only the
         //     width, which is how we know that is the reason.
         // ----------------------------------------------------------------
         do_reset;
         fault = F_SDA_GLITCH; arm_after_bytes = 2; arm_delay = 60;
         req_addr = ADDR; req_read = 1'b0; req_len = 3'd2;
         req_d0 = 8'h01; req_d1 = 8'hFF;
         @(negedge clk); arm_en = 1'b1;
         run_txn;
         @(negedge clk); arm_en = 1'b0;
         $display("T2  a narrow glitch reaches the bus inside a data bit and is survived");
         ck("T2 the injector fired exactly once",  f_ninj, 1);
         ck("T2 the RESOLVED SDA really went low", sda_low_cycles > 0, 1);
         ck("T2 the transfer completed",           c_aacked, 1);
         ck("T2 every byte was acknowledged",      c_nacked, 2);
         ck("T2 the monitor observed the byte intact", mon_last_data, 8'hFF);
         ck("T2 and the target stored it intact",  dut_regs[1*8 +: 8], 8'hFF);
         ck("T2 no internal driver conflict",      dut_conflict, 0);

         // ----------------------------------------------------------------
         // T2b. THE SAME FAULT, AT THE SAME POSITION, 40 CLOCKS WIDE INSTEAD OF 2.
         //
         //      This is the control that stops T2 being vacuous. One variable changes --
         //      width -- so the different outcome can only be caused by width, and T2's
         //      pass is attributable to the glitch being narrower than the sampling window
         //      rather than to the glitch never having reached anything.
         //
         //      The outcome is the most instructive result in the chapter. The wide glitch
         //      spans the rising edge on which the target samples bit 7, so 0xFF arrives as
         //      0x7F -- AND EVERY BYTE IS STILL ACKNOWLEDGED. A protocol-layer checker sees
         //      a flawless transfer. A single PASS bit covering "did the transfer work"
         //      reports success on corrupted data. This is Chapter 20.3's layer separation
         //      turning into a concrete missed bug rather than an argument.
         //
         //      Note carefully which two things are compared. The monitor's observed byte
         //      and the target's stored byte AGREE -- and both differ from what the driver
         //      was asked to send. A scoreboard fed the driver's intent as its expected
         //      value would have called this a failure of the DUT. The bus is the authority:
         //      the target stored exactly what arrived, and what arrived was not what was
         //      sent.
         // ----------------------------------------------------------------
         do_reset;
         fault = F_SDA_GLITCH; arm_after_bytes = 2; arm_delay = 60;
         req_addr = ADDR; req_read = 1'b0; req_len = 3'd2;
         req_d0 = 8'h01; req_d1 = 8'hFF;
         @(negedge clk); armw_en = 1'b1;
         run_txn;
         @(negedge clk); armw_en = 1'b0;
         $display("T2b the same fault, wider: silent corruption under a clean acknowledge");
         ck("T2b the wide injector fired once",     w_ninj, 1);
         ck("T2b the protocol layer is clean: address acked", c_aacked, 1);
         ck("T2b the protocol layer is clean: both data bytes acked", c_nacked, 2);
         ck("T2b but the wire did NOT carry the byte that was requested",
            (mon_last_data !== 8'hFF), 1);
         ck("T2b the observed byte is the one bit-7 corruption predicts",
            mon_last_data, 8'h7F);
         ck("T2b and the target stored exactly what the WIRE carried, not what was sent",
            dut_regs[1*8 +: 8], mon_last_data);
         ck("T2b no internal driver conflict",      dut_conflict, 0);

         // ----------------------------------------------------------------
         // T3. AN EXTRA START MID-TRANSFER. The injector pulls SDA LOW while SCL is HIGH,
         //     which every conforming device must read as a START.
         //
         //     PROOF IT REACHED THE BUS: the MONITOR reports a repeated START. That is an
         //     independent witness -- the monitor has no connection to the injector, so a
         //     restart in its report can only have come from the wire.
         //
         //     THE DUT RESPONSE: it must abandon the transfer it was in and reframe. The
         //     byte in flight must NOT land, because no device received a complete byte.
         // ----------------------------------------------------------------
         do_reset;
         // a known value first, so "did not land" is distinguishable from "was zero"
         fault = F_NONE; arm_en = 1'b0;
         req_addr = ADDR; req_read = 1'b0; req_len = 3'd2;
         req_d0 = 8'h03; req_d1 = 8'hE7;
         run_txn;
         ck("T3 precondition: register 3 holds the known value", dut_regs[3*8 +: 8], 8'hE7);
         // now the same write, with an injected START part-way through
         do_reset;
         fault = F_EXTRA_START;
         req_addr = ADDR; req_read = 1'b0; req_len = 3'd2;
         req_d0 = 8'h04; req_d1 = 8'h99;
         // ARMED AFTER THE ADDRESS BYTE, timed by the monitor -- see the note above.
         arm_after_bytes = 1; arm_delay = 0;
         @(negedge clk); arm_en = 1'b1;
         run_txn;
         $display("T3  an injected START is seen by the monitor and reframes the target");
         ck("T3 the injector fired",              f_ninj > 0, 1);
         // A ONE-SHOT FAULT MUST BE ONE-SHOT. Checked while the injector is still ARMED,
         // because disarming releases the lines anyway and would hide a fault that never
         // let go. An injector that kept holding SDA after its single START would leave the
         // bus permanently unframeable while every check above still passed.
         ck("T3 the one-shot fault let go of SDA again", f_injecting, 0);
         ck("T3 and is not still holding the line while armed", f_sda_low, 0);
         ck("T3 the MONITOR independently saw a repeated START", nrs > 0, 1);
         ck("T3 register 4 was not written",      dut_regs[4*8 +: 8], 8'h00);
         ck("T3 no internal driver conflict",     dut_conflict, 0);

         // ----------------------------------------------------------------
         // T4. SDA HELD LOW: THE WEDGED BUS. The injector holds SDA down for the whole
         //     transfer. A STOP requires SDA to RISE while SCL is high, so the bus cannot
         //     be framed at all -- and the correct DUT behaviour is to give up.
         //
         //     PROOF IT REACHED THE BUS: the resolved line is LOW while the injector is
         //     armed and nobody else is driving it.
         //
         //     THE DUT RESPONSE: `n_aborts` increments. The target releases the bus rather
         //     than holding it forever, which is Chapter 18.11's timeout.
         // ----------------------------------------------------------------
         do_reset;
         fault = F_SDA_STUCK;
         @(negedge clk); arm_en = 1'b1;
         for (n = 0; n < 40; n = n + 1) step;
         $display("T4  SDA held low: the bus is unframeable and the target times out");
         ck("T4 the injector is holding the line", f_injecting, 1);
         ck("T4 and recorded the hold once",       f_ninj, 1);
         ck("T4 the RESOLVED SDA is low",          sda, 0);
         ck("T4 while SCL is free",                scl, 1);
         // give the target long enough to reach its IDLE_CYCLES timeout
         req_addr = ADDR; req_read = 1'b0; req_len = 3'd1; req_d0 = 8'h05;
         run_txn;
         for (n = 0; n < 2500; n = n + 1) step;
         ck("T4 the target aborted rather than holding on", dut_aborts > 0, 1);
         @(negedge clk); arm_en = 1'b0;
         for (n = 0; n < 200; n = n + 1) step;
         ck("T4 and the bus recovers once the fault is removed", scl & sda, 1);

         // ----------------------------------------------------------------
         // T5. SCL HELD LOW: THE DRIVER MUST TIME OUT, NOT HANG.
         //
         //     This is the fault that tests the ENVIRONMENT rather than the DUT. The
         //     controller BFM releases SCL and waits for it to rise; with SCL held down
         //     it never will. A driver with an unbounded wait would hang the run here --
         //     which is why Chapter 20.5's driver has STRETCH_TIMEOUT, and why this test
         //     exists to prove the timeout works.
         // ----------------------------------------------------------------
         do_reset;
         fault = F_SCL_STUCK;
         @(negedge clk); arm_en = 1'b1;
         for (n = 0; n < 20; n = n + 1) step;
         ck("T5 the resolved SCL is low", scl, 0);
         ck("T5 and the hold was recorded", f_ninj, 1);
         req_addr = ADDR; req_read = 1'b0; req_len = 3'd1; req_d0 = 8'h06;
         run_txn;
         $display("T5  SCL held low: the driver reports a timeout instead of hanging");
         ck("T5 the driver reported a timeout", c_timeout, 1);
         ck("T5 the injector was holding SCL",  f_injecting, 1);
         @(negedge clk); arm_en = 1'b0;
         for (n = 0; n < 200; n = n + 1) step;
         ck("T5 and the bus recovers", scl & sda, 1);

         // ----------------------------------------------------------------
         // T5b. A LEGAL STRETCH LOOKS EXACTLY LIKE THE T5 FAULT, AND MUST NOT BE TREATED
         //      AS ONE. This test is what stops T5 being vacuous.
         //
         //      In T5 the injector held SCL low and the correct outcome was a TIMEOUT. Here
         //      the real Module 18 target holds SCL low -- Chapter 18.10's stretch, asked
         //      for through `stall_req` -- and the correct outcome is the opposite: the
         //      driver waits, the transfer completes, and nothing is reported as an error.
         //
         //      The observable is identical in both cases. SCL is low, released by the
         //      controller, and not rising. A driver cannot distinguish them by looking at
         //      the line, because there is nothing to distinguish: the ONLY difference is
         //      how long it lasts. That is why a stretch wait must be bounded rather than
         //      conditional -- there is no condition available. And it is why the bound is a
         //      parameter of the environment rather than a property of the protocol: it
         //      encodes how patient this particular test intends to be.
         //
         //      Without this test, T5's timeout would be evidence that the driver gives up
         //      on a held SCL, which is not the property wanted. With it, the pair shows the
         //      driver waiting when the hold is short and giving up when it is not.
         // ----------------------------------------------------------------
         do_reset;
         fault = F_NONE;
         req_addr = ADDR; req_read = 1'b0; req_len = 3'd2;
         req_d0 = 8'h01; req_d1 = 8'h7E;
         @(negedge clk); stall_arm = 1'b1;
         run_txn;
         @(negedge clk); stall_arm = 1'b0;
         $display("T5b the same held SCL, legally: the driver waits and the transfer lands");
         ck("T5b the target really did hold SCL while asked to", stall_fired, 1);
         ck("T5b the driver observed a stretch",      c_stretched, 1);
         ck("T5b and did NOT report a timeout",       c_timeout, 0);
         ck("T5b the address was acknowledged",       c_aacked, 1);
         ck("T5b both data bytes were acknowledged",  c_nacked, 2);
         ck("T5b the byte landed intact",             dut_regs[1*8 +: 8], 8'h7E);
         ck("T5b no injector was involved",           f_ninj + w_ninj, 0);
         ck("T5b and no driver conflict occurred",    dut_conflict, 0);
         for (n = 0; n < 200; n = n + 1) step;
         ck("T5b the bus returned to idle",           scl & sda, 1);

         // ----------------------------------------------------------------
         // T6. NO CORRECT PARTICIPANT EVER DROVE A LINE HIGH. The injector, the driver
         //     and the DUT all have drive-low-only interfaces, so this is structural --
         //     asserted because it is the property that makes the whole environment safe
         //     to point at real hardware.
         // ----------------------------------------------------------------
         do_reset;
         $display("T6  every participant pulls LOW or releases; none drives HIGH");
         ck("T6 the bus idles high with everyone released", scl & sda, 1);
         ck("T6 no SDA driver conflict was ever recorded",  dut_conflict, 0);

         if (errors == 0) $display("=== i2c_fault: ALL CHECKS PASSED ===");
         else             $display("=== i2c_fault: %0d CHECK(S) FAILED ===", errors);
         $finish;
      end

   endmodule

Three of those deserve comment.

T3's evidence is the monitor, not the injector. The monitor has no connection to the injector, so a repeated START in its report can only have come from the wire. An injector's own counter says a fault was requested and fired; an independent witness says it happened.

T4 and T5 hold lines forever, deliberately. That makes the watchdog and the bounded waits load-bearing rather than defensive: an environment without them would hang on its own stimulus, and the correct outcome of "the bus is wedged" is a report, not a stalled simulation.

T5b is what stops T5 being vacuous. The observable is identical: SCL low, released by the controller, not rising. A driver cannot distinguish a legal stretch from a stuck line by looking at the line, because there is nothing to distinguish — the only difference is duration. With T5 alone, the evidence says the driver gives up on a held SCL, which is the opposite of clock-stretch support. The pair is the claim.

7. Proving the Harness Itself

A mutation applied to a file the build never compiles produces a complete, plausible, entirely fictional survivor table. Module 19 recorded eight survivors of which three were exactly that: the harness resolved dependencies from a shared directory, so files mutated in one place were compiled from another.

8. The One Equivalent Mutation, With Its Proof

Of the ninety-three mutations run across this module, one survives, and it is the good kind of survivor.

The injector's one-shot glitch is guarded by a fired flag so that a single arming produces exactly one glitch. Mutation I07 removes the guard. Nothing changes.

The proof is short. While arm is held, cnt is only ever assigned 16'd1 — at cnt == 0 — or cnt + 1 while cnt is below the width. Once it reaches the width it is never reassigned. It returns to zero only in the reset or disarm branches, both of which also clear fired. So at every evaluation of the guard with fired set, cnt is non-zero, control takes the same branch either way, and that branch re-asserts the same values.

The empirical half: under the mutation, the bench's n_injected check still reads exactly one glitch. The guard made no difference to the count, which is what equivalence means.

The error-injection suite that never injected anything

Pitfall — forcing an internal flag and calling it injection
Buggy Code
// An error-injection suite for an I2C target. Twelve tests, all passing.
//
//    task test_framing_error();
//       dut.inject_framing_err = 1'b1;
//       @(posedge clk);
//       ck("error flagged",   dut.err_framing, 1);
//       ck("transfer aborted", dut.state, S_IDLE);
//    endtask
//
// Every test has this shape. What they establish: the target reacts to its own
// error inputs, and its state machine responds to its own error flags.
//
// What they do NOT establish: whether an illegal edge on SDA would ever SET that
// flag. The detection logic -- the part between "something bad happened on the
// wire" and "the design noticed" -- is never executed. It is the part most
// likely to be wrong, and the suite's twelve green tests say nothing about it.
//
// The target ships. In the field, a noisy bus produces framing violations the
// target does not notice, because its detector requires SCL to be high AND low
// simultaneously. Untestable via the force path, which bypasses it entirely.
Pitfall — a glitch test that was invisible for the wrong reason
Buggy Code
// A glitch-tolerance test. The narrow glitch is survived, so the target is
// declared robust to disturbances below its sampling resolution.
//
//    fault = F_SDA_GLITCH;           // 2 clocks wide
//    arm_after_bytes = 1;            // after the address byte
//    arm_delay       = 0;
//    req_d0 = 8'h01; req_d1 = 8'h3C; // write 0x3C to register 1
//    run_txn;
//    ck("the injector fired",   f_ninj, 1);        // passes
//    ck("the byte still landed", reg1, 8'h3C);     // passes
//
// The mutation run is what exposes it: WIDENING the glitch from 2 clocks to 60
// changes nothing. No result moves at all.
//
// Armed at the first observed byte, the glitch lands in the ACKNOWLEDGE SLOT --
// where the target is pulling SDA low itself and is not sampling it. Pulling an
// already-low line changes nothing. The glitch was harmless for a reason that
// has nothing to do with its width, and the test's conclusion does not follow
// from its result.

9. What 20.9 Settled

Each rule has one right home. Universal protocol rules in the monitor, device decisions in the reference model, policy in the driver, and "no participant drives high" enforced by port structure so it cannot be broken. One rule — exactly one driver in the acknowledge slot — cannot be checked on the bus at all, and needs the bus model's holder count.

Configuration is not injection. A flag forced inside the design proves the design responds to its own input and bypasses the detection path entirely. Evidence is a consequence on the resolved wire, and the strongest form is an independent witness with no connection to the injector.

Position is part of the stimulus. The same fault in the acknowledge slot and in a data bit are different experiments, and the first one is harmless for reasons unrelated to the fault's parameters. A fault whose position is uncontrolled is not a controlled experiment.

Silent corruption is real and it is quiet. 0xFF arrives as 0x7F with every byte acknowledged and the framing perfect. A protocol-layer verdict reports success, and only the application layer sees it.

A harness must prove it compiled what it mutated. Manifest-verified verdicts, with the guard itself shown to be non-vacuous. A fake KILLED is more dangerous than a fake SURVIVED, because nobody investigates it.

10. What Module 20 Settled

Nine chapters, and one argument running through all of them: reason from what the bus did, not from what the environment intended.

Each chapter is a consequence of that. Objectives need falsifiers, because a claim with no falsifying observation is verified by any test that touches the subject. A matrix needs an uncovered column, because the features that are implemented, correct and never executed are invisible to every other kind of review. An intent and an observation are different types, because the fields that exist in only one of them — acknowledgement, and how a transfer ended — are the ones that matter. A monitor must reimplement framing rather than borrow it, because an oracle derived from the thing it checks agrees with it in error. An expected value comes from a contract, because the stimulus agrees with itself. Verdicts stay separated by layer, because a transfer can be protocol-perfect and carry the wrong data. And a fault is a bus event with an observable consequence, because a configured fault is a request.

The evidence, all of it labelled:

Six defects were found in the environment itself while building it, which is the result worth carrying forward. Two clock-stretch implementations that were complete, correct and never executed. A req_restart that had been unreachable since it was written. A glitch test that passed for a reason unrelated to its conclusion. A registered prediction that reported errors against a correct design. A responder state machine that could not represent "somebody else owns this slot". None of these was found by a test failing — they were found by writing a matrix row, by trying to write a test that could not be written, and by mutations that would not die.

What comes next builds this architecture in the form a production environment uses. Module 21 — The I²C UVM Agent takes the same components — driver, responder, monitor, reference model, scoreboard — and expresses them as a methodology's reusable classes, with the sequences, configuration and factory overrides that make an environment something a team can share rather than something one person maintains. Every architectural decision argued for here is a decision that methodology has an opinion about, and the arguments are worth having before the vocabulary arrives.

Continue learning