Skip to content
VLSI Mentor

I²C · Module 17

Timeout, Error Handling and Bus Recovery

UM10204 defines no error codes, so the taxonomy is a design decision — and collapsing two failures into one code makes them indistinguishable to software that could have acted on them differently. Builds six codes, separates protocol outcomes from faults, and implements the nine-pulse recovery including an honest account of what it cannot do.

Chapter 17.10 taught the master to notice when the bus disagrees with it. This chapter decides what to do about it — and the first thing to establish is how little of that decision the specification makes for you.

1. The Specification Defines No Error Codes

There is no status register in UM10204 and no list of failure modes. Every name in this block is this design's choice.

So this chapter is a design-decision chapter, not a derivation chapter — with one exception, in §4, where four sentences of specification force the most consequential decision in the block.

2. The Six Codes, and What Separates Each From Its Neighbours

CodeMeansSeparated from its neighbour because
E_ADDR_NACKthe address was not acknowledgednothing is there or the device is busy — 16.4 shows these are indistinguishable, so they share a code honestly
E_DATA_NACKa data byte was not acknowledgedthe device is there and declined this byte — a write-protect pin, or a read-only location
E_ARB_LOSTanother master wonnot an error in the usual sense — the bus worked exactly as designed
E_STRETCH_TOa stretch outlasted this master's boundthe target may still be conforming — §3.1.6 sets no bound
E_SDA_STUCKSDA held low with no transfer in progress§3.1.16's nine-pulse remedy applies
E_SCL_STUCKSCL held low§3.1.16 offers no protocol remedy

2a. Two of these are not faults

Arbitration loss is normal operation. Another master won; the bus behaved exactly as §3.1.8 specifies. The correct response is to retry when the bus is free, not to report a fault to a user. A driver that surfaces it as an error will report errors on a working multi-master board.

A stretch timeout is a decision, not a discovery. §3.1.6 gives stretching no upper bound, so a target that stretches for longer than this master is willing to wait is still conforming. The flag records that the master gave up — and Chapter 12.4's no-assumption rule is why the bound must be a parameter with a stated default rather than a derived constant.

3. A Low Line Is Not a Stuck Line

Both lines are low most of the time during a normal transfer — that is what a transfer is.

So stuck-line detection is qualified by !in_transfer. Without that qualifier the master reports both lines stuck on every byte it sends, which is mutation M4 below and would make the error register useless rather than merely noisy.

4. The One Decision the Specification Forces

Two different lines, two different remedies — and the reason is structural rather than an omission:

Therefore E_SDA_STUCK and E_SCL_STUCK must be separate codes. A master reporting one "bus stuck" code would send its driver to attempt the impossible procedure half the time. That is the single most consequential decision in this block, and it follows directly from those four sentences.

4a. Diagnose before acting, and let SCL decide

The recovery sequencer's first state is a diagnosis, and the test is ordered rather than parallel:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
if SCL is low          -> escalate. No protocol remedy exists.      (§3.1.16)
else if SDA is high    -> nothing is stuck. Do nothing.
else                   -> SDA is stuck and the clock works: pulse.

SCL is tested first because a stuck clock makes the procedure impossible whatever SDA is doing.

4b. Nine is a bound, not a quota

§3.1.16 says the holder "should release it some time within those nine clocks". So stopping early is correct, and the number of pulses actually needed is worth reporting — it distinguishes a device that let go on the third pulse from one that never let go at all.

4c. And the procedure ends with a STOP

A freed SDA with no framing leaves every device believing a transfer is in progress. The line is unstuck and the bus is not idle — a different and equally unusable state.

So recovery pulls SDA low and releases it while SCL is high, manufacturing a STOP. Mutation M8 removes exactly this, and it survived every end-state check in the original bench; see §7.

5. The Recovery Procedure

A flow diagram. A start-recovery request enters a diagnose state. From diagnose three paths leave. If SCL reads low the flow goes to an escalate state, which reports that no protocol remedy exists and requires a hardware reset or power cycle. If SDA reads high the flow goes to a nothing-stuck state, which drives no lines at all. Otherwise the flow enters a pulse loop of three phases: drive SCL low, release SCL, then sample SDA while SCL is high. From the sample phase, if SDA came free the flow proceeds to a stop state that frames the bus, and if nine pulses have been issued without SDA coming free the flow goes to escalate. All paths converge on a done state.start_recoveryfrom the hostDiagnoseSCL first, then SDAEscalateHW reset or powercycleNothing stuckdrive nothing at allPulse looplow, release, sampleFrame with a STOPunstuck is not idleDonerecovered or escalated12
Figure 1 — the recovery sequencer. The diagnosis is first and is ordered: SCL decides, because the pulses are pulses on SCL. Only one of the two stuck-line faults has a protocol remedy, and the other path exists to say so rather than to attempt it.

Three pulses, the holder releases, then a STOP

10 cycles
Ten intervals at phase resolution. Recovery drives SCL low and releases it three times. During the first two high phases SDA reads low, because the holder is still holding it. During the third high phase SDA reads high: the holder has let go. Recovery then pulls SDA low itself and releases it while SCL is high, which is a stop condition, leaving both lines high and the bus idle.nine-pulse procedurenine-pulse procedureframed — bus idleframed — bus idlepulse 1 — SDA still heldpulse 1 — SDA still heldpulse 3 — SDA came freepulse 3 — SDA came freerecovery pulls SDA lowrecovery pulls SDA lowreleases it: a STOPreleases it: a STOPscl_bussda_bussampled0lowlowlowlowfreefreefreefreefreepulses0112233333t0t1t2t3t4t5t6t7t8t9
Figure 2 — a stuck SDA freed on the third pulse. Each pulse is three phases, because SDA must be sampled while SCL is HIGH: a two-phase pulse samples at the instant of release and reads the value from before the holder let go, costing one extra pulse. Simulation-shaped figure at phase resolution.

6. The Error Manager, in Three Languages

Azvya Education Pvt. Ltd.VLSI Mentor
i2c_err_mgr.sv — six codes, and the one recovery that is possible
   // -----------------------------------------------------------------------------
   // i2c_err_mgr.sv
   // The error taxonomy a master must report, and the recovery it may attempt.
   //
   // UM10204 DEFINES NO ERROR CODES. There is no status register in the specification and
   // no list of failure modes, so every name below is this design's choice -- and the
   // choice matters, because a master that collapses two distinguishable failures into one
   // code makes them indistinguishable to software that could otherwise have acted on
   // them differently.
   //
   // THE SIX, and what separates each from its neighbours:
   //
   //   E_ADDR_NACK   the address was not acknowledged. Nothing is on the bus at that
   //                 address, OR the device is busy -- Chapter 16.4 shows those are
   //                 indistinguishable, so they share a code, honestly.
   //   E_DATA_NACK   a DATA byte was not acknowledged. Different from the above: the
   //                 device is there and declined this byte, which for an EEPROM means a
   //                 write-protect pin or a read-only location (Chapter 16.1 question 4).
   //   E_ARB_LOST    §3.1.8. Another master won. NOT an error in the usual sense -- the
   //                 bus worked exactly as designed -- and the correct response is to
   //                 retry when the bus is free, not to report a fault.
   //   E_STRETCH_TO  a stretch outlasted this master's policy bound. The target may still
   //                 be conforming (§3.1.6 sets no bound), so this reports a decision the
   //                 master made, not a defect it found.
   //   E_SDA_STUCK   SDA is held low with no transfer in progress. §3.1.16's nine-pulse
   //                 remedy applies.
   //   E_SCL_STUCK   SCL is held low. §3.1.16 offers NO protocol remedy, because the nine
   //                 pulses are themselves pulses on SCL.
   //
   // WHY THE LAST TWO MUST BE SEPARATE CODES. Chapter 15.4 established that the remedies
   // are different -- nine pulses for SDA, a hardware reset or a power cycle for SCL --
   // and a master that reported one "bus stuck" code would send its driver to attempt the
   // impossible procedure half the time. That is the single most consequential decision in
   // this block, and it follows from four sentences of specification.
   //
   // RECOVERY IS ATTEMPTED FOR EXACTLY ONE OF THEM. §3.1.16, verbatim: "If the data line
   // (SDA) is stuck LOW, the master should send nine clock pulses." And for a stuck clock:
   // "the preferential procedure is to reset the bus using the HW reset signal". So this
   // block clocks for a stuck SDA and escalates for a stuck SCL, and refuses to clock a
   // line it cannot drive.
   // -----------------------------------------------------------------------------

   module i2c_err_mgr #(
      parameter int N_PULSES = 9,     // §3.1.16: nine, and it is a BOUND not a quota
      parameter int N_HALF   = 4,     // half a recovery clock period, in cycles
      parameter int CNT_W    = 16
   ) (
      input  logic            clk,
      input  logic            rst_n,

      // Reports from the rest of the master.
      input  logic            addr_nack,
      input  logic            data_nack,
      input  logic            arb_lost,
      input  logic            stretch_timeout,
      input  logic            in_transfer,   // a transfer is open, so a low line is normal

      // The lines.
      input  logic            scl_in,
      input  logic            sda_in,

      // Commands.
      input  logic            start_recovery,
      input  logic            clear,

      // Recovery's own bus drive. It owns the lines only while `recovering`.
      output logic            recovering,
      output logic            rec_scl_low,
      output logic            rec_sda_req,
      output logic            rec_sda_bit,

      // The taxonomy. One bit per code, because a master can be in more than one at once
      // and an encoded field would have to choose.
      output logic [5:0]      err,
      output logic            err_valid,
      output logic [CNT_W-1:0] pulses_issued,
      output logic            recovered,     // SDA came free
      output logic            escalate,      // no protocol remedy exists: HW reset or power
      output logic [2:0]      state
   );

      localparam integer E_ADDR_NACK  = 0, E_DATA_NACK = 1, E_ARB_LOST  = 2,
                         E_STRETCH_TO = 3, E_SDA_STUCK = 4, E_SCL_STUCK = 5;

      localparam [2:0] R_IDLE = 3'd0,
                       R_DIAG = 3'd1,   // decide WHICH line is stuck, before acting
                       R_LOW  = 3'd2,   // recovery pulse: drive SCL low
                       R_REL  = 3'd3,   // release SCL
                       R_SMP  = 3'd4,   // sample SDA while SCL is HIGH
                       R_STOP = 3'd5,   // frame the bus before leaving it
                       R_DONE = 3'd6;

      logic [CNT_W-1:0] cnt, pulse_cnt;

      always @(posedge clk or negedge rst_n) begin
         if (!rst_n) begin
            err           <= 6'd0;
            err_valid     <= 1'b0;
            recovering    <= 1'b0;
            rec_scl_low   <= 1'b0;
            rec_sda_req   <= 1'b0;
            rec_sda_bit   <= 1'b1;
            pulses_issued <= {CNT_W{1'b0}};
            recovered     <= 1'b0;
            escalate      <= 1'b0;
            state         <= R_IDLE;
            cnt           <= {CNT_W{1'b0}};
            pulse_cnt     <= {CNT_W{1'b0}};
         end else begin
            err_valid <= 1'b0;

            if (clear) begin
               err       <= 6'd0;
               recovered <= 1'b0;
               escalate  <= 1'b0;
            end else begin
               // ---- the taxonomy, latched as it is reported --------------------
               if (addr_nack)       begin err[E_ADDR_NACK]  <= 1'b1; err_valid <= 1'b1; end
               if (data_nack)       begin err[E_DATA_NACK]  <= 1'b1; err_valid <= 1'b1; end
               if (arb_lost)        begin err[E_ARB_LOST]   <= 1'b1; err_valid <= 1'b1; end
               if (stretch_timeout) begin err[E_STRETCH_TO] <= 1'b1; err_valid <= 1'b1; end
               // A line held low with NO transfer in progress is stuck. The qualifier
               // matters: during a transfer both lines are low most of the time, and a
               // detector without it would report a stuck bus on every byte.
               if (!in_transfer && !recovering) begin
                  if (!sda_in) begin err[E_SDA_STUCK] <= 1'b1; err_valid <= 1'b1; end
                  if (!scl_in) begin err[E_SCL_STUCK] <= 1'b1; err_valid <= 1'b1; end
               end
            end

            // ---- recovery ---------------------------------------------------
            case (state)

               R_IDLE: begin
                  recovering  <= 1'b0;
                  rec_scl_low <= 1'b0;
                  rec_sda_req <= 1'b0;
                  if (start_recovery) begin
                     recovering <= 1'b1;
                     recovered  <= 1'b0;
                     escalate   <= 1'b0;
                     pulse_cnt  <= {CNT_W{1'b0}};
                     cnt        <= {CNT_W{1'b0}};
                     state      <= R_DIAG;
                  end
               end

               R_DIAG: begin
                  // DIAGNOSE BEFORE ACTING, and let SCL decide. A stuck clock makes the
                  // nine-pulse procedure impossible whatever SDA is doing, because the
                  // pulses ARE pulses on SCL -- so this test is ordered, not parallel.
                  if (!scl_in) begin
                     escalate <= 1'b1;        // no protocol remedy exists. §3.1.16.
                     state    <= R_DONE;
                  end else if (sda_in) begin
                     // Nothing is stuck. Running the procedure anyway would clock a bus
                     // that was merely slow, and corrupt a transfer in progress.
                     recovered <= 1'b1;
                     state     <= R_DONE;
                  end else begin
                     rec_scl_low <= 1'b1;
                     cnt         <= {CNT_W{1'b0}};
                     state       <= R_LOW;
                  end
               end

               R_LOW: begin
                  // A three-phase pulse, because SDA must be sampled while SCL is HIGH.
                  // A two-phase pulse samples at the instant of release, which reads the
                  // value from before the holder let go and costs one extra pulse.
                  rec_scl_low <= 1'b1;
                  if (cnt + 1 >= N_HALF) begin
                     rec_scl_low <= 1'b0;
                     cnt         <= {CNT_W{1'b0}};
                     state       <= R_REL;
                  end else cnt <= cnt + 1'b1;
               end

               R_REL: begin
                  // Released. If the clock does not come up, it has failed mid-recovery --
                  // which is a different fault from the one we started on, and the
                  // procedure must abandon rather than keep counting pulses that are not
                  // reaching the wire.
                  if (!scl_in) begin
                     err[E_SCL_STUCK] <= 1'b1;
                     escalate         <= 1'b1;
                     state            <= R_DONE;
                  end else if (cnt + 1 >= N_HALF) begin
                     cnt   <= {CNT_W{1'b0}};
                     state <= R_SMP;
                  end else cnt <= cnt + 1'b1;
               end

               R_SMP: begin
                  pulse_cnt     <= pulse_cnt + 1'b1;
                  pulses_issued <= pulse_cnt + 1'b1;
                  if (sda_in) begin
                     // Free. Nine is a BOUND, not a quota: §3.1.16 says the holder
                     // "should release it some time within those nine clocks", so
                     // stopping early is correct and the count is worth reporting.
                     rec_sda_req <= 1'b1;
                     rec_sda_bit <= 1'b0;      // pull SDA low, to build a STOP
                     cnt         <= {CNT_W{1'b0}};
                     state       <= R_STOP;
                  end else if (pulse_cnt + 1 >= N_PULSES[CNT_W-1:0]) begin
                     escalate <= 1'b1;          // nine were not enough
                     state    <= R_DONE;
                  end else begin
                     rec_scl_low <= 1'b1;
                     cnt         <= {CNT_W{1'b0}};
                     state       <= R_LOW;
                  end
               end

               R_STOP: begin
                  // End with a STOP. A freed SDA with no framing leaves every device on
                  // the bus believing a transfer is in progress -- the line is unstuck and
                  // the bus is not idle, which is a different and equally unusable state.
                  if (cnt + 1 >= N_HALF) begin
                     rec_sda_bit <= 1'b1;      // release SDA while SCL is high: a STOP
                     recovered   <= 1'b1;
                     cnt         <= {CNT_W{1'b0}};
                     state       <= R_DONE;
                  end else cnt <= cnt + 1'b1;
               end

               R_DONE: begin
                  recovering  <= 1'b0;
                  rec_scl_low <= 1'b0;
                  rec_sda_req <= 1'b0;
                  rec_sda_bit <= 1'b1;
                  state       <= R_IDLE;
               end

               default: state <= R_IDLE;
            endcase
         end
      end

   endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
i2c_err_mgr.v — the same design in Verilog-2001
   // -----------------------------------------------------------------------------
   // i2c_err_mgr.sv
   // The error taxonomy a master must report, and the recovery it may attempt.
   //
   // UM10204 DEFINES NO ERROR CODES. There is no status register in the specification and
   // no list of failure modes, so every name below is this design's choice -- and the
   // choice matters, because a master that collapses two distinguishable failures into one
   // code makes them indistinguishable to software that could otherwise have acted on
   // them differently.
   //
   // THE SIX, and what separates each from its neighbours:
   //
   //   E_ADDR_NACK   the address was not acknowledged. Nothing is on the bus at that
   //                 address, OR the device is busy -- Chapter 16.4 shows those are
   //                 indistinguishable, so they share a code, honestly.
   //   E_DATA_NACK   a DATA byte was not acknowledged. Different from the above: the
   //                 device is there and declined this byte, which for an EEPROM means a
   //                 write-protect pin or a read-only location (Chapter 16.1 question 4).
   //   E_ARB_LOST    §3.1.8. Another master won. NOT an error in the usual sense -- the
   //                 bus worked exactly as designed -- and the correct response is to
   //                 retry when the bus is free, not to report a fault.
   //   E_STRETCH_TO  a stretch outlasted this master's policy bound. The target may still
   //                 be conforming (§3.1.6 sets no bound), so this reports a decision the
   //                 master made, not a defect it found.
   //   E_SDA_STUCK   SDA is held low with no transfer in progress. §3.1.16's nine-pulse
   //                 remedy applies.
   //   E_SCL_STUCK   SCL is held low. §3.1.16 offers NO protocol remedy, because the nine
   //                 pulses are themselves pulses on SCL.
   //
   // WHY THE LAST TWO MUST BE SEPARATE CODES. Chapter 15.4 established that the remedies
   // are different -- nine pulses for SDA, a hardware reset or a power cycle for SCL --
   // and a master that reported one "bus stuck" code would send its driver to attempt the
   // impossible procedure half the time. That is the single most consequential decision in
   // this block, and it follows from four sentences of specification.
   //
   // RECOVERY IS ATTEMPTED FOR EXACTLY ONE OF THEM. §3.1.16, verbatim: "If the data line
   // (SDA) is stuck LOW, the master should send nine clock pulses." And for a stuck clock:
   // "the preferential procedure is to reset the bus using the HW reset signal". So this
   // block clocks for a stuck SDA and escalates for a stuck SCL, and refuses to clock a
   // line it cannot drive.
   // -----------------------------------------------------------------------------

   // (Verilog-2001 -- structurally identical to the SystemVerilog above.)
   module i2c_err_mgr #(
      parameter N_PULSES = 9,     // §3.1.16: nine, and it is a BOUND not a quota
      parameter N_HALF   = 4,     // half a recovery clock period, in cycles
      parameter CNT_W    = 16
   ) (
      input  wire            clk,
      input  wire            rst_n,

      // Reports from the rest of the master.
      input  wire            addr_nack,
      input  wire            data_nack,
      input  wire            arb_lost,
      input  wire            stretch_timeout,
      input  wire            in_transfer,   // a transfer is open, so a low line is normal

      // The lines.
      input  wire            scl_in,
      input  wire            sda_in,

      // Commands.
      input  wire            start_recovery,
      input  wire            clear,

      // Recovery's own bus drive. It owns the lines only while `recovering`.
      output reg             recovering,
      output reg             rec_scl_low,
      output reg             rec_sda_req,
      output reg             rec_sda_bit,

      // The taxonomy. One bit per code, because a master can be in more than one at once
      // and an encoded field would have to choose.
      output reg  [5:0]      err,
      output reg             err_valid,
      output reg  [CNT_W-1:0] pulses_issued,
      output reg             recovered,     // SDA came free
      output reg             escalate,      // no protocol remedy exists: HW reset or power
      output reg  [2:0]      state
   );

      localparam integer E_ADDR_NACK  = 0, E_DATA_NACK = 1, E_ARB_LOST  = 2,
                         E_STRETCH_TO = 3, E_SDA_STUCK = 4, E_SCL_STUCK = 5;

      localparam [2:0] R_IDLE = 3'd0,
                       R_DIAG = 3'd1,   // decide WHICH line is stuck, before acting
                       R_LOW  = 3'd2,   // recovery pulse: drive SCL low
                       R_REL  = 3'd3,   // release SCL
                       R_SMP  = 3'd4,   // sample SDA while SCL is HIGH
                       R_STOP = 3'd5,   // frame the bus before leaving it
                       R_DONE = 3'd6;

      reg [CNT_W-1:0] cnt, pulse_cnt;

      always @(posedge clk or negedge rst_n) begin
         if (!rst_n) begin
            err           <= 6'd0;
            err_valid     <= 1'b0;
            recovering    <= 1'b0;
            rec_scl_low   <= 1'b0;
            rec_sda_req   <= 1'b0;
            rec_sda_bit   <= 1'b1;
            pulses_issued <= {CNT_W{1'b0}};
            recovered     <= 1'b0;
            escalate      <= 1'b0;
            state         <= R_IDLE;
            cnt           <= {CNT_W{1'b0}};
            pulse_cnt     <= {CNT_W{1'b0}};
         end else begin
            err_valid <= 1'b0;

            if (clear) begin
               err       <= 6'd0;
               recovered <= 1'b0;
               escalate  <= 1'b0;
            end else begin
               // ---- the taxonomy, latched as it is reported --------------------
               if (addr_nack)       begin err[E_ADDR_NACK]  <= 1'b1; err_valid <= 1'b1; end
               if (data_nack)       begin err[E_DATA_NACK]  <= 1'b1; err_valid <= 1'b1; end
               if (arb_lost)        begin err[E_ARB_LOST]   <= 1'b1; err_valid <= 1'b1; end
               if (stretch_timeout) begin err[E_STRETCH_TO] <= 1'b1; err_valid <= 1'b1; end
               // A line held low with NO transfer in progress is stuck. The qualifier
               // matters: during a transfer both lines are low most of the time, and a
               // detector without it would report a stuck bus on every byte.
               if (!in_transfer && !recovering) begin
                  if (!sda_in) begin err[E_SDA_STUCK] <= 1'b1; err_valid <= 1'b1; end
                  if (!scl_in) begin err[E_SCL_STUCK] <= 1'b1; err_valid <= 1'b1; end
               end
            end

            // ---- recovery ---------------------------------------------------
            case (state)

               R_IDLE: begin
                  recovering  <= 1'b0;
                  rec_scl_low <= 1'b0;
                  rec_sda_req <= 1'b0;
                  if (start_recovery) begin
                     recovering <= 1'b1;
                     recovered  <= 1'b0;
                     escalate   <= 1'b0;
                     pulse_cnt  <= {CNT_W{1'b0}};
                     cnt        <= {CNT_W{1'b0}};
                     state      <= R_DIAG;
                  end
               end

               R_DIAG: begin
                  // DIAGNOSE BEFORE ACTING, and let SCL decide. A stuck clock makes the
                  // nine-pulse procedure impossible whatever SDA is doing, because the
                  // pulses ARE pulses on SCL -- so this test is ordered, not parallel.
                  if (!scl_in) begin
                     escalate <= 1'b1;        // no protocol remedy exists. §3.1.16.
                     state    <= R_DONE;
                  end else if (sda_in) begin
                     // Nothing is stuck. Running the procedure anyway would clock a bus
                     // that was merely slow, and corrupt a transfer in progress.
                     recovered <= 1'b1;
                     state     <= R_DONE;
                  end else begin
                     rec_scl_low <= 1'b1;
                     cnt         <= {CNT_W{1'b0}};
                     state       <= R_LOW;
                  end
               end

               R_LOW: begin
                  // A three-phase pulse, because SDA must be sampled while SCL is HIGH.
                  // A two-phase pulse samples at the instant of release, which reads the
                  // value from before the holder let go and costs one extra pulse.
                  rec_scl_low <= 1'b1;
                  if (cnt + 1 >= N_HALF) begin
                     rec_scl_low <= 1'b0;
                     cnt         <= {CNT_W{1'b0}};
                     state       <= R_REL;
                  end else cnt <= cnt + 1'b1;
               end

               R_REL: begin
                  // Released. If the clock does not come up, it has failed mid-recovery --
                  // which is a different fault from the one we started on, and the
                  // procedure must abandon rather than keep counting pulses that are not
                  // reaching the wire.
                  if (!scl_in) begin
                     err[E_SCL_STUCK] <= 1'b1;
                     escalate         <= 1'b1;
                     state            <= R_DONE;
                  end else if (cnt + 1 >= N_HALF) begin
                     cnt   <= {CNT_W{1'b0}};
                     state <= R_SMP;
                  end else cnt <= cnt + 1'b1;
               end

               R_SMP: begin
                  pulse_cnt     <= pulse_cnt + 1'b1;
                  pulses_issued <= pulse_cnt + 1'b1;
                  if (sda_in) begin
                     // Free. Nine is a BOUND, not a quota: §3.1.16 says the holder
                     // "should release it some time within those nine clocks", so
                     // stopping early is correct and the count is worth reporting.
                     rec_sda_req <= 1'b1;
                     rec_sda_bit <= 1'b0;      // pull SDA low, to build a STOP
                     cnt         <= {CNT_W{1'b0}};
                     state       <= R_STOP;
                  end else if (pulse_cnt + 1 >= N_PULSES[CNT_W-1:0]) begin
                     escalate <= 1'b1;          // nine were not enough
                     state    <= R_DONE;
                  end else begin
                     rec_scl_low <= 1'b1;
                     cnt         <= {CNT_W{1'b0}};
                     state       <= R_LOW;
                  end
               end

               R_STOP: begin
                  // End with a STOP. A freed SDA with no framing leaves every device on
                  // the bus believing a transfer is in progress -- the line is unstuck and
                  // the bus is not idle, which is a different and equally unusable state.
                  if (cnt + 1 >= N_HALF) begin
                     rec_sda_bit <= 1'b1;      // release SDA while SCL is high: a STOP
                     recovered   <= 1'b1;
                     cnt         <= {CNT_W{1'b0}};
                     state       <= R_DONE;
                  end else cnt <= cnt + 1'b1;
               end

               R_DONE: begin
                  recovering  <= 1'b0;
                  rec_scl_low <= 1'b0;
                  rec_sda_req <= 1'b0;
                  rec_sda_bit <= 1'b1;
                  state       <= R_IDLE;
               end

               default: state <= R_IDLE;
            endcase
         end
      end

   endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
i2c_err_mgr.vhd — the same design in VHDL
   -- ---------------------------------------------------------------------------
   -- i2c_err_mgr.vhd
   -- The error taxonomy a master must report, and the recovery it may attempt.
   -- Behavioural twin of i2c_err_mgr.sv / .v.
   --
   -- UM10204 DEFINES NO ERROR CODES. There is no status register in the specification and no
   -- list of failure modes, so every name below is this design's choice -- and the choice
   -- matters, because a master that collapses two distinguishable failures into one code makes
   -- them indistinguishable to software that could otherwise have acted on them differently.
   --
   --   E_ADDR_NACK   the address was not acknowledged. Nothing there, OR the device is busy --
   --                 Chapter 16.4 shows those are indistinguishable, so they share a code.
   --   E_DATA_NACK   a DATA byte was refused. The device is there and declined this byte.
   --   E_ARB_LOST    §3.1.8. Another master won -- NOT an error in the usual sense.
   --   E_STRETCH_TO  a stretch outlasted this master's policy bound. The target may still be
   --                 conforming, so this reports a decision rather than a defect.
   --   E_SDA_STUCK   SDA held low with no transfer in progress. §3.1.16's nine pulses apply.
   --   E_SCL_STUCK   SCL held low. §3.1.16 offers NO protocol remedy, because the nine pulses
   --                 are themselves pulses on SCL.
   --
   -- WHY THE LAST TWO MUST BE SEPARATE CODES. Chapter 15.4 established that the remedies are
   -- different, and a master that reported one "bus stuck" code would send its driver to
   -- attempt the impossible procedure half the time. That is the most consequential decision in
   -- this block, and it follows from four sentences of specification.
   -- ---------------------------------------------------------------------------

   library ieee;
   use ieee.std_logic_1164.all;
   use ieee.numeric_std.all;

   entity i2c_err_mgr is
      generic (
         N_PULSES : integer := 9;   -- §3.1.16: nine, and it is a BOUND not a quota
         N_HALF   : integer := 4;   -- half a recovery clock period, in cycles
         CNT_W    : integer := 16
      );
      port (
         clk   : in std_logic;
         rst_n : in std_logic;

         addr_nack       : in std_logic;
         data_nack       : in std_logic;
         arb_lost        : in std_logic;
         stretch_timeout : in std_logic;
         in_transfer     : in std_logic;

         scl_in : in std_logic;
         sda_in : in std_logic;

         start_recovery : in std_logic;
         clear          : in std_logic;

         recovering  : out std_logic;
         rec_scl_low : out std_logic;
         rec_sda_req : out std_logic;
         rec_sda_bit : out std_logic;

         err           : out std_logic_vector(5 downto 0);
         err_valid     : out std_logic;
         pulses_issued : out unsigned(CNT_W-1 downto 0);
         recovered     : out std_logic;
         escalate      : out std_logic;
         state         : out unsigned(2 downto 0)
      );
   end entity i2c_err_mgr;

   architecture rtl of i2c_err_mgr is

      constant E_ADDR_NACK  : integer := 0;
      constant E_DATA_NACK  : integer := 1;
      constant E_ARB_LOST   : integer := 2;
      constant E_STRETCH_TO : integer := 3;
      constant E_SDA_STUCK  : integer := 4;
      constant E_SCL_STUCK  : integer := 5;

      constant R_IDLE : integer := 0;
      constant R_DIAG : integer := 1;   -- decide WHICH line is stuck, before acting
      constant R_LOW  : integer := 2;
      constant R_REL  : integer := 3;
      constant R_SMP  : integer := 4;   -- sample SDA while SCL is HIGH
      constant R_STOP : integer := 5;
      constant R_DONE : integer := 6;

      signal st        : integer range 0 to 6 := R_IDLE;
      signal cnt       : integer range 0 to 65535 := 0;
      signal pulse_cnt : unsigned(CNT_W-1 downto 0) := (others => '0');
      signal err_i     : std_logic_vector(5 downto 0) := (others => '0');
      signal rec_i, scl_i, req_i, bit_i, rcv_i, esc_i : std_logic := '0';
      signal np        : unsigned(CNT_W-1 downto 0) := (others => '0');

   begin

      recovering    <= rec_i;
      rec_scl_low   <= scl_i;
      rec_sda_req   <= req_i;
      rec_sda_bit   <= bit_i;
      err           <= err_i;
      pulses_issued <= np;
      recovered     <= rcv_i;
      escalate      <= esc_i;
      state         <= to_unsigned(st, 3);

      process (clk, rst_n)
      begin
         if rst_n = '0' then
            err_i     <= (others => '0');
            err_valid <= '0';
            rec_i     <= '0';
            scl_i     <= '0';
            req_i     <= '0';
            bit_i     <= '1';
            np        <= (others => '0');
            rcv_i     <= '0';
            esc_i     <= '0';
            st        <= R_IDLE;
            cnt       <= 0;
            pulse_cnt <= (others => '0');
         elsif rising_edge(clk) then
            err_valid <= '0';

            if clear = '1' then
               err_i <= (others => '0');
               rcv_i <= '0';
               esc_i <= '0';
            else
               -- ---- the taxonomy, latched as it is reported --------------------
               if addr_nack = '1' then
                  err_i(E_ADDR_NACK) <= '1'; err_valid <= '1';
               end if;
               if data_nack = '1' then
                  err_i(E_DATA_NACK) <= '1'; err_valid <= '1';
               end if;
               if arb_lost = '1' then
                  err_i(E_ARB_LOST) <= '1'; err_valid <= '1';
               end if;
               if stretch_timeout = '1' then
                  err_i(E_STRETCH_TO) <= '1'; err_valid <= '1';
               end if;
               -- A line held low with NO transfer in progress is stuck. The qualifier matters:
               -- during a transfer both lines are low most of the time, and a detector without
               -- it would report a stuck bus on every byte.
               if in_transfer = '0' and rec_i = '0' then
                  if sda_in = '0' then err_i(E_SDA_STUCK) <= '1'; err_valid <= '1'; end if;
                  if scl_in = '0' then err_i(E_SCL_STUCK) <= '1'; err_valid <= '1'; end if;
               end if;
            end if;

            -- ---- recovery ---------------------------------------------------
            case st is

               when R_IDLE =>
                  rec_i <= '0';
                  scl_i <= '0';
                  req_i <= '0';
                  if start_recovery = '1' then
                     rec_i     <= '1';
                     rcv_i     <= '0';
                     esc_i     <= '0';
                     pulse_cnt <= (others => '0');
                     cnt       <= 0;
                     st        <= R_DIAG;
                  end if;

               when R_DIAG =>
                  -- DIAGNOSE BEFORE ACTING, and let SCL decide. A stuck clock makes the
                  -- nine-pulse procedure impossible whatever SDA is doing, because the pulses
                  -- ARE pulses on SCL -- so this test is ordered, not parallel.
                  if scl_in = '0' then
                     esc_i <= '1';           -- no protocol remedy exists. §3.1.16.
                     st    <= R_DONE;
                  elsif sda_in = '1' then
                     -- Nothing is stuck. Running the procedure anyway would clock a bus that
                     -- was merely slow, and corrupt a transfer in progress.
                     rcv_i <= '1';
                     st    <= R_DONE;
                  else
                     scl_i <= '1';
                     cnt   <= 0;
                     st    <= R_LOW;
                  end if;

               when R_LOW =>
                  -- A three-phase pulse, because SDA must be sampled while SCL is HIGH. A
                  -- two-phase pulse samples at the instant of release, which reads the value
                  -- from before the holder let go and costs one extra pulse.
                  scl_i <= '1';
                  if cnt + 1 >= N_HALF then
                     scl_i <= '0';
                     cnt   <= 0;
                     st    <= R_REL;
                  else
                     cnt <= cnt + 1;
                  end if;

               when R_REL =>
                  -- Released. If the clock does not come up it has failed mid-recovery, which
                  -- is a different fault from the one we started on, and the procedure must
                  -- abandon rather than keep counting pulses that are not reaching the wire.
                  if scl_in = '0' then
                     err_i(E_SCL_STUCK) <= '1';
                     esc_i <= '1';
                     st    <= R_DONE;
                  elsif cnt + 1 >= N_HALF then
                     cnt <= 0;
                     st  <= R_SMP;
                  else
                     cnt <= cnt + 1;
                  end if;

               when R_SMP =>
                  pulse_cnt <= pulse_cnt + 1;
                  np        <= pulse_cnt + 1;
                  if sda_in = '1' then
                     -- Free. Nine is a BOUND, not a quota: §3.1.16 says the holder "should
                     -- release it some time within those nine clocks".
                     req_i <= '1';
                     bit_i <= '0';           -- pull SDA low, to build a STOP
                     cnt   <= 0;
                     st    <= R_STOP;
                  elsif (pulse_cnt + 1) >= to_unsigned(N_PULSES, CNT_W) then
                     esc_i <= '1';           -- nine were not enough
                     st    <= R_DONE;
                  else
                     scl_i <= '1';
                     cnt   <= 0;
                     st    <= R_LOW;
                  end if;

               when R_STOP =>
                  -- End with a STOP. A freed SDA with no framing leaves every device believing
                  -- a transfer is in progress -- unstuck, and not idle, which is a different
                  -- and equally unusable state.
                  if cnt + 1 >= N_HALF then
                     bit_i <= '1';           -- release SDA while SCL is high: a STOP
                     rcv_i <= '1';
                     cnt   <= 0;
                     st    <= R_DONE;
                  else
                     cnt <= cnt + 1;
                  end if;

               when R_DONE =>
                  rec_i <= '0';
                  scl_i <= '0';
                  req_i <= '0';
                  bit_i <= '1';
                  st    <= R_IDLE;

               when others =>
                  st <= R_IDLE;

            end case;
         end if;
      end process;

   end architecture rtl;

6a. The testbenches

Twelve tests, and note how many of them assert that something does not happen. This block's dangerous failures are all acting when it should not.

#TestProperty
T1the four reported failures are four codesnot one "error" bit
T2arbitration loss and stretch timeout are separate, and neither is a fault§2a
T3a low line during a transfer is not a stuck lineboth lines are low most of the time
T4with no transfer open, the same lines are stuckand the two codes are distinct
T5a stuck SDA: nine pulses, holder lets go on the thirdplus the pulse-width check
T6the procedure ends with a STOPchecked on the wire; see §7
T7a stuck SDA that never lets go: exactly nine, then escalation
T8the central test, and it asserts a zeroa stuck SCL gets no pulses at all
T9diagnose before acting, and SCL decideswith both stuck, the clock wins
T10a clock that fails mid-recoveryabandons rather than counting phantom pulses
T11recovery on a healthy bus drives nothingthe most dangerous case
T12codes clear on requestand a reset manager reports nothing

The bench instantiates a protocol monitor on the resolved lines, because several of these properties are about what reached the wire rather than what state the block ended in.

Azvya Education Pvt. Ltd.VLSI Mentor
i2c_err_mgr_tb.sv — the self-checking testbench
   `timescale 1ns/1ps
   // -----------------------------------------------------------------------------
   // i2c_err_mgr_tb.sv
   // Independent oracle for i2c_err_mgr.
   //
   // Recovery is a fault handler, and every assertion about a fault handler passes
   // vacuously on a bus that never fails. So the bench INJECTS faults, on a real
   // wired-AND bus, by holding lines the way a broken device holds them -- and the
   // central tests are the ones that assert a ZERO: no pulses on a stuck clock, and no
   // procedure at all on a healthy bus.
   // -----------------------------------------------------------------------------
   module i2c_err_mgr_tb;

      localparam integer NP = 9, NH = 4;
      localparam integer E_ADDR = 0, E_DATA = 1, E_ARB = 2, E_STO = 3, E_SDA = 4, E_SCL = 5;
      localparam [2:0] R_IDLE = 3'd0, R_DONE = 3'd6;

      logic clk = 1'b0, rst_n = 1'b0;
      logic addr_nack = 1'b0, data_nack = 1'b0, arb_lost = 1'b0, stretch_to = 1'b0;
      logic in_transfer = 1'b0, start_recovery = 1'b0, clear = 1'b0;

      // The bench as a faulty device: it can hold either line down.
      logic hold_sda = 1'b0, hold_scl = 1'b0;
      // And a device that lets go after a set number of recovery pulses.
      integer release_after;
      integer pulses_seen;

      logic recovering, rec_scl_low, rec_sda_req, rec_sda_bit;
      logic [5:0] err;
      logic err_valid, recovered, escalate;
      logic [15:0] pulses;
      logic [2:0] rstate;

      wire m_sda_low = rec_sda_req ? ~rec_sda_bit : 1'b0;

      logic scl, sda;
      logic [1:0] scl_in, sda_in, scl_rbl, sda_rbl;
      logic [7:0] scl_h, sda_h;

      i2c_line_model #(.N_DEV(2)) bus (
         .scl_drive_low({hold_scl, rec_scl_low}),
         .sda_drive_low({hold_sda, m_sda_low}),
         .scl(scl), .sda(sda), .scl_in(scl_in), .sda_in(sda_in),
         .scl_released_but_low(scl_rbl), .sda_released_but_low(sda_rbl),
         .scl_holders(scl_h), .sda_holders(sda_h));

      // A protocol monitor on the resolved lines. Without it the bench can check the STATE
      // recovery leaves behind but not what it actually put on the wire -- and "the
      // procedure ends with a STOP" is a statement about the wire. A recovery that frees
      // SDA and never frames the bus leaves both lines high and `recovering` clear, which
      // is indistinguishable from success by any end-state check.
      logic mo_start, mo_stop, mo_bitv, mo_byte, mo_ackv, mo_intr, mo_mid, mo_bit, mo_ack;
      logic [7:0] mo_byteval;
      logic [3:0] mo_bidx;
      logic [15:0] mo_nsta, mo_nsto, mo_nbyte, mo_nmid;
      i2c_proto_mon #(.CNT_W(16)) mon (
         .clk(clk), .rst_n(rst_n), .scl(scl), .sda(sda),
         .start_seen(mo_start), .stop_seen(mo_stop), .bit_seen(mo_bit), .bit_val(mo_bitv),
         .byte_seen(mo_byte), .byte_val(mo_byteval), .ack_seen(mo_ackv), .ack_val(mo_ack),
         .in_transfer(mo_intr), .framing_midbyte(mo_mid), .bit_index(mo_bidx),
         .n_starts(mo_nsta), .n_stops(mo_nsto), .n_bytes(mo_nbyte), .n_midbyte(mo_nmid));

      // Did recovery EVER drive SCL low? `pulses_issued` counts completed pulses, so a
      // procedure that drives the clock and then abandons before sampling reports zero
      // pulses while still having touched a line it must not touch. That is the difference
      // between "issued no pulses" and "drove nothing", and only the second is the property
      // §3.1.16 implies for a stuck clock.
      logic ever_drove_scl = 1'b0;
      always @(posedge clk) begin
         if (!rst_n) ever_drove_scl <= 1'b0;
         else if (rec_scl_low) ever_drove_scl <= 1'b1;
      end

      // The recovery pulses must be LEGAL clock pulses, not merely correctly counted.
      // §3.1.16's nine pulses are ordinary SCL pulses on a real bus, so each phase has to
      // last at least as long as the configured half period -- a recovery clock that runs
      // far faster than the bus's timing is not a remedy, it is noise. Nothing else in this
      // bench measures the pulse WIDTH: the pulse count and the final state are both
      // unchanged by a procedure that clocks too fast.
      integer hi_run = 0;
      integer min_hi = 9999;
      reg     armed  = 1'b0;      // only measure high runs BETWEEN pulses
      always @(posedge clk) begin
         if (!rst_n) begin
            hi_run <= 0; min_hi <= 9999; armed <= 1'b0;
         end else if (recovering) begin
            // The high period before the FIRST pulse is not a pulse -- recovery is entered
            // with the clock already high -- so measurement arms on the first falling edge.
            if (!scl) armed <= 1'b1;
            if (scl) hi_run <= hi_run + 1;
            else begin
               if (armed && hi_run > 0 && hi_run < min_hi) min_hi <= hi_run;
               hi_run <= 0;
            end
         end else begin
            hi_run <= 0; armed <= 1'b0;
         end
      end

      i2c_err_mgr #(.N_PULSES(NP), .N_HALF(NH), .CNT_W(16)) dut (
         .clk(clk), .rst_n(rst_n),
         .addr_nack(addr_nack), .data_nack(data_nack), .arb_lost(arb_lost),
         .stretch_timeout(stretch_to), .in_transfer(in_transfer),
         .scl_in(scl_in[0]), .sda_in(sda_in[0]),
         .start_recovery(start_recovery), .clear(clear),
         .recovering(recovering), .rec_scl_low(rec_scl_low),
         .rec_sda_req(rec_sda_req), .rec_sda_bit(rec_sda_bit),
         .err(err), .err_valid(err_valid), .pulses_issued(pulses),
         .recovered(recovered), .escalate(escalate), .state(rstate));

      always #5 clk = ~clk;

      integer errors = 0;
      integer n, k;

      // Count the recovery pulses that actually reach the WIRE, and let go after a
      // configured number of them -- the way a wedged device eventually does.
      logic scl_l;
      always @(negedge clk) begin
         if (rst_n) begin
            if (scl && !scl_l) begin
               pulses_seen = pulses_seen + 1;
               if (release_after != 0 && pulses_seen >= release_after) hold_sda = 1'b0;
            end
            scl_l = scl;
         end
      end

      task step; begin @(posedge clk); @(negedge clk); end endtask

      task do_reset;
         begin
            @(negedge clk);
            rst_n = 1'b0;
            addr_nack = 1'b0; data_nack = 1'b0; arb_lost = 1'b0; stretch_to = 1'b0;
            in_transfer = 1'b0; start_recovery = 1'b0; clear = 1'b0;
            hold_sda = 1'b0; hold_scl = 1'b0;
            release_after = 0; pulses_seen = 0; scl_l = 1'b1;
            repeat (3) @(posedge clk);
            @(negedge clk); rst_n = 1'b1;
            step;
         end
      endtask

      task pulse_recover;
         begin
            @(negedge clk); start_recovery = 1'b1;
            @(posedge clk); @(negedge clk); start_recovery = 1'b0;
         end
      endtask

      task wait_recovery (input integer max_cycles);
         begin
            n = 0;
            while (!(rstate == R_IDLE && !recovering) && n < max_cycles) begin
               step; n = n + 1;
            end
            if (n >= max_cycles) begin
               $display("  FAIL wait_recovery: stuck in state %0d", rstate);
               errors = errors + 1;
            end
         end
      endtask

      task ck_int (input [200*8:1] what, input integer g, input integer e);
         begin
            if (g !== e) begin
               $display("  FAIL %0s: got %0d expected %0d", what, g, e);
               errors = errors + 1;
            end
         end
      endtask

      task ck_bit (input [200*8:1] what, input g, input e);
         begin
            if (g !== e) begin
               $display("  FAIL %0s: got %0b expected %0b", what, g, e);
               errors = errors + 1;
            end
         end
      endtask

      initial begin
         $display("=== i2c_err_mgr: six codes, and a remedy for exactly one of them ===");

         // ----------------------------------------------------------------
         // T1. THE FOUR REPORTED FAILURES ARE FOUR CODES. An address NACK and a data NACK
         //     are different facts -- nothing there versus there and declining -- and a
         //     master that merged them would make them indistinguishable to software.
         // ----------------------------------------------------------------
         do_reset;
         @(negedge clk); in_transfer = 1'b1; addr_nack = 1'b1;
         @(posedge clk); @(negedge clk); addr_nack = 1'b0; step;
         $display("T1  an address NACK and a data NACK are different codes");
         ck_bit("T1 the address NACK is latched", err[E_ADDR], 1'b1);
         ck_bit("T1 and the data NACK is not", err[E_DATA], 1'b0);
         @(negedge clk); data_nack = 1'b1;
         @(posedge clk); @(negedge clk); data_nack = 1'b0; step;
         ck_bit("T1 now the data NACK is latched too", err[E_DATA], 1'b1);
         ck_bit("T1 and both are visible at once", err[E_ADDR] & err[E_DATA], 1'b1);

         // ----------------------------------------------------------------
         // T2. Arbitration loss and a stretch timeout are also separate, and neither is a
         //     defect: §3.1.8's loser lost a fair contest, and §3.1.6 gives stretching no
         //     bound at all, so a timeout reports this master's policy.
         // ----------------------------------------------------------------
         do_reset;
         @(negedge clk); in_transfer = 1'b1; arb_lost = 1'b1;
         @(posedge clk); @(negedge clk); arb_lost = 1'b0;
         @(negedge clk); stretch_to = 1'b1;
         @(posedge clk); @(negedge clk); stretch_to = 1'b0; step;
         $display("T2  arbitration loss and a stretch timeout are separate, and neither is a fault");
         ck_bit("T2 arbitration loss latched", err[E_ARB], 1'b1);
         ck_bit("T2 stretch timeout latched", err[E_STO], 1'b1);
         ck_int("T2 and only those two", err, (1 << E_ARB) | (1 << E_STO));

         // ----------------------------------------------------------------
         // T3. A LOW LINE DURING A TRANSFER IS NOT A STUCK LINE. Both lines are low most
         //     of the time while clocking, and a detector without the qualifier would
         //     report a stuck bus on every byte.
         // ----------------------------------------------------------------
         do_reset;
         @(negedge clk); in_transfer = 1'b1; hold_sda = 1'b1; hold_scl = 1'b1;
         for (k = 0; k < 10; k = k + 1) step;
         $display("T3  a low line during a transfer is not a stuck line");
         ck_bit("T3 SDA is low", sda, 1'b0);
         ck_bit("T3 SCL is low", scl, 1'b0);
         ck_bit("T3 but no stuck-SDA report", err[E_SDA], 1'b0);
         ck_bit("T3 and no stuck-SCL report", err[E_SCL], 1'b0);

         // ----------------------------------------------------------------
         // T4. WITH NO TRANSFER OPEN, THE SAME LINES ARE STUCK. And the two codes are
         //     separate, which is the block's most consequential decision: §3.1.16 gives
         //     the two lines DIFFERENT remedies, so one merged code would send a driver
         //     to attempt the impossible procedure half the time.
         // ----------------------------------------------------------------
         @(negedge clk); in_transfer = 1'b0;
         for (k = 0; k < 6; k = k + 1) step;
         $display("T4  with no transfer open, the same lines are stuck -- as two codes");
         ck_bit("T4 stuck SDA", err[E_SDA], 1'b1);
         ck_bit("T4 stuck SCL", err[E_SCL], 1'b1);

         // ----------------------------------------------------------------
         // T5. A STUCK SDA: nine pulses, and the holder lets go on the third. §3.1.16 says
         //     the device "should release it some time within those nine clocks", so nine
         //     is a BOUND and stopping early is correct.
         // ----------------------------------------------------------------
         do_reset;
         @(negedge clk); hold_sda = 1'b1; release_after = 3;
         step;
         pulse_recover;
         wait_recovery(1000);
         $display("T5  a stuck SDA is freed by clocking, and nine is a bound not a quota");
         ck_bit("T5 recovered", recovered, 1'b1);
         ck_bit("T5 no escalation needed", escalate, 1'b0);
         ck_int("T5 it took three pulses, not nine", pulses, 3);
         ck_bit("T5 SDA is free", sda, 1'b1);

         // ----------------------------------------------------------------
         // T6. AND THE PROCEDURE ENDS WITH A STOP. A freed SDA with no framing leaves
         //     every device believing a transfer is in progress: unstuck, and not idle.
         // ----------------------------------------------------------------
         $display("T6  recovery ends with a STOP, leaving the bus idle rather than unstuck");
         ck_bit("T6 SCL released", rec_scl_low, 1'b0);
         ck_bit("T6 SDA released", m_sda_low, 1'b0);
         ck_bit("T6 both lines high", scl & sda, 1'b1);
         ck_bit("T6 recovery is finished", recovering, 1'b0);
         // And the STOP must have actually appeared on the bus. The four checks above
         // describe the state recovery left behind; a procedure that freed SDA and never
         // framed the bus leaves exactly the same state, with every device still believing
         // a transfer is open.
         if (mo_nsto < 1) begin
            $display("  FAIL T6 recovery freed SDA but never framed the bus");
            errors = errors + 1;
         end

         // ----------------------------------------------------------------
         // T7. A STUCK SDA THAT NEVER LETS GO: exactly nine pulses, then escalation.
         // ----------------------------------------------------------------
         do_reset;
         @(negedge clk); hold_sda = 1'b1; release_after = 0;   // never releases
         step;
         pulse_recover;
         wait_recovery(2000);
         $display("T7  a holder that never lets go: exactly nine pulses, then escalate");
         ck_int("T7 exactly nine pulses", pulses, NP);
         ck_int("T7 and every high phase was at least N_HALF cycles",
                (min_hi >= NH) ? 1 : 0, 1);
         ck_bit("T7 not recovered", recovered, 1'b0);
         ck_bit("T7 escalated", escalate, 1'b1);

         // ----------------------------------------------------------------
         // T8. THE CENTRAL TEST, AND IT ASSERTS A ZERO. A stuck SCL gets NO pulses,
         //     because the nine pulses ARE pulses on SCL and cannot be issued on a line
         //     something else is holding. §3.1.16 offers no protocol remedy at all.
         // ----------------------------------------------------------------
         do_reset;
         @(negedge clk); hold_scl = 1'b1; hold_sda = 1'b1;
         step;
         pulse_recover;
         wait_recovery(1000);
         $display("T8  a stuck SCL gets no pulses at all, because none could reach the wire");
         ck_int("T8 zero pulses issued", pulses, 0);
         // Stronger than the pulse count: the procedure must not have driven the clock at
         // all. A design that starts pulsing and abandons on the first release also reports
         // zero completed pulses, so the counter alone cannot tell the two apart.
         ck_bit("T8 and recovery never drove SCL at all", ever_drove_scl, 1'b0);
         ck_bit("T8 escalated immediately", escalate, 1'b1);
         ck_bit("T8 not recovered", recovered, 1'b0);
         ck_bit("T8 and the stuck clock was diagnosed", err[E_SCL], 1'b1);

         // ----------------------------------------------------------------
         // T9. DIAGNOSE BEFORE ACTING, and SCL decides. With both lines stuck, the clock
         //     is what determines the outcome -- a stuck clock makes the procedure
         //     impossible whatever SDA is doing.
         // ----------------------------------------------------------------
         $display("T9  with both lines stuck, the clock decides the outcome");
         ck_bit("T9 both were diagnosed", err[E_SDA] & err[E_SCL], 1'b1);
         ck_int("T9 and still no pulses were attempted", pulses, 0);

         // ----------------------------------------------------------------
         // T10. A CLOCK THAT FAILS MID-RECOVERY. The procedure starts legitimately on a
         //      stuck SDA and the clock is then seized. Carrying on would be issuing
         //      pulses that never reach the wire and then reporting nine of them.
         // ----------------------------------------------------------------
         do_reset;
         @(negedge clk); hold_sda = 1'b1; release_after = 0;
         step;
         pulse_recover;
         // Let two pulses happen, then seize the clock.
         n = 0;
         while (pulses_seen < 2 && n < 500) begin step; n = n + 1; end
         // One extra step, so the three language variants stay at the same finish time. The
         // VHDL twin needs it because `pulses_seen` there is a SIGNAL assigned by the counting
         // process and read by the stimulus in the same delta, which yields its previous value
         // and runs the loop once more. Here it is a variable-like `integer` assigned blockingly
         // and visible at once. Keeping the step in all three is cheaper than making the three
         // benches differ.
         step;
         @(negedge clk); hold_scl = 1'b1;
         wait_recovery(2000);
         $display("T10 a clock seized mid-recovery abandons the procedure rather than lying");
         ck_bit("T10 escalated", escalate, 1'b1);
         ck_bit("T10 the stuck clock was recorded", err[E_SCL], 1'b1);
         ck_bit("T10 not recovered", recovered, 1'b0);
         if (pulses >= NP) begin
            $display("  FAIL T10 reported %0d pulses after the clock failed", pulses);
            errors = errors + 1;
         end

         // ----------------------------------------------------------------
         // T11. RECOVERY ON A HEALTHY BUS DRIVES NOTHING. This is the most dangerous case
         //      to get wrong: running the procedure on a bus that was merely slow would
         //      clock a transfer to pieces and then report success.
         // ----------------------------------------------------------------
         do_reset;
         pulse_recover;
         wait_recovery(1000);
         $display("T11 recovery on a healthy bus drives nothing and reports it healthy");
         ck_int("T11 no pulses", pulses, 0);
         ck_bit("T11 reported healthy", recovered, 1'b1);
         ck_bit("T11 no escalation", escalate, 1'b0);
         ck_bit("T11 nothing was driven", rec_scl_low | m_sda_low, 1'b0);
         ck_bit("T11 both lines still high", scl & sda, 1'b1);

         // ----------------------------------------------------------------
         // T12. Codes are cleared on request, and a reset manager reports nothing about a
         //      bus it has not seen.
         // ----------------------------------------------------------------
         do_reset;
         @(negedge clk); in_transfer = 1'b1; addr_nack = 1'b1; arb_lost = 1'b1;
         @(posedge clk); @(negedge clk); addr_nack = 1'b0; arb_lost = 1'b0; step;
         ck_int("T12 two codes latched", err, (1 << E_ADDR) | (1 << E_ARB));
         @(negedge clk); clear = 1'b1; step;
         @(negedge clk); clear = 1'b0; step;
         $display("T12 codes clear on request, and a reset manager reports nothing");
         ck_int("T12 cleared", err, 0);
         @(negedge clk); rst_n = 1'b0; step;
         ck_int("T12 reset clears them too", err, 0);
         ck_int("T12 and the pulse count", pulses, 0);
         ck_bit("T12 nothing driven out of reset", rec_scl_low | rec_sda_req, 1'b0);

         if (errors == 0)
            $display("=== i2c_err_mgr: ALL CHECKS PASSED ===");
         else
            $display("=== i2c_err_mgr: %0d CHECK(S) FAILED ===", errors);
         $finish;
      end

   endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
i2c_err_mgr_tb.v — the same tests in Verilog-2001
   `timescale 1ns/1ps
   // -----------------------------------------------------------------------------
   // i2c_err_mgr_tb.sv
   // Independent oracle for i2c_err_mgr.
   //
   // Recovery is a fault handler, and every assertion about a fault handler passes
   // vacuously on a bus that never fails. So the bench INJECTS faults, on a real
   // wired-AND bus, by holding lines the way a broken device holds them -- and the
   // central tests are the ones that assert a ZERO: no pulses on a stuck clock, and no
   // procedure at all on a healthy bus.
   // -----------------------------------------------------------------------------
   // (Verilog-2001 -- structurally identical to the SystemVerilog above.)
   module i2c_err_mgr_tb;

      localparam integer NP = 9, NH = 4;
      localparam integer E_ADDR = 0, E_DATA = 1, E_ARB = 2, E_STO = 3, E_SDA = 4, E_SCL = 5;
      localparam [2:0] R_IDLE = 3'd0, R_DONE = 3'd6;

      reg clk = 1'b0, rst_n = 1'b0;
      reg addr_nack = 1'b0, data_nack = 1'b0, arb_lost = 1'b0, stretch_to = 1'b0;
      reg in_transfer = 1'b0, start_recovery = 1'b0, clear = 1'b0;

      // The bench as a faulty device: it can hold either line down.
      reg hold_sda = 1'b0, hold_scl = 1'b0;
      // And a device that lets go after a set number of recovery pulses.
      integer release_after;
      integer pulses_seen;

      wire recovering, rec_scl_low, rec_sda_req, rec_sda_bit;
      wire [5:0] err;
      wire err_valid, recovered, escalate;
      wire [15:0] pulses;
      wire [2:0] rstate;

      wire m_sda_low = rec_sda_req ? ~rec_sda_bit : 1'b0;

      wire scl, sda;
      wire [1:0] scl_in, sda_in, scl_rbl, sda_rbl;
      wire [7:0] scl_h, sda_h;

      i2c_line_model #(.N_DEV(2)) bus (
         .scl_drive_low({hold_scl, rec_scl_low}),
         .sda_drive_low({hold_sda, m_sda_low}),
         .scl(scl), .sda(sda), .scl_in(scl_in), .sda_in(sda_in),
         .scl_released_but_low(scl_rbl), .sda_released_but_low(sda_rbl),
         .scl_holders(scl_h), .sda_holders(sda_h));

      // A protocol monitor on the resolved lines. Without it the bench can check the STATE
      // recovery leaves behind but not what it actually put on the wire -- and "the
      // procedure ends with a STOP" is a statement about the wire. A recovery that frees
      // SDA and never frames the bus leaves both lines high and `recovering` clear, which
      // is indistinguishable from success by any end-state check.
      wire mo_start, mo_stop, mo_bitv, mo_byte, mo_ackv, mo_intr, mo_mid, mo_bit, mo_ack;
      wire [7:0] mo_byteval;
      wire [3:0] mo_bidx;
      wire [15:0] mo_nsta, mo_nsto, mo_nbyte, mo_nmid;
      i2c_proto_mon #(.CNT_W(16)) mon (
         .clk(clk), .rst_n(rst_n), .scl(scl), .sda(sda),
         .start_seen(mo_start), .stop_seen(mo_stop), .bit_seen(mo_bit), .bit_val(mo_bitv),
         .byte_seen(mo_byte), .byte_val(mo_byteval), .ack_seen(mo_ackv), .ack_val(mo_ack),
         .in_transfer(mo_intr), .framing_midbyte(mo_mid), .bit_index(mo_bidx),
         .n_starts(mo_nsta), .n_stops(mo_nsto), .n_bytes(mo_nbyte), .n_midbyte(mo_nmid));

      // Did recovery EVER drive SCL low? `pulses_issued` counts completed pulses, so a
      // procedure that drives the clock and then abandons before sampling reports zero
      // pulses while still having touched a line it must not touch. That is the difference
      // between "issued no pulses" and "drove nothing", and only the second is the property
      // §3.1.16 implies for a stuck clock.
      reg ever_drove_scl;
      always @(posedge clk) begin
         if (!rst_n) ever_drove_scl <= 1'b0;
         else if (rec_scl_low) ever_drove_scl <= 1'b1;
      end

      // The recovery pulses must be LEGAL clock pulses, not merely correctly counted.
      // §3.1.16's nine pulses are ordinary SCL pulses on a real bus, so each phase has to
      // last at least as long as the configured half period -- a recovery clock that runs
      // far faster than the bus's timing is not a remedy, it is noise. Nothing else in this
      // bench measures the pulse WIDTH: the pulse count and the final state are both
      // unchanged by a procedure that clocks too fast.
      integer hi_run;
      integer min_hi;
      reg     armed;      // only measure high runs BETWEEN pulses
      always @(posedge clk) begin
         if (!rst_n) begin
            hi_run <= 0; min_hi <= 9999; armed <= 1'b0;
         end else if (recovering) begin
            // The high period before the FIRST pulse is not a pulse -- recovery is entered
            // with the clock already high -- so measurement arms on the first falling edge.
            if (!scl) armed <= 1'b1;
            if (scl) hi_run <= hi_run + 1;
            else begin
               if (armed && hi_run > 0 && hi_run < min_hi) min_hi <= hi_run;
               hi_run <= 0;
            end
         end else begin
            hi_run <= 0; armed <= 1'b0;
         end
      end

      i2c_err_mgr #(.N_PULSES(NP), .N_HALF(NH), .CNT_W(16)) dut (
         .clk(clk), .rst_n(rst_n),
         .addr_nack(addr_nack), .data_nack(data_nack), .arb_lost(arb_lost),
         .stretch_timeout(stretch_to), .in_transfer(in_transfer),
         .scl_in(scl_in[0]), .sda_in(sda_in[0]),
         .start_recovery(start_recovery), .clear(clear),
         .recovering(recovering), .rec_scl_low(rec_scl_low),
         .rec_sda_req(rec_sda_req), .rec_sda_bit(rec_sda_bit),
         .err(err), .err_valid(err_valid), .pulses_issued(pulses),
         .recovered(recovered), .escalate(escalate), .state(rstate));

      always #5 clk = ~clk;

      integer errors = 0;
      integer n, k;

      // Count the recovery pulses that actually reach the WIRE, and let go after a
      // configured number of them -- the way a wedged device eventually does.
      reg scl_l;
      always @(negedge clk) begin
         if (rst_n) begin
            if (scl && !scl_l) begin
               pulses_seen = pulses_seen + 1;
               if (release_after != 0 && pulses_seen >= release_after) hold_sda = 1'b0;
            end
            scl_l = scl;
         end
      end

      task step; begin @(posedge clk); @(negedge clk); end endtask

      task do_reset;
         begin
            @(negedge clk);
            rst_n = 1'b0;
            addr_nack = 1'b0; data_nack = 1'b0; arb_lost = 1'b0; stretch_to = 1'b0;
            in_transfer = 1'b0; start_recovery = 1'b0; clear = 1'b0;
            hold_sda = 1'b0; hold_scl = 1'b0;
            release_after = 0; pulses_seen = 0; scl_l = 1'b1;
            repeat (3) @(posedge clk);
            @(negedge clk); rst_n = 1'b1;
            step;
         end
      endtask

      task pulse_recover;
         begin
            @(negedge clk); start_recovery = 1'b1;
            @(posedge clk); @(negedge clk); start_recovery = 1'b0;
         end
      endtask

      task wait_recovery (input integer max_cycles);
         begin
            n = 0;
            while (!(rstate == R_IDLE && !recovering) && n < max_cycles) begin
               step; n = n + 1;
            end
            if (n >= max_cycles) begin
               $display("  FAIL wait_recovery: stuck in state %0d", rstate);
               errors = errors + 1;
            end
         end
      endtask

      task ck_int (input [200*8:1] what, input integer g, input integer e);
         begin
            if (g !== e) begin
               $display("  FAIL %0s: got %0d expected %0d", what, g, e);
               errors = errors + 1;
            end
         end
      endtask

      task ck_bit (input [200*8:1] what, input g, input e);
         begin
            if (g !== e) begin
               $display("  FAIL %0s: got %0b expected %0b", what, g, e);
               errors = errors + 1;
            end
         end
      endtask

      initial begin
         $display("=== i2c_err_mgr: six codes, and a remedy for exactly one of them ===");

         // ----------------------------------------------------------------
         // T1. THE FOUR REPORTED FAILURES ARE FOUR CODES. An address NACK and a data NACK
         //     are different facts -- nothing there versus there and declining -- and a
         //     master that merged them would make them indistinguishable to software.
         // ----------------------------------------------------------------
         do_reset;
         @(negedge clk); in_transfer = 1'b1; addr_nack = 1'b1;
         @(posedge clk); @(negedge clk); addr_nack = 1'b0; step;
         $display("T1  an address NACK and a data NACK are different codes");
         ck_bit("T1 the address NACK is latched", err[E_ADDR], 1'b1);
         ck_bit("T1 and the data NACK is not", err[E_DATA], 1'b0);
         @(negedge clk); data_nack = 1'b1;
         @(posedge clk); @(negedge clk); data_nack = 1'b0; step;
         ck_bit("T1 now the data NACK is latched too", err[E_DATA], 1'b1);
         ck_bit("T1 and both are visible at once", err[E_ADDR] & err[E_DATA], 1'b1);

         // ----------------------------------------------------------------
         // T2. Arbitration loss and a stretch timeout are also separate, and neither is a
         //     defect: §3.1.8's loser lost a fair contest, and §3.1.6 gives stretching no
         //     bound at all, so a timeout reports this master's policy.
         // ----------------------------------------------------------------
         do_reset;
         @(negedge clk); in_transfer = 1'b1; arb_lost = 1'b1;
         @(posedge clk); @(negedge clk); arb_lost = 1'b0;
         @(negedge clk); stretch_to = 1'b1;
         @(posedge clk); @(negedge clk); stretch_to = 1'b0; step;
         $display("T2  arbitration loss and a stretch timeout are separate, and neither is a fault");
         ck_bit("T2 arbitration loss latched", err[E_ARB], 1'b1);
         ck_bit("T2 stretch timeout latched", err[E_STO], 1'b1);
         ck_int("T2 and only those two", err, (1 << E_ARB) | (1 << E_STO));

         // ----------------------------------------------------------------
         // T3. A LOW LINE DURING A TRANSFER IS NOT A STUCK LINE. Both lines are low most
         //     of the time while clocking, and a detector without the qualifier would
         //     report a stuck bus on every byte.
         // ----------------------------------------------------------------
         do_reset;
         @(negedge clk); in_transfer = 1'b1; hold_sda = 1'b1; hold_scl = 1'b1;
         for (k = 0; k < 10; k = k + 1) step;
         $display("T3  a low line during a transfer is not a stuck line");
         ck_bit("T3 SDA is low", sda, 1'b0);
         ck_bit("T3 SCL is low", scl, 1'b0);
         ck_bit("T3 but no stuck-SDA report", err[E_SDA], 1'b0);
         ck_bit("T3 and no stuck-SCL report", err[E_SCL], 1'b0);

         // ----------------------------------------------------------------
         // T4. WITH NO TRANSFER OPEN, THE SAME LINES ARE STUCK. And the two codes are
         //     separate, which is the block's most consequential decision: §3.1.16 gives
         //     the two lines DIFFERENT remedies, so one merged code would send a driver
         //     to attempt the impossible procedure half the time.
         // ----------------------------------------------------------------
         @(negedge clk); in_transfer = 1'b0;
         for (k = 0; k < 6; k = k + 1) step;
         $display("T4  with no transfer open, the same lines are stuck -- as two codes");
         ck_bit("T4 stuck SDA", err[E_SDA], 1'b1);
         ck_bit("T4 stuck SCL", err[E_SCL], 1'b1);

         // ----------------------------------------------------------------
         // T5. A STUCK SDA: nine pulses, and the holder lets go on the third. §3.1.16 says
         //     the device "should release it some time within those nine clocks", so nine
         //     is a BOUND and stopping early is correct.
         // ----------------------------------------------------------------
         do_reset;
         @(negedge clk); hold_sda = 1'b1; release_after = 3;
         step;
         pulse_recover;
         wait_recovery(1000);
         $display("T5  a stuck SDA is freed by clocking, and nine is a bound not a quota");
         ck_bit("T5 recovered", recovered, 1'b1);
         ck_bit("T5 no escalation needed", escalate, 1'b0);
         ck_int("T5 it took three pulses, not nine", pulses, 3);
         ck_bit("T5 SDA is free", sda, 1'b1);

         // ----------------------------------------------------------------
         // T6. AND THE PROCEDURE ENDS WITH A STOP. A freed SDA with no framing leaves
         //     every device believing a transfer is in progress: unstuck, and not idle.
         // ----------------------------------------------------------------
         $display("T6  recovery ends with a STOP, leaving the bus idle rather than unstuck");
         ck_bit("T6 SCL released", rec_scl_low, 1'b0);
         ck_bit("T6 SDA released", m_sda_low, 1'b0);
         ck_bit("T6 both lines high", scl & sda, 1'b1);
         ck_bit("T6 recovery is finished", recovering, 1'b0);
         // And the STOP must have actually appeared on the bus. The checks above describe
         // the state recovery left behind; a procedure that freed SDA and never framed the
         // bus leaves the same state with every device still believing a transfer is open.
         if (mo_nsto < 1) begin
            $display("  FAIL T6 recovery freed SDA but never framed the bus");
            errors = errors + 1;
         end

         // ----------------------------------------------------------------
         // T7. A STUCK SDA THAT NEVER LETS GO: exactly nine pulses, then escalation.
         // ----------------------------------------------------------------
         do_reset;
         @(negedge clk); hold_sda = 1'b1; release_after = 0;   // never releases
         step;
         pulse_recover;
         wait_recovery(2000);
         $display("T7  a holder that never lets go: exactly nine pulses, then escalate");
         ck_int("T7 exactly nine pulses", pulses, NP);
         ck_int("T7 and every high phase was at least N_HALF cycles",
                (min_hi >= NH) ? 1 : 0, 1);
         ck_bit("T7 not recovered", recovered, 1'b0);
         ck_bit("T7 escalated", escalate, 1'b1);

         // ----------------------------------------------------------------
         // T8. THE CENTRAL TEST, AND IT ASSERTS A ZERO. A stuck SCL gets NO pulses,
         //     because the nine pulses ARE pulses on SCL and cannot be issued on a line
         //     something else is holding. §3.1.16 offers no protocol remedy at all.
         // ----------------------------------------------------------------
         do_reset;
         @(negedge clk); hold_scl = 1'b1; hold_sda = 1'b1;
         step;
         pulse_recover;
         wait_recovery(1000);
         $display("T8  a stuck SCL gets no pulses at all, because none could reach the wire");
         ck_int("T8 zero pulses issued", pulses, 0);
         // Stronger than the pulse count: the procedure must not have driven the clock at
         // all. A design that starts pulsing and abandons on the first release also reports
         // zero completed pulses, so the counter alone cannot tell the two apart.
         ck_bit("T8 and recovery never drove SCL at all", ever_drove_scl, 1'b0);
         ck_bit("T8 escalated immediately", escalate, 1'b1);
         ck_bit("T8 not recovered", recovered, 1'b0);
         ck_bit("T8 and the stuck clock was diagnosed", err[E_SCL], 1'b1);

         // ----------------------------------------------------------------
         // T9. DIAGNOSE BEFORE ACTING, and SCL decides. With both lines stuck, the clock
         //     is what determines the outcome -- a stuck clock makes the procedure
         //     impossible whatever SDA is doing.
         // ----------------------------------------------------------------
         $display("T9  with both lines stuck, the clock decides the outcome");
         ck_bit("T9 both were diagnosed", err[E_SDA] & err[E_SCL], 1'b1);
         ck_int("T9 and still no pulses were attempted", pulses, 0);

         // ----------------------------------------------------------------
         // T10. A CLOCK THAT FAILS MID-RECOVERY. The procedure starts legitimately on a
         //      stuck SDA and the clock is then seized. Carrying on would be issuing
         //      pulses that never reach the wire and then reporting nine of them.
         // ----------------------------------------------------------------
         do_reset;
         @(negedge clk); hold_sda = 1'b1; release_after = 0;
         step;
         pulse_recover;
         // Let two pulses happen, then seize the clock.
         n = 0;
         while (pulses_seen < 2 && n < 500) begin step; n = n + 1; end
         // One extra step, so the three language variants stay at the same finish time. The
         // VHDL twin needs it because `pulses_seen` there is a SIGNAL assigned by the counting
         // process and read by the stimulus in the same delta, which yields its previous value
         // and runs the loop once more. Here it is a variable-like `integer` assigned blockingly
         // and visible at once. Keeping the step in all three is cheaper than making the three
         // benches differ.
         step;
         @(negedge clk); hold_scl = 1'b1;
         wait_recovery(2000);
         $display("T10 a clock seized mid-recovery abandons the procedure rather than lying");
         ck_bit("T10 escalated", escalate, 1'b1);
         ck_bit("T10 the stuck clock was recorded", err[E_SCL], 1'b1);
         ck_bit("T10 not recovered", recovered, 1'b0);
         if (pulses >= NP) begin
            $display("  FAIL T10 reported %0d pulses after the clock failed", pulses);
            errors = errors + 1;
         end

         // ----------------------------------------------------------------
         // T11. RECOVERY ON A HEALTHY BUS DRIVES NOTHING. This is the most dangerous case
         //      to get wrong: running the procedure on a bus that was merely slow would
         //      clock a transfer to pieces and then report success.
         // ----------------------------------------------------------------
         do_reset;
         pulse_recover;
         wait_recovery(1000);
         $display("T11 recovery on a healthy bus drives nothing and reports it healthy");
         ck_int("T11 no pulses", pulses, 0);
         ck_bit("T11 reported healthy", recovered, 1'b1);
         ck_bit("T11 no escalation", escalate, 1'b0);
         ck_bit("T11 nothing was driven", rec_scl_low | m_sda_low, 1'b0);
         ck_bit("T11 both lines still high", scl & sda, 1'b1);

         // ----------------------------------------------------------------
         // T12. Codes are cleared on request, and a reset manager reports nothing about a
         //      bus it has not seen.
         // ----------------------------------------------------------------
         do_reset;
         @(negedge clk); in_transfer = 1'b1; addr_nack = 1'b1; arb_lost = 1'b1;
         @(posedge clk); @(negedge clk); addr_nack = 1'b0; arb_lost = 1'b0; step;
         ck_int("T12 two codes latched", err, (1 << E_ADDR) | (1 << E_ARB));
         @(negedge clk); clear = 1'b1; step;
         @(negedge clk); clear = 1'b0; step;
         $display("T12 codes clear on request, and a reset manager reports nothing");
         ck_int("T12 cleared", err, 0);
         @(negedge clk); rst_n = 1'b0; step;
         ck_int("T12 reset clears them too", err, 0);
         ck_int("T12 and the pulse count", pulses, 0);
         ck_bit("T12 nothing driven out of reset", rec_scl_low | rec_sda_req, 1'b0);

         if (errors == 0)
            $display("=== i2c_err_mgr: ALL CHECKS PASSED ===");
         else
            $display("=== i2c_err_mgr: %0d CHECK(S) FAILED ===", errors);
         $finish;
      end

   endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
i2c_err_mgr_tb.vhd — the same tests in VHDL
   -- ---------------------------------------------------------------------------
   -- i2c_err_mgr_tb.vhd
   -- Independent oracle for i2c_err_mgr. Behavioural twin of the SV and Verilog benches.
   --
   -- Recovery is a fault handler, and every assertion about a fault handler passes vacuously on
   -- a bus that never fails. So the bench INJECTS faults, on a real wired-AND bus, by holding
   -- lines the way a broken device holds them -- and the central tests are the ones that assert
   -- a ZERO: no pulses on a stuck clock, and no procedure at all on a healthy bus.
   -- ---------------------------------------------------------------------------

   library ieee;
   use ieee.std_logic_1164.all;
   use ieee.numeric_std.all;

   entity i2c_err_mgr_tb is
   end entity i2c_err_mgr_tb;

   architecture sim of i2c_err_mgr_tb is

      constant TCLK : time := 10 ns;
      constant NP   : integer := 9;
      constant NH   : integer := 4;
      constant E_ADDR : integer := 0;
      constant E_DATA : integer := 1;
      constant E_ARB  : integer := 2;
      constant E_STO  : integer := 3;
      constant E_SDA  : integer := 4;
      constant E_SCL  : integer := 5;
      constant R_IDLE : integer := 0;

      signal clk, rst_n : std_logic := '0';
      signal addr_nack, data_nack, arb_lost, stretch_to : std_logic := '0';
      signal in_transfer, start_recovery, clear : std_logic := '0';
      signal hold_scl : std_logic := '0';

      -- SINGLE OWNER PER SIGNAL. In the SystemVerilog bench `hold_sda` is written by both the
      -- stimulus block and the pulse counter, which Verilog permits because a `reg` has no
      -- resolution. VHDL's `std_logic` IS a resolved type, so two processes driving it give 'X'
      -- -- silently, with no elaboration error, unlike an `integer` where a second driver is
      -- fatal. The symptom is a fault injector that appears to inject nothing.
      --
      -- So the ownership is split: the stimulus REQUESTS a hold, the pulse counter owns the
      -- line and decides when to let go.
      signal want_hold_sda : std_logic := '0';   -- driven by the stimulus only
      signal hold_sda      : std_logic;          -- driven by the pulse counter only
      signal let_go        : std_logic := '0';

      signal recovering, rec_scl_low, rec_sda_req, rec_sda_bit : std_logic;
      signal err : std_logic_vector(5 downto 0);
      signal err_valid, recovered, escalate : std_logic;
      signal pulses : unsigned(15 downto 0);
      signal rstate : unsigned(2 downto 0);

      signal m_sda_low : std_logic;
      signal scl_drv, sda_drv : std_logic_vector(1 downto 0);
      signal scl, sda : std_logic;
      signal scl_in, sda_in, scl_rbl, sda_rbl : std_logic_vector(1 downto 0);
      signal scl_h, sda_h : unsigned(7 downto 0);

      signal release_after : integer := 0;
      signal pulses_seen   : integer := 0;
      signal halt : boolean := false;

      -- A protocol monitor on the resolved lines. Without it the bench can check the STATE
      -- recovery leaves behind but not what it actually put on the wire -- and "the procedure
      -- ends with a STOP" is a statement about the wire. A recovery that frees SDA and never
      -- frames the bus leaves both lines high and recovering clear, which is
      -- indistinguishable from success by any end-state check.
      signal mo_start, mo_stop, mo_bit, mo_bitv, mo_byte, mo_ackv, mo_ack : std_logic;
      signal mo_intr, mo_mid : std_logic;
      signal mo_byteval : std_logic_vector(7 downto 0);
      signal mo_bidx : unsigned(3 downto 0);
      signal mo_nsta, mo_nsto, mo_nbyte, mo_nmid : unsigned(15 downto 0);

      -- Did recovery EVER drive SCL low? pulses_issued counts COMPLETED pulses, so a
      -- procedure that drives the clock and abandons before sampling reports zero pulses
      -- while still having touched a line it must not touch.
      signal ever_drove_scl : std_logic := '0';
      -- The recovery pulses must be LEGAL clock pulses, not merely correctly counted.
      -- §3.1.16's nine pulses are ordinary SCL pulses on a real bus, so each phase has to
      -- last at least the configured half period. Nothing else in this bench measures the
      -- pulse WIDTH: the count and the final state are unchanged by clocking too fast.
      signal hi_run : integer := 0;
      signal min_hi : integer := 9999;
      signal armed  : boolean := false;
   begin

      m_sda_low <= (not rec_sda_bit) when rec_sda_req = '1' else '0';
      scl_drv   <= hold_scl & rec_scl_low;
      sda_drv   <= hold_sda & m_sda_low;

      bus_m : entity work.i2c_line_model
         generic map (N_DEV => 2)
         port map (scl_drive_low => scl_drv, sda_drive_low => sda_drv,
            scl => scl, sda => sda, scl_in => scl_in, sda_in => sda_in,
            scl_released_but_low => scl_rbl, sda_released_but_low => sda_rbl,
            scl_holders => scl_h, sda_holders => sda_h);

      width_obs : process (clk, rst_n)
      begin
         if rst_n = '0' then
            hi_run <= 0; min_hi <= 9999; armed <= false;
         elsif rising_edge(clk) then
            if recovering = '1' then
               -- The high period before the FIRST pulse is not a pulse, so measurement
               -- arms on the first falling edge.
               if scl = '0' then armed <= true; end if;
               if scl = '1' then
                  hi_run <= hi_run + 1;
               else
                  if armed and hi_run > 0 and hi_run < min_hi then min_hi <= hi_run; end if;
                  hi_run <= 0;
               end if;
            else
               hi_run <= 0; armed <= false;
            end if;
         end if;
      end process;

      mon : entity work.i2c_proto_mon
         generic map (CNT_W => 16)
         port map (clk => clk, rst_n => rst_n, scl => scl, sda => sda,
            start_seen => mo_start, stop_seen => mo_stop, bit_seen => mo_bit,
            bit_val => mo_bitv, byte_seen => mo_byte, byte_val => mo_byteval,
            ack_seen => mo_ackv, ack_val => mo_ack, in_transfer => mo_intr,
            framing_midbyte => mo_mid, bit_index => mo_bidx,
            n_starts => mo_nsta, n_stops => mo_nsto, n_bytes => mo_nbyte,
            n_midbyte => mo_nmid);

      drove_obs : process (clk, rst_n)
      begin
         if rst_n = '0' then
            ever_drove_scl <= '0';
         elsif rising_edge(clk) then
            if rec_scl_low = '1' then ever_drove_scl <= '1'; end if;
         end if;
      end process;

      dut : entity work.i2c_err_mgr
         generic map (N_PULSES => NP, N_HALF => NH, CNT_W => 16)
         port map (clk => clk, rst_n => rst_n,
            addr_nack => addr_nack, data_nack => data_nack, arb_lost => arb_lost,
            stretch_timeout => stretch_to, in_transfer => in_transfer,
            scl_in => scl_in(0), sda_in => sda_in(0),
            start_recovery => start_recovery, clear => clear,
            recovering => recovering, rec_scl_low => rec_scl_low,
            rec_sda_req => rec_sda_req, rec_sda_bit => rec_sda_bit,
            err => err, err_valid => err_valid, pulses_issued => pulses,
            recovered => recovered, escalate => escalate, state => rstate);

      clkgen : process
      begin
         while not halt loop
            clk <= '0'; wait for TCLK/2;
            clk <= '1'; wait for TCLK/2;
         end loop;
         wait;
      end process;

      -- Count the recovery pulses that actually reach the WIRE, and let go after a configured
      -- number of them -- the way a wedged device eventually does.
      hold_sda <= want_hold_sda and (not let_go);

      pulsecnt : process (clk, rst_n)
         variable scl_l : std_logic := '1';
      begin
         if rst_n = '0' then
            scl_l := '1';
            pulses_seen <= 0;
            let_go      <= '0';
         elsif falling_edge(clk) then
            if want_hold_sda = '0' then
               -- A new hold request resets the release decision, so a later test is not
               -- affected by an earlier one having let go.
               let_go      <= '0';
               pulses_seen <= 0;
            elsif scl = '1' and scl_l = '0' then
               pulses_seen <= pulses_seen + 1;
               if release_after /= 0 and (pulses_seen + 1) >= release_after then
                  let_go <= '1';
               end if;
            end if;
            scl_l := scl;
         end if;
      end process;

      stim : process
         variable err_n : integer := 0;
         variable n : integer;

         procedure ck_int (what : string; g : integer; e : integer) is
         begin
            if g /= e then
               report "  FAIL " & what & ": got " & integer'image(g)
                      & " expected " & integer'image(e) severity note;
               err_n := err_n + 1;
            end if;
         end procedure;

         procedure ck_bit (what : string; g : std_logic; e : std_logic) is
         begin
            if g /= e then
               report "  FAIL " & what & ": got " & std_logic'image(g)
                      & " expected " & std_logic'image(e) severity note;
               err_n := err_n + 1;
            end if;
         end procedure;

         procedure step is
         begin
            wait until rising_edge(clk); wait until falling_edge(clk);
         end procedure;

         procedure do_reset is
         begin
            wait until falling_edge(clk);
            rst_n <= '0';
            addr_nack <= '0'; data_nack <= '0'; arb_lost <= '0'; stretch_to <= '0';
            in_transfer <= '0'; start_recovery <= '0'; clear <= '0';
            want_hold_sda <= '0'; hold_scl <= '0'; release_after <= 0;
            for i in 0 to 2 loop wait until rising_edge(clk); end loop;
            wait until falling_edge(clk); rst_n <= '1';
            step;
         end procedure;

         procedure pulse_recover is
         begin
            wait until falling_edge(clk); start_recovery <= '1';
            wait until rising_edge(clk); wait until falling_edge(clk); start_recovery <= '0';
         end procedure;

         procedure wait_recovery (max_cycles : integer) is
         begin
            n := 0;
            while not (to_integer(rstate) = R_IDLE and recovering = '0')
                  and n < max_cycles loop
               step; n := n + 1;
            end loop;
            if n >= max_cycles then
               report "  FAIL wait_recovery: stuck in state "
                      & integer'image(to_integer(rstate)) severity note;
               err_n := err_n + 1;
            end if;
         end procedure;

      begin
         report "=== i2c_err_mgr: six codes, and a remedy for exactly one of them ==="
                severity note;

         -- T1. THE FOUR REPORTED FAILURES ARE FOUR CODES. An address NACK and a data NACK are
         --     different facts, and a master that merged them would make them
         --     indistinguishable to software.
         do_reset;
         wait until falling_edge(clk); in_transfer <= '1'; addr_nack <= '1';
         wait until rising_edge(clk); wait until falling_edge(clk); addr_nack <= '0'; step;
         report "T1  an address NACK and a data NACK are different codes" severity note;
         ck_bit("T1 the address NACK is latched", err(E_ADDR), '1');
         ck_bit("T1 and the data NACK is not", err(E_DATA), '0');
         wait until falling_edge(clk); data_nack <= '1';
         wait until rising_edge(clk); wait until falling_edge(clk); data_nack <= '0'; step;
         ck_bit("T1 now the data NACK is latched too", err(E_DATA), '1');
         ck_bit("T1 and both are visible at once", err(E_ADDR) and err(E_DATA), '1');

         -- T2. Arbitration loss and a stretch timeout are also separate, and neither is a
         --     defect: §3.1.8's loser lost a fair contest, and §3.1.6 gives stretching no bound.
         do_reset;
         wait until falling_edge(clk); in_transfer <= '1'; arb_lost <= '1';
         wait until rising_edge(clk); wait until falling_edge(clk); arb_lost <= '0';
         wait until falling_edge(clk); stretch_to <= '1';
         wait until rising_edge(clk); wait until falling_edge(clk); stretch_to <= '0'; step;
         report "T2  arbitration loss and a stretch timeout are separate, and neither is a fault"
                severity note;
         ck_bit("T2 arbitration loss latched", err(E_ARB), '1');
         ck_bit("T2 stretch timeout latched", err(E_STO), '1');
         ck_int("T2 and only those two", to_integer(unsigned(err)),
                2**E_ARB + 2**E_STO);

         -- T3. A LOW LINE DURING A TRANSFER IS NOT A STUCK LINE. Both lines are low most of
         --     the time while clocking.
         do_reset;
         wait until falling_edge(clk); in_transfer <= '1'; want_hold_sda <= '1'; hold_scl <= '1';
         for j in 0 to 9 loop step; end loop;
         report "T3  a low line during a transfer is not a stuck line" severity note;
         ck_bit("T3 SDA is low", sda, '0');
         ck_bit("T3 SCL is low", scl, '0');
         ck_bit("T3 but no stuck-SDA report", err(E_SDA), '0');
         ck_bit("T3 and no stuck-SCL report", err(E_SCL), '0');

         -- T4. WITH NO TRANSFER OPEN, THE SAME LINES ARE STUCK -- as TWO codes, because
         --     §3.1.16 gives the two lines DIFFERENT remedies.
         wait until falling_edge(clk); in_transfer <= '0';
         for j in 0 to 5 loop step; end loop;
         report "T4  with no transfer open, the same lines are stuck -- as two codes"
                severity note;
         ck_bit("T4 stuck SDA", err(E_SDA), '1');
         ck_bit("T4 stuck SCL", err(E_SCL), '1');

         -- T5. A STUCK SDA: the holder lets go on the third pulse, and nine is a BOUND.
         do_reset;
         wait until falling_edge(clk); want_hold_sda <= '1'; release_after <= 3;
         step;
         pulse_recover;
         wait_recovery(1000);
         report "T5  a stuck SDA is freed by clocking, and nine is a bound not a quota"
                severity note;
         ck_bit("T5 recovered", recovered, '1');
         ck_bit("T5 no escalation needed", escalate, '0');
         ck_int("T5 it took three pulses, not nine", to_integer(pulses), 3);
         ck_bit("T5 SDA is free", sda, '1');

         -- T6. AND THE PROCEDURE ENDS WITH A STOP. A freed SDA with no framing leaves every
         --     device believing a transfer is in progress: unstuck, and not idle.
         report "T6  recovery ends with a STOP, leaving the bus idle rather than unstuck"
                severity note;
         ck_bit("T6 SCL released", rec_scl_low, '0');
         ck_bit("T6 SDA released", m_sda_low, '0');
         ck_bit("T6 both lines high", scl and sda, '1');
         ck_bit("T6 recovery is finished", recovering, '0');
         -- And the STOP must have actually appeared on the bus.
         if to_integer(mo_nsto) < 1 then
            report "  FAIL T6 recovery freed SDA but never framed the bus" severity note;
            err_n := err_n + 1;
         end if;

         -- T7. A STUCK SDA THAT NEVER LETS GO: exactly nine pulses, then escalation.
         do_reset;
         wait until falling_edge(clk); want_hold_sda <= '1'; release_after <= 0;
         step;
         pulse_recover;
         wait_recovery(2000);
         report "T7  a holder that never lets go: exactly nine pulses, then escalate"
                severity note;
         ck_int("T7 exactly nine pulses", to_integer(pulses), NP);
         if min_hi < NH then
            report "  FAIL T7 a recovery high phase was shorter than N_HALF" severity note;
            err_n := err_n + 1;
         end if;
         ck_bit("T7 not recovered", recovered, '0');
         ck_bit("T7 escalated", escalate, '1');

         -- T8. THE CENTRAL TEST, AND IT ASSERTS A ZERO. A stuck SCL gets NO pulses, because
         --     the nine pulses ARE pulses on SCL.
         do_reset;
         wait until falling_edge(clk); hold_scl <= '1'; want_hold_sda <= '1';
         step;
         pulse_recover;
         wait_recovery(1000);
         report "T8  a stuck SCL gets no pulses at all, because none could reach the wire"
                severity note;
         ck_int("T8 zero pulses issued", to_integer(pulses), 0);
         -- Stronger than the pulse count: the procedure must not have driven the clock at
         -- all. A design that starts pulsing and abandons on the first release also reports
         -- zero completed pulses, so the counter alone cannot tell the two apart.
         ck_bit("T8 and recovery never drove SCL at all", ever_drove_scl, '0');
         ck_bit("T8 escalated immediately", escalate, '1');
         ck_bit("T8 not recovered", recovered, '0');
         ck_bit("T8 and the stuck clock was diagnosed", err(E_SCL), '1');

         -- T9. DIAGNOSE BEFORE ACTING, and SCL decides.
         report "T9  with both lines stuck, the clock decides the outcome" severity note;
         ck_bit("T9 both were diagnosed", err(E_SDA) and err(E_SCL), '1');
         ck_int("T9 and still no pulses were attempted", to_integer(pulses), 0);

         -- T10. A CLOCK THAT FAILS MID-RECOVERY abandons the procedure rather than reporting
         --      nine pulses that never reached the wire.
         do_reset;
         wait until falling_edge(clk); want_hold_sda <= '1'; release_after <= 0;
         step;
         pulse_recover;
         n := 0;
         while pulses_seen < 2 and n < 500 loop step; n := n + 1; end loop;
         wait until falling_edge(clk); hold_scl <= '1';
         wait_recovery(2000);
         report "T10 a clock seized mid-recovery abandons the procedure rather than lying"
                severity note;
         ck_bit("T10 escalated", escalate, '1');
         ck_bit("T10 the stuck clock was recorded", err(E_SCL), '1');
         ck_bit("T10 not recovered", recovered, '0');
         if to_integer(pulses) >= NP then
            report "  FAIL T10 reported " & integer'image(to_integer(pulses))
                   & " pulses after the clock failed" severity note;
            err_n := err_n + 1;
         end if;

         -- T11. RECOVERY ON A HEALTHY BUS DRIVES NOTHING -- the most dangerous case to get
         --      wrong, because the procedure would clock a live transfer to pieces and then
         --      report success.
         do_reset;
         pulse_recover;
         wait_recovery(1000);
         report "T11 recovery on a healthy bus drives nothing and reports it healthy"
                severity note;
         ck_int("T11 no pulses", to_integer(pulses), 0);
         ck_bit("T11 reported healthy", recovered, '1');
         ck_bit("T11 no escalation", escalate, '0');
         ck_bit("T11 nothing was driven", rec_scl_low or m_sda_low, '0');
         ck_bit("T11 both lines still high", scl and sda, '1');

         -- T12. Codes clear on request, and a reset manager reports nothing.
         do_reset;
         wait until falling_edge(clk); in_transfer <= '1'; addr_nack <= '1'; arb_lost <= '1';
         wait until rising_edge(clk); wait until falling_edge(clk);
         addr_nack <= '0'; arb_lost <= '0'; step;
         ck_int("T12 two codes latched", to_integer(unsigned(err)),
                2**E_ADDR + 2**E_ARB);
         wait until falling_edge(clk); clear <= '1'; step;
         wait until falling_edge(clk); clear <= '0'; step;
         report "T12 codes clear on request, and a reset manager reports nothing"
                severity note;
         ck_int("T12 cleared", to_integer(unsigned(err)), 0);
         wait until falling_edge(clk); rst_n <= '0'; step;
         ck_int("T12 reset clears them too", to_integer(unsigned(err)), 0);
         ck_int("T12 and the pulse count", to_integer(pulses), 0);
         ck_bit("T12 nothing driven out of reset", rec_scl_low or rec_sda_req, '0');

         if err_n = 0 then
            report "=== i2c_err_mgr: ALL CHECKS PASSED ===" severity note;
         else
            report "=== i2c_err_mgr: " & integer'image(err_n)
                   & " CHECK(S) FAILED ===" severity note;
         end if;
         halt <= true;
         wait;
      end process;

   end architecture sim;

6b. Execution

DesignSystemVerilogVerilog-2001VHDLFinish
i2c_err_mgrPASS 12/12PASS 12/12PASS 12/122390 ns, all three

7. Mutation Testing — Four Survivors, Four Different Reasons

Ten defects. Four survived the original bench, and each failed for a different reason — which together make a fair summary of how end-state checking goes wrong.

#Injected defectExpected detectionResult
M1a stuck SCL is clocked anywayT8 after strengtheningKILLED (2)
M2recovery runs on a healthy busT11KILLED (2)
M3the two stuck-line failures share one codeT4KILLED (4)
M4a low line during a transfer reported as stuckT3KILLED (3)
M5ten pulses instead of nineT7KILLED (2)
M6the pulse's high phase is one cycle instead of N_HALFT7 after strengtheningKILLED (2)
M7a clock failing mid-recovery is ignoredT10KILLED (3)
M8the procedure ends without framing the busT6 after strengtheningKILLED (2)
M9an address NACK and a data NACK share one codeT1KILLED (5)
M10recovery claims success after nine pulses failedT7KILLED (2)
Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
baseline: PASS   (verified before injecting anything)
killed: 10   survived: 0   score: 10/10
restored: PASS

M1 — the counter said zero and the master had still driven the clock

T8 asserted pulses_issued == 0 for a stuck SCL, plus escalation and the right code. With the diagnosis bypassed, the sequencer falls through to the pulse loop, drives SCL low for a half period, then finds the clock still low on release and escalates — so all four of T8's assertions still held. Zero completed pulses, escalated, not recovered, correct code.

The difference was one signal nobody looked at: the master had driven SCL low on a bus it must not touch.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
"issued no pulses"  is not the same claim as  "drove nothing"

pulses_issued counts completed pulses, so a procedure that starts and abandons reports zero. T8 now also asserts a latched "did recovery ever drive SCL" flag, which is the property §3.1.16 actually implies for a stuck clock.

M8 — the end state of success and of failure are identical

Removing the STOP framing left: both lines high, rec_scl_low low, m_sda_low low, recovering clear. That is exactly what a successful recovery leaves behind, because the holder had already let go — so all four of T6's checks passed.

The missing STOP is only visible on the wire. The fix was a protocol monitor and a check that n_stops >= 1, which is why §4c matters: unstuck and idle are different states with the same end-of-recovery snapshot.

Collapsing the high phase to a single cycle left the count at nine, the samples valid, and the final state correct. But §3.1.16's nine pulses are ordinary SCL pulses on a real bus, and a recovery clock running far faster than the bus's timing is not a remedy — it is noise that no device will respond to.

Nothing measured pulse width. The bench now measures the high phase of every recovery pulse and requires at least N_HALF cycles — with the measurement armed only on the first falling edge, because the high period before the first pulse is not a pulse and counting it fails the correct design.

M10 — the invalid mutant, and then the real one

The first attempt changed a state transition in the healthy-bus branch and was inert: the target state cleared the same outputs one cycle later. The real defect is claiming success where the design reports failure — setting recovered alongside escalate after nine pulses have not freed the line — and that dies immediately against T7.

8. Verification Connection — Error Injection and Expected-Error Sequences

Azvya Education Pvt. Ltd.VLSI Mentor
error_env.sv — injecting faults, and the scoreboard's hardest problem
   // This block cannot be verified by stimulating the DUT: every one of its inputs is a
   // report about somebody else's misbehaviour. So the environment needs fault injection
   // as a first-class capability, and the injectors are NOT interchangeable:
   //
   //   hold SDA low, no transfer open      -> E_SDA_STUCK, recovery attempted
   //   hold SCL low, no transfer open      -> E_SCL_STUCK, recovery REFUSED
   //   hold SCL low DURING a transfer      -> a stretch, NOT a stuck line (§3)
   //   release the holder after N pulses   -> the §4b early-exit path, per N
   //   hold SDA forever                    -> nine pulses then escalation
   //   hold SCL low from pulse 2 onward    -> the mid-recovery clock failure (T10)
   //
   // THE SCOREBOARD'S HARDEST PROBLEM is that two of the six codes are NOT failures.
   // A scoreboard written as "any error bit set => test fails" will fail every
   // multi-master test, because E_ARB_LOST is the bus working correctly. So the
   // environment needs EXPECTED-ERROR SEQUENCES: a sequence declares the outcome it
   // intends, and the scoreboard compares against that rather than against zero.
   //
   //     class arb_loss_seq extends i2c_base_seq;
   //        // this sequence EXPECTS E_ARB_LOST and expects the master to retry
   //     endclass
   //
   // Without that, the usual outcome is that somebody masks E_ARB_LOST out of the
   // scoreboard's check -- and then a master that reports arbitration loss on every
   // READ (Chapter 17.4 §9's defect) passes silently forever.
   //
   // WHAT MUST NOT BE CHECKED BY MASKING: the difference between E_SDA_STUCK and
   // E_SCL_STUCK. §4 shows the remedies are different, so a scoreboard that accepts
   // either code for a stuck-line test accepts a master that would send its driver to
   // attempt the impossible procedure.
   //
   // COVERAGE, and one bin here is deliberately unreachable:
   //
   //   cover: recovery succeeded on pulse 1..9          -- nine reachable bins
   //   cover: recovery escalated after nine pulses
   //   cover: recovery refused because SCL was stuck
   //   cover: recovery requested on a HEALTHY bus       -- T11, and it must drive nothing
   //   illegal_bin: pulses issued > 9                   -- the bound, not a quota
   //   illegal_bin: pulses issued on a stuck SCL        -- UNREACHABLE by §4

9. FPGA and ASIC Implications

On an FPGA, recovery is the one place the master drives the bus outside a transaction, and the ownership consequence matters: rec_scl_low and rec_sda_req must be muxed into the same pad paths the rest of the design uses, with recovery as Chapter 17.4's owner 3. Giving recovery its own pad drivers would put two drivers on one pin. The recovering output is the mux select, and it is the reason recovery is a requester rather than a special case in the pad logic.

The diagnosis also depends on the synchronised readback, which means a stuck line must be stuck for at least the synchroniser depth before it is diagnosed. That is the right behaviour — a line low for two cycles is not stuck — and it is another case where the latency is a feature.

On an ASIC, §3.1.16's escalation path is a system integration requirement, not an RTL one. The escalate output must reach something that can actually perform a hardware reset or a power cycle: a reset controller, a PMIC sequencer, or at minimum a status bit that firmware polls before deciding the bus is unusable. A design that reports escalate into a register nobody reads has implemented the diagnosis and none of the remedy.

The pulse width is set by N_HALF, and it should be derived from the same mode parameters as Chapter 17.3's generator rather than chosen independently — recovery pulses are ordinary SCL pulses and are subject to the same Table 10 minimums, which is exactly what mutation M6 exploited. Chapter 17.13 ties both to one source.

10. Debugging — The Recovery That Made Things Worse

Symptom

A gateway board has an I2C bus that occasionally locks up: the SoC reports SDA stuck and the bus stops working until the board is power-cycled. A firmware update adds an automatic recovery routine that issues the master's bus-recovery command whenever any transaction times out. After the update, lockups become MORE frequent, and a new symptom appears -- a temperature sensor on the same bus starts returning corrupt readings during periods of heavy bus traffic, which it never did before.

Root Cause

The recovery procedure was invoked while a transfer was in progress, and a transfer in progress looks exactly like a stuck SDA at the instant the diagnosis samples: SCL high between pulses, SDA low because someone is transmitting a zero. The block correctly suppresses stuck-line detection during a transfer -- that is the in_transfer qualifier -- but the host-commanded recovery path did not apply the same qualifier, so a command issued at the wrong moment produced a diagnosis that was wrong for a reason the diagnosis could not see. The firmware was also at fault for invoking recovery on a per-transaction timeout rather than on a bus-level fault, but the hardware accepted a command it should have refused.

Fix
Refuse a recovery request while a transfer is open: the same in_transfer qualifier that guards detection must guard the procedure, because the diagnosis cannot distinguish a stuck line from a live transfer by looking at the lines. Then fix the firmware policy -- recovery is a response to a bus fault, not to one transaction's timeout, and a timed-out transaction on a healthy bus needs a retry rather than nine pulses. For the regression: request recovery mid-transfer and assert that nothing is driven, which is T11's property applied to a new precondition. T11 as written covers a healthy IDLE bus; the case that bit here is a healthy BUSY one.

Three generalisations.

A correct qualifier was applied to detection and not to action. in_transfer guarded the stuck-line report and not the recovery procedure. The two paths need the same guard for the same reason, and the asymmetry is invisible unless someone asks what the diagnosis can actually see.

The diagnosis is a snapshot of a situation that has history. SCL high and SDA low is a stuck bus or an ordinary transmitted zero between clock pulses. No sampling of the two lines can separate them; only knowing whether a transfer is open can.

Automatic recovery on the wrong trigger is worse than no recovery. The procedure worked perfectly and was invoked wrongly, and a procedure that drives nine pulses into a live transfer corrupts a device that had nothing to do with the fault.

11. Common Misconceptions

"The specification defines I²C error codes." It defines none. Every code in a master's status register is a design decision. §1.

"More error codes is over-engineering." Two failures sharing one code are indistinguishable to software forever. E_SDA_STUCK and E_SCL_STUCK have different remedies, and one code sends a driver to attempt the impossible one half the time. §4.

"An address NACK means nothing is there." It means nothing answered — absent device or busy device, and 16.4 shows those are indistinguishable. They share a code honestly. §2.

"Arbitration loss is an error." The bus worked exactly as §3.1.8 specifies. The response is to retry when the bus is free, not to report a fault. §2a.

"A stretch timeout means the target is broken." §3.1.6 sets no bound, so the target may be perfectly conforming. The flag says the master stopped waiting. §2a.

"A low line means a stuck line." Both lines are low most of the time during a transfer. Detection has to be qualified by whether a transfer is open. §3.

"Nine clock pulses can recover any stuck bus." They are pulses on SCL, so they cannot recover a stuck SCL. §3.1.16 sends you to a hardware reset or a power cycle for that. §4.

"Nine pulses means always send nine." Nine is a bound. §3.1.16 says the holder should release within those nine, so stopping early is correct and the count is worth reporting. §4b.

"Once SDA is free, recovery is done." A freed SDA with no framing leaves every device believing a transfer is open — unstuck and not idle. §4c.

"Zero pulses issued proves the clock was never driven." The counter counts completed pulses. A procedure that starts and abandons reports zero having driven the line. §7.

"If the end state is right, the procedure was right." Three of four survivors here left an end state identical to success. §7.

12. Reason It Through

Why must a stuck SDA and a stuck SCL be different error codes?

Because the remedies are different — nine pulses for SDA, a hardware reset or power cycle for SCL — and the nine pulses are themselves pulses on SCL. One code would send a driver to attempt the impossible procedure half the time. §4.

Why does the diagnosis test SCL before SDA rather than in parallel?

Because a stuck clock makes the procedure impossible whatever SDA is doing. The tests are ordered because one of them is disqualifying. §4a.

Two of the six codes are not faults. Which, and what should a driver do with each?

Arbitration loss — retry when the bus is free, because the bus worked as designed. And a stretch timeout — decide whether to wait longer, because the target may still be conforming. §2a.

Why is stuck-line detection qualified by in_transfer?

Because both lines are low most of the time during a transfer, so unqualified detection reports stuck lines on every byte. §3.

A recovery procedure reports zero pulses issued on a stuck clock. Does that prove it never drove SCL?

No. The counter counts completed pulses, so a procedure that drives the clock low and abandons on the first release reports zero having touched the line. A separate flag is needed for "drove anything at all". §7.

Why is "both lines high, recovery finished" not evidence that recovery framed the bus?

Because a recovery that freed SDA and never issued a STOP leaves exactly that state — the holder had already let go. The STOP is only visible on the wire. §7.

Nine recovery pulses are issued at the correct count and the bus does not recover. What else could be wrong with them?

Their width. Recovery pulses are ordinary SCL pulses subject to Table 10's minimums, and a pulse train that is counted correctly but clocked far too fast is noise no device will respond to. §7.

Why did applying in_transfer to detection but not to the recovery command produce a corrupted, innocent device?

Because the diagnosis cannot tell a stuck SDA from a transmitted zero between clock pulses by looking at the lines. Invoked mid-transfer, it diagnosed a stuck line and injected nine pulses into a live byte. §10.

13. Understanding Check

14. Summary

UM10204 defines no error codes, so the taxonomy is a design decision — and two failures sharing one code are indistinguishable to software forever.

Six codes, and two of them are not faults. Arbitration loss is the bus working as specified; a stretch timeout is the master deciding to stop waiting, not a defect it found.

A low line is not a stuck line. Both lines are low most of the time during a transfer, so detection is qualified by whether a transfer is open.

The one decision the specification forces is that SDA-stuck and SCL-stuck are separate codes, because the nine pulses of §3.1.16 are pulses on SCL and therefore cannot rescue a stuck clock.

So the diagnosis is ordered, not parallel — SCL is tested first, because a stuck clock disqualifies the procedure whatever SDA is doing.

Nine is a bound, not a quota, and the count actually needed distinguishes a device that let go on the third pulse from one that never did.

Recovery ends with a STOP, because a freed SDA with no framing leaves the bus unstuck and not idle.

Ten mutants, ten killed — after four survived, and three of the four passed because the end state after the defect is identical to the end state after success: an unframed bus, a briefly-driven clock, a pulse train that was too fast.

So end-state checks are necessary and not sufficient for a block whose subject is a procedure. What reached the wire, and how long it took, are separate observations needing separate instruments — a protocol monitor, a drove-anything latch, and a pulse-width measurement.

And a correct qualifier applied to detection but not to action is a real defect. Recovery invoked mid-transfer diagnoses a transmitted zero as a stuck line and clocks nine pulses into a live byte, corrupting a device that had nothing to do with the fault.

15. What Comes Next

Every block now exists, is verified, and reports what it did. Chapter 17.12 puts them together — and the point of that chapter is the order in which it does so.

The FSM comes last, and its states are not a description of the protocol. They are the coordination between blocks that already work: who is granted SCL, who owns SDA, which block's completion pulse the next transition waits on. Everything that could be a counter, a shift register or a comparator has already been built and tested, so what remains is genuinely small — which is the payoff for the decomposition Chapter 17.1 argued for, and the reason a master written FSM-first gets rewritten.

Continue learning

Related tutorials