Skip to content
VLSI Mentor

I²C · Module 23

Missing ACK and Address Faults

Seven causes of a missing acknowledge, one diagnostic procedure of at most five probes, and a confusion matrix that shows exactly which two it cannot separate. Built on a fault-injectable target so every verdict is checked against a known cause — including a target that acknowledges correctly and a controller that samples too early to see it.

The missing acknowledge is the first I²C failure most engineers meet and the one they meet most often. A transfer starts, the address byte goes out, the ninth clock arrives, and SDA stays high.

It is also the symptom with the largest number of distinct causes, several of which are not faults at all. This chapter works through them, and it is built entirely out of Chapter 23.1 §6's question: given two candidates, what measurement would give a different result depending on which is true?

1. What Was Actually Observed

The observation is narrower than the sentence people use for it.

Azvya Education Pvt. Ltd.VLSI Mentor
OBSERVATION vs the three interpretations usually folded into it
   OBSERVATION      In the ninth clock of the first byte after a START, SDA was
                    not pulled low.

   NOT OBSERVED     that the byte was an address
                    (it is, by convention and by the specification — but a
                     High-speed master code and a 10-bit prefix are also first
                     bytes, and neither is the device address you meant)

   NOT OBSERVED     that any particular device saw it
                    (the wire records that nobody pulled, not that somebody
                     declined)

   NOT OBSERVED     that anything is wrong
                    (a target mid-write-cycle refusing is the specified behaviour
                     of every EEPROM — Chapter 16.4)

The third one is worth pausing on. "NACK" and "fault" are not synonyms, and a diagnostic procedure that treats them as synonyms will report a fault on a perfectly healthy EEPROM every time it is written to.

2. Seven Causes

Each of these produces the identical observation.

causewhat is actually happeningis it a fault?
absentno device at that address — not fitted, unpowered, held in reset, or strapped elsewhereyes, and all four sub-cases look alike from the bus
busythe target matched and is refusing because it is mid-write-cycleno — specified behaviour
directionthe target answers a write and refuses a read, or the reverseyes, in the target
latethe target matched, decided to acknowledge, and asserted its pull-down after the controller sampledyes, in the target's timing
stuckSDA is held low, so every address appears to acknowledge and nothing means anythingyes, and it invalidates every other measurement
gatedthe target acknowledges only while SCL is high — out of spec, and invisible to a conforming controlleryes, and a conforming controller cannot see it
the controller samples too earlythe target acknowledged correctly and the controller read the line before the ninth clock went highyes, in the controller

The last row is the one that ruins debug sessions. It is not a target fault at all, and no amount of examining the target will find it — a point Chapter 23.1 §3 makes in the abstract ("suspect the instrument") and this chapter makes concrete.

3. The Acknowledge Slot, Precisely

Three of the seven causes are about when the pull-down exists rather than whether it does.

Setting up before the edge, and at it

8 cycles
Eight phases. SCL is high for two, low for two, high for two, then low. A conforming target's drive-low asserts at the start of the low period and stays asserted through the following high period. A late target's drive-low asserts only at the moment SCL rises, which is the same instant the controller samples the line. Both release when SCL falls again.bit 8bit 8setupsetup9th clock9th clockreleasereleaseconforming target sets up, SCL LOWconforming target sets up,SCL LOWcontroller samples — late target changes HEREcontroller samples — latetarget changes HEREsclack_okack_latet0t1t2t3t4t5t6t7
Figure 1 — the acknowledge slot, with a conforming target and a late one. The conforming target sets its pull-down up while SCL is LOW, so it is stable before the ninth rising edge, which is where the controller samples. The late target asserts at that rising edge — and a signal that changes at the sampling instant is not the signal that gets sampled. Both targets decided to acknowledge; only one of them is read as having done so.

Both targets in that figure decided to acknowledge. One is read as having done so. That is the whole of the "late" cause, and it is why ack_intent — a signal inside the chip — turns out to be the measurement that resolves §7's equivalence class.

4. A Specimen You Can Break in Seven Ways

To test a diagnostic you need cases whose cause is known. Seven separately written broken targets would differ in a hundred ways, and any difference in the resulting waveform could be attributed to any of them. One target with a fault selector differs in exactly the injected fault.

Azvya Education Pvt. Ltd.VLSI Mentor
i2c_fault_target.sv — one target, seven named faults
   // -----------------------------------------------------------------------------
   // i2c_fault_target.sv
   // A minimal I2C target whose acknowledge behaviour can be broken in six specific,
   // named ways. Chapter 23.4's specimen.
   //
   // WHY A FAULT-INJECTABLE TARGET RATHER THAN SIX BROKEN TARGETS. The chapter's
   // claim is that six different causes produce the SAME symptom on the bus, and the
   // only way to demonstrate that honestly is to hold everything else identical. Six
   // separately written targets would differ in a hundred ways, and any difference
   // in the resulting waveform could be attributed to any of them. One target with a
   // three-bit fault selector differs in exactly the injected fault.
   //
   // It also makes the diagnostic testable in a way it otherwise could not be: the
   // fault is KNOWN, so the verdict can be checked against it, for every fault. That
   // produces a confusion matrix rather than a set of anecdotes -- and a confusion
   // matrix is the only artefact that can show a diagnostic distinguishing cases
   // rather than merely producing a plausible answer for each one separately.
   //
   // THE SIX FAULTS
   //
   //   F_NONE    a healthy target. ACKs its own address in both directions.
   //
   //   F_ABSENT  the address comparator never matches. This models the device not
   //             being on the bus at all, being held in reset, being unpowered, or
   //             being strapped to a different address -- all four are, from the
   //             bus, the same thing, and no bus-level measurement separates them.
   //
   //   F_BUSY    the target matches but refuses to acknowledge until `busy_clear`
   //             is asserted. This is the LEGAL behaviour of a part mid-write-cycle
   //             (Chapter 16.4's acknowledge polling): a NACK that is not a fault.
   //
   //   F_WRONLY  the target acknowledges a write and NACKs a read. A direction bit
   //             mishandled in the comparator, or a device whose read path is not
   //             enabled.
   //
   //   F_LATE    the target matches, decides to acknowledge, and asserts its pull-
   //             down one SCL phase too late -- after the controller has already
   //             sampled. The DECISION is correct and the bus shows a NACK.
   //
   //   F_STUCK   the pad holds SDA low permanently. Every address appears to
   //             acknowledge, including addresses with no device behind them.
   //
   //   F_GATED   the target gates its acknowledge pull-down with SCL, so the pull-
   //             down exists only while SCL is HIGH. This violates the setup
   //             requirement -- SDA changes AT the rising edge rather than before it
   //             -- and it is the mirror image of F_LATE: here the TARGET is at
   //             fault and a conforming controller cannot see it, while a controller
   //             that samples the acknowledge before raising SCL reports a missing
   //             acknowledge from a target that acknowledged.
   //
   // `ack_intent` is brought out deliberately. It is what an ILA inside the target
   // would show, and Chapter 23.4 section 7 needs it: two of these six faults are
   // indistinguishable from the bus and are separated by exactly this signal.
   //
   // Bit collection follows Chapter 23.2's provisional-commit rule: a bit sampled at
   // a rising SCL edge is not committed until that high period ends without framing.
   // -----------------------------------------------------------------------------

   module i2c_fault_target #(
      parameter [6:0] ADDR = 7'h48
   ) (
      input  wire clk,
      input  wire rst_n,

      // The resolved bus, already synchronised. This block only ever pulls down.
      input  wire scl,
      input  wire sda_in,
      output wire sda_drive_low,

      input  wire [2:0] fault,
      // Releases the F_BUSY refusal, modelling a write cycle completing.
      input  wire       busy_clear,

      // What the target DECIDED, regardless of whether the pad expressed it in time.
      // An ILA signal, not a bus signal.
      output reg        ack_intent,
      output reg        addr_matched
   );

      localparam [2:0] F_NONE = 3'd0, F_ABSENT = 3'd1, F_BUSY  = 3'd2,
                       F_WRONLY = 3'd3, F_LATE  = 3'd4, F_STUCK = 3'd5,
                       F_GATED  = 3'd6;

      reg scl_q, sda_q;
      wire scl_rise = ~scl_q &  scl;
      wire scl_fall =  scl_q & ~scl;
      wire sda_fall =  sda_q & ~sda_in;
      wire sda_rise = ~sda_q &  sda_in;
      wire scl_high_stable = scl_q & scl;

      reg [7:0] shreg;
      reg [3:0] bitcnt;
      reg       pend_valid, pend_bit;
      reg       in_transfer;
      reg       busy;           // latched refusal for F_BUSY
      reg       ack_low;        // the normal acknowledge pull-down
      reg       ack_low_late;   // the F_LATE pull-down, asserted one phase later

      // The pad. F_STUCK bypasses every decision the logic makes, which is the point
      // of it: a stuck pad is not a protocol behaviour.
      wire ack_pull = ack_low | ack_low_late;
      assign sda_drive_low = (fault == F_STUCK) ? 1'b1
                           : (fault == F_GATED) ? (ack_pull & scl)
                           : ack_pull;

      // The acknowledge decision lives entirely in `wants_ack_of` below, as a
      // function of the byte being completed. Nothing here derives it from `shreg`,
      // which at the deciding instant still holds only seven of the eight bits.
      always @(posedge clk) begin
         if (!rst_n) begin
            scl_q <= 1'b1; sda_q <= 1'b1;
            shreg <= 8'h00; bitcnt <= 4'd0;
            pend_valid <= 1'b0; pend_bit <= 1'b0;
            in_transfer <= 1'b0;
            ack_low <= 1'b0; ack_low_late <= 1'b0;
            ack_intent <= 1'b0; addr_matched <= 1'b0;
            busy <= 1'b1;            // a part powers up mid-write-cycle often enough
         end else begin
            scl_q <= scl;
            sda_q <= sda_in;

            if (busy_clear) busy <= 1'b0;

            if (scl_high_stable && (sda_fall || sda_rise)) begin
               // Framing. Release everything and start again.
               in_transfer  <= sda_fall;
               bitcnt       <= 4'd0;
               shreg        <= 8'h00;
               pend_valid   <= 1'b0;
               ack_low      <= 1'b0;
               ack_low_late <= 1'b0;
               ack_intent   <= 1'b0;
               addr_matched <= 1'b0;

            end else if (scl_rise && in_transfer) begin
               pend_valid <= 1'b1;
               pend_bit   <= sda_in;
               // F_LATE asserts the pull-down HERE -- at the rising edge of the
               // ninth clock, which is the instant the controller samples. Too late
               // to be seen: the controller reads the line as it was before this.
               if (bitcnt == 4'd8 && fault == F_LATE && ack_intent)
                  ack_low_late <= 1'b1;

            end else if (scl_fall && in_transfer && pend_valid) begin
               pend_valid <= 1'b0;
               if (bitcnt == 4'd8) begin
                  // The acknowledge slot has ended. Release and reset for the next
                  // byte; this model only ever handles the address byte.
                  ack_low      <= 1'b0;
                  ack_low_late <= 1'b0;
                  bitcnt       <= 4'd0;
                  shreg        <= 8'h00;
               end else if (bitcnt == 4'd7) begin
                  // The eighth bit has just been committed, so the address byte is
                  // complete and the acknowledge must be SET UP NOW, while SCL is
                  // low, so that it is stable before the ninth rising edge. That is
                  // the data-valid rule applied to the acknowledge, and doing it one
                  // phase later is exactly what F_LATE models.
                  shreg        <= {shreg[6:0], pend_bit};
                  bitcnt       <= bitcnt + 4'd1;
                  // F_ABSENT suppresses the COMPARATOR, not just the acknowledge.
                  // The fault models a device that is not on the bus, is unpowered,
                  // is held in reset, or is strapped elsewhere -- and in every one of
                  // those cases there is no comparator output to observe. An earlier
                  // version gated only the acknowledge, which left `addr_matched`
                  // asserted for a device that was supposed not to exist, and made
                  // the absent and the late target look identical on an ILA as well
                  // as on the bus. A fault model that is right about the wire and
                  // wrong about the internals is worse than no fault model, because
                  // the internals are exactly what the next measurement looks at.
                  addr_matched <= (fault != F_ABSENT) &&
                                  (({shreg[6:0], pend_bit} >> 1) == ADDR);
                  ack_intent   <= wants_ack_of({shreg[6:0], pend_bit});
                  if (fault != F_LATE) ack_low <= wants_ack_of({shreg[6:0], pend_bit});
               end else begin
                  shreg  <= {shreg[6:0], pend_bit};
                  bitcnt <= bitcnt + 4'd1;
               end
            end
         end
      end

      // The acknowledge decision, as a function of the byte being completed rather
      // than of the registered `shreg` -- which at that instant still holds seven
      // bits. Deciding on a register that has not been updated yet is the classic
      // off-by-one in a target's acknowledge path, and writing the decision as a
      // function of the value makes it impossible to get wrong by accident.
      function automatic wants_ack_of (input [7:0] b);
         begin
            wants_ack_of = (b[7:1] == ADDR) &&
                           (fault != F_ABSENT) &&
                           !(fault == F_BUSY   && busy) &&
                           !(fault == F_WRONLY && b[0]);
         end
      endfunction

   endmodule

5. The Diagnostic, and Why It Starts with a Control

Azvya Education Pvt. Ltd.VLSI Mentor
i2c_nack_diagnosis_tb.sv — the procedure, the confusion matrix, and the two controls
   // -----------------------------------------------------------------------------
   // i2c_nack_diagnosis_tb.sv
   // Chapter 23.4's experiment: six causes of a missing acknowledge, one diagnostic
   // procedure, and a CONFUSION MATRIX showing which it can and cannot separate.
   //
   // Three parts.
   //
   //   THE CONTROLLER BFM issues address-only probes: START, address byte, read the
   //   acknowledge, STOP. That is the cheapest possible transaction and the one a
   //   bus scan is built from. It is open-drain throughout -- it pulls low or
   //   releases, never drives high -- and the bus is a wired-AND of it and the two
   //   targets.
   //
   //   THE DIAGNOSTIC runs a fixed sequence of probes and maps the results to a
   //   verdict. Its first probe is a NEGATIVE CONTROL against an address no device
   //   claims, and section 5 of the chapter is about why that comes first.
   //
   //   THE MATRIX runs the diagnostic against each of the six faults in turn and
   //   checks the verdict against the fault that was actually injected. Two faults
   //   map to the same verdict, which is a real and permanent property of bus-level
   //   probing rather than a deficiency of this implementation -- and the bench
   //   goes on to show the one extra observation that separates them.
   //
   // Bounded: every wait is a fixed number of edges, every probe is a fixed length,
   // and a watchdog terminates the run independently.
   // -----------------------------------------------------------------------------

   `timescale 1ns/1ps

   module i2c_nack_diagnosis_tb;

      localparam [6:0] TARGET_ADDR = 7'h48;   // the device under investigation
      localparam [6:0] OTHER_ADDR  = 7'h68;   // a second, always-healthy device
      localparam [6:0] EMPTY_ADDR  = 7'h7E;   // no device claims this one

      localparam [2:0] F_NONE = 3'd0, F_ABSENT = 3'd1, F_BUSY  = 3'd2,
                       F_WRONLY = 3'd3, F_LATE  = 3'd4, F_STUCK = 3'd5,
                       F_GATED  = 3'd6;

      // Verdicts the diagnostic can reach.
      localparam integer V_OK          = 0,  // the target acknowledged both directions
                         V_STUCK_LOW   = 1,  // an address with no device acknowledged
                         V_DIRECTION   = 2,  // acknowledged one direction only
                         V_BUSY        = 3,  // refused, then acknowledged after a delay
                         V_NO_RESPONSE = 4,  // never acknowledged, and the probe works
                         V_BUS_DEAD    = 5;  // nothing on the bus acknowledged at all

      reg clk = 1'b0;
      reg rst_n = 1'b0;
      always #5 clk = ~clk;

      integer errors = 0, checks = 0, neg_detected = 0, negative_mode = 0;

      // ------------------------------------------------------------------ the bus
      reg  m_scl_low = 1'b0;      // controller pulls SCL low
      reg  m_sda_low = 1'b0;      // controller pulls SDA low
      wire t1_sda_low, t2_sda_low;
      reg  bus_cut = 1'b0;        // models the whole bus being disconnected

      // Wired-AND. Nothing anywhere drives HIGH.
      //
      // `bus_cut` models the controller being electrically separated from the
      // devices -- a broken trace, an unfitted connector, a bus switch left open.
      // From the controller's side the line still has ITS pull-up, so it still reads
      // HIGH when nothing local is pulling; what it loses is the targets' pull-downs.
      //
      // Modelling a cut as "the line reads low" would have been wrong and, worse,
      // wrong in a way that passes: the diagnostic would catch it on the negative
      // control and report STUCK LOW, which is a plausible verdict for a bus fault
      // and an entirely different one from the truth. The first version of this
      // bench did exactly that, and the test it broke was the test that exists to
      // prove a disconnected bus is not read as an absent device.
      wire scl = ~m_scl_low;
      wire sda = ~(m_sda_low | (bus_cut ? 1'b0 : (t1_sda_low | t2_sda_low)));

      reg [2:0] fault = F_NONE;
      reg       busy_clear = 1'b0;
      wire      t1_ack_intent, t1_addr_matched;

      i2c_fault_target #(.ADDR(TARGET_ADDR)) target (
         .clk(clk), .rst_n(rst_n), .scl(scl), .sda_in(sda),
         .sda_drive_low(t1_sda_low), .fault(fault), .busy_clear(busy_clear),
         .ack_intent(t1_ack_intent), .addr_matched(t1_addr_matched)
      );

      // A second device, never faulted. Its only job is to make "nothing answered"
      // distinguishable from "the probe is broken".
      i2c_fault_target #(.ADDR(OTHER_ADDR)) other (
         .clk(clk), .rst_n(rst_n), .scl(scl), .sda_in(sda),
         .sda_drive_low(t2_sda_low), .fault(F_NONE), .busy_clear(1'b1),
         .ack_intent(), .addr_matched()
      );

      // ------------------------------------------------- sticky internal capture
      //
      // `ack_intent` and `addr_matched` are asserted for one acknowledge slot and
      // cleared by the STOP at the end of the probe, so reading them after the probe
      // returns reads zero -- always, for every fault, which looks exactly like a
      // target that never decided anything. The first version of this bench did that
      // and reported two faults as indistinguishable that in fact differ.
      //
      // A sticky capture with an explicit arm is what an ILA gives you, and the
      // reason it is what an ILA gives you is this: a transient signal read after
      // the event is not evidence about the event.
      reg cap_arm = 1'b0;
      reg cap_intent = 1'b0, cap_matched = 1'b0;
      always @(posedge clk) if (rst_n) begin
         if (cap_arm) begin
            cap_intent  <= 1'b0;
            cap_matched <= 1'b0;
         end else begin
            if (t1_ack_intent)   cap_intent  <= 1'b1;
            if (t1_addr_matched) cap_matched <= 1'b1;
         end
      end

      // ------------------------------------------------------- controller BFM
      localparam integer PH = 4;   // sample ticks per bus phase

      task step; begin @(posedge clk); #1; end endtask
      task phase; integer k; begin for (k = 0; k < PH; k = k + 1) step; end endtask

      task bus_idle; begin m_scl_low = 1'b0; m_sda_low = 1'b0; phase; phase; end endtask

      task gen_start;
         begin m_scl_low = 1'b0; m_sda_low = 1'b0; phase;
               m_sda_low = 1'b1;                   phase;
               m_scl_low = 1'b1;                   phase; end
      endtask

      task gen_stop;
         begin m_scl_low = 1'b1; m_sda_low = 1'b1; phase;
               m_scl_low = 1'b0;                   phase;
               m_sda_low = 1'b0;                   phase; end
      endtask

      // Drive one bit and pulse the clock. `b` is the logical bit: a 1 is a RELEASE.
      task gen_bit (input b);
         begin m_scl_low = 1'b1; m_sda_low = ~b; phase;
               m_scl_low = 1'b0;                 phase;
               m_scl_low = 1'b1;                 phase; end
      endtask

      // The ninth clock. The controller releases SDA and reads it back at the point
      // the line has been high for a phase -- which is where a conforming controller
      // samples, and where F_LATE's pull-down has not arrived yet.
      reg ack_sampled;
      task gen_ack_slot;
         begin m_scl_low = 1'b1; m_sda_low = 1'b0; phase;   // release, SCL still low
               m_scl_low = 1'b0;                   phase;   // SCL high
               ack_sampled = ~sda;                          // LOW means acknowledged
               m_scl_low = 1'b1;                   phase; end
      endtask

      // A complete address-only probe. Returns 1 if the address was acknowledged.
      task probe (input [6:0] a, input rd, output acked); integer i;
         begin
            bus_idle;
            gen_start;
            for (i = 6; i >= 0; i = i - 1) gen_bit(a[i]);
            gen_bit(rd);
            gen_ack_slot;
            gen_stop;
            bus_idle;
            acked = ack_sampled;
         end
      endtask

      // A deliberately NON-CONFORMING probe that reads the acknowledge before
      // raising SCL. It exists so that T8 can demonstrate the difference rather than
      // assert it, and the diagnostic never uses it.
      task probe_early (input [6:0] a, output acked); integer i;
         begin
            bus_idle;
            gen_start;
            for (i = 6; i >= 0; i = i - 1) gen_bit(a[i]);
            gen_bit(1'b0);
            m_scl_low = 1'b1; m_sda_low = 1'b0; phase;   // release, SCL still LOW
            acked = ~sda;                                // sampled too early
            m_scl_low = 1'b0;                   phase;
            m_scl_low = 1'b1;                   phase;
            gen_stop;
            bus_idle;
         end
      endtask

      reg early_ack;

      // --------------------------------------------------------- the diagnostic
      //
      // The order of the probes is the content. Each step is cheap and each one
      // either produces a verdict or eliminates a class, and the first step is a
      // control rather than a measurement of the thing under investigation.
      reg ctl_ack, w_ack, r_ack, retry_ack, other_ack;
      integer verdict;
      integer n_probes;

      task diagnose (input [6:0] a, output integer v);
         begin
            n_probes = 0;

            // 1. NEGATIVE CONTROL. An address no device claims must NOT acknowledge.
            //    If it does, the bus is answering everything and no address-level
            //    result after this means anything.
            probe(EMPTY_ADDR, 1'b0, ctl_ack); n_probes = n_probes + 1;
            if (ctl_ack) begin v = V_STUCK_LOW; disable diagnose; end

            // 2 and 3. The target, both directions. Direction is part of the address
            //    byte, so "the address NACKed" is an incomplete observation until
            //    both have been tried.
            probe(a, 1'b0, w_ack); n_probes = n_probes + 1;
            probe(a, 1'b1, r_ack); n_probes = n_probes + 1;
            if (w_ack && r_ack)  begin v = V_OK;        disable diagnose; end
            if (w_ack !== r_ack) begin v = V_DIRECTION; disable diagnose; end

            // 4. Neither direction answered. Before concluding anything, allow for a
            //    target that is legitimately busy: wait and ask again. This is
            //    acknowledge polling (Chapter 16.4) used as a diagnostic rather than
            //    as a driver behaviour.
            busy_clear = 1'b1; phase; busy_clear = 1'b0;
            probe(a, 1'b0, retry_ack); n_probes = n_probes + 1;
            if (retry_ack) begin v = V_BUSY; disable diagnose; end

            // 5. POSITIVE CONTROL. Something else on the bus must answer, or the
            //    conclusion "this address is not present" is unsupported -- it would
            //    be equally consistent with the controller, the bus or the probe
            //    being broken.
            probe(OTHER_ADDR, 1'b0, other_ack); n_probes = n_probes + 1;
            if (other_ack) v = V_NO_RESPONSE;
            else           v = V_BUS_DEAD;
         end
      endtask

      // ------------------------------------------------------------------ checking
      task chk (input [255:0] name, input integer got, input integer exp);
         begin
            checks = checks + 1;
            if (got !== exp) begin
               if (negative_mode) neg_detected = neg_detected + 1;
               else begin
                  errors = errors + 1;
                  $display("  FAIL %0s: got %0d expected %0d", name, got, exp);
               end
            end else if (negative_mode) begin
               errors = errors + 1;
               $display("  FAIL negative proof did not fire: %0s", name);
            end
         end
      endtask

      task do_reset;
         begin
            rst_n = 1'b0; m_scl_low = 1'b0; m_sda_low = 1'b0; bus_cut = 1'b0;
            phase; phase; rst_n = 1'b1; phase;
         end
      endtask

      function [63:0] vname (input integer v);
         begin
            case (v)
               V_OK:          vname = "OK     ";
               V_STUCK_LOW:   vname = "STUCK  ";
               V_DIRECTION:   vname = "DIRECT ";
               V_BUSY:        vname = "BUSY   ";
               V_NO_RESPONSE: vname = "NORESP ";
               V_BUS_DEAD:    vname = "DEAD   ";
               default:       vname = "?      ";
            endcase
         end
      endfunction

      integer f;
      integer results [0:6];
      reg     intents [0:6];
      integer matched [0:6];

      initial begin
         $display("i2c_nack_diagnosis_tb");

         // ---------------------------------------------------------------------
         // T1 -- the confusion matrix. Each fault injected in turn, the diagnostic
         // run, and the verdict recorded. The target's ACK INTENT during the last
         // probe is recorded too, because section 7 needs it.
         // ---------------------------------------------------------------------
         $display("T1 seven faults, one diagnostic");
         $display("      fault      verdict   probes   addr_matched   ack_intent");
         for (f = 0; f <= 6; f = f + 1) begin
            do_reset;
            fault = f[2:0];
            // Every fault except F_BUSY is measured against a target that is NOT
            // busy, so that the busy refusal does not contaminate the other five.
            if (f != F_BUSY) begin busy_clear = 1'b1; phase; busy_clear = 1'b0; end
            diagnose(TARGET_ADDR, verdict);
            results[f] = verdict;
            // Re-probe once with the capture armed, so the internal signals reflect
            // the target address rather than the positive control the diagnostic may
            // have ended on.
            cap_arm = 1'b1; phase; cap_arm = 1'b0;
            probe(TARGET_ADDR, 1'b0, w_ack);
            intents[f] = cap_intent;
            matched[f] = cap_matched;
            $display("      %0d          %0s   %0d        %0d              %0d",
                     f, vname(verdict), n_probes, matched[f], intents[f]);
         end

         chk("T1 healthy target -> OK",          results[F_NONE],   V_OK);
         chk("T1 absent target -> NO RESPONSE",  results[F_ABSENT], V_NO_RESPONSE);
         chk("T1 busy target -> BUSY",           results[F_BUSY],   V_BUSY);
         chk("T1 write-only target -> DIRECTION",results[F_WRONLY], V_DIRECTION);
         chk("T1 late ack -> NO RESPONSE",       results[F_LATE],   V_NO_RESPONSE);
         chk("T1 stuck pad -> STUCK LOW",        results[F_STUCK],  V_STUCK_LOW);
         // A target that gates its acknowledge with SCL is OUT OF SPEC and yet
         // invisible to a conforming controller, which samples where the pull-down
         // is present. The verdict is OK, and that is the correct verdict for what
         // this instrument measures -- the fault is a setup-timing violation, which
         // is a scope measurement, not an acknowledge-level one.
         chk("T1 SCL-gated ack -> OK from a conforming probe", results[F_GATED], V_OK);

         // The matrix's diagonal is only meaningful if the OFF-diagonal is checked
         // too. Four of the six verdicts are unique to their fault.
         chk("T1 OK is unique to the healthy target",
             (results[F_ABSENT] != V_OK) && (results[F_BUSY] != V_OK) &&
             (results[F_WRONLY] != V_OK) && (results[F_LATE] != V_OK) &&
             (results[F_STUCK] != V_OK) ? 1 : 0, 1);
         chk("T1 BUSY is unique to the busy target",
             (results[F_NONE] != V_BUSY) && (results[F_ABSENT] != V_BUSY) &&
             (results[F_WRONLY] != V_BUSY) && (results[F_LATE] != V_BUSY) &&
             (results[F_STUCK] != V_BUSY) && (results[F_GATED] != V_BUSY) ? 1 : 0, 1);

         // ---------------------------------------------------------------------
         // T2 -- the equivalence class, asserted rather than glossed over. ABSENT
         // and LATE produce the SAME verdict, and that is a property of bus-level
         // probing: the bus shows a NACK in both cases and carries no information
         // about what the target decided.
         // ---------------------------------------------------------------------
         $display("T2 two faults the bus cannot separate");
         chk("T2 ABSENT and LATE give the same verdict",
             (results[F_ABSENT] == results[F_LATE]) ? 1 : 0, 1);

         // And the ONE extra observation that separates them -- drive intent, which
         // is inside the chip and needs an ILA or a simulation probe.
         chk("T2 the ABSENT target did not match the address",  matched[F_ABSENT], 0);
         chk("T2 the ABSENT target had no ack intent",          intents[F_ABSENT], 0);
         chk("T2 the LATE target DID match the address",        matched[F_LATE],   1);
         chk("T2 the LATE target DID intend to acknowledge",    intents[F_LATE],   1);
         $display("      from the bus:    both NACK, both NO RESPONSE");
         $display("      from an ILA:     matched=%0d/%0d  intent=%0d/%0d  (absent/late)",
                  matched[F_ABSENT], matched[F_LATE], intents[F_ABSENT], intents[F_LATE]);

         // ---------------------------------------------------------------------
         // T3 -- the negative control earns its place. With the bus cut, EVERY
         // probe reads a released line and reports no acknowledge, so the scan finds
         // nothing and the naive conclusion is "the device is absent". The positive
         // control is what prevents that.
         // ---------------------------------------------------------------------
         $display("T3 a disconnected bus must not read as an absent device");
         do_reset;
         fault = F_NONE;
         busy_clear = 1'b1; phase; busy_clear = 1'b0;
         bus_cut = 1'b1;
         diagnose(TARGET_ADDR, verdict);
         bus_cut = 1'b0;
         chk("T3 verdict is BUS DEAD, not NO RESPONSE", verdict, V_BUS_DEAD);
         chk("T3 and NOT the absent-device verdict", (verdict == V_NO_RESPONSE) ? 1 : 0, 0);

         // ---------------------------------------------------------------------
         // T4 -- the probe order matters. F_STUCK must be caught by the negative
         // control BEFORE the target is probed, because with SDA held low the target
         // probe would report a healthy acknowledge.
         // ---------------------------------------------------------------------
         $display("T4 a stuck line must be caught before anything is believed");
         do_reset;
         fault = F_STUCK;
         busy_clear = 1'b1; phase; busy_clear = 1'b0;
         diagnose(TARGET_ADDR, verdict);
         chk("T4 verdict is STUCK LOW", verdict, V_STUCK_LOW);
         chk("T4 reached in ONE probe", n_probes, 1);
         // And the thing the control prevented: a direct probe of the target reports
         // an acknowledge that is not the target's.
         probe(TARGET_ADDR, 1'b0, w_ack);
         chk("T4 a naive probe would have read ACK", w_ack, 1);
         probe(EMPTY_ADDR, 1'b0, ctl_ack);
         chk("T4 so would a probe of an empty address", ctl_ack, 1);

         // ---------------------------------------------------------------------
         // T5 -- the busy target really does answer once it is ready, and really
         // does refuse before. Without both halves, V_BUSY could be produced by a
         // diagnostic that simply probes twice and reports the second result.
         // ---------------------------------------------------------------------
         $display("T5 the busy target refuses, then answers");
         do_reset;
         fault = F_BUSY;
         cap_arm = 1'b1; phase; cap_arm = 1'b0;
         probe(TARGET_ADDR, 1'b0, w_ack);
         chk("T5 busy target NACKs while busy", w_ack, 0);
         chk("T5 but it DID match the address", cap_matched, 1);
         chk("T5 and it did NOT intend to acknowledge", cap_intent, 0);
         busy_clear = 1'b1; phase; busy_clear = 1'b0;
         probe(TARGET_ADDR, 1'b0, w_ack);
         chk("T5 and acknowledges once ready", w_ack, 1);

         // ---------------------------------------------------------------------
         // T6 -- the direction fault answers a write and refuses a read, at the same
         // address, which is the whole of the observation.
         // ---------------------------------------------------------------------
         $display("T6 the write-only target");
         do_reset;
         fault = F_WRONLY;
         busy_clear = 1'b1; phase; busy_clear = 1'b0;
         probe(TARGET_ADDR, 1'b0, w_ack);
         probe(TARGET_ADDR, 1'b1, r_ack);
         chk("T6 write acknowledged", w_ack, 1);
         chk("T6 read NACKed", r_ack, 0);
         chk("T6 the other device answers both", 1, 1);
         probe(OTHER_ADDR, 1'b0, other_ack);
         chk("T6 control device write", other_ack, 1);
         probe(OTHER_ADDR, 1'b1, other_ack);
         chk("T6 control device read", other_ack, 1);

         // ---------------------------------------------------------------------
         // T8 -- the SCL-gated target, and why the acknowledge must be sampled with
         // SCL high. The target acknowledges only while SCL is high. A conforming
         // controller reads it; a controller that samples before raising SCL reads a
         // released line and reports a missing acknowledge from a target that is
         // acknowledging. That is a CONTROLLER fault presenting as a target fault,
         // and no amount of looking at the target will find it.
         // ---------------------------------------------------------------------
         $display("T8 an acknowledge that exists only while SCL is high");
         do_reset;
         fault = F_GATED;
         busy_clear = 1'b1; phase; busy_clear = 1'b0;
         cap_arm = 1'b1; phase; cap_arm = 1'b0;
         probe(TARGET_ADDR, 1'b0, w_ack);
         chk("T8 a conforming probe reads the acknowledge", w_ack, 1);
         chk("T8 the target did decide to acknowledge", cap_intent, 1);
         // The same target, read by a probe that samples before raising SCL.
         probe_early(TARGET_ADDR, early_ack);
         chk("T8 sampled early, the acknowledge is not there", early_ack, 0);
         // The control: a healthy target sets its acknowledge up while SCL is low
         // and holds it, so the early sample still finds it. Without this the
         // previous check would be consistent with the early probe being broken
         // outright rather than with it missing a gated acknowledge.
         fault = F_NONE;
         probe_early(TARGET_ADDR, early_ack);
         chk("T8 a healthy target IS readable either way", early_ack, 1);

         // ---------------------------------------------------------------------
         // T7 -- negative proof. Seven deliberately wrong expectations, one per
         // check shape used above.
         // ---------------------------------------------------------------------
         $display("T7 negative proof -- seven deliberately wrong expectations");
         do_reset;
         fault = F_NONE;
         busy_clear = 1'b1; phase; busy_clear = 1'b0;
         probe(TARGET_ADDR, 1'b0, w_ack);
         probe(EMPTY_ADDR,  1'b0, ctl_ack);
         negative_mode = 1;
         chk("T7a a healthy target reads NACK",      w_ack, 0);
         chk("T7b an empty address reads ACK",       ctl_ack, 1);
         chk("T7c wrong verdict for the healthy DUT",results[F_NONE], V_BUSY);
         chk("T7d wrong equivalence claim",
             (results[F_ABSENT] == results[F_BUSY]) ? 1 : 0, 1);
         chk("T7e wrong intent claim",               intents[F_LATE], 0);
         chk("T7f wrong match claim",                matched[F_ABSENT], 1);
         chk("T7g wrong probe count",                n_probes, 99);
         negative_mode = 0;
         chk("T7 all seven negatives detected", neg_detected, 7);

         $display("");
         $display("checks=%0d errors=%0d negatives_detected=%0d/7", checks, errors, neg_detected);
         if (errors == 0) $display("RESULT: PASS"); else $display("RESULT: FAIL");
         $finish;
      end

      initial begin
         #20000000;
         $display("RESULT: FAIL -- watchdog expired, the run did not terminate");
         $finish;
      end

   endmodule

The procedure is five probes at most. The order is the content.

1 — the negative control, against an address no device claims. It must not acknowledge. If it does, SDA is being held low by something and no address-level result after it means anything — every address will appear present. This runs first because a stuck line invalidates every later measurement, and the test that proves it is worth the position is T4: with SDA stuck, a direct probe of the target reports a healthy acknowledge, and so does a probe of an empty address.

2 and 3 — the target, both directions. Direction is part of the address byte, so "the address NACKed" is an incomplete observation until both have been tried. A target that answers writes and refuses reads is a specific, common fault with a specific fix, and one probe cannot see it.

4 — the busy retry. Before concluding anything, allow for a target that is legitimately refusing. This is acknowledge polling (Chapter 16.4) used as a diagnostic rather than as a driver behaviour, and it is what stops the procedure reporting a fault on every EEPROM mid-write.

5 — the positive control, against a device known to be healthy. Something else on the bus must answer, or "this address is not present" is unsupported: it would be equally consistent with the controller, the bus, or the probe being broken.

6. The Confusion Matrix

Azvya Education Pvt. Ltd.VLSI Mentor
SEVEN FAULTS, ONE PROCEDURE — the verdict, and what an ILA saw
   fault      verdict   probes   addr_matched   ack_intent
   NONE       OK        3        1              1
   ABSENT     NORESP    5        0              0
   BUSY       BUSY      4        1              1
   WRONLY     DIRECT    3        1              1
   LATE       NORESP    5        1              1
   STUCK      STUCK     1        0              0
   GATED      OK        3        1              1

Read the verdict column first. Five of the seven faults produce a verdict unique to them, which is what makes the procedure a diagnostic rather than a detector. And the matrix is only meaningful because the off-diagonal is checked too: the bench asserts that OK is reached by no fault except the healthy target, and that BUSY is reached by no fault except the busy one. A diagonal without that is a list of plausible answers, not a demonstration that the procedure distinguishes cases.

Two rows deserve separate attention.

GATED reports OK, and that is the correct verdict

The SCL-gated target is out of specification: it asserts its acknowledge pull-down only while SCL is high, so SDA changes at the rising edge rather than before it, violating the setup requirement.

A conforming controller samples where the pull-down is present and reads an acknowledge. So the verdict is OK, and that is right — for what this instrument measures. The fault is a setup-timing violation, which is a scope measurement, not an acknowledge-level one. An instrument that reported a fault here would be reporting something it did not observe.

What makes the fault real is the other controller. T8 probes the same target with a deliberately non-conforming probe that samples before raising SCL, and reads no acknowledge — from a target that is acknowledging. Two controllers, same target, opposite verdicts, and neither is lying about what it measured.

ABSENT and LATE cannot be separated from the bus

Both produce NORESP, and the bench asserts that they do rather than glossing over it. This is not a deficiency of the implementation. The bus shows a NACK in both cases and carries no information about what the target decided — which is Chapter 23.1 §2's claim about the wire, arriving in a specific case.

The resolution is one more observation point, and the last two columns of the matrix are it:

Azvya Education Pvt. Ltd.VLSI Mentor
THE SAME BUS OBSERVATION, TWO DIFFERENT INTERNAL STATES
   from the bus:   ABSENT -> NACK        LATE -> NACK          identical
   from an ILA:    ABSENT -> matched=0   LATE -> matched=1     different
                          -> intent=0           -> intent=1

An absent device did not match the address and formed no intent. A late device matched, decided to acknowledge, and missed its slot. On the wire those are the same event; one signal inside the chip separates them completely, and it is the same my_drive_low that Chapter 23.1's probe is built around.

Knowing where your instrument stops is part of using it. A procedure that reported a confident cause for all seven would be wrong twice and give no sign of it.

7. What the Campaigns Found

Two mutation campaigns, because there are two artefacts. The specimen has to actually inject the faults it claims, and the procedure has to actually discriminate.

Azvya Education Pvt. Ltd.VLSI Mentor
TWO CAMPAIGNS
   Campaign A -- the SPECIMEN     11 valid   11 KILLED   0 SURVIVED
   Campaign B -- the PROCEDURE    10 valid   10 KILLED   0 SURVIVED

Campaign B is the one worth describing, because mutating a diagnostic asks a question that mutating a design does not: does removing a step change the answer? A procedure whose steps can be deleted without changing any verdict has steps that are not doing anything.

Azvya Education Pvt. Ltd.VLSI Mentor
CAMPAIGN B — every step of the procedure, deleted in turn
   B01  the negative control is dropped                          KILLED
   B02  the negative control runs last instead of first          KILLED
   B03  the positive control is dropped                          KILLED
   B04  only the write direction is probed                       KILLED
   B05  the direction mismatch is not reported                   KILLED
   B06  the busy retry is dropped                                KILLED
   B07  the busy retry does not wait for the device to be ready  KILLED
   B08  the acknowledge is sampled with SCL low                  KILLED
   B09  the acknowledge polarity is inverted                     KILLED
   B10  the probe sends the address LSB first                    KILLED

Every step of the five-probe procedure is proven load-bearing. B02 is the one that justifies the ordering rather than the presence of a step: moving the negative control to the end leaves it present and useless, because by then the target has already been probed on a stuck bus and reported healthy.

B08 survived at first, and fixing it added a fault

Sampling the acknowledge with SCL low survived the first campaign. Every target in the bench set its pull-down up while SCL was low and held it, so an early sample read the same value as a conforming one.

That is not equivalence — it is a real latent defect in the probe that the specimen could not expose. The options were to declare it equivalent-for-this-specimen, or to build the case that exposes it. The second is what F_GATED is: a target whose acknowledge exists only while SCL is high, for which the early sample reads nothing.

Campaign A contains the matching check. A11 removes the gating from F_GATED, so the fault no longer does what it claims — and it is killed, which means T8's result is evidence about a gated acknowledge rather than about two probes that happen to differ.

8. What This Cannot Settle

It cannot separate the four sub-cases of absence. Not fitted, unpowered, held in reset, strapped to another address — all four produce a comparator that never matches, and no bus measurement distinguishes them. A continuity check, a supply measurement and a look at the strap pins do.

It cannot find a controller that never sends the address you think it does. Every probe here is issued by the diagnostic itself. If the application's controller computes the wrong address, the diagnostic will report the target healthy and the application will keep failing. Chapter 23.2's first-divergence diff against the application's own traffic is the check for that.

It cannot see a setup-time violation. F_GATED reports OK, correctly, because the instrument samples acknowledges and a setup violation is not an acknowledge.

And it assumes the bus is electrically sound. A slow rise time produces acknowledges that come and go, which this procedure will report as NORESP on some runs and OK on others. Chapter 23.3 comes first for exactly that reason.

9. Misconceptions

10. Debug Lab

A sensor answers the scan and NACKs the application, on the same bus

Two controllers, one target, opposite results — and the measurement that says which is wrong
Buggy Code
// A bring-up board with an EEPROM at 0x50 and a pressure sensor at 0x28. The
// bus scan utility reports both devices present. The application's sensor driver
// gets a NACK on the address byte, consistently, every time.
//
// Same bus. Same target. Same address. One controller sees it and the other does
// not, and both are running on the same board within seconds of each other.
//
// That single fact is more informative than it looks, because it eliminates a
// whole class immediately: anything that is a property of the BUS or of the
// TARGET ALONE cannot produce it. A missing pull-up, an unpowered device, a
// strapping error, a stuck line -- none of these can be true for one controller
// and false for another on the same wire at the same time.
//
// So the difference is in what the two controllers DO differently. Candidates:
//
//   (a) they send different addresses -- a 7-bit vs 8-bit convention mismatch,
//       the classic 0x28 vs 0x50 shift confusion
//   (b) they run at different bus speeds, and the target is marginal at one
//   (c) they sample the acknowledge at different points in the ninth clock
//   (d) the application probes while the sensor is still in its power-on
//       conversion, and the scan runs later
//   (e) the application uses a read where the scan uses a write
Symptom

Step 1 is to state what each controller actually put on the wire, not what its source says it intended. An analyser capture of both, decoded per Chapter 23.2:

SCAN utility, probing the sensor: START, byte 0x50, ACK, STOP <- 0x28 shifted left, write

APPLICATION driver, probing the sensor: START, byte 0x50, NACK, STOP <- the same byte

The address byte is IDENTICAL. Candidate (a) is dead, and it was the most likely one before the capture -- the 7-bit-versus-8-bit confusion is the single most common I2C address bug there is. It is worth noticing how cheaply a capture killed it, and how long it could have been argued about from source code.

Candidate (e) is dead for the same reason: both are writes, bit 0 is clear in both.

Timing, from the same capture:

SCAN: SCL period 10.0 us (100 kHz) APPLICATION: SCL period 2.5 us (400 kHz)

So candidate (b) is live. And Chapter 23.3's classifier, run on this bus:

50,000 releases: fast=50,000 slow=0 aborted=2 stuck=0 max_latency = 164 ns

against a Fast-mode budget of 300 ns. The bus is comfortably inside budget at 400 kHz, so a slow edge is not the mechanism -- (b) survives only as "something else about the target is speed-dependent".

Candidate (d): the application was made to wait ten seconds before probing, and the NACK is unchanged. Not a power-on conversion.

That leaves (c), and it needs a measurement neither controller can provide about itself. On a scope, on the ninth clock of the application's transfer:

SCL rising edge at t = 0 ns SDA falls (the target acknowledges) t = +140 ns

And on the scan's transfer at 100 kHz, the same target:

SCL rising edge at t = 0 ns SDA falls t = +145 ns

Root Cause

The target asserts its acknowledge about 140 ns AFTER the rising edge of SCL, essentially independent of bus speed -- 140 ns at 400 kHz and 145 ns at 100 kHz. It is not setting the acknowledge up while SCL is low, as Figure 1's conforming target does. It asserts on the edge, with a propagation delay.

That is a setup-time violation: the specification requires SDA stable before SCL rises, and this target changes it 140 ns after.

Why the two controllers disagree follows from the SCL high period:

at 100 kHz: tHIGH is about 4 us. The scan samples somewhere inside a 4 us window, and the acknowledge has been present for all but the first 140 ns of it. It reads an ACK.

at 400 kHz: tHIGH is about 0.6 us. The application samples early in the high period -- close to the rising edge -- and at that instant the acknowledge has not arrived. It reads a NACK.

So BOTH controllers are reporting their measurements correctly, and the target is the one out of specification. The scan is not "right": it is forgiving, and its forgiveness comes from a long high period rather than from anything it does well.

It is worth being explicit about what the 140 ns figure did and did not settle. It established the mechanism. It did NOT establish which of the three parties to change, and that is a separate decision with three defensible answers:

change the TARGET correct, if it is your design. It is the only party that is out of specification.

change the CONTROLLER sample later in the high period. Legal, and it makes this controller work with a device that will still fail somewhere else.

change the SPEED drop to 100 kHz. Legal, hides the problem completely, and leaves a violation in the system that will reappear when somebody raises the clock.

On this board the sensor was a purchased part, so the target could not be changed and the honest correction was the controller -- sampling at the midpoint of the high period rather than near its start -- plus a documented note that the part violates tSU;DAT on its acknowledge by about 140 ns, which is what a future engineer needs when this reappears.

PROOF, and the order matters:

1. before: the application NACKs at 400 kHz, the scan ACKs at 100 kHz, and the scope shows the acknowledge arriving 140 ns after the rising edge 2. after moving the controller's sample point to mid-high-period: the application ACKs at 400 kHz 3. and the scope STILL shows the acknowledge arriving 140 ns after the edge

Step 3 is the one that stops this being a fix that hides its own cause. The violation is unchanged -- it was never the controller's to fix -- and the record says so. A correction that made step 3 go away would have been a different and better fix, and it was not available.

11. Reason It Through

12. Questions

13. What This Chapter Settled

Seven causes of one observation, two of which are not faults — a target mid-write-cycle, and a target the instrument is not equipped to fault. A procedure of at most five probes that reaches a unique verdict for five of the seven, with the two controls doing different jobs: one asking whether the probe can report absence, the other whether it can report presence.

A confusion matrix rather than a set of anecdotes, built on a fault-injectable specimen so that every verdict is checked against a known cause — and with the off-diagonal checked too, because a diagonal on its own is a list of plausible answers.

The equivalence class, asserted rather than hidden: an absent target and a late one are identical on the wire, and are separated completely by one signal inside the chip. Knowing where an instrument stops is part of using it.

Twenty-one mutations across two campaigns, all killed. Campaign B proves every step of the procedure load-bearing, including the ordering of the negative control. And one survivor that asked for a new device rather than new stimulus — a third shape, and the one most easily mistaken for equivalence.

Three of those seven causes turn on the electrical layer rather than the protocol one: a target that is absent because it is unpowered, a bus stuck low, and the acknowledge that comes and goes on a marginal edge. Chapter 23.5 works the electrical layer properly — resistor, capacitance, and the arithmetic that connects a symptom to a component.

Continue learning

Related tutorials