Skip to content
VLSI Mentor

I²C · Module 23

Protocol Bug or Electrical Bug? Splitting the Search Space

The highest-value triage decision in I²C debugging, and the evidence that settles it. Why a wrong bit, an intermittent acknowledge and a failure that only appears at 400 kHz are produced identically by a shift-register bug and by a slow edge — and how one measurement, the time from release to a readable HIGH, separates them. Builds the classifier, runs one drive pattern against two rise times, and reports the bench race that produced an off-by-one at some latencies and not others.

Chapter 23.1 puts "establish that the bus is alive" at step 3, before any protocol reasoning, and gives a one-line justification: a bus that cannot idle high, cannot reach a valid low, or rises too slowly will produce protocol symptoms in unlimited variety.

This chapter is that claim made precise, and then made measurable. It is the highest-value single decision in I²C debugging, because getting it the wrong way round does not merely waste time — it sends you to search a region the evidence had already cleared, with symptoms that will keep confirming whatever you suspect.

1. The Symptoms Are the Same

Here is the difficulty, stated as plainly as it can be:

symptom on the busa protocol causean electrical cause
a byte arrives with one bit wrongthe shift register loaded or shifted incorrectlythe receiver sampled a released line before it had finished rising
an acknowledge appears sometimesthe address comparator is marginal, or a state is entered inconsistentlythe acknowledge pull-down is fine and the release is slow, so the next bit is corrupted
it works at 100 kHz and fails at 400 kHza timing-dependent RTL bug that the slower clock hidesthe rise time was always too long, and only Fast-mode's shorter bit period exposes it
it fails on two units out of fiftyan uninitialised register, or a racea resistor at one end of its tolerance plus a board at the other end of its capacitance
it started failing when a device was addedan address conflictthe added device's pin capacitance pushed Cb past what the pull-up can drive

Every row is a real pairing, and in every row the two causes are indistinguishable from the byte values. That is the whole problem. The natural move — look at what came out, work backwards — cannot separate them, because what came out is the same.

2. What Can Separate Them

Not the byte values. The shape of the edge — and specifically one number:

After the last device lets go of a line, how long does it take to read HIGH?

That number is not visible in a protocol decode at all. Chapter 23.2's decoder will happily report 0x5A whether the edges took twenty nanoseconds or six hundred, because it samples on the rising edge of SCL and reports what it found. The information is in the time domain, and it has to be measured deliberately.

Two figures make the point. Both show the same bit slot, in which the receiver samples a zero, and both decode to the same byte.

The receiver samples a zero, and the transmitter sent one

8 cycles
Four phases of context then a bit slot. SCL is high, then low for two phases, then high for two phases, then low. The transmitter's drive-low intent asserts at the start of the low period and stays asserted through the whole of the following high period. SDA is low throughout that window. At the rising edge of SCL the receiver samples a zero, and the transmitter is at that instant actively pulling the line down.bit setupbit setupthe samplethe sampletransmitter pulls LOWtransmitter pulls LOWsampled 0 — and intent is 0sampled 0 — and intent is 0scldrive_lowsdat0t1t2t3t4t5t6t7
Figure 1 — a protocol cause. The transmitter is actively pulling SDA low through the entire high period, so the zero the receiver samples is the zero the transmitter sent. Drive intent and the wire agree at every instant. If that zero was not the bit intended, the fault is upstream of the pin — in the shift register, the bit counter, or whatever chose the bit.

The receiver samples a zero, and nothing is driving

8 cycles
The same four phases of context and the same bit slot. The transmitter's drive-low intent deasserts at the start of the low period, meaning it intends to send a one. SDA nevertheless stays low through the low period and through the first phase of the following high period, only reaching high in the second. At the rising edge of SCL the receiver samples a zero from a line that no device is pulling down.still lowstill lowthe samplethe sampletransmitter RELEASES — intent is 1transmitter RELEASES —intent is 1sampled 0 — nothing is drivingsampled 0 — nothing isdrivingthe line finally rises — too latethe line finally rises —too latescldrive_lowsdat0t1t2t3t4t5t6t7
Figure 2 — an electrical cause, decoding to the same byte. The transmitter released SDA at the start of the low period, intending to send a one. The line had not finished rising when SCL went high, so the receiver sampled a zero from a line nobody was driving. The decoded byte is identical to Figure 1; the discriminator is that drive intent and the wire disagree, and that the rise took three phases.

Put the two sda traces through a protocol decoder and they produce the same bit. Put the two figures in front of an engineer and the difference is obvious — because the figures carry drive_low, which a capture of the bus does not.

That is the same asymmetry as Chapter 23.1 §4: the discriminating information is in the relationship between intent and the wire, and one observation point cannot hold a relationship.

3. From an Impression to a Count

"The edges look bad" is not a finding. It cannot be compared against a budget, it cannot be tracked across a change, and it cannot distinguish one bad edge in ten thousand from a bus that is uniformly slow — which are completely different problems with completely different fixes.

The instrument below measures every release and puts it in one of four bins.

Azvya Education Pvt. Ltd.VLSI Mentor
i2c_rise_classifier.sv — every release, measured and binned
   // -----------------------------------------------------------------------------
   // i2c_rise_classifier.sv
   // Chapter 23.3's instrument: it measures what happens after every RELEASE of a
   // bus line, and turns the electrical-or-protocol question into counts.
   //
   // THE PROBLEM IT SOLVES. A great many I2C symptoms are produced by both a
   // protocol bug and an electrical one, and on the resolved bus they look the same:
   //
   //     a byte with one bit wrong          a shift-register bug, or a slow edge the
   //                                        receiver sampled before it had risen
   //     an acknowledge that comes and goes an address-match bug, or a marginal rise
   //                                        time that only fails at temperature
   //     a transfer that fails at 400 kHz
   //       and works at 100 kHz             a timing-dependent RTL bug, or a bus whose
   //                                        rise time was always too long for Fast-mode
   //
   // Reading the byte values cannot separate those, because the byte values are the
   // same. What separates them is the SHAPE OF THE EDGES, and specifically one
   // number: how long the line takes to read HIGH after the last device lets go.
   //
   // WHAT IT MEASURES. On every transition of `all_released` from 0 to 1 -- the
   // instant the last pull-down is removed -- it starts a counter and stops when
   // `line` reads 1. Each release lands in exactly one bin:
   //
   //   FAST      the line was high within TR_BUDGET clocks. Electrically healthy at
   //             this speed. Any symptom is somewhere else.
   //
   //   SLOW      the line eventually read high, but took longer than TR_BUDGET. The
   //             bus is out of its rise-time budget (Chapter 19.3: tr = 0.8473*Rp*Cb).
   //             Every protocol symptom above this is suspect as a CONSEQUENCE.
   //
   //   ABORTED   a device pulled the line low again before it had risen. Not a
   //             measurement of anything: the pull-up never got to finish. Counted
   //             separately, because folding these into FAST is how an instrument
   //             reports a healthy bus that it never actually measured.
   //
   //   STUCK     everybody stayed released for STUCK_LIMIT clocks and the line never
   //             rose. That is not a slow pull-up, it is something holding the line
   //             that is not visible in `all_released` -- an unmodelled participant,
   //             or a short. Chapter 23.6.
   //
   // THE DISCRIMINATION, stated as a rule:
   //
   //     n_slow = 0 and n_stuck = 0 over a long window
   //       -> the bus is electrically sound at this speed, and a protocol symptom
   //          is a protocol bug. This is a STRONG NEGATIVE and it is the result
   //          worth wanting.
   //
   //     n_slow > 0
   //       -> the bus is out of budget. Fix that BEFORE interpreting any protocol
   //          symptom, because a receiver sampling a line that has not finished
   //          rising reads a zero where a one was sent, and that is indistinguishable
   //          from a data bug at the byte level.
   //
   // WHAT IT CANNOT DO. It cannot tell you WHY a rise is slow -- resistor value, bus
   // capacitance, a leaky device, a level shifter -- and it cannot see anything
   // analogue. It converts "the edges look bad" from an impression into a count, and
   // that is the whole of its contribution.
   //
   // `all_released` is the caller's knowledge that no device it knows about is
   // pulling. In simulation that is the bus model's drive vector. On hardware it is
   // this device's own drive intent, which makes STUCK the bin that catches every
   // participant the caller does not know about.
   // -----------------------------------------------------------------------------

   module i2c_rise_classifier #(
      // Clocks a released line is allowed to take to read HIGH. Derived from the
      // speed mode's tr and the sample clock, NOT chosen for convenience -- the
      // budget is what makes SLOW mean something.
      parameter integer TR_BUDGET   = 4,
      // How long a fully released line may stay low before it is called stuck.
      parameter integer STUCK_LIMIT = 64,
      parameter integer CNT_W       = 16
   ) (
      input  wire clk,
      input  wire rst_n,

      input  wire all_released,   // 1 when no known device is pulling the line low
      input  wire line,           // the resolved bus, read back

      output reg [CNT_W-1:0] n_fast,
      output reg [CNT_W-1:0] n_slow,
      output reg [CNT_W-1:0] n_aborted,
      output reg [CNT_W-1:0] n_stuck,

      // The worst rise seen, in clocks. A single slow edge among thousands of fast
      // ones is a different finding from a bus that is uniformly slow, and a count
      // alone cannot tell them apart.
      output reg [CNT_W-1:0] max_latency,
      // Which release first exceeded the budget. Releases are numbered from ONE, so
      // zero unambiguously means "no slow release has occurred" and no second flag
      // is needed to say so -- two pieces of state that can disagree about the same
      // fact are worse than one that cannot.
      output reg [CNT_W-1:0] first_slow_release,
      output reg [CNT_W-1:0] n_releases,

      output reg measuring
   );

      localparam [CNT_W-1:0] CNT_MAX = {CNT_W{1'b1}};

      reg all_released_q;
      wire release_edge = ~all_released_q & all_released;

      reg [CNT_W-1:0] timer;

      // Every counter increment below is guarded against its maximum. Saturating,
      // not wrapping: a wrapped diagnostic counter reports 0 for a great many events,
      // and 0 is also what a clean bus reports. Written out rather than hidden in a
      // macro, because a compiler directive is global to the compilation and a
      // diagnostic block is the last place to introduce one.

      always @(posedge clk) begin
         if (!rst_n) begin
            all_released_q     <= 1'b1;
            n_fast             <= {CNT_W{1'b0}};
            n_slow             <= {CNT_W{1'b0}};
            n_aborted          <= {CNT_W{1'b0}};
            n_stuck            <= {CNT_W{1'b0}};
            max_latency        <= {CNT_W{1'b0}};
            first_slow_release <= {CNT_W{1'b0}};
            n_releases         <= {CNT_W{1'b0}};
            timer              <= {CNT_W{1'b0}};
            measuring          <= 1'b0;
         end else begin
            all_released_q <= all_released;

            if (measuring) begin
               if (!all_released) begin
                  // Somebody pulled again before the line rose. The pull-up never
                  // finished, so this is not a rise-time measurement.
                  measuring <= 1'b0;
                  if (n_aborted != CNT_MAX) n_aborted <= n_aborted + 1'b1;
               end else if (line) begin
                  measuring <= 1'b0;
                  if (timer > max_latency) max_latency <= timer;
                  if (timer > TR_BUDGET[CNT_W-1:0]) begin
                     if (n_slow != CNT_MAX) n_slow <= n_slow + 1'b1;
                     // Latched on the FIRST slow release only. Guarded on n_slow
                     // being zero rather than on a separate flag, so the two can
                     // never disagree about whether a slow release has happened.
                     if (n_slow == {CNT_W{1'b0}}) first_slow_release <= n_releases;
                  end else begin
                     if (n_fast != CNT_MAX) n_fast <= n_fast + 1'b1;
                  end
               end else if (timer >= STUCK_LIMIT[CNT_W-1:0]) begin
                  measuring <= 1'b0;
                  if (timer > max_latency) max_latency <= timer;
                  if (n_stuck != CNT_MAX) n_stuck <= n_stuck + 1'b1;
               end else begin
                  timer <= timer + 1'b1;
               end

            end else if (release_edge) begin
               // A release only starts a measurement if the line is actually still
               // low. A release onto an already-high line is a zero-length rise and
               // is counted FAST immediately -- it is a real, and common, release.
               if (n_releases != CNT_MAX) n_releases <= n_releases + 1'b1;
               if (line) begin
                  if (n_fast != CNT_MAX) n_fast <= n_fast + 1'b1;
               end else begin
                  measuring <= 1'b1;
                  timer     <= {CNT_W{1'b0}};
               end
            end
         end
      end

   endmodule

Why there are four bins and not two

The obvious design has two bins, fast and slow. It is wrong, and the two extra bins are where the engineering is.

ABORTED is not a measurement. If a device pulls the line low again before it has finished rising, the pull-up never got to complete and the elapsed time means nothing. Folding those into FAST is how an instrument reports a healthy bus it never actually measured — and on a busy bus with back-to-back bits, aborted releases can easily outnumber completed ones. A count that silently includes them is a count of the wrong thing.

STUCK is not a slow rise. If every device the caller knows about has released and the line still has not risen after STUCK_LIMIT, the cause is not a resistor. Something is holding the line that is not in all_released — an unmodelled participant, a short, a device in reset with its pad driven. Reporting that as "very slow" points at the pull-up, which is the one thing that is not the problem. Chapter 23.6 is that case.

The bins therefore sum to the release count, and that sum is checked. A classifier whose bins do not add up is dropping or double-counting events, and neither shows up as a wrong individual bin.

The budget is the whole meaning of SLOW

TR_BUDGET is derived from the speed mode's tr and the sample clock, and it must be derived rather than chosen. Chapter 19.3 gives the relationship: tr = 0.8473 · Rp · Cb, against the specification's maximum for the mode in use — 1000 ns for Standard-mode, 300 ns for Fast-mode, 120 ns for Fast-mode Plus.

A budget picked to make the current bus look acceptable is a number that can never report a problem, and an instrument that can never report a problem is one whose silence gets mistaken for evidence.

4. The Bench and the Experiment

Azvya Education Pvt. Ltd.VLSI Mentor
i2c_rise_classifier_tb.sv — four bins, the budget boundary, and one drive pattern against two buses
   // -----------------------------------------------------------------------------
   // i2c_rise_classifier_tb.sv
   // The oracle, plus the experiment that is the point of Chapter 23.3: the SAME
   // protocol symptom produced twice, once by a protocol bug and once by a slow
   // edge, with the classifier separating them.
   //
   // The bench contains a one-line RC bus model. It is a restatement of Chapter
   // 19.7's `i2c_bus_rc` reduced to a single line and a single parameter, included
   // here so this file runs on its own -- an artifact a reader cannot execute is a
   // listing, not evidence.
   //
   // Bounded: every wait is a fixed number of clock edges, and a watchdog terminates
   // the run independently of anything the DUT does.
   // -----------------------------------------------------------------------------

   `timescale 1ns/1ps

   // A single line with a rise time. Falling is immediate (a transistor discharges
   // the net); rising takes RISE_CLKS after the last device releases (a resistor
   // charges it). That asymmetry is the whole model, and it is the digital
   // consequence of tr = 0.8473 * Rp * Cb rather than an analogue ramp.
   module tb_rc_line #(
      parameter integer RISE_CLKS = 0
   ) (
      input  wire clk,
      input  wire drive_low,      // 1 = a device is pulling this line down
      output wire line
   );
      reg [15:0] rise_cnt = RISE_CLKS[15:0];
      always @(posedge clk) begin
         if (drive_low)                           rise_cnt <= 16'd0;
         else if (rise_cnt < RISE_CLKS[15:0])     rise_cnt <= rise_cnt + 16'd1;
      end
      // Low while anyone pulls, and for RISE_CLKS clocks after the last one lets go.
      assign line = ~drive_low && (rise_cnt >= RISE_CLKS[15:0]);
   endmodule


   module i2c_rise_classifier_tb;

      localparam integer CNT_W = 16;
      localparam integer TR_BUDGET   = 4;
      localparam integer STUCK_LIMIT = 64;

      reg clk = 1'b0;
      reg rst_n = 1'b0;
      always #5 clk = ~clk;

      integer errors = 0, checks = 0, neg_detected = 0, negative_mode = 0;

      // ---------------------------------------------------------------- DUT A
      // Driven directly, so every bin including the unreachable-on-a-real-bus ones
      // can be entered deliberately.
      reg  a_all_released = 1'b1;
      reg  a_line         = 1'b1;
      wire [CNT_W-1:0] a_fast, a_slow, a_abort, a_stuck, a_max, a_first_slow, a_rel;
      wire a_measuring;

      i2c_rise_classifier #(.TR_BUDGET(TR_BUDGET), .STUCK_LIMIT(STUCK_LIMIT), .CNT_W(CNT_W))
      dut_a (
         .clk(clk), .rst_n(rst_n),
         .all_released(a_all_released), .line(a_line),
         .n_fast(a_fast), .n_slow(a_slow), .n_aborted(a_abort), .n_stuck(a_stuck),
         .max_latency(a_max), .first_slow_release(a_first_slow), .n_releases(a_rel),
         .measuring(a_measuring)
      );

      // ---------------------------------------------------------------- DUT B
      // Behind a real RC line, so the classifier is exercised against a bus rather
      // than against a hand-drawn waveform. Two of them, at two rise times, which is
      // the experiment in section 6.
      reg  b_drive = 1'b0;
      wire b_line_fast, b_line_slow;
      wire [CNT_W-1:0] bf_fast, bf_slow, bf_abort, bf_stuck, bf_max, bf_first, bf_rel;
      wire [CNT_W-1:0] bs_fast, bs_slow, bs_abort, bs_stuck, bs_max, bs_first, bs_rel;

      tb_rc_line #(.RISE_CLKS(2)) line_fast (.clk(clk), .drive_low(b_drive), .line(b_line_fast));
      tb_rc_line #(.RISE_CLKS(9)) line_slow (.clk(clk), .drive_low(b_drive), .line(b_line_slow));

      i2c_rise_classifier #(.TR_BUDGET(TR_BUDGET), .STUCK_LIMIT(STUCK_LIMIT), .CNT_W(CNT_W))
      dut_bf (
         .clk(clk), .rst_n(rst_n), .all_released(~b_drive), .line(b_line_fast),
         .n_fast(bf_fast), .n_slow(bf_slow), .n_aborted(bf_abort), .n_stuck(bf_stuck),
         .max_latency(bf_max), .first_slow_release(bf_first), .n_releases(bf_rel),
         .measuring()
      );
      i2c_rise_classifier #(.TR_BUDGET(TR_BUDGET), .STUCK_LIMIT(STUCK_LIMIT), .CNT_W(CNT_W))
      dut_bs (
         .clk(clk), .rst_n(rst_n), .all_released(~b_drive), .line(b_line_slow),
         .n_fast(bs_fast), .n_slow(bs_slow), .n_aborted(bs_abort), .n_stuck(bs_stuck),
         .max_latency(bs_max), .first_slow_release(bs_first), .n_releases(bs_rel),
         .measuring()
      );

      // ---------------------------------------------------------------- DUT C
      // A NARROW instance. At CNT_W = 16 the counters saturate at 65535, which no
      // reasonable run reaches, so the saturation the header claims would be an
      // untested claim and a mutation replacing it with wrapping would survive
      // every check. At CNT_W = 4 the bound is 15 and forty releases reach it.
      // Moving the bound is cheaper than lengthening the run, and the configuration
      // that makes it reachable is a legal one somebody will eventually instantiate.
      reg  c_all_released = 1'b1;
      reg  c_line         = 1'b1;
      wire [3:0] c_fast, c_slow, c_abort, c_stuck, c_max, c_first, c_rel;

      i2c_rise_classifier #(.TR_BUDGET(TR_BUDGET), .STUCK_LIMIT(STUCK_LIMIT), .CNT_W(4))
      dut_c (
         .clk(clk), .rst_n(rst_n),
         .all_released(c_all_released), .line(c_line),
         .n_fast(c_fast), .n_slow(c_slow), .n_aborted(c_abort), .n_stuck(c_stuck),
         .max_latency(c_max), .first_slow_release(c_first), .n_releases(c_rel),
         .measuring()
      );

      // ------------------------------------------------------------------ helpers
      // ── STIMULUS TIMING DISCIPLINE ──────────────────────────────────────────
      // Every wait ends 1 ns AFTER the clock edge, so every stimulus assignment
      // that follows it lands strictly between edges.
      //
      // The first version of this bench did not do that: it drove DUT inputs with
      // blocking assignments at the same simulation instant as the posedge. Whether
      // the DUT's always block saw the old value or the new one then depended on
      // which process the scheduler happened to run first, and the symptom was a
      // measured latency short by one -- for SOME values of the latency and not
      // others. A release programmed for 11 clocks measured 11; one programmed for
      // 14 measured 13; one for 2 measured 1.
      //
      // That pattern is worth recognising because of how it presents. An off-by-one
      // that appears at some values and not others looks exactly like a design bug
      // with an obscure trigger, and invites a search of the design's arithmetic.
      // The tell is that the values with the error have no arithmetic relationship
      // to each other -- they differ by how many edges the preceding stimulus
      // consumed, which is a property of the BENCH.
      task tick (input integer n); integer k;
         begin for (k = 0; k < n; k = k + 1) @(posedge clk); #1; end
      endtask

      // One clock edge, leaving the simulation 1 ns past it.
      task step; begin @(posedge clk); #1; end endtask

      task chk (input [255:0] name, input integer got, input integer exp);
         begin
            checks = checks + 1;
            if (got !== exp) begin
               if (negative_mode) neg_detected = neg_detected + 1;
               else begin
                  errors = errors + 1;
                  $display("  FAIL %0s: got %0d expected %0d", name, got, exp);
               end
            end else if (negative_mode) begin
               errors = errors + 1;
               $display("  FAIL negative proof did not fire: %0s", name);
            end
         end
      endtask

      task do_reset;
         begin
            rst_n = 1'b0; a_all_released = 1'b1; a_line = 1'b1; b_drive = 1'b0;
            tick(3); rst_n = 1'b1; tick(2);
         end
      endtask

      // One release on DUT A, arranged so that `lat` is EXACTLY the latency the
      // classifier will measure.
      //
      // Getting this right needs care, and the first version of it did not. Writing
      // "hold the line low for lat clocks after releasing" produced a measured
      // latency of lat-1, because the edge at which the classifier SEES the release
      // is itself one of the clocks the loop counted. The symptom was three failures
      // -- a max-latency check, and both halves of the budget-boundary test -- none
      // of which was a defect in the classifier. A measurement helper whose argument
      // does not mean what its name says produces failures that all look like
      // off-by-one bugs in the design.
      task a_release_after (input integer lat); integer k;
         begin
            a_all_released = 1'b0; a_line = 1'b0; tick(2);
            a_all_released = 1'b1;
            step;                                    // the release edge itself
            for (k = 0; k < lat; k = k + 1) step;    // lat edges with the line still low
            a_line = 1'b1;                           // read high on the next edge
            tick(2);
         end
      endtask

      // One bus transaction on DUT B: pull the line low for `low_clks`, then release
      // and wait long enough for the slowest model to finish rising.
      task b_pulse (input integer low_clks);
         begin b_drive = 1'b1; tick(low_clks); b_drive = 1'b0; tick(20); end
      endtask

      integer i;
      initial begin
         $display("i2c_rise_classifier_tb");

         // ---------------------------------------------------------------------
         // T1 -- the four bins, each entered deliberately. Two of them cannot occur
         // on a healthy bus, which is exactly why the bench drives the classifier's
         // inputs directly rather than only through a bus model.
         // ---------------------------------------------------------------------
         $display("T1 each bin entered deliberately");
         do_reset;
         a_release_after(2);                         // within the 4-clock budget
         chk("T1 a fast release is FAST", a_fast, 1);
         chk("T1 nothing else counted", a_slow + a_abort + a_stuck, 0);

         a_release_after(9);                         // over budget, but it does rise
         chk("T1 a slow release is SLOW", a_slow, 1);
         chk("T1 fast count unchanged", a_fast, 1);
         chk("T1 max latency is exactly 9", a_max, 9);
         // Releases are numbered from 1, so 0 unambiguously means "no slow release
         // has occurred" without needing a second flag to say so.
         chk("T1 first slow release is the 2nd", a_first_slow, 2);

         // Aborted: release, then pull again before the line has risen.
         a_all_released = 1'b0; a_line = 1'b0; tick(2);
         a_all_released = 1'b1; tick(2);
         a_all_released = 1'b0; tick(2);             // pulled again, still low
         a_line = 1'b0; a_all_released = 1'b1; tick(1); a_line = 1'b1; tick(2);

         chk("T1 an interrupted rise is ABORTED", a_abort, 1);

         // Stuck: everybody released and the line stays low past STUCK_LIMIT.
         a_all_released = 1'b0; a_line = 1'b0; tick(2);
         a_all_released = 1'b1;
         tick(STUCK_LIMIT + 4);
         chk("T1 a line that never rises is STUCK", a_stuck, 1);
         a_line = 1'b1; tick(2);

         // The bins must account for every release. A classifier whose bins do not
         // sum to the release count is dropping or double-counting events, and
         // neither shows up as a wrong individual bin.
         chk("T1 bins sum to releases",
             a_fast + a_slow + a_abort + a_stuck, a_rel);
         $display("      fast=%0d slow=%0d aborted=%0d stuck=%0d releases=%0d max=%0d",
                  a_fast, a_slow, a_abort, a_stuck, a_rel, a_max);

         // ---------------------------------------------------------------------
         // T2 -- the boundary. TR_BUDGET is 4, so exactly 4 clocks must be FAST and
         // exactly 5 must be SLOW. A budget whose boundary is never tested is a
         // number the instrument has never actually applied.
         // ---------------------------------------------------------------------
         $display("T2 the budget boundary, at exactly TR_BUDGET and one beyond");
         do_reset;
         a_release_after(TR_BUDGET);
         chk("T2 exactly at budget is FAST", a_fast, 1);
         chk("T2 exactly at budget is not SLOW", a_slow, 0);
         a_release_after(TR_BUDGET + 1);
         chk("T2 one clock over budget is SLOW", a_slow, 1);
         chk("T2 fast count unchanged", a_fast, 1);

         // ---------------------------------------------------------------------
         // T3 -- a release onto a line that is already high. Zero-length rise, which
         // is what an IDEAL bus model produces on every release, and it must be
         // counted rather than silently ignored.
         // ---------------------------------------------------------------------
         $display("T3 a zero-length rise");
         do_reset;
         a_all_released = 1'b0; a_line = 1'b1; tick(2);   // an ideal bus: never low
         a_all_released = 1'b1;
         step;
         // Sampled ON the release edge, not after it. A classifier that enters
         // measurement and leaves again on the next clock produces the same counts
         // and a one-cycle `measuring` pulse that nothing downstream expects -- and
         // a check taken three cycles later cannot see the difference.
         chk("T3 measurement never started", a_measuring, 0);
         tick(3);
         chk("T3 counted as a release", a_rel, 1);
         chk("T3 counted FAST", a_fast, 1);
         chk("T3 max latency is zero", a_max, 0);
         chk("T3 no measurement left in flight", a_measuring, 0);

         // ---------------------------------------------------------------------
         // T4 -- THE EXPERIMENT. One drive pattern, two buses. The pattern is
         // identical, the classifier is identical, the budget is identical; the only
         // difference is the rise time of the line. A verdict that changes under
         // that single variable is a verdict about the bus.
         // ---------------------------------------------------------------------
         $display("T4 one drive pattern, two rise times");
         do_reset;
         for (i = 0; i < 20; i = i + 1) b_pulse(3 + (i % 4));
         $display("      RISE_CLKS=2:  fast=%0d slow=%0d stuck=%0d max=%0d releases=%0d",
                  bf_fast, bf_slow, bf_stuck, bf_max, bf_rel);
         $display("      RISE_CLKS=9:  fast=%0d slow=%0d stuck=%0d max=%0d releases=%0d",
                  bs_fast, bs_slow, bs_stuck, bs_max, bs_rel);
         chk("T4 both buses saw the same releases", bf_rel, bs_rel);
         chk("T4 the 2-clock bus is entirely FAST", bf_slow, 0);
         chk("T4 the 2-clock bus counted 20", bf_fast, 20);
         chk("T4 the 9-clock bus is entirely SLOW", bs_fast, 0);
         chk("T4 the 9-clock bus counted 20", bs_slow, 20);
         chk("T4 neither bus is stuck", bf_stuck + bs_stuck, 0);
         chk("T4 max latency separates them", (bs_max > bf_max) ? 1 : 0, 1);
         chk("T4 the slow bus was out of budget from release 1", bs_first, 1);
         chk("T4 the fast bus never recorded a slow release", bf_first, 0);

         // ---------------------------------------------------------------------
         // T5 -- the strong negative. The whole value of the fast result is that
         // `n_slow` is zero over a long window, so the window has to be long and it
         // has to be reported. A zero over three releases is not evidence.
         // ---------------------------------------------------------------------
         $display("T5 the strong negative needs a long window");
         do_reset;
         for (i = 0; i < 200; i = i + 1) b_pulse(2 + (i % 5));
         $display("      RISE_CLKS=2 over 200 releases: fast=%0d slow=%0d", bf_fast, bf_slow);
         chk("T5 200 releases observed", bf_rel, 200);
         chk("T5 not one of them was slow", bf_slow, 0);
         chk("T5 bins account for all of them", bf_fast + bf_slow + bf_abort + bf_stuck, bf_rel);

         // ---------------------------------------------------------------------
         // T7 -- max_latency is the GREATEST, not the most recent. Every earlier
         // test released at a constant latency, where "last" and "greatest" are the
         // same number, so the distinction was never under test. A slow release
         // followed by a fast one separates them.
         // ---------------------------------------------------------------------
         $display("T7 max latency is the greatest, not the last");
         do_reset;
         a_release_after(11);
         chk("T7 the slow release set max to 11", a_max, 11);
         a_release_after(1);
         chk("T7 a later FAST release does not lower max", a_max, 11);
         a_release_after(2);
         chk("T7 nor does another", a_max, 11);
         a_release_after(14);
         chk("T7 a larger one does raise it", a_max, 14);

         // ---------------------------------------------------------------------
         // T8 -- the saturation bound, on the narrow instance. A diagnostic counter
         // that wraps reports a small number for a great many events, and a small
         // number is what a nearly-clean bus reports.
         // ---------------------------------------------------------------------
         $display("T8 counter saturation at CNT_W=4 (max 15)");
         do_reset;
         for (i = 0; i < 40; i = i + 1) begin
            c_all_released = 1'b0; c_line = 1'b0; tick(2);
            c_all_released = 1'b1; step;
            c_line = 1'b1; tick(2);
         end
         $display("      narrow instance: releases=%0d fast=%0d", c_rel, c_fast);
         chk("T8 release counter saturated at 15", c_rel, 15);
         chk("T8 fast counter saturated at 15", c_fast, 15);
         chk("T8 neither wrapped to a small value", (c_rel >= 15 && c_fast >= 15) ? 1 : 0, 1);
         chk("T8 the other bins stayed empty", c_slow + c_abort + c_stuck, 0);

         // ---------------------------------------------------------------------
         // T6 -- negative proof. Six deliberately wrong expectations, one per check
         // shape used above.
         // ---------------------------------------------------------------------
         $display("T6 negative proof -- eight deliberately wrong expectations");
         // The negative proof establishes its OWN state rather than borrowing
         // whatever the previous test left behind. An earlier arrangement read
         // counters that a later test's reset had cleared, so two of the eight
         // "deliberately wrong" expectations of zero were accidentally correct and
         // silently stopped proving anything -- while the block's own count still
         // looked healthy. A negative proof built on state it does not control is
         // one refactor away from being decorative.
         do_reset;
         a_release_after(2);    // FAST
         a_release_after(7);    // SLOW
         chk("T6 setup: one fast", a_fast, 1);
         chk("T6 setup: one slow", a_slow, 1);
         chk("T6 setup: max is 7", a_max, 7);
         chk("T6 setup: narrow instance was reset", c_rel, 0);

         negative_mode = 1;
         chk("T6a wrong bin count",        a_fast, 7);
         chk("T6b wrong zero claim",       a_slow, 3);
         chk("T6c wrong release total",    a_rel, 199);
         chk("T6d wrong bin sum",          a_fast + a_slow, 0);
         chk("T6e wrong max latency",      a_max, 0);
         chk("T6f wrong first-slow index", a_first_slow, 19);
         chk("T6g wrong boundary claim",   a_max, 3);
         chk("T6h wrong saturation bound", c_rel, 40);
         negative_mode = 0;
         chk("T6 all eight negatives detected", neg_detected, 8);

         $display("");
         $display("checks=%0d errors=%0d negatives_detected=%0d/8", checks, errors, neg_detected);
         if (errors == 0) $display("RESULT: PASS"); else $display("RESULT: FAIL");
         $finish;
      end

      initial begin
         #2000000;
         $display("RESULT: FAIL -- watchdog expired, the run did not terminate");
         $finish;
      end

   endmodule

The bench drives two things. DUT A has its inputs driven directly, so that every bin — including the two that cannot occur on a healthy bus — can be entered deliberately and the budget boundary can be tested at exactly TR_BUDGET and one clock beyond. DUT B sits behind an actual RC line model, so the classifier is exercised against a bus rather than against a hand-drawn waveform.

T4 is the chapter in one test

Azvya Education Pvt. Ltd.VLSI Mentor
T4 — ONE drive pattern, TWO rise times
   The same twenty pulses, the same classifier, the same TR_BUDGET of 4 clocks.
   The ONLY difference between the two rows is the rise time of the line.

      RISE_CLKS = 2:   fast=20  slow=0   stuck=0  max=1   releases=20
      RISE_CLKS = 9:   fast=0   slow=20  stuck=0  max=8   releases=20

   Both buses saw exactly the same twenty releases. One is entirely inside its
   budget and one is entirely outside it, and nothing about the traffic changed.

That is a controlled experiment in the strict sense: one variable moved, everything else held, and the verdict moved with it. It is worth contrasting with the alternative that feels equally reasonable — run the fast bus, then run a different pattern on the slow bus, and compare. That comparison would confound the bus with the traffic, and any difference could be attributed to either.

T5 is the result you actually want

Azvya Education Pvt. Ltd.VLSI Mentor
T5 — the strong negative
   RISE_CLKS = 2 over 200 releases:   fast=200   slow=0

Zero slow releases over two hundred, with the bins accounting for all two hundred, says the bus is electrically sound at this speed — and therefore that a protocol symptom is a protocol bug. That single negative eliminates an entire layer.

The window length is not decoration. A zero over three releases is not evidence of anything, and the number of releases has to be reported alongside the zero for the zero to mean anything at all. This is the same point as Chapter 23.1's n_consistent: a count of zero events is only informative next to a count of how many opportunities there were.

5. The Bench Race, and Why It Looked Like a Design Bug

The first run reported five failures. Two were a naming problem — a helper whose argument counted one edge more than its name implied — and the interesting one is what remained after that was fixed:

Azvya Education Pvt. Ltd.VLSI Mentor
A MEASURED LATENCY THAT WAS WRONG FOR SOME VALUES AND NOT OTHERS
   programmed 11 clocks  ->  measured 11      correct
   programmed  1 clock   ->  measured  1      correct
   programmed  2 clocks  ->  measured  1      one short
   programmed 14 clocks  ->  measured 13      one short

Read that table before reading on. An off-by-one is a familiar thing and the reflex is to look for an inclusive-versus-exclusive comparison in the design. But the reflex does not survive the table: an inclusive/exclusive error is off by one always, not at 2 and 14 and not at 1 and 11. And 2 and 14 have no arithmetic relationship to each other — not parity in any meaningful sense, not a power of two, nothing the design's comparison could be keying on.

When the values that fail have no arithmetic relationship, the variable is not arithmetic. It is sequence: how many clock edges the preceding stimulus happened to consume.

The cause was the bench driving DUT inputs with blocking assignments at the same simulation instant as the clock edge. Whether the DUT's always @(posedge clk) block saw the old value or the new one then depended on which process the scheduler ran first at that instant — and that ordering shifted with what the previous test had done.

The fix is a discipline rather than a patch: every wait in the bench ends one nanosecond after the clock edge, so every stimulus assignment lands strictly between edges and no DUT input ever changes at a sampling instant.

And one more, found by the mutation campaign's own bookkeeping

Adding tests later in the file moved a do_reset in front of the negative-proof block, which read counters from a DUT that reset had just cleared. Two of the eight deliberately-wrong expectations were expectations of zero — and against a freshly reset DUT they were accidentally correct, so they stopped firing.

The block's own count caught it, because it asserts how many negatives were detected rather than assuming they all were. Without that assertion the negative proof would have quietly dropped from eight cases to six while continuing to print a healthy-looking line.

A negative proof built on state it does not control is one refactor away from being decorative. The fix is that the block now establishes its own state before using it.

6. Mutation

Fifteen mutations, each written to its own path, that path hashed, the compiler invoked on that exact path, and the harness refusing to score any mutation whose anchor did not appear exactly once.

Azvya Education Pvt. Ltd.VLSI Mentor
MUTATION RESULTS
   FIRST CAMPAIGN     15 valid   9 KILLED   4 SURVIVED   (1 anchor did not match)
   AFTER T3, T7, T8   15 valid  15 KILLED   0 SURVIVED

The four survivors were four different holes, and none of them was in the design:

R01 — the budget comparison. Changing > to >= makes a latency of exactly TR_BUDGET count as slow. The bench tested a fast case and a slow case and never the boundary, so the budget was a number that had never actually been applied at the point where it decides anything.

R05 — max latency records the last instead of the greatest. Every test up to that point released at a constant latency, where "last" and "greatest" are the same number. T7 separates them with a slow release followed by two fast ones.

R07 — a release onto an already-high line is ignored. This one is subtle: the mutant still counts the release as FAST, one clock later, via the measurement path. The only observable difference is a one-cycle pulse on measuring, and the original check sampled measuring three cycles afterwards, by which time the pulse was over. T3 now samples it on the release edge.

R13 — counters wrap instead of saturating. CNT_W is 16, the bound is 65535, and no test came near it. T8 adds a narrow instance at CNT_W = 4, where the bound is 15 and forty releases reach it comfortably.

7. What the Split Cannot Settle

It cannot tell you why a rise is slow. Resistor value, bus capacitance, a leaky device, a level shifter with a slow output, a long cable — all produce SLOW, and separating them needs Chapter 23.5.

It cannot see anything analogue. There is no ramp here, no threshold and no curve — only "for this many clocks, a sampler still read low". Two devices may cross their thresholds at different instants, and nothing in simulation models that.

A clean electrical result does not make the protocol right. It eliminates a layer, which is exactly what it is for. The protocol bug still has to be found, and Chapter 23.4 onwards is that work.

all_released is only as good as the caller's knowledge. On hardware it is usually this device's own drive intent, so every participant the caller does not know about lands in STUCK rather than being subtracted out. That is the right failure direction — an unknown holder should be conspicuous — but it means STUCK is a bin about your model of the bus, not only about the bus.

And it cannot model metastability. Nothing in RTL simulation can. A slow edge sampled by a real flip-flop can produce a metastable output whose resolution time is unbounded in principle; the model here resolves to a definite 0 or 1 every time. Chapter 19.4 makes the structural argument; 23.8 returns to what that means for debugging.

8. Misconceptions

9. Debug Lab

Two units in fifty fail, and only after the enclosure is fitted

A symptom with a mechanical trigger, and the measurement that stops it being a mystery
Buggy Code
// A production run of fifty boards. An I2C bus at 400 kHz with an EEPROM, a
// temperature sensor and a port expander. All fifty pass functional test on the
// bench. Two of them fail the final test AFTER the enclosure is fitted, with
// sporadic CRC errors on EEPROM reads.
//
// Removing the enclosure makes the failure go away on both units. Refitting it
// brings it back. That is a strong, repeatable correlation and it is NOT yet a
// cause -- fitting an enclosure changes several things at once:
//
//   (a) it adds capacitance: a grounded metal lid over the bus traces is a
//       plate, and Cb goes up
//   (b) it changes the thermal environment: the enclosed board runs warmer, and
//       resistors, pull-ups and device thresholds all move with temperature
//   (c) it applies mechanical stress: a board screwed into a chassis flexes, and
//       a marginal solder joint can open
//   (d) it changes the grounding: the lid may or may not be bonded, and a
//       partially-bonded lid is an antenna
//
// All four are consistent with "sporadic CRC errors on EEPROM reads", and the
// enclosure correlation does not separate them. Note also that a CRC error is
// several layers above anything physical -- it says a block of bytes did not
// match its checksum, which is downstream of every one of these.
Symptom

Step 3 of Chapter 23.1's workflow, on one failing unit and one passing one, with the classifier instantiated on SDA and TR_BUDGET derived from Fast-mode's 300 ns limit at the sample clock in use.

PASSING unit, enclosure fitted, 50,000 releases: fast=49,997 slow=0 aborted=3 stuck=0 max_latency = 214 ns

FAILING unit, enclosure REMOVED, 50,000 releases: fast=49,996 slow=0 aborted=4 stuck=0 max_latency = 288 ns

FAILING unit, enclosure FITTED, 50,000 releases: fast=48,102 slow=1,894 aborted=4 stuck=0 max_latency = 460 ns first_slow_release = 1

Three numbers do the work here, and two of them are the ones that did NOT change.

stuck = 0 in all three runs. Nothing is holding the line. Candidate (c), a solder joint opening under mechanical stress, would show as STUCK or as a large ABORTED count, and does neither.

aborted is 3 or 4 in all three runs -- unchanged. So the traffic pattern is the same in all three, and the comparison is between like and like. Without this, a change in the slow count could be a change in what was being measured.

max_latency, with the enclosure fitted on the failing unit, is 460 ns against a 300 ns budget. The PASSING unit with the same enclosure is at 214 ns.

And first_slow_release = 1 says the very first release of the run was already over budget. This is not an event that happens occasionally; it is the bus's steady state with the lid on.

Root Cause

The failing unit's rise time is out of Fast-mode budget with the enclosure fitted and inside it without. The passing unit is inside budget either way, at 214 ns against 300.

That localises to the electrical layer and kills the protocol hypotheses outright: the EEPROM, the driver, the CRC routine and the controller RTL are identical across all fifty boards and cannot be a function of whether a lid is on.

It does NOT yet distinguish (a) added capacitance from (b) temperature, and both are still live -- an enclosure does both, and either can add 170 ns.

Two measurements separate them, and each moves ONE variable:

fit the lid but hold the board at bench temperature (forced air) -- if the rise time stays at 460 ns, it is capacitance, not temperature

remove the lid and raise the board to the enclosed operating temperature -- if the rise time stays near 288 ns, temperature is not the mechanism

On these units the first measurement returned 455 ns and the second 291 ns. So it is (a), capacitance, and temperature contributes a few nanoseconds at most.

Why only two units? Because 288 ns with the lid OFF is already at 96% of the budget before the enclosure adds anything. Reading the rest of the batch:

forty-eight units, lid off: max_latency 180 ns to 230 ns the two failures, lid off: 284 ns and 288 ns

The two failures are outliers BEFORE the enclosure is involved. The enclosure is not the cause; it is the last contribution that pushes an already-marginal bus over. That reframing matters, because "the enclosure causes it" leads to mechanical changes and "the bus has no margin" leads to the resistor.

Measuring the pull-ups on the two outliers found 4.7 kilohm where the design called for 2.2 kilohm -- a reel-loading error on one machine, affecting a contiguous run of two boards.

CORRECTION: replace the resistors on the two units, and add a max_latency check to the functional test so the OTHER forty-eight are characterised rather than assumed. The second half is the real correction. Forty-eight units at 180 to 230 ns against a 300 ns budget is a distribution with margin; two at 288 ns was a distribution with an unmeasured tail, and nothing in the test would have caught a third.

PROOF, in an order where each step is attributable:

1. before: max_latency 460 ns with the lid on, 288 ns with it off, and CRC errors with the lid on 2. after the resistor change, lid off: max_latency 198 ns -- inside the range of the other forty-eight, which is the check that the right thing changed 3. after the resistor change, lid on: max_latency 226 ns, slow = 0 over 50,000 releases, with the budget UNCHANGED at 300 ns 4. and the CRC errors stop

Step 4 is last on purpose. The symptom disappearing is the weakest evidence in the list -- it is consistent with the fix and with several other things. Step 3 is the one that proves it, because the budget is the value it always was, so the change in verdict is attributable to the bus rather than to a threshold that was moved until the complaint stopped.

10. Reason It Through

11. Questions

12. What This Chapter Settled

Five pairs of symptoms that a protocol cause and an electrical cause produce identically, and the reason reading bytes cannot separate them: the bytes are the same. The discriminating information is the time from release to a readable HIGH, which no protocol decoder reports.

An instrument that measures it, with four bins rather than two — because an interrupted rise is not a measurement and a line that never rises is not a slow one — and whose bins sum to the release count, checked. Fifty-one checks, zero errors and eight negative proofs, passing under both Verilog standards. Fifteen mutations, fifteen killed.

A controlled experiment: twenty identical releases against two buses differing only in rise time, producing 20 fast / 0 slow against 0 fast / 20 slow. And a strong negative — zero slow releases in two hundred — which is the result worth wanting, because it eliminates a layer rather than nominating a suspect.

Two bench defects worth more than the tests that found them. A stimulus race that produced a latency one short for some programmed values and not others, whose tell was that the failing values had no arithmetic relationship to each other. And a negative-proof block that silently lost two of its eight cases to a do_reset added later, caught only because it asserts its own detection count.

And the four mutation survivors, which cluster where the previous two chapters' did: on boundaries and instants. The logic gets tested because tests are written from what a block does. The boundary and the sampling moment get missed because they are properties of when.

With the electrical layer either cleared or indicted, the next question is the one every I²C engineer meets first and most often: an acknowledge that should have been there and was not. Chapter 23.4 works that symptom through its many causes, and it is built entirely out of discriminating experiments.

Continue learning

Related tutorials