I²C · Module 23
Stuck Bus — Diagnosis and Recovery
A bus held low, who is holding it, and why the standard recovery works for exactly one of its seven causes. Builds a clear sequencer whose pulse count is evidence rather than a boolean, runs it against seven holders including a legally stretching target, and reports the four bugs it took to get there — two in the design and two in the bench.
Three chapters have now deferred the same symptom. 23.3's classifier has a STUCK bin it explicitly declines to interpret; 23.5's profiler has a hold counter for the same reason; 23.4's diagnostic has a verdict that says only "something is answering everything".
All three are pointing here: a line that stays low with nobody who should be driving it. It is the most alarming I²C failure because the bus is completely unusable, and it is one of the more tractable, because the specification provides a recovery — for exactly one of its causes.
1. Who Can Hold a Line Low
The wired-AND (Chapter 2.5) means a line is low if any participant pulls it, and no participant can lift it. So the question is never "why is the line low" but who, and the candidates behave very differently:
| holder | why | does clocking help? |
|---|---|---|
| a target interrupted mid-byte | it is waiting for the clock pulses that will shift out the rest of its byte | yes — this is the case recovery exists for |
| a target whose state machine is wedged | it is not waiting for anything | no |
| a device held in reset with its pad asserted | its output is driven by reset state, not by protocol | no |
| a stuck or shorted pad | electrical | no |
| a short to ground | electrical | no |
| a controller holding SCL | usually its own error path, or a stopped clock | no — and it makes everything else unanswerable |
| a target stretching SCL | perfectly legal, and temporary | nothing to fix |
The last row is why a stuck-bus diagnosis cannot be made instantly. A bus that has been low for 3 µs may be a target stretching; a bus that has been low for 3 ms is not. The difference is a timeout, and choosing it is a real decision — Chapter 12.4 is the argument.
2. Why Clocking Works, When It Works
The mechanism is specific and it is worth stating precisely, because the folk version — "clock the bus to unstick it" — suggests the pulses are doing something physical.
A target transmitting a byte shifts one bit onto SDA per clock pulse, and it changes SDA while SCL is low, because changing it while SCL is high is by definition a START or a STOP. If a controller stops clocking part-way through a byte — a reset, a watchdog, a crash — the target is left holding whatever bit it last shifted out. If that bit is a zero, SDA stays low, and the target will hold it there indefinitely, because it is waiting for a clock pulse that is its entitlement and not its choice.
Give it the pulses and it finishes the byte, reaches its acknowledge slot, releases, and the bus is free.
Three pulses, a release, and the STOP
10 cyclesTwo things in that figure are load-bearing and are usually left out.
The target releases while SCL is LOW. If it released during a high period, the wire would carry a rising SDA with SCL high — which is a STOP condition, generated accidentally, by a device that was only trying to send a one. A conforming target cannot do that; a target with a timing defect can, and Chapter 23.4's F_GATED is its close relative.
The STOP at the end is part of the recovery. When SDA comes free, every device that was mid-transfer is in an arbitrary framing state, and some will read the next START as a continuation. A STOP is the one event that returns every conforming device to idle. A recovery that omits it leaves a bus that is electrically fine and logically inconsistent — which produces a second, stranger failure some transfers later.
3. The Sequencer
// -----------------------------------------------------------------------------
// i2c_bus_clear.sv
// Chapter 23.6's instrument, and the only recovery mechanism I2C has.
//
// Chapter 15.4 established the mechanism: a target that was interrupted part-way
// through transmitting a byte is holding SDA low because it is waiting for the
// clock pulses that will shift the rest of its byte out. Give it those pulses and
// it finishes, releases, and the bus is usable again. Chapter 17.11 built it into
// a master's error path.
//
// This version adds what a DIAGNOSTIC needs, which is not more recovery but more
// evidence about what happened:
//
// result which of three distinguishable outcomes occurred
// pulses_used how many clock pulses the line needed before it let go
//
// WHY pulses_used IS THE INTERESTING OUTPUT. A target interrupted mid-byte needs
// at most nine more pulses -- one per remaining bit, plus the acknowledge slot --
// and the number it actually needs says HOW FAR INTO A BYTE it was when the
// transfer was interrupted. Two pulses means it was almost finished. Eight means
// it had barely started. And a recovery that succeeds after ZERO pulses did not
// recover anything: the line was already released and something else was wrong.
//
// THE THREE OUTCOMES, and only one of them is a success:
//
// R_OK SDA released within MAX_PULSES, and a STOP was then generated so
// every device on the bus resets its framing state.
//
// R_SCL_HELD SCL is being held low by something other than this block. This is
// NOT RECOVERABLE by this mechanism and it is important to say so
// rather than to keep trying: the recovery works by clocking, and
// a master that cannot raise SCL cannot clock. Chapter 15.4 says
// the same thing from the specification's side.
//
// R_SDA_HELD MAX_PULSES were delivered and SDA is still low. Whatever is
// holding it is not a target waiting for clocks -- a stuck pad, a
// device held in reset with its output asserted, a short, or a
// target whose state machine is genuinely wedged rather than
// merely paused.
//
// WHY THE STOP IS PART OF THE RECOVERY AND NOT AN AFTERTHOUGHT. When SDA comes
// free, every device that was mid-transfer is in an arbitrary framing state, and
// some of them will interpret the next START as a continuation. A STOP is the one
// event that returns every conforming device to idle, and a recovery that omits
// it leaves a bus that is electrically fine and logically inconsistent.
//
// OPEN-DRAIN THROUGHOUT. This block pulls low or releases. It never drives high,
// including during recovery -- a recovery that drove SDA high to "unstick" it
// would be contention with whatever is holding the line down, which is the one
// thing an I2C participant must never do.
// -----------------------------------------------------------------------------
module i2c_bus_clear #(
// Clock pulses to deliver before giving up. Nine is the specification's
// figure and the arithmetic behind it is eight data bits plus the
// acknowledge slot -- the most a target can still be waiting for.
parameter integer MAX_PULSES = 9,
// Sample clocks per half-period of the recovery clock. The recovery clock
// should be no faster than the bus's normal rate, because a target that was
// interrupted is not obliged to be fast.
parameter integer PHASE = 4,
// How long to wait for SCL to actually read HIGH after releasing it, before
// concluding that somebody is holding it.
//
// This parameter exists because of a bug. The first version checked that SCL
// had risen IN THE SAME CYCLE it released it -- and reported "SCL is held" on
// a perfectly healthy bus, every time, because a released line does not read
// high until at least the next sample and on a real bus not until tr has
// elapsed. It is the same "released means high" error that Chapter 19.1's
// DebugLab is about, arriving inside the block whose job is to recover from a
// bus that is genuinely held.
//
// The allowance also has to cover CLOCK STRETCHING, because a target that was
// interrupted mid-byte is entitled to hold SCL while it catches up (Chapter
// 12.2). A recovery that treats a legal stretch as a stuck clock abandons
// exactly the case it was built for.
parameter integer STRETCH_LIMIT = 64
) (
input wire clk,
input wire rst_n,
input wire start, // one-cycle request
// The resolved bus, read back.
input wire scl_in,
input wire sda_in,
// This block's contribution to the wired-AND.
output reg scl_drive_low,
output reg sda_drive_low,
output reg busy,
output reg done, // one cycle when a result is available
output reg [3:0] pulses_used,
output reg [1:0] result
);
localparam [1:0] R_OK = 2'd0, R_SCL_HELD = 2'd1, R_SDA_HELD = 2'd2;
localparam [3:0] S_IDLE = 4'd0,
S_CHECK = 4'd1,
S_CLK_LOW = 4'd2,
S_CLK_RISE = 4'd3,
S_CLK_HIGH = 4'd4,
S_STOP_A = 4'd5,
S_STOP_B = 4'd6,
S_STOP_C = 4'd7,
S_DONE = 4'd8;
reg [3:0] state;
reg [15:0] ph;
reg [15:0] wait_cnt;
wire phase_done = (ph + 1) >= PHASE[15:0];
always @(posedge clk) begin
if (!rst_n) begin
state <= S_IDLE;
ph <= 16'd0;
scl_drive_low <= 1'b0;
sda_drive_low <= 1'b0;
busy <= 1'b0;
done <= 1'b0;
pulses_used <= 4'd0;
result <= R_OK;
end else begin
done <= 1'b0;
case (state)
S_IDLE: begin
scl_drive_low <= 1'b0;
sda_drive_low <= 1'b0;
busy <= 1'b0;
if (start) begin
busy <= 1'b1;
pulses_used <= 4'd0;
ph <= 16'd0;
wait_cnt <= 16'd0;
state <= S_CHECK;
end
end
// Before anything else: can this block clock at all? It is not
// driving SCL, so a low SCL means somebody else is holding it, and
// no number of attempted pulses will change that. Reporting it
// immediately is the difference between a diagnosis and a hang.
// Before anything else: can this block clock at all? It is not
// driving SCL, so a low SCL means somebody else is holding it. The
// wait is bounded and generous: the line may simply still be rising,
// or a target may be stretching.
S_CHECK: begin
if (!scl_in) begin
if (wait_cnt >= STRETCH_LIMIT[15:0]) begin
result <= R_SCL_HELD;
state <= S_DONE;
end else wait_cnt <= wait_cnt + 16'd1;
end else if (sda_in) begin
// Already free. Still issue the STOP: the reason for recovery
// was a bus believed to be stuck, and the framing state of
// every device on it is unknown regardless of the line level.
state <= S_STOP_A;
end else begin
state <= S_CLK_LOW;
ph <= 16'd0;
end
end
S_CLK_LOW: begin
scl_drive_low <= 1'b1;
if (phase_done) begin
ph <= 16'd0;
wait_cnt <= 16'd0;
state <= S_CLK_RISE;
end else ph <= ph + 16'd1;
end
// Release SCL and wait for the line to actually read HIGH before the
// high phase begins to count. This is clock synchronisation (Chapter
// 13.2) rather than a delay: the high time a target gets is measured
// from when the clock is high on the WIRE, not from when this block
// stopped pulling.
S_CLK_RISE: begin
scl_drive_low <= 1'b0;
if (scl_in) begin
ph <= 16'd0;
state <= S_CLK_HIGH;
end else if (wait_cnt >= STRETCH_LIMIT[15:0]) begin
result <= R_SCL_HELD;
state <= S_DONE;
end else wait_cnt <= wait_cnt + 16'd1;
end
S_CLK_HIGH: begin
scl_drive_low <= 1'b0;
if (phase_done) begin
ph <= 16'd0;
pulses_used <= pulses_used + 4'd1;
// SDA is sampled at the END of the high phase, after the line
// has had a full phase to rise. Sampling at the start would
// read a line released in this pulse as still held, and the
// recovery would deliver pulses it did not need -- harmless
// for the bus, and it destroys `pulses_used` as evidence.
if (sda_in) state <= S_STOP_A;
else if (pulses_used + 4'd1 >= MAX_PULSES[3:0]) begin
result <= R_SDA_HELD;
state <= S_DONE;
end else state <= S_CLK_LOW;
end else begin
ph <= ph + 16'd1;
// SCL going low again DURING its high phase is another device
// pulling -- a target stretching, or a second controller
// synchronising. Neither is an error, so the block returns to
// waiting for the line rather than reporting a fault, and the
// bounded wait in S_CLK_RISE is what stops that being a hang.
if (!scl_in) begin
wait_cnt <= 16'd0;
state <= S_CLK_RISE;
end
end
end
// A STOP: SDA low while SCL is low, SCL released, then SDA released
// while SCL is high.
S_STOP_A: begin
scl_drive_low <= 1'b1;
sda_drive_low <= 1'b1;
if (phase_done) begin ph <= 16'd0; state <= S_STOP_B; end
else ph <= ph + 16'd1;
end
S_STOP_B: begin
scl_drive_low <= 1'b0;
if (phase_done) begin ph <= 16'd0; state <= S_STOP_C; end
else ph <= ph + 16'd1;
end
S_STOP_C: begin
sda_drive_low <= 1'b0;
if (phase_done) begin
ph <= 16'd0;
result <= R_OK;
state <= S_DONE;
end else ph <= ph + 16'd1;
end
S_DONE: begin
scl_drive_low <= 1'b0;
sda_drive_low <= 1'b0;
busy <= 1'b0;
done <= 1'b1;
state <= S_IDLE;
end
default: state <= S_IDLE;
endcase
end
end
endmoduleThe recovery mechanism itself is Chapter 15.4's and 17.11's. What this version adds is not more recovery but more evidence:
result distinguishes three outcomes, only one of which is a success. "The recovery ran" is not a finding; "the recovery ran and SDA is still held after nine pulses" is.
pulses_used says how far into a byte the transfer was interrupted. A target needs at most nine more pulses — eight bits plus the acknowledge slot — and the number it actually needs is a measurement of where it was. Two means it was nearly finished; eight means it had barely started. And a recovery reporting success after zero pulses did not recover anything: the line was already free, and whatever was wrong is still wrong.
4. The Bug That Made the Recovery Report a Stuck Clock on a Healthy Bus
The first version checked that SCL had risen in the same cycle it released it.
T2 a target interrupted mid-byte
FAIL T2 result OK: got 1 expected 0 (1 is R_SCL_HELD)
FAIL T2 it took exactly three pulses: got 0
FAIL T2 a STOP was issued: got 0
T3 a stuck pad
FAIL T3 result is SDA STILL HELD: got 1 (R_SCL_HELD again)Every case that required clocking reported R_SCL_HELD with zero pulses, on a bus where nothing was holding SCL at all. And the cases that did not require clocking — an already-free bus, a genuinely held SCL — passed, which is the pattern that makes this kind of bug survive review: the tests that fail are the ones exercising the interesting path, and the ones that pass look like good coverage.
The cause is one line. The block released SCL with a non-blocking assignment and tested scl_in in the same cycle, so it read the line before its own release had taken effect — and a released line does not read high until at least the next sample, and on a real bus not until tr has elapsed.
That is the "released means high" error, from Chapter 19.1's DebugLab and 23.1's misconceptions, arriving inside the block whose entire job is recovering from a line that is genuinely held.
The fix is a state, not a delay: release SCL, then wait for the line to actually read high before the high phase begins to count. That is clock synchronisation (Chapter 13.2) rather than a workaround — the high time a target gets is measured from when the clock is high on the wire, not from when this block stopped pulling. And the wait has to be generous, because a target that was interrupted mid-byte is exactly the kind of device that stretches.
5. Five Holders, One Recovery
// -----------------------------------------------------------------------------
// i2c_bus_clear_tb.sv
// Chapter 23.6's experiment: five things that hold a bus low, one recovery, and
// a demonstration of which of them it can and cannot clear.
//
// THE HOLDERS. Each models a real reason a line stays low, and each responds
// differently to clock pulses. That difference is the whole diagnostic:
//
// H_NONE nothing is holding. Included so that a recovery reporting
// success has something to be compared against.
//
// H_MIDBYTE a target interrupted part-way through transmitting. It releases
// after N more clock pulses, where N is how many bits it still had
// to shift. This is the ONE case the recovery is designed for.
//
// H_STUCK a pad asserted permanently. No number of clocks changes it.
//
// H_SCL something holding SCL. The recovery cannot clock, so it cannot
// run at all -- and must say so rather than hang.
//
// H_BOTH both lines held. SCL is reported first, because it is the one
// that makes everything else moot.
//
// Bounded: the recovery has a pulse limit, every wait in the bench is a fixed
// number of edges, and a watchdog terminates the run. A deliberately unclearable
// bus must not hang the runner -- which is the specific reason the pulse limit
// is in the design rather than in the bench.
// -----------------------------------------------------------------------------
`timescale 1ns/1ps
// A holder: something on the bus that is pulling a line down, and a rule for
// when it lets go.
module tb_holder #(
parameter integer KIND = 0, // see the localparams in the bench
parameter integer RELEASE_AT = 3, // H_MIDBYTE: pulses needed before it frees
// The shortest SCL high time this target will accept as a clock pulse. Real
// devices have one -- it is tHIGH, UM10204 Table 10 -- and a recovery clocking
// faster than it delivers pulses the target does not count. Modelling it is
// what makes "the recovery must use a legal clock" testable rather than
// assumed.
parameter integer MIN_HIGH = 4
) (
input wire clk,
input wire rst_n,
input wire scl, // resolved SCL, read back
output wire scl_drive_low,
output wire sda_drive_low
);
localparam integer H_NONE = 0, H_MIDBYTE = 1, H_STUCK = 2, H_SCL = 3, H_BOTH = 4,
H_SCL_LATE = 5, H_STRETCH = 6;
reg scl_q;
reg [7:0] pulses;
// Counted on the FALLING edge, and that is not a detail.
//
// A target shifts its next bit onto SDA while SCL is LOW, because changing SDA
// while SCL is high is by definition a START or a STOP. The first version of
// this model counted rising edges and therefore released SDA during an SCL
// high period -- which the bench's own STOP detector duly counted as a STOP,
// reporting two where one was expected.
//
// That was the MODEL being non-conforming, not the design. It is also a real
// I2C hazard worth keeping in mind: a device that changes SDA at the wrong
// moment does not merely send a wrong bit, it manufactures framing.
wire scl_fall = scl_q & ~scl;
// How long SCL has been continuously high. A pulse counts only if it was high
// for at least MIN_HIGH samples before falling.
reg [7:0] high_run;
always @(posedge clk) begin
if (!rst_n) begin scl_q <= 1'b1; pulses <= 8'd0; high_run <= 8'd0; end
else begin
scl_q <= scl;
if (scl) high_run <= (high_run == 8'hFF) ? high_run : high_run + 8'd1;
else high_run <= 8'd0;
if (scl_fall && high_run >= MIN_HIGH[7:0] && pulses != 8'hFF)
pulses <= pulses + 8'd1;
end
end
// A mid-byte target shifts one bit per clock pulse and lets go when it has
// shifted the last one. It is not "waiting for a timeout"; it is waiting for
// the clock it is entitled to, which is why clocking works and nothing else
// does.
wire midbyte_holding = (pulses < RELEASE_AT[7:0]);
// H_STRETCH pulls SCL low for a bounded interval that begins two samples
// AFTER the clock has risen -- which is legal clock synchronisation (Chapter
// 13.2) and is what a target does when it needs more time part-way through a
// clock's high phase. A recovery that reports it as a stuck clock abandons
// the one case it was built for, because a target interrupted mid-byte is
// exactly the kind of device that stretches.
//
// Positioning it after the rise rather than after the fall is deliberate: a
// stretch that begins while SCL is already low is absorbed by any recovery
// that waits for the line to rise, so it cannot distinguish one that tolerates
// a mid-high stretch from one that does not.
// The stretch is BOUNDED -- two of them, and then the target stops needing
// more time. An unbounded stretch is a different failure (Chapter 23.7) and it
// would hang this test rather than exercise it: a recovery that correctly
// waits for a stretching target would wait forever, which is the correct
// behaviour and an unusable stimulus.
wire scl_rise_h = ~scl_q & scl;
reg [7:0] stretch_arm, stretch_cnt, stretch_budget;
always @(posedge clk) begin
if (!rst_n) begin
stretch_arm <= 8'd0; stretch_cnt <= 8'd0; stretch_budget <= 8'd2;
end else begin
if (scl_rise_h && stretch_budget != 8'd0) stretch_arm <= 8'd2;
else if (stretch_arm != 8'd0) stretch_arm <= stretch_arm - 8'd1;
if (stretch_arm == 8'd1) begin
stretch_cnt <= 8'd2;
stretch_budget <= stretch_budget - 8'd1;
end else if (stretch_cnt != 8'd0) stretch_cnt <= stretch_cnt - 8'd1;
end
end
// H_SCL_LATE leaves SCL alone until the recovery's FIRST clock pulse and then
// holds it forever. SCL is healthy at the moment the recovery checks it, so
// the pre-flight check passes and the fault appears mid-recovery -- which is
// the only way to exercise the bounded wait for the line to rise.
reg scl_late_grab;
always @(posedge clk) begin
if (!rst_n) scl_late_grab <= 1'b0;
else if (scl_fall) scl_late_grab <= 1'b1;
end
assign sda_drive_low = (KIND == H_MIDBYTE) ? midbyte_holding
: (KIND == H_STRETCH) ? midbyte_holding
: (KIND == H_STUCK) ? 1'b1
: (KIND == H_BOTH) ? 1'b1
// Holds SDA too, or the recovery finds the bus already
// free at its pre-flight check and never clocks -- so the
// fault this holder exists to inject is never reached.
: (KIND == H_SCL_LATE) ? 1'b1
: 1'b0;
assign scl_drive_low = (KIND == H_SCL) ? 1'b1
: (KIND == H_BOTH) ? 1'b1
: (KIND == H_SCL_LATE) ? scl_late_grab
: (KIND == H_STRETCH) ? (stretch_cnt != 8'd0)
: 1'b0;
endmodule
module i2c_bus_clear_tb;
localparam integer MAX_PULSES = 9, PHASE = 4;
localparam [1:0] R_OK = 2'd0, R_SCL_HELD = 2'd1, R_SDA_HELD = 2'd2;
localparam integer H_NONE = 0, H_MIDBYTE = 1, H_STUCK = 2, H_SCL = 3, H_BOTH = 4,
H_SCL_LATE = 5, H_STRETCH = 6;
// Both lines have a rise time. Falling is immediate and rising is not, which
// is the asymmetry of Chapter 19.7's bus model reduced to one line each. The
// values are small enough that a release during the low phase completes
// within it -- a longer SDA rise would finish during the following HIGH phase
// and, on the wire, that is a STOP condition rather than a bit, which is a
// real hazard of a bus outside its rise-time budget and not this chapter's
// subject.
localparam integer SDA_RISE = 2;
localparam integer SCL_RISE = 1;
reg clk = 1'b0;
reg rst_n = 1'b0;
always #5 clk = ~clk;
integer errors = 0, checks = 0, neg_detected = 0, negative_mode = 0;
// One bus, five holders, exactly one of which is selected at a time. All five
// are instantiated permanently and gated by `sel`, so every run uses the same
// netlist and the only thing that changes between cases is which holder is
// contributing -- the same reason Chapter 23.4 used one fault-injectable
// target rather than several targets.
reg [2:0] sel = H_NONE;
wire c_scl_low, c_sda_low;
wire [6:0] h_scl_low, h_sda_low;
wire scl_pull = c_scl_low | h_scl_low[sel];
wire sda_pull = c_sda_low | h_sda_low[sel];
// Falling is immediate; rising takes time. The same asymmetry as Chapter
// 19.7's bus model, reduced to one line each.
reg [7:0] scl_rc = 8'd0, sda_rc = 8'd0;
always @(posedge clk) begin
if (scl_pull) scl_rc <= 8'd0;
else if (scl_rc < SCL_RISE[7:0]) scl_rc <= scl_rc + 8'd1;
if (sda_pull) sda_rc <= 8'd0;
else if (sda_rc < SDA_RISE[7:0]) sda_rc <= sda_rc + 8'd1;
end
wire scl = ~scl_pull && (scl_rc >= SCL_RISE[7:0]);
wire sda = ~sda_pull && (sda_rc >= SDA_RISE[7:0]);
genvar g;
generate
for (g = 0; g < 7; g = g + 1) begin : holders
tb_holder #(.KIND(g), .RELEASE_AT(3)) h (
.clk(clk), .rst_n(rst_n), .scl(scl),
.scl_drive_low(h_scl_low[g]), .sda_drive_low(h_sda_low[g])
);
end
endgenerate
reg start = 1'b0;
wire busy, done;
wire [3:0] pulses_used;
wire [1:0] result;
i2c_bus_clear #(.MAX_PULSES(MAX_PULSES), .PHASE(PHASE)) dut (
.clk(clk), .rst_n(rst_n), .start(start),
.scl_in(scl), .sda_in(sda),
.scl_drive_low(c_scl_low), .sda_drive_low(c_sda_low),
.busy(busy), .done(done), .pulses_used(pulses_used), .result(result)
);
// A STOP is SDA rising while SCL is high. Counted continuously, because the
// claim "a STOP was generated" is about an event and an event needs a
// detector, not an inspection afterwards.
reg scl_q2, sda_q2;
integer n_stops = 0;
always @(posedge clk) if (rst_n) begin
scl_q2 <= scl; sda_q2 <= sda;
if (scl_q2 && scl && !sda_q2 && sda) n_stops = n_stops + 1;
end
integer stops_before;
// ------------------------------------------------------------------ helpers
task step; begin @(posedge clk); #1; end endtask
task tick (input integer n); integer k; begin for (k=0;k<n;k=k+1) step; end endtask
task chk (input [255:0] name, input integer got, input integer exp);
begin
checks = checks + 1;
if (got !== exp) begin
if (negative_mode) neg_detected = neg_detected + 1;
else begin
errors = errors + 1;
$display(" FAIL %0s: got %0d expected %0d", name, got, exp);
end
end else if (negative_mode) begin
errors = errors + 1;
$display(" FAIL negative proof did not fire: %0s", name);
end
end
endtask
// `sel` must be set BEFORE calling this. Leaving the previous test's holder
// active across a reset leaves each holder's own `scl_q` initialised to 1
// against a line that holder is pulling low, so the first post-reset edge
// looks like a falling edge and a pulse nobody delivered gets counted. That
// is the same reset-time initial-value mismatch that Chapter 23.5's profiler
// had to guard against, arriving in the bench instead of the design.
task do_reset; begin rst_n = 1'b0; start = 1'b0; tick(3); rst_n = 1'b1; tick(3); end endtask
// Run one recovery to completion, BOUNDED. The wait is a fixed number of
// edges, not a wait on `done` -- a recovery that never finishes must fail the
// test rather than hang it, and a bench that waits on a DUT condition cannot
// tell those apart.
integer timeout_hit;
reg [1:0] r_result;
reg [3:0] r_pulses;
integer r_stops_at_done;
reg r_scl_low_at_done, r_sda_low_at_done;
task run_recovery; integer k;
begin
timeout_hit = 1;
start = 1'b1; step; start = 1'b0;
for (k = 0; k < 4000; k = k + 1) begin
step;
if (done) begin
r_result = result;
r_pulses = pulses_used;
// The STOP count AT the done strobe, not afterwards. A recovery
// that signals completion before its STOP has reached the wire
// has told the rest of the system the bus is usable while it is
// still mid-framing, and a count taken later cannot see that.
r_stops_at_done = n_stops;
// The block's own pull-downs AT the strobe. A block that is still
// pulling a line low while telling the rest of the system it has
// finished has handed over a bus it is itself holding -- and a
// check taken a few cycles later cannot see it, because the idle
// state releases everything on the next clock.
r_scl_low_at_done = c_scl_low;
r_sda_low_at_done = c_sda_low;
timeout_hit = 0;
k = 4000;
end
end
tick(4);
end
endtask
// Did the controller ever drive a line HIGH? It must not, ever, including
// during recovery. Checked structurally: the block has no drive-high output,
// so this watches for the only observable proxy -- a line reading high while
// every holder is still pulling.
integer contention = 0;
always @(posedge clk) if (rst_n)
if ((h_sda_low[sel] && sda) || (h_scl_low[sel] && scl))
contention = contention + 1;
initial begin
$display("i2c_bus_clear_tb");
// ---------------------------------------------------------------------
// T1 -- nothing is holding. The recovery still runs and still issues a
// STOP, and it uses ZERO pulses. That zero is the control for every other
// case: a recovery reporting success after no pulses did not recover
// anything.
// ---------------------------------------------------------------------
$display("T1 an already-free bus");
sel = H_NONE; do_reset; tick(4);
stops_before = n_stops;
run_recovery;
chk("T1 completed", timeout_hit, 0);
chk("T1 result OK", r_result, R_OK);
chk("T1 zero pulses were needed", r_pulses, 0);
chk("T1 a STOP was still issued", n_stops - stops_before, 1);
// ---------------------------------------------------------------------
// T2 -- THE CASE THE RECOVERY EXISTS FOR. A target interrupted mid-byte,
// three bits from the end. It lets go after three pulses, and the count is
// the evidence: it says how far into a byte the transfer was interrupted.
// ---------------------------------------------------------------------
$display("T2 a target interrupted mid-byte");
sel = H_MIDBYTE; do_reset; tick(4);
chk("T2 SDA really is held before recovery", sda, 0);
stops_before = n_stops;
run_recovery;
chk("T2 completed", timeout_hit, 0);
chk("T2 result OK", r_result, R_OK);
chk("T2 it took exactly three pulses", r_pulses, 3);
chk("T2 a STOP was issued", n_stops - stops_before, 1);
chk("T2 the STOP had completed before done was asserted",
r_stops_at_done - stops_before, 1);
tick(6);
chk("T2 the bus is free afterwards", sda, 1);
chk("T2 the block is not pulling SCL", c_scl_low, 0);
chk("T2 nor SDA", c_sda_low, 0);
// ---------------------------------------------------------------------
// T3 -- a stuck pad. Nine pulses, still held, and the recovery says so
// instead of continuing. The distinction from T2 is the whole chapter:
// same symptom, same recovery attempt, different outcome, and the outcome
// is what identifies the cause.
// ---------------------------------------------------------------------
$display("T3 a stuck pad");
sel = H_STUCK; do_reset; tick(4);
stops_before = n_stops;
run_recovery;
chk("T3 completed rather than hanging", timeout_hit, 0);
chk("T3 result is SDA STILL HELD", r_result, R_SDA_HELD);
chk("T3 it used all nine pulses", r_pulses, MAX_PULSES);
chk("T3 and issued NO stop", n_stops - stops_before, 0);
chk("T3 the bus is still held", sda, 0);
// ---------------------------------------------------------------------
// T4 -- SCL held. The recovery cannot clock, so it must report that
// immediately rather than attempt pulses that cannot reach the bus.
// ---------------------------------------------------------------------
$display("T4 SCL held low");
sel = H_SCL; do_reset; tick(4);
stops_before = n_stops;
run_recovery;
chk("T4 completed", timeout_hit, 0);
chk("T4 result is SCL HELD", r_result, R_SCL_HELD);
chk("T4 it did not attempt any pulses", r_pulses, 0);
chk("T4 and issued no stop", n_stops - stops_before, 0);
// ---------------------------------------------------------------------
// T5 -- both lines held. SCL is reported, because it is the condition that
// makes the SDA question unanswerable rather than merely unanswered.
// ---------------------------------------------------------------------
$display("T5 both lines held");
sel = H_BOTH; do_reset; tick(4);
run_recovery;
chk("T5 completed", timeout_hit, 0);
chk("T5 SCL is reported, not SDA", r_result, R_SCL_HELD);
chk("T5 not the SDA verdict", (r_result == R_SDA_HELD) ? 1 : 0, 0);
// ---------------------------------------------------------------------
// T6 -- the pulse count is evidence, not decoration. Three targets
// interrupted at three different points produce three different counts,
// and the count is proportional to how much of the byte was left.
// ---------------------------------------------------------------------
$display("T6 the pulse count says how far into the byte it was");
// The holder needs three pulses, and the recovery reports three. The
// claims worth making about that number, each of which would be false for
// a plausible wrong implementation:
begin : count_claims
integer got_pulses;
sel = H_MIDBYTE; do_reset; tick(4);
run_recovery;
got_pulses = r_pulses;
chk("T6 the count matches what the holder needed", got_pulses, 3);
chk("T6 it is within the specification's nine", (got_pulses <= 9) ? 1 : 0, 1);
// Not zero: a recovery reporting success after no pulses did not
// recover anything, and would report exactly the same R_OK.
chk("T6 it is not zero, so pulses did the work", (got_pulses > 0) ? 1 : 0, 1);
// And not the maximum: a recovery that always delivers nine and then
// checks would also succeed here, and would carry no information about
// how far into the byte the transfer was interrupted.
chk("T6 and not the maximum, so it stopped when the line came free",
(got_pulses < MAX_PULSES) ? 1 : 0, 1);
end
// ---------------------------------------------------------------------
// T9 -- SCL becomes held DURING the recovery. The pre-flight check passes,
// because SCL is healthy at that moment, and the fault appears on the
// first clock pulse. The bounded wait for the line to rise is the only
// thing between this and a permanent hang.
// ---------------------------------------------------------------------
$display("T9 SCL grabbed after the recovery has started");
sel = H_SCL_LATE; do_reset; tick(4);
chk("T9 SCL is healthy before the recovery", scl, 1);
run_recovery;
chk("T9 completed rather than hanging", timeout_hit, 0);
chk("T9 result is SCL HELD", r_result, R_SCL_HELD);
chk("T9 and it got past the pre-flight check", (r_pulses <= 1) ? 1 : 0, 1);
// ---------------------------------------------------------------------
// T10 -- a target that STRETCHES. Holding SCL low after each pulse is
// legal and is exactly what an interrupted target does. The recovery must
// wait rather than report a fault, and must still succeed.
// ---------------------------------------------------------------------
$display("T10 a stretching target must not be reported as a stuck clock");
sel = H_STRETCH; do_reset; tick(4);
chk("T10 SDA is held before recovery", sda, 0);
stops_before = n_stops;
run_recovery;
chk("T10 completed", timeout_hit, 0);
chk("T10 the recovery SUCCEEDED despite the stretching", r_result, R_OK);
chk("T10 not misreported as a stuck clock", (r_result == R_SCL_HELD) ? 1 : 0, 0);
chk("T10 and it still took three pulses", r_pulses, 3);
chk("T10 a STOP was issued", n_stops - stops_before, 1);
// ---------------------------------------------------------------------
// T7 -- NO CONTENTION, ever. Recovery is the moment an implementation is
// most tempted to drive a line high to "unstick" it, and that is the one
// thing an I2C participant must never do. Watched continuously across
// every test above.
// ---------------------------------------------------------------------
$display("T7 the recovery never drove a line high");
chk("T7 no instant where a held line read high", contention, 0);
chk("T7 and the watch was actually running", (n_stops > 0) ? 1 : 0, 1);
// And the block releases everything when it finishes, on a bus with no
// holder at all -- so a recovery cannot leave the bus worse than it found
// it. Checked last, when nothing else is pulling.
sel = H_NONE; do_reset; tick(4);
run_recovery;
tick(8);
chk("T7 SCL released AT the done strobe", r_scl_low_at_done, 0);
chk("T7 SDA released AT the done strobe", r_sda_low_at_done, 0);
chk("T7 SCL still released afterwards", c_scl_low, 0);
chk("T7 SDA still released afterwards", c_sda_low, 0);
chk("T7 and the bus reads idle", (scl && sda) ? 1 : 0, 1);
// ---------------------------------------------------------------------
// T8 -- negative proof. Eight deliberately wrong expectations.
// ---------------------------------------------------------------------
$display("T8 negative proof -- eight deliberately wrong expectations");
sel = H_MIDBYTE; do_reset; tick(4);
stops_before = n_stops;
run_recovery;
chk("T8 setup: recovery succeeded", r_result, R_OK);
chk("T8 setup: three pulses", r_pulses, 3);
negative_mode = 1;
chk("T8a wrong result", r_result, R_SDA_HELD);
chk("T8b wrong pulse count", r_pulses, 9);
chk("T8c wrong stop count", n_stops - stops_before, 0);
chk("T8d wrong completion", timeout_hit, 1);
chk("T8e wrong line state", sda, 0);
chk("T8f wrong contention", contention, 5);
chk("T8g wrong zero-pulse claim", (r_pulses == 0) ? 1 : 0, 1);
chk("T8h wrong verdict distinctness", (r_result == R_SCL_HELD) ? 1 : 0, 1);
negative_mode = 0;
chk("T8 all eight negatives detected", neg_detected, 8);
$display("");
$display("checks=%0d errors=%0d negatives_detected=%0d/8", checks, errors, neg_detected);
if (errors == 0) $display("RESULT: PASS"); else $display("RESULT: FAIL");
$finish;
end
initial begin
#3000000;
$display("RESULT: FAIL -- watchdog expired, the run did not terminate");
$finish;
end
endmodule holder result pulses STOP issued
nothing holding R_OK 0 yes
target interrupted mid-byte R_OK 3 yes
stuck pad R_SDA_HELD 9 no
SCL held from the start R_SCL_HELD 0 no
both lines held R_SCL_HELD 0 no
SCL grabbed mid-recovery R_SCL_HELD 1 no
target that stretches R_OK 3 yesThree rows carry the chapter.
The stuck pad gets nine pulses and a different verdict. Same symptom as the mid-byte target, same recovery attempt, different outcome — and the outcome is what identifies the cause. That is the diagnostic value of running a recovery you expect to fail.
A held SCL is reported in zero pulses, because attempting pulses that cannot reach the bus is not a diagnosis, it is a hang with extra steps. And when both lines are held, SCL is reported rather than SDA, because SCL being held makes the SDA question unanswerable rather than merely unanswered.
A stretching target still succeeds. The recovery waits for the line and completes in the same three pulses. A recovery that reported R_SCL_HELD here would abandon exactly the case it was built for, since a target interrupted mid-byte is a prime candidate for needing more time.
Two things the bench needed that are easy to leave out
A rise time on both lines. Without it, "sample SDA at the end of the high phase" is indistinguishable from "sample it at the start", and a mutation that samples early survives. With a rise time, a line released during the low phase is still rising when the high phase begins.
A minimum clock high time on the target. Real devices have one — it is tHIGH, Table 10 — and a recovery clocking faster than it delivers pulses the target does not count. Modelling it is what makes "the recovery must use a legal clock" testable rather than assumed, and it is what kills the mutation that shortens the high phase to a single sample.
6. The Bench Was Wrong Twice, in Two Different Ways
A holder that released SDA while the clock was high
The mid-byte holder initially counted clock pulses on the rising edge and released SDA there. The bench's own STOP detector counted two STOPs where one was expected.
The detector was right. A device that releases SDA during an SCL high period generates a STOP, on the wire, whatever it intended — and the model was doing exactly that. A real target shifts its next bit while SCL is low, and the fix was to count falling edges.
It is worth keeping the general form: a device that changes SDA at the wrong moment does not merely send a wrong bit, it manufactures framing. That is a much larger failure than a corrupted byte, and it is the reason the data-valid rule is stated the way it is.
A reset that left the previous test's holder pulling
Each test reset the DUT and then selected its holder, so the reset happened with the previous test's holder still active. Every holder's own scl_q register comes out of reset holding 1; against a line the previous holder was pulling low, the first post-reset sample looks like a falling edge, and a pulse nobody delivered gets counted.
The symptom was a recovery that needed three pulses in one test and two in another, with identical stimulus.
This is the same reset-time initial-value mismatch that Chapter 23.5's profiler needed a guard against — arriving, this time, in the bench rather than the design. A register initialised to the idle state is right about an idle bus and wrong about every other kind, and reset is exactly when a bus is least likely to be idle.
7. Mutation
FIRST CAMPAIGN 13 valid 7 KILLED 6 SURVIVED
AFTER T7, T9, T10 13 valid 13 KILLED 0 SURVIVEDThe six genuine survivors were five missing cases and one missing observation:
S04 — SDA sampled at the start of the high phase. Needed both a rise time on SDA and a minimum high time on the target before it became observable.
S08 — a legal stretch reported as a stuck clock. Needed a holder that pulls SCL low after the clock has risen. A stretch that begins while SCL is already low is absorbed by any recovery that waits for the line, so it cannot tell a tolerant implementation from an intolerant one.
S09 and S11 — the unbounded wait and the missing wait. Both needed SCL to become held during the recovery rather than before it, so that the pre-flight check passes and the fault appears on the first pulse. Both mutants then hang, and the bench's bounded wait turns a hang into a failure.
S06 and S13 — outputs checked too late. One leaves SDA pulled for an extra phase; the other keeps pulling SCL in the completion state. Both were invisible because the checks ran several cycles after done, by which time the idle state had tidied up. Capturing the outputs at the done strobe kills both.
8. What Recovery Cannot Do
It cannot clear a held SCL. The mechanism is clocking; a participant that cannot raise SCL cannot clock. The remaining options are outside the protocol.
It cannot clear a stuck pad, a device held in reset, or a short. Nine pulses and the line is still low, which is the verdict R_SDA_HELD — a real finding, and one that points at a device or a board rather than at a protocol.
It cannot tell you which device was holding SDA. The bus resolves every pull-down identically. If two targets could have been mid-byte, the recovery frees the bus and tells you nothing about which one it freed. Attributing it needs a per-device measurement, a physical disconnection, or a bus switch — the same limit Chapter 23.1 §8 names for foreign_hold.
And it does not address why the transfer was interrupted. A bus that needs clearing at power-on is telling you that something stopped mid-byte the last time it ran — a watchdog, a crash, a reset asserted without a STOP. Clearing the bus makes the symptom go away and leaves the cause running.
9. Misconceptions
10. Debug Lab
The bus clears on every boot, and one board in forty does not
A recovery whose success is the symptom, and a pulse count that says where to look
// A product whose firmware runs the bus-clear routine at startup, as a matter of
// course, before touching any device. It has done so for two years.
//
// A new batch shows a different failure: one board in roughly forty comes up with
// the I2C bus unusable, and the clear routine does not fix it. Power-cycling the
// board sometimes helps and sometimes does not.
//
// Two questions are worth separating, and the team had merged them:
//
// why does the clear FAIL on one board in forty?
// why does a clear routine run at every startup at all?
//
// The second one nobody was asking, because the routine had always been there
// and had always appeared to work. "It has always done that" is the most
// effective way for a fault to stay invisible for two years.
//
// For the FAILING board, the candidates for a line that nine clock pulses do not
// free. Section 1's table names seven holders; four of them survive the fact
// that clocking did not help:
//
// (a) a stuck or shorted pad on one device -- electrical, permanent
// (b) a device held in reset with its output asserted -- its pad is driven by
// reset state rather than by protocol, so clocks mean nothing to it
// (c) a target whose state machine is wedged rather than merely paused -- it
// is not waiting for clocks, so it does not respond to them
// (d) a short between SDA and ground on the board itself
//
// All four produce the identical result: nine pulses, line still low. The pulse
// count does NOT separate them, and this is worth saying plainly -- the
// instrument's contribution here is to eliminate the fifth candidate, a target
// mid-byte, which is the only one the recovery could have fixed.The clear routine was instrumented to report its result code and pulse count rather than a boolean, and the numbers were logged across the batch at every boot.
Thirty-nine boards, every boot: result = R_OK pulses_used = 2 (36 boards) result = R_OK pulses_used = 3 (3 boards)
The failing board, every boot: result = R_SDA_HELD pulses_used = 9
Two findings here, and the second is the one nobody was looking for.
FINDING 1, the failing board. Nine pulses and SDA still low is not a target waiting for clocks -- a target mid-byte needs at most nine and this one is unmoved by them. The verdict is a real one: something is holding SDA that is not waiting for anything. That is a stuck pad, a device held in reset with its output asserted, or a short.
FINDING 2, the thirty-nine that "work". They are not recovering from nothing. A pulse count of 2 means SDA WAS held at every single boot, and a target needed two clock pulses to let go -- which means it was two bits from the end of a byte when it was interrupted. Thirty-six boards, every boot, the same count.
A count that is the SAME on thirty-six boards is not random. An interruption at an arbitrary moment would scatter across one to nine; two, repeatedly, is a specific point in a specific transfer.
And the three boards at three pulses are also informative: the population is not uniform, so whatever is being interrupted is not identical everywhere -- which argues for timing rather than for a fixed code path.
Two separate faults, and the log separated them.
THE THIRTY-NINE. Something stops a transfer two bits from the end of a byte, on every boot. Working backwards: the bus is left mid-byte when the previous session ended, so the interruption is at SHUTDOWN or at RESET, not at startup.
Reading the reset circuit: the processor's reset is asserted by the supply supervisor as soon as the rail droops, with no software involvement. Any transfer in progress stops exactly where it is. The last thing the firmware does before power-down is a sensor read, and the consistent count says the reset lands at the same point in it -- which it would, because the droop and the transfer are both paced by the same supply.
So the clear routine is not a defensive measure that has never been needed. It has been needed at every boot for two years, papering over an unclean shutdown, and the count is what says so. A boolean "clear succeeded" could not have.
THE ONE. Nine pulses and no movement is a different fault entirely, and the clear routine's verdict is the evidence that separates it from the fleet's fault -- but NOT the evidence that separates candidates (a) to (d) from each other. Those four are indistinguishable from the bus, and the next measurement has to be physical. On this board the temperature sensor's SDA pin measured 0.06 V with the device unpowered and the rest of the board removed -- a short to ground under the package, from a solder bridge. Reflow fixed it.
Note what the pulse count did for THIS board too: nine, not two. If the failing board had reported two pulses and then failed, the fault would have been somewhere else entirely -- a target that releases and something else that grabs.
CORRECTIONS, and there are two because there were two faults:
the shorted board: reflow, and add an SDA-to-ground continuity check to the production test, because the clear routine found it only after assembly.
the thirty-nine: assert a STOP before allowing the supervisor to reset, or hold reset off until the current transfer completes. The clear routine stays -- it is a reasonable belt-and-braces measure -- but it stops being the thing that makes the product work.
PROOF, and the first two steps are about different faults:
1. after reflow, the failing board reports result = R_OK, pulses_used = 2 -- which is the FLEET's behaviour, not a clean bus. The right outcome: the board now has the same fault as everything else and not one of its own. 2. after the shutdown change, the fleet reports result = R_OK, pulses_used = 0 -- the clear routine finds nothing to clear 3. and pulses_used = 0 stays at 0 across two thousand boots on ten boards
Step 2 is the one that proves the shutdown fault was the cause. A zero pulse count is a much stronger statement than "the bus works": it says the recovery had nothing to do, which is the only outcome consistent with a clean shutdown. Step 3 is there because the fault was intermittent in its consequences and not in its occurrence -- it happened every boot, and only ever mattered when something else went wrong at the same time.
11. Reason It Through
12. Questions
13. What This Chapter Settled
Seven things that hold an I²C line low and exactly one that clocking repairs: a target interrupted mid-byte, waiting for pulses it is entitled to. The other six — a wedged state machine, a device in reset, a stuck pad, a short, a held SCL, and a target legally stretching — are unaffected, and the value of running a recovery against them is the verdict rather than the repair.
A sequencer whose outputs are evidence: three distinguishable outcomes, and a pulse count that says how far into a byte the transfer was interrupted. Fifty-seven checks, zero errors, eight negative proofs, thirteen mutations all killed — against a baseline confirmed passing, after one campaign had to be thrown away for running against a failing one.
Two bugs in the design and two in the bench, all four worth more than the tests that caught them. A recovery that read SCL in the cycle it released it and reported a stuck clock on a healthy bus. A holder that released SDA while the clock was high and manufactured a STOP. A reset taken with the previous test's holder still pulling, which counted a pulse nobody delivered. And two checks that ran a few cycles after the moment they were about.
And the observation the Debug Lab turns on: a recovery that always succeeds is an unexamined finding, visible only if its result is a count rather than a boolean.
One holder on that list was set aside as legal rather than faulty. A target stretching SCL is entitled to, a second controller synchronising is entitled to, and both look exactly like a bus that has stopped. Chapter 23.7 is where those two become failures — and where a timeout stops being a fault report and becomes an ambiguity.
Continue learning
Related tutorials
- Related topic
I²C Stretch Bounds, Timeouts and the No-Assumption Rule
The specification places no limit on how long a slave may hold the clock, so every timeout is a system policy rather than a compliance check. Includes the recovery asymmetry most engineers know only half of.
- Related topic
Bus Clear and Recovery — The Protocol-Level Escape Hatch
Four sentences of specification that say two different things about two different wires. Explains why nine clock pulses free a stuck SDA and can never free a stuck SCL, why that is structural rather than an omission, and what a recovery sequencer must refuse to attempt.
- Related topic
Clock-Stretching and Arbitration Failures in the Field
Why a timeout is the least informative evidence an I²C controller produces, and what to record instead. Six situations that all report the same timeout, a classifier that separates them, and the one failure in the set that never times out at all — a controller that does not honour stretching, corrupting data silently and blaming the target.
- Related topic
The START Condition
START is SDA falling while SCL is high, it is generated only by the controller, and it makes the bus busy. Derive what every device must do in response, then build a detector in three languages and find out why its two guard terms and its reset value are all load-bearing.
