I²C · Module 15
Bus Clear and Recovery — The Protocol-Level Escape Hatch
Four sentences of specification that say two different things about two different wires. Explains why nine clock pulses free a stuck SDA and can never free a stuck SCL, why that is structural rather than an omission, and what a recovery sequencer must refuse to attempt.
Every mechanism in this curriculum assumes the bus works. A master can clock, a target can acknowledge, a line that is released floats high. This chapter is about what to do when that stops being true — and the specification's whole answer is four sentences.
They are worth reading in full before anything else, because the single most important fact about them is that they say two different things about two different wires.
Nine clock pulses for a stuck SDA. For a stuck SCL, nothing at the protocol level at all — straight to a hardware reset or a power cycle.
That is not an omission. The nine pulses are pulses on SCL, and a master cannot clock a line that something else is holding low. The remedy that works for the line the master shares is structurally unavailable for the line the master owns.
A recovery block that misses this will attempt the impossible procedure on a stuck clock, issue nine pulses that never appear on the wire, observe no change, and report a failed recovery — when what it should have reported is that no protocol recovery exists for that line.
1. Which Line, Which Remedy
The two ladders, side by side:
| line stuck low | remedy 1 | remedy 2 | remedy 3 |
|---|---|---|---|
| SDA | nine clock pulses | HW reset | power cycle |
| SCL | — (none exists) | HW reset | power cycle |
And the reason for the asymmetry is Chapter 2.5's open-drain rule, read carefully:
SDA is shared by everybody. A master drives it during an address or a write; a target drives it during a read or an acknowledge. So a master can stop driving SDA entirely and still do something useful — it can clock. Nine pulses give whatever device is stuck in a bit-shift loop nine opportunities to advance past the bit it is holding, and a target wedged mid-byte will typically walk out of its own state machine and release.
SCL is driven by the master, and released by targets that stretch. So a target holding SCL low is holding the only resource the master has. There is no action left that involves the bus — the master's single output on that wire is already being overridden.
On SDA the master retains an instrument it can play. On SCL it does not. Everything in §3.1.16 follows from that one difference.
2. What Recovery Cannot Repair
Worth stating early, because a recovery procedure invites more confidence than it deserves.
The nine-pulse procedure addresses exactly one failure: a device holding SDA low when it should not be. It does nothing about:
| failure | does bus clear help? | why not |
|---|---|---|
| a target wedged mid-byte holding SDA | yes — this is the case | nine pulses walk it out of its frame |
| a target holding SCL (a stuck stretch) | no | the pulses cannot be issued |
| a short to ground on either line | no | nothing releases; escalate |
| a powered-down device pulling a line | no | see §5 |
| two devices at the same address | no | the bus is electrically fine |
| a master that lost track of state | no | the bus is fine; reset the master |
| an addressing or protocol error | no | nothing is stuck |
Most of those escalate to the same two rungs, and the last three are not bus problems at all. A design that runs bus clear on every error will eventually run it on something it cannot fix and learn nothing — which is why §6's design diagnoses first and reports its diagnosis as an output.
And the diagnosis is cheap: read the two lines. A stuck line is one of the few faults on this bus that is directly observable without any protocol context at all, which is a pleasant contrast to Chapter 13.5's attribution problems.
3. The Escalation Ladder
Both ladders end the same way, and the last rung is the one the specification guarantees.
HW reset is "preferential" where it exists. It is a pin, outside the bus, and it does not depend on the stuck device cooperating — which is the property that matters, because the stuck device has already demonstrated that it is not cooperating.
A power cycle "activates the mandatory internal Power-On Reset (POR) circuit". The word doing the work there is mandatory: every I²C device has a POR, so a power cycle always works. It is the bottom of the ladder because it is the most disruptive and the most certain.
Each rung down the ladder requires less cooperation from the broken device and disrupts more of the system. The nine pulses need the device's state machine to be running; a HW reset needs only a pin; a power cycle needs nothing at all.
Which gives a design a real decision to make, and §6's sequencer takes both capabilities as inputs rather than assuming either:
| what the board provides | recovery available |
|---|---|
| HW reset pins on all devices | nine pulses, then HW reset |
| no HW reset, switchable supply | nine pulses, then power cycle |
| neither | nine pulses, then nothing |
That last row is a real configuration and the honest output for it is that the bus cannot be recovered. §6's design reports gave_up and a remedy of impossible, because a recovery block that claimed success there would be the most dangerous kind of wrong.
4. The Procedure, Drawn
Nine pulses are a bound, not a quota: this holder releases on the third and the procedure stops
10 cyclesTwo details in that picture are easy to get wrong in an implementation.
The holder releases during the SCL low phase. That is Chapter 4.2's data-valid rule — SDA may only change while SCL is low — and it means a master must sample SDA while SCL is high to see the release. A master that samples at the instant it releases SCL reads the value from before the change and needs one extra pulse to notice, which it will then report as the bus having been slower to recover than it was. §6a returns to this.
The procedure ends with a STOP. A released SDA with no framing leaves every device on the bus believing a transfer is still in progress. The specification does not spell this out, and it follows from Chapter 5.3: the only thing that returns a device to its idle state is a STOP.
5. The Failure Mode Nobody Plans For
One cause of a stuck line deserves its own section because it is common, it is not a bug in any device, and bus clear cannot fix it.
A device whose supply is off, on a bus whose pull-ups are still powered, can pull a line low through its own ESD protection diodes. The input protection conducts from the pin to a rail that is now at ground, and the pin becomes a low-impedance path to it.
This is why Chapter 14.1 §3 quoted §5.1's requirement:
A conforming device floats when unpowered. A non-conforming one — or a conforming one wired to a supply that droops rather than switching cleanly — holds the bus down, and no number of clock pulses will change that because there is no state machine running to respond to them.
And Chapter 15.1 §3a's precaution is the same hazard from the other direction: a device emerging from a reset while still pulling SDA or SCL low blocks the bus. A general call 06h resets every device at once, so it is forty simultaneous opportunities to create exactly the condition this chapter is about.
The two most likely causes of a stuck line are a device that is off and a device that has just been reset. Neither is repairable by clocking, and one of them was created by the recovery mechanism of a previous chapter.
6. The Recovery Sequencer in Three Languages
The design diagnoses which line is stuck, chooses the remedy that exists for that line, issues at most nine pulses and stops as soon as the line is free, escalates by what the board actually supports, and closes with a STOP.
It also carries one output whose only purpose is to prove a negative: attempted_clocks_on_scl must remain low for the life of the design, and the testbench asserts it in four separate scenarios.
// -----------------------------------------------------------------------------
// i2c_bus_recovery.sv
// Bus clear and recovery sequencer (UM10204 3.1.16).
//
// Section 3.1.16 is four sentences long and it says two different things about two
// different wires. That asymmetry is the whole design:
//
// "In the unlikely event where the CLOCK (SCL) is stuck LOW, the preferential
// procedure is to reset the bus using the HW reset signal if your I2C devices
// have HW reset inputs. If the I2C devices do not have HW reset inputs, cycle
// power to the devices to activate the mandatory internal Power-On Reset (POR)
// circuit."
//
// "If the DATA line (SDA) is stuck LOW, the master should send nine clock
// pulses. The device that held the bus LOW should release it some time within
// those nine clocks. If not, then use the HW reset or cycle power to clear the
// bus."
//
// Nine clock pulses for a stuck SDA. NOTHING at the protocol level for a stuck
// SCL. And the reason is not an omission -- it is that the nine pulses ARE pulses
// on SCL. A master cannot clock a line something else is holding low, so the
// remedy that works for the line the master shares is unavailable for the line
// the master owns.
//
// stuck SDA: nine clocks -> HW reset -> power cycle
// stuck SCL: HW reset -> power cycle
//
// Six obligations:
//
// 1. Diagnose WHICH line is stuck before choosing a remedy. They are not
// interchangeable and the wrong one wastes the only recovery attempt.
// 2. Never attempt the nine-clock procedure on a stuck SCL. The block exposes
// `attempted_clocks_on_scl` for the sole purpose of proving it never does.
// 3. Issue at most nine pulses, and stop as soon as SDA is released -- the
// specification says the holder "should release it some time WITHIN those
// nine clocks", so nine is a bound, not a quota.
// 4. Report how many pulses were actually needed. A bus that needs eight every
// time is telling you something a pass/fail flag would hide.
// 5. Escalate only after the nine pulses have failed, and choose the rung by
// what the devices actually support.
// 6. Finish with a STOP so the bus is left idle rather than merely unstuck.
// A released SDA with no framing leaves every device mid-transaction.
// -----------------------------------------------------------------------------
module i2c_bus_recovery #(
parameter int CNT_W = 8
) (
input logic clk,
input logic rst_n,
// ---- bus observation ----------------------------------------------------
input logic scl_in, // the SCL line as it actually reads
input logic sda_in, // the SDA line as it actually reads
input logic begin_recovery, // pulse: something decided the bus is stuck
// ---- what the board can do ---------------------------------------------
input logic has_hw_reset, // the devices have HW reset inputs
input logic can_cycle_power, // the supply can be cycled
// ---- bus driving --------------------------------------------------------
output logic scl_drive_low, // pull SCL low (generating a pulse)
output logic sda_drive_low, // pull SDA low (for the closing STOP)
// ---- remedy ------------------------------------------------------------
output logic [2:0] remedy,
output logic assert_hw_reset,
output logic request_power_cycle,
// ---- status ------------------------------------------------------------
output logic [2:0] state,
output logic [2:0] diagnosis,
output logic [3:0] pulses_issued,
output logic recovered,
output logic gave_up,
output logic attempted_clocks_on_scl, // must remain 0, always
output logic [CNT_W-1:0] recovery_attempts,
output logic [CNT_W-1:0] successes
);
localparam [2:0] D_NONE = 3'd0, // nothing wrong
D_SDA_STUCK = 3'd1, // SDA low, SCL free
D_SCL_STUCK = 3'd2, // SCL low
D_BOTH = 3'd3; // both low
localparam [2:0] R_NONE = 3'd0,
R_NINE_CLOCKS = 3'd1,
R_HW_RESET = 3'd2,
R_POWER_CYCLE = 3'd3,
R_IMPOSSIBLE = 3'd4; // stuck, and nothing on the board can help
localparam [2:0] S_IDLE = 3'd0,
S_DIAG = 3'd1,
S_CLOCK = 3'd2, // issuing the nine pulses
S_STOP1 = 3'd3, // closing STOP: SDA low, SCL rising
S_STOP2 = 3'd4, // closing STOP: SDA released while SCL high
S_ESC = 3'd5, // escalating past the nine pulses
S_DONE = 3'd6;
// Three phases per pulse, not two. SDA must be SAMPLED while SCL is HIGH --
// Chapter 4.2's data-valid rule -- so releasing SCL and reading SDA cannot be
// the same step. A two-phase pulse reads SDA at the instant of release, before
// a device that let go during the low phase can be seen to have done so, and
// then needs one extra pulse to notice. That extra pulse is an artefact of the
// sampling point, and it would be reported as if the bus had been slower to
// recover than it was.
localparam [1:0] PH_LOW = 2'd0, // pull SCL low: the low phase begins
PH_REL = 2'd1, // release SCL: it rises
PH_SMP = 2'd2; // SCL is high: now read SDA
logic [1:0] pulse_phase;
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
state <= S_IDLE;
scl_drive_low <= 1'b0;
sda_drive_low <= 1'b0;
remedy <= R_NONE;
assert_hw_reset <= 1'b0;
request_power_cycle <= 1'b0;
diagnosis <= D_NONE;
pulses_issued <= 4'd0;
recovered <= 1'b0;
gave_up <= 1'b0;
attempted_clocks_on_scl <= 1'b0;
recovery_attempts <= {CNT_W{1'b0}};
successes <= {CNT_W{1'b0}};
pulse_phase <= PH_LOW;
end else begin
case (state)
// ---------------------------------------------------------------
S_IDLE: begin
scl_drive_low <= 1'b0;
sda_drive_low <= 1'b0;
assert_hw_reset <= 1'b0;
request_power_cycle <= 1'b0;
if (begin_recovery) begin
recovered <= 1'b0;
gave_up <= 1'b0;
pulses_issued <= 4'd0;
pulse_phase <= PH_LOW;
remedy <= R_NONE;
recovery_attempts <= recovery_attempts + 1'b1;
state <= S_DIAG;
end
end
// ---------------------------------------------------------------
// Obligation 1. Diagnose the line, then choose. The order of these
// tests matters: SCL being low dominates, because a stuck SCL makes
// the nine-clock procedure impossible regardless of what SDA is doing.
// ---------------------------------------------------------------
S_DIAG: begin
if (!scl_in && !sda_in) begin
diagnosis <= D_BOTH;
state <= S_ESC; // obligation 2: no clocking
end else if (!scl_in) begin
diagnosis <= D_SCL_STUCK;
state <= S_ESC; // obligation 2: no clocking
end else if (!sda_in) begin
diagnosis <= D_SDA_STUCK;
remedy <= R_NINE_CLOCKS;
state <= S_CLOCK;
end else begin
// Nothing is stuck. Recovery was requested on a healthy bus,
// which is worth reporting rather than "fixing".
diagnosis <= D_NONE;
remedy <= R_NONE;
recovered <= 1'b1;
successes <= successes + 1'b1;
state <= S_DONE;
end
end
// ---------------------------------------------------------------
// Obligation 3 and 4. Up to nine pulses, stopping the instant SDA
// is released. Each pulse is one low half and one high half.
// ---------------------------------------------------------------
S_CLOCK: begin
// A stuck SCL here would mean the diagnosis was wrong; refuse to
// pretend to clock. This is obligation 2 enforced a second time,
// because a line can fail DURING recovery.
if (!scl_in && !scl_drive_low) begin
diagnosis <= D_SCL_STUCK;
scl_drive_low <= 1'b0;
state <= S_ESC;
end else if (pulse_phase == PH_LOW) begin
scl_drive_low <= 1'b1; // pull SCL low
pulse_phase <= PH_REL;
end else if (pulse_phase == PH_REL) begin
scl_drive_low <= 1'b0; // release it; SCL rises
pulse_phase <= PH_SMP;
end else begin
// SCL is high: this is where SDA is valid, so this is where the
// release is observed. The pulse is complete either way.
pulse_phase <= PH_LOW;
pulses_issued <= pulses_issued + 1'b1;
if (sda_in) begin
// Released. Obligation 6: leave the bus framed, not merely
// unstuck.
state <= S_STOP1;
end else if (pulses_issued + 1'b1 >= 4'd9) begin
// Nine pulses spent and still held. Obligation 5.
state <= S_ESC;
end
end
end
// ---------------------------------------------------------------
// Obligation 6. A STOP: SDA low while SCL is high, then SDA rising.
// ---------------------------------------------------------------
S_STOP1: begin
sda_drive_low <= 1'b1;
scl_drive_low <= 1'b0;
state <= S_STOP2;
end
S_STOP2: begin
sda_drive_low <= 1'b0; // SDA rises while SCL is high
recovered <= 1'b1;
successes <= successes + 1'b1;
state <= S_DONE;
end
// ---------------------------------------------------------------
// Obligation 5. Escalate by what the board supports. The order is
// the specification's: HW reset is "preferential", a power cycle is
// the fallback that invokes the mandatory POR.
// ---------------------------------------------------------------
S_ESC: begin
scl_drive_low <= 1'b0;
sda_drive_low <= 1'b0;
if (has_hw_reset) begin
remedy <= R_HW_RESET;
assert_hw_reset <= 1'b1;
end else if (can_cycle_power) begin
remedy <= R_POWER_CYCLE;
request_power_cycle <= 1'b1;
end else begin
// Neither rung exists. Saying so is the only honest output.
remedy <= R_IMPOSSIBLE;
gave_up <= 1'b1;
end
state <= S_DONE;
end
S_DONE: begin
if (begin_recovery) begin
recovered <= 1'b0;
gave_up <= 1'b0;
pulses_issued <= 4'd0;
pulse_phase <= PH_LOW;
remedy <= R_NONE;
assert_hw_reset <= 1'b0;
request_power_cycle <= 1'b0;
recovery_attempts <= recovery_attempts + 1'b1;
state <= S_DIAG;
end
end
default: state <= S_IDLE;
endcase
// Obligation 2, as a standing assertion rather than a comment. If the
// block ever drives SCL while SCL is being held low by somebody else,
// it is attempting the one thing 3.1.16 does not offer.
if (scl_drive_low && (diagnosis == D_SCL_STUCK || diagnosis == D_BOTH))
attempted_clocks_on_scl <= 1'b1;
end
end
endmodule `timescale 1ns/1ps
// -----------------------------------------------------------------------------
// i2c_bus_recovery_tb.sv
// Independent oracle for i2c_bus_recovery.
//
// The bench owns the bus. It models a holder that releases SDA after a chosen
// number of clock pulses, and it computes SCL and SDA as a wired-AND of the
// holder and the DUT -- so what the DUT observes is a value the bench derived,
// never the DUT's own opinion of the line.
//
// The property the suite exists for is a NEGATIVE one: the nine-clock procedure
// must never be attempted on a stuck SCL. Tests 4, 5 and 9 assert it, and one of
// them makes SCL fail in the MIDDLE of a recovery that started legitimately.
// -----------------------------------------------------------------------------
module i2c_bus_recovery_tb;
localparam [2:0] D_NONE = 3'd0, D_SDA_STUCK = 3'd1, D_SCL_STUCK = 3'd2, D_BOTH = 3'd3;
localparam [2:0] R_NONE = 3'd0, R_NINE_CLOCKS = 3'd1, R_HW_RESET = 3'd2,
R_POWER_CYCLE = 3'd3, R_IMPOSSIBLE = 3'd4;
localparam [2:0] S_IDLE = 3'd0, S_DIAG = 3'd1, S_CLOCK = 3'd2, S_STOP1 = 3'd3,
S_STOP2 = 3'd4, S_ESC = 3'd5, S_DONE = 3'd6;
logic clk = 1'b0;
logic rst_n = 1'b0;
logic begin_recovery = 1'b0;
logic has_hw_reset = 1'b1;
logic can_cycle_power = 1'b1;
logic scl_drive_low, sda_drive_low;
logic [2:0] remedy, state, diagnosis;
logic assert_hw_reset, request_power_cycle;
logic [3:0] pulses_issued;
logic recovered, gave_up, attempted_clocks_on_scl;
logic [7:0] recovery_attempts, successes;
// ---- the bench's model of the other devices on the bus -----------------
logic holder_holds_sda = 1'b0; // driven ONLY by the holder block below
logic holder_holds_scl = 1'b0;
logic want_hold_sda = 1'b0; // the stimulus ASKS for a hold
logic clear_count = 1'b0;
integer release_after_pulses = 0; // 0 = never releases
integer pulses_counted = 0;
logic scl_prev = 1'b1;
// THE WIRED-AND. The line is low if the DUT pulls, or the holder pulls.
wire scl_in = !(scl_drive_low || holder_holds_scl);
wire sda_in = !(sda_drive_low || holder_holds_sda);
integer errors = 0;
integer n;
i2c_bus_recovery #(.CNT_W(8)) dut (
.clk(clk), .rst_n(rst_n),
.scl_in(scl_in), .sda_in(sda_in), .begin_recovery(begin_recovery),
.has_hw_reset(has_hw_reset), .can_cycle_power(can_cycle_power),
.scl_drive_low(scl_drive_low), .sda_drive_low(sda_drive_low),
.remedy(remedy), .assert_hw_reset(assert_hw_reset),
.request_power_cycle(request_power_cycle),
.state(state), .diagnosis(diagnosis), .pulses_issued(pulses_issued),
.recovered(recovered), .gave_up(gave_up),
.attempted_clocks_on_scl(attempted_clocks_on_scl),
.recovery_attempts(recovery_attempts), .successes(successes));
always #5 clk = ~clk;
// The holder watches SCL and lets go after the configured number of pulses.
//
// It releases on the FALLING edge of SCL, not the rising one, because that is
// what a real device does: Chapter 4.2's data-valid rule says SDA may only
// change while SCL is LOW. Releasing on the rising edge would be both illegal
// and unobservable -- the master samples SDA on that same edge, so a release
// there costs an extra pulse to notice, and the extra pulse would be an
// artefact of the model rather than a property of the bus.
always @(posedge clk) begin
if (!rst_n) begin
scl_prev <= 1'b1;
pulses_counted = 0;
holder_holds_sda <= want_hold_sda;
end else begin
if (clear_count) begin
pulses_counted = 0;
holder_holds_sda <= want_hold_sda; // re-arm for the next attempt
end else if (!scl_in && scl_prev) begin // SCL falling: the low phase begins
pulses_counted = pulses_counted + 1;
if (release_after_pulses != 0 && pulses_counted >= release_after_pulses)
holder_holds_sda <= 1'b0;
end
scl_prev <= scl_in;
end
end
task step; begin @(posedge clk); @(negedge clk); end endtask
task setup (input holds_sda, input holds_scl, input integer rel);
begin
@(negedge clk);
rst_n = 1'b0;
begin_recovery = 1'b0;
release_after_pulses = rel;
clear_count = 1'b1;
want_hold_sda = holds_sda;
holder_holds_scl = holds_scl;
repeat (3) @(posedge clk);
@(negedge clk); clear_count = 1'b0; rst_n = 1'b1;
step;
end
endtask
task kick;
begin @(negedge clk); begin_recovery = 1'b1; @(posedge clk); @(negedge clk); begin_recovery = 1'b0; end
endtask
task run_to_done (input integer bound);
begin
n = 0;
while (state != S_DONE && n < bound) begin step; n = n + 1; end
if (n >= bound) begin
$display(" FAIL run_to_done: stuck in state %0d", state);
errors = errors + 1;
end
end
endtask
task ck_int (input [200*8:1] what, input integer got, input integer exp);
begin
if (got !== exp) begin
$display(" FAIL %0s: got %0d expected %0d", what, got, exp);
errors = errors + 1;
end
end
endtask
task ck_bit (input [200*8:1] what, input got, input exp);
begin
if (got !== exp) begin
$display(" FAIL %0s: got %0b expected %0b", what, got, exp);
errors = errors + 1;
end
end
endtask
initial begin
$display("=== i2c_bus_recovery: two stuck lines, two different escape hatches ===");
// ----------------------------------------------------------------
// T1. SDA stuck, holder releases after three pulses. The procedure must
// stop as soon as the line is free -- nine is a bound, not a quota.
// ----------------------------------------------------------------
setup(1'b1, 1'b0, 3);
kick;
run_to_done(200);
$display("T1 SDA stuck, released after three pulses");
ck_int("T1 diagnosed SDA", diagnosis, D_SDA_STUCK);
ck_int("T1 chose nine clocks", remedy, R_NINE_CLOCKS);
ck_bit("T1 recovered", recovered, 1'b1);
ck_int("T1 stopped at 3 pulses", pulses_issued, 3);
ck_bit("T1 no HW reset needed", assert_hw_reset, 1'b0);
ck_bit("T1 no power cycle", request_power_cycle, 1'b0);
ck_bit("T1 never clocked a stuck SCL", attempted_clocks_on_scl, 1'b0);
ck_int("T1 one success", successes, 1);
// ----------------------------------------------------------------
// T2. SDA stuck, released on the ninth pulse -- the last one allowed.
// ----------------------------------------------------------------
setup(1'b1, 1'b0, 9);
kick;
run_to_done(300);
$display("T2 released on the ninth pulse: still a success");
ck_bit("T2 recovered", recovered, 1'b1);
ck_int("T2 used all nine", pulses_issued, 9);
ck_int("T2 remedy was the clocks", remedy, R_NINE_CLOCKS);
ck_bit("T2 did not escalate", assert_hw_reset, 1'b0);
// ----------------------------------------------------------------
// T3. SDA stuck and NEVER released. Nine pulses, then escalation. The
// pulse count must stop at nine -- not ten, and not forever.
// ----------------------------------------------------------------
setup(1'b1, 1'b0, 0);
kick;
run_to_done(300);
$display("T3 never released: nine pulses then escalation");
ck_int("T3 exactly nine pulses", pulses_issued, 9);
ck_bit("T3 not recovered", recovered, 1'b0);
ck_int("T3 escalated to HW reset", remedy, R_HW_RESET);
ck_bit("T3 HW reset asserted", assert_hw_reset, 1'b1);
ck_bit("T3 no power cycle yet", request_power_cycle, 1'b0);
// ----------------------------------------------------------------
// T4. SCL STUCK. The nine-clock procedure does not exist for this line.
// The block must diagnose it, escalate immediately, and issue ZERO
// pulses -- because a pulse on a held line is not a pulse.
// ----------------------------------------------------------------
setup(1'b0, 1'b1, 0);
kick;
run_to_done(200);
$display("T4 SCL stuck: no clocking is attempted at all");
ck_int("T4 diagnosed SCL", diagnosis, D_SCL_STUCK);
ck_int("T4 zero pulses issued", pulses_issued, 0);
ck_bit("T4 NEVER clocked a stuck SCL", attempted_clocks_on_scl, 1'b0);
ck_int("T4 went straight to HW reset", remedy, R_HW_RESET);
ck_bit("T4 HW reset asserted", assert_hw_reset, 1'b1);
ck_bit("T4 not recovered by clocking", recovered, 1'b0);
// ----------------------------------------------------------------
// T5. BOTH lines stuck. Still no clocking, and the diagnosis says both
// rather than picking one.
// ----------------------------------------------------------------
setup(1'b1, 1'b1, 0);
kick;
run_to_done(200);
$display("T5 both lines stuck: diagnosed as both, still no clocking");
ck_int("T5 diagnosed both", diagnosis, D_BOTH);
ck_int("T5 zero pulses", pulses_issued, 0);
ck_bit("T5 never clocked", attempted_clocks_on_scl, 1'b0);
ck_int("T5 escalated", remedy, R_HW_RESET);
// ----------------------------------------------------------------
// T6. No HW reset available: the power-cycle rung. The specification calls
// the POR "mandatory", which is what makes this a last resort that
// always works.
// ----------------------------------------------------------------
setup(1'b0, 1'b1, 0);
@(negedge clk); has_hw_reset = 1'b0; can_cycle_power = 1'b1;
kick;
run_to_done(200);
$display("T6 without a HW reset input, cycle power");
ck_int("T6 remedy is a power cycle", remedy, R_POWER_CYCLE);
ck_bit("T6 power cycle requested", request_power_cycle, 1'b1);
ck_bit("T6 no HW reset", assert_hw_reset, 1'b0);
ck_bit("T6 did not give up", gave_up, 1'b0);
// ----------------------------------------------------------------
// T7. Neither rung exists. The only honest answer is that the bus cannot
// be recovered -- a checker that reported success here would be lying.
// ----------------------------------------------------------------
setup(1'b0, 1'b1, 0);
@(negedge clk); has_hw_reset = 1'b0; can_cycle_power = 1'b0;
kick;
run_to_done(200);
$display("T7 with neither remedy available, say so");
ck_int("T7 remedy is impossible", remedy, R_IMPOSSIBLE);
ck_bit("T7 gave up", gave_up, 1'b1);
ck_bit("T7 not recovered", recovered, 1'b0);
@(negedge clk); has_hw_reset = 1'b1; can_cycle_power = 1'b1;
// ----------------------------------------------------------------
// T8. Recovery requested on a HEALTHY bus. Nothing is stuck, so nothing
// should be driven -- and reporting that is more useful than running
// a procedure on a working bus.
// ----------------------------------------------------------------
setup(1'b0, 1'b0, 0);
kick;
run_to_done(200);
$display("T8 recovery on a healthy bus drives nothing");
ck_int("T8 diagnosed nothing wrong", diagnosis, D_NONE);
ck_int("T8 no remedy", remedy, R_NONE);
ck_int("T8 zero pulses", pulses_issued, 0);
ck_bit("T8 reported healthy", recovered, 1'b1);
ck_bit("T8 no HW reset", assert_hw_reset, 1'b0);
// ----------------------------------------------------------------
// T9. SCL FAILS MID-RECOVERY. The recovery starts legitimately on a stuck
// SDA, and after two pulses SCL is seized as well. The block must
// notice and abandon clocking rather than carry on pulsing a dead line.
// ----------------------------------------------------------------
setup(1'b1, 1'b0, 0);
kick;
// let a couple of pulses happen
n = 0;
while (pulses_issued < 2 && n < 100) begin step; n = n + 1; end
ck_int("T9 clocking was under way", diagnosis, D_SDA_STUCK);
@(negedge clk); holder_holds_scl = 1'b1; // SCL seized
run_to_done(300);
$display("T9 SCL failing mid-recovery abandons the clocking");
ck_int("T9 re-diagnosed as SCL stuck", diagnosis, D_SCL_STUCK);
ck_bit("T9 STILL never clocked a stuck SCL", attempted_clocks_on_scl, 1'b0);
ck_int("T9 escalated", remedy, R_HW_RESET);
if (pulses_issued > 9) begin
$display(" FAIL T9 pulses_issued exceeded nine: %0d", pulses_issued);
errors = errors + 1;
end
// ----------------------------------------------------------------
// T10. A successful recovery ends with a STOP. SDA must be pulled low and
// then released while SCL is high, or every device is left believing
// a transfer is still in progress.
// ----------------------------------------------------------------
setup(1'b1, 1'b0, 2);
kick;
// walk to the closing STOP and watch it
n = 0;
while (state != S_STOP1 && n < 200) begin step; n = n + 1; end
ck_int("T10 reached the closing STOP", state, S_STOP1);
step;
ck_bit("T10 SDA pulled low for the STOP", sda_drive_low, 1'b1);
ck_bit("T10 SCL released (high)", scl_in, 1'b1);
step;
$display("T10 a successful recovery ends with a STOP");
ck_bit("T10 SDA released again", sda_drive_low, 1'b0);
ck_bit("T10 bus left idle high", sda_in & scl_in, 1'b1);
ck_bit("T10 recovered", recovered, 1'b1);
// ----------------------------------------------------------------
// T11. Two recoveries in a row. The counters must accumulate and no state
// may leak from the first attempt into the second.
// ----------------------------------------------------------------
setup(1'b1, 1'b0, 4);
kick;
run_to_done(300);
ck_int("T11 first used four pulses", pulses_issued, 4);
@(negedge clk);
want_hold_sda = 1'b1;
release_after_pulses = 2;
clear_count = 1'b1;
step;
@(negedge clk); clear_count = 1'b0;
kick;
run_to_done(300);
$display("T11 two recoveries, counters accumulate, nothing leaks");
ck_int("T11 second used two pulses", pulses_issued, 2);
ck_int("T11 two attempts", recovery_attempts, 2);
ck_int("T11 two successes", successes, 2);
ck_bit("T11 never clocked a stuck SCL", attempted_clocks_on_scl, 1'b0);
// ----------------------------------------------------------------
// T12. The standing property, over every test above: this block has never
// attempted the nine-clock procedure on a held clock.
// ----------------------------------------------------------------
$display("T12 the standing property held throughout");
ck_bit("T12 attempted_clocks_on_scl is still zero", attempted_clocks_on_scl, 1'b0);
if (errors == 0)
$display("=== i2c_bus_recovery: ALL CHECKS PASSED ===");
else
$display("=== i2c_bus_recovery: %0d CHECK(S) FAILED ===", errors);
$finish;
end
endmodule // -----------------------------------------------------------------------------
// i2c_bus_recovery.sv
// Bus clear and recovery sequencer (UM10204 3.1.16).
//
// Section 3.1.16 is four sentences long and it says two different things about two
// different wires. That asymmetry is the whole design:
//
// "In the unlikely event where the CLOCK (SCL) is stuck LOW, the preferential
// procedure is to reset the bus using the HW reset signal if your I2C devices
// have HW reset inputs. If the I2C devices do not have HW reset inputs, cycle
// power to the devices to activate the mandatory internal Power-On Reset (POR)
// circuit."
//
// "If the DATA line (SDA) is stuck LOW, the master should send nine clock
// pulses. The device that held the bus LOW should release it some time within
// those nine clocks. If not, then use the HW reset or cycle power to clear the
// bus."
//
// Nine clock pulses for a stuck SDA. NOTHING at the protocol level for a stuck
// SCL. And the reason is not an omission -- it is that the nine pulses ARE pulses
// on SCL. A master cannot clock a line something else is holding low, so the
// remedy that works for the line the master shares is unavailable for the line
// the master owns.
//
// stuck SDA: nine clocks -> HW reset -> power cycle
// stuck SCL: HW reset -> power cycle
//
// Six obligations:
//
// 1. Diagnose WHICH line is stuck before choosing a remedy. They are not
// interchangeable and the wrong one wastes the only recovery attempt.
// 2. Never attempt the nine-clock procedure on a stuck SCL. The block exposes
// `attempted_clocks_on_scl` for the sole purpose of proving it never does.
// 3. Issue at most nine pulses, and stop as soon as SDA is released -- the
// specification says the holder "should release it some time WITHIN those
// nine clocks", so nine is a bound, not a quota.
// 4. Report how many pulses were actually needed. A bus that needs eight every
// time is telling you something a pass/fail flag would hide.
// 5. Escalate only after the nine pulses have failed, and choose the rung by
// what the devices actually support.
// 6. Finish with a STOP so the bus is left idle rather than merely unstuck.
// A released SDA with no framing leaves every device mid-transaction.
// -----------------------------------------------------------------------------
// (Verilog-2001 -- structurally identical to the SystemVerilog above.)
module i2c_bus_recovery #(
parameter CNT_W = 8
) (
input wire clk,
input wire rst_n,
// ---- bus observation ----------------------------------------------------
input wire scl_in, // the SCL line as it actually reads
input wire sda_in, // the SDA line as it actually reads
input wire begin_recovery, // pulse: something decided the bus is stuck
// ---- what the board can do ---------------------------------------------
input wire has_hw_reset, // the devices have HW reset inputs
input wire can_cycle_power, // the supply can be cycled
// ---- bus driving --------------------------------------------------------
output reg scl_drive_low, // pull SCL low (generating a pulse)
output reg sda_drive_low, // pull SDA low (for the closing STOP)
// ---- remedy ------------------------------------------------------------
output reg [2:0] remedy,
output reg assert_hw_reset,
output reg request_power_cycle,
// ---- status ------------------------------------------------------------
output reg [2:0] state,
output reg [2:0] diagnosis,
output reg [3:0] pulses_issued,
output reg recovered,
output reg gave_up,
output reg attempted_clocks_on_scl, // must remain 0, always
output reg [CNT_W-1:0] recovery_attempts,
output reg [CNT_W-1:0] successes
);
localparam [2:0] D_NONE = 3'd0, // nothing wrong
D_SDA_STUCK = 3'd1, // SDA low, SCL free
D_SCL_STUCK = 3'd2, // SCL low
D_BOTH = 3'd3; // both low
localparam [2:0] R_NONE = 3'd0,
R_NINE_CLOCKS = 3'd1,
R_HW_RESET = 3'd2,
R_POWER_CYCLE = 3'd3,
R_IMPOSSIBLE = 3'd4; // stuck, and nothing on the board can help
localparam [2:0] S_IDLE = 3'd0,
S_DIAG = 3'd1,
S_CLOCK = 3'd2, // issuing the nine pulses
S_STOP1 = 3'd3, // closing STOP: SDA low, SCL rising
S_STOP2 = 3'd4, // closing STOP: SDA released while SCL high
S_ESC = 3'd5, // escalating past the nine pulses
S_DONE = 3'd6;
// Three phases per pulse, not two. SDA must be SAMPLED while SCL is HIGH --
// Chapter 4.2's data-valid rule -- so releasing SCL and reading SDA cannot be
// the same step. A two-phase pulse reads SDA at the instant of release, before
// a device that let go during the low phase can be seen to have done so, and
// then needs one extra pulse to notice. That extra pulse is an artefact of the
// sampling point, and it would be reported as if the bus had been slower to
// recover than it was.
localparam [1:0] PH_LOW = 2'd0, // pull SCL low: the low phase begins
PH_REL = 2'd1, // release SCL: it rises
PH_SMP = 2'd2; // SCL is high: now read SDA
reg [1:0] pulse_phase;
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
state <= S_IDLE;
scl_drive_low <= 1'b0;
sda_drive_low <= 1'b0;
remedy <= R_NONE;
assert_hw_reset <= 1'b0;
request_power_cycle <= 1'b0;
diagnosis <= D_NONE;
pulses_issued <= 4'd0;
recovered <= 1'b0;
gave_up <= 1'b0;
attempted_clocks_on_scl <= 1'b0;
recovery_attempts <= {CNT_W{1'b0}};
successes <= {CNT_W{1'b0}};
pulse_phase <= PH_LOW;
end else begin
case (state)
// ---------------------------------------------------------------
S_IDLE: begin
scl_drive_low <= 1'b0;
sda_drive_low <= 1'b0;
assert_hw_reset <= 1'b0;
request_power_cycle <= 1'b0;
if (begin_recovery) begin
recovered <= 1'b0;
gave_up <= 1'b0;
pulses_issued <= 4'd0;
pulse_phase <= PH_LOW;
remedy <= R_NONE;
recovery_attempts <= recovery_attempts + 1'b1;
state <= S_DIAG;
end
end
// ---------------------------------------------------------------
// Obligation 1. Diagnose the line, then choose. The order of these
// tests matters: SCL being low dominates, because a stuck SCL makes
// the nine-clock procedure impossible regardless of what SDA is doing.
// ---------------------------------------------------------------
S_DIAG: begin
if (!scl_in && !sda_in) begin
diagnosis <= D_BOTH;
state <= S_ESC; // obligation 2: no clocking
end else if (!scl_in) begin
diagnosis <= D_SCL_STUCK;
state <= S_ESC; // obligation 2: no clocking
end else if (!sda_in) begin
diagnosis <= D_SDA_STUCK;
remedy <= R_NINE_CLOCKS;
state <= S_CLOCK;
end else begin
// Nothing is stuck. Recovery was requested on a healthy bus,
// which is worth reporting rather than "fixing".
diagnosis <= D_NONE;
remedy <= R_NONE;
recovered <= 1'b1;
successes <= successes + 1'b1;
state <= S_DONE;
end
end
// ---------------------------------------------------------------
// Obligation 3 and 4. Up to nine pulses, stopping the instant SDA
// is released. Each pulse is one low half and one high half.
// ---------------------------------------------------------------
S_CLOCK: begin
// A stuck SCL here would mean the diagnosis was wrong; refuse to
// pretend to clock. This is obligation 2 enforced a second time,
// because a line can fail DURING recovery.
if (!scl_in && !scl_drive_low) begin
diagnosis <= D_SCL_STUCK;
scl_drive_low <= 1'b0;
state <= S_ESC;
end else if (pulse_phase == PH_LOW) begin
scl_drive_low <= 1'b1; // pull SCL low
pulse_phase <= PH_REL;
end else if (pulse_phase == PH_REL) begin
scl_drive_low <= 1'b0; // release it; SCL rises
pulse_phase <= PH_SMP;
end else begin
// SCL is high: this is where SDA is valid, so this is where the
// release is observed. The pulse is complete either way.
pulse_phase <= PH_LOW;
pulses_issued <= pulses_issued + 1'b1;
if (sda_in) begin
// Released. Obligation 6: leave the bus framed, not merely
// unstuck.
state <= S_STOP1;
end else if (pulses_issued + 1'b1 >= 4'd9) begin
// Nine pulses spent and still held. Obligation 5.
state <= S_ESC;
end
end
end
// ---------------------------------------------------------------
// Obligation 6. A STOP: SDA low while SCL is high, then SDA rising.
// ---------------------------------------------------------------
S_STOP1: begin
sda_drive_low <= 1'b1;
scl_drive_low <= 1'b0;
state <= S_STOP2;
end
S_STOP2: begin
sda_drive_low <= 1'b0; // SDA rises while SCL is high
recovered <= 1'b1;
successes <= successes + 1'b1;
state <= S_DONE;
end
// ---------------------------------------------------------------
// Obligation 5. Escalate by what the board supports. The order is
// the specification's: HW reset is "preferential", a power cycle is
// the fallback that invokes the mandatory POR.
// ---------------------------------------------------------------
S_ESC: begin
scl_drive_low <= 1'b0;
sda_drive_low <= 1'b0;
if (has_hw_reset) begin
remedy <= R_HW_RESET;
assert_hw_reset <= 1'b1;
end else if (can_cycle_power) begin
remedy <= R_POWER_CYCLE;
request_power_cycle <= 1'b1;
end else begin
// Neither rung exists. Saying so is the only honest output.
remedy <= R_IMPOSSIBLE;
gave_up <= 1'b1;
end
state <= S_DONE;
end
S_DONE: begin
if (begin_recovery) begin
recovered <= 1'b0;
gave_up <= 1'b0;
pulses_issued <= 4'd0;
pulse_phase <= PH_LOW;
remedy <= R_NONE;
assert_hw_reset <= 1'b0;
request_power_cycle <= 1'b0;
recovery_attempts <= recovery_attempts + 1'b1;
state <= S_DIAG;
end
end
default: state <= S_IDLE;
endcase
// Obligation 2, as a standing assertion rather than a comment. If the
// block ever drives SCL while SCL is being held low by somebody else,
// it is attempting the one thing 3.1.16 does not offer.
if (scl_drive_low && (diagnosis == D_SCL_STUCK || diagnosis == D_BOTH))
attempted_clocks_on_scl <= 1'b1;
end
end
endmodule `timescale 1ns/1ps
// -----------------------------------------------------------------------------
// i2c_bus_recovery_tb.sv
// Independent oracle for i2c_bus_recovery.
//
// The bench owns the bus. It models a holder that releases SDA after a chosen
// number of clock pulses, and it computes SCL and SDA as a wired-AND of the
// holder and the DUT -- so what the DUT observes is a value the bench derived,
// never the DUT's own opinion of the line.
//
// The property the suite exists for is a NEGATIVE one: the nine-clock procedure
// must never be attempted on a stuck SCL. Tests 4, 5 and 9 assert it, and one of
// them makes SCL fail in the MIDDLE of a recovery that started legitimately.
// -----------------------------------------------------------------------------
// (Verilog-2001 -- structurally identical to the SystemVerilog above.)
module i2c_bus_recovery_tb;
localparam [2:0] D_NONE = 3'd0, D_SDA_STUCK = 3'd1, D_SCL_STUCK = 3'd2, D_BOTH = 3'd3;
localparam [2:0] R_NONE = 3'd0, R_NINE_CLOCKS = 3'd1, R_HW_RESET = 3'd2,
R_POWER_CYCLE = 3'd3, R_IMPOSSIBLE = 3'd4;
localparam [2:0] S_IDLE = 3'd0, S_DIAG = 3'd1, S_CLOCK = 3'd2, S_STOP1 = 3'd3,
S_STOP2 = 3'd4, S_ESC = 3'd5, S_DONE = 3'd6;
reg clk = 1'b0;
reg rst_n = 1'b0;
reg begin_recovery = 1'b0;
reg has_hw_reset = 1'b1;
reg can_cycle_power = 1'b1;
wire scl_drive_low, sda_drive_low;
wire [2:0] remedy, state, diagnosis;
wire assert_hw_reset, request_power_cycle;
wire [3:0] pulses_issued;
wire recovered, gave_up, attempted_clocks_on_scl;
wire [7:0] recovery_attempts, successes;
// ---- the bench's model of the other devices on the bus -----------------
reg holder_holds_sda = 1'b0; // driven ONLY by the holder block below
reg holder_holds_scl = 1'b0;
reg want_hold_sda = 1'b0; // the stimulus ASKS for a hold
reg clear_count = 1'b0;
integer release_after_pulses = 0; // 0 = never releases
integer pulses_counted = 0;
reg scl_prev = 1'b1;
// THE WIRED-AND. The line is low if the DUT pulls, or the holder pulls.
wire scl_in = !(scl_drive_low || holder_holds_scl);
wire sda_in = !(sda_drive_low || holder_holds_sda);
integer errors = 0;
integer n;
i2c_bus_recovery #(.CNT_W(8)) dut (
.clk(clk), .rst_n(rst_n),
.scl_in(scl_in), .sda_in(sda_in), .begin_recovery(begin_recovery),
.has_hw_reset(has_hw_reset), .can_cycle_power(can_cycle_power),
.scl_drive_low(scl_drive_low), .sda_drive_low(sda_drive_low),
.remedy(remedy), .assert_hw_reset(assert_hw_reset),
.request_power_cycle(request_power_cycle),
.state(state), .diagnosis(diagnosis), .pulses_issued(pulses_issued),
.recovered(recovered), .gave_up(gave_up),
.attempted_clocks_on_scl(attempted_clocks_on_scl),
.recovery_attempts(recovery_attempts), .successes(successes));
always #5 clk = ~clk;
// The holder watches SCL and lets go after the configured number of pulses.
//
// It releases on the FALLING edge of SCL, not the rising one, because that is
// what a real device does: Chapter 4.2's data-valid rule says SDA may only
// change while SCL is LOW. Releasing on the rising edge would be both illegal
// and unobservable -- the master samples SDA on that same edge, so a release
// there costs an extra pulse to notice, and the extra pulse would be an
// artefact of the model rather than a property of the bus.
always @(posedge clk) begin
if (!rst_n) begin
scl_prev <= 1'b1;
pulses_counted = 0;
holder_holds_sda <= want_hold_sda;
end else begin
if (clear_count) begin
pulses_counted = 0;
holder_holds_sda <= want_hold_sda; // re-arm for the next attempt
end else if (!scl_in && scl_prev) begin // SCL falling: the low phase begins
pulses_counted = pulses_counted + 1;
if (release_after_pulses != 0 && pulses_counted >= release_after_pulses)
holder_holds_sda <= 1'b0;
end
scl_prev <= scl_in;
end
end
task step; begin @(posedge clk); @(negedge clk); end endtask
task setup (input holds_sda, input holds_scl, input integer rel);
begin
@(negedge clk);
rst_n = 1'b0;
begin_recovery = 1'b0;
release_after_pulses = rel;
clear_count = 1'b1;
want_hold_sda = holds_sda;
holder_holds_scl = holds_scl;
repeat (3) @(posedge clk);
@(negedge clk); clear_count = 1'b0; rst_n = 1'b1;
step;
end
endtask
task kick;
begin @(negedge clk); begin_recovery = 1'b1; @(posedge clk); @(negedge clk); begin_recovery = 1'b0; end
endtask
task run_to_done (input integer bound);
begin
n = 0;
while (state != S_DONE && n < bound) begin step; n = n + 1; end
if (n >= bound) begin
$display(" FAIL run_to_done: stuck in state %0d", state);
errors = errors + 1;
end
end
endtask
task ck_int (input [200*8:1] what, input integer got, input integer exp);
begin
if (got !== exp) begin
$display(" FAIL %0s: got %0d expected %0d", what, got, exp);
errors = errors + 1;
end
end
endtask
task ck_bit (input [200*8:1] what, input got, input exp);
begin
if (got !== exp) begin
$display(" FAIL %0s: got %0b expected %0b", what, got, exp);
errors = errors + 1;
end
end
endtask
initial begin
$display("=== i2c_bus_recovery: two stuck lines, two different escape hatches ===");
// ----------------------------------------------------------------
// T1. SDA stuck, holder releases after three pulses. The procedure must
// stop as soon as the line is free -- nine is a bound, not a quota.
// ----------------------------------------------------------------
setup(1'b1, 1'b0, 3);
kick;
run_to_done(200);
$display("T1 SDA stuck, released after three pulses");
ck_int("T1 diagnosed SDA", diagnosis, D_SDA_STUCK);
ck_int("T1 chose nine clocks", remedy, R_NINE_CLOCKS);
ck_bit("T1 recovered", recovered, 1'b1);
ck_int("T1 stopped at 3 pulses", pulses_issued, 3);
ck_bit("T1 no HW reset needed", assert_hw_reset, 1'b0);
ck_bit("T1 no power cycle", request_power_cycle, 1'b0);
ck_bit("T1 never clocked a stuck SCL", attempted_clocks_on_scl, 1'b0);
ck_int("T1 one success", successes, 1);
// ----------------------------------------------------------------
// T2. SDA stuck, released on the ninth pulse -- the last one allowed.
// ----------------------------------------------------------------
setup(1'b1, 1'b0, 9);
kick;
run_to_done(300);
$display("T2 released on the ninth pulse: still a success");
ck_bit("T2 recovered", recovered, 1'b1);
ck_int("T2 used all nine", pulses_issued, 9);
ck_int("T2 remedy was the clocks", remedy, R_NINE_CLOCKS);
ck_bit("T2 did not escalate", assert_hw_reset, 1'b0);
// ----------------------------------------------------------------
// T3. SDA stuck and NEVER released. Nine pulses, then escalation. The
// pulse count must stop at nine -- not ten, and not forever.
// ----------------------------------------------------------------
setup(1'b1, 1'b0, 0);
kick;
run_to_done(300);
$display("T3 never released: nine pulses then escalation");
ck_int("T3 exactly nine pulses", pulses_issued, 9);
ck_bit("T3 not recovered", recovered, 1'b0);
ck_int("T3 escalated to HW reset", remedy, R_HW_RESET);
ck_bit("T3 HW reset asserted", assert_hw_reset, 1'b1);
ck_bit("T3 no power cycle yet", request_power_cycle, 1'b0);
// ----------------------------------------------------------------
// T4. SCL STUCK. The nine-clock procedure does not exist for this line.
// The block must diagnose it, escalate immediately, and issue ZERO
// pulses -- because a pulse on a held line is not a pulse.
// ----------------------------------------------------------------
setup(1'b0, 1'b1, 0);
kick;
run_to_done(200);
$display("T4 SCL stuck: no clocking is attempted at all");
ck_int("T4 diagnosed SCL", diagnosis, D_SCL_STUCK);
ck_int("T4 zero pulses issued", pulses_issued, 0);
ck_bit("T4 NEVER clocked a stuck SCL", attempted_clocks_on_scl, 1'b0);
ck_int("T4 went straight to HW reset", remedy, R_HW_RESET);
ck_bit("T4 HW reset asserted", assert_hw_reset, 1'b1);
ck_bit("T4 not recovered by clocking", recovered, 1'b0);
// ----------------------------------------------------------------
// T5. BOTH lines stuck. Still no clocking, and the diagnosis says both
// rather than picking one.
// ----------------------------------------------------------------
setup(1'b1, 1'b1, 0);
kick;
run_to_done(200);
$display("T5 both lines stuck: diagnosed as both, still no clocking");
ck_int("T5 diagnosed both", diagnosis, D_BOTH);
ck_int("T5 zero pulses", pulses_issued, 0);
ck_bit("T5 never clocked", attempted_clocks_on_scl, 1'b0);
ck_int("T5 escalated", remedy, R_HW_RESET);
// ----------------------------------------------------------------
// T6. No HW reset available: the power-cycle rung. The specification calls
// the POR "mandatory", which is what makes this a last resort that
// always works.
// ----------------------------------------------------------------
setup(1'b0, 1'b1, 0);
@(negedge clk); has_hw_reset = 1'b0; can_cycle_power = 1'b1;
kick;
run_to_done(200);
$display("T6 without a HW reset input, cycle power");
ck_int("T6 remedy is a power cycle", remedy, R_POWER_CYCLE);
ck_bit("T6 power cycle requested", request_power_cycle, 1'b1);
ck_bit("T6 no HW reset", assert_hw_reset, 1'b0);
ck_bit("T6 did not give up", gave_up, 1'b0);
// ----------------------------------------------------------------
// T7. Neither rung exists. The only honest answer is that the bus cannot
// be recovered -- a checker that reported success here would be lying.
// ----------------------------------------------------------------
setup(1'b0, 1'b1, 0);
@(negedge clk); has_hw_reset = 1'b0; can_cycle_power = 1'b0;
kick;
run_to_done(200);
$display("T7 with neither remedy available, say so");
ck_int("T7 remedy is impossible", remedy, R_IMPOSSIBLE);
ck_bit("T7 gave up", gave_up, 1'b1);
ck_bit("T7 not recovered", recovered, 1'b0);
@(negedge clk); has_hw_reset = 1'b1; can_cycle_power = 1'b1;
// ----------------------------------------------------------------
// T8. Recovery requested on a HEALTHY bus. Nothing is stuck, so nothing
// should be driven -- and reporting that is more useful than running
// a procedure on a working bus.
// ----------------------------------------------------------------
setup(1'b0, 1'b0, 0);
kick;
run_to_done(200);
$display("T8 recovery on a healthy bus drives nothing");
ck_int("T8 diagnosed nothing wrong", diagnosis, D_NONE);
ck_int("T8 no remedy", remedy, R_NONE);
ck_int("T8 zero pulses", pulses_issued, 0);
ck_bit("T8 reported healthy", recovered, 1'b1);
ck_bit("T8 no HW reset", assert_hw_reset, 1'b0);
// ----------------------------------------------------------------
// T9. SCL FAILS MID-RECOVERY. The recovery starts legitimately on a stuck
// SDA, and after two pulses SCL is seized as well. The block must
// notice and abandon clocking rather than carry on pulsing a dead line.
// ----------------------------------------------------------------
setup(1'b1, 1'b0, 0);
kick;
// let a couple of pulses happen
n = 0;
while (pulses_issued < 2 && n < 100) begin step; n = n + 1; end
ck_int("T9 clocking was under way", diagnosis, D_SDA_STUCK);
@(negedge clk); holder_holds_scl = 1'b1; // SCL seized
run_to_done(300);
$display("T9 SCL failing mid-recovery abandons the clocking");
ck_int("T9 re-diagnosed as SCL stuck", diagnosis, D_SCL_STUCK);
ck_bit("T9 STILL never clocked a stuck SCL", attempted_clocks_on_scl, 1'b0);
ck_int("T9 escalated", remedy, R_HW_RESET);
if (pulses_issued > 9) begin
$display(" FAIL T9 pulses_issued exceeded nine: %0d", pulses_issued);
errors = errors + 1;
end
// ----------------------------------------------------------------
// T10. A successful recovery ends with a STOP. SDA must be pulled low and
// then released while SCL is high, or every device is left believing
// a transfer is still in progress.
// ----------------------------------------------------------------
setup(1'b1, 1'b0, 2);
kick;
// walk to the closing STOP and watch it
n = 0;
while (state != S_STOP1 && n < 200) begin step; n = n + 1; end
ck_int("T10 reached the closing STOP", state, S_STOP1);
step;
ck_bit("T10 SDA pulled low for the STOP", sda_drive_low, 1'b1);
ck_bit("T10 SCL released (high)", scl_in, 1'b1);
step;
$display("T10 a successful recovery ends with a STOP");
ck_bit("T10 SDA released again", sda_drive_low, 1'b0);
ck_bit("T10 bus left idle high", sda_in & scl_in, 1'b1);
ck_bit("T10 recovered", recovered, 1'b1);
// ----------------------------------------------------------------
// T11. Two recoveries in a row. The counters must accumulate and no state
// may leak from the first attempt into the second.
// ----------------------------------------------------------------
setup(1'b1, 1'b0, 4);
kick;
run_to_done(300);
ck_int("T11 first used four pulses", pulses_issued, 4);
@(negedge clk);
want_hold_sda = 1'b1;
release_after_pulses = 2;
clear_count = 1'b1;
step;
@(negedge clk); clear_count = 1'b0;
kick;
run_to_done(300);
$display("T11 two recoveries, counters accumulate, nothing leaks");
ck_int("T11 second used two pulses", pulses_issued, 2);
ck_int("T11 two attempts", recovery_attempts, 2);
ck_int("T11 two successes", successes, 2);
ck_bit("T11 never clocked a stuck SCL", attempted_clocks_on_scl, 1'b0);
// ----------------------------------------------------------------
// T12. The standing property, over every test above: this block has never
// attempted the nine-clock procedure on a held clock.
// ----------------------------------------------------------------
$display("T12 the standing property held throughout");
ck_bit("T12 attempted_clocks_on_scl is still zero", attempted_clocks_on_scl, 1'b0);
if (errors == 0)
$display("=== i2c_bus_recovery: ALL CHECKS PASSED ===");
else
$display("=== i2c_bus_recovery: %0d CHECK(S) FAILED ===", errors);
$finish;
end
endmodule -- ---------------------------------------------------------------------------
-- i2c_bus_recovery.vhd
-- Bus clear and recovery sequencer (UM10204 3.1.16).
-- Behavioural twin of i2c_bus_recovery.sv / .v.
--
-- Section 3.1.16 is four sentences long and it says two different things about two
-- different wires. That asymmetry is the whole design:
--
-- "In the unlikely event where the CLOCK (SCL) is stuck LOW, the preferential
-- procedure is to reset the bus using the HW reset signal if your I2C devices
-- have HW reset inputs. If the I2C devices do not have HW reset inputs, cycle
-- power to the devices to activate the mandatory internal Power-On Reset (POR)
-- circuit."
--
-- "If the DATA line (SDA) is stuck LOW, the master should send nine clock
-- pulses. The device that held the bus LOW should release it some time within
-- those nine clocks. If not, then use the HW reset or cycle power to clear the
-- bus."
--
-- Nine clock pulses for a stuck SDA. NOTHING at the protocol level for a stuck
-- SCL -- and the reason is not an omission. The nine pulses ARE pulses on SCL, so
-- a master cannot clock a line something else is holding low. The remedy that
-- works for the line the master shares is unavailable for the line it owns.
--
-- stuck SDA: nine clocks -> HW reset -> power cycle
-- stuck SCL: HW reset -> power cycle
--
-- Six obligations:
-- 1. Diagnose WHICH line is stuck before choosing a remedy.
-- 2. Never attempt the nine-clock procedure on a stuck SCL. The output
-- attempted_clocks_on_scl exists solely to prove it never does.
-- 3. Issue at most nine pulses, and stop as soon as SDA is released -- the
-- release comes "some time WITHIN those nine clocks", so nine is a bound.
-- 4. Report how many pulses were actually needed.
-- 5. Escalate only after the nine pulses have failed, by what is supported.
-- 6. Finish with a STOP so the bus is left idle rather than merely unstuck.
-- ---------------------------------------------------------------------------
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity i2c_bus_recovery is
generic (
CNT_W : integer := 8
);
port (
clk : in std_logic;
rst_n : in std_logic;
-- bus observation
scl_in : in std_logic; -- the SCL line as it actually reads
sda_in : in std_logic; -- the SDA line as it actually reads
begin_recovery : in std_logic; -- pulse: something decided the bus is stuck
-- what the board can do
has_hw_reset : in std_logic;
can_cycle_power : in std_logic;
-- bus driving
scl_drive_low : out std_logic; -- pull SCL low (generating a pulse)
sda_drive_low : out std_logic; -- pull SDA low (for the closing STOP)
-- remedy
remedy : out unsigned(2 downto 0);
assert_hw_reset : out std_logic;
request_power_cycle : out std_logic;
-- status
state : out unsigned(2 downto 0);
diagnosis : out unsigned(2 downto 0);
pulses_issued : out unsigned(3 downto 0);
recovered : out std_logic;
gave_up : out std_logic;
attempted_clocks_on_scl : out std_logic; -- must remain 0, always
recovery_attempts : out unsigned(CNT_W-1 downto 0);
successes : out unsigned(CNT_W-1 downto 0)
);
end entity i2c_bus_recovery;
architecture rtl of i2c_bus_recovery is
constant D_NONE : integer := 0; -- nothing wrong
constant D_SDA_STUCK : integer := 1; -- SDA low, SCL free
constant D_SCL_STUCK : integer := 2; -- SCL low
constant D_BOTH : integer := 3; -- both low
constant R_NONE : integer := 0;
constant R_NINE_CLOCKS : integer := 1;
constant R_HW_RESET : integer := 2;
constant R_POWER_CYCLE : integer := 3;
constant R_IMPOSSIBLE : integer := 4; -- stuck, and nothing on the board helps
constant ST_IDLE : integer := 0;
constant ST_DIAG : integer := 1;
constant ST_CLOCK : integer := 2; -- issuing the nine pulses
constant ST_STOP1 : integer := 3; -- closing STOP: SDA low, SCL high
constant ST_STOP2 : integer := 4; -- closing STOP: SDA released while SCL high
constant ST_ESC : integer := 5; -- escalating past the nine pulses
constant ST_DONE : integer := 6;
-- Three phases per pulse, not two. SDA must be SAMPLED while SCL is HIGH --
-- Chapter 4.2's data-valid rule -- so releasing SCL and reading SDA cannot be
-- the same step. A two-phase pulse reads SDA at the instant of release, before
-- a device that let go during the low phase can be seen to have done so, and
-- then needs one extra pulse to notice. That extra pulse would be an artefact
-- of the sampling point reported as if the bus had been slower to recover.
constant PH_LOW : integer := 0; -- pull SCL low: the low phase begins
constant PH_REL : integer := 1; -- release SCL: it rises
constant PH_SMP : integer := 2; -- SCL is high: now read SDA
signal st : integer := ST_IDLE;
signal diag : integer := D_NONE;
signal ph : integer := PH_LOW;
signal pulses : integer := 0;
signal n_att : integer := 0;
signal n_ok : integer := 0;
signal scl_drv : std_logic := '0';
begin
state <= to_unsigned(st, 3);
diagnosis <= to_unsigned(diag, 3);
pulses_issued <= to_unsigned(pulses, 4);
scl_drive_low <= scl_drv;
process (clk, rst_n)
begin
if rst_n = '0' then
st <= ST_IDLE;
scl_drv <= '0';
sda_drive_low <= '0';
remedy <= to_unsigned(R_NONE, 3);
assert_hw_reset <= '0';
request_power_cycle <= '0';
diag <= D_NONE;
ph <= PH_LOW;
pulses <= 0;
recovered <= '0';
gave_up <= '0';
attempted_clocks_on_scl <= '0';
n_att <= 0;
n_ok <= 0;
recovery_attempts <= (others => '0');
successes <= (others => '0');
elsif rising_edge(clk) then
case st is
when ST_IDLE =>
scl_drv <= '0';
sda_drive_low <= '0';
assert_hw_reset <= '0';
request_power_cycle <= '0';
if begin_recovery = '1' then
recovered <= '0';
gave_up <= '0';
pulses <= 0;
ph <= PH_LOW;
remedy <= to_unsigned(R_NONE, 3);
n_att <= n_att + 1;
recovery_attempts <= to_unsigned(n_att + 1, CNT_W);
st <= ST_DIAG;
end if;
-- Obligation 1. Diagnose the line, then choose. The order matters:
-- SCL being low dominates, because a stuck SCL makes the nine-clock
-- procedure impossible regardless of what SDA is doing.
when ST_DIAG =>
if scl_in = '0' and sda_in = '0' then
diag <= D_BOTH;
st <= ST_ESC; -- obligation 2: no clocking
elsif scl_in = '0' then
diag <= D_SCL_STUCK;
st <= ST_ESC; -- obligation 2: no clocking
elsif sda_in = '0' then
diag <= D_SDA_STUCK;
remedy <= to_unsigned(R_NINE_CLOCKS, 3);
st <= ST_CLOCK;
else
-- Nothing is stuck. Recovery was requested on a healthy bus,
-- which is worth reporting rather than "fixing".
diag <= D_NONE;
remedy <= to_unsigned(R_NONE, 3);
recovered <= '1';
n_ok <= n_ok + 1;
successes <= to_unsigned(n_ok + 1, CNT_W);
st <= ST_DONE;
end if;
-- Obligations 3 and 4. Up to nine pulses, stopping the instant SDA is
-- released. Each pulse is a low phase, a release, and a sample.
when ST_CLOCK =>
-- A stuck SCL here means the diagnosis was wrong; refuse to
-- pretend to clock. Obligation 2 enforced a second time, because a
-- line can fail DURING recovery.
if scl_in = '0' and scl_drv = '0' then
diag <= D_SCL_STUCK;
scl_drv <= '0';
st <= ST_ESC;
elsif ph = PH_LOW then
scl_drv <= '1'; -- pull SCL low
ph <= PH_REL;
elsif ph = PH_REL then
scl_drv <= '0'; -- release it; SCL rises
ph <= PH_SMP;
else
-- SCL is high: this is where SDA is valid, so this is where the
-- release is observed. The pulse is complete either way.
ph <= PH_LOW;
pulses <= pulses + 1;
if sda_in = '1' then
-- Released. Obligation 6: leave the bus framed.
st <= ST_STOP1;
elsif pulses + 1 >= 9 then
-- Nine pulses spent and still held. Obligation 5.
st <= ST_ESC;
end if;
end if;
-- Obligation 6. A STOP: SDA low while SCL is high, then SDA rising.
when ST_STOP1 =>
sda_drive_low <= '1';
scl_drv <= '0';
st <= ST_STOP2;
when ST_STOP2 =>
sda_drive_low <= '0'; -- SDA rises while SCL is high
recovered <= '1';
n_ok <= n_ok + 1;
successes <= to_unsigned(n_ok + 1, CNT_W);
st <= ST_DONE;
-- Obligation 5. Escalate by what the board supports. The order is the
-- specification's: HW reset is "preferential", a power cycle is the
-- fallback that invokes the mandatory POR.
when ST_ESC =>
scl_drv <= '0';
sda_drive_low <= '0';
if has_hw_reset = '1' then
remedy <= to_unsigned(R_HW_RESET, 3);
assert_hw_reset <= '1';
elsif can_cycle_power = '1' then
remedy <= to_unsigned(R_POWER_CYCLE, 3);
request_power_cycle <= '1';
else
-- Neither rung exists. Saying so is the only honest output.
remedy <= to_unsigned(R_IMPOSSIBLE, 3);
gave_up <= '1';
end if;
st <= ST_DONE;
when ST_DONE =>
if begin_recovery = '1' then
recovered <= '0';
gave_up <= '0';
pulses <= 0;
ph <= PH_LOW;
remedy <= to_unsigned(R_NONE, 3);
assert_hw_reset <= '0';
request_power_cycle <= '0';
n_att <= n_att + 1;
recovery_attempts <= to_unsigned(n_att + 1, CNT_W);
st <= ST_DIAG;
end if;
when others =>
st <= ST_IDLE;
end case;
-- Obligation 2, as a standing assertion rather than a comment. If the
-- block ever drives SCL while SCL is being held low by somebody else, it
-- is attempting the one thing 3.1.16 does not offer.
if scl_drv = '1' and (diag = D_SCL_STUCK or diag = D_BOTH) then
attempted_clocks_on_scl <= '1';
end if;
end if;
end process;
end architecture rtl; -- ---------------------------------------------------------------------------
-- i2c_bus_recovery_tb.vhd
-- Independent oracle for i2c_bus_recovery. Behavioural twin of the SystemVerilog
-- and Verilog benches.
--
-- The bench owns the bus. It models a holder that releases SDA after a chosen
-- number of clock pulses, and it computes SCL and SDA as a wired-AND of the holder
-- and the DUT -- so what the DUT observes is a value the bench derived.
--
-- The property the suite exists for is a NEGATIVE one: the nine-clock procedure
-- must never be attempted on a stuck SCL. Tests 4, 5, 9 and 12 assert it, and one
-- of them makes SCL fail in the MIDDLE of a recovery that started legitimately.
-- ---------------------------------------------------------------------------
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity i2c_bus_recovery_tb is
end entity i2c_bus_recovery_tb;
architecture sim of i2c_bus_recovery_tb is
constant TCLK : time := 10 ns;
constant D_NONE : integer := 0;
constant D_SDA_STUCK : integer := 1;
constant D_SCL_STUCK : integer := 2;
constant D_BOTH : integer := 3;
constant R_NONE : integer := 0;
constant R_NINE_CLOCKS : integer := 1;
constant R_HW_RESET : integer := 2;
constant R_POWER_CYCLE : integer := 3;
constant R_IMPOSSIBLE : integer := 4;
constant ST_CLOCK : integer := 2;
constant ST_STOP1 : integer := 3;
constant ST_DONE : integer := 6;
signal clk : std_logic := '0';
signal rst_n : std_logic := '0';
signal begin_recovery : std_logic := '0';
signal has_hw_reset : std_logic := '1';
signal can_cycle_power : std_logic := '1';
signal scl_drive_low, sda_drive_low : std_logic;
signal remedy, st_o, diagnosis : unsigned(2 downto 0);
signal assert_hw_reset, request_power_cycle : std_logic;
signal pulses_issued : unsigned(3 downto 0);
signal recovered, gave_up, attempted_clocks_on_scl : std_logic;
signal recovery_attempts, successes : unsigned(7 downto 0);
-- the bench's model of the other devices on the bus
signal holder_holds_sda : std_logic := '0'; -- driven ONLY by the holder process
signal holder_holds_scl : std_logic := '0';
signal want_hold_sda : std_logic := '0'; -- the stimulus ASKS for a hold
signal release_after : integer := 0; -- 0 = never releases
signal clear_count : std_logic := '0';
-- THE WIRED-AND. The line is low if the DUT pulls, or the holder pulls.
signal scl_in, sda_in : std_logic;
signal halt : boolean := false;
begin
scl_in <= '0' when (scl_drive_low = '1' or holder_holds_scl = '1') else '1';
sda_in <= '0' when (sda_drive_low = '1' or holder_holds_sda = '1') else '1';
dut : entity work.i2c_bus_recovery
generic map (CNT_W => 8)
port map (
clk => clk, rst_n => rst_n,
scl_in => scl_in, sda_in => sda_in, begin_recovery => begin_recovery,
has_hw_reset => has_hw_reset, can_cycle_power => can_cycle_power,
scl_drive_low => scl_drive_low, sda_drive_low => sda_drive_low,
remedy => remedy, assert_hw_reset => assert_hw_reset,
request_power_cycle => request_power_cycle,
state => st_o, diagnosis => diagnosis, pulses_issued => pulses_issued,
recovered => recovered, gave_up => gave_up,
attempted_clocks_on_scl => attempted_clocks_on_scl,
recovery_attempts => recovery_attempts, successes => successes);
clkgen : process
begin
while not halt loop
clk <= '0'; wait for TCLK/2;
clk <= '1'; wait for TCLK/2;
end loop;
wait;
end process;
-- The holder watches SCL and lets go after the configured number of pulses.
--
-- It releases on the FALLING edge of SCL, not the rising one, because that is
-- what a real device does: Chapter 4.2's data-valid rule says SDA may only
-- change while SCL is LOW. Releasing on the rising edge would be both illegal
-- and unobservable -- the master samples SDA in the high phase.
--
-- holder_holds_sda has exactly ONE driver, this process, because VHDL permits
-- only one for a signal of an unresolved type. The stimulus asks for a change
-- through release_after and clear_count instead of writing it directly.
holder : process (clk, rst_n)
variable scl_prev : std_logic := '1';
variable counted : integer := 0;
begin
if rst_n = '0' then
scl_prev := '1';
counted := 0;
holder_holds_sda <= want_hold_sda;
elsif rising_edge(clk) then
if clear_count = '1' then
counted := 0;
holder_holds_sda <= want_hold_sda; -- re-arm for the next attempt
elsif scl_in = '0' and scl_prev = '1' then -- SCL falling
counted := counted + 1;
if release_after /= 0 and counted >= release_after then
holder_holds_sda <= '0';
end if;
end if;
scl_prev := scl_in;
end if;
end process;
stim : process
variable err : integer := 0;
variable n : integer;
procedure ck_int (what : string; got : integer; exp : integer) is
begin
if got /= exp then
report " FAIL " & what & ": got " & integer'image(got)
& " expected " & integer'image(exp) severity note;
err := err + 1;
end if;
end procedure;
procedure ck_bit (what : string; got : std_logic; exp : std_logic) is
begin
if got /= exp then
report " FAIL " & what & ": got " & std_logic'image(got)
& " expected " & std_logic'image(exp) severity note;
err := err + 1;
end if;
end procedure;
procedure step is
begin
wait until rising_edge(clk);
wait until falling_edge(clk);
end procedure;
procedure setup (holds_sda : std_logic; holds_scl : std_logic; rel : integer) is
begin
wait until falling_edge(clk);
rst_n <= '0';
begin_recovery <= '0';
release_after <= rel;
clear_count <= '1';
want_hold_sda <= holds_sda;
holder_holds_scl <= holds_scl;
for k in 0 to 2 loop wait until rising_edge(clk); end loop;
wait until falling_edge(clk);
clear_count <= '0';
rst_n <= '1';
step;
end procedure;
procedure kick is
begin
wait until falling_edge(clk); begin_recovery <= '1';
wait until rising_edge(clk); wait until falling_edge(clk);
begin_recovery <= '0';
end procedure;
procedure run_to_done (bound : integer) is
begin
n := 0;
while to_integer(st_o) /= ST_DONE and n < bound loop
step; n := n + 1;
end loop;
if n >= bound then
report " FAIL run_to_done: stuck in state "
& integer'image(to_integer(st_o)) severity note;
err := err + 1;
end if;
end procedure;
begin
report "=== i2c_bus_recovery: two stuck lines, two different escape hatches ==="
severity note;
-- T1. SDA stuck, released after three pulses: nine is a bound, not a quota.
setup('1', '0', 3);
kick;
run_to_done(200);
report "T1 SDA stuck, released after three pulses" severity note;
ck_int("T1 diagnosed SDA", to_integer(diagnosis), D_SDA_STUCK);
ck_int("T1 chose nine clocks", to_integer(remedy), R_NINE_CLOCKS);
ck_bit("T1 recovered", recovered, '1');
ck_int("T1 stopped at 3 pulses", to_integer(pulses_issued), 3);
ck_bit("T1 no HW reset needed", assert_hw_reset, '0');
ck_bit("T1 no power cycle", request_power_cycle, '0');
ck_bit("T1 never clocked a stuck SCL", attempted_clocks_on_scl, '0');
ck_int("T1 one success", to_integer(successes), 1);
-- T2. Released on the ninth pulse: the last one allowed.
setup('1', '0', 9);
kick;
run_to_done(300);
report "T2 released on the ninth pulse: still a success" severity note;
ck_bit("T2 recovered", recovered, '1');
ck_int("T2 used all nine", to_integer(pulses_issued), 9);
ck_int("T2 remedy was the clocks", to_integer(remedy), R_NINE_CLOCKS);
ck_bit("T2 did not escalate", assert_hw_reset, '0');
-- T3. Never released: nine pulses, then escalation. Not ten, not forever.
setup('1', '0', 0);
kick;
run_to_done(300);
report "T3 never released: nine pulses then escalation" severity note;
ck_int("T3 exactly nine pulses", to_integer(pulses_issued), 9);
ck_bit("T3 not recovered", recovered, '0');
ck_int("T3 escalated to HW reset", to_integer(remedy), R_HW_RESET);
ck_bit("T3 HW reset asserted", assert_hw_reset, '1');
ck_bit("T3 no power cycle yet", request_power_cycle, '0');
-- T4. SCL STUCK. Zero pulses, because a pulse on a held line is not a pulse.
setup('0', '1', 0);
kick;
run_to_done(200);
report "T4 SCL stuck: no clocking is attempted at all" severity note;
ck_int("T4 diagnosed SCL", to_integer(diagnosis), D_SCL_STUCK);
ck_int("T4 zero pulses issued", to_integer(pulses_issued), 0);
ck_bit("T4 NEVER clocked a stuck SCL", attempted_clocks_on_scl, '0');
ck_int("T4 went straight to HW reset", to_integer(remedy), R_HW_RESET);
ck_bit("T4 HW reset asserted", assert_hw_reset, '1');
ck_bit("T4 not recovered by clocking", recovered, '0');
-- T5. Both lines stuck: still no clocking, diagnosed as both.
setup('1', '1', 0);
kick;
run_to_done(200);
report "T5 both lines stuck: diagnosed as both, still no clocking" severity note;
ck_int("T5 diagnosed both", to_integer(diagnosis), D_BOTH);
ck_int("T5 zero pulses", to_integer(pulses_issued), 0);
ck_bit("T5 never clocked", attempted_clocks_on_scl, '0');
ck_int("T5 escalated", to_integer(remedy), R_HW_RESET);
-- T6. No HW reset available: the power-cycle rung.
setup('0', '1', 0);
wait until falling_edge(clk);
has_hw_reset <= '0'; can_cycle_power <= '1';
kick;
run_to_done(200);
report "T6 without a HW reset input, cycle power" severity note;
ck_int("T6 remedy is a power cycle", to_integer(remedy), R_POWER_CYCLE);
ck_bit("T6 power cycle requested", request_power_cycle, '1');
ck_bit("T6 no HW reset", assert_hw_reset, '0');
ck_bit("T6 did not give up", gave_up, '0');
-- T7. Neither rung exists: say so.
setup('0', '1', 0);
wait until falling_edge(clk);
has_hw_reset <= '0'; can_cycle_power <= '0';
kick;
run_to_done(200);
report "T7 with neither remedy available, say so" severity note;
ck_int("T7 remedy is impossible", to_integer(remedy), R_IMPOSSIBLE);
ck_bit("T7 gave up", gave_up, '1');
ck_bit("T7 not recovered", recovered, '0');
wait until falling_edge(clk);
has_hw_reset <= '1'; can_cycle_power <= '1';
-- T8. Recovery on a HEALTHY bus drives nothing.
setup('0', '0', 0);
kick;
run_to_done(200);
report "T8 recovery on a healthy bus drives nothing" severity note;
ck_int("T8 diagnosed nothing wrong", to_integer(diagnosis), D_NONE);
ck_int("T8 no remedy", to_integer(remedy), R_NONE);
ck_int("T8 zero pulses", to_integer(pulses_issued), 0);
ck_bit("T8 reported healthy", recovered, '1');
ck_bit("T8 no HW reset", assert_hw_reset, '0');
-- T9. SCL FAILS MID-RECOVERY. Started on a stuck SDA; SCL is then seized.
setup('1', '0', 0);
kick;
n := 0;
while to_integer(pulses_issued) < 2 and n < 100 loop step; n := n + 1; end loop;
ck_int("T9 clocking was under way", to_integer(diagnosis), D_SDA_STUCK);
wait until falling_edge(clk); holder_holds_scl <= '1';
run_to_done(300);
report "T9 SCL failing mid-recovery abandons the clocking" severity note;
ck_int("T9 re-diagnosed as SCL stuck", to_integer(diagnosis), D_SCL_STUCK);
ck_bit("T9 STILL never clocked a stuck SCL", attempted_clocks_on_scl, '0');
ck_int("T9 escalated", to_integer(remedy), R_HW_RESET);
if to_integer(pulses_issued) > 9 then
report " FAIL T9 pulses_issued exceeded nine" severity note;
err := err + 1;
end if;
-- T10. A successful recovery ends with a STOP.
setup('1', '0', 2);
kick;
n := 0;
while to_integer(st_o) /= ST_STOP1 and n < 200 loop step; n := n + 1; end loop;
ck_int("T10 reached the closing STOP", to_integer(st_o), ST_STOP1);
step;
ck_bit("T10 SDA pulled low for the STOP", sda_drive_low, '1');
ck_bit("T10 SCL released (high)", scl_in, '1');
step;
report "T10 a successful recovery ends with a STOP" severity note;
ck_bit("T10 SDA released again", sda_drive_low, '0');
ck_bit("T10 bus left idle high", sda_in and scl_in, '1');
ck_bit("T10 recovered", recovered, '1');
-- T11. Two recoveries in a row: counters accumulate, nothing leaks.
setup('1', '0', 4);
kick;
run_to_done(300);
ck_int("T11 first used four pulses", to_integer(pulses_issued), 4);
wait until falling_edge(clk);
want_hold_sda <= '1';
release_after <= 2;
clear_count <= '1';
step;
wait until falling_edge(clk);
clear_count <= '0';
kick;
run_to_done(300);
report "T11 two recoveries, counters accumulate, nothing leaks" severity note;
ck_int("T11 second used two pulses", to_integer(pulses_issued), 2);
ck_int("T11 two attempts", to_integer(recovery_attempts), 2);
ck_int("T11 two successes", to_integer(successes), 2);
ck_bit("T11 never clocked a stuck SCL", attempted_clocks_on_scl, '0');
-- T12. The standing property, over every test above.
report "T12 the standing property held throughout" severity note;
ck_bit("T12 attempted_clocks_on_scl is still zero",
attempted_clocks_on_scl, '0');
if err = 0 then
report "=== i2c_bus_recovery: ALL CHECKS PASSED ===" severity note;
else
report "=== i2c_bus_recovery: " & integer'image(err)
& " CHECK(S) FAILED ===" severity note;
end if;
halt <= true;
wait;
end process;
end architecture sim;6a. Seven Decisions Worth Defending
The diagnosis comes before the remedy, and SCL dominates. S_DIAG tests both lines and orders the tests so that a low clock decides the outcome regardless of what SDA is doing — because a stuck clock makes the nine-pulse procedure impossible whatever else is true. Mutation Z2 sends a both-lines-stuck bus into the clocking state and is killed.
There are two independent guards against clocking a held SCL, and they are not redundant with each other. One is in S_DIAG, at the point the remedy is chosen. The other is the first statement of S_CLOCK, which re-tests the clock on every pulse — because a line can fail during a recovery. §7 shows that the second guard subsumes the first for an initially-stuck clock but not the reverse, which is why removing the second one is caught and removing the first one is not.
attempted_clocks_on_scl is a standing assertion written as hardware. It latches if the block ever drives SCL while its own diagnosis says SCL is held. It is not a functional output and nothing reads it in a real system; it exists so a testbench can assert the one property this chapter is about, over an entire run rather than at an instant.
SDA is sampled while SCL is high, which needs a three-phase pulse. A two-phase pulse — pull low, release — reads SDA at the instant of release, before a device that let go during the low phase can be seen to have done so, and then needs one extra pulse to notice. §4's second detail, and it is not a theoretical concern: an early version of this design had exactly that shape, and the testbench caught it by asserting the pulse count rather than only the outcome. Mutation Z6 restores the two-phase form and sixteen checks fail.
Nine is a bound and the count is reported. The procedure stops the instant SDA reads high, and pulses_issued says how many it took. Mutation Z5 makes it always spend nine and is killed by the tests that assert three and four; mutation Z4 allows ten and is killed by the test that asserts exactly nine on a line that never releases.
Escalation is ordered by the specification and gated on what exists. HW reset is "preferential", so it is tried first; the power cycle is the fallback. Mutation Z7 inverts the order and seven checks fail. And when neither is available the design reports R_IMPOSSIBLE and gave_up rather than either remedy — mutation Z8 claims success there and is killed.
A recovery request on a healthy bus drives nothing. It is reported as a success with a diagnosis of D_NONE and zero pulses, because running the procedure on a working bus is both unnecessary and a way to corrupt a transfer that was merely slow. Mutation Z10 runs it anyway and two checks fail.
6b. Verified Execution
$ iverilog -g2012 -o d i2c_bus_recovery.sv i2c_bus_recovery_tb.sv && ./d
=== i2c_bus_recovery: two stuck lines, two different escape hatches ===
T1 SDA stuck, released after three pulses
T2 released on the ninth pulse: still a success
T3 never released: nine pulses then escalation
T4 SCL stuck: no clocking is attempted at all
T5 both lines stuck: diagnosed as both, still no clocking
T6 without a HW reset input, cycle power
T7 with neither remedy available, say so
T8 recovery on a healthy bus drives nothing
T9 SCL failing mid-recovery abandons the clocking
T10 a successful recovery ends with a STOP
T11 two recoveries, counters accumulate, nothing leaks
T12 the standing property held throughout
=== i2c_bus_recovery: ALL CHECKS PASSED ===
i2c_bus_recovery_tb.sv:336: $finish called at 2090000 (1ps)
$ iverilog -g2005 -o v i2c_bus_recovery.v i2c_bus_recovery_tb.v && ./v
=== i2c_bus_recovery: two stuck lines, two different escape hatches ===
T1 SDA stuck, released after three pulses
T2 released on the ninth pulse: still a success
T3 never released: nine pulses then escalation
T4 SCL stuck: no clocking is attempted at all
T5 both lines stuck: diagnosed as both, still no clocking
T6 without a HW reset input, cycle power
T7 with neither remedy available, say so
T8 recovery on a healthy bus drives nothing
T9 SCL failing mid-recovery abandons the clocking
T10 a successful recovery ends with a STOP
T11 two recoveries, counters accumulate, nothing leaks
T12 the standing property held throughout
=== i2c_bus_recovery: ALL CHECKS PASSED ===
i2c_bus_recovery_tb.v:337: $finish called at 2090000 (1ps)
$ nvc --std=2008 -a i2c_bus_recovery.vhd i2c_bus_recovery_tb.vhd
$ nvc --std=2008 -e i2c_bus_recovery_tb && nvc --std=2008 -r i2c_bus_recovery_tb --stop-time=500us
** Note: 0ms+0: === i2c_bus_recovery: two stuck lines, two different escape hatches ===
Process :i2c_bus_recovery_tb:stim at i2c_bus_recovery_tb.vhd:126
** Note: 190ns+1: T1 SDA stuck, released after three pulses
Process :i2c_bus_recovery_tb:stim at i2c_bus_recovery_tb.vhd:126
** Note: 560ns+1: T2 released on the ninth pulse: still a success
Process :i2c_bus_recovery_tb:stim at i2c_bus_recovery_tb.vhd:126
** Note: 920ns+1: T3 never released: nine pulses then escalation
Process :i2c_bus_recovery_tb:stim at i2c_bus_recovery_tb.vhd:126
** Note: 1010ns+1: T4 SCL stuck: no clocking is attempted at all
Process :i2c_bus_recovery_tb:stim at i2c_bus_recovery_tb.vhd:126
** Note: 1100ns+1: T5 both lines stuck: diagnosed as both, still no clocking
Process :i2c_bus_recovery_tb:stim at i2c_bus_recovery_tb.vhd:126
** Note: 1200ns+1: T6 without a HW reset input, cycle power
Process :i2c_bus_recovery_tb:stim at i2c_bus_recovery_tb.vhd:126
** Note: 1300ns+1: T7 with neither remedy available, say so
Process :i2c_bus_recovery_tb:stim at i2c_bus_recovery_tb.vhd:126
** Note: 1390ns+1: T8 recovery on a healthy bus drives nothing
Process :i2c_bus_recovery_tb:stim at i2c_bus_recovery_tb.vhd:126
** Note: 1570ns+1: T9 SCL failing mid-recovery abandons the clocking
Process :i2c_bus_recovery_tb:stim at i2c_bus_recovery_tb.vhd:126
** Note: 1730ns+1: T10 a successful recovery ends with a STOP
Process :i2c_bus_recovery_tb:stim at i2c_bus_recovery_tb.vhd:126
** Note: 2090ns+1: T11 two recoveries, counters accumulate, nothing leaks
Process :i2c_bus_recovery_tb:stim at i2c_bus_recovery_tb.vhd:126
** Note: 2090ns+1: T12 the standing property held throughout
Process :i2c_bus_recovery_tb:stim at i2c_bus_recovery_tb.vhd:126
** Note: 2090ns+1: === i2c_bus_recovery: ALL CHECKS PASSED ===
Process :i2c_bus_recovery_tb:stim at i2c_bus_recovery_tb.vhd:126All three at 2090 ns. The VHDL bench differs from the two Verilog ones in one structural way worth noting: VHDL permits only a single driver for a signal of an unresolved type, so the modelled holder owns holder_holds_sda outright and the stimulus requests a hold through a separate signal. Getting that wrong was this module's most instructive bug — std_logic is resolved, so two drivers elaborated silently and resolved to 'X', which made every stuck-line test read as a healthy bus. The Verilog benches were then restructured to match, which is why all three now agree on the cycle.
6c. What The Testbench Proves
| # | scenario | what it establishes |
|---|---|---|
| 1 | SDA stuck, released after three pulses | stops at three — nine is a bound |
| 2 | released on the ninth pulse | still a success, using all nine |
| 3 | never released | exactly nine, then escalation |
| 4 | SCL stuck | zero pulses; straight to HW reset |
| 5 | both lines stuck | diagnosed as both; still zero pulses |
| 6 | no HW reset available | the power cycle rung |
| 7 | neither rung available | impossible, and says so |
| 8 | a healthy bus | drives nothing; reports healthy |
| 9 | SCL fails mid-recovery | re-diagnosed; clocking abandoned |
| 10 | a successful recovery | ends with a STOP; bus left idle |
| 11 | two recoveries in a row | counters accumulate; nothing leaks |
| 12 | the standing property, over the whole run | attempted_clocks_on_scl is still zero |
Tests 1, 2 and 3 are the same procedure with the release at three different points, and together they pin down that nine is an upper bound: stop early when you can, use all nine when you must, and never exceed it.
Test 4 is the chapter's central claim and its assertion is a zero. A stuck clock must produce no pulses at all. Asserting the count rather than only the final remedy is what makes it a test of the procedure rather than of the outcome — a design that issued nine futile pulses and then escalated would reach the same remedy by the wrong route.
Test 9 is the reason the second guard exists. The recovery begins legitimately on a stuck SDA, and after two pulses the clock is seized as well. A design that only checked the clock at diagnosis time carries on pulsing a dead line for seven more pulses. Mutation Z3 is that design.
Test 12 asserts the standing property one final time, after everything else has run. It is not redundant with tests 4, 5 and 9: those check it at the end of their own scenarios, and test 12 checks that nothing in the whole run — including the healthy-bus case and the two-recovery case — ever tripped it.
7. Mutation Testing
Eleven defects injected into the SystemVerilog sequencer.
| # | injected defect | outcome |
|---|---|---|
| Z1 | the in-flight guard that refuses to clock a held SCL removed | killed — test 9 |
| Z2 | a bus with both lines stuck sent to the clocking state | killed — test 5 |
| Z3 | clocking continued when SCL fails mid-recovery | killed — test 9 |
| Z4 | ten pulses issued instead of nine | killed — test 3 |
| Z5 | always spend nine rather than stopping on release | killed — 3 checks |
| Z6 | SDA sampled at the instant SCL is released | killed — 16 checks |
| Z7 | a power cycle preferred over the HW reset | killed — 7 checks |
| Z8 | success claimed when neither remedy exists | killed — 3 checks |
| Z9 | the closing STOP skipped | killed — 4 checks |
| Z10 | the procedure run on a healthy bus | killed — 2 checks |
| Z11 | SDA released before SCL is high in the closing STOP | killed — test 10 |
Eleven of eleven, and one mutation was proved equivalent and kept out of the count.
The two guards are not symmetric, and a mutation showed it. A twelfth mutation sent a stuck-SCL diagnosis straight into the clocking state — breaking the guard in S_DIAG — and survived. The reason is that S_CLOCK's first act is to re-test the clock, so the in-flight guard catches it before any pulse is driven and the observable outcome is unchanged: zero pulses, a stuck-SCL diagnosis, an escalated remedy.
The in-flight guard subsumes the diagnosis-time guard for an initially held clock, and additionally covers a clock that fails mid-recovery. The reverse is not true — which is why removing the in-flight guard alone (Z1) is killed by test 9, and removing the diagnosis-time guard alone is not killed by anything.
So the diagnosis-time guard is redundant for correctness and it is kept deliberately: it puts the refusal at the point where the remedy is chosen, which is where someone reading the code looks for it. The equivalence and its proof are recorded in the suite, and the mutation was replaced by Z1, which breaks the guard that does the work.
Z6's sixteen checks are the sampling point. Sampling SDA at the instant of release shifts every pulse count by one, so tests 1, 2, 3 and 11 all fail on their counts, plus the knock-on assertions. That breadth is the signature of an off-by-one in a measurement convention rather than a local logic error — the same shape as Chapter 14.1's S6.
Z10 kills with only two checks, and it is the most dangerous mutation in the table. Running the procedure on a healthy bus produces a successful recovery — the line is already free, so the first pulse finds it released — and the only evidence is the pulse count and the remedy. A suite that checked "did recovery succeed" would pass this mutation while the design corrupted a transfer that was merely slow.
8. Verification Connection — Asserting That Something Never Happens
This design's central property is a prohibition, and prohibitions are verified differently from behaviours: the assertion is cheap and the coverage is where the work is.
// The property: the recovery block must never drive SCL while SCL is being held
// low by another device. It is a SAFETY property -- something that must never
// happen -- so the assertion is trivial and proves nothing on its own. A bus on
// which the clock is never stuck satisfies it vacuously, forever.
interface i2c_recovery_if (input logic clk);
logic scl_in, sda_in; // the lines as they actually read
logic scl_drive_low; // the recovery block's SCL output
logic sda_drive_low;
logic [2:0] diagnosis;
logic [3:0] pulses_issued;
logic recovered, gave_up;
endinterface
module i2c_recovery_checker (i2c_recovery_if bus);
localparam D_SCL_STUCK = 3'd2, D_BOTH = 3'd3;
// 1. THE PROHIBITION. Never drive the clock while somebody else holds it.
// Expressed on the OBSERVED line rather than on the diagnosis, so a wrong
// diagnosis cannot make the assertion vacuous.
a_never_clock_held_scl : assert property (
@(posedge bus.clk) !(bus.scl_drive_low && $past(!bus.scl_in && !bus.scl_drive_low)))
else $error("drove SCL while another device was holding it low");
// 2. Nine is a bound. Not eight, not ten.
a_pulse_bound : assert property (
@(posedge bus.clk) bus.pulses_issued <= 4'd9)
else $error("issued more than nine clock pulses");
// 3. A stuck clock must produce NO pulses. This is the asymmetry of §1, and
// it is the assertion that distinguishes a correct design from one that
// reaches the right remedy by attempting the wrong procedure first.
a_no_pulses_on_stuck_scl : assert property (
@(posedge bus.clk)
(bus.diagnosis == D_SCL_STUCK || bus.diagnosis == D_BOTH) |-> bus.pulses_issued == 4'd0)
else $error("pulses were issued despite a stuck-clock diagnosis");
// 4. Never both outcomes at once.
a_exclusive : assert property (
@(posedge bus.clk) !(bus.recovered && bus.gave_up))
else $error("reported both recovery and giving up");
// ---- COVERAGE: where the actual work is --------------------------------
// Every assertion above passes on a bus that never gets stuck. These bins
// are the evidence that the scenarios existed at all.
covergroup cg_recovery @(posedge bus.clk);
option.per_instance = 1;
// Which fault was actually injected. All four must be reached, and the
// stuck-clock bins are the ones a happy-path regression never builds.
cp_fault : coverpoint bus.diagnosis {
bins healthy = {3'd0};
bins sda_stuck = {3'd1};
bins scl_stuck = {3'd2}; // the case assertion 3 exists for
bins both = {3'd3};
}
// How long the release took. The boundary bins matter most: a release on
// the ninth pulse and a line that never releases are different outcomes
// one pulse apart.
cp_pulses : coverpoint bus.pulses_issued {
bins none = {0};
bins early = {[1:3]};
bins late = {[4:8]};
bins last_chance = {9}; // released exactly at the bound, or not at all
}
// And the combination that proves the asymmetry was exercised rather than
// merely asserted: a stuck clock together with a pulse count of zero.
x_asymmetry : cross cp_fault, cp_pulses {
bins stuck_clock_no_pulses =
binsof(cp_fault.scl_stuck) && binsof(cp_pulses.none);
}
endgroup
cg_recovery cg = new();
endmodule9. FPGA and ASIC Implications
Recovery needs to read the lines, which means an input path on SCL. A master that only ever drives its clock has no way to know it is stuck. Reading SCL back is also what clock synchronization (Chapter 13.2) and stretch detection (Chapter 12.2) require, so on any design that supports those the path already exists — but on a minimal master it is an addition, and without it §3.1.16 cannot be implemented at all.
The nine pulses must be generated without the protocol engine. The state machine is, by hypothesis, wedged — that is why recovery is running. So the pulse generator has to be reachable from a reset or supervisory path, not from the transfer FSM, and that is a structural requirement rather than a coding preference.
A HW reset input is cheap and is the rung that matters. §3: it needs no cooperation from the broken device, and it is "preferential" in the specification's own word. A device without one forces every recovery to the power-cycle rung, which on a system where the I²C bus manages the power sequencing may not be available at all.
Watch the ordering trap in a power-managed system. If the devices on the bus are powered through a regulator that is itself configured over that bus, a power cycle is not obviously possible — and the design needs either a HW reset or an out-of-band path. That is worth checking at architecture time rather than discovering during a field failure.
A conforming device floats its pins when unpowered. §5, and §5.1 states it as a requirement. A design whose I/O pulls low with the supply off will hold the bus down for every other device on it and cannot be recovered by any protocol means.
Report the pulse count, and log it. pulses_issued distinguishes a bus that occasionally needs one pulse from one that routinely needs eight. The second is a system that is failing regularly and being rescued regularly, which is a different problem from an isolated event and is invisible in a pass/fail flag.
10. Debugging — The Recovery That Always Failed
A product implements bus recovery: on any transfer timeout it issues nine clock pulses and retries. The recovery works reliably in the lab. In the field a small number of units enter a state where the recovery runs, reports failure, runs again on the next timeout, and never succeeds — indefinitely, until power is cycled by the user.
Two separate faults, and the recovery code could not see either. The sensor violates §5.1 by pulling the bus low when its rail droops, which produces a stuck SCL -- and §3.1.16 provides no protocol remedy for that line, because the nine pulses are themselves pulses on SCL and cannot be issued on a line something else is holding. The recovery code never read SCL, so it diagnosed a stuck SDA, attempted the one procedure that cannot work on a held clock, and correctly observed no change. It then reported failure, which was accurate and useless: the failure was not that nine pulses were insufficient but that zero pulses were issued.
Read both lines and branch on which is stuck, as §3.1.16 requires. For a stuck SCL escalate immediately rather than clocking: assert a HW reset if one exists, otherwise request a power cycle. Report the diagnosis, not just the outcome, so a field log distinguishes 'nine pulses did not free SDA' from 'SCL was held and no protocol remedy exists'. Separately, fix the rail: the sensor's brownout behaviour violates §5.1 and is the actual root cause.Three things generalise.
The recovery code was not wrong about SDA; it was blind to SCL. It observed a true fact — SDA was low — and drew a conclusion the specification does not support, because it never asked the question that decides which remedy applies.
"Recovery failed" was accurate and told nobody anything. The useful report is the diagnosis: SCL held, no protocol remedy, escalation required. That is the difference between a log line an engineer can act on and one that prompts a power cycle.
The real bug was a device violating §5.1, two layers away from the symptom, and it produced a fault class the recovery mechanism is structurally unable to address. Recovery bought nothing here, which is the honest limit of §2's table.
11. Common Misconceptions
"Nine clock pulses clear a stuck bus." They clear a stuck SDA. For a stuck SCL the specification offers no protocol remedy at all. §1.
"The specification forgot to cover a stuck SCL." It covers it — with a HW reset or a power cycle. There is no protocol remedy because the nine pulses are pulses on SCL and cannot be issued on a held line. §1.
"Nine pulses must always be sent." The device "should release it some time within those nine clocks", so the procedure stops as soon as the line is free. Nine is a bound. §1.
"Why nine and not eight?" A byte on this bus is nine bits — eight data plus the acknowledge — so nine pulses carry a wedged device through any position in its frame. §1.
"A power cycle is a last resort that might not work." Every I²C device has a mandatory internal POR, so a power cycle always works. It is the last resort because it is the most disruptive, not because it is the least reliable. §3.
"Bus clear fixes a hung bus." It fixes one cause: a device holding SDA. Shorts, unpowered devices, address conflicts and confused masters are all untouched. §2.
"Recovery does not need to read SCL." Without reading SCL there is no way to know which remedy applies, and the wrong one is attempted. §10.
"A recovery that ends with the line released is complete." A released line with no framing leaves every device mid-transfer. The procedure ends with a STOP. §4.
"An unpowered device is harmless on the bus." Only if it floats its pins, which §5.1 requires and not every device does. §5.
"If recovery reports success the bus is healthy." A recovery run on a bus that was merely slow reports success having corrupted a transfer. Mutation Z10, §7.
12. Reason It Through
A master finds SDA low and SCL low. What should it do?
Escalate immediately. A stuck SCL has no protocol remedy, and the nine pulses cannot be issued on a line something else is holding. Assert a HW reset if one exists, otherwise request a power cycle — and do not issue clock pulses, because they will not appear on the wire. §1 and §10.
Why is nine the number, and why is it a bound rather than a quota?
Nine is the length of a byte including the acknowledge, so it carries a device stuck anywhere in a frame through to the end of it. It is a bound because §3.1.16 says the release comes "some time within" those nine clocks — so stopping early is correct, and reporting how many were needed is more informative than always spending nine. §1.
A recovery block issues nine pulses on a bus where SCL is held, observes no change, and reports failure. What is wrong with that report?
It is accurate and useless. Zero pulses actually reached the wire, so the report "nine pulses did not free the line" describes something that did not happen. The correct report is the diagnosis: SCL is held, and no protocol remedy exists for that line. §10.
Why must SDA be sampled while SCL is high rather than at the moment SCL is released?
Because the holder releases during the low phase — the data-valid rule requires it — and a master sampling at the release instant reads the pre-change value. It then needs one more pulse to notice, and reports a pulse count one higher than the truth. §4 and mutation Z6.
Why does a successful recovery end with a STOP?
Because a released SDA with no framing leaves every device believing a transfer is in progress. A STOP is the only thing that returns them to idle, so the bus ends the procedure idle rather than merely unstuck. §4.
A general call 06h resets forty devices and the bus goes dead. What happened, and what does this chapter say about it?
One or more devices came out of reset still pulling SDA or SCL low — the hazard Chapter 15.1 §3a warns about, and the reason it is warned about in the general call section is that the broadcast multiplies it by forty. If a clock is among the held lines there is no protocol remedy, so the recovery mechanism of the previous chapter created a fault this chapter cannot fix. §5.
Why is fault injection mandatory for a recovery block rather than merely desirable?
Because every assertion about a recovery procedure is a safety property that passes vacuously on a bus that never fails. The fault the block handles is the one the rest of the design exists to prevent, so it will not occur in ordinary stimulus — and a green regression with an empty stuck-clock cover bin has tested nothing. §8.
13. Understanding Check
14. Summary
§3.1.16 is four sentences and says two different things about two different wires. Nine clock pulses for a stuck SDA; nothing at the protocol level for a stuck SCL.
The asymmetry is structural. The nine pulses are pulses on SCL, so the remedy that works for the line the master shares cannot be applied to the line the master owns.
Nine is the length of a byte, acknowledge included, and it is a bound rather than a quota — the release comes "some time within" those nine clocks, so stopping early is correct and the count is worth reporting.
Both ladders end the same way. HW reset is "preferential" because it needs no cooperation from the broken device; a power cycle invokes a mandatory POR and therefore always works. Each rung down needs less cooperation and disrupts more.
And a board may have neither rung, in which case the honest output is that the bus cannot be recovered.
Bus clear repairs exactly one fault: a device holding SDA. Shorts, unpowered devices, address conflicts and confused masters are all untouched, so diagnose before acting and report the diagnosis rather than only the outcome.
Two of the likeliest causes of a stuck line are a device that is off and a device that has just been reset — and §5.1 requires an unpowered device to float its pins, which not every device honours.
Sample SDA while SCL is high. The holder releases during the low phase, so sampling at the release instant costs a pulse and misreports the count.
End with a STOP, or the bus is unstuck and not idle.
And a recovery block's assertions all pass on a working bus. Fault injection is not optional here — the fault it handles is the one the rest of the design exists to prevent, so a green regression with an empty stuck-clock cover bin has verified nothing.
15. What Comes Next
Module 15 is complete, and it has been about the parts of the specification that sit outside an ordinary transfer.
The through-line turned out to be what a master can and cannot know. A general call reaches every device and tells the master only that somebody was listening, because the wired-AND flattens a broadcast's acknowledge into a single OR. A Device ID confirms an identity you already expected and cannot enumerate a bus, because an unimplemented optional feature is indistinguishable from an absent device. An SMBus limit can be violated by a perfectly conforming I²C part, because the protocols are identical and only the numbers differ. And a stuck clock admits no protocol remedy at all, because the one instrument the master has is the thing being held.
Four chapters, four different ways the bus declines to answer a question — and in three of them the correct engineering response was to get the information from somewhere else entirely: a per-device read, a datasheet, a scope on a device pin.
Module 16 turns from the bus to the devices on it. Every chapter so far has described a generic target — one that acknowledges, holds registers, and can be read and written. Real devices are not generic. They have register maps with auto-increment behaviour the designer chose, paged addressing that silently wraps, write-protect pins, internal write cycles during which they will not respond at all, and power-on defaults that may or may not match the datasheet's front page.
Which is where the acknowledge stops being a protocol mechanism and becomes a device mechanism. A target that NACKs because it is busy, one that NACKs because the address is wrong, and one that NACKs because a write-protect pin is asserted are indistinguishable on the wire — and each needs a different response from the master.
Continue learning
Related tutorials
- Related topic
Stuck Bus — Diagnosis and Recovery
A bus held low, who is holding it, and why the standard recovery works for exactly one of its seven causes. Builds a clear sequencer whose pulse count is evidence rather than a boolean, runs it against seven holders including a legally stretching target, and reports the four bugs it took to get there — two in the design and two in the bench.
- Related topic
I²C Stretch Bounds, Timeouts and the No-Assumption Rule
The specification places no limit on how long a slave may hold the clock, so every timeout is a system policy rather than a compliance check. Includes the recovery asymmetry most engineers know only half of.
- Related topic
Timeout, Error Handling and Bus Recovery
UM10204 defines no error codes, so the taxonomy is a design decision — and collapsing two failures into one code makes them indistinguishable to software that could have acted on them differently. Builds six codes, separates protocol outcomes from faults, and implements the nine-pulse recovery including an honest account of what it cannot do.
- Related topic
Real Bus Electrical Behavior — Edges, Levels and Level Shifting
Take the idealised node onto a real board. Edges acquire shape and become asymmetric, levels turn out to be thresholds with a region between them, capacitance comes from identifiable places a designer controls, and two devices on one bus may not share a supply — which makes a bidirectional level shifter structural.
