I²C · Module 12
I²C Stretch Bounds, Timeouts and the No-Assumption Rule
The specification places no limit on how long a slave may hold the clock, so every timeout is a system policy rather than a compliance check. Includes the recovery asymmetry most engineers know only half of.
Three chapters have treated a stretch as something that ends. This one takes the case where it does not, and the specification's position turns out to be more uncomfortable than most designs assume.
There is no limit. Not a generous one, not an implicit one. The I²C specification declines to bound how long a slave may hold the clock, and says so explicitly.
Which means every timeout in every I²C driver ever written is a policy — a statement about what this product can tolerate — and not a compliance check. A design that reports one as a protocol violation is claiming a rule the specification refuses to state.
1. What the Specification Actually Says
That single paragraph contains the whole chapter. Set the two buses side by side:
| I²C | SMBus | |
|---|---|---|
| maximum stretch | none | 35 ms |
| minimum clock frequency | none — a "DC bus" | 10 kHz |
| a very long stretch is | legal | a fault: reset everything |
| who decides a bound | the system designer | the specification |
"I²C can be a 'DC' bus" is the phrase to remember. DC — direct current — means the clock may stop indefinitely and the bus is still functioning correctly. Not degraded, not marginal: correct. And the reason the two specifications differ here is not oversight. SMBus's 10 kHz floor exists because it chose to bound the stretch: once you assert that nothing may take longer than 35 ms, you have implied a minimum rate, and you have to state it. I²C declined the bound and therefore has no floor.
So the two design positions are coherent and incompatible. SMBus buys recoverability and pays with a rate requirement no software slave can meet. I²C buys arbitrary slaves — including a microcontroller servicing the bus in an interrupt handler (Chapter 12.3 §11) — and pays by having no protocol-level answer when a slave never lets go.
2. A Timeout Is a Policy, and the Naming Matters
Every other checker in Module 11 and this module compares a measurement against a limit the specification states. This one cannot, and the consequence runs all the way into the port list.
3. The Recovery Asymmetry
Now the part most engineers know only half of. §3.1.16 gives two remedies for two stuck lines, and they are not equivalent.
| stuck line | remedy | is it a protocol action? |
|---|---|---|
| SDA | send nine clock pulses | yes — the master does it on the bus |
| SCL | hardware reset, or cycle power | no — out of band entirely |
The reason for the asymmetry is almost tautological once stated: clocking your way out requires the clock.
The nine-clock recovery works on SDA because the master still owns SCL. It drives nine pulses, and whichever device was mid-byte finishes its byte, sees a ninth bit it can interpret as a NACK, and releases SDA. Nine is the right number because a byte is nine bits — no device can be more than eight bits into a transfer, so nine pulses guarantees every device reaches a byte boundary.
On SCL there is nothing to try. The master cannot generate pulses on a line another device is holding down; that is precisely what the wired-AND guarantees (Chapter 12.2 §2 — a pull-down is a command). There is no sequence of bus activity that helps, which is why the specification goes straight to hardware reset and then to power cycling.
An unbounded stretch and a stuck SCL are the same electrical condition. The only difference is whether the slave intends to release it — and the bus cannot tell you that.
That is the uncomfortable fact at the centre of this chapter. A master waiting on a legal 3 ms EEPROM write and a master waiting on a wedged slave are in identical states. No measurement distinguishes them, because §1 says there is no duration that is too long. Only a policy can separate them, and a policy is a guess informed by a datasheet.
4. The Escalation, Drawn
Over policy is still legal; stuck is a decision, not a measurement
10 cyclesThe middle phase is the chapter in one label: over budget, still legal. Nothing has been violated. A product has decided it cannot wait this long, which is a different statement — and the figure keeps them visually distinct so that a reader does not collapse them.
5. The No-Assumption Rule
Chapter 12.1 §4 quoted Table 2's footnote [2]:
"Clock stretching is a feature of some slaves. If no slaves in a system can stretch the clock (hold SCL LOW), the master need not be designed to handle this procedure."
Read as permission, that sentence is a trap. Read as a contract, it is the rule this module closes on:
Never assume a target will not stretch unless the device or system contract guarantees it.
Three reasons the assumption is more dangerous than it looks, each one established earlier in the module.
Its failure mode is electrical, not logical. A master that assumed no stretching may have a push-pull SCL (Chapter 12.1 §4, §3.1.1), and a push-pull driver meeting a stretching slave is contention — every device on the bus loses its clock simultaneously, which is §11 of that chapter.
The assumption is invalidated by a board revision, not by a code change. Nothing in the RTL breaks when a stretching slave is added. The contract was made at schematic time and violated at schematic time, years apart, by people who had no way to see it.
And FPGA slaves stretch by default. Stretching costs a fixed-function part an SCL pad driver, which is why most do not have it — but in an FPGA the driver is free, so the obvious slave implementation uses it (Chapter 12.1 §10). The single most likely thing to appear on a mature bus is an FPGA, and it is the single most likely thing to stretch.
The cheap insurance is to tolerate stretching even where nothing stretches. Chapter 12.2's generator costs one extra state and sequences on the observed line; on a bus with no stretching devices it behaves identically. There is no runtime cost and no configuration bit, and it converts a load-bearing assumption into a non-issue.
6. The Stretch Guard in Three Languages
// STRETCH BOUNDS, AND WHY THEY ARE NOT PROTOCOL CHECKS. Every other checker in Modules 11 and 12
// compares a measurement against a limit the specification states. This one cannot, because the
// specification states the opposite:
//
// UM10204 section 4.2.2: "I2C can be a 'DC' bus, meaning that a slave device stretches the
// master clock when performing some routine while the master is accessing it. This notifies the
// master that the slave is busy but does not want to lose the communication. The slave device
// will allow continuation after its task is complete. THERE IS NO LIMIT IN THE I2C-BUS PROTOCOL
// AS TO HOW LONG THIS DELAY CAN BE, whereas for a SMBus system, it would be limited to 35 ms."
//
// So a long stretch is not a violation of anything. A timeout on it is a SYSTEM POLICY -- a
// statement that this product cannot wait longer than X -- and conflating the two is the mistake
// this block is shaped to prevent. The output is called `warn_over_limit`, not `viol_*`, and
// T_LIMIT = 0 means NO TIMEOUT, which is the I2C default and must be the parameter's neutral
// value. That mirrors Chapter 11.8's T_SP = 0 for the same reason: a configuration the
// specification explicitly permits must be expressible, and must cost nothing.
//
// The second half of the block is the ESCALATION, and it exists because section 3.1.16 gives two
// different remedies for two stuck lines and most people know only one of them:
//
// "In the unlikely event where the clock (SCL) is stuck LOW, the preferential procedure is to
// reset the bus using the HW reset signal ... If the I2C devices do not have HW reset inputs,
// cycle power ... If the data line (SDA) is stuck LOW, the master should send nine clock
// pulses. The device that held the bus LOW should release it some time within those nine
// clocks."
//
// The asymmetry is the whole point, and it is not arbitrary:
//
// SDA stuck LOW -> send nine clock pulses. A protocol-level recovery EXISTS, because the thing
// you need in order to perform it (the clock) is still yours.
// SCL stuck LOW -> HW reset, or cycle power. There is NO protocol-level recovery, because
// clocking your way out requires the clock you have just lost.
//
// An unbounded stretch is an SCL-stuck-low condition. So the honest thing for a guard to report is
// not "violation" but "this needs a reset line or a power cycle, and no amount of bus activity
// will help" -- which is a very different message to send to a driver.
module i2c_stretch_guard #(
parameter int TICK_W = 24,
// The POLICY limit. 0 = no timeout, which is what plain I2C specifies. A system that needs a
// bound sets it; nothing in the protocol does.
parameter int T_LIMIT = 0,
// The ESCALATION threshold, beyond which the line is treated as stuck rather than slow.
// Also 0 = never escalate.
parameter int T_STUCK = 0
)(
input logic clk,
input logic rst_n,
// intent and observation, for both lines
input logic scl_release,
input logic sda_release,
input logic scl_in,
input logic sda_in,
// ---- the stretch, as Chapter 12.1 defines it ----
output logic stretching,
output logic [TICK_W-1:0] stretch_ticks,
output logic [TICK_W-1:0] max_stretch_seen,
// ---- POLICY, not protocol. Named so that nobody reads it as a compliance failure. ----
output logic warn_over_limit,
output logic [TICK_W-1:0] n_warn,
// ---- escalation, and the remedy section 3.1.16 prescribes for each line ----
output logic scl_stuck_low,
output logic sda_stuck_low,
output logic [1:0] recovery_action,
output logic [TICK_W-1:0] n_escalations
);
// Section 3.1.16's two remedies, as an enumeration so a driver cannot confuse them.
localparam logic [1:0] RECOV_NONE = 2'd0;
localparam logic [1:0] RECOV_NINE_CLOCKS = 2'd1; // SDA stuck: a protocol recovery exists
localparam logic [1:0] RECOV_HW_RESET = 2'd2; // SCL stuck: it does not
// A limit of zero means "no limit", and it must be impossible for any comparison to fire.
// Writing this as a wire rather than testing T_LIMIT at each use keeps the neutral case in
// one place; mutation D2 removes the guard and is killed by a 5000-tick stretch at T_LIMIT=0.
wire limit_enabled = (T_LIMIT != 0);
wire stuck_enabled = (T_STUCK != 0);
// Sized ONCE. A part-select of a parameter -- `LIMIT_W` -- is not portable:
// Icarus evaluates it as zero, which would make every comparison fire immediately. Design 3
// of this module lost time to exactly that, so both thresholds are widened here instead.
localparam logic [TICK_W-1:0] LIMIT_W = T_LIMIT;
localparam logic [TICK_W-1:0] STUCK_W = T_STUCK;
wire scl_held = scl_release && !scl_in; // somebody else is holding SCL low
wire sda_held = sda_release && !sda_in; // somebody else is holding SDA low
logic [TICK_W-1:0] ticks_q, sda_ticks_q;
wire [TICK_W-1:0] ticks_now = ticks_q + 1'b1;
wire [TICK_W-1:0] sda_ticks_now = sda_ticks_q + 1'b1;
// SCL dominance, decided COMBINATIONALLY on the condition rather than on the registered flag.
// Both lines can cross the escalation threshold in the SAME cycle, and in that cycle
// `scl_stuck_low` still holds its old value -- so a test of the flag reads 0, the SDA branch
// runs, and because it is written later in the same block its assignment to recovery_action
// wins. The result is a recommendation to send nine clock pulses on a clock line that is being
// held down, which is a loop a driver cannot exit.
//
// The general rule: a PRIORITY decision between two conditions must be made on the
// conditions, never on the flags that record them.
wire scl_stuck_c = stuck_enabled
&& (scl_stuck_low || (scl_held && stretching && (ticks_now > STUCK_W)));
always_ff @(posedge clk) begin
if (!rst_n) begin
stretching <= 1'b0;
ticks_q <= '0;
sda_ticks_q <= '0;
max_stretch_seen <= '0;
warn_over_limit <= 1'b0;
n_warn <= '0;
scl_stuck_low <= 1'b0;
sda_stuck_low <= 1'b0;
recovery_action <= RECOV_NONE;
n_escalations <= '0;
end else begin
// ---- SCL: the stretch ----
if (scl_held) begin
if (!stretching) begin
stretching <= 1'b1;
ticks_q <= {{(TICK_W-1){1'b0}}, 1'b1}; // "including this cycle"
end else begin
ticks_q <= ticks_now;
if (ticks_now > max_stretch_seen) max_stretch_seen <= ticks_now;
// POLICY. Raised once per stretch, on the cycle the limit is passed, and only
// if a limit was configured at all.
if (limit_enabled && !warn_over_limit && ticks_now > LIMIT_W) begin
warn_over_limit <= 1'b1;
n_warn <= n_warn + 1'b1;
end
// ESCALATION. Past this point the line is not slow, it is stuck, and section
// 3.1.16 says the remedy is out-of-band.
if (stuck_enabled && !scl_stuck_low && ticks_now > STUCK_W) begin
scl_stuck_low <= 1'b1;
recovery_action <= RECOV_HW_RESET;
n_escalations <= n_escalations + 1'b1;
end
end
end else begin
// Released. Everything about THIS stretch clears; the totals do not.
stretching <= 1'b0;
ticks_q <= '0;
warn_over_limit <= 1'b0;
if (scl_stuck_low) begin
scl_stuck_low <= 1'b0;
// The SCL remedy no longer applies. If SDA is still held, the surviving
// condition is the one with a protocol recovery.
recovery_action <= sda_stuck_low ? RECOV_NINE_CLOCKS : RECOV_NONE;
end
end
// ---- SDA: stuck low, which has its own and much better remedy ----
if (sda_held) begin
sda_ticks_q <= sda_ticks_now;
if (stuck_enabled && !sda_stuck_low && sda_ticks_now > STUCK_W) begin
sda_stuck_low <= 1'b1;
n_escalations <= n_escalations + 1'b1;
// SCL dominates if both are stuck: you cannot send nine clock pulses on a
// clock line somebody else is holding down. Reporting the nine-clock remedy
// while SCL is stuck would send a driver into a loop that cannot terminate.
if (!scl_stuck_c) recovery_action <= RECOV_NINE_CLOCKS;
end
end else begin
sda_ticks_q <= '0;
if (sda_stuck_low) begin
sda_stuck_low <= 1'b0;
recovery_action <= scl_stuck_low ? RECOV_HW_RESET : RECOV_NONE;
end
end
end
end
assign stretch_ticks = ticks_q;
endmodule `timescale 1ns/1ps
// 100 MHz sample clock. THREE instances, because the whole point of this block is that the
// interesting configuration is the DEGENERATE one:
//
// dut_none -- T_LIMIT = 0, T_STUCK = 0 : plain I2C. Section 4.2.2 puts no limit on a stretch,
// so this instance must NEVER warn, however long the stretch.
// dut_pol -- T_LIMIT = 500, T_STUCK = 0 : a system policy, no escalation.
// dut_full -- T_LIMIT = 500, T_STUCK = 2000 : policy plus section 3.1.16 escalation.
//
// All three see identical stimulus on the same wires in one simulation, which is the control
// structure Chapter 11.8 section 2 argues for: a control in a separate run can diverge in its
// stimulus without anyone noticing.
module i2c_stretch_guard_tb;
localparam int TICK_W = 24;
logic clk = 1'b0, rst_n = 1'b0;
always #5 clk = ~clk;
logic m_scl_release = 1'b1, m_sda_release = 1'b1;
logic s_scl_hold = 1'b0, s_sda_hold = 1'b0;
wire scl_line = m_scl_release && !s_scl_hold;
wire sda_line = m_sda_release && !s_sda_hold;
// ---- the no-timeout instance: plain I2C ----
logic n_stretching, n_warn_over;
logic [TICK_W-1:0] n_ticks, n_max, n_nwarn, n_nesc;
logic n_scl_stuck, n_sda_stuck;
logic [1:0] n_recov;
i2c_stretch_guard #(.TICK_W(TICK_W), .T_LIMIT(0), .T_STUCK(0)) dut_none (
.clk(clk), .rst_n(rst_n),
.scl_release(m_scl_release), .sda_release(m_sda_release),
.scl_in(scl_line), .sda_in(sda_line),
.stretching(n_stretching), .stretch_ticks(n_ticks), .max_stretch_seen(n_max),
.warn_over_limit(n_warn_over), .n_warn(n_nwarn),
.scl_stuck_low(n_scl_stuck), .sda_stuck_low(n_sda_stuck),
.recovery_action(n_recov), .n_escalations(n_nesc));
// ---- policy only ----
logic p_stretching, p_warn_over;
logic [TICK_W-1:0] p_ticks, p_max, p_nwarn, p_nesc;
logic p_scl_stuck, p_sda_stuck;
logic [1:0] p_recov;
i2c_stretch_guard #(.TICK_W(TICK_W), .T_LIMIT(500), .T_STUCK(0)) dut_pol (
.clk(clk), .rst_n(rst_n),
.scl_release(m_scl_release), .sda_release(m_sda_release),
.scl_in(scl_line), .sda_in(sda_line),
.stretching(p_stretching), .stretch_ticks(p_ticks), .max_stretch_seen(p_max),
.warn_over_limit(p_warn_over), .n_warn(p_nwarn),
.scl_stuck_low(p_scl_stuck), .sda_stuck_low(p_sda_stuck),
.recovery_action(p_recov), .n_escalations(p_nesc));
// ---- policy plus escalation ----
logic f_stretching, f_warn_over;
logic [TICK_W-1:0] f_ticks, f_max, f_nwarn, f_nesc;
logic f_scl_stuck, f_sda_stuck;
logic [1:0] f_recov;
i2c_stretch_guard #(.TICK_W(TICK_W), .T_LIMIT(500), .T_STUCK(2000)) dut_full (
.clk(clk), .rst_n(rst_n),
.scl_release(m_scl_release), .sda_release(m_sda_release),
.scl_in(scl_line), .sda_in(sda_line),
.stretching(f_stretching), .stretch_ticks(f_ticks), .max_stretch_seen(f_max),
.warn_over_limit(f_warn_over), .n_warn(f_nwarn),
.scl_stuck_low(f_scl_stuck), .sda_stuck_low(f_sda_stuck),
.recovery_action(f_recov), .n_escalations(f_nesc));
localparam logic [1:0] RECOV_NONE = 2'd0, RECOV_NINE = 2'd1, RECOV_HW = 2'd2;
int errors = 0;
task automatic tick(input int n); begin repeat (n) @(negedge clk); end endtask
// A slave holds SCL low for `len` ticks while the master has released it.
task automatic hold_scl(input int len);
begin
m_scl_release = 1'b0; tick(20);
s_scl_hold = 1'b1; m_scl_release = 1'b1;
tick(len);
s_scl_hold = 1'b0; tick(20);
end
endtask
task automatic hold_sda(input int len);
begin
s_sda_hold = 1'b1; m_sda_release = 1'b1;
tick(len);
s_sda_hold = 1'b0; tick(20);
end
endtask
initial begin
tick(4); rst_n = 1'b1; tick(4);
// ---- 1. reset ------------------------------------------------------------------------
if (n_max !== '0 || p_max !== '0 || f_max !== '0) begin
$display("FAIL: max_stretch_seen nonzero out of reset"); errors++; end
if (n_recov !== RECOV_NONE || f_recov !== RECOV_NONE) begin
$display("FAIL: a recovery action was recommended out of reset"); errors++; end
if (n_warn_over !== 1'b0 || p_warn_over !== 1'b0) begin
$display("FAIL: warn asserted out of reset"); errors++; end
// ---- 2. a short stretch: nobody warns ------------------------------------------------
hold_scl(100);
if (n_nwarn !== '0 || p_nwarn !== '0 || f_nwarn !== '0) begin
$display("FAIL: a 100-tick stretch warned (none=%0d pol=%0d full=%0d)",
n_nwarn, p_nwarn, f_nwarn); errors++; end
if (n_max != 100 || p_max != 100) begin
$display("FAIL: a 100-tick stretch measured none=%0d pol=%0d", n_max, p_max);
errors++; end
// ---- 3. THE test of this chapter: a 5000-tick stretch with NO timeout configured ------
// Section 4.2.2: "There is no limit in the I2C-bus protocol as to how long this delay can
// be." The T_LIMIT = 0 instance must stay silent. If it warns, the block has invented a
// protocol rule the specification explicitly declines to state.
hold_scl(5000);
if (n_nwarn !== '0) begin
$display("FAIL: with T_LIMIT = 0 a 5000-tick stretch produced %0d warning(s) -- the block invented a limit the protocol does not have",
n_nwarn); errors++; end
if (n_scl_stuck !== 1'b0 || n_recov !== RECOV_NONE) begin
$display("FAIL: with T_STUCK = 0 a 5000-tick stretch escalated"); errors++; end
if (n_max != 5000) begin
$display("FAIL: the no-timeout instance measured %0d, expected 5000", n_max);
errors++; end
// ---- 4. the policy instance DOES warn on the same stimulus ----------------------------
// The positive control for test 3. Without it, a block that never warned at all would
// pass test 3 perfectly.
if (p_nwarn !== 1) begin
$display("FAIL: with T_LIMIT = 500 a 5000-tick stretch produced %0d warnings, expected 1",
p_nwarn); errors++; end
// ---- 5. escalation, and the remedy section 3.1.16 prescribes for SCL ------------------
// 5000 > T_STUCK of 2000, so the full instance must escalate -- and the recommendation
// must be HW RESET, because you cannot clock your way out of a held clock line.
if (f_nesc !== 1) begin
$display("FAIL: the escalating instance recorded %0d escalations, expected 1", f_nesc);
errors++; end
// ---- 6. the warning and the escalation both CLEAR when the line is released ----------
// A stretch that ended is over. A guard that latched forever would report a stuck bus for
// the rest of time after one slow device.
if (p_warn_over !== 1'b0) begin
$display("FAIL: warn_over_limit still set after the line was released"); errors++; end
if (f_scl_stuck !== 1'b0) begin
$display("FAIL: scl_stuck_low still set after the line was released"); errors++; end
if (f_recov !== RECOV_NONE) begin
$display("FAIL: a recovery action is still recommended after recovery (%0d)", f_recov);
errors++; end
// ---- 7. exactly AT the limit is not over it ------------------------------------------
// The comparison is strictly greater-than, so a stretch of exactly T_LIMIT is within
// policy. One tick more is not. Both halves are needed: a block that warned at the limit
// passes the second check and fails the first.
begin
int w0; w0 = p_nwarn;
hold_scl(500);
if (p_nwarn != w0) begin
$display("FAIL: a stretch of EXACTLY the 500-tick limit warned"); errors++; end
hold_scl(501);
if (p_nwarn != w0 + 1) begin
$display("FAIL: a stretch of limit+1 did not warn"); errors++; end
end
// ---- 8. SDA stuck low: the OTHER remedy ---------------------------------------------
// Section 3.1.16: "If the data line (SDA) is stuck LOW, the master should send nine clock
// pulses." A protocol recovery exists here, and reporting the SCL remedy instead would
// send a driver to reset hardware it does not need to touch.
begin
int e0; e0 = f_nesc;
// Hold SDA past the threshold and check the remedy WHILE it is still held. Checking
// only after release misses the whole point: the remedy is what a driver acts on, and
// it is only meaningful during the fault. Mutation M5 survived the first suite for
// exactly this reason -- the escalation count advanced either way.
s_sda_hold = 1'b1; m_sda_release = 1'b1;
tick(2100);
if (f_sda_stuck !== 1'b1) begin
$display("FAIL: SDA held past the threshold did not set sda_stuck_low"); errors++; end
if (f_recov !== RECOV_NINE) begin
$display("FAIL: SDA stuck alone recommended remedy %0d, expected nine clock pulses (%0d) -- section 3.1.16 gives SDA a protocol recovery that SCL does not have",
f_recov, RECOV_NINE); errors++; end
s_sda_hold = 1'b0; tick(20);
if (f_sda_stuck !== 1'b0) begin
$display("FAIL: sda_stuck_low still set after SDA was released"); errors++; end
if (f_recov !== RECOV_NONE) begin
$display("FAIL: a remedy is still recommended after SDA was released"); errors++; end
if (f_nesc != e0 + 1) begin
$display("FAIL: an SDA stuck-low condition did not escalate"); errors++; end
end
// ---- 9. both lines stuck: SCL dominates ---------------------------------------------
// You cannot send nine clock pulses on a clock line somebody else is holding down, so a
// guard that recommended the nine-clock recovery here would put a driver into a loop that
// cannot terminate.
begin
s_scl_hold = 1'b1; s_sda_hold = 1'b1;
m_scl_release = 1'b1; m_sda_release = 1'b1;
tick(2500);
if (f_recov !== RECOV_HW) begin
$display("FAIL: with BOTH lines stuck the remedy was %0d, expected HW reset (%0d) -- nine clock pulses are impossible without the clock",
f_recov, RECOV_HW); errors++; end
if (f_scl_stuck !== 1'b1 || f_sda_stuck !== 1'b1) begin
$display("FAIL: both lines stuck but scl=%b sda=%b", f_scl_stuck, f_sda_stuck);
errors++; end
// Release SCL only. SDA is still stuck, so the remedy must fall back to nine clocks.
s_scl_hold = 1'b0; tick(40);
if (f_recov !== RECOV_NINE) begin
$display("FAIL: with SCL freed and SDA still stuck the remedy was %0d, expected nine clocks (%0d)",
f_recov, RECOV_NINE); errors++; end
s_sda_hold = 1'b0; tick(40);
if (f_recov !== RECOV_NONE) begin
$display("FAIL: both lines free but a remedy is still recommended (%0d)", f_recov);
errors++; end
end
// ---- 10. this device driving SCL low is never a stretch ------------------------------
begin
int w0; w0 = p_nwarn;
m_scl_release = 1'b0;
tick(3000);
m_scl_release = 1'b1;
tick(20);
if (p_nwarn != w0) begin
$display("FAIL: this device holding SCL low for 3000 ticks warned"); errors++; end
end
// ---- 11. the worst case is a MAXIMUM and survives recovery ---------------------------
if (p_max != 5000) begin
$display("FAIL: max_stretch_seen = %0d, expected the 5000-tick peak to persist", p_max);
errors++; end
if (errors == 0)
$display("PASS: T_LIMIT=0 never warns because the protocol states no limit, a policy limit is not a protocol violation, exactly at the limit is within it, and section 3.1.16's two remedies are kept apart with SCL dominating");
else
$display("FAIL: %0d error(s)", errors);
$finish;
end
initial begin
#6000000;
$display("FAIL: watchdog expired");
$finish;
end
endmodule // STRETCH BOUNDS, AND WHY THEY ARE NOT PROTOCOL CHECKS. Every other checker in Modules 11 and 12
// compares a measurement against a limit the specification states. This one cannot, because the
// specification states the opposite:
//
// UM10204 section 4.2.2: "I2C can be a 'DC' bus, meaning that a slave device stretches the
// master clock when performing some routine while the master is accessing it. This notifies the
// master that the slave is busy but does not want to lose the communication. The slave device
// will allow continuation after its task is complete. THERE IS NO LIMIT IN THE I2C-BUS PROTOCOL
// AS TO HOW LONG THIS DELAY CAN BE, whereas for a SMBus system, it would be limited to 35 ms."
//
// So a long stretch is not a violation of anything. A timeout on it is a SYSTEM POLICY -- a
// statement that this product cannot wait longer than X -- and conflating the two is the mistake
// this block is shaped to prevent. The output is called `warn_over_limit`, not `viol_*`, and
// T_LIMIT = 0 means NO TIMEOUT, which is the I2C default and must be the parameter's neutral
// value. That mirrors Chapter 11.8's T_SP = 0 for the same reason: a configuration the
// specification explicitly permits must be expressible, and must cost nothing.
//
// The second half of the block is the ESCALATION, and it exists because section 3.1.16 gives two
// different remedies for two stuck lines and most people know only one of them:
//
// "In the unlikely event where the clock (SCL) is stuck LOW, the preferential procedure is to
// reset the bus using the HW reset signal ... If the I2C devices do not have HW reset inputs,
// cycle power ... If the data line (SDA) is stuck LOW, the master should send nine clock
// pulses. The device that held the bus LOW should release it some time within those nine
// clocks."
//
// The asymmetry is the whole point, and it is not arbitrary:
//
// SDA stuck LOW -> send nine clock pulses. A protocol-level recovery EXISTS, because the thing
// you need in order to perform it (the clock) is still yours.
// SCL stuck LOW -> HW reset, or cycle power. There is NO protocol-level recovery, because
// clocking your way out requires the clock you have just lost.
//
// An unbounded stretch is an SCL-stuck-low condition. So the honest thing for a guard to report is
// not "violation" but "this needs a reset line or a power cycle, and no amount of bus activity
// will help" -- which is a very different message to send to a driver.
// (Verilog-2001 -- structurally identical to the SystemVerilog above.)
module i2c_stretch_guard #(
parameter TICK_W = 24,
// The POLICY limit. 0 = no timeout, which is what plain I2C specifies. A system that needs a
// bound sets it; nothing in the protocol does.
parameter T_LIMIT = 0,
// The ESCALATION threshold, beyond which the line is treated as stuck rather than slow.
// Also 0 = never escalate.
parameter T_STUCK = 0
)(
input wire clk,
input wire rst_n,
// intent and observation, for both lines
input wire scl_release,
input wire sda_release,
input wire scl_in,
input wire sda_in,
// ---- the stretch, as Chapter 12.1 defines it ----
output reg stretching,
output wire [TICK_W-1:0] stretch_ticks,
output reg [TICK_W-1:0] max_stretch_seen,
// ---- POLICY, not protocol. Named so that nobody reads it as a compliance failure. ----
output reg warn_over_limit,
output reg [TICK_W-1:0] n_warn,
// ---- escalation, and the remedy section 3.1.16 prescribes for each line ----
output reg scl_stuck_low,
output reg sda_stuck_low,
output reg [1:0] recovery_action,
output reg [TICK_W-1:0] n_escalations
);
// Section 3.1.16's two remedies, as an enumeration so a driver cannot confuse them.
localparam [1:0] RECOV_NONE = 2'd0;
localparam [1:0] RECOV_NINE_CLOCKS = 2'd1; // SDA stuck: a protocol recovery exists
localparam [1:0] RECOV_HW_RESET = 2'd2; // SCL stuck: it does not
// A limit of zero means "no limit", and it must be impossible for any comparison to fire.
// Writing this as a wire rather than testing T_LIMIT at each use keeps the neutral case in
// one place; mutation D2 removes the guard and is killed by a 5000-tick stretch at T_LIMIT=0.
wire limit_enabled = (T_LIMIT != 0);
wire stuck_enabled = (T_STUCK != 0);
// Sized ONCE. A part-select of a parameter -- `LIMIT_W` -- is not portable:
// Icarus evaluates it as zero, which would make every comparison fire immediately. Design 3
// of this module lost time to exactly that, so both thresholds are widened here instead.
localparam [TICK_W-1:0] LIMIT_W = T_LIMIT;
localparam [TICK_W-1:0] STUCK_W = T_STUCK;
wire scl_held = scl_release && !scl_in; // somebody else is holding SCL low
wire sda_held = sda_release && !sda_in; // somebody else is holding SDA low
reg [TICK_W-1:0] ticks_q, sda_ticks_q;
wire [TICK_W-1:0] ticks_now = ticks_q + 1'b1;
wire [TICK_W-1:0] sda_ticks_now = sda_ticks_q + 1'b1;
// SCL dominance, decided COMBINATIONALLY on the condition rather than on the registered flag.
// Both lines can cross the escalation threshold in the SAME cycle, and in that cycle
// `scl_stuck_low` still holds its old value -- so a test of the flag reads 0, the SDA branch
// runs, and because it is written later in the same block its assignment to recovery_action
// wins. The result is a recommendation to send nine clock pulses on a clock line that is being
// held down, which is a loop a driver cannot exit.
//
// The general rule: a PRIORITY decision between two conditions must be made on the
// conditions, never on the flags that record them.
wire scl_stuck_c = stuck_enabled
&& (scl_stuck_low || (scl_held && stretching && (ticks_now > STUCK_W)));
always @(posedge clk) begin
if (!rst_n) begin
stretching <= 1'b0;
ticks_q <= {TICK_W{1'b0}};
sda_ticks_q <= {TICK_W{1'b0}};
max_stretch_seen <= {TICK_W{1'b0}};
warn_over_limit <= 1'b0;
n_warn <= {TICK_W{1'b0}};
scl_stuck_low <= 1'b0;
sda_stuck_low <= 1'b0;
recovery_action <= RECOV_NONE;
n_escalations <= {TICK_W{1'b0}};
end else begin
// ---- SCL: the stretch ----
if (scl_held) begin
if (!stretching) begin
stretching <= 1'b1;
ticks_q <= {{(TICK_W-1){1'b0}}, 1'b1}; // "including this cycle"
end else begin
ticks_q <= ticks_now;
if (ticks_now > max_stretch_seen) max_stretch_seen <= ticks_now;
// POLICY. Raised once per stretch, on the cycle the limit is passed, and only
// if a limit was configured at all.
if (limit_enabled && !warn_over_limit && ticks_now > LIMIT_W) begin
warn_over_limit <= 1'b1;
n_warn <= n_warn + 1'b1;
end
// ESCALATION. Past this point the line is not slow, it is stuck, and section
// 3.1.16 says the remedy is out-of-band.
if (stuck_enabled && !scl_stuck_low && ticks_now > STUCK_W) begin
scl_stuck_low <= 1'b1;
recovery_action <= RECOV_HW_RESET;
n_escalations <= n_escalations + 1'b1;
end
end
end else begin
// Released. Everything about THIS stretch clears; the totals do not.
stretching <= 1'b0;
ticks_q <= {TICK_W{1'b0}};
warn_over_limit <= 1'b0;
if (scl_stuck_low) begin
scl_stuck_low <= 1'b0;
// The SCL remedy no longer applies. If SDA is still held, the surviving
// condition is the one with a protocol recovery.
recovery_action <= sda_stuck_low ? RECOV_NINE_CLOCKS : RECOV_NONE;
end
end
// ---- SDA: stuck low, which has its own and much better remedy ----
if (sda_held) begin
sda_ticks_q <= sda_ticks_now;
if (stuck_enabled && !sda_stuck_low && sda_ticks_now > STUCK_W) begin
sda_stuck_low <= 1'b1;
n_escalations <= n_escalations + 1'b1;
// SCL dominates if both are stuck: you cannot send nine clock pulses on a
// clock line somebody else is holding down. Reporting the nine-clock remedy
// while SCL is stuck would send a driver into a loop that cannot terminate.
if (!scl_stuck_c) recovery_action <= RECOV_NINE_CLOCKS;
end
end else begin
sda_ticks_q <= {TICK_W{1'b0}};
if (sda_stuck_low) begin
sda_stuck_low <= 1'b0;
recovery_action <= scl_stuck_low ? RECOV_HW_RESET : RECOV_NONE;
end
end
end
end
assign stretch_ticks = ticks_q;
endmodule `timescale 1ns/1ps
// 100 MHz sample clock. THREE instances, because the whole point of this block is that the
// interesting configuration is the DEGENERATE one:
//
// dut_none -- T_LIMIT = 0, T_STUCK = 0 : plain I2C. Section 4.2.2 puts no limit on a stretch,
// so this instance must NEVER warn, however long the stretch.
// dut_pol -- T_LIMIT = 500, T_STUCK = 0 : a system policy, no escalation.
// dut_full -- T_LIMIT = 500, T_STUCK = 2000 : policy plus section 3.1.16 escalation.
//
// All three see identical stimulus on the same wires in one simulation, which is the control
// structure Chapter 11.8 section 2 argues for: a control in a separate run can diverge in its
// stimulus without anyone noticing.
// (Verilog-2001 testbench -- same stimulus, same checks.)
module i2c_stretch_guard_tb;
localparam TICK_W = 24;
reg clk = 1'b0, rst_n = 1'b0;
always #5 clk = ~clk;
reg m_scl_release = 1'b1, m_sda_release = 1'b1;
reg s_scl_hold = 1'b0, s_sda_hold = 1'b0;
wire scl_line = m_scl_release && !s_scl_hold;
wire sda_line = m_sda_release && !s_sda_hold;
// ---- the no-timeout instance: plain I2C ----
wire n_stretching, n_warn_over;
wire [TICK_W-1:0] n_max;
wire [TICK_W-1:0] n_nwarn;
wire [TICK_W-1:0] n_ticks, n_nesc;
wire n_scl_stuck, n_sda_stuck;
wire [1:0] n_recov;
i2c_stretch_guard #(.TICK_W(TICK_W), .T_LIMIT(0), .T_STUCK(0)) dut_none (
.clk(clk), .rst_n(rst_n),
.scl_release(m_scl_release), .sda_release(m_sda_release),
.scl_in(scl_line), .sda_in(sda_line),
.stretching(n_stretching), .stretch_ticks(n_ticks), .max_stretch_seen(n_max),
.warn_over_limit(n_warn_over), .n_warn(n_nwarn),
.scl_stuck_low(n_scl_stuck), .sda_stuck_low(n_sda_stuck),
.recovery_action(n_recov), .n_escalations(n_nesc));
// ---- policy only ----
wire p_stretching, p_warn_over;
wire [TICK_W-1:0] p_max;
wire [TICK_W-1:0] p_nwarn;
wire [TICK_W-1:0] p_ticks, p_nesc;
wire p_scl_stuck, p_sda_stuck;
wire [1:0] p_recov;
i2c_stretch_guard #(.TICK_W(TICK_W), .T_LIMIT(500), .T_STUCK(0)) dut_pol (
.clk(clk), .rst_n(rst_n),
.scl_release(m_scl_release), .sda_release(m_sda_release),
.scl_in(scl_line), .sda_in(sda_line),
.stretching(p_stretching), .stretch_ticks(p_ticks), .max_stretch_seen(p_max),
.warn_over_limit(p_warn_over), .n_warn(p_nwarn),
.scl_stuck_low(p_scl_stuck), .sda_stuck_low(p_sda_stuck),
.recovery_action(p_recov), .n_escalations(p_nesc));
// ---- policy plus escalation ----
wire f_stretching, f_warn_over;
wire [TICK_W-1:0] f_max;
wire [TICK_W-1:0] f_nwarn;
wire [TICK_W-1:0] f_ticks, f_nesc;
wire f_scl_stuck, f_sda_stuck;
wire [1:0] f_recov;
i2c_stretch_guard #(.TICK_W(TICK_W), .T_LIMIT(500), .T_STUCK(2000)) dut_full (
.clk(clk), .rst_n(rst_n),
.scl_release(m_scl_release), .sda_release(m_sda_release),
.scl_in(scl_line), .sda_in(sda_line),
.stretching(f_stretching), .stretch_ticks(f_ticks), .max_stretch_seen(f_max),
.warn_over_limit(f_warn_over), .n_warn(f_nwarn),
.scl_stuck_low(f_scl_stuck), .sda_stuck_low(f_sda_stuck),
.recovery_action(f_recov), .n_escalations(f_nesc));
localparam [1:0] RECOV_NONE = 2'd0, RECOV_NINE = 2'd1, RECOV_HW = 2'd2;
integer errors = 0;
// Hoisted to module scope: Verilog-2001 permits a variable declaration only at
// module level or in a NAMED block, and every call site below is sequential.
integer w0 = 0;
integer e0 = 0;
task tick(input integer n); begin repeat (n) @(negedge clk); end endtask
// A slave holds SCL low for `len` ticks while the master has released it.
task hold_scl(input integer len);
begin
m_scl_release = 1'b0; tick(20);
s_scl_hold = 1'b1; m_scl_release = 1'b1;
tick(len);
s_scl_hold = 1'b0; tick(20);
end
endtask
task hold_sda(input integer len);
begin
s_sda_hold = 1'b1; m_sda_release = 1'b1;
tick(len);
s_sda_hold = 1'b0; tick(20);
end
endtask
initial begin
tick(4); rst_n = 1'b1; tick(4);
// ---- 1. reset ------------------------------------------------------------------------
if (n_max !== {TICK_W{1'b0}} || p_max !== {TICK_W{1'b0}} || f_max !== {TICK_W{1'b0}}) begin
$display("FAIL: max_stretch_seen nonzero out of reset"); errors = errors + 1; end
if (n_recov !== RECOV_NONE || f_recov !== RECOV_NONE) begin
$display("FAIL: a recovery action was recommended out of reset"); errors = errors + 1; end
if (n_warn_over !== 1'b0 || p_warn_over !== 1'b0) begin
$display("FAIL: warn asserted out of reset"); errors = errors + 1; end
// ---- 2. a short stretch: nobody warns ------------------------------------------------
hold_scl(100);
if (n_nwarn !== {TICK_W{1'b0}} || p_nwarn !== {TICK_W{1'b0}} || f_nwarn !== {TICK_W{1'b0}}) begin
$display("FAIL: a 100-tick stretch warned (none=%0d pol=%0d full=%0d)",
n_nwarn, p_nwarn, f_nwarn); errors = errors + 1; end
if (n_max != 100 || p_max != 100) begin
$display("FAIL: a 100-tick stretch measured none=%0d pol=%0d", n_max, p_max);
errors = errors + 1; end
// ---- 3. THE test of this chapter: a 5000-tick stretch with NO timeout configured ------
// Section 4.2.2: "There is no limit in the I2C-bus protocol as to how long this delay can
// be." The T_LIMIT = 0 instance must stay silent. If it warns, the block has invented a
// protocol rule the specification explicitly declines to state.
hold_scl(5000);
if (n_nwarn !== {TICK_W{1'b0}}) begin
$display("FAIL: with T_LIMIT = 0 a 5000-tick stretch produced %0d warning(s) -- the block invented a limit the protocol does not have",
n_nwarn); errors = errors + 1; end
if (n_scl_stuck !== 1'b0 || n_recov !== RECOV_NONE) begin
$display("FAIL: with T_STUCK = 0 a 5000-tick stretch escalated"); errors = errors + 1; end
if (n_max != 5000) begin
$display("FAIL: the no-timeout instance measured %0d, expected 5000", n_max);
errors = errors + 1; end
// ---- 4. the policy instance DOES warn on the same stimulus ----------------------------
// The positive control for test 3. Without it, a block that never warned at all would
// pass test 3 perfectly.
if (p_nwarn !== 1) begin
$display("FAIL: with T_LIMIT = 500 a 5000-tick stretch produced %0d warnings, expected 1",
p_nwarn); errors = errors + 1; end
// ---- 5. escalation, and the remedy section 3.1.16 prescribes for SCL ------------------
// 5000 > T_STUCK of 2000, so the full instance must escalate -- and the recommendation
// must be HW RESET, because you cannot clock your way out of a held clock line.
if (f_nesc !== 1) begin
$display("FAIL: the escalating instance recorded %0d escalations, expected 1", f_nesc);
errors = errors + 1; end
// ---- 6. the warning and the escalation both CLEAR when the line is released ----------
// A stretch that ended is over. A guard that latched forever would report a stuck bus for
// the rest of time after one slow device.
if (p_warn_over !== 1'b0) begin
$display("FAIL: warn_over_limit still set after the line was released"); errors = errors + 1; end
if (f_scl_stuck !== 1'b0) begin
$display("FAIL: scl_stuck_low still set after the line was released"); errors = errors + 1; end
if (f_recov !== RECOV_NONE) begin
$display("FAIL: a recovery action is still recommended after recovery (%0d)", f_recov);
errors = errors + 1; end
// ---- 7. exactly AT the limit is not over it ------------------------------------------
// The comparison is strictly greater-than, so a stretch of exactly T_LIMIT is within
// policy. One tick more is not. Both halves are needed: a block that warned at the limit
// passes the second check and fails the first.
begin
w0 = p_nwarn;
hold_scl(500);
if (p_nwarn != w0) begin
$display("FAIL: a stretch of EXACTLY the 500-tick limit warned"); errors = errors + 1; end
hold_scl(501);
if (p_nwarn != w0 + 1) begin
$display("FAIL: a stretch of limit+1 did not warn"); errors = errors + 1; end
end
// ---- 8. SDA stuck low: the OTHER remedy ---------------------------------------------
// Section 3.1.16: "If the data line (SDA) is stuck LOW, the master should send nine clock
// pulses." A protocol recovery exists here, and reporting the SCL remedy instead would
// send a driver to reset hardware it does not need to touch.
begin
e0 = f_nesc;
// Hold SDA past the threshold and check the remedy WHILE it is still held. Checking
// only after release misses the whole point: the remedy is what a driver acts on, and
// it is only meaningful during the fault. Mutation M5 survived the first suite for
// exactly this reason -- the escalation count advanced either way.
s_sda_hold = 1'b1; m_sda_release = 1'b1;
tick(2100);
if (f_sda_stuck !== 1'b1) begin
$display("FAIL: SDA held past the threshold did not set sda_stuck_low"); errors = errors + 1; end
if (f_recov !== RECOV_NINE) begin
$display("FAIL: SDA stuck alone recommended remedy %0d, expected nine clock pulses (%0d) -- section 3.1.16 gives SDA a protocol recovery that SCL does not have",
f_recov, RECOV_NINE); errors = errors + 1; end
s_sda_hold = 1'b0; tick(20);
if (f_sda_stuck !== 1'b0) begin
$display("FAIL: sda_stuck_low still set after SDA was released"); errors = errors + 1; end
if (f_recov !== RECOV_NONE) begin
$display("FAIL: a remedy is still recommended after SDA was released"); errors = errors + 1; end
if (f_nesc != e0 + 1) begin
$display("FAIL: an SDA stuck-low condition did not escalate"); errors = errors + 1; end
end
// ---- 9. both lines stuck: SCL dominates ---------------------------------------------
// You cannot send nine clock pulses on a clock line somebody else is holding down, so a
// guard that recommended the nine-clock recovery here would put a driver into a loop that
// cannot terminate.
begin
s_scl_hold = 1'b1; s_sda_hold = 1'b1;
m_scl_release = 1'b1; m_sda_release = 1'b1;
tick(2500);
if (f_recov !== RECOV_HW) begin
$display("FAIL: with BOTH lines stuck the remedy was %0d, expected HW reset (%0d) -- nine clock pulses are impossible without the clock",
f_recov, RECOV_HW); errors = errors + 1; end
if (f_scl_stuck !== 1'b1 || f_sda_stuck !== 1'b1) begin
$display("FAIL: both lines stuck but scl=%b sda=%b", f_scl_stuck, f_sda_stuck);
errors = errors + 1; end
// Release SCL only. SDA is still stuck, so the remedy must fall back to nine clocks.
s_scl_hold = 1'b0; tick(40);
if (f_recov !== RECOV_NINE) begin
$display("FAIL: with SCL freed and SDA still stuck the remedy was %0d, expected nine clocks (%0d)",
f_recov, RECOV_NINE); errors = errors + 1; end
s_sda_hold = 1'b0; tick(40);
if (f_recov !== RECOV_NONE) begin
$display("FAIL: both lines free but a remedy is still recommended (%0d)", f_recov);
errors = errors + 1; end
end
// ---- 10. this device driving SCL low is never a stretch ------------------------------
begin
w0 = p_nwarn;
m_scl_release = 1'b0;
tick(3000);
m_scl_release = 1'b1;
tick(20);
if (p_nwarn != w0) begin
$display("FAIL: this device holding SCL low for 3000 ticks warned"); errors = errors + 1; end
end
// ---- 11. the worst case is a MAXIMUM and survives recovery ---------------------------
if (p_max != 5000) begin
$display("FAIL: max_stretch_seen = %0d, expected the 5000-tick peak to persist", p_max);
errors = errors + 1; end
if (errors == 0)
$display("PASS: T_LIMIT=0 never warns because the protocol states no limit, a policy limit is not a protocol violation, exactly at the limit is within it, and section 3.1.16's two remedies are kept apart with SCL dominating");
else
$display("FAIL: %0d error(s)", errors);
$finish;
end
initial begin
#6000000;
$display("FAIL: watchdog expired");
$finish;
end
endmodule -- STRETCH BOUNDS, AND WHY THEY ARE NOT PROTOCOL CHECKS -- the VHDL form.
--
-- UM10204 section 4.2.2: "There is no limit in the I2C-bus protocol as to how long this delay can
-- be, whereas for a SMBus system, it would be limited to 35 ms."
--
-- So a timeout here is a SYSTEM POLICY, never a compliance check, and T_LIMIT = 0 -- no timeout --
-- is the neutral value the protocol itself specifies.
--
-- UM10204 section 3.1.16: "In the unlikely event where the clock (SCL) is stuck LOW, the
-- preferential procedure is to reset the bus using the HW reset signal ... If the data line (SDA)
-- is stuck LOW, the master should send nine clock pulses."
--
-- Two lines, two remedies, and only one of them is a protocol action -- because clocking your way
-- out of trouble requires the clock you have just lost.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity i2c_stretch_guard is
generic (
TICK_W : natural := 24;
T_LIMIT : natural := 0; -- policy warn threshold; 0 = no timeout (plain I2C)
T_STUCK : natural := 0 -- escalation threshold; 0 = never escalate
);
port (
clk : in std_logic;
rst_n : in std_logic;
scl_release : in std_logic;
sda_release : in std_logic;
scl_in : in std_logic;
sda_in : in std_logic;
stretching : out std_logic;
stretch_ticks : out unsigned(TICK_W-1 downto 0);
max_stretch_seen : out unsigned(TICK_W-1 downto 0);
-- POLICY, not protocol. Named so nobody reads it as a compliance failure.
warn_over_limit : out std_logic;
n_warn : out unsigned(TICK_W-1 downto 0);
scl_stuck_low : out std_logic;
sda_stuck_low : out std_logic;
recovery_action : out unsigned(1 downto 0);
n_escalations : out unsigned(TICK_W-1 downto 0)
);
end entity;
architecture rtl of i2c_stretch_guard is
-- Section 3.1.16's two remedies, enumerated so a driver cannot confuse them.
constant RECOV_NONE : unsigned(1 downto 0) := "00";
constant RECOV_NINE_CLOCKS : unsigned(1 downto 0) := "01"; -- SDA: a protocol recovery exists
constant RECOV_HW_RESET : unsigned(1 downto 0) := "10"; -- SCL: it does not
constant LIMIT_W : unsigned(TICK_W-1 downto 0) := to_unsigned(T_LIMIT, TICK_W);
constant STUCK_W : unsigned(TICK_W-1 downto 0) := to_unsigned(T_STUCK, TICK_W);
signal limit_enabled : boolean := (T_LIMIT /= 0);
signal stuck_enabled : boolean := (T_STUCK /= 0);
signal scl_held, sda_held : std_logic;
signal s_stretching : std_logic := '0';
signal ticks_q : unsigned(TICK_W-1 downto 0) := (others => '0');
signal sda_ticks_q : unsigned(TICK_W-1 downto 0) := (others => '0');
signal ticks_now, sda_ticks_now : unsigned(TICK_W-1 downto 0);
signal r_max : unsigned(TICK_W-1 downto 0) := (others => '0');
signal r_nwarn : unsigned(TICK_W-1 downto 0) := (others => '0');
signal r_nesc : unsigned(TICK_W-1 downto 0) := (others => '0');
signal s_warn : std_logic := '0';
signal s_scl_stuck : std_logic := '0';
signal s_sda_stuck : std_logic := '0';
signal s_recov : unsigned(1 downto 0) := RECOV_NONE;
-- SCL dominance, decided COMBINATIONALLY on the condition and not on the registered flag.
-- Both lines can cross the threshold in the SAME cycle, and in that cycle s_scl_stuck still
-- holds its old value -- so a test of the flag reads '0', the SDA branch runs, and because it
-- is written later in the same process its assignment wins. The result is a recommendation to
-- send nine clock pulses on a clock line that is being held down: a loop with no exit.
--
-- The general rule: a PRIORITY decision between two conditions must be made on the conditions,
-- never on the flags that record them.
signal scl_stuck_c : boolean;
begin
scl_held <= '1' when (scl_release = '1' and scl_in = '0') else '0';
sda_held <= '1' when (sda_release = '1' and sda_in = '0') else '0';
ticks_now <= ticks_q + 1;
sda_ticks_now <= sda_ticks_q + 1;
scl_stuck_c <= stuck_enabled and
(s_scl_stuck = '1' or
(scl_held = '1' and s_stretching = '1' and ticks_now > STUCK_W));
stretching <= s_stretching;
stretch_ticks <= ticks_q;
max_stretch_seen <= r_max;
warn_over_limit <= s_warn;
n_warn <= r_nwarn;
scl_stuck_low <= s_scl_stuck;
sda_stuck_low <= s_sda_stuck;
recovery_action <= s_recov;
n_escalations <= r_nesc;
process (clk) is
begin
if rising_edge(clk) then
if rst_n = '0' then
s_stretching <= '0';
ticks_q <= (others => '0');
sda_ticks_q <= (others => '0');
r_max <= (others => '0');
s_warn <= '0';
r_nwarn <= (others => '0');
s_scl_stuck <= '0';
s_sda_stuck <= '0';
s_recov <= RECOV_NONE;
r_nesc <= (others => '0');
else
-- ---- SCL: the stretch ----
if scl_held = '1' then
if s_stretching = '0' then
s_stretching <= '1';
ticks_q <= to_unsigned(1, TICK_W); -- "including this cycle"
else
ticks_q <= ticks_now;
if ticks_now > r_max then
r_max <= ticks_now;
end if;
-- POLICY: once per stretch, and only if a limit was configured at all.
if limit_enabled and s_warn = '0' and ticks_now > LIMIT_W then
s_warn <= '1';
r_nwarn <= r_nwarn + 1;
end if;
-- ESCALATION: past this point the line is stuck, not slow, and section
-- 3.1.16 says the remedy is out of band.
if stuck_enabled and s_scl_stuck = '0' and ticks_now > STUCK_W then
s_scl_stuck <= '1';
s_recov <= RECOV_HW_RESET;
r_nesc <= r_nesc + 1;
end if;
end if;
else
-- Released. This stretch's state clears; the totals do not.
s_stretching <= '0';
ticks_q <= (others => '0');
s_warn <= '0';
if s_scl_stuck = '1' then
s_scl_stuck <= '0';
if s_sda_stuck = '1' then
s_recov <= RECOV_NINE_CLOCKS;
else
s_recov <= RECOV_NONE;
end if;
end if;
end if;
-- ---- SDA: stuck low, which has its own and much better remedy ----
if sda_held = '1' then
sda_ticks_q <= sda_ticks_now;
if stuck_enabled and s_sda_stuck = '0' and sda_ticks_now > STUCK_W then
s_sda_stuck <= '1';
r_nesc <= r_nesc + 1;
if not scl_stuck_c then
s_recov <= RECOV_NINE_CLOCKS;
end if;
end if;
else
sda_ticks_q <= (others => '0');
if s_sda_stuck = '1' then
s_sda_stuck <= '0';
if s_scl_stuck = '1' then
s_recov <= RECOV_HW_RESET;
else
s_recov <= RECOV_NONE;
end if;
end if;
end if;
end if;
end if;
end process;
end architecture; -- The VHDL testbench for the stretch guard. THREE instances on the same wires in one simulation,
-- because the interesting configuration here is the DEGENERATE one:
--
-- dut_none : T_LIMIT = 0, T_STUCK = 0 -- plain I2C: must NEVER warn, however long the stretch
-- dut_pol : T_LIMIT = 500, T_STUCK = 0 -- a system policy, no escalation
-- dut_full : T_LIMIT = 500, T_STUCK = 2000 -- policy plus section 3.1.16 escalation
--
-- One simulation, identical stimulus: a control in a separate run can diverge without anyone
-- noticing, which is the argument Chapter 11.8 section 2 makes for an in-simulation control.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity i2c_stretch_guard_tb is
end entity;
architecture tb of i2c_stretch_guard_tb is
constant TICK_W : natural := 24;
constant RECOV_NONE : unsigned(1 downto 0) := "00";
constant RECOV_NINE : unsigned(1 downto 0) := "01";
constant RECOV_HW : unsigned(1 downto 0) := "10";
signal clk : std_logic := '0';
signal rst_n : std_logic := '0';
signal m_scl_release : std_logic := '1';
signal m_sda_release : std_logic := '1';
signal s_scl_hold : std_logic := '0';
signal s_sda_hold : std_logic := '0';
signal scl_line, sda_line : std_logic;
signal n_stretching, n_warn_over, n_scl_stuck, n_sda_stuck : std_logic;
signal n_ticks, n_max, n_nwarn, n_nesc : unsigned(TICK_W-1 downto 0);
signal n_recov : unsigned(1 downto 0);
signal p_stretching, p_warn_over, p_scl_stuck, p_sda_stuck : std_logic;
signal p_ticks, p_max, p_nwarn, p_nesc : unsigned(TICK_W-1 downto 0);
signal p_recov : unsigned(1 downto 0);
signal f_stretching, f_warn_over, f_scl_stuck, f_sda_stuck : std_logic;
signal f_ticks, f_max, f_nwarn, f_nesc : unsigned(TICK_W-1 downto 0);
signal f_recov : unsigned(1 downto 0);
signal done : boolean := false;
signal errors : integer := 0;
begin
scl_line <= m_scl_release and (not s_scl_hold);
sda_line <= m_sda_release and (not s_sda_hold);
clk_gen : process is
begin
while not done loop
clk <= '0'; wait for 5 ns;
clk <= '1'; wait for 5 ns;
end loop;
wait;
end process;
dut_none : entity work.i2c_stretch_guard
generic map (TICK_W => TICK_W, T_LIMIT => 0, T_STUCK => 0)
port map (clk => clk, rst_n => rst_n,
scl_release => m_scl_release, sda_release => m_sda_release,
scl_in => scl_line, sda_in => sda_line,
stretching => n_stretching, stretch_ticks => n_ticks, max_stretch_seen => n_max,
warn_over_limit => n_warn_over, n_warn => n_nwarn,
scl_stuck_low => n_scl_stuck, sda_stuck_low => n_sda_stuck,
recovery_action => n_recov, n_escalations => n_nesc);
dut_pol : entity work.i2c_stretch_guard
generic map (TICK_W => TICK_W, T_LIMIT => 500, T_STUCK => 0)
port map (clk => clk, rst_n => rst_n,
scl_release => m_scl_release, sda_release => m_sda_release,
scl_in => scl_line, sda_in => sda_line,
stretching => p_stretching, stretch_ticks => p_ticks, max_stretch_seen => p_max,
warn_over_limit => p_warn_over, n_warn => p_nwarn,
scl_stuck_low => p_scl_stuck, sda_stuck_low => p_sda_stuck,
recovery_action => p_recov, n_escalations => p_nesc);
dut_full : entity work.i2c_stretch_guard
generic map (TICK_W => TICK_W, T_LIMIT => 500, T_STUCK => 2000)
port map (clk => clk, rst_n => rst_n,
scl_release => m_scl_release, sda_release => m_sda_release,
scl_in => scl_line, sda_in => sda_line,
stretching => f_stretching, stretch_ticks => f_ticks, max_stretch_seen => f_max,
warn_over_limit => f_warn_over, n_warn => f_nwarn,
scl_stuck_low => f_scl_stuck, sda_stuck_low => f_sda_stuck,
recovery_action => f_recov, n_escalations => f_nesc);
stim : process is
procedure tick(n : in integer) is
begin
for i in 1 to n loop
wait until falling_edge(clk);
end loop;
end procedure;
procedure hold_scl(len : in integer) is
begin
m_scl_release <= '0'; tick(20);
s_scl_hold <= '1'; m_scl_release <= '1';
tick(len);
s_scl_hold <= '0'; tick(20);
end procedure;
procedure hold_sda(len : in integer) is
begin
s_sda_hold <= '1'; m_sda_release <= '1';
tick(len);
s_sda_hold <= '0'; tick(20);
end procedure;
procedure chk(cond : in boolean; msg : in string) is
begin
if not cond then
report "FAIL: " & msg severity error;
errors <= errors + 1;
wait for 0 ns;
end if;
end procedure;
variable w0, e0 : integer;
begin
tick(4); rst_n <= '1'; tick(4);
-- 1. reset
chk(n_max = 0 and p_max = 0 and f_max = 0, "max_stretch_seen nonzero out of reset");
chk(n_recov = RECOV_NONE and f_recov = RECOV_NONE,
"a recovery action was recommended out of reset");
chk(n_warn_over = '0' and p_warn_over = '0', "warn asserted out of reset");
-- 2. a short stretch: nobody warns
hold_scl(100);
chk(n_nwarn = 0 and p_nwarn = 0 and f_nwarn = 0, "a 100-tick stretch warned");
chk(to_integer(n_max) = 100 and to_integer(p_max) = 100,
"a 100-tick stretch was measured wrongly");
-- 3. THE test of this chapter: 5000 ticks with NO timeout configured. Section 4.2.2 puts
-- no limit on a stretch, so the T_LIMIT = 0 instance must stay silent.
hold_scl(5000);
chk(n_nwarn = 0,
"with T_LIMIT = 0 a 5000-tick stretch warned -- the block invented a limit the protocol does not have");
chk(n_scl_stuck = '0' and n_recov = RECOV_NONE,
"with T_STUCK = 0 a 5000-tick stretch escalated");
chk(to_integer(n_max) = 5000, "the no-timeout instance measured the wrong peak");
-- 4. the positive control: the policy instance DOES warn on the same stimulus
chk(to_integer(p_nwarn) = 1, "with T_LIMIT = 500 a 5000-tick stretch did not warn once");
-- 5. escalation on the full instance
chk(to_integer(f_nesc) = 1, "the escalating instance did not record one escalation");
-- 6. warning and escalation both CLEAR when the line is released
chk(p_warn_over = '0', "warn_over_limit still set after release");
chk(f_scl_stuck = '0', "scl_stuck_low still set after release");
chk(f_recov = RECOV_NONE, "a recovery action is still recommended after recovery");
-- 7. exactly AT the limit is not over it
w0 := to_integer(p_nwarn);
hold_scl(500);
chk(to_integer(p_nwarn) = w0, "a stretch of EXACTLY the limit warned");
hold_scl(501);
chk(to_integer(p_nwarn) = w0 + 1, "a stretch of limit+1 did not warn");
-- 8. SDA stuck low: the OTHER remedy (section 3.1.16 -- nine clock pulses). The remedy is
-- checked WHILE the line is held, because that is when a driver would act on it.
e0 := to_integer(f_nesc);
s_sda_hold <= '1'; m_sda_release <= '1';
tick(2100);
chk(f_sda_stuck = '1', "SDA held past the threshold did not set sda_stuck_low");
chk(f_recov = RECOV_NINE,
"SDA stuck alone did not recommend nine clock pulses -- section 3.1.16 gives SDA a protocol recovery that SCL does not have");
s_sda_hold <= '0'; tick(20);
chk(f_sda_stuck = '0', "sda_stuck_low still set after SDA was released");
chk(f_recov = RECOV_NONE, "a remedy is still recommended after SDA was released");
chk(to_integer(f_nesc) = e0 + 1, "an SDA stuck-low condition did not escalate");
-- 9. both lines stuck: SCL dominates, because nine clock pulses need the clock
s_scl_hold <= '1'; s_sda_hold <= '1';
m_scl_release <= '1'; m_sda_release <= '1';
tick(2500);
chk(f_recov = RECOV_HW,
"with BOTH lines stuck the remedy was not HW reset -- nine clock pulses are impossible without the clock");
chk(f_scl_stuck = '1' and f_sda_stuck = '1', "both lines stuck but the flags disagree");
s_scl_hold <= '0'; tick(40);
chk(f_recov = RECOV_NINE,
"with SCL freed and SDA still stuck the remedy was not nine clocks");
s_sda_hold <= '0'; tick(40);
chk(f_recov = RECOV_NONE, "both lines free but a remedy is still recommended");
-- 10. this device driving SCL low is never a stretch
w0 := to_integer(p_nwarn);
m_scl_release <= '0';
tick(3000);
m_scl_release <= '1';
tick(20);
chk(to_integer(p_nwarn) = w0, "this device holding SCL low for 3000 ticks warned");
-- 11. the worst case is a MAXIMUM and survives recovery
chk(to_integer(p_max) = 5000, "max_stretch_seen did not keep the 5000-tick peak");
if errors = 0 then
report "i2c_stretch_guard self-check complete: T_LIMIT=0 never warns because the protocol states no limit, a policy limit is not a protocol violation, exactly at the limit is within it, and section 3.1.16's two remedies are kept apart with SCL dominating" severity note;
else
report "FAILURES in i2c_stretch_guard" severity error;
end if;
done <= true;
wait;
end process;
end architecture;6a. Five Decisions Worth Defending
T_LIMIT = 0 means no timeout, and it is the default. §2's callout. The specification's position is that there is no limit, so the parameter's neutral value must express it — and the comparison must be incapable of firing, not merely unlikely to. Mutation M1 makes zero active.
The output is warn_over_limit, never viol_*. A name is an interface. §2's table is what a reader does with each.
The two remedies of §3.1.16 are an enumeration, not a boolean. RECOV_NINE_CLOCKS and RECOV_HW_RESET are distinct values because they are distinct actions, and a driver that conflated them would either reset hardware unnecessarily or attempt an impossible recovery. §7's tests 8 and 9 exercise both.
SCL dominates when both lines are stuck, and the decision is made combinationally. Nine clock pulses are impossible on a held clock line, so recommending them would put a driver in a loop that cannot terminate. The dominance is computed from the condition rather than from the registered flag, and §8's mutation M3 is what happens otherwise — a bug this design had and the testbench caught.
The warning and the escalation both clear on release; the totals do not. A stretch that ended is over. A guard that latched forever would report a stuck bus for the rest of time after one slow device — while n_warn, n_escalations and max_stretch_seen must persist, because they are the record.
6b. Verified Execution
$ iverilog -g2012 -o d4 i2c_stretch_guard.sv i2c_stretch_guard_tb.sv && ./d4
PASS: T_LIMIT=0 never warns because the protocol states no limit, a policy limit is not a
protocol violation, exactly at the limit is within it, and section 3.1.16's two remedies are
kept apart with SCL dominating
i2c_stretch_guard_tb.sv:235: $finish called at 139890000 (1ps)
$ iverilog -g2005 -o v4 i2c_stretch_guard.v i2c_stretch_guard_tb.v && ./v4
PASS: T_LIMIT=0 never warns because the protocol states no limit, a policy limit is not a
protocol violation, exactly at the limit is within it, and section 3.1.16's two remedies are
kept apart with SCL dominating
i2c_stretch_guard_tb.v:247: $finish called at 139890000 (1ps)
$ nvc -a i2c_stretch_guard.vhd i2c_stretch_guard_tb.vhd
$ nvc -e i2c_stretch_guard_tb && nvc -r i2c_stretch_guard_tb --stop-time=3000us
** Note: 139890ns+1: i2c_stretch_guard self-check complete: T_LIMIT=0 never warns because
the protocol states no limit, a policy limit is not a protocol violation, exactly at the
limit is within it, and section 3.1.16's two remedies are kept apart with SCL dominatingAll three at 139890 ns.
7. What the Testbench Proves
The suite instantiates the guard three times on the same wires in one simulation, because the interesting configuration here is the degenerate one:
| instance | T_LIMIT | T_STUCK | represents |
|---|---|---|---|
dut_none | 0 | 0 | plain I²C — must never warn |
dut_pol | 500 | 0 | a system policy, no escalation |
dut_full | 500 | 2000 | policy plus §3.1.16 escalation |
| # | stimulus | what it establishes |
|---|---|---|
| 1 | reset | no remedy recommended; no warning |
| 2 | a 100-tick stretch | nobody warns; measured as 100 |
| 3 | a 5000-tick stretch, T_LIMIT = 0 | silence — the protocol states no limit |
| 4 | the same stretch, T_LIMIT = 500 | one warning — the positive control |
| 5 | the same stretch, T_STUCK = 2000 | one escalation |
| 6 | the line released | warning and escalation both clear |
| 7 | exactly at the limit, then one over | within policy, then over it |
| 8 | SDA held alone past the threshold | nine clocks recommended, and it clears |
| 9 | both lines stuck | HW reset — then nine clocks when SCL frees, then none |
| 10 | this device holding SCL low | never a warning |
| 11 | after recovery | max_stretch_seen keeps its 5000-tick peak |
Test 3 is the most important test in the chapter, and test 4 is what makes it mean anything. Test 3 asserts silence, and a guard that never warned at all would pass it perfectly. Test 4 drives the identical stimulus into an instance with a limit configured and requires exactly one warning. The two together say the block discriminates — which is Chapter 11.8 §2's positive-control argument, and the reason all three instances share one simulation rather than three runs.
Test 9 is the escalation's priority, and it is three assertions rather than one. Both lines stuck must recommend HW reset; freeing SCL while SDA remains stuck must fall back to nine clocks; freeing both must recommend nothing. A guard that only recommended HW reset would pass the first and fail the second — and the fall-back is the useful part, because it is the transition from an unrecoverable state to a recoverable one.
Test 8 was strengthened after a mutation survived, and §8 records why. It originally checked only that the escalation counted and the flag cleared; it now checks the remedy while the line is still held, because the remedy is what a driver acts on and it is only meaningful during the fault.
Test 7 is the boundary, and the comparison is strictly greater-than. A stretch of exactly the policy limit is within policy. Both halves are needed: a guard warning at the limit passes the second check and fails the first.
8. Mutation Testing
Six defects injected into the SystemVerilog guard.
| # | injected defect | outcome |
|---|---|---|
| M1 | T_LIMIT = 0 is treated as an active limit of zero | killed — test 3 |
| M2 | the policy comparison is >= rather than > | killed — test 7 |
| M3 | SCL dominance is decided on the registered flag | killed — test 9 |
| M4 | the warning does not clear when the line is released | killed — test 6 |
| M5 | an SDA stuck-low condition recommends the SCL remedy | killed — test 8, after strengthening |
| M6 | the maximum tracker is initialised like a minimum tracker | killed — test 1 |
Six injected, six killed. Three worth recording.
M1 is the mutation this chapter exists to make impossible. Forcing limit_enabled true makes a zero limit active, so every stretch exceeds it and the block warns on all of them. The failure message is a 100-tick stretch warned (none=1 pol=0 full=0) — and the none=1 is the point: the instance configured to have no timeout produced one. A block that did this would be asserting a protocol rule that §4.2.2 explicitly declines to state.
M3 was a real bug in this design, found by the testbench rather than injected into a correct one. The first version tested the registered scl_stuck_low flag when deciding whether to recommend the nine-clock remedy. Both lines can cross the threshold in the same cycle, and in that cycle the flag still holds its old value — so the test read zero, the SDA branch ran, and because it is written later in the same block its assignment to recovery_action won. The guard recommended nine clock pulses on a clock line being held down: a loop a driver cannot exit.
The fix was to compute the dominance combinationally from the condition. The lesson generalises past this block:
A priority decision between two conditions must be made on the conditions, never on the flags that record them. In the cycle both arise, the flags are still stale, and whichever branch is written last wins.
M5 survived the first run, and it is the module's one genuine test gap. Changing the SDA branch to recommend RECOV_HW_RESET left the escalation count and the flags identical — and the original test 8 checked only those. Nothing asserted the remedy, which is the one output a driver actually consumes.
That is a precise instance of a general failure: the output that matters was not the output being checked. The escalation count says a fault was detected; the remedy says what to do about it, and §3's whole argument is that getting the remedy wrong sends a driver to reset hardware it need not touch — or, worse, to attempt a recovery that cannot work. Test 8 now checks the remedy while the line is held, and the mutation dies.
9. Verification Connection — Asserting a Policy Without Asserting a Rule
// There is deliberately NO property of the form:
//
// @(posedge clk) $rose(stretching) |-> ##[1:N] !stretching;
//
// for any N. Section 4.2.2 states that "there is no limit in the I2C-bus protocol as to how long
// this delay can be", so any bounded form asserts a rule the specification declines to state.
// Writing one turns a system policy into a compliance claim -- which is mutation M1 in hardware
// and a false bug report in practice.
// What IS assertable is unbounded liveness: a stretch must eventually end, with no bound on when.
property p_stretch_eventually_ends;
@(posedge clk) $rose(stretching) |-> s_eventually (!stretching);
endproperty
assert property (p_stretch_eventually_ends)
else $error("a stretch never ended -- SCL is stuck, and section 3.1.16's remedy is a hardware reset, not a protocol action");
// The POLICY is expressed as a coverage-style observation or a warning, never as an assertion.
// This is the shape: a non-fatal note, so that a long stretch shows up in a report without
// failing a regression that is testing something else.
always @(posedge clk)
if (warn_over_limit && !$past(warn_over_limit))
$info("stretch exceeded the %0d-tick POLICY limit (not a protocol violation)", T_LIMIT);
// The remedy must match the fault, which is section 3.1.16's asymmetry as a property. Nine clock
// pulses require the clock, so recommending them while SCL is stuck is never correct.
property p_remedy_matches_fault;
@(posedge clk) scl_stuck_low |-> (recovery_action != RECOV_NINE_CLOCKS);
endproperty
assert property (p_remedy_matches_fault)
else $error("nine clock pulses were recommended while SCL is stuck -- that recovery needs the clock it has lost");
// And the degenerate configuration, asserted rather than assumed -- the same discipline
// Chapter 11.8 section 9 applies to its T_SP = 0 instance.
if (T_LIMIT == 0) begin : g_no_timeout
assert property (@(posedge clk) !warn_over_limit)
else $error("T_LIMIT = 0 must mean NO timeout: the protocol states no limit");
end covergroup i2c_guard_cg with function sample(int ticks, int limit, bit scl_stuck,
bit sda_stuck, int remedy);
// The stretch length RELATIVE to the configured policy, because the absolute length means
// nothing without it -- 5000 ticks is comfortable for one product and fatal for another.
// A zero-limit configuration is its own bin, since no relative figure exists there.
margin: coverpoint (limit == 0 ? -1 : ticks - limit) {
bins no_policy = {-1}; // T_LIMIT = 0: the protocol's own position
bins well_under = {[$:-100]};
bins near_under = {[-99:-1]};
bins exactly = {0}; // at the limit: within policy
bins just_over = {[1:99]};
bins well_over = {[100:$]};
}
// THE cross of section 3: the fault and the remedy must correspond. The cell that must stay
// EMPTY is (scl stuck, nine clocks) -- an impossible recovery -- and covering the others
// proves the enumeration is actually being driven rather than defaulted.
fault: coverpoint {scl_stuck, sda_stuck} {
bins clean = {2'b00};
bins sda_only = {2'b01};
bins scl_only = {2'b10};
bins both = {2'b11};
}
action: coverpoint remedy {
bins none = {0};
bins nine = {1};
bins reset = {2};
}
fault_x_action: cross fault, action;
// Recovery TRANSITIONS, because section 7's test 9 shows the interesting behaviour is the
// fall-back: both stuck, then SCL freed, then a recoverable state. A suite that only ever
// samples steady states never sees it.
recovery_path: coverpoint remedy {
bins escalate = (0 => 2);
bins fallback = (2 => 1); // SCL freed while SDA still stuck
bins cleared = (1 => 0);
}
endgroup10. FPGA and ASIC Implications
The guard is two counters, three compares and a two-bit enumeration — around 110 flops at TICK_W = 24. Its cost is trivial and its value is entirely in what it lets a driver say.
Choose T_LIMIT from the slowest device's datasheet, not from the bus rate. This is the concrete design rule of the chapter. An EEPROM page write is milliseconds; a bus-rate-derived timeout is microseconds; the two differ by three orders of magnitude, and a timeout set from the wrong one fires on every legal write. The number has to come from the part, and if no part on the bus documents its worst case, the honest configuration is T_LIMIT = 0 with a much larger T_STUCK.
Size the counter for the policy, and then some. §1 means no width is provably sufficient (Chapter 12.1 §10 makes the same point), but a guard has a defensible bound that a detector does not: it only needs to count as far as T_STUCK. That makes the width a design choice rather than a guess — one of the few places in this module where a bound is legitimately available.
A master that can assert a hardware reset to its slaves should wire it. §3.1.16 names it as the preferential remedy for a stuck SCL, and it is the only remedy that does not involve cycling power. A board with I²C slaves that have reset inputs and no reset line to them has given up the specification's first recommendation for the sake of one net.
Report a diagnosis, not a timeout. "SCL held 4.2 ms, SDA free, remedy: hardware reset" is actionable. "I²C timeout" is not, and the difference is a handful of registers this block already has.
And the escalation threshold is a product decision that should be reviewable. T_STUCK is the moment a design stops believing a slave will recover. Putting it in a parameter — rather than in a magic number in a driver — is what makes that decision visible to the person who has to justify it.
11. Debugging — The Timeout That Was Set From the Bus Rate
Pitfall — deriving a stretch timeout from the clock frequency instead of the slowest device
// A Fast-mode driver for a board with sensors, a GPIO expander and a 256 kbit EEPROM. The
// bus-level timeout was derived, reasonably enough, from the bus rate:
//
// // A byte is 9 bits at 400 kHz = 22.5 us. Allow 40x for stretching. Generous.
// #define I2C_BYTE_TIMEOUT_US 900
//
// // and in the transfer loop:
// if (elapsed_us > I2C_BYTE_TIMEOUT_US) {
// log_error("I2C protocol violation: slave held clock too long");
// i2c_abort(); // drive a STOP, discard the transfer, return -EIO
// }
//
// 900 us against a 22.5 us byte is a factor of forty, which felt like a wide margin -- and for
// the sensors and the expander it was: none of them stretched at all, and every transfer to them
// completed in well under 30 us.
//
// Note the log message. It says "protocol violation", which is the framing that did the damage.EEPROM writes failed. Not always: single-byte writes succeeded, and page writes succeeded when the page happened to be already erased. Anything that triggered a real programming cycle returned -EIO about 80 % of the time.
The driver logged "I2C protocol violation: slave held clock too long" every time, so the investigation began where the log pointed: at the EEPROM. Its datasheet was checked for compliance, a second manufacturer's part was fitted, and the bus was scoped for signal integrity. All fine.
A bug was filed against the EEPROM vendor, with captures attached. The vendor's reply was short and correct: the part holds SCL low for up to 5 ms during a page programming cycle, this is documented on page 4 of the datasheet, and UM10204 section 4.2.2 places no limit on the length of a stretch.
Which is when somebody read section 4.2.2, and the timeout, and did the arithmetic:
the driver's tolerance 900 us the EEPROM's documented need 5 ms -- 5.6x longer the protocol's limit none
The EEPROM was compliant. The driver was aborting legal transfers and then blaming the slave for them. The 80 % failure rate was simply how often a programming cycle exceeded 900 us, and the successes were the writes that did not program anything.
The damage from the log message was larger than the damage from the timeout. It sent two engineers and a vendor after a part that was working correctly, because it asserted a violation of a rule that does not exist.
The timeout was derived from the bus rate, and a stretch has nothing to do with the bus rate. Section 4.2.2's whole point is that I2C "can be a 'DC' bus" -- a slave may hold the clock for as long as its internal work takes, and that work is set by physics inside the part rather than by the clock outside it. An EEPROM programming cycle is milliseconds because charge pumps are slow, at any bus frequency.
A factor of forty on the wrong quantity is still the wrong quantity.
The second and worse fault was calling the result a protocol violation. A timeout on an I2C stretch is a SYSTEM POLICY -- a statement about what this product can wait for -- and the specification explicitly declines to provide the rule the message claimed had been broken. That framing is what turned a one-line configuration error into a vendor escalation.
12. Common Misconceptions
"There must be some maximum stretch." §4.2.2: "There is no limit in the I²C-bus protocol as to how long this delay can be." I²C "can be a 'DC' bus".
"A long stretch is a protocol violation." It is legal. A timeout is a policy — a statement about what this product tolerates — and reporting it as a violation sends people after a compliant part.
"I²C and SMBus are the same here." SMBus bounds the stretch at 35 ms and therefore mandates a 10 kHz minimum clock. I²C does neither, and the two choices are coherent and incompatible.
"Nine clock pulses recover a stuck bus." They recover a stuck SDA. A stuck SCL has no protocol-level recovery, because generating pulses requires the clock that has been lost — §3.1.16 goes straight to hardware reset or power cycling.
"A timeout can distinguish a slow slave from a wedged one." It cannot: the two are the identical electrical condition, and §1 says no duration is too long. A policy is a guess informed by a datasheet, not a measurement.
"Derive the timeout from the bus rate." A stretch has nothing to do with the bus rate. §11 is a factor-of-forty margin on the wrong quantity.
"A master can always drive a STOP to abandon a transfer." Not while SCL is held low: a STOP is an SDA transition while SCL is high, and the master cannot raise a line somebody else is holding down.
"T_LIMIT = 0 is a meaningless configuration." It is the protocol's own position and must mean no timeout. Mutation M1 makes it an active limit of zero and warns on every stretch.
"If nothing on this bus stretches, the master need not handle it." True today, and it is a contract over every device that will ever be on the bus — invalidated by a board revision, with an electrical failure mode.
13. Reason It Through
What is the maximum time a slave may hold SCL low, according to the specification?
There is none. §4.2.2 says so explicitly and calls I²C a "DC bus" — the clock may stop indefinitely and the bus is still operating correctly. Any bound a design uses is a policy it has chosen.
Why does SMBus have a 10 kHz minimum clock frequency and I²C none?
Because SMBus chose to bound the stretch at 35 ms. Once you assert that nothing may take longer than a fixed time, you have implied a minimum rate and must state it. I²C declined the bound and therefore needs no floor — which is what lets a software slave in an interrupt handler participate at all.
Why can nine clock pulses clear a stuck SDA but not a stuck SCL?
Because the recovery is performed by clocking, and that requires the master to own SCL. On a stuck SDA it does. On a stuck SCL it does not, and the wired-AND guarantees it cannot take the line back — so there is no bus activity that helps, and §3.1.16 goes out of band.
Why is nine the right number of pulses?
Because a byte is nine bits, so no device can be more than eight bits into a transfer. Nine pulses guarantees every device reaches a byte boundary, where it sees a ninth bit it can read as a NACK and releases SDA.
A driver times out after 900 µs and logs "protocol violation". Two things are wrong. What are they, and which cost more?
The timeout is derived from the bus rate rather than from the slowest device's datasheet, so it fires on a legal 5 ms EEPROM write. And the message asserts a rule that does not exist. The message cost more: it directed two engineers and a vendor at a compliant part, which a warning about a local budget would not have done.
A driver's abort path drives a STOP when it times out. Why does that not work, and what does it leave behind?
A STOP is an SDA transition while SCL is high, and the timeout happened because SCL is being held low. The master cannot produce the edge, so the abort silently does nothing — and the driver returns an error while the bus is still mid-transfer, so the next transfer begins from an unknown state.
Both SCL and SDA are stuck. Why must the guard recommend a hardware reset rather than nine clock pulses?
Because nine clock pulses require the clock, which is stuck. Recommending them puts a driver in a loop it cannot exit. And the decision must be made on the conditions rather than the registered flags, because both can cross their thresholds in the same cycle — which is the bug §8's mutation M3 records.
14. Understanding Check
15. Summary
There is no limit. §4.2.2 states it outright and calls I²C a "DC bus": the clock may stop indefinitely and the bus is still correct.
So every timeout is a policy, and naming matters. warn_over_limit, not viol_*. The first keeps open the question of whether the budget is right; the second asserts a rule that does not exist and sends people after compliant parts.
T_LIMIT = 0 must mean no timeout, because that is the protocol's own position and the neutral value has to be expressible.
The two recovery paths are not equivalent. Nine clock pulses clear a stuck SDA because the master still owns the clock. A stuck SCL has no protocol-level recovery at all, because clocking your way out needs the clock — so §3.1.16 goes to hardware reset, then to power cycling.
An unbounded stretch and a stuck SCL are the same electrical condition, and no measurement separates them. Only a policy does.
But a master with a stuck clock still owns SDA, so it can diagnose rather than merely time out — and it can assert the reset line the specification prefers, if the board has one.
Derive the limit from the slowest device, never from the bus rate. They differ by orders of magnitude, and a factor of forty on the wrong quantity is still the wrong quantity.
And never assume a target will not stretch unless a contract guarantees it. The assumption's failure mode is electrical, it is invalidated by a board revision rather than a code change, and the cheapest insurance — a generator that sequences on the observed line — costs one state and no runtime.
16. What Comes Next
Module 12 is complete, and with it the last mechanism that can pause a transfer without ending it.
The module's arc was short and it inverted an assumption every earlier chapter made. Chapter 12.1 established that the slave has a way to push back on a clock it does not own. Chapter 12.2 showed the bus already contained the machinery — §3.1.7's clock synchronization is clock stretching with one noun changed — so tolerating it costs a master one state and the discipline to sequence on the line it observes rather than the line it intended. Chapter 12.3 priced it, and found that one of its two levels taxes transfers the slow device never took part in. This chapter found that the protocol sets no bound at all, and that the remedy for the worst case is not a protocol action.
What every chapter in this module took on faith is that the master holding the bus is the only master. arb_lost has been an input since Chapter 10.3, and Module 13 finally builds the mechanism behind it: how two masters that start transmitting simultaneously discover the fact, why the loser always finds out before it corrupts anything, and why on a wired-AND bus the winner is decided bit by bit with no arbitration protocol whatsoever.
It is the same substitution this module's §3.1.7 performed, run in the other direction — and it closes the causal chain Chapter 2.5 opened, where a single electrical rule about pulling low turned out to be the whole reason the bus works.
Continue learning
Related tutorials
- Related topic
Stuck Bus — Diagnosis and Recovery
A bus held low, who is holding it, and why the standard recovery works for exactly one of its seven causes. Builds a clear sequencer whose pulse count is evidence rather than a boolean, runs it against seven holders including a legally stretching target, and reports the four bugs it took to get there — two in the design and two in the bench.
- Related topic
Clock-Stretching and Arbitration Failures in the Field
Why a timeout is the least informative evidence an I²C controller produces, and what to record instead. Six situations that all report the same timeout, a classifier that separates them, and the one failure in the set that never times out at all — a controller that does not honour stretching, corrupting data silently and blaming the target.
- Related topic
I²C Transaction Atomicity and Bus Ownership Across Phases
What is and is not atomic on an I²C bus, stated precisely. Three things end bus ownership and one that looks like it should does not — and telling them apart needs one input the wire cannot supply.
- Related topic
Building an I²C Timing Budget
Eight chapters established sixteen parameters; this one adds them up and finds that per-parameter compliance is necessary and not sufficient. Builds the budget as synthesisable hardware and closes every deficit the module uncovered.
