I²C · Module 17
Timeout, Error Handling and Bus Recovery
UM10204 defines no error codes, so the taxonomy is a design decision — and collapsing two failures into one code makes them indistinguishable to software that could have acted on them differently. Builds six codes, separates protocol outcomes from faults, and implements the nine-pulse recovery including an honest account of what it cannot do.
Chapter 17.10 taught the master to notice when the bus disagrees with it. This chapter decides what to do about it — and the first thing to establish is how little of that decision the specification makes for you.
1. The Specification Defines No Error Codes
There is no status register in UM10204 and no list of failure modes. Every name in this block is this design's choice.
So this chapter is a design-decision chapter, not a derivation chapter — with one exception, in §4, where four sentences of specification force the most consequential decision in the block.
2. The Six Codes, and What Separates Each From Its Neighbours
| Code | Means | Separated from its neighbour because |
|---|---|---|
E_ADDR_NACK | the address was not acknowledged | nothing is there or the device is busy — 16.4 shows these are indistinguishable, so they share a code honestly |
E_DATA_NACK | a data byte was not acknowledged | the device is there and declined this byte — a write-protect pin, or a read-only location |
E_ARB_LOST | another master won | not an error in the usual sense — the bus worked exactly as designed |
E_STRETCH_TO | a stretch outlasted this master's bound | the target may still be conforming — §3.1.6 sets no bound |
E_SDA_STUCK | SDA held low with no transfer in progress | §3.1.16's nine-pulse remedy applies |
E_SCL_STUCK | SCL held low | §3.1.16 offers no protocol remedy |
2a. Two of these are not faults
Arbitration loss is normal operation. Another master won; the bus behaved exactly as §3.1.8 specifies. The correct response is to retry when the bus is free, not to report a fault to a user. A driver that surfaces it as an error will report errors on a working multi-master board.
A stretch timeout is a decision, not a discovery. §3.1.6 gives stretching no upper bound, so a target that stretches for longer than this master is willing to wait is still conforming. The flag records that the master gave up — and Chapter 12.4's no-assumption rule is why the bound must be a parameter with a stated default rather than a derived constant.
3. A Low Line Is Not a Stuck Line
Both lines are low most of the time during a normal transfer — that is what a transfer is.
So stuck-line detection is qualified by !in_transfer. Without that qualifier the master reports both lines stuck on every byte it sends, which is mutation M4 below and would make the error register useless rather than merely noisy.
4. The One Decision the Specification Forces
Two different lines, two different remedies — and the reason is structural rather than an omission:
Therefore E_SDA_STUCK and E_SCL_STUCK must be separate codes. A master reporting one "bus stuck" code would send its driver to attempt the impossible procedure half the time. That is the single most consequential decision in this block, and it follows directly from those four sentences.
4a. Diagnose before acting, and let SCL decide
The recovery sequencer's first state is a diagnosis, and the test is ordered rather than parallel:
if SCL is low -> escalate. No protocol remedy exists. (§3.1.16)
else if SDA is high -> nothing is stuck. Do nothing.
else -> SDA is stuck and the clock works: pulse.SCL is tested first because a stuck clock makes the procedure impossible whatever SDA is doing.
4b. Nine is a bound, not a quota
§3.1.16 says the holder "should release it some time within those nine clocks". So stopping early is correct, and the number of pulses actually needed is worth reporting — it distinguishes a device that let go on the third pulse from one that never let go at all.
4c. And the procedure ends with a STOP
A freed SDA with no framing leaves every device believing a transfer is in progress. The line is unstuck and the bus is not idle — a different and equally unusable state.
So recovery pulls SDA low and releases it while SCL is high, manufacturing a STOP. Mutation M8 removes exactly this, and it survived every end-state check in the original bench; see §7.
5. The Recovery Procedure
Three pulses, the holder releases, then a STOP
10 cycles6. The Error Manager, in Three Languages
// -----------------------------------------------------------------------------
// i2c_err_mgr.sv
// The error taxonomy a master must report, and the recovery it may attempt.
//
// UM10204 DEFINES NO ERROR CODES. There is no status register in the specification and
// no list of failure modes, so every name below is this design's choice -- and the
// choice matters, because a master that collapses two distinguishable failures into one
// code makes them indistinguishable to software that could otherwise have acted on
// them differently.
//
// THE SIX, and what separates each from its neighbours:
//
// E_ADDR_NACK the address was not acknowledged. Nothing is on the bus at that
// address, OR the device is busy -- Chapter 16.4 shows those are
// indistinguishable, so they share a code, honestly.
// E_DATA_NACK a DATA byte was not acknowledged. Different from the above: the
// device is there and declined this byte, which for an EEPROM means a
// write-protect pin or a read-only location (Chapter 16.1 question 4).
// E_ARB_LOST §3.1.8. Another master won. NOT an error in the usual sense -- the
// bus worked exactly as designed -- and the correct response is to
// retry when the bus is free, not to report a fault.
// E_STRETCH_TO a stretch outlasted this master's policy bound. The target may still
// be conforming (§3.1.6 sets no bound), so this reports a decision the
// master made, not a defect it found.
// E_SDA_STUCK SDA is held low with no transfer in progress. §3.1.16's nine-pulse
// remedy applies.
// E_SCL_STUCK SCL is held low. §3.1.16 offers NO protocol remedy, because the nine
// pulses are themselves pulses on SCL.
//
// WHY THE LAST TWO MUST BE SEPARATE CODES. Chapter 15.4 established that the remedies
// are different -- nine pulses for SDA, a hardware reset or a power cycle for SCL --
// and a master that reported one "bus stuck" code would send its driver to attempt the
// impossible procedure half the time. That is the single most consequential decision in
// this block, and it follows from four sentences of specification.
//
// RECOVERY IS ATTEMPTED FOR EXACTLY ONE OF THEM. §3.1.16, verbatim: "If the data line
// (SDA) is stuck LOW, the master should send nine clock pulses." And for a stuck clock:
// "the preferential procedure is to reset the bus using the HW reset signal". So this
// block clocks for a stuck SDA and escalates for a stuck SCL, and refuses to clock a
// line it cannot drive.
// -----------------------------------------------------------------------------
module i2c_err_mgr #(
parameter int N_PULSES = 9, // §3.1.16: nine, and it is a BOUND not a quota
parameter int N_HALF = 4, // half a recovery clock period, in cycles
parameter int CNT_W = 16
) (
input logic clk,
input logic rst_n,
// Reports from the rest of the master.
input logic addr_nack,
input logic data_nack,
input logic arb_lost,
input logic stretch_timeout,
input logic in_transfer, // a transfer is open, so a low line is normal
// The lines.
input logic scl_in,
input logic sda_in,
// Commands.
input logic start_recovery,
input logic clear,
// Recovery's own bus drive. It owns the lines only while `recovering`.
output logic recovering,
output logic rec_scl_low,
output logic rec_sda_req,
output logic rec_sda_bit,
// The taxonomy. One bit per code, because a master can be in more than one at once
// and an encoded field would have to choose.
output logic [5:0] err,
output logic err_valid,
output logic [CNT_W-1:0] pulses_issued,
output logic recovered, // SDA came free
output logic escalate, // no protocol remedy exists: HW reset or power
output logic [2:0] state
);
localparam integer E_ADDR_NACK = 0, E_DATA_NACK = 1, E_ARB_LOST = 2,
E_STRETCH_TO = 3, E_SDA_STUCK = 4, E_SCL_STUCK = 5;
localparam [2:0] R_IDLE = 3'd0,
R_DIAG = 3'd1, // decide WHICH line is stuck, before acting
R_LOW = 3'd2, // recovery pulse: drive SCL low
R_REL = 3'd3, // release SCL
R_SMP = 3'd4, // sample SDA while SCL is HIGH
R_STOP = 3'd5, // frame the bus before leaving it
R_DONE = 3'd6;
logic [CNT_W-1:0] cnt, pulse_cnt;
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
err <= 6'd0;
err_valid <= 1'b0;
recovering <= 1'b0;
rec_scl_low <= 1'b0;
rec_sda_req <= 1'b0;
rec_sda_bit <= 1'b1;
pulses_issued <= {CNT_W{1'b0}};
recovered <= 1'b0;
escalate <= 1'b0;
state <= R_IDLE;
cnt <= {CNT_W{1'b0}};
pulse_cnt <= {CNT_W{1'b0}};
end else begin
err_valid <= 1'b0;
if (clear) begin
err <= 6'd0;
recovered <= 1'b0;
escalate <= 1'b0;
end else begin
// ---- the taxonomy, latched as it is reported --------------------
if (addr_nack) begin err[E_ADDR_NACK] <= 1'b1; err_valid <= 1'b1; end
if (data_nack) begin err[E_DATA_NACK] <= 1'b1; err_valid <= 1'b1; end
if (arb_lost) begin err[E_ARB_LOST] <= 1'b1; err_valid <= 1'b1; end
if (stretch_timeout) begin err[E_STRETCH_TO] <= 1'b1; err_valid <= 1'b1; end
// A line held low with NO transfer in progress is stuck. The qualifier
// matters: during a transfer both lines are low most of the time, and a
// detector without it would report a stuck bus on every byte.
if (!in_transfer && !recovering) begin
if (!sda_in) begin err[E_SDA_STUCK] <= 1'b1; err_valid <= 1'b1; end
if (!scl_in) begin err[E_SCL_STUCK] <= 1'b1; err_valid <= 1'b1; end
end
end
// ---- recovery ---------------------------------------------------
case (state)
R_IDLE: begin
recovering <= 1'b0;
rec_scl_low <= 1'b0;
rec_sda_req <= 1'b0;
if (start_recovery) begin
recovering <= 1'b1;
recovered <= 1'b0;
escalate <= 1'b0;
pulse_cnt <= {CNT_W{1'b0}};
cnt <= {CNT_W{1'b0}};
state <= R_DIAG;
end
end
R_DIAG: begin
// DIAGNOSE BEFORE ACTING, and let SCL decide. A stuck clock makes the
// nine-pulse procedure impossible whatever SDA is doing, because the
// pulses ARE pulses on SCL -- so this test is ordered, not parallel.
if (!scl_in) begin
escalate <= 1'b1; // no protocol remedy exists. §3.1.16.
state <= R_DONE;
end else if (sda_in) begin
// Nothing is stuck. Running the procedure anyway would clock a bus
// that was merely slow, and corrupt a transfer in progress.
recovered <= 1'b1;
state <= R_DONE;
end else begin
rec_scl_low <= 1'b1;
cnt <= {CNT_W{1'b0}};
state <= R_LOW;
end
end
R_LOW: begin
// A three-phase pulse, because SDA must be sampled while SCL is HIGH.
// A two-phase pulse samples at the instant of release, which reads the
// value from before the holder let go and costs one extra pulse.
rec_scl_low <= 1'b1;
if (cnt + 1 >= N_HALF) begin
rec_scl_low <= 1'b0;
cnt <= {CNT_W{1'b0}};
state <= R_REL;
end else cnt <= cnt + 1'b1;
end
R_REL: begin
// Released. If the clock does not come up, it has failed mid-recovery --
// which is a different fault from the one we started on, and the
// procedure must abandon rather than keep counting pulses that are not
// reaching the wire.
if (!scl_in) begin
err[E_SCL_STUCK] <= 1'b1;
escalate <= 1'b1;
state <= R_DONE;
end else if (cnt + 1 >= N_HALF) begin
cnt <= {CNT_W{1'b0}};
state <= R_SMP;
end else cnt <= cnt + 1'b1;
end
R_SMP: begin
pulse_cnt <= pulse_cnt + 1'b1;
pulses_issued <= pulse_cnt + 1'b1;
if (sda_in) begin
// Free. Nine is a BOUND, not a quota: §3.1.16 says the holder
// "should release it some time within those nine clocks", so
// stopping early is correct and the count is worth reporting.
rec_sda_req <= 1'b1;
rec_sda_bit <= 1'b0; // pull SDA low, to build a STOP
cnt <= {CNT_W{1'b0}};
state <= R_STOP;
end else if (pulse_cnt + 1 >= N_PULSES[CNT_W-1:0]) begin
escalate <= 1'b1; // nine were not enough
state <= R_DONE;
end else begin
rec_scl_low <= 1'b1;
cnt <= {CNT_W{1'b0}};
state <= R_LOW;
end
end
R_STOP: begin
// End with a STOP. A freed SDA with no framing leaves every device on
// the bus believing a transfer is in progress -- the line is unstuck and
// the bus is not idle, which is a different and equally unusable state.
if (cnt + 1 >= N_HALF) begin
rec_sda_bit <= 1'b1; // release SDA while SCL is high: a STOP
recovered <= 1'b1;
cnt <= {CNT_W{1'b0}};
state <= R_DONE;
end else cnt <= cnt + 1'b1;
end
R_DONE: begin
recovering <= 1'b0;
rec_scl_low <= 1'b0;
rec_sda_req <= 1'b0;
rec_sda_bit <= 1'b1;
state <= R_IDLE;
end
default: state <= R_IDLE;
endcase
end
end
endmodule // -----------------------------------------------------------------------------
// i2c_err_mgr.sv
// The error taxonomy a master must report, and the recovery it may attempt.
//
// UM10204 DEFINES NO ERROR CODES. There is no status register in the specification and
// no list of failure modes, so every name below is this design's choice -- and the
// choice matters, because a master that collapses two distinguishable failures into one
// code makes them indistinguishable to software that could otherwise have acted on
// them differently.
//
// THE SIX, and what separates each from its neighbours:
//
// E_ADDR_NACK the address was not acknowledged. Nothing is on the bus at that
// address, OR the device is busy -- Chapter 16.4 shows those are
// indistinguishable, so they share a code, honestly.
// E_DATA_NACK a DATA byte was not acknowledged. Different from the above: the
// device is there and declined this byte, which for an EEPROM means a
// write-protect pin or a read-only location (Chapter 16.1 question 4).
// E_ARB_LOST §3.1.8. Another master won. NOT an error in the usual sense -- the
// bus worked exactly as designed -- and the correct response is to
// retry when the bus is free, not to report a fault.
// E_STRETCH_TO a stretch outlasted this master's policy bound. The target may still
// be conforming (§3.1.6 sets no bound), so this reports a decision the
// master made, not a defect it found.
// E_SDA_STUCK SDA is held low with no transfer in progress. §3.1.16's nine-pulse
// remedy applies.
// E_SCL_STUCK SCL is held low. §3.1.16 offers NO protocol remedy, because the nine
// pulses are themselves pulses on SCL.
//
// WHY THE LAST TWO MUST BE SEPARATE CODES. Chapter 15.4 established that the remedies
// are different -- nine pulses for SDA, a hardware reset or a power cycle for SCL --
// and a master that reported one "bus stuck" code would send its driver to attempt the
// impossible procedure half the time. That is the single most consequential decision in
// this block, and it follows from four sentences of specification.
//
// RECOVERY IS ATTEMPTED FOR EXACTLY ONE OF THEM. §3.1.16, verbatim: "If the data line
// (SDA) is stuck LOW, the master should send nine clock pulses." And for a stuck clock:
// "the preferential procedure is to reset the bus using the HW reset signal". So this
// block clocks for a stuck SDA and escalates for a stuck SCL, and refuses to clock a
// line it cannot drive.
// -----------------------------------------------------------------------------
// (Verilog-2001 -- structurally identical to the SystemVerilog above.)
module i2c_err_mgr #(
parameter N_PULSES = 9, // §3.1.16: nine, and it is a BOUND not a quota
parameter N_HALF = 4, // half a recovery clock period, in cycles
parameter CNT_W = 16
) (
input wire clk,
input wire rst_n,
// Reports from the rest of the master.
input wire addr_nack,
input wire data_nack,
input wire arb_lost,
input wire stretch_timeout,
input wire in_transfer, // a transfer is open, so a low line is normal
// The lines.
input wire scl_in,
input wire sda_in,
// Commands.
input wire start_recovery,
input wire clear,
// Recovery's own bus drive. It owns the lines only while `recovering`.
output reg recovering,
output reg rec_scl_low,
output reg rec_sda_req,
output reg rec_sda_bit,
// The taxonomy. One bit per code, because a master can be in more than one at once
// and an encoded field would have to choose.
output reg [5:0] err,
output reg err_valid,
output reg [CNT_W-1:0] pulses_issued,
output reg recovered, // SDA came free
output reg escalate, // no protocol remedy exists: HW reset or power
output reg [2:0] state
);
localparam integer E_ADDR_NACK = 0, E_DATA_NACK = 1, E_ARB_LOST = 2,
E_STRETCH_TO = 3, E_SDA_STUCK = 4, E_SCL_STUCK = 5;
localparam [2:0] R_IDLE = 3'd0,
R_DIAG = 3'd1, // decide WHICH line is stuck, before acting
R_LOW = 3'd2, // recovery pulse: drive SCL low
R_REL = 3'd3, // release SCL
R_SMP = 3'd4, // sample SDA while SCL is HIGH
R_STOP = 3'd5, // frame the bus before leaving it
R_DONE = 3'd6;
reg [CNT_W-1:0] cnt, pulse_cnt;
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
err <= 6'd0;
err_valid <= 1'b0;
recovering <= 1'b0;
rec_scl_low <= 1'b0;
rec_sda_req <= 1'b0;
rec_sda_bit <= 1'b1;
pulses_issued <= {CNT_W{1'b0}};
recovered <= 1'b0;
escalate <= 1'b0;
state <= R_IDLE;
cnt <= {CNT_W{1'b0}};
pulse_cnt <= {CNT_W{1'b0}};
end else begin
err_valid <= 1'b0;
if (clear) begin
err <= 6'd0;
recovered <= 1'b0;
escalate <= 1'b0;
end else begin
// ---- the taxonomy, latched as it is reported --------------------
if (addr_nack) begin err[E_ADDR_NACK] <= 1'b1; err_valid <= 1'b1; end
if (data_nack) begin err[E_DATA_NACK] <= 1'b1; err_valid <= 1'b1; end
if (arb_lost) begin err[E_ARB_LOST] <= 1'b1; err_valid <= 1'b1; end
if (stretch_timeout) begin err[E_STRETCH_TO] <= 1'b1; err_valid <= 1'b1; end
// A line held low with NO transfer in progress is stuck. The qualifier
// matters: during a transfer both lines are low most of the time, and a
// detector without it would report a stuck bus on every byte.
if (!in_transfer && !recovering) begin
if (!sda_in) begin err[E_SDA_STUCK] <= 1'b1; err_valid <= 1'b1; end
if (!scl_in) begin err[E_SCL_STUCK] <= 1'b1; err_valid <= 1'b1; end
end
end
// ---- recovery ---------------------------------------------------
case (state)
R_IDLE: begin
recovering <= 1'b0;
rec_scl_low <= 1'b0;
rec_sda_req <= 1'b0;
if (start_recovery) begin
recovering <= 1'b1;
recovered <= 1'b0;
escalate <= 1'b0;
pulse_cnt <= {CNT_W{1'b0}};
cnt <= {CNT_W{1'b0}};
state <= R_DIAG;
end
end
R_DIAG: begin
// DIAGNOSE BEFORE ACTING, and let SCL decide. A stuck clock makes the
// nine-pulse procedure impossible whatever SDA is doing, because the
// pulses ARE pulses on SCL -- so this test is ordered, not parallel.
if (!scl_in) begin
escalate <= 1'b1; // no protocol remedy exists. §3.1.16.
state <= R_DONE;
end else if (sda_in) begin
// Nothing is stuck. Running the procedure anyway would clock a bus
// that was merely slow, and corrupt a transfer in progress.
recovered <= 1'b1;
state <= R_DONE;
end else begin
rec_scl_low <= 1'b1;
cnt <= {CNT_W{1'b0}};
state <= R_LOW;
end
end
R_LOW: begin
// A three-phase pulse, because SDA must be sampled while SCL is HIGH.
// A two-phase pulse samples at the instant of release, which reads the
// value from before the holder let go and costs one extra pulse.
rec_scl_low <= 1'b1;
if (cnt + 1 >= N_HALF) begin
rec_scl_low <= 1'b0;
cnt <= {CNT_W{1'b0}};
state <= R_REL;
end else cnt <= cnt + 1'b1;
end
R_REL: begin
// Released. If the clock does not come up, it has failed mid-recovery --
// which is a different fault from the one we started on, and the
// procedure must abandon rather than keep counting pulses that are not
// reaching the wire.
if (!scl_in) begin
err[E_SCL_STUCK] <= 1'b1;
escalate <= 1'b1;
state <= R_DONE;
end else if (cnt + 1 >= N_HALF) begin
cnt <= {CNT_W{1'b0}};
state <= R_SMP;
end else cnt <= cnt + 1'b1;
end
R_SMP: begin
pulse_cnt <= pulse_cnt + 1'b1;
pulses_issued <= pulse_cnt + 1'b1;
if (sda_in) begin
// Free. Nine is a BOUND, not a quota: §3.1.16 says the holder
// "should release it some time within those nine clocks", so
// stopping early is correct and the count is worth reporting.
rec_sda_req <= 1'b1;
rec_sda_bit <= 1'b0; // pull SDA low, to build a STOP
cnt <= {CNT_W{1'b0}};
state <= R_STOP;
end else if (pulse_cnt + 1 >= N_PULSES[CNT_W-1:0]) begin
escalate <= 1'b1; // nine were not enough
state <= R_DONE;
end else begin
rec_scl_low <= 1'b1;
cnt <= {CNT_W{1'b0}};
state <= R_LOW;
end
end
R_STOP: begin
// End with a STOP. A freed SDA with no framing leaves every device on
// the bus believing a transfer is in progress -- the line is unstuck and
// the bus is not idle, which is a different and equally unusable state.
if (cnt + 1 >= N_HALF) begin
rec_sda_bit <= 1'b1; // release SDA while SCL is high: a STOP
recovered <= 1'b1;
cnt <= {CNT_W{1'b0}};
state <= R_DONE;
end else cnt <= cnt + 1'b1;
end
R_DONE: begin
recovering <= 1'b0;
rec_scl_low <= 1'b0;
rec_sda_req <= 1'b0;
rec_sda_bit <= 1'b1;
state <= R_IDLE;
end
default: state <= R_IDLE;
endcase
end
end
endmodule -- ---------------------------------------------------------------------------
-- i2c_err_mgr.vhd
-- The error taxonomy a master must report, and the recovery it may attempt.
-- Behavioural twin of i2c_err_mgr.sv / .v.
--
-- UM10204 DEFINES NO ERROR CODES. There is no status register in the specification and no
-- list of failure modes, so every name below is this design's choice -- and the choice
-- matters, because a master that collapses two distinguishable failures into one code makes
-- them indistinguishable to software that could otherwise have acted on them differently.
--
-- E_ADDR_NACK the address was not acknowledged. Nothing there, OR the device is busy --
-- Chapter 16.4 shows those are indistinguishable, so they share a code.
-- E_DATA_NACK a DATA byte was refused. The device is there and declined this byte.
-- E_ARB_LOST §3.1.8. Another master won -- NOT an error in the usual sense.
-- E_STRETCH_TO a stretch outlasted this master's policy bound. The target may still be
-- conforming, so this reports a decision rather than a defect.
-- E_SDA_STUCK SDA held low with no transfer in progress. §3.1.16's nine pulses apply.
-- E_SCL_STUCK SCL held low. §3.1.16 offers NO protocol remedy, because the nine pulses
-- are themselves pulses on SCL.
--
-- WHY THE LAST TWO MUST BE SEPARATE CODES. Chapter 15.4 established that the remedies are
-- different, and a master that reported one "bus stuck" code would send its driver to
-- attempt the impossible procedure half the time. That is the most consequential decision in
-- this block, and it follows from four sentences of specification.
-- ---------------------------------------------------------------------------
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity i2c_err_mgr is
generic (
N_PULSES : integer := 9; -- §3.1.16: nine, and it is a BOUND not a quota
N_HALF : integer := 4; -- half a recovery clock period, in cycles
CNT_W : integer := 16
);
port (
clk : in std_logic;
rst_n : in std_logic;
addr_nack : in std_logic;
data_nack : in std_logic;
arb_lost : in std_logic;
stretch_timeout : in std_logic;
in_transfer : in std_logic;
scl_in : in std_logic;
sda_in : in std_logic;
start_recovery : in std_logic;
clear : in std_logic;
recovering : out std_logic;
rec_scl_low : out std_logic;
rec_sda_req : out std_logic;
rec_sda_bit : out std_logic;
err : out std_logic_vector(5 downto 0);
err_valid : out std_logic;
pulses_issued : out unsigned(CNT_W-1 downto 0);
recovered : out std_logic;
escalate : out std_logic;
state : out unsigned(2 downto 0)
);
end entity i2c_err_mgr;
architecture rtl of i2c_err_mgr is
constant E_ADDR_NACK : integer := 0;
constant E_DATA_NACK : integer := 1;
constant E_ARB_LOST : integer := 2;
constant E_STRETCH_TO : integer := 3;
constant E_SDA_STUCK : integer := 4;
constant E_SCL_STUCK : integer := 5;
constant R_IDLE : integer := 0;
constant R_DIAG : integer := 1; -- decide WHICH line is stuck, before acting
constant R_LOW : integer := 2;
constant R_REL : integer := 3;
constant R_SMP : integer := 4; -- sample SDA while SCL is HIGH
constant R_STOP : integer := 5;
constant R_DONE : integer := 6;
signal st : integer range 0 to 6 := R_IDLE;
signal cnt : integer range 0 to 65535 := 0;
signal pulse_cnt : unsigned(CNT_W-1 downto 0) := (others => '0');
signal err_i : std_logic_vector(5 downto 0) := (others => '0');
signal rec_i, scl_i, req_i, bit_i, rcv_i, esc_i : std_logic := '0';
signal np : unsigned(CNT_W-1 downto 0) := (others => '0');
begin
recovering <= rec_i;
rec_scl_low <= scl_i;
rec_sda_req <= req_i;
rec_sda_bit <= bit_i;
err <= err_i;
pulses_issued <= np;
recovered <= rcv_i;
escalate <= esc_i;
state <= to_unsigned(st, 3);
process (clk, rst_n)
begin
if rst_n = '0' then
err_i <= (others => '0');
err_valid <= '0';
rec_i <= '0';
scl_i <= '0';
req_i <= '0';
bit_i <= '1';
np <= (others => '0');
rcv_i <= '0';
esc_i <= '0';
st <= R_IDLE;
cnt <= 0;
pulse_cnt <= (others => '0');
elsif rising_edge(clk) then
err_valid <= '0';
if clear = '1' then
err_i <= (others => '0');
rcv_i <= '0';
esc_i <= '0';
else
-- ---- the taxonomy, latched as it is reported --------------------
if addr_nack = '1' then
err_i(E_ADDR_NACK) <= '1'; err_valid <= '1';
end if;
if data_nack = '1' then
err_i(E_DATA_NACK) <= '1'; err_valid <= '1';
end if;
if arb_lost = '1' then
err_i(E_ARB_LOST) <= '1'; err_valid <= '1';
end if;
if stretch_timeout = '1' then
err_i(E_STRETCH_TO) <= '1'; err_valid <= '1';
end if;
-- A line held low with NO transfer in progress is stuck. The qualifier matters:
-- during a transfer both lines are low most of the time, and a detector without
-- it would report a stuck bus on every byte.
if in_transfer = '0' and rec_i = '0' then
if sda_in = '0' then err_i(E_SDA_STUCK) <= '1'; err_valid <= '1'; end if;
if scl_in = '0' then err_i(E_SCL_STUCK) <= '1'; err_valid <= '1'; end if;
end if;
end if;
-- ---- recovery ---------------------------------------------------
case st is
when R_IDLE =>
rec_i <= '0';
scl_i <= '0';
req_i <= '0';
if start_recovery = '1' then
rec_i <= '1';
rcv_i <= '0';
esc_i <= '0';
pulse_cnt <= (others => '0');
cnt <= 0;
st <= R_DIAG;
end if;
when R_DIAG =>
-- DIAGNOSE BEFORE ACTING, and let SCL decide. A stuck clock makes the
-- nine-pulse procedure impossible whatever SDA is doing, because the pulses
-- ARE pulses on SCL -- so this test is ordered, not parallel.
if scl_in = '0' then
esc_i <= '1'; -- no protocol remedy exists. §3.1.16.
st <= R_DONE;
elsif sda_in = '1' then
-- Nothing is stuck. Running the procedure anyway would clock a bus that
-- was merely slow, and corrupt a transfer in progress.
rcv_i <= '1';
st <= R_DONE;
else
scl_i <= '1';
cnt <= 0;
st <= R_LOW;
end if;
when R_LOW =>
-- A three-phase pulse, because SDA must be sampled while SCL is HIGH. A
-- two-phase pulse samples at the instant of release, which reads the value
-- from before the holder let go and costs one extra pulse.
scl_i <= '1';
if cnt + 1 >= N_HALF then
scl_i <= '0';
cnt <= 0;
st <= R_REL;
else
cnt <= cnt + 1;
end if;
when R_REL =>
-- Released. If the clock does not come up it has failed mid-recovery, which
-- is a different fault from the one we started on, and the procedure must
-- abandon rather than keep counting pulses that are not reaching the wire.
if scl_in = '0' then
err_i(E_SCL_STUCK) <= '1';
esc_i <= '1';
st <= R_DONE;
elsif cnt + 1 >= N_HALF then
cnt <= 0;
st <= R_SMP;
else
cnt <= cnt + 1;
end if;
when R_SMP =>
pulse_cnt <= pulse_cnt + 1;
np <= pulse_cnt + 1;
if sda_in = '1' then
-- Free. Nine is a BOUND, not a quota: §3.1.16 says the holder "should
-- release it some time within those nine clocks".
req_i <= '1';
bit_i <= '0'; -- pull SDA low, to build a STOP
cnt <= 0;
st <= R_STOP;
elsif (pulse_cnt + 1) >= to_unsigned(N_PULSES, CNT_W) then
esc_i <= '1'; -- nine were not enough
st <= R_DONE;
else
scl_i <= '1';
cnt <= 0;
st <= R_LOW;
end if;
when R_STOP =>
-- End with a STOP. A freed SDA with no framing leaves every device believing
-- a transfer is in progress -- unstuck, and not idle, which is a different
-- and equally unusable state.
if cnt + 1 >= N_HALF then
bit_i <= '1'; -- release SDA while SCL is high: a STOP
rcv_i <= '1';
cnt <= 0;
st <= R_DONE;
else
cnt <= cnt + 1;
end if;
when R_DONE =>
rec_i <= '0';
scl_i <= '0';
req_i <= '0';
bit_i <= '1';
st <= R_IDLE;
when others =>
st <= R_IDLE;
end case;
end if;
end process;
end architecture rtl;6a. The testbenches
Twelve tests, and note how many of them assert that something does not happen. This block's dangerous failures are all acting when it should not.
| # | Test | Property |
|---|---|---|
| T1 | the four reported failures are four codes | not one "error" bit |
| T2 | arbitration loss and stretch timeout are separate, and neither is a fault | §2a |
| T3 | a low line during a transfer is not a stuck line | both lines are low most of the time |
| T4 | with no transfer open, the same lines are stuck | and the two codes are distinct |
| T5 | a stuck SDA: nine pulses, holder lets go on the third | plus the pulse-width check |
| T6 | the procedure ends with a STOP | checked on the wire; see §7 |
| T7 | a stuck SDA that never lets go: exactly nine, then escalation | |
| T8 | the central test, and it asserts a zero | a stuck SCL gets no pulses at all |
| T9 | diagnose before acting, and SCL decides | with both stuck, the clock wins |
| T10 | a clock that fails mid-recovery | abandons rather than counting phantom pulses |
| T11 | recovery on a healthy bus drives nothing | the most dangerous case |
| T12 | codes clear on request | and a reset manager reports nothing |
The bench instantiates a protocol monitor on the resolved lines, because several of these properties are about what reached the wire rather than what state the block ended in.
`timescale 1ns/1ps
// -----------------------------------------------------------------------------
// i2c_err_mgr_tb.sv
// Independent oracle for i2c_err_mgr.
//
// Recovery is a fault handler, and every assertion about a fault handler passes
// vacuously on a bus that never fails. So the bench INJECTS faults, on a real
// wired-AND bus, by holding lines the way a broken device holds them -- and the
// central tests are the ones that assert a ZERO: no pulses on a stuck clock, and no
// procedure at all on a healthy bus.
// -----------------------------------------------------------------------------
module i2c_err_mgr_tb;
localparam integer NP = 9, NH = 4;
localparam integer E_ADDR = 0, E_DATA = 1, E_ARB = 2, E_STO = 3, E_SDA = 4, E_SCL = 5;
localparam [2:0] R_IDLE = 3'd0, R_DONE = 3'd6;
logic clk = 1'b0, rst_n = 1'b0;
logic addr_nack = 1'b0, data_nack = 1'b0, arb_lost = 1'b0, stretch_to = 1'b0;
logic in_transfer = 1'b0, start_recovery = 1'b0, clear = 1'b0;
// The bench as a faulty device: it can hold either line down.
logic hold_sda = 1'b0, hold_scl = 1'b0;
// And a device that lets go after a set number of recovery pulses.
integer release_after;
integer pulses_seen;
logic recovering, rec_scl_low, rec_sda_req, rec_sda_bit;
logic [5:0] err;
logic err_valid, recovered, escalate;
logic [15:0] pulses;
logic [2:0] rstate;
wire m_sda_low = rec_sda_req ? ~rec_sda_bit : 1'b0;
logic scl, sda;
logic [1:0] scl_in, sda_in, scl_rbl, sda_rbl;
logic [7:0] scl_h, sda_h;
i2c_line_model #(.N_DEV(2)) bus (
.scl_drive_low({hold_scl, rec_scl_low}),
.sda_drive_low({hold_sda, m_sda_low}),
.scl(scl), .sda(sda), .scl_in(scl_in), .sda_in(sda_in),
.scl_released_but_low(scl_rbl), .sda_released_but_low(sda_rbl),
.scl_holders(scl_h), .sda_holders(sda_h));
// A protocol monitor on the resolved lines. Without it the bench can check the STATE
// recovery leaves behind but not what it actually put on the wire -- and "the
// procedure ends with a STOP" is a statement about the wire. A recovery that frees
// SDA and never frames the bus leaves both lines high and `recovering` clear, which
// is indistinguishable from success by any end-state check.
logic mo_start, mo_stop, mo_bitv, mo_byte, mo_ackv, mo_intr, mo_mid, mo_bit, mo_ack;
logic [7:0] mo_byteval;
logic [3:0] mo_bidx;
logic [15:0] mo_nsta, mo_nsto, mo_nbyte, mo_nmid;
i2c_proto_mon #(.CNT_W(16)) mon (
.clk(clk), .rst_n(rst_n), .scl(scl), .sda(sda),
.start_seen(mo_start), .stop_seen(mo_stop), .bit_seen(mo_bit), .bit_val(mo_bitv),
.byte_seen(mo_byte), .byte_val(mo_byteval), .ack_seen(mo_ackv), .ack_val(mo_ack),
.in_transfer(mo_intr), .framing_midbyte(mo_mid), .bit_index(mo_bidx),
.n_starts(mo_nsta), .n_stops(mo_nsto), .n_bytes(mo_nbyte), .n_midbyte(mo_nmid));
// Did recovery EVER drive SCL low? `pulses_issued` counts completed pulses, so a
// procedure that drives the clock and then abandons before sampling reports zero
// pulses while still having touched a line it must not touch. That is the difference
// between "issued no pulses" and "drove nothing", and only the second is the property
// §3.1.16 implies for a stuck clock.
logic ever_drove_scl = 1'b0;
always @(posedge clk) begin
if (!rst_n) ever_drove_scl <= 1'b0;
else if (rec_scl_low) ever_drove_scl <= 1'b1;
end
// The recovery pulses must be LEGAL clock pulses, not merely correctly counted.
// §3.1.16's nine pulses are ordinary SCL pulses on a real bus, so each phase has to
// last at least as long as the configured half period -- a recovery clock that runs
// far faster than the bus's timing is not a remedy, it is noise. Nothing else in this
// bench measures the pulse WIDTH: the pulse count and the final state are both
// unchanged by a procedure that clocks too fast.
integer hi_run = 0;
integer min_hi = 9999;
reg armed = 1'b0; // only measure high runs BETWEEN pulses
always @(posedge clk) begin
if (!rst_n) begin
hi_run <= 0; min_hi <= 9999; armed <= 1'b0;
end else if (recovering) begin
// The high period before the FIRST pulse is not a pulse -- recovery is entered
// with the clock already high -- so measurement arms on the first falling edge.
if (!scl) armed <= 1'b1;
if (scl) hi_run <= hi_run + 1;
else begin
if (armed && hi_run > 0 && hi_run < min_hi) min_hi <= hi_run;
hi_run <= 0;
end
end else begin
hi_run <= 0; armed <= 1'b0;
end
end
i2c_err_mgr #(.N_PULSES(NP), .N_HALF(NH), .CNT_W(16)) dut (
.clk(clk), .rst_n(rst_n),
.addr_nack(addr_nack), .data_nack(data_nack), .arb_lost(arb_lost),
.stretch_timeout(stretch_to), .in_transfer(in_transfer),
.scl_in(scl_in[0]), .sda_in(sda_in[0]),
.start_recovery(start_recovery), .clear(clear),
.recovering(recovering), .rec_scl_low(rec_scl_low),
.rec_sda_req(rec_sda_req), .rec_sda_bit(rec_sda_bit),
.err(err), .err_valid(err_valid), .pulses_issued(pulses),
.recovered(recovered), .escalate(escalate), .state(rstate));
always #5 clk = ~clk;
integer errors = 0;
integer n, k;
// Count the recovery pulses that actually reach the WIRE, and let go after a
// configured number of them -- the way a wedged device eventually does.
logic scl_l;
always @(negedge clk) begin
if (rst_n) begin
if (scl && !scl_l) begin
pulses_seen = pulses_seen + 1;
if (release_after != 0 && pulses_seen >= release_after) hold_sda = 1'b0;
end
scl_l = scl;
end
end
task step; begin @(posedge clk); @(negedge clk); end endtask
task do_reset;
begin
@(negedge clk);
rst_n = 1'b0;
addr_nack = 1'b0; data_nack = 1'b0; arb_lost = 1'b0; stretch_to = 1'b0;
in_transfer = 1'b0; start_recovery = 1'b0; clear = 1'b0;
hold_sda = 1'b0; hold_scl = 1'b0;
release_after = 0; pulses_seen = 0; scl_l = 1'b1;
repeat (3) @(posedge clk);
@(negedge clk); rst_n = 1'b1;
step;
end
endtask
task pulse_recover;
begin
@(negedge clk); start_recovery = 1'b1;
@(posedge clk); @(negedge clk); start_recovery = 1'b0;
end
endtask
task wait_recovery (input integer max_cycles);
begin
n = 0;
while (!(rstate == R_IDLE && !recovering) && n < max_cycles) begin
step; n = n + 1;
end
if (n >= max_cycles) begin
$display(" FAIL wait_recovery: stuck in state %0d", rstate);
errors = errors + 1;
end
end
endtask
task ck_int (input [200*8:1] what, input integer g, input integer e);
begin
if (g !== e) begin
$display(" FAIL %0s: got %0d expected %0d", what, g, e);
errors = errors + 1;
end
end
endtask
task ck_bit (input [200*8:1] what, input g, input e);
begin
if (g !== e) begin
$display(" FAIL %0s: got %0b expected %0b", what, g, e);
errors = errors + 1;
end
end
endtask
initial begin
$display("=== i2c_err_mgr: six codes, and a remedy for exactly one of them ===");
// ----------------------------------------------------------------
// T1. THE FOUR REPORTED FAILURES ARE FOUR CODES. An address NACK and a data NACK
// are different facts -- nothing there versus there and declining -- and a
// master that merged them would make them indistinguishable to software.
// ----------------------------------------------------------------
do_reset;
@(negedge clk); in_transfer = 1'b1; addr_nack = 1'b1;
@(posedge clk); @(negedge clk); addr_nack = 1'b0; step;
$display("T1 an address NACK and a data NACK are different codes");
ck_bit("T1 the address NACK is latched", err[E_ADDR], 1'b1);
ck_bit("T1 and the data NACK is not", err[E_DATA], 1'b0);
@(negedge clk); data_nack = 1'b1;
@(posedge clk); @(negedge clk); data_nack = 1'b0; step;
ck_bit("T1 now the data NACK is latched too", err[E_DATA], 1'b1);
ck_bit("T1 and both are visible at once", err[E_ADDR] & err[E_DATA], 1'b1);
// ----------------------------------------------------------------
// T2. Arbitration loss and a stretch timeout are also separate, and neither is a
// defect: §3.1.8's loser lost a fair contest, and §3.1.6 gives stretching no
// bound at all, so a timeout reports this master's policy.
// ----------------------------------------------------------------
do_reset;
@(negedge clk); in_transfer = 1'b1; arb_lost = 1'b1;
@(posedge clk); @(negedge clk); arb_lost = 1'b0;
@(negedge clk); stretch_to = 1'b1;
@(posedge clk); @(negedge clk); stretch_to = 1'b0; step;
$display("T2 arbitration loss and a stretch timeout are separate, and neither is a fault");
ck_bit("T2 arbitration loss latched", err[E_ARB], 1'b1);
ck_bit("T2 stretch timeout latched", err[E_STO], 1'b1);
ck_int("T2 and only those two", err, (1 << E_ARB) | (1 << E_STO));
// ----------------------------------------------------------------
// T3. A LOW LINE DURING A TRANSFER IS NOT A STUCK LINE. Both lines are low most
// of the time while clocking, and a detector without the qualifier would
// report a stuck bus on every byte.
// ----------------------------------------------------------------
do_reset;
@(negedge clk); in_transfer = 1'b1; hold_sda = 1'b1; hold_scl = 1'b1;
for (k = 0; k < 10; k = k + 1) step;
$display("T3 a low line during a transfer is not a stuck line");
ck_bit("T3 SDA is low", sda, 1'b0);
ck_bit("T3 SCL is low", scl, 1'b0);
ck_bit("T3 but no stuck-SDA report", err[E_SDA], 1'b0);
ck_bit("T3 and no stuck-SCL report", err[E_SCL], 1'b0);
// ----------------------------------------------------------------
// T4. WITH NO TRANSFER OPEN, THE SAME LINES ARE STUCK. And the two codes are
// separate, which is the block's most consequential decision: §3.1.16 gives
// the two lines DIFFERENT remedies, so one merged code would send a driver
// to attempt the impossible procedure half the time.
// ----------------------------------------------------------------
@(negedge clk); in_transfer = 1'b0;
for (k = 0; k < 6; k = k + 1) step;
$display("T4 with no transfer open, the same lines are stuck -- as two codes");
ck_bit("T4 stuck SDA", err[E_SDA], 1'b1);
ck_bit("T4 stuck SCL", err[E_SCL], 1'b1);
// ----------------------------------------------------------------
// T5. A STUCK SDA: nine pulses, and the holder lets go on the third. §3.1.16 says
// the device "should release it some time within those nine clocks", so nine
// is a BOUND and stopping early is correct.
// ----------------------------------------------------------------
do_reset;
@(negedge clk); hold_sda = 1'b1; release_after = 3;
step;
pulse_recover;
wait_recovery(1000);
$display("T5 a stuck SDA is freed by clocking, and nine is a bound not a quota");
ck_bit("T5 recovered", recovered, 1'b1);
ck_bit("T5 no escalation needed", escalate, 1'b0);
ck_int("T5 it took three pulses, not nine", pulses, 3);
ck_bit("T5 SDA is free", sda, 1'b1);
// ----------------------------------------------------------------
// T6. AND THE PROCEDURE ENDS WITH A STOP. A freed SDA with no framing leaves
// every device believing a transfer is in progress: unstuck, and not idle.
// ----------------------------------------------------------------
$display("T6 recovery ends with a STOP, leaving the bus idle rather than unstuck");
ck_bit("T6 SCL released", rec_scl_low, 1'b0);
ck_bit("T6 SDA released", m_sda_low, 1'b0);
ck_bit("T6 both lines high", scl & sda, 1'b1);
ck_bit("T6 recovery is finished", recovering, 1'b0);
// And the STOP must have actually appeared on the bus. The four checks above
// describe the state recovery left behind; a procedure that freed SDA and never
// framed the bus leaves exactly the same state, with every device still believing
// a transfer is open.
if (mo_nsto < 1) begin
$display(" FAIL T6 recovery freed SDA but never framed the bus");
errors = errors + 1;
end
// ----------------------------------------------------------------
// T7. A STUCK SDA THAT NEVER LETS GO: exactly nine pulses, then escalation.
// ----------------------------------------------------------------
do_reset;
@(negedge clk); hold_sda = 1'b1; release_after = 0; // never releases
step;
pulse_recover;
wait_recovery(2000);
$display("T7 a holder that never lets go: exactly nine pulses, then escalate");
ck_int("T7 exactly nine pulses", pulses, NP);
ck_int("T7 and every high phase was at least N_HALF cycles",
(min_hi >= NH) ? 1 : 0, 1);
ck_bit("T7 not recovered", recovered, 1'b0);
ck_bit("T7 escalated", escalate, 1'b1);
// ----------------------------------------------------------------
// T8. THE CENTRAL TEST, AND IT ASSERTS A ZERO. A stuck SCL gets NO pulses,
// because the nine pulses ARE pulses on SCL and cannot be issued on a line
// something else is holding. §3.1.16 offers no protocol remedy at all.
// ----------------------------------------------------------------
do_reset;
@(negedge clk); hold_scl = 1'b1; hold_sda = 1'b1;
step;
pulse_recover;
wait_recovery(1000);
$display("T8 a stuck SCL gets no pulses at all, because none could reach the wire");
ck_int("T8 zero pulses issued", pulses, 0);
// Stronger than the pulse count: the procedure must not have driven the clock at
// all. A design that starts pulsing and abandons on the first release also reports
// zero completed pulses, so the counter alone cannot tell the two apart.
ck_bit("T8 and recovery never drove SCL at all", ever_drove_scl, 1'b0);
ck_bit("T8 escalated immediately", escalate, 1'b1);
ck_bit("T8 not recovered", recovered, 1'b0);
ck_bit("T8 and the stuck clock was diagnosed", err[E_SCL], 1'b1);
// ----------------------------------------------------------------
// T9. DIAGNOSE BEFORE ACTING, and SCL decides. With both lines stuck, the clock
// is what determines the outcome -- a stuck clock makes the procedure
// impossible whatever SDA is doing.
// ----------------------------------------------------------------
$display("T9 with both lines stuck, the clock decides the outcome");
ck_bit("T9 both were diagnosed", err[E_SDA] & err[E_SCL], 1'b1);
ck_int("T9 and still no pulses were attempted", pulses, 0);
// ----------------------------------------------------------------
// T10. A CLOCK THAT FAILS MID-RECOVERY. The procedure starts legitimately on a
// stuck SDA and the clock is then seized. Carrying on would be issuing
// pulses that never reach the wire and then reporting nine of them.
// ----------------------------------------------------------------
do_reset;
@(negedge clk); hold_sda = 1'b1; release_after = 0;
step;
pulse_recover;
// Let two pulses happen, then seize the clock.
n = 0;
while (pulses_seen < 2 && n < 500) begin step; n = n + 1; end
// One extra step, so the three language variants stay at the same finish time. The
// VHDL twin needs it because `pulses_seen` there is a SIGNAL assigned by the counting
// process and read by the stimulus in the same delta, which yields its previous value
// and runs the loop once more. Here it is a variable-like `integer` assigned blockingly
// and visible at once. Keeping the step in all three is cheaper than making the three
// benches differ.
step;
@(negedge clk); hold_scl = 1'b1;
wait_recovery(2000);
$display("T10 a clock seized mid-recovery abandons the procedure rather than lying");
ck_bit("T10 escalated", escalate, 1'b1);
ck_bit("T10 the stuck clock was recorded", err[E_SCL], 1'b1);
ck_bit("T10 not recovered", recovered, 1'b0);
if (pulses >= NP) begin
$display(" FAIL T10 reported %0d pulses after the clock failed", pulses);
errors = errors + 1;
end
// ----------------------------------------------------------------
// T11. RECOVERY ON A HEALTHY BUS DRIVES NOTHING. This is the most dangerous case
// to get wrong: running the procedure on a bus that was merely slow would
// clock a transfer to pieces and then report success.
// ----------------------------------------------------------------
do_reset;
pulse_recover;
wait_recovery(1000);
$display("T11 recovery on a healthy bus drives nothing and reports it healthy");
ck_int("T11 no pulses", pulses, 0);
ck_bit("T11 reported healthy", recovered, 1'b1);
ck_bit("T11 no escalation", escalate, 1'b0);
ck_bit("T11 nothing was driven", rec_scl_low | m_sda_low, 1'b0);
ck_bit("T11 both lines still high", scl & sda, 1'b1);
// ----------------------------------------------------------------
// T12. Codes are cleared on request, and a reset manager reports nothing about a
// bus it has not seen.
// ----------------------------------------------------------------
do_reset;
@(negedge clk); in_transfer = 1'b1; addr_nack = 1'b1; arb_lost = 1'b1;
@(posedge clk); @(negedge clk); addr_nack = 1'b0; arb_lost = 1'b0; step;
ck_int("T12 two codes latched", err, (1 << E_ADDR) | (1 << E_ARB));
@(negedge clk); clear = 1'b1; step;
@(negedge clk); clear = 1'b0; step;
$display("T12 codes clear on request, and a reset manager reports nothing");
ck_int("T12 cleared", err, 0);
@(negedge clk); rst_n = 1'b0; step;
ck_int("T12 reset clears them too", err, 0);
ck_int("T12 and the pulse count", pulses, 0);
ck_bit("T12 nothing driven out of reset", rec_scl_low | rec_sda_req, 1'b0);
if (errors == 0)
$display("=== i2c_err_mgr: ALL CHECKS PASSED ===");
else
$display("=== i2c_err_mgr: %0d CHECK(S) FAILED ===", errors);
$finish;
end
endmodule `timescale 1ns/1ps
// -----------------------------------------------------------------------------
// i2c_err_mgr_tb.sv
// Independent oracle for i2c_err_mgr.
//
// Recovery is a fault handler, and every assertion about a fault handler passes
// vacuously on a bus that never fails. So the bench INJECTS faults, on a real
// wired-AND bus, by holding lines the way a broken device holds them -- and the
// central tests are the ones that assert a ZERO: no pulses on a stuck clock, and no
// procedure at all on a healthy bus.
// -----------------------------------------------------------------------------
// (Verilog-2001 -- structurally identical to the SystemVerilog above.)
module i2c_err_mgr_tb;
localparam integer NP = 9, NH = 4;
localparam integer E_ADDR = 0, E_DATA = 1, E_ARB = 2, E_STO = 3, E_SDA = 4, E_SCL = 5;
localparam [2:0] R_IDLE = 3'd0, R_DONE = 3'd6;
reg clk = 1'b0, rst_n = 1'b0;
reg addr_nack = 1'b0, data_nack = 1'b0, arb_lost = 1'b0, stretch_to = 1'b0;
reg in_transfer = 1'b0, start_recovery = 1'b0, clear = 1'b0;
// The bench as a faulty device: it can hold either line down.
reg hold_sda = 1'b0, hold_scl = 1'b0;
// And a device that lets go after a set number of recovery pulses.
integer release_after;
integer pulses_seen;
wire recovering, rec_scl_low, rec_sda_req, rec_sda_bit;
wire [5:0] err;
wire err_valid, recovered, escalate;
wire [15:0] pulses;
wire [2:0] rstate;
wire m_sda_low = rec_sda_req ? ~rec_sda_bit : 1'b0;
wire scl, sda;
wire [1:0] scl_in, sda_in, scl_rbl, sda_rbl;
wire [7:0] scl_h, sda_h;
i2c_line_model #(.N_DEV(2)) bus (
.scl_drive_low({hold_scl, rec_scl_low}),
.sda_drive_low({hold_sda, m_sda_low}),
.scl(scl), .sda(sda), .scl_in(scl_in), .sda_in(sda_in),
.scl_released_but_low(scl_rbl), .sda_released_but_low(sda_rbl),
.scl_holders(scl_h), .sda_holders(sda_h));
// A protocol monitor on the resolved lines. Without it the bench can check the STATE
// recovery leaves behind but not what it actually put on the wire -- and "the
// procedure ends with a STOP" is a statement about the wire. A recovery that frees
// SDA and never frames the bus leaves both lines high and `recovering` clear, which
// is indistinguishable from success by any end-state check.
wire mo_start, mo_stop, mo_bitv, mo_byte, mo_ackv, mo_intr, mo_mid, mo_bit, mo_ack;
wire [7:0] mo_byteval;
wire [3:0] mo_bidx;
wire [15:0] mo_nsta, mo_nsto, mo_nbyte, mo_nmid;
i2c_proto_mon #(.CNT_W(16)) mon (
.clk(clk), .rst_n(rst_n), .scl(scl), .sda(sda),
.start_seen(mo_start), .stop_seen(mo_stop), .bit_seen(mo_bit), .bit_val(mo_bitv),
.byte_seen(mo_byte), .byte_val(mo_byteval), .ack_seen(mo_ackv), .ack_val(mo_ack),
.in_transfer(mo_intr), .framing_midbyte(mo_mid), .bit_index(mo_bidx),
.n_starts(mo_nsta), .n_stops(mo_nsto), .n_bytes(mo_nbyte), .n_midbyte(mo_nmid));
// Did recovery EVER drive SCL low? `pulses_issued` counts completed pulses, so a
// procedure that drives the clock and then abandons before sampling reports zero
// pulses while still having touched a line it must not touch. That is the difference
// between "issued no pulses" and "drove nothing", and only the second is the property
// §3.1.16 implies for a stuck clock.
reg ever_drove_scl;
always @(posedge clk) begin
if (!rst_n) ever_drove_scl <= 1'b0;
else if (rec_scl_low) ever_drove_scl <= 1'b1;
end
// The recovery pulses must be LEGAL clock pulses, not merely correctly counted.
// §3.1.16's nine pulses are ordinary SCL pulses on a real bus, so each phase has to
// last at least as long as the configured half period -- a recovery clock that runs
// far faster than the bus's timing is not a remedy, it is noise. Nothing else in this
// bench measures the pulse WIDTH: the pulse count and the final state are both
// unchanged by a procedure that clocks too fast.
integer hi_run;
integer min_hi;
reg armed; // only measure high runs BETWEEN pulses
always @(posedge clk) begin
if (!rst_n) begin
hi_run <= 0; min_hi <= 9999; armed <= 1'b0;
end else if (recovering) begin
// The high period before the FIRST pulse is not a pulse -- recovery is entered
// with the clock already high -- so measurement arms on the first falling edge.
if (!scl) armed <= 1'b1;
if (scl) hi_run <= hi_run + 1;
else begin
if (armed && hi_run > 0 && hi_run < min_hi) min_hi <= hi_run;
hi_run <= 0;
end
end else begin
hi_run <= 0; armed <= 1'b0;
end
end
i2c_err_mgr #(.N_PULSES(NP), .N_HALF(NH), .CNT_W(16)) dut (
.clk(clk), .rst_n(rst_n),
.addr_nack(addr_nack), .data_nack(data_nack), .arb_lost(arb_lost),
.stretch_timeout(stretch_to), .in_transfer(in_transfer),
.scl_in(scl_in[0]), .sda_in(sda_in[0]),
.start_recovery(start_recovery), .clear(clear),
.recovering(recovering), .rec_scl_low(rec_scl_low),
.rec_sda_req(rec_sda_req), .rec_sda_bit(rec_sda_bit),
.err(err), .err_valid(err_valid), .pulses_issued(pulses),
.recovered(recovered), .escalate(escalate), .state(rstate));
always #5 clk = ~clk;
integer errors = 0;
integer n, k;
// Count the recovery pulses that actually reach the WIRE, and let go after a
// configured number of them -- the way a wedged device eventually does.
reg scl_l;
always @(negedge clk) begin
if (rst_n) begin
if (scl && !scl_l) begin
pulses_seen = pulses_seen + 1;
if (release_after != 0 && pulses_seen >= release_after) hold_sda = 1'b0;
end
scl_l = scl;
end
end
task step; begin @(posedge clk); @(negedge clk); end endtask
task do_reset;
begin
@(negedge clk);
rst_n = 1'b0;
addr_nack = 1'b0; data_nack = 1'b0; arb_lost = 1'b0; stretch_to = 1'b0;
in_transfer = 1'b0; start_recovery = 1'b0; clear = 1'b0;
hold_sda = 1'b0; hold_scl = 1'b0;
release_after = 0; pulses_seen = 0; scl_l = 1'b1;
repeat (3) @(posedge clk);
@(negedge clk); rst_n = 1'b1;
step;
end
endtask
task pulse_recover;
begin
@(negedge clk); start_recovery = 1'b1;
@(posedge clk); @(negedge clk); start_recovery = 1'b0;
end
endtask
task wait_recovery (input integer max_cycles);
begin
n = 0;
while (!(rstate == R_IDLE && !recovering) && n < max_cycles) begin
step; n = n + 1;
end
if (n >= max_cycles) begin
$display(" FAIL wait_recovery: stuck in state %0d", rstate);
errors = errors + 1;
end
end
endtask
task ck_int (input [200*8:1] what, input integer g, input integer e);
begin
if (g !== e) begin
$display(" FAIL %0s: got %0d expected %0d", what, g, e);
errors = errors + 1;
end
end
endtask
task ck_bit (input [200*8:1] what, input g, input e);
begin
if (g !== e) begin
$display(" FAIL %0s: got %0b expected %0b", what, g, e);
errors = errors + 1;
end
end
endtask
initial begin
$display("=== i2c_err_mgr: six codes, and a remedy for exactly one of them ===");
// ----------------------------------------------------------------
// T1. THE FOUR REPORTED FAILURES ARE FOUR CODES. An address NACK and a data NACK
// are different facts -- nothing there versus there and declining -- and a
// master that merged them would make them indistinguishable to software.
// ----------------------------------------------------------------
do_reset;
@(negedge clk); in_transfer = 1'b1; addr_nack = 1'b1;
@(posedge clk); @(negedge clk); addr_nack = 1'b0; step;
$display("T1 an address NACK and a data NACK are different codes");
ck_bit("T1 the address NACK is latched", err[E_ADDR], 1'b1);
ck_bit("T1 and the data NACK is not", err[E_DATA], 1'b0);
@(negedge clk); data_nack = 1'b1;
@(posedge clk); @(negedge clk); data_nack = 1'b0; step;
ck_bit("T1 now the data NACK is latched too", err[E_DATA], 1'b1);
ck_bit("T1 and both are visible at once", err[E_ADDR] & err[E_DATA], 1'b1);
// ----------------------------------------------------------------
// T2. Arbitration loss and a stretch timeout are also separate, and neither is a
// defect: §3.1.8's loser lost a fair contest, and §3.1.6 gives stretching no
// bound at all, so a timeout reports this master's policy.
// ----------------------------------------------------------------
do_reset;
@(negedge clk); in_transfer = 1'b1; arb_lost = 1'b1;
@(posedge clk); @(negedge clk); arb_lost = 1'b0;
@(negedge clk); stretch_to = 1'b1;
@(posedge clk); @(negedge clk); stretch_to = 1'b0; step;
$display("T2 arbitration loss and a stretch timeout are separate, and neither is a fault");
ck_bit("T2 arbitration loss latched", err[E_ARB], 1'b1);
ck_bit("T2 stretch timeout latched", err[E_STO], 1'b1);
ck_int("T2 and only those two", err, (1 << E_ARB) | (1 << E_STO));
// ----------------------------------------------------------------
// T3. A LOW LINE DURING A TRANSFER IS NOT A STUCK LINE. Both lines are low most
// of the time while clocking, and a detector without the qualifier would
// report a stuck bus on every byte.
// ----------------------------------------------------------------
do_reset;
@(negedge clk); in_transfer = 1'b1; hold_sda = 1'b1; hold_scl = 1'b1;
for (k = 0; k < 10; k = k + 1) step;
$display("T3 a low line during a transfer is not a stuck line");
ck_bit("T3 SDA is low", sda, 1'b0);
ck_bit("T3 SCL is low", scl, 1'b0);
ck_bit("T3 but no stuck-SDA report", err[E_SDA], 1'b0);
ck_bit("T3 and no stuck-SCL report", err[E_SCL], 1'b0);
// ----------------------------------------------------------------
// T4. WITH NO TRANSFER OPEN, THE SAME LINES ARE STUCK. And the two codes are
// separate, which is the block's most consequential decision: §3.1.16 gives
// the two lines DIFFERENT remedies, so one merged code would send a driver
// to attempt the impossible procedure half the time.
// ----------------------------------------------------------------
@(negedge clk); in_transfer = 1'b0;
for (k = 0; k < 6; k = k + 1) step;
$display("T4 with no transfer open, the same lines are stuck -- as two codes");
ck_bit("T4 stuck SDA", err[E_SDA], 1'b1);
ck_bit("T4 stuck SCL", err[E_SCL], 1'b1);
// ----------------------------------------------------------------
// T5. A STUCK SDA: nine pulses, and the holder lets go on the third. §3.1.16 says
// the device "should release it some time within those nine clocks", so nine
// is a BOUND and stopping early is correct.
// ----------------------------------------------------------------
do_reset;
@(negedge clk); hold_sda = 1'b1; release_after = 3;
step;
pulse_recover;
wait_recovery(1000);
$display("T5 a stuck SDA is freed by clocking, and nine is a bound not a quota");
ck_bit("T5 recovered", recovered, 1'b1);
ck_bit("T5 no escalation needed", escalate, 1'b0);
ck_int("T5 it took three pulses, not nine", pulses, 3);
ck_bit("T5 SDA is free", sda, 1'b1);
// ----------------------------------------------------------------
// T6. AND THE PROCEDURE ENDS WITH A STOP. A freed SDA with no framing leaves
// every device believing a transfer is in progress: unstuck, and not idle.
// ----------------------------------------------------------------
$display("T6 recovery ends with a STOP, leaving the bus idle rather than unstuck");
ck_bit("T6 SCL released", rec_scl_low, 1'b0);
ck_bit("T6 SDA released", m_sda_low, 1'b0);
ck_bit("T6 both lines high", scl & sda, 1'b1);
ck_bit("T6 recovery is finished", recovering, 1'b0);
// And the STOP must have actually appeared on the bus. The checks above describe
// the state recovery left behind; a procedure that freed SDA and never framed the
// bus leaves the same state with every device still believing a transfer is open.
if (mo_nsto < 1) begin
$display(" FAIL T6 recovery freed SDA but never framed the bus");
errors = errors + 1;
end
// ----------------------------------------------------------------
// T7. A STUCK SDA THAT NEVER LETS GO: exactly nine pulses, then escalation.
// ----------------------------------------------------------------
do_reset;
@(negedge clk); hold_sda = 1'b1; release_after = 0; // never releases
step;
pulse_recover;
wait_recovery(2000);
$display("T7 a holder that never lets go: exactly nine pulses, then escalate");
ck_int("T7 exactly nine pulses", pulses, NP);
ck_int("T7 and every high phase was at least N_HALF cycles",
(min_hi >= NH) ? 1 : 0, 1);
ck_bit("T7 not recovered", recovered, 1'b0);
ck_bit("T7 escalated", escalate, 1'b1);
// ----------------------------------------------------------------
// T8. THE CENTRAL TEST, AND IT ASSERTS A ZERO. A stuck SCL gets NO pulses,
// because the nine pulses ARE pulses on SCL and cannot be issued on a line
// something else is holding. §3.1.16 offers no protocol remedy at all.
// ----------------------------------------------------------------
do_reset;
@(negedge clk); hold_scl = 1'b1; hold_sda = 1'b1;
step;
pulse_recover;
wait_recovery(1000);
$display("T8 a stuck SCL gets no pulses at all, because none could reach the wire");
ck_int("T8 zero pulses issued", pulses, 0);
// Stronger than the pulse count: the procedure must not have driven the clock at
// all. A design that starts pulsing and abandons on the first release also reports
// zero completed pulses, so the counter alone cannot tell the two apart.
ck_bit("T8 and recovery never drove SCL at all", ever_drove_scl, 1'b0);
ck_bit("T8 escalated immediately", escalate, 1'b1);
ck_bit("T8 not recovered", recovered, 1'b0);
ck_bit("T8 and the stuck clock was diagnosed", err[E_SCL], 1'b1);
// ----------------------------------------------------------------
// T9. DIAGNOSE BEFORE ACTING, and SCL decides. With both lines stuck, the clock
// is what determines the outcome -- a stuck clock makes the procedure
// impossible whatever SDA is doing.
// ----------------------------------------------------------------
$display("T9 with both lines stuck, the clock decides the outcome");
ck_bit("T9 both were diagnosed", err[E_SDA] & err[E_SCL], 1'b1);
ck_int("T9 and still no pulses were attempted", pulses, 0);
// ----------------------------------------------------------------
// T10. A CLOCK THAT FAILS MID-RECOVERY. The procedure starts legitimately on a
// stuck SDA and the clock is then seized. Carrying on would be issuing
// pulses that never reach the wire and then reporting nine of them.
// ----------------------------------------------------------------
do_reset;
@(negedge clk); hold_sda = 1'b1; release_after = 0;
step;
pulse_recover;
// Let two pulses happen, then seize the clock.
n = 0;
while (pulses_seen < 2 && n < 500) begin step; n = n + 1; end
// One extra step, so the three language variants stay at the same finish time. The
// VHDL twin needs it because `pulses_seen` there is a SIGNAL assigned by the counting
// process and read by the stimulus in the same delta, which yields its previous value
// and runs the loop once more. Here it is a variable-like `integer` assigned blockingly
// and visible at once. Keeping the step in all three is cheaper than making the three
// benches differ.
step;
@(negedge clk); hold_scl = 1'b1;
wait_recovery(2000);
$display("T10 a clock seized mid-recovery abandons the procedure rather than lying");
ck_bit("T10 escalated", escalate, 1'b1);
ck_bit("T10 the stuck clock was recorded", err[E_SCL], 1'b1);
ck_bit("T10 not recovered", recovered, 1'b0);
if (pulses >= NP) begin
$display(" FAIL T10 reported %0d pulses after the clock failed", pulses);
errors = errors + 1;
end
// ----------------------------------------------------------------
// T11. RECOVERY ON A HEALTHY BUS DRIVES NOTHING. This is the most dangerous case
// to get wrong: running the procedure on a bus that was merely slow would
// clock a transfer to pieces and then report success.
// ----------------------------------------------------------------
do_reset;
pulse_recover;
wait_recovery(1000);
$display("T11 recovery on a healthy bus drives nothing and reports it healthy");
ck_int("T11 no pulses", pulses, 0);
ck_bit("T11 reported healthy", recovered, 1'b1);
ck_bit("T11 no escalation", escalate, 1'b0);
ck_bit("T11 nothing was driven", rec_scl_low | m_sda_low, 1'b0);
ck_bit("T11 both lines still high", scl & sda, 1'b1);
// ----------------------------------------------------------------
// T12. Codes are cleared on request, and a reset manager reports nothing about a
// bus it has not seen.
// ----------------------------------------------------------------
do_reset;
@(negedge clk); in_transfer = 1'b1; addr_nack = 1'b1; arb_lost = 1'b1;
@(posedge clk); @(negedge clk); addr_nack = 1'b0; arb_lost = 1'b0; step;
ck_int("T12 two codes latched", err, (1 << E_ADDR) | (1 << E_ARB));
@(negedge clk); clear = 1'b1; step;
@(negedge clk); clear = 1'b0; step;
$display("T12 codes clear on request, and a reset manager reports nothing");
ck_int("T12 cleared", err, 0);
@(negedge clk); rst_n = 1'b0; step;
ck_int("T12 reset clears them too", err, 0);
ck_int("T12 and the pulse count", pulses, 0);
ck_bit("T12 nothing driven out of reset", rec_scl_low | rec_sda_req, 1'b0);
if (errors == 0)
$display("=== i2c_err_mgr: ALL CHECKS PASSED ===");
else
$display("=== i2c_err_mgr: %0d CHECK(S) FAILED ===", errors);
$finish;
end
endmodule -- ---------------------------------------------------------------------------
-- i2c_err_mgr_tb.vhd
-- Independent oracle for i2c_err_mgr. Behavioural twin of the SV and Verilog benches.
--
-- Recovery is a fault handler, and every assertion about a fault handler passes vacuously on
-- a bus that never fails. So the bench INJECTS faults, on a real wired-AND bus, by holding
-- lines the way a broken device holds them -- and the central tests are the ones that assert
-- a ZERO: no pulses on a stuck clock, and no procedure at all on a healthy bus.
-- ---------------------------------------------------------------------------
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity i2c_err_mgr_tb is
end entity i2c_err_mgr_tb;
architecture sim of i2c_err_mgr_tb is
constant TCLK : time := 10 ns;
constant NP : integer := 9;
constant NH : integer := 4;
constant E_ADDR : integer := 0;
constant E_DATA : integer := 1;
constant E_ARB : integer := 2;
constant E_STO : integer := 3;
constant E_SDA : integer := 4;
constant E_SCL : integer := 5;
constant R_IDLE : integer := 0;
signal clk, rst_n : std_logic := '0';
signal addr_nack, data_nack, arb_lost, stretch_to : std_logic := '0';
signal in_transfer, start_recovery, clear : std_logic := '0';
signal hold_scl : std_logic := '0';
-- SINGLE OWNER PER SIGNAL. In the SystemVerilog bench `hold_sda` is written by both the
-- stimulus block and the pulse counter, which Verilog permits because a `reg` has no
-- resolution. VHDL's `std_logic` IS a resolved type, so two processes driving it give 'X'
-- -- silently, with no elaboration error, unlike an `integer` where a second driver is
-- fatal. The symptom is a fault injector that appears to inject nothing.
--
-- So the ownership is split: the stimulus REQUESTS a hold, the pulse counter owns the
-- line and decides when to let go.
signal want_hold_sda : std_logic := '0'; -- driven by the stimulus only
signal hold_sda : std_logic; -- driven by the pulse counter only
signal let_go : std_logic := '0';
signal recovering, rec_scl_low, rec_sda_req, rec_sda_bit : std_logic;
signal err : std_logic_vector(5 downto 0);
signal err_valid, recovered, escalate : std_logic;
signal pulses : unsigned(15 downto 0);
signal rstate : unsigned(2 downto 0);
signal m_sda_low : std_logic;
signal scl_drv, sda_drv : std_logic_vector(1 downto 0);
signal scl, sda : std_logic;
signal scl_in, sda_in, scl_rbl, sda_rbl : std_logic_vector(1 downto 0);
signal scl_h, sda_h : unsigned(7 downto 0);
signal release_after : integer := 0;
signal pulses_seen : integer := 0;
signal halt : boolean := false;
-- A protocol monitor on the resolved lines. Without it the bench can check the STATE
-- recovery leaves behind but not what it actually put on the wire -- and "the procedure
-- ends with a STOP" is a statement about the wire. A recovery that frees SDA and never
-- frames the bus leaves both lines high and recovering clear, which is
-- indistinguishable from success by any end-state check.
signal mo_start, mo_stop, mo_bit, mo_bitv, mo_byte, mo_ackv, mo_ack : std_logic;
signal mo_intr, mo_mid : std_logic;
signal mo_byteval : std_logic_vector(7 downto 0);
signal mo_bidx : unsigned(3 downto 0);
signal mo_nsta, mo_nsto, mo_nbyte, mo_nmid : unsigned(15 downto 0);
-- Did recovery EVER drive SCL low? pulses_issued counts COMPLETED pulses, so a
-- procedure that drives the clock and abandons before sampling reports zero pulses
-- while still having touched a line it must not touch.
signal ever_drove_scl : std_logic := '0';
-- The recovery pulses must be LEGAL clock pulses, not merely correctly counted.
-- §3.1.16's nine pulses are ordinary SCL pulses on a real bus, so each phase has to
-- last at least the configured half period. Nothing else in this bench measures the
-- pulse WIDTH: the count and the final state are unchanged by clocking too fast.
signal hi_run : integer := 0;
signal min_hi : integer := 9999;
signal armed : boolean := false;
begin
m_sda_low <= (not rec_sda_bit) when rec_sda_req = '1' else '0';
scl_drv <= hold_scl & rec_scl_low;
sda_drv <= hold_sda & m_sda_low;
bus_m : entity work.i2c_line_model
generic map (N_DEV => 2)
port map (scl_drive_low => scl_drv, sda_drive_low => sda_drv,
scl => scl, sda => sda, scl_in => scl_in, sda_in => sda_in,
scl_released_but_low => scl_rbl, sda_released_but_low => sda_rbl,
scl_holders => scl_h, sda_holders => sda_h);
width_obs : process (clk, rst_n)
begin
if rst_n = '0' then
hi_run <= 0; min_hi <= 9999; armed <= false;
elsif rising_edge(clk) then
if recovering = '1' then
-- The high period before the FIRST pulse is not a pulse, so measurement
-- arms on the first falling edge.
if scl = '0' then armed <= true; end if;
if scl = '1' then
hi_run <= hi_run + 1;
else
if armed and hi_run > 0 and hi_run < min_hi then min_hi <= hi_run; end if;
hi_run <= 0;
end if;
else
hi_run <= 0; armed <= false;
end if;
end if;
end process;
mon : entity work.i2c_proto_mon
generic map (CNT_W => 16)
port map (clk => clk, rst_n => rst_n, scl => scl, sda => sda,
start_seen => mo_start, stop_seen => mo_stop, bit_seen => mo_bit,
bit_val => mo_bitv, byte_seen => mo_byte, byte_val => mo_byteval,
ack_seen => mo_ackv, ack_val => mo_ack, in_transfer => mo_intr,
framing_midbyte => mo_mid, bit_index => mo_bidx,
n_starts => mo_nsta, n_stops => mo_nsto, n_bytes => mo_nbyte,
n_midbyte => mo_nmid);
drove_obs : process (clk, rst_n)
begin
if rst_n = '0' then
ever_drove_scl <= '0';
elsif rising_edge(clk) then
if rec_scl_low = '1' then ever_drove_scl <= '1'; end if;
end if;
end process;
dut : entity work.i2c_err_mgr
generic map (N_PULSES => NP, N_HALF => NH, CNT_W => 16)
port map (clk => clk, rst_n => rst_n,
addr_nack => addr_nack, data_nack => data_nack, arb_lost => arb_lost,
stretch_timeout => stretch_to, in_transfer => in_transfer,
scl_in => scl_in(0), sda_in => sda_in(0),
start_recovery => start_recovery, clear => clear,
recovering => recovering, rec_scl_low => rec_scl_low,
rec_sda_req => rec_sda_req, rec_sda_bit => rec_sda_bit,
err => err, err_valid => err_valid, pulses_issued => pulses,
recovered => recovered, escalate => escalate, state => rstate);
clkgen : process
begin
while not halt loop
clk <= '0'; wait for TCLK/2;
clk <= '1'; wait for TCLK/2;
end loop;
wait;
end process;
-- Count the recovery pulses that actually reach the WIRE, and let go after a configured
-- number of them -- the way a wedged device eventually does.
hold_sda <= want_hold_sda and (not let_go);
pulsecnt : process (clk, rst_n)
variable scl_l : std_logic := '1';
begin
if rst_n = '0' then
scl_l := '1';
pulses_seen <= 0;
let_go <= '0';
elsif falling_edge(clk) then
if want_hold_sda = '0' then
-- A new hold request resets the release decision, so a later test is not
-- affected by an earlier one having let go.
let_go <= '0';
pulses_seen <= 0;
elsif scl = '1' and scl_l = '0' then
pulses_seen <= pulses_seen + 1;
if release_after /= 0 and (pulses_seen + 1) >= release_after then
let_go <= '1';
end if;
end if;
scl_l := scl;
end if;
end process;
stim : process
variable err_n : integer := 0;
variable n : integer;
procedure ck_int (what : string; g : integer; e : integer) is
begin
if g /= e then
report " FAIL " & what & ": got " & integer'image(g)
& " expected " & integer'image(e) severity note;
err_n := err_n + 1;
end if;
end procedure;
procedure ck_bit (what : string; g : std_logic; e : std_logic) is
begin
if g /= e then
report " FAIL " & what & ": got " & std_logic'image(g)
& " expected " & std_logic'image(e) severity note;
err_n := err_n + 1;
end if;
end procedure;
procedure step is
begin
wait until rising_edge(clk); wait until falling_edge(clk);
end procedure;
procedure do_reset is
begin
wait until falling_edge(clk);
rst_n <= '0';
addr_nack <= '0'; data_nack <= '0'; arb_lost <= '0'; stretch_to <= '0';
in_transfer <= '0'; start_recovery <= '0'; clear <= '0';
want_hold_sda <= '0'; hold_scl <= '0'; release_after <= 0;
for i in 0 to 2 loop wait until rising_edge(clk); end loop;
wait until falling_edge(clk); rst_n <= '1';
step;
end procedure;
procedure pulse_recover is
begin
wait until falling_edge(clk); start_recovery <= '1';
wait until rising_edge(clk); wait until falling_edge(clk); start_recovery <= '0';
end procedure;
procedure wait_recovery (max_cycles : integer) is
begin
n := 0;
while not (to_integer(rstate) = R_IDLE and recovering = '0')
and n < max_cycles loop
step; n := n + 1;
end loop;
if n >= max_cycles then
report " FAIL wait_recovery: stuck in state "
& integer'image(to_integer(rstate)) severity note;
err_n := err_n + 1;
end if;
end procedure;
begin
report "=== i2c_err_mgr: six codes, and a remedy for exactly one of them ==="
severity note;
-- T1. THE FOUR REPORTED FAILURES ARE FOUR CODES. An address NACK and a data NACK are
-- different facts, and a master that merged them would make them
-- indistinguishable to software.
do_reset;
wait until falling_edge(clk); in_transfer <= '1'; addr_nack <= '1';
wait until rising_edge(clk); wait until falling_edge(clk); addr_nack <= '0'; step;
report "T1 an address NACK and a data NACK are different codes" severity note;
ck_bit("T1 the address NACK is latched", err(E_ADDR), '1');
ck_bit("T1 and the data NACK is not", err(E_DATA), '0');
wait until falling_edge(clk); data_nack <= '1';
wait until rising_edge(clk); wait until falling_edge(clk); data_nack <= '0'; step;
ck_bit("T1 now the data NACK is latched too", err(E_DATA), '1');
ck_bit("T1 and both are visible at once", err(E_ADDR) and err(E_DATA), '1');
-- T2. Arbitration loss and a stretch timeout are also separate, and neither is a
-- defect: §3.1.8's loser lost a fair contest, and §3.1.6 gives stretching no bound.
do_reset;
wait until falling_edge(clk); in_transfer <= '1'; arb_lost <= '1';
wait until rising_edge(clk); wait until falling_edge(clk); arb_lost <= '0';
wait until falling_edge(clk); stretch_to <= '1';
wait until rising_edge(clk); wait until falling_edge(clk); stretch_to <= '0'; step;
report "T2 arbitration loss and a stretch timeout are separate, and neither is a fault"
severity note;
ck_bit("T2 arbitration loss latched", err(E_ARB), '1');
ck_bit("T2 stretch timeout latched", err(E_STO), '1');
ck_int("T2 and only those two", to_integer(unsigned(err)),
2**E_ARB + 2**E_STO);
-- T3. A LOW LINE DURING A TRANSFER IS NOT A STUCK LINE. Both lines are low most of
-- the time while clocking.
do_reset;
wait until falling_edge(clk); in_transfer <= '1'; want_hold_sda <= '1'; hold_scl <= '1';
for j in 0 to 9 loop step; end loop;
report "T3 a low line during a transfer is not a stuck line" severity note;
ck_bit("T3 SDA is low", sda, '0');
ck_bit("T3 SCL is low", scl, '0');
ck_bit("T3 but no stuck-SDA report", err(E_SDA), '0');
ck_bit("T3 and no stuck-SCL report", err(E_SCL), '0');
-- T4. WITH NO TRANSFER OPEN, THE SAME LINES ARE STUCK -- as TWO codes, because
-- §3.1.16 gives the two lines DIFFERENT remedies.
wait until falling_edge(clk); in_transfer <= '0';
for j in 0 to 5 loop step; end loop;
report "T4 with no transfer open, the same lines are stuck -- as two codes"
severity note;
ck_bit("T4 stuck SDA", err(E_SDA), '1');
ck_bit("T4 stuck SCL", err(E_SCL), '1');
-- T5. A STUCK SDA: the holder lets go on the third pulse, and nine is a BOUND.
do_reset;
wait until falling_edge(clk); want_hold_sda <= '1'; release_after <= 3;
step;
pulse_recover;
wait_recovery(1000);
report "T5 a stuck SDA is freed by clocking, and nine is a bound not a quota"
severity note;
ck_bit("T5 recovered", recovered, '1');
ck_bit("T5 no escalation needed", escalate, '0');
ck_int("T5 it took three pulses, not nine", to_integer(pulses), 3);
ck_bit("T5 SDA is free", sda, '1');
-- T6. AND THE PROCEDURE ENDS WITH A STOP. A freed SDA with no framing leaves every
-- device believing a transfer is in progress: unstuck, and not idle.
report "T6 recovery ends with a STOP, leaving the bus idle rather than unstuck"
severity note;
ck_bit("T6 SCL released", rec_scl_low, '0');
ck_bit("T6 SDA released", m_sda_low, '0');
ck_bit("T6 both lines high", scl and sda, '1');
ck_bit("T6 recovery is finished", recovering, '0');
-- And the STOP must have actually appeared on the bus.
if to_integer(mo_nsto) < 1 then
report " FAIL T6 recovery freed SDA but never framed the bus" severity note;
err_n := err_n + 1;
end if;
-- T7. A STUCK SDA THAT NEVER LETS GO: exactly nine pulses, then escalation.
do_reset;
wait until falling_edge(clk); want_hold_sda <= '1'; release_after <= 0;
step;
pulse_recover;
wait_recovery(2000);
report "T7 a holder that never lets go: exactly nine pulses, then escalate"
severity note;
ck_int("T7 exactly nine pulses", to_integer(pulses), NP);
if min_hi < NH then
report " FAIL T7 a recovery high phase was shorter than N_HALF" severity note;
err_n := err_n + 1;
end if;
ck_bit("T7 not recovered", recovered, '0');
ck_bit("T7 escalated", escalate, '1');
-- T8. THE CENTRAL TEST, AND IT ASSERTS A ZERO. A stuck SCL gets NO pulses, because
-- the nine pulses ARE pulses on SCL.
do_reset;
wait until falling_edge(clk); hold_scl <= '1'; want_hold_sda <= '1';
step;
pulse_recover;
wait_recovery(1000);
report "T8 a stuck SCL gets no pulses at all, because none could reach the wire"
severity note;
ck_int("T8 zero pulses issued", to_integer(pulses), 0);
-- Stronger than the pulse count: the procedure must not have driven the clock at
-- all. A design that starts pulsing and abandons on the first release also reports
-- zero completed pulses, so the counter alone cannot tell the two apart.
ck_bit("T8 and recovery never drove SCL at all", ever_drove_scl, '0');
ck_bit("T8 escalated immediately", escalate, '1');
ck_bit("T8 not recovered", recovered, '0');
ck_bit("T8 and the stuck clock was diagnosed", err(E_SCL), '1');
-- T9. DIAGNOSE BEFORE ACTING, and SCL decides.
report "T9 with both lines stuck, the clock decides the outcome" severity note;
ck_bit("T9 both were diagnosed", err(E_SDA) and err(E_SCL), '1');
ck_int("T9 and still no pulses were attempted", to_integer(pulses), 0);
-- T10. A CLOCK THAT FAILS MID-RECOVERY abandons the procedure rather than reporting
-- nine pulses that never reached the wire.
do_reset;
wait until falling_edge(clk); want_hold_sda <= '1'; release_after <= 0;
step;
pulse_recover;
n := 0;
while pulses_seen < 2 and n < 500 loop step; n := n + 1; end loop;
wait until falling_edge(clk); hold_scl <= '1';
wait_recovery(2000);
report "T10 a clock seized mid-recovery abandons the procedure rather than lying"
severity note;
ck_bit("T10 escalated", escalate, '1');
ck_bit("T10 the stuck clock was recorded", err(E_SCL), '1');
ck_bit("T10 not recovered", recovered, '0');
if to_integer(pulses) >= NP then
report " FAIL T10 reported " & integer'image(to_integer(pulses))
& " pulses after the clock failed" severity note;
err_n := err_n + 1;
end if;
-- T11. RECOVERY ON A HEALTHY BUS DRIVES NOTHING -- the most dangerous case to get
-- wrong, because the procedure would clock a live transfer to pieces and then
-- report success.
do_reset;
pulse_recover;
wait_recovery(1000);
report "T11 recovery on a healthy bus drives nothing and reports it healthy"
severity note;
ck_int("T11 no pulses", to_integer(pulses), 0);
ck_bit("T11 reported healthy", recovered, '1');
ck_bit("T11 no escalation", escalate, '0');
ck_bit("T11 nothing was driven", rec_scl_low or m_sda_low, '0');
ck_bit("T11 both lines still high", scl and sda, '1');
-- T12. Codes clear on request, and a reset manager reports nothing.
do_reset;
wait until falling_edge(clk); in_transfer <= '1'; addr_nack <= '1'; arb_lost <= '1';
wait until rising_edge(clk); wait until falling_edge(clk);
addr_nack <= '0'; arb_lost <= '0'; step;
ck_int("T12 two codes latched", to_integer(unsigned(err)),
2**E_ADDR + 2**E_ARB);
wait until falling_edge(clk); clear <= '1'; step;
wait until falling_edge(clk); clear <= '0'; step;
report "T12 codes clear on request, and a reset manager reports nothing"
severity note;
ck_int("T12 cleared", to_integer(unsigned(err)), 0);
wait until falling_edge(clk); rst_n <= '0'; step;
ck_int("T12 reset clears them too", to_integer(unsigned(err)), 0);
ck_int("T12 and the pulse count", to_integer(pulses), 0);
ck_bit("T12 nothing driven out of reset", rec_scl_low or rec_sda_req, '0');
if err_n = 0 then
report "=== i2c_err_mgr: ALL CHECKS PASSED ===" severity note;
else
report "=== i2c_err_mgr: " & integer'image(err_n)
& " CHECK(S) FAILED ===" severity note;
end if;
halt <= true;
wait;
end process;
end architecture sim;6b. Execution
| Design | SystemVerilog | Verilog-2001 | VHDL | Finish |
|---|---|---|---|---|
i2c_err_mgr | PASS 12/12 | PASS 12/12 | PASS 12/12 | 2390 ns, all three |
7. Mutation Testing — Four Survivors, Four Different Reasons
Ten defects. Four survived the original bench, and each failed for a different reason — which together make a fair summary of how end-state checking goes wrong.
| # | Injected defect | Expected detection | Result |
|---|---|---|---|
| M1 | a stuck SCL is clocked anyway | T8 after strengthening | KILLED (2) |
| M2 | recovery runs on a healthy bus | T11 | KILLED (2) |
| M3 | the two stuck-line failures share one code | T4 | KILLED (4) |
| M4 | a low line during a transfer reported as stuck | T3 | KILLED (3) |
| M5 | ten pulses instead of nine | T7 | KILLED (2) |
| M6 | the pulse's high phase is one cycle instead of N_HALF | T7 after strengthening | KILLED (2) |
| M7 | a clock failing mid-recovery is ignored | T10 | KILLED (3) |
| M8 | the procedure ends without framing the bus | T6 after strengthening | KILLED (2) |
| M9 | an address NACK and a data NACK share one code | T1 | KILLED (5) |
| M10 | recovery claims success after nine pulses failed | T7 | KILLED (2) |
baseline: PASS (verified before injecting anything)
killed: 10 survived: 0 score: 10/10
restored: PASSM1 — the counter said zero and the master had still driven the clock
T8 asserted pulses_issued == 0 for a stuck SCL, plus escalation and the right code. With the diagnosis bypassed, the sequencer falls through to the pulse loop, drives SCL low for a half period, then finds the clock still low on release and escalates — so all four of T8's assertions still held. Zero completed pulses, escalated, not recovered, correct code.
The difference was one signal nobody looked at: the master had driven SCL low on a bus it must not touch.
"issued no pulses" is not the same claim as "drove nothing"pulses_issued counts completed pulses, so a procedure that starts and abandons reports zero. T8 now also asserts a latched "did recovery ever drive SCL" flag, which is the property §3.1.16 actually implies for a stuck clock.
M8 — the end state of success and of failure are identical
Removing the STOP framing left: both lines high, rec_scl_low low, m_sda_low low, recovering clear. That is exactly what a successful recovery leaves behind, because the holder had already let go — so all four of T6's checks passed.
The missing STOP is only visible on the wire. The fix was a protocol monitor and a check that n_stops >= 1, which is why §4c matters: unstuck and idle are different states with the same end-of-recovery snapshot.
M6 — the pulses were counted, and were not legal pulses
Collapsing the high phase to a single cycle left the count at nine, the samples valid, and the final state correct. But §3.1.16's nine pulses are ordinary SCL pulses on a real bus, and a recovery clock running far faster than the bus's timing is not a remedy — it is noise that no device will respond to.
Nothing measured pulse width. The bench now measures the high phase of every recovery pulse and requires at least N_HALF cycles — with the measurement armed only on the first falling edge, because the high period before the first pulse is not a pulse and counting it fails the correct design.
M10 — the invalid mutant, and then the real one
The first attempt changed a state transition in the healthy-bus branch and was inert: the target state cleared the same outputs one cycle later. The real defect is claiming success where the design reports failure — setting recovered alongside escalate after nine pulses have not freed the line — and that dies immediately against T7.
8. Verification Connection — Error Injection and Expected-Error Sequences
// This block cannot be verified by stimulating the DUT: every one of its inputs is a
// report about somebody else's misbehaviour. So the environment needs fault injection
// as a first-class capability, and the injectors are NOT interchangeable:
//
// hold SDA low, no transfer open -> E_SDA_STUCK, recovery attempted
// hold SCL low, no transfer open -> E_SCL_STUCK, recovery REFUSED
// hold SCL low DURING a transfer -> a stretch, NOT a stuck line (§3)
// release the holder after N pulses -> the §4b early-exit path, per N
// hold SDA forever -> nine pulses then escalation
// hold SCL low from pulse 2 onward -> the mid-recovery clock failure (T10)
//
// THE SCOREBOARD'S HARDEST PROBLEM is that two of the six codes are NOT failures.
// A scoreboard written as "any error bit set => test fails" will fail every
// multi-master test, because E_ARB_LOST is the bus working correctly. So the
// environment needs EXPECTED-ERROR SEQUENCES: a sequence declares the outcome it
// intends, and the scoreboard compares against that rather than against zero.
//
// class arb_loss_seq extends i2c_base_seq;
// // this sequence EXPECTS E_ARB_LOST and expects the master to retry
// endclass
//
// Without that, the usual outcome is that somebody masks E_ARB_LOST out of the
// scoreboard's check -- and then a master that reports arbitration loss on every
// READ (Chapter 17.4 §9's defect) passes silently forever.
//
// WHAT MUST NOT BE CHECKED BY MASKING: the difference between E_SDA_STUCK and
// E_SCL_STUCK. §4 shows the remedies are different, so a scoreboard that accepts
// either code for a stuck-line test accepts a master that would send its driver to
// attempt the impossible procedure.
//
// COVERAGE, and one bin here is deliberately unreachable:
//
// cover: recovery succeeded on pulse 1..9 -- nine reachable bins
// cover: recovery escalated after nine pulses
// cover: recovery refused because SCL was stuck
// cover: recovery requested on a HEALTHY bus -- T11, and it must drive nothing
// illegal_bin: pulses issued > 9 -- the bound, not a quota
// illegal_bin: pulses issued on a stuck SCL -- UNREACHABLE by §49. FPGA and ASIC Implications
On an FPGA, recovery is the one place the master drives the bus outside a transaction, and the ownership consequence matters: rec_scl_low and rec_sda_req must be muxed into the same pad paths the rest of the design uses, with recovery as Chapter 17.4's owner 3. Giving recovery its own pad drivers would put two drivers on one pin. The recovering output is the mux select, and it is the reason recovery is a requester rather than a special case in the pad logic.
The diagnosis also depends on the synchronised readback, which means a stuck line must be stuck for at least the synchroniser depth before it is diagnosed. That is the right behaviour — a line low for two cycles is not stuck — and it is another case where the latency is a feature.
On an ASIC, §3.1.16's escalation path is a system integration requirement, not an RTL one. The escalate output must reach something that can actually perform a hardware reset or a power cycle: a reset controller, a PMIC sequencer, or at minimum a status bit that firmware polls before deciding the bus is unusable. A design that reports escalate into a register nobody reads has implemented the diagnosis and none of the remedy.
The pulse width is set by N_HALF, and it should be derived from the same mode parameters as Chapter 17.3's generator rather than chosen independently — recovery pulses are ordinary SCL pulses and are subject to the same Table 10 minimums, which is exactly what mutation M6 exploited. Chapter 17.13 ties both to one source.
10. Debugging — The Recovery That Made Things Worse
A gateway board has an I2C bus that occasionally locks up: the SoC reports SDA stuck and the bus stops working until the board is power-cycled. A firmware update adds an automatic recovery routine that issues the master's bus-recovery command whenever any transaction times out. After the update, lockups become MORE frequent, and a new symptom appears -- a temperature sensor on the same bus starts returning corrupt readings during periods of heavy bus traffic, which it never did before.
The recovery procedure was invoked while a transfer was in progress, and a transfer in progress looks exactly like a stuck SDA at the instant the diagnosis samples: SCL high between pulses, SDA low because someone is transmitting a zero. The block correctly suppresses stuck-line detection during a transfer -- that is the in_transfer qualifier -- but the host-commanded recovery path did not apply the same qualifier, so a command issued at the wrong moment produced a diagnosis that was wrong for a reason the diagnosis could not see. The firmware was also at fault for invoking recovery on a per-transaction timeout rather than on a bus-level fault, but the hardware accepted a command it should have refused.
Refuse a recovery request while a transfer is open: the same in_transfer qualifier that guards detection must guard the procedure, because the diagnosis cannot distinguish a stuck line from a live transfer by looking at the lines. Then fix the firmware policy -- recovery is a response to a bus fault, not to one transaction's timeout, and a timed-out transaction on a healthy bus needs a retry rather than nine pulses. For the regression: request recovery mid-transfer and assert that nothing is driven, which is T11's property applied to a new precondition. T11 as written covers a healthy IDLE bus; the case that bit here is a healthy BUSY one.Three generalisations.
A correct qualifier was applied to detection and not to action. in_transfer guarded the stuck-line report and not the recovery procedure. The two paths need the same guard for the same reason, and the asymmetry is invisible unless someone asks what the diagnosis can actually see.
The diagnosis is a snapshot of a situation that has history. SCL high and SDA low is a stuck bus or an ordinary transmitted zero between clock pulses. No sampling of the two lines can separate them; only knowing whether a transfer is open can.
Automatic recovery on the wrong trigger is worse than no recovery. The procedure worked perfectly and was invoked wrongly, and a procedure that drives nine pulses into a live transfer corrupts a device that had nothing to do with the fault.
11. Common Misconceptions
"The specification defines I²C error codes." It defines none. Every code in a master's status register is a design decision. §1.
"More error codes is over-engineering." Two failures sharing one code are indistinguishable to software forever. E_SDA_STUCK and E_SCL_STUCK have different remedies, and one code sends a driver to attempt the impossible one half the time. §4.
"An address NACK means nothing is there." It means nothing answered — absent device or busy device, and 16.4 shows those are indistinguishable. They share a code honestly. §2.
"Arbitration loss is an error." The bus worked exactly as §3.1.8 specifies. The response is to retry when the bus is free, not to report a fault. §2a.
"A stretch timeout means the target is broken." §3.1.6 sets no bound, so the target may be perfectly conforming. The flag says the master stopped waiting. §2a.
"A low line means a stuck line." Both lines are low most of the time during a transfer. Detection has to be qualified by whether a transfer is open. §3.
"Nine clock pulses can recover any stuck bus." They are pulses on SCL, so they cannot recover a stuck SCL. §3.1.16 sends you to a hardware reset or a power cycle for that. §4.
"Nine pulses means always send nine." Nine is a bound. §3.1.16 says the holder should release within those nine, so stopping early is correct and the count is worth reporting. §4b.
"Once SDA is free, recovery is done." A freed SDA with no framing leaves every device believing a transfer is open — unstuck and not idle. §4c.
"Zero pulses issued proves the clock was never driven." The counter counts completed pulses. A procedure that starts and abandons reports zero having driven the line. §7.
"If the end state is right, the procedure was right." Three of four survivors here left an end state identical to success. §7.
12. Reason It Through
Why must a stuck SDA and a stuck SCL be different error codes?
Because the remedies are different — nine pulses for SDA, a hardware reset or power cycle for SCL — and the nine pulses are themselves pulses on SCL. One code would send a driver to attempt the impossible procedure half the time. §4.
Why does the diagnosis test SCL before SDA rather than in parallel?
Because a stuck clock makes the procedure impossible whatever SDA is doing. The tests are ordered because one of them is disqualifying. §4a.
Two of the six codes are not faults. Which, and what should a driver do with each?
Arbitration loss — retry when the bus is free, because the bus worked as designed. And a stretch timeout — decide whether to wait longer, because the target may still be conforming. §2a.
Why is stuck-line detection qualified by in_transfer?
Because both lines are low most of the time during a transfer, so unqualified detection reports stuck lines on every byte. §3.
A recovery procedure reports zero pulses issued on a stuck clock. Does that prove it never drove SCL?
No. The counter counts completed pulses, so a procedure that drives the clock low and abandons on the first release reports zero having touched the line. A separate flag is needed for "drove anything at all". §7.
Why is "both lines high, recovery finished" not evidence that recovery framed the bus?
Because a recovery that freed SDA and never issued a STOP leaves exactly that state — the holder had already let go. The STOP is only visible on the wire. §7.
Nine recovery pulses are issued at the correct count and the bus does not recover. What else could be wrong with them?
Their width. Recovery pulses are ordinary SCL pulses subject to Table 10's minimums, and a pulse train that is counted correctly but clocked far too fast is noise no device will respond to. §7.
Why did applying in_transfer to detection but not to the recovery command produce a corrupted, innocent device?
Because the diagnosis cannot tell a stuck SDA from a transmitted zero between clock pulses by looking at the lines. Invoked mid-transfer, it diagnosed a stuck line and injected nine pulses into a live byte. §10.
13. Understanding Check
14. Summary
UM10204 defines no error codes, so the taxonomy is a design decision — and two failures sharing one code are indistinguishable to software forever.
Six codes, and two of them are not faults. Arbitration loss is the bus working as specified; a stretch timeout is the master deciding to stop waiting, not a defect it found.
A low line is not a stuck line. Both lines are low most of the time during a transfer, so detection is qualified by whether a transfer is open.
The one decision the specification forces is that SDA-stuck and SCL-stuck are separate codes, because the nine pulses of §3.1.16 are pulses on SCL and therefore cannot rescue a stuck clock.
So the diagnosis is ordered, not parallel — SCL is tested first, because a stuck clock disqualifies the procedure whatever SDA is doing.
Nine is a bound, not a quota, and the count actually needed distinguishes a device that let go on the third pulse from one that never did.
Recovery ends with a STOP, because a freed SDA with no framing leaves the bus unstuck and not idle.
Ten mutants, ten killed — after four survived, and three of the four passed because the end state after the defect is identical to the end state after success: an unframed bus, a briefly-driven clock, a pulse train that was too fast.
So end-state checks are necessary and not sufficient for a block whose subject is a procedure. What reached the wire, and how long it took, are separate observations needing separate instruments — a protocol monitor, a drove-anything latch, and a pulse-width measurement.
And a correct qualifier applied to detection but not to action is a real defect. Recovery invoked mid-transfer diagnoses a transmitted zero as a stuck line and clocks nine pulses into a live byte, corrupting a device that had nothing to do with the fault.
15. What Comes Next
Every block now exists, is verified, and reports what it did. Chapter 17.12 puts them together — and the point of that chapter is the order in which it does so.
The FSM comes last, and its states are not a description of the protocol. They are the coordination between blocks that already work: who is granted SCL, who owns SDA, which block's completion pulse the next transition waits on. Everything that could be a counter, a shift register or a comparator has already been built and tested, so what remains is genuinely small — which is the payoff for the decomposition Chapter 17.1 argued for, and the reason a master written FSM-first gets rewritten.
Continue learning
Related tutorials
- Related topic
Bus Clear and Recovery — The Protocol-Level Escape Hatch
Four sentences of specification that say two different things about two different wires. Explains why nine clock pulses free a stuck SDA and can never free a stuck SCL, why that is structural rather than an omission, and what a recovery sequencer must refuse to attempt.
- Related topic
Decomposing an I²C Master — From Requirements to Architecture
An I²C master is not one state machine, and the reason is structural rather than stylistic: the protocol imposes four independent time bases that change on four unrelated events. Derives the block structure from the normative obligations, establishes the wired-AND bus model every later chapter is written against, and shows why the framing generator cannot live inside the bit engine.
- Related topic
The Master Command Interface — Register Model and On-Chip Bus
The one block in an I²C master that UM10204 says nothing about, which makes it harder rather than easier. Derives what software must be able to express and what the master must report back, why done and ok are two bits rather than one, why only the command register may start a transfer, and why an on-chip register bus forces a post-then-poll handshake.
- Related topic
The SCL Timing Generator — Phases, Strobes and the Readback Rule
Where Table 10's microseconds become counts of system-clock cycles. Derives the period budget that must include rise and fall time, shows why rounding down is always illegal and rounding up always legal, and builds a generator that leaves its low phase only when the line actually reads back high — which implements clock stretching and clock synchronization with no extra logic.
