I²C · Module 19
Hardware Bring-Up and On-Chip Debug
What to do on the morning the board arrives, in what order, with what instrument. Builds a synthesizable bus-health block that answers 'is the bus even there' at a glance, works through a probe set that spans layers rather than concentrating on the FSM, and shows why an analyzer and an ILA disagreeing is the most localising evidence available.
Chapter 19.7 catalogued what can differ between simulation and a board, and ordered the questions. This chapter is the procedure: the board is on the bench, the bitstream is loaded, and nothing works.
The single most important idea is the order.
1. The Order
Each step establishes a precondition for the next, and each has a definite instrument and a definite pass criterion.
| # | question | instrument | pass criterion |
|---|---|---|---|
| 1 | is there power, and a common ground? | multimeter | rails at nominal; continuity between the two boards' grounds |
| 2 | do the pins go where the constraints say? | multimeter / scope | continuity from FPGA ball to bus net |
| 3 | are the pull-ups fitted, and sized? | multimeter, then scope | correct resistance; rise time within 19.3's budget |
| 4 | does the bus idle HIGH? | scope, or i2c_bus_health.idle | both lines at VDD with nothing driving |
| 5 | is the FPGA clocked and out of reset? | ILA, or an LED | the design's own heartbeat |
| 6 | do the raw pins reach the fabric? | ILA on scl_pin, sda_pin | they move when the controller talks |
| 7 | does synchronization work? | ILA on scl_q, sda_q | follow the pins, 19.4's latency later |
| 8 | does the filter pass real edges? | ILA on the filtered level and n_rejected | edges arrive; rejects near zero |
| 9 | is framing detected? | ILA on start_pulse, stop_pulse | one pulse per framing event |
| 10 | does the address match? | ILA on selected, dir_read | asserts for your address only |
| 11 | is the acknowledge ours? | scope on the ninth bit plus ILA | the wire goes low, and the target intended it |
| 12 | does application state advance? | ILA on the pointer and register file | writes land where the model says |
Steps 1 to 4 involve no FPGA logic at all, and in 19.7's experience they are where first-power-on failures usually are. Steps 6 to 8 are the front end this module built. Steps 9 to 12 are Module 18, which is the part least likely to be wrong.
2. Four Bits That Answer Step 4 Without a Scope
Steps 4 to 6 are the ones you re-check constantly — after every wiring change, every resistor swap, every "did that help". Doing it with an ILA means arming a capture and reading waveforms each time. Doing it with four bits means looking.
// -----------------------------------------------------------------------------
// i2c_bus_health.sv
// A bring-up aid: what is wrong with this bus, before any protocol is attempted.
//
// WHY THIS EXISTS. On the morning a board arrives, the first question is not "does
// my state machine work" but "is the bus even there". Answering it with an ILA means
// capturing waveforms and reading them; answering it with four LEDs or four register
// bits means looking once. This block is small enough to leave in a shipping design
// and useful enough to be worth it.
//
// WHAT IT OBSERVES, and it observes ONLY -- it never drives either line:
//
// idle both lines have read HIGH continuously for IDLE_CLKS clocks.
// This is the healthy resting state, and it is the single most
// informative bit: if it never asserts, nothing else matters yet.
//
// scl_stuck_low SCL has read LOW continuously for STUCK_CLKS clocks. A controller
// sda_stuck_low SDA likewise. Distinguishes "no pull-up fitted", "a device holding
// the line", and "the pin is not connected to what you think" from
// each other -- see the chapter's step 4.
//
// activity at least one edge has been seen on either line since reset. The
// difference between "the bus is idle because nothing is talking"
// and "the bus is idle because it is not connected".
//
// WHY `activity` IS A LATCHED STICKY BIT rather than a level: bring-up happens at
// human speed, and a level that is only true for one clock in ten million is not
// observable on an LED. Once an edge has been seen, the fact is kept until reset.
//
// WHAT IT DELIBERATELY DOES NOT DO: decode the protocol. There is no framing, no
// address, no acknowledge. Every one of those requires the bus to be working first,
// and a block that reported "no START seen" on a bus with no pull-up would be
// answering a question the engineer has not reached yet. Debug from the physical
// boundary inward, one layer at a time.
//
// IT TAKES THE SYNCHRONISED LEVELS, not the pins -- Chapter 19.4. A diagnostic that
// sampled the raw pins would be the one block in the design with a metastability
// path, which would be an unusually poor trade for a status LED.
// -----------------------------------------------------------------------------
module i2c_bus_health #(
// Clocks both lines must read HIGH before the bus is called idle. Long enough
// that a real inter-transfer gap does not count -- tBUF is 1.3 us in Fast mode,
// so at 50 MHz anything above ~100 clocks distinguishes "resting" from "between
// bytes". The default is deliberately generous.
parameter int IDLE_CLKS = 256,
// Clocks a line must read LOW before it is called stuck. Must exceed the longest
// legal LOW: a clock-stretching target can hold SCL down for a long time, so this
// is the parameter to raise if a healthy bus reports stuck.
parameter int STUCK_CLKS = 4096,
// Width of the edge counter. A PARAMETER rather than a fixed 16, because the
// saturation behaviour has to be testable: reaching the top of a 16-bit counter
// takes 65535 edges, which no reasonable simulation drives, so a bench that
// could not narrow it could not test saturation at all. Mutation G10 -- which
// removes the saturation guard -- survived the first version of this bench for
// exactly that reason.
parameter int EDGE_W = 16
) (
input logic clk,
input logic rst_n,
// The SYNCHRONISED levels from Chapter 19.4, never the pins.
input logic scl_q,
input logic sda_q,
// ---- the four bits worth putting on an LED -------------------------------
output logic idle,
output logic scl_stuck_low,
output logic sda_stuck_low,
output logic activity,
// ---- and one number worth putting on a register --------------------------
// Edges seen on either line. Saturating, because a counter that wrapped would
// read 0 on a busy bus and be indistinguishable from a dead one.
output logic [EDGE_W-1:0] n_edges
);
logic scl_d, sda_d;
logic [15:0] idle_cnt, scl_low_cnt, sda_low_cnt;
wire any_edge = (scl_q != scl_d) || (sda_q != sda_d);
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
// Reset to the values that mean "I have not seen anything yet", which is
// NOT the same as the values that mean "the bus is healthy". A diagnostic
// that reset to `idle = 1` would report a healthy bus before it had looked
// at one, which is the worst possible default for a bring-up aid.
scl_d <= 1'b1;
sda_d <= 1'b1;
idle <= 1'b0;
scl_stuck_low <= 1'b0;
sda_stuck_low <= 1'b0;
activity <= 1'b0;
idle_cnt <= 16'd0;
scl_low_cnt <= 16'd0;
sda_low_cnt <= 16'd0;
n_edges <= {EDGE_W{1'b0}};
end else begin
scl_d <= scl_q;
sda_d <= sda_q;
// ---- activity: sticky, and saturating ----------------------------
if (any_edge) begin
activity <= 1'b1;
if (n_edges != {EDGE_W{1'b1}}) n_edges <= n_edges + 1'b1;
end
// ---- idle: both high, continuously ------------------------------
if (scl_q && sda_q) begin
if (idle_cnt >= IDLE_CLKS[15:0]) idle <= 1'b1;
else idle_cnt <= idle_cnt + 16'd1;
end else begin
idle <= 1'b0;
idle_cnt <= 16'd0;
end
// ---- stuck low: one line, continuously --------------------------
if (!scl_q) begin
if (scl_low_cnt >= STUCK_CLKS[15:0]) scl_stuck_low <= 1'b1;
else scl_low_cnt <= scl_low_cnt + 16'd1;
end else begin
scl_stuck_low <= 1'b0;
scl_low_cnt <= 16'd0;
end
if (!sda_q) begin
if (sda_low_cnt >= STUCK_CLKS[15:0]) sda_stuck_low <= 1'b1;
else sda_low_cnt <= sda_low_cnt + 16'd1;
end else begin
sda_stuck_low <= 1'b0;
sda_low_cnt <= 16'd0;
end
end
end
endmoduleThree design decisions in there are the chapter.
It resets to "I have not seen anything yet", not to "healthy". idle comes out of reset at 0. A diagnostic that reset to idle = 1 would report a good bus before it had looked at one, which is the worst possible default for a block whose whole job is to be believed at power-on. Mutation G01 makes that change and fails nine checks.
idle and activity answer different questions, and together they answer the one that matters. A disconnected bus with pull-ups fitted looks exactly like a healthy resting bus: both lines high, nothing moving. The only thing that distinguishes them is whether anything has ever happened.
idle=1 activity=0 → the bus is quiet and has NEVER been used.
Likely: not connected, or the controller is not running.
idle=1 activity=1 → healthy, resting between transfers.
idle=0 activity=1 → busy, or something is holding a line.
idle=0 activity=0 → a line is stuck from power-on, and nothing has moved.
Look at the two stuck bits.Test T3 is the one that makes activity worth having: it asserts idle == 1 and activity == 0 simultaneously, which is the dead-bus signature, and would fail on any design where one implied the other.
The two stuck bits are reported separately, because which line is stuck is the entire diagnostic value:
| meaning | |
|---|---|
| SDA stuck low, SCL healthy | a target is holding the bus — a wedged device, or a stretch that never ended |
| SCL stuck low, SDA healthy | a controller not releasing, or a target stretching for ever |
| both stuck low | usually no pull-up fitted, or no power to the pull-up rail |
| neither, and not idle | normal traffic |
Mutation G06 wires the two together and fails five checks.
3. Verifying a Diagnostic
A diagnostic's tests are about discrimination: each bit must distinguish the condition it names from the conditions it does not. A status bit that is right about a healthy bus and right about a dead one is not a diagnostic.
// -----------------------------------------------------------------------------
// i2c_bus_health_tb.sv
// Independent oracle for i2c_bus_health.
//
// THE FOUR OUTPUTS ARE A DIAGNOSTIC, so the tests are about DISCRIMINATION: each
// one must distinguish the condition it names from the conditions it does not. A
// status bit that is right about a healthy bus and also right about a dead one is
// not a diagnostic, and T3 and T7 are the tests that separate those.
//
// Small parameters (IDLE=8, STUCK=12) so the thresholds are reachable in a short
// simulation. The chapter body gives the real values and the reason for them.
//
// Every wait is a fixed number of clocks; nothing waits on the DUT.
// -----------------------------------------------------------------------------
`timescale 1ns/1ps
module i2c_bus_health_tb;
localparam int IDLE = 8;
localparam int STUCK = 12;
logic clk = 1'b0, rst_n = 1'b0;
logic scl_q = 1'b1, sda_q = 1'b1;
logic idle, scl_stuck, sda_stuck, activity;
logic [15:0] n_edges;
// A SECOND INSTANCE with a 3-bit edge counter, so the saturation behaviour is
// reachable in a short simulation. Reaching the top of the default 16-bit counter
// takes 65535 edges; at 3 bits it takes 7, so T9 can drive past it and require
// the count to STOP rather than wrap. Mutation G10 survived until this existed.
logic nn_idle, nn_scls, nn_sdas, nn_act;
logic [2:0] nn_edges;
integer errors = 0;
integer n;
i2c_bus_health #(.IDLE_CLKS(IDLE), .STUCK_CLKS(STUCK)) dut (
.clk(clk), .rst_n(rst_n), .scl_q(scl_q), .sda_q(sda_q),
.idle(idle), .scl_stuck_low(scl_stuck), .sda_stuck_low(sda_stuck),
.activity(activity), .n_edges(n_edges));
i2c_bus_health #(.IDLE_CLKS(IDLE), .STUCK_CLKS(STUCK), .EDGE_W(3)) narrow (
.clk(clk), .rst_n(rst_n), .scl_q(scl_q), .sda_q(sda_q),
.idle(nn_idle), .scl_stuck_low(nn_scls), .sda_stuck_low(nn_sdas),
.activity(nn_act), .n_edges(nn_edges));
always #5 clk = ~clk;
task step; begin @(posedge clk); @(negedge clk); end endtask
task do_reset;
begin
@(negedge clk); rst_n = 1'b0; scl_q = 1'b1; sda_q = 1'b1;
step; step;
@(negedge clk); rst_n = 1'b1;
end
endtask
// One SCL pulse, as a working controller would produce.
task scl_pulse;
begin
@(negedge clk); scl_q = 1'b0; step; step;
@(negedge clk); scl_q = 1'b1; step; step;
end
endtask
task ck (input [200*8:1] what, input integer g, input integer e);
begin
if (g !== e) begin
$display(" FAIL %0s: got %0d expected %0d", what, g, e);
errors = errors + 1;
end
end
endtask
task ck_idx (input [200*8:1] what, input integer idx,
input integer g, input integer e);
begin
if (g !== e) begin
$display(" FAIL %0s[%0d]: got %0d expected %0d", what, idx, g, e);
errors = errors + 1;
end
end
endtask
initial begin
$display("=== i2c_bus_health: four bits that each distinguish something ===");
// ----------------------------------------------------------------
// T1. RESET REPORTS IGNORANCE, NOT HEALTH. Every bit is 0 out of reset,
// including `idle`. A diagnostic that reset to "healthy" would report a
// good bus before it had looked at one, which is the worst possible
// default for a block whose entire job is to be believed on power-on.
// ----------------------------------------------------------------
@(negedge clk); rst_n = 1'b0; scl_q = 1'b1; sda_q = 1'b1; step; step;
$display("T1 out of reset every bit says 'I have not seen anything yet'");
ck("T1 not idle yet", idle, 0);
ck("T1 no scl stuck", scl_stuck, 0);
ck("T1 no sda stuck", sda_stuck, 0);
ck("T1 no activity yet", activity, 0);
ck("T1 no edges counted", n_edges, 0);
// ----------------------------------------------------------------
// T2. AN IDLE BUS IS RECOGNISED, AFTER EXACTLY IDLE_CLKS. Not before: a bus
// that has been high for two clocks is between bytes, not resting.
// ----------------------------------------------------------------
do_reset;
for (n = 1; n <= IDLE + 2; n = n + 1) begin
step;
ck_idx("T2 idle asserts only after IDLE_CLKS", n, idle,
(n > IDLE) ? 1 : 0);
end
$display("T2 an idle bus is recognised after exactly IDLE_CLKS clocks");
// ----------------------------------------------------------------
// T3. AND AN IDLE BUS IS NOT THE SAME AS A DEAD ONE. This is the test that
// makes `idle` worth having: a disconnected bus with pull-ups fitted looks
// exactly like a healthy resting bus on the two lines, and the ONLY thing
// that separates them is whether anything has ever happened. `activity`
// is that bit, and here it must still be 0 while `idle` is 1.
// ----------------------------------------------------------------
$display("T3 idle and dead look identical on the wires -- 'activity' separates them");
ck("T3 the bus reports idle", idle, 1);
ck("T3 but nothing has ever happened", activity, 0);
ck("T3 and no edges were counted", n_edges, 0);
// ----------------------------------------------------------------
// T4. ONE EDGE IS ENOUGH TO PROVE THE BUS IS CONNECTED, AND IT STICKS. Bring-up
// happens at human speed, so a level true for one clock in ten million is
// not observable. The fact is latched.
// ----------------------------------------------------------------
@(negedge clk); sda_q = 1'b0; step;
$display("T4 one edge proves connectivity, and the fact is kept");
ck("T4 activity asserted", activity, 1);
ck("T4 one edge counted", n_edges, 1);
@(negedge clk); sda_q = 1'b1; step; step; step;
ck("T4 two edges now", n_edges, 2);
ck("T4 and activity is sticky", activity, 1);
// ----------------------------------------------------------------
// T5. A LINE HELD LOW IS REPORTED, AFTER EXACTLY STUCK_CLKS, AND THE TWO LINES
// ARE REPORTED SEPARATELY. Which line is stuck is the whole diagnostic
// value: SDA stuck low with SCL healthy is a device holding the bus, while
// both stuck low is usually no pull-up or no power.
// ----------------------------------------------------------------
do_reset;
@(negedge clk); sda_q = 1'b0;
for (n = 1; n <= STUCK + 2; n = n + 1) begin
step;
ck_idx("T5 sda_stuck asserts only after STUCK_CLKS", n, sda_stuck,
(n > STUCK) ? 1 : 0);
ck_idx("T5 and SCL is never implicated", n, scl_stuck, 0);
end
$display("T5 a stuck line is named individually, after exactly STUCK_CLKS");
ck("T5 and the bus is not idle", idle, 0);
// ----------------------------------------------------------------
// T6. THE OTHER LINE, SYMMETRICALLY. A diagnostic that only worked on one line
// would be half a diagnostic, and the half that is missing is the one that
// matters most: SCL stuck low is the failure that stops everything.
// ----------------------------------------------------------------
do_reset;
@(negedge clk); scl_q = 1'b0;
for (n = 1; n <= STUCK + 2; n = n + 1) begin
step;
ck_idx("T6 scl_stuck asserts only after STUCK_CLKS", n, scl_stuck,
(n > STUCK) ? 1 : 0);
ck_idx("T6 and SDA is never implicated", n, sda_stuck, 0);
end
$display("T6 and the same for SCL, independently");
// ----------------------------------------------------------------
// T7. A WORKING BUS IS NOT REPORTED AS STUCK. The discrimination that matters
// in the other direction: a controller clocking normally holds SCL low for
// part of every bit, and a threshold that counted that would report a
// healthy bus as broken. Eight pulses, and neither stuck bit may assert.
// ----------------------------------------------------------------
do_reset;
for (n = 0; n < 8; n = n + 1) scl_pulse;
$display("T7 normal clocking is not mistaken for a stuck line");
ck("T7 SCL not reported stuck", scl_stuck, 0);
ck("T7 SDA not reported stuck", sda_stuck, 0);
ck("T7 but activity was seen", activity, 1);
ck("T7 with sixteen edges", n_edges, 16);
// ----------------------------------------------------------------
// T8. A STUCK REPORT CLEARS WHEN THE LINE RECOVERS. A latched fault would be
// indistinguishable from a present one, so the engineer could not tell
// whether the fix worked without a reset -- and on a board being probed,
// "did that help" is the question being asked every few seconds.
// ----------------------------------------------------------------
do_reset;
@(negedge clk); sda_q = 1'b0;
for (n = 0; n < STUCK + 2; n = n + 1) step;
ck("T8 reported stuck", sda_stuck, 1);
@(negedge clk); sda_q = 1'b1; step; step;
$display("T8 a stuck report clears when the line recovers");
ck("T8 and cleared when the line came back", sda_stuck, 0);
// ... and the bus then becomes idle again, which is the recovery confirmed.
for (n = 0; n < IDLE + 2; n = n + 1) step;
ck("T8 the bus is idle again", idle, 1);
// ----------------------------------------------------------------
// T9. THE EDGE COUNTER SATURATES RATHER THAN WRAPPING. A counter that wrapped
// would read 0 on a very busy bus and be indistinguishable from a dead one
// -- the exact confusion this block exists to remove. Checked by forcing it
// near the top and driving past it.
// ----------------------------------------------------------------
do_reset;
ck("T9 both counters start at zero", n_edges, 0);
ck("T9 including the narrow one", nn_edges, 0);
// Three pulses = six edges, one short of the 3-bit maximum of 7.
for (n = 0; n < 3; n = n + 1) scl_pulse;
ck("T9 the narrow counter reached six", nn_edges, 6);
scl_pulse; // two more edges: 7, then it must STOP
ck("T9 and saturated at its maximum", nn_edges, 7);
for (n = 0; n < 6; n = n + 1) scl_pulse;
ck("T9 twelve further edges do not wrap it", nn_edges, 7);
// The wide counter, meanwhile, is still counting normally -- which proves the
// saturation is at the parameterised maximum and not at a hard-coded value.
ck("T9 while the wide counter kept counting", n_edges, 20);
// ----------------------------------------------------------------
// T10. `idle` DEASSERTS THE MOMENT THE BUS STOPS RESTING.
//
// Checked from an ASSERTED idle state, because that is the only state in
// which the deassertion is observable -- every stuck-line test above
// begins with a reset, where idle is already 0, so none of them can see
// it. Mutation G11, which never clears idle, survived every other test in
// this list. Placed last rather than inside T3 because T4 depends on T3
// leaving `activity` still deasserted, and driving edges here would have
// spoiled that.
// ----------------------------------------------------------------
do_reset;
for (n = 0; n < IDLE + 2; n = n + 1) step;
ck("T10 precondition: the bus reports idle", idle, 1);
@(negedge clk); scl_q = 1'b0; step;
$display("T10 idle deasserts as soon as the bus stops resting");
ck("T10 idle cleared immediately", idle, 0);
@(negedge clk); scl_q = 1'b1;
for (n = 0; n < IDLE + 2; n = n + 1) step;
ck("T10 and returns after IDLE_CLKS of quiet again", idle, 1);
if (errors == 0) $display("=== i2c_bus_health: ALL CHECKS PASSED ===");
else $display("=== i2c_bus_health: %0d CHECK(S) FAILED ===", errors);
$finish;
end
endmoduleTwo tests are the discrimination in each direction. T3 proves a dead bus is not reported as healthy. T7 proves a working bus is not reported as stuck — a controller clocking normally holds SCL low for part of every bit, and a threshold that counted that would report a healthy bus as broken. Both directions have to be tested, because a threshold error in either one produces a diagnostic that is confidently wrong.
// -----------------------------------------------------------------------------
// i2c_bus_health.v
// A bring-up aid: what is wrong with this bus, before any protocol is attempted.
//
// WHY THIS EXISTS. On the morning a board arrives, the first question is not "does
// my state machine work" but "is the bus even there". Answering it with an ILA means
// capturing waveforms and reading them; answering it with four LEDs or four register
// bits means looking once. This block is small enough to leave in a shipping design
// and useful enough to be worth it.
//
// WHAT IT OBSERVES, and it observes ONLY -- it never drives either line:
//
// idle both lines have read HIGH continuously for IDLE_CLKS clocks.
// This is the healthy resting state, and it is the single most
// informative bit: if it never asserts, nothing else matters yet.
//
// scl_stuck_low SCL has read LOW continuously for STUCK_CLKS clocks. A controller
// sda_stuck_low SDA likewise. Distinguishes "no pull-up fitted", "a device holding
// the line", and "the pin is not connected to what you think" from
// each other -- see the chapter's step 4.
//
// activity at least one edge has been seen on either line since reset. The
// difference between "the bus is idle because nothing is talking"
// and "the bus is idle because it is not connected".
//
// WHY `activity` IS A LATCHED STICKY BIT rather than a level: bring-up happens at
// human speed, and a level that is only true for one clock in ten million is not
// observable on an LED. Once an edge has been seen, the fact is kept until reset.
//
// WHAT IT DELIBERATELY DOES NOT DO: decode the protocol. There is no framing, no
// address, no acknowledge. Every one of those requires the bus to be working first,
// and a block that reported "no START seen" on a bus with no pull-up would be
// answering a question the engineer has not reached yet. Debug from the physical
// boundary inward, one layer at a time.
//
// IT TAKES THE SYNCHRONISED LEVELS, not the pins -- Chapter 19.4. A diagnostic that
// sampled the raw pins would be the one block in the design with a metastability
// path, which would be an unusually poor trade for a status LED.
// -----------------------------------------------------------------------------
module i2c_bus_health #(
// Clocks both lines must read HIGH before the bus is called idle. Long enough
// that a real inter-transfer gap does not count -- tBUF is 1.3 us in Fast mode,
// so at 50 MHz anything above ~100 clocks distinguishes "resting" from "between
// bytes". The default is deliberately generous.
parameter IDLE_CLKS = 256,
// Clocks a line must read LOW before it is called stuck. Must exceed the longest
// legal LOW: a clock-stretching target can hold SCL down for a long time, so this
// is the parameter to raise if a healthy bus reports stuck.
parameter STUCK_CLKS = 4096,
// Width of the edge counter. A PARAMETER rather than a fixed 16, because the
// saturation behaviour has to be testable: reaching the top of a 16-bit counter
// takes 65535 edges, which no reasonable simulation drives, so a bench that
// could not narrow it could not test saturation at all. Mutation G10 -- which
// removes the saturation guard -- survived the first version of this bench for
// exactly that reason.
parameter EDGE_W = 16
) (
input wire clk,
input wire rst_n,
// The SYNCHRONISED levels from Chapter 19.4, never the pins.
input wire scl_q,
input wire sda_q,
// ---- the four bits worth putting on an LED -------------------------------
output reg idle,
output reg scl_stuck_low,
output reg sda_stuck_low,
output reg activity,
// ---- and one number worth putting on a register --------------------------
// Edges seen on either line. Saturating, because a counter that wrapped would
// read 0 on a busy bus and be indistinguishable from a dead one.
output reg [EDGE_W-1:0] n_edges
);
reg scl_d, sda_d;
reg [15:0] idle_cnt, scl_low_cnt, sda_low_cnt;
wire any_edge = (scl_q != scl_d) || (sda_q != sda_d);
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
// Reset to the values that mean "I have not seen anything yet", which is
// NOT the same as the values that mean "the bus is healthy". A diagnostic
// that reset to `idle = 1` would report a healthy bus before it had looked
// at one, which is the worst possible default for a bring-up aid.
scl_d <= 1'b1;
sda_d <= 1'b1;
idle <= 1'b0;
scl_stuck_low <= 1'b0;
sda_stuck_low <= 1'b0;
activity <= 1'b0;
idle_cnt <= 16'd0;
scl_low_cnt <= 16'd0;
sda_low_cnt <= 16'd0;
n_edges <= {EDGE_W{1'b0}};
end else begin
scl_d <= scl_q;
sda_d <= sda_q;
// ---- activity: sticky, and saturating ----------------------------
if (any_edge) begin
activity <= 1'b1;
if (n_edges != {EDGE_W{1'b1}}) n_edges <= n_edges + 1'b1;
end
// ---- idle: both high, continuously ------------------------------
if (scl_q && sda_q) begin
if (idle_cnt >= IDLE_CLKS[15:0]) idle <= 1'b1;
else idle_cnt <= idle_cnt + 16'd1;
end else begin
idle <= 1'b0;
idle_cnt <= 16'd0;
end
// ---- stuck low: one line, continuously --------------------------
if (!scl_q) begin
if (scl_low_cnt >= STUCK_CLKS[15:0]) scl_stuck_low <= 1'b1;
else scl_low_cnt <= scl_low_cnt + 16'd1;
end else begin
scl_stuck_low <= 1'b0;
scl_low_cnt <= 16'd0;
end
if (!sda_q) begin
if (sda_low_cnt >= STUCK_CLKS[15:0]) sda_stuck_low <= 1'b1;
else sda_low_cnt <= sda_low_cnt + 16'd1;
end else begin
sda_stuck_low <= 1'b0;
sda_low_cnt <= 16'd0;
end
end
end
endmodule // -----------------------------------------------------------------------------
// i2c_bus_health_tb.v
// Independent oracle for i2c_bus_health.
//
// THE FOUR OUTPUTS ARE A DIAGNOSTIC, so the tests are about DISCRIMINATION: each
// one must distinguish the condition it names from the conditions it does not. A
// status bit that is right about a healthy bus and also right about a dead one is
// not a diagnostic, and T3 and T7 are the tests that separate those.
//
// Small parameters (IDLE=8, STUCK=12) so the thresholds are reachable in a short
// simulation. The chapter body gives the real values and the reason for them.
//
// Every wait is a fixed number of clocks; nothing waits on the DUT.
//
// (Verilog-2001 -- the same tests as the SystemVerilog bench.)
// -----------------------------------------------------------------------------
`timescale 1ns/1ps
module i2c_bus_health_tb;
localparam IDLE = 8;
localparam STUCK = 12;
reg clk = 1'b0, rst_n = 1'b0;
reg scl_q = 1'b1, sda_q = 1'b1;
wire idle, scl_stuck, sda_stuck, activity;
wire [15:0] n_edges;
// A SECOND INSTANCE with a 3-bit edge counter, so the saturation behaviour is
// reachable in a short simulation. Reaching the top of the default 16-bit counter
// takes 65535 edges; at 3 bits it takes 7, so T9 can drive past it and require
// the count to STOP rather than wrap. Mutation G10 survived until this existed.
wire nn_idle, nn_scls, nn_sdas, nn_act;
wire [2:0] nn_edges;
integer errors = 0;
integer n;
i2c_bus_health #(.IDLE_CLKS(IDLE), .STUCK_CLKS(STUCK)) dut (
.clk(clk), .rst_n(rst_n), .scl_q(scl_q), .sda_q(sda_q),
.idle(idle), .scl_stuck_low(scl_stuck), .sda_stuck_low(sda_stuck),
.activity(activity), .n_edges(n_edges));
i2c_bus_health #(.IDLE_CLKS(IDLE), .STUCK_CLKS(STUCK), .EDGE_W(3)) narrow (
.clk(clk), .rst_n(rst_n), .scl_q(scl_q), .sda_q(sda_q),
.idle(nn_idle), .scl_stuck_low(nn_scls), .sda_stuck_low(nn_sdas),
.activity(nn_act), .n_edges(nn_edges));
always #5 clk = ~clk;
task step; begin @(posedge clk); @(negedge clk); end endtask
task do_reset;
begin
@(negedge clk); rst_n = 1'b0; scl_q = 1'b1; sda_q = 1'b1;
step; step;
@(negedge clk); rst_n = 1'b1;
end
endtask
// One SCL pulse, as a working controller would produce.
task scl_pulse;
begin
@(negedge clk); scl_q = 1'b0; step; step;
@(negedge clk); scl_q = 1'b1; step; step;
end
endtask
task ck (input [200*8:1] what, input integer g, input integer e);
begin
if (g !== e) begin
$display(" FAIL %0s: got %0d expected %0d", what, g, e);
errors = errors + 1;
end
end
endtask
task ck_idx (input [200*8:1] what, input integer idx,
input integer g, input integer e);
begin
if (g !== e) begin
$display(" FAIL %0s[%0d]: got %0d expected %0d", what, idx, g, e);
errors = errors + 1;
end
end
endtask
initial begin
$display("=== i2c_bus_health: four bits that each distinguish something ===");
// ----------------------------------------------------------------
// T1. RESET REPORTS IGNORANCE, NOT HEALTH. Every bit is 0 out of reset,
// including `idle`. A diagnostic that reset to "healthy" would report a
// good bus before it had looked at one, which is the worst possible
// default for a block whose entire job is to be believed on power-on.
// ----------------------------------------------------------------
@(negedge clk); rst_n = 1'b0; scl_q = 1'b1; sda_q = 1'b1; step; step;
$display("T1 out of reset every bit says 'I have not seen anything yet'");
ck("T1 not idle yet", idle, 0);
ck("T1 no scl stuck", scl_stuck, 0);
ck("T1 no sda stuck", sda_stuck, 0);
ck("T1 no activity yet", activity, 0);
ck("T1 no edges counted", n_edges, 0);
// ----------------------------------------------------------------
// T2. AN IDLE BUS IS RECOGNISED, AFTER EXACTLY IDLE_CLKS. Not before: a bus
// that has been high for two clocks is between bytes, not resting.
// ----------------------------------------------------------------
do_reset;
for (n = 1; n <= IDLE + 2; n = n + 1) begin
step;
ck_idx("T2 idle asserts only after IDLE_CLKS", n, idle,
(n > IDLE) ? 1 : 0);
end
$display("T2 an idle bus is recognised after exactly IDLE_CLKS clocks");
// ----------------------------------------------------------------
// T3. AND AN IDLE BUS IS NOT THE SAME AS A DEAD ONE. This is the test that
// makes `idle` worth having: a disconnected bus with pull-ups fitted looks
// exactly like a healthy resting bus on the two lines, and the ONLY thing
// that separates them is whether anything has ever happened. `activity`
// is that bit, and here it must still be 0 while `idle` is 1.
// ----------------------------------------------------------------
$display("T3 idle and dead look identical on the wires -- 'activity' separates them");
ck("T3 the bus reports idle", idle, 1);
ck("T3 but nothing has ever happened", activity, 0);
ck("T3 and no edges were counted", n_edges, 0);
// ----------------------------------------------------------------
// T4. ONE EDGE IS ENOUGH TO PROVE THE BUS IS CONNECTED, AND IT STICKS. Bring-up
// happens at human speed, so a level true for one clock in ten million is
// not observable. The fact is latched.
// ----------------------------------------------------------------
@(negedge clk); sda_q = 1'b0; step;
$display("T4 one edge proves connectivity, and the fact is kept");
ck("T4 activity asserted", activity, 1);
ck("T4 one edge counted", n_edges, 1);
@(negedge clk); sda_q = 1'b1; step; step; step;
ck("T4 two edges now", n_edges, 2);
ck("T4 and activity is sticky", activity, 1);
// ----------------------------------------------------------------
// T5. A LINE HELD LOW IS REPORTED, AFTER EXACTLY STUCK_CLKS, AND THE TWO LINES
// ARE REPORTED SEPARATELY. Which line is stuck is the whole diagnostic
// value: SDA stuck low with SCL healthy is a device holding the bus, while
// both stuck low is usually no pull-up or no power.
// ----------------------------------------------------------------
do_reset;
@(negedge clk); sda_q = 1'b0;
for (n = 1; n <= STUCK + 2; n = n + 1) begin
step;
ck_idx("T5 sda_stuck asserts only after STUCK_CLKS", n, sda_stuck,
(n > STUCK) ? 1 : 0);
ck_idx("T5 and SCL is never implicated", n, scl_stuck, 0);
end
$display("T5 a stuck line is named individually, after exactly STUCK_CLKS");
ck("T5 and the bus is not idle", idle, 0);
// ----------------------------------------------------------------
// T6. THE OTHER LINE, SYMMETRICALLY. A diagnostic that only worked on one line
// would be half a diagnostic, and the half that is missing is the one that
// matters most: SCL stuck low is the failure that stops everything.
// ----------------------------------------------------------------
do_reset;
@(negedge clk); scl_q = 1'b0;
for (n = 1; n <= STUCK + 2; n = n + 1) begin
step;
ck_idx("T6 scl_stuck asserts only after STUCK_CLKS", n, scl_stuck,
(n > STUCK) ? 1 : 0);
ck_idx("T6 and SDA is never implicated", n, sda_stuck, 0);
end
$display("T6 and the same for SCL, independently");
// ----------------------------------------------------------------
// T7. A WORKING BUS IS NOT REPORTED AS STUCK. The discrimination that matters
// in the other direction: a controller clocking normally holds SCL low for
// part of every bit, and a threshold that counted that would report a
// healthy bus as broken. Eight pulses, and neither stuck bit may assert.
// ----------------------------------------------------------------
do_reset;
for (n = 0; n < 8; n = n + 1) scl_pulse;
$display("T7 normal clocking is not mistaken for a stuck line");
ck("T7 SCL not reported stuck", scl_stuck, 0);
ck("T7 SDA not reported stuck", sda_stuck, 0);
ck("T7 but activity was seen", activity, 1);
ck("T7 with sixteen edges", n_edges, 16);
// ----------------------------------------------------------------
// T8. A STUCK REPORT CLEARS WHEN THE LINE RECOVERS. A latched fault would be
// indistinguishable from a present one, so the engineer could not tell
// whether the fix worked without a reset -- and on a board being probed,
// "did that help" is the question being asked every few seconds.
// ----------------------------------------------------------------
do_reset;
@(negedge clk); sda_q = 1'b0;
for (n = 0; n < STUCK + 2; n = n + 1) step;
ck("T8 reported stuck", sda_stuck, 1);
@(negedge clk); sda_q = 1'b1; step; step;
$display("T8 a stuck report clears when the line recovers");
ck("T8 and cleared when the line came back", sda_stuck, 0);
// ... and the bus then becomes idle again, which is the recovery confirmed.
for (n = 0; n < IDLE + 2; n = n + 1) step;
ck("T8 the bus is idle again", idle, 1);
// ----------------------------------------------------------------
// T9. THE EDGE COUNTER SATURATES RATHER THAN WRAPPING. A counter that wrapped
// would read 0 on a very busy bus and be indistinguishable from a dead one
// -- the exact confusion this block exists to remove. Checked by forcing it
// near the top and driving past it.
// ----------------------------------------------------------------
do_reset;
ck("T9 both counters start at zero", n_edges, 0);
ck("T9 including the narrow one", nn_edges, 0);
// Three pulses = six edges, one short of the 3-bit maximum of 7.
for (n = 0; n < 3; n = n + 1) scl_pulse;
ck("T9 the narrow counter reached six", nn_edges, 6);
scl_pulse; // two more edges: 7, then it must STOP
ck("T9 and saturated at its maximum", nn_edges, 7);
for (n = 0; n < 6; n = n + 1) scl_pulse;
ck("T9 twelve further edges do not wrap it", nn_edges, 7);
// The wide counter, meanwhile, is still counting normally -- which proves the
// saturation is at the parameterised maximum and not at a hard-coded value.
ck("T9 while the wide counter kept counting", n_edges, 20);
// ----------------------------------------------------------------
// T10. `idle` DEASSERTS THE MOMENT THE BUS STOPS RESTING.
//
// Checked from an ASSERTED idle state, because that is the only state in
// which the deassertion is observable -- every stuck-line test above
// begins with a reset, where idle is already 0, so none of them can see
// it. Mutation G11, which never clears idle, survived every other test in
// this list. Placed last rather than inside T3 because T4 depends on T3
// leaving `activity` still deasserted, and driving edges here would have
// spoiled that.
// ----------------------------------------------------------------
do_reset;
for (n = 0; n < IDLE + 2; n = n + 1) step;
ck("T10 precondition: the bus reports idle", idle, 1);
@(negedge clk); scl_q = 1'b0; step;
$display("T10 idle deasserts as soon as the bus stops resting");
ck("T10 idle cleared immediately", idle, 0);
@(negedge clk); scl_q = 1'b1;
for (n = 0; n < IDLE + 2; n = n + 1) step;
ck("T10 and returns after IDLE_CLKS of quiet again", idle, 1);
if (errors == 0) $display("=== i2c_bus_health: ALL CHECKS PASSED ===");
else $display("=== i2c_bus_health: %0d CHECK(S) FAILED ===", errors);
$finish;
end
endmodule -- -----------------------------------------------------------------------------
-- i2c_bus_health.vhd
-- A bring-up aid: what is wrong with this bus, before any protocol is attempted.
-- Behavioural twin of the SystemVerilog and Verilog designs.
--
-- On the morning a board arrives the first question is not "does my state machine
-- work" but "is the bus even there". Four bits answer it at a glance, and the block
-- is small enough to leave in a shipping design.
--
-- IT OBSERVES ONLY and never drives either line:
-- idle both lines HIGH continuously for IDLE_CLKS
-- scl_stuck_low SCL LOW continuously for STUCK_CLKS
-- sda_stuck_low SDA likewise, reported SEPARATELY -- which line is stuck is the
-- whole diagnostic value
-- activity at least one edge seen since reset; STICKY, because bring-up
-- happens at human speed and a one-clock level is not observable
--
-- IT TAKES THE SYNCHRONISED LEVELS, not the pins (Chapter 19.4). A diagnostic that
-- sampled the raw pins would be the one block in the design with a metastability
-- path, which is a poor trade for a status LED.
--
-- IT DELIBERATELY DOES NOT DECODE THE PROTOCOL. Every protocol fact requires the bus
-- to be working first, and a block reporting "no START seen" on a bus with no pull-up
-- answers a question the engineer has not reached. Debug from the boundary inward.
-- -----------------------------------------------------------------------------
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity i2c_bus_health is
generic (
-- Clocks both lines must read HIGH before the bus is called idle. Long enough
-- that a real inter-transfer gap does not count.
IDLE_CLKS : positive := 256;
-- Clocks a line must read LOW before it is called stuck. Must exceed the
-- longest legal LOW, since a stretching target holds SCL down for a long time.
STUCK_CLKS : positive := 4096;
-- Width of the edge counter. A GENERIC rather than a fixed 16, because the
-- saturation behaviour has to be testable: reaching the top of a 16-bit counter
-- takes 65535 edges, which no reasonable simulation drives. Mutation G10, which
-- removes the saturation guard, survived until a narrow instance existed.
EDGE_W : positive := 16
);
port (
clk : in std_logic;
rst_n : in std_logic;
-- The SYNCHRONISED levels from Chapter 19.4, never the pins.
scl_q : in std_logic;
sda_q : in std_logic;
idle : out std_logic;
scl_stuck_low : out std_logic;
sda_stuck_low : out std_logic;
activity : out std_logic;
-- Edges seen on either line. Saturating: a counter that wrapped would read 0 on
-- a busy bus and be indistinguishable from a dead one.
n_edges : out unsigned(EDGE_W-1 downto 0)
);
end entity i2c_bus_health;
architecture rtl of i2c_bus_health is
signal scl_d, sda_d : std_logic := '1';
signal idle_cnt, scl_low_cnt, sda_low_cnt : unsigned(15 downto 0) := (others => '0');
signal edges_i : unsigned(EDGE_W-1 downto 0) := (others => '0');
signal any_edge : boolean;
begin
any_edge <= (scl_q /= scl_d) or (sda_q /= sda_d);
n_edges <= edges_i;
process (clk, rst_n)
begin
if rst_n = '0' then
-- Reset to the values that mean "I have not seen anything yet", which is NOT
-- the same as the values that mean "the bus is healthy". A diagnostic that
-- reset to idle = '1' would report a good bus before it had looked at one.
scl_d <= '1';
sda_d <= '1';
idle <= '0';
scl_stuck_low <= '0';
sda_stuck_low <= '0';
activity <= '0';
idle_cnt <= (others => '0');
scl_low_cnt <= (others => '0');
sda_low_cnt <= (others => '0');
edges_i <= (others => '0');
elsif rising_edge(clk) then
scl_d <= scl_q;
sda_d <= sda_q;
if any_edge then
activity <= '1';
if edges_i /= (edges_i'range => '1') then
edges_i <= edges_i + 1;
end if;
end if;
if scl_q = '1' and sda_q = '1' then
if idle_cnt >= IDLE_CLKS then
idle <= '1';
else
idle_cnt <= idle_cnt + 1;
end if;
else
idle <= '0';
idle_cnt <= (others => '0');
end if;
if scl_q = '0' then
if scl_low_cnt >= STUCK_CLKS then
scl_stuck_low <= '1';
else
scl_low_cnt <= scl_low_cnt + 1;
end if;
else
scl_stuck_low <= '0';
scl_low_cnt <= (others => '0');
end if;
if sda_q = '0' then
if sda_low_cnt >= STUCK_CLKS then
sda_stuck_low <= '1';
else
sda_low_cnt <= sda_low_cnt + 1;
end if;
else
sda_stuck_low <= '0';
sda_low_cnt <= (others => '0');
end if;
end if;
end process;
end architecture rtl; -- -----------------------------------------------------------------------------
-- i2c_bus_health_tb.vhd
-- Independent oracle for i2c_bus_health.
-- Behavioural twin of the SystemVerilog and Verilog benches.
--
-- THE FOUR OUTPUTS ARE A DIAGNOSTIC, so the tests are about DISCRIMINATION: each must
-- distinguish the condition it names from the conditions it does not. A status bit
-- right about a healthy bus and also right about a dead one is not a diagnostic --
-- T3 and T7 are the tests that separate those.
--
-- TWO INSTANCES: the main one, and a second with a 3-bit edge counter so the
-- saturation behaviour is reachable in a short simulation.
-- -----------------------------------------------------------------------------
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity i2c_bus_health_tb is
end entity i2c_bus_health_tb;
architecture sim of i2c_bus_health_tb is
-- NOTE THE NAMES. VHDL is CASE-INSENSITIVE, so a constant called `T_IDLE` is the
-- same identifier as the signal `idle` and the analyser rejects the file with a
-- type error pointing at the signal. The SystemVerilog and Verilog benches use
-- `T_IDLE` and `idle` side by side without complaint, because those languages are
-- case-sensitive -- so this is a porting hazard with no analogue in the original.
constant T_IDLE : integer := 8;
constant T_STUCK : integer := 12;
signal clk : std_logic := '0';
signal rst_n : std_logic := '0';
signal scl_q : std_logic := '1';
signal sda_q : std_logic := '1';
signal idle, scl_stuck, sda_stuck, activity : std_logic;
signal n_edges : unsigned(15 downto 0);
signal nn_idle, nn_scls, nn_sdas, nn_act : std_logic;
signal nn_edges : unsigned(2 downto 0);
signal halt : boolean := false;
begin
clkgen : process
begin
while not halt loop
clk <= '0'; wait for 5 ns;
clk <= '1'; wait for 5 ns;
end loop;
wait;
end process;
dut : entity work.i2c_bus_health
generic map (IDLE_CLKS => T_IDLE, STUCK_CLKS => T_STUCK, EDGE_W => 16)
port map (clk => clk, rst_n => rst_n, scl_q => scl_q, sda_q => sda_q,
idle => idle, scl_stuck_low => scl_stuck, sda_stuck_low => sda_stuck,
activity => activity, n_edges => n_edges);
narrow : entity work.i2c_bus_health
generic map (IDLE_CLKS => T_IDLE, STUCK_CLKS => T_STUCK, EDGE_W => 3)
port map (clk => clk, rst_n => rst_n, scl_q => scl_q, sda_q => sda_q,
idle => nn_idle, scl_stuck_low => nn_scls, sda_stuck_low => nn_sdas,
activity => nn_act, n_edges => nn_edges);
stim : process
variable err : integer := 0;
function b2i (b : std_logic) return integer is
begin
if b = '1' then return 1; else return 0; end if;
end function;
procedure step is
begin
wait until rising_edge(clk);
wait until falling_edge(clk);
end procedure;
procedure do_reset is
begin
wait until falling_edge(clk);
rst_n <= '0'; scl_q <= '1'; sda_q <= '1';
step; step;
wait until falling_edge(clk);
rst_n <= '1';
end procedure;
-- One SCL pulse, as a working controller would produce.
procedure scl_pulse is
begin
wait until falling_edge(clk); scl_q <= '0'; step; step;
wait until falling_edge(clk); scl_q <= '1'; step; step;
end procedure;
procedure ck (what : string; g : integer; e : integer) is
begin
if g /= e then
report " FAIL " & what & ": got " & integer'image(g)
& " expected " & integer'image(e) severity note;
err := err + 1;
end if;
end procedure;
procedure ck_idx (what : string; idx : integer; g : integer; e : integer) is
begin
if g /= e then
report " FAIL " & what & "[" & integer'image(idx) & "]: got "
& integer'image(g) & " expected " & integer'image(e) severity note;
err := err + 1;
end if;
end procedure;
begin
report "=== i2c_bus_health: four bits that each distinguish something ==="
severity note;
-- T1. Reset reports IGNORANCE, not health. A diagnostic that reset to "healthy"
-- would report a good bus before it had looked at one.
wait until falling_edge(clk);
rst_n <= '0'; scl_q <= '1'; sda_q <= '1'; step; step;
report "T1 out of reset every bit says 'I have not seen anything yet'"
severity note;
ck("T1 not idle yet", b2i(idle), 0);
ck("T1 no scl stuck", b2i(scl_stuck), 0);
ck("T1 no sda stuck", b2i(sda_stuck), 0);
ck("T1 no activity yet", b2i(activity), 0);
ck("T1 no edges counted", to_integer(n_edges), 0);
-- T2. An idle bus is recognised after exactly IDLE_CLKS, not before: a bus high
-- for two clocks is between bytes, not resting.
do_reset;
for n in 1 to T_IDLE + 2 loop
step;
if n > T_IDLE then ck_idx("T2 idle asserts only after IDLE_CLKS", n, b2i(idle), 1);
else ck_idx("T2 idle asserts only after IDLE_CLKS", n, b2i(idle), 0); end if;
end loop;
report "T2 an idle bus is recognised after exactly IDLE_CLKS clocks"
severity note;
-- T3. And an idle bus is not the same as a DEAD one. A disconnected bus with
-- pull-ups fitted looks exactly like a healthy resting bus on the two
-- lines; the only thing that separates them is whether anything has ever
-- happened, and `activity` is that bit.
report "T3 idle and dead look identical on the wires -- 'activity' separates them"
severity note;
ck("T3 the bus reports idle", b2i(idle), 1);
ck("T3 but nothing has ever happened", b2i(activity), 0);
ck("T3 and no edges were counted", to_integer(n_edges), 0);
-- T4. One edge proves the bus is connected, and the fact STICKS: bring-up
-- happens at human speed, so a one-clock level is not observable.
wait until falling_edge(clk); sda_q <= '0'; step;
report "T4 one edge proves connectivity, and the fact is kept" severity note;
ck("T4 activity asserted", b2i(activity), 1);
ck("T4 one edge counted", to_integer(n_edges), 1);
wait until falling_edge(clk); sda_q <= '1'; step; step; step;
ck("T4 two edges now", to_integer(n_edges), 2);
ck("T4 and activity is sticky", b2i(activity), 1);
-- T5. A line held low is reported after exactly STUCK_CLKS, and the two lines
-- are reported SEPARATELY -- which line is stuck is the diagnostic value.
do_reset;
wait until falling_edge(clk); sda_q <= '0';
for n in 1 to T_STUCK + 2 loop
step;
if n > T_STUCK then ck_idx("T5 sda_stuck asserts only after STUCK_CLKS", n, b2i(sda_stuck), 1);
else ck_idx("T5 sda_stuck asserts only after STUCK_CLKS", n, b2i(sda_stuck), 0); end if;
ck_idx("T5 and SCL is never implicated", n, b2i(scl_stuck), 0);
end loop;
report "T5 a stuck line is named individually, after exactly STUCK_CLKS"
severity note;
ck("T5 and the bus is not idle", b2i(idle), 0);
-- T6. The other line, symmetrically. SCL stuck low is the failure that stops
-- everything, so a diagnostic that only worked on SDA would be missing the
-- half that matters most.
do_reset;
wait until falling_edge(clk); scl_q <= '0';
for n in 1 to T_STUCK + 2 loop
step;
if n > T_STUCK then ck_idx("T6 scl_stuck asserts only after STUCK_CLKS", n, b2i(scl_stuck), 1);
else ck_idx("T6 scl_stuck asserts only after STUCK_CLKS", n, b2i(scl_stuck), 0); end if;
ck_idx("T6 and SDA is never implicated", n, b2i(sda_stuck), 0);
end loop;
report "T6 and the same for SCL, independently" severity note;
-- T7. A WORKING bus is not reported as stuck. A controller clocking normally
-- holds SCL low for part of every bit, and a threshold that counted that
-- would report a healthy bus as broken.
do_reset;
for n in 1 to 8 loop scl_pulse; end loop;
report "T7 normal clocking is not mistaken for a stuck line" severity note;
ck("T7 SCL not reported stuck", b2i(scl_stuck), 0);
ck("T7 SDA not reported stuck", b2i(sda_stuck), 0);
ck("T7 but activity was seen", b2i(activity), 1);
ck("T7 with sixteen edges", to_integer(n_edges), 16);
-- T8. A stuck report CLEARS when the line recovers. A latched fault would be
-- indistinguishable from a present one, so the engineer could not tell
-- whether the fix worked -- and "did that help" is the question being asked
-- every few seconds on a board being probed.
do_reset;
wait until falling_edge(clk); sda_q <= '0';
for n in 1 to T_STUCK + 2 loop step; end loop;
ck("T8 reported stuck", b2i(sda_stuck), 1);
wait until falling_edge(clk); sda_q <= '1'; step; step;
report "T8 a stuck report clears when the line recovers" severity note;
ck("T8 and cleared when the line came back", b2i(sda_stuck), 0);
for n in 1 to T_IDLE + 2 loop step; end loop;
ck("T8 the bus is idle again", b2i(idle), 1);
-- T9. The edge counter SATURATES rather than wrapping, checked on the 3-bit
-- instance because the 16-bit one cannot be driven to its maximum. A counter
-- that wrapped would read 0 on a busy bus -- the exact confusion this block
-- exists to remove.
do_reset;
ck("T9 both counters start at zero", to_integer(n_edges), 0);
ck("T9 including the narrow one", to_integer(nn_edges), 0);
for n in 1 to 3 loop scl_pulse; end loop;
ck("T9 the narrow counter reached six", to_integer(nn_edges), 6);
scl_pulse;
ck("T9 and saturated at its maximum", to_integer(nn_edges), 7);
for n in 1 to 6 loop scl_pulse; end loop;
ck("T9 twelve further edges do not wrap it", to_integer(nn_edges), 7);
ck("T9 while the wide counter kept counting", to_integer(n_edges), 20);
-- T10. `idle` deasserts the moment the bus stops resting. Checked from an
-- ASSERTED idle state, the only state in which the deassertion is
-- observable: every stuck-line test above begins with a reset, where idle
-- is already 0. Mutation G11, which never clears idle, survived every
-- other test in this list.
do_reset;
for n in 1 to T_IDLE + 2 loop step; end loop;
ck("T10 precondition: the bus reports idle", b2i(idle), 1);
wait until falling_edge(clk); scl_q <= '0'; step;
report "T10 idle deasserts as soon as the bus stops resting" severity note;
ck("T10 idle cleared immediately", b2i(idle), 0);
wait until falling_edge(clk); scl_q <= '1';
for n in 1 to T_IDLE + 2 loop step; end loop;
ck("T10 and returns after IDLE_CLKS of quiet again", b2i(idle), 1);
if err = 0 then
report "=== i2c_bus_health: ALL CHECKS PASSED ===" severity note;
else
report "=== i2c_bus_health: " & integer'image(err) & " CHECK(S) FAILED ==="
severity note;
end if;
halt <= true;
wait;
end process;
end architecture sim;Three survivors, and what each one changed
| # | mutation | verdict |
|---|---|---|
| G01 | resets to idle = 1 | KILLED (9) |
| G02 | activity resets asserted | KILLED (2) |
| G03 | idle needs only one high clock | KILLED (8) |
| G04 | idle ignores SDA | KILLED (1) |
| G05 | stuck threshold ignored | KILLED (12) |
| G06 | the two stuck bits wired together | KILLED (5) |
| G07 | a stuck report latches for ever | KILLED (1) |
| G08 | activity is a level, not sticky | EQUIVALENT as first written, then killed (2) |
| G09 | edge detector sees only SCL | KILLED (4) |
| G10 | edge counter wraps instead of saturating | survived, then killed (2) |
| G11 | idle never clears | survived, then killed (1) |
G08 was an equivalent mutation I had written badly. activity <= 1'b1 inside if (any_edge) became activity <= any_edge — but inside that branch any_edge is 1, so the two are identical and no stimulus could distinguish them. The mutation did not express "not sticky" at all. Re-expressed by moving the assignment outside the condition, it dies in two checks. A mutation that is accidentally equivalent is not evidence of a strong bench; it is a mutation that tested nothing.
G10 exposed a property that could not be tested, and the fix was in the design. The edge counter saturates rather than wrapping — because a counter that wrapped would read 0 on a busy bus and be indistinguishable from a dead one, which is the exact confusion this block exists to remove. Reaching the top of a 16-bit counter takes 65535 edges, which no reasonable simulation drives, so the property was untestable and the mutation survived.
The repair was to make the counter width a parameter, and instantiate a second copy at EDGE_W = 3 alongside the first. Seven edges then reach the maximum, and T9 drives twelve more and requires the count to stay there — while the wide instance keeps counting, which proves the saturation is at the parameterised maximum rather than a hard-coded value.
G11 survived because of where the test looked. idle never clearing was undetectable, because every stuck-line test begins with do_reset, where idle is already 0 — so none of them could observe a failure to clear. The repair is T10, which establishes idle = 1 first and then drives a line low. It had to go at the end of the bench rather than inside T3, because T4 depends on T3 leaving activity still deasserted, and driving edges earlier would have spoiled that.
4. What to Probe, and Why Not Just the State Machine
An ILA has a finite capture width and depth, so the probe set is a choice. The temptation is to capture the FSM state, because that is where the design's logic is.
A useful set, ordered by layer, and roughly what it costs:
| probe | layer | what its absence would hide |
|---|---|---|
scl_pin, sda_pin (synchronized) | boundary | the pins are not connected, or not clocked |
scl_q, sda_q | 19.4 | synchronization is broken |
filtered level, n_rejected | 19.5 | the filter is eating real edges, or the bus is noisy |
scl_rise, scl_fall, sda_rise, sda_fall | 18.2 | edges are not one cycle, or are double-counted |
start_pulse, stop_pulse | 18.3 | framing is not detected |
selected, dir_read | 18.4 | the address does not match, or direction is wrong |
sda_drive_low, scl_drive_low | 19.1 | the design's intent, which is what a pin-level disagreement is measured against |
n_sda_conflict | 18.11 | two internal drivers fighting |
stretching, n_aborts | 18.10 / 18.11 | a stretch or a timeout nobody noticed |
pointer, n_writes, n_refused | 18.9 | application state is not advancing |
The row in bold is the one people leave out, and it is the most valuable single probe in the list. Section 5 is why.
5. The External Analyzer and the ILA Disagree — and That Is the Point
Two instruments, two different questions:
external analyser → what HAPPENED ON THE WIRE
ILA → what the FPGA's LOGIC BELIEVED happenedBoth can be right while disagreeing, and the disagreement is where the fault is.
| analyser says | ILA says | where the fault is |
|---|---|---|
| clean I²C, NACK on the ninth bit | sda_drive_low was asserted | between intent and pin: OE polarity, pin constraint, I/O standard (19.1, 19.2) |
| clean I²C | the FSM never left IDLE | the front end: pins, synchronizer or filter (19.4, 19.5) |
| a START | start_pulse never fired | framing input, or the filter ate the edge |
| SDA held low for ever | nothing internally is driving | another device on the bus, or no pull-up |
| nothing at all on the bus | start_pulse fired | a phantom event from internal state — 19.4's reset-value bug |
| correct data | occasional wrong byte, internally consistent | metastability — 19.6's false-path failure |
The last two rows are the ones that are only findable this way. A phantom START and a metastability failure both produce internally coherent logic behaviour, so an ILA alone shows a design doing something reasonable, and an analyser alone shows a bus that is fine. Only the contradiction between them names the fault.
6. What This Chapter Cannot Do
7. ASIC Contrast, Briefly
An ASIC has no ILA, and the substitution is not like-for-like. Internal visibility comes from infrastructure decided long before bring-up: scan chains, a debug bus, trace macrocells — and whatever was not planned in is not observable at any price. An FPGA lets you add a probe and rebuild in minutes; an ASIC lets you add a probe in the next revision.
The practical consequence is that the ASIC equivalent of this chapter happens during design: deciding which signals reach a debug register, and keeping the diagnostic counters this module has been accumulating — n_rejected, n_sda_conflict, n_aborts, n_edges — because a counter readable over the functional interface is the one piece of internal visibility that survives into silicon.
Which is a good argument for those counters even on an FPGA, where the ILA makes them feel redundant. They are not redundant: they work in the field, on a customer's board, over the existing interface, with no capture and no cable.
8. Focused Verification Insight
A diagnostic block needs discrimination tests in both directions, and this is the transferable point. A status bit only earns trust if a test shows it asserting when it should and a test shows it staying quiet when it should. T3 and T7 are that pair, and a bench with only one of them is measuring half a diagnostic.
Module 20's environment can assert on these counters as invariants. n_sda_conflict == 0 and n_rejected == 0 should hold across a clean run; a soak test that leaves either non-zero has found something, and it is worth knowing whether that something is the DUT or the bench's own stimulus integrity.
Coverage worth collecting is the four-way cross of idle × activity, plus each stuck bit asserted and cleared. That is sixteen cells of which the interesting ones are the two that look identical on the wire — which is exactly the table in Section 2.
9. Misconceptions
10. Debugging
The board works on the bench and fails in the rack, and the ILA says everything is fine
Pitfall — a diagnostic counter nobody read, and a margin nobody measured
// A target that passes every RTL test, every bring-up step, and a week of soak
// testing on the bench. Deployed into a rack alongside a switching power supply and
// three motor drives. In the rack it fails: roughly one transfer in fifty thousand
// returns a wrong byte. No NACKs, no protocol errors, no timeouts.
//
// The design instantiates the filter from Chapter 19.5 with N_SAMP = 4 at 50 MHz,
// which is correct: it rejects disturbances up to 50 ns as the specification allows,
// and 4 clocks is 80 ns of latency against a 2500 ns bit period.
//
// It also brings out n_rejected, exactly as Chapter 19.5 argues it should.
//
// Nobody connected n_rejected to anything. It is an output on the block, wired to an
// unused net at the top level, optimised away by synthesis. The information the
// design was built to provide was never made readable.On the bench: flawless. Millions of transfers, no errors, n_rejected -- had anyone been able to read it -- would have been 0.
In the rack: about one transfer in fifty thousand returns a wrong byte. Internally consistent: the target's own view of the transaction is coherent, the register pointer is where it should be, no error bit is set anywhere. The controller does not retry because nothing told it to.
An ILA capture triggered on a data mismatch shows nothing unusual. The framing is right, the address matched, the byte arrived -- just the wrong byte. Every probe in the capture agrees with every other probe.
That signature -- correct protocol, occasional wrong data, internally coherent logic, no error condition, environment-dependent -- has two candidates from Chapter 19.7: metastability (Class 4/5), or a disturbance being accepted as a real edge (Class 7). Both are statistical and neither sets a flag.
The instrument that separates them is the one that was not connected.
The bus was picking up disturbances from the motor drives -- some of them wider than 50 ns, and therefore wider than the filter was designed to reject. A disturbance of 4 or more clocks at 50 MHz is 80 ns or more, which the filter ACCEPTS, because that is what N_SAMP = 4 means. The filter was working exactly as specified; the specification's 50 ns allowance is not a promise about this rack.
Two things are worth separating here, because they are different failures.
The ENGINEERING problem is that the installation is noisier than the electrical assumptions behind the filter threshold, and the fixes are electrical and architectural: better routing and shielding, a lower bus speed with more filtering (the window from Chapter 19.5 is 4..30 at 50 MHz, so there is plenty of room), or a bus buffer to isolate the segment.
The PROCESS problem is the one that cost the time, and it is the transferable one: the design had already been built to measure this, and the measurement was thrown away. n_rejected would have been non-zero on the very first rack test -- large and obviously non-zero -- which would have pointed at Class 7 immediately instead of at a week of metastability hypotheses. On the bench it would have read 0, which is the comparison that makes the rack figure mean something.
The fix is two lines of top-level wiring into a readable register, and the general rule is: A DIAGNOSTIC OUTPUT THAT IS NOT CONNECTED TO ANYTHING IS NOT A DIAGNOSTIC. Synthesis will remove it silently and the design will look identical. Every counter this module argues for -- n_rejected, n_sda_conflict, n_aborts, n_edges -- is worth exactly as much as the path from it to somewhere a human can read it, which is also why Chapter 19.7's ASIC contrast matters: on real hardware in the field, a readable counter is the only internal visibility you have.
11. Reason It Through
12. Questions
13. What This Chapter Settled
Bring-up runs from the physical boundary inward, and the first four steps involve no FPGA logic. Four status bits answer the steps you re-check constantly, and they are designed around discrimination rather than around reporting: idle and activity together separate a healthy bus from a dead one, and the two stuck bits are reported individually because which line is stuck is the diagnosis.
The block is verified in three languages, nineteen mutations killed. Three of them survived first: one was an equivalent mutation written badly, one needed the design changed to make a property testable, and one needed a test that observed a deassertion from an asserted state. Three survivors, three different correct responses.
The probe set spans layers rather than concentrating on the interesting logic, and sda_drive_low is the most valuable single probe because it is the only one that lets a pin observation be compared against an intent. When the analyser and the ILA disagree, the fault is usually between them — and two failure modes are findable no other way.
What remains is composition. Every piece now exists — output stage, wrapper, pull-up, synchronizer, filter, constraints, and a diagnostic — and nothing has yet wired them to the verified cores of Modules 17 and 18 in one place, in both directions. Chapter 19.9 is that assembly.
Continue learning
Related tutorials
- Related topic
FPGA Debug, ILA Capture and Board Bring-Up
A capture buffer whose trigger schedules the stop rather than the start, the off-by-one in the ring unwrap that only unique probe data exposes, and a bring-up sequence that orders the measurements so each one is interpretable.
- Related topic
FPGA as Master and as Target — Integration Patterns
Assembles everything: Module 18's verified target behind Module 19's pad, synchronizer and filter, proven end to end on a wired-AND bus in three languages. Shows the controller-side shape on Module 17's real interface, where a soft CPU attaches, and closes with two mutations that cannot be killed in simulation — the module's thesis stated as evidence.
- Related topic
Simulation vs Synthesized Hardware — Where Behavior Diverges
Eight classes of mismatch that pass in RTL simulation and fail on a board, each with why simulation passes, what hardware does instead, and what evidence exposes it. Includes a bus model with a rise time, and an experiment running Module 18's verified target against it whose result is not the expected one.
- Related topic
Oversampling, Edge Detection and Spike Filtering
A synchronizer reports every disturbance faithfully, including the ones the specification lets you ignore. Builds an agreement-counter filter whose threshold is one number with one meaning, works out what that number must be at a given system clock, and shows why the upper bound — not rejecting legal traffic — matters as much as the lower one.
