I²C · Module 23
A Systematic I²C Debug Workflow
Turns “it doesn’t work” into an ordered narrowing: separate observation from interpretation, suspect the instrument before the design, eliminate a layer, then form competing hypotheses and pick the measurement that discriminates them. Builds the instrument step four depends on — three observation points, one layer verdict — in SystemVerilog and VHDL, and reports the saturation bound a mutation proved had never been tested.
Twenty-two modules of I²C have produced a lot of knowledge and one significant gap. You know the electricals, the framing, the timing table, the transactions, the master and slave architectures, the FPGA reality, the verification architecture and where each protocol rule is asserted. What you have not been given is a procedure — the thing to do when a board, or a simulation, or a customer's log, says the bus does not work.
This module is that procedure, and this chapter is its skeleton. Its subject is not I²C. Its subject is how to narrow a failure from an entire system to one cause using evidence, and how to know when you have actually done so rather than merely found something to blame.
The distinction that makes this possible is small enough to state in two lines and hard enough to practise that most of the module is about practising it.
1. Two Sentences That Are Not the Same Sentence
Here are two statements about the same bus.
OBSERVATION SDA remains LOW after the point at which a STOP was expected.
INTERPRETATION The target's state machine is stuck.The first is a measurement. Given the instrument, it is either true or false, and repeating the measurement gives the same answer. The second is a hypothesis — one of several explanations that all produce that measurement, and on any given board probably not the right one.
How many explanations? Here are eight, all consistent with the observation, none eliminated by it:
| candidate cause | layer |
|---|---|
| the controller never released SDA, so no STOP was generated at all | controller logic |
| the controller released SDA but its pad did not — an output-enable polarity error | controller output path |
| the target is mid-byte and still driving a zero from its shift register | target logic |
| the target's FSM is genuinely stuck in a state that asserts drive-low | target logic |
| a third participant on the bus is holding SDA down | another device |
| the pull-up is absent or far too weak, so the line never rises | electrical |
| the bus capacitance is high enough that the rise has not completed yet | electrical |
| the analyser's ground reference is wrong and the line is actually high | the instrument |
Eight candidates across five layers, and the last one is worth sitting with: the instrument is a candidate. A debug procedure that cannot suspect its own measurement apparatus will eventually spend a week on a design that was correct.
2. What "The Bus" Actually Is
The reason a single observation is consistent with so many causes is structural. Every I²C symptom is observed on a wire, and the wire is the last element in a chain:
One arrow is deliberately missing from that figure, and its absence is worth naming. Nothing in the diagram returns from the wire to the engine — but a device's only way to learn the line's state is to read it back through its own input path, synchronizer and filter (Chapter 19.4, 19.5). That path is a fourth place a fault can hide, and it has an awkward property for a debug procedure: if you use the design's own input path as your observation of line, the instrument cannot detect a fault in that input path, because the fault is now upstream of the measurement. Section 8 returns to it.
Read the figure as a set of seams, because a seam is where a fault can hide:
engine → intent. A logic bug. The design decided the wrong thing. This is the region every RTL engineer searches first, and it is one of four.
intent → pad. The decision did not reach the pin. Inverted output enable, a pin constraint naming the wrong package pin, a pad shared with another function, a bank voltage that cannot pull the far end below its threshold. Chapter 19.1 and 19.2 built this seam; nothing in Modules 17, 18, 20 or 21 simulates it.
pad → wire. The pull-down did not reach the net, or the net did not respond in the time assumed. A broken connection; a rise time longer than the sampler's patience (Chapter 19.3).
wire → everyone. The level is the wired-AND of every participant, so the wire cannot attribute itself. This is not a defect of I²C, it is the mechanism clock stretching and arbitration are built on (Chapter 13.3) — and it is the reason "SDA is low" names no device.
3. The Workflow
Here is the procedure. It is ordered, and the order is the content — each step is cheap, and each one eliminates a region that the next step would otherwise have to search.
Step 1 — state the observation
Write down what was measured, in terms of a level, an edge, a byte value or a count, with the instrument named. If the sentence contains a mechanism, a component name or the word "because", it is an interpretation and it belongs in step 5.
A useful test: could someone who has never seen this design reproduce the measurement from your sentence? "SDA remains low 4 µs after the falling edge of the ninth SCL pulse, measured on the target's SDA pin with a 100 MHz analyser" passes. "The slave hangs" does not.
Step 2 — suspect the instrument first
This step is placed second, before anything about the design, because it is the cheapest and because a wrong instrument invalidates every later step. The questions are concrete:
- Is the ground reference shared between the probe and the board?
- Is the analyser's threshold appropriate for the bus voltage? A 1.8 V bus read with a 3.3 V threshold reads as permanently low.
- Is the sample rate high enough to see the shortest pulse of interest? An analyser at 4× the bit rate cannot distinguish a spike from an edge.
- In simulation: is the monitor sampling on the edge the protocol defines, or on the one that was convenient? (Chapter 20.7 is the long version.)
- In simulation: is the file being run the file that was edited? A build system that silently uses a stale object is the most demoralising instrument fault there is.
Step 3 — establish that the bus is alive
Before any protocol reasoning, three measurements that take a minute and eliminate an entire layer:
1. With no traffic, do SDA and SCL both sit at the supply rail?
NO -> something is holding a line, or there is no pull-up. Stop here;
no protocol reasoning applies to a bus that cannot idle.
YES -> a pull-up exists on each line and nothing is stuck.
2. When the controller drives a line low, does it reach a valid low level?
NO -> the pull-up is too strong, the pad cannot sink enough, or the
pull-down is not electrically reaching the net.
YES -> the pull-down path works.
3. On release, how long does the line take to reach a valid high?
Compare against tr for the speed mode in use (Table 10, UM10204).
LONGER -> an electrical problem, and every timing symptom above this
layer is a consequence rather than a cause.Note what step 3 does not do: it does not look at a single byte. A bus that cannot idle high, or cannot reach a valid low, or rises too slowly, will produce protocol symptoms in unlimited variety, and analysing those symptoms is wasted work until this layer is clean. Chapter 23.5 is this step in full.
Step 4 — compare drive intent with the resolved wire
This is the step that partitions the remaining space, and it needs an instrument. It is the subject of the rest of this chapter.
Step 5 — form at least two hypotheses
Not one. The discipline is mechanical: write the candidates down, and require a second one even when the first feels obvious. A single hypothesis cannot be tested, only confirmed — you will find evidence consistent with it, because evidence consistent with a plausible explanation is abundant.
Step 6 — choose the discriminating measurement
Given two hypotheses, the useful question is not "what else can I measure?" but:
What measurement would give a different result depending on which hypothesis is true?
A measurement both hypotheses predict identically costs time and returns nothing, however interesting it looks. This is the single highest-leverage habit in the module, and Chapter 23.4 is built entirely out of it.
Step 7 — root cause, correction, and proof
A correction is not proven by the symptom disappearing. The symptom disappearing is consistent with the fix, and also with having perturbed the timing enough to hide it. Proof requires a check that fails before the change and passes after, for a reason connected to the root cause — which is exactly what the negative proof discipline in Modules 17 through 22 has been building.
4. The Instrument Step 4 Needs
Step 4 asks whether drive intent matches the resolved wire. Answering it requires three observation points, and the whole value of the step is that the disagreements between them are layer-specific:
| observation point | where it comes from | what a disagreement with the next point means |
|---|---|---|
my_drive_low — the engine's decision | an ILA signal inside the FPGA; a simulation probe | intent is not reaching the pad: output path |
pin_drive_low — what the pad asserts | a scope on the pin; a loopback input | the pull-down is not reaching the net: electrical |
line — the resolved bus, read back | an analyser; the design's own input path | another participant is holding the line |
So the block below is not a checker. It knows nothing about I²C, asserts nothing about correctness, and would be equally useful on any open-drain bus. It answers exactly one question — do the three points agree, and if not, where do they first stop agreeing — and reports a class, which is a statement about a layer rather than about a device.
// -----------------------------------------------------------------------------
// i2c_intent_probe.sv
// The instrument Chapter 23.1's workflow runs on: one bus line, three observation
// points, and a verdict that says WHICH LAYER the evidence implicates.
//
// A debug session begins with a symptom on the wire. The wire is the last thing in
// a chain, so a symptom there is consistent with a fault anywhere along it:
//
// protocol engine -> drive intent -> pad / OE / pin -> the shared wire
// ^
// other participants pull here
//
// Reading the wire alone cannot tell those apart. Reading THREE points can, because
// the disagreements between them are layer-specific:
//
// my_drive_low what this device's logic decided (an ILA signal, inside the FPGA)
// pin_drive_low what the pad is actually asserting (a scope or a loopback pin)
// line the resolved bus, read back (an analyser, or the input path)
//
// This block is not a checker for a design under test. It has no idea what the
// protocol should be doing and never asserts anything about correctness. It answers
// one question -- do the three observation points agree, and if not, WHERE do they
// first stop agreeing -- and the answer partitions the search space.
//
// THE THREE VERDICTS, in the order the block resolves them, which is the order of
// the chain above. The ordering is not arbitrary: a fault nearer the logic makes
// every observation further out misleading, so the innermost disagreement must be
// reported rather than the symptom it produces.
//
// 1. out_path_fault intent != pin. The logic's decision is not reaching the
// pad. Inverted output enable, wrong pin constraint, a pad
// held by another function. STRICTLY LOCAL, and nothing about
// the protocol engine is implicated.
//
// 2. pulldown_fault the pad asserts LOW, the wire reads HIGH. No legal I2C
// configuration produces this: a device pulling down that the
// wire does not follow means the pull-down is not electrically
// reaching the net. Broken connection, a pin that is not the
// one on the bus, a bank voltage that cannot pull below the
// far end's threshold.
//
// 3. foreign_hold this device has released, the pad has released, the wire is
// nonetheless LOW, and the settling window has expired. Some
// OTHER participant is holding the line. This is a legal and
// frequent I2C state -- clock stretching and arbitration are
// both built on it -- so it is a verdict about OWNERSHIP, not
// an error.
//
// WHY settle_window EXISTS, and why omitting it is the classic false positive here.
// "Released and still low" is the primitive that clock stretching, arbitration and
// a stuck bus are all detected with. It is also exactly what a healthy bus looks
// like for tr after the last device lets go, because a pull-up charges a capacitor
// and that takes time (Chapter 19.3's tr = 0.8473 * Rp * Cb). An instrument without
// this input reports a foreign holder on every single release of a real bus. The
// caller drives it from whatever it knows about the rise time -- the bus model's
// `rising` output in simulation, a counter sized from tr on hardware.
//
// FIRST DIVERGENCE. The latched cycle and class are the point of the instrument.
// The first cycle at which the observation points disagree is far more localising
// than any later one, because by the time a protocol failure is visible the design
// has usually made several further decisions on top of the wrong one.
//
// SYNTHESISABLE. This is intended to be instantiated in an FPGA next to the block
// being brought up, with its counters read over a debug interface -- the same role
// as Chapter 19.8's bus-health block. It contains no delays, no initial blocks and
// no non-synthesisable constructs. It is also useful in simulation, where all four
// inputs can be driven independently to reach combinations a correct bus model can
// never produce, which is how its own test exhausts them.
// -----------------------------------------------------------------------------
module i2c_intent_probe #(
// Width of the saturating event counters and of the cycle stamp. Saturating,
// not wrapping: a counter that wraps to zero reports "no events" for a fault
// that occurred a great many times, which is the worst possible failure for a
// diagnostic register.
parameter integer CNT_W = 16
) (
input wire clk,
input wire rst_n,
// Observation point 1 -- the protocol engine's decision. 1 = intends to pull LOW.
input wire my_drive_low,
// Observation point 2 -- what the pad is asserting. 1 = the pad is pulling LOW.
input wire pin_drive_low,
// Observation point 3 -- the resolved wire, read back. 0 = the bus is LOW.
input wire line,
// 1 while a released line is allowed to still read LOW (the pull-up is charging).
input wire settle_window,
// Sticky verdicts. Sticky because a debug register that clears itself reports the
// state of the bus at the moment it was read rather than what happened.
output reg out_path_fault,
output reg pulldown_fault,
output reg foreign_hold,
// Where the three points FIRST stopped agreeing.
output reg [CNT_W-1:0] first_div_cyc,
output reg [1:0] first_div_class, // 0 none, 1 out-path, 2 pull-down, 3 foreign
// Saturating event counts, one per verdict, plus the cycles on which all three
// points agreed. `n_consistent` is what makes a zero in the others meaningful:
// three zero fault counts and a zero consistent count means the instrument was
// never clocked, which is a different finding from a clean bus.
output reg [CNT_W-1:0] n_out_path,
output reg [CNT_W-1:0] n_pulldown,
output reg [CNT_W-1:0] n_foreign,
output reg [CNT_W-1:0] n_consistent,
// Free-running cycle counter, exposed so a reader can tell how long the window
// was and place first_div_cyc in it. Saturating for the same reason.
output reg [CNT_W-1:0] cyc
);
localparam [CNT_W-1:0] CNT_MAX = {CNT_W{1'b1}};
// ---------------------------------------------------------------------------
// The classification, as conditions over the three observation points. Written
// as combinational conditions rather than as tests on the sticky flags, because
// a sticky flag says "this happened once" and the classification is about THIS
// cycle. Deciding a priority on a latched flag classifies every later cycle by
// the first fault ever seen.
// ---------------------------------------------------------------------------
// An XOR, not a !== comparison. `!==` is a simulation-only operator: it would
// make this block unsynthesisable, and it would change the VHDL and Verilog-2001
// ports of the same design into something that is not the same design.
wire c_out_path = my_drive_low ^ pin_drive_low;
wire c_pulldown = pin_drive_low && line;
wire c_foreign = !pin_drive_low && !my_drive_low && !line && !settle_window;
// Priority: innermost layer first. See the header -- a fault between the logic
// and the pad makes the pad-to-wire and wire-to-others observations meaningless,
// so it must be the reported class.
wire [1:0] cls = c_out_path ? 2'd1
: c_pulldown ? 2'd2
: c_foreign ? 2'd3
: 2'd0;
always @(posedge clk) begin
if (!rst_n) begin
out_path_fault <= 1'b0;
pulldown_fault <= 1'b0;
foreign_hold <= 1'b0;
first_div_cyc <= {CNT_W{1'b0}};
first_div_class <= 2'd0;
n_out_path <= {CNT_W{1'b0}};
n_pulldown <= {CNT_W{1'b0}};
n_foreign <= {CNT_W{1'b0}};
n_consistent <= {CNT_W{1'b0}};
cyc <= {CNT_W{1'b0}};
end else begin
if (cyc != CNT_MAX) cyc <= cyc + 1'b1;
// One counter increments per cycle, chosen by the same priority as the
// class. The counts therefore sum to the number of cycles observed, which
// is a property the test checks -- a classifier whose bins do not sum to
// the sample count is either double-counting or dropping cycles.
case (cls)
2'd1: begin
out_path_fault <= 1'b1;
if (n_out_path != CNT_MAX) n_out_path <= n_out_path + 1'b1;
end
2'd2: begin
pulldown_fault <= 1'b1;
if (n_pulldown != CNT_MAX) n_pulldown <= n_pulldown + 1'b1;
end
2'd3: begin
foreign_hold <= 1'b1;
if (n_foreign != CNT_MAX) n_foreign <= n_foreign + 1'b1;
end
default: begin
if (n_consistent != CNT_MAX) n_consistent <= n_consistent + 1'b1;
end
endcase
// Latch the FIRST divergence only. The guard is on first_div_class, not on
// a separate "seen" flag, because class 0 is exactly the "nothing yet"
// state and a second flag could disagree with it.
if (first_div_class == 2'd0 && cls != 2'd0) begin
first_div_class <= cls;
first_div_cyc <= cyc;
end
end
end
endmoduleWhy the priority order is the design
Three conditions can be true at once, and the block reports one class. Which one it reports is the entire engineering content of the block, because the class is what a reader acts on.
The order is innermost seam first: out-path, then pull-down, then foreign hold. The justification is causal rather than aesthetic — a fault nearer the logic makes every observation further out misleading. If the pad is not doing what the engine asked, then what the wire shows is a consequence of the output-path fault, and reporting "another device is holding the line" would send a reader to the bus when the fault is inside the chip.
There is a second, subtler property of that ordering, and a mutation found it. Because out_path_fault takes priority, the pull-down and foreign conditions are only ever used when intent and pin agree — which makes several ways of writing them indistinguishable. Section 6 works that through, because it is the more interesting result of the two.
Why settle_window exists
"Released and still low" is the primitive on which clock stretching, arbitration loss and a stuck bus are all detected — Chapter 17.10 builds all three out of one comparator. It is also precisely what a healthy bus looks like for tr after the last device lets go, because a pull-up charges a capacitance and that takes time.
An instrument without this input therefore reports a foreign holder on every single release of every real bus. That is not a small inconvenience: a diagnostic that fires constantly gets switched off, and a diagnostic that is switched off is worth less than none, because its silence is mistaken for evidence.
The caller drives it from whatever it knows about the rise time — a bus model's settling output in simulation, a counter sized from tr = 0.8473 · Rp · Cb on hardware.
5. The Oracle
The bench has five parts, and two of them exist because of things that went wrong while writing it.
// -----------------------------------------------------------------------------
// i2c_intent_probe_tb.sv
// An independent oracle for the probe, in four parts.
//
// T1 the exhaustive truth table -- all sixteen combinations of the four inputs,
// each checked against a classification written out INDEPENDENTLY as a case
// statement over the input vector rather than by re-deriving the design's
// expression. A reference model that restates the design's condition proves
// only that the expression equals itself.
//
// T2 the bin sum -- over a long pseudo-random stimulus, the four counters must
// sum to the number of clocked cycles. A classifier whose bins do not sum to
// the sample count is double-counting or dropping cycles, and neither shows
// up as a wrong individual count.
//
// T3 first divergence -- that the latch takes the FIRST anomaly and not a later
// one, including the case where a different, higher-priority class occurs
// afterwards. This is the property the instrument exists for.
//
// T4 negative proof -- that this bench can fail. Every check is exercised once
// against a deliberately wrong expectation, and the count of detections is
// itself asserted. A checker whose ability to fail is not demonstrated is
// indistinguishable from no checker at all.
//
// Bounded: every wait is a fixed number of clock edges. There is no wait on a DUT
// condition anywhere, so a broken probe cannot hang the run.
// -----------------------------------------------------------------------------
`timescale 1ns/1ps
module i2c_intent_probe_tb;
localparam integer CNT_W = 16;
reg clk = 1'b0;
reg rst_n = 1'b0;
reg my_drive_low = 1'b0, pin_drive_low = 1'b0, line = 1'b1, settle_window = 1'b0;
wire out_path_fault, pulldown_fault, foreign_hold;
wire [CNT_W-1:0] first_div_cyc, n_out_path, n_pulldown, n_foreign, n_consistent, cyc;
wire [1:0] first_div_class;
integer errors = 0;
integer checks = 0;
integer neg_detected = 0;
integer neg_expected = 0;
integer negative_mode = 0; // 1 while T4 deliberately checks a wrong expectation
always #5 clk = ~clk;
// A SECOND, NARROW instance. The wide instance's counters saturate at 65535, a
// bound no reasonable test length reaches -- so the saturation the header claims
// was, in the first version of this bench, an entirely untested claim, and a
// mutation that replaced saturation with wrapping survived every check. CNT_W is
// a parameter boundary (nothing else in the design depends on its value), and the
// cheapest way to test a bound is to move the bound rather than lengthen the run.
// At CNT_W = 4 the counters saturate at 15, which 40 cycles reaches comfortably.
reg nq_rst_n = 1'b0;
reg nq_my = 1'b0, nq_pin = 1'b0, nq_line = 1'b1;
wire [3:0] nq_n_out_path, nq_n_consistent, nq_cyc;
wire nq_out_path_fault;
i2c_intent_probe #(.CNT_W(4)) narrow (
.clk(clk), .rst_n(nq_rst_n),
.my_drive_low(nq_my), .pin_drive_low(nq_pin),
.line(nq_line), .settle_window(1'b0),
.out_path_fault(nq_out_path_fault), .pulldown_fault(), .foreign_hold(),
.first_div_cyc(), .first_div_class(),
.n_out_path(nq_n_out_path), .n_pulldown(), .n_foreign(),
.n_consistent(nq_n_consistent), .cyc(nq_cyc)
);
i2c_intent_probe #(.CNT_W(CNT_W)) dut (
.clk(clk), .rst_n(rst_n),
.my_drive_low(my_drive_low), .pin_drive_low(pin_drive_low),
.line(line), .settle_window(settle_window),
.out_path_fault(out_path_fault), .pulldown_fault(pulldown_fault),
.foreign_hold(foreign_hold),
.first_div_cyc(first_div_cyc), .first_div_class(first_div_class),
.n_out_path(n_out_path), .n_pulldown(n_pulldown), .n_foreign(n_foreign),
.n_consistent(n_consistent), .cyc(cyc)
);
// A check that records rather than aborts, and that in negative mode counts a
// detection instead of an error -- so T4 can assert that detection happened.
task chk (input [255:0] name, input integer got, input integer exp);
begin
checks = checks + 1;
if (got !== exp) begin
if (negative_mode) begin
neg_detected = neg_detected + 1;
end else begin
errors = errors + 1;
$display(" FAIL %0s: got %0d expected %0d (t=%0t)", name, got, exp, $time);
end
end else if (negative_mode) begin
errors = errors + 1;
$display(" FAIL negative proof did not fire: %0s", name);
end
end
endtask
task do_reset;
begin
rst_n = 1'b0;
@(posedge clk); @(posedge clk);
rst_n = 1'b1;
end
endtask
// ---------------------------------------------------------------------------
// The INDEPENDENT oracle. Written as an explicit enumeration of the sixteen
// input combinations rather than as the design's boolean expression, so a
// mutation to that expression cannot also mutate the expectation.
//
// Vector order: {my_drive_low, pin_drive_low, line, settle_window}
// ---------------------------------------------------------------------------
function [1:0] oracle (input [3:0] v);
begin
case (v)
// my pin line settle
4'b0000: oracle = 2'd3; // 0 0 0 0 released, released, wire LOW, settled -> foreign
4'b0001: oracle = 2'd0; // 0 0 0 1 released, wire LOW but still charging -> agree
4'b0010: oracle = 2'd0; // 0 0 1 0 released, wire HIGH -> agree
4'b0011: oracle = 2'd0; // 0 0 1 1 released, wire HIGH, in window -> agree
4'b0100: oracle = 2'd1; // 0 1 0 0 pad pulls, logic did not -> out-path
4'b0101: oracle = 2'd1; // 0 1 0 1
4'b0110: oracle = 2'd1; // 0 1 1 0 out-path takes priority over pull-down
4'b0111: oracle = 2'd1; // 0 1 1 1
4'b1000: oracle = 2'd1; // 1 0 0 0 logic pulls, pad does not -> out-path
4'b1001: oracle = 2'd1; // 1 0 0 1
4'b1010: oracle = 2'd1; // 1 0 1 0
4'b1011: oracle = 2'd1; // 1 0 1 1
4'b1100: oracle = 2'd0; // 1 1 0 0 pulling and wire LOW -> agree (self-hold)
4'b1101: oracle = 2'd0; // 1 1 0 1
4'b1110: oracle = 2'd2; // 1 1 1 0 pad pulls, wire HIGH -> pull-down fault
4'b1111: oracle = 2'd2; // 1 1 1 1 settling cannot explain a HIGH while pulling
endcase
end
endfunction
// Drive one input vector, clock it in, and return the class the probe assigned
// by reading which counter moved. Reading the COUNTERS rather than a class
// output means the check covers the counting path as well as the decode.
reg [CNT_W-1:0] p_op, p_pd, p_fg, p_cn;
reg [1:0] observed;
task apply_and_classify (input [3:0] v);
begin
p_op = n_out_path; p_pd = n_pulldown; p_fg = n_foreign; p_cn = n_consistent;
my_drive_low = v[3];
pin_drive_low = v[2];
line = v[1];
settle_window = v[0];
@(posedge clk);
#1; // settle the non-blocking updates before reading them
if (n_out_path != p_op) observed = 2'd1;
else if (n_pulldown != p_pd) observed = 2'd2;
else if (n_foreign != p_fg) observed = 2'd3;
else if (n_consistent != p_cn) observed = 2'd0;
else observed = 2'd3 + 2'd1; // no counter moved at all
end
endtask
integer i, j;
reg [31:0] lfsr;
integer exp_total;
integer got_total;
initial begin
$display("i2c_intent_probe_tb");
// -----------------------------------------------------------------------
// T1 -- exhaustive truth table, all sixteen input combinations.
// -----------------------------------------------------------------------
$display("T1 exhaustive truth table (16 combinations)");
do_reset;
for (i = 0; i < 16; i = i + 1) begin
apply_and_classify(i[3:0]);
chk({"T1 class v=", 8'd48 + i[7:0]}, observed, oracle(i[3:0]));
end
$display(" 16 combinations classified, %0d errors so far", errors);
// -----------------------------------------------------------------------
// T2 -- bins sum to the sample count, over pseudo-random stimulus.
// -----------------------------------------------------------------------
$display("T2 bin sum over 400 random vectors");
do_reset;
lfsr = 32'h1234_5678;
for (i = 0; i < 400; i = i + 1) begin
lfsr = {lfsr[30:0], lfsr[31] ^ lfsr[21] ^ lfsr[1] ^ lfsr[0]};
my_drive_low = lfsr[4];
pin_drive_low = lfsr[9];
line = lfsr[14];
settle_window = lfsr[19];
@(posedge clk);
end
#1;
got_total = n_out_path + n_pulldown + n_foreign + n_consistent;
chk("T2 bins sum to cycles", got_total, cyc);
chk("T2 sample count is 400", cyc, 400);
// And that the random stimulus actually reached every bin -- a sum check over
// a stimulus that only ever hit one bin would pass while proving nothing.
chk("T2 out-path bin reached", (n_out_path > 0) ? 1 : 0, 1);
chk("T2 pull-down bin reached", (n_pulldown > 0) ? 1 : 0, 1);
chk("T2 foreign bin reached", (n_foreign > 0) ? 1 : 0, 1);
chk("T2 agree bin reached", (n_consistent > 0) ? 1 : 0, 1);
$display(" out_path=%0d pulldown=%0d foreign=%0d agree=%0d total=%0d cyc=%0d",
n_out_path, n_pulldown, n_foreign, n_consistent, got_total, cyc);
// -----------------------------------------------------------------------
// T3 -- first divergence is the FIRST one, including when a higher-priority
// class arrives later. The probe must not upgrade the latched class.
// -----------------------------------------------------------------------
$display("T3 first divergence latching");
do_reset;
// Five clean cycles: released, released, wire high.
my_drive_low = 0; pin_drive_low = 0; line = 1; settle_window = 0;
for (i = 0; i < 5; i = i + 1) @(posedge clk);
#1;
chk("T3 no divergence yet", first_div_class, 0);
chk("T3 five agreeing cycles", n_consistent, 5);
// Cycle 5: a foreign hold (class 3, the LOWEST priority).
line = 0;
@(posedge clk); #1;
chk("T3 class latched as foreign", first_div_class, 3);
chk("T3 cycle stamp is 5", first_div_cyc, 5);
// Now an out-path fault (class 1, the HIGHEST priority) three cycles later.
// The sticky flag must set, the counter must move, and the LATCH MUST NOT
// CHANGE -- the first divergence is a fact about the past.
line = 1; my_drive_low = 1; pin_drive_low = 0;
for (i = 0; i < 3; i = i + 1) @(posedge clk);
#1;
chk("T3 out-path flag set", out_path_fault, 1);
chk("T3 out-path counted", (n_out_path == 3) ? 1 : 0, 1);
chk("T3 latched class UNCHANGED", first_div_class, 3);
chk("T3 latched cycle UNCHANGED", first_div_cyc, 5);
chk("T3 foreign flag still set", foreign_hold, 1);
// Reset clears the record, which a debug register must do or a second
// bring-up attempt reports the first one's faults.
do_reset;
#1;
chk("T3 reset clears class", first_div_class, 0);
chk("T3 reset clears cycle", first_div_cyc, 0);
chk("T3 reset clears flags", {out_path_fault, pulldown_fault, foreign_hold}, 0);
chk("T3 reset clears counts", n_out_path + n_pulldown + n_foreign + n_consistent, 0);
// -----------------------------------------------------------------------
// T5 -- the saturation bound, on the narrow instance. Three claims: the
// counter reaches its maximum, it STAYS there under further events, and it
// does not wrap to zero. The third is the one that matters for a diagnostic
// register: a wrapped counter reporting 0 for 65536 faults is worse than no
// counter, because 0 is also what a clean bus reports.
// -----------------------------------------------------------------------
$display("T5 counter saturation at CNT_W=4 (max 15)");
nq_rst_n = 1'b0; nq_my = 1'b0; nq_pin = 1'b0; nq_line = 1'b1;
@(posedge clk); @(posedge clk);
nq_rst_n = 1'b1;
// An out-path fault every cycle for 40 cycles -- well past the bound of 15.
nq_my = 1'b1; nq_pin = 1'b0;
for (i = 0; i < 40; i = i + 1) @(posedge clk);
#1;
chk("T5 out-path counter saturated at 15", nq_n_out_path, 15);
chk("T5 counter did not wrap to a small value", (nq_n_out_path >= 15) ? 1 : 0, 1);
chk("T5 cycle counter saturated at 15", nq_cyc, 15);
chk("T5 flag still set at saturation", nq_out_path_fault, 1);
// And that the OTHER bin stayed at zero throughout -- saturation must not leak
// between counters, which a shared-increment bug would produce.
chk("T5 agree bin untouched", nq_n_consistent, 0);
// -----------------------------------------------------------------------
// T4 -- negative proof. Each of the four T1/T2/T3 check shapes is run once
// against an expectation known to be wrong. Every one must be detected.
// -----------------------------------------------------------------------
$display("T4 negative proof -- five deliberately wrong expectations");
do_reset;
negative_mode = 1;
neg_expected = 5;
// (a) a classification check with the wrong expected class
apply_and_classify(4'b0000); // truly foreign, class 3
chk("T4a wrong class", observed, 0);
// (b) a sticky-flag check with the flag inverted
chk("T4b wrong flag", foreign_hold, 0);
// (c) a counter check off by one
chk("T4c wrong count", n_foreign, 0);
// (d) a first-divergence stamp that is wrong
chk("T4d wrong stamp", first_div_class, 1);
// (e) a saturation check with the wrong bound -- so T5's shape is proven able
// to fail too, not only T1's and T3's.
chk("T4e wrong saturation bound", nq_n_out_path, 40);
negative_mode = 0;
chk("T4 all five negatives detected", neg_detected, neg_expected);
// -----------------------------------------------------------------------
$display("");
$display("checks=%0d errors=%0d negatives_detected=%0d/%0d",
checks, errors, neg_detected, neg_expected);
if (errors == 0) $display("RESULT: PASS");
else $display("RESULT: FAIL");
$finish;
end
// Watchdog -- bounded completion, independent of anything the DUT does.
initial begin
#200000;
$display("RESULT: FAIL -- watchdog expired, the run did not terminate");
$finish;
end
endmoduleFour points about its construction, each of which is a transferable rule:
The oracle is an enumeration, not an expression. oracle() is a sixteen-entry case statement, written from the specification of the classes rather than from the design's boolean conditions. Had it been written as the same conditions, a mutation to the design's expression would be a mutation to the expectation as well, and the test would prove that an expression equals itself.
The check reads the counters, not a class output. apply_and_classify determines the class by observing which counter moved, which makes every check cover the counting path as well as the decode. A class output would have been easier to check and would have left the counters — the part a human actually reads on hardware — unverified.
The bins must sum to the sample count. T2 asserts n_out_path + n_pulldown + n_foreign + n_consistent == cyc. A classifier that double-counts or drops a cycle passes every individual count check and fails this one. T2 also asserts that all four bins were reached, because a sum check over stimulus that only ever hit one bin is arithmetic about nothing.
Every wait is a fixed number of edges. There is no wait on a DUT condition anywhere, and a watchdog terminates the run independently. A diagnostic instrument is exactly the kind of block you point at a broken design, so its own bench must not be able to hang.
The VHDL implementation
Written independently rather than transliterated, because the point of a second implementation is that it disagrees with the first if either is wrong. Two naming and typing decisions had to be made, and both are the kind of thing that costs an afternoon when discovered at analysis time rather than at design time.
-- -----------------------------------------------------------------------------
-- i2c_intent_probe.vhd
-- The VHDL implementation of Chapter 23.1's instrument. Behaviourally identical to
-- i2c_intent_probe.sv, and written independently rather than transliterated, because
-- the point of a second implementation is to disagree with the first if either is
-- wrong.
--
-- Two naming decisions that a port of this block has to make, and both are the kind
-- of thing that costs an afternoon if discovered at analysis time:
--
-- * the Verilog port `line` becomes `bus_line`. `line` is a type declared in
-- std.textio (an access type over string), and a port whose name collides with
-- a visible type name is a hazard rather than an error -- it analyses, and then
-- misbehaves in whichever direction the visibility rules happen to resolve.
-- VHDL is also case-INSENSITIVE, so `Line` would not have helped.
--
-- * the counters are `unsigned`, not `integer`. An integer with a range
-- constraint does not saturate, it raises an exception and kills the
-- simulation, so a saturation test written against a range-constrained integer
-- reports a crash rather than a wrong value. `unsigned` with an explicit
-- comparison against the maximum expresses the intended behaviour instead of
-- relying on the type system to object to the unintended one.
-- -----------------------------------------------------------------------------
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity i2c_intent_probe is
generic (
CNT_W : positive := 16
);
port (
clk : in std_logic;
rst_n : in std_logic;
my_drive_low : in std_logic; -- observation 1: the engine's decision
pin_drive_low : in std_logic; -- observation 2: what the pad asserts
bus_line : in std_logic; -- observation 3: the resolved wire
settle_window : in std_logic; -- 1 while a released line may legally read LOW
out_path_fault : out std_logic;
pulldown_fault : out std_logic;
foreign_hold : out std_logic;
first_div_cyc : out unsigned(CNT_W - 1 downto 0);
first_div_class : out unsigned(1 downto 0);
n_out_path : out unsigned(CNT_W - 1 downto 0);
n_pulldown : out unsigned(CNT_W - 1 downto 0);
n_foreign : out unsigned(CNT_W - 1 downto 0);
n_consistent : out unsigned(CNT_W - 1 downto 0);
cyc : out unsigned(CNT_W - 1 downto 0)
);
end entity i2c_intent_probe;
architecture rtl of i2c_intent_probe is
constant CNT_MAX : unsigned(CNT_W - 1 downto 0) := (others => '1');
constant CNT_ZERO : unsigned(CNT_W - 1 downto 0) := (others => '0');
-- Internal copies. A VHDL-93 `out` port cannot be read, and every one of these
-- is read by the logic that updates it.
signal r_out_path_fault : std_logic := '0';
signal r_pulldown_fault : std_logic := '0';
signal r_foreign_hold : std_logic := '0';
signal r_first_div_cyc : unsigned(CNT_W - 1 downto 0) := CNT_ZERO;
signal r_first_div_class : unsigned(1 downto 0) := "00";
signal r_n_out_path : unsigned(CNT_W - 1 downto 0) := CNT_ZERO;
signal r_n_pulldown : unsigned(CNT_W - 1 downto 0) := CNT_ZERO;
signal r_n_foreign : unsigned(CNT_W - 1 downto 0) := CNT_ZERO;
signal r_n_consistent : unsigned(CNT_W - 1 downto 0) := CNT_ZERO;
signal r_cyc : unsigned(CNT_W - 1 downto 0) := CNT_ZERO;
-- The three conditions, as concurrent signals over the observation points. Same
-- structure as the Verilog: conditions on inputs, never on the sticky flags.
signal c_out_path : std_logic;
signal c_pulldown : std_logic;
signal c_foreign : std_logic;
signal cls : unsigned(1 downto 0);
begin
c_out_path <= my_drive_low xor pin_drive_low;
c_pulldown <= pin_drive_low and bus_line;
c_foreign <= (not pin_drive_low) and (not my_drive_low)
and (not bus_line) and (not settle_window);
-- The priority chain, innermost layer first.
cls <= "01" when c_out_path = '1' else
"10" when c_pulldown = '1' else
"11" when c_foreign = '1' else
"00";
process (clk)
begin
if rising_edge(clk) then
if rst_n = '0' then
r_out_path_fault <= '0';
r_pulldown_fault <= '0';
r_foreign_hold <= '0';
r_first_div_cyc <= CNT_ZERO;
r_first_div_class <= "00";
r_n_out_path <= CNT_ZERO;
r_n_pulldown <= CNT_ZERO;
r_n_foreign <= CNT_ZERO;
r_n_consistent <= CNT_ZERO;
r_cyc <= CNT_ZERO;
else
if r_cyc /= CNT_MAX then
r_cyc <= r_cyc + 1;
end if;
case cls is
when "01" =>
r_out_path_fault <= '1';
if r_n_out_path /= CNT_MAX then
r_n_out_path <= r_n_out_path + 1;
end if;
when "10" =>
r_pulldown_fault <= '1';
if r_n_pulldown /= CNT_MAX then
r_n_pulldown <= r_n_pulldown + 1;
end if;
when "11" =>
r_foreign_hold <= '1';
if r_n_foreign /= CNT_MAX then
r_n_foreign <= r_n_foreign + 1;
end if;
when others =>
if r_n_consistent /= CNT_MAX then
r_n_consistent <= r_n_consistent + 1;
end if;
end case;
if r_first_div_class = "00" and cls /= "00" then
r_first_div_class <= cls;
r_first_div_cyc <= r_cyc;
end if;
end if;
end if;
end process;
out_path_fault <= r_out_path_fault;
pulldown_fault <= r_pulldown_fault;
foreign_hold <= r_foreign_hold;
first_div_cyc <= r_first_div_cyc;
first_div_class <= r_first_div_class;
n_out_path <= r_n_out_path;
n_pulldown <= r_n_pulldown;
n_foreign <= r_n_foreign;
n_consistent <= r_n_consistent;
cyc <= r_cyc;
end architecture rtl; -- -----------------------------------------------------------------------------
-- i2c_intent_probe_tb.vhd
-- The VHDL oracle. Same five parts as the Verilog bench, with the truth table
-- written out as an explicit 16-entry table rather than as the design's boolean
-- expression, for the same reason.
-- -----------------------------------------------------------------------------
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity i2c_intent_probe_tb is
end entity i2c_intent_probe_tb;
architecture tb of i2c_intent_probe_tb is
signal clk : std_logic := '0';
signal rst_n : std_logic := '0';
signal my_drive_low, pin_drive_low, settle_window : std_logic := '0';
signal bus_line : std_logic := '1';
signal out_path_fault, pulldown_fault, foreign_hold : std_logic;
signal first_div_cyc : unsigned(15 downto 0);
signal first_div_class : unsigned(1 downto 0);
signal n_out_path, n_pulldown, n_foreign, n_consistent, cyc : unsigned(15 downto 0);
-- The narrow instance, for the saturation bound.
signal nq_rst_n : std_logic := '0';
signal nq_my, nq_pin : std_logic := '0';
signal nq_line : std_logic := '1';
signal nq_n_out_path, nq_n_consistent, nq_cyc : unsigned(3 downto 0);
signal nq_out_path_fault : std_logic;
signal done : boolean := false;
-- Counters live in the single stimulus process, not as shared variables. VHDL-2008
-- requires a shared variable to have a protected type, and nvc enforces it; since
-- exactly one process writes these, a process variable is both legal and the more
-- honest statement about ownership.
-- The independent oracle: a table indexed by the four observation points,
-- {my, pin, bus_line, settle}, giving the class the priority chain must produce.
type cls_table_t is array (0 to 15) of integer;
constant ORACLE : cls_table_t := (
0 => 3, -- 0 0 0 0 released, wire LOW, settled -> foreign
1 => 0, -- 0 0 0 1 released, wire LOW, still charging -> agree
2 => 0, -- 0 0 1 0
3 => 0, -- 0 0 1 1
4 => 1, -- 0 1 0 0 pad pulls, logic did not -> out-path
5 => 1,
6 => 1, -- 0 1 1 0 out-path outranks pull-down
7 => 1,
8 => 1, -- 1 0 0 0 logic pulls, pad did not -> out-path
9 => 1,
10 => 1,
11 => 1,
12 => 0, -- 1 1 0 0 pulling, wire LOW -> agree
13 => 0,
14 => 2, -- 1 1 1 0 pad pulls, wire HIGH -> pull-down
15 => 2
);
begin
clk_gen : process
begin
while not done loop
clk <= '0'; wait for 5 ns;
clk <= '1'; wait for 5 ns;
end loop;
wait;
end process;
dut : entity work.i2c_intent_probe
generic map (CNT_W => 16)
port map (
clk => clk, rst_n => rst_n,
my_drive_low => my_drive_low, pin_drive_low => pin_drive_low,
bus_line => bus_line, settle_window => settle_window,
out_path_fault => out_path_fault, pulldown_fault => pulldown_fault,
foreign_hold => foreign_hold,
first_div_cyc => first_div_cyc, first_div_class => first_div_class,
n_out_path => n_out_path, n_pulldown => n_pulldown, n_foreign => n_foreign,
n_consistent => n_consistent, cyc => cyc
);
narrow : entity work.i2c_intent_probe
generic map (CNT_W => 4)
port map (
clk => clk, rst_n => nq_rst_n,
my_drive_low => nq_my, pin_drive_low => nq_pin,
bus_line => nq_line, settle_window => '0',
out_path_fault => nq_out_path_fault, pulldown_fault => open,
foreign_hold => open,
first_div_cyc => open, first_div_class => open,
n_out_path => nq_n_out_path, n_pulldown => open, n_foreign => open,
n_consistent => nq_n_consistent, cyc => nq_cyc
);
stim : process
variable errors : integer := 0;
variable checks : integer := 0;
variable neg_detected : integer := 0;
variable negative_mode : boolean := false;
-- std_logic'pos is NOT the bit value. The enumeration is ('U','X','0','1','Z',
-- 'W','L','H','-'), so 'pos('1') is 3 and 'pos('0') is 2 -- and a check written
-- as chk(..., std_logic'pos(flag), 1) fails against a perfectly correct '1'.
-- The first version of this bench had exactly that, and reported five failures
-- against a design whose every counter already matched the Verilog run byte for
-- byte. That is a bench bug wearing a design bug's clothes.
function bitval (s : std_logic) return integer is
begin
if s = '1' then return 1; else return 0; end if;
end function;
procedure chk (name : in string; got : in integer; exp : in integer) is
begin
checks := checks + 1;
if got /= exp then
if negative_mode then
neg_detected := neg_detected + 1;
else
errors := errors + 1;
report " FAIL " & name & ": got " & integer'image(got)
& " expected " & integer'image(exp) severity warning;
end if;
elsif negative_mode then
errors := errors + 1;
report " FAIL negative proof did not fire: " & name severity warning;
end if;
end procedure;
procedure do_reset is
begin
rst_n <= '0';
wait until rising_edge(clk);
wait until rising_edge(clk);
rst_n <= '1';
end procedure;
variable p_op, p_pd, p_fg, p_cn : unsigned(15 downto 0);
variable observed : integer;
variable lfsr : unsigned(31 downto 0);
variable v : integer;
variable got_total : integer;
procedure apply_and_classify (vec : in integer) is
begin
p_op := n_out_path; p_pd := n_pulldown; p_fg := n_foreign; p_cn := n_consistent;
my_drive_low <= '1' when (vec / 8) mod 2 = 1 else '0';
pin_drive_low <= '1' when (vec / 4) mod 2 = 1 else '0';
bus_line <= '1' when (vec / 2) mod 2 = 1 else '0';
settle_window <= '1' when vec mod 2 = 1 else '0';
wait until rising_edge(clk);
wait for 1 ns;
if n_out_path /= p_op then observed := 1;
elsif n_pulldown /= p_pd then observed := 2;
elsif n_foreign /= p_fg then observed := 3;
elsif n_consistent /= p_cn then observed := 0;
else observed := 4; -- no counter moved at all
end if;
end procedure;
begin
report "i2c_intent_probe_tb (VHDL)";
-- T1 -- exhaustive truth table.
report "T1 exhaustive truth table (16 combinations)";
do_reset;
for i in 0 to 15 loop
apply_and_classify(i);
chk("T1 class v=" & integer'image(i), observed, ORACLE(i));
end loop;
-- T2 -- bins sum to the sample count.
report "T2 bin sum over 400 pseudo-random vectors";
do_reset;
lfsr := x"12345678";
for i in 0 to 399 loop
lfsr := lfsr(30 downto 0) & (lfsr(31) xor lfsr(21) xor lfsr(1) xor lfsr(0));
my_drive_low <= lfsr(4);
pin_drive_low <= lfsr(9);
bus_line <= lfsr(14);
settle_window <= lfsr(19);
wait until rising_edge(clk);
end loop;
wait for 1 ns;
got_total := to_integer(n_out_path) + to_integer(n_pulldown)
+ to_integer(n_foreign) + to_integer(n_consistent);
chk("T2 bins sum to cycles", got_total, to_integer(cyc));
chk("T2 sample count is 400", to_integer(cyc), 400);
chk("T2 out-path bin reached", boolean'pos(n_out_path > 0), 1);
chk("T2 pull-down bin reached", boolean'pos(n_pulldown > 0), 1);
chk("T2 foreign bin reached", boolean'pos(n_foreign > 0), 1);
chk("T2 agree bin reached", boolean'pos(n_consistent > 0), 1);
report " out_path=" & integer'image(to_integer(n_out_path))
& " pulldown=" & integer'image(to_integer(n_pulldown))
& " foreign=" & integer'image(to_integer(n_foreign))
& " agree=" & integer'image(to_integer(n_consistent));
-- T3 -- first divergence.
report "T3 first divergence latching";
do_reset;
my_drive_low <= '0'; pin_drive_low <= '0'; bus_line <= '1'; settle_window <= '0';
for i in 0 to 4 loop wait until rising_edge(clk); end loop;
wait for 1 ns;
chk("T3 no divergence yet", to_integer(first_div_class), 0);
chk("T3 five agreeing cycles", to_integer(n_consistent), 5);
bus_line <= '0';
wait until rising_edge(clk); wait for 1 ns;
chk("T3 class latched as foreign", to_integer(first_div_class), 3);
chk("T3 cycle stamp is 5", to_integer(first_div_cyc), 5);
bus_line <= '1'; my_drive_low <= '1'; pin_drive_low <= '0';
for i in 0 to 2 loop wait until rising_edge(clk); end loop;
wait for 1 ns;
chk("T3 out-path flag set", bitval(out_path_fault), 1);
chk("T3 out-path counted", to_integer(n_out_path), 3);
chk("T3 latched class UNCHANGED", to_integer(first_div_class), 3);
chk("T3 latched cycle UNCHANGED", to_integer(first_div_cyc), 5);
chk("T3 foreign flag still set", bitval(foreign_hold), 1);
do_reset;
wait for 1 ns;
chk("T3 reset clears class", to_integer(first_div_class), 0);
chk("T3 reset clears cycle", to_integer(first_div_cyc), 0);
chk("T3 reset clears out-path flag", bitval(out_path_fault), 0);
chk("T3 reset clears foreign flag", bitval(foreign_hold), 0);
chk("T3 reset clears counts",
to_integer(n_out_path) + to_integer(n_pulldown)
+ to_integer(n_foreign) + to_integer(n_consistent), 0);
-- T5 -- the saturation bound, on the narrow instance.
report "T5 counter saturation at CNT_W=4 (max 15)";
nq_rst_n <= '0'; nq_my <= '0'; nq_pin <= '0'; nq_line <= '1';
wait until rising_edge(clk); wait until rising_edge(clk);
nq_rst_n <= '1';
nq_my <= '1'; nq_pin <= '0';
for i in 0 to 39 loop wait until rising_edge(clk); end loop;
wait for 1 ns;
chk("T5 out-path counter saturated at 15", to_integer(nq_n_out_path), 15);
chk("T5 cycle counter saturated at 15", to_integer(nq_cyc), 15);
chk("T5 flag still set at saturation", bitval(nq_out_path_fault), 1);
chk("T5 agree bin untouched", to_integer(nq_n_consistent), 0);
-- T4 -- negative proof.
report "T4 negative proof -- five deliberately wrong expectations";
do_reset;
negative_mode := true;
apply_and_classify(0); -- truly class 3
chk("T4a wrong class", observed, 0);
chk("T4b wrong flag", bitval(foreign_hold), 0);
chk("T4c wrong count", to_integer(n_foreign), 0);
chk("T4d wrong stamp", to_integer(first_div_class), 1);
chk("T4e wrong saturation bound", to_integer(nq_n_out_path), 40);
negative_mode := false;
chk("T4 all five negatives detected", neg_detected, 5);
report "checks=" & integer'image(checks) & " errors=" & integer'image(errors)
& " negatives_detected=" & integer'image(neg_detected) & "/5";
if errors = 0 then
report "RESULT: PASS";
else
report "RESULT: FAIL";
end if;
done <= true;
wait;
end process;
-- Watchdog. Bounded completion independent of the DUT.
wd : process
begin
wait for 200 us;
if not done then
report "RESULT: FAIL -- watchdog expired" severity failure;
end if;
wait;
end process;
end architecture tb;Three things in that pair are worth naming explicitly, because each one is a rule rather than a detail:
The Verilog port line became bus_line. line is a type declared in std.textio — an access type over string — and a port whose name collides with a visible type name is a hazard rather than an error: it analyses, and then resolves whichever way the visibility rules happen to fall. VHDL is also case-insensitive, so Line would not have helped.
The counters are unsigned, not a range-constrained integer. A range-constrained integer does not saturate; it raises an exception and kills the simulation. A saturation test written against one reports a crash rather than a wrong value, which is a materially worse diagnostic — and at that point the type system is objecting to the unintended behaviour instead of the design expressing the intended one.
The bench's counters are process variables, not shared variables. VHDL-2008 requires a shared variable to have a protected type, and nvc enforces it. Since exactly one process writes them, a process variable is both legal and the more honest statement about ownership.
6. What the Runs Say
Both languages, both Verilog standards:
iverilog -g2012 i2c_intent_probe.sv i2c_intent_probe_tb.sv
checks=46 errors=0 negatives_detected=5/5 RESULT: PASS
iverilog -g2001 i2c_intent_probe.sv i2c_intent_probe_tb.sv
checks=46 errors=0 negatives_detected=5/5 RESULT: PASS
nvc --std=2008 -a i2c_intent_probe.vhd i2c_intent_probe_tb.vhd
nvc --std=2008 -e -r i2c_intent_probe_tb
checks=46 errors=0 negatives_detected=5/5 RESULT: PASS
T2 bin populations, 400 pseudo-random vectors, identical in both languages:
out_path=214 pulldown=38 foreign=26 agree=122 total=400 cyc=400The bin populations matching exactly across two independently written implementations, driven by two independently written LFSRs that happen to produce the same sequence, is the strongest single piece of parity evidence available here. It says the two implementations classify the same 400 vectors the same way, bin by bin, not merely that both benches passed.
The Verilog-2001 result is one file, not a copy
The design and its bench compile and pass unchanged under -g2001 as well as -g2012. That is reported as a property of one source, not as a second implementation, because the source contains no SystemVerilog-only construct — the only logic in the file is in a comment. Shipping a duplicate .v differing in its header would have inflated a matrix without adding a check.
The design was written that way deliberately: !== was the natural operator for "intent differs from pin", and it is a simulation-only construct that would have made the block unsynthesisable and would have had no VHDL equivalent. An XOR is the same comparison over two-valued signals, portable, and synthesisable.
The VHDL port found a bench bug, not a design bug
The first VHDL run reported five failures. All five were of the form:
FAIL T3 out-path flag set: got 3 expected 1
FAIL T3 foreign flag still set: got 3 expected 1
FAIL T3 reset clears out-path flag: got 2 expected 0
FAIL T3 reset clears foreign flag: got 2 expected 0
FAIL T5 flag still set at saturation: got 3 expected 1
Every one of them was std_logic'pos(). The enumeration is
('U', 'X', '0', '1', 'Z', 'W', 'L', 'H', '-')
0 1 2 3 4 5 6 7 8
so 'pos('1') is 3 and 'pos('0') is 2. A check written as
chk(..., std_logic'pos(flag), 1) fails against a perfectly correct '1',
and reports "got 3" -- a number that looks like corrupt data rather than
like a units error.This is worth dwelling on as debug practice, because it is a small instance of the whole chapter. Five failures appeared, all in the same design, all reporting values that were not remotely close to expectation. The tempting interpretation was a design bug in the VHDL port.
The observation that settled it took one line and cost nothing: the counter values in the same run already matched the Verilog run exactly — 214, 38, 26, 122. A design whose every event count agrees bit for bit with an independent implementation is not a design that has corrupted five flags. The failing checks were all of one kind, and the one thing that kind had in common was the conversion, not the design.
Generalised: when a set of failures shares a shape, suspect the thing they share. And when part of a run agrees with an independent reference and part does not, the disagreement is far more likely to be in the part that differs — which here was the bench's own type conversion.
7. Mutation — and the Bound Nobody Tested
Sixteen mutations on the SystemVerilog, eight on the VHDL, each written to its own path, that path hashed, and iverilog/nvc invoked on that exact path, with the harness asserting that every sed actually changed the file. A mutation applied to a copy that is not the one compiled scores a verdict against the baseline source and looks entirely normal doing it.
SystemVerilog: 16 valid 13 KILLED 3 SURVIVED
M02 the pull-down condition reads drive INTENT instead of the pad SURVIVED
M04 the foreign condition drops its own-intent term SURVIVED
M08 the counters WRAP instead of saturating SURVIVEDThree survivors, and they are not the same kind of thing at all. Taking the third one first, because it is the one that was a real gap.
M08 was a verification gap, and the parameter boundary was the reason
The block's header claims the counters saturate, and gives a reason: a wrapped counter reporting 0 after 65536 faults is worse than no counter, because 0 is also what a clean bus reports. Replacing saturation with wrapping changed nothing in any test.
Of course it did not. CNT_W is 16, the counters saturate at 65535, and the longest run in the bench is 400 cycles. The saturation was an entirely untested claim, asserted confidently in a comment, and the only reason nobody noticed is that the default configuration never approaches the bound.
That is the parameter-boundary problem in its purest form. CNT_W is the one parameter the block has; nothing else depends on its value; and the behaviour that depends on it was never exercised, because exercising it at the default width needs 65536 cycles of a single fault class.
The fix is not a longer run. It is to move the bound instead of lengthening the test — a second instance at CNT_W = 4, where saturation is at 15 and forty cycles reaches it comfortably:
A second probe instance at CNT_W = 4, driven with an out-path fault for 40
consecutive cycles. Three claims, none of which the wide instance can reach:
the counter reaches 15
it STAYS at 15 under further events
it does not wrap to a small value
plus one that catches a shared-increment bug: the agree bin must still read 0.
After T5: SystemVerilog 16 valid 14 KILLED 2 SURVIVEDM02 and M04 are equivalent, and the equivalence is fragile
The remaining two survivors are a different animal. Both remove a term:
baseline c_pulldown = pin_drive_low && line
M02 c_pulldown = my_drive_low && line
baseline c_foreign = !pin_drive_low && !my_drive_low && !line && !settle_window
M04 c_foreign = !pin_drive_low && !line && !settle_windowNeither can be killed, and adding stimulus until one dies would be the wrong move — the question is why they cannot be killed, and the answer is in the priority chain.
c_out_path is my_drive_low ^ pin_drive_low and takes priority 1. So whenever my_drive_low and pin_drive_low differ, the class is 1 and the values of c_pulldown and c_foreign are never used. And whenever they agree, my_drive_low and pin_drive_low are interchangeable, so the mutated and baseline expressions are identical. There is no input at which the mutation can be observed:
Enumerating {my_drive_low, pin_drive_low, line, settle_window} and comparing the
CLASS leaving the priority chain -- not the masked term, which is where the
argument has to be made:
M02: differing input vectors out of 16 = 0
M04: differing input vectors out of 16 = 0
The VHDL campaign reproduces it independently: V07, the same mutation on the
independently written VHDL implementation, also survives. 7 KILLED, 1 SURVIVED.So both are EQUIVALENT, proven exhaustively rather than argued. But the more useful finding is what happens if the priority is changed:
Reorder the chain so pull-down precedes out-path, and re-run the same comparison:
M02 under a REORDERED priority: differing vectors = 4
(my, pin, line, settle) = (0,1,1,0) (0,1,1,1) (1,0,1,0) (1,0,1,1)
Exactly the vectors where intent and pin DISAGREE and the wire is HIGH -- the
region the out-path priority was masking.This is why the baseline is written with pin_drive_low even though the current priority makes the choice unobservable. The pull-down claim is a claim about the pad: the pad asserts a pull-down and the wire does not follow. Writing it in terms of the engine's intent happens to be equivalent today and stops being equivalent the moment somebody reorders the chain — and a latent non-equivalence behind a priority decision is a bug waiting for an unrelated edit.
8. What This Workflow Cannot Do
Being specific about the limits is part of the method.
It cannot model metastability. Nothing in RTL simulation can. A simulated flip-flop has no aperture and no resolution time, so driving an edge into a clock edge produces whichever value the event ordering picks, deterministically, every run. Chapter 19.4 argues the structural case for a synchronizer; no simulation, including this one, is evidence about it. Chapter 23.8 returns to this.
It cannot tell you which device is holding a line. A foreign_hold verdict says "not this device", and that is all the resolved bus contains. Attributing it needs a second measurement — a per-device probe, a physical disconnection, or a segment switch. Chapter 23.6 is that problem.
It cannot distinguish a legal hold from an illegal one. Clock stretching is a foreign hold. Arbitration loss is a foreign hold. A stuck bus is a foreign hold. The class is about ownership, not legality, and deciding legality needs protocol context the block deliberately does not have. Chapter 23.7 supplies that context.
On hardware, the three observation points are not equally easy to get. my_drive_low needs an ILA or a spare pin. pin_drive_low needs a scope or a loopback. line needs an analyser or the design's own input path — and if you use the design's input path for line, the probe cannot detect a fault in that input path, because the fault is now upstream of the instrument. Naming that limitation is more useful than pretending three independent observations are always available.
9. Misconceptions
10. Debug Lab
A controller reports arbitration loss on a bus with exactly one controller
Step 4 in one measurement — and why three of the four candidates die before any RTL is read
// A single-controller I2C bus. One FPGA controller, one EEPROM target, no second
// master anywhere on the board. The controller's error register reports ARBITRATION
// LOST, intermittently, on perhaps one transfer in twenty.
//
// Arbitration loss is detected the way Chapter 17.10 builds it: the controller
// releases SDA intending to transmit a 1, reads the line back, and finds a 0.
// Someone else is transmitting a 0 in the same bit slot, so this controller has
// lost and must withdraw.
//
// The detection logic is almost certainly correct -- it is one comparator, it is
// verified, and it is the same comparator that detects clock stretching, which
// works. What is NOT established is that its INPUT is correct.
//
// Four candidates, and they live in four different layers:
//
// (a) there is in fact a second controller -- a bootloader, a management
// processor, a debug probe left attached -- and the board diagram is wrong
// (b) the readback is sampled too early, so the controller reads its own
// release before the pull-up has finished charging, and sees its own
// pull-down as somebody else's
// (c) the output enable is inverted for part of a bit slot, so the pad is
// briefly pulling low while the engine intends to release
// (d) the arbitration comparator is enabled during a slot where it should be
// masked -- the acknowledge slot, where the TARGET legitimately drives a 0
//
// Note that (b), (c) and (d) all produce "released, read back low" with no second
// device present at all, and all three are intermittent in a way that looks like
// contention. The symptom does not discriminate them.FIRST, the instrument (step 2). The error register is read over a debug interface after the fact, so what it reports is a LATCHED flag with no timestamp -- which means "one transfer in twenty" is a count of flags, not of events, and a single transfer could be setting it several times. Reading a counter instead of a flag changes the number: it is not one event per twenty transfers, it is roughly one event per twenty transfers CONCENTRATED in the address phase.
That is already a discriminating observation, and it costs one register read. It does not yet name a cause, but it says the events are not spread uniformly over the transfer, which is what candidate (a) -- a genuine second controller starting a transfer whenever it feels like it -- would predict.
SECOND, liveness (step 3). The bus idles at 3.3 V on both lines. A driven low reaches 180 mV. Release to a valid high takes 640 ns, measured on a scope, against a tr budget of 300 ns for the 400 kHz Fast-mode configuration in use.
-> the bus is alive, and the rise time is TWICE its budget.
THIRD, step 4, with the probe instantiated on the controller's SDA path and settle_window driven from a counter sized for 300 ns -- the budget, not the measured value, because the budget is what the design was built against:
n_foreign = 1034 first_div_class = 3 (foreign hold) n_out_path = 0 first_div_cyc = 41 n_pulldown = 0 n_consistent = 2871
Read those five numbers carefully, because four of them are zeros and near-zeros doing most of the work.
n_out_path = 0 over 3905 observed cycles says drive intent and the pad agreed on EVERY cycle. Candidate (c) is dead -- not "unlikely", dead, because an inverted output enable cannot be silent in this counter.
n_pulldown = 0 says the pad never asserted a pull-down the wire failed to follow. The output path is electrically fine.
n_foreign = 1034 says the line was low, with this controller released and its pad released, outside the 300 ns settling window, on 1034 cycles.
Candidate (b). The 640 ns rise time exceeds the 300 ns the design's settling allowance was built for, so on every release the controller reads back a low line for roughly 340 ns longer than it expects. Inside the arbitration window that reads as another device transmitting a zero.
The probe did not prove this on its own, and it is worth being exact about what it did and did not establish:
the probe eliminated (c) and the output path entirely -- two zero counters over 3905 cycles, which is a strong negative and the whole reason step 4 comes before hypothesis formation
the probe CANNOT distinguish (a) from (b) from (d), because all three produce a foreign-hold class. The class is about OWNERSHIP, not legality, and it says only "not this device"
What discriminated (b) from (a) was the settle_window input itself. Re-running with settle_window sized from the MEASURED 640 ns rather than the 300 ns budget takes n_foreign from 1034 to 0. A second controller's pull-down is not a function of this controller's rise-time allowance; a misjudged settling window is exactly that function. One parameter change, and the verdict moves from 1034 to 0.
And (d) dies on the concentration observation from step 2: a comparator wrongly enabled during the acknowledge slot would put events at the ninth bit of every byte, not in the address phase.
CORRECTION, and this is where it gets interesting -- there are two, and they fix different things:
the SYMPTOM is fixed by widening the controller's settling allowance to cover 640 ns. Arbitration-loss reports stop.
the ROOT CAUSE is that the bus rise time is twice its budget, which means a 2.2 kilohm pull-up where the capacitance needs about 1 kilohm. Widening the allowance leaves a bus that is still out of specification for Fast-mode, still eating 640 ns of every bit period, and still one added device away from failing something else.
Fixing the resistor is the correction. Widening the allowance is defensive engineering worth doing anyway, because a controller should not report contention on a slow edge -- but shipping only the allowance change would be a fix that hides its own root cause, and step 7 exists to catch exactly that.
PROOF, and note that "the error stopped" is not on the list:
1. before: n_foreign = 1034 with settle_window at the 300 ns budget 2. after the resistor change: rise time measured at 210 ns, inside budget 3. after: n_foreign = 0 with settle_window UNCHANGED at the 300 ns budget
Step 3 is the one that proves it. The allowance is the value it always was, so the change in verdict is attributable to the bus rather than to a threshold that was moved until the complaint stopped.
11. Reason It Through
12. Questions
13. What This Chapter Settled
A procedure, in seven steps, whose order is the content: state the observation, suspect the instrument, establish liveness, compare intent against the wire, form competing hypotheses, choose the discriminating measurement, and prove the correction. Steps 2 and 4 come before any hypothesis because they are cheap, mechanical, and eliminate regions rather than nominating suspects.
An instrument for step 4, in two languages, that classifies a disagreement by layer rather than by device. Forty-six checks, zero errors and five negative proofs in each language, with the 400-vector bin populations identical across two independently written implementations. The design and its bench are one source that passes under both Verilog standards, which is reported as a property of that source rather than as a second entry in a matrix.
Two findings from the mutation campaign that are worth more than the campaign's score. A saturation bound asserted in a comment and never tested, because the default parameter made it unreachable — fixed by shrinking the bound, not by lengthening the run. And two survivors that are exhaustively equivalent under the current priority chain and stop being equivalent when that chain is reordered, which is why the code is written in the terms that stay correct.
And one bench bug, found by the third language rather than by a validator: five failures that looked like corrupt design state and were a type conversion, identified by noticing that the same run's counters already agreed with the Verilog exactly.
What the chapter has deliberately not done is read a capture. Step 1 says "state the observation", and on a real bus the observation usually arrives as a few thousand samples of two signals with no annotation on them at all. Chapter 23.2 is how you turn that into a sentence — and how you find the first byte in it that is not what anybody intended.
Continue learning
Related tutorials
- Related topic
Reading a Capture — Framing, Address, ACK and the First Bad Byte
How to turn a few thousand samples of two wires into a sentence, and then into a number: the decode rules in the order they apply, the provisional-bit rule that stops every STOP decoding as a truncated byte, and a first-divergence diff that answers where a capture stops matching what was asked for. Includes a decoder whose first version had that exact bug, and the ten failures from one cause that found it.
- Related topic
Protocol Bug or Electrical Bug? Splitting the Search Space
The highest-value triage decision in I²C debugging, and the evidence that settles it. Why a wrong bit, an intermittent acknowledge and a failure that only appears at 400 kHz are produced identically by a shift-register bug and by a slow edge — and how one measurement, the time from release to a readable HIGH, separates them. Builds the classifier, runs one drive pattern against two rise times, and reports the bench race that produced an off-by-one at some latencies and not others.
- Related topic
The START Condition
START is SDA falling while SCL is high, it is generated only by the controller, and it makes the bus busy. Derive what every device must do in response, then build a detector in three languages and find out why its two guard terms and its reset value are all load-bearing.
- Related topic
Missing ACK and Address Faults
Seven causes of a missing acknowledge, one diagnostic procedure of at most five probes, and a confusion matrix that shows exactly which two it cannot separate. Built on a fault-injectable target so every verdict is checked against a known cause — including a target that acknowledges correctly and a controller that samples too early to see it.
