I²C · Module 23
RTL and FPGA Bugs That Only Appear on Hardware
The last class of I²C failure: bugs whose evidence is not on the bus. Separates what simulation genuinely cannot show — metastability — from the larger set nobody modelled, turns a frequency-dependent filter bug into a measured boundary, and closes the module by naming, for each failure class, the observation point where its evidence actually lives.
Seven chapters of this module have worked from evidence on the bus, or from evidence inside a controller that an ILA can reach. This one is about the residue: failures that appear on hardware, pass every simulation, and whose evidence is somewhere neither a capture nor a testbench has been looking.
Chapter 19.7 catalogues those mismatches from the design side. This chapter takes them from the debugging side, and it starts by dividing them honestly — because the most common mistake in this territory is treating "simulation cannot show it" as one category when it is two.
1. Two Categories, Not One
CATEGORY A -- outside simulation in principle
metastability and its resolution time
threshold variation between devices
the moment two receivers cross a threshold on one slow edge
temperature, ageing, supply droop
CATEGORY B -- outside THIS simulation, and only because nobody modelled it
a filter threshold that is in clocks while its obligation is in nanoseconds
a synchroniser's LATENCY, as distinct from its protection
a pin constraint, an I/O standard, a bank voltage
a reset value that is wrong for the signal it resets
a rise timeCategory A is a genuine boundary and the honest response is to say so. No result in this chapter, or anywhere in this module, is evidence about whether a synchroniser works.
Category B is where the money is, and the distinction matters because these two get the same sentence — "you can't simulate that" — and only one of them deserves it. A filter that deletes legal bits at 100 MHz and not at 25 MHz is completely deterministic, completely simulable, and was missed because nobody wrote the test.
2. Where the Evidence Lives
The input path is where most of Category B collects, because it is the one place where an asynchronous, analogue, externally-driven signal becomes a synchronous digital one.
// -----------------------------------------------------------------------------
// i2c_input_conditioner.sv
// Chapter 23.8's instrument, and the last block in this module.
//
// It is the standard input path an I2C block needs on real hardware: a
// synchroniser, then a filter that rejects spikes, then the conditioned line the
// protocol engine actually reads. Chapters 19.4 and 19.5 built both halves; this
// version exists to be MEASURED rather than to be used, because the two most
// common hardware-only I2C bugs live in exactly here:
//
// a spike that is not rejected, which manufactures framing
// a LEGAL pulse that IS rejected, which deletes a bit
//
// and the second one is the interesting case, because it is a filter working
// exactly as designed.
//
// ── WHAT THIS BLOCK CAN AND CANNOT BE EVIDENCE ABOUT ────────────────────────
//
// The SYNCHRONISER's purpose is to bound the probability that a metastable value
// propagates. Nothing in RTL simulation can be evidence about that. A simulated
// flip-flop has no aperture and no resolution time; driving an edge into a clock
// edge yields whichever value the event ordering picks, deterministically, every
// run. The synchroniser here is structural insurance and this file's testbench
// says nothing whatever about whether it works.
//
// What IS simulable, and what the bench measures, is the synchroniser's other
// consequence: LATENCY. Every stage delays the protocol engine's view of the bus
// by one sample, and that delay is real, deterministic, and shows up in timing
// margins. It is also the part people forget, because the flops were added for
// metastability and the latency arrives as a side effect.
//
// The FILTER is entirely simulable, and its threshold is a number that should be
// measured rather than assumed. That is the point of the bench: the filter's
// actual rejection width is a property of the implementation, and a comment
// claiming it is FILT_LEN is a claim, not a measurement.
//
// ── THE FILTER, AND ITS ONE DANGEROUS PROPERTY ──────────────────────────────
//
// It accepts a new level only after seeing it on FILT_LEN consecutive samples.
// Anything shorter is discarded. That threshold is expressed in SAMPLE CLOCKS,
// and the obligation it is meant to satisfy -- reject spikes up to tSP, do not
// reject a legal bus pulse -- is expressed in NANOSECONDS.
//
// So the same RTL, at a different sample clock, rejects a different set of
// pulses. A filter sized for 50 ns at 25 MHz rejects 120 ns of pulse at 100 MHz,
// and a bus running Fast-mode Plus has legal low periods of 500 ns. Move the
// design to a faster clock and the filter starts eating real bits, silently, with
// no line of RTL changed. Chapter 19.5's DebugLab is the same arithmetic; this is
// what it looks like as a measurement.
// -----------------------------------------------------------------------------
module i2c_input_conditioner #(
// Synchroniser stages. Two is the usual choice; this is a parameter so the
// bench can MEASURE the latency each one adds rather than assume it.
parameter integer SYNC_DEPTH = 2,
// Consecutive samples at a new level before the filter accepts it.
parameter integer FILT_LEN = 3
) (
input wire clk,
input wire rst_n,
// The raw pin. Asynchronous to clk: that is the whole reason this block exists.
input wire pin,
// The conditioned line the protocol engine reads.
output reg line_out,
// Diagnostics, for a status register or an ILA.
output reg [15:0] n_rejected, // level changes discarded as too short
output reg [15:0] n_accepted // level changes that survived the filter
);
// ── synchroniser ────────────────────────────────────────────────────────
// Reset to 1, the idle state of an I2C line. A synchroniser resetting to 0 on
// a bus that idles high manufactures a falling edge at every reset release --
// which, if SDA, is a START condition invented by the reset (Chapter 19.4's
// T2). The reset value of a synchroniser is part of its specification.
reg [SYNC_DEPTH-1:0] sync;
wire synced = sync[SYNC_DEPTH-1];
// ── filter ──────────────────────────────────────────────────────────────
reg [15:0] run; // consecutive samples at a level differing from line_out
always @(posedge clk) begin
if (!rst_n) begin
sync <= {SYNC_DEPTH{1'b1}};
line_out <= 1'b1;
run <= 16'd0;
n_rejected <= 16'd0;
n_accepted <= 16'd0;
end else begin
sync <= {sync[SYNC_DEPTH-2:0], pin};
if (synced != line_out) begin
if (run + 16'd1 >= FILT_LEN[15:0]) begin
// Sustained long enough. Accept it.
line_out <= synced;
run <= 16'd0;
if (n_accepted != 16'hFFFF) n_accepted <= n_accepted + 16'd1;
end else begin
run <= run + 16'd1;
end
end else if (run != 16'd0) begin
// The level went back before the filter accepted it: a spike.
//
// Counted here rather than at the moment it appeared, because until
// it goes back there is no way to know whether it was a spike or the
// beginning of a real transition. That is not an implementation
// detail -- it is why a filter necessarily costs latency.
run <= 16'd0;
if (n_rejected != 16'hFFFF) n_rejected <= n_rejected + 16'd1;
end
end
end
endmodule3. The Filter's Dangerous Property
The filter accepts a new level only after seeing it on FILT_LEN consecutive samples. That threshold is in sample clocks. The obligation it exists to satisfy is in nanoseconds: reject spikes up to tSP (50 ns), do not reject a legal bus pulse.
Those two units are joined by the sample clock, and nothing in the RTL records that. Move the design to a faster clock — a new FPGA family, a shared PLL, a build-time parameter someone changed for an unrelated block — and the filter's threshold in nanoseconds grows in proportion, with no line of RTL altered.
A bit that exists on the pin and not in the engine
12 cyclesThe consequence on a bus is a deleted bit, and a deleted bit is not a corrupted one. The target's shift register never advanced, so every bit after it is shifted — which is the same signature as Chapter 23.7's ignored stretch, from a completely different mechanism, and it will be attributed to the same wrong places.
4. Measuring Rather Than Asserting
// -----------------------------------------------------------------------------
// i2c_input_conditioner_tb.sv
// Measures the input path rather than exercising it. Four quantities, each of
// which is normally a comment:
//
// the filter's actual rejection threshold, in samples
// the latency the whole chain adds, in samples
// the shortest pulse that survives it
// the bus frequency at which the filter begins deleting legal bits
//
// The last one is the point. It turns a hardware-only, frequency-dependent bug
// into a number produced by a sweep, and the sweep is what makes the boundary
// visible rather than argued.
//
// WHAT THIS BENCH IS NOT EVIDENCE ABOUT. Metastability. A simulated flip-flop has
// no aperture and no resolution time, so the synchroniser's actual purpose is
// outside everything measured here. The latency it adds is measured; its
// protection is asserted structurally and nowhere demonstrated, and no result in
// this file should be read as saying otherwise.
//
// Bounded: every wait is a fixed number of edges, and a watchdog terminates the
// run independently.
// -----------------------------------------------------------------------------
`timescale 1ns/1ps
module i2c_input_conditioner_tb;
localparam integer SYNC_DEPTH = 2;
localparam integer FILT_LEN = 3;
reg clk = 1'b0;
reg rst_n = 1'b0;
always #5 clk = ~clk;
reg pin = 1'b1;
wire line_out;
wire [15:0] n_rejected, n_accepted;
// A SECOND instance with a longer filter, representing the same RTL running
// on a faster sample clock: the filter's threshold is in samples, so a faster
// clock means more samples per nanosecond and a longer effective filter.
reg fast_pin = 1'b1;
wire fast_line;
wire [15:0] fast_rej, fast_acc;
integer errors = 0, checks = 0, neg_detected = 0, negative_mode = 0;
i2c_input_conditioner #(.SYNC_DEPTH(SYNC_DEPTH), .FILT_LEN(FILT_LEN)) dut (
.clk(clk), .rst_n(rst_n), .pin(pin),
.line_out(line_out), .n_rejected(n_rejected), .n_accepted(n_accepted)
);
i2c_input_conditioner #(.SYNC_DEPTH(SYNC_DEPTH), .FILT_LEN(12)) fast (
.clk(clk), .rst_n(rst_n), .pin(fast_pin),
.line_out(fast_line), .n_rejected(fast_rej), .n_accepted(fast_acc)
);
// ------------------------------------------------------------------ helpers
task step; begin @(posedge clk); #1; end endtask
task tick (input integer n); integer k; begin for (k=0;k<n;k=k+1) step; end endtask
task chk (input [255:0] name, input integer got, input integer exp);
begin
checks = checks + 1;
if (got !== exp) begin
if (negative_mode) neg_detected = neg_detected + 1;
else begin
errors = errors + 1;
$display(" FAIL %0s: got %0d expected %0d", name, got, exp);
end
end else if (negative_mode) begin
errors = errors + 1;
$display(" FAIL negative proof did not fire: %0s", name);
end
end
endtask
task do_reset;
begin rst_n = 1'b0; pin = 1'b1; fast_pin = 1'b1; tick(3); rst_n = 1'b1; tick(4); end
endtask
// A low pulse `w` samples wide on an otherwise idle pin.
task pulse_low (input integer w); integer k;
begin
pin = 1'b1; tick(FILT_LEN + SYNC_DEPTH + 4);
pin = 1'b0;
for (k = 0; k < w; k = k + 1) step;
pin = 1'b1;
tick(FILT_LEN + SYNC_DEPTH + 6);
end
endtask
// Does a low pulse `w` samples wide reach the protocol engine at all?
integer saw_low;
task probe_width (input integer w); integer k;
begin
do_reset;
saw_low = 0;
pin = 1'b1; tick(FILT_LEN + SYNC_DEPTH + 4);
pin = 1'b0;
for (k = 0; k < w; k = k + 1) begin step; if (!line_out) saw_low = 1; end
pin = 1'b1;
for (k = 0; k < FILT_LEN + SYNC_DEPTH + 6; k = k + 1)
begin step; if (!line_out) saw_low = 1; end
end
endtask
integer i, w, threshold, latency, fast_latency, boundary;
initial begin
$display("i2c_input_conditioner_tb");
// ---------------------------------------------------------------------
// T1 -- reset state. A synchroniser resetting to 0 on a line that idles
// HIGH manufactures a falling edge at every reset release, which on SDA is
// a START condition invented by the reset. Checked before anything else
// because every later measurement assumes the chain starts idle.
// ---------------------------------------------------------------------
$display("T1 the chain resets to the line's idle state");
rst_n = 1'b0; pin = 1'b1; tick(3);
chk("T1 conditioned line is HIGH during reset", line_out, 1);
rst_n = 1'b1; tick(1);
chk("T1 and immediately after release", line_out, 1);
tick(6);
chk("T1 with no accepted transition", n_accepted, 0);
chk("T1 and none rejected", n_rejected, 0);
// ---------------------------------------------------------------------
// T2 -- the rejection threshold, MEASURED. Sweep the pulse width upward
// and record the narrowest one that reaches the protocol engine. The
// header claims the filter accepts a level after FILT_LEN consecutive
// samples; this is the check of that claim.
// ---------------------------------------------------------------------
$display("T2 the rejection threshold, measured by sweep");
threshold = 0;
for (w = 1; w <= 10; w = w + 1) begin
probe_width(w);
if (saw_low && threshold == 0) threshold = w;
$display(" pulse %0d samples -> %0s", w, saw_low ? "PASSES" : "rejected");
end
chk("T2 the narrowest pulse that passes is FILT_LEN", threshold, FILT_LEN);
chk("T2 and it is not zero, so the sweep found something", (threshold > 0) ? 1 : 0, 1);
// ---------------------------------------------------------------------
// T3 -- the latency, measured. The synchroniser adds one sample per stage
// and the filter adds its threshold, and the sum is what the protocol
// engine's view of the bus is delayed by. This is the part that is
// simulable about a synchroniser; its metastability protection is not.
// ---------------------------------------------------------------------
$display("T3 the latency the chain adds");
do_reset;
pin = 1'b1; tick(6);
latency = 0;
pin = 1'b0;
for (i = 0; i < 30; i = i + 1) begin
step;
if (line_out) latency = latency + 1;
else i = 30;
end
$display(" SYNC_DEPTH=%0d FILT_LEN=%0d -> latency %0d samples",
SYNC_DEPTH, FILT_LEN, latency);
chk("T3 and the edge did eventually arrive", line_out, 0);
// The same measurement on the instance with a different FILT_LEN.
//
// ONE measurement cannot test a relationship -- any formula can be made to
// fit a single point, and the first version of this test simply asserted
// SYNC_DEPTH + FILT_LEN and was wrong by one. Two points at different
// parameters test the formula rather than the constant, which is the whole
// difference between measuring a design and recording its output.
do_reset;
fast_pin = 1'b1; tick(20);
fast_latency = 0;
fast_pin = 1'b0;
for (i = 0; i < 40; i = i + 1) begin
step;
if (fast_line) fast_latency = fast_latency + 1;
else i = 40;
end
$display(" SYNC_DEPTH=%0d FILT_LEN=12 -> latency %0d samples",
SYNC_DEPTH, fast_latency);
chk("T3 slow instance matches SYNC_DEPTH + FILT_LEN - 1",
latency, SYNC_DEPTH + FILT_LEN - 1);
chk("T3 fast instance matches the SAME formula",
fast_latency, SYNC_DEPTH + 12 - 1);
chk("T3 and the two differ by the filter-length difference",
fast_latency - latency, 12 - FILT_LEN);
// ---------------------------------------------------------------------
// T4 -- the filter's purpose: a spike is rejected AND counted. A filter
// that silently swallows spikes is doing its job and reporting nothing,
// and a spike count is the difference between a clean bus and a bus whose
// noise is being absorbed.
// ---------------------------------------------------------------------
$display("T4 spikes rejected, and counted");
do_reset;
for (i = 0; i < 5; i = i + 1) pulse_low(FILT_LEN - 1);
chk("T4 five spikes rejected", n_rejected, 5);
chk("T4 none accepted", n_accepted, 0);
chk("T4 the engine never saw a low", line_out, 1);
do_reset;
for (i = 0; i < 5; i = i + 1) pulse_low(FILT_LEN + 2);
chk("T4 five real pulses accepted, both edges each", n_accepted, 10);
chk("T4 and none counted as spikes", n_rejected, 0);
// ---------------------------------------------------------------------
// T5 -- THE FREQUENCY BOUNDARY. The same RTL on a faster sample clock has
// a longer filter in nanoseconds. Sweep the pulse width against the longer
// filter and find where a pulse that is legal for the bus stops surviving.
//
// This is the hardware-only bug made into a number. Nothing in the RTL
// changed; the clock did.
// ---------------------------------------------------------------------
$display("T5 the same filter on a faster clock");
boundary = 0;
for (w = 1; w <= 20; w = w + 1) begin
do_reset;
fast_pin = 1'b1; tick(20);
fast_pin = 1'b0;
for (i = 0; i < w; i = i + 1) step;
fast_pin = 1'b1;
tick(20);
if (fast_acc > 0 && boundary == 0) boundary = w;
end
$display(" FILT_LEN=12 -> narrowest surviving pulse is %0d samples", boundary);
chk("T5 the threshold scaled with the filter length", boundary, 12);
chk("T5 which is four times the slower instance's", boundary, FILT_LEN * 4);
// And the consequence, stated as the comparison that matters: a pulse that
// the slow instance passes is deleted by the fast one.
do_reset;
pin = 1'b0; fast_pin = 1'b0;
tick(FILT_LEN + SYNC_DEPTH + 2);
chk("T5 a FILT_LEN pulse reaches the slow engine", line_out, 0);
chk("T5 and does NOT reach the fast one", fast_line, 1);
// ---------------------------------------------------------------------
// T6 -- a run that goes back and forth. A level that flaps must not
// accumulate credit toward acceptance across the gaps, or a noisy line
// eventually passes a transition nobody sent.
// ---------------------------------------------------------------------
$display("T6 a flapping line does not accumulate credit");
do_reset;
for (i = 0; i < 8; i = i + 1) begin
pin = 1'b0; tick(FILT_LEN - 1);
pin = 1'b1; tick(FILT_LEN + SYNC_DEPTH + 2);
end
chk("T6 eight spikes rejected", n_rejected, 8);
chk("T6 nothing was ever accepted", n_accepted, 0);
chk("T6 the engine's view never changed", line_out, 1);
// ---------------------------------------------------------------------
// T7 -- negative proof. Eight deliberately wrong expectations.
// ---------------------------------------------------------------------
$display("T7 negative proof -- nine deliberately wrong expectations");
do_reset;
pulse_low(FILT_LEN - 1);
pulse_low(FILT_LEN + 2);
chk("T7 setup: one rejected", n_rejected, 1);
chk("T7 setup: two accepted", n_accepted, 2);
negative_mode = 1;
chk("T7a wrong reject count", n_rejected, 0);
chk("T7b wrong accept count", n_accepted, 0);
chk("T7c wrong line state", line_out, 0);
chk("T7d wrong threshold", threshold, 1);
chk("T7e wrong latency", latency, 0);
chk("T7i wrong latency relationship", fast_latency, latency);
chk("T7f wrong boundary", boundary, FILT_LEN);
chk("T7g wrong fast-line state",fast_line, 0);
chk("T7h wrong sweep claim", (threshold == boundary) ? 1 : 0, 1);
negative_mode = 0;
chk("T7 all nine negatives detected", neg_detected, 9);
$display("");
$display("checks=%0d errors=%0d negatives_detected=%0d/9", checks, errors, neg_detected);
if (errors == 0) $display("RESULT: PASS"); else $display("RESULT: FAIL");
$finish;
end
initial begin
#2000000;
$display("RESULT: FAIL -- watchdog expired, the run did not terminate");
$finish;
end
endmoduleFour things the bench measures, each of which is usually a claim in a header:
T2 the rejection threshold, by sweeping pulse width upward
pulse 1 samples -> rejected
pulse 2 samples -> rejected
pulse 3 samples -> PASSES threshold = FILT_LEN, confirmed
T3 the latency the chain adds
SYNC_DEPTH=2 FILT_LEN=3 -> 4 samples
SYNC_DEPTH=2 FILT_LEN=12 -> 13 samples
T5 the boundary on a longer filter
FILT_LEN=12 -> narrowest surviving pulse is 12 samples
a pulse the 3-sample instance passes does NOT reach the 12-sample one
T4 spikes rejected AND counted, because a filter doing its job silently
is indistinguishable from a bus with no noise on itOne measurement cannot test a relationship
The latency test originally asserted SYNC_DEPTH + FILT_LEN against a single instance, and was wrong by one. The measured value is SYNC_DEPTH + FILT_LEN − 1.
The error is trivial; the fix is not. A single data point can be matched by any number of formulas, so asserting one against it tests a constant rather than a relationship. The corrected test measures two instances with different FILT_LEN and checks that both satisfy the same formula, plus that their difference equals the difference in filter length.
5. What the Synchroniser Section Is Not Evidence About
SYNC_DEPTH is a parameter here and the bench measures what it does to latency. It says nothing at all about what it is for.
A simulated flip-flop has no aperture and no resolution time. Driving an edge into a clock edge yields whichever value the event ordering picks, deterministically, every run — and the same run repeated gives the same answer. Metastability is a probability distribution over resolution times, and the simulator does not have one.
So:
A mutation that reduces SYNC_DEPTH to one stage would not be killed by any test that could be written here, and the bench does not pretend otherwise. What is tested is that the synchroniser is in the path at all — mutation I05 bypasses it and I06 reads the raw pin past it, and both are killed by the latency and filtering behaviour, not by anything about metastability.
The argument for two stages is structural, from Chapter 19.4: an MTBF calculation and a constraint that keeps the internal path timed. Neither is a simulation result, and a debugging session that "proves the synchroniser is fine" by running the testbench has proved nothing.
And the reset value is testable, and matters. A synchroniser resetting to 0 on a line that idles high manufactures a falling edge at every reset release, which on SDA is a START condition invented by the reset. That is Category B, entirely deterministic, and I04 is killed by T1.
6. Mutation
12 valid 12 KILLED 0 SURVIVED (first campaign)The only campaign in this module to come back clean on the first attempt, and the reason is worth naming: the bench was written as a set of measurements rather than as a set of behaviours. A test that sweeps a threshold necessarily probes both sides of it; a test that measures a latency necessarily fails if the latency changes; a test that counts rejections and acceptances separately necessarily notices when one is credited to the other.
Every previous chapter's survivors clustered on boundaries and instants — the things a behaviour-driven test does not naturally contain. A measurement-driven test contains them by construction, because a measurement is a number and a number has neighbours.
7. Where Each Failure Class's Evidence Lives
This is the module's closing summary, and it is the single table worth keeping.
| failure | the evidence is at | what you need |
|---|---|---|
| a protocol bug in the engine | the bus | an analyser, and 23.2's decode |
| an inverted output enable, a wrong pin | intent vs the pad | 23.1's probe — an ILA plus a scope |
| a slow rise, a missing pull-up | the shape of the edge | 23.5's profile — a scope, or an RC model |
| noise | a low pulse with no driver | the same profile, and the tSP boundary |
| a stuck bus | who releases when clocked | 23.6's recovery, and its pulse count |
| a controller ignoring stretching | phase advance AND the line | a counter inside the controller — nothing else can see it |
| arbitration not acted on | intent vs the line, during transmit | 23.7's classifier |
| a filter deleting legal bits | the filter's threshold vs the bus period | a sweep, in simulation |
| a setup-time violation on an acknowledge | the scope, at the SCL edge | 23.4 |
| metastability | nowhere you can measure | structure, constraints, and an MTBF argument |
Read the middle column. In only one row is the evidence on the bus. In five it is a relationship between two signals, which no single observation point holds. In one it is in the time domain rather than the value domain. And in one it is not available at all.
That is the module's thesis, and it explains why almost every instrument built in it takes drive intent as an input. The bus is where symptoms appear and it is almost never where the evidence is.
8. Misconceptions
9. Debug Lab
A design moves to a new FPGA and the bus stops working above 100 kHz
Nothing in the RTL changed, and the RTL is the problem
// An I2C target block, in production for three years on one FPGA family, ported
// to a newer part. The port is a recompile: no RTL changes, same constraints
// file structure, same testbench, all thirty simulation tests pass.
//
// On the new hardware it works at 100 kHz and fails at 400 kHz. At 400 kHz the
// controller reads bytes that look like the right data shifted along by one or
// two positions.
//
// Candidates, and the first three are where everyone looks:
//
// (a) the new part's I/O timing differs -- a pad delay, a different input
// standard, a different bank voltage
// (b) the new part's routing is slower and something misses timing at the
// higher rate
// (c) the board is different and the bus is electrically worse
// (d) something about the design is frequency-dependent in a way that was
// previously inside its margin
//
// Note that "the RTL did not change" is being used as evidence that the RTL is
// not the problem, and it is not evidence of that at all. It is evidence that
// the RTL is the same, which is a different statement when its ENVIRONMENT has
// changed.Step 3 first. Chapter 23.5's profiler on the new board's bus at 400 kHz:
SDA, 40,000 releases: b0=39,991 b3=0 glitch=0 hold=0 max=151 ns SCL, 40,000 releases: b0=39,995 b3=0 glitch=0 hold=0 max=144 ns
Inside a 300 ns Fast-mode budget on both lines, no spikes. Candidate (c) is dead.
Timing closure on the new part: all paths met, with 1.8 ns of slack on the worst. Candidate (b) is dead, and it was the one the team had spent two days on.
Chapter 23.7's stall classifier, in the controller on the other side:
n_stalls = 0 n_ignored_stretch = 0 n_arb_loss = 0 n_stretch = 0
The controller is not ignoring anything and the target never stretches. So the controller is delivering clock pulses correctly and the target is not asking for more time -- which narrows the fault to the target's own handling of the pulses it does receive.
The shifted-data signature says a bit was DELETED rather than flipped: a flip changes one position, a deletion moves every position after it. So the target's shift register failed to advance on at least one clock pulse that reached its pin.
That points at the input path, and the input path has one frequency-dependent number in it. Reading the project file for the new part:
old family: the block's sample clock was 25 MHz (40 ns per sample) new family: the same clock is derived from a shared PLL at 100 MHz because another block in the new design needed it
FILT_LEN = 3, unchanged.
old: 3 samples x 40 ns = 120 ns rejection threshold new: 3 samples x 10 ns = 30 ns rejection threshold
That is the opposite direction from the failure -- a SHORTER threshold rejects less, not more. So the arithmetic as stated does not explain it, and the hypothesis is wrong as written.
Reading further: the port also raised FILT_LEN, in a commit whose message says "restore spike rejection after clock change".
new: FILT_LEN = 12, 12 samples x 10 ns = 120 ns
which restores the 120 ns threshold exactly, and is the correct intent.
120 ns of rejection was inside the margin at 100 kHz and is not at 400 kHz.
The arithmetic that nobody did, in either direction:
Standard-mode, 100 kHz: tLOW minimum is 4700 ns. 120 ns is 2.6% of it. Fast-mode, 400 kHz: tLOW minimum is 1300 ns. 120 ns is 9.2% of it.
Neither of those deletes a bit on its own, and that is the part worth being careful about: 120 ns does not eat a 1300 ns pulse. What it eats is the MARGIN.
The controller on this board generates SCL with a low period of 1340 ns at 400 kHz -- legal, 3% above the minimum. The target's input path delays its view of every edge by SYNC_DEPTH + FILT_LEN - 1 = 13 samples = 130 ns, and delays the RISING edge by the same amount. The low period the protocol engine SEES is therefore the same 1340 ns, shifted.
The deletion comes from somewhere else, and it took a sweep to find. Running the bench's T5 with the bus's actual pulse widths: an SCL high period of 600 ns at 400 kHz is 60 samples at 100 MHz and survives easily. But the controller's SCL has a rise time of 144 ns, measured above -- and the input path sees the line as LOW for that entire rise. The high period the target's ENGINE sees is
600 ns - 144 ns (rise) - 130 ns (input path latency) = 326 ns
at 400 kHz, against 4000 - 144 - 130 = 3726 ns at 100 kHz. Still not zero, and still not a deletion -- but the target's own bit engine requires its sampled high period to be at least 400 ns to register a clock, which at 100 kHz it always was.
So: not one cause, but three numbers that each moved and one requirement that did not.
the sample clock rose, so FILT_LEN had to rise to keep 120 ns the bus speed rose, so the high period fell the rise time is what it always was, and now it is a large fraction
CORRECTION. Reducing FILT_LEN to 6 gives a 60 ns rejection threshold -- still above tSP's 50 ns, which is the actual obligation -- and returns 60 ns of the budget. Combined with reducing the pull-ups from 4.7k to 2.2k, which takes the rise time from 144 ns to 71 ns, the engine's sampled high period becomes
600 - 71 - 70 = 459 ns against a 400 ns requirement
PROOF:
1. before: reads shifted at 400 kHz, correct at 100 kHz 2. after FILT_LEN = 6 alone: still fails, at a lower rate -- which is the result that says the fix is incomplete rather than wrong 3. after the pull-up change as well: correct at 400 kHz over 10^6 transfers 4. and the measured sampled high period is 461 ns, within 2 ns of the arithmetic
Step 2 is the one worth keeping. A partial fix that reduces a failure RATE is usually reported as an improvement and shipped; here it was the evidence that more than one quantity had moved, and stopping there would have left a design with 59 ns of margin.
10. Reason It Through
11. Questions
12. What This Chapter and This Module Settled
Two categories of "you can't simulate that", only one of which deserves the sentence. Metastability, threshold variation and temperature are genuinely outside simulation. Filter thresholds, synchroniser latency, reset values, pin configuration and rise times are outside a particular simulation because nobody modelled them — and that is the larger set.
An input path built to be measured rather than used: a rejection threshold found by sweep rather than asserted, a latency verified at two parameterisations rather than one, a spike count that makes a clean result mean something, and a frequency boundary that turns a hardware-only bug into a number. Thirty-four checks, zero errors, nine negative proofs, twelve mutations killed on the first campaign — the only clean first pass in the module, and the reason is that the bench measures rather than exercises.
And the table in §7, which is what the module has been building toward. Across ten failure classes, the evidence is on the bus exactly once. Five times it is a relationship between two signals that no single observation point holds. Once it is in the time domain rather than the value domain. Once it is not available at all, and the honest response is structure and an argument rather than a test.
That is why every instrument in these eight chapters takes drive intent as an input, why every one of them reports counts rather than booleans, and why the workflow in Chapter 23.1 puts "does intent match the wire" before forming any hypothesis. The bus is where symptoms appear. It is almost never where the evidence is — and an engineer who has internalised that will reach for the right measurement first, which is the whole of what this module was for.
What remains is judgement: whether a design is ready, whether a testbench has proven what it claims, whether the trade-offs were the right ones. Chapter 24.1 begins that.
Continue learning
Related tutorials
- Related topic
Asynchronous SDA/SCL Inputs — Metastability and Synchronization
SDA and SCL have no relationship to the FPGA's clock, so every sample of them lands somewhere in a flip-flop's aperture. Explains metastability as settling time rather than a propagating X, derives what each synchronizer stage buys and costs, and is explicit about the one claim RTL simulation can never support.
- Related topic
The START Condition
START is SDA falling while SCL is high, it is generated only by the controller, and it makes the bus busy. Derive what every device must do in response, then build a detector in three languages and find out why its two guard terms and its reset value are all load-bearing.
- Related topic
tSU;DAT and tHD;DAT — The I²C Data Window
Two parameters every bit of every byte must satisfy, measured against two different edges. One has a specified minimum of zero and is in practice among the tightest constraints on the bus — this chapter resolves that contradiction.
- Related topic
FPGA as Master and as Target — Integration Patterns
Assembles everything: Module 18's verified target behind Module 19's pad, synchronizer and filter, proven end to end on a wired-AND bus in three languages. Shows the controller-side shape on Module 17's real interface, where a soft CPU attaches, and closes with two mutations that cannot be killed in simulation — the module's thesis stated as evidence.
