I²C · Module 19
Asynchronous SDA/SCL Inputs — Metastability and Synchronization
SDA and SCL have no relationship to the FPGA's clock, so every sample of them lands somewhere in a flip-flop's aperture. Explains metastability as settling time rather than a propagating X, derives what each synchronizer stage buys and costs, and is explicit about the one claim RTL simulation can never support.
Chapter 18.2 built a sampling front end out of two flip-flops per line, called the depth an interface contract, and said plainly that the argument for the number belongs to this chapter. Here it is.
The chapter has an unusual shape, because the central claim is one that simulation cannot check. Most of this module's chapters end with evidence; this one ends with evidence and a clear statement of where the evidence stops.
1. Why the Lines Are Asynchronous, and Why That Is Not Fixable
An FPGA I²C target runs on a system clock — say 50 MHz — that it generates or receives from its own board. SCL arrives from a controller that generates it from its clock. Nothing relates the two.
Worse, from this chapter's point of view: even a controller sharing your crystal would not help, because Chapter 19.3 established that the line's rising edge is an RC curve hundreds of nanoseconds long. The moment the line crosses your input threshold is set by a resistor and a capacitance, not by anybody's clock.
So the lines are asynchronous, permanently, by construction. The only available move is to handle it.
2. Metastability Is a Settling-Time Problem
The most common wrong picture is that metastability is a mysterious X that propagates through your design and corrupts it. That is a simulation artifact, not what hardware does.
That reframing is what makes the fix obvious. If the problem is that the output may be late, the fix is to not look at it yet.
SDA pin ──▶ ┌─────┐ ┌─────┐
│ FF1 │ ───▶ │ FF2 │ ───▶ sda_q, safe to use
└─────┘ └─────┘
▲ ▲
may be late │ │ samples FF1 a whole clock later,
at this edge │ │ by which time FF1 has settledFF2 does not prevent FF1 from going metastable — nothing does. It gives FF1 an entire clock period to finish resolving before anything reads it. The synchronizer buys settling time, and that is all it buys.
3. Three Different Jobs That Get Confused
This is the distinction most worth taking away, because the three are routinely treated as one thing and they have different costs and different tests.
| what it does | what it costs | who owns it | |
|---|---|---|---|
| synchronization | gives a flop time to settle before anything reads it | one clock of latency per stage | this chapter |
| edge detection | turns a level into a one-cycle event | one clock, one register | 18.2 |
| spike filtering | rejects pulses shorter than a threshold | latency equal to the threshold | 19.5 |
A synchronizer does not filter. A one-clock spike that happens to be sampled arrives at the output as a clean one-clock pulse, at every depth. Test T8 below proves exactly that, and mutation D09 — which makes the synchronizer quietly suppress it — is killed by eight checks because suppressing it would be doing 19.5's job in 19.4's module.
4. The Design, With the Depth as the Subject
// -----------------------------------------------------------------------------
// i2c_line_sync.sv
// A depth-parameterised two-line synchroniser, and the argument for the depth.
//
// Module 18.2 fixed SYNC_DEPTH at 2 and called it an interface contract, deferring
// the reason to this chapter. This module is the same structure with the depth as
// the subject rather than a constant, so the cost of each stage is measurable.
//
// WHAT A SYNCHRONISER DOES: it gives a flip-flop's output TIME TO SETTLE before
// anything downstream uses it. A flip-flop whose input changes inside its setup/hold
// aperture may take longer than usual to resolve to a valid level. Adding a second
// flip-flop does not prevent that -- nothing prevents it -- it gives the first one a
// whole clock period to finish resolving before the second one samples it.
//
// WHAT IT DOES NOT DO, and this is the part most descriptions get wrong:
//
// It does not "remove metastability". The probability of an unresolved level
// propagating is reduced, not eliminated, and the reduction is exponential in the
// settling time each stage allows.
//
// It does not filter. A one-clock spike that happens to be sampled arrives at the
// output as a clean one-clock pulse. Rejecting spikes is a separate mechanism with
// a separate cost, and it is Chapter 19.5.
//
// It does not align the two lines. SDA and SCL are synchronised independently and
// each may be delayed by a different amount relative to the true bus event,
// because each crosses the aperture at its own moment. T7 proves the independence.
//
// WHAT SIMULATION CAN PROVE ABOUT THIS MODULE: that it is a shift register of the
// stated depth, that its latency equals its depth, that it resets to bus-idle, and
// that the two lines do not interfere. That is the whole list.
//
// WHAT SIMULATION CANNOT PROVE: anything about metastability. An RTL simulator has
// no aperture, no resolution time constant and no analogue behaviour -- its
// flip-flops sample instantaneously and always produce a clean 0 or 1. The bench
// below states this in a comment at the point where a reader might expect such a
// test, because a test that appeared to prove it would be worse than no test.
// -----------------------------------------------------------------------------
module i2c_line_sync #(
// Stages in each synchroniser chain. TWO is the common engineering default and
// the value Module 18 uses. The chapter body derives what the number buys and
// what it costs; this module only has to make both measurable.
parameter int SYNC_DEPTH = 2
) (
input logic clk,
input logic rst_n,
// The two bus lines, straight off the pads. Asynchronous to clk: that is the
// entire problem this module exists for.
input logic scl_pin,
input logic sda_pin,
// The synchronised levels -- what the lines ARE, as far as this device can know,
// SYNC_DEPTH clocks after they were that on the wire.
output logic scl_q,
output logic sda_q,
// How many clocks of observation latency this configuration costs. Exposed as a
// port so the number appears in a bench's output and in a synthesised design's
// documentation rather than only in a comment.
output logic [3:0] latency_clocks
);
// One chain per line. RESET VALUE IS 1, not 0: an idle I²C line is HIGH, so a
// chain that reset to 0 would manufacture a falling edge at every release of
// reset -- which the framing detector of Chapter 18.3 would read as a START on a
// bus nobody is using. T2 is that test.
logic [SYNC_DEPTH-1:0] scl_chain;
logic [SYNC_DEPTH-1:0] sda_chain;
integer i;
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
scl_chain <= {SYNC_DEPTH{1'b1}};
sda_chain <= {SYNC_DEPTH{1'b1}};
end else begin
// Shift toward bit 0: the pin enters at the top, bit 0 is the settled end.
//
// WRITTEN AS AN EXPLICIT LOOP, not as `{pin, chain[DEPTH-1:1]}`. The
// concatenation form is shorter and is what the first draft used, and it
// fails to elaborate at SYNC_DEPTH = 1, where `chain[0:1]` is a part select
// in the wrong order. The loop is correct at every depth from 1 upward: at
// depth 1 the body never executes and the chain is the single sampling
// flop. The bug was found only because the bench instantiates depths 1, 2
// and 3 -- at the default of 2 the broken form works perfectly.
scl_chain[SYNC_DEPTH-1] <= scl_pin;
sda_chain[SYNC_DEPTH-1] <= sda_pin;
for (i = 0; i < SYNC_DEPTH-1; i = i + 1) begin
scl_chain[i] <= scl_chain[i+1];
sda_chain[i] <= sda_chain[i+1];
end
end
end
// The settled end of each chain. NOTHING ELSE IN THE DESIGN MAY READ A HIGHER
// BIT: bit SYNC_DEPTH-1 is the flip-flop that sampled the asynchronous pin and is
// the one whose output may not be settled. Reading it is the bug this module
// exists to prevent, and mutation D03 is exactly that.
assign scl_q = scl_chain[0];
assign sda_q = sda_chain[0];
assign latency_clocks = SYNC_DEPTH[3:0];
endmoduleThree details are worth drawing out, and one of them was a bug found by the bench.
The reset value is 1, not 0. An idle I²C line is HIGH. A chain resetting to zero presents a falling edge on SDA while SCL is high at the instant reset is released — and Chapter 18.3 defines exactly that as a START. A zero-reset synchronizer invents a START condition on a bus nobody is using, every time the device comes out of reset. Mutation D01 is one character and fails ten checks.
Nothing downstream may read the top bit. Bit SYNC_DEPTH-1 is the flop that sampled the asynchronous pin; it is the one whose output may not have settled. The whole mechanism is that this bit is read only by the next flop. Mutation D03 exports it instead, which is a synchronizer that has been structurally defeated while still looking like one — eleven failing checks.
The shift is written as a loop, because the concatenation form is broken at depth 1.
5. How Many Stages?
What the equation does tell you, without any numbers at all, is the shape of the trade:
t_settle is in the exponent. Doubling the settling time does not double the MTBF; it squares it. This is why two stages is such a common default: the improvement from one stage to two is enormous, and the improvement from two to three is enormous again but starting from an already-large number.
t_settle for a two-stage synchronizer is one clock period. So a slower clock is a more reliable synchronizer, and the same design is less safe when someone raises the system clock — a change nobody associates with the I²C block.
f_data is how often the line moves. I²C at 400 kHz changes a line at most a few hundred thousand times a second, against a system clock in the tens of megahertz. That ratio is what makes two stages comfortable here; it would not be for a source-synchronous interface changing every clock.
And the cost, which is a number you do have:
| depth | settling time allowed | observation latency | what it costs I²C |
|---|---|---|---|
| 1 | none — not a synchronizer | 1 clock | the unsettled flop feeds your logic |
| 2 | one clock period | 2 clocks | 40 ns at 50 MHz |
| 3 | two clock periods | 3 clocks | 60 ns at 50 MHz |
At 50 MHz, the difference between two and three stages is 20 ns of extra latency against a Fast-mode bit period of 2500 ns. The latency is essentially free here, which is why the honest answer for an I²C block is two, and three if a CDC report asks for it — and why the constraint work in 19.6 matters more than the exact depth.
6. Verifying What Can Be Verified
// -----------------------------------------------------------------------------
// i2c_line_sync_tb.sv
// Independent oracle for i2c_line_sync, at three depths.
//
// THREE INSTANCES, and that is the point. A bench that only ever ran at
// SYNC_DEPTH = 2 could not tell a design whose latency is the depth from one whose
// latency is a constant 2 -- the two are indistinguishable at the default. So depths
// 1, 2 and 3 run side by side on the same stimulus, and the latency of each is
// MEASURED rather than assumed.
//
// WHAT THIS BENCH DELIBERATELY DOES NOT CLAIM: anything about metastability. See the
// note above T6. Every wait is a fixed number of clocks; nothing waits on the DUT.
// -----------------------------------------------------------------------------
`timescale 1ns/1ps
module i2c_line_sync_tb;
logic clk = 1'b0, rst_n = 1'b0;
logic scl_pin = 1'b1, sda_pin = 1'b1;
logic scl_q1, sda_q1, scl_q2, sda_q2, scl_q3, sda_q3;
logic [3:0] lat1, lat2, lat3;
integer errors = 0;
integer n, d, measured;
i2c_line_sync #(.SYNC_DEPTH(1)) u1 (
.clk(clk), .rst_n(rst_n), .scl_pin(scl_pin), .sda_pin(sda_pin),
.scl_q(scl_q1), .sda_q(sda_q1), .latency_clocks(lat1));
i2c_line_sync #(.SYNC_DEPTH(2)) u2 (
.clk(clk), .rst_n(rst_n), .scl_pin(scl_pin), .sda_pin(sda_pin),
.scl_q(scl_q2), .sda_q(sda_q2), .latency_clocks(lat2));
i2c_line_sync #(.SYNC_DEPTH(3)) u3 (
.clk(clk), .rst_n(rst_n), .scl_pin(scl_pin), .sda_pin(sda_pin),
.scl_q(scl_q3), .sda_q(sda_q3), .latency_clocks(lat3));
always #5 clk = ~clk;
task step; begin @(posedge clk); @(negedge clk); end endtask
task do_reset;
begin
@(negedge clk); rst_n = 1'b0; scl_pin = 1'b1; sda_pin = 1'b1;
step; step;
@(negedge clk); rst_n = 1'b1; step;
end
endtask
task ck (input [200*8:1] what, input integer g, input integer e);
begin
if (g !== e) begin
$display(" FAIL %0s: got %0d expected %0d", what, g, e);
errors = errors + 1;
end
end
endtask
task ck_idx (input [200*8:1] what, input integer idx,
input integer g, input integer e);
begin
if (g !== e) begin
$display(" FAIL %0s[%0d]: got %0d expected %0d", what, idx, g, e);
errors = errors + 1;
end
end
endtask
// Drive one line to a value and count how many rising edges pass before the
// named synchronised output follows. Bounded at 12 so a broken DUT fails rather
// than hanging.
task measure_latency (input which, input v, output integer clocks);
begin
clocks = 0;
@(negedge clk);
if (which == 1'b0) scl_pin = v; else sda_pin = v;
while (clocks < 12) begin
@(posedge clk); #1;
clocks = clocks + 1;
if (which == 1'b0) begin
if (scl_q2 === v) disable measure_latency;
end else begin
if (sda_q2 === v) disable measure_latency;
end
end
end
endtask
initial begin
$display("=== i2c_line_sync: what a synchroniser costs, and what it cannot promise ===");
// ----------------------------------------------------------------
// T1. THE DEPTH IS REPORTED, AND IT IS THE PARAMETER. A design that ignored
// SYNC_DEPTH would still pass every latency test at one depth, so the
// reported value is checked against all three instances first.
// ----------------------------------------------------------------
do_reset;
$display("T1 each instance reports its own configured depth");
ck("T1 depth 1 reports 1", lat1, 1);
ck("T1 depth 2 reports 2", lat2, 2);
ck("T1 depth 3 reports 3", lat3, 3);
// ----------------------------------------------------------------
// T2. RESET IS BUS-IDLE, NOT ZERO. An I²C line idles HIGH. A chain that reset
// to 0 would present a falling edge on both lines at the release of reset,
// and Chapter 18.3's framing detector reads SDA falling while SCL is high
// as a START -- so a zero-reset synchroniser invents a START on an idle
// bus. This is the single most consequential line in the module.
// ----------------------------------------------------------------
@(negedge clk); rst_n = 1'b0; scl_pin = 1'b1; sda_pin = 1'b1; step; step;
$display("T2 reset presents an IDLE bus, so no edge is manufactured");
ck("T2 depth 1 SCL high in reset", scl_q1, 1);
ck("T2 depth 1 SDA high in reset", sda_q1, 1);
ck("T2 depth 2 SCL high in reset", scl_q2, 1);
ck("T2 depth 2 SDA high in reset", sda_q2, 1);
ck("T2 depth 3 SCL high in reset", scl_q3, 1);
ck("T2 depth 3 SDA high in reset", sda_q3, 1);
@(negedge clk); rst_n = 1'b1; step;
ck("T2 and still high after release", sda_q2, 1);
ck("T2 with no glitch on SCL either", scl_q2, 1);
// ----------------------------------------------------------------
// T3. LATENCY EQUALS DEPTH, MEASURED. Not asserted from the parameter: the
// bench drives an edge and counts clocks until each output follows.
// ----------------------------------------------------------------
do_reset;
sda_pin = 1'b0;
for (n = 1; n <= 4; n = n + 1) begin
@(posedge clk); #1;
// After n clocks, a depth-d chain has propagated if n >= d.
ck_idx("T3 depth 1 follows after 1 clock", n, sda_q1, (n >= 1) ? 0 : 1);
ck_idx("T3 depth 2 follows after 2 clocks", n, sda_q2, (n >= 2) ? 0 : 1);
ck_idx("T3 depth 3 follows after 3 clocks", n, sda_q3, (n >= 3) ? 0 : 1);
end
$display("T3 latency is the depth: 1, 2 and 3 clocks respectively");
// ----------------------------------------------------------------
// T4. AND THE SAME ON THE WAY BACK UP. A chain that was asymmetric -- fast on
// one edge and slow on the other -- would pass T3 and corrupt every bit
// period, because I²C uses both edges of SCL for different purposes.
// ----------------------------------------------------------------
sda_pin = 1'b1;
for (n = 1; n <= 4; n = n + 1) begin
@(posedge clk); #1;
ck_idx("T4 depth 1 rises after 1 clock", n, sda_q1, (n >= 1) ? 1 : 0);
ck_idx("T4 depth 2 rises after 2 clocks", n, sda_q2, (n >= 2) ? 1 : 0);
ck_idx("T4 depth 3 rises after 3 clocks", n, sda_q3, (n >= 3) ? 1 : 0);
end
$display("T4 the latency is symmetric: both edges cost the same");
// ----------------------------------------------------------------
// T5. MEASURED INDEPENDENTLY, VIA A COUNTING TASK. T3 checks the shape; this
// checks the number, through a different mechanism, so a mistake in one
// does not hide a mistake in the other.
// ----------------------------------------------------------------
do_reset;
measure_latency(1'b1, 1'b0, measured);
ck("T5 SDA falling took 2 clocks at depth 2", measured, 2);
measure_latency(1'b1, 1'b1, measured);
ck("T5 SDA rising took 2 clocks at depth 2", measured, 2);
measure_latency(1'b0, 1'b0, measured);
ck("T5 SCL falling took 2 clocks at depth 2", measured, 2);
$display("T5 measured latency agrees with the shape test");
// ----------------------------------------------------------------
// T6. WHAT IS NOT TESTED HERE, AND WHY.
//
// A reader arriving from the chapter body expects a metastability test at
// this point. There is none, and its absence is deliberate rather than an
// omission.
//
// An RTL simulator's flip-flop has no setup/hold aperture, no resolution
// time constant and no analogue behaviour. It samples instantaneously and
// always produces a clean 0 or 1. Driving a pin to change "at the same
// time" as the clock edge does not produce a metastable flop -- it
// produces whichever value the simulator's event ordering happens to pick,
// deterministically, every run.
//
// So a test that drove an edge into the clock edge and then asserted that
// the output was clean would PASS -- and would prove nothing whatsoever
// about the hardware, while looking exactly like evidence. That is worse
// than having no test, which is why there is none.
//
// What CAN be tested is the structural property that makes the
// synchroniser work: nothing downstream reads the first flop. T6 checks
// the consequence of that -- the output changes only through the chain, so
// a pin change is never visible at the output in the same clock.
// ----------------------------------------------------------------
do_reset;
@(negedge clk);
sda_pin = 1'b0; // change the pin between edges
#1;
$display("T6 a pin change is not visible at the output in the same clock");
ck("T6 depth 2 has not moved yet", sda_q2, 1);
ck("T6 depth 3 has not moved yet", sda_q3, 1);
@(posedge clk); #1;
ck("T6 depth 2 still not settled after one clock", sda_q2, 1);
@(posedge clk); #1;
ck("T6 and arrives on the second", sda_q2, 0);
// ----------------------------------------------------------------
// T7. THE TWO LINES ARE INDEPENDENT. Each crosses the sampling aperture at its
// own moment, so each is synchronised on its own chain. A shared chain, or
// one line's value leaking into the other's, would corrupt framing -- which
// is defined by the RELATIONSHIP between an SDA edge and the SCL level.
// ----------------------------------------------------------------
do_reset;
@(negedge clk); sda_pin = 1'b0; // SDA falls, SCL stays high
step; step; step;
$display("T7 the lines do not interfere: SDA moved, SCL did not");
ck("T7 SDA followed", sda_q2, 0);
ck("T7 SCL did not move", scl_q2, 1);
@(negedge clk); scl_pin = 1'b0; // now SCL falls, SDA stays low
step; step; step;
ck("T7 SCL followed", scl_q2, 0);
ck("T7 SDA stayed where it was", sda_q2, 0);
@(negedge clk); sda_pin = 1'b1; // SDA rises, SCL stays low
step; step; step;
ck("T7 SDA rose", sda_q2, 1);
ck("T7 SCL still low", scl_q2, 0);
// ----------------------------------------------------------------
// T8. A ONE-CLOCK SPIKE PASSES THROUGH. This is the property that separates
// this chapter from the next one. A synchroniser is not a filter: a pulse
// narrow enough to be sampled exactly once arrives at the output as a
// clean one-clock pulse, at every depth. Chapter 19.5 exists because of
// this test, and a design that passed T1-T7 and also suppressed this
// pulse would be a filter mislabelled as a synchroniser.
// ----------------------------------------------------------------
do_reset;
@(negedge clk); sda_pin = 1'b0; // low across exactly one rising edge
@(negedge clk); sda_pin = 1'b1; // back high
// The rising edge inside the spike loaded 0 into the sampling flop. The NEXT
// rising edge shifts it to the output -- so the spike is visible one clock
// after it has already ended on the wire, and for exactly one clock.
@(posedge clk); #1;
$display("T8 a one-clock spike survives synchronisation -- 19.5 owns rejecting it");
ck("T8 the spike reached the output", sda_q2, 0);
@(posedge clk); #1;
ck("T8 and it was exactly one clock wide", sda_q2, 1);
if (errors == 0) $display("=== i2c_line_sync: ALL CHECKS PASSED ===");
else $display("=== i2c_line_sync: %0d CHECK(S) FAILED ===", errors);
$finish;
end
endmoduleThe bench instantiates depths 1, 2 and 3 on the same stimulus, because a bench running only at the default cannot distinguish a design whose latency is the depth from one whose latency is a constant 2. Mutation D05 — which reports a constant 2 — is caught by T1 for exactly that reason, and D06, which ignores the depth in the shift, by twelve.
Two tests are the chapter's structural argument rather than ordinary checks.
T6 is where a metastability test would go, and it explains why there isn't one. The comment is in the published file, not just here, because that is where a reader looks for it.
T8 proves the synchronizer is not a filter. A one-clock spike goes in and a clean one-clock pulse comes out, one clock later than the spike ended on the wire. This test is the reason Chapter 19.5 exists, and it also guards the boundary in the other direction: a future "improvement" that suppressed the spike here would fail T8.
// -----------------------------------------------------------------------------
// i2c_line_sync.v
// A depth-parameterised two-line synchroniser, and the argument for the depth.
//
// Module 18.2 fixed SYNC_DEPTH at 2 and called it an interface contract, deferring
// the reason to this chapter. This module is the same structure with the depth as
// the subject rather than a constant, so the cost of each stage is measurable.
//
// WHAT A SYNCHRONISER DOES: it gives a flip-flop's output TIME TO SETTLE before
// anything downstream uses it. A flip-flop whose input changes inside its setup/hold
// aperture may take longer than usual to resolve to a valid level. Adding a second
// flip-flop does not prevent that -- nothing prevents it -- it gives the first one a
// whole clock period to finish resolving before the second one samples it.
//
// WHAT IT DOES NOT DO, and this is the part most descriptions get wrong:
//
// It does not "remove metastability". The probability of an unresolved level
// propagating is reduced, not eliminated, and the reduction is exponential in the
// settling time each stage allows.
//
// It does not filter. A one-clock spike that happens to be sampled arrives at the
// output as a clean one-clock pulse. Rejecting spikes is a separate mechanism with
// a separate cost, and it is Chapter 19.5.
//
// It does not align the two lines. SDA and SCL are synchronised independently and
// each may be delayed by a different amount relative to the true bus event,
// because each crosses the aperture at its own moment. T7 proves the independence.
//
// WHAT SIMULATION CAN PROVE ABOUT THIS MODULE: that it is a shift register of the
// stated depth, that its latency equals its depth, that it resets to bus-idle, and
// that the two lines do not interfere. That is the whole list.
//
// WHAT SIMULATION CANNOT PROVE: anything about metastability. An RTL simulator has
// no aperture, no resolution time constant and no analogue behaviour -- its
// flip-flops sample instantaneously and always produce a clean 0 or 1. The bench
// below states this in a comment at the point where a reader might expect such a
// test, because a test that appeared to prove it would be worse than no test.
// -----------------------------------------------------------------------------
module i2c_line_sync #(
// Stages in each synchroniser chain. TWO is the common engineering default and
// the value Module 18 uses. The chapter body derives what the number buys and
// what it costs; this module only has to make both measurable.
parameter SYNC_DEPTH = 2
) (
input wire clk,
input wire rst_n,
// The two bus lines, straight off the pads. Asynchronous to clk: that is the
// entire problem this module exists for.
input wire scl_pin,
input wire sda_pin,
// The synchronised levels -- what the lines ARE, as far as this device can know,
// SYNC_DEPTH clocks after they were that on the wire.
output wire scl_q,
output wire sda_q,
// How many clocks of observation latency this configuration costs. Exposed as a
// port so the number appears in a bench's output and in a synthesised design's
// documentation rather than only in a comment.
output wire [3:0] latency_clocks
);
// One chain per line. RESET VALUE IS 1, not 0: an idle I²C line is HIGH, so a
// chain that reset to 0 would manufacture a falling edge at every release of
// reset -- which the framing detector of Chapter 18.3 would read as a START on a
// bus nobody is using. T2 is that test.
reg [SYNC_DEPTH-1:0] scl_chain;
reg [SYNC_DEPTH-1:0] sda_chain;
integer i;
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
scl_chain <= {SYNC_DEPTH{1'b1}};
sda_chain <= {SYNC_DEPTH{1'b1}};
end else begin
// Shift toward bit 0: the pin enters at the top, bit 0 is the settled end.
//
// WRITTEN AS AN EXPLICIT LOOP, not as `{pin, chain[DEPTH-1:1]}`. The
// concatenation form is shorter and is what the first draft used, and it
// fails to elaborate at SYNC_DEPTH = 1, where `chain[0:1]` is a part select
// in the wrong order. The loop is correct at every depth from 1 upward: at
// depth 1 the body never executes and the chain is the single sampling
// flop. The bug was found only because the bench instantiates depths 1, 2
// and 3 -- at the default of 2 the broken form works perfectly.
scl_chain[SYNC_DEPTH-1] <= scl_pin;
sda_chain[SYNC_DEPTH-1] <= sda_pin;
for (i = 0; i < SYNC_DEPTH-1; i = i + 1) begin
scl_chain[i] <= scl_chain[i+1];
sda_chain[i] <= sda_chain[i+1];
end
end
end
// The settled end of each chain. NOTHING ELSE IN THE DESIGN MAY READ A HIGHER
// BIT: bit SYNC_DEPTH-1 is the flip-flop that sampled the asynchronous pin and is
// the one whose output may not be settled. Reading it is the bug this module
// exists to prevent, and mutation D03 is exactly that.
assign scl_q = scl_chain[0];
assign sda_q = sda_chain[0];
assign latency_clocks = SYNC_DEPTH[3:0];
endmodule // -----------------------------------------------------------------------------
// i2c_line_sync_tb.v
// Independent oracle for i2c_line_sync, at three depths.
//
// THREE INSTANCES, and that is the point. A bench that only ever ran at
// SYNC_DEPTH = 2 could not tell a design whose latency is the depth from one whose
// latency is a constant 2 -- the two are indistinguishable at the default. So depths
// 1, 2 and 3 run side by side on the same stimulus, and the latency of each is
// MEASURED rather than assumed.
//
// (Verilog-2001 -- the same tests as the SystemVerilog bench.)
//
// WHAT THIS BENCH DELIBERATELY DOES NOT CLAIM: anything about metastability. See the
// note above T6. Every wait is a fixed number of clocks; nothing waits on the DUT.
// -----------------------------------------------------------------------------
`timescale 1ns/1ps
module i2c_line_sync_tb;
reg clk = 1'b0, rst_n = 1'b0;
reg scl_pin = 1'b1, sda_pin = 1'b1;
wire scl_q1, sda_q1, scl_q2, sda_q2, scl_q3, sda_q3;
wire [3:0] lat1, lat2, lat3;
integer errors = 0;
integer n, d, measured;
i2c_line_sync #(.SYNC_DEPTH(1)) u1 (
.clk(clk), .rst_n(rst_n), .scl_pin(scl_pin), .sda_pin(sda_pin),
.scl_q(scl_q1), .sda_q(sda_q1), .latency_clocks(lat1));
i2c_line_sync #(.SYNC_DEPTH(2)) u2 (
.clk(clk), .rst_n(rst_n), .scl_pin(scl_pin), .sda_pin(sda_pin),
.scl_q(scl_q2), .sda_q(sda_q2), .latency_clocks(lat2));
i2c_line_sync #(.SYNC_DEPTH(3)) u3 (
.clk(clk), .rst_n(rst_n), .scl_pin(scl_pin), .sda_pin(sda_pin),
.scl_q(scl_q3), .sda_q(sda_q3), .latency_clocks(lat3));
always #5 clk = ~clk;
task step; begin @(posedge clk); @(negedge clk); end endtask
task do_reset;
begin
@(negedge clk); rst_n = 1'b0; scl_pin = 1'b1; sda_pin = 1'b1;
step; step;
@(negedge clk); rst_n = 1'b1; step;
end
endtask
task ck (input [200*8:1] what, input integer g, input integer e);
begin
if (g !== e) begin
$display(" FAIL %0s: got %0d expected %0d", what, g, e);
errors = errors + 1;
end
end
endtask
task ck_idx (input [200*8:1] what, input integer idx,
input integer g, input integer e);
begin
if (g !== e) begin
$display(" FAIL %0s[%0d]: got %0d expected %0d", what, idx, g, e);
errors = errors + 1;
end
end
endtask
// Drive one line to a value and count how many rising edges pass before the
// named synchronised output follows. Bounded at 12 so a broken DUT fails rather
// than hanging.
task measure_latency (input which, input v, output integer clocks);
begin
clocks = 0;
@(negedge clk);
if (which == 1'b0) scl_pin = v; else sda_pin = v;
while (clocks < 12) begin
@(posedge clk); #1;
clocks = clocks + 1;
if (which == 1'b0) begin
if (scl_q2 === v) disable measure_latency;
end else begin
if (sda_q2 === v) disable measure_latency;
end
end
end
endtask
initial begin
$display("=== i2c_line_sync: what a synchroniser costs, and what it cannot promise ===");
// ----------------------------------------------------------------
// T1. THE DEPTH IS REPORTED, AND IT IS THE PARAMETER. A design that ignored
// SYNC_DEPTH would still pass every latency test at one depth, so the
// reported value is checked against all three instances first.
// ----------------------------------------------------------------
do_reset;
$display("T1 each instance reports its own configured depth");
ck("T1 depth 1 reports 1", lat1, 1);
ck("T1 depth 2 reports 2", lat2, 2);
ck("T1 depth 3 reports 3", lat3, 3);
// ----------------------------------------------------------------
// T2. RESET IS BUS-IDLE, NOT ZERO. An I²C line idles HIGH. A chain that reset
// to 0 would present a falling edge on both lines at the release of reset,
// and Chapter 18.3's framing detector reads SDA falling while SCL is high
// as a START -- so a zero-reset synchroniser invents a START on an idle
// bus. This is the single most consequential line in the module.
// ----------------------------------------------------------------
@(negedge clk); rst_n = 1'b0; scl_pin = 1'b1; sda_pin = 1'b1; step; step;
$display("T2 reset presents an IDLE bus, so no edge is manufactured");
ck("T2 depth 1 SCL high in reset", scl_q1, 1);
ck("T2 depth 1 SDA high in reset", sda_q1, 1);
ck("T2 depth 2 SCL high in reset", scl_q2, 1);
ck("T2 depth 2 SDA high in reset", sda_q2, 1);
ck("T2 depth 3 SCL high in reset", scl_q3, 1);
ck("T2 depth 3 SDA high in reset", sda_q3, 1);
@(negedge clk); rst_n = 1'b1; step;
ck("T2 and still high after release", sda_q2, 1);
ck("T2 with no glitch on SCL either", scl_q2, 1);
// ----------------------------------------------------------------
// T3. LATENCY EQUALS DEPTH, MEASURED. Not asserted from the parameter: the
// bench drives an edge and counts clocks until each output follows.
// ----------------------------------------------------------------
do_reset;
sda_pin = 1'b0;
for (n = 1; n <= 4; n = n + 1) begin
@(posedge clk); #1;
// After n clocks, a depth-d chain has propagated if n >= d.
ck_idx("T3 depth 1 follows after 1 clock", n, sda_q1, (n >= 1) ? 0 : 1);
ck_idx("T3 depth 2 follows after 2 clocks", n, sda_q2, (n >= 2) ? 0 : 1);
ck_idx("T3 depth 3 follows after 3 clocks", n, sda_q3, (n >= 3) ? 0 : 1);
end
$display("T3 latency is the depth: 1, 2 and 3 clocks respectively");
// ----------------------------------------------------------------
// T4. AND THE SAME ON THE WAY BACK UP. A chain that was asymmetric -- fast on
// one edge and slow on the other -- would pass T3 and corrupt every bit
// period, because I²C uses both edges of SCL for different purposes.
// ----------------------------------------------------------------
sda_pin = 1'b1;
for (n = 1; n <= 4; n = n + 1) begin
@(posedge clk); #1;
ck_idx("T4 depth 1 rises after 1 clock", n, sda_q1, (n >= 1) ? 1 : 0);
ck_idx("T4 depth 2 rises after 2 clocks", n, sda_q2, (n >= 2) ? 1 : 0);
ck_idx("T4 depth 3 rises after 3 clocks", n, sda_q3, (n >= 3) ? 1 : 0);
end
$display("T4 the latency is symmetric: both edges cost the same");
// ----------------------------------------------------------------
// T5. MEASURED INDEPENDENTLY, VIA A COUNTING TASK. T3 checks the shape; this
// checks the number, through a different mechanism, so a mistake in one
// does not hide a mistake in the other.
// ----------------------------------------------------------------
do_reset;
measure_latency(1'b1, 1'b0, measured);
ck("T5 SDA falling took 2 clocks at depth 2", measured, 2);
measure_latency(1'b1, 1'b1, measured);
ck("T5 SDA rising took 2 clocks at depth 2", measured, 2);
measure_latency(1'b0, 1'b0, measured);
ck("T5 SCL falling took 2 clocks at depth 2", measured, 2);
$display("T5 measured latency agrees with the shape test");
// ----------------------------------------------------------------
// T6. WHAT IS NOT TESTED HERE, AND WHY.
//
// A reader arriving from the chapter body expects a metastability test at
// this point. There is none, and its absence is deliberate rather than an
// omission.
//
// An RTL simulator's flip-flop has no setup/hold aperture, no resolution
// time constant and no analogue behaviour. It samples instantaneously and
// always produces a clean 0 or 1. Driving a pin to change "at the same
// time" as the clock edge does not produce a metastable flop -- it
// produces whichever value the simulator's event ordering happens to pick,
// deterministically, every run.
//
// So a test that drove an edge into the clock edge and then asserted that
// the output was clean would PASS -- and would prove nothing whatsoever
// about the hardware, while looking exactly like evidence. That is worse
// than having no test, which is why there is none.
//
// What CAN be tested is the structural property that makes the
// synchroniser work: nothing downstream reads the first flop. T6 checks
// the consequence of that -- the output changes only through the chain, so
// a pin change is never visible at the output in the same clock.
// ----------------------------------------------------------------
do_reset;
@(negedge clk);
sda_pin = 1'b0; // change the pin between edges
#1;
$display("T6 a pin change is not visible at the output in the same clock");
ck("T6 depth 2 has not moved yet", sda_q2, 1);
ck("T6 depth 3 has not moved yet", sda_q3, 1);
@(posedge clk); #1;
ck("T6 depth 2 still not settled after one clock", sda_q2, 1);
@(posedge clk); #1;
ck("T6 and arrives on the second", sda_q2, 0);
// ----------------------------------------------------------------
// T7. THE TWO LINES ARE INDEPENDENT. Each crosses the sampling aperture at its
// own moment, so each is synchronised on its own chain. A shared chain, or
// one line's value leaking into the other's, would corrupt framing -- which
// is defined by the RELATIONSHIP between an SDA edge and the SCL level.
// ----------------------------------------------------------------
do_reset;
@(negedge clk); sda_pin = 1'b0; // SDA falls, SCL stays high
step; step; step;
$display("T7 the lines do not interfere: SDA moved, SCL did not");
ck("T7 SDA followed", sda_q2, 0);
ck("T7 SCL did not move", scl_q2, 1);
@(negedge clk); scl_pin = 1'b0; // now SCL falls, SDA stays low
step; step; step;
ck("T7 SCL followed", scl_q2, 0);
ck("T7 SDA stayed where it was", sda_q2, 0);
@(negedge clk); sda_pin = 1'b1; // SDA rises, SCL stays low
step; step; step;
ck("T7 SDA rose", sda_q2, 1);
ck("T7 SCL still low", scl_q2, 0);
// ----------------------------------------------------------------
// T8. A ONE-CLOCK SPIKE PASSES THROUGH. This is the property that separates
// this chapter from the next one. A synchroniser is not a filter: a pulse
// narrow enough to be sampled exactly once arrives at the output as a
// clean one-clock pulse, at every depth. Chapter 19.5 exists because of
// this test, and a design that passed T1-T7 and also suppressed this
// pulse would be a filter mislabelled as a synchroniser.
// ----------------------------------------------------------------
do_reset;
@(negedge clk); sda_pin = 1'b0; // low across exactly one rising edge
@(negedge clk); sda_pin = 1'b1; // back high
// The rising edge inside the spike loaded 0 into the sampling flop. The NEXT
// rising edge shifts it to the output -- so the spike is visible one clock
// after it has already ended on the wire, and for exactly one clock.
@(posedge clk); #1;
$display("T8 a one-clock spike survives synchronisation -- 19.5 owns rejecting it");
ck("T8 the spike reached the output", sda_q2, 0);
@(posedge clk); #1;
ck("T8 and it was exactly one clock wide", sda_q2, 1);
if (errors == 0) $display("=== i2c_line_sync: ALL CHECKS PASSED ===");
else $display("=== i2c_line_sync: %0d CHECK(S) FAILED ===", errors);
$finish;
end
endmodule -- -----------------------------------------------------------------------------
-- i2c_line_sync.vhd
-- A depth-parameterised two-line synchroniser, and the argument for the depth.
-- Behavioural twin of the SystemVerilog and Verilog designs.
--
-- Module 18.2 fixed SYNC_DEPTH at 2 and called it an interface contract, deferring
-- the reason to this chapter. Here the depth is the subject, so the cost of each
-- stage is measurable.
--
-- WHAT A SYNCHRONISER DOES: it gives a flip-flop's output time to settle before
-- anything downstream uses it. It does NOT remove metastability, it does NOT filter
-- spikes (Chapter 19.5), and it does NOT align the two lines with each other.
--
-- WHAT SIMULATION CAN PROVE HERE: the chain depth, the latency, the reset value and
-- the independence of the two lines. Nothing about metastability -- an RTL simulator
-- has no sampling aperture and no resolution time constant.
--
-- NOTE THE NULL RANGE. The shift loop is `for i in 0 to SYNC_DEPTH-2`, which at
-- SYNC_DEPTH = 1 is the null range `0 to -1` and correctly does nothing. The
-- SystemVerilog original first used a part select instead and failed to elaborate at
-- depth 1; the bench instantiates depths 1, 2 and 3 precisely so that the boundary
-- value is exercised rather than assumed.
-- -----------------------------------------------------------------------------
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity i2c_line_sync is
generic (
-- Stages in each synchroniser chain. TWO is the common engineering default and
-- the value Module 18 uses.
SYNC_DEPTH : positive := 2
);
port (
clk : in std_logic;
rst_n : in std_logic;
-- The two bus lines, straight off the pads. Asynchronous to clk.
scl_pin : in std_logic;
sda_pin : in std_logic;
-- The synchronised levels, SYNC_DEPTH clocks behind the wire.
scl_q : out std_logic;
sda_q : out std_logic;
-- How many clocks of observation latency this configuration costs.
latency_clocks : out unsigned(3 downto 0)
);
end entity i2c_line_sync;
architecture rtl of i2c_line_sync is
-- RESET VALUE IS '1', not '0': an idle I²C line is HIGH, so a chain resetting to
-- '0' would manufacture a falling edge at every release of reset -- which
-- Chapter 18.3's framing detector reads as a START on a bus nobody is using.
signal scl_chain : std_logic_vector(SYNC_DEPTH-1 downto 0) := (others => '1');
signal sda_chain : std_logic_vector(SYNC_DEPTH-1 downto 0) := (others => '1');
begin
process (clk, rst_n)
begin
if rst_n = '0' then
scl_chain <= (others => '1');
sda_chain <= (others => '1');
elsif rising_edge(clk) then
-- The pin enters at the top; bit 0 is the settled end.
scl_chain(SYNC_DEPTH-1) <= scl_pin;
sda_chain(SYNC_DEPTH-1) <= sda_pin;
for i in 0 to SYNC_DEPTH-2 loop
scl_chain(i) <= scl_chain(i+1);
sda_chain(i) <= sda_chain(i+1);
end loop;
end if;
end process;
-- The settled end of each chain. NOTHING ELSE MAY READ A HIGHER BIT: the top bit
-- is the flop that sampled the asynchronous pin and is the one whose output may
-- not have settled. Mutation D03 reads it, and is killed by eleven checks.
scl_q <= scl_chain(0);
sda_q <= sda_chain(0);
latency_clocks <= to_unsigned(SYNC_DEPTH, 4);
end architecture rtl; -- -----------------------------------------------------------------------------
-- i2c_line_sync_tb.vhd
-- Independent oracle for i2c_line_sync, at three depths.
-- Behavioural twin of the SystemVerilog and Verilog benches.
--
-- THREE INSTANCES, and that is the point. A bench that only ever ran at
-- SYNC_DEPTH = 2 could not tell a design whose latency is the depth from one whose
-- latency is a constant 2. Depths 1, 2 and 3 run on the same stimulus, and the
-- latency of each is MEASURED rather than assumed.
--
-- WHAT THIS BENCH DELIBERATELY DOES NOT CLAIM: anything about metastability. See the
-- note at T6.
-- -----------------------------------------------------------------------------
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity i2c_line_sync_tb is
end entity i2c_line_sync_tb;
architecture sim of i2c_line_sync_tb is
signal clk : std_logic := '0';
signal rst_n : std_logic := '0';
signal scl_pin : std_logic := '1';
signal sda_pin : std_logic := '1';
signal scl_q1, sda_q1, scl_q2, sda_q2, scl_q3, sda_q3 : std_logic;
signal lat1, lat2, lat3 : unsigned(3 downto 0);
signal halt : boolean := false;
begin
clkgen : process
begin
while not halt loop
clk <= '0'; wait for 5 ns;
clk <= '1'; wait for 5 ns;
end loop;
wait;
end process;
u1 : entity work.i2c_line_sync
generic map (SYNC_DEPTH => 1)
port map (clk => clk, rst_n => rst_n, scl_pin => scl_pin, sda_pin => sda_pin,
scl_q => scl_q1, sda_q => sda_q1, latency_clocks => lat1);
u2 : entity work.i2c_line_sync
generic map (SYNC_DEPTH => 2)
port map (clk => clk, rst_n => rst_n, scl_pin => scl_pin, sda_pin => sda_pin,
scl_q => scl_q2, sda_q => sda_q2, latency_clocks => lat2);
u3 : entity work.i2c_line_sync
generic map (SYNC_DEPTH => 3)
port map (clk => clk, rst_n => rst_n, scl_pin => scl_pin, sda_pin => sda_pin,
scl_q => scl_q3, sda_q => sda_q3, latency_clocks => lat3);
stim : process
variable err : integer := 0;
variable measured : integer;
function b2i (b : std_logic) return integer is
begin
if b = '1' then return 1; else return 0; end if;
end function;
procedure step is
begin
wait until rising_edge(clk);
wait until falling_edge(clk);
end procedure;
procedure do_reset is
begin
wait until falling_edge(clk);
rst_n <= '0'; scl_pin <= '1'; sda_pin <= '1';
step; step;
wait until falling_edge(clk);
rst_n <= '1';
step;
end procedure;
procedure ck (what : string; g : integer; e : integer) is
begin
if g /= e then
report " FAIL " & what & ": got " & integer'image(g)
& " expected " & integer'image(e) severity note;
err := err + 1;
end if;
end procedure;
procedure ck_idx (what : string; idx : integer; g : integer; e : integer) is
begin
if g /= e then
report " FAIL " & what & "[" & integer'image(idx) & "]: got "
& integer'image(g) & " expected " & integer'image(e) severity note;
err := err + 1;
end if;
end procedure;
begin
report "=== i2c_line_sync: what a synchroniser costs, and what it cannot promise ==="
severity note;
-- T1. The depth is reported, and it is the parameter. A design that ignored
-- SYNC_DEPTH would still pass every latency test at one depth.
do_reset;
report "T1 each instance reports its own configured depth" severity note;
ck("T1 depth 1 reports 1", to_integer(lat1), 1);
ck("T1 depth 2 reports 2", to_integer(lat2), 2);
ck("T1 depth 3 reports 3", to_integer(lat3), 3);
-- T2. Reset is bus-idle, not zero. An I²C line idles HIGH. A chain resetting
-- to '0' presents a falling edge on both lines at the release of reset, and
-- Chapter 18.3 reads SDA falling while SCL is high as a START -- so a
-- zero-reset synchroniser invents a START on an idle bus.
wait until falling_edge(clk);
rst_n <= '0'; scl_pin <= '1'; sda_pin <= '1'; step; step;
report "T2 reset presents an IDLE bus, so no edge is manufactured" severity note;
ck("T2 depth 1 SCL high in reset", b2i(scl_q1), 1);
ck("T2 depth 1 SDA high in reset", b2i(sda_q1), 1);
ck("T2 depth 2 SCL high in reset", b2i(scl_q2), 1);
ck("T2 depth 2 SDA high in reset", b2i(sda_q2), 1);
ck("T2 depth 3 SCL high in reset", b2i(scl_q3), 1);
ck("T2 depth 3 SDA high in reset", b2i(sda_q3), 1);
wait until falling_edge(clk);
rst_n <= '1'; step;
ck("T2 and still high after release", b2i(sda_q2), 1);
ck("T2 with no glitch on SCL either", b2i(scl_q2), 1);
-- T3. Latency equals depth, measured rather than asserted from the parameter.
do_reset;
sda_pin <= '0';
for n in 1 to 4 loop
wait until rising_edge(clk); wait for 1 ns;
if n >= 1 then ck_idx("T3 depth 1 follows after 1 clock", n, b2i(sda_q1), 0);
else ck_idx("T3 depth 1 follows after 1 clock", n, b2i(sda_q1), 1); end if;
if n >= 2 then ck_idx("T3 depth 2 follows after 2 clocks", n, b2i(sda_q2), 0);
else ck_idx("T3 depth 2 follows after 2 clocks", n, b2i(sda_q2), 1); end if;
if n >= 3 then ck_idx("T3 depth 3 follows after 3 clocks", n, b2i(sda_q3), 0);
else ck_idx("T3 depth 3 follows after 3 clocks", n, b2i(sda_q3), 1); end if;
end loop;
report "T3 latency is the depth: 1, 2 and 3 clocks respectively" severity note;
-- T4. And the same on the way back up. An asymmetric chain -- fast on one edge
-- and slow on the other -- would pass T3 and corrupt every bit period,
-- because I²C uses both edges of SCL for different purposes.
sda_pin <= '1';
for n in 1 to 4 loop
wait until rising_edge(clk); wait for 1 ns;
if n >= 1 then ck_idx("T4 depth 1 rises after 1 clock", n, b2i(sda_q1), 1);
else ck_idx("T4 depth 1 rises after 1 clock", n, b2i(sda_q1), 0); end if;
if n >= 2 then ck_idx("T4 depth 2 rises after 2 clocks", n, b2i(sda_q2), 1);
else ck_idx("T4 depth 2 rises after 2 clocks", n, b2i(sda_q2), 0); end if;
if n >= 3 then ck_idx("T4 depth 3 rises after 3 clocks", n, b2i(sda_q3), 1);
else ck_idx("T4 depth 3 rises after 3 clocks", n, b2i(sda_q3), 0); end if;
end loop;
report "T4 the latency is symmetric: both edges cost the same" severity note;
-- T5. Measured independently, by counting. T3 checks the shape; this checks the
-- number through a different mechanism, so a mistake in one does not hide a
-- mistake in the other. The loop is bounded at 12 so a broken DUT fails.
do_reset;
sda_pin <= '0';
measured := 0;
for k in 1 to 12 loop
wait until rising_edge(clk); wait for 1 ns;
measured := measured + 1;
exit when sda_q2 = '0';
end loop;
ck("T5 SDA falling took 2 clocks at depth 2", measured, 2);
sda_pin <= '1';
measured := 0;
for k in 1 to 12 loop
wait until rising_edge(clk); wait for 1 ns;
measured := measured + 1;
exit when sda_q2 = '1';
end loop;
ck("T5 SDA rising took 2 clocks at depth 2", measured, 2);
scl_pin <= '0';
measured := 0;
for k in 1 to 12 loop
wait until rising_edge(clk); wait for 1 ns;
measured := measured + 1;
exit when scl_q2 = '0';
end loop;
ck("T5 SCL falling took 2 clocks at depth 2", measured, 2);
report "T5 measured latency agrees with the shape test" severity note;
-- T6. What is NOT tested here, and why.
--
-- A reader arriving from the chapter body expects a metastability test at
-- this point. There is none, and its absence is deliberate.
--
-- An RTL simulator's flip-flop has no setup/hold aperture, no resolution
-- time constant and no analogue behaviour. Driving a pin to change "at the
-- same time" as the clock edge does not produce a metastable flop -- it
-- produces whichever value the simulator's event ordering picks,
-- deterministically, every run. A test that drove an edge into the clock
-- edge and then asserted the output was clean would PASS and prove nothing,
-- while looking exactly like evidence. That is worse than no test.
--
-- What CAN be tested is the structural property: nothing downstream reads
-- the first flop, so a pin change is never visible at the output in the
-- same clock.
do_reset;
wait until falling_edge(clk);
sda_pin <= '0';
wait for 1 ns;
report "T6 a pin change is not visible at the output in the same clock"
severity note;
ck("T6 depth 2 has not moved yet", b2i(sda_q2), 1);
ck("T6 depth 3 has not moved yet", b2i(sda_q3), 1);
wait until rising_edge(clk); wait for 1 ns;
ck("T6 depth 2 still not settled after one clock", b2i(sda_q2), 1);
wait until rising_edge(clk); wait for 1 ns;
ck("T6 and arrives on the second", b2i(sda_q2), 0);
-- T7. The two lines are independent. Each crosses the sampling aperture at its
-- own moment, so each is synchronised on its own chain. A shared chain
-- would corrupt framing, which is defined by the RELATIONSHIP between an
-- SDA edge and the SCL level.
do_reset;
wait until falling_edge(clk); sda_pin <= '0';
step; step; step;
report "T7 the lines do not interfere: SDA moved, SCL did not" severity note;
ck("T7 SDA followed", b2i(sda_q2), 0);
ck("T7 SCL did not move", b2i(scl_q2), 1);
wait until falling_edge(clk); scl_pin <= '0';
step; step; step;
ck("T7 SCL followed", b2i(scl_q2), 0);
ck("T7 SDA stayed where it was", b2i(sda_q2), 0);
wait until falling_edge(clk); sda_pin <= '1';
step; step; step;
ck("T7 SDA rose", b2i(sda_q2), 1);
ck("T7 SCL still low", b2i(scl_q2), 0);
-- T8. A one-clock spike passes through. This is the property that separates
-- this chapter from the next one. A synchroniser is not a filter: a pulse
-- narrow enough to be sampled exactly once arrives at the output as a
-- clean one-clock pulse. Chapter 19.5 exists because of this test.
do_reset;
wait until falling_edge(clk); sda_pin <= '0'; -- low across one rising edge
wait until falling_edge(clk); sda_pin <= '1'; -- back high
-- The rising edge inside the spike loaded '0' into the sampling flop. The NEXT
-- rising edge shifts it to the output -- so the spike is visible one clock
-- after it has already ended on the wire, and for exactly one clock.
wait until rising_edge(clk); wait for 1 ns;
report "T8 a one-clock spike survives synchronisation -- 19.5 owns rejecting it"
severity note;
ck("T8 the spike reached the output", b2i(sda_q2), 0);
wait until rising_edge(clk); wait for 1 ns;
ck("T8 and it was exactly one clock wide", b2i(sda_q2), 1);
if err = 0 then
report "=== i2c_line_sync: ALL CHECKS PASSED ===" severity note;
else
report "=== i2c_line_sync: " & integer'image(err) & " CHECK(S) FAILED ==="
severity note;
end if;
halt <= true;
wait;
end process;
end architecture sim;7. What the Mutations Found, Including One That Does Not Count
| # | mutation | verdict | what caught it |
|---|---|---|---|
| D01 | chains reset to 0 | KILLED (10) | T2 — a START invented on an idle bus |
| D02 | only SCL resets to 0 | KILLED (4) | T2, on one line |
| D03 | output taken from the first flop | KILLED (11) | T6 — the pin becomes visible too early |
| D04 | output taken straight from the pin | KILLED (12) | T3/T6 — no latency at all |
| D05 | latency report a constant 2 | KILLED (2) | T1 — only visible at depths 1 and 3 |
| D06 | depth ignored, chain always 2 deep | KILLED (12) | T1/T3 at depths 1 and 3 |
| D07 | the two lines share one chain | KILLED (3) | T7 — SDA leaking into SCL |
| D08 | shift direction reversed | KILLED (17) | T3/T4 |
| D09 | synchronizer also filters | KILLED (8) | T8 — 19.5's job done in 19.4's module |
8. Where This Sits Relative to Module 18
Chapter 18.2's i2c_slave_sync bundles synchronization and edge detection into one block with the depth fixed at 2, because the rest of Module 18 needed a settled level and one-cycle events and did not need the depth to be a variable.
This chapter's i2c_line_sync is deliberately narrower: synchronization only, depth as the subject. The two are not competitors — 18.2 is what you instantiate, and this is the argument for why its SYNC_DEPTH port exists and what happens if you change it. The parameter is exposed on i2c_slave all the way up to the top level for precisely this reason.
9. Focused Verification Insight
A monitor must not read the pin. A verification monitor sampling the raw asynchronous line in an RTL simulation gets a clean value, because the simulator has no aperture — so the monitor appears to work and is measuring something the hardware does not have. Module 20's monitor should sample the synchronized level, and should account for the latency when it timestamps an event.
Coverage worth collecting at this layer is about the latency and the boundary, not the protocol: a transition observed at each instantiated depth; both lines moving independently; a change arriving in the clock immediately before a protocol decision, which is the case where the latency matters most; and a one-clock spike passing through, so the synchronizer's non-filtering is a covered fact rather than an assumed one.
Assertions, as concepts — Icarus supports no concurrent assertions, so nothing here is executed:
// The output changes only through the chain: it can never equal a pin value that
// the pin took on less than SYNC_DEPTH clocks ago. Written for depth 2.
property p_output_lags_by_depth;
@(posedge clk) disable iff (!rst_n)
$changed(sda_q) |-> $past(sda_pin, 2) == sda_q;
endproperty
// Reset presents an idle bus, so no framing event can be manufactured.
property p_reset_is_idle;
@(posedge clk) (!rst_n) |-> (sda_q && scl_q);
endproperty
// THE PROPERTY THAT CANNOT BE WRITTEN, and is worth naming so nobody thinks it was
// forgotten: "the first flop always settles within one clock period". It is not
// expressible in SVA and not checkable in RTL simulation, because the simulator has
// no aperture and no resolution time constant. It is an argument about silicon,
// supported by the MTBF equation and by a CDC report -- never by a passing test.10. Misconceptions
11. Debugging
The target answers a START that never happened
Pitfall — a synchroniser chain that resets to zero
// A synchroniser written from a generic CDC template, where the template's reset
// value was 0 because the signal it was originally written for was active-high:
//
// always @(posedge clk or negedge rst_n)
// if (!rst_n) begin
// sda_sync <= 2'b00; // <-- wrong for a bus that idles HIGH
// scl_sync <= 2'b00;
// end else begin
// sda_sync <= {sda_pin, sda_sync[1]};
// scl_sync <= {scl_pin, scl_sync[1]};
// end
//
// Nothing about this is wrong as a synchroniser. The latency is right, the depth is
// right, the two lines are independent. The reset VALUE is wrong for this signal,
// and the reason it matters is entirely downstream: Chapter 18.3 defines a START as
// SDA falling while SCL is high.
//
// At the release of reset, with a genuinely idle bus (both lines pulled high):
//
// cycle 0 rst_n low sda_q = 0, scl_q = 0 (the reset values)
// cycle 1 rst_n high the chains begin loading the real pin values
// cycle 2 scl_q -> 1 and sda_q -> 1
//
// SCL and SDA both rise, and depending on which chain the framing block sees move
// first, it may observe SDA rising while SCL is high -- a STOP -- or, if the pins
// are not both exactly high, SDA falling while SCL is high, which is a START.The target's transaction counter increments once, immediately after every reset, with no controller attached to the board at all. The bus is idle; both lines measure a clean 3.3 V; nothing is clocking.
Worse, and this is what made it a two-day bug: the phantom event sometimes leaves the target SELECTED. Its address comparator ran on whatever the shift register held after the manufactured START, and on one board in five that garbage matched. A selected target then answers the next real transfer's address byte with an ACK it has no right to send, so the FIRST genuine transaction after power-on fails -- and retrying always works, because by then the phantom START has been cleared by a real STOP.
An ILA triggered on the framing block's start_pulse shows it firing two clocks after rst_n rises, with the pins both high the whole time. That is the giveaway: the event is generated from internal state, not from anything on the wire.
The chains reset to 0 while the bus they model is idle at 1, so the release of reset produces a transition on both synchronised lines that never happened on the pins. Chapter 18.3's framing detector is doing its job correctly -- it is faithfully reporting an edge in the data it was given. The data was manufactured by the reset value.
The general statement is the useful one: A SYNCHRONISER'S RESET VALUE MUST BE THE IDLE STATE OF THE SIGNAL IT CARRIES, because the release of reset is itself an event and every edge-sensitive consumer downstream will see it. For an I2C line that idle state is HIGH. For an active-high request it would be 0. A generic CDC template cannot know which, so this is the one line of a template you must always look at.
Note what makes it hard to find: it is not a protocol bug, not a timing bug, and not a bug in either module involved. The synchroniser is a correct synchroniser and the framing detector is a correct framing detector. The defect lives in the initial condition at the seam between them, which is exactly the class of fault that survives unit testing of both blocks.
The fix is two characters per line. The test that catches it is T2, which checks both synchronised outputs during reset AND immediately after its release, at all three depths -- and mutation D01 is this exact bug, killed by ten checks.
12. Reason It Through
13. Questions
14. What This Chapter Settled
The two bus lines are asynchronous to the system clock permanently, and no clock frequency changes that. Metastability is a settling-time problem, not a propagating X, so the fix is a chain that gives the sampling flop a clock period to resolve before anything reads it. Two stages is the common default; the cost is one clock of observation latency per stage, which at 50 MHz against a Fast-mode bit period is negligible.
Synchronization, edge detection and spike filtering are three different jobs with three different costs. T8 pins the boundary: a one-clock spike passes straight through a synchronizer of any depth.
Which is the problem Chapter 19.5 takes up. The bus now arrives as a settled level, and that level faithfully reports every disturbance on the wire — including the ones the specification allows a device to ignore. Rejecting them costs latency, and spending too much of it deletes the edges you needed.
Continue learning
Related tutorials
- Related topic
RTL and FPGA Bugs That Only Appear on Hardware
The last class of I²C failure: bugs whose evidence is not on the bus. Separates what simulation genuinely cannot show — metastability — from the larger set nobody modelled, turns a frequency-dependent filter bug into a measured boundary, and closes the module by naming, for each failure class, the observation point where its evidence actually lives.
- Related topic
FPGA as Master and as Target — Integration Patterns
Assembles everything: Module 18's verified target behind Module 19's pad, synchronizer and filter, proven end to end on a wired-AND bus in three languages. Shows the controller-side shape on Module 17's real interface, where a soft CPU attaches, and closes with two mutations that cannot be killed in simulation — the module's thesis stated as evidence.
- Related topic
Hardware Bring-Up and On-Chip Debug
What to do on the morning the board arrives, in what order, with what instrument. Builds a synthesizable bus-health block that answers 'is the bus even there' at a glance, works through a probe set that spans layers rather than concentrating on the FSM, and shows why an analyzer and an ILA disagreeing is the most localising evidence available.
- Related topic
Simulation vs Synthesized Hardware — Where Behavior Diverges
Eight classes of mismatch that pass in RTL simulation and fail on a board, each with why simulation passes, what hardware does instead, and what evidence exposes it. Includes a bus model with a rise time, and an experiment running Module 18's verified target against it whose result is not the expected one.
