Skip to content
VLSI Mentor

I²C · Module 22

Corner Cases — Stuck Lines, Excessive Stretch and Reset Mid-Transaction

The conditions that break real controllers are hostile without being illegal, so no assertion catches them. Covers why a stuck line and a legal stretch are the same observation, and states plainly what reset mid-transfer this curriculum does not verify.

The conditions that actually break shipped I²C controllers have an awkward property in common: they are not protocol violations. A line held down forever is legal at every instant. A stretch that lasts a second breaks no rule. A reset in the middle of a transfer is not a bus event at all.

So there is nothing for an assertion to catch, and this chapter is about what catches them instead.

This is the central difficulty and it is not solvable by better observation.

During a clock stretch: SCL is low, the controller has released it, and it is not rising. During a stuck SCL: SCL is low, the controller has released it, and it is not rising.

There is no difference. Not a subtle one — none. The only thing that distinguishes them is duration, and duration is not a protocol fact: the specification places no upper bound on a stretch.

So the "check" for a stuck line is a policy, not a property: the environment chooses how patient to be, and that choice is a parameter rather than a derived fact. Chapter 20.5's driver takes STRETCH_TIMEOUT for exactly this reason.

2. What a Wedged Bus Does to an Environment

A stuck line is the corner case with the widest blast radius, because it breaks the machinery that reports failures.

what holdsconsequence
SDA held lowa STOP requires SDA to rise while SCL is high — so the bus cannot be framed at all
SCL held lowno data moves, and every bounded wait in the environment expires

And the target's own defence matters as much: Module 18 abandons a transfer after IDLE_CYCLES of a held bus rather than waiting forever. 20.9's T4 confirms it — with SDA held, n_aborts increments and the target releases, so the bus recovers the moment the fault is removed.

3. Excessive Stretch Is a Budget Question

A target that stretches for longer than the system can tolerate is not misbehaving. It is doing something legal that the rest of the system cannot afford.

That makes it a budget rather than a rule, and budgets are checked differently:

  • the driver's STRETCH_TIMEOUT is the environment's patience, chosen and documented;
  • the target's IDLE_CYCLES is its own patience before abandoning a transfer;
  • the two must be consistent — a driver more patient than the target's abort threshold will see the target give up mid-transfer and report it as a target fault.

4. Reset Mid-Transaction

The case most likely to be missing from a plan, and the one this curriculum covers least.

A reset asserted between transfers is trivial and every bench does it. A reset asserted during a transfer is different in kind: the target must release both lines, forget its transfer state, and — critically — not leave the bus held, because the controller is still clocking and will keep clocking until it frames a STOP.

The questions a plan needs, which are design decisions rather than protocol rules:

questionwhy it is not derivable
does the register pointer survive a reset?Module 18's D3 says framing does not clear it; reset is not framing
do the diagnostic counters clear?a counter that clears loses the evidence of what happened
does the target release both lines immediately?the alternative wedges the bus for the next transfer
what does the controller see?a transfer that stops being acknowledged part way through

5. Address Mismatch Is the Corner Case That Looks Boring

One more, because it is routinely under-tested in a specific way.

A target must not answer an address that is not its own. Testing that with an address four bits away proves almost nothing: a comparator missing any single bit still rejects it. Chapter 20.8's mutation P08 dropped one address bit and survived a foreign-address test at 0x21 against a target at 0x50 — four bits of difference.

Only a one-bit near miss demonstrates that a specific bit participates. The test that kills it uses 0x10, and the general rule is that a near-miss test's value is inversely proportional to its Hamming distance.

Four consecutive tests failed and the culprit passed

Pitfall — a test that left the bus wedged
Buggy Code
// An error-injection test that holds SDA low, then ends:
//
//    fault <= F_SDA_STUCK;
//    arm   <= '1';
//    for i in 1 to 40 loop step; end loop;
//    -- checks pass: the line is held, the target aborted, all good
//    -- ... and the test ends here, with arm still '1'
//
// The test PASSES. Its own checks were satisfied.
//
// The next four tests fail identically: the target never acknowledges. Because
// SDA is still held low by an injector nobody disarmed, so no START can be framed
// -- a START needs an SDA FALL, and the line is already low.
//
// The failure report names four innocent tests. The first one debugged is the one
// AFTER the culprit, and it looks like an addressing fault.
Root Cause

A held line disables the framing mechanism itself, so every subsequent test fails at its first START and the reports all look like addressing faults. The culprit passes because its own assertions were about the fault it injected, not about the state it left behind. Releasing on disarm removes the possibility, and a bus-idle check at the end of every test makes any remaining leak attributable to the test that caused it.

Fix
// Two defences, and the second one is what makes failures attributable.
//
// 1. THE INJECTOR RELEASES ON DISARM, unconditionally, so a fault cannot outlive
//    the test that asked for it:
//
//       if arm = '0' then
//          scl_drive_low <= '0';
//          sda_drive_low <= '0';
//          injecting     <= '0';
//          fired         <= '0';       -- and its one-shot state resets too
//       end if;
//
// 2. EVERY TEST ENDS WITH A BUS-IDLE CHECK:
//
//       if scl /= '1' or sda /= '1' then
//          report "test finished with a line still held" severity error;
//       end if;
//
// Now the test that breaks the bus is the test that reports, instead of the four
// that inherit it.
//
// THE DIAGNOSTIC for the symptom -- several consecutive identical failures with a
// pass before them: SUSPECT THE TEST BEFORE THE FIRST FAILURE. A wedged bus
// produces a cascade whose head is innocent, and the per-test idle check turns
// that cascade into one message.
Pitfall — a stretch timeout shorter than a half period
Buggy Code
// An environment configured for a fast bus, with the stretch patience copied
// from an older project:
//
//    cfg.half_period     = 64;      // a slow bus this time
//    cfg.stretch_timeout = 40;      // ... and the old value
//
// EVERY transfer now times out. The driver releases SCL, waits 40 cycles for it
// to rise, and gives up -- but a half period is 64 cycles, so on this bus SCL is
// still legitimately low when the patience expires.
//
// The symptom is "the target never responds". Every transfer reports a stretch
// timeout, the target looks dead, and the investigation goes to the target's
// clock, its reset, and its address decode -- none of which is involved.
//
// Nothing illegal happens on the bus at any point, so no property fires.
Root Cause

A misconfigured timeout produces a bus that is legal at every instant, so no property can fire and the symptom — a target that never responds — points at the target. Checking relationships between configuration values at elaboration catches it before any stimulus runs, and phrasing the message as the symptom rather than the rule is what makes it findable by someone who has not read the configuration code.

Fix
// Check the RELATIONSHIP at elaboration, where the topology is final and nothing
// has run yet:
//
//    function void check_valid(string ctx);
//       if (vif == null)
//          uvm_fatal_msg({ctx, ": virtual interface is null"});
//       if (half_period < 4)
//          uvm_fatal_msg({ctx, ": half_period below 4 leaves no room to move SDA while SCL is low"});
//       if (stretch_timeout <= half_period)
//          uvm_fatal_msg({ctx, ": stretch_timeout must exceed one half period or every transfer times out"});
//    endfunction
//
// Each message names the SYMPTOM the misconfiguration produces, because that is
// what somebody will be searching for.
//
// WHY A PROPERTY CANNOT DO THIS: the bus is entirely legal throughout. The fault
// is a relationship between two configuration values, and it exists before any
// waveform. Assertions watch waveforms; this needs a check where the numbers are.
//
// AND THE PAIR THAT MATTERS: the driver's patience must also exceed the TARGET's
// abort threshold, or the target gives up mid-transfer and the driver reports it
// as a target fault.

6. What 22.8 Settled

The conditions that break real controllers are legal, so assertions cannot catch them and the checks are policies, budgets and configuration relationships instead.

A stuck line and a legal stretch are the same observation. Only duration separates them, duration is not a protocol fact, and the environment's patience is therefore a chosen parameter needing a test on each side.

A wedged bus breaks the reporting machinery, so bounded waits and a per-test bus-idle check are what keep a failure attributable to the test that caused it.

Excessive stretch is a budget with two ends — the driver's patience and the target's abort threshold — and they must be consistent or each blames the other.

Reset mid-transaction is a stated gap. The design has a reset path; no bench exercises it during a transfer, and that is recorded rather than implied.

A near-miss address test is worth less the further away it is. One bit, or it demonstrates nothing about which bits participate.

Next, the last problem: making these behaviours happen every time rather than occasionally. Chapter 22.9 — Verifying Stretching and Arbitration Without Flaky Tests.

Continue learning