Skip to content
VLSI Mentor

I²C · Module 24

I²C Misconceptions Engineers Still Get Wrong

Twelve wrong mental models that survive because they are right almost everywhere — each with the region that confirms it, the failure signature it uniquely produces, and the correct model. Includes a symptom-to-model table for reading a failure backwards, and why the most dangerous signatures are absences.

A misconception is not ignorance. Nobody holds a mental model that fails immediately — it would not survive the first day. The models in this chapter survive for years because they produce correct answers almost everywhere, and are wrong in one specific place that is hard to reach.

That is why "here is the right answer" rarely fixes one. What fixes one is seeing the failure it uniquely produces, because that failure is the only evidence that distinguishes the wrong model from the right one. So each entry below has five parts:

  1. the wrong model, stated as somebody would actually think it;
  2. why it sounds plausible — the region where it gives correct answers, which is most of them;
  3. the failure signature — the symptom it produces and no other model does;
  4. the correct model;
  5. where the evidence lives.

The fifth part matters because a corrected model with no evidence attached is just a different thing to believe.

1. "A pull-up resistor is how the bus makes a logic HIGH"

Why it sounds plausible. It is functionally true: release the line and it goes high, and every logic analyzer shows a clean 1. For a bus with one talker, "the pull-up drives high" and "a driver drives high" are interchangeable descriptions of identical waveforms.

Failure signature. A design works with one device and fails intermittently the moment a second is fitted — with the analyzer showing a level that is neither 0 nor 1, often around half the supply, during acknowledge slots. Chapter 24.1 §3 produces the RTL form of this: the submitted output stage passes every single-participant test and puts the bus into x in the one case where the stage has released and somebody else is pulling low.

Correct model. The pull-up is not a driver; it is the default in the absence of any driver. The bus has no source of 1 at all — only devices that pull to 0 and devices that let go. That is what makes the wired-AND, and the wired-AND is what makes acknowledge, clock stretching and arbitration possible without contention. The resistor is not there to produce a high. It is there so that "nobody is pulling" has a defined voltage.

Evidence. 2.3, 2.5, and 2.2 for what happens when a driver does source a high.

2. "SDA must never change while SCL is HIGH"

This one is stated in nearly every introduction to I²C, and taken literally it makes the protocol impossible.

Why it sounds plausible. It is the data-valid rule, and it is true of every one of the eight data bits and the acknowledge bit. For the overwhelming majority of the transitions on a real bus, it is correct.

Failure signature. An assertion written from it fires on every START and every STOP, so it gets switched off within a day — and switching it off removes the check on the rule's actual domain as well. Downstream, an environment reports no data-valid violations because nothing is checking.

Two signals, SCL and SDA, over ten cells. SDA falls while SCL is high at cell one, marked START. SDA changes while SCL is low at cell three, marked as a legal data change. SDA falls again while SCL is high at cell five, mid-byte, marked as having the same shape as a START. SDA rises while SCL is high at cell eight, marked STOP.START — SDA falls, SCL highSTART — SDA falls, SCL highdata change, SCL low: legaldata change, SCL low: legalmid-byte: same shape as STARTmid-byte: same shape asSTARTSTOPSTOPSCLSDAt0t1t2t3t4t5t6t7t8t9
The rule and its two deliberate exceptions, on one picture.
Figure 1 — four SDA transitions. A change between two adjacent cells that are both SCL-high is a change during the high phase. The first and last are legal, and are how framing is signalled at all. The third is illegal — and has exactly the same shape as the first, so a receiver applying the specification reads it as a repeated START. That is why one corrupted bit can desynchronize a transfer rather than merely corrupt a byte.

Correct model. SDA changing while SCL is high is the signalling mechanism for framing. The rule is: inside a byte, SDA is stable while SCL is high. The exception is not an annoyance the specification tolerates — it is the entire reason START and STOP are distinguishable from data on a wire that carries both.

The consequence worth internalizing is the one in the figure's caption. Because a mid-byte violation is shaped identically to framing, its effect is not a wrong bit. It is a receiver that believes a new transfer has begun, discards the partial byte, and starts interpreting the following bits as an address. One electrical glitch becomes a structural desynchronization.

Evidence. 4.2, 5.2, 5.5.

3. "Clock stretching is a target's problem, and a slow target just makes things slower"

Why it sounds plausible. From the controller's point of view a stretch looks like latency, and latency is an inconvenience rather than a hazard. A controller that handles stretching correctly does indeed just go slower.

Failure signature. A bus that works with every device except one, where the failure is a corrupted byte rather than a timeout — because a controller that does not honour stretching does not hang; it clocks on regardless, and the target, still busy, misses bits. The characteristic capture shows nine clocks with the target's data arriving one bit late throughout.

Correct model. Stretching is an obligation on the controller, not a capability of the target. The target is not requesting anything; it is holding SCL low, and SCL is a wired-AND like SDA, so the clock genuinely does not rise. A controller that "does not support stretching" is a controller that drives SCL push-pull, or that counts time instead of watching the line — and both are defects rather than a simpler feature set.

The second half, which 24.6 develops: the specification places no upper bound on how long a target may stretch. Every timeout a controller implements is therefore a system policy that declares some legal target non-compliant.

Evidence. 12.1, 12.2, 12.4.

4. "More filtering is safer"

Why it sounds plausible. Filtering exists to reject disturbances, and rejecting more disturbances sounds strictly better. The axis appears to have only one direction.

Failure signature. Protocol that disappears with no error reported anywhere. A device stops responding at the higher speed mode while the bus looks electrically perfect on a scope, because a legal clock high phase is being deleted before any logic sees it.

Correct model. A filter has a stop-band and a pass-band, and the specification fixes both ends: tSP = 50 ns must be rejected, and the narrowest legal high phase of the chosen speed mode must survive. 24.1 §5 measures the resulting window directly — at a 100 MHz sampling clock and Fast-mode Plus, N_SAMP must be between 6 and 26. Depth is not margin; it is a position between two bounds, and it costs observation latency on every edge whether or not there is anything to filter.

Evidence. 19.5, 11.8, and 24.3 §3 for what the latency is spent against.

5. "Two synchronizer stages solve the asynchronous-input problem"

Why it sounds plausible. They solve a problem, completely and correctly — the propagation of an unsettled value into logic — and that is the problem the phrase "asynchronous input" usually refers to.

Failure signature. Two distinct ones, which is the tell that one phrase is covering two problems. Either a design that is architecturally correct and fails at one temperature corner on one board (the synchronizer is there and the timing that makes it work was constrained away — 24.3's second DebugLab case), or a design whose observations are correct but late, missing timing windows at speed.

Correct model. Three separate jobs get bundled under "synchronization", and depth addresses only the first:

jobaddressed bywhat depth does
stop an unsettled value reaching logicchain depththis is what depth is for
reject disturbances too narrow to be protocola filternothing
see the edge in time to act on itthe latency budgetmakes it worse

The third row is the one that surprises people, because depth and latency move in opposite directions and both are called "being careful about the inputs".

Evidence. 19.4 §3 separates the three jobs; 24.3 §3 budgets the third.

6. "Simulation passed, so the hardware should work"

Why it sounds plausible. It is nearly true for synchronous logic with a clean interface, which is most of what most engineers verify. For an internal pipeline it is close enough to true that the exceptions are rare.

Failure signature. The signature is not a particular symptom — it is the pattern of a design that is functionally flawless and fails in ways that correlate with temperature, board, cable length or speed mode rather than with data.

Correct model. A simulated bus has no rise time, no threshold, no supply, no temperature and no coupling. Everything that distinguishes a working board from a failing one at the same RTL lives in exactly those five. 19.7 catalogues the divergences, and 24.3's first DebugLab case is the canonical instance: 4.7 kΩ works at 100 kHz and fails at 400 kHz because the rise time did not change while the time available for it shrank by a factor of six.

Evidence. 19.7, and 24.2 §4's evidence hierarchy for what each claim needs.

7. "100 % coverage means verification is done"

Why it sounds plausible. Coverage is a real measurement of a real thing, it is hard to get to 100 %, and getting there genuinely means the stimulus reached every condition somebody wrote down.

Failure signature. A defect that escapes to a customer while the sign-off pack is still green on a rerun — because nothing in the pack was measuring the thing that broke.

Correct model. A behavior is verified when it is stimulated, observed and checked. Coverage measures the first only. 24.2 §3 runs a design with an active defect at 100 % of bins hit and zero mismatches; adding a single transfer catches it and leaves the coverage report byte-for-byte identical. The gap is not a matter of degree — the two quantities are independent.

The sharper version, for a bin that is genuinely well chosen: a covered behavior usually has several consequences, and the bin is satisfied by observing any one of them. "A write to a read-only register was refused" has two consequences — the register does not change, and the pointer does not move — and the bin is green having observed either.

Evidence. 24.2 §3, 20.1.

8. "If a mutant survives, the RTL must be wrong"

Why it sounds plausible. Mutation testing is presented as a measure of test quality, so a survivor reads as a failure of the tests, and the instinctive fix is a test.

Failure signature. A suite that grows a layer of tests checking implementation details — each one written to kill a specific mutant, each one breaking on the next legitimate refactor, and none of them checking a property anybody stated.

Correct model. A survivor says the environment did not detect that change, and nothing yet about whether it should have. Of the four classifications, three end with no test written: equivalent (the finding is a redundancy in the design), outside the published contract (the finding is in the plan), and invalid (the finding is in the harness). Only the fourth is a verification hole — and even then the test to write is the one that checks the violated property, not the one that kills the mutant.

Evidence. 24.2 §6, 20.1.

9. "The monitor can use the design's internal state — it is much simpler"

Why it sounds plausible. It is much simpler, it works, and the signal is right there. Reconstructing an address match from bus edges is genuinely harder than reading addr_match.

Failure signature. An environment that never reports a mismatch in the affected area, for years, and a bug that escapes in exactly that area. 24.2 §2 gives the general measurement: a scoreboard whose expected value comes from inside the loop performs its comparisons and reports zero mismatches with the device removed entirely.

Correct model. Every claim that depends on a signal taken from inside the design is an assumption about that signal, not a check on it. A monitor reading addr_match cannot detect a wrong address match, and every downstream check inherits that blindness silently.

There is a defensible use: a bring-up or triage configuration, whose purpose is localizing a known failure rather than establishing correctness. The condition is that it cannot be the configuration that signs the block off, and that the test plan records which objectives depend on it.

Evidence. 24.2 §2, 20.7, 21.5.

10. "Parameterization makes RTL reusable"

Why it sounds plausible. It makes RTL configurable, which is a prerequisite for reuse and feels like the whole of it.

Failure signature. An integrator reaches for a boundary value — the smallest, cheapest, most obvious one — and it either fails to elaborate or elaborates into something quietly wrong. This curriculum's own target has a documented SYNC_DEPTH whose most useful value for an integrator, 1, does not elaborate at all (19.9).

Correct model. A parameter is a promise that a range of values works, and the promise is nearly always untested, because a module's own bench instantiates the default. Every added parameter multiplies the configuration space and divides the fraction of it anybody has run. Boundary values are where the risk concentrates, because that is where general-case expressions degenerate: a shift chain of length 1 is not a shift, an agreement counter with a threshold of 1 is unconditional, a FIFO of depth 1 has no distinct full and empty.

The reviewable question is not "is it parameterized" but which values have been elaborated, as opposed to declared legal.

Evidence. 19.9, 17.13, 24.3 §5.

11. "An address NACK means the device is not there"

Why it sounds plausible. It is the commonest cause by a wide margin, so the inference is usually right — which is exactly what makes it expensive when it is not.

Failure signature. Hours spent on wiring, addresses and power rails for a device that is present, powered, correctly addressed, and busy. Or the opposite: a device that answers a bus scan and then fails every real transfer.

Correct model. A NACK on the address byte means nobody pulled SDA low during the ninth clock. The causes that produce it are a set, not a single fact: wrong address, address strapping not what the schematic says, device still in its power-on sequence, device busy with an internal write cycle and deliberately not acknowledging (16.4), general-call or reserved-address handling, the target never seeing the START because a filter or synchronizer swallowed it, or the acknowledge arriving outside the window because observation latency ate the budget (24.3 §3).

The useful discipline is to notice that "device absent" and "device busy" are distinguishable: a busy device answers if you retry, which takes one line of script and eliminates half the set.

Evidence. 7.2, 16.4, 6.4.

12. "Repeated START is just a STOP followed by a START, saving a little time"

Why it sounds plausible. The bus activity is nearly the same and the data transferred is identical, so on a capture the two look like a minor optimization of each other.

Failure signature. A driver that works perfectly on a single-controller bus and corrupts data occasionally on a multi-controller one — with the corruption being a whole transfer's worth of wrong data, from the wrong register, rather than a bad byte.

Correct model. The difference is bus ownership, and it is categorical rather than quantitative. A STOP releases the bus; any other controller may then win arbitration and start a transfer of its own. A repeated START does not release it. For the register-pointer pattern — write the pointer, then read — that is the whole point: the two phases must be one indivisible occupancy, because between them the target holds state that another controller's transfer would overwrite.

Saved time is a side effect. Atomicity is the feature.

Evidence. 10.2, 10.3, 5.4.

13. Reading a Failure Backwards

The value of the list above is not in memorizing it. It is that several common symptoms map to a small set of candidate wrong models, so a symptom can be used to narrow which assumption to question first.

symptomwrong models that produce itthe cheapest discriminating check
works with one device, fails with two§1 pull-up as driversearch the source for any assignment of 1 to SDA or SCL
works at 100 kHz, fails at 400 kHz§6 simulation implies hardware; §4 deeper filtermeasure the rise time; compute the observation-latency headroom
one device on the bus never works, others fine§3 stretching as latency; §11 NACK means absentretry the address; check whether SCL is held low after the eighth clock
receiver desynchronizes mid-transfer§2 the data-valid rule without its exceptionlook for an SDA edge during an SCL high phase, mid-byte
regression green, customer finds it in a week§7 coverage as completeness; §9 monitor reads internalsremove the device from the bus; does the environment still pass?
works in the lab, fails on one board at one corner§5 two stages is enoughread the constraint file for a false path through the synchronizer
fails only when another controller is active§12 repeated START as an optimizationcheck whether the pointer write and the read are one occupancy
integrator reports it will not build§10 parameterized means reusableelaborate every boundary value of every parameter

The final row of that table is the one that gets a happy ending, and it is worth noticing why: an elaboration failure is loud, immediate and points at a line. Most of the others are quiet.

14. Reason It Through

A. A colleague says "we don't need to worry about clock stretching — our targets are fast." What is wrong with the sentence as an engineering statement, and what would you need to know before agreeing to anything?

Two things are wrong and they are different. The first is that "our targets" describes today's bill of materials, not the controller, and the controller is the thing being designed — a controller that cannot handle stretching is defective independently of what is currently on the bus, and it will be reused. The second is subtler: "fast" is a claim about typical behavior, and stretching is not a performance characteristic but something a device does at specific moments — after a write that starts an internal cycle, on a read that needs a conversion, when a FIFO is momentarily full. A device that never stretches in the lab may stretch when its buffer is contended. Before agreeing to anything, the question is not how fast the targets are but whether any datasheet on the bus mentions stretching at all, and whether the controller watches SCL or counts time — because only the second of those is a property of the design under discussion.

B. An engineer proposes removing the glitch filter entirely, arguing that the board is clean, the signals are short, and the filter only adds latency. Where is that argument strong, and where does it fail?

It is strong on the latency point, which is real, quantified, and paid on every edge: 24.3 §3 shows the observation chain consuming a measurable fraction of the acknowledge budget, and at a low clock frequency it is the dominant term. It is also strong that filter depth is often chosen by habit rather than by calculation. It fails on the premise. Spike suppression is not a board-quality mitigation the designer may decline — tSP is a specification requirement on a compliant input, so a device without it is out of spec regardless of how clean this particular board is. And "this board is clean" is a claim about one board at one moment, which is precisely the kind of claim that stops being true after a mezzanine, a longer harness, or a switching regulator moves. The defensible version of the proposal is not removal but sizing: compute the window from tSP and the narrowest legal pulse, and take the smallest value in it rather than the largest.

C. Why is "SDA must never change while SCL is high" a more dangerous thing to half-believe than "a pull-up drives the bus high"?

Because the pull-up misconception fails loudly the first time a second device is fitted — a contended level, an intermittent bus, a symptom that demands investigation. The data-valid misconception fails by making a correct check impossible to write, and its symptom is an absence: an assertion that fires on every START, gets disabled, and takes the real check with it. Nothing then reports anything, and the resulting environment looks identical to one that is correctly checking the rule. A model that produces a wrong measurement gets caught; a model that produces no measurement does not.

D. Several entries above share a structure: the wrong model gives correct answers in a region, and the region is exactly where testing happens. Which entries, and what does that suggest about how to design a test suite?

At least §1, §7 and §9 have that shape precisely. The pull-up misconception is correct for every single-participant test, and single-participant benches are what get built first because they run first. The coverage misconception is correct whenever stimulus and checking happen to coincide, which is the common case. The internal-signal misconception is correct whenever the internal signal is correct, which is most of the time and always in the configuration where it was debugged. The suggestion is uncomfortable: the region where a suite is easiest to build is systematically the region where these models agree with reality, so a suite that grows by the path of least resistance grows entirely inside the blind spot. The counter-measure is not more tests but deliberately chosen ones — put a second participant on the modelled bus, remove the device and require failure, compute a boundary value and elaborate it. Each is one test, and each covers a region that no amount of growth in the easy direction would ever reach.

15. Understanding Check

16. What 24.4 Settled

A misconception survives because it is right almost everywhere. None of the twelve is a naive error; each gives correct answers across the region where most work happens, which is why being told the right answer rarely dislodges one.

The failure signature is the useful handle. A wrong model is identified by the symptom it uniquely produces — works with one device and not two, works at 100 kHz and not 400 kHz, green regression and an escaped bug. §13 runs that mapping backwards, from symptom to the assumption worth questioning first.

The most dangerous signatures are absences. A disabled assertion, a scoreboard that cannot fail, a coverage bin standing in for a check: each produces silence, and silence is indistinguishable from correctness. The only remedy is to demonstrate that the check is capable of firing before its silence is allowed to mean anything.

Several of these models are correct exactly where testing is cheapest. A suite that grows in the easy direction grows inside the blind spot, so the corrective tests have to be chosen rather than accumulated — a second participant on the bus, a device removed, a boundary value elaborated.

The next three chapters take three of the mechanisms behind these models and work them end to end, starting with the one the whole protocol rests on. Chapter 24.5 — Reasoning About Open-Drain, Pull-Ups and Rise Time.

Continue learning