Skip to content
VLSI Mentor

I²C · Module 24

Electrical, Timing and Bring-Up Readiness Review

The audit before power is applied for the first time: sorting every claim into established, calculable, deferred to the board, or unsettleable here. Computes the observation-latency budget against tVD;ACK and finds where it closes — 11.1 MHz for Fast-mode, 22.2 MHz for Fast-mode Plus — and gives a bring-up order that makes each failure diagnostic.

Chapter 24.2 ended by ruling several claims out of simulation's reach — rise time, capacitance, synchronization reliability, timing closure. Those claims do not stop mattering. They arrive all at once, on the day somebody applies power to a board for the first time, and if nothing was done about them beforehand they arrive as a single undifferentiated symptom: it doesn't work.

The readiness review exists to make that day diagnostic rather than mysterious. Its question is not "will it work" — nobody can answer that — but:

For each thing that must be true, what evidence do we already have, and in what order will the missing evidence arrive?

The second half is the part people skip, and it is what turns bring-up from a search into a sequence.

1. The Readiness Table

Every claim an I²C block depends on falls into one of four evidence states. Sorting them is most of the review.

statemeaningwhat the review does
ESTABLISHEDevidence exists and has been seenrecord which run, which report
CALCULABLEnobody has done the arithmetic yetdo it now — it costs minutes
DEFERREDonly the board can settle itplan the measurement and its order
UNSETTLEABLE HEREno available tool addresses itstate the reasoning and the residual risk

The failure mode is not having claims in the last two states. It is having claims in the second state that were treated as the first — a number nobody computed, presented as a property somebody verified. That is a two-minute calculation standing between a design and a week of bring-up.

2. What the Board Decides, and What It Cannot

Three questions get confused during bring-up because they produce the same symptom, and separating them in advance is the single highest-value thing the review does.

Is the line physically able to reach a valid high? This is a DC question and the pull-up's lower bound answers it. Fitted too small, the device pulling low cannot reach VOL while sinking only its guaranteed current, and the "low" is not low enough for some receiver on the bus. 24.5 works this end.

Does it get there in time? An AC question, set by Rp × Cb against the mode's tr(max). This is the one that produces the classic bring-up experience of a bus that works at 100 kHz and fails at 400 kHz — because the rise time did not change, and the time available for it did.

Did the design see it? A digital question, and the one this chapter owns. The line can be electrically perfect and the block can still miss the edge, because between the pin and the protocol decision there are synchronizer stages, filter samples and input registers, each costing clocks.

Those three fail differently and are fixed differently, and a bring-up that has not separated them will try a smaller resistor on a problem that is a filter depth.

3. The Observation-Latency Budget

A target does not respond to an edge on the bus. It responds to an edge it has observed, and observation costs clocks. Chapter 19.9 measured the chain on this curriculum's assembled target:

Azvya Education Pvt. Ltd.VLSI Mentor
19.9's measured latency chain
  total observation latency = SYNC_DEPTH (front end)  2
                            + N_SAMP     (filter)     3
                            + 2          (the slave's own synchroniser)
                            = 7 clocks

and reported it as 5.6 % of a Fast-mode bit period at 50 MHz. For a readiness review the bit period is the wrong denominator. What actually has to be met is tVD;ACK — the time from SCL going low until the acknowledge is valid on SDA — and the observation latency is spent before the target has even seen the edge, so it comes off the front of that budget.

Azvya Education Pvt. Ltd.VLSI Mentor
readiness.py — the budget, and the frequency at which it closes
   LATENCY_CLKS = 7            # 19.9: sync 2 + filter 3 + the slave's own sync 2

   # UM10204 Table 10. tVD;ACK is a MAXIMUM: the acknowledge must be valid by then.
   MODES = [("Standard-mode  100k", 3.45e-6, 10.0e-6),
            ("Fast-mode      400k", 0.90e-6,  2.5e-6),
            ("Fast-mode Plus 1M",   0.45e-6,  1.0e-6)]

   # What the target still has to do AFTER it has observed the falling edge:
   # compare the address against its own, decide, and drive the pad.
   DECIDE_CLKS = 3
Azvya Education Pvt. Ltd.VLSI Mentor
Result — headroom against tVD;ACK
system clock 50 MHz  (20.0 ns per clock)
  mode                     observe   decide     total   tVD;ACK   headroom
  Standard-mode  100k        140ns     60ns     200ns    3450ns       94 %
  Fast-mode      400k        140ns     60ns     200ns     900ns       78 %
  Fast-mode Plus 1M          140ns     60ns     200ns     450ns       56 %

system clock 25 MHz  (40.0 ns per clock)
  mode                     observe   decide     total   tVD;ACK   headroom
  Standard-mode  100k        280ns    120ns     400ns    3450ns       88 %
  Fast-mode      400k        280ns    120ns     400ns     900ns       56 %
  Fast-mode Plus 1M          280ns    120ns     400ns     450ns       11 %  <-- thin

system clock 12 MHz  (83.3 ns per clock)
  mode                     observe   decide     total   tVD;ACK   headroom
  Standard-mode  100k        583ns    250ns     833ns    3450ns       76 %
  Fast-mode      400k        583ns    250ns     833ns     900ns        7 %  <-- thin
  Fast-mode Plus 1M          583ns    250ns     833ns     450ns      -85 %  <-- DOES NOT FIT

Solving for the boundary directly:

Azvya Education Pvt. Ltd.VLSI Mentor
Result — where each mode stops being reachable
  Standard-mode  100k    fits at all above   2.90 MHz; half the budget still free above   5.80 MHz
  Fast-mode      400k    fits at all above  11.11 MHz; half the budget still free above  22.22 MHz
  Fast-mode Plus 1M      fits at all above  22.22 MHz; half the budget still free above  44.44 MHz

Three readable conclusions, and the third is the one worth carrying.

The oversampling ratio is not a free parameter. A design intended for Fast-mode Plus needs better than 22 MHz just to acknowledge in time with this latency chain, and comfortably more than that to have margin. The often-repeated guidance to "oversample by 8×" or "16×" is a rule of thumb whose real content is this inequality; stated as a ratio it survives a change of speed mode unchanged, and this inequality does not.

The latency is a design output, and it moves when the filter moves. Raising N_SAMP — 24.1 §5's specimen — adds directly to the first column. At 25 MHz and Fast-mode Plus, 11 % headroom means roughly three more filter samples before the acknowledge budget is gone.

A margin this thin is not really about protocol legality. It is the number that decides whether a board that works at room temperature works at the corners, and no simulation in this curriculum bears on it. It is CALCULABLE, and the review's job is to make sure somebody calculated it.

4. The Assumptions Underneath That Table

Every number above rests on things that are decisions rather than facts, and a readiness review that does not surface them has produced a spreadsheet rather than an argument.

assumptionvalue usedwhat breaks if it is wrong
observation latency7 clocksmeasured in 19.9 for that assembly; a different composition has a different number
decision time after observation3 clocksan unverified placeholder — an address compare against a configurable address, a register read, or a FIFO status lookup can each be longer
tVD;ACKUM10204 Table 10fixed by specification; safe
the clock is the one in the constraintas configureda derived or gated clock makes every ns figure wrong
the filter runs at the system clockyesa divided sampling clock multiplies the first column

The second row is the one to argue about in a review. Three clocks is what this chapter assumed to produce a number; it is not a property of any design. An address compare that has to consult a strapping input, or a target whose acknowledge depends on whether a receive FIFO has room (16.4), can spend considerably more — and at 25 MHz Fast-mode Plus there were only about 11 % of the budget spare, which is four clocks.

5. The Parameter-Boundary Audit

A parameter is a promise that a range of values works. It is nearly always a promise nobody tested.

The audit is mechanical and takes minutes: for each parameter, list its declared legal range, the values that have actually been elaborated, and the values the product might plausibly use. Where those three sets differ, there is an untested promise.

This curriculum's own target failed exactly that audit. Chapter 19.9 went to integrate the Module 18 slave behind a front end that had already synchronized and filtered the lines, so the sensible setting for the slave's own synchronizer was SYNC_DEPTH = 1 — one register, which is what an already-synchronous signal needs.

Azvya Education Pvt. Ltd.VLSI Mentor
19.9's elaboration result
  i2c_slave_sync  at SYNC_DEPTH=1 -> FAILS        i2c_line_sync at 1 -> ELABORATES
  i2c_slave_sync  at SYNC_DEPTH=2 -> ELABORATES   i2c_line_sync at 2 -> ELABORATES
  i2c_slave_sync  at SYNC_DEPTH=3 -> ELABORATES   i2c_line_sync at 3 -> ELABORATES

The chain shifts with {pin, chain[SYNC_DEPTH-1:1]}, which at depth 1 is the part select chain[0:1] — descending where the declaration ascends, and rejected. A documented, configurable parameter has a legal-looking value that cannot be used, and it was found only because a later module's bench instantiated three depths while the owning module's bench instantiated the default.

Three things generalize from that, and none of them is about part selects.

The value a parameter is tested at is the one its own bench uses, which is the default. A bench that sweeps is the only bench that audits.

Boundary values are where expressions degenerate. Depth 1 makes a shift chain a single register; N_SAMP = 1 makes an agreement counter unconditional; a FIFO of depth 1 has no distinct full and empty. Degenerate cases are where an expression written for the general case stops being well-formed — and they are the values an integrator reaches for, because they are the cheapest.

Elaboration failure is the good outcome. It is loud, immediate, and points at a line. The same audit's dangerous finding is a boundary value that elaborates and is quietly wrong — which is what 24.1 §5's N_SAMP = 32 is.

6. Constraints and Clock Domains

Chapter 19.6 settled what this design's constraint story is: there is one clock, scl appears in no sensitivity list anywhere, and the bus inputs are asynchronous levels that cannot be given an arrival time. The review does not redo that analysis. It checks three things that go wrong afterwards.

Is the asynchronous input actually declared asynchronous, and nothing more? The correct constraint says these inputs have no timing relationship to the clock. The tempting one says the paths from them are false paths — which is a stronger statement that also removes the internal path from the first synchronizer stage to the second, and that path is real and must be met. A blanket false path on a synchronizer is how a design with a correct architecture ships without the timing that makes it work.

Is the synchronizer's intermediate node used anywhere else? A two-stage chain whose first stage also feeds a comparator has one stage of synchronization and one path carrying a possibly-unsettled value into logic. This is a structural question, answerable by one search, and it is not visible in simulation at all.

Does the timing report cover the clock the design actually runs on? A report on a clock the design does not use is a green report that says nothing. This sounds too obvious to check and is one of the commonest bring-up surprises after a clock-source change.

And one claim that must stay in the fourth state of §1's table. Synchronizer reliability cannot be established here. Chapter 19.4 makes that case: RTL simulation has no model of settling time, so no simulation produces evidence about metastability. What a review can do is record the depth, the reasoning behind it, and the fact that the residual failure rate is a calculation requiring library data nobody in this project has. That is an honest entry. "Verified by simulation" is not.

7. A Bring-Up Order That Makes Failures Diagnostic

The order matters because each step, if it passes, removes a class of cause from the next step's search space. Reversing two steps turns a five-minute answer into an afternoon.

A flowchart of eight bring-up steps in sequence. Power and idle levels, then the rise time on a scope, then the clock frequency, then a single address transfer with a logic analyzer, then acknowledge, then a register read, then the full speed mode, then clock stretching and multi-device traffic. Each step has a branch to a named cause set when it fails.noyesnoyesnoyesnoyesnoyesnoyesPower, design inresetBoth linesidle high?Pull-ups, stuckdriver, resetpolarityRise timewithintr(max)?Rp too large, CbunderestimatedSystem clockat itsconstrainedrate?Clock source,divider, PLL lockOne addressbyte appearson the wire?Pin mapping, OEpolarity, SCLgenerationTargetacknowledges?Address,observationlatency, ACK windowRegister readreturns thedatasheetvalue?Pointer semantics,read path,endiannessRaise to the targetspeed modeStretching andmulti-devicetraffic
Figure 1 — bring-up as a sequence of eliminations. Each step is chosen so that passing it removes an entire class of cause from everything below, and each step's failure has a small, named set of causes. The first three steps happen with the design held in reset and no traffic on the bus, which is why they are first: they are the only steps whose result is unambiguous.

Two properties of that order are worth naming, because they are what make it work rather than merely look tidy.

The first three steps happen with no traffic. Idle levels, edge shape and clock rate are all measurable on a quiet bus with the design held in reset, and each has an unambiguous pass. Any step that requires traffic has a larger cause set by construction, because traffic involves both ends.

Speed comes last, deliberately. Running at 100 kHz first is not timidity; it is removing the rise-time and observation-latency terms from the search space so that a failure at that point is a protocol or address problem. Then raising the speed re-introduces exactly those two terms, so a failure that appears only after the speed change has a cause set of two items rather than twelve. Starting at the target speed conflates them and is the commonest way a bring-up loses a day.

8. The Readiness Checklist

Electrical

  • Is Rp inside the window at the estimated Cb, and how much can Cb grow before it leaves? (24.5 measures the headroom)
  • Where does the Cb estimate come from, and does it include the connector, the cable and anything added after the schematic was frozen?
  • Is the pull-up supply the same rail as every device's supply, and does anything power up in a different order?
  • Does any device have an internal pull-up enabled as well? (19.3)

Timing and observation

  • Observation latency in clocks, and in ns at the actual clock rate.
  • Headroom against tVD;ACK at the fastest supported mode — and what decision time that assumed.
  • Does the timing report cover the clock the design runs on, at the frequency it runs at?

Clock domains

  • Synchronizer depth on each bus input, and no intermediate tap.
  • Asynchronous inputs declared asynchronous — not false-pathed through the synchronizer's internal stage.
  • scl absent from every sensitivity list, or a written justification for why not.

Parameters

  • For each: declared range, elaborated values, values the product will use.
  • Any boundary value that fails to elaborate, and any that elaborates and degenerates.

Reset and recovery

  • Does reset leave both lines released?
  • If the design is reset mid-transfer, what does the bus look like afterwards, and can the controller recover it? (15.4)

Observability

  • Can the pins be probed on the assembled board, or are they under a connector?
  • Is there an on-chip trace of the resolved lines, not just the drive intents? (19.8)
  • Is there a counter for each error the block can report, readable without a debugger attached?

That last group is the one that gets cut for area and is regretted within a week. An error the block detects but cannot report is indistinguishable, from outside, from an error it never detected.

9. Common Misconceptions

"Simulation passed, so the board should work." Simulation has no rise time, no threshold, no supply and no temperature. 19.7 catalogues the divergences; the readiness review's job is to make the list of things simulation did not cover explicit before it is discovered one item at a time.

"A two-stage synchronizer solves the asynchronous input problem." It bounds the probability of an unsettled value propagating. It does not make the input synchronous, it does not remove the need for the internal path to meet timing, and it says nothing at all about whether the observation arrived in time — which is §3's budget and a completely separate question.

"It works at 100 kHz, so the hardware is fine and this is a software problem." Almost the opposite. Working at 100 kHz and failing at 400 kHz is the signature of a term that did not change while the budget shrank — rise time or observation latency. It is among the most diagnostic results bring-up produces, and reading it as a software problem discards that.

"A false path on the asynchronous inputs is the correct CDC constraint." It is stronger than intended and removes a real path. The constraint should say the inputs have no timing relationship to the clock; the flop-to-flop path inside the synchronizer must still be met, and it is the path the whole arrangement depends on.

"The parameter is documented as configurable, so it is configurable." Documented is a claim; elaborated is evidence. This curriculum's own target has a documented parameter whose most useful value does not elaborate.

Two boards where the readiness question had not been asked

1The bus that worked at 100 kHz and failed at 400 kHz
Buggy Code
// Board bring-up. At 100 kHz every device answers, every register reads
// correctly, a 4-hour soak is clean. The product ships at 400 kHz.
//
// At 400 kHz: the first device on the bus works. The one at the far end of
// the board NACKs intermittently -- perhaps one transfer in fifty.
//
// What was checked before power-on:
//    RTL simulation        clean, including the 400 kHz configuration
//    pull-up value         4.7k, "the standard value"
//    Cb                    not estimated
//    rise time             not measured
Symptom

A scope on SDA at the far device shows a rise that is still climbing when SCL goes high. The near device's input has already crossed its threshold; the far one has not.

Root Cause

The rise time did not change between 100 kHz and 400 kHz. The time available for it did: tHIGH(min) falls from 4.0 us to 0.6 us, and the whole rise has to complete before the line is sampled.

4.7k into an unestimated Cb is the defect. Chapter 24.5's audit on a board of this shape reports the window as [967, 3978] ohm at 89 pF -- 4.7k is ABOVE Rp(max) even at that optimistic capacitance, so it was never legal for Fast-mode. It passed at 100 kHz because Standard-mode allows 1000 ns of rise and there was time to spare.

This is why the bring-up order in Section 7 puts rise time at step 2, before any traffic: it is measurable on a quiet bus, it takes a minute, and it removes an entire class of cause from every step after it.

Fix
Compute the window before choosing the resistor, at the ESTIMATED Cb and at a
pessimistic one:

  Rp(min) = (Vdd - VOL) / IOL          = (3.3 - 0.4) / 3 mA  =  967 ohm
  Rp(max) = tr(max) / (0.8473 * Cb)

then fit toward the LOW end of the window rather than the middle, because
Rp(min) does not move with Cb and Rp(max) does -- so all the risk from a bad
capacitance estimate is on one side. Measure the edge at step 2 of bring-up and
compare it against the calculation; a measured rise time that disagrees with
the estimate is itself the finding, because it means Cb is not what was assumed
and every other number derived from Cb is also wrong.
2The synchronizer that was constrained out of existence
Buggy Code
// A target on an FPGA. The architecture is correct: SDA and SCL each go
// through a two-stage synchroniser before any logic uses them.
//
// The constraint file, written to silence 200 unconstrained-path warnings:
//
//    set_false_path -from [get_ports {sda scl}]
//    set_false_path -through [get_pins u_sync/chain_reg[*]/Q]
//
// The first line is defensible. The second removes the path from the first
// synchroniser stage to the second -- which is the ONE path the whole
// arrangement depends on meeting.
Symptom

Timing closes with large margin. The design works on three boards and fails on the fourth, in one temperature corner, with occasional byte corruption that no simulation reproduces. Rebuilding with a different seed changes which corruption appears.

Root Cause

A synchroniser works by giving the first stage's output a full clock period to settle before the second stage samples it. Excluding that path tells the tool it need not deliver the period -- so placement may put the two flops far apart, and the settling time available becomes whatever the routing happened to give. The architecture is still correct and the implementation no longer honours it.

Note what this failure is NOT visible to. RTL simulation has no notion of settling time, so every simulation passes. It is not visible to functional coverage. It is not visible to a code review of the RTL, because the RTL is right. The only artifacts that carry the evidence are the constraint file and the timing report, and the review that reads them is the one this chapter is about.

Fix
Say the true thing, which is narrower:

  set_false_path -from [get_ports {sda scl}] -to [get_pins u_sync/chain_reg[0]/D]

-- the ASYNCHRONOUS arrival has no relationship to the clock, and that is all.
Everything downstream of the first flop is ordinary synchronous logic and must
meet timing. Then add the placement constraint that makes the intent explicit,
so a later tool version cannot separate them:

  set_max_delay -from [get_pins u_sync/chain_reg[0]/C] \
                -to   [get_pins u_sync/chain_reg[1]/D] <one period>

And review the structural question alongside it: does anything ELSE read
chain_reg[0]? A tap on the intermediate stage has the same effect as the bad
constraint, is invisible in simulation for the same reason, and is found by one
search of the source.

10. Reason It Through

A. A design is moving from a 50 MHz FPGA to a 20 MHz one to save power, and the requirement is unchanged: Fast-mode, 400 kHz. Using §3's chain, what must be re-examined, and what is the first thing you would compute?

The observation-latency budget, because it is the only term in the whole design that scales directly with the clock period, and the requirement did not move with it. At 20 MHz the 7-clock chain is 350 ns and the 3-clock decision another 150 ns, against a tVD;ACK of 900 ns — so roughly 44 % headroom, which fits but is materially tighter than the 78 % at 50 MHz. The boundary table says Fast-mode stops fitting at all below 11.1 MHz, so 20 MHz is not marginal in the sense of being nearly illegal; it is marginal in the sense that the assumed 3-clock decision time now costs 150 ns, and if the real decision path is six clocks the headroom halves again. So the first computation is the budget, and the immediate follow-up is to replace the assumed decision time with the measured one — because at 50 MHz that assumption was cheap and at 20 MHz it is the dominant uncertainty.

B. A reviewer finds that N_SAMP was raised from 3 to 8 during bring-up, and the commit message says "fixes intermittent NACKs on board 4". What is your response?

That the change may be correct and the evidence is not, and those are separable. Raising the filter depth suppresses narrow disturbances, so it will make a noise symptom go away — and it will equally make a symptom go away whose cause is a marginal rise time, a ground problem, or a reflection, because a slow, ringing edge presents as narrow disturbances at the input threshold. The commit has removed the symptom that would have led to the cause. Concretely, ask for three things: the measured edge on board 4 compared with a working board, the new observation latency and its headroom against §3's budget, and whether the value still passes the upper bound for the fastest supported mode. If the edge is clean and the noise is real, the change is right and needs the latency recorded; if the edge is not clean, the filter is hiding an electrical fault that will come back at a corner the filter cannot reach.

C. Which claims in §8's checklist can a simulation close, which need a calculation, and which need the board? Give one of each and say why.

Simulation closes "does reset leave both lines released" — it is a functional property of the RTL, observable directly in any run that resets mid-transfer, and it is the kind of claim simulation is actually good for. Calculation closes "headroom against tVD;ACK": every input is a published limit or a counted clock, nothing needs to be run, and it takes minutes. Only the board closes "is Cb what we estimated" — capacitance is a property of copper, connectors and whatever was added after the schematic was frozen, and the first honest measurement of it is the rise time at step 2 of bring-up. The reason the sorting matters is that each category has a characteristic way of going wrong: simulation claims get closed by a test that does not check the consequence, calculation claims get closed by nobody doing them, and board claims get closed by assuming a number.

D. The block reports six distinct error codes internally, but only a single error flag is brought out to the register interface. The area saving is negligible. What is the argument for the six, phrased as a readiness question?

That a bring-up failure's cost is dominated by the size of the cause set, and the six codes are exactly a partition of that set which the design has already computed and is discarding at the boundary. With one flag, an arbitration loss, a stretch timeout, an address NACK and a data NACK are indistinguishable from outside — so the first hour of every bring-up failure is spent re-deriving, with a scope, information the block already knew. The readiness question is "what can be observed without a debugger attached", and the answer here is one bit where six were available for free. There is a second argument that matters more in the field than at bring-up: a counter per error class turns an intermittent fault into a rate, and a rate is something you can correlate with temperature, with traffic, or with which board it is. A single flag gives you an anecdote.

11. Understanding Check

12. What 24.3 Settled

Sort every claim by the evidence that could settle it, before asking whether it holds. Established, calculable, deferred to the board, or unsettleable here. The dangerous category is the second, because a number nobody computed is easily mistaken for a property somebody verified.

Observation latency is a first-class budget, and it has a boundary. Seven clocks against tVD;ACK fits comfortably at 50 MHz and stops fitting at all below 11 MHz for Fast-mode and 22 MHz for Fast-mode Plus. The familiar "oversample by 8×" guidance is a rule of thumb whose real content is that inequality — and unlike the ratio, the inequality moves when the speed mode does.

Three questions produce the same bring-up symptom and have different fixes. Can the line reach a valid level; does it get there in time; did the design see it. A bring-up that has not separated them will try a different resistor on a filter-depth problem.

The bring-up order is an elimination sequence, not a ritual. Quiet-bus measurements first because their results are unambiguous, speed last because it re-introduces exactly the two terms everything else was arranged to exclude.

Some claims stay open, and saying so is the review's output. Metastability is not settleable with anything in this project. Recording the depth, the reasoning and the missing data is a complete entry. Calling it verified is not.

The next chapter turns from the evidence to the engineer. Several of the failures in this chapter began with a mental model that was plausible, widely held, and wrong — and the fastest way to recognise one is by the shape of the failure it produces. Chapter 24.4 — I²C Misconceptions Engineers Still Get Wrong.

Continue learning