I²C · Module 24
Reasoning About Clock Stretching
From a device’s reason to stretch, through the wired-AND on SCL, to the obligation it creates on somebody else’s controller. Measures the result that the timeout number is not the policy: two architectures given the same value differ by 10× in the legal stretch they tolerate, and one lets a four-byte transfer hold the bus for 17 784 clocks.
Clock stretching is where an I²C conversation stops being about wires and becomes about obligations. A device does something entirely local — it is not ready — and the consequence lands on a different device, designed by a different team, possibly years earlier.
Three questions get merged into one and they have different answers:
- Why does a device stretch? A statement about that device's internals.
- How does stretching work? A statement about the electrical layer.
- Who must do what? A statement about the controller, which is the only design the first two constrain.
Most confusion about stretching is an answer to one of these offered as an answer to another.
1. Why a Device Stretches
The reasons are varied and the shape is always the same: the device owes a response at a moment fixed by somebody else's clock, and it cannot produce it yet.
| situation | what it owes | why it cannot yet |
|---|---|---|
| a byte just arrived | an acknowledge | the receive buffer is full, or the write has a side effect that takes time |
| a read was requested | the next data byte | the value needs a conversion, a flash access, or a lock on shared state |
| an address just matched | an acknowledge | the device was in a low-power state and its logic is still waking |
| a register write landed | an acknowledge for the next byte | an internal erase or program cycle has begun (16.3) |
| a slow controller | — | a microcontroller bit-banging the bus and losing the CPU to an interrupt |
The last row is the one people forget, and it is a useful corrective: stretching is not only a target behavior. Any participant holding SCL low stretches, including a controller that has been interrupted mid-transfer.
Notice what the table does not contain: a request. The device is not asking for time. It is taking it, by a mechanism that cannot be refused.
2. The Mechanism Is the Same One as Acknowledge
The whole of stretching follows from a fact established in 2.5 and then applied to the other wire: SCL is a wired-AND too.
The controller does not drive SCL high. It releases SCL and the pull-up raises it — exactly as with SDA. So if any device is still pulling SCL low, the line does not rise, and the controller's clock generator, whatever it intended, has not produced a clock edge. The bus decides.
Two consequences follow immediately, and both are design constraints rather than observations.
A controller that drives SCL push-pull cannot be stretched. It can be fought — the target pulls low, the controller drives high, and the result is a contended level and current through both devices. This is not a controller that "does not support stretching"; it is a controller that violates the electrical layer, and it will damage margins on every bus it is fitted to.
A controller that counts time instead of watching the line cannot be stretched either. Its output stage may be perfectly open-drain, but if its state machine advances on an internal counter rather than on the observed rise of SCL, it will sample the next bit while the line is still low. The target is stretching correctly; the controller simply is not looking. This is the commonest real form of the defect, and its signature is data corruption rather than a hang — the transfer completes, with bits misaligned.
That second point is why 17.3 builds the generator as a request and 17.10 closes the loop with the observed line: the timing generator proposes, and the bus disposes.
3. The Obligation, Stated Precisely
For a controller, the rule is short:
After releasing SCL, do not treat the high phase as begun until SCL is observed high. Every subsequent timing measurement starts from that observation, not from the release.
The second sentence is the one that gets dropped, and dropping it produces a subtler defect than ignoring stretching entirely. A controller can correctly wait for SCL to rise and then time tHIGH from the moment it released, in which case a stretch shortens the high phase it actually delivers — possibly below tHIGH(min). The transfer looks stretched and handled, and the clock it produces is out of specification.
There is a matching obligation on the target that is easy to overlook: it must release. A target that stretches indefinitely has not slowed the bus down; it has taken the bus away from every other device on it, including devices it has nothing to do with. 18.11 §5 calls this the one failure a target can inflict on everybody.
4. The Specification Sets No Upper Bound
This is the fact that turns stretching from a protocol topic into an engineering judgment, and it is worth stating starkly: there is no maximum stretch. A target may hold SCL low for as long as it needs, and remain compliant.
Which means:
Every timeout a controller implements declares some legal target non-compliant. The only question is which ones.
That is not a criticism of timeouts. A controller with no timeout is the 24.1 §6 specimen: a single misbehaving device hangs the bus permanently, the failure produces no wrong data, and no data-integrity check can see it. A timeout is necessary. It is simply not a protocol requirement — it is a system policy, and it belongs in the same conversation as the rest of the system's tolerances.
Some profiles built on I²C do bound it. SMBus specifies a 35 ms clock-low timeout precisely because a general-purpose system bus cannot tolerate an unbounded hold (15.3). That is exactly the move being described: a profile adding a constraint the base specification declines to make, because the profile knows something about the system that I²C does not.
5. The Number Is Not the Policy
Here is where the reasoning becomes measurable, and the result is not obvious.
Two timeout blocks. Same parameter, same name, same value. They differ in one line: where the counter is cleared.
module to_per_bit #(parameter integer LIMIT = 500)
(input wire clk, input wire rst_n, input wire in_txn, input wire stretched,
output reg fired);
integer cnt;
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin cnt <= 0; fired <= 0; end
else if (!in_txn) begin cnt <= 0; fired <= 0; end
else if (!stretched) cnt <= 0; // released: restart
else if (cnt == LIMIT - 1) fired <= 1;
else cnt <= cnt + 1;
end
endmodule
module to_per_txn #(parameter integer LIMIT = 500)
(input wire clk, input wire rst_n, input wire in_txn, input wire stretched,
output reg fired);
integer cnt;
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin cnt <= 0; fired <= 0; end
else if (!in_txn) begin cnt <= 0; fired <= 0; end
else if (cnt == LIMIT - 1) fired <= 1;
else cnt <= cnt + 1; // never restarts
end
endmoduleSweeping the per-bit stretch length against both, with LIMIT = 500 in each:
LIMIT = 500 clocks, identical in both blocks
per-bit counter : tolerates a per-bit stretch up to 490 clocks; reports a stuck bus after 500 clocks
per-txn counter : tolerates a per-bit stretch up to 50 clocks; reports a stuck bus after 500 clocksA factor of nearly ten in the legal behavior each tolerates, from the same parameter value. And identical performance on the thing the timeout is nominally for — both detect a genuinely stuck bus in 500 clocks.
So the per-bit version is strictly better? No, and this is the part that makes it a real tradeoff. Ask the other question: how long may a transfer occupy the bus before each intervenes?
9 bits (1 bytes), each stretched 490 clocks: per-bit completed (never fired), per-txn fired at 500 clocks [bus would have been held 4446 clocks]
18 bits (2 bytes), each stretched 490 clocks: per-bit completed (never fired), per-txn fired at 500 clocks [bus would have been held 8892 clocks]
27 bits (3 bytes), each stretched 490 clocks: per-bit completed (never fired), per-txn fired at 500 clocks [bus would have been held 13338 clocks]
36 bits (4 bytes), each stretched 490 clocks: per-bit completed (never fired), per-txn fired at 500 clocks [bus would have been held 17784 clocks]The per-bit counter never fires, at any transfer length. A four-byte transfer holds the bus for 17 784 clocks — thirty-five times the parameter — and every individual stretch is legal, so there is nothing to object to bit by bit. The per-transfer counter caps occupancy at 500 regardless of length, and pays for it by killing a legal 50-clock-per-bit stretch.
| per-bit counter | per-transfer counter | |
|---|---|---|
| bounds | one stretch | total bus occupancy |
| legal per-bit stretch tolerated (LIMIT 500) | 490 clocks | 50 clocks |
| stuck-bus detection | 500 clocks | 500 clocks |
| worst-case occupancy, 4-byte transfer | 17 784 clocks | 500 clocks |
| grows with transfer length | yes | no |
The parameter value is not the policy. Where the counter is cleared is the policy. A review that asks "is there a timeout?" gets a yes and learns nothing; a review that asks "what does it bound?" gets an answer with consequences.
6. Choosing a Policy From Requirements
Neither architecture wins in general. Which one is right follows from what the rest of the system cannot tolerate — and that is a question about the system, not about I²C.
Bound the stretch (per-bit) when the bus carries devices whose legal stretches are long and variable — an EEPROM mid-write-cycle, a sensor mid-conversion — and when nothing else in the system has a hard deadline on the I²C transaction. The risk accepted is unbounded occupancy, which is acceptable when this controller is the only user.
Bound the transaction (per-transfer) when something upstream has a deadline: a watchdog, a control loop, a driver with a timeout of its own, or another initiator waiting for the bus. The risk accepted is aborting a legal transfer, and the mitigation is that LIMIT must be sized against the longest legal transfer the bus can produce — which means it is a function of the device list and the transfer lengths, not a round number.
Both, with different values, is a defensible third option and often the right one: a per-bit bound sized to the slowest device's worst legal stretch, and a per-transfer bound sized to the system's deadline. They catch different faults, and the cost is one more counter.
Whichever is chosen, two things belong in the block's documentation next to the parameter, because neither is recoverable from the code: what the number bounds, and which legal target behavior it declares out of scope.
7. What Can and Cannot Be Verified About Stretching
Stretching is an unusually good example of a feature where a passing test says less than it appears to, because a correctly handled stretch produces a transfer that is byte-for-byte identical to one with no stretch at all. The data is the same. The only difference is timing.
Verifiable by simulation. That the controller waits — measurable by holding SCL and checking no sample is taken. That the timeout fires at the right count. That the state machine resumes correctly at every point a stretch can be inserted, which is the coverage question worth crossing: stretch position × transfer phase (24.2 §5).
Requires a deliberate check. That the transfer completed. A controller stuck in a stretch corrupts nothing, so a data-oriented scoreboard reports zero mismatches (24.1 §6 measures exactly this: zero mismatches, zero transfers). The check must be on completion, and the bench's own waits must be bounded or the bench hangs alongside the design.
Requires the right stimulus to mean anything. A coverage bin for "stretching occurred" can be hit thousands of times by traffic no check distinguishes from unstretched traffic. What makes a stretched transfer checkable is a stretch that changes something observable — one long enough to trip or nearly trip the timeout, one placed at a byte boundary versus mid-byte, one during the acknowledge slot.
Not settleable in simulation. Whether a real device on a real bus will stretch longer than the chosen bound. That is a property of a part and its datasheet, and the honest entry is the list of devices and their worst documented stretch, with the margin stated.
8. Common Misconceptions
"Clock stretching is a target feature." It is a target action and a controller obligation. The only design it constrains is the controller's, and a controller that cannot be stretched is defective rather than simpler.
"A controller that doesn't support stretching just hangs." Almost never. It clocks on regardless and the data is corrupted — the target, still busy, misses bits, and the capture shows nine clocks with everything arriving a bit late. Expecting a hang sends the investigation toward the wrong symptom.
"Set the timeout long enough and legal targets are safe." Long enough for what? The specification sets no bound, so there is no value that is safe for every legal target. What a long value does buy is a worse stuck-bus detection time, which is the other half of the trade.
"The timeout value is the timeout policy." Measured above: the same 500 tolerates a 490-clock stretch or a 50-clock stretch depending on one line of code, and the worst-case bus occupancy differs by 35×.
"Stretching just makes things slower, so it costs throughput and nothing else." It also costs determinism, which is often the scarcer resource. A control loop that samples a sensor every millisecond does not care about average throughput; it cares that no single transaction can take longer than its period, and that is a per-transaction bound.
"We tested stretching — there's a directed test and a coverage bin." A correctly handled stretch produces identical data, so the test passes for reasons that may have nothing to do with stretching. Ask what the test would still pass with, and whether any check distinguishes the stretched transfer from the unstretched one.
Two systems where stretching was handled, and the handling was the problem
1The watchdog that reset a subsystem every few minutes
// An SoC with an I2C controller, an EEPROM, and two sensors. A supervisory
// watchdog resets the subsystem if a task misses its 20 ms deadline.
//
// The I2C block has a timeout. It is a documented parameter. It has NEVER
// FIRED in the field, which is taken as evidence that I2C is not involved.
//
// else if (!stretched) cnt <= 0; // <-- restarts at every release
//
// so TIMEOUT_CLKS bounds ONE STRETCH. The EEPROM, mid-page-write, stretches
// by just under the limit at every bit of every byte -- each stretch legal,
// each one resetting the counter.Watchdog resets every few minutes, correlated with logging activity. The I2C error counters are all zero. Bus captures look legal: every stretch is within the configured limit, every byte transfers correctly, every transaction eventually completes.
Everything in the capture is legal, which is exactly why it was hard. The defect is not in any bit -- it is in the SUM, and the per-bit counter cannot see a sum by construction.
Measured on the specimen in Section 5: with LIMIT = 500, a four-byte transfer whose every bit is stretched by 490 clocks holds the bus for 17 784 clocks and the per-bit counter never fires. That is 35x the number written in the parameter, and it grows linearly with transfer length -- so a page write is worse than a register read by a factor nobody computed.
The mismatch is between what the block bounds (one stretch) and what the system needs bounded (one transaction, against a 20 ms deadline).
Add a per-transaction counter alongside the per-bit one. They catch different
faults and the cost is a second counter:
per-bit LIMIT_BIT sized to the slowest device's worst legal stretch
-- from the datasheets, not from a round number
per-txn LIMIT_TXN sized to the system deadline, 20 ms here, minus the
rest of the task's budget
Then write both numbers, and what each BOUNDS, next to the parameters. Neither
fact is recoverable from the code by the next person, and the failure above is
what it costs to rediscover them.
Finally, add the check that would have caught it in simulation: assert a bound
on TOTAL transaction duration, not on individual stretches. The per-bit
behaviour was already covered and already passing.2The controller that waited for SCL and still corrupted data
// A bit-banged controller. The author knew about stretching and handled it:
//
// scl_release();
// while (!scl_read()) ; // wait for the target to let go. CORRECT.
// delay_ns(T_HIGH); // <-- measured from... when?
// sample_sda();
// scl_drive_low();
//
// The wait is right. The delay is started from the moment the loop EXITS,
// which is correct here -- but the sibling routine got it wrong:
//
// scl_release();
// t0 = now(); // <-- started at RELEASE
// while (!scl_read()) ;
// delay_until(t0 + T_HIGH); // stretch is SUBTRACTED from tHIGH
// sample_sda();Works with every device except one sensor, which stretches roughly 2 us after its address byte. With that device, one bit in every few hundred transfers is wrong. A capture shows all nine clocks present and the stretch correctly honoured -- the controller genuinely waited.
The controller waited and then delivered a SHORT high phase. Timing tHIGH from the release rather than from the observed rise means the stretch is deducted from the high phase: a 2 us stretch against a 2.5 us tHIGH leaves 0.5 us, well under tHIGH(min) for the mode.
So the target is stretching correctly, the controller is waiting correctly, and the clock is out of specification. Both halves of the obligation are needed and only one was implemented:
1. do not proceed until SCL is observed high -- done 2. measure every subsequent interval FROM THAT OBSERVATION -- not done
This failure is invisible to a test whose targets do not stretch, because with no stretch the two reference points coincide.
scl_release();
while (!scl_read()) ;
t0 = now(); // the clock starts when the LINE rises
delay_until(t0 + T_HIGH);
sample_sda();
scl_drive_low();
and the test that distinguishes them, which no unstretched test can: insert a
stretch of a swept length and MEASURE the delivered high phase. Correct
behaviour holds tHIGH constant as the stretch grows; this defect shows tHIGH
falling linearly until it crosses tHIGH(min).
That sweep is the general shape of a stretching test worth having. A single
stretch either side of a threshold tells you almost nothing, because a
correctly handled stretch produces a transfer identical to an unstretched one
-- the information is in how a measured quantity VARIES with stretch length.9. Reason It Through
A. A controller is specified with a 25 ms timeout "because SMBus uses 35 ms and we want margin." What is wrong with the reasoning, and what would the right derivation look like?
Two things. First, the direction is backwards: a shorter timeout is not more margin, it is less tolerance of legal targets — margin against a stuck bus and tolerance of a slow device pull in opposite directions, and "we want margin" does not say which. Second, and more fundamental, SMBus's 35 ms is a bound that profile chose for its own systems, and importing it imports an assumption about a different bus. The right derivation has two independent inputs. From below: the longest stretch any device on this bus can legally produce, taken from datasheets — an EEPROM's write cycle, a sensor's conversion time — plus margin. From above: the longest the rest of the system can tolerate the bus being unavailable, from the watchdog, the control loop, or the other initiator. If the second is smaller than the first, the design has a real conflict that no timeout value resolves, and the fix is elsewhere: a different device, an acknowledge-polling protocol instead of stretching (16.4), or a separate bus segment.
B. Why does a controller that ignores stretching produce corrupted data rather than a hang, and why does that matter for debugging?
Because it never waits for anything. It drives or releases SCL on its own schedule and samples on its own schedule, so it always completes — it simply samples a target that has not produced the bit yet. The transfer finishes, with a plausible-looking wrong value. It matters for debugging because "hang" and "corruption" send investigations in opposite directions: a hang points at a wait and a lock-up, which is where people look for timing problems, while corruption points at data paths, endianness, and register maps. The diagnostic that separates them is in the capture: count the clocks. Nine clocks per byte with the target's data arriving one bit late throughout is the stretching signature, and it looks nothing like a data-path bug once you know to look for it.
C. A verification engineer reports "stretching coverage: 100 %, 12 000 stretched transactions, zero failures." What would you ask?
Which check distinguished a stretched transaction from an unstretched one. A correctly handled stretch produces byte-identical data, so 12 000 stretched transactions can all pass through a data scoreboard that is entirely blind to stretching — the bin is measuring the stimulus generator. Then three follow-ups. What was the distribution of stretch lengths, and did any approach the timeout, since a stretch far from every boundary discriminates nothing? Where in the transfer were they placed — after the address, mid-byte, in the acknowledge slot — because the interesting failures are position-dependent? And is there a check on transaction completion, or only on data, since the failure this feature actually produces is one that transfers nothing and therefore mismatches nothing?
D. A team proposes removing stretching support from their target to simplify it, arguing they can always respond in time. What does that decision actually commit them to?
To a guarantee about worst-case internal latency for every response the device owes, under every condition, forever — including at the lowest supported clock frequency, with the cache cold, during a low-power wake, and after any future firmware change. That is a strong and easily invalidated claim, and the failure mode when it breaks is not a stretch but silent data corruption, because a device that is late and cannot stretch simply supplies the wrong bit. It also commits them to it being true at every speed mode the bus may be run at, since a response that fits inside a 10 µs bit period may not fit inside 1 µs. The decision can be right — a simple register target with a single-cycle read path genuinely has no reason to stretch — but the reviewable form of it is not "we can always respond in time"; it is a worst-case latency number, the fastest supported bit period, and the comparison between them.
10. Understanding Check
11. What 24.6 Settled
Three questions, three different answers. Why a device stretches is about that device; how stretching works is the wired-AND on SCL; who must do what constrains only the controller. Most confusion is one of these offered as an answer to another.
A released line is not a driven line. The controller's clock generator proposes and the bus disposes — which is why a push-pull SCL cannot be stretched but fought, and why a controller that counts time rather than watching the line corrupts data instead of hanging.
The obligation has two halves, and the second is usually missing. Wait for the observed rise, and time everything after it from that observation. Doing only the first delivers a short high phase that a stretch-free test can never expose.
Every timeout is a policy, because the specification sets no bound. A controller without one can be hung permanently by one device, silently, with nothing corrupted for any data check to notice. A controller with one has declared some legal target out of scope. Both facts have to be on the table at once.
The number is not the policy. Measured: the same LIMIT = 500 tolerates a 490-clock or a 50-clock per-bit stretch depending on where the counter clears, with identical stuck-bus detection — and the per-bit architecture allows a four-byte transfer to hold the bus for 17 784 clocks while never firing.
What a timeout bounds, and which legal behavior it excludes, belong in the documentation. Neither is recoverable from the code, and rediscovering them costs a field failure.
The last of the three mechanisms is the one where no participant has a complete view of what is happening. Chapter 24.7 — Reasoning About Multi-Master Arbitration.
Continue learning
Related tutorials
- Related topic
Why I²C Clock Stretching Exists
The one mechanism that lets a target push back on a clock it does not own. Covers the mismatch it solves, the specification's two stretching levels, and why 'optional' makes a non-stretching bus an electrical contract.
- Related topic
The I²C Stretching Mechanism — Holding SCL Low
Stretching needed no new mechanism: the specification already described it for multi-master synchronization. One sentence decides whether a master survives it — and getting it wrong collapses the high phase on the bit a stretch ended.
- Related topic
Wired-AND — Many Drivers, One Line, No Contention
Put several open-drain devices on one node and a logic function appears in the wiring: any participant asserting LOW wins, and HIGH requires unanimous release. Derive dominant LOW, see why two devices pulling together is agreement rather than conflict, and watch three of the protocol's mechanisms become predictable.
- Related topic
Address Allocation, Strapping, Conflicts and Bus Switches
112 addresses are available and boards still collide constantly, because almost no part lets you choose freely. How real devices move their address, what a strap pin is in silicon, how strapping can land a device in reserved space, and how muxes make one address appear twice.
