Skip to content
VLSI Mentor

I²C · Module 24

Designing an I²C Master at the Whiteboard

Producing the architecture under time pressure and defending it: the order to derive the blocks in, why the framing sequencer is structurally necessary, the line-ownership handoff that separates a real answer from a shift register, and impact analysis for eight requirement changes. Closes the I²C curriculum.

"Design an I²C master." Forty minutes, a whiteboard, no simulator, and somebody watching.

What is being tested is not recall. A memorised block diagram is worth very little, because the follow-up question is always a change — now make it multi-controller, now put it on an FPGA, now the target stretches — and a remembered picture has no way to answer that. What is being tested is whether the blocks on the board are there for reasons you can state, because only then can you say what happens to each one when the requirement moves.

This chapter is the last in the I²C curriculum. It does not introduce new architecture; Module 17 built all of it. It is about producing that architecture from first principles, in order, under pressure, and defending each piece.

1. The First Ninety Seconds

Do not draw anything yet. Four questions, and each one changes the answer materially.

askwhy it changes the design
What speed mode?sets tr(max), the pull-up window, and the minimum system clock (24.3 §3)
One controller or several?arbitration detection and loss recovery are either present or absent — not a feature to add later
Do any targets stretch?decides whether SCL is a request or an assertion, which is structural
FPGA or ASIC, and what system clock?decides the oversampling ratio, the I/O primitives, and whether the latency budget closes

Two more worth asking if there is time: does the host drive this through registers or an on-chip bus, and is there a deadline on a transaction from anything upstream.

Then state your assumptions out loud, because you will be designing against them:

"I'll assume Fast-mode at 400 kHz, a single controller for now but I'll say where multi-controller support attaches, targets that may stretch, an FPGA at 50 MHz, and a register interface. Tell me if any of those is wrong."

That sentence does more work than the next ten minutes of drawing. It shows the decisions are conscious, it invites correction before effort is spent, and — the part that matters most in a review rather than an interview — it produces a written record of what the design was built against.

2. Derive the Blocks, Do Not Recall Them

The decomposition follows from two observations, and deriving it takes about three minutes. Doing that instead of reproducing a diagram is the difference between an answer that survives the follow-up question and one that does not.

Observation one: four unrelated clocks. A master must simultaneously track four kinds of state, and they advance on four unrelated events (17.1 §2):

stateadvances onevent source
TIME — where in tLOW/tHIGH the bus isa divider tickthe internal system clock
BIT — which of the nine pulses of a bytean SCL edge read back from the busthe bus, not this master
BYTE — which byte, and its directionthe ninth pulse completingthe bit engine
TRANSACTION — which phase of a combined transfera host commandsoftware, asynchronously

A single state machine holding all four takes the product of their states and needs four sets of transitions out of every one. The counters that inevitably get bolted on to make that tractable are the other three machines, admitted late and without their own reset or tests. So the four time bases are the four blocks — not as a style preference, but because any correct master contains them whether or not anyone drew boxes.

Observation two: framing and data have contradictory obligations. A bit engine exists to guarantee SDA is stable while SCL is high. A START requires SDA to move while SCL is high. Those are opposite postconditions on the same wire in the same phase, so the block that generates framing cannot be the block that transmits bits (17.1 §3). That is a fifth block, derived with certainty before any code exists.

Then the wire itself. A device drives low or releases; nothing sources a high; the only way to learn the line's state is to read it back. That gives the feedback path — and one comparator serves two protocol features, because "released and still low" means a stretch on SCL and an arbitration loss on SDA (17.10).

Draw it in that order and the diagram assembles itself:

Azvya Education Pvt. Ltd.VLSI Mentor
The order to draw it in, and what each step is forced by
   1.  the two wires, as three signals each     drive_low out, line_in back, oe to the pad
   2.  SCL generator          TIME              phase counts from the divider
   3.  bit engine             BIT               drive point, sample point, one bit
   4.  byte engine            BYTE              eight bits + the ACK slot, ownership handoff
   5.  transaction controller TRANSACTION       write / read / combined, ACK policy
   6.  framing sequencer      the contradiction START, repeated START, STOP
   7.  SDA ownership arbiter  four claimants    priority, and overlap reported as an error
   8.  bus feedback           read-back         stretch on SCL, arbitration loss on SDA
   9.  error manager          what to report    timeouts, taxonomy, recovery
   10. host interface         the request       registers or an on-chip bus

Steps 1 and 8 are the ones candidates leave out, and leaving them out is what turns the answer into a shift register with a clock generator — a design that would ignore stretching, miss arbitration, and never recover a stuck line.

3. The FSM, Derived Last

The transaction controller's state machine is the only FSM worth drawing on a whiteboard, and it is drawn last, because its states are transitions between things the other blocks already do.

A state machine drawn as a ring of eight states with a ninth, ABORT, in the centre. IDLE leads to START on a host command, then ADDR, then AACK where the acknowledge arrives, then DATA, then DACK. DACK returns to DATA while bytes remain and goes to STOP on the last byte. STOP leads to DONE and DONE returns to IDLE. ADDR and DATA both lead inward to ABORT when arbitration is lost, and ABORT returns to IDLE once the bus is free, so the transfer can be retried from the beginning.IDLES_STARTS_ADDRS_AACKS_DATAS_DACKS_STOPS_DONES_ABORThost commandhost commandSTART issuedSTART issued8 bits8 bitsACKACK8 bits8 bitsbytes remainbytes remainlast bytelast byteSTOP issuedSTOP issuedhost clearshost clearsarb lostarb lostarb lostarb lostbus free: retrybus free: retry
Figure 1 — the transaction controller's states, using this curriculum's names from Chapter 17.8. Every state is a phase in which some other block is working: the framer in START and STOP, the byte engine in ADDR and DATA, and the bit engine's ninth slot in AACK and DACK. ABORT sits in the middle because it is reachable from every state that holds the bus — arbitration is lost on a bit, not at a boundary. The NACK path out of AACK is described in the text rather than drawn, to keep the ring readable.

Three things an interviewer will look for on this drawing, and all three are about arrows rather than circles.

S_ABORT is reachable from every state that has the bus. Not from a byte boundary — arbitration is lost on a bit, and the loser must stop driving on that bit (24.7 §4). A candidate who draws arbitration loss as a transition out of S_DACK has understood it as an error code rather than as a bus event.

The NACK path out of S_AACK goes to S_STOP. It is not drawn above because it is the one edge that crosses the ring, and it is worth saying rather than drawing: a target that does not acknowledge its address is not an error condition needing its own machinery — it is a normal outcome, and the controller's response is the same STOP it would issue at the end of a successful transfer. The same edge exists from S_DACK for a data NACK.

There is no state for stretching. Stretching does not advance or block this machine; it prevents the SCL generator from observing a rise, so the BIT state does not advance, so the byte boundary never arrives. It is absent from this diagram because it belongs to a different time base — and being able to say that is worth more than adding a S_WAIT circle.

S_TOFRM, or an equivalent, is missing here on purpose. The real controller has a state between the last data byte and the STOP, for taking SCL back from the generator, because if SCL rises between owners the framer's next act — pulling SDA low — is a START rather than the intended STOP (17.12 §2). Saying "there is a handover state here and it exists because the ownership transfer must overlap" is a strong whiteboard moment. Drawing nine tidy states and not knowing why the tenth exists is the weaker version of the same picture.

4. The Thing Most Candidates Get Wrong: Line Ownership

If there is time for one more figure, this is the one — because it is where a plausible answer and a correct one diverge.

Four rows over eight cells: SCL, SDA, the owner of SDA, and the transaction controller's state. The owner changes from none while idle, to the framer during START, to the bit engine during the address byte, to the target during the acknowledge slot, and back to the framer for the STOP in the final cell.framer drives the STARTframer drives the STARTcontroller releases; target pullscontroller releases; targetpullsSCLSDASDA ownernoneframebitbitbitbittgtframestateIDLESTARTADDRADDRADDRADDRAACKSTOPt0t1t2t3t4t5t6t7
Three owners inside one transfer, and one of them is not this device.
Figure 2 — who owns SDA through a short write, alongside the controller's state. The line has three different owners inside one transfer, and each handoff is a moment where two blocks must not both drive. The acknowledge slot is the important one: the controller releases and the target pulls, which is the same mechanism as clock stretching and arbitration and is why nothing may ever drive SDA high.

Four claimants want SDA — the framer, the bit engine, the byte engine's acknowledge (which drives through the bit engine), and the recovery sequence — so there is an arbiter with a priority order, and the framer is highest because a START must pre-empt a data bit. The detail worth volunteering is that the arbiter should report an overlap as an error rather than resolve it silently: two claimants at once is a design fault, and a priority encoder that quietly picks one hides it forever (17.4).

And the pin is three signals, not one: drive_low out, line_in back, and the pad's output enable derived from the first. An inout in synthesizable RTL can be neither synthesized nor simulated against a second driver (17.1 §4).

5. Error Handling, and Where Each Case Is Caught

An interviewer who asks "what can go wrong" is asking whether the design has somewhere to put each answer. Six cases, and the useful structure is that each is detected in a different place.

what goes wrongdetected whereresponse
address NACKbyte engine, ninth slotSTOP, report — it is a normal outcome, not a fault
data NACK from the targetbyte engine, ninth slotSTOP, report the byte count that completed
arbitration lostbus feedback, on any bit this master releasedstop driving on that bit, no STOP, retry after bus-free
target stretches beyond the boundSCL generator + timeoutabort, report, recover — and see §6 on which bound
SDA stuck lowbus feedback, when idlethe nine-pulse recovery sequence (15.4)
SCL stuck lowbus feedback, when idlenothing this master can do — recovery clocks SCL, and SCL is the line that is stuck

The last row is the one worth saying out loud, because it is the one that demonstrates the electrical model is real rather than memorised. The recovery procedure works by clocking SCL until a target that is holding SDA releases it. If SCL itself is held low by a faulty device, no other device can lift it — the line is low if anyone pulls it low — so there is no in-band recovery. The honest design reports it and escalates to whatever can cycle power.

The other one to volunteer: recovery must never run during a transfer. Both lines are low most of the time while clocking, so a stuck-line diagnosis taken mid-transfer reads a transmitted zero as a stuck line and clocks nine pulses into a live byte. It is one AND gate, and it is the difference between a recovery feature and a corruption source (17.12 §2).

6. Defending the Decisions

The follow-ups are predictable. Short answers that name the mechanism.

"Why not one state machine?" Four unrelated event sources. One FSM takes the product of four state spaces and needs four sets of transitions out of every state; the counters added to make that tractable are the other three machines without their invariants.

"Why is the framer separate from the bit engine?" Contradictory postconditions on the same wire in the same phase: the bit engine guarantees SDA is stable while SCL is high, and a START requires it to move while SCL is high.

"Why read SCL back instead of using your own clock?" Because the controller does not drive SCL high — it releases it, and the pull-up raises it if nobody else is pulling. A design that advances on its intended edge rather than the observed one samples while the line is still low and corrupts data without hanging (24.6 §2).

"Where does arbitration detection live, and what does it cost?" The same comparator as stretch detection: released and read back low. On SCL that is a stretch; on SDA it is a loss. It must run on every bit this master released, including the address (24.7 §5 measures what skipping the address costs: 9.3 % of losses undetected).

"What's your timeout?" The strong answer names what it bounds before naming a number. A per-bit counter bounds one stretch and permits unbounded total occupancy; a per-transaction counter bounds the transfer and will abort a legal long stretch. Same value, opposite failure modes (24.6 §5).

"How do you know the design works?" Not "it passes the tests". Which claims are supported by which observations, and what a defect would have to look like to survive them — including the single cheapest experiment available, which is to remove the target and require the environment to fail (24.2 §2).

7. Requirement Changes: The Real Exercise

This is where forty minutes are actually spent, and it is the part a memorised diagram cannot survive. For each change: what breaks, and what moves in RTL, in verification, in constraints, and in debug.

the requirement changes towhat breaksRTLverificationimplementationdebug
400 kHz → 1 MHz (Fm+)the pull-up window and the latency budgetdivider values; filter depth must stay under the narrower tHIGH(min)re-run every timing-sensitive test at the new divider; add the Fm+ boundarytr(max) 120 ns needs a 20 mA sink or a current source — a BOM change, not a resistor changeedges become the first suspect for everything
single → multi-controllernothing adds cleanly — detection must be on every bitarbitration comparator on every released bit; S_ABORT from every bus-holding state; retry-from-starta second controller agent; the oracle becomes the wired-AND of intents, not any device's viewnoneper-controller checkers can no longer attribute failures (24.7 §8)
no stretching → stretching targetsSCL stops being an assertiongenerator becomes a request; every interval timed from the observed risesweep stretch length and check a derived quantity, not datanonea controller that ignores it corrupts rather than hangs
FPGA → ASICthe I/O and the debug storypad cell replaces the FPGA primitive; the open-drain intent is unchangedunchangedCDC signoff replaces a constraint file; STA on a real libraryno ILA — trace infrastructure must be designed in, or it does not exist
50 MHz → 20 MHz system clockthe observation-latency headroompossibly fewer filter samples, which is the wrong lever to pull firstre-verify at the new divider; boundary values movere-check the budget: Fast-mode stops fitting at all below 11.1 MHz (24.3 §3)marginal timing appears as speed-dependent failures
fixed → configurable target addressthe acknowledge decision path lengthensaddress compare now reads a register or strappingthe address space becomes a swept dimensionthe decision time in the tVD;ACK budget is no longer the assumed 3 clocksa late ACK looks exactly like a wrong address
simple registers → side-effect registersatomicity assumptionsthe register interface needs a commit point; reads may now need to stretcha reference model that models the side effect, not just storagenonea read that changes state cannot be repeated to reproduce a failure
add a hard transaction deadlinethe timeout policya per-transaction counter alongside the per-bit oneassert a bound on total transaction duration, which per-bit coverage never exercisednonewatchdog resets with every I²C error counter at zero (24.6)

Two rows deserve emphasis because they are the ones most often answered too lightly.

Multi-controller is not a feature you add later. It is a property of how every bit is transmitted. Retrofitting it means touching the bit engine, the transaction controller, the error taxonomy, the environment's oracle and the scoreboard's structure. The right whiteboard answer is to say where it attaches now, even while designing single-controller.

FPGA → ASIC changes almost no RTL and almost all of the evidence. The open-drain intent, the FSM and the timing are unchanged. What changes is that the constraint file becomes a CDC signoff, the timing report becomes STA against a real library, and the ILA — which was the entire bring-up plan — does not exist. A candidate who answers "swap the I/O primitive" has answered the RTL question and missed the project.

8. What a Strong Answer Actually Demonstrates

Not completeness. Ten dimensions, and an answer that does six of them well is stronger than one that mentions all ten.

Requirements clarified before drawing. Assumptions stated and revisited when they change. Clean interfaces — each block's inputs are things somebody can produce. Ownership named — which block drives which line in which phase. State and data separated — the FSM decides, the datapath moves bytes. Failure handled — each of §5's six cases has a home. Verification strategy — what would be checked and against what oracle. Implementation awareness — what changes between FPGA and ASIC. Observability — what a bring-up engineer can see. Tradeoffs reasoned — a choice presented with what it costs.

The honest note is that under time pressure nobody covers all of them, and an interviewer knows it. What separates answers is not coverage but whether the parts that are present are connected by reasons.

9. How to Handle What You Do Not Know

The question you cannot answer is not the failure. Pretending is.

"I don't know the exact tSU;DAT for Fast-mode — it's a couple of hundred nanoseconds. What I'd do is take it from Table 10 and add it to the transmitter's valid time and the rise time, and check the sum fits inside tLOW(min). That sum is tight at Fast-mode and doesn't close at Fm+ with a resistive pull-up, which is why Fm+ assumes a different pull-up."

That answer contains one admission and three pieces of engineering: where the number lives, what it is combined with, and what the combination implies. It is a stronger answer than the correct number recited alone — and it is a more accurate picture of real work, where the number is always available and the structure of the budget is what you have to carry with you.

The same move works for a design question. "I'd need to know the worst legal stretch across the devices on this bus before choosing that bound" is not evasion; it is naming the input the decision depends on, which is exactly what §1 was for.

10. Reason It Through

A. You have drawn the architecture and the interviewer says: "Good. Now the host is a CPU running an RTOS, and the I²C transaction must complete within 5 ms or a watchdog fires." Work the change.

It converts the timeout from a bus-protection feature into a system deadline, and those need different counters. The per-bit bound protects the bus from a hung device; it cannot bound a transaction, because every individual stretch resets it and a transfer can accumulate arbitrarily many legal stretches — measured at 17 784 clocks for a four-byte transfer with a 500-clock per-bit limit. So a per-transaction counter is added alongside, sized from 5 ms minus the rest of the task's budget. Then two consequences fall out. In verification, the assertion to add is a bound on total transaction duration, which nothing in a per-bit coverage model exercises. And in the design, S_ABORT now has a second cause with different semantics: arbitration loss means retry, deadline abort means give up and tell software — so the error taxonomy grows a code rather than reusing one.

B. The interviewer points at the bit engine and asks why the acknowledge slot is inside the byte engine rather than being a ninth iteration of the bit engine. Answer it.

Because SDA changes hands in it and nothing else about the byte does. For the eight data bits the master owns SDA throughout — it drives, the target reads. In the ninth the master releases and the target pulls, so the direction of ownership inverts for exactly one bit. A bit engine parameterised to sometimes-drive-sometimes-release has to be told which, by something that knows where it is in the byte — and that something is the byte engine, which means the knowledge has moved there anyway while the mechanism stayed in the bit engine. Putting the slot in the byte engine keeps the ownership decision where the byte position is known. The deeper reason is the one worth adding: the ninth slot is where this master learns whether the target exists, so it is the only bit of the nine that produces information rather than transmitting it.

C. A candidate draws a WAIT_FOR_SCL state in the transaction FSM for clock stretching. Is it wrong, and what would you say?

It is not wrong in behavior and it is wrong in placement, which is worth separating because the candidate has understood stretching and misplaced it. Stretching acts on the TIME and BIT bases: the SCL generator releases, the line does not rise, the bit engine never sees the edge, so the byte boundary never arrives and the transaction controller simply does not advance. The waiting already happens, one level down, without a state. Putting a state in the transaction FSM means the transaction level has to know about SCL edges, which reintroduces the coupling the decomposition removed — and it will be wrong in a specific way: a stretch during the address byte and a stretch during a data byte would need separate wait states, or one wait state with a return-path register, which is a state machine growing a stack. The thing to say is that it is a correct behavior placed in the wrong time base, and to ask which block observes the SCL rise.

D. Forty minutes are nearly up and you have drawn the blocks but no FSM. The interviewer asks what you would do with ten more minutes. What is the highest-value thing to add?

The feedback path and the line-ownership handoff, not the FSM. The FSM is the most derivable part of the design — given the blocks, its states are nearly forced, and an interviewer can see that they are. The feedback path is the part that distinguishes an I²C master from a shift register with a clock generator, and it is the part most commonly missing: read SCL back to detect stretching, read SDA back on released bits to detect arbitration loss, one comparator serving both. The ownership handoff is the second: which block drives SDA in each phase, that the acknowledge slot belongs to the target, and that an overlap between claimants should be reported rather than silently prioritised. Ten minutes on those two says more about whether the design would work on a real bus than a complete FSM would.

11. Understanding Check

12. What 24.8 Settled

Derive, do not recall. Four unrelated event sources give four blocks; contradictory obligations on SDA give a fifth; the wired-AND gives the feedback path. A diagram produced that way can answer the follow-up question, and a remembered one cannot.

Two blocks are the ones that get left out. The pin modelled as three signals, and the read-back path. Without them the answer is a shift register with a clock generator — a design that ignores stretching, misses arbitration, and cannot recover a stuck line.

The FSM comes last and is nearly forced. What an interviewer reads on it is the arrows: abort reachable from every state that holds the bus, no state for stretching, and a handover state before the STOP that exists because SCL must not rise between owners.

Line ownership is where plausible and correct diverge. Three owners inside one transfer, one of which is not this device, and an arbiter that reports overlap rather than hiding it.

The requirement change is the real exercise. Eight of them in §7, each with consequences in RTL, verification, implementation and debug — and two where the easy answer is wrong: multi-controller cannot be added later, and FPGA to ASIC changes the evidence rather than the code.

13. Module 24, and the End of the Track

Module 24 asked a different question from the twenty-three modules before it. Those asked what I²C does and how to build it. This one asked how to decide, on incomplete evidence, whether something is right.

The eight chapters were one argument. 24.1 established that a review question is worth asking only if some observation answers it either way, and measured four cases where a purposeful test passed a defective design. 24.2 turned that on the testbench and measured an environment that passes with the device removed, and a defect surviving at 100 % coverage. 24.3 sorted the claims simulation cannot reach and found where the latency budget closes. 24.4 catalogued the models that survive because they are right almost everywhere. 24.5, 24.6 and 24.7 worked the three mechanisms that rest on the wired-AND, each ending in a number rather than a definition. This chapter put them back together under time pressure.

What connects them is a single discipline, and it is the thing worth carrying out of the track:

Know what has been proven, and by what observation.

Everything else in Module 24 is a consequence. A review question is a demand for an observation. A coverage bin is not one. A passing regression is consistent with an environment that cannot fail, and one parameter separates the two. A timeout number says nothing without saying what it bounds. A pull-up value is a decision made against an estimate, and the margin worth quoting is the capacitance at which it stops working. A controller cannot know it is alone on the bus — only that it has not yet been contradicted.

That last one is a fair description of engineering generally. You do not get certainty; you get evidence that has not yet been contradicted, and the professional skill is knowing exactly how strong it is.

The I²C track ends here — one hundred and fifty chapters from two wires and a pull-up to defending an architecture under pressure. The mechanisms are specific to I²C. The habits are not: state the assumption, name the evidence, distinguish the specification from the implementation choice, and be exact about what remains unproven.

Those transfer to every bus you will ever work on.

Continue learning