Skip to content
VLSI Mentor

SPI · Module 1

Electrical and Board-Level Limits

Why an SPI link with identical logic runs at one clock rate and fails at a higher one. Push-pull drivers into real capacitance, the four delays inside a bit time, the round trip a returned bit must complete, and why usable SCLK is a property of the board.

Every chapter in this module has treated the wires as ideal. A bit driven at one end appeared at the other; a clock edge produced by the master reached the slave; Chapter 1.3's ring shifted as though the two registers were adjacent on a die.

They are not. They are on separate packages, connected by copper with resistance, inductance and — most importantly here — capacitance, and separated by a distance that light itself takes measurable time to cross. For most of what the preceding chapters taught, ignoring that was correct: ownership, the shift model and full duplex are all true regardless of how fast the link runs.

This chapter is where the abstraction has to be paid for, because it answers the question that ends more SPI bring-up sessions than any protocol misunderstanding:

The logic is unchanged. The data is unchanged. Why does it work at one clock rate and fail at a higher one?

The answer is not that SPI has a maximum frequency. It is that a bit time is a budget, and the things that spend it live on your board.

1. What a Push-Pull Driver Has to Do

Chapter 1.2 introduced push-pull outputs to explain contention. Here the same structure matters for a different reason: what it takes to change a wire's voltage at all.

A driver's job is to move the net between two voltage levels. That net is not an abstract connection — it is a conductor with capacitance to its surroundings, contributed by the trace itself, by the input pin of every device attached to it, by vias and connectors, and by any test point or probe. To change the net's voltage, the driver must move charge into or out of that capacitance, and it can only supply so much current.

That gives the first physical limit, and it is the one most people already half-know:

A stronger driver or a smaller capacitance means faster edges. A weaker driver or a larger capacitance means slower edges. The transition is not instantaneous; it takes a rise time and a fall time, during which the net's voltage is somewhere between the two valid logic levels.

And here is why that matters for clock rate rather than for correctness: a receiver's input is only guaranteed to interpret the net correctly when the voltage is beyond its threshold with margin. The time spent in transition is time during which the net is not carrying a usable value. Halve the bit time and the transition occupies twice the fraction of it. Keep halving and eventually the net never fully arrives before it is asked to change again — and what a receiver sees is a signal that no longer reaches valid levels.

2. Four Delays Inside One Bit Time

Take a single bit travelling from slave to master — the harder direction, for reasons §3 makes clear — and enumerate what must happen between the clock edge that starts it and the moment the master can trust it.

Trace delay on SCLK. The master's clock edge is produced at the master's pin. It has to travel along a trace to reach the slave's pin. Propagation along a PCB trace is fast but finite, and it is a function of the trace's length and the board material.

The slave's clock-to-output time. Once the edge arrives, the slave's internal logic must respond and its output driver must place the new bit on MISO. A datasheet names this — often as an output valid or output delay parameter — and it includes the driver's own transition time into a specified load capacitance. Crucially, that specified load is part of the number: if your board presents more capacitance than the datasheet's test condition, the real delay is larger than the published figure.

Trace delay on MISO. The bit now travels back from the slave's pin to the master's pin — a second traverse of the board, in the opposite direction.

The master's setup requirement. The master's capture flip-flop needs the data stable for a defined interval before the capturing edge. Arriving is not enough; arriving early enough is the requirement.

Four contributions, all positive, all additive. None of them is a protocol parameter. Every one of them is a property of two specific parts and one specific board.

3. The Round Trip

Now assemble them, because the order in which they occur is the point.

The five stages of the SPI return path. The master launches an SCLK edge. The edge propagates along the board and arrives at the slave. The slave responds after its clock to output delay by driving MISO. The MISO signal propagates back along the board to the master pin. The master then requires setup time before its capture edge. The four delays are serial and additive.Master launches an SCLK edget = 0, at the master's pinEdge arrives at the slaveafter board propagationSlave drives the bit onto MISOafter its clock-to-output delayBit arrives at the master's pinafter board propagation, againMaster captures itneeds setup time before its edgeSCLK trace delayclock-to-outputMISO trace delaysetup needed12
Figure 1 — the round trip a returned bit must complete. The clock travels out, the slave responds, the data travels back, and the master still needs setup time — all before the capture edge. These four delays are serial, so they add.

Two things about this chain explain most of what surprises people about SPI speed.

The board is crossed twice. The clock goes out and the data comes back, so trace delay is paid once in each direction. Doubling the distance to a peripheral does not add one trace delay to the budget; it adds two. This is why a device at the far corner of a board can be the one that limits the whole bus.

The master's own output path is not in this chain. The master launching MOSI has an easier job: its data travels one way and the slave's setup requirement is met at the far end, with the clock arriving at the same place at roughly the same time. There is no return leg. MISO is structurally harder than MOSI, and on a fast link it is almost always MISO that fails first — which is worth remembering, because the symptom (received data wrong, transmitted data fine) is often misread as a device problem.

4. Writing the Budget Down

Express the constraint symbolically. Let the four contributions be, in order, t_pd_sclk, t_v_slave, t_pd_miso and t_su_master. Let t_window be the interval the master allows between the edge that causes the slave to launch and the edge on which the master captures. Then the link works only if:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
t_pd_sclk  +  t_v_slave  +  t_pd_miso  +  t_su_master   ≤   t_window

Everything interesting follows from reading that inequality carefully.

The left-hand side is almost fixed; the right-hand side is what you control. Trace delays are set by layout. The slave's clock-to-output is set by the part and its actual load. The master's setup is set by the controller. None of them shrinks because you want a faster link. t_window, however, is inversely proportional to SCLK: raise the clock and the window closes. This is the entire mechanism behind "works slow, fails fast."

t_window is usually a half period, not a whole one. In the common arrangement, the slave launches on one clock edge and the master captures on the next — half a period later. So the budget an engineer actually has is roughly T/2, and the maximum usable SCLK is roughly half what a naive "the round trip fits in a bit time" estimate suggests. Which edge plays which role is configuration, and deriving that properly is Module 2 and Module 3's work — but the shape of the constraint is the same whatever arrangement you choose.

Some controllers can widen the window. Because this constraint is so commonly the limit, many SPI masters offer a configurable sampling delay — capturing MISO a half cycle later than the naive position, which trades away some hold margin to buy round-trip time. If a controller offers such an option and a link fails only at high clock rates, that setting is worth understanding before redesigning the board.

No term in that inequality is specified by SPI. Not one. This is why the next section's claim is not a hedge.

5. Why There Is No Universal Maximum SCLK

A datasheet will tell you the maximum clock frequency its part supports. That number is necessary and not sufficient, because it describes one device in isolation under the test conditions the vendor chose.

The link's usable rate is bounded by the worst of several things: the lowest maximum among all devices on the bus; the round-trip budget of §4 for the physically furthest or slowest-responding device; and whatever the board's signal integrity permits at the resulting edge rates. A bus of parts each rated for a high frequency can be limited well below all of their individual ratings by a long trace or a heavily loaded net.

So the honest statement is: the usable SCLK is a property of a system, established from datasheets and layout, not a property of SPI. A tutorial that quoted a number here would be describing someone else's board. When you need the figure for your design, you compute it from §4's inequality using the parameters your parts actually publish — which is the datasheet skill Module 10 develops and the throughput argument Module 9 builds on.

6. Why Adding a Device Can Slow the Whole Bus

Chapter 1.5 showed that SCLK, MOSI and MISO are shared across every peripheral. Electrically, sharing means every device's input capacitance is attached to those nets, all the time — including devices that are not selected.

Add a peripheral and the shared nets gain capacitance. Every driver on them now has more charge to move per transition, so edges become slower. Slower edges mean more of each bit time spent in transition and less spent at a valid level — and, on the MISO net, they lengthen the effective clock-to-output of whichever slave is driving, because that parameter includes the driver's transition into its load.

This produces a genuinely counter-intuitive result worth stating plainly: adding a device that is never selected can reduce the rate at which the rest of the bus works. It contributes capacitance whether or not it participates. Selection is a logical mechanism; capacitance is not selective.

The same reasoning explains two other familiar board-level effects. A long stub to a distant peripheral adds both propagation delay and capacitance to a shared net. And attaching a scope probe adds capacitance too — which is why a marginal link occasionally behaves differently when you measure it, and why a link that only works with the probe attached is telling you something real about margin rather than about the probe.

7. Ringing, Overshoot and Source Termination

One more physical effect deserves naming, because it produces a symptom that looks like a logic fault.

When a driver switches quickly into a trace, the trace behaves less like a lumped capacitor and more like a transmission line: the edge travels along it, and where the trace's characteristic impedance does not match what it meets at the far end, part of the energy reflects back. The reflection travels to the driver, may reflect again, and the result is ringing — the signal overshooting its target level and oscillating around it before settling.

Two consequences matter for SPI specifically.

Ringing on a data line costs settling time. The net eventually reaches a valid level, but later than a clean edge would, which spends budget from §4 that the accounting did not include.

Ringing on the clock line can be much worse. A clock edge that overshoots and rings back through a receiver's threshold can be seen as more than one edge. A slave that counts an extra clock shifts one bit too many, and the symptom is a received word offset by a bit — which looks exactly like the one-bit shift Chapter 1.3 attributed to an edge-alignment error. Two very different causes, one signature, which is why Module 18 treats clock integrity as a thing to rule out before re-reading a mode table.

The standard mitigation is source-series termination: a small resistor placed at the driver, in series with the trace, close to the driver's pin. It slows the edge slightly and absorbs the returning reflection rather than letting it re-reflect. It is a routine, cheap measure on fast SPI clock lines, and boards intended to run near their limits usually have a footprint for one whether or not it is populated.

That is as far as this chapter goes into signal integrity. The point is not to teach transmission-line theory; it is that a "digital" signal is an analogue waveform that we have agreed to interpret digitally, and near the limit that agreement needs help.

8. Why an RTL Designer Cares

It is tempting to file all of this under board design. Three parts of it reach the RTL.

The logic is almost never what limits an SPI link. The datapath of Chapter 1.3 is a shift register — a handful of flip-flops with no arithmetic, which will close timing comfortably on any device you are likely to target. When an SPI link fails at a higher clock, the cause is overwhelmingly in the paths this chapter describes, not in the shift logic. Knowing that stops you from optimising the wrong thing.

Configurability is what buys margin later. A master whose clock divider is fixed at design time gives an integrator no options when a board turns out to be marginal. A master with a programmable divider — and, better, a configurable sampling point of the kind §4 described — lets the same RTL work across boards whose electrical properties differ. That is a design decision made early and regretted late, and it is one reason Module 13 treats the divider as a first-class block.

Where you capture the input decides how much of the budget you control. t_su_master in §4's inequality is not a fixed property of the chip; it depends on the path from the input pin to the capturing flip-flop. Capturing close to the pin makes that path short and predictable; capturing after a long stretch of fabric routing adds delay you did not budget for and that changes between builds. The RTL choice is where you place that register, and §10 is the FPGA-specific version of the same point.

9. Why a Verification Engineer Cares

This chapter marks the boundary of what functional simulation can tell you, and being explicit about that boundary is itself a verification skill.

A zero-delay simulation cannot fail this way. In RTL simulation, the round trip of §3 takes no time: the clock edge, the slave's response and the master's capture all happen in the same simulation instant, ordered only by event scheduling. Every test passes at every clock rate, because the simulation's notion of clock rate has no relationship to the physical budget. A design can be functionally perfect and electrically unusable, and the functional testbench will report success.

So the checks that matter here are not in the same environment. Round-trip margin is established by static timing analysis against the constraints the design declares, and by measurement on hardware — not by simulation. What a testbench can usefully do is verify that the mechanisms which buy margin actually work: that the clock divider produces every rate it claims, that a configurable sampling-point setting genuinely moves the capture edge, and that the design behaves correctly at its slowest and fastest configured divider settings. Those are functional properties of margin-buying features.

Gate-level simulation with annotated delays occupies the middle ground. It can expose setup and hold violations inside a device that a zero-delay run hides, and it is the right tool for some of this — but it still knows nothing about your board's trace delay or the capacitance of a peripheral the vendor did not model. Knowing which questions each tool can answer, and refusing to accept a green functional run as evidence about electrical margin, is the transferable lesson. Module 15 covers the constraints and Module 18 the hardware measurements.

10. Why an FPGA Engineer Cares

This is the chapter with the most direct FPGA consequences in Module 1, and they are constraint-shaped rather than code-shaped.

Your half of the budget is declared, not discovered. t_su_master and the master's own output timing are not properties the tool infers — they follow from where the registers sit and what you have told the tool about the outside world. The mechanism is input and output delay constraints, which describe the board's contribution so the tool can close the on-chip part of the path against it. A design with no such constraints has not met timing on these paths; it has simply not been asked to, and the result is a build that works or fails depending on where the router happened to put things. Module 15 is entirely about getting this right.

I/O-block registers are the cheapest margin you will ever buy. Most FPGAs provide flip-flops in the I/O block, right at the pad. Capturing MISO there makes the pin-to-flop path short, fixed and predictable rather than a routed path that varies between builds; driving MOSI and SCLK from there does the same for the output side. It costs nothing but an attribute or a constraint, and it converts a variable term in §4's inequality into a nearly constant one.

As a slave, the budget changes shape entirely. Everything above assumed the FPGA is the master and owns SCLK. As a slave it receives SCLK from outside, so the clock's arrival is not something the design controls, and the relevant question becomes how that clock reaches the capture logic and how its domain relates to the system clock. That is a different and harder problem — the one Chapter 1.2 flagged and Module 15 solves.

11. Common Misconceptions

12. Reason It Through

Work this before reading the answers.

A design reads an ADC over SPI at a modest clock rate and works perfectly. To meet a higher sample rate, the clock is raised. Above a certain frequency the received samples become intermittently wrong, while the commands the ADC receives continue to be interpreted correctly — the converter performs the right conversions and simply returns values the master mis-reads. The RTL and the data are unchanged.

Why does the asymmetry between the two directions point at the mechanism immediately? Because MOSI and MISO have structurally different timing. MOSI travels one way, and the slave's setup requirement is met at the far end with the clock arriving alongside it. MISO requires the full round trip of §3 — clock out, slave response, data back, master setup — so it has roughly twice the board delay plus the slave's clock-to-output in its budget. MISO fails first. Commands being obeyed correctly while returned data is corrupt is the signature of a return-path timing limit, not of a protocol or encoding error.

Which term in §4's inequality changed when the clock was raised? Only t_window. The four left-hand contributions are set by the parts and the layout and did not move. Raising SCLK shortens the interval between the slave's launch edge and the master's capture edge until the round trip no longer fits — at which point the master samples MISO before the data has arrived and settled.

Why intermittent rather than a clean failure at a threshold? Because the terms are not single values. Propagation and clock-to-output vary with temperature, supply and process, and the sampled level near the threshold depends on the previous bit through the settling of the edge. Just past the limit, some bits still make it and some do not, and which ones depends on data and conditions — so the failure is data-dependent and looks random.

What should be measured, and what is the decisive observation? Look at MISO at the master's pin relative to the capture edge. If the data is transitioning at or just before the moment the master samples — rather than sitting stable — the round trip has run out of room. That is a direct confirmation, and it distinguishes this from a device or command problem in one capture.

What are the fixes, in order of cost? First, check whether the controller offers a delayed sampling point; widening t_window by half a cycle is a configuration change and often sufficient. Second, reduce the left-hand side: shorten the MISO and SCLK routing to this device, and remove unnecessary load from those nets. Third, consider source-series termination if the edges are ringing and stealing settling time. Only after those, accept a lower SCLK and find the sample rate elsewhere — by reading more per transaction, for instance, which is the efficiency argument Module 9 develops.

Why is "add a delay in the RTL" usually wrong here? Because the shortage is in a path that leaves the chip, and inserting logic delay in the capture path typically consumes margin rather than creating it. The exception is precisely the controller feature named above — moving the capture to a later edge, which is a different operation from adding combinational delay, and which trades hold margin for setup margin deliberately rather than accidentally.

13. Understanding Check

14. Summary

Every earlier chapter in this module treated the wires as ideal, and for ownership, the shift model and full duplex that was correct. Here the abstraction is paid for.

A push-pull driver must move charge into a real capacitance to change a net's voltage, so edges take time, and time spent in transition is time the net carries no usable value. Inside each bit time four delays must fit, in series: clock propagation to the slave, the slave's clock-to-output, data propagation back, and the master's setup requirement. That chain is the round trip, it crosses the board twice, and it applies to MISO but not to MOSI — which is why MISO fails first and why "commands obeyed, returned data wrong" is a timing signature rather than a device fault.

Written as a budget, the four terms must fit inside the window the master allows between the slave's launch edge and its own capture edge — commonly about half an SCLK period. The left-hand side is fixed by parts and layout; the window is what the clock rate controls. That is the entire mechanism behind a link that works slowly and fails quickly, and it is why the failure is intermittent and data-dependent near the limit rather than clean.

Because none of those terms is specified by SPI, there is no universal maximum SCLK — only a rate your parts and your board can sustain, computed from published parameters. And because shared nets carry every attached device's capacitance whether or not it is selected, adding a peripheral that is never accessed can reduce the rate the rest of the bus achieves. Ringing adds a further hazard that is specifically dangerous on the clock line, where an edge crossing a threshold twice is counted twice and produces a word offset by a bit — the same signature a mode error produces, from a completely different cause.

15. What Comes Next

That closes Module 1. You now have the architecture and the mental model: why the interface exists and what it trades, who owns each wire, the shift ring that moves the bits, the exchange that ring makes unavoidable, how the topology scales, and where physical reality limits the abstraction.

Two things have been deferred at every step, and Module 2 — SPI Timing Fundamentals collects them. This chapter kept saying "the window the master allows between the launch edge and the capture edge" without ever pinning down which edges those are. That is next: SCLK generation, the leading and trailing edge vocabulary, launch and sample edges as a contract, setup and hold in SPI's own terms, CS-to-clock timing, and the slave-output and master-input parameters this chapter treated symbolically. Module 3 then derives CPOL, CPHA and the four modes from that contract rather than asking you to memorise a table.

Browse the path on the SPI curriculum index, or revisit Bus Topologies and Daisy-Chain for the shared nets whose loading this chapter costed. For the same physical reasoning applied to a bus that deliberately uses weak pull-ups instead of push-pull drivers — and pays for it in exactly the edge-rate currency described here — see Why I²C Exists.

Continue learning