SPI · Module 2
Setup, Hold, and Timing Margin
What makes a bit electrically safe to capture. The stability window a receiver demands, why violating it produces undefined rather than merely wrong results, and which engineering layer owns each part of the answer.
Chapter 2.3 established that a receiver captures on an edge, and that the transmitter must not be changing the line at that instant. It justified the separation with a phrase it did not define: the receiver "requires its input to be stable for an interval before and after the capturing edge."
This chapter defines that interval, and it is where SPI stops being a protocol and starts being electronics.
What makes a bit electrically safe to capture?
The answer is a window, not an instant — and the reason it is a window rather than an instant is physical, not conventional. Once you have it, the half-period budget from Chapter 2.3 becomes something you can actually spend, and the remaining chapters of this module are the itemised bill.
1. A Flip-Flop Does Not Sample at a Point
It is convenient to draw a capture as an arrow at an instant. The hardware does not work that way.
A flip-flop captures by allowing its input to influence an internal storage node and then cutting it off and letting positive feedback drive that node to a solid 0 or 1. Both of those take time. The input must have been present long enough to push the storage node decisively toward one state before the cut-off, and it must remain present briefly through the cut-off so the decision is not reversed while it is being made.
That gives two requirements, and they are separate because they describe different intervals on opposite sides of the edge:
Setup time — the minimum interval the input must be stable before the capturing edge. Call it t_su.
Hold time — the minimum interval the input must remain stable after the capturing edge. Call it t_h.
Together they define an aperture window of width t_su + t_h straddling the edge. The data must not change anywhere inside it. Outside it, the data may do whatever it likes.
2. What Violation Actually Does
This is where intuition usually fails, and the failure is expensive.
The natural assumption is that violating setup or hold makes the flip-flop capture the old value instead of the new one, or vice versa — a wrong bit, deterministic and debuggable. That does happen, but it is the benign outcome. The real possibility is worse: the storage node can be left balanced between the two states, and the feedback that is supposed to resolve it can take an arbitrarily long time to do so. This is metastability.
Three consequences matter, and each one changes how you interpret a symptom.
The resolution time is unbounded, not merely long. There is no maximum. There is only a probability distribution: the chance of still being unresolved decays exponentially with the time allowed. This is why metastability is characterised statistically, as a mean time between failures, rather than as a worst case you can design against absolutely.
A metastable output can propagate. While unresolved, the node sits at an intermediate voltage. Downstream logic reading it may interpret it as 0, as 1, or — for gates with different thresholds — differently in different places, so a single flip-flop's indecision can produce a state that is not merely wrong but internally inconsistent.
It is data- and condition-dependent, so it hides. The probability of an aperture violation depends on exactly where the data edge lands relative to the clock edge, which depends on the data pattern, the temperature, the supply voltage and the individual part. A link can pass every test on the bench and fail in the field, and the rate can be one transfer in millions.
For an SPI link this is the mechanism behind a specific and very common report: "it works, but occasionally a byte is wrong, and we cannot reproduce it." That is not a software bug. It is an aperture violation somewhere on the link, and no amount of retrying the transfer fixes the cause.
3. The Window, Drawn
Setup and hold — the window the data must not move in
10 cyclesRead the figure as a prohibition rather than a requirement. The shaded band is the interval in which the data line must not transition. The safe trace changes early and settles; the violating trace is mid-transition when the window opens and is still moving at the edge. Note that the violating trace does eventually reach a perfectly good logic 1 — arriving at the right value is not the issue. Arriving in time to be still is.
The x regions are drawn deliberately: during a transition the line genuinely has no defined logic value, and that is the whole point.
4. Margin: What Is Left Over
Now connect the window to Chapter 2.3's budget.
The receiver needs its window. The system provides an interval between the launch edge and the sample edge — half a period. Margin is the difference:
setup margin = (time from data becoming valid to the sampling edge) − t_su
hold margin = (time the data stays valid after the sampling edge) − t_hA link is safe when both margins are positive, with something left over for the variation §5 describes. Note immediately that they are separate quantities that fail in opposite directions:
Setup margin shrinks when data arrives late. Raising the clock frequency shortens the half period and therefore squeezes setup margin. This is the failure mode of a fast link, and it is what Chapter 2.7's round trip consumes.
Hold margin shrinks when data arrives too early — or changes too soon after the edge. Raising the frequency does not directly hurt hold margin, which is why hold problems are frequency-independent and therefore often survive every speed-related test anyone thinks to run. A design that fails hold fails at every frequency, including slow ones, which is a useful diagnostic asymmetry.
That asymmetry is worth carrying: setup failures are fixed by slowing down; hold failures are not. If a link misbehaves identically at 1 MHz and 20 MHz, slowing it further is not going to help and the investigation should be looking at hold, at the resting level (Chapter 2.2), or at edge roles (Chapter 2.3).
A worked budget
Hypothetical numbers, chosen to make the arithmetic visible. They are not from any device.
SCLK is 20 MHz, so the period is 50 ns and the half period is 25 ns. A receiver requires
t_su = 6 ns. The data becomes valid 14 ns after the launch edge.
The data is valid for the remainder of the half period: 25 − 14 = 11 ns before the sampling edge. Setup margin is therefore 11 − 6 = **5 ns** — positive, so the capture is nominally safe.
Now double the clock to 40 MHz. The half period becomes 12.5 ns. If the data still takes 14 ns to become valid, it is not valid until after the sampling edge has already passed. Setup margin is (12.5 − 14) − 6 = −7.5 ns, and the link cannot work at that rate no matter how clean the board is. The 14 ns did not change; the budget did.
This is the arithmetic behind every "works slow, fails fast" report, and it is worth doing explicitly the first time a link is designed rather than after it fails.
5. Margin Is Not a Single Number
A budget computed once, from typical values, is a statement about one part at one temperature. Real margin must survive variation, and there are four sources worth naming.
Process. Individual devices differ. A parameter's published maximum is a guarantee across the distribution, not a measurement of the part on your bench — which is exactly why a design that works on three prototypes can fail in volume.
Voltage. Propagation and output delays lengthen as supply voltage falls. A margin measured at nominal supply can vanish at the low end of the tolerance band.
Temperature. Delays generally increase with temperature in the regimes most digital designs operate in. A link characterised at room temperature has not been characterised.
Clock uncertainty. The sampling edge does not arrive at a mathematically exact time. Jitter and, within a chip, skew between the clock's arrival at different flip-flops both subtract from the margin you calculated.
The practical consequence is a discipline rather than a formula: use the worst-case numbers a datasheet publishes, not the typical ones, and leave margin beyond zero. How much beyond is an engineering judgement that depends on volume, environment and the cost of a field failure — but a design sitting at 0.2 ns of calculated margin using typical values has, in truth, no margin at all.
6. Which Layer Owns What
This is the most valuable distinction in the chapter, and blurring it is how engineers end up arguing past each other. Several different activities all get called "timing," and they answer different questions with different tools.
| Layer | What it constrains | Who states it | What checks it |
|---|---|---|---|
| Protocol timing | That launch and sample are separate edges, that CS brackets the transfer | The SPI convention and the device datasheet | Reading the specification; protocol assertions |
| RTL architecture | Which internal event drives each strobe; where registers sit | Your design | Simulation, code review |
| Synthesis / implementation | On-chip delay from register to pin and pin to register | The tool, against your constraints | Static timing analysis |
| I/O constraints | What the design must guarantee at its pins, and require of its inputs | You, in an SDC or equivalent | Static timing analysis |
| Device timing | The peripheral's t_su, t_h, output delay, maximum clock | The peripheral's datasheet | Reading it; bench measurement |
| PCB | Propagation delay, loading, edge quality | Layout | Field solver, oscilloscope |
| Receiver electrical | The aperture window itself | Silicon physics | Nothing you can run — it is the ground truth |
Two rules follow from the table, and both are violated routinely.
A functional simulation cannot fail a setup or hold check. In a zero-delay RTL simulation, the round trip takes no time and every frequency works. Reporting "simulation passes" as evidence about margin is a category error — it is evidence about logic, which is a different question. Chapter 2.3's exclusivity assertion proves two events were in different simulation cycles, not that a signal had settled at a pin.
Static timing analysis only covers the paths you constrained. STA is rigorous and exhaustive within its model, and its model is the constraints you wrote plus the library's delays. It knows nothing about the trace to the peripheral or that peripheral's internal delay unless an input or output delay constraint told it. An unconstrained I/O path is not a passing path; it is an unexamined one. Module 15 covers writing those constraints.
7. The Vocabulary in a Datasheet
The terms above appear in device specifications, usually in a timing table with a diagram beside it. Recognising them is the bridge from this chapter to real work.
- A data setup parameter — how long before the sampling edge the device needs its input stable. This is the
t_suyour master must satisfy on MOSI. - A data hold parameter — how long after the sampling edge the device needs it held. Your master must not change MOSI too soon.
- An output valid, output delay or clock-to-output parameter — how long after an edge the device takes to present data on MISO. This is the device's contribution to your budget, and Chapter 2.6 is devoted to it.
- Clock high and clock low times, and a maximum frequency — the constraints Chapter 2.1 showed you must check separately.
- CS setup and CS hold parameters — Chapter 2.5's subject.
Two cautions about reading them. Parameters are stated with a condition: an output delay is valid for a specified load capacitance, and your board is probably not that load. And a parameter is a minimum requirement or a maximum guarantee, never a typical value to design against — the direction matters, and mixing them up inverts the whole calculation.
Full datasheet interpretation, including how to find these when a vendor names them something unexpected, is Module 10.
8. What Each Tool Can Actually Tell You
A short but load-bearing section, because reaching for the wrong instrument wastes more time than any other habit in timing debug.
RTL simulation shows event ordering and logical correctness. It is the right tool for "does my state machine issue the strobes in the right order." It is the wrong tool for anything involving nanoseconds on a board, because it has no model of them.
Static timing analysis shows whether the implemented on-chip paths meet the constraints you declared, across the corners you asked for. It is the right tool for "will my FPGA present MOSI early enough, given what I told it about the board." It is silent about anything you did not constrain.
A logic analyser shows decoded, threshold-crossed digital activity over long windows. It is the right tool for "was CS framed correctly, was the command what I expected, how many clocks were there." It is the wrong tool for measuring a 3 ns setup margin, because its own sampling granularity and threshold behaviour are comparable to what you are trying to measure.
An oscilloscope shows the actual voltage. It is the only instrument that can answer "was the line settled when the edge arrived," and the only one that shows ringing, slow edges and intermediate levels — the things §2's aperture violations are made of. Measure at the receiver's pin, because that is where the requirement applies; the same signal at the driver looks better than what the receiver sees.
Nothing in this list is a substitute for another. An engineer who says "the logic analyser shows it is fine" has established that the digital decoding worked on the transfers captured, which is worth knowing and is not a margin measurement.
9. Why There Is No RTL in This Chapter
Deliberate, and worth explaining rather than leaving as a gap.
Setup and hold are properties of a receiving flip-flop and the signal reaching it — a characteristic of silicon and of the path between two pins. There is no Verilog, SystemVerilog or VHDL that implements them; a behavioural model can check them in a timing-annotated simulation, but writing such a check would teach a simulation feature rather than the engineering idea.
What RTL genuinely owns here is a different question — where you place the capture register, which determines how much of the path you control. Chapter 2.1 made that point for the output side and Chapter 2.7 makes it for the input side, which is where it is decisive. Adding a code example here would be code for the sake of a quota, and would push the reader toward thinking margin is something you write rather than something you budget.
The correct representation for this chapter is the window, the arithmetic and the layer table. That is what §3 to §6 provide.
10. Failure Signature — Works at 10 MHz, Fails at 50 MHz
Symptom. A link is reliable at a low clock rate. Raised to a high one, it returns occasional wrong bytes — not consistently, not on every transfer, and with the error rate increasing as the clock rises further.
Candidate mechanisms. Setup margin exhausted somewhere on the link is the leading hypothesis, because the symptom tracks frequency and is intermittent. The competing explanations are an over-clocked device (Chapter 2.1 — check the datasheet maximum first, it is free), degraded edge quality from loading or missing termination (Chapter 1.6), and contention on a shared return line (Chapter 1.2), which is also intermittent but is not frequency-dependent in the same monotonic way.
The discriminating observations. Take them in order.
First, is the behaviour monotonic in frequency — solid below some rate, worsening steadily above it, with no recovery at higher rates? A margin shortfall behaves exactly like that. Contention does not: it depends on data patterns rather than on rate.
Second, which direction fails? If the master receives badly while the slave receives correctly, the round trip is the constraint and Chapter 2.7 owns the analysis. If both directions fail together, look at clock quality — a degraded clock hurts every capture on the link.
Third, put a scope on the failing line at the receiving pin and look at where it settles relative to the sampling edge. This is the measurement that converts a hypothesis into a fact, and it is why §8 insists on the instrument distinction.
Why "slow it down until it works" is a diagnosis, not a fix. Lowering the clock restores setup margin and the symptom disappears — which confirms the mechanism and tells you nothing about which delay consumed the budget. The link is then running slower than it needs to, with an unknown amount of margin, and the same board will fail again at a different temperature or with a different part lot. Use the frequency sweep to confirm the class of fault, then measure to find the term.
11. Common Misconceptions
12. Reason It Through
Work this before reading the answers.
An SPI link to a sensor is reliable at 8 MHz. At 24 MHz roughly one transfer in a few thousand returns a corrupt byte; at 32 MHz the rate rises noticeably. The sensor's datasheet permits 50 MHz. A colleague argues that since 24 MHz is well under the device's rating, the fault cannot be timing, and proposes adding a CRC and retrying failed transfers.
Why does "well under the rated maximum" not exonerate timing? Because the device's maximum frequency is a statement about the device, not about the link. Chapter 1.6 showed the link's budget also contains board propagation in both directions and the receiver's own setup requirement, none of which appear in the sensor's rating. A part rated to 50 MHz on the vendor's evaluation board can easily be limited to a fraction of that by a long trace, a heavily loaded net, or a controller with a large input path.
What does the error-rate behaviour tell you? It is the strongest evidence in the statement. An error rate that is zero below a threshold and increases monotonically with frequency is the signature of a shrinking margin: as the window closes, a larger fraction of the naturally occurring variation pushes a capture into the aperture. Faults with other mechanisms do not behave this way — contention tracks data patterns, a configuration error is rate-independent and total, and a software bug does not care about the clock at all.
Is the CRC-and-retry proposal reasonable? It is a reasonable mitigation for residual errors and a poor response to this one, for two reasons. It leaves the link operating with negative or near-zero margin, so the error rate will move with temperature, supply and part lot — and a rate that is tolerable in the lab may not be in the field. And metastability's resolution time is unbounded, so a marginal capture can occasionally produce something worse than a wrong byte: a value read inconsistently by downstream logic. Retrying handles detected corruption; it does not handle a system that is operating outside its electrical requirements.
What should actually be measured? Establish which direction fails first — if the master's receive path is the one corrupting, the round trip is the constraint. Then scope MISO at the master's pin against the master's sampling edge and see where the signal settles. That single measurement distinguishes "the sensor responds too slowly for this rate" from "the board's propagation is eating the budget" from "the edge is so degraded that the line never settles," and each has a different fix.
What is the correct engineering outcome? Either find and reduce the dominant term — a delayed sampling point if the controller offers one, shorter or lighter routing, better clock integrity — or set the clock to a rate with genuine, calculated margin against worst-case numbers, and document what limits it. Choosing a lower frequency deliberately, with the budget written down, is a different act from lowering it until symptoms stop.
13. Understanding Check
14. Summary
A receiver does not capture at an instant. It requires its input to be stable for t_su before the sampling edge and t_h after it — an aperture window straddling the edge inside which the line must not transition. The requirement is about stillness, not about correctness: arriving at the right value too late, or still moving, both violate it.
Violation does not reliably give the old value. It can leave the capturing flip-flop metastable, with an unbounded resolution time and the possibility of being read inconsistently downstream. That is why the resulting failures are intermittent, condition-dependent and hard to reproduce — and why "occasionally one byte is wrong and we cannot repeat it" is a timing report rather than a software one.
Margin is what the system provides minus what the receiver requires, and setup and hold margins fail in opposite directions. Setup margin shrinks as the clock speeds up; hold margin does not depend on the period at all. So a fault that is identical at 1 MHz and 20 MHz is not a setup problem and will not be improved by slowing down — a diagnostic asymmetry worth using before any instrument is connected.
Margin must survive process, voltage, temperature and clock uncertainty, so it is budgeted from worst-case published numbers rather than typical ones, with headroom beyond zero.
Finally, several distinct activities are all called "timing," and they answer different questions: protocol convention, RTL architecture, implementation and STA, I/O constraints, device datasheet parameters, PCB propagation, and the receiver's physical aperture. A simulation cannot fail a margin check, and STA only covers the paths you constrained. Keeping those layers apart is what makes a timing discussion productive.
15. What Comes Next
The window is defined and the budget is understood. Chapter 2.5 — CS-to-SCLK and SCLK-to-CS Timing applies the same thinking to the signal this module has so far treated as a simple enable: chip select has its own setup and hold requirements relative to the clock burst, and violating them breaks a transfer in which every single SCLK edge is perfectly correct. It is the most commonly overlooked timing requirement in SPI, and it explains a failure signature — the first byte wrong, the rest fine — that looks nothing like a margin problem.
Browse the path on the SPI curriculum index, or revisit Launch and Sample Edges for the half-period budget this chapter spends.
Continue learning
Related tutorials
- Related topic
MISO Valid Timing
When returned data may be trusted relative to the launch edge: the round trip separating arrival from validity, why the sampling window has two edges, the configurable input sampler in three HDLs, and the margin no delay can create.
- Related topic
SCLK Generation, Period, and Frequency
Where SCLK comes from and what one period buys. Dividing a system clock to a bus clock, why the divisor is an integer and what that costs, and how a period in nanoseconds becomes the budget every later timing parameter is spent from.
- Related topic
Leading and Trailing Edges
Why rising and falling are the wrong words for an SPI transfer. How the clock's resting level decides which physical edge comes first, and the vocabulary every later timing chapter depends on.
- Related topic
Launch and Sample Edges
One edge of each bit time places a bit on the wire, the other captures it, and they must never be the same edge. Why the separation is forced, why it buys half a period, and how RTL maps physical edges onto those roles.
