Skip to content
VLSI Mentor

SPI · Module 2

Master Input Capture and Round-Trip Delay

The complete return-path budget: clock out, peripheral response, data back, and the master's setup requirement, all inside half a period. Why maximum SCLK is a property of a whole system.

Every term is now on the table. Chapter 2.1 set the period and showed it is reached by integer division. Chapter 2.3 reduced the usable budget to a half period. Chapter 2.4 named the receiver's aperture requirement. Chapter 2.6 supplied the peripheral's output delay and showed it is conditional on your board.

This chapter assembles them, and in doing so answers the question the whole module has been building toward.

Why does the practical maximum SCLK depend on the complete controller → PCB → peripheral → PCB → controller path?

Because that path is a loop, and every stage of it is spent from one half period. Chapter 1.6 introduced the loop qualitatively and deliberately left it symbolic. Here it becomes an inequality you can evaluate, a worked example you can follow, and a set of interventions ordered by what they actually cost.

1. The Asymmetry That Makes This Hard

Start by being precise about why the return direction is the difficult one, because engineers routinely analyse the wrong path.

MOSI is a one-way trip. The master launches a bit and the clock edge that will capture it travels alongside it, over comparable trace lengths, to the same device. Both signals experience similar propagation. The slave's setup requirement must be met, but the geometry is favourable: the data and the clock arrive together, having taken similar routes.

MISO is a round trip. The clock edge must first travel out to the peripheral. The peripheral then takes its output valid time to respond. The data must then travel back. And only then must it satisfy the master's own setup requirement — all before the master's sampling edge, which is one half period after the launch edge that started the sequence.

So the board is crossed twice on the return path and once on the outbound path, and the peripheral's output delay sits in the middle of the return path with nothing comparable on the outbound side. The two directions are not symmetric, and the practical consequence is blunt: MISO fails first, essentially always.

That gives the most valuable single diagnostic in this module. Commands accepted correctly while returned data is corrupt is not a device fault and not a protocol misunderstanding — it is the signature of a return-path budget that has run out.

2. The Loop, Stage by Stage

Five stages, in order, each consuming time.

1 — Master clock output delay. From the internal register that generates SCLK to the master's pin. On an FPGA this is the I/O path Chapter 2.1 argued should be an I/O-block register; on an ASIC it is pad delay. Call it t_co_m.

2 — Outbound propagation. SCLK travels from the master's pin to the peripheral's pin. A function of trace length and board material. Call it t_pd_sclk.

3 — Peripheral output valid. The device responds, taking Chapter 2.6's t_v — internal propagation plus the driver's transition into your load, not the datasheet's test load.

4 — Return propagation. MISO travels from the peripheral's pin back to the master's pin. Call it t_pd_miso.

5 — Master setup requirement. The signal must be stable for t_su_m before the master's capture register samples it — and that register may be in the I/O block or deep in the fabric, which changes the number considerably.

Clock out, respond, data back, capture — inside one half period

10 cycles
One SPI bit time showing the complete return path. The master's clock edge is generated and propagates outward. The peripheral responds after its output valid delay. The data propagates back across the board. The master then requires setup time before its sampling edge. All five intervals occur inside the half period between the launch edge and the sampling edge.outoutt_vt_vbackbacksetupsetuplaunchlaunchsamplesamplesclk @mastersclk @slavemiso @slaveXXXmiso @masterXXXt0t1t2t3t4t5t6t7t8t9
Figure 1 — the round trip inside one bit time. The launch edge starts the sequence; the sampling edge half a period later ends it. The five stages are serial, so they add, and the board contributes twice.

Read the figure by following one bit down the rows. The master's clock edge at its own pin (row 1) reaches the slave's pin slightly later (row 2). The slave's MISO (row 3) is undefined for t_v and then settles. That settled value reaches the master's pin (row 4) later still — and only then does the master's setup window have to be satisfied, before the sampling edge at the far right.

The two miso rows are the point of the figure: what the peripheral drives and what the master sees are not the same waveform, and the difference is propagation you must budget for.

3. The Inequality

Written out:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
t_co_m  +  t_pd_sclk  +  t_v  +  t_pd_miso  +  t_su_m   ≤   t_window

where t_window is the interval the master allows between the launch edge and its own sampling edge — half a period in the standard arrangement.

Four observations about this expression, and they are the chapter.

The left side is nearly fixed; the right side is what frequency controls. None of the five terms shrinks because you want a faster link. t_window, however, is T/2 = 1/(2·f). Raise the frequency and the budget closes while the bill stays the same. This single fact explains every "works slow, fails fast" report in SPI.

The board appears twice. t_pd_sclk and t_pd_miso are both paid. Doubling the distance to a peripheral adds two propagation delays, not one, which is why a device at the far corner of a board frequently limits the whole bus.

t_v is usually the largest term, and it is the one Chapter 2.6 showed is conditional on your load. A budget built from a published figure without checking the test-load condition has probably understated its biggest contributor.

This is a teaching inequality, not an STA equation. It omits clock jitter, it treats propagation as a single number rather than a min/max pair, it ignores duty-cycle distortion through buffers, and it assumes the symmetric clock Chapter 2.1 warned is not guaranteed with an odd divisor. Real signoff uses worst-case corners on every term and a proper static timing analysis of the on-chip portions. Use this form to understand where the time goes and to size a design; use STA and measurement to sign it off.

4. A Worked Example

Hypothetical numbers throughout, chosen so the arithmetic is followable. None of these are from a real device — the point is the method.

A master drives a peripheral across a board. Suppose t_co_m = 3 ns, t_pd_sclk = 1 ns, t_v = 12 ns at the load actually present, t_pd_miso = 1 ns, and t_su_m = 4 ns.

Total round trip: 3 + 1 + 12 + 1 + 4 = 21 ns.

What frequency does that permit? The budget is the half period, so T/2 ≥ 21 ns, giving T ≥ 42 ns and f ≤ 23.8 MHz. Round down to a frequency the divider can actually reach (Chapter 2.1) — from a 100 MHz system clock the options near there are 25 MHz (DIV = 4, over budget) and 20 MHz (DIV = 5, odd, so check the pulse widths). 20 MHz is the workable answer, with T/2 = 25 ns against a 21 ns requirement: 4 ns of margin.

Now notice how thin that is. Four nanoseconds must absorb everything §3 said the inequality omits — temperature and voltage variation on t_v and t_co_m, clock jitter, and the fact that every published figure is a worst-case that individual parts approach differently. On a design that must work across a temperature range, 4 ns is not obviously enough, and the honest response is to either re-derive with worst-case corner values or step down to the next divisor.

Where would you attack it if 20 MHz were not enough? t_v is 12 ns of a 21 ns total — 57% of the budget. Halving the trace lengths would recover 1 ns. Improving the master's setup path might recover 2 or 3. Reducing t_v, by lightening the MISO net or choosing a faster peripheral, is the only change that moves the number substantially. Attack the dominant term, and do the arithmetic before choosing where to spend effort.

What if the peripheral's datasheet says it supports 50 MHz? It does — the device tolerates a 50 MHz clock. The link cannot run there, because the device's rating says nothing about your board's propagation or your master's setup requirement. That gap between a device rating and a system capability is the single most common misunderstanding this module exists to correct.

5. The Master's Own Contribution

Two of the five terms belong to the master, and both are more controllable than engineers usually assume.

t_su_m depends on where the capture register sits. The requirement is not a fixed property of the chip; it is the setup requirement of a specific flip-flop plus the delay from the input pin to that flip-flop. Capture in an I/O-block register and that path is short and, crucially, constant across builds. Capture after a stretch of fabric routing and you have added delay that varies every time the tool runs — so a design can meet timing on one build and fail on the next with no source change. This is the single most effective thing an FPGA engineer can do about the input side.

t_co_m depends on the same choice for the output. Chapter 2.1 argued for an I/O-block register on SCLK, and this is the quantified reason: it is a term in the budget.

Both need constraints, or they are not bounded at all. An input delay constraint tells the tool what the outside world will do, so it can verify the on-chip portion of the path; an output delay constraint does the same for the output side. Without them the paths are unanalysed, and a clean timing report on an unconstrained interface means the tool was never asked. Module 15 covers writing them.

For an ASIC the structure is identical with pad delays in place of I/O blocks, and the interface constrained at the boundary for signoff.

6. What a Delayed Sampling Point Buys

This is the intervention most worth understanding, because it is often free and is frequently the difference between a working link and a redesign.

Many SPI masters offer a configuration that moves the sampling instant later — typically by half a cycle, sometimes by a programmable number of system-clock cycles. The effect on the inequality is direct: t_window grows, often from T/2 to something approaching T.

In §4's example that would change the constraint from T/2 ≥ 21 ns to roughly T ≥ 21 ns, permitting nearly double the frequency from the same physical link. That is a very large gain for a register setting.

What it costs. Hold margin. Chapter 2.4 §4 established that setup and hold fail in opposite directions: moving the sampling point later increases the time available before the capture, and decreases the time the data remains valid after it. Push far enough and the next bit's transition arrives before the hold requirement is satisfied. So the option trades one margin for the other, and the trade must be checked in the direction it weakens — which is exactly the check people forget, because the symptom that prompted the change was a setup symptom.

Why it is not cheating. The peripheral's data remains valid until its next launch edge, which is a full period after it became valid. So there genuinely is more valid time available than the half period uses, and a delayed sample harvests it. The technique is legitimate and widely implemented; it simply must be applied with the hold side verified rather than assumed.

7. Why There Is No RTL in This Chapter

t_co_m, t_pd_sclk, t_v, t_pd_miso and t_su_m are delays through silicon and copper. No Verilog, SystemVerilog or VHDL implements any of them.

What RTL genuinely owns here has already been built: Chapter 2.1's divider sets t_window through the divisor, and Chapter 2.3's strobe generator determines which edge samples. A delayed sampling point is a variation on that strobe selection, and adding it here would duplicate Chapter 2.3's module for a one-line change — exactly the code inflation this curriculum avoids. The production version, with the sampling delay as a configurable field, is Module 13.

The correct representations for this chapter are the loop diagram, the inequality, the worked budget and the intervention ordering. Writing HDL would suggest that round-trip margin is something you code, when it is something you budget, constrain, measure and — when necessary — buy with a slower clock.

8. What Each Tool Can Tell You Here

Chapter 2.4 §8 drew the general distinction; this chapter is where applying it matters most, because the failure lives outside every simulator's model.

RTL simulation cannot see this at all. In a zero-delay simulation the entire loop takes no time: the clock edge, the peripheral's response and the master's capture occur in one instant, ordered only by event scheduling. Every frequency passes. A design can be functionally perfect and electrically unusable with a completely green regression — and this is the specific failure that makes "the simulation passes" a category error rather than merely incomplete evidence.

Static timing analysis covers the on-chip halves, if you constrained them. It bounds t_co_m and verifies the path to your capture register against the input delay you declared. It knows nothing about t_pd_sclk, t_v or t_pd_miso unless your constraints described them — those are the board and the other device, which is why an input delay constraint is a statement about the outside world rather than a property the tool can discover.

An oscilloscope is the only instrument that closes the loop. Probe SCLK and MISO at the master's pin, trigger on the master's sampling edge, and look at where MISO settles. That single measurement is the direct test of the inequality, and it is the reason §10's diagnostic converges quickly.

A logic analyser will mislead you here. It will decode the transaction and show you plausible bytes, because its thresholding recovers a digital value from a signal that was marginal. It cannot measure a few nanoseconds of settling. Using one to investigate a round-trip problem produces the frustrating experience of a capture that looks fine alongside a link that does not work.

9. Verification: Closing the Timing Monitor

Chapter 2.5 introduced a timing monitor as a stopwatch recording CS and edge times. This chapter supplies its last check and its honest limitation.

The monitor can record the interval from each launch edge to the moment MISO settles, and compare it against the configured budget — which is a genuine check in a timing-annotated simulation, where the slave model of Chapter 2.6 imposes a realistic t_v and the interconnect carries delay. It catches a master configured to sample too early for the modelled peripheral.

What it cannot do is establish the board's contribution. t_pd_sclk and t_pd_miso are not in any simulation unless somebody put them there, and putting them there means encoding an estimate of a physical layout. A monitor reporting adequate margin has verified the model you built, and the model's fidelity to the board is a separate question that measurement answers.

The genuinely useful verification strategy is therefore a sweep: parameterise the slave model's t_v and the interconnect delay, sweep them upward, and find where the master breaks. Compare that breaking point against the inequality from §3. When they agree, your analysis and your environment are consistent and both are probably right. When they disagree, one of them is wrong and finding out which is worth the afternoon.

10. Failure Signature — RTL Passes, Hardware Fails Above a Frequency

Symptom. The design simulates perfectly at every clock rate. On hardware it is reliable below some frequency and intermittently corrupts received data above it, with the error rate rising as the clock rises. Transmitted data is unaffected.

Candidate mechanisms. A round-trip shortfall is the leading hypothesis and fits every element of the description. The competitors are an over-clocked peripheral (Chapter 2.1 — check the datasheet maximum first, it costs nothing), degraded clock integrity from loading or missing termination (Chapter 1.6), and contention on the shared MISO net (Chapter 1.2).

The discriminating observations, in order of cost.

First, which direction fails? Receive-only failure is the return-path signature, from §1's asymmetry. If both directions fail together, suspect clock quality instead — a degraded clock damages every capture on the link, not just one direction.

Second, is the error rate monotonic in frequency? A closing budget produces steadily worsening behaviour above a threshold with no recovery higher up. Contention tracks data patterns rather than rate, and a configuration error is rate-independent and total.

Third, scope MISO at the master's pin against the master's sampling edge. If the line is still settling when the edge arrives, the inequality is violated and you have converted a hypothesis into a measurement.

Why simulation passing is evidence, not an alibi. It tells you the logic and sequencing are right, which usefully eliminates a large class of bugs. It cannot exonerate the timing, because the simulator has no model of the quantity that failed. An engineer who treats a green regression as evidence about board margin has confused two layers that Chapter 2.4's table exists to separate.

11. Common Misconceptions

12. Reason It Through

Work this before reading the answers.

A peripheral's datasheet permits 40 MHz SCLK. A controller is configured for 25 MHz. Transactions still occasionally return the previous bit on MISO. A colleague proposes trying the other SPI mode, reasoning that a one-position offset means a mode problem.

Why is the mode hypothesis inconsistent with "occasionally"? Because a mode misconfiguration is deterministic. A wrong resting level or inverted edge role displaces the sampling instant identically on every transfer at every frequency, so it would corrupt every transaction, not some of them, and it would do so at 1 MHz as readily as at 25. Intermittency rules it out before any instrument is connected — the reproducibility heuristic Chapter 2.2 established.

Why does "25 is well under 40" not settle the question? Because 40 MHz is a property of the device in the vendor's test conditions, and the link's limit is the §3 inequality evaluated on your system. The device's rating contains no term for your board propagation, your master's setup requirement, or the additional capacitance your MISO net presents beyond the datasheet's test load — and Chapter 2.6 showed that last one inflates the largest term in the budget.

Which terms should be examined, and in what order? Start with the cheapest information. Confirm the failure is receive-only, which points at the return path. Confirm it is frequency-dependent by lowering the clock — if it cleans up, the budget is the mechanism. Then obtain the numbers: the peripheral's t_v and its test-load condition, the actual capacitance on your MISO net, your trace lengths, and your master's input setup requirement including where the capture register sits. Evaluate the inequality. Frequently the arithmetic alone reveals that 25 MHz was never achievable.

What is the decisive measurement? Scope MISO at the master's pin, triggered on the master's sampling edge. If MISO is still settling when the edge arrives, the loop does not close and you have measured it rather than inferred it.

What should be done, in order? First check whether the controller offers a delayed sampling point — §6's near-doubling of the window is the cheapest fix available and may resolve it entirely, with the hold side then verified. Next reduce the dominant term: lighten the MISO net, shorten its routing, or move a device off the shared net. Only then lower the clock, and if you do, document the budget so the next person does not raise it back.

What is the trap in the colleague's proposal? Changing the mode may appear to help, because a different edge-role assignment moves the sampling instant and can land it somewhere survivable by accident. The result is a link that is both misconfigured and marginal, working for reasons nobody understands, and certain to fail on the next board revision.

13. Understanding Check

14. Summary

The return path is a loop, and the whole loop must fit inside the interval between the launch edge and the sampling edge — half a period in the standard arrangement. Five delays are spent from it: the master's clock output delay, outbound propagation, the peripheral's output valid time, return propagation, and the master's own setup requirement.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
t_co_m  +  t_pd_sclk  +  t_v  +  t_pd_miso  +  t_su_m   ≤   t_window

The left side is fixed by silicon and copper; the right side is what frequency controls. That is the whole mechanism behind "works slow, fails fast." The board is crossed twice, so distance costs double, and t_v is usually the largest single term — and the one most often under-estimated, because Chapter 2.6 showed the published figure is conditional on a load your board probably exceeds.

Because MOSI is a one-way trip and MISO is a round trip, MISO fails first. Commands accepted while returned data is corrupt is the return-path signature, and it is the most useful single diagnostic in this module.

The master owns two of the five terms, and both improve by putting the capture and clock registers in I/O blocks — shortening the paths and, more importantly, making them constant across builds. Both need input and output delay constraints, without which the paths are unanalysed rather than passing.

A delayed sampling point widens the window toward a full period and can nearly double the achievable frequency for free, at the cost of hold margin that must then be checked. And no simulator can fail this: a zero-delay run collapses the loop into an instant, and STA sees only the on-chip portion of a world you described to it. The closing measurement is an oscilloscope at the master's pin.

15. What Comes Next

That closes Module 2. You now have the complete timing vocabulary: where SCLK comes from and what a period costs to reach (2.1); why edges are named by position rather than direction (2.2); the launch/sample contract and its half-period budget (2.3); the aperture window and what margin means (2.4); chip select as a timed signal (2.5); the peripheral's conditional output delay (2.6); and the loop that ties them together.

Module 3 now spends it. Two configuration bits have been deferred throughout: the level SCLK rests at, and which logical edge samples. Module 3 names them, derives all four combinations, and shows that the four standard SPI modes are not four facts to memorise but a two-bit truth table you can reconstruct from the vocabulary this module built. Everything it needs — leading and trailing, launch and sample, setup and hold, the half-period budget — is now in place.

Browse the path on the SPI curriculum index, or revisit Slave Output Valid Timing for the dominant term in this chapter's budget.

Continue learning