SPI · Module 20
Timing, Constraints, and CDC Considerations
The MISO round trip decides the maximum SCLK rate and static timing analysis never checks it, plus why this architecture has no internal clock-domain crossing and what simulation cannot establish about reset release.
Chapter 20.3 measured everything in cycles. A device on a board does not have cycles; it has nanoseconds. This chapter converts one into the other and finds that the controller's own logic is nowhere near the limit.
For the delays used below,
cfg_div= 0 is unusable — and nothing in the RTL, the simulation, or the static timing report says so. The binding constraint is a path that leaves the chip and comes back.
1. Which Signals Are Synchronous To What
Four clocks get confused in SPI discussions. Separating them is the whole of this chapter's first half.
SYSTEM CLOCK `clk`. Every register in this controller. 100 MHz in what follows.
SCLK An OUTPUT WAVEFORM. Generated by toggling a register on `clk`.
Nothing inside this design is clocked by it -- Chapter 20.2,
commitment 1.
THE DEVICE'S CLOCK
SCLK, after a board delay. The device's internal timing is
referenced to the edges it actually receives, not the ones the
controller emitted.
MISO's TIMING Referenced to the device's received SCLK, delayed again coming
back. Not asynchronous, and not a clock domain -- a computable
arrival time. See section 5.The consequence of commitment 1 is worth stating in timing language: there is exactly one clock in this design, so every path a static timing tool analyses is either clk to clk, clk to a pin, or a pin to clk. There is no clock-to-clock relationship to declare and no synchroniser to review.
2. The Two Paths That Leave The Chip
Every number below is hypothetical and illustrative. They are chosen to be the right order of magnitude for a mid-range FPGA and a general-purpose SPI peripheral; they are not any real part's datasheet figures, and no real device is quoted anywhere in this module.
f_clk = 100 MHz T_clk = 10.0 ns
t_co = 3.0 ns controller register to package pin
t_pcb = 0.9 ns one board trace, each direction
t_dev_co = 12.0 ns device's clock-to-output on MISO (hypothetical)
t_dev_su = 6.0 ns device's setup requirement on MOSI (hypothetical)
t_su_in = 2.2 ns MISO setup at this design's capture registerThe outbound path: MOSI
MOSI is launched on one transition and sampled by the device on the next — one half-period later.
MOSI: controller launch edge to device sampling edge
timing pathThe return path: MISO, and it is the one that matters
MISO cannot be launched until the device has received an edge, which means the path goes out and back before it is sampled — and the same one half-period has to cover all of it.
MISO: SCLK launch edge to controller capture edge
timing path3. The Divider Falls Out Of The Arithmetic
required t_half >= t_co + t_pcb + t_dev_co + t_pcb + t_su_in
= 3.0 + 0.9 + 12.0 + 0.9 + 2.2
= 19.0 ns
available t_half = (cfg_div + 1) x T_clk
= (cfg_div + 1) x 10.0 ns
so cfg_div + 1 >= 1.9 -> cfg_div >= 1cfg_div | t_half | SCLK | MISO slack | Verdict |
|---|---|---|---|---|
| 0 | 10 ns | 50 MHz | −9.0 ns | Fails. The controller samples before the data arrives |
| 1 | 20 ns | 25 MHz | +1.0 ns | Legal, and 1.0 ns is not a margin anyone should ship |
| 2 | 30 ns | 16.7 MHz | +11.0 ns | Comfortable |
| 3 | 40 ns | 12.5 MHz | +21.0 ns | The setting Chapter 20.1 §7 costed |
Three things are worth taking from this table.
The device dominates. Of the 19.0 ns required, 12.0 ns is the device's own clock-to-output — 63% of the budget, and not a number any decision in this project can change. The controller contributes 5.2 ns and the board 1.8 ns.
The return path is the constraint, not the outbound one. MOSI needs 9.9 ns and MISO needs 19.0 ns, because MISO's path includes the outbound SCLK and the return. Every SPI link is limited by its read path, which is why a write-only device can usually be clocked roughly twice as fast as a read-write one on the same board.
Halving the SCLK rate does not halve the throughput problem. Chapter 20.1 §7 showed that the framing overhead is a fixed 15.8% of a 16-bit frame at any divider, because every term scales with the half-period. So moving from cfg_div = 1 to 3 halves the data rate and changes the efficiency not at all.
cfg_div = 2: MISO settles one cycle before the capture edge
12 cycles4. The CS Setup Limitation, Quantified
Chapter 20.1 §8 promised this answer. cfg_lead is counted in half-periods, and a device's chip-select setup requirement is in nanoseconds. So the delivered lead time changes when the divider changes, even though cfg_lead did not.
t_lead = (cfg_lead + 2) x t_half measured in Chapter 20.3
cfg_lead = 2, cfg_div = 3 -> 4 x 40 ns = 160 ns
cfg_lead = 2, cfg_div = 1 -> 4 x 20 ns = 80 ns
cfg_lead = 2, cfg_div = 0 -> 4 x 10 ns = 40 nsA configuration validated against a device needing 100 ns of CS setup at cfg_div = 3 silently violates it at cfg_div = 1, with no change to cfg_lead and no error from anything. That is a genuine limitation of expressing a time requirement in clock periods, it was recorded as such in Chapter 20.1, and the mitigation is a rule rather than a gate: whoever changes the divider must recheck the lead.
5. What Is And Is Not A Clock-Domain Crossing
This section exists because the temptation to put a synchroniser on MISO is strong and it would be a mistake.
| Boundary | Is it a CDC? | What it actually is | Correct treatment |
|---|---|---|---|
clk to internal registers | No | one domain | nothing |
clk to SCLK, MOSI, cs_n | No | registered outputs | set_output_delay |
| MISO to the capture register | No | a computable arrival time | set_input_delay, and a divider that satisfies §3 |
rst_n release to every register | Yes, in effect | asynchronous control, assumed pre-synchronised | a reset synchroniser, outside this block |
cfg_* from software | No | same domain, and captured at acceptance | nothing |
Why MISO is not a CDC
A clock-domain crossing is two clocks with no fixed phase relationship, where a receiving flip-flop can be clocked arbitrarily close to a transition and go metastable. The mitigation is a synchroniser, and the cost is latency.
MISO's timing is derived from a clock this controller generated. Its arrival relative to the sampling edge is not unknown — it is the 19.0 ns computed in §2. That makes it a static timing problem with a setup requirement, and the mitigation is a slower divider.
Putting a two-flop synchroniser on MISO would be actively harmful. It would add two cycles of latency per bit and fix nothing: a synchroniser samples exactly as late as the register it replaces, and lateness rather than metastability is the failure here.
The one real asynchronous boundary
rst_n is asynchronous and this design does not synchronise it. Chapter 20.1 records that as an assumption on the integrator, and it is the one thing in this chapter that matters most and can be verified least.
6. Constraints, And What They Do Not Check
The timing model first
Before any syntax: in this architecture, SCLK is a data output. That single fact decides the whole constraint strategy, because a timing tool constrains data outputs against the clock that launched them — and every pin here is launched by clk.
sclk, mosi, cs_n[3:0] clk -> pin set_output_delay
miso pin -> clk set_input_delayThere is no create_generated_clock for SCLK, because nothing is clocked by it. A design that did clock its shift register from a divided SCLK would need one, plus a declared relationship between the two clocks, plus a review of every path that crossed between them. That is the constraint-side cost of the architecture Chapter 20.2 did not choose.
An illustrative example, not a sign-off set
# ---- educational example. NOT a sign-off constraint set. ----
create_clock -name clk -period 10.000 [get_ports clk]
# Outputs. The numbers are the board delay plus the device's requirement,
# expressed as how much of the period they consume.
set_output_delay -clock clk -max 6.9 [get_ports {sclk mosi cs_n[*]}]
set_output_delay -clock clk -min 0.5 [get_ports {sclk mosi cs_n[*]}]
# Input. MISO's arrival, relative to the clock that captures it.
set_input_delay -clock clk -max 16.8 [get_ports miso]
set_input_delay -clock clk -min 3.0 [get_ports miso]
# The reset release is asynchronous; check recovery and removal, not setup.
set_false_path -from [get_ports rst_n]Four things are deliberately missing, and a sign-off set needs all of them: clock uncertainty, the I/O standard and drive strength that make t_co meaningful, per-pin board delays rather than one number for all, and a documented derivation of every figure from a real datasheet.
7. Why There Is No New RTL In This Chapter
Intentional absence — standalone synthesizable RTL would not honestly represent this concept.
Nothing in this chapter changes the hardware. The controller of Chapter 20.3 already has one clock, registered outputs, and a divider whose legal range this chapter has now bounded from below. The chapter's product is a number — cfg_div ≥ 1 for the delays given — plus a constraint strategy and a list of three things that must be true outside this block.
Writing RTL to accompany it would mean inventing something: an I/O register that the tool already infers, or a synchroniser for MISO that §5 argues against. Both would make the chapter look more substantial and would be wrong.
The honest deliverable of a timing chapter is arithmetic, a constraint file, and an explicit statement of what no tool in the flow will check for you.
8. Summary
The controller's logic was never the limit. The MISO round trip — out through a pin and a trace, through the device's clock-to-output, back through a trace and into a capture register — needs 19.0 ns for the illustrative delays used here, of which the device alone contributes 12.0 ns. One half-period has to cover all of it, so cfg_div ≥ 1, and cfg_div = 0 is unusable at 50 MHz SCLK despite passing every simulation and every static timing check.
MOSI needs only 9.9 ns, because its path is one-way. That asymmetry is why a write-only device on the same board can be clocked about twice as fast, and why a link that will not run at its headline rate should have its read path examined first.
Expressing cfg_lead in half-periods means the delivered CS setup time changes by 4× between cfg_div = 3 and cfg_div = 0 with no change to the configuration field — a real limitation, recorded rather than hidden, whose mitigation is a rule for whoever changes the divider.
There is no internal clock-domain crossing, and MISO is not one either: its arrival time is computable, so it is a static timing problem and a synchroniser on it would add latency and fix nothing. The one genuine asynchronous boundary is the reset release, which this block assumes is already synchronised — and that assumption is the thing in the whole module that simulation can least establish, because a zero-delay simulator releases every flip-flop together on every run and will show a clean start-up no matter how badly it is violated.
The sentence to carry out of the chapter: static timing analysis never checks the SPI link. The round-trip budget is a calculation the designer owns, and its output is a minimum divider that belongs in the driver and in the bring-up document.
9. What Comes Next
Chapter 20.5 builds the checking layer: a reference model that predicts from the specification rather than from the RTL, a scoreboard that compares transactions, and eight properties — three of which, as first written, fired on a controller that was behaving perfectly.
Continue learning
Related tutorials
- Related topic
Input and Output Delay Constraints
A constraint cannot be simulated — set_input_delay describes a board. So the RTL is one register per pin and no logic, the bench checks only what the constraints assume, and the file is written three times for three tools. Adding that pad register also moves three numbers Module 14 published.
- Related topic
MISO Valid Timing
When returned data may be trusted relative to the launch edge: the round trip separating arrival from validity, why the sampling window has two edges, the configurable input sampler in three HDLs, and the margin no delay can create.
- Related topic
RTL, CDC, Constraint, or Testbench?
Three perturbations — the clock ratio, the bench's data placement, and the payload — map four failure layers to four signatures. And one of the four is provably invisible to any RTL regression.
- Related topic
Electrical, Timing and Bring-Up Readiness Review
The audit before power is applied for the first time: sorting every claim into established, calculable, deferred to the board, or unsettleable here. Computes the observation-latency budget against tVD;ACK and finds where it closes — 11.1 MHz for Fast-mode, 22.2 MHz for Fast-mode Plus — and gives a bring-up order that makes each failure diagnostic.
