Skip to content
VLSI Mentor

SPI · Module 20

Timing, Constraints, and CDC Considerations

The MISO round trip decides the maximum SCLK rate and static timing analysis never checks it, plus why this architecture has no internal clock-domain crossing and what simulation cannot establish about reset release.

Chapter 20.3 measured everything in cycles. A device on a board does not have cycles; it has nanoseconds. This chapter converts one into the other and finds that the controller's own logic is nowhere near the limit.

For the delays used below, cfg_div = 0 is unusable — and nothing in the RTL, the simulation, or the static timing report says so. The binding constraint is a path that leaves the chip and comes back.

1. Which Signals Are Synchronous To What

Four clocks get confused in SPI discussions. Separating them is the whole of this chapter's first half.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   SYSTEM CLOCK    `clk`. Every register in this controller. 100 MHz in what follows.

   SCLK            An OUTPUT WAVEFORM. Generated by toggling a register on `clk`.
                   Nothing inside this design is clocked by it -- Chapter 20.2,
                   commitment 1.

   THE DEVICE'S CLOCK
                   SCLK, after a board delay. The device's internal timing is
                   referenced to the edges it actually receives, not the ones the
                   controller emitted.

   MISO's TIMING   Referenced to the device's received SCLK, delayed again coming
                   back. Not asynchronous, and not a clock domain -- a computable
                   arrival time. See section 5.

The consequence of commitment 1 is worth stating in timing language: there is exactly one clock in this design, so every path a static timing tool analyses is either clk to clk, clk to a pin, or a pin to clk. There is no clock-to-clock relationship to declare and no synchroniser to review.

2. The Two Paths That Leave The Chip

Every number below is hypothetical and illustrative. They are chosen to be the right order of magnitude for a mid-range FPGA and a general-purpose SPI peripheral; they are not any real part's datasheet figures, and no real device is quoted anywhere in this module.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   f_clk      = 100 MHz            T_clk = 10.0 ns

   t_co       =  3.0 ns            controller register to package pin
   t_pcb      =  0.9 ns            one board trace, each direction
   t_dev_co   = 12.0 ns            device's clock-to-output on MISO   (hypothetical)
   t_dev_su   =  6.0 ns            device's setup requirement on MOSI (hypothetical)
   t_su_in    =  2.2 ns            MISO setup at this design's capture register

The outbound path: MOSI

MOSI is launched on one transition and sampled by the device on the next — one half-period later.

MOSI: controller launch edge to device sampling edge

timing path
A left to right path: controller register, package pin at 3.0 nanoseconds, board trace at 0.9 nanoseconds, device input, and the device's setup requirement of 6.0 nanoseconds. Arrival 9.9 nanoseconds against a required 10.0 nanoseconds at a divider of zero, slack 0.1 nanoseconds.mosi registerlaunchpackage pin3.0 nsboard trace0.9 nsdevice input6.0 ns surequired: t_half = 10.0 ns at cfg_div = 0arrival: 9.9 nsslack +0.1 ns
One half-period is available. The requirement is fixed by the board and the device; only t_half moves.
Figure 1 — the MOSI path. The controller launches MOSI on a transition and the device samples it one half-period later, so the available time is t_half. At a divider of zero that is 10 ns against a 9.9 ns requirement: it passes, with 0.1 ns of margin, which is another way of saying it does not pass.

The return path: MISO, and it is the one that matters

MISO cannot be launched until the device has received an edge, which means the path goes out and back before it is sampled — and the same one half-period has to cover all of it.

MISO: SCLK launch edge to controller capture edge

timing path
A left to right path: SCLK register, package pin at 3.0 nanoseconds, board trace out at 0.9 nanoseconds, device clock to output at 12.0 nanoseconds, board trace back at 0.9 nanoseconds, and a capture setup of 2.2 nanoseconds. Arrival 19.0 nanoseconds against a required 10.0 nanoseconds at a divider of zero, slack minus 9.0 nanoseconds, flagged as a violation.sclk registerlaunchpin + trace3.9 nsdeviceclk-to-q12.0 nstrace back0.9 nscapture setup2.2 nsrequired: t_half = 10.0 ns at cfg_div = 0arrival: 19.0 nsslack -9.0 ns
The requirement is 19.0 ns and is set entirely outside this design. cfg_div is the only term the designer controls.
Figure 2 — the MISO round trip. The controller's SCLK edge has to reach the device, the device has to respond, and the response has to come back and settle, all inside one half-period. At a divider of zero the available time is 10.0 ns against a 19.0 ns requirement: 9 ns short, and no change to the RTL can recover it.

3. The Divider Falls Out Of The Arithmetic

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   required   t_half  >=  t_co + t_pcb + t_dev_co + t_pcb + t_su_in
                      =   3.0 + 0.9 + 12.0 + 0.9 + 2.2
                      =   19.0 ns

   available  t_half  =   (cfg_div + 1) x T_clk
                      =   (cfg_div + 1) x 10.0 ns

   so         cfg_div + 1  >=  1.9      ->      cfg_div >= 1
cfg_divt_halfSCLKMISO slackVerdict
010 ns50 MHz−9.0 nsFails. The controller samples before the data arrives
120 ns25 MHz+1.0 nsLegal, and 1.0 ns is not a margin anyone should ship
230 ns16.7 MHz+11.0 nsComfortable
340 ns12.5 MHz+21.0 nsThe setting Chapter 20.1 §7 costed

Three things are worth taking from this table.

The device dominates. Of the 19.0 ns required, 12.0 ns is the device's own clock-to-output — 63% of the budget, and not a number any decision in this project can change. The controller contributes 5.2 ns and the board 1.8 ns.

The return path is the constraint, not the outbound one. MOSI needs 9.9 ns and MISO needs 19.0 ns, because MISO's path includes the outbound SCLK and the return. Every SPI link is limited by its read path, which is why a write-only device can usually be clocked roughly twice as fast as a read-write one on the same board.

Halving the SCLK rate does not halve the throughput problem. Chapter 20.1 §7 showed that the framing overhead is a fixed 15.8% of a 16-bit frame at any divider, because every term scales with the half-period. So moving from cfg_div = 1 to 3 halves the data rate and changes the efficiency not at all.

cfg_div = 2: MISO settles one cycle before the capture edge

12 cycles
Twelve cycles at ten nanoseconds each. SCLK launches at cycle zero and its next transition is at cycle three. A MISO valid row is invalid for cycles zero to one and valid from cycle two. A marker at cycle one is labelled as where a divider of zero would sample, and a marker at cycle three marks the safe capture point.in flightin flightsettledsettledcapturecapturenext half-periodnext half-periodcfg_div = 0 would sample herecfg_div = 0 would sampleherecfg_div = 2 samples herecfg_div = 2 samples hereclk 10nssclkmiso??DDDD??DDDDcapturet0t1t2t3t4t5t6t7t8t9t10t11
Figure 3 — the same round trip at the cycle level, with one cycle drawn as 10 ns. At cfg_div = 2 the half-period spans three cycles, MISO settles during the third, and the capture edge finds it stable. The marker at cycle 1 is where a cfg_div = 0 design would have sampled: one cycle after the launch edge, while MISO is still in flight.

4. The CS Setup Limitation, Quantified

Chapter 20.1 §8 promised this answer. cfg_lead is counted in half-periods, and a device's chip-select setup requirement is in nanoseconds. So the delivered lead time changes when the divider changes, even though cfg_lead did not.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   t_lead = (cfg_lead + 2) x t_half           measured in Chapter 20.3

   cfg_lead = 2, cfg_div = 3   ->  4 x 40 ns  = 160 ns
   cfg_lead = 2, cfg_div = 1   ->  4 x 20 ns  =  80 ns
   cfg_lead = 2, cfg_div = 0   ->  4 x 10 ns  =  40 ns

A configuration validated against a device needing 100 ns of CS setup at cfg_div = 3 silently violates it at cfg_div = 1, with no change to cfg_lead and no error from anything. That is a genuine limitation of expressing a time requirement in clock periods, it was recorded as such in Chapter 20.1, and the mitigation is a rule rather than a gate: whoever changes the divider must recheck the lead.

5. What Is And Is Not A Clock-Domain Crossing

This section exists because the temptation to put a synchroniser on MISO is strong and it would be a mistake.

BoundaryIs it a CDC?What it actually isCorrect treatment
clk to internal registersNoone domainnothing
clk to SCLK, MOSI, cs_nNoregistered outputsset_output_delay
MISO to the capture registerNoa computable arrival timeset_input_delay, and a divider that satisfies §3
rst_n release to every registerYes, in effectasynchronous control, assumed pre-synchroniseda reset synchroniser, outside this block
cfg_* from softwareNosame domain, and captured at acceptancenothing

Why MISO is not a CDC

A clock-domain crossing is two clocks with no fixed phase relationship, where a receiving flip-flop can be clocked arbitrarily close to a transition and go metastable. The mitigation is a synchroniser, and the cost is latency.

MISO's timing is derived from a clock this controller generated. Its arrival relative to the sampling edge is not unknown — it is the 19.0 ns computed in §2. That makes it a static timing problem with a setup requirement, and the mitigation is a slower divider.

Putting a two-flop synchroniser on MISO would be actively harmful. It would add two cycles of latency per bit and fix nothing: a synchroniser samples exactly as late as the register it replaces, and lateness rather than metastability is the failure here.

The one real asynchronous boundary

rst_n is asynchronous and this design does not synchronise it. Chapter 20.1 records that as an assumption on the integrator, and it is the one thing in this chapter that matters most and can be verified least.

6. Constraints, And What They Do Not Check

The timing model first

Before any syntax: in this architecture, SCLK is a data output. That single fact decides the whole constraint strategy, because a timing tool constrains data outputs against the clock that launched them — and every pin here is launched by clk.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   sclk, mosi, cs_n[3:0]     clk -> pin        set_output_delay
   miso                      pin -> clk        set_input_delay

There is no create_generated_clock for SCLK, because nothing is clocked by it. A design that did clock its shift register from a divided SCLK would need one, plus a declared relationship between the two clocks, plus a review of every path that crossed between them. That is the constraint-side cost of the architecture Chapter 20.2 did not choose.

An illustrative example, not a sign-off set

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   # ---- educational example. NOT a sign-off constraint set. ----
   create_clock -name clk -period 10.000 [get_ports clk]

   # Outputs. The numbers are the board delay plus the device's requirement,
   # expressed as how much of the period they consume.
   set_output_delay -clock clk -max 6.9 [get_ports {sclk mosi cs_n[*]}]
   set_output_delay -clock clk -min 0.5 [get_ports {sclk mosi cs_n[*]}]

   # Input. MISO's arrival, relative to the clock that captures it.
   set_input_delay  -clock clk -max 16.8 [get_ports miso]
   set_input_delay  -clock clk -min  3.0 [get_ports miso]

   # The reset release is asynchronous; check recovery and removal, not setup.
   set_false_path -from [get_ports rst_n]

Four things are deliberately missing, and a sign-off set needs all of them: clock uncertainty, the I/O standard and drive strength that make t_co meaningful, per-pin board delays rather than one number for all, and a documented derivation of every figure from a real datasheet.

7. Why There Is No New RTL In This Chapter

Intentional absence — standalone synthesizable RTL would not honestly represent this concept.

Nothing in this chapter changes the hardware. The controller of Chapter 20.3 already has one clock, registered outputs, and a divider whose legal range this chapter has now bounded from below. The chapter's product is a number — cfg_div ≥ 1 for the delays given — plus a constraint strategy and a list of three things that must be true outside this block.

Writing RTL to accompany it would mean inventing something: an I/O register that the tool already infers, or a synchroniser for MISO that §5 argues against. Both would make the chapter look more substantial and would be wrong.

The honest deliverable of a timing chapter is arithmetic, a constraint file, and an explicit statement of what no tool in the flow will check for you.

8. Summary

The controller's logic was never the limit. The MISO round trip — out through a pin and a trace, through the device's clock-to-output, back through a trace and into a capture register — needs 19.0 ns for the illustrative delays used here, of which the device alone contributes 12.0 ns. One half-period has to cover all of it, so cfg_div ≥ 1, and cfg_div = 0 is unusable at 50 MHz SCLK despite passing every simulation and every static timing check.

MOSI needs only 9.9 ns, because its path is one-way. That asymmetry is why a write-only device on the same board can be clocked about twice as fast, and why a link that will not run at its headline rate should have its read path examined first.

Expressing cfg_lead in half-periods means the delivered CS setup time changes by 4× between cfg_div = 3 and cfg_div = 0 with no change to the configuration field — a real limitation, recorded rather than hidden, whose mitigation is a rule for whoever changes the divider.

There is no internal clock-domain crossing, and MISO is not one either: its arrival time is computable, so it is a static timing problem and a synchroniser on it would add latency and fix nothing. The one genuine asynchronous boundary is the reset release, which this block assumes is already synchronised — and that assumption is the thing in the whole module that simulation can least establish, because a zero-delay simulator releases every flip-flop together on every run and will show a clean start-up no matter how badly it is violated.

The sentence to carry out of the chapter: static timing analysis never checks the SPI link. The round-trip budget is a calculation the designer owns, and its output is a minimum divider that belongs in the driver and in the bring-up document.

9. What Comes Next

Chapter 20.5 builds the checking layer: a reference model that predicts from the specification rather than from the RTL, a scoreboard that compares transactions, and eight properties — three of which, as first written, fired on a controller that was behaving perfectly.

Continue learning