Skip to content
VLSI Mentor

SPI · Module 20

Capstone Requirements and Specification

One configurable SPI controller, specified before it is designed: nineteen numbered requirements, their corner cases, and the non-goals that keep the project finishable.

Nineteen modules of SPI end here, in one deliverable: a configurable controller carried from a written specification to a design review. This chapter writes the specification. No RTL appears in it, and that is the point.

Verification cannot be stronger than the specification it checks. Every ambiguity left in this chapter becomes a bug that no testbench is able to call a bug.

1. Why The Specification Comes First, With A Concrete Reason

The usual argument for specifying first is discipline. There is a better one, and it is mechanical.

A testbench compares what happened against what should have happened. The second half has to come from somewhere. If it comes from the design, the comparison is a tautology — the design agrees with itself. If it comes from a sentence like "the controller shall transfer data correctly", there is nothing to compute. The only thing a checker can actually use is a statement precise enough to evaluate:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   "correctly"                              → a checker cannot use this
   "rx_data holds the received word,        → a checker can compute this,
    right-aligned, valid from `done`"          and can disagree with the design

Every requirement below is written so that a checker can compute it. Chapter 20.5 builds those checkers, and each one names the requirement it protects. Chapter 20.7 deliberately breaks the design and confirms that the requirement's checker fires.

There is also a negative version of this chapter's job, which is stated in §9: the things the controller is not required to do. A project without non-goals does not finish.

2. The Use Case, And Where The Specification Stops

One FPGA or SoC block drives up to four SPI devices that do not agree with one another. They have different modes, different word widths, different maximum clock rates, and different setup requirements. Software — or a state machine acting for it — configures the controller per transfer and posts a request.

Left to right: software or logic writes configuration and posts a request to spi_capstone_ctrl, which drives four pins, which reach one of four devices. Below, two amber blocks marked as assumptions point inward: a reset already synchronised to the system clock, and a board delay budget.software or logicwrites configuration,posts a requestspi_capstone_ctrlTHIS specificationfour pinssclk, mosi, miso,cs_n[3:0]one of fourdevicesits own mode and timingreset, alreadysyncedassumed, not providedhereboard delay budgetassumed, quantified in20.4cfg + startdrivenwiresassumptionassumption12
Figure 1 — the boundary of this specification. The controller in the centre is what Chapters 20.2 and 20.3 build; the two amber blocks are assumptions this document makes and does not deliver. A specification that does not say what it assumes is not smaller, it is only vaguer — and both assumptions below reappear as real engineering work in Chapter 20.4.

3. The External Interfaces

Three groups of signals, and the grouping is itself a design statement: configuration is separate from the request, and both are separate from the pins.

Configuration

SignalWidthMeaning
cfg_cpol1SCLK level while no transfer is active
cfg_cpha1Which edge kind launches and which samples
cfg_lsb_first10 = bit width-1 first, 1 = bit 0 first
cfg_width5Bits per transfer, 4 to 16 inclusive
cfg_div8SCLK half-period = cfg_div + 1 system clocks
cfg_dev2Which of the four chip selects to assert
cfg_lead4Minimum CS-low to first SCLK edge, in half-periods
cfg_lag4Minimum last SCLK edge to CS-high, in half-periods
cfg_idle4Minimum enforced turnaround between transfers

Lead, lag and turnaround are counted in half-periods, not system clocks, and that is deliberate: a device's requirements scale with the clock it is being given, so expressing them in half-periods means one configuration works at every divider. A reviewer should ask whether that is true of real devices — §8 answers it.

Request and response

SignalDirMeaning
startinOne-cycle pulse. Ignored unless the controller is idle
tx_datainThe word to send, right-aligned in 16 bits
abortinOne-cycle pulse. Abandons a transfer in flight
busyoutHigh from acceptance until the enforced turnaround has elapsed
doneoutOne cycle, exactly once per completed transfer
cfg_erroutOne cycle. The request was refused, nothing happened
rx_dataoutThe received word, right-aligned, valid from done
bits_doneoutBits actually sampled. Survives an abort

Pins

PinDirMeaning
sclkoutGenerated clock. Not a clock this design uses — see 20.2
mosioutController to device
misoinDevice to controller
cs_n[3:0]outActive-low selects. At most one low at any time

4. The Configuration Model, And The One Decision That Defines It

cfg_* are live wires. Software may change them at any moment, including halfway through a transfer. So the specification has to answer one question before anything else can be verified:

Does a transfer use the configuration that was present when it was accepted, or whatever the configuration happens to be at each instant?

This specification chooses the first, and states it as a requirement rather than leaving it to the implementation:

REQ-CFG-001 — The configuration in force for a transfer is the value of every cfg_* input at the clock edge on which the request is accepted. Changes after that edge affect the next accepted transfer and have no effect on the one in flight.

5. The Four Modes, Derived Rather Than Tabulated

SPI modes are usually given as a four-row table to memorise. They are two independent one-bit decisions, and writing them out as such is what makes the specification checkable.

Every SCLK transition in a frame is one of two kinds:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   LEADING    the transition AWAY from the idle level
   TRAILING   the transition BACK to the idle level

cfg_cpol decides what "the idle level" is, and therefore which physical direction is leading. cfg_cpha decides which kind of transition the data moves on. Those are the whole of it:

cfg_cpolcfg_cphaModeIdle levelLaunch onSample on
000lowtrailing (1→0)leading (0→1)
011lowleading (0→1)trailing (1→0)
102hightrailing (0→1)leading (1→0)
113highleading (1→0)trailing (0→1)

The table is a consequence, not a definition. The definition is:

REQ-MODE-002 — With cfg_cpha = 0, the controller samples MISO on every leading transition and changes MOSI on every trailing transition. Because the first sampling edge arrives before any trailing transition exists, the first bit is presented on MOSI at the same edge the chip select is asserted.

REQ-MODE-003 — With cfg_cpha = 1, the controller changes MOSI on every leading transition and samples MISO on every trailing transition. MOSI holds a defined low level from the assertion of the chip select until the first leading transition.

Note what REQ-MODE-002's second sentence is doing. It is not describing an implementation choice — it is stating the only thing that can happen. A device that samples on the first leading edge must be given a valid bit before that edge, and the only earlier event in the frame is the chip select. Specifying it is what lets Chapter 20.5 assert that MOSI changes when CS asserts and call that legal rather than a glitch.

6. The Requirements

Nineteen, grouped. Each carries a rationale, the behaviour that makes it observable, its corner cases, and how it is verified — because a requirement with no stated means of observation is a wish.

Functional

IDRequirement
REQ-FUNC-001The controller supports all four SPI modes, selected by cfg_cpol and cfg_cpha
REQ-FUNC-002Bit order is selectable: MSB-first or LSB-first, independently of mode
REQ-FUNC-003Transfer width is selectable per request, 4 to 16 bits inclusive
REQ-FUNC-004Exactly one chip select is asserted for the duration of a transfer, and never more than one at any instant
REQ-FUNC-005rx_data holds the received word right-aligned, and is valid from the cycle done is asserted
REQ-FUNC-006done is asserted for exactly one cycle, exactly once, per completed transfer
REQ-FUNC-007A start pulse while busy is asserted has no effect of any kind

Timing

IDRequirement
REQ-TIM-001The SCLK half-period is exactly cfg_div + 1 system clock cycles
REQ-TIM-002The chip select is low for at least cfg_lead half-periods before the first SCLK transition
REQ-TIM-003The chip select remains low for at least cfg_lag half-periods after the last SCLK transition
REQ-TIM-004At least cfg_idle half-periods elapse between the release of one chip select and the assertion of the next

Mode and level

IDRequirement
REQ-MODE-001Outside an active frame, SCLK sits at the level given by cfg_cpol
REQ-MODE-002cfg_cpha = 0: sample on leading transitions, launch on trailing, first bit presented with the chip select
REQ-MODE-003cfg_cpha = 1: launch on leading transitions, sample on trailing, MOSI low until the first launch

Reset, error, abort

IDRequirement
REQ-RST-001While reset is asserted: every chip select high, SCLK at a defined low level, busy and done low, all counters zero
REQ-RST-002Reset during a frame abandons it immediately. The chip select is released and done is not asserted
REQ-ERR-001A request whose cfg_width lies outside 4 to 16 is refused: cfg_err for one cycle, no chip select, no done, nothing captured
REQ-ABT-001abort during a frame stops it. done is not asserted, rx_data is unchanged, bits_done holds the count reached, and the chip select still receives its full lag

Configuration capture

IDRequirement
REQ-CFG-001Configuration is captured at acceptance; later changes affect only the next transfer

7. Three Transactions, With Numbers

A specification that has never been evaluated on a concrete case is a draft. Take a 100 MHz system clock, so T_clk = 10 ns throughout.

A 16-bit mode-0 transfer at cfg_div = 3

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   t_half   = (cfg_div + 1) x T_clk = 4 x 10 ns          = 40 ns
   f_sclk   = 1 / (2 x t_half) = 1 / 80 ns               = 12.5 MHz
   frame    = (lead + lag + 2N + 2) half-periods
            = (2 + 2 + 32 + 2) = 38 half-periods
            = 38 x 40 ns                                 = 1520 ns
   payload  = 16 bits / 1520 ns                          = 10.5 Mbit/s
   overhead = 6 of 38 half-periods are not clocking data  = 15.8%

The +2 in the frame expression is the pair of half-periods the state machine spends entering and leaving its clocking phase. It is margin above REQ-TIM-002 and REQ-TIM-003, not part of them — the requirements say at least, and Chapter 20.3 measures what the implementation actually delivers rather than assuming the two coincide.

The same transfer at cfg_div = 0

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   t_half   = 1 x 10 ns                                  = 10 ns
   f_sclk                                                = 50 MHz
   frame    = 38 x 10 ns                                 = 380 ns
   payload  = 16 bits / 380 ns                           = 42.1 Mbit/s

Four times the divider gives four times the rate, exactly, because every term in the frame expression is a multiple of the half-period. That is a property worth noticing: it means the overhead fraction is independent of the divider, so a designer cannot improve efficiency by slowing down, and cannot lose it by speeding up. The only way to change 15.8% is to change the lead, lag, turnaround, or width.

A 4-bit transfer at cfg_div = 0

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   frame    = (2 + 2 + 8 + 2) = 14 half-periods x 10 ns  = 140 ns
   payload  = 4 bits / 140 ns                            = 28.6 Mbit/s
   overhead = 6 of 14 half-periods                       = 42.9%

Short transfers are dominated by their own framing. A status register polled as four separate 4-bit reads moves the same 16 bits in 560 ns against 380 ns for one 16-bit transfer — 47% slower, with no change to any clock rate. All the numbers in this section are derived from the requirements alone, before any hardware exists, which is what makes them useful for choosing a width in the first place.

8. The Corner Cases That Make The Requirements Testable

A requirement is only as good as the edge of its range. These are the cases Chapter 20.3's directed suite is built from, and each one exists because the requirement above it has a boundary.

CornerRequirement under testWhat would fail without it
cfg_width = 4REQ-FUNC-003The narrow end of the alignment shift
cfg_width = 16REQ-FUNC-003A mask computed as (1 << w) - 1 in 16 bits yields zero
cfg_width = 3 and 17REQ-ERR-001Both sides of the legal range, refused
cfg_div = 0REQ-TIM-001A one-cycle half-period, where the divider never counts
cfg_lead/cfg_lag = 0REQ-TIM-002/003A phase counter that must terminate at zero
cfg_lead/cfg_lag = 15REQ-TIM-002/003The top of a 4-bit phase counter
All four cfg_devREQ-FUNC-004Every decoder output, including the last
Reset during a frameREQ-RST-002A frame that completes after its reset
Abort at the first bit and near the lastREQ-ABT-001A partial bits_done, and a stale rx_data
start while busyREQ-FUNC-007A second frame launched on top of a live one
Configuration rewritten mid-frameREQ-CFG-001The shadow that makes it harmless
A palindromic data patternREQ-FUNC-002Covered in §10 — it is a trap, not a corner

The one about lead and lag scaling with the divider deserves the answer promised in §3. Real devices specify CS setup and hold in nanoseconds, not in clock periods. Expressing them in half-periods means that at a slow divider the controller delivers far more than the device needs, and at the fastest divider it may deliver less. That is a real limitation of this interface, it is recorded as such, and Chapter 20.4 computes the divider below which the configuration stops being safe.

9. Non-Goals, Stated So The Project Ends

Everything here is a thing a commercial SPI block does and this one deliberately does not. Each line is a decision, and the reason is the same in every case: a feature costs RTL state, verification state space, coverage obligation, and debug burden — four costs, paid on every future change.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   no request queue          one transfer in flight; `start` while busy is ignored
   no DMA, no FIFO           the caller presents one word and collects one word
   no dual / quad / octal    one data pin each way
   no slave mode             this is a controller
   no 3-wire / bidirectional MOSI and MISO are separate pins
   no interrupt controller   `done` is a pulse; whoever wants an interrupt makes one
   no register bus           Chapter 19.3 already wraps a controller in APB
   no automatic CS toggling  one transfer, one assertion; bursts are the caller's job
   no clock gating           power is out of scope for this module

The absence of a queue has a consequence that Chapter 20.3 measures and Chapter 20.7 puts to the review: because the next request cannot be posted until busy falls, the actual gap between two frames is the enforced turnaround plus the requester's own reaction time. REQ-TIM-004 is a floor, not a promise of back-to-back throughput. A queue would fix that, and would add a second configuration shadow, a second set of coverage bins, and a new class of ordering bug.

10. The Verification Plan, Previewed

Each requirement is claimed only when something can be pointed at. The full matrix is closed in Chapter 20.7; this is the shape of it.

Requirement groupDirectedReference modelPropertyCoverageBug injected
REQ-FUNC-001/002/003all 4 modes x 2 orders x 4 widthspredicted rx_data—mode, order, width bins and crosseswrong sample edge; bit order inverted
REQ-FUNC-004all four devicesobserved select indexat most one lowdevice binwrong select asserted
REQ-FUNC-005/006every transferpredicted worddone implies a complete word—extra edges per frame
REQ-FUNC-007request while busyone frame, not twoa select implies busy——
REQ-TIM-001..004measured at the pinspredicted frame duration—divider binslag removed
REQ-MODE-001parked level after each frame—SCLK parked, and quiet—wrong idle polarity
REQ-MODE-002/003modes 0-3predicted MOSI streamMOSI changes only on an edgemode x ordershift-then-present
REQ-RST-001/002reset in idle and mid-frame————
REQ-ERR-001widths 3, 4, 16, 17——illegal bin, asserted emptyillegal width accepted
REQ-ABT-001abort mid-framerx_data unchanged——done on abort
REQ-CFG-001rewrite mid-framepredicted from captured values——live width in the edge count

11. Why There Is No RTL In This Chapter

Intentional absence — standalone synthesizable RTL would not honestly represent this concept.

There is nothing here to implement. A specification is a set of claims about observable behaviour, and any RTL written to accompany it would be either a fragment too small to be the controller, or the controller itself — which is Chapter 20.3's subject and which would arrive here without the architecture Chapter 20.2 has to choose first.

The temptation is to include a skeleton — a port list, an empty state machine — and it should be resisted, because a skeleton makes architectural commitments silently. Where the state boundary falls, whether configuration is registered, whether SCLK is generated or gated: a port list implies answers to all three, and Chapter 20.2 exists to argue about them in the open.

What this chapter produces instead is the thing the next five chapters consume: nineteen numbered, observable, corner-cased requirements, and a list of what the controller will not do.

12. Summary

This chapter produced a contract: nineteen numbered requirements over three groups of signals, with corner cases at every boundary and a stated list of what the controller will not do.

Four of the decisions will be argued about for the rest of the module. Configuration is captured at acceptance, so a mid-frame rewrite is harmless and software cannot change its mind. The four modes are derived from two independent bits rather than tabulated, which is what makes them checkable one at a time. Lead, lag and turnaround are counted in half-periods, which makes one configuration work at every divider and creates a real limitation at the fastest one. And reset puts SCLK low rather than at the configured idle level, because a reset state that reads configuration depends on a value that may not exist yet.

The timing arithmetic is already useful before any hardware exists: at 100 MHz with a divider of 3, a 16-bit frame takes 1520 ns and delivers 10.5 Mbit/s, with 15.8% of the frame spent not clocking data — and that overhead fraction does not change with the divider, so it cannot be tuned away by choosing a clock rate. Only the width, lead, lag and turnaround move it.

The one trap already visible is in the test data rather than the design: a5 and 5a are both eight-bit palindromes, so neither can tell MSB-first from LSB-first. Every pattern from here on is asymmetric under reversal.

13. What Comes Next

Chapter 20.2 turns these requirements into hardware, and the argument is not about what the blocks are — a divider, a shift register and a counter are not in doubt. It is about where the boundaries between them fall, whether SCLK is a clock this design uses or a waveform it emits, and which of those choices the final review will be able to defend.

Continue learning