SPI · Module 20
Microarchitecture and Design Decisions
Turning nineteen requirements into five blocks, and defending four architectural commitments the design review will attack: a generated clock, a captured configuration, one shift direction, and timing split from control.
Chapter 20.1 produced nineteen requirements. This chapter turns them into hardware, and the interesting part is not which blocks exist — a divider, a shift register and a counter were never in doubt — but where the boundaries between them fall and which choices survive being questioned.
Four commitments are made here. Each has a named alternative, a stated cost, and a place in Chapter 20.7 where it is attacked.
1. The Blocks
| Block | Responsibility | State it owns | Reads | Drives |
|---|---|---|---|---|
| request accept | Decide whether a start becomes a transfer | none | start, busy, width legality | the capture pulse |
| config capture | Freeze the configuration for one frame | 9 shadow fields | live cfg_* | every other block |
| clock divider | Produce one tick per SCLK half-period | down-counter | captured cfg_div | tick |
| control FSM | Sequence the five phases of a frame | state, phase counter, edge index | tick, abort | SCLK level, phase |
| mode decoder | Classify each transition as launch or sample | none — pure function of the edge index | edge index, captured CPHA | launch, sample |
| shift datapath | Move bits, one direction, and align at the boundaries | tx_sr, rx_sr, two indices | launch, sample, MISO | MOSI, rx_data |
| select decoder | One index to one active-low output | none | captured device index | cs_n[3:0] |
Every block above is clocked by clk. Not one of them is clocked by SCLK, and that is the first commitment.
2. Commitment 1 — SCLK Is An Output Waveform, Not A Clock
The controller generates SCLK by toggling a register. Nothing inside the design is clocked by it.
The alternative is real and is used in practice: divide the system clock down, treat the result as a clock, and clock the shift registers with it. That gives a shift register running at exactly the data rate, with no tick to distribute and no divider in the datapath's way.
GENERATED CLOCK, USED INTERNALLY CLOCK ENABLE, SCLK AS OUTPUT (chosen)
clk ──► divider ──► sclk_int ──┐ clk ──┬──► everything
│ │
shift reg ◄┘ └──► divider ──► tick ──► enables
│ │
└──► pins shift reg ◄┘
│
sclk_r ──┴──► pinsWhat the chosen architecture buys:
There is no internal clock-domain crossing. Two clocks, even when one is derived from the other, need the synthesis and timing tools told how they relate; get it wrong and the tool either reports failures that do not exist or misses ones that do. With one clock there is nothing to relate.
The divider is not on a timing path to a pin. SCLK comes out of a flip-flop, so its output delay is one register-to-pin path, not a path through a counter's comparator.
Every state is visible on one clock. A waveform, a simulation, and a scan chain all see the whole design on clk. With an internal generated clock, the shift register's state advances at a rate the rest of the design cannot observe directly.
What it costs, stated plainly:
Every register toggles at the system clock rate whether a transfer is active or not — the divider counts, the FSM is evaluated, the pins are re-driven. A clock-gated or internally-clocked design switches far less. For a low-power target that is the decisive argument the other way, and this module does not pretend otherwise: power is a stated non-goal in Chapter 20.1 §9.
The datapath is paced by an enable, not by a clock, so every sequential element needs tick routed to it, and every piece of logic has to remember that a cycle without tick is a cycle where nothing moves. That is a whole class of bug — an update that forgets its enable — and Chapter 20.7 injects one.
3. Commitment 2 — Configuration Is Captured, Not Read Live
REQ-CFG-001 requires it, so the question here is only where the shadow lives and what it costs.
It lives in nine registers written by one pulse. The alternative — reading cfg_* live wherever needed — costs no registers at all and is the reason it is tempting.
LIVE CAPTURED (chosen)
cfg_width ──────────────► edge count cfg_width ──► [ shadow ] ──► edge count
cfg_div ──────────────► divider cfg_div ──► [ shadow ] ──► divider
cfg_cpha ──────────────► mode decoder cfg_cpha ──► [ shadow ] ──► mode decoder
a write mid-frame changes the frame a write mid-frame changes the NEXT frameThe cost is nine registers and one rule that has to be obeyed everywhere: inside a frame, every read is of a shadow. A single line that reads the live wire instead re-introduces the whole problem, invisibly, for one field only. That is Chapter 20.7's bug B6 — the edge count reading live cfg_width — and it is worth noting now that the only test that catches it is one that rewrites the configuration during a transfer. A suite that configures, transfers, and then reconfigures will never see it.
4. Commitment 3 — One Shift Direction, With Alignment At The Boundaries
MSB-first and LSB-first are usually implemented as two shift directions selected by a multiplexer. This design has one: always left, MOSI always the top bit, MISO always shifted into the bottom.
Bit order is produced entirely by how the word is aligned as it enters and leaves:
MSB-first load: tx << (16 - width) the top bit is bit width-1
store: mask to width bits first bit received is already at width-1
LSB-first load: reverse(tx) the top bit is now bit 0
store: reverse, then shift first bit received belongs at 0The second shift direction would have cost a mux in the datapath and, more importantly, a second set of everything that has to be proved about it: a second wrap condition, a second end-of-word test, and a doubling of the state space every property has to hold over. One reversal function called at two boundaries has neither.
5. Commitment 4 — Timing Generation Split From Datapath Control
Three concerns, three places, and the split is the reason the mode handling ends up being three lines.
DIVIDER counts system clocks, knows nothing about SPI
│
▼ tick, once per half-period
FSM counts half-periods and transitions, knows nothing about CPOL/CPHA
│
▼ edge index
MODE DECODER classifies the transition, knows nothing about data
│
▼ launch / sample
DATAPATH shifts, knows nothing about clocks or modesThe alternative is one state machine that counts clocks, tracks polarity, decides edges and drives the shift register — which is how a first version usually looks, and it works. What it does not do is let any part be reasoned about alone: a question about CPHA becomes a question about the whole machine.
The mode decoder, in its entirety
This is the intellectual payoff of the whole SPI track, and in this architecture it is three assignments:
edge_i the index of the transition about to be made, 0 .. 2N-1
leading = NOT edge_i[0] even transitions leave the idle level
sample_event = cpha ? trailing : leading
launch_event = cpha ? leading : trailingCPOL does not appear. It cannot, because the edge index already counts from the idle level — whatever that level is. CPOL's only job is to set the level SCLK sits at, and once SCLK starts there, transition zero is by definition the one that leaves it. That is why the four modes are two bits and not four cases, and why a design that gets the parity right gets all four modes right at once.
An even total of 2N transitions then gives REQ-MODE-001 for free: the clock ends a frame at the level it started, because it moved an even number of times. Nothing re-parks it.
6. The Five Phases, And Why One Of Them Is Inside busy
S_IDLE nothing selected, SCLK parked at the live CPOL, waiting for `start`
S_LEAD select asserted, counting cfg_lead half-periods, SCLK still
S_XFER 2N transitions, each one a launch, a sample, or the end
S_LAG SCLK back at the idle level, select still asserted
S_GAP select released, enforcing the turnaroundbusy is defined as not S_IDLE, so it stays high through S_GAP. That is a deliberate interface decision: REQ-TIM-004's turnaround becomes structurally enforced rather than a timing hope the caller has to honour. A controller that reported itself idle during the turnaround would push that obligation onto software, which is exactly where such requirements get lost — and it would make the requirement unverifiable from the controller's own interface.
The cost was named in Chapter 20.1 and is worth repeating because Chapter 20.7 asks about it: with no request queue, the caller cannot post the next transfer until busy falls, so the measured gap between frames is the enforced turnaround plus the caller's reaction time. The controller guarantees a floor and cannot guarantee the gap.
7. The Architectural Invariants
Written down before the RTL, because these are what Chapter 20.5 turns into checkers. Each one is a claim this architecture intends to make true.
| Invariant | Comes from | |
|---|---|---|
| INV-1 | At most one chip select is low at any instant | one index through one decoder |
| INV-2 | Outside a frame, SCLK equals the configured idle level | S_IDLE drives it every cycle |
| INV-3 | The bit index never exceeds the captured width | the edge count bounds the sample events |
| INV-4 | The captured configuration does not change while busy | the shadow is written only at acceptance |
| INV-5 | MOSI changes only at a launch event or with the chip select | the datapath is driven from launch_event alone |
| INV-6 | A sample event occurs only while a device is selected | S_XFER is reachable only from S_LEAD |
| INV-7 | done implies the full word was transferred | done is asserted on leaving S_LAG |
| INV-8 | busy is high from acceptance until the turnaround expires | busy is decoded from the state |
| INV-9 | An accepted request eventually completes, absent reset or abort | every phase has a terminating counter |
| INV-10 | The chip select releases only after the lag has elapsed | S_LAG precedes the release |
8. What This Architecture Makes Easy, And What It Makes Hard
Easy:
Adding a mode. The mode decoder is a function of one bit and a parity. A hypothetical CPHA-like variant would be a third term in two expressions, and nothing in the datapath would change.
Changing the width range. The width is a captured number used by one comparator and two alignment shifts. Extending to 32 bits is a parameter change and a wider shift register.
Reasoning about the pins. Every pin comes out of a register clocked by clk, so the output-delay story in Chapter 20.4 is one path per pin.
Hard:
Back-to-back throughput, for the reason in §6 — there is no queue, so the caller is inside the loop.
Very high SCLK rates. The half-period is an integer number of system clocks, so the achievable rates are f_clk / 2, f_clk / 4, f_clk / 6 and so on. There is no way to ask for a rate between two of those, and at cfg_div = 0 the design is generating a full-rate output from a flip-flop — which is where Chapter 20.4's output-delay budget becomes the limit rather than the logic.
Anything per-byte inside a frame. The frame is one word. A command-then-data transaction that holds the select across several words has to be built by the caller, and cannot be, because the select releases at the end of every frame. That is the most significant functional limitation of this architecture, it follows directly from the non-goals, and it is the first thing Chapter 20.7's review raises.
9. Summary
Nineteen requirements became seven blocks and four commitments.
SCLK is a waveform this design emits, not a clock it uses, which removes every internal clock-domain crossing at the cost of switching every register at the system rate and threading an enable through the datapath. Configuration is captured at acceptance into nine shadow registers, which makes a mid-frame rewrite harmless and creates one rule that must hold everywhere — inside a frame, every read is of a shadow. Bit order is alignment, not direction, so there is one shift path instead of two, which concentrates bit-order correctness into a single flag whose inversion a loopback test cannot see. And timing generation is split from datapath control, which is what reduces the four SPI modes to a parity and two conditional expressions with no mention of CPOL.
Ten invariants were written down, nine of them safety properties that a single cycle can refute and one — an accepted request eventually completes — that no per-cycle observation can check at all.
The architecture makes modes, widths and pin timing easy, and makes back-to-back throughput, arbitrary clock rates and multi-word transactions hard. All three of those limitations trace back to stated non-goals rather than to oversights, which is the property that lets the final review be about trade-offs instead of about mistakes.
10. What Comes Next
Chapter 20.3 implements this architecture in SystemVerilog, Verilog-2001 and VHDL, and the mode decoder turns out to be three lines in all three. The chapter also documents four defects found while writing it — one of which was correct for two modes and wrong for the other two, and was caught because the three languages disagreed.
Continue learning
Related tutorials
- Related topic
Master Microarchitecture
Partitioning an SPI master into blocks that each earn their place: why control and datapath separate along the line of what changes per bit, why configuration must be snapshotted at transfer start, and what a mid-transfer write must do instead of being ignored.
- Related topic
Capstone Requirements and Specification
One configurable SPI controller, specified before it is designed: nineteen numbered requirements, their corner cases, and the non-goals that keep the project finishable.
- Related topic
RTL Implementation and Mode Handling
The capstone controller in SystemVerilog, Verilog-2001 and VHDL with all four SPI modes derived from a parity, plus five defects found by running it — three of them because the three languages disagreed.
- Related topic
Timing, Constraints, and CDC Considerations
The MISO round trip decides the maximum SCLK rate and static timing analysis never checks it, plus why this architecture has no internal clock-domain crossing and what simulation cannot establish about reset release.
