Skip to content
VLSI Mentor

SPI · Module 20

Microarchitecture and Design Decisions

Turning nineteen requirements into five blocks, and defending four architectural commitments the design review will attack: a generated clock, a captured configuration, one shift direction, and timing split from control.

Chapter 20.1 produced nineteen requirements. This chapter turns them into hardware, and the interesting part is not which blocks exist — a divider, a shift register and a counter were never in doubt — but where the boundaries between them fall and which choices survive being questioned.

Four commitments are made here. Each has a named alternative, a stated cost, and a place in Chapter 20.7 where it is attacked.

1. The Blocks

Upper row left to right: config capture, clock divider, control FSM, mode decoder, shift datapath. Lower row: request accept feeding config capture, select decoder fed by the FSM, and registered pins fed by the datapath.config captureshadow written atacceptanceclock dividerone tick perhalf-periodcontrol FSMidle, lead, xfer,lag, gapmode decodertick to launch orsampleshift datapathone direction, bothordersrequest acceptstart, and only whenidleselect decoderone index, fouroutputsregistered pinssclk, mosi, cs_ncfg_divtickedge indexlaunch, sampleacceptphasemosi12
Figure 1 — the microarchitecture. The upper row is the path a transfer takes: configuration is captured, the divider paces it, the state machine sequences it, the mode decoder turns SCLK transitions into launch and sample events, and the datapath shifts. The lower row holds the three things that are deliberately NOT in that path. Every block shown is clocked by the system clock; SCLK appears only as an output of the registered pins.
BlockResponsibilityState it ownsReadsDrives
request acceptDecide whether a start becomes a transfernonestart, busy, width legalitythe capture pulse
config captureFreeze the configuration for one frame9 shadow fieldslive cfg_*every other block
clock dividerProduce one tick per SCLK half-perioddown-countercaptured cfg_divtick
control FSMSequence the five phases of a framestate, phase counter, edge indextick, abortSCLK level, phase
mode decoderClassify each transition as launch or samplenone — pure function of the edge indexedge index, captured CPHAlaunch, sample
shift datapathMove bits, one direction, and align at the boundariestx_sr, rx_sr, two indiceslaunch, sample, MISOMOSI, rx_data
select decoderOne index to one active-low outputnonecaptured device indexcs_n[3:0]

Every block above is clocked by clk. Not one of them is clocked by SCLK, and that is the first commitment.

2. Commitment 1 — SCLK Is An Output Waveform, Not A Clock

The controller generates SCLK by toggling a register. Nothing inside the design is clocked by it.

The alternative is real and is used in practice: divide the system clock down, treat the result as a clock, and clock the shift registers with it. That gives a shift register running at exactly the data rate, with no tick to distribute and no divider in the datapath's way.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   GENERATED CLOCK, USED INTERNALLY          CLOCK ENABLE, SCLK AS OUTPUT (chosen)

   clk ──► divider ──► sclk_int ──┐          clk ──┬──► everything
                                  │                │
                        shift reg ◄┘                └──► divider ──► tick ──► enables
                             │                                                  │
                             └──► pins                                shift reg ◄┘
                                                                          │
                                                                 sclk_r ──┴──► pins

What the chosen architecture buys:

There is no internal clock-domain crossing. Two clocks, even when one is derived from the other, need the synthesis and timing tools told how they relate; get it wrong and the tool either reports failures that do not exist or misses ones that do. With one clock there is nothing to relate.

The divider is not on a timing path to a pin. SCLK comes out of a flip-flop, so its output delay is one register-to-pin path, not a path through a counter's comparator.

Every state is visible on one clock. A waveform, a simulation, and a scan chain all see the whole design on clk. With an internal generated clock, the shift register's state advances at a rate the rest of the design cannot observe directly.

What it costs, stated plainly:

Every register toggles at the system clock rate whether a transfer is active or not — the divider counts, the FSM is evaluated, the pins are re-driven. A clock-gated or internally-clocked design switches far less. For a low-power target that is the decisive argument the other way, and this module does not pretend otherwise: power is a stated non-goal in Chapter 20.1 §9.

The datapath is paced by an enable, not by a clock, so every sequential element needs tick routed to it, and every piece of logic has to remember that a cycle without tick is a cycle where nothing moves. That is a whole class of bug — an update that forgets its enable — and Chapter 20.7 injects one.

3. Commitment 2 — Configuration Is Captured, Not Read Live

REQ-CFG-001 requires it, so the question here is only where the shadow lives and what it costs.

It lives in nine registers written by one pulse. The alternative — reading cfg_* live wherever needed — costs no registers at all and is the reason it is tempting.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   LIVE                                      CAPTURED (chosen)

   cfg_width ──────────────► edge count      cfg_width ──► [ shadow ] ──► edge count
   cfg_div   ──────────────► divider         cfg_div   ──► [ shadow ] ──► divider
   cfg_cpha  ──────────────► mode decoder    cfg_cpha  ──► [ shadow ] ──► mode decoder

   a write mid-frame changes the frame       a write mid-frame changes the NEXT frame

The cost is nine registers and one rule that has to be obeyed everywhere: inside a frame, every read is of a shadow. A single line that reads the live wire instead re-introduces the whole problem, invisibly, for one field only. That is Chapter 20.7's bug B6 — the edge count reading live cfg_width — and it is worth noting now that the only test that catches it is one that rewrites the configuration during a transfer. A suite that configures, transfers, and then reconfigures will never see it.

4. Commitment 3 — One Shift Direction, With Alignment At The Boundaries

MSB-first and LSB-first are usually implemented as two shift directions selected by a multiplexer. This design has one: always left, MOSI always the top bit, MISO always shifted into the bottom.

Bit order is produced entirely by how the word is aligned as it enters and leaves:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   MSB-first   load:  tx << (16 - width)     the top bit is bit width-1
               store: mask to width bits     first bit received is already at width-1

   LSB-first   load:  reverse(tx)            the top bit is now bit 0
               store: reverse, then shift    first bit received belongs at 0

The second shift direction would have cost a mux in the datapath and, more importantly, a second set of everything that has to be proved about it: a second wrap condition, a second end-of-word test, and a doubling of the state space every property has to hold over. One reversal function called at two boundaries has neither.

5. Commitment 4 — Timing Generation Split From Datapath Control

Three concerns, three places, and the split is the reason the mode handling ends up being three lines.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   DIVIDER          counts system clocks, knows nothing about SPI
                       │
                       ▼  tick, once per half-period
   FSM              counts half-periods and transitions, knows nothing about CPOL/CPHA
                       │
                       ▼  edge index
   MODE DECODER     classifies the transition, knows nothing about data
                       │
                       ▼  launch / sample
   DATAPATH         shifts, knows nothing about clocks or modes

The alternative is one state machine that counts clocks, tracks polarity, decides edges and drives the shift register — which is how a first version usually looks, and it works. What it does not do is let any part be reasoned about alone: a question about CPHA becomes a question about the whole machine.

The mode decoder, in its entirety

This is the intellectual payoff of the whole SPI track, and in this architecture it is three assignments:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   edge_i         the index of the transition about to be made, 0 .. 2N-1
   leading        = NOT edge_i[0]            even transitions leave the idle level
   sample_event   = cpha ? trailing : leading
   launch_event   = cpha ? leading  : trailing

CPOL does not appear. It cannot, because the edge index already counts from the idle level — whatever that level is. CPOL's only job is to set the level SCLK sits at, and once SCLK starts there, transition zero is by definition the one that leaves it. That is why the four modes are two bits and not four cases, and why a design that gets the parity right gets all four modes right at once.

An even total of 2N transitions then gives REQ-MODE-001 for free: the clock ends a frame at the level it started, because it moved an even number of times. Nothing re-parks it.

6. The Five Phases, And Why One Of Them Is Inside busy

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   S_IDLE   nothing selected, SCLK parked at the live CPOL, waiting for `start`
   S_LEAD   select asserted, counting cfg_lead half-periods, SCLK still
   S_XFER   2N transitions, each one a launch, a sample, or the end
   S_LAG    SCLK back at the idle level, select still asserted
   S_GAP    select released, enforcing the turnaround

busy is defined as not S_IDLE, so it stays high through S_GAP. That is a deliberate interface decision: REQ-TIM-004's turnaround becomes structurally enforced rather than a timing hope the caller has to honour. A controller that reported itself idle during the turnaround would push that obligation onto software, which is exactly where such requirements get lost — and it would make the requirement unverifiable from the controller's own interface.

The cost was named in Chapter 20.1 and is worth repeating because Chapter 20.7 asks about it: with no request queue, the caller cannot post the next transfer until busy falls, so the measured gap between frames is the enforced turnaround plus the caller's reaction time. The controller guarantees a floor and cannot guarantee the gap.

7. The Architectural Invariants

Written down before the RTL, because these are what Chapter 20.5 turns into checkers. Each one is a claim this architecture intends to make true.

InvariantComes from
INV-1At most one chip select is low at any instantone index through one decoder
INV-2Outside a frame, SCLK equals the configured idle levelS_IDLE drives it every cycle
INV-3The bit index never exceeds the captured widththe edge count bounds the sample events
INV-4The captured configuration does not change while busythe shadow is written only at acceptance
INV-5MOSI changes only at a launch event or with the chip selectthe datapath is driven from launch_event alone
INV-6A sample event occurs only while a device is selectedS_XFER is reachable only from S_LEAD
INV-7done implies the full word was transferreddone is asserted on leaving S_LAG
INV-8busy is high from acceptance until the turnaround expiresbusy is decoded from the state
INV-9An accepted request eventually completes, absent reset or abortevery phase has a terminating counter
INV-10The chip select releases only after the lag has elapsedS_LAG precedes the release

8. What This Architecture Makes Easy, And What It Makes Hard

Easy:

Adding a mode. The mode decoder is a function of one bit and a parity. A hypothetical CPHA-like variant would be a third term in two expressions, and nothing in the datapath would change.

Changing the width range. The width is a captured number used by one comparator and two alignment shifts. Extending to 32 bits is a parameter change and a wider shift register.

Reasoning about the pins. Every pin comes out of a register clocked by clk, so the output-delay story in Chapter 20.4 is one path per pin.

Hard:

Back-to-back throughput, for the reason in §6 — there is no queue, so the caller is inside the loop.

Very high SCLK rates. The half-period is an integer number of system clocks, so the achievable rates are f_clk / 2, f_clk / 4, f_clk / 6 and so on. There is no way to ask for a rate between two of those, and at cfg_div = 0 the design is generating a full-rate output from a flip-flop — which is where Chapter 20.4's output-delay budget becomes the limit rather than the logic.

Anything per-byte inside a frame. The frame is one word. A command-then-data transaction that holds the select across several words has to be built by the caller, and cannot be, because the select releases at the end of every frame. That is the most significant functional limitation of this architecture, it follows directly from the non-goals, and it is the first thing Chapter 20.7's review raises.

9. Summary

Nineteen requirements became seven blocks and four commitments.

SCLK is a waveform this design emits, not a clock it uses, which removes every internal clock-domain crossing at the cost of switching every register at the system rate and threading an enable through the datapath. Configuration is captured at acceptance into nine shadow registers, which makes a mid-frame rewrite harmless and creates one rule that must hold everywhere — inside a frame, every read is of a shadow. Bit order is alignment, not direction, so there is one shift path instead of two, which concentrates bit-order correctness into a single flag whose inversion a loopback test cannot see. And timing generation is split from datapath control, which is what reduces the four SPI modes to a parity and two conditional expressions with no mention of CPOL.

Ten invariants were written down, nine of them safety properties that a single cycle can refute and one — an accepted request eventually completes — that no per-cycle observation can check at all.

The architecture makes modes, widths and pin timing easy, and makes back-to-back throughput, arbitrary clock rates and multi-word transactions hard. All three of those limitations trace back to stated non-goals rather than to oversights, which is the property that lets the final review be about trade-offs instead of about mistakes.

10. What Comes Next

Chapter 20.3 implements this architecture in SystemVerilog, Verilog-2001 and VHDL, and the mode decoder turns out to be three lines in all three. The chapter also documents four defects found while writing it — one of which was correct for two modes and wrong for the other two, and was caught because the three languages disagreed.

Continue learning