SPI · Module 1
The Shift-Register Mental Model
The hardware underneath every SPI transfer: two shift registers wired into a ring, one bit leaving each end and one arriving on every enabled edge. Traced cycle by cycle, built in Verilog, SystemVerilog and VHDL, and separated from the meaning a device layers on top.
Chapter 1.2 settled who is allowed to drive each wire. It deliberately said nothing about where the bit on MOSI comes from, or what happens to the bit arriving on MISO. That is this chapter, and it is the most load-bearing model in the whole curriculum: almost every later result — full duplex, dummy cycles, why a "read" still transmits, why a one-bit shift is the signature of a mode error — is this one picture, applied.
The model is small enough to hold in your head and precise enough to predict real waveforms. Each end of an SPI link holds a register that shifts. The two registers are wired into a ring. Every enabled clock edge pushes one bit out of each register and pulls one bit in. After the register's width in edges, the two ends have swapped contents.
Everything below earns that sentence, then builds it in hardware.
1. The Problem: a Register Is Parallel, a Wire Is Not
Start from the mismatch. Inside a chip, a byte lives in a register — eight flip-flops, read and written all at once across eight parallel nets. On the board between two chips, Chapter 1.1 established there is one data wire in each direction, because width is what costs pins.
So the interface needs a structure that converts between the two: something that accepts a whole word in parallel, emits it one bit at a time, and simultaneously assembles the arriving bits back into a whole word. That structure is a shift register, and it is not a special SPI invention — it is the standard answer to the parallel/serial mismatch, and the reason SPI hardware is small enough to put in a sensor.
2. What a Shift Register Does
A shift register is a row of flip-flops in which each stage's output feeds the next stage's input. On an enabled clock edge, every stage simultaneously takes the value of its neighbour. Three consequences follow, and all three matter.
One bit leaves. The stage at the end of the row has no downstream neighbour inside the register, so its value is available as a serial output. This is the bit driven onto the wire.
One bit enters. The stage at the other end has no upstream neighbour, so its input must come from outside — a serial input. This is the bit captured from the wire.
The width is conserved. The register does not grow or shrink. A bit leaving and a bit entering happen on the same edge, so the register always holds exactly its width in bits — a mixture of what is left of the original word and what has arrived so far.
That third point is where the interesting behaviour comes from, and it is the part newcomers tend not to notice. A shift register is not a queue that drains. It is a fixed-size window that the outgoing word slides out of while the incoming word slides in behind it.
3. Two Registers, One Ring
Now connect two of them, using the ownership rules from Chapter 1.2.
The master's shift register drives MOSI from its serial output. MOSI reaches the selected slave's serial input. The slave's shift register drives MISO from its serial output, and MISO reaches the master's serial input. Both registers are shifted by the same clock — SCLK, which the master owns.
The result is a ring: master → MOSI → slave → MISO → master. Not two channels; one loop.
Two properties of the ring are worth stating explicitly, because later chapters lean on both.
It is symmetric in structure, asymmetric in control. Both ends hold the same kind of register and do the same thing on each edge. What differs is that one end supplies the clock and the selection — the ownership asymmetry from Chapter 1.2. The data path is a ring between equals; the control path is not.
It conserves bits. Nothing in the ring creates or destroys a bit. Over N enabled edges, exactly N bits cross MOSI and exactly N cross MISO. That conservation is what makes the exchange predictable enough to verify against a reference model.
4. One Edge, Two Halves
Zoom in on a single bit time, because this is where a model that is too coarse starts producing wrong predictions.
A bit time has two jobs to do, and they cannot be simultaneous. The bit being sent has to be presented on the wire and held there long enough for the other end to see it reliably. Then it has to be captured into the receiving register. If both ends tried to present and capture at the same instant, the receiver would be sampling a line that is changing.
So SPI splits the bit time: one clock edge is used to launch new data onto the wire, and the other edge is used to capture it. A bit presented on one edge is sampled by the far end on the next, while the line is stable.
This chapter deliberately stops there. Which edge launches and which captures, what the clock does between transfers, and how those two choices produce four named modes is the subject of Module 2 and Module 3. What you need here is only the structural fact: within one bit time, launch and capture are separated, which is why both directions can move a bit per clock without either end sampling a moving signal.
For the rest of this chapter, "an enabled edge" means one complete bit time — one bit out and one bit in at each end. That abstraction is exactly what the RTL in §7 makes concrete with a shift_en pulse.
5. Eight Edges, Traced
Abstractions are cheap; let us run one. Take an 8-bit exchange. The master has loaded 0xA5 (1010_0101) into its register and the slave has loaded 0x3C (0011_1100) into its. Both shift most-significant bit first — the common arrangement, and one we will interrogate in §6.
Each row shows the state before that edge: what each register holds, and therefore what each end is presenting on its wire.
| Edge | Master register (before) | MOSI | MISO | Slave register (before) |
|---|---|---|---|---|
| 1 | 1010_0101 | 1 | 0 | 0011_1100 |
| 2 | 0100_1010 | 0 | 0 | 0111_1001 |
| 3 | 1001_0100 | 1 | 1 | 1111_0010 |
| 4 | 0010_1001 | 0 | 1 | 1110_0101 |
| 5 | 0101_0011 | 0 | 1 | 1100_1010 |
| 6 | 1010_0111 | 1 | 1 | 1001_0100 |
| 7 | 0100_1111 | 0 | 0 | 0010_1001 |
| 8 | 1001_1110 | 1 | 0 | 0101_0010 |
| after 8 | 0011_1100 = 0x3C | — | — | 1010_0101 = 0xA5 |
Read three things off that table.
The MOSI column, top to bottom, is 1010_0101 — the master's original word, MSB first. Nothing assembled it; it is simply the successive most-significant bits as the word slides out.
The MISO column is 0011_1100 — the slave's original word, MSB first. The same mechanism, in the other direction, on the same edges.
After eight edges the registers hold each other's words. The master's register contains 0x3C and the slave's contains 0xA5. This is the sense in which an SPI transfer is an exchange rather than a send or a receive: the hardware swapped two words, and any interpretation more specific than that is something a device layered on top.
One 8-bit exchange — 0xA5 out, 0x3C in
8 cyclesNote what the waveform does not show: it draws one cycle per bit time rather than separating the launch edge from the capture edge, because that separation belongs to Module 2. Treat each column as one complete bit time.
6. Shift Direction and Which Bit Goes First
The trace above shifted left and emitted the most significant bit first. Both halves of that sentence deserve a moment, because they are the same decision seen from two sides.
Emitting the MSB first requires the serial output to be taken from the register's most-significant stage, and each shift to move every bit one position toward that stage — a left shift with the arriving bit entering at the least-significant end. Emitting the LSB first requires the mirror image: output from the least-significant stage, a right shift, arriving bits entering at the top.
Two observations that save real debugging time later.
The direction is a property of the implementation, not of SPI. MSB-first is much the more common convention and is what most devices specify, but it is a device and controller configuration rather than something the interface fixes. A configurable master implements both — how, and what that costs, is Module 13.
Transmit and receive share the decision. One register does both jobs, so the end that shifts out MSB-first necessarily assembles the incoming word MSB-first too. You do not get to choose independently for each direction on a single register. This is why a bit-order mismatch corrupts both directions at once, and why its waveform signature — a byte arriving bit-reversed rather than shifted — differs from a mode error. Module 4 covers ordering as configuration; Module 18 covers telling the two failure signatures apart.
7. Building It — the Exchange Core in Three HDLs
Now make the model concrete. What follows is a deliberately narrow module: the shift register at one end of the ring, with a shift_en pulse standing in for "one bit time has elapsed."
Architecture
One register of WIDTH flip-flops. Its most-significant stage drives the serial output sdo. On an enabled edge it shifts left, with the serial input sdi entering at the least-significant end. A separate load input captures a parallel word to be sent, and the register contents are continuously visible as rx_data.
State. Exactly one register, shreg. There is no separate transmit and receive register — and that is the point of the chapter. tx_data is what you put in; rx_data is what has arrived; they are the same flip-flops at different moments.
Combinational logic. Two assignments, both trivial: sdo is the top bit of shreg, and rx_data is shreg itself. All the behaviour lives in the clocked block.
Clocking. Everything advances on the rising edge of clk. Note carefully that clk here is the system clock, not SCLK: shift_en is a one-cycle pulse that a real master would generate at the appropriate SCLK edge. Structuring it this way keeps the example honest about what it is — the datapath, not the clock generator.
Reset. Asynchronous, active-low, clearing the register to zero. Chosen for familiarity rather than because SPI requires it; reset strategy for a real SPI block is Module 13 for a master and Module 14 for a slave, where it interacts with CS and an externally supplied clock.
Ownership. sdi is an input, sdo an unconditional output. There is no output enable here — that belongs to the slave's MISO path, built in Module 14, for exactly the reasons Chapter 1.2 gave.
Timing assumption. That shift_en is asserted for exactly one clk cycle per bit time, and that sdi is stable when it is. A real implementation earns that by generating shift_en from the correct SCLK edge.
Limitation — read this before reusing the code. This is an educational fragment, not SPI IP. It has no SCLK generation, no clock divider, no CS handling, no bit counter or transfer-complete signal, no mode configuration, no bit-order option, and no clock-domain crossing for the slave case. Adding those is precisely the work of Modules 13 to 15. What it does have is the exchange mechanism, isolated so you can see it.
// WIDTH must be >= 2: the shift concatenation slices [WIDTH-2:0].
module spi_shift_core #(
parameter int WIDTH = 8
) (
input logic clk, // SYSTEM clock, not SCLK
input logic rst_n, // asynchronous, active-low
input logic load, // capture tx_data into the register
input logic [WIDTH-1:0] tx_data, // word to send
input logic shift_en, // one pulse = one bit time
input logic sdi, // serial in — the far end's sdo
output logic sdo, // serial out — this end's MSB
output logic [WIDTH-1:0] rx_data // what has shifted in so far
);
logic [WIDTH-1:0] shreg;
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) shreg <= '0;
else if (load) shreg <= tx_data; // load wins over shift
else if (shift_en) shreg <= {shreg[WIDTH-2:0], sdi}; // MSB out, sdi in at LSB
end
assign sdo = shreg[WIDTH-1]; // the bit currently presented on the wire
assign rx_data = shreg; // same flip-flops, read as the received word
endmoduleThe testbench is the chapter's claim made executable: instantiate two cores, wire them into the ring of Figure 1, and check both that each wire carries the right bit at the right time and that the two words end up swapped. It also holds shift_en low for a cycle mid-transfer, because a paused clock must not corrupt the exchange — the SPI reality that SCLK need not be continuous.
module spi_shift_core_tb;
localparam int WIDTH = 8;
localparam logic [WIDTH-1:0] M_TX = 8'hA5; // 1010_0101
localparam logic [WIDTH-1:0] S_TX = 8'h3C; // 0011_1100
logic clk = 1'b0, rst_n, load, shift_en;
logic m_sdo, s_sdo;
logic [WIDTH-1:0] m_rx, s_rx;
int errors = 0;
// The ring: master sdo -> slave sdi is MOSI; slave sdo -> master sdi is MISO.
spi_shift_core #(.WIDTH(WIDTH)) master (
.clk(clk), .rst_n(rst_n), .load(load), .tx_data(M_TX),
.shift_en(shift_en), .sdi(s_sdo), .sdo(m_sdo), .rx_data(m_rx));
spi_shift_core #(.WIDTH(WIDTH)) slave (
.clk(clk), .rst_n(rst_n), .load(load), .tx_data(S_TX),
.shift_en(shift_en), .sdi(m_sdo), .sdo(s_sdo), .rx_data(s_rx));
always #5 clk = ~clk;
initial begin
rst_n = 1'b0; load = 1'b0; shift_en = 1'b0;
@(posedge clk); #1; rst_n = 1'b1;
@(posedge clk); #1; load = 1'b1; // both ends capture their word
@(posedge clk); #1; load = 1'b0;
for (int i = 0; i < WIDTH; i++) begin
// Before the edge, each end presents its next bit, MSB-first.
if (m_sdo !== M_TX[WIDTH-1-i]) begin
$error("bit %0d: MOSI=%b expected %b", i, m_sdo, M_TX[WIDTH-1-i]); errors++;
end
if (s_sdo !== S_TX[WIDTH-1-i]) begin
$error("bit %0d: MISO=%b expected %b", i, s_sdo, S_TX[WIDTH-1-i]); errors++;
end
// Boundary case: a bit time with no shift_en must change nothing.
if (i == 3) begin
@(posedge clk); #1;
if (m_sdo !== M_TX[WIDTH-1-i] || s_sdo !== S_TX[WIDTH-1-i]) begin
$error("paused cycle shifted the registers"); errors++;
end
end
shift_en = 1'b1; @(posedge clk); #1; shift_en = 1'b0;
end
// The whole point: after WIDTH bit times the words have swapped.
if (m_rx !== S_TX) begin $error("master rx=%h expected %h", m_rx, S_TX); errors++; end
if (s_rx !== M_TX) begin $error("slave rx=%h expected %h", s_rx, M_TX); errors++; end
if (errors == 0) $display("PASS: MOSI=%h, MISO=%h, registers exchanged", M_TX, S_TX);
else $display("FAIL: %0d mismatches", errors);
$finish;
end
endmoduleThe Verilog form is the same hardware with reg/wire typing, {WIDTH{1'b0}} for the reset constant, and $display in place of $error — the latter is a SystemVerilog task, so a Verilog-2001 toolchain will not accept it.
// WIDTH must be >= 2: the shift concatenation slices [WIDTH-2:0].
module spi_shift_core #(
parameter WIDTH = 8
) (
input clk, // SYSTEM clock, not SCLK
input rst_n, // asynchronous, active-low
input load,
input [WIDTH-1:0] tx_data,
input shift_en, // one pulse = one bit time
input sdi,
output sdo,
output [WIDTH-1:0] rx_data
);
reg [WIDTH-1:0] shreg;
always @(posedge clk or negedge rst_n) begin
if (!rst_n) shreg <= {WIDTH{1'b0}};
else if (load) shreg <= tx_data;
else if (shift_en) shreg <= {shreg[WIDTH-2:0], sdi};
end
assign sdo = shreg[WIDTH-1];
assign rx_data = shreg;
endmodule module spi_shift_core_tb;
parameter WIDTH = 8;
localparam [WIDTH-1:0] M_TX = 8'hA5;
localparam [WIDTH-1:0] S_TX = 8'h3C;
reg clk = 1'b0;
reg rst_n, load, shift_en;
wire m_sdo, s_sdo;
wire [WIDTH-1:0] m_rx, s_rx;
integer i;
integer errors = 0;
spi_shift_core #(.WIDTH(WIDTH)) master (
.clk(clk), .rst_n(rst_n), .load(load), .tx_data(M_TX),
.shift_en(shift_en), .sdi(s_sdo), .sdo(m_sdo), .rx_data(m_rx));
spi_shift_core #(.WIDTH(WIDTH)) slave (
.clk(clk), .rst_n(rst_n), .load(load), .tx_data(S_TX),
.shift_en(shift_en), .sdi(m_sdo), .sdo(s_sdo), .rx_data(s_rx));
always #5 clk = ~clk;
initial begin
rst_n = 1'b0; load = 1'b0; shift_en = 1'b0;
@(posedge clk); #1; rst_n = 1'b1;
@(posedge clk); #1; load = 1'b1;
@(posedge clk); #1; load = 1'b0;
for (i = 0; i < WIDTH; i = i + 1) begin
if (m_sdo !== M_TX[WIDTH-1-i]) begin
$display("ERROR bit %0d: MOSI=%b expected %b", i, m_sdo, M_TX[WIDTH-1-i]);
errors = errors + 1;
end
if (s_sdo !== S_TX[WIDTH-1-i]) begin
$display("ERROR bit %0d: MISO=%b expected %b", i, s_sdo, S_TX[WIDTH-1-i]);
errors = errors + 1;
end
if (i == 3) begin
@(posedge clk); #1;
if (m_sdo !== M_TX[WIDTH-1-i] || s_sdo !== S_TX[WIDTH-1-i]) begin
$display("ERROR: paused cycle shifted the registers");
errors = errors + 1;
end
end
shift_en = 1'b1; @(posedge clk); #1; shift_en = 1'b0;
end
if (m_rx !== S_TX) begin
$display("ERROR master rx=%h expected %h", m_rx, S_TX); errors = errors + 1;
end
if (s_rx !== M_TX) begin
$display("ERROR slave rx=%h expected %h", s_rx, M_TX); errors = errors + 1;
end
if (errors == 0) $display("PASS: MOSI=%h, MISO=%h, registers exchanged", M_TX, S_TX);
else $display("FAIL: %0d mismatches", errors);
$finish;
end
endmoduleThe VHDL form uses a std_logic_vector for the register and expresses the shift as a concatenation of a slice with the incoming bit — the same operation the two other languages perform, written with VHDL's downto indexing.
library ieee;
use ieee.std_logic_1164.all;
entity spi_shift_core is
generic (
WIDTH : positive := 8 -- WIDTH must be >= 2
);
port (
clk : in std_logic; -- SYSTEM clock, not SCLK
rst_n : in std_logic; -- asynchronous, active-low
load : in std_logic;
tx_data : in std_logic_vector(WIDTH-1 downto 0);
shift_en : in std_logic; -- one pulse = one bit time
sdi : in std_logic;
sdo : out std_logic;
rx_data : out std_logic_vector(WIDTH-1 downto 0)
);
end entity spi_shift_core;
architecture rtl of spi_shift_core is
signal shreg : std_logic_vector(WIDTH-1 downto 0);
begin
shift_proc : process (clk, rst_n)
begin
if rst_n = '0' then
shreg <= (others => '0');
elsif rising_edge(clk) then
if load = '1' then
shreg <= tx_data; -- load wins over shift
elsif shift_en = '1' then
shreg <= shreg(WIDTH-2 downto 0) & sdi; -- MSB out, sdi in at LSB
end if;
end if;
end process shift_proc;
sdo <= shreg(WIDTH-1);
rx_data <= shreg;
end architecture rtl; library ieee;
use ieee.std_logic_1164.all;
entity spi_shift_core_tb is
end entity spi_shift_core_tb;
architecture sim of spi_shift_core_tb is
constant WIDTH : positive := 8;
constant M_TX : std_logic_vector(WIDTH-1 downto 0) := "10100101"; -- 0xA5
constant S_TX : std_logic_vector(WIDTH-1 downto 0) := "00111100"; -- 0x3C
constant TP : time := 10 ns;
signal clk : std_logic := '0';
signal rst_n : std_logic := '0';
signal load : std_logic := '0';
signal shift_en : std_logic := '0';
signal m_sdo, s_sdo : std_logic;
signal m_rx, s_rx : std_logic_vector(WIDTH-1 downto 0);
signal done : boolean := false;
begin
clk <= not clk after TP/2 when not done else '0';
-- The ring: master sdo -> slave sdi (MOSI); slave sdo -> master sdi (MISO).
master : entity work.spi_shift_core
generic map (WIDTH => WIDTH)
port map (clk => clk, rst_n => rst_n, load => load, tx_data => M_TX,
shift_en => shift_en, sdi => s_sdo, sdo => m_sdo, rx_data => m_rx);
slave : entity work.spi_shift_core
generic map (WIDTH => WIDTH)
port map (clk => clk, rst_n => rst_n, load => load, tx_data => S_TX,
shift_en => shift_en, sdi => m_sdo, sdo => s_sdo, rx_data => s_rx);
stim : process
variable errors : natural := 0;
begin
wait until rising_edge(clk); wait for 1 ns; rst_n <= '1';
wait until rising_edge(clk); wait for 1 ns; load <= '1';
wait until rising_edge(clk); wait for 1 ns; load <= '0';
for i in 0 to WIDTH-1 loop
if m_sdo /= M_TX(WIDTH-1-i) then
report "MOSI mismatch at bit " & integer'image(i) severity error;
errors := errors + 1;
end if;
if s_sdo /= S_TX(WIDTH-1-i) then
report "MISO mismatch at bit " & integer'image(i) severity error;
errors := errors + 1;
end if;
-- Boundary case: a bit time with no shift_en must change nothing.
if i = 3 then
wait until rising_edge(clk); wait for 1 ns;
if m_sdo /= M_TX(WIDTH-1-i) or s_sdo /= S_TX(WIDTH-1-i) then
report "paused cycle shifted the registers" severity error;
errors := errors + 1;
end if;
end if;
shift_en <= '1';
wait until rising_edge(clk); wait for 1 ns;
shift_en <= '0';
end loop;
if m_rx /= S_TX then
report "master rx mismatch" severity error; errors := errors + 1;
end if;
if s_rx /= M_TX then
report "slave rx mismatch" severity error; errors := errors + 1;
end if;
if errors = 0 then
report "PASS: registers exchanged" severity note;
else
report "FAIL: " & integer'image(errors) & " mismatches" severity error;
end if;
done <= true;
wait;
end process stim;
end architecture sim;What the three versions agree on, and where they differ
All three infer the same hardware: WIDTH flip-flops with an asynchronous active-low reset, a load path, and a shift path taking the top bit out and the serial input in at the bottom. The differences are linguistic rather than structural — logic and always_ff versus reg and always, '0 versus {WIDTH{1'b0}} versus (others => '0'), and the concatenation operator {a, b} versus a & b. The SystemVerilog version states its intent to the tool (always_ff is a promise that this block infers flip-flops); the Verilog version relies on the coding pattern alone; the VHDL version makes the reset structure explicit in the if/elsif rising_edge shape.
One genuine semantic note on the testbenches: $error and assert are SystemVerilog, so the Verilog testbench counts errors and reports with $display, while VHDL uses report … severity. All three are self-checking and print a single PASS or FAIL line.
8. Why This Model Explains Full Duplex
You now have everything needed to see why SPI cannot help being full duplex, so it is worth stating before the next chapter develops it.
There is one register per end, and it shifts. A bit leaves and a bit arrives on the same edge because those are the two ends of one shift operation, not two separate features. A design that wanted to transmit without receiving would have to actively discard what shifted in — and would still have shifted it in.
That is the whole mechanism. Chapter 1.4 takes it to its consequences: why a device "read" still transmits bits, where dummy bytes come from, why the received bytes during a command phase are meaningless but present, and why a transaction object in a testbench naturally holds both directions rather than being typed read or write.
9. What This Model Does Not Tell You
A good mental model is defined as much by its boundaries as by its content. This one is silent on four things, and assuming otherwise causes real errors.
It does not say which edge does what. Launch and capture are separated within a bit time (§4), but which physical edge plays which role is configuration — Modules 2 and 3.
It does not say how many bits a transfer contains. The register has a width; a transaction may be many registers' worth, and what delimits it is CS, not the register. Chapter 1.2 established CS as the boundary; Module 4 covers transfer width and framing.
It does not say what the bits mean. The ring swaps words. That a particular outgoing word is an opcode and a particular incoming word is a status byte is a device contract, from the datasheet — the ownership-versus-meaning separation from Chapter 1.2, applied again.
It does not describe a real slave's clocking. The RTL above shifts on a system clock with an enable. A slave has no system clock of its own relative to SCLK — the clock arrives from outside, which turns this simple register into a clock-domain problem. That is Module 15, and it is the single biggest gap between this fragment and a working slave.
10. Why an RTL Designer Cares
This chapter is the datapath of every SPI block you will ever write, and three design consequences follow directly from it.
Transmit and receive are one register, not two. The strong temptation when starting an SPI master is to build a TX shift register and a separate RX shift register. It works, and it costs twice the flip-flops for no benefit in the common case, because the outgoing word has vacated a bit position exactly as an incoming bit needs one. Knowing why one register suffices is the difference between copying a design and choosing one. Where two registers genuinely help — decoupling the next transmit word from the received one so software can be slower than the bus — is a buffering decision, not a shift decision, and it is Module 13.
The register is the natural place for width and order configuration. Transfer width becomes the count of enabled edges and the register's parameterisation; bit order becomes the choice of which stage drives sdo and which direction the shift runs. Both are small changes to this core rather than new blocks — which is why a parameterised core is worth building once.
shift_en is where the real work goes. Notice how little logic the core needed. In a real master, essentially all the remaining complexity is in producing shift_en at the right moments: dividing the system clock to SCLK, placing launch and capture on the configured edges, counting bits, and framing with CS. This chapter deliberately handed you the easy half so that Module 13 can spend its time on the hard half.
11. Why a Verification Engineer Cares
The shift model is what makes SPI checkable at all, because it supplies a reference model that needs no DUT internals.
A monitor can reconstruct both words from the pins alone. Sample MOSI and MISO on the capture edges between a CS assertion and its deassertion, assemble each into a word using the configured bit order, and you have the complete exchange — without probing a single internal register. That property is what makes a passive SPI monitor possible, and it is a direct consequence of the ring conserving bits.
The strongest early check is a structural one. For a transfer of N bit times, exactly N bits must cross in each direction. A monitor that counts capture edges between CS edges and compares against the configured width catches short frames, extra clocks and dropped edges — a whole family of bugs — before any data comparison happens. That is a bit-count assertion, and it is cheap.
A scoreboard must expect an exchange, not a transfer. Because every transfer produces a received word, a scoreboard that models only the direction the test cared about will either ignore real evidence or trip over meaningless bytes. The right model is the one the hardware implements: two words in, two words out, with the device's contract deciding which of them carry meaning. Chapter 1.4 shows what that does to a transaction object, and Modules 16 and 17 build the environment properly.
12. Why an FPGA Engineer Cares
Two practical points, both consequences of the model rather than restatements of it.
The datapath is nearly free; the control is not. A shift register of a byte or two is a handful of flip-flops and no arithmetic — it will meet timing on any device you are likely to use. When an SPI block fails timing on an FPGA, the failure is essentially never in the shifting. It is in the clock path, the I/O registers, or the path from an externally supplied SCLK into the fabric — which is why Module 15 exists and why this chapter's RTL deliberately abstracts SCLK behind shift_en.
Where you place the capture register decides your margin. The model says a bit is captured at some point in the bit time; the physical question is how much delay sits between the pin and the capturing flip-flop. Capturing in an I/O-block register places that flip-flop as close to the pad as the device allows and gives a predictable, constrainable path; capturing deep in the fabric adds routing delay that varies with each build. The model does not care. Timing closure does.
13. Common Misconceptions
14. Reason It Through
Work this before reading the answer. It is the kind of question a bring-up log actually poses.
A master and a slave are exchanging 8-bit words. The master intends to send
0xA5and the slave intends to return0x3C, exactly as in §5. On the logic analyser, MOSI shows0xA5correctly, but the master's received byte is0x78instead of0x3C— and MISO on the capture shows the slave's bits are correct. What single mechanism explains this?
First, what is ruled out. MOSI is correct, so the master's shift-out path, its bit order and its clocking of the outgoing bit are all fine. MISO on the wire is correct, so the slave produced the right bits at the right times. The fault is therefore in what the master does with MISO — the capture side only.
Compare the words bit by bit. Expected 0x3C is 0011_1100. Observed 0x78 is 0111_1000. The observed word is the expected word shifted left by one, with a zero entering at the bottom. The master assembled eight bits, but they are offset by one position relative to the slave's stream.
What mechanism produces exactly a one-position offset? The master's register performed one shift too many relative to the arriving data — equivalently, it began capturing one bit time before the slave began presenting, so its first captured bit was not the slave's first data bit and the last real bit fell off the top. In the language of this chapter, the two registers were not stepping in lockstep: the ring conserves bits only if both ends agree on which edge is bit one.
What would cause that in practice? The common causes are a capture edge chosen one half-cycle early relative to the slave's launch edge — a clock-phase disagreement — or a master that starts shifting on the CS assertion edge rather than on the first clock. Both are real, and both produce this identical signature.
Why does this chapter stop here rather than name the culprit? Because distinguishing them requires the launch/capture edge vocabulary that Module 2 builds and the mode derivation that Module 3 supplies. What the shift model gives you is the more valuable half of the diagnosis: the failure is a position error, not a data error, so look at edge alignment and not at either device's data path. That narrowing is the reasoning skill — Module 18 turns it into a method.
15. Understanding Check
16. Summary
A shift register resolves the mismatch between a parallel register inside a chip and a single wire between chips. Each end of an SPI link holds one, and they are wired into a ring: the master's serial output drives MOSI into the slave's serial input, the slave's serial output drives MISO back into the master's, and SCLK shifts both.
On each enabled edge, one bit leaves and one bit arrives at every end, because those are the two ends of a single shift rather than two separate operations. The register never empties — it is a fixed-width window holding what remains of the outgoing word plus what has arrived of the incoming one. After the width in edges, the two ends have exchanged words: in §5's trace, 0xA5 and 0x3C swap places, with MOSI carrying one and MISO the other, MSB first.
The RTL makes the claim testable in eleven lines of logic: one register, a load path, a shift path taking the top bit out and the serial input in at the bottom. Two of them wired into a ring provably exchange their contents — and the smallness of that datapath is itself the lesson, because it tells you that the real work in an SPI block is generating the shift enable at the right moments, not shifting.
The model's boundaries matter as much as its content: it says nothing about which edge launches or captures, how many bits a transaction contains, what the bits mean, or how a slave copes with a clock it does not own. Each of those is a later module. What it does give you is a prediction of the contents of both registers at every instant of a transfer — which is exactly what you need to read a waveform, write a reference model, or recognise that a byte is offset rather than wrong.
17. What Comes Next
The ring makes one consequence unavoidable, and Chapter 1.4 — Full-Duplex Exchange is about facing it: every transfer moves information both ways, whether or not the software wanted it to. That chapter explains where dummy transmit bytes come from, why the bytes received during a command phase exist but mean nothing, and why "read" and "write" are interpretations a device imposes rather than modes the hardware has.
Browse the path on the SPI curriculum index, or revisit Master, Slave, and Signal Ownership for the driver rules this chapter's ring depends on. For the same shift-register structure solving the asynchronous version of the problem — no shared clock, timing recovered from the data — see What a UART Actually Is.
Continue learning
Related tutorials
- Related topic
SCLK Generation, Period, and Frequency
Where SCLK comes from and what one period buys. Dividing a system clock to a bus clock, why the divisor is an integer and what that costs, and how a period in nanoseconds becomes the budget every later timing parameter is spent from.
- Related topic
CPOL — Clock Polarity
The first of the two bits that define an SPI mode. What clock polarity specifies, why it is a property of the idle state rather than of the transfer, and the RTL change that makes a divider polarity-aware.
- Related topic
Transfer Width
SPI has no native word size. What a 12-bit ADC or 16-bit codec requires of a master, why chip select rather than a bit count delimits a frame, and a width-parameterized transfer engine in three HDLs.
- Related topic
TX and RX Shift Registers
The shift datapath, and why there are two registers rather than the one SPI's symmetry seems to permit: the three failures a circulating register cannot express, why transmit needs alignment and receive does not, and why MOSI must be a flop.
