Skip to content
VLSI Mentor

SPI · Module 1

The Shift-Register Mental Model

The hardware underneath every SPI transfer: two shift registers wired into a ring, one bit leaving each end and one arriving on every enabled edge. Traced cycle by cycle, built in Verilog, SystemVerilog and VHDL, and separated from the meaning a device layers on top.

Chapter 1.2 settled who is allowed to drive each wire. It deliberately said nothing about where the bit on MOSI comes from, or what happens to the bit arriving on MISO. That is this chapter, and it is the most load-bearing model in the whole curriculum: almost every later result — full duplex, dummy cycles, why a "read" still transmits, why a one-bit shift is the signature of a mode error — is this one picture, applied.

The model is small enough to hold in your head and precise enough to predict real waveforms. Each end of an SPI link holds a register that shifts. The two registers are wired into a ring. Every enabled clock edge pushes one bit out of each register and pulls one bit in. After the register's width in edges, the two ends have swapped contents.

Everything below earns that sentence, then builds it in hardware.

1. The Problem: a Register Is Parallel, a Wire Is Not

Start from the mismatch. Inside a chip, a byte lives in a register — eight flip-flops, read and written all at once across eight parallel nets. On the board between two chips, Chapter 1.1 established there is one data wire in each direction, because width is what costs pins.

So the interface needs a structure that converts between the two: something that accepts a whole word in parallel, emits it one bit at a time, and simultaneously assembles the arriving bits back into a whole word. That structure is a shift register, and it is not a special SPI invention — it is the standard answer to the parallel/serial mismatch, and the reason SPI hardware is small enough to put in a sensor.

2. What a Shift Register Does

A shift register is a row of flip-flops in which each stage's output feeds the next stage's input. On an enabled clock edge, every stage simultaneously takes the value of its neighbour. Three consequences follow, and all three matter.

One bit leaves. The stage at the end of the row has no downstream neighbour inside the register, so its value is available as a serial output. This is the bit driven onto the wire.

One bit enters. The stage at the other end has no upstream neighbour, so its input must come from outside — a serial input. This is the bit captured from the wire.

The width is conserved. The register does not grow or shrink. A bit leaving and a bit entering happen on the same edge, so the register always holds exactly its width in bits — a mixture of what is left of the original word and what has arrived so far.

That third point is where the interesting behaviour comes from, and it is the part newcomers tend not to notice. A shift register is not a queue that drains. It is a fixed-size window that the outgoing word slides out of while the incoming word slides in behind it.

3. Two Registers, One Ring

Now connect two of them, using the ownership rules from Chapter 1.2.

The master's shift register drives MOSI from its serial output. MOSI reaches the selected slave's serial input. The slave's shift register drives MISO from its serial output, and MISO reaches the master's serial input. Both registers are shifted by the same clock — SCLK, which the master owns.

The result is a ring: master → MOSI → slave → MISO → master. Not two channels; one loop.

Two shift registers wired into a ring. The master register's serial output drives MOSI along the upper path to the slave register's serial input. The slave register's serial output drives MISO along the lower path back to the master's serial input. A shared SCLK from the master shifts both registers on the same edge.Master registershifts on every edgeMOSImaster out, slave inSCLKshifts both endsMISOslave out, master inSlave registershifts on every edgeserial outserial inserial outserial in12
Figure 1 — the shift ring. The master's serial output feeds the slave's serial input over MOSI; the slave's serial output feeds the master's serial input over MISO. One clock shifts both registers, so a bit leaves and a bit arrives at each end on the same edge.

Two properties of the ring are worth stating explicitly, because later chapters lean on both.

It is symmetric in structure, asymmetric in control. Both ends hold the same kind of register and do the same thing on each edge. What differs is that one end supplies the clock and the selection — the ownership asymmetry from Chapter 1.2. The data path is a ring between equals; the control path is not.

It conserves bits. Nothing in the ring creates or destroys a bit. Over N enabled edges, exactly N bits cross MOSI and exactly N cross MISO. That conservation is what makes the exchange predictable enough to verify against a reference model.

4. One Edge, Two Halves

Zoom in on a single bit time, because this is where a model that is too coarse starts producing wrong predictions.

A bit time has two jobs to do, and they cannot be simultaneous. The bit being sent has to be presented on the wire and held there long enough for the other end to see it reliably. Then it has to be captured into the receiving register. If both ends tried to present and capture at the same instant, the receiver would be sampling a line that is changing.

So SPI splits the bit time: one clock edge is used to launch new data onto the wire, and the other edge is used to capture it. A bit presented on one edge is sampled by the far end on the next, while the line is stable.

This chapter deliberately stops there. Which edge launches and which captures, what the clock does between transfers, and how those two choices produce four named modes is the subject of Module 2 and Module 3. What you need here is only the structural fact: within one bit time, launch and capture are separated, which is why both directions can move a bit per clock without either end sampling a moving signal.

For the rest of this chapter, "an enabled edge" means one complete bit time — one bit out and one bit in at each end. That abstraction is exactly what the RTL in §7 makes concrete with a shift_en pulse.

5. Eight Edges, Traced

Abstractions are cheap; let us run one. Take an 8-bit exchange. The master has loaded 0xA5 (1010_0101) into its register and the slave has loaded 0x3C (0011_1100) into its. Both shift most-significant bit first — the common arrangement, and one we will interrogate in §6.

Each row shows the state before that edge: what each register holds, and therefore what each end is presenting on its wire.

EdgeMaster register (before)MOSIMISOSlave register (before)
11010_0101100011_1100
20100_1010000111_1001
31001_0100111111_0010
40010_1001011110_0101
50101_0011011100_1010
61010_0111111001_0100
70100_1111000010_1001
81001_1110100101_0010
after 80011_1100 = 0x3C——1010_0101 = 0xA5

Read three things off that table.

The MOSI column, top to bottom, is 1010_0101 — the master's original word, MSB first. Nothing assembled it; it is simply the successive most-significant bits as the word slides out.

The MISO column is 0011_1100 — the slave's original word, MSB first. The same mechanism, in the other direction, on the same edges.

After eight edges the registers hold each other's words. The master's register contains 0x3C and the slave's contains 0xA5. This is the sense in which an SPI transfer is an exchange rather than a send or a receive: the hardware swapped two words, and any interpretation more specific than that is something a device layered on top.

One 8-bit exchange — 0xA5 out, 0x3C in

8 cycles
An eight-bit SPI exchange. SCLK runs for eight bit times. MOSI carries the bit pattern one, zero, one, zero, zero, one, zero, one, which is the byte A5 most significant bit first. MISO simultaneously carries zero, zero, one, one, one, one, zero, zero, which is the byte 3C most significant bit first. Chip select is asserted low across all eight bit times.first bit — both MSBsfirst bit — both MSBseighth bit — registers now swappedeighth bit — registers nowswappedsclkcs_nmosimisobitb7b6b5b4b3b2b1b0t0t1t2t3t4t5t6t7
Figure 2 — the same eight bit times as a waveform. MOSI carries 0xA5 and MISO carries 0x3C, MSB first, on the same edges. The edge arrangement shown is one common choice; Module 3 derives all four.

Note what the waveform does not show: it draws one cycle per bit time rather than separating the launch edge from the capture edge, because that separation belongs to Module 2. Treat each column as one complete bit time.

6. Shift Direction and Which Bit Goes First

The trace above shifted left and emitted the most significant bit first. Both halves of that sentence deserve a moment, because they are the same decision seen from two sides.

Emitting the MSB first requires the serial output to be taken from the register's most-significant stage, and each shift to move every bit one position toward that stage — a left shift with the arriving bit entering at the least-significant end. Emitting the LSB first requires the mirror image: output from the least-significant stage, a right shift, arriving bits entering at the top.

Two observations that save real debugging time later.

The direction is a property of the implementation, not of SPI. MSB-first is much the more common convention and is what most devices specify, but it is a device and controller configuration rather than something the interface fixes. A configurable master implements both — how, and what that costs, is Module 13.

Transmit and receive share the decision. One register does both jobs, so the end that shifts out MSB-first necessarily assembles the incoming word MSB-first too. You do not get to choose independently for each direction on a single register. This is why a bit-order mismatch corrupts both directions at once, and why its waveform signature — a byte arriving bit-reversed rather than shifted — differs from a mode error. Module 4 covers ordering as configuration; Module 18 covers telling the two failure signatures apart.

7. Building It — the Exchange Core in Three HDLs

Now make the model concrete. What follows is a deliberately narrow module: the shift register at one end of the ring, with a shift_en pulse standing in for "one bit time has elapsed."

Architecture

One register of WIDTH flip-flops. Its most-significant stage drives the serial output sdo. On an enabled edge it shifts left, with the serial input sdi entering at the least-significant end. A separate load input captures a parallel word to be sent, and the register contents are continuously visible as rx_data.

State. Exactly one register, shreg. There is no separate transmit and receive register — and that is the point of the chapter. tx_data is what you put in; rx_data is what has arrived; they are the same flip-flops at different moments.

Combinational logic. Two assignments, both trivial: sdo is the top bit of shreg, and rx_data is shreg itself. All the behaviour lives in the clocked block.

Clocking. Everything advances on the rising edge of clk. Note carefully that clk here is the system clock, not SCLK: shift_en is a one-cycle pulse that a real master would generate at the appropriate SCLK edge. Structuring it this way keeps the example honest about what it is — the datapath, not the clock generator.

Reset. Asynchronous, active-low, clearing the register to zero. Chosen for familiarity rather than because SPI requires it; reset strategy for a real SPI block is Module 13 for a master and Module 14 for a slave, where it interacts with CS and an externally supplied clock.

Ownership. sdi is an input, sdo an unconditional output. There is no output enable here — that belongs to the slave's MISO path, built in Module 14, for exactly the reasons Chapter 1.2 gave.

Timing assumption. That shift_en is asserted for exactly one clk cycle per bit time, and that sdi is stable when it is. A real implementation earns that by generating shift_en from the correct SCLK edge.

Limitation — read this before reusing the code. This is an educational fragment, not SPI IP. It has no SCLK generation, no clock divider, no CS handling, no bit counter or transfer-complete signal, no mode configuration, no bit-order option, and no clock-domain crossing for the slave case. Adding those is precisely the work of Modules 13 to 15. What it does have is the exchange mechanism, isolated so you can see it.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_shift_core.sv — one end of the shift ring: MSB out, LSB in, one register doing both
   // WIDTH must be >= 2: the shift concatenation slices [WIDTH-2:0].
   module spi_shift_core #(
       parameter int WIDTH = 8
   ) (
       input  logic             clk,       // SYSTEM clock, not SCLK
       input  logic             rst_n,     // asynchronous, active-low
       input  logic             load,      // capture tx_data into the register
       input  logic [WIDTH-1:0] tx_data,   // word to send
       input  logic             shift_en,  // one pulse = one bit time
       input  logic             sdi,       // serial in  — the far end's sdo
       output logic             sdo,       // serial out — this end's MSB
       output logic [WIDTH-1:0] rx_data    // what has shifted in so far
   );
       logic [WIDTH-1:0] shreg;

       always_ff @(posedge clk or negedge rst_n) begin
           if (!rst_n)        shreg <= '0;
           else if (load)     shreg <= tx_data;                    // load wins over shift
           else if (shift_en) shreg <= {shreg[WIDTH-2:0], sdi};    // MSB out, sdi in at LSB
       end

       assign sdo     = shreg[WIDTH-1];   // the bit currently presented on the wire
       assign rx_data = shreg;            // same flip-flops, read as the received word
   endmodule

The testbench is the chapter's claim made executable: instantiate two cores, wire them into the ring of Figure 1, and check both that each wire carries the right bit at the right time and that the two words end up swapped. It also holds shift_en low for a cycle mid-transfer, because a paused clock must not corrupt the exchange — the SPI reality that SCLK need not be continuous.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_shift_core_tb.sv — two cores wired into a ring; self-checks the bit stream, a paused bit time, and the swap
   module spi_shift_core_tb;
       localparam int WIDTH = 8;
       localparam logic [WIDTH-1:0] M_TX = 8'hA5;   // 1010_0101
       localparam logic [WIDTH-1:0] S_TX = 8'h3C;   // 0011_1100

       logic clk = 1'b0, rst_n, load, shift_en;
       logic m_sdo, s_sdo;
       logic [WIDTH-1:0] m_rx, s_rx;
       int   errors = 0;

       // The ring: master sdo -> slave sdi is MOSI; slave sdo -> master sdi is MISO.
       spi_shift_core #(.WIDTH(WIDTH)) master (
           .clk(clk), .rst_n(rst_n), .load(load), .tx_data(M_TX),
           .shift_en(shift_en), .sdi(s_sdo), .sdo(m_sdo), .rx_data(m_rx));

       spi_shift_core #(.WIDTH(WIDTH)) slave (
           .clk(clk), .rst_n(rst_n), .load(load), .tx_data(S_TX),
           .shift_en(shift_en), .sdi(m_sdo), .sdo(s_sdo), .rx_data(s_rx));

       always #5 clk = ~clk;

       initial begin
           rst_n = 1'b0; load = 1'b0; shift_en = 1'b0;
           @(posedge clk); #1; rst_n = 1'b1;
           @(posedge clk); #1; load  = 1'b1;     // both ends capture their word
           @(posedge clk); #1; load  = 1'b0;

           for (int i = 0; i < WIDTH; i++) begin
               // Before the edge, each end presents its next bit, MSB-first.
               if (m_sdo !== M_TX[WIDTH-1-i]) begin
                   $error("bit %0d: MOSI=%b expected %b", i, m_sdo, M_TX[WIDTH-1-i]); errors++;
               end
               if (s_sdo !== S_TX[WIDTH-1-i]) begin
                   $error("bit %0d: MISO=%b expected %b", i, s_sdo, S_TX[WIDTH-1-i]); errors++;
               end

               // Boundary case: a bit time with no shift_en must change nothing.
               if (i == 3) begin
                   @(posedge clk); #1;
                   if (m_sdo !== M_TX[WIDTH-1-i] || s_sdo !== S_TX[WIDTH-1-i]) begin
                       $error("paused cycle shifted the registers"); errors++;
                   end
               end

               shift_en = 1'b1; @(posedge clk); #1; shift_en = 1'b0;
           end

           // The whole point: after WIDTH bit times the words have swapped.
           if (m_rx !== S_TX) begin $error("master rx=%h expected %h", m_rx, S_TX); errors++; end
           if (s_rx !== M_TX) begin $error("slave  rx=%h expected %h", s_rx, M_TX); errors++; end

           if (errors == 0) $display("PASS: MOSI=%h, MISO=%h, registers exchanged", M_TX, S_TX);
           else             $display("FAIL: %0d mismatches", errors);
           $finish;
       end
   endmodule

The Verilog form is the same hardware with reg/wire typing, {WIDTH{1'b0}} for the reset constant, and $display in place of $error — the latter is a SystemVerilog task, so a Verilog-2001 toolchain will not accept it.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_shift_core.v — the same exchange core in Verilog-2001
   // WIDTH must be >= 2: the shift concatenation slices [WIDTH-2:0].
   module spi_shift_core #(
       parameter WIDTH = 8
   ) (
       input                  clk,        // SYSTEM clock, not SCLK
       input                  rst_n,      // asynchronous, active-low
       input                  load,
       input  [WIDTH-1:0]     tx_data,
       input                  shift_en,   // one pulse = one bit time
       input                  sdi,
       output                 sdo,
       output [WIDTH-1:0]     rx_data
   );
       reg [WIDTH-1:0] shreg;

       always @(posedge clk or negedge rst_n) begin
           if (!rst_n)        shreg <= {WIDTH{1'b0}};
           else if (load)     shreg <= tx_data;
           else if (shift_en) shreg <= {shreg[WIDTH-2:0], sdi};
       end

       assign sdo     = shreg[WIDTH-1];
       assign rx_data = shreg;
   endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_shift_core_tb.v — Verilog-2001 self-checking ring: same three checks, integer loop, $display reporting
   module spi_shift_core_tb;
       parameter WIDTH = 8;
       localparam [WIDTH-1:0] M_TX = 8'hA5;
       localparam [WIDTH-1:0] S_TX = 8'h3C;

       reg  clk = 1'b0;
       reg  rst_n, load, shift_en;
       wire m_sdo, s_sdo;
       wire [WIDTH-1:0] m_rx, s_rx;
       integer i;
       integer errors = 0;

       spi_shift_core #(.WIDTH(WIDTH)) master (
           .clk(clk), .rst_n(rst_n), .load(load), .tx_data(M_TX),
           .shift_en(shift_en), .sdi(s_sdo), .sdo(m_sdo), .rx_data(m_rx));

       spi_shift_core #(.WIDTH(WIDTH)) slave (
           .clk(clk), .rst_n(rst_n), .load(load), .tx_data(S_TX),
           .shift_en(shift_en), .sdi(m_sdo), .sdo(s_sdo), .rx_data(s_rx));

       always #5 clk = ~clk;

       initial begin
           rst_n = 1'b0; load = 1'b0; shift_en = 1'b0;
           @(posedge clk); #1; rst_n = 1'b1;
           @(posedge clk); #1; load  = 1'b1;
           @(posedge clk); #1; load  = 1'b0;

           for (i = 0; i < WIDTH; i = i + 1) begin
               if (m_sdo !== M_TX[WIDTH-1-i]) begin
                   $display("ERROR bit %0d: MOSI=%b expected %b", i, m_sdo, M_TX[WIDTH-1-i]);
                   errors = errors + 1;
               end
               if (s_sdo !== S_TX[WIDTH-1-i]) begin
                   $display("ERROR bit %0d: MISO=%b expected %b", i, s_sdo, S_TX[WIDTH-1-i]);
                   errors = errors + 1;
               end

               if (i == 3) begin
                   @(posedge clk); #1;
                   if (m_sdo !== M_TX[WIDTH-1-i] || s_sdo !== S_TX[WIDTH-1-i]) begin
                       $display("ERROR: paused cycle shifted the registers");
                       errors = errors + 1;
                   end
               end

               shift_en = 1'b1; @(posedge clk); #1; shift_en = 1'b0;
           end

           if (m_rx !== S_TX) begin
               $display("ERROR master rx=%h expected %h", m_rx, S_TX); errors = errors + 1;
           end
           if (s_rx !== M_TX) begin
               $display("ERROR slave  rx=%h expected %h", s_rx, M_TX); errors = errors + 1;
           end

           if (errors == 0) $display("PASS: MOSI=%h, MISO=%h, registers exchanged", M_TX, S_TX);
           else             $display("FAIL: %0d mismatches", errors);
           $finish;
       end
   endmodule

The VHDL form uses a std_logic_vector for the register and expresses the shift as a concatenation of a slice with the incoming bit — the same operation the two other languages perform, written with VHDL's downto indexing.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_shift_core.vhd — the same exchange core in VHDL
   library ieee;
   use ieee.std_logic_1164.all;

   entity spi_shift_core is
       generic (
           WIDTH : positive := 8               -- WIDTH must be >= 2
       );
       port (
           clk      : in  std_logic;           -- SYSTEM clock, not SCLK
           rst_n    : in  std_logic;           -- asynchronous, active-low
           load     : in  std_logic;
           tx_data  : in  std_logic_vector(WIDTH-1 downto 0);
           shift_en : in  std_logic;           -- one pulse = one bit time
           sdi      : in  std_logic;
           sdo      : out std_logic;
           rx_data  : out std_logic_vector(WIDTH-1 downto 0)
       );
   end entity spi_shift_core;

   architecture rtl of spi_shift_core is
       signal shreg : std_logic_vector(WIDTH-1 downto 0);
   begin
       shift_proc : process (clk, rst_n)
       begin
           if rst_n = '0' then
               shreg <= (others => '0');
           elsif rising_edge(clk) then
               if load = '1' then
                   shreg <= tx_data;                              -- load wins over shift
               elsif shift_en = '1' then
                   shreg <= shreg(WIDTH-2 downto 0) & sdi;        -- MSB out, sdi in at LSB
               end if;
           end if;
       end process shift_proc;

       sdo     <= shreg(WIDTH-1);
       rx_data <= shreg;
   end architecture rtl;
Azvya Education Pvt. Ltd.VLSI Mentor
spi_shift_core_tb.vhd — VHDL self-checking ring: bit stream, paused bit time, and the swap
   library ieee;
   use ieee.std_logic_1164.all;

   entity spi_shift_core_tb is
   end entity spi_shift_core_tb;

   architecture sim of spi_shift_core_tb is
       constant WIDTH : positive := 8;
       constant M_TX  : std_logic_vector(WIDTH-1 downto 0) := "10100101";  -- 0xA5
       constant S_TX  : std_logic_vector(WIDTH-1 downto 0) := "00111100";  -- 0x3C
       constant TP    : time := 10 ns;

       signal clk      : std_logic := '0';
       signal rst_n    : std_logic := '0';
       signal load     : std_logic := '0';
       signal shift_en : std_logic := '0';
       signal m_sdo, s_sdo : std_logic;
       signal m_rx, s_rx   : std_logic_vector(WIDTH-1 downto 0);
       signal done     : boolean := false;
   begin
       clk <= not clk after TP/2 when not done else '0';

       -- The ring: master sdo -> slave sdi (MOSI); slave sdo -> master sdi (MISO).
       master : entity work.spi_shift_core
           generic map (WIDTH => WIDTH)
           port map (clk => clk, rst_n => rst_n, load => load, tx_data => M_TX,
                     shift_en => shift_en, sdi => s_sdo, sdo => m_sdo, rx_data => m_rx);

       slave : entity work.spi_shift_core
           generic map (WIDTH => WIDTH)
           port map (clk => clk, rst_n => rst_n, load => load, tx_data => S_TX,
                     shift_en => shift_en, sdi => m_sdo, sdo => s_sdo, rx_data => s_rx);

       stim : process
           variable errors : natural := 0;
       begin
           wait until rising_edge(clk); wait for 1 ns; rst_n <= '1';
           wait until rising_edge(clk); wait for 1 ns; load  <= '1';
           wait until rising_edge(clk); wait for 1 ns; load  <= '0';

           for i in 0 to WIDTH-1 loop
               if m_sdo /= M_TX(WIDTH-1-i) then
                   report "MOSI mismatch at bit " & integer'image(i) severity error;
                   errors := errors + 1;
               end if;
               if s_sdo /= S_TX(WIDTH-1-i) then
                   report "MISO mismatch at bit " & integer'image(i) severity error;
                   errors := errors + 1;
               end if;

               -- Boundary case: a bit time with no shift_en must change nothing.
               if i = 3 then
                   wait until rising_edge(clk); wait for 1 ns;
                   if m_sdo /= M_TX(WIDTH-1-i) or s_sdo /= S_TX(WIDTH-1-i) then
                       report "paused cycle shifted the registers" severity error;
                       errors := errors + 1;
                   end if;
               end if;

               shift_en <= '1';
               wait until rising_edge(clk); wait for 1 ns;
               shift_en <= '0';
           end loop;

           if m_rx /= S_TX then
               report "master rx mismatch" severity error; errors := errors + 1;
           end if;
           if s_rx /= M_TX then
               report "slave rx mismatch"  severity error; errors := errors + 1;
           end if;

           if errors = 0 then
               report "PASS: registers exchanged" severity note;
           else
               report "FAIL: " & integer'image(errors) & " mismatches" severity error;
           end if;

           done <= true;
           wait;
       end process stim;
   end architecture sim;

What the three versions agree on, and where they differ

All three infer the same hardware: WIDTH flip-flops with an asynchronous active-low reset, a load path, and a shift path taking the top bit out and the serial input in at the bottom. The differences are linguistic rather than structural — logic and always_ff versus reg and always, '0 versus {WIDTH{1'b0}} versus (others => '0'), and the concatenation operator {a, b} versus a & b. The SystemVerilog version states its intent to the tool (always_ff is a promise that this block infers flip-flops); the Verilog version relies on the coding pattern alone; the VHDL version makes the reset structure explicit in the if/elsif rising_edge shape.

One genuine semantic note on the testbenches: $error and assert are SystemVerilog, so the Verilog testbench counts errors and reports with $display, while VHDL uses report … severity. All three are self-checking and print a single PASS or FAIL line.

8. Why This Model Explains Full Duplex

You now have everything needed to see why SPI cannot help being full duplex, so it is worth stating before the next chapter develops it.

There is one register per end, and it shifts. A bit leaves and a bit arrives on the same edge because those are the two ends of one shift operation, not two separate features. A design that wanted to transmit without receiving would have to actively discard what shifted in — and would still have shifted it in.

That is the whole mechanism. Chapter 1.4 takes it to its consequences: why a device "read" still transmits bits, where dummy bytes come from, why the received bytes during a command phase are meaningless but present, and why a transaction object in a testbench naturally holds both directions rather than being typed read or write.

9. What This Model Does Not Tell You

A good mental model is defined as much by its boundaries as by its content. This one is silent on four things, and assuming otherwise causes real errors.

It does not say which edge does what. Launch and capture are separated within a bit time (§4), but which physical edge plays which role is configuration — Modules 2 and 3.

It does not say how many bits a transfer contains. The register has a width; a transaction may be many registers' worth, and what delimits it is CS, not the register. Chapter 1.2 established CS as the boundary; Module 4 covers transfer width and framing.

It does not say what the bits mean. The ring swaps words. That a particular outgoing word is an opcode and a particular incoming word is a status byte is a device contract, from the datasheet — the ownership-versus-meaning separation from Chapter 1.2, applied again.

It does not describe a real slave's clocking. The RTL above shifts on a system clock with an enable. A slave has no system clock of its own relative to SCLK — the clock arrives from outside, which turns this simple register into a clock-domain problem. That is Module 15, and it is the single biggest gap between this fragment and a working slave.

10. Why an RTL Designer Cares

This chapter is the datapath of every SPI block you will ever write, and three design consequences follow directly from it.

Transmit and receive are one register, not two. The strong temptation when starting an SPI master is to build a TX shift register and a separate RX shift register. It works, and it costs twice the flip-flops for no benefit in the common case, because the outgoing word has vacated a bit position exactly as an incoming bit needs one. Knowing why one register suffices is the difference between copying a design and choosing one. Where two registers genuinely help — decoupling the next transmit word from the received one so software can be slower than the bus — is a buffering decision, not a shift decision, and it is Module 13.

The register is the natural place for width and order configuration. Transfer width becomes the count of enabled edges and the register's parameterisation; bit order becomes the choice of which stage drives sdo and which direction the shift runs. Both are small changes to this core rather than new blocks — which is why a parameterised core is worth building once.

shift_en is where the real work goes. Notice how little logic the core needed. In a real master, essentially all the remaining complexity is in producing shift_en at the right moments: dividing the system clock to SCLK, placing launch and capture on the configured edges, counting bits, and framing with CS. This chapter deliberately handed you the easy half so that Module 13 can spend its time on the hard half.

11. Why a Verification Engineer Cares

The shift model is what makes SPI checkable at all, because it supplies a reference model that needs no DUT internals.

A monitor can reconstruct both words from the pins alone. Sample MOSI and MISO on the capture edges between a CS assertion and its deassertion, assemble each into a word using the configured bit order, and you have the complete exchange — without probing a single internal register. That property is what makes a passive SPI monitor possible, and it is a direct consequence of the ring conserving bits.

The strongest early check is a structural one. For a transfer of N bit times, exactly N bits must cross in each direction. A monitor that counts capture edges between CS edges and compares against the configured width catches short frames, extra clocks and dropped edges — a whole family of bugs — before any data comparison happens. That is a bit-count assertion, and it is cheap.

A scoreboard must expect an exchange, not a transfer. Because every transfer produces a received word, a scoreboard that models only the direction the test cared about will either ignore real evidence or trip over meaningless bytes. The right model is the one the hardware implements: two words in, two words out, with the device's contract deciding which of them carry meaning. Chapter 1.4 shows what that does to a transaction object, and Modules 16 and 17 build the environment properly.

12. Why an FPGA Engineer Cares

Two practical points, both consequences of the model rather than restatements of it.

The datapath is nearly free; the control is not. A shift register of a byte or two is a handful of flip-flops and no arithmetic — it will meet timing on any device you are likely to use. When an SPI block fails timing on an FPGA, the failure is essentially never in the shifting. It is in the clock path, the I/O registers, or the path from an externally supplied SCLK into the fabric — which is why Module 15 exists and why this chapter's RTL deliberately abstracts SCLK behind shift_en.

Where you place the capture register decides your margin. The model says a bit is captured at some point in the bit time; the physical question is how much delay sits between the pin and the capturing flip-flop. Capturing in an I/O-block register places that flip-flop as close to the pad as the device allows and gives a predictable, constrainable path; capturing deep in the fabric adds routing delay that varies with each build. The model does not care. Timing closure does.

13. Common Misconceptions

14. Reason It Through

Work this before reading the answer. It is the kind of question a bring-up log actually poses.

A master and a slave are exchanging 8-bit words. The master intends to send 0xA5 and the slave intends to return 0x3C, exactly as in §5. On the logic analyser, MOSI shows 0xA5 correctly, but the master's received byte is 0x78 instead of 0x3C — and MISO on the capture shows the slave's bits are correct. What single mechanism explains this?

First, what is ruled out. MOSI is correct, so the master's shift-out path, its bit order and its clocking of the outgoing bit are all fine. MISO on the wire is correct, so the slave produced the right bits at the right times. The fault is therefore in what the master does with MISO — the capture side only.

Compare the words bit by bit. Expected 0x3C is 0011_1100. Observed 0x78 is 0111_1000. The observed word is the expected word shifted left by one, with a zero entering at the bottom. The master assembled eight bits, but they are offset by one position relative to the slave's stream.

What mechanism produces exactly a one-position offset? The master's register performed one shift too many relative to the arriving data — equivalently, it began capturing one bit time before the slave began presenting, so its first captured bit was not the slave's first data bit and the last real bit fell off the top. In the language of this chapter, the two registers were not stepping in lockstep: the ring conserves bits only if both ends agree on which edge is bit one.

What would cause that in practice? The common causes are a capture edge chosen one half-cycle early relative to the slave's launch edge — a clock-phase disagreement — or a master that starts shifting on the CS assertion edge rather than on the first clock. Both are real, and both produce this identical signature.

Why does this chapter stop here rather than name the culprit? Because distinguishing them requires the launch/capture edge vocabulary that Module 2 builds and the mode derivation that Module 3 supplies. What the shift model gives you is the more valuable half of the diagnosis: the failure is a position error, not a data error, so look at edge alignment and not at either device's data path. That narrowing is the reasoning skill — Module 18 turns it into a method.

15. Understanding Check

16. Summary

A shift register resolves the mismatch between a parallel register inside a chip and a single wire between chips. Each end of an SPI link holds one, and they are wired into a ring: the master's serial output drives MOSI into the slave's serial input, the slave's serial output drives MISO back into the master's, and SCLK shifts both.

On each enabled edge, one bit leaves and one bit arrives at every end, because those are the two ends of a single shift rather than two separate operations. The register never empties — it is a fixed-width window holding what remains of the outgoing word plus what has arrived of the incoming one. After the width in edges, the two ends have exchanged words: in §5's trace, 0xA5 and 0x3C swap places, with MOSI carrying one and MISO the other, MSB first.

The RTL makes the claim testable in eleven lines of logic: one register, a load path, a shift path taking the top bit out and the serial input in at the bottom. Two of them wired into a ring provably exchange their contents — and the smallness of that datapath is itself the lesson, because it tells you that the real work in an SPI block is generating the shift enable at the right moments, not shifting.

The model's boundaries matter as much as its content: it says nothing about which edge launches or captures, how many bits a transaction contains, what the bits mean, or how a slave copes with a clock it does not own. Each of those is a later module. What it does give you is a prediction of the contents of both registers at every instant of a transfer — which is exactly what you need to read a waveform, write a reference model, or recognise that a byte is offset rather than wrong.

17. What Comes Next

The ring makes one consequence unavoidable, and Chapter 1.4 — Full-Duplex Exchange is about facing it: every transfer moves information both ways, whether or not the software wanted it to. That chapter explains where dummy transmit bytes come from, why the bytes received during a command phase exist but mean nothing, and why "read" and "write" are interpretations a device imposes rather than modes the hardware has.

Browse the path on the SPI curriculum index, or revisit Master, Slave, and Signal Ownership for the driver rules this chapter's ring depends on. For the same shift-register structure solving the asynchronous version of the problem — no shared clock, timing recovered from the data — see What a UART Actually Is.

Continue learning