SPI · Module 4
Dummy Phases and Read Latency
Why a device needs turnaround before it can answer, why dummy is counted in clock cycles rather than bytes, how its length grows with frequency, and the one-byte data offset a mismatch produces.
Chapter 4.4 drew a read as opcode, address, then data — with the device answering immediately after the final address byte. Real devices frequently cannot.
The master has finished sending the address and starts clocking for data. Where does the device get that first bit from, and what if it is not ready?
The answer is that it does not get it from anywhere, and so the transaction must contain a phase whose entire purpose is to pass time. That phase is where a surprising amount of SPI's practical complexity lives.
1. Why a Device Cannot Answer Immediately
Consider the instant the last address bit is clocked in. From the master's perspective the address is complete and the next edge should produce data. From the device's perspective, at that same instant:
- The address has only just finished arriving — the final bit landed on this very edge.
- Nothing has been fetched yet, because until this edge the address was incomplete.
- Its output register is empty, and the next edge is already coming.
Chapter 2.6 established that a slave must present a valid bit within its clock-to-output delay after the launch edge. That budget assumes the data exists. Here it does not: it has to be fetched from a memory array, and that takes time the bus has not allowed for.
So the device needs the master to keep clocking without expecting meaningful data. Those edges are the dummy phase. They carry nothing, and their only function is to let a fixed amount of wall-clock time elapse.
2. Two Different Reasons for Turnaround
Dummy cycles are usually explained with one reason. There are two, they are independent, and conflating them causes real confusion when reading datasheets.
Access time. The device must physically fetch data — sense amplifiers, an array access, an internal pipeline. This is a delay measured in nanoseconds, is a property of the silicon, and does not change when you change the clock rate.
Bus turnaround. On a shared bidirectional data line — the Dual and Quad modes of Module 12 — the same wire that carried the address inbound must now carry data outbound. Both ends must release and re-drive it, and a gap is required so they never drive simultaneously (Chapter 1.2's contention, arriving as a protocol requirement).
On standard four-wire SPI with separate MOSI and MISO there is no turnaround requirement at all — the lines are unidirectional and nothing has to switch direction. Every dummy cycle on a standard SPI read is buying access time. That is why the conventional low-speed read command has no dummy phase while the fast read does: it is the same array, and the difference is entirely how much time a clock period provides.
3. Dummy Is Counted in Cycles, Not Bytes
This follows directly, and it is the detail most often got wrong in a first implementation.
If a device needs 20 ns of access time and the clock period is 10 ns, it needs two cycles — not one byte, not "a bit more than none". Rounding that up to a whole byte would waste six cycles on every read, and rounding down would break the device.
So a dummy phase of 6, 8 or 10 cycles is entirely normal, and 8 is common only because it happens to coincide with a byte on a byte-oriented controller. Several widely-used devices specify dummy counts that are not multiples of eight, and a driver that can only express turnaround in whole bytes cannot talk to them correctly.
The arithmetic is just the conversion:
dummy_cycles = ceil( t_access / T_sclk )
Hypothetical device: t_access = 24 ns
SCLK 20 MHz → T = 50.0 ns → ceil(24/50.0) = 1 cycle
SCLK 50 MHz → T = 20.0 ns → ceil(24/20.0) = 2 cycles
SCLK 104 MHz → T = 9.6 ns → ceil(24/9.6) = 3 cyclesTwo consequences worth holding onto. More dummy cycles are needed as the clock goes faster, which is counter-intuitive until you see that the device's requirement is fixed in nanoseconds while each cycle buys less time. And the dummy count is a property of a frequency range, not of a command — which is why it is a configuration register rather than a constant on parts designed to run over a wide range.
4. The Transaction, With Turnaround
5. The Phases on the Wire
The dummy phase, and what MISO does during it
9 cyclesThe dummy phase occupies one byte time in this figure only because the example uses eight dummy cycles. Drawn at six or ten cycles it would not align to a byte boundary at all — which is §3's point made visible.
6. What Turnaround Costs
The dummy phase is pure overhead, and quantifying it explains a great deal about how SPI is used in practice.
For a fast read with a 1-byte opcode, a 3-byte address and 8 dummy cycles, the edges spent before the first data bit are:
opcode 8 cycles
address 24 cycles
dummy 8 cycles
─────────────────────
overhead 40 cycles, every readEfficiency — data bits as a fraction of all clocked bits — therefore depends entirely on how much data follows:
payload total cycles efficiency
1 byte 40 + 8 = 48 16.7 %
4 bytes 40 + 32 = 72 44.4 %
16 bytes 40 + 128 = 168 76.2 %
256 bytes 40 + 2048 = 2088 98.1 %Read one byte at a time and five sixths of the bus is overhead. The same bus reading 256-byte bursts is 98% efficient. That ratio is why sequential-read and continuous-read modes exist, and it is the subject of Module 7.
It also reframes clock rate. Doubling SCLK from 50 to 100 MHz does not double the throughput of single-byte reads: the overhead doubles in frequency too, and the dummy count may increase by a cycle or two as §3 showed. Burst length buys more than clock rate for small transfers, which is a genuinely useful thing to know before optimising the wrong parameter.
7. Building the Dummy Phase — Three HDLs
The circuit
Circuit. Chapter 4.4's phase sequencer with one additional state and one additional counter.
State. The phase, the latched opcode, the accumulating address, the address-byte count, and a dummy-cycle counter.
Datapath. Unchanged from Chapter 4.4. The dummy counter is control only — it moves no data, which is the point of the phase.
Control. Two inputs now advance the machine: byte_done in the command and address phases, and bit_stb — one pulse per SPI bit — in the dummy phase. That difference is §3 expressed directly in the RTL: the dummy phase counts a different unit from every other phase, so it consumes a different strobe.
Clock and reset. As before: system clock, asynchronous active-low reset, cs_n forcing IDLE unconditionally.
Enables. dummy_cycles is an input rather than a parameter, because the required count varies with clock frequency and real parts make it programmable.
Timing. drive_miso is a decode of the registered phase, so the output enable asserts one cycle after the transition into the data phase. On a real slave that cycle is part of the clock-to-output budget and must be accounted for, not discovered.
Synthesis. Three state bits, a five-bit counter, a comparator against a configuration register, and the Chapter 4.4 logic unchanged. The comparator now has a register rather than a constant on one input, so it cannot be folded away — the runtime-configurability cost Chapter 4.2 §8 described, appearing again.
Limitations. One dummy length shared by all commands that have one, and a three-command decode. A real device maps different dummy counts to different commands, which is a table lookup rather than a new idea.
// spi_dummy_phase.sv — Chapter 4.4's sequencer with a dummy phase added.
//
// The dummy phase is counted in SCLK CYCLES, not bytes. That is not a detail:
// a device buys a fixed amount of TIME, and time is measured in clock periods.
// Expressing it in bytes would only work when the requirement happened to be
// a multiple of eight, and would break the moment the clock rate changed.
//
// `dummy_cycles` is an input rather than a parameter because real devices make
// it programmable: a faster SCLK needs more cycles to cover the same delay.
module spi_dummy_phase #(
parameter int ADDR_BYTES = 3
) (
input logic clk,
input logic rst_n,
input logic cs_n,
input logic bit_stb, // one pulse per SPI bit
input logic byte_done, // one pulse per completed byte
input logic [7:0] rx_byte,
input logic [4:0] dummy_cycles, // 0..31, runtime configurable
output logic [2:0] phase,
output logic [7:0] cmd,
output logic [8*ADDR_BYTES-1:0] addr,
output logic in_data,
output logic drive_miso // output enable for the slave
);
localparam logic [2:0] PH_IDLE = 3'd0,
PH_CMD = 3'd1,
PH_ADDR = 3'd2,
PH_DUMMY = 3'd3,
PH_DATA = 3'd4;
// Representative serial-flash conventions, not SPI definitions.
localparam logic [7:0] CMD_READ = 8'h03, // no dummy: low-speed read
CMD_FAST_READ = 8'h0B, // dummy phase, then data
CMD_RD_STATUS = 8'h05; // no address, no dummy
localparam int CW = (ADDR_BYTES > 1) ? $clog2(ADDR_BYTES) : 1;
logic [CW-1:0] addr_cnt;
logic [4:0] dummy_cnt;
function automatic logic has_addr(input logic [7:0] op);
return (op == CMD_READ) || (op == CMD_FAST_READ);
endfunction
// Only the fast read pays for turnaround; the slow read is specified at a
// low enough clock that the device can answer immediately.
function automatic logic has_dummy(input logic [7:0] op);
return (op == CMD_FAST_READ);
endfunction
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
phase <= PH_IDLE;
cmd <= 8'h00;
addr <= '0;
addr_cnt <= '0;
dummy_cnt <= 5'd0;
end else if (cs_n) begin
phase <= PH_IDLE;
addr_cnt <= '0;
dummy_cnt <= 5'd0;
end else begin
case (phase)
PH_IDLE: begin
phase <= PH_CMD;
addr_cnt <= '0;
end
PH_CMD: if (byte_done) begin
cmd <= rx_byte;
if (has_addr(rx_byte)) begin
phase <= PH_ADDR;
addr_cnt <= '0;
end else begin
phase <= PH_DATA;
end
end
PH_ADDR: if (byte_done) begin
addr <= {addr[8*ADDR_BYTES-9:0], rx_byte};
if (addr_cnt == CW'(ADDR_BYTES - 1)) begin
dummy_cnt <= 5'd0;
// A zero-length dummy phase must be skipped entirely,
// not entered and exited -- entering would consume one
// cycle that belongs to the data phase.
// `cmd` was registered back in PH_CMD, so it already
// holds this transaction's opcode.
if (has_dummy(cmd) && dummy_cycles != 5'd0)
phase <= PH_DUMMY;
else
phase <= PH_DATA;
end else begin
addr_cnt <= addr_cnt + 1'b1;
end
end
PH_DUMMY: if (bit_stb) begin
if (dummy_cnt == dummy_cycles - 5'd1) phase <= PH_DATA;
else dummy_cnt <= dummy_cnt + 5'd1;
end
PH_DATA: ; // remain until CS rises
default: phase <= PH_IDLE;
endcase
end
end
assign in_data = (phase == PH_DATA);
// The slave must not drive MISO before the data phase -- everything
// earlier belongs to the master's half of the exchange.
assign drive_miso = (phase == PH_DATA);
endmoduleOne decision in that code is worth defending explicitly.
A zero-length dummy phase is skipped, not entered. The transition out of PH_ADDR checks dummy_cycles != 0 and goes straight to PH_DATA when it is zero. Entering PH_DUMMY and leaving it on the next strobe would consume one cycle that belongs to the data phase, producing an output offset by exactly one bit — a bug that appears only in the zero-dummy configuration and is invisible in every other. It is the classic off-by-one of this design, which is why the testbenches check dummy_cycles = 0 as a distinct case.
// spi_dummy_phase_tb.sv — dummy lengths, a zero-length dummy, and an abort.
`timescale 1ns/1ps
module spi_dummy_phase_tb;
logic clk = 0, rst_n = 0;
always #5 clk = ~clk;
localparam int AB = 3;
logic cs_n = 1, bit_stb = 0, byte_done = 0;
logic [7:0] rx_byte = 8'h00;
logic [4:0] dummy_cycles = 5'd8;
logic [2:0] phase; logic [7:0] cmd; logic [8*AB-1:0] addr;
logic in_data, drive_miso;
spi_dummy_phase #(.ADDR_BYTES(AB)) dut (
.clk, .rst_n, .cs_n, .bit_stb, .byte_done, .rx_byte, .dummy_cycles,
.phase, .cmd, .addr, .in_data, .drive_miso);
localparam logic [2:0] PH_IDLE = 3'd0, PH_CMD = 3'd1, PH_ADDR = 3'd2,
PH_DUMMY = 3'd3, PH_DATA = 3'd4;
int errors = 0;
task automatic chk(input string what, input int got, input int exp);
if (got !== exp) begin $display("FAIL %s: got %0d exp %0d", what, got, exp); errors++; end
endtask
// One SPI bit. A byte is eight of these, with byte_done on the last.
task automatic one_bit(input logic last);
@(negedge clk); bit_stb = 1; byte_done = last;
@(negedge clk); bit_stb = 0; byte_done = 0;
@(negedge clk);
endtask
task automatic send_byte(input logic [7:0] b);
rx_byte = b;
for (int i = 0; i < 8; i++) one_bit(i == 7);
endtask
task automatic open_frame(); cs_n = 0; @(negedge clk); @(negedge clk); endtask
task automatic close_frame(); cs_n = 1; @(negedge clk); @(negedge clk); endtask
// Count bits spent in the dummy phase, so its LENGTH is actually checked.
int dummy_bits;
always @(posedge clk) if (rst_n && bit_stb && phase == PH_DUMMY) dummy_bits++;
initial begin
repeat (3) @(negedge clk); rst_n = 1; @(negedge clk);
// --- FAST READ 0x0B with 8 dummy cycles ---
dummy_cycles = 5'd8; dummy_bits = 0;
open_frame();
send_byte(8'h0B);
chk("fast read: phase after opcode", phase, PH_ADDR);
send_byte(8'h12); send_byte(8'h34); send_byte(8'h56);
chk("fast read: enters DUMMY", phase, PH_DUMMY);
chk("fast read: addr", addr, 24'h123456);
chk("fast read: MISO not driven yet", drive_miso, 0);
// Exactly eight bit strobes must pass before DATA.
for (int i = 0; i < 8; i++) one_bit(1'b0);
chk("fast read: dummy length", dummy_bits, 8);
chk("fast read: now DATA", phase, PH_DATA);
chk("fast read: MISO driven", drive_miso, 1);
close_frame();
// --- The same command with a DIFFERENT programmed dummy length ---
dummy_cycles = 5'd4; dummy_bits = 0;
open_frame();
send_byte(8'h0B); send_byte(8'h00); send_byte(8'h00); send_byte(8'h10);
chk("dummy=4: enters DUMMY", phase, PH_DUMMY);
for (int i = 0; i < 4; i++) one_bit(1'b0);
chk("dummy=4: dummy length", dummy_bits, 4);
chk("dummy=4: now DATA", phase, PH_DATA);
close_frame();
// --- READ 0x03: address, but NO dummy phase at all ---
dummy_cycles = 5'd8; dummy_bits = 0;
open_frame();
send_byte(8'h03); send_byte(8'hAB); send_byte(8'hCD); send_byte(8'hEF);
chk("slow read: straight to DATA", phase, PH_DATA);
chk("slow read: no dummy bits", dummy_bits, 0);
chk("slow read: addr", addr, 24'hABCDEF);
close_frame();
// --- dummy_cycles = 0 must SKIP the dummy phase, not enter it ---
dummy_cycles = 5'd0; dummy_bits = 0;
open_frame();
send_byte(8'h0B); send_byte(8'h00); send_byte(8'h00); send_byte(8'h01);
chk("dummy=0: skips DUMMY entirely", phase, PH_DATA);
chk("dummy=0: no dummy bits", dummy_bits, 0);
close_frame();
// --- STATUS 0x05: no address, no dummy ---
open_frame();
send_byte(8'h05);
chk("status: straight to DATA", phase, PH_DATA);
close_frame();
// --- Abort during the dummy phase: CS must resynchronise ---
dummy_cycles = 5'd8;
open_frame();
send_byte(8'h0B); send_byte(8'h11); send_byte(8'h22); send_byte(8'h33);
one_bit(1'b0); one_bit(1'b0);
chk("abort: in DUMMY", phase, PH_DUMMY);
close_frame();
chk("abort: back to IDLE", phase, PH_IDLE);
chk("abort: MISO released", drive_miso, 0);
open_frame();
send_byte(8'h05);
chk("recovery after abort", phase, PH_DATA);
close_frame();
if (errors == 0)
$display("PASS: dummy length is programmable and counted in SCLK cycles; a zero-length dummy is skipped; MISO is driven only in the data phase; CS resynchronises from DUMMY");
else
$display("FAILED with %0d error(s)", errors);
$finish;
end
initial begin #500000; $display("FAIL: watchdog timeout"); $finish; end
endmodule// spi_dummy_phase.v — the same sequencer with a dummy phase, in Verilog-2001.
module spi_dummy_phase #(
parameter ADDR_BYTES = 3
) (
input wire clk,
input wire rst_n,
input wire cs_n,
input wire bit_stb,
input wire byte_done,
input wire [7:0] rx_byte,
input wire [4:0] dummy_cycles,
output reg [2:0] phase,
output reg [7:0] cmd,
output reg [8*ADDR_BYTES-1:0] addr,
output wire in_data,
output wire drive_miso
);
localparam PH_IDLE = 3'd0,
PH_CMD = 3'd1,
PH_ADDR = 3'd2,
PH_DUMMY = 3'd3,
PH_DATA = 3'd4;
localparam CMD_READ = 8'h03,
CMD_FAST_READ = 8'h0B,
CMD_RD_STATUS = 8'h05;
function integer clogb2;
input integer value;
integer v;
begin
v = value - 1;
for (clogb2 = 0; v > 0; clogb2 = clogb2 + 1) v = v >> 1;
end
endfunction
localparam CW = (ADDR_BYTES > 1) ? clogb2(ADDR_BYTES) : 1;
reg [CW-1:0] addr_cnt;
reg [4:0] dummy_cnt;
function has_addr;
input [7:0] op;
begin
has_addr = (op == CMD_READ) || (op == CMD_FAST_READ);
end
endfunction
// Only the fast read pays for turnaround; the slow read is specified at a
// low enough clock that the device can answer immediately.
function has_dummy;
input [7:0] op;
begin
has_dummy = (op == CMD_FAST_READ);
end
endfunction
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
phase <= PH_IDLE;
cmd <= 8'h00;
addr <= {(8*ADDR_BYTES){1'b0}};
addr_cnt <= {CW{1'b0}};
dummy_cnt <= 5'd0;
end else if (cs_n) begin
phase <= PH_IDLE;
addr_cnt <= {CW{1'b0}};
dummy_cnt <= 5'd0;
end else begin
case (phase)
PH_IDLE: begin
phase <= PH_CMD;
addr_cnt <= {CW{1'b0}};
end
PH_CMD: if (byte_done) begin
cmd <= rx_byte;
if (has_addr(rx_byte)) begin
phase <= PH_ADDR;
addr_cnt <= {CW{1'b0}};
end else begin
phase <= PH_DATA;
end
end
PH_ADDR: if (byte_done) begin
addr <= {addr[8*ADDR_BYTES-9:0], rx_byte};
if (addr_cnt == (ADDR_BYTES - 1)) begin
dummy_cnt <= 5'd0;
// `cmd` was registered back in PH_CMD, so it already
// holds this transaction's opcode. A zero-length dummy
// must be SKIPPED, not entered and exited.
if (has_dummy(cmd) && dummy_cycles != 5'd0)
phase <= PH_DUMMY;
else
phase <= PH_DATA;
end else begin
addr_cnt <= addr_cnt + 1'b1;
end
end
PH_DUMMY: if (bit_stb) begin
if (dummy_cnt == (dummy_cycles - 5'd1)) phase <= PH_DATA;
else dummy_cnt <= dummy_cnt + 5'd1;
end
PH_DATA: ;
default: phase <= PH_IDLE;
endcase
end
end
assign in_data = (phase == PH_DATA);
assign drive_miso = (phase == PH_DATA);
endmodule// spi_dummy_phase_tb.v — the same checks in Verilog-2001.
`timescale 1ns/1ps
module spi_dummy_phase_tb;
reg clk = 0, rst_n = 0;
always #5 clk = ~clk;
parameter AB = 3;
reg cs_n = 1, bit_stb = 0, byte_done = 0;
reg [7:0] rx_byte = 8'h00;
reg [4:0] dummy_cycles = 5'd8;
wire [2:0] phase; wire [7:0] cmd; wire [8*AB-1:0] addr;
wire in_data, drive_miso;
spi_dummy_phase #(.ADDR_BYTES(AB)) dut (
.clk(clk), .rst_n(rst_n), .cs_n(cs_n), .bit_stb(bit_stb), .byte_done(byte_done),
.rx_byte(rx_byte), .dummy_cycles(dummy_cycles), .phase(phase), .cmd(cmd),
.addr(addr), .in_data(in_data), .drive_miso(drive_miso));
localparam PH_IDLE = 3'd0, PH_CMD = 3'd1, PH_ADDR = 3'd2,
PH_DUMMY = 3'd3, PH_DATA = 3'd4;
integer errors = 0, dummy_bits = 0, i;
task chk;
input [80*8-1:0] what;
input [31:0] got, exp;
begin
if (got !== exp) begin
$display("FAIL %0s: got %0d exp %0d", what, got, exp);
errors = errors + 1;
end
end
endtask
task one_bit;
input last;
begin
@(negedge clk); bit_stb = 1; byte_done = last;
@(negedge clk); bit_stb = 0; byte_done = 0;
@(negedge clk);
end
endtask
task send_byte;
input [7:0] b;
integer j;
begin
rx_byte = b;
for (j = 0; j < 8; j = j + 1) one_bit(j == 7);
end
endtask
task open_frame; begin cs_n = 0; @(negedge clk); @(negedge clk); end endtask
task close_frame; begin cs_n = 1; @(negedge clk); @(negedge clk); end endtask
always @(posedge clk)
if (rst_n && bit_stb && phase == PH_DUMMY) dummy_bits = dummy_bits + 1;
initial begin
repeat (3) @(negedge clk); rst_n = 1; @(negedge clk);
dummy_cycles = 5'd8; dummy_bits = 0;
open_frame();
send_byte(8'h0B);
chk("fast read: phase after opcode", phase, PH_ADDR);
send_byte(8'h12); send_byte(8'h34); send_byte(8'h56);
chk("fast read: enters DUMMY", phase, PH_DUMMY);
chk("fast read: addr", addr, 24'h123456);
chk("fast read: MISO not driven yet", drive_miso, 0);
for (i = 0; i < 8; i = i + 1) one_bit(1'b0);
chk("fast read: dummy length", dummy_bits, 8);
chk("fast read: now DATA", phase, PH_DATA);
chk("fast read: MISO driven", drive_miso, 1);
close_frame();
dummy_cycles = 5'd4; dummy_bits = 0;
open_frame();
send_byte(8'h0B); send_byte(8'h00); send_byte(8'h00); send_byte(8'h10);
chk("dummy=4: enters DUMMY", phase, PH_DUMMY);
for (i = 0; i < 4; i = i + 1) one_bit(1'b0);
chk("dummy=4: dummy length", dummy_bits, 4);
chk("dummy=4: now DATA", phase, PH_DATA);
close_frame();
dummy_cycles = 5'd8; dummy_bits = 0;
open_frame();
send_byte(8'h03); send_byte(8'hAB); send_byte(8'hCD); send_byte(8'hEF);
chk("slow read: straight to DATA", phase, PH_DATA);
chk("slow read: no dummy bits", dummy_bits, 0);
chk("slow read: addr", addr, 24'hABCDEF);
close_frame();
dummy_cycles = 5'd0; dummy_bits = 0;
open_frame();
send_byte(8'h0B); send_byte(8'h00); send_byte(8'h00); send_byte(8'h01);
chk("dummy=0: skips DUMMY entirely", phase, PH_DATA);
chk("dummy=0: no dummy bits", dummy_bits, 0);
close_frame();
open_frame();
send_byte(8'h05);
chk("status: straight to DATA", phase, PH_DATA);
close_frame();
dummy_cycles = 5'd8;
open_frame();
send_byte(8'h0B); send_byte(8'h11); send_byte(8'h22); send_byte(8'h33);
one_bit(1'b0); one_bit(1'b0);
chk("abort: in DUMMY", phase, PH_DUMMY);
close_frame();
chk("abort: back to IDLE", phase, PH_IDLE);
chk("abort: MISO released", drive_miso, 0);
open_frame();
send_byte(8'h05);
chk("recovery after abort", phase, PH_DATA);
close_frame();
if (errors == 0)
$display("PASS: dummy length is programmable and counted in SCLK cycles; a zero-length dummy is skipped; MISO is driven only in the data phase; CS resynchronises from DUMMY");
else
$display("FAILED with %0d error(s)", errors);
$finish;
end
initial begin #500000; $display("FAIL: watchdog timeout"); $finish; end
endmodule-- spi_dummy_phase.vhd — the same sequencer with a dummy phase, in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_dummy_phase is
generic (
ADDR_BYTES : positive := 3
);
port (
clk : in std_logic;
rst_n : in std_logic;
cs_n : in std_logic;
bit_stb : in std_logic;
byte_done : in std_logic;
rx_byte : in std_logic_vector(7 downto 0);
dummy_cycles : in unsigned(4 downto 0);
phase : out std_logic_vector(2 downto 0);
cmd : out std_logic_vector(7 downto 0);
addr : out std_logic_vector(8 * ADDR_BYTES - 1 downto 0);
in_data : out std_logic;
drive_miso : out std_logic
);
end entity spi_dummy_phase;
architecture rtl of spi_dummy_phase is
type phase_t is (PH_IDLE, PH_CMD, PH_ADDR, PH_DUMMY, PH_DATA);
signal state : phase_t;
constant CMD_READ : std_logic_vector(7 downto 0) := x"03";
constant CMD_FAST_READ : std_logic_vector(7 downto 0) := x"0B";
constant CMD_RD_STATUS : std_logic_vector(7 downto 0) := x"05";
signal addr_cnt : integer range 0 to ADDR_BYTES - 1;
signal dummy_cnt : unsigned(4 downto 0);
signal addr_r : std_logic_vector(8 * ADDR_BYTES - 1 downto 0);
signal cmd_r : std_logic_vector(7 downto 0);
function has_addr (op : std_logic_vector(7 downto 0)) return boolean is
begin
return (op = CMD_READ) or (op = CMD_FAST_READ);
end function;
-- Only the fast read pays for turnaround; the slow read is specified at a
-- low enough clock that the device can answer immediately.
function has_dummy (op : std_logic_vector(7 downto 0)) return boolean is
begin
return op = CMD_FAST_READ;
end function;
function encode (s : phase_t) return std_logic_vector is
begin
case s is
when PH_IDLE => return "000";
when PH_CMD => return "001";
when PH_ADDR => return "010";
when PH_DUMMY => return "011";
when PH_DATA => return "100";
end case;
end function;
begin
process (clk, rst_n) is
begin
if rst_n = '0' then
state <= PH_IDLE;
cmd_r <= (others => '0');
addr_r <= (others => '0');
addr_cnt <= 0;
dummy_cnt <= (others => '0');
elsif rising_edge(clk) then
if cs_n = '1' then
state <= PH_IDLE;
addr_cnt <= 0;
dummy_cnt <= (others => '0');
else
case state is
when PH_IDLE =>
state <= PH_CMD;
addr_cnt <= 0;
when PH_CMD =>
if byte_done = '1' then
cmd_r <= rx_byte;
if has_addr(rx_byte) then
state <= PH_ADDR;
addr_cnt <= 0;
else
state <= PH_DATA;
end if;
end if;
when PH_ADDR =>
if byte_done = '1' then
addr_r <= addr_r(8 * ADDR_BYTES - 9 downto 0) & rx_byte;
if addr_cnt = ADDR_BYTES - 1 then
dummy_cnt <= (others => '0');
-- cmd_r was registered back in PH_CMD. A
-- zero-length dummy must be SKIPPED, not
-- entered and exited.
if has_dummy(cmd_r) and dummy_cycles /= 0 then
state <= PH_DUMMY;
else
state <= PH_DATA;
end if;
else
addr_cnt <= addr_cnt + 1;
end if;
end if;
when PH_DUMMY =>
if bit_stb = '1' then
if dummy_cnt = dummy_cycles - 1 then
state <= PH_DATA;
else
dummy_cnt <= dummy_cnt + 1;
end if;
end if;
when PH_DATA =>
null; -- remain until CS rises
end case;
end if;
end if;
end process;
phase <= encode(state);
cmd <= cmd_r;
addr <= addr_r;
in_data <= '1' when state = PH_DATA else '0';
-- The slave must not drive MISO before the data phase.
drive_miso <= '1' when state = PH_DATA else '0';
end architecture rtl;-- spi_dummy_phase_tb.vhd — the same checks in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_dummy_phase_tb is
end entity spi_dummy_phase_tb;
architecture tb of spi_dummy_phase_tb is
constant AB : positive := 3;
signal clk : std_logic := '0';
signal rst_n : std_logic := '0';
signal cs_n : std_logic := '1';
signal bit_stb : std_logic := '0';
signal byte_done : std_logic := '0';
signal rx_byte : std_logic_vector(7 downto 0) := (others => '0');
signal dummy_cycles : unsigned(4 downto 0) := to_unsigned(8, 5);
signal halt : boolean := false;
signal phase : std_logic_vector(2 downto 0);
signal cmd : std_logic_vector(7 downto 0);
signal addr : std_logic_vector(8 * AB - 1 downto 0);
signal in_data : std_logic;
signal drive_miso : std_logic;
signal errors : natural := 0;
signal dummy_bits : natural := 0;
constant PH_IDLE : std_logic_vector(2 downto 0) := "000";
constant PH_CMD : std_logic_vector(2 downto 0) := "001";
constant PH_ADDR : std_logic_vector(2 downto 0) := "010";
constant PH_DUMMY : std_logic_vector(2 downto 0) := "011";
constant PH_DATA : std_logic_vector(2 downto 0) := "100";
begin
clk <= not clk after 5 ns when not halt else '0';
dut : entity work.spi_dummy_phase
generic map (ADDR_BYTES => AB)
port map (clk => clk, rst_n => rst_n, cs_n => cs_n, bit_stb => bit_stb,
byte_done => byte_done, rx_byte => rx_byte,
dummy_cycles => dummy_cycles, phase => phase, cmd => cmd,
addr => addr, in_data => in_data, drive_miso => drive_miso);
-- Count bits spent in the dummy phase, so its LENGTH is actually checked.
counter : process (clk) is
begin
if rising_edge(clk) then
if rst_n = '1' and bit_stb = '1' and phase = PH_DUMMY then
dummy_bits <= dummy_bits + 1;
end if;
end if;
end process;
stim : process is
variable base : natural;
procedure chk_v (what : string; got, exp : std_logic_vector) is
begin
if got /= exp then
report "FAIL " & what & ": got 0x" & to_hstring(got)
& " exp 0x" & to_hstring(exp) severity error;
errors <= errors + 1;
end if;
end procedure;
procedure chk_b (what : string; got, exp : std_logic) is
begin
if got /= exp then
report "FAIL " & what severity error;
errors <= errors + 1;
end if;
end procedure;
procedure chk_n (what : string; got, exp : natural) is
begin
if got /= exp then
report "FAIL " & what & ": got " & integer'image(got)
& " exp " & integer'image(exp) severity error;
errors <= errors + 1;
end if;
end procedure;
procedure one_bit (last : std_logic) is
begin
wait until falling_edge(clk); bit_stb <= '1'; byte_done <= last;
wait until falling_edge(clk); bit_stb <= '0'; byte_done <= '0';
wait until falling_edge(clk);
end procedure;
procedure send_byte (b : std_logic_vector(7 downto 0)) is
begin
rx_byte <= b;
for j in 0 to 7 loop
if j = 7 then one_bit('1'); else one_bit('0'); end if;
end loop;
end procedure;
procedure open_frame is
begin
cs_n <= '0';
wait until falling_edge(clk);
wait until falling_edge(clk);
end procedure;
procedure close_frame is
begin
cs_n <= '1';
wait until falling_edge(clk);
wait until falling_edge(clk);
end procedure;
begin
for i in 0 to 2 loop
wait until falling_edge(clk);
end loop;
rst_n <= '1';
wait until falling_edge(clk);
-- FAST READ 0x0B with 8 dummy cycles
dummy_cycles <= to_unsigned(8, 5);
wait until falling_edge(clk);
base := dummy_bits;
open_frame;
send_byte(x"0B");
chk_v("fast read: phase after opcode", phase, PH_ADDR);
send_byte(x"12"); send_byte(x"34"); send_byte(x"56");
chk_v("fast read: enters DUMMY", phase, PH_DUMMY);
chk_v("fast read: addr", addr, x"123456");
chk_b("fast read: MISO not driven yet", drive_miso, '0');
for i in 0 to 7 loop one_bit('0'); end loop;
chk_n("fast read: dummy length", dummy_bits - base, 8);
chk_v("fast read: now DATA", phase, PH_DATA);
chk_b("fast read: MISO driven", drive_miso, '1');
close_frame;
-- The same command with a different programmed dummy length
dummy_cycles <= to_unsigned(4, 5);
wait until falling_edge(clk);
base := dummy_bits;
open_frame;
send_byte(x"0B"); send_byte(x"00"); send_byte(x"00"); send_byte(x"10");
chk_v("dummy=4: enters DUMMY", phase, PH_DUMMY);
for i in 0 to 3 loop one_bit('0'); end loop;
chk_n("dummy=4: dummy length", dummy_bits - base, 4);
chk_v("dummy=4: now DATA", phase, PH_DATA);
close_frame;
-- READ 0x03: address, but no dummy phase at all
dummy_cycles <= to_unsigned(8, 5);
wait until falling_edge(clk);
base := dummy_bits;
open_frame;
send_byte(x"03"); send_byte(x"AB"); send_byte(x"CD"); send_byte(x"EF");
chk_v("slow read: straight to DATA", phase, PH_DATA);
chk_n("slow read: no dummy bits", dummy_bits - base, 0);
chk_v("slow read: addr", addr, x"ABCDEF");
close_frame;
-- dummy_cycles = 0 must SKIP the dummy phase
dummy_cycles <= to_unsigned(0, 5);
wait until falling_edge(clk);
base := dummy_bits;
open_frame;
send_byte(x"0B"); send_byte(x"00"); send_byte(x"00"); send_byte(x"01");
chk_v("dummy=0: skips DUMMY entirely", phase, PH_DATA);
chk_n("dummy=0: no dummy bits", dummy_bits - base, 0);
close_frame;
-- STATUS 0x05: no address, no dummy
open_frame;
send_byte(x"05");
chk_v("status: straight to DATA", phase, PH_DATA);
close_frame;
-- Abort during the dummy phase
dummy_cycles <= to_unsigned(8, 5);
wait until falling_edge(clk);
open_frame;
send_byte(x"0B"); send_byte(x"11"); send_byte(x"22"); send_byte(x"33");
one_bit('0'); one_bit('0');
chk_v("abort: in DUMMY", phase, PH_DUMMY);
close_frame;
chk_v("abort: back to IDLE", phase, PH_IDLE);
chk_b("abort: MISO released", drive_miso, '0');
open_frame;
send_byte(x"05");
chk_v("recovery after abort", phase, PH_DATA);
close_frame;
wait until falling_edge(clk);
if errors = 0 then
report "PASS: dummy length is programmable and counted in SCLK cycles; "
& "a zero-length dummy is skipped; MISO is driven only in the data "
& "phase; CS resynchronises from DUMMY" severity note;
else
report "FAILED with " & integer'image(errors) & " error(s)" severity error;
end if;
halt <= true;
wait;
end process;
watchdog : process is
begin
wait for 500 us;
if not halt then
report "FAIL: watchdog timeout" severity failure;
end if;
wait;
end process;
end architecture tb;Parity
All three describe the same machine: identical ports with dummy_cycles as a runtime input, asynchronous active-low reset, cs_n forcing IDLE with priority, the opcode decode selecting whether an address and a dummy phase follow, the dummy phase counting bit_stb rather than byte_done, a zero-length dummy skipped entirely, and drive_miso asserted only in the data phase. The VHDL carries dummy_cycles as unsigned where the Verilog dialects use a plain vector — a typing difference, not a behavioural one. All three testbenches run the same six scenarios and count the bits actually spent in the dummy phase, so its length is checked rather than merely its existence.
8. Why a Verification Engineer Cares
The tri-state boundary is the assertion that matters most, because getting it wrong causes contention rather than merely wrong data:
// Driving MISO before the data phase is not a data error -- it is bus
// CONTENTION against whatever else shares the line (Chapter 1.2), and on
// a Dual/Quad bus it is contention against the MASTER.
property p_miso_only_in_data;
@(posedge clk) disable iff (!rst_n)
drive_miso |-> (phase == PH_DATA);
endproperty
a_miso_only_in_data : assert property (p_miso_only_in_data)
else $error("slave drove MISO outside the data phase");
// And it must release promptly when the frame ends, or it will still be
// driving when the next device is selected.
property p_miso_released_on_cs;
@(posedge clk) disable iff (!rst_n)
cs_n |=> !drive_miso;
endproperty
a_miso_released : assert property (p_miso_released_on_cs)
else $error("slave still driving MISO after CS deasserted");
// The dummy phase lasted exactly as long as configured -- the property
// that catches an off-by-one in either direction.
property p_dummy_length;
@(posedge clk) disable iff (!rst_n || cs_n)
$rose(phase == PH_DUMMY) |-> ##1 (dummy_bits_counted == dummy_cycles)[->1]
within (phase == PH_DUMMY)[*1:$];
endpropertyWhat these prove. That the output enable is confined to the data phase and released at frame end, and that the turnaround lasted the configured number of cycles. What they do not prove. That the configured number is the number the device requires — a datasheet fact, and the Chapter 4.1 §7 boundary once more. Nor do they prove the analogue reality: an enable deasserting on time in RTL says nothing about how long the pad actually takes to stop driving, which is a Chapter 1.6 concern and measurable only with an oscilloscope.
Coverage should treat the dummy length as a small numeric axis with meaningful boundaries:
covergroup spi_dummy_cg @(posedge cs_rose);
cp_dummy : coverpoint cfg.dummy_cycles {
bins none = {0}; // MUST be hit: the skip path
bins one = {1}; // the minimum non-zero count
bins sub_byte = {[2:7]}; // NOT a whole byte -- §3
bins one_byte = {8};
bins multi_byte = {[9:31]};
}
// The relationship the chapter is about: the same command needs a
// different count at a different clock rate.
cp_rate : coverpoint cfg.sclk_hz {
bins slow = {[0:25_000_000]};
bins fast = {[25_000_001:$]};
}
x_dummy_rate : cross cp_dummy, cp_rate;
// Aborting during turnaround: the phase that exists purely to wait is
// also the one a master is most likely to give up in.
cp_abort_in_dummy : coverpoint (phase_at_cs_rise == PH_DUMMY);
endgroupThe none bin is the one to insist on. It is the only bin that exercises the skip path of §7, that path is a distinct branch in the RTL, and a suite that always configures a non-zero dummy length never executes it — while the bug it contains produces a one-bit offset that looks like a timing problem.
9. Why an FPGA or ASIC Engineer Cares
On a slave, the dummy phase is your fetch budget and you should spend it deliberately. The cycles exist so the device can retrieve data. If the internal path — a block-RAM read, a register-file access, a CDC crossing into another domain — takes longer than the configured count, the first data byte is wrong. Knowing the count converts a vague "be fast enough" into a hard number of cycles to design against, which is precisely what makes it a tractable timing problem.
On a master, dummy cycles are the cheapest thing to get wrong and the easiest to make configurable. A controller that hard-codes eight cannot talk to a device needing six or ten, and cannot follow a device whose requirement changes with frequency. Making the count a register costs one comparator input.
The output-enable path deserves attention. drive_miso controls a tri-state buffer, and on an FPGA the enable should be registered in the I/O block alongside the data so that both change together. If the enable is a combinational decode reaching the pad by a different route than the data, the two can skew — asserting the driver before the data is valid, or releasing it after it should have gone high impedance. Both produce marginal, temperature-dependent faults of exactly the kind Chapter 2.4 warns are hardest to find.
On an ASIC, the pad's disable time is a real number and it is not zero. A pad may take a nanosecond or more to stop driving after its enable deasserts. On a shared line that time is part of the turnaround requirement, which is one of the reasons Dual and Quad modes specify more dummy cycles than the array access alone would justify.
10. Failure Signature — Data Offset by Exactly One Byte
Symptom. A fast read returns data that is recognisably close to correct: the values are the right kind of data and in the right sequence, but shifted — the first byte returned is the byte that should have been second, or the whole stream is displaced by a fixed amount. Re-reading gives the same offset. Slow reads using the other opcode work perfectly.
Why "slow reads work" is decisive. The low-speed read command has no dummy phase. That it works proves the mode, bit order, width, framing, command decode and address handling are all correct — the entire chain except turnaround. The fault is confined to the one thing fast reads add, which is a remarkably clean partition for a single observation.
Plausible mechanisms.
- A dummy-length disagreement: the driver sends 8 dummy cycles where the device expects 6, or the reverse. The offset equals the difference.
- A zero-dummy skip bug of the §7 kind, if the configuration happens to be zero.
- The wrong dummy count for the clock rate — correct at 20 MHz, insufficient at 80 MHz, which is §3 arriving as a rate-dependent fault.
- The device configured for a different dummy count than the driver assumes, on a part where it is a register with a power-on default.
The discriminating observations. Measure the offset in bits, not bytes. The difference between the sent and expected dummy counts is the offset, so a two-bit displacement means a two-cycle disagreement — which points straight at a frequency-table mismatch, since those differ by small numbers of cycles. A displacement of exactly eight bits instead suggests a whole-byte assumption in the driver.
Then vary the clock rate. If the offset appears above some frequency and vanishes below it, the dummy count is right for the low rate and too small for the high one, and the device's frequency table is the thing to consult. That behaviour is distinctive: unlike a margin problem, which degrades progressively, this switches cleanly at the threshold where the required count increments.
Why the investigation goes wrong. Because rate-dependent corruption reads as a signal-integrity problem, and the team starts probing edges. The giveaway is that a genuine margin failure produces random corruption that worsens gradually, while this produces a fixed, reproducible offset that appears at a threshold. Reproducibility is the discriminator, and it is available from the data alone before any probe is attached.
11. Common Misconceptions
12. Reason It Through
Work this before reading the answer.
A SPI flash driver performs fast reads correctly at 20 MHz. At 80 MHz every read returns data displaced by two bits — consistently, reproducibly, with the same displacement every time. An engineer concludes the board cannot support 80 MHz and proposes adding series termination and shortening the traces.
Is that the right conclusion?
No, and the evidence rules it out before any board work is considered.
Start with reproducibility. A displacement that is identical on every read is not a signal-integrity failure. Marginal timing produces corruption that varies between reads, varies with temperature, and affects different bits on different attempts (Chapter 2.4). A fixed two-bit offset is a logical error that happens to be frequency-dependent.
Now read the number. The offset is two bits, and the dummy phase is the only part of the transaction measured in individual cycles. A two-bit displacement means the device consumed two more cycles of turnaround than the driver supplied — so the driver is sending a dummy count that was correct at 20 MHz and is two cycles short at 80 MHz.
Why does that happen at exactly this boundary? Because the device's access time is fixed in nanoseconds while the clock period shrank by a factor of four — 50 ns to 12.5 ns. The required count is ceil(t_access / T), so it increases as the period falls, in steps. The driver is presumably using one hard-coded value, which is correct in the frequency band it was tested in and wrong above it.
What confirms it in one step? Increase the dummy count by two and retry at 80 MHz. If the data comes back correct, the diagnosis is complete. That is a one-line driver change against a board respin — and the asymmetry of those two costs is the reason to test the cheap hypothesis first.
What would have made this rate-dependence obvious? Sweeping the clock and recording where the offset appears. A margin problem has no threshold — it degrades. This has a sharp boundary at the frequency where the required dummy count increments, and finding that boundary both confirms the mechanism and tells you which row of the device's frequency table you have crossed.
The general lesson. "It breaks when I go faster" is not synonymous with "the board cannot go faster". Configuration that is a function of frequency fails at a threshold and reproduces exactly; analogue margin fails progressively and randomly. Distinguishing them costs one observation and can save a respin.
13. Understanding Check
14. Summary
A device cannot answer the instant an address completes: the address only just arrived, nothing has been fetched, and the next edge is imminent. The dummy phase is edges clocked with no meaningful data, and its only function is to let time pass.
There are two independent reasons for it. Access time is the silicon's fetch delay, fixed in nanoseconds. Bus turnaround is the gap needed when a shared bidirectional line changes direction — which does not exist on four-wire SPI, where every dummy cycle is buying access time.
Because the requirement is in nanoseconds and the bus offers clock periods, the count is a conversion: ceil(t_access / T_sclk). So it is measured in cycles, not bytes, counts of 6 or 10 are normal, and the number grows as the clock speeds up — which is why it is a programmable register on parts covering a wide frequency range.
The phase is pure overhead. A fast read with a 3-byte address and 8 dummy cycles spends 40 cycles before the first data bit, making single-byte reads about 17% efficient and 256-byte bursts about 98%. Burst length, not clock rate, is what recovers that for small transfers.
In RTL it is one state and one counter, notable for consuming a per-bit strobe where every other phase consumes a per-byte one, and for skipping a zero-length dummy rather than entering it — since entering would steal a cycle from the data phase and offset the output by one bit.
For verification the critical assertions are that the slave drives MISO only in the data phase and releases it at frame end, because driving early is contention rather than merely wrong data. And the coverage bin most worth insisting on is dummy_cycles == 0, the only one that exercises the skip path.
When data comes back displaced, measure the offset in bits: it equals the disagreement in cycles. And a fault that appears at a frequency threshold and reproduces exactly is configuration, not signal integrity — which fails progressively and randomly.
15. What Comes Next
That closes Module 4. Chapter 4.1 established that SPI defines signalling and nothing above it, and the four chapters since have each taken one field it declines to specify — width, bit order, transaction phases, and turnaround — and built the hardware a device's choice about each one requires.
Module 3 taught which edge makes a bit valid; Module 4 taught what the bits mean, and that the meaning is always a per-device agreement. Module 5 puts the two together and follows a complete write transaction end to end: what a device does with data once it has been framed and received, why a write is acknowledged by nothing at all, and how a master knows a write has finished when the bus provides no completion signal.
Browse the path on the SPI curriculum index, or revisit Command, Address, and Data Phases for the sequencer this chapter extends.
Continue learning
Related tutorials
- Related topic
Back-to-Back Transactions and Inter-Frame Gap
How soon chip select may fall again after it rises: the minimum deselect time, why a device needs it, what the gap costs in throughput, and the hardware enforcement that keeps software from violating it.
- Related topic
Raw vs Effective Payload Throughput
The three throughput numbers a link has — raw bit rate, protocol rate and effective payload rate — why they differ by four times on an ordinary flash read, and the performance counter that measures the gap instead of estimating it.
- Related topic
Address Fields, Dummy Cycles, and Burst Behaviour
Why three address bytes make 16 MB a category boundary, why dummy latency is counted in cycles and never rounds to bytes, what a device does at the end of a burst, and the planner that turns those numbers into a cycle-accurate schedule.
- Related topic
Normal Read, Fast Read, and Dummy Cycles
Why a flash offers two read commands returning the same data, why the faster one must buy array access time in dummy cycles, why fast read is not always faster, and the selector that chooses by transfer time rather than clock rate.
