SPI · Module 6
Anatomy of a Read Transaction
The four phases of a device read, why a read cannot be a write reversed, when the slave takes and releases MISO, the obligation to have the first data bit valid before any edge can launch it, and the slave read datapath in three HDLs.
Module 5 followed a write from chip select to commit. Its defining property was that nothing came back: the master drove everything and received no confirmation.
A read inverts that, and the inversion is not symmetric.
The master has asked for data. The device has to produce it, place it on a wire it does not own by default, and have the first bit ready before there is any edge to launch it on. How?
Every difficulty in this module comes from that sentence, and this chapter establishes the shape of it.
1. A Read Is Not a Write Reversed
It is tempting to think of a read as a write with the data travelling the other way. Three things break that symmetry.
The device must be told what to read before it can read it. A read begins as a transmission from the master — an opcode and usually an address — and only becomes a reception partway through. Chapter 6.2 is devoted to this: every read is a write first.
The data does not exist when the request finishes. At the instant the last address bit arrives, the device has not fetched anything. It needs time, and the bus only offers clock periods — which is the dummy phase, now appearing as a transaction-level necessity rather than a framing curiosity.
The device does not own MISO by default. On a multi-slave bus MISO is shared, so a device must take the line, drive it for exactly the data phase, and release it — three actions a write never requires.
2. The Four Phases
In order, with what each accomplishes:
- Command — the master names the operation. The device decodes it and learns the shape of the rest (Chapter 4.4 §5).
- Address — the master says where. Accumulated most-significant byte first.
- Dummy — the master clocks edges carrying nothing while the device fetches (Chapter 4.5). Absent on slow-read commands and on devices with no fetch delay.
- Data — the device drives MISO, one bit per launch edge, until the master releases CS.
Phases 1 and 2 travel master → device on MOSI. Phase 4 travels device → master on MISO. Phase 3 carries nothing in either direction and exists purely to separate them in time.
Note what still holds from Module 4: no phase boundary is marked on the bus, and both ends track position by counting from CS (Chapter 4.4 §2).
3. The Transaction
4. What MISO Does
The slave takes MISO, uses it, and gives it back
8 cyclesThree obligations are visible in that lane, and a device that gets any of them wrong causes trouble beyond its own transaction:
Do not drive early. Before the data phase another device may own MISO, or the master may, on a shared-line bus. Driving early is contention, which is an electrical fault rather than a data error.
Do not drive late. After CS rises the device must release promptly, or it will still be driving when the next device is selected.
Do not leave a gap. Between taking the line and the first launch edge, MISO must already carry a valid bit — which is §5.
5. The First-Bit Obligation
This is the read's version of a problem Chapter 3.4 §3 raised for writes, and it is sharper here.
Consider the moment the dummy phase ends. The next launch edge will shift the device's output register by one position — so whatever is to be the first data bit must already be sitting at the output tap before that edge. There is no earlier edge on which to place it, exactly as there was none for a CPHA = 0 master's first transmitted bit.
So the device cannot "start shifting when the data phase begins". It must load the output register at the transition into the data phase, with the first byte already fetched. That is why the RTL in §6 loads on data_phase && !data_phase_q rather than on the first launch strobe, and it is why the dummy phase's length is a hard requirement rather than a performance tuning knob: the fetch must have completed before that transition.
The same obligation repeats at every byte boundary during a burst. When the last bit of byte n is launched, byte n+1 must already be available — so a streaming read requires the device to prefetch, running one byte ahead of the bus for the whole transfer. The fetch_req strobe in §6 is exactly that prefetch, and it fires one byte early by design.
6. Building the Slave Read Datapath — Three HDLs
The circuit
Circuit. An output shift register, a bit counter, a prefetch strobe generator, and an output-enable decode.
State. The word being shifted out, the bit position within it, and a registered copy of data_phase so the entry into the phase can be detected.
Datapath. On entry to the data phase the register is loaded from fetch_data. On each launch strobe it shifts left by one, presenting the next bit at the MSB tap. At the last bit position it reloads instead of shifting.
Control. Three mutually exclusive conditions, in priority order: outside the data phase, entering it, and shifting within it. The entry case is the one that carries §5's obligation.
Clock. The system clock. launch_stb is a strobe derived from SCLK edges, not a clock.
Reset. Asynchronous, active-low, to a cleared register with the prefetch strobe low.
Enables. Nothing moves except on entry or on a launch strobe, so the module is idle-safe between transfers.
Timing. fetch_req is a single-cycle strobe issued one byte in advance. The memory has a full byte time — eight launch edges — to respond, which is what makes streaming possible at speed.
Synthesis. WIDTH flip-flops for the shift register, ceil(log2(WIDTH)) for the counter, one comparator, and a small amount of control. miso_oe is a direct decode of data_phase, deliberately simple because it drives a pad.
Limitations. No memory model, no address generation (Chapter 5.3 built the pointer), and MSB-first fixed rather than configurable — Chapter 4.3 showed what making it configurable costs.
// spi_read_path.sv — the slave-side read datapath.
//
// A write ends inside the device (Chapter 5.1). A read must come back out,
// and that creates an obligation a write never has: the FIRST data bit must
// already be on MISO before the first launch edge of the data phase. There is
// no edge before it on which to place it.
//
// The dummy phase (Chapter 4.5) exists to buy the time this module needs
// between "the address is known" and "the first bit is on the pin".
module spi_read_path #(
parameter int WIDTH = 8
) (
input logic clk,
input logic rst_n,
input logic data_phase, // level, from the phase sequencer
input logic launch_stb, // one pulse per launch edge
input logic [WIDTH-1:0] fetch_data, // byte presented by the memory
output logic fetch_req, // pulse: hand me the next byte
output logic miso,
output logic miso_oe // output enable for the pad
);
localparam int CW = (WIDTH > 1) ? $clog2(WIDTH) : 1;
logic [WIDTH-1:0] shreg;
logic [CW-1:0] bit_cnt;
logic data_phase_q;
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
shreg <= '0;
bit_cnt <= '0;
fetch_req <= 1'b0;
data_phase_q <= 1'b0;
end else begin
data_phase_q <= data_phase;
fetch_req <= 1'b0; // a strobe, not a level
if (!data_phase) begin
// Outside the data phase nothing is shifted and nothing is
// driven. Holding the counter cleared means the first launch
// edge after entry is bit 0 of a fresh word.
bit_cnt <= '0;
end else if (!data_phase_q) begin
// ENTERING the data phase: the first byte must be loaded NOW,
// because the next launch edge will already be shifting it.
shreg <= fetch_data;
bit_cnt <= '0;
fetch_req <= 1'b1; // prefetch the byte after this one
end else if (launch_stb) begin
if (bit_cnt == CW'(WIDTH - 1)) begin
// Last bit of this word has just left; the next word must
// already be available on fetch_data.
shreg <= fetch_data;
bit_cnt <= '0;
fetch_req <= 1'b1;
end else begin
shreg <= {shreg[WIDTH-2:0], 1'b0};
bit_cnt <= bit_cnt + 1'b1;
end
end
end
end
assign miso = shreg[WIDTH-1]; // MSB-first (Chapter 4.3)
// The slave drives MISO ONLY in the data phase. Everything earlier belongs
// to the master's half of the exchange, and on a multi-slave bus driving
// early is contention rather than merely wrong data.
assign miso_oe = data_phase;
endmoduleThe ordering of the three branches is the design. !data_phase clears the counter so that entry always begins a fresh word; !data_phase_q is the entry case that loads; and only then does launch_stb shift. Checking launch_stb first — the natural way to write it — would consume the first launch edge shifting a register that had not been loaded, losing the first bit of every read.
// spi_read_path_tb.sv — first-bit timing, multi-byte streaming, and the
// output enable staying confined to the data phase.
`timescale 1ns/1ps
module spi_read_path_tb;
logic clk = 0, rst_n = 0;
always #5 clk = ~clk;
localparam int W = 8;
logic data_phase = 0, launch = 0;
logic [W-1:0] fetch_data = 8'h00;
logic fetch_req, miso, miso_oe;
spi_read_path #(.WIDTH(W)) dut (
.clk, .rst_n, .data_phase, .launch_stb(launch),
.fetch_data, .fetch_req, .miso, .miso_oe);
int errors = 0, fetches = 0;
task automatic chk(input string what, input int got, input int exp);
if (got !== exp) begin $display("FAIL %s: got 0x%0h exp 0x%0h", what, got, exp); errors++; end
endtask
// A tiny memory model: hand over the next byte whenever asked.
logic [W-1:0] mem [4] = '{8'hA5, 8'h3C, 8'h81, 8'h00};
int mem_idx = 0;
always @(posedge clk) begin
if (rst_n && fetch_req) begin
fetches++;
mem_idx <= (mem_idx + 1) % 4;
end
end
always_comb fetch_data = mem[mem_idx];
// Collect MISO as it is launched.
logic [W-1:0] got_byte;
task automatic shift_bits(input int n);
for (int i = 0; i < n; i++) begin
got_byte = {got_byte[W-2:0], miso}; // sample BEFORE the launch edge
@(negedge clk); launch = 1; @(negedge clk); launch = 0; @(negedge clk);
end
endtask
initial begin
repeat (3) @(negedge clk); rst_n = 1; @(negedge clk);
chk("idle: miso_oe low", miso_oe, 0);
// --- Enter the data phase. The first bit must be valid IMMEDIATELY,
// before any launch edge occurs. ---
data_phase = 1; @(negedge clk); @(negedge clk);
chk("data phase: miso_oe high", miso_oe, 1);
chk("first bit already valid", miso, 1'b1); // 0xA5 MSB = 1
got_byte = '0;
shift_bits(W);
chk("byte 0 streamed", got_byte, 8'hA5);
got_byte = '0;
shift_bits(W);
chk("byte 1 streamed", got_byte, 8'h3C);
got_byte = '0;
shift_bits(W);
chk("byte 2 streamed", got_byte, 8'h81);
// --- Leaving the data phase must release the pad ---
data_phase = 0; @(negedge clk); @(negedge clk);
chk("released: miso_oe low", miso_oe, 0);
// --- Re-entering starts a fresh word, not mid-byte ---
mem_idx = 0; @(negedge clk);
data_phase = 1; @(negedge clk); @(negedge clk);
chk("re-entry: first bit valid again", miso, 1'b1);
got_byte = '0;
shift_bits(W);
chk("re-entry byte", got_byte, 8'hA5);
data_phase = 0; @(negedge clk);
if (errors == 0)
$display("PASS: the first data bit is valid on entry to the data phase, bytes stream back to back, and MISO is driven only while the data phase is active");
else
$display("FAILED with %0d error(s)", errors);
$finish;
end
initial begin #500000; $display("FAIL: watchdog timeout"); $finish; end
endmoduleThe check that matters most is chk("first bit already valid", miso, 1'b1) — performed before any launch strobe. A design that loaded on the first strobe would still stream the remaining bits correctly and pass every byte comparison, while being wrong about the one bit that has no edge to carry it.
// spi_read_path.v — the same slave-side read datapath in Verilog-2001.
module spi_read_path #(
parameter WIDTH = 8
) (
input wire clk,
input wire rst_n,
input wire data_phase,
input wire launch_stb,
input wire [WIDTH-1:0] fetch_data,
output reg fetch_req,
output wire miso,
output wire miso_oe
);
function integer clogb2;
input integer value;
integer v;
begin
v = value - 1;
for (clogb2 = 0; v > 0; clogb2 = clogb2 + 1) v = v >> 1;
end
endfunction
localparam CW = (WIDTH > 1) ? clogb2(WIDTH) : 1;
reg [WIDTH-1:0] shreg;
reg [CW-1:0] bit_cnt;
reg data_phase_q;
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
shreg <= {WIDTH{1'b0}};
bit_cnt <= {CW{1'b0}};
fetch_req <= 1'b0;
data_phase_q <= 1'b0;
end else begin
data_phase_q <= data_phase;
fetch_req <= 1'b0;
if (!data_phase) begin
bit_cnt <= {CW{1'b0}};
end else if (!data_phase_q) begin
// ENTERING the data phase: load now, because the next launch
// edge will already be shifting this word.
shreg <= fetch_data;
bit_cnt <= {CW{1'b0}};
fetch_req <= 1'b1;
end else if (launch_stb) begin
if (bit_cnt == (WIDTH - 1)) begin
shreg <= fetch_data;
bit_cnt <= {CW{1'b0}};
fetch_req <= 1'b1;
end else begin
shreg <= {shreg[WIDTH-2:0], 1'b0};
bit_cnt <= bit_cnt + 1'b1;
end
end
end
end
assign miso = shreg[WIDTH-1];
assign miso_oe = data_phase;
endmodule// spi_read_path_tb.v — the same checks in Verilog-2001.
`timescale 1ns/1ps
module spi_read_path_tb;
reg clk = 0, rst_n = 0;
always #5 clk = ~clk;
parameter W = 8;
reg data_phase = 0, launch = 0;
wire [W-1:0] fetch_data;
wire fetch_req, miso, miso_oe;
spi_read_path #(.WIDTH(W)) dut (
.clk(clk), .rst_n(rst_n), .data_phase(data_phase), .launch_stb(launch),
.fetch_data(fetch_data), .fetch_req(fetch_req), .miso(miso), .miso_oe(miso_oe));
integer errors = 0, fetches = 0, mem_idx = 0, i;
reg [W-1:0] mem [0:3];
reg [W-1:0] got_byte;
initial begin
mem[0] = 8'hA5; mem[1] = 8'h3C; mem[2] = 8'h81; mem[3] = 8'h00;
end
assign fetch_data = mem[mem_idx];
task chk;
input [80*8-1:0] what;
input [31:0] got, exp;
begin
if (got !== exp) begin
$display("FAIL %0s: got 0x%0h exp 0x%0h", what, got, exp);
errors = errors + 1;
end
end
endtask
always @(posedge clk) if (rst_n && fetch_req) begin
fetches = fetches + 1;
mem_idx = (mem_idx + 1) % 4;
end
task shift_bits;
input integer n;
integer k;
begin
for (k = 0; k < n; k = k + 1) begin
got_byte = {got_byte[W-2:0], miso};
@(negedge clk); launch = 1; @(negedge clk); launch = 0; @(negedge clk);
end
end
endtask
initial begin
repeat (3) @(negedge clk); rst_n = 1; @(negedge clk);
chk("idle: miso_oe low", miso_oe, 0);
data_phase = 1; @(negedge clk); @(negedge clk);
chk("data phase: miso_oe high", miso_oe, 1);
chk("first bit already valid", miso, 1'b1);
got_byte = 0; shift_bits(W); chk("byte 0 streamed", got_byte, 8'hA5);
got_byte = 0; shift_bits(W); chk("byte 1 streamed", got_byte, 8'h3C);
got_byte = 0; shift_bits(W); chk("byte 2 streamed", got_byte, 8'h81);
data_phase = 0; @(negedge clk); @(negedge clk);
chk("released: miso_oe low", miso_oe, 0);
mem_idx = 0; @(negedge clk);
data_phase = 1; @(negedge clk); @(negedge clk);
chk("re-entry: first bit valid again", miso, 1'b1);
got_byte = 0; shift_bits(W);
chk("re-entry byte", got_byte, 8'hA5);
data_phase = 0; @(negedge clk);
if (errors == 0)
$display("PASS: the first data bit is valid on entry to the data phase, bytes stream back to back, and MISO is driven only while the data phase is active");
else
$display("FAILED with %0d error(s)", errors);
$finish;
end
initial begin #500000; $display("FAIL: watchdog timeout"); $finish; end
endmodule-- spi_read_path.vhd — the same slave-side read datapath in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_read_path is
generic (
WIDTH : positive := 8
);
port (
clk : in std_logic;
rst_n : in std_logic;
data_phase : in std_logic; -- level
launch_stb : in std_logic; -- one pulse per launch edge
fetch_data : in std_logic_vector(WIDTH - 1 downto 0);
fetch_req : out std_logic; -- pulse
miso : out std_logic;
miso_oe : out std_logic
);
end entity spi_read_path;
architecture rtl of spi_read_path is
signal shreg : std_logic_vector(WIDTH - 1 downto 0);
signal bit_cnt : integer range 0 to WIDTH - 1;
signal data_phase_q : std_logic;
begin
process (clk, rst_n) is
begin
if rst_n = '0' then
shreg <= (others => '0');
bit_cnt <= 0;
fetch_req <= '0';
data_phase_q <= '0';
elsif rising_edge(clk) then
data_phase_q <= data_phase;
fetch_req <= '0';
if data_phase = '0' then
bit_cnt <= 0;
elsif data_phase_q = '0' then
-- ENTERING the data phase: load now, because the next launch
-- edge will already be shifting this word.
shreg <= fetch_data;
bit_cnt <= 0;
fetch_req <= '1';
elsif launch_stb = '1' then
if bit_cnt = WIDTH - 1 then
shreg <= fetch_data;
bit_cnt <= 0;
fetch_req <= '1';
else
shreg <= shreg(WIDTH - 2 downto 0) & '0';
bit_cnt <= bit_cnt + 1;
end if;
end if;
end if;
end process;
miso <= shreg(WIDTH - 1);
-- Driven ONLY in the data phase: driving early is contention, not merely
-- wrong data.
miso_oe <= data_phase;
end architecture rtl;-- spi_read_path_tb.vhd — the same checks in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_read_path_tb is
end entity spi_read_path_tb;
architecture tb of spi_read_path_tb is
constant W : positive := 8;
signal clk : std_logic := '0';
signal rst_n : std_logic := '0';
signal data_phase : std_logic := '0';
signal launch : std_logic := '0';
signal halt : boolean := false;
signal fetch_data : std_logic_vector(W - 1 downto 0);
signal fetch_req : std_logic;
signal miso : std_logic;
signal miso_oe : std_logic;
signal errors : natural := 0;
signal mem_idx : natural := 0;
signal mem_rst : std_logic := '0'; -- stim requests; memproc owns mem_idx
signal got_byte : std_logic_vector(W - 1 downto 0) := (others => '0');
type mem_t is array (0 to 3) of std_logic_vector(W - 1 downto 0);
constant MEM : mem_t := (x"A5", x"3C", x"81", x"00");
begin
clk <= not clk after 5 ns when not halt else '0';
fetch_data <= MEM(mem_idx);
dut : entity work.spi_read_path
generic map (WIDTH => W)
port map (clk => clk, rst_n => rst_n, data_phase => data_phase,
launch_stb => launch, fetch_data => fetch_data,
fetch_req => fetch_req, miso => miso, miso_oe => miso_oe);
-- Tiny memory model: advance on every fetch request.
memproc : process (clk) is
begin
if rising_edge(clk) then
if mem_rst = '1' then
mem_idx <= 0;
elsif rst_n = '1' and fetch_req = '1' then
mem_idx <= (mem_idx + 1) mod 4;
end if;
end if;
end process;
stim : process is
procedure chk_v (what : string; got, exp : std_logic_vector) is
begin
if got /= exp then
report "FAIL " & what & ": got 0x" & to_hstring(got)
& " exp 0x" & to_hstring(exp) severity error;
errors <= errors + 1;
end if;
end procedure;
procedure chk_b (what : string; got, exp : std_logic) is
begin
if got /= exp then
report "FAIL " & what severity error;
errors <= errors + 1;
end if;
end procedure;
procedure shift_bits (n : positive) is
begin
for k in 0 to n - 1 loop
got_byte <= got_byte(W - 2 downto 0) & miso;
wait until falling_edge(clk); launch <= '1';
wait until falling_edge(clk); launch <= '0';
wait until falling_edge(clk);
end loop;
end procedure;
begin
for i in 0 to 2 loop
wait until falling_edge(clk);
end loop;
rst_n <= '1';
wait until falling_edge(clk);
chk_b("idle: miso_oe low", miso_oe, '0');
data_phase <= '1';
wait until falling_edge(clk);
wait until falling_edge(clk);
chk_b("data phase: miso_oe high", miso_oe, '1');
chk_b("first bit already valid", miso, '1');
got_byte <= (others => '0'); wait until falling_edge(clk);
shift_bits(W); chk_v("byte 0 streamed", got_byte, x"A5");
got_byte <= (others => '0'); wait until falling_edge(clk);
shift_bits(W); chk_v("byte 1 streamed", got_byte, x"3C");
got_byte <= (others => '0'); wait until falling_edge(clk);
shift_bits(W); chk_v("byte 2 streamed", got_byte, x"81");
data_phase <= '0';
wait until falling_edge(clk);
wait until falling_edge(clk);
chk_b("released: miso_oe low", miso_oe, '0');
mem_rst <= '1';
wait until falling_edge(clk);
mem_rst <= '0';
wait until falling_edge(clk);
data_phase <= '1';
wait until falling_edge(clk);
wait until falling_edge(clk);
chk_b("re-entry: first bit valid again", miso, '1');
got_byte <= (others => '0'); wait until falling_edge(clk);
shift_bits(W);
chk_v("re-entry byte", got_byte, x"A5");
data_phase <= '0';
wait until falling_edge(clk);
if errors = 0 then
report "PASS: the first data bit is valid on entry to the data phase, "
& "bytes stream back to back, and MISO is driven only while the "
& "data phase is active" severity note;
else
report "FAILED with " & integer'image(errors) & " error(s)" severity error;
end if;
halt <= true;
wait;
end process;
watchdog : process is
begin
wait for 500 us;
if not halt then
report "FAIL: watchdog timeout" severity failure;
end if;
wait;
end process;
end architecture tb;Parity
All three describe the same hardware: identical ports, asynchronous active-low reset, a registered data_phase giving an entry detect, load on entry, shift on launch, reload at the last bit position, a prefetch strobe issued one byte ahead, and miso_oe as a direct decode of data_phase. All three testbenches drive the same four-byte memory model, verify the first bit before any strobe, stream three bytes back to back, confirm the pad is released on exit, and confirm re-entry starts a fresh word.
7. The UVM Read Transaction
A read is where the transaction object stops being a container for bytes and starts having to represent a direction change.
// A write item can be a byte queue. A read item cannot: the same frame
// carries master-driven bytes and device-driven bytes, and the boundary
// between them is a DEVICE fact (Chapter 4.1 §7) rather than an
// observable one. The item therefore records both streams plus the
// configuration that says where to split them.
class spi_read_item extends uvm_sequence_item;
`uvm_object_utils(spi_read_item)
// --- what the master sent ---
rand bit [7:0] opcode;
rand bit [7:0] addr_bytes[];
rand int unsigned dummy_cycles;
// --- what the device returned (filled by the monitor) ---
bit [7:0] rx_data[$];
// --- what the monitor needs in order to split the frame at all ---
int unsigned cmd_len;
int unsigned addr_len;
// The full MOSI and MISO streams, kept UNSPLIT. Keeping the raw
// streams alongside the interpretation means a scoreboard can be
// re-run against a corrected configuration without recapturing.
bit [7:0] mosi_raw[$];
bit [7:0] miso_raw[$];
constraint c_addr { addr_bytes.size() inside {0, 1, 2, 3, 4}; }
constraint c_dummy { dummy_cycles inside {[0:32]}; }
endclassTwo decisions in that class are worth defending.
The raw streams are kept. A read's interpretation depends on cmd_len, addr_len and dummy_cycles, none of which are observable (Chapter 6.5 makes this concrete). Keeping the unsplit streams means a wrong configuration produces a re-analysable item rather than a lost one — and when a test fails because the dummy count was wrong, you can prove it from the stored item instead of re-running.
dummy_cycles is randomised, not fixed. Chapter 4.5 §3 showed the required count varies with clock rate, so a sequence that always uses eight never exercises the alignment logic at any other value.
8. Why a Verification Engineer Cares
The bus-ownership properties are the ones that matter, because violating them is an electrical fault rather than a data error:
// 1. Never drive outside the data phase. On a multi-slave bus this is
// contention with another device; on a shared-line bus, with the master.
a_no_early_drive : assert property (
@(posedge clk) disable iff (!rst_n) miso_oe |-> data_phase)
else $error("slave drove MISO outside the data phase");
// 2. The first bit must be valid on entry, BEFORE any launch edge.
// This is §5 as a checkable statement.
property p_first_bit_ready;
@(posedge clk) disable iff (!rst_n)
$rose(data_phase) |-> ##1 (miso === fetch_data_at_entry[WIDTH-1]);
endproperty
// 3. Prefetch must lead by a full word, or a streaming read will stall.
// Catches a design that requests the next byte too late.
property p_prefetch_leads;
@(posedge clk) disable iff (!rst_n || !data_phase)
fetch_req |-> ##[1:$] (bit_cnt == 0);
endproperty
// 4. Release on frame close. A slave still driving when the next device
// is selected produces contention the NEXT transaction will suffer.
a_release_on_close : assert property (
@(posedge clk) disable iff (!rst_n) $fell(data_phase) |=> !miso_oe)
else $error("slave still driving MISO after the data phase ended");What these prove. That the slave's ownership of MISO is confined to the data phase and that the datapath meets its first-bit and prefetch obligations. What they do not prove is anything about the analogue reality of the pad: an output enable deasserting on time in RTL says nothing about how long the driver takes to reach high impedance, which is a Chapter 1.6 concern measurable only with an oscilloscope. A design can satisfy every property here and still overlap another driver on a real board.
Coverage should target the read's structural variables:
covergroup spi_read_cg @(posedge cs_rose);
cp_data_bytes : coverpoint rx_byte_count {
bins none = {0}; // a read aborted before any data
bins one = {1}; // no byte boundary exercised at all
bins two = {2}; // the first reload -- where prefetch bugs live
bins burst = {[3:255]};
bins long = {[256:$]};
}
// §5: a burst is where prefetch matters. A single-byte read never
// exercises the reload path, so it can pass with prefetch broken.
cp_abort_in_data : coverpoint aborted_mid_byte {
bins clean_boundary = {0};
bins mid_byte = {1}; // CS rose part-way through a word
}
x_bytes_abort : cross cp_data_bytes, cp_abort_in_data;
endgroupThe one and two bins are the pair that matters. A suite reading single bytes never crosses a byte boundary, so the reload-and-prefetch path — the part most likely to be wrong — is never executed. Two bytes is the smallest read that tests it.
9. Why an FPGA or ASIC Engineer Cares
The output register should live in the pad. miso comes straight from a flip-flop's output in this design, which is what allows an FPGA to pack that register into the I/O block and achieve a clean, short clock-to-output (Chapter 2.6). Placing any logic between the register and the pin — a multiplexer for configurable bit order, for instance — defeats the packing and lengthens t_v. If both are needed, select before the register.
The output enable must be registered too, and packed alongside. If miso_oe reaches the pad by a different route than the data, the two can skew: the driver may turn on before the data is valid or release after it should have gone high impedance. Both produce marginal, temperature-dependent contention. Most FPGA I/O blocks provide a dedicated output-enable register for exactly this reason.
The prefetch budget is a real timing constraint. fetch_req fires one byte early, so the memory has eight SCLK periods to respond. At 50 MHz that is 160 ns — ample for block RAM, tight for a slow external resource, and a genuine problem if the fetch crosses into another clock domain. Working out that number during design converts "the memory must be fast enough" into a checkable constraint.
On an ASIC, the pad's disable time is part of the bus turnaround. A pad may take a nanosecond or more to stop driving after its enable deasserts, and on a shared bus that time must be inside the dummy phase or the CS lag. It is one reason Dual and Quad modes specify more dummy cycles than the array access alone would justify.
10. Failure Signature — The First Byte Is Wrong and Every Byte After It Is Right
Symptom. A read returns a corrupted first byte followed by entirely correct data. Reproducible, identical every time, unaffected by clock rate. Writes to the same device work perfectly.
What the correct later bytes establish. This is the strongest evidence available and it eliminates most of the hypothesis space in one step. If bytes 2 onward are right, then the mode is correct — polarity, phase and bit order all produce correct bytes — and the framing, address handling and command decode are all correct too. A mode mismatch would corrupt every byte equally.
So the fault is specific to the first byte of the data phase, which is precisely §5's obligation.
Plausible mechanisms.
- The slave loads its output register on the first launch edge rather than on entry to the data phase, consuming that edge and shifting out a stale bit. This is the §6 branch-ordering bug.
- The dummy phase is one cycle short, so the master samples a bit before the device has presented it (Chapter 4.5 §7's zero-length case, and the off-by-one either side of it).
- The fetch did not complete before the data phase opened, so the register was loaded with stale or undefined content.
- The master is sampling one edge early on the first bit only, which is unusual but occurs with a misconfigured receive-delay setting (Chapter 2.7).
The discriminating observation. Lengthen the dummy phase by one cycle and re-read. If the first byte becomes correct and everything shifts into place, the fault was timing around the data-phase entry — either the fetch or the load. If the first byte is still wrong but now differently wrong, the device is presenting data correctly and the master's capture is at fault.
That test is decisive because the dummy count is the one parameter that moves the data-phase boundary without changing anything about the mode, framing or addressing.
Why the investigation goes wrong. Because a single wrong byte reads as noise or a glitch, and people look for an electrical cause — while the perfectly correct remainder is the very thing that rules electrical causes out. Reproducible corruption confined to exactly one byte position is a structural fault, and its position names the mechanism.
11. Common Misconceptions
12. Reason It Through
Work this before reading the answer.
A slave streams reads correctly at 1 MHz. At 25 MHz the first byte of every read is correct, and every byte from the second onward is corrupted — and the corruption gets worse the longer the burst runs. Writes work at both speeds.
What is happening?
Read the asymmetry carefully, because it is the reverse of §10. The first byte is right and later ones are wrong. So the data-phase entry, the load, the dummy count and the mode are all correct — every one of those would corrupt byte one. The fault appears only once the transfer starts crossing byte boundaries.
What happens at a byte boundary? §5: the device must have byte n+1 available when the last bit of byte n is launched. The fetch_req strobe fires one byte early precisely to allow this. So the suspect is the prefetch.
Why is it rate-dependent? Because the prefetch budget is measured in bus time, not system time. One byte at 1 MHz SCLK is 8 µs; at 25 MHz it is 320 ns. If the memory — or a clock-domain crossing to it, or an arbiter granting access to a shared resource — takes, say, 500 ns to return a byte, it comfortably meets the first deadline and misses every subsequent one.
Why does it get worse as the burst runs? Because the deficit accumulates. If each fetch arrives slightly late, the register is reloaded with stale data more often as the transfer proceeds, and once the device is a full byte behind it never recovers within the frame.
Why do writes still work? Because a write has no fetch at all — data flows into the device and is staged (Chapter 5.1). Nothing has to be produced on a deadline, so the write path is indifferent to memory latency. That writes work while reads fail is itself strong evidence pointing at the fetch path rather than anything about SPI.
What would confirm it? Measure the actual fetch turnaround in the design — from fetch_req to valid fetch_data — in system clock cycles, and compare against one byte time at 25 MHz. If the fetch takes longer than eight SCLK periods, the diagnosis is complete and arithmetic settles it without a capture.
The fixes, in order. Deepen the prefetch to more than one byte, using a small FIFO between the memory and the shift register so the fetch has several byte times to respond. Or reduce the maximum SCLK the device advertises, which is what a datasheet's maximum read frequency often is. Or move the fetch to a faster resource or a faster clock domain.
The general lesson. A streaming read imposes a continuous throughput requirement that a single-byte read does not reveal, and the budget shrinks as SCLK rises. "It works at low speed" on a burst read is evidence about the fetch path, not about signal integrity.
13. Understanding Check
14. Summary
A read has four phases — command, address, dummy, data — and reverses the direction of meaning once, at the end of the dummy phase. That reversal is what makes it structurally harder than a write.
It is not a write reversed. The device must be told what to read before it can read it, the data does not exist when the request finishes, and MISO is a line the device does not own by default.
The slave's obligation is continuous rather than atomic: from the moment the data phase opens until CS rises, the right bit must be at the output tap on every launch edge. In particular the first bit must be present before any edge, because the first launch edge shifts rather than presents — so the output register is loaded on entry to the data phase, not on the first strobe. Getting that branch ordering wrong loses exactly one bit per read and nothing else.
The same obligation recurs at every byte boundary, so a streaming read requires prefetch: the next byte is requested a full word ahead, giving the memory eight launch edges to respond. That converts memory speed into a checkable number of SCLK periods, and it is why burst reads impose a throughput requirement single-byte reads never reveal.
Bus ownership is the third obligation. Drive for exactly the data phase — early is contention, late corrupts the next transaction — and register both the data and the enable in the pad so they cannot skew.
For verification the properties worth writing are about ownership and the first-bit and prefetch obligations, while remembering that RTL cannot speak to the pad's analogue disable time. For coverage, the decisive bins are one byte versus two, because only the second crosses a boundary.
15. What Comes Next
This chapter took the command and address phases for granted and concentrated on what the device does once it knows what was asked. Chapter 6.2 — Command-Then-Read Sequences examines the first half: why every SPI read is a write first, what that costs in bus time and in master-side hardware, why the command and the read must usually occupy one frame rather than two, and the master-side sequencer that has to transmit a request and then receive a response without ever stopping the clock.
Continue learning
Related tutorials
- Related topic
CS-to-SCLK and SCLK-to-CS Timing
Chip select has timing requirements of its own: the lead before the first clock edge, the lag after the last, and the minimum deselect between transactions. Why violating them breaks a transfer whose every SCLK edge was correct.
- Related topic
Read Waveform Analysis
Decode an unlabelled read capture including the dummy phase: why the MISO tri-state transition is the most informative event on the bus, how to measure read latency without a datasheet, and checking versus discovery.
- Related topic
Continuous Transfers Under One CS
Holding chip select low across many bytes and what the device assumes: why the frame is the transaction, why the master may legally stop the clock when its data runs dry, and the streaming controller in three HDLs.
- Related topic
Slave Output Enable and MISO Tri-State
When a slave may drive MISO and when it must let go: the three independent reasons to stop, why release must be faster than assert, and the output-enable controller that makes the dead gap structural.
