SPI · Module 6
MISO Valid Timing
When returned data may be trusted relative to the launch edge: the round trip separating arrival from validity, why the sampling window has two edges, the configurable input sampler in three HDLs, and the margin no delay can create.
Chapter 6.2 established when the device's data arrives in the transaction — which byte times carry a response. This chapter asks a narrower and harder question.
The device has launched a bit. When may the master believe what it sees on the pin?
Those are different instants, and the gap between them is the whole subject. Module 2 built the arithmetic; here it becomes a configuration decision with hardware behind it, and one that has a wrong answer in both directions.
1. Arrival Is Not the Same as Valid
Trace one bit from the edge that launches it.
- The master's clock generator produces an SCLK edge at its pin.
- That edge propagates down the board to the device — flight time out.
- The device's output register responds, driving MISO after its clock-to-output delay —
t_v(Chapter 2.6). - The new level propagates back — flight time home.
- Only now is the bit present at the master's pin.
Call the total the round trip:
t_round = t_flight_out + t_v(slave) + t_flight_homeThe protocol says the master samples on the opposite edge — half a bit period after the launch. So the question of whether a read works at all is a comparison:
T_sclk / 2 versus t_round + t_setup(master)If the half period is comfortably larger, the nominal edge works. If it is not, the master is sampling a bit that has not arrived yet, and it captures the previous one.
2. The Window Has Two Edges
This is the part that gets missed, and the testbench in §4 asserts both edges rather than only the near one.
The near edge. Sampling before the bit arrives captures the previous bit — a one-position displacement across the whole word, indistinguishable at first glance from a phase mismatch §3.
The far edge. The bit is only on the wire for one bit period. Delay the sample past that and the device has already launched the next bit, so the capture belongs to the following bit position — the same one-position displacement, in the other direction.
So the usable window is bounded on both sides:
window opens at t_round + t_setup
window closes at t_round + T_sclk
width = T_sclk - t_setupThe width is one bit period minus the master's setup requirement — and notably it does not shrink as the round trip grows. A longer round trip moves the window later; it does not narrow it. That is precisely why a delay setting works: you are not buying margin, you are relocating a window that was always the same size.
What a longer round trip does cost is which bit the window belongs to. Once t_round exceeds a full bit period the window has moved past the bit boundary entirely, and the master must delay by more than a bit period and account for the resulting offset in its byte assembly. Controllers that support this describe it as sampling "one cycle late" — and Chapter 6.4 makes that accounting explicit.
3. A Worked Example
A master at 40 MHz, a device specifying t_v = 12 ns, a board contributing 2 ns each way, and a master needing 3 ns of setup:
T_sclk = 25 ns (40 MHz)
half period = 12.5 ns
t_round = 2 + 12 + 2 = 16 ns
required = 16 + 3 = 19 ns
19 ns > 12.5 ns → the nominal edge is 6.5 ns TOO EARLYThe link fails at the nominal sampling point. Now apply a delay measured in the master's own system clock — say 200 MHz, a 5 ns period:
delay 0 → sample at 12.5 ns → too early (19 ns needed)
delay 1 → 17.5 ns → still too early
delay 2 → 22.5 ns → inside the window ✓
delay 3 → 27.5 ns → past 12.5 + 25 = 37.5? no -- still inside
...
window closes at t_round + T_sclk = 41 ns → delay 6 (42.5 ns) is too lateSo delays 2 through 5 work, and the design should choose the middle of that range rather than its edge — giving margin against temperature, voltage and part-to-part variation in t_v.
Two conclusions worth carrying. The window is several system clocks wide, so the granularity of the delay setting matters: a controller whose delay steps are one SCLK period rather than one system clock may not be able to land inside it. And choosing the centre is a design decision, not a detail — picking the first delay that works leaves no margin on the near side.
4. Seeing It
The master samples later than the protocol edge
10 cyclesRead the marker as where the protocol says to look, and the sample pulses as where this master actually looks. At the nominal edge MISO is still carrying the previous bit; two cycles later it has settled.
5. Moving the Sampling Point — Three HDLs
The circuit
Circuit. A two-flop synchroniser on MISO, a shift-register delay line for the sampling strobe, and a multiplexer selecting which delayed strobe to use.
State. Two synchroniser flops and 2^DELAY_W strobe-pipeline flops.
Datapath. MISO is synchronised once and held; the strobe is what gets delayed. Delaying the strobe rather than the data is deliberate — it keeps exactly one synchronised copy of MISO and makes the delay a pure control choice, whereas a data delay line would need a full pipeline of MISO samples and would change the metastability analysis.
Control. sample_delay == 0 uses the nominal strobe directly; any other value selects a pipeline stage.
Clock. The master's system clock, which is what makes the delay granularity a system-clock period rather than an SCLK period.
Reset. Asynchronous, active-low, clearing both the synchroniser and the pipeline.
Enables. The pipeline runs continuously; it must, since it has to be holding strobes in flight.
Timing. bit_valid is a single-cycle strobe one clock after the delayed strobe. The synchroniser contributes two clocks of latency, which is part of the window arithmetic and not an implementation detail — §2's "window opens at t_round + t_setup" becomes, in system clocks, arrival + synchroniser latency.
Synthesis. 2 + 2^DELAY_W flip-flops and a multiplexer. The pipeline is the cost: DELAY_W = 5 means 32 flops, which is nothing, but the multiplexer feeding delayed_stb grows with it and sits in the control path.
Limitations. Delay granularity is one system clock, so a master whose system clock is not comfortably faster than SCLK cannot position the sample finely. That ratio is the real constraint and it is a Module 15 topic.
// spi_miso_sample.sv — the master's input capture, with a movable sampling point.
//
// The nominal sampling edge is where the PROTOCOL says to look. The round
// trip -- the master's clock out, the device's output delay, and the flight
// time back (Chapter 2.7) -- means the data may not have arrived yet. Real
// controllers therefore expose a receive-sample delay, and this is it.
//
// Note what the delay does NOT do: it cannot create margin that does not
// exist. It moves the sampling instant later within the bit period, which
// helps only while there is still bit period left to move into.
module spi_miso_sample #(
parameter int DELAY_W = 4 // sample_delay is 0 .. 2**DELAY_W-1
) (
input logic clk,
input logic rst_n,
input logic miso_async, // straight from the pin
input logic sample_stb, // the nominal sampling edge
input logic [DELAY_W-1:0] sample_delay, // extra system clocks
output logic bit_out,
output logic bit_valid // pulse: bit_out is fresh
);
localparam int MAXD = (1 << DELAY_W);
// MISO is asynchronous to this clock, so it crosses through a
// synchroniser before anything looks at it (Chapter 5.2 §5).
logic [1:0] miso_sync;
// The nominal strobe, delayed by up to MAXD-1 clocks. Delaying the STROBE
// rather than the data is what keeps a single synchronised copy of MISO
// and makes the delay a pure control choice.
logic [MAXD-1:0] stb_pipe;
logic delayed_stb;
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
miso_sync <= 2'b00;
stb_pipe <= '0;
bit_out <= 1'b0;
bit_valid <= 1'b0;
end else begin
miso_sync <= {miso_sync[0], miso_async};
stb_pipe <= {stb_pipe[MAXD-2:0], sample_stb};
bit_valid <= 1'b0;
if (delayed_stb) begin
bit_out <= miso_sync[1];
bit_valid <= 1'b1;
end
end
end
// Delay 0 uses the strobe directly; delay N uses the Nth pipeline stage.
assign delayed_stb = (sample_delay == '0) ? sample_stb
: stb_pipe[sample_delay - 1];
endmodule// spi_miso_sample_tb.sv — the sampling WINDOW has two edges.
//
// Stimulus models a real round trip: the strobe marks the nominal sampling
// edge, and MISO only settles ARRIVAL clocks later. The synchroniser adds a
// further SYNC_LAT clocks before the master can act on it, so the earliest
// delay that can possibly capture the intended bit is ARRIVAL + SYNC_LAT.
`timescale 1ns/1ps
module spi_miso_sample_tb;
logic clk = 0, rst_n = 0;
always #5 clk = ~clk;
localparam int DW = 5; // wide enough to reach the far edge
localparam int BIT_PERIOD = 10; // system clocks per SPI bit
localparam int ARRIVAL = 3; // clocks after the strobe that MISO settles
localparam int SYNC_LAT = 2; // the DUT's two-flop synchroniser
// The window opens once MISO has arrived AND cleared the synchroniser.
localparam int WIN_LO = ARRIVAL + SYNC_LAT;
// It stays open for one whole bit period. At WIN_LO + BIT_PERIOD the
// capture and the synchroniser update on the SAME edge, so the capture
// still sees the old (intended) value -- the first genuinely late delay
// is one beyond that.
localparam int WIN_HI = WIN_LO + BIT_PERIOD; // last usable delay
localparam int TOO_LATE = WIN_HI + 1;
logic miso_async = 0, sample_stb = 0;
logic [DW-1:0] sample_delay = '0;
logic bit_out, bit_valid;
spi_miso_sample #(.DELAY_W(DW)) dut (
.clk, .rst_n, .miso_async, .sample_stb, .sample_delay, .bit_out, .bit_valid);
int errors = 0;
logic armed, first_val, first_seen;
task automatic chk(input string what, input int got, input int exp);
if (got !== exp) begin $display("FAIL %s: got %0d exp %0d", what, got, exp); errors++; end
endtask
// Record the FIRST capture after arming, so the result is independent of
// how long the stimulus runs afterwards.
always @(posedge clk) if (rst_n && armed && bit_valid && !first_seen) begin
first_val <= bit_out;
first_seen <= 1'b1;
end
task automatic one_bit(input logic new_val);
@(negedge clk); sample_stb = 1;
@(negedge clk); sample_stb = 0;
repeat (ARRIVAL - 1) @(negedge clk);
miso_async = new_val;
repeat (BIT_PERIOD - ARRIVAL) @(negedge clk);
endtask
// Drive: previous bit = 0, bit A = 1, bit B = 0. Report what the first
// capture after bit A's strobe returned.
task automatic probe(input int d, output logic got);
sample_delay = d[DW-1:0];
miso_async = 1'b0;
armed = 0; first_seen = 0; first_val = 1'bx;
repeat (2) @(negedge clk);
armed = 1;
one_bit(1'b1); // bit A = 1
one_bit(1'b0); // bit B = 0
repeat (BIT_PERIOD) @(negedge clk); // let any late strobe land
armed = 0;
got = first_val;
endtask
logic got;
initial begin
repeat (3) @(negedge clk); rst_n = 1; @(negedge clk);
$display("window for these parameters: delays %0d..%0d inclusive", WIN_LO, WIN_HI);
// --- TOO EARLY: MISO has not arrived (or not cleared the
// synchroniser), so the capture returns the PREVIOUS bit. ---
for (int d = 0; d < WIN_LO; d++) begin
probe(d, got);
chk($sformatf("delay %0d is too early -> stale bit", d), got, 0);
end
// --- INSIDE THE WINDOW: the intended bit. ---
for (int d = WIN_LO; d < WIN_LO + 4; d++) begin
probe(d, got);
chk($sformatf("delay %0d is in the window -> correct bit", d), got, 1);
end
// --- TOO LATE: the strobe has run past the bit boundary and lands
// inside the NEXT bit's window. More delay is not always better. ---
// The last delay that still works, to pin the far edge exactly.
probe(WIN_HI, got);
chk($sformatf("delay %0d is the last usable one", WIN_HI), got, 1);
probe(TOO_LATE, got);
chk($sformatf("delay %0d is too late -> next bit", TOO_LATE), got, 0);
if (errors == 0)
$display("PASS: the sampling window has both edges -- delays below arrival+sync capture the previous bit, delays past the bit boundary capture the next one, and only delays inside the window return the intended bit");
else
$display("FAILED with %0d error(s)", errors);
$finish;
end
initial begin #900000; $display("FAIL: watchdog timeout"); $finish; end
endmoduleThat testbench is worth reading as a measurement rather than a check. It models a real round trip — the strobe marks the nominal edge, MISO settles ARRIVAL clocks later — then sweeps the delay and asserts three regions: below ARRIVAL + SYNC_LAT the capture returns the previous bit, inside the window it returns the intended one, and past WIN_LO + BIT_PERIOD it returns the next. All three languages report the same window for the same parameters, which is the parity claim made concrete.
The far edge is subtle enough to be worth stating precisely: at exactly WIN_LO + BIT_PERIOD the capture and the synchroniser update on the same clock edge, so the capture still sees the intended (pre-edge) value. The first genuinely late delay is one beyond that, and the testbench pins both.
// spi_miso_sample.v — the same movable input capture in Verilog-2001.
module spi_miso_sample #(
parameter DELAY_W = 4
) (
input wire clk,
input wire rst_n,
input wire miso_async,
input wire sample_stb,
input wire [DELAY_W-1:0] sample_delay,
output reg bit_out,
output reg bit_valid
);
localparam MAXD = (1 << DELAY_W);
// MISO is asynchronous to this clock and crosses through a synchroniser
// before anything looks at it.
reg [1:0] miso_sync;
// Delaying the STROBE rather than the data keeps one synchronised copy of
// MISO and makes the delay a pure control choice.
reg [MAXD-1:0] stb_pipe;
wire delayed_stb;
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
miso_sync <= 2'b00;
stb_pipe <= {MAXD{1'b0}};
bit_out <= 1'b0;
bit_valid <= 1'b0;
end else begin
miso_sync <= {miso_sync[0], miso_async};
stb_pipe <= {stb_pipe[MAXD-2:0], sample_stb};
bit_valid <= 1'b0;
if (delayed_stb) begin
bit_out <= miso_sync[1];
bit_valid <= 1'b1;
end
end
end
assign delayed_stb = (sample_delay == {DELAY_W{1'b0}})
? sample_stb
: stb_pipe[sample_delay - 1'b1];
endmodule// spi_miso_sample_tb.v — the same window characterisation in Verilog-2001.
`timescale 1ns/1ps
module spi_miso_sample_tb;
reg clk = 0, rst_n = 0;
always #5 clk = ~clk;
parameter DW = 5;
parameter BIT_PERIOD = 10;
parameter ARRIVAL = 3;
parameter SYNC_LAT = 2;
localparam WIN_LO = ARRIVAL + SYNC_LAT;
localparam WIN_HI = WIN_LO + BIT_PERIOD;
localparam TOO_LATE = WIN_HI + 1;
reg miso_async = 0, sample_stb = 0;
reg [DW-1:0] sample_delay = 0;
wire bit_out, bit_valid;
spi_miso_sample #(.DELAY_W(DW)) dut (
.clk(clk), .rst_n(rst_n), .miso_async(miso_async), .sample_stb(sample_stb),
.sample_delay(sample_delay), .bit_out(bit_out), .bit_valid(bit_valid));
integer errors = 0, d;
reg armed, first_val, first_seen, got;
task chk;
input integer delay_n;
input [31:0] got_v, exp_v;
begin
if (got_v !== exp_v) begin
$display("FAIL delay %0d: got %0d exp %0d", delay_n, got_v, exp_v);
errors = errors + 1;
end
end
endtask
always @(posedge clk) if (rst_n && armed && bit_valid && !first_seen) begin
first_val <= bit_out;
first_seen <= 1'b1;
end
task one_bit;
input new_val;
begin
@(negedge clk); sample_stb = 1;
@(negedge clk); sample_stb = 0;
repeat (ARRIVAL - 1) @(negedge clk);
miso_async = new_val;
repeat (BIT_PERIOD - ARRIVAL) @(negedge clk);
end
endtask
task probe;
input integer dly;
begin
sample_delay = dly[DW-1:0];
miso_async = 1'b0;
armed = 0; first_seen = 0; first_val = 1'bx;
repeat (2) @(negedge clk);
armed = 1;
one_bit(1'b1);
one_bit(1'b0);
repeat (BIT_PERIOD) @(negedge clk);
armed = 0;
got = first_val;
end
endtask
initial begin
repeat (3) @(negedge clk); rst_n = 1; @(negedge clk);
$display("window for these parameters: delays %0d..%0d inclusive", WIN_LO, WIN_HI);
for (d = 0; d < WIN_LO; d = d + 1) begin
probe(d); chk(d, got, 0);
end
for (d = WIN_LO; d < WIN_LO + 4; d = d + 1) begin
probe(d); chk(d, got, 1);
end
probe(WIN_HI); chk(WIN_HI, got, 1);
probe(TOO_LATE); chk(TOO_LATE, got, 0);
if (errors == 0)
$display("PASS: the sampling window has both edges -- delays below arrival+sync capture the previous bit, delays past the bit boundary capture the next one, and only delays inside the window return the intended bit");
else
$display("FAILED with %0d error(s)", errors);
$finish;
end
initial begin #900000; $display("FAIL: watchdog timeout"); $finish; end
endmodule-- spi_miso_sample.vhd — the same movable input capture in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_miso_sample is
generic (
DELAY_W : positive := 4 -- sample_delay is 0 .. 2**DELAY_W-1
);
port (
clk : in std_logic;
rst_n : in std_logic;
miso_async : in std_logic; -- straight from the pin
sample_stb : in std_logic; -- nominal sampling edge
sample_delay : in unsigned(DELAY_W - 1 downto 0); -- extra system clocks
bit_out : out std_logic;
bit_valid : out std_logic -- pulse
);
end entity spi_miso_sample;
architecture rtl of spi_miso_sample is
constant MAXD : positive := 2 ** DELAY_W;
-- MISO is asynchronous to this clock and crosses through a synchroniser
-- before anything looks at it.
signal miso_sync : std_logic_vector(1 downto 0);
-- Delaying the STROBE rather than the data keeps one synchronised copy of
-- MISO and makes the delay a pure control choice.
signal stb_pipe : std_logic_vector(MAXD - 1 downto 0);
signal delayed_stb : std_logic;
begin
delayed_stb <= sample_stb when sample_delay = 0
else stb_pipe(to_integer(sample_delay) - 1);
process (clk, rst_n) is
begin
if rst_n = '0' then
miso_sync <= (others => '0');
stb_pipe <= (others => '0');
bit_out <= '0';
bit_valid <= '0';
elsif rising_edge(clk) then
miso_sync <= miso_sync(0) & miso_async;
stb_pipe <= stb_pipe(MAXD - 2 downto 0) & sample_stb;
bit_valid <= '0';
if delayed_stb = '1' then
bit_out <= miso_sync(1);
bit_valid <= '1';
end if;
end if;
end process;
end architecture rtl;-- spi_miso_sample_tb.vhd — the same window characterisation in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_miso_sample_tb is
end entity spi_miso_sample_tb;
architecture tb of spi_miso_sample_tb is
constant DW : positive := 5;
constant BIT_PERIOD : positive := 10;
constant ARRIVAL : positive := 3;
constant SYNC_LAT : positive := 2;
-- The window opens once MISO has arrived AND cleared the synchroniser,
-- and stays open for one bit period. At WIN_HI the capture and the
-- synchroniser update on the same edge, so WIN_HI is still usable.
constant WIN_LO : positive := ARRIVAL + SYNC_LAT;
constant WIN_HI : positive := WIN_LO + BIT_PERIOD;
constant TOO_LATE : positive := WIN_HI + 1;
signal clk : std_logic := '0';
signal rst_n : std_logic := '0';
signal miso_async : std_logic := '0';
signal sample_stb : std_logic := '0';
signal sample_delay : unsigned(DW - 1 downto 0) := (others => '0');
signal halt : boolean := false;
signal bit_out : std_logic;
signal bit_valid : std_logic;
signal errors : natural := 0;
signal armed : std_logic := '0';
signal first_val : std_logic := 'X';
signal first_seen : std_logic := '0';
signal arm_clr : std_logic := '0';
begin
clk <= not clk after 5 ns when not halt else '0';
dut : entity work.spi_miso_sample
generic map (DELAY_W => DW)
port map (clk => clk, rst_n => rst_n, miso_async => miso_async,
sample_stb => sample_stb, sample_delay => sample_delay,
bit_out => bit_out, bit_valid => bit_valid);
-- One process owns first_val/first_seen; stim only requests a clear.
capture : process (clk) is
begin
if rising_edge(clk) then
if arm_clr = '1' then
first_seen <= '0';
first_val <= 'X';
elsif rst_n = '1' and armed = '1' and bit_valid = '1'
and first_seen = '0' then
first_val <= bit_out;
first_seen <= '1';
end if;
end if;
end process;
stim : process is
variable got : std_logic;
procedure chk (delay_n : natural; got_v, exp_v : std_logic) is
begin
if got_v /= exp_v then
report "FAIL delay " & integer'image(delay_n)
& ": got " & std_logic'image(got_v)
& " exp " & std_logic'image(exp_v) severity error;
errors <= errors + 1;
end if;
end procedure;
procedure one_bit (new_val : std_logic) is
begin
wait until falling_edge(clk); sample_stb <= '1';
wait until falling_edge(clk); sample_stb <= '0';
for i in 0 to ARRIVAL - 2 loop
wait until falling_edge(clk);
end loop;
miso_async <= new_val;
for i in 0 to BIT_PERIOD - ARRIVAL - 1 loop
wait until falling_edge(clk);
end loop;
end procedure;
procedure probe (dly : natural) is
begin
sample_delay <= to_unsigned(dly, DW);
miso_async <= '0';
armed <= '0';
arm_clr <= '1';
wait until falling_edge(clk);
arm_clr <= '0';
wait until falling_edge(clk);
armed <= '1';
one_bit('1'); -- bit A = 1
one_bit('0'); -- bit B = 0
for i in 0 to BIT_PERIOD - 1 loop
wait until falling_edge(clk);
end loop;
armed <= '0';
got := first_val;
end procedure;
begin
for i in 0 to 2 loop
wait until falling_edge(clk);
end loop;
rst_n <= '1';
wait until falling_edge(clk);
report "window for these parameters: delays " & integer'image(WIN_LO)
& ".." & integer'image(WIN_HI) & " inclusive" severity note;
for d in 0 to WIN_LO - 1 loop
probe(d); chk(d, got, '0');
end loop;
for d in WIN_LO to WIN_LO + 3 loop
probe(d); chk(d, got, '1');
end loop;
probe(WIN_HI); chk(WIN_HI, got, '1');
probe(TOO_LATE); chk(TOO_LATE, got, '0');
if errors = 0 then
report "PASS: the sampling window has both edges -- delays below "
& "arrival+sync capture the previous bit, delays past the bit "
& "boundary capture the next one, and only delays inside the "
& "window return the intended bit" severity note;
else
report "FAILED with " & integer'image(errors) & " error(s)" severity error;
end if;
halt <= true;
wait;
end process;
watchdog : process is
begin
wait for 900 us;
if not halt then
report "FAIL: watchdog timeout" severity failure;
end if;
wait;
end process;
end architecture tb;Parity
All three implement the same hardware: identical ports, asynchronous active-low reset, a two-flop synchroniser on MISO, a 2^DELAY_W-stage strobe pipeline, delay zero bypassing the pipeline, and a single-cycle valid strobe. The VHDL takes sample_delay as unsigned where the Verilog dialects use a plain vector — a typing difference only. All three testbenches sweep the same delays with the same stimulus model and report the identical window, delays 5 to 15 inclusive, with 16 the first that is too late.
6. What the Delay Cannot Do
Worth being blunt about, because "add sample delay" gets applied as a universal remedy.
It cannot create margin. §2 showed the window's width is T_sclk - t_setup, independent of the round trip. If the bit period itself is too short — if t_setup approaches T_sclk — the window is narrow or empty at every delay value, and no setting helps. The fix there is a slower clock.
It cannot fix a mode error. A phase mismatch also displaces data by one position, and it is tempting to "correct" it by shifting the sample point. Sometimes that appears to work, which is the worst possible outcome: the link now depends on two errors cancelling, and it will stop working at a different frequency or on a different board. Check the idle level and the transition clustering before touching the delay.
It cannot compensate for a wrong dummy count. That displaces data by whole bits or bytes at the transaction level; the sample delay operates within a single bit period. Using one to hide the other produces a design that works at exactly one clock rate.
The honest use of a sample delay is narrow and important: the round trip is real, measurable, and larger than half a bit period, and the window has simply moved. That is a genuine and common condition at high SCLK, and it is what the setting exists for.
7. Why a Verification Engineer Cares
The properties here are about the strobe discipline, because the analogue reality is out of reach in RTL:
// 1. One capture per nominal strobe -- no more, no fewer. A pipeline bug
// that produced two captures would double-shift the receive register.
property p_one_capture_per_strobe;
@(posedge clk) disable iff (!rst_n)
sample_stb |-> ##[1:MAXD] bit_valid;
endproperty
// 2. bit_valid is a pulse. Held for two cycles it would insert a bit.
a_valid_is_pulse : assert property (
@(posedge clk) disable iff (!rst_n) bit_valid |=> !bit_valid)
else $error("bit_valid held for more than one cycle");
// 3. Nothing may read the raw pin. This is a STRUCTURAL rule and the most
// valuable check here -- see below.
// (enforced by lint, not by a property)
// 4. The delay must not change mid-transfer. Moving the sampling point
// between bits of the same word reorders them.
property p_delay_stable_in_frame;
@(posedge clk) disable iff (!rst_n)
!cs_n |=> $stable(sample_delay);
endproperty
a_delay_stable : assert property (p_delay_stable_in_frame)
else $error("sample_delay changed while a frame was open");What these prove. That the sampler produces exactly one well-formed capture per nominal strobe and that the configuration is stable across a frame.
What they emphatically do not prove is that the chosen delay is correct. The window's position depends on t_v, board flight time and the master's setup requirement — all analogue quantities that an RTL simulation does not model. A simulation with a zero-delay device model will show delay 0 working perfectly and every non-zero delay failing, which is the exact opposite of the truth on a real board.
That gap is worth naming because it misleads reliably. The way to verify a sample-delay choice is a timing analysis against the datasheet's t_v and the board's extracted flight time, or a bench sweep on real hardware. What simulation can contribute is the check that the mechanism works and that the window is where the arithmetic says — which is exactly what the §5 testbenches do, by injecting an explicit ARRIVAL delay into the model.
covergroup spi_miso_delay_cg @(posedge transfer_done);
cp_delay : coverpoint cfg.sample_delay {
bins zero = {0}; // the nominal edge
bins small = {[1:3]};
bins large = {[4:15]};
bins beyond_bit = {[16:$]}; // past a whole bit period -- needs
// transaction-level re-alignment
}
// The delay is only meaningful relative to the round trip the model
// is injecting. Covering the delay alone proves nothing.
cp_arrival : coverpoint env_model.arrival_clocks {
bins fast = {[0:2]}; // nominal edge would have worked
bins medium = {[3:8]};
bins slow = {[9:$]}; // past a bit period
}
x_delay_arrival : cross cp_delay, cp_arrival;
endgroupThe cross is the point, and it generalises. A configuration value that exists to compensate for an environmental condition is meaningless to cover on its own — the bin that matters is this delay with that round trip, and in particular a large delay with a fast arrival, which is the over-correction case that captures the next bit.
8. Why an FPGA or ASIC Engineer Cares
The system-clock-to-SCLK ratio decides whether you can position the sample at all. Delay granularity is one system clock. At 100 MHz system and 25 MHz SCLK there are four steps per bit period — usable but coarse. At 50 MHz system and 25 MHz SCLK there are two, and the window may fall between them. Designing the interface into a fast clock domain is what buys positioning resolution, and it is the same ratio argument as Chapter 5.2 §12's synchroniser latency.
Some FPGA families offer sub-cycle delay primitives. An input delay element on the pin can shift MISO by picosecond-scale taps, giving far finer positioning than a system-clock pipeline. It is the right tool when the window is narrower than a system clock, and it is a different mechanism with its own calibration requirements — not a replacement for understanding where the window is.
Constrain the input path. set_input_delay on MISO, derived from the device's t_v and the board flight time, is what lets static timing check the capture. Without it the tool has no idea when MISO arrives and will report the path as met regardless. A sample-delay setting chosen by experiment and never constrained is a latent failure waiting for a faster part or a hotter day.
Do not read the raw pin anywhere. MISO is asynchronous. The synchroniser exists so that exactly one place samples the pin; every other consumer must take miso_sync. This is the same structural rule as Chapter 5.2 §8, and it is enforceable by lint rather than by simulation, because a simulator will never show the difference.
9. Failure Signature — Correct Below a Threshold Frequency, Shifted Above It
Symptom. Reads are perfect at 10 MHz. At 30 MHz every returned byte is displaced by one bit position. The transition is sharp — 25 MHz works, 30 MHz does not — and it is reproducible rather than intermittent. Writes work at every rate.
What the sharp threshold establishes. This is the key discriminator and it is available before any instrument is attached. A margin problem degrades progressively: the error rate rises gradually, varies with temperature, and produces random corruption. A clean transition at a specific frequency, with a fixed displacement on the far side, is a structural boundary being crossed.
Why writes are unaffected. A write has no return path. The master launches MOSI and the device samples it — one direction, no round trip (Chapter 6.1 §1). That reads fail while writes succeed at the same clock rate is strong evidence for a return-path cause and rules out the mode, the framing and the clock generation in one observation.
Plausible mechanisms.
- The round trip has exceeded half a bit period, so the nominal sampling edge is now too early — this chapter's subject, and the leading candidate given the threshold behaviour.
- The device's
t_vis specified only up to a lower maximum frequency, so above it the part is out of spec. - A dummy count that should have increased with frequency (Chapter 4.5 §3) — but that displaces by whole cycles at the transaction level and would usually be a larger offset.
The discriminating observation. Compute the threshold. If the round trip is t_round, the nominal edge fails when T_sclk / 2 < t_round + t_setup, so the predicted failure frequency is f = 1 / (2 × (t_round + t_setup)). Substituting the datasheet's t_v and an estimate of board delay gives a number — and if it lands between the working and failing rates, the diagnosis is confirmed arithmetically before touching the board.
Then apply a receive-sample delay of one or two system clocks and retry at 30 MHz. If the link comes back, the window had simply moved.
Why the investigation goes wrong. Because "fails at higher frequency" reads as signal integrity, and the team starts probing edges and adding termination. Signal integrity produces progressive degradation and random errors; this produces a sharp threshold and a fixed displacement. Distinguishing them costs one observation — plot the error rate against frequency and look at the shape — and it redirects the whole investigation.
10. Common Misconceptions
11. Reason It Through
Work this before reading the answer.
A board reads a flash device at 50 MHz. It works at room temperature on every unit built so far. A new batch of the flash — same part number, different date code — fails on about a third of boards, always with data displaced by one bit position, always at 50 MHz, always working at 20 MHz.
What is going on, and what is the right fix?
Start with what varies and what does not. The board is unchanged, the master is unchanged, and the failure correlates with the device batch. So the variable is inside the flash — and the parameter of the flash that governs return-path timing is t_v, its clock-to-output delay.
Why would a new batch differ? Because t_v is specified as a maximum, not a value. A datasheet promising t_v ≤ 8 ns permits any part that meets it, and different process lots legitimately land at different points under that ceiling. An earlier batch clustering near 5 ns and a new one near the full 8 ns are both entirely in specification.
Now the arithmetic. At 50 MHz the half period is 10 ns. With 2 ns of flight each way and 3 ns of master setup:
old batch: 2 + 5 + 2 + 3 = 12 ns > 10 ns → should have failed too
new batch: 2 + 8 + 2 + 3 = 15 ns > 10 ns → failsAnd that is the real finding. Even the old batch exceeded the half period. The original design never had margin — it worked because the actual t_v happened to sit far enough below its maximum, and because a marginal capture often still resolves correctly. The new batch did not introduce the fault; it revealed a design that was out of budget from the start and passing by luck.
Why a third of boards? Because the population spreads. Board-to-board flight time, supply voltage, temperature and the individual die all vary, so a design sitting just past the boundary fails for the part of the distribution that lands on the wrong side. A fault affecting some units at a fixed frequency is the signature of a budget that is marginal rather than broken.
The right fix. Apply a receive-sample delay sized from the worst-case t_v — the datasheet maximum, not the measured typical — so that the sampling point is inside the window for every legal part. That is exactly what the setting is for, it costs nothing, and it makes the design correct by specification rather than by lot.
The wrong fixes, and why. Screening the flash by date code makes the design dependent on a supplier's process distribution. Dropping to 20 MHz works and gives up more than half the read bandwidth for a problem that has a free fix. And "it worked before" is the reasoning that produced the situation.
The general lesson. When a design that "always worked" fails on a new batch, the first hypothesis should be that it was never inside its budget and was passing on the strength of typical values. Recompute the budget with datasheet maxima; if it does not close, the batch is the messenger.
12. Understanding Check
13. Summary
A bit's arrival and its validity at the master are separated by the round trip: flight out, the device's clock-to-output delay, and flight home. The nominal sampling edge — half a bit period after the launch — works only if the half period exceeds that round trip plus the master's setup requirement.
The master's capture instant is local and configurable. The device launches on an SCLK edge and cannot see when the master looks, which is why a receive-sample delay is legitimate rather than a hack.
The window has two edges: too early captures the previous bit, too late captures the next. Its width is T_sclk − t_setup and is independent of the round trip — a long round trip moves the window rather than narrowing it, which is exactly why relocating the sample point works.
In RTL the mechanism is a synchroniser on MISO plus a delay line on the strobe — delaying the strobe rather than the data keeps one crossing point and one synchronised copy. The synchroniser's own latency is part of the window arithmetic: it opens at arrival + sync latency and stays open for one bit period, a result all three implementations reproduce identically.
What the delay cannot do is create margin, fix a mode error, or compensate for a wrong dummy count. Using it for any of those produces a link that depends on two errors cancelling.
And simulation cannot confirm a delay choice, because the quantities that place the window are analogue. Verify by timing analysis against datasheet maxima, or on hardware — and when a design that always worked fails on a new batch, suspect that it was never inside its budget.
14. What Comes Next
This chapter measured latency within one bit. Chapter 6.4 — Read Latency Accounting zooms out to the whole transaction and counts every cycle between asking for data and holding it: the command, the address, the dummy phase, the payload itself, and the round trip this chapter just characterised — in SCLK cycles, with the arithmetic that tells you what a read actually costs and where the master must realign its receive buffer when the sampling point has moved past a bit boundary.
Continue learning
Related tutorials
- Related topic
Setup, Hold, and Timing Margin
What makes a bit electrically safe to capture. The stability window a receiver demands, why violating it produces undefined rather than merely wrong results, and which engineering layer owns each part of the answer.
- Related topic
Master Input Capture and Round-Trip Delay
The complete return-path budget: clock out, peripheral response, data back, and the master's setup requirement, all inside half a period. Why maximum SCLK is a property of a whole system.
- Related topic
Timing, Constraints, and CDC Considerations
The MISO round trip decides the maximum SCLK rate and static timing analysis never checks it, plus why this architecture has no internal clock-domain crossing and what simulation cannot establish about reset release.
- Related topic
SCLK Generation, Period, and Frequency
Where SCLK comes from and what one period buys. Dividing a system clock to a bus clock, why the divisor is an integer and what that costs, and how a period in nanoseconds becomes the budget every later timing parameter is spent from.
