SPI · Module 6
Read Latency Accounting
Every cycle between asking for data and holding it: command, address, dummy, sampling skew and round trip counted in SCLK cycles, why receive alignment reduces to one number, and the one-byte displacement a miscount produces.
Chapter 6.3 measured latency inside one bit. This chapter zooms out to the whole transaction and counts everything.
From the instant the master decides it wants a byte, how many SCLK cycles pass before that byte is in its receive register — and what does the master have to know in order to find the byte at all?
The second half is the practical one. The master must discard exactly the right number of received bits before the payload starts, and that number is a sum of contributions from five different places. Get it wrong by one and every byte of the response is displaced.
1. The Five Contributions
Each contribution comes from a different chapter, which is why this is the module's accounting chapter.
| Contribution | Measured in | Source |
|---|---|---|
| Command | bits | opcode length × 8 (Chapter 4.4) |
| Address | bits | address bytes × 8 (Chapter 4.4) |
| Dummy | cycles | the device's fetch time ÷ T_sclk (Chapter 4.5) |
| Sampling skew | cycles | whether the round trip pushed the sample past a bit boundary (Chapter 6.3) |
| Frame overhead | time, not cycles | CS lead and lag (Chapter 2.5) |
Note the units. The first four are bit or cycle counts on SCLK; the last is wall-clock time that does not correspond to any clocked bit. Mixing them is the commonest arithmetic error in this area, and it is why the master's alignment number counts only the first four.
2. Counting in SCLK Cycles
Where the cycles go before the first payload bit
8 cyclesThe cycles lane is not a signal — it is the running total, drawn because it is the quantity this chapter is about. At the boundary where the payload starts it reads 40, which is exactly the number the master must use.
3. The Master's Single Number
Everything reduces to:
skip_bits = cmd_bytes × 8
+ addr_bytes × 8
+ dummy_cycles
+ sampling_skew_bitsFor the figure above, with no sampling skew:
skip_bits = 1×8 + 3×8 + 8 + 0 = 40The master discards 40 received bits, then assembles payload words from bit 41 onward. That is the whole of §5's module.
Three observations about this formula are worth more than the formula.
Dummy is added in cycles, not bytes. Chapter 4.5 §3 established this, and here is where it bites: a device needing 6 dummy cycles contributes 6, not 8. A master that expresses everything in bytes cannot represent it and will be wrong by two bits.
The sampling skew is usually zero and occasionally not. Chapter 6.3 §2 showed that when the round trip exceeds a full bit period, the master's sample point moves past a bit boundary — and then it is receiving bit n while the protocol says bit n+1, so one extra bit must be discarded. Controllers describe this as sampling "one cycle late", and it is the contribution most often forgotten because it comes from board physics rather than from the datasheet's command table.
Everything is on SCLK. CS lead and lag are real time and real latency, but they clock no bits, so they do not appear here. They appear in §4's time calculation and not in the alignment.
4. What a Read Actually Costs
Now the same sum as a duration. For the transaction above at 50 MHz, with a 100 ns CS lead and 100 ns lag:
T_sclk = 20 ns
request = 40 cycles × 20 ns = 800 ns
one payload byte = 8 cycles × 20 ns = 160 ns
CS lead + lag = 200 ns
───────────────────────────────────────────────────
single-byte read = 1160 nsLatency and throughput diverge sharply here, and conflating them is the usual error.
payload total time effective throughput
1 byte 1160 ns 0.86 MB/s
4 bytes 1640 ns 2.44 MB/s
32 bytes 6120 ns 5.23 MB/s
256 bytes 41960 ns 6.10 MB/sThe bus is nominally 50 Mbit/s = 6.25 MB/s. A 256-byte burst reaches 98% of that; a single-byte read reaches 14%. Same bus, same clock, same device.
The random-access case is what most systems actually do, and it is the worst one. Reading one status byte at a time from a sensor spends 800 ns asking and 160 ns receiving. If the application does that in a loop, the effective data rate is under a megabyte per second on a 50 MHz bus — and raising SCLK does not change the ratio, as Chapter 6.2 §4 showed.
The lever that works is amortisation: fewer, larger reads. If the device supports auto-incrementing reads, reading 32 bytes in one transaction instead of 32 transactions of one byte is a six-fold improvement in effective throughput, for no change in clock rate, no board work and no risk.
5. Building the Receive Aligner — Three HDLs
The circuit
Circuit. A skip counter, a shift register, a bit counter, and a registered status flag.
State. How many bits have been discarded so far, the partially-assembled word, the bit position within it, and whether alignment is complete.
Datapath. Every captured bit either increments the skip counter or is shifted into the word register — never both. That mutual exclusion is the module.
Control. frame_start restarts the accounting unconditionally. This is the master-side mirror of Chapter 4.4 §7: CS is the only resynchronising event at either end.
Clock. The master's system clock; bit_valid is the strobe from Chapter 6.3's sampler.
Reset. Asynchronous, active-low, to a cleared counter with alignment false.
Enables. Nothing moves except on frame_start or bit_valid.
Timing. aligned is registered rather than a continuous comparison. A continuous compare against an undefined counter evaluates before reset and is flagged by some simulators; keeping the decision inside the clocked process means it is only evaluated once the counter is defined. All three implementations were restructured together so the semantics stay identical.
Synthesis. A CNT_W-bit counter and comparator, a WIDTH-bit shift register, a small bit counter. The comparator's second input is a register rather than a constant — skip_bits is runtime-configurable because the dummy count and the skew both vary — so it cannot be folded away, the same trade as Chapter 4.2 §8.
Limitations. Fixed word width, no FIFO, and one outstanding frame.
// spi_rx_align.sv — the master's receive alignment.
//
// Everything this module does reduces to one number: how many received bits
// to throw away before the payload begins. That number is the SUM of every
// latency contribution in the transaction --
//
// skip_bits = command bits + address bits + dummy cycles + sampling skew
//
// -- and getting it wrong by one displaces every byte of the response. This
// is read latency accounting expressed as hardware: the arithmetic of
// Chapter 6.4 is literally the value driven onto `skip_bits`.
module spi_rx_align #(
parameter int WIDTH = 8,
parameter int CNT_W = 12 // enough for a long command+address+dummy
) (
input logic clk,
input logic rst_n,
input logic frame_start, // pulse: CS asserted, restart counting
input logic [CNT_W-1:0] skip_bits, // bits to discard before the payload
input logic bit_valid, // pulse: bit_in is a captured bit
input logic bit_in,
output logic [WIDTH-1:0] byte_out,
output logic byte_valid, // pulse: byte_out is a payload word
output logic aligned // level: the skip is complete
);
localparam int BW = (WIDTH > 1) ? $clog2(WIDTH) : 1;
logic [CNT_W-1:0] skipped;
logic [WIDTH-1:0] shreg;
logic [BW-1:0] bit_cnt;
// `aligned` is REGISTERED rather than a continuous compare. A continuous
// comparison against an undefined counter evaluates before reset, which
// some simulators flag; keeping the decision inside the clocked process
// means it is only evaluated once the counter is defined.
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
skipped <= '0;
shreg <= '0;
bit_cnt <= '0;
byte_out <= '0;
byte_valid <= 1'b0;
aligned <= 1'b0;
end else begin
byte_valid <= 1'b0;
if (frame_start) begin
// A new frame restarts the whole accounting. This is the only
// resynchronising event the master has, exactly as CS is the
// only one the device has (Chapter 4.4 §7).
skipped <= '0;
bit_cnt <= '0;
shreg <= '0;
aligned <= (skip_bits == '0);
end else if (bit_valid) begin
if (skipped < skip_bits) begin
// Still inside the command / address / dummy region.
// The bit is counted and discarded, not shifted.
skipped <= skipped + 1'b1;
if (skipped + 1'b1 >= skip_bits) aligned <= 1'b1;
end else begin
shreg <= {shreg[WIDTH-2:0], bit_in}; // MSB-first
if (bit_cnt == BW'(WIDTH - 1)) begin
bit_cnt <= '0;
byte_out <= {shreg[WIDTH-2:0], bit_in};
byte_valid <= 1'b1;
end else begin
bit_cnt <= bit_cnt + 1'b1;
end
end
end
end
end
endmodule// spi_rx_align_tb.sv — the skip total, an off-by-one, and a re-armed frame.
`timescale 1ns/1ps
module spi_rx_align_tb;
logic clk = 0, rst_n = 0;
always #5 clk = ~clk;
localparam int W = 8, CW = 12;
logic frame_start = 0, bit_valid = 0, bit_in = 0;
logic [CW-1:0] skip_bits = '0;
logic [W-1:0] byte_out;
logic byte_valid, aligned;
spi_rx_align #(.WIDTH(W), .CNT_W(CW)) dut (
.clk, .rst_n, .frame_start, .skip_bits, .bit_valid, .bit_in,
.byte_out, .byte_valid, .aligned);
int errors = 0;
logic [W-1:0] got [$];
task automatic chk(input string what, input int g, input int e);
if (g !== e) begin $display("FAIL %s: got 0x%0h exp 0x%0h", what, g, e); errors++; end
endtask
always @(posedge clk) if (rst_n && byte_valid) got.push_back(byte_out);
task automatic send_bit(input logic b);
bit_in = b; @(negedge clk); bit_valid = 1; @(negedge clk); bit_valid = 0; @(negedge clk);
endtask
task automatic send_byte(input logic [W-1:0] v);
for (int i = W - 1; i >= 0; i--) send_bit(v[i]); // MSB-first
endtask
task automatic open_frame(input int skip);
skip_bits = skip[CW-1:0];
frame_start = 1; @(negedge clk); frame_start = 0; @(negedge clk);
got.delete();
endtask
initial begin
repeat (3) @(negedge clk); rst_n = 1; @(negedge clk);
// --- A read: 1 command byte + 3 address bytes + 8 dummy cycles = 40
// bits to discard, then the payload. ---
open_frame(8 + 24 + 8);
chk("not aligned at frame start", aligned, 0);
send_byte(8'hFF); send_byte(8'hFF); send_byte(8'hFF); send_byte(8'hFF);
repeat (8) send_bit(1'b1); // the dummy cycles
chk("aligned after the skip", aligned, 1);
send_byte(8'hA5);
send_byte(8'h3C);
chk("two payload bytes", got.size(), 2);
chk("payload 0", got[0], 8'hA5);
chk("payload 1", got[1], 8'h3C);
// --- The same stream with skip_bits ONE TOO SMALL. Every payload
// byte is displaced by one bit position -- the signature of a
// miscounted latency. ---
open_frame(8 + 24 + 8 - 1);
send_byte(8'hFF); send_byte(8'hFF); send_byte(8'hFF); send_byte(8'hFF);
repeat (8) send_bit(1'b1);
send_byte(8'hA5);
send_byte(8'h3C);
// One extra leading 1 is absorbed, so the first byte is
// {1, A5[7:1]} = 0xD2 and the next inherits A5's last bit.
chk("off-by-one: first byte displaced", got[0], 8'hD2);
chk("off-by-one: not the intended value", (got[0] == 8'hA5) ? 1 : 0, 0);
// --- Zero skip: the payload begins immediately. ---
open_frame(0);
chk("zero skip is aligned at once", aligned, 1);
send_byte(8'h5A);
chk("zero skip byte", got[0], 8'h5A);
// --- A sampling skew that pushed past a bit boundary adds ONE more
// bit to the skip total (Chapter 6.3 §2). ---
open_frame(8 + 24 + 8 + 1);
send_byte(8'hFF); send_byte(8'hFF); send_byte(8'hFF); send_byte(8'hFF);
repeat (8) send_bit(1'b1);
send_bit(1'b1); // the skew bit
send_byte(8'h81);
chk("skew-corrected payload", got[0], 8'h81);
// --- A new frame must restart the accounting even mid-payload. ---
open_frame(8);
send_bit(1'b1); send_bit(1'b1); // partial skip, then re-arm
open_frame(8);
send_byte(8'hFF);
send_byte(8'h7E);
chk("re-armed frame realigns", got[0], 8'h7E);
if (errors == 0)
$display("PASS: the payload begins after exactly skip_bits received bits; a skip one too small displaces every byte; and frame_start restarts the accounting");
else
$display("FAILED with %0d error(s)", errors);
$finish;
end
initial begin #900000; $display("FAIL: watchdog timeout"); $finish; end
endmoduleThe off-by-one case is the most valuable test here, and it is worth reading closely. With skip_bits one too small, one extra leading 1 from the dummy region is absorbed into the payload, so the first byte becomes {1, A5[7:1]} = 0xD2 and every subsequent byte inherits the displacement. The testbench asserts the specific wrong value rather than merely asserting inequality — which means a future change that produces a different wrong answer is also caught, and it documents the failure signature that §9 describes.
// spi_rx_align.v — the same receive alignment in Verilog-2001.
module spi_rx_align #(
parameter WIDTH = 8,
parameter CNT_W = 12
) (
input wire clk,
input wire rst_n,
input wire frame_start,
input wire [CNT_W-1:0] skip_bits,
input wire bit_valid,
input wire bit_in,
output reg [WIDTH-1:0] byte_out,
output reg byte_valid,
output reg aligned
);
function integer clogb2;
input integer value;
integer v;
begin
v = value - 1;
for (clogb2 = 0; v > 0; clogb2 = clogb2 + 1) v = v >> 1;
end
endfunction
localparam BW = (WIDTH > 1) ? clogb2(WIDTH) : 1;
// `aligned` is REGISTERED rather than a continuous compare, so the
// comparison is only evaluated once the counter is defined.
reg [CNT_W-1:0] skipped;
reg [WIDTH-1:0] shreg;
reg [BW-1:0] bit_cnt;
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
skipped <= {CNT_W{1'b0}};
shreg <= {WIDTH{1'b0}};
bit_cnt <= {BW{1'b0}};
byte_out <= {WIDTH{1'b0}};
byte_valid <= 1'b0;
aligned <= 1'b0;
end else begin
byte_valid <= 1'b0;
if (frame_start) begin
// A new frame restarts the whole accounting.
skipped <= {CNT_W{1'b0}};
bit_cnt <= {BW{1'b0}};
shreg <= {WIDTH{1'b0}};
aligned <= (skip_bits == {CNT_W{1'b0}});
end else if (bit_valid) begin
if (skipped < skip_bits) begin
// Command / address / dummy region: counted, discarded.
skipped <= skipped + 1'b1;
if ((skipped + 1'b1) >= skip_bits) aligned <= 1'b1;
end else begin
shreg <= {shreg[WIDTH-2:0], bit_in};
if (bit_cnt == (WIDTH - 1)) begin
bit_cnt <= {BW{1'b0}};
byte_out <= {shreg[WIDTH-2:0], bit_in};
byte_valid <= 1'b1;
end else begin
bit_cnt <= bit_cnt + 1'b1;
end
end
end
end
end
endmodule// spi_rx_align_tb.v — the same checks in Verilog-2001.
`timescale 1ns/1ps
module spi_rx_align_tb;
reg clk = 0, rst_n = 0;
always #5 clk = ~clk;
parameter W = 8, CW = 12;
reg frame_start = 0, bit_valid = 0, bit_in = 0;
reg [CW-1:0] skip_bits = 0;
wire [W-1:0] byte_out;
wire byte_valid, aligned;
spi_rx_align #(.WIDTH(W), .CNT_W(CW)) dut (
.clk(clk), .rst_n(rst_n), .frame_start(frame_start), .skip_bits(skip_bits),
.bit_valid(bit_valid), .bit_in(bit_in), .byte_out(byte_out),
.byte_valid(byte_valid), .aligned(aligned));
integer errors = 0, got_n = 0, i, k;
reg [W-1:0] got [0:15];
task chk;
input [80*8-1:0] what;
input [31:0] g, e;
begin
if (g !== e) begin
$display("FAIL %0s: got 0x%0h exp 0x%0h", what, g, e);
errors = errors + 1;
end
end
endtask
always @(posedge clk) if (rst_n && byte_valid) begin
got[got_n] = byte_out; got_n = got_n + 1;
end
task send_bit;
input b;
begin
bit_in = b; @(negedge clk); bit_valid = 1; @(negedge clk); bit_valid = 0; @(negedge clk);
end
endtask
task send_byte;
input [W-1:0] v;
integer j;
begin
for (j = W - 1; j >= 0; j = j - 1) send_bit(v[j]);
end
endtask
task open_frame;
input integer skip;
begin
skip_bits = skip[CW-1:0];
frame_start = 1; @(negedge clk); frame_start = 0; @(negedge clk);
got_n = 0;
end
endtask
initial begin
repeat (3) @(negedge clk); rst_n = 1; @(negedge clk);
open_frame(8 + 24 + 8);
chk("not aligned at frame start", aligned, 0);
send_byte(8'hFF); send_byte(8'hFF); send_byte(8'hFF); send_byte(8'hFF);
repeat (8) send_bit(1'b1);
chk("aligned after the skip", aligned, 1);
send_byte(8'hA5);
send_byte(8'h3C);
chk("two payload bytes", got_n, 2);
chk("payload 0", got[0], 8'hA5);
chk("payload 1", got[1], 8'h3C);
open_frame(8 + 24 + 8 - 1);
send_byte(8'hFF); send_byte(8'hFF); send_byte(8'hFF); send_byte(8'hFF);
repeat (8) send_bit(1'b1);
send_byte(8'hA5);
send_byte(8'h3C);
chk("off-by-one: first byte displaced", got[0], 8'hD2);
chk("off-by-one: not the intended value", (got[0] == 8'hA5) ? 1 : 0, 0);
open_frame(0);
chk("zero skip is aligned at once", aligned, 1);
send_byte(8'h5A);
chk("zero skip byte", got[0], 8'h5A);
open_frame(8 + 24 + 8 + 1);
send_byte(8'hFF); send_byte(8'hFF); send_byte(8'hFF); send_byte(8'hFF);
repeat (8) send_bit(1'b1);
send_bit(1'b1);
send_byte(8'h81);
chk("skew-corrected payload", got[0], 8'h81);
open_frame(8);
send_bit(1'b1); send_bit(1'b1);
open_frame(8);
send_byte(8'hFF);
send_byte(8'h7E);
chk("re-armed frame realigns", got[0], 8'h7E);
if (errors == 0)
$display("PASS: the payload begins after exactly skip_bits received bits; a skip one too small displaces every byte; and frame_start restarts the accounting");
else
$display("FAILED with %0d error(s)", errors);
$finish;
end
initial begin #900000; $display("FAIL: watchdog timeout"); $finish; end
endmodule-- spi_rx_align.vhd — the same receive alignment in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_rx_align is
generic (
WIDTH : positive := 8;
CNT_W : positive := 12
);
port (
clk : in std_logic;
rst_n : in std_logic;
frame_start : in std_logic; -- pulse
skip_bits : in unsigned(CNT_W - 1 downto 0); -- bits to discard
bit_valid : in std_logic; -- pulse
bit_in : in std_logic;
byte_out : out std_logic_vector(WIDTH - 1 downto 0);
byte_valid : out std_logic; -- pulse
aligned : out std_logic -- level
);
end entity spi_rx_align;
architecture rtl of spi_rx_align is
signal skipped : unsigned(CNT_W - 1 downto 0);
signal shreg : std_logic_vector(WIDTH - 1 downto 0);
signal bit_cnt : integer range 0 to WIDTH - 1;
signal aligned_r : std_logic;
begin
aligned <= aligned_r;
process (clk, rst_n) is
variable next_word : std_logic_vector(WIDTH - 1 downto 0);
begin
if rst_n = '0' then
skipped <= (others => '0');
shreg <= (others => '0');
bit_cnt <= 0;
byte_out <= (others => '0');
byte_valid <= '0';
aligned_r <= '0';
elsif rising_edge(clk) then
byte_valid <= '0';
if frame_start = '1' then
-- A new frame restarts the whole accounting.
skipped <= (others => '0');
bit_cnt <= 0;
shreg <= (others => '0');
if skip_bits = 0 then
aligned_r <= '1';
else
aligned_r <= '0';
end if;
elsif bit_valid = '1' then
if skipped < skip_bits then
-- Command / address / dummy region: counted, discarded.
skipped <= skipped + 1;
if skipped + 1 >= skip_bits then
aligned_r <= '1';
end if;
else
next_word := shreg(WIDTH - 2 downto 0) & bit_in;
shreg <= next_word;
if bit_cnt = WIDTH - 1 then
bit_cnt <= 0;
byte_out <= next_word;
byte_valid <= '1';
else
bit_cnt <= bit_cnt + 1;
end if;
end if;
end if;
end if;
end process;
end architecture rtl;-- spi_rx_align_tb.vhd — the same checks in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_rx_align_tb is
end entity spi_rx_align_tb;
architecture tb of spi_rx_align_tb is
constant W : positive := 8;
constant CW : positive := 12;
signal clk : std_logic := '0';
signal rst_n : std_logic := '0';
signal frame_start : std_logic := '0';
signal bit_valid : std_logic := '0';
signal bit_in : std_logic := '0';
signal skip_bits : unsigned(CW - 1 downto 0) := (others => '0');
signal halt : boolean := false;
signal byte_out : std_logic_vector(W - 1 downto 0);
signal byte_valid : std_logic;
signal aligned : std_logic;
signal errors : natural := 0;
type bytes_t is array (0 to 15) of std_logic_vector(W - 1 downto 0);
signal got : bytes_t := (others => (others => '0'));
signal got_n : natural := 0;
signal got_clr : std_logic := '0';
begin
clk <= not clk after 5 ns when not halt else '0';
dut : entity work.spi_rx_align
generic map (WIDTH => W, CNT_W => CW)
port map (clk => clk, rst_n => rst_n, frame_start => frame_start,
skip_bits => skip_bits, bit_valid => bit_valid, bit_in => bit_in,
byte_out => byte_out, byte_valid => byte_valid, aligned => aligned);
-- One process owns got/got_n; stim only requests a clear.
collect : process (clk) is
begin
if rising_edge(clk) then
if got_clr = '1' then
got_n <= 0;
elsif rst_n = '1' and byte_valid = '1' then
got(got_n) <= byte_out;
got_n <= got_n + 1;
end if;
end if;
end process;
stim : process is
procedure chk_v (what : string; g, e : std_logic_vector) is
begin
if g /= e then
report "FAIL " & what & ": got 0x" & to_hstring(g)
& " exp 0x" & to_hstring(e) severity error;
errors <= errors + 1;
end if;
end procedure;
procedure chk_n (what : string; g, e : natural) is
begin
if g /= e then
report "FAIL " & what & ": got " & integer'image(g)
& " exp " & integer'image(e) severity error;
errors <= errors + 1;
end if;
end procedure;
procedure chk_b (what : string; g, e : std_logic) is
begin
if g /= e then
report "FAIL " & what severity error;
errors <= errors + 1;
end if;
end procedure;
procedure send_bit (b : std_logic) is
begin
bit_in <= b;
wait until falling_edge(clk); bit_valid <= '1';
wait until falling_edge(clk); bit_valid <= '0';
wait until falling_edge(clk);
end procedure;
procedure send_byte (v : std_logic_vector(W - 1 downto 0)) is
begin
for j in W - 1 downto 0 loop
send_bit(v(j));
end loop;
end procedure;
procedure open_frame (skip : natural) is
begin
skip_bits <= to_unsigned(skip, CW);
frame_start <= '1';
got_clr <= '1';
wait until falling_edge(clk);
frame_start <= '0';
got_clr <= '0';
wait until falling_edge(clk);
end procedure;
begin
for i in 0 to 2 loop
wait until falling_edge(clk);
end loop;
rst_n <= '1';
wait until falling_edge(clk);
open_frame(8 + 24 + 8);
chk_b("not aligned at frame start", aligned, '0');
send_byte(x"FF"); send_byte(x"FF"); send_byte(x"FF"); send_byte(x"FF");
for i in 0 to 7 loop send_bit('1'); end loop;
chk_b("aligned after the skip", aligned, '1');
send_byte(x"A5");
send_byte(x"3C");
chk_n("two payload bytes", got_n, 2);
chk_v("payload 0", got(0), x"A5");
chk_v("payload 1", got(1), x"3C");
open_frame(8 + 24 + 8 - 1);
send_byte(x"FF"); send_byte(x"FF"); send_byte(x"FF"); send_byte(x"FF");
for i in 0 to 7 loop send_bit('1'); end loop;
send_byte(x"A5");
send_byte(x"3C");
chk_v("off-by-one: first byte displaced", got(0), x"D2");
open_frame(0);
chk_b("zero skip is aligned at once", aligned, '1');
send_byte(x"5A");
chk_v("zero skip byte", got(0), x"5A");
open_frame(8 + 24 + 8 + 1);
send_byte(x"FF"); send_byte(x"FF"); send_byte(x"FF"); send_byte(x"FF");
for i in 0 to 7 loop send_bit('1'); end loop;
send_bit('1');
send_byte(x"81");
chk_v("skew-corrected payload", got(0), x"81");
open_frame(8);
send_bit('1'); send_bit('1');
open_frame(8);
send_byte(x"FF");
send_byte(x"7E");
chk_v("re-armed frame realigns", got(0), x"7E");
if errors = 0 then
report "PASS: the payload begins after exactly skip_bits received bits; "
& "a skip one too small displaces every byte; and frame_start "
& "restarts the accounting" severity note;
else
report "FAILED with " & integer'image(errors) & " error(s)" severity error;
end if;
halt <= true;
wait;
end process;
watchdog : process is
begin
wait for 900 us;
if not halt then
report "FAIL: watchdog timeout" severity failure;
end if;
wait;
end process;
end architecture tb;Parity
All three implement the same counter: identical ports, asynchronous active-low reset, frame_start restarting the accounting and pre-computing alignment for a zero skip, each captured bit either counted or shifted, MSB-first assembly, a single-cycle byte_valid, and a registered aligned. The VHDL carries skip_bits as unsigned where the Verilog dialects use a plain vector. All three testbenches run the same five scenarios and produce the same values, including the 0xD2 displacement.
6. When the Sampling Point Crosses a Bit Boundary
This is the contribution that surprises people, and it is where Chapters 6.3 and 6.4 meet.
Chapter 6.3 established that the master may delay its sampling instant to land inside the valid window. If the round trip is long enough, the required delay exceeds one full bit period — and at that point the master is sampling, at protocol-bit-position n, a bit the device launched for position n−1.
Nothing is wrong electrically. The captures are clean and the data is correct. It is offset by one bit position, and the correction is a single extra bit in skip_bits.
round trip < 1 bit period → skew = 0
round trip < 2 bit periods → skew = 1
round trip < 3 bit periods → skew = 2Two consequences worth holding.
The skew is a function of frequency. At 10 MHz a 30 ns round trip is well under one 100 ns bit period, so skew is zero. At 50 MHz the same 30 ns round trip exceeds the 20 ns bit period, so skew becomes 1. The same board and the same device need a different skip_bits at a different clock rate — which is exactly the shape of Chapter 4.5's dummy-cycle table, arriving from the board rather than from the silicon.
It is invisible in simulation. A functional model has no round trip, so skew is always zero and the alignment always appears correct. This is the same gap Chapter 6.3 §7 named: the quantity is analogue, and only timing analysis or hardware can establish it.
7. Why a Verification Engineer Cares
The alignment is a checkable relationship between a configuration and an outcome:
// 1. No payload byte may be emitted before the skip is complete. This is
// the property the whole module exists to hold.
a_no_early_payload : assert property (
@(posedge clk) disable iff (!rst_n) byte_valid |-> aligned)
else $error("payload byte emitted before alignment completed");
// 2. Exactly skip_bits captures are discarded -- no more, no fewer.
property p_skip_exact;
@(posedge clk) disable iff (!rst_n)
frame_start |-> (bit_valid[->skip_bits] ##0 !aligned)
or (skip_bits == 0);
endproperty
// 3. A new frame always restarts the accounting, even mid-payload.
// Without this, a frame aborted mid-byte poisons the next one.
a_frame_restarts : assert property (
@(posedge clk) disable iff (!rst_n)
frame_start |=> (skip_bits != 0) -> !aligned)
else $error("frame_start did not restart the alignment");
// 4. byte_valid is a pulse. Held two cycles it would duplicate a byte.
a_valid_is_pulse : assert property (
@(posedge clk) disable iff (!rst_n) byte_valid |=> !byte_valid)
else $error("byte_valid held for more than one cycle");What these prove. That the module discards exactly skip_bits captures and emits well-formed payload words afterwards. What they cannot prove is that skip_bits is the right number — that depends on the device's command structure and dummy count, both datasheet facts, plus the sampling skew, which is a board fact. A perfectly-verified aligner fed a wrong number produces perfectly-aligned garbage.
A reference model is the only way to catch a wrong number, and this is where a scoreboard earns its place:
// The scoreboard must NOT derive its expectation from the DUT's own
// skip_bits, or it will agree with the DUT about a wrong value. It builds
// the expectation from the DEVICE MODEL's structure -- which is the
// independent-predictor discipline: never let the checker inherit the
// implementation's assumptions.
function automatic int expected_skip(spi_device_model dev, int sclk_hz);
int skew;
// The device's own structure -- from its datasheet, not from the DUT.
int base = (dev.cmd_bytes + dev.addr_bytes) * 8 + dev.dummy_for(sclk_hz);
// The board's contribution -- from the timing model, not from the DUT.
skew = board.round_trip_ns * sclk_hz / 1_000_000_000;
return base + skew;
endfunction
task check_read(spi_read_item item);
int exp_skip = expected_skip(dev_model, cfg.sclk_hz);
if (item.cfg_skip != exp_skip)
`uvm_error("SKIP", $sformatf(
"master configured skip=%0d but device+board require %0d",
item.cfg_skip, exp_skip))
// Only then compare the payload itself.
foreach (item.rx_data[i])
if (item.rx_data[i] != dev_model.mem[item.addr + i])
`uvm_error("DATA", "payload mismatch")
endtaskThe first check is the one that matters. Comparing payload alone would report a data mismatch — true but unhelpful — whereas comparing the configured skip against an independently computed one names the cause. This is §28's principle in its most concrete form: a predictor that reads the DUT's configuration reproduces the DUT's bug.
Coverage should target the accounting, not the addresses:
covergroup spi_latency_cg @(posedge transfer_done);
cp_skip : coverpoint cfg.skip_bits {
bins zero = {0}; // no command at all
bins byte_only = {8, 16, 24, 32};
bins with_dummy = {[33:63]};
bins non_byte = {34, 38, 41, 46}; // dummy not a multiple of 8
bins large = {[64:$]};
}
// §6: the skew is a function of clock rate, so the SAME device needs
// a different skip at a different frequency. Covering skip alone
// cannot express that.
cp_skew : coverpoint cfg.sampling_skew {
bins none = {0};
bins one = {1};
bins more = {[2:$]};
}
cp_rate : coverpoint cfg.sclk_hz {
bins slow = {[0:20_000_000]};
bins fast = {[20_000_001:$]};
}
x_skew_rate : cross cp_skew, cp_rate;
endgroupThe non_byte bin is the one that catches byte-oriented drivers: a skip of 41 bits cannot be expressed as a whole number of bytes, and a driver that computes latency in bytes is wrong by up to seven bits without any indication.
8. Why an FPGA or ASIC Engineer Cares
Make skip_bits a register, not a parameter. It depends on the dummy count, which varies with frequency, and on the skew, which varies with frequency and board. A design that fixes it at elaboration works at one clock rate on one board. The cost is a comparator that cannot be folded — small, and the alternative is a part that cannot be re-clocked.
Size the counter for the worst case. The largest plausible skip is a 1-byte command plus a 4-byte address plus 32 dummy cycles plus a few bits of skew — about 72 bits, so seven counter bits. Sizing to a nominal case and then meeting a device with a longer address is an overflow that silently mis-aligns.
The aligner is where a FIFO belongs. byte_valid is a single-cycle strobe and the payload arrives continuously at SCLK rate. If the consumer is a bus interface that can stall, the byte must be buffered — and the depth needed follows from how long the consumer can stall relative to a byte time, the same arithmetic as Chapter 6.1 §9's prefetch budget seen from the other end.
Latency is not throughput, and the system cares about both. A design reading a sensor on an interrupt cares about the 1160 ns latency of §4; a design streaming a display buffer cares only about the 6.10 MB/s. Knowing which number the application is sensitive to determines whether to optimise the request overhead or the burst length — and they are different optimisations.
9. Failure Signature — Every Byte Shifted by Exactly One Byte Position
Symptom. A read returns data that is recognisably correct but displaced: the byte expected at offset 0 appears at offset 1, and the first byte is something else entirely — often 0xFF or 0x00. Reproducible. Writes work. The displacement is exactly one byte, not one bit.
What a whole-byte displacement tells you. It is the most informative feature of this symptom. A bit-level miscount — a wrong dummy count, a missed skew bit — displaces by one to seven bits and produces bytes that look like nothing at all, because each output byte is a blend of two source bytes. A displacement of exactly eight bits keeps every byte intact and merely moves it, which points at a whole byte-sized unit being miscounted.
Plausible mechanisms.
- The address length is wrong by one byte — the master sends 3 where the device expects 4, or the reverse (Chapter 4.4 §10).
- The dummy count is wrong by exactly 8 cycles, which happens when a driver expresses dummy in bytes and the device specifies it in cycles.
- The master is counting the command byte twice, or not at all, in its skip total.
- The device returns a leading status or dummy byte that the datasheet documents and the driver did not account for.
The discriminating observation. Look at what landed in the first position. If it is 0xFF or 0x00, it is filler from the dummy region being treated as payload — so the skip is too small, and by exactly one byte. If instead the first expected byte is missing entirely and the sequence starts at the second, the skip is too large. That single observation gives both the direction and the magnitude of the correction.
Then reconcile the arithmetic explicitly: write out cmd × 8 + addr × 8 + dummy + skew from the datasheet and compare it against what the driver configured. A discrepancy of 8 confirms it and identifies which term is wrong.
Why the investigation goes wrong. Because displaced-but-intact bytes read as a data problem, so people compare payload values and conclude the device returned the wrong data. The bytes are right; their position is wrong. Recognising that a whole-byte shift is an accounting error rather than a data error is what turns this into a two-minute fix — and it is the same reasoning as Chapter 5.3 §9's displaced-but-intact rule, applied to time rather than to address.
10. Common Misconceptions
11. Reason It Through
Work this before reading the answer.
A design reads 4-byte samples from an ADC at 20 MHz using a 1-byte command, a 2-byte address and 4 dummy cycles. It works. The team moves to a faster board revision and raises SCLK to 60 MHz to triple the sample rate. Reads now return data displaced by one bit. The ADC's datasheet specifies 4 dummy cycles below 25 MHz and 10 above it.
What is the correct
skip_bitsat 60 MHz, and is the dummy table the whole story?
Compute the original. At 20 MHz:
skip = 1×8 + 2×8 + 4 + skew = 8 + 16 + 4 + 0 = 28 bitsSkew is zero at 20 MHz because a 50 ns bit period comfortably exceeds any plausible round trip.
Apply the datasheet's table. At 60 MHz the device requires 10 dummy cycles, not 4:
skip = 1×8 + 2×8 + 10 + skew = 8 + 16 + 10 + skew = 34 + skewNow the question the problem is really asking. The observed displacement is one bit. If only the dummy count had been missed, the error would be 10 − 4 = 6 bits — a six-bit displacement, which would make every byte a blend of two and look like noise, not a clean one-bit shift.
So the team has evidently already applied the dummy table, and a one-bit residue remains. That is the signature of the sampling skew, and it is the term the datasheet cannot give them because it depends on their board.
Why does skew appear at 60 MHz and not at 20? The bit period falls from 50 ns to 16.7 ns. A round trip that was a fraction of a bit period is now comparable to or larger than one:
20 MHz: T_sclk = 50.0 ns round trip ~20 ns → skew = 0
60 MHz: T_sclk = 16.7 ns round trip ~20 ns → skew = 1So the correct value is skip = 8 + 16 + 10 + 1 = 35 bits.
Is the dummy table the whole story? No — and that is the lesson. The datasheet describes the device: how many cycles it needs to fetch. It cannot describe the link: how long a bit takes to travel out and back on this particular board. Two boards with the same ADC at the same clock can legitimately need different skip totals.
What would confirm it? Chapter 6.3's approach: compute the round trip from the device's t_v and the board's flight time, compare against the 16.7 ns bit period, and check that the resulting skew is 1. If a receive-sample delay is also being applied, its magnitude relative to a bit period gives the same answer directly — a delay exceeding one bit period is a skew of one.
The general lesson. Read latency has a device component and a link component. The datasheet supplies the first and is silent on the second, so a design that changes clock rate must recompute both. A residual displacement after applying the datasheet's table is the link component announcing itself.
12. Understanding Check
13. Summary
Read latency is a sum of five contributions: command bits, address bits, dummy cycles, sampling skew, and frame overhead. The first four are counted on SCLK; the last is real time that clocks no bits.
The same arithmetic serves two purposes. As a duration it says what a read costs. As a bit count it is skip_bits — the number of received bits the master discards before the payload — and the RTL takes it as a single input, because nothing on the bus marks where the response begins.
skip_bits = cmd_bytes×8 + addr_bytes×8 + dummy_cycles + skew_bitsIt cannot be computed in bytes. Dummy is specified in cycles and is often not a multiple of eight; skew is a single bit.
It is a function of frequency. The dummy count rises with clock rate because the device's fetch time is fixed in nanoseconds, and the skew rises because the round trip is fixed while the bit period shrinks. The same device on the same board needs different totals at different rates.
As a cost, a 40-cycle request makes a single-byte read about 14% efficient and a 256-byte burst about 98%. Raising SCLK does not change the ratio; amortising the request over a larger burst does, often several-fold.
In RTL the module is a counter that either discards a bit or shifts it, never both, restarted unconditionally by frame_start — the master-side mirror of CS being the only resynchronising event. aligned is registered rather than continuously compared, so the comparison is never evaluated against an undefined counter.
For verification, the module can be provably correct while fed a wrong number, so the check that matters is an independently computed expectation built from the device model and the board model — never from the DUT's own configuration.
And when a read comes back displaced by exactly one byte, that is an accounting error, not a data error: the byte in the first position tells you the direction, and reconciling the four terms against the datasheet tells you which one is wrong.
14. What Comes Next
Chapter 6.5 — Read Waveform Analysis closes the module by inverting everything in it. Given an unlabelled capture of a read — no datasheet, no configuration — recover the mode, locate the frame, find the turnaround by watching where MISO stops being high impedance, measure the dummy phase directly from the capture, and reconstruct the transaction. It is the practical pay-off of this chapter's arithmetic: once you can count the cycles in a capture, you can read a device's latency off an oscilloscope without its datasheet.
Continue learning
Related tutorials
- Related topic
Dummy Phases and Read Latency
Why a device needs turnaround before it can answer, why dummy is counted in clock cycles rather than bytes, how its length grows with frequency, and the one-byte data offset a mismatch produces.
- Related topic
Back-to-Back Transactions and Inter-Frame Gap
How soon chip select may fall again after it rises: the minimum deselect time, why a device needs it, what the gap costs in throughput, and the hardware enforcement that keeps software from violating it.
- Related topic
Raw vs Effective Payload Throughput
The three throughput numbers a link has — raw bit rate, protocol rate and effective payload rate — why they differ by four times on an ordinary flash read, and the performance counter that measures the gap instead of estimating it.
- Related topic
SDR, DDR Transfers, and Throughput
Why data on both edges is independent of lane width, why the byte assembly is identical in both modes, why the dummy phase does not halve so DDR gives under two times, and why DDR parts specify more dummy cycles than their SDR modes.
