SPI · Module 9
Protocol Overhead
Overhead split into the part SCLK pays for and the part it does not: why the non-clocked share grows from 5% to half the transaction as the clock rises, and the frame timer that partitions a transfer into gap, lead, active and lag.
Chapter 9.1 established that effective throughput is a fraction of the raw rate. This chapter takes the fraction apart.
A transaction takes 1340 ns and delivers 16 bits of payload. Where did the other time go, line by line — and which parts would change if the payload were bigger?
The second half is what makes this more than bookkeeping: overhead divides into two kinds with completely different consequences, and telling them apart decides which optimisation is worth doing.
1. Two Kinds of Overhead
Clocked overhead — command bits, address bits, dummy cycles. These occupy SCLK edges, so they scale inversely with clock rate: double the clock and they take half the time.
Non-clocked overhead — the CS lead, the CS lag, and the inter-frame gap. No bit moves during these, and critically they do not scale with the clock at all. They are fixed in nanoseconds because they are set by device requirements: the lead and lag by the device's internal setup, the gap by its recovery time.
That distinction is the chapter's centre of gravity:
clocked overhead shrinks when you raise the clock
non-clocked overhead does not2. Where a Frame's Time Goes
Two clock pulses inside a much longer frame
10 cyclesThe moving lane marks the interval in which bits are actually being carried. Everything outside it, inside the frame, is non-clocked overhead — and in this figure it is the majority.
3. The Per-Access Budget
For a read with a 1-byte command, 3-byte address and 8 dummy cycles, at 50 MHz with a 100 ns lead, 20 ns lag and 100 ns gap:
CLOCKED, scales with 1/f_sclk
command 8 cycles × 20 ns = 160 ns
address 24 cycles × 20 ns = 480 ns
dummy 8 cycles × 20 ns = 160 ns
─────────────────────────────────────────
subtotal 40 cycles = 800 ns
NON-CLOCKED, fixed in nanoseconds
CS lead = 100 ns
CS lag = 20 ns
inter-frame gap = 100 ns
─────────────────────────────────────────
subtotal = 220 ns
TOTAL OVERHEAD PER ACCESS = 1020 nsAgainst a 1-byte payload (160 ns of clocking) that is 1020 ns of overhead for 160 ns of data — 86% overhead.
4. What Happens When You Raise the Clock
This is where the two kinds separate visibly. Take the same transaction at three clock rates:
f_sclk clocked ovh non-clocked total non-clocked share
10 MHz 4000 ns 220 ns 4220 ns 5 %
50 MHz 800 ns 220 ns 1020 ns 22 %
100 MHz 400 ns 220 ns 620 ns 35 %
200 MHz 200 ns 220 ns 420 ns 52 %At 200 MHz more than half the overhead is time in which nothing is clocked — and no further increase in clock rate can touch it.
Two conclusions follow.
There is a knee beyond which clock rate stops helping. It arrives when clocked overhead falls to roughly the size of the non-clocked overhead, and for typical devices that is somewhere between 100 and 200 MHz — which is, not coincidentally, around where SPI devices stop getting faster.
Reducing the number of accesses attacks both kinds at once. Merging two transactions into one removes a full command, address, dummy phase, lead, lag and gap. Raising the clock removes a fraction of one of those categories. This is the quantitative form of the advice Chapters 7.1 and 7.2 gave qualitatively.
5. Measuring the Part a Bit Counter Cannot See
Chapter 9.1's counter counts bits. The non-clocked overhead carries no bits, so it is invisible to it — the counter can tell you 40 of 48 bits were overhead, and nothing at all about the 220 ns in which no bit moved.
That is what the frame timer in §6 is for. It splits a frame into lead, active and lag, and separately measures the gap before it, which is the four-way breakdown of §3 made measurable.
Its defining property is that the buckets partition the frame: every in-frame cycle lands in exactly one of lead, active or lag. Without that, a breakdown is a collection of plausible numbers that need not sum to anything.
6. Building the Frame Timer — Three HDLs
The circuit
Circuit. Four accumulators, a published-output register set, and a small state bit.
State. The four running counts, a registered copy of the frame level, and whether any bit has occurred in this frame.
Datapath. None.
Control. Three cases while in-frame — the entry cycle, before the first bit, and after it — plus the gap accumulation while out of frame. The subtlety is the lag bucket: a cycle with no bit is provisionally lag, and if another bit follows it is folded back into active. Only the tail that reaches the end of the frame is genuinely lag.
Clock and reset. System clock; asynchronous active-low reset clearing everything.
Enables. gap_acc is deliberately not cleared when a frame opens — it has been accumulating since the previous frame closed and is this frame's preceding gap. Clearing it there is the natural-looking mistake that loses the measurement.
Timing. The four outputs are published together on the CS rising edge with a single frame_valid strobe, so a consumer reads a coherent set rather than a skewed snapshot (Chapter 9.1 §8).
Synthesis. Eight T_W counters — four accumulating and four published — plus a little control. The doubling is what allows a completed frame to be read while the next is in progress.
Limitations. It measures time, not bits; the two are complementary and a real controller carries both.
// spi_frame_timer.sv — where a frame's time actually goes.
//
// Chapter 9.1 counted BITS. This counts TIME, and the two disagree in a way
// that matters: a frame contains intervals during which chip select is
// asserted and no bit moves at all -- the CS lead and lag (Chapter 2.5) --
// plus the gap between frames (Chapter 7.3). None of those carry a bit, and
// all of them are paid per access.
//
// Splitting a frame into lead / active / lag, plus the preceding gap, turns
// "framing overhead" from a hand-waved term into four measured numbers.
module spi_frame_timer #(
parameter int T_W = 16
) (
input logic clk,
input logic rst_n,
input logic cs_active, // level: a frame is open
input logic bit_stb, // one pulse per SPI bit
output logic [T_W-1:0] gap_cycles, // CS high, before this frame
output logic [T_W-1:0] lead_cycles, // CS low, before the first bit
output logic [T_W-1:0] active_cycles, // first bit to last bit
output logic [T_W-1:0] lag_cycles, // last bit to CS high
output logic frame_valid // pulse: the four outputs are a frame
);
logic cs_q;
logic seen_bit; // has any bit occurred this frame?
// Running accumulators. Kept separate from the outputs so a consumer can
// read a completed frame's breakdown while the next one is in progress.
logic [T_W-1:0] gap_acc, lead_acc, active_acc, lag_acc;
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
cs_q <= 1'b0;
seen_bit <= 1'b0;
gap_acc <= '0;
lead_acc <= '0;
active_acc <= '0;
lag_acc <= '0;
gap_cycles <= '0;
lead_cycles <= '0;
active_cycles <= '0;
lag_cycles <= '0;
frame_valid <= 1'b0;
end else begin
cs_q <= cs_active;
frame_valid <= 1'b0;
if (!cs_active) begin
if (cs_q) begin
// CS just rose: publish the frame that has ended.
gap_cycles <= gap_acc;
lead_cycles <= lead_acc;
active_cycles <= active_acc;
lag_cycles <= lag_acc;
frame_valid <= 1'b1;
gap_acc <= '0; // the next gap starts now
end else begin
gap_acc <= gap_acc + 1'b1;
end
end else begin
if (!cs_q) begin
// CS just fell: a new frame. Reset the in-frame buckets
// but NOT gap_acc, which has been accumulating and is
// this frame's preceding gap.
//
// This entry cycle is itself in-frame and must be counted,
// or the buckets do not partition the frame.
seen_bit <= bit_stb;
lead_acc <= bit_stb ? '0 : {{(T_W-1){1'b0}}, 1'b1};
active_acc <= bit_stb ? {{(T_W-1){1'b0}}, 1'b1} : '0;
lag_acc <= '0;
end else if (!seen_bit) begin
// Before the first bit: lead time. The cycle carrying the
// FIRST bit belongs to active, not lead.
if (bit_stb) begin
seen_bit <= 1'b1;
active_acc <= active_acc + 1'b1;
end else begin
lead_acc <= lead_acc + 1'b1;
end
end else begin
// After the first bit. A cycle is "active" if a bit has
// occurred recently and "lag" otherwise -- so lag is the
// tail that turns out to have had no further bits.
if (bit_stb) begin
// Any accumulated lag was actually part of the active
// interval: fold it in and keep going.
active_acc <= active_acc + lag_acc + 1'b1;
lag_acc <= '0;
end else begin
lag_acc <= lag_acc + 1'b1;
end
end
end
end
end
endmoduleThe fold-back in the final else branch is the design's one subtle line. A cycle without a bit cannot be classified when it occurs — it is lag only if the frame ends before another bit arrives. Accumulating it provisionally and adding it to active on the next bit resolves that without needing to look ahead.
// spi_frame_timer_tb.sv — a frame with known lead, active and lag, and the
// check that the four buckets account for every cycle.
`timescale 1ns/1ps
module spi_frame_timer_tb;
logic clk = 0, rst_n = 0;
always #5 clk = ~clk;
localparam int TW = 16;
logic cs_active = 0, bit_stb = 0;
logic [TW-1:0] gap_cycles, lead_cycles, active_cycles, lag_cycles;
logic frame_valid;
spi_frame_timer #(.T_W(TW)) dut (
.clk, .rst_n, .cs_active, .bit_stb,
.gap_cycles, .lead_cycles, .active_cycles, .lag_cycles, .frame_valid);
int errors = 0, frames = 0;
int g_gap, g_lead, g_active, g_lag;
int in_frame_cycles = 0, measured_in_frame = 0;
// Independently count the cycles CS was asserted, so the DUT's buckets
// can be checked against a number it did not produce.
always @(posedge clk) if (rst_n) begin
if (cs_active) in_frame_cycles++;
if (frame_valid) begin
measured_in_frame <= in_frame_cycles;
in_frame_cycles <= 0;
end
end
task automatic chk(input string what, input int g, input int e);
if (g !== e) begin $display("FAIL %s: got %0d exp %0d", what, g, e); errors++; end
endtask
// Latch a published frame.
always @(posedge clk) if (rst_n && frame_valid) begin
frames++;
g_gap <= gap_cycles;
g_lead <= lead_cycles;
g_active <= active_cycles;
g_lag <= lag_cycles;
end
// One bit occupies exactly two system cycles: one with bit_stb high.
task automatic send_bit();
@(negedge clk); bit_stb = 1;
@(negedge clk); bit_stb = 0;
endtask
task automatic frame(input int lead, input int n_bits, input int lag);
cs_active = 1;
repeat (lead) @(negedge clk);
for (int i = 0; i < n_bits; i++) send_bit();
repeat (lag) @(negedge clk);
cs_active = 0;
@(negedge clk);
endtask
task automatic idle(input int n); repeat (n) @(negedge clk); endtask
initial begin
repeat (3) @(negedge clk); rst_n = 1; @(negedge clk);
// --- a frame with 6 lead cycles, 8 bits, 5 lag cycles, after a gap ---
idle(12);
frame(6, 8, 5);
idle(4);
chk("one frame published", frames, 1);
$display(" gap=%0d lead=%0d active=%0d lag=%0d",
g_gap, g_lead, g_active, g_lag);
// The cycle on which CS falls is in-frame and carries no bit, so it
// is lead -- hence requested + 1.
chk("lead", g_lead, 6 + 1);
// The lag is the tail after the last bit strobe, before CS rises.
chk("lag", g_lag, 5);
// Active is whatever remains, which is the partition property stated
// as an equality rather than a guessed constant.
chk("active = frame - lead - lag", g_active, measured_in_frame - g_lead - g_lag);
chk("buckets partition the frame", g_lead + g_active + g_lag, measured_in_frame);
// The gap is everything CS was high before this frame, including the
// reset settling -- so check it is at least the idle we inserted.
if (g_gap < 12) begin
$display("FAIL: gap %0d should be at least 12", g_gap); errors++;
end
// --- a frame with NO lead and NO lag: all time is active ---
idle(8);
frame(0, 4, 0);
idle(4);
chk("two frames", frames, 2);
// With zero requested lead, only the CS-fall cycle is lead.
chk("minimal lead", g_lead, 1);
chk("no lag", g_lag, 0);
chk("active = frame - lead - lag", g_active, measured_in_frame - g_lead - g_lag);
chk("buckets partition (no lead/lag)", g_lead + g_active + g_lag, measured_in_frame);
if (g_active <= g_lead) begin
$display("FAIL: a 4-bit frame with no lead should be mostly active"); errors++;
end
// --- a frame with no bits at all: everything is lead ---
idle(8);
frame(7, 0, 0);
idle(4);
chk("three frames", frames, 3);
// With no bits at all, every in-frame cycle is lead. Stating it as
// an identity is stronger than a constant and says what it means.
chk("empty: all in-frame time is lead", g_lead, measured_in_frame);
chk("empty: active", g_active, 0);
chk("empty: lag", g_lag, 0);
chk("buckets partition (empty)", g_lead + g_active + g_lag, measured_in_frame);
// --- overhead ratio: a short frame is mostly framing ---
idle(8);
frame(6, 2, 5);
idle(4);
$display(" short frame: lead=%0d active=%0d lag=%0d -> framing is %0d of %0d cycles",
g_lead, g_active, g_lag, g_lead + g_lag, g_lead + g_active + g_lag);
if (g_lead + g_lag <= g_active) begin
$display("FAIL: for a 2-bit frame, framing should exceed active time");
errors++;
end
if (errors == 0)
$display("PASS: lead, active and lag are measured separately, a frame with no bits is all lead, and on a short frame the framing overhead exceeds the clocked time");
else
$display("FAILED with %0d error(s)", errors);
$finish;
end
initial begin #900000; $display("FAIL: watchdog timeout"); $finish; end
endmoduleThe testbench counts in-frame cycles independently and asserts that lead + active + lag equals that count. That is stronger than checking the three values individually: it verifies the breakdown is a genuine partition rather than three plausible numbers, and it holds for every frame shape without the test needing to know the expected split.
It also reports a result worth quoting: for a 2-bit frame, framing consumes 12 of 15 cycles.
// spi_frame_timer.v — the same frame-time breakdown in Verilog-2001.
module spi_frame_timer #(
parameter T_W = 16
) (
input wire clk,
input wire rst_n,
input wire cs_active,
input wire bit_stb,
output reg [T_W-1:0] gap_cycles,
output reg [T_W-1:0] lead_cycles,
output reg [T_W-1:0] active_cycles,
output reg [T_W-1:0] lag_cycles,
output reg frame_valid
);
reg cs_q, seen_bit;
reg [T_W-1:0] gap_acc, lead_acc, active_acc, lag_acc;
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
cs_q <= 1'b0;
seen_bit <= 1'b0;
gap_acc <= {T_W{1'b0}};
lead_acc <= {T_W{1'b0}};
active_acc <= {T_W{1'b0}};
lag_acc <= {T_W{1'b0}};
gap_cycles <= {T_W{1'b0}};
lead_cycles <= {T_W{1'b0}};
active_cycles <= {T_W{1'b0}};
lag_cycles <= {T_W{1'b0}};
frame_valid <= 1'b0;
end else begin
cs_q <= cs_active;
frame_valid <= 1'b0;
if (!cs_active) begin
if (cs_q) begin
// CS just rose: publish the frame that has ended.
gap_cycles <= gap_acc;
lead_cycles <= lead_acc;
active_cycles <= active_acc;
lag_cycles <= lag_acc;
frame_valid <= 1'b1;
gap_acc <= {T_W{1'b0}};
end else begin
gap_acc <= gap_acc + 1'b1;
end
end else begin
if (!cs_q) begin
// The CS-fall cycle is in-frame and must be counted, or
// the buckets do not partition the frame.
seen_bit <= bit_stb;
lead_acc <= bit_stb ? {T_W{1'b0}} : {{(T_W-1){1'b0}}, 1'b1};
active_acc <= bit_stb ? {{(T_W-1){1'b0}}, 1'b1} : {T_W{1'b0}};
lag_acc <= {T_W{1'b0}};
end else if (!seen_bit) begin
// Lead time. The cycle carrying the FIRST bit is active.
if (bit_stb) begin
seen_bit <= 1'b1;
active_acc <= active_acc + 1'b1;
end else begin
lead_acc <= lead_acc + 1'b1;
end
end else begin
// Accumulated lag turns out to be active if another bit
// follows; it is only lag once the frame ends.
if (bit_stb) begin
active_acc <= active_acc + lag_acc + 1'b1;
lag_acc <= {T_W{1'b0}};
end else begin
lag_acc <= lag_acc + 1'b1;
end
end
end
end
end
endmodule// spi_frame_timer_tb.v — the same checks in Verilog-2001.
`timescale 1ns/1ps
module spi_frame_timer_tb;
reg clk = 0, rst_n = 0;
always #5 clk = ~clk;
parameter TW = 16;
reg cs_active = 0, bit_stb = 0;
wire [TW-1:0] gap_cycles, lead_cycles, active_cycles, lag_cycles;
wire frame_valid;
spi_frame_timer #(.T_W(TW)) dut (
.clk(clk), .rst_n(rst_n), .cs_active(cs_active), .bit_stb(bit_stb),
.gap_cycles(gap_cycles), .lead_cycles(lead_cycles),
.active_cycles(active_cycles), .lag_cycles(lag_cycles),
.frame_valid(frame_valid));
integer errors = 0, frames = 0, i;
integer g_gap = 0, g_lead = 0, g_active = 0, g_lag = 0;
integer in_frame_cycles = 0, measured_in_frame = 0;
task chk;
input [80*8-1:0] what;
input [31:0] g, e;
begin
if (g !== e) begin
$display("FAIL %0s: got %0d exp %0d", what, g, e);
errors = errors + 1;
end
end
endtask
always @(posedge clk) if (rst_n) begin
if (cs_active) in_frame_cycles = in_frame_cycles + 1;
if (frame_valid) begin
frames = frames + 1;
g_gap = gap_cycles;
g_lead = lead_cycles;
g_active = active_cycles;
g_lag = lag_cycles;
measured_in_frame = in_frame_cycles;
in_frame_cycles = 0;
end
end
task send_bit;
begin
@(negedge clk); bit_stb = 1;
@(negedge clk); bit_stb = 0;
end
endtask
task frame;
input integer lead, n_bits, lag;
integer k;
begin
cs_active = 1;
repeat (lead) @(negedge clk);
for (k = 0; k < n_bits; k = k + 1) send_bit;
repeat (lag) @(negedge clk);
cs_active = 0;
@(negedge clk);
end
endtask
task idle; input integer n; begin repeat (n) @(negedge clk); end endtask
initial begin
repeat (3) @(negedge clk); rst_n = 1; @(negedge clk);
idle(12);
frame(6, 8, 5);
idle(4);
chk("one frame published", frames, 1);
$display(" gap=%0d lead=%0d active=%0d lag=%0d", g_gap, g_lead, g_active, g_lag);
chk("lead", g_lead, 6 + 1);
chk("lag", g_lag, 5);
chk("active = frame - lead - lag", g_active, measured_in_frame - g_lead - g_lag);
chk("buckets partition the frame", g_lead + g_active + g_lag, measured_in_frame);
if (g_gap < 12) begin
$display("FAIL: gap %0d should be at least 12", g_gap); errors = errors + 1;
end
idle(8);
frame(0, 4, 0);
idle(4);
chk("two frames", frames, 2);
chk("minimal lead", g_lead, 1);
chk("no lag", g_lag, 0);
chk("active = frame - lead - lag", g_active, measured_in_frame - g_lead - g_lag);
chk("buckets partition (no lead/lag)", g_lead + g_active + g_lag, measured_in_frame);
if (g_active <= g_lead) begin
$display("FAIL: a 4-bit frame with no lead should be mostly active");
errors = errors + 1;
end
idle(8);
frame(7, 0, 0);
idle(4);
chk("three frames", frames, 3);
chk("empty: all in-frame time is lead", g_lead, measured_in_frame);
chk("empty: active", g_active, 0);
chk("empty: lag", g_lag, 0);
chk("buckets partition (empty)", g_lead + g_active + g_lag, measured_in_frame);
idle(8);
frame(6, 2, 5);
idle(4);
$display(" short frame: lead=%0d active=%0d lag=%0d -> framing is %0d of %0d cycles",
g_lead, g_active, g_lag, g_lead + g_lag, g_lead + g_active + g_lag);
if (g_lead + g_lag <= g_active) begin
$display("FAIL: for a 2-bit frame, framing should exceed active time");
errors = errors + 1;
end
if (errors == 0)
$display("PASS: lead, active and lag are measured separately, a frame with no bits is all lead, and on a short frame the framing overhead exceeds the clocked time");
else
$display("FAILED with %0d error(s)", errors);
$finish;
end
initial begin #900000; $display("FAIL: watchdog timeout"); $finish; end
endmodule-- spi_frame_timer.vhd — the same frame-time breakdown in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_frame_timer is
generic (
T_W : positive := 16
);
port (
clk : in std_logic;
rst_n : in std_logic;
cs_active : in std_logic; -- level
bit_stb : in std_logic; -- one pulse per SPI bit
gap_cycles : out unsigned(T_W - 1 downto 0); -- CS high, before this frame
lead_cycles : out unsigned(T_W - 1 downto 0); -- CS low, before the first bit
active_cycles : out unsigned(T_W - 1 downto 0); -- first bit to last bit
lag_cycles : out unsigned(T_W - 1 downto 0); -- last bit to CS high
frame_valid : out std_logic -- pulse
);
end entity spi_frame_timer;
architecture rtl of spi_frame_timer is
signal cs_q, seen_bit : std_logic;
signal gap_acc, lead_acc, active_acc, lag_acc : unsigned(T_W - 1 downto 0);
begin
process (clk, rst_n) is
begin
if rst_n = '0' then
cs_q <= '0';
seen_bit <= '0';
gap_acc <= (others => '0');
lead_acc <= (others => '0');
active_acc <= (others => '0');
lag_acc <= (others => '0');
gap_cycles <= (others => '0');
lead_cycles <= (others => '0');
active_cycles <= (others => '0');
lag_cycles <= (others => '0');
frame_valid <= '0';
elsif rising_edge(clk) then
cs_q <= cs_active;
frame_valid <= '0';
if cs_active = '0' then
if cs_q = '1' then
-- CS just rose: publish the frame that has ended.
gap_cycles <= gap_acc;
lead_cycles <= lead_acc;
active_cycles <= active_acc;
lag_cycles <= lag_acc;
frame_valid <= '1';
gap_acc <= (others => '0');
else
gap_acc <= gap_acc + 1;
end if;
else
if cs_q = '0' then
-- The CS-fall cycle is in-frame and must be counted, or
-- the buckets do not partition the frame.
seen_bit <= bit_stb;
lag_acc <= (others => '0');
if bit_stb = '1' then
lead_acc <= (others => '0');
active_acc <= to_unsigned(1, T_W);
else
lead_acc <= to_unsigned(1, T_W);
active_acc <= (others => '0');
end if;
elsif seen_bit = '0' then
-- Lead time. The cycle carrying the FIRST bit is active.
if bit_stb = '1' then
seen_bit <= '1';
active_acc <= active_acc + 1;
else
lead_acc <= lead_acc + 1;
end if;
else
-- Accumulated lag turns out to be active if another bit
-- follows; it is only lag once the frame ends.
if bit_stb = '1' then
active_acc <= active_acc + lag_acc + 1;
lag_acc <= (others => '0');
else
lag_acc <= lag_acc + 1;
end if;
end if;
end if;
end if;
end process;
end architecture rtl;-- spi_frame_timer_tb.vhd — the same checks in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_frame_timer_tb is
end entity spi_frame_timer_tb;
architecture tb of spi_frame_timer_tb is
constant TW : positive := 16;
signal clk : std_logic := '0';
signal rst_n : std_logic := '0';
signal cs_active : std_logic := '0';
signal bit_stb : std_logic := '0';
signal halt : boolean := false;
signal gap_cycles : unsigned(TW - 1 downto 0);
signal lead_cycles : unsigned(TW - 1 downto 0);
signal active_cycles : unsigned(TW - 1 downto 0);
signal lag_cycles : unsigned(TW - 1 downto 0);
signal frame_valid : std_logic;
signal errors : natural := 0;
signal frames : natural := 0;
signal g_gap, g_lead, g_active, g_lag : natural := 0;
signal measured_in_frame : natural := 0;
begin
clk <= not clk after 5 ns when not halt else '0';
dut : entity work.spi_frame_timer
generic map (T_W => TW)
port map (clk => clk, rst_n => rst_n, cs_active => cs_active, bit_stb => bit_stb,
gap_cycles => gap_cycles, lead_cycles => lead_cycles,
active_cycles => active_cycles, lag_cycles => lag_cycles,
frame_valid => frame_valid);
-- One process owns the observation signals.
observe : process (clk) is
variable in_frame : natural := 0;
begin
if rising_edge(clk) and rst_n = '1' then
if cs_active = '1' then
in_frame := in_frame + 1;
end if;
if frame_valid = '1' then
frames <= frames + 1;
g_gap <= to_integer(gap_cycles);
g_lead <= to_integer(lead_cycles);
g_active <= to_integer(active_cycles);
g_lag <= to_integer(lag_cycles);
measured_in_frame <= in_frame;
in_frame := 0;
end if;
end if;
end process;
stim : process is
procedure chk_n (what : string; g, e : natural) is
begin
if g /= e then
report "FAIL " & what & ": got " & integer'image(g)
& " exp " & integer'image(e) severity error;
errors <= errors + 1;
end if;
end procedure;
procedure send_bit is
begin
wait until falling_edge(clk); bit_stb <= '1';
wait until falling_edge(clk); bit_stb <= '0';
end procedure;
procedure frame (lead, n_bits, lag : natural) is
begin
cs_active <= '1';
for i in 1 to lead loop wait until falling_edge(clk); end loop;
for i in 1 to n_bits loop send_bit; end loop;
for i in 1 to lag loop wait until falling_edge(clk); end loop;
cs_active <= '0';
wait until falling_edge(clk);
end procedure;
procedure idle (n : natural) is
begin
for i in 1 to n loop wait until falling_edge(clk); end loop;
end procedure;
begin
for i in 0 to 2 loop wait until falling_edge(clk); end loop;
rst_n <= '1';
wait until falling_edge(clk);
idle(12);
frame(6, 8, 5);
idle(4);
chk_n("one frame published", frames, 1);
report " gap=" & integer'image(g_gap) & " lead=" & integer'image(g_lead)
& " active=" & integer'image(g_active) & " lag=" & integer'image(g_lag)
severity note;
chk_n("lead", g_lead, 6 + 1);
chk_n("lag", g_lag, 5);
chk_n("active = frame - lead - lag", g_active, measured_in_frame - g_lead - g_lag);
chk_n("buckets partition the frame", g_lead + g_active + g_lag, measured_in_frame);
if g_gap < 12 then
report "FAIL: gap too small" severity error;
errors <= errors + 1;
end if;
idle(8);
frame(0, 4, 0);
idle(4);
chk_n("two frames", frames, 2);
chk_n("minimal lead", g_lead, 1);
chk_n("no lag", g_lag, 0);
chk_n("active = frame - lead - lag", g_active, measured_in_frame - g_lead - g_lag);
chk_n("buckets partition (no lead/lag)", g_lead + g_active + g_lag, measured_in_frame);
if g_active <= g_lead then
report "FAIL: a 4-bit frame with no lead should be mostly active" severity error;
errors <= errors + 1;
end if;
idle(8);
frame(7, 0, 0);
idle(4);
chk_n("three frames", frames, 3);
chk_n("empty: all in-frame time is lead", g_lead, measured_in_frame);
chk_n("empty: active", g_active, 0);
chk_n("empty: lag", g_lag, 0);
idle(8);
frame(6, 2, 5);
idle(4);
report " short frame: lead=" & integer'image(g_lead)
& " active=" & integer'image(g_active)
& " lag=" & integer'image(g_lag)
& " -> framing is " & integer'image(g_lead + g_lag)
& " of " & integer'image(g_lead + g_active + g_lag) & " cycles"
severity note;
if g_lead + g_lag <= g_active then
report "FAIL: for a 2-bit frame, framing should exceed active time" severity error;
errors <= errors + 1;
end if;
if errors = 0 then
report "PASS: lead, active and lag are measured separately, a frame with no "
& "bits is all lead, and on a short frame the framing overhead exceeds "
& "the clocked time" severity note;
else
report "FAILED with " & integer'image(errors) & " error(s)" severity error;
end if;
halt <= true;
wait;
end process;
watchdog : process is
begin
wait for 900 us;
if not halt then
report "FAIL: watchdog timeout" severity failure;
end if;
wait;
end process;
end architecture tb;Parity
All three implement the same timer: identical ports and generics, asynchronous active-low reset, the CS-fall cycle counted as in-frame, the first bit's cycle attributed to active rather than lead, provisional lag folded back on a subsequent bit, the gap accumulator preserved across a frame opening, and all four outputs published together. All three testbenches report identical numbers — gap=13, lead=7, active=15, lag=5 for the reference frame, and 12 of 15 cycles framing for the short one.
7. Why a Verification Engineer Cares
// 1. THE property. The three in-frame buckets partition the frame, so the
// breakdown is a genuine decomposition rather than three estimates.
a_partition : assert property (
@(posedge clk) disable iff (!rst_n)
frame_valid |-> (lead_cycles + active_cycles + lag_cycles)
== in_frame_cycles_ref)
else $error("lead + active + lag does not equal the frame length");
// 2. A frame with no bits is entirely lead. Catches an implementation
// that leaves such a frame's time unattributed.
a_empty_is_all_lead : assert property (
@(posedge clk) disable iff (!rst_n)
(frame_valid && active_cycles == 0) |-> (lag_cycles == 0))
else $error("a frame with no active time reported lag");
// 3. The gap is preserved across the frame opening. Clearing it on CS
// fall is the natural-looking mistake, and it silently reports zero.
a_gap_nonzero_after_idle : assert property (
@(posedge clk) disable iff (!rst_n)
(frame_valid && idle_preceded) |-> (gap_cycles > 0))
else $error("gap reported as zero after an idle interval");
// 4. All four outputs update together, so a reader never sees a mix of
// two frames.
a_atomic_publish : assert property (
@(posedge clk) disable iff (!rst_n)
!frame_valid |=> ($stable(lead_cycles) && $stable(active_cycles) &&
$stable(lag_cycles) && $stable(gap_cycles)))
else $error("an output changed outside a publish");Property 1 is the one to write first, and it illustrates a general point about measurement hardware: the useful assertions are accounting identities, not behaviours. A timer that mis-attributes a few cycles still produces plausible numbers, and only a sum that must balance catches it.
What these prove. That the breakdown is coherent and atomically published. What they cannot prove is that the measured lead and lag meet the device's requirements — that is a datasheet comparison, and a design can measure its own violation perfectly.
Coverage should target the frame shapes whose breakdowns differ:
covergroup spi_overhead_cg @(posedge frame_valid);
// The ratio that matters: is this frame mostly moving data or mostly
// framing? A suite of long frames never sees the interesting case.
cp_framing_share : coverpoint
((lead_cycles + lag_cycles) * 100 / (lead_cycles + active_cycles + lag_cycles)) {
bins negligible = {[0:10]}; // long burst
bins moderate = {[11:40]};
bins dominant = {[41:100]}; // single-register access
}
cp_active : coverpoint active_cycles {
bins none = {0}; // an empty frame (Chapter 8.6)
bins tiny = {[1:16]};
bins large = {[17:$]};
}
// Back-to-back frames have a small gap; isolated ones have a large
// one. The overhead per access differs enormously between them.
cp_gap : coverpoint gap_cycles {
bins minimal = {[0:4]};
bins modest = {[5:64]};
bins idle = {[65:$]};
}
x_framing_active : cross cp_framing_share, cp_active;
endgroup8. Why an FPGA or ASIC Engineer Cares
Shortening the lead and lag is a real optimisation and a bounded one. Both are device requirements with datasheet minima, and many masters default to values well above them — a conservative controller may insert microseconds where the device needs tens of nanoseconds. Reading the actual requirement and configuring to it, with margin, is free throughput on small transfers. But it cannot go below the device's number, and on an FPGA slave part of that number is your own synchroniser latency §12.
The gap is often the largest single item and the easiest to overlook. Chapter 7.3 §7 showed why it must be enforced in hardware; §3 above shows it can exceed the CS lead and lag combined. A controller that enforces a conservative gap in hardware is safe and may be leaving half the small-transfer throughput unclaimed.
Measure before optimising the clock. The table in §4 makes the case quantitatively: above the knee, clock rate buys very little. Knowing where your design sits on that table requires the frame timer, and it is the difference between a justified decision and a guess.
Instrument the gap separately from the frame. It is tempting to fold the gap into "overhead", but it has a different cause and a different fix — the lead and lag are per-device setup, while the gap is per-device recovery and is only paid when transactions are back to back. Measuring them separately tells you which to attack.
9. Failure Signature — Throughput That Improves Far Less Than the Clock
Symptom. A design doubles its SPI clock from 25 MHz to 50 MHz expecting roughly twice the data rate. Measured throughput improves by about 20%. Everything is functionally correct, and doubling again to 100 MHz improves it by a further 10%.
What the diminishing pattern establishes. Each doubling helps less than the last, which is the exact signature of a fixed cost becoming dominant. If the bottleneck scaled with the clock, every doubling would give the same proportional gain; instead the gains are shrinking toward a limit, so something in the transaction does not depend on the clock at all.
Plausible mechanisms.
- Non-clocked overhead dominating — lead, lag and gap are a large share of each transaction, and none of them scales. This is §4's table being lived.
- Software issue rate, so the bus is idle between transactions regardless of clock (Chapter 9.1 §9).
- A rate limit clamping the effective divisor below what was configured (Chapter 9.4).
- Per-transaction device latency, such as a programming or conversion time, which is fixed in microseconds.
The discriminating observation. Measure the frame breakdown at both clock rates. If active_cycles halves while lead + lag + gap stays constant, the diagnosis is complete and §4's table predicts exactly how much further clocking will help — which is usually "not much".
If active_cycles does not halve, the clock is not actually doubling: check the effective divisor against the requested one, because a per-slave rate limit will silently clamp it.
And if both scale properly but throughput does not, the bus is idle and the problem is upstream.
Why the investigation goes wrong. Because a 20% gain from a 2× clock increase reads as something being broken, and the search goes to timing margins and signal integrity. Nothing is broken — the arithmetic simply says that when half the transaction is fixed in nanoseconds, doubling the other half cannot do better than 1.33×, and with framing dominant it does considerably worse.
10. Common Misconceptions
11. Reason It Through
Work this before reading the answer.
A design polls a status register once per millisecond: a 1-byte command and a 1-byte response, with a 200 ns CS lead, a 50 ns lag and a 500 ns enforced inter-frame gap. The bus runs at 25 MHz.
An engineer proposes moving to 100 MHz to "reduce the polling overhead". How much does that actually save, and what would save more?
Compute the transaction at 25 MHz. T_sclk = 40 ns, and the poll is 16 clocked bits.
clocked 16 × 40 ns = 640 ns
CS lead = 200 ns
CS lag = 50 ns
gap = 500 ns
─────────────────────────────────────
total per poll = 1390 nsNow at 100 MHz. T_sclk = 10 ns:
clocked 16 × 10 ns = 160 ns
non-clocked unchanged = 750 ns
─────────────────────────────────────
total per poll = 910 nsThe saving is 480 ns per poll — 35%, for a 4× clock increase. And the reason is stark: at 100 MHz, 750 of 910 ns (82%) is non-clocked. The clock now governs less than a fifth of the transaction, so even an infinitely fast bus could not get below 750 ns.
What would save more? Three things, in increasing order of effect.
Reduce the gap. At 500 ns it is the largest single item — bigger than the entire clocked portion at 100 MHz. If the device's actual requirement is 100 ns and the controller is enforcing 500 conservatively, that is 400 ns saved for a configuration change, comparable to the entire benefit of quadrupling the clock.
Reduce the CS lead. Same argument at 200 ns, and the same caution: it has a datasheet minimum that must be respected.
Stop polling. At one poll per millisecond, 1390 ns is 0.14% of the time — the polling is not a throughput problem at all. If the concern is CPU or power rather than bandwidth, a data-ready interrupt removes 100% of it, which no clock change can approach.
And the question worth asking before any of this. At 0.14% bus utilisation, what problem is being solved? If the answer is "the bus looks slow", the measurement in Chapter 9.1 §9 would have shown 99.86% idle and redirected the effort. Optimising a resource that is idle almost all the time is the most common wasted performance work there is.
The general lesson. When a transaction is dominated by fixed costs, clock rate is nearly irrelevant, and the fixed costs are usually configuration values chosen conservatively rather than physical limits. Read the datasheet minima, compare them against what the controller is actually inserting, and the saving is often larger than anything the clock can offer — and free.
12. Understanding Check
13. Summary
Overhead divides into two kinds with different behaviour. Clocked overhead — command, address, dummy — occupies SCLK edges and shrinks as the clock rises. Non-clocked overhead — CS lead, CS lag, inter-frame gap — carries no bits and is fixed in nanoseconds, because it is set by device requirements.
For a typical read at 50 MHz the budget is 800 ns clocked and 220 ns non-clocked, totalling about a microsecond of overhead per access.
Raising the clock attacks only the first. At 200 MHz more than half the overhead is time in which nothing is clocked, which is why there is a knee beyond which clock rate stops helping — and it lands, not coincidentally, near where SPI devices stop getting faster.
Reducing the number of accesses attacks both kinds at once, which is the quantitative case for the merging advice of Module 7.
A bit counter cannot see the non-clocked half. Measuring it needs a frame timer that splits a frame into lead, active and lag and separately records the preceding gap — with the defining property that the buckets partition the frame, so the breakdown is a decomposition rather than three estimates.
Two implementation subtleties carry the design: a bitless cycle is only provisionally lag and folds back into active if another bit follows, and the gap accumulator must survive the frame opening or every gap reads as zero.
And when a clock increase disappoints, the diminishing pattern itself is the diagnosis — with the fixed costs usually turning out to be conservative configuration, not physics.
14. What Comes Next
The overhead is per access, so the obvious remedy is fewer, larger accesses. Chapter 9.3 — Burst Efficiency and Transfer Sizing makes that precise: exactly how efficiency rises with burst length, where the curve flattens and why pushing past that point buys nothing, what limits burst length in practice, and the hardware that splits a transfer into legal bursts — in all three HDLs.
Continue learning
Related tutorials
- Related topic
Raw vs Effective Payload Throughput
The three throughput numbers a link has — raw bit rate, protocol rate and effective payload rate — why they differ by four times on an ordinary flash read, and the performance counter that measures the gap instead of estimating it.
- Related topic
Launch and Sample Edges
One edge of each bit time places a bit on the wire, the other captures it, and they must never be the same edge. Why the separation is forced, why it buys half a period, and how RTL maps physical edges onto those roles.
- Related topic
Deriving Mode Behaviour from CPOL and CPHA
The four SPI modes are a two-bit truth table you can rebuild in seconds. The standard numbering, the derivation, the complete mode decoder in three HDLs, and the assertions that keep a configurable design honest.
- Related topic
Command, Address, and Data Phases
How a device layers a transaction onto a raw byte stream: why the opcode decides the shape of everything after it, how a slave tracks phases with no phase marker, and the sequencer that requires in three HDLs.
