SPI · Module 9
Burst Efficiency and Transfer Sizing
How burst length amortises overhead and where the returns flatten: the knee of the efficiency curve, the four constraints that really limit burst length, and the splitter that divides a transfer at page boundaries and caps.
Chapter 9.2 showed that overhead is paid per access. The obvious remedy is fewer, larger accesses — and this chapter makes that precise, because "larger is better" stops being true long before most people expect.
Overhead is about a microsecond per access. How long does a burst have to be before that stops mattering — and is there any reason not to make every burst as long as possible?
There is a knee, it arrives sooner than intuition suggests, and past it the remaining gains are small enough that other considerations take over.
1. The Curve
Take Chapter 9.2's budget: 40 cycles of clocked overhead plus 220 ns of non-clocked overhead, at 50 MHz where a cycle is 20 ns.
overhead time = 40 × 20 ns + 220 ns = 1020 ns per access
payload time = N bytes × 8 × 20 ns = 160 N nsEfficiency is payload time over total time:
N bytes payload total efficiency
1 160 ns 1180 ns 13.6 %
2 320 ns 1340 ns 23.9 %
4 640 ns 1660 ns 38.6 %
8 1280 ns 2300 ns 55.7 %
16 2560 ns 3580 ns 71.5 %
32 5120 ns 6140 ns 83.4 %
64 10240 ns 11260 ns 90.9 %
128 20480 ns 21500 ns 95.3 %
256 40960 ns 41980 ns 97.6 %
512 81920 ns 82940 ns 98.8 %2. Where the Returns Flatten
Read the increments, not the values:
1 → 2 bytes +10.3 points
2 → 4 +14.7
4 → 8 +17.1
8 → 16 +15.8
16 → 32 +11.9
32 → 64 +7.5
64 → 128 +4.4
128 → 256 +2.3
256 → 512 +1.2The gain per doubling peaks around 4-to-8 bytes and then falls away steadily. By 64 bytes a doubling buys 4 points; by 256 it buys 2.
The knee is where payload time equals overhead time. Here that is 1020 / 160 ≈ 6.4 bytes — and sure enough, the largest increments straddle it. That gives a rule worth carrying:
knee (bytes) = overhead_time / (8 × T_sclk)Below the knee, doubling the burst is transformative. Above it, you are chasing an asymptote.
3. Two Bursts, Compared
Identical overhead, different payload
7 cycles4. What Actually Limits Burst Length
If longer is better, why not always burst the maximum? Four real constraints, and only one of them is about the bus.
Page boundaries. A write burst cannot cross one (Chapter 7.2 §4), so the page size is a hard cap on write bursts regardless of what the master would prefer. Read bursts are usually unconstrained.
Buffer memory. A burst needs somewhere to go. A 4 KB read into a 256-byte buffer is not one transaction, and on a small microcontroller the buffer is often the binding constraint rather than anything about SPI.
Latency. A long burst occupies the bus for its whole duration, and SPI has no preemption. A 4 KB burst at 50 MHz takes 655 µs, during which no other device can be accessed. If another device needs servicing within 100 µs, the maximum burst is set by that deadline, not by efficiency.
Error granularity. A failure anywhere in a burst usually means retrying the whole burst. Longer bursts amortise overhead and amplify the cost of a retry — and on a noisy link the optimum is shorter than the efficiency curve alone suggests.
The third is the one most often missed, and it inverts the usual advice: on a shared bus with real-time requirements, bursts should be as short as efficiency permits, not as long as the buffer allows.
5. Choosing a Size
Putting the four constraints together:
burst = min( page_remaining, ← hard, for writes
buffer_space, ← hard
deadline_budget, ← hard, if other devices have deadlines
comfortably_above_knee ) ← soft, and smallThe last term is the point of §2: once you are a few times past the knee, the efficiency argument is finished. Choosing 64 bytes over 32 buys 7 points; choosing 512 over 256 buys 1.2. If any of the hard constraints suggests a smaller number, take it — the efficiency cost is negligible and the latency benefit is not.
Which is why the hardware in §6 takes max_burst as an input alongside the boundary: the cap is a system decision, not a protocol one.
6. Building the Burst Sizer — Three HDLs
The circuit
Circuit. Two registers and a three-way minimum.
State. The current address and the remaining length.
Datapath. Each burst length is the smallest of three limits: what remains, the configured cap, and the distance to the next boundary. Only the third depends on the current address, and it is the one drivers get wrong (Chapter 7.2 §8).
Control. A valid/ack handshake per burst; done pulses when the transfer is fully split. A zero-length transfer completes immediately rather than emitting a zero-length burst.
Clock and reset. System clock; asynchronous active-low reset.
Enables. Distance to a boundary is computed as 2^BND_BITS − (addr & mask) — no division and no modulo, which matters because a modulo on a 24-bit address would infer a divider.
Timing. burst_valid, burst_addr and burst_len are combinational from the registered state, so a consumer sees a coherent burst whenever valid is asserted.
Synthesis. An address register, a length register, a small subtractor for the boundary distance and two comparators. The boundary arithmetic is a mask and a subtract, not a divide.
Limitations. One boundary size, fixed at elaboration. A device supporting several page sizes makes BND_BITS an input, turning the mask into a variable one.
// spi_burst_sizer.sv — splitting a transfer into legal bursts.
//
// Chapter 7.2 §9 said splitting at page boundaries is the driver's job
// because the controller cannot know the device's page size. That is true of
// a generic controller; a controller told the page size can do it in
// hardware, and doing so removes an entire class of driver bug.
//
// Each emitted burst is the SMALLEST of three limits:
//
// remaining what is left of the transfer
// max_burst a controller or device cap on one burst
// to_boundary bytes until the next alignment boundary
//
// The third is the one drivers get wrong (Chapter 7.2 §8), and it is the
// only one that depends on the current address rather than on configuration.
module spi_burst_sizer #(
parameter int ADDR_W = 24,
parameter int LEN_W = 16,
parameter int BND_BITS = 8 // boundary = 2**BND_BITS bytes
) (
input logic clk,
input logic rst_n,
input logic start, // pulse: begin splitting
input logic [ADDR_W-1:0] start_addr,
input logic [LEN_W-1:0] total_len,
input logic [LEN_W-1:0] max_burst,
output logic burst_valid, // level: a burst is offered
output logic [ADDR_W-1:0] burst_addr,
output logic [LEN_W-1:0] burst_len,
input logic burst_ack, // pulse: burst consumed
output logic done // pulse: transfer fully split
);
logic [ADDR_W-1:0] addr;
logic [LEN_W-1:0] remaining;
logic active;
// Bytes from the current address to the next boundary. For a boundary of
// 2**BND_BITS this is just the complement of the low bits, plus one --
// no division, no modulo.
localparam logic [ADDR_W-1:0] BND_MASK = ADDR_W'((1 << BND_BITS) - 1);
logic [LEN_W-1:0] to_boundary;
assign to_boundary = LEN_W'((1 << BND_BITS) - (addr & BND_MASK));
// The burst length is the minimum of the three limits. Written as two
// comparisons so the intent survives reading.
logic [LEN_W-1:0] cap;
always_comb begin
cap = remaining;
if (max_burst != '0 && max_burst < cap) cap = max_burst;
if (to_boundary < cap) cap = to_boundary;
end
assign burst_valid = active && (remaining != '0);
assign burst_addr = addr;
assign burst_len = cap;
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
addr <= '0;
remaining <= '0;
active <= 1'b0;
done <= 1'b0;
end else begin
done <= 1'b0;
if (start) begin
addr <= start_addr;
remaining <= total_len;
active <= (total_len != '0);
// A zero-length transfer completes immediately rather than
// emitting a zero-length burst.
if (total_len == '0) done <= 1'b1;
end else if (active && burst_valid && burst_ack) begin
addr <= addr + ADDR_W'(cap);
remaining <= remaining - cap;
if (remaining == cap) begin
active <= 1'b0;
done <= 1'b1;
end
end
end
end
endmodule// spi_burst_sizer_tb.sv — splitting at a boundary, at a cap, and at neither.
`timescale 1ns/1ps
module spi_burst_sizer_tb;
logic clk = 0, rst_n = 0;
always #5 clk = ~clk;
localparam int AW = 24, LW = 16, BB = 8; // 256-byte boundary
logic start = 0, burst_ack = 0;
logic [AW-1:0] start_addr = '0;
logic [LW-1:0] total_len = '0, max_burst = '0;
logic burst_valid, done;
logic [AW-1:0] burst_addr;
logic [LW-1:0] burst_len;
spi_burst_sizer #(.ADDR_W(AW), .LEN_W(LW), .BND_BITS(BB)) dut (
.clk, .rst_n, .start, .start_addr, .total_len, .max_burst,
.burst_valid, .burst_addr, .burst_len, .burst_ack, .done);
int errors = 0, n_bursts = 0;
int total, bad; // assigned procedurally, never initialised statically
int got_addr [$], got_len [$];
task automatic chk(input string what, input int g, input int e);
if (g !== e) begin $display("FAIL %s: got 0x%0h exp 0x%0h", what, g, e); errors++; end
endtask
// Run a whole split, collecting every burst offered.
task automatic split(input logic [AW-1:0] a, input logic [LW-1:0] len,
input logic [LW-1:0] cap);
got_addr.delete(); got_len.delete(); n_bursts = 0;
start_addr = a; total_len = len; max_burst = cap;
start = 1; @(negedge clk); start = 0; @(negedge clk);
while (burst_valid && n_bursts < 32) begin
got_addr.push_back(burst_addr);
got_len.push_back(burst_len);
n_bursts++;
burst_ack = 1; @(negedge clk); burst_ack = 0; @(negedge clk);
end
repeat (2) @(negedge clk);
endtask
initial begin
repeat (3) @(negedge clk); rst_n = 1; @(negedge clk);
// --- crosses a 256-byte boundary: 0xFE + 8 bytes ---
split(24'h0000FE, 16'd8, 16'd256);
chk("crossing: two bursts", n_bursts, 2);
chk("crossing: burst 0 addr", got_addr[0], 24'h0000FE);
chk("crossing: burst 0 len", got_len[0], 2); // to the boundary
chk("crossing: burst 1 addr", got_addr[1], 24'h000100);
chk("crossing: burst 1 len", got_len[1], 6); // the remainder
// --- entirely inside one boundary: no split ---
split(24'h000110, 16'd16, 16'd256);
chk("inside: one burst", n_bursts, 1);
chk("inside: len", got_len[0], 16);
// --- ends exactly ON the boundary: still one burst ---
split(24'h0000F0, 16'd16, 16'd256);
chk("ends at boundary: one burst", n_bursts, 1);
chk("ends at boundary: len", got_len[0], 16);
// --- one byte past the boundary: two bursts of 16 and 1 ---
split(24'h0000F0, 16'd17, 16'd256);
chk("one past: two bursts", n_bursts, 2);
chk("one past: first len", got_len[0], 16);
chk("one past: second len", got_len[1], 1);
// --- a max_burst smaller than the boundary distance dominates ---
split(24'h000000, 16'd100, 16'd32);
chk("capped: four bursts", n_bursts, 4);
chk("capped: len 0", got_len[0], 32);
chk("capped: len 1", got_len[1], 32);
chk("capped: len 2", got_len[2], 32);
chk("capped: len 3", got_len[3], 4);
// --- a long transfer crossing several boundaries ---
split(24'h000000, 16'd600, 16'd256);
chk("long: three bursts", n_bursts, 3);
chk("long: len 0", got_len[0], 256);
chk("long: len 1", got_len[1], 256);
chk("long: len 2", got_len[2], 88);
// --- the lengths must always sum to the requested total ---
begin
total = 0;
foreach (got_len[i]) total += got_len[i];
chk("lengths sum to the total", total, 600);
end
// --- no burst may cross a boundary: check every one ---
begin
bad = 0;
foreach (got_addr[i])
if (((got_addr[i] & 24'hFF) + got_len[i]) > 256) bad++;
chk("no burst crosses a boundary", bad, 0);
end
// --- a zero-length transfer emits nothing ---
split(24'h000000, 16'd0, 16'd256);
chk("zero length: no bursts", n_bursts, 0);
if (errors == 0)
$display("PASS: every burst stops at the boundary or the cap, no burst crosses a boundary, the lengths sum to the request, and a zero-length transfer emits nothing");
else
$display("FAILED with %0d error(s)", errors);
$finish;
end
initial begin #900000; $display("FAIL: watchdog timeout"); $finish; end
endmoduleTwo of that testbench's checks are worth more than the individual cases. The lengths must sum to the request — a splitter that loses or duplicates a byte would otherwise pass every per-burst check. And no burst may cross a boundary, verified by examining every emitted burst rather than the ones the test expected to be interesting. Together they make the splitter's contract checkable without enumerating scenarios.
The 0x0000F0 + 16 and 0x0000F0 + 17 pair is the deliberate off-by-one probe: the first ends exactly on the boundary and must not split, the second exceeds it by one byte and must. A comparison written with >= instead of > fails exactly one of them.
// spi_burst_sizer.v — the same burst splitter in Verilog-2001.
module spi_burst_sizer #(
parameter ADDR_W = 24,
parameter LEN_W = 16,
parameter BND_BITS = 8 // boundary = 2**BND_BITS bytes
) (
input wire clk,
input wire rst_n,
input wire start,
input wire [ADDR_W-1:0] start_addr,
input wire [LEN_W-1:0] total_len,
input wire [LEN_W-1:0] max_burst,
output wire burst_valid,
output wire [ADDR_W-1:0] burst_addr,
output wire [LEN_W-1:0] burst_len,
input wire burst_ack,
output reg done
);
localparam [ADDR_W-1:0] BND_MASK = ((1 << BND_BITS) - 1);
reg [ADDR_W-1:0] addr;
reg [LEN_W-1:0] remaining;
reg active;
// Bytes to the next boundary: no division, no modulo.
wire [LEN_W-1:0] to_boundary = (1 << BND_BITS) - (addr & BND_MASK);
// The burst length is the minimum of the three limits.
reg [LEN_W-1:0] cap;
always @(*) begin
cap = remaining;
if (max_burst != {LEN_W{1'b0}} && max_burst < cap) cap = max_burst;
if (to_boundary < cap) cap = to_boundary;
end
assign burst_valid = active && (remaining != {LEN_W{1'b0}});
assign burst_addr = addr;
assign burst_len = cap;
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
addr <= {ADDR_W{1'b0}};
remaining <= {LEN_W{1'b0}};
active <= 1'b0;
done <= 1'b0;
end else begin
done <= 1'b0;
if (start) begin
addr <= start_addr;
remaining <= total_len;
active <= (total_len != {LEN_W{1'b0}});
if (total_len == {LEN_W{1'b0}}) done <= 1'b1;
end else if (active && burst_valid && burst_ack) begin
addr <= addr + cap;
remaining <= remaining - cap;
if (remaining == cap) begin
active <= 1'b0;
done <= 1'b1;
end
end
end
end
endmodule// spi_burst_sizer_tb.v — the same checks in Verilog-2001.
`timescale 1ns/1ps
module spi_burst_sizer_tb;
reg clk = 0, rst_n = 0;
always #5 clk = ~clk;
parameter AW = 24, LW = 16, BB = 8;
reg start = 0, burst_ack = 0;
reg [AW-1:0] start_addr = 0;
reg [LW-1:0] total_len = 0, max_burst = 0;
wire burst_valid, done;
wire [AW-1:0] burst_addr;
wire [LW-1:0] burst_len;
spi_burst_sizer #(.ADDR_W(AW), .LEN_W(LW), .BND_BITS(BB)) dut (
.clk(clk), .rst_n(rst_n), .start(start), .start_addr(start_addr),
.total_len(total_len), .max_burst(max_burst), .burst_valid(burst_valid),
.burst_addr(burst_addr), .burst_len(burst_len), .burst_ack(burst_ack),
.done(done));
integer errors = 0, n_bursts = 0, total = 0, bad = 0, i;
reg [AW-1:0] got_addr [0:31];
reg [LW-1:0] got_len [0:31];
task chk;
input [80*8-1:0] what;
input [31:0] g, e;
begin
if (g !== e) begin
$display("FAIL %0s: got 0x%0h exp 0x%0h", what, g, e);
errors = errors + 1;
end
end
endtask
task split;
input [AW-1:0] a;
input [LW-1:0] len, capv;
begin
n_bursts = 0;
start_addr = a; total_len = len; max_burst = capv;
start = 1; @(negedge clk); start = 0; @(negedge clk);
while (burst_valid && n_bursts < 32) begin
got_addr[n_bursts] = burst_addr;
got_len[n_bursts] = burst_len;
n_bursts = n_bursts + 1;
burst_ack = 1; @(negedge clk); burst_ack = 0; @(negedge clk);
end
repeat (2) @(negedge clk);
end
endtask
initial begin
repeat (3) @(negedge clk); rst_n = 1; @(negedge clk);
split(24'h0000FE, 16'd8, 16'd256);
chk("crossing: two bursts", n_bursts, 2);
chk("crossing: burst 0 addr", got_addr[0], 24'h0000FE);
chk("crossing: burst 0 len", got_len[0], 2);
chk("crossing: burst 1 addr", got_addr[1], 24'h000100);
chk("crossing: burst 1 len", got_len[1], 6);
split(24'h000110, 16'd16, 16'd256);
chk("inside: one burst", n_bursts, 1);
chk("inside: len", got_len[0], 16);
split(24'h0000F0, 16'd16, 16'd256);
chk("ends at boundary: one burst", n_bursts, 1);
chk("ends at boundary: len", got_len[0], 16);
split(24'h0000F0, 16'd17, 16'd256);
chk("one past: two bursts", n_bursts, 2);
chk("one past: first len", got_len[0], 16);
chk("one past: second len", got_len[1], 1);
split(24'h000000, 16'd100, 16'd32);
chk("capped: four bursts", n_bursts, 4);
chk("capped: len 0", got_len[0], 32);
chk("capped: len 1", got_len[1], 32);
chk("capped: len 2", got_len[2], 32);
chk("capped: len 3", got_len[3], 4);
split(24'h000000, 16'd600, 16'd256);
chk("long: three bursts", n_bursts, 3);
chk("long: len 0", got_len[0], 256);
chk("long: len 1", got_len[1], 256);
chk("long: len 2", got_len[2], 88);
total = 0;
for (i = 0; i < n_bursts; i = i + 1) total = total + got_len[i];
chk("lengths sum to the total", total, 600);
bad = 0;
for (i = 0; i < n_bursts; i = i + 1)
if (((got_addr[i] & 24'hFF) + got_len[i]) > 256) bad = bad + 1;
chk("no burst crosses a boundary", bad, 0);
split(24'h000000, 16'd0, 16'd256);
chk("zero length: no bursts", n_bursts, 0);
if (errors == 0)
$display("PASS: every burst stops at the boundary or the cap, no burst crosses a boundary, the lengths sum to the request, and a zero-length transfer emits nothing");
else
$display("FAILED with %0d error(s)", errors);
$finish;
end
initial begin #900000; $display("FAIL: watchdog timeout"); $finish; end
endmodule-- spi_burst_sizer.vhd — the same burst splitter in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_burst_sizer is
generic (
ADDR_W : positive := 24;
LEN_W : positive := 16;
BND_BITS : positive := 8 -- boundary = 2**BND_BITS bytes
);
port (
clk : in std_logic;
rst_n : in std_logic;
start : in std_logic; -- pulse
start_addr : in unsigned(ADDR_W - 1 downto 0);
total_len : in unsigned(LEN_W - 1 downto 0);
max_burst : in unsigned(LEN_W - 1 downto 0);
burst_valid : out std_logic; -- level
burst_addr : out unsigned(ADDR_W - 1 downto 0);
burst_len : out unsigned(LEN_W - 1 downto 0);
burst_ack : in std_logic; -- pulse
done : out std_logic -- pulse
);
end entity spi_burst_sizer;
architecture rtl of spi_burst_sizer is
constant BND_SIZE : natural := 2 ** BND_BITS;
-- Declaration initialisers keep the combinational minimum below from
-- comparing 'U' before the first reset. They are simulation hygiene, not
-- a substitute for reset: rst_n still loads these registers.
signal addr : unsigned(ADDR_W - 1 downto 0) := (others => '0');
signal remaining : unsigned(LEN_W - 1 downto 0) := (others => '0');
signal active : std_logic := '0';
signal to_boundary : unsigned(LEN_W - 1 downto 0) := (others => '0');
signal cap : unsigned(LEN_W - 1 downto 0) := (others => '0');
signal valid_i : std_logic;
begin
-- Bytes to the next boundary: no division, no modulo.
to_boundary <= to_unsigned(BND_SIZE, LEN_W)
- resize(addr(BND_BITS - 1 downto 0), LEN_W);
-- The burst length is the minimum of the three limits.
minimum : process (remaining, max_burst, to_boundary) is
variable c : unsigned(LEN_W - 1 downto 0);
begin
c := remaining;
if max_burst /= 0 and max_burst < c then
c := max_burst;
end if;
if to_boundary < c then
c := to_boundary;
end if;
cap <= c;
end process;
valid_i <= '1' when active = '1' and remaining /= 0 else '0';
burst_valid <= valid_i;
burst_addr <= addr;
burst_len <= cap;
process (clk, rst_n) is
begin
if rst_n = '0' then
addr <= (others => '0');
remaining <= (others => '0');
active <= '0';
done <= '0';
elsif rising_edge(clk) then
done <= '0';
if start = '1' then
addr <= start_addr;
remaining <= total_len;
if total_len = 0 then
active <= '0';
done <= '1'; -- a zero-length transfer emits no burst
else
active <= '1';
end if;
elsif active = '1' and valid_i = '1' and burst_ack = '1' then
addr <= addr + resize(cap, ADDR_W);
remaining <= remaining - cap;
if remaining = cap then
active <= '0';
done <= '1';
end if;
end if;
end if;
end process;
end architecture rtl;-- spi_burst_sizer_tb.vhd — the same checks in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_burst_sizer_tb is
end entity spi_burst_sizer_tb;
architecture tb of spi_burst_sizer_tb is
constant AW : positive := 24;
constant LW : positive := 16;
constant BB : positive := 8;
signal clk : std_logic := '0';
signal rst_n : std_logic := '0';
signal start : std_logic := '0';
signal burst_ack : std_logic := '0';
signal start_addr : unsigned(AW - 1 downto 0) := (others => '0');
signal total_len : unsigned(LW - 1 downto 0) := (others => '0');
signal max_burst : unsigned(LW - 1 downto 0) := (others => '0');
signal halt : boolean := false;
signal burst_valid : std_logic;
signal burst_addr : unsigned(AW - 1 downto 0);
signal burst_len : unsigned(LW - 1 downto 0);
signal done : std_logic;
signal errors : natural := 0;
type addr_arr is array (0 to 31) of natural;
type len_arr is array (0 to 31) of natural;
begin
clk <= not clk after 5 ns when not halt else '0';
dut : entity work.spi_burst_sizer
generic map (ADDR_W => AW, LEN_W => LW, BND_BITS => BB)
port map (clk => clk, rst_n => rst_n, start => start, start_addr => start_addr,
total_len => total_len, max_burst => max_burst,
burst_valid => burst_valid, burst_addr => burst_addr,
burst_len => burst_len, burst_ack => burst_ack, done => done);
stim : process is
variable got_addr : addr_arr := (others => 0);
variable got_len : len_arr := (others => 0);
variable n_bursts : natural;
variable total : natural;
variable bad : natural;
procedure chk_n (what : string; g, e : natural) is
begin
if g /= e then
report "FAIL " & what & ": got " & integer'image(g)
& " exp " & integer'image(e) severity error;
errors <= errors + 1;
end if;
end procedure;
procedure split (a : natural; len, capv : natural) is
begin
n_bursts := 0;
start_addr <= to_unsigned(a, AW);
total_len <= to_unsigned(len, LW);
max_burst <= to_unsigned(capv, LW);
start <= '1'; wait until falling_edge(clk);
start <= '0'; wait until falling_edge(clk);
while burst_valid = '1' and n_bursts < 32 loop
got_addr(n_bursts) := to_integer(burst_addr);
got_len(n_bursts) := to_integer(burst_len);
n_bursts := n_bursts + 1;
burst_ack <= '1'; wait until falling_edge(clk);
burst_ack <= '0'; wait until falling_edge(clk);
end loop;
for i in 0 to 1 loop wait until falling_edge(clk); end loop;
end procedure;
begin
for i in 0 to 2 loop wait until falling_edge(clk); end loop;
rst_n <= '1';
wait until falling_edge(clk);
split(16#0000FE#, 8, 256);
chk_n("crossing: two bursts", n_bursts, 2);
chk_n("crossing: burst 0 addr", got_addr(0), 16#0000FE#);
chk_n("crossing: burst 0 len", got_len(0), 2);
chk_n("crossing: burst 1 addr", got_addr(1), 16#000100#);
chk_n("crossing: burst 1 len", got_len(1), 6);
split(16#000110#, 16, 256);
chk_n("inside: one burst", n_bursts, 1);
chk_n("inside: len", got_len(0), 16);
split(16#0000F0#, 16, 256);
chk_n("ends at boundary: one burst", n_bursts, 1);
chk_n("ends at boundary: len", got_len(0), 16);
split(16#0000F0#, 17, 256);
chk_n("one past: two bursts", n_bursts, 2);
chk_n("one past: first len", got_len(0), 16);
chk_n("one past: second len", got_len(1), 1);
split(0, 100, 32);
chk_n("capped: four bursts", n_bursts, 4);
chk_n("capped: len 0", got_len(0), 32);
chk_n("capped: len 3", got_len(3), 4);
split(0, 600, 256);
chk_n("long: three bursts", n_bursts, 3);
chk_n("long: len 0", got_len(0), 256);
chk_n("long: len 1", got_len(1), 256);
chk_n("long: len 2", got_len(2), 88);
total := 0;
for i in 0 to n_bursts - 1 loop total := total + got_len(i); end loop;
chk_n("lengths sum to the total", total, 600);
bad := 0;
for i in 0 to n_bursts - 1 loop
if ((got_addr(i) mod 256) + got_len(i)) > 256 then
bad := bad + 1;
end if;
end loop;
chk_n("no burst crosses a boundary", bad, 0);
split(0, 0, 256);
chk_n("zero length: no bursts", n_bursts, 0);
if errors = 0 then
report "PASS: every burst stops at the boundary or the cap, no burst "
& "crosses a boundary, the lengths sum to the request, and a "
& "zero-length transfer emits nothing" severity note;
else
report "FAILED with " & integer'image(errors) & " error(s)" severity error;
end if;
halt <= true;
wait;
end process;
watchdog : process is
begin
wait for 900 us;
if not halt then
report "FAIL: watchdog timeout" severity failure;
end if;
wait;
end process;
end architecture tb;Parity
All three implement the same splitter: identical ports and generics, asynchronous active-low reset, a burst length that is the minimum of remaining, cap and boundary distance, a valid/ack handshake, a done pulse, and a zero-length transfer emitting no burst. All three testbenches run the same seven scenarios and produce identical splits — including 0xFE + 8 splitting into 2 and 6, and 600 bytes from zero splitting into 256, 256 and 88.
7. Why a Verification Engineer Cares
// 1. CONSERVATION. The emitted lengths sum to the requested total.
// A splitter that drops or duplicates a byte passes every per-burst
// check and fails only this one.
a_conservation : assert property (
@(posedge clk) disable iff (!rst_n)
done |-> (emitted_total == requested_total))
else $error("emitted lengths do not sum to the request");
// 2. CONTAINMENT. No burst crosses a boundary.
a_no_crossing : assert property (
@(posedge clk) disable iff (!rst_n)
(burst_valid) |-> (((burst_addr & BND_MASK) + burst_len) <= (1 << BND_BITS)))
else $error("a burst crosses a boundary");
// 3. Progress. Every acknowledged burst advances the address by its own
// length -- catches an off-by-one that would overlap or skip bytes.
a_progress : assert property (
@(posedge clk) disable iff (!rst_n)
(burst_valid && burst_ack) |=> (burst_addr == $past(burst_addr + burst_len)))
else $error("address did not advance by the burst length");
// 4. Termination. A non-zero transfer eventually completes -- a liveness
// property, and the one that catches a zero-length burst looping.
a_terminates : assert property (
@(posedge clk) disable iff (!rst_n)
(start && total_len != 0) |-> ##[1:$] done)
else $error("splitting never completed");Properties 1 and 2 are the pair worth internalising: conservation and containment. Almost every splitting, packing or fragmenting design has exactly this shape, and those two properties together are close to a complete specification — everything else is performance.
Property 4 matters because the natural failure of a splitter is not a wrong answer but an infinite loop: a zero-length burst that is acknowledged without advancing produces a design that hangs rather than corrupts, and only a liveness property catches it.
What these prove. That the splitting is correct and terminates. What they cannot prove is that BND_BITS matches the device's page size — a datasheet fact, and a perfectly correct splitter fed the wrong boundary produces perfectly-contained bursts that cross the real one.
Coverage should target the relationship to the boundary, as always:
covergroup spi_burst_cg @(posedge burst_valid);
cp_span : coverpoint burst_class {
bins whole_inside = {INSIDE};
bins ends_at_edge = {ENDS_AT_EDGE}; // must NOT split
bins one_over = {ONE_OVER}; // must split
bins spans_many = {SPANS_MANY};
}
// Which of the three limits decided this burst's length. A suite in
// which the cap always wins never tests the boundary logic at all.
cp_limiter : coverpoint which_limit {
bins by_remaining = {LIM_REMAIN};
bins by_cap = {LIM_CAP};
bins by_boundary = {LIM_BOUNDARY};
}
cp_len : coverpoint burst_len {
bins one = {1};
bins small = {[2:15]};
bins near_knee = {[4:16]}; // where the curve is steepest
bins large = {[17:$]};
}
x_span_limiter : cross cp_span, cp_limiter;
endgroupcp_limiter is the coverpoint that matters and the one usually absent. Three limits produce the length, and a test suite in which one always dominates has verified a third of the logic while reporting full coverage of the lengths.
8. Why an FPGA or ASIC Engineer Cares
Put the splitter in hardware if the boundary is known. Chapter 7.2 §9 said splitting is the driver's job because a generic controller cannot know the page size. A controller told the page size can do it, and doing so removes a whole class of driver bug — the one where a block-write helper is reused for reads, or the page size changes with a device variant.
The boundary arithmetic must not infer a divider. addr % page looks natural and is expensive; addr & mask is the same thing for a power-of-two boundary and costs nothing. Since page sizes are always powers of two, there is no reason to write the first.
Size max_burst from the latency budget, not from the buffer. §4's third constraint is the one that bites in a real system: the longest burst determines the worst-case wait for every other device on the bus. Computing it from the tightest deadline, rather than from whatever the buffer allows, is the difference between a design that meets its timing and one that meets it most of the time.
Expose which limit applied. A limited_by output costing three bits tells software whether a short burst was caused by the boundary, the cap or the remaining length — which turns a performance investigation from guesswork into a read of a register.
9. Failure Signature — Throughput That Collapses at Certain Buffer Addresses
Symptom. A block transfer achieves its expected rate most of the time. For particular source or destination addresses it runs at roughly half speed, reproducibly. The data is always correct, and the addresses that are slow have no obvious pattern until someone plots them.
What "data always correct" establishes. This is not a splitting bug in the correctness sense — conservation and containment are holding. It is purely a sizing problem, so the transfers are being split more finely than necessary.
Plausible mechanisms.
- Misalignment. A transfer starting just after a boundary is split into a tiny first burst and then full ones. A 256-byte transfer starting at offset 255 becomes a 1-byte burst and a 255-byte burst, paying two full overheads instead of one.
- A cap smaller than the boundary, so every burst is capped and the boundary logic never matters — the transfer is split more often than it needs to be.
- The buffer, not the bus, limiting the burst.
- The boundary being misconfigured smaller than the device's real page.
The discriminating observation. Compute the start address modulo the boundary size. If the slow cases cluster near the end of a boundary — offsets just below the page size — misalignment is confirmed, because those are exactly the transfers whose first burst is tiny.
Then read the per-burst limiter if one is exposed: if the boundary is the limiter on transfers that should have been capped, the boundary is set too small.
The fix, and why it is cheap. Align the buffers. A transfer starting on a boundary splits into whole pages with no runt burst, and on most systems allocating buffers page-aligned costs nothing. That single change removes the entire effect.
Why the investigation goes wrong. Because address-dependent performance looks like a memory or cache effect, and the search goes to the DMA or the memory controller. The SPI layer is behaving correctly and efficiently given the addresses it was handed — the inefficiency was decided by whoever allocated the buffer, several layers away.
10. Common Misconceptions
11. Reason It Through
Work this before reading the answer.
A system streams 4 KB blocks from flash and also services a motor-control sensor that must be read every 200 µs. The SPI bus runs at 40 MHz with 1 µs of overhead per access. An engineer sets the flash burst length to 4 KB "for maximum efficiency".
What happens, and what is the right burst length?
Compute how long a 4 KB burst occupies the bus.
4096 bytes × 8 bits × 25 ns = 819 µsThat is four times the sensor's deadline. While the flash burst runs, the bus is unavailable — SPI has no preemption, and deasserting CS mid-burst would abort the transfer (Chapter 8.7). So the sensor is read every 819 µs instead of every 200 µs, and the motor control fails.
What burst length does the deadline permit? The sensor read itself takes time — say 1 µs of overhead plus a couple of bytes — so allow 195 µs for the flash burst:
195 µs / (8 × 25 ns) = 975 bytesRound down to a page-friendly 512 bytes, which takes 102 µs and leaves comfortable margin.
What does that cost in efficiency? Using §1's method with 1 µs of overhead:
4096 bytes: 819 µs payload / 820 µs total = 99.88 %
512 bytes: 102 µs payload / 103 µs total = 99.03 %Less than one percentage point. The "maximum efficiency" choice bought 0.85 points and broke the system.
And the deeper reading. Both numbers are above 99%, which means the efficiency argument was settled long before either candidate. Going back to §1's curve: at 40 MHz with 1 µs overhead the knee is 1000 / 200 = 5 bytes, so anything above about 32 bytes is already in the flat region. The entire range from 64 bytes to 4 KB differs by under 2 points of efficiency, and within that range the choice should be made entirely on latency.
What if the deadline were tighter still — say 20 µs? Then the burst is about 100 bytes, still 96% efficient. Even a 10 µs deadline permits 50 bytes at 92%. The curve is forgiving in exactly the region where real-time constraints bite, which is a genuinely useful property of the arithmetic.
The general lesson. When two candidate designs both sit in the flat part of an efficiency curve, efficiency is not the deciding criterion — and continuing to optimise it means trading a real constraint for an imaginary gain. Compute where the knee is first; if both options are past it, decide on something else.
12. Understanding Check
13. Summary
Overhead is paid per access, so efficiency rises with burst length — but as a hyperbola, and hyperbolas flatten.
The knee is where payload time equals overhead time:
knee (bytes) = overhead_time / (8 × T_sclk)For a typical device at 50 MHz that is about six bytes. The gain per doubling peaks around 4-to-8 bytes and falls steadily after: by 64 bytes a doubling buys 4 points, by 256 it buys 2.
Burst length is limited by four things and only one is the bus: the page boundary for writes, the buffer, the latency imposed on other devices, and the cost of retrying on error. The third inverts the usual advice — on a shared bus with deadlines, bursts should be as short as efficiency permits rather than as long as memory allows.
In hardware the splitter is a three-way minimum of remaining, cap and boundary distance, with the distance computed by mask and subtract rather than a modulo, since page sizes are powers of two.
For verification the near-complete specification is two properties — conservation and containment — plus termination, because a splitter's natural failure is an infinite loop rather than a wrong answer. Coverage must record which limit decided each burst, or a suite where one dominates reports full coverage of a third of the logic.
And when two candidate sizes both sit past the knee, efficiency is not the deciding criterion: compute the knee first, and if both options clear it, decide on latency.
14. What Comes Next
Every chapter in this module has treated SCLK as a number that could be raised if only it were worth it. Chapter 9.4 — Maximum Practical SCLK asks what actually sets that number: the device's own maximum, the round trip that makes a link fail above a threshold, the board's loading and edge rates, and the divisor granularity that means the achievable rate is rarely the one you asked for — with the per-slave rate limiter that keeps a mixed bus honest, in all three HDLs.
Continue learning
Related tutorials
- Related topic
Address Fields, Dummy Cycles, and Burst Behaviour
Why three address bytes make 16 MB a category boundary, why dummy latency is counted in cycles and never rounds to bytes, what a device does at the end of a burst, and the planner that turns those numbers into a cycle-accurate schedule.
- Related topic
Chip-Select Generation
Chip select is a state machine, not a wire: the three ways deriving it from a busy signal fails, why the between-frames pause and the between-transactions pause are opposites, and why a select for a slave that is not fitted must be refused.
- Related topic
Launch and Sample Edges
One edge of each bit time places a bit on the wire, the other captures it, and they must never be the same edge. Why the separation is forced, why it buys half a period, and how RTL maps physical edges onto those roles.
- Related topic
Deriving Mode Behaviour from CPOL and CPHA
The four SPI modes are a two-bit truth table you can rebuild in seconds. The standard numbering, the derivation, the complete mode decoder in three HDLs, and the assertions that keep a configurable design honest.
