SPI · Module 7
Burst Transfers and Auto-Increment
How a device advances its own address through a burst, the three wrap policies real devices implement, why read bursts cross page boundaries while write bursts wrap inside one, and the address generator in three HDLs.
Chapter 7.1 held a frame open and streamed bytes without asking where they were going. This chapter supplies that, and the answer contains an asymmetry most engineers meet the hard way.
One address was sent. A thousand bytes follow. What address does byte 500 land on — and is the answer the same for a read as for a write?
It is not, and the reason is physical rather than arbitrary.
1. Who Advances the Address
The master sends one address, at the start of the frame (Chapter 4.4). Every byte after the first is placed by an internal pointer inside the device, advancing once per byte.
That pointer is invisible. The master cannot read it, cannot set it mid-burst, and cannot resynchronise it — Chapter 7.1 §1's "the frame is the transaction" seen from the address side. The master's only means of control is where it started and when it stops.
So a burst's correctness rests entirely on the master and the device agreeing about a rule the master cannot observe. Getting that rule wrong produces data in the wrong place, silently, which is the failure in §8.
2. Three Wrap Policies
Real devices implement three, and the choice is per-operation rather than per-device.
| Policy | Behaviour | Where it applies |
|---|---|---|
| Free | increments across the whole address space | read bursts, on nearly every device |
| Page-bounded | increments within a page, wrapping to its base | write bursts to page-organised memory |
| Window | increments within a small aligned window | wrap-around burst reads |
Free is the intuitive one: address 0x0000FF is followed by 0x000100, crossing into the next page without ceremony.
Page-bounded is Chapter 5.3's behaviour: 0x0000FF is followed by 0x000000, back to the base of the same page.
Window is less common and worth recognising. Some devices support a burst read that wraps inside an aligned window of 8, 16, 32 or 64 bytes, so a cache line can be fetched starting at the word that was actually needed. The remaining words then arrive by wrapping around to the start of the line.
3. The Read/Write Asymmetry
Same start, same steps, different destinations
7 cycles4. Why the Asymmetry Exists
It looks arbitrary and is not. It follows from what the device physically does with the bytes.
A write accumulates into a buffer. Chapter 5.3 §4 established that flash collects incoming bytes in a page buffer and commits the whole buffer in one programming operation. That buffer is one page and has no carry-out — so the pointer indexing it cannot leave the page. The wrap is a description of a counter addressing a fixed-size RAM.
A read does not accumulate anything. Bytes are fetched and shifted out; nothing is being filled. There is no buffer to be bounded by, so the address simply increments. A read burst can run from one end of the device to the other in a single frame, and on serial flash that is exactly how firmware is loaded.
So the asymmetry is not a protocol quirk — it is the presence or absence of a page buffer, visible at the interface.
A useful consequence: the constraint is on the operation, not the device. The same flash part has a page-bounded write and a free read, and a driver that applies one rule to both is wrong half the time — and wrong in the safe direction for reads (unnecessary splitting, costing throughput) and the unsafe direction for writes (silent corruption).
5. The Transaction
6. Building the Burst Address Generator — Three HDLs
The circuit
Circuit. A loadable counter whose increment width is selected at runtime.
State. One address register of ADDR_W bits.
Datapath. Three mutually exclusive increment behaviours. Free mode increments the whole register; page mode increments only addr[PAGE_BITS-1:0]; window mode only addr[WRAP_BITS-1:0]. In the two bounded modes the bits above the window are never written, which is what guarantees the burst cannot escape.
Control. load has priority over step, so an address phase always wins over a stray data strobe. wrap_mode selects the policy and is expected to change between transactions, never within one.
Clock. The system clock; step is one pulse per transferred byte.
Reset. Asynchronous, active-low, to address zero with both strobes low.
Enables. The register holds unless load or step is asserted.
Timing. wrapped and crossed_page are single-cycle strobes. crossed_page is not an error — it reports that a free burst left a page, which is legal and is worth knowing because some devices take a timing penalty at the boundary and because it is a useful coverage event.
Synthesis. ADDR_W flip-flops plus three increment paths. The free path needs a full-width adder; the bounded paths need only PAGE_BITS and WRAP_BITS adders respectively, and Chapter 5.3 §8's observation still holds — the bounded increments have no carry chain out of their window. Supporting all three costs the full-width adder plus a multiplexer.
Limitations. Page and window sizes are fixed at elaboration. A device supporting several page sizes makes them inputs, turning the all-ones detectors into masked comparisons.
// spi_burst_addr.sv — the burst address generator, with all three real
// wrap behaviours.
//
// Chapter 5.3 built a page-bounded pointer for WRITES. That is only one of
// the three policies real devices implement, and which one applies depends
// on the operation rather than on the device:
//
// WRAP_NONE free increment across the whole address space. What a READ
// burst does on nearly every device -- there is no buffer to
// be bounded by, because nothing is being accumulated.
// WRAP_PAGE increment confined to a page. What a WRITE burst does,
// because the device is filling a page-sized buffer (5.3 §4).
// WRAP_BUF increment confined to a small aligned window. Used by
// wrap-around burst read modes, where a cache line is fetched
// starting at the critical word.
//
// The asymmetry is the point: the SAME device auto-increments differently
// for a read than for a write, and assuming otherwise corrupts data.
module spi_burst_addr #(
parameter int ADDR_W = 24,
parameter int PAGE_BITS = 8, // page = 2**PAGE_BITS bytes
parameter int WRAP_BITS = 4 // window = 2**WRAP_BITS bytes
) (
input logic clk,
input logic rst_n,
input logic load, // latch the start address
input logic [ADDR_W-1:0] load_addr,
input logic step, // one pulse per transferred byte
input logic [1:0] wrap_mode, // see the localparams below
output logic [ADDR_W-1:0] addr,
output logic wrapped, // pulse: pointer rolled inside its window
output logic crossed_page // pulse: a free increment left a page
);
localparam logic [1:0] WRAP_NONE = 2'd0,
WRAP_PAGE = 2'd1,
WRAP_BUF = 2'd2;
// All-ones in the low bits means the next step rolls within that window.
logic page_full, buf_full;
assign page_full = &addr[PAGE_BITS-1:0];
assign buf_full = &addr[WRAP_BITS-1:0];
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
addr <= '0;
wrapped <= 1'b0;
crossed_page <= 1'b0;
end else begin
wrapped <= 1'b0; // both outputs are strobes
crossed_page <= 1'b0;
if (load) begin
addr <= load_addr;
end else if (step) begin
case (wrap_mode)
WRAP_PAGE: begin
if (page_full) begin
// Back to the page base, NOT into the next page.
addr[PAGE_BITS-1:0] <= '0;
wrapped <= 1'b1;
end else begin
addr[PAGE_BITS-1:0] <= addr[PAGE_BITS-1:0] + 1'b1;
end
end
WRAP_BUF: begin
if (buf_full) begin
addr[WRAP_BITS-1:0] <= '0;
wrapped <= 1'b1;
end else begin
addr[WRAP_BITS-1:0] <= addr[WRAP_BITS-1:0] + 1'b1;
end
end
default: begin // WRAP_NONE
addr <= addr + 1'b1;
// Not an error -- reads cross pages freely. Reported
// because some devices take a timing penalty at the
// boundary, and because it is useful in coverage.
if (page_full) crossed_page <= 1'b1;
end
endcase
end
end
end
endmodule// spi_burst_addr_tb.sv — the three policies, and the read/write asymmetry
// demonstrated on the SAME start address.
`timescale 1ns/1ps
module spi_burst_addr_tb;
logic clk = 0, rst_n = 0;
always #5 clk = ~clk;
localparam int AW = 24, PB = 8, WB = 4;
localparam logic [1:0] WRAP_NONE = 2'd0, WRAP_PAGE = 2'd1, WRAP_BUF = 2'd2;
logic load = 0, step = 0;
logic [1:0] wrap_mode = WRAP_NONE;
logic [AW-1:0] load_addr = '0, addr;
logic wrapped, crossed_page;
spi_burst_addr #(.ADDR_W(AW), .PAGE_BITS(PB), .WRAP_BITS(WB)) dut (
.clk, .rst_n, .load, .load_addr, .step, .wrap_mode,
.addr, .wrapped, .crossed_page);
int errors = 0, wraps = 0, crossings = 0;
int w0, c0; // per-section baselines
task automatic chk(input string what, input int g, input int e);
if (g !== e) begin $display("FAIL %s: got 0x%0h exp 0x%0h", what, g, e); errors++; end
endtask
always @(posedge clk) if (rst_n) begin
if (wrapped) wraps++;
if (crossed_page) crossings++;
end
task automatic do_load(input logic [AW-1:0] a);
load_addr = a; load = 1; @(negedge clk); load = 0; @(negedge clk);
endtask
task automatic do_step();
step = 1; @(negedge clk); step = 0; @(negedge clk);
endtask
initial begin
repeat (3) @(negedge clk); rst_n = 1; @(negedge clk);
// ================= READ burst: free increment =================
w0 = wraps; c0 = crossings;
wrap_mode = WRAP_NONE;
do_load(24'h0000FE);
do_step(); chk("read: FE -> FF", addr, 24'h0000FF);
chk("read: no crossing yet", crossings - c0, 0);
do_step();
// The decisive check: a read burst CROSSES into the next page.
chk("read: FF -> 0x000100 (crossed)", addr, 24'h000100);
chk("read: crossing reported", crossings - c0, 1);
chk("read: nothing wrapped", wraps - w0, 0);
// ================= WRITE burst: same address, page-bounded =======
w0 = wraps; c0 = crossings;
wrap_mode = WRAP_PAGE;
do_load(24'h0000FE);
do_step(); chk("write: FE -> FF", addr, 24'h0000FF);
do_step();
// Same start, same step count, DIFFERENT destination.
chk("write: FF -> page base 0x000000", addr, 24'h000000);
chk("write: wrap reported", wraps - w0, 1);
chk("write: no page crossing", crossings - c0, 0);
// The page base must survive a whole page of steps.
do_load(24'h0A0700);
for (int i = 0; i < 256; i++) begin
do_step();
if (addr[AW-1:PB] !== 24'h0A07) begin
$display("FAIL: page base moved at step %0d -> 0x%0h", i, addr);
errors++;
end
end
chk("write: full page returns to base", addr, 24'h0A0700);
// ================= WRAP-AROUND read: 16-byte window ==============
w0 = wraps; c0 = crossings;
wrap_mode = WRAP_BUF;
do_load(24'h00200E);
do_step(); chk("wrap: 0E -> 0F", addr, 24'h00200F);
do_step();
// Wraps within the aligned 16-byte window, not the page.
chk("wrap: 0F -> window base 0x002000", addr, 24'h002000);
chk("wrap: reported", wraps - w0, 1);
// A window step must not disturb anything above WRAP_BITS.
do_load(24'h00FFF8);
for (int i = 0; i < 16; i++) begin
do_step();
if (addr[AW-1:WB] !== 24'h00FFF) begin
$display("FAIL: window base moved at step %0d -> 0x%0h", i, addr);
errors++;
end
end
// ================= load beats step ===============================
wrap_mode = WRAP_NONE;
load_addr = 24'h123456; load = 1; step = 1; @(negedge clk);
load = 0; step = 0; @(negedge clk);
chk("load takes priority over step", addr, 24'h123456);
if (errors == 0)
$display("PASS: a read burst crosses the page, a write burst wraps to the page base, a wrap-around burst wraps within its window, and the same start address gives three different destinations");
else
$display("FAILED with %0d error(s)", errors);
$finish;
end
initial begin #900000; $display("FAIL: watchdog timeout"); $finish; end
endmoduleThe structure of that testbench is the argument of §3. It loads the same address, 0x0000FE, under free and then page-bounded mode, takes the same number of steps, and asserts two different destinations. A test that exercised each mode from a different address would verify both behaviours and demonstrate nothing about the asymmetry.
The full-page loop also re-asserts Chapter 5.3's invariant on every step rather than only at the boundary, and the window loop does the same for its narrower window — catching a design that wraps correctly once and then drifts.
// spi_burst_addr.v — the same burst address generator in Verilog-2001.
module spi_burst_addr #(
parameter ADDR_W = 24,
parameter PAGE_BITS = 8,
parameter WRAP_BITS = 4
) (
input wire clk,
input wire rst_n,
input wire load,
input wire [ADDR_W-1:0] load_addr,
input wire step,
input wire [1:0] wrap_mode,
output reg [ADDR_W-1:0] addr,
output reg wrapped,
output reg crossed_page
);
localparam WRAP_NONE = 2'd0,
WRAP_PAGE = 2'd1,
WRAP_BUF = 2'd2;
wire page_full = &addr[PAGE_BITS-1:0];
wire buf_full = &addr[WRAP_BITS-1:0];
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
addr <= {ADDR_W{1'b0}};
wrapped <= 1'b0;
crossed_page <= 1'b0;
end else begin
wrapped <= 1'b0;
crossed_page <= 1'b0;
if (load) begin
addr <= load_addr;
end else if (step) begin
case (wrap_mode)
WRAP_PAGE: begin
if (page_full) begin
// Back to the page base, NOT into the next page.
addr[PAGE_BITS-1:0] <= {PAGE_BITS{1'b0}};
wrapped <= 1'b1;
end else begin
addr[PAGE_BITS-1:0] <= addr[PAGE_BITS-1:0] + 1'b1;
end
end
WRAP_BUF: begin
if (buf_full) begin
addr[WRAP_BITS-1:0] <= {WRAP_BITS{1'b0}};
wrapped <= 1'b1;
end else begin
addr[WRAP_BITS-1:0] <= addr[WRAP_BITS-1:0] + 1'b1;
end
end
default: begin // WRAP_NONE
addr <= addr + 1'b1;
// Not an error -- reads cross pages freely.
if (page_full) crossed_page <= 1'b1;
end
endcase
end
end
end
endmodule// spi_burst_addr_tb.v — the same checks in Verilog-2001.
`timescale 1ns/1ps
module spi_burst_addr_tb;
reg clk = 0, rst_n = 0;
always #5 clk = ~clk;
parameter AW = 24, PB = 8, WB = 4;
localparam WRAP_NONE = 2'd0, WRAP_PAGE = 2'd1, WRAP_BUF = 2'd2;
reg load = 0, step = 0;
reg [1:0] wrap_mode = WRAP_NONE;
reg [AW-1:0] load_addr = 0;
wire [AW-1:0] addr;
wire wrapped, crossed_page;
spi_burst_addr #(.ADDR_W(AW), .PAGE_BITS(PB), .WRAP_BITS(WB)) dut (
.clk(clk), .rst_n(rst_n), .load(load), .load_addr(load_addr), .step(step),
.wrap_mode(wrap_mode), .addr(addr), .wrapped(wrapped),
.crossed_page(crossed_page));
integer errors = 0, wraps = 0, crossings = 0, w0 = 0, c0 = 0, i;
task chk;
input [80*8-1:0] what;
input [31:0] g, e;
begin
if (g !== e) begin
$display("FAIL %0s: got 0x%0h exp 0x%0h", what, g, e);
errors = errors + 1;
end
end
endtask
always @(posedge clk) if (rst_n) begin
if (wrapped) wraps = wraps + 1;
if (crossed_page) crossings = crossings + 1;
end
task do_load;
input [AW-1:0] a;
begin
load_addr = a; load = 1; @(negedge clk); load = 0; @(negedge clk);
end
endtask
task do_step;
begin
step = 1; @(negedge clk); step = 0; @(negedge clk);
end
endtask
initial begin
repeat (3) @(negedge clk); rst_n = 1; @(negedge clk);
// READ burst: free increment
w0 = wraps; c0 = crossings;
wrap_mode = WRAP_NONE;
do_load(24'h0000FE);
do_step; chk("read: FE -> FF", addr, 24'h0000FF);
chk("read: no crossing yet", crossings - c0, 0);
do_step;
chk("read: FF -> 0x000100 (crossed)", addr, 24'h000100);
chk("read: crossing reported", crossings - c0, 1);
chk("read: nothing wrapped", wraps - w0, 0);
// WRITE burst: same address, page-bounded
w0 = wraps; c0 = crossings;
wrap_mode = WRAP_PAGE;
do_load(24'h0000FE);
do_step; chk("write: FE -> FF", addr, 24'h0000FF);
do_step;
chk("write: FF -> page base 0x000000", addr, 24'h000000);
chk("write: wrap reported", wraps - w0, 1);
chk("write: no page crossing", crossings - c0, 0);
do_load(24'h0A0700);
for (i = 0; i < 256; i = i + 1) begin
do_step;
if (addr[AW-1:PB] !== 24'h0A07) begin
$display("FAIL: page base moved at step %0d -> 0x%0h", i, addr);
errors = errors + 1;
end
end
chk("write: full page returns to base", addr, 24'h0A0700);
// WRAP-AROUND read: 16-byte window
w0 = wraps; c0 = crossings;
wrap_mode = WRAP_BUF;
do_load(24'h00200E);
do_step; chk("wrap: 0E -> 0F", addr, 24'h00200F);
do_step;
chk("wrap: 0F -> window base 0x002000", addr, 24'h002000);
chk("wrap: reported", wraps - w0, 1);
do_load(24'h00FFF8);
for (i = 0; i < 16; i = i + 1) begin
do_step;
if (addr[AW-1:WB] !== 24'h00FFF) begin
$display("FAIL: window base moved at step %0d -> 0x%0h", i, addr);
errors = errors + 1;
end
end
// load beats step
wrap_mode = WRAP_NONE;
load_addr = 24'h123456; load = 1; step = 1; @(negedge clk);
load = 0; step = 0; @(negedge clk);
chk("load takes priority over step", addr, 24'h123456);
if (errors == 0)
$display("PASS: a read burst crosses the page, a write burst wraps to the page base, a wrap-around burst wraps within its window, and the same start address gives three different destinations");
else
$display("FAILED with %0d error(s)", errors);
$finish;
end
initial begin #900000; $display("FAIL: watchdog timeout"); $finish; end
endmodule-- spi_burst_addr.vhd — the same burst address generator in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_burst_addr is
generic (
ADDR_W : positive := 24;
PAGE_BITS : positive := 8; -- page = 2**PAGE_BITS bytes
WRAP_BITS : positive := 4 -- window = 2**WRAP_BITS bytes
);
port (
clk : in std_logic;
rst_n : in std_logic;
load : in std_logic;
load_addr : in std_logic_vector(ADDR_W - 1 downto 0);
step : in std_logic;
wrap_mode : in std_logic_vector(1 downto 0);
addr : out std_logic_vector(ADDR_W - 1 downto 0);
wrapped : out std_logic; -- pulse
crossed_page : out std_logic -- pulse
);
end entity spi_burst_addr;
architecture rtl of spi_burst_addr is
constant WRAP_NONE : std_logic_vector(1 downto 0) := "00";
constant WRAP_PAGE : std_logic_vector(1 downto 0) := "01";
constant WRAP_BUF : std_logic_vector(1 downto 0) := "10";
constant PAGE_ONES : unsigned(PAGE_BITS - 1 downto 0) := (others => '1');
constant BUF_ONES : unsigned(WRAP_BITS - 1 downto 0) := (others => '1');
signal addr_r : unsigned(ADDR_W - 1 downto 0);
begin
addr <= std_logic_vector(addr_r);
process (clk, rst_n) is
begin
if rst_n = '0' then
addr_r <= (others => '0');
wrapped <= '0';
crossed_page <= '0';
elsif rising_edge(clk) then
wrapped <= '0';
crossed_page <= '0';
if load = '1' then
addr_r <= unsigned(load_addr);
elsif step = '1' then
if wrap_mode = WRAP_PAGE then
if addr_r(PAGE_BITS - 1 downto 0) = PAGE_ONES then
-- Back to the page base, NOT into the next page.
addr_r(PAGE_BITS - 1 downto 0) <= (others => '0');
wrapped <= '1';
else
addr_r(PAGE_BITS - 1 downto 0) <=
addr_r(PAGE_BITS - 1 downto 0) + 1;
end if;
elsif wrap_mode = WRAP_BUF then
if addr_r(WRAP_BITS - 1 downto 0) = BUF_ONES then
addr_r(WRAP_BITS - 1 downto 0) <= (others => '0');
wrapped <= '1';
else
addr_r(WRAP_BITS - 1 downto 0) <=
addr_r(WRAP_BITS - 1 downto 0) + 1;
end if;
else -- WRAP_NONE
addr_r <= addr_r + 1;
-- Not an error: reads cross pages freely.
if addr_r(PAGE_BITS - 1 downto 0) = PAGE_ONES then
crossed_page <= '1';
end if;
end if;
end if;
end if;
end process;
end architecture rtl;-- spi_burst_addr_tb.vhd — the same checks in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_burst_addr_tb is
end entity spi_burst_addr_tb;
architecture tb of spi_burst_addr_tb is
constant AW : positive := 24;
constant PB : positive := 8;
constant WB : positive := 4;
constant WRAP_NONE : std_logic_vector(1 downto 0) := "00";
constant WRAP_PAGE : std_logic_vector(1 downto 0) := "01";
constant WRAP_BUF : std_logic_vector(1 downto 0) := "10";
signal clk : std_logic := '0';
signal rst_n : std_logic := '0';
signal load : std_logic := '0';
signal step : std_logic := '0';
signal wrap_mode : std_logic_vector(1 downto 0) := WRAP_NONE;
signal load_addr : std_logic_vector(AW - 1 downto 0) := (others => '0');
signal halt : boolean := false;
signal addr : std_logic_vector(AW - 1 downto 0);
signal wrapped : std_logic;
signal crossed_page : std_logic;
signal errors : natural := 0;
signal wraps : natural := 0;
signal crossings : natural := 0;
begin
clk <= not clk after 5 ns when not halt else '0';
dut : entity work.spi_burst_addr
generic map (ADDR_W => AW, PAGE_BITS => PB, WRAP_BITS => WB)
port map (clk => clk, rst_n => rst_n, load => load, load_addr => load_addr,
step => step, wrap_mode => wrap_mode, addr => addr,
wrapped => wrapped, crossed_page => crossed_page);
counters : process (clk) is
begin
if rising_edge(clk) and rst_n = '1' then
if wrapped = '1' then
wraps <= wraps + 1;
end if;
if crossed_page = '1' then
crossings <= crossings + 1;
end if;
end if;
end process;
stim : process is
variable w0, c0 : natural;
procedure chk_v (what : string; g, e : std_logic_vector) is
begin
if g /= e then
report "FAIL " & what & ": got 0x" & to_hstring(g)
& " exp 0x" & to_hstring(e) severity error;
errors <= errors + 1;
end if;
end procedure;
procedure chk_n (what : string; g, e : natural) is
begin
if g /= e then
report "FAIL " & what & ": got " & integer'image(g)
& " exp " & integer'image(e) severity error;
errors <= errors + 1;
end if;
end procedure;
procedure do_load (a : std_logic_vector(AW - 1 downto 0)) is
begin
load_addr <= a; load <= '1';
wait until falling_edge(clk); load <= '0';
wait until falling_edge(clk);
end procedure;
procedure do_step is
begin
step <= '1';
wait until falling_edge(clk); step <= '0';
wait until falling_edge(clk);
end procedure;
begin
for i in 0 to 2 loop
wait until falling_edge(clk);
end loop;
rst_n <= '1';
wait until falling_edge(clk);
-- READ burst: free increment
w0 := wraps; c0 := crossings;
wrap_mode <= WRAP_NONE;
do_load(x"0000FE");
do_step; chk_v("read: FE -> FF", addr, x"0000FF");
chk_n("read: no crossing yet", crossings - c0, 0);
do_step;
chk_v("read: FF -> 0x000100 (crossed)", addr, x"000100");
chk_n("read: crossing reported", crossings - c0, 1);
chk_n("read: nothing wrapped", wraps - w0, 0);
-- WRITE burst: same address, page-bounded
w0 := wraps; c0 := crossings;
wrap_mode <= WRAP_PAGE;
wait until falling_edge(clk);
do_load(x"0000FE");
do_step; chk_v("write: FE -> FF", addr, x"0000FF");
do_step;
chk_v("write: FF -> page base 0x000000", addr, x"000000");
chk_n("write: wrap reported", wraps - w0, 1);
chk_n("write: no page crossing", crossings - c0, 0);
do_load(x"0A0700");
for i in 0 to 255 loop
do_step;
if addr(AW - 1 downto PB) /= x"0A07" then
report "FAIL: page base moved at step " & integer'image(i) severity error;
errors <= errors + 1;
end if;
end loop;
chk_v("write: full page returns to base", addr, x"0A0700");
-- WRAP-AROUND read: 16-byte window
w0 := wraps; c0 := crossings;
wrap_mode <= WRAP_BUF;
wait until falling_edge(clk);
do_load(x"00200E");
do_step; chk_v("wrap: 0E -> 0F", addr, x"00200F");
do_step;
chk_v("wrap: 0F -> window base 0x002000", addr, x"002000");
chk_n("wrap: reported", wraps - w0, 1);
do_load(x"00FFF8");
for i in 0 to 15 loop
do_step;
if addr(AW - 1 downto WB) /= x"00FFF" then
report "FAIL: window base moved at step " & integer'image(i) severity error;
errors <= errors + 1;
end if;
end loop;
-- load beats step
wrap_mode <= WRAP_NONE;
wait until falling_edge(clk);
load_addr <= x"123456"; load <= '1'; step <= '1';
wait until falling_edge(clk);
load <= '0'; step <= '0';
wait until falling_edge(clk);
chk_v("load takes priority over step", addr, x"123456");
if errors = 0 then
report "PASS: a read burst crosses the page, a write burst wraps to the page "
& "base, a wrap-around burst wraps within its window, and the same start "
& "address gives three different destinations" severity note;
else
report "FAILED with " & integer'image(errors) & " error(s)" severity error;
end if;
halt <= true;
wait;
end process;
watchdog : process is
begin
wait for 900 us;
if not halt then
report "FAIL: watchdog timeout" severity failure;
end if;
wait;
end process;
end architecture tb;Parity
All three implement the same counter: identical ports and generics, asynchronous active-low reset to zero, load taking priority over step, a full-width increment in free mode, increments confined to PAGE_BITS and WRAP_BITS in the bounded modes with the bits above never written, and single-cycle wrapped and crossed_page strobes. The VHDL holds the address as unsigned internally and converts at the port. All three testbenches run the same scenarios from the same addresses and produce the same values.
7. Why a Verification Engineer Cares
The invariants differ per mode, which is what makes this worth asserting rather than inspecting:
// 1. Page mode: the base is latched by `load` and nothing else may move it.
// This is Chapter 5.3's invariant, now conditioned on the mode.
a_page_base_fixed : assert property (
@(posedge clk) disable iff (!rst_n)
(!load && wrap_mode == WRAP_PAGE) |=> $stable(addr[ADDR_W-1:PAGE_BITS]))
else $error("page-bounded burst escaped its page");
// 2. Window mode: same idea, narrower window.
a_window_base_fixed : assert property (
@(posedge clk) disable iff (!rst_n)
(!load && wrap_mode == WRAP_BUF) |=> $stable(addr[ADDR_W-1:WRAP_BITS]))
else $error("wrap-around burst escaped its window");
// 3. Free mode must NEVER wrap. A free burst that wrapped would silently
// re-read data it had already returned -- the mirror of the write bug.
a_free_never_wraps : assert property (
@(posedge clk) disable iff (!rst_n)
(wrap_mode == WRAP_NONE) |-> !wrapped)
else $error("free burst wrapped -- page bound applied to a read");
// 4. A step always moves the pointer. Catches a stalled increment, which
// writes every byte of a burst to one address.
a_step_advances : assert property (
@(posedge clk) disable iff (!rst_n)
(step && !load) |=> (addr != $past(addr)))
else $error("step did not advance the pointer");
// 5. The mode must not change inside a frame. Switching policy mid-burst
// produces an address sequence that corresponds to no device at all.
a_mode_stable_in_frame : assert property (
@(posedge clk) disable iff (!rst_n) !cs_n |=> $stable(wrap_mode))
else $error("wrap_mode changed during a burst");Property 3 is the one this chapter adds to Chapter 5.3's set, and it catches the more likely engineering mistake: applying the page bound to a read. That error is invisible in a short burst, costs throughput in a long one, and silently returns duplicate data if the wrap is actually implemented.
What these prove. That each policy's boundary behaviour is correct and that the policy is stable within a frame. What they cannot prove is that the right policy was selected for the operation — a datasheet fact, and the failure in §8.
Coverage must cross the policy with the boundary relationship:
covergroup spi_burst_cg @(posedge cs_rose);
cp_mode : coverpoint cfg.wrap_mode {
bins free = {WRAP_NONE};
bins page = {WRAP_PAGE};
bins window = {WRAP_BUF};
}
// Chapter 5.3 §7's classification, now crossed with the policy.
cp_span : coverpoint burst_span_class {
bins inside_one = {INSIDE}; // never reaches a boundary
bins ends_at_edge = {ENDS_AT_EDGE}; // off-by-one territory
bins crosses_once = {CROSSES_ONCE};
bins crosses_many = {CROSSES_MANY};
}
// THE cross. A policy exercised only on bursts that never reach a
// boundary has not been exercised at all -- every mode behaves
// identically inside a page.
x_mode_span : cross cp_mode, cp_span;
cp_direction : coverpoint cfg.is_read { bins read = {1}; bins write = {0}; }
x_mode_dir : cross cp_mode, cp_direction;
endgroupx_mode_span is the cross that matters and the reason coverage on cp_mode alone is misleading: all three policies are identical for a burst that never reaches a boundary. A suite of short, page-aligned bursts covers every mode bin and distinguishes none of them.
x_mode_dir is the second, and it catches the configuration error directly: the bin for page-bounded read should be empty on a correct design, and its presence is a bug report.
8. Failure Signature — Burst Reads Work and Burst Writes Corrupt
Symptom. A driver reads large blocks from a flash device perfectly — firmware loads, checksums match, arbitrarily long transfers succeed. Writing large blocks corrupts data, and the corruption appears only for writes above a certain size or starting at certain addresses. Small writes always work.
What the working reads establish, and it is a lot. Mode, bit order, framing, command decode, address handling and the entire round-trip path are all correct — a long, correct read exercises every one of them. The fault is confined to something writes do and reads do not.
Plausible mechanisms.
- The driver applies the read's free-increment assumption to writes, sending bursts that cross page boundaries. The device wraps and the tail of the burst overwrites its head (Chapter 5.3 §9). This is §4's asymmetry, met the hard way, and it matches the size and alignment dependence exactly.
- A write-enable interlock not reissued per transaction (Chapter 5.1 §5) — but that produces missing writes, not corrupted ones.
- The programming time is not being respected between writes, so later transactions are ignored while the device is busy — again missing rather than corrupted.
The discriminating observation. Compute the corrupted write addresses modulo the device's page size. If corruption begins at a page base and the misplaced data came from the end of the burst, the page wrap is confirmed. That is arithmetic on addresses already in hand.
The second observation is even cheaper: check whether writes that fit entirely within one page always succeed. If they do, and only page-crossing writes fail, the diagnosis is complete — that is the precise signature of the policy asymmetry and nothing else produces it.
Why the investigation goes wrong. Because reads working perfectly is taken as proof that the link and the driver are sound, so the search moves to the write path's timing, the enable sequence, or the device's programming behaviour. The link is sound. The bug is that a single spi_transfer(addr, buf, len) helper was written once and used for both directions, and only one direction is allowed to cross pages.
9. Why an FPGA or ASIC Engineer Cares
Splitting is the driver's job and cannot be delegated. The controller has no knowledge of the device's page size, so a master cannot detect a crossing. A block-write routine must take the page size as a parameter, split the buffer at page boundaries, and issue one transaction per page — while a block-read routine must not, because splitting a read costs a full request per page for no benefit (Chapter 6.4 §4's arithmetic).
The first and last fragments are the awkward ones. A write starting mid-page and spanning several pages splits into a short leading fragment, some full pages, and a short trailing fragment. That is where off-by-one errors live, and it is why the ends_at_edge coverage bin exists: a burst whose last byte lands exactly on the final address of a page is correct and must not wrap, while one byte more must.
On a slave, the policy must match the buffer. If the address pointer wraps at 256 but the page buffer is addressed with more bits, the device writes past its buffer — memory corruption inside the device rather than the benign in-page overwrite the protocol implies. Deriving both from one parameter makes the mismatch impossible by construction.
Supporting all three policies costs a full-width adder. Free mode needs one; the bounded modes do not. If a design only ever performs one kind of burst, fixing the policy at elaboration removes the multiplexer and, for a bounded-only design, most of the adder.
10. Common Misconceptions
11. Reason It Through
Work this before reading the answer.
A device supports a wrap-around burst read with an aligned 32-byte window. A driver issues a burst read starting at address
0x1234and clocks 32 bytes, expecting0x1234through0x1253.What does it actually receive, and why might a design want this behaviour?
Find the window. The window is 32 bytes and aligned, so it spans 0x1220 to 0x123F — the addresses whose upper bits match and whose low five bits run 0 to 31. The start address 0x1234 has low bits 0x14 = 20, so it sits at offset 20 within the window.
Walk the burst. Bytes arrive from 0x1234 upward: 0x1234 … 0x123F is 12 bytes. The pointer then wraps to the window base and continues: 0x1220 … 0x1233 is the remaining 20 bytes.
So the driver receives the 32 bytes of the window 0x1220–0x123F, rotated so that the requested address comes first. It does not receive 0x1234–0x1253; the addresses above 0x123F were never read.
Why would anyone want that? Because it is exactly what a cache line fill needs. A processor stalls on one word, and that word is at 0x1234. The critical-word-first ordering delivers the stalled-on word immediately — the core can resume after the first transfer — while the rest of the line arrives behind it. Fetching the line in address order would make the core wait for up to a full line time to get the byte it actually needed.
That is why the feature exists on devices intended for execute-in-place, where the SPI flash is backing a processor's instruction fetch.
What breaks if the driver does not know? It gets 32 correct bytes belonging to the right window, in the wrong order, for the wrong range. Every byte is real data from the device, so nothing looks corrupt in the usual sense — checksums fail, but no individual byte is wrong. That is a particularly confusing signature, and it is Chapter 5.3 §9's "displaced but intact" rule appearing in a third form: rotated but intact.
How would you recognise it from a dump? Look for the data being present but rotated — the expected content appearing at a consistent offset, wrapping around a power-of-two boundary. Computing the start address modulo the window size gives the rotation amount directly, and if that matches, the diagnosis is complete.
The general lesson. The wrap policy is part of the command, not just the device. A device can offer a free-increment read and a wrap-around read as different opcodes, and choosing the wrong one returns real data in an unexpected arrangement rather than an error.
12. Understanding Check
13. Summary
During a burst the address is advanced by an internal pointer inside the device, invisible to the master. The master sends one address and controls only where the burst starts and when it stops, so correctness depends on both ends agreeing about a rule the master cannot observe.
There are three policies: free increment across the address space, page-bounded increment wrapping to the page base, and window increment wrapping inside a small aligned region. Which applies depends on the operation, not the device.
The read/write asymmetry is physical. A write accumulates into a page buffer with no carry-out, so its pointer cannot leave the page. A read accumulates nothing, so there is no bound and it crosses pages freely. The same part therefore places bytes differently for the two directions.
The practical consequence is that one transfer helper cannot serve both: writes must be split at page boundaries and reads must not, and sharing the routine is wrong in the unsafe direction for writes.
In RTL all three policies are one loadable counter whose increment width is selected at runtime, with the bits above the active window never written — which is what makes escape impossible rather than merely unlikely.
For verification, each policy needs its own boundary invariant, plus the property that a free burst never wraps, which catches the page bound being wrongly applied to a read. Coverage must cross the policy with the burst's relationship to the boundary, because all three policies are identical inside a page.
And when long reads work while long writes corrupt, the diagnosis is almost always the asymmetry: check whether writes confined to a single page always succeed.
14. What Comes Next
Chapters 7.1 and 7.2 kept one frame open for as long as possible. Chapter 7.3 — Back-to-Back Transactions and Inter-Frame Gap closes the module by asking the opposite question: when a design must issue many separate frames in quick succession, how quickly may chip select fall again after it rises? There is a minimum, it is a device property that appears in no SPI document, and violating it produces a device that misses transactions entirely — with the master, as always, receiving no indication.
Continue learning
Related tutorials
- Related topic
Deriving Mode Behaviour from CPOL and CPHA
The four SPI modes are a two-bit truth table you can rebuild in seconds. The standard numbering, the derivation, the complete mode decoder in three HDLs, and the assertions that keep a configurable design honest.
- Related topic
Transfer Width
SPI has no native word size. What a 12-bit ADC or 16-bit codec requires of a master, why chip select rather than a bit count delimits a frame, and a width-parameterized transfer engine in three HDLs.
- Related topic
Bit Ordering — MSB-First and LSB-First
Which end of the shift register goes out first, the two multiplexers that make the order configurable in three HDLs, and why a bit-order bug is perfectly deterministic and yet invisible on certain data.
- Related topic
Command, Address, and Data Phases
How a device layers a transaction onto a raw byte stream: why the opcode decides the shape of everything after it, how a slave tracks phases with no phase marker, and the sequencer that requires in three HDLs.
