SPI · Module 5
Register and Multi-Byte Writes
The auto-incrementing write pointer, why that increment is bounded by a page so a burst crossing the boundary wraps and overwrites data already written, and the page-bounded pointer in three HDLs.
Chapter 5.1 wrote one register: opcode, address, one data byte, commit. But nothing in SPI ends a transaction after one byte — chip select does, and CS can stay asserted as long as the master likes.
The address phase named one location. The master keeps clocking data. Where does each byte after the first one go?
The answer is an internal pointer, and the interesting part is not that it increments but where it stops — because on most memory devices it does not carry into the next page. It wraps.
1. One Frame, Many Bytes
A multi-byte write is structurally identical to Chapter 5.1's single write with more data bytes before CS rises:
CS↓ opcode address D0 D1 D2 D3 … Dn CS↑Nothing separates the data bytes, nothing counts them, and there is no length field (Chapter 4.1 §2). The device takes bytes until the frame closes. The master decides the length by deciding when to deassert CS — which makes the length an implicit parameter the device learns only after the fact.
Two device conventions cover nearly everything you will meet:
Fixed-length writes. The device expects exactly N data bytes. Extra bytes are ignored, or overwrite (Chapter 5.1's last-byte-wins), and fewer means the transaction is incomplete.
Auto-incrementing writes. The device advances an internal pointer per byte, so one transaction fills a range. This is what memory devices and most multi-register peripherals do, and it is the subject of the rest of this chapter.
2. The Auto-Incrementing Pointer
The mechanism is simple: latch the address from the address phase, then advance it once per received data byte.
The value of this is real. Writing 256 bytes as 256 separate transactions costs 256 opcodes and 256 addresses — by Chapter 4.5 §6's arithmetic, roughly a third of the bus spent on overhead. As one auto-incrementing transaction it costs one opcode and one address, and efficiency approaches 100%.
The complication is equally real, and it is that the pointer is usually not a plain counter.
3. The Page Boundary
A burst crossing the page end wraps back to the page base
8 cyclesRead the bottom lane as the answer to "where did that byte go". The first two bytes land where the master intended. The third and fourth land at the start of the page, on top of data that may have been written moments earlier in the same transaction.
The corruption has three properties that together make it nasty:
- It is silent. No error, no status bit, no indication on the bus (Chapter 5.1 §5).
- It is data-dependent. A write that fits inside a page is perfectly correct, so the bug appears only for particular start addresses and lengths.
- It destroys data the master already wrote, so the corruption is inside the region the master believes it just successfully filled.
4. Why Pages Exist
The wrap looks like a design flaw until you see where it comes from, and then it looks inevitable.
A flash device does not write bytes individually. It accumulates incoming data in an internal page buffer, and when the transaction ends it commits the whole buffer to the array in one programming operation — a slow, high-voltage process that works on a whole page at a time for physical reasons.
That buffer is exactly one page and has no carry-out. The low address bits index into it; the high bits select which page it will be committed to and are latched once. So "the pointer wraps at the page boundary" is not a policy decision — it is a description of a counter addressing a fixed-size buffer.
Two consequences worth carrying:
The page size is a property of the silicon, published in the datasheet, and commonly 256 bytes on serial flash. It is not configurable.
A write that crosses a page boundary is not a single operation at any level. Even if the device did carry, it would need two separate programming operations. Splitting the write at page boundaries is therefore not a workaround for a quirk — it matches what the hardware actually does.
5. The Transaction
6. Building the Page-Bounded Pointer — Three HDLs
The circuit
Circuit. A loadable counter in which only the low bits count.
State. One address register of ADDR_W bits. The low PAGE_BITS are a counter; the upper bits are a latch.
Datapath. On load, the whole register takes the start address. On step, only addr[PAGE_BITS-1:0] changes — and when those bits are all ones, they return to zero instead of carrying.
Control. load has priority over step, so an address phase always wins over a stray data strobe in the same cycle. The testbenches assert this explicitly.
Clock. The system clock; step is one pulse per received byte.
Reset. Asynchronous, active-low, to address zero with the wrap strobe low.
Enables. The register holds unless load or step is asserted.
Timing. page_wrapped is a single-cycle strobe coincident with the wrapping step.
Synthesis. ADDR_W flip-flops, a PAGE_BITS-wide incrementer, and an all-ones detector. Notably the incrementer is only PAGE_BITS wide, not ADDR_W — a 24-bit address with 256-byte pages needs an 8-bit incrementer, not a 24-bit one, because there is no carry path out of the page. That is a real saving on a wide address and it falls out of the semantics rather than being an optimisation.
Limitations. One page size, fixed at elaboration. Devices supporting several page sizes make PAGE_BITS a configuration input, which turns the all-ones detector into a masked comparison.
// spi_addr_autoinc.sv — the auto-incrementing write pointer, and its page bound.
//
// A device that accepts many data bytes in one frame must decide where each
// one goes. The usual answer is an internal pointer that advances per byte.
// The pointer is almost never a plain counter: on page-organised memory only
// the LOW bits advance, so a burst running past the end of a page WRAPS to
// the start of that same page and overwrites bytes already written.
//
// Real devices do this silently. `page_wrapped` is exported here purely so
// the behaviour is observable -- it is a teaching and verification aid, and
// on a real part you would not get one.
module spi_addr_autoinc #(
parameter int ADDR_W = 16,
parameter int PAGE_BITS = 8 // page size = 2**PAGE_BITS bytes
) (
input logic clk,
input logic rst_n,
input logic load, // latch the start address
input logic [ADDR_W-1:0] load_addr,
input logic step, // one pulse per data byte
output logic [ADDR_W-1:0] addr,
output logic page_wrapped // pulse: the pointer rolled within its page
);
// The page base -- the high bits -- is fixed for the whole transaction.
// Nothing in this module can change it, which is the point: a burst
// cannot escape the page it started in.
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
addr <= '0;
page_wrapped <= 1'b0;
end else begin
page_wrapped <= 1'b0;
if (load) begin
addr <= load_addr;
end else if (step) begin
if (&addr[PAGE_BITS-1:0]) begin
// Low bits are all ones: the next byte lands back at the
// page base, NOT in the following page.
addr[PAGE_BITS-1:0] <= '0;
page_wrapped <= 1'b1;
end else begin
addr[PAGE_BITS-1:0] <= addr[PAGE_BITS-1:0] + 1'b1;
end
end
end
end
endmoduleThe page_wrapped output deserves comment because it is not something a real device gives you. It exists here so the behaviour is observable — for the testbench, for an assertion, and for a design that wants to flag the condition internally. A real part wraps in silence, which is exactly what makes §9 hard.
// spi_addr_autoinc_tb.sv — stepping, the wrap, and the invariant that the
// page base never changes.
`timescale 1ns/1ps
module spi_addr_autoinc_tb;
logic clk = 0, rst_n = 0;
always #5 clk = ~clk;
localparam int AW = 16, PB = 8;
localparam int PAGE = 1 << PB;
logic load = 0, step = 0;
logic [AW-1:0] load_addr = '0, addr;
logic page_wrapped;
spi_addr_autoinc #(.ADDR_W(AW), .PAGE_BITS(PB)) dut (
.clk, .rst_n, .load, .load_addr, .step, .addr, .page_wrapped);
int errors = 0, wraps = 0;
task automatic chk(input string what, input int got, input int exp);
if (got !== exp) begin $display("FAIL %s: got 0x%0h exp 0x%0h", what, got, exp); errors++; end
endtask
always @(posedge clk) if (rst_n && page_wrapped) wraps++;
task automatic do_load(input logic [AW-1:0] a);
load_addr = a; load = 1; @(negedge clk); load = 0; @(negedge clk);
endtask
task automatic do_step();
step = 1; @(negedge clk); step = 0; @(negedge clk);
endtask
logic [AW-1:0] base;
initial begin
repeat (3) @(negedge clk); rst_n = 1; @(negedge clk);
// --- ordinary stepping inside a page ---
do_load(16'h0140);
chk("loaded", addr, 16'h0140);
do_step(); chk("step 1", addr, 16'h0141);
do_step(); chk("step 2", addr, 16'h0142);
chk("no wrap yet", wraps, 0);
// --- step up to the page boundary and over it ---
do_load(16'h01FE);
do_step(); chk("to 0x01FF", addr, 16'h01FF);
chk("still no wrap", wraps, 0);
do_step();
// The decisive check: it must return to the PAGE BASE 0x0100,
// NOT advance to 0x0200.
chk("wrapped to page base", addr, 16'h0100);
chk("wrap reported", wraps, 1);
// --- the page base must be untouched across a whole page of steps ---
do_load(16'h0700);
base = addr & ~((1 << PB) - 1);
for (int i = 0; i < PAGE; i++) begin
do_step();
if ((addr & ~((1 << PB) - 1)) !== base) begin
$display("FAIL page base moved at step %0d: 0x%0h", i, addr);
errors++;
end
end
// A full page of steps from the base returns exactly to the base.
chk("full page returns to base", addr, 16'h0700);
// --- a page-aligned load followed by one step ---
do_load(16'h0800);
do_step(); chk("aligned +1", addr, 16'h0801);
// --- load takes priority over step in the same cycle ---
load_addr = 16'h0ABC; load = 1; step = 1; @(negedge clk);
load = 0; step = 0; @(negedge clk);
chk("load beats step", addr, 16'h0ABC);
if (errors == 0)
$display("PASS: the pointer advances within its page, wraps to the page base rather than crossing, and never alters the page base");
else
$display("FAILED with %0d error(s)", errors);
$finish;
end
initial begin #500000; $display("FAIL: watchdog timeout"); $finish; end
endmoduleThe strongest check in that testbench is the loop that steps a full page and verifies the high bits never change on any step. Checking only the wrap point would pass a design that carried correctly once and then drifted; asserting the invariant on every step is what actually pins down "the burst cannot escape its page".
// spi_addr_autoinc.v — the same page-bounded write pointer in Verilog-2001.
module spi_addr_autoinc #(
parameter ADDR_W = 16,
parameter PAGE_BITS = 8
) (
input wire clk,
input wire rst_n,
input wire load,
input wire [ADDR_W-1:0] load_addr,
input wire step,
output reg [ADDR_W-1:0] addr,
output reg page_wrapped
);
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
addr <= {ADDR_W{1'b0}};
page_wrapped <= 1'b0;
end else begin
page_wrapped <= 1'b0;
if (load) begin
addr <= load_addr;
end else if (step) begin
if (&addr[PAGE_BITS-1:0]) begin
// Low bits all ones: the next byte lands back at the page
// base, NOT in the following page.
addr[PAGE_BITS-1:0] <= {PAGE_BITS{1'b0}};
page_wrapped <= 1'b1;
end else begin
addr[PAGE_BITS-1:0] <= addr[PAGE_BITS-1:0] + 1'b1;
end
end
end
end
endmodule// spi_addr_autoinc_tb.v — the same checks in Verilog-2001.
`timescale 1ns/1ps
module spi_addr_autoinc_tb;
reg clk = 0, rst_n = 0;
always #5 clk = ~clk;
parameter AW = 16, PB = 8;
parameter PAGE = 1 << PB;
reg load = 0, step = 0;
reg [AW-1:0] load_addr = 0;
wire [AW-1:0] addr;
wire page_wrapped;
spi_addr_autoinc #(.ADDR_W(AW), .PAGE_BITS(PB)) dut (
.clk(clk), .rst_n(rst_n), .load(load), .load_addr(load_addr),
.step(step), .addr(addr), .page_wrapped(page_wrapped));
integer errors = 0, wraps = 0, i;
reg [AW-1:0] base;
task chk;
input [80*8-1:0] what;
input [31:0] got, exp;
begin
if (got !== exp) begin
$display("FAIL %0s: got 0x%0h exp 0x%0h", what, got, exp);
errors = errors + 1;
end
end
endtask
always @(posedge clk) if (rst_n && page_wrapped) wraps = wraps + 1;
task do_load;
input [AW-1:0] a;
begin
load_addr = a; load = 1; @(negedge clk); load = 0; @(negedge clk);
end
endtask
task do_step;
begin
step = 1; @(negedge clk); step = 0; @(negedge clk);
end
endtask
initial begin
repeat (3) @(negedge clk); rst_n = 1; @(negedge clk);
do_load(16'h0140);
chk("loaded", addr, 16'h0140);
do_step; chk("step 1", addr, 16'h0141);
do_step; chk("step 2", addr, 16'h0142);
chk("no wrap yet", wraps, 0);
do_load(16'h01FE);
do_step; chk("to 0x01FF", addr, 16'h01FF);
chk("still no wrap", wraps, 0);
do_step;
chk("wrapped to page base", addr, 16'h0100);
chk("wrap reported", wraps, 1);
do_load(16'h0700);
base = addr & ~((1 << PB) - 1);
for (i = 0; i < PAGE; i = i + 1) begin
do_step;
if ((addr & ~((1 << PB) - 1)) !== base) begin
$display("FAIL page base moved at step %0d: 0x%0h", i, addr);
errors = errors + 1;
end
end
chk("full page returns to base", addr, 16'h0700);
do_load(16'h0800);
do_step; chk("aligned +1", addr, 16'h0801);
load_addr = 16'h0ABC; load = 1; step = 1; @(negedge clk);
load = 0; step = 0; @(negedge clk);
chk("load beats step", addr, 16'h0ABC);
if (errors == 0)
$display("PASS: the pointer advances within its page, wraps to the page base rather than crossing, and never alters the page base");
else
$display("FAILED with %0d error(s)", errors);
$finish;
end
initial begin #500000; $display("FAIL: watchdog timeout"); $finish; end
endmodule-- spi_addr_autoinc.vhd — the same page-bounded write pointer in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_addr_autoinc is
generic (
ADDR_W : positive := 16;
PAGE_BITS : positive := 8 -- page size = 2**PAGE_BITS bytes
);
port (
clk : in std_logic;
rst_n : in std_logic;
load : in std_logic;
load_addr : in std_logic_vector(ADDR_W - 1 downto 0);
step : in std_logic;
addr : out std_logic_vector(ADDR_W - 1 downto 0);
page_wrapped : out std_logic
);
end entity spi_addr_autoinc;
architecture rtl of spi_addr_autoinc is
signal addr_r : unsigned(ADDR_W - 1 downto 0);
-- All-ones in the low bits means the next step rolls within the page.
constant LOW_ONES : unsigned(PAGE_BITS - 1 downto 0) := (others => '1');
begin
process (clk, rst_n) is
begin
if rst_n = '0' then
addr_r <= (others => '0');
page_wrapped <= '0';
elsif rising_edge(clk) then
page_wrapped <= '0';
if load = '1' then
addr_r <= unsigned(load_addr);
elsif step = '1' then
if addr_r(PAGE_BITS - 1 downto 0) = LOW_ONES then
-- Back to the page base, NOT on into the next page. The
-- high bits are deliberately left untouched.
addr_r(PAGE_BITS - 1 downto 0) <= (others => '0');
page_wrapped <= '1';
else
addr_r(PAGE_BITS - 1 downto 0) <=
addr_r(PAGE_BITS - 1 downto 0) + 1;
end if;
end if;
end if;
end process;
addr <= std_logic_vector(addr_r);
end architecture rtl;-- spi_addr_autoinc_tb.vhd — the same checks in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_addr_autoinc_tb is
end entity spi_addr_autoinc_tb;
architecture tb of spi_addr_autoinc_tb is
constant AW : positive := 16;
constant PB : positive := 8;
constant PAGE : positive := 2 ** PB;
signal clk : std_logic := '0';
signal rst_n : std_logic := '0';
signal load : std_logic := '0';
signal step : std_logic := '0';
signal load_addr : std_logic_vector(AW - 1 downto 0) := (others => '0');
signal halt : boolean := false;
signal addr : std_logic_vector(AW - 1 downto 0);
signal page_wrapped : std_logic;
signal errors : natural := 0;
signal wraps : natural := 0;
begin
clk <= not clk after 5 ns when not halt else '0';
dut : entity work.spi_addr_autoinc
generic map (ADDR_W => AW, PAGE_BITS => PB)
port map (clk => clk, rst_n => rst_n, load => load, load_addr => load_addr,
step => step, addr => addr, page_wrapped => page_wrapped);
counter : process (clk) is
begin
if rising_edge(clk) and rst_n = '1' and page_wrapped = '1' then
wraps <= wraps + 1;
end if;
end process;
stim : process is
variable base : unsigned(AW - 1 downto 0);
procedure chk_v (what : string; got, exp : std_logic_vector) is
begin
if got /= exp then
report "FAIL " & what & ": got 0x" & to_hstring(got)
& " exp 0x" & to_hstring(exp) severity error;
errors <= errors + 1;
end if;
end procedure;
procedure chk_n (what : string; got, exp : natural) is
begin
if got /= exp then
report "FAIL " & what & ": got " & integer'image(got)
& " exp " & integer'image(exp) severity error;
errors <= errors + 1;
end if;
end procedure;
procedure do_load (a : std_logic_vector(AW - 1 downto 0)) is
begin
load_addr <= a; load <= '1';
wait until falling_edge(clk); load <= '0';
wait until falling_edge(clk);
end procedure;
procedure do_step is
begin
step <= '1';
wait until falling_edge(clk); step <= '0';
wait until falling_edge(clk);
end procedure;
begin
for i in 0 to 2 loop
wait until falling_edge(clk);
end loop;
rst_n <= '1';
wait until falling_edge(clk);
do_load(x"0140");
chk_v("loaded", addr, x"0140");
do_step; chk_v("step 1", addr, x"0141");
do_step; chk_v("step 2", addr, x"0142");
chk_n("no wrap yet", wraps, 0);
do_load(x"01FE");
do_step; chk_v("to 0x01FF", addr, x"01FF");
chk_n("still no wrap", wraps, 0);
do_step;
chk_v("wrapped to page base", addr, x"0100");
chk_n("wrap reported", wraps, 1);
do_load(x"0700");
base := unsigned(addr);
base(PB - 1 downto 0) := (others => '0');
for i in 0 to PAGE - 1 loop
do_step;
if unsigned(addr(AW - 1 downto PB)) /= base(AW - 1 downto PB) then
report "FAIL page base moved at step " & integer'image(i) severity error;
errors <= errors + 1;
end if;
end loop;
chk_v("full page returns to base", addr, x"0700");
do_load(x"0800");
do_step; chk_v("aligned +1", addr, x"0801");
load_addr <= x"0ABC"; load <= '1'; step <= '1';
wait until falling_edge(clk);
load <= '0'; step <= '0';
wait until falling_edge(clk);
chk_v("load beats step", addr, x"0ABC");
wait until falling_edge(clk);
if errors = 0 then
report "PASS: the pointer advances within its page, wraps to the page base "
& "rather than crossing, and never alters the page base" severity note;
else
report "FAILED with " & integer'image(errors) & " error(s)" severity error;
end if;
halt <= true;
wait;
end process;
watchdog : process is
begin
wait for 500 us;
if not halt then
report "FAIL: watchdog timeout" severity failure;
end if;
wait;
end process;
end architecture tb;Parity
All three implement the same counter: identical ports and generics, asynchronous active-low reset to zero, load taking priority over step, an increment confined to the low PAGE_BITS, a wrap to the page base when those bits are all ones, and a single-cycle wrap strobe. The VHDL holds the address as unsigned internally and converts at the port, which is idiomatic and behaviourally identical. All three testbenches run the same scenarios including the full-page invariant loop.
7. Why a Verification Engineer Cares
The invariant is more valuable than the wrap check, and it is what belongs in an assertion:
// 1. THE invariant. The page base is latched by `load` and nothing else
// may change it. This is the property the whole design exists to hold,
// and it catches a carry bug on ANY step, not just at the boundary.
property p_page_base_fixed;
@(posedge clk) disable iff (!rst_n)
(!load) |=> $stable(addr[ADDR_W-1:PAGE_BITS]);
endproperty
a_page_base_fixed : assert property (p_page_base_fixed)
else $error("page base changed without a load -- the burst escaped its page");
// 2. The wrap strobe means what it says: low bits were all ones and are
// now all zeros.
property p_wrap_is_real;
@(posedge clk) disable iff (!rst_n)
page_wrapped |-> ($past(&addr[PAGE_BITS-1:0]) && (addr[PAGE_BITS-1:0] == '0));
endproperty
// 3. A step always moves the pointer. Catches a stalled increment, which
// would silently write every byte to the same location.
property p_step_advances;
@(posedge clk) disable iff (!rst_n)
(step && !load) |=> (addr != $past(addr));
endproperty
a_step_advances : assert property (p_step_advances)
else $error("step did not advance the pointer");Property 3 is worth calling out. A pointer that fails to increment writes every byte of the burst to the same address — the last byte wins and the rest vanish. That produces "only the last byte of my buffer was written", a symptom easily mistaken for a transfer-length problem, and no wrap-focused check would catch it.
What these prove. That the pointer's arithmetic obeys page-bounded semantics. What they do not prove is that PAGE_BITS matches the real device — a datasheet fact, and the Chapter 4.1 §7 boundary again. A model with the wrong page size is internally consistent and wrong.
Coverage must target the boundary relationship, not the addresses:
covergroup spi_page_cg @(posedge cs_rose); // sample per transaction
// Absolute addresses are a huge and mostly uninteresting space. What
// matters is where the burst sits RELATIVE to the page boundary.
cp_burst_vs_page : coverpoint burst_class {
bins entirely_inside = {INSIDE}; // the always-works case
bins ends_exactly_at = {ENDS_AT_EDGE}; // off-by-one territory
bins starts_at_base = {STARTS_AT_BASE};
bins crosses_once = {CROSSES_ONCE}; // the bug
bins crosses_multiple = {CROSSES_MANY}; // burst longer than a page
}
cp_len : coverpoint burst_len {
bins one = {1};
bins short = {[2:15]};
bins near_page = {[240:256]}; // page size and just under
bins over_page = {[257:1024]}; // guaranteed to wrap
}
x_class_len : cross cp_burst_vs_page, cp_len;
endgroupends_exactly_at is the bin that catches off-by-one errors — a burst whose last byte lands on the final address of the page is correct and must not wrap, while one byte more must. Those two adjacent cases are where a wrap condition written as >= instead of == fails, and a random address generator hits them rarely.
8. Why an FPGA or ASIC Engineer Cares
The incrementer is page-wide, not address-wide. Only PAGE_BITS bits ever change, so a 24-bit address with 256-byte pages needs an 8-bit incrementer. Writing addr <= addr + 1 and masking afterwards infers a 24-bit adder and a wider carry chain for no benefit, and on a fast interface that carry chain can matter.
If you implement the page buffer, its wrap must match the pointer's. A design where the pointer wraps at 256 but the buffer is addressed with more bits will write past the buffer — memory corruption inside the device rather than the benign in-page overwrite the protocol implies. The two must be the same modulus by construction, ideally by deriving both from one parameter.
Runtime-configurable page size costs more than it looks. Making PAGE_BITS an input means the all-ones detector becomes a comparison against a mask, and the incrementer must be sized for the largest page. Worth it only if the device genuinely supports several.
On a master, split the transfer in software. The controller cannot detect a page crossing — it has no idea what the device's page size is. The driver must know the page size, split buffers at page boundaries, and issue one transaction per page. That is the correct fix and the only one available, since the device will not report the problem.
9. Failure Signature — The End of the Buffer Overwrote the Beginning
Symptom. A firmware image is written to flash in large blocks. Verification shows most of it correct, but scattered regions differ — and on inspection, the corrupted bytes at the start of certain regions contain data that belongs near the end of the same region. Re-running gives identical corruption. Small writes during development always worked.
Why "small writes always worked" is the key evidence. It says the mechanism is length- or alignment-dependent, not a signalling fault. A mode, bit-order or framing problem would corrupt short writes just as reliably. Something only goes wrong when a write is large enough or positioned to reach some boundary.
Plausible mechanisms.
- Page wrap: the write crossed a page boundary and the tail overwrote the head of the same page. The observation that the corrupted head contains data from the tail is the signature, and it is essentially diagnostic on its own.
- A write-enable interlock not reissued per transaction (Chapter 5.1 §5) — but that causes missing writes, not displaced data.
- Insufficient delay for the programming time between transactions, so later writes are ignored while the device is busy — again missing rather than displaced.
- A driver buffer bug wrapping its own pointer.
The discriminating observation. Compute the corrupted addresses modulo the device's page size. If the corruption always begins at a page base and the displaced data came from an offset that equals the burst's overrun past the page end, page wrap is confirmed. That is arithmetic on addresses you already have — no capture, no instrument.
To separate it from a driver-side bug, check whether the transaction boundaries align with pages: a driver that splits correctly at page boundaries cannot produce this, so if the transactions are page-aligned and the corruption persists, the fault is elsewhere.
Why the investigation goes wrong. Because the data is present and wrong rather than missing, which reads as corruption and sends people to signal integrity and timing. Displaced-but-intact data is almost never an analogue problem — intact bytes in the wrong place mean an addressing fault, and that distinction is available immediately from the dump.
10. Common Misconceptions
11. Reason It Through
Work this before reading the answer.
A device has 256-byte pages. A driver writes a 512-byte buffer in one transaction starting at address
0x1000, which is page-aligned. Verification then reads back 512 bytes from0x1000and reports that the first 256 bytes are wrong and the second 256 bytes are... also wrong, but differently.What is in the device, and what does the read-back actually show?
Work out where each byte went. The burst starts at 0x1000, a page base. Bytes 0–255 fill the page 0x1000–0x10FF exactly. Byte 256 is the one that would cross — and instead of landing at 0x1100 it wraps to 0x1000. Bytes 256–511 therefore overwrite 0x1000–0x10FF a second time.
So what does the page contain? The second half of the buffer — bytes 256–511 — since they were written last and the commit takes the final state of the page buffer. The first half was written and then entirely overwritten before the page was committed.
Now the read-back. Reading 512 bytes from 0x1000 spans two pages. The first 256 bytes return the page's contents: buffer bytes 256–511, where the verifier expected 0–255. Wrong. The next 256 bytes come from page 0x1100, which this transaction never touched — so they hold whatever was there before, most likely erased 0xFF. Also wrong, and wrong in a completely different way.
That asymmetry is the giveaway. One region holds real data in the wrong place; the other holds no data at all. A corruption mechanism — noise, timing, a bad bit — would not produce a clean split at exactly a page boundary with displaced-but-intact data on one side and untouched memory on the other.
Why did the aligned start not save it? Because alignment only guarantees the burst begins at a page base. What matters is the length: any burst longer than one page wraps regardless of alignment. Page-aligning is necessary and not sufficient; the split must also be per page.
The fix. The driver must split the buffer into page-sized transactions: 0x1000 for 256 bytes, then 0x1100 for 256 bytes, each its own CS frame — and, on flash, with the device's programming time respected between them, because the second write will be ignored if the device is still busy committing the first.
The general lesson. When verification of a large write fails, compute what the device would have done with the addresses before assuming the data was corrupted in transit. Displaced-but-intact data and untouched memory are both addressing symptoms, and both are visible in the dump without touching the bus.
12. Understanding Check
13. Summary
A multi-byte write is one frame carrying an opcode, an address and as many data bytes as the master chooses to clock before releasing CS. There is no length field — the master sets the length implicitly by deciding when to deassert.
Devices that accept bursts advance an auto-incrementing pointer per byte, which is what makes bulk writing efficient: one opcode and one address for the whole range instead of per byte.
That pointer is bounded by a page, because it is a counter over the device's internal page buffer rather than over memory. The high address bits are latched for the transaction and there is no carry out of the page, so a burst reaching the page end wraps to the page base and overwrites bytes already written — silently, with no error, and only for particular start addresses and lengths.
Pages exist because flash commits a whole page in one programming operation, so the wrap is a description of the hardware rather than a quirk. Page size is a property of the silicon and is not configurable.
In RTL the design is a loadable counter whose increment touches only PAGE_BITS, which also means the incrementer is page-wide, not address-wide. The property worth asserting is the invariant — the page base never changes without a load — because it holds on every step rather than only at the boundary, and a companion check that a step always advances the pointer catches the stalled-increment bug that writes every byte to one address.
For coverage, the bins that matter describe a burst's relationship to the page edge: entirely inside, ending exactly at the edge, crossing once, crossing many times. Random addresses reach the off-by-one cases rarely.
And when a large write verifies wrong, displaced-but-intact data is an addressing symptom, not a corruption one — computable from the dump modulo the page size, without touching the bus.
14. What Comes Next
Every chapter so far has explained a transaction you already knew the shape of. Chapter 5.4 — Write Waveform Analysis reverses the exercise and closes the module: given an unlabelled capture of four signals and no documentation, reconstruct the transaction — recover the mode from the idle level and the edges, find the frame boundaries, segment the bytes, and identify which are opcode, address and data. It is the skill every previous chapter has been quietly building toward, and the one that turns a waveform on a screen into a diagnosis.
Continue learning
Related tutorials
- Related topic
Deriving Mode Behaviour from CPOL and CPHA
The four SPI modes are a two-bit truth table you can rebuild in seconds. The standard numbering, the derivation, the complete mode decoder in three HDLs, and the assertions that keep a configurable design honest.
- Related topic
Transfer Width
SPI has no native word size. What a 12-bit ADC or 16-bit codec requires of a master, why chip select rather than a bit count delimits a frame, and a width-parameterized transfer engine in three HDLs.
- Related topic
Bit Ordering — MSB-First and LSB-First
Which end of the shift register goes out first, the two multiplexers that make the order configurable in three HDLs, and why a bit-order bug is perfectly deterministic and yet invisible on certain data.
- Related topic
Command, Address, and Data Phases
How a device layers a transaction onto a raw byte stream: why the opcode decides the shape of everything after it, how a slave tracks phases with no phase marker, and the sequencer that requires in three HDLs.
