SPI · Module 10
Address Fields, Dummy Cycles, and Burst Behaviour
Why three address bytes make 16 MB a category boundary, why dummy latency is counted in cycles and never rounds to bytes, what a device does at the end of a burst, and the planner that turns those numbers into a cycle-accurate schedule.
Chapter 10.4 packed the register address into the command byte. That works while the map is small. A 128 Mbit flash has sixteen million addresses and seven spare bits, so the address becomes a phase of its own — and three new numbers appear with it.
A command table entry reads:
0x0B — Fast Read — 3 address bytes, 8 dummy clocks. Another reads0xEB — Fast Read Quad I/O — 3 address bytes, 6 dummy clocks. Why is the second number 6, and what happens if you send a byte?
Six is not a mistake and it does not round. Sending a byte where six cycles were asked for shifts every payload bit by two positions, and the data that comes back looks like data.
1. Address Width — and the Ceiling It Creates
The command table states the number of address bytes, and it follows directly from the device's capacity:
address bytes reachable addresses capacity
1 256 256 B
2 65 536 64 KB
3 16 777 216 16 MB ← the common case
4 4 294 967 296 4 GBThree bytes reaching exactly 16 MB is why 128 Mbit is a boundary in this whole product category. A 128 Mbit part is 16 MB, which is precisely what three address bytes address. Everything larger needs a fourth byte, and the industry's response was not to move everyone to four — it was to add a mode.
So parts above 128 Mbit typically support both:
- 3-byte mode, reaching the bottom 16 MB only, for compatibility with every existing controller and boot ROM.
- 4-byte mode, reaching the whole device, entered by a command or by a non-volatile configuration bit.
- Sometimes 4-byte opcodes — a parallel set of commands that take four address bytes without changing any mode.
Two consequences matter more than the table.
An address that does not fit is not rejected. The device receives whatever bytes arrive and uses them. Sending a 4-byte address to a part in 3-byte mode means the device consumes the first three as the address and the fourth as the first byte of data or the dummy phase — so a write corrupts the wrong location and a read returns the wrong data, with no error anywhere. The check must live in the controller, which is what spec_error does in §6.
Mode is persistent state. A part left in 4-byte mode by previous software, then reset without its interface being reset, greets the boot ROM with an address phase one byte longer than expected. This is a genuinely common bring-up failure and it survives a warm reset.
2. Byte Order
SPI devices send addresses most-significant byte first, essentially without exception. The table rarely says so, because the timing diagram shows it.
This is worth stating explicitly for one reason: it is the opposite of the byte order most host processors store addresses in. A little-endian 32-bit address in memory has its least significant byte first, and a driver that memcpy's the address into a transmit buffer sends it backwards. The result is an access to a wildly wrong address that is nonetheless a perfectly valid one — so the device answers, and the data is wrong in a way that looks like corruption rather than like a bug.
3. Dummy Latency — Counted in Cycles
This is the number that does the most damage when misread.
Why it exists is Chapter 4.5's subject: the device needs internal time between receiving the address and producing data, and the dummy phase is where that time is spent. What this chapter adds is how the number is specified — and it is specified in SCLK cycles, not bytes.
0x03 Read 0 dummy cycles
0x0B Fast Read 8 dummy cycles
0x3B Dual Output 8 dummy cycles
0xBB Dual I/O 4 dummy cycles
0xEB Quad I/O 6 dummy cycles ← not a byte
0x6B Quad Output 8 dummy cyclesThe values that are not multiples of eight are not exotic. They appear because the dummy count is set by the device's internal timing measured in clock periods, and there is no reason for that to land on a byte boundary. On the multi-line commands it frequently does not, because those commands move more bits per cycle and so need fewer cycles to cover the same internal delay.
What happens when you send a byte instead of six cycles:
Six dummy cycles versus a dummy byte
8 cyclesThe lower lane is the important one. It does not return garbage — it returns bytes assembled from the last two bits of one real byte and the first six of the next. Those bytes are plausible values. A driver checking for a non-zero response is satisfied, a checksum over a block fails, and the investigation starts at the memory rather than at the dummy count.
4. Burst Behaviour — What the Device Does at the End
The command table's last contribution is what happens when a burst runs past something.
read past the top of the device → wrap to address 0, or return undefined
write past a page boundary → WRAP within the page (Chapter 7.2)
burst longer than a stated limit → wrap, stall, or repeat one location
CS released mid-burst → abort (Chapter 8.7)Reads and writes differ, and the asymmetry is the point. A read burst normally runs freely across the whole device — the internal pointer just advances. A write burst wraps inside the page, because a page is the program granularity, and writing 300 bytes to a 256-byte page does not write 300 bytes: it writes the first 256 and then overwrites the first 44 of them with the remainder. Chapter 9.3 built the splitter that prevents this; the table is where you find the page size to give it.
The burst limit, where one is stated, often has no stated mechanism. A part specifying "up to 16 consecutive registers" may wrap to the block start, stop advancing and repeat one location, or enter reserved space. All three occur, the difference is invisible on the bench, and the only reliable answer is to test the boundary deliberately.
5. The Whole Shape, in Order
Four phases, and the two in the middle are optional per command. A status read has neither; a sector erase has an address and no dummy and no data; a write-enable is the opcode alone. A planner that emits a zero-length phase rather than skipping it produces a state nothing downstream expects — which is the first property §6's testbench checks.
6. Building the Latency Planner — Three HDLs
The circuit
Circuit. A phase walker that converts the three datasheet numbers into a cycle-accurate schedule.
State. The remaining length of the current phase, plus the three phase lengths captured at start.
Datapath. All three lengths are converted to cycles once, at start — address bytes multiplied by eight, dummy taken as given, data bytes multiplied by eight — so nothing downstream has to remember which datasheet numbers were bytes and which were cycles. That single conversion point is the design's main defence against the error of §3.
Control. A four-state walk with empty phases skipped, not entered for zero cycles.
Clock and reset. One tick per SCLK cycle; asynchronous active-low reset.
Enables. latency is published at start — command bits plus address bits plus dummy cycles — so a consumer knows when to begin capturing before the transfer has run.
Timing. Per-phase cycle counters are maintained as the walk proceeds, and they must sum to the total. That is the property worth asserting: a walker that loses or double-counts a single cycle produces a latency that is almost right, which is the hardest kind of error to notice.
Synthesis. Three length registers, a down-counter, four accounting counters and a small state machine.
Limitations. It plans; it does not drive. Turning the plan into pins is Chapter 10.6's job, and keeping the two separate means one driver serves every device.
// spi_latency_plan.sv
//
// Chapter 10.5 -- address width, dummy cycles and burst length, turned
// into a cycle-accurate plan.
//
// Three datasheet numbers decide when the first payload bit appears:
//
// the command length -- almost always 8 bits,
// the address width -- in BYTES, so 8 bits each,
// the dummy latency -- in CYCLES, and NOT necessarily a
// multiple of eight.
//
// That last asymmetry is the one that catches people. A part specifying
// "6 dummy cycles" does not want a dummy byte, and a driver that sends one
// shifts every payload byte by two bit positions -- the failure signature
// of Chapter 9.4 arriving from a completely different cause.
//
// This block walks the phases, counts the SCLK cycles each one occupies,
// and reports the LATENCY -- the number of cycles before the first data
// bit -- along with a per-phase accounting that must sum to the total.
//
// Empty phases are SKIPPED, not entered for zero cycles. A device with no
// address phase and no dummy phase goes straight from command to data, and
// a planner that enters a zero-length phase emits a state nothing
// downstream expects.
module spi_latency_plan #(
parameter int CNT_W = 16,
parameter int CMD_BITS = 8
) (
input logic clk, // one tick per SCLK cycle
input logic rst_n,
input logic start,
input logic [2:0] addr_bytes, // 0..4
input logic [5:0] dummy_cycles, // CYCLES, not bytes
input logic [CNT_W-1:0] data_bytes,
output logic [1:0] phase, // 0=cmd 1=addr 2=dummy 3=data
output logic busy,
output logic done,
output logic [CNT_W-1:0] latency, // cycles before the first data bit
output logic [CNT_W-1:0] total_cycles,
// The plan's own accounting. These must sum to total_cycles, which is
// the property worth asserting: a phase walker that loses or
// double-counts a cycle produces a latency that is almost right.
output logic [CNT_W-1:0] cyc_cmd,
output logic [CNT_W-1:0] cyc_addr,
output logic [CNT_W-1:0] cyc_dummy,
output logic [CNT_W-1:0] cyc_data
);
localparam logic [1:0] P_CMD = 2'd0;
localparam logic [1:0] P_ADDR = 2'd1;
localparam logic [1:0] P_DUMMY = 2'd2;
localparam logic [1:0] P_DATA = 2'd3;
logic [CNT_W-1:0] addr_len, dummy_len, data_len;
logic [CNT_W-1:0] rem;
logic [1:0] nxt_phase;
logic [CNT_W-1:0] nxt_rem;
logic nxt_last;
// Where to go when the current phase runs out, skipping every phase the
// profile gives a length of zero.
always_comb begin
nxt_phase = phase;
nxt_rem = {CNT_W{1'b0}};
nxt_last = 1'b0;
case (phase)
P_CMD: begin
if (addr_len != 0) begin nxt_phase = P_ADDR; nxt_rem = addr_len; end
else if (dummy_len != 0) begin nxt_phase = P_DUMMY; nxt_rem = dummy_len; end
else if (data_len != 0) begin nxt_phase = P_DATA; nxt_rem = data_len; end
else nxt_last = 1'b1;
end
P_ADDR: begin
if (dummy_len != 0) begin nxt_phase = P_DUMMY; nxt_rem = dummy_len; end
else if (data_len != 0) begin nxt_phase = P_DATA; nxt_rem = data_len; end
else nxt_last = 1'b1;
end
P_DUMMY: begin
if (data_len != 0) begin nxt_phase = P_DATA; nxt_rem = data_len; end
else nxt_last = 1'b1;
end
default: nxt_last = 1'b1;
endcase
end
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
phase <= P_CMD;
busy <= 1'b0;
done <= 1'b0;
addr_len <= {CNT_W{1'b0}};
dummy_len <= {CNT_W{1'b0}};
data_len <= {CNT_W{1'b0}};
rem <= {CNT_W{1'b0}};
latency <= {CNT_W{1'b0}};
total_cycles <= {CNT_W{1'b0}};
cyc_cmd <= {CNT_W{1'b0}};
cyc_addr <= {CNT_W{1'b0}};
cyc_dummy <= {CNT_W{1'b0}};
cyc_data <= {CNT_W{1'b0}};
end else begin
done <= 1'b0;
if (start && !busy) begin
// Every length is converted to CYCLES here, once, so that
// nothing downstream has to remember which datasheet
// numbers were bytes and which were cycles.
addr_len <= {{(CNT_W-6){1'b0}}, addr_bytes, 3'b000}; // x8
dummy_len <= {{(CNT_W-6){1'b0}}, dummy_cycles}; // as given
data_len <= data_bytes << 3; // x8
latency <= CNT_W'(CMD_BITS)
+ {{(CNT_W-6){1'b0}}, addr_bytes, 3'b000}
+ {{(CNT_W-6){1'b0}}, dummy_cycles};
total_cycles <= CNT_W'(CMD_BITS)
+ {{(CNT_W-6){1'b0}}, addr_bytes, 3'b000}
+ {{(CNT_W-6){1'b0}}, dummy_cycles}
+ (data_bytes << 3);
phase <= P_CMD;
rem <= CNT_W'(CMD_BITS);
busy <= 1'b1;
cyc_cmd <= {CNT_W{1'b0}};
cyc_addr <= {CNT_W{1'b0}};
cyc_dummy <= {CNT_W{1'b0}};
cyc_data <= {CNT_W{1'b0}};
end else if (busy) begin
// Account for the cycle being spent right now, before
// deciding whether it was the phase's last.
case (phase)
P_CMD: cyc_cmd <= cyc_cmd + 1'b1;
P_ADDR: cyc_addr <= cyc_addr + 1'b1;
P_DUMMY: cyc_dummy <= cyc_dummy + 1'b1;
default: cyc_data <= cyc_data + 1'b1;
endcase
if (rem == CNT_W'(1)) begin
if (nxt_last) begin
busy <= 1'b0;
done <= 1'b1;
end else begin
phase <= nxt_phase;
rem <= nxt_rem;
end
end else begin
rem <= rem - 1'b1;
end
end
end
end
endmodule// spi_latency_plan_tb.sv
//
// Every expectation here is computed from the datasheet numbers by the
// testbench's own arithmetic, not read back from the DUT. The two
// properties that matter are that the first data cycle lands exactly at
// the reported latency, and that the per-phase cycle counts sum to the
// total -- a walker that loses or double-counts a cycle produces a
// latency that is almost right, which is the hardest kind to notice.
`timescale 1ns/1ps
module spi_latency_plan_tb;
localparam int CNT_W = 16;
localparam int CMD_BITS = 8;
localparam logic [1:0] P_CMD = 2'd0;
localparam logic [1:0] P_ADDR = 2'd1;
localparam logic [1:0] P_DUMMY = 2'd2;
localparam logic [1:0] P_DATA = 2'd3;
logic clk = 1'b0;
logic rst_n = 1'b0;
always #5 clk = ~clk;
logic start = 1'b0;
logic [2:0] addr_bytes = 3'd0;
logic [5:0] dummy_cycles = 6'd0;
logic [CNT_W-1:0] data_bytes = {CNT_W{1'b0}};
logic [1:0] phase;
logic busy, done;
logic [CNT_W-1:0] latency, total_cycles;
logic [CNT_W-1:0] cyc_cmd, cyc_addr, cyc_dummy, cyc_data;
int errors = 0;
spi_latency_plan #(.CNT_W(CNT_W), .CMD_BITS(CMD_BITS)) dut (
.clk(clk), .rst_n(rst_n), .start(start),
.addr_bytes(addr_bytes), .dummy_cycles(dummy_cycles),
.data_bytes(data_bytes),
.phase(phase), .busy(busy), .done(done),
.latency(latency), .total_cycles(total_cycles),
.cyc_cmd(cyc_cmd), .cyc_addr(cyc_addr),
.cyc_dummy(cyc_dummy), .cyc_data(cyc_data)
);
task automatic run_plan(input string name,
input int ab, input int dc, input int db,
input bit want_addr, input bit want_dummy,
input bit want_data);
int n, first_data, exp_lat, exp_tot;
bit saw_addr, saw_dummy, saw_data;
begin
exp_lat = CMD_BITS + ab*8 + dc;
exp_tot = exp_lat + db*8;
@(negedge clk);
addr_bytes = 3'(ab);
dummy_cycles = 6'(dc);
data_bytes = CNT_W'(db);
start = 1'b1;
@(negedge clk);
start = 1'b0;
n = 0; first_data = -1;
saw_addr = 1'b0; saw_dummy = 1'b0; saw_data = 1'b0;
while (busy) begin
if (phase == P_ADDR) saw_addr = 1'b1;
if (phase == P_DUMMY) saw_dummy = 1'b1;
if (phase == P_DATA) begin
saw_data = 1'b1;
if (first_data < 0) first_data = n;
end
@(negedge clk);
n++;
end
// 1. The reported latency matches the arithmetic.
if (latency !== CNT_W'(exp_lat)) begin
$display(" FAIL: %s latency reported %0d, expected %0d",
name, latency, exp_lat);
errors++;
end
if (total_cycles !== CNT_W'(exp_tot)) begin
$display(" FAIL: %s total reported %0d, expected %0d",
name, total_cycles, exp_tot);
errors++;
end
// 2. The walk actually took that many cycles.
if (n != exp_tot) begin
$display(" FAIL: %s walked %0d cycles, planned %0d",
name, n, exp_tot);
errors++;
end
// 3. The first data cycle lands exactly at the latency. This is
// what the number is FOR -- a latency that is right on paper
// and wrong in the walk is worse than no latency at all.
if (want_data) begin
if (first_data != exp_lat) begin
$display(" FAIL: %s first data cycle at %0d, latency says %0d",
name, first_data, exp_lat);
errors++;
end
end else if (saw_data) begin
$display(" FAIL: %s entered the data phase with no data", name);
errors++;
end
// 4. Empty phases are skipped, not entered for zero cycles.
if (saw_addr != want_addr) begin
$display(" FAIL: %s address phase %0s", name,
saw_addr ? "entered when empty" : "skipped when needed");
errors++;
end
if (saw_dummy != want_dummy) begin
$display(" FAIL: %s dummy phase %0s", name,
saw_dummy ? "entered when empty" : "skipped when needed");
errors++;
end
// 5. CONSERVATION. The per-phase counts sum to the total.
if ((cyc_cmd + cyc_addr + cyc_dummy + cyc_data) !== CNT_W'(exp_tot)) begin
$display(" FAIL: %s phases sum to %0d, total is %0d",
name, cyc_cmd + cyc_addr + cyc_dummy + cyc_data, exp_tot);
errors++;
end
if (cyc_cmd !== CNT_W'(CMD_BITS) || cyc_addr !== CNT_W'(ab*8) ||
cyc_dummy !== CNT_W'(dc) || cyc_data !== CNT_W'(db*8)) begin
$display(" FAIL: %s phase split cmd=%0d addr=%0d dummy=%0d data=%0d",
name, cyc_cmd, cyc_addr, cyc_dummy, cyc_data);
errors++;
end
$display(" %-22s cmd=%0d addr=%0d dummy=%0d data=%0d latency=%0d total=%0d",
name, cyc_cmd, cyc_addr, cyc_dummy, cyc_data,
latency, total_cycles);
end
endtask
initial begin
repeat (3) @(negedge clk);
rst_n = 1'b1;
@(negedge clk);
// A flash fast read: 3 address bytes, 8 dummy cycles, 4 data bytes.
run_plan("fast read 3B/8d/4B", 3, 8, 4, 1'b1, 1'b1, 1'b1);
// The same part with a 6-cycle dummy latency. Six is not a byte,
// and a driver that rounds it up to eight shifts every payload bit
// by two positions.
run_plan("6 dummy cycles", 3, 6, 4, 1'b1, 1'b1, 1'b1);
// A dummy latency that is not even a whole number of nibbles.
run_plan("5 dummy cycles", 3, 5, 2, 1'b1, 1'b1, 1'b1);
// A four-byte address -- the part above 128 Mbit, where the third
// address byte stops being enough.
run_plan("4-byte address", 4, 8, 4, 1'b1, 1'b1, 1'b1);
// A status read: no address, no dummy. The address and dummy phases
// must be skipped entirely.
run_plan("status read", 0, 0, 1, 1'b0, 1'b0, 1'b1);
// A command with an address but no dummy and no data -- a sector
// erase. Only the data phase is absent.
run_plan("erase 3B/0d/0B", 3, 0, 0, 1'b1, 1'b0, 1'b0);
// A bare command: write-enable. Command phase only.
run_plan("write enable", 0, 0, 0, 1'b0, 1'b0, 1'b0);
// A read with dummy but no address -- rarer, but real on parts
// whose read command continues from an internal pointer.
run_plan("no address, 4 dummy", 0, 4, 2, 1'b0, 1'b1, 1'b1);
if (errors == 0)
$display("PASS: latency equals command plus address bits plus dummy CYCLES, the first data cycle lands exactly there, empty phases are skipped rather than entered, and the per-phase cycle counts always sum to the total");
else
$display("FAIL: %0d error(s)", errors);
$finish;
end
endmoduleThe testbench computes every expectation from the datasheet numbers itself rather than reading anything back from the planner, and then checks three separate things about each of eight profiles.
The reported latency matches the arithmetic. That is the easy one.
The first data cycle actually lands there. This is what the number is for — a latency that is right on paper and wrong in the walk is worse than no latency at all, because a consumer will trust it.
The per-phase counts sum to the total, and each equals its own phase's length. Conservation again, in the form Chapter 9.3 used for bytes and Chapter 9.2 used for frame time: when a design partitions something, the partition summing correctly is close to a complete specification of the accounting.
The eight profiles are chosen to cover the shapes that exist rather than a range of numbers: a fast read; the same part with six and then five dummy cycles; four-byte addressing; a status read with no address and no dummy; a sector erase with an address and no data; a bare write-enable; and a read with dummy but no address. Between them, every phase is present in some and absent in others.
// spi_latency_plan.v
//
// Chapter 10.5 -- address width, dummy cycles and burst length turned
// into a cycle-accurate plan, in Verilog-2001.
//
// Three datasheet numbers decide when the first payload bit appears: the
// command length (almost always 8 bits), the address width in BYTES, and
// the dummy latency in CYCLES -- which is NOT necessarily a multiple of
// eight. A part specifying "6 dummy cycles" does not want a dummy byte,
// and a driver that sends one shifts every payload byte by two bit
// positions.
//
// Empty phases are SKIPPED, not entered for zero cycles: a device with no
// address and no dummy goes straight from command to data, and a planner
// that enters a zero-length phase emits a state nothing downstream
// expects.
module spi_latency_plan #(
parameter CNT_W = 16,
parameter CMD_BITS = 8
) (
input wire clk, // one tick per SCLK cycle
input wire rst_n,
input wire start,
input wire [2:0] addr_bytes, // 0..4
input wire [5:0] dummy_cycles, // CYCLES, not bytes
input wire [CNT_W-1:0] data_bytes,
output reg [1:0] phase, // 0=cmd 1=addr 2=dummy 3=data
output reg busy,
output reg done,
output reg [CNT_W-1:0] latency, // cycles before the first data bit
output reg [CNT_W-1:0] total_cycles,
// The plan's own accounting. These must sum to total_cycles: a phase
// walker that loses or double-counts a cycle produces a latency that
// is almost right.
output reg [CNT_W-1:0] cyc_cmd,
output reg [CNT_W-1:0] cyc_addr,
output reg [CNT_W-1:0] cyc_dummy,
output reg [CNT_W-1:0] cyc_data
);
localparam [1:0] P_CMD = 2'd0;
localparam [1:0] P_ADDR = 2'd1;
localparam [1:0] P_DUMMY = 2'd2;
localparam [1:0] P_DATA = 2'd3;
reg [CNT_W-1:0] addr_len, dummy_len, data_len;
reg [CNT_W-1:0] rem;
reg [1:0] nxt_phase;
reg [CNT_W-1:0] nxt_rem;
reg nxt_last;
// Lengths of the phases as they are about to be loaded, in cycles.
wire [CNT_W-1:0] addr_cyc = {{(CNT_W-6){1'b0}}, addr_bytes, 3'b000};
wire [CNT_W-1:0] dummy_cyc = {{(CNT_W-6){1'b0}}, dummy_cycles};
wire [CNT_W-1:0] data_cyc = data_bytes << 3;
// Where to go when the current phase runs out, skipping every phase
// the profile gives a length of zero.
always @(*) begin
nxt_phase = phase;
nxt_rem = {CNT_W{1'b0}};
nxt_last = 1'b0;
case (phase)
P_CMD: begin
if (addr_len != 0) begin nxt_phase = P_ADDR; nxt_rem = addr_len; end
else if (dummy_len != 0) begin nxt_phase = P_DUMMY; nxt_rem = dummy_len; end
else if (data_len != 0) begin nxt_phase = P_DATA; nxt_rem = data_len; end
else nxt_last = 1'b1;
end
P_ADDR: begin
if (dummy_len != 0) begin nxt_phase = P_DUMMY; nxt_rem = dummy_len; end
else if (data_len != 0) begin nxt_phase = P_DATA; nxt_rem = data_len; end
else nxt_last = 1'b1;
end
P_DUMMY: begin
if (data_len != 0) begin nxt_phase = P_DATA; nxt_rem = data_len; end
else nxt_last = 1'b1;
end
default: nxt_last = 1'b1;
endcase
end
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
phase <= P_CMD;
busy <= 1'b0;
done <= 1'b0;
addr_len <= {CNT_W{1'b0}};
dummy_len <= {CNT_W{1'b0}};
data_len <= {CNT_W{1'b0}};
rem <= {CNT_W{1'b0}};
latency <= {CNT_W{1'b0}};
total_cycles <= {CNT_W{1'b0}};
cyc_cmd <= {CNT_W{1'b0}};
cyc_addr <= {CNT_W{1'b0}};
cyc_dummy <= {CNT_W{1'b0}};
cyc_data <= {CNT_W{1'b0}};
end else begin
done <= 1'b0;
if (start && !busy) begin
// Every length is converted to CYCLES here, once, so
// nothing downstream has to remember which datasheet
// numbers were bytes and which were cycles.
addr_len <= addr_cyc;
dummy_len <= dummy_cyc;
data_len <= data_cyc;
latency <= CMD_BITS + addr_cyc + dummy_cyc;
total_cycles <= CMD_BITS + addr_cyc + dummy_cyc + data_cyc;
phase <= P_CMD;
rem <= CMD_BITS;
busy <= 1'b1;
cyc_cmd <= {CNT_W{1'b0}};
cyc_addr <= {CNT_W{1'b0}};
cyc_dummy <= {CNT_W{1'b0}};
cyc_data <= {CNT_W{1'b0}};
end else if (busy) begin
// Account for the cycle being spent right now, before
// deciding whether it was the phase's last.
case (phase)
P_CMD: cyc_cmd <= cyc_cmd + 1'b1;
P_ADDR: cyc_addr <= cyc_addr + 1'b1;
P_DUMMY: cyc_dummy <= cyc_dummy + 1'b1;
default: cyc_data <= cyc_data + 1'b1;
endcase
if (rem == 1) begin
if (nxt_last) begin
busy <= 1'b0;
done <= 1'b1;
end else begin
phase <= nxt_phase;
rem <= nxt_rem;
end
end else begin
rem <= rem - 1'b1;
end
end
end
end
endmodule// spi_latency_plan_tb.v
//
// The same checks as the SystemVerilog testbench: the reported latency
// matches the arithmetic, the first data cycle lands exactly there, empty
// phases are skipped rather than entered, and the per-phase cycle counts
// always sum to the total.
`timescale 1ns/1ps
module spi_latency_plan_tb;
parameter CNT_W = 16;
parameter CMD_BITS = 8;
localparam [1:0] P_CMD = 2'd0;
localparam [1:0] P_ADDR = 2'd1;
localparam [1:0] P_DUMMY = 2'd2;
localparam [1:0] P_DATA = 2'd3;
reg clk;
reg rst_n;
reg start;
reg [2:0] addr_bytes;
reg [5:0] dummy_cycles;
reg [CNT_W-1:0] data_bytes;
wire [1:0] phase;
wire busy, done;
wire [CNT_W-1:0] latency, total_cycles;
wire [CNT_W-1:0] cyc_cmd, cyc_addr, cyc_dummy, cyc_data;
integer errors;
initial begin
clk = 1'b0; rst_n = 1'b0; start = 1'b0;
addr_bytes = 3'd0; dummy_cycles = 6'd0; data_bytes = {CNT_W{1'b0}};
errors = 0;
end
always #5 clk = ~clk;
spi_latency_plan #(.CNT_W(CNT_W), .CMD_BITS(CMD_BITS)) dut (
.clk(clk), .rst_n(rst_n), .start(start),
.addr_bytes(addr_bytes), .dummy_cycles(dummy_cycles),
.data_bytes(data_bytes),
.phase(phase), .busy(busy), .done(done),
.latency(latency), .total_cycles(total_cycles),
.cyc_cmd(cyc_cmd), .cyc_addr(cyc_addr),
.cyc_dummy(cyc_dummy), .cyc_data(cyc_data)
);
task run_plan;
input [8*24:1] name;
input integer ab;
input integer dc;
input integer db;
input want_addr;
input want_dummy;
input want_data;
integer n, first_data, exp_lat, exp_tot;
reg saw_addr, saw_dummy, saw_data;
begin
exp_lat = CMD_BITS + ab*8 + dc;
exp_tot = exp_lat + db*8;
@(negedge clk);
addr_bytes = ab[2:0];
dummy_cycles = dc[5:0];
data_bytes = db[CNT_W-1:0];
start = 1'b1;
@(negedge clk);
start = 1'b0;
n = 0; first_data = -1;
saw_addr = 1'b0; saw_dummy = 1'b0; saw_data = 1'b0;
while (busy) begin
if (phase == P_ADDR) saw_addr = 1'b1;
if (phase == P_DUMMY) saw_dummy = 1'b1;
if (phase == P_DATA) begin
saw_data = 1'b1;
if (first_data < 0) first_data = n;
end
@(negedge clk);
n = n + 1;
end
// 1. The reported latency matches the arithmetic.
if (latency !== exp_lat[CNT_W-1:0]) begin
$display(" FAIL: %0s latency reported %0d, expected %0d",
name, latency, exp_lat);
errors = errors + 1;
end
if (total_cycles !== exp_tot[CNT_W-1:0]) begin
$display(" FAIL: %0s total reported %0d, expected %0d",
name, total_cycles, exp_tot);
errors = errors + 1;
end
// 2. The walk actually took that many cycles.
if (n != exp_tot) begin
$display(" FAIL: %0s walked %0d cycles, planned %0d",
name, n, exp_tot);
errors = errors + 1;
end
// 3. The first data cycle lands exactly at the latency.
if (want_data) begin
if (first_data != exp_lat) begin
$display(" FAIL: %0s first data cycle at %0d, latency says %0d",
name, first_data, exp_lat);
errors = errors + 1;
end
end else if (saw_data) begin
$display(" FAIL: %0s entered the data phase with no data", name);
errors = errors + 1;
end
// 4. Empty phases are skipped, not entered for zero cycles.
if (saw_addr !== want_addr) begin
$display(" FAIL: %0s address phase handled wrongly", name);
errors = errors + 1;
end
if (saw_dummy !== want_dummy) begin
$display(" FAIL: %0s dummy phase handled wrongly", name);
errors = errors + 1;
end
// 5. CONSERVATION. The per-phase counts sum to the total.
if ((cyc_cmd + cyc_addr + cyc_dummy + cyc_data) !== exp_tot[CNT_W-1:0]) begin
$display(" FAIL: %0s phases sum to %0d, total is %0d",
name, cyc_cmd + cyc_addr + cyc_dummy + cyc_data, exp_tot);
errors = errors + 1;
end
if (cyc_cmd !== CMD_BITS || cyc_addr !== (ab*8) ||
cyc_dummy !== dc || cyc_data !== (db*8)) begin
$display(" FAIL: %0s phase split cmd=%0d addr=%0d dummy=%0d data=%0d",
name, cyc_cmd, cyc_addr, cyc_dummy, cyc_data);
errors = errors + 1;
end
$display(" %0s cmd=%0d addr=%0d dummy=%0d data=%0d latency=%0d total=%0d",
name, cyc_cmd, cyc_addr, cyc_dummy, cyc_data,
latency, total_cycles);
end
endtask
initial begin
repeat (3) @(negedge clk);
rst_n = 1'b1;
@(negedge clk);
// A flash fast read: 3 address bytes, 8 dummy cycles, 4 data bytes.
run_plan("fast read 3B/8d/4B ", 3, 8, 4, 1'b1, 1'b1, 1'b1);
// The same part with a 6-cycle dummy latency. Six is not a byte.
run_plan("6 dummy cycles ", 3, 6, 4, 1'b1, 1'b1, 1'b1);
// A dummy latency that is not even a whole number of nibbles.
run_plan("5 dummy cycles ", 3, 5, 2, 1'b1, 1'b1, 1'b1);
// A four-byte address -- the part above 128 Mbit.
run_plan("4-byte address ", 4, 8, 4, 1'b1, 1'b1, 1'b1);
// A status read: no address, no dummy.
run_plan("status read ", 0, 0, 1, 1'b0, 1'b0, 1'b1);
// A sector erase: address but no dummy and no data.
run_plan("erase 3B/0d/0B ", 3, 0, 0, 1'b1, 1'b0, 1'b0);
// A bare command: write-enable.
run_plan("write enable ", 0, 0, 0, 1'b0, 1'b0, 1'b0);
// A read with dummy but no address.
run_plan("no address, 4 dummy", 0, 4, 2, 1'b0, 1'b1, 1'b1);
if (errors == 0)
$display("PASS: latency equals command plus address bits plus dummy CYCLES, the first data cycle lands exactly there, empty phases are skipped rather than entered, and the per-phase cycle counts always sum to the total");
else
$display("FAIL: %0d error(s)", errors);
$finish;
end
endmodule-- spi_latency_plan.vhd
--
-- Chapter 10.5 -- address width, dummy cycles and burst length turned
-- into a cycle-accurate plan, in VHDL.
--
-- Three datasheet numbers decide when the first payload bit appears: the
-- command length (almost always 8 bits), the address width in BYTES, and
-- the dummy latency in CYCLES -- which is NOT necessarily a multiple of
-- eight. A part specifying "6 dummy cycles" does not want a dummy byte,
-- and a driver that sends one shifts every payload byte by two bit
-- positions.
--
-- Empty phases are SKIPPED, not entered for zero cycles.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_latency_plan is
generic (
CNT_W : positive := 16;
CMD_BITS : natural := 8
);
port (
clk : in std_logic; -- one tick per SCLK cycle
rst_n : in std_logic;
start : in std_logic;
addr_bytes : in unsigned(2 downto 0); -- 0..4
dummy_cycles : in unsigned(5 downto 0); -- CYCLES, not bytes
data_bytes : in unsigned(CNT_W - 1 downto 0);
phase : out unsigned(1 downto 0); -- 0=cmd 1=addr 2=dummy 3=data
busy : out std_logic;
done : out std_logic;
latency : out unsigned(CNT_W - 1 downto 0);
total_cycles : out unsigned(CNT_W - 1 downto 0);
-- The plan's own accounting. These must sum to total_cycles.
cyc_cmd : out unsigned(CNT_W - 1 downto 0);
cyc_addr : out unsigned(CNT_W - 1 downto 0);
cyc_dummy : out unsigned(CNT_W - 1 downto 0);
cyc_data : out unsigned(CNT_W - 1 downto 0)
);
end entity;
architecture rtl of spi_latency_plan is
constant P_CMD : unsigned(1 downto 0) := "00";
constant P_ADDR : unsigned(1 downto 0) := "01";
constant P_DUMMY : unsigned(1 downto 0) := "10";
constant P_DATA : unsigned(1 downto 0) := "11";
signal phase_r : unsigned(1 downto 0) := P_CMD;
signal busy_r : std_logic := '0';
signal done_r : std_logic := '0';
signal addr_len : unsigned(CNT_W - 1 downto 0) := (others => '0');
signal dummy_len : unsigned(CNT_W - 1 downto 0) := (others => '0');
signal data_len : unsigned(CNT_W - 1 downto 0) := (others => '0');
signal rem_cnt : unsigned(CNT_W - 1 downto 0) := (others => '0');
signal lat_r : unsigned(CNT_W - 1 downto 0) := (others => '0');
signal tot_r : unsigned(CNT_W - 1 downto 0) := (others => '0');
signal c_cmd : unsigned(CNT_W - 1 downto 0) := (others => '0');
signal c_addr : unsigned(CNT_W - 1 downto 0) := (others => '0');
signal c_dummy : unsigned(CNT_W - 1 downto 0) := (others => '0');
signal c_data : unsigned(CNT_W - 1 downto 0) := (others => '0');
signal nxt_phase : unsigned(1 downto 0) := P_CMD;
signal nxt_rem : unsigned(CNT_W - 1 downto 0) := (others => '0');
signal nxt_last : std_logic := '0';
-- Lengths of the phases as they are about to be loaded, in cycles.
signal addr_cyc : unsigned(CNT_W - 1 downto 0);
signal dummy_cyc : unsigned(CNT_W - 1 downto 0);
signal data_cyc : unsigned(CNT_W - 1 downto 0);
begin
addr_cyc <= resize(addr_bytes, CNT_W - 3) & "000";
dummy_cyc <= resize(dummy_cycles, CNT_W);
data_cyc <= shift_left(data_bytes, 3);
-- Where to go when the current phase runs out, skipping every phase
-- the profile gives a length of zero.
nxt : process (phase_r, addr_len, dummy_len, data_len)
begin
nxt_phase <= phase_r;
nxt_rem <= (others => '0');
nxt_last <= '0';
case phase_r is
when P_CMD =>
if addr_len /= 0 then
nxt_phase <= P_ADDR; nxt_rem <= addr_len;
elsif dummy_len /= 0 then
nxt_phase <= P_DUMMY; nxt_rem <= dummy_len;
elsif data_len /= 0 then
nxt_phase <= P_DATA; nxt_rem <= data_len;
else
nxt_last <= '1';
end if;
when P_ADDR =>
if dummy_len /= 0 then
nxt_phase <= P_DUMMY; nxt_rem <= dummy_len;
elsif data_len /= 0 then
nxt_phase <= P_DATA; nxt_rem <= data_len;
else
nxt_last <= '1';
end if;
when P_DUMMY =>
if data_len /= 0 then
nxt_phase <= P_DATA; nxt_rem <= data_len;
else
nxt_last <= '1';
end if;
when others =>
nxt_last <= '1';
end case;
end process;
walk : process (clk, rst_n)
begin
if rst_n = '0' then
phase_r <= P_CMD;
busy_r <= '0';
done_r <= '0';
addr_len <= (others => '0');
dummy_len <= (others => '0');
data_len <= (others => '0');
rem_cnt <= (others => '0');
lat_r <= (others => '0');
tot_r <= (others => '0');
c_cmd <= (others => '0');
c_addr <= (others => '0');
c_dummy <= (others => '0');
c_data <= (others => '0');
elsif rising_edge(clk) then
done_r <= '0';
if start = '1' and busy_r = '0' then
-- Every length is converted to CYCLES here, once, so
-- nothing downstream has to remember which datasheet
-- numbers were bytes and which were cycles.
addr_len <= addr_cyc;
dummy_len <= dummy_cyc;
data_len <= data_cyc;
lat_r <= to_unsigned(CMD_BITS, CNT_W) + addr_cyc + dummy_cyc;
tot_r <= to_unsigned(CMD_BITS, CNT_W) + addr_cyc + dummy_cyc
+ data_cyc;
phase_r <= P_CMD;
rem_cnt <= to_unsigned(CMD_BITS, CNT_W);
busy_r <= '1';
c_cmd <= (others => '0');
c_addr <= (others => '0');
c_dummy <= (others => '0');
c_data <= (others => '0');
elsif busy_r = '1' then
-- Account for the cycle being spent right now, before
-- deciding whether it was the phase's last.
case phase_r is
when P_CMD => c_cmd <= c_cmd + 1;
when P_ADDR => c_addr <= c_addr + 1;
when P_DUMMY => c_dummy <= c_dummy + 1;
when others => c_data <= c_data + 1;
end case;
if rem_cnt = 1 then
if nxt_last = '1' then
busy_r <= '0';
done_r <= '1';
else
phase_r <= nxt_phase;
rem_cnt <= nxt_rem;
end if;
else
rem_cnt <= rem_cnt - 1;
end if;
end if;
end if;
end process;
phase <= phase_r;
busy <= busy_r;
done <= done_r;
latency <= lat_r;
total_cycles <= tot_r;
cyc_cmd <= c_cmd;
cyc_addr <= c_addr;
cyc_dummy <= c_dummy;
cyc_data <= c_data;
end architecture;-- spi_latency_plan_tb.vhd
--
-- The same checks as the SystemVerilog and Verilog testbenches: the
-- reported latency matches the arithmetic, the first data cycle lands
-- exactly there, empty phases are skipped rather than entered, and the
-- per-phase cycle counts always sum to the total.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_latency_plan_tb is
end entity;
architecture sim of spi_latency_plan_tb is
constant CNT_W : positive := 16;
constant CMD_BITS : natural := 8;
constant P_ADDR : unsigned(1 downto 0) := "01";
constant P_DUMMY : unsigned(1 downto 0) := "10";
constant P_DATA : unsigned(1 downto 0) := "11";
signal clk : std_logic := '0';
signal rst_n : std_logic := '0';
signal halt : boolean := false;
signal start : std_logic := '0';
signal addr_bytes : unsigned(2 downto 0) := (others => '0');
signal dummy_cycles : unsigned(5 downto 0) := (others => '0');
signal data_bytes : unsigned(CNT_W - 1 downto 0) := (others => '0');
signal phase : unsigned(1 downto 0);
signal busy, done : std_logic;
signal latency : unsigned(CNT_W - 1 downto 0);
signal total_cycles : unsigned(CNT_W - 1 downto 0);
signal cyc_cmd : unsigned(CNT_W - 1 downto 0);
signal cyc_addr : unsigned(CNT_W - 1 downto 0);
signal cyc_dummy : unsigned(CNT_W - 1 downto 0);
signal cyc_data : unsigned(CNT_W - 1 downto 0);
signal errors : natural := 0;
begin
clk <= not clk after 5 ns when not halt else '0';
dut : entity work.spi_latency_plan
generic map (CNT_W => CNT_W, CMD_BITS => CMD_BITS)
port map (
clk => clk, rst_n => rst_n, start => start,
addr_bytes => addr_bytes, dummy_cycles => dummy_cycles,
data_bytes => data_bytes,
phase => phase, busy => busy, done => done,
latency => latency, total_cycles => total_cycles,
cyc_cmd => cyc_cmd, cyc_addr => cyc_addr,
cyc_dummy => cyc_dummy, cyc_data => cyc_data
);
stim : process
variable errs : natural := 0;
procedure run_plan(name : string; ab : natural; dc : natural;
db : natural; want_addr : boolean;
want_dummy : boolean; want_data : boolean) is
variable n : natural;
variable first_data : integer;
variable exp_lat : natural;
variable exp_tot : natural;
variable saw_addr : boolean;
variable saw_dummy : boolean;
variable saw_data : boolean;
begin
exp_lat := CMD_BITS + ab * 8 + dc;
exp_tot := exp_lat + db * 8;
wait until falling_edge(clk);
addr_bytes <= to_unsigned(ab, 3);
dummy_cycles <= to_unsigned(dc, 6);
data_bytes <= to_unsigned(db, CNT_W);
start <= '1';
wait until falling_edge(clk);
start <= '0';
n := 0; first_data := -1;
saw_addr := false; saw_dummy := false; saw_data := false;
while busy = '1' loop
if phase = P_ADDR then saw_addr := true; end if;
if phase = P_DUMMY then saw_dummy := true; end if;
if phase = P_DATA then
saw_data := true;
if first_data < 0 then first_data := n; end if;
end if;
wait until falling_edge(clk);
n := n + 1;
end loop;
-- 1. The reported latency matches the arithmetic.
if to_integer(latency) /= exp_lat then
report " FAIL: " & name & " latency reported " &
integer'image(to_integer(latency)) & ", expected " &
integer'image(exp_lat);
errs := errs + 1;
end if;
if to_integer(total_cycles) /= exp_tot then
report " FAIL: " & name & " total reported " &
integer'image(to_integer(total_cycles)) &
", expected " & integer'image(exp_tot);
errs := errs + 1;
end if;
-- 2. The walk actually took that many cycles.
if n /= exp_tot then
report " FAIL: " & name & " walked " & integer'image(n) &
" cycles, planned " & integer'image(exp_tot);
errs := errs + 1;
end if;
-- 3. The first data cycle lands exactly at the latency.
if want_data then
if first_data /= exp_lat then
report " FAIL: " & name & " first data cycle at " &
integer'image(first_data) & ", latency says " &
integer'image(exp_lat);
errs := errs + 1;
end if;
elsif saw_data then
report " FAIL: " & name & " entered the data phase with no data";
errs := errs + 1;
end if;
-- 4. Empty phases are skipped, not entered for zero cycles.
if saw_addr /= want_addr then
report " FAIL: " & name & " address phase handled wrongly";
errs := errs + 1;
end if;
if saw_dummy /= want_dummy then
report " FAIL: " & name & " dummy phase handled wrongly";
errs := errs + 1;
end if;
-- 5. CONSERVATION. The per-phase counts sum to the total.
if to_integer(cyc_cmd) + to_integer(cyc_addr) +
to_integer(cyc_dummy) + to_integer(cyc_data) /= exp_tot then
report " FAIL: " & name & " phases do not sum to the total";
errs := errs + 1;
end if;
if to_integer(cyc_cmd) /= CMD_BITS or
to_integer(cyc_addr) /= ab * 8 or
to_integer(cyc_dummy) /= dc or
to_integer(cyc_data) /= db * 8 then
report " FAIL: " & name & " phase split is wrong";
errs := errs + 1;
end if;
report " " & name & " cmd=" &
integer'image(to_integer(cyc_cmd)) & " addr=" &
integer'image(to_integer(cyc_addr)) & " dummy=" &
integer'image(to_integer(cyc_dummy)) & " data=" &
integer'image(to_integer(cyc_data)) & " latency=" &
integer'image(to_integer(latency)) & " total=" &
integer'image(to_integer(total_cycles));
end procedure;
begin
for i in 0 to 2 loop
wait until falling_edge(clk);
end loop;
rst_n <= '1';
wait until falling_edge(clk);
-- A flash fast read: 3 address bytes, 8 dummy cycles, 4 data bytes.
run_plan("fast read 3B/8d/4B ", 3, 8, 4, true, true, true);
-- The same part with a 6-cycle dummy latency. Six is not a byte.
run_plan("6 dummy cycles ", 3, 6, 4, true, true, true);
-- A dummy latency that is not even a whole number of nibbles.
run_plan("5 dummy cycles ", 3, 5, 2, true, true, true);
-- A four-byte address -- the part above 128 Mbit.
run_plan("4-byte address ", 4, 8, 4, true, true, true);
-- A status read: no address, no dummy.
run_plan("status read ", 0, 0, 1, false, false, true);
-- A sector erase: address but no dummy and no data.
run_plan("erase 3B/0d/0B ", 3, 0, 0, true, false, false);
-- A bare command: write-enable.
run_plan("write enable ", 0, 0, 0, false, false, false);
-- A read with dummy but no address.
run_plan("no address, 4 dummy", 0, 4, 2, false, true, true);
errors <= errs;
if errs = 0 then
report "PASS: latency equals command plus address bits plus dummy CYCLES, the first data cycle lands exactly there, empty phases are skipped rather than entered, and the per-phase cycle counts always sum to the total";
else
report "FAIL: " & integer'image(errs) & " error(s)" severity error;
end if;
halt <= true;
wait;
end process;
end architecture;Parity
All three implement the same planner: identical ports and generics, a single conversion of every length to cycles at start, empty phases skipped rather than entered, a published latency and total, and per-phase accounting that sums to the total. All three testbenches run the same eight profiles and report identical numbers — a latency of 40 for the standard fast read, 38 with six dummy cycles, 37 with five, 48 with four address bytes, and 8 for a bare command.
7. Why a Verification Engineer Cares
// 1. CONSERVATION. The per-phase cycle counts sum to the total. A
// walker that loses or double-counts one cycle yields a latency
// that is ALMOST right, which no spot check reliably catches.
a_conservation : assert property (
@(posedge clk) disable iff (!rst_n)
(done) |-> ((cyc_cmd + cyc_addr + cyc_dummy + cyc_data)
== total_cycles))
else $error("the phase cycle counts do not sum to the total");
// 2. THE LATENCY IS HONOURED. The first cycle spent in the data phase
// is the cycle the published latency named. This is what the number
// exists for, and a consumer will trust it.
a_latency_honoured : assert property (
@(posedge clk) disable iff (!rst_n)
($rose(phase == P_DATA)) |-> (elapsed == latency))
else $error("the data phase did not begin at the published latency");
// 3. NO EMPTY PHASES. A phase with zero length is skipped, never
// entered -- a zero-cycle state is one nothing downstream expects.
a_no_empty_phase : assert property (
@(posedge clk) disable iff (!rst_n)
((phase == P_DUMMY) |-> (dummy_len != 0)) and
((phase == P_ADDR) |-> (addr_len != 0)))
else $error("a zero-length phase was entered");
// 4. TERMINATION. Every started plan completes. A phase walker's
// natural failure is a stall rather than a wrong answer, and only a
// liveness property catches it.
a_terminates : assert property (
@(posedge clk) disable iff (!rst_n)
(start && !busy) |-> ##[1:$] done)
else $error("the plan never completed");
// 5. DUMMY IS NOT ROUNDED. The cycles spent in the dummy phase equal
// the number the profile gave -- exactly, including odd values.
a_dummy_exact : assert property (
@(posedge clk) disable iff (!rst_n)
(done) |-> (cyc_dummy == $past(dummy_cycles_at_start)))
else $error("the dummy phase was not exactly the specified length");Property 5 exists because rounding is such a natural thing to do. A designer who thinks in bytes will write the dummy phase as a byte count without noticing, and every profile whose dummy count happens to be a multiple of eight will pass. The assertion makes the odd values part of the specification rather than part of the test data.
Property 4 is the liveness property, and it is the one that catches the specific failure mode of a phase walker: not a wrong answer but a stall, produced by a next-phase decision that skips into nothing.
Coverage must cross presence with length, because the absent cases are the interesting ones:
covergroup spi_phase_plan_cg @(posedge clk iff start);
cp_abytes : coverpoint addr_bytes {
bins none = {0}; // status reads -- phase must be SKIPPED
bins one = {1};
bins three = {3}; // the common case
bins four = {4}; // 4-byte mode
}
// The values that are not multiples of eight are the point of this
// whole chapter, and a suite using only 0 and 8 never tests them.
cp_dummy : coverpoint dummy_cycles {
bins none = {0};
bins odd_small = {[1:7]}; // 4, 5, 6 -- real values
bins one_byte = {8};
bins odd_large = {[9:15]}; // 10, 12, 14 -- also real
bins large = {[16:32]};
}
cp_data : coverpoint data_bytes {
bins none = {0}; // erase, write-enable
bins one = {1};
bins many = {[2:$]};
}
// The shape of the transfer, which is what the walker actually
// implements. Eight combinations, and a suite that only ever sends
// full four-phase transfers has tested one of them.
x_shape : cross cp_abytes, cp_dummy, cp_data;
endgroupcp_dummy's odd_small and odd_large bins are the ones to insist on. A suite whose dummy counts are all 0 or 8 will pass a planner that quietly works in bytes.
8. Why an FPGA or ASIC Engineer Cares
Convert to cycles once, at the boundary. The byte-versus-cycle confusion is a units error, and units errors are prevented the same way everywhere: convert on entry and keep one unit internally. A design that carries "dummy" around as a byte count in one module and a cycle count in another will eventually mix them.
Make the dummy count a profile field with odd values reachable. A field wide enough only for multiples of eight is a design that has decided the datasheet is wrong.
Check the address against the address width in hardware. The device cannot detect an over-wide address — it simply consumes the bytes it receives — so a 4-byte address sent to a part in 3-byte mode shifts the entire rest of the transfer. Three gates and a comparator prevent it.
Track 3-byte/4-byte mode as state, and re-establish it after reset. It is persistent, it survives a warm reset, and previous software can leave it set. A controller that issues an explicit mode command at initialisation rather than assuming the default removes a bring-up failure that is otherwise very hard to see.
Publish the latency, do not recompute it. Every consumer that needs to know when data begins should read the same number from the same place. Two modules computing it independently is how a design comes to disagree with itself about a value that is right in both.
9. Failure Signature — A Flash That Reads Correctly Below 16 MB
Symptom. A 256 Mbit flash works perfectly. Every read verifies, the file system is stable, and the product ships. Some months later a build grows past a certain size and reads from the upper half of the device return data that belongs to the lower half — correct-looking data from entirely the wrong place. Below the halfway mark everything is still perfect.
What "correct data from the wrong place" establishes. The link is fine. The mode, the rate, the dummy count and the command encoding are all right, because the data is intact — only its address is wrong. So this is an addressing fault, not a timing or protocol fault.
Plausible mechanisms.
- The device is in 3-byte mode and the controller sends three address bytes. A 256 Mbit part is 32 MB and needs 25 address bits; three bytes give 24, so the top bit is simply absent and every address above 16 MB aliases to the same address 16 MB lower.
- The controller sends four bytes while the device is in 3-byte mode, which would shift the whole transfer and corrupt the data — so this does not fit intact data.
- A driver truncating the address to a 24-bit type somewhere in its call chain, which produces the identical aliasing.
- A file system or wear-levelling layer with its own 24-bit assumption.
- The part genuinely being 128 Mbit and mislabelled, which fits the symptom exactly and is worth eliminating with an ID read.
The discriminating observation. Aliasing has a precise signature: address X and address X + 16 MB return identical data. Test that directly. Write a known pattern at 0x000100, read 0x1000100, and if the pattern appears there the top address bit is being lost.
Then determine where it is lost. Read the device's configuration register to see whether it reports 3-byte or 4-byte mode. If the device says 4-byte and the aliasing persists, the truncation is in the controller or the driver; if the device says 3-byte, the device never received the fourth byte.
The fix. Enter 4-byte mode explicitly at initialisation — or use the 4-byte opcodes, which need no mode at all and are the more robust choice precisely because they carry no persistent state. Then widen whatever type truncated the address.
Why this ships. Because the entire lower 16 MB works perfectly, and that is all anybody tested. Address-space bugs are invisible until the address space is used, and a product whose image fits in the bottom half will pass every test for years. The lesson is to test the top of any address space deliberately, at bring-up, when the fix is a configuration line rather than a field update.
10. Common Misconceptions
11. Reason It Through
Work this before reading the answer.
A driver reads 256 bytes from a flash at address 0x7FFF80 using
0x0Bwith three address bytes and eight dummy cycles, at 20 MHz. The first 128 bytes verify. The last 128 are wrong — and, oddly, they match the contents of address 0x000000 onwards.The part is 8 MB. What happened, and would a faster or slower clock change anything?
Start with the arithmetic. The part is 8 MB, so its top address is 0x7FFFFF. The read starts at 0x7FFF80, which is 128 bytes below the top.
0x7FFF80 + 128 = 0x800000So the read runs off the end of the device exactly halfway through.
What does the device do? Per §4, a read burst that passes the top of the device either wraps to address 0 or returns undefined data — and this one wraps. That is why the last 128 bytes match address 0x000000 onwards. The device is behaving exactly as its datasheet says.
So nothing is broken. The link is correct, the dummy count is correct, the address is correct, and the device did what it documents. The driver issued a read that crosses the end of the address space, and nothing in the protocol prevents that.
Would changing the clock help? No — and this is the part worth being sure about. The failure is an addressing behaviour, not a timing one. Per the warning in §3, rate dependence is what distinguishes a timing fault from a logical one: this one would reproduce identically at 1 MHz and at 50 MHz, because the device wraps at the same address either way.
What is the fix? Bound the read at the device's top address, exactly as Chapter 9.3's splitter bounds a burst at a page boundary — the mechanism is the same three-way minimum, with the distance to the end of the device as one more limit. A read that would cross the end is split into one read that stops at the top and, if the caller really wanted the wrap, a second read from zero.
And the deeper reading. The end-of-device boundary behaves like a page boundary for reads, even though pages are usually discussed only for writes. Any burst has several limits — the page, the device's end, the buffer, the deadline — and a splitter that knows about only one of them works until the first transfer that meets another. When you build the splitter of Chapter 9.3, the end of the device belongs in its minimum.
12. Understanding Check
13. Summary
Once a device has more addresses than a command byte has spare bits, the address becomes its own phase, and three numbers come with it.
Address width is stated in bytes and creates the ceiling: three bytes reach exactly 16 MB, which is why 128 Mbit is a category boundary and why larger parts carry a persistent 3-byte/4-byte mode that survives a warm reset and that previous software can leave set. An address that does not fit is never rejected — the device consumes the bytes it receives and everything after shifts.
Byte order is most-significant first, which is the opposite of how a little-endian host stores it — so a copied address reaches a valid but wildly wrong location.
Dummy latency is counted in CYCLES, and 4, 5, 6 and 10 are ordinary values. Sending a byte instead shifts the payload, and the result is plausible bytes spliced from adjacent real ones rather than obvious garbage. The signature matches a round-trip violation exactly, and rate dependence is what separates them.
Burst behaviour differs by direction: reads usually run freely and wrap at the end of the device; writes wrap inside the page, overwriting what was just written.
In hardware, convert every length to cycles once, at the boundary, skip empty phases rather than entering them, publish the latency so no two modules compute it separately, and assert that the per-phase counts sum to the total — because a walker that loses one cycle produces a latency that is almost right.
14. What Comes Next
Every number the route asked for is now in hand: the mode, the rate, the timing, the command encoding, the address width, the dummy count and the burst rules.
Chapter 10.6 — From Datasheet to Transaction Specification assembles them. It works a complete device end to end, turns the result into the exact element sequence a controller must issue, and builds the specification engine that produces it from the profile — so that supporting a new part becomes a table entry rather than a driver. It is also where the full UVM environment for a datasheet-derived device comes together, in all three HDLs.
Continue learning
Related tutorials
- Related topic
Dummy Phases and Read Latency
Why a device needs turnaround before it can answer, why dummy is counted in clock cycles rather than bytes, how its length grows with frequency, and the one-byte data offset a mismatch produces.
- Related topic
Burst Efficiency and Transfer Sizing
How burst length amortises overhead and where the returns flatten: the knee of the efficiency curve, the four constraints that really limit burst length, and the splitter that divides a transfer at page boundaries and caps.
- Related topic
Normal Read, Fast Read, and Dummy Cycles
Why a flash offers two read commands returning the same data, why the faster one must buy array access time in dummy cycles, why fast read is not always faster, and the selector that chooses by transfer time rather than clock rate.
- Related topic
Chip-Select Generation
Chip select is a state machine, not a wire: the three ways deriving it from a busy signal fails, why the between-frames pause and the between-transactions pause are opposites, and why a select for a slave that is not fitted must be refused.
