SPI · Module 12
Why Wider SPI Exists
The bandwidth pressure that pushed flash past one data line, why four data lanes deliver far less than four times the speed, why the dummy phase becomes more prominent as the bus widens, and the model that computes the gain for any width combination.
Every ceiling this track has met is a limit on how fast one line can be clocked: the round trip, the array access time, the board's loading. This module takes the other direction.
When frequency cannot rise, width can. So why does a device with four data lanes instead of one deliver about three times the throughput rather than four — and on a short transfer, barely one and a half?
Because widening the bus widens the data phase and leaves the rest alone. What is left over is a fixed cost, and fixed costs behave in a way that has a name.
1. The Pressure That Caused It
Single-lane SPI moves one bit per clock edge. A 100 MHz link therefore delivers 12.5 MB/s at best, and Chapter 9.4 showed that 100 MHz is itself optimistic on a real board — the round trip often caps a link near 40 MHz, or 5 MB/s.
Meanwhile the thing at the other end grew. A microcontroller booting a 256 KB image from a 5 MB/s link spends 50 ms doing it. An FPGA loading a 4 MB bitstream spends most of a second. And Chapter 12.5's execute-in-place, where a processor fetches instructions directly from flash, needs bandwidth comparable to RAM rather than to a peripheral bus.
Three routes were available and only one was viable.
Raise the clock. Already at its limit for the reasons Module 9 gave, and every increase costs signal-integrity margin on a bus that must still work with several devices on it.
Use a parallel bus. A parallel flash interface with 16 data pins works and costs 16 pins, which is exactly what SPI was chosen to avoid (Chapter 1.1).
Reuse the pins already there. A standard serial flash package has pins that are only used in specific circumstances — a write-protect and a hold pin. Repurposing them as data lanes gives four bidirectional data pins in the same package, with no extra pins and no new connector. That is what won, and its cost is that those pins become bidirectional, which is Chapter 12.3's subject.
2. What Widens, and What Does Not
A transaction has four phases (Chapter 10.5), and widening the bus does not affect them equally:
command 8 bits may widen, if the device supports it
address 8N bits may widen, if the device supports it
dummy N CYCLES CANNOT widen -- it is not measured in bits
data 8L bits widens, and this is the pointThe dummy phase is the one that matters and the reason is worth stating precisely: it is specified in clock cycles, not in bits. Widening the bus moves more bits per cycle, so it shortens anything measured in bits — and leaves anything measured in cycles exactly as it was.
So the cycle count is:
cycles = 8 / cmd_lanes
+ 8 × addr_bytes / addr_lanes
+ dummy ← unchanged by width
+ 8 × length / data_lanesAnd the notation for which phases widen is the x-y-z form Chapter 12.3 covers in full:
1-1-1 single-lane SPI
1-1-4 quad OUTPUT — only the data phase widens
1-4-4 quad I/O — address and data widen
4-4-4 full quad — everything but the dummy phase
8-8-8 octal3. The Numbers
Take a fast read of 64 bytes with three address bytes and eight dummy cycles, and count cycles at each width:
width fixed data total vs 1-1-1
1-1-1 40 512 552 1.00×
1-1-4 40 128 168 3.28×
1-4-4 22 128 150 3.68×
4-4-4 16 128 144 3.83×
8-8-8 12 64 76 7.26×Four data lanes give 3.28×, not 4×. Eight give 7.26×, not 8×. The gap is the fixed phases, and the table shows exactly where it goes: at 1-1-4 the fixed cost is 40 cycles out of 168 — a quarter of the transfer spent on phases that did not widen at all.
Widening those phases recovers most of it. Going from 1-1-4 to 4-4-4 does nothing to the data phase and still improves the total by 14%, purely by shrinking the command and address.
The dummy phase is what remains. At 8-8-8 the fixed cost is 12 cycles, of which 8 are the dummy phase — two thirds of everything width could not remove. That is why it is the term to look at when a wide link underperforms.
4. The Short-Transfer Result
Now the same widths on a four-byte read:
width fixed data total vs 1-1-1
1-1-1 40 32 72 1.00×
1-1-4 40 8 48 1.50×
4-4-4 16 8 24 3.00×
8-8-8 12 4 16 4.50×Quad output gives 1.50×. Four data lanes, half again the speed. And octal gives 4.50× rather than 8.
This is Amdahl's argument, and it arrives here in a form worth internalising: the speedup from widening the data phase is bounded by the fraction of the transfer that was data. On a 64-byte read the data phase was 93% of a single-lane transfer, so there was a great deal to win. On a four-byte read it was 44%, so there was not.
Two consequences follow directly, and both shape the rest of the module.
Widening the command and address phases matters more than it looks. At four bytes, 1-1-4 gives 1.50× and 4-4-4 gives 3.00× — the same data width, double the throughput, entirely from the phases the notation's first two digits describe. This is why 4-4-4 exists as a distinct mode rather than a refinement.
Short accesses need the overhead removed rather than the data widened. That is precisely what continuous read and XIP do, and it is why Chapter 12.5 is the module's destination rather than an appendix.
5. Where the Cycles Go
The three dummy 8 nodes across the figure are the point. Every other row narrows somewhere; that one never does.
6. Building the Width Model — Three HDLs
The circuit
Circuit. A cycle-count calculator over a per-phase width set.
State. The computed counts, registered on a calc pulse.
Datapath. Each phase's cycle count is its bit count divided by its lane count — and because lane counts are always 1, 2, 4 or 8, every one of those divisions is a shift. A design that wrote bits / lanes would infer a divider for a quantity that is always a power of two.
Control. None; one pulse, one result.
Clock and reset. System clock; asynchronous active-low reset.
Enables. width_err reports a lane count that is not 1, 2, 4 or 8 — three lanes is not a width any device has, and treating it as one lane silently would give a plausible answer to an impossible question.
Timing. Registered, so the result is stable one cycle after calc.
Synthesis. Three shifters, an adder tree and a comparator chain. Small.
Limitations. It counts cycles, not time. Converting to time needs the divisor of Chapter 9.4, and comparing widths at different clock rates needs both — which is a real consideration, because some devices reduce their maximum rate in wide modes.
What it deliberately does not compute. The ratio. A divider costs far more than the insight, and the difference — cycles saved — is a subtract and is the more useful number for a latency budget anyway. The design publishes both cycle counts and lets software divide if it wants to.
And the split that explains the result. data_cycles and fixed_cycles are published separately, because their sum is the total and only the first shrinks with width. That decomposition is the chapter's argument expressed as two output ports.
// spi_width_model.sv
//
// Chapter 12.1 -- what widening the bus actually buys.
//
// A wider SPI moves more bits per clock on the DATA phase. It does not
// widen the phases that carry the opcode, the address or the dummy
// latency unless the device supports doing so -- and those phases are a
// fixed cost that width cannot reduce.
//
// So the speedup from N data lanes is NOT N. It is bounded by how much of
// the transfer was data in the first place, which is Amdahl's argument
// arriving in a serial bus:
//
// cycles = 8/cmd_lanes + 8*addr_bytes/addr_lanes + dummy
// + 8*len/data_lanes
//
// The dummy phase is the interesting term. It is quoted in CLOCK CYCLES
// (Chapter 10.5), so widening the bus does not shorten it at all -- and as
// the data phase shrinks, the dummy phase becomes a larger share of what
// remains. Widening the bus makes the latency you cannot remove more
// prominent, not less.
//
// This block computes the cycle count for any per-phase width combination
// and, alongside it, the same transfer entirely on one lane -- so the two
// can be compared. It deliberately does NOT compute the ratio: a divider
// costs far more than the insight, and the difference (cycles saved) is a
// subtract and is the more useful number for a latency budget anyway.
//
// DIVISION BY LANES IS A SHIFT. Lane counts are 1, 2, 4 or 8 -- always
// powers of two -- so every division here is a shift by the lane count's
// logarithm. Anything else would infer a divider for no reason.
module spi_width_model #(
parameter int CNT_W = 32,
parameter int CMD_BITS = 8
) (
input logic clk,
input logic rst_n,
input logic calc, // pulse: evaluate this shape
// Per-phase lane widths. Legal values are 1, 2, 4 and 8.
input logic [3:0] cmd_lanes,
input logic [3:0] addr_lanes,
input logic [3:0] data_lanes,
input logic [2:0] addr_bytes,
input logic [5:0] dummy_cycles, // CYCLES -- width cannot shrink these
input logic [15:0] data_bytes,
output logic [CNT_W-1:0] cycles, // this width combination
output logic [CNT_W-1:0] base_cycles, // the same transfer on one lane
output logic [CNT_W-1:0] saved_cycles, // base - cycles, a subtract
output logic [CNT_W-1:0] data_cycles, // the part width can reduce
output logic [CNT_W-1:0] fixed_cycles, // the part it cannot
output logic width_err // a lane count that is not 1/2/4/8
);
// Lane count to shift amount. A case rather than arithmetic, because
// the legal set is small and this is also where an illegal value is
// caught -- one table serving both jobs.
function automatic logic [2:0] lane_shift(input logic [3:0] lanes);
case (lanes)
4'd1: lane_shift = 3'd0;
4'd2: lane_shift = 3'd1;
4'd4: lane_shift = 3'd2;
4'd8: lane_shift = 3'd3;
default: lane_shift = 3'd0; // flagged by width_err
endcase
endfunction
function automatic logic lane_bad(input logic [3:0] lanes);
lane_bad = !((lanes == 4'd1) || (lanes == 4'd2) ||
(lanes == 4'd4) || (lanes == 4'd8));
endfunction
logic [CNT_W-1:0] cmd_bits, addr_bits, data_bits;
logic [CNT_W-1:0] c_cmd, c_addr, c_data;
always_comb begin
cmd_bits = CNT_W'(CMD_BITS);
addr_bits = CNT_W'(addr_bytes) << 3;
data_bits = CNT_W'(data_bytes) << 3;
// Each phase's cycle count is its bit count divided by its lanes,
// which is a right shift because the lane count is a power of two.
c_cmd = cmd_bits >> lane_shift(cmd_lanes);
c_addr = addr_bits >> lane_shift(addr_lanes);
c_data = data_bits >> lane_shift(data_lanes);
end
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
cycles <= {CNT_W{1'b0}};
base_cycles <= {CNT_W{1'b0}};
saved_cycles <= {CNT_W{1'b0}};
data_cycles <= {CNT_W{1'b0}};
fixed_cycles <= {CNT_W{1'b0}};
width_err <= 1'b0;
end else if (calc) begin
width_err <= lane_bad(cmd_lanes) || lane_bad(addr_lanes) ||
lane_bad(data_lanes);
cycles <= c_cmd + c_addr + CNT_W'(dummy_cycles) + c_data;
// The same transfer entirely on one lane: every phase at its
// full bit count, and the dummy phase UNCHANGED -- because it
// is cycles, not bits, and that is the whole point.
base_cycles <= cmd_bits + addr_bits + CNT_W'(dummy_cycles)
+ data_bits;
saved_cycles <= (cmd_bits + addr_bits + data_bits)
- (c_cmd + c_addr + c_data);
// The split that explains the diminishing return: only
// data_cycles shrinks with the data width, and fixed_cycles is
// the floor no amount of widening can go below.
data_cycles <= c_data;
fixed_cycles <= c_cmd + c_addr + CNT_W'(dummy_cycles);
end
end
endmodule// spi_width_model_tb.sv
//
// The testbench computes every expected cycle count from the arithmetic
// itself, then uses the model to demonstrate the result the chapter is
// about: the speedup from N data lanes is bounded by how much of the
// transfer was data, and at short lengths it is far below N.
`timescale 1ns/1ps
module spi_width_model_tb;
localparam int CNT_W = 32;
logic clk = 1'b0;
logic rst_n = 1'b0;
always #5 clk = ~clk;
logic calc = 1'b0;
logic [3:0] cmd_lanes = 4'd1;
logic [3:0] addr_lanes = 4'd1;
logic [3:0] data_lanes = 4'd1;
logic [2:0] addr_bytes = 3'd3;
logic [5:0] dummy_cycles = 6'd8;
logic [15:0] data_bytes = 16'd64;
logic [CNT_W-1:0] cycles, base_cycles, saved_cycles;
logic [CNT_W-1:0] data_cycles, fixed_cycles;
logic width_err;
int errors = 0;
int m_base, m_quad;
spi_width_model #(.CNT_W(CNT_W), .CMD_BITS(8)) dut (
.clk(clk), .rst_n(rst_n), .calc(calc),
.cmd_lanes(cmd_lanes), .addr_lanes(addr_lanes),
.data_lanes(data_lanes),
.addr_bytes(addr_bytes), .dummy_cycles(dummy_cycles),
.data_bytes(data_bytes),
.cycles(cycles), .base_cycles(base_cycles),
.saved_cycles(saved_cycles), .data_cycles(data_cycles),
.fixed_cycles(fixed_cycles), .width_err(width_err)
);
// The testbench's own model, from the arithmetic rather than the design.
function automatic int model_cycles(input int cl, input int al,
input int dl, input int ab,
input int dum, input int len);
begin
model_cycles = (8 / cl) + ((ab * 8) / al) + dum + ((len * 8) / dl);
end
endfunction
task automatic evaluate(input int cl, input int al, input int dl,
input int ab, input int dum, input int len);
begin
@(negedge clk);
cmd_lanes = 4'(cl); addr_lanes = 4'(al); data_lanes = 4'(dl);
addr_bytes = 3'(ab); dummy_cycles = 6'(dum);
data_bytes = 16'(len);
calc = 1'b1;
@(negedge clk);
calc = 1'b0;
@(negedge clk);
end
endtask
task automatic check(input string name, input int cl, input int al,
input int dl, input int len);
int want;
begin
evaluate(cl, al, dl, 3, 8, len);
want = model_cycles(cl, al, dl, 3, 8, len);
if (cycles !== CNT_W'(want)) begin
$display(" FAIL: %s len=%0d reported %0d cycles, model says %0d",
name, len, cycles, want);
errors++;
end
if (width_err) begin
$display(" FAIL: %s reported a width error on legal lanes", name);
errors++;
end
// The split must reconstruct the total -- a partition property.
if ((data_cycles + fixed_cycles) !== cycles) begin
$display(" FAIL: %s data(%0d) + fixed(%0d) != total(%0d)",
name, data_cycles, fixed_cycles, cycles);
errors++;
end
end
endtask
initial begin
repeat (3) @(negedge clk);
rst_n = 1'b1;
@(negedge clk);
// 1. The single-lane baseline: a fast read of 64 bytes.
check("1-1-1", 1, 1, 1, 64);
$display(" 1-1-1 len=64: %0d cycles (fixed %0d + data %0d)",
cycles, fixed_cycles, data_cycles);
if (cycles !== base_cycles) begin
$display(" FAIL: an all-one-lane transfer must equal its own baseline");
errors++;
end
m_base = cycles;
// 2. Quad OUTPUT -- 1-1-4. Only the data phase widens, which is
// what a device supporting 0x6B offers. Four times the data
// lanes, and nothing like four times the speed.
check("1-1-4", 1, 1, 4, 64);
m_quad = cycles;
$display(" 1-1-4 len=64: %0d cycles (fixed %0d + data %0d) -- %0d.%02dx",
cycles, fixed_cycles, data_cycles,
m_base / cycles, ((m_base * 100) / cycles) % 100);
if (cycles >= m_base) begin
$display(" FAIL: widening the data phase did not reduce the cycle count");
errors++;
end
// THE POINT: four data lanes, less than four times the speedup.
if ((m_base * 100) / cycles >= 400) begin
$display(" FAIL: 1-1-4 achieved 4x or better -- the fixed phases must cost something");
errors++;
end
// 3. Quad I/O -- 1-4-4. The address widens too, so the fixed cost
// falls and the speedup rises.
check("1-4-4", 1, 4, 4, 64);
$display(" 1-4-4 len=64: %0d cycles (fixed %0d + data %0d) -- %0d.%02dx",
cycles, fixed_cycles, data_cycles,
m_base / cycles, ((m_base * 100) / cycles) % 100);
if (cycles >= m_quad) begin
$display(" FAIL: widening the address phase did not help");
errors++;
end
// 4. Full quad -- 4-4-4. Everything but the dummy phase widens.
check("4-4-4", 4, 4, 4, 64);
$display(" 4-4-4 len=64: %0d cycles (fixed %0d + data %0d) -- %0d.%02dx",
cycles, fixed_cycles, data_cycles,
m_base / cycles, ((m_base * 100) / cycles) % 100);
// 5. Octal -- 8-8-8. Eight data lanes, and still not eight times,
// because the dummy phase is CYCLES and did not move.
check("8-8-8", 8, 8, 8, 64);
$display(" 8-8-8 len=64: %0d cycles (fixed %0d + data %0d) -- %0d.%02dx",
cycles, fixed_cycles, data_cycles,
m_base / cycles, ((m_base * 100) / cycles) % 100);
if ((m_base * 100) / cycles >= 800) begin
$display(" FAIL: 8-8-8 achieved 8x or better -- the dummy phase must cost something");
errors++;
end
// 6. THE AMDAHL RESULT. The same widths on a SHORT transfer, where
// the fixed phases dominate. Four data lanes buys very little.
check("1-1-1", 1, 1, 1, 4);
m_base = cycles;
$display(" 1-1-1 len=4: %0d cycles", cycles);
check("1-1-4", 1, 1, 4, 4);
$display(" 1-1-4 len=4: %0d cycles -- %0d.%02dx (four data lanes!)",
cycles, m_base / cycles, ((m_base * 100) / cycles) % 100);
// On four bytes, quad output must give less than 2x -- the whole
// reason 4-4-4 and XIP exist.
if ((m_base * 100) / cycles >= 200) begin
$display(" FAIL: on a 4-byte transfer quad output gave 2x or more");
errors++;
end
check("4-4-4", 4, 4, 4, 4);
$display(" 4-4-4 len=4: %0d cycles -- %0d.%02dx",
cycles, m_base / cycles, ((m_base * 100) / cycles) % 100);
// 7. THE FLOOR. With the data phase free, the fixed phases remain.
// A zero-length transfer shows exactly what width cannot remove.
check("4-4-4", 4, 4, 4, 0);
if (data_cycles !== CNT_W'(0)) begin
$display(" FAIL: a zero-length transfer reported %0d data cycles",
data_cycles);
errors++;
end
$display(" 4-4-4 len=0: %0d cycles -- the floor width cannot go below",
fixed_cycles);
// 8. The dummy phase is in CYCLES and never shrinks. Two identical
// shapes differing only in dummy count must differ by exactly
// that, at every width.
begin
int c_lo, c_hi;
evaluate(4, 4, 4, 3, 0, 64); c_lo = cycles;
evaluate(4, 4, 4, 3, 8, 64); c_hi = cycles;
if ((c_hi - c_lo) != 8) begin
$display(" FAIL: 8 dummy cycles changed the total by %0d at 4-4-4",
c_hi - c_lo);
errors++;
end
evaluate(8, 8, 8, 3, 0, 64); c_lo = cycles;
evaluate(8, 8, 8, 3, 8, 64); c_hi = cycles;
if ((c_hi - c_lo) != 8) begin
$display(" FAIL: 8 dummy cycles changed the total by %0d at 8-8-8",
c_hi - c_lo);
errors++;
end
$display(" dummy phase costs 8 cycles at 1-1-1, 4-4-4 and 8-8-8 alike");
end
// 9. An illegal lane count is reported rather than silently treated
// as one lane -- three lanes is not a width any device has.
evaluate(1, 1, 3, 3, 8, 64);
if (!width_err) begin
$display(" FAIL: three data lanes was not reported as illegal");
errors++;
end
evaluate(1, 1, 8, 3, 8, 64);
if (width_err) begin
$display(" FAIL: eight data lanes was reported as illegal");
errors++;
end
$display(" lane validation: 3 rejected, 8 accepted");
// 10. MONOTONICITY. Across a sweep of lengths and widths, a wider
// data phase never costs MORE cycles -- and the partition
// always reconstructs the total.
for (int len = 0; len <= 256; len += 8) begin
int c1, c2, c4, c8;
evaluate(1, 1, 1, 3, 8, len); c1 = cycles;
evaluate(1, 1, 2, 3, 8, len); c2 = cycles;
evaluate(1, 1, 4, 3, 8, len); c4 = cycles;
evaluate(1, 1, 8, 3, 8, len); c8 = cycles;
if (!(c8 <= c4 && c4 <= c2 && c2 <= c1)) begin
$display(" FAIL: len=%0d widths are not monotone (%0d,%0d,%0d,%0d)",
len, c1, c2, c4, c8);
errors++;
end
if ((data_cycles + fixed_cycles) !== cycles) begin
$display(" FAIL: len=%0d the partition does not reconstruct the total", len);
errors++;
end
end
$display(" 33 lengths swept: wider is never slower, and the partition always reconstructs");
if (errors == 0)
$display("PASS: the cycle count matches the arithmetic at every width, the data and fixed phases partition the total exactly, widening is never slower, the dummy phase costs the same at every width because it is measured in cycles, four data lanes give well under four times the speedup and under two times on a short transfer, and an illegal lane count is reported rather than treated as one lane");
else
$display("FAIL: %0d error(s)", errors);
$finish;
end
endmoduleThe testbench computes every expected count from the arithmetic itself, and then does something more interesting than checking values: it asserts the bounds.
1-1-4 must give less than 4× and 8-8-8 less than 8× — not approximately, but strictly. Those two assertions encode the chapter's claim as a test, so a design that somehow reported a 4× gain would fail rather than being believed.
On a four-byte transfer, quad output must give less than 2×. That is the short-transfer result made into a check, and it is the one that would catch a model that had quietly ignored the fixed phases.
Two structural properties hold across a sweep of thirty-three lengths. Monotonicity: a wider data phase is never slower, at any length. And the partition reconstructs — data_cycles + fixed_cycles always equals the total, which catches a decomposition whose two halves are each plausible and mutually inconsistent.
Finally the dummy phase is checked directly: changing only the dummy count must change the total by exactly that amount at 1-1-1, 4-4-4 and 8-8-8 alike. Three widths, the same eight cycles, which is the clearest possible statement that width and cycle-denominated latency are independent.
// spi_width_model.v
//
// Chapter 12.1 -- what widening the bus actually buys, in Verilog-2001.
//
// A wider SPI moves more bits per clock on the DATA phase. It does not
// widen the opcode, address or dummy phases unless the device supports
// doing so, and those are a fixed cost width cannot reduce:
//
// cycles = 8/cmd_lanes + 8*addr_bytes/addr_lanes + dummy
// + 8*len/data_lanes
//
// The dummy phase is the interesting term. It is quoted in CLOCK CYCLES,
// so widening the bus does not shorten it at all -- and as the data phase
// shrinks, the dummy phase becomes a larger share of what remains.
//
// This block reports the cycle count and, alongside it, the same transfer
// entirely on one lane. It deliberately does NOT compute the ratio: a
// divider costs more than the insight, and the difference is a subtract and
// the more useful number for a latency budget.
//
// DIVISION BY LANES IS A SHIFT, because lane counts are always 1, 2, 4 or 8.
module spi_width_model #(
parameter CNT_W = 32,
parameter CMD_BITS = 8
) (
input wire clk,
input wire rst_n,
input wire calc, // pulse: evaluate this shape
// Per-phase lane widths. Legal values are 1, 2, 4 and 8.
input wire [3:0] cmd_lanes,
input wire [3:0] addr_lanes,
input wire [3:0] data_lanes,
input wire [2:0] addr_bytes,
input wire [5:0] dummy_cycles, // CYCLES -- width cannot shrink these
input wire [15:0] data_bytes,
output reg [CNT_W-1:0] cycles, // this width combination
output reg [CNT_W-1:0] base_cycles, // the same transfer on one lane
output reg [CNT_W-1:0] saved_cycles, // base - cycles, a subtract
output reg [CNT_W-1:0] data_cycles, // the part width can reduce
output reg [CNT_W-1:0] fixed_cycles, // the part it cannot
output reg width_err // a lane count that is not 1/2/4/8
);
// Lane count to shift amount. A case rather than arithmetic, because
// the legal set is small and this is also where an illegal value is
// caught -- one table serving both jobs.
function [2:0] lane_shift;
input [3:0] lanes;
begin
case (lanes)
4'd1: lane_shift = 3'd0;
4'd2: lane_shift = 3'd1;
4'd4: lane_shift = 3'd2;
4'd8: lane_shift = 3'd3;
default: lane_shift = 3'd0; // flagged by width_err
endcase
end
endfunction
function lane_bad;
input [3:0] lanes;
begin
lane_bad = !((lanes == 4'd1) || (lanes == 4'd2) ||
(lanes == 4'd4) || (lanes == 4'd8));
end
endfunction
reg [CNT_W-1:0] cmd_bits, addr_bits, data_bits;
reg [CNT_W-1:0] c_cmd, c_addr, c_data;
always @(*) begin
cmd_bits = CMD_BITS;
addr_bits = {{(CNT_W-3){1'b0}}, addr_bytes} << 3;
data_bits = {{(CNT_W-16){1'b0}}, data_bytes} << 3;
// Each phase's cycle count is its bit count divided by its lanes,
// which is a right shift because the lane count is a power of two.
c_cmd = cmd_bits >> lane_shift(cmd_lanes);
c_addr = addr_bits >> lane_shift(addr_lanes);
c_data = data_bits >> lane_shift(data_lanes);
end
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
cycles <= {CNT_W{1'b0}};
base_cycles <= {CNT_W{1'b0}};
saved_cycles <= {CNT_W{1'b0}};
data_cycles <= {CNT_W{1'b0}};
fixed_cycles <= {CNT_W{1'b0}};
width_err <= 1'b0;
end else if (calc) begin
width_err <= lane_bad(cmd_lanes) || lane_bad(addr_lanes) ||
lane_bad(data_lanes);
cycles <= c_cmd + c_addr + dummy_cycles + c_data;
// The same transfer entirely on one lane: every phase at its
// full bit count, and the dummy phase UNCHANGED -- because it
// is cycles, not bits, and that is the whole point.
base_cycles <= cmd_bits + addr_bits + dummy_cycles + data_bits;
saved_cycles <= (cmd_bits + addr_bits + data_bits)
- (c_cmd + c_addr + c_data);
// The split that explains the diminishing return: only
// data_cycles shrinks with the data width, and fixed_cycles is
// the floor no amount of widening can go below.
data_cycles <= c_data;
fixed_cycles <= c_cmd + c_addr + dummy_cycles;
end
end
endmodule// spi_width_model_tb.v
//
// The same checks as the SystemVerilog testbench: every expected cycle
// count computed from the arithmetic, then the result the chapter is about
// -- the speedup from N data lanes is bounded by how much of the transfer
// was data, and at short lengths it is far below N.
`timescale 1ns/1ps
module spi_width_model_tb;
parameter CNT_W = 32;
reg clk;
reg rst_n;
reg calc;
reg [3:0] cmd_lanes;
reg [3:0] addr_lanes;
reg [3:0] data_lanes;
reg [2:0] addr_bytes;
reg [5:0] dummy_cycles;
reg [15:0] data_bytes;
wire [CNT_W-1:0] cycles, base_cycles, saved_cycles;
wire [CNT_W-1:0] data_cycles, fixed_cycles;
wire width_err;
integer errors;
integer m_base, m_quad;
integer c_lo, c_hi;
integer c1, c2, c4, c8;
integer len_i;
initial begin
clk = 1'b0; rst_n = 1'b0; calc = 1'b0;
cmd_lanes = 4'd1; addr_lanes = 4'd1; data_lanes = 4'd1;
addr_bytes = 3'd3; dummy_cycles = 6'd8; data_bytes = 16'd64;
errors = 0;
end
always #5 clk = ~clk;
spi_width_model #(.CNT_W(CNT_W), .CMD_BITS(8)) dut (
.clk(clk), .rst_n(rst_n), .calc(calc),
.cmd_lanes(cmd_lanes), .addr_lanes(addr_lanes),
.data_lanes(data_lanes),
.addr_bytes(addr_bytes), .dummy_cycles(dummy_cycles),
.data_bytes(data_bytes),
.cycles(cycles), .base_cycles(base_cycles),
.saved_cycles(saved_cycles), .data_cycles(data_cycles),
.fixed_cycles(fixed_cycles), .width_err(width_err)
);
// The testbench's own model, from the arithmetic rather than the design.
function integer model_cycles;
input integer cl;
input integer al;
input integer dl;
input integer ab;
input integer dum;
input integer len;
begin
model_cycles = (8 / cl) + ((ab * 8) / al) + dum + ((len * 8) / dl);
end
endfunction
task evaluate;
input integer cl;
input integer al;
input integer dl;
input integer ab;
input integer dum;
input integer len;
begin
@(negedge clk);
cmd_lanes = cl[3:0]; addr_lanes = al[3:0]; data_lanes = dl[3:0];
addr_bytes = ab[2:0]; dummy_cycles = dum[5:0];
data_bytes = len[15:0];
calc = 1'b1;
@(negedge clk);
calc = 1'b0;
@(negedge clk);
end
endtask
task check;
input [8*8:1] name;
input integer cl;
input integer al;
input integer dl;
input integer len;
integer want;
begin
evaluate(cl, al, dl, 3, 8, len);
want = model_cycles(cl, al, dl, 3, 8, len);
if (cycles !== want) begin
$display(" FAIL: %0s len=%0d reported %0d cycles, model says %0d",
name, len, cycles, want);
errors = errors + 1;
end
if (width_err) begin
$display(" FAIL: %0s reported a width error on legal lanes", name);
errors = errors + 1;
end
// The split must reconstruct the total -- a partition property.
if ((data_cycles + fixed_cycles) !== cycles) begin
$display(" FAIL: %0s data(%0d) + fixed(%0d) != total(%0d)",
name, data_cycles, fixed_cycles, cycles);
errors = errors + 1;
end
end
endtask
initial begin
repeat (3) @(negedge clk);
rst_n = 1'b1;
@(negedge clk);
// 1. The single-lane baseline: a fast read of 64 bytes.
check("1-1-1 ", 1, 1, 1, 64);
$display(" 1-1-1 len=64: %0d cycles (fixed %0d + data %0d)",
cycles, fixed_cycles, data_cycles);
if (cycles !== base_cycles) begin
$display(" FAIL: an all-one-lane transfer must equal its own baseline");
errors = errors + 1;
end
m_base = cycles;
// 2. Quad OUTPUT -- 1-1-4. Only the data phase widens.
check("1-1-4 ", 1, 1, 4, 64);
m_quad = cycles;
$display(" 1-1-4 len=64: %0d cycles (fixed %0d + data %0d) -- %0d.%02dx",
cycles, fixed_cycles, data_cycles,
m_base / cycles, ((m_base * 100) / cycles) % 100);
if (cycles >= m_base) begin
$display(" FAIL: widening the data phase did not reduce the cycle count");
errors = errors + 1;
end
// THE POINT: four data lanes, less than four times the speedup.
if ((m_base * 100) / cycles >= 400) begin
$display(" FAIL: 1-1-4 achieved 4x or better -- the fixed phases must cost something");
errors = errors + 1;
end
// 3. Quad I/O -- 1-4-4.
check("1-4-4 ", 1, 4, 4, 64);
$display(" 1-4-4 len=64: %0d cycles (fixed %0d + data %0d) -- %0d.%02dx",
cycles, fixed_cycles, data_cycles,
m_base / cycles, ((m_base * 100) / cycles) % 100);
if (cycles >= m_quad) begin
$display(" FAIL: widening the address phase did not help");
errors = errors + 1;
end
// 4. Full quad -- 4-4-4.
check("4-4-4 ", 4, 4, 4, 64);
$display(" 4-4-4 len=64: %0d cycles (fixed %0d + data %0d) -- %0d.%02dx",
cycles, fixed_cycles, data_cycles,
m_base / cycles, ((m_base * 100) / cycles) % 100);
// 5. Octal -- 8-8-8. Eight data lanes, still not eight times,
// because the dummy phase is CYCLES and did not move.
check("8-8-8 ", 8, 8, 8, 64);
$display(" 8-8-8 len=64: %0d cycles (fixed %0d + data %0d) -- %0d.%02dx",
cycles, fixed_cycles, data_cycles,
m_base / cycles, ((m_base * 100) / cycles) % 100);
if ((m_base * 100) / cycles >= 800) begin
$display(" FAIL: 8-8-8 achieved 8x or better -- the dummy phase must cost something");
errors = errors + 1;
end
// 6. THE AMDAHL RESULT on a SHORT transfer.
check("1-1-1 ", 1, 1, 1, 4);
m_base = cycles;
$display(" 1-1-1 len=4: %0d cycles", cycles);
check("1-1-4 ", 1, 1, 4, 4);
$display(" 1-1-4 len=4: %0d cycles -- %0d.%02dx (four data lanes!)",
cycles, m_base / cycles, ((m_base * 100) / cycles) % 100);
if ((m_base * 100) / cycles >= 200) begin
$display(" FAIL: on a 4-byte transfer quad output gave 2x or more");
errors = errors + 1;
end
check("4-4-4 ", 4, 4, 4, 4);
$display(" 4-4-4 len=4: %0d cycles -- %0d.%02dx",
cycles, m_base / cycles, ((m_base * 100) / cycles) % 100);
// 7. THE FLOOR.
check("4-4-4 ", 4, 4, 4, 0);
if (data_cycles !== 0) begin
$display(" FAIL: a zero-length transfer reported %0d data cycles",
data_cycles);
errors = errors + 1;
end
$display(" 4-4-4 len=0: %0d cycles -- the floor width cannot go below",
fixed_cycles);
// 8. The dummy phase is in CYCLES and never shrinks.
evaluate(4, 4, 4, 3, 0, 64); c_lo = cycles;
evaluate(4, 4, 4, 3, 8, 64); c_hi = cycles;
if ((c_hi - c_lo) != 8) begin
$display(" FAIL: 8 dummy cycles changed the total by %0d at 4-4-4",
c_hi - c_lo);
errors = errors + 1;
end
evaluate(8, 8, 8, 3, 0, 64); c_lo = cycles;
evaluate(8, 8, 8, 3, 8, 64); c_hi = cycles;
if ((c_hi - c_lo) != 8) begin
$display(" FAIL: 8 dummy cycles changed the total by %0d at 8-8-8",
c_hi - c_lo);
errors = errors + 1;
end
$display(" dummy phase costs 8 cycles at 1-1-1, 4-4-4 and 8-8-8 alike");
// 9. An illegal lane count is reported rather than treated as one.
evaluate(1, 1, 3, 3, 8, 64);
if (!width_err) begin
$display(" FAIL: three data lanes was not reported as illegal");
errors = errors + 1;
end
evaluate(1, 1, 8, 3, 8, 64);
if (width_err) begin
$display(" FAIL: eight data lanes was reported as illegal");
errors = errors + 1;
end
$display(" lane validation: 3 rejected, 8 accepted");
// 10. MONOTONICITY across a sweep.
for (len_i = 0; len_i <= 256; len_i = len_i + 8) begin
evaluate(1, 1, 1, 3, 8, len_i); c1 = cycles;
evaluate(1, 1, 2, 3, 8, len_i); c2 = cycles;
evaluate(1, 1, 4, 3, 8, len_i); c4 = cycles;
evaluate(1, 1, 8, 3, 8, len_i); c8 = cycles;
if (!(c8 <= c4 && c4 <= c2 && c2 <= c1)) begin
$display(" FAIL: len=%0d widths are not monotone (%0d,%0d,%0d,%0d)",
len_i, c1, c2, c4, c8);
errors = errors + 1;
end
if ((data_cycles + fixed_cycles) !== cycles) begin
$display(" FAIL: len=%0d the partition does not reconstruct the total",
len_i);
errors = errors + 1;
end
end
$display(" 33 lengths swept: wider is never slower, and the partition always reconstructs");
if (errors == 0)
$display("PASS: the cycle count matches the arithmetic at every width, the data and fixed phases partition the total exactly, widening is never slower, the dummy phase costs the same at every width because it is measured in cycles, four data lanes give well under four times the speedup and under two times on a short transfer, and an illegal lane count is reported rather than treated as one lane");
else
$display("FAIL: %0d error(s)", errors);
$finish;
end
endmodule-- spi_width_model.vhd
--
-- Chapter 12.1 -- what widening the bus actually buys, in VHDL.
--
-- A wider SPI moves more bits per clock on the DATA phase. It does not
-- widen the opcode, address or dummy phases unless the device supports
-- doing so, and those are a fixed cost width cannot reduce:
--
-- cycles = 8/cmd_lanes + 8*addr_bytes/addr_lanes + dummy
-- + 8*len/data_lanes
--
-- The dummy phase is the interesting term. It is quoted in CLOCK CYCLES,
-- so widening the bus does not shorten it at all -- and as the data phase
-- shrinks, the dummy phase becomes a larger share of what remains.
--
-- This block reports the cycle count and, alongside it, the same transfer
-- entirely on one lane. It deliberately does NOT compute the ratio: a
-- divider costs more than the insight, and the difference is a subtract and
-- the more useful number for a latency budget.
--
-- DIVISION BY LANES IS A SHIFT, because lane counts are always 1, 2, 4 or 8.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_width_model is
generic (
CNT_W : positive := 32;
CMD_BITS : natural := 8
);
port (
clk : in std_logic;
rst_n : in std_logic;
calc : in std_logic; -- pulse: evaluate this shape
-- Per-phase lane widths. Legal values are 1, 2, 4 and 8.
cmd_lanes : in unsigned(3 downto 0);
addr_lanes : in unsigned(3 downto 0);
data_lanes : in unsigned(3 downto 0);
addr_bytes : in unsigned(2 downto 0);
dummy_cycles : in unsigned(5 downto 0); -- CYCLES, not bits
data_bytes : in unsigned(15 downto 0);
cycles : out unsigned(CNT_W - 1 downto 0);
base_cycles : out unsigned(CNT_W - 1 downto 0);
saved_cycles : out unsigned(CNT_W - 1 downto 0);
data_cycles : out unsigned(CNT_W - 1 downto 0);
fixed_cycles : out unsigned(CNT_W - 1 downto 0);
width_err : out std_logic
);
end entity;
architecture rtl of spi_width_model is
-- Lane count to shift amount. A case rather than arithmetic, because
-- the legal set is small and this is also where an illegal value is
-- caught -- one table serving both jobs.
function lane_shift(lanes : unsigned(3 downto 0)) return natural is
begin
case to_integer(lanes) is
when 1 => return 0;
when 2 => return 1;
when 4 => return 2;
when 8 => return 3;
when others => return 0; -- flagged by width_err
end case;
end function;
function lane_bad(lanes : unsigned(3 downto 0)) return boolean is
begin
return not (to_integer(lanes) = 1 or to_integer(lanes) = 2 or
to_integer(lanes) = 4 or to_integer(lanes) = 8);
end function;
signal cyc_r : unsigned(CNT_W - 1 downto 0) := (others => '0');
signal base_r : unsigned(CNT_W - 1 downto 0) := (others => '0');
signal saved_r : unsigned(CNT_W - 1 downto 0) := (others => '0');
signal data_r : unsigned(CNT_W - 1 downto 0) := (others => '0');
signal fixed_r : unsigned(CNT_W - 1 downto 0) := (others => '0');
signal err_r : std_logic := '0';
begin
model : process (clk, rst_n)
-- These are NOT named after the generic. VHDL is case-insensitive,
-- so a variable called cmd_bits would be the SAME identifier as the
-- generic CMD_BITS and would shadow it inside this process -- making
-- the assignment a self-assignment of an uninitialised natural.
-- Legal, silent, and zero.
variable n_cmd, n_addr, n_data : natural;
variable c_cmd, c_addr, c_data : natural;
begin
if rst_n = '0' then
cyc_r <= (others => '0');
base_r <= (others => '0');
saved_r <= (others => '0');
data_r <= (others => '0');
fixed_r <= (others => '0');
err_r <= '0';
elsif rising_edge(clk) then
if calc = '1' then
n_cmd := CMD_BITS;
n_addr := to_integer(addr_bytes) * 8;
n_data := to_integer(data_bytes) * 8;
-- Each phase's cycle count is its bit count divided by its
-- lanes, which is a right shift because the lane count is a
-- power of two.
c_cmd := n_cmd / (2 ** lane_shift(cmd_lanes));
c_addr := n_addr / (2 ** lane_shift(addr_lanes));
c_data := n_data / (2 ** lane_shift(data_lanes));
if lane_bad(cmd_lanes) or lane_bad(addr_lanes) or
lane_bad(data_lanes) then
err_r <= '1';
else
err_r <= '0';
end if;
cyc_r <= to_unsigned(c_cmd + c_addr +
to_integer(dummy_cycles) + c_data, CNT_W);
-- The same transfer entirely on one lane: every phase at its
-- full bit count, and the dummy phase UNCHANGED -- because
-- it is cycles, not bits, and that is the whole point.
base_r <= to_unsigned(n_cmd + n_addr +
to_integer(dummy_cycles) + n_data,
CNT_W);
saved_r <= to_unsigned((n_cmd + n_addr + n_data) -
(c_cmd + c_addr + c_data), CNT_W);
-- The split that explains the diminishing return: only
-- data_cycles shrinks with the data width, and fixed_cycles
-- is the floor no amount of widening can go below.
data_r <= to_unsigned(c_data, CNT_W);
fixed_r <= to_unsigned(c_cmd + c_addr +
to_integer(dummy_cycles), CNT_W);
end if;
end if;
end process;
cycles <= cyc_r;
base_cycles <= base_r;
saved_cycles <= saved_r;
data_cycles <= data_r;
fixed_cycles <= fixed_r;
width_err <= err_r;
end architecture;-- spi_width_model_tb.vhd
--
-- The same checks as the SystemVerilog and Verilog testbenches: every
-- expected cycle count computed from the arithmetic, then the result the
-- chapter is about -- the speedup from N data lanes is bounded by how much
-- of the transfer was data, and at short lengths it is far below N.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_width_model_tb is
end entity;
architecture sim of spi_width_model_tb is
constant CNT_W : positive := 32;
signal clk : std_logic := '0';
signal rst_n : std_logic := '0';
signal halt : boolean := false;
signal calc : std_logic := '0';
signal cmd_lanes : unsigned(3 downto 0) := to_unsigned(1, 4);
signal addr_lanes : unsigned(3 downto 0) := to_unsigned(1, 4);
signal data_lanes : unsigned(3 downto 0) := to_unsigned(1, 4);
signal addr_bytes : unsigned(2 downto 0) := to_unsigned(3, 3);
signal dummy_cycles : unsigned(5 downto 0) := to_unsigned(8, 6);
signal data_bytes : unsigned(15 downto 0) := to_unsigned(64, 16);
signal cycles : unsigned(CNT_W - 1 downto 0);
signal base_cycles : unsigned(CNT_W - 1 downto 0);
signal saved_cycles : unsigned(CNT_W - 1 downto 0);
signal data_cycles : unsigned(CNT_W - 1 downto 0);
signal fixed_cycles : unsigned(CNT_W - 1 downto 0);
signal width_err : std_logic;
signal errors : natural := 0;
-- The testbench's own model, from the arithmetic rather than the design.
function model_cycles(cl : natural; al : natural; dl : natural;
ab : natural; dum : natural;
len : natural) return natural is
begin
return (8 / cl) + ((ab * 8) / al) + dum + ((len * 8) / dl);
end function;
begin
clk <= not clk after 5 ns when not halt else '0';
dut : entity work.spi_width_model
generic map (CNT_W => CNT_W, CMD_BITS => 8)
port map (
clk => clk, rst_n => rst_n, calc => calc,
cmd_lanes => cmd_lanes, addr_lanes => addr_lanes,
data_lanes => data_lanes,
addr_bytes => addr_bytes, dummy_cycles => dummy_cycles,
data_bytes => data_bytes,
cycles => cycles, base_cycles => base_cycles,
saved_cycles => saved_cycles, data_cycles => data_cycles,
fixed_cycles => fixed_cycles, width_err => width_err
);
stim : process
variable errs : natural := 0;
variable m_base, m_quad : natural;
variable c_lo, c_hi : natural;
variable v1, v2, v4, v8 : natural;
procedure evaluate(cl : natural; al : natural; dl : natural;
ab : natural; dum : natural; len : natural) is
begin
wait until falling_edge(clk);
cmd_lanes <= to_unsigned(cl, 4);
addr_lanes <= to_unsigned(al, 4);
data_lanes <= to_unsigned(dl, 4);
addr_bytes <= to_unsigned(ab, 3);
dummy_cycles <= to_unsigned(dum, 6);
data_bytes <= to_unsigned(len, 16);
calc <= '1';
wait until falling_edge(clk);
calc <= '0';
wait until falling_edge(clk);
end procedure;
procedure check(name : string; cl : natural; al : natural;
dl : natural; len : natural) is
variable want : natural;
begin
evaluate(cl, al, dl, 3, 8, len);
want := model_cycles(cl, al, dl, 3, 8, len);
if to_integer(cycles) /= want then
report " FAIL: " & name & " reported " &
integer'image(to_integer(cycles)) &
" cycles, model says " & integer'image(want);
errs := errs + 1;
end if;
if width_err = '1' then
report " FAIL: " & name & " reported a width error on legal lanes";
errs := errs + 1;
end if;
-- The split must reconstruct the total -- a partition property.
if data_cycles + fixed_cycles /= cycles then
report " FAIL: " & name & " the partition does not reconstruct the total";
errs := errs + 1;
end if;
end procedure;
begin
for k in 0 to 2 loop
wait until falling_edge(clk);
end loop;
rst_n <= '1';
wait until falling_edge(clk);
-- 1. The single-lane baseline: a fast read of 64 bytes.
check("1-1-1", 1, 1, 1, 64);
report " 1-1-1 len=64: " & integer'image(to_integer(cycles)) &
" cycles (fixed " & integer'image(to_integer(fixed_cycles)) &
" + data " & integer'image(to_integer(data_cycles)) & ")";
if cycles /= base_cycles then
report " FAIL: an all-one-lane transfer must equal its own baseline";
errs := errs + 1;
end if;
m_base := to_integer(cycles);
-- 2. Quad OUTPUT -- 1-1-4. Only the data phase widens.
check("1-1-4", 1, 1, 4, 64);
m_quad := to_integer(cycles);
report " 1-1-4 len=64: " & integer'image(to_integer(cycles)) &
" cycles (fixed " & integer'image(to_integer(fixed_cycles)) &
" + data " & integer'image(to_integer(data_cycles)) & ") -- x100 = " &
integer'image((m_base * 100) / to_integer(cycles));
if to_integer(cycles) >= m_base then
report " FAIL: widening the data phase did not reduce the cycle count";
errs := errs + 1;
end if;
-- THE POINT: four data lanes, less than four times the speedup.
if (m_base * 100) / to_integer(cycles) >= 400 then
report " FAIL: 1-1-4 achieved 4x or better";
errs := errs + 1;
end if;
-- 3. Quad I/O -- 1-4-4.
check("1-4-4", 1, 4, 4, 64);
report " 1-4-4 len=64: " & integer'image(to_integer(cycles)) &
" cycles (fixed " & integer'image(to_integer(fixed_cycles)) &
" + data " & integer'image(to_integer(data_cycles)) & ") -- x100 = " &
integer'image((m_base * 100) / to_integer(cycles));
if to_integer(cycles) >= m_quad then
report " FAIL: widening the address phase did not help";
errs := errs + 1;
end if;
-- 4. Full quad -- 4-4-4.
check("4-4-4", 4, 4, 4, 64);
report " 4-4-4 len=64: " & integer'image(to_integer(cycles)) &
" cycles (fixed " & integer'image(to_integer(fixed_cycles)) &
" + data " & integer'image(to_integer(data_cycles)) & ") -- x100 = " &
integer'image((m_base * 100) / to_integer(cycles));
-- 5. Octal -- 8-8-8. Eight data lanes, still not eight times.
check("8-8-8", 8, 8, 8, 64);
report " 8-8-8 len=64: " & integer'image(to_integer(cycles)) &
" cycles (fixed " & integer'image(to_integer(fixed_cycles)) &
" + data " & integer'image(to_integer(data_cycles)) & ") -- x100 = " &
integer'image((m_base * 100) / to_integer(cycles));
if (m_base * 100) / to_integer(cycles) >= 800 then
report " FAIL: 8-8-8 achieved 8x or better";
errs := errs + 1;
end if;
-- 6. THE AMDAHL RESULT on a SHORT transfer.
check("1-1-1", 1, 1, 1, 4);
m_base := to_integer(cycles);
report " 1-1-1 len=4: " & integer'image(to_integer(cycles)) & " cycles";
check("1-1-4", 1, 1, 4, 4);
report " 1-1-4 len=4: " & integer'image(to_integer(cycles)) &
" cycles -- x100 = " &
integer'image((m_base * 100) / to_integer(cycles)) &
" (four data lanes!)";
if (m_base * 100) / to_integer(cycles) >= 200 then
report " FAIL: on a 4-byte transfer quad output gave 2x or more";
errs := errs + 1;
end if;
check("4-4-4", 4, 4, 4, 4);
report " 4-4-4 len=4: " & integer'image(to_integer(cycles)) &
" cycles -- x100 = " &
integer'image((m_base * 100) / to_integer(cycles));
-- 7. THE FLOOR.
check("4-4-4", 4, 4, 4, 0);
if data_cycles /= 0 then
report " FAIL: a zero-length transfer reported data cycles";
errs := errs + 1;
end if;
report " 4-4-4 len=0: " & integer'image(to_integer(fixed_cycles)) &
" cycles -- the floor width cannot go below";
-- 8. The dummy phase is in CYCLES and never shrinks.
evaluate(4, 4, 4, 3, 0, 64); c_lo := to_integer(cycles);
evaluate(4, 4, 4, 3, 8, 64); c_hi := to_integer(cycles);
if (c_hi - c_lo) /= 8 then
report " FAIL: the dummy phase did not cost 8 cycles at 4-4-4";
errs := errs + 1;
end if;
evaluate(8, 8, 8, 3, 0, 64); c_lo := to_integer(cycles);
evaluate(8, 8, 8, 3, 8, 64); c_hi := to_integer(cycles);
if (c_hi - c_lo) /= 8 then
report " FAIL: the dummy phase did not cost 8 cycles at 8-8-8";
errs := errs + 1;
end if;
report " dummy phase costs 8 cycles at 1-1-1, 4-4-4 and 8-8-8 alike";
-- 9. An illegal lane count is reported rather than treated as one.
evaluate(1, 1, 3, 3, 8, 64);
if width_err /= '1' then
report " FAIL: three data lanes was not reported as illegal";
errs := errs + 1;
end if;
evaluate(1, 1, 8, 3, 8, 64);
if width_err = '1' then
report " FAIL: eight data lanes was reported as illegal";
errs := errs + 1;
end if;
report " lane validation: 3 rejected, 8 accepted";
-- 10. MONOTONICITY across a sweep.
for k in 0 to 32 loop
evaluate(1, 1, 1, 3, 8, k * 8); v1 := to_integer(cycles);
evaluate(1, 1, 2, 3, 8, k * 8); v2 := to_integer(cycles);
evaluate(1, 1, 4, 3, 8, k * 8); v4 := to_integer(cycles);
evaluate(1, 1, 8, 3, 8, k * 8); v8 := to_integer(cycles);
if not (v8 <= v4 and v4 <= v2 and v2 <= v1) then
report " FAIL: the widths are not monotone"; errs := errs + 1;
end if;
if data_cycles + fixed_cycles /= cycles then
report " FAIL: the partition does not reconstruct the total";
errs := errs + 1;
end if;
end loop;
report " 33 lengths swept: wider is never slower, and the partition always reconstructs";
errors <= errs;
if errs = 0 then
report "PASS: the cycle count matches the arithmetic at every width, the data and fixed phases partition the total exactly, widening is never slower, the dummy phase costs the same at every width because it is measured in cycles, four data lanes give well under four times the speedup and under two times on a short transfer, and an illegal lane count is reported rather than treated as one lane";
else
report "FAIL: " & integer'image(errs) & " error(s)" severity error;
end if;
halt <= true;
wait;
end process;
end architecture;Parity
All three implement the same model: identical ports and generics, division by lanes as a shift, a single-lane baseline computed alongside, the data and fixed phases published separately, and an illegal lane count reported rather than treated as one lane. All three testbenches compute the same expectations independently and report identical results — 552, 168, 150, 144 and 76 cycles at the five widths, and 1.50× for quad output on a four-byte transfer.
One VHDL detail is worth naming because it produced a silent wrong answer rather than an error. A local variable named cmd_bits and the generic CMD_BITS are the same identifier, because VHDL is case-insensitive — so cmd_bits := CMD_BITS; is a self-assignment of an uninitialised natural, which is zero. It analyses cleanly, simulates without a warning, and reports every cycle count eight short. The variables are named n_cmd, n_addr and n_data for exactly that reason.
7. Why a Verification Engineer Cares
// 1. THE PARTITION. The data and fixed phases sum to the total. This is
// the decomposition the whole chapter rests on, and it catches two
// outputs that are each plausible and mutually inconsistent.
a_partition : assert property (
@(posedge clk) disable iff (!rst_n)
(valid) |-> ((data_cycles + fixed_cycles) == cycles))
else $error("the data and fixed phases do not sum to the total");
// 2. MONOTONICITY. A wider data phase is never slower. Stated as a
// property because it is a claim about the FUNCTION rather than about
// any particular input, and a sign-error in the shift breaks it.
a_monotone : assert property (
@(posedge clk) disable iff (!rst_n)
(valid && same_shape_wider_data) |-> (cycles <= $past(cycles)))
else $error("widening the data phase increased the cycle count");
// 3. THE AMDAHL BOUND. The gain from N data lanes is strictly less than
// N whenever any fixed phase is non-zero. A design reporting N would
// be claiming the fixed phases cost nothing.
a_bound : assert property (
@(posedge clk) disable iff (!rst_n)
(valid && (fixed_cycles != 0)) |->
((base_cycles * 1) < (cycles * data_lanes)))
else $error("the reported gain reached or exceeded the lane count");
// 4. THE DUMMY PHASE IS WIDTH-INVARIANT. Two shapes differing only in
// dummy count differ in total by exactly that -- at every width.
a_dummy_invariant : assert property (
@(posedge clk) disable iff (!rst_n)
(valid) |-> (fixed_cycles >= dummy_cycles))
else $error("the fixed cost is smaller than the dummy count alone");
// 5. AN ILLEGAL WIDTH IS REPORTED, not silently treated as one lane.
// Three lanes is not a width any device has, and a plausible answer
// to an impossible question is worse than a refusal.
a_width_checked : assert property (
@(posedge clk) disable iff (!rst_n)
(valid && !is_pow2_lane(data_lanes)) |-> width_err)
else $error("an illegal lane count was accepted");Property 3 is the one worth copying, and not for SPI. When a design's value is a claimed improvement, the bound on that improvement is the specification — stronger and more durable than any table of expected values, because it holds for inputs nobody enumerated. A model that returned exactly 4× for four lanes would pass a spot check against a hand-computed table that happened to omit the fixed phases; it cannot pass property 3.
Property 1 is the partition property this track has now used for bytes (Chapter 9.3), for frame time (Chapter 9.2), for protocol phases (Chapter 10.5) and now for width. The pattern is general: whenever a design splits a quantity, the split summing correctly is close to a complete specification of the accounting.
Coverage must cross the width with the length, because the interesting result lives at one corner:
covergroup spi_width_cg @(posedge clk iff calc);
cp_data_lanes : coverpoint data_lanes {
bins one = {1};
bins two = {2};
bins four = {4};
bins eight = {8};
illegal_bins impossible = {0, 3, [5:7], [9:15]};
}
// The shape of the width combination, which is what the x-y-z
// notation names -- and a suite testing only 1-1-4 has tested one.
cp_shape : coverpoint shape_class {
bins single = {S_111};
bins quad_out = {S_114}; // only data widens
bins quad_io = {S_144}; // address widens too
bins full_quad = {S_444}; // and the command
bins octal = {S_888};
}
// Length decides how much of the transfer was data, and therefore
// how much width can win. Covering it alone proves nothing.
cp_len : coverpoint data_bytes {
bins zero = {0}; // the floor, pure overhead
bins tiny = {[1:8]}; // width barely helps
bins medium = {[9:64]};
bins large = {[65:$]}; // width nearly saturates
}
// The dummy count, because it is the term width cannot reduce and a
// suite that always uses 8 never shows that.
cp_dummy : coverpoint dummy_cycles {
bins none = {0}; // where the gain is largest
bins odd = {[1:7]};
bins byte = {8};
bins many = {[9:32]};
}
// THE CROSS THAT MATTERS. The gain is a function of both, so a
// suite covering each axis separately has characterised neither.
x_shape_len : cross cp_shape, cp_len;
x_shape_dummy : cross cp_shape, cp_dummy;
endgroupx_shape_len is the coverage goal, and cp_dummy's none bin is the one most often absent. With no dummy phase the gain is at its maximum, and a suite that always specifies eight dummy cycles never sees the upper bound of the design's own behaviour.
8. Why an FPGA or ASIC Engineer Cares
Budget in cycles before choosing a width. The arithmetic runs on a datasheet in a minute. Discovering after a board spin that quad output bought 1.5× on your access pattern is a discovery worth making earlier.
Widen the command and address phases if the device supports it. At short lengths that is worth more than widening the data phase — 1.50× against 3.00× in §4, from the same data width. A controller that implements only 1-1-4 has left the larger half on the table.
Divide by lane counts with a shift. Lane counts are 1, 2, 4 and 8 without exception, so bits / lanes is a shift. Writing the division infers a divider for a quantity that is always a power of two.
Validate the lane count. Three lanes does not exist. Accepting it and treating it as one produces a plausible cycle count for an impossible configuration, which is worse than a refusal.
Report the fixed and data cycles separately. Two registers turn "why is the quad link slower than expected?" into a reading. If the fixed count dominates, widening the data phase further will not help and the answer is elsewhere.
Check whether the device derates its clock in wide modes. Some parts specify a lower maximum SCLK for quad than for single-lane operation, because more simultaneously switching outputs means more supply noise. A 4× width at 0.8× the clock is 3.2×, and the cycle model alone will not tell you.
9. Failure Signature — A Quad Upgrade That Delivers Almost Nothing
Symptom. A design is moved from single-lane fast read to quad output to speed up a configuration load. The link works correctly — every byte verifies — and the measured improvement is about 15%, against an expected 4×.
What "every byte verifies" establishes. The lane mapping, the dummy count and the command are all correct. This is not a correctness problem at all, which is the first thing to settle because it rules out everything Modules 10 and 11 were about.
Plausible mechanisms.
- The transfers are short. If the load is issued as many small accesses rather than a few large ones, §4's result applies to each of them and 1.5× is the ceiling — before any other effect. This fits a modest gain exactly.
- The command and address phases were not widened. 1-1-4 rather than 4-4-4, which at short lengths is half the available gain.
- The device derates its clock in quad mode, so the width gain is partly cancelled by a rate loss.
- The bottleneck is elsewhere — a software copy loop, a DMA that cannot keep up, or the destination memory. In that case the SPI link is no longer the limit and no amount of widening helps.
- The dummy count was increased when quad mode was enabled, as many parts require, adding fixed cost exactly where §3 showed it hurts.
The discriminating observation. Measure the transfer length distribution, not the throughput. That single measurement splits the case:
- If the accesses are tens of bytes or fewer, the answer is §4 and the fix is to batch them — no amount of protocol work will help otherwise.
- If the accesses are kilobytes and the gain is still 15%, the SPI link is not the bottleneck, and the next thing to instrument is the consumer.
Then, if the accesses really are large, compute the expected cycle count with the model of §6 and compare against the measured time converted to cycles. A large discrepancy points at the clock rate — check whether quad mode derated it — and a close match confirms the bottleneck is outside the link.
The fix depends entirely on which it was, which is the point of measuring first. But the common root is that a 4× width was expected to give 4× without anyone computing the fixed share, and that arithmetic takes a minute.
Why the investigation goes wrong. Because "we widened the bus and it barely helped" sounds like a wiring or configuration fault, so the search goes to the lane mapping and the dummy count — both of which the verifying data has already exonerated. The failure is not in the implementation; it is in the expectation, and no amount of debugging the implementation will correct an expectation.
10. Common Misconceptions
11. Reason It Through
Work this before reading the answer.
A design reads 16-byte records from flash, one at a time, at 50 MHz with three address bytes and eight dummy cycles. An engineer proposes two changes and asks which to do first:
(a) move from 1-1-1 to 4-4-4, or (b) keep 1-1-1 and batch sixteen records into one 256-byte read.
Which wins, and what does that tell you?
Compute the baseline. One 16-byte record at 1-1-1:
8 + 24 + 8 + 128 = 168 cycles per record
16 records = 2688 cyclesOption (a): 4-4-4, still one record per transaction.
2 + 6 + 8 + 32 = 48 cycles per record
16 records = 768 cycles → 3.50×Option (b): 1-1-1, one 256-byte read.
8 + 24 + 8 + 2048 = 2088 cycles → 1.29×Option (a) wins decisively, 3.50× against 1.29×. Batching helps far less than it might seem, because at 1-1-1 the data phase already dominated — 2048 of 2088 cycles — so there was almost no overhead left to amortise.
Now do both. 4-4-4 with one 256-byte read:
2 + 6 + 8 + 512 = 528 cycles → 5.09×And here is the result worth having. Batching added 1.29× on its own and 1.45× on top of 4-4-4 (768 → 528). It is worth more after widening than before, because widening shrank the data phase and thereby made the per-record overhead a larger share — so there was suddenly something for batching to amortise.
The general lesson, and it is the module's organising idea. Width and batching attack different terms: width shrinks the bit-denominated phases, batching amortises the per-transaction ones. Which matters depends on which currently dominates, and applying one changes the answer for the other. So the order to work in is: compute the split of fixed to data cycles, attack whichever is larger, then recompute — because the split will have moved.
One caveat worth stating. Option (b) assumes the sixteen records are contiguous and that 256 bytes of buffer exist. If the records are scattered, batching is unavailable and (a) is the only option — which is Chapter 9.3's point about buffer and locality constraints arriving in a new form.
12. Understanding Check
13. Summary
Width exists because frequency stopped scaling and pins were not available. Repurposing the write-protect and hold pins gives four bidirectional data lanes in the same package, and it leaves every board timing budget unchanged — which is why width reached eight lanes while frequency stalled.
But widening the bus shortens only what is measured in bits. The dummy phase is measured in cycles, so it does not shrink at all, and as the data phase narrows it becomes the dominant irreducible term — eight of the twelve fixed cycles at 8-8-8.
So the speedup is bounded by the data share:
64 bytes: 1-1-4 → 3.28× 8-8-8 → 7.26×
4 bytes: 1-1-4 → 1.50× 8-8-8 → 4.50×That is Amdahl's argument on a serial bus, and it has two consequences. Widening the command and address phases matters more than it looks — 1.50× against 3.00× on a short access from the same data width, which is why 4-4-4 is a distinct mode. And short accesses need the overhead removed rather than the data widened, which is what continuous read and XIP do and why Chapter 12.5 is this module's destination.
In hardware every division by a lane count is a shift, because lane counts are always powers of two; an illegal width is reported rather than rounded; and the data and fixed cycles are published separately, because their sum is the total and only one of them shrinks.
For verification, the bound is the specification — strictly less than N for N lanes — which is stronger than any table because it holds for inputs nobody enumerated. And the partition property joins the ones this track has used for bytes, frame time and protocol phases: whenever a design splits a quantity, the split summing correctly is nearly a complete specification.
14. What Comes Next
The gain is now computable. What has not been said is how a byte actually crosses four wires.
Chapter 12.2 — Dual, Quad, and Octal SPI answers the question single-lane SPI never had to ask: which bit goes on which lane. There is a uniform convention, getting it backwards turns 0xA5 into 0x5A — a plausible value rather than an obvious error — and the same pins that carry data outbound must carry it inbound, which makes the output enable a safety signal rather than a convenience. It ends with the gearbox that serialises a byte across any width and deserialises it back, verified exhaustively in all three HDLs.
Continue learning
Related tutorials
- Related topic
Dual, Quad, and Octal SPI
How dedicated pins become bidirectional lanes, the bit-to-lane convention every multi-lane device shares and what getting it backwards produces, why the output enable becomes a safety signal, and the gearbox that serialises a byte across any width.
- Related topic
Lane Widths per Phase and Direction Changes
Reading the 1-1-4 / 1-4-4 / 4-4-4 notation properly, the direction it does not state, why a multi-lane read reverses four wires at once, and why that is the second reason a quad read needs dummy cycles.
- Related topic
Launch and Sample Edges
One edge of each bit time places a bit on the wire, the other captures it, and they must never be the same edge. Why the separation is forced, why it buys half a period, and how RTL maps physical edges onto those roles.
- Related topic
Deriving Mode Behaviour from CPOL and CPHA
The four SPI modes are a two-bit truth table you can rebuild in seconds. The standard numbering, the derivation, the complete mode decoder in three HDLs, and the assertions that keep a configurable design honest.
