SPI · Module 12
Lane Widths per Phase and Direction Changes
Reading the 1-1-4 / 1-4-4 / 4-4-4 notation properly, the direction it does not state, why a multi-lane read reverses four wires at once, and why that is the second reason a quad read needs dummy cycles.
Chapter 12.2 established that the data pins become bidirectional. This chapter is about what follows from that, and it resolves something Module 11 left half-explained.
A single-lane read never changes a wire's direction — MOSI is always the master's and MISO always the device's. A quad read reverses four wires at once. Where does that happen, and what pays for it?
The dummy phase pays for it. Chapter 11.2 said dummy cycles buy array access time, which is true and is only half the reason.
1. Reading the Notation
The x-y-z form names the lane width of three phases, in order:
1-1-1 command 1 lane address 1 lane data 1 lane
1-1-2 command 1 address 1 data 2 (dual output)
1-2-2 command 1 address 2 data 2 (dual I/O)
1-1-4 command 1 address 1 data 4 (quad output, 0x6B)
1-4-4 command 1 address 4 data 4 (quad I/O, 0xEB)
4-4-4 command 4 address 4 data 4 (QPI mode)
8-8-8 all eight lanes (octal / OPI)Three things the notation encodes, and one large thing it does not.
The opcode is always sent on the width the device is currently in. That is why 4-4-4 needs the device to have been put into a quad-command mode first: a controller in single-lane mode cannot send a four-lane opcode, because the device would not be listening on four lanes. So 4-4-4 is a mode, while 1-1-4 and 1-4-4 are ordinary commands issued from single-lane mode.
The dummy phase is not in the notation. It has no width — it is measured in cycles (Chapter 12.1) — so there is nothing for a digit to say.
Each combination is a distinct opcode, not a configuration. 0x6B is quad output and 0xEB is quad I/O; a device supports the ones its command table lists and no others. There is no 1-2-4, because no device implements one.
And what the notation does not say is the DIRECTION of each phase. That is where the cost lives.
2. The Direction Nobody Writes Down
Take a quad-output read, 1-1-4. Walk the phases and ask who drives:
phase lanes who drives
command 1 the MASTER -- on IO0
address 1 the MASTER -- on IO0
dummy 4 NOBODY -- the master has released, the device
has not yet started
data 4 the DEVICE -- on IO0..IO3The bus reverses exactly once, between the address and the data. And the reversal cannot be instantaneous, for the reason Chapter 8.4 gave: the master's drivers take time to release, the device's take time to assert, and any overlap is contention.
So something must separate them. On a single-lane read nothing had to, because MOSI and MISO are different wires and the master never stops driving MOSI. On a multi-lane read the same four wires carry both directions, and the separation has to be built into the transaction.
It is the dummy phase. That is the second reason it exists, and it unifies two things this track has treated separately:
Chapter 11.2: the dummy phase gives the ARRAY time to produce data
Chapter 12.3: the dummy phase gives the BUS time to reverseBoth are true, both are paid for by the same cycles, and neither is redundant. A device whose array is fast enough to need no latency still needs a turnaround window if its read is multi-lane.
3. The Read With No Window
Now the configuration that cannot work.
A read specified with zero dummy cycles on multiple lanes has no phase between the address and the data. So on one clock edge the master must stop driving four lanes and the device must start driving them — with no interval between.
That is contention by construction. Not a marginal timing risk but a specification error, and it is why no real device offers a multi-lane read with zero dummy cycles. The plain single-lane read 0x03 has none and needs none; every multi-lane read has some.
The walker in §6 reports that configuration rather than walking it, because it is a specification error rather than a runtime one — the transaction cannot be made safe by being careful, only by being changed.
4. Watching the Reversal
Who drives, phase by phase
8 cyclesThe two Z cells are the whole chapter. They are not idle time and they are not waiting for the array — they are the interval in which neither side drives, and removing them does not make the transfer faster, it makes it broken.
5. The Sequence, End to End
Note the ordering of the two middle messages. The master's release comes first, and the dummy cycles are clocked with nobody driving. A controller that released only on the last dummy cycle would have driven against the device for all the earlier ones.
6. Building the Phase Walker — Three HDLs
The circuit
Circuit. A phase walker carrying a width and a direction per phase.
State. The remaining length of the current phase, the three phase lengths, and the last driving direction.
Datapath. Each phase's length is its bit count shifted right by the lane count's logarithm — the same shift-not-divide rule as Chapter 12.1. The direction is a small function of the phase and the transfer's direction.
Control. Four phases with empty ones skipped, which is Chapter 10.5's rule now carrying two more attributes per phase.
Clock and reset. System clock; asynchronous active-low reset.
Enables. turnaround pulses on a reversal and turnarounds counts them. no_turn_window reports §3's specification error.
Timing. One cycle per phase cycle; cur_lanes and cur_dir describe the phase in progress.
Synthesis. A down-counter, three length registers, a small state machine and a direction comparator.
Limitations. It walks and reports; it does not drive pins. The output-enable generation is Chapter 12.2's gearbox, and keeping them separate means one gearbox serves every phase shape.
The direction encoding is three-valued, and that is the design's one real subtlety. D_OUT, D_IN and D_Z — driving, receiving, and released. A two-valued encoding could not express the dummy phase at all, and the released state is the entire point.
Which makes the turnaround count non-obvious. A reversal is a change between driving and receiving; the released interval between them is not itself a change. So the comparison ignores D_Z — because counting transitions into and out of it would report two turnarounds where the bus reverses once. That is the kind of off-by-one that produces a plausible number.
// spi_phase_lanes.sv
//
// Chapter 12.3 -- per-phase lane widths, and the direction changes they
// imply.
//
// The x-y-z notation names the lane width of three phases: 1-1-4 is a
// command on one lane, an address on one lane and data on four. 1-4-4 and
// 4-4-4 widen more of the transfer, and the notation exists because a
// device supports some combinations and not others.
//
// What the notation does NOT say, and what this block makes explicit, is
// the DIRECTION of each phase -- and that is where the cost hides. On a
// single-lane read, MOSI and MISO are separate wires, so no line ever
// changes direction. On a multi-lane read every data pin is bidirectional:
// the master drives the command and address on lanes it must then RELEASE
// so the device can drive the data back.
//
// That release is the turnaround of Chapter 8.4, arriving on four lanes at
// once -- and it is the second reason a multi-lane read specifies dummy
// cycles. Chapter 11.2 gave the first: the array needs time. The dummy
// phase pays for both at once, which is why a quad read has dummy cycles
// even on parts whose arrays are fast enough not to need them.
//
// So a read with a dummy count of ZERO has no turnaround window at all:
// the master must stop driving and the device must start driving on the
// same edge, which is contention. The block reports that configuration
// rather than walking it, because it is a specification error and not a
// runtime one.
module spi_phase_lanes #(
parameter int CNT_W = 16,
parameter int CMD_BITS = 8
) (
input logic clk,
input logic rst_n,
input logic start,
input logic [3:0] cmd_lanes,
input logic [3:0] addr_lanes,
input logic [3:0] data_lanes,
input logic [2:0] addr_bytes,
input logic [5:0] dummy_cycles,
input logic [CNT_W-1:0] data_bytes,
input logic data_is_read, // direction of the data phase
output logic [1:0] phase, // 0 cmd 1 addr 2 dummy 3 data
output logic [3:0] cur_lanes, // width of the current phase
output logic [1:0] cur_dir, // 0 drive 1 receive 2 released
output logic turnaround, // pulse: direction changed
output logic [3:0] turnarounds, // count for this transfer
output logic busy,
output logic done,
output logic [CNT_W-1:0] cycles_elapsed,
output logic no_turn_window // a read with no dummy phase
);
localparam logic [1:0] P_CMD = 2'd0;
localparam logic [1:0] P_ADDR = 2'd1;
localparam logic [1:0] P_DUMMY = 2'd2;
localparam logic [1:0] P_DATA = 2'd3;
localparam logic [1:0] D_OUT = 2'd0; // master drives
localparam logic [1:0] D_IN = 2'd1; // device drives
localparam logic [1:0] D_Z = 2'd2; // nobody drives -- the window
logic [CNT_W-1:0] rem;
logic [CNT_W-1:0] len_addr, len_dummy, len_data;
logic [3:0] w_addr, w_data;
logic is_read;
logic [1:0] last_drive; // the last OUT or IN, ignoring Z
// Bit count divided by lanes, as a shift -- lane counts are powers of
// two, so this is never a divider.
function automatic logic [2:0] lane_shift(input logic [3:0] n);
case (n)
4'd1: lane_shift = 3'd0;
4'd2: lane_shift = 3'd1;
4'd4: lane_shift = 3'd2;
4'd8: lane_shift = 3'd3;
default: lane_shift = 3'd0;
endcase
endfunction
// The direction a phase runs in. Only the data phase can go either
// way; the dummy phase is RELEASED on a read, because that is what it
// is for, and driven on a write, because there is nothing to release.
function automatic logic [1:0] phase_dir(input logic [1:0] p,
input logic rd);
case (p)
P_CMD: phase_dir = D_OUT;
P_ADDR: phase_dir = D_OUT;
P_DUMMY: phase_dir = rd ? D_Z : D_OUT;
default: phase_dir = rd ? D_IN : D_OUT;
endcase
endfunction
logic [1:0] nxt_phase;
logic [CNT_W-1:0] nxt_len;
logic nxt_last;
// Where to go when the current phase runs out, skipping every phase of
// zero length -- the same rule as Chapter 10.5, now carrying a width
// and a direction with each phase.
always_comb begin
nxt_phase = phase;
nxt_len = {CNT_W{1'b0}};
nxt_last = 1'b0;
case (phase)
P_CMD: begin
if (len_addr != 0) begin nxt_phase = P_ADDR; nxt_len = len_addr; end
else if (len_dummy != 0) begin nxt_phase = P_DUMMY; nxt_len = len_dummy; end
else if (len_data != 0) begin nxt_phase = P_DATA; nxt_len = len_data; end
else nxt_last = 1'b1;
end
P_ADDR: begin
if (len_dummy != 0) begin nxt_phase = P_DUMMY; nxt_len = len_dummy; end
else if (len_data != 0) begin nxt_phase = P_DATA; nxt_len = len_data; end
else nxt_last = 1'b1;
end
P_DUMMY: begin
if (len_data != 0) begin nxt_phase = P_DATA; nxt_len = len_data; end
else nxt_last = 1'b1;
end
default: nxt_last = 1'b1;
endcase
end
// The width of whichever phase is current.
always_comb begin
case (phase)
P_CMD: cur_lanes = cmd_lanes;
P_ADDR: cur_lanes = w_addr;
P_DUMMY: cur_lanes = w_data; // the window is on the DATA lanes
default: cur_lanes = w_data;
endcase
end
assign cur_dir = phase_dir(phase, is_read);
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
phase <= P_CMD;
busy <= 1'b0;
done <= 1'b0;
rem <= {CNT_W{1'b0}};
len_addr <= {CNT_W{1'b0}};
len_dummy <= {CNT_W{1'b0}};
len_data <= {CNT_W{1'b0}};
w_addr <= 4'd1;
w_data <= 4'd1;
is_read <= 1'b0;
last_drive <= D_OUT;
turnaround <= 1'b0;
turnarounds <= 4'd0;
cycles_elapsed <= {CNT_W{1'b0}};
no_turn_window <= 1'b0;
end else begin
done <= 1'b0;
turnaround <= 1'b0;
if (start && !busy) begin
len_addr <= (CNT_W'(addr_bytes) << 3) >> lane_shift(addr_lanes);
len_dummy <= CNT_W'(dummy_cycles);
len_data <= (data_bytes << 3) >> lane_shift(data_lanes);
w_addr <= addr_lanes;
w_data <= data_lanes;
is_read <= data_is_read;
phase <= P_CMD;
rem <= CNT_W'(CMD_BITS) >> lane_shift(cmd_lanes);
busy <= 1'b1;
last_drive <= D_OUT; // the command always goes out
turnarounds <= 4'd0;
cycles_elapsed <= {CNT_W{1'b0}};
// A read whose data phase follows no dummy phase has no
// window in which the master can release the lanes before
// the device drives them. That is a specification error --
// reported here rather than walked.
no_turn_window <= data_is_read && (dummy_cycles == 6'd0);
end else if (busy) begin
cycles_elapsed <= cycles_elapsed + 1'b1;
if (rem == CNT_W'(1)) begin
if (nxt_last) begin
busy <= 1'b0;
done <= 1'b1;
end else begin
phase <= nxt_phase;
rem <= nxt_len;
// A turnaround is a change between DRIVING and
// RECEIVING. The released window between them is
// not itself a change, which is why the comparison
// ignores D_Z -- otherwise every read would count
// two turnarounds where the bus reverses once.
if (phase_dir(nxt_phase, is_read) != D_Z &&
phase_dir(nxt_phase, is_read) != last_drive) begin
turnaround <= 1'b1;
turnarounds <= turnarounds + 4'd1;
last_drive <= phase_dir(nxt_phase, is_read);
end
end
end else begin
rem <= rem - 1'b1;
end
end
end
end
endmodule// spi_phase_lanes_tb.sv
//
// The checks that matter are about DIRECTION, not cycle counts -- those
// were settled in Chapter 12.1. Here the properties are: a read reverses
// the bus exactly once, a write never reverses it, the master never drives
// during a read's dummy phase, and a read with no dummy phase is reported
// as having no turnaround window.
`timescale 1ns/1ps
module spi_phase_lanes_tb;
localparam int CNT_W = 16;
localparam int CMD_BITS = 8;
localparam logic [1:0] P_CMD = 2'd0;
localparam logic [1:0] P_ADDR = 2'd1;
localparam logic [1:0] P_DUMMY = 2'd2;
localparam logic [1:0] P_DATA = 2'd3;
localparam logic [1:0] D_OUT = 2'd0;
localparam logic [1:0] D_IN = 2'd1;
localparam logic [1:0] D_Z = 2'd2;
logic clk = 1'b0;
logic rst_n = 1'b0;
always #5 clk = ~clk;
logic start = 1'b0;
logic [3:0] cmd_lanes = 4'd1;
logic [3:0] addr_lanes = 4'd1;
logic [3:0] data_lanes = 4'd1;
logic [2:0] addr_bytes = 3'd3;
logic [5:0] dummy_cycles = 6'd8;
logic [CNT_W-1:0] data_bytes = 16'd4;
logic data_is_read = 1'b1;
logic [1:0] phase;
logic [3:0] cur_lanes;
logic [1:0] cur_dir;
logic turnaround;
logic [3:0] turnarounds;
logic busy, done;
logic [CNT_W-1:0] cycles_elapsed;
logic no_turn_window;
int errors = 0;
spi_phase_lanes #(.CNT_W(CNT_W), .CMD_BITS(CMD_BITS)) dut (
.clk(clk), .rst_n(rst_n), .start(start),
.cmd_lanes(cmd_lanes), .addr_lanes(addr_lanes),
.data_lanes(data_lanes), .addr_bytes(addr_bytes),
.dummy_cycles(dummy_cycles), .data_bytes(data_bytes),
.data_is_read(data_is_read),
.phase(phase), .cur_lanes(cur_lanes), .cur_dir(cur_dir),
.turnaround(turnaround), .turnarounds(turnarounds),
.busy(busy), .done(done), .cycles_elapsed(cycles_elapsed),
.no_turn_window(no_turn_window)
);
// Observations accumulated over a walk.
int n_cmd, n_addr, n_dummy, n_data;
int n_turn_pulses;
bit drove_in_dummy; // the safety violation
bit saw_in, saw_out;
int data_lane_seen;
always_ff @(posedge clk) begin
if (rst_n && busy) begin
case (phase)
P_CMD: n_cmd <= n_cmd + 1;
P_ADDR: n_addr <= n_addr + 1;
P_DUMMY: begin
n_dummy <= n_dummy + 1;
// THE SAFETY OBSERVATION. On a read the dummy phase is
// the window in which the master must have released the
// lanes. Driving through it is contention on four lanes.
if (data_is_read && cur_dir == D_OUT)
drove_in_dummy <= 1'b1;
end
default: begin
n_data <= n_data + 1;
data_lane_seen <= cur_lanes;
if (cur_dir == D_IN) saw_in <= 1'b1;
if (cur_dir == D_OUT) saw_out <= 1'b1;
end
endcase
if (turnaround) n_turn_pulses <= n_turn_pulses + 1;
end
end
task automatic walk(input int cl, input int al, input int dl,
input int ab, input int dum, input int len,
input bit rd);
int guard;
begin
@(negedge clk);
cmd_lanes = 4'(cl); addr_lanes = 4'(al); data_lanes = 4'(dl);
addr_bytes = 3'(ab); dummy_cycles = 6'(dum);
data_bytes = CNT_W'(len); data_is_read = rd;
n_cmd = 0; n_addr = 0; n_dummy = 0; n_data = 0;
n_turn_pulses = 0; drove_in_dummy = 1'b0;
saw_in = 1'b0; saw_out = 1'b0; data_lane_seen = 0;
start = 1'b1;
@(negedge clk);
start = 1'b0;
guard = 0;
while (busy && guard < 4000) begin
@(negedge clk);
guard++;
end
@(negedge clk);
end
endtask
initial begin
n_cmd = 0; n_addr = 0; n_dummy = 0; n_data = 0;
n_turn_pulses = 0; drove_in_dummy = 1'b0;
saw_in = 1'b0; saw_out = 1'b0; data_lane_seen = 0;
repeat (3) @(negedge clk);
rst_n = 1'b1;
@(negedge clk);
// 1. A SINGLE-LANE READ. One reversal, and the dummy phase is
// released rather than driven.
walk(1, 1, 1, 3, 8, 4, 1'b1);
if (turnarounds !== 4'd1) begin
$display(" FAIL: 1-1-1 read reported %0d turnarounds, expected 1",
turnarounds);
errors++;
end
if (drove_in_dummy) begin
$display(" FAIL: 1-1-1 read drove the lanes during the dummy phase");
errors++;
end
if (!saw_in || saw_out) begin
$display(" FAIL: 1-1-1 read's data phase direction was wrong");
errors++;
end
$display(" 1-1-1 read: cmd=%0d addr=%0d dummy=%0d data=%0d cycles, %0d turnaround(s)",
n_cmd, n_addr, n_dummy, n_data, turnarounds);
// 2. QUAD OUTPUT -- 1-1-4. The same single reversal, but now the
// released window is four lanes wide.
walk(1, 1, 4, 3, 8, 4, 1'b1);
if (turnarounds !== 4'd1) begin
$display(" FAIL: 1-1-4 read reported %0d turnarounds, expected 1",
turnarounds);
errors++;
end
if (data_lane_seen != 4) begin
$display(" FAIL: 1-1-4 data phase ran on %0d lanes, expected 4",
data_lane_seen);
errors++;
end
if (drove_in_dummy) begin
$display(" FAIL: 1-1-4 read drove four lanes during the dummy phase");
errors++;
end
$display(" 1-1-4 read: cmd=%0d addr=%0d dummy=%0d data=%0d cycles, data on %0d lanes",
n_cmd, n_addr, n_dummy, n_data, data_lane_seen);
// 3. QUAD I/O -- 1-4-4. The address widens too, so its phase is
// shorter, and the direction pattern is unchanged.
walk(1, 4, 4, 3, 8, 4, 1'b1);
if (n_addr != 6) begin
$display(" FAIL: 1-4-4 address phase took %0d cycles, expected 6",
n_addr);
errors++;
end
if (turnarounds !== 4'd1) begin
$display(" FAIL: 1-4-4 read reported %0d turnarounds, expected 1",
turnarounds);
errors++;
end
$display(" 1-4-4 read: cmd=%0d addr=%0d dummy=%0d data=%0d cycles",
n_cmd, n_addr, n_dummy, n_data);
// 4. FULL QUAD -- 4-4-4. Even the command is four lanes wide, so
// every phase is narrow and the reversal count is still one.
walk(4, 4, 4, 3, 8, 4, 1'b1);
if (n_cmd != 2) begin
$display(" FAIL: 4-4-4 command phase took %0d cycles, expected 2",
n_cmd);
errors++;
end
if (turnarounds !== 4'd1) begin
$display(" FAIL: 4-4-4 read reported %0d turnarounds, expected 1",
turnarounds);
errors++;
end
$display(" 4-4-4 read: cmd=%0d addr=%0d dummy=%0d data=%0d cycles",
n_cmd, n_addr, n_dummy, n_data);
// 5. A WRITE. The data phase goes the same way as the command, so
// the bus NEVER reverses -- which is why a page program has no
// dummy phase and needs none.
walk(1, 1, 4, 3, 0, 4, 1'b0);
if (turnarounds !== 4'd0) begin
$display(" FAIL: a write reported %0d turnarounds, expected 0",
turnarounds);
errors++;
end
if (n_dummy != 0) begin
$display(" FAIL: a write with no dummy count entered the dummy phase");
errors++;
end
if (!saw_out || saw_in) begin
$display(" FAIL: a write's data phase was not outbound");
errors++;
end
if (no_turn_window) begin
$display(" FAIL: a write was flagged as having no turnaround window");
errors++;
end
$display(" 1-1-4 write: cmd=%0d addr=%0d data=%0d cycles, %0d turnarounds -- the bus never reverses",
n_cmd, n_addr, n_data, turnarounds);
// 6. A write WITH a dummy phase still never reverses, and the dummy
// phase is DRIVEN rather than released -- there is nothing to
// release.
walk(4, 4, 4, 3, 4, 4, 1'b0);
if (turnarounds !== 4'd0) begin
$display(" FAIL: a write with a dummy phase reported %0d turnarounds",
turnarounds);
errors++;
end
if (n_dummy != 4) begin
$display(" FAIL: a write's 4-cycle dummy phase took %0d cycles",
n_dummy);
errors++;
end
$display(" 4-4-4 write with dummy: %0d dummy cycles, still %0d turnarounds",
n_dummy, turnarounds);
// 7. THE SPECIFICATION ERROR. A read whose data phase follows no
// dummy phase has no window: the master must stop driving and
// the device start driving on the same edge.
walk(1, 1, 4, 3, 0, 4, 1'b1);
if (!no_turn_window) begin
$display(" FAIL: a read with no dummy phase was not flagged");
errors++;
end
if (turnarounds !== 4'd1) begin
$display(" FAIL: the reversal still happens -- expected 1 turnaround, got %0d",
turnarounds);
errors++;
end
$display(" 1-1-4 read, 0 dummy: flagged no turnaround window, and the reversal still occurs");
// 8. A read WITH a dummy phase must not be flagged.
walk(1, 1, 4, 3, 1, 4, 1'b1);
if (no_turn_window) begin
$display(" FAIL: a read with one dummy cycle was flagged as having no window");
errors++;
end
$display(" 1-1-4 read, 1 dummy: not flagged -- one cycle is a window");
// 9. Every turnaround PULSE is matched by an increment, so the
// count and the pulses cannot disagree.
walk(1, 1, 4, 3, 8, 16, 1'b1);
if (n_turn_pulses != int'(turnarounds)) begin
$display(" FAIL: %0d turnaround pulses but a count of %0d",
n_turn_pulses, turnarounds);
errors++;
end
// 10. SWEEP. Across every width combination and both directions:
// a read reverses exactly once and a write never, the master
// never drives a read's dummy phase, and the per-phase cycles
// sum to the elapsed total.
for (int wi = 0; wi < 4; wi++) begin
int lw;
lw = (wi == 0) ? 1 : (wi == 1) ? 2 : (wi == 2) ? 4 : 8;
for (int rd = 0; rd < 2; rd++) begin
walk(lw, lw, lw, 3, 8, 8, rd[0]);
if (turnarounds !== (rd[0] ? 4'd1 : 4'd0)) begin
$display(" FAIL: %0d lanes, read=%0b gave %0d turnarounds",
lw, rd[0], turnarounds);
errors++;
end
if (drove_in_dummy) begin
$display(" FAIL: %0d lanes, read=%0b drove the dummy phase",
lw, rd[0]);
errors++;
end
if ((n_cmd + n_addr + n_dummy + n_data) != int'(cycles_elapsed)) begin
$display(" FAIL: %0d lanes, read=%0b phases sum to %0d, elapsed %0d",
lw, rd[0], n_cmd + n_addr + n_dummy + n_data,
cycles_elapsed);
errors++;
end
end
end
$display(" 8 width/direction combinations swept: reads reverse once, writes never, the dummy phase is never driven on a read, and the phases always sum to the elapsed total");
if (errors == 0)
$display("PASS: a read reverses the bus exactly once at every width and a write never reverses it, the master releases the lanes through a read's dummy phase rather than driving them, a write's dummy phase is driven because there is nothing to release, a read specified with no dummy phase is reported as having no turnaround window, and the per-phase cycle counts always sum to the elapsed total");
else
$display("FAIL: %0d error(s)", errors);
$finish;
end
endmoduleThe testbench's checks are about direction rather than cycle counts, because Chapter 12.1 settled the counts. Four properties carry it.
A read reverses exactly once, at every width. One turnaround for 1-1-1, 1-1-4, 1-4-4 and 4-4-4 alike — the count does not depend on how many lanes reverse, only on how many times.
A write never reverses. Zero turnarounds, and a write with a dummy phase still zero — with the dummy phase driven rather than released, because there is nothing to release. That pair of cases is what proves the direction logic is reasoning about the transfer rather than about the phase.
The master never drives during a read's dummy phase. This is the safety property, and it is checked by observation across every width: an assertion that the direction is D_Z there, in a testbench that would see D_OUT if the release were late.
A read with no dummy phase is flagged, and — the detail that makes the test complete — the reversal still happens. The walker does not pretend the transaction is safe by suppressing the turnaround; it reports the missing window and the reversal, because both are true.
A sweep over all eight width-and-direction combinations then confirms the counts, the safety property, and that the per-phase cycles sum to the elapsed total.
// spi_phase_lanes.v
//
// Chapter 12.3 -- per-phase lane widths and the direction changes they
// imply, in Verilog-2001.
//
// The x-y-z notation names the lane width of three phases: 1-1-4 is a
// command on one lane, an address on one lane and data on four. What the
// notation does NOT say is the DIRECTION of each phase, and that is where
// the cost hides.
//
// On a single-lane read, MOSI and MISO are separate wires, so no line ever
// changes direction. On a multi-lane read every data pin is bidirectional:
// the master drives the command and address on lanes it must then RELEASE
// so the device can drive the data back. That release is the turnaround of
// Chapter 8.4 arriving on four lanes at once -- and it is the second reason
// a multi-lane read specifies dummy cycles. Chapter 11.2 gave the first:
// the array needs time. The dummy phase pays for both.
//
// So a read with a dummy count of ZERO has no turnaround window: the
// master must stop driving and the device start driving on the same edge,
// which is contention. That is reported rather than walked, because it is a
// specification error and not a runtime one.
module spi_phase_lanes #(
parameter CNT_W = 16,
parameter CMD_BITS = 8
) (
input wire clk,
input wire rst_n,
input wire start,
input wire [3:0] cmd_lanes,
input wire [3:0] addr_lanes,
input wire [3:0] data_lanes,
input wire [2:0] addr_bytes,
input wire [5:0] dummy_cycles,
input wire [CNT_W-1:0] data_bytes,
input wire data_is_read, // direction of the data phase
output reg [1:0] phase, // 0 cmd 1 addr 2 dummy 3 data
output reg [3:0] cur_lanes, // width of the current phase
output wire [1:0] cur_dir, // 0 drive 1 receive 2 released
output reg turnaround, // pulse: direction changed
output reg [3:0] turnarounds, // count for this transfer
output reg busy,
output reg done,
output reg [CNT_W-1:0] cycles_elapsed,
output reg no_turn_window // a read with no dummy phase
);
localparam [1:0] P_CMD = 2'd0;
localparam [1:0] P_ADDR = 2'd1;
localparam [1:0] P_DUMMY = 2'd2;
localparam [1:0] P_DATA = 2'd3;
localparam [1:0] D_OUT = 2'd0; // master drives
localparam [1:0] D_IN = 2'd1; // device drives
localparam [1:0] D_Z = 2'd2; // nobody drives -- the window
reg [CNT_W-1:0] rem;
reg [CNT_W-1:0] len_addr, len_dummy, len_data;
reg [3:0] w_addr, w_data;
reg is_read;
reg [1:0] last_drive; // the last OUT or IN, ignoring Z
// Bit count divided by lanes, as a shift -- lane counts are powers of
// two, so this is never a divider.
function [2:0] lane_shift;
input [3:0] n;
begin
case (n)
4'd1: lane_shift = 3'd0;
4'd2: lane_shift = 3'd1;
4'd4: lane_shift = 3'd2;
4'd8: lane_shift = 3'd3;
default: lane_shift = 3'd0;
endcase
end
endfunction
// The direction a phase runs in. Only the data phase can go either way;
// the dummy phase is RELEASED on a read, because that is what it is
// for, and driven on a write, because there is nothing to release.
function [1:0] phase_dir;
input [1:0] p;
input rd;
begin
case (p)
P_CMD: phase_dir = D_OUT;
P_ADDR: phase_dir = D_OUT;
P_DUMMY: phase_dir = rd ? D_Z : D_OUT;
default: phase_dir = rd ? D_IN : D_OUT;
endcase
end
endfunction
reg [1:0] nxt_phase;
reg [CNT_W-1:0] nxt_len;
reg nxt_last;
// Where to go when the current phase runs out, skipping every phase of
// zero length -- now carrying a width and a direction with each phase.
always @(*) begin
nxt_phase = phase;
nxt_len = {CNT_W{1'b0}};
nxt_last = 1'b0;
case (phase)
P_CMD: begin
if (len_addr != 0) begin nxt_phase = P_ADDR; nxt_len = len_addr; end
else if (len_dummy != 0) begin nxt_phase = P_DUMMY; nxt_len = len_dummy; end
else if (len_data != 0) begin nxt_phase = P_DATA; nxt_len = len_data; end
else nxt_last = 1'b1;
end
P_ADDR: begin
if (len_dummy != 0) begin nxt_phase = P_DUMMY; nxt_len = len_dummy; end
else if (len_data != 0) begin nxt_phase = P_DATA; nxt_len = len_data; end
else nxt_last = 1'b1;
end
P_DUMMY: begin
if (len_data != 0) begin nxt_phase = P_DATA; nxt_len = len_data; end
else nxt_last = 1'b1;
end
default: nxt_last = 1'b1;
endcase
end
// The width of whichever phase is current.
always @(*) begin
case (phase)
P_CMD: cur_lanes = cmd_lanes;
P_ADDR: cur_lanes = w_addr;
P_DUMMY: cur_lanes = w_data; // the window is on the DATA lanes
default: cur_lanes = w_data;
endcase
end
assign cur_dir = phase_dir(phase, is_read);
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
phase <= P_CMD;
busy <= 1'b0;
done <= 1'b0;
rem <= {CNT_W{1'b0}};
len_addr <= {CNT_W{1'b0}};
len_dummy <= {CNT_W{1'b0}};
len_data <= {CNT_W{1'b0}};
w_addr <= 4'd1;
w_data <= 4'd1;
is_read <= 1'b0;
last_drive <= D_OUT;
turnaround <= 1'b0;
turnarounds <= 4'd0;
cycles_elapsed <= {CNT_W{1'b0}};
no_turn_window <= 1'b0;
end else begin
done <= 1'b0;
turnaround <= 1'b0;
if (start && !busy) begin
len_addr <= ({{(CNT_W-3){1'b0}}, addr_bytes} << 3) >> lane_shift(addr_lanes);
len_dummy <= {{(CNT_W-6){1'b0}}, dummy_cycles};
len_data <= (data_bytes << 3) >> lane_shift(data_lanes);
w_addr <= addr_lanes;
w_data <= data_lanes;
is_read <= data_is_read;
phase <= P_CMD;
rem <= CMD_BITS >> lane_shift(cmd_lanes);
busy <= 1'b1;
last_drive <= D_OUT; // the command always goes out
turnarounds <= 4'd0;
cycles_elapsed <= {CNT_W{1'b0}};
// A read whose data phase follows no dummy phase has no
// window in which the master can release the lanes before
// the device drives them -- a specification error.
no_turn_window <= data_is_read && (dummy_cycles == 6'd0);
end else if (busy) begin
cycles_elapsed <= cycles_elapsed + 1'b1;
if (rem == 1) begin
if (nxt_last) begin
busy <= 1'b0;
done <= 1'b1;
end else begin
phase <= nxt_phase;
rem <= nxt_len;
// A turnaround is a change between DRIVING and
// RECEIVING. The released window between them is not
// itself a change, which is why the comparison
// ignores D_Z -- otherwise every read would count
// two turnarounds where the bus reverses once.
if (phase_dir(nxt_phase, is_read) != D_Z &&
phase_dir(nxt_phase, is_read) != last_drive) begin
turnaround <= 1'b1;
turnarounds <= turnarounds + 4'd1;
last_drive <= phase_dir(nxt_phase, is_read);
end
end
end else begin
rem <= rem - 1'b1;
end
end
end
end
endmodule// spi_phase_lanes_tb.v
//
// The same checks as the SystemVerilog testbench, all about DIRECTION: a
// read reverses the bus exactly once, a write never reverses it, the master
// never drives during a read's dummy phase, and a read with no dummy phase
// is reported as having no turnaround window.
`timescale 1ns/1ps
module spi_phase_lanes_tb;
parameter CNT_W = 16;
parameter CMD_BITS = 8;
localparam [1:0] P_CMD = 2'd0;
localparam [1:0] P_ADDR = 2'd1;
localparam [1:0] P_DUMMY = 2'd2;
localparam [1:0] P_DATA = 2'd3;
localparam [1:0] D_OUT = 2'd0;
localparam [1:0] D_IN = 2'd1;
localparam [1:0] D_Z = 2'd2;
reg clk;
reg rst_n;
reg start;
reg [3:0] cmd_lanes;
reg [3:0] addr_lanes;
reg [3:0] data_lanes;
reg [2:0] addr_bytes;
reg [5:0] dummy_cycles;
reg [CNT_W-1:0] data_bytes;
reg data_is_read;
wire [1:0] phase;
wire [3:0] cur_lanes;
wire [1:0] cur_dir;
wire turnaround;
wire [3:0] turnarounds;
wire busy, done;
wire [CNT_W-1:0] cycles_elapsed;
wire no_turn_window;
integer errors;
integer guard;
integer wi, rd, lw;
initial begin
clk = 1'b0; rst_n = 1'b0; start = 1'b0;
cmd_lanes = 4'd1; addr_lanes = 4'd1; data_lanes = 4'd1;
addr_bytes = 3'd3; dummy_cycles = 6'd8;
data_bytes = 16'd4; data_is_read = 1'b1;
errors = 0;
end
always #5 clk = ~clk;
spi_phase_lanes #(.CNT_W(CNT_W), .CMD_BITS(CMD_BITS)) dut (
.clk(clk), .rst_n(rst_n), .start(start),
.cmd_lanes(cmd_lanes), .addr_lanes(addr_lanes),
.data_lanes(data_lanes), .addr_bytes(addr_bytes),
.dummy_cycles(dummy_cycles), .data_bytes(data_bytes),
.data_is_read(data_is_read),
.phase(phase), .cur_lanes(cur_lanes), .cur_dir(cur_dir),
.turnaround(turnaround), .turnarounds(turnarounds),
.busy(busy), .done(done), .cycles_elapsed(cycles_elapsed),
.no_turn_window(no_turn_window)
);
// Observations accumulated over a walk.
integer n_cmd, n_addr, n_dummy, n_data;
integer n_turn_pulses;
reg drove_in_dummy; // the safety violation
reg saw_in, saw_out;
integer data_lane_seen;
initial begin
n_cmd = 0; n_addr = 0; n_dummy = 0; n_data = 0;
n_turn_pulses = 0; drove_in_dummy = 1'b0;
saw_in = 1'b0; saw_out = 1'b0; data_lane_seen = 0;
end
always @(posedge clk) begin
if (rst_n && busy) begin
case (phase)
P_CMD: n_cmd <= n_cmd + 1;
P_ADDR: n_addr <= n_addr + 1;
P_DUMMY: begin
n_dummy <= n_dummy + 1;
// THE SAFETY OBSERVATION. On a read the dummy phase is
// the window in which the master must have released the
// lanes. Driving through it is contention on four lanes.
if (data_is_read && cur_dir == D_OUT)
drove_in_dummy <= 1'b1;
end
default: begin
n_data <= n_data + 1;
data_lane_seen <= cur_lanes;
if (cur_dir == D_IN) saw_in <= 1'b1;
if (cur_dir == D_OUT) saw_out <= 1'b1;
end
endcase
if (turnaround) n_turn_pulses <= n_turn_pulses + 1;
end
end
task walk;
input integer cl;
input integer al;
input integer dl;
input integer ab;
input integer dum;
input integer len;
input rdi;
begin
@(negedge clk);
cmd_lanes = cl[3:0]; addr_lanes = al[3:0]; data_lanes = dl[3:0];
addr_bytes = ab[2:0]; dummy_cycles = dum[5:0];
data_bytes = len[CNT_W-1:0]; data_is_read = rdi;
n_cmd = 0; n_addr = 0; n_dummy = 0; n_data = 0;
n_turn_pulses = 0; drove_in_dummy = 1'b0;
saw_in = 1'b0; saw_out = 1'b0; data_lane_seen = 0;
start = 1'b1;
@(negedge clk);
start = 1'b0;
guard = 0;
while (busy && guard < 4000) begin
@(negedge clk);
guard = guard + 1;
end
@(negedge clk);
end
endtask
initial begin
repeat (3) @(negedge clk);
rst_n = 1'b1;
@(negedge clk);
// 1. A SINGLE-LANE READ. One reversal, dummy phase released.
walk(1, 1, 1, 3, 8, 4, 1'b1);
if (turnarounds !== 4'd1) begin
$display(" FAIL: 1-1-1 read reported %0d turnarounds, expected 1",
turnarounds);
errors = errors + 1;
end
if (drove_in_dummy) begin
$display(" FAIL: 1-1-1 read drove the lanes during the dummy phase");
errors = errors + 1;
end
if (!saw_in || saw_out) begin
$display(" FAIL: 1-1-1 read's data phase direction was wrong");
errors = errors + 1;
end
$display(" 1-1-1 read: cmd=%0d addr=%0d dummy=%0d data=%0d cycles, %0d turnaround(s)",
n_cmd, n_addr, n_dummy, n_data, turnarounds);
// 2. QUAD OUTPUT -- 1-1-4.
walk(1, 1, 4, 3, 8, 4, 1'b1);
if (turnarounds !== 4'd1) begin
$display(" FAIL: 1-1-4 read reported %0d turnarounds, expected 1",
turnarounds);
errors = errors + 1;
end
if (data_lane_seen != 4) begin
$display(" FAIL: 1-1-4 data phase ran on %0d lanes, expected 4",
data_lane_seen);
errors = errors + 1;
end
if (drove_in_dummy) begin
$display(" FAIL: 1-1-4 read drove four lanes during the dummy phase");
errors = errors + 1;
end
$display(" 1-1-4 read: cmd=%0d addr=%0d dummy=%0d data=%0d cycles, data on %0d lanes",
n_cmd, n_addr, n_dummy, n_data, data_lane_seen);
// 3. QUAD I/O -- 1-4-4.
walk(1, 4, 4, 3, 8, 4, 1'b1);
if (n_addr != 6) begin
$display(" FAIL: 1-4-4 address phase took %0d cycles, expected 6",
n_addr);
errors = errors + 1;
end
if (turnarounds !== 4'd1) begin
$display(" FAIL: 1-4-4 read reported %0d turnarounds, expected 1",
turnarounds);
errors = errors + 1;
end
$display(" 1-4-4 read: cmd=%0d addr=%0d dummy=%0d data=%0d cycles",
n_cmd, n_addr, n_dummy, n_data);
// 4. FULL QUAD -- 4-4-4.
walk(4, 4, 4, 3, 8, 4, 1'b1);
if (n_cmd != 2) begin
$display(" FAIL: 4-4-4 command phase took %0d cycles, expected 2",
n_cmd);
errors = errors + 1;
end
if (turnarounds !== 4'd1) begin
$display(" FAIL: 4-4-4 read reported %0d turnarounds, expected 1",
turnarounds);
errors = errors + 1;
end
$display(" 4-4-4 read: cmd=%0d addr=%0d dummy=%0d data=%0d cycles",
n_cmd, n_addr, n_dummy, n_data);
// 5. A WRITE. The data phase goes the same way as the command, so
// the bus NEVER reverses -- which is why a page program has no
// dummy phase and needs none.
walk(1, 1, 4, 3, 0, 4, 1'b0);
if (turnarounds !== 4'd0) begin
$display(" FAIL: a write reported %0d turnarounds, expected 0",
turnarounds);
errors = errors + 1;
end
if (n_dummy != 0) begin
$display(" FAIL: a write with no dummy count entered the dummy phase");
errors = errors + 1;
end
if (!saw_out || saw_in) begin
$display(" FAIL: a write's data phase was not outbound");
errors = errors + 1;
end
if (no_turn_window) begin
$display(" FAIL: a write was flagged as having no turnaround window");
errors = errors + 1;
end
$display(" 1-1-4 write: cmd=%0d addr=%0d data=%0d cycles, %0d turnarounds -- the bus never reverses",
n_cmd, n_addr, n_data, turnarounds);
// 6. A write WITH a dummy phase still never reverses, and that
// phase is DRIVEN -- there is nothing to release.
walk(4, 4, 4, 3, 4, 4, 1'b0);
if (turnarounds !== 4'd0) begin
$display(" FAIL: a write with a dummy phase reported %0d turnarounds",
turnarounds);
errors = errors + 1;
end
if (n_dummy != 4) begin
$display(" FAIL: a write's 4-cycle dummy phase took %0d cycles",
n_dummy);
errors = errors + 1;
end
$display(" 4-4-4 write with dummy: %0d dummy cycles, still %0d turnarounds",
n_dummy, turnarounds);
// 7. THE SPECIFICATION ERROR: a read with no dummy phase.
walk(1, 1, 4, 3, 0, 4, 1'b1);
if (!no_turn_window) begin
$display(" FAIL: a read with no dummy phase was not flagged");
errors = errors + 1;
end
if (turnarounds !== 4'd1) begin
$display(" FAIL: the reversal still happens -- expected 1, got %0d",
turnarounds);
errors = errors + 1;
end
$display(" 1-1-4 read, 0 dummy: flagged no turnaround window, and the reversal still occurs");
// 8. A read WITH a dummy phase must not be flagged.
walk(1, 1, 4, 3, 1, 4, 1'b1);
if (no_turn_window) begin
$display(" FAIL: a read with one dummy cycle was flagged");
errors = errors + 1;
end
$display(" 1-1-4 read, 1 dummy: not flagged -- one cycle is a window");
// 9. Every turnaround PULSE is matched by an increment.
walk(1, 1, 4, 3, 8, 16, 1'b1);
if (n_turn_pulses != turnarounds) begin
$display(" FAIL: %0d turnaround pulses but a count of %0d",
n_turn_pulses, turnarounds);
errors = errors + 1;
end
// 10. SWEEP over every width and both directions.
for (wi = 0; wi < 4; wi = wi + 1) begin
if (wi == 0) lw = 1;
else if (wi == 1) lw = 2;
else if (wi == 2) lw = 4;
else lw = 8;
for (rd = 0; rd < 2; rd = rd + 1) begin
walk(lw, lw, lw, 3, 8, 8, rd[0]);
if (turnarounds !== (rd[0] ? 4'd1 : 4'd0)) begin
$display(" FAIL: %0d lanes, read=%0b gave %0d turnarounds",
lw, rd[0], turnarounds);
errors = errors + 1;
end
if (drove_in_dummy) begin
$display(" FAIL: %0d lanes, read=%0b drove the dummy phase",
lw, rd[0]);
errors = errors + 1;
end
if ((n_cmd + n_addr + n_dummy + n_data) != cycles_elapsed) begin
$display(" FAIL: %0d lanes, read=%0b phases sum to %0d, elapsed %0d",
lw, rd[0], n_cmd + n_addr + n_dummy + n_data,
cycles_elapsed);
errors = errors + 1;
end
end
end
$display(" 8 width/direction combinations swept: reads reverse once, writes never, the dummy phase is never driven on a read, and the phases always sum to the elapsed total");
if (errors == 0)
$display("PASS: a read reverses the bus exactly once at every width and a write never reverses it, the master releases the lanes through a read's dummy phase rather than driving them, a write's dummy phase is driven because there is nothing to release, a read specified with no dummy phase is reported as having no turnaround window, and the per-phase cycle counts always sum to the elapsed total");
else
$display("FAIL: %0d error(s)", errors);
$finish;
end
endmodule-- spi_phase_lanes.vhd
--
-- Chapter 12.3 -- per-phase lane widths and the direction changes they
-- imply, in VHDL.
--
-- The x-y-z notation names the lane width of three phases: 1-1-4 is a
-- command on one lane, an address on one lane and data on four. What the
-- notation does NOT say is the DIRECTION of each phase, and that is where
-- the cost hides.
--
-- On a single-lane read, MOSI and MISO are separate wires, so no line ever
-- changes direction. On a multi-lane read every data pin is bidirectional:
-- the master drives the command and address on lanes it must then RELEASE
-- so the device can drive the data back. That release is the turnaround of
-- Chapter 8.4 arriving on four lanes at once -- and it is the second reason
-- a multi-lane read specifies dummy cycles. Chapter 11.2 gave the first:
-- the array needs time. The dummy phase pays for both.
--
-- So a read with a dummy count of ZERO has no turnaround window, which is a
-- specification error and is reported rather than walked.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_phase_lanes is
generic (
CNT_W : positive := 16;
CMD_BITS : natural := 8
);
port (
clk : in std_logic;
rst_n : in std_logic;
start : in std_logic;
cmd_lanes : in unsigned(3 downto 0);
addr_lanes : in unsigned(3 downto 0);
data_lanes : in unsigned(3 downto 0);
addr_bytes : in unsigned(2 downto 0);
dummy_cycles : in unsigned(5 downto 0);
data_bytes : in unsigned(CNT_W - 1 downto 0);
data_is_read : in std_logic;
phase : out unsigned(1 downto 0); -- 0 cmd 1 addr 2 dummy 3 data
cur_lanes : out unsigned(3 downto 0);
cur_dir : out unsigned(1 downto 0); -- 0 drive 1 receive 2 released
turnaround : out std_logic;
turnarounds : out unsigned(3 downto 0);
busy : out std_logic;
done : out std_logic;
cycles_elapsed : out unsigned(CNT_W - 1 downto 0);
no_turn_window : out std_logic
);
end entity;
architecture rtl of spi_phase_lanes is
constant P_CMD : unsigned(1 downto 0) := "00";
constant P_ADDR : unsigned(1 downto 0) := "01";
constant P_DUMMY : unsigned(1 downto 0) := "10";
constant P_DATA : unsigned(1 downto 0) := "11";
constant D_OUT : unsigned(1 downto 0) := "00"; -- master drives
constant D_IN : unsigned(1 downto 0) := "01"; -- device drives
constant D_Z : unsigned(1 downto 0) := "10"; -- nobody -- the window
-- Bit count divided by lanes, as a power of two -- lane counts are
-- always 1, 2, 4 or 8, so this is never a divider.
function lane_div(n : unsigned(3 downto 0)) return natural is
begin
case to_integer(n) is
when 1 => return 1;
when 2 => return 2;
when 4 => return 4;
when 8 => return 8;
when others => return 1;
end case;
end function;
-- The direction a phase runs in. Only the data phase can go either way;
-- the dummy phase is RELEASED on a read, because that is what it is for,
-- and driven on a write, because there is nothing to release.
function phase_dir(p : unsigned(1 downto 0);
rd : std_logic) return unsigned is
begin
if p = P_CMD or p = P_ADDR then
return D_OUT;
elsif p = P_DUMMY then
if rd = '1' then return D_Z; else return D_OUT; end if;
else
if rd = '1' then return D_IN; else return D_OUT; end if;
end if;
end function;
signal ph_r : unsigned(1 downto 0) := P_CMD;
signal rem_cnt : unsigned(CNT_W - 1 downto 0) := (others => '0');
signal len_addr : unsigned(CNT_W - 1 downto 0) := (others => '0');
signal len_dum : unsigned(CNT_W - 1 downto 0) := (others => '0');
signal len_dat : unsigned(CNT_W - 1 downto 0) := (others => '0');
signal w_addr : unsigned(3 downto 0) := to_unsigned(1, 4);
signal w_data : unsigned(3 downto 0) := to_unsigned(1, 4);
signal is_read : std_logic := '0';
signal last_drv : unsigned(1 downto 0) := D_OUT;
signal busy_r : std_logic := '0';
signal done_r : std_logic := '0';
signal turn_r : std_logic := '0';
signal tcnt_r : unsigned(3 downto 0) := (others => '0');
signal elap_r : unsigned(CNT_W - 1 downto 0) := (others => '0');
signal ntw_r : std_logic := '0';
signal nxt_phase : unsigned(1 downto 0) := P_CMD;
signal nxt_len : unsigned(CNT_W - 1 downto 0) := (others => '0');
signal nxt_last : std_logic := '0';
begin
-- Where to go when the current phase runs out, skipping every phase of
-- zero length -- now carrying a width and a direction with each phase.
nxt : process (ph_r, len_addr, len_dum, len_dat)
begin
nxt_phase <= ph_r;
nxt_len <= (others => '0');
nxt_last <= '0';
if ph_r = P_CMD then
if len_addr /= 0 then
nxt_phase <= P_ADDR; nxt_len <= len_addr;
elsif len_dum /= 0 then
nxt_phase <= P_DUMMY; nxt_len <= len_dum;
elsif len_dat /= 0 then
nxt_phase <= P_DATA; nxt_len <= len_dat;
else
nxt_last <= '1';
end if;
elsif ph_r = P_ADDR then
if len_dum /= 0 then
nxt_phase <= P_DUMMY; nxt_len <= len_dum;
elsif len_dat /= 0 then
nxt_phase <= P_DATA; nxt_len <= len_dat;
else
nxt_last <= '1';
end if;
elsif ph_r = P_DUMMY then
if len_dat /= 0 then
nxt_phase <= P_DATA; nxt_len <= len_dat;
else
nxt_last <= '1';
end if;
else
nxt_last <= '1';
end if;
end process;
-- The width of whichever phase is current.
width : process (ph_r, cmd_lanes, w_addr, w_data)
begin
if ph_r = P_CMD then
cur_lanes <= cmd_lanes;
elsif ph_r = P_ADDR then
cur_lanes <= w_addr;
else
cur_lanes <= w_data; -- the window is on the DATA lanes
end if;
end process;
cur_dir <= phase_dir(ph_r, is_read);
walk : process (clk, rst_n)
begin
if rst_n = '0' then
ph_r <= P_CMD;
busy_r <= '0';
done_r <= '0';
rem_cnt <= (others => '0');
len_addr <= (others => '0');
len_dum <= (others => '0');
len_dat <= (others => '0');
w_addr <= to_unsigned(1, 4);
w_data <= to_unsigned(1, 4);
is_read <= '0';
last_drv <= D_OUT;
turn_r <= '0';
tcnt_r <= (others => '0');
elap_r <= (others => '0');
ntw_r <= '0';
elsif rising_edge(clk) then
done_r <= '0';
turn_r <= '0';
if start = '1' and busy_r = '0' then
len_addr <= to_unsigned((to_integer(addr_bytes) * 8) /
lane_div(addr_lanes), CNT_W);
len_dum <= resize(dummy_cycles, CNT_W);
len_dat <= to_unsigned((to_integer(data_bytes) * 8) /
lane_div(data_lanes), CNT_W);
w_addr <= addr_lanes;
w_data <= data_lanes;
is_read <= data_is_read;
ph_r <= P_CMD;
rem_cnt <= to_unsigned(CMD_BITS / lane_div(cmd_lanes), CNT_W);
busy_r <= '1';
last_drv <= D_OUT; -- the command always goes out
tcnt_r <= (others => '0');
elap_r <= (others => '0');
-- A read whose data phase follows no dummy phase has no
-- window in which the master can release the lanes before
-- the device drives them -- a specification error.
if data_is_read = '1' and dummy_cycles = 0 then
ntw_r <= '1';
else
ntw_r <= '0';
end if;
elsif busy_r = '1' then
elap_r <= elap_r + 1;
if rem_cnt = 1 then
if nxt_last = '1' then
busy_r <= '0';
done_r <= '1';
else
ph_r <= nxt_phase;
rem_cnt <= nxt_len;
-- A turnaround is a change between DRIVING and
-- RECEIVING. The released window between them is not
-- itself a change, which is why the comparison
-- ignores D_Z -- otherwise every read would count
-- two turnarounds where the bus reverses once.
if phase_dir(nxt_phase, is_read) /= D_Z and
phase_dir(nxt_phase, is_read) /= last_drv then
turn_r <= '1';
tcnt_r <= tcnt_r + 1;
last_drv <= phase_dir(nxt_phase, is_read);
end if;
end if;
else
rem_cnt <= rem_cnt - 1;
end if;
end if;
end if;
end process;
phase <= ph_r;
turnaround <= turn_r;
turnarounds <= tcnt_r;
busy <= busy_r;
done <= done_r;
cycles_elapsed <= elap_r;
no_turn_window <= ntw_r;
end architecture;-- spi_phase_lanes_tb.vhd
--
-- The same checks as the SystemVerilog and Verilog testbenches, all about
-- DIRECTION: a read reverses the bus exactly once, a write never reverses
-- it, the master never drives during a read's dummy phase, and a read with
-- no dummy phase is reported as having no turnaround window.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_phase_lanes_tb is
end entity;
architecture sim of spi_phase_lanes_tb is
constant CNT_W : positive := 16;
constant CMD_BITS : natural := 8;
constant P_CMD : unsigned(1 downto 0) := "00";
constant P_ADDR : unsigned(1 downto 0) := "01";
constant P_DUMMY : unsigned(1 downto 0) := "10";
constant P_DATA : unsigned(1 downto 0) := "11";
constant D_OUT : unsigned(1 downto 0) := "00";
constant D_IN : unsigned(1 downto 0) := "01";
constant D_Z : unsigned(1 downto 0) := "10";
signal clk : std_logic := '0';
signal rst_n : std_logic := '0';
signal halt : boolean := false;
signal start : std_logic := '0';
signal cmd_lanes : unsigned(3 downto 0) := to_unsigned(1, 4);
signal addr_lanes : unsigned(3 downto 0) := to_unsigned(1, 4);
signal data_lanes : unsigned(3 downto 0) := to_unsigned(1, 4);
signal addr_bytes : unsigned(2 downto 0) := to_unsigned(3, 3);
signal dummy_cycles : unsigned(5 downto 0) := to_unsigned(8, 6);
signal data_bytes : unsigned(CNT_W - 1 downto 0) := to_unsigned(4, CNT_W);
signal data_is_read : std_logic := '1';
signal phase : unsigned(1 downto 0);
signal cur_lanes : unsigned(3 downto 0);
signal cur_dir : unsigned(1 downto 0);
signal turnaround : std_logic;
signal turnarounds : unsigned(3 downto 0);
signal busy, done : std_logic;
signal cycles_elapsed : unsigned(CNT_W - 1 downto 0);
signal no_turn_window : std_logic;
-- Observations accumulated over a walk.
signal n_cmd, n_addr, n_dummy, n_data : natural := 0;
signal n_turn_pulses : natural := 0;
signal drove_in_dummy : std_logic := '0';
signal saw_in, saw_out : std_logic := '0';
signal data_lane_seen : natural := 0;
signal clr_obs : std_logic := '0';
signal errors : natural := 0;
begin
clk <= not clk after 5 ns when not halt else '0';
dut : entity work.spi_phase_lanes
generic map (CNT_W => CNT_W, CMD_BITS => CMD_BITS)
port map (
clk => clk, rst_n => rst_n, start => start,
cmd_lanes => cmd_lanes, addr_lanes => addr_lanes,
data_lanes => data_lanes, addr_bytes => addr_bytes,
dummy_cycles => dummy_cycles, data_bytes => data_bytes,
data_is_read => data_is_read,
phase => phase, cur_lanes => cur_lanes, cur_dir => cur_dir,
turnaround => turnaround, turnarounds => turnarounds,
busy => busy, done => done, cycles_elapsed => cycles_elapsed,
no_turn_window => no_turn_window
);
-- Observer. The clear arrives on its own signal because two processes
-- driving one signal is a multiple-driver error in VHDL.
observe : process (clk)
begin
if rising_edge(clk) then
if clr_obs = '1' then
n_cmd <= 0; n_addr <= 0; n_dummy <= 0; n_data <= 0;
n_turn_pulses <= 0; drove_in_dummy <= '0';
saw_in <= '0'; saw_out <= '0'; data_lane_seen <= 0;
elsif rst_n = '1' and busy = '1' then
if phase = P_CMD then
n_cmd <= n_cmd + 1;
elsif phase = P_ADDR then
n_addr <= n_addr + 1;
elsif phase = P_DUMMY then
n_dummy <= n_dummy + 1;
-- THE SAFETY OBSERVATION. On a read the dummy phase is
-- the window in which the master must have released the
-- lanes. Driving through it is contention on four lanes.
if data_is_read = '1' and cur_dir = D_OUT then
drove_in_dummy <= '1';
end if;
else
n_data <= n_data + 1;
data_lane_seen <= to_integer(cur_lanes);
if cur_dir = D_IN then saw_in <= '1'; end if;
if cur_dir = D_OUT then saw_out <= '1'; end if;
end if;
if turnaround = '1' then
n_turn_pulses <= n_turn_pulses + 1;
end if;
end if;
end if;
end process;
stim : process
variable errs : natural := 0;
variable guard : natural;
variable lw : natural;
procedure walk(cl : natural; al : natural; dl : natural;
ab : natural; dum : natural; len : natural;
rd : std_logic) is
begin
wait until falling_edge(clk);
clr_obs <= '1';
wait until falling_edge(clk);
clr_obs <= '0';
cmd_lanes <= to_unsigned(cl, 4);
addr_lanes <= to_unsigned(al, 4);
data_lanes <= to_unsigned(dl, 4);
addr_bytes <= to_unsigned(ab, 3);
dummy_cycles <= to_unsigned(dum, 6);
data_bytes <= to_unsigned(len, CNT_W);
data_is_read <= rd;
start <= '1';
wait until falling_edge(clk);
start <= '0';
guard := 0;
while busy = '1' and guard < 4000 loop
wait until falling_edge(clk);
guard := guard + 1;
end loop;
wait until falling_edge(clk);
end procedure;
begin
for k in 0 to 2 loop
wait until falling_edge(clk);
end loop;
rst_n <= '1';
wait until falling_edge(clk);
-- 1. A SINGLE-LANE READ. One reversal, dummy phase released.
walk(1, 1, 1, 3, 8, 4, '1');
if turnarounds /= 1 then
report " FAIL: 1-1-1 read did not report exactly one turnaround";
errs := errs + 1;
end if;
if drove_in_dummy = '1' then
report " FAIL: 1-1-1 read drove the lanes during the dummy phase";
errs := errs + 1;
end if;
if saw_in /= '1' or saw_out = '1' then
report " FAIL: 1-1-1 read's data phase direction was wrong";
errs := errs + 1;
end if;
report " 1-1-1 read: cmd=" & integer'image(n_cmd) & " addr=" &
integer'image(n_addr) & " dummy=" & integer'image(n_dummy) &
" data=" & integer'image(n_data) & " cycles, " &
integer'image(to_integer(turnarounds)) & " turnaround(s)";
-- 2. QUAD OUTPUT -- 1-1-4.
walk(1, 1, 4, 3, 8, 4, '1');
if turnarounds /= 1 then
report " FAIL: 1-1-4 read did not report exactly one turnaround";
errs := errs + 1;
end if;
if data_lane_seen /= 4 then
report " FAIL: 1-1-4 data phase did not run on four lanes";
errs := errs + 1;
end if;
if drove_in_dummy = '1' then
report " FAIL: 1-1-4 read drove four lanes during the dummy phase";
errs := errs + 1;
end if;
report " 1-1-4 read: cmd=" & integer'image(n_cmd) & " addr=" &
integer'image(n_addr) & " dummy=" & integer'image(n_dummy) &
" data=" & integer'image(n_data) & " cycles, data on " &
integer'image(data_lane_seen) & " lanes";
-- 3. QUAD I/O -- 1-4-4.
walk(1, 4, 4, 3, 8, 4, '1');
if n_addr /= 6 then
report " FAIL: 1-4-4 address phase did not take six cycles";
errs := errs + 1;
end if;
if turnarounds /= 1 then
report " FAIL: 1-4-4 read did not report exactly one turnaround";
errs := errs + 1;
end if;
report " 1-4-4 read: cmd=" & integer'image(n_cmd) & " addr=" &
integer'image(n_addr) & " dummy=" & integer'image(n_dummy) &
" data=" & integer'image(n_data) & " cycles";
-- 4. FULL QUAD -- 4-4-4.
walk(4, 4, 4, 3, 8, 4, '1');
if n_cmd /= 2 then
report " FAIL: 4-4-4 command phase did not take two cycles";
errs := errs + 1;
end if;
if turnarounds /= 1 then
report " FAIL: 4-4-4 read did not report exactly one turnaround";
errs := errs + 1;
end if;
report " 4-4-4 read: cmd=" & integer'image(n_cmd) & " addr=" &
integer'image(n_addr) & " dummy=" & integer'image(n_dummy) &
" data=" & integer'image(n_data) & " cycles";
-- 5. A WRITE. The bus NEVER reverses -- which is why a page program
-- has no dummy phase and needs none.
walk(1, 1, 4, 3, 0, 4, '0');
if turnarounds /= 0 then
report " FAIL: a write reported a turnaround"; errs := errs + 1;
end if;
if n_dummy /= 0 then
report " FAIL: a write with no dummy count entered the dummy phase";
errs := errs + 1;
end if;
if saw_out /= '1' or saw_in = '1' then
report " FAIL: a write's data phase was not outbound";
errs := errs + 1;
end if;
if no_turn_window = '1' then
report " FAIL: a write was flagged as having no turnaround window";
errs := errs + 1;
end if;
report " 1-1-4 write: cmd=" & integer'image(n_cmd) & " addr=" &
integer'image(n_addr) & " data=" & integer'image(n_data) &
" cycles, 0 turnarounds -- the bus never reverses";
-- 6. A write WITH a dummy phase still never reverses, and that
-- phase is DRIVEN -- there is nothing to release.
walk(4, 4, 4, 3, 4, 4, '0');
if turnarounds /= 0 then
report " FAIL: a write with a dummy phase reported a turnaround";
errs := errs + 1;
end if;
if n_dummy /= 4 then
report " FAIL: a write's four-cycle dummy phase was mis-counted";
errs := errs + 1;
end if;
report " 4-4-4 write with dummy: " & integer'image(n_dummy) &
" dummy cycles, still 0 turnarounds";
-- 7. THE SPECIFICATION ERROR: a read with no dummy phase.
walk(1, 1, 4, 3, 0, 4, '1');
if no_turn_window /= '1' then
report " FAIL: a read with no dummy phase was not flagged";
errs := errs + 1;
end if;
if turnarounds /= 1 then
report " FAIL: the reversal should still occur"; errs := errs + 1;
end if;
report " 1-1-4 read, 0 dummy: flagged no turnaround window, and the reversal still occurs";
-- 8. A read WITH a dummy phase must not be flagged.
walk(1, 1, 4, 3, 1, 4, '1');
if no_turn_window = '1' then
report " FAIL: a read with one dummy cycle was flagged";
errs := errs + 1;
end if;
report " 1-1-4 read, 1 dummy: not flagged -- one cycle is a window";
-- 9. Every turnaround PULSE is matched by an increment.
walk(1, 1, 4, 3, 8, 16, '1');
if n_turn_pulses /= to_integer(turnarounds) then
report " FAIL: the turnaround pulses and count disagree";
errs := errs + 1;
end if;
-- 10. SWEEP over every width and both directions.
for wi in 0 to 3 loop
case wi is
when 0 => lw := 1;
when 1 => lw := 2;
when 2 => lw := 4;
when others => lw := 8;
end case;
for rd in 0 to 1 loop
if rd = 1 then
walk(lw, lw, lw, 3, 8, 8, '1');
if turnarounds /= 1 then
report " FAIL: a read did not reverse exactly once";
errs := errs + 1;
end if;
else
walk(lw, lw, lw, 3, 8, 8, '0');
if turnarounds /= 0 then
report " FAIL: a write reversed the bus";
errs := errs + 1;
end if;
end if;
if drove_in_dummy = '1' then
report " FAIL: the dummy phase was driven on a read";
errs := errs + 1;
end if;
if (n_cmd + n_addr + n_dummy + n_data) /=
to_integer(cycles_elapsed) then
report " FAIL: the phases do not sum to the elapsed total";
errs := errs + 1;
end if;
end loop;
end loop;
report " 8 width/direction combinations swept: reads reverse once, writes never, the dummy phase is never driven on a read, and the phases always sum to the elapsed total";
errors <= errs;
if errs = 0 then
report "PASS: a read reverses the bus exactly once at every width and a write never reverses it, the master releases the lanes through a read's dummy phase rather than driving them, a write's dummy phase is driven because there is nothing to release, a read specified with no dummy phase is reported as having no turnaround window, and the per-phase cycle counts always sum to the elapsed total";
else
report "FAIL: " & integer'image(errs) & " error(s)" severity error;
end if;
halt <= true;
wait;
end process;
end architecture;Parity
All three implement the same walker: identical ports and generics, a three-valued direction, empty phases skipped, a turnaround count that ignores the released state, and no_turn_window for a read specified with no dummy phase. All three testbenches run the same ten scenarios and the same eight-combination sweep, reporting one turnaround for every read and zero for every write at every width.
7. Why a Verification Engineer Cares
// 1. THE SAFETY PROPERTY. The master must not drive during a read's
// dummy phase -- that interval IS the turnaround window, and driving
// through it is contention on every active lane at once.
a_released_in_dummy : assert property (
@(posedge clk) disable iff (!rst_n)
(busy && (phase == P_DUMMY) && is_read) |-> (cur_dir == D_Z))
else $error("the master drove the lanes during a read's dummy phase");
// 2. ONE REVERSAL PER READ, at every width. The count depends on how
// many TIMES the bus reverses, not on how many lanes reverse -- so it
// must be one whether the data phase is one lane or eight.
a_one_reversal : assert property (
@(posedge clk) disable iff (!rst_n)
(done && is_read) |-> (turnarounds == 1))
else $error("a read did not reverse the bus exactly once");
// 3. NO REVERSAL ON A WRITE. Everything is outbound, so there is nothing
// to turn around -- which is why a write's dummy count of zero is
// correct rather than an omission.
a_no_write_reversal : assert property (
@(posedge clk) disable iff (!rst_n)
(done && !is_read) |-> (turnarounds == 0))
else $error("a write reversed the bus");
// 4. THE COUNT IGNORES THE RELEASED STATE. Entering and leaving D_Z is
// one reversal, not two -- an off-by-one that would double every
// read's count and look plausible.
a_z_not_counted : assert property (
@(posedge clk) disable iff (!rst_n)
(turnaround) |-> (cur_dir != D_Z))
else $error("a turnaround was counted on entering the released state");
// 5. THE SPECIFICATION ERROR IS REPORTED, and the reversal is still
// reported too -- because both are true, and suppressing the
// turnaround would make an unsafe transaction look safe.
a_window_reported : assert property (
@(posedge clk) disable iff (!rst_n)
(start && data_is_read && (dummy_cycles == 0)) |=> no_turn_window)
else $error("a read with no dummy phase was not reported");Property 4 is the one worth extracting, because it is about modelling rather than about SPI. A three-valued signal invites counting transitions, and transitions through an intermediate state are not transitions between the endpoints. Any design that models a tri-state, a handover, or a mode change through a neutral position has this trap, and the symptom is always a count that is exactly double and entirely plausible.
Property 5 encodes a discipline worth naming. When a design detects an unsafe configuration, the temptation is to also suppress the unsafe behaviour — here, to not report a turnaround because the transaction should not happen. That makes the report and the behaviour disagree, and a consumer reading only one of them is misled. Report the error and describe what would happen.
Coverage must cross the width with the direction, because the interesting cases are asymmetric:
covergroup spi_phase_lanes_cg @(posedge clk iff start);
// The named width combinations, because those are what devices
// actually implement -- arbitrary triples do not exist.
cp_shape : coverpoint shape_class {
bins single = {S_111};
bins dual_out = {S_112};
bins dual_io = {S_122};
bins quad_out = {S_114};
bins quad_io = {S_144};
bins qpi = {S_444};
bins octal = {S_888};
}
cp_dir : coverpoint data_is_read { bins write = {0}; bins read = {1}; }
// The dummy count, because ZERO on a read is the specification error
// and a suite that always uses eight never reaches it.
cp_dummy : coverpoint dummy_cycles {
bins none = {0}; // legal on a write, an ERROR on a read
bins one = {1}; // the minimum viable window
bins typical = {[2:8]};
bins many = {[9:32]};
}
// The turnaround count, which must be exactly the direction bit.
cp_turns : coverpoint turnarounds {
bins none = {0};
bins one = {1};
illegal_bins impossible = {[2:15]};
}
// THE CROSSES THAT MATTER. Width against direction, because a suite
// testing reads at every width and writes only at one has tested the
// write path once. And dummy-zero against direction, because it is
// legal in one and an error in the other.
x_shape_dir : cross cp_shape, cp_dir;
x_dummy_dir : cross cp_dummy, cp_dir;
endgroupcp_turns' illegal_bins are doing real work here. Two or more turnarounds is structurally impossible for any transaction this walker describes, so declaring it illegal converts property 4's off-by-one into a coverage failure as well as an assertion failure — two independent detectors, which is the same technique Chapter 11.1 used for the impossible alignment combinations.
And x_dummy_dir is the cross that separates the legal zero from the illegal one. A dummy count of zero is correct on a write and a specification error on a read, so covering the dummy count without crossing it against direction cannot distinguish them.
8. Why an FPGA or ASIC Engineer Cares
Release the lanes at the START of the dummy phase, not at its end. The whole phase is the window. A controller that drives until the last dummy cycle has contended for every earlier one, and the symptom is intermittent corruption that worsens as the dummy count grows — the opposite of what intuition suggests.
Model direction with three values, not two. Driving, receiving and released. A two-valued encoding cannot represent the dummy phase, so a design using one must be driving something through it.
Count reversals ignoring the released state. Entering and leaving high-impedance is one reversal. Counting it as two produces a number that is plausible and double.
Reject a multi-lane read with no dummy cycles at configuration time. It cannot be made safe by careful timing — there is no interval. Catching it when the profile is loaded is Chapter 10.1's validation argument applied to a combination of fields rather than to one.
Do not add a dummy phase to a write. It costs cycles and buys nothing, because the bus never reverses. A controller that inserts one uniformly for all commands is slower for no reason.
Expose the turnaround count. One small counter answers "is this transaction shaped the way I think?" — and a read reporting zero reversals is a read whose data phase is going the wrong way.
9. Failure Signature — Corruption That Gets Worse With More Dummy Cycles
Symptom. A quad-output read returns occasional corrupted bytes. Increasing the dummy count from six to eight — the usual first response to a marginal read — makes it worse, not better. Decreasing it to four reduces the corruption. The device is rated well above the clock in use.
What "more dummy cycles makes it worse" establishes. This is the observation that solves the case, and it does so by contradiction. Dummy cycles exist to give the array time and the bus time; adding them can only help a transfer that is short of either. A transfer that degrades when given more of them is not short of time at all — something about the dummy phase itself is harmful.
Plausible mechanisms.
- The master drives the lanes through the dummy phase, releasing only at its end. Then every dummy cycle is a cycle of contention, and adding cycles adds contention. This fits the inverted response exactly.
- The master releases correctly but the device starts driving early, so the overlap is at the beginning. That would not scale with the dummy count.
- Exceeding the round trip, which would not improve with fewer dummy cycles.
- A wrong dummy count, which shifts data rather than corrupting it and would not respond to the count in this direction — a shift would change which bytes are wrong, not how many.
- Supply noise from four simultaneously switching outputs, which is real but would not correlate with the dummy count.
The discriminating observation. The inverse correlation is nearly conclusive on its own, but it can be confirmed without a scope. Set the dummy count to zero — if the corruption largely disappears while the data becomes shifted instead, the contention hypothesis is confirmed: with no dummy phase there is no interval to contend over, so the electrical fault is replaced by a pure counting fault.
With a scope, look at the current on the supply during the dummy cycles, or simply at the lane voltage: a contended lane sits at neither rail.
The fix. Release the output enable on the first cycle of the dummy phase. The gearbox of Chapter 12.2 already derives the enable from the width; what is missing is that the phase must gate it too — which is precisely the D_Z state §6's walker publishes.
Why the intuitive response made it worse. Because "marginal read, add dummy cycles" is correct for a single-lane read and for a multi-lane read whose turnaround is handled properly. It is exactly wrong when the dummy phase is being driven, and the engineer following good instinct is led further from the answer with each step. That inversion — a standard remedy that worsens the symptom — is itself the diagnostic, and recognising it is worth more than any measurement.
10. Common Misconceptions
11. Reason It Through
Work this before reading the answer.
A controller implements
1-1-4quad-output reads correctly. An engineer wants to move to4-4-4for the short-transfer gain Chapter 12.1 demonstrated, and reasons: "the data phase already works on four lanes, so widening the command and address is just using the same lanes for more of the transfer."What has been missed, and what new failure becomes possible?
Start with what is the same. The data phase is unchanged — four lanes, device-driven. The dummy phase is unchanged. The single reversal between them is unchanged.
Now what is different. In 1-1-4 the command and address go out on IO0 alone. In 4-4-4 they go out on all four lanes. So the master now drives all four lanes during the command and address, where before it drove one.
And here is what was missed. In 1-1-4, lanes IO1, IO2 and IO3 were never driven by the master at all — they were released for the entire transaction until the device took them. A controller could be sloppy about releasing them, because it never asserted them.
In 4-4-4 the master drives all four and must release all four. So a release bug that was latent in 1-1-4 becomes active: if the controller releases IO0 correctly at the dummy phase but leaves IO1..IO3 driven — because in the old mode they were never driven and the release logic was only ever exercised on IO0 — three of the four lanes now contend.
Which produces a specific and confusing symptom. One quarter of the bits are correct. The data on IO0 arrives cleanly and the data on the other three is contended, so every nibble has one good bit and three bad — and the corruption is regular, not random, which makes it look like a mapping error rather than an electrical one.
There is a second, subtler issue. 4-4-4 requires the device to be in a quad-command mode, entered by a command. So the sequence is:
1. in single-lane mode, send the enter-QPI command
2. from now on, EVERY command is four-lane -- including the one
that leaves QPI modeA controller that cannot send a four-lane command cannot get out again. If the controller enters QPI mode and then resets to single-lane (Chapter 12.2's recommended reset state), it can no longer talk to the device at all — it is sending one-lane commands to a device listening on four. The device is not broken and the controller is not broken, and nothing works.
Recovery requires either a controller that can issue four-lane commands from reset (undoing the safe default), or the device's hardware reset — which Chapter 12.2 noted may have been repurposed as IO3, or a power cycle.
The general lesson, and it is the one worth carrying. Widening a phase does not only change its width — it changes who drives it, and therefore which release paths are exercised. A mode that drives more lanes activates release logic that a narrower mode never tested. And a mode that changes how commands are framed changes how you leave it, which makes entering it a one-way door unless the exit was designed first.
12. Understanding Check
13. Summary
The x-y-z notation gives three widths and no directions, and the directions are where the cost is.
A single-lane read never reverses a wire, because MOSI and MISO are separate. A multi-lane read reverses all the active lanes at once, exactly once, between the address and the data — and the count is how many times rather than how many lanes.
The dummy phase is the turnaround window. That is the second reason it exists, alongside Chapter 11.2's array access time, and it explains the asymmetry in every flash command table: a write never reverses, so its dummy count of zero is correct rather than missing.
A multi-lane read specified with zero dummy cycles has no window at all, which is a specification error rather than a tight margin — and no real device offers one.
In hardware the direction needs three values — driving, receiving, released — because two cannot express the dummy phase. The turnaround count must ignore the released state, or every read reports twice the reversals it performs. And the lanes must be released at the start of the dummy phase: driving until its last cycle contends for every earlier one, which produces corruption that worsens as the dummy count grows — a standard remedy making the symptom worse, which is itself the diagnostic.
For verification, the safety property is that the master is released through a read's dummy phase, and the discipline worth naming is to report an unsafe configuration and still describe what would happen — because suppressing the behaviour makes the report and the reality disagree.
14. What Comes Next
Width multiplies the bits per edge. There is one more axis, and it multiplies the edges per period.
Chapter 12.4 — SDR, DDR Transfers, and Throughput covers double data rate: data on both edges of every SCLK period, independent of the width, so an octal DDR part moves sixteen bits per period. The chapter's finding is that the byte assembly is identical in both modes — the same group sequence must produce the same bytes — so DDR changes exactly one thing: how many SCLK periods those groups occupy. And because the dummy phase is counted in periods, it does not halve either, which puts DDR under the same bound Chapter 12.1 found for width.
Continue learning
Related tutorials
- Related topic
Bus Turnaround and the Contention Window
The overlap between one driver releasing and the next asserting: how long the gap really is once pad turn-off counts, why re-selecting the same device needs none, and the guard that enforces dead time only on a handover.
- Related topic
Why Wider SPI Exists
The bandwidth pressure that pushed flash past one data line, why four data lanes deliver far less than four times the speed, why the dummy phase becomes more prominent as the bus widens, and the model that computes the gain for any width combination.
- Related topic
Dual, Quad, and Octal SPI
How dedicated pins become bidirectional lanes, the bit-to-lane convention every multi-lane device shares and what getting it backwards produces, why the output enable becomes a safety signal, and the gearbox that serialises a byte across any width.
- Related topic
Launch and Sample Edges
One edge of each bit time places a bit on the wire, the other captures it, and they must never be the same edge. Why the separation is forced, why it buys half a period, and how RTL maps physical edges onto those roles.
