SPI · Module 12
Dual, Quad, and Octal SPI
How dedicated pins become bidirectional lanes, the bit-to-lane convention every multi-lane device shares and what getting it backwards produces, why the output enable becomes a safety signal, and the gearbox that serialises a byte across any width.
Chapter 12.1 established what width buys. This chapter is about what it costs, and the first cost appears before any data moves.
Single-lane SPI only ever had to decide bit order. With four lanes there is a second question — which bit goes on which lane — and getting it backwards turns
0xA5into0x5A. Why is that worse than a wrong answer that looks wrong?
Because 0x5A is a perfectly ordinary byte. A driver checking for a plausible value is satisfied, a checksum fails later, and the investigation begins somewhere else entirely.
1. Where the Lanes Come From
A standard serial flash package has eight pins, and in single-lane mode two of them do almost nothing:
single-lane use multi-lane use
───────────────── ────────────────
CS# CS# unchanged
SCLK SCLK unchanged
MOSI (DI) IO0 bidirectional
MISO (DO) IO1 bidirectional
WP# (write protect) IO2 bidirectional
HOLD# (or RESET#) IO3 bidirectional
VCC, GND VCC, GNDThe write-protect and hold pins become data lanes. That is the whole trick, and it explains two things that otherwise look arbitrary.
Why the modes must be enabled explicitly. Repurposing WP# means the write-protect function is gone while quad mode is active, so a device cannot simply default to it — a part that powered up in quad mode would have no write protection and would interpret a legacy controller's idle WP# level as data. Enabling quad mode is a deliberate act, usually a status-register bit, and on many parts a non-volatile one that survives power cycling.
Why IO0 and IO1 are MOSI and MISO. The lane numbering is not arbitrary: IO0 is the pin that was MOSI, so a single-lane transaction on a quad-capable part uses IO0 outbound and IO1 inbound exactly as before. Backwards compatibility is a property of the numbering.
2. The Bit-to-Lane Convention
With one lane, a byte is eight clocks and the only question is which bit goes first — most significant, by universal convention.
With four lanes, a byte is two clocks and there are two questions. The convention is uniform across every multi-lane device and has two parts:
1. the most significant GROUP goes first
2. within a group, the most significant bit goes on the
HIGHEST-numbered laneSo 0xA5 = 1010_0101 on four lanes:
clock 1: IO3=1 IO2=0 IO1=1 IO0=0 the high nibble, 0xA
clock 2: IO3=0 IO2=1 IO1=0 IO0=1 the low nibble, 0x5Both halves matter and they fail differently.
Getting the group order backwards sends the low nibble first, so 0xA5 becomes 0x5A. Every byte's nibbles are swapped.
Getting the within-group order backwards puts bit 7 on IO0, so the high nibble 1010 is transmitted as 0101 — and 0xA5 again reads as 0x5A, by a completely different mechanism. The two errors are distinguishable only on a byte whose nibbles are not each other's bit-reverse.
That is why §6's testbench checks 0x13 as well as 0xA5: 0x13 gives 1 then 3, and both error modes produce something other than that.
3. Time, Compressed
One byte, one lane and four
8 cyclesThe idle cells are not decoration. They are where Chapter 12.1's gain comes from and, equally, where its limit lives: those six clocks are recovered only if there is more data to send. On a transfer whose data phase was already short, most of what the figure shows as saved was never there to save.
4. What Each Width Buys — and Costs
width clocks/byte pins used as data what is given up
1 8 2 (MOSI, MISO) nothing
2 4 2 (IO0, IO1) full duplex
4 2 4 (IO0..IO3) WP#, HOLD#/RESET#
8 1 8 a wider packageThe second column is the gain and the fourth is the price. Two entries deserve comment.
Dual mode gives up full duplex and gains nothing in pins. It uses the same two pins as single-lane SPI, but both become unidirectional-at-a-time rather than one each way. So a dual transfer is half the clocks of a single-lane one and cannot send and receive simultaneously — which for a flash is irrelevant, because flash transactions are never symmetric anyway. Dual exists mainly as a fallback for parts or boards where IO2 and IO3 are unavailable.
Octal needs more pins than a standard package has. Eight data lanes plus clock, select and power do not fit an 8-pin part, so octal devices come in larger packages — and at that point the pin argument that motivated SPI in the first place (Chapter 1.1) is considerably weaker. Octal is used where the alternative is a parallel bus, not where the alternative is single-lane SPI.
5. The Output Enable Becomes a Safety Signal
Here is the consequence that separates multi-lane SPI from everything earlier in this track.
In single-lane SPI, MOSI is always driven by the master and MISO always by the slave. Neither line ever changes direction, so there is no question of who drives what.
In quad mode all four lanes are bidirectional. The master drives them for the command and address; the device drives them for the data. So a lane driven by the master while the device is also driving it is contention, with the currents Chapter 8.5 described — now on four lanes at once.
Two rules follow, and §6's gearbox enforces both:
Drive exactly as many lanes as the width calls for. A single-lane transfer on a quad-capable controller must drive IO0 and nothing else. Driving all four because the port is four wide means IO1 — the device's output — is being fought over on every transfer.
Drive nothing when idle. A controller that holds the lanes between transfers prevents the device from answering at all, and the symptom is a device that appears dead rather than one that appears contended.
The single Lane mask node feeding both the data path and the output enable is the design decision. Deriving the enable from the same mask that narrows the data makes it structurally impossible to drive a lane the width does not include — which is stronger than checking for it.
6. Building the Lane Gearbox — Three HDLs
The circuit
Circuit. One shift register serving every width, in both directions.
State. A transmit shift register, a receive assembly register, and a group counter.
Datapath. Transmit shifts left by the lane count each clock and presents the top lanes bits, aligned so bit 7 lands on the highest active lane. Receive shifts left by the same amount and admits the masked input at the bottom. The two are exact inverses, which is what the round trip in the testbench exploits.
Control. A group counter of 8 / lanes, held as a table rather than a division — four legal values, and the same table is where an illegal width is caught.
Clock and reset. System clock; asynchronous active-low reset.
Enables. io_oe carries exactly lanes bits while transmitting and zero when idle. It is derived from the same mask that narrows the data, so the two cannot disagree.
Timing. One group per clock; tx_done pulses on the last group and rx_valid when a byte completes.
Synthesis. Two 8-bit shift registers, a small barrel shifter for the alignment, and two counters. Tens of flip-flops.
Limitations. Byte-granular. A device transferring an odd number of nibbles — which Chapter 12.4 shows DDR can produce — needs the group count decoupled from the byte boundary.
// spi_lane_serdes.sv
//
// Chapter 12.2 -- the gearbox between a byte and a group of lanes.
//
// Single-lane SPI moves one bit per clock, so a byte is eight clocks and
// the bit order is the only question. With two, four or eight lanes a byte
// becomes four, two or one clock -- and a second question appears that has
// no single-lane equivalent: WHICH BIT GOES ON WHICH LANE.
//
// The convention is uniform across every multi-lane SPI device:
//
// * the most significant GROUP goes first, and
// * within a group, the most significant bit goes on the
// highest-numbered lane.
//
// So for 0xA5 = 1010_0101 on four lanes, the first clock carries the high
// nibble 1010 with IO3=1, IO2=0, IO1=1, IO0=0, and the second carries
// 0101. Getting the within-group order backwards gives a byte whose
// nibbles are bit-reversed -- 0xA5 becomes 0x5A -- which is a plausible
// value and not an obvious error.
//
// OUTPUT ENABLE IS THE SAFETY OUTPUT. In quad and octal mode every data
// pin is bidirectional, so a lane the master drives while the device is
// also driving it is contention with real current behind it (Chapter 8.5).
// A four-lane transfer must drive exactly four lanes -- never the eight the
// port is wide.
module spi_lane_serdes #(
parameter int MAX_LANES = 8
) (
input logic clk,
input logic rst_n,
input logic [3:0] lanes, // 1, 2, 4 or 8
// Transmit: load a byte, shift it out most significant group first.
input logic load,
input logic [7:0] tx_byte,
output logic [MAX_LANES-1:0] io_out,
output logic [MAX_LANES-1:0] io_oe, // exactly `lanes` bits set
output logic tx_active,
output logic tx_done,
// Receive: capture groups and assemble a byte.
input logic [MAX_LANES-1:0] io_in,
input logic rx_en,
output logic [7:0] rx_byte,
output logic rx_valid,
output logic [3:0] groups, // 8 / lanes, for this width
output logic lane_err
);
logic [7:0] tx_sr;
logic [3:0] tx_left;
logic [7:0] rx_sr;
logic [3:0] rx_left;
// 8 / lanes, as a table. The legal set is four values, so a table is
// both cheaper than a divider and the place an illegal width is caught.
function automatic logic [3:0] group_count(input logic [3:0] n);
case (n)
4'd1: group_count = 4'd8;
4'd2: group_count = 4'd4;
4'd4: group_count = 4'd2;
4'd8: group_count = 4'd1;
default: group_count = 4'd8; // flagged by lane_err
endcase
endfunction
function automatic logic bad_lanes(input logic [3:0] n);
bad_lanes = !((n == 4'd1) || (n == 4'd2) ||
(n == 4'd4) || (n == 4'd8));
endfunction
// The active-lane mask, and the alignment shift that brings the most
// significant group down to the low bits.
wire [MAX_LANES-1:0] lane_mask = MAX_LANES'((1 << lanes) - 1);
wire [3:0] top_shift = 4'd8 - lanes;
// The group currently presented: the top `lanes` bits of the shift
// register, with bit 7 landing on the highest active lane.
assign io_out = MAX_LANES'(tx_sr >> top_shift) & lane_mask;
assign io_oe = tx_active ? lane_mask : {MAX_LANES{1'b0}};
assign lane_err = bad_lanes(lanes);
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
tx_sr <= 8'h00;
tx_left <= 4'd0;
tx_active <= 1'b0;
tx_done <= 1'b0;
rx_sr <= 8'h00;
rx_left <= 4'd0;
rx_byte <= 8'h00;
rx_valid <= 1'b0;
groups <= 4'd8;
end else begin
tx_done <= 1'b0;
rx_valid <= 1'b0;
groups <= group_count(lanes);
if (load) begin
tx_sr <= tx_byte;
tx_left <= group_count(lanes);
tx_active <= 1'b1;
// A load also restarts the receive assembly, because a
// multi-lane transfer is full duplex only in the sense that
// the same clocks carry both directions -- the group count
// is shared.
rx_sr <= 8'h00;
rx_left <= group_count(lanes);
end else if (tx_active) begin
// Shift by a whole group. The bits that leave the top are
// the ones just presented on the lanes.
tx_sr <= tx_sr << lanes;
if (tx_left == 4'd1) begin
tx_active <= 1'b0;
tx_done <= 1'b1;
end else begin
tx_left <= tx_left - 4'd1;
end
end
// Receive assembles with the same convention in reverse: each
// group arrives on the low `lanes` bits of io_in and is shifted
// in below the ones already captured.
if (rx_en && (rx_left != 4'd0)) begin
rx_sr <= (rx_sr << lanes) | 8'(io_in & lane_mask);
if (rx_left == 4'd1) begin
rx_byte <= 8'((rx_sr << lanes) | 8'(io_in & lane_mask));
rx_valid <= 1'b1;
rx_left <= 4'd0;
end else begin
rx_left <= rx_left - 4'd1;
end
end
end
end
endmodule// spi_lane_serdes_tb.sv
//
// Two kinds of check. First the lane MAPPING is verified against values
// computed by hand from the convention, at every width -- because a
// mapping error produces a plausible byte rather than an obvious fault.
// Then the gearbox is looped back on itself and every byte value is
// round-tripped at every width, which is the property that catches a
// mapping error the examples happen to miss.
`timescale 1ns/1ps
module spi_lane_serdes_tb;
localparam int MAX_LANES = 8;
logic clk = 1'b0;
logic rst_n = 1'b0;
always #5 clk = ~clk;
logic [3:0] lanes = 4'd1;
logic load = 1'b0;
logic [7:0] tx_byte = 8'h00;
logic [MAX_LANES-1:0] io_out;
logic [MAX_LANES-1:0] io_oe;
logic tx_active, tx_done;
logic [MAX_LANES-1:0] io_in;
logic rx_en;
logic [7:0] rx_byte;
logic rx_valid;
logic [3:0] groups;
logic lane_err;
int errors = 0;
// A named value rather than a literal: a bit-select of a sized literal
// is not portable, and the pattern is referred to often enough to be
// worth naming.
localparam logic [7:0] PAT = 8'hA5;
// Loopback: what the gearbox drives is what it receives. rx_en follows
// tx_active so the two run on the same clocks.
assign io_in = io_out;
assign rx_en = tx_active;
spi_lane_serdes #(.MAX_LANES(MAX_LANES)) dut (
.clk(clk), .rst_n(rst_n), .lanes(lanes),
.load(load), .tx_byte(tx_byte),
.io_out(io_out), .io_oe(io_oe),
.tx_active(tx_active), .tx_done(tx_done),
.io_in(io_in), .rx_en(rx_en),
.rx_byte(rx_byte), .rx_valid(rx_valid),
.groups(groups), .lane_err(lane_err)
);
// Captured groups, for the mapping check.
logic [MAX_LANES-1:0] seen [0:7];
int n_seen;
bit saw_rx;
logic [7:0] got_byte;
logic [MAX_LANES-1:0] oe_union;
always_ff @(posedge clk) begin
if (rst_n && tx_active && n_seen < 8) begin
seen[n_seen] <= io_out;
n_seen <= n_seen + 1;
oe_union <= oe_union | io_oe;
end
if (rst_n && rx_valid) begin
saw_rx <= 1'b1;
got_byte <= rx_byte;
end
end
task automatic send(input int n, input logic [7:0] b);
int guard;
begin
@(negedge clk);
lanes = 4'(n); tx_byte = b;
n_seen = 0; saw_rx = 1'b0; oe_union = {MAX_LANES{1'b0}};
load = 1'b1;
@(negedge clk);
load = 1'b0;
guard = 0;
while (tx_active && guard < 32) begin
@(negedge clk);
guard++;
end
@(negedge clk);
end
endtask
initial begin
n_seen = 0; saw_rx = 1'b0; got_byte = 8'h00;
oe_union = {MAX_LANES{1'b0}};
repeat (3) @(negedge clk);
rst_n = 1'b1;
@(negedge clk);
// 1. SINGLE LANE. 0xA5 = 1010_0101, most significant bit first --
// the same order every earlier module used.
send(1, PAT);
if (n_seen != 8) begin
$display(" FAIL: one lane produced %0d groups, expected 8", n_seen);
errors++;
end
for (int i = 0; i < 8; i++) begin
if (seen[i][0] !== PAT[7-i]) begin
$display(" FAIL: 1 lane group %0d carried %0b, expected %0b",
i, seen[i][0], PAT[7-i]);
errors++;
end
end
$display(" 1 lane: 8 groups, MSB first: %0b%0b%0b%0b%0b%0b%0b%0b",
seen[0][0], seen[1][0], seen[2][0], seen[3][0],
seen[4][0], seen[5][0], seen[6][0], seen[7][0]);
// 2. FOUR LANES. The high nibble first, and within it bit 7 on the
// highest lane: 0xA5 gives 4'hA then 4'h5. A within-group
// reversal would give 4'h5 then 4'hA -- which reads as 0x5A and
// looks like a perfectly ordinary byte.
send(4, PAT);
if (n_seen != 2) begin
$display(" FAIL: four lanes produced %0d groups, expected 2", n_seen);
errors++;
end
if (seen[0][3:0] !== 4'hA || seen[1][3:0] !== 4'h5) begin
$display(" FAIL: 4 lanes gave 0x%01h then 0x%01h, expected A then 5",
seen[0][3:0], seen[1][3:0]);
errors++;
end
$display(" 4 lanes: 2 groups, 0x%01h then 0x%01h (IO3..IO0 per group)",
seen[0][3:0], seen[1][3:0]);
// 3. TWO LANES. Four groups of two bits, most significant pair
// first: 10, 10, 01, 01.
send(2, PAT);
if (n_seen != 4) begin
$display(" FAIL: two lanes produced %0d groups, expected 4", n_seen);
errors++;
end
if (seen[0][1:0] !== 2'b10 || seen[1][1:0] !== 2'b10 ||
seen[2][1:0] !== 2'b01 || seen[3][1:0] !== 2'b01) begin
$display(" FAIL: 2 lanes gave %0b %0b %0b %0b, expected 10 10 01 01",
seen[0][1:0], seen[1][1:0], seen[2][1:0], seen[3][1:0]);
errors++;
end
$display(" 2 lanes: 4 groups, %02b %02b %02b %02b",
seen[0][1:0], seen[1][1:0], seen[2][1:0], seen[3][1:0]);
// 4. EIGHT LANES. One group carrying the whole byte.
send(8, PAT);
if (n_seen != 1) begin
$display(" FAIL: eight lanes produced %0d groups, expected 1", n_seen);
errors++;
end
if (seen[0] !== 8'hA5) begin
$display(" FAIL: 8 lanes gave 0x%02h, expected 0xA5", seen[0]);
errors++;
end
$display(" 8 lanes: 1 group, 0x%02h", seen[0]);
// 5. OUTPUT ENABLE. Exactly `lanes` bits driven, never more. In
// quad and octal mode every data pin is bidirectional, so a lane
// driven when it should not be is contention with real current.
send(1, 8'hFF);
if (oe_union !== 8'b0000_0001) begin
$display(" FAIL: one lane drove oe=%08b, expected 00000001", oe_union);
errors++;
end
send(2, 8'hFF);
if (oe_union !== 8'b0000_0011) begin
$display(" FAIL: two lanes drove oe=%08b, expected 00000011", oe_union);
errors++;
end
send(4, 8'hFF);
if (oe_union !== 8'b0000_1111) begin
$display(" FAIL: four lanes drove oe=%08b, expected 00001111", oe_union);
errors++;
end
send(8, 8'hFF);
if (oe_union !== 8'b1111_1111) begin
$display(" FAIL: eight lanes drove oe=%08b, expected 11111111", oe_union);
errors++;
end
$display(" output enable: exactly 1, 2, 4 and 8 lanes driven -- never more");
// 6. The enable must be clear when idle, or the master holds the
// bus between transfers and the device cannot answer at all.
@(negedge clk);
if (io_oe !== {MAX_LANES{1'b0}}) begin
$display(" FAIL: lanes still driven while idle: %08b", io_oe);
errors++;
end
// 7. EXHAUSTIVE ROUND TRIP. For every width and every byte value,
// serialising and deserialising must return the byte. This is
// what catches a mapping error the hand-checked examples above
// happen to agree with -- 0xA5 and 0xFF are both symmetric under
// some wrong mappings.
for (int w = 0; w < 4; w++) begin
int nl;
nl = (w == 0) ? 1 : (w == 1) ? 2 : (w == 2) ? 4 : 8;
for (int b = 0; b < 256; b++) begin
send(nl, 8'(b));
if (!saw_rx) begin
$display(" FAIL: %0d lanes, 0x%02h produced no received byte",
nl, b);
errors++;
end else if (got_byte !== 8'(b)) begin
$display(" FAIL: %0d lanes round trip 0x%02h returned 0x%02h",
nl, b, got_byte);
errors++;
end
end
end
$display(" round trip: 1024 (width, byte) pairs recovered exactly");
// 8. A byte with NO symmetry, checked by hand at four lanes, to
// pin the mapping independently of the round trip -- which would
// pass even if BOTH directions were reversed consistently.
send(4, 8'h13);
if (seen[0][3:0] !== 4'h1 || seen[1][3:0] !== 4'h3) begin
$display(" FAIL: 0x13 gave 0x%01h then 0x%01h, expected 1 then 3",
seen[0][3:0], seen[1][3:0]);
errors++;
end
$display(" 0x13 at 4 lanes: 0x%01h then 0x%01h -- high nibble first",
seen[0][3:0], seen[1][3:0]);
// 9. An illegal width is reported rather than silently treated as
// one lane.
@(negedge clk);
lanes = 4'd3;
@(negedge clk);
if (!lane_err) begin
$display(" FAIL: three lanes was not reported as illegal"); errors++;
end
lanes = 4'd4;
@(negedge clk);
if (lane_err) begin
$display(" FAIL: four lanes was reported as illegal"); errors++;
end
$display(" lane validation: 3 rejected, 4 accepted");
if (errors == 0)
$display("PASS: the most significant group goes first and within a group the most significant bit goes on the highest-numbered lane, at one, two, four and eight lanes; exactly as many lanes are driven as the width calls for and none while idle; and serialising then deserialising returns the byte for all 256 values at all four widths");
else
$display("FAIL: %0d error(s)", errors);
$finish;
end
endmoduleThe testbench has two layers and both are necessary, which is the point worth taking from it.
The hand-checked mapping pins the convention against values computed from §2 rather than from the design: 0xA5 gives 0xA then 0x5 at four lanes, 10 10 01 01 at two, and the whole byte at eight. Those checks would catch a group-order error.
The exhaustive round trip then loops the gearbox's output back into its input and requires every one of the 256 byte values to survive at every one of the four widths — 1024 pairs. That catches a mapping error the examples happen to agree with, which is a real risk here because 0xA5 and 0xFF are both symmetric under some wrong mappings.
But a round trip alone would be insufficient, and the testbench says so: it would pass even if both directions were reversed consistently, because the two errors cancel. So the asymmetric byte 0x13 is checked by hand at four lanes — 1 then 3 — to pin the absolute direction that the round trip cannot.
Two properties then cover the safety side. The output enable must carry exactly one, two, four or eight bits according to the width — never the eight the port is wide — and it must be clear when idle, because a controller holding the lanes prevents the device from answering at all.
// spi_lane_serdes.v
//
// Chapter 12.2 -- the gearbox between a byte and a group of lanes, in
// Verilog-2001.
//
// With two, four or eight lanes a byte becomes four, two or one clock, and
// a question appears that has no single-lane equivalent: WHICH BIT GOES ON
// WHICH LANE. The convention is uniform across every multi-lane device:
//
// * the most significant GROUP goes first, and
// * within a group, the most significant bit goes on the
// highest-numbered lane.
//
// So 0xA5 on four lanes is 0xA then 0x5. Getting the within-group order
// backwards turns 0xA5 into 0x5A -- a plausible value, not an obvious
// error.
//
// OUTPUT ENABLE IS THE SAFETY OUTPUT. In quad and octal mode every data pin
// is bidirectional, so driving a lane the device is also driving is
// contention with real current behind it. A four-lane transfer must drive
// exactly four lanes, never the eight the port is wide.
module spi_lane_serdes #(
parameter MAX_LANES = 8
) (
input wire clk,
input wire rst_n,
input wire [3:0] lanes, // 1, 2, 4 or 8
// Transmit: load a byte, shift it out most significant group first.
input wire load,
input wire [7:0] tx_byte,
output wire [MAX_LANES-1:0] io_out,
output wire [MAX_LANES-1:0] io_oe, // exactly `lanes` bits set
output reg tx_active,
output reg tx_done,
// Receive: capture groups and assemble a byte.
input wire [MAX_LANES-1:0] io_in,
input wire rx_en,
output reg [7:0] rx_byte,
output reg rx_valid,
output reg [3:0] groups, // 8 / lanes, for this width
output wire lane_err
);
reg [7:0] tx_sr;
reg [3:0] tx_left;
reg [7:0] rx_sr;
reg [3:0] rx_left;
// 8 / lanes, as a table. The legal set is four values, so a table is
// both cheaper than a divider and the place an illegal width is caught.
function [3:0] group_count;
input [3:0] n;
begin
case (n)
4'd1: group_count = 4'd8;
4'd2: group_count = 4'd4;
4'd4: group_count = 4'd2;
4'd8: group_count = 4'd1;
default: group_count = 4'd8; // flagged by lane_err
endcase
end
endfunction
function bad_lanes;
input [3:0] n;
begin
bad_lanes = !((n == 4'd1) || (n == 4'd2) ||
(n == 4'd4) || (n == 4'd8));
end
endfunction
// The active-lane mask, and the alignment shift that brings the most
// significant group down to the low bits.
wire [MAX_LANES-1:0] lane_mask = (1 << lanes) - 1;
wire [3:0] top_shift = 4'd8 - lanes;
// The group currently presented: the top `lanes` bits of the shift
// register, with bit 7 landing on the highest active lane.
assign io_out = (tx_sr >> top_shift) & lane_mask;
assign io_oe = tx_active ? lane_mask : {MAX_LANES{1'b0}};
assign lane_err = bad_lanes(lanes);
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
tx_sr <= 8'h00;
tx_left <= 4'd0;
tx_active <= 1'b0;
tx_done <= 1'b0;
rx_sr <= 8'h00;
rx_left <= 4'd0;
rx_byte <= 8'h00;
rx_valid <= 1'b0;
groups <= 4'd8;
end else begin
tx_done <= 1'b0;
rx_valid <= 1'b0;
groups <= group_count(lanes);
if (load) begin
tx_sr <= tx_byte;
tx_left <= group_count(lanes);
tx_active <= 1'b1;
// A load also restarts the receive assembly: the group count
// is shared between the two directions.
rx_sr <= 8'h00;
rx_left <= group_count(lanes);
end else if (tx_active) begin
// Shift by a whole group. The bits that leave the top are
// the ones just presented on the lanes.
tx_sr <= tx_sr << lanes;
if (tx_left == 4'd1) begin
tx_active <= 1'b0;
tx_done <= 1'b1;
end else begin
tx_left <= tx_left - 4'd1;
end
end
// Receive assembles with the same convention in reverse: each
// group arrives on the low `lanes` bits of io_in and is shifted
// in below the ones already captured.
if (rx_en && (rx_left != 4'd0)) begin
rx_sr <= (rx_sr << lanes) | (io_in & lane_mask);
if (rx_left == 4'd1) begin
rx_byte <= (rx_sr << lanes) | (io_in & lane_mask);
rx_valid <= 1'b1;
rx_left <= 4'd0;
end else begin
rx_left <= rx_left - 4'd1;
end
end
end
end
endmodule// spi_lane_serdes_tb.v
//
// The same checks as the SystemVerilog testbench: the lane mapping
// verified against values computed by hand from the convention at every
// width, then an exhaustive loopback round trip of all 256 byte values at
// all four widths.
`timescale 1ns/1ps
module spi_lane_serdes_tb;
parameter MAX_LANES = 8;
// A named value rather than a literal: a bit-select of a sized literal
// is not portable, and the pattern is referred to often enough to name.
localparam [7:0] PAT = 8'hA5;
reg clk;
reg rst_n;
reg [3:0] lanes;
reg load;
reg [7:0] tx_byte;
wire [MAX_LANES-1:0] io_out;
wire [MAX_LANES-1:0] io_oe;
wire tx_active, tx_done;
wire [MAX_LANES-1:0] io_in;
wire rx_en;
wire [7:0] rx_byte;
wire rx_valid;
wire [3:0] groups;
wire lane_err;
integer errors;
integer i, w, b, nl, guard;
// Loopback: what the gearbox drives is what it receives.
assign io_in = io_out;
assign rx_en = tx_active;
initial begin
clk = 1'b0; rst_n = 1'b0;
lanes = 4'd1; load = 1'b0; tx_byte = 8'h00;
errors = 0;
end
always #5 clk = ~clk;
spi_lane_serdes #(.MAX_LANES(MAX_LANES)) dut (
.clk(clk), .rst_n(rst_n), .lanes(lanes),
.load(load), .tx_byte(tx_byte),
.io_out(io_out), .io_oe(io_oe),
.tx_active(tx_active), .tx_done(tx_done),
.io_in(io_in), .rx_en(rx_en),
.rx_byte(rx_byte), .rx_valid(rx_valid),
.groups(groups), .lane_err(lane_err)
);
// Captured groups, for the mapping check.
reg [MAX_LANES-1:0] seen [0:7];
integer n_seen;
reg saw_rx;
reg [7:0] got_byte;
reg [MAX_LANES-1:0] oe_union;
initial begin
n_seen = 0; saw_rx = 1'b0; got_byte = 8'h00;
oe_union = {MAX_LANES{1'b0}};
end
always @(posedge clk) begin
if (rst_n && tx_active && n_seen < 8) begin
seen[n_seen] <= io_out;
n_seen <= n_seen + 1;
oe_union <= oe_union | io_oe;
end
if (rst_n && rx_valid) begin
saw_rx <= 1'b1;
got_byte <= rx_byte;
end
end
task send;
input integer n;
input [7:0] bb;
begin
@(negedge clk);
lanes = n[3:0]; tx_byte = bb;
n_seen = 0; saw_rx = 1'b0; oe_union = {MAX_LANES{1'b0}};
load = 1'b1;
@(negedge clk);
load = 1'b0;
guard = 0;
while (tx_active && guard < 32) begin
@(negedge clk);
guard = guard + 1;
end
@(negedge clk);
end
endtask
initial begin
repeat (3) @(negedge clk);
rst_n = 1'b1;
@(negedge clk);
// 1. SINGLE LANE, most significant bit first.
send(1, PAT);
if (n_seen != 8) begin
$display(" FAIL: one lane produced %0d groups, expected 8", n_seen);
errors = errors + 1;
end
for (i = 0; i < 8; i = i + 1) begin
if (seen[i][0] !== PAT[7-i]) begin
$display(" FAIL: 1 lane group %0d carried %0b, expected %0b",
i, seen[i][0], PAT[7-i]);
errors = errors + 1;
end
end
$display(" 1 lane: 8 groups, MSB first: %0b%0b%0b%0b%0b%0b%0b%0b",
seen[0][0], seen[1][0], seen[2][0], seen[3][0],
seen[4][0], seen[5][0], seen[6][0], seen[7][0]);
// 2. FOUR LANES. High nibble first, bit 7 on the highest lane.
// A within-group reversal would give 5 then A -- reading as
// 0x5A, a perfectly ordinary byte.
send(4, PAT);
if (n_seen != 2) begin
$display(" FAIL: four lanes produced %0d groups, expected 2", n_seen);
errors = errors + 1;
end
if (seen[0][3:0] !== 4'hA || seen[1][3:0] !== 4'h5) begin
$display(" FAIL: 4 lanes gave 0x%01h then 0x%01h, expected A then 5",
seen[0][3:0], seen[1][3:0]);
errors = errors + 1;
end
$display(" 4 lanes: 2 groups, 0x%01h then 0x%01h (IO3..IO0 per group)",
seen[0][3:0], seen[1][3:0]);
// 3. TWO LANES: four groups of two bits, 10 10 01 01.
send(2, PAT);
if (n_seen != 4) begin
$display(" FAIL: two lanes produced %0d groups, expected 4", n_seen);
errors = errors + 1;
end
if (seen[0][1:0] !== 2'b10 || seen[1][1:0] !== 2'b10 ||
seen[2][1:0] !== 2'b01 || seen[3][1:0] !== 2'b01) begin
$display(" FAIL: 2 lanes gave %0b %0b %0b %0b, expected 10 10 01 01",
seen[0][1:0], seen[1][1:0], seen[2][1:0], seen[3][1:0]);
errors = errors + 1;
end
$display(" 2 lanes: 4 groups, %02b %02b %02b %02b",
seen[0][1:0], seen[1][1:0], seen[2][1:0], seen[3][1:0]);
// 4. EIGHT LANES: one group carrying the whole byte.
send(8, PAT);
if (n_seen != 1) begin
$display(" FAIL: eight lanes produced %0d groups, expected 1", n_seen);
errors = errors + 1;
end
if (seen[0] !== 8'hA5) begin
$display(" FAIL: 8 lanes gave 0x%02h, expected 0xA5", seen[0]);
errors = errors + 1;
end
$display(" 8 lanes: 1 group, 0x%02h", seen[0]);
// 5. OUTPUT ENABLE: exactly `lanes` bits driven, never more.
send(1, 8'hFF);
if (oe_union !== 8'b0000_0001) begin
$display(" FAIL: one lane drove oe=%08b, expected 00000001", oe_union);
errors = errors + 1;
end
send(2, 8'hFF);
if (oe_union !== 8'b0000_0011) begin
$display(" FAIL: two lanes drove oe=%08b, expected 00000011", oe_union);
errors = errors + 1;
end
send(4, 8'hFF);
if (oe_union !== 8'b0000_1111) begin
$display(" FAIL: four lanes drove oe=%08b, expected 00001111", oe_union);
errors = errors + 1;
end
send(8, 8'hFF);
if (oe_union !== 8'b1111_1111) begin
$display(" FAIL: eight lanes drove oe=%08b, expected 11111111", oe_union);
errors = errors + 1;
end
$display(" output enable: exactly 1, 2, 4 and 8 lanes driven -- never more");
// 6. Clear when idle, or the master holds the bus between transfers
// and the device cannot answer at all.
@(negedge clk);
if (io_oe !== {MAX_LANES{1'b0}}) begin
$display(" FAIL: lanes still driven while idle: %08b", io_oe);
errors = errors + 1;
end
// 7. EXHAUSTIVE ROUND TRIP over every width and every byte value.
for (w = 0; w < 4; w = w + 1) begin
if (w == 0) nl = 1;
else if (w == 1) nl = 2;
else if (w == 2) nl = 4;
else nl = 8;
for (b = 0; b < 256; b = b + 1) begin
send(nl, b[7:0]);
if (!saw_rx) begin
$display(" FAIL: %0d lanes, 0x%02h produced no received byte",
nl, b);
errors = errors + 1;
end else if (got_byte !== b[7:0]) begin
$display(" FAIL: %0d lanes round trip 0x%02h returned 0x%02h",
nl, b, got_byte);
errors = errors + 1;
end
end
end
$display(" round trip: 1024 (width, byte) pairs recovered exactly");
// 8. A byte with NO symmetry, to pin the mapping independently of
// the round trip -- which would pass even if BOTH directions
// were reversed consistently.
send(4, 8'h13);
if (seen[0][3:0] !== 4'h1 || seen[1][3:0] !== 4'h3) begin
$display(" FAIL: 0x13 gave 0x%01h then 0x%01h, expected 1 then 3",
seen[0][3:0], seen[1][3:0]);
errors = errors + 1;
end
$display(" 0x13 at 4 lanes: 0x%01h then 0x%01h -- high nibble first",
seen[0][3:0], seen[1][3:0]);
// 9. An illegal width is reported rather than treated as one lane.
@(negedge clk);
lanes = 4'd3;
@(negedge clk);
if (!lane_err) begin
$display(" FAIL: three lanes was not reported as illegal");
errors = errors + 1;
end
lanes = 4'd4;
@(negedge clk);
if (lane_err) begin
$display(" FAIL: four lanes was reported as illegal");
errors = errors + 1;
end
$display(" lane validation: 3 rejected, 4 accepted");
if (errors == 0)
$display("PASS: the most significant group goes first and within a group the most significant bit goes on the highest-numbered lane, at one, two, four and eight lanes; exactly as many lanes are driven as the width calls for and none while idle; and serialising then deserialising returns the byte for all 256 values at all four widths");
else
$display("FAIL: %0d error(s)", errors);
$finish;
end
endmodule-- spi_lane_serdes.vhd
--
-- Chapter 12.2 -- the gearbox between a byte and a group of lanes, in VHDL.
--
-- With two, four or eight lanes a byte becomes four, two or one clock, and
-- a question appears that has no single-lane equivalent: WHICH BIT GOES ON
-- WHICH LANE. The convention is uniform across every multi-lane device:
--
-- * the most significant GROUP goes first, and
-- * within a group, the most significant bit goes on the
-- highest-numbered lane.
--
-- So 0xA5 on four lanes is 0xA then 0x5. Getting the within-group order
-- backwards turns 0xA5 into 0x5A -- a plausible value, not an obvious
-- error.
--
-- OUTPUT ENABLE IS THE SAFETY OUTPUT. In quad and octal mode every data pin
-- is bidirectional, so driving a lane the device is also driving is
-- contention with real current behind it. A four-lane transfer must drive
-- exactly four lanes, never the eight the port is wide.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_lane_serdes is
generic (
MAX_LANES : positive := 8
);
port (
clk : in std_logic;
rst_n : in std_logic;
lanes : in unsigned(3 downto 0); -- 1, 2, 4 or 8
-- Transmit: load a byte, shift it out most significant group first.
load : in std_logic;
tx_byte : in unsigned(7 downto 0);
io_out : out unsigned(MAX_LANES - 1 downto 0);
io_oe : out unsigned(MAX_LANES - 1 downto 0);
tx_active : out std_logic;
tx_done : out std_logic;
-- Receive: capture groups and assemble a byte.
io_in : in unsigned(MAX_LANES - 1 downto 0);
rx_en : in std_logic;
rx_byte : out unsigned(7 downto 0);
rx_valid : out std_logic;
groups : out unsigned(3 downto 0); -- 8 / lanes, for this width
lane_err : out std_logic
);
end entity;
architecture rtl of spi_lane_serdes is
-- 8 / lanes, as a table. The legal set is four values, so a table is
-- both cheaper than a divider and the place an illegal width is caught.
function group_count(n : unsigned(3 downto 0)) return unsigned is
begin
case to_integer(n) is
when 1 => return to_unsigned(8, 4);
when 2 => return to_unsigned(4, 4);
when 4 => return to_unsigned(2, 4);
when 8 => return to_unsigned(1, 4);
when others => return to_unsigned(8, 4); -- flagged by lane_err
end case;
end function;
function bad_lanes(n : unsigned(3 downto 0)) return boolean is
begin
return not (to_integer(n) = 1 or to_integer(n) = 2 or
to_integer(n) = 4 or to_integer(n) = 8);
end function;
signal tx_sr : unsigned(7 downto 0) := (others => '0');
signal tx_left : unsigned(3 downto 0) := (others => '0');
signal rx_sr : unsigned(7 downto 0) := (others => '0');
signal rx_left : unsigned(3 downto 0) := (others => '0');
signal act_r : std_logic := '0';
signal done_r : std_logic := '0';
signal rxb_r : unsigned(7 downto 0) := (others => '0');
signal rxv_r : std_logic := '0';
signal grp_r : unsigned(3 downto 0) := to_unsigned(8, 4);
-- The active-lane mask and the alignment shift that brings the most
-- significant group down to the low bits. Initialised so the
-- combinational outputs below are defined before the first reset.
signal lane_mask : unsigned(MAX_LANES - 1 downto 0) := (others => '0');
signal top_shift : natural := 7;
begin
lane_mask <= to_unsigned(2 ** to_integer(lanes) - 1, MAX_LANES)
when not bad_lanes(lanes)
else to_unsigned(0, MAX_LANES);
top_shift <= 8 - to_integer(lanes) when not bad_lanes(lanes) else 7;
-- The group currently presented: the top `lanes` bits of the shift
-- register, with bit 7 landing on the highest active lane.
io_out <= resize(shift_right(tx_sr, top_shift), MAX_LANES) and lane_mask;
io_oe <= lane_mask when act_r = '1' else to_unsigned(0, MAX_LANES);
lane_err <= '1' when bad_lanes(lanes) else '0';
tx_active <= act_r;
tx_done <= done_r;
rx_byte <= rxb_r;
rx_valid <= rxv_r;
groups <= grp_r;
gear : process (clk, rst_n)
variable nxt : unsigned(7 downto 0);
begin
if rst_n = '0' then
tx_sr <= (others => '0');
tx_left <= (others => '0');
act_r <= '0';
done_r <= '0';
rx_sr <= (others => '0');
rx_left <= (others => '0');
rxb_r <= (others => '0');
rxv_r <= '0';
grp_r <= to_unsigned(8, 4);
elsif rising_edge(clk) then
done_r <= '0';
rxv_r <= '0';
grp_r <= group_count(lanes);
if load = '1' then
tx_sr <= tx_byte;
tx_left <= group_count(lanes);
act_r <= '1';
-- A load also restarts the receive assembly: the group count
-- is shared between the two directions.
rx_sr <= (others => '0');
rx_left <= group_count(lanes);
elsif act_r = '1' then
-- Shift by a whole group. The bits that leave the top are
-- the ones just presented on the lanes.
tx_sr <= shift_left(tx_sr, to_integer(lanes));
if tx_left = 1 then
act_r <= '0';
done_r <= '1';
else
tx_left <= tx_left - 1;
end if;
end if;
-- Receive assembles with the same convention in reverse: each
-- group arrives on the low `lanes` bits of io_in and is shifted
-- in below the ones already captured.
if rx_en = '1' and rx_left /= 0 then
nxt := shift_left(rx_sr, to_integer(lanes)) or
resize(io_in and lane_mask, 8);
rx_sr <= nxt;
if rx_left = 1 then
rxb_r <= nxt;
rxv_r <= '1';
rx_left <= (others => '0');
else
rx_left <= rx_left - 1;
end if;
end if;
end if;
end process;
end architecture;-- spi_lane_serdes_tb.vhd
--
-- The same checks as the SystemVerilog and Verilog testbenches: the lane
-- mapping verified against values computed by hand from the convention at
-- every width, then an exhaustive loopback round trip of all 256 byte
-- values at all four widths.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_lane_serdes_tb is
end entity;
architecture sim of spi_lane_serdes_tb is
constant MAX_LANES : positive := 8;
constant PAT : unsigned(7 downto 0) := x"A5";
signal clk : std_logic := '0';
signal rst_n : std_logic := '0';
signal halt : boolean := false;
signal lanes : unsigned(3 downto 0) := to_unsigned(1, 4);
signal load : std_logic := '0';
signal tx_byte : unsigned(7 downto 0) := (others => '0');
signal io_out : unsigned(MAX_LANES - 1 downto 0);
signal io_oe : unsigned(MAX_LANES - 1 downto 0);
signal tx_active : std_logic;
signal tx_done : std_logic;
signal io_in : unsigned(MAX_LANES - 1 downto 0);
signal rx_en : std_logic;
signal rx_byte : unsigned(7 downto 0);
signal rx_valid : std_logic;
signal groups : unsigned(3 downto 0);
signal lane_err : std_logic;
-- Captured groups, for the mapping check.
type grp_arr is array (0 to 7) of unsigned(MAX_LANES - 1 downto 0);
signal seen : grp_arr;
signal n_seen : natural := 0;
signal saw_rx : std_logic := '0';
signal got_byte : unsigned(7 downto 0) := (others => '0');
signal oe_union : unsigned(MAX_LANES - 1 downto 0) := (others => '0');
signal clr_cap : std_logic := '0';
signal errors : natural := 0;
begin
clk <= not clk after 5 ns when not halt else '0';
-- Loopback: what the gearbox drives is what it receives.
io_in <= io_out;
rx_en <= tx_active;
dut : entity work.spi_lane_serdes
generic map (MAX_LANES => MAX_LANES)
port map (
clk => clk, rst_n => rst_n, lanes => lanes,
load => load, tx_byte => tx_byte,
io_out => io_out, io_oe => io_oe,
tx_active => tx_active, tx_done => tx_done,
io_in => io_in, rx_en => rx_en,
rx_byte => rx_byte, rx_valid => rx_valid,
groups => groups, lane_err => lane_err
);
-- Capture. The clear arrives on its own signal because two processes
-- driving one signal is a multiple-driver error in VHDL.
capture : process (clk)
begin
if rising_edge(clk) then
if clr_cap = '1' then
n_seen <= 0;
saw_rx <= '0';
oe_union <= (others => '0');
elsif rst_n = '1' then
if tx_active = '1' and n_seen < 8 then
seen(n_seen) <= io_out;
n_seen <= n_seen + 1;
oe_union <= oe_union or io_oe;
end if;
if rx_valid = '1' then
saw_rx <= '1';
got_byte <= rx_byte;
end if;
end if;
end if;
end process;
stim : process
variable errs : natural := 0;
variable guard : natural;
variable nl : natural;
procedure send(n : natural; bb : unsigned(7 downto 0)) is
begin
wait until falling_edge(clk);
clr_cap <= '1';
wait until falling_edge(clk);
clr_cap <= '0';
lanes <= to_unsigned(n, 4);
tx_byte <= bb;
load <= '1';
wait until falling_edge(clk);
load <= '0';
guard := 0;
while tx_active = '1' and guard < 32 loop
wait until falling_edge(clk);
guard := guard + 1;
end loop;
wait until falling_edge(clk);
end procedure;
begin
for k in 0 to 2 loop
wait until falling_edge(clk);
end loop;
rst_n <= '1';
wait until falling_edge(clk);
-- 1. SINGLE LANE, most significant bit first.
send(1, PAT);
if n_seen /= 8 then
report " FAIL: one lane produced " & integer'image(n_seen) &
" groups, expected 8";
errs := errs + 1;
end if;
for i in 0 to 7 loop
if seen(i)(0) /= PAT(7 - i) then
report " FAIL: one-lane group " & integer'image(i) &
" carried the wrong bit";
errs := errs + 1;
end if;
end loop;
report " 1 lane: 8 groups, MSB first";
-- 2. FOUR LANES. High nibble first, bit 7 on the highest lane.
-- A within-group reversal would give 5 then A -- reading as
-- 0x5A, a perfectly ordinary byte.
send(4, PAT);
if n_seen /= 2 then
report " FAIL: four lanes produced " & integer'image(n_seen) &
" groups, expected 2";
errs := errs + 1;
end if;
if seen(0)(3 downto 0) /= x"A" or seen(1)(3 downto 0) /= x"5" then
report " FAIL: four lanes did not give A then 5";
errs := errs + 1;
end if;
-- integer'image prints decimal, so the values are reported as
-- decimal rather than given a misleading 0x prefix: the high nibble
-- of 0xA5 is 10, the low nibble 5.
report " 4 lanes: 2 groups, high nibble " &
integer'image(to_integer(seen(0)(3 downto 0))) &
" then low nibble " &
integer'image(to_integer(seen(1)(3 downto 0)));
-- 3. TWO LANES: four groups of two bits, 10 10 01 01.
send(2, PAT);
if n_seen /= 4 then
report " FAIL: two lanes produced " & integer'image(n_seen) &
" groups, expected 4";
errs := errs + 1;
end if;
if seen(0)(1 downto 0) /= "10" or seen(1)(1 downto 0) /= "10" or
seen(2)(1 downto 0) /= "01" or seen(3)(1 downto 0) /= "01" then
report " FAIL: two lanes did not give 10 10 01 01";
errs := errs + 1;
end if;
report " 2 lanes: 4 groups, 10 10 01 01";
-- 4. EIGHT LANES: one group carrying the whole byte.
send(8, PAT);
if n_seen /= 1 then
report " FAIL: eight lanes produced " & integer'image(n_seen) &
" groups, expected 1";
errs := errs + 1;
end if;
if seen(0) /= x"A5" then
report " FAIL: eight lanes did not give 0xA5"; errs := errs + 1;
end if;
report " 8 lanes: 1 group, 0xA5";
-- 5. OUTPUT ENABLE: exactly `lanes` bits driven, never more.
send(1, x"FF");
if oe_union /= "00000001" then
report " FAIL: one lane drove the wrong enable mask";
errs := errs + 1;
end if;
send(2, x"FF");
if oe_union /= "00000011" then
report " FAIL: two lanes drove the wrong enable mask";
errs := errs + 1;
end if;
send(4, x"FF");
if oe_union /= "00001111" then
report " FAIL: four lanes drove the wrong enable mask";
errs := errs + 1;
end if;
send(8, x"FF");
if oe_union /= "11111111" then
report " FAIL: eight lanes drove the wrong enable mask";
errs := errs + 1;
end if;
report " output enable: exactly 1, 2, 4 and 8 lanes driven -- never more";
-- 6. Clear when idle, or the master holds the bus between transfers.
wait until falling_edge(clk);
if io_oe /= 0 then
report " FAIL: lanes still driven while idle"; errs := errs + 1;
end if;
-- 7. EXHAUSTIVE ROUND TRIP over every width and every byte value.
for w in 0 to 3 loop
case w is
when 0 => nl := 1;
when 1 => nl := 2;
when 2 => nl := 4;
when others => nl := 8;
end case;
for b in 0 to 255 loop
send(nl, to_unsigned(b, 8));
if saw_rx /= '1' then
report " FAIL: a transfer produced no received byte";
errs := errs + 1;
elsif to_integer(got_byte) /= b then
report " FAIL: a round trip returned the wrong byte";
errs := errs + 1;
end if;
end loop;
end loop;
report " round trip: 1024 (width, byte) pairs recovered exactly";
-- 8. A byte with NO symmetry, to pin the mapping independently of
-- the round trip -- which would pass even if BOTH directions
-- were reversed consistently.
send(4, x"13");
if seen(0)(3 downto 0) /= x"1" or seen(1)(3 downto 0) /= x"3" then
report " FAIL: 0x13 did not give 1 then 3"; errs := errs + 1;
end if;
report " 0x13 at 4 lanes: 0x1 then 0x3 -- high nibble first";
-- 9. An illegal width is reported rather than treated as one lane.
wait until falling_edge(clk);
lanes <= to_unsigned(3, 4);
wait until falling_edge(clk);
if lane_err /= '1' then
report " FAIL: three lanes was not reported as illegal";
errs := errs + 1;
end if;
lanes <= to_unsigned(4, 4);
wait until falling_edge(clk);
if lane_err = '1' then
report " FAIL: four lanes was reported as illegal";
errs := errs + 1;
end if;
report " lane validation: 3 rejected, 4 accepted";
errors <= errs;
if errs = 0 then
report "PASS: the most significant group goes first and within a group the most significant bit goes on the highest-numbered lane, at one, two, four and eight lanes; exactly as many lanes are driven as the width calls for and none while idle; and serialising then deserialising returns the byte for all 256 values at all four widths";
else
report "FAIL: " & integer'image(errs) & " error(s)" severity error;
end if;
halt <= true;
wait;
end process;
end architecture;Parity
All three implement the same gearbox: identical ports and generics, most significant group first with bit 7 on the highest active lane, a group count held as a table, an output enable derived from the lane mask and cleared when idle, and an illegal width reported. All three testbenches check the same hand-computed mappings at all four widths, run the same 1024-pair exhaustive round trip, and verify the same asymmetric byte — reporting identical results throughout.
7. Why a Verification Engineer Cares
// 1. ROUND TRIP. Serialising then deserialising returns the byte. Cheap,
// exhaustive over 256 values, and it catches field overlap -- the same
// property Chapter 10.4 used for a command codec.
a_round_trip : assert property (
@(posedge clk) disable iff (!rst_n)
(rx_valid) |-> (rx_byte == $past(tx_byte, tx_latency)))
else $error("the byte did not survive the round trip");
// 2. ...AND WHAT IT CANNOT PROVE. A round trip passes if BOTH directions
// are reversed consistently, because the errors cancel. So the
// absolute mapping needs its own property, on a value whose nibbles
// are not each other's bit-reverse.
a_absolute_mapping : assert property (
@(posedge clk) disable iff (!rst_n)
(first_group && lanes == 4) |->
(io_out[3:0] == tx_byte[7:4]))
else $error("the first group is not the high nibble");
// 3. THE ENABLE MATCHES THE WIDTH, exactly. Driving a lane the width does
// not include is contention on a bidirectional pin -- and because the
// enable is derived from the same mask as the data, this should be
// structurally impossible rather than merely checked.
a_enable_exact : assert property (
@(posedge clk) disable iff (!rst_n)
(tx_active) |-> ($countones(io_oe) == lanes))
else $error("the number of driven lanes does not match the width");
// 4. NOTHING DRIVEN WHEN IDLE. A controller holding the lanes between
// transfers prevents the device from answering, and the symptom is a
// device that looks dead rather than one that looks contended.
a_idle_released : assert property (
@(posedge clk) disable iff (!rst_n)
(!tx_active) |-> (io_oe == '0))
else $error("lanes were driven while idle");
// 5. GROUP COUNT. A byte takes exactly 8 / lanes groups -- so a width
// change mid-byte, which is a configuration error, cannot silently
// produce a byte assembled from mismatched group sizes.
a_group_count : assert property (
@(posedge clk) disable iff (!rst_n)
(rx_valid) |-> (groups_this_byte == (8 / lanes)))
else $error("a byte was assembled from the wrong number of groups");Properties 1 and 2 together are the lesson, and it generalises well beyond this block. A round trip is a powerful and incomplete specification: it proves the two directions are mutually consistent and says nothing about whether either is right. Any design with an inverse has this shape — encoders and decoders, packers and unpackers, serialisers and deserialisers — and every one of them needs at least one absolute check alongside the round trip.
Property 3 is worth contrasting with its own implementation. The gearbox derives the enable from the mask that narrows the data, so the property should be unfalsifiable by construction — and that is the right relationship between an assertion and a design. An assertion that cannot fail because the structure prevents it is not a wasted assertion; it is a statement that the structure is doing the work, and it will fail loudly if someone later separates the two.
Coverage must cross the width with the data pattern, because the symmetric values hide mapping errors:
covergroup spi_lane_cg @(posedge clk iff load);
cp_lanes : coverpoint lanes {
bins one = {1};
bins two = {2};
bins four = {4};
bins eight = {8};
illegal_bins impossible = {0, 3, [5:7], [9:15]};
}
// The SYMMETRY of the payload is what decides whether a mapping
// error is visible. 0xA5 and 0xFF are symmetric under some wrong
// mappings; 0x13 is not. A suite using only round numbers tests the
// gearbox against the cases most likely to hide its own bugs.
cp_symmetry : coverpoint byte_class {
bins all_zeros = {B_00}; // invisible to everything
bins all_ones = {B_FF}; // invisible to everything
bins nibble_sym = {B_NIBBLE_SYM};// 0xA5 -- hides some errors
bins asymmetric = {B_ASYM}; // 0x13 -- hides nothing
}
cp_dir : coverpoint transfer_dir { bins tx = {0}; bins rx = {1}; }
// Exactly the driven-lane counts that are legal, and nothing else.
cp_oe : coverpoint $countones(io_oe) {
bins idle = {0};
bins one = {1};
bins two = {2};
bins four = {4};
bins eight = {8};
illegal_bins wrong = {3, [5:7]};
}
x_lanes_symmetry : cross cp_lanes, cp_symmetry;
endgroupcp_symmetry is the coverpoint this chapter argues for and the one almost never written. A test suite naturally reaches for 0x00, 0xFF, 0xAA and 0x55 — and every one of those is invisible to at least one mapping error. The asymmetric bin is the only one that is not, and crossing it with the width is what proves the mapping rather than merely exercising it.
8. Why an FPGA or ASIC Engineer Cares
Derive the output enable from the lane mask. One mask feeding both the data and the enable makes it structurally impossible to drive a lane the width does not include. Two independent computations of the same thing will eventually disagree, and the disagreement is contention.
Release the lanes when idle. A held bus produces a device that appears dead, which is a much harder symptom to trace than a contended one.
Hold 8 / lanes as a table. Four legal values, and the table is also where an illegal width is caught — one structure doing both jobs, and no divider.
Test with an asymmetric byte. 0xAA, 0x55, 0x00 and 0xFF are all invisible to at least one mapping error. 0x13 is not, and one such value in a test list is worth more than a dozen round numbers.
Remember what enabling quad mode gives up. WP# and HOLD#/RESET# become data lanes, so write protection and — on many parts — the hardware reset are gone while the mode is active. If the board depends on either, that dependency must be re-examined before the mode is enabled, not after.
Treat the mode bit as persistent state. On many parts the quad-enable bit is non-volatile, so it survives power cycling and a controller that assumes single-lane at reset will be wrong. This is Chapter 11.5's warm-boot problem with a different bit.
9. Failure Signature — Nibble-Swapped Data on Exactly One Board Variant
Symptom. A quad-mode read returns data whose bytes are correct in content but have their nibbles exchanged — 0x12 0x34 reads back as 0x21 0x43. It happens on one board variant and not another, with identical firmware and identical flash part numbers.
What "correct content, exchanged nibbles" establishes. Every bit arrived. Nothing was lost, added or corrupted — only rearranged, and rearranged in a completely regular way. So this is a wiring or mapping fault rather than a timing or protocol one, and that single observation eliminates most of what the previous three modules were about.
Plausible mechanisms.
- The
IO2/IO3pair swapped on the board. Exchanging two lanes within a nibble does not swap the nibbles, so this does not fit — but exchangingIO0↔IO2andIO1↔IO3would produce a bit permutation within each nibble rather than a nibble swap either. - A group-order error in the controller, sending the low nibble first. This produces exactly a nibble swap — but it would affect both board variants, since it is in the firmware.
- The board variants differ in lane routing, and one has
IO0..IO3connected in a different order — which, combined with a controller that assumes the standard order, produces a fixed permutation. - One variant's flash is in a different mode — DDR, or a different address width — which would shift rather than permute.
- A level shifter or buffer on two of the four lanes only, which would produce timing-dependent corruption rather than a clean swap.
The discriminating observation. "Identical firmware, one variant affected" is decisive: a firmware-side group-order error cannot be variant-specific, so the fault must be on the board. That immediately rules out the controller and points at the routing.
Then determine which permutation, and there is a clean way to do it without a scope. Read a byte whose nibbles are distinguishable and whose bits within each nibble are too — 0x13 gives 0001 and 0011, and each of the plausible mis-routings maps it to a different value. Compare against a short table computed by hand and the permutation identifies itself.
The stronger test, if the flash can be written on the affected variant: write and read back on the same board. If a round trip on the bad variant succeeds — the value written comes back — the error is symmetric, which means it is a routing permutation applied identically in both directions. That is the point §7's property 2 makes: a round trip passing is consistent with both directions being wrong in the same way.
The fix. For a wiring error already in hardware, the lane order becomes a configuration of the controller: a per-board lane permutation applied in the gearbox. That is a small amount of logic and it turns a board respin into a table entry — which is the same argument Chapter 10.1 made about device profiles, arriving at the board level.
Why it reaches production. Because quad lane routing is often done for signal-integrity reasons — swapping lanes to shorten a trace or avoid a via — on the correct understanding that the lanes are interchangeable electrically. They are not interchangeable logically, and the layout engineer has no reason to know that unless the constraint is written down.
10. Common Misconceptions
11. Reason It Through
Work this before reading the answer.
A controller supports 1-1-1 and 1-1-4. A board has a quad-capable flash, and the firmware enables quad mode by setting the non-volatile quad-enable bit, then reads with
0x6B(quad output).It works. Some weeks later the same firmware is loaded onto a fresh board from the same batch, and the very first read — before quad mode is enabled — returns garbage. What happened, and what does the order of events tell you?
Start with the difference. On the working board, quad mode was enabled at some point in the past and the bit is non-volatile, so it is still set. On the fresh board it is not.
But the failing read happens BEFORE quad mode is enabled, and on the fresh board that read is a single-lane read on a part in its default state. That should work — so the firmware cannot simply be relying on the quad bit.
So what is different about the first read on the fresh board? Nothing about the flash. The difference must be in the controller.
And there it is. On the working board, the controller was left configured for quad mode by the previous run — its lane count, and more importantly its output enable mask, are still set to four lanes. A single-lane read issued with a four-lane enable mask drives IO1, IO2 and IO3 as well as IO0. On a part in single-lane mode, IO1 is the device's output — so the controller is fighting the device on every returned bit.
Why did it ever work, then? Because on the working board the quad bit was set, so the device also treated the lanes as bidirectional and released them during the command phase. The contention window was small enough to survive. On the fresh board the device is driving IO1 continuously as MISO, and the controller is driving it too.
What does the ORDER of events tell you? That the fault is in the controller's reset state, not in its configured state. The working board hid a controller bug because the device was in a matching mode; the fresh board exposed it because the device was in its default one. This is the general shape worth remembering: a controller and a device that are both mis-configured in the same direction can appear to work, and the failure surfaces when only one of them is reset.
The fix, in two parts.
The controller must reset to single-lane with a one-bit enable mask — the same argument Chapter 11.2 made for resetting to the plain read at the slowest divisor, now applied to width. Reset to the configuration that works against a device in its default state.
And the firmware must not treat the quad-enable bit as something it set once. It should read it back and configure the controller to match what the device actually reports, rather than what the firmware believes it wrote — which is Chapter 11.4's "confirm the latch" discipline applied to a mode bit.
One more observation worth making. The non-volatile quad bit means a board can arrive from manufacturing in either state depending on what test software ran on it. A design that works only in one of those states will fail on some fraction of units, apparently at random — and the fraction will depend on the factory's process rather than on anything in the design.
12. Understanding Check
13. Summary
Multi-lane SPI repurposes the write-protect and hold pins as IO2 and IO3, which is why it needs no new pins and why enabling it gives up write protection and often the hardware reset. The mode bit is frequently non-volatile, so it survives power cycling.
Single-lane SPI had one convention to get right; multi-lane has two. The most significant group goes first, and within a group the most significant bit goes on the highest-numbered lane. Both halves produce 0x5A from 0xA5 when reversed, by different mechanisms, and are distinguishable only on an asymmetric byte.
Every data pin becomes bidirectional, so the output enable stops being a convenience and becomes a safety signal: exactly as many lanes as the width calls for, and none when idle. Deriving it from the same mask that narrows the data makes over-driving structurally impossible rather than merely checked.
In hardware, one shift register serves every width in both directions; 8 / lanes is a table, not a division; and the alignment shift is the only thing that changes with the width.
For verification, the round trip is powerful and incomplete — it proves the two directions agree and not that either is right, so an absolute check belongs beside it. And the coverage axis that matters is the payload's symmetry, because 0x00, 0xFF, 0xAA and 0x55 are each invisible to at least one mapping error while 0x13 is invisible to none.
14. What Comes Next
The lanes are now understood as wires and as a convention. What has not been examined is the consequence of their being bidirectional.
Chapter 12.3 — Lane Widths per Phase and Direction Changes reads the 1-1-4 / 1-4-4 / 4-4-4 notation properly, and then asks the question the notation does not answer: which phases the master drives and which the device does. A single-lane read never reverses a wire; a multi-lane read reverses four at once — and the chapter's finding is that this is the second reason a quad read specifies dummy cycles, alongside Chapter 11.2's array access time. A quad read specified with no dummy phase has no turnaround window at all, and the walker in that chapter reports it.
Continue learning
Related tutorials
- Related topic
Why Wider SPI Exists
The bandwidth pressure that pushed flash past one data line, why four data lanes deliver far less than four times the speed, why the dummy phase becomes more prominent as the bus widens, and the model that computes the gain for any width combination.
- Related topic
Lane Widths per Phase and Direction Changes
Reading the 1-1-4 / 1-4-4 / 4-4-4 notation properly, the direction it does not state, why a multi-lane read reverses four wires at once, and why that is the second reason a quad read needs dummy cycles.
- Related topic
Launch and Sample Edges
One edge of each bit time places a bit on the wire, the other captures it, and they must never be the same edge. Why the separation is forced, why it buys half a period, and how RTL maps physical edges onto those roles.
- Related topic
Deriving Mode Behaviour from CPOL and CPHA
The four SPI modes are a two-bit truth table you can rebuild in seconds. The standard numbering, the derivation, the complete mode decoder in three HDLs, and the assertions that keep a configurable design honest.
