Skip to content
VLSI Mentor

SPI · Module 12

Dual, Quad, and Octal SPI

How dedicated pins become bidirectional lanes, the bit-to-lane convention every multi-lane device shares and what getting it backwards produces, why the output enable becomes a safety signal, and the gearbox that serialises a byte across any width.

Chapter 12.1 established what width buys. This chapter is about what it costs, and the first cost appears before any data moves.

Single-lane SPI only ever had to decide bit order. With four lanes there is a second question — which bit goes on which lane — and getting it backwards turns 0xA5 into 0x5A. Why is that worse than a wrong answer that looks wrong?

Because 0x5A is a perfectly ordinary byte. A driver checking for a plausible value is satisfied, a checksum fails later, and the investigation begins somewhere else entirely.

1. Where the Lanes Come From

A standard serial flash package has eight pins, and in single-lane mode two of them do almost nothing:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   single-lane use        multi-lane use
   ─────────────────      ────────────────
   CS#                    CS#          unchanged
   SCLK                   SCLK         unchanged
   MOSI  (DI)             IO0          bidirectional
   MISO  (DO)             IO1          bidirectional
   WP#   (write protect)  IO2          bidirectional
   HOLD# (or RESET#)      IO3          bidirectional
   VCC, GND               VCC, GND

The write-protect and hold pins become data lanes. That is the whole trick, and it explains two things that otherwise look arbitrary.

Why the modes must be enabled explicitly. Repurposing WP# means the write-protect function is gone while quad mode is active, so a device cannot simply default to it — a part that powered up in quad mode would have no write protection and would interpret a legacy controller's idle WP# level as data. Enabling quad mode is a deliberate act, usually a status-register bit, and on many parts a non-volatile one that survives power cycling.

Why IO0 and IO1 are MOSI and MISO. The lane numbering is not arbitrary: IO0 is the pin that was MOSI, so a single-lane transaction on a quad-capable part uses IO0 outbound and IO1 inbound exactly as before. Backwards compatibility is a property of the numbering.

2. The Bit-to-Lane Convention

With one lane, a byte is eight clocks and the only question is which bit goes first — most significant, by universal convention.

With four lanes, a byte is two clocks and there are two questions. The convention is uniform across every multi-lane device and has two parts:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   1. the most significant GROUP goes first
   2. within a group, the most significant bit goes on the
      HIGHEST-numbered lane

So 0xA5 = 1010_0101 on four lanes:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   clock 1:   IO3=1  IO2=0  IO1=1  IO0=0      the high nibble, 0xA
   clock 2:   IO3=0  IO2=1  IO1=0  IO0=1      the low nibble,  0x5

Both halves matter and they fail differently.

Getting the group order backwards sends the low nibble first, so 0xA5 becomes 0x5A. Every byte's nibbles are swapped.

Getting the within-group order backwards puts bit 7 on IO0, so the high nibble 1010 is transmitted as 0101 — and 0xA5 again reads as 0x5A, by a completely different mechanism. The two errors are distinguishable only on a byte whose nibbles are not each other's bit-reverse.

That is why §6's testbench checks 0x13 as well as 0xA5: 0x13 gives 1 then 3, and both error modes produce something other than that.

3. Time, Compressed

One byte, one lane and four

8 cycles
Two lanes over eight byte times. The single-lane row sends bits seven down to zero, one per clock. The four-lane row sends the high nibble then the low nibble in the first two clocks and is idle thereafter.quad byte completequad byte complete1 laneb7b6b5b4b3b2b1b04 lanesA5idleidleidleidleidleidlet0t1t2t3t4t5t6t7
Figure 1 — the same byte at one lane and at four. Eight clocks become two, and the high nibble goes first with bit 7 on the highest-numbered lane. The four-lane transfer is finished while the single-lane one is a quarter through.

The idle cells are not decoration. They are where Chapter 12.1's gain comes from and, equally, where its limit lives: those six clocks are recovered only if there is more data to send. On a transfer whose data phase was already short, most of what the figure shows as saved was never there to save.

4. What Each Width Buys — and Costs

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   width    clocks/byte    pins used as data    what is given up
   1         8             2 (MOSI, MISO)       nothing
   2         4             2 (IO0, IO1)         full duplex
   4         2             4 (IO0..IO3)         WP#, HOLD#/RESET#
   8         1             8                    a wider package

The second column is the gain and the fourth is the price. Two entries deserve comment.

Dual mode gives up full duplex and gains nothing in pins. It uses the same two pins as single-lane SPI, but both become unidirectional-at-a-time rather than one each way. So a dual transfer is half the clocks of a single-lane one and cannot send and receive simultaneously — which for a flash is irrelevant, because flash transactions are never symmetric anyway. Dual exists mainly as a fallback for parts or boards where IO2 and IO3 are unavailable.

Octal needs more pins than a standard package has. Eight data lanes plus clock, select and power do not fit an 8-pin part, so octal devices come in larger packages — and at that point the pin argument that motivated SPI in the first place (Chapter 1.1) is considerably weaker. Octal is used where the alternative is a parallel bus, not where the alternative is single-lane SPI.

5. The Output Enable Becomes a Safety Signal

Here is the consequence that separates multi-lane SPI from everything earlier in this track.

In single-lane SPI, MOSI is always driven by the master and MISO always by the slave. Neither line ever changes direction, so there is no question of who drives what.

In quad mode all four lanes are bidirectional. The master drives them for the command and address; the device drives them for the data. So a lane driven by the master while the device is also driving it is contention, with the currents Chapter 8.5 described — now on four lanes at once.

Two rules follow, and §6's gearbox enforces both:

Drive exactly as many lanes as the width calls for. A single-lane transfer on a quad-capable controller must drive IO0 and nothing else. Driving all four because the port is four wide means IO1 — the device's output — is being fought over on every transfer.

Drive nothing when idle. A controller that holds the lanes between transfers prevents the device from answering at all, and the symptom is a device that appears dead rather than one that appears contended.

A lane gearbox. A byte register feeds a shift-by-group stage whose amount is the lane count. The top group is extracted and masked by a lane mask derived from the width, producing the lane outputs. The same mask produces the output enable. A group counter derived from the width counts the groups per byte. On the receive side, incoming groups are masked by the same lane mask and shifted into an assembly register that reproduces the byte.Byte registerloaded once per byteShift by groupamount = the lane countGroup counter8 / lanes — a table, not adividerAssembly registerreceive: the same shift,reversedLane maskone bit per active laneTop groupbit 7 on the highest laneInput maskthe SAME mask, appliedinboundIO0..IO7bidirectional lanesOutput enablefrom the mask — cannotexceed the widthReceived byteidentical at every widthextractnarrowsdrivesderiveshow many groupscapturesshifts in12
Figure 2 — one gearbox, four widths. The byte register and the group counter are shared; only the shift amount, the lane mask and the group count change with the width. The output enable is derived from the same mask, which is what makes it impossible to drive a lane the width does not include.

The single Lane mask node feeding both the data path and the output enable is the design decision. Deriving the enable from the same mask that narrows the data makes it structurally impossible to drive a lane the width does not include — which is stronger than checking for it.

6. Building the Lane Gearbox — Three HDLs

The circuit

Circuit. One shift register serving every width, in both directions.

State. A transmit shift register, a receive assembly register, and a group counter.

Datapath. Transmit shifts left by the lane count each clock and presents the top lanes bits, aligned so bit 7 lands on the highest active lane. Receive shifts left by the same amount and admits the masked input at the bottom. The two are exact inverses, which is what the round trip in the testbench exploits.

Control. A group counter of 8 / lanes, held as a table rather than a division — four legal values, and the same table is where an illegal width is caught.

Clock and reset. System clock; asynchronous active-low reset.

Enables. io_oe carries exactly lanes bits while transmitting and zero when idle. It is derived from the same mask that narrows the data, so the two cannot disagree.

Timing. One group per clock; tx_done pulses on the last group and rx_valid when a byte completes.

Synthesis. Two 8-bit shift registers, a small barrel shifter for the alignment, and two counters. Tens of flip-flops.

Limitations. Byte-granular. A device transferring an odd number of nibbles — which Chapter 12.4 shows DDR can produce — needs the group count decoupled from the byte boundary.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_lane_serdes.sv — one gearbox, four widths, both directions
// spi_lane_serdes.sv
//
// Chapter 12.2 -- the gearbox between a byte and a group of lanes.
//
// Single-lane SPI moves one bit per clock, so a byte is eight clocks and
// the bit order is the only question. With two, four or eight lanes a byte
// becomes four, two or one clock -- and a second question appears that has
// no single-lane equivalent: WHICH BIT GOES ON WHICH LANE.
//
// The convention is uniform across every multi-lane SPI device:
//
//   * the most significant GROUP goes first, and
//   * within a group, the most significant bit goes on the
//     highest-numbered lane.
//
// So for 0xA5 = 1010_0101 on four lanes, the first clock carries the high
// nibble 1010 with IO3=1, IO2=0, IO1=1, IO0=0, and the second carries
// 0101. Getting the within-group order backwards gives a byte whose
// nibbles are bit-reversed -- 0xA5 becomes 0x5A -- which is a plausible
// value and not an obvious error.
//
// OUTPUT ENABLE IS THE SAFETY OUTPUT. In quad and octal mode every data
// pin is bidirectional, so a lane the master drives while the device is
// also driving it is contention with real current behind it (Chapter 8.5).
// A four-lane transfer must drive exactly four lanes -- never the eight the
// port is wide.

module spi_lane_serdes #(
    parameter int MAX_LANES = 8
) (
    input  logic                 clk,
    input  logic                 rst_n,

    input  logic [3:0]           lanes,      // 1, 2, 4 or 8

    // Transmit: load a byte, shift it out most significant group first.
    input  logic                 load,
    input  logic [7:0]           tx_byte,
    output logic [MAX_LANES-1:0] io_out,
    output logic [MAX_LANES-1:0] io_oe,      // exactly `lanes` bits set
    output logic                 tx_active,
    output logic                 tx_done,

    // Receive: capture groups and assemble a byte.
    input  logic [MAX_LANES-1:0] io_in,
    input  logic                 rx_en,
    output logic [7:0]           rx_byte,
    output logic                 rx_valid,

    output logic [3:0]           groups,     // 8 / lanes, for this width
    output logic                 lane_err
);

    logic [7:0] tx_sr;
    logic [3:0] tx_left;
    logic [7:0] rx_sr;
    logic [3:0] rx_left;

    // 8 / lanes, as a table. The legal set is four values, so a table is
    // both cheaper than a divider and the place an illegal width is caught.
    function automatic logic [3:0] group_count(input logic [3:0] n);
        case (n)
            4'd1:    group_count = 4'd8;
            4'd2:    group_count = 4'd4;
            4'd4:    group_count = 4'd2;
            4'd8:    group_count = 4'd1;
            default: group_count = 4'd8;   // flagged by lane_err
        endcase
    endfunction

    function automatic logic bad_lanes(input logic [3:0] n);
        bad_lanes = !((n == 4'd1) || (n == 4'd2) ||
                      (n == 4'd4) || (n == 4'd8));
    endfunction

    // The active-lane mask, and the alignment shift that brings the most
    // significant group down to the low bits.
    wire [MAX_LANES-1:0] lane_mask = MAX_LANES'((1 << lanes) - 1);
    wire [3:0]           top_shift = 4'd8 - lanes;

    // The group currently presented: the top `lanes` bits of the shift
    // register, with bit 7 landing on the highest active lane.
    assign io_out   = MAX_LANES'(tx_sr >> top_shift) & lane_mask;
    assign io_oe    = tx_active ? lane_mask : {MAX_LANES{1'b0}};
    assign lane_err = bad_lanes(lanes);

    always_ff @(posedge clk or negedge rst_n) begin
        if (!rst_n) begin
            tx_sr     <= 8'h00;
            tx_left   <= 4'd0;
            tx_active <= 1'b0;
            tx_done   <= 1'b0;
            rx_sr     <= 8'h00;
            rx_left   <= 4'd0;
            rx_byte   <= 8'h00;
            rx_valid  <= 1'b0;
            groups    <= 4'd8;
        end else begin
            tx_done  <= 1'b0;
            rx_valid <= 1'b0;
            groups   <= group_count(lanes);

            if (load) begin
                tx_sr     <= tx_byte;
                tx_left   <= group_count(lanes);
                tx_active <= 1'b1;
                // A load also restarts the receive assembly, because a
                // multi-lane transfer is full duplex only in the sense that
                // the same clocks carry both directions -- the group count
                // is shared.
                rx_sr     <= 8'h00;
                rx_left   <= group_count(lanes);
            end else if (tx_active) begin
                // Shift by a whole group. The bits that leave the top are
                // the ones just presented on the lanes.
                tx_sr <= tx_sr << lanes;
                if (tx_left == 4'd1) begin
                    tx_active <= 1'b0;
                    tx_done   <= 1'b1;
                end else begin
                    tx_left <= tx_left - 4'd1;
                end
            end

            // Receive assembles with the same convention in reverse: each
            // group arrives on the low `lanes` bits of io_in and is shifted
            // in below the ones already captured.
            if (rx_en && (rx_left != 4'd0)) begin
                rx_sr <= (rx_sr << lanes) | 8'(io_in & lane_mask);
                if (rx_left == 4'd1) begin
                    rx_byte  <= 8'((rx_sr << lanes) | 8'(io_in & lane_mask));
                    rx_valid <= 1'b1;
                    rx_left  <= 4'd0;
                end else begin
                    rx_left <= rx_left - 4'd1;
                end
            end
        end
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_lane_serdes_tb.sv — the mapping by hand, then exhaustively
// spi_lane_serdes_tb.sv
//
// Two kinds of check. First the lane MAPPING is verified against values
// computed by hand from the convention, at every width -- because a
// mapping error produces a plausible byte rather than an obvious fault.
// Then the gearbox is looped back on itself and every byte value is
// round-tripped at every width, which is the property that catches a
// mapping error the examples happen to miss.

`timescale 1ns/1ps

module spi_lane_serdes_tb;

    localparam int MAX_LANES = 8;

    logic clk = 1'b0;
    logic rst_n = 1'b0;
    always #5 clk = ~clk;

    logic [3:0]           lanes = 4'd1;
    logic                 load = 1'b0;
    logic [7:0]           tx_byte = 8'h00;
    logic [MAX_LANES-1:0] io_out;
    logic [MAX_LANES-1:0] io_oe;
    logic                 tx_active, tx_done;
    logic [MAX_LANES-1:0] io_in;
    logic                 rx_en;
    logic [7:0]           rx_byte;
    logic                 rx_valid;
    logic [3:0]           groups;
    logic                 lane_err;

    int errors = 0;

    // A named value rather than a literal: a bit-select of a sized literal
    // is not portable, and the pattern is referred to often enough to be
    // worth naming.
    localparam logic [7:0] PAT = 8'hA5;

    // Loopback: what the gearbox drives is what it receives. rx_en follows
    // tx_active so the two run on the same clocks.
    assign io_in = io_out;
    assign rx_en = tx_active;

    spi_lane_serdes #(.MAX_LANES(MAX_LANES)) dut (
        .clk(clk), .rst_n(rst_n), .lanes(lanes),
        .load(load), .tx_byte(tx_byte),
        .io_out(io_out), .io_oe(io_oe),
        .tx_active(tx_active), .tx_done(tx_done),
        .io_in(io_in), .rx_en(rx_en),
        .rx_byte(rx_byte), .rx_valid(rx_valid),
        .groups(groups), .lane_err(lane_err)
    );

    // Captured groups, for the mapping check.
    logic [MAX_LANES-1:0] seen [0:7];
    int                   n_seen;
    bit                   saw_rx;
    logic [7:0]           got_byte;
    logic [MAX_LANES-1:0] oe_union;

    always_ff @(posedge clk) begin
        if (rst_n && tx_active && n_seen < 8) begin
            seen[n_seen] <= io_out;
            n_seen       <= n_seen + 1;
            oe_union     <= oe_union | io_oe;
        end
        if (rst_n && rx_valid) begin
            saw_rx   <= 1'b1;
            got_byte <= rx_byte;
        end
    end

    task automatic send(input int n, input logic [7:0] b);
        int guard;
        begin
            @(negedge clk);
            lanes = 4'(n); tx_byte = b;
            n_seen = 0; saw_rx = 1'b0; oe_union = {MAX_LANES{1'b0}};
            load = 1'b1;
            @(negedge clk);
            load = 1'b0;
            guard = 0;
            while (tx_active && guard < 32) begin
                @(negedge clk);
                guard++;
            end
            @(negedge clk);
        end
    endtask

    initial begin
        n_seen = 0; saw_rx = 1'b0; got_byte = 8'h00;
        oe_union = {MAX_LANES{1'b0}};
        repeat (3) @(negedge clk);
        rst_n = 1'b1;
        @(negedge clk);

        // 1. SINGLE LANE. 0xA5 = 1010_0101, most significant bit first --
        //    the same order every earlier module used.
        send(1, PAT);
        if (n_seen != 8) begin
            $display("  FAIL: one lane produced %0d groups, expected 8", n_seen);
            errors++;
        end
        for (int i = 0; i < 8; i++) begin
            if (seen[i][0] !== PAT[7-i]) begin
                $display("  FAIL: 1 lane group %0d carried %0b, expected %0b",
                         i, seen[i][0], PAT[7-i]);
                errors++;
            end
        end
        $display("  1 lane:  8 groups, MSB first: %0b%0b%0b%0b%0b%0b%0b%0b",
                 seen[0][0], seen[1][0], seen[2][0], seen[3][0],
                 seen[4][0], seen[5][0], seen[6][0], seen[7][0]);

        // 2. FOUR LANES. The high nibble first, and within it bit 7 on the
        //    highest lane: 0xA5 gives 4'hA then 4'h5. A within-group
        //    reversal would give 4'h5 then 4'hA -- which reads as 0x5A and
        //    looks like a perfectly ordinary byte.
        send(4, PAT);
        if (n_seen != 2) begin
            $display("  FAIL: four lanes produced %0d groups, expected 2", n_seen);
            errors++;
        end
        if (seen[0][3:0] !== 4'hA || seen[1][3:0] !== 4'h5) begin
            $display("  FAIL: 4 lanes gave 0x%01h then 0x%01h, expected A then 5",
                     seen[0][3:0], seen[1][3:0]);
            errors++;
        end
        $display("  4 lanes: 2 groups, 0x%01h then 0x%01h  (IO3..IO0 per group)",
                 seen[0][3:0], seen[1][3:0]);

        // 3. TWO LANES. Four groups of two bits, most significant pair
        //    first: 10, 10, 01, 01.
        send(2, PAT);
        if (n_seen != 4) begin
            $display("  FAIL: two lanes produced %0d groups, expected 4", n_seen);
            errors++;
        end
        if (seen[0][1:0] !== 2'b10 || seen[1][1:0] !== 2'b10 ||
            seen[2][1:0] !== 2'b01 || seen[3][1:0] !== 2'b01) begin
            $display("  FAIL: 2 lanes gave %0b %0b %0b %0b, expected 10 10 01 01",
                     seen[0][1:0], seen[1][1:0], seen[2][1:0], seen[3][1:0]);
            errors++;
        end
        $display("  2 lanes: 4 groups, %02b %02b %02b %02b",
                 seen[0][1:0], seen[1][1:0], seen[2][1:0], seen[3][1:0]);

        // 4. EIGHT LANES. One group carrying the whole byte.
        send(8, PAT);
        if (n_seen != 1) begin
            $display("  FAIL: eight lanes produced %0d groups, expected 1", n_seen);
            errors++;
        end
        if (seen[0] !== 8'hA5) begin
            $display("  FAIL: 8 lanes gave 0x%02h, expected 0xA5", seen[0]);
            errors++;
        end
        $display("  8 lanes: 1 group,  0x%02h", seen[0]);

        // 5. OUTPUT ENABLE. Exactly `lanes` bits driven, never more. In
        //    quad and octal mode every data pin is bidirectional, so a lane
        //    driven when it should not be is contention with real current.
        send(1, 8'hFF);
        if (oe_union !== 8'b0000_0001) begin
            $display("  FAIL: one lane drove oe=%08b, expected 00000001", oe_union);
            errors++;
        end
        send(2, 8'hFF);
        if (oe_union !== 8'b0000_0011) begin
            $display("  FAIL: two lanes drove oe=%08b, expected 00000011", oe_union);
            errors++;
        end
        send(4, 8'hFF);
        if (oe_union !== 8'b0000_1111) begin
            $display("  FAIL: four lanes drove oe=%08b, expected 00001111", oe_union);
            errors++;
        end
        send(8, 8'hFF);
        if (oe_union !== 8'b1111_1111) begin
            $display("  FAIL: eight lanes drove oe=%08b, expected 11111111", oe_union);
            errors++;
        end
        $display("  output enable: exactly 1, 2, 4 and 8 lanes driven -- never more");

        // 6. The enable must be clear when idle, or the master holds the
        //    bus between transfers and the device cannot answer at all.
        @(negedge clk);
        if (io_oe !== {MAX_LANES{1'b0}}) begin
            $display("  FAIL: lanes still driven while idle: %08b", io_oe);
            errors++;
        end

        // 7. EXHAUSTIVE ROUND TRIP. For every width and every byte value,
        //    serialising and deserialising must return the byte. This is
        //    what catches a mapping error the hand-checked examples above
        //    happen to agree with -- 0xA5 and 0xFF are both symmetric under
        //    some wrong mappings.
        for (int w = 0; w < 4; w++) begin
            int nl;
            nl = (w == 0) ? 1 : (w == 1) ? 2 : (w == 2) ? 4 : 8;
            for (int b = 0; b < 256; b++) begin
                send(nl, 8'(b));
                if (!saw_rx) begin
                    $display("  FAIL: %0d lanes, 0x%02h produced no received byte",
                             nl, b);
                    errors++;
                end else if (got_byte !== 8'(b)) begin
                    $display("  FAIL: %0d lanes round trip 0x%02h returned 0x%02h",
                             nl, b, got_byte);
                    errors++;
                end
            end
        end
        $display("  round trip: 1024 (width, byte) pairs recovered exactly");

        // 8. A byte with NO symmetry, checked by hand at four lanes, to
        //    pin the mapping independently of the round trip -- which would
        //    pass even if BOTH directions were reversed consistently.
        send(4, 8'h13);
        if (seen[0][3:0] !== 4'h1 || seen[1][3:0] !== 4'h3) begin
            $display("  FAIL: 0x13 gave 0x%01h then 0x%01h, expected 1 then 3",
                     seen[0][3:0], seen[1][3:0]);
            errors++;
        end
        $display("  0x13 at 4 lanes: 0x%01h then 0x%01h -- high nibble first",
                 seen[0][3:0], seen[1][3:0]);

        // 9. An illegal width is reported rather than silently treated as
        //    one lane.
        @(negedge clk);
        lanes = 4'd3;
        @(negedge clk);
        if (!lane_err) begin
            $display("  FAIL: three lanes was not reported as illegal"); errors++;
        end
        lanes = 4'd4;
        @(negedge clk);
        if (lane_err) begin
            $display("  FAIL: four lanes was reported as illegal"); errors++;
        end
        $display("  lane validation: 3 rejected, 4 accepted");

        if (errors == 0)
            $display("PASS: the most significant group goes first and within a group the most significant bit goes on the highest-numbered lane, at one, two, four and eight lanes; exactly as many lanes are driven as the width calls for and none while idle; and serialising then deserialising returns the byte for all 256 values at all four widths");
        else
            $display("FAIL: %0d error(s)", errors);
        $finish;
    end

endmodule

The testbench has two layers and both are necessary, which is the point worth taking from it.

The hand-checked mapping pins the convention against values computed from §2 rather than from the design: 0xA5 gives 0xA then 0x5 at four lanes, 10 10 01 01 at two, and the whole byte at eight. Those checks would catch a group-order error.

The exhaustive round trip then loops the gearbox's output back into its input and requires every one of the 256 byte values to survive at every one of the four widths — 1024 pairs. That catches a mapping error the examples happen to agree with, which is a real risk here because 0xA5 and 0xFF are both symmetric under some wrong mappings.

But a round trip alone would be insufficient, and the testbench says so: it would pass even if both directions were reversed consistently, because the two errors cancel. So the asymmetric byte 0x13 is checked by hand at four lanes — 1 then 3 — to pin the absolute direction that the round trip cannot.

Two properties then cover the safety side. The output enable must carry exactly one, two, four or eight bits according to the width — never the eight the port is wide — and it must be clear when idle, because a controller holding the lanes prevents the device from answering at all.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_lane_serdes.v — the same gearbox in Verilog-2001
// spi_lane_serdes.v
//
// Chapter 12.2 -- the gearbox between a byte and a group of lanes, in
// Verilog-2001.
//
// With two, four or eight lanes a byte becomes four, two or one clock, and
// a question appears that has no single-lane equivalent: WHICH BIT GOES ON
// WHICH LANE. The convention is uniform across every multi-lane device:
//
//   * the most significant GROUP goes first, and
//   * within a group, the most significant bit goes on the
//     highest-numbered lane.
//
// So 0xA5 on four lanes is 0xA then 0x5. Getting the within-group order
// backwards turns 0xA5 into 0x5A -- a plausible value, not an obvious
// error.
//
// OUTPUT ENABLE IS THE SAFETY OUTPUT. In quad and octal mode every data pin
// is bidirectional, so driving a lane the device is also driving is
// contention with real current behind it. A four-lane transfer must drive
// exactly four lanes, never the eight the port is wide.

module spi_lane_serdes #(
    parameter MAX_LANES = 8
) (
    input  wire                 clk,
    input  wire                 rst_n,

    input  wire [3:0]           lanes,      // 1, 2, 4 or 8

    // Transmit: load a byte, shift it out most significant group first.
    input  wire                 load,
    input  wire [7:0]           tx_byte,
    output wire [MAX_LANES-1:0] io_out,
    output wire [MAX_LANES-1:0] io_oe,      // exactly `lanes` bits set
    output reg                  tx_active,
    output reg                  tx_done,

    // Receive: capture groups and assemble a byte.
    input  wire [MAX_LANES-1:0] io_in,
    input  wire                 rx_en,
    output reg  [7:0]           rx_byte,
    output reg                  rx_valid,

    output reg  [3:0]           groups,     // 8 / lanes, for this width
    output wire                 lane_err
);

    reg [7:0] tx_sr;
    reg [3:0] tx_left;
    reg [7:0] rx_sr;
    reg [3:0] rx_left;

    // 8 / lanes, as a table. The legal set is four values, so a table is
    // both cheaper than a divider and the place an illegal width is caught.
    function [3:0] group_count;
        input [3:0] n;
        begin
            case (n)
                4'd1:    group_count = 4'd8;
                4'd2:    group_count = 4'd4;
                4'd4:    group_count = 4'd2;
                4'd8:    group_count = 4'd1;
                default: group_count = 4'd8;   // flagged by lane_err
            endcase
        end
    endfunction

    function bad_lanes;
        input [3:0] n;
        begin
            bad_lanes = !((n == 4'd1) || (n == 4'd2) ||
                          (n == 4'd4) || (n == 4'd8));
        end
    endfunction

    // The active-lane mask, and the alignment shift that brings the most
    // significant group down to the low bits.
    wire [MAX_LANES-1:0] lane_mask = (1 << lanes) - 1;
    wire [3:0]           top_shift = 4'd8 - lanes;

    // The group currently presented: the top `lanes` bits of the shift
    // register, with bit 7 landing on the highest active lane.
    assign io_out   = (tx_sr >> top_shift) & lane_mask;
    assign io_oe    = tx_active ? lane_mask : {MAX_LANES{1'b0}};
    assign lane_err = bad_lanes(lanes);

    always @(posedge clk or negedge rst_n) begin
        if (!rst_n) begin
            tx_sr     <= 8'h00;
            tx_left   <= 4'd0;
            tx_active <= 1'b0;
            tx_done   <= 1'b0;
            rx_sr     <= 8'h00;
            rx_left   <= 4'd0;
            rx_byte   <= 8'h00;
            rx_valid  <= 1'b0;
            groups    <= 4'd8;
        end else begin
            tx_done  <= 1'b0;
            rx_valid <= 1'b0;
            groups   <= group_count(lanes);

            if (load) begin
                tx_sr     <= tx_byte;
                tx_left   <= group_count(lanes);
                tx_active <= 1'b1;
                // A load also restarts the receive assembly: the group count
                // is shared between the two directions.
                rx_sr     <= 8'h00;
                rx_left   <= group_count(lanes);
            end else if (tx_active) begin
                // Shift by a whole group. The bits that leave the top are
                // the ones just presented on the lanes.
                tx_sr <= tx_sr << lanes;
                if (tx_left == 4'd1) begin
                    tx_active <= 1'b0;
                    tx_done   <= 1'b1;
                end else begin
                    tx_left <= tx_left - 4'd1;
                end
            end

            // Receive assembles with the same convention in reverse: each
            // group arrives on the low `lanes` bits of io_in and is shifted
            // in below the ones already captured.
            if (rx_en && (rx_left != 4'd0)) begin
                rx_sr <= (rx_sr << lanes) | (io_in & lane_mask);
                if (rx_left == 4'd1) begin
                    rx_byte  <= (rx_sr << lanes) | (io_in & lane_mask);
                    rx_valid <= 1'b1;
                    rx_left  <= 4'd0;
                end else begin
                    rx_left <= rx_left - 4'd1;
                end
            end
        end
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_lane_serdes_tb.v — the same 1024 round trips in Verilog-2001
// spi_lane_serdes_tb.v
//
// The same checks as the SystemVerilog testbench: the lane mapping
// verified against values computed by hand from the convention at every
// width, then an exhaustive loopback round trip of all 256 byte values at
// all four widths.

`timescale 1ns/1ps

module spi_lane_serdes_tb;

    parameter MAX_LANES = 8;

    // A named value rather than a literal: a bit-select of a sized literal
    // is not portable, and the pattern is referred to often enough to name.
    localparam [7:0] PAT = 8'hA5;

    reg clk;
    reg rst_n;

    reg  [3:0]           lanes;
    reg                  load;
    reg  [7:0]           tx_byte;
    wire [MAX_LANES-1:0] io_out;
    wire [MAX_LANES-1:0] io_oe;
    wire                 tx_active, tx_done;
    wire [MAX_LANES-1:0] io_in;
    wire                 rx_en;
    wire [7:0]           rx_byte;
    wire                 rx_valid;
    wire [3:0]           groups;
    wire                 lane_err;

    integer errors;
    integer i, w, b, nl, guard;

    // Loopback: what the gearbox drives is what it receives.
    assign io_in = io_out;
    assign rx_en = tx_active;

    initial begin
        clk = 1'b0; rst_n = 1'b0;
        lanes = 4'd1; load = 1'b0; tx_byte = 8'h00;
        errors = 0;
    end
    always #5 clk = ~clk;

    spi_lane_serdes #(.MAX_LANES(MAX_LANES)) dut (
        .clk(clk), .rst_n(rst_n), .lanes(lanes),
        .load(load), .tx_byte(tx_byte),
        .io_out(io_out), .io_oe(io_oe),
        .tx_active(tx_active), .tx_done(tx_done),
        .io_in(io_in), .rx_en(rx_en),
        .rx_byte(rx_byte), .rx_valid(rx_valid),
        .groups(groups), .lane_err(lane_err)
    );

    // Captured groups, for the mapping check.
    reg [MAX_LANES-1:0] seen [0:7];
    integer             n_seen;
    reg                 saw_rx;
    reg [7:0]           got_byte;
    reg [MAX_LANES-1:0] oe_union;

    initial begin
        n_seen = 0; saw_rx = 1'b0; got_byte = 8'h00;
        oe_union = {MAX_LANES{1'b0}};
    end

    always @(posedge clk) begin
        if (rst_n && tx_active && n_seen < 8) begin
            seen[n_seen] <= io_out;
            n_seen       <= n_seen + 1;
            oe_union     <= oe_union | io_oe;
        end
        if (rst_n && rx_valid) begin
            saw_rx   <= 1'b1;
            got_byte <= rx_byte;
        end
    end

    task send;
        input integer n;
        input [7:0]   bb;
        begin
            @(negedge clk);
            lanes = n[3:0]; tx_byte = bb;
            n_seen = 0; saw_rx = 1'b0; oe_union = {MAX_LANES{1'b0}};
            load = 1'b1;
            @(negedge clk);
            load = 1'b0;
            guard = 0;
            while (tx_active && guard < 32) begin
                @(negedge clk);
                guard = guard + 1;
            end
            @(negedge clk);
        end
    endtask

    initial begin
        repeat (3) @(negedge clk);
        rst_n = 1'b1;
        @(negedge clk);

        // 1. SINGLE LANE, most significant bit first.
        send(1, PAT);
        if (n_seen != 8) begin
            $display("  FAIL: one lane produced %0d groups, expected 8", n_seen);
            errors = errors + 1;
        end
        for (i = 0; i < 8; i = i + 1) begin
            if (seen[i][0] !== PAT[7-i]) begin
                $display("  FAIL: 1 lane group %0d carried %0b, expected %0b",
                         i, seen[i][0], PAT[7-i]);
                errors = errors + 1;
            end
        end
        $display("  1 lane:  8 groups, MSB first: %0b%0b%0b%0b%0b%0b%0b%0b",
                 seen[0][0], seen[1][0], seen[2][0], seen[3][0],
                 seen[4][0], seen[5][0], seen[6][0], seen[7][0]);

        // 2. FOUR LANES. High nibble first, bit 7 on the highest lane.
        //    A within-group reversal would give 5 then A -- reading as
        //    0x5A, a perfectly ordinary byte.
        send(4, PAT);
        if (n_seen != 2) begin
            $display("  FAIL: four lanes produced %0d groups, expected 2", n_seen);
            errors = errors + 1;
        end
        if (seen[0][3:0] !== 4'hA || seen[1][3:0] !== 4'h5) begin
            $display("  FAIL: 4 lanes gave 0x%01h then 0x%01h, expected A then 5",
                     seen[0][3:0], seen[1][3:0]);
            errors = errors + 1;
        end
        $display("  4 lanes: 2 groups, 0x%01h then 0x%01h  (IO3..IO0 per group)",
                 seen[0][3:0], seen[1][3:0]);

        // 3. TWO LANES: four groups of two bits, 10 10 01 01.
        send(2, PAT);
        if (n_seen != 4) begin
            $display("  FAIL: two lanes produced %0d groups, expected 4", n_seen);
            errors = errors + 1;
        end
        if (seen[0][1:0] !== 2'b10 || seen[1][1:0] !== 2'b10 ||
            seen[2][1:0] !== 2'b01 || seen[3][1:0] !== 2'b01) begin
            $display("  FAIL: 2 lanes gave %0b %0b %0b %0b, expected 10 10 01 01",
                     seen[0][1:0], seen[1][1:0], seen[2][1:0], seen[3][1:0]);
            errors = errors + 1;
        end
        $display("  2 lanes: 4 groups, %02b %02b %02b %02b",
                 seen[0][1:0], seen[1][1:0], seen[2][1:0], seen[3][1:0]);

        // 4. EIGHT LANES: one group carrying the whole byte.
        send(8, PAT);
        if (n_seen != 1) begin
            $display("  FAIL: eight lanes produced %0d groups, expected 1", n_seen);
            errors = errors + 1;
        end
        if (seen[0] !== 8'hA5) begin
            $display("  FAIL: 8 lanes gave 0x%02h, expected 0xA5", seen[0]);
            errors = errors + 1;
        end
        $display("  8 lanes: 1 group,  0x%02h", seen[0]);

        // 5. OUTPUT ENABLE: exactly `lanes` bits driven, never more.
        send(1, 8'hFF);
        if (oe_union !== 8'b0000_0001) begin
            $display("  FAIL: one lane drove oe=%08b, expected 00000001", oe_union);
            errors = errors + 1;
        end
        send(2, 8'hFF);
        if (oe_union !== 8'b0000_0011) begin
            $display("  FAIL: two lanes drove oe=%08b, expected 00000011", oe_union);
            errors = errors + 1;
        end
        send(4, 8'hFF);
        if (oe_union !== 8'b0000_1111) begin
            $display("  FAIL: four lanes drove oe=%08b, expected 00001111", oe_union);
            errors = errors + 1;
        end
        send(8, 8'hFF);
        if (oe_union !== 8'b1111_1111) begin
            $display("  FAIL: eight lanes drove oe=%08b, expected 11111111", oe_union);
            errors = errors + 1;
        end
        $display("  output enable: exactly 1, 2, 4 and 8 lanes driven -- never more");

        // 6. Clear when idle, or the master holds the bus between transfers
        //    and the device cannot answer at all.
        @(negedge clk);
        if (io_oe !== {MAX_LANES{1'b0}}) begin
            $display("  FAIL: lanes still driven while idle: %08b", io_oe);
            errors = errors + 1;
        end

        // 7. EXHAUSTIVE ROUND TRIP over every width and every byte value.
        for (w = 0; w < 4; w = w + 1) begin
            if (w == 0)      nl = 1;
            else if (w == 1) nl = 2;
            else if (w == 2) nl = 4;
            else             nl = 8;
            for (b = 0; b < 256; b = b + 1) begin
                send(nl, b[7:0]);
                if (!saw_rx) begin
                    $display("  FAIL: %0d lanes, 0x%02h produced no received byte",
                             nl, b);
                    errors = errors + 1;
                end else if (got_byte !== b[7:0]) begin
                    $display("  FAIL: %0d lanes round trip 0x%02h returned 0x%02h",
                             nl, b, got_byte);
                    errors = errors + 1;
                end
            end
        end
        $display("  round trip: 1024 (width, byte) pairs recovered exactly");

        // 8. A byte with NO symmetry, to pin the mapping independently of
        //    the round trip -- which would pass even if BOTH directions
        //    were reversed consistently.
        send(4, 8'h13);
        if (seen[0][3:0] !== 4'h1 || seen[1][3:0] !== 4'h3) begin
            $display("  FAIL: 0x13 gave 0x%01h then 0x%01h, expected 1 then 3",
                     seen[0][3:0], seen[1][3:0]);
            errors = errors + 1;
        end
        $display("  0x13 at 4 lanes: 0x%01h then 0x%01h -- high nibble first",
                 seen[0][3:0], seen[1][3:0]);

        // 9. An illegal width is reported rather than treated as one lane.
        @(negedge clk);
        lanes = 4'd3;
        @(negedge clk);
        if (!lane_err) begin
            $display("  FAIL: three lanes was not reported as illegal");
            errors = errors + 1;
        end
        lanes = 4'd4;
        @(negedge clk);
        if (lane_err) begin
            $display("  FAIL: four lanes was reported as illegal");
            errors = errors + 1;
        end
        $display("  lane validation: 3 rejected, 4 accepted");

        if (errors == 0)
            $display("PASS: the most significant group goes first and within a group the most significant bit goes on the highest-numbered lane, at one, two, four and eight lanes; exactly as many lanes are driven as the width calls for and none while idle; and serialising then deserialising returns the byte for all 256 values at all four widths");
        else
            $display("FAIL: %0d error(s)", errors);
        $finish;
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_lane_serdes.vhd — the same gearbox in VHDL
-- spi_lane_serdes.vhd
--
-- Chapter 12.2 -- the gearbox between a byte and a group of lanes, in VHDL.
--
-- With two, four or eight lanes a byte becomes four, two or one clock, and
-- a question appears that has no single-lane equivalent: WHICH BIT GOES ON
-- WHICH LANE. The convention is uniform across every multi-lane device:
--
--   * the most significant GROUP goes first, and
--   * within a group, the most significant bit goes on the
--     highest-numbered lane.
--
-- So 0xA5 on four lanes is 0xA then 0x5. Getting the within-group order
-- backwards turns 0xA5 into 0x5A -- a plausible value, not an obvious
-- error.
--
-- OUTPUT ENABLE IS THE SAFETY OUTPUT. In quad and octal mode every data pin
-- is bidirectional, so driving a lane the device is also driving is
-- contention with real current behind it. A four-lane transfer must drive
-- exactly four lanes, never the eight the port is wide.

library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

entity spi_lane_serdes is
    generic (
        MAX_LANES : positive := 8
    );
    port (
        clk       : in  std_logic;
        rst_n     : in  std_logic;

        lanes     : in  unsigned(3 downto 0);   -- 1, 2, 4 or 8

        -- Transmit: load a byte, shift it out most significant group first.
        load      : in  std_logic;
        tx_byte   : in  unsigned(7 downto 0);
        io_out    : out unsigned(MAX_LANES - 1 downto 0);
        io_oe     : out unsigned(MAX_LANES - 1 downto 0);
        tx_active : out std_logic;
        tx_done   : out std_logic;

        -- Receive: capture groups and assemble a byte.
        io_in     : in  unsigned(MAX_LANES - 1 downto 0);
        rx_en     : in  std_logic;
        rx_byte   : out unsigned(7 downto 0);
        rx_valid  : out std_logic;

        groups    : out unsigned(3 downto 0);   -- 8 / lanes, for this width
        lane_err  : out std_logic
    );
end entity;

architecture rtl of spi_lane_serdes is

    -- 8 / lanes, as a table. The legal set is four values, so a table is
    -- both cheaper than a divider and the place an illegal width is caught.
    function group_count(n : unsigned(3 downto 0)) return unsigned is
    begin
        case to_integer(n) is
            when 1      => return to_unsigned(8, 4);
            when 2      => return to_unsigned(4, 4);
            when 4      => return to_unsigned(2, 4);
            when 8      => return to_unsigned(1, 4);
            when others => return to_unsigned(8, 4);   -- flagged by lane_err
        end case;
    end function;

    function bad_lanes(n : unsigned(3 downto 0)) return boolean is
    begin
        return not (to_integer(n) = 1 or to_integer(n) = 2 or
                    to_integer(n) = 4 or to_integer(n) = 8);
    end function;

    signal tx_sr   : unsigned(7 downto 0) := (others => '0');
    signal tx_left : unsigned(3 downto 0) := (others => '0');
    signal rx_sr   : unsigned(7 downto 0) := (others => '0');
    signal rx_left : unsigned(3 downto 0) := (others => '0');

    signal act_r   : std_logic := '0';
    signal done_r  : std_logic := '0';
    signal rxb_r   : unsigned(7 downto 0) := (others => '0');
    signal rxv_r   : std_logic := '0';
    signal grp_r   : unsigned(3 downto 0) := to_unsigned(8, 4);

    -- The active-lane mask and the alignment shift that brings the most
    -- significant group down to the low bits. Initialised so the
    -- combinational outputs below are defined before the first reset.
    signal lane_mask : unsigned(MAX_LANES - 1 downto 0) := (others => '0');
    signal top_shift : natural := 7;

begin

    lane_mask <= to_unsigned(2 ** to_integer(lanes) - 1, MAX_LANES)
                 when not bad_lanes(lanes)
                 else to_unsigned(0, MAX_LANES);
    top_shift <= 8 - to_integer(lanes) when not bad_lanes(lanes) else 7;

    -- The group currently presented: the top `lanes` bits of the shift
    -- register, with bit 7 landing on the highest active lane.
    io_out   <= resize(shift_right(tx_sr, top_shift), MAX_LANES) and lane_mask;
    io_oe    <= lane_mask when act_r = '1' else to_unsigned(0, MAX_LANES);
    lane_err <= '1' when bad_lanes(lanes) else '0';

    tx_active <= act_r;
    tx_done   <= done_r;
    rx_byte   <= rxb_r;
    rx_valid  <= rxv_r;
    groups    <= grp_r;

    gear : process (clk, rst_n)
        variable nxt : unsigned(7 downto 0);
    begin
        if rst_n = '0' then
            tx_sr   <= (others => '0');
            tx_left <= (others => '0');
            act_r   <= '0';
            done_r  <= '0';
            rx_sr   <= (others => '0');
            rx_left <= (others => '0');
            rxb_r   <= (others => '0');
            rxv_r   <= '0';
            grp_r   <= to_unsigned(8, 4);
        elsif rising_edge(clk) then
            done_r <= '0';
            rxv_r  <= '0';
            grp_r  <= group_count(lanes);

            if load = '1' then
                tx_sr   <= tx_byte;
                tx_left <= group_count(lanes);
                act_r   <= '1';
                -- A load also restarts the receive assembly: the group count
                -- is shared between the two directions.
                rx_sr   <= (others => '0');
                rx_left <= group_count(lanes);
            elsif act_r = '1' then
                -- Shift by a whole group. The bits that leave the top are
                -- the ones just presented on the lanes.
                tx_sr <= shift_left(tx_sr, to_integer(lanes));
                if tx_left = 1 then
                    act_r  <= '0';
                    done_r <= '1';
                else
                    tx_left <= tx_left - 1;
                end if;
            end if;

            -- Receive assembles with the same convention in reverse: each
            -- group arrives on the low `lanes` bits of io_in and is shifted
            -- in below the ones already captured.
            if rx_en = '1' and rx_left /= 0 then
                nxt := shift_left(rx_sr, to_integer(lanes)) or
                       resize(io_in and lane_mask, 8);
                rx_sr <= nxt;
                if rx_left = 1 then
                    rxb_r   <= nxt;
                    rxv_r   <= '1';
                    rx_left <= (others => '0');
                else
                    rx_left <= rx_left - 1;
                end if;
            end if;
        end if;
    end process;

end architecture;
Azvya Education Pvt. Ltd.VLSI Mentor
spi_lane_serdes_tb.vhd — the same 1024 round trips in VHDL
-- spi_lane_serdes_tb.vhd
--
-- The same checks as the SystemVerilog and Verilog testbenches: the lane
-- mapping verified against values computed by hand from the convention at
-- every width, then an exhaustive loopback round trip of all 256 byte
-- values at all four widths.

library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

entity spi_lane_serdes_tb is
end entity;

architecture sim of spi_lane_serdes_tb is

    constant MAX_LANES : positive := 8;
    constant PAT : unsigned(7 downto 0) := x"A5";

    signal clk   : std_logic := '0';
    signal rst_n : std_logic := '0';
    signal halt  : boolean   := false;

    signal lanes     : unsigned(3 downto 0) := to_unsigned(1, 4);
    signal load      : std_logic := '0';
    signal tx_byte   : unsigned(7 downto 0) := (others => '0');
    signal io_out    : unsigned(MAX_LANES - 1 downto 0);
    signal io_oe     : unsigned(MAX_LANES - 1 downto 0);
    signal tx_active : std_logic;
    signal tx_done   : std_logic;
    signal io_in     : unsigned(MAX_LANES - 1 downto 0);
    signal rx_en     : std_logic;
    signal rx_byte   : unsigned(7 downto 0);
    signal rx_valid  : std_logic;
    signal groups    : unsigned(3 downto 0);
    signal lane_err  : std_logic;

    -- Captured groups, for the mapping check.
    type grp_arr is array (0 to 7) of unsigned(MAX_LANES - 1 downto 0);
    signal seen     : grp_arr;
    signal n_seen   : natural := 0;
    signal saw_rx   : std_logic := '0';
    signal got_byte : unsigned(7 downto 0) := (others => '0');
    signal oe_union : unsigned(MAX_LANES - 1 downto 0) := (others => '0');
    signal clr_cap  : std_logic := '0';

    signal errors : natural := 0;

begin

    clk <= not clk after 5 ns when not halt else '0';

    -- Loopback: what the gearbox drives is what it receives.
    io_in <= io_out;
    rx_en <= tx_active;

    dut : entity work.spi_lane_serdes
        generic map (MAX_LANES => MAX_LANES)
        port map (
            clk => clk, rst_n => rst_n, lanes => lanes,
            load => load, tx_byte => tx_byte,
            io_out => io_out, io_oe => io_oe,
            tx_active => tx_active, tx_done => tx_done,
            io_in => io_in, rx_en => rx_en,
            rx_byte => rx_byte, rx_valid => rx_valid,
            groups => groups, lane_err => lane_err
        );

    -- Capture. The clear arrives on its own signal because two processes
    -- driving one signal is a multiple-driver error in VHDL.
    capture : process (clk)
    begin
        if rising_edge(clk) then
            if clr_cap = '1' then
                n_seen   <= 0;
                saw_rx   <= '0';
                oe_union <= (others => '0');
            elsif rst_n = '1' then
                if tx_active = '1' and n_seen < 8 then
                    seen(n_seen) <= io_out;
                    n_seen       <= n_seen + 1;
                    oe_union     <= oe_union or io_oe;
                end if;
                if rx_valid = '1' then
                    saw_rx   <= '1';
                    got_byte <= rx_byte;
                end if;
            end if;
        end if;
    end process;

    stim : process
        variable errs  : natural := 0;
        variable guard : natural;
        variable nl    : natural;

        procedure send(n : natural; bb : unsigned(7 downto 0)) is
        begin
            wait until falling_edge(clk);
            clr_cap <= '1';
            wait until falling_edge(clk);
            clr_cap <= '0';
            lanes   <= to_unsigned(n, 4);
            tx_byte <= bb;
            load    <= '1';
            wait until falling_edge(clk);
            load    <= '0';
            guard := 0;
            while tx_active = '1' and guard < 32 loop
                wait until falling_edge(clk);
                guard := guard + 1;
            end loop;
            wait until falling_edge(clk);
        end procedure;
    begin
        for k in 0 to 2 loop
            wait until falling_edge(clk);
        end loop;
        rst_n <= '1';
        wait until falling_edge(clk);

        -- 1. SINGLE LANE, most significant bit first.
        send(1, PAT);
        if n_seen /= 8 then
            report "  FAIL: one lane produced " & integer'image(n_seen) &
                   " groups, expected 8";
            errs := errs + 1;
        end if;
        for i in 0 to 7 loop
            if seen(i)(0) /= PAT(7 - i) then
                report "  FAIL: one-lane group " & integer'image(i) &
                       " carried the wrong bit";
                errs := errs + 1;
            end if;
        end loop;
        report "  1 lane:  8 groups, MSB first";

        -- 2. FOUR LANES. High nibble first, bit 7 on the highest lane.
        --    A within-group reversal would give 5 then A -- reading as
        --    0x5A, a perfectly ordinary byte.
        send(4, PAT);
        if n_seen /= 2 then
            report "  FAIL: four lanes produced " & integer'image(n_seen) &
                   " groups, expected 2";
            errs := errs + 1;
        end if;
        if seen(0)(3 downto 0) /= x"A" or seen(1)(3 downto 0) /= x"5" then
            report "  FAIL: four lanes did not give A then 5";
            errs := errs + 1;
        end if;
        -- integer'image prints decimal, so the values are reported as
        -- decimal rather than given a misleading 0x prefix: the high nibble
        -- of 0xA5 is 10, the low nibble 5.
        report "  4 lanes: 2 groups, high nibble " &
               integer'image(to_integer(seen(0)(3 downto 0))) &
               " then low nibble " &
               integer'image(to_integer(seen(1)(3 downto 0)));

        -- 3. TWO LANES: four groups of two bits, 10 10 01 01.
        send(2, PAT);
        if n_seen /= 4 then
            report "  FAIL: two lanes produced " & integer'image(n_seen) &
                   " groups, expected 4";
            errs := errs + 1;
        end if;
        if seen(0)(1 downto 0) /= "10" or seen(1)(1 downto 0) /= "10" or
           seen(2)(1 downto 0) /= "01" or seen(3)(1 downto 0) /= "01" then
            report "  FAIL: two lanes did not give 10 10 01 01";
            errs := errs + 1;
        end if;
        report "  2 lanes: 4 groups, 10 10 01 01";

        -- 4. EIGHT LANES: one group carrying the whole byte.
        send(8, PAT);
        if n_seen /= 1 then
            report "  FAIL: eight lanes produced " & integer'image(n_seen) &
                   " groups, expected 1";
            errs := errs + 1;
        end if;
        if seen(0) /= x"A5" then
            report "  FAIL: eight lanes did not give 0xA5"; errs := errs + 1;
        end if;
        report "  8 lanes: 1 group,  0xA5";

        -- 5. OUTPUT ENABLE: exactly `lanes` bits driven, never more.
        send(1, x"FF");
        if oe_union /= "00000001" then
            report "  FAIL: one lane drove the wrong enable mask";
            errs := errs + 1;
        end if;
        send(2, x"FF");
        if oe_union /= "00000011" then
            report "  FAIL: two lanes drove the wrong enable mask";
            errs := errs + 1;
        end if;
        send(4, x"FF");
        if oe_union /= "00001111" then
            report "  FAIL: four lanes drove the wrong enable mask";
            errs := errs + 1;
        end if;
        send(8, x"FF");
        if oe_union /= "11111111" then
            report "  FAIL: eight lanes drove the wrong enable mask";
            errs := errs + 1;
        end if;
        report "  output enable: exactly 1, 2, 4 and 8 lanes driven -- never more";

        -- 6. Clear when idle, or the master holds the bus between transfers.
        wait until falling_edge(clk);
        if io_oe /= 0 then
            report "  FAIL: lanes still driven while idle"; errs := errs + 1;
        end if;

        -- 7. EXHAUSTIVE ROUND TRIP over every width and every byte value.
        for w in 0 to 3 loop
            case w is
                when 0      => nl := 1;
                when 1      => nl := 2;
                when 2      => nl := 4;
                when others => nl := 8;
            end case;
            for b in 0 to 255 loop
                send(nl, to_unsigned(b, 8));
                if saw_rx /= '1' then
                    report "  FAIL: a transfer produced no received byte";
                    errs := errs + 1;
                elsif to_integer(got_byte) /= b then
                    report "  FAIL: a round trip returned the wrong byte";
                    errs := errs + 1;
                end if;
            end loop;
        end loop;
        report "  round trip: 1024 (width, byte) pairs recovered exactly";

        -- 8. A byte with NO symmetry, to pin the mapping independently of
        --    the round trip -- which would pass even if BOTH directions
        --    were reversed consistently.
        send(4, x"13");
        if seen(0)(3 downto 0) /= x"1" or seen(1)(3 downto 0) /= x"3" then
            report "  FAIL: 0x13 did not give 1 then 3"; errs := errs + 1;
        end if;
        report "  0x13 at 4 lanes: 0x1 then 0x3 -- high nibble first";

        -- 9. An illegal width is reported rather than treated as one lane.
        wait until falling_edge(clk);
        lanes <= to_unsigned(3, 4);
        wait until falling_edge(clk);
        if lane_err /= '1' then
            report "  FAIL: three lanes was not reported as illegal";
            errs := errs + 1;
        end if;
        lanes <= to_unsigned(4, 4);
        wait until falling_edge(clk);
        if lane_err = '1' then
            report "  FAIL: four lanes was reported as illegal";
            errs := errs + 1;
        end if;
        report "  lane validation: 3 rejected, 4 accepted";

        errors <= errs;
        if errs = 0 then
            report "PASS: the most significant group goes first and within a group the most significant bit goes on the highest-numbered lane, at one, two, four and eight lanes; exactly as many lanes are driven as the width calls for and none while idle; and serialising then deserialising returns the byte for all 256 values at all four widths";
        else
            report "FAIL: " & integer'image(errs) & " error(s)" severity error;
        end if;
        halt <= true;
        wait;
    end process;

end architecture;

Parity

All three implement the same gearbox: identical ports and generics, most significant group first with bit 7 on the highest active lane, a group count held as a table, an output enable derived from the lane mask and cleared when idle, and an illegal width reported. All three testbenches check the same hand-computed mappings at all four widths, run the same 1024-pair exhaustive round trip, and verify the same asymmetric byte — reporting identical results throughout.

7. Why a Verification Engineer Cares

Azvya Education Pvt. Ltd.VLSI Mentor
spi_lane_serdes.sva — the round trip, and what it cannot prove
   // 1. ROUND TRIP. Serialising then deserialising returns the byte. Cheap,
   //    exhaustive over 256 values, and it catches field overlap -- the same
   //    property Chapter 10.4 used for a command codec.
   a_round_trip : assert property (
       @(posedge clk) disable iff (!rst_n)
           (rx_valid) |-> (rx_byte == $past(tx_byte, tx_latency)))
       else $error("the byte did not survive the round trip");

   // 2. ...AND WHAT IT CANNOT PROVE. A round trip passes if BOTH directions
   //    are reversed consistently, because the errors cancel. So the
   //    absolute mapping needs its own property, on a value whose nibbles
   //    are not each other's bit-reverse.
   a_absolute_mapping : assert property (
       @(posedge clk) disable iff (!rst_n)
           (first_group && lanes == 4) |->
               (io_out[3:0] == tx_byte[7:4]))
       else $error("the first group is not the high nibble");

   // 3. THE ENABLE MATCHES THE WIDTH, exactly. Driving a lane the width does
   //    not include is contention on a bidirectional pin -- and because the
   //    enable is derived from the same mask as the data, this should be
   //    structurally impossible rather than merely checked.
   a_enable_exact : assert property (
       @(posedge clk) disable iff (!rst_n)
           (tx_active) |-> ($countones(io_oe) == lanes))
       else $error("the number of driven lanes does not match the width");

   // 4. NOTHING DRIVEN WHEN IDLE. A controller holding the lanes between
   //    transfers prevents the device from answering, and the symptom is a
   //    device that looks dead rather than one that looks contended.
   a_idle_released : assert property (
       @(posedge clk) disable iff (!rst_n)
           (!tx_active) |-> (io_oe == '0))
       else $error("lanes were driven while idle");

   // 5. GROUP COUNT. A byte takes exactly 8 / lanes groups -- so a width
   //    change mid-byte, which is a configuration error, cannot silently
   //    produce a byte assembled from mismatched group sizes.
   a_group_count : assert property (
       @(posedge clk) disable iff (!rst_n)
           (rx_valid) |-> (groups_this_byte == (8 / lanes)))
       else $error("a byte was assembled from the wrong number of groups");

Properties 1 and 2 together are the lesson, and it generalises well beyond this block. A round trip is a powerful and incomplete specification: it proves the two directions are mutually consistent and says nothing about whether either is right. Any design with an inverse has this shape — encoders and decoders, packers and unpackers, serialisers and deserialisers — and every one of them needs at least one absolute check alongside the round trip.

Property 3 is worth contrasting with its own implementation. The gearbox derives the enable from the mask that narrows the data, so the property should be unfalsifiable by construction — and that is the right relationship between an assertion and a design. An assertion that cannot fail because the structure prevents it is not a wasted assertion; it is a statement that the structure is doing the work, and it will fail loudly if someone later separates the two.

Coverage must cross the width with the data pattern, because the symmetric values hide mapping errors:

Azvya Education Pvt. Ltd.VLSI Mentor
spi_lane_cg.sv — symmetry is the axis that matters
   covergroup spi_lane_cg @(posedge clk iff load);
       cp_lanes : coverpoint lanes {
           bins one   = {1};
           bins two   = {2};
           bins four  = {4};
           bins eight = {8};
           illegal_bins impossible = {0, 3, [5:7], [9:15]};
       }

       // The SYMMETRY of the payload is what decides whether a mapping
       // error is visible. 0xA5 and 0xFF are symmetric under some wrong
       // mappings; 0x13 is not. A suite using only round numbers tests the
       // gearbox against the cases most likely to hide its own bugs.
       cp_symmetry : coverpoint byte_class {
           bins all_zeros    = {B_00};        // invisible to everything
           bins all_ones     = {B_FF};        // invisible to everything
           bins nibble_sym   = {B_NIBBLE_SYM};// 0xA5 -- hides some errors
           bins asymmetric   = {B_ASYM};      // 0x13 -- hides nothing
       }

       cp_dir : coverpoint transfer_dir { bins tx = {0}; bins rx = {1}; }

       // Exactly the driven-lane counts that are legal, and nothing else.
       cp_oe : coverpoint $countones(io_oe) {
           bins idle  = {0};
           bins one   = {1};
           bins two   = {2};
           bins four  = {4};
           bins eight = {8};
           illegal_bins wrong = {3, [5:7]};
       }

       x_lanes_symmetry : cross cp_lanes, cp_symmetry;
   endgroup

cp_symmetry is the coverpoint this chapter argues for and the one almost never written. A test suite naturally reaches for 0x00, 0xFF, 0xAA and 0x55 — and every one of those is invisible to at least one mapping error. The asymmetric bin is the only one that is not, and crossing it with the width is what proves the mapping rather than merely exercising it.

8. Why an FPGA or ASIC Engineer Cares

Derive the output enable from the lane mask. One mask feeding both the data and the enable makes it structurally impossible to drive a lane the width does not include. Two independent computations of the same thing will eventually disagree, and the disagreement is contention.

Release the lanes when idle. A held bus produces a device that appears dead, which is a much harder symptom to trace than a contended one.

Hold 8 / lanes as a table. Four legal values, and the table is also where an illegal width is caught — one structure doing both jobs, and no divider.

Test with an asymmetric byte. 0xAA, 0x55, 0x00 and 0xFF are all invisible to at least one mapping error. 0x13 is not, and one such value in a test list is worth more than a dozen round numbers.

Remember what enabling quad mode gives up. WP# and HOLD#/RESET# become data lanes, so write protection and — on many parts — the hardware reset are gone while the mode is active. If the board depends on either, that dependency must be re-examined before the mode is enabled, not after.

Treat the mode bit as persistent state. On many parts the quad-enable bit is non-volatile, so it survives power cycling and a controller that assumes single-lane at reset will be wrong. This is Chapter 11.5's warm-boot problem with a different bit.

9. Failure Signature — Nibble-Swapped Data on Exactly One Board Variant

Symptom. A quad-mode read returns data whose bytes are correct in content but have their nibbles exchanged — 0x12 0x34 reads back as 0x21 0x43. It happens on one board variant and not another, with identical firmware and identical flash part numbers.

What "correct content, exchanged nibbles" establishes. Every bit arrived. Nothing was lost, added or corrupted — only rearranged, and rearranged in a completely regular way. So this is a wiring or mapping fault rather than a timing or protocol one, and that single observation eliminates most of what the previous three modules were about.

Plausible mechanisms.

  • The IO2/IO3 pair swapped on the board. Exchanging two lanes within a nibble does not swap the nibbles, so this does not fit — but exchanging IO0↔IO2 and IO1↔IO3 would produce a bit permutation within each nibble rather than a nibble swap either.
  • A group-order error in the controller, sending the low nibble first. This produces exactly a nibble swap — but it would affect both board variants, since it is in the firmware.
  • The board variants differ in lane routing, and one has IO0..IO3 connected in a different order — which, combined with a controller that assumes the standard order, produces a fixed permutation.
  • One variant's flash is in a different mode — DDR, or a different address width — which would shift rather than permute.
  • A level shifter or buffer on two of the four lanes only, which would produce timing-dependent corruption rather than a clean swap.

The discriminating observation. "Identical firmware, one variant affected" is decisive: a firmware-side group-order error cannot be variant-specific, so the fault must be on the board. That immediately rules out the controller and points at the routing.

Then determine which permutation, and there is a clean way to do it without a scope. Read a byte whose nibbles are distinguishable and whose bits within each nibble are too — 0x13 gives 0001 and 0011, and each of the plausible mis-routings maps it to a different value. Compare against a short table computed by hand and the permutation identifies itself.

The stronger test, if the flash can be written on the affected variant: write and read back on the same board. If a round trip on the bad variant succeeds — the value written comes back — the error is symmetric, which means it is a routing permutation applied identically in both directions. That is the point §7's property 2 makes: a round trip passing is consistent with both directions being wrong in the same way.

The fix. For a wiring error already in hardware, the lane order becomes a configuration of the controller: a per-board lane permutation applied in the gearbox. That is a small amount of logic and it turns a board respin into a table entry — which is the same argument Chapter 10.1 made about device profiles, arriving at the board level.

Why it reaches production. Because quad lane routing is often done for signal-integrity reasons — swapping lanes to shorten a trace or avoid a via — on the correct understanding that the lanes are interchangeable electrically. They are not interchangeable logically, and the layout engineer has no reason to know that unless the constraint is written down.

10. Common Misconceptions

11. Reason It Through

Work this before reading the answer.

A controller supports 1-1-1 and 1-1-4. A board has a quad-capable flash, and the firmware enables quad mode by setting the non-volatile quad-enable bit, then reads with 0x6B (quad output).

It works. Some weeks later the same firmware is loaded onto a fresh board from the same batch, and the very first read — before quad mode is enabled — returns garbage. What happened, and what does the order of events tell you?

Start with the difference. On the working board, quad mode was enabled at some point in the past and the bit is non-volatile, so it is still set. On the fresh board it is not.

But the failing read happens BEFORE quad mode is enabled, and on the fresh board that read is a single-lane read on a part in its default state. That should work — so the firmware cannot simply be relying on the quad bit.

So what is different about the first read on the fresh board? Nothing about the flash. The difference must be in the controller.

And there it is. On the working board, the controller was left configured for quad mode by the previous run — its lane count, and more importantly its output enable mask, are still set to four lanes. A single-lane read issued with a four-lane enable mask drives IO1, IO2 and IO3 as well as IO0. On a part in single-lane mode, IO1 is the device's output — so the controller is fighting the device on every returned bit.

Why did it ever work, then? Because on the working board the quad bit was set, so the device also treated the lanes as bidirectional and released them during the command phase. The contention window was small enough to survive. On the fresh board the device is driving IO1 continuously as MISO, and the controller is driving it too.

What does the ORDER of events tell you? That the fault is in the controller's reset state, not in its configured state. The working board hid a controller bug because the device was in a matching mode; the fresh board exposed it because the device was in its default one. This is the general shape worth remembering: a controller and a device that are both mis-configured in the same direction can appear to work, and the failure surfaces when only one of them is reset.

The fix, in two parts.

The controller must reset to single-lane with a one-bit enable mask — the same argument Chapter 11.2 made for resetting to the plain read at the slowest divisor, now applied to width. Reset to the configuration that works against a device in its default state.

And the firmware must not treat the quad-enable bit as something it set once. It should read it back and configure the controller to match what the device actually reports, rather than what the firmware believes it wrote — which is Chapter 11.4's "confirm the latch" discipline applied to a mode bit.

One more observation worth making. The non-volatile quad bit means a board can arrive from manufacturing in either state depending on what test software ran on it. A design that works only in one of those states will fail on some fraction of units, apparently at random — and the fraction will depend on the factory's process rather than on anything in the design.

12. Understanding Check

13. Summary

Multi-lane SPI repurposes the write-protect and hold pins as IO2 and IO3, which is why it needs no new pins and why enabling it gives up write protection and often the hardware reset. The mode bit is frequently non-volatile, so it survives power cycling.

Single-lane SPI had one convention to get right; multi-lane has two. The most significant group goes first, and within a group the most significant bit goes on the highest-numbered lane. Both halves produce 0x5A from 0xA5 when reversed, by different mechanisms, and are distinguishable only on an asymmetric byte.

Every data pin becomes bidirectional, so the output enable stops being a convenience and becomes a safety signal: exactly as many lanes as the width calls for, and none when idle. Deriving it from the same mask that narrows the data makes over-driving structurally impossible rather than merely checked.

In hardware, one shift register serves every width in both directions; 8 / lanes is a table, not a division; and the alignment shift is the only thing that changes with the width.

For verification, the round trip is powerful and incomplete — it proves the two directions agree and not that either is right, so an absolute check belongs beside it. And the coverage axis that matters is the payload's symmetry, because 0x00, 0xFF, 0xAA and 0x55 are each invisible to at least one mapping error while 0x13 is invisible to none.

14. What Comes Next

The lanes are now understood as wires and as a convention. What has not been examined is the consequence of their being bidirectional.

Chapter 12.3 — Lane Widths per Phase and Direction Changes reads the 1-1-4 / 1-4-4 / 4-4-4 notation properly, and then asks the question the notation does not answer: which phases the master drives and which the device does. A single-lane read never reverses a wire; a multi-lane read reverses four at once — and the chapter's finding is that this is the second reason a quad read specifies dummy cycles, alongside Chapter 11.2's array access time. A quad read specified with no dummy phase has no turnaround window at all, and the walker in that chapter reports it.

Continue learning