Skip to content
VLSI Mentor

SPI · Module 10

Address Fields, Dummy Cycles, and Burst Behaviour

Why three address bytes make 16 MB a category boundary, why dummy latency is counted in cycles and never rounds to bytes, what a device does at the end of a burst, and the planner that turns those numbers into a cycle-accurate schedule.

Chapter 10.4 packed the register address into the command byte. That works while the map is small. A 128 Mbit flash has sixteen million addresses and seven spare bits, so the address becomes a phase of its own — and three new numbers appear with it.

A command table entry reads: 0x0B — Fast Read — 3 address bytes, 8 dummy clocks. Another reads 0xEB — Fast Read Quad I/O — 3 address bytes, 6 dummy clocks. Why is the second number 6, and what happens if you send a byte?

Six is not a mistake and it does not round. Sending a byte where six cycles were asked for shifts every payload bit by two positions, and the data that comes back looks like data.

1. Address Width — and the Ceiling It Creates

The command table states the number of address bytes, and it follows directly from the device's capacity:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   address bytes    reachable addresses    capacity
        1                   256             256 B
        2                 65 536             64 KB
        3             16 777 216             16 MB    ← the common case
        4          4 294 967 296              4 GB

Three bytes reaching exactly 16 MB is why 128 Mbit is a boundary in this whole product category. A 128 Mbit part is 16 MB, which is precisely what three address bytes address. Everything larger needs a fourth byte, and the industry's response was not to move everyone to four — it was to add a mode.

So parts above 128 Mbit typically support both:

  • 3-byte mode, reaching the bottom 16 MB only, for compatibility with every existing controller and boot ROM.
  • 4-byte mode, reaching the whole device, entered by a command or by a non-volatile configuration bit.
  • Sometimes 4-byte opcodes — a parallel set of commands that take four address bytes without changing any mode.

Two consequences matter more than the table.

An address that does not fit is not rejected. The device receives whatever bytes arrive and uses them. Sending a 4-byte address to a part in 3-byte mode means the device consumes the first three as the address and the fourth as the first byte of data or the dummy phase — so a write corrupts the wrong location and a read returns the wrong data, with no error anywhere. The check must live in the controller, which is what spec_error does in §6.

Mode is persistent state. A part left in 4-byte mode by previous software, then reset without its interface being reset, greets the boot ROM with an address phase one byte longer than expected. This is a genuinely common bring-up failure and it survives a warm reset.

2. Byte Order

SPI devices send addresses most-significant byte first, essentially without exception. The table rarely says so, because the timing diagram shows it.

This is worth stating explicitly for one reason: it is the opposite of the byte order most host processors store addresses in. A little-endian 32-bit address in memory has its least significant byte first, and a driver that memcpy's the address into a transmit buffer sends it backwards. The result is an access to a wildly wrong address that is nonetheless a perfectly valid one — so the device answers, and the data is wrong in a way that looks like corruption rather than like a bug.

3. Dummy Latency — Counted in Cycles

This is the number that does the most damage when misread.

Why it exists is Chapter 4.5's subject: the device needs internal time between receiving the address and producing data, and the dummy phase is where that time is spent. What this chapter adds is how the number is specified — and it is specified in SCLK cycles, not bytes.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   0x03  Read              0 dummy cycles
   0x0B  Fast Read         8 dummy cycles
   0x3B  Dual Output       8 dummy cycles
   0xBB  Dual I/O          4 dummy cycles
   0xEB  Quad I/O          6 dummy cycles    ← not a byte
   0x6B  Quad Output       8 dummy cycles

The values that are not multiples of eight are not exotic. They appear because the dummy count is set by the device's internal timing measured in clock periods, and there is no reason for that to land on a byte boundary. On the multi-line commands it frequently does not, because those commands move more bits per cycle and so need fewer cycles to cover the same internal delay.

What happens when you send a byte instead of six cycles:

Six dummy cycles versus a dummy byte

8 cycles
Two byte-time lanes for the same read command. Both send a command byte and three address bytes. The first spends six dummy cycles and then returns correct data bytes. The second spends eight dummy cycles and returns corrupted bytes because capture began two cycles late.payload beginspayload beginsas specifiedCMDA2A1A06 cycD0D1D2dummy byteCMDA2A1A08 cyc??D0?D1?t0t1t2t3t4t5t6t7
Figure 1 — the same read, with the dummy phase right and wrong. The upper lane spends the six cycles the datasheet asked for and the payload lands where expected. The lower lane spends eight, so capture begins two cycles late and every byte returned is a splice of two adjacent ones.

The lower lane is the important one. It does not return garbage — it returns bytes assembled from the last two bits of one real byte and the first six of the next. Those bytes are plausible values. A driver checking for a non-zero response is satisfied, a checksum over a block fails, and the investigation starts at the memory rather than at the dummy count.

4. Burst Behaviour — What the Device Does at the End

The command table's last contribution is what happens when a burst runs past something.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   read past the top of the device   →  wrap to address 0, or return undefined
   write past a page boundary        →  WRAP within the page  (Chapter 7.2)
   burst longer than a stated limit  →  wrap, stall, or repeat one location
   CS released mid-burst             →  abort (Chapter 8.7)

Reads and writes differ, and the asymmetry is the point. A read burst normally runs freely across the whole device — the internal pointer just advances. A write burst wraps inside the page, because a page is the program granularity, and writing 300 bytes to a 256-byte page does not write 300 bytes: it writes the first 256 and then overwrites the first 44 of them with the remainder. Chapter 9.3 built the splitter that prevents this; the table is where you find the page size to give it.

The burst limit, where one is stated, often has no stated mechanism. A part specifying "up to 16 consecutive registers" may wrap to the block start, stop advancing and repeat one location, or enter reserved space. All three occur, the difference is invisible on the bench, and the only reliable answer is to test the boundary deliberately.

5. The Whole Shape, in Order

A sequence diagram of a fast read. The master asserts chip select, sends a one-byte opcode, sends three address bytes most significant first, clocks a dummy phase of a stated number of cycles during which neither side sends meaningful data, then receives payload bytes before releasing chip select.Opcode, address, dummy, datamasterdeviceCS assertsopcode — 8 bitsaddress — MSB bytefirstdummy — N CYCLES,not bytespayload beginsexactly here… pointer advancesCS releases
Figure 2 — the four phases of a device-defined read. Each phase's length comes from a different column of the command table, and any of the middle two may be absent entirely.

Four phases, and the two in the middle are optional per command. A status read has neither; a sector erase has an address and no dummy and no data; a write-enable is the opcode alone. A planner that emits a zero-length phase rather than skipping it produces a state nothing downstream expects — which is the first property §6's testbench checks.

6. Building the Latency Planner — Three HDLs

The circuit

Circuit. A phase walker that converts the three datasheet numbers into a cycle-accurate schedule.

State. The remaining length of the current phase, plus the three phase lengths captured at start.

Datapath. All three lengths are converted to cycles once, at start — address bytes multiplied by eight, dummy taken as given, data bytes multiplied by eight — so nothing downstream has to remember which datasheet numbers were bytes and which were cycles. That single conversion point is the design's main defence against the error of §3.

Control. A four-state walk with empty phases skipped, not entered for zero cycles.

Clock and reset. One tick per SCLK cycle; asynchronous active-low reset.

Enables. latency is published at start — command bits plus address bits plus dummy cycles — so a consumer knows when to begin capturing before the transfer has run.

Timing. Per-phase cycle counters are maintained as the walk proceeds, and they must sum to the total. That is the property worth asserting: a walker that loses or double-counts a single cycle produces a latency that is almost right, which is the hardest kind of error to notice.

Synthesis. Three length registers, a down-counter, four accounting counters and a small state machine.

Limitations. It plans; it does not drive. Turning the plan into pins is Chapter 10.6's job, and keeping the two separate means one driver serves every device.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_latency_plan.sv — three numbers in, a cycle-accurate schedule out
// spi_latency_plan.sv
//
// Chapter 10.5 -- address width, dummy cycles and burst length, turned
// into a cycle-accurate plan.
//
// Three datasheet numbers decide when the first payload bit appears:
//
//   the command length        -- almost always 8 bits,
//   the address width         -- in BYTES, so 8 bits each,
//   the dummy latency         -- in CYCLES, and NOT necessarily a
//                                multiple of eight.
//
// That last asymmetry is the one that catches people. A part specifying
// "6 dummy cycles" does not want a dummy byte, and a driver that sends one
// shifts every payload byte by two bit positions -- the failure signature
// of Chapter 9.4 arriving from a completely different cause.
//
// This block walks the phases, counts the SCLK cycles each one occupies,
// and reports the LATENCY -- the number of cycles before the first data
// bit -- along with a per-phase accounting that must sum to the total.
//
// Empty phases are SKIPPED, not entered for zero cycles. A device with no
// address phase and no dummy phase goes straight from command to data, and
// a planner that enters a zero-length phase emits a state nothing
// downstream expects.

module spi_latency_plan #(
    parameter int CNT_W    = 16,
    parameter int CMD_BITS = 8
) (
    input  logic               clk,          // one tick per SCLK cycle
    input  logic               rst_n,

    input  logic               start,
    input  logic [2:0]         addr_bytes,   // 0..4
    input  logic [5:0]         dummy_cycles, // CYCLES, not bytes
    input  logic [CNT_W-1:0]   data_bytes,

    output logic [1:0]         phase,        // 0=cmd 1=addr 2=dummy 3=data
    output logic               busy,
    output logic               done,

    output logic [CNT_W-1:0]   latency,      // cycles before the first data bit
    output logic [CNT_W-1:0]   total_cycles,

    // The plan's own accounting. These must sum to total_cycles, which is
    // the property worth asserting: a phase walker that loses or
    // double-counts a cycle produces a latency that is almost right.
    output logic [CNT_W-1:0]   cyc_cmd,
    output logic [CNT_W-1:0]   cyc_addr,
    output logic [CNT_W-1:0]   cyc_dummy,
    output logic [CNT_W-1:0]   cyc_data
);

    localparam logic [1:0] P_CMD   = 2'd0;
    localparam logic [1:0] P_ADDR  = 2'd1;
    localparam logic [1:0] P_DUMMY = 2'd2;
    localparam logic [1:0] P_DATA  = 2'd3;

    logic [CNT_W-1:0] addr_len, dummy_len, data_len;
    logic [CNT_W-1:0] rem;

    logic [1:0]       nxt_phase;
    logic [CNT_W-1:0] nxt_rem;
    logic             nxt_last;

    // Where to go when the current phase runs out, skipping every phase the
    // profile gives a length of zero.
    always_comb begin
        nxt_phase = phase;
        nxt_rem   = {CNT_W{1'b0}};
        nxt_last  = 1'b0;
        case (phase)
            P_CMD: begin
                if      (addr_len  != 0) begin nxt_phase = P_ADDR;  nxt_rem = addr_len;  end
                else if (dummy_len != 0) begin nxt_phase = P_DUMMY; nxt_rem = dummy_len; end
                else if (data_len  != 0) begin nxt_phase = P_DATA;  nxt_rem = data_len;  end
                else                           nxt_last  = 1'b1;
            end
            P_ADDR: begin
                if      (dummy_len != 0) begin nxt_phase = P_DUMMY; nxt_rem = dummy_len; end
                else if (data_len  != 0) begin nxt_phase = P_DATA;  nxt_rem = data_len;  end
                else                           nxt_last  = 1'b1;
            end
            P_DUMMY: begin
                if      (data_len  != 0) begin nxt_phase = P_DATA;  nxt_rem = data_len;  end
                else                           nxt_last  = 1'b1;
            end
            default:                           nxt_last  = 1'b1;
        endcase
    end

    always_ff @(posedge clk or negedge rst_n) begin
        if (!rst_n) begin
            phase        <= P_CMD;
            busy         <= 1'b0;
            done         <= 1'b0;
            addr_len     <= {CNT_W{1'b0}};
            dummy_len    <= {CNT_W{1'b0}};
            data_len     <= {CNT_W{1'b0}};
            rem          <= {CNT_W{1'b0}};
            latency      <= {CNT_W{1'b0}};
            total_cycles <= {CNT_W{1'b0}};
            cyc_cmd      <= {CNT_W{1'b0}};
            cyc_addr     <= {CNT_W{1'b0}};
            cyc_dummy    <= {CNT_W{1'b0}};
            cyc_data     <= {CNT_W{1'b0}};
        end else begin
            done <= 1'b0;

            if (start && !busy) begin
                // Every length is converted to CYCLES here, once, so that
                // nothing downstream has to remember which datasheet
                // numbers were bytes and which were cycles.
                addr_len  <= {{(CNT_W-6){1'b0}}, addr_bytes, 3'b000};  // x8
                dummy_len <= {{(CNT_W-6){1'b0}}, dummy_cycles};        // as given
                data_len  <= data_bytes << 3;                          // x8

                latency <= CNT_W'(CMD_BITS)
                         + {{(CNT_W-6){1'b0}}, addr_bytes, 3'b000}
                         + {{(CNT_W-6){1'b0}}, dummy_cycles};
                total_cycles <= CNT_W'(CMD_BITS)
                         + {{(CNT_W-6){1'b0}}, addr_bytes, 3'b000}
                         + {{(CNT_W-6){1'b0}}, dummy_cycles}
                         + (data_bytes << 3);

                phase     <= P_CMD;
                rem       <= CNT_W'(CMD_BITS);
                busy      <= 1'b1;
                cyc_cmd   <= {CNT_W{1'b0}};
                cyc_addr  <= {CNT_W{1'b0}};
                cyc_dummy <= {CNT_W{1'b0}};
                cyc_data  <= {CNT_W{1'b0}};
            end else if (busy) begin
                // Account for the cycle being spent right now, before
                // deciding whether it was the phase's last.
                case (phase)
                    P_CMD:   cyc_cmd   <= cyc_cmd   + 1'b1;
                    P_ADDR:  cyc_addr  <= cyc_addr  + 1'b1;
                    P_DUMMY: cyc_dummy <= cyc_dummy + 1'b1;
                    default: cyc_data  <= cyc_data  + 1'b1;
                endcase

                if (rem == CNT_W'(1)) begin
                    if (nxt_last) begin
                        busy <= 1'b0;
                        done <= 1'b1;
                    end else begin
                        phase <= nxt_phase;
                        rem   <= nxt_rem;
                    end
                end else begin
                    rem <= rem - 1'b1;
                end
            end
        end
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_latency_plan_tb.sv — eight real device shapes, checked against their own arithmetic
// spi_latency_plan_tb.sv
//
// Every expectation here is computed from the datasheet numbers by the
// testbench's own arithmetic, not read back from the DUT. The two
// properties that matter are that the first data cycle lands exactly at
// the reported latency, and that the per-phase cycle counts sum to the
// total -- a walker that loses or double-counts a cycle produces a
// latency that is almost right, which is the hardest kind to notice.

`timescale 1ns/1ps

module spi_latency_plan_tb;

    localparam int CNT_W    = 16;
    localparam int CMD_BITS = 8;

    localparam logic [1:0] P_CMD   = 2'd0;
    localparam logic [1:0] P_ADDR  = 2'd1;
    localparam logic [1:0] P_DUMMY = 2'd2;
    localparam logic [1:0] P_DATA  = 2'd3;

    logic clk = 1'b0;
    logic rst_n = 1'b0;
    always #5 clk = ~clk;

    logic             start = 1'b0;
    logic [2:0]       addr_bytes = 3'd0;
    logic [5:0]       dummy_cycles = 6'd0;
    logic [CNT_W-1:0] data_bytes = {CNT_W{1'b0}};

    logic [1:0]       phase;
    logic             busy, done;
    logic [CNT_W-1:0] latency, total_cycles;
    logic [CNT_W-1:0] cyc_cmd, cyc_addr, cyc_dummy, cyc_data;

    int errors = 0;

    spi_latency_plan #(.CNT_W(CNT_W), .CMD_BITS(CMD_BITS)) dut (
        .clk(clk), .rst_n(rst_n), .start(start),
        .addr_bytes(addr_bytes), .dummy_cycles(dummy_cycles),
        .data_bytes(data_bytes),
        .phase(phase), .busy(busy), .done(done),
        .latency(latency), .total_cycles(total_cycles),
        .cyc_cmd(cyc_cmd), .cyc_addr(cyc_addr),
        .cyc_dummy(cyc_dummy), .cyc_data(cyc_data)
    );

    task automatic run_plan(input string name,
                            input int ab, input int dc, input int db,
                            input bit want_addr, input bit want_dummy,
                            input bit want_data);
        int n, first_data, exp_lat, exp_tot;
        bit saw_addr, saw_dummy, saw_data;
        begin
            exp_lat = CMD_BITS + ab*8 + dc;
            exp_tot = exp_lat + db*8;

            @(negedge clk);
            addr_bytes   = 3'(ab);
            dummy_cycles = 6'(dc);
            data_bytes   = CNT_W'(db);
            start = 1'b1;
            @(negedge clk);
            start = 1'b0;

            n = 0; first_data = -1;
            saw_addr = 1'b0; saw_dummy = 1'b0; saw_data = 1'b0;
            while (busy) begin
                if (phase == P_ADDR)  saw_addr  = 1'b1;
                if (phase == P_DUMMY) saw_dummy = 1'b1;
                if (phase == P_DATA) begin
                    saw_data = 1'b1;
                    if (first_data < 0) first_data = n;
                end
                @(negedge clk);
                n++;
            end

            // 1. The reported latency matches the arithmetic.
            if (latency !== CNT_W'(exp_lat)) begin
                $display("  FAIL: %s latency reported %0d, expected %0d",
                         name, latency, exp_lat);
                errors++;
            end
            if (total_cycles !== CNT_W'(exp_tot)) begin
                $display("  FAIL: %s total reported %0d, expected %0d",
                         name, total_cycles, exp_tot);
                errors++;
            end

            // 2. The walk actually took that many cycles.
            if (n != exp_tot) begin
                $display("  FAIL: %s walked %0d cycles, planned %0d",
                         name, n, exp_tot);
                errors++;
            end

            // 3. The first data cycle lands exactly at the latency. This is
            //    what the number is FOR -- a latency that is right on paper
            //    and wrong in the walk is worse than no latency at all.
            if (want_data) begin
                if (first_data != exp_lat) begin
                    $display("  FAIL: %s first data cycle at %0d, latency says %0d",
                             name, first_data, exp_lat);
                    errors++;
                end
            end else if (saw_data) begin
                $display("  FAIL: %s entered the data phase with no data", name);
                errors++;
            end

            // 4. Empty phases are skipped, not entered for zero cycles.
            if (saw_addr != want_addr) begin
                $display("  FAIL: %s address phase %0s", name,
                         saw_addr ? "entered when empty" : "skipped when needed");
                errors++;
            end
            if (saw_dummy != want_dummy) begin
                $display("  FAIL: %s dummy phase %0s", name,
                         saw_dummy ? "entered when empty" : "skipped when needed");
                errors++;
            end

            // 5. CONSERVATION. The per-phase counts sum to the total.
            if ((cyc_cmd + cyc_addr + cyc_dummy + cyc_data) !== CNT_W'(exp_tot)) begin
                $display("  FAIL: %s phases sum to %0d, total is %0d",
                         name, cyc_cmd + cyc_addr + cyc_dummy + cyc_data, exp_tot);
                errors++;
            end
            if (cyc_cmd !== CNT_W'(CMD_BITS) || cyc_addr !== CNT_W'(ab*8) ||
                cyc_dummy !== CNT_W'(dc) || cyc_data !== CNT_W'(db*8)) begin
                $display("  FAIL: %s phase split cmd=%0d addr=%0d dummy=%0d data=%0d",
                         name, cyc_cmd, cyc_addr, cyc_dummy, cyc_data);
                errors++;
            end

            $display("  %-22s cmd=%0d addr=%0d dummy=%0d data=%0d  latency=%0d total=%0d",
                     name, cyc_cmd, cyc_addr, cyc_dummy, cyc_data,
                     latency, total_cycles);
        end
    endtask

    initial begin
        repeat (3) @(negedge clk);
        rst_n = 1'b1;
        @(negedge clk);

        // A flash fast read: 3 address bytes, 8 dummy cycles, 4 data bytes.
        run_plan("fast read 3B/8d/4B", 3, 8, 4, 1'b1, 1'b1, 1'b1);

        // The same part with a 6-cycle dummy latency. Six is not a byte,
        // and a driver that rounds it up to eight shifts every payload bit
        // by two positions.
        run_plan("6 dummy cycles", 3, 6, 4, 1'b1, 1'b1, 1'b1);

        // A dummy latency that is not even a whole number of nibbles.
        run_plan("5 dummy cycles", 3, 5, 2, 1'b1, 1'b1, 1'b1);

        // A four-byte address -- the part above 128 Mbit, where the third
        // address byte stops being enough.
        run_plan("4-byte address", 4, 8, 4, 1'b1, 1'b1, 1'b1);

        // A status read: no address, no dummy. The address and dummy phases
        // must be skipped entirely.
        run_plan("status read", 0, 0, 1, 1'b0, 1'b0, 1'b1);

        // A command with an address but no dummy and no data -- a sector
        // erase. Only the data phase is absent.
        run_plan("erase 3B/0d/0B", 3, 0, 0, 1'b1, 1'b0, 1'b0);

        // A bare command: write-enable. Command phase only.
        run_plan("write enable", 0, 0, 0, 1'b0, 1'b0, 1'b0);

        // A read with dummy but no address -- rarer, but real on parts
        // whose read command continues from an internal pointer.
        run_plan("no address, 4 dummy", 0, 4, 2, 1'b0, 1'b1, 1'b1);

        if (errors == 0)
            $display("PASS: latency equals command plus address bits plus dummy CYCLES, the first data cycle lands exactly there, empty phases are skipped rather than entered, and the per-phase cycle counts always sum to the total");
        else
            $display("FAIL: %0d error(s)", errors);
        $finish;
    end

endmodule

The testbench computes every expectation from the datasheet numbers itself rather than reading anything back from the planner, and then checks three separate things about each of eight profiles.

The reported latency matches the arithmetic. That is the easy one.

The first data cycle actually lands there. This is what the number is for — a latency that is right on paper and wrong in the walk is worse than no latency at all, because a consumer will trust it.

The per-phase counts sum to the total, and each equals its own phase's length. Conservation again, in the form Chapter 9.3 used for bytes and Chapter 9.2 used for frame time: when a design partitions something, the partition summing correctly is close to a complete specification of the accounting.

The eight profiles are chosen to cover the shapes that exist rather than a range of numbers: a fast read; the same part with six and then five dummy cycles; four-byte addressing; a status read with no address and no dummy; a sector erase with an address and no data; a bare write-enable; and a read with dummy but no address. Between them, every phase is present in some and absent in others.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_latency_plan.v — the same planner in Verilog-2001
// spi_latency_plan.v
//
// Chapter 10.5 -- address width, dummy cycles and burst length turned
// into a cycle-accurate plan, in Verilog-2001.
//
// Three datasheet numbers decide when the first payload bit appears: the
// command length (almost always 8 bits), the address width in BYTES, and
// the dummy latency in CYCLES -- which is NOT necessarily a multiple of
// eight. A part specifying "6 dummy cycles" does not want a dummy byte,
// and a driver that sends one shifts every payload byte by two bit
// positions.
//
// Empty phases are SKIPPED, not entered for zero cycles: a device with no
// address and no dummy goes straight from command to data, and a planner
// that enters a zero-length phase emits a state nothing downstream
// expects.

module spi_latency_plan #(
    parameter CNT_W    = 16,
    parameter CMD_BITS = 8
) (
    input  wire               clk,          // one tick per SCLK cycle
    input  wire               rst_n,

    input  wire               start,
    input  wire [2:0]         addr_bytes,   // 0..4
    input  wire [5:0]         dummy_cycles, // CYCLES, not bytes
    input  wire [CNT_W-1:0]   data_bytes,

    output reg  [1:0]         phase,        // 0=cmd 1=addr 2=dummy 3=data
    output reg                busy,
    output reg                done,

    output reg  [CNT_W-1:0]   latency,      // cycles before the first data bit
    output reg  [CNT_W-1:0]   total_cycles,

    // The plan's own accounting. These must sum to total_cycles: a phase
    // walker that loses or double-counts a cycle produces a latency that
    // is almost right.
    output reg  [CNT_W-1:0]   cyc_cmd,
    output reg  [CNT_W-1:0]   cyc_addr,
    output reg  [CNT_W-1:0]   cyc_dummy,
    output reg  [CNT_W-1:0]   cyc_data
);

    localparam [1:0] P_CMD   = 2'd0;
    localparam [1:0] P_ADDR  = 2'd1;
    localparam [1:0] P_DUMMY = 2'd2;
    localparam [1:0] P_DATA  = 2'd3;

    reg [CNT_W-1:0] addr_len, dummy_len, data_len;
    reg [CNT_W-1:0] rem;

    reg [1:0]       nxt_phase;
    reg [CNT_W-1:0] nxt_rem;
    reg             nxt_last;

    // Lengths of the phases as they are about to be loaded, in cycles.
    wire [CNT_W-1:0] addr_cyc  = {{(CNT_W-6){1'b0}}, addr_bytes, 3'b000};
    wire [CNT_W-1:0] dummy_cyc = {{(CNT_W-6){1'b0}}, dummy_cycles};
    wire [CNT_W-1:0] data_cyc  = data_bytes << 3;

    // Where to go when the current phase runs out, skipping every phase
    // the profile gives a length of zero.
    always @(*) begin
        nxt_phase = phase;
        nxt_rem   = {CNT_W{1'b0}};
        nxt_last  = 1'b0;
        case (phase)
            P_CMD: begin
                if      (addr_len  != 0) begin nxt_phase = P_ADDR;  nxt_rem = addr_len;  end
                else if (dummy_len != 0) begin nxt_phase = P_DUMMY; nxt_rem = dummy_len; end
                else if (data_len  != 0) begin nxt_phase = P_DATA;  nxt_rem = data_len;  end
                else                           nxt_last  = 1'b1;
            end
            P_ADDR: begin
                if      (dummy_len != 0) begin nxt_phase = P_DUMMY; nxt_rem = dummy_len; end
                else if (data_len  != 0) begin nxt_phase = P_DATA;  nxt_rem = data_len;  end
                else                           nxt_last  = 1'b1;
            end
            P_DUMMY: begin
                if      (data_len  != 0) begin nxt_phase = P_DATA;  nxt_rem = data_len;  end
                else                           nxt_last  = 1'b1;
            end
            default:                           nxt_last  = 1'b1;
        endcase
    end

    always @(posedge clk or negedge rst_n) begin
        if (!rst_n) begin
            phase        <= P_CMD;
            busy         <= 1'b0;
            done         <= 1'b0;
            addr_len     <= {CNT_W{1'b0}};
            dummy_len    <= {CNT_W{1'b0}};
            data_len     <= {CNT_W{1'b0}};
            rem          <= {CNT_W{1'b0}};
            latency      <= {CNT_W{1'b0}};
            total_cycles <= {CNT_W{1'b0}};
            cyc_cmd      <= {CNT_W{1'b0}};
            cyc_addr     <= {CNT_W{1'b0}};
            cyc_dummy    <= {CNT_W{1'b0}};
            cyc_data     <= {CNT_W{1'b0}};
        end else begin
            done <= 1'b0;

            if (start && !busy) begin
                // Every length is converted to CYCLES here, once, so
                // nothing downstream has to remember which datasheet
                // numbers were bytes and which were cycles.
                addr_len  <= addr_cyc;
                dummy_len <= dummy_cyc;
                data_len  <= data_cyc;

                latency      <= CMD_BITS + addr_cyc + dummy_cyc;
                total_cycles <= CMD_BITS + addr_cyc + dummy_cyc + data_cyc;

                phase     <= P_CMD;
                rem       <= CMD_BITS;
                busy      <= 1'b1;
                cyc_cmd   <= {CNT_W{1'b0}};
                cyc_addr  <= {CNT_W{1'b0}};
                cyc_dummy <= {CNT_W{1'b0}};
                cyc_data  <= {CNT_W{1'b0}};
            end else if (busy) begin
                // Account for the cycle being spent right now, before
                // deciding whether it was the phase's last.
                case (phase)
                    P_CMD:   cyc_cmd   <= cyc_cmd   + 1'b1;
                    P_ADDR:  cyc_addr  <= cyc_addr  + 1'b1;
                    P_DUMMY: cyc_dummy <= cyc_dummy + 1'b1;
                    default: cyc_data  <= cyc_data  + 1'b1;
                endcase

                if (rem == 1) begin
                    if (nxt_last) begin
                        busy <= 1'b0;
                        done <= 1'b1;
                    end else begin
                        phase <= nxt_phase;
                        rem   <= nxt_rem;
                    end
                end else begin
                    rem <= rem - 1'b1;
                end
            end
        end
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_latency_plan_tb.v — the same eight profiles in Verilog-2001
// spi_latency_plan_tb.v
//
// The same checks as the SystemVerilog testbench: the reported latency
// matches the arithmetic, the first data cycle lands exactly there, empty
// phases are skipped rather than entered, and the per-phase cycle counts
// always sum to the total.

`timescale 1ns/1ps

module spi_latency_plan_tb;

    parameter CNT_W    = 16;
    parameter CMD_BITS = 8;

    localparam [1:0] P_CMD   = 2'd0;
    localparam [1:0] P_ADDR  = 2'd1;
    localparam [1:0] P_DUMMY = 2'd2;
    localparam [1:0] P_DATA  = 2'd3;

    reg clk;
    reg rst_n;

    reg               start;
    reg  [2:0]        addr_bytes;
    reg  [5:0]        dummy_cycles;
    reg  [CNT_W-1:0]  data_bytes;

    wire [1:0]        phase;
    wire              busy, done;
    wire [CNT_W-1:0]  latency, total_cycles;
    wire [CNT_W-1:0]  cyc_cmd, cyc_addr, cyc_dummy, cyc_data;

    integer errors;

    initial begin
        clk = 1'b0; rst_n = 1'b0; start = 1'b0;
        addr_bytes = 3'd0; dummy_cycles = 6'd0; data_bytes = {CNT_W{1'b0}};
        errors = 0;
    end
    always #5 clk = ~clk;

    spi_latency_plan #(.CNT_W(CNT_W), .CMD_BITS(CMD_BITS)) dut (
        .clk(clk), .rst_n(rst_n), .start(start),
        .addr_bytes(addr_bytes), .dummy_cycles(dummy_cycles),
        .data_bytes(data_bytes),
        .phase(phase), .busy(busy), .done(done),
        .latency(latency), .total_cycles(total_cycles),
        .cyc_cmd(cyc_cmd), .cyc_addr(cyc_addr),
        .cyc_dummy(cyc_dummy), .cyc_data(cyc_data)
    );

    task run_plan;
        input [8*24:1] name;
        input integer  ab;
        input integer  dc;
        input integer  db;
        input          want_addr;
        input          want_dummy;
        input          want_data;
        integer n, first_data, exp_lat, exp_tot;
        reg saw_addr, saw_dummy, saw_data;
        begin
            exp_lat = CMD_BITS + ab*8 + dc;
            exp_tot = exp_lat + db*8;

            @(negedge clk);
            addr_bytes   = ab[2:0];
            dummy_cycles = dc[5:0];
            data_bytes   = db[CNT_W-1:0];
            start = 1'b1;
            @(negedge clk);
            start = 1'b0;

            n = 0; first_data = -1;
            saw_addr = 1'b0; saw_dummy = 1'b0; saw_data = 1'b0;
            while (busy) begin
                if (phase == P_ADDR)  saw_addr  = 1'b1;
                if (phase == P_DUMMY) saw_dummy = 1'b1;
                if (phase == P_DATA) begin
                    saw_data = 1'b1;
                    if (first_data < 0) first_data = n;
                end
                @(negedge clk);
                n = n + 1;
            end

            // 1. The reported latency matches the arithmetic.
            if (latency !== exp_lat[CNT_W-1:0]) begin
                $display("  FAIL: %0s latency reported %0d, expected %0d",
                         name, latency, exp_lat);
                errors = errors + 1;
            end
            if (total_cycles !== exp_tot[CNT_W-1:0]) begin
                $display("  FAIL: %0s total reported %0d, expected %0d",
                         name, total_cycles, exp_tot);
                errors = errors + 1;
            end

            // 2. The walk actually took that many cycles.
            if (n != exp_tot) begin
                $display("  FAIL: %0s walked %0d cycles, planned %0d",
                         name, n, exp_tot);
                errors = errors + 1;
            end

            // 3. The first data cycle lands exactly at the latency.
            if (want_data) begin
                if (first_data != exp_lat) begin
                    $display("  FAIL: %0s first data cycle at %0d, latency says %0d",
                             name, first_data, exp_lat);
                    errors = errors + 1;
                end
            end else if (saw_data) begin
                $display("  FAIL: %0s entered the data phase with no data", name);
                errors = errors + 1;
            end

            // 4. Empty phases are skipped, not entered for zero cycles.
            if (saw_addr !== want_addr) begin
                $display("  FAIL: %0s address phase handled wrongly", name);
                errors = errors + 1;
            end
            if (saw_dummy !== want_dummy) begin
                $display("  FAIL: %0s dummy phase handled wrongly", name);
                errors = errors + 1;
            end

            // 5. CONSERVATION. The per-phase counts sum to the total.
            if ((cyc_cmd + cyc_addr + cyc_dummy + cyc_data) !== exp_tot[CNT_W-1:0]) begin
                $display("  FAIL: %0s phases sum to %0d, total is %0d",
                         name, cyc_cmd + cyc_addr + cyc_dummy + cyc_data, exp_tot);
                errors = errors + 1;
            end
            if (cyc_cmd !== CMD_BITS || cyc_addr !== (ab*8) ||
                cyc_dummy !== dc || cyc_data !== (db*8)) begin
                $display("  FAIL: %0s phase split cmd=%0d addr=%0d dummy=%0d data=%0d",
                         name, cyc_cmd, cyc_addr, cyc_dummy, cyc_data);
                errors = errors + 1;
            end

            $display("  %0s cmd=%0d addr=%0d dummy=%0d data=%0d  latency=%0d total=%0d",
                     name, cyc_cmd, cyc_addr, cyc_dummy, cyc_data,
                     latency, total_cycles);
        end
    endtask

    initial begin
        repeat (3) @(negedge clk);
        rst_n = 1'b1;
        @(negedge clk);

        // A flash fast read: 3 address bytes, 8 dummy cycles, 4 data bytes.
        run_plan("fast read 3B/8d/4B ", 3, 8, 4, 1'b1, 1'b1, 1'b1);

        // The same part with a 6-cycle dummy latency. Six is not a byte.
        run_plan("6 dummy cycles     ", 3, 6, 4, 1'b1, 1'b1, 1'b1);

        // A dummy latency that is not even a whole number of nibbles.
        run_plan("5 dummy cycles     ", 3, 5, 2, 1'b1, 1'b1, 1'b1);

        // A four-byte address -- the part above 128 Mbit.
        run_plan("4-byte address     ", 4, 8, 4, 1'b1, 1'b1, 1'b1);

        // A status read: no address, no dummy.
        run_plan("status read        ", 0, 0, 1, 1'b0, 1'b0, 1'b1);

        // A sector erase: address but no dummy and no data.
        run_plan("erase 3B/0d/0B     ", 3, 0, 0, 1'b1, 1'b0, 1'b0);

        // A bare command: write-enable.
        run_plan("write enable       ", 0, 0, 0, 1'b0, 1'b0, 1'b0);

        // A read with dummy but no address.
        run_plan("no address, 4 dummy", 0, 4, 2, 1'b0, 1'b1, 1'b1);

        if (errors == 0)
            $display("PASS: latency equals command plus address bits plus dummy CYCLES, the first data cycle lands exactly there, empty phases are skipped rather than entered, and the per-phase cycle counts always sum to the total");
        else
            $display("FAIL: %0d error(s)", errors);
        $finish;
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_latency_plan.vhd — the same planner in VHDL
-- spi_latency_plan.vhd
--
-- Chapter 10.5 -- address width, dummy cycles and burst length turned
-- into a cycle-accurate plan, in VHDL.
--
-- Three datasheet numbers decide when the first payload bit appears: the
-- command length (almost always 8 bits), the address width in BYTES, and
-- the dummy latency in CYCLES -- which is NOT necessarily a multiple of
-- eight. A part specifying "6 dummy cycles" does not want a dummy byte,
-- and a driver that sends one shifts every payload byte by two bit
-- positions.
--
-- Empty phases are SKIPPED, not entered for zero cycles.

library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

entity spi_latency_plan is
    generic (
        CNT_W    : positive := 16;
        CMD_BITS : natural  := 8
    );
    port (
        clk          : in  std_logic;   -- one tick per SCLK cycle
        rst_n        : in  std_logic;

        start        : in  std_logic;
        addr_bytes   : in  unsigned(2 downto 0);   -- 0..4
        dummy_cycles : in  unsigned(5 downto 0);   -- CYCLES, not bytes
        data_bytes   : in  unsigned(CNT_W - 1 downto 0);

        phase        : out unsigned(1 downto 0);   -- 0=cmd 1=addr 2=dummy 3=data
        busy         : out std_logic;
        done         : out std_logic;

        latency      : out unsigned(CNT_W - 1 downto 0);
        total_cycles : out unsigned(CNT_W - 1 downto 0);

        -- The plan's own accounting. These must sum to total_cycles.
        cyc_cmd      : out unsigned(CNT_W - 1 downto 0);
        cyc_addr     : out unsigned(CNT_W - 1 downto 0);
        cyc_dummy    : out unsigned(CNT_W - 1 downto 0);
        cyc_data     : out unsigned(CNT_W - 1 downto 0)
    );
end entity;

architecture rtl of spi_latency_plan is

    constant P_CMD   : unsigned(1 downto 0) := "00";
    constant P_ADDR  : unsigned(1 downto 0) := "01";
    constant P_DUMMY : unsigned(1 downto 0) := "10";
    constant P_DATA  : unsigned(1 downto 0) := "11";

    signal phase_r : unsigned(1 downto 0) := P_CMD;
    signal busy_r  : std_logic := '0';
    signal done_r  : std_logic := '0';

    signal addr_len  : unsigned(CNT_W - 1 downto 0) := (others => '0');
    signal dummy_len : unsigned(CNT_W - 1 downto 0) := (others => '0');
    signal data_len  : unsigned(CNT_W - 1 downto 0) := (others => '0');
    signal rem_cnt   : unsigned(CNT_W - 1 downto 0) := (others => '0');

    signal lat_r   : unsigned(CNT_W - 1 downto 0) := (others => '0');
    signal tot_r   : unsigned(CNT_W - 1 downto 0) := (others => '0');
    signal c_cmd   : unsigned(CNT_W - 1 downto 0) := (others => '0');
    signal c_addr  : unsigned(CNT_W - 1 downto 0) := (others => '0');
    signal c_dummy : unsigned(CNT_W - 1 downto 0) := (others => '0');
    signal c_data  : unsigned(CNT_W - 1 downto 0) := (others => '0');

    signal nxt_phase : unsigned(1 downto 0) := P_CMD;
    signal nxt_rem   : unsigned(CNT_W - 1 downto 0) := (others => '0');
    signal nxt_last  : std_logic := '0';

    -- Lengths of the phases as they are about to be loaded, in cycles.
    signal addr_cyc  : unsigned(CNT_W - 1 downto 0);
    signal dummy_cyc : unsigned(CNT_W - 1 downto 0);
    signal data_cyc  : unsigned(CNT_W - 1 downto 0);

begin

    addr_cyc  <= resize(addr_bytes, CNT_W - 3) & "000";
    dummy_cyc <= resize(dummy_cycles, CNT_W);
    data_cyc  <= shift_left(data_bytes, 3);

    -- Where to go when the current phase runs out, skipping every phase
    -- the profile gives a length of zero.
    nxt : process (phase_r, addr_len, dummy_len, data_len)
    begin
        nxt_phase <= phase_r;
        nxt_rem   <= (others => '0');
        nxt_last  <= '0';
        case phase_r is
            when P_CMD =>
                if addr_len /= 0 then
                    nxt_phase <= P_ADDR;  nxt_rem <= addr_len;
                elsif dummy_len /= 0 then
                    nxt_phase <= P_DUMMY; nxt_rem <= dummy_len;
                elsif data_len /= 0 then
                    nxt_phase <= P_DATA;  nxt_rem <= data_len;
                else
                    nxt_last <= '1';
                end if;
            when P_ADDR =>
                if dummy_len /= 0 then
                    nxt_phase <= P_DUMMY; nxt_rem <= dummy_len;
                elsif data_len /= 0 then
                    nxt_phase <= P_DATA;  nxt_rem <= data_len;
                else
                    nxt_last <= '1';
                end if;
            when P_DUMMY =>
                if data_len /= 0 then
                    nxt_phase <= P_DATA;  nxt_rem <= data_len;
                else
                    nxt_last <= '1';
                end if;
            when others =>
                nxt_last <= '1';
        end case;
    end process;

    walk : process (clk, rst_n)
    begin
        if rst_n = '0' then
            phase_r   <= P_CMD;
            busy_r    <= '0';
            done_r    <= '0';
            addr_len  <= (others => '0');
            dummy_len <= (others => '0');
            data_len  <= (others => '0');
            rem_cnt   <= (others => '0');
            lat_r     <= (others => '0');
            tot_r     <= (others => '0');
            c_cmd     <= (others => '0');
            c_addr    <= (others => '0');
            c_dummy   <= (others => '0');
            c_data    <= (others => '0');
        elsif rising_edge(clk) then
            done_r <= '0';

            if start = '1' and busy_r = '0' then
                -- Every length is converted to CYCLES here, once, so
                -- nothing downstream has to remember which datasheet
                -- numbers were bytes and which were cycles.
                addr_len  <= addr_cyc;
                dummy_len <= dummy_cyc;
                data_len  <= data_cyc;

                lat_r <= to_unsigned(CMD_BITS, CNT_W) + addr_cyc + dummy_cyc;
                tot_r <= to_unsigned(CMD_BITS, CNT_W) + addr_cyc + dummy_cyc
                         + data_cyc;

                phase_r <= P_CMD;
                rem_cnt <= to_unsigned(CMD_BITS, CNT_W);
                busy_r  <= '1';
                c_cmd   <= (others => '0');
                c_addr  <= (others => '0');
                c_dummy <= (others => '0');
                c_data  <= (others => '0');
            elsif busy_r = '1' then
                -- Account for the cycle being spent right now, before
                -- deciding whether it was the phase's last.
                case phase_r is
                    when P_CMD   => c_cmd   <= c_cmd   + 1;
                    when P_ADDR  => c_addr  <= c_addr  + 1;
                    when P_DUMMY => c_dummy <= c_dummy + 1;
                    when others  => c_data  <= c_data  + 1;
                end case;

                if rem_cnt = 1 then
                    if nxt_last = '1' then
                        busy_r <= '0';
                        done_r <= '1';
                    else
                        phase_r <= nxt_phase;
                        rem_cnt <= nxt_rem;
                    end if;
                else
                    rem_cnt <= rem_cnt - 1;
                end if;
            end if;
        end if;
    end process;

    phase        <= phase_r;
    busy         <= busy_r;
    done         <= done_r;
    latency      <= lat_r;
    total_cycles <= tot_r;
    cyc_cmd      <= c_cmd;
    cyc_addr     <= c_addr;
    cyc_dummy    <= c_dummy;
    cyc_data     <= c_data;

end architecture;
Azvya Education Pvt. Ltd.VLSI Mentor
spi_latency_plan_tb.vhd — the same eight profiles in VHDL
-- spi_latency_plan_tb.vhd
--
-- The same checks as the SystemVerilog and Verilog testbenches: the
-- reported latency matches the arithmetic, the first data cycle lands
-- exactly there, empty phases are skipped rather than entered, and the
-- per-phase cycle counts always sum to the total.

library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

entity spi_latency_plan_tb is
end entity;

architecture sim of spi_latency_plan_tb is

    constant CNT_W    : positive := 16;
    constant CMD_BITS : natural  := 8;

    constant P_ADDR  : unsigned(1 downto 0) := "01";
    constant P_DUMMY : unsigned(1 downto 0) := "10";
    constant P_DATA  : unsigned(1 downto 0) := "11";

    signal clk   : std_logic := '0';
    signal rst_n : std_logic := '0';
    signal halt  : boolean   := false;

    signal start        : std_logic := '0';
    signal addr_bytes   : unsigned(2 downto 0) := (others => '0');
    signal dummy_cycles : unsigned(5 downto 0) := (others => '0');
    signal data_bytes   : unsigned(CNT_W - 1 downto 0) := (others => '0');

    signal phase        : unsigned(1 downto 0);
    signal busy, done   : std_logic;
    signal latency      : unsigned(CNT_W - 1 downto 0);
    signal total_cycles : unsigned(CNT_W - 1 downto 0);
    signal cyc_cmd      : unsigned(CNT_W - 1 downto 0);
    signal cyc_addr     : unsigned(CNT_W - 1 downto 0);
    signal cyc_dummy    : unsigned(CNT_W - 1 downto 0);
    signal cyc_data     : unsigned(CNT_W - 1 downto 0);

    signal errors : natural := 0;

begin

    clk <= not clk after 5 ns when not halt else '0';

    dut : entity work.spi_latency_plan
        generic map (CNT_W => CNT_W, CMD_BITS => CMD_BITS)
        port map (
            clk => clk, rst_n => rst_n, start => start,
            addr_bytes => addr_bytes, dummy_cycles => dummy_cycles,
            data_bytes => data_bytes,
            phase => phase, busy => busy, done => done,
            latency => latency, total_cycles => total_cycles,
            cyc_cmd => cyc_cmd, cyc_addr => cyc_addr,
            cyc_dummy => cyc_dummy, cyc_data => cyc_data
        );

    stim : process
        variable errs : natural := 0;

        procedure run_plan(name : string; ab : natural; dc : natural;
                           db : natural; want_addr : boolean;
                           want_dummy : boolean; want_data : boolean) is
            variable n          : natural;
            variable first_data : integer;
            variable exp_lat    : natural;
            variable exp_tot    : natural;
            variable saw_addr   : boolean;
            variable saw_dummy  : boolean;
            variable saw_data   : boolean;
        begin
            exp_lat := CMD_BITS + ab * 8 + dc;
            exp_tot := exp_lat + db * 8;

            wait until falling_edge(clk);
            addr_bytes   <= to_unsigned(ab, 3);
            dummy_cycles <= to_unsigned(dc, 6);
            data_bytes   <= to_unsigned(db, CNT_W);
            start        <= '1';
            wait until falling_edge(clk);
            start        <= '0';

            n := 0; first_data := -1;
            saw_addr := false; saw_dummy := false; saw_data := false;
            while busy = '1' loop
                if phase = P_ADDR  then saw_addr  := true; end if;
                if phase = P_DUMMY then saw_dummy := true; end if;
                if phase = P_DATA then
                    saw_data := true;
                    if first_data < 0 then first_data := n; end if;
                end if;
                wait until falling_edge(clk);
                n := n + 1;
            end loop;

            -- 1. The reported latency matches the arithmetic.
            if to_integer(latency) /= exp_lat then
                report "  FAIL: " & name & " latency reported " &
                       integer'image(to_integer(latency)) & ", expected " &
                       integer'image(exp_lat);
                errs := errs + 1;
            end if;
            if to_integer(total_cycles) /= exp_tot then
                report "  FAIL: " & name & " total reported " &
                       integer'image(to_integer(total_cycles)) &
                       ", expected " & integer'image(exp_tot);
                errs := errs + 1;
            end if;

            -- 2. The walk actually took that many cycles.
            if n /= exp_tot then
                report "  FAIL: " & name & " walked " & integer'image(n) &
                       " cycles, planned " & integer'image(exp_tot);
                errs := errs + 1;
            end if;

            -- 3. The first data cycle lands exactly at the latency.
            if want_data then
                if first_data /= exp_lat then
                    report "  FAIL: " & name & " first data cycle at " &
                           integer'image(first_data) & ", latency says " &
                           integer'image(exp_lat);
                    errs := errs + 1;
                end if;
            elsif saw_data then
                report "  FAIL: " & name & " entered the data phase with no data";
                errs := errs + 1;
            end if;

            -- 4. Empty phases are skipped, not entered for zero cycles.
            if saw_addr /= want_addr then
                report "  FAIL: " & name & " address phase handled wrongly";
                errs := errs + 1;
            end if;
            if saw_dummy /= want_dummy then
                report "  FAIL: " & name & " dummy phase handled wrongly";
                errs := errs + 1;
            end if;

            -- 5. CONSERVATION. The per-phase counts sum to the total.
            if to_integer(cyc_cmd) + to_integer(cyc_addr) +
               to_integer(cyc_dummy) + to_integer(cyc_data) /= exp_tot then
                report "  FAIL: " & name & " phases do not sum to the total";
                errs := errs + 1;
            end if;
            if to_integer(cyc_cmd) /= CMD_BITS or
               to_integer(cyc_addr) /= ab * 8 or
               to_integer(cyc_dummy) /= dc or
               to_integer(cyc_data) /= db * 8 then
                report "  FAIL: " & name & " phase split is wrong";
                errs := errs + 1;
            end if;

            report "  " & name & " cmd=" &
                   integer'image(to_integer(cyc_cmd)) & " addr=" &
                   integer'image(to_integer(cyc_addr)) & " dummy=" &
                   integer'image(to_integer(cyc_dummy)) & " data=" &
                   integer'image(to_integer(cyc_data)) & "  latency=" &
                   integer'image(to_integer(latency)) & " total=" &
                   integer'image(to_integer(total_cycles));
        end procedure;
    begin
        for i in 0 to 2 loop
            wait until falling_edge(clk);
        end loop;
        rst_n <= '1';
        wait until falling_edge(clk);

        -- A flash fast read: 3 address bytes, 8 dummy cycles, 4 data bytes.
        run_plan("fast read 3B/8d/4B ", 3, 8, 4, true,  true,  true);
        -- The same part with a 6-cycle dummy latency. Six is not a byte.
        run_plan("6 dummy cycles     ", 3, 6, 4, true,  true,  true);
        -- A dummy latency that is not even a whole number of nibbles.
        run_plan("5 dummy cycles     ", 3, 5, 2, true,  true,  true);
        -- A four-byte address -- the part above 128 Mbit.
        run_plan("4-byte address     ", 4, 8, 4, true,  true,  true);
        -- A status read: no address, no dummy.
        run_plan("status read        ", 0, 0, 1, false, false, true);
        -- A sector erase: address but no dummy and no data.
        run_plan("erase 3B/0d/0B     ", 3, 0, 0, true,  false, false);
        -- A bare command: write-enable.
        run_plan("write enable       ", 0, 0, 0, false, false, false);
        -- A read with dummy but no address.
        run_plan("no address, 4 dummy", 0, 4, 2, false, true,  true);

        errors <= errs;
        if errs = 0 then
            report "PASS: latency equals command plus address bits plus dummy CYCLES, the first data cycle lands exactly there, empty phases are skipped rather than entered, and the per-phase cycle counts always sum to the total";
        else
            report "FAIL: " & integer'image(errs) & " error(s)" severity error;
        end if;
        halt <= true;
        wait;
    end process;

end architecture;

Parity

All three implement the same planner: identical ports and generics, a single conversion of every length to cycles at start, empty phases skipped rather than entered, a published latency and total, and per-phase accounting that sums to the total. All three testbenches run the same eight profiles and report identical numbers — a latency of 40 for the standard fast read, 38 with six dummy cycles, 37 with five, 48 with four address bytes, and 8 for a bare command.

7. Why a Verification Engineer Cares

Azvya Education Pvt. Ltd.VLSI Mentor
spi_latency_plan.sva — the accounting properties
   // 1. CONSERVATION. The per-phase cycle counts sum to the total. A
   //    walker that loses or double-counts one cycle yields a latency
   //    that is ALMOST right, which no spot check reliably catches.
   a_conservation : assert property (
       @(posedge clk) disable iff (!rst_n)
           (done) |-> ((cyc_cmd + cyc_addr + cyc_dummy + cyc_data)
                        == total_cycles))
       else $error("the phase cycle counts do not sum to the total");

   // 2. THE LATENCY IS HONOURED. The first cycle spent in the data phase
   //    is the cycle the published latency named. This is what the number
   //    exists for, and a consumer will trust it.
   a_latency_honoured : assert property (
       @(posedge clk) disable iff (!rst_n)
           ($rose(phase == P_DATA)) |-> (elapsed == latency))
       else $error("the data phase did not begin at the published latency");

   // 3. NO EMPTY PHASES. A phase with zero length is skipped, never
   //    entered -- a zero-cycle state is one nothing downstream expects.
   a_no_empty_phase : assert property (
       @(posedge clk) disable iff (!rst_n)
           ((phase == P_DUMMY) |-> (dummy_len != 0)) and
           ((phase == P_ADDR)  |-> (addr_len  != 0)))
       else $error("a zero-length phase was entered");

   // 4. TERMINATION. Every started plan completes. A phase walker's
   //    natural failure is a stall rather than a wrong answer, and only a
   //    liveness property catches it.
   a_terminates : assert property (
       @(posedge clk) disable iff (!rst_n)
           (start && !busy) |-> ##[1:$] done)
       else $error("the plan never completed");

   // 5. DUMMY IS NOT ROUNDED. The cycles spent in the dummy phase equal
   //    the number the profile gave -- exactly, including odd values.
   a_dummy_exact : assert property (
       @(posedge clk) disable iff (!rst_n)
           (done) |-> (cyc_dummy == $past(dummy_cycles_at_start)))
       else $error("the dummy phase was not exactly the specified length");

Property 5 exists because rounding is such a natural thing to do. A designer who thinks in bytes will write the dummy phase as a byte count without noticing, and every profile whose dummy count happens to be a multiple of eight will pass. The assertion makes the odd values part of the specification rather than part of the test data.

Property 4 is the liveness property, and it is the one that catches the specific failure mode of a phase walker: not a wrong answer but a stall, produced by a next-phase decision that skips into nothing.

Coverage must cross presence with length, because the absent cases are the interesting ones:

Azvya Education Pvt. Ltd.VLSI Mentor
spi_phase_plan_cg.sv — which phases exist, not just how long they are
   covergroup spi_phase_plan_cg @(posedge clk iff start);
       cp_abytes : coverpoint addr_bytes {
           bins none  = {0};       // status reads -- phase must be SKIPPED
           bins one   = {1};
           bins three = {3};       // the common case
           bins four  = {4};       // 4-byte mode
       }

       // The values that are not multiples of eight are the point of this
       // whole chapter, and a suite using only 0 and 8 never tests them.
       cp_dummy : coverpoint dummy_cycles {
           bins none        = {0};
           bins odd_small   = {[1:7]};    // 4, 5, 6 -- real values
           bins one_byte    = {8};
           bins odd_large   = {[9:15]};   // 10, 12, 14 -- also real
           bins large       = {[16:32]};
       }

       cp_data : coverpoint data_bytes {
           bins none  = {0};       // erase, write-enable
           bins one   = {1};
           bins many  = {[2:$]};
       }

       // The shape of the transfer, which is what the walker actually
       // implements. Eight combinations, and a suite that only ever sends
       // full four-phase transfers has tested one of them.
       x_shape : cross cp_abytes, cp_dummy, cp_data;
   endgroup

cp_dummy's odd_small and odd_large bins are the ones to insist on. A suite whose dummy counts are all 0 or 8 will pass a planner that quietly works in bytes.

8. Why an FPGA or ASIC Engineer Cares

Convert to cycles once, at the boundary. The byte-versus-cycle confusion is a units error, and units errors are prevented the same way everywhere: convert on entry and keep one unit internally. A design that carries "dummy" around as a byte count in one module and a cycle count in another will eventually mix them.

Make the dummy count a profile field with odd values reachable. A field wide enough only for multiples of eight is a design that has decided the datasheet is wrong.

Check the address against the address width in hardware. The device cannot detect an over-wide address — it simply consumes the bytes it receives — so a 4-byte address sent to a part in 3-byte mode shifts the entire rest of the transfer. Three gates and a comparator prevent it.

Track 3-byte/4-byte mode as state, and re-establish it after reset. It is persistent, it survives a warm reset, and previous software can leave it set. A controller that issues an explicit mode command at initialisation rather than assuming the default removes a bring-up failure that is otherwise very hard to see.

Publish the latency, do not recompute it. Every consumer that needs to know when data begins should read the same number from the same place. Two modules computing it independently is how a design comes to disagree with itself about a value that is right in both.

9. Failure Signature — A Flash That Reads Correctly Below 16 MB

Symptom. A 256 Mbit flash works perfectly. Every read verifies, the file system is stable, and the product ships. Some months later a build grows past a certain size and reads from the upper half of the device return data that belongs to the lower half — correct-looking data from entirely the wrong place. Below the halfway mark everything is still perfect.

What "correct data from the wrong place" establishes. The link is fine. The mode, the rate, the dummy count and the command encoding are all right, because the data is intact — only its address is wrong. So this is an addressing fault, not a timing or protocol fault.

Plausible mechanisms.

  • The device is in 3-byte mode and the controller sends three address bytes. A 256 Mbit part is 32 MB and needs 25 address bits; three bytes give 24, so the top bit is simply absent and every address above 16 MB aliases to the same address 16 MB lower.
  • The controller sends four bytes while the device is in 3-byte mode, which would shift the whole transfer and corrupt the data — so this does not fit intact data.
  • A driver truncating the address to a 24-bit type somewhere in its call chain, which produces the identical aliasing.
  • A file system or wear-levelling layer with its own 24-bit assumption.
  • The part genuinely being 128 Mbit and mislabelled, which fits the symptom exactly and is worth eliminating with an ID read.

The discriminating observation. Aliasing has a precise signature: address X and address X + 16 MB return identical data. Test that directly. Write a known pattern at 0x000100, read 0x1000100, and if the pattern appears there the top address bit is being lost.

Then determine where it is lost. Read the device's configuration register to see whether it reports 3-byte or 4-byte mode. If the device says 4-byte and the aliasing persists, the truncation is in the controller or the driver; if the device says 3-byte, the device never received the fourth byte.

The fix. Enter 4-byte mode explicitly at initialisation — or use the 4-byte opcodes, which need no mode at all and are the more robust choice precisely because they carry no persistent state. Then widen whatever type truncated the address.

Why this ships. Because the entire lower 16 MB works perfectly, and that is all anybody tested. Address-space bugs are invisible until the address space is used, and a product whose image fits in the bottom half will pass every test for years. The lesson is to test the top of any address space deliberately, at bring-up, when the fix is a configuration line rather than a field update.

10. Common Misconceptions

11. Reason It Through

Work this before reading the answer.

A driver reads 256 bytes from a flash at address 0x7FFF80 using 0x0B with three address bytes and eight dummy cycles, at 20 MHz. The first 128 bytes verify. The last 128 are wrong — and, oddly, they match the contents of address 0x000000 onwards.

The part is 8 MB. What happened, and would a faster or slower clock change anything?

Start with the arithmetic. The part is 8 MB, so its top address is 0x7FFFFF. The read starts at 0x7FFF80, which is 128 bytes below the top.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   0x7FFF80 + 128 = 0x800000

So the read runs off the end of the device exactly halfway through.

What does the device do? Per §4, a read burst that passes the top of the device either wraps to address 0 or returns undefined data — and this one wraps. That is why the last 128 bytes match address 0x000000 onwards. The device is behaving exactly as its datasheet says.

So nothing is broken. The link is correct, the dummy count is correct, the address is correct, and the device did what it documents. The driver issued a read that crosses the end of the address space, and nothing in the protocol prevents that.

Would changing the clock help? No — and this is the part worth being sure about. The failure is an addressing behaviour, not a timing one. Per the warning in §3, rate dependence is what distinguishes a timing fault from a logical one: this one would reproduce identically at 1 MHz and at 50 MHz, because the device wraps at the same address either way.

What is the fix? Bound the read at the device's top address, exactly as Chapter 9.3's splitter bounds a burst at a page boundary — the mechanism is the same three-way minimum, with the distance to the end of the device as one more limit. A read that would cross the end is split into one read that stops at the top and, if the caller really wanted the wrap, a second read from zero.

And the deeper reading. The end-of-device boundary behaves like a page boundary for reads, even though pages are usually discussed only for writes. Any burst has several limits — the page, the device's end, the buffer, the deadline — and a splitter that knows about only one of them works until the first transfer that meets another. When you build the splitter of Chapter 9.3, the end of the device belongs in its minimum.

12. Understanding Check

13. Summary

Once a device has more addresses than a command byte has spare bits, the address becomes its own phase, and three numbers come with it.

Address width is stated in bytes and creates the ceiling: three bytes reach exactly 16 MB, which is why 128 Mbit is a category boundary and why larger parts carry a persistent 3-byte/4-byte mode that survives a warm reset and that previous software can leave set. An address that does not fit is never rejected — the device consumes the bytes it receives and everything after shifts.

Byte order is most-significant first, which is the opposite of how a little-endian host stores it — so a copied address reaches a valid but wildly wrong location.

Dummy latency is counted in CYCLES, and 4, 5, 6 and 10 are ordinary values. Sending a byte instead shifts the payload, and the result is plausible bytes spliced from adjacent real ones rather than obvious garbage. The signature matches a round-trip violation exactly, and rate dependence is what separates them.

Burst behaviour differs by direction: reads usually run freely and wrap at the end of the device; writes wrap inside the page, overwriting what was just written.

In hardware, convert every length to cycles once, at the boundary, skip empty phases rather than entering them, publish the latency so no two modules compute it separately, and assert that the per-phase counts sum to the total — because a walker that loses one cycle produces a latency that is almost right.

14. What Comes Next

Every number the route asked for is now in hand: the mode, the rate, the timing, the command encoding, the address width, the dummy count and the burst rules.

Chapter 10.6 — From Datasheet to Transaction Specification assembles them. It works a complete device end to end, turns the result into the exact element sequence a controller must issue, and builds the specification engine that produces it from the profile — so that supporting a new part becomes a table entry rather than a driver. It is also where the full UVM environment for a datasheet-derived device comes together, in all three HDLs.

Continue learning