Skip to content
VLSI Mentor

SPI · Module 11

Normal Read, Fast Read, and Dummy Cycles

Why a flash offers two read commands returning the same data, why the faster one must buy array access time in dummy cycles, why fast read is not always faster, and the selector that chooses by transfer time rather than clock rate.

Chapter 11.1 noted that reads are the only unconstrained flash operation — any address, any length. They are also the only operation with two commands that do exactly the same thing.

0x03 and 0x0B both read data. 0x0B is called Fast Read and costs eight dummy cycles that 0x03 does not. Why does the slower command still exist, and when is it actually the faster choice?

The answer is not "backwards compatibility". The two commands trade clock rate against cycle count, and which wins depends on how much data you are moving.

1. Why Two Commands Exist

The plain read, 0x03, is the simplest possible transaction: opcode, address, data. No latency, nothing to configure, and every 25-series device implements it identically. It is the command a boot ROM uses before it knows anything about the part.

Its limitation is rate. The device must fetch the first byte from the array and present it on MISO within half an SCLK period of the last address bit — and the internal array access takes a fixed time in nanoseconds regardless of the clock. So 0x03 carries a maximum frequency substantially below the device's headline number, commonly 50 MHz on a part rated 104.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   0x03  Read        no dummy cycles      max ~50 MHz
   0x0B  Fast Read   8 dummy cycles       max ~104 MHz

The fast read solves it in the only way available: by buying time with clock cycles. Eight dummy cycles after the address give the array eight SCLK periods to produce the first byte instead of half of one. At 104 MHz that is about 77 ns of internal access time, which is comfortable — and the cost is eight cycles of nothing.

2. Where the Time Goes

0x03 and 0x0B, byte time by byte time

8 cycles
Two byte-time lanes for the same read. The plain read sends a command byte, three address bytes, then four data bytes. The fast read sends a command byte, three address bytes, one dummy byte, then three data bytes in the same span.0x03 payload starts0x03 payload startsread 0x03CMDA2A1A0D0D1D2D3fast 0x0BCMDA2A1A0dumD0D1D2t0t1t2t3t4t5t6t7
Figure 1 — the same read, both commands. Both spend one byte time on the opcode and three on the address. The fast read then spends one more byte time doing nothing, so its payload begins a byte later — but it may be clocked at twice the rate, which more than repays the delay on a long transfer.

Per byte time the plain read is strictly ahead — it is one byte further into the payload throughout. But the diagram deliberately hides the variable that decides the outcome: the byte times are not the same width. If the fast read is clocked twice as fast, each of its columns is half as wide, and it finishes a long transfer well ahead despite starting the payload later.

So comparing the commands means comparing times, not cycles.

3. The Arithmetic

Both commands take the same shape, differing only in the dummy phase and the divisor each permits:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   sclk_cycles = 8 (opcode) + 8 × addr_bytes + dummy + 8 × length
   time        = sclk_cycles × divisor

Take an 80 MHz system clock, a part whose plain read needs divisor 8 or slower and whose fast read runs at divisor 2, and three address bytes.

At divisor 8 — a rate both commands can meet:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   0x03:  (8 + 24 + 0 + 8 × 64) × 8  =  (32 + 512) × 8  =  4352 → wait

Careful — at divisor 8 both commands run at divisor 8, so:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   0x03:  (8 + 24 + 0  + 512) × 8  =  544 × 8  =  4352 cycles
   0x0B:  (8 + 24 + 8  + 512) × 8  =  552 × 8  =  4416 cycles

The plain read wins. Same rate, eight fewer cycles. This is the case the usual advice gets wrong: when the requested rate is one the plain read can meet, the fast read is strictly worse — it pays for latency it does not need.

At divisor 2 — a rate only the fast read can meet:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   0x03:  forced to divisor 8   →  544 × 8  =  4352 cycles
   0x0B:  runs at divisor 2     →  552 × 2  =  1104 cycles

The fast read wins by a factor of four. Eight extra cycles against a quarter of the period is not a close contest.

4. The Crossover

The interesting region is a divisor just below the plain read's ceiling, where the plain read must slow down to 8 while the fast read runs at 7.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   at divisor 7, three address bytes:

   length    0x03 (div 8)     0x0B (div 7)     winner
      0        256 cycles        280            plain
      1        320               336            plain
      2        384               392            plain
      3        448               448            tie
      4        512               504            FAST
      8        768               728            fast
     64       4352              3976            fast

The crossover is at four bytes. Below it the plain read's smaller cycle count wins even though it runs slower; above it the fast read's higher rate takes over. The §6 testbench finds that length from the arithmetic and then checks the selector on both sides of it.

Two things make this worth knowing rather than merely cute.

The crossover exists at all. It means "use fast read" is not a rule but a default whose exceptions are short transfers near the plain read's ceiling — which is exactly what a status-register-adjacent read or a single-byte probe looks like.

The margin is narrow there. At divisor 7 and four bytes the two differ by eight cycles in five hundred. Nothing is lost by choosing either, which is itself the useful conclusion: near the crossover the choice does not matter, so choose the simpler command. That is Chapter 9.3's lesson again — when two candidates sit in a flat region, decide on something other than the metric.

5. The Read Transaction, End to End

A sequence diagram of a fast read. The master asserts chip select, sends the fast read opcode, sends three address bytes most significant first, clocks eight dummy cycles during which nothing meaningful is exchanged, then receives payload bytes as the device advances its internal pointer, before releasing chip select.Fast read, phase by phasemasterflashCS asserts0x0B — fast readaddress, MSB bytefirst8 dummy cycles —array access timefirst payload byte… pointer advancesfreelyCS releases
Figure 2 — a fast read as a sequence. The dummy phase is clocked time in which neither side sends anything meaningful; the master must clock it and must not treat what appears on MISO during it as data.

Three properties of that sequence matter for the rest of the module.

The read pointer advances without limit. Unlike a program, a read burst is not page-bounded — it runs to the end of the device and then, on most parts, wraps to zero. Chapter 10.5's worked exercise turned on exactly that.

The dummy phase must be clocked, not waited out. Deasserting CS and pausing does not work: the device needs SCLK edges, not elapsed time. A master that inserts a delay instead of clocks gets no data.

Whatever MISO shows during the dummy phase is not data. It may be the previous byte, high impedance, or the first bits arriving early. Treating it as payload is the off-by-one that shifts the whole transfer, and Chapter 10.5 showed the signature.

6. Building the Read Selector — Three HDLs

The circuit

Circuit. A two-way comparison of estimated transfer times.

State. The chosen opcode, dummy count, effective divisor and both estimates, registered on a select pulse.

Datapath. Each command's divisor is clamped to the fastest it allows — larger divisor means slower clock, so the floor is a maximum on the divisor, the same inversion Chapter 9.4's rate limiter turns on. Then each command's cycle count is multiplied by its divisor and the smaller product wins.

Control. One pulse in, a decision out. No sequencing.

Clock and reset. System clock; asynchronous active-low reset. Reset selects the plain read at the slowest divisor — the command every device implements identically, with no latency to get wrong, which is precisely what a boot ROM needs before it has read an ID.

Enables. too_fast reports that the request exceeded even the fast read, so the clamp was applied rather than the request honoured.

Timing. Registered, so the decision is stable one cycle after select.

Synthesis. Two multipliers, two comparators and a handful of registers. The multipliers are the cost, and they are the honest answer: sharing one and taking two cycles is the right trade if area matters, since a decision made once per transfer has cycles to spare.

Limitations. Two commands. A part offering dual and quad reads has four or five candidates, which turns the two-way comparison into a minimum over a small table — the same shape, more entries.

The tie-break is deliberate. On equal estimates the plain read wins, because it has fewer cycles and therefore less exposure to a latency mismatch. Preferring the simpler command when the metric cannot distinguish them is an engineering choice, not an arbitrary one.

Azvya Education Pvt. Ltd.VLSI Mentor
flash_read_sel.sv — choose by time, not by rate
// flash_read_sel.sv
//
// Chapter 11.2 -- choosing between the plain read and the fast read.
//
// A serial flash offers two read commands that return the same data:
//
//   0x03  Read        no dummy cycles, but a LOW maximum clock rate
//   0x0B  Fast Read   a high maximum rate, paid for with dummy cycles
//
// The usual summary -- "use fast read, it is faster" -- is wrong often
// enough to matter, because the two commands trade CLOCK RATE against
// CYCLE COUNT and which wins depends on the transfer LENGTH.
//
// So this block does not choose by rate. It computes the time each command
// would take, in system-clock cycles, and picks the smaller:
//
//   time = sclk_cycles x divisor
//   sclk_cycles = 8 (opcode) + 8 x addr_bytes + dummy + 8 x length
//
// At a requested divisor the plain read can meet, it wins outright -- same
// rate, eight fewer cycles. Below its ceiling the fast read's higher rate
// usually wins. But just below the plain read's ceiling the two CROSS: the
// plain read, forced to a slower divisor, still beats the fast read for
// short transfers and loses for long ones. The testbench finds that
// crossover and checks it.
//
// Both estimates are published, not just the winner, so software can see
// what the choice cost.

module flash_read_sel #(
    parameter int DIV_W        = 8,
    parameter int LEN_W        = 16,
    parameter int TIME_W       = 32,
    parameter logic [7:0] OP_READ = 8'h03,
    parameter logic [7:0] OP_FAST = 8'h0B,
    parameter int MIN_DIV_READ = 8,   // plain read: slowest ceiling
    parameter int MIN_DIV_FAST = 2,   // fast read: can go much faster
    parameter int DUMMY_READ   = 0,
    parameter int DUMMY_FAST   = 8
) (
    input  logic              clk,
    input  logic              rst_n,

    input  logic              select,       // pulse: decide for this request
    input  logic [DIV_W-1:0]  req_div,      // the divisor the system wants
    input  logic [LEN_W-1:0]  req_len,      // payload bytes
    input  logic [2:0]        addr_bytes,

    output logic [7:0]        opcode,
    output logic [5:0]        dummy_cycles,
    output logic [DIV_W-1:0]  eff_div,      // possibly slower than requested
    output logic              fast_chosen,
    output logic              too_fast,     // even fast read cannot go this fast
    output logic [TIME_W-1:0] est_cycles,   // time of the chosen command
    output logic [TIME_W-1:0] alt_cycles    // time of the rejected one
);

    // Each command runs at the fastest divisor IT allows, which may be
    // slower than the one requested. Larger divisor means slower clock, so
    // the floor is a MAXIMUM on the divisor -- the inversion that catches
    // people, and the same one Chapter 9.4's rate limiter turns on.
    logic [DIV_W-1:0] div_read, div_fast;
    logic [TIME_W-1:0] sclk_read, sclk_fast;
    logic [TIME_W-1:0] time_read, time_fast;
    logic [TIME_W-1:0] overhead;

    always_comb begin
        div_read = (int'(req_div) < MIN_DIV_READ) ? DIV_W'(MIN_DIV_READ) : req_div;
        div_fast = (int'(req_div) < MIN_DIV_FAST) ? DIV_W'(MIN_DIV_FAST) : req_div;

        // Cycles common to both: the opcode and the address phase.
        overhead = TIME_W'(8) + (TIME_W'(addr_bytes) << 3);

        sclk_read = overhead + TIME_W'(DUMMY_READ) + (TIME_W'(req_len) << 3);
        sclk_fast = overhead + TIME_W'(DUMMY_FAST) + (TIME_W'(req_len) << 3);

        time_read = sclk_read * TIME_W'(div_read);
        time_fast = sclk_fast * TIME_W'(div_fast);
    end

    always_ff @(posedge clk or negedge rst_n) begin
        if (!rst_n) begin
            // Reset picks the plain read at the slowest divisor. It is the
            // command every 25-series device implements identically and the
            // one with no latency to get wrong -- which is exactly what a
            // boot ROM wants before it knows anything about the part.
            opcode       <= OP_READ;
            dummy_cycles <= 6'(DUMMY_READ);
            eff_div      <= {DIV_W{1'b1}};
            fast_chosen  <= 1'b0;
            too_fast     <= 1'b0;
            est_cycles   <= {TIME_W{1'b0}};
            alt_cycles   <= {TIME_W{1'b0}};
        end else if (select) begin
            // Strictly less: on a tie the plain read wins, because it has
            // fewer cycles and therefore less exposure to a latency
            // mismatch. A tie-break toward the simpler command is a real
            // engineering preference, not an arbitrary one.
            if (time_fast < time_read) begin
                opcode       <= OP_FAST;
                dummy_cycles <= 6'(DUMMY_FAST);
                eff_div      <= div_fast;
                fast_chosen  <= 1'b1;
                est_cycles   <= time_fast;
                alt_cycles   <= time_read;
            end else begin
                opcode       <= OP_READ;
                dummy_cycles <= 6'(DUMMY_READ);
                eff_div      <= div_read;
                fast_chosen  <= 1'b0;
                est_cycles   <= time_read;
                alt_cycles   <= time_fast;
            end

            // The request was faster than any read command supports. The
            // choice above already clamped; this reports that it did.
            too_fast <= (int'(req_div) < MIN_DIV_FAST);
        end
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
flash_read_sel_tb.sv — the crossover, found and checked
// flash_read_sel_tb.sv
//
// The testbench computes both commands' times itself and requires the
// block to pick the smaller, then searches for the CROSSOVER length and
// checks it lies where the arithmetic says.

`timescale 1ns/1ps

module flash_read_sel_tb;

    localparam int DIV_W  = 8;
    localparam int LEN_W  = 16;
    localparam int TIME_W = 32;
    localparam int MIN_DIV_READ = 8;
    localparam int MIN_DIV_FAST = 2;
    localparam int DUMMY_FAST   = 8;

    logic clk = 1'b0;
    logic rst_n = 1'b0;
    always #5 clk = ~clk;

    logic              select = 1'b0;
    logic [DIV_W-1:0]  req_div = 8'd8;
    logic [LEN_W-1:0]  req_len = 16'd0;
    logic [2:0]        addr_bytes = 3'd3;

    logic [7:0]        opcode;
    logic [5:0]        dummy_cycles;
    logic [DIV_W-1:0]  eff_div;
    logic              fast_chosen, too_fast;
    logic [TIME_W-1:0] est_cycles, alt_cycles;

    int errors = 0;

    // Declared at module scope. A declaration inside an unnamed begin block
    // is not portable, and one carrying an initialiser inside a procedural
    // block is a STATIC initialiser evaluated once at time zero.
    int xover;
    int sw_read, sw_fast, sw_want;

    flash_read_sel #(
        .DIV_W(DIV_W), .LEN_W(LEN_W), .TIME_W(TIME_W),
        .MIN_DIV_READ(MIN_DIV_READ), .MIN_DIV_FAST(MIN_DIV_FAST),
        .DUMMY_READ(0), .DUMMY_FAST(DUMMY_FAST)
    ) dut (
        .clk(clk), .rst_n(rst_n), .select(select),
        .req_div(req_div), .req_len(req_len), .addr_bytes(addr_bytes),
        .opcode(opcode), .dummy_cycles(dummy_cycles), .eff_div(eff_div),
        .fast_chosen(fast_chosen), .too_fast(too_fast),
        .est_cycles(est_cycles), .alt_cycles(alt_cycles)
    );

    // The testbench's own model, written from the datasheet numbers rather
    // than from the design.
    function automatic int model_time(input int div, input int len,
                                      input int ab, input int dummy,
                                      input int min_div);
        int d;
        begin
            d = (div < min_div) ? min_div : div;
            model_time = (8 + ab*8 + dummy + len*8) * d;
        end
    endfunction

    task automatic decide(input int div, input int len);
        begin
            @(negedge clk);
            req_div = DIV_W'(div); req_len = LEN_W'(len);
            select = 1'b1;
            @(negedge clk);
            select = 1'b0;
            @(negedge clk);
        end
    endtask

    task automatic expect_choice(input int div, input int len,
                                 input bit want_fast);
        int t_read, t_fast;
        begin
            decide(div, len);
            t_read = model_time(div, len, 3, 0,          MIN_DIV_READ);
            t_fast = model_time(div, len, 3, DUMMY_FAST, MIN_DIV_FAST);
            if (fast_chosen !== want_fast) begin
                $display("  FAIL: div=%0d len=%0d chose %s (read %0d, fast %0d cycles)",
                         div, len, fast_chosen ? "fast" : "plain", t_read, t_fast);
                errors++;
            end
            if (est_cycles !== TIME_W'(want_fast ? t_fast : t_read)) begin
                $display("  FAIL: div=%0d len=%0d reported %0d cycles, model says %0d",
                         div, len, est_cycles, want_fast ? t_fast : t_read);
                errors++;
            end
            if (opcode !== (want_fast ? 8'h0B : 8'h03)) begin
                $display("  FAIL: div=%0d len=%0d opcode 0x%02h", div, len, opcode);
                errors++;
            end
            if (dummy_cycles !== (want_fast ? 6'(DUMMY_FAST) : 6'd0)) begin
                $display("  FAIL: div=%0d len=%0d dummy %0d", div, len, dummy_cycles);
                errors++;
            end
        end
    endtask

    initial begin
        repeat (3) @(negedge clk);
        rst_n = 1'b1;
        @(negedge clk);

        // 1. Reset picks the plain read at the slowest divisor -- what a
        //    boot ROM wants before it knows anything about the part.
        if (opcode !== 8'h03 || eff_div !== 8'hFF || fast_chosen) begin
            $display("  FAIL: reset did not select the plain read at the slowest divisor");
            errors++;
        end
        $display("  reset: opcode=0x%02h div=%0d dummy=%0d", opcode, eff_div,
                 dummy_cycles);

        // 2. At a divisor the plain read can meet, it WINS -- same rate,
        //    eight fewer cycles. This is the case the usual advice gets
        //    wrong.
        expect_choice(8,  64, 1'b0);
        expect_choice(16, 64, 1'b0);
        expect_choice(8,   1, 1'b0);
        $display("  div=8  len=64: plain read chosen, %0d cycles vs %0d for fast",
                 est_cycles, alt_cycles);

        // 3. Well below the plain read's ceiling, the fast read's higher
        //    rate beats its eight extra cycles comfortably.
        expect_choice(2, 64, 1'b1);
        expect_choice(4, 64, 1'b1);
        expect_choice(2,  1, 1'b1);
        decide(2, 64);
        $display("  div=2  len=64: fast read chosen, %0d cycles vs %0d for plain",
                 est_cycles, alt_cycles);

        // 4. THE CROSSOVER. At divisor 7 the plain read must slow to 8
        //    while the fast read runs at 7. Short transfers favour the
        //    plain read's smaller cycle count; long ones favour the fast
        //    read's higher rate. Find where they xover, from the model.
        xover = -1;
        for (int len = 0; len <= 64; len++) begin
            sw_read = model_time(7, len, 3, 0,          MIN_DIV_READ);
            sw_fast = model_time(7, len, 3, DUMMY_FAST, MIN_DIV_FAST);
            if (sw_fast < sw_read && xover < 0) xover = len;
        end
        if (xover < 1) begin
            $display("  FAIL: no crossover found at divisor 7 -- the premise is wrong");
            errors++;
        end else begin
            $display("  crossover at divisor 7: plain read wins below %0d bytes, fast read from %0d up",
                     xover, xover);
            // Just below the crossover the plain read must win.
            expect_choice(7, xover - 1, 1'b0);
            // At and above it, the fast read must win.
            expect_choice(7, xover,     1'b1);
            expect_choice(7, xover + 8, 1'b1);
        end

        // 5. A request faster than any read command supports is clamped and
        //    reported, not silently honoured.
        decide(1, 64);
        if (!too_fast) begin
            $display("  FAIL: divisor 1 not reported as beyond both commands");
            errors++;
        end
        if (int'(eff_div) < MIN_DIV_FAST) begin
            $display("  FAIL: effective divisor %0d is faster than the fast read allows",
                     eff_div);
            errors++;
        end
        $display("  div=1: clamped to %0d and reported (too_fast=%0b)", eff_div,
                 too_fast);

        // 6. A legal request must never be reported as too fast.
        decide(8, 64);
        if (too_fast) begin
            $display("  FAIL: a legal divisor was reported as too fast"); errors++;
        end

        // 7. SWEEP. Over a wide range the block must always pick the
        //    smaller of the two modelled times, and never return a divisor
        //    faster than the chosen command allows.
        for (int div = 1; div <= 20; div++) begin
            for (int len = 0; len <= 40; len += 4) begin
                decide(div, len);
                sw_read = model_time(div, len, 3, 0,          MIN_DIV_READ);
                sw_fast = model_time(div, len, 3, DUMMY_FAST, MIN_DIV_FAST);
                sw_want = (sw_fast < sw_read) ? sw_fast : sw_read;
                if (est_cycles !== TIME_W'(sw_want)) begin
                    $display("  FAIL: div=%0d len=%0d chose %0d cycles, best is %0d",
                             div, len, est_cycles, sw_want);
                    errors++;
                end
                if (fast_chosen && int'(eff_div) < MIN_DIV_FAST) begin
                    $display("  FAIL: fast read at divisor %0d", eff_div);
                    errors++;
                end
                if (!fast_chosen && int'(eff_div) < MIN_DIV_READ) begin
                    $display("  FAIL: plain read at divisor %0d", eff_div);
                    errors++;
                end
            end
        end
        $display("  220 (divisor, length) pairs swept: the faster command is always chosen and never overclocked");

        if (errors == 0)
            $display("PASS: the command is chosen by transfer TIME rather than clock rate, the plain read wins whenever it can meet the requested divisor, the crossover at a divisor just below its ceiling falls where the arithmetic says, and neither command is ever run faster than it allows");
        else
            $display("FAIL: %0d error(s)", errors);
        $finish;
    end

endmodule

The testbench models both commands' times itself, from the datasheet numbers, and never reads an estimate back from the design to compare against. Three of its checks carry the weight.

At a divisor the plain read can meet, it must be chosen — and the printed numbers show the margin: 320 cycles against 384, an 18 per cent saving from not using the fast command.

The crossover is located by the model, not asserted. The testbench searches for the first length at which the fast read wins at divisor 7, prints it, and then requires the selector to agree on both sides. Had the premise been wrong — had no crossover existed — the search would have reported that instead of quietly passing.

A sweep over 220 (divisor, length) pairs requires the smaller of the two modelled times every time, and separately requires that neither command is ever run faster than it allows. That second check is the safety one: a selector that picked correctly but forgot to clamp would pass the first.

Azvya Education Pvt. Ltd.VLSI Mentor
flash_read_sel.v — the same selector in Verilog-2001
// flash_read_sel.v
//
// Chapter 11.2 -- choosing between the plain read and the fast read, in
// Verilog-2001.
//
//   0x03  Read        no dummy cycles, but a LOW maximum clock rate
//   0x0B  Fast Read   a high maximum rate, paid for with dummy cycles
//
// The two commands trade CLOCK RATE against CYCLE COUNT, so which is
// faster depends on the transfer LENGTH. This block computes the time each
// would take, in system-clock cycles, and picks the smaller:
//
//   time = sclk_cycles x divisor
//   sclk_cycles = 8 (opcode) + 8 x addr_bytes + dummy + 8 x length
//
// At a divisor the plain read can meet it wins outright -- same rate,
// eight fewer cycles. Below its ceiling the fast read's higher rate
// usually wins. Just below the plain read's ceiling the two CROSS.

module flash_read_sel #(
    parameter DIV_W        = 8,
    parameter LEN_W        = 16,
    parameter TIME_W       = 32,
    parameter [7:0] OP_READ = 8'h03,
    parameter [7:0] OP_FAST = 8'h0B,
    parameter MIN_DIV_READ = 8,   // plain read: slowest ceiling
    parameter MIN_DIV_FAST = 2,   // fast read: can go much faster
    parameter DUMMY_READ   = 0,
    parameter DUMMY_FAST   = 8
) (
    input  wire              clk,
    input  wire              rst_n,

    input  wire              select,       // pulse: decide for this request
    input  wire [DIV_W-1:0]  req_div,      // the divisor the system wants
    input  wire [LEN_W-1:0]  req_len,      // payload bytes
    input  wire [2:0]        addr_bytes,

    output reg  [7:0]        opcode,
    output reg  [5:0]        dummy_cycles,
    output reg  [DIV_W-1:0]  eff_div,      // possibly slower than requested
    output reg               fast_chosen,
    output reg               too_fast,     // beyond even the fast read
    output reg  [TIME_W-1:0] est_cycles,   // time of the chosen command
    output reg  [TIME_W-1:0] alt_cycles    // time of the rejected one
);

    // Each command runs at the fastest divisor IT allows, which may be
    // slower than the one requested. Larger divisor means slower clock, so
    // the floor is a MAXIMUM on the divisor -- the same inversion Chapter
    // 9.4's rate limiter turns on.
    reg [DIV_W-1:0]  div_read, div_fast;
    reg [TIME_W-1:0] sclk_read, sclk_fast;
    reg [TIME_W-1:0] time_read, time_fast;
    reg [TIME_W-1:0] overhead;

    always @(*) begin
        if (req_div < MIN_DIV_READ) div_read = MIN_DIV_READ;
        else                        div_read = req_div;
        if (req_div < MIN_DIV_FAST) div_fast = MIN_DIV_FAST;
        else                        div_fast = req_div;

        // Cycles common to both: the opcode and the address phase.
        overhead = 8 + ({{(TIME_W-3){1'b0}}, addr_bytes} << 3);

        sclk_read = overhead + DUMMY_READ + ({{(TIME_W-LEN_W){1'b0}}, req_len} << 3);
        sclk_fast = overhead + DUMMY_FAST + ({{(TIME_W-LEN_W){1'b0}}, req_len} << 3);

        time_read = sclk_read * div_read;
        time_fast = sclk_fast * div_fast;
    end

    always @(posedge clk or negedge rst_n) begin
        if (!rst_n) begin
            // Reset picks the plain read at the slowest divisor: the command
            // every 25-series device implements identically and the one with
            // no latency to get wrong -- what a boot ROM wants before it
            // knows anything about the part.
            opcode       <= OP_READ;
            dummy_cycles <= DUMMY_READ;
            eff_div      <= {DIV_W{1'b1}};
            fast_chosen  <= 1'b0;
            too_fast     <= 1'b0;
            est_cycles   <= {TIME_W{1'b0}};
            alt_cycles   <= {TIME_W{1'b0}};
        end else if (select) begin
            // Strictly less: on a tie the plain read wins, because it has
            // fewer cycles and therefore less exposure to a latency
            // mismatch. A tie-break toward the simpler command is a real
            // engineering preference.
            if (time_fast < time_read) begin
                opcode       <= OP_FAST;
                dummy_cycles <= DUMMY_FAST;
                eff_div      <= div_fast;
                fast_chosen  <= 1'b1;
                est_cycles   <= time_fast;
                alt_cycles   <= time_read;
            end else begin
                opcode       <= OP_READ;
                dummy_cycles <= DUMMY_READ;
                eff_div      <= div_read;
                fast_chosen  <= 1'b0;
                est_cycles   <= time_read;
                alt_cycles   <= time_fast;
            end

            too_fast <= (req_div < MIN_DIV_FAST);
        end
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
flash_read_sel_tb.v — the same crossover in Verilog-2001
// flash_read_sel_tb.v
//
// The same checks as the SystemVerilog testbench: both commands' times
// modelled independently, the smaller required, and the crossover length
// found from the model and verified on either side.

`timescale 1ns/1ps

module flash_read_sel_tb;

    parameter DIV_W  = 8;
    parameter LEN_W  = 16;
    parameter TIME_W = 32;
    parameter MIN_DIV_READ = 8;
    parameter MIN_DIV_FAST = 2;
    parameter DUMMY_FAST   = 8;

    reg clk;
    reg rst_n;

    reg              select;
    reg  [DIV_W-1:0] req_div;
    reg  [LEN_W-1:0] req_len;
    reg  [2:0]       addr_bytes;

    wire [7:0]        opcode;
    wire [5:0]        dummy_cycles;
    wire [DIV_W-1:0]  eff_div;
    wire              fast_chosen, too_fast;
    wire [TIME_W-1:0] est_cycles, alt_cycles;

    integer errors;
    integer xover;
    integer sw_read, sw_fast, sw_want;
    integer div_i, len_i;

    initial begin
        clk = 1'b0; rst_n = 1'b0; select = 1'b0;
        req_div = 8'd8; req_len = 16'd0; addr_bytes = 3'd3;
        errors = 0;
    end
    always #5 clk = ~clk;

    flash_read_sel #(
        .DIV_W(DIV_W), .LEN_W(LEN_W), .TIME_W(TIME_W),
        .MIN_DIV_READ(MIN_DIV_READ), .MIN_DIV_FAST(MIN_DIV_FAST),
        .DUMMY_READ(0), .DUMMY_FAST(DUMMY_FAST)
    ) dut (
        .clk(clk), .rst_n(rst_n), .select(select),
        .req_div(req_div), .req_len(req_len), .addr_bytes(addr_bytes),
        .opcode(opcode), .dummy_cycles(dummy_cycles), .eff_div(eff_div),
        .fast_chosen(fast_chosen), .too_fast(too_fast),
        .est_cycles(est_cycles), .alt_cycles(alt_cycles)
    );

    // The testbench's own model, from the datasheet numbers rather than
    // from the design.
    function integer model_time;
        input integer div;
        input integer len;
        input integer ab;
        input integer dummy;
        input integer min_div;
        integer d;
        begin
            if (div < min_div) d = min_div; else d = div;
            model_time = (8 + ab*8 + dummy + len*8) * d;
        end
    endfunction

    task decide;
        input integer div;
        input integer len;
        begin
            @(negedge clk);
            req_div = div[DIV_W-1:0]; req_len = len[LEN_W-1:0];
            select = 1'b1;
            @(negedge clk);
            select = 1'b0;
            @(negedge clk);
        end
    endtask

    task expect_choice;
        input integer div;
        input integer len;
        input         want_fast;
        integer t_read, t_fast;
        begin
            decide(div, len);
            t_read = model_time(div, len, 3, 0,          MIN_DIV_READ);
            t_fast = model_time(div, len, 3, DUMMY_FAST, MIN_DIV_FAST);
            if (fast_chosen !== want_fast) begin
                $display("  FAIL: div=%0d len=%0d chose the wrong command (read %0d, fast %0d cycles)",
                         div, len, t_read, t_fast);
                errors = errors + 1;
            end
            if (est_cycles !== (want_fast ? t_fast : t_read)) begin
                $display("  FAIL: div=%0d len=%0d reported %0d cycles, model says %0d",
                         div, len, est_cycles, want_fast ? t_fast : t_read);
                errors = errors + 1;
            end
            if (opcode !== (want_fast ? 8'h0B : 8'h03)) begin
                $display("  FAIL: div=%0d len=%0d opcode 0x%02h", div, len, opcode);
                errors = errors + 1;
            end
            if (dummy_cycles !== (want_fast ? DUMMY_FAST : 0)) begin
                $display("  FAIL: div=%0d len=%0d dummy %0d", div, len, dummy_cycles);
                errors = errors + 1;
            end
        end
    endtask

    initial begin
        repeat (3) @(negedge clk);
        rst_n = 1'b1;
        @(negedge clk);

        // 1. Reset picks the plain read at the slowest divisor.
        if (opcode !== 8'h03 || eff_div !== 8'hFF || fast_chosen) begin
            $display("  FAIL: reset did not select the plain read at the slowest divisor");
            errors = errors + 1;
        end
        $display("  reset: opcode=0x%02h div=%0d dummy=%0d", opcode, eff_div,
                 dummy_cycles);

        // 2. At a divisor the plain read can meet, it WINS -- same rate,
        //    eight fewer cycles.
        expect_choice(8,  64, 1'b0);
        expect_choice(16, 64, 1'b0);
        expect_choice(8,   1, 1'b0);
        $display("  div=8  len=64: plain read chosen, %0d cycles vs %0d for fast",
                 est_cycles, alt_cycles);

        // 3. Well below the plain read's ceiling, the fast read wins.
        expect_choice(2, 64, 1'b1);
        expect_choice(4, 64, 1'b1);
        expect_choice(2,  1, 1'b1);
        decide(2, 64);
        $display("  div=2  len=64: fast read chosen, %0d cycles vs %0d for plain",
                 est_cycles, alt_cycles);

        // 4. THE CROSSOVER at divisor 7, found from the model.
        xover = -1;
        for (len_i = 0; len_i <= 64; len_i = len_i + 1) begin
            sw_read = model_time(7, len_i, 3, 0,          MIN_DIV_READ);
            sw_fast = model_time(7, len_i, 3, DUMMY_FAST, MIN_DIV_FAST);
            if (sw_fast < sw_read && xover < 0) xover = len_i;
        end
        if (xover < 1) begin
            $display("  FAIL: no crossover found at divisor 7 -- the premise is wrong");
            errors = errors + 1;
        end else begin
            $display("  crossover at divisor 7: plain read wins below %0d bytes, fast read from %0d up",
                     xover, xover);
            expect_choice(7, xover - 1, 1'b0);
            expect_choice(7, xover,     1'b1);
            expect_choice(7, xover + 8, 1'b1);
        end

        // 5. A request faster than any read command is clamped and reported.
        decide(1, 64);
        if (!too_fast) begin
            $display("  FAIL: divisor 1 not reported as beyond both commands");
            errors = errors + 1;
        end
        if (eff_div < MIN_DIV_FAST) begin
            $display("  FAIL: effective divisor %0d is faster than the fast read allows",
                     eff_div);
            errors = errors + 1;
        end
        $display("  div=1: clamped to %0d and reported (too_fast=%0b)", eff_div,
                 too_fast);

        // 6. A legal request must never be reported as too fast.
        decide(8, 64);
        if (too_fast) begin
            $display("  FAIL: a legal divisor was reported as too fast");
            errors = errors + 1;
        end

        // 7. SWEEP.
        for (div_i = 1; div_i <= 20; div_i = div_i + 1) begin
            for (len_i = 0; len_i <= 40; len_i = len_i + 4) begin
                decide(div_i, len_i);
                sw_read = model_time(div_i, len_i, 3, 0,          MIN_DIV_READ);
                sw_fast = model_time(div_i, len_i, 3, DUMMY_FAST, MIN_DIV_FAST);
                if (sw_fast < sw_read) sw_want = sw_fast; else sw_want = sw_read;
                if (est_cycles !== sw_want) begin
                    $display("  FAIL: div=%0d len=%0d chose %0d cycles, best is %0d",
                             div_i, len_i, est_cycles, sw_want);
                    errors = errors + 1;
                end
                if (fast_chosen && eff_div < MIN_DIV_FAST) begin
                    $display("  FAIL: fast read at divisor %0d", eff_div);
                    errors = errors + 1;
                end
                if (!fast_chosen && eff_div < MIN_DIV_READ) begin
                    $display("  FAIL: plain read at divisor %0d", eff_div);
                    errors = errors + 1;
                end
            end
        end
        $display("  220 (divisor, length) pairs swept: the faster command is always chosen and never overclocked");

        if (errors == 0)
            $display("PASS: the command is chosen by transfer TIME rather than clock rate, the plain read wins whenever it can meet the requested divisor, the crossover at a divisor just below its ceiling falls where the arithmetic says, and neither command is ever run faster than it allows");
        else
            $display("FAIL: %0d error(s)", errors);
        $finish;
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
flash_read_sel.vhd — the same selector in VHDL
-- flash_read_sel.vhd
--
-- Chapter 11.2 -- choosing between the plain read and the fast read, in
-- VHDL.
--
--   0x03  Read        no dummy cycles, but a LOW maximum clock rate
--   0x0B  Fast Read   a high maximum rate, paid for with dummy cycles
--
-- The two commands trade CLOCK RATE against CYCLE COUNT, so which is
-- faster depends on the transfer LENGTH. This block computes the time each
-- would take, in system-clock cycles, and picks the smaller:
--
--   time = sclk_cycles * divisor
--   sclk_cycles = 8 (opcode) + 8 * addr_bytes + dummy + 8 * length
--
-- At a divisor the plain read can meet it wins outright -- same rate,
-- eight fewer cycles. Below its ceiling the fast read's higher rate
-- usually wins. Just below the plain read's ceiling the two CROSS.

library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

entity flash_read_sel is
    generic (
        DIV_W        : positive := 8;
        LEN_W        : positive := 16;
        TIME_W       : positive := 32;
        OP_READ      : natural  := 16#03#;
        OP_FAST      : natural  := 16#0B#;
        MIN_DIV_READ : natural  := 8;   -- plain read: slowest ceiling
        MIN_DIV_FAST : natural  := 2;   -- fast read: can go much faster
        DUMMY_READ   : natural  := 0;
        DUMMY_FAST   : natural  := 8
    );
    port (
        clk          : in  std_logic;
        rst_n        : in  std_logic;

        select_req   : in  std_logic;   -- pulse: decide for this request
        req_div      : in  unsigned(DIV_W - 1 downto 0);
        req_len      : in  unsigned(LEN_W - 1 downto 0);
        addr_bytes   : in  unsigned(2 downto 0);

        opcode       : out unsigned(7 downto 0);
        dummy_cycles : out unsigned(5 downto 0);
        eff_div      : out unsigned(DIV_W - 1 downto 0);
        fast_chosen  : out std_logic;
        too_fast     : out std_logic;
        est_cycles   : out unsigned(TIME_W - 1 downto 0);
        alt_cycles   : out unsigned(TIME_W - 1 downto 0)
    );
end entity;

architecture rtl of flash_read_sel is

    signal op_r    : unsigned(7 downto 0) := to_unsigned(OP_READ, 8);
    signal dum_r   : unsigned(5 downto 0) := to_unsigned(DUMMY_READ, 6);
    signal div_r   : unsigned(DIV_W - 1 downto 0) := (others => '1');
    signal fast_r  : std_logic := '0';
    signal tf_r    : std_logic := '0';
    signal est_r   : unsigned(TIME_W - 1 downto 0) := (others => '0');
    signal alt_r   : unsigned(TIME_W - 1 downto 0) := (others => '0');

    -- Declaration initialisers keep the comparison below from testing 'U'
    -- before the first reset; rst_n still loads every register.
    signal div_read  : natural := MIN_DIV_READ;
    signal div_fast  : natural := MIN_DIV_FAST;
    signal time_read : natural := 0;
    signal time_fast : natural := 0;

begin

    -- Each command runs at the fastest divisor IT allows, which may be
    -- slower than the one requested. Larger divisor means slower clock, so
    -- the floor is a MAXIMUM on the divisor -- the same inversion Chapter
    -- 9.4's rate limiter turns on.
    estimate : process (req_div, req_len, addr_bytes)
        variable dr, df   : natural;
        variable overhead : natural;
        variable sr, sf   : natural;
    begin
        if to_integer(req_div) < MIN_DIV_READ then dr := MIN_DIV_READ;
        else                                      dr := to_integer(req_div);
        end if;
        if to_integer(req_div) < MIN_DIV_FAST then df := MIN_DIV_FAST;
        else                                      df := to_integer(req_div);
        end if;

        -- Cycles common to both: the opcode and the address phase.
        overhead := 8 + to_integer(addr_bytes) * 8;

        sr := overhead + DUMMY_READ + to_integer(req_len) * 8;
        sf := overhead + DUMMY_FAST + to_integer(req_len) * 8;

        div_read  <= dr;
        div_fast  <= df;
        time_read <= sr * dr;
        time_fast <= sf * df;
    end process;

    choose : process (clk, rst_n)
    begin
        if rst_n = '0' then
            -- Reset picks the plain read at the slowest divisor: the command
            -- every 25-series device implements identically and the one with
            -- no latency to get wrong -- what a boot ROM wants before it
            -- knows anything about the part.
            op_r   <= to_unsigned(OP_READ, 8);
            dum_r  <= to_unsigned(DUMMY_READ, 6);
            div_r  <= (others => '1');
            fast_r <= '0';
            tf_r   <= '0';
            est_r  <= (others => '0');
            alt_r  <= (others => '0');
        elsif rising_edge(clk) then
            if select_req = '1' then
                -- Strictly less: on a tie the plain read wins, because it
                -- has fewer cycles and therefore less exposure to a latency
                -- mismatch.
                if time_fast < time_read then
                    op_r   <= to_unsigned(OP_FAST, 8);
                    dum_r  <= to_unsigned(DUMMY_FAST, 6);
                    div_r  <= to_unsigned(div_fast, DIV_W);
                    fast_r <= '1';
                    est_r  <= to_unsigned(time_fast, TIME_W);
                    alt_r  <= to_unsigned(time_read, TIME_W);
                else
                    op_r   <= to_unsigned(OP_READ, 8);
                    dum_r  <= to_unsigned(DUMMY_READ, 6);
                    div_r  <= to_unsigned(div_read, DIV_W);
                    fast_r <= '0';
                    est_r  <= to_unsigned(time_read, TIME_W);
                    alt_r  <= to_unsigned(time_fast, TIME_W);
                end if;

                if to_integer(req_div) < MIN_DIV_FAST then tf_r <= '1';
                else                                       tf_r <= '0';
                end if;
            end if;
        end if;
    end process;

    opcode       <= op_r;
    dummy_cycles <= dum_r;
    eff_div      <= div_r;
    fast_chosen  <= fast_r;
    too_fast     <= tf_r;
    est_cycles   <= est_r;
    alt_cycles   <= alt_r;

end architecture;
Azvya Education Pvt. Ltd.VLSI Mentor
flash_read_sel_tb.vhd — the same crossover in VHDL
-- flash_read_sel_tb.vhd
--
-- The same checks as the SystemVerilog and Verilog testbenches: both
-- commands' times modelled independently, the smaller required, and the
-- crossover length found from the model and verified on either side.

library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

entity flash_read_sel_tb is
end entity;

architecture sim of flash_read_sel_tb is

    constant DIV_W        : positive := 8;
    constant LEN_W        : positive := 16;
    constant TIME_W       : positive := 32;
    constant MIN_DIV_READ : natural  := 8;
    constant MIN_DIV_FAST : natural  := 2;
    constant DUMMY_FAST   : natural  := 8;

    signal clk   : std_logic := '0';
    signal rst_n : std_logic := '0';
    signal halt  : boolean   := false;

    signal select_req : std_logic := '0';
    signal req_div    : unsigned(DIV_W - 1 downto 0) := to_unsigned(8, DIV_W);
    signal req_len    : unsigned(LEN_W - 1 downto 0) := (others => '0');
    signal addr_bytes : unsigned(2 downto 0) := to_unsigned(3, 3);

    signal opcode       : unsigned(7 downto 0);
    signal dummy_cycles : unsigned(5 downto 0);
    signal eff_div      : unsigned(DIV_W - 1 downto 0);
    signal fast_chosen  : std_logic;
    signal too_fast     : std_logic;
    signal est_cycles   : unsigned(TIME_W - 1 downto 0);
    signal alt_cycles   : unsigned(TIME_W - 1 downto 0);

    signal errors : natural := 0;

    -- The testbench's own model, from the datasheet numbers rather than
    -- from the design.
    function model_time(div : natural; len : natural; ab : natural;
                        dummy : natural; min_div : natural) return natural is
        variable d : natural;
    begin
        if div < min_div then d := min_div; else d := div; end if;
        return (8 + ab * 8 + dummy + len * 8) * d;
    end function;

begin

    clk <= not clk after 5 ns when not halt else '0';

    dut : entity work.flash_read_sel
        generic map (DIV_W => DIV_W, LEN_W => LEN_W, TIME_W => TIME_W,
                     MIN_DIV_READ => MIN_DIV_READ, MIN_DIV_FAST => MIN_DIV_FAST,
                     DUMMY_READ => 0, DUMMY_FAST => DUMMY_FAST)
        port map (
            clk => clk, rst_n => rst_n, select_req => select_req,
            req_div => req_div, req_len => req_len, addr_bytes => addr_bytes,
            opcode => opcode, dummy_cycles => dummy_cycles, eff_div => eff_div,
            fast_chosen => fast_chosen, too_fast => too_fast,
            est_cycles => est_cycles, alt_cycles => alt_cycles
        );

    stim : process
        variable errs  : natural := 0;
        variable xover : integer;
        variable sw_read, sw_fast, sw_want : natural;

        procedure decide(div : natural; len : natural) is
        begin
            wait until falling_edge(clk);
            req_div    <= to_unsigned(div, DIV_W);
            req_len    <= to_unsigned(len, LEN_W);
            select_req <= '1';
            wait until falling_edge(clk);
            select_req <= '0';
            wait until falling_edge(clk);
        end procedure;

        procedure expect_choice(div : natural; len : natural;
                                want_fast : boolean) is
            variable t_read, t_fast : natural;
        begin
            decide(div, len);
            t_read := model_time(div, len, 3, 0,          MIN_DIV_READ);
            t_fast := model_time(div, len, 3, DUMMY_FAST, MIN_DIV_FAST);
            if (fast_chosen = '1') /= want_fast then
                report "  FAIL: div=" & integer'image(div) & " len=" &
                       integer'image(len) & " chose the wrong command";
                errs := errs + 1;
            end if;
            if want_fast then
                if to_integer(est_cycles) /= t_fast then
                    report "  FAIL: reported cycles do not match the model";
                    errs := errs + 1;
                end if;
                if to_integer(opcode) /= 16#0B# or
                   to_integer(dummy_cycles) /= DUMMY_FAST then
                    report "  FAIL: fast read opcode or dummy count wrong";
                    errs := errs + 1;
                end if;
            else
                if to_integer(est_cycles) /= t_read then
                    report "  FAIL: reported cycles do not match the model";
                    errs := errs + 1;
                end if;
                if to_integer(opcode) /= 16#03# or dummy_cycles /= 0 then
                    report "  FAIL: plain read opcode or dummy count wrong";
                    errs := errs + 1;
                end if;
            end if;
        end procedure;
    begin
        for i in 0 to 2 loop
            wait until falling_edge(clk);
        end loop;
        rst_n <= '1';
        wait until falling_edge(clk);

        -- 1. Reset picks the plain read at the slowest divisor.
        if to_integer(opcode) /= 16#03# or eff_div /= (eff_div'range => '1') or
           fast_chosen = '1' then
            report "  FAIL: reset did not select the plain read at the slowest divisor";
            errs := errs + 1;
        end if;
        report "  reset: opcode=0x03 div=" &
               integer'image(to_integer(eff_div)) & " dummy=" &
               integer'image(to_integer(dummy_cycles));

        -- 2. At a divisor the plain read can meet, it WINS.
        expect_choice(8,  64, false);
        expect_choice(16, 64, false);
        expect_choice(8,   1, false);
        report "  div=8  len=64: plain read chosen, " &
               integer'image(to_integer(est_cycles)) & " cycles vs " &
               integer'image(to_integer(alt_cycles)) & " for fast";

        -- 3. Well below the plain read's ceiling, the fast read wins.
        expect_choice(2, 64, true);
        expect_choice(4, 64, true);
        expect_choice(2,  1, true);
        decide(2, 64);
        report "  div=2  len=64: fast read chosen, " &
               integer'image(to_integer(est_cycles)) & " cycles vs " &
               integer'image(to_integer(alt_cycles)) & " for plain";

        -- 4. THE CROSSOVER at divisor 7, found from the model.
        xover := -1;
        for len in 0 to 64 loop
            sw_read := model_time(7, len, 3, 0,          MIN_DIV_READ);
            sw_fast := model_time(7, len, 3, DUMMY_FAST, MIN_DIV_FAST);
            if sw_fast < sw_read and xover < 0 then
                xover := len;
            end if;
        end loop;
        if xover < 1 then
            report "  FAIL: no crossover found at divisor 7 -- the premise is wrong";
            errs := errs + 1;
        else
            report "  crossover at divisor 7: plain read wins below " &
                   integer'image(xover) & " bytes, fast read from " &
                   integer'image(xover) & " up";
            expect_choice(7, xover - 1, false);
            expect_choice(7, xover,     true);
            expect_choice(7, xover + 8, true);
        end if;

        -- 5. A request faster than any read command is clamped and reported.
        decide(1, 64);
        if too_fast /= '1' then
            report "  FAIL: divisor 1 not reported as beyond both commands";
            errs := errs + 1;
        end if;
        if to_integer(eff_div) < MIN_DIV_FAST then
            report "  FAIL: the effective divisor is faster than the fast read allows";
            errs := errs + 1;
        end if;
        report "  div=1: clamped to " & integer'image(to_integer(eff_div)) &
               " and reported";

        -- 6. A legal request must never be reported as too fast.
        decide(8, 64);
        if too_fast = '1' then
            report "  FAIL: a legal divisor was reported as too fast";
            errs := errs + 1;
        end if;

        -- 7. SWEEP.
        for div_i in 1 to 20 loop
            for k in 0 to 10 loop
                decide(div_i, k * 4);
                sw_read := model_time(div_i, k * 4, 3, 0,          MIN_DIV_READ);
                sw_fast := model_time(div_i, k * 4, 3, DUMMY_FAST, MIN_DIV_FAST);
                if sw_fast < sw_read then sw_want := sw_fast;
                else                      sw_want := sw_read; end if;
                if to_integer(est_cycles) /= sw_want then
                    report "  FAIL: a swept pair did not choose the faster command";
                    errs := errs + 1;
                end if;
                if fast_chosen = '1' and to_integer(eff_div) < MIN_DIV_FAST then
                    report "  FAIL: fast read overclocked"; errs := errs + 1;
                end if;
                if fast_chosen = '0' and to_integer(eff_div) < MIN_DIV_READ then
                    report "  FAIL: plain read overclocked"; errs := errs + 1;
                end if;
            end loop;
        end loop;
        report "  220 (divisor, length) pairs swept: the faster command is always chosen and never overclocked";

        errors <= errs;
        if errs = 0 then
            report "PASS: the command is chosen by transfer TIME rather than clock rate, the plain read wins whenever it can meet the requested divisor, the crossover at a divisor just below its ceiling falls where the arithmetic says, and neither command is ever run faster than it allows";
        else
            report "FAIL: " & integer'image(errs) & " error(s)" severity error;
        end if;
        halt <= true;
        wait;
    end process;

end architecture;

Parity

All three implement the same selector: identical ports and generics, a per-command divisor clamp that is a maximum on the divisor, a decision by estimated time with the plain read winning ties, reset to the plain read at the slowest divisor, and a too_fast report. All three testbenches find the crossover at four bytes at divisor 7 and report identical estimates throughout — 320 against 384 cycles at divisor 8, and 1104 against 4352 at divisor 2.

7. Why a Verification Engineer Cares

Azvya Education Pvt. Ltd.VLSI Mentor
flash_read_sel.sva — optimality and safety are separate
   // 1. OPTIMALITY. The chosen command is the faster of the two. This is
   //    what the block is for, and it is checked against an independently
   //    computed model rather than against the design's own estimate.
   a_optimal : assert property (
       @(posedge clk) disable iff (!rst_n)
           (decided) |-> (est_cycles == ((model_fast < model_read)
                                          ? model_fast : model_read)))
       else $error("a slower command was chosen");

   // 2. SAFETY, which is INDEPENDENT of optimality. Neither command may be
   //    run faster than it allows. A selector that always picked correctly
   //    but forgot to clamp would satisfy property 1 and destroy the link.
   a_never_overclocked : assert property (
       @(posedge clk) disable iff (!rst_n)
           (decided) |-> (fast_chosen ? (eff_div >= MIN_DIV_FAST)
                                      : (eff_div >= MIN_DIV_READ)))
       else $error("the chosen command was run faster than it permits");

   // 3. CONSISTENCY. The opcode, the dummy count and the flag agree. A
   //    fast-read opcode with a dummy count of zero is a link that fails
   //    with a one-byte shift -- and every output looks individually sane.
   a_consistent : assert property (
       @(posedge clk) disable iff (!rst_n)
           (decided) |-> ((fast_chosen && opcode == OP_FAST &&
                           dummy_cycles == DUMMY_FAST) ||
                          (!fast_chosen && opcode == OP_READ &&
                           dummy_cycles == DUMMY_READ)))
       else $error("the opcode and dummy count disagree with the choice");

   // 4. RESET SAFETY. Out of reset the plain read at the slowest divisor --
   //    the command every device implements identically, with no latency
   //    to get wrong. A boot ROM depends on this before it knows anything.
   a_reset_safe : assert property (
       @(posedge clk) (!rst_n) |=> (opcode == OP_READ &&
                                    dummy_cycles == 0 &&
                                    eff_div == '1))
       else $error("reset did not select the safe read");

   // 5. The tie goes to the plain read -- the simpler command when the
   //    metric cannot distinguish them.
   a_tie_to_plain : assert property (
       @(posedge clk) disable iff (!rst_n)
           (decided && (model_fast == model_read)) |-> !fast_chosen)
       else $error("a tie did not go to the plain read");

Properties 1 and 2 being separate is the point worth carrying. Optimality and safety are independent claims, and the natural bug satisfies one while violating the other: a selector that compares correctly but applies the wrong divisor picks the right command and then overclocks it. A single combined assertion would pass.

Property 3 catches the mismatch that produces the module's signature failure. A fast-read opcode issued with a dummy count of zero shifts the entire payload by one byte, and every output is individually plausible — the opcode is valid, the dummy count is a legal number, the divisor is in range. Only their agreement is wrong.

Coverage must cross rate with length, because the crossover lives at one particular pair:

Azvya Education Pvt. Ltd.VLSI Mentor
flash_read_cg.sv — the interesting region is a corner
   covergroup flash_read_cg @(posedge clk iff select);
       // Where the requested divisor sits relative to the plain read's
       // ceiling is what decides whether there is a choice to make at all.
       cp_div : coverpoint div_class {
           bins below_fast_min  = {D_TOO_FAST};    // clamped
           bins fast_only       = {D_FAST_ONLY};   // below the plain ceiling
           bins just_below      = {D_ONE_UNDER};   // the crossover region
           bins at_plain_min    = {D_AT_PLAIN};    // plain read wins
           bins well_above      = {D_SLOW};        // plain read wins
       }

       // Length matters ONLY in the crossover region, which is exactly why
       // it must be crossed with the divisor rather than covered alone.
       cp_len : coverpoint req_len {
           bins zero    = {0};
           bins tiny    = {[1:3]};        // below the crossover
           bins at_knee = {[4:8]};        // at and just above it
           bins medium  = {[9:64]};
           bins large   = {[65:$]};
       }

       cp_choice : coverpoint fast_chosen { bins plain = {0}; bins fast = {1}; }

       // The cross is the coverage goal. A suite that only ever requests
       // the maximum rate always chooses the fast read and never exercises
       // the comparison at all.
       x_div_len    : cross cp_div, cp_len;
       x_div_choice : cross cp_div, cp_choice;
   endgroup

x_div_choice is the cross that catches the lazy suite. A testbench that always requests the fastest divisor gets fast every time, reports full coverage of the opcode and dummy outputs, and has never once exercised the comparison the block exists to perform.

8. Why an FPGA or ASIC Engineer Cares

Do not hard-code the fast read. It is the right default and the wrong universal. At a divisor the plain read can meet it is strictly worse, and on short transfers near the crossover it is worse still.

Reset to the plain read at the slowest divisor. It costs a few slow transactions during boot and removes the window in which an un-initialised controller issues a fast read with a dummy count it has not been told.

The dummy count and the opcode must be set together, atomically. This is Chapter 10.1's atomic profile argument at its sharpest: a fast-read opcode with the plain read's dummy count produces a one-byte shift in every payload, and neither field is individually wrong.

Clock the dummy phase; do not delay through it. The device needs SCLK edges. A controller that implements the dummy phase as a timer produces no data and a very confusing capture.

Publish both estimates. A register showing what the chosen command costs and what the rejected one would have cost turns "is the controller choosing well?" into a register read. It is two words and it prevents an argument.

If the part offers dual or quad reads, the comparison generalises. Four or five candidates, each with its own dummy count and rate ceiling, and the same minimum-over-estimates structure. Do not special-case them into an if-chain.

9. Failure Signature — A Read That Works at Low Speed and Returns Shifted Data at High Speed

Symptom. A flash read returns correct data with the controller configured at 10 MHz. Raised to 50 MHz, every byte comes back shifted by exactly eight bit positions — byte n returns what byte n−1 should have. No errors are reported. Lowering the clock fixes it completely.

What "shifted by exactly eight bits" establishes. A whole-byte shift is a counting error, not an electrical one. Electrical problems corrupt bits; a clean one-byte offset means the master began capturing one byte time too early or too late, so the transfer's structure is wrong rather than its signal integrity.

Plausible mechanisms.

  • A fast-read opcode with no dummy phase. The controller sends 0x0B and begins capturing immediately, so the byte it takes first is the dummy byte and everything shifts by one. This fits a whole-byte shift exactly.
  • A plain-read opcode with a dummy phase, the same error inverted: the controller clocks eight cycles the device did not expect and the first payload byte is discarded.
  • Exceeding the round-trip limit, which also produces a shift and is also rate-dependent — the strongest competing explanation.
  • A wrong divisor so the link is above the plain read's ceiling while still using 0x03.
  • Signal integrity, which is rate-dependent but corrupts rather than shifts.

The discriminating observation. Rate dependence does not separate the dummy-count error from the round-trip error here, because a controller reconfigured for high speed may have switched commands at the same time. So check which opcode is actually on the wire.

That is the decisive test and it takes one capture: look at the first byte of the frame. 0x03 at 50 MHz on a part rated 50 MHz for plain read is marginal-to-over; 0x0B at 50 MHz is comfortable but requires the dummy phase. Then count byte times between the last address bit and the first payload bit. If the count is 0 with 0x0B, the dummy phase is missing and the diagnosis is complete; if it is 8 and the data is still shifted, the problem is the round trip.

Why the shift is exactly a byte matters. A round-trip violation shifts by a fraction of a bit time that manifests as a one-bit or scrambled error, not a clean eight-bit one. A whole-byte shift is almost diagnostic of a phase-count mismatch on its own — and that observation alone should send the investigation to the command table before the scope.

The fix. Set the opcode and its dummy count from one profile field, atomically, per Chapter 10.1. The reason this failure is common is that raising the clock and switching to fast read are done in the same change, so a dummy count that was correct for the old command is silently wrong for the new one — and nothing in the controller couples the two.

10. Common Misconceptions

11. Reason It Through

Work this before reading the answer.

A controller has a 120 MHz system clock. The flash permits 50 MHz for 0x03 and 104 MHz for 0x0B. A driver reads 4-byte status structures thousands of times per second, and occasionally reads a 64 KB image.

What divisor and command should each use, and what would a single global setting cost?

First find the legal divisors. With a 120 MHz system clock:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   divisor    SCLK        0x03 legal?     0x0B legal?
      1      120.0 MHz        no              no
      2       60.0 MHz        no              yes
      3       40.0 MHz        yes             yes
      4       30.0 MHz        yes             yes

So 0x03 needs divisor 3 or slower; 0x0B can use divisor 2.

The 64 KB image. Length dominates everything, so compare at each command's best rate:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   0x03 at div 3:  (8 + 24 + 0 + 524288) × 3  =  1 572 960 cycles
   0x0B at div 2:  (8 + 24 + 8 + 524288) × 2  =  1 048 656 cycles

Fast read at divisor 2 — a third faster. For a long transfer the rate is everything and the eight cycles are noise.

The 4-byte status structure. Now the overhead dominates:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   0x03 at div 3:  (8 + 24 + 0 + 32) × 3  =  192 cycles
   0x0B at div 2:  (8 + 24 + 8 + 32) × 2  =  144 cycles

Fast read still wins, by 25 per cent. Note that this is not the crossover case — divisor 2 is well below the plain read's ceiling, so the fast read's rate advantage is a full 1.5× and easily covers eight cycles even on four bytes.

So both should use fast read at divisor 2? For throughput, yes. But look at what a single global setting costs, which is the actual question:

  • Global 0x0B at divisor 2: image 1 048 656, status 144. Optimal for both.
  • Global 0x03 at divisor 3: image 1 572 960 (+50%), status 192 (+33%).

So the global fast-read setting is optimal here and there is nothing to trade. That is worth stating plainly rather than manufacturing a dilemma: when the fast read's rate advantage is large, it wins at every length and the choice collapses.

When would it not? Make the system clock 100 MHz and the plain read's ceiling 50 MHz. Then 0x03 runs at divisor 2 (50 MHz) and 0x0B also at divisor 2 — the same rate — and the fast read's eight cycles are pure loss:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   status, 0x03 at div 2:  (8 + 24 + 0 + 32) × 2  =  128 cycles
   status, 0x0B at div 2:  (8 + 24 + 8 + 32) × 2  =  144 cycles

The plain read wins by 12 per cent, and on the image it loses by only 0.03 per cent. With this clock, 0x03 is the better global choice — the reverse of the first case, from nothing but a different system clock.

The general lesson, and the reason the selector in §6 takes both the divisor and the length. The right command depends on the ratio between the two rate ceilings and the system clock's divisor ladder, and that ratio changes when the system clock changes. A design that hard-codes the command is right for one clock configuration and quietly wrong after a power-mode change — which is Chapter 9.4's warning about system-clock changes arriving in a new form.

12. Understanding Check

13. Summary

A flash offers two read commands returning identical data because they trade clock rate against cycle count.

The plain read 0x03 has no dummy phase and a low rate ceiling, because the device must produce the first byte within half a period of the last address bit. The fast read 0x0B buys that array access time in eight clock cycles and runs perhaps twice as fast.

So the comparison is of times, not cycles:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   time = (8 + 8 × addr_bytes + dummy + 8 × length) × divisor

At a divisor the plain read can meet, it wins — same rate, eight fewer cycles. Well below its ceiling the fast read wins by the rate ratio. Just below its ceiling the two cross, at four bytes in the worked case — and near the crossover the difference is eight cycles in five hundred, so the simpler command should win.

The dummy phase must be clocked, not waited out, and MISO during it is undefined — capturing it shifts the whole payload by a byte, which is a counting error and distinguishable from a round-trip violation by the shift being exactly one byte.

In hardware the selector clamps each command to the divisor it allows — a maximum on the divisor — compares estimated times, and gives ties to the plain read. Reset selects the plain read at the slowest divisor, which is what a boot ROM needs.

For verification, optimality and safety are separate properties, because a selector can choose the right command and then overclock it. And the opcode with its dummy count must be one atomic fact: a fast read with the plain read's dummy count shifts every payload while every individual output looks sane.

14. What Comes Next

Reads are the easy half. Writing is where flash stops resembling memory at all.

Chapter 11.3 — Page Program, Page Boundaries, and Erase takes the asymmetry Chapter 11.1 introduced and makes it concrete: why a program cannot cross a page boundary, what the device actually does when you try — it does not fail, it overwrites what you just wrote — how erase granularity differs from program granularity, and the guard that reports not just that an operation is unsafe but exactly which bytes it would destroy, in all three HDLs.

Continue learning