Skip to content
VLSI Mentor

SPI · Module 12

XIP and Memory-Mapped Controllers

Continuous-read mode, execute-in-place, and the latency a CPU actually sees: why a sequential access pays nothing while a random one pays everything, why removing overhead helps short accesses most, and why a stale read stream returns wrong data silently.

Chapter 12.1 widened the phases measured in bits. Chapter 12.4 halved the phases measured in periods. Neither could remove a phase, and that is the one thing left.

A CPU issues a load from 0x6000_0004. The controller turns it into a flash transaction. Does that cost 14 cycles before data appears, or zero?

It depends entirely on whether the previous access ended at 0x6000_0004. If it did, the flash is still clocking out bytes and there is no command, no address and no dummy phase — just data.

1. What Memory-Mapped Means

A memory-mapped flash controller claims a region of the CPU's address space and serves reads from it out of flash:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   CPU address                controller action
   0x6000_0000 .. 0x60FF_FFFF  translate and fetch from flash
   anything else               not ours

To the CPU this is memory: a load instruction, a few cycles of stall, a value in a register. There is no driver, no transaction API and no software involved at all — which is what makes executing code directly from flash possible, and what "execute in place" names.

The cost is that the stall is not a few cycles. A random access costs the full command, address and dummy latency of Chapter 10.5, and a CPU stalled for 22 cycles on every instruction fetch would run at a small fraction of its clock.

So XIP is not viable on random access alone. It is viable because instruction fetch is overwhelmingly sequential, and sequential access can be made nearly free.

2. Continuous Read — Why Sequential Is Free

A flash read command does not end when the master stops asking. Chapter 7.2 established that the device's internal pointer advances as long as clocks arrive and chip select stays asserted.

Continuous-read mode turns that into a persistent state: the device keeps its pointer between accesses and expects more clocks rather than a new command. So if the controller's next request is for exactly the address the last one ended at, it can simply clock more — no new transaction at all.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   random access:      CMD → ADDR → dummy → data     22 + data cycles
   sequential access:                       data      0 + data cycles

That is the entire mechanism, and it explains the shape of every number in this chapter.

XIP goes one step further and removes the command phase from even a random access. The device is told once — usually by a mode bit set inside a read command — that every subsequent access is that command. So the controller sends an address and nothing else:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   random access, XIP:       ADDR → dummy → data     14 + data cycles

3. The Two Cases, Side by Side

Random and sequential, from the CPU's side

8 cycles
Two rows over eight byte times. The random access row shows a command byte, an address, a dummy interval and then five data bytes. The sequential access row shows eight data bytes starting immediately.sequential: data nowsequential: data nowrandom: data hererandom: data hererandomCMDADDRdumD0D1D2D3D4sequentialD0D1D2D3D4D5D6D7t0t1t2t3t4t5t6t7
Figure 1 — the same amount of data, two access patterns. A random access spends its first three byte times on command, address and dummy; a sequential one is already delivering payload in the first. The data phases are identical — only the start differs.

4. The Numbers, and the Inversion

Take a quad-I/O read — 1-4-4, three address bytes, eight dummy cycles — and a 32-byte cache line:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
                        latency   data    total
   random, no XIP         22       64       86
   random, XIP            14       64       78
   sequential              0       64       64

Sequential saves 22 of 86 cycles — 1.34×. Which is modest, and on its own would not justify any of this.

Now a four-byte access:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
                        latency   data    total
   random, XIP            14        8       22
   sequential              0        8        8

2.75×. And that is the inversion this whole module has been building towards.

Chapter 12.1 found that fixed overhead limits what widening can win, and hurts short transfers most. Here, removing that overhead helps short transfers most — by exactly the same arithmetic, read in the other direction. The fixed cost that was a ceiling on width is a floor that overhead-removal can lift.

Which is why XIP works for a CPU. An instruction fetch is a handful of bytes, sequential, issued constantly — the precise access pattern for which width and DDR do least and overhead-removal does most.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   technique              32-byte line   4-byte fetch
   quad output (12.1)        3.28×          1.50×
   DDR (12.4)                1.79×          1.11×
   sequential (this)         1.34×          2.75×

Read the two columns. Every technique from the earlier chapters is worth less on the short access; this one is worth more. They are complements, not alternatives, and a controller wanting both large-block throughput and small-fetch latency needs all of them.

5. The Controller

A memory-mapped flash controller. A CPU bus interface feeds a window decoder that checks the address against a base and size and produces a flash address. A stream tracker compares that address against where the open read stream has reached. A sequential hit goes straight to the data path with no latency; a miss goes to a transaction builder that issues address and dummy phases first. A stream-close input from any other bus user forces the next access to be treated as random.CPU busa load, and a stallWindow decoderbase and size — a shift,not a compareStream trackerwhere the open read hasreachedStream closeany other bus userinvalidates itSequential hitno command, address ordummyRandom missaddress + dummy, andcommand unless XIPData pathidentical in both casesSerial flashpointer kept betweenaccessesCPU stall0 cycles, or 14, or 22addressflash addresscontinuesdoes notinvalidateafter latencyclocksno stallfull stall12
Figure 2 — a memory-mapped flash controller. The window decoder translates the CPU address; the stream tracker decides whether it continues the open read; and the latency the CPU sees is whichever of the two paths the decision selects. Anything that takes the bus away must close the stream.

6. The Stream Must Close

Everything above depends on the controller's belief about where the device's pointer is. And that belief can be wrong in one specific way.

Anything else that uses the bus invalidates it. If a second device is selected, or software issues a status read, or a DMA writes to the flash, the device's read stream has been interrupted — its chip select was deasserted — and the pointer the controller remembers is no longer where the device is.

If the controller then treats the next sequential-looking access as a stream hit, it clocks the device expecting data from address X and receives whatever the device produces from wherever it actually is. The data is wrong and nothing reports it, because a stream hit performs no command and therefore has nothing that could fail.

So the rule is absolute: any access outside the window, and any other use of the bus, closes the stream. The next in-window access then pays full latency, which is a few cycles lost against silently wrong data.

7. Building the Window Decoder — Three HDLs

The circuit

Circuit. A window decoder, a stream tracker and a latency calculator.

State. The address the open stream has reached, and whether one is open.

Datapath. The window test is (addr − base) >> size_bits == 0, which is a shift because window sizes are powers of two — the same rule as everywhere else in this module. The flash address is the offset itself. The latency is a sum of three terms, two of which are conditionally zero.

Control. One decision per request.

Clock and reset. System clock; asynchronous active-low reset. Reset closes the stream, so the first access after reset is always a miss — which is correct, because nothing is open.

Enables. out_of_window distinguishes "not ours" from "ours and slow", and the two need different handling upstream.

Timing. Registered, one cycle after the request.

Synthesis. A subtractor, a shifter, a comparator and an adder.

Limitations. It decides and reports; it does not issue the transaction. It also assumes the stream advances by exactly line_bytes per access, which is true for a fixed-line-size cache and not for a CPU issuing variable-width loads — a real controller tracks the actual byte count consumed.

The stream-closing behaviour is the design's safety property, and it is worth noting where it lives: in the else branch of the window test, not in a separate guard. An out-of-window access closes the stream as part of not serving it, which means there is no path that both skips the window and leaves the stream open.

Azvya Education Pvt. Ltd.VLSI Mentor
xip_window.sv — translate, decide, and report the latency
// xip_window.sv
//
// Chapter 12.5 -- execute in place, and the latency a CPU actually sees.
//
// A memory-mapped flash controller makes a serial device look like memory:
// the CPU issues a load from an address in a window, and the controller
// turns it into a flash transaction. Whether that costs 14 cycles or 0
// before data appears depends on one thing:
//
//   * a SEQUENTIAL access continues an already-open read stream. The flash
//     is still clocking out bytes from where the last access stopped, so
//     there is no command, no address and no dummy phase -- just data.
//
//   * a RANDOM access must start a new transaction, paying the full
//     command, address and dummy latency before the first byte.
//
// Continuous-read mode is what makes the first case possible: the device
// keeps its internal pointer and expects more clocks rather than a new
// command. XIP goes one step further and removes the COMMAND phase from
// even a random access -- the controller sends an address and nothing else,
// because the device has been told once what command every access is.
//
// The result is the inverse of Chapter 12.1's. There, fixed overhead hurt
// SHORT transfers most. Here, REMOVING that overhead helps short transfers
// most -- which is exactly the instruction-fetch pattern a CPU generates,
// and the reason XIP exists at all rather than being a curiosity.
//
// This block decodes the window, translates the address, decides whether
// the access continues the open stream, and reports the latency and the
// total for each case so the two can be compared.

module xip_window #(
    parameter int ADDR_W   = 32,
    parameter int FLASH_AW = 24,
    parameter int CNT_W    = 16,
    parameter int CMD_BITS = 8
) (
    input  logic                clk,
    input  logic                rst_n,

    input  logic                req,
    input  logic [ADDR_W-1:0]   cpu_addr,

    // The memory-mapped window: base and size, the size as an exponent.
    input  logic [ADDR_W-1:0]   win_base,
    input  logic [4:0]          win_bits,

    // The device profile, as always.
    input  logic [3:0]          cmd_lanes,
    input  logic [3:0]          addr_lanes,
    input  logic [3:0]          data_lanes,
    input  logic [2:0]          addr_bytes,
    input  logic [5:0]          dummy_cycles,
    input  logic [CNT_W-1:0]    line_bytes,     // bytes fetched per access
    input  logic                xip_enabled,    // no command phase at all

    output logic                hit_window,
    output logic                out_of_window,
    output logic [FLASH_AW-1:0] flash_addr,
    output logic                stream_hit,     // continues the open stream
    output logic [CNT_W-1:0]    latency_cycles, // before the first data bit
    output logic [CNT_W-1:0]    line_cycles,    // to clock the line out
    output logic [CNT_W-1:0]    total_cycles,
    output logic                valid           // a decision is published
);

    logic [FLASH_AW-1:0] next_addr;    // where the open stream has reached
    logic                stream_open;

    function automatic logic [2:0] lane_shift(input logic [3:0] n);
        case (n)
            4'd1:    lane_shift = 3'd0;
            4'd2:    lane_shift = 3'd1;
            4'd4:    lane_shift = 3'd2;
            4'd8:    lane_shift = 3'd3;
            default: lane_shift = 3'd0;
        endcase
    endfunction

    // Window decode. The offset is the flash address, and the window size
    // is a power of two so the range test is a shift rather than a compare
    // against a computed limit.
    wire [ADDR_W-1:0] offset  = cpu_addr - win_base;
    wire              in_win  = (cpu_addr >= win_base) &&
                                ((offset >> win_bits) == {ADDR_W{1'b0}});
    wire [FLASH_AW-1:0] f_addr = offset[FLASH_AW-1:0];

    // Does this access continue the stream the device already has open?
    wire seq = stream_open && in_win && (f_addr == next_addr);

    // Latency, in the three cases that matter.
    wire [CNT_W-1:0] c_cmd  = CNT_W'(CMD_BITS) >> lane_shift(cmd_lanes);
    wire [CNT_W-1:0] c_addr = (CNT_W'(addr_bytes) << 3) >> lane_shift(addr_lanes);
    wire [CNT_W-1:0] c_data = (line_bytes << 3) >> lane_shift(data_lanes);

    // A sequential access pays nothing before its data. A random access
    // pays the address and dummy phases, and the command phase too unless
    // XIP has removed it.
    wire [CNT_W-1:0] lat = seq ? {CNT_W{1'b0}}
                               : ((xip_enabled ? {CNT_W{1'b0}} : c_cmd)
                                  + c_addr + CNT_W'(dummy_cycles));

    always_ff @(posedge clk or negedge rst_n) begin
        if (!rst_n) begin
            next_addr      <= {FLASH_AW{1'b0}};
            stream_open    <= 1'b0;
            hit_window     <= 1'b0;
            out_of_window  <= 1'b0;
            flash_addr     <= {FLASH_AW{1'b0}};
            stream_hit     <= 1'b0;
            latency_cycles <= {CNT_W{1'b0}};
            line_cycles    <= {CNT_W{1'b0}};
            total_cycles   <= {CNT_W{1'b0}};
            valid          <= 1'b0;
        end else begin
            valid <= 1'b0;

            if (req) begin
                valid         <= 1'b1;
                hit_window    <= in_win;
                out_of_window <= !in_win;
                flash_addr    <= in_win ? f_addr : {FLASH_AW{1'b0}};
                stream_hit    <= seq;
                latency_cycles <= in_win ? lat : {CNT_W{1'b0}};
                line_cycles    <= in_win ? c_data : {CNT_W{1'b0}};
                total_cycles   <= in_win ? (lat + c_data) : {CNT_W{1'b0}};

                if (in_win) begin
                    // The stream now reaches to the end of this line, so the
                    // next access at exactly that address continues it.
                    next_addr   <= f_addr + FLASH_AW'(line_bytes);
                    stream_open <= 1'b1;
                end else begin
                    // An access outside the window is not served from flash
                    // at all, and it must CLOSE the stream: something else
                    // has used the bus, so the device's pointer can no
                    // longer be assumed. Leaving the stream open here is the
                    // bug that returns data from the wrong address.
                    stream_open <= 1'b0;
                end
            end
        end
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
xip_window_tb.sv — the latency, and the stream's lifetime
// xip_window_tb.sv
//
// The properties under test are about LATENCY and about the stream's
// lifetime. A sequential access must cost nothing before its data; a random
// one must pay the full command, address and dummy phases; XIP must remove
// exactly the command phase; and anything that takes the bus away must
// close the stream, because a stream wrongly believed open returns data
// from the wrong address.

`timescale 1ns/1ps

module xip_window_tb;

    localparam int ADDR_W   = 32;
    localparam int FLASH_AW = 24;
    localparam int CNT_W    = 16;

    logic clk = 1'b0;
    logic rst_n = 1'b0;
    always #5 clk = ~clk;

    logic              req = 1'b0;
    logic [ADDR_W-1:0] cpu_addr = 32'h0;
    logic [ADDR_W-1:0] win_base = 32'h6000_0000;
    logic [4:0]        win_bits = 5'd24;          // a 16 MB window
    logic [3:0]        cmd_lanes = 4'd1;
    logic [3:0]        addr_lanes = 4'd4;
    logic [3:0]        data_lanes = 4'd4;
    logic [2:0]        addr_bytes = 3'd3;
    logic [5:0]        dummy_cycles = 6'd8;
    logic [CNT_W-1:0]  line_bytes = 16'd32;
    logic              xip_enabled = 1'b0;

    logic                hit_window, out_of_window;
    logic [FLASH_AW-1:0] flash_addr;
    logic                stream_hit;
    logic [CNT_W-1:0]    latency_cycles, line_cycles, total_cycles;
    logic                valid;

    int errors = 0;
    int t_miss, t_hit;

    xip_window #(.ADDR_W(ADDR_W), .FLASH_AW(FLASH_AW), .CNT_W(CNT_W),
                 .CMD_BITS(8)) dut (
        .clk(clk), .rst_n(rst_n),
        .req(req), .cpu_addr(cpu_addr),
        .win_base(win_base), .win_bits(win_bits),
        .cmd_lanes(cmd_lanes), .addr_lanes(addr_lanes),
        .data_lanes(data_lanes), .addr_bytes(addr_bytes),
        .dummy_cycles(dummy_cycles), .line_bytes(line_bytes),
        .xip_enabled(xip_enabled),
        .hit_window(hit_window), .out_of_window(out_of_window),
        .flash_addr(flash_addr), .stream_hit(stream_hit),
        .latency_cycles(latency_cycles), .line_cycles(line_cycles),
        .total_cycles(total_cycles), .valid(valid)
    );

    task automatic access(input logic [ADDR_W-1:0] a);
        begin
            @(negedge clk);
            cpu_addr = a;
            req = 1'b1;
            @(negedge clk);
            req = 1'b0;
            @(negedge clk);
        end
    endtask

    initial begin
        repeat (3) @(negedge clk);
        rst_n = 1'b1;
        @(negedge clk);

        // 1. THE FIRST ACCESS is always a miss -- no stream is open yet.
        //    Latency is command + address + dummy: 8 + 6 + 8 = 22 at
        //    1-4-4 without XIP.
        access(32'h6000_0000);
        if (!hit_window || stream_hit) begin
            $display("  FAIL: the first access should hit the window and miss the stream");
            errors++;
        end
        if (latency_cycles !== 16'd22) begin
            $display("  FAIL: first-access latency %0d, expected 22 (8 cmd + 6 addr + 8 dummy)",
                     latency_cycles);
            errors++;
        end
        if (line_cycles !== 16'd64) begin
            $display("  FAIL: a 32-byte line on four lanes is %0d cycles, expected 64",
                     line_cycles);
            errors++;
        end
        t_miss = total_cycles;
        $display("  first access:      latency=%0d line=%0d total=%0d  (stream miss)",
                 latency_cycles, line_cycles, total_cycles);

        // 2. THE NEXT SEQUENTIAL LINE continues the stream, so there is no
        //    command, no address and no dummy phase -- only data.
        access(32'h6000_0020);
        if (!stream_hit) begin
            $display("  FAIL: the sequential access did not hit the stream");
            errors++;
        end
        if (latency_cycles !== 16'd0) begin
            $display("  FAIL: a stream hit had a latency of %0d, expected 0",
                     latency_cycles);
            errors++;
        end
        t_hit = total_cycles;
        $display("  sequential access: latency=%0d line=%0d total=%0d  (stream HIT)",
                 latency_cycles, line_cycles, total_cycles);
        $display("  sequential saves %0d of %0d cycles -- %0d.%02dx on a 32-byte line",
                 t_miss - t_hit, t_miss, t_miss / t_hit,
                 ((t_miss * 100) / t_hit) % 100);

        // 3. And the one after it, to show the stream keeps advancing rather
        //    than matching only once.
        access(32'h6000_0040);
        if (!stream_hit || latency_cycles !== 16'd0) begin
            $display("  FAIL: the stream did not continue advancing");
            errors++;
        end

        // 4. A RANDOM JUMP breaks the stream and pays full latency again.
        access(32'h6010_0000);
        if (stream_hit) begin
            $display("  FAIL: a non-sequential address hit the stream");
            errors++;
        end
        if (latency_cycles !== 16'd22) begin
            $display("  FAIL: a random access had latency %0d, expected 22",
                     latency_cycles);
            errors++;
        end
        if (flash_addr !== 24'h100000) begin
            $display("  FAIL: flash address 0x%06h, expected 0x100000", flash_addr);
            errors++;
        end
        $display("  random jump:       latency=%0d, flash address 0x%06h",
                 latency_cycles, flash_addr);

        // 5. XIP removes the COMMAND phase from a random access -- the
        //    controller sends an address and nothing else.
        @(negedge clk); xip_enabled = 1'b1;
        access(32'h6020_0000);
        if (latency_cycles !== 16'd14) begin
            $display("  FAIL: XIP random latency %0d, expected 14 (6 addr + 8 dummy)",
                     latency_cycles);
            errors++;
        end
        $display("  XIP random:        latency=%0d -- the command phase is gone",
                 latency_cycles);

        // 6. XIP does NOT remove the dummy phase. Confirm by changing only
        //    the dummy count and seeing the latency track it exactly.
        @(negedge clk); dummy_cycles = 6'd0;
        access(32'h6030_0000);
        if (latency_cycles !== 16'd6) begin
            $display("  FAIL: XIP with no dummy phase gave latency %0d, expected 6",
                     latency_cycles);
            errors++;
        end
        @(negedge clk); dummy_cycles = 6'd8;
        $display("  XIP, 0 dummy:      latency=6 -- the dummy phase is a separate cost");

        // 7. THE SMALL-ACCESS RESULT. On a short line the saved latency is a
        //    much larger share, which is why XIP suits instruction fetch.
        @(negedge clk); line_bytes = 16'd4; xip_enabled = 1'b1;
        access(32'h6040_0000);            // a miss: full latency
        t_miss = total_cycles;
        access(32'h6040_0004);            // sequential: no latency
        if (!stream_hit) begin
            $display("  FAIL: a 4-byte sequential access did not hit the stream");
            errors++;
        end
        t_hit = total_cycles;
        $display("  4-byte line:       %0d -> %0d cycles, %0d.%02dx",
                 t_miss, t_hit, t_miss / t_hit, ((t_miss * 100) / t_hit) % 100);
        // The speedup on a short access must EXCEED the 32-byte case, which
        // is the whole point of the chapter.
        if ((t_miss * 100) / t_hit < 250) begin
            $display("  FAIL: a 4-byte sequential access saved less than 2.5x");
            errors++;
        end

        // 8. OUT OF WINDOW. An address below the base or past the size is
        //    not served from flash at all.
        @(negedge clk); line_bytes = 16'd32;
        access(32'h5FFF_FFFF);
        if (hit_window || !out_of_window) begin
            $display("  FAIL: an address below the window base was accepted");
            errors++;
        end
        access(32'h6100_0000);
        if (hit_window || !out_of_window) begin
            $display("  FAIL: an address past the 16 MB window was accepted");
            errors++;
        end
        $display("  out of window:     both below-base and past-size rejected");

        // 9. THE STREAM-CLOSING PROPERTY. An out-of-window access means
        //    something else used the bus, so the device's pointer can no
        //    longer be assumed. The next in-window access at what WOULD
        //    have been the sequential address must MISS.
        access(32'h6050_0000);            // opens a stream, next is 0x50_0020
        access(32'h7000_0000);            // out of window -- closes it
        access(32'h6050_0020);            // would have been sequential
        if (stream_hit) begin
            $display("  FAIL: the stream survived an out-of-window access -- this returns data from the wrong address");
            errors++;
        end
        if (latency_cycles === 16'd0) begin
            $display("  FAIL: the access after a closed stream paid no latency");
            errors++;
        end
        $display("  stream closing:    an out-of-window access closes the stream, so the next access pays full latency");

        // 10. The window boundaries, exactly. The last byte inside is in and
        //     the first byte outside is out -- one apart.
        @(negedge clk); win_bits = 5'd12;      // a 4 KB window
        access(32'h6000_0FFF);
        if (!hit_window) begin
            $display("  FAIL: the last address in a 4 KB window was rejected");
            errors++;
        end
        access(32'h6000_1000);
        if (hit_window) begin
            $display("  FAIL: the first address past a 4 KB window was accepted");
            errors++;
        end
        $display("  window boundary:   0x...0FFF in, 0x...1000 out -- exact");

        // 11. INVARIANT SWEEP. Over a run of sequential accesses, every one
        //     after the first hits the stream and pays zero latency; and
        //     every address translates to its offset from the base.
        @(negedge clk); win_bits = 5'd24; line_bytes = 16'd32;
        access(32'h6000_0000);
        for (int i = 1; i < 64; i++) begin
            access(32'h6000_0000 + ADDR_W'(i * 32));
            if (!stream_hit) begin
                $display("  FAIL: sequential access %0d missed the stream", i);
                errors++;
            end
            if (latency_cycles !== 16'd0) begin
                $display("  FAIL: sequential access %0d paid %0d cycles of latency",
                         i, latency_cycles);
                errors++;
            end
            if (flash_addr !== FLASH_AW'(i * 32)) begin
                $display("  FAIL: access %0d translated to 0x%06h, expected 0x%06h",
                         i, flash_addr, i * 32);
                errors++;
            end
            if (total_cycles !== line_cycles) begin
                $display("  FAIL: a stream hit's total %0d differs from its line %0d",
                         total_cycles, line_cycles);
                errors++;
            end
        end
        $display("  63 sequential accesses: all hit the stream, all zero latency, all translated exactly");

        if (errors == 0)
            $display("PASS: the first access to a window pays the full command, address and dummy latency and every sequential access after it pays none, XIP removes exactly the command phase and leaves the dummy phase, the saving is a far larger share on a short access than on a long one, addresses outside the window are rejected at both boundaries exactly, and an out-of-window access closes the stream so the next access cannot wrongly be treated as sequential");
        else
            $display("FAIL: %0d error(s)", errors);
        $finish;
    end

endmodule

The testbench checks latency values against §4's arithmetic, and then four properties that are about behaviour rather than numbers.

The first access is always a miss, because nothing is open — and reset must produce that state rather than an arbitrary one.

XIP removes exactly the command phase. Latency falls from 22 to 14, and then setting the dummy count to zero gives 6 — which confirms the three terms are independent and that XIP has not accidentally removed the dummy phase too.

The small-access result is asserted as a bound. A four-byte sequential access must save at least 2.5×, which is §4's inversion encoded as a check rather than a claim.

And the stream-closing property is tested directly: open a stream, issue an out-of-window access, then request exactly the address that would have been sequential. It must miss. That test is the one that matters most, because the failure it guards against returns wrong data silently.

A sweep of sixty-three sequential accesses then confirms that every one hits, every one pays zero latency, every one translates to its exact offset, and every hit's total equals its data phase — the last being the arithmetic statement that a hit has no overhead at all.

Azvya Education Pvt. Ltd.VLSI Mentor
xip_window.v — the same decoder in Verilog-2001
// xip_window.v
//
// Chapter 12.5 -- execute in place, and the latency a CPU actually sees,
// in Verilog-2001.
//
// A memory-mapped flash controller makes a serial device look like memory.
// Whether an access costs 14 cycles or 0 before data appears depends on one
// thing:
//
//   * a SEQUENTIAL access continues an already-open read stream, so there is
//     no command, no address and no dummy phase -- just data;
//   * a RANDOM access must start a new transaction and pay all three.
//
// Continuous-read mode makes the first case possible. XIP goes further and
// removes the COMMAND phase from even a random access, because the device
// has been told once what command every access is.
//
// The result is the inverse of Chapter 12.1's. There, fixed overhead hurt
// SHORT transfers most. Here, REMOVING that overhead helps short transfers
// most -- which is the instruction-fetch pattern a CPU generates, and the
// reason XIP exists rather than being a curiosity.

module xip_window #(
    parameter ADDR_W   = 32,
    parameter FLASH_AW = 24,
    parameter CNT_W    = 16,
    parameter CMD_BITS = 8
) (
    input  wire                clk,
    input  wire                rst_n,

    input  wire                req,
    input  wire [ADDR_W-1:0]   cpu_addr,

    // The memory-mapped window: base and size, the size as an exponent.
    input  wire [ADDR_W-1:0]   win_base,
    input  wire [4:0]          win_bits,

    // The device profile, as always.
    input  wire [3:0]          cmd_lanes,
    input  wire [3:0]          addr_lanes,
    input  wire [3:0]          data_lanes,
    input  wire [2:0]          addr_bytes,
    input  wire [5:0]          dummy_cycles,
    input  wire [CNT_W-1:0]    line_bytes,     // bytes fetched per access
    input  wire                xip_enabled,    // no command phase at all

    output reg                 hit_window,
    output reg                 out_of_window,
    output reg  [FLASH_AW-1:0] flash_addr,
    output reg                 stream_hit,     // continues the open stream
    output reg  [CNT_W-1:0]    latency_cycles, // before the first data bit
    output reg  [CNT_W-1:0]    line_cycles,    // to clock the line out
    output reg  [CNT_W-1:0]    total_cycles,
    output reg                 valid           // a decision is published
);

    reg [FLASH_AW-1:0] next_addr;    // where the open stream has reached
    reg                stream_open;

    function [2:0] lane_shift;
        input [3:0] n;
        begin
            case (n)
                4'd1:    lane_shift = 3'd0;
                4'd2:    lane_shift = 3'd1;
                4'd4:    lane_shift = 3'd2;
                4'd8:    lane_shift = 3'd3;
                default: lane_shift = 3'd0;
            endcase
        end
    endfunction

    // Window decode. The offset is the flash address, and the window size is
    // a power of two so the range test is a shift rather than a compare
    // against a computed limit.
    wire [ADDR_W-1:0]   offset = cpu_addr - win_base;
    wire                in_win = (cpu_addr >= win_base) &&
                                 ((offset >> win_bits) == {ADDR_W{1'b0}});
    wire [FLASH_AW-1:0] f_addr = offset[FLASH_AW-1:0];

    // Does this access continue the stream the device already has open?
    wire seq = stream_open && in_win && (f_addr == next_addr);

    // Latency, in the three cases that matter.
    wire [CNT_W-1:0] c_cmd  = CMD_BITS >> lane_shift(cmd_lanes);
    wire [CNT_W-1:0] c_addr = ({{(CNT_W-3){1'b0}}, addr_bytes} << 3)
                              >> lane_shift(addr_lanes);
    wire [CNT_W-1:0] c_data = (line_bytes << 3) >> lane_shift(data_lanes);

    // A sequential access pays nothing before its data. A random access pays
    // the address and dummy phases, and the command phase too unless XIP has
    // removed it.
    wire [CNT_W-1:0] lat = seq ? {CNT_W{1'b0}}
                               : ((xip_enabled ? {CNT_W{1'b0}} : c_cmd)
                                  + c_addr + dummy_cycles);

    always @(posedge clk or negedge rst_n) begin
        if (!rst_n) begin
            next_addr      <= {FLASH_AW{1'b0}};
            stream_open    <= 1'b0;
            hit_window     <= 1'b0;
            out_of_window  <= 1'b0;
            flash_addr     <= {FLASH_AW{1'b0}};
            stream_hit     <= 1'b0;
            latency_cycles <= {CNT_W{1'b0}};
            line_cycles    <= {CNT_W{1'b0}};
            total_cycles   <= {CNT_W{1'b0}};
            valid          <= 1'b0;
        end else begin
            valid <= 1'b0;

            if (req) begin
                valid          <= 1'b1;
                hit_window     <= in_win;
                out_of_window  <= !in_win;
                flash_addr     <= in_win ? f_addr : {FLASH_AW{1'b0}};
                stream_hit     <= seq;
                latency_cycles <= in_win ? lat : {CNT_W{1'b0}};
                line_cycles    <= in_win ? c_data : {CNT_W{1'b0}};
                total_cycles   <= in_win ? (lat + c_data) : {CNT_W{1'b0}};

                if (in_win) begin
                    // The stream now reaches to the end of this line, so the
                    // next access at exactly that address continues it.
                    // Zero-extended explicitly: line_bytes is CNT_W wide
                    // and the address is FLASH_AW wide, so a part-select of
                    // FLASH_AW bits from it would read past its end and
                    // yield X -- which propagates into the stream
                    // comparison and makes every later access undecidable.
                    next_addr   <= f_addr +
                                   {{(FLASH_AW-CNT_W){1'b0}}, line_bytes};
                    stream_open <= 1'b1;
                end else begin
                    // An access outside the window is not served from flash,
                    // and it must CLOSE the stream: something else has used
                    // the bus, so the device's pointer can no longer be
                    // assumed. Leaving it open is the bug that returns data
                    // from the wrong address.
                    stream_open <= 1'b0;
                end
            end
        end
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
xip_window_tb.v — the same stream-lifetime checks in Verilog-2001
// xip_window_tb.v
//
// The same checks as the SystemVerilog testbench. The properties are about
// LATENCY and about the stream's lifetime: a sequential access costs nothing
// before its data, a random one pays all three phases, XIP removes exactly
// the command phase, and anything that takes the bus away closes the stream.

`timescale 1ns/1ps

module xip_window_tb;

    parameter ADDR_W   = 32;
    parameter FLASH_AW = 24;
    parameter CNT_W    = 16;

    reg clk;
    reg rst_n;

    reg              req;
    reg [ADDR_W-1:0] cpu_addr;
    reg [ADDR_W-1:0] win_base;
    reg [4:0]        win_bits;
    reg [3:0]        cmd_lanes;
    reg [3:0]        addr_lanes;
    reg [3:0]        data_lanes;
    reg [2:0]        addr_bytes;
    reg [5:0]        dummy_cycles;
    reg [CNT_W-1:0]  line_bytes;
    reg              xip_enabled;

    wire                hit_window, out_of_window;
    wire [FLASH_AW-1:0] flash_addr;
    wire                stream_hit;
    wire [CNT_W-1:0]    latency_cycles, line_cycles, total_cycles;
    wire                valid;

    integer errors;
    integer t_miss, t_hit, i;

    initial begin
        clk = 1'b0; rst_n = 1'b0; req = 1'b0; cpu_addr = 32'h0;
        win_base = 32'h6000_0000; win_bits = 5'd24;
        cmd_lanes = 4'd1; addr_lanes = 4'd4; data_lanes = 4'd4;
        addr_bytes = 3'd3; dummy_cycles = 6'd8;
        line_bytes = 16'd32; xip_enabled = 1'b0;
        errors = 0;
    end
    always #5 clk = ~clk;

    xip_window #(.ADDR_W(ADDR_W), .FLASH_AW(FLASH_AW), .CNT_W(CNT_W),
                 .CMD_BITS(8)) dut (
        .clk(clk), .rst_n(rst_n),
        .req(req), .cpu_addr(cpu_addr),
        .win_base(win_base), .win_bits(win_bits),
        .cmd_lanes(cmd_lanes), .addr_lanes(addr_lanes),
        .data_lanes(data_lanes), .addr_bytes(addr_bytes),
        .dummy_cycles(dummy_cycles), .line_bytes(line_bytes),
        .xip_enabled(xip_enabled),
        .hit_window(hit_window), .out_of_window(out_of_window),
        .flash_addr(flash_addr), .stream_hit(stream_hit),
        .latency_cycles(latency_cycles), .line_cycles(line_cycles),
        .total_cycles(total_cycles), .valid(valid)
    );

    task access;
        input [ADDR_W-1:0] a;
        begin
            @(negedge clk);
            cpu_addr = a;
            req = 1'b1;
            @(negedge clk);
            req = 1'b0;
            @(negedge clk);
        end
    endtask

    initial begin
        repeat (3) @(negedge clk);
        rst_n = 1'b1;
        @(negedge clk);

        // 1. THE FIRST ACCESS is always a miss -- no stream is open yet.
        access(32'h6000_0000);
        if (!hit_window || stream_hit) begin
            $display("  FAIL: the first access should hit the window and miss the stream");
            errors = errors + 1;
        end
        if (latency_cycles !== 16'd22) begin
            $display("  FAIL: first-access latency %0d, expected 22 (8 cmd + 6 addr + 8 dummy)",
                     latency_cycles);
            errors = errors + 1;
        end
        if (line_cycles !== 16'd64) begin
            $display("  FAIL: a 32-byte line on four lanes is %0d cycles, expected 64",
                     line_cycles);
            errors = errors + 1;
        end
        t_miss = total_cycles;
        $display("  first access:      latency=%0d line=%0d total=%0d  (stream miss)",
                 latency_cycles, line_cycles, total_cycles);

        // 2. THE NEXT SEQUENTIAL LINE continues the stream.
        access(32'h6000_0020);
        if (!stream_hit) begin
            $display("  FAIL: the sequential access did not hit the stream");
            errors = errors + 1;
        end
        if (latency_cycles !== 16'd0) begin
            $display("  FAIL: a stream hit had a latency of %0d, expected 0",
                     latency_cycles);
            errors = errors + 1;
        end
        t_hit = total_cycles;
        $display("  sequential access: latency=%0d line=%0d total=%0d  (stream HIT)",
                 latency_cycles, line_cycles, total_cycles);
        $display("  sequential saves %0d of %0d cycles -- %0d.%02dx on a 32-byte line",
                 t_miss - t_hit, t_miss, t_miss / t_hit,
                 ((t_miss * 100) / t_hit) % 100);

        // 3. And the one after it, to show the stream keeps advancing.
        access(32'h6000_0040);
        if (!stream_hit || latency_cycles !== 16'd0) begin
            $display("  FAIL: the stream did not continue advancing");
            errors = errors + 1;
        end

        // 4. A RANDOM JUMP breaks the stream and pays full latency again.
        access(32'h6010_0000);
        if (stream_hit) begin
            $display("  FAIL: a non-sequential address hit the stream");
            errors = errors + 1;
        end
        if (latency_cycles !== 16'd22) begin
            $display("  FAIL: a random access had latency %0d, expected 22",
                     latency_cycles);
            errors = errors + 1;
        end
        if (flash_addr !== 24'h100000) begin
            $display("  FAIL: flash address 0x%06h, expected 0x100000", flash_addr);
            errors = errors + 1;
        end
        $display("  random jump:       latency=%0d, flash address 0x%06h",
                 latency_cycles, flash_addr);

        // 5. XIP removes the COMMAND phase from a random access.
        @(negedge clk); xip_enabled = 1'b1;
        access(32'h6020_0000);
        if (latency_cycles !== 16'd14) begin
            $display("  FAIL: XIP random latency %0d, expected 14 (6 addr + 8 dummy)",
                     latency_cycles);
            errors = errors + 1;
        end
        $display("  XIP random:        latency=%0d -- the command phase is gone",
                 latency_cycles);

        // 6. XIP does NOT remove the dummy phase.
        @(negedge clk); dummy_cycles = 6'd0;
        access(32'h6030_0000);
        if (latency_cycles !== 16'd6) begin
            $display("  FAIL: XIP with no dummy phase gave latency %0d, expected 6",
                     latency_cycles);
            errors = errors + 1;
        end
        @(negedge clk); dummy_cycles = 6'd8;
        $display("  XIP, 0 dummy:      latency=6 -- the dummy phase is a separate cost");

        // 7. THE SMALL-ACCESS RESULT.
        @(negedge clk); line_bytes = 16'd4; xip_enabled = 1'b1;
        access(32'h6040_0000);            // a miss: full latency
        t_miss = total_cycles;
        access(32'h6040_0004);            // sequential: no latency
        if (!stream_hit) begin
            $display("  FAIL: a 4-byte sequential access did not hit the stream");
            errors = errors + 1;
        end
        t_hit = total_cycles;
        $display("  4-byte line:       %0d -> %0d cycles, %0d.%02dx",
                 t_miss, t_hit, t_miss / t_hit, ((t_miss * 100) / t_hit) % 100);
        if ((t_miss * 100) / t_hit < 250) begin
            $display("  FAIL: a 4-byte sequential access saved less than 2.5x");
            errors = errors + 1;
        end

        // 8. OUT OF WINDOW at both ends.
        @(negedge clk); line_bytes = 16'd32;
        access(32'h5FFF_FFFF);
        if (hit_window || !out_of_window) begin
            $display("  FAIL: an address below the window base was accepted");
            errors = errors + 1;
        end
        access(32'h6100_0000);
        if (hit_window || !out_of_window) begin
            $display("  FAIL: an address past the 16 MB window was accepted");
            errors = errors + 1;
        end
        $display("  out of window:     both below-base and past-size rejected");

        // 9. THE STREAM-CLOSING PROPERTY.
        access(32'h6050_0000);            // opens a stream, next is 0x50_0020
        access(32'h7000_0000);            // out of window -- closes it
        access(32'h6050_0020);            // would have been sequential
        if (stream_hit) begin
            $display("  FAIL: the stream survived an out-of-window access");
            errors = errors + 1;
        end
        if (latency_cycles === 16'd0) begin
            $display("  FAIL: the access after a closed stream paid no latency");
            errors = errors + 1;
        end
        $display("  stream closing:    an out-of-window access closes the stream, so the next access pays full latency");

        // 10. The window boundaries, exactly.
        @(negedge clk); win_bits = 5'd12;      // a 4 KB window
        access(32'h6000_0FFF);
        if (!hit_window) begin
            $display("  FAIL: the last address in a 4 KB window was rejected");
            errors = errors + 1;
        end
        access(32'h6000_1000);
        if (hit_window) begin
            $display("  FAIL: the first address past a 4 KB window was accepted");
            errors = errors + 1;
        end
        $display("  window boundary:   0x...0FFF in, 0x...1000 out -- exact");

        // 11. INVARIANT SWEEP over a run of sequential accesses.
        @(negedge clk); win_bits = 5'd24; line_bytes = 16'd32;
        access(32'h6000_0000);
        for (i = 1; i < 64; i = i + 1) begin
            access(32'h6000_0000 + (i * 32));
            if (!stream_hit) begin
                $display("  FAIL: sequential access %0d missed the stream", i);
                errors = errors + 1;
            end
            if (latency_cycles !== 16'd0) begin
                $display("  FAIL: sequential access %0d paid %0d cycles of latency",
                         i, latency_cycles);
                errors = errors + 1;
            end
            if (flash_addr !== (i * 32)) begin
                $display("  FAIL: access %0d translated to 0x%06h, expected 0x%06h",
                         i, flash_addr, i * 32);
                errors = errors + 1;
            end
            if (total_cycles !== line_cycles) begin
                $display("  FAIL: a stream hit's total %0d differs from its line %0d",
                         total_cycles, line_cycles);
                errors = errors + 1;
            end
        end
        $display("  63 sequential accesses: all hit the stream, all zero latency, all translated exactly");

        if (errors == 0)
            $display("PASS: the first access to a window pays the full command, address and dummy latency and every sequential access after it pays none, XIP removes exactly the command phase and leaves the dummy phase, the saving is a far larger share on a short access than on a long one, addresses outside the window are rejected at both boundaries exactly, and an out-of-window access closes the stream so the next access cannot wrongly be treated as sequential");
        else
            $display("FAIL: %0d error(s)", errors);
        $finish;
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
xip_window.vhd — the same decoder in VHDL
-- xip_window.vhd
--
-- Chapter 12.5 -- execute in place, and the latency a CPU actually sees,
-- in VHDL.
--
-- A memory-mapped flash controller makes a serial device look like memory.
-- Whether an access costs 14 cycles or 0 before data appears depends on one
-- thing:
--
--   * a SEQUENTIAL access continues an already-open read stream, so there is
--     no command, no address and no dummy phase -- just data;
--   * a RANDOM access must start a new transaction and pay all three.
--
-- Continuous-read mode makes the first case possible. XIP goes further and
-- removes the COMMAND phase from even a random access, because the device
-- has been told once what command every access is.
--
-- The result is the inverse of Chapter 12.1's. There, fixed overhead hurt
-- SHORT transfers most. Here, REMOVING that overhead helps short transfers
-- most -- which is the instruction-fetch pattern a CPU generates.

library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

entity xip_window is
    generic (
        ADDR_W   : positive := 32;
        FLASH_AW : positive := 24;
        CNT_W    : positive := 16;
        CMD_BITS : natural  := 8
    );
    port (
        clk            : in  std_logic;
        rst_n          : in  std_logic;

        req            : in  std_logic;
        cpu_addr       : in  unsigned(ADDR_W - 1 downto 0);

        -- The memory-mapped window: base and size, the size as an exponent.
        win_base       : in  unsigned(ADDR_W - 1 downto 0);
        win_bits       : in  unsigned(4 downto 0);

        -- The device profile, as always.
        cmd_lanes      : in  unsigned(3 downto 0);
        addr_lanes     : in  unsigned(3 downto 0);
        data_lanes     : in  unsigned(3 downto 0);
        addr_bytes     : in  unsigned(2 downto 0);
        dummy_cycles   : in  unsigned(5 downto 0);
        line_bytes     : in  unsigned(CNT_W - 1 downto 0);
        xip_enabled    : in  std_logic;

        hit_window     : out std_logic;
        out_of_window  : out std_logic;
        flash_addr     : out unsigned(FLASH_AW - 1 downto 0);
        stream_hit     : out std_logic;
        latency_cycles : out unsigned(CNT_W - 1 downto 0);
        line_cycles    : out unsigned(CNT_W - 1 downto 0);
        total_cycles   : out unsigned(CNT_W - 1 downto 0);
        valid          : out std_logic
    );
end entity;

architecture rtl of xip_window is

    function lane_div(n : unsigned(3 downto 0)) return natural is
    begin
        case to_integer(n) is
            when 1      => return 1;
            when 2      => return 2;
            when 4      => return 4;
            when 8      => return 8;
            when others => return 1;
        end case;
    end function;

    signal next_addr   : unsigned(FLASH_AW - 1 downto 0) := (others => '0');
    signal stream_open : std_logic := '0';

    signal hw_r  : std_logic := '0';
    signal oow_r : std_logic := '0';
    signal fa_r  : unsigned(FLASH_AW - 1 downto 0) := (others => '0');
    signal sh_r  : std_logic := '0';
    signal lat_r : unsigned(CNT_W - 1 downto 0) := (others => '0');
    signal lc_r  : unsigned(CNT_W - 1 downto 0) := (others => '0');
    signal tc_r  : unsigned(CNT_W - 1 downto 0) := (others => '0');
    signal v_r   : std_logic := '0';

begin

    decode : process (clk, rst_n)
        variable offs   : unsigned(ADDR_W - 1 downto 0);
        variable in_win : boolean;
        variable f_addr : unsigned(FLASH_AW - 1 downto 0);
        variable seq    : boolean;
        variable c_cmd, c_addr, c_data, lat : natural;
    begin
        if rst_n = '0' then
            next_addr   <= (others => '0');
            stream_open <= '0';
            hw_r        <= '0';
            oow_r       <= '0';
            fa_r        <= (others => '0');
            sh_r        <= '0';
            lat_r       <= (others => '0');
            lc_r        <= (others => '0');
            tc_r        <= (others => '0');
            v_r         <= '0';
        elsif rising_edge(clk) then
            v_r <= '0';

            if req = '1' then
                -- Window decode. The offset is the flash address, and the
                -- window size is a power of two so the range test is a shift
                -- rather than a compare against a computed limit.
                offs   := cpu_addr - win_base;
                in_win := (cpu_addr >= win_base) and
                          (shift_right(offs, to_integer(win_bits)) = 0);
                f_addr := resize(offs, FLASH_AW);

                -- Does this access continue the stream already open?
                seq := (stream_open = '1') and in_win and (f_addr = next_addr);

                c_cmd  := CMD_BITS / lane_div(cmd_lanes);
                c_addr := (to_integer(addr_bytes) * 8) / lane_div(addr_lanes);
                c_data := (to_integer(line_bytes) * 8) / lane_div(data_lanes);

                -- A sequential access pays nothing before its data. A random
                -- access pays the address and dummy phases, and the command
                -- phase too unless XIP has removed it.
                if seq then
                    lat := 0;
                elsif xip_enabled = '1' then
                    lat := c_addr + to_integer(dummy_cycles);
                else
                    lat := c_cmd + c_addr + to_integer(dummy_cycles);
                end if;

                v_r <= '1';
                if in_win then
                    hw_r  <= '1';
                    oow_r <= '0';
                    fa_r  <= f_addr;
                    lat_r <= to_unsigned(lat, CNT_W);
                    lc_r  <= to_unsigned(c_data, CNT_W);
                    tc_r  <= to_unsigned(lat + c_data, CNT_W);

                    -- The stream now reaches to the end of this line, so the
                    -- next access at exactly that address continues it.
                    next_addr   <= f_addr + resize(line_bytes, FLASH_AW);
                    stream_open <= '1';
                else
                    hw_r  <= '0';
                    oow_r <= '1';
                    fa_r  <= (others => '0');
                    lat_r <= (others => '0');
                    lc_r  <= (others => '0');
                    tc_r  <= (others => '0');

                    -- An access outside the window is not served from flash,
                    -- and it must CLOSE the stream: something else has used
                    -- the bus, so the device's pointer can no longer be
                    -- assumed. Leaving it open is the bug that returns data
                    -- from the wrong address.
                    stream_open <= '0';
                end if;

                if seq then sh_r <= '1'; else sh_r <= '0'; end if;
            end if;
        end if;
    end process;

    hit_window     <= hw_r;
    out_of_window  <= oow_r;
    flash_addr     <= fa_r;
    stream_hit     <= sh_r;
    latency_cycles <= lat_r;
    line_cycles    <= lc_r;
    total_cycles   <= tc_r;
    valid          <= v_r;

end architecture;
Azvya Education Pvt. Ltd.VLSI Mentor
xip_window_tb.vhd — the same stream-lifetime checks in VHDL
-- xip_window_tb.vhd
--
-- The same checks as the SystemVerilog and Verilog testbenches. The
-- properties are about LATENCY and about the stream's lifetime: a sequential
-- access costs nothing before its data, a random one pays all three phases,
-- XIP removes exactly the command phase, and anything that takes the bus
-- away closes the stream.

library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

entity xip_window_tb is
end entity;

architecture sim of xip_window_tb is

    constant ADDR_W   : positive := 32;
    constant FLASH_AW : positive := 24;
    constant CNT_W    : positive := 16;

    signal clk   : std_logic := '0';
    signal rst_n : std_logic := '0';
    signal halt  : boolean   := false;

    signal req          : std_logic := '0';
    signal cpu_addr     : unsigned(ADDR_W - 1 downto 0) := (others => '0');
    signal win_base     : unsigned(ADDR_W - 1 downto 0) :=
        to_unsigned(16#60000000#, ADDR_W);
    signal win_bits     : unsigned(4 downto 0) := to_unsigned(24, 5);
    signal cmd_lanes    : unsigned(3 downto 0) := to_unsigned(1, 4);
    signal addr_lanes   : unsigned(3 downto 0) := to_unsigned(4, 4);
    signal data_lanes   : unsigned(3 downto 0) := to_unsigned(4, 4);
    signal addr_bytes   : unsigned(2 downto 0) := to_unsigned(3, 3);
    signal dummy_cycles : unsigned(5 downto 0) := to_unsigned(8, 6);
    signal line_bytes   : unsigned(CNT_W - 1 downto 0) := to_unsigned(32, CNT_W);
    signal xip_enabled  : std_logic := '0';

    signal hit_window     : std_logic;
    signal out_of_window  : std_logic;
    signal flash_addr     : unsigned(FLASH_AW - 1 downto 0);
    signal stream_hit     : std_logic;
    signal latency_cycles : unsigned(CNT_W - 1 downto 0);
    signal line_cycles    : unsigned(CNT_W - 1 downto 0);
    signal total_cycles   : unsigned(CNT_W - 1 downto 0);
    signal valid          : std_logic;

    signal errors : natural := 0;

begin

    clk <= not clk after 5 ns when not halt else '0';

    dut : entity work.xip_window
        generic map (ADDR_W => ADDR_W, FLASH_AW => FLASH_AW, CNT_W => CNT_W,
                     CMD_BITS => 8)
        port map (
            clk => clk, rst_n => rst_n,
            req => req, cpu_addr => cpu_addr,
            win_base => win_base, win_bits => win_bits,
            cmd_lanes => cmd_lanes, addr_lanes => addr_lanes,
            data_lanes => data_lanes, addr_bytes => addr_bytes,
            dummy_cycles => dummy_cycles, line_bytes => line_bytes,
            xip_enabled => xip_enabled,
            hit_window => hit_window, out_of_window => out_of_window,
            flash_addr => flash_addr, stream_hit => stream_hit,
            latency_cycles => latency_cycles, line_cycles => line_cycles,
            total_cycles => total_cycles, valid => valid
        );

    stim : process
        variable errs : natural := 0;
        variable t_miss, t_hit : natural;

        -- Not named `access`: that is a VHDL reserved word for access
        -- types, so it cannot be a subprogram name.
        procedure do_access(a : natural) is
        begin
            wait until falling_edge(clk);
            cpu_addr <= to_unsigned(a, ADDR_W);
            req      <= '1';
            wait until falling_edge(clk);
            req      <= '0';
            wait until falling_edge(clk);
        end procedure;
    begin
        for k in 0 to 2 loop
            wait until falling_edge(clk);
        end loop;
        rst_n <= '1';
        wait until falling_edge(clk);

        -- 1. THE FIRST ACCESS is always a miss -- no stream is open yet.
        do_access(16#60000000#);
        if hit_window /= '1' or stream_hit = '1' then
            report "  FAIL: the first access should hit the window and miss the stream";
            errs := errs + 1;
        end if;
        if to_integer(latency_cycles) /= 22 then
            report "  FAIL: first-access latency is not 22 (8 cmd + 6 addr + 8 dummy)";
            errs := errs + 1;
        end if;
        if to_integer(line_cycles) /= 64 then
            report "  FAIL: a 32-byte line on four lanes is not 64 cycles";
            errs := errs + 1;
        end if;
        t_miss := to_integer(total_cycles);
        report "  first access:      latency=" &
               integer'image(to_integer(latency_cycles)) & " line=" &
               integer'image(to_integer(line_cycles)) & " total=" &
               integer'image(t_miss) & "  (stream miss)";

        -- 2. THE NEXT SEQUENTIAL LINE continues the stream.
        do_access(16#60000020#);
        if stream_hit /= '1' then
            report "  FAIL: the sequential access did not hit the stream";
            errs := errs + 1;
        end if;
        if latency_cycles /= 0 then
            report "  FAIL: a stream hit had a non-zero latency";
            errs := errs + 1;
        end if;
        t_hit := to_integer(total_cycles);
        report "  sequential access: latency=0 line=" &
               integer'image(to_integer(line_cycles)) & " total=" &
               integer'image(t_hit) & "  (stream HIT)";
        report "  sequential saves " & integer'image(t_miss - t_hit) &
               " of " & integer'image(t_miss) & " cycles -- x100 = " &
               integer'image((t_miss * 100) / t_hit) & " on a 32-byte line";

        -- 3. And the one after it, to show the stream keeps advancing.
        do_access(16#60000040#);
        if stream_hit /= '1' or latency_cycles /= 0 then
            report "  FAIL: the stream did not continue advancing";
            errs := errs + 1;
        end if;

        -- 4. A RANDOM JUMP breaks the stream and pays full latency again.
        do_access(16#60100000#);
        if stream_hit = '1' then
            report "  FAIL: a non-sequential address hit the stream";
            errs := errs + 1;
        end if;
        if to_integer(latency_cycles) /= 22 then
            report "  FAIL: a random access did not pay full latency";
            errs := errs + 1;
        end if;
        if to_integer(flash_addr) /= 16#100000# then
            report "  FAIL: the flash address translation is wrong";
            errs := errs + 1;
        end if;
        report "  random jump:       latency=22, flash address 0x100000";

        -- 5. XIP removes the COMMAND phase from a random access.
        wait until falling_edge(clk);
        xip_enabled <= '1';
        do_access(16#60200000#);
        if to_integer(latency_cycles) /= 14 then
            report "  FAIL: XIP random latency is not 14 (6 addr + 8 dummy)";
            errs := errs + 1;
        end if;
        report "  XIP random:        latency=14 -- the command phase is gone";

        -- 6. XIP does NOT remove the dummy phase.
        wait until falling_edge(clk);
        dummy_cycles <= to_unsigned(0, 6);
        do_access(16#60300000#);
        if to_integer(latency_cycles) /= 6 then
            report "  FAIL: XIP with no dummy phase did not give latency 6";
            errs := errs + 1;
        end if;
        wait until falling_edge(clk);
        dummy_cycles <= to_unsigned(8, 6);
        report "  XIP, 0 dummy:      latency=6 -- the dummy phase is a separate cost";

        -- 7. THE SMALL-ACCESS RESULT.
        wait until falling_edge(clk);
        line_bytes  <= to_unsigned(4, CNT_W);
        xip_enabled <= '1';
        do_access(16#60400000#);            -- a miss: full latency
        t_miss := to_integer(total_cycles);
        do_access(16#60400004#);            -- sequential: no latency
        if stream_hit /= '1' then
            report "  FAIL: a 4-byte sequential access did not hit the stream";
            errs := errs + 1;
        end if;
        t_hit := to_integer(total_cycles);
        report "  4-byte line:       " & integer'image(t_miss) & " -> " &
               integer'image(t_hit) & " cycles, x100 = " &
               integer'image((t_miss * 100) / t_hit);
        if (t_miss * 100) / t_hit < 250 then
            report "  FAIL: a 4-byte sequential access saved less than 2.5x";
            errs := errs + 1;
        end if;

        -- 8. OUT OF WINDOW at both ends.
        wait until falling_edge(clk);
        line_bytes <= to_unsigned(32, CNT_W);
        do_access(16#5FFFFFFF#);
        if hit_window = '1' or out_of_window /= '1' then
            report "  FAIL: an address below the window base was accepted";
            errs := errs + 1;
        end if;
        do_access(16#61000000#);
        if hit_window = '1' or out_of_window /= '1' then
            report "  FAIL: an address past the 16 MB window was accepted";
            errs := errs + 1;
        end if;
        report "  out of window:     both below-base and past-size rejected";

        -- 9. THE STREAM-CLOSING PROPERTY.
        do_access(16#60500000#);            -- opens a stream
        do_access(16#70000000#);            -- out of window -- closes it
        do_access(16#60500020#);            -- would have been sequential
        if stream_hit = '1' then
            report "  FAIL: the stream survived an out-of-window access";
            errs := errs + 1;
        end if;
        if latency_cycles = 0 then
            report "  FAIL: the access after a closed stream paid no latency";
            errs := errs + 1;
        end if;
        report "  stream closing:    an out-of-window access closes the stream, so the next access pays full latency";

        -- 10. The window boundaries, exactly.
        wait until falling_edge(clk);
        win_bits <= to_unsigned(12, 5);      -- a 4 KB window
        do_access(16#60000FFF#);
        if hit_window /= '1' then
            report "  FAIL: the last address in a 4 KB window was rejected";
            errs := errs + 1;
        end if;
        do_access(16#60001000#);
        if hit_window = '1' then
            report "  FAIL: the first address past a 4 KB window was accepted";
            errs := errs + 1;
        end if;
        report "  window boundary:   0x...0FFF in, 0x...1000 out -- exact";

        -- 11. INVARIANT SWEEP over a run of sequential accesses.
        wait until falling_edge(clk);
        win_bits   <= to_unsigned(24, 5);
        line_bytes <= to_unsigned(32, CNT_W);
        do_access(16#60000000#);
        for i in 1 to 63 loop
            do_access(16#60000000# + i * 32);
            if stream_hit /= '1' then
                report "  FAIL: a sequential access missed the stream";
                errs := errs + 1;
            end if;
            if latency_cycles /= 0 then
                report "  FAIL: a sequential access paid latency";
                errs := errs + 1;
            end if;
            if to_integer(flash_addr) /= i * 32 then
                report "  FAIL: a sequential access translated wrongly";
                errs := errs + 1;
            end if;
            if total_cycles /= line_cycles then
                report "  FAIL: a stream hit's total differs from its line";
                errs := errs + 1;
            end if;
        end loop;
        report "  63 sequential accesses: all hit the stream, all zero latency, all translated exactly";

        errors <= errs;
        if errs = 0 then
            report "PASS: the first access to a window pays the full command, address and dummy latency and every sequential access after it pays none, XIP removes exactly the command phase and leaves the dummy phase, the saving is a far larger share on a short access than on a long one, addresses outside the window are rejected at both boundaries exactly, and an out-of-window access closes the stream so the next access cannot wrongly be treated as sequential";
        else
            report "FAIL: " & integer'image(errs) & " error(s)" severity error;
        end if;
        halt <= true;
        wait;
    end process;

end architecture;

Parity

All three implement the same decoder: identical ports and generics, a window test by shift, a stream tracker advancing by the line size, latency as three independently-zeroable terms, and an out-of-window access closing the stream as part of not serving it. All three testbenches report identical figures — 22 cycles of latency without XIP, 14 with it, 6 with XIP and no dummy phase, 1.34× on a 32-byte line and 2.75× on four bytes — and all three confirm the window boundary is exact at 0x...0FFF and 0x...1000.

Two naming notes, both language constraints rather than choices. In Verilog, line_bytes[FLASH_AW-1:0] selects 24 bits from a 16-bit value and yields X, which propagates into the stream comparison and makes every subsequent access undecidable — SystemVerilog's cast zero-extends, so the explicit extension is needed only in Verilog. And in VHDL, access is a reserved word for access types, so the testbench's request procedure is named do_access.

8. Why a Verification Engineer Cares

Azvya Education Pvt. Ltd.VLSI Mentor
xip_window.sva — the stream's lifetime is the specification
   // 1. THE SAFETY PROPERTY. A stream hit is only ever claimed when the
   //    address continues the open stream. This is the one failure in the
   //    module that returns WRONG DATA rather than slow data, because a hit
   //    performs no command and so has nothing that could fail.
   a_hit_is_real : assert property (
       @(posedge clk) disable iff (!rst_n)
           (valid && stream_hit) |-> (flash_addr == $past(next_addr)))
       else $error("a stream hit was claimed for a non-continuing address");

   // 2. ANYTHING ELSE ON THE BUS CLOSES THE STREAM. The controller's belief
   //    about the device's pointer is only valid while nothing else has
   //    selected the device.
   a_foreign_access_closes : assert property (
       @(posedge clk) disable iff (!rst_n)
           (valid && out_of_window) |=> !stream_open)
       else $error("an out-of-window access left the stream open");

   // 3. A HIT HAS NO LATENCY, and a miss always has some. The two are
   //    exclusive, and a hit whose latency were non-zero would mean the
   //    overhead was paid without being needed.
   a_hit_free : assert property (
       @(posedge clk) disable iff (!rst_n)
           (valid && hit_window) |->
               ((stream_hit && latency_cycles == 0) ||
                (!stream_hit && latency_cycles != 0)))
       else $error("a hit paid latency, or a miss did not");

   // 4. XIP REMOVES EXACTLY THE COMMAND PHASE. Not the address phase and
   //    not the dummy phase -- a controller that removed more would read
   //    from the wrong place or sample before the data was valid.
   a_xip_scope : assert property (
       @(posedge clk) disable iff (!rst_n)
           (valid && !stream_hit && hit_window) |->
               (latency_cycles == ((xip_enabled ? 0 : cmd_cycles)
                                   + addr_cycles + dummy_cycles)))
       else $error("XIP changed more or less than the command phase");

   // 5. THE WINDOW BOUNDARY IS EXACT. The last address inside and the first
   //    outside are one apart and must behave differently -- the off-by-one
   //    that would serve a neighbouring peripheral's address from flash.
   a_window_exact : assert property (
       @(posedge clk) disable iff (!rst_n)
           (valid) |-> (hit_window == ((cpu_addr - win_base) < (1 << win_bits))))
       else $error("the window boundary is off by one");

Property 1 is the one this chapter exists for, and it has a shape worth naming: it is a property about the validity of an optimisation's precondition. The optimisation is "skip the command phase"; its precondition is "the device's pointer is where we think". Every fast path in every design has this structure, and the precondition is almost always the thing that goes stale rather than the fast path itself being wrong.

Property 2 is the same property expressed as a lifetime rather than a check. Rather than verifying the belief at use, it verifies that the belief is discarded whenever it could have become invalid — which is the stronger form, because it does not depend on enumerating the ways a use could be wrong.

Coverage must reach the sequences, because the fault is a sequence:

Azvya Education Pvt. Ltd.VLSI Mentor
xip_window_cg.sv — access patterns, not addresses
   covergroup xip_window_cg @(posedge clk iff valid);
       // The access CLASS is what the design branches on, and covering raw
       // addresses cannot express it.
       cp_class : coverpoint access_class {
           bins first_after_reset = {A_FIRST};     // must miss
           bins sequential        = {A_SEQ};       // must hit
           bins random_in_window  = {A_RANDOM};    // must miss
           bins out_of_window     = {A_OUT};       // not served
       }

       // The SEQUENCE is the fault's shape. A suite issuing only sequential
       // runs, or only random accesses, never reaches the case where a
       // foreign access sits between two sequential ones.
       cp_sequence : coverpoint seq_class {
           bins seq_run           = {S_RUN};       // hit after hit
           bins random_then_seq   = {S_RESTART};   // miss then hit
           bins foreign_then_seq  = {S_STALE};     // THE dangerous case
           bins seq_then_foreign  = {S_INTERRUPT};
       }

       cp_xip : coverpoint xip_enabled { bins off = {0}; bins on = {1}; }

       // Line size decides how much the saving is worth, and the small bins
       // are where it is worth most -- the opposite of a throughput suite's
       // instinct.
       cp_line : coverpoint line_bytes {
           bins word    = {4};           // 2.75x -- the CPU fetch case
           bins small   = {[5:16]};
           bins cache   = {32, 64};
           bins large   = {[65:$]};      // 1.34x and falling
       }

       // The window boundary, exactly. One apart, opposite outcomes.
       cp_boundary : coverpoint boundary_class {
           bins last_inside  = {B_LAST_IN};
           bins first_outside = {B_FIRST_OUT};
           bins below_base   = {B_BELOW};
           bins well_inside  = {B_MIDDLE};
       }

       x_class_xip  : cross cp_class, cp_xip;
       x_seq_line   : cross cp_sequence, cp_line;
   endgroup

cp_sequence's foreign_then_seq bin is the coverage goal of the entire chapter. It is the only bin that can expose the stale-stream fault, it requires a three-access sequence to reach, and a suite built from sequential runs and random accesses never produces it — because nothing in either pattern puts a foreign access between two sequential ones.

That is the general lesson about coverage worth taking from this module: when a fault is a property of a sequence, a coverage model over individual transactions cannot express it, however finely the transactions are binned.

9. Why an FPGA or ASIC Engineer Cares

Close the stream on anything you do not serve. Make it part of not serving the access rather than a separate guard, so no path can both skip the window and leave the stream open. This is the difference between a performance bug and a data-corruption bug.

Reset with the stream closed. The first access after reset must be a miss, because nothing is open. A controller resetting with a stream believed open reads from wherever the device happens to be.

Make the window size a power of two and test it with a shift. A comparison against a computed limit is larger and gets the boundary wrong more often.

Distinguish "not ours" from "ours and slow". They need different responses upstream — one is a bus error or another decoder's business, the other is a stall.

Share the bus reluctantly. Every other use of the flash costs a stream close and therefore a full-latency access afterwards. On an XIP system the flash is effectively code memory, and treating it as a peripheral that anything may poll is expensive in a way that does not show up in a throughput measurement.

Publish the hit rate. Two counters — accesses and stream hits — turn "why is the CPU slow?" into a ratio. A hit rate far below what the code's locality predicts is usually something else on the bus, and nothing else will tell you that.

Do not assume a fixed line size. A CPU issuing variable-width loads advances the stream by different amounts, so the tracker must add the actual bytes consumed rather than a constant.

10. Failure Signature — XIP Code That Runs Slowly Only When a Sensor Is Polled

Symptom. A processor executes from XIP flash. A benchmark loop runs at the expected speed. In the full application the same loop runs about three times slower — and the slowdown correlates with a background task that reads a sensor on the same SPI bus once a millisecond.

What "correlates with the sensor task" establishes. The code and the flash are unchanged, so the slowdown is caused by the other bus user. That is almost the whole diagnosis, and the remaining question is by what mechanism — because there are two, and they need different fixes.

Plausible mechanisms.

  • Every sensor access closes the stream, so the instruction fetch after it pays full latency instead of none. One close per millisecond costs one full-latency fetch per millisecond, which is negligible — so this alone does not explain 3×.
  • The stream is closed far more often than the sensor is read. If the driver polls a status register in a loop, or the bus is arbitrated per-byte rather than per-transaction, there may be hundreds of closes per sensor reading.
  • Bus arbitration inserts waits on the instruction fetch itself, independent of the stream.
  • The sensor task holds chip select or a lock across its whole transaction, stalling fetches for its duration.
  • Cache thrash: the sensor task's data evicts the loop's instructions, so fetches that were hits become misses — a mechanism that has nothing to do with SPI at all.

The discriminating observation. Measure the stream hit rate, which §9 recommends publishing for exactly this reason. Three outcomes, three different answers:

  • Hit rate near 100% — the stream is not the problem, and the slowdown is arbitration or cache. Instrument those.
  • Hit rate around 99% — consistent with one close per millisecond, which cannot produce 3×. Look elsewhere.
  • Hit rate collapsed to a few per cent — the stream is being closed constantly, and the next question is how many times per sensor reading.

For the third case, count the closes and divide by the sensor readings. A ratio near 1 means the arbitration is per-transaction and something else is wrong; a ratio in the hundreds means the flash is losing the bus per byte or the sensor driver is polling, and the fix is at that layer.

The fix depends on the ratio, which is why measuring it comes first. But the structural answer, if the flash is genuinely code memory, is to give it its own bus — a dedicated SPI controller for the XIP flash and a second for peripherals. That costs a controller and removes an entire class of interaction whose symptom appears in code that has nothing to do with SPI.

Why this is hard to find. Because the slow code is not the code that caused it. An engineer profiling the loop finds the loop is slow, with no indication that a task elsewhere is responsible — and XIP makes instruction fetch invisible, so there is nothing in the profile that looks like I/O. The hit-rate counter is the only cheap instrument that points at the real layer, and it has to have been built in before the question was asked.

11. Common Misconceptions

12. Reason It Through

Work this before reading the answer.

A processor executes from XIP quad flash with a 32-byte line, 14 cycles of random latency and 64 cycles per line. Its instruction cache is disabled during early boot, so every instruction fetch goes to flash.

The boot code contains a loop of 40 sequential instructions that runs 1000 times. How many full-latency accesses occur, and what would enabling the cache change?

First, what makes an access sequential. The stream continues only if the next request is for the address the last one ended at. A loop's body is sequential — but the branch back to the top is not.

Count the accesses per iteration. Forty instructions of, say, four bytes is 160 bytes, which at a 32-byte line is five lines. Within one iteration, lines 2 to 5 are sequential and line 1 follows the backward branch:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   line 1:  random  — the branch target, not where line 5 ended
   lines 2-5: sequential

So one full-latency access per iteration, and 1000 iterations give 1000 full-latency accesses — plus 4000 sequential ones.

The cost:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   1000 × (14 + 64)  =  78 000 cycles     the branch-target lines
   4000 × (0  + 64)  = 256 000 cycles     the sequential lines
                        ─────────
                        334 000 cycles

What if every access were random? 5000 × 78 = 390 000. So the stream saves 56 000 cycles — 1.17×, which is much less than §4's 1.34× because four accesses in five were already sequential and only the fifth could be improved.

Now enable the cache. The loop body is 160 bytes. If the cache holds it, the first iteration fetches all five lines and the remaining 999 fetch nothing:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   iteration 1:  14 + 64 + 4 × 64  =  334 cycles
   iterations 2-1000:                    0 flash cycles

334 cycles instead of 334 000 — a factor of a thousand, and it dwarfs every technique in this module combined.

Which is the answer worth having. XIP, continuous read, quad width and DDR together might improve this loop by 3× or 4×. A cache that holds the working set improves it by 1000×, because it removes the accesses rather than accelerating them.

So what is XIP actually for? Two things, and both matter:

Code that does not fit the cache. A large, cold, mostly-sequential path — boot code, an error handler, a rarely-used driver — is fetched once and evicted. XIP makes that fetch cheap, and there is no cache behaviour to exploit because there is no reuse.

Cache line fills themselves. With a cache enabled, every miss is a line fill — a 32-byte sequential burst, which is exactly what continuous read serves best. XIP does not compete with the cache; it makes the cache's misses cheaper.

The general lesson. Measure where the accesses actually are before optimising how fast they are. A technique that makes each access 3× faster loses to one that removes 99.9% of them — and the two compose, so the right order is to remove what you can and accelerate what remains.

13. Understanding Check

14. Summary

A memory-mapped controller turns a CPU load into a flash transaction, and the stall it causes depends on one question: does this access continue the stream the device already has open?

Continuous read removes the command, address and dummy phases — for sequential accesses. XIP removes the command phase — for all of them. Width and DDR shrink what remains. Three mechanisms, and they compose.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   random, no XIP    22 + data
   random, XIP       14 + data
   sequential         0 + data

The result inverts the module. Sequential access saves 1.34× on a 32-byte line and 2.75× on a four-byte fetch — the exact reverse of quad output's 3.28× and 1.50×. Fixed overhead limited what width could win on short transfers; removing it helps short transfers most, and a CPU's instruction fetch is precisely that case.

The stream must close on anything else that uses the bus — a foreign access, a status poll, a second slave. A stale stream returns wrong data silently, because a hit performs no command and so has nothing that can fail. That makes it the only failure in this module that is a correctness bug rather than a performance one, and it belongs in the else branch of the window test so no path can skip the window and leave the stream open.

In hardware the window test is a shift, reset leaves the stream closed, and "not ours" is distinguished from "ours and slow" because they need different responses.

For verification, the property that matters is about the validity of an optimisation's precondition — every fast path has one, and it is the precondition that goes stale. And the coverage insight is that a fault which is a property of a sequence cannot be expressed by a model over individual transactions, however finely they are binned.

Finally, the measurement that puts all of it in proportion: a cache that holds the working set removes accesses rather than accelerating them, and removing 99.9% of them beats making each one three times faster. XIP's real value is the code that does not fit — and making the cache's own misses cheap.

15. What Comes Next

This closes Module 12, and with it the description of what SPI is. Twelve modules have covered the clocking, the framing, the modes, the transactions, the topology, the performance, the datasheet, the canonical device, and the widening that keeps a serial interface relevant against parallel alternatives.

Module 13 — Designing an SPI Master in RTL turns from reading the protocol to building it. It is the track's flagship design module: taking a device transaction specification and converting it into testable RTL requirements, partitioning control from datapath and justifying every block before writing it, generating mode-aware SCLK with launch and capture strobes, building a shift datapath that serves all four modes without four copies of itself, and arriving at a configurable master that survives review. Everything this track has established becomes a requirement there — and the chapters that follow are where the requirements become a design.

Continue learning