SPI · Module 13
Multi-Byte and Back-to-Back Transactions
Keeping the wire busy between frames with one register's worth of flops: why one deep is enough and what a deeper FIFO actually buys, why a stall must be reported rather than absorbed, and why an overrun is lost data.
Everything so far transfers one frame correctly. A real transaction is a stream, and the gap between its frames is where the throughput arithmetic of Module 9 is won or lost.
A flash page program is 4 command and address bytes followed by 256 data bytes. Between which two of those 260 frames is the wire idle?
With a single transmit register, between all 259 boundaries — and for however long the CPU takes to notice each one. The fix is one register deep, and it is worth being precise about why one is enough.
1. The Problem With A Single Transmit Register
The datapath of Chapter 13.6 holds the word being shifted. If that is the only place a word can live, the next word cannot be written until the current one has left — and "has left" means the frame is over:
shift word N ... frame ends ... CPU notices ... CPU writes word N+1
... controller starts ... shift word N+1Every one of those ellipses is dead time on SCLK, and none of it is under the hardware's control: it is however long the CPU takes to notice. At 50 MHz SCLK and a 100 MHz CPU, a byte takes 160 ns and an interrupt-driven driver can easily spend more time between bytes than inside them. The transfer is correct and the wire is half idle.
8-bit frame at div=4, 100 MHz clock 320 ns
interrupt latency + handler 1000 ns, typically
wire utilisation 24%That number is not a pathological case. It is what an interrupt-per-byte driver achieves, and it is why the register below exists.
2. One Register Deep, And What Deeper Buys
Add a shadow register in front of the datapath. The CPU writes the shadow; the controller moves the shadow into the datapath between frames. The two are decoupled, so the CPU can write word N+1 at any point while word N is still shifting — an entire frame's worth of time instead of an inter-frame gap.
without: CPU has [inter-frame gap] to produce the next word
with: CPU has [a whole frame] to produce the next wordOne deep is enough to remove the dependency, and this is worth stating carefully because it is where intuition goes wrong:
A deeper FIFO buys latency TOLERANCE, not throughput.
Once the producer can stay one word ahead, the wire is already saturated and nothing deeper makes it faster. What depth buys is survival of a longer stall: a four-deep FIFO lets the CPU disappear for four frame times instead of one without the wire going idle. That is a real and useful property — it is what lets a driver batch its work — and it is not a throughput property.
depth 1 producer must return within 1 frame time wire saturated
depth 4 producer must return within 4 frame times wire saturated
depth 16 producer must return within 16 frame times wire saturatedAll three saturate the wire. They differ in how long the producer may be absent, which is a question about the system, not about the interface. So the right depth is decided by the worst-case interrupt latency divided by the frame time — and for a design where that ratio is below one, depth one is the correct answer and anything more is flops spent on nothing.
3. What The Block Reports
A design that silently absorbs a slow producer is worse than one that says so, because the lost throughput is invisible. So two flags:
stalled — the transaction is open, the clock is not running, and nothing is queued. The wire is idle because the CPU was late. Every cycle of this is throughput that was paid for and not used. It is not an error: it is the exact measure of how late the producer was, and a driver that is late by design — a low-priority background transfer, say — will assert it constantly and correctly.
rx_overrun — a received word was overwritten before it was read. Unlike a stall this is lost data, and the flag must be sticky: a receive path that drops a word and carries on is how a flash image ends up with one wrong word in it, and a driver that only checks at the end of a transaction must still find out.
The distinction between the two is worth keeping in mind as a general pattern:
a performance flag reports a COST -- not an error, clear on read is fine
a data-loss flag reports a LOSS -- an error, must be sticky4. The Tail Is Not A Stall
stalled has a subtlety that produced a false report in the first version of this block, and it generalises.
After the final frame of a transaction, the chip-select controller of Chapter 13.7 walks out through LAG and GAP. Throughout that, the transaction is still open — busy is asserted — the clock is not running, and nothing is queued, because by design there is nothing more to send.
That is exactly the condition stalled tests for. Without a guard, every transaction ends with a burst of stall cycles equal to the lag plus the gap, and a driver reading the counter concludes it is always late. A flag that fires on correct behaviour is a flag that gets ignored, and then it is worse than absent.
So the block tracks tail: set when the final frame of a transaction starts, cleared on the next load, and ANDed out of stalled. Two flops, and the flag now means what it says.
5. The Two Cases
Read the sh_full row against the prompt sclk row. The slot empties at cycle 3 — the load — and refills at cycle 5, giving the CPU a window that spans the rest of the frame. And read the tail row at the far right: the final frame has started, so the idle cycles that follow are the transaction closing, not the producer being late.
6. Building the Streaming Controller — Three HDLs
The circuit
A one-word shadow with a valid bit, an armed flag, a one-word receive holding register, and the two flags. The decisions:
The datapath may only be loaded while the clock is stopped. Loading it mid-frame would overwrite the word being shifted — the one failure a single-register design cannot even express, and that a shadow design has to be careful about. The guard is !shift_en.
armed stops a second load landing on the first. It means "the datapath is loaded and a frame has been requested, but the frame has not started yet". Without it, a shadow refilled during the lead time would overwrite the word the controller is about to send.
A push and a load on the same cycle both succeed. The shadow is freed by the load and filled by the push, and the two are ordered so neither is lost. That matters more than it looks: the CPU polling tx_ready and writing the instant it goes high will hit exactly this cycle, every time, on a tight loop.
flush throws queued work away. A word queued for a transaction that is being abandoned must not become the first word of the next one, and an armed request left standing would start that next transaction by itself the moment the bus came free. Chapter 13.10 drives it; here it is simply the one input that can discard work.
clr_flags clears the sticky overrun, and the clear is applied first. A clear arriving on the same cycle as a fresh overrun must not swallow it — losing a report of lost data is the same bug twice.
// spi_multibyte.sv
//
// Chapter 13.9 -- keeping the wire busy between frames.
//
// Everything built so far transfers ONE frame correctly. A real transaction is
// a stream: a flash page program is a command byte, three address bytes and
// then 256 data bytes, all under one chip select. Getting each of those bytes
// right individually is not the same as getting the stream right, and the gap
// between them is where the throughput of Chapter 9 is won or lost.
//
// THE PROBLEM WITH A SINGLE TRANSMIT REGISTER.
//
// The datapath of 13.6 holds the word being shifted. If that is the only place
// a word can live, then the next byte cannot be written until the current one
// has left -- and "has left" means the frame is over. So the sequence is:
//
// shift byte N ... frame ends ... CPU notices ... CPU writes byte N+1
// ... controller starts ... shift byte N+1
//
// Every one of those dots is dead time on SCLK, and none of it is under the
// hardware's control: it is however long the CPU takes to notice. At 50 MHz
// SCLK and a 100 MHz CPU, an interrupt-driven driver can easily spend more
// time between bytes than inside them. The transfer is correct and the wire is
// half idle.
//
// THE FIX IS ONE REGISTER DEEP.
//
// Add a SHADOW register in front of the datapath. The CPU writes into the
// shadow; the controller moves the shadow into the datapath between frames.
// The two are decoupled, so the CPU can write byte N+1 at any point while byte
// N is still shifting -- it has an entire frame's worth of time to do it,
// instead of the inter-frame gap.
//
// One deep is enough to remove the dependency, and that is worth being clear
// about: a deeper FIFO buys latency TOLERANCE (surviving a longer CPU stall),
// not extra throughput. Once the producer can stay one byte ahead, the wire is
// already saturated and nothing deeper makes it faster.
//
// WHAT THIS BLOCK REPORTS.
//
// A design that silently absorbs a slow producer is worse than one that says
// so, because the lost throughput is invisible. So:
//
// stalled the transaction is open, the clock is not running, and there is
// nothing queued -- the wire is idle because the CPU was late.
// Every cycle of this is throughput that was paid for and not
// used.
// rx_overrun a received word was overwritten before it was read. Unlike a
// stall this is not a performance problem, it is LOST DATA, and
// it must be sticky: a receive path that drops a byte and carries
// on is how a flash image ends up with one wrong word in it.
module spi_multibyte #(
parameter int MAX_W = 32,
parameter int LEN_W = 6
) (
input wire clk,
input wire rst_n,
// --- CPU side -------------------------------------------------------
input wire tx_push, // write a word into the shadow
input wire [MAX_W-1:0] tx_wdata,
input wire tx_last, // ... and end the transaction after it
output wire tx_ready, // the shadow is free
input wire flush, // drop whatever is queued
input wire rx_pop, // read the received word
output wire [MAX_W-1:0] rx_rdata,
output wire rx_ready, // a received word is waiting
output reg rx_overrun, // sticky: a word was lost
// Clearing a sticky flag is the REGISTER INTERFACE's business, not this
// block's -- but the flag lives here, because this is where the loss is
// detected. So each flag has exactly one owner and exactly one clear path,
// and Chapter 13.11 wires this to a write of the command register.
input wire clr_flags,
// --- transfer engine side (13.7 and 13.6/13.8) ----------------------
input wire core_busy, // a transaction is open
input wire shift_en, // the clock is running
input wire start_stb, // the engine took a frame
input wire rx_valid_stb,
input wire [MAX_W-1:0] rx_word,
output wire req, // to the chip-select controller
output wire hold,
output reg [MAX_W-1:0] tx_data, // to the shift datapath
output wire load_stb,
output wire stalled // the wire is idle for want of data
);
// The shadow: one word, its end-of-transaction flag, and a valid bit.
reg [MAX_W-1:0] sh_data;
reg sh_last;
reg sh_valid;
// `armed` means the datapath has been loaded and the controller has been
// asked for a frame, but the frame has not started yet. It stops a second
// load from landing on top of the first.
reg armed;
reg armed_last;
reg load_r;
// `tail` means the final frame of the transaction has started, so the
// controller is walking out through its lag and gap with nothing more to
// send. Without it the closing states look exactly like a producer that
// went quiet, and `stalled` would accuse the CPU of being late at the end
// of every single transaction.
reg tail;
assign tx_ready = ~sh_valid;
assign req = armed;
assign hold = ~armed_last;
assign load_stb = load_r;
// The datapath may only be loaded while the clock is stopped. Loading it
// mid-frame would overwrite the word being shifted -- the one failure
// that a single-register design cannot even express, and that a shadow
// design has to be careful about.
wire can_load = sh_valid & ~armed & ~shift_en & ~flush;
// Idle inside an open transaction with nothing to send. Note it is NOT an
// error: it is the exact measure of how late the producer was.
assign stalled = core_busy & ~shift_en & ~armed & ~sh_valid & ~tail;
reg [MAX_W-1:0] rx_hold;
reg rx_full;
assign rx_rdata = rx_hold;
assign rx_ready = rx_full;
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
sh_data <= {MAX_W{1'b0}};
sh_last <= 1'b0;
sh_valid <= 1'b0;
armed <= 1'b0;
armed_last <= 1'b0;
load_r <= 1'b0;
tail <= 1'b0;
tx_data <= {MAX_W{1'b0}};
rx_hold <= {MAX_W{1'b0}};
rx_full <= 1'b0;
rx_overrun <= 1'b0;
end else begin
load_r <= 1'b0;
// --- the CPU writing into the shadow ---
// Accepted whenever the shadow is free, which includes the same
// cycle it is being emptied into the datapath below: the two are
// ordered so a push and a load on one cycle both succeed.
// FLUSH takes priority over everything on the transmit side. A
// word queued for a transaction that is being abandoned must not
// become the first word of the next one, and an `armed` request
// left standing would start that next transaction by itself the
// moment the bus came free. Chapter 13.10 is where this is driven
// from; here it is simply the one input that can throw work away.
if (flush) begin
sh_valid <= 1'b0;
armed <= 1'b0;
armed_last <= 1'b0;
tail <= 1'b0;
end else if (tx_push && tx_ready) begin
sh_data <= tx_wdata;
sh_last <= tx_last;
sh_valid <= 1'b1;
end
// --- the shadow moving into the datapath ---
if (can_load && !flush) begin
tx_data <= sh_data;
armed_last <= sh_last;
armed <= 1'b1;
load_r <= 1'b1;
tail <= 1'b0;
// Freed immediately, so the CPU may refill it on the very
// next cycle rather than waiting for the frame to finish.
// This single line is what makes the stream back-to-back.
if (!(tx_push && tx_ready))
sh_valid <= 1'b0;
end
// --- the engine taking the frame ---
if (start_stb) begin
armed <= 1'b0;
if (armed_last)
tail <= 1'b1;
end
// --- the receive side ---
// The clear is applied FIRST so that a clear arriving on the same
// cycle as a fresh overrun does not swallow it: the new event wins,
// because losing a report of lost data is the same bug twice.
if (clr_flags)
rx_overrun <= 1'b0;
if (rx_valid_stb) begin
rx_hold <= rx_word;
rx_full <= 1'b1;
// A word arriving on top of an unread one is DATA LOST, and
// the flag is sticky so a driver that only checks at the end
// of the transaction still finds out.
if (rx_full && !rx_pop)
rx_overrun <= 1'b1;
end else if (rx_pop) begin
rx_full <= 1'b0;
end
end
end
endmodule// spi_multibyte_tb.sv
//
// This is the first testbench that runs the whole engine: the divider of 13.4,
// the mode logic of 13.5, the datapath of 13.6 wrapped by 13.8, the chip-select
// controller of 13.7, and the streaming shadow of 13.9. Only the register
// interface is missing, and that is Chapter 13.11.
//
// The point of the chapter is throughput, so the measurements are about TIME as
// well as data:
//
// - a prompt producer must stream N bytes under ONE chip select with no
// stall cycles at all, and with the inter-frame idle equal to the gap the
// controller was told to insert -- not a cycle more;
// - a deliberately late producer must still be CORRECT, and must be
// reported: the stall counter is how the lost throughput becomes visible;
// - a receiver that does not read fast enough must set a STICKY overrun,
// because a dropped word is lost data, not a slow day.
//
// Data is checked in both directions against an independent slave, byte by
// byte and in order, so a stream that gets every byte right but in the wrong
// sequence fails.
`timescale 1ns/1ps
module spi_multibyte_tb;
localparam int MAX_W = 32;
localparam int LEN_W = 6;
localparam int DIV_W = 8;
localparam int N_CS = 2;
localparam int SEL_W = 1;
localparam int CNT_W = 8;
logic clk = 1'b0;
logic rst_n = 1'b0;
always #5 clk = ~clk;
// --- 13.9, the block under test ---------------------------------------
logic tx_push = 1'b0;
logic [MAX_W-1:0] tx_wdata = 32'h0;
logic tx_last = 1'b0;
wire tx_ready;
logic flush = 1'b0;
logic clr_flags = 1'b0;
logic rx_pop = 1'b0;
wire [MAX_W-1:0] rx_rdata;
wire rx_ready, rx_overrun;
wire req, hold, load_stb, stalled;
wire [MAX_W-1:0] tx_data;
// --- 13.7 -------------------------------------------------------------
logic [SEL_W-1:0] sel = 1'b0;
logic [CNT_W-1:0] lead_cyc = 8'd3;
logic [CNT_W-1:0] lag_cyc = 8'd3;
logic [CNT_W-1:0] gap_cyc = 8'd2;
wire [N_CS-1:0] cs_n;
wire shift_en, start_stb, busy, sel_err;
wire [2:0] state_id;
// --- 13.4 -------------------------------------------------------------
logic [DIV_W-1:0] div = 8'd4;
logic cpol = 1'b0;
wire sclk, edge_a_stb, edge_b_stb, bit_done, div_err;
wire [DIV_W-1:0] half_a, half_b;
// --- 13.5 / 13.8 ------------------------------------------------------
logic cpha = 1'b0;
logic [LEN_W-1:0] len = 6'd8;
logic lsb_first = 1'b0;
wire preload_stb, launch_stb, capture_stb, frame_done;
wire [LEN_W-1:0] bit_idx;
wire miso, mosi;
wire [MAX_W-1:0] rx_word;
wire rx_valid_stb, len_err;
spi_clkdiv_strobe #(.DIV_W(DIV_W)) u_div (
.clk(clk), .rst_n(rst_n), .en(shift_en), .div(div), .cpol(cpol),
.sclk(sclk), .edge_a_stb(edge_a_stb), .edge_b_stb(edge_b_stb),
.bit_done(bit_done), .half_a(half_a), .half_b(half_b),
.div_err(div_err)
);
spi_mode_edges #(.LEN_W(LEN_W)) u_mode (
.clk(clk), .rst_n(rst_n), .cpha(cpha), .len(len),
.active(shift_en), .start_stb(start_stb),
.edge_a_stb(edge_a_stb), .edge_b_stb(edge_b_stb),
.preload_stb(preload_stb), .launch_stb(launch_stb),
.capture_stb(capture_stb), .bit_idx(bit_idx), .frame_done(frame_done)
);
spi_width_order #(.MAX_W(MAX_W), .LEN_W(LEN_W)) u_data (
.clk(clk), .rst_n(rst_n),
.tx_data(tx_data), .len(len), .lsb_first(lsb_first),
.load_stb(load_stb),
.preload_stb(preload_stb), .launch_stb(launch_stb),
.capture_stb(capture_stb),
.miso(miso), .mosi(mosi),
.rx_data(rx_word), .rx_valid_stb(rx_valid_stb), .len_err(len_err)
);
spi_cs_ctrl #(.N_CS(N_CS), .SEL_W(SEL_W), .CNT_W(CNT_W)) u_cs (
.clk(clk), .rst_n(rst_n),
.req(req), .hold(hold), .sel(sel),
.lead_cyc(lead_cyc), .lag_cyc(lag_cyc), .gap_cyc(gap_cyc),
.core_done(frame_done), .abort(1'b0),
.cs_n(cs_n), .shift_en(shift_en), .start_stb(start_stb),
.busy(busy), .sel_err(sel_err), .state_id(state_id)
);
spi_multibyte #(.MAX_W(MAX_W), .LEN_W(LEN_W)) dut (
.clk(clk), .rst_n(rst_n),
.tx_push(tx_push), .tx_wdata(tx_wdata), .tx_last(tx_last),
.tx_ready(tx_ready),
.flush(flush), .clr_flags(clr_flags),
.rx_pop(rx_pop), .rx_rdata(rx_rdata), .rx_ready(rx_ready),
.rx_overrun(rx_overrun),
.core_busy(busy), .shift_en(shift_en), .start_stb(start_stb),
.rx_valid_stb(rx_valid_stb), .rx_word(rx_word),
.req(req), .hold(hold), .tx_data(tx_data), .load_stb(load_stb),
.stalled(stalled)
);
// --- the independent slave, word by word ------------------------------
logic [MAX_W-1:0] slv_tx_q [0:127];
logic [MAX_W-1:0] slv_rx_q [0:127];
integer slv_tx_i, slv_rx_i;
logic [MAX_W-1:0] slv_sr, slv_rx_sr;
logic slave_bit_r;
assign miso = slave_bit_r;
always_ff @(posedge clk) begin
if (load_stb) begin
slv_sr <= slv_tx_q[slv_tx_i] << (MAX_W - len);
slv_tx_i <= slv_tx_i + 1;
end else if (preload_stb || launch_stb) begin
slave_bit_r <= slv_sr[MAX_W-1];
slv_sr <= {slv_sr[MAX_W-2:0], 1'b0};
end
// Cleared at each load, so the recorded word holds THIS frame's bits
// and not a running history of every frame before it.
if (load_stb)
slv_rx_sr <= {MAX_W{1'b0}};
else if (capture_stb)
slv_rx_sr <= {slv_rx_sr[MAX_W-2:0], mosi};
// `rx_valid_stb` is one cycle after the final capture, so the slave's
// own shift register is already complete when it is recorded.
if (rx_valid_stb) begin
slv_rx_q[slv_rx_i] <= slv_rx_sr;
slv_rx_i <= slv_rx_i + 1;
end
end
// --- the throughput monitor -------------------------------------------
integer stall_cycles;
integer idle_in_txn; // cycles inside a transaction with no clock
integer frames_started;
integer cs_falls;
integer slack_cycles; // room in the shadow WHILE the wire is busy
integer valid_pulses;
integer n_loads, n_caps, n_pre, n_lau, n_race;
logic busy_q;
always_ff @(posedge clk) begin
if (!rst_n) begin
busy_q <= 1'b0;
end else begin
busy_q <= busy;
if (stalled) stall_cycles <= stall_cycles + 1;
if (busy && !shift_en) idle_in_txn <= idle_in_txn + 1;
if (start_stb) frames_started <= frames_started + 1;
if (busy && !busy_q) cs_falls <= cs_falls + 1;
if (tx_ready && shift_en) slack_cycles <= slack_cycles + 1;
if (rx_valid_stb) valid_pulses <= valid_pulses + 1;
if (load_stb) n_loads <= n_loads + 1;
if (capture_stb) n_caps <= n_caps + 1;
if (preload_stb) n_pre <= n_pre + 1;
if (launch_stb) n_lau <= n_lau + 1;
// The datapath is loaded by one strobe and told to drive its first
// bit by another. If they ever land on the same cycle the load is
// overwritten by the shift in the same always block and the frame
// sends the PREVIOUS word. The handover is built so they cannot,
// and this counts the cycles on which that claim would be false.
if (load_stb && start_stb) n_race <= n_race + 1;
end
end
integer errors = 0;
task automatic clear_counts;
begin
stall_cycles = 0; idle_in_txn = 0; frames_started = 0;
cs_falls = 0; slack_cycles = 0; valid_pulses = 0;
n_loads = 0; n_caps = 0; n_pre = 0; n_lau = 0; n_race = 0;
end
endtask
// The CPU writing a word. `late` cycles of dithering before the write is
// how a slow driver is modelled.
task automatic push(input [MAX_W-1:0] w, input bit last,
input integer late);
integer guard;
begin
repeat (late) @(negedge clk);
guard = 8000;
while (!tx_ready && guard > 0) begin
@(negedge clk);
guard = guard - 1;
end
if (guard == 0) begin
$display(" FAIL: the shadow never became free");
errors = errors + 1;
end
tx_wdata = w;
tx_last = last;
tx_push = 1'b1;
@(negedge clk);
tx_push = 1'b0;
end
endtask
// The CPU reading. `drain` = 0 leaves words unread on purpose.
integer got_n;
logic [MAX_W-1:0] got_q [0:127];
task automatic drain_rx;
begin
while (rx_ready) begin
got_q[got_n] = rx_rdata;
got_n = got_n + 1;
rx_pop = 1'b1;
@(negedge clk);
rx_pop = 1'b0;
@(negedge clk);
end
end
endtask
task automatic wait_idle;
integer guard;
begin
guard = 20000;
while (busy && guard > 0) begin
drain_rx();
@(negedge clk);
guard = guard - 1;
end
repeat (4) @(negedge clk);
drain_rx();
end
endtask
logic [MAX_W-1:0] sent_q [0:127];
integer i, k, nb, late, seed, bad;
initial begin
clear_counts();
slv_tx_i = 0; slv_rx_i = 0; got_n = 0;
slv_sr = 32'h0; slv_rx_sr = 32'h0; slave_bit_r = 1'b0;
seed = 32'h5EED_1234;
for (i = 0; i < 128; i = i + 1) begin
seed = (seed * 32'h0019_660D) + 32'h3C6E_F35F;
slv_tx_q[i] = (seed >> 7) & 32'hFF;
end
repeat (3) @(negedge clk);
rst_n = 1'b1;
@(negedge clk);
// 1. A PROMPT PRODUCER, sixteen bytes, one chip select. This is the
// case the shadow register exists for.
clear_counts();
nb = 16;
for (i = 0; i < nb; i = i + 1) begin
seed = (seed * 32'h0019_660D) + 32'h3C6E_F35F;
sent_q[i] = (seed >> 11) & 32'hFF;
end
got_n = 0;
for (i = 0; i < nb; i = i + 1) begin
push(sent_q[i], (i == nb - 1), 0);
drain_rx();
end
wait_idle();
if (cs_falls != 1) begin
$display(" FAIL: a %0d-byte stream took %0d chip selects",
nb, cs_falls);
errors = errors + 1;
end
if (frames_started != nb) begin
$display(" FAIL: %0d bytes produced %0d frames", nb,
frames_started);
errors = errors + 1;
end
if (stall_cycles != 0) begin
$display(" FAIL: a prompt producer still stalled the wire for %0d cycles",
stall_cycles);
errors = errors + 1;
end
// Everything the slave received, in order.
bad = 0;
for (i = 0; i < nb; i = i + 1)
if (slv_rx_q[i] !== sent_q[i]) bad = bad + 1;
if (bad != 0) begin
$display(" FAIL: the slave received %0d of %0d bytes wrongly",
bad, nb);
errors = errors + 1;
end
// Everything the master received, in order.
bad = 0;
if (got_n != nb) begin
$display(" FAIL: the master read back %0d of %0d bytes",
got_n, nb);
errors = errors + 1;
end else begin
for (i = 0; i < nb; i = i + 1)
if (got_q[i] !== slv_tx_q[i]) bad = bad + 1;
if (bad != 0) begin
$display(" FAIL: the master received %0d of %0d bytes wrongly",
bad, nb);
errors = errors + 1;
end
end
if (rx_overrun) begin
$display(" FAIL: a fully drained receiver reported an overrun");
errors = errors + 1;
end
// The per-frame bookkeeping, counted over the whole stream rather
// than trusted: one load and one preload per frame, len captures per
// frame, and len-1 launches because CPHA=0's preload is the first
// drive (Chapter 13.5).
if (n_loads != nb || n_pre != nb) begin
$display(" FAIL: %0d frames produced %0d loads and %0d preloads",
nb, n_loads, n_pre);
errors = errors + 1;
end
if (n_caps != nb * 8 || n_lau != nb * 7) begin
$display(" FAIL: %0d eight-bit frames produced %0d captures and %0d launches, expected %0d and %0d",
nb, n_caps, n_lau, nb * 8, nb * 7);
errors = errors + 1;
end
if (n_race != 0) begin
$display(" FAIL: a load landed on a frame start %0d times",
n_race);
errors = errors + 1;
end
if (valid_pulses != nb) begin
$display(" FAIL: %0d frames raised %0d received-word pulses",
nb, valid_pulses);
errors = errors + 1;
end
$display(" per-frame bookkeeping over the stream: %0d loads, %0d preloads, %0d launches, %0d captures, %0d received words, and no load ever landed on a frame start",
n_loads, n_pre, n_lau, n_caps, valid_pulses);
$display(" 16 bytes, one chip select, %0d frames, %0d stall cycles, %0d idle cycles inside the transaction -- %0d per byte boundary",
frames_started, stall_cycles, idle_in_txn,
idle_in_txn / (nb - 1));
// 2. THE IDLE BETWEEN FRAMES IS THE GAP THAT WAS ASKED FOR, and not a
// cycle more. This is the check that would catch a shadow register
// that works but hands over one cycle late.
if (idle_in_txn > (nb - 1) * (gap_cyc + 3)) begin
$display(" FAIL: %0d idle cycles across %0d boundaries, more than the %0d-cycle gap explains",
idle_in_txn, nb - 1, gap_cyc);
errors = errors + 1;
end
$display(" the inter-frame idle is the programmed %0d-cycle hold and nothing else",
gap_cyc);
// 3. A LATE PRODUCER. Still correct, and now visibly costly.
clear_counts();
nb = 8;
for (i = 0; i < nb; i = i + 1) begin
seed = (seed * 32'h0019_660D) + 32'h3C6E_F35F;
sent_q[i] = (seed >> 11) & 32'hFF;
end
k = slv_tx_i;
got_n = 0;
for (i = 0; i < nb; i = i + 1) begin
push(sent_q[i], (i == nb - 1), 40);
drain_rx();
end
wait_idle();
if (stall_cycles == 0) begin
$display(" FAIL: a producer 40 cycles late stalled the wire not at all -- the counter cannot be working");
errors = errors + 1;
end
bad = 0;
for (i = 0; i < nb; i = i + 1)
if (slv_rx_q[k + i] !== sent_q[i]) bad = bad + 1;
if (bad != 0) begin
$display(" FAIL: a late producer corrupted %0d of %0d bytes",
bad, nb);
errors = errors + 1;
end
if (cs_falls != 1) begin
$display(" FAIL: a late producer broke the transaction into %0d chip selects",
cs_falls);
errors = errors + 1;
end
$display(" the same stream with the producer 40 cycles late: every byte still correct, one chip select, but %0d stall cycles the fast run did not have",
stall_cycles);
// 4. AN OVERRUN IS STICKY. Stream four bytes and never read one.
clear_counts();
nb = 4;
for (i = 0; i < nb; i = i + 1)
push(32'h5A, (i == nb - 1), 0);
// Deliberately no drain_rx here.
i = 20000;
while (busy && i > 0) begin
@(negedge clk);
i = i - 1;
end
repeat (4) @(negedge clk);
if (!rx_overrun) begin
$display(" FAIL: four unread received words did not raise an overrun");
errors = errors + 1;
end
$display(" four received words left unread: overrun raised and latched");
// It must STAY raised through a clean transfer -- a flag that clears
// itself is a flag that hides the one byte that was lost.
drain_rx();
clear_counts();
for (i = 0; i < 2; i = i + 1) begin
push(32'h33, (i == 1), 0);
drain_rx();
end
wait_idle();
if (!rx_overrun) begin
$display(" FAIL: the overrun flag cleared itself on the next clean transfer");
errors = errors + 1;
end
$display(" a clean transfer afterwards does not clear it -- the lost byte stays reported");
// 5. THE CPU CAN REFILL WHILE THE WIRE IS BUSY. Without that the
// shadow buys nothing, so it is checked directly: at least one push
// must be accepted with the clock running.
clear_counts();
// Reset the sticky flag by resetting the block, since nothing else
// clears it -- which is itself the design decision under test.
rst_n = 1'b0;
repeat (3) @(negedge clk);
rst_n = 1'b1;
repeat (2) @(negedge clk);
slv_tx_i = 0; slv_rx_i = 0; got_n = 0;
// What matters is not the instant the testbench happens to write, but
// HOW LONG the shadow stays free while the wire is busy -- that window
// is the slack the producer is given, and without the shadow it would
// be zero.
for (i = 0; i < 8; i = i + 1) begin
push(32'hC5, (i == 7), 0);
drain_rx();
end
wait_idle();
if (slack_cycles == 0) begin
$display(" FAIL: the shadow was never free while the clock was running -- it is not decoupling anything");
errors = errors + 1;
end
if (stall_cycles != 0) begin
$display(" FAIL: %0d stall cycles while refilling ahead",
stall_cycles);
errors = errors + 1;
end
$display(" the shadow stood free for %0d of the cycles the wire was busy -- that window is the slack the producer gets, and it is a whole frame wide",
slack_cycles);
if (errors == 0)
$display("PASS: a one-deep shadow register decouples the producer from the wire -- sixteen bytes stream under a single chip select with every byte correct in both directions and in order, zero stall cycles, and an inter-frame idle equal to the programmed hold and nothing more -- words are accepted while the previous frame is still shifting, which is the whole reason the shadow exists -- a producer forty cycles late still transfers every byte correctly under one chip select but now reports the stall cycles it cost, so the lost throughput is visible instead of silent -- and a received word overwritten before it was read raises a sticky overrun that a later clean transfer does not clear");
else
$display("FAIL: %0d error(s)", errors);
$finish;
end
endmodule// spi_multibyte.v
//
// Chapter 13.9 -- keeping the wire busy between frames.
//
// Everything built so far transfers ONE frame correctly. A real transaction is
// a stream: a flash page program is a command byte, three address bytes and
// then 256 data bytes, all under one chip select. Getting each of those bytes
// right individually is not the same as getting the stream right, and the gap
// between them is where the throughput of Chapter 9 is won or lost.
//
// THE PROBLEM WITH A SINGLE TRANSMIT REGISTER.
//
// The datapath of 13.6 holds the word being shifted. If that is the only place
// a word can live, then the next byte cannot be written until the current one
// has left -- and "has left" means the frame is over. So the sequence is:
//
// shift byte N ... frame ends ... CPU notices ... CPU writes byte N+1
// ... controller starts ... shift byte N+1
//
// Every one of those dots is dead time on SCLK, and none of it is under the
// hardware's control: it is however long the CPU takes to notice. At 50 MHz
// SCLK and a 100 MHz CPU, an interrupt-driven driver can easily spend more
// time between bytes than inside them. The transfer is correct and the wire is
// half idle.
//
// THE FIX IS ONE REGISTER DEEP.
//
// Add a SHADOW register in front of the datapath. The CPU writes into the
// shadow; the controller moves the shadow into the datapath between frames.
// The two are decoupled, so the CPU can write byte N+1 at any point while byte
// N is still shifting -- it has an entire frame's worth of time to do it,
// instead of the inter-frame gap.
//
// One deep is enough to remove the dependency, and that is worth being clear
// about: a deeper FIFO buys latency TOLERANCE (surviving a longer CPU stall),
// not extra throughput. Once the producer can stay one byte ahead, the wire is
// already saturated and nothing deeper makes it faster.
//
// WHAT THIS BLOCK REPORTS.
//
// A design that silently absorbs a slow producer is worse than one that says
// so, because the lost throughput is invisible. So:
//
// stalled the transaction is open, the clock is not running, and there is
// nothing queued -- the wire is idle because the CPU was late.
// Every cycle of this is throughput that was paid for and not
// used.
// rx_overrun a received word was overwritten before it was read. Unlike a
// stall this is not a performance problem, it is LOST DATA, and
// it must be sticky: a receive path that drops a byte and carries
// on is how a flash image ends up with one wrong word in it.
module spi_multibyte #(
parameter MAX_W = 32,
parameter LEN_W = 6
) (
input wire clk,
input wire rst_n,
// --- CPU side -------------------------------------------------------
input wire tx_push, // write a word into the shadow
input wire [MAX_W-1:0] tx_wdata,
input wire tx_last, // ... and end the transaction after it
output wire tx_ready, // the shadow is free
input wire flush, // drop whatever is queued
input wire rx_pop, // read the received word
output wire [MAX_W-1:0] rx_rdata,
output wire rx_ready, // a received word is waiting
output reg rx_overrun, // sticky: a word was lost
// Clearing a sticky flag is the REGISTER INTERFACE's business, not this
// block's -- but the flag lives here, because this is where the loss is
// detected. So each flag has exactly one owner and exactly one clear path,
// and Chapter 13.11 wires this to a write of the command register.
input wire clr_flags,
// --- transfer engine side (13.7 and 13.6/13.8) ----------------------
input wire core_busy, // a transaction is open
input wire shift_en, // the clock is running
input wire start_stb, // the engine took a frame
input wire rx_valid_stb,
input wire [MAX_W-1:0] rx_word,
output wire req, // to the chip-select controller
output wire hold,
output reg [MAX_W-1:0] tx_data, // to the shift datapath
output wire load_stb,
output wire stalled // the wire is idle for want of data
);
// The shadow: one word, its end-of-transaction flag, and a valid bit.
reg [MAX_W-1:0] sh_data;
reg sh_last;
reg sh_valid;
// `armed` means the datapath has been loaded and the controller has been
// asked for a frame, but the frame has not started yet. It stops a second
// load from landing on top of the first.
reg armed;
reg armed_last;
reg load_r;
// `tail` means the final frame of the transaction has started, so the
// controller is walking out through its lag and gap with nothing more to
// send. Without it the closing states look exactly like a producer that
// went quiet, and `stalled` would accuse the CPU of being late at the end
// of every single transaction.
reg tail;
assign tx_ready = ~sh_valid;
assign req = armed;
assign hold = ~armed_last;
assign load_stb = load_r;
// The datapath may only be loaded while the clock is stopped. Loading it
// mid-frame would overwrite the word being shifted -- the one failure
// that a single-register design cannot even express, and that a shadow
// design has to be careful about.
wire can_load = sh_valid & ~armed & ~shift_en & ~flush;
// Idle inside an open transaction with nothing to send. Note it is NOT an
// error: it is the exact measure of how late the producer was.
assign stalled = core_busy & ~shift_en & ~armed & ~sh_valid & ~tail;
reg [MAX_W-1:0] rx_hold;
reg rx_full;
assign rx_rdata = rx_hold;
assign rx_ready = rx_full;
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
sh_data <= {MAX_W{1'b0}};
sh_last <= 1'b0;
sh_valid <= 1'b0;
armed <= 1'b0;
armed_last <= 1'b0;
load_r <= 1'b0;
tail <= 1'b0;
tx_data <= {MAX_W{1'b0}};
rx_hold <= {MAX_W{1'b0}};
rx_full <= 1'b0;
rx_overrun <= 1'b0;
end else begin
load_r <= 1'b0;
// --- the CPU writing into the shadow ---
// Accepted whenever the shadow is free, which includes the same
// cycle it is being emptied into the datapath below: the two are
// ordered so a push and a load on one cycle both succeed.
// FLUSH takes priority over everything on the transmit side. A
// word queued for a transaction that is being abandoned must not
// become the first word of the next one, and an `armed` request
// left standing would start that next transaction by itself the
// moment the bus came free. Chapter 13.10 is where this is driven
// from; here it is simply the one input that can throw work away.
if (flush) begin
sh_valid <= 1'b0;
armed <= 1'b0;
armed_last <= 1'b0;
tail <= 1'b0;
end else if (tx_push && tx_ready) begin
sh_data <= tx_wdata;
sh_last <= tx_last;
sh_valid <= 1'b1;
end
// --- the shadow moving into the datapath ---
if (can_load && !flush) begin
tx_data <= sh_data;
armed_last <= sh_last;
armed <= 1'b1;
load_r <= 1'b1;
tail <= 1'b0;
// Freed immediately, so the CPU may refill it on the very
// next cycle rather than waiting for the frame to finish.
// This single line is what makes the stream back-to-back.
if (!(tx_push && tx_ready))
sh_valid <= 1'b0;
end
// --- the engine taking the frame ---
if (start_stb) begin
armed <= 1'b0;
if (armed_last)
tail <= 1'b1;
end
// --- the receive side ---
// The clear is applied FIRST so that a clear arriving on the same
// cycle as a fresh overrun does not swallow it: the new event wins,
// because losing a report of lost data is the same bug twice.
if (clr_flags)
rx_overrun <= 1'b0;
if (rx_valid_stb) begin
rx_hold <= rx_word;
rx_full <= 1'b1;
// A word arriving on top of an unread one is DATA LOST, and
// the flag is sticky so a driver that only checks at the end
// of the transaction still finds out.
if (rx_full && !rx_pop)
rx_overrun <= 1'b1;
end else if (rx_pop) begin
rx_full <= 1'b0;
end
end
end
endmodule// spi_multibyte_tb.v
//
// This is the first testbench that runs the whole engine: the divider of 13.4,
// the mode logic of 13.5, the datapath of 13.6 wrapped by 13.8, the chip-select
// controller of 13.7, and the streaming shadow of 13.9. Only the register
// interface is missing, and that is Chapter 13.11.
//
// The point of the chapter is throughput, so the measurements are about TIME as
// well as data:
//
// - a prompt producer must stream N bytes under ONE chip select with no
// stall cycles at all, and with the inter-frame idle equal to the gap the
// controller was told to insert -- not a cycle more;
// - a deliberately late producer must still be CORRECT, and must be
// reported: the stall counter is how the lost throughput becomes visible;
// - a receiver that does not read fast enough must set a STICKY overrun,
// because a dropped word is lost data, not a slow day.
//
// Data is checked in both directions against an independent slave, byte by
// byte and in order, so a stream that gets every byte right but in the wrong
// sequence fails.
`timescale 1ns/1ps
module spi_multibyte_tb;
localparam MAX_W = 32;
localparam LEN_W = 6;
localparam DIV_W = 8;
localparam N_CS = 2;
localparam SEL_W = 1;
localparam CNT_W = 8;
reg clk;
reg rst_n;
always #5 clk = ~clk;
// --- 13.9, the block under test ---------------------------------------
reg tx_push;
reg [MAX_W-1:0] tx_wdata;
reg tx_last;
wire tx_ready;
reg flush;
reg clr_flags;
reg rx_pop;
wire [MAX_W-1:0] rx_rdata;
wire rx_ready, rx_overrun;
wire req, hold, load_stb, stalled;
wire [MAX_W-1:0] tx_data;
// --- 13.7 -------------------------------------------------------------
reg [SEL_W-1:0] sel;
reg [CNT_W-1:0] lead_cyc;
reg [CNT_W-1:0] lag_cyc;
reg [CNT_W-1:0] gap_cyc;
wire [N_CS-1:0] cs_n;
wire shift_en, start_stb, busy, sel_err;
wire [2:0] state_id;
// --- 13.4 -------------------------------------------------------------
reg [DIV_W-1:0] div;
reg cpol;
wire sclk, edge_a_stb, edge_b_stb, bit_done, div_err;
wire [DIV_W-1:0] half_a, half_b;
// --- 13.5 / 13.8 ------------------------------------------------------
reg cpha;
reg [LEN_W-1:0] len;
reg lsb_first;
wire preload_stb, launch_stb, capture_stb, frame_done;
wire [LEN_W-1:0] bit_idx;
wire miso, mosi;
wire [MAX_W-1:0] rx_word;
wire rx_valid_stb, len_err;
spi_clkdiv_strobe #(.DIV_W(DIV_W)) u_div (
.clk(clk), .rst_n(rst_n), .en(shift_en), .div(div), .cpol(cpol),
.sclk(sclk), .edge_a_stb(edge_a_stb), .edge_b_stb(edge_b_stb),
.bit_done(bit_done), .half_a(half_a), .half_b(half_b),
.div_err(div_err)
);
spi_mode_edges #(.LEN_W(LEN_W)) u_mode (
.clk(clk), .rst_n(rst_n), .cpha(cpha), .len(len),
.active(shift_en), .start_stb(start_stb),
.edge_a_stb(edge_a_stb), .edge_b_stb(edge_b_stb),
.preload_stb(preload_stb), .launch_stb(launch_stb),
.capture_stb(capture_stb), .bit_idx(bit_idx), .frame_done(frame_done)
);
spi_width_order #(.MAX_W(MAX_W), .LEN_W(LEN_W)) u_data (
.clk(clk), .rst_n(rst_n),
.tx_data(tx_data), .len(len), .lsb_first(lsb_first),
.load_stb(load_stb),
.preload_stb(preload_stb), .launch_stb(launch_stb),
.capture_stb(capture_stb),
.miso(miso), .mosi(mosi),
.rx_data(rx_word), .rx_valid_stb(rx_valid_stb), .len_err(len_err)
);
spi_cs_ctrl #(.N_CS(N_CS), .SEL_W(SEL_W), .CNT_W(CNT_W)) u_cs (
.clk(clk), .rst_n(rst_n),
.req(req), .hold(hold), .sel(sel),
.lead_cyc(lead_cyc), .lag_cyc(lag_cyc), .gap_cyc(gap_cyc),
.core_done(frame_done), .abort(1'b0),
.cs_n(cs_n), .shift_en(shift_en), .start_stb(start_stb),
.busy(busy), .sel_err(sel_err), .state_id(state_id)
);
spi_multibyte #(.MAX_W(MAX_W), .LEN_W(LEN_W)) dut (
.clk(clk), .rst_n(rst_n),
.tx_push(tx_push), .tx_wdata(tx_wdata), .tx_last(tx_last),
.tx_ready(tx_ready),
.flush(flush), .clr_flags(clr_flags),
.rx_pop(rx_pop), .rx_rdata(rx_rdata), .rx_ready(rx_ready),
.rx_overrun(rx_overrun),
.core_busy(busy), .shift_en(shift_en), .start_stb(start_stb),
.rx_valid_stb(rx_valid_stb), .rx_word(rx_word),
.req(req), .hold(hold), .tx_data(tx_data), .load_stb(load_stb),
.stalled(stalled)
);
// --- the independent slave, word by word ------------------------------
reg [MAX_W-1:0] slv_tx_q [0:127];
reg [MAX_W-1:0] slv_rx_q [0:127];
integer slv_tx_i, slv_rx_i;
reg [MAX_W-1:0] slv_sr, slv_rx_sr;
reg slave_bit_r;
assign miso = slave_bit_r;
always @(posedge clk) begin
if (load_stb) begin
slv_sr <= slv_tx_q[slv_tx_i] << (MAX_W - len);
slv_tx_i <= slv_tx_i + 1;
end else if (preload_stb || launch_stb) begin
slave_bit_r <= slv_sr[MAX_W-1];
slv_sr <= {slv_sr[MAX_W-2:0], 1'b0};
end
// Cleared at each load, so the recorded word holds THIS frame's bits
// and not a running history of every frame before it.
if (load_stb)
slv_rx_sr <= {MAX_W{1'b0}};
else if (capture_stb)
slv_rx_sr <= {slv_rx_sr[MAX_W-2:0], mosi};
// `rx_valid_stb` is one cycle after the final capture, so the slave's
// own shift register is already complete when it is recorded.
if (rx_valid_stb) begin
slv_rx_q[slv_rx_i] <= slv_rx_sr;
slv_rx_i <= slv_rx_i + 1;
end
end
// --- the throughput monitor -------------------------------------------
integer stall_cycles;
integer idle_in_txn; // cycles inside a transaction with no clock
integer frames_started;
integer cs_falls;
integer slack_cycles; // room in the shadow WHILE the wire is busy
integer valid_pulses;
integer n_loads, n_caps, n_pre, n_lau, n_race;
reg busy_q;
always @(posedge clk) begin
if (!rst_n) begin
busy_q <= 1'b0;
end else begin
busy_q <= busy;
if (stalled) stall_cycles <= stall_cycles + 1;
if (busy && !shift_en) idle_in_txn <= idle_in_txn + 1;
if (start_stb) frames_started <= frames_started + 1;
if (busy && !busy_q) cs_falls <= cs_falls + 1;
if (tx_ready && shift_en) slack_cycles <= slack_cycles + 1;
if (rx_valid_stb) valid_pulses <= valid_pulses + 1;
if (load_stb) n_loads <= n_loads + 1;
if (capture_stb) n_caps <= n_caps + 1;
if (preload_stb) n_pre <= n_pre + 1;
if (launch_stb) n_lau <= n_lau + 1;
// The datapath is loaded by one strobe and told to drive its first
// bit by another. If they ever land on the same cycle the load is
// overwritten by the shift in the same always block and the frame
// sends the PREVIOUS word. The handover is built so they cannot,
// and this counts the cycles on which that claim would be false.
if (load_stb && start_stb) n_race <= n_race + 1;
end
end
integer errors;
task clear_counts;
begin
stall_cycles = 0; idle_in_txn = 0; frames_started = 0;
cs_falls = 0; slack_cycles = 0; valid_pulses = 0;
n_loads = 0; n_caps = 0; n_pre = 0; n_lau = 0; n_race = 0;
end
endtask
// The CPU writing a word. `late` cycles of dithering before the write is
// how a slow driver is modelled.
task push;
input [MAX_W-1:0] w;
input last;
input integer late;
integer guard;
begin
repeat (late) @(negedge clk);
guard = 8000;
while (!tx_ready && guard > 0) begin
@(negedge clk);
guard = guard - 1;
end
if (guard == 0) begin
$display(" FAIL: the shadow never became free");
errors = errors + 1;
end
tx_wdata = w;
tx_last = last;
tx_push = 1'b1;
@(negedge clk);
tx_push = 1'b0;
end
endtask
// The CPU reading. `drain` = 0 leaves words unread on purpose.
integer got_n;
reg [MAX_W-1:0] got_q [0:127];
task drain_rx;
begin
while (rx_ready) begin
got_q[got_n] = rx_rdata;
got_n = got_n + 1;
rx_pop = 1'b1;
@(negedge clk);
rx_pop = 1'b0;
@(negedge clk);
end
end
endtask
task wait_idle;
integer guard;
begin
guard = 20000;
while (busy && guard > 0) begin
drain_rx();
@(negedge clk);
guard = guard - 1;
end
repeat (4) @(negedge clk);
drain_rx();
end
endtask
reg [MAX_W-1:0] sent_q [0:127];
integer i, k, nb, late, seed, bad;
initial begin
clear_counts();
slv_tx_i = 0; slv_rx_i = 0; got_n = 0;
slv_sr = 32'h0; slv_rx_sr = 32'h0; slave_bit_r = 1'b0;
seed = 32'h5EED_1234;
for (i = 0; i < 128; i = i + 1) begin
seed = (seed * 32'h0019_660D) + 32'h3C6E_F35F;
slv_tx_q[i] = (seed >> 7) & 32'hFF;
end
repeat (3) @(negedge clk);
rst_n = 1'b1;
@(negedge clk);
// 1. A PROMPT PRODUCER, sixteen bytes, one chip select. This is the
// case the shadow register exists for.
clear_counts();
nb = 16;
for (i = 0; i < nb; i = i + 1) begin
seed = (seed * 32'h0019_660D) + 32'h3C6E_F35F;
sent_q[i] = (seed >> 11) & 32'hFF;
end
got_n = 0;
for (i = 0; i < nb; i = i + 1) begin
push(sent_q[i], (i == nb - 1), 0);
drain_rx();
end
wait_idle();
if (cs_falls != 1) begin
$display(" FAIL: a %0d-byte stream took %0d chip selects",
nb, cs_falls);
errors = errors + 1;
end
if (frames_started != nb) begin
$display(" FAIL: %0d bytes produced %0d frames", nb,
frames_started);
errors = errors + 1;
end
if (stall_cycles != 0) begin
$display(" FAIL: a prompt producer still stalled the wire for %0d cycles",
stall_cycles);
errors = errors + 1;
end
// Everything the slave received, in order.
bad = 0;
for (i = 0; i < nb; i = i + 1)
if (slv_rx_q[i] !== sent_q[i]) bad = bad + 1;
if (bad != 0) begin
$display(" FAIL: the slave received %0d of %0d bytes wrongly",
bad, nb);
errors = errors + 1;
end
// Everything the master received, in order.
bad = 0;
if (got_n != nb) begin
$display(" FAIL: the master read back %0d of %0d bytes",
got_n, nb);
errors = errors + 1;
end else begin
for (i = 0; i < nb; i = i + 1)
if (got_q[i] !== slv_tx_q[i]) bad = bad + 1;
if (bad != 0) begin
$display(" FAIL: the master received %0d of %0d bytes wrongly",
bad, nb);
errors = errors + 1;
end
end
if (rx_overrun) begin
$display(" FAIL: a fully drained receiver reported an overrun");
errors = errors + 1;
end
// The per-frame bookkeeping, counted over the whole stream rather
// than trusted: one load and one preload per frame, len captures per
// frame, and len-1 launches because CPHA=0's preload is the first
// drive (Chapter 13.5).
if (n_loads != nb || n_pre != nb) begin
$display(" FAIL: %0d frames produced %0d loads and %0d preloads",
nb, n_loads, n_pre);
errors = errors + 1;
end
if (n_caps != nb * 8 || n_lau != nb * 7) begin
$display(" FAIL: %0d eight-bit frames produced %0d captures and %0d launches, expected %0d and %0d",
nb, n_caps, n_lau, nb * 8, nb * 7);
errors = errors + 1;
end
if (n_race != 0) begin
$display(" FAIL: a load landed on a frame start %0d times",
n_race);
errors = errors + 1;
end
if (valid_pulses != nb) begin
$display(" FAIL: %0d frames raised %0d received-word pulses",
nb, valid_pulses);
errors = errors + 1;
end
$display(" per-frame bookkeeping over the stream: %0d loads, %0d preloads, %0d launches, %0d captures, %0d received words, and no load ever landed on a frame start",
n_loads, n_pre, n_lau, n_caps, valid_pulses);
$display(" 16 bytes, one chip select, %0d frames, %0d stall cycles, %0d idle cycles inside the transaction -- %0d per byte boundary",
frames_started, stall_cycles, idle_in_txn,
idle_in_txn / (nb - 1));
// 2. THE IDLE BETWEEN FRAMES IS THE GAP THAT WAS ASKED FOR, and not a
// cycle more. This is the check that would catch a shadow register
// that works but hands over one cycle late.
if (idle_in_txn > (nb - 1) * (gap_cyc + 3)) begin
$display(" FAIL: %0d idle cycles across %0d boundaries, more than the %0d-cycle gap explains",
idle_in_txn, nb - 1, gap_cyc);
errors = errors + 1;
end
$display(" the inter-frame idle is the programmed %0d-cycle hold and nothing else",
gap_cyc);
// 3. A LATE PRODUCER. Still correct, and now visibly costly.
clear_counts();
nb = 8;
for (i = 0; i < nb; i = i + 1) begin
seed = (seed * 32'h0019_660D) + 32'h3C6E_F35F;
sent_q[i] = (seed >> 11) & 32'hFF;
end
k = slv_tx_i;
got_n = 0;
for (i = 0; i < nb; i = i + 1) begin
push(sent_q[i], (i == nb - 1), 40);
drain_rx();
end
wait_idle();
if (stall_cycles == 0) begin
$display(" FAIL: a producer 40 cycles late stalled the wire not at all -- the counter cannot be working");
errors = errors + 1;
end
bad = 0;
for (i = 0; i < nb; i = i + 1)
if (slv_rx_q[k + i] !== sent_q[i]) bad = bad + 1;
if (bad != 0) begin
$display(" FAIL: a late producer corrupted %0d of %0d bytes",
bad, nb);
errors = errors + 1;
end
if (cs_falls != 1) begin
$display(" FAIL: a late producer broke the transaction into %0d chip selects",
cs_falls);
errors = errors + 1;
end
$display(" the same stream with the producer 40 cycles late: every byte still correct, one chip select, but %0d stall cycles the fast run did not have",
stall_cycles);
// 4. AN OVERRUN IS STICKY. Stream four bytes and never read one.
clear_counts();
nb = 4;
for (i = 0; i < nb; i = i + 1)
push(32'h5A, (i == nb - 1), 0);
// Deliberately no drain_rx here.
i = 20000;
while (busy && i > 0) begin
@(negedge clk);
i = i - 1;
end
repeat (4) @(negedge clk);
if (!rx_overrun) begin
$display(" FAIL: four unread received words did not raise an overrun");
errors = errors + 1;
end
$display(" four received words left unread: overrun raised and latched");
// It must STAY raised through a clean transfer -- a flag that clears
// itself is a flag that hides the one byte that was lost.
drain_rx();
clear_counts();
for (i = 0; i < 2; i = i + 1) begin
push(32'h33, (i == 1), 0);
drain_rx();
end
wait_idle();
if (!rx_overrun) begin
$display(" FAIL: the overrun flag cleared itself on the next clean transfer");
errors = errors + 1;
end
$display(" a clean transfer afterwards does not clear it -- the lost byte stays reported");
// 5. THE CPU CAN REFILL WHILE THE WIRE IS BUSY. Without that the
// shadow buys nothing, so it is checked directly: at least one push
// must be accepted with the clock running.
clear_counts();
// Reset the sticky flag by resetting the block, since nothing else
// clears it -- which is itself the design decision under test.
rst_n = 1'b0;
repeat (3) @(negedge clk);
rst_n = 1'b1;
repeat (2) @(negedge clk);
slv_tx_i = 0; slv_rx_i = 0; got_n = 0;
// What matters is not the instant the testbench happens to write, but
// HOW LONG the shadow stays free while the wire is busy -- that window
// is the slack the producer is given, and without the shadow it would
// be zero.
for (i = 0; i < 8; i = i + 1) begin
push(32'hC5, (i == 7), 0);
drain_rx();
end
wait_idle();
if (slack_cycles == 0) begin
$display(" FAIL: the shadow was never free while the clock was running -- it is not decoupling anything");
errors = errors + 1;
end
if (stall_cycles != 0) begin
$display(" FAIL: %0d stall cycles while refilling ahead",
stall_cycles);
errors = errors + 1;
end
$display(" the shadow stood free for %0d of the cycles the wire was busy -- that window is the slack the producer gets, and it is a whole frame wide",
slack_cycles);
if (errors == 0)
$display("PASS: a one-deep shadow register decouples the producer from the wire -- sixteen bytes stream under a single chip select with every byte correct in both directions and in order, zero stall cycles, and an inter-frame idle equal to the programmed hold and nothing more -- words are accepted while the previous frame is still shifting, which is the whole reason the shadow exists -- a producer forty cycles late still transfers every byte correctly under one chip select but now reports the stall cycles it cost, so the lost throughput is visible instead of silent -- and a received word overwritten before it was read raises a sticky overrun that a later clean transfer does not clear");
else
$display("FAIL: %0d error(s)", errors);
$finish;
end
initial begin
clk = 1'b0;
rst_n = 1'b0;
tx_push = 1'b0;
tx_wdata = 32'h0;
tx_last = 1'b0;
flush = 1'b0;
clr_flags = 1'b0;
rx_pop = 1'b0;
sel = 1'b0;
lead_cyc = 8'd3;
lag_cyc = 8'd3;
gap_cyc = 8'd2;
div = 8'd4;
cpol = 1'b0;
cpha = 1'b0;
len = 6'd8;
lsb_first = 1'b0;
errors = 0;
end
endmodule-- spi_multibyte.vhd
--
-- Chapter 13.9 -- keeping the wire busy between frames.
--
-- Everything built so far transfers ONE frame correctly. A real transaction is
-- a stream: a flash page program is a command byte, three address bytes and
-- then 256 data bytes, all under one chip select. The gap between those bytes
-- is where the throughput of Chapter 9 is won or lost.
--
-- THE PROBLEM WITH A SINGLE TRANSMIT REGISTER. If the datapath's shift
-- register is the only place a word can live, the next byte cannot be written
-- until the current one has left -- and "has left" means the frame is over:
--
-- shift byte N ... frame ends ... CPU notices ... CPU writes byte N+1
-- ... controller starts ... shift byte N+1
--
-- Every one of those dots is dead time on SCLK, and none of it is under the
-- hardware's control. The transfer is correct and the wire is half idle.
--
-- THE FIX IS ONE REGISTER DEEP. A SHADOW register in front of the datapath:
-- the CPU writes the shadow, the controller moves the shadow into the datapath
-- between frames. The CPU now has an entire FRAME to produce the next word
-- instead of an inter-frame gap.
--
-- One deep is enough, and that is worth being clear about: a deeper FIFO buys
-- latency TOLERANCE, not throughput. Once the producer can stay one word
-- ahead, the wire is saturated and nothing deeper makes it faster.
--
-- WHAT THIS BLOCK REPORTS. A design that silently absorbs a slow producer is
-- worse than one that says so, because the lost throughput is invisible:
--
-- stalled the transaction is open, the clock is not running, and nothing
-- is queued -- the wire is idle because the CPU was late.
-- rx_overrun a received word was overwritten before it was read. Unlike a
-- stall this is LOST DATA, and the flag is sticky: a receive
-- path that drops a word and carries on is how a flash image
-- ends up with one wrong word in it.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_multibyte is
generic (
MAX_W : positive := 32;
LEN_W : positive := 6
);
port (
clk : in std_logic;
rst_n : in std_logic;
-- CPU side
tx_push : in std_logic; -- write a word into the shadow
tx_wdata : in std_logic_vector(MAX_W - 1 downto 0);
tx_last : in std_logic; -- .. and end the transaction after it
tx_ready : out std_logic; -- the shadow is free
flush : in std_logic; -- drop whatever is queued
rx_pop : in std_logic; -- read the received word
rx_rdata : out std_logic_vector(MAX_W - 1 downto 0);
rx_ready : out std_logic; -- a received word is waiting
rx_overrun : out std_logic; -- sticky: a word was lost
-- Clearing a sticky flag is the REGISTER INTERFACE's business, not this
-- block's -- but the flag lives here, because this is where the loss is
-- detected. So each flag has exactly one owner and one clear path, and
-- Chapter 13.11 wires this to a write of the command register.
clr_flags : in std_logic;
-- transfer engine side (13.7 and 13.6/13.8)
core_busy : in std_logic; -- a transaction is open
shift_en : in std_logic; -- the clock is running
start_stb : in std_logic; -- the engine took a frame
rx_valid_stb : in std_logic;
rx_word : in std_logic_vector(MAX_W - 1 downto 0);
req : out std_logic; -- to the chip-select controller
hold : out std_logic;
tx_data : out std_logic_vector(MAX_W - 1 downto 0);
load_stb : out std_logic;
stalled : out std_logic -- the wire is idle for want of data
);
end entity;
architecture rtl of spi_multibyte is
-- The shadow: one word, its end-of-transaction flag, and a valid bit.
signal sh_data : std_logic_vector(MAX_W - 1 downto 0) := (others => '0');
signal sh_last : std_logic := '0';
signal sh_valid : std_logic := '0';
-- `armed` means the datapath has been loaded and the controller has been
-- asked for a frame, but the frame has not started yet. It stops a second
-- load from landing on top of the first.
signal armed : std_logic := '0';
signal armed_last : std_logic := '0';
signal load_r : std_logic := '0';
-- `tail` means the final frame of the transaction has started, so the
-- controller is walking out through its lag and gap with nothing more to
-- send. Without it the closing states look exactly like a producer that
-- went quiet, and `stalled` would accuse the CPU of being late at the end
-- of every single transaction.
signal tail : std_logic := '0';
signal tx_data_r : std_logic_vector(MAX_W - 1 downto 0) := (others => '0');
signal rx_hold : std_logic_vector(MAX_W - 1 downto 0) := (others => '0');
signal rx_full : std_logic := '0';
signal ovr_r : std_logic := '0';
signal can_load : std_logic;
signal ready_i : std_logic;
begin
ready_i <= not sh_valid;
tx_ready <= ready_i;
req <= armed;
hold <= not armed_last;
load_stb <= load_r;
tx_data <= tx_data_r;
rx_rdata <= rx_hold;
rx_ready <= rx_full;
rx_overrun <= ovr_r;
-- The datapath may only be loaded while the clock is stopped. Loading it
-- mid-frame would overwrite the word being shifted -- the one failure a
-- single-register design cannot even express, and that a shadow design has
-- to be careful about.
can_load <= sh_valid and (not armed) and (not shift_en) and (not flush);
-- Idle inside an open transaction with nothing to send. Note it is NOT an
-- error: it is the exact measure of how late the producer was.
stalled <= core_busy and (not shift_en) and (not armed) and
(not sh_valid) and (not tail);
seq : process (clk, rst_n)
begin
if rst_n = '0' then
sh_data <= (others => '0');
sh_last <= '0';
sh_valid <= '0';
armed <= '0';
armed_last <= '0';
load_r <= '0';
tail <= '0';
tx_data_r <= (others => '0');
rx_hold <= (others => '0');
rx_full <= '0';
ovr_r <= '0';
elsif rising_edge(clk) then
load_r <= '0';
-- The CPU writing into the shadow. Accepted whenever the shadow is
-- free, which includes the same cycle it is being emptied into the
-- datapath below: the two are ordered so a push and a load on one
-- cycle both succeed.
-- FLUSH takes priority over everything on the transmit side. A
-- word queued for a transaction that is being abandoned must not
-- become the first word of the next one, and an `armed` request
-- left standing would start that next transaction by itself the
-- moment the bus came free. Chapter 13.10 drives this; here it is
-- simply the one input that can throw work away.
if flush = '1' then
sh_valid <= '0';
armed <= '0';
armed_last <= '0';
tail <= '0';
elsif tx_push = '1' and ready_i = '1' then
sh_data <= tx_wdata;
sh_last <= tx_last;
sh_valid <= '1';
end if;
-- The shadow moving into the datapath.
if can_load = '1' and flush = '0' then
tx_data_r <= sh_data;
armed_last <= sh_last;
armed <= '1';
load_r <= '1';
tail <= '0';
-- Freed immediately, so the CPU may refill it on the very next
-- cycle rather than waiting for the frame to finish. This
-- single line is what makes the stream back-to-back.
if not (tx_push = '1' and ready_i = '1') then
sh_valid <= '0';
end if;
end if;
-- The engine taking the frame.
if start_stb = '1' then
armed <= '0';
if armed_last = '1' then
tail <= '1';
end if;
end if;
-- The receive side. The clear is applied FIRST so that a clear
-- arriving on the same cycle as a fresh overrun does not swallow
-- it: the new event wins, because losing a report of lost data is
-- the same bug twice.
if clr_flags = '1' then
ovr_r <= '0';
end if;
if rx_valid_stb = '1' then
rx_hold <= rx_word;
rx_full <= '1';
-- A word arriving on top of an unread one is DATA LOST, and
-- the flag is sticky so a driver that only checks at the end
-- of the transaction still finds out.
if rx_full = '1' and rx_pop = '0' then
ovr_r <= '1';
end if;
elsif rx_pop = '1' then
rx_full <= '0';
end if;
end if;
end process;
end architecture;-- spi_multibyte_tb.vhd
--
-- This is the first testbench that runs the whole engine: the divider of 13.4,
-- the mode logic of 13.5, the datapath of 13.6 wrapped by 13.8, the chip-select
-- controller of 13.7, and the streaming shadow of 13.9. Only the register
-- interface is missing, and that is Chapter 13.11.
--
-- The point of the chapter is throughput, so the measurements are about TIME as
-- well as data:
--
-- - a prompt producer must stream N words under ONE chip select with no stall
-- cycles at all, and with the inter-frame idle equal to the gap the
-- controller was told to insert -- not a cycle more;
-- - a deliberately late producer must still be CORRECT, and must be reported;
-- - a receiver that does not read fast enough must set a STICKY overrun.
--
-- Data is checked in both directions against an independent slave, word by word
-- and in order, so a stream that gets every word right but in the wrong
-- sequence fails.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_multibyte_tb is
end entity;
architecture sim of spi_multibyte_tb is
constant MAX_W : positive := 32;
constant LEN_W : positive := 6;
constant DIV_W : positive := 8;
constant N_CS : positive := 2;
constant SEL_W : positive := 1;
constant CNT_W : positive := 8;
type word_array is array (0 to 127) of std_logic_vector(MAX_W - 1 downto 0);
signal clk : std_logic := '0';
signal rst_n : std_logic := '0';
signal halt : boolean := false;
-- 13.9, the block under test
signal tx_push : std_logic := '0';
signal tx_wdata : std_logic_vector(MAX_W - 1 downto 0) := (others => '0');
signal tx_last : std_logic := '0';
signal tx_ready : std_logic;
signal flush : std_logic := '0';
signal clr_flags : std_logic := '0';
signal rx_pop : std_logic := '0';
signal rx_rdata : std_logic_vector(MAX_W - 1 downto 0);
signal rx_ready, rx_overrun : std_logic;
signal req, hold, load_stb, stalled : std_logic;
signal tx_data : std_logic_vector(MAX_W - 1 downto 0);
-- 13.7
signal sel : unsigned(SEL_W - 1 downto 0) := (others => '0');
signal lead_cyc : unsigned(CNT_W - 1 downto 0) := to_unsigned(3, CNT_W);
signal lag_cyc : unsigned(CNT_W - 1 downto 0) := to_unsigned(3, CNT_W);
signal gap_cyc : unsigned(CNT_W - 1 downto 0) := to_unsigned(2, CNT_W);
signal cs_n : std_logic_vector(N_CS - 1 downto 0);
signal shift_en, start_stb, busy, sel_err : std_logic;
signal state_id : unsigned(2 downto 0);
-- 13.4
signal div : unsigned(DIV_W - 1 downto 0) := to_unsigned(4, DIV_W);
signal cpol : std_logic := '0';
signal sclk, edge_a_stb, edge_b_stb, bit_done, div_err : std_logic;
signal half_a, half_b : unsigned(DIV_W - 1 downto 0);
-- 13.5 / 13.8
signal cpha : std_logic := '0';
signal len : unsigned(LEN_W - 1 downto 0) := to_unsigned(8, LEN_W);
signal lsb_first : std_logic := '0';
signal preload_stb, launch_stb, capture_stb, frame_done : std_logic;
signal bit_idx : unsigned(LEN_W - 1 downto 0);
signal miso, mosi : std_logic;
signal rx_word : std_logic_vector(MAX_W - 1 downto 0);
signal rx_valid_stb, len_err : std_logic;
-- the independent slave, word by word
signal slv_tx_q : word_array := (others => (others => '0'));
signal slv_rx_q : word_array := (others => (others => '0'));
signal slv_tx_i : natural := 0;
signal slv_rx_i : natural := 0;
signal slv_sr, slv_rx_sr : std_logic_vector(MAX_W - 1 downto 0)
:= (others => '0');
signal slave_bit_r : std_logic := '0';
signal slv_rst : std_logic := '0';
-- the throughput monitor
signal stall_cycles, idle_in_txn, frames_started, cs_falls : natural := 0;
signal slack_cycles, valid_pulses : natural := 0;
signal n_loads, n_caps, n_pre, n_lau, n_race : natural := 0;
signal busy_q : std_logic := '0';
signal clear_stb : std_logic := '0';
signal errors : natural := 0;
function hex8(v : std_logic_vector) return string is
constant DIGITS : string(1 to 16) := "0123456789abcdef";
variable u : unsigned(MAX_W - 1 downto 0);
variable r : string(1 to 8);
begin
u := unsigned(v);
for k in 8 downto 1 loop
r(k) := DIGITS(to_integer(u(3 downto 0)) + 1);
u := shift_right(u, 4);
end loop;
return r;
end function;
begin
clk <= not clk after 5 ns when not halt else '0';
u_div : entity work.spi_clkdiv_strobe
generic map (DIV_W => DIV_W)
port map (clk => clk, rst_n => rst_n, en => shift_en, div => div,
cpol => cpol, sclk => sclk,
edge_a_stb => edge_a_stb, edge_b_stb => edge_b_stb,
bit_done => bit_done, half_a => half_a, half_b => half_b,
div_err => div_err);
u_mode : entity work.spi_mode_edges
generic map (LEN_W => LEN_W)
port map (clk => clk, rst_n => rst_n, cpha => cpha, len => len,
active => shift_en, start_stb => start_stb,
edge_a_stb => edge_a_stb, edge_b_stb => edge_b_stb,
preload_stb => preload_stb, launch_stb => launch_stb,
capture_stb => capture_stb, bit_idx => bit_idx,
frame_done => frame_done);
u_data : entity work.spi_width_order
generic map (MAX_W => MAX_W, LEN_W => LEN_W)
port map (clk => clk, rst_n => rst_n,
tx_data => tx_data, len => len, lsb_first => lsb_first,
load_stb => load_stb,
preload_stb => preload_stb, launch_stb => launch_stb,
capture_stb => capture_stb,
miso => miso, mosi => mosi,
rx_data => rx_word, rx_valid_stb => rx_valid_stb,
len_err => len_err);
u_cs : entity work.spi_cs_ctrl
generic map (N_CS => N_CS, SEL_W => SEL_W, CNT_W => CNT_W)
port map (clk => clk, rst_n => rst_n,
req => req, hold => hold, sel => sel,
lead_cyc => lead_cyc, lag_cyc => lag_cyc, gap_cyc => gap_cyc,
core_done => frame_done, abort => '0',
cs_n => cs_n, shift_en => shift_en, start_stb => start_stb,
busy => busy, sel_err => sel_err, state_id => state_id);
dut : entity work.spi_multibyte
generic map (MAX_W => MAX_W, LEN_W => LEN_W)
port map (clk => clk, rst_n => rst_n,
tx_push => tx_push, tx_wdata => tx_wdata, tx_last => tx_last,
tx_ready => tx_ready,
flush => flush, clr_flags => clr_flags,
rx_pop => rx_pop, rx_rdata => rx_rdata, rx_ready => rx_ready,
rx_overrun => rx_overrun,
core_busy => busy, shift_en => shift_en,
start_stb => start_stb,
rx_valid_stb => rx_valid_stb, rx_word => rx_word,
req => req, hold => hold, tx_data => tx_data,
load_stb => load_stb,
stalled => stalled);
miso <= slave_bit_r;
slave_p : process (clk)
begin
if rising_edge(clk) then
if slv_rst = '1' then
slv_tx_i <= 0;
slv_rx_i <= 0;
slv_sr <= (others => '0');
slv_rx_sr <= (others => '0');
slave_bit_r <= '0';
else
if load_stb = '1' then
slv_sr <= std_logic_vector(
shift_left(unsigned(slv_tx_q(slv_tx_i)),
MAX_W - to_integer(len)));
slv_tx_i <= slv_tx_i + 1;
elsif preload_stb = '1' or launch_stb = '1' then
slave_bit_r <= slv_sr(MAX_W - 1);
slv_sr <= slv_sr(MAX_W - 2 downto 0) & '0';
end if;
-- Cleared at each load, so the recorded word holds THIS
-- frame's bits and not a running history of every frame
-- before it.
if load_stb = '1' then
slv_rx_sr <= (others => '0');
elsif capture_stb = '1' then
slv_rx_sr <= slv_rx_sr(MAX_W - 2 downto 0) & mosi;
end if;
-- `rx_valid_stb` is one cycle after the final capture, so the
-- slave's own shift register is already complete here.
if rx_valid_stb = '1' then
slv_rx_q(slv_rx_i) <= slv_rx_sr;
slv_rx_i <= slv_rx_i + 1;
end if;
end if;
end if;
end process;
monitor : process (clk)
begin
if rising_edge(clk) then
if rst_n = '1' then
busy_q <= busy;
if clear_stb = '1' then
stall_cycles <= 0; idle_in_txn <= 0; frames_started <= 0;
cs_falls <= 0; slack_cycles <= 0; valid_pulses <= 0;
n_loads <= 0; n_caps <= 0; n_pre <= 0; n_lau <= 0;
n_race <= 0;
else
if stalled = '1' then
stall_cycles <= stall_cycles + 1;
end if;
if busy = '1' and shift_en = '0' then
idle_in_txn <= idle_in_txn + 1;
end if;
if start_stb = '1' then
frames_started <= frames_started + 1;
end if;
if busy = '1' and busy_q = '0' then
cs_falls <= cs_falls + 1;
end if;
if tx_ready = '1' and shift_en = '1' then
slack_cycles <= slack_cycles + 1;
end if;
if rx_valid_stb = '1' then
valid_pulses <= valid_pulses + 1;
end if;
if load_stb = '1' then n_loads <= n_loads + 1; end if;
if capture_stb = '1' then n_caps <= n_caps + 1; end if;
if preload_stb = '1' then n_pre <= n_pre + 1; end if;
if launch_stb = '1' then n_lau <= n_lau + 1; end if;
-- The datapath is loaded by one strobe and told to drive
-- its first bit by another. If they ever land on the same
-- cycle the load is overwritten by the shift in the same
-- process and the frame sends the PREVIOUS word. The
-- handover is built so they cannot, and this counts the
-- cycles on which that claim would be false.
if load_stb = '1' and start_stb = '1' then
n_race <= n_race + 1;
end if;
end if;
end if;
end if;
end process;
stim : process
variable errs : natural := 0;
variable guard : natural;
variable seed : unsigned(31 downto 0) := x"5EED1234";
variable sent_q : word_array;
variable got_q : word_array;
variable got_n : natural := 0;
variable nb, k, bad : natural;
variable accepted_free : natural;
procedure clear_counts is
begin
clear_stb <= '1';
wait until falling_edge(clk);
clear_stb <= '0';
end procedure;
procedure next_rand(variable v : out unsigned(31 downto 0)) is
begin
seed := resize(seed * x"0019660D", 32) + x"3C6EF35F";
v := seed;
end procedure;
-- The CPU reading. Leaving this uncalled is how an overrun is made.
procedure drain_rx is
begin
while rx_ready = '1' loop
got_q(got_n) := rx_rdata;
got_n := got_n + 1;
rx_pop <= '1';
wait until falling_edge(clk);
rx_pop <= '0';
wait until falling_edge(clk);
end loop;
end procedure;
-- The CPU writing a word. `late` cycles of dithering before the write
-- is how a slow driver is modelled.
procedure push(w : std_logic_vector(MAX_W - 1 downto 0);
last : std_logic; late : natural) is
begin
for j in 1 to late loop wait until falling_edge(clk); end loop;
guard := 8000;
while tx_ready = '0' and guard > 0 loop
wait until falling_edge(clk);
guard := guard - 1;
end loop;
if guard = 0 then
report " FAIL: the shadow never became free";
errs := errs + 1;
end if;
tx_wdata <= w;
tx_last <= last;
tx_push <= '1';
wait until falling_edge(clk);
tx_push <= '0';
end procedure;
procedure wait_idle is
begin
guard := 20000;
while busy = '1' and guard > 0 loop
drain_rx;
wait until falling_edge(clk);
guard := guard - 1;
end loop;
for j in 1 to 4 loop wait until falling_edge(clk); end loop;
drain_rx;
end procedure;
variable r : unsigned(31 downto 0);
begin
for i in 0 to 127 loop
next_rand(r);
slv_tx_q(i) <= std_logic_vector(resize(shift_right(r, 7) and x"000000FF",
MAX_W));
end loop;
for k2 in 1 to 3 loop wait until falling_edge(clk); end loop;
rst_n <= '1';
wait until falling_edge(clk);
-- 1. A PROMPT PRODUCER, sixteen words, one chip select. This is the
-- case the shadow register exists for.
clear_counts;
nb := 16;
for i in 0 to nb - 1 loop
next_rand(r);
sent_q(i) := std_logic_vector(resize(shift_right(r, 11) and x"000000FF",
MAX_W));
end loop;
got_n := 0;
for i in 0 to nb - 1 loop
if i = nb - 1 then push(sent_q(i), '1', 0);
else push(sent_q(i), '0', 0); end if;
drain_rx;
end loop;
wait_idle;
if cs_falls /= 1 then
report " FAIL: a " & integer'image(nb) & "-word stream took " &
integer'image(cs_falls) & " chip selects";
errs := errs + 1;
end if;
if frames_started /= nb then
report " FAIL: " & integer'image(nb) & " words produced " &
integer'image(frames_started) & " frames";
errs := errs + 1;
end if;
if stall_cycles /= 0 then
report " FAIL: a prompt producer still stalled the wire for " &
integer'image(stall_cycles) & " cycles";
errs := errs + 1;
end if;
bad := 0;
for i in 0 to nb - 1 loop
if slv_rx_q(i) /= sent_q(i) then bad := bad + 1; end if;
end loop;
if bad /= 0 then
report " FAIL: the slave received " & integer'image(bad) & " of " &
integer'image(nb) & " words wrongly";
errs := errs + 1;
end if;
bad := 0;
if got_n /= nb then
report " FAIL: the master read back " & integer'image(got_n) &
" of " & integer'image(nb) & " words";
errs := errs + 1;
else
for i in 0 to nb - 1 loop
if got_q(i) /= slv_tx_q(i) then bad := bad + 1; end if;
end loop;
if bad /= 0 then
report " FAIL: the master received " & integer'image(bad) &
" of " & integer'image(nb) & " words wrongly";
errs := errs + 1;
end if;
end if;
if rx_overrun = '1' then
report " FAIL: a fully drained receiver reported an overrun";
errs := errs + 1;
end if;
-- The per-frame bookkeeping, counted over the whole stream rather than
-- trusted: one load and one preload per frame, len captures per frame,
-- and len-1 launches because CPHA=0's preload is the first drive.
if n_loads /= nb or n_pre /= nb then
report " FAIL: " & integer'image(nb) & " frames produced " &
integer'image(n_loads) & " loads and " &
integer'image(n_pre) & " preloads";
errs := errs + 1;
end if;
if n_caps /= nb * 8 or n_lau /= nb * 7 then
report " FAIL: " & integer'image(nb) &
" eight-bit frames produced " & integer'image(n_caps) &
" captures and " & integer'image(n_lau) & " launches";
errs := errs + 1;
end if;
if n_race /= 0 then
report " FAIL: a load landed on a frame start";
errs := errs + 1;
end if;
if valid_pulses /= nb then
report " FAIL: " & integer'image(nb) & " frames raised " &
integer'image(valid_pulses) & " received-word pulses";
errs := errs + 1;
end if;
report " per-frame bookkeeping over the stream: " &
integer'image(n_loads) & " loads, " & integer'image(n_pre) &
" preloads, " & integer'image(n_lau) & " launches, " &
integer'image(n_caps) & " captures, " &
integer'image(valid_pulses) &
" received words, and no load ever landed on a frame start";
report " 16 words, one chip select, " &
integer'image(frames_started) & " frames, " &
integer'image(stall_cycles) & " stall cycles, " &
integer'image(idle_in_txn) &
" idle cycles inside the transaction -- " &
integer'image(idle_in_txn / (nb - 1)) & " per word boundary";
-- 2. THE IDLE BETWEEN FRAMES IS THE GAP THAT WAS ASKED FOR, and not a
-- cycle more. This would catch a shadow register that works but
-- hands over one cycle late.
if idle_in_txn > (nb - 1) * (to_integer(gap_cyc) + 3) then
report " FAIL: more idle cycles than the programmed gap explains";
errs := errs + 1;
end if;
report " the inter-frame idle is the programmed " &
integer'image(to_integer(gap_cyc)) &
"-cycle hold and nothing else";
-- 3. A LATE PRODUCER. Still correct, and now visibly costly.
clear_counts;
nb := 8;
for i in 0 to nb - 1 loop
next_rand(r);
sent_q(i) := std_logic_vector(resize(shift_right(r, 11) and x"000000FF",
MAX_W));
end loop;
k := slv_rx_i;
got_n := 0;
for i in 0 to nb - 1 loop
if i = nb - 1 then push(sent_q(i), '1', 40);
else push(sent_q(i), '0', 40); end if;
drain_rx;
end loop;
wait_idle;
if stall_cycles = 0 then
report " FAIL: a producer 40 cycles late stalled the wire not at all -- the counter cannot be working";
errs := errs + 1;
end if;
bad := 0;
for i in 0 to nb - 1 loop
if slv_rx_q(k + i) /= sent_q(i) then bad := bad + 1; end if;
end loop;
if bad /= 0 then
report " FAIL: a late producer corrupted " & integer'image(bad) &
" of " & integer'image(nb) & " words";
errs := errs + 1;
end if;
if cs_falls /= 1 then
report " FAIL: a late producer broke the transaction into " &
integer'image(cs_falls) & " chip selects";
errs := errs + 1;
end if;
report " the same stream with the producer 40 cycles late: every word still correct, one chip select, but " &
integer'image(stall_cycles) &
" stall cycles the fast run did not have";
-- 4. AN OVERRUN IS STICKY. Stream four words and never read one.
clear_counts;
nb := 4;
for i in 0 to nb - 1 loop
if i = nb - 1 then push(x"0000005A", '1', 0);
else push(x"0000005A", '0', 0); end if;
end loop;
guard := 20000;
while busy = '1' and guard > 0 loop
wait until falling_edge(clk);
guard := guard - 1;
end loop;
for j in 1 to 4 loop wait until falling_edge(clk); end loop;
if rx_overrun /= '1' then
report " FAIL: four unread received words did not raise an overrun";
errs := errs + 1;
end if;
report " four received words left unread: overrun raised and latched";
-- It must STAY raised through a clean transfer -- a flag that clears
-- itself is a flag that hides the one word that was lost.
drain_rx;
clear_counts;
for i in 0 to 1 loop
if i = 1 then push(x"00000033", '1', 0);
else push(x"00000033", '0', 0); end if;
drain_rx;
end loop;
wait_idle;
if rx_overrun /= '1' then
report " FAIL: the overrun flag cleared itself on the next clean transfer";
errs := errs + 1;
end if;
report " a clean transfer afterwards does not clear it -- the lost word stays reported";
-- 5. HOW LONG THE SHADOW STAYS FREE WHILE THE WIRE IS BUSY. That
-- window is the slack the producer is given, and without the shadow
-- it would be zero. Reset first, because nothing else clears the
-- sticky overrun -- which is itself the design decision under test.
rst_n <= '0';
slv_rst <= '1';
for j in 1 to 3 loop wait until falling_edge(clk); end loop;
rst_n <= '1';
slv_rst <= '0';
for j in 1 to 2 loop wait until falling_edge(clk); end loop;
clear_counts;
got_n := 0;
for i in 0 to 7 loop
if i = 7 then push(x"000000C5", '1', 0);
else push(x"000000C5", '0', 0); end if;
drain_rx;
end loop;
wait_idle;
if slack_cycles = 0 then
report " FAIL: the shadow was never free while the clock was running -- it is not decoupling anything";
errs := errs + 1;
end if;
if stall_cycles /= 0 then
report " FAIL: " & integer'image(stall_cycles) &
" stall cycles while refilling ahead";
errs := errs + 1;
end if;
report " the shadow stood free for " & integer'image(slack_cycles) &
" of the cycles the wire was busy -- that window is the slack the producer gets, and it is a whole frame wide";
errors <= errs;
if errs = 0 then
report "PASS: a one-deep shadow register decouples the producer from the wire -- sixteen words stream under a single chip select with every word correct in both directions and in order, zero stall cycles, and an inter-frame idle equal to the programmed hold and nothing more -- words are accepted while the previous frame is still shifting, which is the whole reason the shadow exists -- a producer forty cycles late still transfers every word correctly under one chip select but now reports the stall cycles it cost, so the lost throughput is visible instead of silent -- and a received word overwritten before it was read raises a sticky overrun that a later clean transfer does not clear";
else
report "FAIL: " & integer'image(errs) & " error(s)" severity error;
end if;
halt <= true;
wait;
end process;
end architecture;Parity
All three implementations stream sixteen words under a single chip select with zero stall cycles, 38 idle cycles across fifteen boundaries — the programmed two-cycle hold and nothing more — and exactly 16 loads, 16 preloads, 112 launches and 128 captures. The same stream with the producer forty cycles late transfers every word correctly under one chip select and reports 65 stall cycles that the fast run did not have. And the shadow stands free for 31 of the cycles the wire is busy, which at a divisor of four and an eight-bit frame is a whole frame wide.
7. Why a Verification Engineer Cares
// The properties divide cleanly: correctness properties about the handover, and
// a performance property about the slack. The second is unusual in an assertion
// set and belongs there -- a design that is correct and slow fails its
// requirement just as surely as one that is fast and wrong.
module spi_multibyte_sva #(parameter int MAX_W = 32) (
input logic clk,
input logic rst_n,
input logic tx_push,
input logic tx_ready,
input logic flush,
input logic clr_flags,
input logic rx_pop,
input logic rx_ready,
input logic rx_overrun,
input logic core_busy,
input logic shift_en,
input logic start_stb,
input logic rx_valid_stb,
input logic req,
input logic load_stb,
input logic stalled
);
default clocking cb @(posedge clk); endclocking
default disable iff (!rst_n);
// THE handover rule. The datapath may never be loaded while the clock runs,
// because that would overwrite the word being shifted.
a_no_load_while_shifting: assert property (load_stb |-> !shift_en);
// And a load must never coincide with a frame start, because the datapath's
// load and its first drive are both sampled on the same edge and the drive
// wins -- so a coinciding load is silently discarded and the frame sends the
// PREVIOUS word.
a_no_load_at_start: assert property (!(load_stb && start_stb));
// A push is accepted exactly when there is room, and acceptance must not
// depend on anything else -- a design that also required the bus to be idle
// would remove the whole benefit.
a_push_accepted: assert property (tx_push && tx_ready |=> !tx_ready || flush);
// THE slack property. The shadow must be free for some part of every frame,
// or the producer has no more time than it had without the shadow. Stated as
// a cover because it is about the stimulus AND the design together.
c_slack_exists: cover property (tx_ready && shift_en);
// A stall means the wire is idle inside a transaction with nothing queued.
// It must NOT fire while the transaction is closing, which is the tail guard.
a_stall_implies_idle: assert property (
stalled |-> core_busy && !shift_en && !req
);
// The receive side. An overrun is data loss and must be sticky.
a_overrun_sticky: assert property (
rx_overrun && !clr_flags |=> rx_overrun
);
a_overrun_cause: assert property (
rx_valid_stb && rx_ready && !rx_pop |=> rx_overrun
);
// And it must not fire when the consumer kept up, or it means nothing.
a_no_false_overrun: assert property (
!rx_overrun && rx_valid_stb && !rx_ready |=> !rx_overrun
);
// A pop on the same cycle as an arriving word keeps the new word: the pop
// took the old one and the new one has not been read.
a_pop_and_valid: assert property (
rx_pop && rx_valid_stb |=> rx_ready
);
// Flush must leave nothing armed, or the abandoned transaction's word starts
// the next transaction by itself.
a_flush_disarms: assert property (flush |=> !req && tx_ready);
endmodule// The axis that matters is WHEN the producer wrote, relative to the frame. The
// data is irrelevant here -- Chapter 13.6 covered data -- and the interesting
// bins are the boundaries: writing on the load cycle, and writing too late.
covergroup cg_multibyte @(posedge clk);
// Where in the frame the push landed, as a fraction of the frame.
push_when: coverpoint push_phase iff (tx_push && tx_ready) {
bins on_load_cycle = {0}; // the simultaneous push-and-load case
bins early_in_frame = {1};
bins late_in_frame = {2};
bins in_the_gap = {3};
bins too_late = {4}; // after the gap: a stall occurred
}
// Stall cycles per boundary. Zero is the design goal; the others must be
// reachable or the stall counter is untested.
stall_len: coverpoint stalls_this_boundary iff (frame_start) {
bins none = {0};
bins one = {1};
bins few = {[2:8]};
bins many = {[9:$]};
}
// Transaction length in frames. One frame has no boundary at all, so it
// cannot exercise the shadow; two is the minimum that can.
frames: coverpoint frames_this_txn iff (txn_end) {
bins one = {1};
bins two = {2};
bins few = {[3:16]};
bins many = {[17:$]};
}
// The receive side's drain behaviour, which is what produces or avoids an
// overrun. "Never" must be reachable, because that is the overrun test.
drain: coverpoint drain_policy iff (txn_end) {
bins every_word = {0};
bins every_other = {1};
bins never = {2};
}
// The tail: a transaction closing with nothing queued, which must NOT be
// counted as a stall. This bin existing is how the guard stays tested.
tail_seen: coverpoint tail iff (core_busy && !shift_en) {
bins closing = {1};
bins running = {0};
}
x_when_stall: cross push_when, stall_len;
x_frames_drain: cross frames, drain;
endgroup8. Why an FPGA or ASIC Engineer Cares
The cost is one MAX_W register plus five flops. At MAX_W = 32 the shadow is 32 flops, plus sh_last, sh_valid, armed, armed_last and tail. The receive holding register is another 32 plus rx_full and rx_overrun. That is the whole feature: about 70 flops, no arithmetic, no wide combinational logic.
Nothing here is on a timing-critical path. can_load is a four-term AND of registered signals; the handover is a register-to-register transfer evaluated once per frame. The CPU-side interface is a write enable and a read enable.
The receive holding register is what makes the read interface safe. Without it, software would read the shift register directly — which is being shifted during the next frame, so a read landing mid-frame returns a partially-shifted word. The holding register is loaded once per frame at a known moment and held, which means the read path has no timing relationship to the SCLK rate at all.
If depth is increased, the shadow becomes a FIFO and the handover does not change. can_load becomes "the FIFO is not empty" and the free-on-load becomes a read pointer increment. That is worth noting because it means the depth decision can be deferred: build one deep, measure the stall counter on real traffic, and add depth only if the measurement says the producer cannot keep up. The interface to the rest of the design is identical either way.
Cost of getting the depth wrong in the other direction. A 16-deep FIFO on a design whose producer always keeps up is 512 flops doing nothing. On a small FPGA that is a block RAM that could have been something else, and it is very commonly spent because "deeper is faster" is an intuition that survives contact with almost no evidence.
9. Failure Signature — A Flash Page Program That Takes Four Times Too Long
Symptom. A flash driver programs a 256-byte page. The operation takes 2.1 ms where the arithmetic says 0.55 ms. The flash's own program time is 0.4 ms of that, and the SPI transfer should be 150 µs. SCLK is running at the configured rate and every byte arrives correctly; the page verifies.
What that rules out. Correct data and a correct SCLK frequency rule out the divider, the mode, the framing and the datapath. The transfer is taking four times as long as its byte count implies, which means the wire is idle most of the time — and the only thing that idles a correct wire is a producer that is not keeping up.
The measurement that identifies it. stalled, accumulated over the transaction, against the 259 frame boundaries:
transfer time 150 us expected, 1700 us measured
frames 260
idle per frame (1700 - 150) / 259 = 6.0 usSix microseconds per byte, at a 100 MHz clock, is 600 cycles. That is not a hardware interval — it is an interrupt latency plus a handler, which puts the fault firmly in the driver's structure rather than in the master.
The mechanism. The driver takes an interrupt per byte. Each interrupt writes one byte and returns; the next byte is not queued until the next interrupt, which fires when the current frame completes. So the shadow is empty at every boundary and the wire waits for the interrupt every time — the shadow register is present and unused.
Why the shadow did not help. Because the driver never writes ahead. A one-deep shadow gives the producer a frame's worth of slack, and a producer that is only ever invoked at the boundary cannot use slack that exists before the boundary. The hardware feature is correct and the software does not exploit it.
The fix, in the driver. Write the next byte inside the same interrupt that writes the current one, while tx_ready is still asserted — which it is, because the shadow was freed at load. That single change takes the driver from one byte per interrupt to one byte per frame with the interrupt off the critical path, and the measured stall drops to zero.
Why a deeper FIFO would also have fixed it, and why that is the wrong fix. A 16-deep FIFO lets the interrupt-per-byte driver fall 16 frames behind without the wire idling, so the symptom disappears. It also spends 512 flops to paper over a driver structure that will cause the same problem on the next interface, and it does nothing for a driver that falls 17 behind. The stall counter pointed at the driver; the fix belongs there.
The general lesson. A throughput shortfall with correct data is always a gap problem, and the gap is measurable. stalled turns "it is slower than it should be" into "the producer is 600 cycles late per frame", which is a number that identifies the layer at fault.
10. Common Misconceptions
"A deeper FIFO is faster." It is not. Once the producer can stay one word ahead the wire is saturated, and no depth improves that. Depth buys tolerance of a longer producer absence, which is a different and also useful property — and the right depth is the worst-case producer latency divided by the frame time.
"A stall is an error." It is a cost. A low-priority background transfer that stalls constantly is behaving correctly. Treating it as an error produces a flag that gets masked, and then the one transfer where the stall matters is silent too.
"An overrun and a stall are the same kind of problem." A stall costs time; an overrun loses data. The first can clear on read; the second must be sticky, because a driver that checks once at the end of a transaction has to find out about a word dropped two hundred frames earlier.
"The shadow should be freed when the frame completes — that is when the word has actually gone." Freeing it at frame completion means it is full for the whole of every frame, so the producer gets only the inter-frame gap and the decoupling is worth nothing. Freeing it at load is the entire feature.
"Loading the datapath mid-frame is harmless if the load only changes the upper bits." It overwrites the shift register, which is mid-shift, so the bits still to be sent are replaced by bits from a word that was meant for the next frame. The guard is !shift_en and it is not negotiable.
11. Reason It Through
Why must a push and a load on the same cycle both succeed?
Because that cycle is the common case, not a corner. The CPU polling tx_ready and writing the instant it rises will hit exactly the cycle after the load, and a tight loop will hit the load cycle itself. A design that lost one of the two would drop a word on most fast streams, and the symptom would be a transaction one byte short with no flag set.
Why does stalled need the tail guard, and what would happen without it?
Because after the final frame the transaction is still open, the clock is stopped, and nothing is queued — by design. Without the guard every transaction ends with lag-plus-gap stall cycles, so a driver reading the counter concludes it is always late. A flag that fires on correct behaviour gets masked, and then it is worse than absent.
What is the right depth for the transmit buffer, and how would you determine it?
The worst-case producer latency divided by the frame time, rounded up. Determine it by measurement rather than by argument: build one deep, instrument stalled, and run real traffic. If the stall count is zero the depth is right; if it is not, the counter tells you how many frame times the producer is behind, which is the depth needed.
Why is there a separate receive holding register rather than letting software read the shift register?
Because the shift register is being shifted during the next frame, so a read landing mid-frame returns a partially-shifted word. The holding register is loaded once at a known moment and held, which removes any timing relationship between the read path and the SCLK rate — and it is what makes the overrun flag meaningful, since there is now a definite word that can be overwritten.
A design has a 16-deep FIFO, its wire never idles, and its throughput is still half what the arithmetic predicts. Where would you look?
Not at the buffer — a wire that never idles is saturated, so the shortfall is not a gap problem. The remaining candidates are the intervals: a lead, lag or gap programmed larger than the slave requires, which Chapter 13.7's two-sided interval checks exist to catch; or a frame width smaller than it needs to be, so the fixed per-frame overhead is amortised over fewer bits, which is Chapter 9.2's arithmetic. Measure the idle cycles inside a transaction: if they equal the programmed gap times the boundary count, the gap is the answer.
12. Understanding Check
13. Summary
With a single transmit register the wire is idle at every frame boundary for however long the CPU takes to notice — 24% utilisation for an interrupt-per-byte driver at a plausible clock ratio. A one-deep shadow removes the dependency by giving the producer a whole frame instead of a gap.
A deeper FIFO buys latency tolerance, not throughput. Once the producer stays one word ahead the wire is saturated; depth determines how long the producer may be absent. The right depth is the worst-case producer latency divided by the frame time, and it should be measured rather than argued — build one deep, instrument the stall counter, add depth if the numbers say so.
The shadow is freed at load, not at frame completion. One line, and it is the difference between a shadow register and a delay.
Two flags, and the distinction between them is a general pattern: stalled reports a cost and is not an error — a low-priority transfer stalls constantly and correctly — while rx_overrun reports a loss and must be sticky, because a driver may only check at the end of a two-hundred-frame transaction.
stalled needs the tail guard, because after the final frame the transaction is still open with nothing queued by design. Without it every transaction ends with a burst of false stalls, and a flag that fires on correct behaviour gets masked.
The datapath may only be loaded while the clock is stopped, a push and a load on the same cycle must both succeed — that cycle is the common case for a tight loop, not a corner — and flush must leave nothing armed, or an abandoned transaction's word starts the next one by itself.
Integrating this block found a stale-level bug in Chapter 13.5 that no unit test could have found: a level meaning "the previous frame finished" was still true on the next frame's first cycle. Any consumer that restarts and observes in the same cycle needs the start term.
For verification, the driver's pacing is the feature under test — a driver that waits for completion before writing the next word exercises none of it — and the scoreboard must check stalled against the delay the sequence asked for, or both a broken flag and a broken handover are silent.
14. What Comes Next
The master streams correctly and reports what it costs. Every path through it so far assumes the transfer finishes.
Chapter 13.10 — Busy/Done/Valid Interface, Reset, and Abort handles the case where it does not, and the two ways of stopping turn out to be opposites: one is orderly and costs time, the other is instant and costs protocol correctness — and knowing which is which decides whether the slave is still usable afterwards.
Continue learning
Related tutorials
- Related topic
Dummy Phases and Read Latency
Why a device needs turnaround before it can answer, why dummy is counted in clock cycles rather than bytes, how its length grows with frequency, and the one-byte data offset a mismatch produces.
- Related topic
Back-to-Back Transactions and Inter-Frame Gap
How soon chip select may fall again after it rises: the minimum deselect time, why a device needs it, what the gap costs in throughput, and the hardware enforcement that keeps software from violating it.
- Related topic
Raw vs Effective Payload Throughput
The three throughput numbers a link has — raw bit rate, protocol rate and effective payload rate — why they differ by four times on an ordinary flash read, and the performance counter that measures the gap instead of estimating it.
- Related topic
SDR, DDR Transfers, and Throughput
Why data on both edges is independent of lane width, why the byte assembly is identical in both modes, why the dummy phase does not halve so DDR gives under two times, and why DDR parts specify more dummy cycles than their SDR modes.
