SPI · Module 10
From Datasheet to Transaction Specification
A converter datasheet worked end to end into a mode, a divisor, a phase schedule and an executable element sequence — with the specification engine that generates it from a profile and the UVM environment that verifies a device described entirely by data.
Five chapters have each extracted one thing. This one puts them together on a single real part and ends where the module promised: with a specification a controller can execute.
You have read the datasheet. What, precisely, do you hand to the RTL team — and what makes it a specification rather than a description?
A specification is something that can be executed and checked. Prose cannot. The artifact this chapter ends on is a table of numbers and an element sequence, and both are machine-readable.
1. The Device
Work a 12-bit successive-approximation ADC — the most common shape of SPI converter, and deliberately not a flash, because a converter exercises the phases that memory chapters skip.
What a datasheet for such a part typically tells you, spread across four sections:
Pin table SCLK, SDI, SDO, CS 4-wire, CS active low, 3.3 V
Timing diagram SCLK idles low; SDO changes on falling edges
Timing table f_SCLK max 20 MHz at 20 pF; t_V 20 ns; t_SU 10 ns
Conversion acquisition begins on CS fall; result available
after 4 SCLK cycles
Output format 12 bits, MSB first, in a 16-bit frame,
left-aligned with four trailing zerosNothing there says "mode 0" and nothing says "dummy cycles". Both have to be derived.
2. Deriving the Mode
Apply Chapter 10.2's two observations.
CPOL. SCLK idles low → CPOL = 0.
CPHA. The sentence given is "SDO changes on falling edges". SDO is the device's output — our MISO. With CPOL=0 the leading edge is rising, so falling edges are the trailing ones, the even edges. The device launches on even edges, which means the master must sample on the odd ones.
Sampling on odd edges is CPHA = 0.
So the mode is 0 — and note how the derivation went. The datasheet described the device's output edge, exactly the trap of Chapter 10.2's worked example, and the conversion needed one extra step: launch on even means sample on odd.
3. Deriving the Rate
Chapter 10.3's extraction, then Chapter 9.4's budget.
The table gives 20 MHz at 20 pF and t_V of 20 ns. Take a board presenting about 40 pF, a controller with 2 ns clock-to-out and 2 ns input setup, and 0.5 ns of trace each way. Derate t_V for double the quoted load — call it 26 ns.
controller clock-to-out .... 2.0 ns
trace out .................. 0.5 ns
device t_V (derated) ....... 26.0 ns
trace back ................. 0.5 ns
controller input setup ..... 2.0 ns
──────────────────────────────────
total ...................... 31.0 ns → T ≥ 62 ns → f_max ≈ 16.1 MHzThe device's own 20 MHz never binds — the round trip stops us at 16. With an 80 MHz system clock the divisor ladder offers 20, 16 and 13.3 MHz; round down to divisor 5, giving 16 MHz, which sits just under the ceiling.
4. Deriving the Phase Schedule
The datasheet says acquisition begins on CS fall and the result is available after 4 SCLK cycles. It never uses the word "dummy". But per Chapter 10.5, that is precisely what a dummy phase is: clocked time in which no meaningful data moves, spent waiting for the device.
command phase 0 bits — this part has no opcode at all
address phase 0 bytes — there is nothing to address
dummy phase 4 cycles — the acquisition time, in CYCLES
data phase 2 bytes — 12 bits left-aligned in 16Two things are worth noticing.
Four is not eight. A driver that sends a leading dummy byte clocks four extra cycles, and since the result is left-aligned the first four bits of the payload are lost and the value is shifted — returning a number that is plausibly in range and wrong by a factor of sixteen. This is exactly the failure Chapter 10.5 drew, on a part where the wrong answer still looks like a measurement.
The command phase is empty. Many converters have no opcode — the act of asserting CS is the command. A planner that assumes at least one command byte cannot describe this part, which is why Chapter 10.5's walker skips empty phases rather than entering them.
5. The Specification
Everything above collapses into this:
── interface ────────────────────────────
mode 0 (CPOL=0, CPHA=0)
divisor 5 16 MHz from an 80 MHz clock
bit order MSB first
── transaction: read conversion ─────────
command none
address none
dummy 4 cycles
data 2 bytes in
first data cycle 4
total cycles 20
── framing ──────────────────────────────
CS asserted for the whole transaction
CS released between conversions (starts the next acquisition)
── post-processing ──────────────────────
result (raw >> 4) & 0x0FFFThat is the deliverable. It is executable, every number in it is checkable, and nothing in it is prose. The RTL team does not need the datasheet to implement it, and — more usefully — a reviewer can compare it against the datasheet line by line without reading any code.
The (raw >> 4) & 0x0FFF line is easy to skip and belongs in the specification as firmly as the rest. A left-aligned result read as though it were right-aligned is wrong by a factor of sixteen while remaining a legal 16-bit number, so nothing downstream can detect it.
6. Building the Specification Engine — Three HDLs
The circuit
Circuit. A generator that turns a profile and a request into the element sequence a controller must issue.
State. The captured profile fields, a pre-shifted address register, and a small phase machine.
Datapath. The opcode is selected by direction; the address is emitted most-significant byte first; the dummy phase is emitted as one element carrying a cycle count, never as bytes; and the data phase is one element carrying a length. The first data cycle is computed with Chapter 10.5's arithmetic and published at start.
Control. A five-state walk that skips absent phases entirely. An absent phase emits no element, rather than an element with a length of zero — a consumer that sees a dummy element will wait, so an empty one costs real cycles.
Clock and reset. System clock; asynchronous active-low reset.
Enables. spec_error reports an address too wide for the configured address phase — the 16 MB ceiling of Chapter 10.5, which the device cannot detect because the bits simply are not in the transfer.
Timing. One element per cycle while busy. The consumer is a pin-level master that expands each element into clocked bits.
Address order. The address is pre-shifted once at start so the first byte to send sits at the top of the register; every element after that is a fixed top-byte extraction and a shift by eight. That costs one barrel shifter at load instead of one per byte, and it is the standard answer whenever a variable-width field must be emitted from a fixed-width register.
Synthesis. An address register with a load-time shifter, three small counters and a five-state machine.
Limitations. It deliberately stops at the element level and does not drive pins. A pin-level master already exists earlier in this track; what was missing was the thing that decides what to send. Keeping them separate is what makes one master serve every device.
// spi_txn_spec.sv
//
// Chapter 10.6 -- from datasheet to transaction specification.
//
// This is the block the whole module has been building towards. It takes
// the device profile of Chapter 10.1, the command convention of Chapter
// 10.4 and the latency arithmetic of Chapter 10.5, and emits the exact
// sequence a controller must issue: the opcode byte, the address bytes in
// the order the device expects, the dummy phase with its cycle count, and
// the data phase with its length -- plus the cycle at which returned data
// begins.
//
// It deliberately stops at the ELEMENT level rather than driving pins. A
// pin-level master already exists earlier in this track; what was missing
// was the thing that says WHAT to send, derived from datasheet numbers
// rather than hard-coded. Separating the two means one master serves every
// device, and adding a part means adding a profile rather than a driver.
//
// ADDRESS ORDER. SPI devices send addresses most-significant byte first,
// essentially without exception. The address is pre-shifted once at start
// so that the first byte to send sits at the top of the register; every
// element after that is a fixed top-byte extraction and a shift by eight.
// That costs one barrel shifter at load instead of one per byte.
module spi_txn_spec #(
parameter int ADDR_W = 32,
parameter int LEN_W = 16,
parameter int MAX_ABYTES = 4,
parameter logic [7:0] CMD_READ = 8'h0B, // fast read
parameter logic [7:0] CMD_WRITE = 8'h02 // page program
) (
input logic clk,
input logic rst_n,
input logic start,
input logic req_read,
input logic [ADDR_W-1:0] req_addr,
input logic [LEN_W-1:0] req_len,
// From the device profile.
input logic [2:0] addr_bytes,
input logic [5:0] dummy_cycles,
output logic elem_valid,
output logic [1:0] elem_kind, // 0=cmd 1=addr 2=dummy 3=data
output logic [7:0] elem_data, // opcode / address byte / dummy count
output logic [LEN_W-1:0] elem_len, // data phase: bytes; otherwise 1
output logic spec_error, // address does not fit addr_bytes
output logic busy,
output logic done,
output logic [15:0] first_data_cycle
);
localparam logic [1:0] K_CMD = 2'd0;
localparam logic [1:0] K_ADDR = 2'd1;
localparam logic [1:0] K_DUMMY = 2'd2;
localparam logic [1:0] K_DATA = 2'd3;
localparam logic [2:0] S_IDLE = 3'd0;
localparam logic [2:0] S_CMD = 3'd1;
localparam logic [2:0] S_ADDR = 3'd2;
localparam logic [2:0] S_DUMMY = 3'd3;
localparam logic [2:0] S_DATA = 3'd4;
logic [2:0] state;
logic [ADDR_W-1:0] addr_sr;
logic [2:0] abytes_r;
logic [5:0] dummy_r;
logic [LEN_W-1:0] len_r;
logic read_r;
// An address wider than the device's address phase cannot be expressed.
// This is the arithmetic behind the 16 MB ceiling on three-byte
// addressing: a part with a 3-byte address phase cannot be told to read
// above 0xFFFFFF, and a controller that truncates silently wraps the
// access to the bottom of the device -- reading valid-looking data from
// entirely the wrong place.
// Width of the address phase in bits. The multiply is done on a
// widened value: addr_bytes is three bits, and "addr_bytes * 8" in
// three-bit arithmetic overflows for anything above 0 -- the same
// width-truncation trap that makes a bound check always pass.
wire [31:0] addr_bits = 32'(addr_bytes) << 3;
wire [ADDR_W-1:0] addr_limit_mask =
(addr_bits >= 32'(ADDR_W)) ? {ADDR_W{1'b1}}
: ((ADDR_W'(1) << addr_bits) - ADDR_W'(1));
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
state <= S_IDLE;
elem_valid <= 1'b0;
elem_kind <= K_CMD;
elem_data <= 8'h00;
elem_len <= {LEN_W{1'b0}};
spec_error <= 1'b0;
busy <= 1'b0;
done <= 1'b0;
first_data_cycle <= 16'd0;
addr_sr <= {ADDR_W{1'b0}};
abytes_r <= 3'd0;
dummy_r <= 6'd0;
len_r <= {LEN_W{1'b0}};
read_r <= 1'b0;
end else begin
elem_valid <= 1'b0;
done <= 1'b0;
case (state)
S_IDLE: begin
if (start) begin
abytes_r <= addr_bytes;
dummy_r <= dummy_cycles;
len_r <= req_len;
read_r <= req_read;
// Pre-shift so the most significant address byte
// that the device expects sits at the top.
addr_sr <= req_addr << (8 * (MAX_ABYTES - addr_bytes));
spec_error <= ((req_addr & ~addr_limit_mask) != 0);
// The same arithmetic as Chapter 10.5: command bits
// plus address bits plus dummy CYCLES.
first_data_cycle <= 16'd8
+ (16'(addr_bytes) << 3)
+ 16'(dummy_cycles);
// The opcode is chosen by direction. A device whose
// read and write opcodes differ in a single bit can
// be described by Chapter 10.4's codec instead; a
// flash, whose opcodes are unrelated values, needs
// a table like this one.
elem_valid <= 1'b1;
elem_kind <= K_CMD;
elem_data <= req_read ? CMD_READ : CMD_WRITE;
elem_len <= LEN_W'(1);
busy <= 1'b1;
state <= S_CMD;
end
end
S_CMD: begin
if (abytes_r != 3'd0) begin
elem_valid <= 1'b1;
elem_kind <= K_ADDR;
elem_data <= addr_sr[ADDR_W-1 -: 8];
elem_len <= LEN_W'(1);
addr_sr <= addr_sr << 8;
abytes_r <= abytes_r - 3'd1;
state <= S_ADDR;
end else if (dummy_r != 6'd0) begin
state <= S_DUMMY;
end else if (len_r != 0) begin
state <= S_DATA;
end else begin
busy <= 1'b0;
done <= 1'b1;
state <= S_IDLE;
end
end
S_ADDR: begin
if (abytes_r != 3'd0) begin
elem_valid <= 1'b1;
elem_kind <= K_ADDR;
elem_data <= addr_sr[ADDR_W-1 -: 8];
elem_len <= LEN_W'(1);
addr_sr <= addr_sr << 8;
abytes_r <= abytes_r - 3'd1;
end else if (dummy_r != 6'd0) begin
state <= S_DUMMY;
end else if (len_r != 0) begin
state <= S_DATA;
end else begin
busy <= 1'b0;
done <= 1'b1;
state <= S_IDLE;
end
end
S_DUMMY: begin
// One element carrying the cycle count -- not a byte.
// Emitting a dummy BYTE is the error this whole module
// exists to prevent.
elem_valid <= 1'b1;
elem_kind <= K_DUMMY;
elem_data <= {2'b00, dummy_r};
elem_len <= LEN_W'(1);
if (len_r != 0) begin
state <= S_DATA;
end else begin
busy <= 1'b0;
done <= 1'b1;
state <= S_IDLE;
end
end
default: begin // S_DATA
elem_valid <= 1'b1;
elem_kind <= K_DATA;
elem_data <= 8'h00;
elem_len <= len_r;
busy <= 1'b0;
done <= 1'b1;
state <= S_IDLE;
end
endcase
end
end
endmodule// spi_txn_spec_tb.sv
//
// The testbench collects the emitted element stream and compares it,
// element for element, against a sequence computed by hand from a real
// datasheet reading. A specification builder is only worth having if its
// output is exactly what the device expects, so nothing here is checked
// against the DUT's own arithmetic.
`timescale 1ns/1ps
module spi_txn_spec_tb;
localparam int ADDR_W = 32;
localparam int LEN_W = 16;
localparam logic [1:0] K_CMD = 2'd0;
localparam logic [1:0] K_ADDR = 2'd1;
localparam logic [1:0] K_DUMMY = 2'd2;
localparam logic [1:0] K_DATA = 2'd3;
logic clk = 1'b0;
logic rst_n = 1'b0;
always #5 clk = ~clk;
logic start = 1'b0;
logic req_read = 1'b0;
logic [ADDR_W-1:0] req_addr = {ADDR_W{1'b0}};
logic [LEN_W-1:0] req_len = {LEN_W{1'b0}};
logic [2:0] addr_bytes = 3'd0;
logic [5:0] dummy_cycles = 6'd0;
logic elem_valid;
logic [1:0] elem_kind;
logic [7:0] elem_data;
logic [LEN_W-1:0] elem_len;
logic spec_error, busy, done;
logic [15:0] first_data_cycle;
int errors = 0;
// Collected stream.
logic [1:0] got_kind [0:15];
logic [7:0] got_data [0:15];
logic [LEN_W-1:0] got_len [0:15];
int got_n;
spi_txn_spec #(
.ADDR_W(ADDR_W), .LEN_W(LEN_W), .MAX_ABYTES(4),
.CMD_READ(8'h0B), .CMD_WRITE(8'h02)
) dut (
.clk(clk), .rst_n(rst_n), .start(start),
.req_read(req_read), .req_addr(req_addr), .req_len(req_len),
.addr_bytes(addr_bytes), .dummy_cycles(dummy_cycles),
.elem_valid(elem_valid), .elem_kind(elem_kind),
.elem_data(elem_data), .elem_len(elem_len),
.spec_error(spec_error), .busy(busy), .done(done),
.first_data_cycle(first_data_cycle)
);
// Collect every emitted element.
always @(posedge clk) begin
if (rst_n && elem_valid && got_n < 16) begin
got_kind[got_n] <= elem_kind;
got_data[got_n] <= elem_data;
got_len[got_n] <= elem_len;
got_n <= got_n + 1;
end
end
task automatic issue(input logic rd, input logic [ADDR_W-1:0] a,
input int len, input int ab, input int dc);
begin
@(negedge clk);
got_n = 0;
req_read = rd; req_addr = a; req_len = LEN_W'(len);
addr_bytes = 3'(ab); dummy_cycles = 6'(dc);
start = 1'b1;
@(negedge clk);
start = 1'b0;
while (busy) @(negedge clk);
@(negedge clk);
end
endtask
task automatic expect_elem(input string name, input int i,
input logic [1:0] k, input logic [7:0] d);
begin
if (i >= got_n) begin
$display(" FAIL: %s has only %0d elements, wanted one at %0d",
name, got_n, i);
errors++;
end else if (got_kind[i] !== k || got_data[i] !== d) begin
$display(" FAIL: %s element %0d is kind=%0d data=0x%02h, expected kind=%0d data=0x%02h",
name, i, got_kind[i], got_data[i], k, d);
errors++;
end
end
endtask
task automatic expect_count(input string name, input int n);
begin
if (got_n != n) begin
$display(" FAIL: %s emitted %0d elements, expected %0d",
name, got_n, n);
errors++;
end
end
endtask
task automatic show(input string name);
string s;
begin
s = "";
for (int i = 0; i < got_n; i++) begin
case (got_kind[i])
K_CMD: s = {s, $sformatf("CMD 0x%02h | ", got_data[i])};
K_ADDR: s = {s, $sformatf("ADDR 0x%02h | ", got_data[i])};
K_DUMMY: s = {s, $sformatf("DUMMY %0dc | ", got_data[i])};
default: s = {s, $sformatf("DATA %0dB", got_len[i])};
endcase
end
$display(" %-20s %s (data begins at cycle %0d)",
name, s, first_data_cycle);
end
endtask
initial begin
got_n = 0;
repeat (3) @(negedge clk);
rst_n = 1'b1;
@(negedge clk);
// 1. A flash fast read, 0x0B, three address bytes, eight dummy
// cycles, four bytes of payload. Every byte here is what a
// logic analyser on a real part would show.
issue(1'b1, 32'h00123456, 4, 3, 8);
expect_count("fast read", 6);
expect_elem("fast read", 0, K_CMD, 8'h0B);
expect_elem("fast read", 1, K_ADDR, 8'h12); // MSB first
expect_elem("fast read", 2, K_ADDR, 8'h34);
expect_elem("fast read", 3, K_ADDR, 8'h56);
expect_elem("fast read", 4, K_DUMMY, 8'd8);
expect_elem("fast read", 5, K_DATA, 8'h00);
if (got_len[5] !== LEN_W'(4)) begin
$display(" FAIL: fast read data length %0d, expected 4", got_len[5]);
errors++;
end
if (first_data_cycle !== 16'd40) begin
$display(" FAIL: fast read data begins at %0d, expected 40",
first_data_cycle);
errors++;
end
if (spec_error) begin
$display(" FAIL: a 24-bit address in a 3-byte phase was rejected");
errors++;
end
show("fast read");
// 2. A page program: a different opcode, no dummy phase at all.
// The dummy element must be ABSENT, not present with a count of
// zero -- a consumer that sees a dummy element will wait.
issue(1'b0, 32'h00000100, 4, 3, 0);
expect_count("page program", 5);
expect_elem("page program", 0, K_CMD, 8'h02);
expect_elem("page program", 1, K_ADDR, 8'h00);
expect_elem("page program", 2, K_ADDR, 8'h01);
expect_elem("page program", 3, K_ADDR, 8'h00);
expect_elem("page program", 4, K_DATA, 8'h00);
if (first_data_cycle !== 16'd32) begin
$display(" FAIL: page program data begins at %0d, expected 32",
first_data_cycle);
errors++;
end
show("page program");
// 3. Four-byte addressing, the mode a part above 128 Mbit needs.
issue(1'b1, 32'h01234567, 2, 4, 8);
expect_count("4-byte address", 7);
expect_elem("4-byte address", 1, K_ADDR, 8'h01);
expect_elem("4-byte address", 2, K_ADDR, 8'h23);
expect_elem("4-byte address", 3, K_ADDR, 8'h45);
expect_elem("4-byte address", 4, K_ADDR, 8'h67);
expect_elem("4-byte address", 5, K_DUMMY, 8'd8);
if (first_data_cycle !== 16'd48) begin
$display(" FAIL: 4-byte address data begins at %0d, expected 48",
first_data_cycle);
errors++;
end
show("4-byte address");
// 4. THE 16 MB CEILING. An address above 0xFFFFFF cannot be
// expressed in three address bytes. Truncating silently would
// wrap the access to the bottom of the device and return
// perfectly valid data from entirely the wrong place.
issue(1'b1, 32'h01000000, 4, 3, 8);
if (!spec_error) begin
$display(" FAIL: 0x01000000 does not fit three address bytes and was accepted");
errors++;
end
issue(1'b1, 32'h00FFFFFF, 4, 3, 8);
if (spec_error) begin
$display(" FAIL: 0x00FFFFFF fits three address bytes exactly and was rejected");
errors++;
end
$display(" address ceiling: 0x00FFFFFF accepted, 0x01000000 rejected -- the boundary is exact");
// 5. A status read: no address phase, no dummy. Straight from
// command to data.
issue(1'b1, 32'h00000000, 1, 0, 0);
expect_count("status read", 2);
expect_elem("status read", 0, K_CMD, 8'h0B);
expect_elem("status read", 1, K_DATA, 8'h00);
if (first_data_cycle !== 16'd8) begin
$display(" FAIL: status read data begins at %0d, expected 8",
first_data_cycle);
errors++;
end
show("status read");
// 6. AN ADC. No address phase at all -- the conversion result is
// simply clocked out -- but a leading latency while the sample
// is acquired, then a two-byte result. This is the shape most
// converters have, and it is the worked example of this
// chapter: the dummy phase exists without any address before it.
issue(1'b1, 32'h00000000, 2, 0, 4);
expect_count("adc read", 3);
expect_elem("adc read", 0, K_CMD, 8'h0B);
expect_elem("adc read", 1, K_DUMMY, 8'd4);
expect_elem("adc read", 2, K_DATA, 8'h00);
if (got_len[2] !== LEN_W'(2)) begin
$display(" FAIL: adc read data length %0d, expected 2", got_len[2]);
errors++;
end
if (first_data_cycle !== 16'd12) begin
$display(" FAIL: adc read data begins at %0d, expected 12",
first_data_cycle);
errors++;
end
show("adc read");
// 7. A bare command with neither address nor data -- write enable.
issue(1'b0, 32'h00000000, 0, 0, 0);
expect_count("write enable", 1);
expect_elem("write enable", 0, K_CMD, 8'h02);
show("write enable");
if (errors == 0)
$display("PASS: every element stream matches the sequence a real device expects, addresses are emitted most-significant byte first, absent phases emit no element at all, the data cycle agrees with the latency arithmetic, and an address too wide for the address phase is rejected rather than truncated");
else
$display("FAIL: %0d error(s)", errors);
$finish;
end
endmoduleEvery expected sequence in that testbench is what a logic analyser on the real part would show, written down from the command table rather than derived from the design. That distinction is the whole value of the test: a specification engine checked against its own arithmetic proves only that it is self-consistent.
The seven shapes are chosen so that each phase is present in some and absent in others — a flash fast read with all four; a page program with no dummy; four-byte addressing; the ADC of §1–5 with no command address and a dummy phase but no address before it; a status read with neither; and a bare write-enable. The adc read line in the output is the specification of §5, produced by hardware.
Two checks carry most of the weight. An absent phase must emit no element at all — the page program produces five elements, not six with a zero-length dummy. And the address ceiling is exact: 0x00FFFFFF is accepted and 0x01000000 is rejected, one apart, which is the boundary a comparison written with the wrong relational operator gets wrong in exactly one direction.
// spi_txn_spec.v
//
// Chapter 10.6 -- from datasheet to transaction specification, in
// Verilog-2001.
//
// Takes the device profile of Chapter 10.1, the command convention of
// Chapter 10.4 and the latency arithmetic of Chapter 10.5, and emits the
// exact sequence a controller must issue: the opcode byte, the address
// bytes in the order the device expects, the dummy phase with its CYCLE
// count, and the data phase with its length -- plus the cycle at which
// returned data begins.
//
// It stops at the ELEMENT level rather than driving pins, so one
// pin-level master serves every device and adding a part means adding a
// profile rather than a driver.
//
// ADDRESS ORDER. SPI devices send addresses most-significant byte first.
// The address is pre-shifted once at start so the first byte to send sits
// at the top of the register; every element after that is a fixed
// top-byte extraction and a shift by eight -- one barrel shifter at load
// instead of one per byte.
module spi_txn_spec #(
parameter ADDR_W = 32,
parameter LEN_W = 16,
parameter MAX_ABYTES = 4,
parameter [7:0] CMD_READ = 8'h0B, // fast read
parameter [7:0] CMD_WRITE = 8'h02 // page program
) (
input wire clk,
input wire rst_n,
input wire start,
input wire req_read,
input wire [ADDR_W-1:0] req_addr,
input wire [LEN_W-1:0] req_len,
// From the device profile.
input wire [2:0] addr_bytes,
input wire [5:0] dummy_cycles,
output reg elem_valid,
output reg [1:0] elem_kind, // 0=cmd 1=addr 2=dummy 3=data
output reg [7:0] elem_data, // opcode / address byte / dummy count
output reg [LEN_W-1:0] elem_len, // data phase: bytes; otherwise 1
output reg spec_error, // address does not fit addr_bytes
output reg busy,
output reg done,
output reg [15:0] first_data_cycle
);
localparam [1:0] K_CMD = 2'd0;
localparam [1:0] K_ADDR = 2'd1;
localparam [1:0] K_DUMMY = 2'd2;
localparam [1:0] K_DATA = 2'd3;
localparam [2:0] S_IDLE = 3'd0;
localparam [2:0] S_CMD = 3'd1;
localparam [2:0] S_ADDR = 3'd2;
localparam [2:0] S_DUMMY = 3'd3;
localparam [2:0] S_DATA = 3'd4;
reg [2:0] state;
reg [ADDR_W-1:0] addr_sr;
reg [2:0] abytes_r;
reg [5:0] dummy_r;
reg [LEN_W-1:0] len_r;
reg read_r;
// Width of the address phase in bits. The multiply is done on a
// widened value: addr_bytes is three bits, and "addr_bytes * 8" in
// three-bit arithmetic overflows -- the same width-truncation trap
// that makes a bound check always pass.
wire [31:0] addr_bits = {29'd0, addr_bytes} << 3;
wire [31:0] shift_load = (MAX_ABYTES * 8) - addr_bits;
// An address wider than the device's address phase cannot be
// expressed. This is the arithmetic behind the 16 MB ceiling on
// three-byte addressing: a controller that truncates silently wraps
// the access to the bottom of the device, returning valid-looking data
// from entirely the wrong place.
wire [ADDR_W-1:0] addr_limit_mask =
(addr_bits >= ADDR_W) ? {ADDR_W{1'b1}}
: ((({{(ADDR_W-1){1'b0}}, 1'b1}) << addr_bits) - 1'b1);
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
state <= S_IDLE;
elem_valid <= 1'b0;
elem_kind <= K_CMD;
elem_data <= 8'h00;
elem_len <= {LEN_W{1'b0}};
spec_error <= 1'b0;
busy <= 1'b0;
done <= 1'b0;
first_data_cycle <= 16'd0;
addr_sr <= {ADDR_W{1'b0}};
abytes_r <= 3'd0;
dummy_r <= 6'd0;
len_r <= {LEN_W{1'b0}};
read_r <= 1'b0;
end else begin
elem_valid <= 1'b0;
done <= 1'b0;
case (state)
S_IDLE: begin
if (start) begin
abytes_r <= addr_bytes;
dummy_r <= dummy_cycles;
len_r <= req_len;
read_r <= req_read;
// Pre-shift so the most significant address byte
// the device expects sits at the top.
addr_sr <= req_addr << shift_load;
spec_error <= ((req_addr & ~addr_limit_mask) != 0);
// The same arithmetic as Chapter 10.5: command
// bits plus address bits plus dummy CYCLES.
first_data_cycle <= 16'd8
+ ({13'd0, addr_bytes} << 3)
+ {10'd0, dummy_cycles};
// The opcode is chosen by direction. A device whose
// read and write opcodes differ in one bit can use
// Chapter 10.4's codec instead; a flash, whose
// opcodes are unrelated values, needs a table.
elem_valid <= 1'b1;
elem_kind <= K_CMD;
elem_data <= req_read ? CMD_READ : CMD_WRITE;
elem_len <= {{(LEN_W-1){1'b0}}, 1'b1};
busy <= 1'b1;
state <= S_CMD;
end
end
S_CMD: begin
if (abytes_r != 3'd0) begin
elem_valid <= 1'b1;
elem_kind <= K_ADDR;
elem_data <= addr_sr[ADDR_W-1 -: 8];
elem_len <= {{(LEN_W-1){1'b0}}, 1'b1};
addr_sr <= addr_sr << 8;
abytes_r <= abytes_r - 3'd1;
state <= S_ADDR;
end else if (dummy_r != 6'd0) begin
state <= S_DUMMY;
end else if (len_r != 0) begin
state <= S_DATA;
end else begin
busy <= 1'b0;
done <= 1'b1;
state <= S_IDLE;
end
end
S_ADDR: begin
if (abytes_r != 3'd0) begin
elem_valid <= 1'b1;
elem_kind <= K_ADDR;
elem_data <= addr_sr[ADDR_W-1 -: 8];
elem_len <= {{(LEN_W-1){1'b0}}, 1'b1};
addr_sr <= addr_sr << 8;
abytes_r <= abytes_r - 3'd1;
end else if (dummy_r != 6'd0) begin
state <= S_DUMMY;
end else if (len_r != 0) begin
state <= S_DATA;
end else begin
busy <= 1'b0;
done <= 1'b1;
state <= S_IDLE;
end
end
S_DUMMY: begin
// One element carrying the CYCLE count -- not a byte.
// Emitting a dummy BYTE is the error this whole module
// exists to prevent.
elem_valid <= 1'b1;
elem_kind <= K_DUMMY;
elem_data <= {2'b00, dummy_r};
elem_len <= {{(LEN_W-1){1'b0}}, 1'b1};
if (len_r != 0) begin
state <= S_DATA;
end else begin
busy <= 1'b0;
done <= 1'b1;
state <= S_IDLE;
end
end
default: begin // S_DATA
elem_valid <= 1'b1;
elem_kind <= K_DATA;
elem_data <= 8'h00;
elem_len <= len_r;
busy <= 1'b0;
done <= 1'b1;
state <= S_IDLE;
end
endcase
end
end
endmodule// spi_txn_spec_tb.v
//
// The same checks as the SystemVerilog testbench: the element stream is
// compared element for element against a sequence computed by hand from a
// real datasheet reading, not against the DUT's own arithmetic.
`timescale 1ns/1ps
module spi_txn_spec_tb;
parameter ADDR_W = 32;
parameter LEN_W = 16;
localparam [1:0] K_CMD = 2'd0;
localparam [1:0] K_ADDR = 2'd1;
localparam [1:0] K_DUMMY = 2'd2;
localparam [1:0] K_DATA = 2'd3;
reg clk;
reg rst_n;
reg start;
reg req_read;
reg [ADDR_W-1:0] req_addr;
reg [LEN_W-1:0] req_len;
reg [2:0] addr_bytes;
reg [5:0] dummy_cycles;
wire elem_valid;
wire [1:0] elem_kind;
wire [7:0] elem_data;
wire [LEN_W-1:0] elem_len;
wire spec_error, busy, done;
wire [15:0] first_data_cycle;
integer errors;
integer i;
reg [1:0] got_kind [0:15];
reg [7:0] got_data [0:15];
reg [LEN_W-1:0] got_len [0:15];
integer got_n;
initial begin
clk = 1'b0; rst_n = 1'b0; start = 1'b0;
req_read = 1'b0; req_addr = {ADDR_W{1'b0}}; req_len = {LEN_W{1'b0}};
addr_bytes = 3'd0; dummy_cycles = 6'd0;
errors = 0; got_n = 0;
end
always #5 clk = ~clk;
spi_txn_spec #(
.ADDR_W(ADDR_W), .LEN_W(LEN_W), .MAX_ABYTES(4),
.CMD_READ(8'h0B), .CMD_WRITE(8'h02)
) dut (
.clk(clk), .rst_n(rst_n), .start(start),
.req_read(req_read), .req_addr(req_addr), .req_len(req_len),
.addr_bytes(addr_bytes), .dummy_cycles(dummy_cycles),
.elem_valid(elem_valid), .elem_kind(elem_kind),
.elem_data(elem_data), .elem_len(elem_len),
.spec_error(spec_error), .busy(busy), .done(done),
.first_data_cycle(first_data_cycle)
);
// Collect every emitted element.
always @(posedge clk) begin
if (rst_n && elem_valid && got_n < 16) begin
got_kind[got_n] <= elem_kind;
got_data[got_n] <= elem_data;
got_len[got_n] <= elem_len;
got_n <= got_n + 1;
end
end
task issue;
input rd;
input [ADDR_W-1:0] a;
input integer len;
input integer ab;
input integer dc;
begin
@(negedge clk);
got_n = 0;
req_read = rd; req_addr = a; req_len = len[LEN_W-1:0];
addr_bytes = ab[2:0]; dummy_cycles = dc[5:0];
start = 1'b1;
@(negedge clk);
start = 1'b0;
while (busy) @(negedge clk);
@(negedge clk);
end
endtask
task expect_elem;
input [8*20:1] name;
input integer idx;
input [1:0] k;
input [7:0] d;
begin
if (idx >= got_n) begin
$display(" FAIL: %0s has only %0d elements, wanted one at %0d",
name, got_n, idx);
errors = errors + 1;
end else if (got_kind[idx] !== k || got_data[idx] !== d) begin
$display(" FAIL: %0s element %0d is kind=%0d data=0x%02h, expected kind=%0d data=0x%02h",
name, idx, got_kind[idx], got_data[idx], k, d);
errors = errors + 1;
end
end
endtask
task expect_count;
input [8*20:1] name;
input integer n;
begin
if (got_n != n) begin
$display(" FAIL: %0s emitted %0d elements, expected %0d",
name, got_n, n);
errors = errors + 1;
end
end
endtask
task show;
input [8*20:1] name;
integer j;
begin
$write(" %0s ", name);
for (j = 0; j < got_n; j = j + 1) begin
case (got_kind[j])
K_CMD: $write("CMD 0x%02h | ", got_data[j]);
K_ADDR: $write("ADDR 0x%02h | ", got_data[j]);
K_DUMMY: $write("DUMMY %0dc | ", got_data[j]);
default: $write("DATA %0dB", got_len[j]);
endcase
end
$display(" (data begins at cycle %0d)", first_data_cycle);
end
endtask
initial begin
repeat (3) @(negedge clk);
rst_n = 1'b1;
@(negedge clk);
// 1. A flash fast read, 0x0B, three address bytes, eight dummy
// cycles, four bytes of payload.
issue(1'b1, 32'h00123456, 4, 3, 8);
expect_count("fast read", 6);
expect_elem("fast read", 0, K_CMD, 8'h0B);
expect_elem("fast read", 1, K_ADDR, 8'h12); // MSB first
expect_elem("fast read", 2, K_ADDR, 8'h34);
expect_elem("fast read", 3, K_ADDR, 8'h56);
expect_elem("fast read", 4, K_DUMMY, 8'd8);
expect_elem("fast read", 5, K_DATA, 8'h00);
if (got_len[5] !== 4) begin
$display(" FAIL: fast read data length %0d, expected 4", got_len[5]);
errors = errors + 1;
end
if (first_data_cycle !== 16'd40) begin
$display(" FAIL: fast read data begins at %0d, expected 40",
first_data_cycle);
errors = errors + 1;
end
if (spec_error) begin
$display(" FAIL: a 24-bit address in a 3-byte phase was rejected");
errors = errors + 1;
end
show("fast read ");
// 2. A page program: a different opcode, no dummy phase at all.
// The dummy element must be ABSENT, not present with a count of
// zero -- a consumer that sees a dummy element will wait.
issue(1'b0, 32'h00000100, 4, 3, 0);
expect_count("page program", 5);
expect_elem("page program", 0, K_CMD, 8'h02);
expect_elem("page program", 1, K_ADDR, 8'h00);
expect_elem("page program", 2, K_ADDR, 8'h01);
expect_elem("page program", 3, K_ADDR, 8'h00);
expect_elem("page program", 4, K_DATA, 8'h00);
if (first_data_cycle !== 16'd32) begin
$display(" FAIL: page program data begins at %0d, expected 32",
first_data_cycle);
errors = errors + 1;
end
show("page program ");
// 3. Four-byte addressing, the mode a part above 128 Mbit needs.
issue(1'b1, 32'h01234567, 2, 4, 8);
expect_count("4-byte address", 7);
expect_elem("4-byte address", 1, K_ADDR, 8'h01);
expect_elem("4-byte address", 2, K_ADDR, 8'h23);
expect_elem("4-byte address", 3, K_ADDR, 8'h45);
expect_elem("4-byte address", 4, K_ADDR, 8'h67);
expect_elem("4-byte address", 5, K_DUMMY, 8'd8);
if (first_data_cycle !== 16'd48) begin
$display(" FAIL: 4-byte address data begins at %0d, expected 48",
first_data_cycle);
errors = errors + 1;
end
show("4-byte address ");
// 4. THE 16 MB CEILING. An address above 0xFFFFFF cannot be
// expressed in three address bytes.
issue(1'b1, 32'h01000000, 4, 3, 8);
if (!spec_error) begin
$display(" FAIL: 0x01000000 does not fit three address bytes and was accepted");
errors = errors + 1;
end
issue(1'b1, 32'h00FFFFFF, 4, 3, 8);
if (spec_error) begin
$display(" FAIL: 0x00FFFFFF fits three address bytes exactly and was rejected");
errors = errors + 1;
end
$display(" address ceiling: 0x00FFFFFF accepted, 0x01000000 rejected -- the boundary is exact");
// 5. A status read: no address phase, no dummy.
issue(1'b1, 32'h00000000, 1, 0, 0);
expect_count("status read", 2);
expect_elem("status read", 0, K_CMD, 8'h0B);
expect_elem("status read", 1, K_DATA, 8'h00);
if (first_data_cycle !== 16'd8) begin
$display(" FAIL: status read data begins at %0d, expected 8",
first_data_cycle);
errors = errors + 1;
end
show("status read ");
// 6. AN ADC. No address phase at all -- the conversion result is
// simply clocked out -- but a leading latency while the sample
// is acquired, then a two-byte result. This is the shape most
// converters have: the dummy phase exists without any address
// before it.
issue(1'b1, 32'h00000000, 2, 0, 4);
expect_count("adc read", 3);
expect_elem("adc read", 0, K_CMD, 8'h0B);
expect_elem("adc read", 1, K_DUMMY, 8'd4);
expect_elem("adc read", 2, K_DATA, 8'h00);
if (got_len[2] !== 2) begin
$display(" FAIL: adc read data length %0d, expected 2", got_len[2]);
errors = errors + 1;
end
if (first_data_cycle !== 16'd12) begin
$display(" FAIL: adc read data begins at %0d, expected 12",
first_data_cycle);
errors = errors + 1;
end
show("adc read ");
// 7. A bare command with neither address nor data.
issue(1'b0, 32'h00000000, 0, 0, 0);
expect_count("write enable", 1);
expect_elem("write enable", 0, K_CMD, 8'h02);
show("write enable ");
if (errors == 0)
$display("PASS: every element stream matches the sequence a real device expects, addresses are emitted most-significant byte first, absent phases emit no element at all, the data cycle agrees with the latency arithmetic, and an address too wide for the address phase is rejected rather than truncated");
else
$display("FAIL: %0d error(s)", errors);
$finish;
end
endmodule-- spi_txn_spec.vhd
--
-- Chapter 10.6 -- from datasheet to transaction specification, in VHDL.
--
-- Takes the device profile of Chapter 10.1, the command convention of
-- Chapter 10.4 and the latency arithmetic of Chapter 10.5, and emits the
-- exact sequence a controller must issue: the opcode byte, the address
-- bytes in the order the device expects, the dummy phase with its CYCLE
-- count, and the data phase with its length -- plus the cycle at which
-- returned data begins.
--
-- ADDRESS ORDER. SPI devices send addresses most-significant byte first.
-- The address is pre-shifted once at start so the first byte to send sits
-- at the top of the register; every element after that is a fixed
-- top-byte extraction and a shift by eight.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_txn_spec is
generic (
ADDR_W : positive := 32;
LEN_W : positive := 16;
MAX_ABYTES : positive := 4;
CMD_READ : natural := 16#0B#; -- fast read
CMD_WRITE : natural := 16#02# -- page program
);
port (
clk : in std_logic;
rst_n : in std_logic;
start : in std_logic;
req_read : in std_logic;
req_addr : in unsigned(ADDR_W - 1 downto 0);
req_len : in unsigned(LEN_W - 1 downto 0);
-- From the device profile.
addr_bytes : in unsigned(2 downto 0);
dummy_cycles : in unsigned(5 downto 0);
elem_valid : out std_logic;
elem_kind : out unsigned(1 downto 0); -- 0=cmd 1=addr 2=dummy 3=data
elem_data : out unsigned(7 downto 0);
elem_len : out unsigned(LEN_W - 1 downto 0);
spec_error : out std_logic;
busy : out std_logic;
done : out std_logic;
first_data_cycle : out unsigned(15 downto 0)
);
end entity;
architecture rtl of spi_txn_spec is
constant K_CMD : unsigned(1 downto 0) := "00";
constant K_ADDR : unsigned(1 downto 0) := "01";
constant K_DUMMY : unsigned(1 downto 0) := "10";
constant K_DATA : unsigned(1 downto 0) := "11";
type state_t is (S_IDLE, S_CMD, S_ADDR, S_DUMMY, S_DATA);
signal state : state_t := S_IDLE;
signal addr_sr : unsigned(ADDR_W - 1 downto 0) := (others => '0');
signal abytes_r : unsigned(2 downto 0) := (others => '0');
signal dummy_r : unsigned(5 downto 0) := (others => '0');
signal len_r : unsigned(LEN_W - 1 downto 0) := (others => '0');
signal ev_r : std_logic := '0';
signal ek_r : unsigned(1 downto 0) := K_CMD;
signal ed_r : unsigned(7 downto 0) := (others => '0');
signal el_r : unsigned(LEN_W - 1 downto 0) := (others => '0');
signal err_r : std_logic := '0';
signal busy_r : std_logic := '0';
signal done_r : std_logic := '0';
signal fdc_r : unsigned(15 downto 0) := (others => '0');
begin
build : process (clk, rst_n)
variable abits : natural;
variable limit_mask : unsigned(ADDR_W - 1 downto 0);
begin
if rst_n = '0' then
state <= S_IDLE;
ev_r <= '0';
ek_r <= K_CMD;
ed_r <= (others => '0');
el_r <= (others => '0');
err_r <= '0';
busy_r <= '0';
done_r <= '0';
fdc_r <= (others => '0');
addr_sr <= (others => '0');
abytes_r <= (others => '0');
dummy_r <= (others => '0');
len_r <= (others => '0');
elsif rising_edge(clk) then
ev_r <= '0';
done_r <= '0';
case state is
when S_IDLE =>
if start = '1' then
abytes_r <= addr_bytes;
dummy_r <= dummy_cycles;
len_r <= req_len;
-- Pre-shift so the most significant address byte
-- the device expects sits at the top.
abits := to_integer(addr_bytes) * 8;
addr_sr <= shift_left(req_addr, MAX_ABYTES * 8 - abits);
-- An address wider than the device's address phase
-- cannot be expressed. This is the arithmetic
-- behind the 16 MB ceiling on three-byte
-- addressing: a controller that truncates silently
-- wraps the access to the bottom of the device,
-- returning valid-looking data from the wrong
-- place.
if abits >= ADDR_W then
limit_mask := (others => '1');
else
limit_mask := shift_left(
to_unsigned(1, ADDR_W), abits) - 1;
end if;
if (req_addr and not limit_mask) /= 0 then
err_r <= '1';
else
err_r <= '0';
end if;
-- The same arithmetic as Chapter 10.5: command
-- bits plus address bits plus dummy CYCLES.
fdc_r <= to_unsigned(8 + abits +
to_integer(dummy_cycles), 16);
-- The opcode is chosen by direction. A device whose
-- read and write opcodes differ in one bit can use
-- Chapter 10.4's codec instead; a flash, whose
-- opcodes are unrelated values, needs a table.
ev_r <= '1';
ek_r <= K_CMD;
if req_read = '1' then
ed_r <= to_unsigned(CMD_READ, 8);
else
ed_r <= to_unsigned(CMD_WRITE, 8);
end if;
el_r <= to_unsigned(1, LEN_W);
busy_r <= '1';
state <= S_CMD;
end if;
when S_CMD =>
if abytes_r /= 0 then
ev_r <= '1';
ek_r <= K_ADDR;
ed_r <= addr_sr(ADDR_W - 1 downto ADDR_W - 8);
el_r <= to_unsigned(1, LEN_W);
addr_sr <= shift_left(addr_sr, 8);
abytes_r <= abytes_r - 1;
state <= S_ADDR;
elsif dummy_r /= 0 then
state <= S_DUMMY;
elsif len_r /= 0 then
state <= S_DATA;
else
busy_r <= '0';
done_r <= '1';
state <= S_IDLE;
end if;
when S_ADDR =>
if abytes_r /= 0 then
ev_r <= '1';
ek_r <= K_ADDR;
ed_r <= addr_sr(ADDR_W - 1 downto ADDR_W - 8);
el_r <= to_unsigned(1, LEN_W);
addr_sr <= shift_left(addr_sr, 8);
abytes_r <= abytes_r - 1;
elsif dummy_r /= 0 then
state <= S_DUMMY;
elsif len_r /= 0 then
state <= S_DATA;
else
busy_r <= '0';
done_r <= '1';
state <= S_IDLE;
end if;
when S_DUMMY =>
-- One element carrying the CYCLE count -- not a byte.
-- Emitting a dummy BYTE is the error this whole module
-- exists to prevent.
ev_r <= '1';
ek_r <= K_DUMMY;
ed_r <= resize(dummy_r, 8);
el_r <= to_unsigned(1, LEN_W);
if len_r /= 0 then
state <= S_DATA;
else
busy_r <= '0';
done_r <= '1';
state <= S_IDLE;
end if;
when others => -- S_DATA
ev_r <= '1';
ek_r <= K_DATA;
ed_r <= (others => '0');
el_r <= len_r;
busy_r <= '0';
done_r <= '1';
state <= S_IDLE;
end case;
end if;
end process;
elem_valid <= ev_r;
elem_kind <= ek_r;
elem_data <= ed_r;
elem_len <= el_r;
spec_error <= err_r;
busy <= busy_r;
done <= done_r;
first_data_cycle <= fdc_r;
end architecture;-- spi_txn_spec_tb.vhd
--
-- The same checks as the SystemVerilog and Verilog testbenches: the
-- element stream is compared element for element against a sequence
-- computed by hand from a real datasheet reading, not against the DUT's
-- own arithmetic.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_txn_spec_tb is
end entity;
architecture sim of spi_txn_spec_tb is
constant ADDR_W : positive := 32;
constant LEN_W : positive := 16;
constant K_CMD : unsigned(1 downto 0) := "00";
constant K_ADDR : unsigned(1 downto 0) := "01";
constant K_DUMMY : unsigned(1 downto 0) := "10";
constant K_DATA : unsigned(1 downto 0) := "11";
signal clk : std_logic := '0';
signal rst_n : std_logic := '0';
signal halt : boolean := false;
signal start : std_logic := '0';
signal req_read : std_logic := '0';
signal req_addr : unsigned(ADDR_W - 1 downto 0) := (others => '0');
signal req_len : unsigned(LEN_W - 1 downto 0) := (others => '0');
signal addr_bytes : unsigned(2 downto 0) := (others => '0');
signal dummy_cycles : unsigned(5 downto 0) := (others => '0');
signal elem_valid : std_logic;
signal elem_kind : unsigned(1 downto 0);
signal elem_data : unsigned(7 downto 0);
signal elem_len : unsigned(LEN_W - 1 downto 0);
signal spec_error : std_logic;
signal busy : std_logic;
signal done : std_logic;
signal first_data_cycle : unsigned(15 downto 0);
type kind_arr is array (0 to 15) of unsigned(1 downto 0);
type data_arr is array (0 to 15) of unsigned(7 downto 0);
type len_arr is array (0 to 15) of unsigned(LEN_W - 1 downto 0);
signal got_kind : kind_arr;
signal got_data : data_arr;
signal got_len : len_arr;
signal got_n : natural := 0;
signal clr_n : std_logic := '0';
signal errors : natural := 0;
begin
clk <= not clk after 5 ns when not halt else '0';
dut : entity work.spi_txn_spec
generic map (ADDR_W => ADDR_W, LEN_W => LEN_W, MAX_ABYTES => 4,
CMD_READ => 16#0B#, CMD_WRITE => 16#02#)
port map (
clk => clk, rst_n => rst_n, start => start,
req_read => req_read, req_addr => req_addr, req_len => req_len,
addr_bytes => addr_bytes, dummy_cycles => dummy_cycles,
elem_valid => elem_valid, elem_kind => elem_kind,
elem_data => elem_data, elem_len => elem_len,
spec_error => spec_error, busy => busy, done => done,
first_data_cycle => first_data_cycle
);
-- Collect every emitted element. The clear request arrives on its own
-- signal rather than the stimulus process writing got_n directly: two
-- processes driving one signal is a multiple-driver error in VHDL.
collect : process (clk)
begin
if rising_edge(clk) then
if clr_n = '1' then
got_n <= 0;
elsif rst_n = '1' and elem_valid = '1' and got_n < 16 then
got_kind(got_n) <= elem_kind;
got_data(got_n) <= elem_data;
got_len(got_n) <= elem_len;
got_n <= got_n + 1;
end if;
end if;
end process;
stim : process
variable errs : natural := 0;
procedure issue(rd : std_logic; a : natural; len : natural;
ab : natural; dc : natural) is
begin
wait until falling_edge(clk);
clr_n <= '1';
wait until falling_edge(clk);
clr_n <= '0';
req_read <= rd;
req_addr <= to_unsigned(a, ADDR_W);
req_len <= to_unsigned(len, LEN_W);
addr_bytes <= to_unsigned(ab, 3);
dummy_cycles <= to_unsigned(dc, 6);
start <= '1';
wait until falling_edge(clk);
start <= '0';
while busy = '1' loop
wait until falling_edge(clk);
end loop;
wait until falling_edge(clk);
end procedure;
procedure expect_elem(name : string; idx : natural;
k : unsigned(1 downto 0); d : natural) is
begin
if idx >= got_n then
report " FAIL: " & name & " has only " &
integer'image(got_n) & " elements, wanted one at " &
integer'image(idx);
errs := errs + 1;
elsif got_kind(idx) /= k or to_integer(got_data(idx)) /= d then
report " FAIL: " & name & " element " & integer'image(idx) &
" is kind=" &
integer'image(to_integer(got_kind(idx))) & " data=" &
integer'image(to_integer(got_data(idx))) &
", expected kind=" &
integer'image(to_integer(k)) & " data=" &
integer'image(d);
errs := errs + 1;
end if;
end procedure;
procedure expect_count(name : string; n : natural) is
begin
if got_n /= n then
report " FAIL: " & name & " emitted " &
integer'image(got_n) & " elements, expected " &
integer'image(n);
errs := errs + 1;
end if;
end procedure;
procedure expect_fdc(name : string; n : natural) is
begin
if to_integer(first_data_cycle) /= n then
report " FAIL: " & name & " data begins at " &
integer'image(to_integer(first_data_cycle)) &
", expected " & integer'image(n);
errs := errs + 1;
end if;
end procedure;
begin
for i in 0 to 2 loop
wait until falling_edge(clk);
end loop;
rst_n <= '1';
wait until falling_edge(clk);
-- 1. A flash fast read, 0x0B, three address bytes, eight dummy
-- cycles, four bytes of payload.
issue('1', 16#123456#, 4, 3, 8);
expect_count("fast read", 6);
expect_elem("fast read", 0, K_CMD, 16#0B#);
expect_elem("fast read", 1, K_ADDR, 16#12#); -- MSB first
expect_elem("fast read", 2, K_ADDR, 16#34#);
expect_elem("fast read", 3, K_ADDR, 16#56#);
expect_elem("fast read", 4, K_DUMMY, 8);
expect_elem("fast read", 5, K_DATA, 0);
if to_integer(got_len(5)) /= 4 then
report " FAIL: fast read data length is wrong"; errs := errs + 1;
end if;
expect_fdc("fast read", 40);
if spec_error = '1' then
report " FAIL: a 24-bit address in a 3-byte phase was rejected";
errs := errs + 1;
end if;
report " fast read: CMD 0x0B | ADDR 0x12 | ADDR 0x34 | ADDR 0x56 | DUMMY 8c | DATA 4B (data begins at cycle 40)";
-- 2. A page program: a different opcode, no dummy phase at all.
-- The dummy element must be ABSENT, not present with a count of
-- zero -- a consumer that sees a dummy element will wait.
issue('0', 16#000100#, 4, 3, 0);
expect_count("page program", 5);
expect_elem("page program", 0, K_CMD, 16#02#);
expect_elem("page program", 1, K_ADDR, 16#00#);
expect_elem("page program", 2, K_ADDR, 16#01#);
expect_elem("page program", 3, K_ADDR, 16#00#);
expect_elem("page program", 4, K_DATA, 0);
expect_fdc("page program", 32);
report " page program: CMD 0x02 | ADDR 0x00 | ADDR 0x01 | ADDR 0x00 | DATA 4B (data begins at cycle 32)";
-- 3. Four-byte addressing, the mode a part above 128 Mbit needs.
issue('1', 16#01234567#, 2, 4, 8);
expect_count("4-byte address", 7);
expect_elem("4-byte address", 1, K_ADDR, 16#01#);
expect_elem("4-byte address", 2, K_ADDR, 16#23#);
expect_elem("4-byte address", 3, K_ADDR, 16#45#);
expect_elem("4-byte address", 4, K_ADDR, 16#67#);
expect_elem("4-byte address", 5, K_DUMMY, 8);
expect_fdc("4-byte address", 48);
report " 4-byte address: CMD 0x0B | ADDR 0x01 | ADDR 0x23 | ADDR 0x45 | ADDR 0x67 | DUMMY 8c | DATA 2B (data begins at cycle 48)";
-- 4. THE 16 MB CEILING. An address above 0xFFFFFF cannot be
-- expressed in three address bytes.
issue('1', 16#01000000#, 4, 3, 8);
if spec_error /= '1' then
report " FAIL: 0x01000000 does not fit three address bytes and was accepted";
errs := errs + 1;
end if;
issue('1', 16#00FFFFFF#, 4, 3, 8);
if spec_error = '1' then
report " FAIL: 0x00FFFFFF fits three address bytes exactly and was rejected";
errs := errs + 1;
end if;
report " address ceiling: 0x00FFFFFF accepted, 0x01000000 rejected -- the boundary is exact";
-- 5. A status read: no address phase, no dummy.
issue('1', 0, 1, 0, 0);
expect_count("status read", 2);
expect_elem("status read", 0, K_CMD, 16#0B#);
expect_elem("status read", 1, K_DATA, 0);
expect_fdc("status read", 8);
report " status read: CMD 0x0B | DATA 1B (data begins at cycle 8)";
-- 6. AN ADC. No address phase at all -- the conversion result is
-- simply clocked out -- but a leading latency while the sample
-- is acquired, then a two-byte result. This is the shape most
-- converters have: the dummy phase exists without any address
-- before it.
issue('1', 0, 2, 0, 4);
expect_count("adc read", 3);
expect_elem("adc read", 0, K_CMD, 16#0B#);
expect_elem("adc read", 1, K_DUMMY, 4);
expect_elem("adc read", 2, K_DATA, 0);
if to_integer(got_len(2)) /= 2 then
report " FAIL: adc read data length is wrong"; errs := errs + 1;
end if;
expect_fdc("adc read", 12);
report " adc read: CMD 0x0B | DUMMY 4c | DATA 2B (data begins at cycle 12)";
-- 7. A bare command with neither address nor data.
issue('0', 0, 0, 0, 0);
expect_count("write enable", 1);
expect_elem("write enable", 0, K_CMD, 16#02#);
report " write enable: CMD 0x02 (command only)";
errors <= errs;
if errs = 0 then
report "PASS: every element stream matches the sequence a real device expects, addresses are emitted most-significant byte first, absent phases emit no element at all, the data cycle agrees with the latency arithmetic, and an address too wide for the address phase is rejected rather than truncated";
else
report "FAIL: " & integer'image(errs) & " error(s)" severity error;
end if;
halt <= true;
wait;
end process;
end architecture;Parity
All three implement the same engine: identical ports and generics, an address pre-shifted once at start and emitted most-significant byte first, a dummy phase emitted as a cycle count rather than bytes, absent phases emitting no element, a published first-data cycle, and an exact address-width check. All three testbenches drive the same seven requests and produce identical element sequences — including the ADC read as three elements with data beginning at cycle 12, and the address ceiling rejecting 0x01000000 while accepting 0x00FFFFFF.
7. Why a Verification Engineer Cares
This is where the UVM environment for a data-described device comes together. Chapter 10.1 built the configuration object; here everything else hangs off it.
// The reference model is the point of the whole module. It predicts the
// element sequence from the PROFILE -- the same seven numbers the RTL
// is configured from -- so a new device means a new config object and
// no new scoreboard code at all.
class spi_spec_model extends uvm_component;
`uvm_component_utils(spi_spec_model)
spi_device_cfg cfg;
function void build_phase(uvm_phase phase);
super.build_phase(phase);
if (!uvm_config_db #(spi_device_cfg)::get(this, "", "cfg", cfg))
`uvm_fatal("CFG", "no device profile supplied")
endfunction
// Predict, from the profile alone, what the controller must issue.
function void predict(spi_req_item req, ref spi_elem_q exp);
exp.delete();
if (cfg.has_opcode)
exp.push_back(elem_cmd(req.read ? cfg.op_read : cfg.op_write));
for (int i = cfg.addr_bytes - 1; i >= 0; i--) // MSB first
exp.push_back(elem_addr((req.addr >> (8*i)) & 8'hFF));
if (cfg.dummy_cycles != 0)
exp.push_back(elem_dummy(cfg.dummy_cycles)); // CYCLES
if (req.len != 0)
exp.push_back(elem_data(req.len));
endfunction
endclass
// The scoreboard compares the observed stream against the prediction.
// Note what it does NOT contain: any knowledge of any particular
// device. Every device-specific fact lives in the config object.
class spi_spec_sb extends uvm_scoreboard;
`uvm_component_utils(spi_spec_sb)
function void write_observed(spi_elem_item obs);
if (exp_q.size() == 0)
`uvm_error("SB", "an element was issued that the profile does not call for")
else begin
spi_elem_item e = exp_q.pop_front();
if (!obs.compare(e))
`uvm_error("SB", $sformatf("expected %s, observed %s",
e.convert2string(),
obs.convert2string()))
end
endfunction
// The check that catches a MISSING element -- the failure a
// per-element comparison structurally cannot see, because it only
// ever compares the elements that arrived.
function void check_phase(uvm_phase phase);
if (exp_q.size() != 0)
`uvm_error("SB", $sformatf("%0d expected elements never issued",
exp_q.size()))
endfunction
endclassThe check_phase method is the part worth copying. A scoreboard that compares each observed item against the head of a queue detects wrong items perfectly and missing items not at all — because a missing item produces no call. The end-of-test queue check is what closes that hole, and its absence is one of the most common real defects in a working UVM environment.
// A layered sequence. The test asks for a register read; the sequence
// library turns that into a request item; the spec engine turns THAT
// into elements. No layer above the engine knows the dummy count.
class spi_read_seq extends uvm_sequence #(spi_req_item);
`uvm_object_utils(spi_read_seq)
rand bit [31:0] addr;
rand int len;
spi_device_cfg cfg;
// The address must fit the profile's address phase. Constraining it
// HERE means the ceiling is enforced by construction in normal
// tests, and an error test can deliberately relax it.
constraint c_fits {
cfg.addr_bytes == 3 -> addr < 32'h0100_0000;
cfg.addr_bytes == 4 -> addr <= 32'hFFFF_FFFF;
len inside {[1:256]};
}
task body();
spi_req_item req = spi_req_item::type_id::create("req");
start_item(req);
req.read = 1'b1; req.addr = addr; req.len = len;
finish_item(req);
endtask
endclassThe discipline this buys is that adding a device is a new spi_device_cfg and nothing else. No new sequence, no new driver, no new scoreboard — and, crucially, no new opportunity to encode one device's dummy count somewhere a second device will inherit it.
// 1. The element sequence is exactly what the profile describes: one
// opcode, addr_bytes address elements, a dummy element only if the
// count is non-zero, and a data element only if the length is.
a_element_count : assert property (
@(posedge clk) disable iff (!rst_n)
(done) |-> (elems_issued == 1 + addr_bytes_r
+ (dummy_r != 0)
+ (len_r != 0)))
else $error("the element count does not match the profile");
// 2. ADDRESS ORDER. The first address element is the most significant
// byte. Reversed order reaches a valid but wildly wrong location,
// so the device answers and the data merely looks corrupted.
a_msb_first : assert property (
@(posedge clk) disable iff (!rst_n)
(first_addr_elem) |-> (elem_data == req_addr_at_start[
8*addr_bytes_r-1 -: 8]))
else $error("the address was not emitted most significant byte first");
// 3. NO EMPTY ELEMENTS. An absent phase emits nothing -- an element
// with a zero count would make the consumer wait for zero cycles,
// which several consumers implement as waiting forever.
a_no_empty_elem : assert property (
@(posedge clk) disable iff (!rst_n)
(elem_valid && elem_kind == K_DUMMY) |-> (elem_data != 0))
else $error("a dummy element was issued with a count of zero");
// 4. The published data cycle agrees with the elements actually
// emitted -- the two must never be computed independently.
a_latency_agrees : assert property (
@(posedge clk) disable iff (!rst_n)
(done) |-> (first_data_cycle == 8 + 8*addr_bytes_r + dummy_r))
else $error("the published latency disagrees with the emitted phases");Coverage, at this level, is about device shapes rather than values:
covergroup spi_spec_cg @(posedge start);
// The four phases, present or absent. Sixteen combinations exist on
// paper; the eight reachable ones are all real devices, and a suite
// that only sends full four-phase transfers has tested one.
cp_shape : coverpoint {has_opcode, addr_bytes != 0,
dummy_cycles != 0, len != 0} {
bins full_read = {4'b1111}; // flash fast read
bins no_dummy = {4'b1101}; // page program
bins no_addr = {4'b1011}; // ADC, buffered read
bins cmd_only = {4'b1000}; // write enable
bins status = {4'b1001}; // status read
bins erase = {4'b1100}; // address, no data
}
cp_abytes : coverpoint addr_bytes { bins b[] = {0, 1, 2, 3, 4}; }
// The address ceiling, exactly. One apart, opposite outcomes.
cp_ceiling : coverpoint addr_class {
bins fits_exactly = {ADDR_MAX_FOR_WIDTH};
bins one_over = {ADDR_MAX_FOR_WIDTH + 1};
bins well_under = {ADDR_TYPICAL};
bins well_over = {ADDR_FAR_OVER};
}
x_shape_abytes : cross cp_shape, cp_abytes;
endgroup8. Why an FPGA or ASIC Engineer Cares
Separate what to send from how to send it. The element-level engine and the pin-level master are different blocks with a clean interface between them. One master then serves every device on the board, and a new part is a profile.
Publish the latency once. Every consumer that needs to know when data begins reads the same number from the same register. Two modules computing it separately is how a design comes to disagree with itself about a value that is correct in both.
Emit nothing for an absent phase. A zero-count element is worse than no element, because a consumer written to wait for the count will wait for zero cycles — and several natural implementations of "wait for N" treat zero as "wait forever".
Pre-shift the address once. One barrel shifter at load beats a variable shift per byte, and the pattern generalises to any variable-width field emitted from a fixed-width register.
Check the address width in hardware. The device cannot: the bits are not in the transfer. This is the last line of defence against the aliasing failure of Chapter 10.5, and it costs a mask and a comparator.
Make the post-processing part of the specification. The (raw >> 4) & 0x0FFF of §5 is as much a part of the interface as the dummy count. Left out, the first integration reads a left-aligned result as right-aligned and is wrong by a factor of sixteen while remaining a legal number.
9. Failure Signature — A Second Device That Needs "Just a Small Driver Change"
Symptom. A working controller supports one flash. A second part is added — a different vendor, similar capacity. The driver is extended with a few conditionals. The new part works. Some weeks later the original part begins failing intermittently on a fraction of boards, with no change to its own code path.
What "no change to its own code path" establishes. If the original part's code is genuinely untouched, the shared state between the two paths is where the fault must be. In a driver full of conditionals, that shared state is larger than it appears.
Plausible mechanisms.
- A shared configuration register — divisor, mode, or dummy count — set for whichever device was accessed last and not restored. The original part then runs with the second part's dummy count, which shifts its payload.
- A shared buffer or descriptor whose size assumption differs between the parts.
- A 3-byte/4-byte mode left set by the second device's initialisation, which Chapter 10.5 noted is persistent and survives a warm reset.
- A genuine timing marginality revealed by the added bus load of the second device — Chapter 9.4's observation that every device lowers the ceiling for all of them.
- Ordering, where the failure depends on which part was accessed first after boot.
The discriminating observation. "On a fraction of boards" and "intermittent" point away from a pure logic error and towards either a race or a marginality. But the strongest single test is an ordering experiment: access the second device, then the first, and see whether the first fails deterministically. If it does, the fault is shared state and not timing — and the intermittency was simply the natural variation in access order.
If ordering makes no difference, measure: the monitor of Chapter 10.3 will say whether the setup and hold margins collapsed when the second device was added, which is the signature of the added-load mechanism.
The fix, and the architectural point. Give each device a profile and apply it as part of selecting the device, atomically, the way Chapter 10.1's block does. The conditionals in the driver are the disease, not the treatment: every if (device == B) is a place where one device's fact can leak into the other's path, and they multiply with each new part.
Why this pattern recurs. Because adding the second device by extending the first device's driver is always the smaller change at the time. The cost is paid later, by the first device, in a form that does not look related — which is exactly why the whole module argues for a profile rather than a driver.
10. Common Misconceptions
11. Reason It Through
Work this before reading the answer.
The ADC of §1–5 is integrated. The specification is followed exactly. Readings are stable and plausible — but when the input is shorted to ground the converter reports about 200 counts instead of 0, and at full scale it reports about 4095 as expected.
The link is correct. What is wrong, and what does the full-scale reading being right tell you?
Start with what full scale being correct rules out. A shifted payload, a wrong dummy count or a wrong alignment would corrupt both ends of the range, not one. A left-aligned result read as right-aligned would report 65 520 at full scale, not 4095. So the transaction specification is being executed correctly, and this is not a protocol fault at all.
So the error is an offset, roughly 200 counts, present at zero and invisible at full scale. That framing is the key move: an offset that vanishes at the top of the range is not an offset in the digital path, because a digital offset would shift everything equally and full scale would read 4295 — which is unrepresentable and would clip to 4095.
And there is the answer to the second question. Full scale reading 4095 may not mean "correct" at all; it may mean clipped. A uniform offset of 200 counts would be invisible at the top precisely because the result saturates. So the full-scale reading tells you much less than it appears to, and taking it as evidence that the high end is fine is the trap.
What produces a uniform offset? Not the SPI link — it reproduces whatever the converter produced. The candidates are all analogue or configuration:
- A reference voltage lower than assumed, which scales rather than offsets, so it fits poorly.
- An input offset in the driving amplifier, or a ground difference between the source and the ADC's ground pin — which fits exactly.
- A configured input range that is bipolar rather than unipolar, so ground sits partway up the scale. This fits very well and is a configuration error, not a hardware one.
- Insufficient acquisition time, so the sampling capacitor has not settled to the input. This produces an error that depends on the previous sample — and that is a testable prediction.
How to discriminate, in order.
- Alternate the input between ground and full scale and see whether the zero reading depends on what preceded it. If shorting to ground after a full-scale reading gives 200 counts but after another ground reading gives 5, the sampling capacitor is not settling — and the cause is back in this module's territory, because §4's four-cycle acquisition may not be enough at 16 MHz for a high-impedance source.
- Check the configured input range. A bipolar setting puts ground at mid-scale, not 200 counts, so this is quickly confirmed or eliminated.
- Measure the voltage at the ADC's input pin, not at the signal source, to settle the ground-difference case.
And the connection back to this chapter. Case 1 is the one worth taking seriously, because the acquisition time in §4 was read as "4 SCLK cycles" and converted into a dummy phase — a cycle count. But acquisition is a physical settling time, in nanoseconds, and four cycles at 16 MHz is 250 ns while four cycles at 1 MHz is 4 µs. A number specified in cycles at one clock rate does not remain the same number of nanoseconds at another.
The general lesson, and it is the last one of the module. Check whether each datasheet number is genuinely a count of cycles or a time that the datasheet happened to express in cycles at some assumed rate. Dummy cycles for an internal pipeline really are cycles. An acquisition time really is nanoseconds. Treating the second as the first produces a design that works at the rate the datasheet assumed and degrades as you speed up — with an error that looks analogue and lives in the interface.
12. Understanding Check
13. Summary
The module's route ends in an artifact: a mode, a divisor, a phase schedule, an element sequence and a post-processing rule, all of them numbers.
Working the converter end to end showed each derivation in turn. The mode came from the idle level plus one conversion step, because the datasheet described the device's output edge. The rate came from the round-trip budget, where the device's own 20 MHz never bound and the divisor ladder rounded 16.1 down to 16. The phase schedule came from a sentence that never used the word dummy — four cycles of acquisition, which is not a byte and would lose the top four bits of a left-aligned result if it were sent as one.
The deliverable is a specification because it can be executed and checked, and because a reviewer can compare it against the datasheet without reading code. The (raw >> 4) & 0x0FFF belongs in it as firmly as the dummy count.
In hardware, the engine turns a profile and a request into elements, skipping absent phases entirely, emitting the address most-significant byte first from a register pre-shifted once, and checking the address against a width the device physically cannot check.
In UVM the same profile is the configuration object, the reference model predicts from it, and adding a device becomes a new config object and nothing else — with the scoreboard's end-of-test queue check closing the hole that per-item comparison structurally cannot see.
And the last caution of the module: confirm whether each datasheet number is a count of cycles or a time the datasheet happened to express in cycles. The two behave identically until you change the clock.
14. What Comes Next
This closes Module 10, and with it the transferable half of the track. The route through a datasheet, the two observations that give the mode, the timing table's owners and conditions, the command encoding's polarity, the address and dummy arithmetic, and the specification they assemble into — all of it applies to parts that have nothing else in common.
Module 11 — SPI Flash and Boot turns from the general to the specific. Serial flash is the device most engineers meet first and the one with the most behaviour hiding behind a simple interface: erase before write, program granularity, busy polling, protection registers, and a set of commands whose durations vary by four orders of magnitude between them. Everything this module built is what makes that tractable — a flash is simply a device whose profile has more entries, and whose boot path depends on a controller getting every one of them right before any software runs.
Continue learning
Related tutorials
- Related topic
How to Read an SPI Device Datasheet
A repeatable four-pass route through any SPI peripheral datasheet — pins, timing, commands, registers — ending in the seven numbers a controller needs, and the register block that validates them atomically so no half-applied profile can ever be in force.
- Related topic
CS-to-SCLK and SCLK-to-CS Timing
Chip select has timing requirements of its own: the lead before the first clock edge, the lag after the last, and the minimum deselect between transactions. Why violating them breaks a transfer whose every SCLK edge was correct.
- Related topic
Anatomy of a Read Transaction
The four phases of a device read, why a read cannot be a write reversed, when the slave takes and releases MISO, the obligation to have the first data bit valid before any edge can launch it, and the slave read datapath in three HDLs.
- Related topic
Continuous Transfers Under One CS
Holding chip select low across many bytes and what the device assumes: why the frame is the transaction, why the master may legally stop the clock when its data runs dry, and the streaming controller in three HDLs.
