SPI · Module 15
Choosing an Architecture
Both architectures on the same pins with the same frame width, capture edge and stimulus, so the table differs in exactly one thing. It has three regions, not two — the third stops 'A is faster' from becoming 'A has no limit' — plus two warnings a ratio table structurally cannot carry.
Chapter 15.2 built Architecture A and measured its per-word requirement. Chapter 15.3 took Architecture B's per-bit precondition apart. Neither chapter said which to choose.
Two designs, same function, same frame width, same capture edge, same pins, same stimulus. Which one should a given system use?
The answer is arithmetic rather than preference, and this chapter computes it — but the table that computes it needs two warnings printed beside it, because a ratio table structurally cannot carry them.
1. Making The Comparison Fair
Module 14's slave is Architecture B and it is nine blocks. Putting it next to Chapter 15.2's single block would compare a product with a mechanism.
So this chapter builds the smallest honest Architecture B slave: recover SCLK's edges on the system clock, shift MOSI in on the capture edge, publish a word every LEN_FIXED captures. Same ports as Architecture A's, same frame width, same reversal, same reported overrun — and one clock domain instead of two.
What is deliberately the same, so that the comparison is about the clocking:
* the frame width and capture edge are PARAMETERS in both, even though
Architecture B could make them run-time inputs for two gates (Chapter 14.6).
* the word is published through the same kind of register, and `rx_overrun`
means the same thing.Fixing B's mode removes a genuine Architecture B advantage from the measurement. That is done on purpose, so the ratio result stands on its own — and §5 puts the advantage back in words, where it can be weighed rather than smuggled into a table.
What is necessarily different, and is the whole result:
* B has no second clock domain, so no synchroniser on the data path, no
hand-off lifetime question, and no reset that has to cross anything;
* and B's SCLK must be slow enough to oversample, which Chapter 15.3 measured
exactly -- below one system clock per half-period the edges are gone or aliased.There is also an asymmetry in how each one fails, and it matters for diagnosis: Architecture B fails by not seeing the bus; Architecture A fails by not being read in time.
2. The Block Diagram
3. The Table, And Its Three Regions
Four words per run at a fixed 10 ns system clock, SYNC_N = 2, an 8-bit frame:
sclk_half clocks/half word_int archA archB A_ovr B_ovr
40 ns 4.0 640 ok ok 0 0
15 ns 1.5 240 ok ok 0 0
10 ns 1.0 160 ok ok 0 0
5 ns 0.5 80 ok FAIL 0 0
2 ns 0.2 32 ok FAIL 0 0
1 ns 0.1 16 FAIL FAIL 1 0Region 1 — both work. The top three rows. Both preconditions hold, both deliver four correct words, and the choice is made on things a table cannot show.
Region 2 — only A works. Rows four and five. Architecture B's edges are gone or aliased (Chapter 15.3 measured exactly this), and Architecture A keeps working — at a 4 ns SCLK period, which is SCLK two and a half times faster than the system clock.
Region 3 — neither works. The last row. A word arrives every 16 ns and Architecture A needs SYNC_N × 10 = 20 ns, so its per-word requirement fails too.
Region 3 is in the table on purpose. Without it, "Architecture A is faster" gets remembered as "Architecture A has no limit", and the per-word requirement is then discovered by a customer.
The crossover, as two numbers
against a 10 ns system clock, HALF_MIN = 3, SYNC_N = 2, an 8-bit frame:
Architecture B needs SCLK period >= 60 ns
Architecture A needs SCLK period >= 2.5 ns
a ratio of 2 x HALF_MIN x LEN / SYNC_N = 24That factor is where the whole choice is decided, and note what supplies most of it: the frame width. At a one-bit frame the relief would be 3, and Architecture A's advantage would largely evaporate — which is worth knowing before choosing it for a bit-oriented protocol.
4. Two Warnings The Table Cannot Carry
5. What The Table Structurally Cannot Show
Both designs were given a fixed mode and frame width so the comparison is only about the clocking. Putting the advantage back:
Architecture A Architecture B
------------------------ ------------------------ ------------------------
mode at run time NO -- a mux on a clock yes, two gates (14.6)
frame width at run time a crossing into SCLK yes, a register read
clock domains two one
data-path synchroniser none (hand-off + flag) three, equal depth
reset crossings one (15.2 section 4) none
SCLK in the timing setup a clock: tree, constraint data: input delay only
SCLK pin requirement clock-capable any I/O pin
fails by not being READ in time not SEEING the bus
flags its own failure yes (rate proxy) no -- needs min_halfRead that table next to §3's and the decision becomes concrete:
Compute both preconditions from the two clock frequencies and the frame width. If B's holds, prefer B — for its single domain, its run-time mode and its absent hand-off. If it does not, A is not a preference but the only option.
The asymmetry in the last two rows is worth dwelling on. Architecture A's failure is reported (albeit by a rate proxy, Chapter 15.2 §3), and Architecture B's is not — B needs a separate measurement that a designer has to remember to publish. So a system that must diagnose its own faults in the field has a reason to prefer A that has nothing to do with frequency.
6. Building The Comparison — Three HDLs
The circuit
The minimal Architecture B slave, written to be comparable rather than complete: three equal-depth synchronisers, a capture strobe defined against the mode parameter, one shift register and one publication register.
// spi_slave_oversampled.sv
//
// Chapter 15.4 -- Architecture B, reduced to exactly what Architecture A does, so
// that the two can be compared without the comparison being about anything else.
//
// Module 14's slave is Architecture B and it is nine blocks. Putting it next to
// Chapter 15.2's single block would compare a product with a mechanism. So this
// file is the smallest honest Architecture B slave: recover SCLK's edges on the
// system clock, shift MOSI in on the capture edge, publish a word every LEN_FIXED
// captures. Same ports as Chapter 15.2's, same frame width, same reversal, same
// reported overrun -- and one clock domain instead of two.
//
// WHAT IS DELIBERATELY THE SAME, so the comparison is fair:
//
// * the frame width and the capture edge are PARAMETERS here too, even though
// Architecture B could make them run-time inputs for two gates (Chapter 14.6).
// Leaving them fixed removes a genuine Architecture B advantage from the
// measurement so that the ratio result stands on its own -- and Section 5 of
// the chapter puts the advantage back in words, where it can be weighed rather
// than smuggled into a table.
// * the word is published through the same kind of register, and `rx_overrun`
// means the same thing.
//
// WHAT IS NECESSARILY DIFFERENT, and is the whole result:
//
// * there is no second clock domain, so no synchroniser on the data path, no
// hand-off lifetime question, and no reset that has to cross anything;
// * and SCLK must be slow enough to oversample, which Chapter 15.3 measured
// exactly: below one system clock per half-period the edges are not merely
// late, they are gone or aliased.
//
// The overrun here can only be caused by the system side, never by the ratio,
// because there is no ratio to be caught out by -- which is the asymmetry the
// comparison is about. Architecture B fails by not seeing the bus; Architecture A
// fails by not being read in time.
module spi_slave_oversampled #(
parameter int MAX_W = 32,
parameter int LEN_W = 6,
parameter int SYNC_N = 2,
parameter bit CAP_ON_RISING = 1'b1,
parameter bit LSB_FIRST = 1'b0,
parameter int LEN_FIXED = 8
) (
// --- the pins. sclk_pin is DATA here, not a clock. ---------------------
input wire sclk_pin,
input wire cs_n_pin,
input wire mosi_pin,
// --- one domain -------------------------------------------------------
input wire clk,
input wire rst_n,
output reg [MAX_W-1:0] rx_data,
output reg rx_valid_stb,
output reg rx_overrun,
input wire clr_flags
);
localparam [LEN_W-1:0] LAST_BIT = LEN_FIXED - 1;
// Three synchronisers of EQUAL depth, which is Chapter 14.1's delay-matching
// requirement and is the one structural obligation Architecture B has that
// Architecture A does not: here the data is sampled by the system clock, so it
// must arrive with the same latency as the edge that says to sample it.
reg [SYNC_N-1:0] sclk_sr, cs_sr, mosi_sr;
reg sclk_d;
wire sclk_q = sclk_sr[SYNC_N-1];
wire cs_act = ~cs_sr[SYNC_N-1];
wire mosi_q = mosi_sr[SYNC_N-1];
// The capture edge, defined against the parameter rather than against rising
// and falling -- the same convention as Chapter 14.1, and here it costs a
// constant rather than a mux because the mode is fixed.
wire cap_stb = cs_act & (CAP_ON_RISING ? (sclk_q & ~sclk_d)
: (~sclk_q & sclk_d));
reg [MAX_W-1:0] sr;
reg [LEN_W-1:0] bit_idx;
reg taken; // the published word has been retired
wire [MAX_W-1:0] sr_next = {sr[MAX_W-2:0], mosi_q};
wire word_done = cap_stb & (bit_idx == LAST_BIT);
function automatic [MAX_W-1:0] reverse_low(input [MAX_W-1:0] v);
integer i;
begin
reverse_low = {MAX_W{1'b0}};
for (i = 0; i < LEN_FIXED; i = i + 1)
reverse_low[i] = v[LEN_FIXED - 1 - i];
end
endfunction
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
sclk_sr <= {SYNC_N{1'b0}};
cs_sr <= {SYNC_N{1'b1}}; // deselected, not selected
mosi_sr <= {SYNC_N{1'b0}};
sclk_d <= 1'b0;
sr <= {MAX_W{1'b0}};
bit_idx <= {LEN_W{1'b0}};
taken <= 1'b1;
rx_data <= {MAX_W{1'b0}};
rx_valid_stb <= 1'b0;
rx_overrun <= 1'b0;
end else begin
sclk_sr <= {sclk_sr[SYNC_N-2:0], sclk_pin};
cs_sr <= {cs_sr[SYNC_N-2:0], cs_n_pin};
mosi_sr <= {mosi_sr[SYNC_N-2:0], mosi_pin};
sclk_d <= sclk_q;
if (clr_flags)
rx_overrun <= 1'b0;
// The bit counter takes its zero from the SELECT, which is Chapter
// 14.3's rule and is the one thing both architectures share exactly.
if (!cs_act) begin
sr <= {MAX_W{1'b0}};
bit_idx <= {LEN_W{1'b0}};
end else if (cap_stb) begin
sr <= sr_next;
if (bit_idx == LAST_BIT)
bit_idx <= {LEN_W{1'b0}};
else
bit_idx <= bit_idx + 1'b1;
end
rx_valid_stb <= word_done;
if (word_done) begin
rx_data <= LSB_FIRST ? reverse_low(sr_next) : sr_next;
// An overrun here means the SYSTEM did not retire the previous
// word, never that the bus was too fast -- there is no bus-side
// way to overrun a design that samples the bus.
if (!taken)
rx_overrun <= 1'b1;
taken <= 1'b0;
end else if (rx_valid_stb) begin
taken <= 1'b1;
end
end
end
endmodule// spi_slave_oversampled.v
//
// Chapter 15.4 -- Architecture B, reduced to exactly what Architecture A does, so
// that the two can be compared without the comparison being about anything else.
//
// Module 14's slave is Architecture B and it is nine blocks. Putting it next to
// Chapter 15.2's single block would compare a product with a mechanism. So this
// file is the smallest honest Architecture B slave: recover SCLK's edges on the
// system clock, shift MOSI in on the capture edge, publish a word every LEN_FIXED
// captures. Same ports as Chapter 15.2's, same frame width, same reversal, same
// reported overrun -- and one clock domain instead of two.
//
// WHAT IS DELIBERATELY THE SAME, so the comparison is fair:
//
// * the frame width and the capture edge are PARAMETERS here too, even though
// Architecture B could make them run-time inputs for two gates (Chapter 14.6).
// Leaving them fixed removes a genuine Architecture B advantage from the
// measurement so that the ratio result stands on its own -- and Section 5 of
// the chapter puts the advantage back in words, where it can be weighed rather
// than smuggled into a table.
// * the word is published through the same kind of register, and `rx_overrun`
// means the same thing.
//
// WHAT IS NECESSARILY DIFFERENT, and is the whole result:
//
// * there is no second clock domain, so no synchroniser on the data path, no
// hand-off lifetime question, and no reset that has to cross anything;
// * and SCLK must be slow enough to oversample, which Chapter 15.3 measured
// exactly: below one system clock per half-period the edges are not merely
// late, they are gone or aliased.
//
// The overrun here can only be caused by the system side, never by the ratio,
// because there is no ratio to be caught out by -- which is the asymmetry the
// comparison is about. Architecture B fails by not seeing the bus; Architecture A
// fails by not being read in time.
module spi_slave_oversampled #(
parameter MAX_W = 32,
parameter LEN_W = 6,
parameter SYNC_N = 2,
parameter CAP_ON_RISING = 1'b1,
parameter LSB_FIRST = 1'b0,
parameter LEN_FIXED = 8
) (
// --- the pins. sclk_pin is DATA here, not a clock. ---------------------
input wire sclk_pin,
input wire cs_n_pin,
input wire mosi_pin,
// --- one domain -------------------------------------------------------
input wire clk,
input wire rst_n,
output reg [MAX_W-1:0] rx_data,
output reg rx_valid_stb,
output reg rx_overrun,
input wire clr_flags
);
localparam [LEN_W-1:0] LAST_BIT = LEN_FIXED - 1;
// Three synchronisers of EQUAL depth, which is Chapter 14.1's delay-matching
// requirement and is the one structural obligation Architecture B has that
// Architecture A does not: here the data is sampled by the system clock, so it
// must arrive with the same latency as the edge that says to sample it.
reg [SYNC_N-1:0] sclk_sr, cs_sr, mosi_sr;
reg sclk_d;
wire sclk_q = sclk_sr[SYNC_N-1];
wire cs_act = ~cs_sr[SYNC_N-1];
wire mosi_q = mosi_sr[SYNC_N-1];
// The capture edge, defined against the parameter rather than against rising
// and falling -- the same convention as Chapter 14.1, and here it costs a
// constant rather than a mux because the mode is fixed.
wire cap_stb = cs_act & (CAP_ON_RISING ? (sclk_q & ~sclk_d)
: (~sclk_q & sclk_d));
reg [MAX_W-1:0] sr;
reg [LEN_W-1:0] bit_idx;
reg taken; // the published word has been retired
wire [MAX_W-1:0] sr_next = {sr[MAX_W-2:0], mosi_q};
wire word_done = cap_stb & (bit_idx == LAST_BIT);
function [MAX_W-1:0] reverse_low;
input [MAX_W-1:0] v;
integer i;
begin
reverse_low = {MAX_W{1'b0}};
for (i = 0; i < LEN_FIXED; i = i + 1)
reverse_low[i] = v[LEN_FIXED - 1 - i];
end
endfunction
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
sclk_sr <= {SYNC_N{1'b0}};
cs_sr <= {SYNC_N{1'b1}}; // deselected, not selected
mosi_sr <= {SYNC_N{1'b0}};
sclk_d <= 1'b0;
sr <= {MAX_W{1'b0}};
bit_idx <= {LEN_W{1'b0}};
taken <= 1'b1;
rx_data <= {MAX_W{1'b0}};
rx_valid_stb <= 1'b0;
rx_overrun <= 1'b0;
end else begin
sclk_sr <= {sclk_sr[SYNC_N-2:0], sclk_pin};
cs_sr <= {cs_sr[SYNC_N-2:0], cs_n_pin};
mosi_sr <= {mosi_sr[SYNC_N-2:0], mosi_pin};
sclk_d <= sclk_q;
if (clr_flags)
rx_overrun <= 1'b0;
// The bit counter takes its zero from the SELECT, which is Chapter
// 14.3's rule and is the one thing both architectures share exactly.
if (!cs_act) begin
sr <= {MAX_W{1'b0}};
bit_idx <= {LEN_W{1'b0}};
end else if (cap_stb) begin
sr <= sr_next;
if (bit_idx == LAST_BIT)
bit_idx <= {LEN_W{1'b0}};
else
bit_idx <= bit_idx + 1'b1;
end
rx_valid_stb <= word_done;
if (word_done) begin
rx_data <= LSB_FIRST ? reverse_low(sr_next) : sr_next;
// An overrun here means the SYSTEM did not retire the previous
// word, never that the bus was too fast -- there is no bus-side
// way to overrun a design that samples the bus.
if (!taken)
rx_overrun <= 1'b1;
taken <= 1'b0;
end else if (rx_valid_stb) begin
taken <= 1'b1;
end
end
end
endmodule-- spi_slave_oversampled.vhd
--
-- Chapter 15.4 -- Architecture B, reduced to exactly what Architecture A does, so
-- that the two can be compared without the comparison being about anything else.
--
-- Module 14's slave is Architecture B and it is nine blocks. Putting it next to
-- Chapter 15.2's single block would compare a product with a mechanism. So this
-- file is the smallest honest Architecture B slave: recover SCLK's edges on the
-- system clock, shift MOSI in on the capture edge, publish a word every LEN_FIXED
-- captures. Same ports as Chapter 15.2's, same frame width, same reversal, same
-- reported overrun -- and one clock domain instead of two.
--
-- WHAT IS DELIBERATELY THE SAME, so the comparison is fair:
--
-- * the frame width and the capture edge are PARAMETERS here too, even though
-- Architecture B could make them run-time inputs for two gates (Chapter 14.6).
-- Leaving them fixed removes a genuine Architecture B advantage from the
-- measurement so that the ratio result stands on its own -- and Section 5 of
-- the chapter puts the advantage back in words, where it can be weighed rather
-- than smuggled into a table.
-- * the word is published through the same kind of register, and `rx_overrun`
-- means the same thing.
--
-- WHAT IS NECESSARILY DIFFERENT, and is the whole result:
--
-- * there is no second clock domain, so no synchroniser on the data path, no
-- hand-off lifetime question, and no reset that has to cross anything;
-- * and SCLK must be slow enough to oversample, which Chapter 15.3 measured
-- exactly: below one system clock per half-period the edges are not merely
-- late, they are gone or aliased.
--
-- The overrun here can only be caused by the system side, never by the ratio,
-- because there is no ratio to be caught out by -- which is the asymmetry the
-- comparison is about. Architecture B fails by not seeing the bus; Architecture A
-- fails by not being read in time.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_slave_oversampled is
generic (
MAX_W : positive := 32;
LEN_W : positive := 6;
SYNC_N : positive := 2;
CAP_ON_RISING : std_logic := '1';
LSB_FIRST : std_logic := '0';
LEN_FIXED : positive := 8
);
port (
-- the pins. sclk_pin is DATA here, not a clock.
sclk_pin : in std_logic;
cs_n_pin : in std_logic;
mosi_pin : in std_logic;
-- one domain
clk : in std_logic;
rst_n : in std_logic;
rx_data : out std_logic_vector(MAX_W - 1 downto 0);
rx_valid_stb : out std_logic;
rx_overrun : out std_logic;
clr_flags : in std_logic
);
end entity;
architecture rtl of spi_slave_oversampled is
constant LAST_BIT : unsigned(LEN_W - 1 downto 0) :=
to_unsigned(LEN_FIXED - 1, LEN_W);
-- Three synchronisers of EQUAL depth, which is Chapter 14.1's delay-matching
-- requirement and the one structural obligation Architecture B has that
-- Architecture A does not: here the data is sampled by the system clock, so it
-- must arrive with the same latency as the edge that says to sample it.
signal sclk_sr, cs_sr, mosi_sr : std_logic_vector(SYNC_N - 1 downto 0);
signal sclk_d : std_logic := '0';
signal sclk_q, cs_act, mosi_q, cap_stb : std_logic;
signal sr : std_logic_vector(MAX_W - 1 downto 0) := (others => '0');
signal sr_next : std_logic_vector(MAX_W - 1 downto 0);
signal bit_idx : unsigned(LEN_W - 1 downto 0) := (others => '0');
signal word_done : std_logic;
signal taken : std_logic := '1';
signal rx_data_r : std_logic_vector(MAX_W - 1 downto 0) := (others => '0');
signal rx_vld_r : std_logic := '0';
signal rx_ovr_r : std_logic := '0';
function reverse_low(v : std_logic_vector) return std_logic_vector is
variable r : std_logic_vector(v'range) := (others => '0');
begin
for i in 0 to LEN_FIXED - 1 loop
r(i) := v(LEN_FIXED - 1 - i);
end loop;
return r;
end function;
begin
sclk_q <= sclk_sr(SYNC_N - 1);
cs_act <= not cs_sr(SYNC_N - 1);
mosi_q <= mosi_sr(SYNC_N - 1);
-- The capture edge, defined against the generic rather than against rising and
-- falling -- the same convention as Chapter 14.1, and here it costs a constant
-- rather than a mux because the mode is fixed.
cap_stb <= cs_act and (sclk_q and not sclk_d) when CAP_ON_RISING = '1'
else cs_act and ((not sclk_q) and sclk_d);
sr_next <= sr(MAX_W - 2 downto 0) & mosi_q;
word_done <= cap_stb when bit_idx = LAST_BIT else '0';
rx_data <= rx_data_r;
rx_valid_stb <= rx_vld_r;
rx_overrun <= rx_ovr_r;
main : process (clk, rst_n)
begin
if rst_n = '0' then
sclk_sr <= (others => '0');
cs_sr <= (others => '1'); -- deselected, not selected
mosi_sr <= (others => '0');
sclk_d <= '0';
sr <= (others => '0');
bit_idx <= (others => '0');
taken <= '1';
rx_data_r <= (others => '0');
rx_vld_r <= '0';
rx_ovr_r <= '0';
elsif rising_edge(clk) then
sclk_sr <= sclk_sr(SYNC_N - 2 downto 0) & sclk_pin;
cs_sr <= cs_sr(SYNC_N - 2 downto 0) & cs_n_pin;
mosi_sr <= mosi_sr(SYNC_N - 2 downto 0) & mosi_pin;
sclk_d <= sclk_q;
if clr_flags = '1' then
rx_ovr_r <= '0';
end if;
-- The bit counter takes its zero from the SELECT, which is Chapter
-- 14.3's rule and is the one thing both architectures share exactly.
if cs_act = '0' then
sr <= (others => '0');
bit_idx <= (others => '0');
elsif cap_stb = '1' then
sr <= sr_next;
if bit_idx = LAST_BIT then
bit_idx <= (others => '0');
else
bit_idx <= bit_idx + 1;
end if;
end if;
rx_vld_r <= word_done;
if word_done = '1' then
if LSB_FIRST = '1' then
rx_data_r <= reverse_low(sr_next);
else
rx_data_r <= sr_next;
end if;
-- An overrun here means the SYSTEM did not retire the previous
-- word, never that the bus was too fast -- there is no bus-side
-- way to overrun a design that samples the bus.
if taken = '0' then
rx_ovr_r <= '1';
end if;
taken <= '0';
elsif rx_vld_r = '1' then
taken <= '1';
end if;
end if;
end process;
end architecture;The harness
Both designs instantiated on the same pins, driven by the same pin-level master, with independent collectors so that neither result can be contaminated by the other's timing.
// spi_arch_compare_tb.sv
//
// Both architectures, one stimulus, one table.
//
// The two slaves have the same ports, the same frame width, the same capture edge
// and the same meaning for `rx_overrun`. They are wired to the SAME pins and driven
// by the same pin-level master. The only difference between them is where the shift
// register's clock comes from.
//
// So every row of the table is a like-for-like comparison, and the crossover in it
// is the answer to "which architecture should I use?" -- which is a question with a
// numerical answer that depends on one ratio, and not a matter of preference.
//
// THE TABLE'S TWO REGIONS, PREDICTED BEFORE IT IS PRINTED.
//
// SCLK slow relative to the system clock
// B works: the half-period spans enough system clocks to oversample.
// A works: a word interval spans far more than SYNC_N system clocks.
// -> both work, and the choice is made on the things a table cannot show.
//
// SCLK fast relative to the system clock
// B fails: below one system clock per half-period the edges are gone or
// aliased (Chapter 15.3 measured this exactly).
// A works: until a word interval falls below SYNC_N system clocks, which is
// LEN_FIXED times further away.
// -> only A works, and the margin between the two failures is the factor the
// chapter quotes.
//
// AND THE REGION NEITHER SURVIVES, which is worth having in the table so that
// "Architecture A is faster" does not get remembered as "Architecture A has no
// limit": once a word interval falls below SYNC_N system clocks, A fails too.
`timescale 1ns/1ps
module spi_arch_compare_tb;
localparam int MAX_W = 32;
localparam int LEN_W = 6;
localparam int SYNC_N = 2;
localparam int LEN = 8;
reg clk = 1'b0;
reg rst_n = 1'b1;
always #5 clk = ~clk; // a 10 ns system clock, fixed throughout
reg sclk_pin = 1'b0;
reg cs_n_pin = 1'b1;
reg mosi_pin = 1'b0;
reg clr = 1'b0;
// --- Architecture A: the shift path clocked on SCLK ---------------------
wire [MAX_W-1:0] a_data;
wire a_vld, a_ovr;
spi_slave_sclk_domain #(.MAX_W(MAX_W), .LEN_W(LEN_W), .SYNC_N(SYNC_N),
.CAP_ON_RISING(1'b1), .LSB_FIRST(1'b0),
.LEN_FIXED(LEN)) u_a (
.sclk_pin(sclk_pin), .cs_n_pin(cs_n_pin), .mosi_pin(mosi_pin),
.clk(clk), .rst_n(rst_n),
.rx_data(a_data), .rx_valid_stb(a_vld), .rx_overrun(a_ovr),
.clr_flags(clr)
);
// --- Architecture B: everything on the system clock ---------------------
wire [MAX_W-1:0] b_data;
wire b_vld, b_ovr;
spi_slave_oversampled #(.MAX_W(MAX_W), .LEN_W(LEN_W), .SYNC_N(SYNC_N),
.CAP_ON_RISING(1'b1), .LSB_FIRST(1'b0),
.LEN_FIXED(LEN)) u_b (
.sclk_pin(sclk_pin), .cs_n_pin(cs_n_pin), .mosi_pin(mosi_pin),
.clk(clk), .rst_n(rst_n),
.rx_data(b_data), .rx_valid_stb(b_vld), .rx_overrun(b_ovr),
.clr_flags(clr)
);
integer errors = 0;
initial begin
#4_000_000;
$display("FAIL: the simulation did not finish within its time limit");
$finish;
end
// Two independent collectors, so that neither architecture's result can be
// contaminated by the other's timing.
reg [MAX_W-1:0] a_got [0:15];
reg [MAX_W-1:0] b_got [0:15];
integer a_n = 0, b_n = 0;
always @(posedge clk) if (rst_n && a_vld) begin
if (a_n < 16) a_got[a_n] = a_data;
a_n = a_n + 1;
end
always @(posedge clk) if (rst_n && b_vld) begin
if (b_n < 16) b_got[b_n] = b_data;
b_n = b_n + 1;
end
task automatic send_word(input [LEN-1:0] d, input integer half_ns);
integer i;
begin
mosi_pin = d[LEN-1];
#(half_ns);
for (i = 0; i < LEN; i = i + 1) begin
sclk_pin = 1'b1;
#(half_ns);
sclk_pin = 1'b0;
if (i < LEN-1) mosi_pin = d[LEN-2-i];
#(half_ns);
end
end
endtask
task automatic restart;
begin
cs_n_pin = 1'b1;
sclk_pin = 1'b0;
a_n = 0; b_n = 0;
rst_n = 1'b1;
repeat (2) @(posedge clk);
rst_n = 1'b0; // a real edge: see Chapter 15.2's bench
repeat (6) @(posedge clk);
rst_n = 1'b1;
repeat (6) @(posedge clk);
a_n = 0; b_n = 0;
end
endtask
integer halves [0:5];
integer h, k, a_bad, b_bad;
reg [LEN-1:0] pat [0:3];
// "works" means: four words arrived and all four were right.
function automatic [8*5-1:0] verdict(input integer n, input integer bad);
verdict = (n == 4 && bad == 0) ? " ok " : " FAIL";
endfunction
initial begin
pat[0] = 8'hA5; pat[1] = 8'h3C; pat[2] = 8'hFF; pat[3] = 8'h01;
halves[0] = 40; // SCLK period 80 ns: 8 system clocks per half-period
halves[1] = 15; // 30 ns: 1.5 per half-period
halves[2] = 10; // 20 ns: 1.0
halves[3] = 5; // 10 ns: 0.5 -- B's aliasing point
halves[4] = 2; // 4 ns: 0.2
halves[5] = 1; // 2 ns: 0.1 -- below A's word requirement too
$display(" a 10 ns system clock, SYNC_N = %0d, an %0d-bit frame, four words per run",
SYNC_N, LEN);
$display(" B needs a half-period of at least %0d system clocks; A needs a word interval of at least %0d",
SYNC_N + 1, SYNC_N);
$display(" sclk_half clocks/half word_int archA archB A_ovr B_ovr");
for (h = 0; h <= 5; h = h + 1) begin
restart();
cs_n_pin = 1'b0;
#(halves[h] * 2);
for (k = 0; k < 4; k = k + 1)
send_word(pat[k], halves[h]);
#(halves[h] * 2);
cs_n_pin = 1'b1;
#(halves[h] * 4);
repeat (40) @(posedge clk);
a_bad = 0; b_bad = 0;
for (k = 0; k < 4; k = k + 1) begin
if (k < a_n && a_got[k][LEN-1:0] !== pat[k]) a_bad = a_bad + 1;
if (k < b_n && b_got[k][LEN-1:0] !== pat[k]) b_bad = b_bad + 1;
end
if (a_n != 4) a_bad = a_bad + (4 - a_n);
if (b_n != 4) b_bad = b_bad + (4 - b_n);
$display(" %8d ns %6d.%0d %7d %5s %5s %5b %5b",
halves[h], halves[h]/10, halves[h]%10,
LEN * 2 * halves[h],
verdict(a_n, a_bad), verdict(b_n, b_bad), a_ovr, b_ovr);
// 1. WHERE BOTH PRECONDITIONS HOLD, BOTH ARCHITECTURES WORK. If this
// ever fails, the comparison is measuring a bug rather than a
// trade-off and nothing else in the table means anything.
if (halves[h] >= (SYNC_N+1)*10 &&
LEN*2*halves[h] > SYNC_N*10) begin
if (a_n != 4 || a_bad != 0) begin
$display(" FAIL: Architecture A failed where its precondition holds");
errors = errors + 1;
end
if (b_n != 4 || b_bad != 0) begin
$display(" FAIL: Architecture B failed where its precondition holds");
errors = errors + 1;
end
end
// 2. BELOW ONE SYSTEM CLOCK PER HALF-PERIOD, B MUST FAIL. This is the
// half of the table that decides the architecture, and without the
// assertion the table could be produced by two identical designs.
if (halves[h] < 10) begin
if (b_n == 4 && b_bad == 0) begin
$display(" FAIL: Architecture B delivered four correct words at a half-period of %0d ns, which is below one system clock -- Chapter 15.3 measured that the edges are gone or aliased there",
halves[h]);
errors = errors + 1;
end
end
// 3. AND WHERE ONLY A'S PRECONDITION HOLDS, A MUST STILL WORK. The
// other half of the decision.
if (halves[h] < 10 && LEN*2*halves[h] > SYNC_N*10) begin
if (a_n != 4 || a_bad != 0) begin
$display(" FAIL: Architecture A failed at a half-period of %0d ns where a word interval of %0d ns still exceeds %0d ns",
halves[h], LEN*2*halves[h], SYNC_N*10);
errors = errors + 1;
end
end
// 4. AND WHERE NEITHER HOLDS, NEITHER WORKS. Included so that
// "Architecture A is faster" cannot be remembered as "Architecture A
// has no limit".
if (LEN*2*halves[h] < SYNC_N*10) begin
if (a_n == 4 && a_bad == 0) begin
$display(" FAIL: Architecture A delivered four correct words with a word interval of %0d ns against a requirement of %0d ns",
LEN*2*halves[h], SYNC_N*10);
errors = errors + 1;
end
end
end
// 4b. THE ROWS THAT ARE GREEN AND UNSHIPPABLE. Architecture B passes at
// 1.5 and 1.0 system clocks per half-period, both of which are BELOW its
// stated rule of SYNC_N + 1. That is Chapter 15.3's result appearing in
// a comparison table, and it is the single most dangerous thing here: a
// reader who takes the table as the boundary will conclude B works down
// to one system clock per half-period, and it does not -- it works there
// in every simulation and loses edges on silicon.
//
// The bench asserts the green result, because that IS what a simulator
// does, and prints the warning beside it -- because a table cannot carry
// a caveat and a regression log can.
restart();
cs_n_pin = 1'b0; #20;
for (k = 0; k < 4; k = k + 1) send_word(pat[k], 10);
#20; cs_n_pin = 1'b1; #40;
repeat (40) @(posedge clk);
b_bad = 0;
for (k = 0; k < 4; k = k + 1)
if (k < b_n && b_got[k][LEN-1:0] !== pat[k]) b_bad = b_bad + 1;
if (b_n != 4 || b_bad != 0) begin
$display(" FAIL: Architecture B should pass IN SIMULATION at one system clock per half-period");
errors = errors + 1;
end
$display(" Architecture B delivered four correct words at 1.0 system clocks per half-period, which is below its own stated rule of %0d -- the two rows above the failure boundary are GREEN AND UNSHIPPABLE, for the reason Chapter 15.3 measured: simulation cannot lose an edge to metastability",
SYNC_N + 1);
// 4c. EACH ARCHITECTURE'S OVERRUN FLAG IS BLIND TO THE OTHER'S FAILURE, and
// the table shows it: at the fastest row Architecture A reports an
// overrun and Architecture B reports nothing at all while delivering
// wrong data. That is not a bug in B -- its overrun is about the SYSTEM
// side, and a ratio violation is invisible to it by construction. B
// needs the separate ratio measurement of Chapter 14.1, which is
// precisely why that front end publishes `min_half` and `ratio_err`.
restart();
cs_n_pin = 1'b0; #4;
for (k = 0; k < 4; k = k + 1) send_word(pat[k], 1);
#4; cs_n_pin = 1'b1; #8;
repeat (40) @(posedge clk);
if (b_ovr) begin
$display(" FAIL: Architecture B's overrun fired for a ratio violation, which it cannot observe");
errors = errors + 1;
end
$display(" at the fastest ratio Architecture A reports an overrun and Architecture B reports nothing while delivering wrong data -- each flag describes its own architecture's failure mode and is blind to the other's, which is why Architecture B needs the separate ratio measurement of Chapter 14.1 rather than an overrun");
// 5. THE CROSSOVER, STATED AS TWO NUMBERS. The smallest SCLK period each
// architecture tolerates against this 10 ns system clock, from the
// preconditions rather than from the table -- and the table above is
// consistent with both.
$display(" against a 10 ns system clock: Architecture B needs an SCLK period of at least %0d ns and Architecture A needs at least %0d ns, a ratio of %0d -- which is 2 x HALF_MIN x LEN / SYNC_N and is where the whole choice is decided",
2*(SYNC_N+1)*10, (SYNC_N*10)/LEN, (2*(SYNC_N+1)*10*LEN)/(SYNC_N*10));
// 6. THE ONE THING THE TABLE CANNOT SHOW, asserted as a structural fact
// rather than measured: Architecture B's mode and width COULD be
// run-time inputs, and Architecture A's cannot. Both are parameters here
// so that the ratio result is not contaminated, and the bench records
// that this was a deliberate handicap rather than a property.
if (LEN != 8) begin
$display(" FAIL: the comparison assumes an 8-bit frame in its arithmetic");
errors = errors + 1;
end
$display(" both designs were given a FIXED mode and frame width so the table compares only the clocking; in a real Architecture B both are run-time inputs costing two gates (Chapter 14.6), and in Architecture A they cannot be without a mux on a clock -- an advantage to B that a ratio table structurally cannot show");
if (errors == 0)
$display("PASS: the two architectures were wired to the same pins, given the same frame width, the same capture edge and the same stimulus, so the table differs in exactly one thing: where the shift register's clock comes from -- and it has three regions rather than two. Where both preconditions hold, both deliver four correct words and the choice is made on things a table cannot show. Below one system clock per SCLK half-period Architecture B fails, because its edges are gone or aliased, while Architecture A keeps working for a further factor of 2 x HALF_MIN x LEN / SYNC_N -- twenty-four at these parameters, an SCLK period of 60 ns against 2.5 ns. And below SYNC_N system clocks per WORD both fail, which is the region that stops 'A is faster' from being remembered as 'A has no limit'. So the decision is arithmetic: compute both preconditions from the two clock frequencies and the frame width, and if B's holds, prefer B for its single domain, its run-time mode and its absent hand-off; if it does not, A is not a preference but the only option -- and two warnings come with the table. B passes two rows BELOW its own stated rule, green and unshippable, because no simulator can lose an edge to metastability; and each architecture's overrun flag describes only its own failure mode, so B reports nothing at all at the fastest ratio while delivering wrong data, which is exactly why Chapter 14.1's front end publishes a measured half-period alongside its flags");
else
$display("FAIL: %0d error(s)", errors);
$finish;
end
endmodule// spi_arch_compare_tb.v
//
// Both architectures, one stimulus, one table.
//
// The two slaves have the same ports, the same frame width, the same capture edge
// and the same meaning for `rx_overrun`. They are wired to the SAME pins and driven
// by the same pin-level master. The only difference between them is where the shift
// register's clock comes from.
//
// So every row of the table is a like-for-like comparison, and the crossover in it
// is the answer to "which architecture should I use?" -- which is a question with a
// numerical answer that depends on one ratio, and not a matter of preference.
//
// THE TABLE'S TWO REGIONS, PREDICTED BEFORE IT IS PRINTED.
//
// SCLK slow relative to the system clock
// B works: the half-period spans enough system clocks to oversample.
// A works: a word interval spans far more than SYNC_N system clocks.
// -> both work, and the choice is made on the things a table cannot show.
//
// SCLK fast relative to the system clock
// B fails: below one system clock per half-period the edges are gone or
// aliased (Chapter 15.3 measured this exactly).
// A works: until a word interval falls below SYNC_N system clocks, which is
// LEN_FIXED times further away.
// -> only A works, and the margin between the two failures is the factor the
// chapter quotes.
//
// AND THE REGION NEITHER SURVIVES, which is worth having in the table so that
// "Architecture A is faster" does not get remembered as "Architecture A has no
// limit": once a word interval falls below SYNC_N system clocks, A fails too.
`timescale 1ns/1ps
module spi_arch_compare_tb;
localparam MAX_W = 32;
localparam LEN_W = 6;
localparam SYNC_N = 2;
localparam LEN = 8;
reg clk;
reg rst_n;
always #5 clk = ~clk; // a 10 ns system clock, fixed throughout
reg sclk_pin;
reg cs_n_pin;
reg mosi_pin;
reg clr;
// --- Architecture A: the shift path clocked on SCLK ---------------------
wire [MAX_W-1:0] a_data;
wire a_vld, a_ovr;
spi_slave_sclk_domain #(.MAX_W(MAX_W), .LEN_W(LEN_W), .SYNC_N(SYNC_N),
.CAP_ON_RISING(1'b1), .LSB_FIRST(1'b0),
.LEN_FIXED(LEN)) u_a (
.sclk_pin(sclk_pin), .cs_n_pin(cs_n_pin), .mosi_pin(mosi_pin),
.clk(clk), .rst_n(rst_n),
.rx_data(a_data), .rx_valid_stb(a_vld), .rx_overrun(a_ovr),
.clr_flags(clr)
);
// --- Architecture B: everything on the system clock ---------------------
wire [MAX_W-1:0] b_data;
wire b_vld, b_ovr;
spi_slave_oversampled #(.MAX_W(MAX_W), .LEN_W(LEN_W), .SYNC_N(SYNC_N),
.CAP_ON_RISING(1'b1), .LSB_FIRST(1'b0),
.LEN_FIXED(LEN)) u_b (
.sclk_pin(sclk_pin), .cs_n_pin(cs_n_pin), .mosi_pin(mosi_pin),
.clk(clk), .rst_n(rst_n),
.rx_data(b_data), .rx_valid_stb(b_vld), .rx_overrun(b_ovr),
.clr_flags(clr)
);
integer errors;
initial begin
#4_000_000;
$display("FAIL: the simulation did not finish within its time limit");
$finish;
end
// Two independent collectors, so that neither architecture's result can be
// contaminated by the other's timing.
reg [MAX_W-1:0] a_got [0:15];
reg [MAX_W-1:0] b_got [0:15];
integer a_n, b_n;
always @(posedge clk) if (rst_n && a_vld) begin
if (a_n < 16) a_got[a_n] = a_data;
a_n = a_n + 1;
end
always @(posedge clk) if (rst_n && b_vld) begin
if (b_n < 16) b_got[b_n] = b_data;
b_n = b_n + 1;
end
task send_word;
input [LEN-1:0] d;
input integer half_ns;
integer i;
begin
mosi_pin = d[LEN-1];
#(half_ns);
for (i = 0; i < LEN; i = i + 1) begin
sclk_pin = 1'b1;
#(half_ns);
sclk_pin = 1'b0;
if (i < LEN-1) mosi_pin = d[LEN-2-i];
#(half_ns);
end
end
endtask
task restart;
begin
cs_n_pin = 1'b1;
sclk_pin = 1'b0;
a_n = 0; b_n = 0;
rst_n = 1'b1;
repeat (2) @(posedge clk);
rst_n = 1'b0; // a real edge: see Chapter 15.2's bench
repeat (6) @(posedge clk);
rst_n = 1'b1;
repeat (6) @(posedge clk);
a_n = 0; b_n = 0;
end
endtask
integer halves [0:5];
integer h, k, a_bad, b_bad;
reg [LEN-1:0] pat [0:3];
// "works" means: four words arrived and all four were right.
function [8*5-1:0] verdict;
input integer n;
input integer bad;
verdict = (n == 4 && bad == 0) ? " ok " : " FAIL";
endfunction
initial begin
pat[0] = 8'hA5; pat[1] = 8'h3C; pat[2] = 8'hFF; pat[3] = 8'h01;
halves[0] = 40; // SCLK period 80 ns: 8 system clocks per half-period
halves[1] = 15; // 30 ns: 1.5 per half-period
halves[2] = 10; // 20 ns: 1.0
halves[3] = 5; // 10 ns: 0.5 -- B's aliasing point
halves[4] = 2; // 4 ns: 0.2
halves[5] = 1; // 2 ns: 0.1 -- below A's word requirement too
$display(" a 10 ns system clock, SYNC_N = %0d, an %0d-bit frame, four words per run",
SYNC_N, LEN);
$display(" B needs a half-period of at least %0d system clocks; A needs a word interval of at least %0d",
SYNC_N + 1, SYNC_N);
$display(" sclk_half clocks/half word_int archA archB A_ovr B_ovr");
for (h = 0; h <= 5; h = h + 1) begin
restart();
cs_n_pin = 1'b0;
#(halves[h] * 2);
for (k = 0; k < 4; k = k + 1)
send_word(pat[k], halves[h]);
#(halves[h] * 2);
cs_n_pin = 1'b1;
#(halves[h] * 4);
repeat (40) @(posedge clk);
a_bad = 0; b_bad = 0;
for (k = 0; k < 4; k = k + 1) begin
if (k < a_n && a_got[k][LEN-1:0] !== pat[k]) a_bad = a_bad + 1;
if (k < b_n && b_got[k][LEN-1:0] !== pat[k]) b_bad = b_bad + 1;
end
if (a_n != 4) a_bad = a_bad + (4 - a_n);
if (b_n != 4) b_bad = b_bad + (4 - b_n);
$display(" %8d ns %6d.%0d %7d %0s %0s %5b %5b",
halves[h], halves[h]/10, halves[h]%10,
LEN * 2 * halves[h],
verdict(a_n, a_bad), verdict(b_n, b_bad), a_ovr, b_ovr);
// 1. WHERE BOTH PRECONDITIONS HOLD, BOTH ARCHITECTURES WORK. If this
// ever fails, the comparison is measuring a bug rather than a
// trade-off and nothing else in the table means anything.
if (halves[h] >= (SYNC_N+1)*10 &&
LEN*2*halves[h] > SYNC_N*10) begin
if (a_n != 4 || a_bad != 0) begin
$display(" FAIL: Architecture A failed where its precondition holds");
errors = errors + 1;
end
if (b_n != 4 || b_bad != 0) begin
$display(" FAIL: Architecture B failed where its precondition holds");
errors = errors + 1;
end
end
// 2. BELOW ONE SYSTEM CLOCK PER HALF-PERIOD, B MUST FAIL. This is the
// half of the table that decides the architecture, and without the
// assertion the table could be produced by two identical designs.
if (halves[h] < 10) begin
if (b_n == 4 && b_bad == 0) begin
$display(" FAIL: Architecture B delivered four correct words at a half-period of %0d ns, which is below one system clock -- Chapter 15.3 measured that the edges are gone or aliased there",
halves[h]);
errors = errors + 1;
end
end
// 3. AND WHERE ONLY A'S PRECONDITION HOLDS, A MUST STILL WORK. The
// other half of the decision.
if (halves[h] < 10 && LEN*2*halves[h] > SYNC_N*10) begin
if (a_n != 4 || a_bad != 0) begin
$display(" FAIL: Architecture A failed at a half-period of %0d ns where a word interval of %0d ns still exceeds %0d ns",
halves[h], LEN*2*halves[h], SYNC_N*10);
errors = errors + 1;
end
end
// 4. AND WHERE NEITHER HOLDS, NEITHER WORKS. Included so that
// "Architecture A is faster" cannot be remembered as "Architecture A
// has no limit".
if (LEN*2*halves[h] < SYNC_N*10) begin
if (a_n == 4 && a_bad == 0) begin
$display(" FAIL: Architecture A delivered four correct words with a word interval of %0d ns against a requirement of %0d ns",
LEN*2*halves[h], SYNC_N*10);
errors = errors + 1;
end
end
end
// 4b. THE ROWS THAT ARE GREEN AND UNSHIPPABLE. Architecture B passes at
// 1.5 and 1.0 system clocks per half-period, both of which are BELOW its
// stated rule of SYNC_N + 1. That is Chapter 15.3's result appearing in
// a comparison table, and it is the single most dangerous thing here: a
// reader who takes the table as the boundary will conclude B works down
// to one system clock per half-period, and it does not -- it works there
// in every simulation and loses edges on silicon.
//
// The bench asserts the green result, because that IS what a simulator
// does, and prints the warning beside it -- because a table cannot carry
// a caveat and a regression log can.
restart();
cs_n_pin = 1'b0; #20;
for (k = 0; k < 4; k = k + 1) send_word(pat[k], 10);
#20; cs_n_pin = 1'b1; #40;
repeat (40) @(posedge clk);
b_bad = 0;
for (k = 0; k < 4; k = k + 1)
if (k < b_n && b_got[k][LEN-1:0] !== pat[k]) b_bad = b_bad + 1;
if (b_n != 4 || b_bad != 0) begin
$display(" FAIL: Architecture B should pass IN SIMULATION at one system clock per half-period");
errors = errors + 1;
end
$display(" Architecture B delivered four correct words at 1.0 system clocks per half-period, which is below its own stated rule of %0d -- the two rows above the failure boundary are GREEN AND UNSHIPPABLE, for the reason Chapter 15.3 measured: simulation cannot lose an edge to metastability",
SYNC_N + 1);
// 4c. EACH ARCHITECTURE'S OVERRUN FLAG IS BLIND TO THE OTHER'S FAILURE, and
// the table shows it: at the fastest row Architecture A reports an
// overrun and Architecture B reports nothing at all while delivering
// wrong data. That is not a bug in B -- its overrun is about the SYSTEM
// side, and a ratio violation is invisible to it by construction. B
// needs the separate ratio measurement of Chapter 14.1, which is
// precisely why that front end publishes `min_half` and `ratio_err`.
restart();
cs_n_pin = 1'b0; #4;
for (k = 0; k < 4; k = k + 1) send_word(pat[k], 1);
#4; cs_n_pin = 1'b1; #8;
repeat (40) @(posedge clk);
if (b_ovr) begin
$display(" FAIL: Architecture B's overrun fired for a ratio violation, which it cannot observe");
errors = errors + 1;
end
$display(" at the fastest ratio Architecture A reports an overrun and Architecture B reports nothing while delivering wrong data -- each flag describes its own architecture's failure mode and is blind to the other's, which is why Architecture B needs the separate ratio measurement of Chapter 14.1 rather than an overrun");
// 5. THE CROSSOVER, STATED AS TWO NUMBERS. The smallest SCLK period each
// architecture tolerates against this 10 ns system clock, from the
// preconditions rather than from the table -- and the table above is
// consistent with both.
$display(" against a 10 ns system clock: Architecture B needs an SCLK period of at least %0d ns and Architecture A needs at least %0d ns, a ratio of %0d -- which is 2 x HALF_MIN x LEN / SYNC_N and is where the whole choice is decided",
2*(SYNC_N+1)*10, (SYNC_N*10)/LEN, (2*(SYNC_N+1)*10*LEN)/(SYNC_N*10));
// 6. THE ONE THING THE TABLE CANNOT SHOW, asserted as a structural fact
// rather than measured: Architecture B's mode and width COULD be
// run-time inputs, and Architecture A's cannot. Both are parameters here
// so that the ratio result is not contaminated, and the bench records
// that this was a deliberate handicap rather than a property.
if (LEN != 8) begin
$display(" FAIL: the comparison assumes an 8-bit frame in its arithmetic");
errors = errors + 1;
end
$display(" both designs were given a FIXED mode and frame width so the table compares only the clocking; in a real Architecture B both are run-time inputs costing two gates (Chapter 14.6), and in Architecture A they cannot be without a mux on a clock -- an advantage to B that a ratio table structurally cannot show");
if (errors == 0)
$display("PASS: the two architectures were wired to the same pins, given the same frame width, the same capture edge and the same stimulus, so the table differs in exactly one thing: where the shift register's clock comes from -- and it has three regions rather than two. Where both preconditions hold, both deliver four correct words and the choice is made on things a table cannot show. Below one system clock per SCLK half-period Architecture B fails, because its edges are gone or aliased, while Architecture A keeps working for a further factor of 2 x HALF_MIN x LEN / SYNC_N -- twenty-four at these parameters, an SCLK period of 60 ns against 2.5 ns. And below SYNC_N system clocks per WORD both fail, which is the region that stops 'A is faster' from being remembered as 'A has no limit'. So the decision is arithmetic: compute both preconditions from the two clock frequencies and the frame width, and if B's holds, prefer B for its single domain, its run-time mode and its absent hand-off; if it does not, A is not a preference but the only option -- and two warnings come with the table. B passes two rows BELOW its own stated rule, green and unshippable, because no simulator can lose an edge to metastability; and each architecture's overrun flag describes only its own failure mode, so B reports nothing at all at the fastest ratio while delivering wrong data, which is exactly why Chapter 14.1's front end publishes a measured half-period alongside its flags");
else
$display("FAIL: %0d error(s)", errors);
$finish;
end
initial begin
a_n = 0;
b_n = 0;
clk = 1'b0;
rst_n = 1'b1;
sclk_pin = 1'b0;
cs_n_pin = 1'b1;
mosi_pin = 1'b0;
clr = 1'b0;
errors = 0;
end
endmodule-- spi_arch_compare_tb.vhd
--
-- Both architectures, one stimulus, one table.
--
-- The two slaves have the same ports, the same frame width, the same capture edge
-- and the same meaning for `rx_overrun`. They are wired to the SAME pins and driven
-- by the same pin-level master. The only difference between them is where the shift
-- register's clock comes from.
--
-- So every row of the table is a like-for-like comparison, and the crossover in it
-- is the answer to "which architecture should I use?" -- which is a question with a
-- numerical answer that depends on one ratio, and not a matter of preference.
--
-- THE TABLE'S TWO REGIONS, PREDICTED BEFORE IT IS PRINTED.
--
-- SCLK slow relative to the system clock
-- B works: the half-period spans enough system clocks to oversample.
-- A works: a word interval spans far more than SYNC_N system clocks.
-- -> both work, and the choice is made on the things a table cannot show.
--
-- SCLK fast relative to the system clock
-- B fails: below one system clock per half-period the edges are gone or
-- aliased (Chapter 15.3 measured this exactly).
-- A works: until a word interval falls below SYNC_N system clocks, which is
-- LEN_FIXED times further away.
-- -> only A works, and the margin between the two failures is the factor the
-- chapter quotes.
--
-- AND THE REGION NEITHER SURVIVES, which is worth having in the table so that
-- "Architecture A is faster" does not get remembered as "Architecture A has no
-- limit": once a word interval falls below SYNC_N system clocks, A fails too.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_arch_compare_tb is
end entity;
architecture sim of spi_arch_compare_tb is
constant MAX_W : positive := 32;
constant LEN_W : positive := 6;
constant SYNC_N : positive := 2;
constant LEN : positive := 8;
signal clk : std_logic := '0';
signal rst_n : std_logic := '1';
signal halt : boolean := false;
signal sclk_pin : std_logic := '0';
signal cs_n_pin : std_logic := '1';
signal mosi_pin : std_logic := '0';
signal clr : std_logic := '0';
signal a_data, b_data : std_logic_vector(MAX_W - 1 downto 0);
signal a_vld, a_ovr, b_vld, b_ovr : std_logic;
-- Two independent collectors, so that neither architecture's result can be
-- contaminated by the other's timing. Each owns its own array and counter, and
-- the stimulus asks for a clear through a strobe rather than writing them --
-- two drivers on a signal resolve to 'X' in VHDL.
type word_arr is array (0 to 15) of std_logic_vector(MAX_W - 1 downto 0);
signal a_got, b_got : word_arr := (others => (others => '0'));
signal a_n, b_n : natural := 0;
signal got_clr : std_logic := '0';
begin
sysclk : process
begin
while not halt loop
clk <= '0'; wait for 5 ns; -- a 10 ns system clock, fixed
clk <= '1'; wait for 5 ns;
end loop;
wait;
end process;
-- Architecture A: the shift path clocked on SCLK
u_a : entity work.spi_slave_sclk_domain
generic map (MAX_W => MAX_W, LEN_W => LEN_W, SYNC_N => SYNC_N,
CAP_ON_RISING => '1', LSB_FIRST => '0', LEN_FIXED => LEN)
port map (sclk_pin => sclk_pin, cs_n_pin => cs_n_pin, mosi_pin => mosi_pin,
clk => clk, rst_n => rst_n,
rx_data => a_data, rx_valid_stb => a_vld, rx_overrun => a_ovr,
clr_flags => clr);
-- Architecture B: everything on the system clock
u_b : entity work.spi_slave_oversampled
generic map (MAX_W => MAX_W, LEN_W => LEN_W, SYNC_N => SYNC_N,
CAP_ON_RISING => '1', LSB_FIRST => '0', LEN_FIXED => LEN)
port map (sclk_pin => sclk_pin, cs_n_pin => cs_n_pin, mosi_pin => mosi_pin,
clk => clk, rst_n => rst_n,
rx_data => b_data, rx_valid_stb => b_vld, rx_overrun => b_ovr,
clr_flags => clr);
collect_a : process (clk)
begin
if rising_edge(clk) then
if got_clr = '1' then
a_n <= 0;
elsif rst_n = '1' and a_vld = '1' then
if a_n < 16 then a_got(a_n) <= a_data; end if;
a_n <= a_n + 1;
end if;
end if;
end process;
collect_b : process (clk)
begin
if rising_edge(clk) then
if got_clr = '1' then
b_n <= 0;
elsif rst_n = '1' and b_vld = '1' then
if b_n < 16 then b_got(b_n) <= b_data; end if;
b_n <= b_n + 1;
end if;
end if;
end process;
watchdog : process
begin
wait for 4 ms;
if not halt then
report "FAIL: the simulation did not finish within its time limit"
severity failure;
end if;
wait;
end process;
stim : process
variable errs : natural := 0;
variable a_bad, b_bad : natural;
type half_arr is array (0 to 5) of time;
constant HALVES : half_arr := (40 ns, 15 ns, 10 ns, 5 ns, 2 ns, 1 ns);
type pat_arr is array (0 to 3) of std_logic_vector(LEN - 1 downto 0);
constant PAT : pat_arr := (x"A5", x"3C", x"FF", x"01");
-- "works" means: four words arrived and all four were right.
function verdict(n : natural; bad : natural) return string is
begin
if n = 4 and bad = 0 then return " ok "; else return " FAIL"; end if;
end function;
procedure send_word(d : std_logic_vector(LEN - 1 downto 0); half : time) is
begin
mosi_pin <= d(LEN - 1);
wait for half;
for i in 0 to LEN - 1 loop
sclk_pin <= '1';
wait for half;
sclk_pin <= '0';
if i < LEN - 1 then mosi_pin <= d(LEN - 2 - i); end if;
wait for half;
end loop;
end procedure;
procedure restart is
begin
cs_n_pin <= '1';
sclk_pin <= '0';
got_clr <= '1';
rst_n <= '1';
for i in 1 to 2 loop wait until rising_edge(clk); end loop;
got_clr <= '0';
rst_n <= '0'; -- a real edge: see Chapter 15.2's bench
for i in 1 to 6 loop wait until rising_edge(clk); end loop;
rst_n <= '1';
for i in 1 to 6 loop wait until rising_edge(clk); end loop;
got_clr <= '1';
wait until rising_edge(clk);
got_clr <= '0';
wait until rising_edge(clk);
end procedure;
procedure run_words(half : time) is
begin
cs_n_pin <= '0';
wait for 2 * half;
for k in 0 to 3 loop send_word(PAT(k), half); end loop;
wait for 2 * half;
cs_n_pin <= '1';
wait for 4 * half;
for i in 1 to 40 loop wait until rising_edge(clk); end loop;
end procedure;
procedure score(variable ab : out natural; variable bb : out natural) is
variable x, y : natural := 0;
begin
for k in 0 to 3 loop
if k < a_n and a_got(k)(LEN-1 downto 0) /= PAT(k) then x := x + 1; end if;
if k < b_n and b_got(k)(LEN-1 downto 0) /= PAT(k) then y := y + 1; end if;
end loop;
if a_n /= 4 then x := x + (4 - a_n); end if;
if b_n /= 4 then y := y + (4 - b_n); end if;
ab := x; bb := y;
end procedure;
begin
report " a 10 ns system clock, SYNC_N = " & integer'image(SYNC_N) &
", an " & integer'image(LEN) & "-bit frame, four words per run";
report " B needs a half-period of at least " & integer'image(SYNC_N + 1) &
" system clocks; A needs a word interval of at least " &
integer'image(SYNC_N);
report " sclk_half clocks/half word_int archA archB A_ovr B_ovr";
for h in HALVES'range loop
restart;
run_words(HALVES(h));
score(a_bad, b_bad);
report " " & integer'image(HALVES(h) / 1 ns) & " ns " &
integer'image((HALVES(h) / 1 ns) / 10) & "." &
integer'image((HALVES(h) / 1 ns) mod 10) & " " &
integer'image(LEN * 2 * (HALVES(h) / 1 ns)) & " " &
verdict(a_n, a_bad) & " " & verdict(b_n, b_bad) & " " &
std_logic'image(a_ovr) & " " & std_logic'image(b_ovr);
-- 1. WHERE BOTH PRECONDITIONS HOLD, BOTH ARCHITECTURES WORK.
if HALVES(h) >= (SYNC_N + 1) * 10 ns
and LEN * 2 * HALVES(h) > SYNC_N * 10 ns then
if a_n /= 4 or a_bad /= 0 then
report " FAIL: Architecture A failed where its precondition holds";
errs := errs + 1;
end if;
if b_n /= 4 or b_bad /= 0 then
report " FAIL: Architecture B failed where its precondition holds";
errs := errs + 1;
end if;
end if;
-- 2. BELOW ONE SYSTEM CLOCK PER HALF-PERIOD, B MUST FAIL.
if HALVES(h) < 10 ns and a_n = 4 and b_bad = 0 and b_n = 4 then
report " FAIL: Architecture B delivered four correct words below one system clock per half-period";
errs := errs + 1;
end if;
-- 3. AND WHERE ONLY A'S PRECONDITION HOLDS, A MUST STILL WORK.
if HALVES(h) < 10 ns and LEN * 2 * HALVES(h) > SYNC_N * 10 ns then
if a_n /= 4 or a_bad /= 0 then
report " FAIL: Architecture A failed where a word interval still exceeds the requirement";
errs := errs + 1;
end if;
end if;
-- 4. AND WHERE NEITHER HOLDS, NEITHER WORKS -- so that "A is faster"
-- cannot be remembered as "A has no limit".
if LEN * 2 * HALVES(h) < SYNC_N * 10 ns then
if a_n = 4 and a_bad = 0 then
report " FAIL: Architecture A delivered four correct words below its word requirement";
errs := errs + 1;
end if;
end if;
end loop;
-- 4b. THE ROWS THAT ARE GREEN AND UNSHIPPABLE.
restart;
run_words(10 ns);
score(a_bad, b_bad);
if b_n /= 4 or b_bad /= 0 then
report " FAIL: Architecture B should pass IN SIMULATION at one system clock per half-period";
errs := errs + 1;
end if;
report " Architecture B delivered four correct words at 1.0 system clocks per half-period, which is below its own stated rule of " &
integer'image(SYNC_N + 1) &
" -- the two rows above the failure boundary are GREEN AND UNSHIPPABLE, for the reason Chapter 15.3 measured: simulation cannot lose an edge to metastability";
-- 4c. EACH ARCHITECTURE'S OVERRUN FLAG IS BLIND TO THE OTHER'S FAILURE.
restart;
run_words(1 ns);
if b_ovr = '1' then
report " FAIL: Architecture B's overrun fired for a ratio violation, which it cannot observe";
errs := errs + 1;
end if;
report " at the fastest ratio Architecture A reports an overrun and Architecture B reports nothing while delivering wrong data -- each flag describes its own architecture's failure mode and is blind to the other's, which is why Architecture B needs the separate ratio measurement of Chapter 14.1 rather than an overrun";
-- 5. THE CROSSOVER, STATED AS TWO NUMBERS.
report " against a 10 ns system clock: Architecture B needs an SCLK period of at least " &
integer'image(2 * (SYNC_N + 1) * 10) &
" ns and Architecture A needs at least " &
integer'image((SYNC_N * 10) / LEN) & " ns, a ratio of " &
integer'image((2 * (SYNC_N + 1) * 10 * LEN) / (SYNC_N * 10)) &
" -- which is 2 x HALF_MIN x LEN / SYNC_N and is where the whole choice is decided";
report " both designs were given a FIXED mode and frame width so the table compares only the clocking; in a real Architecture B both are run-time inputs costing two gates (Chapter 14.6), and in Architecture A they cannot be without a mux on a clock -- an advantage to B that a ratio table structurally cannot show";
if errs = 0 then
report "PASS: the two architectures were wired to the same pins, given the same frame width, the same capture edge and the same stimulus, so the table differs in exactly one thing: where the shift register's clock comes from -- and it has three regions rather than two. Where both preconditions hold, both deliver four correct words and the choice is made on things a table cannot show. Below one system clock per SCLK half-period Architecture B fails, because its edges are gone or aliased, while Architecture A keeps working for a further factor of 2 x HALF_MIN x LEN / SYNC_N -- twenty-four at these parameters, an SCLK period of 60 ns against 2.5 ns. And below SYNC_N system clocks per WORD both fail, which is the region that stops 'A is faster' from being remembered as 'A has no limit'. So the decision is arithmetic: compute both preconditions from the two clock frequencies and the frame width, and if B's holds, prefer B for its single domain, its run-time mode and its absent hand-off; if it does not, A is not a preference but the only option -- and two warnings come with the table. B passes two rows BELOW its own stated rule, green and unshippable, because no simulator can lose an edge to metastability; and each architecture's overrun flag describes only its own failure mode, so B reports nothing at all at the fastest ratio while delivering wrong data, which is exactly why Chapter 14.1's front end publishes a measured half-period alongside its flags";
else
report "FAIL: " & integer'image(errs) & " error(s)" severity error;
end if;
halt <= true;
wait;
end process;
end architecture;7. Why a Verification Engineer Cares
A comparison bench needs one stimulus and two collectors. Sharing the stimulus is what makes the difference attributable; separating the collectors is what stops one design's timing perturbing the other's record. Both halves are easy to get wrong in the direction that produces a plausible table.
Derive expectations from the specification, not from a previous run. This is the discipline that makes a comparison table trustworthy: every assertion in the harness is an inequality in the two preconditions, so the table cannot drift into recording whatever the designs happen to do.
Assert the negative regions. "A fails here too" and "B passes below its own rule" are both assertions about things going wrong, and both protect the reader. A suite that only checks the rows where everything works produces a table whose boundaries are decoration.
And put the unshippable-but-green rows in the log. §4's first callout is a bench $display, deliberately — the one place a caveat travels with the numbers.
// Properties for the comparison. All of them are about the RELATIONSHIP between the
// two designs, which is why none of them can live inside either one.
property p_agree_where_both_preconditions_hold;
// Region 1: the two architectures must deliver the same words. If they ever
// disagree here, the comparison is measuring a bug rather than a trade-off.
@(posedge clk) disable iff (!rst_n)
(a_vld && both_preconditions_hold) |-> (a_data == b_data_expected);
endproperty
property p_b_fails_below_the_hard_limit;
// Region 2, as an assertion about a FAILURE -- which is unusual and necessary,
// because without it the table could be produced by two identical designs.
@(posedge clk) disable iff (!rst_n)
(end_of_run && sclk_half_below_one_dst_cycle) |-> (b_words != 4 || b_bad != 0);
endproperty
property p_a_fails_below_its_word_limit;
// Region 3. This is the assertion that stops "A is faster" from being
// remembered as "A has no limit".
@(posedge clk) disable iff (!rst_n)
(end_of_run && word_interval_below_sync_n) |-> (a_words != 4 || a_bad != 0);
endproperty
property p_b_overrun_never_reports_a_ratio_fault;
// Section 4's second callout: B's overrun is about the system side and cannot
// observe a ratio violation. Asserting that it stays CLEAR is what documents the
// gap, rather than leaving a reader to infer it from a zero in a table.
@(posedge clk) disable iff (!rst_n)
sclk_half_below_one_dst_cycle |-> !b_overrun;
endproperty// Coverage. The axes are the two preconditions, and what matters is the CROSS --
// because the table's three regions are three cells of it and a suite that visits
// only two has not made the comparison.
covergroup cg_choice @(posedge clk iff end_of_run);
option.per_instance = 1;
// Architecture B's precondition: SCLK half-period in system clocks, tenths.
b_pre: coverpoint sclk_half_tenths {
bins violated = {[1:9]};
bins green_bad = {[10:29]}; // passes in simulation, below the rule
bins met = {[30:$]};
}
// Architecture A's precondition: word interval in destination cycles.
a_pre: coverpoint word_interval_dst_cycles {
bins violated = {[0:1]};
bins at = {2};
bins met = {[3:$]};
}
// The three regions are cells of this cross. `b_pre.violated & a_pre.met` is
// region 2 -- the only reason to choose Architecture A -- and a suite that never
// reaches it has not tested the decision.
x_regions: cross b_pre, a_pre;
endgroup8. Why an FPGA or ASIC Engineer Cares
The two architectures make opposite demands of the same pin. Architecture B needs SCLK on any I/O pin with a set_input_delay. Architecture A needs it on a clock-capable pin reaching a global clock buffer, with a create_clock. That is a pinout constraint arriving from an RTL decision, and it is discovered late if the choice is made late — a board already routed to a general-purpose pin cannot be given Architecture A without a respin.
Architecture A's clock is not free-running, and most tools assume clocks are. Clock-gating checks, CDC reports and anything that reasons about "cycles" must cope with a clock that stops for milliseconds. In practice the SCLK domain is constrained against the SCLK period and every path leaving it is declared asynchronous.
Architecture B's synchroniser depth is load-bearing in a way that has a datasheet consequence. Deepening the chain raises HALF_MIN and lowers the maximum supported SPI frequency, so a change made for metastability margin changes a published number (Chapter 15.3 §7).
Architecture A's hand-off has no meaningful timing constraint, so it gets a CDC exception and its correctness lives in an inequality rather than in a report. B has no such path at all — which is the strongest implementation argument for B and the reason to prefer it whenever its precondition holds.
9. Failure Signature — A Design That Works Until The Clock Tree Is Rebalanced
The symptom:
"The slave worked for two years. A timing-driven rebuild for an unrelated block changed the placement and now the SPI interface fails intermittently at the same SPI frequency it always used."
What is happening: an Architecture B design whose SCLK half-period was between one and SYNC_N + 1 system clocks all along — one of §4's green-but-unshippable rows. It worked because the particular placement happened to give enough margin, and the rebuild took it away. Nothing about the SPI interface changed.
Why the two-year history makes it harder rather than easier: everyone's first hypothesis is the change, and the change is innocent. The design has been outside its stated precondition since it shipped, and simulation has been green throughout.
How to find it in one read: compute SCLK half-period ÷ system clock period and compare against SYNC_N + 1. If it is between 1 and SYNC_N + 1, that is the fault, and no amount of placement work will fix it — the options are a slower SPI clock, a faster system clock, or Architecture A.
10. Common Misconceptions
"Architecture A is faster, so it is better." It is faster by a factor of 2 × HALF_MIN × LEN / SYNC_N and it costs a second clock domain, a synthesis-time mode, a hand-off lifetime requirement, a reset crossing and a clock-capable pin. Where B's precondition holds, B is the better design.
"Architecture A has no ratio requirement." It has a per-word one, and region 3 of the table is it failing.
"The table's lowest green row for B is B's limit." It is two rows below B's stated rule. This is §4's first warning and the single most quotable wrong number in the module.
"If a design's overrun flag is clear, the crossing is healthy." Each flag describes its own architecture's failure mode. B's overrun is blind to a ratio violation by construction, and at the fastest row it reports nothing while delivering wrong data.
"The choice is a matter of house style." It is an inequality in two clock frequencies and a frame width. Region 2 of the table is the region where there is no choice at all.
"Fixing B's mode to make the comparison fair understates B." It does, deliberately and by exactly one known amount — §5 restores it in words. The alternative, letting B's run-time mode into a ratio table, would mix an advantage that is not about frequency into a number that is.
11. Reason It Through
Q. A 50 MHz system clock and a 25 MHz SPI interface, SYNC_N = 2, 8-bit frames. Which architectures are available?
The system period is 20 ns and the SCLK half-period is 20 ns — exactly one system clock. Architecture B's rule needs 3, so B is unavailable: this is precisely a §4 green-but-unshippable row. Architecture A needs 2 × 20 = 40 ns to fit inside 8 × 40 = 320 ns, which it does comfortably. So A is the only option, and a design that shipped B here would be in §9's failure signature from day one.
Q. Why does a one-bit frame nearly remove Architecture A's advantage?
Because the relief is 2 × HALF_MIN × LEN / SYNC_N and LEN is a direct factor. At LEN = 1, HALF_MIN = 3 and SYNC_N = 2 the factor is 3 rather than 24 — so A tolerates an SCLK three times faster rather than twenty-four times. Since A costs a clock domain, a hand-off, a reset crossing and a clock-capable pin, a threefold relief is a much weaker case, and a bit-oriented protocol is a poor fit for the architecture.
Q. The harness fixes both designs' mode to a parameter. Name a measurement that decision makes impossible, and say why it is worth losing.
It makes it impossible to measure the cost of B's run-time mode selection — the two gates of Chapter 14.6 — against A's inability to have one at all. That is worth losing because the cost is not a frequency cost: putting it into a ratio table would make the table's numbers depend on a feature difference, so a reader could no longer attribute any row to the clocking. The advantage is real, it is stated in §5's table, and it is weighed there rather than encoded in a number that looks like a measurement.
Q. Both architectures deliver four correct words in region 1. On what basis should a team choose?
On everything in §5's table: one clock domain against two, a run-time mode against a build per mode, no reset crossing against one, an ordinary I/O pin against a clock-capable one, and a fault that reports itself against one that needs a separate measurement published. The first four favour B and the last favours A. In practice B wins region 1 decisively, because a design with no crossing at all has no Chapter 15.5 and no Chapter 15.6 to get wrong.
Q. Why is it necessary to assert that Architecture B PASSES at a ratio of 1.0, rather than simply omitting that row?
Because omitting it leaves the sweep's lowest green row somewhere else, and a reader then treats that as the boundary — the problem does not go away, it moves. Asserting the pass and printing the warning keeps the row, its result and its caveat together in one place. It also protects against a future contributor "fixing" the apparent inconsistency by making the design fail there, which would be making the RTL model something a simulator cannot model.
12. Understanding Check
13. Summary
Both architectures were wired to the same pins with the same frame width, capture edge and stimulus, so the table differs in exactly one thing: where the shift register's clock comes from.
It has three regions. Both work where both preconditions hold. Only A works below one system clock per SCLK half-period, where B's edges are gone or aliased — including at a 4 ns SCLK period against a 10 ns system clock, which is SCLK two and a half times faster. And neither works once a word arrives in fewer than SYNC_N destination cycles, which is the region that stops "A is faster" from being remembered as "A has no limit".
The crossover is 2 × HALF_MIN × LEN / SYNC_N — twenty-four here, an SCLK period of 60 ns against 2.5 ns — and most of it comes from the frame width, so a bit-oriented protocol is a poor case for Architecture A.
Two warnings the table cannot carry. B passes two rows below its own rule, green in every simulation and losing edges on silicon — the most quotable wrong number in the module. And each flag is blind to the other architecture's failure: at the fastest row A reports an overrun and B reports nothing while delivering wrong data, because B's overrun is about the system side by construction.
So the decision is arithmetic: compute both preconditions from the two clock frequencies and the frame width. If B's holds, prefer B — one domain, a run-time mode, no hand-off, no reset crossing, an ordinary I/O pin. If it does not, A is not a preference but the only option.
For verification: one stimulus, two collectors; expectations derived from the preconditions rather than tabulated; the negative regions asserted, because a table whose boundaries are unasserted has decoration where it should have measurements; and the unshippable-green rows printed in the log, which is the only place a caveat travels with a number.
For implementation: the two architectures make opposite demands of the same pin — any I/O with an input delay, against a clock-capable pin with a generated clock — which is a pinout constraint arriving from an RTL decision and a respin if the choice is made late.
14. What Comes Next
The architecture is chosen. If it is Architecture A, or Architecture B with a system-side buffer, there is a crossing to build — and the obvious way to build one is wrong.
Chapter 15.5 — Synchronizers and Their Limits measures what a two-flop synchroniser actually guarantees, which is less than it is usually credited with: not which of two values you get, not a pulse, and emphatically not a bus. Driven with the per-bit skew ordinary routing produces, N independent synchronisers on an 8-bit counter observe values the source never held — repeatedly, in many distinct forms, every one a well-formed pattern no downstream check could reject, and still wrong when the source is slowed, which is exactly how the crossing passes a regression and reaches silicon.
Continue learning
Related tutorials
- Related topic
Architecture A — SCLK as a Clock Domain
Clocking the shift path on SCLK removes the per-bit ratio precondition and works at SCLK faster than the system clock, which oversampling cannot. It replaces it with a per-word requirement — a measured twenty-four-fold relief — and costs a synthesis-time mode, a hand-off register that must outlive its transaction, and a reset edge chip select does not supply.
- Related topic
Back-to-Back Transactions and Inter-Frame Gap
How soon chip select may fall again after it rises: the minimum deselect time, why a device needs it, what the gap costs in throughput, and the hardware enforcement that keeps software from violating it.
- Related topic
Slave Microarchitecture and Clocking Assumptions
The two architectures available to a slave that does not own its clock, the one precondition the chosen architecture rests on and why no simulation can test it, why MOSI must pass through exactly as many flops as SCLK, and a delay-matched front end verified in three HDLs.
- Related topic
Why an SPI Slave Is a Clock-Domain Problem
An externally generated SCLK turns a shift register into a CDC question, and there are only four ways to carry an event across: a level, a synchronised level, a toggle, and a handshake. Each is limited by something different, no scheme creates bandwidth, and the difference that matters most is the one no simulation can show.
