SPI · Module 9
Maximum Practical SCLK
The four ceilings on clock rate and which one actually binds: the round-trip budget dominated by the slave's valid time, signal integrity on a shared net, divisor granularity, and the per-slave rate limiter that keeps a mixed bus honest.
Every chapter of this module has treated SCLK as a number that could simply be raised. This one asks what stops you.
The flash is rated 104 MHz, the FPGA can generate 100 MHz, and the link fails above 42 MHz. Nothing is broken. What decided 42?
Four separate limits produce that number, and only one of them is in a datasheet.
1. Four Ceilings, and Which One Binds
f_max = min( device maximum, ← from the datasheet
round-trip limit, ← from the board + the device's t_v
signal integrity, ← from edge rates and loading
achievable divisor ) ← from the master's clock and dividerThe interesting property of that list is that the first one almost never binds. A 104 MHz flash rarely runs at 104 MHz on a real board, because the second and third limits arrive first — and the fourth then rounds you down.
Take them in turn.
2. The Device Maximum, and Its Fine Print
A datasheet's headline frequency is a conditional number. It typically assumes a specified load capacitance (often 15 or 30 pF), a supply voltage at the top of its range, a temperature range, and — critically — a particular command. A flash rated 104 MHz for fast-read is frequently rated 50 MHz or less for the plain read opcode, which is exactly why the dummy cycles of Chapter 4.5 exist.
Three practical consequences:
- Check which opcode the rating applies to. A driver that uses the basic read command gets the basic read's limit no matter what the front page says.
- Check the load. A rating at 15 pF on a board presenting 40 pF is not the same rating; the output transition takes longer and
t_vgrows. - Derate for the voltage and temperature corner you actually ship, not the typical one.
This limit is real but rarely the binding one. The next two arrive first.
3. The Round Trip — the Limit That Usually Decides
This is Chapter 6.3's argument, now turned into the number in the question above, and the budget Chapter 6.4 counts in cycles.
The master launches an SCLK edge. That edge travels to the slave, the slave's output responds after t_v, the response travels back, and the master must see it stable before its own sampling edge. On a mode-0 link the sampling edge is half a period after the launching edge:
T/2 ≥ t_pd(out) + t_v(slave) + t_pd(back) + t_su(master)
f_max = 1 / ( 2 × [ t_pd_out + t_v + t_pd_back + t_su ] )Insert representative numbers — a short board, an ordinary flash:
master clock-to-out ..... 2.0 ns
trace out ............... 0.6 ns
slave t_v ............... 7.0 ns
trace back .............. 0.6 ns
master setup ............ 1.5 ns
─────────────────────────────────
total ................... 11.7 ns → T ≥ 23.4 ns → f_max = 42.7 MHzThere is the 42 MHz. The 104 MHz device is honest, the FPGA is honest, and the link still fails above 42 — because the half-period budget ran out long before either component did.
And the reason the half-period is the budget rather than the full period is mode-0's launch-and-sample-on-opposite-edges structure. There are two escapes — sampling a half cycle later than the nominal edge, or clocking the return path with a clock the slave sends back — and they turn T/2 into T or remove the round trip from the budget altogether. Without one of those, the arithmetic above is the ceiling.
4. The Failure Looks Like a Bit Shift
MISO below and above the round-trip ceiling
7 cyclesThat shift is the signature worth memorising. A link pushed past its round-trip limit does not produce random noise — it produces data that is consistently wrong in a structured way, most often every byte shifted one bit, which reads as a plausible-looking value rather than obvious garbage. A read of 0x9F returns 0x3E or 0x4F depending on the neighbouring bit, and a driver checking only for a non-zero response is satisfied.
5. Signal Integrity — the Ceiling Nobody Calculates
Even inside the timing budget, the electrical link has its own limit.
Rise and fall times. A CMOS output driving a capacitive load has a transition time roughly proportional to that load. At 50 MHz a period is 20 ns, and a 6 ns rise time already consumes 30% of it. The usual guideline is to keep the transition under about 10% of the period; past that, the receiver's input threshold is crossed so late that the effective t_su grows and the §3 budget quietly shrinks.
Load capacitance. Every additional slave on the bus adds its input capacitance — a few pF each — plus trace. A bus that works with one device can fail with four, having changed nothing else. This is the mechanism behind Chapter 8.2's observation that fan-out costs speed.
Reflections. Above roughly 25 MHz, an unterminated trace of any length starts to ring. A series resistor of 22 to 47 Ω at the driver end is the standard, nearly free fix, and it is why so many SPI schematics have one in each line.
Crosstalk. SCLK is the aggressor — it is the only line switching every cycle. A clock routed adjacent to MISO for a long run couples into it, and the coupling is worst exactly where MISO is settling.
None of these produce a clean number the way §3 does, which is precisely why they are skipped: the calculation has no closed form, so it gets replaced by "it worked on the bench". Then the fourth board, with a longer trace and a warmer enclosure, does not work.
6. Divisor Granularity — You Cannot Have the Rate You Computed
Suppose §2 through §5 yield a ceiling of 42 MHz. You still cannot run at 42 MHz.
SPI masters derive SCLK by integer division of a system clock. With an 80 MHz system clock:
divisor SCLK vs 42 MHz ceiling
1 80.0 MHz over ✗
2 40.0 MHz under ✓ ← the best available
3 26.7 MHz under ✓
4 20.0 MHz under ✓You get 40 MHz, not 42 — a 5% loss to rounding. And note how coarse the ladder is at the top: the first three options are 80, 40 and 26.7. Between 40 and 80 there is nothing at all.
Two consequences that matter more than the 5%:
- Always round down. Selecting the divisor whose rate is nearest the target rather than the largest rate not exceeding it produces a link that is over its ceiling by a few percent — which, per §4, fails structurally rather than obviously.
- The ladder is coarsest where you most want resolution. Many masters restrict divisors to powers of two, in which case the choices above 26.7 MHz are 40 and 80 and nothing else. A ceiling of 39 MHz costs you a drop to 26.7.
This is also why changing the system clock — for a power mode, say — silently changes SCLK on every device unless the divisors are recomputed. A system clock halved for a low-power mode halves every SCLK, which is safe; one raised for a performance mode raises every SCLK, which is not.
7. Building the Rate Limiter — Three HDLs
The circuit
Circuit. A per-slave divisor floor with a clamp and a flag.
State. The selected slave's minimum divisor, looked up from a table loaded at reset or by software; the effective divisor; a limited flag.
Datapath. A larger divisor means a slower clock, so the safe-rate clamp is a maximum on the divisor, not a minimum on it. This inversion is the single most common bug in rate-limiting logic, and it fails in the dangerous direction — a clamp written as a minimum lets the slowest device run at the fastest rate.
Control. No sequencing. The lookup and comparison are a single cycle from a change of selection or requested divisor.
Clock and reset. System clock; asynchronous active-low reset. Reset loads the slowest divisor, so an un-initialised master cannot overclock anything.
Enables. limited reports that the request was overridden, which is what lets software discover that a device is capping the bus without having to read back a rate.
Timing. eff_div is registered, so it settles one cycle after a change of selection. A master must therefore not begin clocking in the same cycle it changes slaves — which is already guaranteed by the bus turnaround of Chapter 8.4, and is stated in the header for anyone integrating it without that.
Synthesis. A small lookup table, one comparator, two registers. The table dominates, and for eight slaves at an 8-bit divisor that is 64 bits.
Limitations. The table is per-slave, not per-opcode. A device with a fast-read limit and a slower basic-read limit (§2) needs the opcode folded into the lookup, which widens the table but changes nothing structurally.
// spi_rate_limit.sv — per-slave clock rate limiting.
//
// A shared SPI bus runs at one rate at a time, and every device on it has its
// own maximum. The naive arrangement picks one divisor for the whole system
// -- necessarily the one the SLOWEST device can take -- and every fast device
// then runs far below its capability (Chapter 8.2 §10).
//
// The fix is to carry a per-slave floor and clamp the request against the
// floor of whichever slave is selected. Note the direction: a LARGER divisor
// is a SLOWER clock, so the limit is a MAXIMUM on frequency and therefore a
// MINIMUM on the divisor -- and clamping means taking the larger value.
// Getting that inversion backwards over-clocks every slow device on the bus.
module spi_rate_limit #(
parameter int DIV_W = 16,
parameter int N_SLAVES = 4,
parameter int IDX_W = 2
) (
input logic clk,
input logic rst_n,
input logic [IDX_W-1:0] sel_idx,
input logic [DIV_W-1:0] req_div, // what software asked for
// Per-slave minimum divisor, flattened so the port is portable across
// all three languages. Entry k occupies bits [k*DIV_W +: DIV_W].
input logic [N_SLAVES*DIV_W-1:0] min_div_flat,
output logic [DIV_W-1:0] eff_div, // what the divider gets
output logic limited // the request was clamped
);
// TIMING CONSTRAINT. `eff_div` is REGISTERED, so for one cycle after
// `sel_idx` changes it still holds the previous slave's divisor. Selecting
// a slower slave and clocking it immediately would over-clock it for that
// cycle.
//
// This is not a defect to be worked around: the bus turnaround interval
// (Chapter 8.4) already has nothing selected, and the divisor settles
// inside it. The rule is simply that the selection must be presented
// before the frame opens -- the same rule the turnaround guard enforces
// for its own reasons.
//
// A combinational output would remove the constraint and put a comparator
// and a mux in the divider's clock-control path, which is the wrong trade
// at high system clocks.
logic [DIV_W-1:0] this_min;
always_comb begin
this_min = '0;
for (int k = 0; k < N_SLAVES; k++)
if (k == int'(sel_idx))
this_min = min_div_flat[k*DIV_W +: DIV_W];
end
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
// Reset to the slowest representable rate. A divider that comes
// out of reset fast would clock a device above its maximum before
// software has configured anything.
eff_div <= '1;
limited <= 1'b0;
end else begin
// Larger divisor = slower clock, so the clamp is a maximum.
if (this_min > req_div) begin
eff_div <= this_min;
limited <= 1'b1;
end else begin
eff_div <= req_div;
limited <= 1'b0;
end
end
end
endmodule// spi_rate_limit_tb.sv — a mixed-speed bus, and the invariant that no slave
// is ever clocked faster than its own maximum.
`timescale 1ns/1ps
module spi_rate_limit_tb;
logic clk = 0, rst_n = 0;
always #5 clk = ~clk;
localparam int DW = 16, N = 4, IW = 2;
// A realistic mixed bus: a fast flash, a slow EEPROM, a display, and a
// sensor with no particular limit.
localparam logic [DW-1:0] MIN0 = 16'd2; // flash -- fastest
localparam logic [DW-1:0] MIN1 = 16'd50; // EEPROM -- slowest
localparam logic [DW-1:0] MIN2 = 16'd8; // display
localparam logic [DW-1:0] MIN3 = 16'd1; // sensor -- no real limit
logic [IW-1:0] sel_idx = '0;
logic [DW-1:0] req_div = '0;
logic [N*DW-1:0] min_div_flat = {MIN3, MIN2, MIN1, MIN0};
logic [DW-1:0] eff_div;
logic limited;
spi_rate_limit #(.DIV_W(DW), .N_SLAVES(N), .IDX_W(IW)) dut (
.clk, .rst_n, .sel_idx, .req_div, .min_div_flat, .eff_div, .limited);
int errors = 0;
logic [DW-1:0] mins [4];
task automatic chk(input string what, input int g, input int e);
if (g !== e) begin $display("FAIL %s: got %0d exp %0d", what, g, e); errors++; end
endtask
// CONTINUOUS INVARIANT: while a selection is STABLE -- which is the only
// time a frame may be open -- the effective divisor is never below that
// slave's minimum. The one-cycle settling after a change is excluded
// deliberately: it is the documented constraint, not a violation, and the
// turnaround interval covers it.
logic [IW-1:0] sel_q;
always @(posedge clk) if (rst_n) begin
sel_q <= sel_idx;
if ((sel_idx == sel_q) && eff_div < mins[sel_idx]) begin
$display("FAIL: slave %0d clocked too fast (eff=%0d min=%0d)",
sel_idx, eff_div, mins[sel_idx]);
errors++;
end
end
task automatic settle(); repeat (3) @(negedge clk); endtask
initial begin
mins[0] = MIN0; mins[1] = MIN1; mins[2] = MIN2; mins[3] = MIN3;
repeat (3) @(negedge clk); rst_n = 1;
// Give the invariant checker a defined selection before it runs.
sel_idx = 2'd0; req_div = 16'd50; settle();
// --- a fast request: only the slaves that need it are clamped ---
req_div = 16'd4;
sel_idx = 2'd0; settle();
chk("flash: not limited", limited, 0);
chk("flash: div", eff_div, 4);
sel_idx = 2'd1; settle();
chk("EEPROM: limited", limited, 1);
chk("EEPROM: clamped to 50", eff_div, 50);
sel_idx = 2'd2; settle();
chk("display: limited", limited, 1);
chk("display: clamped to 8", eff_div, 8);
sel_idx = 2'd3; settle();
chk("sensor: not limited", limited, 0);
chk("sensor: div", eff_div, 4);
// --- a slow request is never SPED UP, even for a fast slave ---
req_div = 16'd100;
for (int i = 0; i < N; i++) begin
sel_idx = i[IW-1:0]; settle();
chk($sformatf("slave %0d: slow request honoured", i), eff_div, 100);
chk($sformatf("slave %0d: not limited", i), limited, 0);
end
// --- a request exactly at a slave's floor is not "limited" ---
req_div = 16'd8;
sel_idx = 2'd2; settle();
chk("display at exactly its floor: div", eff_div, 8);
chk("display at exactly its floor: not limited", limited, 0);
// --- one below the floor IS limited ---
req_div = 16'd7;
sel_idx = 2'd2; settle();
chk("one below the floor: clamped", eff_div, 8);
chk("one below the floor: limited", limited, 1);
// --- switching to the slow slave mid-stream must clamp immediately ---
req_div = 16'd2;
sel_idx = 2'd0; settle();
chk("flash at full speed", eff_div, 2);
sel_idx = 2'd1; settle();
chk("switch to EEPROM clamps", eff_div, 50);
sel_idx = 2'd0; settle();
chk("switch back restores", eff_div, 2);
if (errors == 0)
$display("PASS: each slave is clamped to its own floor, a slow request is never sped up, a request exactly at the floor is not reported as limited, and switching slaves re-clamps immediately");
else
$display("FAILED with %0d error(s)", errors);
$finish;
end
initial begin #900000; $display("FAIL: watchdog timeout"); $finish; end
endmoduleThe testbench's most valuable check is the safety invariant: for a stable selection, eff_div is never smaller than the selected slave's minimum. That single property catches the inverted-comparison bug regardless of which slave, which request or which table — and it is far stronger than checking the individual clamping cases.
The qualification "for a stable selection" is not a convenience. It is the timing constraint from the header, made explicit: during the cycle in which selection changes, the registered output still holds the previous slave's divisor. The testbench asserting the invariant unconditionally would fail on correct hardware, which is how a real design constraint gets mistaken for a bug — and then "fixed" by removing the register.
// spi_rate_limit.v — the same per-slave rate limiter in Verilog-2001.
module spi_rate_limit #(
parameter DIV_W = 16,
parameter N_SLAVES = 4,
parameter IDX_W = 2
) (
input wire clk,
input wire rst_n,
input wire [IDX_W-1:0] sel_idx,
input wire [DIV_W-1:0] req_div,
// Per-slave minimum divisor, flattened. Entry k is bits [k*DIV_W +: DIV_W].
input wire [N_SLAVES*DIV_W-1:0] min_div_flat,
output reg [DIV_W-1:0] eff_div,
output reg limited
);
// TIMING CONSTRAINT: eff_div is registered, so it settles one cycle after
// a selection change. The bus turnaround (Chapter 8.4) covers that.
integer k;
reg [DIV_W-1:0] this_min;
always @(*) begin
this_min = {DIV_W{1'b0}};
for (k = 0; k < N_SLAVES; k = k + 1)
if (k == sel_idx)
this_min = min_div_flat[k*DIV_W +: DIV_W];
end
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
// Slowest representable rate out of reset: a divider that came
// out fast would over-clock a slow device before configuration.
eff_div <= {DIV_W{1'b1}};
limited <= 1'b0;
end else begin
// Larger divisor = slower clock, so the clamp is a maximum.
if (this_min > req_div) begin
eff_div <= this_min;
limited <= 1'b1;
end else begin
eff_div <= req_div;
limited <= 1'b0;
end
end
end
endmodule// spi_rate_limit_tb.v — the same checks in Verilog-2001.
`timescale 1ns/1ps
module spi_rate_limit_tb;
reg clk = 0, rst_n = 0;
always #5 clk = ~clk;
parameter DW = 16, N = 4, IW = 2;
localparam [DW-1:0] MIN0 = 16'd2, MIN1 = 16'd50, MIN2 = 16'd8, MIN3 = 16'd1;
reg [IW-1:0] sel_idx = 0;
reg [DW-1:0] req_div = 0;
reg [N*DW-1:0] min_div_flat = {MIN3, MIN2, MIN1, MIN0};
wire [DW-1:0] eff_div;
wire limited;
spi_rate_limit #(.DIV_W(DW), .N_SLAVES(N), .IDX_W(IW)) dut (
.clk(clk), .rst_n(rst_n), .sel_idx(sel_idx), .req_div(req_div),
.min_div_flat(min_div_flat), .eff_div(eff_div), .limited(limited));
integer errors = 0, i;
reg [DW-1:0] mins [0:3];
reg [IW-1:0] sel_q;
task chk;
input [80*8-1:0] what;
input [31:0] g, e;
begin
if (g !== e) begin
$display("FAIL %0s: got %0d exp %0d", what, g, e);
errors = errors + 1;
end
end
endtask
// Invariant, qualified by a stable selection -- the one-cycle settling is
// the documented constraint, covered by the bus turnaround.
always @(posedge clk) if (rst_n) begin
sel_q <= sel_idx;
if ((sel_idx == sel_q) && eff_div < mins[sel_idx]) begin
$display("FAIL: slave %0d clocked too fast (eff=%0d min=%0d)",
sel_idx, eff_div, mins[sel_idx]);
errors = errors + 1;
end
end
task settle; begin repeat (3) @(negedge clk); end endtask
initial begin
mins[0] = MIN0; mins[1] = MIN1; mins[2] = MIN2; mins[3] = MIN3;
repeat (3) @(negedge clk); rst_n = 1;
sel_idx = 2'd0; req_div = 16'd50; settle;
req_div = 16'd4;
sel_idx = 2'd0; settle;
chk("flash: not limited", limited, 0);
chk("flash: div", eff_div, 4);
sel_idx = 2'd1; settle;
chk("EEPROM: limited", limited, 1);
chk("EEPROM: clamped to 50", eff_div, 50);
sel_idx = 2'd2; settle;
chk("display: limited", limited, 1);
chk("display: clamped to 8", eff_div, 8);
sel_idx = 2'd3; settle;
chk("sensor: not limited", limited, 0);
chk("sensor: div", eff_div, 4);
req_div = 16'd100;
for (i = 0; i < N; i = i + 1) begin
sel_idx = i[IW-1:0]; settle;
chk("slow request honoured", eff_div, 100);
chk("not limited", limited, 0);
end
req_div = 16'd8;
sel_idx = 2'd2; settle;
chk("display at exactly its floor: div", eff_div, 8);
chk("display at exactly its floor: not limited", limited, 0);
req_div = 16'd7;
sel_idx = 2'd2; settle;
chk("one below the floor: clamped", eff_div, 8);
chk("one below the floor: limited", limited, 1);
req_div = 16'd2;
sel_idx = 2'd0; settle;
chk("flash at full speed", eff_div, 2);
sel_idx = 2'd1; settle;
chk("switch to EEPROM clamps", eff_div, 50);
sel_idx = 2'd0; settle;
chk("switch back restores", eff_div, 2);
if (errors == 0)
$display("PASS: each slave is clamped to its own floor, a slow request is never sped up, a request exactly at the floor is not reported as limited, and switching slaves re-clamps immediately");
else
$display("FAILED with %0d error(s)", errors);
$finish;
end
initial begin #900000; $display("FAIL: watchdog timeout"); $finish; end
endmodule-- spi_rate_limit.vhd — the same per-slave rate limiter in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_rate_limit is
generic (
DIV_W : positive := 16;
N_SLAVES : positive := 4;
IDX_W : positive := 2
);
port (
clk : in std_logic;
rst_n : in std_logic;
sel_idx : in unsigned(IDX_W - 1 downto 0);
req_div : in unsigned(DIV_W - 1 downto 0);
-- Per-slave minimum divisor, flattened so the port shape matches the
-- Verilog versions. Entry k is bits [(k+1)*DIV_W-1 downto k*DIV_W].
min_div_flat : in std_logic_vector(N_SLAVES * DIV_W - 1 downto 0);
eff_div : out unsigned(DIV_W - 1 downto 0);
limited : out std_logic
);
end entity spi_rate_limit;
architecture rtl of spi_rate_limit is
-- TIMING CONSTRAINT: eff_div is registered, so it settles one cycle after
-- a selection change. The bus turnaround interval covers that, and a
-- combinational output would put a comparator in the divider's control
-- path instead.
signal this_min : unsigned(DIV_W - 1 downto 0);
begin
select_min : process (sel_idx, min_div_flat) is
variable m : unsigned(DIV_W - 1 downto 0);
begin
m := (others => '0');
for k in 0 to N_SLAVES - 1 loop
if k = to_integer(sel_idx) then
m := unsigned(min_div_flat((k + 1) * DIV_W - 1 downto k * DIV_W));
end if;
end loop;
this_min <= m;
end process;
process (clk, rst_n) is
begin
if rst_n = '0' then
-- Slowest representable rate out of reset.
eff_div <= (others => '1');
limited <= '0';
elsif rising_edge(clk) then
-- Larger divisor = slower clock, so the clamp is a maximum.
if this_min > req_div then
eff_div <= this_min;
limited <= '1';
else
eff_div <= req_div;
limited <= '0';
end if;
end if;
end process;
end architecture rtl;-- spi_rate_limit_tb.vhd — the same checks in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_rate_limit_tb is
end entity spi_rate_limit_tb;
architecture tb of spi_rate_limit_tb is
constant DW : positive := 16;
constant N : positive := 4;
constant IW : positive := 2;
constant MIN0 : natural := 2; -- flash -- fastest
constant MIN1 : natural := 50; -- EEPROM -- slowest
constant MIN2 : natural := 8; -- display
constant MIN3 : natural := 1; -- sensor
signal clk : std_logic := '0';
signal rst_n : std_logic := '0';
signal sel_idx : unsigned(IW - 1 downto 0) := (others => '0');
signal req_div : unsigned(DW - 1 downto 0) := (others => '0');
signal halt : boolean := false;
signal min_div_flat : std_logic_vector(N * DW - 1 downto 0) :=
std_logic_vector(to_unsigned(MIN3, DW)) &
std_logic_vector(to_unsigned(MIN2, DW)) &
std_logic_vector(to_unsigned(MIN1, DW)) &
std_logic_vector(to_unsigned(MIN0, DW));
signal eff_div : unsigned(DW - 1 downto 0);
signal limited : std_logic;
signal errors : natural := 0;
signal inv_err : natural := 0;
type min_arr is array (0 to 3) of natural;
constant MINS : min_arr := (MIN0, MIN1, MIN2, MIN3);
begin
clk <= not clk after 5 ns when not halt else '0';
dut : entity work.spi_rate_limit
generic map (DIV_W => DW, N_SLAVES => N, IDX_W => IW)
port map (clk => clk, rst_n => rst_n, sel_idx => sel_idx, req_div => req_div,
min_div_flat => min_div_flat, eff_div => eff_div, limited => limited);
-- Invariant, qualified by a stable selection.
invariant : process (clk) is
variable sel_q : unsigned(IW - 1 downto 0) := (others => '0');
begin
if rising_edge(clk) and rst_n = '1' then
if sel_idx = sel_q and to_integer(eff_div) < MINS(to_integer(sel_idx)) then
report "FAIL: a slave was clocked faster than its floor" severity error;
inv_err <= inv_err + 1;
end if;
sel_q := sel_idx;
end if;
end process;
stim : process is
procedure chk_n (what : string; g, e : natural) is
begin
if g /= e then
report "FAIL " & what & ": got " & integer'image(g)
& " exp " & integer'image(e) severity error;
errors <= errors + 1;
end if;
end procedure;
procedure chk_b (what : string; g, e : std_logic) is
begin
if g /= e then
report "FAIL " & what severity error;
errors <= errors + 1;
end if;
end procedure;
procedure settle is
begin
for i in 0 to 2 loop wait until falling_edge(clk); end loop;
end procedure;
begin
for i in 0 to 2 loop wait until falling_edge(clk); end loop;
rst_n <= '1';
sel_idx <= to_unsigned(0, IW);
req_div <= to_unsigned(50, DW);
settle;
req_div <= to_unsigned(4, DW);
sel_idx <= to_unsigned(0, IW); settle;
chk_b("flash: not limited", limited, '0');
chk_n("flash: div", to_integer(eff_div), 4);
sel_idx <= to_unsigned(1, IW); settle;
chk_b("EEPROM: limited", limited, '1');
chk_n("EEPROM: clamped to 50", to_integer(eff_div), 50);
sel_idx <= to_unsigned(2, IW); settle;
chk_b("display: limited", limited, '1');
chk_n("display: clamped to 8", to_integer(eff_div), 8);
sel_idx <= to_unsigned(3, IW); settle;
chk_b("sensor: not limited", limited, '0');
chk_n("sensor: div", to_integer(eff_div), 4);
req_div <= to_unsigned(100, DW);
for i in 0 to N - 1 loop
sel_idx <= to_unsigned(i, IW); settle;
chk_n("slow request honoured", to_integer(eff_div), 100);
chk_b("not limited", limited, '0');
end loop;
req_div <= to_unsigned(8, DW);
sel_idx <= to_unsigned(2, IW); settle;
chk_n("display at exactly its floor: div", to_integer(eff_div), 8);
chk_b("display at exactly its floor: not limited", limited, '0');
req_div <= to_unsigned(7, DW);
sel_idx <= to_unsigned(2, IW); settle;
chk_n("one below the floor: clamped", to_integer(eff_div), 8);
chk_b("one below the floor: limited", limited, '1');
req_div <= to_unsigned(2, DW);
sel_idx <= to_unsigned(0, IW); settle;
chk_n("flash at full speed", to_integer(eff_div), 2);
sel_idx <= to_unsigned(1, IW); settle;
chk_n("switch to EEPROM clamps", to_integer(eff_div), 50);
sel_idx <= to_unsigned(0, IW); settle;
chk_n("switch back restores", to_integer(eff_div), 2);
chk_n("invariant never violated", inv_err, 0);
if errors = 0 then
report "PASS: each slave is clamped to its own floor, a slow request is "
& "never sped up, a request exactly at the floor is not reported as "
& "limited, and switching slaves re-clamps immediately" severity note;
else
report "FAILED with " & integer'image(errors) & " error(s)" severity error;
end if;
halt <= true;
wait;
end process;
watchdog : process is
begin
wait for 900 us;
if not halt then
report "FAIL: watchdog timeout" severity failure;
end if;
wait;
end process;
end architecture tb;Parity
All three implement the same limiter: identical ports and generics, asynchronous active-low reset loading the slowest divisor, a per-slave minimum-divisor table, a clamp that is a maximum on the divisor, and a limited flag. All three testbenches run the same scenarios — a fast request to a slow slave being clamped, a slow request to a fast slave passing through, reset safety, and the invariant under stable selection — and report identical results.
8. Why a Verification Engineer Cares
// 1. SAFETY. For a stable selection, the effective divisor is never
// smaller than the selected slave's minimum. Larger divisor = slower
// clock, so "never smaller" means "never faster than allowed".
// This one property catches the inverted comparison in every case.
a_never_too_fast : assert property (
@(posedge clk) disable iff (!rst_n)
(sel_stable) |-> (eff_div >= min_div_table[sel]))
else $error("effective divisor is faster than the slave permits");
// 2. TRANSPARENCY. A request at or below the limit is not altered --
// a limiter that clamps everything is "safe" and useless.
a_no_false_clamp : assert property (
@(posedge clk) disable iff (!rst_n)
(sel_stable && req_div >= min_div_table[sel])
|=> (eff_div == $past(req_div) && !limited))
else $error("a legal request was clamped");
// 3. The flag tells the truth -- software reads it to discover the cap.
a_flag_honest : assert property (
@(posedge clk) disable iff (!rst_n)
(limited) |-> (eff_div != $past(req_div)))
else $error("limited asserted without a clamp having occurred");
// 4. RESET SAFETY. Out of reset the divisor is the slowest available,
// so an un-initialised master cannot overclock a device.
a_reset_safe : assert property (
@(posedge clk) (!rst_n) |=> (eff_div == '1))
else $error("reset did not load the safe divisor");Properties 1 and 2 are a matched pair and neither is sufficient alone. Safety says the limiter never permits a rate that is too fast; transparency says it does not clamp what it should not. A limiter that drives eff_div to all-ones unconditionally satisfies safety perfectly and is worthless — which is the general shape of the trap: a safety property alone is always satisfiable by doing nothing.
Property 4 is the one that catches a real integration failure rather than a logic bug. The window between reset release and software configuring the divisors is short, and on most designs nothing happens in it — until a boot ROM issues a flash read before the rate table is loaded.
What these prove. That the clamp is correct, honest and safe at reset. What they cannot prove is that the table's values are right — those come from §2 through §6 for each device, and a perfectly correct limiter loaded with optimistic numbers enforces the wrong ceiling flawlessly.
Coverage should target the relationship, not the values:
covergroup spi_rate_cg @(posedge clk iff sel_stable);
// The relationship between request and limit is the whole design.
cp_rel : coverpoint rel {
bins far_below = {REQ_MUCH_SLOWER}; // passes through
bins just_below = {REQ_ONE_SLOWER}; // passes through, boundary
bins exact = {REQ_EQUAL}; // must NOT set limited
bins just_above = {REQ_ONE_FASTER}; // must clamp, boundary
bins far_above = {REQ_MUCH_FASTER}; // must clamp
}
// Which slave -- a suite exercising only the fastest device never
// clamps anything, and reports full coverage of the divisors.
cp_sel : coverpoint sel;
// Selection changes are where the registered output is stale.
cp_change : coverpoint sel_changed { bins stable = {0}; bins changed = {1}; }
x_rel_sel : cross cp_rel, cp_sel;
endgroupThe exact and just_above bins are the pair that matters: they are one divisor step apart and they must behave differently. A comparison written with >= instead of > gets exact wrong — it clamps a legal request and asserts limited — while every other bin passes.
9. Why an FPGA or ASIC Engineer Cares
Put the ceiling in hardware, not in a comment. Every device's limit is known at design time. A table in the controller makes it impossible for a driver to overclock a device by omission, and costs 64 bits for eight slaves.
Reset to the slowest rate, always. The safe direction is unambiguous and the cost is a few slow transactions during boot. The alternative — resetting to the fastest and relying on software to slow down — fails in exactly the window where nothing else is running to notice.
Budget the round trip before choosing the device. §3's arithmetic runs on a datasheet, before any board exists. Discovering after layout that a 104 MHz flash tops out at 42 MHz on your board is a discovery worth making in the schematic review.
Add the series resistors. 22 to 47 Ω at the driver end of SCLK, MOSI and MISO costs three components and removes the most common cause of an unexplained frequency ceiling. Populating them and shorting them later is far cheaper than finding room for them on a respin.
Expose the effective rate. A readable register showing the divisor actually in use, plus the limited flag, turns "the bus seems slow" into a one-line diagnosis. Without it, software believes whatever it last wrote.
10. Failure Signature — A Link That Works on Three Boards and Fails on the Fourth
Symptom. A design ships. Three prototype boards run the flash at 40 MHz for weeks without error. The fourth, assembled from the same files, returns corrupted data within minutes. Dropping to 20 MHz fixes it. The boards are electrically identical.
What "identical boards behave differently" establishes. The design is sitting at the edge of a limit rather than inside or outside it. A design comfortably inside works on every board; one outside fails on every board. Board-to-board variation only becomes visible when the margin is smaller than the variation.
Plausible mechanisms.
- The round-trip budget has no margin. §3's 42 MHz is a typical-corner number; the slave's
t_vvaries with process, voltage and temperature, and a device at the slow corner in a warm enclosure has a lower ceiling than the calculation suggested. - A slower device population. A different date code, or a second source, with a longer
t_vinside its own specification. - Assembly variation — a different solder profile, a component seated differently — altering load capacitance enough to matter at 40 MHz.
- Missing or mis-stuffed series resistors on the boards that fail, leaving the lines unterminated.
- Supply or temperature at a different point in the range on the failing unit.
The discriminating observation. Heat the working boards. If a board that passes at room temperature fails at 70 °C, the margin is the problem and the device population is a red herring — t_v grows with temperature, so a marginal design fails hot. If a heated board keeps working while the fourth fails cold, the fourth board's components or assembly differ, and the investigation moves to the date codes and the resistors.
Then read the corrupted data: per §4, a round-trip violation shifts bits, so a consistent one-bit shift confirms timing rather than noise.
The fix, and what it costs. Drop to the next divisor down and check the margin against the slow corner, not the typical one. Per Chapter 9.3's §1, the efficiency cost of halving SCLK is a doubling of transfer time but no change at all to the overhead fraction — the link is simply slower, not less efficient. Or sample a half cycle later than the nominal edge, which doubles the budget and usually recovers the rate outright.
Why the investigation goes wrong. Because "three out of four boards work" reads as a manufacturing defect, and the fourth board gets inspected, reworked and replaced while the design stays at 40 MHz. The fourth board is not defective — it is the first one to reveal that the design never had margin. Every board is at risk; only one has shown it yet.
11. Common Misconceptions
12. Reason It Through
Work this before reading the answer.
A board runs a 50 MHz-rated sensor at 25 MHz with no errors. A second, identical sensor is added to the same bus, on its own chip select. Now both sensors return corrupted data above 12 MHz — including the one that worked fine before. No firmware changed.
What happened, and what are the options?
Start with what changed. One device was added. The original sensor's own t_v did not change, and neither did its trace. Yet its ceiling halved.
So the change must be on a shared line. MISO and SCLK are shared; only the chip selects are not. Adding a device adds its input capacitance to SCLK and MOSI, and its output-stage capacitance to MISO even while tri-stated — a tri-stated output still presents its pad and ESD capacitance to the net.
Work out the mechanism. Per §5, transition time scales roughly with load. Doubling the capacitance on SCLK roughly doubles its rise time. The receiving device's input threshold is therefore crossed later, which delays its internal clock edge — and per §3 that delay comes straight out of the half-period budget. A ceiling that was 25 MHz with margin becomes 12 MHz without.
Why does the original sensor fail too? Because the degradation is on the shared net, not in either device. This is the observation that solves it: a fault in the new sensor could not affect the old one's transfers, since the new one is deselected during them. Only a shared resource can, and the only shared resources are the three signal lines and the supply.
What are the options, best first?
- Series termination — 22 to 47 Ω at each driver. Cheapest, often sufficient on its own, and it addresses ringing as well as the edge.
- Shorten the stub to the added device. A long branch to the second sensor is worse than the capacitance itself; bringing it close to the trunk helps disproportionately.
- Run at 12 MHz. It works and costs throughput. Per Chapter 9.3, the efficiency is unchanged — the link is simply slower.
- Split the bus — a second SPI controller, or a buffer per branch. This is what Chapter 8.2 means by fan-out costing speed, and it is the structural fix when a bus must carry many devices fast.
- Stronger drive strength, if the master offers it. It reduces the transition time but worsens ringing and emissions, so it is the option to try after termination, not before.
The general lesson. Adding a device to an SPI bus lowers the ceiling for every device on it. The bus is a shared electrical resource, and its speed is a property of the whole net rather than of any one device. The number of slaves belongs in the frequency budget from the start, alongside t_v and the trace delay — and a design that must support "up to eight slaves" must be budgeted at eight, not at the two on the prototype.
13. Understanding Check
14. Summary
Four limits set SCLK, and the device's rated maximum — the only one on the front page of a datasheet — is usually the loosest:
f_max = min( device max, round-trip, signal integrity, achievable divisor )The round trip normally binds. On a mode-0 link the budget is a half period, and it is dominated by the slave's t_v rather than by trace delay — which is why looking up t_v beats shortening traces. A 104 MHz device on a short board can top out near 42 MHz with nothing broken.
Exceeding it produces a consistent bit shift, not obvious garbage — plausible values that pass a non-zero check.
Signal integrity has no closed form and so gets skipped: edges should stay under about a tenth of the period, every added slave loads the shared net, and series resistors of 22 to 47 Ω at the driver are the cheapest fix in the whole protocol. Adding a device lowers the ceiling for every device, so the slave count belongs in the budget from the start.
Divisor granularity then rounds the result down, and the ladder is coarsest exactly at the top. Round down, never to the nearest.
In hardware the limiter is a per-slave divisor table with a clamp that is a maximum on the divisor, reset to the slowest rate so an un-initialised master cannot overclock anything.
For verification, safety must be paired with transparency — a limiter that clamps everything is safe and useless — with the safety property qualified by a stable selection, because the registered output is deliberately one cycle behind a change of slave.
15. What Comes Next
This closes Module 9. Performance is now a set of numbers you can compute rather than a rate you hope for: what the link delivers after overhead, where the burst-length curve flattens, and what the board will actually permit.
Module 10 — Real Devices and Datasheet Interpretation turns that outward. Every number this module used came from a datasheet — t_v, the dummy count, the rated maximum and the opcode it applies to — and Module 10 is the repeatable route through one: finding the interface section, reading the timing table into a budget like §3's, decoding the command table and the register map, and arriving at the exact transaction the RTL must issue.
It is also where the other direction opens. Every ceiling in this chapter is a limit on one data line, and when frequency cannot rise, width can — the multi-line variants multiply throughput without touching SCLK at all. But the first step is reading what a given device actually supports, which is what the next module teaches.
Continue learning
Related tutorials
- Related topic
Clock Divider and SCLK Generation
A divider that produces SCLK and the two edge strobes a datapath needs: why the strobes must not be called launch and capture, why an odd divisor puts the longer half first, and why SCLK must be a registered output rather than a gated clock.
- Related topic
Electrical and Board-Level Limits
Why an SPI link with identical logic runs at one clock rate and fails at a higher one. Push-pull drivers into real capacitance, the four delays inside a bit time, the round trip a returned bit must complete, and why usable SCLK is a property of the board.
- Related topic
SCLK Generation, Period, and Frequency
Where SCLK comes from and what one period buys. Dividing a system clock to a bus clock, why the divisor is an integer and what that costs, and how a period in nanoseconds becomes the budget every later timing parameter is spent from.
- Related topic
Launch and Sample Edges
One edge of each bit time places a bit on the wire, the other captures it, and they must never be the same edge. Why the separation is forced, why it buys half a period, and how RTL maps physical edges onto those roles.
