Skip to content
VLSI Mentor

SPI · Module 9

Maximum Practical SCLK

The four ceilings on clock rate and which one actually binds: the round-trip budget dominated by the slave's valid time, signal integrity on a shared net, divisor granularity, and the per-slave rate limiter that keeps a mixed bus honest.

Every chapter of this module has treated SCLK as a number that could simply be raised. This one asks what stops you.

The flash is rated 104 MHz, the FPGA can generate 100 MHz, and the link fails above 42 MHz. Nothing is broken. What decided 42?

Four separate limits produce that number, and only one of them is in a datasheet.

1. Four Ceilings, and Which One Binds

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   f_max  =  min( device maximum,        ← from the datasheet
                  round-trip limit,      ← from the board + the device's t_v
                  signal integrity,      ← from edge rates and loading
                  achievable divisor )   ← from the master's clock and divider

The interesting property of that list is that the first one almost never binds. A 104 MHz flash rarely runs at 104 MHz on a real board, because the second and third limits arrive first — and the fourth then rounds you down.

Take them in turn.

2. The Device Maximum, and Its Fine Print

A datasheet's headline frequency is a conditional number. It typically assumes a specified load capacitance (often 15 or 30 pF), a supply voltage at the top of its range, a temperature range, and — critically — a particular command. A flash rated 104 MHz for fast-read is frequently rated 50 MHz or less for the plain read opcode, which is exactly why the dummy cycles of Chapter 4.5 exist.

Three practical consequences:

  • Check which opcode the rating applies to. A driver that uses the basic read command gets the basic read's limit no matter what the front page says.
  • Check the load. A rating at 15 pF on a board presenting 40 pF is not the same rating; the output transition takes longer and t_v grows.
  • Derate for the voltage and temperature corner you actually ship, not the typical one.

This limit is real but rarely the binding one. The next two arrive first.

3. The Round Trip — the Limit That Usually Decides

This is Chapter 6.3's argument, now turned into the number in the question above, and the budget Chapter 6.4 counts in cycles.

The master launches an SCLK edge. That edge travels to the slave, the slave's output responds after t_v, the response travels back, and the master must see it stable before its own sampling edge. On a mode-0 link the sampling edge is half a period after the launching edge:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   T/2  ≥  t_pd(out) + t_v(slave) + t_pd(back) + t_su(master)

   f_max  =  1 / ( 2 × [ t_pd_out + t_v + t_pd_back + t_su ] )

Insert representative numbers — a short board, an ordinary flash:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   master clock-to-out .....   2.0 ns
   trace out ...............   0.6 ns
   slave t_v ...............   7.0 ns
   trace back ..............   0.6 ns
   master setup ............   1.5 ns
   ─────────────────────────────────
   total ...................  11.7 ns   →  T ≥ 23.4 ns  →  f_max = 42.7 MHz

There is the 42 MHz. The 104 MHz device is honest, the FPGA is honest, and the link still fails above 42 — because the half-period budget ran out long before either component did.

And the reason the half-period is the budget rather than the full period is mode-0's launch-and-sample-on-opposite-edges structure. There are two escapes — sampling a half cycle later than the nominal edge, or clocking the return path with a clock the slave sends back — and they turn T/2 into T or remove the round trip from the budget altogether. Without one of those, the arithmetic above is the ceiling.

4. The Failure Looks Like a Bit Shift

MISO below and above the round-trip ceiling

7 cycles
Two MISO lanes sampled at the same point. Below the ceiling the first sample carries D7. Above the ceiling the first sample is unknown and the data is shifted one position later.first samplefirst sampleat 20 MHz?D7D6D5D4D3D2at 60 MHz??D7D6D5D4D3t0t1t2t3t4t5t6
Figure 1 — what exceeding the round-trip limit looks like on MISO. Below the ceiling the first sample already carries D7; above it the round trip has not completed, the first sample catches the idle line, and every byte arrives shifted by one bit position.

That shift is the signature worth memorising. A link pushed past its round-trip limit does not produce random noise — it produces data that is consistently wrong in a structured way, most often every byte shifted one bit, which reads as a plausible-looking value rather than obvious garbage. A read of 0x9F returns 0x3E or 0x4F depending on the neighbouring bit, and a driver checking only for a non-zero response is satisfied.

5. Signal Integrity — the Ceiling Nobody Calculates

Even inside the timing budget, the electrical link has its own limit.

Rise and fall times. A CMOS output driving a capacitive load has a transition time roughly proportional to that load. At 50 MHz a period is 20 ns, and a 6 ns rise time already consumes 30% of it. The usual guideline is to keep the transition under about 10% of the period; past that, the receiver's input threshold is crossed so late that the effective t_su grows and the §3 budget quietly shrinks.

Load capacitance. Every additional slave on the bus adds its input capacitance — a few pF each — plus trace. A bus that works with one device can fail with four, having changed nothing else. This is the mechanism behind Chapter 8.2's observation that fan-out costs speed.

Reflections. Above roughly 25 MHz, an unterminated trace of any length starts to ring. A series resistor of 22 to 47 Ω at the driver end is the standard, nearly free fix, and it is why so many SPI schematics have one in each line.

Crosstalk. SCLK is the aggressor — it is the only line switching every cycle. A clock routed adjacent to MISO for a long run couples into it, and the coupling is worst exactly where MISO is settling.

None of these produce a clean number the way §3 does, which is precisely why they are skipped: the calculation has no closed form, so it gets replaced by "it worked on the bench". Then the fourth board, with a longer trace and a warmer enclosure, does not work.

6. Divisor Granularity — You Cannot Have the Rate You Computed

Suppose §2 through §5 yield a ceiling of 42 MHz. You still cannot run at 42 MHz.

SPI masters derive SCLK by integer division of a system clock. With an 80 MHz system clock:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   divisor   SCLK        vs 42 MHz ceiling
      1     80.0 MHz     over  ✗
      2     40.0 MHz     under ✓   ← the best available
      3     26.7 MHz     under ✓
      4     20.0 MHz     under ✓

You get 40 MHz, not 42 — a 5% loss to rounding. And note how coarse the ladder is at the top: the first three options are 80, 40 and 26.7. Between 40 and 80 there is nothing at all.

Two consequences that matter more than the 5%:

  • Always round down. Selecting the divisor whose rate is nearest the target rather than the largest rate not exceeding it produces a link that is over its ceiling by a few percent — which, per §4, fails structurally rather than obviously.
  • The ladder is coarsest where you most want resolution. Many masters restrict divisors to powers of two, in which case the choices above 26.7 MHz are 40 and 80 and nothing else. A ceiling of 39 MHz costs you a drop to 26.7.

This is also why changing the system clock — for a power mode, say — silently changes SCLK on every device unless the divisors are recomputed. A system clock halved for a low-power mode halves every SCLK, which is safe; one raised for a performance mode raises every SCLK, which is not.

7. Building the Rate Limiter — Three HDLs

The circuit

Circuit. A per-slave divisor floor with a clamp and a flag.

State. The selected slave's minimum divisor, looked up from a table loaded at reset or by software; the effective divisor; a limited flag.

Datapath. A larger divisor means a slower clock, so the safe-rate clamp is a maximum on the divisor, not a minimum on it. This inversion is the single most common bug in rate-limiting logic, and it fails in the dangerous direction — a clamp written as a minimum lets the slowest device run at the fastest rate.

Control. No sequencing. The lookup and comparison are a single cycle from a change of selection or requested divisor.

Clock and reset. System clock; asynchronous active-low reset. Reset loads the slowest divisor, so an un-initialised master cannot overclock anything.

Enables. limited reports that the request was overridden, which is what lets software discover that a device is capping the bus without having to read back a rate.

Timing. eff_div is registered, so it settles one cycle after a change of selection. A master must therefore not begin clocking in the same cycle it changes slaves — which is already guaranteed by the bus turnaround of Chapter 8.4, and is stated in the header for anyone integrating it without that.

Synthesis. A small lookup table, one comparator, two registers. The table dominates, and for eight slaves at an 8-bit divisor that is 64 bits.

Limitations. The table is per-slave, not per-opcode. A device with a fast-read limit and a slower basic-read limit (§2) needs the opcode folded into the lookup, which widens the table but changes nothing structurally.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_rate_limit.sv — the clamp is a maximum on the divisor
// spi_rate_limit.sv — per-slave clock rate limiting.
//
// A shared SPI bus runs at one rate at a time, and every device on it has its
// own maximum. The naive arrangement picks one divisor for the whole system
// -- necessarily the one the SLOWEST device can take -- and every fast device
// then runs far below its capability (Chapter 8.2 §10).
//
// The fix is to carry a per-slave floor and clamp the request against the
// floor of whichever slave is selected. Note the direction: a LARGER divisor
// is a SLOWER clock, so the limit is a MAXIMUM on frequency and therefore a
// MINIMUM on the divisor -- and clamping means taking the larger value.
// Getting that inversion backwards over-clocks every slow device on the bus.
module spi_rate_limit #(
    parameter int DIV_W    = 16,
    parameter int N_SLAVES = 4,
    parameter int IDX_W    = 2
) (
    input  logic                        clk,
    input  logic                        rst_n,
    input  logic [IDX_W-1:0]            sel_idx,
    input  logic [DIV_W-1:0]            req_div,      // what software asked for
    // Per-slave minimum divisor, flattened so the port is portable across
    // all three languages. Entry k occupies bits [k*DIV_W +: DIV_W].
    input  logic [N_SLAVES*DIV_W-1:0]   min_div_flat,
    output logic [DIV_W-1:0]            eff_div,      // what the divider gets
    output logic                        limited       // the request was clamped
);
    // TIMING CONSTRAINT. `eff_div` is REGISTERED, so for one cycle after
    // `sel_idx` changes it still holds the previous slave's divisor. Selecting
    // a slower slave and clocking it immediately would over-clock it for that
    // cycle.
    //
    // This is not a defect to be worked around: the bus turnaround interval
    // (Chapter 8.4) already has nothing selected, and the divisor settles
    // inside it. The rule is simply that the selection must be presented
    // before the frame opens -- the same rule the turnaround guard enforces
    // for its own reasons.
    //
    // A combinational output would remove the constraint and put a comparator
    // and a mux in the divider's clock-control path, which is the wrong trade
    // at high system clocks.
    logic [DIV_W-1:0] this_min;

    always_comb begin
        this_min = '0;
        for (int k = 0; k < N_SLAVES; k++)
            if (k == int'(sel_idx))
                this_min = min_div_flat[k*DIV_W +: DIV_W];
    end

    always_ff @(posedge clk or negedge rst_n) begin
        if (!rst_n) begin
            // Reset to the slowest representable rate. A divider that comes
            // out of reset fast would clock a device above its maximum before
            // software has configured anything.
            eff_div <= '1;
            limited <= 1'b0;
        end else begin
            // Larger divisor = slower clock, so the clamp is a maximum.
            if (this_min > req_div) begin
                eff_div <= this_min;
                limited <= 1'b1;
            end else begin
                eff_div <= req_div;
                limited <= 1'b0;
            end
        end
    end
endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_rate_limit_tb.sv — clamping, passthrough, reset safety and the transition window
// spi_rate_limit_tb.sv — a mixed-speed bus, and the invariant that no slave
// is ever clocked faster than its own maximum.
`timescale 1ns/1ps
module spi_rate_limit_tb;
    logic clk = 0, rst_n = 0;
    always #5 clk = ~clk;

    localparam int DW = 16, N = 4, IW = 2;

    // A realistic mixed bus: a fast flash, a slow EEPROM, a display, and a
    // sensor with no particular limit.
    localparam logic [DW-1:0] MIN0 = 16'd2;    // flash   -- fastest
    localparam logic [DW-1:0] MIN1 = 16'd50;   // EEPROM  -- slowest
    localparam logic [DW-1:0] MIN2 = 16'd8;    // display
    localparam logic [DW-1:0] MIN3 = 16'd1;    // sensor  -- no real limit

    logic [IW-1:0] sel_idx = '0;
    logic [DW-1:0] req_div = '0;
    logic [N*DW-1:0] min_div_flat = {MIN3, MIN2, MIN1, MIN0};
    logic [DW-1:0] eff_div;
    logic limited;

    spi_rate_limit #(.DIV_W(DW), .N_SLAVES(N), .IDX_W(IW)) dut (
        .clk, .rst_n, .sel_idx, .req_div, .min_div_flat, .eff_div, .limited);

    int errors = 0;
    logic [DW-1:0] mins [4];

    task automatic chk(input string what, input int g, input int e);
        if (g !== e) begin $display("FAIL %s: got %0d exp %0d", what, g, e); errors++; end
    endtask

    // CONTINUOUS INVARIANT: while a selection is STABLE -- which is the only
    // time a frame may be open -- the effective divisor is never below that
    // slave's minimum. The one-cycle settling after a change is excluded
    // deliberately: it is the documented constraint, not a violation, and the
    // turnaround interval covers it.
    logic [IW-1:0] sel_q;
    always @(posedge clk) if (rst_n) begin
        sel_q <= sel_idx;
        if ((sel_idx == sel_q) && eff_div < mins[sel_idx]) begin
            $display("FAIL: slave %0d clocked too fast (eff=%0d min=%0d)",
                     sel_idx, eff_div, mins[sel_idx]);
            errors++;
        end
    end

    task automatic settle(); repeat (3) @(negedge clk); endtask

    initial begin
        mins[0] = MIN0; mins[1] = MIN1; mins[2] = MIN2; mins[3] = MIN3;
        repeat (3) @(negedge clk); rst_n = 1;
        // Give the invariant checker a defined selection before it runs.
        sel_idx = 2'd0; req_div = 16'd50; settle();

        // --- a fast request: only the slaves that need it are clamped ---
        req_div = 16'd4;
        sel_idx = 2'd0; settle();
        chk("flash: not limited",   limited, 0);
        chk("flash: div",           eff_div, 4);

        sel_idx = 2'd1; settle();
        chk("EEPROM: limited",      limited, 1);
        chk("EEPROM: clamped to 50", eff_div, 50);

        sel_idx = 2'd2; settle();
        chk("display: limited",     limited, 1);
        chk("display: clamped to 8", eff_div, 8);

        sel_idx = 2'd3; settle();
        chk("sensor: not limited",  limited, 0);
        chk("sensor: div",          eff_div, 4);

        // --- a slow request is never SPED UP, even for a fast slave ---
        req_div = 16'd100;
        for (int i = 0; i < N; i++) begin
            sel_idx = i[IW-1:0]; settle();
            chk($sformatf("slave %0d: slow request honoured", i), eff_div, 100);
            chk($sformatf("slave %0d: not limited",            i), limited, 0);
        end

        // --- a request exactly at a slave's floor is not "limited" ---
        req_div = 16'd8;
        sel_idx = 2'd2; settle();
        chk("display at exactly its floor: div", eff_div, 8);
        chk("display at exactly its floor: not limited", limited, 0);

        // --- one below the floor IS limited ---
        req_div = 16'd7;
        sel_idx = 2'd2; settle();
        chk("one below the floor: clamped", eff_div, 8);
        chk("one below the floor: limited", limited, 1);

        // --- switching to the slow slave mid-stream must clamp immediately ---
        req_div = 16'd2;
        sel_idx = 2'd0; settle();
        chk("flash at full speed", eff_div, 2);
        sel_idx = 2'd1; settle();
        chk("switch to EEPROM clamps", eff_div, 50);
        sel_idx = 2'd0; settle();
        chk("switch back restores", eff_div, 2);

        if (errors == 0)
            $display("PASS: each slave is clamped to its own floor, a slow request is never sped up, a request exactly at the floor is not reported as limited, and switching slaves re-clamps immediately");
        else
            $display("FAILED with %0d error(s)", errors);
        $finish;
    end

    initial begin #900000; $display("FAIL: watchdog timeout"); $finish; end
endmodule

The testbench's most valuable check is the safety invariant: for a stable selection, eff_div is never smaller than the selected slave's minimum. That single property catches the inverted-comparison bug regardless of which slave, which request or which table — and it is far stronger than checking the individual clamping cases.

The qualification "for a stable selection" is not a convenience. It is the timing constraint from the header, made explicit: during the cycle in which selection changes, the registered output still holds the previous slave's divisor. The testbench asserting the invariant unconditionally would fail on correct hardware, which is how a real design constraint gets mistaken for a bug — and then "fixed" by removing the register.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_rate_limit.v — the same clamp in Verilog-2001
// spi_rate_limit.v — the same per-slave rate limiter in Verilog-2001.
module spi_rate_limit #(
    parameter DIV_W    = 16,
    parameter N_SLAVES = 4,
    parameter IDX_W    = 2
) (
    input  wire                      clk,
    input  wire                      rst_n,
    input  wire [IDX_W-1:0]          sel_idx,
    input  wire [DIV_W-1:0]          req_div,
    // Per-slave minimum divisor, flattened. Entry k is bits [k*DIV_W +: DIV_W].
    input  wire [N_SLAVES*DIV_W-1:0] min_div_flat,
    output reg  [DIV_W-1:0]          eff_div,
    output reg                       limited
);
    // TIMING CONSTRAINT: eff_div is registered, so it settles one cycle after
    // a selection change. The bus turnaround (Chapter 8.4) covers that.

    integer k;
    reg [DIV_W-1:0] this_min;

    always @(*) begin
        this_min = {DIV_W{1'b0}};
        for (k = 0; k < N_SLAVES; k = k + 1)
            if (k == sel_idx)
                this_min = min_div_flat[k*DIV_W +: DIV_W];
    end

    always @(posedge clk or negedge rst_n) begin
        if (!rst_n) begin
            // Slowest representable rate out of reset: a divider that came
            // out fast would over-clock a slow device before configuration.
            eff_div <= {DIV_W{1'b1}};
            limited <= 1'b0;
        end else begin
            // Larger divisor = slower clock, so the clamp is a maximum.
            if (this_min > req_div) begin
                eff_div <= this_min;
                limited <= 1'b1;
            end else begin
                eff_div <= req_div;
                limited <= 1'b0;
            end
        end
    end
endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_rate_limit_tb.v — the same checks in Verilog-2001
// spi_rate_limit_tb.v — the same checks in Verilog-2001.
`timescale 1ns/1ps
module spi_rate_limit_tb;
    reg clk = 0, rst_n = 0;
    always #5 clk = ~clk;

    parameter DW = 16, N = 4, IW = 2;
    localparam [DW-1:0] MIN0 = 16'd2, MIN1 = 16'd50, MIN2 = 16'd8, MIN3 = 16'd1;

    reg [IW-1:0] sel_idx = 0;
    reg [DW-1:0] req_div = 0;
    reg [N*DW-1:0] min_div_flat = {MIN3, MIN2, MIN1, MIN0};
    wire [DW-1:0] eff_div;
    wire limited;

    spi_rate_limit #(.DIV_W(DW), .N_SLAVES(N), .IDX_W(IW)) dut (
        .clk(clk), .rst_n(rst_n), .sel_idx(sel_idx), .req_div(req_div),
        .min_div_flat(min_div_flat), .eff_div(eff_div), .limited(limited));

    integer errors = 0, i;
    reg [DW-1:0] mins [0:3];
    reg [IW-1:0] sel_q;

    task chk;
        input [80*8-1:0] what;
        input [31:0] g, e;
        begin
            if (g !== e) begin
                $display("FAIL %0s: got %0d exp %0d", what, g, e);
                errors = errors + 1;
            end
        end
    endtask

    // Invariant, qualified by a stable selection -- the one-cycle settling is
    // the documented constraint, covered by the bus turnaround.
    always @(posedge clk) if (rst_n) begin
        sel_q <= sel_idx;
        if ((sel_idx == sel_q) && eff_div < mins[sel_idx]) begin
            $display("FAIL: slave %0d clocked too fast (eff=%0d min=%0d)",
                     sel_idx, eff_div, mins[sel_idx]);
            errors = errors + 1;
        end
    end

    task settle; begin repeat (3) @(negedge clk); end endtask

    initial begin
        mins[0] = MIN0; mins[1] = MIN1; mins[2] = MIN2; mins[3] = MIN3;
        repeat (3) @(negedge clk); rst_n = 1;
        sel_idx = 2'd0; req_div = 16'd50; settle;

        req_div = 16'd4;
        sel_idx = 2'd0; settle;
        chk("flash: not limited",    limited, 0);
        chk("flash: div",            eff_div, 4);
        sel_idx = 2'd1; settle;
        chk("EEPROM: limited",       limited, 1);
        chk("EEPROM: clamped to 50", eff_div, 50);
        sel_idx = 2'd2; settle;
        chk("display: limited",      limited, 1);
        chk("display: clamped to 8", eff_div, 8);
        sel_idx = 2'd3; settle;
        chk("sensor: not limited",   limited, 0);
        chk("sensor: div",           eff_div, 4);

        req_div = 16'd100;
        for (i = 0; i < N; i = i + 1) begin
            sel_idx = i[IW-1:0]; settle;
            chk("slow request honoured", eff_div, 100);
            chk("not limited",           limited, 0);
        end

        req_div = 16'd8;
        sel_idx = 2'd2; settle;
        chk("display at exactly its floor: div",         eff_div, 8);
        chk("display at exactly its floor: not limited", limited, 0);

        req_div = 16'd7;
        sel_idx = 2'd2; settle;
        chk("one below the floor: clamped", eff_div, 8);
        chk("one below the floor: limited", limited, 1);

        req_div = 16'd2;
        sel_idx = 2'd0; settle;
        chk("flash at full speed", eff_div, 2);
        sel_idx = 2'd1; settle;
        chk("switch to EEPROM clamps", eff_div, 50);
        sel_idx = 2'd0; settle;
        chk("switch back restores", eff_div, 2);

        if (errors == 0)
            $display("PASS: each slave is clamped to its own floor, a slow request is never sped up, a request exactly at the floor is not reported as limited, and switching slaves re-clamps immediately");
        else
            $display("FAILED with %0d error(s)", errors);
        $finish;
    end

    initial begin #900000; $display("FAIL: watchdog timeout"); $finish; end
endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_rate_limit.vhd — the same clamp in VHDL
-- spi_rate_limit.vhd — the same per-slave rate limiter in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

entity spi_rate_limit is
    generic (
        DIV_W    : positive := 16;
        N_SLAVES : positive := 4;
        IDX_W    : positive := 2
    );
    port (
        clk          : in  std_logic;
        rst_n        : in  std_logic;
        sel_idx      : in  unsigned(IDX_W - 1 downto 0);
        req_div      : in  unsigned(DIV_W - 1 downto 0);
        -- Per-slave minimum divisor, flattened so the port shape matches the
        -- Verilog versions. Entry k is bits [(k+1)*DIV_W-1 downto k*DIV_W].
        min_div_flat : in  std_logic_vector(N_SLAVES * DIV_W - 1 downto 0);
        eff_div      : out unsigned(DIV_W - 1 downto 0);
        limited      : out std_logic
    );
end entity spi_rate_limit;

architecture rtl of spi_rate_limit is
    -- TIMING CONSTRAINT: eff_div is registered, so it settles one cycle after
    -- a selection change. The bus turnaround interval covers that, and a
    -- combinational output would put a comparator in the divider's control
    -- path instead.
    signal this_min : unsigned(DIV_W - 1 downto 0);
begin

    select_min : process (sel_idx, min_div_flat) is
        variable m : unsigned(DIV_W - 1 downto 0);
    begin
        m := (others => '0');
        for k in 0 to N_SLAVES - 1 loop
            if k = to_integer(sel_idx) then
                m := unsigned(min_div_flat((k + 1) * DIV_W - 1 downto k * DIV_W));
            end if;
        end loop;
        this_min <= m;
    end process;

    process (clk, rst_n) is
    begin
        if rst_n = '0' then
            -- Slowest representable rate out of reset.
            eff_div <= (others => '1');
            limited <= '0';
        elsif rising_edge(clk) then
            -- Larger divisor = slower clock, so the clamp is a maximum.
            if this_min > req_div then
                eff_div <= this_min;
                limited <= '1';
            else
                eff_div <= req_div;
                limited <= '0';
            end if;
        end if;
    end process;

end architecture rtl;
Azvya Education Pvt. Ltd.VLSI Mentor
spi_rate_limit_tb.vhd — the same checks in VHDL
-- spi_rate_limit_tb.vhd — the same checks in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

entity spi_rate_limit_tb is
end entity spi_rate_limit_tb;

architecture tb of spi_rate_limit_tb is
    constant DW : positive := 16;
    constant N  : positive := 4;
    constant IW : positive := 2;

    constant MIN0 : natural := 2;    -- flash   -- fastest
    constant MIN1 : natural := 50;   -- EEPROM  -- slowest
    constant MIN2 : natural := 8;    -- display
    constant MIN3 : natural := 1;    -- sensor

    signal clk     : std_logic := '0';
    signal rst_n   : std_logic := '0';
    signal sel_idx : unsigned(IW - 1 downto 0) := (others => '0');
    signal req_div : unsigned(DW - 1 downto 0) := (others => '0');
    signal halt    : boolean := false;

    signal min_div_flat : std_logic_vector(N * DW - 1 downto 0) :=
        std_logic_vector(to_unsigned(MIN3, DW)) &
        std_logic_vector(to_unsigned(MIN2, DW)) &
        std_logic_vector(to_unsigned(MIN1, DW)) &
        std_logic_vector(to_unsigned(MIN0, DW));

    signal eff_div : unsigned(DW - 1 downto 0);
    signal limited : std_logic;

    signal errors  : natural := 0;
    signal inv_err : natural := 0;

    type min_arr is array (0 to 3) of natural;
    constant MINS : min_arr := (MIN0, MIN1, MIN2, MIN3);
begin

    clk <= not clk after 5 ns when not halt else '0';

    dut : entity work.spi_rate_limit
        generic map (DIV_W => DW, N_SLAVES => N, IDX_W => IW)
        port map (clk => clk, rst_n => rst_n, sel_idx => sel_idx, req_div => req_div,
                  min_div_flat => min_div_flat, eff_div => eff_div, limited => limited);

    -- Invariant, qualified by a stable selection.
    invariant : process (clk) is
        variable sel_q : unsigned(IW - 1 downto 0) := (others => '0');
    begin
        if rising_edge(clk) and rst_n = '1' then
            if sel_idx = sel_q and to_integer(eff_div) < MINS(to_integer(sel_idx)) then
                report "FAIL: a slave was clocked faster than its floor" severity error;
                inv_err <= inv_err + 1;
            end if;
            sel_q := sel_idx;
        end if;
    end process;

    stim : process is
        procedure chk_n (what : string; g, e : natural) is
        begin
            if g /= e then
                report "FAIL " & what & ": got " & integer'image(g)
                    & " exp " & integer'image(e) severity error;
                errors <= errors + 1;
            end if;
        end procedure;

        procedure chk_b (what : string; g, e : std_logic) is
        begin
            if g /= e then
                report "FAIL " & what severity error;
                errors <= errors + 1;
            end if;
        end procedure;

        procedure settle is
        begin
            for i in 0 to 2 loop wait until falling_edge(clk); end loop;
        end procedure;
    begin
        for i in 0 to 2 loop wait until falling_edge(clk); end loop;
        rst_n <= '1';
        sel_idx <= to_unsigned(0, IW);
        req_div <= to_unsigned(50, DW);
        settle;

        req_div <= to_unsigned(4, DW);
        sel_idx <= to_unsigned(0, IW); settle;
        chk_b("flash: not limited",    limited, '0');
        chk_n("flash: div",            to_integer(eff_div), 4);
        sel_idx <= to_unsigned(1, IW); settle;
        chk_b("EEPROM: limited",       limited, '1');
        chk_n("EEPROM: clamped to 50", to_integer(eff_div), 50);
        sel_idx <= to_unsigned(2, IW); settle;
        chk_b("display: limited",      limited, '1');
        chk_n("display: clamped to 8", to_integer(eff_div), 8);
        sel_idx <= to_unsigned(3, IW); settle;
        chk_b("sensor: not limited",   limited, '0');
        chk_n("sensor: div",           to_integer(eff_div), 4);

        req_div <= to_unsigned(100, DW);
        for i in 0 to N - 1 loop
            sel_idx <= to_unsigned(i, IW); settle;
            chk_n("slow request honoured", to_integer(eff_div), 100);
            chk_b("not limited",           limited, '0');
        end loop;

        req_div <= to_unsigned(8, DW);
        sel_idx <= to_unsigned(2, IW); settle;
        chk_n("display at exactly its floor: div",         to_integer(eff_div), 8);
        chk_b("display at exactly its floor: not limited", limited, '0');

        req_div <= to_unsigned(7, DW);
        sel_idx <= to_unsigned(2, IW); settle;
        chk_n("one below the floor: clamped", to_integer(eff_div), 8);
        chk_b("one below the floor: limited", limited, '1');

        req_div <= to_unsigned(2, DW);
        sel_idx <= to_unsigned(0, IW); settle;
        chk_n("flash at full speed", to_integer(eff_div), 2);
        sel_idx <= to_unsigned(1, IW); settle;
        chk_n("switch to EEPROM clamps", to_integer(eff_div), 50);
        sel_idx <= to_unsigned(0, IW); settle;
        chk_n("switch back restores", to_integer(eff_div), 2);

        chk_n("invariant never violated", inv_err, 0);

        if errors = 0 then
            report "PASS: each slave is clamped to its own floor, a slow request is "
                 & "never sped up, a request exactly at the floor is not reported as "
                 & "limited, and switching slaves re-clamps immediately" severity note;
        else
            report "FAILED with " & integer'image(errors) & " error(s)" severity error;
        end if;
        halt <= true;
        wait;
    end process;

    watchdog : process is
    begin
        wait for 900 us;
        if not halt then
            report "FAIL: watchdog timeout" severity failure;
        end if;
        wait;
    end process;

end architecture tb;

Parity

All three implement the same limiter: identical ports and generics, asynchronous active-low reset loading the slowest divisor, a per-slave minimum-divisor table, a clamp that is a maximum on the divisor, and a limited flag. All three testbenches run the same scenarios — a fast request to a slow slave being clamped, a slow request to a fast slave passing through, reset safety, and the invariant under stable selection — and report identical results.

8. Why a Verification Engineer Cares

Azvya Education Pvt. Ltd.VLSI Mentor
spi_rate_limit.sva — the safety property and its qualification
   // 1. SAFETY. For a stable selection, the effective divisor is never
   //    smaller than the selected slave's minimum. Larger divisor = slower
   //    clock, so "never smaller" means "never faster than allowed".
   //    This one property catches the inverted comparison in every case.
   a_never_too_fast : assert property (
       @(posedge clk) disable iff (!rst_n)
           (sel_stable) |-> (eff_div >= min_div_table[sel]))
       else $error("effective divisor is faster than the slave permits");

   // 2. TRANSPARENCY. A request at or below the limit is not altered --
   //    a limiter that clamps everything is "safe" and useless.
   a_no_false_clamp : assert property (
       @(posedge clk) disable iff (!rst_n)
           (sel_stable && req_div >= min_div_table[sel])
               |=> (eff_div == $past(req_div) && !limited))
       else $error("a legal request was clamped");

   // 3. The flag tells the truth -- software reads it to discover the cap.
   a_flag_honest : assert property (
       @(posedge clk) disable iff (!rst_n)
           (limited) |-> (eff_div != $past(req_div)))
       else $error("limited asserted without a clamp having occurred");

   // 4. RESET SAFETY. Out of reset the divisor is the slowest available,
   //    so an un-initialised master cannot overclock a device.
   a_reset_safe : assert property (
       @(posedge clk) (!rst_n) |=> (eff_div == '1))
       else $error("reset did not load the safe divisor");

Properties 1 and 2 are a matched pair and neither is sufficient alone. Safety says the limiter never permits a rate that is too fast; transparency says it does not clamp what it should not. A limiter that drives eff_div to all-ones unconditionally satisfies safety perfectly and is worthless — which is the general shape of the trap: a safety property alone is always satisfiable by doing nothing.

Property 4 is the one that catches a real integration failure rather than a logic bug. The window between reset release and software configuring the divisors is short, and on most designs nothing happens in it — until a boot ROM issues a flash read before the rate table is loaded.

What these prove. That the clamp is correct, honest and safe at reset. What they cannot prove is that the table's values are right — those come from §2 through §6 for each device, and a perfectly correct limiter loaded with optimistic numbers enforces the wrong ceiling flawlessly.

Coverage should target the relationship, not the values:

Azvya Education Pvt. Ltd.VLSI Mentor
spi_rate_cg.sv — the cases that matter are relational
   covergroup spi_rate_cg @(posedge clk iff sel_stable);
       // The relationship between request and limit is the whole design.
       cp_rel : coverpoint rel {
           bins far_below  = {REQ_MUCH_SLOWER};   // passes through
           bins just_below = {REQ_ONE_SLOWER};    // passes through, boundary
           bins exact      = {REQ_EQUAL};         // must NOT set limited
           bins just_above = {REQ_ONE_FASTER};    // must clamp, boundary
           bins far_above  = {REQ_MUCH_FASTER};   // must clamp
       }

       // Which slave -- a suite exercising only the fastest device never
       // clamps anything, and reports full coverage of the divisors.
       cp_sel : coverpoint sel;

       // Selection changes are where the registered output is stale.
       cp_change : coverpoint sel_changed { bins stable = {0}; bins changed = {1}; }

       x_rel_sel : cross cp_rel, cp_sel;
   endgroup

The exact and just_above bins are the pair that matters: they are one divisor step apart and they must behave differently. A comparison written with >= instead of > gets exact wrong — it clamps a legal request and asserts limited — while every other bin passes.

9. Why an FPGA or ASIC Engineer Cares

Put the ceiling in hardware, not in a comment. Every device's limit is known at design time. A table in the controller makes it impossible for a driver to overclock a device by omission, and costs 64 bits for eight slaves.

Reset to the slowest rate, always. The safe direction is unambiguous and the cost is a few slow transactions during boot. The alternative — resetting to the fastest and relying on software to slow down — fails in exactly the window where nothing else is running to notice.

Budget the round trip before choosing the device. §3's arithmetic runs on a datasheet, before any board exists. Discovering after layout that a 104 MHz flash tops out at 42 MHz on your board is a discovery worth making in the schematic review.

Add the series resistors. 22 to 47 Ω at the driver end of SCLK, MOSI and MISO costs three components and removes the most common cause of an unexplained frequency ceiling. Populating them and shorting them later is far cheaper than finding room for them on a respin.

Expose the effective rate. A readable register showing the divisor actually in use, plus the limited flag, turns "the bus seems slow" into a one-line diagnosis. Without it, software believes whatever it last wrote.

Symptom. A design ships. Three prototype boards run the flash at 40 MHz for weeks without error. The fourth, assembled from the same files, returns corrupted data within minutes. Dropping to 20 MHz fixes it. The boards are electrically identical.

What "identical boards behave differently" establishes. The design is sitting at the edge of a limit rather than inside or outside it. A design comfortably inside works on every board; one outside fails on every board. Board-to-board variation only becomes visible when the margin is smaller than the variation.

Plausible mechanisms.

  • The round-trip budget has no margin. §3's 42 MHz is a typical-corner number; the slave's t_v varies with process, voltage and temperature, and a device at the slow corner in a warm enclosure has a lower ceiling than the calculation suggested.
  • A slower device population. A different date code, or a second source, with a longer t_v inside its own specification.
  • Assembly variation — a different solder profile, a component seated differently — altering load capacitance enough to matter at 40 MHz.
  • Missing or mis-stuffed series resistors on the boards that fail, leaving the lines unterminated.
  • Supply or temperature at a different point in the range on the failing unit.

The discriminating observation. Heat the working boards. If a board that passes at room temperature fails at 70 °C, the margin is the problem and the device population is a red herring — t_v grows with temperature, so a marginal design fails hot. If a heated board keeps working while the fourth fails cold, the fourth board's components or assembly differ, and the investigation moves to the date codes and the resistors.

Then read the corrupted data: per §4, a round-trip violation shifts bits, so a consistent one-bit shift confirms timing rather than noise.

The fix, and what it costs. Drop to the next divisor down and check the margin against the slow corner, not the typical one. Per Chapter 9.3's §1, the efficiency cost of halving SCLK is a doubling of transfer time but no change at all to the overhead fraction — the link is simply slower, not less efficient. Or sample a half cycle later than the nominal edge, which doubles the budget and usually recovers the rate outright.

Why the investigation goes wrong. Because "three out of four boards work" reads as a manufacturing defect, and the fourth board gets inspected, reworked and replaced while the design stays at 40 MHz. The fourth board is not defective — it is the first one to reveal that the design never had margin. Every board is at risk; only one has shown it yet.

11. Common Misconceptions

12. Reason It Through

Work this before reading the answer.

A board runs a 50 MHz-rated sensor at 25 MHz with no errors. A second, identical sensor is added to the same bus, on its own chip select. Now both sensors return corrupted data above 12 MHz — including the one that worked fine before. No firmware changed.

What happened, and what are the options?

Start with what changed. One device was added. The original sensor's own t_v did not change, and neither did its trace. Yet its ceiling halved.

So the change must be on a shared line. MISO and SCLK are shared; only the chip selects are not. Adding a device adds its input capacitance to SCLK and MOSI, and its output-stage capacitance to MISO even while tri-stated — a tri-stated output still presents its pad and ESD capacitance to the net.

Work out the mechanism. Per §5, transition time scales roughly with load. Doubling the capacitance on SCLK roughly doubles its rise time. The receiving device's input threshold is therefore crossed later, which delays its internal clock edge — and per §3 that delay comes straight out of the half-period budget. A ceiling that was 25 MHz with margin becomes 12 MHz without.

Why does the original sensor fail too? Because the degradation is on the shared net, not in either device. This is the observation that solves it: a fault in the new sensor could not affect the old one's transfers, since the new one is deselected during them. Only a shared resource can, and the only shared resources are the three signal lines and the supply.

What are the options, best first?

  1. Series termination — 22 to 47 Ω at each driver. Cheapest, often sufficient on its own, and it addresses ringing as well as the edge.
  2. Shorten the stub to the added device. A long branch to the second sensor is worse than the capacitance itself; bringing it close to the trunk helps disproportionately.
  3. Run at 12 MHz. It works and costs throughput. Per Chapter 9.3, the efficiency is unchanged — the link is simply slower.
  4. Split the bus — a second SPI controller, or a buffer per branch. This is what Chapter 8.2 means by fan-out costing speed, and it is the structural fix when a bus must carry many devices fast.
  5. Stronger drive strength, if the master offers it. It reduces the transition time but worsens ringing and emissions, so it is the option to try after termination, not before.

The general lesson. Adding a device to an SPI bus lowers the ceiling for every device on it. The bus is a shared electrical resource, and its speed is a property of the whole net rather than of any one device. The number of slaves belongs in the frequency budget from the start, alongside t_v and the trace delay — and a design that must support "up to eight slaves" must be budgeted at eight, not at the two on the prototype.

13. Understanding Check

14. Summary

Four limits set SCLK, and the device's rated maximum — the only one on the front page of a datasheet — is usually the loosest:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   f_max = min( device max, round-trip, signal integrity, achievable divisor )

The round trip normally binds. On a mode-0 link the budget is a half period, and it is dominated by the slave's t_v rather than by trace delay — which is why looking up t_v beats shortening traces. A 104 MHz device on a short board can top out near 42 MHz with nothing broken.

Exceeding it produces a consistent bit shift, not obvious garbage — plausible values that pass a non-zero check.

Signal integrity has no closed form and so gets skipped: edges should stay under about a tenth of the period, every added slave loads the shared net, and series resistors of 22 to 47 Ω at the driver are the cheapest fix in the whole protocol. Adding a device lowers the ceiling for every device, so the slave count belongs in the budget from the start.

Divisor granularity then rounds the result down, and the ladder is coarsest exactly at the top. Round down, never to the nearest.

In hardware the limiter is a per-slave divisor table with a clamp that is a maximum on the divisor, reset to the slowest rate so an un-initialised master cannot overclock anything.

For verification, safety must be paired with transparency — a limiter that clamps everything is safe and useless — with the safety property qualified by a stable selection, because the registered output is deliberately one cycle behind a change of slave.

15. What Comes Next

This closes Module 9. Performance is now a set of numbers you can compute rather than a rate you hope for: what the link delivers after overhead, where the burst-length curve flattens, and what the board will actually permit.

Module 10 — Real Devices and Datasheet Interpretation turns that outward. Every number this module used came from a datasheet — t_v, the dummy count, the rated maximum and the opcode it applies to — and Module 10 is the repeatable route through one: finding the interface section, reading the timing table into a budget like §3's, decoding the command table and the register map, and arriving at the exact transaction the RTL must issue.

It is also where the other direction opens. Every ceiling in this chapter is a limit on one data line, and when frequency cannot rise, width can — the multi-line variants multiply throughput without touching SCLK at all. But the first step is reading what a given device actually supports, which is what the next module teaches.

Continue learning