Skip to content
VLSI Mentor

SPI · Module 9

Protocol Overhead

Overhead split into the part SCLK pays for and the part it does not: why the non-clocked share grows from 5% to half the transaction as the clock rises, and the frame timer that partitions a transfer into gap, lead, active and lag.

Chapter 9.1 established that effective throughput is a fraction of the raw rate. This chapter takes the fraction apart.

A transaction takes 1340 ns and delivers 16 bits of payload. Where did the other time go, line by line — and which parts would change if the payload were bigger?

The second half is what makes this more than bookkeeping: overhead divides into two kinds with completely different consequences, and telling them apart decides which optimisation is worth doing.

1. Two Kinds of Overhead

Clocked overhead — command bits, address bits, dummy cycles. These occupy SCLK edges, so they scale inversely with clock rate: double the clock and they take half the time.

Non-clocked overhead — the CS lead, the CS lag, and the inter-frame gap. No bit moves during these, and critically they do not scale with the clock at all. They are fixed in nanoseconds because they are set by device requirements: the lead and lag by the device's internal setup, the gap by its recovery time.

That distinction is the chapter's centre of gravity:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   clocked overhead      shrinks when you raise the clock
   non-clocked overhead  does not

2. Where a Frame's Time Goes

Two clock pulses inside a much longer frame

10 cycles
An SPI frame in which chip select falls, several idle cycles pass, two clock pulses occur, several more idle cycles pass, and chip select rises.CS leadCS leadCS lagCS lagcs_nsclkmovingt0t1t2t3t4t5t6t7t8t9
Figure 1 — a frame with only two clock pulses. Chip select is asserted well before the first edge and released well after the last, and during those intervals no bit moves. For a short transaction these two intervals dominate everything else.

The moving lane marks the interval in which bits are actually being carried. Everything outside it, inside the frame, is non-clocked overhead — and in this figure it is the majority.

3. The Per-Access Budget

For a read with a 1-byte command, 3-byte address and 8 dummy cycles, at 50 MHz with a 100 ns lead, 20 ns lag and 100 ns gap:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   CLOCKED, scales with 1/f_sclk
     command      8 cycles ×  20 ns  =  160 ns
     address     24 cycles ×  20 ns  =  480 ns
     dummy        8 cycles ×  20 ns  =  160 ns
     ─────────────────────────────────────────
     subtotal    40 cycles           =  800 ns

   NON-CLOCKED, fixed in nanoseconds
     CS lead                         =  100 ns
     CS lag                          =   20 ns
     inter-frame gap                 =  100 ns
     ─────────────────────────────────────────
     subtotal                        =  220 ns

   TOTAL OVERHEAD PER ACCESS         = 1020 ns

Against a 1-byte payload (160 ns of clocking) that is 1020 ns of overhead for 160 ns of data — 86% overhead.

4. What Happens When You Raise the Clock

This is where the two kinds separate visibly. Take the same transaction at three clock rates:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   f_sclk    clocked ovh    non-clocked    total    non-clocked share
    10 MHz      4000 ns        220 ns     4220 ns          5 %
    50 MHz       800 ns        220 ns     1020 ns         22 %
   100 MHz       400 ns        220 ns      620 ns         35 %
   200 MHz       200 ns        220 ns      420 ns         52 %

At 200 MHz more than half the overhead is time in which nothing is clocked — and no further increase in clock rate can touch it.

Two conclusions follow.

There is a knee beyond which clock rate stops helping. It arrives when clocked overhead falls to roughly the size of the non-clocked overhead, and for typical devices that is somewhere between 100 and 200 MHz — which is, not coincidentally, around where SPI devices stop getting faster.

Reducing the number of accesses attacks both kinds at once. Merging two transactions into one removes a full command, address, dummy phase, lead, lag and gap. Raising the clock removes a fraction of one of those categories. This is the quantitative form of the advice Chapters 7.1 and 7.2 gave qualitatively.

5. Measuring the Part a Bit Counter Cannot See

Chapter 9.1's counter counts bits. The non-clocked overhead carries no bits, so it is invisible to it — the counter can tell you 40 of 48 bits were overhead, and nothing at all about the 220 ns in which no bit moved.

That is what the frame timer in §6 is for. It splits a frame into lead, active and lag, and separately measures the gap before it, which is the four-way breakdown of §3 made measurable.

Its defining property is that the buckets partition the frame: every in-frame cycle lands in exactly one of lead, active or lag. Without that, a breakdown is a collection of plausible numbers that need not sum to anything.

6. Building the Frame Timer — Three HDLs

The circuit

Circuit. Four accumulators, a published-output register set, and a small state bit.

State. The four running counts, a registered copy of the frame level, and whether any bit has occurred in this frame.

Datapath. None.

Control. Three cases while in-frame — the entry cycle, before the first bit, and after it — plus the gap accumulation while out of frame. The subtlety is the lag bucket: a cycle with no bit is provisionally lag, and if another bit follows it is folded back into active. Only the tail that reaches the end of the frame is genuinely lag.

Clock and reset. System clock; asynchronous active-low reset clearing everything.

Enables. gap_acc is deliberately not cleared when a frame opens — it has been accumulating since the previous frame closed and is this frame's preceding gap. Clearing it there is the natural-looking mistake that loses the measurement.

Timing. The four outputs are published together on the CS rising edge with a single frame_valid strobe, so a consumer reads a coherent set rather than a skewed snapshot (Chapter 9.1 §8).

Synthesis. Eight T_W counters — four accumulating and four published — plus a little control. The doubling is what allows a completed frame to be read while the next is in progress.

Limitations. It measures time, not bits; the two are complementary and a real controller carries both.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_frame_timer.sv — lead, active, lag, and the gap before
// spi_frame_timer.sv — where a frame's time actually goes.
//
// Chapter 9.1 counted BITS. This counts TIME, and the two disagree in a way
// that matters: a frame contains intervals during which chip select is
// asserted and no bit moves at all -- the CS lead and lag (Chapter 2.5) --
// plus the gap between frames (Chapter 7.3). None of those carry a bit, and
// all of them are paid per access.
//
// Splitting a frame into lead / active / lag, plus the preceding gap, turns
// "framing overhead" from a hand-waved term into four measured numbers.
module spi_frame_timer #(
    parameter int T_W = 16
) (
    input  logic           clk,
    input  logic           rst_n,
    input  logic           cs_active,     // level: a frame is open
    input  logic           bit_stb,       // one pulse per SPI bit

    output logic [T_W-1:0] gap_cycles,    // CS high, before this frame
    output logic [T_W-1:0] lead_cycles,   // CS low, before the first bit
    output logic [T_W-1:0] active_cycles, // first bit to last bit
    output logic [T_W-1:0] lag_cycles,    // last bit to CS high
    output logic           frame_valid    // pulse: the four outputs are a frame
);
    logic cs_q;
    logic seen_bit;                        // has any bit occurred this frame?

    // Running accumulators. Kept separate from the outputs so a consumer can
    // read a completed frame's breakdown while the next one is in progress.
    logic [T_W-1:0] gap_acc, lead_acc, active_acc, lag_acc;

    always_ff @(posedge clk or negedge rst_n) begin
        if (!rst_n) begin
            cs_q          <= 1'b0;
            seen_bit      <= 1'b0;
            gap_acc       <= '0;
            lead_acc      <= '0;
            active_acc    <= '0;
            lag_acc       <= '0;
            gap_cycles    <= '0;
            lead_cycles   <= '0;
            active_cycles <= '0;
            lag_cycles    <= '0;
            frame_valid   <= 1'b0;
        end else begin
            cs_q        <= cs_active;
            frame_valid <= 1'b0;

            if (!cs_active) begin
                if (cs_q) begin
                    // CS just rose: publish the frame that has ended.
                    gap_cycles    <= gap_acc;
                    lead_cycles   <= lead_acc;
                    active_cycles <= active_acc;
                    lag_cycles    <= lag_acc;
                    frame_valid   <= 1'b1;
                    gap_acc       <= '0;     // the next gap starts now
                end else begin
                    gap_acc <= gap_acc + 1'b1;
                end
            end else begin
                if (!cs_q) begin
                    // CS just fell: a new frame. Reset the in-frame buckets
                    // but NOT gap_acc, which has been accumulating and is
                    // this frame's preceding gap.
                    //
                    // This entry cycle is itself in-frame and must be counted,
                    // or the buckets do not partition the frame.
                    seen_bit   <= bit_stb;
                    lead_acc   <= bit_stb ? '0 : {{(T_W-1){1'b0}}, 1'b1};
                    active_acc <= bit_stb ? {{(T_W-1){1'b0}}, 1'b1} : '0;
                    lag_acc    <= '0;
                end else if (!seen_bit) begin
                    // Before the first bit: lead time. The cycle carrying the
                    // FIRST bit belongs to active, not lead.
                    if (bit_stb) begin
                        seen_bit   <= 1'b1;
                        active_acc <= active_acc + 1'b1;
                    end else begin
                        lead_acc <= lead_acc + 1'b1;
                    end
                end else begin
                    // After the first bit. A cycle is "active" if a bit has
                    // occurred recently and "lag" otherwise -- so lag is the
                    // tail that turns out to have had no further bits.
                    if (bit_stb) begin
                        // Any accumulated lag was actually part of the active
                        // interval: fold it in and keep going.
                        active_acc <= active_acc + lag_acc + 1'b1;
                        lag_acc    <= '0;
                    end else begin
                        lag_acc <= lag_acc + 1'b1;
                    end
                end
            end
        end
    end
endmodule

The fold-back in the final else branch is the design's one subtle line. A cycle without a bit cannot be classified when it occurs — it is lag only if the frame ends before another bit arrives. Accumulating it provisionally and adding it to active on the next bit resolves that without needing to look ahead.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_frame_timer_tb.sv — a known frame, and the partition property
// spi_frame_timer_tb.sv — a frame with known lead, active and lag, and the
// check that the four buckets account for every cycle.
`timescale 1ns/1ps
module spi_frame_timer_tb;
    logic clk = 0, rst_n = 0;
    always #5 clk = ~clk;

    localparam int TW = 16;
    logic cs_active = 0, bit_stb = 0;
    logic [TW-1:0] gap_cycles, lead_cycles, active_cycles, lag_cycles;
    logic frame_valid;

    spi_frame_timer #(.T_W(TW)) dut (
        .clk, .rst_n, .cs_active, .bit_stb,
        .gap_cycles, .lead_cycles, .active_cycles, .lag_cycles, .frame_valid);

    int errors = 0, frames = 0;
    int g_gap, g_lead, g_active, g_lag;
    int in_frame_cycles = 0, measured_in_frame = 0;

    // Independently count the cycles CS was asserted, so the DUT's buckets
    // can be checked against a number it did not produce.
    always @(posedge clk) if (rst_n) begin
        if (cs_active) in_frame_cycles++;
        if (frame_valid) begin
            measured_in_frame <= in_frame_cycles;
            in_frame_cycles   <= 0;
        end
    end

    task automatic chk(input string what, input int g, input int e);
        if (g !== e) begin $display("FAIL %s: got %0d exp %0d", what, g, e); errors++; end
    endtask

    // Latch a published frame.
    always @(posedge clk) if (rst_n && frame_valid) begin
        frames++;
        g_gap    <= gap_cycles;
        g_lead   <= lead_cycles;
        g_active <= active_cycles;
        g_lag    <= lag_cycles;
    end

    // One bit occupies exactly two system cycles: one with bit_stb high.
    task automatic send_bit();
        @(negedge clk); bit_stb = 1;
        @(negedge clk); bit_stb = 0;
    endtask

    task automatic frame(input int lead, input int n_bits, input int lag);
        cs_active = 1;
        repeat (lead) @(negedge clk);
        for (int i = 0; i < n_bits; i++) send_bit();
        repeat (lag) @(negedge clk);
        cs_active = 0;
        @(negedge clk);
    endtask

    task automatic idle(input int n); repeat (n) @(negedge clk); endtask

    initial begin
        repeat (3) @(negedge clk); rst_n = 1; @(negedge clk);

        // --- a frame with 6 lead cycles, 8 bits, 5 lag cycles, after a gap ---
        idle(12);
        frame(6, 8, 5);
        idle(4);
        chk("one frame published", frames, 1);
        $display("  gap=%0d lead=%0d active=%0d lag=%0d",
                 g_gap, g_lead, g_active, g_lag);

        // The cycle on which CS falls is in-frame and carries no bit, so it
        // is lead -- hence requested + 1.
        chk("lead", g_lead, 6 + 1);
        // The lag is the tail after the last bit strobe, before CS rises.
        chk("lag",  g_lag,  5);
        // Active is whatever remains, which is the partition property stated
        // as an equality rather than a guessed constant.
        chk("active = frame - lead - lag", g_active, measured_in_frame - g_lead - g_lag);
        chk("buckets partition the frame", g_lead + g_active + g_lag, measured_in_frame);
        // The gap is everything CS was high before this frame, including the
        // reset settling -- so check it is at least the idle we inserted.
        if (g_gap < 12) begin
            $display("FAIL: gap %0d should be at least 12", g_gap); errors++;
        end

        // --- a frame with NO lead and NO lag: all time is active ---
        idle(8);
        frame(0, 4, 0);
        idle(4);
        chk("two frames", frames, 2);
        // With zero requested lead, only the CS-fall cycle is lead.
        chk("minimal lead", g_lead, 1);
        chk("no lag",       g_lag,  0);
        chk("active = frame - lead - lag", g_active, measured_in_frame - g_lead - g_lag);
        chk("buckets partition (no lead/lag)", g_lead + g_active + g_lag, measured_in_frame);
        if (g_active <= g_lead) begin
            $display("FAIL: a 4-bit frame with no lead should be mostly active"); errors++;
        end

        // --- a frame with no bits at all: everything is lead ---
        idle(8);
        frame(7, 0, 0);
        idle(4);
        chk("three frames",  frames, 3);
        // With no bits at all, every in-frame cycle is lead. Stating it as
        // an identity is stronger than a constant and says what it means.
        chk("empty: all in-frame time is lead", g_lead, measured_in_frame);
        chk("empty: active", g_active, 0);
        chk("empty: lag",    g_lag,    0);
        chk("buckets partition (empty)", g_lead + g_active + g_lag, measured_in_frame);

        // --- overhead ratio: a short frame is mostly framing ---
        idle(8);
        frame(6, 2, 5);
        idle(4);
        $display("  short frame: lead=%0d active=%0d lag=%0d  -> framing is %0d of %0d cycles",
                 g_lead, g_active, g_lag, g_lead + g_lag, g_lead + g_active + g_lag);
        if (g_lead + g_lag <= g_active) begin
            $display("FAIL: for a 2-bit frame, framing should exceed active time");
            errors++;
        end

        if (errors == 0)
            $display("PASS: lead, active and lag are measured separately, a frame with no bits is all lead, and on a short frame the framing overhead exceeds the clocked time");
        else
            $display("FAILED with %0d error(s)", errors);
        $finish;
    end

    initial begin #900000; $display("FAIL: watchdog timeout"); $finish; end
endmodule

The testbench counts in-frame cycles independently and asserts that lead + active + lag equals that count. That is stronger than checking the three values individually: it verifies the breakdown is a genuine partition rather than three plausible numbers, and it holds for every frame shape without the test needing to know the expected split.

It also reports a result worth quoting: for a 2-bit frame, framing consumes 12 of 15 cycles.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_frame_timer.v — the same timer in Verilog-2001
// spi_frame_timer.v — the same frame-time breakdown in Verilog-2001.
module spi_frame_timer #(
    parameter T_W = 16
) (
    input  wire           clk,
    input  wire           rst_n,
    input  wire           cs_active,
    input  wire           bit_stb,
    output reg  [T_W-1:0] gap_cycles,
    output reg  [T_W-1:0] lead_cycles,
    output reg  [T_W-1:0] active_cycles,
    output reg  [T_W-1:0] lag_cycles,
    output reg            frame_valid
);
    reg cs_q, seen_bit;
    reg [T_W-1:0] gap_acc, lead_acc, active_acc, lag_acc;

    always @(posedge clk or negedge rst_n) begin
        if (!rst_n) begin
            cs_q          <= 1'b0;
            seen_bit      <= 1'b0;
            gap_acc       <= {T_W{1'b0}};
            lead_acc      <= {T_W{1'b0}};
            active_acc    <= {T_W{1'b0}};
            lag_acc       <= {T_W{1'b0}};
            gap_cycles    <= {T_W{1'b0}};
            lead_cycles   <= {T_W{1'b0}};
            active_cycles <= {T_W{1'b0}};
            lag_cycles    <= {T_W{1'b0}};
            frame_valid   <= 1'b0;
        end else begin
            cs_q        <= cs_active;
            frame_valid <= 1'b0;

            if (!cs_active) begin
                if (cs_q) begin
                    // CS just rose: publish the frame that has ended.
                    gap_cycles    <= gap_acc;
                    lead_cycles   <= lead_acc;
                    active_cycles <= active_acc;
                    lag_cycles    <= lag_acc;
                    frame_valid   <= 1'b1;
                    gap_acc       <= {T_W{1'b0}};
                end else begin
                    gap_acc <= gap_acc + 1'b1;
                end
            end else begin
                if (!cs_q) begin
                    // The CS-fall cycle is in-frame and must be counted, or
                    // the buckets do not partition the frame.
                    seen_bit   <= bit_stb;
                    lead_acc   <= bit_stb ? {T_W{1'b0}} : {{(T_W-1){1'b0}}, 1'b1};
                    active_acc <= bit_stb ? {{(T_W-1){1'b0}}, 1'b1} : {T_W{1'b0}};
                    lag_acc    <= {T_W{1'b0}};
                end else if (!seen_bit) begin
                    // Lead time. The cycle carrying the FIRST bit is active.
                    if (bit_stb) begin
                        seen_bit   <= 1'b1;
                        active_acc <= active_acc + 1'b1;
                    end else begin
                        lead_acc <= lead_acc + 1'b1;
                    end
                end else begin
                    // Accumulated lag turns out to be active if another bit
                    // follows; it is only lag once the frame ends.
                    if (bit_stb) begin
                        active_acc <= active_acc + lag_acc + 1'b1;
                        lag_acc    <= {T_W{1'b0}};
                    end else begin
                        lag_acc <= lag_acc + 1'b1;
                    end
                end
            end
        end
    end
endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_frame_timer_tb.v — the same checks in Verilog-2001
// spi_frame_timer_tb.v — the same checks in Verilog-2001.
`timescale 1ns/1ps
module spi_frame_timer_tb;
    reg clk = 0, rst_n = 0;
    always #5 clk = ~clk;

    parameter TW = 16;
    reg cs_active = 0, bit_stb = 0;
    wire [TW-1:0] gap_cycles, lead_cycles, active_cycles, lag_cycles;
    wire frame_valid;

    spi_frame_timer #(.T_W(TW)) dut (
        .clk(clk), .rst_n(rst_n), .cs_active(cs_active), .bit_stb(bit_stb),
        .gap_cycles(gap_cycles), .lead_cycles(lead_cycles),
        .active_cycles(active_cycles), .lag_cycles(lag_cycles),
        .frame_valid(frame_valid));

    integer errors = 0, frames = 0, i;
    integer g_gap = 0, g_lead = 0, g_active = 0, g_lag = 0;
    integer in_frame_cycles = 0, measured_in_frame = 0;

    task chk;
        input [80*8-1:0] what;
        input [31:0] g, e;
        begin
            if (g !== e) begin
                $display("FAIL %0s: got %0d exp %0d", what, g, e);
                errors = errors + 1;
            end
        end
    endtask

    always @(posedge clk) if (rst_n) begin
        if (cs_active) in_frame_cycles = in_frame_cycles + 1;
        if (frame_valid) begin
            frames   = frames + 1;
            g_gap    = gap_cycles;
            g_lead   = lead_cycles;
            g_active = active_cycles;
            g_lag    = lag_cycles;
            measured_in_frame = in_frame_cycles;
            in_frame_cycles   = 0;
        end
    end

    task send_bit;
        begin
            @(negedge clk); bit_stb = 1;
            @(negedge clk); bit_stb = 0;
        end
    endtask

    task frame;
        input integer lead, n_bits, lag;
        integer k;
        begin
            cs_active = 1;
            repeat (lead) @(negedge clk);
            for (k = 0; k < n_bits; k = k + 1) send_bit;
            repeat (lag) @(negedge clk);
            cs_active = 0;
            @(negedge clk);
        end
    endtask

    task idle; input integer n; begin repeat (n) @(negedge clk); end endtask

    initial begin
        repeat (3) @(negedge clk); rst_n = 1; @(negedge clk);

        idle(12);
        frame(6, 8, 5);
        idle(4);
        chk("one frame published", frames, 1);
        $display("  gap=%0d lead=%0d active=%0d lag=%0d", g_gap, g_lead, g_active, g_lag);
        chk("lead", g_lead, 6 + 1);
        chk("lag",  g_lag,  5);
        chk("active = frame - lead - lag", g_active, measured_in_frame - g_lead - g_lag);
        chk("buckets partition the frame", g_lead + g_active + g_lag, measured_in_frame);
        if (g_gap < 12) begin
            $display("FAIL: gap %0d should be at least 12", g_gap); errors = errors + 1;
        end

        idle(8);
        frame(0, 4, 0);
        idle(4);
        chk("two frames", frames, 2);
        chk("minimal lead", g_lead, 1);
        chk("no lag",       g_lag,  0);
        chk("active = frame - lead - lag", g_active, measured_in_frame - g_lead - g_lag);
        chk("buckets partition (no lead/lag)", g_lead + g_active + g_lag, measured_in_frame);
        if (g_active <= g_lead) begin
            $display("FAIL: a 4-bit frame with no lead should be mostly active");
            errors = errors + 1;
        end

        idle(8);
        frame(7, 0, 0);
        idle(4);
        chk("three frames", frames, 3);
        chk("empty: all in-frame time is lead", g_lead, measured_in_frame);
        chk("empty: active", g_active, 0);
        chk("empty: lag",    g_lag,    0);
        chk("buckets partition (empty)", g_lead + g_active + g_lag, measured_in_frame);

        idle(8);
        frame(6, 2, 5);
        idle(4);
        $display("  short frame: lead=%0d active=%0d lag=%0d  -> framing is %0d of %0d cycles",
                 g_lead, g_active, g_lag, g_lead + g_lag, g_lead + g_active + g_lag);
        if (g_lead + g_lag <= g_active) begin
            $display("FAIL: for a 2-bit frame, framing should exceed active time");
            errors = errors + 1;
        end

        if (errors == 0)
            $display("PASS: lead, active and lag are measured separately, a frame with no bits is all lead, and on a short frame the framing overhead exceeds the clocked time");
        else
            $display("FAILED with %0d error(s)", errors);
        $finish;
    end

    initial begin #900000; $display("FAIL: watchdog timeout"); $finish; end
endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_frame_timer.vhd — the same timer in VHDL
-- spi_frame_timer.vhd — the same frame-time breakdown in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

entity spi_frame_timer is
    generic (
        T_W : positive := 16
    );
    port (
        clk           : in  std_logic;
        rst_n         : in  std_logic;
        cs_active     : in  std_logic;                  -- level
        bit_stb       : in  std_logic;                  -- one pulse per SPI bit
        gap_cycles    : out unsigned(T_W - 1 downto 0); -- CS high, before this frame
        lead_cycles   : out unsigned(T_W - 1 downto 0); -- CS low, before the first bit
        active_cycles : out unsigned(T_W - 1 downto 0); -- first bit to last bit
        lag_cycles    : out unsigned(T_W - 1 downto 0); -- last bit to CS high
        frame_valid   : out std_logic                   -- pulse
    );
end entity spi_frame_timer;

architecture rtl of spi_frame_timer is
    signal cs_q, seen_bit : std_logic;
    signal gap_acc, lead_acc, active_acc, lag_acc : unsigned(T_W - 1 downto 0);
begin

    process (clk, rst_n) is
    begin
        if rst_n = '0' then
            cs_q          <= '0';
            seen_bit      <= '0';
            gap_acc       <= (others => '0');
            lead_acc      <= (others => '0');
            active_acc    <= (others => '0');
            lag_acc       <= (others => '0');
            gap_cycles    <= (others => '0');
            lead_cycles   <= (others => '0');
            active_cycles <= (others => '0');
            lag_cycles    <= (others => '0');
            frame_valid   <= '0';
        elsif rising_edge(clk) then
            cs_q        <= cs_active;
            frame_valid <= '0';

            if cs_active = '0' then
                if cs_q = '1' then
                    -- CS just rose: publish the frame that has ended.
                    gap_cycles    <= gap_acc;
                    lead_cycles   <= lead_acc;
                    active_cycles <= active_acc;
                    lag_cycles    <= lag_acc;
                    frame_valid   <= '1';
                    gap_acc       <= (others => '0');
                else
                    gap_acc <= gap_acc + 1;
                end if;
            else
                if cs_q = '0' then
                    -- The CS-fall cycle is in-frame and must be counted, or
                    -- the buckets do not partition the frame.
                    seen_bit <= bit_stb;
                    lag_acc  <= (others => '0');
                    if bit_stb = '1' then
                        lead_acc   <= (others => '0');
                        active_acc <= to_unsigned(1, T_W);
                    else
                        lead_acc   <= to_unsigned(1, T_W);
                        active_acc <= (others => '0');
                    end if;

                elsif seen_bit = '0' then
                    -- Lead time. The cycle carrying the FIRST bit is active.
                    if bit_stb = '1' then
                        seen_bit   <= '1';
                        active_acc <= active_acc + 1;
                    else
                        lead_acc <= lead_acc + 1;
                    end if;

                else
                    -- Accumulated lag turns out to be active if another bit
                    -- follows; it is only lag once the frame ends.
                    if bit_stb = '1' then
                        active_acc <= active_acc + lag_acc + 1;
                        lag_acc    <= (others => '0');
                    else
                        lag_acc <= lag_acc + 1;
                    end if;
                end if;
            end if;
        end if;
    end process;

end architecture rtl;
Azvya Education Pvt. Ltd.VLSI Mentor
spi_frame_timer_tb.vhd — the same checks in VHDL
-- spi_frame_timer_tb.vhd — the same checks in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

entity spi_frame_timer_tb is
end entity spi_frame_timer_tb;

architecture tb of spi_frame_timer_tb is
    constant TW : positive := 16;

    signal clk       : std_logic := '0';
    signal rst_n     : std_logic := '0';
    signal cs_active : std_logic := '0';
    signal bit_stb   : std_logic := '0';
    signal halt      : boolean := false;

    signal gap_cycles    : unsigned(TW - 1 downto 0);
    signal lead_cycles   : unsigned(TW - 1 downto 0);
    signal active_cycles : unsigned(TW - 1 downto 0);
    signal lag_cycles    : unsigned(TW - 1 downto 0);
    signal frame_valid   : std_logic;

    signal errors : natural := 0;
    signal frames : natural := 0;
    signal g_gap, g_lead, g_active, g_lag : natural := 0;
    signal measured_in_frame : natural := 0;
begin

    clk <= not clk after 5 ns when not halt else '0';

    dut : entity work.spi_frame_timer
        generic map (T_W => TW)
        port map (clk => clk, rst_n => rst_n, cs_active => cs_active, bit_stb => bit_stb,
                  gap_cycles => gap_cycles, lead_cycles => lead_cycles,
                  active_cycles => active_cycles, lag_cycles => lag_cycles,
                  frame_valid => frame_valid);

    -- One process owns the observation signals.
    observe : process (clk) is
        variable in_frame : natural := 0;
    begin
        if rising_edge(clk) and rst_n = '1' then
            if cs_active = '1' then
                in_frame := in_frame + 1;
            end if;
            if frame_valid = '1' then
                frames            <= frames + 1;
                g_gap             <= to_integer(gap_cycles);
                g_lead            <= to_integer(lead_cycles);
                g_active          <= to_integer(active_cycles);
                g_lag             <= to_integer(lag_cycles);
                measured_in_frame <= in_frame;
                in_frame          := 0;
            end if;
        end if;
    end process;

    stim : process is
        procedure chk_n (what : string; g, e : natural) is
        begin
            if g /= e then
                report "FAIL " & what & ": got " & integer'image(g)
                    & " exp " & integer'image(e) severity error;
                errors <= errors + 1;
            end if;
        end procedure;

        procedure send_bit is
        begin
            wait until falling_edge(clk); bit_stb <= '1';
            wait until falling_edge(clk); bit_stb <= '0';
        end procedure;

        procedure frame (lead, n_bits, lag : natural) is
        begin
            cs_active <= '1';
            for i in 1 to lead loop wait until falling_edge(clk); end loop;
            for i in 1 to n_bits loop send_bit; end loop;
            for i in 1 to lag loop wait until falling_edge(clk); end loop;
            cs_active <= '0';
            wait until falling_edge(clk);
        end procedure;

        procedure idle (n : natural) is
        begin
            for i in 1 to n loop wait until falling_edge(clk); end loop;
        end procedure;
    begin
        for i in 0 to 2 loop wait until falling_edge(clk); end loop;
        rst_n <= '1';
        wait until falling_edge(clk);

        idle(12);
        frame(6, 8, 5);
        idle(4);
        chk_n("one frame published", frames, 1);
        report "  gap=" & integer'image(g_gap) & " lead=" & integer'image(g_lead)
             & " active=" & integer'image(g_active) & " lag=" & integer'image(g_lag)
             severity note;
        chk_n("lead", g_lead, 6 + 1);
        chk_n("lag",  g_lag,  5);
        chk_n("active = frame - lead - lag", g_active, measured_in_frame - g_lead - g_lag);
        chk_n("buckets partition the frame", g_lead + g_active + g_lag, measured_in_frame);
        if g_gap < 12 then
            report "FAIL: gap too small" severity error;
            errors <= errors + 1;
        end if;

        idle(8);
        frame(0, 4, 0);
        idle(4);
        chk_n("two frames",   frames, 2);
        chk_n("minimal lead", g_lead, 1);
        chk_n("no lag",       g_lag,  0);
        chk_n("active = frame - lead - lag", g_active, measured_in_frame - g_lead - g_lag);
        chk_n("buckets partition (no lead/lag)", g_lead + g_active + g_lag, measured_in_frame);
        if g_active <= g_lead then
            report "FAIL: a 4-bit frame with no lead should be mostly active" severity error;
            errors <= errors + 1;
        end if;

        idle(8);
        frame(7, 0, 0);
        idle(4);
        chk_n("three frames", frames, 3);
        chk_n("empty: all in-frame time is lead", g_lead, measured_in_frame);
        chk_n("empty: active", g_active, 0);
        chk_n("empty: lag",    g_lag,    0);

        idle(8);
        frame(6, 2, 5);
        idle(4);
        report "  short frame: lead=" & integer'image(g_lead)
             & " active=" & integer'image(g_active)
             & " lag=" & integer'image(g_lag)
             & "  -> framing is " & integer'image(g_lead + g_lag)
             & " of " & integer'image(g_lead + g_active + g_lag) & " cycles"
             severity note;
        if g_lead + g_lag <= g_active then
            report "FAIL: for a 2-bit frame, framing should exceed active time" severity error;
            errors <= errors + 1;
        end if;

        if errors = 0 then
            report "PASS: lead, active and lag are measured separately, a frame with no "
                 & "bits is all lead, and on a short frame the framing overhead exceeds "
                 & "the clocked time" severity note;
        else
            report "FAILED with " & integer'image(errors) & " error(s)" severity error;
        end if;
        halt <= true;
        wait;
    end process;

    watchdog : process is
    begin
        wait for 900 us;
        if not halt then
            report "FAIL: watchdog timeout" severity failure;
        end if;
        wait;
    end process;

end architecture tb;

Parity

All three implement the same timer: identical ports and generics, asynchronous active-low reset, the CS-fall cycle counted as in-frame, the first bit's cycle attributed to active rather than lead, provisional lag folded back on a subsequent bit, the gap accumulator preserved across a frame opening, and all four outputs published together. All three testbenches report identical numbers — gap=13, lead=7, active=15, lag=5 for the reference frame, and 12 of 15 cycles framing for the short one.

7. Why a Verification Engineer Cares

Azvya Education Pvt. Ltd.VLSI Mentor
spi_frame_timer.sva — a breakdown that adds up
   // 1. THE property. The three in-frame buckets partition the frame, so the
   //    breakdown is a genuine decomposition rather than three estimates.
   a_partition : assert property (
       @(posedge clk) disable iff (!rst_n)
           frame_valid |-> (lead_cycles + active_cycles + lag_cycles)
                           == in_frame_cycles_ref)
       else $error("lead + active + lag does not equal the frame length");

   // 2. A frame with no bits is entirely lead. Catches an implementation
   //    that leaves such a frame's time unattributed.
   a_empty_is_all_lead : assert property (
       @(posedge clk) disable iff (!rst_n)
           (frame_valid && active_cycles == 0) |-> (lag_cycles == 0))
       else $error("a frame with no active time reported lag");

   // 3. The gap is preserved across the frame opening. Clearing it on CS
   //    fall is the natural-looking mistake, and it silently reports zero.
   a_gap_nonzero_after_idle : assert property (
       @(posedge clk) disable iff (!rst_n)
           (frame_valid && idle_preceded) |-> (gap_cycles > 0))
       else $error("gap reported as zero after an idle interval");

   // 4. All four outputs update together, so a reader never sees a mix of
   //    two frames.
   a_atomic_publish : assert property (
       @(posedge clk) disable iff (!rst_n)
           !frame_valid |=> ($stable(lead_cycles) && $stable(active_cycles) &&
                             $stable(lag_cycles)  && $stable(gap_cycles)))
       else $error("an output changed outside a publish");

Property 1 is the one to write first, and it illustrates a general point about measurement hardware: the useful assertions are accounting identities, not behaviours. A timer that mis-attributes a few cycles still produces plausible numbers, and only a sum that must balance catches it.

What these prove. That the breakdown is coherent and atomically published. What they cannot prove is that the measured lead and lag meet the device's requirements — that is a datasheet comparison, and a design can measure its own violation perfectly.

Coverage should target the frame shapes whose breakdowns differ:

Azvya Education Pvt. Ltd.VLSI Mentor
spi_overhead_cg.sv — the shapes where overhead dominates
   covergroup spi_overhead_cg @(posedge frame_valid);
       // The ratio that matters: is this frame mostly moving data or mostly
       // framing? A suite of long frames never sees the interesting case.
       cp_framing_share : coverpoint
           ((lead_cycles + lag_cycles) * 100 / (lead_cycles + active_cycles + lag_cycles)) {
           bins negligible = {[0:10]};    // long burst
           bins moderate   = {[11:40]};
           bins dominant   = {[41:100]};  // single-register access
       }

       cp_active : coverpoint active_cycles {
           bins none  = {0};              // an empty frame (Chapter 8.6)
           bins tiny  = {[1:16]};
           bins large = {[17:$]};
       }

       // Back-to-back frames have a small gap; isolated ones have a large
       // one. The overhead per access differs enormously between them.
       cp_gap : coverpoint gap_cycles {
           bins minimal = {[0:4]};
           bins modest  = {[5:64]};
           bins idle    = {[65:$]};
       }

       x_framing_active : cross cp_framing_share, cp_active;
   endgroup

8. Why an FPGA or ASIC Engineer Cares

Shortening the lead and lag is a real optimisation and a bounded one. Both are device requirements with datasheet minima, and many masters default to values well above them — a conservative controller may insert microseconds where the device needs tens of nanoseconds. Reading the actual requirement and configuring to it, with margin, is free throughput on small transfers. But it cannot go below the device's number, and on an FPGA slave part of that number is your own synchroniser latency §12.

The gap is often the largest single item and the easiest to overlook. Chapter 7.3 §7 showed why it must be enforced in hardware; §3 above shows it can exceed the CS lead and lag combined. A controller that enforces a conservative gap in hardware is safe and may be leaving half the small-transfer throughput unclaimed.

Measure before optimising the clock. The table in §4 makes the case quantitatively: above the knee, clock rate buys very little. Knowing where your design sits on that table requires the frame timer, and it is the difference between a justified decision and a guess.

Instrument the gap separately from the frame. It is tempting to fold the gap into "overhead", but it has a different cause and a different fix — the lead and lag are per-device setup, while the gap is per-device recovery and is only paid when transactions are back to back. Measuring them separately tells you which to attack.

9. Failure Signature — Throughput That Improves Far Less Than the Clock

Symptom. A design doubles its SPI clock from 25 MHz to 50 MHz expecting roughly twice the data rate. Measured throughput improves by about 20%. Everything is functionally correct, and doubling again to 100 MHz improves it by a further 10%.

What the diminishing pattern establishes. Each doubling helps less than the last, which is the exact signature of a fixed cost becoming dominant. If the bottleneck scaled with the clock, every doubling would give the same proportional gain; instead the gains are shrinking toward a limit, so something in the transaction does not depend on the clock at all.

Plausible mechanisms.

  • Non-clocked overhead dominating — lead, lag and gap are a large share of each transaction, and none of them scales. This is §4's table being lived.
  • Software issue rate, so the bus is idle between transactions regardless of clock (Chapter 9.1 §9).
  • A rate limit clamping the effective divisor below what was configured (Chapter 9.4).
  • Per-transaction device latency, such as a programming or conversion time, which is fixed in microseconds.

The discriminating observation. Measure the frame breakdown at both clock rates. If active_cycles halves while lead + lag + gap stays constant, the diagnosis is complete and §4's table predicts exactly how much further clocking will help — which is usually "not much".

If active_cycles does not halve, the clock is not actually doubling: check the effective divisor against the requested one, because a per-slave rate limit will silently clamp it.

And if both scale properly but throughput does not, the bus is idle and the problem is upstream.

Why the investigation goes wrong. Because a 20% gain from a 2× clock increase reads as something being broken, and the search goes to timing margins and signal integrity. Nothing is broken — the arithmetic simply says that when half the transaction is fixed in nanoseconds, doubling the other half cannot do better than 1.33×, and with framing dominant it does considerably worse.

10. Common Misconceptions

11. Reason It Through

Work this before reading the answer.

A design polls a status register once per millisecond: a 1-byte command and a 1-byte response, with a 200 ns CS lead, a 50 ns lag and a 500 ns enforced inter-frame gap. The bus runs at 25 MHz.

An engineer proposes moving to 100 MHz to "reduce the polling overhead". How much does that actually save, and what would save more?

Compute the transaction at 25 MHz. T_sclk = 40 ns, and the poll is 16 clocked bits.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   clocked      16 × 40 ns   =  640 ns
   CS lead                   =  200 ns
   CS lag                    =   50 ns
   gap                       =  500 ns
   ─────────────────────────────────────
   total per poll            = 1390 ns

Now at 100 MHz. T_sclk = 10 ns:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   clocked      16 × 10 ns   =  160 ns
   non-clocked  unchanged    =  750 ns
   ─────────────────────────────────────
   total per poll            =  910 ns

The saving is 480 ns per poll — 35%, for a 4× clock increase. And the reason is stark: at 100 MHz, 750 of 910 ns (82%) is non-clocked. The clock now governs less than a fifth of the transaction, so even an infinitely fast bus could not get below 750 ns.

What would save more? Three things, in increasing order of effect.

Reduce the gap. At 500 ns it is the largest single item — bigger than the entire clocked portion at 100 MHz. If the device's actual requirement is 100 ns and the controller is enforcing 500 conservatively, that is 400 ns saved for a configuration change, comparable to the entire benefit of quadrupling the clock.

Reduce the CS lead. Same argument at 200 ns, and the same caution: it has a datasheet minimum that must be respected.

Stop polling. At one poll per millisecond, 1390 ns is 0.14% of the time — the polling is not a throughput problem at all. If the concern is CPU or power rather than bandwidth, a data-ready interrupt removes 100% of it, which no clock change can approach.

And the question worth asking before any of this. At 0.14% bus utilisation, what problem is being solved? If the answer is "the bus looks slow", the measurement in Chapter 9.1 §9 would have shown 99.86% idle and redirected the effort. Optimising a resource that is idle almost all the time is the most common wasted performance work there is.

The general lesson. When a transaction is dominated by fixed costs, clock rate is nearly irrelevant, and the fixed costs are usually configuration values chosen conservatively rather than physical limits. Read the datasheet minima, compare them against what the controller is actually inserting, and the saving is often larger than anything the clock can offer — and free.

12. Understanding Check

13. Summary

Overhead divides into two kinds with different behaviour. Clocked overhead — command, address, dummy — occupies SCLK edges and shrinks as the clock rises. Non-clocked overhead — CS lead, CS lag, inter-frame gap — carries no bits and is fixed in nanoseconds, because it is set by device requirements.

For a typical read at 50 MHz the budget is 800 ns clocked and 220 ns non-clocked, totalling about a microsecond of overhead per access.

Raising the clock attacks only the first. At 200 MHz more than half the overhead is time in which nothing is clocked, which is why there is a knee beyond which clock rate stops helping — and it lands, not coincidentally, near where SPI devices stop getting faster.

Reducing the number of accesses attacks both kinds at once, which is the quantitative case for the merging advice of Module 7.

A bit counter cannot see the non-clocked half. Measuring it needs a frame timer that splits a frame into lead, active and lag and separately records the preceding gap — with the defining property that the buckets partition the frame, so the breakdown is a decomposition rather than three estimates.

Two implementation subtleties carry the design: a bitless cycle is only provisionally lag and folds back into active if another bit follows, and the gap accumulator must survive the frame opening or every gap reads as zero.

And when a clock increase disappoints, the diminishing pattern itself is the diagnosis — with the fixed costs usually turning out to be conservative configuration, not physics.

14. What Comes Next

The overhead is per access, so the obvious remedy is fewer, larger accesses. Chapter 9.3 — Burst Efficiency and Transfer Sizing makes that precise: exactly how efficiency rises with burst length, where the curve flattens and why pushing past that point buys nothing, what limits burst length in practice, and the hardware that splits a transfer into legal bursts — in all three HDLs.

Continue learning