SPI · Module 9
Raw vs Effective Payload Throughput
The three throughput numbers a link has — raw bit rate, protocol rate and effective payload rate — why they differ by four times on an ordinary flash read, and the performance counter that measures the gap instead of estimating it.
Modules 4 through 8 established what SPI does. This module asks what it costs, and it begins with the number everyone quotes and almost nobody checks.
A datasheet says 50 MHz. An engineer says "so, about 6 MB/s". How wrong is that, and what does the application actually get?
Usually wrong by a factor of two to six, and occasionally by a factor of thirty — and the size of the error depends entirely on how the bus is used rather than on how fast it runs.
1. Three Different Numbers
The confusion comes from one word — throughput — being used for three quantities that can differ by an order of magnitude.
Raw bit rate. The clock frequency. At 50 MHz the bus moves 50 million bits per second, counting every bit including commands, addresses and dummy cycles. This is what a datasheet advertises and it is a true statement about the wire.
Protocol throughput. Payload bits divided by clocked bits. It excludes the overhead bits but still assumes the clock never stops, so it measures how much of the traffic was useful.
Effective throughput. Payload bits divided by elapsed time. This includes the clock being stopped: CS lead and lag (Chapter 2.5), inter-frame gaps (Chapter 7.3), and any interval where the bus sat idle because software had nothing ready.
2. Where the Bits Go
Seven byte times, two of them useful
8 cycles3. The Arithmetic
Take the read above at 50 MHz: a 1-byte opcode, a 3-byte address, 8 dummy cycles and 2 payload bytes.
clocked bits = 8 + 24 + 8 + 16 = 56
payload bits = 16
raw bit rate = 50 Mbit/s = 6.25 MB/s
protocol throughput = 50 × (16/56) = 14.3 Mbit/s = 1.79 MB/sNow add framing. With a 100 ns CS lead, a 20 ns lag and a 100 ns inter-frame gap:
clocking time = 56 bits × 20 ns = 1120 ns
framing time = 100 + 20 + 100 = 220 ns
────────────────────────────────────────────────
per transaction = 1340 ns
effective throughput = 16 bits / 1340 ns = 11.9 Mbit/s = 1.49 MB/sRaw 6.25 MB/s, effective 1.49 MB/s — a factor of 4.2. And nothing is wrong: no margin is being violated, no device is misbehaving, and the bus is running at exactly its advertised speed.
4. Why the Quoted Number Is the Least Useful
Raw bit rate is a property of the clock. Effective throughput is a property of the workload. They coincide only when the payload dominates, which for a bus used the way SPI usually is — small register accesses — it never does.
Two consequences follow, and they are the practical content of the whole module.
Doubling the clock does not double the throughput of small transfers. Clocking time halves, but framing time does not: CS lead and lag are fixed in nanoseconds, set by device requirements rather than by the clock. In the example above, doubling to 100 MHz gives:
clocking = 560 ns (halved)
framing = 220 ns (unchanged)
total = 780 ns (not 670)
effective = 16 bits / 780 ns = 20.5 Mbit/sThat is 1.72× for a 2× clock increase — and the ratio gets worse the smaller the payload, because framing becomes a larger share of what remains.
Comparing two designs by clock rate is meaningless. A 20 MHz bus doing 256-byte bursts beats a 50 MHz bus doing single-byte reads by a wide margin, and no amount of staring at the datasheets reveals that.
5. Measuring Rather Than Estimating
Everything above is arithmetic on assumed numbers. On a real system the assumptions are usually wrong — software adds gaps the designer did not model, transactions are shorter than intended, and retries inflate the count.
That is what the counter in §6 is for, and the three quantities it separates are exactly the three of §1:
clocked_bitsgives the raw utilisation.payload_bits / clocked_bitsgives protocol throughput.payload_bits / (busy_cycles + idle_cycles)gives effective throughput.
The third requires counting idle time as well as busy time, which is why the counter buckets every cycle into one or the other. Without that, a measurement can only report how efficiently the bus ran while it was running — which is the question nobody is asking.
6. Building the Performance Counter — Three HDLs
The circuit
Circuit. Five counters and an edge detector.
State. Busy cycles, idle cycles, clocked bits, payload bits, frame count, and a registered copy of the frame-active level.
Datapath. None — this observes.
Control. Every cycle increments exactly one of the two time buckets, so busy + idle is the elapsed window by construction. That is what makes the wall-clock ratio trustworthy rather than an estimate.
Clock and reset. System clock; asynchronous active-low reset clearing everything.
Enables. clear restarts the window. A ratio computed across a clear would mix two workloads, so the counters must all zero together.
Timing. Purely observational; nothing here is in a critical path.
Synthesis. Five CNT_W counters. At 32 bits that is 160 flip-flops, which is real but small, and the width matters: a 32-bit counter at 100 MHz wraps in 43 seconds, so a long measurement needs either wider counters or software that reads and accumulates.
Limitations. It needs a payload_bit qualifier from the protocol layer — the counter cannot know which bits are payload, because nothing on the bus says so. That signal comes from the sequencer's phase (Chapter 4.4).
// spi_perf_counter.sv — a hardware performance monitor for an SPI link.
//
// Throughput arguments are usually made on paper and are usually wrong,
// because the paper version counts the transaction someone intended rather
// than the traffic that actually occurred. This counts the traffic.
//
// It separates three quantities that are routinely conflated:
//
// clocked_bits every bit the bus carried, useful or not
// payload_bits only the bits that were data
// busy_cycles system cycles with a frame open, INCLUDING the CS lead
// and lag in which no bit moves at all
//
// From those, two different efficiencies fall out -- and they answer
// different questions (Chapter 9.1 §3):
//
// payload_bits / clocked_bits -> protocol efficiency
// payload_bits / elapsed -> what the application actually gets
module spi_perf_counter #(
parameter int CNT_W = 32
) (
input logic clk,
input logic rst_n,
input logic clear, // pulse: restart the measurement window
input logic cs_active, // level: a frame is open
input logic bit_stb, // one pulse per SPI bit
input logic payload_bit, // level: this bit is payload
output logic [CNT_W-1:0] busy_cycles, // cycles with a frame open
output logic [CNT_W-1:0] idle_cycles, // cycles with no frame open
output logic [CNT_W-1:0] clocked_bits,
output logic [CNT_W-1:0] payload_bits,
output logic [CNT_W-1:0] frames
);
logic cs_active_q;
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
busy_cycles <= '0;
idle_cycles <= '0;
clocked_bits <= '0;
payload_bits <= '0;
frames <= '0;
cs_active_q <= 1'b0;
end else if (clear) begin
// A measurement window must start from zero, or a ratio computed
// over it mixes two workloads.
busy_cycles <= '0;
idle_cycles <= '0;
clocked_bits <= '0;
payload_bits <= '0;
frames <= '0;
cs_active_q <= cs_active;
end else begin
cs_active_q <= cs_active;
// Time accounting. Every cycle lands in exactly one bucket, so
// busy + idle is the elapsed window by construction -- which is
// what makes the wall-clock ratio trustworthy.
if (cs_active) busy_cycles <= busy_cycles + 1'b1;
else idle_cycles <= idle_cycles + 1'b1;
// Bit accounting, only while a frame is open.
if (cs_active && bit_stb) begin
clocked_bits <= clocked_bits + 1'b1;
if (payload_bit) payload_bits <= payload_bits + 1'b1;
end
// A frame begins on the rising edge of cs_active.
if (cs_active && !cs_active_q) frames <= frames + 1'b1;
end
end
endmodule// spi_perf_counter_tb.sv — count a known workload and check the ratios the
// chapter computes by hand.
`timescale 1ns/1ps
module spi_perf_counter_tb;
logic clk = 0, rst_n = 0;
always #5 clk = ~clk;
localparam int CW = 32;
logic clear = 0, cs_active = 0, bit_stb = 0, payload_bit = 0;
logic [CW-1:0] busy_cycles, idle_cycles, clocked_bits, payload_bits, frames;
spi_perf_counter #(.CNT_W(CW)) dut (
.clk, .rst_n, .clear, .cs_active, .bit_stb, .payload_bit,
.busy_cycles, .idle_cycles, .clocked_bits, .payload_bits, .frames);
int errors = 0;
int elapsed; // assigned procedurally, never initialised statically
task automatic chk(input string what, input int g, input int e);
if (g !== e) begin $display("FAIL %s: got %0d exp %0d", what, g, e); errors++; end
endtask
// One SPI bit: two system cycles, with payload_bit already set.
task automatic send_bit(input logic is_payload);
payload_bit = is_payload;
@(negedge clk); bit_stb = 1;
@(negedge clk); bit_stb = 0;
endtask
// A transaction: `lead` idle-in-frame cycles, then cmd/addr/dummy bits,
// then payload bits, then `lag` idle-in-frame cycles.
task automatic transaction(input int lead, input int overhead_bits,
input int pay_bits, input int lag);
cs_active = 1;
repeat (lead) @(negedge clk);
for (int i = 0; i < overhead_bits; i++) send_bit(1'b0);
for (int i = 0; i < pay_bits; i++) send_bit(1'b1);
payload_bit = 0;
repeat (lag) @(negedge clk);
cs_active = 0;
endtask
task automatic gap(input int n); repeat (n) @(negedge clk); endtask
initial begin
repeat (3) @(negedge clk); rst_n = 1; @(negedge clk);
clear = 1; @(negedge clk); clear = 0; @(negedge clk);
// --- one read: 8 command + 24 address + 8 dummy = 40 overhead bits,
// then 8 payload bits. The Chapter 6.4 example, measured. ---
transaction(4, 40, 8, 4);
gap(20);
chk("one frame", frames, 1);
chk("clocked bits", clocked_bits, 48);
chk("payload bits", payload_bits, 8);
// Protocol efficiency: payload / clocked = 8/48 = 16.7%.
chk("overhead bits", clocked_bits - payload_bits, 40);
// --- a second, larger transaction: same overhead, 256 payload bits ---
clear = 1; @(negedge clk); clear = 0; @(negedge clk);
transaction(4, 40, 256, 4);
gap(20);
chk("large: clocked", clocked_bits, 296);
chk("large: payload", payload_bits, 256);
// --- two frames in one window, to show frames counts correctly ---
clear = 1; @(negedge clk); clear = 0; @(negedge clk);
transaction(4, 8, 8, 4);
gap(10);
transaction(4, 8, 8, 4);
gap(10);
chk("two frames", frames, 2);
chk("two frames: bits", clocked_bits, 32);
chk("two frames: payload", payload_bits, 16);
// --- the time buckets must partition the window exactly ---
begin
elapsed = busy_cycles + idle_cycles;
if (elapsed == 0) begin
$display("FAIL: no time accounted"); errors++;
end
// Busy must exceed the clocked time, because the CS lead and lag
// are in-frame cycles during which no bit moves -- that gap is
// exactly the overhead the chapter is about.
if (busy_cycles <= clocked_bits) begin
$display("FAIL: busy (%0d) should exceed clocked bits (%0d)",
busy_cycles, clocked_bits);
errors++;
end
$display(" window: busy=%0d idle=%0d elapsed=%0d clocked=%0d payload=%0d",
busy_cycles, idle_cycles, elapsed, clocked_bits, payload_bits);
end
// --- clear must reset every counter, or a second measurement window
// silently includes the first ---
clear = 1; @(negedge clk); clear = 0; @(negedge clk);
chk("clear: frames", frames, 0);
chk("clear: clocked", clocked_bits, 0);
chk("clear: payload", payload_bits, 0);
chk("clear: busy", busy_cycles, 0);
if (errors == 0)
$display("PASS: payload, overhead and frame counts match the workload, busy time exceeds clocked time by the CS lead and lag, and clear restarts the window");
else
$display("FAILED with %0d error(s)", errors);
$finish;
end
initial begin #900000; $display("FAIL: watchdog timeout"); $finish; end
endmoduleThe check worth reading is busy_cycles > clocked_bits. Those two count different things — cycles with a frame open versus bits carried — and the gap between them is the CS lead and lag. A design where they were equal would be one with no framing overhead at all, which is not achievable, so the inequality is a sanity check on the measurement itself.
// spi_perf_counter.v — the same performance monitor in Verilog-2001.
module spi_perf_counter #(
parameter CNT_W = 32
) (
input wire clk,
input wire rst_n,
input wire clear,
input wire cs_active,
input wire bit_stb,
input wire payload_bit,
output reg [CNT_W-1:0] busy_cycles,
output reg [CNT_W-1:0] idle_cycles,
output reg [CNT_W-1:0] clocked_bits,
output reg [CNT_W-1:0] payload_bits,
output reg [CNT_W-1:0] frames
);
reg cs_active_q;
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
busy_cycles <= {CNT_W{1'b0}};
idle_cycles <= {CNT_W{1'b0}};
clocked_bits <= {CNT_W{1'b0}};
payload_bits <= {CNT_W{1'b0}};
frames <= {CNT_W{1'b0}};
cs_active_q <= 1'b0;
end else if (clear) begin
busy_cycles <= {CNT_W{1'b0}};
idle_cycles <= {CNT_W{1'b0}};
clocked_bits <= {CNT_W{1'b0}};
payload_bits <= {CNT_W{1'b0}};
frames <= {CNT_W{1'b0}};
cs_active_q <= cs_active;
end else begin
cs_active_q <= cs_active;
// Every cycle lands in exactly one bucket, so busy + idle is the
// elapsed window by construction.
if (cs_active) busy_cycles <= busy_cycles + 1'b1;
else idle_cycles <= idle_cycles + 1'b1;
if (cs_active && bit_stb) begin
clocked_bits <= clocked_bits + 1'b1;
if (payload_bit) payload_bits <= payload_bits + 1'b1;
end
if (cs_active && !cs_active_q) frames <= frames + 1'b1;
end
end
endmodule// spi_perf_counter_tb.v — the same checks in Verilog-2001.
`timescale 1ns/1ps
module spi_perf_counter_tb;
reg clk = 0, rst_n = 0;
always #5 clk = ~clk;
parameter CW = 32;
reg clear = 0, cs_active = 0, bit_stb = 0, payload_bit = 0;
wire [CW-1:0] busy_cycles, idle_cycles, clocked_bits, payload_bits, frames;
spi_perf_counter #(.CNT_W(CW)) dut (
.clk(clk), .rst_n(rst_n), .clear(clear), .cs_active(cs_active),
.bit_stb(bit_stb), .payload_bit(payload_bit),
.busy_cycles(busy_cycles), .idle_cycles(idle_cycles),
.clocked_bits(clocked_bits), .payload_bits(payload_bits), .frames(frames));
integer errors = 0, elapsed = 0, i;
task chk;
input [80*8-1:0] what;
input [31:0] g, e;
begin
if (g !== e) begin
$display("FAIL %0s: got %0d exp %0d", what, g, e);
errors = errors + 1;
end
end
endtask
task send_bit;
input is_payload;
begin
payload_bit = is_payload;
@(negedge clk); bit_stb = 1;
@(negedge clk); bit_stb = 0;
end
endtask
task transaction;
input integer lead, overhead_bits, pay_bits, lag;
integer k;
begin
cs_active = 1;
repeat (lead) @(negedge clk);
for (k = 0; k < overhead_bits; k = k + 1) send_bit(1'b0);
for (k = 0; k < pay_bits; k = k + 1) send_bit(1'b1);
payload_bit = 0;
repeat (lag) @(negedge clk);
cs_active = 0;
end
endtask
task gap; input integer n; begin repeat (n) @(negedge clk); end endtask
initial begin
repeat (3) @(negedge clk); rst_n = 1; @(negedge clk);
clear = 1; @(negedge clk); clear = 0; @(negedge clk);
transaction(4, 40, 8, 4);
gap(20);
chk("one frame", frames, 1);
chk("clocked bits", clocked_bits, 48);
chk("payload bits", payload_bits, 8);
chk("overhead bits", clocked_bits - payload_bits, 40);
clear = 1; @(negedge clk); clear = 0; @(negedge clk);
transaction(4, 40, 256, 4);
gap(20);
chk("large: clocked", clocked_bits, 296);
chk("large: payload", payload_bits, 256);
clear = 1; @(negedge clk); clear = 0; @(negedge clk);
transaction(4, 8, 8, 4);
gap(10);
transaction(4, 8, 8, 4);
gap(10);
chk("two frames", frames, 2);
chk("two frames: bits", clocked_bits, 32);
chk("two frames: payload", payload_bits, 16);
elapsed = busy_cycles + idle_cycles;
if (elapsed == 0) begin
$display("FAIL: no time accounted"); errors = errors + 1;
end
if (busy_cycles <= clocked_bits) begin
$display("FAIL: busy (%0d) should exceed clocked bits (%0d)",
busy_cycles, clocked_bits);
errors = errors + 1;
end
$display(" window: busy=%0d idle=%0d elapsed=%0d clocked=%0d payload=%0d",
busy_cycles, idle_cycles, elapsed, clocked_bits, payload_bits);
clear = 1; @(negedge clk); clear = 0; @(negedge clk);
chk("clear: frames", frames, 0);
chk("clear: clocked", clocked_bits, 0);
chk("clear: payload", payload_bits, 0);
chk("clear: busy", busy_cycles, 0);
if (errors == 0)
$display("PASS: payload, overhead and frame counts match the workload, busy time exceeds clocked time by the CS lead and lag, and clear restarts the window");
else
$display("FAILED with %0d error(s)", errors);
$finish;
end
initial begin #900000; $display("FAIL: watchdog timeout"); $finish; end
endmodule-- spi_perf_counter.vhd — the same performance monitor in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_perf_counter is
generic (
CNT_W : positive := 32
);
port (
clk : in std_logic;
rst_n : in std_logic;
clear : in std_logic; -- pulse
cs_active : in std_logic; -- level
bit_stb : in std_logic; -- one pulse per SPI bit
payload_bit : in std_logic; -- level
busy_cycles : out unsigned(CNT_W - 1 downto 0);
idle_cycles : out unsigned(CNT_W - 1 downto 0);
clocked_bits : out unsigned(CNT_W - 1 downto 0);
payload_bits : out unsigned(CNT_W - 1 downto 0);
frames : out unsigned(CNT_W - 1 downto 0)
);
end entity spi_perf_counter;
architecture rtl of spi_perf_counter is
signal busy_r : unsigned(CNT_W - 1 downto 0);
signal idle_r : unsigned(CNT_W - 1 downto 0);
signal clocked_r : unsigned(CNT_W - 1 downto 0);
signal payload_r : unsigned(CNT_W - 1 downto 0);
signal frames_r : unsigned(CNT_W - 1 downto 0);
signal cs_q : std_logic;
begin
busy_cycles <= busy_r;
idle_cycles <= idle_r;
clocked_bits <= clocked_r;
payload_bits <= payload_r;
frames <= frames_r;
process (clk, rst_n) is
begin
if rst_n = '0' then
busy_r <= (others => '0');
idle_r <= (others => '0');
clocked_r <= (others => '0');
payload_r <= (others => '0');
frames_r <= (others => '0');
cs_q <= '0';
elsif rising_edge(clk) then
if clear = '1' then
-- A measurement window must start from zero, or a ratio
-- computed over it mixes two workloads.
busy_r <= (others => '0');
idle_r <= (others => '0');
clocked_r <= (others => '0');
payload_r <= (others => '0');
frames_r <= (others => '0');
cs_q <= cs_active;
else
cs_q <= cs_active;
-- Every cycle lands in exactly one bucket, so busy + idle is
-- the elapsed window by construction.
if cs_active = '1' then
busy_r <= busy_r + 1;
else
idle_r <= idle_r + 1;
end if;
if cs_active = '1' and bit_stb = '1' then
clocked_r <= clocked_r + 1;
if payload_bit = '1' then
payload_r <= payload_r + 1;
end if;
end if;
if cs_active = '1' and cs_q = '0' then
frames_r <= frames_r + 1;
end if;
end if;
end if;
end process;
end architecture rtl;-- spi_perf_counter_tb.vhd — the same checks in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_perf_counter_tb is
end entity spi_perf_counter_tb;
architecture tb of spi_perf_counter_tb is
constant CW : positive := 32;
signal clk : std_logic := '0';
signal rst_n : std_logic := '0';
signal clear : std_logic := '0';
signal cs_active : std_logic := '0';
signal bit_stb : std_logic := '0';
signal payload_bit : std_logic := '0';
signal halt : boolean := false;
signal busy_cycles : unsigned(CW - 1 downto 0);
signal idle_cycles : unsigned(CW - 1 downto 0);
signal clocked_bits : unsigned(CW - 1 downto 0);
signal payload_bits : unsigned(CW - 1 downto 0);
signal frames : unsigned(CW - 1 downto 0);
signal errors : natural := 0;
begin
clk <= not clk after 5 ns when not halt else '0';
dut : entity work.spi_perf_counter
generic map (CNT_W => CW)
port map (clk => clk, rst_n => rst_n, clear => clear, cs_active => cs_active,
bit_stb => bit_stb, payload_bit => payload_bit,
busy_cycles => busy_cycles, idle_cycles => idle_cycles,
clocked_bits => clocked_bits, payload_bits => payload_bits,
frames => frames);
stim : process is
variable elapsed : natural;
procedure chk_n (what : string; g, e : natural) is
begin
if g /= e then
report "FAIL " & what & ": got " & integer'image(g)
& " exp " & integer'image(e) severity error;
errors <= errors + 1;
end if;
end procedure;
procedure send_bit (is_payload : std_logic) is
begin
payload_bit <= is_payload;
wait until falling_edge(clk); bit_stb <= '1';
wait until falling_edge(clk); bit_stb <= '0';
end procedure;
procedure transaction (lead, overhead_bits, pay_bits, lag : natural) is
begin
cs_active <= '1';
for i in 1 to lead loop wait until falling_edge(clk); end loop;
for i in 1 to overhead_bits loop send_bit('0'); end loop;
for i in 1 to pay_bits loop send_bit('1'); end loop;
payload_bit <= '0';
for i in 1 to lag loop wait until falling_edge(clk); end loop;
cs_active <= '0';
end procedure;
procedure gap (n : natural) is
begin
for i in 1 to n loop wait until falling_edge(clk); end loop;
end procedure;
procedure do_clear is
begin
clear <= '1'; wait until falling_edge(clk);
clear <= '0'; wait until falling_edge(clk);
end procedure;
begin
for i in 0 to 2 loop wait until falling_edge(clk); end loop;
rst_n <= '1';
wait until falling_edge(clk);
do_clear;
transaction(4, 40, 8, 4);
gap(20);
chk_n("one frame", to_integer(frames), 1);
chk_n("clocked bits", to_integer(clocked_bits), 48);
chk_n("payload bits", to_integer(payload_bits), 8);
chk_n("overhead bits", to_integer(clocked_bits) - to_integer(payload_bits), 40);
do_clear;
transaction(4, 40, 256, 4);
gap(20);
chk_n("large: clocked", to_integer(clocked_bits), 296);
chk_n("large: payload", to_integer(payload_bits), 256);
do_clear;
transaction(4, 8, 8, 4);
gap(10);
transaction(4, 8, 8, 4);
gap(10);
chk_n("two frames", to_integer(frames), 2);
chk_n("two frames: bits", to_integer(clocked_bits), 32);
chk_n("two frames: payload", to_integer(payload_bits), 16);
elapsed := to_integer(busy_cycles) + to_integer(idle_cycles);
if elapsed = 0 then
report "FAIL: no time accounted" severity error;
errors <= errors + 1;
end if;
if to_integer(busy_cycles) <= to_integer(clocked_bits) then
report "FAIL: busy should exceed clocked bits" severity error;
errors <= errors + 1;
end if;
report " window: busy=" & integer'image(to_integer(busy_cycles))
& " idle=" & integer'image(to_integer(idle_cycles))
& " elapsed=" & integer'image(elapsed)
& " clocked=" & integer'image(to_integer(clocked_bits))
& " payload=" & integer'image(to_integer(payload_bits)) severity note;
do_clear;
chk_n("clear: frames", to_integer(frames), 0);
chk_n("clear: clocked", to_integer(clocked_bits), 0);
chk_n("clear: payload", to_integer(payload_bits), 0);
chk_n("clear: busy", to_integer(busy_cycles), 0);
if errors = 0 then
report "PASS: payload, overhead and frame counts match the workload, busy "
& "time exceeds clocked time by the CS lead and lag, and clear "
& "restarts the window" severity note;
else
report "FAILED with " & integer'image(errors) & " error(s)" severity error;
end if;
halt <= true;
wait;
end process;
watchdog : process is
begin
wait for 900 us;
if not halt then
report "FAIL: watchdog timeout" severity failure;
end if;
wait;
end process;
end architecture tb;Parity
All three implement the same counters: identical ports and generics, asynchronous active-low reset, a clear that zeros every counter together, exactly one time bucket incremented per cycle, bit counting qualified by the frame being open, and a frame count from the rising edge of cs_active. All three testbenches drive the same workload and report the same numbers — busy=80, idle=21, clocked=32, payload=16 for the two-frame window.
7. Why a Verification Engineer Cares
// 1. THE property that makes the wall-clock ratio meaningful: every cycle
// lands in exactly one time bucket, so busy + idle is the window.
a_time_partitions : assert property (
@(posedge clk) disable iff (!rst_n || clear)
(busy_cycles + idle_cycles) == ($past(busy_cycles) + $past(idle_cycles) + 1))
else $error("a cycle was counted twice or not at all");
// 2. Payload can never exceed clocked -- a containment property that
// catches a payload qualifier asserted outside a frame.
a_payload_subset : assert property (
@(posedge clk) disable iff (!rst_n) payload_bits <= clocked_bits)
else $error("more payload bits than clocked bits");
// 3. No bit is counted while the bus is idle.
a_no_bits_when_idle : assert property (
@(posedge clk) disable iff (!rst_n)
(!cs_active && bit_stb) |=> $stable(clocked_bits))
else $error("counted a bit outside a frame");
// 4. clear zeros everything. A partial clear silently blends windows.
a_clear_is_total : assert property (
@(posedge clk) disable iff (!rst_n)
clear |=> (busy_cycles == 0 && idle_cycles == 0 &&
clocked_bits == 0 && payload_bits == 0 && frames == 0))
else $error("clear left a counter non-zero");Property 1 is the one that matters, and it is unusual: it asserts an accounting identity rather than a behaviour. A counter that occasionally double-counts or drops a cycle still looks plausible — the numbers are in the right range — and produces a ratio that is quietly wrong. Only an identity catches it.
What these prove. That the counters are self-consistent. What they cannot prove is that payload_bit is asserted on the right bits — that comes from the protocol layer, and a measurement built on a wrong qualifier is confidently wrong.
Coverage should target the workloads whose ratios differ:
covergroup spi_perf_cg @(posedge frame_end);
cp_payload_share : coverpoint (payload_bits * 100 / clocked_bits) {
bins tiny = {[0:20]}; // single-register access
bins low = {[21:50]};
bins good = {[51:90]};
bins dominant = {[91:100]}; // long burst
}
// Idle share of the window. A benchmark with no idle time measures
// the bus; a real workload measures the system.
cp_idle_share : coverpoint (idle_cycles * 100 / (busy_cycles + idle_cycles)) {
bins saturated = {[0:10]};
bins moderate = {[11:70]};
bins mostly_idle = {[71:100]}; // a polled sensor
}
x_payload_idle : cross cp_payload_share, cp_idle_share;
endgroupThe cross is the point. A design can have excellent protocol efficiency and dreadful effective throughput — long bursts issued rarely — and the two coverpoints separately would both look healthy.
8. Why an FPGA or ASIC Engineer Cares
Size the counters for the measurement window. A 32-bit cycle counter at 100 MHz wraps in about 43 seconds. Measuring over minutes needs 40+ bits, or software that reads and accumulates before each wrap — and a counter that wraps silently produces a ratio that is not merely wrong but arbitrary.
The payload qualifier must come from the sequencer, not be inferred. The counter cannot know which bits are payload. Wiring payload_bit to the phase decode of Chapter 4.4's sequencer is the only correct source, and it is worth a comment in the design, because the signal looks like something that could be derived locally and cannot.
Read the counters atomically. Five counters read over five bus cycles are a skewed snapshot — the ratio computed from them may correspond to no instant that ever existed. A shadow-register capture triggered by a single read, or a clear-and-read discipline, avoids it.
Instrument early. Performance counters cost almost nothing and are the difference between a throughput discussion based on measurement and one based on assertion. Adding them after a performance problem appears means the first measurement happens under pressure.
9. Failure Signature — A Bus That Is Fast and a System That Is Slow
Symptom. A design must move a fixed amount of data per second and does not. The SPI bus is configured at its maximum rate, every transaction is verified correct, and a scope shows clean signalling. Raising the clock further is not possible, and the team concludes SPI is inadequate for the application.
What "every transaction is correct" establishes. This is not a functional problem, so the entire protocol and electrical hypothesis space is empty. The question is purely one of accounting: the bus is delivering fewer payload bits per second than required, and the reason must be visible in how the time is spent.
Plausible mechanisms.
- Transaction size. Many small transfers, each paying full framing and request overhead (Chapter 6.4 §4). This is by far the most common.
- Idle time. The bus is mostly not transferring at all, because software is slow to issue the next transaction. Protocol efficiency is fine and effective throughput is terrible.
- A rate limit. The bus is running at the slowest device's maximum rather than at each device's own (Chapter 9.4).
- Retries or polling. Traffic that is correct but unnecessary, inflating the clocked-bit count without adding payload.
The discriminating observation, and it is what the counter exists for. Measure payload_bits, clocked_bits and idle_cycles over a representative second.
- Low payload/clocked with low idle → the bus is busy doing overhead. Fix transaction size.
- High payload/clocked with high idle → the bus is starved. Fix the software or the DMA.
- Both poor → both, and the transaction-size fix usually helps more.
That single measurement partitions the problem in a way no amount of reasoning about the datasheet can, and it points at a different fix in each case.
Why the investigation goes wrong. Because "the bus is at maximum clock" is treated as proof the bus is saturated. It proves the clock is at maximum. Whether the bus is carrying payload is a separate question, and on a system doing small transfers the answer is usually that it is idle or doing overhead for most of every second.
10. Common Misconceptions
11. Reason It Through
Work this before reading the answer.
A system reads a 12-byte sensor record every millisecond over a 20 MHz SPI bus. Each read is a 1-byte command, a 2-byte address, 4 dummy cycles and 12 payload bytes. Measured over one second, the performance counter reports:
clocked_bits = 2,200,000,payload_bits = 960,000,busy_cycles = 6,000,000,idle_cycles = 94,000,000,frames = 10,000.The system clock is 100 MHz. What is going on?
Start with the frame count, because it contradicts the specification. One read per millisecond for one second is 1,000 frames. The counter says 10,000. The system is issuing ten times as many transactions as the design intends.
Check that against the payload. 12 bytes × 1,000 reads = 96,000 bits expected. The counter says 960,000 — exactly ten times more, confirming the frame count rather than contradicting it. So this is not a miscount; the system really is doing ten reads per millisecond.
Why ten? Almost certainly polling: the driver reads the record, finds a status or sequence field unchanged, and reads again. Ten reads per sample means nine are wasted — and all nine are correct transactions, which is why nothing is failing.
Now the efficiency numbers, which look surprisingly healthy.
protocol throughput = 960,000 / 2,200,000 = 43.6 %That is respectable, because a 12-byte payload against 4 bytes of request is a decent ratio (Chapter 9.3 shows why). The protocol is being used well. That is exactly what makes this case instructive: the obvious efficiency metric is fine, and the system is still wasting 90% of its work.
The time numbers tell the real story.
busy fraction = 6,000,000 / 100,000,000 = 6 %
effective throughput = 960,000 bits / 1 s = 960 kbit/s = 120 kB/sThe bus is idle 94% of the time and still doing ten times the necessary work. Both facts are true simultaneously, which is the part that confuses people: there is enormous headroom and enormous waste.
So what should change? Not the clock, and not the transaction size — both are fine. The fix is to stop issuing nine unnecessary reads. If the sensor has a data-ready interrupt, use it. If not, poll a one-byte status register instead of re-reading the whole record: that reduces the wasted traffic by a factor of twelve at a stroke, and Chapter 9.2 shows what the remaining polls then cost.
And note what a naive optimisation would have done. Seeing 43.6% protocol efficiency, an engineer might lengthen the bursts to improve it — reading 24 bytes instead of 12. That would raise the efficiency percentage and double the wasted bandwidth, because the problem was never efficiency per transaction.
The general lesson. Efficiency ratios measure how well each transaction is constructed. They say nothing about whether the transaction should have happened. A system can score well on every efficiency metric while doing ten times the necessary work, and only the absolute counts — frames and payload bits against what the specification requires — reveal it.
12. Understanding Check
13. Summary
Throughput is three numbers, not one. Raw bit rate is the clock and counts every bit. Protocol throughput is payload over clocked bits and measures how well the protocol is used. Effective throughput is payload over elapsed time, including idle, and is the only one an application experiences.
A 50 MHz bus doing a typical register read delivers roughly 1.5 MB/s of payload against a 6.25 MB/s raw rate — a factor of four, with nothing wrong.
Doubling the clock does not double throughput for small transfers, because framing time is fixed in nanoseconds while clocking time halves. And clock rate cannot rank two designs: a 20 MHz bus doing long bursts beats a 50 MHz bus doing single-byte reads.
Measurement beats estimation, because real systems add gaps and transactions the designer did not model. The counter separates the three quantities by bucketing every cycle into busy or idle — an accounting identity that makes the wall-clock ratio trustworthy, and the property worth asserting.
The payload_bit qualifier must come from the sequencer, since nothing on the bus marks payload.
And when a system is slow on a maximally-clocked bus, one measurement partitions it: busy with low payload share means fix the transaction size; idle with good payload share means fix the software. Efficiency ratios describe how well each transaction is built — only absolute counts reveal whether it needed to happen at all.
14. What Comes Next
This chapter established that overhead is large. Chapter 9.2 — Protocol Overhead takes it apart: exactly how many cycles go to the command, the address, the dummy phase and the chip-select framing, which of those scale with payload and which are fixed per access, and the frame timer that measures the non-clocking overhead the bit counters of this chapter cannot see.
Continue learning
Related tutorials
- Related topic
Dummy Phases and Read Latency
Why a device needs turnaround before it can answer, why dummy is counted in clock cycles rather than bytes, how its length grows with frequency, and the one-byte data offset a mismatch produces.
- Related topic
Back-to-Back Transactions and Inter-Frame Gap
How soon chip select may fall again after it rises: the minimum deselect time, why a device needs it, what the gap costs in throughput, and the hardware enforcement that keeps software from violating it.
- Related topic
Protocol Overhead
Overhead split into the part SCLK pays for and the part it does not: why the non-clocked share grows from 5% to half the transaction as the clock rises, and the frame timer that partitions a transfer into gap, lead, active and lag.
- Related topic
SDR, DDR Transfers, and Throughput
Why data on both edges is independent of lane width, why the byte assembly is identical in both modes, why the dummy phase does not halve so DDR gives under two times, and why DDR parts specify more dummy cycles than their SDR modes.
