SPI · Module 5
MOSI Data Flow and CS Framing
What the master drives on MOSI through every region of a transfer, what the two chip-select edges bracket, and why an asynchronous CS must cross into the slave's clock domain before any edge is derived from it.
Chapter 5.1 treated chip select as a clean framing signal and took its rising edge as the commit event. Both of those are assumptions, and this chapter examines them.
CS is low and the clock is running. What is the master obliged to put on MOSI at each moment — and what, exactly, do the two CS edges bracket?
The answer divides a frame into three regions with different rules, and it ends somewhere unexpected: with the observation that CS is an asynchronous input to the slave, and that the edge Chapter 5.1 committed on cannot be taken directly from the pin.
1. The Three Regions of a Frame
A transfer is not simply "CS low". It has three intervals, and Chapter 2.5 named the outer two:
The CS lead — from CS asserting to the first clock edge. No data moves. The device uses this time to wake its interface, and under CPHA = 0 the master must already have the first bit on MOSI before it ends (Chapter 3.4 §3).
The active interval — the clock edges themselves. Every edge moves a bit in each direction.
The CS lag — from the last clock edge to CS deasserting. No data moves, but the interval is not empty of meaning: the device needs the last bit to be safely captured before the frame closes, and on most devices this is where the commit happens.
The rules differ in each. Conflating them produces a specific and common bug, which is §10.
2. What the Master Drives, Region by Region
Where MOSI must be valid, and where it need not be
10 cyclesRead MOSI as three held values — 1, 0, 1 — each straddling one rising edge. The x regions at either end are the honest representation of "the master may drive anything here, and the device must not care".
That last clause is a device obligation, and it is worth stating because devices occasionally get it wrong. A slave that samples MOSI outside its expected bit positions, or that treats a transition during the CS lag as meaningful, will misbehave with a perfectly correct master.
3. MOSI During a Read — the Byte That Isn't Data
Here is where full duplex stops being an abstraction and becomes a driver-level decision.
Chapter 1.4 established that every edge moves a bit in both directions. During a read's data phase the device is sending meaningful bytes on MISO — but the master is still shifting something out on MOSI, because it must clock the transfer and the clock moves both registers.
So the master has to drive something. The conventional choices:
0x00— the most common default.0xFF— preferred where the bus idles high, or where the device treats an all-ones byte as a guaranteed no-op.- Whatever was in the buffer — what you get by accident if the driver reuses a transmit buffer without clearing it.
In principle the device ignores MOSI during a read, and most do. In practice some do not, and the cases where it matters are worth knowing: devices where a byte sent during the data phase is interpreted as a subsequent command, and devices where particular values on MOSI during a read abort the operation. The datasheet will say which, and if it specifies a value, sending a different one is a real bug that will look like data corruption.
4. What the CS Edges Actually Bracket
Both edges do more than delimit, and the asymmetry between them is the useful part.
The falling edge starts everything. It opens the frame, resets the device's phase sequencer to expect a command (Chapter 4.4), and starts the CS lead interval. On a multi-slave bus it is also the point at which this device — and only this device — begins driving MISO (Chapter 1.5).
The rising edge ends everything, and is the only atomic event. It closes the frame, commits a staged write, releases MISO to high impedance, and resynchronises the sequencer. Every one of those is a state change that must happen exactly once.
That asymmetry explains a practical rule: the rising edge is the dangerous one. A spurious falling edge starts a frame that will be discarded harmlessly because no valid command follows. A spurious rising edge can commit a partial write, release the bus mid-transfer, or truncate a transaction — and because nothing acknowledges on SPI, none of that is reported.
Which is why the next section matters more than it first appears.
5. CS Is Asynchronous to the Slave
Chapter 5.1 §8 flagged this and deferred it. Here it is.
CS is driven by an external master whose clock has no defined relationship to the slave's. It can therefore change arbitrarily close to the slave's clock edge, violating setup or hold on the first flip-flop that samples it. That flop can go metastable — settling to an unpredictable value after an unpredictable delay.
The standard remedy is a two-flop synchroniser: sample the asynchronous signal, then sample the result again a cycle later, giving the first flop a full clock period to resolve. It does not eliminate metastability — nothing does — but it reduces the probability of an unresolved value propagating to a rate measured in years.
The part specific to this chapter is what you must do with the result. Every edge must be derived from the synchronised signal, never the raw pin. Taking an edge directly from an asynchronous input is worse than sampling its level, because a metastable sample that resolves the "wrong" way for one cycle produces a spurious edge — and §4 just established that a spurious rising edge commits a write.
The cost is latency: two or three clock periods between the real CS edge and the slave reacting. That is not free, and it is spent out of the CS lead interval (Chapter 2.5). A slave needing three cycles to recognise CS, clocked at 50 MHz, consumes 60 ns of lead time before it has done anything at all — which is a real constraint on how soon the master may start clocking.
6. Building the Frame Detector — Three HDLs
The circuit
Circuit. A synchroniser chain, one delay flop, and three combinational decodes.
State. SYNC_STAGES flip-flops holding the in-flight samples of CS, plus one more holding the previous synchronised value so an edge can be detected.
Datapath. None — this module carries no data. It converts an asynchronous level into two safe, single-cycle events.
Control. frame_start is the falling edge of the synchronised CS, frame_end the rising edge, and in_frame its inverted level.
Clock. The slave's system clock. Nothing here is clocked by SCLK, which is the whole point: SCLK is an external signal that stops between transfers.
Reset. Asynchronous, active-low, and — importantly — the synchroniser resets high. Deasserted is the safe idle state, so a slave leaving reset must not believe it is mid-frame. Resetting the chain to zero would manufacture a frame out of reset.
Enables. None; the chain is free-running, which is required — it must be sampling before CS arrives.
Timing. frame_start and frame_end are single-cycle strobes. Both appear SYNC_STAGES + 1 clocks after the real pin edge, and that latency is deterministic even though the resolution of the first flop is not.
Synthesis. SYNC_STAGES + 1 flip-flops and a handful of gates. The synchroniser flops should be constrained as a synchroniser so the tools do not optimise the chain away or retime across it — §9.
Limitations. No glitch filtering, which §7 treats as a feature to be understood rather than a gap to be patched.
// spi_cs_frame.sv — bringing an asynchronous chip select into the slave's
// clock domain, and deriving the frame boundaries from it.
//
// CS is driven by an external master with no relationship to this clock. Two
// consequences follow, and both are the reason this module exists:
//
// 1. Sampling it directly risks METASTABILITY, so it crosses through a
// synchroniser first.
// 2. Every edge must be derived from the SYNCHRONISED signal, never the
// raw pin -- an edge taken from a metastable sample can be spurious,
// and in a write path a spurious frame_end is a spurious commit.
//
// What this does NOT do is filter glitches. A synchroniser resolves
// metastability; it does not decide whether a short pulse was real. A CS
// glitch wider than one clock period will produce a frame boundary.
module spi_cs_frame #(
parameter int SYNC_STAGES = 2
) (
input logic clk,
input logic rst_n,
input logic cs_n_async, // straight from the pin
output logic cs_n_sync, // safe to use in logic
output logic frame_start, // one pulse: CS asserted
output logic frame_end, // one pulse: CS deasserted
output logic in_frame // level: between the two
);
// Reset HIGH: deasserted is the safe idle state, so a slave coming out of
// reset must not believe it is mid-frame.
logic [SYNC_STAGES-1:0] sync_q;
logic cs_n_d;
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
sync_q <= {SYNC_STAGES{1'b1}};
cs_n_d <= 1'b1;
end else begin
sync_q <= {sync_q[SYNC_STAGES-2:0], cs_n_async};
cs_n_d <= sync_q[SYNC_STAGES-1];
end
end
assign cs_n_sync = sync_q[SYNC_STAGES-1];
// Edges come from the synchronised signal and its delayed copy only.
assign frame_start = !cs_n_sync && cs_n_d; // falling edge of cs_n
assign frame_end = cs_n_sync && !cs_n_d; // rising edge of cs_n
assign in_frame = !cs_n_sync;
endmodule// spi_cs_frame_tb.sv — asynchronous edges, one pulse each, and the honest
// demonstration that a synchroniser is not a glitch filter.
`timescale 1ns/1ps
module spi_cs_frame_tb;
logic clk = 0, rst_n = 0;
always #5 clk = ~clk; // 10 ns period
logic cs_n_async = 1;
logic cs_n_sync, frame_start, frame_end, in_frame;
spi_cs_frame #(.SYNC_STAGES(2)) dut (
.clk, .rst_n, .cs_n_async, .cs_n_sync, .frame_start, .frame_end, .in_frame);
int errors = 0, starts = 0, ends = 0;
int s0, e0; // glitch-check baselines, ASSIGNED at use
task automatic chk(input string what, input int got, input int exp);
if (got !== exp) begin $display("FAIL %s: got %0d exp %0d", what, got, exp); errors++; end
endtask
always @(posedge clk) if (rst_n) begin
if (frame_start) starts++;
if (frame_end) ends++;
end
initial begin
repeat (3) @(negedge clk); rst_n = 1; @(negedge clk);
chk("out of reset: not in frame", in_frame, 0);
chk("out of reset: cs_n_sync high", cs_n_sync, 1);
// --- Assert CS deliberately OFF the clock edge (async stimulus) ---
#3 cs_n_async = 0;
repeat (6) @(posedge clk);
chk("frame_start pulsed once", starts, 1);
chk("in_frame asserted", in_frame, 1);
chk("no frame_end yet", ends, 0);
// Holding CS low must not produce more pulses.
repeat (10) @(posedge clk);
chk("no extra starts while held", starts, 1);
#7 cs_n_async = 1;
repeat (6) @(posedge clk);
chk("frame_end pulsed once", ends, 1);
chk("in_frame deasserted", in_frame, 0);
// --- A second frame, asserted at a different phase of the clock ---
#1 cs_n_async = 0;
repeat (6) @(posedge clk);
chk("second frame_start", starts, 2);
#9 cs_n_async = 1;
repeat (6) @(posedge clk);
chk("second frame_end", ends, 2);
// --- HONESTY CHECK: a glitch WIDER than a clock period is seen as a
// real frame. The synchroniser resolves metastability; it does not
// decide whether a pulse was intended. Asserting the observed
// behaviour keeps the documentation truthful.
begin
s0 = starts; e0 = ends;
#2 cs_n_async = 0;
#25 cs_n_async = 1; // 25 ns low = 2.5 clock periods
repeat (6) @(posedge clk);
chk("wide glitch produces a frame_start", starts - s0, 1);
chk("wide glitch produces a frame_end", ends - e0, 1);
end
if (errors == 0)
$display("PASS: CS crosses through the synchroniser, each edge yields exactly one pulse, and a glitch wider than a clock period is indistinguishable from a real frame");
else
$display("FAILED with %0d error(s)", errors);
$finish;
end
initial begin #200000; $display("FAIL: watchdog timeout"); $finish; end
endmoduleNote how that testbench drives CS: #3, #7, #1, #9 — deliberately off the clock edge and at different phases each time. A stimulus that changed CS on a clock edge would exercise the one case a synchroniser is not needed for.
// spi_cs_frame.v — the same CS synchroniser and frame detector in Verilog-2001.
module spi_cs_frame #(
parameter SYNC_STAGES = 2
) (
input wire clk,
input wire rst_n,
input wire cs_n_async,
output wire cs_n_sync,
output wire frame_start,
output wire frame_end,
output wire in_frame
);
// Reset HIGH: deasserted is the safe idle state.
reg [SYNC_STAGES-1:0] sync_q;
reg cs_n_d;
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
sync_q <= {SYNC_STAGES{1'b1}};
cs_n_d <= 1'b1;
end else begin
sync_q <= {sync_q[SYNC_STAGES-2:0], cs_n_async};
cs_n_d <= sync_q[SYNC_STAGES-1];
end
end
assign cs_n_sync = sync_q[SYNC_STAGES-1];
assign frame_start = !cs_n_sync && cs_n_d;
assign frame_end = cs_n_sync && !cs_n_d;
assign in_frame = !cs_n_sync;
endmodule// spi_cs_frame_tb.v — the same checks in Verilog-2001.
`timescale 1ns/1ps
module spi_cs_frame_tb;
reg clk = 0, rst_n = 0;
always #5 clk = ~clk;
reg cs_n_async = 1;
wire cs_n_sync, frame_start, frame_end, in_frame;
spi_cs_frame #(.SYNC_STAGES(2)) dut (
.clk(clk), .rst_n(rst_n), .cs_n_async(cs_n_async), .cs_n_sync(cs_n_sync),
.frame_start(frame_start), .frame_end(frame_end), .in_frame(in_frame));
integer errors = 0, starts = 0, ends = 0, s0 = 0, e0 = 0, i;
task chk;
input [80*8-1:0] what;
input [31:0] got, exp;
begin
if (got !== exp) begin
$display("FAIL %0s: got %0d exp %0d", what, got, exp);
errors = errors + 1;
end
end
endtask
always @(posedge clk) if (rst_n) begin
if (frame_start) starts = starts + 1;
if (frame_end) ends = ends + 1;
end
initial begin
repeat (3) @(negedge clk); rst_n = 1; @(negedge clk);
chk("out of reset: not in frame", in_frame, 0);
chk("out of reset: cs_n_sync high", cs_n_sync, 1);
#3 cs_n_async = 0;
repeat (6) @(posedge clk);
chk("frame_start pulsed once", starts, 1);
chk("in_frame asserted", in_frame, 1);
chk("no frame_end yet", ends, 0);
repeat (10) @(posedge clk);
chk("no extra starts while held", starts, 1);
#7 cs_n_async = 1;
repeat (6) @(posedge clk);
chk("frame_end pulsed once", ends, 1);
chk("in_frame deasserted", in_frame, 0);
#1 cs_n_async = 0;
repeat (6) @(posedge clk);
chk("second frame_start", starts, 2);
#9 cs_n_async = 1;
repeat (6) @(posedge clk);
chk("second frame_end", ends, 2);
// A glitch wider than a clock period is indistinguishable from a frame.
s0 = starts; e0 = ends;
#2 cs_n_async = 0;
#25 cs_n_async = 1;
repeat (6) @(posedge clk);
chk("wide glitch produces a frame_start", starts - s0, 1);
chk("wide glitch produces a frame_end", ends - e0, 1);
if (errors == 0)
$display("PASS: CS crosses through the synchroniser, each edge yields exactly one pulse, and a glitch wider than a clock period is indistinguishable from a real frame");
else
$display("FAILED with %0d error(s)", errors);
$finish;
end
initial begin #200000; $display("FAIL: watchdog timeout"); $finish; end
endmodule-- spi_cs_frame.vhd — the same CS synchroniser and frame detector in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_cs_frame is
generic (
SYNC_STAGES : positive := 2
);
port (
clk : in std_logic;
rst_n : in std_logic;
cs_n_async : in std_logic; -- straight from the pin
cs_n_sync : out std_logic; -- safe to use in logic
frame_start : out std_logic; -- one pulse: CS asserted
frame_end : out std_logic; -- one pulse: CS deasserted
in_frame : out std_logic
);
end entity spi_cs_frame;
architecture rtl of spi_cs_frame is
-- Reset HIGH: deasserted is the safe idle state, so a slave leaving reset
-- must not believe it is mid-frame.
signal sync_q : std_logic_vector(SYNC_STAGES - 1 downto 0);
signal cs_n_d : std_logic;
signal cs_n_sync_i : std_logic;
begin
process (clk, rst_n) is
begin
if rst_n = '0' then
sync_q <= (others => '1');
cs_n_d <= '1';
elsif rising_edge(clk) then
sync_q <= sync_q(SYNC_STAGES - 2 downto 0) & cs_n_async;
cs_n_d <= sync_q(SYNC_STAGES - 1);
end if;
end process;
cs_n_sync_i <= sync_q(SYNC_STAGES - 1);
cs_n_sync <= cs_n_sync_i;
-- Edges come from the synchronised signal and its delayed copy only.
frame_start <= '1' when cs_n_sync_i = '0' and cs_n_d = '1' else '0';
frame_end <= '1' when cs_n_sync_i = '1' and cs_n_d = '0' else '0';
in_frame <= '1' when cs_n_sync_i = '0' else '0';
end architecture rtl;-- spi_cs_frame_tb.vhd — the same checks in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_cs_frame_tb is
end entity spi_cs_frame_tb;
architecture tb of spi_cs_frame_tb is
signal clk : std_logic := '0';
signal rst_n : std_logic := '0';
signal cs_n_async : std_logic := '1';
signal halt : boolean := false;
signal cs_n_sync : std_logic;
signal frame_start : std_logic;
signal frame_end : std_logic;
signal in_frame : std_logic;
signal errors : natural := 0;
signal starts : natural := 0;
signal ends : natural := 0;
begin
clk <= not clk after 5 ns when not halt else '0';
dut : entity work.spi_cs_frame
generic map (SYNC_STAGES => 2)
port map (clk => clk, rst_n => rst_n, cs_n_async => cs_n_async,
cs_n_sync => cs_n_sync, frame_start => frame_start,
frame_end => frame_end, in_frame => in_frame);
counter : process (clk) is
begin
if rising_edge(clk) and rst_n = '1' then
if frame_start = '1' then
starts <= starts + 1;
end if;
if frame_end = '1' then
ends <= ends + 1;
end if;
end if;
end process;
stim : process is
variable s0, e0 : natural;
procedure chk_n (what : string; got, exp : natural) is
begin
if got /= exp then
report "FAIL " & what & ": got " & integer'image(got)
& " exp " & integer'image(exp) severity error;
errors <= errors + 1;
end if;
end procedure;
procedure chk_b (what : string; got, exp : std_logic) is
begin
if got /= exp then
report "FAIL " & what severity error;
errors <= errors + 1;
end if;
end procedure;
begin
for i in 0 to 2 loop
wait until falling_edge(clk);
end loop;
rst_n <= '1';
wait until falling_edge(clk);
chk_b("out of reset: not in frame", in_frame, '0');
chk_b("out of reset: cs_n_sync high", cs_n_sync, '1');
-- Assert CS deliberately off the clock edge.
wait for 3 ns; cs_n_async <= '0';
for i in 0 to 5 loop
wait until rising_edge(clk);
end loop;
chk_n("frame_start pulsed once", starts, 1);
chk_b("in_frame asserted", in_frame, '1');
chk_n("no frame_end yet", ends, 0);
for i in 0 to 9 loop
wait until rising_edge(clk);
end loop;
chk_n("no extra starts while held", starts, 1);
wait for 7 ns; cs_n_async <= '1';
for i in 0 to 5 loop
wait until rising_edge(clk);
end loop;
chk_n("frame_end pulsed once", ends, 1);
chk_b("in_frame deasserted", in_frame, '0');
wait for 1 ns; cs_n_async <= '0';
for i in 0 to 5 loop
wait until rising_edge(clk);
end loop;
chk_n("second frame_start", starts, 2);
wait for 9 ns; cs_n_async <= '1';
for i in 0 to 5 loop
wait until rising_edge(clk);
end loop;
chk_n("second frame_end", ends, 2);
-- A glitch wider than a clock period is indistinguishable from a frame.
s0 := starts; e0 := ends;
wait for 2 ns; cs_n_async <= '0';
wait for 25 ns; cs_n_async <= '1';
for i in 0 to 5 loop
wait until rising_edge(clk);
end loop;
chk_n("wide glitch produces a frame_start", starts - s0, 1);
chk_n("wide glitch produces a frame_end", ends - e0, 1);
wait until falling_edge(clk);
if errors = 0 then
report "PASS: CS crosses through the synchroniser, each edge yields exactly "
& "one pulse, and a glitch wider than a clock period is indistinguishable "
& "from a real frame" severity note;
else
report "FAILED with " & integer'image(errors) & " error(s)" severity error;
end if;
halt <= true;
wait;
end process;
watchdog : process is
begin
wait for 200 us;
if not halt then
report "FAIL: watchdog timeout" severity failure;
end if;
wait;
end process;
end architecture tb;Parity
All three implement the same hardware: a parameterised synchroniser chain resetting high, one delay flop, edges taken only from the synchronised value, and single-cycle strobes. All three testbenches assert CS at several different phases relative to the clock, confirm exactly one pulse per edge, confirm no additional pulses while CS is held, and confirm that a wide glitch produces a full frame.
7. A Synchroniser Is Not a Glitch Filter
This distinction is confused often enough to be worth its own section, and the testbenches assert it rather than merely mentioning it.
A synchroniser answers one question: given that this signal changed near my clock edge, what value did it have? It resolves the metastability of sampling. It does not, and cannot, answer was that pulse intended?
So a glitch on CS wider than a clock period passes through and produces a complete frame: a frame_start, and a frame_end behind it. In a write path that is a frame that opens and closes — harmlessly, since no valid command arrives, but it is a real event the slave acts on. A glitch narrower than a clock period may or may not be seen at all, depending on where it lands; that nondeterminism is itself the reason not to rely on the behaviour.
If glitches are a genuine concern — a long cable, a noisy connector, a hot-plugged board — filtering is a separate mechanism: a digital debounce requiring N consecutive identical samples, or analogue conditioning at the pin. Adding synchroniser stages does not help, because more stages add latency without adding rejection.
The reason to be precise here is that "add another sync stage" is a common and useless response to a glitch problem. It makes the system slower and no more robust, and it obscures the actual fault.
8. Why a Verification Engineer Cares
The properties worth asserting are about the strobes, because everything downstream trusts them:
// 1. The two strobes are mutually exclusive: an edge is one or the other.
a_start_xor_end : assert property (
@(posedge clk) disable iff (!rst_n) !(frame_start && frame_end))
else $error("frame_start and frame_end asserted together");
// 2. Strobes are single-cycle. A frame_end held for two cycles would
// commit a Chapter 5.1 write twice.
a_start_is_pulse : assert property (
@(posedge clk) disable iff (!rst_n) frame_start |=> !frame_start)
else $error("frame_start held for more than one cycle");
// 3. Events alternate: no two starts without an end between them. This is
// what "in_frame is a reliable level" actually means.
a_alternating : assert property (
@(posedge clk) disable iff (!rst_n)
frame_start |=> (!frame_start throughout frame_end[->1]))
else $error("two frame_starts with no intervening frame_end");
// 4. Out of reset the slave must not believe it is mid-frame.
a_reset_idle : assert property (
@(posedge clk) $rose(rst_n) |-> !in_frame)
else $error("in_frame asserted immediately after reset");What these prove. That the module converts a level into well-formed, non-overlapping, alternating events. What they cannot prove is the thing the module exists for: that metastability was handled. A simulator samples cleanly and never produces a metastable value, so an RTL test bench cannot distinguish this design from one that uses the raw pin directly. Both pass every property above.
That gap is worth being explicit about, because it is a place where passing simulation actively misleads. The synchroniser's correctness is established by structure and constraints, not by simulation: the chain exists, edges derive only from its output, and the timing tool is told it is a synchroniser. Verification's contribution is a lint or structural check that no logic reads cs_n_async other than the first synchroniser stage — which is a rule a tool can enforce and a waveform cannot show.
Coverage should focus on the phase relationship, since that is the axis the design exists to tolerate:
covergroup spi_cs_phase_cg @(posedge clk);
// Where in the clock period the CS edge landed. A suite that only
// changes CS on clock edges tests the one case needing no synchroniser.
cp_edge_phase : coverpoint cs_edge_offset_ps {
bins near_setup = {[0:200]}; // close to the edge -- the hard case
bins mid_period = {[201:9800]};
bins near_hold = {[9801:10000]}; // also the hard case
}
cp_frame_len : coverpoint frame_length_cycles {
bins glitch = {[1:2]}; // §7: shorter than a real frame
bins short = {[3:16]};
bins normal = {[17:1024]};
}
endgroup9. Why an FPGA or ASIC Engineer Cares
Tell the tools it is a synchroniser. Synthesis will happily retime across a two-flop chain or merge the flops if it believes it is optimising a shift register — destroying the very thing that makes the design safe. Every toolchain has a mechanism: an attribute such as ASYNC_REG on the flops, a false path or max-delay constraint on the crossing, and a directive preventing the flops being packed into a shift-register primitive. This is not optional polish; a synchroniser without constraints may not survive implementation.
Place the synchroniser flops close together. The time available for the first flop to resolve is one clock period minus the routing delay to the second flop. A long route between them directly reduces the metastability margin, which is why tools provide a way to keep them adjacent.
The latency lands in the CS lead budget. Three clocks at 50 MHz is 60 ns, and the master's CS-lead-to-first-edge requirement must exceed it or the slave will still be recognising CS when the first bit is clocked. This is exactly the sort of constraint that does not appear in RTL simulation and then appears on a board — worth computing during design rather than discovering.
Do not clock the slave's logic with CS. Using CS as a clock or an asynchronous reset for the shift register is a tempting shortcut and a genuine hazard: it creates a clock domain that exists only while an external device holds a pin low, with no defined relationship to anything, and it is not analysable by static timing. The strobe-based architecture used throughout this track keeps everything in one always-running domain.
On an ASIC the same rules apply, with the addition that the synchroniser flops should be a characterised cell. Standard-cell libraries usually provide flops specified for metastability resolution, and the MTBF calculation depends on using them.
10. Failure Signature — The First Bit of Every Transfer Is Wrong
Symptom. Every transfer is off by one bit position: the received data looks like the expected value shifted, and the first bit is garbage. It is perfectly reproducible, identical on every transaction, and unaffected by clock rate.
What reproducibility rules out immediately. Chapter 2.4 established that a margin problem improves as the clock slows. This does not change at all, so it is logical, not analogue.
Plausible mechanisms.
- The slave misses the first bit because it has not finished recognising CS when the first clock edge arrives — the synchroniser latency of §5 exceeding the master's CS lead interval. This is the mechanism specific to this chapter.
- A CPHA disagreement (Chapter 3.8), which also displaces by one position.
- A first-bit launch bug in the master (Chapter 3.4 §3), where the first bit is never placed before the leading edge.
The discriminating observation, and it is decisive. Increase the CS lead time — the delay between asserting CS and starting the clock — without changing anything else. If the fault disappears, it is the synchroniser latency and the fix is a longer lead (or fewer sync stages, if the margin analysis supports it). If it is unaffected, the lead interval is not the constraint and the fault is in the mode or the launch path.
That test is valuable because it varies exactly one parameter that only the first mechanism depends on. A mode mismatch and a launch bug are both indifferent to how long CS has been asserted.
Why the investigation goes wrong. Because "off by one bit" points everyone at CPHA, which is the famous cause, and the mode gets checked repeatedly while the lead interval is never questioned — it is not usually thought of as a parameter at all. On an FPGA slave with a slow system clock relative to SCLK, it very much is one.
11. Common Misconceptions
12. Reason It Through
Work this before reading the answer.
An FPGA SPI slave runs from a 25 MHz system clock and uses a two-flop synchroniser on CS. The master asserts CS and begins clocking at 10 MHz after a 40 ns lead. Transfers work. The team then raises the SPI clock to 20 MHz, keeping the same 40 ns lead, and the first bit of every transfer becomes wrong.
Why, and what are the options?
Work out the slave's reaction time first. A 25 MHz system clock is a 40 ns period. A two-flop synchroniser plus the edge-detect delay flop means the frame_start strobe appears three system-clock periods after the CS pin falls — 120 ns in the worst case.
Now compare against the lead. The master waits 40 ns and then starts clocking. So the first SCLK edge arrives roughly 80 ns before the slave has even registered that CS is asserted. That was already true at 10 MHz.
So why did it work at 10 MHz? Because of what happens next, and this is the part worth getting right. At 10 MHz the first sampling edge comes half a bit time after clocking starts — 50 ns — for a total of 90 ns from CS. Still inside the slave's 120 ns. But under CPHA = 0 the first bit only has to be captured by the slave's shift logic, and the slave's own sampling is driven by its detection of SCLK edges, which are themselves synchronised. At the slower rate, the first SCLK edge is still being processed when the frame is finally recognised, so the bit is not lost. At 20 MHz the edges arrive twice as fast and the first one has come and gone before the frame opens.
The general statement. The slave must recognise CS before the first SCLK edge it needs to act on. That gives a requirement:
CS lead > (SYNC_STAGES + 1) x T_system
Here: (2 + 1) x 40 ns = 120 ns required
40 ns supplied → violated at any SPI rateThe options, in order of preference.
Increase the CS lead to more than 120 ns. Free, if the master can be configured for it — many controllers expose a CS-to-clock delay, and this is what it is for.
Raise the slave's system clock. At 100 MHz the same three stages cost 30 ns and the existing 40 ns lead is sufficient. This is often the real fix on an FPGA, where the SPI interface is needlessly clocked from a slow domain.
Reduce to a single synchroniser stage. Cheapest in latency and the worst idea: it trades a deterministic functional requirement for a probabilistic failure that appears in the field rather than on the bench.
Clock the slave's shift register directly from SCLK, removing the crossing for the data path. This is a legitimate and widely-used architecture — it is how many production slaves are built — but it is a different design with its own domain-crossing problem at the boundary to the system clock, and it belongs to Module 15 rather than being a patch.
The lesson worth keeping. Synchroniser latency is a protocol-level timing parameter, not an implementation detail. It belongs in the same budget as CS lead and clock-to-output, and the ratio between the slave's system clock and SCLK is what decides whether it fits.
13. Understanding Check
14. Summary
A frame has three regions — CS lead, active interval, CS lag — and the rules differ in each. Under CPHA = 0 the first bit must be on MOSI before the lead ends; during the active interval every edge moves a bit in both directions; during the lag no data moves but the write commits.
There is no "not sending" on SPI. The master drives MOSI on every clock edge, so a read is an exchange whose outbound half is meaningless by agreement — an agreement recorded in the datasheet, since a few devices do care what they receive during a read.
The two CS edges are not symmetric. The falling edge opens a frame and starts MISO driving; the rising edge commits, releases and resynchronises. A spurious falling edge is mostly harmless; a spurious rising edge can commit a partial write, and nothing reports it.
Because CS comes from an external master, it is asynchronous to the slave and must cross through a synchroniser before any edge is taken from it — deriving an edge from a metastable sample can manufacture one. The chain must reset high, so a slave leaving reset does not believe it is mid-frame.
A synchroniser is not a glitch filter: it resolves metastability, not intent, and a glitch wider than a clock period becomes a complete frame. Adding stages adds latency without rejection.
That latency is a protocol timing parameter: CS lead > (SYNC_STAGES + 1) × T_system. Violating it produces a lost first bit that looks exactly like a CPHA mismatch, and the measurement that separates them is to lengthen the CS lead and see whether the fault clears.
Finally, an RTL simulation cannot verify this crossing — simulators never go metastable, so the raw-pin version passes identically. Correctness comes from structure, constraints, and a lint rule that nothing but the first stage reads the pin.
15. What Comes Next
Chapters 5.1 and 5.2 covered a write of a single register. Chapter 5.3 — Register and Multi-Byte Writes extends the frame: what happens when data keeps arriving after the first byte, how a device auto-increments its address so one transaction can fill a buffer, and why that increment is usually bounded by a page — so a write crossing a page boundary wraps to the start of the page and overwrites data the master never intended to touch, with no error and no indication.
Continue learning
Related tutorials
- Related topic
Mode 0 (CPOL=0, CPHA=0)
The most widely used SPI mode, and the one carrying a real implementation problem: why CPHA=0 forces the first bit onto the line before any clock edge, and the first-bit launch path in Verilog, SystemVerilog and VHDL.
- Related topic
Bus Topologies and Daisy-Chain
What the four-wire model costs as devices are added: shared clock and data with one select per device, or a daisy chain that turns several peripherals into one long shift ring. Pin arithmetic, ownership consequences, and why chaining is a device property.
- Related topic
CS-to-SCLK and SCLK-to-CS Timing
Chip select has timing requirements of its own: the lead before the first clock edge, the lag after the last, and the minimum deselect between transactions. Why violating them breaks a transfer whose every SCLK edge was correct.
- Related topic
Bit Ordering — MSB-First and LSB-First
Which end of the shift register goes out first, the two multiplexers that make the order configurable in three HDLs, and why a bit-order bug is perfectly deterministic and yet invisible on certain data.
