SPI · Module 4
Bit Ordering — MSB-First and LSB-First
Which end of the shift register goes out first, the two multiplexers that make the order configurable in three HDLs, and why a bit-order bug is perfectly deterministic and yet invisible on certain data.
Chapter 4.2 settled how many bits a transfer carries. This chapter settles which of them leaves first.
A shift register has two ends. Which one feeds the pin — and what does the choice cost in hardware, in verification, and when it is wrong?
The hardware answer is two multiplexers. The verification answer is considerably more interesting, because a bit-order bug has the unusual property of being completely deterministic and frequently invisible at the same time.
1. Which End Goes First
Chapter 1.3 built the exchange as a ring of shift registers, and drew the transmit register shifting left with the bit leaving from the top. That was a choice, not a necessity.
MSB-first means the most significant bit is transmitted first: the register shifts left, and the pin is tapped at bit WIDTH-1. LSB-first means the least significant bit leads: the register shifts right, and the pin is tapped at bit 0.
The receive side mirrors it exactly. If bits arrive most-significant-first, each new bit enters at the bottom of the receive register and everything shifts up. If they arrive least-significant-first, each new bit enters at the top and everything shifts down. In both cases the rule is the same: the receive register fills from the end opposite the one the transmit register empties toward, so a bit's position is preserved across the link.
Nothing about SPI prefers either. The bus clocks edges; which bit you chose to place on the pin is entirely a property of the two endpoints, and — exactly as in Chapter 4.1 — it is not observable on the wire. A capture shows a sequence of bits; it does not show which one you considered significant.
2. Why MSB-First Became Usual
MSB-first dominates, and the reasons are worth knowing because they explain where the exceptions live.
Numeric values arrive usefully. With MSB-first, a receiver accumulating bits has a progressively refined approximation of the value: after three bits of a 12-bit ADC sample it already knows the top three bits, so it knows the magnitude. Under LSB-first, no useful information exists until the final bit arrives. For a converter that matters; for a truncating receiver it matters a great deal.
It matches how values are written. 0xA5 written out is 10100101, and MSB-first puts those bits on the wire in that order. A logic analyzer trace then reads the same way as the datasheet, which removes a whole category of transcription error.
Addresses compose. A flash command followed by a 24-bit address, sent MSB-first, is a byte sequence you can read directly as the address. Under LSB-first the bytes are individually reversed and the concatenation stops being legible.
The exceptions cluster where those reasons do not apply: some shift-register expanders, where the bits are independent outputs and no value is being represented; some communication peripherals, where LSB-first matches UART convention; and a number of microcontroller SPI blocks that simply expose the choice because it is nearly free to offer.
3. The Same Byte, Both Ways
0x2C on the wire, both orders
10 cyclesRead the two data rows as the same byte. 0x2C is 00101100; MSB-first walks it left to right, LSB-first walks it right to left. Six of the eight bit times happen to agree, which is not a coincidence worth relying on — it is the first hint of §5.
4. What Changes in the Hardware — Three HDLs
Configurable bit order costs two multiplexers, and being precise about which two is the whole design.
The circuit
Circuit. Chapter 4.2's engine with a direction control added to both shift registers and a tap selector on the output.
State. Unchanged — transmit register, receive register, bit counter, registered CS.
Datapath. Two decisions. The shift direction: left for MSB-first, right for LSB-first, on both registers. The output tap: bit WIDTH-1 for MSB-first, bit 0 for LSB-first. On the receive side the entry point moves correspondingly — bottom for MSB-first, top for LSB-first.
Control. A single msb_first input selects both. Critically, it must be stable for the duration of a frame: changing it mid-word leaves the register half-shifted in each direction, producing a result that corresponds to no bit order at all. §6 makes that an assertion.
Clock and reset. Unchanged from Chapter 4.2 — system clock, asynchronous active-low reset to constants.
Timing. The multiplexers are combinational and sit between the register and the pin, so they add one level of logic to the sdo path — relevant only at the output register, and not on the SCLK path at all.
Synthesis. Two WIDTH-bit 2:1 multiplexers and one 1-bit 2:1 multiplexer. No additional storage: configurable bit order costs no flip-flops, which is why so many controllers offer it.
Limitations. The engine still assumes a fixed frame width per Chapter 4.2, and takes launch/sample strobes from Chapter 3.3's decoder rather than generating them.
// spi_bitorder.sv — a transfer engine whose bit order is configurable.
//
// Bit order costs exactly two multiplexers: one choosing which END of the
// transmit register feeds the pin, and one choosing which DIRECTION each
// register shifts. Everything else is Chapter 4.2's engine unchanged.
module spi_bitorder #(
parameter int WIDTH = 8
) (
input logic clk,
input logic rst_n,
input logic cs_n,
input logic launch_stb,
input logic sample_stb,
input logic msb_first, // 1 = MSB-first, 0 = LSB-first
input logic [WIDTH-1:0] tx_data,
input logic sdi,
output logic sdo,
output logic [WIDTH-1:0] rx_data,
output logic done
);
localparam int CW = (WIDTH > 1) ? $clog2(WIDTH) : 1;
logic [WIDTH-1:0] tx_sh, rx_sh, rx_next;
logic [CW-1:0] bit_cnt;
logic cs_n_q;
// The receive register fills from the opposite end it empties toward, so
// the first bit in ends up where the first bit out came from.
assign rx_next = msb_first ? {rx_sh[WIDTH-2:0], sdi} // in at the bottom
: {sdi, rx_sh[WIDTH-1:1]}; // in at the top
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
tx_sh <= '0;
rx_sh <= '0;
bit_cnt <= '0;
rx_data <= '0;
done <= 1'b0;
cs_n_q <= 1'b1;
end else begin
cs_n_q <= cs_n;
done <= 1'b0;
if (cs_n) begin
tx_sh <= tx_data;
bit_cnt <= '0;
end else begin
if (cs_n_q) begin
tx_sh <= tx_data;
bit_cnt <= '0;
end
if (sample_stb) begin
rx_sh <= rx_next;
if (bit_cnt == CW'(WIDTH - 1)) begin
bit_cnt <= '0;
rx_data <= rx_next;
done <= 1'b1;
end else begin
bit_cnt <= bit_cnt + 1'b1;
end
end
if (launch_stb)
tx_sh <= msb_first ? {tx_sh[WIDTH-2:0], 1'b0} // shift left
: {1'b0, tx_sh[WIDTH-1:1]}; // shift right
end
end
end
// The output tap follows the shift direction: the bit about to leave.
assign sdo = msb_first ? tx_sh[WIDTH-1] : tx_sh[0];
endmoduleThe receive path deserves a second look. rx_next is computed once and used twice — for the running register and for the final rx_data latch — so the last bit is captured by the same expression that captures every other bit. Writing the final capture separately is a classic way to introduce an off-by-one that appears only in the last bit position.
// spi_bitorder_tb.sv — both orders, and the values that HIDE an order bug.
`timescale 1ns/1ps
module spi_bitorder_tb;
logic clk = 0, rst_n = 0;
always #5 clk = ~clk;
localparam int W = 8;
logic cs_n = 1, launch = 0, sample = 0, sdi = 0, msb_first = 1;
logic [W-1:0] tx_data;
logic sdo; logic [W-1:0] rx_data; logic done;
spi_bitorder #(.WIDTH(W)) dut (.clk, .rst_n, .cs_n, .launch_stb(launch), .sample_stb(sample),
.msb_first, .tx_data, .sdi, .sdo, .rx_data, .done);
int errors = 0;
task automatic chk(input string what, input int got, input int exp);
if (got !== exp) begin $display("FAIL %s: got 0x%0h exp 0x%0h", what, got, exp); errors++; end
endtask
function automatic logic [W-1:0] reverse(input logic [W-1:0] v);
for (int i = 0; i < W; i++) reverse[i] = v[W-1-i];
endfunction
// Drive one bit time and collect sdo (sampled before the launch edge).
logic [W-1:0] got_sdo;
// Values whose bit-reversal equals themselves. MSB-first and LSB-first
// produce IDENTICAL wire patterns for each, so a suite built only from
// them passes with the bit order wired backwards.
logic [W-1:0] blind [4] = '{8'h00, 8'hFF, 8'hA5, 8'h18};
task automatic run_frame(input logic [W-1:0] din);
got_sdo = '0;
cs_n = 0; @(negedge clk);
for (int i = 0; i < W; i++) begin
got_sdo = {got_sdo[W-2:0], sdo};
sdi = din[W-1-i]; // feed MSB-first on the wire
@(negedge clk); sample = 1; @(negedge clk); sample = 0;
@(negedge clk); launch = 1; @(negedge clk); launch = 0;
end
cs_n = 1; @(negedge clk);
endtask
initial begin
repeat (3) @(negedge clk); rst_n = 1; @(negedge clk);
// --- MSB-first: the wire carries the word as written ---
msb_first = 1; tx_data = 8'h2C;
run_frame(8'h9B);
chk("MSB-first sdo", got_sdo, 8'h2C);
chk("MSB-first rx", rx_data, 8'h9B);
// --- LSB-first: the wire carries the word bit-reversed ---
msb_first = 0; tx_data = 8'h2C;
run_frame(8'h9B);
chk("LSB-first sdo", got_sdo, reverse(8'h2C)); // 0x2C -> 0x34
chk("LSB-first rx", rx_data, reverse(8'h9B)); // 0x9B -> 0xD9
// --- The values that HIDE the bug. For each of these, MSB-first and
// LSB-first produce IDENTICAL wire patterns, so a test using only
// them passes with the order wired backwards. This is the point
// of the chapter, asserted rather than asserted-about. ---
begin
foreach (blind[k]) begin
logic [W-1:0] wire_msb, wire_lsb;
msb_first = 1; tx_data = blind[k]; run_frame(8'h00); wire_msb = got_sdo;
msb_first = 0; tx_data = blind[k]; run_frame(8'h00); wire_lsb = got_sdo;
chk($sformatf("blind value 0x%02h is a palindrome", blind[k]), wire_lsb, wire_msb);
end
end
if (errors == 0)
$display("PASS: both orders transmit and receive correctly; the four palindromic values produce identical wire patterns in both orders");
else
$display("FAILED with %0d error(s)", errors);
$finish;
end
initial begin #200000; $display("FAIL: watchdog timeout"); $finish; end
endmodule// spi_bitorder.v — the same configurable-order engine in Verilog-2001.
module spi_bitorder #(
parameter WIDTH = 8
) (
input wire clk,
input wire rst_n,
input wire cs_n,
input wire launch_stb,
input wire sample_stb,
input wire msb_first,
input wire [WIDTH-1:0] tx_data,
input wire sdi,
output wire sdo,
output reg [WIDTH-1:0] rx_data,
output reg done
);
function integer clogb2;
input integer value;
integer v;
begin
v = value - 1;
for (clogb2 = 0; v > 0; clogb2 = clogb2 + 1) v = v >> 1;
end
endfunction
localparam CW = (WIDTH > 1) ? clogb2(WIDTH) : 1;
reg [WIDTH-1:0] tx_sh, rx_sh;
reg [CW-1:0] bit_cnt;
reg cs_n_q;
wire [WIDTH-1:0] rx_next;
assign rx_next = msb_first ? {rx_sh[WIDTH-2:0], sdi}
: {sdi, rx_sh[WIDTH-1:1]};
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
tx_sh <= {WIDTH{1'b0}};
rx_sh <= {WIDTH{1'b0}};
bit_cnt <= {CW{1'b0}};
rx_data <= {WIDTH{1'b0}};
done <= 1'b0;
cs_n_q <= 1'b1;
end else begin
cs_n_q <= cs_n;
done <= 1'b0;
if (cs_n) begin
tx_sh <= tx_data;
bit_cnt <= {CW{1'b0}};
end else begin
if (cs_n_q) begin
tx_sh <= tx_data;
bit_cnt <= {CW{1'b0}};
end
if (sample_stb) begin
rx_sh <= rx_next;
if (bit_cnt == (WIDTH - 1)) begin
bit_cnt <= {CW{1'b0}};
rx_data <= rx_next;
done <= 1'b1;
end else begin
bit_cnt <= bit_cnt + 1'b1;
end
end
if (launch_stb)
tx_sh <= msb_first ? {tx_sh[WIDTH-2:0], 1'b0}
: {1'b0, tx_sh[WIDTH-1:1]};
end
end
end
assign sdo = msb_first ? tx_sh[WIDTH-1] : tx_sh[0];
endmodule// spi_bitorder_tb.v — the same checks in Verilog-2001.
`timescale 1ns/1ps
module spi_bitorder_tb;
reg clk = 0, rst_n = 0;
always #5 clk = ~clk;
parameter W = 8;
reg cs_n = 1, launch = 0, sample = 0, sdi = 0, msb_first = 1;
reg [W-1:0] tx_data;
wire sdo; wire [W-1:0] rx_data; wire done;
spi_bitorder #(.WIDTH(W)) dut (
.clk(clk), .rst_n(rst_n), .cs_n(cs_n), .launch_stb(launch), .sample_stb(sample),
.msb_first(msb_first), .tx_data(tx_data), .sdi(sdi), .sdo(sdo),
.rx_data(rx_data), .done(done));
integer errors = 0, i, k;
reg [W-1:0] got_sdo, din_r, wire_msb, wire_lsb;
// Values whose bit-reversal equals themselves: both orders look identical.
reg [W-1:0] blind [0:3];
initial begin
blind[0] = 8'h00; blind[1] = 8'hFF; blind[2] = 8'hA5; blind[3] = 8'h18;
end
function [W-1:0] reverse;
input [W-1:0] v;
integer j;
begin
for (j = 0; j < W; j = j + 1) reverse[j] = v[W-1-j];
end
endfunction
task chk;
input [80*8-1:0] what;
input [W-1:0] got, exp;
begin
if (got !== exp) begin
$display("FAIL %0s: got 0x%0h exp 0x%0h", what, got, exp);
errors = errors + 1;
end
end
endtask
task run_frame;
input [W-1:0] din;
begin
got_sdo = {W{1'b0}};
cs_n = 0; @(negedge clk);
for (i = 0; i < W; i = i + 1) begin
got_sdo = {got_sdo[W-2:0], sdo};
sdi = din[W-1-i];
@(negedge clk); sample = 1; @(negedge clk); sample = 0;
@(negedge clk); launch = 1; @(negedge clk); launch = 0;
end
cs_n = 1; @(negedge clk);
end
endtask
initial begin
repeat (3) @(negedge clk); rst_n = 1; @(negedge clk);
msb_first = 1; tx_data = 8'h2C; run_frame(8'h9B);
chk("MSB-first sdo", got_sdo, 8'h2C);
chk("MSB-first rx", rx_data, 8'h9B);
msb_first = 0; tx_data = 8'h2C; run_frame(8'h9B);
chk("LSB-first sdo", got_sdo, reverse(8'h2C));
chk("LSB-first rx", rx_data, reverse(8'h9B));
for (k = 0; k < 4; k = k + 1) begin
msb_first = 1; tx_data = blind[k]; run_frame(8'h00); wire_msb = got_sdo;
msb_first = 0; tx_data = blind[k]; run_frame(8'h00); wire_lsb = got_sdo;
chk("blind palindromic value", wire_lsb, wire_msb);
end
if (errors == 0)
$display("PASS: both orders transmit and receive correctly; the four palindromic values produce identical wire patterns in both orders");
else
$display("FAILED with %0d error(s)", errors);
$finish;
end
initial begin #200000; $display("FAIL: watchdog timeout"); $finish; end
endmoduleVHDL expresses the reversal helper most directly, because 'range, 'high and 'low let the function work for any vector without a width parameter.
-- spi_bitorder.vhd — the same configurable-order engine in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_bitorder is
generic (
WIDTH : positive := 8
);
port (
clk : in std_logic;
rst_n : in std_logic;
cs_n : in std_logic;
launch_stb : in std_logic;
sample_stb : in std_logic;
msb_first : in std_logic;
tx_data : in std_logic_vector(WIDTH - 1 downto 0);
sdi : in std_logic;
sdo : out std_logic;
rx_data : out std_logic_vector(WIDTH - 1 downto 0);
done : out std_logic
);
end entity spi_bitorder;
architecture rtl of spi_bitorder is
signal bit_cnt : integer range 0 to WIDTH - 1;
signal tx_sh : std_logic_vector(WIDTH - 1 downto 0);
signal rx_sh : std_logic_vector(WIDTH - 1 downto 0);
signal rx_next : std_logic_vector(WIDTH - 1 downto 0);
signal cs_n_q : std_logic;
begin
-- The receive register fills from the end opposite the one it empties
-- toward, so the first bit in lands where the first bit out came from.
rx_next <= rx_sh(WIDTH - 2 downto 0) & sdi when msb_first = '1'
else sdi & rx_sh(WIDTH - 1 downto 1);
process (clk, rst_n) is
begin
if rst_n = '0' then
tx_sh <= (others => '0');
rx_sh <= (others => '0');
bit_cnt <= 0;
rx_data <= (others => '0');
done <= '0';
cs_n_q <= '1';
elsif rising_edge(clk) then
cs_n_q <= cs_n;
done <= '0';
if cs_n = '1' then
tx_sh <= tx_data;
bit_cnt <= 0;
else
if cs_n_q = '1' then
tx_sh <= tx_data;
bit_cnt <= 0;
end if;
if sample_stb = '1' then
rx_sh <= rx_next;
if bit_cnt = WIDTH - 1 then
bit_cnt <= 0;
rx_data <= rx_next;
done <= '1';
else
bit_cnt <= bit_cnt + 1;
end if;
end if;
if launch_stb = '1' then
if msb_first = '1' then
tx_sh <= tx_sh(WIDTH - 2 downto 0) & '0'; -- shift left
else
tx_sh <= '0' & tx_sh(WIDTH - 1 downto 1); -- shift right
end if;
end if;
end if;
end if;
end process;
-- The output tap follows the shift direction: the bit about to leave.
sdo <= tx_sh(WIDTH - 1) when msb_first = '1' else tx_sh(0);
end architecture rtl;-- spi_bitorder_tb.vhd — the same checks in VHDL.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity spi_bitorder_tb is
end entity spi_bitorder_tb;
architecture tb of spi_bitorder_tb is
constant W : positive := 8;
signal clk : std_logic := '0';
signal rst_n : std_logic := '0';
signal cs_n : std_logic := '1';
signal launch : std_logic := '0';
signal sample : std_logic := '0';
signal sdi : std_logic := '0';
signal msb_first : std_logic := '1';
signal halt : boolean := false;
signal tx_data : std_logic_vector(W - 1 downto 0) := (others => '0');
signal sdo : std_logic;
signal rx_data : std_logic_vector(W - 1 downto 0);
signal done : std_logic;
signal errors : natural := 0;
signal got_sdo : std_logic_vector(W - 1 downto 0) := (others => '0');
type blind_arr is array (0 to 3) of std_logic_vector(W - 1 downto 0);
-- Values whose bit-reversal equals themselves: both orders look identical.
constant BLIND : blind_arr := (x"00", x"FF", x"A5", x"18");
function reverse (v : std_logic_vector) return std_logic_vector is
variable r : std_logic_vector(v'range);
begin
for i in v'range loop
r(v'high - i + v'low) := v(i);
end loop;
return r;
end function;
begin
clk <= not clk after 5 ns when not halt else '0';
dut : entity work.spi_bitorder
generic map (WIDTH => W)
port map (clk => clk, rst_n => rst_n, cs_n => cs_n, launch_stb => launch,
sample_stb => sample, msb_first => msb_first, tx_data => tx_data,
sdi => sdi, sdo => sdo, rx_data => rx_data, done => done);
stim : process is
variable wire_msb : std_logic_vector(W - 1 downto 0);
variable wire_lsb : std_logic_vector(W - 1 downto 0);
procedure chk (what : string; got, exp : std_logic_vector) is
begin
if got /= exp then
report "FAIL " & what & ": got 0x" & to_hstring(got)
& " exp 0x" & to_hstring(exp) severity error;
errors <= errors + 1;
end if;
end procedure;
procedure run_frame (din : std_logic_vector(W - 1 downto 0)) is
begin
got_sdo <= (others => '0');
cs_n <= '0';
wait until falling_edge(clk);
for i in 0 to W - 1 loop
got_sdo <= got_sdo(W - 2 downto 0) & sdo;
sdi <= din(W - 1 - i);
wait until falling_edge(clk); sample <= '1';
wait until falling_edge(clk); sample <= '0';
wait until falling_edge(clk); launch <= '1';
wait until falling_edge(clk); launch <= '0';
end loop;
cs_n <= '1';
wait until falling_edge(clk);
end procedure;
begin
for i in 0 to 2 loop
wait until falling_edge(clk);
end loop;
rst_n <= '1';
wait until falling_edge(clk);
msb_first <= '1'; tx_data <= x"2C";
wait until falling_edge(clk);
run_frame(x"9B");
chk("MSB-first sdo", got_sdo, x"2C");
chk("MSB-first rx", rx_data, x"9B");
msb_first <= '0'; tx_data <= x"2C";
wait until falling_edge(clk);
run_frame(x"9B");
chk("LSB-first sdo", got_sdo, reverse(x"2C"));
chk("LSB-first rx", rx_data, reverse(x"9B"));
for k in BLIND'range loop
msb_first <= '1'; tx_data <= BLIND(k);
wait until falling_edge(clk);
run_frame(x"00");
wire_msb := got_sdo;
msb_first <= '0'; tx_data <= BLIND(k);
wait until falling_edge(clk);
run_frame(x"00");
wire_lsb := got_sdo;
chk("blind palindromic value", wire_lsb, wire_msb);
end loop;
wait until falling_edge(clk);
if errors = 0 then
report "PASS: both orders transmit and receive correctly; the four "
& "palindromic values produce identical wire patterns in both orders"
severity note;
else
report "FAILED with " & integer'image(errors) & " error(s)" severity error;
end if;
halt <= true;
wait;
end process;
watchdog : process is
begin
wait for 200 us;
if not halt then
report "FAIL: watchdog timeout" severity failure;
end if;
wait;
end process;
end architecture tb;Parity
All three implement identical hardware: the same ports plus msb_first, asynchronous active-low reset to constants, CS reload with priority, a counter advancing on sample_stb, a direction multiplexer on each shift register, and an output tap selected between the top and bottom bit. All three testbenches run the same vectors — 0x2C out and 0x9B in under both orders — and all three then run the same four palindromic values.
5. The Values That Hide the Bug
This is the part worth remembering after the RTL is forgotten.
A bit-order error is deterministic: the same input always produces the same wrong output. That should make it easy to catch. But consider what happens for particular values:
| Value | Binary | Reversed | Same? |
|---|---|---|---|
0x00 | 00000000 | 00000000 | yes |
0xFF | 11111111 | 11111111 | yes |
0xA5 | 10100101 | 10100101 | yes |
0x18 | 00011000 | 00011000 | yes |
0x2C | 00101100 | 00110100 | no |
For any bit-palindromic value, MSB-first and LSB-first put identical patterns on the wire. A test suite built from 0x00, 0xFF and 0xA5 — three of the most commonly chosen test values in existence — will pass with the bit order wired backwards.
0xA5 is the trap worth naming explicitly. It is chosen constantly because its alternating bits look like a good pattern for catching stuck-at and shorted-pin faults, which it is. It is simultaneously a bit palindrome, and therefore blind to exactly the bug this chapter is about.
Of the 256 byte values, 16 are palindromic — one in sixteen. Random data finds the bug almost immediately; hand-picked "obvious" test values find it remarkably often never, because the values engineers reach for are disproportionately symmetric.
That is why the testbenches above assert the blindness rather than merely mentioning it: they check that all four palindromic values produce identical wire patterns under both orders. If someone later changes the engine so that palindromic values differ between orders, the engine is wrong, and the test says so.
6. Why a Verification Engineer Cares
Two things belong in the environment, and neither is about data comparison.
The configuration must not move during a frame. §4 noted that changing msb_first mid-word leaves the register partially shifted in each direction. That is a temporal invariant, which is what assertions are for:
// A mid-frame change produces a word that corresponds to NO bit order --
// not a reversed word, a meaningless one. Worth catching at the source
// rather than debugging from corrupted data.
property p_order_stable_in_frame;
@(posedge clk) disable iff (!rst_n)
!cs_n |=> $stable(msb_first);
endproperty
a_order_stable : assert property (p_order_stable_in_frame)
else $error("msb_first changed while CS was asserted");What it proves. That the order was constant across every frame. What it does not prove. That the order is the one the device expects — again the Chapter 4.1 §7 boundary. It also says nothing about the interval between frames, which is deliberate: reconfiguring between transfers is legal and is how a controller serves two devices with different conventions.
The stimulus must not be symmetric. This is the coverage consequence of §5, and it is a constraint on the sequence rather than a coverpoint:
covergroup spi_bitorder_cg @(posedge transfer_done);
cp_order : coverpoint cfg.msb_first { bins msb = {1}; bins lsb = {0}; }
// The bug is only OBSERVABLE on data that is not its own reversal.
// Covering "both orders ran" is not enough if every word was 0xA5.
cp_asymmetric : coverpoint (tx_word != bit_reverse(tx_word)) {
bins symmetric = {0}; // the blind spot -- fine to hit, fatal to hit ONLY
bins asymmetric = {1}; // the bin that can actually fail
}
x_order_asym : cross cp_order, cp_asymmetric;
endgroupThe cross is the point. Four bins, and the two that matter are MSB-first with asymmetric data and LSB-first with asymmetric data. A report showing both orders covered while cp_asymmetric sits entirely in symmetric is a report saying nothing has been tested, and it will not look like one.
This generalises past bit order: whenever a fault can only manifest on data with a particular property, covering the configuration is not enough — the data property must be covered too, and crossed with the configuration.
7. Why an FPGA or ASIC Engineer Cares
It is free in flip-flops and not quite free in timing. Configurable order adds no state, but it does insert a multiplexer between the shift register and the output pin. On an FPGA, if the output register is packed into the I/O block for a clean clock-to-out (Chapter 2.6), a multiplexer after that register cannot be packed with it, and the tool will either pull the register out of the IOB — degrading clock-to-out — or place the multiplexer before it. Selecting the tap before the output register, so that the register holds the already-selected bit, keeps the IOB packing intact. That is a real and commonly-missed implementation detail, and it is why the tap selection in §4's RTL sits on a register output feeding a pin rather than being buried mid-path.
Reversal is free at elaboration and not at runtime. A fixed LSB-first design needs no multiplexers at all — the tap and direction are constants, and synthesis optimises the choice away entirely. Only runtime configurability costs logic. If a design only ever talks to one device, making the order a parameter rather than a port is strictly cheaper.
Do not reverse in software. It is tempting to leave the hardware MSB-first and bit-reverse bytes in the driver. On a processor without a bit-reverse instruction that is roughly eight operations per byte, which at any real data rate dominates the transfer cost. Most ARM cores have RBIT, but the assumption is worth checking rather than inheriting.
8. Failure Signature — Every Byte Is Bit-Reversed
Symptom. A device returns values that are consistently wrong but not random. Repeating the read gives the same wrong value. Clock rate makes no difference. Occasionally — and confusingly — a particular register reads back correctly.
Why the occasional correct read is the clue. That is §5 arriving as evidence. A register whose value happens to be bit-palindromic reads correctly under the wrong bit order, and a 0x00, 0xFF or 0xA5 default register value is extremely common. So "most registers wrong, some right" is not two faults or an intermittent one — it is a single deterministic fault plus symmetric data, and recognising that collapses the hypothesis space immediately.
Plausible mechanisms. A bit-order disagreement is the leading candidate. Competitors are a mode mismatch, which also corrupts deterministically, and a width or alignment error.
The discriminating observation, and it is decisive. Take a wrong value and bit-reverse it. If the result is the expected value, the diagnosis is complete — no equipment, no further capture, one operation. Nothing else produces an exact bit-reversal: a mode mismatch produces a one-position shift (Chapter 3.8 §3), not a reversal, and a width error produces misalignment across a word boundary.
That test is worth doing early precisely because it is so cheap and so specific. If the reversal does not match, bit order is eliminated outright and you have lost thirty seconds.
Why the investigation goes wrong. Because a deterministic-but-wrong value reads as a plausible-looking number, and engineers reach for the bus rather than the arithmetic. The fault is fully diagnosable from a single captured value and a calculator, and people routinely attach an analyzer first — which then displays the bytes under its own configured bit order and confirms whatever it was told, exactly as in Chapter 4.1 §9.
9. Common Misconceptions
10. Reason It Through
Work this before reading the answer.
A driver for an 8-bit I/O expander is being brought up. Writing
0x01lights the expander's output 7 instead of output 0. Writing0x80lights output 0. Writing0xFFlights all eight, and writing0x00lights none. The engineer concludes the expander's outputs are wired backwards on the board and prepares a layout change.Is that the right conclusion?
No, and the evidence already rules it out — but not for the reason most people reach for first.
Start with the two clean data points. 0x01 is 00000001 and lights output 7; 0x80 is 10000000 and lights output 0. Each is the other's bit-reversal, and the outputs are exactly swapped. That is a bit-order disagreement: the master is sending LSB-first where the expander latches MSB-first, or the reverse.
Why 0xFF and 0x00 contribute nothing. Both are palindromic, so they behave identically under either order. They are consistent with the bug, consistent with correct operation, and consistent with reversed wiring — they discriminate nothing, and the fact that they "work" is what makes the engineer believe the link is basically healthy. This is §5 operating on a real board.
So why is reversed wiring the wrong conclusion? Because it and a bit-order disagreement produce identical symptoms at the outputs. The evidence so far genuinely cannot separate them — which means the proposed board change is being made on an untested hypothesis, and if the true cause is configuration, the respin will not fix it and will instead break any unit that gets its software corrected.
What measurement separates them? Look at the wire, not the outputs. Send a strongly asymmetric value such as 0x2C and capture MOSI with the analyzer set to raw bits rather than decoded bytes. If the line carries 00110100 — the reversal — the master is transmitting LSB-first and the fault is in configuration. If it carries 00101100 as written, the master is correct and the mismatch is downstream, at which point reversed wiring or an expander configured for the other order become the live candidates.
That distinction costs one capture and decides between a software change and a board change.
The general lesson. Two faults at different layers can produce indistinguishable observations at the output, and the way out is to observe at a layer between them. Choosing an asymmetric value is what makes that observation informative at all — with 0xFF the capture would have told you nothing.
11. Understanding Check
12. Summary
Bit order decides which end of the shift register feeds the pin: MSB-first shifts left and taps the top bit, LSB-first shifts right and taps the bottom one, with the receive register filling from the opposite end in each case. SPI does not specify it, and it is not observable on the wire.
MSB-first dominates because numeric values arrive usefully under truncation, traces read the same way as datasheets, and multi-byte addresses compose legibly. The exceptions are devices where none of those reasons applies — shift-register expanders, UART-adjacent peripherals, and controllers that expose the option because it is nearly free.
In hardware the cost is two multiplexers and no flip-flops: a direction select on each shift register and a tap select on the output. The configuration must be stable across a frame, since changing it mid-word yields a result belonging to no order at all.
The property worth carrying away is that a bit-order fault is deterministic and often invisible. Bit-palindromic values — 0x00, 0xFF, 0xA5, 0x18 and twelve others — are identical under both orders, so a suite built from the usual test constants passes with the order backwards. Coverage must therefore target the data property, crossed with the configuration, not the configuration alone.
And because the fault is exactly reversible, it is diagnosable without instruments: bit-reverse a wrong value, and if the expected value appears, the question is settled. Nothing else produces a reversal — a mode mismatch shifts by one position, a width error misaligns across a word.
13. What Comes Next
Chapters 4.2 and 4.3 settled the shape of a single word: how many bits, and in what order. Chapter 4.4 — Command, Address, and Data Phases moves up a level and asks how devices build a transaction out of several of them — why a flash read looks like an opcode followed by an address followed by a data burst, how a slave tracks which phase it is in when the bus gives it no phase marker whatsoever, and the phase sequencer that structure requires in RTL.
Continue learning
Related tutorials
- Related topic
Master Input Capture and Round-Trip Delay
The complete return-path budget: clock out, peripheral response, data back, and the master's setup requirement, all inside half a period. Why maximum SCLK is a property of a whole system.
- Related topic
Deriving Mode Behaviour from CPOL and CPHA
The four SPI modes are a two-bit truth table you can rebuild in seconds. The standard numbering, the derivation, the complete mode decoder in three HDLs, and the assertions that keep a configurable design honest.
- Related topic
Mode 0 (CPOL=0, CPHA=0)
The most widely used SPI mode, and the one carrying a real implementation problem: why CPHA=0 forces the first bit onto the line before any clock edge, and the first-bit launch path in Verilog, SystemVerilog and VHDL.
- Related topic
Mode Mismatch and Its Failure Signature
What happens when the two ends disagree about the mode. The distinct signature each mismatch produces, how to tell polarity from phase disagreement from the data alone, and the monitor and coverage work that catches it.
