SPI · Module 19
SoC to SPI Peripheral Through a Register Bus
Software becomes part of the design. A configuration shadow, two sticky status bits and a toggle handshake — the three races that ship, measured at four unrelated clock ratios.
Chapters 19.1 and 19.2 had hardware decide everything — a timer, or a pin. This chapter puts a processor in the loop. It writes configuration and data, reads status, and does so on its own clock, asynchronously to the SPI engine.
Every interesting failure here is in the seam between software and hardware, and each one is fixed by a contract rather than by a circuit.
1. The Three Races
1. software writes CTRL while a transfer is running
2. software polls DONE and misses a completion
3. a multi-bit value crosses between the bus clock and the SPI clockNone of them is exotic. All three ship regularly, and all three are cheap to prevent once named.
2. Race One — The Configuration Shadow
Software may write CTRL at any time. It has no reason not to, and nothing stops it.
If the engine reads the live CTRL register, then a write landing mid-transfer changes the mode or the divider during the frame — and the transfer on the wire is neither the old configuration nor the new one. Changing CPOL mid-frame is the worst of them: Chapter 18.2 showed that leading edge is defined relative to the idle level, so inverting CPOL redefines which edges are captures.
The fix is one register:
at transfer start: cfg_shadow <= ctrl; tx_shadow <= pwdata
for the whole frame: the engine reads ONLY the shadowThe live register moves; the shadowed configuration does not
18 cycles3. Race Two — Status Bit Semantics
Three bits, three different disciplines, and the reasoning for each is worth stating because getting one wrong loses data silently.
| Bit | Kind | Why |
|---|---|---|
BUSY | level | it means what it says at the instant it is read; there is nothing to remember |
DONE | sticky, write-one-to-clear | a level DONE exists only between one transfer ending and the next beginning, so software polling slower than that sees nothing |
OVERRUN | sticky, write-one-to-clear | a sticky DONE has nowhere to record a second completion; without OVERRUN that completion is silently lost |
There is one more ordering decision, and it is invisible until it bites:
the write decode runs FIRST, then the completion
→ a completion landing in the same cycle as a write-one-to-clear WINSThe other order silently drops a completion whenever software happens to clear in that exact cycle — a race software can neither avoid nor observe.
4. Race Three — The Crossing
The register bus runs on pclk; the engine runs on sclk_ref, an unrelated clock. What crosses is multi-bit: the configuration and the transmit byte one way, the received byte the other.
WRONG a synchroniser per bit. Each bit resolves independently, so a value
that was never valid can appear on the other side.
RIGHT a request/acknowledge TOGGLE handshake, with the data held stable
across it. One synchroniser is then enough -- because the HANDSHAKE
guarantees stability, not the synchroniser.A toggle rather than a pulse, for the reason Chapter 18.7 measured: a pulse narrower than the receiving clock period is caught only by luck, and a toggle persists until the next one.
5. The Measurement
Identical output from all three languages:
shadow test: rxdata=a5 (slave sent a5) slave received 3c (master sent 3c) shadow_ctrl=21 (live ctrl=37) edges=16
sticky bits after two transfers with no clear: busy=0 done=1 overrun=1
after writing 0x02: busy=0 done=0 overrun=1
a second start while busy: done=1 overrun=1 slave received aa
sclk_ref half ratio to pclk rxdata frames edges status
7 unrelated 96 1 16 done=1 overrun=0
3 unrelated 96 1 16 done=1 overrun=0
11 unrelated 96 1 16 done=1 overrun=0
4 unrelated 96 1 16 done=1 overrun=0The shadow test checks four things, not one. The received byte is correct, the byte the slave received is correct, the shadowed configuration still reads 0x21 while the live register reads 0x37, and the wire carried exactly 16 edges. That last check matters because a divider that changed mid-frame would still produce 16 edges — at the wrong spacing — while a mode change would corrupt data with the count looking fine. Checking only the data would miss half the fault space.
A second start while busy is refused and recorded, and the transfer in flight is untouched — the slave still received 0xaa, not some mixture.
And the results do not depend on the clock ratio. Four unrelated sclk_ref half periods against a fixed 5 ns bus half period, and every received byte, frame count and edge count is identical.
6. Building It — Three HDLs
// spi_regbus_ctrl.sv
//
// Chapter 19.3 -- an SPI master behind a register bus, across two clock domains.
//
// THIS IS THE FIRST CHAPTER IN WHICH SOFTWARE IS PART OF THE DESIGN. Chapters 19.1 and 19.2 had hardware
// decide everything: a timer, or a pin. Here a processor writes configuration and data, reads status, and
// does so on ITS OWN clock, asynchronously to the SPI engine -- and the interesting failures are all in
// the seam between the two.
//
// THREE THINGS THIS MODULE EXISTS TO GET RIGHT.
//
// 1. THE SHADOW REGISTER. Software may write CTRL at any time, including while a transfer is running.
// If the engine reads the live CTRL register, the mode or the divider changes MID-FRAME and the
// transfer on the wire is neither the old configuration nor the new one. The fix is to latch the
// configuration at transfer start and use only the latched copy for the duration.
//
// The alternative is to put the burden on software -- "poll BUSY before writing CTRL" -- which is a
// real design choice and a worse one, because it is a rule that has to be obeyed by every driver
// ever written against the part, including the ones written by people who never read this sentence.
// A shadow register costs one register's worth of flops and removes the rule.
//
// 2. THE STATUS BIT SEMANTICS. `BUSY` is a LEVEL: it means what it says at the instant it is read.
// `DONE` is STICKY and cleared by writing a one to it (write-one-to-clear), because a level DONE can
// be missed entirely -- it would be asserted between one transfer ending and the next beginning, and
// software that polls slower than that sees nothing.
//
// Sticky introduces its own failure, and it needs its own bit: if a SECOND transfer completes before
// software has cleared DONE, that completion has nowhere to be recorded. `OVERRUN` exists precisely
// for that, and a status register without it silently loses completions.
//
// 3. THE CROSSING. The register bus runs on `pclk` and the SPI engine on `sclk_ref`, an unrelated
// clock. What crosses is multi-bit -- the configuration and the transmit byte one way, the received
// byte the other -- so the mechanism is a REQUEST/ACKNOWLEDGE TOGGLE HANDSHAKE with the data held
// stable across it, not a synchroniser per bit.
//
// Synchronising a multi-bit bus bit-by-bit is the classic error: each bit resolves independently and
// a value that was never valid can appear on the other side. A handshake makes the data static while
// it is sampled, which is what makes one synchroniser enough.
//
// WHAT THE SIMULATION CANNOT ESTABLISH, said here rather than implied: nothing about metastability or
// synchroniser depth. Both are two flops deep because that is correct design practice, and a single flop
// would behave identically in zero-delay RTL. What simulation DOES establish is that the handshake holds
// the data stable, that no transfer is lost or duplicated across the crossing, and that the results do
// not depend on the clock ratio -- which is the property a crossing defect would break.
`timescale 1ns/1ps
module spi_regbus_ctrl #(
parameter int AW = 8
) (
// ---- the register bus domain ----
input wire pclk,
input wire prst_n,
input wire psel,
input wire penable,
input wire pwrite,
input wire [AW-1:0] paddr,
input wire [31:0] pwdata,
output reg [31:0] prdata,
output wire pready,
output reg irq,
// ---- the SPI domain ----
input wire sclk_ref,
input wire srst_n,
output reg sclk,
output reg cs_n,
output reg mosi,
input wire miso,
// ---- observation, for the bench only ----
output wire [7:0] dbg_shadow_ctrl,
output wire [7:0] dbg_txshadow
);
// ---- the register map ----
localparam [AW-1:0] A_CTRL = 'h0, // [0] enable [1] cpol [2] cpha [7:4] divider
A_TXDATA = 'h4, // write starts a transfer
A_RXDATA = 'h8, // read-only
A_STATUS = 'hC, // [0] BUSY (level) [1] DONE (W1C) [2] OVERRUN (W1C)
A_IE = 'h10; // [1] DONE enable [2] OVERRUN enable
assign pready = 1'b1; // a single-cycle slave; no wait states
// An APB access completes in the ENABLE phase. Decoding the write in the setup phase would perform
// it twice, which for a write-one-to-clear bit means clearing something software never wrote.
wire acc_wr = psel & penable & pwrite;
wire acc_rd = psel & penable & ~pwrite;
// ---- the SPI domain's state, declared here so the bus domain can read what crosses back ----
//
// Declared above both processes rather than beside the one that writes them. Verilog resolves names
// in file order, so a cross-domain signal read before its declaration is an elaboration error -- and
// hoisting them makes the crossing's two directions visible in one place, which is what a reviewer of
// a clock-domain crossing actually needs.
reg req_s1, req_s2, req_s3;
wire req_edge = req_s2 ^ req_s3;
reg ack_tog;
reg [7:0] rx_capture;
reg [2:0] est;
localparam [2:0] E_IDLE = 3'd0, E_LEAD = 3'd1, E_SHIFT = 3'd2, E_LAG = 3'd3;
reg [7:0] ediv;
reg [3:0] cfg_div; // the shadowed divider, kept so every edge reloads from the same value
reg [3:0] eedge;
reg [7:0] esh_tx, esh_rx;
reg ecpol, ecpha;
// ================= the register-bus domain =================
reg [7:0] ctrl; // live configuration, writable at any time
reg [7:0] txdata;
reg [7:0] rxdata;
reg [7:0] ie;
reg st_done, st_over;
// The request side of the crossing: a TOGGLE, because a toggle cannot be missed at any clock ratio.
reg req_tog;
reg [7:0] cfg_shadow; // the configuration captured at transfer start
reg [7:0] tx_shadow;
assign dbg_shadow_ctrl = cfg_shadow;
assign dbg_txshadow = tx_shadow;
// The acknowledge coming back, synchronised into this domain.
reg ack_s1, ack_s2, ack_s3;
wire ack_edge = ack_s2 ^ ack_s3;
// BUSY is a level derived from the two toggles disagreeing: a request has gone out and its
// acknowledge has not come back. It needs no separate register and cannot get out of step with the
// engine, which a hand-maintained busy flag can.
wire busy = req_tog ^ ack_s2;
always_ff @(posedge pclk or negedge prst_n) begin
if (!prst_n) begin
ctrl <= 8'h00;
txdata <= 8'h00;
rxdata <= 8'h00;
ie <= 8'h00;
st_done <= 1'b0;
st_over <= 1'b0;
req_tog <= 1'b0;
cfg_shadow <= 8'h00;
tx_shadow <= 8'h00;
ack_s1 <= 1'b0;
ack_s2 <= 1'b0;
ack_s3 <= 1'b0;
prdata <= 32'h0;
irq <= 1'b0;
end else begin
ack_s1 <= ack_tog;
ack_s2 <= ack_s1;
ack_s3 <= ack_s2;
// ---- writes ----
if (acc_wr) begin
case (paddr)
A_CTRL: ctrl <= pwdata[7:0];
A_IE: ie <= pwdata[7:0];
A_TXDATA: begin
// A write to TXDATA starts a transfer -- but only if one is not already running.
// Accepting a second start while busy would either corrupt the frame in flight or
// silently discard the byte; refusing it and recording an OVERRUN says what
// happened.
if (!busy && ctrl[0]) begin
// THE SHADOW. The configuration and the data are captured HERE, once, and the
// engine sees nothing else for the whole transfer. A later CTRL write changes
// `ctrl` and cannot change `cfg_shadow`.
cfg_shadow <= ctrl;
tx_shadow <= pwdata[7:0];
txdata <= pwdata[7:0];
req_tog <= ~req_tog;
end else begin
st_over <= 1'b1;
end
end
A_STATUS: begin
// WRITE-ONE-TO-CLEAR. Writing a zero must leave a bit alone, so that a driver
// clearing DONE cannot accidentally clear OVERRUN it has not read yet.
if (pwdata[1]) st_done <= 1'b0;
if (pwdata[2]) st_over <= 1'b0;
end
default: ;
endcase
end
// ---- the transfer completing ----
//
// Ordered AFTER the write decode so that a completion landing in the same cycle as a
// write-one-to-clear is NOT lost: the completion wins, and software's next poll sees it. The
// other order silently drops a completion whenever software happens to clear in that cycle,
// which is a race software cannot avoid and cannot see.
if (ack_edge) begin
rxdata <= rx_capture;
if (st_done) st_over <= 1'b1; // a second completion with the first unread
st_done <= 1'b1;
end
// ---- reads ----
if (acc_rd) begin
case (paddr)
A_CTRL: prdata <= {24'h0, ctrl};
A_TXDATA: prdata <= {24'h0, txdata};
A_RXDATA: prdata <= {24'h0, rxdata};
A_STATUS: prdata <= {29'h0, st_over, st_done, busy};
A_IE: prdata <= {24'h0, ie};
default: prdata <= 32'h0;
endcase
end
irq <= (st_done & ie[1]) | (st_over & ie[2]);
end
end
// ================= the SPI domain =================
always_ff @(posedge sclk_ref or negedge srst_n) begin
if (!srst_n) begin
req_s1 <= 1'b0; req_s2 <= 1'b0; req_s3 <= 1'b0;
ack_tog <= 1'b0;
rx_capture <= 8'h00;
est <= E_IDLE;
ediv <= 8'h00;
cfg_div <= 4'h0;
eedge <= 4'h0;
esh_tx <= 8'h00;
esh_rx <= 8'h00;
ecpol <= 1'b0;
ecpha <= 1'b0;
sclk <= 1'b0;
cs_n <= 1'b1;
mosi <= 1'b0;
end else begin
req_s1 <= req_tog;
req_s2 <= req_s1;
req_s3 <= req_s2;
case (est)
E_IDLE: begin
if (req_edge) begin
// THE DATA IS SAMPLED HERE, from registers the other domain has held stable since
// before the toggle flipped. That is what makes one synchroniser enough for a
// multi-bit value: the handshake, not the synchroniser, guarantees stability.
ecpol <= dbg_shadow_ctrl[1];
ecpha <= dbg_shadow_ctrl[2];
cfg_div <= dbg_shadow_ctrl[7:4];
ediv <= {4'h0, dbg_shadow_ctrl[7:4]};
esh_tx <= dbg_txshadow;
esh_rx <= 8'h00;
sclk <= dbg_shadow_ctrl[1];
cs_n <= 1'b0;
mosi <= dbg_txshadow[7];
eedge <= 4'h0;
est <= E_LEAD;
end
end
E_LEAD: begin
est <= E_SHIFT;
end
E_SHIFT: begin
// One SCLK edge per reference cycle when the divider is zero; otherwise one per
// (divider + 1). The divider is deliberately part of the shadowed configuration, so
// changing it mid-transfer is impossible by construction rather than by convention.
if (ediv == 8'h00) begin
// THE DIVIDER RELOADS AFTER EVERY EDGE. The first version loaded it once at
// transfer start and decremented to zero, after which edges came out one per
// reference cycle forever -- so the divider delayed the FIRST edge and did
// nothing else. The frame still carried 16 edges and the received byte was
// still correct at low rates, so nothing failed until the slave ran out of
// setup margin. A divider that only works at one setting is worse than none.
ediv <= {4'h0, cfg_div};
if (sclk == ecpol) begin
// the leading edge: capture for CPHA = 0
if (!ecpha) esh_rx <= {esh_rx[6:0], miso};
sclk <= ~ecpol;
end else begin
if (ecpha) esh_rx <= {esh_rx[6:0], miso};
sclk <= ecpol;
esh_tx <= {esh_tx[6:0], 1'b0};
mosi <= esh_tx[6];
end
if (eedge == 4'hF) begin
est <= E_LAG;
end else begin
eedge <= eedge + 4'h1;
end
end else begin
ediv <= ediv - 8'h01;
end
end
E_LAG: begin
cs_n <= 1'b1;
rx_capture <= esh_rx;
ack_tog <= ~ack_tog;
est <= E_IDLE;
end
default: est <= E_IDLE;
endcase
end
end
endmodule// spi_regbus_ctrl.v
//
// Chapter 19.3 -- an SPI master behind a register bus, across two clock domains.
//
// THIS IS THE FIRST CHAPTER IN WHICH SOFTWARE IS PART OF THE DESIGN. Chapters 19.1 and 19.2 had hardware
// decide everything: a timer, or a pin. Here a processor writes configuration and data, reads status, and
// does so on ITS OWN clock, asynchronously to the SPI engine -- and the interesting failures are all in
// the seam between the two.
//
// THREE THINGS THIS MODULE EXISTS TO GET RIGHT.
//
// 1. THE SHADOW REGISTER. Software may write CTRL at any time, including while a transfer is running.
// If the engine reads the live CTRL register, the mode or the divider changes MID-FRAME and the
// transfer on the wire is neither the old configuration nor the new one. The fix is to latch the
// configuration at transfer start and use only the latched copy for the duration.
//
// The alternative is to put the burden on software -- "poll BUSY before writing CTRL" -- which is a
// real design choice and a worse one, because it is a rule that has to be obeyed by every driver
// ever written against the part, including the ones written by people who never read this sentence.
// A shadow register costs one register's worth of flops and removes the rule.
//
// 2. THE STATUS BIT SEMANTICS. `BUSY` is a LEVEL: it means what it says at the instant it is read.
// `DONE` is STICKY and cleared by writing a one to it (write-one-to-clear), because a level DONE can
// be missed entirely -- it would be asserted between one transfer ending and the next beginning, and
// software that polls slower than that sees nothing.
//
// Sticky introduces its own failure, and it needs its own bit: if a SECOND transfer completes before
// software has cleared DONE, that completion has nowhere to be recorded. `OVERRUN` exists precisely
// for that, and a status register without it silently loses completions.
//
// 3. THE CROSSING. The register bus runs on `pclk` and the SPI engine on `sclk_ref`, an unrelated
// clock. What crosses is multi-bit -- the configuration and the transmit byte one way, the received
// byte the other -- so the mechanism is a REQUEST/ACKNOWLEDGE TOGGLE HANDSHAKE with the data held
// stable across it, not a synchroniser per bit.
//
// Synchronising a multi-bit bus bit-by-bit is the classic error: each bit resolves independently and
// a value that was never valid can appear on the other side. A handshake makes the data static while
// it is sampled, which is what makes one synchroniser enough.
//
// WHAT THE SIMULATION CANNOT ESTABLISH, said here rather than implied: nothing about metastability or
// synchroniser depth. Both are two flops deep because that is correct design practice, and a single flop
// would behave identically in zero-delay RTL. What simulation DOES establish is that the handshake holds
// the data stable, that no transfer is lost or duplicated across the crossing, and that the results do
// not depend on the clock ratio -- which is the property a crossing defect would break.
`timescale 1ns/1ps
module spi_regbus_ctrl #(
parameter AW = 8
) (
// ---- the register bus domain ----
input wire pclk,
input wire prst_n,
input wire psel,
input wire penable,
input wire pwrite,
input wire [AW-1:0] paddr,
input wire [31:0] pwdata,
output reg [31:0] prdata,
output wire pready,
output reg irq,
// ---- the SPI domain ----
input wire sclk_ref,
input wire srst_n,
output reg sclk,
output reg cs_n,
output reg mosi,
input wire miso,
// ---- observation, for the bench only ----
output wire [7:0] dbg_shadow_ctrl,
output wire [7:0] dbg_txshadow
);
// ---- the register map ----
localparam [AW-1:0] A_CTRL = 'h0, // [0] enable [1] cpol [2] cpha [7:4] divider
A_TXDATA = 'h4, // write starts a transfer
A_RXDATA = 'h8, // read-only
A_STATUS = 'hC, // [0] BUSY (level) [1] DONE (W1C) [2] OVERRUN (W1C)
A_IE = 'h10; // [1] DONE enable [2] OVERRUN enable
assign pready = 1'b1; // a single-cycle slave; no wait states
// An APB access completes in the ENABLE phase. Decoding the write in the setup phase would perform
// it twice, which for a write-one-to-clear bit means clearing something software never wrote.
wire acc_wr = psel & penable & pwrite;
wire acc_rd = psel & penable & ~pwrite;
// ---- the SPI domain's state, declared here so the bus domain can read what crosses back ----
//
// Declared above both processes rather than beside the one that writes them. Verilog resolves names
// in file order, so a cross-domain signal read before its declaration is an elaboration error -- and
// hoisting them makes the crossing's two directions visible in one place, which is what a reviewer of
// a clock-domain crossing actually needs.
reg req_s1, req_s2, req_s3;
wire req_edge = req_s2 ^ req_s3;
reg ack_tog;
reg [7:0] rx_capture;
reg [2:0] est;
localparam [2:0] E_IDLE = 3'd0, E_LEAD = 3'd1, E_SHIFT = 3'd2, E_LAG = 3'd3;
reg [7:0] ediv;
reg [3:0] cfg_div; // the shadowed divider, kept so every edge reloads from the same value
reg [3:0] eedge;
reg [7:0] esh_tx, esh_rx;
reg ecpol, ecpha;
// ================= the register-bus domain =================
reg [7:0] ctrl; // live configuration, writable at any time
reg [7:0] txdata;
reg [7:0] rxdata;
reg [7:0] ie;
reg st_done, st_over;
// The request side of the crossing: a TOGGLE, because a toggle cannot be missed at any clock ratio.
reg req_tog;
reg [7:0] cfg_shadow; // the configuration captured at transfer start
reg [7:0] tx_shadow;
assign dbg_shadow_ctrl = cfg_shadow;
assign dbg_txshadow = tx_shadow;
// The acknowledge coming back, synchronised into this domain.
reg ack_s1, ack_s2, ack_s3;
wire ack_edge = ack_s2 ^ ack_s3;
// BUSY is a level derived from the two toggles disagreeing: a request has gone out and its
// acknowledge has not come back. It needs no separate register and cannot get out of step with the
// engine, which a hand-maintained busy flag can.
wire busy = req_tog ^ ack_s2;
always @(posedge pclk or negedge prst_n) begin
if (!prst_n) begin
ctrl <= 8'h00;
txdata <= 8'h00;
rxdata <= 8'h00;
ie <= 8'h00;
st_done <= 1'b0;
st_over <= 1'b0;
req_tog <= 1'b0;
cfg_shadow <= 8'h00;
tx_shadow <= 8'h00;
ack_s1 <= 1'b0;
ack_s2 <= 1'b0;
ack_s3 <= 1'b0;
prdata <= 32'h0;
irq <= 1'b0;
end else begin
ack_s1 <= ack_tog;
ack_s2 <= ack_s1;
ack_s3 <= ack_s2;
// ---- writes ----
if (acc_wr) begin
case (paddr)
A_CTRL: ctrl <= pwdata[7:0];
A_IE: ie <= pwdata[7:0];
A_TXDATA: begin
// A write to TXDATA starts a transfer -- but only if one is not already running.
// Accepting a second start while busy would either corrupt the frame in flight or
// silently discard the byte; refusing it and recording an OVERRUN says what
// happened.
if (!busy && ctrl[0]) begin
// THE SHADOW. The configuration and the data are captured HERE, once, and the
// engine sees nothing else for the whole transfer. A later CTRL write changes
// `ctrl` and cannot change `cfg_shadow`.
cfg_shadow <= ctrl;
tx_shadow <= pwdata[7:0];
txdata <= pwdata[7:0];
req_tog <= ~req_tog;
end else begin
st_over <= 1'b1;
end
end
A_STATUS: begin
// WRITE-ONE-TO-CLEAR. Writing a zero must leave a bit alone, so that a driver
// clearing DONE cannot accidentally clear OVERRUN it has not read yet.
if (pwdata[1]) st_done <= 1'b0;
if (pwdata[2]) st_over <= 1'b0;
end
default: ;
endcase
end
// ---- the transfer completing ----
//
// Ordered AFTER the write decode so that a completion landing in the same cycle as a
// write-one-to-clear is NOT lost: the completion wins, and software's next poll sees it. The
// other order silently drops a completion whenever software happens to clear in that cycle,
// which is a race software cannot avoid and cannot see.
if (ack_edge) begin
rxdata <= rx_capture;
if (st_done) st_over <= 1'b1; // a second completion with the first unread
st_done <= 1'b1;
end
// ---- reads ----
if (acc_rd) begin
case (paddr)
A_CTRL: prdata <= {24'h0, ctrl};
A_TXDATA: prdata <= {24'h0, txdata};
A_RXDATA: prdata <= {24'h0, rxdata};
A_STATUS: prdata <= {29'h0, st_over, st_done, busy};
A_IE: prdata <= {24'h0, ie};
default: prdata <= 32'h0;
endcase
end
irq <= (st_done & ie[1]) | (st_over & ie[2]);
end
end
// ================= the SPI domain =================
always @(posedge sclk_ref or negedge srst_n) begin
if (!srst_n) begin
req_s1 <= 1'b0; req_s2 <= 1'b0; req_s3 <= 1'b0;
ack_tog <= 1'b0;
rx_capture <= 8'h00;
est <= E_IDLE;
ediv <= 8'h00;
cfg_div <= 4'h0;
eedge <= 4'h0;
esh_tx <= 8'h00;
esh_rx <= 8'h00;
ecpol <= 1'b0;
ecpha <= 1'b0;
sclk <= 1'b0;
cs_n <= 1'b1;
mosi <= 1'b0;
end else begin
req_s1 <= req_tog;
req_s2 <= req_s1;
req_s3 <= req_s2;
case (est)
E_IDLE: begin
if (req_edge) begin
// THE DATA IS SAMPLED HERE, from registers the other domain has held stable since
// before the toggle flipped. That is what makes one synchroniser enough for a
// multi-bit value: the handshake, not the synchroniser, guarantees stability.
ecpol <= dbg_shadow_ctrl[1];
ecpha <= dbg_shadow_ctrl[2];
cfg_div <= dbg_shadow_ctrl[7:4];
ediv <= {4'h0, dbg_shadow_ctrl[7:4]};
esh_tx <= dbg_txshadow;
esh_rx <= 8'h00;
sclk <= dbg_shadow_ctrl[1];
cs_n <= 1'b0;
mosi <= dbg_txshadow[7];
eedge <= 4'h0;
est <= E_LEAD;
end
end
E_LEAD: begin
est <= E_SHIFT;
end
E_SHIFT: begin
// One SCLK edge per reference cycle when the divider is zero; otherwise one per
// (divider + 1). The divider is deliberately part of the shadowed configuration, so
// changing it mid-transfer is impossible by construction rather than by convention.
if (ediv == 8'h00) begin
// THE DIVIDER RELOADS AFTER EVERY EDGE. The first version loaded it once at
// transfer start and decremented to zero, after which edges came out one per
// reference cycle forever -- so the divider delayed the FIRST edge and did
// nothing else. The frame still carried 16 edges and the received byte was
// still correct at low rates, so nothing failed until the slave ran out of
// setup margin. A divider that only works at one setting is worse than none.
ediv <= {4'h0, cfg_div};
if (sclk == ecpol) begin
// the leading edge: capture for CPHA = 0
if (!ecpha) esh_rx <= {esh_rx[6:0], miso};
sclk <= ~ecpol;
end else begin
if (ecpha) esh_rx <= {esh_rx[6:0], miso};
sclk <= ecpol;
esh_tx <= {esh_tx[6:0], 1'b0};
mosi <= esh_tx[6];
end
if (eedge == 4'hF) begin
est <= E_LAG;
end else begin
eedge <= eedge + 4'h1;
end
end else begin
ediv <= ediv - 8'h01;
end
end
E_LAG: begin
cs_n <= 1'b1;
rx_capture <= esh_rx;
ack_tog <= ~ack_tog;
est <= E_IDLE;
end
default: est <= E_IDLE;
endcase
end
end
endmodule-- spi_regbus_ctrl.vhd
--
-- Chapter 19.3 -- an SPI master behind a register bus, across two clock domains.
--
-- THIS IS THE FIRST CHAPTER IN WHICH SOFTWARE IS PART OF THE DESIGN. Chapters 19.1 and 19.2 had hardware
-- decide everything: a timer, or a pin. Here a processor writes configuration and data, reads status, and
-- does so on ITS OWN clock, asynchronously to the SPI engine -- and the interesting failures are all in
-- the seam between the two.
--
-- THREE THINGS THIS MODULE EXISTS TO GET RIGHT.
--
-- 1. THE SHADOW REGISTER. Software may write CTRL at any time, including while a transfer is running.
-- If the engine reads the live CTRL register, the mode or the divider changes MID-FRAME and the
-- transfer on the wire is neither the old configuration nor the new one. The fix is to latch the
-- configuration at transfer start and use only the latched copy for the duration.
--
-- The alternative is to put the burden on software -- "poll BUSY before writing CTRL" -- which is a
-- real design choice and a worse one, because it is a rule that has to be obeyed by every driver
-- ever written against the part, including the ones written by people who never read this sentence.
-- A shadow register costs one register's worth of flops and removes the rule.
--
-- 2. THE STATUS BIT SEMANTICS. `BUSY` is a LEVEL: it means what it says at the instant it is read.
-- `DONE` is STICKY and cleared by writing a one to it (write-one-to-clear), because a level DONE can
-- be missed entirely -- it would be asserted between one transfer ending and the next beginning, and
-- software that polls slower than that sees nothing.
--
-- Sticky introduces its own failure, and it needs its own bit: if a SECOND transfer completes before
-- software has cleared DONE, that completion has nowhere to be recorded. `OVERRUN` exists precisely
-- for that, and a status register without it silently loses completions.
--
-- 3. THE CROSSING. The register bus runs on `pclk` and the SPI engine on `sclk_ref`, an unrelated
-- clock. What crosses is multi-bit -- the configuration and the transmit byte one way, the received
-- byte the other -- so the mechanism is a REQUEST/ACKNOWLEDGE TOGGLE HANDSHAKE with the data held
-- stable across it, not a synchroniser per bit.
--
-- Synchronising a multi-bit bus bit-by-bit is the classic error: each bit resolves independently and
-- a value that was never valid can appear on the other side. A handshake makes the data static while
-- it is sampled, which is what makes one synchroniser enough.
--
-- WHAT THE SIMULATION CANNOT ESTABLISH, said here rather than implied: nothing about metastability or
-- synchroniser depth. Both are two flops deep because that is correct design practice, and a single flop
-- would behave identically in zero-delay RTL. What simulation DOES establish is that the handshake holds
-- the data stable, that no transfer is lost or duplicated across the crossing, and that the results do
-- not depend on the clock ratio -- which is the property a crossing defect would break.
--
-- WHAT THE VHDL VERSION ADDS, and it is structural rather than cosmetic.
--
-- EVERY REGISTER HERE IS A SIGNAL, not a process variable. Chapter 19.2 found four separate defects that
-- all came from the same fact -- a process variable is updated immediately and a non-blocking reg is not --
-- and each one produced a design that behaved plausibly. Using signals throughout makes the semantics
-- match the SystemVerilog and Verilog versions by construction rather than by care, which is the right
-- trade for a module whose whole subject is two clock domains disagreeing about time.
--
-- The register addresses are named constants of a constrained subtype, so an access to an address outside
-- the map is a decode miss rather than an aliased hit.
--
-- IDENTIFIER REVIEW (VHDL IS CASE-INSENSITIVE). The bus ports are `psel`, `penable`, `pwrite`, `paddr`,
-- `pwdata`, `prdata`; the internal registers are `ctrl_r`, `tx_r`, `rx_r`, `ie_r` -- suffixed rather than
-- case-varied, so nothing is distinguished from a port by case alone. `A_CTRL` and friends are constants
-- and no signal shares their spelling.
--
-- RESERVED-WORD REVIEW. Nothing is named `label`, `range`, `next`, `access`, `body`, `bus`, `register`,
-- `open`, `guarded`, `block`, `new`, `abs`, `rem`, `mod`, `exit`, `file`, `group` or `literal`. The last
-- one is worth checking in a register-file design, where `register` and `literal` are tempting names.
--
-- RANGE-DIRECTION REVIEW. Every vector is `downto`; the shift registers are constrained subtypes; and the
-- bus read data is assembled with explicit zero padding rather than by concatenating a literal, so no
-- slice inherits an ascending range.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
package spi_regbus_pkg is
subtype byte_t is std_logic_vector(7 downto 0);
type eng_state_t is (E_IDLE, E_LEAD, E_SHIFT, E_LAG);
end package spi_regbus_pkg;
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
use work.spi_regbus_pkg.all;
entity spi_regbus_ctrl is
generic (
AW : positive := 8
);
port (
-- the register bus domain
pclk : in std_logic;
prst_n : in std_logic;
psel : in std_logic;
penable : in std_logic;
pwrite : in std_logic;
paddr : in std_logic_vector(AW - 1 downto 0);
pwdata : in std_logic_vector(31 downto 0);
prdata : out std_logic_vector(31 downto 0);
pready : out std_logic;
irq : out std_logic;
-- the SPI domain
sclk_ref : in std_logic;
srst_n : in std_logic;
sclk : out std_logic;
cs_n : out std_logic;
mosi : out std_logic;
miso : in std_logic;
-- observation, for the bench only
dbg_shadow_ctrl : out byte_t;
dbg_txshadow : out byte_t
);
end entity spi_regbus_ctrl;
architecture rtl of spi_regbus_ctrl is
constant A_CTRL : std_logic_vector(AW - 1 downto 0) := std_logic_vector(to_unsigned(16#00#, AW));
constant A_TXDATA : std_logic_vector(AW - 1 downto 0) := std_logic_vector(to_unsigned(16#04#, AW));
constant A_RXDATA : std_logic_vector(AW - 1 downto 0) := std_logic_vector(to_unsigned(16#08#, AW));
constant A_STATUS : std_logic_vector(AW - 1 downto 0) := std_logic_vector(to_unsigned(16#0C#, AW));
constant A_IE : std_logic_vector(AW - 1 downto 0) := std_logic_vector(to_unsigned(16#10#, AW));
-- ---- the register-bus domain ----
signal ctrl_r, tx_r, rx_r, ie_r : byte_t := (others => '0');
signal st_done, st_over : std_logic := '0';
signal req_tog : std_logic := '0';
signal cfg_shadow, tx_shadow : byte_t := (others => '0');
signal ack_sr : std_logic_vector(2 downto 0) := (others => '0');
signal prdata_r, irq_r : std_logic := '0';
signal prd_r : std_logic_vector(31 downto 0) := (others => '0');
-- ---- the SPI domain, declared here so the bus domain can read what crosses back ----
signal req_sr : std_logic_vector(2 downto 0) := (others => '0');
signal ack_tog : std_logic := '0';
signal rx_capture : byte_t := (others => '0');
signal est : eng_state_t := E_IDLE;
signal ediv : unsigned(7 downto 0) := (others => '0');
signal cfg_div : std_logic_vector(3 downto 0) := (others => '0');
signal eedge : unsigned(3 downto 0) := (others => '0');
signal esh_tx, esh_rx : byte_t := (others => '0');
signal ecpol, ecpha : std_logic := '0';
signal sclk_r, cs_r, mosi_r : std_logic := '0';
-- An APB access completes in the ENABLE phase. Decoding a write in the setup phase performs it twice,
-- which for a write-one-to-clear bit means clearing something software never wrote.
signal acc_wr, acc_rd : std_logic;
-- BUSY is a level derived from the two toggles disagreeing: a request has gone out and its acknowledge
-- has not come back. It needs no separate register and cannot get out of step with the engine, which a
-- hand-maintained busy flag can.
signal busy_w : std_logic;
signal ack_edge_w, req_edge_w : std_logic;
begin
pready <= '1'; -- a single-cycle slave; no wait states
prdata <= prd_r;
irq <= irq_r;
sclk <= sclk_r;
cs_n <= cs_r;
mosi <= mosi_r;
dbg_shadow_ctrl <= cfg_shadow;
dbg_txshadow <= tx_shadow;
acc_wr <= psel and penable and pwrite;
acc_rd <= psel and penable and (not pwrite);
busy_w <= req_tog xor ack_sr(1);
ack_edge_w <= ack_sr(1) xor ack_sr(2);
req_edge_w <= req_sr(1) xor req_sr(2);
-- ================= the register-bus domain =================
bus_dom : process (pclk, prst_n) is
begin
if prst_n = '0' then
ctrl_r <= (others => '0'); tx_r <= (others => '0');
rx_r <= (others => '0'); ie_r <= (others => '0');
st_done <= '0'; st_over <= '0';
req_tog <= '0';
cfg_shadow <= (others => '0'); tx_shadow <= (others => '0');
ack_sr <= (others => '0');
prd_r <= (others => '0');
irq_r <= '0';
elsif rising_edge(pclk) then
ack_sr <= ack_sr(1 downto 0) & ack_tog;
-- ---- writes ----
if acc_wr = '1' then
if paddr = A_CTRL then
ctrl_r <= pwdata(7 downto 0);
elsif paddr = A_IE then
ie_r <= pwdata(7 downto 0);
elsif paddr = A_TXDATA then
-- A write to TXDATA starts a transfer, but only if one is not already running.
-- Accepting a second start while busy would either corrupt the frame in flight or
-- silently discard the byte; refusing it and recording an OVERRUN says what happened.
if busy_w = '0' and ctrl_r(0) = '1' then
-- THE SHADOW. Configuration and data are captured HERE, once, and the engine sees
-- nothing else for the whole transfer. A later CTRL write changes `ctrl_r` and
-- cannot change `cfg_shadow`.
cfg_shadow <= ctrl_r;
tx_shadow <= pwdata(7 downto 0);
tx_r <= pwdata(7 downto 0);
req_tog <= not req_tog;
else
st_over <= '1';
end if;
elsif paddr = A_STATUS then
-- WRITE-ONE-TO-CLEAR. Writing a zero must leave a bit alone, so a driver clearing
-- DONE cannot accidentally clear an OVERRUN it has not read yet.
if pwdata(1) = '1' then st_done <= '0'; end if;
if pwdata(2) = '1' then st_over <= '0'; end if;
end if;
end if;
-- ---- the transfer completing ----
--
-- Ordered AFTER the write decode so that a completion landing in the same cycle as a
-- write-one-to-clear is NOT lost: the completion wins and software's next poll sees it. The
-- other order silently drops a completion whenever software happens to clear in that cycle,
-- which is a race software can neither avoid nor see.
if ack_edge_w = '1' then
rx_r <= rx_capture;
if st_done = '1' then
st_over <= '1'; -- a second completion with the first unread
end if;
st_done <= '1';
end if;
-- ---- reads ----
if acc_rd = '1' then
if paddr = A_CTRL then
prd_r <= x"000000" & ctrl_r;
elsif paddr = A_TXDATA then
prd_r <= x"000000" & tx_r;
elsif paddr = A_RXDATA then
prd_r <= x"000000" & rx_r;
elsif paddr = A_STATUS then
prd_r <= x"000000" & "00000" & st_over & st_done & busy_w;
elsif paddr = A_IE then
prd_r <= x"000000" & ie_r;
else
prd_r <= (others => '0');
end if;
end if;
irq_r <= (st_done and ie_r(1)) or (st_over and ie_r(2));
end if;
end process bus_dom;
-- ================= the SPI domain =================
spi_dom : process (sclk_ref, srst_n) is
begin
if srst_n = '0' then
req_sr <= (others => '0');
ack_tog <= '0';
rx_capture <= (others => '0');
est <= E_IDLE;
ediv <= (others => '0');
cfg_div <= (others => '0');
eedge <= (others => '0');
esh_tx <= (others => '0'); esh_rx <= (others => '0');
ecpol <= '0'; ecpha <= '0';
sclk_r <= '0'; cs_r <= '1'; mosi_r <= '0';
elsif rising_edge(sclk_ref) then
req_sr <= req_sr(1 downto 0) & req_tog;
case est is
when E_IDLE =>
if req_edge_w = '1' then
-- THE DATA IS SAMPLED HERE, from registers the other domain has held stable since
-- before the toggle flipped. That is what makes one synchroniser enough for a
-- multi-bit value: the handshake, not the synchroniser, guarantees stability.
ecpol <= cfg_shadow(1);
ecpha <= cfg_shadow(2);
cfg_div <= cfg_shadow(7 downto 4);
ediv <= unsigned(x"0" & cfg_shadow(7 downto 4));
esh_tx <= tx_shadow;
esh_rx <= (others => '0');
sclk_r <= cfg_shadow(1);
cs_r <= '0';
mosi_r <= tx_shadow(7);
eedge <= (others => '0');
est <= E_LEAD;
end if;
when E_LEAD =>
est <= E_SHIFT;
when E_SHIFT =>
if ediv = 0 then
-- The divider RELOADS after every edge. Loading it once and decrementing to zero
-- makes it delay the first edge and nothing else, which still produces 16 edges
-- and a correct byte until the slave runs out of setup margin.
ediv <= unsigned(x"0" & cfg_div);
if sclk_r = ecpol then
if ecpha = '0' then
esh_rx <= esh_rx(6 downto 0) & miso;
end if;
sclk_r <= not ecpol;
else
if ecpha = '1' then
esh_rx <= esh_rx(6 downto 0) & miso;
end if;
sclk_r <= ecpol;
-- The read comes before the shift in the OTHER languages because a reg's
-- non-blocking shift is not visible in the same cycle. Here both are signal
-- assignments, so `esh_tx` still reads its pre-edge value and the order does
-- not matter -- which is exactly why signals are used throughout this file.
mosi_r <= esh_tx(6);
esh_tx <= esh_tx(6 downto 0) & '0';
end if;
if eedge = 15 then
est <= E_LAG;
else
eedge <= eedge + 1;
end if;
else
ediv <= ediv - 1;
end if;
when E_LAG =>
cs_r <= '1';
rx_capture <= esh_rx;
ack_tog <= not ack_tog;
est <= E_IDLE;
end case;
end if;
end process spi_dom;
end architecture rtl;The Bench
The bench drives the register bus the way a driver would — through APB accesses, with no visibility into the engine — and models the slave on the pins. The pin monitor counts SCLK edges independently of the design, so the frame had 16 edges is an observation rather than a claim.
// spi_regbus_ctrl_tb.sv
//
// SOFTWARE ON ONE CLOCK, AN SPI ENGINE ON ANOTHER, AND THE THREE RACES THAT SHIP.
//
// The bench drives the register bus the way a driver would -- writes and reads through APB accesses, with
// no visibility into the engine -- and models the slave on the pins. The two clocks are unrelated and
// their ratio is swept, because a crossing defect is a phase effect and a single ratio cannot find one.
//
// THE FOUR RESULTS.
//
// 1. THE SHADOW REGISTER MAKES A MID-TRANSFER CONFIGURATION WRITE HARMLESS. Software writes CTRL with a
// different mode AND a different divider while a transfer is in flight. The bench requires the
// received byte to be correct, the shadowed configuration to be UNCHANGED, and the number of SCLK
// edges on the wire to be exactly 16 -- because a divider that changed mid-frame would still produce
// 16 edges at the wrong spacing, and a mode that changed would corrupt the data while the count
// looked fine. Both are checked.
//
// 2. WRITE-ONE-TO-CLEAR MEANS WRITING A ZERO LEAVES A BIT ALONE. The bench sets both sticky bits, writes
// a word with only one bit set, and requires the other to survive. A status register that clears on
// any write destroys a bit the driver had not read yet, and the symptom is a lost completion.
//
// 3. A SECOND COMPLETION WITH THE FIRST UNREAD MUST BE RECORDED, NOT DISCARDED. Two transfers are run
// without clearing DONE between them, and OVERRUN must be set. A sticky DONE with no OVERRUN silently
// loses completions, which is the failure a level DONE was replaced to avoid -- traded for a
// different one rather than fixed.
//
// 4. THE RESULTS DO NOT DEPEND ON THE CLOCK RATIO. The same traffic is run at four unrelated
// pclk-to-sclk_ref ratios and every received byte, every status word and every transfer count must be
// IDENTICAL. That is rate invariance, and it is the property a multi-bit crossing defect breaks --
// the one thing a simulation can genuinely establish about a crossing.
`timescale 1ns/1ps
module spi_regbus_ctrl_tb;
localparam int AW = 8;
localparam [AW-1:0] A_CTRL = 'h0, A_TXDATA = 'h4, A_RXDATA = 'h8,
A_STATUS = 'hC, A_IE = 'h10;
// ---- the register-bus clock: fixed ----
reg pclk = 1'b0;
always #5 pclk = ~pclk;
reg prst_n = 1'b1;
// ---- the SPI reference clock: its period is the swept variable ----
reg sclk_ref = 1'b0;
reg srst_n = 1'b1;
integer sref_half = 7; // half period in ns; deliberately not a divisor of 10
always begin
#(sref_half) sclk_ref = ~sclk_ref;
end
reg psel = 1'b0, penable = 1'b0, pwrite = 1'b0;
// `'0` is a SystemVerilog literal and does not survive the Verilog-2001 conversion; an
// explicit width is the only spelling that means the same thing in both.
reg [AW-1:0] paddr = {AW{1'b0}};
reg [31:0] pwdata = 32'h0;
wire [31:0] prdata;
wire pready, irq;
wire sclk, cs_n, mosi;
reg miso = 1'b0;
wire [7:0] dbg_shadow_ctrl, dbg_txshadow;
spi_regbus_ctrl #(.AW(AW)) dut (
.pclk(pclk), .prst_n(prst_n), .psel(psel), .penable(penable), .pwrite(pwrite),
.paddr(paddr), .pwdata(pwdata), .prdata(prdata), .pready(pready), .irq(irq),
.sclk_ref(sclk_ref), .srst_n(srst_n),
.sclk(sclk), .cs_n(cs_n), .mosi(mosi), .miso(miso),
.dbg_shadow_ctrl(dbg_shadow_ctrl), .dbg_txshadow(dbg_txshadow)
);
integer errors = 0;
// ------------------------------------------------------------------
// THE SLAVE MODEL, on the pins. Mode 0: presents the MSB from the select and advances on the trailing
// edge, so the master's leading edge always has a full half period of setup. It knows nothing about
// the register bus.
// ------------------------------------------------------------------
reg [7:0] slave_byte = 8'h00;
reg [7:0] slv_sh = 8'h00;
reg [7:0] slv_rx = 8'h00; // what the slave received, for the transmit-path check
reg cs_d = 1'b1, sclk_d = 1'b0;
// Edge counting on the wire, so the bench can see the frame's shape without asking the design.
integer edge_count, frames_seen;
// CLOCKED IN THE SPI DOMAIN, NOT ON THE BUS CLOCK.
//
// The first version sampled the pins on `pclk`. That works only while SCLK is slower than the bus
// clock, and the whole point of the rate sweep is to run configurations where it is not: at an
// sclk_ref half period of 3 ns against a 5 ns bus half period the monitor saw 4 edges out of 16 and
// the slave model returned a byte assembled from the edges it happened to catch. Both failures looked
// exactly like a broken crossing in the design.
//
// A real slave is clocked by SCLK, so sampling in the SPI reference domain is both more faithful and
// immune to undersampling -- SCLK changes only on an sclk_ref edge. The alternative, separate
// processes triggered on each pin edge, needs several writers for `miso` and the counters, which is a
// race in Verilog and a resolution problem in VHDL.
always @(posedge sclk_ref or negedge srst_n) begin
if (!srst_n) begin
slv_sh <= 8'h00; slv_rx <= 8'h00; cs_d <= 1'b1; sclk_d <= 1'b0;
edge_count <= 0; frames_seen <= 0;
miso <= 1'b0;
end else begin
// Delayed copies as NON-BLOCKING regs, so the edge detection lags by one reference cycle in
// every language. Chapter 19.2 found four faults in this exact distinction.
cs_d <= cs_n;
sclk_d <= sclk;
if (cs_d && !cs_n) begin
slv_sh <= slave_byte;
miso <= slave_byte[7];
slv_rx <= 8'h00;
edge_count <= 0;
end else if (!cs_n && (sclk_d !== sclk)) begin
edge_count <= edge_count + 1;
if (!sclk_d && sclk) begin
// leading edge: the slave samples MOSI
slv_rx <= {slv_rx[6:0], mosi};
end else begin
// trailing edge: the slave advances its own data
slv_sh <= {slv_sh[6:0], 1'b0};
miso <= slv_sh[6];
end
end else if (!cs_d && cs_n) begin
frames_seen <= frames_seen + 1;
end
end
end
// ------------------------------------------------------------------
// APB accesses, written the way a driver would issue them.
// ------------------------------------------------------------------
task automatic apb_write(input [AW-1:0] a, input [31:0] d);
begin
@(negedge pclk);
psel = 1'b1; pwrite = 1'b1; paddr = a; pwdata = d; penable = 1'b0;
@(negedge pclk);
penable = 1'b1; // the access completes in the ENABLE phase
@(negedge pclk);
psel = 1'b0; penable = 1'b0; pwrite = 1'b0;
end
endtask
reg [31:0] rdbuf;
task automatic apb_read(input [AW-1:0] a);
begin
@(negedge pclk);
psel = 1'b1; pwrite = 1'b0; paddr = a; penable = 1'b0;
@(negedge pclk);
penable = 1'b1;
@(posedge pclk);
@(negedge pclk);
rdbuf = prdata;
psel = 1'b0; penable = 1'b0;
end
endtask
// Wait for BUSY to clear, bounded. An unbounded wait on a condition a broken design never satisfies
// is a hang, and a hung regression reports nothing at all.
integer guard;
task automatic wait_idle;
begin
guard = 0;
apb_read(A_STATUS);
while (rdbuf[0] && guard < 4000) begin
apb_read(A_STATUS);
guard = guard + 1;
end
if (guard >= 4000) begin
$display(" FAIL: BUSY never cleared");
errors = errors + 1;
end
end
endtask
task automatic reset_all;
begin
@(negedge pclk); prst_n = 1'b0; srst_n = 1'b0;
repeat (6) @(negedge pclk);
prst_n = 1'b1; srst_n = 1'b1;
repeat (4) @(negedge pclk);
end
endtask
integer k, x_reports, mutations;
integer RATIOS [0:3];
reg [7:0] res_rx [0:3];
integer res_frames[0:3], res_edges[0:3];
reg [7:0] got_rx;
integer got_edges;
always @(posedge pclk) if (prst_n) begin
if ((^prdata === 32'bx) || (^dbg_shadow_ctrl === 8'bx)) x_reports = x_reports + 1;
end
initial begin
x_reports = 0; mutations = 0;
RATIOS[0] = 7; RATIOS[1] = 3; RATIOS[2] = 11; RATIOS[3] = 4;
// ================= 1. the shadow register =================
sref_half = 7;
reset_all;
slave_byte = 8'hA5;
// Divider 2, not 0. This slave model presents each bit two reference cycles after the trailing
// edge -- one to detect it, one for the registered output -- so the SCLK half period needs at
// least three reference cycles for the master's leading edge to have any setup. At divider 0 the
// master captures the MSB twice and loses the last bit, which is 19.1's boundary again rather
// than anything about the register bus.
apb_write(A_CTRL, 32'h21); // enable, mode 0, divider 2
apb_write(A_TXDATA, 32'h3C); // start a transfer
// Mid-transfer, software writes a completely different configuration: mode 3 and divider 3.
repeat (8) @(negedge pclk);
apb_write(A_CTRL, 32'h37); // enable, cpol=1, cpha=1, divider=3
wait_idle;
apb_read(A_RXDATA); got_rx = rdbuf[7:0];
got_edges = edge_count;
$display(" shadow test: rxdata=%02h (slave sent %02h) slave received %02h (master sent %02h) shadow_ctrl=%02h (live ctrl=%02h) edges=%0d",
got_rx, slave_byte, slv_rx, 8'h3C, dbg_shadow_ctrl, 8'h37, got_edges);
if (got_rx !== slave_byte) begin
$display(" FAIL: the mid-transfer CTRL write corrupted the received byte");
errors = errors + 1;
end
if (slv_rx !== 8'h3C) begin
$display(" FAIL: the slave received %02h where the master was told to send 3c", slv_rx);
errors = errors + 1;
end
if (dbg_shadow_ctrl !== 8'h21) begin
$display(" FAIL: the shadowed configuration changed to %02h mid-transfer", dbg_shadow_ctrl);
errors = errors + 1;
end
if (got_edges != 16) begin
$display(" FAIL: %0d SCLK edges on the wire where 16 were expected", got_edges);
errors = errors + 1;
end
// ================= 2. write-one-to-clear =================
reset_all;
slave_byte = 8'h5A;
apb_write(A_CTRL, 32'h21);
apb_write(A_TXDATA, 32'h11);
wait_idle;
apb_write(A_TXDATA, 32'h22); // a second start with DONE still set -> OVERRUN
wait_idle;
apb_read(A_STATUS);
$display(" sticky bits after two transfers with no clear: busy=%b done=%b overrun=%b",
rdbuf[0], rdbuf[1], rdbuf[2]);
if (!(rdbuf[1] && rdbuf[2])) begin
$display(" FAIL: DONE and OVERRUN were not both set after an unread completion");
errors = errors + 1;
end
// Writing a ZERO must leave a bit alone.
apb_write(A_STATUS, 32'h00000002); // clear DONE only
apb_read(A_STATUS);
$display(" after writing 0x02: busy=%b done=%b overrun=%b",
rdbuf[0], rdbuf[1], rdbuf[2]);
if (rdbuf[1]) begin
$display(" FAIL: writing a one to DONE did not clear it");
errors = errors + 1;
end
if (!rdbuf[2]) begin
$display(" FAIL: writing a one to DONE also cleared OVERRUN; a zero must leave a bit alone");
errors = errors + 1;
end
apb_write(A_STATUS, 32'h00000004); // now clear OVERRUN
apb_read(A_STATUS);
if (rdbuf[2]) begin
$display(" FAIL: writing a one to OVERRUN did not clear it");
errors = errors + 1;
end
// ================= 3. a start while busy is refused and recorded =================
reset_all;
slave_byte = 8'h3C;
apb_write(A_CTRL, 32'h21);
apb_write(A_TXDATA, 32'hAA);
apb_write(A_TXDATA, 32'hBB); // issued while the first is still running
wait_idle;
apb_read(A_STATUS);
$display(" a second start while busy: done=%b overrun=%b slave received %02h",
rdbuf[1], rdbuf[2], slv_rx);
if (!rdbuf[2]) begin
$display(" FAIL: a start issued while busy was not recorded as an overrun");
errors = errors + 1;
end
if (slv_rx !== 8'hAA) begin
$display(" FAIL: the slave received %02h; the refused start should not have disturbed the transfer in flight",
slv_rx);
errors = errors + 1;
end
// ================= 4. rate invariance =================
$display("");
$display(" sclk_ref half ratio to pclk rxdata frames edges status");
for (k = 0; k < 4; k = k + 1) begin
sref_half = RATIOS[k];
reset_all;
slave_byte = 8'h96;
apb_write(A_CTRL, 32'h21);
apb_write(A_TXDATA, 32'h69);
wait_idle;
apb_read(A_RXDATA); res_rx[k] = rdbuf[7:0];
res_frames[k] = frames_seen;
res_edges[k] = edge_count;
apb_read(A_STATUS);
// NO `%Ns` ON A STRING. Icarus pads it in one language mode and not the other, so the two
// transcripts differed by whitespace alone -- which is exactly the difference that makes a
// cross-language comparison worthless. Fixed-width literals only.
$display(" %13d unrelated %02h %6d %5d done=%b overrun=%b",
sref_half,
res_rx[k], res_frames[k], res_edges[k], rdbuf[1], rdbuf[2]);
if (res_rx[k] !== 8'h96) begin
$display(" FAIL: at a half period of %0d the received byte was %02h, not 96",
sref_half, res_rx[k]);
errors = errors + 1;
end
if (res_edges[k] != 16) begin
$display(" FAIL: at a half period of %0d the wire carried %0d edges, not 16",
sref_half, res_edges[k]);
errors = errors + 1;
end
if (rdbuf[2]) begin
$display(" FAIL: at a half period of %0d an overrun was reported for a single transfer",
sref_half);
errors = errors + 1;
end
end
if (!(res_rx[0] == res_rx[1] && res_rx[1] == res_rx[2] && res_rx[2] == res_rx[3])) begin
$display(" FAIL: the received byte depended on the clock ratio (%02h %02h %02h %02h)",
res_rx[0], res_rx[1], res_rx[2], res_rx[3]);
errors = errors + 1;
end
if (!(res_frames[0] == res_frames[1] && res_frames[1] == res_frames[2]
&& res_frames[2] == res_frames[3])) begin
$display(" FAIL: the frame count depended on the clock ratio");
errors = errors + 1;
end
// ================= conclusions =================
$display("");
$display(" 1. a mid-transfer CTRL write changed the live register from 21 to 37 -- a different MODE and a different DIVIDER -- and the transfer in flight was untouched: the slave received 3c, the master received %02h, the shadowed configuration still read 21, and the wire carried exactly 16 edges. The shadow costs one register's worth of flops and removes a rule every future driver would otherwise have to obey. The alternative contract -- `poll BUSY before writing CTRL` -- is a rule that has to be honoured by people who will never read the datasheet section that states it",
8'h96);
$display(" 2. write-one-to-clear means what it says: after both sticky bits were set, a write of 0x02 cleared DONE and LEFT OVERRUN alone. A status register that clears on any write destroys a bit the driver has not read yet, and because the destroyed bit is the record of a lost completion the symptom is a stream that is quietly short");
$display(" 3. a second completion with the first unread set OVERRUN rather than being discarded, and a second START issued while busy was refused without disturbing the transfer in flight -- the slave still received aa. A sticky DONE without an OVERRUN companion silently loses completions, which is the failure a level DONE was replaced to avoid rather than a different one");
$display(" 4. and the results do not depend on the clock ratio. At sclk_ref half periods of %0d, %0d, %0d and %0d nanoseconds against a fixed 5 ns bus half period -- none of them an integer relationship -- every received byte was 96, every frame count was identical and every wire carried 16 edges. That is rate invariance, and it is the property a multi-bit crossing defect breaks. It is also the ONLY thing this simulation establishes about the crossing: both synchronisers are two flops deep because that is correct practice, and a single flop would behave identically here",
RATIOS[0], RATIOS[1], RATIOS[2], RATIOS[3]);
// ================= BENCH INTEGRITY =================
if (dbg_shadow_ctrl !== 8'hDE) mutations = mutations + 1;
if (res_edges[0] != 99) mutations = mutations + 1;
if (mutations != 2) begin
$display(" FAIL: a deliberately wrong expectation did not mismatch (%0d of 2)", mutations);
errors = errors + 1;
end
if (frames_seen == 0) begin
$display(" FAIL: the pin monitor never saw a frame, so nothing above was observed on the wire");
errors = errors + 1;
end
if (x_reports != 0) begin
$display(" FAIL: %0d bus reads returned an X", x_reports);
errors = errors + 1;
end
if (errors == 0) begin
$display("");
$display(" and the bench proved itself: two deliberately wrong expectations mismatched, the pin monitor observed every frame independently of the design's own status bits, every bus read carried a known value, and the slave model knows nothing about the register bus");
$display("PASS: putting an SPI master behind a register bus makes software part of the design, and the three failures that ship are all in the seam. A CONFIGURATION SHADOW makes a mid-transfer CTRL write harmless -- the live register went from 21 to 37 while the transfer in flight kept its mode, its divider and its 16 edges -- and it replaces a rule every future driver would have to obey with one register's worth of flops. STATUS SEMANTICS need two bits and a discipline: BUSY is a level, DONE is sticky and write-one-to-clear because a level DONE can be missed between transfers, and OVERRUN exists because a sticky DONE with no companion silently loses the second completion. Writing a zero left a bit alone, a second completion set OVERRUN, and a second start while busy was refused without disturbing the transfer on the wire. And the CROSSING is multi-bit, so it is a request/acknowledge TOGGLE handshake with the data held stable across it rather than a synchroniser per bit: at four unrelated clock ratios every received byte, frame count and edge count was identical. Rate invariance is the one property a simulation can genuinely establish about a crossing -- it says nothing whatever about synchroniser depth, because metastability is not representable and a single flop would behave identically here");
end else begin
$display("FAIL: %0d error(s)", errors);
end
$finish;
end
endmodule// spi_regbus_ctrl_tb.v
//
// SOFTWARE ON ONE CLOCK, AN SPI ENGINE ON ANOTHER, AND THE THREE RACES THAT SHIP.
//
// The bench drives the register bus the way a driver would -- writes and reads through APB accesses, with
// no visibility into the engine -- and models the slave on the pins. The two clocks are unrelated and
// their ratio is swept, because a crossing defect is a phase effect and a single ratio cannot find one.
//
// THE FOUR RESULTS.
//
// 1. THE SHADOW REGISTER MAKES A MID-TRANSFER CONFIGURATION WRITE HARMLESS. Software writes CTRL with a
// different mode AND a different divider while a transfer is in flight. The bench requires the
// received byte to be correct, the shadowed configuration to be UNCHANGED, and the number of SCLK
// edges on the wire to be exactly 16 -- because a divider that changed mid-frame would still produce
// 16 edges at the wrong spacing, and a mode that changed would corrupt the data while the count
// looked fine. Both are checked.
//
// 2. WRITE-ONE-TO-CLEAR MEANS WRITING A ZERO LEAVES A BIT ALONE. The bench sets both sticky bits, writes
// a word with only one bit set, and requires the other to survive. A status register that clears on
// any write destroys a bit the driver had not read yet, and the symptom is a lost completion.
//
// 3. A SECOND COMPLETION WITH THE FIRST UNREAD MUST BE RECORDED, NOT DISCARDED. Two transfers are run
// without clearing DONE between them, and OVERRUN must be set. A sticky DONE with no OVERRUN silently
// loses completions, which is the failure a level DONE was replaced to avoid -- traded for a
// different one rather than fixed.
//
// 4. THE RESULTS DO NOT DEPEND ON THE CLOCK RATIO. The same traffic is run at four unrelated
// pclk-to-sclk_ref ratios and every received byte, every status word and every transfer count must be
// IDENTICAL. That is rate invariance, and it is the property a multi-bit crossing defect breaks --
// the one thing a simulation can genuinely establish about a crossing.
`timescale 1ns/1ps
module spi_regbus_ctrl_tb;
localparam AW = 8;
localparam [AW-1:0] A_CTRL = 'h0, A_TXDATA = 'h4, A_RXDATA = 'h8,
A_STATUS = 'hC, A_IE = 'h10;
// ---- the register-bus clock: fixed ----
reg pclk;
always #5 pclk = ~pclk;
reg prst_n;
// ---- the SPI reference clock: its period is the swept variable ----
reg sclk_ref;
reg srst_n;
integer sref_half; // half period in ns; deliberately not a divisor of 10
always begin
#(sref_half) sclk_ref = ~sclk_ref;
end
reg psel, penable, pwrite;
// `'0` is a SystemVerilog literal and does not survive the Verilog-2001 conversion; an
// explicit width is the only spelling that means the same thing in both.
reg [AW-1:0] paddr;
reg [31:0] pwdata;
wire [31:0] prdata;
wire pready, irq;
wire sclk, cs_n, mosi;
reg miso;
wire [7:0] dbg_shadow_ctrl, dbg_txshadow;
spi_regbus_ctrl #(.AW(AW)) dut (
.pclk(pclk), .prst_n(prst_n), .psel(psel), .penable(penable), .pwrite(pwrite),
.paddr(paddr), .pwdata(pwdata), .prdata(prdata), .pready(pready), .irq(irq),
.sclk_ref(sclk_ref), .srst_n(srst_n),
.sclk(sclk), .cs_n(cs_n), .mosi(mosi), .miso(miso),
.dbg_shadow_ctrl(dbg_shadow_ctrl), .dbg_txshadow(dbg_txshadow)
);
integer errors;
// ------------------------------------------------------------------
// THE SLAVE MODEL, on the pins. Mode 0: presents the MSB from the select and advances on the trailing
// edge, so the master's leading edge always has a full half period of setup. It knows nothing about
// the register bus.
// ------------------------------------------------------------------
reg [7:0] slave_byte;
reg [7:0] slv_sh;
reg [7:0] slv_rx; // what the slave received, for the transmit-path check
reg cs_d, sclk_d;
// Edge counting on the wire, so the bench can see the frame's shape without asking the design.
integer edge_count, frames_seen;
// CLOCKED IN THE SPI DOMAIN, NOT ON THE BUS CLOCK.
//
// The first version sampled the pins on `pclk`. That works only while SCLK is slower than the bus
// clock, and the whole point of the rate sweep is to run configurations where it is not: at an
// sclk_ref half period of 3 ns against a 5 ns bus half period the monitor saw 4 edges out of 16 and
// the slave model returned a byte assembled from the edges it happened to catch. Both failures looked
// exactly like a broken crossing in the design.
//
// A real slave is clocked by SCLK, so sampling in the SPI reference domain is both more faithful and
// immune to undersampling -- SCLK changes only on an sclk_ref edge. The alternative, separate
// processes triggered on each pin edge, needs several writers for `miso` and the counters, which is a
// race in Verilog and a resolution problem in VHDL.
always @(posedge sclk_ref or negedge srst_n) begin
if (!srst_n) begin
slv_sh <= 8'h00; slv_rx <= 8'h00; cs_d <= 1'b1; sclk_d <= 1'b0;
edge_count <= 0; frames_seen <= 0;
miso <= 1'b0;
end else begin
// Delayed copies as NON-BLOCKING regs, so the edge detection lags by one reference cycle in
// every language. Chapter 19.2 found four faults in this exact distinction.
cs_d <= cs_n;
sclk_d <= sclk;
if (cs_d && !cs_n) begin
slv_sh <= slave_byte;
miso <= slave_byte[7];
slv_rx <= 8'h00;
edge_count <= 0;
end else if (!cs_n && (sclk_d !== sclk)) begin
edge_count <= edge_count + 1;
if (!sclk_d && sclk) begin
// leading edge: the slave samples MOSI
slv_rx <= {slv_rx[6:0], mosi};
end else begin
// trailing edge: the slave advances its own data
slv_sh <= {slv_sh[6:0], 1'b0};
miso <= slv_sh[6];
end
end else if (!cs_d && cs_n) begin
frames_seen <= frames_seen + 1;
end
end
end
// ------------------------------------------------------------------
// APB accesses, written the way a driver would issue them.
// ------------------------------------------------------------------
task apb_write;
input [AW-1:0] a;
input [31:0] d;
begin
@(negedge pclk);
psel = 1'b1; pwrite = 1'b1; paddr = a; pwdata = d; penable = 1'b0;
@(negedge pclk);
penable = 1'b1; // the access completes in the ENABLE phase
@(negedge pclk);
psel = 1'b0; penable = 1'b0; pwrite = 1'b0;
end
endtask
reg [31:0] rdbuf;
task apb_read;
input [AW-1:0] a;
begin
@(negedge pclk);
psel = 1'b1; pwrite = 1'b0; paddr = a; penable = 1'b0;
@(negedge pclk);
penable = 1'b1;
@(posedge pclk);
@(negedge pclk);
rdbuf = prdata;
psel = 1'b0; penable = 1'b0;
end
endtask
// Wait for BUSY to clear, bounded. An unbounded wait on a condition a broken design never satisfies
// is a hang, and a hung regression reports nothing at all.
integer guard;
task wait_idle;
begin
guard = 0;
apb_read(A_STATUS);
while (rdbuf[0] && guard < 4000) begin
apb_read(A_STATUS);
guard = guard + 1;
end
if (guard >= 4000) begin
$display(" FAIL: BUSY never cleared");
errors = errors + 1;
end
end
endtask
task reset_all;
begin
@(negedge pclk); prst_n = 1'b0; srst_n = 1'b0;
repeat (6) @(negedge pclk);
prst_n = 1'b1; srst_n = 1'b1;
repeat (4) @(negedge pclk);
end
endtask
integer k, x_reports, mutations;
integer RATIOS [0:3];
reg [7:0] res_rx [0:3];
integer res_frames[0:3], res_edges[0:3];
reg [7:0] got_rx;
integer got_edges;
always @(posedge pclk) if (prst_n) begin
if ((^prdata === 32'bx) || (^dbg_shadow_ctrl === 8'bx)) x_reports = x_reports + 1;
end
initial begin
x_reports = 0; mutations = 0;
RATIOS[0] = 7; RATIOS[1] = 3; RATIOS[2] = 11; RATIOS[3] = 4;
// ================= 1. the shadow register =================
sref_half = 7;
reset_all;
slave_byte = 8'hA5;
// Divider 2, not 0. This slave model presents each bit two reference cycles after the trailing
// edge -- one to detect it, one for the registered output -- so the SCLK half period needs at
// least three reference cycles for the master's leading edge to have any setup. At divider 0 the
// master captures the MSB twice and loses the last bit, which is 19.1's boundary again rather
// than anything about the register bus.
apb_write(A_CTRL, 32'h21); // enable, mode 0, divider 2
apb_write(A_TXDATA, 32'h3C); // start a transfer
// Mid-transfer, software writes a completely different configuration: mode 3 and divider 3.
repeat (8) @(negedge pclk);
apb_write(A_CTRL, 32'h37); // enable, cpol=1, cpha=1, divider=3
wait_idle;
apb_read(A_RXDATA); got_rx = rdbuf[7:0];
got_edges = edge_count;
$display(" shadow test: rxdata=%02h (slave sent %02h) slave received %02h (master sent %02h) shadow_ctrl=%02h (live ctrl=%02h) edges=%0d",
got_rx, slave_byte, slv_rx, 8'h3C, dbg_shadow_ctrl, 8'h37, got_edges);
if (got_rx !== slave_byte) begin
$display(" FAIL: the mid-transfer CTRL write corrupted the received byte");
errors = errors + 1;
end
if (slv_rx !== 8'h3C) begin
$display(" FAIL: the slave received %02h where the master was told to send 3c", slv_rx);
errors = errors + 1;
end
if (dbg_shadow_ctrl !== 8'h21) begin
$display(" FAIL: the shadowed configuration changed to %02h mid-transfer", dbg_shadow_ctrl);
errors = errors + 1;
end
if (got_edges != 16) begin
$display(" FAIL: %0d SCLK edges on the wire where 16 were expected", got_edges);
errors = errors + 1;
end
// ================= 2. write-one-to-clear =================
reset_all;
slave_byte = 8'h5A;
apb_write(A_CTRL, 32'h21);
apb_write(A_TXDATA, 32'h11);
wait_idle;
apb_write(A_TXDATA, 32'h22); // a second start with DONE still set -> OVERRUN
wait_idle;
apb_read(A_STATUS);
$display(" sticky bits after two transfers with no clear: busy=%b done=%b overrun=%b",
rdbuf[0], rdbuf[1], rdbuf[2]);
if (!(rdbuf[1] && rdbuf[2])) begin
$display(" FAIL: DONE and OVERRUN were not both set after an unread completion");
errors = errors + 1;
end
// Writing a ZERO must leave a bit alone.
apb_write(A_STATUS, 32'h00000002); // clear DONE only
apb_read(A_STATUS);
$display(" after writing 0x02: busy=%b done=%b overrun=%b",
rdbuf[0], rdbuf[1], rdbuf[2]);
if (rdbuf[1]) begin
$display(" FAIL: writing a one to DONE did not clear it");
errors = errors + 1;
end
if (!rdbuf[2]) begin
$display(" FAIL: writing a one to DONE also cleared OVERRUN; a zero must leave a bit alone");
errors = errors + 1;
end
apb_write(A_STATUS, 32'h00000004); // now clear OVERRUN
apb_read(A_STATUS);
if (rdbuf[2]) begin
$display(" FAIL: writing a one to OVERRUN did not clear it");
errors = errors + 1;
end
// ================= 3. a start while busy is refused and recorded =================
reset_all;
slave_byte = 8'h3C;
apb_write(A_CTRL, 32'h21);
apb_write(A_TXDATA, 32'hAA);
apb_write(A_TXDATA, 32'hBB); // issued while the first is still running
wait_idle;
apb_read(A_STATUS);
$display(" a second start while busy: done=%b overrun=%b slave received %02h",
rdbuf[1], rdbuf[2], slv_rx);
if (!rdbuf[2]) begin
$display(" FAIL: a start issued while busy was not recorded as an overrun");
errors = errors + 1;
end
if (slv_rx !== 8'hAA) begin
$display(" FAIL: the slave received %02h; the refused start should not have disturbed the transfer in flight",
slv_rx);
errors = errors + 1;
end
// ================= 4. rate invariance =================
$display("");
$display(" sclk_ref half ratio to pclk rxdata frames edges status");
for (k = 0; k < 4; k = k + 1) begin
sref_half = RATIOS[k];
reset_all;
slave_byte = 8'h96;
apb_write(A_CTRL, 32'h21);
apb_write(A_TXDATA, 32'h69);
wait_idle;
apb_read(A_RXDATA); res_rx[k] = rdbuf[7:0];
res_frames[k] = frames_seen;
res_edges[k] = edge_count;
apb_read(A_STATUS);
// NO `%Ns` ON A STRING. Icarus pads it in one language mode and not the other, so the two
// transcripts differed by whitespace alone -- which is exactly the difference that makes a
// cross-language comparison worthless. Fixed-width literals only.
$display(" %13d unrelated %02h %6d %5d done=%b overrun=%b",
sref_half,
res_rx[k], res_frames[k], res_edges[k], rdbuf[1], rdbuf[2]);
if (res_rx[k] !== 8'h96) begin
$display(" FAIL: at a half period of %0d the received byte was %02h, not 96",
sref_half, res_rx[k]);
errors = errors + 1;
end
if (res_edges[k] != 16) begin
$display(" FAIL: at a half period of %0d the wire carried %0d edges, not 16",
sref_half, res_edges[k]);
errors = errors + 1;
end
if (rdbuf[2]) begin
$display(" FAIL: at a half period of %0d an overrun was reported for a single transfer",
sref_half);
errors = errors + 1;
end
end
if (!(res_rx[0] == res_rx[1] && res_rx[1] == res_rx[2] && res_rx[2] == res_rx[3])) begin
$display(" FAIL: the received byte depended on the clock ratio (%02h %02h %02h %02h)",
res_rx[0], res_rx[1], res_rx[2], res_rx[3]);
errors = errors + 1;
end
if (!(res_frames[0] == res_frames[1] && res_frames[1] == res_frames[2]
&& res_frames[2] == res_frames[3])) begin
$display(" FAIL: the frame count depended on the clock ratio");
errors = errors + 1;
end
// ================= conclusions =================
$display("");
$display(" 1. a mid-transfer CTRL write changed the live register from 21 to 37 -- a different MODE and a different DIVIDER -- and the transfer in flight was untouched: the slave received 3c, the master received %02h, the shadowed configuration still read 21, and the wire carried exactly 16 edges. The shadow costs one register's worth of flops and removes a rule every future driver would otherwise have to obey. The alternative contract -- `poll BUSY before writing CTRL` -- is a rule that has to be honoured by people who will never read the datasheet section that states it",
8'h96);
$display(" 2. write-one-to-clear means what it says: after both sticky bits were set, a write of 0x02 cleared DONE and LEFT OVERRUN alone. A status register that clears on any write destroys a bit the driver has not read yet, and because the destroyed bit is the record of a lost completion the symptom is a stream that is quietly short");
$display(" 3. a second completion with the first unread set OVERRUN rather than being discarded, and a second START issued while busy was refused without disturbing the transfer in flight -- the slave still received aa. A sticky DONE without an OVERRUN companion silently loses completions, which is the failure a level DONE was replaced to avoid rather than a different one");
$display(" 4. and the results do not depend on the clock ratio. At sclk_ref half periods of %0d, %0d, %0d and %0d nanoseconds against a fixed 5 ns bus half period -- none of them an integer relationship -- every received byte was 96, every frame count was identical and every wire carried 16 edges. That is rate invariance, and it is the property a multi-bit crossing defect breaks. It is also the ONLY thing this simulation establishes about the crossing: both synchronisers are two flops deep because that is correct practice, and a single flop would behave identically here",
RATIOS[0], RATIOS[1], RATIOS[2], RATIOS[3]);
// ================= BENCH INTEGRITY =================
if (dbg_shadow_ctrl !== 8'hDE) mutations = mutations + 1;
if (res_edges[0] != 99) mutations = mutations + 1;
if (mutations != 2) begin
$display(" FAIL: a deliberately wrong expectation did not mismatch (%0d of 2)", mutations);
errors = errors + 1;
end
if (frames_seen == 0) begin
$display(" FAIL: the pin monitor never saw a frame, so nothing above was observed on the wire");
errors = errors + 1;
end
if (x_reports != 0) begin
$display(" FAIL: %0d bus reads returned an X", x_reports);
errors = errors + 1;
end
if (errors == 0) begin
$display("");
$display(" and the bench proved itself: two deliberately wrong expectations mismatched, the pin monitor observed every frame independently of the design's own status bits, every bus read carried a known value, and the slave model knows nothing about the register bus");
$display("PASS: putting an SPI master behind a register bus makes software part of the design, and the three failures that ship are all in the seam. A CONFIGURATION SHADOW makes a mid-transfer CTRL write harmless -- the live register went from 21 to 37 while the transfer in flight kept its mode, its divider and its 16 edges -- and it replaces a rule every future driver would have to obey with one register's worth of flops. STATUS SEMANTICS need two bits and a discipline: BUSY is a level, DONE is sticky and write-one-to-clear because a level DONE can be missed between transfers, and OVERRUN exists because a sticky DONE with no companion silently loses the second completion. Writing a zero left a bit alone, a second completion set OVERRUN, and a second start while busy was refused without disturbing the transfer on the wire. And the CROSSING is multi-bit, so it is a request/acknowledge TOGGLE handshake with the data held stable across it rather than a synchroniser per bit: at four unrelated clock ratios every received byte, frame count and edge count was identical. Rate invariance is the one property a simulation can genuinely establish about a crossing -- it says nothing whatever about synchroniser depth, because metastability is not representable and a single flop would behave identically here");
end else begin
$display("FAIL: %0d error(s)", errors);
end
$finish;
end
initial begin
psel = 1'b0;
penable = 1'b0;
pwrite = 1'b0;
cs_d = 1'b1;
sclk_d = 1'b0;
pclk = 1'b0;
prst_n = 1'b1;
sclk_ref = 1'b0;
srst_n = 1'b1;
sref_half = 7;
paddr = {AW{1'b0}};
pwdata = 32'h0;
miso = 1'b0;
errors = 0;
slave_byte = 8'h00;
slv_sh = 8'h00;
slv_rx = 8'h00;
end
endmodule-- spi_regbus_ctrl_tb.vhd
--
-- SOFTWARE ON ONE CLOCK, AN SPI ENGINE ON ANOTHER, AND THE THREE RACES THAT SHIP.
--
-- The bench drives the register bus the way a driver would -- writes and reads through APB accesses, with
-- no visibility into the engine -- and models the slave on the pins. The two clocks are unrelated and their
-- ratio is swept, because a crossing defect is a phase effect and a single ratio cannot find one.
--
-- The slave model and the pin monitor are clocked in the SPI REFERENCE domain, not on the bus clock.
-- Sampling the pins on the bus clock works only while SCLK is slower than it, and the whole point of the
-- sweep is to run configurations where it is not -- at which point the monitor undersamples and both
-- failures look exactly like a broken crossing in the design.
--
-- The same four results as the other two languages, with the same numbers.
--
-- IDENTIFIER REVIEW (VHDL IS CASE-INSENSITIVE). `AW_C`, `A_CTRL_C` and friends carry suffixes; the bus
-- drive signals are `s_psel`, `s_paddr` and so on. Nothing collides with a reserved word -- in particular
-- nothing is named `register` or `literal`, which are tempting in a register-file bench.
--
-- RANGE DIRECTION: every vector is `downto` and every subprogram formal is constrained.
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
use std.textio.all;
use work.spi_regbus_pkg.all;
entity spi_regbus_ctrl_tb is
end entity spi_regbus_ctrl_tb;
architecture tb of spi_regbus_ctrl_tb is
constant AW_C : positive := 8;
function addr_of (v : natural) return std_logic_vector is
begin
return std_logic_vector(to_unsigned(v, AW_C));
end function addr_of;
constant A_CTRL_C : std_logic_vector(AW_C - 1 downto 0) := addr_of(16#00#);
constant A_TXDATA_C : std_logic_vector(AW_C - 1 downto 0) := addr_of(16#04#);
constant A_RXDATA_C : std_logic_vector(AW_C - 1 downto 0) := addr_of(16#08#);
constant A_STATUS_C : std_logic_vector(AW_C - 1 downto 0) := addr_of(16#0C#);
signal pclk : std_logic := '0';
signal prst_n : std_logic := '1';
signal run : boolean := true;
-- the SPI reference clock: its half period is the swept variable
signal sclk_ref : std_logic := '0';
signal srst_n : std_logic := '1';
signal sref_half : time := 7 ns;
signal s_psel, s_penable, s_pwrite : std_logic := '0';
signal s_paddr : std_logic_vector(AW_C - 1 downto 0) := (others => '0');
signal s_pwdata : std_logic_vector(31 downto 0) := (others => '0');
signal prdata : std_logic_vector(31 downto 0);
signal pready, irq : std_logic;
signal sclk, cs_n, mosi : std_logic;
signal miso : std_logic := '0';
signal dbg_shadow_ctrl, dbg_txshadow : byte_t;
-- the slave model and pin monitor, driven by ONE process in the SPI reference domain
signal slave_byte : byte_t := (others => '0');
signal slv_rx : byte_t := (others => '0');
signal edge_count, frames_seen : natural := 0;
-- Delayed copies as SIGNALS, driven by ONE process, so the edge detection lags by exactly one
-- reference cycle in every language -- the distinction that produced four separate defects in
-- Chapter 19.2.
signal cs_d : std_logic := '1';
signal sclk_d : std_logic := '0';
signal x_reports : natural := 0;
type nat_arr is array (natural range <>) of natural;
type byte_arr is array (natural range <>) of byte_t;
function i2s (v : integer; w : natural) return string is
constant S : string := integer'image(v);
constant P : string(1 to 40) := (others => ' ');
begin
if S'length >= w then return S; end if;
return P(1 to w - S'length) & S;
end function i2s;
function hex2 (v : byte_t) return string is
constant D : string := "0123456789abcdef";
variable u : natural := to_integer(unsigned(v));
variable r : string(1 to 2);
begin
r(1) := D(u / 16 + 1);
r(2) := D(u mod 16 + 1);
return r;
end function hex2;
function b2s (b : std_logic) return string is
begin
if b = '1' then return "1"; else return "0"; end if;
end function b2s;
begin
pclk_gen : process is
begin
while run loop
pclk <= '0'; wait for 5 ns;
pclk <= '1'; wait for 5 ns;
end loop;
wait;
end process pclk_gen;
sref_gen : process is
begin
while run loop
sclk_ref <= '0'; wait for sref_half;
sclk_ref <= '1'; wait for sref_half;
end loop;
wait;
end process sref_gen;
dut : entity work.spi_regbus_ctrl
generic map (AW => AW_C)
port map (
pclk => pclk, prst_n => prst_n, psel => s_psel, penable => s_penable,
pwrite => s_pwrite, paddr => s_paddr, pwdata => s_pwdata,
prdata => prdata, pready => pready, irq => irq,
sclk_ref => sclk_ref, srst_n => srst_n,
sclk => sclk, cs_n => cs_n, mosi => mosi, miso => miso,
dbg_shadow_ctrl => dbg_shadow_ctrl, dbg_txshadow => dbg_txshadow
);
-- THE SLAVE MODEL AND PIN MONITOR, in the SPI reference domain. Mode 0: presents the MSB from the
-- select and advances on the trailing edge, so the master's leading edge always has setup. It knows
-- nothing about the register bus.
--
-- The delayed copies are SIGNALS, so the edge detection lags by one reference cycle in every language.
slave : process (sclk_ref, srst_n) is
variable slv_sh : byte_t := (others => '0');
begin
if srst_n = '0' then
slv_sh := (others => '0');
slv_rx <= (others => '0');
edge_count <= 0;
frames_seen <= 0;
miso <= '0';
elsif rising_edge(sclk_ref) then
if cs_d = '1' and cs_n = '0' then
slv_sh := slave_byte;
miso <= slave_byte(7);
slv_rx <= (others => '0');
edge_count <= 0;
elsif cs_n = '0' and sclk_d /= sclk then
edge_count <= edge_count + 1;
if sclk_d = '0' and sclk = '1' then
slv_rx <= slv_rx(6 downto 0) & mosi; -- leading edge: the slave samples MOSI
else
miso <= slv_sh(6); -- trailing edge: the slave advances
slv_sh := slv_sh(6 downto 0) & '0';
end if;
elsif cs_d = '0' and cs_n = '1' then
frames_seen <= frames_seen + 1;
end if;
end if;
end process slave;
dly : process (sclk_ref, srst_n) is
begin
if srst_n = '0' then
cs_d <= '1'; sclk_d <= '0';
elsif rising_edge(sclk_ref) then
cs_d <= cs_n;
sclk_d <= sclk;
end if;
end process dly;
xchk : process (pclk) is
begin
if rising_edge(pclk) and prst_n = '1' then
for i in 0 to 31 loop
if prdata(i) /= '0' and prdata(i) /= '1' then
x_reports <= x_reports + 1;
end if;
end loop;
end if;
end process xchk;
stim : process is
variable e, mutations : natural := 0;
variable rdbuf : std_logic_vector(31 downto 0);
variable got_rx : byte_t;
variable got_edges, guard : natural := 0;
constant RATIOS_C : nat_arr(0 to 3) := (7, 3, 11, 4);
variable res_rx : byte_arr(0 to 3);
variable res_frames, res_edges : nat_arr(0 to 3);
variable ln : line;
procedure apb_write (a : std_logic_vector(AW_C - 1 downto 0);
d : std_logic_vector(31 downto 0)) is
begin
wait until falling_edge(pclk);
s_psel <= '1'; s_pwrite <= '1'; s_paddr <= a; s_pwdata <= d; s_penable <= '0';
wait until falling_edge(pclk);
s_penable <= '1'; -- the access completes in the ENABLE phase
wait until falling_edge(pclk);
s_psel <= '0'; s_penable <= '0'; s_pwrite <= '0';
end procedure apb_write;
procedure apb_read (a : std_logic_vector(AW_C - 1 downto 0)) is
begin
wait until falling_edge(pclk);
s_psel <= '1'; s_pwrite <= '0'; s_paddr <= a; s_penable <= '0';
wait until falling_edge(pclk);
s_penable <= '1';
wait until rising_edge(pclk);
wait until falling_edge(pclk);
rdbuf := prdata;
s_psel <= '0'; s_penable <= '0';
end procedure apb_read;
-- Bounded. An unbounded wait on a condition a broken design never satisfies is a hang, and a hung
-- regression reports nothing at all.
procedure wait_idle is
begin
guard := 0;
apb_read(A_STATUS_C);
while rdbuf(0) = '1' and guard < 4000 loop
apb_read(A_STATUS_C);
guard := guard + 1;
end loop;
if guard >= 4000 then
write(ln, string'(" FAIL: BUSY never cleared")); writeline(output, ln); e := e + 1;
end if;
end procedure wait_idle;
procedure reset_all is
begin
wait until falling_edge(pclk);
prst_n <= '0'; srst_n <= '0';
for i in 1 to 6 loop wait until falling_edge(pclk); end loop;
prst_n <= '1'; srst_n <= '1';
for i in 1 to 4 loop wait until falling_edge(pclk); end loop;
end procedure reset_all;
begin
-- ============ 1. the shadow register ============
sref_half <= 7 ns;
reset_all;
slave_byte <= x"A5";
-- Divider 2, not 0. This slave model presents each bit two reference cycles after the trailing
-- edge, so the SCLK half period needs at least three reference cycles for the master's leading edge
-- to have any setup. That boundary belongs to Chapter 19.1, not to the register bus.
apb_write(A_CTRL_C, x"00000021");
apb_write(A_TXDATA_C, x"0000003C");
-- Mid-transfer, software writes a completely different configuration: mode 3 and divider 3.
for i in 1 to 8 loop wait until falling_edge(pclk); end loop;
apb_write(A_CTRL_C, x"00000037");
wait_idle;
apb_read(A_RXDATA_C); got_rx := rdbuf(7 downto 0);
got_edges := edge_count;
write(ln, string'(" shadow test: rxdata=") & hex2(got_rx) & string'(" (slave sent ")
& hex2(slave_byte) & string'(") slave received ") & hex2(slv_rx)
& string'(" (master sent 3c) shadow_ctrl=") & hex2(dbg_shadow_ctrl)
& string'(" (live ctrl=37) edges=") & i2s(got_edges, 1));
writeline(output, ln);
if got_rx /= slave_byte then
write(ln, string'(" FAIL: the mid-transfer CTRL write corrupted the received byte"));
writeline(output, ln); e := e + 1;
end if;
if slv_rx /= x"3C" then
write(ln, string'(" FAIL: the slave received the wrong byte"));
writeline(output, ln); e := e + 1;
end if;
if dbg_shadow_ctrl /= x"21" then
write(ln, string'(" FAIL: the shadowed configuration changed mid-transfer"));
writeline(output, ln); e := e + 1;
end if;
if got_edges /= 16 then
write(ln, string'(" FAIL: ") & i2s(got_edges, 1)
& string'(" SCLK edges on the wire where 16 were expected"));
writeline(output, ln); e := e + 1;
end if;
-- ============ 2. write-one-to-clear ============
reset_all;
slave_byte <= x"5A";
apb_write(A_CTRL_C, x"00000021");
apb_write(A_TXDATA_C, x"00000011");
wait_idle;
apb_write(A_TXDATA_C, x"00000022"); -- a second start with DONE still set -> OVERRUN
wait_idle;
apb_read(A_STATUS_C);
write(ln, string'(" sticky bits after two transfers with no clear: busy=") & b2s(rdbuf(0))
& string'(" done=") & b2s(rdbuf(1)) & string'(" overrun=") & b2s(rdbuf(2)));
writeline(output, ln);
if not (rdbuf(1) = '1' and rdbuf(2) = '1') then
write(ln, string'(" FAIL: DONE and OVERRUN were not both set after an unread completion"));
writeline(output, ln); e := e + 1;
end if;
apb_write(A_STATUS_C, x"00000002"); -- clear DONE only
apb_read(A_STATUS_C);
write(ln, string'(" after writing 0x02: busy=") & b2s(rdbuf(0))
& string'(" done=") & b2s(rdbuf(1)) & string'(" overrun=") & b2s(rdbuf(2)));
writeline(output, ln);
if rdbuf(1) = '1' then
write(ln, string'(" FAIL: writing a one to DONE did not clear it"));
writeline(output, ln); e := e + 1;
end if;
if rdbuf(2) = '0' then
write(ln, string'(" FAIL: writing a one to DONE also cleared OVERRUN; a zero must leave a bit alone"));
writeline(output, ln); e := e + 1;
end if;
apb_write(A_STATUS_C, x"00000004");
apb_read(A_STATUS_C);
if rdbuf(2) = '1' then
write(ln, string'(" FAIL: writing a one to OVERRUN did not clear it"));
writeline(output, ln); e := e + 1;
end if;
-- ============ 3. a start while busy is refused and recorded ============
reset_all;
slave_byte <= x"3C";
apb_write(A_CTRL_C, x"00000021");
apb_write(A_TXDATA_C, x"000000AA");
apb_write(A_TXDATA_C, x"000000BB"); -- issued while the first is still running
wait_idle;
apb_read(A_STATUS_C);
write(ln, string'(" a second start while busy: done=") & b2s(rdbuf(1))
& string'(" overrun=") & b2s(rdbuf(2)) & string'(" slave received ") & hex2(slv_rx));
writeline(output, ln);
if rdbuf(2) = '0' then
write(ln, string'(" FAIL: a start issued while busy was not recorded as an overrun"));
writeline(output, ln); e := e + 1;
end if;
if slv_rx /= x"AA" then
write(ln, string'(" FAIL: the refused start disturbed the transfer in flight"));
writeline(output, ln); e := e + 1;
end if;
-- ============ 4. rate invariance ============
write(ln, string'(""));
writeline(output, ln);
write(ln, string'(" sclk_ref half ratio to pclk rxdata frames edges status"));
writeline(output, ln);
for k in 0 to 3 loop
sref_half <= RATIOS_C(k) * 1 ns;
reset_all;
slave_byte <= x"96";
apb_write(A_CTRL_C, x"00000021");
apb_write(A_TXDATA_C, x"00000069");
wait_idle;
apb_read(A_RXDATA_C); res_rx(k) := rdbuf(7 downto 0);
res_frames(k) := frames_seen;
res_edges(k) := edge_count;
apb_read(A_STATUS_C);
write(ln, string'(" ") & i2s(RATIOS_C(k), 13) & string'(" unrelated ")
& hex2(res_rx(k)) & string'(" ") & i2s(res_frames(k), 6)
& string'(" ") & i2s(res_edges(k), 5) & string'(" done=") & b2s(rdbuf(1))
& string'(" overrun=") & b2s(rdbuf(2)));
writeline(output, ln);
if res_rx(k) /= x"96" then
write(ln, string'(" FAIL: at a half period of ") & i2s(RATIOS_C(k), 1)
& string'(" the received byte was wrong"));
writeline(output, ln); e := e + 1;
end if;
if res_edges(k) /= 16 then
write(ln, string'(" FAIL: at a half period of ") & i2s(RATIOS_C(k), 1)
& string'(" the wire carried ") & i2s(res_edges(k), 1)
& string'(" edges, not 16"));
writeline(output, ln); e := e + 1;
end if;
if rdbuf(2) = '1' then
write(ln, string'(" FAIL: an overrun was reported for a single transfer"));
writeline(output, ln); e := e + 1;
end if;
end loop;
if not (res_rx(0) = res_rx(1) and res_rx(1) = res_rx(2) and res_rx(2) = res_rx(3)) then
write(ln, string'(" FAIL: the received byte depended on the clock ratio"));
writeline(output, ln); e := e + 1;
end if;
if not (res_frames(0) = res_frames(1) and res_frames(1) = res_frames(2)
and res_frames(2) = res_frames(3)) then
write(ln, string'(" FAIL: the frame count depended on the clock ratio"));
writeline(output, ln); e := e + 1;
end if;
-- ============ conclusions ============
write(ln, string'(""));
writeline(output, ln);
write(ln, string'(" 1. a mid-transfer CTRL write changed the live register from 21 to 37 -- a different MODE and a different DIVIDER -- and the transfer in flight was untouched: the slave received 3c, the master received 96, the shadowed configuration still read 21, and the wire carried exactly 16 edges. The shadow costs one register's worth of flops and removes a rule every future driver would otherwise have to obey. The alternative contract -- `poll BUSY before writing CTRL` -- is a rule that has to be honoured by people who will never read the datasheet section that states it"));
writeline(output, ln);
write(ln, string'(" 2. write-one-to-clear means what it says: after both sticky bits were set, a write of 0x02 cleared DONE and LEFT OVERRUN alone. A status register that clears on any write destroys a bit the driver has not read yet, and because the destroyed bit is the record of a lost completion the symptom is a stream that is quietly short"));
writeline(output, ln);
write(ln, string'(" 3. a second completion with the first unread set OVERRUN rather than being discarded, and a second START issued while busy was refused without disturbing the transfer in flight -- the slave still received aa. A sticky DONE without an OVERRUN companion silently loses completions, which is the failure a level DONE was replaced to avoid rather than a different one"));
writeline(output, ln);
write(ln, string'(" 4. and the results do not depend on the clock ratio. At sclk_ref half periods of ")
& i2s(RATIOS_C(0), 1) & string'(", ") & i2s(RATIOS_C(1), 1) & string'(", ")
& i2s(RATIOS_C(2), 1) & string'(" and ") & i2s(RATIOS_C(3), 1)
& string'(" nanoseconds against a fixed 5 ns bus half period -- none of them an integer relationship -- every received byte was 96, every frame count was identical and every wire carried 16 edges. That is rate invariance, and it is the property a multi-bit crossing defect breaks. It is also the ONLY thing this simulation establishes about the crossing: both synchronisers are two flops deep because that is correct practice, and a single flop would behave identically here"));
writeline(output, ln);
-- ============ BENCH INTEGRITY ============
if dbg_shadow_ctrl /= x"DE" then mutations := mutations + 1; end if;
if res_edges(0) /= 99 then mutations := mutations + 1; end if;
if mutations /= 2 then
write(ln, string'(" FAIL: a deliberately wrong expectation did not mismatch ("
)) ; write(ln, i2s(mutations, 1) & string'(" of 2)"));
writeline(output, ln); e := e + 1;
end if;
if frames_seen = 0 then
write(ln, string'(" FAIL: the pin monitor never saw a frame, so nothing above was observed on the wire"));
writeline(output, ln); e := e + 1;
end if;
if x_reports /= 0 then
write(ln, string'(" FAIL: bus reads returned a metavalue"));
writeline(output, ln); e := e + 1;
end if;
if e = 0 then
write(ln, string'(""));
writeline(output, ln);
write(ln, string'(" and the bench proved itself: two deliberately wrong expectations mismatched, the pin monitor observed every frame independently of the design's own status bits, every bus read carried a known value, and the slave model knows nothing about the register bus"));
writeline(output, ln);
write(ln, string'("PASS: putting an SPI master behind a register bus makes software part of the design, and the three failures that ship are all in the seam. A CONFIGURATION SHADOW makes a mid-transfer CTRL write harmless -- the live register went from 21 to 37 while the transfer in flight kept its mode, its divider and its 16 edges -- and it replaces a rule every future driver would have to obey with one register's worth of flops. STATUS SEMANTICS need two bits and a discipline: BUSY is a level, DONE is sticky and write-one-to-clear because a level DONE can be missed between transfers, and OVERRUN exists because a sticky DONE with no companion silently loses the second completion. Writing a zero left a bit alone, a second completion set OVERRUN, and a second start while busy was refused without disturbing the transfer on the wire. And the CROSSING is multi-bit, so it is a request/acknowledge TOGGLE handshake with the data held stable across it rather than a synchroniser per bit: at four unrelated clock ratios every received byte, frame count and edge count was identical. Rate invariance is the one property a simulation can genuinely establish about a crossing -- it says nothing whatever about synchroniser depth, because metastability is not representable and a single flop would behave identically here"));
writeline(output, ln);
else
write(ln, string'("FAIL: ") & i2s(e, 1) & string'(" error(s)"));
writeline(output, ln);
end if;
run <= false;
wait;
end process stim;
end architecture tb;7. Where UVM RAL Genuinely Belongs
This is the first chapter in the module with a register map, and a register abstraction layer is the right tool here rather than a decoration.
a register model mirrors CTRL, TXDATA, RXDATA, STATUS and IE
frontdoor access drives real APB transactions
the model PREDICTS the mirror after each access8. The Assertions Worth Writing
BUSY is high if and only if the request and acknowledge toggles disagree
no frame starts while BUSY is high
the shadowed configuration does not change between the start and the end
of a frame
every frame carries exactly 16 SCLK edges9. FPGA And ASIC Implementation
On an FPGA, the two clocks need to be declared asynchronous to each other — set_clock_groups -asynchronous or the vendor equivalent — or static timing analysis will try to close a path between them and report a violation on a crossing that is deliberately unconstrained. The synchroniser flops should carry whatever the vendor's false path or max delay convention is, so the tool does not optimise them into one.
The register file is small enough to live in logic. What is worth thinking about is prdata: a wide read mux across five registers is a combinational cone that can become the bus's critical path. Registering it — as this design does — costs one cycle of read latency that APB is happy to absorb.
On an ASIC, the same crossing needs a set_clock_groups and a CDC tool run, and the synchroniser cells usually come from a library with a documented mean-time-between-failure rather than from inferred flops. That MTBF number is the thing simulation cannot give you and the reason the tool exists.
And the register map is a specification deliverable. Field widths, reset values, access types and volatility go in a machine-readable description that generates both the RTL and the register model — because a map maintained in two places diverges, and the divergence appears as a driver that works against the documentation and not against the silicon.
10. Failure Signature — "The Driver Loses Every Fourth Byte"
Symptom a streaming driver loses roughly one byte in four at high rates
and none at low rates
Checked the SPI pins on a scope -- every frame present and correct
Checked the driver's own accounting -- one write per byte, as expected
Concluded the hardware drops transfers under load
Actual the driver polled DONE, saw it set, read RXDATA, and cleared
STATUS with a write of 0xFF -- which also cleared an OVERRUN
that had been set by a completion arriving during the read
Found by an engineer who noticed OVERRUN was never once observed as
set, in a system that was demonstrably overrunningThe scope was right, the driver's accounting was right, and the conclusion followed. What made it invisible is that the evidence was being destroyed by the act of reading: a blanket 0xFF write to a status register clears bits the driver has not looked at, and the bit it destroyed was the only record that anything had gone wrong.
A counter that is never observed to be non-zero in a system that is demonstrably failing is not evidence of correctness — it is evidence that something is clearing it.
11. Common Misconceptions
| Misconception | What is actually true |
|---|---|
| Software should just poll BUSY before writing CTRL | That is a rule every future driver must obey; a shadow register removes it |
| A sticky DONE fixes missed completions | It trades them for lost second completions, which is why OVERRUN exists |
| Writing 0xFF to a status register is a safe way to clear it | It destroys bits the driver has not read, including the record of a loss |
| A multi-bit value can be synchronised bit by bit | Each bit resolves independently; a value that was never valid can appear |
| A busy flag is simpler than deriving BUSY from the toggles | A separate flag can disagree with the engine; a derived one cannot |
| A divider that produces the right number of edges is working | It may be delaying only the first edge and running free afterwards |
| A passing rate sweep proves the crossing is safe | It proves the handshake's shape; it says nothing about synchroniser depth |
12. Reason It Through
13. Understanding Check
14. Summary
Putting an SPI master behind a register bus makes software part of the design, and the three failures that ship are all in the seam.
A configuration shadow makes a mid-transfer CTRL write harmless: the live register went from 0x21 to 0x37 while the transfer in flight kept its mode, its divider and its 16 edges. It replaces a rule every future driver would have to obey with one register's worth of flops — prefer a hardware invariant to a software obligation, because the invariant is enforced once and the obligation is re-tested by every integration.
Status semantics need two bits and a discipline. BUSY is a level, derived from the crossing's own toggles rather than maintained separately, so it cannot disagree with the engine. DONE is sticky and write-one-to-clear, because a level DONE exists only between transfers and is missed by any driver doing work. And OVERRUN exists because a sticky DONE has nowhere to record a second completion — it is the other half of the design, not a diagnostic. Writing a zero left a bit alone, a second completion set OVERRUN, and a second start while busy was refused without disturbing the transfer on the wire.
The crossing is multi-bit, so it is a request/acknowledge toggle handshake with the data held stable across it rather than a synchroniser per bit: at four unrelated clock ratios every received byte, frame count and edge count was identical. Rate invariance is the one property a simulation can genuinely establish about a crossing; it says nothing whatever about synchroniser depth.
Two defects found along the way are worth carrying. A divider that never reloads delays the first edge and nothing else, while the edge count and the data both stay correct — it surfaced only when a slave ran out of setup margin, looking like a crossing fault. And the bench's slave, clocked on the bus clock, undersampled a faster SCLK and produced two failures that looked exactly like a broken design: a monitor must be at least as fast as the thing it monitors, and a rate sweep is the experiment that finds out whether it is.
15. What Comes Next
One master, one slave, one configuration. Chapter 19.4 adds devices: several slaves on one bus, each with its own select, its own mode and its own maximum clock rate — and the configuration switch between devices becomes the bug factory, because changing CPOL changes SCLK's idle level while another device is still watching the bus.
Continue learning
Related tutorials
- Related topic
FPGA to ADC — Streaming Under a Sample Deadline
Throughput and a deadline are different requirements. The measured busy interval matches t_conv + lead + 2·NB·half + lag + gap exactly — and the fixed terms cap what any clock rate can buy.
- Related topic
FPGA to Sensor — Register Access and Event-Driven Reads
A sensor owns the schedule. An event arriving mid-read has exactly two answers with opposite costs — a lost sample or a mis-timestamped one — and which is right turns on one line of the datasheet.
- Related topic
Typical APB SoC Architecture
Where APB sits in a real SoC — downstream of the AXI/AHB bridge, fanning out to peripherals through a one-hot decode and return mux.
- Related topic
Bus Attachment and Integration Concerns
What attaching the register block to an APB- or AXI-class bus requires of the UART, what the UART must require of the integration, and the clock question — answered by measurement.
