Skip to content
VLSI Mentor

SPI · Module 3

Mode Mismatch and Its Failure Signature

What happens when the two ends disagree about the mode. The distinct signature each mismatch produces, how to tell polarity from phase disagreement from the data alone, and the monitor and coverage work that catches it.

Every chapter in this module has assumed both ends of the link agree about the mode. This one removes that assumption, and in doing so turns the module's theory into a diagnostic method.

When a master and a device disagree about the mode, what exactly goes wrong — and can you tell which bit is wrong from the corrupted data alone?

The answer to the second question is mostly yes, and that makes this the most immediately useful chapter in the module. A mode mismatch is one of the most common SPI bring-up faults, it produces symptoms that are frequently misattributed to electrical problems, and the distinguishing evidence is usually already in front of you.

1. Three Kinds of Disagreement

Two bits, two ends. If the ends differ, they differ in the polarity bit, the phase bit, or both — and the three cases are genuinely different failures.

DisagreementModes involvedWhat is displaced
Phase only (CPHA differs)0↔1, 2↔3sampling instant moves by half a bit time
Polarity only (CPOL differs)0↔2, 1↔3leading/trailing roles swap; sampling moves by half a bit time
Both differ0↔3, 1↔2the two displacements cancel

The third row is the surprising one and it deserves its own section, because it is the case that produces the hardest bugs.

2. Why Disagreeing About Both Bits Often "Works"

Chapter 3.3 §4 and Chapter 3.7 §3 established that Modes 0 and 3 present the same physical edge directions — both launch on falling and sample on rising. Chapter 3.6 §5 showed the same for Modes 1 and 2.

So if a master is configured for Mode 0 and a device expects Mode 3, both are launching on falling and sampling on rising. The bits are exchanged correctly. The only disagreement is about the level the clock rests at between transfers — and if neither end cares, the link works.

That is not a bug that hides; it is a genuine equivalence for the transfer itself. Where it stops being equivalent is exactly where Chapter 3.7 §4 said it does:

The first bit. Mode 0 is CPHA = 0 and expects the first bit launched at CS; Mode 3 is CPHA = 1 and launches it on the first edge. A 0↔3 disagreement therefore corrupts bit one and nothing else — if the first bit carries meaning at all.

The idle level. A device that does care where the clock rests will see the wrong level, and may treat the master's first clock transition as spurious.

So a 0↔3 mismatch presents as either perfect operation or a single wrong first bit — never as the wholesale corruption the other two cases produce. Recognising that narrow signature is what stops an engineer from chasing a mode problem when they have a first-bit problem, or vice versa.

3. Phase Mismatch: the Half-Bit Displacement

The most common genuine mismatch, and the one with the clearest signature.

If the master samples on the leading edge and the device launches on the leading edge — a CPHA disagreement — then the master reads the line at the instant the device is changing it. Two things follow.

The captured value is the previous bit. In practice the device's output has not yet changed when the master samples, so the master captures bit k−1 where it expected bit k. Across a whole word that produces a one-position shift, with the first captured bit being whatever was on the line before the transfer and the last transmitted bit never captured at all.

The margin is gone. The master is sampling at the moment of transition, so whether it gets the old or the new bit depends on the device's output-valid delay (Chapter 2.6) against nothing at all. This is the aperture violation of Chapter 2.4 §2, arrived at by configuration rather than by frequency.

That second point explains something that otherwise looks contradictory: a phase mismatch can produce intermittent corruption rather than a clean shift. If the device's output happens to settle before the master samples, the master gets the new bit and the word looks correct. At a lower clock rate this becomes more likely, not less — so a phase mismatch can appear to improve when you slow the link down, which is exactly the behaviour engineers associate with a timing problem.

Phase mismatch — the master samples while the device is launching

10 cycles
An SPI transfer with a phase mismatch. The device launches each bit on the leading rising edge. The master also samples on the rising edge, at the same instant the line is changing, so each captured value is uncertain and the received word is shifted by one position relative to the transmitted one.three bit timesthree bit timesdevice launchesdevice launchesmaster captures old bitmaster captures old bitsclkmisoXXXXmaster getsXXXXXXt0t1t2t3t4t5t6t7t8t9
Figure 1 — a phase mismatch. The device launches on the leading edge while the master samples on it, so the master reads the line as it changes. The captured word is shifted by one position, and the value at each capture depends on the device's output delay rather than on any margin.

Read the third row as what the master ends up with. Each capture lands on a transition, so the value it records is the previous bit — visible as the whole row being displaced one bit time to the right of the row above it.

4. Polarity Mismatch: Roles Swap

If only CPOL differs, something subtler happens: the two ends still agree on which logical edge samples, but they disagree about which physical edge is leading.

Suppose the master is configured CPOL = 0, CPHA = 0 (Mode 0) and the device expects CPOL = 1, CPHA = 0 (Mode 2). The master treats rising as leading and therefore samples on rising. The device treats falling as leading and therefore samples on falling — and launches on rising.

So the master samples on rising while the device launches on rising. That is the same collision as a phase mismatch, arrived at differently, and it produces the same one-position displacement with the same absent margin.

This is why the table in §1 lists both single-bit disagreements as displacing the sampling instant by half a bit time. Their mechanism differs — one swaps roles, the other redefines which direction is leading — but their signature in the data is the same.

There is one additional symptom that belongs to polarity alone, and it is the discriminator: the idle level is visibly wrong. A phase mismatch leaves SCLK resting where the device expects it; a polarity mismatch does not. That observation is free, requires no data interpretation, and separates the two cases immediately — the habit Chapter 2.2 §5 established, now earning its keep.

5. The Diagnostic Procedure

Assemble the whole module into a sequence. Each step is cheap and each one partitions what remains.

Step 1 — is it reproducible at every clock rate? If corruption is identical at 1 MHz and 20 MHz, it is configuration, not margin (Chapter 2.4). If it improves monotonically as you slow down and becomes perfect, it is margin, not configuration. If it improves erratically and never becomes reliable, suspect a phase or polarity mismatch, for the reason §3 gave.

Step 2 — read the idle level. Capture SCLK with CS deasserted and compare against the device's datasheet. Wrong level means a polarity disagreement, and you are done. Right level means polarity agrees, leaving phase.

Step 3 — is the whole word displaced, or just one bit? A whole-word one-position shift is a phase or polarity mismatch. A single wrong bit in the most significant position is the first-bit path (Chapter 3.4) or a 0↔3 style disagreement — different faults, both narrow.

Step 4 — is the transmit direction affected too? A mode mismatch is symmetric: the device misreads MOSI just as the master misreads MISO. If only the received direction is corrupt while commands are obeyed correctly, that is not a mode mismatch — it is a return-path timing problem (Chapter 2.7), and this step is what prevents the two being confused.

Step 4 is the one most often skipped and the most valuable. Mode mismatches and round-trip shortfalls both produce corrupt received data; only one of them also corrupts the command the device receives.

6. Detecting It in Verification

A mode mismatch is trivially detectable in simulation if the environment is built to look for it, and invisible if it is not.

The monitor must not adapt. A monitor that samples on whichever edge the DUT uses will reconstruct transactions perfectly regardless of configuration — and will therefore never detect a mismatch. The monitor must sample according to the configuration object (Chapter 3.2 §6), because that object represents what the device expects. When the DUT disagrees with it, the monitor's reconstruction diverges from the scoreboard's expectation and the test fails, which is the desired behaviour.

A checker can detect the collision directly. The characteristic of a mismatch is that the line transitions at or immediately before the sampling edge. That is observable:

Azvya Education Pvt. Ltd.VLSI Mentor
spi_mode_mismatch.sva — the line must be stable across the sampling edge
   // A mode mismatch makes the transmitter change the line at the instant the
   // receiver samples it. Checking that the sampled value is the same one
   // cycle either side of the sampling strobe catches that collision.
   //
   // ASSUMPTION: `clk` is fast enough relative to SCLK that a bit time spans
   // several clk cycles -- true for an oversampling monitor, NOT true if the
   // monitor is clocked by SCLK itself. Stated because the property is
   // meaningless otherwise, not as a footnote.
   property p_stable_across_sample;
       @(posedge clk) disable iff (!rst_n || cs_n)
           sample_stb |-> (miso == $past(miso, 1));
   endproperty

   a_stable_across_sample : assert property (p_stable_across_sample)
       else $error("MISO changed at the sampling edge -- mode mismatch or t_v violation");

What it proves and what it does not. It proves the line was not transitioning across the sampling instant, which catches both a mode mismatch and a t_v shortfall — those are genuinely the same observable event, and the assertion correctly cannot distinguish them. Separating them is §5's frequency test, not a property. It also cannot prove the configured mode is right for the device, which is a datasheet comparison no signal-level check can perform.

A predictor makes the mismatch loud rather than subtle. If the environment holds the device's expected mode and the DUT's configured mode as separate values, a mismatch is detectable by comparison before any traffic runs — a configuration check rather than a functional one. That is the cheapest possible detection and it belongs in the environment's build phase.

7. Coverage That Would Have Caught It

Chapter 3.3 §7 built a mode coverage model. This chapter adds the cross that specifically targets mismatch.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_mode_match_cg.sv — cover agreement AND disagreement between the two ends
   // Covering the DUT's mode alone cannot reveal a mismatch, because a mismatch
   // is a relationship between two configurations. Cover the PAIR.
   covergroup spi_mode_match_cg @(posedge transfer_done);
       cp_dut_mode : coverpoint dut_cfg.mode  { bins m[] = {[0:3]}; }
       cp_dev_mode : coverpoint dev_cfg.mode  { bins m[] = {[0:3]}; }

       // The interesting axis is the RELATIONSHIP, not either value.
       cp_relation : coverpoint mode_relation {
           bins matched        = {MATCH};
           bins phase_only     = {PHASE_DIFF};
           bins polarity_only  = {POLARITY_DIFF};
           bins both_differ    = {BOTH_DIFF};     // the 0<->3 / 1<->2 case
       }

       // Chapter 3.7 §7: modes 0 and 3 are the SAME physical configuration, so
       // binning by mode number alone can report broad coverage while only one
       // physical arrangement has ever run. Cover that explicitly.
       cp_physical : coverpoint dut_cfg.mode {
           bins launch_falling = {0, 3};
           bins launch_rising  = {1, 2};
       }

       x_relation_first_bit : cross cp_relation, cp_first_bit_meaningful;
   endgroup

Three points about why these bins and not others.

The relationship is the coverage target, not the mode. A mismatch is a property of a pair of configurations, so a coverpoint on the DUT's mode cannot express it. This is the general lesson — cover the thing that can be wrong.

BOTH_DIFF needs its own bin precisely because it usually passes. §2 showed a 0↔3 disagreement transfers data correctly and fails only at the first bit. A coverage model that lumps all disagreements together will report the mismatch space as covered while never exercising the case whose failure is narrowest and hardest to spot.

The physical-arrangement coverpoint closes Chapter 3.7 §7's trap. Modes 0 and 3 are one physical configuration; 1 and 2 are the other. Binning them that way makes it impossible to claim broad mode coverage while having exercised only launch-on-falling.

8. Why This Chapter Has No New RTL

Deliberate, and the justification is the chapter's subject.

A mode mismatch is not something hardware implements — it is a disagreement between two correctly-implemented endpoints. There is no module to write, because neither end is doing anything wrong internally; each is faithfully executing a configuration, and the fault lives in the relationship between them.

What is buildable here is verification and diagnosis, and §6 and §7 supply both. Writing a synthesizable module would require inventing hardware that deliberately misbehaves, which teaches nothing an engineer will use.

The correct representations for this chapter are the signature table of §1, the waveform of §3, the diagnostic procedure of §5 and the coverage relationship of §7 — the things that let someone with a broken board reach the cause in a few minutes rather than an afternoon.

9. Failure Signature — Nothing Works, and It Looks Electrical

Symptom. A newly assembled board's SPI link returns data that is wrong in a way that looks random. Slowing the clock changes the corruption pattern but never fixes it. The team begins investigating signal integrity, trace lengths and termination.

Plausible mechanisms. A phase or polarity mismatch is the leading candidate, and the frequency behaviour is why. Competing explanations are genuine margin exhaustion, MISO contention, and an over-clocked device.

The discriminating observations, in §5's order. The frequency response is already suggestive: a real margin problem becomes perfect below some rate, and this one does not. That single observation should redirect the investigation away from signal integrity before any probe is attached.

Then read the idle level — free, and it settles polarity outright. Then ask whether the device is also misreading commands: if the device executes the wrong operation, or none, the corruption is symmetric and a mode mismatch is confirmed. A return-path timing problem never corrupts the command direction (Chapter 2.7 §1).

Why the investigation goes wrong so often. Because "changes with clock rate but never becomes correct" reads as analogue to most engineers, and the first hypotheses are physical. The distinguishing feature is subtle — monotonic improvement to perfection versus erratic change without convergence — and it is easy to miss when you are watching a corruption pattern rather than an error rate. Measuring the error rate at three or four clock frequencies and plotting it takes ten minutes and eliminates half the hypothesis space.

10. Common Misconceptions

11. Reason It Through

Work this before reading the answers.

A board is brought up with a new sensor. Reads return plausible-looking but wrong values. At 8 MHz roughly one byte in three is wrong; at 1 MHz roughly one in twenty; at 500 kHz still roughly one in twenty. The sensor's configuration registers, written over the same link, are demonstrably taking effect — the device's behaviour changes as expected when they are written. The team suspects a marginal board.

What does the frequency response establish? That this is not a margin problem. A genuine margin shortfall improves monotonically to perfection as the period lengthens — Chapter 2.4 §4's arithmetic guarantees it, because the budget grows without limit while the delays stay fixed. This error rate falls from 8 MHz to 1 MHz and then stops falling. A floor that persists as the clock keeps dropping means something is wrong that time does not fix.

What does "configuration writes take effect" tell you? This is the strongest clue in the statement and it points away from the obvious conclusion. The device is correctly receiving data the master transmits — so MOSI is being sampled correctly by the device. A mode mismatch is symmetric: if the master and device disagreed about which edge samples, the device would misread MOSI just as the master misreads MISO, and the configuration writes would not work.

So what is left? Something that corrupts the return direction only. That is Chapter 2.7's territory — but the frequency behaviour rules out a plain round-trip shortfall too, for the same reason it ruled out margin.

The remaining candidate that fits every observation is a fault in the device's transmit path or the master's capture that is not time-dependent — for example the master sampling MISO on the wrong edge while sampling MOSI-side timing correctly, which happens when a controller has separate transmit and receive edge configuration and only one is wrong. Some controllers genuinely expose these separately; others have a receive-sampling-delay option (Chapter 2.7 §6) that can be misconfigured independently of the mode.

Why does the rate still improve between 8 MHz and 1 MHz, then floor? Because two effects are present. Slowing down removes a genuine margin component — real, and it accounts for the improvement. The floor is the configuration error underneath, which no amount of slowing fixes. Two mechanisms at once is the normal case on a real board, and expecting a single cause is why the frequency evidence gets misread.

What would you do next? Read the controller's receive-path configuration and compare it against the transmit path's. Then scope MISO at the master's pin against the master's actual sampling edge — not the edge you believe it uses — because the whole hypothesis is that those differ. The asymmetry between directions is the finding; preserving it through the investigation is what stops the team from re-laying out a board that is electrically fine.

12. Understanding Check

13. Summary

A mode mismatch is a disagreement between two correctly-implemented endpoints, and its form depends on which bit differs.

Phase only or polarity only both displace the sampling instant by half a bit time, so the master samples while the device is launching. The data arrives shifted by one position, and — because the master is reading a transitioning line with no margin — the corruption is rate-dependent and erratic. That makes a mismatch look like a timing problem, and the discriminator is that a genuine margin shortfall improves monotonically to perfection while a mismatch floors at a non-zero error rate.

Both bits differing — 0↔3 or 1↔2 — is the dangerous case, because the two displacements cancel. The ends still agree on physical edge directions, so data transfers correctly and the link fails only at the first bit, and only if that bit carries meaning.

The diagnostic sequence is four cheap steps: frequency response separates configuration from margin; idle level separates polarity from phase; whole word versus one bit separates a mismatch from a first-bit-path fault; and whether the transmit direction is also corrupt separates a mode mismatch from a return-path timing problem, since a mismatch is symmetric and a round-trip shortfall is not. That last step is the most valuable and the most often skipped.

In verification, a monitor must sample by the configuration object rather than by the DUT, or it adapts to the fault and can never detect it. And coverage must target the relationship between the two ends' configurations, with the both-differ case binned separately because it usually passes — plus a physical-arrangement coverpoint, because modes 0 and 3 are one arrangement rather than two.

14. What Comes Next

That closes Module 3. Both configuration bits have names, the four modes are derived rather than memorised, each has been examined in working depth, and their disagreements have a diagnostic procedure.

Module 2 taught when a bit is valid; Module 3 taught which edge makes it so. Module 4 now asks what the bits mean: transfer width, bit ordering, and the command, address, dummy and data phases a device layers onto the stream — beginning from the principle Chapter 1.4 established, that SPI defines the signalling and the device defines everything above it.

Browse the path on the SPI curriculum index, or revisit Deriving Mode Behaviour for the method this chapter inverts into a diagnosis.

Continue learning