SPI · Module 1
Full-Duplex Exchange
Every SPI transfer moves a bit in both directions on every edge, whether the software wanted it to or not. Where dummy bytes come from, why bytes received during a command phase exist but mean nothing, and why read and write are interpretations rather than modes.
Chapter 1.3 built the ring: one shift register per end, wired so that each end's serial output feeds the other's serial input, both stepped by the same clock. That structure has a consequence the previous chapter named but deliberately did not develop, and it is the one that reshapes how you read every SPI datasheet afterwards.
There is no transmit-only transfer and no receive-only transfer. A bit leaves and a bit arrives at each end on every edge, because those are the two ends of one shift. An SPI link does not offer a direction; it offers an exchange, and everything a device calls a read or a write is a convention layered on top of that exchange.
This chapter takes that from a structural fact to a working habit: where the extra bytes in a read command come from, why a received byte is never by itself evidence of anything, and what it does to a verification environment.
1. The Consequence the Ring Forces
Run the argument once, slowly, because it is short and everything else depends on it.
A shift register has one serial output and one serial input. On an enabled edge it shifts once: the bit at the output end leaves, every other bit moves one position, and the bit at the input end is filled from the serial input. The width is fixed, so a departing bit and an arriving bit are not two events that happen to coincide — the position one vacates is the position the other fills.
Both ends of an SPI link contain such a register, and the same clock steps both. Therefore, on every enabled edge:
- the master presents a bit on MOSI and captures one from MISO,
- the slave presents a bit on MISO and captures one from MOSI.
There is no control signal in the structure that could switch one of those off. To "not transmit," a master would have to present nothing on MOSI — but the register's output stage always holds some value, and Chapter 1.2 established the master owns that driver unconditionally. To "not receive," it would have to leave a bit position unfilled, which a fixed-width shift cannot do.
2. There Is No Receive-Disable Bit
It is worth being blunt about this, because the mental habit imported from other buses is strong.
On a bus with separate request and response transactions, a write is a thing you do and a read is a different thing you do; there is no sense in which a write "also reads." On SPI there is. Every byte a master sends produces a byte it receives, at the same time, on the same edges — and every byte a slave returns was clocked out by a master that was simultaneously sending something.
So a master's driver code always ends up with two buffers, even when it wants one. A transmit-only operation discards the received buffer. A receive-only operation must invent a transmit buffer, because the clock cannot run without the master presenting bits. Neither operation avoided the exchange; each one ignored half of it.
That invention has a name, and it is the subject of the next two sections.
3. A Real Read, Traced
Take an ordinary register read from a peripheral: the master sends a one-byte command that identifies the register and marks the access as a read, then reads back one byte of data. Two byte times, sixteen bit times, CS held low across both.
| Byte time | MOSI carries | MISO carries |
|---|---|---|
| 1 (bit times 1–8) | 0x8B — the opcode, meaningful | undefined — the device has not been told what to answer yet |
| 2 (bit times 9–16) | 0x00 — filler, ignored by the device | 0x42 — the register contents, meaningful |
Sixteen bits crossed MOSI and sixteen crossed MISO. In each byte time exactly one of those directions carried meaning, and the other carried bits anyway. Zoom into the second byte time, where that is most visible: the master is driving MOSI for all eight bit times with a value the device will discard, purely because driving the clock is the only way to get the answer out.
Response byte — MISO answers, MOSI still carries filler
8 cyclesFour observations, and none of them are about the device.
Both lines are active in both byte times. The hardware did not do less work during the half of each byte time that carries no meaning. A "read" cost the master exactly as much MOSI activity as a write of the same length would have.
MISO during the command byte is undefined, not zero. The slave has not yet been told what to answer. What it presents there is device-specific — it may hold a fixed level, echo something, or present the tail of a previous transfer. The honest engineering position is that the specification does not oblige it to be anything, so a design that depends on its value depends on undefined behaviour.
MOSI during the data byte is filler. The master had nothing left to say, but it had to keep clocking to get the answer out, and Chapter 1.2's ownership rule says it is driving MOSI the whole time regardless. So it drives something — here all zeros.
Neither "phase" is a protocol feature. SPI has no notion of a command byte or a data byte. The device defines that the first byte it receives after selection is an opcode and that it will answer from the ninth bit time onward. Read a different part and the split lands somewhere else.
4. Where Dummy Bytes Come From
The filler in the second byte time is what datasheets call a dummy byte — and once you have the ring, it stops being a quirk and becomes arithmetic.
The master needs N bits of answer. Bits only move when the clock runs. The clock only runs while the master drives it, and while it drives it the master is also presenting bits on MOSI. Therefore getting N bits of answer costs N bit times of MOSI traffic that the device will ignore. Dummy bytes are not overhead the protocol adds; they are the MOSI side of the clock the master had to generate anyway.
Two practical consequences that catch people out.
The dummy value is usually free, but not always. Most devices ignore MOSI entirely during a response phase, so 0x00 and 0xFF are equally fine. Some do not — a device may allow a new command to be pipelined into those bit times, or may treat a particular pattern specially. The value is a device question, and the datasheet answers it.
Dummy bytes are distinct from dummy cycles. The filler discussed here exists because the master must clock to receive. Some devices additionally require a number of idle bit times before their data is valid, because they need internal time to fetch it — a latency requirement, not a clocking artefact. The two often coincide on the wire and are frequently conflated. They have different causes: one is the consequence of this chapter, the other is a device timing parameter. Module 4 separates them properly, and Module 11 shows the case where it matters most, fast reads from serial flash.
5. Read and Write Are Interpretations
Now name the layering explicitly, because it is the chapter's durable output.
Read the figure as a claim about where information lives. The hardware layer is symmetric and unconditional: it moves the same number of bits both ways in both cases. The asymmetry that makes one exchange a "read" and another a "write" is introduced entirely by the device contract — and a contract is a document, not a signal.
This is the same ownership-versus-meaning separation Chapter 1.2 drew, now applied one level up. There it distinguished who drives a wire from what the bits on it mean. Here it distinguishes what the hardware did from what we agree to call it.
6. The Vocabulary Trap
Because the words are unavoidable, it is worth stating exactly how far they can be trusted.
"Read" and "write" are perfectly good names for operations a driver performs against a device. They are not names for anything the SPI link does. When a datasheet says "read register 0x0B," it is describing an exchange whose MOSI half the device will interpret as an opcode and whose MISO half will, after a defined point, carry the register's contents.
The failure this produces is specific and common: an engineer reasons "this is a write, so MISO is irrelevant," configures a controller or a testbench accordingly, and is then unable to explain a symptom that MISO was showing all along — a slave that never released the line, an unselected device answering, or a return path stuck at a constant level. The bits were there. The model said not to look.
Keep the two vocabularies separate and the problem disappears. At the wire, there are exchanges. At the device, there are reads and writes. Translating between them is your job, and doing it consciously is what Module 10 trains.
7. Why an RTL Designer Cares
Three consequences reach the hardware, and the first is the one that changes an interface definition.
A transfer primitive has two data ports, not one. The natural interface for an SPI master is tx_data in, rx_data out, one start, one done — because that is what the shift core in Chapter 1.3 actually does. A design that exposes separate "send byte" and "receive byte" operations is not describing the hardware; it is describing two use cases, and it will need an invented transmit value for one of them anyway. Making both directions explicit at the interface pushes the filler decision up to the caller, which is where the device knowledge lives.
Received data must be captured even when nobody asked for it. There is nowhere to put a bit other than the register. A master that "optimises" by not updating rx_data during a write has saved nothing — the shift happened regardless — and has removed the evidence a debugger needs. Capture it; let software ignore it.
The filler value belongs in configuration, not in the datapath. Because a few devices care what is sent during a response phase, a reusable master lets the caller choose the byte it transmits while receiving rather than hard-wiring zeros. It is a small parameter that prevents a whole class of "works on part A, fails on part B" integration problems. Where it sits in a real master's register map is Module 13.
8. Why a Verification Engineer Cares
This chapter has a sharper effect on a testbench than on a design, because it changes the shape of the transaction object at the centre of the environment.
A transfer is not typed by direction. The instinct is enum { READ, WRITE } plus one data array. The hardware says a transfer has two equal-length byte streams, always, and that which one matters is a device-layer question the agent should not be encoding. Model what happened, not what it was for:
class spi_item extends uvm_sequence_item;
`uvm_object_utils(spi_item)
// ONE transfer = TWO byte streams of EQUAL length. There is no
// direction field, because the hardware has no direction.
rand byte unsigned mosi_bytes[]; // what the master drove
byte unsigned miso_bytes[]; // what came back — OBSERVED, never randomised
rand int unsigned n_bits; // bit times between the CS edges
int unsigned cs_index; // which slave was selected (Chapter 1.5)
constraint c_whole_bytes { n_bits % 8 == 0;
mosi_bytes.size() == n_bits / 8; }
function new(string name = "spi_item");
super.new(name);
endfunction
function string convert2string();
return $sformatf("cs=%0d bits=%0d mosi=%p miso=%p",
cs_index, n_bits, mosi_bytes, miso_bytes);
endfunction
endclassNote the two deliberate asymmetries. miso_bytes is not rand: a monitor observes it and a slave model produces it, but no stimulus generator invents it. And there is no is_read field — a sequence that wants to perform a device read constructs the right mosi_bytes (opcode, then filler) and reads its answer out of miso_bytes. The device semantics live in the sequence layer, where the datasheet knowledge belongs, not in the item.
A monitor gets the second direction for free. Because both lines are sampled on the same edges, reconstructing both streams costs one extra push_back — there is no separate receive monitor, because there was no separate receive:
// Fragment: assumes a capture-on-rising-edge configuration and ignores
// reset. A complete monitor — mode-aware sampling, CS glitch handling,
// partial frames — is Module 16.
task automatic collect_transfer(output spi_item item);
bit mosi_bits[$], miso_bits[$];
@(negedge vif.cs_n); // transaction opens
fork
forever begin
@(posedge vif.sclk); // the configured capture edge
mosi_bits.push_back(vif.mosi);
miso_bits.push_back(vif.miso); // same edge — the bit is already there
end
join_none
@(posedge vif.cs_n); // transaction closes
disable fork;
item = spi_item::type_id::create("item");
item.n_bits = mosi_bits.size(); // bit COUNT is itself a check (Ch. 1.3)
pack_msb_first(mosi_bits, item.mosi_bytes);
pack_msb_first(miso_bits, item.miso_bytes);
endtaskAn assertion can prove the line is driven, but never that it is meaningful. The invariant this chapter exposes is that the master drives MOSI for the entire transfer, including the bit times whose content the device ignores:
// The master owns MOSI unconditionally (Chapter 1.2), so an unknown value
// at a capture edge is a driver bug — not a legitimate "don't care".
property p_mosi_driven;
@(posedge sclk) disable iff (cs_n)
!$isunknown(mosi);
endproperty
a_mosi_driven : assert property (p_mosi_driven)
else $error("MOSI unknown at a capture edge while CS is asserted");Be precise about what that proves. It proves the master is driving a defined level at every sampled edge — catching a real bug class where a filler byte is left undriven or a tri-state is mistakenly enabled on the master side. It proves nothing about whether those bits are the right bits, because "right" is defined by a datasheet and no assertion on the pins has access to one. The same distinction applies to MISO: an assertion can require it to be known during the bit times the device is specified to answer, but only the sequence layer knows when that is.
Everything above is a fragment chosen to teach the exchange. The complete environment — interface and clocking blocks, driver, mode-aware monitor, reference model, coverage, and the agent that holds them — is Modules 16 and 17.
9. Why an FPGA Engineer Cares
The practical consequence is about buffering and about what you can see.
Both directions consume buffer capacity. A design streaming from an ADC needs somewhere to put the samples, which is expected. What surprises people is that a design streaming to a DAC is also receiving a byte per byte sent, and if that byte is captured into a FIFO that nobody drains, the FIFO overflows and — depending on how the overflow is handled — can stall the transmit side. The received stream is real traffic even when it is meaningless traffic, and it must be either drained or explicitly discarded in hardware.
The ignored direction is your best diagnostic. When an SPI link misbehaves on a board, the half of the exchange the application discards is frequently where the evidence is. MISO sitting at a constant level through a transfer where the device should have answered, or showing activity during a write where nothing should be driving, tells you about selection and output-enable problems that the application-visible data never will. Capture both directions in your integrated logic analyser setup even when the design only uses one — the cost is a few block RAM bits and it repeatedly pays for itself. Module 18 turns that into a debugging method.
10. Common Misconceptions
11. Reason It Through
Work this before reading the answer.
A driver performs a device write: one command byte, then two data bytes, with CS held low across all three. The author reasons that MISO is irrelevant to a write and leaves the controller's receive FIFO undrained. On the bench, the first few writes succeed and later ones stop taking effect, with no error reported anywhere.
What is the hardware doing on MISO during those three bytes? Exactly what it does during any transfer: shifting in twenty-four bits, one per bit time, and delivering three received bytes into the controller's receive path. The operation being a "write" changed nothing about that.
Why does an undrained receive FIFO stop the writes? Because the receive path fills. Three bytes arrive per write whether or not anyone reads them, so after a small number of transfers the FIFO is full. What happens next is controller-specific and both common behaviours are bad here: some controllers stall the transfer engine until space is available, which halts writes silently; others drop the incoming byte and set an overrun status bit that this driver never reads. Either way the symptom is writes that quietly stop working — and no error in the sense the author was looking for, because from the device's point of view nothing illegal occurred.
Why did the first few succeed? Because a FIFO has depth. This is the signature worth memorising: a fault that appears only after a number of operations, with that number proportional to a buffer size, is a drain problem rather than a protocol problem. It also explains why it survived every short test.
What is the correct fix, and what is the tempting wrong one? The correct fix is to drain the receive path — read and discard the bytes, or enable a discard mechanism if the controller provides one. The tempting wrong one is to reduce the transfer length or add delays until the symptom disappears, which hides the drain problem behind timing and guarantees its return under different load.
Where should this have been caught? In review, from the interface: a transfer primitive whose signature returns received data makes the obligation visible, while one that hides it invites exactly this bug. That is §7's point about two data ports, arriving as a concrete failure rather than a style preference.
12. Understanding Check
13. Summary
An SPI transfer is an exchange, and not by choice. One shift register per end, wired into a ring and stepped by one clock, means the bit position each departing bit vacates is filled by an arriving bit — so a bit leaves and a bit arrives at each end on every edge. Nothing in the structure could disable one direction.
From that, the vocabulary follows. A device read is an exchange whose MOSI half the device interprets as a command and whose MISO half carries the answer; the master still drove bits for the whole transfer, and those bits are the dummy byte — the MOSI side of a clock it had to generate to receive anything. A device write is an exchange whose MISO half the driver discards; a byte still arrived on every byte time. Read and write are interpretations a datasheet imposes on identical hardware behaviour, not modes the link has.
Two habits follow and are worth carrying. A received byte is never, by itself, evidence that a device responded — it only proves that bit times elapsed, and interpreting it requires knowing that the right device was selected and that those bit times fell inside its defined response phase. And the half of the exchange the application ignores is still real: it consumes buffer capacity that must be drained, and it carries the diagnostic evidence that selection and output-enable faults leave behind.
In verification, the shape of the transaction object is the tell. A transfer holds two equal-length byte streams and no direction field, because that is what the hardware produces; a monitor reconstructs both from one edge loop because there was only ever one shift. Assertions on the pins can prove a line is driven; only a reference model holding the device's contract can say whether the bits are right.
14. What Comes Next
So far every chapter has assumed one master and one slave. Chapter 1.5 — Bus Topologies and Daisy-Chain removes that assumption and asks what the four-wire model costs as devices are added: what can be shared, what cannot, how the select lines scale, and what the daisy-chain topology does differently by turning several devices into one long shift ring — the ring of Chapter 1.3, extended across chips.
Browse the path on the SPI curriculum index, or revisit The Shift-Register Mental Model for the hardware this chapter's consequence comes from. For a link where the two directions genuinely are separate wires with independent timing, see The UART Link: TX, RX, Idle and Full Duplex — a useful contrast, because there full duplex really is two independent channels rather than one shared shift.
Continue learning
Related tutorials
- Related topic
Extracting Protocol Rules and the Verification Plan
Eight pin-observable SPI rules, each with a checker and an exercised counter, because a checker alone cannot tell never-broken from never-reached. Legal traffic violates nothing and exercises all eight; eight injected faults produce a diagonal violation matrix; and one plan row is proved to have no checker at all.
- Related topic
Driver Architecture
A driver owns every timing number in the protocol, so it is the one component that must be checked against something not written to agree with it. Thirty-two legal transactions violate nothing and exercise all eight rules; seven injected faults fire exactly the rules predicted.
- Related topic
Monitor and Transaction Reconstruction
Two instances of one monitor on one set of pins, differing only in where they got CPHA. The specification's view catches an injected fault; the implementation's view reports the intended word with total confidence, and passivity is measured with a pin hash.
- Related topic
Reference Model and Scoreboard
Two predictions of the same traffic through one scoreboard: one computed from the transaction, one produced by a second copy of the design. Both report a clean run on a correct design, and with a fault injected the copy records a pass on every transaction.
