USB · Module 29
USB on FPGA Development Boards
The connector on the board does not reach the FPGA — it reaches a bridge chip, and what arrives on the pins is a byte FIFO with two active-low flags and a bus turnaround. Built as a synchronous FIFO bus master, where the bug that matters starves one direction forever and corrupts nothing.
The last case study, and the one where USB is barely present. Everything in this track so far has been about what happens on the wire. On a development board the wire usually stops at a chip that is not yours, and what reaches the FPGA is something else entirely — with its own flags, its own clock, its own latency, and its own ways of being got wrong.
1. The Connector On The Board Is Not Your USB
Plug a development board into a laptop and a lot appears to happen. A serial port shows up. A programmer finds the device. A console prints. It is easy — and almost universal — to conclude that the FPGA is speaking USB.
It is not. On the overwhelming majority of boards the USB connector goes to a bridge chip, and the FPGA never sees a USB signal at all.
WHAT THE LEARNER PICTURES WHAT IS ON THE BOARD
------------------------- -----------------------------------
host <-- USB --> FPGA host <-- USB --> bridge chip
|
a byte FIFO,
two flags and
a shared bus
|
FPGAThe bridge is a complete USB device. It enumerates, it has descriptors, it has endpoints, it does the whole of chapters 6 through 13 — and it does all of that before the FPGA is even configured. By the time your RTL runs, USB has already happened.
What is actually between the host and your RTL
2. What Is Actually On The Pins
Six signals and a clock, and the clock is the first surprise.
clk OUTPUT of the bridge, INPUT to the FPGA. Typically 60 MHz.
The FPGA is the slave of somebody else's clock, which means
this interface is a clock-domain boundary before it is
anything else.
data[7:0] BIDIRECTIONAL. One bus, both directions, so it has to be
turned around and only one end may drive it at a time.
rxf_n active low: the bridge is presenting a byte to be read
txe_n active low: the bridge has room to be given a byte
rd_n active low: take the presented byte
wr_n active low: here is a byte
oe_n active low: the bridge may drive the data busTwo flags in, three strobes out, one bus shared. That is the entire interface, and it is a good deal more subtle than its size suggests.
3. The Latency Timer, And Why A Round Trip Takes Sixteen Milliseconds
Before the RTL, the single most common false conclusion about USB on a dev board, because it sends people looking in the wrong place for days.
A user writes a four-byte command from the host, the FPGA answers in a microsecond, and the host sees the answer around sixteen milliseconds later. Every time. The natural reading is that USB is slow, or that the FPGA is slow, or that the driver is broken.
None of those. The bridge has a latency timer.
THE BRIDGE'S PROBLEM
It is a USB device. It cannot send anything until the host asks, and
the host asks on a schedule. When the FPGA hands it four bytes, the
bridge has a choice:
send them now one USB transaction for four bytes -- correct,
and catastrophic for throughput if the FPGA is
streaming, because every byte pays a full
transaction's overhead
wait a bit in case more bytes arrive and can be packed
into the same transaction
THE COMPROMISE
A timer, restarted on every byte from the FPGA. When it expires, send
whatever is buffered. The default on the common parts is a few
milliseconds up to about sixteen, and it is CONFIGURABLE from the host
driver.That is the whole of section 3, and there is no RTL in it, because there is no RTL decision in it. It is here because it is the first thing that goes wrong on a board and the last thing anyone suspects.
4. Three Decisions The Bus Master Has To Make
Now the part that is yours. Six signals, and three decisions.
The bus is bidirectional, so changing direction costs a cycle
One data bus, two drivers. The bridge drives it during a read and the FPGA drives
it during a write, and between those two states there must be a cycle in which
neither drives — because the bridge's output drivers do not turn off the
instant oe_n goes high, and if the FPGA starts driving into them the result is
not a wrong bit, it is a current path between two output stages.
... read read read | TAIL | TURN | write write ...
| |
| +-- nobody drives. No data moves here
| and the cycle is not optional.
+-- the strobe is off but the bridge's output
enable is still assertedThat dead cycle is pure cost: at 60 MHz, alternating single-byte transfers would spend a third of the bus on turnarounds. It is also why the design reads and writes in bursts rather than alternating byte by byte — and why the burst length is a parameter with a trade-off in it rather than a constant somebody picked.
The strobes leave the chip, so they are registered — and that costs a byte
rd_n goes off-chip to a part with a setup requirement measured in nanoseconds.
It cannot be the output of a deep cone of logic inside the FPGA. In this design it
is one gate from an input pin:
assign rd_n = ~(state == S_RD && ~rxf_n);state is registered and rxf_n is a pin. Nothing else is allowed in that
expression — and in particular the user's rx_ready is not, because
rx_ready comes from whatever the user built, which might be a FIFO almost-full
flag five levels of logic deep.
Which means the decision to read is made before it is known whether the user can take the byte. The byte arrives anyway. It has to go somewhere.
Read and write compete for one bus, and the obvious priority is a bug
This is the decision the chapter is really about.
Only one direction can use the bus at a time. So when both want it, something has to choose, and there is an obvious answer: give it to the reader. The argument is good — a byte the bridge is holding may be lost if the FPGA is slow to take it, whereas a byte the FPGA wants to send is sitting safely in the FPGA. Read is the urgent one.
Implement that, and the design is correct. Every byte that arrives is delivered, in order, intact. Every byte that is sent is sent correctly. The bus is never contended. The skid never overflows. Nothing is dropped, duplicated or reordered.
And if the host is streaming data in continuously, the FPGA transmits nothing, ever.
5. The Design (Verilog-2005)
The contract, written down before any code, because three languages have to implement the same thing:
PURPOSE master a synchronous FIFO bus to a USB bridge chip:
schedule the two directions, honour the flow-control
flags, turn the bus around safely, and absorb the byte
that a registered strobe commits to before the user's
readiness is known
INPUTS clk sourced by the BRIDGE, not by the FPGA
rst_n
rxf_n a byte is presented
txe_n there is room for a byte
data_in the presented byte
rx_ready the user can take one this cycle
tx_valid the user is offering one
tx_data
OUTPUTS rd_n, wr_n, oe_n, bus_drive, data_out
rx_valid, rx_data
tx_ready
state_o and five counters observation only
AUTHORITATIVE STATE
state one of six
burst bytes moved in this grant
last_was_rd which direction went last
sk0, sk1, wptr, rptr, cnt the two-deep skid
DERIVED, WITH NO STATE OF THEIR OWN
every strobe, from the state and at most one input pin
rx_valid, rx_data, tx_ready
PRIORITY, SAME CYCLE
rst_n beats everything
reads and writes ALTERNATE when both want the bus
a burst ends on the cycle its last permitted byte moves
the turnaround is unconditional between directions
LATENCY the read strobe is one cycle behind the decision, which
is what the skid exists to pay for. Everything else is
one clock.
BOUNDARIES the last byte in the bridge
the user stalling mid-burst
the bridge filling mid-burst
the user running dry mid-burst
both directions wanting the bus on the same cycle
a burst reaching exactly MAX_BURST
ASSUMPTIONS the bridge model in section 2
one clock domain: clk is the bridge's, and the user side
is assumed to be in it. See section 16.
OMISSIONS the bridge's USB side entirely, the latency timer, the
second channel, JTAG, the tristate pad itself, and any
clock-domain crossing to the user's logicThe datapath, and the one place the two directions meet
Six states, and only two of them move data
// =====================================================================
// ft245_sync_if -- the FPGA side of a USB-to-FIFO bridge in synchronous
// mode. This is what "USB" actually is on most FPGA development boards:
// not a USB controller in the fabric, but a bus master for a bridge
// chip that speaks USB on the far side and a byte FIFO on this one.
//
// CLASSIFICATION: simplified synthesisable teaching RTL.
// It is NOT an FT2232H / FT600 driver and it is not a USB controller.
// There is no USB here at all -- no PHY, no packets, no endpoints, no
// enumeration. Everything on the other side of the bridge has already
// happened. What is left is the part that runs in the FPGA, and the
// three decisions it has to get right:
//
// 1. THE BUS IS BIDIRECTIONAL, so changing direction costs a dead
// cycle in which no data moves and nobody drives.
// 2. THE STROBES GO OFF-CHIP, so they are registered. The decision to
// read is therefore made a cycle before the byte arrives, which is
// why there is a skid buffer and why it is two deep.
// 3. READ AND WRITE COMPETE FOR ONE BUS, so a priority that always
// favours one direction starves the other FOREVER. That is a
// liveness bug, and no amount of checking that the data is correct
// will find it.
//
// THE BRIDGE MODEL THIS ASSUMES (stated, because RTL cannot see a chip)
// --------------------------------------------------------------------
// rxf_n low the bridge is presenting a byte on data_in, and will go
// on presenting it until it is read. It does NOT withdraw
// a byte, so a read taken while rxf_n is low is always
// valid.
// rd_n low consumes the presented byte. The next cycle presents the
// following byte, or deasserts rxf_n.
// txe_n low the bridge has room for a byte.
// wr_n low hands it one, from data_out.
// oe_n must be asserted at least one cycle BEFORE rd_n, so the
// bridge's drivers are on before the strobe.
// =====================================================================
module ft245_sync_if #(
// The longest run of bytes in one direction before the machine must
// hand the bus to the other. This is the bound that turns "read has
// priority" into "read has priority for a while".
parameter integer MAX_BURST = 4
) (
input wire clk, // sourced by the BRIDGE, not by the FPGA
input wire rst_n,
// ---- bridge side -----------------------------------------------
input wire rxf_n, // active low: a byte is presented
input wire txe_n, // active low: there is room for a byte
input wire [7:0] data_in,
output wire rd_n,
output wire wr_n,
output wire oe_n,
output wire bus_drive, // we are driving data_out onto the bus
output wire [7:0] data_out,
// ---- user side, receive ----------------------------------------
output wire rx_valid,
output wire [7:0] rx_data,
input wire rx_ready,
// ---- user side, transmit ---------------------------------------
input wire tx_valid,
input wire [7:0] tx_data,
output wire tx_ready,
// ---- observation ------------------------------------------------
output wire [2:0] state_o,
output wire [15:0] n_rx,
output wire [15:0] n_tx,
output wire [15:0] n_turn,
// A byte arrived with nowhere to put it. Should be impossible; the
// skid depth is chosen so that it is. A non-zero value here means the
// depth argument in section 4 is wrong.
output wire [15:0] n_lost,
// We drove the bus within one cycle of the bridge's output enable. The
// electrical consequence is not simulable; the SCHEDULING error is.
output wire [15:0] n_contend
);
localparam [2:0] S_IDLE = 3'd0,
S_RDOE = 3'd1, // output enable leads the strobe
S_RD = 3'd2, // rd_n gated by rxf_n
S_TAIL = 3'd3, // the last committed byte settles
S_TURN = 3'd4, // the dead cycle between directions
S_WR = 3'd5;
reg [2:0] state;
reg [15:0] burst;
reg last_was_rd; // for the alternation in S_IDLE
reg [15:0] c_rx, c_tx, c_turn, c_lost, c_cont;
reg oe_n_d;
// ---- the skid buffer -------------------------------------------
// TWO deep, and the depth is the whole argument. rd_n is registered
// because it leaves the chip, so the decision to take a byte is made
// one cycle before the byte exists. One entry absorbs that byte; the
// second is what lets the machine keep reading on consecutive cycles
// while the user is draining, instead of every other cycle.
reg [7:0] sk0, sk1;
reg wptr, rptr;
reg [1:0] cnt;
// ---- combinational strobes -------------------------------------
// rd_n is gated by rxf_n and by NOTHING ELSE from outside this module.
// One gate from an input pin to an output pin closes timing on a
// bridge-clocked interface; a path from the user's rx_ready would not,
// which is exactly why the skid exists instead.
assign rd_n = ~(state == S_RD && ~rxf_n);
assign wr_n = ~(state == S_WR && ~txe_n && tx_valid);
assign oe_n = ~(state == S_RDOE || state == S_RD || state == S_TAIL);
assign bus_drive = (state == S_WR);
assign data_out = tx_data;
assign rx_valid = (cnt != 2'd0);
assign rx_data = rptr ? sk1 : sk0;
assign tx_ready = (state == S_WR) && ~txe_n && tx_valid;
wire take = ~rd_n; // a byte is consumed this cycle
wire drain = rx_valid & rx_ready; // a byte leaves the skid
// What the occupancy will be after this edge. The read decision uses
// this and not rx_ready, because rx_ready next cycle is not knowable.
wire [2:0] cnt_after = {1'b0, cnt} + (take ? 3'd1 : 3'd0)
- (drain ? 3'd1 : 3'd0);
// Room for the byte a further read would commit to.
wire room_next = (cnt_after <= 3'd1);
// ---- who gets the bus ------------------------------------------
wire want_rd = ~rxf_n & room_next;
wire want_wr = ~txe_n & tx_valid;
// Both want it: give it to whichever did not go last. Without this the
// design still passes every check that data is correct.
wire grant_rd = want_rd & (~want_wr | !last_was_rd);
wire grant_wr = want_wr & ~grant_rd;
// Keep going only while there is something to move and somewhere to
// put it. The BUDGET is deliberately not tested here: it is enforced
// one place only, in the exit condition inside each state, which fires
// on the cycle the last permitted byte moves. Testing it here as well
// is dead code -- the exit always fires first -- and a mutation that
// removed it scored zero in three languages across 187,957 checks
// before the duplication was noticed.
wire stay_rd = ~rxf_n & room_next;
wire stay_wr = ~txe_n & tx_valid;
always @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
state <= S_IDLE;
burst <= 16'd0;
last_was_rd <= 1'b0;
sk0 <= 8'd0;
sk1 <= 8'd0;
wptr <= 1'b0;
rptr <= 1'b0;
cnt <= 2'd0;
oe_n_d <= 1'b1;
c_rx <= 16'd0;
c_tx <= 16'd0;
c_turn <= 16'd0;
c_lost <= 16'd0;
c_cont <= 16'd0;
end else begin
oe_n_d <= oe_n;
// ---- the skid, independent of the state machine ----
if (take) begin
if (cnt == 2'd2 && !drain) begin
c_lost <= c_lost + 16'd1; // must never happen
end else begin
if (wptr) sk1 <= data_in; else sk0 <= data_in;
wptr <= ~wptr;
c_rx <= c_rx + 16'd1;
end
end
if (drain) rptr <= ~rptr;
case ({take, drain})
2'b10: if (cnt != 2'd2) cnt <= cnt + 2'd1;
2'b01: cnt <= cnt - 2'd1;
default: ; // 00 and 11 both leave it alone
endcase
if (tx_ready) c_tx <= c_tx + 16'd1;
if (bus_drive && !oe_n_d) c_cont <= c_cont + 16'd1;
if (state == S_TURN) c_turn <= c_turn + 16'd1;
// ---- the state machine ----
case (state)
S_IDLE: begin
burst <= 16'd0;
if (grant_rd) state <= S_RDOE;
else if (grant_wr) begin state <= S_WR; last_was_rd <= 1'b0; end
end
// One cycle of output-enable lead, with no strobe. The bridge is
// turning its drivers on; nothing is read here.
S_RDOE: begin
state <= S_RD;
last_was_rd <= 1'b1;
end
S_RD: begin
if (take) burst <= burst + 16'd1;
// Leaving because the budget ran out is DIFFERENT from leaving
// because there is nothing there, and only one of the two is
// about fairness.
if (!stay_rd || (take && (burst + 16'd1 >= MAX_BURST)))
state <= S_TAIL;
end
// The strobe is off but the output enable is still asserted: the
// bridge has not let go of the bus yet.
S_TAIL: state <= S_TURN;
// Nobody drives. This cycle moves no data and it is not optional.
S_TURN: state <= S_IDLE;
S_WR: begin
if (tx_ready) burst <= burst + 16'd1;
if (!stay_wr || (tx_ready && (burst + 16'd1 >= MAX_BURST)))
state <= S_TURN;
end
default: state <= S_IDLE;
endcase
end
end
assign state_o = state;
assign n_rx = c_rx;
assign n_tx = c_tx;
assign n_turn = c_turn;
assign n_lost = c_lost;
assign n_contend = c_cont;
endmodule6. What The Machine Actually Does
Three traces. In all of them a value shown in a cycle is sampled by the rising edge at the end of that cycle, and the registered result appears in the cycle after it.
A read burst, and the two cycles at each end that move no data
Four bytes read, and the four cycles of overhead around them
The turnaround, and the one cycle nobody may drive
Handing the bus from the bridge to the FPGA
The user stalls, and the skid earns its two entries
A byte already committed to, arriving at a user who stopped being ready
7. The Testbench (Verilog)
The checker here is a different shape from the one in 29.5, and the difference is the point.
29.5's design is a small pile of state, so a cycle-accurate model of that state is the strongest thing to compare against: predict every register, every cycle. This design is a bus scheduler. What matters is not what its burst counter holds on cycle 400 — it is that no byte is lost, no byte is reordered, the bus is never driven from both ends, and neither direction can be starved. A cycle-accurate twin of this state machine would mostly be a second copy of the state machine, and a copy agrees with its original about the wrong things.
So the checker is a set of properties plus two byte streams, and only the output equations are compared directly.
P1 rd_n low implies rxf_n low never over-read
P2 wr_n low implies txe_n low and tx_valid never write garbage
P3 bus_drive implies oe_n was high last cycle the turnaround
P4 rd_n low implies oe_n was low last cycle the OE lead
P5 the skid never overflows the depth argument
P6 every byte the bridge gave is delivered once, in order
P7 every byte the user offered is written once, in order
P8 a transmit never waits longer than a bound LIVENESS
P9 no more than MAX_BURST bytes without a turnaround
P10 the delivered byte sequence does not depend on the rx_ready
pattern RELATIONAL PHASE 1 STATE SWEEP 6 states x 16 input combinations, exhaustively,
each state CONSTRUCTED and then PROVED
PHASE 2 BOUNDARY 8 named scenarios, one per edge worth naming
PHASE 3 STREAM 64 bytes each way, simultaneously, with
backpressure on both sides
PHASE 4 INDEPENDENCE the same stream under 8 ready patterns
PHASE 5 RANDOM supplementary, and audited for what it reaches// =====================================================================
// tb_ft245_sync_if -- Verilog-2005 testbench for ft245_sync_if.
//
// The checker here is a different shape from the one in 29.5, and the
// difference is deliberate. 29.5's design is a small pile of state, so
// a cycle-accurate model of the state is the strongest thing to compare
// against. This design is a bus scheduler: what matters is not what its
// counter holds on cycle 400, it is that no byte is lost, no byte is
// reordered, the bus is never driven from both ends, and neither
// direction can be starved. So the checker is a set of PROPERTIES plus
// two byte streams, and only the output equations are compared directly.
//
// P1 rd_n low implies rxf_n low -- never over-read
// P2 wr_n low implies txe_n low and tx_valid
// P3 bus_drive implies oe_n was high last cycle -- the turnaround
// P4 rd_n low implies oe_n was low last cycle -- OE leads
// P5 the skid never overflows -- n_lost stays 0
// P6 every byte the bridge gave is delivered once, in order
// P7 every byte the user offered is written once, in order
// P8 a transmit never waits longer than a bound -- LIVENESS
// P9 no more than MAX_BURST bytes without a turnaround
// P10 the delivered byte sequence does not depend on the rx_ready
// pattern -- RELATIONAL
//
// PHASES
// 1 STATE SWEEP 6 states x 16 input combinations, exhaustively
// 2 BOUNDARY the named edges, one scenario each
// 3 STREAM a known byte stream each way, with backpressure
// 4 INDEPENDENCE the same stream under 8 different ready patterns
// 5 RANDOM supplementary, and audited for what it reaches
// =====================================================================
`timescale 1ns/1ps
module tb_ft245_sync_if;
localparam integer MAX_BURST = 4;
localparam [2:0] S_IDLE = 3'd0, S_RDOE = 3'd1, S_RD = 3'd2,
S_TAIL = 3'd3, S_TURN = 3'd4, S_WR = 3'd5;
reg clk = 1'b0;
reg rst_n;
reg rxf_n, txe_n;
reg [7:0] data_in;
reg rx_ready, tx_valid;
reg [7:0] tx_data;
wire rd_n, wr_n, oe_n, bus_drive;
wire [7:0] data_out;
wire rx_valid, tx_ready;
wire [7:0] rx_data;
wire [2:0] state_o;
wire [15:0] n_rx, n_tx, n_turn, n_lost, n_contend;
ft245_sync_if #(.MAX_BURST(MAX_BURST)) dut (
.clk(clk), .rst_n(rst_n),
.rxf_n(rxf_n), .txe_n(txe_n), .data_in(data_in),
.rd_n(rd_n), .wr_n(wr_n), .oe_n(oe_n),
.bus_drive(bus_drive), .data_out(data_out),
.rx_valid(rx_valid), .rx_data(rx_data), .rx_ready(rx_ready),
.tx_valid(tx_valid), .tx_data(tx_data), .tx_ready(tx_ready),
.state_o(state_o),
.n_rx(n_rx), .n_tx(n_tx), .n_turn(n_turn),
.n_lost(n_lost), .n_contend(n_contend)
);
always #5 clk = ~clk;
// ---- bookkeeping -------------------------------------------------
integer chk_dir, chk_rnd, err, in_random;
integer oe_n_prev, rd_n_prev;
integer burst_run; // consecutive strobes, checked
integer m_rd, m_wr, m_turn, m_skid2, m_burstmax, m_bothwant, m_altern,
m_setupfail, m_txwait_max, m_rxstall;
integer i, j;
task bump; begin
if (in_random) chk_rnd = chk_rnd + 1; else chk_dir = chk_dir + 1;
end endtask
task ck;
input [255:0] what;
input [31:0] got;
input [31:0] exp;
begin
bump;
if (got !== exp) begin
err = err + 1;
if (!in_random && err <= 40)
$display(" ** %0s: got %0d expected %0d (t=%0t, state=%0d)",
what, got, exp, $time, state_o);
end
end
endtask
// ---- the bridge model (a VERIFICATION MODEL, not hardware) -------
// Holds the bytes it is going to hand over and the bytes it has been
// handed. Presents a byte for as long as it is unread, which is the
// property the design's read-safety argument rests on.
reg [7:0] br_rx [0:1023];
integer br_rx_wr, br_rx_rd;
reg [7:0] br_tx [0:4095];
integer br_tx_n;
integer br_room; // how many more bytes it can take
integer br_room_pat; // 0 = always room, else a pattern
task bridge_present; begin
rxf_n = (br_rx_rd >= br_rx_wr);
data_in = (br_rx_rd < br_rx_wr) ? br_rx[br_rx_rd] : 8'hxx;
txe_n = (br_room == 0);
end endtask
// Called at every clock edge: the bridge reacts to the strobes.
task bridge_step; begin
if (!rd_n) begin
if (br_rx_rd < br_rx_wr) br_rx_rd = br_rx_rd + 1;
else begin
err = err + 1;
$display(" ** BRIDGE: read strobe with nothing to give (t=%0t)", $time);
end
end
if (!wr_n) begin
if (br_room > 0 && br_tx_n < 4096) begin
br_tx[br_tx_n] = data_out; br_tx_n = br_tx_n + 1;
br_room = br_room - 1;
end else if (br_room > 0) begin
err = err + 1;
$display(" ** BRIDGE: transmit log full at %0d (t=%0t)", br_tx_n, $time);
end else begin
err = err + 1;
$display(" ** BRIDGE: write strobe with no room (t=%0t)", $time);
end
end
end endtask
// ---- the expected-value side of the stream checks -----------------
reg [7:0] exp_rx [0:1023]; // what the user must receive
integer exp_rx_n, got_rx_n;
reg [7:0] src_tx [0:1023]; // what the user offered
integer src_tx_n, src_tx_sent;
integer tx_wait; // cycles this byte has waited
integer tx_wait_bound;
// ---- the per-cycle property checks -------------------------------
// Only the output EQUATIONS are compared directly. Everything else is
// a property that does not know how the design is built.
task props;
reg exp_rd, exp_wr, exp_oe, exp_bd, exp_tr;
begin
exp_rd = ~((state_o == S_RD) && ~rxf_n);
exp_wr = ~((state_o == S_WR) && ~txe_n && tx_valid);
exp_oe = ~((state_o == S_RDOE) || (state_o == S_RD) ||
(state_o == S_TAIL));
exp_bd = (state_o == S_WR);
exp_tr = (state_o == S_WR) && ~txe_n && tx_valid;
ck("rd_n", {31'd0, rd_n}, {31'd0, exp_rd});
ck("wr_n", {31'd0, wr_n}, {31'd0, exp_wr});
ck("oe_n", {31'd0, oe_n}, {31'd0, exp_oe});
ck("bus_drive", {31'd0, bus_drive}, {31'd0, exp_bd});
ck("tx_ready", {31'd0, tx_ready}, {31'd0, exp_tr});
// P1 never read an empty bridge
bump; if (!rd_n && rxf_n) begin
err = err + 1;
if (!in_random) $display(" ** P1 over-read (t=%0t)", $time);
end
// P2 never write to a full bridge, and never without a byte
bump; if (!wr_n && (txe_n || !tx_valid)) begin
err = err + 1;
if (!in_random) $display(" ** P2 bad write (t=%0t)", $time);
end
// P3 the turnaround: we must not drive within one cycle of the
// bridge's output enable
bump; if (bus_drive && oe_n_prev === 1'b0) begin
err = err + 1;
if (!in_random) $display(" ** P3 no turnaround (t=%0t)", $time);
end
// P4 output enable leads the strobe
bump; if (!rd_n && oe_n_prev !== 1'b0) begin
err = err + 1;
if (!in_random) $display(" ** P4 strobe without OE lead (t=%0t)", $time);
end
// P5 the skid never overflows, and the bus is never contended
ck("P5 n_lost", {16'd0, n_lost}, 32'd0);
ck("P3 n_contend", {16'd0, n_contend}, 32'd0);
// P9 the burst bound
bump; if (burst_run > MAX_BURST) begin
err = err + 1;
if (!in_random) $display(" ** P9 burst %0d > %0d (t=%0t)",
burst_run, MAX_BURST, $time);
end
end
endtask
// Advance every model that tracks the design. Called right at the edge.
task edge_models;
begin
// P6: a byte leaving the skid must be the next expected one
if (rx_valid && rx_ready) begin
bump;
if (got_rx_n >= exp_rx_n) begin
err = err + 1;
if (err <= 40)
$display(" ** P6 extra byte delivered (t=%0t)", $time);
end else if (rx_data !== exp_rx[got_rx_n]) begin
err = err + 1;
if (err <= 40)
$display(" ** P6 byte %0d: got %02h expected %02h (t=%0t)",
got_rx_n, rx_data, exp_rx[got_rx_n], $time);
got_rx_n = got_rx_n + 1;
end else got_rx_n = got_rx_n + 1;
end
// P7 is checked at the end of a stream by comparing br_tx to src_tx
if (tx_ready) src_tx_sent = src_tx_sent + 1;
// P8: how long has a transmit been waiting FOR THE DESIGN. Cycles
// where the bridge has no room are the environment refusing, not
// the arbiter starving, and counting them measures the bridge.
if (tx_valid && !txe_n && !tx_ready) tx_wait = tx_wait + 1;
else tx_wait = 0;
if (tx_wait > m_txwait_max) m_txwait_max = tx_wait;
// P9's run counter: consecutive strobes without a turnaround
if (state_o == S_TURN) burst_run = 0;
else if (!rd_n || !wr_n) burst_run = burst_run + 1;
// reachability tallies
if (!rd_n) m_rd = m_rd + 1;
if (!wr_n) m_wr = m_wr + 1;
if (state_o == S_TURN) m_turn = m_turn + 1;
if (!rxf_n && !txe_n && tx_valid) m_bothwant = m_bothwant + 1;
if (rx_valid && !rx_ready) m_rxstall = m_rxstall + 1;
if (burst_run == MAX_BURST) m_burstmax = m_burstmax + 1;
oe_n_prev = oe_n;
rd_n_prev = rd_n;
bridge_step;
end
endtask
task step; begin
#1;
props;
@(posedge clk);
edge_models;
#1;
bridge_present;
end endtask
task hard_reset; begin
rst_n = 0; rxf_n = 1; txe_n = 1; data_in = 8'h00;
rx_ready = 0; tx_valid = 0; tx_data = 8'h00;
br_rx_wr = 0; br_rx_rd = 0; br_tx_n = 0; br_room = 0;
exp_rx_n = 0; got_rx_n = 0; src_tx_n = 0; src_tx_sent = 0;
tx_wait = 0; burst_run = 0;
oe_n_prev = 1; rd_n_prev = 1;
repeat (3) @(posedge clk);
#1; rst_n = 1;
@(posedge clk); #1;
bridge_present;
end endtask
// -----------------------------------------------------------------
// Build one of the six states and PROVE it was built.
// -----------------------------------------------------------------
task setup_state;
input [2:0] want;
integer guard;
begin
hard_reset;
case (want)
S_IDLE: ; // reset leaves us here
S_RDOE: begin
br_rx[br_rx_wr] = 8'hA5; br_rx_wr = br_rx_wr + 1;
exp_rx[exp_rx_n] = 8'hA5; exp_rx_n = exp_rx_n + 1;
bridge_present; rx_ready = 1;
guard = 0;
while (state_o !== S_RDOE && guard < 20) begin step; guard = guard + 1; end
end
S_RD: begin
br_rx[br_rx_wr] = 8'hA5; br_rx_wr = br_rx_wr + 1;
exp_rx[exp_rx_n] = 8'hA5; exp_rx_n = exp_rx_n + 1;
bridge_present; rx_ready = 1;
guard = 0;
while (state_o !== S_RD && guard < 20) begin step; guard = guard + 1; end
end
S_TAIL: begin
br_rx[br_rx_wr] = 8'hA5; br_rx_wr = br_rx_wr + 1;
exp_rx[exp_rx_n] = 8'hA5; exp_rx_n = exp_rx_n + 1;
bridge_present; rx_ready = 1;
guard = 0;
while (state_o !== S_TAIL && guard < 20) begin step; guard = guard + 1; end
end
S_TURN: begin
br_rx[br_rx_wr] = 8'hA5; br_rx_wr = br_rx_wr + 1;
exp_rx[exp_rx_n] = 8'hA5; exp_rx_n = exp_rx_n + 1;
bridge_present; rx_ready = 1;
guard = 0;
while (state_o !== S_TURN && guard < 20) begin step; guard = guard + 1; end
end
S_WR: begin
br_room = 8; bridge_present;
tx_valid = 1; tx_data = 8'h5A;
src_tx[src_tx_n] = 8'h5A; src_tx_n = src_tx_n + 1;
guard = 0;
while (state_o !== S_WR && guard < 20) begin step; guard = guard + 1; end
end
default: ;
endcase
bump;
if (state_o !== want) begin
err = err + 1; m_setupfail = m_setupfail + 1;
$display(" ** setup: wanted state %0d, reached %0d", want, state_o);
end
end
endtask
// -----------------------------------------------------------------
// PHASE 1: every state against every input combination.
// -----------------------------------------------------------------
integer st_i, in_i;
task phase_sweep;
begin
for (st_i = 0; st_i <= 5; st_i = st_i + 1)
for (in_i = 0; in_i <= 15; in_i = in_i + 1) begin
setup_state(st_i[2:0]);
// in_i bits: [0] a byte is available [1] the bridge has room
// [2] the user can take [3] the user has a byte
if (in_i[0] && br_rx_rd >= br_rx_wr) begin
br_rx[br_rx_wr] = 8'h30 + in_i[3:0];
exp_rx[exp_rx_n] = 8'h30 + in_i[3:0];
br_rx_wr = br_rx_wr + 1; exp_rx_n = exp_rx_n + 1;
end else if (!in_i[0]) begin
br_rx_rd = br_rx_wr; // nothing on offer
end
br_room = in_i[1] ? 8 : 0;
rx_ready = in_i[2];
tx_valid = in_i[3];
if (in_i[3]) begin
tx_data = 8'h80 + in_i[3:0];
src_tx[src_tx_n] = tx_data; src_tx_n = src_tx_n + 1;
end
bridge_present;
step; // the transition itself
step; // and one cycle later
end
end
endtask
// -----------------------------------------------------------------
// PHASE 2: the named boundaries.
// -----------------------------------------------------------------
task load_rx;
input integer n;
integer k;
begin
for (k = 0; k < n; k = k + 1) begin
br_rx[br_rx_wr] = 8'h40 + k[7:0];
exp_rx[exp_rx_n] = 8'h40 + k[7:0];
br_rx_wr = br_rx_wr + 1; exp_rx_n = exp_rx_n + 1;
end
bridge_present;
end
endtask
task phase_boundary;
integer k, seen_turn, first_burst;
begin
// B1 a long run of available bytes must be broken by a turnaround
// at MAX_BURST. Without the bound the bus is never released.
hard_reset; load_rx(40); rx_ready = 1; br_room = 0;
first_burst = 0; seen_turn = 0;
for (k = 0; k < 60; k = k + 1) begin
step;
if (state_o == S_TURN) seen_turn = seen_turn + 1;
if (burst_run > first_burst) first_burst = burst_run;
end
ck("B1 burst never exceeds MAX", first_burst, MAX_BURST);
bump; if (seen_turn < 4) begin
err = err + 1;
$display(" ** B1 only %0d turnarounds in 60 cycles", seen_turn);
end
// B2 the user stalls mid-burst: no byte may be lost, and the skid
// must be what absorbs the one already committed.
hard_reset; load_rx(12); rx_ready = 1; br_room = 0;
repeat (4) step;
rx_ready = 0;
repeat (8) step; // the committed byte lands in the skid
ck("B2 nothing lost", {16'd0, n_lost}, 32'd0);
rx_ready = 1;
repeat (40) step;
ck("B2 every byte delivered", got_rx_n, exp_rx_n);
// B3 the LAST byte. rxf_n deasserts only after it is read, so a
// machine that gives up when the count reaches one still gets it.
hard_reset; load_rx(1); rx_ready = 1; br_room = 0;
repeat (12) step;
ck("B3 the last byte arrived", got_rx_n, 32'd1);
ck("B3 rxf_n now high", {31'd0, rxf_n}, 32'd1);
// B4 transmit only: every offered byte reaches the bridge, in order
hard_reset; br_room = 16; bridge_present;
tx_valid = 1;
for (k = 0; k < 8; k = k + 1) begin
tx_data = 8'hC0 + k[7:0];
src_tx[src_tx_n] = tx_data; src_tx_n = src_tx_n + 1;
while (!tx_ready) step;
step;
end
tx_valid = 0;
repeat (4) step;
ck("B4 eight bytes written", br_tx_n, 32'd8);
for (k = 0; k < 8; k = k + 1)
ck("B4 byte order", {24'd0, br_tx[k]}, {24'd0, 8'hC0 + k[7:0]});
// B5 THE LIVENESS CASE. The bridge always has a byte to give and
// the user always has a byte to send. A design that lets read
// win unconditionally transmits NOTHING, forever, and every
// check above still passes.
hard_reset; load_rx(200); rx_ready = 1;
br_room = 64; bridge_present;
tx_valid = 1; tx_data = 8'hEE;
src_tx[src_tx_n] = 8'hEE; src_tx_n = src_tx_n + 1;
for (k = 0; k < 300; k = k + 1) step;
bump; if (br_tx_n == 0) begin
err = err + 1;
$display(" ** B5 STARVED: 300 cycles of contention, 0 bytes sent");
end
ck("B5 reads happened too", (n_rx > 0) ? 1 : 0, 32'd1);
tx_wait_bound = 24;
bump; if (m_txwait_max > tx_wait_bound) begin
err = err + 1;
$display(" ** B5 a transmit waited %0d cycles (bound %0d)",
m_txwait_max, tx_wait_bound);
end
// B6 the turnaround is not optional: after a read burst the bus
// must be idle for a cycle before we drive it.
hard_reset; load_rx(6); rx_ready = 1;
br_room = 8; bridge_present;
tx_valid = 1; tx_data = 8'h77;
src_tx[src_tx_n] = 8'h77; src_tx_n = src_tx_n + 1;
repeat (40) step;
ck("B6 no contention", {16'd0, n_contend}, 32'd0);
bump; if (n_turn == 0) begin
err = err + 1; $display(" ** B6 no turnaround cycles at all");
end
// B8 the user runs out of bytes in the middle of a write burst,
// while the bridge still has room. The machine must let go of
// the bus: a receive behind it has to make progress. This is
// the mirror of B5 and it is the ONLY directed scenario that
// constructs it.
hard_reset; br_room = 16; bridge_present;
tx_valid = 1; tx_data = 8'h11;
src_tx[src_tx_n] = 8'h11; src_tx_n = src_tx_n + 1;
k = 0;
while (state_o !== S_WR && k < 20) begin step; k = k + 1; end
ck("B8 reached the write state", {29'd0, state_o}, {29'd0, S_WR});
tx_valid = 0;
load_rx(8); rx_ready = 1;
repeat (60) step;
ck("B8 the receive behind it completed", got_rx_n, exp_rx_n);
// B7 reset in the middle of a read burst returns everything, and
// the skid does not deliver a stale byte afterwards.
hard_reset; load_rx(10); rx_ready = 1; br_room = 0;
repeat (5) step;
// Drop rx_ready and empty the bridge before releasing reset, so
// that the state observed is the one reset produced and not the
// first grant after it. Setting rxf_n by hand is not enough:
// bridge_present drives it from the model at the end of every step.
rx_ready = 0; br_rx_rd = br_rx_wr; bridge_present;
rst_n = 0; step; step; rst_n = 1; step;
ck("B7 state is idle", {29'd0, state_o}, {29'd0, S_IDLE});
ck("B7 skid empty", {31'd0, rx_valid}, 32'd0);
ck("B7 counters clear",{16'd0, n_rx}, 32'd0);
// the bytes the bridge already gave up are gone; resynchronise
exp_rx_n = 0; got_rx_n = 0;
br_rx_rd = br_rx_wr;
end
endtask
// -----------------------------------------------------------------
// PHASE 3: a known stream each way, under backpressure.
// -----------------------------------------------------------------
integer rdy_pat [0:7];
integer room_pat [0:7];
task phase_stream;
integer k, cyc;
begin
rdy_pat[0]=1; rdy_pat[1]=1; rdy_pat[2]=0; rdy_pat[3]=1;
rdy_pat[4]=0; rdy_pat[5]=0; rdy_pat[6]=1; rdy_pat[7]=1;
room_pat[0]=1; room_pat[1]=0; room_pat[2]=1; room_pat[3]=1;
room_pat[4]=0; room_pat[5]=1; room_pat[6]=0; room_pat[7]=1;
hard_reset;
for (k = 0; k < 64; k = k + 1) begin
br_rx[br_rx_wr] = 8'h10 + k[7:0];
exp_rx[exp_rx_n] = 8'h10 + k[7:0];
br_rx_wr = br_rx_wr + 1; exp_rx_n = exp_rx_n + 1;
end
for (k = 0; k < 64; k = k + 1) src_tx[k] = 8'hA0 + k[7:0];
src_tx_n = 64; src_tx_sent = 0;
br_room = 0; bridge_present;
tx_valid = 1; tx_data = src_tx[0];
for (cyc = 0; cyc < 1400; cyc = cyc + 1) begin
rx_ready = rdy_pat[cyc % 8];
if (room_pat[cyc % 8] && br_room < 8) br_room = br_room + 1;
if (src_tx_sent < 64) tx_data = src_tx[src_tx_sent];
tx_valid = (src_tx_sent < 64);
bridge_present;
step;
end
ck("P6 all 64 bytes received", got_rx_n, 32'd64);
ck("P7 all 64 bytes written", br_tx_n, 32'd64);
for (k = 0; k < 64; k = k + 1)
ck("P7 order", {24'd0, br_tx[k]}, {24'd0, 8'hA0 + k[7:0]});
// the two directions really did overlap
bump; if (m_bothwant < 100) begin
err = err + 1;
$display(" ** stream: the two directions contended only %0d times",
m_bothwant);
end
end
endtask
// -----------------------------------------------------------------
// PHASE 4: the relational property. The SAME bridge stimulus under
// eight different rx_ready patterns must deliver the same byte
// sequence. Backpressure is allowed to change WHEN; it is not allowed
// to change WHAT.
// -----------------------------------------------------------------
reg [7:0] seq_ref [0:127];
integer seq_ref_n;
integer pat_runs;
task phase_independence;
integer p, k, cyc, mask;
begin
pat_runs = 0;
for (p = 0; p < 8; p = p + 1) begin
hard_reset;
for (k = 0; k < 48; k = k + 1) begin
br_rx[br_rx_wr] = 8'h60 + k[7:0];
exp_rx[exp_rx_n] = 8'h60 + k[7:0];
br_rx_wr = br_rx_wr + 1; exp_rx_n = exp_rx_n + 1;
end
br_room = 0; bridge_present; // receive only
mask = 0;
for (cyc = 0; cyc < 900; cyc = cyc + 1) begin
// eight genuinely different duty cycles, from "always ready"
// to "ready one cycle in eight"
rx_ready = (p == 0) ? 1'b1 : ((cyc % (p + 1)) == 0);
if (rx_valid && rx_ready && mask < 48) begin
seq_ref[mask] = rx_data; // recorded, then
mask = mask + 1; // compared below
end
step;
end
ck("P10 byte count", mask, 32'd48);
for (k = 0; k < 48; k = k + 1)
ck("P10 sequence", {24'd0, seq_ref[k]}, {24'd0, 8'h60 + k[7:0]});
pat_runs = pat_runs + 1;
if (p > 0) m_rxstall = m_rxstall; // measured in edge_models
end
end
endtask
// -----------------------------------------------------------------
// PHASE 5: random, audited.
// -----------------------------------------------------------------
task phase_random;
integer cyc, k;
begin
in_random = 1;
hard_reset;
for (k = 0; k < 900; k = k + 1) begin
br_rx[br_rx_wr] = ({$random} % 256);
exp_rx[exp_rx_n] = br_rx[br_rx_wr]; // the SAME bytes, expected
br_rx_wr = br_rx_wr + 1; exp_rx_n = exp_rx_n + 1;
end
src_tx_n = 0;
br_room = 0; bridge_present;
for (cyc = 0; cyc < 6000; cyc = cyc + 1) begin
rx_ready = (({$random} % 100) < 62);
if ((({$random} % 100) < 45) && br_room < 8) br_room = br_room + 1;
tx_valid = (({$random} % 100) < 58);
if (tx_valid) tx_data = ({$random} % 256);
bridge_present;
step;
end
in_random = 0;
end
endtask
initial begin
chk_dir = 0; chk_rnd = 0; err = 0; in_random = 0;
m_rd=0; m_wr=0; m_turn=0; m_skid2=0; m_burstmax=0; m_bothwant=0;
m_altern=0; m_setupfail=0; m_txwait_max=0; m_rxstall=0;
phase_sweep;
$display(" phase 1 state sweep : %0d checks, %0d errors (96 transitions)",
chk_dir, err);
phase_boundary;
$display(" phase 2 boundary : %0d checks, %0d errors", chk_dir, err);
phase_stream;
$display(" phase 3 stream : %0d checks, %0d errors", chk_dir, err);
phase_independence;
$display(" phase 4 independence : %0d checks, %0d errors (%0d ready patterns)",
chk_dir, err, pat_runs);
$display(" ---- DIRECTED-ONLY TOTAL : %0d checks, %0d errors ----", chk_dir, err);
phase_random;
$display("");
$display(" measured reachability (all phases)");
$display(" read strobes ........... %0d", m_rd);
$display(" write strobes .......... %0d", m_wr);
$display(" turnaround cycles ...... %0d", m_turn);
$display(" both wanted the bus .... %0d", m_bothwant);
$display(" bursts reaching MAX .... %0d", m_burstmax);
$display(" user stalled a byte .... %0d", m_rxstall);
$display(" longest transmit wait .. %0d cycles", m_txwait_max);
$display(" setup failures ......... %0d", m_setupfail);
$display("");
$display(" directed checks .......... %0d", chk_dir);
$display(" random checks ............ %0d", chk_rnd);
$display(" TOTAL checks ............. %0d", chk_dir + chk_rnd);
$display(" ERRORS ................... %0d", err);
if (err == 0) $display(" PASS"); else $display(" FAIL");
$finish;
end
endmodule8. The Measurement
VERILOG SYSTEMVERILOG VHDL-2008
phase 1 sweep 4,932 4,932 4,932
phase 2 boundary 11,692 11,692 11,692
phase 3 stream 28,623 28,623 28,623
phase 4 independence 115,799 115,799 115,799
---- DIRECTED 115,799 115,799 115,799
errors 0 0 0
phase 5 random 72,900 72,900 72,900
TOTAL 188,699 188,699 188,699
errors 0 0 0Every column is identical, including the random one — which in 29.3, 29.4 and 29.5 it was not. The reason is worth a sentence, because a coincidence that looks like a result is worse than no result: the random phase here drives a fixed number of cycles and the checks run once per cycle, so the check count does not depend on the random values. The check contents do, and the three generators produce three different stimuli — visible in the reachability table below, where the three languages report different strobe counts.
The exhaustive part, and where its denominator comes from
STATES IDLE, RDOE, RD, TAIL, TURN, WR 6
INPUTS rxf_n, txe_n, rx_ready, tx_valid 2^4 = 16
---------------------------------------------------------------------
6 x 16 = 96Ninety-six, and all ninety-six are meaningful: unlike 29.5, no combination here is unreachable. Every state can be entered, and every one of the sixteen input vectors can be presented in every state, because the four inputs are driven by two independent agents — the bridge and the user — neither of which is constrained by the machine's state.
Each of the 96 is entered from a state that was built and then verified, and
is compared on the transition itself and one cycle later. The burst counter is
deliberately not an axis of the sweep: it is a bounded counter whose only
interesting values are zero and MAX_BURST, and both are covered by boundary
scenario B1, which sweeps a 60-cycle read against a bridge that never runs out.
The relational part
Phase 4 runs the same 48-byte bridge stimulus eight times, under eight different
rx_ready duty cycles — from "ready every cycle" to "ready one cycle in eight" —
and requires the delivered byte sequence to be identical every time.
ALLOWED TO CHANGE when each byte arrives, how many turnarounds
happen, how long the run takes
NOT ALLOWED TO CHANGE which bytes arrive, and in what orderThat is a statement about eight different executions, so it is not an assertion; it is a testbench structure, and it is the third shape of relational property in this module. 29.4 compared a run against its mirror image; 29.5 compared a run against the same run with the two buffers relabelled; this one compares a run against itself under different timing. All three are exhaustively checkable because the transformation is finite and total, and none of the three can be written in SVA.
It is also the property the skid exists to provide, which makes it the right headline check: backpressure is allowed to change when, never what.
What the phases actually reached
read strobes ............... 1,558
write strobes .............. 1,144
turnaround cycles .......... 1,560
cycles where BOTH directions
wanted the bus ........... 2,141
bursts reaching MAX_BURST .. 478
cycles the user stalled a
byte it had been offered .. 1,940
longest transmit wait ...... 10 cycles
setups that failed to build
the requested state ....... 0Two of those rows are the ones that make the suite mean anything.
9. SystemVerilog
The same contract. One type does the work here, and it does it in a way that is specific to state machines.
typedef enum logic [2:0] {
S_IDLE = 3'd0,
S_RDOE = 3'd1, // output enable leads the strobe
S_RD = 3'd2, // rd_n gated by rxf_n
S_TAIL = 3'd3, // the last committed byte settles
S_TURN = 3'd4, // the dead cycle between directions
S_WR = 3'd5
} state_e;Three bits hold eight values and the machine uses six. In the Verilog version the
other two are reachable in principle — a state that is a plain reg [2:0] can
hold 3'd6, and the default branch in the case statement is there because of
it. With an enum there is nothing to recover from: an assignment of a value the
type does not define is a compile error, and the two spare encodings are not
states the machine can be in.
That is a smaller claim than it sounds, and worth being precise about. It does not protect against an upset that flips a state bit in silicon — a type is a compile-time thing and a single-event upset is not. What it removes is the class of bug where a designer writes a transition to a state that does not exist, or renumbers the encoding and misses one place.
// =====================================================================
// ft245_sync_if (SystemVerilog) -- the same hardware contract as the
// Verilog-2005 module: same ports, same states, same strobe equations,
// same burst bound, same skid depth, same latency.
//
// What the types add here is that the STATE is an enum rather than six
// localparams, so a transition to a value the machine does not define
// is a type error rather than a silent default; and the strobes are
// written once, in always_comb, where a missing assignment infers a
// latch and the tool says so.
//
// Everything below the line is the Verilog module's documentation and
// applies unchanged.
// ------------------------------------------------------------------
// The FPGA side of a USB-to-FIFO bridge in synchronous mode. This is
// what "USB" actually is on most FPGA development boards:
// not a USB controller in the fabric, but a bus master for a bridge
// chip that speaks USB on the far side and a byte FIFO on this one.
//
// CLASSIFICATION: simplified synthesisable teaching RTL.
// It is NOT an FT2232H / FT600 driver and it is not a USB controller.
// There is no USB here at all -- no PHY, no packets, no endpoints, no
// enumeration. Everything on the other side of the bridge has already
// happened. What is left is the part that runs in the FPGA, and the
// three decisions it has to get right:
//
// 1. THE BUS IS BIDIRECTIONAL, so changing direction costs a dead
// cycle in which no data moves and nobody drives.
// 2. THE STROBES GO OFF-CHIP, so they are registered. The decision to
// read is therefore made a cycle before the byte arrives, which is
// why there is a skid buffer and why it is two deep.
// 3. READ AND WRITE COMPETE FOR ONE BUS, so a priority that always
// favours one direction starves the other FOREVER. That is a
// liveness bug, and no amount of checking that the data is correct
// will find it.
//
// THE BRIDGE MODEL THIS ASSUMES (stated, because RTL cannot see a chip)
// --------------------------------------------------------------------
// rxf_n low the bridge is presenting a byte on data_in, and will go
// on presenting it until it is read. It does NOT withdraw
// a byte, so a read taken while rxf_n is low is always
// valid.
// rd_n low consumes the presented byte. The next cycle presents the
// following byte, or deasserts rxf_n.
// txe_n low the bridge has room for a byte.
// wr_n low hands it one, from data_out.
// oe_n must be asserted at least one cycle BEFORE rd_n, so the
// bridge's drivers are on before the strobe.
// =====================================================================
module ft245_sync_if_sv #(
// The longest run of bytes in one direction before the machine must
// hand the bus to the other. This is the bound that turns "read has
// priority" into "read has priority for a while".
parameter int MAX_BURST = 4
) (
input logic clk, // sourced by the BRIDGE, not by the FPGA
input logic rst_n,
// ---- bridge side -----------------------------------------------
input logic rxf_n, // active low: a byte is presented
input logic txe_n, // active low: there is room for a byte
input logic [7:0] data_in,
output logic rd_n,
output logic wr_n,
output logic oe_n,
output logic bus_drive, // we are driving data_out onto the bus
output logic [7:0] data_out,
// ---- user side, receive ----------------------------------------
output logic rx_valid,
output logic [7:0] rx_data,
input logic rx_ready,
// ---- user side, transmit ---------------------------------------
input logic tx_valid,
input logic [7:0] tx_data,
output logic tx_ready,
// ---- observation ------------------------------------------------
output logic [2:0] state_o,
output logic [15:0] n_rx,
output logic [15:0] n_tx,
output logic [15:0] n_turn,
// A byte arrived with nowhere to put it. Should be impossible; the
// skid depth is chosen so that it is. A non-zero value here means the
// depth argument in section 4 is wrong.
output logic [15:0] n_lost,
// We drove the bus within one cycle of the bridge's output enable. The
// electrical consequence is not simulable; the SCHEDULING error is.
output logic [15:0] n_contend
);
// The six states the bus can be in, as a type. A transition to
// anything else is now a compile error rather than a default branch.
typedef enum logic [2:0] {
S_IDLE = 3'd0,
S_RDOE = 3'd1, // output enable leads the strobe
S_RD = 3'd2, // rd_n gated by rxf_n
S_TAIL = 3'd3, // the last committed byte settles
S_TURN = 3'd4, // the dead cycle between directions
S_WR = 3'd5
} state_e;
state_e state;
logic [15:0] burst;
logic last_was_rd; // for the alternation in S_IDLE
logic [15:0] c_rx, c_tx, c_turn, c_lost, c_cont;
logic oe_n_d;
// ---- the skid buffer -------------------------------------------
// TWO deep, and the depth is the whole argument. rd_n is registered
// because it leaves the chip, so the decision to take a byte is made
// one cycle before the byte exists. One entry absorbs that byte; the
// second is what lets the machine keep reading on consecutive cycles
// while the user is draining, instead of every other cycle.
logic [7:0] sk0, sk1;
logic wptr, rptr;
logic [1:0] cnt;
// ---- combinational strobes -------------------------------------
// rd_n is gated by rxf_n and by NOTHING ELSE from outside this module.
// One gate from an input pin to an output pin closes timing on a
// bridge-clocked interface; a path from the user's rx_ready would not,
// which is exactly why the skid exists instead.
assign rd_n = ~((state == S_RD) && ~rxf_n);
assign wr_n = ~((state == S_WR) && ~txe_n && tx_valid);
assign oe_n = ~((state == S_RDOE) || (state == S_RD) || (state == S_TAIL));
assign bus_drive = (state == S_WR);
assign data_out = tx_data;
assign rx_valid = (cnt != 2'd0);
assign rx_data = rptr ? sk1 : sk0;
assign tx_ready = (state == S_WR) && ~txe_n && tx_valid;
logic take, drain;
assign take = ~rd_n; // a byte is consumed this cycle
assign drain = rx_valid & rx_ready; // a byte leaves the skid
// What the occupancy will be after this edge. The read decision uses
// this and not rx_ready, because rx_ready next cycle is not knowable.
logic [2:0] cnt_after;
assign cnt_after = {1'b0, cnt} + (take ? 3'd1 : 3'd0)
- (drain ? 3'd1 : 3'd0);
// Room for the byte a further read would commit to.
logic room_next;
assign room_next = (cnt_after <= 3'd1);
// ---- who gets the bus ------------------------------------------
logic want_rd, want_wr, grant_rd, grant_wr, stay_rd, stay_wr;
assign want_rd = ~rxf_n & room_next;
assign want_wr = ~txe_n & tx_valid;
// Both want it: give it to whichever did not go last. Without this the
// design still passes every check that data is correct.
assign grant_rd = want_rd & (~want_wr | !last_was_rd);
assign grant_wr = want_wr & ~grant_rd;
// Keep going only while there is something to move and somewhere to
// put it. The BUDGET is deliberately not tested here: it is enforced
// one place only, in the exit condition inside each state, which fires
// on the cycle the last permitted byte moves. Testing it here as well
// is dead code -- the exit always fires first.
assign stay_rd = ~rxf_n & room_next;
assign stay_wr = ~txe_n & tx_valid;
always_ff @(posedge clk or negedge rst_n) begin
if (!rst_n) begin
state <= S_IDLE;
burst <= 16'd0;
last_was_rd <= 1'b0;
sk0 <= 8'd0;
sk1 <= 8'd0;
wptr <= 1'b0;
rptr <= 1'b0;
cnt <= 2'd0;
oe_n_d <= 1'b1;
c_rx <= 16'd0;
c_tx <= 16'd0;
c_turn <= 16'd0;
c_lost <= 16'd0;
c_cont <= 16'd0;
end else begin
oe_n_d <= oe_n;
// ---- the skid, independent of the state machine ----
if (take) begin
if (cnt == 2'd2 && !drain) begin
c_lost <= c_lost + 16'd1; // must never happen
end else begin
if (wptr) sk1 <= data_in; else sk0 <= data_in;
wptr <= ~wptr;
c_rx <= c_rx + 16'd1;
end
end
if (drain) rptr <= ~rptr;
case ({take, drain})
2'b10: if (cnt != 2'd2) cnt <= cnt + 2'd1;
2'b01: cnt <= cnt - 2'd1;
default: ; // 00 and 11 both leave it alone
endcase
if (tx_ready) c_tx <= c_tx + 16'd1;
if (bus_drive && !oe_n_d) c_cont <= c_cont + 16'd1;
if (state == S_TURN) c_turn <= c_turn + 16'd1;
// ---- the state machine ----
case (state)
S_IDLE: begin
burst <= 16'd0;
if (grant_rd) state <= S_RDOE;
else if (grant_wr) begin state <= S_WR; last_was_rd <= 1'b0; end
end
// One cycle of output-enable lead, with no strobe. The bridge is
// turning its drivers on; nothing is read here.
S_RDOE: begin
state <= S_RD;
last_was_rd <= 1'b1;
end
S_RD: begin
if (take) burst <= burst + 16'd1;
// Leaving because the budget ran out is DIFFERENT from leaving
// because there is nothing there, and only one of the two is
// about fairness.
if (!stay_rd || (take && (burst + 16'd1 >= MAX_BURST)))
state <= S_TAIL;
end
// The strobe is off but the output enable is still asserted: the
// bridge has not let go of the bus yet.
S_TAIL: state <= S_TURN;
// Nobody drives. This cycle moves no data and it is not optional.
S_TURN: state <= S_IDLE;
S_WR: begin
if (tx_ready) burst <= burst + 16'd1;
if (!stay_wr || (tx_ready && (burst + 16'd1 >= MAX_BURST)))
state <= S_TURN;
end
default: state <= S_IDLE;
endcase
end
end
assign state_o = state;
assign n_rx = c_rx;
assign n_tx = c_tx;
assign n_turn = c_turn;
assign n_lost = c_lost;
assign n_contend = c_cont;
`ifdef SVA_ON
// ---------------------------------------------------------------
// Concurrent assertions. Icarus Verilog 13.0 rejects SVA outright,
// so these compile only where a tool supports them; under Icarus
// each is enforced by the procedural property named beside it.
// ---------------------------------------------------------------
// SAFETY. Never strobe a read at a bridge that is presenting nothing.
// The byte would be whatever was last on the pins.
property p_no_overread;
@(posedge clk) disable iff (!rst_n) !rd_n |-> !rxf_n;
endproperty
a_no_overread: assert property (p_no_overread);
// SAFETY. Never strobe a write at a full bridge, or without a byte.
property p_no_badwrite;
@(posedge clk) disable iff (!rst_n) !wr_n |-> (!txe_n && tx_valid);
endproperty
a_no_badwrite: assert property (p_no_badwrite);
// SAFETY. The turnaround. We must not drive the bus in the cycle
// after the bridge's output enable was asserted: its drivers are
// still on. This is the property whose violation is a current spike
// rather than a wrong bit, and simulation cannot see the consequence.
property p_turnaround;
@(posedge clk) disable iff (!rst_n) bus_drive |-> $past(oe_n);
endproperty
a_turnaround: assert property (p_turnaround);
// SAFETY. The output enable leads the strobe by at least one cycle.
property p_oe_leads;
@(posedge clk) disable iff (!rst_n) !rd_n |-> $past(!oe_n);
endproperty
a_oe_leads: assert property (p_oe_leads);
// SAFETY. The two strobes are mutually exclusive by construction; a
// state encoding mistake is what would break it.
property p_strobes_exclusive;
@(posedge clk) disable iff (!rst_n) !(!rd_n && !wr_n);
endproperty
a_strobes_exclusive: assert property (p_strobes_exclusive);
// BOUNDS. The skid never overflows, which is the claim the depth-two
// argument makes, and the burst never exceeds its budget.
property p_skid_bounded;
@(posedge clk) disable iff (!rst_n) cnt <= 2'd2;
endproperty
a_skid_bounded: assert property (p_skid_bounded);
property p_burst_bounded;
@(posedge clk) disable iff (!rst_n) burst <= MAX_BURST;
endproperty
a_burst_bounded: assert property (p_burst_bounded);
// ORDERING. The read setup state always leads to the read state, so
// the output-enable cycle is never spent for nothing.
property p_rdoe_then_rd;
@(posedge clk) disable iff (!rst_n) (state == S_RDOE) |=> (state == S_RD);
endproperty
a_rdoe_then_rd: assert property (p_rdoe_then_rd);
// PROGRESS, and the reason this module has a burst bound at all.
//
// A transmit that the bridge has room for is accepted within a bounded
// number of cycles. The bound is the longest a read burst can hold the
// bus: one output-enable cycle, MAX_BURST strobes, one tail, one
// turnaround and one idle, plus slack.
//
// This is the ONLY property in the set that fails when the arbiter
// gives reads unconditional priority. Every safety property above
// still passes on that design, and so does every byte-for-byte data
// check, because a starved transmitter corrupts nothing.
property p_tx_not_starved;
@(posedge clk) disable iff (!rst_n)
(tx_valid && !txe_n) |-> ##[1:(MAX_BURST + 8)] tx_ready;
endproperty
a_tx_not_starved: assert property (p_tx_not_starved);
// PROGRESS. A read state with a byte on offer actually strobes,
// rather than sitting in the state doing nothing.
property p_rd_strobes;
@(posedge clk) disable iff (!rst_n) ((state == S_RD) && !rxf_n) |-> !rd_n;
endproperty
a_rd_strobes: assert property (p_rd_strobes);
// COVER, so that none of the above can pass by never happening.
c_read_burst: cover property (@(posedge clk) burst == MAX_BURST);
c_turnaround: cover property (@(posedge clk) state == S_TURN);
c_both_want: cover property (@(posedge clk) want_rd && want_wr);
c_skid_full: cover property (@(posedge clk) cnt == 2'd2);
c_user_stalls: cover property (@(posedge clk) rx_valid && !rx_ready);
c_bridge_full: cover property (@(posedge clk) tx_valid && txe_n);
`endif
endmodule// =====================================================================
// tb_ft245_sync_if_sv -- SystemVerilog testbench for ft245_sync_if_sv.
//
// Phases 1-4 present the SAME directed stimulus, in the same order, as
// the Verilog-2005 bench, so their directed check counts must agree to
// the digit. Phase 5 uses its own generator and is not expected to.
//
// The checker here is a different shape from the one in 29.5, and the
// difference is deliberate. 29.5's design is a small pile of state, so
// a cycle-accurate model of the state is the strongest thing to compare
// against. This design is a bus scheduler: what matters is not what its
// counter holds on cycle 400, it is that no byte is lost, no byte is
// reordered, the bus is never driven from both ends, and neither
// direction can be starved. So the checker is a set of PROPERTIES plus
// two byte streams, and only the output equations are compared directly.
//
// P1 rd_n low implies rxf_n low -- never over-read
// P2 wr_n low implies txe_n low and tx_valid
// P3 bus_drive implies oe_n was high last cycle -- the turnaround
// P4 rd_n low implies oe_n was low last cycle -- OE leads
// P5 the skid never overflows -- n_lost stays 0
// P6 every byte the bridge gave is delivered once, in order
// P7 every byte the user offered is written once, in order
// P8 a transmit never waits longer than a bound -- LIVENESS
// P9 no more than MAX_BURST bytes without a turnaround
// P10 the delivered byte sequence does not depend on the rx_ready
// pattern -- RELATIONAL
//
// PHASES
// 1 STATE SWEEP 6 states x 16 input combinations, exhaustively
// 2 BOUNDARY the named edges, one scenario each
// 3 STREAM a known byte stream each way, with backpressure
// 4 INDEPENDENCE the same stream under 8 different ready patterns
// 5 RANDOM supplementary, and audited for what it reaches
// =====================================================================
`timescale 1ns/1ps
module tb_ft245_sync_if_sv;
localparam int MAX_BURST = 4;
localparam [2:0] S_IDLE = 3'd0, S_RDOE = 3'd1, S_RD = 3'd2,
S_TAIL = 3'd3, S_TURN = 3'd4, S_WR = 3'd5;
logic clk = 1'b0;
logic rst_n;
logic rxf_n, txe_n;
logic [7:0] data_in;
logic rx_ready, tx_valid;
logic [7:0] tx_data;
wire rd_n, wr_n, oe_n, bus_drive;
wire [7:0] data_out;
wire rx_valid, tx_ready;
wire [7:0] rx_data;
wire [2:0] state_o;
wire [15:0] n_rx, n_tx, n_turn, n_lost, n_contend;
ft245_sync_if_sv #(.MAX_BURST(MAX_BURST)) dut (
.clk(clk), .rst_n(rst_n),
.rxf_n(rxf_n), .txe_n(txe_n), .data_in(data_in),
.rd_n(rd_n), .wr_n(wr_n), .oe_n(oe_n),
.bus_drive(bus_drive), .data_out(data_out),
.rx_valid(rx_valid), .rx_data(rx_data), .rx_ready(rx_ready),
.tx_valid(tx_valid), .tx_data(tx_data), .tx_ready(tx_ready),
.state_o(state_o),
.n_rx(n_rx), .n_tx(n_tx), .n_turn(n_turn),
.n_lost(n_lost), .n_contend(n_contend)
);
always #5 clk = ~clk;
// ---- bookkeeping -------------------------------------------------
int chk_dir, chk_rnd, err;
bit in_random;
logic oe_n_prev, rd_n_prev;
int burst_run; // consecutive strobes, checked
int m_rd, m_wr, m_turn, m_skid2, m_burstmax, m_bothwant, m_altern,
m_setupfail, m_txwait_max, m_rxstall;
int i, j;
task bump; begin
if (in_random) chk_rnd = chk_rnd + 1; else chk_dir = chk_dir + 1;
end endtask
task ck(string what, logic [31:0] got, logic [31:0] exp);
begin
bump;
if (got !== exp) begin
err = err + 1;
if (!in_random && err <= 40)
$display(" ** %s: got %0d expected %0d (t=%0t, state=%0d)",
what, got, exp, $time, state_o);
end
end
endtask
// ---- the bridge model (a VERIFICATION MODEL, not hardware) -------
// Holds the bytes it is going to hand over and the bytes it has been
// handed. Presents a byte for as long as it is unread, which is the
// property the design's read-safety argument rests on.
logic [7:0] br_rx [1024];
int br_rx_wr, br_rx_rd;
logic [7:0] br_tx [4096];
int br_tx_n;
int br_room; // how many more bytes it can take
int br_room_pat; // 0 = always room, else a pattern
task bridge_present; begin
rxf_n = (br_rx_rd >= br_rx_wr);
data_in = (br_rx_rd < br_rx_wr) ? br_rx[br_rx_rd] : 8'hxx;
txe_n = (br_room == 0);
end endtask
// Called at every clock edge: the bridge reacts to the strobes.
task bridge_step; begin
if (!rd_n) begin
if (br_rx_rd < br_rx_wr) br_rx_rd = br_rx_rd + 1;
else begin
err = err + 1;
$display(" ** BRIDGE: read strobe with nothing to give (t=%0t)", $time);
end
end
if (!wr_n) begin
if (br_room > 0 && br_tx_n < 4096) begin
br_tx[br_tx_n] = data_out; br_tx_n = br_tx_n + 1;
br_room = br_room - 1;
end else if (br_room > 0) begin
err = err + 1;
$display(" ** BRIDGE: transmit log full at %0d (t=%0t)", br_tx_n, $time);
end else begin
err = err + 1;
$display(" ** BRIDGE: write strobe with no room (t=%0t)", $time);
end
end
end endtask
// ---- the expected-value side of the stream checks -----------------
logic [7:0] exp_rx [1024]; // what the user must receive
int exp_rx_n, got_rx_n;
logic [7:0] src_tx [1024]; // what the user offered
int src_tx_n, src_tx_sent;
int tx_wait; // cycles this byte has waited
int tx_wait_bound;
// ---- the per-cycle property checks -------------------------------
// Only the output EQUATIONS are compared directly. Everything else is
// a property that does not know how the design is built.
task props;
logic exp_rd, exp_wr, exp_oe, exp_bd, exp_tr;
begin
exp_rd = ~((state_o == S_RD) && ~rxf_n);
exp_wr = ~((state_o == S_WR) && ~txe_n && tx_valid);
exp_oe = ~((state_o == S_RDOE) || (state_o == S_RD) ||
(state_o == S_TAIL));
exp_bd = (state_o == S_WR);
exp_tr = (state_o == S_WR) && ~txe_n && tx_valid;
ck("rd_n", {31'd0, rd_n}, {31'd0, exp_rd});
ck("wr_n", {31'd0, wr_n}, {31'd0, exp_wr});
ck("oe_n", {31'd0, oe_n}, {31'd0, exp_oe});
ck("bus_drive", {31'd0, bus_drive}, {31'd0, exp_bd});
ck("tx_ready", {31'd0, tx_ready}, {31'd0, exp_tr});
// P1 never read an empty bridge
bump; if (!rd_n && rxf_n) begin
err = err + 1;
if (!in_random) $display(" ** P1 over-read (t=%0t)", $time);
end
// P2 never write to a full bridge, and never without a byte
bump; if (!wr_n && (txe_n || !tx_valid)) begin
err = err + 1;
if (!in_random) $display(" ** P2 bad write (t=%0t)", $time);
end
// P3 the turnaround: we must not drive within one cycle of the
// bridge's output enable
bump; if (bus_drive && oe_n_prev === 1'b0) begin
err = err + 1;
if (!in_random) $display(" ** P3 no turnaround (t=%0t)", $time);
end
// P4 output enable leads the strobe
bump; if (!rd_n && oe_n_prev !== 1'b0) begin
err = err + 1;
if (!in_random) $display(" ** P4 strobe without OE lead (t=%0t)", $time);
end
// P5 the skid never overflows, and the bus is never contended
ck("P5 n_lost", {16'd0, n_lost}, 32'd0);
ck("P3 n_contend", {16'd0, n_contend}, 32'd0);
// P9 the burst bound
bump; if (burst_run > MAX_BURST) begin
err = err + 1;
if (!in_random) $display(" ** P9 burst %0d > %0d (t=%0t)",
burst_run, MAX_BURST, $time);
end
end
endtask
// Advance every model that tracks the design. Called right at the edge.
task edge_models;
begin
// P6: a byte leaving the skid must be the next expected one
if (rx_valid && rx_ready) begin
bump;
if (got_rx_n >= exp_rx_n) begin
err = err + 1;
if (err <= 40)
$display(" ** P6 extra byte delivered (t=%0t)", $time);
end else if (rx_data !== exp_rx[got_rx_n]) begin
err = err + 1;
if (err <= 40)
$display(" ** P6 byte %0d: got %02h expected %02h (t=%0t)",
got_rx_n, rx_data, exp_rx[got_rx_n], $time);
got_rx_n = got_rx_n + 1;
end else got_rx_n = got_rx_n + 1;
end
// P7 is checked at the end of a stream by comparing br_tx to src_tx
if (tx_ready) src_tx_sent = src_tx_sent + 1;
// P8: how long has a transmit been waiting FOR THE DESIGN. Cycles
// where the bridge has no room are the environment refusing, not
// the arbiter starving, and counting them measures the bridge.
if (tx_valid && !txe_n && !tx_ready) tx_wait = tx_wait + 1;
else tx_wait = 0;
if (tx_wait > m_txwait_max) m_txwait_max = tx_wait;
// P9's run counter: consecutive strobes without a turnaround
if (state_o == S_TURN) burst_run = 0;
else if (!rd_n || !wr_n) burst_run = burst_run + 1;
// reachability tallies
if (!rd_n) m_rd = m_rd + 1;
if (!wr_n) m_wr = m_wr + 1;
if (state_o == S_TURN) m_turn = m_turn + 1;
if (!rxf_n && !txe_n && tx_valid) m_bothwant = m_bothwant + 1;
if (rx_valid && !rx_ready) m_rxstall = m_rxstall + 1;
if (burst_run == MAX_BURST) m_burstmax = m_burstmax + 1;
oe_n_prev = oe_n;
rd_n_prev = rd_n;
bridge_step;
end
endtask
task step; begin
#1;
props;
@(posedge clk);
edge_models;
#1;
bridge_present;
end endtask
task hard_reset; begin
rst_n = 0; rxf_n = 1; txe_n = 1; data_in = 8'h00;
rx_ready = 0; tx_valid = 0; tx_data = 8'h00;
br_rx_wr = 0; br_rx_rd = 0; br_tx_n = 0; br_room = 0;
exp_rx_n = 0; got_rx_n = 0; src_tx_n = 0; src_tx_sent = 0;
tx_wait = 0; burst_run = 0;
oe_n_prev = 1; rd_n_prev = 1;
repeat (3) @(posedge clk);
#1; rst_n = 1;
@(posedge clk); #1;
bridge_present;
end endtask
// -----------------------------------------------------------------
// Build one of the six states and PROVE it was built.
// -----------------------------------------------------------------
task setup_state(logic [2:0] want);
int guard;
begin
hard_reset;
case (want)
S_IDLE: ; // reset leaves us here
S_RDOE: begin
br_rx[br_rx_wr] = 8'hA5; br_rx_wr = br_rx_wr + 1;
exp_rx[exp_rx_n] = 8'hA5; exp_rx_n = exp_rx_n + 1;
bridge_present; rx_ready = 1;
guard = 0;
while (state_o !== S_RDOE && guard < 20) begin step; guard = guard + 1; end
end
S_RD: begin
br_rx[br_rx_wr] = 8'hA5; br_rx_wr = br_rx_wr + 1;
exp_rx[exp_rx_n] = 8'hA5; exp_rx_n = exp_rx_n + 1;
bridge_present; rx_ready = 1;
guard = 0;
while (state_o !== S_RD && guard < 20) begin step; guard = guard + 1; end
end
S_TAIL: begin
br_rx[br_rx_wr] = 8'hA5; br_rx_wr = br_rx_wr + 1;
exp_rx[exp_rx_n] = 8'hA5; exp_rx_n = exp_rx_n + 1;
bridge_present; rx_ready = 1;
guard = 0;
while (state_o !== S_TAIL && guard < 20) begin step; guard = guard + 1; end
end
S_TURN: begin
br_rx[br_rx_wr] = 8'hA5; br_rx_wr = br_rx_wr + 1;
exp_rx[exp_rx_n] = 8'hA5; exp_rx_n = exp_rx_n + 1;
bridge_present; rx_ready = 1;
guard = 0;
while (state_o !== S_TURN && guard < 20) begin step; guard = guard + 1; end
end
S_WR: begin
br_room = 8; bridge_present;
tx_valid = 1; tx_data = 8'h5A;
src_tx[src_tx_n] = 8'h5A; src_tx_n = src_tx_n + 1;
guard = 0;
while (state_o !== S_WR && guard < 20) begin step; guard = guard + 1; end
end
default: ;
endcase
bump;
if (state_o !== want) begin
err = err + 1; m_setupfail = m_setupfail + 1;
$display(" ** setup: wanted state %0d, reached %0d", want, state_o);
end
end
endtask
// -----------------------------------------------------------------
// PHASE 1: every state against every input combination.
// -----------------------------------------------------------------
int st_i, in_i;
task phase_sweep;
begin
for (st_i = 0; st_i <= 5; st_i = st_i + 1)
for (in_i = 0; in_i <= 15; in_i = in_i + 1) begin
setup_state(3'(st_i));
// in_i bits: [0] a byte is available [1] the bridge has room
// [2] the user can take [3] the user has a byte
if (in_i[0] && br_rx_rd >= br_rx_wr) begin
br_rx[br_rx_wr] = 8'h30 + 8'(in_i);
exp_rx[exp_rx_n] = 8'h30 + 8'(in_i);
br_rx_wr = br_rx_wr + 1; exp_rx_n = exp_rx_n + 1;
end else if (!in_i[0]) begin
br_rx_rd = br_rx_wr; // nothing on offer
end
br_room = in_i[1] ? 8 : 0;
rx_ready = in_i[2];
tx_valid = in_i[3];
if (in_i[3]) begin
tx_data = 8'h80 + 8'(in_i);
src_tx[src_tx_n] = tx_data; src_tx_n = src_tx_n + 1;
end
bridge_present;
step; // the transition itself
step; // and one cycle later
end
end
endtask
// -----------------------------------------------------------------
// PHASE 2: the named boundaries.
// -----------------------------------------------------------------
task load_rx(int n);
int k;
begin
for (k = 0; k < n; k = k + 1) begin
br_rx[br_rx_wr] = 8'h40 + 8'(k);
exp_rx[exp_rx_n] = 8'h40 + 8'(k);
br_rx_wr = br_rx_wr + 1; exp_rx_n = exp_rx_n + 1;
end
bridge_present;
end
endtask
task phase_boundary;
int k, seen_turn, first_burst;
begin
// B1 a long run of available bytes must be broken by a turnaround
// at MAX_BURST. Without the bound the bus is never released.
hard_reset; load_rx(40); rx_ready = 1; br_room = 0;
first_burst = 0; seen_turn = 0;
for (k = 0; k < 60; k = k + 1) begin
step;
if (state_o == S_TURN) seen_turn = seen_turn + 1;
if (burst_run > first_burst) first_burst = burst_run;
end
ck("B1 burst never exceeds MAX", first_burst, MAX_BURST);
bump; if (seen_turn < 4) begin
err = err + 1;
$display(" ** B1 only %0d turnarounds in 60 cycles", seen_turn);
end
// B2 the user stalls mid-burst: no byte may be lost, and the skid
// must be what absorbs the one already committed.
hard_reset; load_rx(12); rx_ready = 1; br_room = 0;
repeat (4) step;
rx_ready = 0;
repeat (8) step; // the committed byte lands in the skid
ck("B2 nothing lost", {16'd0, n_lost}, 32'd0);
rx_ready = 1;
repeat (40) step;
ck("B2 every byte delivered", got_rx_n, exp_rx_n);
// B3 the LAST byte. rxf_n deasserts only after it is read, so a
// machine that gives up when the count reaches one still gets it.
hard_reset; load_rx(1); rx_ready = 1; br_room = 0;
repeat (12) step;
ck("B3 the last byte arrived", got_rx_n, 32'd1);
ck("B3 rxf_n now high", {31'd0, rxf_n}, 32'd1);
// B4 transmit only: every offered byte reaches the bridge, in order
hard_reset; br_room = 16; bridge_present;
tx_valid = 1;
for (k = 0; k < 8; k = k + 1) begin
tx_data = 8'hC0 + 8'(k);
src_tx[src_tx_n] = tx_data; src_tx_n = src_tx_n + 1;
while (!tx_ready) step;
step;
end
tx_valid = 0;
repeat (4) step;
ck("B4 eight bytes written", br_tx_n, 32'd8);
for (k = 0; k < 8; k = k + 1)
ck("B4 byte order", {24'd0, br_tx[k]}, {24'd0, 8'hC0 + 8'(k)});
// B5 THE LIVENESS CASE. The bridge always has a byte to give and
// the user always has a byte to send. A design that lets read
// win unconditionally transmits NOTHING, forever, and every
// check above still passes.
hard_reset; load_rx(200); rx_ready = 1;
br_room = 64; bridge_present;
tx_valid = 1; tx_data = 8'hEE;
src_tx[src_tx_n] = 8'hEE; src_tx_n = src_tx_n + 1;
for (k = 0; k < 300; k = k + 1) step;
bump; if (br_tx_n == 0) begin
err = err + 1;
$display(" ** B5 STARVED: 300 cycles of contention, 0 bytes sent");
end
ck("B5 reads happened too", (n_rx > 0) ? 1 : 0, 32'd1);
tx_wait_bound = 24;
bump; if (m_txwait_max > tx_wait_bound) begin
err = err + 1;
$display(" ** B5 a transmit waited %0d cycles (bound %0d)",
m_txwait_max, tx_wait_bound);
end
// B6 the turnaround is not optional: after a read burst the bus
// must be idle for a cycle before we drive it.
hard_reset; load_rx(6); rx_ready = 1;
br_room = 8; bridge_present;
tx_valid = 1; tx_data = 8'h77;
src_tx[src_tx_n] = 8'h77; src_tx_n = src_tx_n + 1;
repeat (40) step;
ck("B6 no contention", {16'd0, n_contend}, 32'd0);
bump; if (n_turn == 0) begin
err = err + 1; $display(" ** B6 no turnaround cycles at all");
end
// B8 the user runs out of bytes in the middle of a write burst,
// while the bridge still has room. The machine must let go of
// the bus: a receive behind it has to make progress. This is
// the mirror of B5 and it is the ONLY directed scenario that
// constructs it.
hard_reset; br_room = 16; bridge_present;
tx_valid = 1; tx_data = 8'h11;
src_tx[src_tx_n] = 8'h11; src_tx_n = src_tx_n + 1;
k = 0;
while (state_o !== S_WR && k < 20) begin step; k = k + 1; end
ck("B8 reached the write state", {29'd0, state_o}, {29'd0, S_WR});
tx_valid = 0;
load_rx(8); rx_ready = 1;
repeat (60) step;
ck("B8 the receive behind it completed", got_rx_n, exp_rx_n);
// B7 reset in the middle of a read burst returns everything, and
// the skid does not deliver a stale byte afterwards.
hard_reset; load_rx(10); rx_ready = 1; br_room = 0;
repeat (5) step;
// Drop rx_ready and empty the bridge before releasing reset, so
// that the state observed is the one reset produced and not the
// first grant after it. Setting rxf_n by hand is not enough:
// bridge_present drives it from the model at the end of every step.
rx_ready = 0; br_rx_rd = br_rx_wr; bridge_present;
rst_n = 0; step; step; rst_n = 1; step;
ck("B7 state is idle", {29'd0, state_o}, {29'd0, S_IDLE});
ck("B7 skid empty", {31'd0, rx_valid}, 32'd0);
ck("B7 counters clear",{16'd0, n_rx}, 32'd0);
// the bytes the bridge already gave up are gone; resynchronise
exp_rx_n = 0; got_rx_n = 0;
br_rx_rd = br_rx_wr;
end
endtask
// -----------------------------------------------------------------
// PHASE 3: a known stream each way, under backpressure.
// -----------------------------------------------------------------
int rdy_pat [8];
int room_pat [8];
task phase_stream;
int k, cyc;
begin
rdy_pat[0]=1; rdy_pat[1]=1; rdy_pat[2]=0; rdy_pat[3]=1;
rdy_pat[4]=0; rdy_pat[5]=0; rdy_pat[6]=1; rdy_pat[7]=1;
room_pat[0]=1; room_pat[1]=0; room_pat[2]=1; room_pat[3]=1;
room_pat[4]=0; room_pat[5]=1; room_pat[6]=0; room_pat[7]=1;
hard_reset;
for (k = 0; k < 64; k = k + 1) begin
br_rx[br_rx_wr] = 8'h10 + 8'(k);
exp_rx[exp_rx_n] = 8'h10 + 8'(k);
br_rx_wr = br_rx_wr + 1; exp_rx_n = exp_rx_n + 1;
end
for (k = 0; k < 64; k = k + 1) src_tx[k] = 8'hA0 + 8'(k);
src_tx_n = 64; src_tx_sent = 0;
br_room = 0; bridge_present;
tx_valid = 1; tx_data = src_tx[0];
for (cyc = 0; cyc < 1400; cyc = cyc + 1) begin
rx_ready = rdy_pat[cyc % 8];
if (room_pat[cyc % 8] && br_room < 8) br_room = br_room + 1;
if (src_tx_sent < 64) tx_data = src_tx[src_tx_sent];
tx_valid = (src_tx_sent < 64);
bridge_present;
step;
end
ck("P6 all 64 bytes received", got_rx_n, 32'd64);
ck("P7 all 64 bytes written", br_tx_n, 32'd64);
for (k = 0; k < 64; k = k + 1)
ck("P7 order", {24'd0, br_tx[k]}, {24'd0, 8'hA0 + 8'(k)});
// the two directions really did overlap
bump; if (m_bothwant < 100) begin
err = err + 1;
$display(" ** stream: the two directions contended only %0d times",
m_bothwant);
end
end
endtask
// -----------------------------------------------------------------
// PHASE 4: the relational property. The SAME bridge stimulus under
// eight different rx_ready patterns must deliver the same byte
// sequence. Backpressure is allowed to change WHEN; it is not allowed
// to change WHAT.
// -----------------------------------------------------------------
logic [7:0] seq_ref [128];
int seq_ref_n;
int pat_runs;
task phase_independence;
int p, k, cyc, mask;
begin
pat_runs = 0;
for (p = 0; p < 8; p = p + 1) begin
hard_reset;
for (k = 0; k < 48; k = k + 1) begin
br_rx[br_rx_wr] = 8'h60 + 8'(k);
exp_rx[exp_rx_n] = 8'h60 + 8'(k);
br_rx_wr = br_rx_wr + 1; exp_rx_n = exp_rx_n + 1;
end
br_room = 0; bridge_present; // receive only
mask = 0;
for (cyc = 0; cyc < 900; cyc = cyc + 1) begin
// eight genuinely different duty cycles, from "always ready"
// to "ready one cycle in eight"
rx_ready = (p == 0) ? 1'b1 : ((cyc % (p + 1)) == 0);
if (rx_valid && rx_ready && mask < 48) begin
seq_ref[mask] = rx_data; // recorded, then
mask = mask + 1; // compared below
end
step;
end
ck("P10 byte count", mask, 32'd48);
for (k = 0; k < 48; k = k + 1)
ck("P10 sequence", {24'd0, seq_ref[k]}, {24'd0, 8'h60 + 8'(k)});
pat_runs = pat_runs + 1;
if (p > 0) m_rxstall = m_rxstall; // measured in edge_models
end
end
endtask
// -----------------------------------------------------------------
// PHASE 5: random, audited.
// -----------------------------------------------------------------
task phase_random;
int cyc, k;
begin
in_random = 1;
hard_reset;
for (k = 0; k < 900; k = k + 1) begin
br_rx[br_rx_wr] = $urandom_range(255);
exp_rx[exp_rx_n] = br_rx[br_rx_wr]; // the SAME bytes, expected
br_rx_wr = br_rx_wr + 1; exp_rx_n = exp_rx_n + 1;
end
src_tx_n = 0;
br_room = 0; bridge_present;
for (cyc = 0; cyc < 6000; cyc = cyc + 1) begin
rx_ready = ($urandom_range(99) < 62);
if (($urandom_range(99) < 45) && br_room < 8) br_room = br_room + 1;
tx_valid = ($urandom_range(99) < 58);
if (tx_valid) tx_data = $urandom_range(255);
bridge_present;
step;
end
in_random = 0;
end
endtask
initial begin
chk_dir = 0; chk_rnd = 0; err = 0; in_random = 0;
m_rd=0; m_wr=0; m_turn=0; m_skid2=0; m_burstmax=0; m_bothwant=0;
m_altern=0; m_setupfail=0; m_txwait_max=0; m_rxstall=0;
phase_sweep;
$display(" phase 1 state sweep : %0d checks, %0d errors (96 transitions)",
chk_dir, err);
phase_boundary;
$display(" phase 2 boundary : %0d checks, %0d errors", chk_dir, err);
phase_stream;
$display(" phase 3 stream : %0d checks, %0d errors", chk_dir, err);
phase_independence;
$display(" phase 4 independence : %0d checks, %0d errors (%0d ready patterns)",
chk_dir, err, pat_runs);
$display(" ---- DIRECTED-ONLY TOTAL : %0d checks, %0d errors ----", chk_dir, err);
phase_random;
$display("");
$display(" measured reachability (all phases)");
$display(" read strobes ........... %0d", m_rd);
$display(" write strobes .......... %0d", m_wr);
$display(" turnaround cycles ...... %0d", m_turn);
$display(" both wanted the bus .... %0d", m_bothwant);
$display(" bursts reaching MAX .... %0d", m_burstmax);
$display(" user stalled a byte .... %0d", m_rxstall);
$display(" longest transmit wait .. %0d cycles", m_txwait_max);
$display(" setup failures ......... %0d", m_setupfail);
$display("");
$display(" directed checks .......... %0d", chk_dir);
$display(" random checks ............ %0d", chk_rnd);
$display(" TOTAL checks ............. %0d", chk_dir + chk_rnd);
$display(" ERRORS ................... %0d", err);
if (err == 0) $display(" PASS"); else $display(" FAIL");
$finish;
end
endmodule10. VHDL-2008
The same contract again, and here the state type is not an encoding at all:
type state_t is (S_IDLE, S_RDOE, S_RD, S_TAIL, S_TURN, S_WR);There are six values, there is no seventh, and "the state machine fell into an undefined value" is not a sentence that can be written about this file. The synthesiser chooses an encoding; the source does not have one.
The second thing VHDL insists on here is the arithmetic:
signal cnt : natural range 0 to 2;
signal cnt_after : integer range -1 to 3;The Verilog version computes cnt + take - drain in three bits and relies on the
result staying small — which section 4 argues it does. Here the range is
declared, and a design that broke the depth argument stops the simulation at
the instant the value goes out of bounds rather than wrapping quietly. That is
not a stylistic preference: it converts a silent wrong answer into a located
failure, and the bridge-model overflow in section 7 is an example where it did
exactly that to a bug in the testbench that two other languages hid.
-- =====================================================================
-- ft245_sync_if (VHDL-2008) -- the same hardware contract as the
-- Verilog-2005 and SystemVerilog modules: same ports, same six states,
-- same strobe equations, same burst bound, same two-deep skid, same
-- one-cycle latency.
--
-- CLASSIFICATION: simplified synthesisable teaching RTL. It is not an
-- FT2232H / FT600 driver and there is no USB in it; everything on the
-- far side of the bridge has already happened.
--
-- Two things VHDL makes explicit that the other two do not.
--
-- The state is an ENUMERATION with no encoding, so "the state machine
-- fell into an undefined value" is not a sentence that can be written
-- about this file -- there is no undefined value to fall into.
--
-- The skid occupancy arithmetic is done on a RANGE-CONSTRAINED integer.
-- The Verilog version computes cnt + take - drain in three bits and
-- relies on the result staying small; here the range is declared, and
-- a design that broke the depth argument would stop the simulation at
-- the instant it went out of bounds rather than wrapping quietly.
-- =====================================================================
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
entity ft245_sync_if is
generic (
MAX_BURST : natural := 4
);
port (
clk : in std_logic; -- sourced by the BRIDGE
rst_n : in std_logic;
-- bridge side
rxf_n : in std_logic;
txe_n : in std_logic;
data_in : in std_logic_vector(7 downto 0);
rd_n : out std_logic;
wr_n : out std_logic;
oe_n : out std_logic;
bus_drive : out std_logic;
data_out : out std_logic_vector(7 downto 0);
-- user side, receive
rx_valid : out std_logic;
rx_data : out std_logic_vector(7 downto 0);
rx_ready : in std_logic;
-- user side, transmit
tx_valid : in std_logic;
tx_data : in std_logic_vector(7 downto 0);
tx_ready : out std_logic;
-- observation
state_o : out unsigned(2 downto 0);
n_rx : out unsigned(15 downto 0);
n_tx : out unsigned(15 downto 0);
n_turn : out unsigned(15 downto 0);
n_lost : out unsigned(15 downto 0);
n_contend : out unsigned(15 downto 0)
);
end entity ft245_sync_if;
architecture rtl of ft245_sync_if is
type state_t is (S_IDLE, S_RDOE, S_RD, S_TAIL, S_TURN, S_WR);
signal state : state_t := S_IDLE;
signal burst : natural range 0 to 65535 := 0;
signal last_was_rd : std_logic := '0';
signal oe_n_d : std_logic := '1';
signal c_rx, c_tx, c_turn, c_lost, c_cont : unsigned(15 downto 0)
:= (others => '0');
-- the two-deep skid
signal sk0, sk1 : std_logic_vector(7 downto 0) := (others => '0');
signal wptr : std_logic := '0';
signal rptr : std_logic := '0';
signal cnt : natural range 0 to 2 := 0;
-- internal copies of the strobes, because they are read as well as driven
signal rd_n_i, wr_n_i, oe_n_i, bd_i : std_logic;
signal rxv_i : std_logic;
signal take, drain : std_logic;
signal cnt_after : integer range -1 to 3;
signal room_next : std_logic;
signal want_rd, want_wr, grant_rd, grant_wr : std_logic;
signal stay_rd, stay_wr : std_logic;
begin
-- ---- combinational strobes -------------------------------------
-- rd_n is gated by rxf_n and by NOTHING ELSE from outside this
-- entity. One gate from an input pin to an output pin closes timing
-- on a bridge-clocked interface; a path from the user's rx_ready
-- would not, which is exactly why the skid exists instead.
rd_n_i <= '0' when (state = S_RD and rxf_n = '0') else '1';
wr_n_i <= '0' when (state = S_WR and txe_n = '0' and tx_valid = '1')
else '1';
oe_n_i <= '0' when (state = S_RDOE or state = S_RD or state = S_TAIL)
else '1';
bd_i <= '1' when state = S_WR else '0';
rxv_i <= '1' when cnt /= 0 else '0';
rd_n <= rd_n_i;
wr_n <= wr_n_i;
oe_n <= oe_n_i;
bus_drive <= bd_i;
data_out <= tx_data;
rx_valid <= rxv_i;
rx_data <= sk1 when rptr = '1' else sk0;
tx_ready <= '1' when (state = S_WR and txe_n = '0' and tx_valid = '1')
else '0';
take <= not rd_n_i;
drain <= rxv_i and rx_ready;
-- What the occupancy will be after this edge. The read decision uses
-- this and not rx_ready, because rx_ready next cycle is not knowable.
cnt_after <= cnt + 1 when (take = '1' and drain = '0') else
cnt - 1 when (take = '0' and drain = '1') else
cnt;
room_next <= '1' when cnt_after <= 1 else '0';
-- ---- who gets the bus ------------------------------------------
want_rd <= (not rxf_n) and room_next;
want_wr <= (not txe_n) and tx_valid;
-- Both want it: give it to whichever did not go last. Without this
-- the design still passes every check that the data is correct.
grant_rd <= '1' when (want_rd = '1' and
(want_wr = '0' or last_was_rd = '0')) else '0';
grant_wr <= want_wr and (not grant_rd);
-- Keep going only while there is something to move and somewhere to
-- put it. The BUDGET is deliberately not tested here: it is enforced
-- one place only, in the exit condition inside each state, which fires
-- on the cycle the last permitted byte moves. Testing it here as well
-- is dead code -- the exit always fires first.
stay_rd <= '1' when (rxf_n = '0' and room_next = '1') else '0';
stay_wr <= '1' when (txe_n = '0' and tx_valid = '1') else '0';
seq : process (clk, rst_n)
begin
if rst_n = '0' then
state <= S_IDLE;
burst <= 0;
last_was_rd <= '0';
sk0 <= (others => '0');
sk1 <= (others => '0');
wptr <= '0';
rptr <= '0';
cnt <= 0;
oe_n_d <= '1';
c_rx <= (others => '0');
c_tx <= (others => '0');
c_turn <= (others => '0');
c_lost <= (others => '0');
c_cont <= (others => '0');
elsif rising_edge(clk) then
oe_n_d <= oe_n_i;
-- ---- the skid, independent of the state machine ----
if take = '1' then
if cnt = 2 and drain = '0' then
c_lost <= c_lost + 1; -- must never happen
else
if wptr = '1' then sk1 <= data_in; else sk0 <= data_in; end if;
wptr <= not wptr;
c_rx <= c_rx + 1;
end if;
end if;
if drain = '1' then
rptr <= not rptr;
end if;
if take = '1' and drain = '0' then
if cnt /= 2 then cnt <= cnt + 1; end if;
elsif take = '0' and drain = '1' then
cnt <= cnt - 1;
end if;
if (state = S_WR and txe_n = '0' and tx_valid = '1') then
c_tx <= c_tx + 1;
end if;
if bd_i = '1' and oe_n_d = '0' then
c_cont <= c_cont + 1;
end if;
if state = S_TURN then
c_turn <= c_turn + 1;
end if;
-- ---- the state machine ----
case state is
when S_IDLE =>
burst <= 0;
if grant_rd = '1' then
state <= S_RDOE;
elsif grant_wr = '1' then
state <= S_WR;
last_was_rd <= '0';
end if;
-- One cycle of output-enable lead, with no strobe. The bridge is
-- turning its drivers on; nothing is read here.
when S_RDOE =>
state <= S_RD;
last_was_rd <= '1';
when S_RD =>
if take = '1' then
burst <= burst + 1;
end if;
-- Leaving because the budget ran out is DIFFERENT from leaving
-- because there is nothing there, and only one is about fairness.
if stay_rd = '0' or (take = '1' and (burst + 1 >= MAX_BURST)) then
state <= S_TAIL;
end if;
-- The strobe is off but the output enable is still asserted: the
-- bridge has not let go of the bus yet.
when S_TAIL =>
state <= S_TURN;
-- Nobody drives. This cycle moves no data and it is not optional.
when S_TURN =>
state <= S_IDLE;
when S_WR =>
if (txe_n = '0' and tx_valid = '1') then
burst <= burst + 1;
end if;
if stay_wr = '0' or
((txe_n = '0' and tx_valid = '1') and (burst + 1 >= MAX_BURST))
then
state <= S_TURN;
end if;
end case;
end if;
end process seq;
state_o <= to_unsigned(state_t'pos(state), 3);
n_rx <= c_rx;
n_tx <= c_tx;
n_turn <= c_turn;
n_lost <= c_lost;
n_contend <= c_cont;
end architecture rtl;-- =====================================================================
-- tb_ft245_sync_if -- VHDL-2008 testbench for ft245_sync_if.
--
-- Phases 1-4 present the SAME directed stimulus, in the same order, as
-- the Verilog-2005 and SystemVerilog benches, so their directed check
-- counts must agree to the digit. Phase 5 uses its own generator.
--
-- The checker is a set of PROPERTIES plus two byte streams, not a
-- cycle-accurate twin of the state machine. See the Verilog bench for
-- the list; the ten properties are the same ten.
-- =====================================================================
library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;
use ieee.math_real.all;
entity tb_ft245_sync_if is
end entity tb_ft245_sync_if;
architecture sim of tb_ft245_sync_if is
constant MAX_BURST : natural := 4;
constant HALF : time := 10 ns;
constant S_IDLE : unsigned(2 downto 0) := "000";
constant S_RDOE : unsigned(2 downto 0) := "001";
constant S_RD : unsigned(2 downto 0) := "010";
constant S_TAIL : unsigned(2 downto 0) := "011";
constant S_TURN : unsigned(2 downto 0) := "100";
constant S_WR : unsigned(2 downto 0) := "101";
signal clk : std_logic := '0';
signal rst_n : std_logic := '0';
signal rxf_n : std_logic := '1';
signal txe_n : std_logic := '1';
signal data_in : std_logic_vector(7 downto 0) := (others => '0');
signal rx_ready : std_logic := '0';
signal tx_valid : std_logic := '0';
signal tx_data : std_logic_vector(7 downto 0) := (others => '0');
signal rd_n, wr_n, oe_n, bus_drive : std_logic;
signal data_out : std_logic_vector(7 downto 0);
signal rx_valid, tx_ready : std_logic;
signal rx_data : std_logic_vector(7 downto 0);
signal state_o : unsigned(2 downto 0);
signal n_rx, n_tx, n_turn, n_lost, n_contend : unsigned(15 downto 0);
signal done : boolean := false;
type byte_arr is array (natural range <>) of std_logic_vector(7 downto 0);
function b2i (s : std_logic) return integer is
begin
if s = '1' then return 1; else return 0; end if;
end function b2i;
begin
dut : entity work.ft245_sync_if
generic map (MAX_BURST => MAX_BURST)
port map (
clk => clk, rst_n => rst_n,
rxf_n => rxf_n, txe_n => txe_n, data_in => data_in,
rd_n => rd_n, wr_n => wr_n, oe_n => oe_n,
bus_drive => bus_drive, data_out => data_out,
rx_valid => rx_valid, rx_data => rx_data, rx_ready => rx_ready,
tx_valid => tx_valid, tx_data => tx_data, tx_ready => tx_ready,
state_o => state_o,
n_rx => n_rx, n_tx => n_tx, n_turn => n_turn,
n_lost => n_lost, n_contend => n_contend
);
clkgen : process
begin
while not done loop
clk <= '0'; wait for HALF;
clk <= '1'; wait for HALF;
end loop;
wait;
end process clkgen;
stim : process
variable chk_dir, chk_rnd, errs : natural := 0;
variable in_random : boolean := false;
variable shown : natural := 0;
variable oe_n_prev, rd_n_prev : std_logic := '1';
variable burst_run : natural := 0;
variable m_rd, m_wr, m_turn, m_burstmax, m_bothwant : natural := 0;
variable m_setupfail, m_txwait_max, m_rxstall : natural := 0;
-- the bridge model (a VERIFICATION MODEL, not hardware)
variable br_rx : byte_arr(0 to 1023) := (others => (others => '0'));
variable br_rx_wr : natural := 0;
variable br_rx_rd : natural := 0;
variable br_tx : byte_arr(0 to 4095) := (others => (others => '0'));
variable br_tx_n : natural := 0;
variable br_room : natural := 0;
variable exp_rx : byte_arr(0 to 1023) := (others => (others => '0'));
variable exp_rx_n : natural := 0;
variable got_rx_n : natural := 0;
variable src_tx : byte_arr(0 to 1023) := (others => (others => '0'));
variable src_tx_n : natural := 0;
variable src_tx_sent : natural := 0;
variable tx_wait : natural := 0;
variable tx_wait_bound: natural := 24;
variable seq_ref : byte_arr(0 to 127) := (others => (others => '0'));
variable pat_runs : natural := 0;
variable seed1 : positive := 774_113;
variable seed2 : positive := 38_209;
procedure bump is
begin
if in_random then chk_rnd := chk_rnd + 1;
else chk_dir := chk_dir + 1;
end if;
end procedure bump;
procedure ck (what : string; got : integer; exp : integer) is
begin
bump;
if got /= exp then
errs := errs + 1;
if (not in_random) and shown < 40 then
shown := shown + 1;
report " ** " & what & ": got " & integer'image(got) &
" expected " & integer'image(exp) severity warning;
end if;
end if;
end procedure ck;
procedure fail (what : string) is
begin
bump;
errs := errs + 1;
if (not in_random) and shown < 40 then
shown := shown + 1;
report " ** " & what severity warning;
end if;
end procedure fail;
procedure bridge_present is
begin
if br_rx_rd < br_rx_wr then
rxf_n <= '0';
data_in <= br_rx(br_rx_rd);
else
rxf_n <= '1';
data_in <= (others => '0');
end if;
if br_room = 0 then txe_n <= '1'; else txe_n <= '0'; end if;
end procedure bridge_present;
procedure bridge_step is
begin
if rd_n = '0' then
if br_rx_rd < br_rx_wr then
br_rx_rd := br_rx_rd + 1;
else
errs := errs + 1;
report " ** BRIDGE: read strobe with nothing to give"
severity warning;
end if;
end if;
if wr_n = '0' then
if br_room > 0 and br_tx_n < 4096 then
br_tx(br_tx_n) := data_out;
br_tx_n := br_tx_n + 1;
br_room := br_room - 1;
elsif br_room > 0 then
errs := errs + 1;
report " ** BRIDGE: transmit log full" severity warning;
else
errs := errs + 1;
report " ** BRIDGE: write strobe with no room" severity warning;
end if;
end if;
end procedure bridge_step;
-- Only the output EQUATIONS are compared directly; the rest are
-- properties that do not know how the design is built.
procedure props is
variable e_rd, e_wr, e_oe, e_bd, e_tr : std_logic;
begin
if state_o = S_RD and rxf_n = '0' then e_rd := '0'; else e_rd := '1'; end if;
if state_o = S_WR and txe_n = '0' and tx_valid = '1' then
e_wr := '0'; e_tr := '1';
else
e_wr := '1'; e_tr := '0';
end if;
if state_o = S_RDOE or state_o = S_RD or state_o = S_TAIL then
e_oe := '0';
else
e_oe := '1';
end if;
if state_o = S_WR then e_bd := '1'; else e_bd := '0'; end if;
ck("rd_n", b2i(rd_n), b2i(e_rd));
ck("wr_n", b2i(wr_n), b2i(e_wr));
ck("oe_n", b2i(oe_n), b2i(e_oe));
ck("bus_drive", b2i(bus_drive), b2i(e_bd));
ck("tx_ready", b2i(tx_ready), b2i(e_tr));
bump;
if rd_n = '0' and rxf_n = '1' then
errs := errs + 1;
if (not in_random) then report " ** P1 over-read" severity warning; end if;
end if;
bump;
if wr_n = '0' and (txe_n = '1' or tx_valid = '0') then
errs := errs + 1;
if (not in_random) then report " ** P2 bad write" severity warning; end if;
end if;
bump;
if bus_drive = '1' and oe_n_prev = '0' then
errs := errs + 1;
if (not in_random) then report " ** P3 no turnaround" severity warning; end if;
end if;
bump;
if rd_n = '0' and oe_n_prev /= '0' then
errs := errs + 1;
if (not in_random) then report " ** P4 no OE lead" severity warning; end if;
end if;
ck("P5 n_lost", to_integer(n_lost), 0);
ck("P3 n_contend", to_integer(n_contend), 0);
bump;
if burst_run > MAX_BURST then
errs := errs + 1;
if (not in_random) then report " ** P9 burst bound" severity warning; end if;
end if;
end procedure props;
procedure edge_models is
begin
if rx_valid = '1' and rx_ready = '1' then
bump;
if got_rx_n >= exp_rx_n then
errs := errs + 1;
if shown < 40 then
shown := shown + 1;
report " ** P6 extra byte delivered" severity warning;
end if;
elsif rx_data /= exp_rx(got_rx_n) then
errs := errs + 1;
if shown < 40 then
shown := shown + 1;
report " ** P6 byte " & integer'image(got_rx_n) & " wrong"
severity warning;
end if;
got_rx_n := got_rx_n + 1;
else
got_rx_n := got_rx_n + 1;
end if;
end if;
if tx_ready = '1' then src_tx_sent := src_tx_sent + 1; end if;
-- P8: cycles where the bridge HAD room and the design did not move
if tx_valid = '1' and txe_n = '0' and tx_ready = '0' then
tx_wait := tx_wait + 1;
else
tx_wait := 0;
end if;
if tx_wait > m_txwait_max then m_txwait_max := tx_wait; end if;
if state_o = S_TURN then
burst_run := 0;
elsif rd_n = '0' or wr_n = '0' then
burst_run := burst_run + 1;
end if;
if rd_n = '0' then m_rd := m_rd + 1; end if;
if wr_n = '0' then m_wr := m_wr + 1; end if;
if state_o = S_TURN then m_turn := m_turn + 1; end if;
if rxf_n = '0' and txe_n = '0' and tx_valid = '1' then
m_bothwant := m_bothwant + 1;
end if;
if rx_valid = '1' and rx_ready = '0' then m_rxstall := m_rxstall + 1; end if;
if burst_run = MAX_BURST then m_burstmax := m_burstmax + 1; end if;
oe_n_prev := oe_n;
rd_n_prev := rd_n;
bridge_step;
end procedure edge_models;
procedure step is
begin
wait for 1 ns;
props;
wait until rising_edge(clk);
edge_models;
wait for 1 ns;
bridge_present;
end procedure step;
procedure hard_reset is
begin
rst_n <= '0'; rxf_n <= '1'; txe_n <= '1';
data_in <= (others => '0');
rx_ready <= '0'; tx_valid <= '0'; tx_data <= (others => '0');
br_rx_wr := 0; br_rx_rd := 0; br_tx_n := 0; br_room := 0;
exp_rx_n := 0; got_rx_n := 0; src_tx_n := 0; src_tx_sent := 0;
tx_wait := 0; burst_run := 0;
oe_n_prev := '1'; rd_n_prev := '1';
for i in 0 to 2 loop wait until rising_edge(clk); end loop;
wait for 1 ns;
rst_n <= '1';
wait until rising_edge(clk);
wait for 1 ns;
bridge_present;
end procedure hard_reset;
procedure setup_state (want : unsigned(2 downto 0)) is
variable guard : natural := 0;
begin
hard_reset;
if want = S_IDLE then
null;
elsif want = S_WR then
br_room := 8; bridge_present;
tx_valid <= '1'; tx_data <= x"5A";
src_tx(src_tx_n) := x"5A"; src_tx_n := src_tx_n + 1;
guard := 0;
while state_o /= S_WR and guard < 20 loop
step; guard := guard + 1;
end loop;
else
br_rx(br_rx_wr) := x"A5"; br_rx_wr := br_rx_wr + 1;
exp_rx(exp_rx_n) := x"A5"; exp_rx_n := exp_rx_n + 1;
bridge_present;
rx_ready <= '1';
guard := 0;
while state_o /= want and guard < 20 loop
step; guard := guard + 1;
end loop;
end if;
bump;
if state_o /= want then
errs := errs + 1; m_setupfail := m_setupfail + 1;
report " ** setup: wanted state " &
integer'image(to_integer(want)) & ", reached " &
integer'image(to_integer(state_o)) severity warning;
end if;
end procedure setup_state;
procedure load_rx (n : natural) is
begin
for k in 0 to n - 1 loop
br_rx(br_rx_wr) := std_logic_vector(to_unsigned(16#40# + k, 8));
exp_rx(exp_rx_n) := std_logic_vector(to_unsigned(16#40# + k, 8));
br_rx_wr := br_rx_wr + 1; exp_rx_n := exp_rx_n + 1;
end loop;
bridge_present;
end procedure load_rx;
-- ---- PHASE 1 ----
procedure phase_sweep is
variable inb : std_logic_vector(3 downto 0);
begin
for st_i in 0 to 5 loop
for in_i in 0 to 15 loop
setup_state(to_unsigned(st_i, 3));
inb := std_logic_vector(to_unsigned(in_i, 4));
if inb(0) = '1' and br_rx_rd >= br_rx_wr then
br_rx(br_rx_wr) := std_logic_vector(to_unsigned(16#30# + in_i, 8));
exp_rx(exp_rx_n) := std_logic_vector(to_unsigned(16#30# + in_i, 8));
br_rx_wr := br_rx_wr + 1; exp_rx_n := exp_rx_n + 1;
elsif inb(0) = '0' then
br_rx_rd := br_rx_wr;
end if;
if inb(1) = '1' then br_room := 8; else br_room := 0; end if;
rx_ready <= inb(2);
tx_valid <= inb(3);
if inb(3) = '1' then
tx_data <= std_logic_vector(to_unsigned(16#80# + in_i, 8));
src_tx(src_tx_n) := std_logic_vector(to_unsigned(16#80# + in_i, 8));
src_tx_n := src_tx_n + 1;
end if;
bridge_present;
step;
step;
end loop;
end loop;
end procedure phase_sweep;
-- ---- PHASE 2 ----
procedure phase_boundary is
variable seen_turn, first_burst : natural := 0;
variable gk : natural := 0;
begin
-- B1 a long run of available bytes must be broken at MAX_BURST
hard_reset; load_rx(40); rx_ready <= '1'; br_room := 0;
first_burst := 0; seen_turn := 0;
for k in 0 to 59 loop
step;
if state_o = S_TURN then seen_turn := seen_turn + 1; end if;
if burst_run > first_burst then first_burst := burst_run; end if;
end loop;
ck("B1 burst never exceeds MAX", first_burst, MAX_BURST);
bump;
if seen_turn < 4 then
errs := errs + 1;
report " ** B1 too few turnarounds" severity warning;
end if;
-- B2 the user stalls mid-burst: no byte may be lost
hard_reset; load_rx(12); rx_ready <= '1'; br_room := 0;
for k in 0 to 3 loop step; end loop;
rx_ready <= '0';
for k in 0 to 7 loop step; end loop;
ck("B2 nothing lost", to_integer(n_lost), 0);
rx_ready <= '1';
for k in 0 to 39 loop step; end loop;
ck("B2 every byte delivered", got_rx_n, exp_rx_n);
-- B3 the LAST byte
hard_reset; load_rx(1); rx_ready <= '1'; br_room := 0;
for k in 0 to 11 loop step; end loop;
ck("B3 the last byte arrived", got_rx_n, 1);
ck("B3 rxf_n now high", b2i(rxf_n), 1);
-- B4 transmit only, in order
hard_reset; br_room := 16; bridge_present;
tx_valid <= '1';
for k in 0 to 7 loop
tx_data <= std_logic_vector(to_unsigned(16#C0# + k, 8));
src_tx(src_tx_n) := std_logic_vector(to_unsigned(16#C0# + k, 8));
src_tx_n := src_tx_n + 1;
while tx_ready /= '1' loop step; end loop;
step;
end loop;
tx_valid <= '0';
for k in 0 to 3 loop step; end loop;
ck("B4 eight bytes written", br_tx_n, 8);
for k in 0 to 7 loop
ck("B4 byte order", to_integer(unsigned(br_tx(k))), 16#C0# + k);
end loop;
-- B5 THE LIVENESS CASE
hard_reset; load_rx(200); rx_ready <= '1';
br_room := 64; bridge_present;
tx_valid <= '1'; tx_data <= x"EE";
src_tx(src_tx_n) := x"EE"; src_tx_n := src_tx_n + 1;
for k in 0 to 299 loop step; end loop;
bump;
if br_tx_n = 0 then
errs := errs + 1;
report " ** B5 STARVED: 300 cycles of contention, 0 bytes sent"
severity warning;
end if;
if n_rx > 0 then ck("B5 reads happened too", 1, 1);
else ck("B5 reads happened too", 0, 1); end if;
bump;
if m_txwait_max > tx_wait_bound then
errs := errs + 1;
report " ** B5 a transmit waited " & integer'image(m_txwait_max) &
" cycles" severity warning;
end if;
-- B6 the turnaround is not optional
hard_reset; load_rx(6); rx_ready <= '1';
br_room := 8; bridge_present;
tx_valid <= '1'; tx_data <= x"77";
src_tx(src_tx_n) := x"77"; src_tx_n := src_tx_n + 1;
for k in 0 to 39 loop step; end loop;
ck("B6 no contention", to_integer(n_contend), 0);
bump;
if n_turn = 0 then
errs := errs + 1;
report " ** B6 no turnaround cycles at all" severity warning;
end if;
-- B8 the user runs out of bytes in the middle of a write burst,
-- while the bridge still has room. The machine must let go of
-- the bus: a receive behind it has to make progress. This is
-- the mirror of B5 and it is the ONLY directed scenario that
-- constructs it.
hard_reset; br_room := 16; bridge_present;
tx_valid <= '1'; tx_data <= x"11";
src_tx(src_tx_n) := x"11"; src_tx_n := src_tx_n + 1;
gk := 0;
while state_o /= S_WR and gk < 20 loop step; gk := gk + 1; end loop;
ck("B8 reached the write state", to_integer(state_o), 5);
tx_valid <= '0';
load_rx(8); rx_ready <= '1';
for k in 0 to 59 loop step; end loop;
ck("B8 the receive behind it completed", got_rx_n, exp_rx_n);
-- B7 reset in the middle of a read burst
hard_reset; load_rx(10); rx_ready <= '1'; br_room := 0;
for k in 0 to 4 loop step; end loop;
rx_ready <= '0'; br_rx_rd := br_rx_wr; bridge_present;
rst_n <= '0'; step; step; rst_n <= '1'; step;
ck("B7 state is idle", to_integer(state_o), 0);
ck("B7 skid empty", b2i(rx_valid), 0);
ck("B7 counters clear",to_integer(n_rx), 0);
exp_rx_n := 0; got_rx_n := 0;
br_rx_rd := br_rx_wr;
end procedure phase_boundary;
-- ---- PHASE 3 ----
procedure phase_stream is
type ipat is array (0 to 7) of natural;
constant rdy_pat : ipat := (1, 1, 0, 1, 0, 0, 1, 1);
constant room_pat : ipat := (1, 0, 1, 1, 0, 1, 0, 1);
begin
hard_reset;
for k in 0 to 63 loop
br_rx(br_rx_wr) := std_logic_vector(to_unsigned(16#10# + k, 8));
exp_rx(exp_rx_n) := std_logic_vector(to_unsigned(16#10# + k, 8));
br_rx_wr := br_rx_wr + 1; exp_rx_n := exp_rx_n + 1;
end loop;
for k in 0 to 63 loop
src_tx(k) := std_logic_vector(to_unsigned(16#A0# + k, 8));
end loop;
src_tx_n := 64; src_tx_sent := 0;
br_room := 0; bridge_present;
tx_valid <= '1'; tx_data <= src_tx(0);
for cyc in 0 to 1399 loop
if rdy_pat(cyc mod 8) = 1 then rx_ready <= '1';
else rx_ready <= '0'; end if;
if room_pat(cyc mod 8) = 1 and br_room < 8 then
br_room := br_room + 1;
end if;
if src_tx_sent < 64 then
tx_data <= src_tx(src_tx_sent);
tx_valid <= '1';
else
tx_valid <= '0';
end if;
bridge_present;
step;
end loop;
ck("P6 all 64 bytes received", got_rx_n, 64);
ck("P7 all 64 bytes written", br_tx_n, 64);
for k in 0 to 63 loop
ck("P7 order", to_integer(unsigned(br_tx(k))), 16#A0# + k);
end loop;
bump;
if m_bothwant < 100 then
errs := errs + 1;
report " ** stream: the two directions barely contended"
severity warning;
end if;
end procedure phase_stream;
-- ---- PHASE 4 : the relational property ----
procedure phase_independence is
variable mask : natural;
-- The ready value has to exist as a VARIABLE as well as a signal.
-- A signal assignment does not take effect until the next wait, so
-- testing rx_ready here would test the PREVIOUS cycle's value --
-- which in Verilog, where the same line is a blocking assignment,
-- it does not. One byte goes missing from the recording and the
-- design is not involved.
variable rdy : std_logic;
begin
pat_runs := 0;
for p in 0 to 7 loop
hard_reset;
for k in 0 to 47 loop
br_rx(br_rx_wr) := std_logic_vector(to_unsigned(16#60# + k, 8));
exp_rx(exp_rx_n) := std_logic_vector(to_unsigned(16#60# + k, 8));
br_rx_wr := br_rx_wr + 1; exp_rx_n := exp_rx_n + 1;
end loop;
br_room := 0; bridge_present;
mask := 0;
for cyc in 0 to 899 loop
if p = 0 then
rdy := '1';
elsif (cyc mod (p + 1)) = 0 then
rdy := '1';
else
rdy := '0';
end if;
rx_ready <= rdy;
if rx_valid = '1' and rdy = '1' and mask < 48 then
seq_ref(mask) := rx_data;
mask := mask + 1;
end if;
step;
end loop;
ck("P10 byte count", mask, 48);
for k in 0 to 47 loop
ck("P10 sequence", to_integer(unsigned(seq_ref(k))), 16#60# + k);
end loop;
pat_runs := pat_runs + 1;
end loop;
end procedure phase_independence;
-- ---- PHASE 5 ----
impure function rnd (n : positive) return natural is
variable x : real;
begin
uniform(seed1, seed2, x);
return natural(real(n - 1) * x);
end function rnd;
procedure phase_random is
begin
in_random := true;
hard_reset;
for k in 0 to 899 loop
br_rx(br_rx_wr) := std_logic_vector(to_unsigned(rnd(256), 8));
exp_rx(exp_rx_n) := br_rx(br_rx_wr);
br_rx_wr := br_rx_wr + 1; exp_rx_n := exp_rx_n + 1;
end loop;
src_tx_n := 0;
br_room := 0; bridge_present;
for cyc in 0 to 5999 loop
if rnd(100) < 62 then rx_ready <= '1'; else rx_ready <= '0'; end if;
if rnd(100) < 45 and br_room < 8 then br_room := br_room + 1; end if;
if rnd(100) < 58 then
tx_valid <= '1';
tx_data <= std_logic_vector(to_unsigned(rnd(256), 8));
else
tx_valid <= '0';
end if;
bridge_present;
step;
end loop;
in_random := false;
end procedure phase_random;
begin
phase_sweep;
report " phase 1 state sweep : " & integer'image(chk_dir) &
" checks, " & integer'image(errs) & " errors (96 transitions)";
phase_boundary;
report " phase 2 boundary : " & integer'image(chk_dir) &
" checks, " & integer'image(errs) & " errors";
phase_stream;
report " phase 3 stream : " & integer'image(chk_dir) &
" checks, " & integer'image(errs) & " errors";
phase_independence;
report " phase 4 independence : " & integer'image(chk_dir) &
" checks, " & integer'image(errs) & " errors (" &
integer'image(pat_runs) & " ready patterns)";
report " ---- DIRECTED-ONLY TOTAL : " & integer'image(chk_dir) &
" checks, " & integer'image(errs) & " errors ----";
phase_random;
report " measured reachability (all phases)";
report " read strobes ........... " & integer'image(m_rd);
report " write strobes .......... " & integer'image(m_wr);
report " turnaround cycles ...... " & integer'image(m_turn);
report " both wanted the bus .... " & integer'image(m_bothwant);
report " bursts reaching MAX .... " & integer'image(m_burstmax);
report " user stalled a byte .... " & integer'image(m_rxstall);
report " longest transmit wait .. " & integer'image(m_txwait_max);
report " setup failures ......... " & integer'image(m_setupfail);
report " directed checks .......... " & integer'image(chk_dir);
report " random checks ............ " & integer'image(chk_rnd);
report " TOTAL checks ............. " & integer'image(chk_dir + chk_rnd);
report " ERRORS ................... " & integer'image(errs);
if errs = 0 then report " PASS"; else report " FAIL" severity failure; end if;
done <= true;
wait;
end process stim;
end architecture sim;11. Assertions
Ten properties in five categories, written as SVA for a tool that supports concurrent assertions. Icarus Verilog 13.0 rejects them, so each is listed with the procedural check that enforces it in the runs above — and every one of those named checks executes and is counted in the 115,799.
CATEGORY PROPERTY ENFORCED BY
----------- ---------------------- ------------------------------
safety p_no_overread P1, every cycle of every phase,
and the bridge model independently
safety p_no_badwrite P2, and the bridge model
safety p_turnaround P3, and the n_contend comparison
safety p_oe_leads P4
safety p_strobes_exclusive the rd_n and wr_n equations,
which cannot both be low
bounds p_skid_bounded P5, the n_lost comparison
bounds p_burst_bounded P9, and boundary B1
ordering p_rdoe_then_rd the state sweep, all 16 inputs
progress p_tx_not_starved boundary B5 -- and ONLY B5
progress p_rd_strobes the rd_n equationThe source is in the SystemVerilog module above, behind SVA_ON. One of the ten
is unlike anything earlier in this module.
12. Where UVM Fits
The case for UVM here is narrower than in 29.5, and saying why is more useful than building an environment that does not earn its place.
29.5 had two agents writing to shared state, and the interesting bugs were in the overlap of their timing — a space that grows combinatorially with the register set and that a directed bench has to enumerate by hand. This design also has two agents, but they do not share state: the bridge and the user each talk to the master through their own signals, and the master serialises them. There is no "both wrote the same bit" case to construct.
What there is is a performance envelope, and that is a genuinely different reason to reach for a constrained-random environment.
// =====================================================================
// ILLUSTRATIVE. Not compiled or run in this chapter: Icarus Verilog
// cannot run UVM, and every number in sections 8 and 13 comes from the
// procedural benches above.
//
// The item is not a transaction. It is a PAIR OF OFFERED RATES -- how
// hard each side of the bus is pushing -- held for a while. The design
// has no bug that appears at one rate and not another; it has a
// THROUGHPUT and a LATENCY that are functions of both rates, and the
// question a random environment answers is whether they hold across the
// whole envelope.
// =====================================================================
class bus_load_item extends uvm_sequence_item;
`uvm_object_utils(bus_load_item)
rand int unsigned rx_duty; // percent of cycles the bridge offers a byte
rand int unsigned tx_duty; // percent of cycles the user offers a byte
rand int unsigned rdy_duty; // percent of cycles the user can accept
rand int unsigned room_duty; // percent of cycles the bridge has room
rand int unsigned hold_cycles; // how long this operating point is held
// The corners are where the arbiter is decided, and uniform random
// spends almost all of its time in a middle where nothing is contended.
constraint c_corners {
rx_duty dist { 0 :/ 5, [1:40] :/ 20, [41:95] :/ 30, 100 :/ 45 };
tx_duty dist { 0 :/ 5, [1:40] :/ 20, [41:95] :/ 30, 100 :/ 45 };
rdy_duty dist { [10:50] :/ 30, [51:99] :/ 30, 100 :/ 40 };
}
// Held long enough for a steady state to exist. An operating point
// sampled for three cycles measures the transient, not the envelope.
constraint c_hold { hold_cycles inside { [40 : 400] }; }
endclass
// =====================================================================
// The scoreboard is a LATENCY and THROUGHPUT monitor, not a data
// comparator. The data comparison belongs in the monitors, where it is
// a byte queue and three lines long; there is no reference model to
// build because the design does not transform anything.
// =====================================================================
class bus_sb extends uvm_scoreboard;
`uvm_component_utils(bus_sb)
int unsigned worst_tx_wait; // the liveness bound, measured
int unsigned worst_rx_wait;
int unsigned bytes_rd, bytes_wr, bus_cycles;
// The obligation, and the whole reason the environment exists.
function void check_phase(uvm_phase phase);
if (worst_tx_wait > TX_BOUND)
`uvm_error("ARB", $sformatf(
"a transmit waited %0d cycles with the bridge ready; bound is %0d",
worst_tx_wait, TX_BOUND))
// A direction that moved NO bytes while the other moved thousands is
// starvation even if no single wait exceeded the bound -- which can
// happen if the starved side was rarely offered anything.
if (bytes_wr == 0 && bytes_rd > 100)
`uvm_error("ARB", "the transmit direction never moved a byte")
// And the efficiency floor: turnarounds are overhead, and a design
// that alternates byte-by-byte passes every other check here.
if (bus_cycles > 0 &&
(bytes_rd + bytes_wr) * 100 / bus_cycles < EFFICIENCY_FLOOR)
`uvm_error("PERF", "bus efficiency below the floor")
endfunction
endclass
// =====================================================================
// Coverage over the OPERATING POINT, because that is what the design's
// correctness is parameterised by.
// =====================================================================
class bus_cov extends uvm_subscriber #(bus_load_item);
`uvm_component_utils(bus_cov)
covergroup cg;
cp_rx : coverpoint rx_bucket { bins b[] = {IDLE_D, LOW_D, MID_D, SAT_D}; }
cp_tx : coverpoint tx_bucket { bins b[] = {IDLE_D, LOW_D, MID_D, SAT_D}; }
cp_rdy : coverpoint rdy_bucket { bins b[] = {SPARSE, MID_D, ALWAYS}; }
// The corner the chapter is about: both sides saturated at once.
x_contention : cross cp_rx, cp_tx {
bins both_saturated = binsof(cp_rx) intersect {SAT_D} &&
binsof(cp_tx) intersect {SAT_D};
}
// And the corner the skid is about: reading hard into a stalling user.
x_backpressure : cross cp_rx, cp_rdy {
bins read_into_stall = binsof(cp_rx) intersect {SAT_D} &&
binsof(cp_rdy) intersect {SPARSE};
}
endgroup
endclass13. Mutation Testing
Twelve mutations, each a plausible single mistake, each generated by a script that asserts its replacement applied.
MUT V-ALL V-DIR S-ALL S-DIR H-ALL H-DIR
BASE 0 0 0 0 0 0
M1 2 2 2 2 2 2
M2 215 91 209 91 237 91
M3 9449 8611 9369 8611 9439 8611
M4 792 468 788 468 775 468
M5 15005 8192 14954 8192 14986 8192
M6 220 216 220 216 220 216
M7 1450 10 1454 10 1452 10
M8 1129 473 1133 473 1146 473
M9 1355 782 1337 782 1334 782
M10 1 1 1 1 1 1
M11 1 1 1 1 1 1
M12 4 4 4 4 4 4 M1 reads take unconditional priority, so a busy receive stream
starves transmit forever
M2 the read burst bound is off by one, so a burst runs one long
M3 the bus turnaround cycle is skipped
M4 the output enable no longer leads the read strobe
M5 the skid over-commits by one, so a byte arrives with nowhere
to go
M6 the read strobe is not gated by the bridge having a byte
M7 the write strobe is not gated by the user actually having one
M8 the skid read pointer advances when a byte ARRIVES rather than
when one leaves
M9 the skid write pointer never advances, so every byte lands in
the same slot
M10 the burst budget is never refilled, so every grant moves
exactly one byte
M11 the write state is held without checking the user still has a
byte -- the mirror of M1
M12 the user handshake ignores whether the bridge had room, so a
byte is consumed and never sentBASE reads zero in all six columns, and every directed column is identical
across all three languages — twelve mutations, three languages, thirty-six
measurements, and the directed score depends only on the mutation.
The one the data checks cannot see
M10, M11 and M12 score one, one and four, and each is worth a sentence because thin scores are where a suite's luck shows.
M10 never refills the burst budget, so after the first grant every subsequent
one moves a single byte. The design is entirely correct and roughly four times
slower. The check that catches it is B1 burst never exceeds MAX, written as an
equality rather than a bound — so it fails when the burst is too short as well
as too long. Writing a bound check as == where the target is reachable turns it
into a reachability assertion for free, and that is the only reason M10 is not a
survivor.
M11 is M1 in the other direction: the write state is held without checking the
user still has a byte, so once the user runs dry the machine sits in S_WR
forever and the receive path starves. It survived the first campaign — scoring
zero in all six columns — because no directed scenario deasserted tx_valid while
the bridge still had room and a read was waiting behind it. That is a stimulus
hole, not an equivalent mutant, and the fix is boundary scenario B8, which
constructs exactly that. It now scores one.
M12 consumes the user's byte without writing it. Four checks: the tx_ready
equation and three order comparisons in phase 3.
The survivor that was a design finding
Survivors, equivalents, duplicates
survivors, final none
equivalent mutants found 1 (M2 as first written -- retargeted)
stimulus holes found 1 (M11 -- closed by scenario B8)
duplicate mutants noneThe near-duplicate worth recording: M1 and M11 are the same architectural mistake in opposite directions, and they are not duplicates — they break different properties (M1 fails the transmit bound, M11 fails a receive completeness check), they are caught by different scenarios, and a design could plausibly have one without the other. Two mutations that are mirror images of each other are usually worth keeping both, because the asymmetry in how they are caught is itself information: M1 is caught by a property written in advance, M11 only by a scenario written after it survived.
14. Debugging A Board Bring-Up
The symptom on a development board is almost always the same sentence — "USB isn't working" — and it covers at least six unrelated faults across four layers, only one of which is in your RTL.
Which evidence answers which question
EVIDENCE ANSWERS CANNOT ANSWER
------------------------- ---------------------------- --------------
does the host enumerate whether the BRIDGE is alive, anything about
the bridge at all? powered and connected the FPGA
host-side throughput and whether the latency timer or which side of
round-trip latency the packet size is the limit the bridge is
the bottleneck
an integrated logic the strobes, the flags, the what the bridge
analyser in the FPGA state machine, the skid did with them
occupancy -- everything in
this chapter
the bridge's own status whether ITS buffers are anything about
via the host driver filling your timing
n_lost, n_contend and whether the interface broke which of the two
the burst counters the contract, and when sides broke it
timing report on the
strobe paths whether rd_n and wr_n close whether they are
at the bridge's clock logically rightThe third row is the one people reach for last and should reach for first. Every signal in this chapter is inside the FPGA and costs a few hundred flip-flops to capture. The bridge is opaque; your master is not.
The ladder, in the order that costs least
1 Does the host see the bridge?
no -> not the FPGA. Power, the cable, the bridge's own EEPROM,
or the board's USB connector. The FPGA may not even be
configured.
yes -> continue.
2 Do bytes move at all, in either direction?
neither -> is clk arriving? The bridge sources it, and an
unconfigured or wrongly-constrained clock input is
the commonest bring-up fault. Check the state
machine is leaving IDLE.
one only -> ARBITRATION. See step 4. This is the interesting one.
both -> continue to 3.
3 Are bytes CORRUPT, or merely LATE?
corrupt -> read n_lost. Non-zero means the skid overflowed, so
the depth argument is wrong for this build: check
MAX_BURST, and check whether rx_ready is coming from
a different clock domain than section 16 assumes.
Zero means the corruption is not here -- suspect the
bus turnaround and read n_contend.
late -> the latency timer. Section 3. Do not change the RTL.
4 One direction moves and the other does not.
Read the burst counters and watch the state. If the machine
never enters S_WR while tx_valid is high and txe_n is low, the
arbiter is not alternating -- either last_was_rd is stuck, or
the grant expression lost its alternation term, or the burst
budget is never refilled so the other direction is never
reached. All three are in section 13 as M1, M11 and M10.
5 Intermittent loss at high rates only.
Read n_contend. A non-zero count means the bus was driven within
one cycle of the bridge's output enable, which is a scheduling
error with an ELECTRICAL consequence -- and the electrical
consequence is a current spike that damages nothing visible and
corrupts a byte occasionally. It will not reproduce at low rates
because at low rates the turnaround is never the tight case.First divergence, applied here
The general rule — find the first event after which expected and observed architectural state differ — is cheap to apply to this design because the expected state is a queue and a direction.
EXPECTED, from the host's byte stream alone:
the sequence of bytes the host sent, and the sequence it received
OBSERVED, from a capture of rxf_n, rd_n, txe_n, wr_n and the data bus:
the sequence of bytes actually strobed in each direction
The first index at which the two sequences differ localises the fault to
one side of the bridge. Before that index everything is consistent; after
it, everything is a consequence.15. Four Mistakes That Ship
MODEL "the board has USB, so the FPGA speaks USB"
BUG time spent implementing or debugging USB behaviour that is
entirely inside a chip you did not write
SHIPS AS a week lost on the wrong layer, and a fix applied to the
FPGA for a problem in a driver setting
CAUGHT BY looking at the schematic before the specification.
MODEL "a flow-control flag can be sampled when the data is needed"
BUG rd_n depends combinationally on the user's ready signal
SHIPS AS a design that fails timing at the bridge's clock, or --
worse -- passes timing on one build and not the next, so
the failure arrives with an unrelated change
CAUGHT BY the skid, and by the reasoning that produces its depth.
Mutation M5 is the depth being wrong: 8,192 directed
failures.
MODEL "the bus turnaround is a formality"
BUG driving the bus in the cycle after the bridge's output enable
SHIPS AS occasional corruption at high rates that does not
reproduce at low rates and looks like noise
CAUGHT BY a property, not a data check -- the consequence is
electrical and simulation cannot see it. Mutation M3,
8,611 directed failures, all from the scheduling.
MODEL "the urgent direction should have priority"
BUG unconditional read priority
SHIPS AS a transmit path that works until the day somebody streams
data inward, and then never works again
CAUGHT BY one scenario and one bounded-liveness property. Mutation
M1, 2 directed failures out of 115,799 checks, and zero
without that scenario.16. What This Does Not Cover
NOT MODELLED WHY IT IS OUT OF SCOPE
------------------------------- ------------------------------------
the bridge's USB side entirely it is a finished USB device and its
behaviour is chapters 6 to 13,
already done, inside a chip
the latency timer a property of the bridge and its
host driver; section 3 explains it
and there is no RTL decision in it
the tristate pad bus_drive stands in for an IOBUF's
output enable. The pad's turn-off
time is what the turnaround pays
for, and RTL cannot see it.
the bridge's second channel many bridges present two: one for
and JTAG-over-USB a console and one for programming.
They share one USB device and one
upstream pipe, which is why they
interfere -- a bandwidth question,
not a protocol one
asynchronous FIFO mode an older, slower mode with no clock
from the bridge and a different set
of timing rules
the FPGA's own USB hard block a genuine USB controller, where
29.5 is the relevant chapter
CLOCK DOMAIN CROSSING see below. This is the real one.17. Exercises
1 TRACE
The machine is in S_RD with burst = 2 and skid occupancy 1, and
MAX_BURST is 4. Write out the next six cycles of state, rd_n,
oe_n and occupancy for each of:
(a) rxf_n stays low, rx_ready stays high
(b) rxf_n stays low, rx_ready goes low and stays low
(c) rxf_n goes high on the next cycle
Then say which of the three ends the burst for a reason that has
nothing to do with the budget.
2 THE DEPTH ARGUMENT
Re-derive the skid depth for a master whose read strobe is
registered TWICE -- one flop for the decision and one more to
close timing to a distant pin. What depth is needed, and what is
the general formula in terms of the pipeline depth?
3 THE BOUND
Section 11's liveness bound is MAX_BURST + 8. Derive it from the
state machine, cycle by cycle, and say which cycles the +8 covers.
Then work out the bound for a THREE-direction arbiter and say
whether alternation still suffices as a fairness rule.
4 VERILOG
Add a second read channel: two bridges, one user. Decide first
what fairness means with three contenders, then what the bound
becomes, then write it.
5 SYSTEMVERILOG
Implement the same three-contender contract. Use the enum to make
the extra states explicit, and say whether the exit conditions
are still expressible in one line each.
6 VHDL
Implement it, and declare the range of every counter. Then
deliberately break the depth argument and confirm the simulation
stops at the instant the occupancy goes out of range rather than
producing a wrong answer.
7 TESTBENCH
Phase 4 proves the delivered sequence is independent of the
rx_ready pattern. Write the corresponding phase for the TRANSMIT
direction: the written sequence must be independent of the txe_n
pattern. Then say why that one is easier, and what that tells you
about where the skid is.
8 SVA
Write the receive-side liveness property: a byte the bridge is
presenting is read within a bounded number of cycles, given the
user is willing to take it. State the bound and derive it. Then
say which mutation in section 13 it would catch that P8 does not.
9 UVM
Section 12 randomises duty cycles rather than transactions. Write
the sequence that walks the envelope deterministically -- a sweep
rather than a random walk -- and say what each approach finds that
the other does not.
10 MUTATION
Predict, for each of the twelve mutations, whether it is caught by
a DATA check, a PROPERTY, or a SCENARIO. Then check your three
least confident predictions. M3, M7 and M10 are the interesting
ones.
11 DEBUG
A board streams data inward at full rate and the FPGA's replies
arrive in bursts of one byte every few milliseconds. n_lost is
zero, n_contend is zero, and every byte that arrives is correct.
Name the two faults consistent with all of that -- one in this
chapter and one not -- and the single measurement that separates
them.18. The Interview Answer
"Your FPGA board talks to a PC over USB. Walk me through what is actually between them, and tell me about a bug in that path that testing usually misses."
The first half should be quick and it is where most answers stop too early. The FPGA almost certainly does not speak USB. The connector goes to a bridge chip which is a complete USB device — it enumerates, it has descriptors, it handles endpoints — and what reaches the FPGA is a synchronous FIFO interface: a bidirectional byte bus, two active-low flow-control flags, three strobes, and a clock that the bridge sources. Everything I know about USB stops at that chip. Everything I have to design starts after it.
Worth adding unprompted, because it saves a week on every board: if the complaint is that USB is slow, check the bridge's latency timer before anything else. A small reply arriving milliseconds late is the bridge waiting to see whether more bytes are coming, and no change to the FPGA fixes it.
The second half is the real question. Three things in that interface are easy to get wrong, and they fail differently:
The bus turnaround. One bus, two drivers, so there must be a cycle where neither drives. Skip it and the failure is electrical — a current path between two output stages — which means it corrupts a byte occasionally at high rates and never reproduces on the bench. Simulation cannot show the consequence, so you check the scheduling and count the violations.
The registered strobe, and the skid it forces. The read strobe goes off-chip and has to be one gate from a pin, so it cannot depend on the user's ready signal. That means the decision to read is made a cycle before you know whether the byte has anywhere to go — so the byte arrives regardless and you need somewhere to put it. Two entries, and the two is derived: one for the byte already committed to, one so consecutive reads are possible while the user drains.
And the one testing misses: the arbiter. Read and write share the bus, and the obvious priority is to favour the reader, because a byte the bridge is holding is the one at risk. Do that and the design is perfectly correct — every byte delivered, in order, intact, no contention, no overflow — and it transmits nothing at all whenever the host is streaming inward.
That failure has no wrong output. There is nothing to compare against anything. Every data check, every ordering check, every safety property passes. I measured it: with the arbitration fix removed, the only thing in a 115,000-check suite that fails is one scenario that deliberately loads both directions at once, and a bounded-liveness property that says a transmit must be accepted within twelve cycles. Remove that one scenario and the bug survives the entire suite.
So the thing I would want a team to take from it is that safety and progress are separate obligations and need separate work. "Nothing bad happens" is verified by comparing data. "Something good eventually happens" is not verified by comparing anything — it needs a bound written down as a property, and a scenario that creates the conditions under which the bound must hold. Neither comes for free with more stimulus.
19. What Carries Forward
THE MECHANISM
o the USB connector on a dev board reaches a BRIDGE, not the FPGA,
and the bridge has already finished being a USB device by the time
your RTL runs
o what is on the pins is a byte FIFO: two active-low flags, three
strobes, one shared bus, and a clock you do not own
o a bidirectional bus costs a TURNAROUND cycle, which is why the
design moves bytes in bursts and why the burst length is a
parameter with a trade-off in it
o an off-chip strobe must be registered, so the read decision is made
a cycle early, so a skid is required -- and its depth of two is
derived from that one cycle, not chosen
o the latency timer converts a latency question into a throughput
question, and there is no RTL decision in it
THE RESULT
o 96 transitions -- 6 states x 16 input combinations -- exhausted in
three languages with identical directed counts, and none of the 96
unreachable because the four inputs come from two agents that the
state machine does not constrain
o the delivered byte sequence is identical under all eight rx_ready
duty cycles: backpressure changes WHEN, never WHAT
o the longest a transmit ever waited is 10 cycles, against a derived
bound of 12
THE METHOD
o choose the CHECKER SHAPE to fit the design: a cycle-accurate twin
for a pile of state (29.5), properties and streams for a scheduler
(here). A twin of a scheduler is a second copy of the scheduler.
o SAFETY and PROGRESS are separate obligations. A starved direction
corrupts nothing, so no data check can see it.
o a liveness property is useful only as a BOUNDED one, the bound must
be DERIVED rather than measured, and its antecedent must exclude
the environment legitimately refusing
o M1 survives 188,699 checks in three languages once one eleven-line
scenario is deleted -- measured, not argued
o a surviving mutation is a question about the DESIGN first: M2
scored zero because the bound it removed was stated twice and one
statement was dead code, proven by instrumenting the term
o the other survivor was a stimulus hole, not an equivalence, and the
two look identical until you investigate
o a bound check written as == is a reachability assertion for free,
and it is the only thing that catches M10
o a verification model that COMPLAINS is a second independent
statement of the contract; one that stays quiet hides a bug class
o a declared RANGE is a checker: VHDL stopped on a testbench array
overflow that Verilog and SystemVerilog discarded silently
o a VHDL signal assignment is not visible on the next line, and a
Verilog engineer's first VHDL bench loses exactly one item to it
o randomise the thing your design's correctness is a FUNCTION of --
scenarios for 29.5, operating points hereThat is Module 29. Five case studies and a sixth that turned out not to be about USB at all, and between them a mass-storage pipe, an isochronous video stream, a feedback-controlled audio clock, a reversible connector's attach machine, an embedded endpoint's ownership seam, and a bus master for a chip that does the USB so you do not have to.
The next module stops building and starts reviewing. Six checklists — RTL, verification, compliance, integration, debug and interview — each one a list of the questions somebody should have asked before the thing went out. Several of the entries on them are findings from this module: the reset-scope question from 29.5, the turnaround from here, the liveness obligation from both, and the habit that produced all three — that a green run is a claim, and a claim has to say what it measured.
Continue learning
Related tutorials
- Related topic
Endpoint RTL
Fixed priority starves the interrupt endpoint exactly under the load its deadline was specified for — round robin replaces fairness-as-a-feeling with bounded waiting, a number you can put in a latency budget.
- Related topic
USB vs UART
UART spends zero wires on synchronisation and pays a tolerance budget that shrinks as the frame grows; USB spends a SYNC field, an encoding rule and a PLL to buy that budget away — measured across 5376 exhaustive points, not quoted.
- Related topic
USB vs SPI
SPI selects a peripheral with a wire routed at layout time and USB with an address the host assigned — so a chip-select contention is invisible to every slave (0 of 11) while a duplicate USB address is detected every time (274 of 274).
- Related topic
USB vs Ethernet
USB has one authority that assigns every address; Ethernet has none, so a switch infers the topology from traffic — and an inferred table is wrong 294 times out of 1065 where an assigned one is wrong 0 times out of 130.
Standards & specifications
- Governing standard
- USB-IF (Universal Serial Bus Specification)(opens USB Implementers Forum (USB-IF) in a new tab)
Defines the USB bus — its electrical signalling, connectors, packet and transaction model, device framework and the descriptors a device must expose — together with the device-class specifications layered on it. It does not define host-controller register interfaces (xHCI and EHCI are separate documents) nor any operating system's driver architecture.
This page also covers RTL structure, verification approach and debugging technique. Those are engineering practice built on the standard, not requirements the standard itself imposes.
Where this fits
Part of the USB curriculum.
