Skip to content
VLSI Mentor

SPI · Module 15

Input and Output Delay Constraints

A constraint cannot be simulated — set_input_delay describes a board. So the RTL is one register per pin and no logic, the bench checks only what the constraints assume, and the file is written three times for three tools. Adding that pad register also moves three numbers Module 14 published.

Modules 13, 14 and the first half of 15 have deferred the constraints file four times. The reason is worth stating before the file appears:

It cannot be written from inside the RTL, and it cannot be checked by simulation.

set_input_delay says how long after the master's clock edge its data is valid at this device's pin. set_output_delay says how long before the master's next edge this device's data must be valid at its pin. Neither number is in any HDL file, neither is in any simulation, and both come from the board and the other device's datasheet.

So this chapter has an unusual shape: a very small piece of RTL, a bench that checks only the things the constraints assume, and the constraints themselves written three times for three tools.

1. The Module A Constraints File Is Written About

A constraint is a claim about a path, and the path runs from a pin to a register. If the first register on that path is three levels down inside a hierarchy, behind a synchroniser someone may deepen and a mux someone may add, then:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   * `set_input_delay` still applies -- it is about the PORT -- but the slack it
     produces depends on logic that has nothing to do with the board;
   * a tool asked to pull the register into the pad cannot, because there is logic
     between the pad and the register;
   * and a report showing the path failing does not say whether the fault is the
     board's arrival time or somebody's combinational logic.

So the design is a shell: exactly one register between each pin and everything else, and nothing else at all. Not an edge detector, not a synchroniser, not a polarity comparison — because anything that is not a register is what stops a register reaching the pad.

Ten cycles across six rows. Three pin rows change at cycle 2. Their core-side copies all follow one cycle later at cycle 3. A MISO enable and value are asserted at cycle 5 and the pad drives from cycle 6, releasing together at cycle 8.pins changepins changecore sees them: one cycle, all threecore sees them: one cycle,all threevalue and enable togethervalue and enable togetherclksclk_pinsclk_i (core)mosi_pinmosi_i (core)miso_padt0t1t2t3t4t5t6t7t8t9
Figure 1 — the shell's one obligation, drawn at the pin. Every input follows its pin after exactly one cycle and the SAME one cycle, which is what makes `set_input_delay` a claim about the board rather than about somebody's logic. The output value and its enable are registered together, so the pad's turn-on is a flop's clock-to-out and `set_output_delay` can describe it.

2. What The Shell Costs, Which Is A Datasheet Entry

The shell adds one cycle to every input before the rest of the design sees it. So the slave's recovered-edge latency becomes SYNC_N + 1 rather than SYNC_N, and every derived number in Module 14 moves by one:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   LEAD  >= SYNC_N + 3     chip select to the first SCLK edge   (was SYNC_N + 2, 14.4)
   HALF  >= SYNC_N + 2     one SCLK half-period                 (was SYNC_N + 1)
   GAP   >= SYNC_N + 2     chip select high between transactions (was SYNC_N + 1, 14.5)

3. What The Bench Can And Cannot Check

A constraints file cannot be simulated. So the bench checks the two things the constraints assume:

The latency is exactly one cycle, on every input, and the SAME one cycle. A constraints file written against a one-cycle shell, on a design whose SCLK path is one cycle and whose MOSI path is two, is a design whose delay-matching (Chapter 14.1) is broken in a place no timing report will mention — because both paths are short and both pass. Measured: sclk 1, cs_n 1, mosi 1.

The output enable comes up released, and is registered alongside the value. The first is Chapter 14.5's obligation and is checked while reset is asserted, which is the only time it can be checked at all. The second is what makes set_output_delay describe a flop's clock-to-out rather than a fabric path.

And then it says, in its own log, what it cannot check:

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
   NOT CHECKED, and not checkable:
     whether set_input_delay on the three inputs describes the board
     whether set_output_delay on MISO describes it
     whether SCLK is declared as DATA rather than as a clock
     whether the pad registers were actually placed in the pads

A green run of that bench is evidence about the RTL only, and saying so inside the regression is the only place the limitation travels with the result.

4. The Constraints, Three Ways

The numbers below assume a 100 MHz system clock, a master whose SCLK is 10 MHz, and a board where the longest trace adds about 1 ns. Every one of them is a board number and has to be replaced with the real one.

Vivado — XDC

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
# spi_slave.xdc -- Xilinx Vivado
#
# The system clock. This is the ONLY clock in the design, which is the whole point
# of Architecture B (Chapter 15.3): SCLK is data.
create_clock -period 10.000 -name sys_clk [get_ports clk]

# ---------------------------------------------------------------------------
# SCLK, CS# and MOSI are INPUTS, not clocks. There is deliberately NO
# create_clock on sclk_pin -- a design review that finds one has found the
# problem, because it means somebody has told the tool to build a clock tree for
# a signal the RTL samples as data.
# ---------------------------------------------------------------------------
#
# The numbers: the master launches SCLK and MOSI from its own clock, and they
# arrive here after the master's clock-to-out plus the board delay. A 10 MHz
# master with a 3 ns output delay and a 1 ns trace gives a 4 ns arrival, with a
# minimum of 1 ns if the master is fast and the trace short.
#
# These are stated relative to sys_clk because the design samples them with
# sys_clk. That is not a fiction -- it is exactly what the pad register does.
set_input_delay -clock sys_clk -max 4.000 [get_ports {sclk_pin cs_n_pin mosi_pin}]
set_input_delay -clock sys_clk -min 1.000 [get_ports {sclk_pin cs_n_pin mosi_pin}]

# MISO is driven from a registered value AND a registered enable (see the RTL).
# The master samples it on its own edge, so the requirement is that it is valid
# early enough for the master's setup plus the return trace.
set_output_delay -clock sys_clk -max 5.000 [get_ports miso_pad]
set_output_delay -clock sys_clk -min 0.500 [get_ports miso_pad]

# ---------------------------------------------------------------------------
# The synchroniser chains. ASYNC_REG does two jobs and both matter:
#   1. it stops the tool retiming or SRL-inferring the chain, whose DEPTH the
#      design's correctness depends on (Chapter 14.1);
#   2. it tells timing analysis that the first flop's input is an asynchronous
#      arrival rather than a path to report as failing forever.
#
# Applied to the CELLS rather than to a hierarchy, so that a refactor that moves
# them does not silently drop the attribute.
# ---------------------------------------------------------------------------
set_property ASYNC_REG TRUE [get_cells -hierarchical -filter {NAME =~ *sync_sr_reg*}]
set_property SHREG_EXTRACT NO [get_cells -hierarchical -filter {NAME =~ *sync_sr_reg*}]

# Ask for the pad registers. This is a REQUEST: the tool honours it when nothing
# sits between the pin and the flop, and WARNS when something does (Chapter 15.8).
# The warning is in a log of ten thousand lines, so the mechanism is the structure
# and this only makes the tool's refusal visible.
set_property IOB TRUE [get_ports {sclk_pin cs_n_pin mosi_pin miso_pad}]

# The one path that crosses domains in an Architecture A design (Chapter 15.2).
# Not needed in Architecture B, which has one domain -- and the fact that this
# list is SHORT is the reason to prefer a FIFO over an ad-hoc crossing: two
# exceptions can be reviewed.
# set_max_delay -datapath_only 10.000 -from [get_cells u_a/word_src_reg*] \
#                                     -to   [get_cells u_a/rx_data_reg*]

Quartus — SDC

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
# spi_slave.sdc -- Intel Quartus Prime
#
# Same three claims in a different dialect. The differences are mechanical; the
# one that is NOT mechanical is the synchroniser identification, which Quartus
# needs in order to exempt the chain from its own CDC analysis.

create_clock -period 10.000 -name sys_clk [get_ports clk]

# SCLK, CS# and MOSI as inputs. Again: no create_clock on sclk_pin.
set_input_delay -clock sys_clk -max 4.000 [get_ports {sclk_pin cs_n_pin mosi_pin}]
set_input_delay -clock sys_clk -min 1.000 [get_ports {sclk_pin cs_n_pin mosi_pin}]

set_output_delay -clock sys_clk -max 5.000 [get_ports miso_pad]
set_output_delay -clock sys_clk -min 0.500 [get_ports miso_pad]

# Quartus's equivalent of ASYNC_REG is an assignment rather than an SDC property,
# so it lives in the QSF -- which is worth knowing, because a team that puts
# everything in the SDC will find the chain retimed and the attribute absent.
#
#   set_instance_assignment -name SYNCHRONIZER_IDENTIFICATION \
#       "FORCED IF ASYNCHRONOUS" -to "*sync_sr[0]"
#
# And shift-register inference is disabled per-entity in the QSF too:
#   set_instance_assignment -name AUTO_SHIFT_REGISTER_RECOGNITION OFF \
#       -to "spi_slave_frontend"

# The Architecture A crossing, if present.
# set_max_skew -from [get_registers {*word_src*}] -to [get_registers {*rx_data*}] 2.0
# set_net_delay -from [get_registers {*word_src*}] -to [get_registers {*rx_data*}] -max 8.0

Generic SDC — what survives between tools

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
# spi_slave_generic.sdc
#
# The portable subset, and it is worth knowing how small it is: the CLOCK and the
# DELAYS are portable, and everything that protects a synchroniser is not.
#
# That asymmetry is the practical reason a CDC review cannot be replaced by a
# constraints review -- half of what protects the crossings lives in
# vendor-specific attributes, in a different file, in a different language.

create_clock -period 10.000 -name sys_clk [get_ports clk]

set_input_delay  -clock sys_clk -max 4.000 [get_ports sclk_pin]
set_input_delay  -clock sys_clk -min 1.000 [get_ports sclk_pin]
set_input_delay  -clock sys_clk -max 4.000 [get_ports cs_n_pin]
set_input_delay  -clock sys_clk -min 1.000 [get_ports cs_n_pin]
set_input_delay  -clock sys_clk -max 4.000 [get_ports mosi_pin]
set_input_delay  -clock sys_clk -min 1.000 [get_ports mosi_pin]

set_output_delay -clock sys_clk -max 5.000 [get_ports miso_pad]
set_output_delay -clock sys_clk -min 0.500 [get_ports miso_pad]

# A false path is the WRONG exception for a data crossing, and it appears here as
# a warning rather than as advice. `set_false_path` tells the tool the path does
# not matter; a data crossing's path DOES matter -- its skew has to be bounded, or
# Chapter 15.5's incoherence follows. The right exception is a bounded one:
# set_max_delay -datapath_only, or set_net_delay, or the vendor's skew constraint.
#
#   set_false_path -from [get_clocks sclk_dom] -to [get_clocks sys_clk]   # NO

5. Building the Shell — Three HDLs

The circuit

One register per input, a registered value and enable for the output, and the tri-state at the top driving a port. No reset on the input registers except CS, and that exception is argued in the code.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_io_shell.sv — one register per pin and no logic, because anything else is what stops a register reaching the pad
// spi_io_shell.sv
//
// Chapter 15.7 -- the module a constraints file is written about.
//
// Modules 13, 14 and the first half of 15 have deferred the constraints file three
// times, and the reason is that it cannot be written from inside the RTL. It is
// about the BOARD: how long after the master's clock edge its data is valid at this
// device's pin, and how long after this device's clock edge its data is valid at the
// master's. Neither number is in any HDL file, and no simulation contains either.
//
// So this file is small on purpose. It is the SHELL -- the pad registers and nothing
// else -- and it exists to give the constraints file something with named ports and
// a known latency to be written against.
//
// WHAT THE SHELL IS FOR, BEYOND TIDINESS.
//
// A constraint is a claim about a PATH, and the path runs from a pin to a register.
// If the first register on that path is three levels down inside a hierarchy, behind
// a synchroniser that someone may deepen and a mux someone may add, then:
//
//   * `set_input_delay` still applies -- it is about the port -- but the slack it
//     produces depends on logic that has nothing to do with the board;
//   * a tool asked to pull the register into the pad (`IOB = TRUE`) cannot, because
//     there is logic between the pad and the register;
//   * and a report that shows the path failing does not say whether the fault is the
//     board's arrival time or somebody's combinational logic.
//
// Putting exactly one register between each pin and everything else makes the path a
// constraints file talks about the same path the tool reports on.
//
// THE LATENCY IS THE NUMBER THE CONSTRAINTS ASSUME, AND IT IS CHECKABLE.
//
// This shell adds ONE cycle to every input before the rest of the design sees it. So
// the slave's recovered-edge latency becomes SYNC_N + 1 rather than SYNC_N, and
// every derived number in Module 14 moves by one:
//
//     LEAD  >= SYNC_N + 3     (was SYNC_N + 2, Chapter 14.4)
//     HALF  >= SYNC_N + 2     (was SYNC_N + 1)
//     GAP   >= SYNC_N + 2     (was SYNC_N + 1)
//
// That is the thing worth taking from this file: ADDING A PAD REGISTER CHANGES A
// DATASHEET. It is the right thing to do for timing closure and it is not free, and
// a team that adds it late discovers the cost as a customer's master no longer
// working. The testbench measures the latency so that the number in the constraints
// file and the number in the datasheet are the same number.
//
// WHAT IS DELIBERATELY NOT HERE. No logic. Not an edge detector, not a synchroniser,
// not a polarity comparison. Anything that is not a register belongs behind the
// shell, because anything that is not a register is what stops the register reaching
// the pad.

module spi_io_shell #(
    // The output enable is registered too, and its reset value matters more than any
    // other bit in the design: a shell that comes up driving makes every other
    // device on the MISO net unreadable (Chapter 14.5).
    parameter bit OE_RESET = 1'b0
) (
    input  wire  clk,
    input  wire  rst_n,

    // --- the pins ---------------------------------------------------------
    input  wire  sclk_pin,
    input  wire  cs_n_pin,
    input  wire  mosi_pin,
    output wire  miso_pad,

    // --- the core side: one cycle later, and nothing else different -------
    output reg   sclk_i,
    output reg   cs_n_i,
    output reg   mosi_i,
    input  wire  miso_o,
    input  wire  miso_oe
);

    // Input pad registers. One flop, no logic, no reset -- and the missing reset is
    // deliberate rather than an omission worth flagging: resetting an input pad
    // register adds a mux in front of it on some devices, which is exactly the logic
    // that stops it being placed in the pad. The values are meaningless until the
    // core's own reset has released anyway, and the core resets its synchronisers.
    //
    // The exception is CS, which is reset to DESELECTED. A slave that wakes up
    // believing it is selected is worse than one that wakes up believing it is not,
    // and one mux is a fair price for that (Chapter 14.8).
    always_ff @(posedge clk or negedge rst_n) begin
        if (!rst_n) begin
            sclk_i <= 1'b0;
            cs_n_i <= 1'b1;        // deselected
            mosi_i <= 1'b0;
        end else begin
            sclk_i <= sclk_pin;
            cs_n_i <= cs_n_pin;
            mosi_i <= mosi_pin;
        end
    end

    // Output pad registers. Both the value AND the enable are registered, and they
    // must be registered TOGETHER -- a design that registers the value and drives
    // the enable combinationally has a pad whose turn-on time depends on fabric
    // logic, which is the one thing `set_output_delay` cannot describe.
    reg miso_r, oe_r;

    always_ff @(posedge clk or negedge rst_n) begin
        if (!rst_n) begin
            miso_r <= 1'b0;
            oe_r   <= OE_RESET;    // released, unless a design asks otherwise
        end else begin
            miso_r <= miso_o;
            oe_r   <= miso_oe;
        end
    end

    // The tri-state, at the top of the shell and driving a port directly, which is
    // what makes an OBUFT inferrable (Chapter 14.5).
    assign miso_pad = oe_r ? miso_r : 1'bz;

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_io_shell.v — the same design in Verilog-2001
// spi_io_shell.v
//
// Chapter 15.7 -- the module a constraints file is written about.
//
// Modules 13, 14 and the first half of 15 have deferred the constraints file three
// times, and the reason is that it cannot be written from inside the RTL. It is
// about the BOARD: how long after the master's clock edge its data is valid at this
// device's pin, and how long after this device's clock edge its data is valid at the
// master's. Neither number is in any HDL file, and no simulation contains either.
//
// So this file is small on purpose. It is the SHELL -- the pad registers and nothing
// else -- and it exists to give the constraints file something with named ports and
// a known latency to be written against.
//
// WHAT THE SHELL IS FOR, BEYOND TIDINESS.
//
// A constraint is a claim about a PATH, and the path runs from a pin to a register.
// If the first register on that path is three levels down inside a hierarchy, behind
// a synchroniser that someone may deepen and a mux someone may add, then:
//
//   * `set_input_delay` still applies -- it is about the port -- but the slack it
//     produces depends on logic that has nothing to do with the board;
//   * a tool asked to pull the register into the pad (`IOB = TRUE`) cannot, because
//     there is logic between the pad and the register;
//   * and a report that shows the path failing does not say whether the fault is the
//     board's arrival time or somebody's combinational logic.
//
// Putting exactly one register between each pin and everything else makes the path a
// constraints file talks about the same path the tool reports on.
//
// THE LATENCY IS THE NUMBER THE CONSTRAINTS ASSUME, AND IT IS CHECKABLE.
//
// This shell adds ONE cycle to every input before the rest of the design sees it. So
// the slave's recovered-edge latency becomes SYNC_N + 1 rather than SYNC_N, and
// every derived number in Module 14 moves by one:
//
//     LEAD  >= SYNC_N + 3     (was SYNC_N + 2, Chapter 14.4)
//     HALF  >= SYNC_N + 2     (was SYNC_N + 1)
//     GAP   >= SYNC_N + 2     (was SYNC_N + 1)
//
// That is the thing worth taking from this file: ADDING A PAD REGISTER CHANGES A
// DATASHEET. It is the right thing to do for timing closure and it is not free, and
// a team that adds it late discovers the cost as a customer's master no longer
// working. The testbench measures the latency so that the number in the constraints
// file and the number in the datasheet are the same number.
//
// WHAT IS DELIBERATELY NOT HERE. No logic. Not an edge detector, not a synchroniser,
// not a polarity comparison. Anything that is not a register belongs behind the
// shell, because anything that is not a register is what stops the register reaching
// the pad.

module spi_io_shell #(
    // The output enable is registered too, and its reset value matters more than any
    // other bit in the design: a shell that comes up driving makes every other
    // device on the MISO net unreadable (Chapter 14.5).
    parameter OE_RESET = 1'b0
) (
    input  wire  clk,
    input  wire  rst_n,

    // --- the pins ---------------------------------------------------------
    input  wire  sclk_pin,
    input  wire  cs_n_pin,
    input  wire  mosi_pin,
    output wire  miso_pad,

    // --- the core side: one cycle later, and nothing else different -------
    output reg   sclk_i,
    output reg   cs_n_i,
    output reg   mosi_i,
    input  wire  miso_o,
    input  wire  miso_oe
);

    // Input pad registers. One flop, no logic, no reset -- and the missing reset is
    // deliberate rather than an omission worth flagging: resetting an input pad
    // register adds a mux in front of it on some devices, which is exactly the logic
    // that stops it being placed in the pad. The values are meaningless until the
    // core's own reset has released anyway, and the core resets its synchronisers.
    //
    // The exception is CS, which is reset to DESELECTED. A slave that wakes up
    // believing it is selected is worse than one that wakes up believing it is not,
    // and one mux is a fair price for that (Chapter 14.8).
    always @(posedge clk or negedge rst_n) begin
        if (!rst_n) begin
            sclk_i <= 1'b0;
            cs_n_i <= 1'b1;        // deselected
            mosi_i <= 1'b0;
        end else begin
            sclk_i <= sclk_pin;
            cs_n_i <= cs_n_pin;
            mosi_i <= mosi_pin;
        end
    end

    // Output pad registers. Both the value AND the enable are registered, and they
    // must be registered TOGETHER -- a design that registers the value and drives
    // the enable combinationally has a pad whose turn-on time depends on fabric
    // logic, which is the one thing `set_output_delay` cannot describe.
    reg miso_r, oe_r;

    always @(posedge clk or negedge rst_n) begin
        if (!rst_n) begin
            miso_r <= 1'b0;
            oe_r   <= OE_RESET;    // released, unless a design asks otherwise
        end else begin
            miso_r <= miso_o;
            oe_r   <= miso_oe;
        end
    end

    // The tri-state, at the top of the shell and driving a port directly, which is
    // what makes an OBUFT inferrable (Chapter 14.5).
    assign miso_pad = oe_r ? miso_r : 1'bz;

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_io_shell.vhd — the same design in VHDL
-- spi_io_shell.vhd
--
-- Chapter 15.7 -- the module a constraints file is written about.
--
-- Modules 13, 14 and the first half of 15 have deferred the constraints file three
-- times, and the reason is that it cannot be written from inside the RTL. It is
-- about the BOARD: how long after the master's clock edge its data is valid at this
-- device's pin, and how long after this device's clock edge its data is valid at the
-- master's. Neither number is in any HDL file, and no simulation contains either.
--
-- So this file is small on purpose. It is the SHELL -- the pad registers and nothing
-- else -- and it exists to give the constraints file something with named ports and
-- a known latency to be written against.
--
-- WHAT THE SHELL IS FOR, BEYOND TIDINESS.
--
-- A constraint is a claim about a PATH, and the path runs from a pin to a register.
-- If the first register on that path is three levels down inside a hierarchy, behind
-- a synchroniser that someone may deepen and a mux someone may add, then:
--
--   * `set_input_delay` still applies -- it is about the port -- but the slack it
--     produces depends on logic that has nothing to do with the board;
--   * a tool asked to pull the register into the pad (`IOB = TRUE`) cannot, because
--     there is logic between the pad and the register;
--   * and a report that shows the path failing does not say whether the fault is the
--     board's arrival time or somebody's combinational logic.
--
-- Putting exactly one register between each pin and everything else makes the path a
-- constraints file talks about the same path the tool reports on.
--
-- THE LATENCY IS THE NUMBER THE CONSTRAINTS ASSUME, AND IT IS CHECKABLE.
--
-- This shell adds ONE cycle to every input before the rest of the design sees it. So
-- the slave's recovered-edge latency becomes SYNC_N + 1 rather than SYNC_N, and
-- every derived number in Module 14 moves by one:
--
--     LEAD  >= SYNC_N + 3     (was SYNC_N + 2, Chapter 14.4)
--     HALF  >= SYNC_N + 2     (was SYNC_N + 1)
--     GAP   >= SYNC_N + 2     (was SYNC_N + 1)
--
-- That is the thing worth taking from this file: ADDING A PAD REGISTER CHANGES A
-- DATASHEET. It is the right thing to do for timing closure and it is not free, and
-- a team that adds it late discovers the cost as a customer's master no longer
-- working. The testbench measures the latency so that the number in the constraints
-- file and the number in the datasheet are the same number.
--
-- WHAT IS DELIBERATELY NOT HERE. No logic. Not an edge detector, not a synchroniser,
-- not a polarity comparison. Anything that is not a register belongs behind the
-- shell, because anything that is not a register is what stops the register reaching
-- the pad.

library ieee;
use ieee.std_logic_1164.all;

entity spi_io_shell is
    generic (
        -- The output enable is registered too, and its reset value matters more than
        -- any other bit in the design: a shell that comes up driving makes every
        -- other device on the MISO net unreadable (Chapter 14.5).
        OE_RESET : std_logic := '0'
    );
    port (
        clk      : in  std_logic;
        rst_n    : in  std_logic;

        -- the pins
        sclk_pin : in  std_logic;
        cs_n_pin : in  std_logic;
        mosi_pin : in  std_logic;
        miso_pad : out std_logic;

        -- the core side: one cycle later, and nothing else different
        sclk_i   : out std_logic;
        cs_n_i   : out std_logic;
        mosi_i   : out std_logic;
        miso_o   : in  std_logic;
        miso_oe  : in  std_logic
    );
end entity;

architecture rtl of spi_io_shell is
    signal sclk_r : std_logic := '0';
    signal cs_n_r : std_logic := '1';
    signal mosi_r : std_logic := '0';
    signal miso_r : std_logic := '0';
    signal oe_r   : std_logic := '0';
begin

    sclk_i <= sclk_r;
    cs_n_i <= cs_n_r;
    mosi_i <= mosi_r;

    -- Input pad registers. One flop, no logic, and no reset except on CS -- the
    -- missing resets are deliberate rather than an omission: resetting an input pad
    -- register adds a mux in front of it on some devices, which is exactly the logic
    -- that stops it being placed in the pad. CS is the exception, reset to
    -- DESELECTED, because a slave that wakes up believing it is selected is worse
    -- than one that wakes up believing it is not (Chapter 14.8).
    inreg : process (clk, rst_n)
    begin
        if rst_n = '0' then
            sclk_r <= '0';
            cs_n_r <= '1';        -- deselected
            mosi_r <= '0';
        elsif rising_edge(clk) then
            sclk_r <= sclk_pin;
            cs_n_r <= cs_n_pin;
            mosi_r <= mosi_pin;
        end if;
    end process;

    -- Output pad registers. The value AND the enable, registered TOGETHER -- a
    -- design that registers the value and drives the enable combinationally has a
    -- pad whose turn-on time depends on fabric logic, which is the one thing
    -- `set_output_delay` cannot describe.
    outreg : process (clk, rst_n)
    begin
        if rst_n = '0' then
            miso_r <= '0';
            oe_r   <= OE_RESET;   -- released, unless a design asks otherwise
        elsif rising_edge(clk) then
            miso_r <= miso_o;
            oe_r   <= miso_oe;
        end if;
    end process;

    -- The tri-state, at the top of the shell and driving a port directly, which is
    -- what makes an OBUFT inferrable (Chapter 14.5).
    miso_pad <= miso_r when oe_r = '1' else 'Z';

end architecture;

The testbench

Five checks and two statements of what cannot be checked.

Azvya Education Pvt. Ltd.VLSI Mentor
spi_io_shell_tb.sv — five checks, and two statements of what cannot be checked
// spi_io_shell_tb.sv
//
// A constraints file cannot be simulated. That is the chapter's premise and it is
// not a limitation of this bench -- `set_input_delay` describes a board, and no
// board is present. So this bench checks the two things about the shell that a
// constraints file ASSUMES and that simulation can settle:
//
//   1  THE LATENCY IS EXACTLY ONE CYCLE, on every input, and the same one cycle on
//      every input. A constraints file written against a one-cycle shell and a
//      design whose SCLK path is one cycle and whose MOSI path is two is a design
//      whose delay-matching is broken (Chapter 14.1) in a place no timing report
//      will mention, because both paths pass.
//
//   2  THE OUTPUT ENABLE COMES UP RELEASED, and is registered alongside the value
//      rather than ahead of it. Both halves matter: the first is the obligation of
//      Chapter 14.5, and the second is what makes `set_output_delay` describe the
//      real turn-on time rather than a fabric path.
//
// And then it states the thing it cannot check, in the run's own log, so that a
// green regression is not mistaken for a constrained design.

`timescale 1ns/1ps

module spi_io_shell_tb;

    reg clk   = 1'b0;
    reg rst_n = 1'b1;
    always #5 clk = ~clk;

    reg sclk_pin = 1'b0;
    reg cs_n_pin = 1'b1;
    reg mosi_pin = 1'b0;
    reg miso_o   = 1'b0;
    reg miso_oe  = 1'b0;

    wire sclk_i, cs_n_i, mosi_i, miso_pad;

    spi_io_shell #(.OE_RESET(1'b0)) dut (
        .clk(clk), .rst_n(rst_n),
        .sclk_pin(sclk_pin), .cs_n_pin(cs_n_pin), .mosi_pin(mosi_pin),
        .miso_pad(miso_pad),
        .sclk_i(sclk_i), .cs_n_i(cs_n_i), .mosi_i(mosi_i),
        .miso_o(miso_o), .miso_oe(miso_oe)
    );

    integer errors = 0;

    initial begin
        #200_000;
        $display("FAIL: the simulation did not finish within its time limit");
        $finish;
    end

    // Measures how many clock edges pass between a pin changing and the core-side
    // copy following it. The answer must be 1 and must be the same for all three.
    task automatic measure(input integer which, output integer cycles);
        integer n;
        begin
            cycles = -1;
            @(negedge clk);
            case (which)
                0: sclk_pin = ~sclk_pin;
                1: cs_n_pin = ~cs_n_pin;
                2: mosi_pin = ~mosi_pin;
            endcase
            for (n = 1; n <= 8 && cycles < 0; n = n + 1) begin
                @(posedge clk);
                #1;
                case (which)
                    0: if (sclk_i === sclk_pin) cycles = n;
                    1: if (cs_n_i === cs_n_pin) cycles = n;
                    2: if (mosi_i === mosi_pin) cycles = n;
                endcase
            end
        end
    endtask

    integer c_sclk, c_cs, c_mosi, k;
    reg [2:0] pat;

    initial begin
        // Power-on: reset released, then asserted, then released -- the model
        // Chapter 15.2 established, so that an asynchronous reset actually fires.
        rst_n = 1'b1;
        repeat (2) @(posedge clk);
        rst_n = 1'b0;

        // 1. THE ENABLE COMES UP RELEASED, checked WHILE reset is asserted, which is
        //    the only time it can be checked at all. A slave driving a shared bus
        //    during reset makes every other device on the net unreadable, and after
        //    reset has gone the evidence has gone with it.
        #1;
        if (miso_pad !== 1'bz) begin
            $display("  FAIL: the shell drives MISO during reset (pad reads %b) -- every other device on the net is unreadable while this one is held in reset",
                     miso_pad);
            errors = errors + 1;
        end
        $display("  during reset the MISO pad reads high-impedance, so a device held in reset does not make the net unusable for everything else");

        repeat (4) @(posedge clk);
        rst_n = 1'b1;
        repeat (4) @(posedge clk);

        // 2. CS COMES UP DESELECTED. The one input pad register that carries a reset
        //    value, and the reason it earns the mux that a reset costs.
        if (cs_n_i !== 1'b1) begin
            $display("  FAIL: the shell's chip select came out of reset asserted");
            errors = errors + 1;
        end
        $display("  chip select comes out of reset DESELECTED, which is the one input worth spending a reset mux on: a slave that wakes up believing it is selected is worse than one that wakes up believing it is not");

        // 3. EVERY INPUT IS EXACTLY ONE CYCLE, AND THE SAME ONE CYCLE. Measured
        //    rather than assumed, because delay-matching is a lint property and no
        //    timing report will mention an asymmetry here.
        measure(0, c_sclk);
        measure(1, c_cs);
        measure(2, c_mosi);

        $display("  pad-to-core latency:  sclk %0d cycle(s)   cs_n %0d   mosi %0d",
                 c_sclk, c_cs, c_mosi);

        if (c_sclk != 1 || c_cs != 1 || c_mosi != 1) begin
            $display("  FAIL: the shell's latency must be exactly one cycle on every input, got %0d, %0d and %0d",
                     c_sclk, c_cs, c_mosi);
            errors = errors + 1;
        end
        if (!(c_sclk == c_cs && c_cs == c_mosi)) begin
            $display("  FAIL: the three inputs have different latencies, which breaks Chapter 14.1's delay matching in a place no timing report will mention");
            errors = errors + 1;
        end

        // 4. THE OUTPUT VALUE AND ITS ENABLE ARE REGISTERED TOGETHER. Driving the
        //    enable a cycle ahead of the value would put a known-wrong bit on the
        //    bus for one cycle; driving it a cycle behind would release the bus late.
        //    Both are checked by asserting they change on the same edge.
        @(negedge clk);
        miso_o  = 1'b1;
        miso_oe = 1'b1;
        @(posedge clk); #1;
        if (miso_pad !== 1'b1) begin
            $display("  FAIL: the value and the enable did not take effect together; the pad reads %b one cycle after both were asserted",
                     miso_pad);
            errors = errors + 1;
        end
        @(negedge clk);
        miso_oe = 1'b0;
        @(posedge clk); #1;
        if (miso_pad !== 1'bz) begin
            $display("  FAIL: the pad did not release one cycle after the enable was dropped");
            errors = errors + 1;
        end
        $display("  the value and the enable are registered together, so the pad turns on and off one cycle after both -- which is what lets an output delay constraint describe a flop's clock-to-out rather than a fabric path");

        // 5. AND THE SHELL IS TRANSPARENT: whatever goes in comes out, once. Checked
        //    over a pattern rather than a single bit, because a shell that inverted
        //    or held a value would pass every test above.
        pat = 3'b000;
        for (k = 0; k < 8; k = k + 1) begin
            @(negedge clk);
            sclk_pin = k[0];
            cs_n_pin = k[1];
            mosi_pin = k[2];
            @(posedge clk); #1;
            if (sclk_i !== k[0] || cs_n_i !== k[1] || mosi_i !== k[2]) begin
                $display("  FAIL: at pattern %0d the core side read %b%b%b, expected %b%b%b",
                         k, mosi_i, cs_n_i, sclk_i, k[2], k[1], k[0]);
                errors = errors + 1;
            end
        end
        $display("  all eight input combinations pass through unchanged after exactly one cycle, so the shell adds latency and nothing else");

        // 6. THE THING THIS BENCH CANNOT CHECK, printed in the log on purpose.
        $display("  NOT CHECKED HERE, and not checkable: whether `set_input_delay` on the three inputs and `set_output_delay` on MISO describe the board, whether SCLK is correctly declared as DATA rather than as a clock, and whether the pad registers were actually placed in the pads. Those are a constraints file and a place-and-route report, and a green run of this bench is evidence about the RTL only");

        // 7. AND THE NUMBER THIS SHELL CHANGES. Adding it moves every requirement
        //    Module 14 derived by one cycle, which is a datasheet change rather than
        //    an implementation detail.
        $display("  the shell adds one cycle, so Chapter 14.4's lead requirement becomes SYNC_N + 3, the half-period SYNC_N + 2 and Chapter 14.5's gap SYNC_N + 2 -- adding a pad register is the right thing for timing closure and it changes the numbers a master must respect, which is why it belongs in the datasheet and not only in the constraints file");

        if (errors == 0)
            $display("PASS: the shell is one register per pin and no logic, and that is the whole design -- because anything that is not a register is what stops a register being placed in the pad, and a constraint is a claim about the path from a pin to its first register. Measured: every input follows its pin after exactly ONE cycle and all three follow after the SAME one cycle, which is Chapter 14.1's delay-matching obligation made checkable rather than left to a lint rule that no timing report enforces; the MISO pad reads high-impedance throughout reset, so a device held in reset does not make the net unusable for everything else; chip select comes out of reset deselected, which is the one input worth the mux a reset costs; the output value and its enable are registered together, so the pad's turn-on is a flop's clock-to-out and an output delay constraint can describe it; and all eight input combinations pass through unchanged, so the shell adds latency and nothing else. What this bench cannot check is everything the chapter is actually about -- whether the delays describe the board, whether SCLK is declared as data, and whether the registers reached the pads -- and the shell's one cycle moves Module 14's lead, half-period and gap requirements each by one, which makes adding it a datasheet change rather than an implementation detail");
        else
            $display("FAIL: %0d error(s)", errors);
        $finish;
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_io_shell_tb.v — the same bench in Verilog-2001
// spi_io_shell_tb.v
//
// A constraints file cannot be simulated. That is the chapter's premise and it is
// not a limitation of this bench -- `set_input_delay` describes a board, and no
// board is present. So this bench checks the two things about the shell that a
// constraints file ASSUMES and that simulation can settle:
//
//   1  THE LATENCY IS EXACTLY ONE CYCLE, on every input, and the same one cycle on
//      every input. A constraints file written against a one-cycle shell and a
//      design whose SCLK path is one cycle and whose MOSI path is two is a design
//      whose delay-matching is broken (Chapter 14.1) in a place no timing report
//      will mention, because both paths pass.
//
//   2  THE OUTPUT ENABLE COMES UP RELEASED, and is registered alongside the value
//      rather than ahead of it. Both halves matter: the first is the obligation of
//      Chapter 14.5, and the second is what makes `set_output_delay` describe the
//      real turn-on time rather than a fabric path.
//
// And then it states the thing it cannot check, in the run's own log, so that a
// green regression is not mistaken for a constrained design.

`timescale 1ns/1ps

module spi_io_shell_tb;

    reg clk;
    reg rst_n;
    always #5 clk = ~clk;

    reg sclk_pin;
    reg cs_n_pin;
    reg mosi_pin;
    reg miso_o;
    reg miso_oe;

    wire sclk_i, cs_n_i, mosi_i, miso_pad;

    spi_io_shell #(.OE_RESET(1'b0)) dut (
        .clk(clk), .rst_n(rst_n),
        .sclk_pin(sclk_pin), .cs_n_pin(cs_n_pin), .mosi_pin(mosi_pin),
        .miso_pad(miso_pad),
        .sclk_i(sclk_i), .cs_n_i(cs_n_i), .mosi_i(mosi_i),
        .miso_o(miso_o), .miso_oe(miso_oe)
    );

    integer errors;

    initial begin
        #200_000;
        $display("FAIL: the simulation did not finish within its time limit");
        $finish;
    end

    // Measures how many clock edges pass between a pin changing and the core-side
    // copy following it. The answer must be 1 and must be the same for all three.
        task measure;
        input integer which;
        output integer cycles;
        integer n;
        begin
            cycles = -1;
            @(negedge clk);
            case (which)
                0: sclk_pin = ~sclk_pin;
                1: cs_n_pin = ~cs_n_pin;
                2: mosi_pin = ~mosi_pin;
            endcase
            for (n = 1; n <= 8 && cycles < 0; n = n + 1) begin
                @(posedge clk);
                #1;
                case (which)
                    0: if (sclk_i === sclk_pin) cycles = n;
                    1: if (cs_n_i === cs_n_pin) cycles = n;
                    2: if (mosi_i === mosi_pin) cycles = n;
                endcase
            end
        end
    endtask

    integer c_sclk, c_cs, c_mosi, k;
    reg [2:0] pat;

    initial begin
        // Power-on: reset released, then asserted, then released -- the model
        // Chapter 15.2 established, so that an asynchronous reset actually fires.
        rst_n = 1'b1;
        repeat (2) @(posedge clk);
        rst_n = 1'b0;

        // 1. THE ENABLE COMES UP RELEASED, checked WHILE reset is asserted, which is
        //    the only time it can be checked at all. A slave driving a shared bus
        //    during reset makes every other device on the net unreadable, and after
        //    reset has gone the evidence has gone with it.
        #1;
        if (miso_pad !== 1'bz) begin
            $display("  FAIL: the shell drives MISO during reset (pad reads %b) -- every other device on the net is unreadable while this one is held in reset",
                     miso_pad);
            errors = errors + 1;
        end
        $display("  during reset the MISO pad reads high-impedance, so a device held in reset does not make the net unusable for everything else");

        repeat (4) @(posedge clk);
        rst_n = 1'b1;
        repeat (4) @(posedge clk);

        // 2. CS COMES UP DESELECTED. The one input pad register that carries a reset
        //    value, and the reason it earns the mux that a reset costs.
        if (cs_n_i !== 1'b1) begin
            $display("  FAIL: the shell's chip select came out of reset asserted");
            errors = errors + 1;
        end
        $display("  chip select comes out of reset DESELECTED, which is the one input worth spending a reset mux on: a slave that wakes up believing it is selected is worse than one that wakes up believing it is not");

        // 3. EVERY INPUT IS EXACTLY ONE CYCLE, AND THE SAME ONE CYCLE. Measured
        //    rather than assumed, because delay-matching is a lint property and no
        //    timing report will mention an asymmetry here.
        measure(0, c_sclk);
        measure(1, c_cs);
        measure(2, c_mosi);

        $display("  pad-to-core latency:  sclk %0d cycle(s)   cs_n %0d   mosi %0d",
                 c_sclk, c_cs, c_mosi);

        if (c_sclk != 1 || c_cs != 1 || c_mosi != 1) begin
            $display("  FAIL: the shell's latency must be exactly one cycle on every input, got %0d, %0d and %0d",
                     c_sclk, c_cs, c_mosi);
            errors = errors + 1;
        end
        if (!(c_sclk == c_cs && c_cs == c_mosi)) begin
            $display("  FAIL: the three inputs have different latencies, which breaks Chapter 14.1's delay matching in a place no timing report will mention");
            errors = errors + 1;
        end

        // 4. THE OUTPUT VALUE AND ITS ENABLE ARE REGISTERED TOGETHER. Driving the
        //    enable a cycle ahead of the value would put a known-wrong bit on the
        //    bus for one cycle; driving it a cycle behind would release the bus late.
        //    Both are checked by asserting they change on the same edge.
        @(negedge clk);
        miso_o  = 1'b1;
        miso_oe = 1'b1;
        @(posedge clk); #1;
        if (miso_pad !== 1'b1) begin
            $display("  FAIL: the value and the enable did not take effect together; the pad reads %b one cycle after both were asserted",
                     miso_pad);
            errors = errors + 1;
        end
        @(negedge clk);
        miso_oe = 1'b0;
        @(posedge clk); #1;
        if (miso_pad !== 1'bz) begin
            $display("  FAIL: the pad did not release one cycle after the enable was dropped");
            errors = errors + 1;
        end
        $display("  the value and the enable are registered together, so the pad turns on and off one cycle after both -- which is what lets an output delay constraint describe a flop's clock-to-out rather than a fabric path");

        // 5. AND THE SHELL IS TRANSPARENT: whatever goes in comes out, once. Checked
        //    over a pattern rather than a single bit, because a shell that inverted
        //    or held a value would pass every test above.
        pat = 3'b000;
        for (k = 0; k < 8; k = k + 1) begin
            @(negedge clk);
            sclk_pin = k[0];
            cs_n_pin = k[1];
            mosi_pin = k[2];
            @(posedge clk); #1;
            if (sclk_i !== k[0] || cs_n_i !== k[1] || mosi_i !== k[2]) begin
                $display("  FAIL: at pattern %0d the core side read %b%b%b, expected %b%b%b",
                         k, mosi_i, cs_n_i, sclk_i, k[2], k[1], k[0]);
                errors = errors + 1;
            end
        end
        $display("  all eight input combinations pass through unchanged after exactly one cycle, so the shell adds latency and nothing else");

        // 6. THE THING THIS BENCH CANNOT CHECK, printed in the log on purpose.
        $display("  NOT CHECKED HERE, and not checkable: whether `set_input_delay` on the three inputs and `set_output_delay` on MISO describe the board, whether SCLK is correctly declared as DATA rather than as a clock, and whether the pad registers were actually placed in the pads. Those are a constraints file and a place-and-route report, and a green run of this bench is evidence about the RTL only");

        // 7. AND THE NUMBER THIS SHELL CHANGES. Adding it moves every requirement
        //    Module 14 derived by one cycle, which is a datasheet change rather than
        //    an implementation detail.
        $display("  the shell adds one cycle, so Chapter 14.4's lead requirement becomes SYNC_N + 3, the half-period SYNC_N + 2 and Chapter 14.5's gap SYNC_N + 2 -- adding a pad register is the right thing for timing closure and it changes the numbers a master must respect, which is why it belongs in the datasheet and not only in the constraints file");

        if (errors == 0)
            $display("PASS: the shell is one register per pin and no logic, and that is the whole design -- because anything that is not a register is what stops a register being placed in the pad, and a constraint is a claim about the path from a pin to its first register. Measured: every input follows its pin after exactly ONE cycle and all three follow after the SAME one cycle, which is Chapter 14.1's delay-matching obligation made checkable rather than left to a lint rule that no timing report enforces; the MISO pad reads high-impedance throughout reset, so a device held in reset does not make the net unusable for everything else; chip select comes out of reset deselected, which is the one input worth the mux a reset costs; the output value and its enable are registered together, so the pad's turn-on is a flop's clock-to-out and an output delay constraint can describe it; and all eight input combinations pass through unchanged, so the shell adds latency and nothing else. What this bench cannot check is everything the chapter is actually about -- whether the delays describe the board, whether SCLK is declared as data, and whether the registers reached the pads -- and the shell's one cycle moves Module 14's lead, half-period and gap requirements each by one, which makes adding it a datasheet change rather than an implementation detail");
        else
            $display("FAIL: %0d error(s)", errors);
        $finish;
    end


    initial begin
        clk = 1'b0;
        rst_n = 1'b1;
        sclk_pin = 1'b0;
        cs_n_pin = 1'b1;
        mosi_pin = 1'b0;
        miso_o = 1'b0;
        miso_oe = 1'b0;
        errors = 0;
    end

endmodule
Azvya Education Pvt. Ltd.VLSI Mentor
spi_io_shell_tb.vhd — the same bench in VHDL
-- spi_io_shell_tb.vhd
--
-- A constraints file cannot be simulated. That is the chapter's premise and it is
-- not a limitation of this bench -- `set_input_delay` describes a board, and no
-- board is present. So this bench checks the two things about the shell that a
-- constraints file ASSUMES and that simulation can settle:
--
--   1  THE LATENCY IS EXACTLY ONE CYCLE, on every input, and the same one cycle on
--      every input. A constraints file written against a one-cycle shell and a
--      design whose SCLK path is one cycle and whose MOSI path is two is a design
--      whose delay-matching is broken (Chapter 14.1) in a place no timing report
--      will mention, because both paths pass.
--
--   2  THE OUTPUT ENABLE COMES UP RELEASED, and is registered alongside the value
--      rather than ahead of it. Both halves matter: the first is the obligation of
--      Chapter 14.5, and the second is what makes `set_output_delay` describe the
--      real turn-on time rather than a fabric path.
--
-- And then it states the thing it cannot check, in the run's own log, so that a
-- green regression is not mistaken for a constrained design.

library ieee;
use ieee.std_logic_1164.all;
use ieee.numeric_std.all;

entity spi_io_shell_tb is
end entity;

architecture sim of spi_io_shell_tb is

    signal clk   : std_logic := '0';
    signal rst_n : std_logic := '1';
    signal halt  : boolean   := false;

    signal sclk_pin : std_logic := '0';
    signal cs_n_pin : std_logic := '1';
    signal mosi_pin : std_logic := '0';
    signal miso_o   : std_logic := '0';
    signal miso_oe  : std_logic := '0';

    signal sclk_i, cs_n_i, mosi_i, miso_pad : std_logic;

begin

    clkgen : process
    begin
        while not halt loop
            clk <= '0'; wait for 5 ns; clk <= '1'; wait for 5 ns;
        end loop;
        wait;
    end process;

    dut : entity work.spi_io_shell
        generic map (OE_RESET => '0')
        port map (clk => clk, rst_n => rst_n,
                  sclk_pin => sclk_pin, cs_n_pin => cs_n_pin, mosi_pin => mosi_pin,
                  miso_pad => miso_pad,
                  sclk_i => sclk_i, cs_n_i => cs_n_i, mosi_i => mosi_i,
                  miso_o => miso_o, miso_oe => miso_oe);

    watchdog : process
    begin
        wait for 200 us;
        if not halt then
            report "FAIL: the simulation did not finish within its time limit"
                severity failure;
        end if;
        wait;
    end process;

    stim : process
        variable errs : natural := 0;
        variable c_sclk, c_cs, c_mosi : integer;

        -- Measures how many clock edges pass between a pin changing and the core-side
        -- copy following it. The answer must be 1, and the same 1 for all three.
        procedure measure(which : natural; cycles : out integer) is
            variable n : integer := -1;
            variable target : std_logic;
        begin
            n := -1;
            wait until falling_edge(clk);
            case which is
                when 0 => sclk_pin <= not sclk_pin; target := not sclk_pin;
                when 1 => cs_n_pin <= not cs_n_pin; target := not cs_n_pin;
                when others => mosi_pin <= not mosi_pin; target := not mosi_pin;
            end case;
            for i in 1 to 8 loop
                wait until rising_edge(clk);
                wait for 1 ns;
                if n < 0 then
                    case which is
                        when 0 => if sclk_i = sclk_pin then n := i; end if;
                        when 1 => if cs_n_i = cs_n_pin then n := i; end if;
                        when others => if mosi_i = mosi_pin then n := i; end if;
                    end case;
                end if;
            end loop;
            cycles := n;
        end procedure;
    begin
        -- Power-on: reset released, then asserted, then released -- the model
        -- Chapter 15.2 established, so an asynchronous reset actually fires.
        rst_n <= '1';
        for i in 1 to 2 loop wait until rising_edge(clk); end loop;
        rst_n <= '0';

        -- 1. THE ENABLE COMES UP RELEASED, checked WHILE reset is asserted, which is
        --    the only time it can be checked at all.
        wait for 1 ns;
        if miso_pad /= 'Z' then
            report "  FAIL: the shell drives MISO during reset -- every other device on the net is unreadable while this one is held in reset";
            errs := errs + 1;
        end if;
        report "  during reset the MISO pad reads high-impedance, so a device held in reset does not make the net unusable for everything else";

        for i in 1 to 4 loop wait until rising_edge(clk); end loop;
        rst_n <= '1';
        for i in 1 to 4 loop wait until rising_edge(clk); end loop;

        -- 2. CS COMES UP DESELECTED.
        if cs_n_i /= '1' then
            report "  FAIL: the shell's chip select came out of reset asserted";
            errs := errs + 1;
        end if;
        report "  chip select comes out of reset DESELECTED, which is the one input worth spending a reset mux on: a slave that wakes up believing it is selected is worse than one that wakes up believing it is not";

        -- 3. EVERY INPUT IS EXACTLY ONE CYCLE, AND THE SAME ONE CYCLE.
        measure(0, c_sclk);
        measure(1, c_cs);
        measure(2, c_mosi);

        report "  pad-to-core latency:  sclk " & integer'image(c_sclk) &
               " cycle(s)   cs_n " & integer'image(c_cs) &
               "   mosi " & integer'image(c_mosi);

        if c_sclk /= 1 or c_cs /= 1 or c_mosi /= 1 then
            report "  FAIL: the shell's latency must be exactly one cycle on every input";
            errs := errs + 1;
        end if;
        if not (c_sclk = c_cs and c_cs = c_mosi) then
            report "  FAIL: the three inputs have different latencies, which breaks Chapter 14.1's delay matching in a place no timing report will mention";
            errs := errs + 1;
        end if;

        -- 4. THE OUTPUT VALUE AND ITS ENABLE ARE REGISTERED TOGETHER.
        wait until falling_edge(clk);
        miso_o  <= '1';
        miso_oe <= '1';
        wait until rising_edge(clk);
        wait for 1 ns;
        if miso_pad /= '1' then
            report "  FAIL: the value and the enable did not take effect together";
            errs := errs + 1;
        end if;
        wait until falling_edge(clk);
        miso_oe <= '0';
        wait until rising_edge(clk);
        wait for 1 ns;
        if miso_pad /= 'Z' then
            report "  FAIL: the pad did not release one cycle after the enable was dropped";
            errs := errs + 1;
        end if;
        report "  the value and the enable are registered together, so the pad turns on and off one cycle after both -- which is what lets an output delay constraint describe a flop's clock-to-out rather than a fabric path";

        -- 5. AND THE SHELL IS TRANSPARENT over every input combination.
        for k in 0 to 7 loop
            wait until falling_edge(clk);
            sclk_pin <= '1' when (k mod 2) = 1 else '0';
            cs_n_pin <= '1' when ((k / 2) mod 2) = 1 else '0';
            mosi_pin <= '1' when (k / 4) = 1 else '0';
            wait until rising_edge(clk);
            wait for 1 ns;
            if sclk_i /= sclk_pin or cs_n_i /= cs_n_pin or mosi_i /= mosi_pin then
                report "  FAIL: at pattern " & integer'image(k) &
                       " the core side did not match the pins";
                errs := errs + 1;
            end if;
        end loop;
        report "  all eight input combinations pass through unchanged after exactly one cycle, so the shell adds latency and nothing else";

        -- 6. THE THING THIS BENCH CANNOT CHECK, printed in the log on purpose.
        report "  NOT CHECKED HERE, and not checkable: whether `set_input_delay` on the three inputs and `set_output_delay` on MISO describe the board, whether SCLK is correctly declared as DATA rather than as a clock, and whether the pad registers were actually placed in the pads. Those are a constraints file and a place-and-route report, and a green run of this bench is evidence about the RTL only";

        -- 7. AND THE NUMBER THIS SHELL CHANGES.
        report "  the shell adds one cycle, so Chapter 14.4's lead requirement becomes SYNC_N + 3, the half-period SYNC_N + 2 and Chapter 14.5's gap SYNC_N + 2 -- adding a pad register is the right thing for timing closure and it changes the numbers a master must respect, which is why it belongs in the datasheet and not only in the constraints file";

        if errs = 0 then
            report "PASS: the shell is one register per pin and no logic, and that is the whole design -- because anything that is not a register is what stops a register being placed in the pad, and a constraint is a claim about the path from a pin to its first register. Measured: every input follows its pin after exactly ONE cycle and all three follow after the SAME one cycle, which is Chapter 14.1's delay-matching obligation made checkable rather than left to a lint rule that no timing report enforces; the MISO pad reads high-impedance throughout reset, so a device held in reset does not make the net unusable for everything else; chip select comes out of reset deselected, which is the one input worth the mux a reset costs; the output value and its enable are registered together, so the pad's turn-on is a flop's clock-to-out and an output delay constraint can describe it; and all eight input combinations pass through unchanged, so the shell adds latency and nothing else. What this bench cannot check is everything the chapter is actually about -- whether the delays describe the board, whether SCLK is declared as data, and whether the registers reached the pads -- and the shell's one cycle moves Module 14's lead, half-period and gap requirements each by one, which makes adding it a datasheet change rather than an implementation detail";
        else
            report "FAIL: " & integer'image(errs) & " error(s)" severity error;
        end if;

        halt <= true;
        wait;
    end process;

end architecture;

6. Why a Verification Engineer Cares

Check the tri-state state during reset, not after it. A slave driving a shared bus while held in reset makes every other device unreadable, and after reset has gone the evidence has gone with it. That check has to be an assertion live from time zero or a stimulus process that inspects the DUT before releasing reset — and most bench structures cannot do either.

Measure the pad-to-core latency rather than assuming it. Delay-matching is a lint property that no timing report enforces, so an asymmetry between the three inputs is invisible everywhere except in a measurement. The bench measures all three and asserts they are equal and that they equal one.

Print what the bench cannot check. This is the chapter's central verification idea and it generalises well beyond constraints: when a green run is evidence about a narrower thing than a reader will assume, the run should say so. The alternative is that the limitation lives in someone's memory and leaves with them.

And check the shell is transparent over every input combination, not just over one bit. A shell that inverted or held a value would pass a latency measurement and a reset check.

Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// Properties for an I/O shell. Short, because the shell is short -- and the
// shortage is the point: everything interesting about a constraints file is
// outside the RTL.

property p_input_latency_exactly_one;
    // The claim a constraints file makes about the path from the pin.
    @(posedge clk) disable iff (!rst_n)
        1 |=> (sclk_i == $past(sclk_pin));
endproperty

property p_all_inputs_same_latency;
    // Chapter 14.1's delay matching, as a property rather than a lint rule. No
    // timing report will mention an asymmetry here, because both paths pass.
    @(posedge clk) disable iff (!rst_n)
        1 |=> ((sclk_i == $past(sclk_pin)) &&
               (cs_n_i == $past(cs_n_pin)) &&
               (mosi_i == $past(mosi_pin)));
endproperty

property p_released_during_reset;
    // Deliberately NOT inside a disable iff: this property is about the reset state
    // itself, so disabling it during reset would remove its only content.
    @(posedge clk)
        !rst_n |-> (miso_pad === 1'bz);
endproperty

property p_value_and_enable_together;
    // What makes set_output_delay describe a flop's clock-to-out rather than a
    // fabric path. A design that registers one and not the other has a pad whose
    // turn-on depends on logic the constraint cannot see.
    @(posedge clk) disable iff (!rst_n)
        1 |=> (miso_pad === ($past(miso_oe) ? $past(miso_o) : 1'bz));
endproperty

property p_cs_comes_up_deselected;
    // The one input worth spending a reset mux on.
    @(posedge clk)
        $rose(rst_n) |-> cs_n_i;
endproperty
Azvya Education Pvt. Ltd.VLSI Mentor
Snippet
// Coverage. There is little to cover in the shell itself, and the honest coverage
// model says so -- the axes that matter are in the constraints and in the
// place-and-route report, neither of which a covergroup can reach.

covergroup cg_shell @(posedge clk);
    option.per_instance = 1;

    // Every combination of the three inputs, because a shell that held or inverted
    // one would pass a single-bit test.
    combo: coverpoint {mosi_pin, cs_n_pin, sclk_pin} { bins all[] = {[0:7]}; }

    // The output's three states. `released` during reset is the one that matters and
    // the one a bench structure usually cannot reach.
    pad: coverpoint pad_state {
        bins driving_0 = {0};
        bins driving_1 = {1};
        bins released  = {2};
    }

    // Whether reset was asserted while the pad was observed -- which is how the
    // released-during-reset bin gets reached at all.
    in_reset: coverpoint rst_asserted { bins yes = {1}; bins no = {0}; }

    x_pad_reset: cross pad, in_reset;

endgroup

7. Why an FPGA or ASIC Engineer Cares

The absence of a create_clock on SCLK is the single most important line in the file, and it is a line that is not there. A review looks for what is missing: no generated clock on sclk_pin, and three set_input_delay entries that are present. A design with a create_clock on SCLK has been told to build a clock tree for a signal the RTL samples as data, and the resulting reports are about a circuit nobody designed.

The synchroniser protection is not portable and does not live in the SDC. Vivado takes ASYNC_REG as a cell property; Quartus takes SYNCHRONIZER_IDENTIFICATION as a QSF assignment. So the portable subset of the constraints is the clock and the delays, and everything that protects a crossing is vendor-specific and in another file — which is the practical reason a constraints review cannot replace a CDC review.

Attribute the cells, not the hierarchy. get_cells -hierarchical -filter {NAME =~ *sync_sr_reg*} survives a refactor that moves the module; a path-based selector does not, and the failure mode is an attribute that silently matches nothing.

And set_false_path is the wrong exception for anything carrying data. §4's callout has the argument. The short form: a false path says arrival time does not matter, and a data crossing's skew requirement is exactly a statement about arrival time.

8. Failure Signature — A Design That Closes Timing And Fails On The Board

The symptom:

"Timing closes with 2 ns of slack on every path. The interface does not work on hardware at all — not intermittently, never. Simulation is perfect."

What is happening: there is no set_input_delay on the SPI inputs. With no arrival time declared, most tools assume zero — the data arrives exactly at the clock edge — which is the most optimistic possible assumption and produces enormous slack on a path that has none in reality.

Why it fails completely rather than intermittently: the real arrival is several nanoseconds after the edge, and the pad register's setup window closed before it. Every bit is wrong, every time, so the interface never works — which is actually the easy version, because an intermittent version of the same fault is the one that ships.

How to confirm in one command: report the input path's timing and look at what arrival time the tool used. A zero arrival on an asynchronous input is the fault, and the fix is a number from the master's datasheet plus a trace estimate, not an RTL change.

9. Common Misconceptions

"The constraints follow from the RTL." They follow from the board and the other device's datasheet. Nothing in any HDL file contains the master's output delay or the trace length.

"A timing-clean build means the interface will work." With no input delay declared the tool assumes zero and reports enormous slack on a path that has none. §8 is that build.

"SCLK is a clock, so it needs a create_clock." In Architecture B it is data, and a create_clock on it tells the tool to build a clock tree for something the RTL samples. In Architecture A it is a clock and does need one — so the correct answer depends on Chapter 15.4's choice, which is the clearest illustration of why that choice reaches beyond the RTL.

"ASYNC_REG makes a synchroniser correct." It stops the tool destroying a property the structure already has, and tells timing analysis not to report the path forever. The two flops are what make it correct.

"set_false_path between two clock domains is the standard CDC exception." It is correct for a control bit and wrong for anything carrying data, because a data crossing has a skew requirement and a false path removes the only thing bounding it.

"Adding a pad register is an implementation detail." It moves three published requirements by one cycle each. A team that adds it after the datasheet is written has changed the datasheet without noticing.

10. Reason It Through

Q. A master with a 6 ns clock-to-out drives a 2 ns trace into a slave on a 100 MHz system clock. What set_input_delay values, and what do they mean physically?

-max 8.000 and -min something smaller — say 3 ns if the master's fast-corner clock-to-out is 1 ns and the trace's minimum is 2 ns. The max tells the tool the data may not be valid until 8 ns after the reference edge, which is what its setup analysis must accommodate; the min tells it the data may arrive as early as 3 ns, which is what its hold analysis needs. Omitting the min is the more common error and it leaves hold unchecked on a path where an early arrival is entirely possible.

Q. Why is the latency of all three inputs a property that no timing report will ever flag?

Because the property is a relationship between three paths, and timing analysis reports each path against its own constraint. Three paths that are each one flop deep all pass; three paths of depths 1, 1 and 2 also all pass, because each is short. Nothing in the tool's model says the three must be equal — that is Chapter 14.1's requirement, it is structural, and the places it is enforceable are a review, a lint rule that counts flops per input, and a measurement in a bench.

Q. The shell has no reset on sclk_i or mosi_i but does on cs_n_i. Justify both halves.

Resetting an input pad register adds a mux in front of it on some devices, which is exactly the logic that prevents it being placed in the pad — so the default is no reset, and the values are meaningless until the core's own reset has released anyway, because the core resets its synchronisers. CS is the exception because a slave that wakes up believing it is selected is worse than one that wakes up believing it is not (Chapter 14.8): it would drive MISO and process a transaction it has no position in. One mux is a fair price for that, and it is the only input where the price is worth paying.

Q. A reviewer proposes set_false_path from the SCLK domain to the system clock in an Architecture A design, to clear the remaining timing violations. What breaks?

The hand-off register's skew bound. Chapter 15.2's crossing is a multi-bit value read on a flag, and its correctness depends on the bits arriving together enough that the destination samples a coherent word — which is Chapter 15.5's measurement. A false path frees the tool to route one bit locally and another the long way round, and the resulting skew produces exactly the impossible values that chapter counted. The right exception is set_max_delay -datapath_only or a skew constraint, which removes the setup requirement and keeps the bound.

Q. Why does the bench print the list of things it cannot check, rather than the chapter simply saying so?

Because the bench's output is what a future engineer reads when the regression goes green, and the chapter is what nobody reads at that moment. A limitation that lives only in prose is a limitation that is lost at the first handover. Putting it in the $display means the evidence and the scope of the evidence arrive together — which is the same reasoning as Chapter 15.3's asserting that a dangerous ratio passes and printing why.

11. Understanding Check

12. Summary

A constraints file cannot be written from the RTL and cannot be simulated: set_input_delay describes a board and no board is present. So this chapter is a very small piece of RTL, a bench that checks only what the constraints assume, and the file itself written three times.

The RTL is a shell: one register between each pin and everything else, and nothing else at all — because anything that is not a register both lengthens the constrained path with logic unrelated to the board and stops the register reaching the pad.

Measured: every input follows its pin after exactly one cycle and all three follow after the same one cycle, which is Chapter 14.1's delay-matching obligation made checkable rather than left to a lint rule no timing report enforces. The MISO pad reads high-impedance throughout reset. CS comes out of reset deselected, the one input worth the mux a reset costs. The value and its enable are registered together, so the pad's turn-on is a flop's clock-to-out. And all eight input combinations pass through unchanged.

And the shell's one cycle moves three published numbers: lead to SYNC_N + 3, half-period and gap to SYNC_N + 2. Adding a pad register changes a datasheet, and a team that adds it late discovers the cost as a customer's master no longer working.

The constraints themselves: a create_clock on the system clock and deliberately none on SCLK, which is the most important line in the file and is a line that is not there; set_input_delay with both max and min on all three inputs; set_output_delay on MISO; ASYNC_REG and an anti-shift-register attribute applied to the cells; and an IOB request that is a request rather than a mechanism.

Two warnings. The portable subset is small — the clock and the delays travel between tools and everything protecting a synchroniser does not, which is why a constraints review cannot replace a CDC review. And set_false_path is the wrong exception for anything carrying data, because a data crossing has a skew requirement and a false path removes the only thing bounding it.

For verification: check the tri-state during reset; measure the pad-to-core latency on all three inputs rather than assuming it; check transparency over every combination; and print what the bench cannot check, so the evidence and its scope arrive together.

13. What Comes Next

The shell exists and the constraints ask for pad registers. Whether the tool can grant that request is a structural question, and two designs that compute the same function can differ on it.

Chapter 15.8 — FPGA I/O Registers and Timing Closure builds both: the same selection with the mux after the register and with it before, functionally equivalent under a static select and bit-for-bit identical across eight hundred samples. They are not the same circuit, and finding that took two attempts — compared at clock edges they agree on any stimulus, because the difference lives between edges. So "we diffed them cycle by cycle" turns out not to mean "they are the same circuit".

Continue learning