Skip to content
VLSI Mentor

Python for VLSI · Module 1 · Python in the VLSI Engineering Workflow

Why Flow Scripts Moved to Python

Your project almost certainly contains working scripts written in csh, Bash, awk, sed or Perl, and some of them have been correct for a decade. This chapter is not an argument that they were wrong. It is an explanation of what changed in the work itself, because the reason teams reach for Python when they write something new is not fashion and not language aesthetics — it is that flow automation stopped being about filtering lines and started being about holding data. Six hundred results with a test name, a seed, a status, a runtime and a failure category are a structure, and the older tools were built to stream text, not to hold structures. By the end of this chapter you will be able to say which of your existing scripts have no reason to change and which ones are the kind that grew past the tool they are written in.

Foundation10 min readPython for VLSIEDA AutomationScripting LanguagesMaintainabilityEngineering Reasoning

Module 1 · Chapter 1.2 · Why Flow Scripts Moved to Python

1. The Engineering Problem

Here is a real script, and there is nothing wrong with it:

Azvya Education Pvt. Ltd.VLSI Mentor
terminal — count the errors in one log. This is fine. Leave it alone.
grep -c UVM_ERROR sim.log

One line. Instantly readable. No file to maintain, no dependency, nothing to test. If someone rewrote this in Python they would have made the project worse.

Now the same team needs the thing from Chapter 1.1: for 600 runs, collect the test name, the seed, the status, the runtime and the failure category; group failures by cause; compare against last night; and write a summary.

That is not a bigger version of the one-liner. It is a different kind of problem. The one-liner streams text and prints a number. The new job has to hold six fields about six hundred things, keep them associated, and then reason across them.

2. What Actually Changed

Three pressures, in the order they usually bite a team.

Pressure 1 — The data grew fields

A regression result used to be a line of text. Now it is a record:

Azvya Education Pvt. Ltd.VLSI Mentor
one regression result — five fields that must stay associated
test    = axi_burst_rw
seed    = 4210
status  = FAIL
runtime = 182.4 s
cause   = scoreboard mismatch

In awk you can carry that as five parallel arrays indexed by a made-up key, and people do. It works until the day one array gets updated and another does not, and then two tests swap their runtimes and nobody notices for a week.

Python has a built-in way to keep the five fields in one thing:

Azvya Education Pvt. Ltd.VLSI Mentor
preview — the same result, as one value that keeps its fields together
result = {"test": "axi_burst_rw", "seed": 4210, "status": "FAIL", "runtime_s": 182.4}

Read it as English: the curly braces { } group things into one collection. Inside, each entry is a name, then a colon, then its value — "test" has the value "axi_burst_rw". The name result now refers to the whole group of five, so you can pass all five around together and they cannot come apart.

That is called a dictionary, and it gets its own chapter in Module 6. You are not expected to write one yet. The only point here is that the fields cannot drift apart, because there is one thing, not five.

Pressure 2 — Scripts started being read by other people

A flow script that decides whether the nightly regression passed will be read by:

  • the person who wrote it, six months later, having forgotten it;
  • whoever inherits it when that person changes project;
  • somebody at 02:00 trying to work out why it reported PASS.

Compare these two, both of which extract a slack number from a timing report line:

Azvya Education Pvt. Ltd.VLSI Mentor
awk — compact, and this is genuinely a strength for a one-off
awk '/slack/ {print $NF}' timing.rpt
Azvya Education Pvt. Ltd.VLSI Mentor
preview — the same job, written to be read by a stranger
if "slack" in line:
    slack_ns = float(line.split()[-1])

The awk version is shorter and an experienced engineer reads it immediately. The Python version says out loud what $NF means: line.split() breaks the line into its whitespace-separated fields, [-1] takes the last one, float(...) converts that text into a number, and the result is called slack_ns so the next reader knows the units.

Pressure 3 — Somebody asked "how do you know it's right?"

The script from Chapter 1.1 decides pass or fail. Sooner or later a lead asks how you know it is correct.

The answer has to be: I gave it a log I know is failing, and it said FAIL. I gave it a truncated log, and it said the run never finished.

That requires taking the part that reads text and calling it with text you chose. Python is comfortable here: the parsing can be a plain function that takes log text and returns a result, so a test can call it directly with no simulator and no licence. That is the whole of Module 17, and it is the capability that most changes how much a team trusts its automation.

This is achievable in older scripting languages — Perl in particular has a long testing tradition. It is simply easier and more common in Python, and "easier" is what decides what actually gets done on a busy project.

3. What Did Not Change

Now the correction, because the honest version of this story is less dramatic than the usual one.

And plenty of jobs should still not be Python:

The jobThe natural toolWhy
Count matching lines in one filegrep -cone command, nothing to maintain
Set up an environment and launch one commandshellthat is what a shell is for
Rebuild only what changedmakedependency tracking is the tool's whole purpose
Query a loaded design databasethe tool's own Tclonly the tool can see inside itself
Collect, group and compare 600 resultsPythonit is a data problem

Chapter 1.3 takes the first four rows of that table seriously and builds a way of reasoning about the boundary.

4. The Misconception Worth Naming

There is a belief that quietly decides how much of this an engineer bothers to learn: Python is for software engineers.

It is worth being precise about why this is wrong, because the usual rebuttal ("everyone uses Python now!") is not an argument.

The real answer is that the skill in the 600-log problem was never a programming skill. It is knowing:

  • what counts as a failure in your environment, and at what severity it is reported;
  • that a log with no end-of-simulation marker is a different outcome from a log with errors;
  • that two failures with the same first error message are probably one bug;
  • what the team actually needs to see at 09:00.

Every one of those is verification judgment. You already have it. A software engineer writing this tool would have to come and ask you all four questions.

5. A Common Wrong Approach

The wrong lesson to take from this chapter is: our old scripts are legacy, let us port them to Python.

This usually goes badly. A working 300-line Perl script that has processed every regression for six years is a specification of every corner case the team has hit — including ones nobody remembers. A rewrite reproduces the obvious behaviour and quietly loses the corner cases, and you find out which ones over the following months.

The pattern that works is the opposite: leave the working scripts alone, and write the next new thing in Python. Over a couple of projects the centre of gravity shifts on its own, and nothing ever regressed.

6. Exercises

Reasoning only.

Exercise 1 — Sort your own scripts

List four scripts in your flow. For each, mark it data problem (holds records with several fields, reasons across them) or text/command problem (filters, launches, one job). Be honest about the small ones — most projects have more text/command scripts than they think, and those are fine as they are.

Exercise 2 — Find one that grew

Find a script that started as a text problem and grew into a data problem: it now builds parallel arrays, or passes six things around by position, or has a comment explaining which field is which. That script is the one this curriculum is for.

Exercise 3 — Answer the lead's question

Take a script in your flow that decides pass or fail. Write down, in plain English, how you would demonstrate to a lead that it is correct. If the answer is "it has been running for a year and nobody complained", note what that does and does not prove.

Exercise 4 — Defend a one-liner

Pick a one-line shell or awk command in your flow and write one sentence explaining why rewriting it in Python would make the project worse. Being able to argue this direction is what keeps you from over-applying everything that follows.

What the parsers in this track will be reading:

  • The UVM Report Server — the severities a log-triage script keys on, and why choosing the wrong one silently changes the verdict.
  • Verbosity Control — why a simulation log grew large enough that reading it by hand stopped being realistic.

Automation as an engineering practice:

  • Automating GLS Triage — a worked example of the data problem described in §2, in a different corner of the flow.

8. Summary

  • A one-liner that filters text was never the problem. grep -c UVM_ERROR sim.log is the right tool, and this curriculum is not an argument against it.
  • Flow automation grew data. A regression result is five fields about six hundred things, and the older tools were built to stream text rather than hold structures. Keeping fields associated is the pressure that moves work to Python.
  • Two more pressures follow: scripts get read by strangers, and someone eventually asks how you know it is right. Being explicit costs nothing at two hundred lines, and testing a parsing function is easy when the parsing is a function.
  • Python did not replace the older scripts. Established flows still run plenty of csh, awk, sed and Perl, correctly. The real shift is in what teams reach for when writing something new.
  • The hard part is domain judgment, and you already have it. Knowing what counts as a failure is verification knowledge; Python is only the notation.

You should now be able to say which scripts in your own flow have no reason to change, and recognise the ones that have outgrown the tool they are written in.

That leaves the boundary question unanswered, and it deserves more than a table. Chapter 1.3 — Python, Tcl and Bash: Who Owns Which Layer builds a way of reasoning about it that survives the fact that every project draws the lines differently.

Where this fits

Part of the Python curriculum.