Skip to content
VLSI Mentor

Python for VLSI · Module 1 · Python in the VLSI Engineering Workflow

What Is Worth Automating — and What Is Not

The most expensive automation on a project is rarely the script that was too slow or too ugly. It is the script that should never have been written: the one that automated a judgment call, or produced a number nobody could check, or was abandoned by its author and is now a load-bearing mystery. This chapter is the filter that prevents those. It gives you seven questions to ask about a candidate task before you write a line of code, works them through on four real jobs from a verification flow, and then does the harder half — naming the categories of task where automating reliably makes things worse. The arithmetic of hours saved is in here, but it is the shallowest of the seven questions, and treating it as the whole answer is the classic way to end up with a script nobody trusts.

Foundation11 min readPython for VLSIAutomation StrategyEngineering JudgmentMaintainabilityVerification Flow

Module 1 · Chapter 1.4 · What Is Worth Automating — and What Is Not

1. The Engineering Problem

An engineer spends two days writing a script that saves twenty minutes a week. Four months later he changes project. The script keeps running.

Eight months after that, it starts producing a subtly wrong coverage number — a tool upgrade renamed a field in the report. Nobody notices for three weeks, because the number looks plausible and the person who would have recognised it as wrong left.

Work out the real cost of that script. It is not two days. It is two days, plus three weeks of decisions made on a wrong number, plus the afternoon somebody spent finding out why.

Now consider a different script on the same project: forty lines, written in an afternoon, that reads 600 regression logs and produces the morning summary. It has saved a morning a day for two years, and when it breaks it breaks loudly.

So: before writing anything, there are questions to ask.

2. The Seven Questions

Ask these in order. The first two are about value, the middle three about feasibility, and the last two about survival — and the later ones are the ones people skip.

1. How often does it actually happen? Not how often it feels like it happens. Per night, per week, per project.

2. How likely is a human to get it wrong? A task you do twice a year but get wrong half the time may be worth automating even though the time saved is nil. Correctness is a separate reason from speed.

3. Is the input well defined? Can you say precisely what goes in — this file, in this format, from this stage? "Whatever is in the run area" is not a defined input.

4. Can the output be checked? Given a known input, can you state the correct output in advance? If you cannot, you cannot test it, and an untestable script that produces numbers is a liability. This is the question most often skipped, and the most expensive to skip.

5. How bad is a silent wrong answer? If the script is wrong and nobody notices, what happens? A wrongly-formatted report is an annoyance. A regression falsely reported green is a bug reaching silicon.

6. Will anyone else use it? A script used by one person can assume that person's habits. A script three people use is an interface, and changing it breaks their work.

7. Who maintains it in six months? Including the case where that person is not you. If the answer is "nobody", the script's real lifetime is however long until the first tool upgrade.

The arithmetic, and why it is the weakest question

Question 1 has an obvious arithmetic form, and it is worth writing down:

Azvya Education Pvt. Ltd.VLSI Mentor
preview — payback arithmetic. Useful, and the least important of the seven.
hours_saved_per_year = runs_per_week * minutes_each / 60 * 48

Read it as English: take how many times a week the task runs, multiply by the minutes each one costs, divide by 60 to get hours, multiply by roughly 48 working weeks. Everything to the right of = is worked out first, and the answer is stored in the name on the left. Multiplication is * and division is /, the same as everywhere else.

Compare that number against the days it would take to write, test and own the script, and you get a payback time.

3. Working the Questions

Four candidate jobs from a real verification flow, put through the questions:

CandidateOften?Human error?Input defined?Output checkable?Verdict
Collect 600 regression results into a summarynightlyvery likely by handyes — 600 logs in known placesyes — hand-check 10 and compareautomate
Triage which failures share a causenightlylikelyyes — the failure messagesmostly — a human confirms the groupingautomate, with a human reviewing
Decide whether coverage is good enough to sign offonce per milestoneit is a judgment, not an errorno — depends on risk and schedulenodo not automate
Rename a signal across three files, onceoncepossibleyesyes — read the diffdo it by hand

The two clear yeses are the ones from Chapter 1.1. Look at why they pass, because it is not mainly the time saved:

  • The input is a known set of files in known places.
  • The output can be checked: read ten logs by hand and compare against what the script said.
  • The failure is visible: a script that cannot find a log says so.

And look at why the third one fails. It is not hard to compute a coverage percentage — Module 18 does exactly that. What cannot be automated is the decision that 94.2% is enough given the remaining holes, the schedule and the risk. A script that returns sign_off = True has not automated a judgment, it has hidden one.

4. Good Candidates

The pattern behind every job worth automating: the same well-defined thing, many times, with a checkable answer.

Concretely, in a verification and implementation flow:

  • Repeated simulation launches — one test, one seed, the same way every time.
  • Regression result collection — many runs into one summary.
  • Log triage — grouping failures by cause so a review reads five buckets instead of four hundred lines.
  • Coverage aggregation and comparison — the numbers, and the delta from the last run.
  • Report extraction — pulling slack, area or warning counts out of tool reports.
  • Filelist generation — the compile list from the source tree, deterministically.
  • Repetitive configuration and register-map generation — one spec producing several consistent outputs.

Every one of those is a chapter later in this track. They are not examples chosen to look good; they are the list of things the curriculum builds, because they are the jobs that pass the seven questions.

5. Poor Candidates

This half matters more, because nobody warns you about it.

Genuine one-off tasks. Rename a signal in three files. Do it by hand. The script takes longer, and you will never run it again.

Engineering judgment. Is this coverage hole worth chasing? Is this timing violation real? Is the design ready? Automation can supply every input to those decisions and should not make them.

Anything whose output you cannot check. If you cannot say what the right answer is for a known input, you cannot tell a working script from a broken one. You have not automated the task, you have moved it somewhere you cannot see it.

Work that hides simple hardware intent. This one is specific to our field, and it is worth stating now even though the generation chapters are far away in Module 19. A Python loop that emits sixty-four identical channel instances is good automation — it is repetitive structure. A Python loop whose cleverness is the design of a state machine has made the RTL unreviewable: the next engineer has to read Python to understand hardware behaviour. Module 19.1 takes this up properly.

Work you do not understand yet. If you cannot do the task reliably by hand, you cannot specify the script. Automating a process you have not understood produces a fast, confident, wrong answer.

6. Proportional Effort

One more calibration, because this curriculum is about to spend twenty modules on making automation robust, and applying all of it to everything would be its own mistake.

The engineering you invest should be proportional to what the script decides and how many people depend on it:

The scriptProportional effort
Five lines you run once, on your own machinewrite it, use it, delete it
A check you run yourself most daysa clear name, a defined input, a real exit status
A script your team depends on nightlythe above, plus tests, error handling and a documented contract
A script that decides whether a regression passedall of it — and tests that can actually fail

The last row is not enthusiasm. It is the direct consequence of Chapter 1.1: whatever decides pass or fail is part of your verification.

7. Exercises

Reasoning only.

Exercise 1 — Run the seven questions

Pick three repetitive tasks from your own flow. For each, answer all seven questions in one line each, then give a verdict. Expect at least one to fail on Question 4 — that is the useful outcome of this exercise.

Exercise 2 — Define input, output, failure

For the strongest candidate from Exercise 1, write three sentences: exactly what goes in, exactly what comes out, and exactly what should happen when the input is missing or malformed. Keep this — Chapter 2.2 asks you to do the same thing before writing your first script, and Chapter 1.5 is about why the third sentence is the one that matters.

Exercise 3 — Find something not worth automating

Find a task in your flow that feels automatable and is not, and name which question it fails. Being able to argue against automation is what stops this curriculum turning into a hammer.

Exercise 4 — Check the checkability

Take a script your team already relies on. Can you state, for one known input, what its correct output is? If you can, you have its first test. If you cannot, you have found something worth being uneasy about.

Good candidates, already treated as engineering topics:

  • Automating GLS Triage — a task that passes all seven questions, and what automating it actually buys.
  • First-Mismatch Triage — why the first failure is the evidence worth extracting, which is part of what makes log triage checkable.

A judgment that should not be automated:

9. Summary

  • Seven questions before any code: how often, how error-prone, is the input defined, is the output checkable, how bad is a silent wrong answer, who else uses it, who maintains it.
  • Payback arithmetic is the shallowest of the seven. It tells you whether automation is worth the effort; Questions 3 to 7 tell you whether it will work. Worth it but uncheckable is the most expensive combination.
  • Automate the evidence, not the judgment. Gather the numbers and group the failures; let a person decide whether the block is ready. A script returning a verdict on a judgment call has hidden the decision, not removed it.
  • Good candidates are boring and mechanical: repeated launches, result collection, log triage, coverage aggregation, report extraction, filelist and register-map generation. Those are the chapters of this track.
  • Poor candidates are one-offs, judgments, anything uncheckable, and anything that hides hardware intent — plus anything you cannot yet do reliably by hand.
  • Match the engineering to the responsibility. Five lines for yourself needs nothing. Something that decides whether a regression passed needs all of it.

You should now be able to look at a candidate task and predict whether automating it will help — and be willing to say no.

That leaves the question the whole module has been building toward. Suppose the task passes all seven questions and you write the script. What actually makes it something a team can rely on, rather than something that works on your machine? Chapter 1.5 — What Separates a Script From a Tool Your Team Trusts is the contract that the remaining 22 modules exist to satisfy.

Where this fits

Part of the Python curriculum.