PONYλM2Modula-2

Bash.CodeCompared.To/Python

An interactive executable cheatsheet comparing Bash and Python

Bash 5.3 Python 3.14
Running a Script
Hello, World
The one row where nothing has changed. Both are a single line, both go to standard output, and both add the newline for you.
echo "Hello, World!"
print("Hello, World!")
Worth noticing before anything else diverges: print takes arguments rather than a line of text, so print("a", "b") puts a space between them the way echo a b does. The difference that will matter later is that echo hands the shell a string and Python hands print objects — a list prints as [1, 2, 3] rather than as three words, because it is still a list right up until it is displayed.
The shebang and running the file
Identical mechanics: a shebang line, the executable bit, and the kernel picking an interpreter. The habit worth carrying across is /usr/bin/env, which finds the interpreter on the path rather than assuming where it lives.
#!/usr/bin/env bash # Saved as report.sh, then: # chmod +x report.sh # ./report.sh set -euo pipefail echo "started"
#!/usr/bin/env python3 # Saved as report.py, then either: # chmod +x report.py && ./report.py # python3 report.py print("started")
The line to look at is set -euo pipefail, and its absence on the right. Those three options are the shell's attempt to behave the way Python already does: stop on an error, treat an unset variable as an error, and let a failure in the middle of a pipeline count. Python needs no equivalent because an unset name raises NameError, an error raises rather than continuing, and there is no pipeline to swallow anything.
Indentation replaces do/done
Six lines become three, and the two that disappear are done and fi. Python has no closing keyword for any block — the indentation is the syntax, and getting it wrong is an error rather than a silent change in meaning.
# The block markers are words, and every one has a partner. for name in alpha beta; do if [[ "$name" == "beta" ]]; then echo "found $name" fi done
# The block IS the indentation. There is nothing to close. for name in ["alpha", "beta"]: if name == "beta": print("found", name)
Shell programmers are usually the ones who take to this fastest, because a deeply nested script is exactly where fi/done/esac stop helping. The compensating rule is that indentation must be consistent: four spaces is the convention, tabs and spaces cannot be mixed, and a misaligned line is a syntax error rather than a line that quietly joins the wrong block.
Variables & Quoting
Assignment, and the spaces that break it
The first thing every shell programmer types wrong in Python is nothing — and the first thing every Python programmer types wrong in shell is a space around the equals sign. In Python the spaces are normal; in Bash they turn assignment into a command invocation.
name="widget" count=3 # name = "widget" ← this is a COMMAND called name, and it fails echo "$name $count"
name = "widget" count = 3 # Spaces are required by convention and ignored by the parser. print(name, count)
The deeper change is that count = 3 in Python is the number three, and count=3 in Bash is the one-character string "3". Everything in a shell is text; the shell converts to a number only inside $(( )) or [[ -eq ]], and converts back immediately. That is why shell arithmetic is a separate sub-language and why Python does not need one.
🚨 The quoting problem simply ends
This is the single best reason to move a script to Python, and it deserves the loudest row on the page. Watch the Bash loop run twice for one filename.
file="my report.txt" # Unquoted, the shell splits this into TWO words and both go wrong: for word in $file; do echo "got: [$word]" done # Quoted, it is one value. Getting this wrong is the classic shell bug. echo "quoted: [$file]"
file = "my report.txt" # There is no splitting. A string is one value, everywhere, always. print("got: [" + file + "]") for character_count in [len(file)]: print("length:", character_count)
Word splitting is not a bug in Bash, it is the design: an unquoted expansion is re-split on whitespace and then glob-expanded, because that is what made the shell convenient for interactive use in 1978. Every "$variable" you have ever typed is you switching it off. In Python a string is a value like any other — passing it around, storing it in a list, handing it to a function and printing it never splits it, so an entire category of bug that shell programmers spend years learning to avoid does not exist to be avoided.
Default values
Bash packs these into the expansion syntax — :- reads with a fallback and := also assigns. Python spells the same two operations as two different method calls, which is longer and unambiguous.
unset LEVEL echo "${LEVEL:-info}" # use a default, do not assign echo "${LEVEL:=warn}" # assign the default if unset echo "level is now: $LEVEL"
import os level = os.environ.get("LEVEL", "info") # use a default print(level) configuration = {} value = configuration.setdefault("level", "warn") # assign if missing print(value) print("level is now:", configuration["level"])
The shell's parameter expansions are genuinely compact and there is no shame in missing them; ${VAR:?message} in particular, which errors out with your message when unset, is nicer than anything Python offers in one expression. What you gain is that the fallback is ordinary code — os.environ.get(name) or compute_it() can call a function, catch an exception or read a file, where the shell form is limited to a word.
Arithmetic stops being a special mode
Bash has integers and nothing else, so anything with a decimal point has to leave the shell entirely and come back as text — which is what the bc pipeline is doing.
count=7 half=$(( count / 2 )) echo "half: $half" # Floating point is not available to the shell at all. echo "with bc: $(echo "scale=3; 7 / 2" | bc)"
count = 7 print("half:", count // 2) # floor division, an integer print("exact:", count / 2) # true division, a float print("rounded:", round(count / 3, 3))
This is the second-most common reason a script has to leave the shell. $(( )) is integer-only, bc means shelling out and parsing text back, and every intermediate result is a string that could be empty. Python has integers of unlimited size, floats, decimal for money and fractions for exact ratios, none of which needs a subprocess. Note the two division operators: // matches what $(( )) does, / always gives you a float.
Strings
Interpolation
An f-string is the shell's double-quoted string with the same job and clearer rules: the braces say exactly where the name ends, so there is no ${name}_suffix ambiguity to work around.
name="widget" count=3 echo "$count x $name" printf '%-10s|%5d|\n' "$name" "$count"
name = "widget" count = 3 print(f"{count} x {name}") print(f"{name:<10}|{count:>5}|")
printf maps almost directly onto the format specification after the colon — %-10s is {:<10}, %5d is {:>5}, %.2f is {:.2f}. What f-strings add is that the thing being formatted is an expression rather than a positional argument, so f"{count * 2}" and f"{name.upper()}" both work and there is no chance of the arguments getting out of step with the placeholders.
Slicing and substrings
The shell's expansions are terse and famously hard to remember which way round: one # is a short match from the front, two is a long match, and % is the same from the back. Python spells each one out as a named method.
path="/var/log/system.log" echo "${path##*/}" # basename: everything after the last slash echo "${path%/*}" # dirname: everything before the last slash echo "${path%.log}" # strip a suffix echo "${path:5:3}" # offset and length
path = "/var/log/system.log" print(path.rsplit("/", 1)[-1]) # basename print(path.rsplit("/", 1)[0]) # dirname print(path.removesuffix(".log")) # strip a suffix print(path[5:8]) # start and end, not start and length
🚨 The slice is the one to be careful with when porting: ${path:5:3} is offset and length, while path[5:8] is start and end. They agree here only because 5 + 3 is 8. Also note removesuffix rather than rstrip — rstrip(".log") strips any of those four characters repeatedly from the end, which is a real and common bug. For paths specifically, use pathlib rather than either, as the Files section shows.
Testing what a string looks like
[[ ]] is doing three different jobs here: a glob comparison with ==, a regular-expression comparison with =~, and case conversion in the expansion. Python separates them into a string method, the re module, and another string method.
value="Widget-42" if [[ "$value" == Widget-* ]]; then echo "glob match" fi if [[ "$value" =~ ^[A-Za-z]+-[0-9]+$ ]]; then echo "regex match: ${BASH_REMATCH[0]}" fi echo "${value,,}" # lowercase, bash 4+
import re value = "Widget-42" if value.startswith("Widget-"): print("glob match") match = re.fullmatch(r"[A-Za-z]+-[0-9]+", value) if match: print("regex match:", match.group(0)) print(value.lower())
The trap worth naming: inside [[ ]] the right-hand side of == is a glob, not a regular expression, and quoting it turns the glob off — so [[ "$value" == "Widget-*" ]] silently tests for a literal asterisk. Python has no such mode switch. It also has no equivalent of BASH_REMATCH as a magic global: the match object is returned, so it cannot be clobbered by an unrelated comparison two lines later.
Arrays & Dictionaries
Arrays
🚨 The quoting in "${names[@]}" is not optional decoration — without the quotes each element is word-split again, so an array holding my report.txt silently becomes two elements when you loop over it.
names=(alpha beta gamma) echo "count: ${#names[@]}" echo "first: ${names[0]}" echo "all: ${names[@]}" names+=(delta) for name in "${names[@]}"; do echo "- $name" done
names = ["alpha", "beta", "gamma"] print("count:", len(names)) print("first:", names[0]) print("all: ", " ".join(names)) names.append("delta") for name in names: print("-", name)
Bash arrays work, and they are the point at which most scripts should have stopped. They cannot be nested, cannot be passed to a function or returned from one, cannot be stored in another array, and every reference to one needs bracket-and-quote syntax that nobody remembers without looking it up. A Python list has none of those limits: it holds anything, including other lists, and is passed around like any other value.
Associative arrays and dictionaries
declare -A is required and easy to forget: without it Bash treats the subscript as arithmetic, evaluates apple as an unset variable worth zero, and every key silently becomes index 0.
declare -A counts counts[apple]=2 counts[pear]=5 echo "apple: ${counts[apple]}" echo "keys: ${!counts[@]}" for key in "${!counts[@]}"; do echo "$key = ${counts[$key]}" done
counts = {"apple": 2, "pear": 5} print("apple:", counts["apple"]) print("keys:", list(counts)) for key, value in counts.items(): print(f"{key} = {value}")
Note ${!counts[@]} for the keys against ${counts[@]} for the values — the exclamation mark is the difference, which is the kind of thing that makes shell scripts hard for the next reader. Python's .items() hands you both at once. Two limits that do not carry over: Bash associative arrays need version 4 or later, which rules out macOS's system /bin/bash, and their values can only be strings, so a dictionary of lists has no shell equivalent at all.
The structure the shell cannot hold
A dictionary whose values are dictionaries — the most ordinary structure imaginable — and the shell has no way to express it. The left column is what people actually write instead.
# There is no nesting. The usual workaround is to flatten the key # and parse it back out, which works until a value contains the # separator you chose. declare -A machine_role machine_role["web1:role"]="frontend" machine_role["web1:port"]="8080" machine_role["db1:role"]="database" for key in "${!machine_role[@]}"; do host="${key%%:*}" field="${key##*:}" echo "$host $field = ${machine_role[$key]}" done | sort
machines = { "web1": {"role": "frontend", "port": 8080}, "db1": {"role": "database"}, } for host, fields in sorted(machines.items()): for field, value in sorted(fields.items()): print(host, field, "=", value)
This is usually the row that settles the argument. Once a script has more than one thing to say about each item, the shell forces you to invent an encoding, and every encoding has a character it cannot survive. Note also that the port on the right is a real integer that can be compared and added to, while everything in the left column is a string that has to be converted at each use.
Conditionals & Loops
if, and what counts as true
The two languages mean different things by "true". In the shell an if runs a command and looks at its exit status, so zero is success; in Python it evaluates an expression, and zero is false. Getting these the wrong way round is the classic first-week mistake in both directions.
count=0 name="" # Truth in the shell is an EXIT STATUS: zero means success. if [[ "$count" -eq 0 ]]; then echo "count is zero"; fi if [[ -z "$name" ]]; then echo "name is empty"; fi if grep -q alpha <<< "alpha beta"; then echo "found it"; fi
count = 0 name = "" text = "alpha beta" # Truth is a property of the VALUE. 0 and "" are both falsy. if count == 0: print("count is zero") if not name: print("name is empty") if "alpha" in text: print("found it")
The comparison operators swap too, and they swap confusingly: in [[ ]], -eq compares numbers and == compares strings, so [[ "08" == "8" ]] is false while [[ "08" -eq "8" ]] is true. Python has one ==, and "08" == 8 is simply false because a string is not a number — no silent conversion in either direction.
case becomes match
The correspondence is unusually direct: | means "or" in both, and _ is the catch-all that * is. Python's match arrived in 3.10, so a script targeting an older interpreter uses if/elif instead.
for value in start stop restart other; do case "$value" in start) echo "$value: starting" ;; stop|restart) echo "$value: stopping" ;; *) echo "$value: unknown" ;; esac done
for value in ["start", "stop", "restart", "other"]: match value: case "start": print(f"{value}: starting") case "stop" | "restart": print(f"{value}: stopping") case _: print(f"{value}: unknown")
The similarity stops at what may be matched. A case pattern is a glob against a string; a match case can destructure a list or a dictionary, bind parts of it to names, and add an if guard — case {"role": role, "port": port} if port > 1024: is one statement. That is the direction the whole page runs in: the shell matches text, Python matches structure.
Counting loops
Note that {1..3} includes 3, and range(1, 4) stops before 4. Python's ranges are half-open everywhere, which is why range(len(items)) lines up with the valid indices.
for index in {1..3}; do echo "index $index" done for (( index = 0; index < 3; index++ )); do echo "c-style $index" done count=0 while (( count < 2 )); do echo "while $count" count=$(( count + 1 )) done
for index in range(1, 4): print("index", index) for index in range(3): print("c-style", index) count = 0 while count < 2: print("while", count) count += 1
{1..3} is brace expansion, which happens before anything else runs — which is why {1..$count} does not work in Bash and quietly produces the literal text instead. range is an ordinary object taking ordinary arguments, so range(1, count + 1) is unremarkable. The C-style loop exists in both; in Python you almost never write it, because iterating the collection directly is available and clearer.
Functions & Return Values
Parameters get names
A shell function receives positional variables — $1, $2, $@ — and naming them is a convention you perform by hand on the first lines. Python's parameters have names in the signature, which is also the documentation.
greet() { local name="$1" local greeting="${2:-Hello}" echo "$greeting, $name!" } greet "world" greet "world" "Goodbye"
def greet(name, greeting="Hello"): print(f"{greeting}, {name}!") greet("world") greet("world", "Goodbye") greet("world", greeting="Hi") # by name, in any order
🚨 local is the most important word in the left column, and leaving it out is a silent bug: without it, name is global and a function three levels up finds its variable overwritten. Python has the opposite default — a name assigned in a function is local unless you say global — which is the safer way round. Keyword arguments have no shell equivalent at all, and they are what makes a five-parameter Python function readable at the call site.
🚨 return does not return a value
This is the row that explains most of what looks strange in shell scripts. A function cannot hand a value back, so it prints one and the caller captures the output — which means every "return" is a subprocess, a fork and a round trip through text.
# A shell function returns an EXIT STATUS: 0-255, and nothing else. double() { echo $(( $1 * 2 )) # the "return value" is printed } result="$(double 21)" # and captured by running it in a subshell echo "result: $result" is_even() { return $(( $1 % 2 )) # 0 means true, which reads backwards } if is_even 4; then echo "4 is even"; fi
def double(value): return value * 2 result = double(21) print("result:", result) def is_even(value): return value % 2 == 0 if is_even(4): print("4 is even")
Three consequences disappear at once on the right. A function may return a list, a dictionary or an object rather than text. It costs nothing to call, so a script can be decomposed into small functions without paying a fork each time. And return means what it says: no more return 0 for true and return 1 for false, which reads backwards to everyone and is the source of endless confusion when a shell function is used both ways.
All the arguments
"$@" with the quotes is one of the few pieces of shell punctuation worth memorizing: it expands to each argument as a separate word, with the internal spaces preserved. $* and unquoted $@ both mangle the second argument here.
show_all() { echo "count: $#" for argument in "$@"; do echo "- [$argument]" done } show_all one "two words" three
def show_all(*arguments): print("count:", len(arguments)) for argument in arguments: print(f"- [{argument}]") show_all("one", "two words", "three")
Python's *arguments collects the extras into an ordinary tuple, so it can be measured, sliced, passed onward and stored. There is also **keyword_arguments for named ones, which the shell has no concept of. The pattern function(*some_list) spreads a list back out into separate arguments — the reverse operation, and the equivalent of writing "${array[@]}" at a call site.
Pipelines Become Loops
A pipeline, rewritten
The classic four-stage counting pipeline against the one object that does the whole job. Counter is a dictionary that counts, and most_common() is the sort | uniq -c | sort -rn at the end.
printf '%s\n' banana apple cherry apple | sort | uniq -c | sort -rn | awk '{ printf "%7d %s\n", $1, $2 }'
from collections import Counter fruit = ["banana", "apple", "cherry", "apple"] counts = Counter(fruit) for name, count in counts.most_common(): print(f"{count:>7} {name}")
Be honest about what is lost: the pipeline streams, so it counts a file larger than memory, and it is four words long. What is gained is that the counts are numbers rather than a column of text that has to be re-parsed by whatever comes next, and that adding a second condition does not mean adding a fifth stage. The rule of thumb that holds up: a pipeline that ends by printing is fine where it is, and a pipeline whose output another stage has to parse is asking to be rewritten.
while read becomes a for loop
The incantation on the left is not superstition: IFS= stops leading and trailing whitespace being trimmed, and -r stops backslashes being interpreted. Leave either out and the data changes on the way in.
# IFS= and -r are both required, and both are easy to forget: # without them, leading spaces are eaten and backslashes are consumed. printf '%s\n' " alpha" "beta\\gamma" | while IFS= read -r line; do echo "[$line]" done
text = " alpha\nbeta\\gamma\n" for line in text.splitlines(): print(f"[{line}]")
🚨 There is a second, worse trap in the left column that this example sidesteps: a while read loop on the right-hand side of a pipe runs in a subshell, so any variable it sets is gone when the loop ends. Countless scripts count something in a loop and then print zero. Python has no subshell and no such rule — a name assigned in a loop is still there afterwards.
grep and awk become a comprehension
The comprehension reads in the same order as the awk program: a condition, and something to do with what survives it. The difference is that the numbers stay numbers instead of being re-parsed from text at every stage.
printf '%s\n' 3 14 15 92 65 | awk '$1 > 10 { print $1 * 2 }'
values = [3, 14, 15, 92, 65] doubled = [value * 2 for value in values if value > 10] for value in doubled: print(value)
A comprehension composes where a pipeline chains, and the two hit different walls. Chaining three awk stages is easy and reading them a year later is not; nesting three comprehensions is possible and immediately unreadable, at which point you write a loop. The genuinely useful thing Python adds is that the intermediate list is a value: it can be measured, reused, sorted differently twice, or passed to two functions, none of which a pipeline can do without running the work again.
Keeping the streaming, without the pipeline
A generator is Python's pipeline stage: yield hands one value to whoever is looping and suspends until the next is asked for, so nothing accumulates. The parentheses around the second expression make it a generator rather than a list.
# The pipeline's real advantage: nothing is ever fully in memory. seq 1 100000 | awk '$1 % 9999 == 0' | head -3
def numbers(limit): for value in range(1, limit + 1): yield value # produced one at a time interesting = (value for value in numbers(100000) if value % 9999 == 0) for index, value in enumerate(interesting): if index == 3: break print(value)
This is worth knowing before anyone tells you Python cannot handle a large file. for line in open(path) reads one line at a time, generators chain the way pipeline stages do, and the itertools module supplies the equivalents of head, uniq and paste. What the shell still wins on is that its stages run as separate processes on separate cores; a generator chain is one thread doing one thing at a time.
sed, awk, grep
sed s/// becomes re.sub
When the pattern is a fixed string, str.replace is the answer and no regular expression is involved at all — which is faster and cannot be broken by a special character in the data.
echo "hello world" | sed 's/o/0/g' echo "2026-09-09" | sed -E 's/([0-9]{4})-([0-9]{2})-([0-9]{2})/\3\/\2\/\1/'
import re print("hello world".replace("o", "0")) print(re.sub(r"(\d{4})-(\d{2})-(\d{2})", r"\3/\2/\1", "2026-09-09"))
Three differences that bite when porting. Python's backreferences are \1 in a raw string rather than \1 after shell quoting has had a go at it, which removes a whole layer of escaping. The delimiter is not part of the syntax, so a pattern containing a slash needs no s|…|…| trick. And re uses Perl-style syntax throughout, so \d and + work without -E and without wondering which of the four regular-expression dialects on the machine you are talking to.
awk fields become split
Unpacking straight into three named variables is the move that makes this readable. line.split() with no argument splits on any run of whitespace, exactly as awk does by default.
printf '%s\n' "web1 frontend 8080" "db1 database 5432" | awk '{ print $2 " on " $1 ":" $3 }'
lines = ["web1 frontend 8080", "db1 database 5432"] for line in lines: host, role, port = line.split() print(f"{role} on {host}:{port}")
awk is genuinely excellent at this and there is no need to be sniffy about it — a one-line field extraction should stay a one-line field extraction. The case for moving is names: $2 is a mystery on line 40 of a script and role is not, and if a column is inserted upstream the Python version breaks loudly at the unpacking rather than quietly printing the wrong field. Use line.split(",") for a fixed delimiter, and the csv module the moment quoting is involved.
grep, and what comes after it
Three stages become one expression, and the order matches: filter, transform, sort. sorted accepts any iterable, so the generator inside it never becomes a list of its own.
printf '%s\n' "INFO ok" "ERROR disk full" "INFO fine" "ERROR timeout" | grep ERROR | sed 's/^ERROR //' | sort
lines = ["INFO ok", "ERROR disk full", "INFO fine", "ERROR timeout"] problems = sorted( line.removeprefix("ERROR ") for line in lines if line.startswith("ERROR") ) for problem in problems: print(problem)
This is where the balance genuinely tips. One grep is better as a grep. By the third stage you are maintaining a small program written in three languages with two quoting regimes, and every value in it is text. The Python version keeps one language and one set of rules, and the moment you need to count the errors as well as list them, it is a two-line change rather than a redesign.
Files & Paths
Writing and reading a file
Redirection is the shell's single best feature and Python has nothing as short. What it has instead is pathlib, where the path is an object that knows how to read and write itself.
printf '%s\n' alpha beta > /tmp/example.txt echo gamma >> /tmp/example.txt cat /tmp/example.txt echo "lines: $(wc -l < /tmp/example.txt | tr -d '[:space:]')"
from pathlib import Path path = Path("/tmp/example.txt") path.write_text("alpha\nbeta\n") with path.open("a") as handle: handle.write("gamma\n") print(path.read_text(), end="") print("lines:", len(path.read_text().splitlines()))
The with block is the equivalent of the shell closing the file when the redirection ends: the handle is closed at the end of the block, including when an exception is raised inside it. For anything small, read_text and write_text skip the block entirely. Both columns here are writing to a temporary filesystem that starts empty on every run — the browser gives each of these runtimes its own, which is why every example on this page creates what it reads.
Paths stop being strings
The slash operator is overloaded to join paths, so directory / "system.log" is a path rather than a division. It collapses duplicate separators, which is the bug the left column is one stray slash away from.
path="/var/log/system.log" echo "$(basename "$path")" echo "$(dirname "$path")" echo "${path##*.}" # Joining, with the classic double-slash risk: directory="/var/log/" echo "$directory/system.log"
from pathlib import Path path = Path("/var/log/system.log") print(path.name) print(path.parent) print(path.suffix.lstrip(".")) # Joining with / handles the separators for you: directory = Path("/var/log/") print(directory / "system.log")
A Path also answers questions the shell answers with test operators: path.exists(), path.is_dir(), path.stat().st_size. And it carries the whole path around as one value, so a filename containing a space, a newline or a leading dash is simply a filename — where in the shell each of those is a documented hazard with its own workaround. Prefer pathlib over the older os.path functions in new code.
Globbing
Both expand a pattern into filenames. The difference is what happens when nothing matches: the shell hands the loop the literal pattern *.txt as if it were a filename, and glob yields nothing.
mkdir -p /tmp/globdemo : > /tmp/globdemo/one.txt : > /tmp/globdemo/two.txt : > /tmp/globdemo/three.log for file in /tmp/globdemo/*.txt; do echo "found: $(basename "$file")" done
from pathlib import Path directory = Path("/tmp/globdemo") directory.mkdir(parents=True, exist_ok=True) for name in ["one.txt", "two.txt", "three.log"]: (directory / name).touch() for file in sorted(directory.glob("*.txt")): print("found:", file.name)
That difference has ended more scripts than it should have — a backup loop that matches nothing goes on to process a file literally named *.txt, and shopt -s nullglob is the fix nobody remembers to set. Two more: glob returns paths in filesystem order, so sorted is worth writing when the order matters, and rglob("*.txt") or glob("**/*.txt") recurses without needing find.
Running Other Programs
Running a command
🚨 The list is the whole security story: ["sort"] passes arguments directly to the program with no shell in between, so nothing in the data can be interpreted as a command. Adding shell=True hands the string to a shell and reintroduces every injection risk you have ever quoted against.
output="$(printf '%s\n' gamma alpha beta | sort)" echo "$output" echo "exit status: $?"
import subprocess result = subprocess.run( ["sort"], input="gamma\nalpha\nbeta\n", capture_output=True, text=True, check=True, ) print(result.stdout, end="") print("exit status:", result.returncode)
This is the row where the shell keeps its advantage and it is worth saying plainly: running a program is what a shell is for, and six lines against one is a real cost. The compensation is that result holds the output, the error output and the status as three separate values you can examine, rather than one string plus a magic variable that the next command overwrites. check=True raises on a non-zero status, which is set -e for one command.
Two programs, piped together
A pipe is one character in the shell and a variable handoff in Python. This is the clearest case where the shell is simply the better tool, and the honest advice is to notice it rather than to argue.
printf '%s\n' 3 1 4 1 5 | sort -n | uniq | tr '\n' ' ' echo
import subprocess first = subprocess.run( ["sort", "-n"], input="3\n1\n4\n1\n5\n", capture_output=True, text=True, check=True, ) second = subprocess.run( ["uniq"], input=first.stdout, capture_output=True, text=True, check=True, ) print(second.stdout.replace("\n", " ").strip())
When a script is mostly this, leave it in the shell. When a script has one pipeline and four hundred lines of logic around it, move the logic and either keep the pipeline as a single subprocess.run with shell=True on a string you wrote yourself, or replace it with the twenty characters of Python that do the same thing — sorted(set(values)) covers this particular example. Wiring real pipes between processes is possible with Popen and stdout handoff, and it is more code than it is usually worth.
The environment
os.environ is a dictionary, so everything you know about dictionaries applies: .get with a default, in to test, iteration over the keys.
export REPORT_LEVEL=verbose echo "level: $REPORT_LEVEL" echo "set? $([[ -n "${REPORT_LEVEL:-}" ]] && echo yes || echo no)" echo "unset var reads as: [${NOT_SET_ANYWHERE:-}]"
import os os.environ["REPORT_LEVEL"] = "verbose" print("level:", os.environ["REPORT_LEVEL"]) print("set?", "yes" if os.environ.get("REPORT_LEVEL") else "no") print("unset var reads as: [" + os.environ.get("NOT_SET_ANYWHERE", "") + "]")
The shell distinction between a shell variable and an exported one has no Python equivalent — every entry in os.environ is exported by definition, and an ordinary Python variable is not in the environment at all. Changing os.environ affects processes this program starts afterwards and nothing else, which is exactly the shell's rule too: neither can alter the environment of the process that started it.
Errors & Exit Codes
set -e becomes not needing set -e
An error in Python stops the program by default and there is no flag to set. The shell's default is the opposite — a failed command sets a status nobody reads and the script continues — which is what set -e is trying to patch.
set -euo pipefail # Without those options, this script would carry on after the failure # and report success at the end. if ! false; then echo "the command failed, and we noticed" fi echo "still running"
def might_fail(): raise ValueError("the command failed") try: might_fail() except ValueError as error: print(f"{error}, and we noticed") print("still running")
set -euo pipefail is the right first line of any shell script and it is still not enough: -e is famously full of exceptions, ignoring failures inside if conditions, in any command followed by &&, and in most of a function called in a condition. Nobody can hold the full rule in their head. Python's version is one sentence: an exception propagates until something catches it.
Exiting with a status
Both write the message to standard error rather than standard output — >&2 on the left, file=sys.stderr on the right — which is the rule that lets a caller separate a diagnostic from a result.
check() { if [[ "$1" == "bad" ]]; then echo "fatal: bad input" >&2 return 2 fi echo "ok" } check good check bad || echo "caller saw status $?"
import sys def check(value): if value == "bad": raise ValueError("bad input") print("ok") check("good") try: check("bad") except ValueError as error: print(f"caller saw: {error}", file=sys.stderr) # sys.exit(2) here would end the program with that status.
The important difference is how much the failure can carry. An exit status is a number from 0 to 255, so everything the caller learns is one of 256 values, and the message is unstructured text on another stream. An exception carries a type, a message, any attributes you put on it, and the stack it came from — so the caller can respond differently to different failures without parsing anything. sys.exit(2) still exists for when the program genuinely has to hand a status back to a shell.
trap becomes finally
trap … RETURN and finally make the same promise: this runs whether the body succeeded or not. Python's version is a block rather than a string of shell code evaluated later, so the cleanup is checked at compile time and can see local variables directly.
work() { local scratch scratch="$(mktemp)" trap 'rm -f "$scratch"; echo "cleaned up"' RETURN echo "working with a scratch file" } work
import tempfile, os def work(): handle = tempfile.NamedTemporaryFile(delete=False) path = handle.name handle.close() try: print("working with a scratch file") finally: os.unlink(path) print("cleaned up") work()
The idiomatic Python is shorter still — with tempfile.NamedTemporaryFile() as handle: deletes the file at the end of the block with no finally written at all, because the cleanup belongs to the object rather than to the call site. That is the pattern to look for whenever you would have reached for trap: an object that tidies up after itself is available for files, locks, sockets, database connections and directories.
Command-Line Arguments
Positional arguments
sys.argv is one list with the program name at index 0, which is $0, and the arguments after it. $# counts the arguments without the program name, which is why the right-hand side subtracts one.
bash script.sh alpha beta
# $0 is the script's own name, and is left out here because it differs # between the two runtimes on this page rather than between languages. echo "count: $#" echo "first: $1" echo "all: $*"
python main.py alpha beta
import sys # sys.argv[0] is the script's own name, the equivalent of $0. print("count:", len(sys.argv) - 1) print("first:", sys.argv[1]) print("all:", " ".join(sys.argv[1:]))
It being a list rather than a set of numbered variables is the practical difference: it can be sliced (sys.argv[1:] is "$@"), passed to a function, and iterated without the quoting rules that make "$@" and $* behave differently. But for anything with options, do not parse this by hand — the next row is what you want.
getopts becomes argparse
Each option is declared once, with its short form, its long form and its help text together. The parser generates --help, rejects unknown options with a usage message, and converts values to the type you asked for.
bash script.sh -v -o report.txt
verbose=0 output="" while getopts "vo:" option; do case "$option" in v) verbose=1 ;; o) output="$OPTARG" ;; *) echo "usage: report.sh [-v] [-o FILE]" >&2; exit 1 ;; esac done echo "verbose=$verbose output=${output:-none}" # getopts handles short options only. Long options such as --output # mean writing the parsing loop by hand.
python main.py --verbose --output report.txt
import argparse parser = argparse.ArgumentParser(description="Summarize a log file") parser.add_argument("-v", "--verbose", action="store_true", help="print more detail") parser.add_argument("-o", "--output", metavar="FILE", help="where to write the report") options = parser.parse_args() # reads sys.argv print(f"verbose={int(options.verbose)} output={options.output or 'none'}") # --help is generated from the definitions above, at no extra cost.
This is the row that quietly justifies the whole move for anything another person will run. getopts handles single-letter options and nothing else, so every script that accepts --dry-run contains a hand-written loop with its own bugs, and the usage message is a string that drifts out of date the first time an option is added. argparse also gives you type=int, choices=[...], required=True, subcommands and mutually exclusive groups, none of which you write yourself.
JSON & Structured Data
jq becomes json.loads
jq is excellent and is the right answer for a one-off query at a terminal. What it cannot do is hand the result back as anything but text, so every value that comes out of it has to be re-parsed by whatever uses it next.
read -r -d '' payload <<'JSON' || true {"host": "web1", "ports": [80, 443], "tags": {"role": "frontend"}} JSON echo "$payload" | jq -r '.host' echo "$payload" | jq -r '.ports[1]' echo "$payload" | jq -r '.tags.role'
import json payload = '{"host": "web1", "ports": [80, 443], "tags": {"role": "frontend"}}' data = json.loads(payload) print(data["host"]) print(data["ports"][1]) print(data["tags"]["role"])
json.loads returns real Python objects: data["ports"] is a list of integers that can be summed, sorted or appended to, and data["ports"][1] + 1 is 444 rather than a string concatenation. The reverse direction is json.dumps, which is where the shell struggles most — building JSON by string concatenation is how invalid output gets shipped, and jq -n with arguments is the shell-side fix.
Producing JSON safely
Look at the first line of output on the left: the embedded quotes break the JSON, and nothing warns you. This is the most common way a shell script produces output that a downstream consumer rejects.
name='say "hello"' count=3 # Concatenating by hand produces INVALID json the moment a value # contains a quote, a newline or a backslash: echo "{\"name\": \"$name\", \"count\": $count}" # jq -n is the correct way, and it is worth the extra words: jq -nc --arg name "$name" --argjson count "$count" \ '{name: $name, count: $count}'
import json name = 'say "hello"' count = 3 print(json.dumps({"name": name, "count": count})) # Escaping is not something the caller can get wrong. print(json.dumps({"nested": {"values": [1, 2, 3]}}, indent=2))
The rule in the shell is that JSON must be built by jq -n with --arg for every value, never by interpolation — and it is a rule that is skipped under deadline pressure precisely because the naive version works on the test data. json.dumps has no unsafe mode to fall into: it takes objects rather than text, so quoting is not a decision the caller makes. The same argument applies to building SQL, HTML, shell commands and URLs.
Comma-separated data
The quoted comma is the whole row. cut -d, counts commas and knows nothing about quoting, so any real comma-separated file — one exported from a spreadsheet, say — is one embedded comma away from silently wrong output.
printf '%s\n' 'name,note' 'widget,"contains, a comma"' | tail -n +2 | cut -d, -f1 # cut has no idea what a quoted field is, so the second column of # that row would be split in half by any -f2 you asked for.
import csv, io text = 'name,note\nwidget,"contains, a comma"\n' reader = csv.DictReader(io.StringIO(text)) for row in reader: print(row["name"]) print(" note:", row["note"])
csv.DictReader handles quoting, embedded commas, embedded newlines and the header row, and hands back each record as a dictionary keyed by column name. That last part matters as much as the parsing: row["note"] keeps working when a column is inserted upstream, where -f2 silently starts reading a different column. There is a matching csv.DictWriter for the other direction.
When The Script Grows Up
Splitting the script up
This is the difference that decides whether a script can grow past one file. source is textual inclusion into one shared namespace; import creates a namespace and binds names from it.
# One file, or "source" another and hope: # source ./helpers.sh # # Everything sourced shares one global namespace, so two helper files # that both define a function called "log" silently overwrite each # other, and the winner depends on the order of the source lines. log() { echo "[log] $*"; } log "started"
# Each module has its own namespace, and the import says where a # name came from: # # from helpers import log # or: import helpers; helpers.log(...) # # Two modules may both define log without either being disturbed. def log(*parts): print("[log]", *parts) log("started")
Everything else about large-program structure follows from this. A Python module is loaded once no matter how many files import it, so there is no include-guard problem; a name collision is impossible to have by accident; and the import list at the top of a file is an honest inventory of its dependencies, which a pile of source lines is not. A shell script that has reached three files is telling you something.
Testing
A test runner is in the standard library, and so is a second one — unittest here, and doctest for examples embedded in documentation. Neither is installed; both are simply there.
# The shell has no assertion and no test runner in the box. Testing # means either a third-party framework (bats, shunit2) or this: assert_equal() { if [[ "$1" != "$2" ]]; then echo "FAIL: expected [$2], got [$1]" >&2 return 1 fi echo "ok" } double() { echo $(( $1 * 2 )); } assert_equal "$(double 21)" "42"
import unittest def double(value): return value * 2 class DoubleTests(unittest.TestCase): def test_doubles_a_number(self): self.assertEqual(double(21), 42) def test_handles_zero(self): self.assertEqual(double(0), 0) unittest.main(argv=["ignored"], exit=False, verbosity=2)
Testability is the quiet reason scripts get rewritten. A shell function can only be tested through its output text, so a test asserts on formatting as well as behavior and breaks when a message is reworded. A Python function returns a value, so the test can check the value and ignore the presentation. The wider ecosystem is the other half: pytest, coverage measurement, type checking with mypy and formatting with ruff all expect ordinary functions in importable modules.
When to stay in the shell
A page like this one can read as an argument that shell scripts are bad. They are not, and the last row should say so: this script is exactly right as it stands.
#!/usr/bin/env bash # This script should NOT be rewritten. It is glue, it is short, and # every line is a program being run — which is what a shell is for. set -euo pipefail mkdir -p /tmp/release printf '%s\n' one two three > /tmp/release/manifest.txt sort /tmp/release/manifest.txt | tee /tmp/release/sorted.txt | wc -l
# The same thing in Python: longer, no clearer, and it gives up the # one thing the shell is best at. Rewrite when the script has grown # LOGIC — data structures, arithmetic, error handling, options, # anything a second person must modify — not merely because it is long. from pathlib import Path directory = Path("/tmp/release") directory.mkdir(parents=True, exist_ok=True) (directory / "manifest.txt").write_text("one\ntwo\nthree\n") lines = sorted((directory / "manifest.txt").read_text().splitlines()) (directory / "sorted.txt").write_text("\n".join(lines) + "\n") print(len(lines))
The honest test is not length. It is whether the script has logic — a data structure with more than one level, arithmetic that is not integer counting, more than about three options, error handling that has to distinguish between failures, or anything another person will have to modify under time pressure. Orchestration, glue and one-shot pipelines belong in the shell, where they are shorter and clearer than anything Python will give you.

Thank you — anything else?